vix.ing · top · new · best · stats · spec

Philip S. Thomas

  1. Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
    2016/04/04 by Philip S. Thomas, Thomas, Philip S., Emma Brunskill +1 · 16 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Simulation Techniques and Applications #Software Reliability and Analysis Research
  2. Is the Policy Gradient a Gradient?
    2019/06/17 by Chris Nota, Philip S. Thomas, Nota, Chris +1 · 3 voices · 1 citation
    #cs.LG #stat.ML
  3. Data-Efficient Policy Evaluation Through Behavior Policy Search
    2017/06/12 by Josiah P. Hanna, Philip S. Thomas, Hanna, Josiah P. +5 · 2 citations
    Computer Science · Business, Management and Accounting · #Reinforcement Learning in Robotics #Supply Chain and Inventory Management #Software Reliability and Analysis Research
  4. High-Confidence Off-Policy (or Counterfactual) Variance Estimation
    2021/01/25 by Yash Chandak, Chandak, Yash, Shiv Shankar +3 · 2 citations
    Computer Science · #Age of Information Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel Computing and Optimization Techniques #Software System Performance and Reliability
  5. Position: Benchmarking is Limited in Reinforcement Learning Research
    2024/06/23 by Scott M. Jordan, Adam White, Jordan, Scott M. +7 · 3 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Methodology (stat.ME) #Open Source Software Innovations
  6. Optimizing for the Future in Non-Stationary MDPs
    2020/05/17 by Yash Chandak, Georgios Theocharous, Chandak, Yash +9 · 1 citation
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Advanced Multi-Objective Optimization Algorithms
  7. Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
    2025/06/09 by Yaswanth Chittepu, Blossom Metevier, Chittepu, Yaswanth +9 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Applications (stat.AP) #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling