Thomas, Philip S.
- Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
2016/04/04 by Philip S. Thomas, Emma Brunskill, Thomas, Philip S. +1 · 19 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Simulation Techniques and Applications #Software Reliability and Analysis Research
- Evaluating the Performance of Reinforcement Learning Algorithms
2020/06/30 by Jordan, Scott M., Chandak, Yash, Cohen, Daniel +2 · 5 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Learning Action Representations for Reinforcement Learning
2019/02/01 by Chandak, Yash, Theocharous, Georgios, Kostas, James +2 · 4 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- A New Confidence Interval for the Mean of a Bounded Random Variable
2019/05/15 by Learned-Miller, Erik, Thomas, Philip S. · 3 citations
#FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Probability (math.PR) #Statistics Theory (math.ST)
- Behavior Alignment via Reward Function Optimization
2023/10/29 by Gupta, Dhawal, Chandak, Yash, Jordan, Scott M. +2 · 5 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Is the Policy Gradient a Gradient?
2019/06/17 by Chris Nota, Philip S. Thomas, Nota, Chris +1 · 3 voices · 1 citation
#cs.LG #stat.ML
- Lifelong Learning with a Changing Action Set
2019/06/05 by Chandak, Yash, Theocharous, Georgios, Nota, Chris +1 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Data-Efficient Policy Evaluation Through Behavior Policy Search
2017/06/12 by Josiah P. Hanna, Hanna, Josiah P., Philip S. Thomas +5 · 2 citations
Computer Science · Business, Management and Accounting · #Reinforcement Learning in Robotics #Supply Chain and Inventory Management #Software Reliability and Analysis Research
- High-Confidence Off-Policy (or Counterfactual) Variance Estimation
2021/01/25 by Yash Chandak, Shiv Shankar, Chandak, Yash +3 · 2 citations
Computer Science · #Age of Information Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel Computing and Optimization Techniques #Software System Performance and Reliability
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
2021/05/31 by Satija, Harsh, Thomas, Philip S., Pineau, Joelle +1 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Position: Benchmarking is Limited in Reinforcement Learning Research
2024/06/23 by Scott M. Jordan, Jordan, Scott M., Adam White +7 · 3 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Methodology (stat.ME) #Open Source Software Innovations
- Importance Sampling with Unequal Support
2016/11/10 by Thomas, Philip S., Brunskill, Emma · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- A Compression-Inspired Framework for Macro Discovery
2017/11/24 by Garcia, Francisco M., da Silva, Bruno C., Thomas, Philip S. · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Electrical engineering #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
- A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning
2019/02/03 by Garcia, Francisco M., Thomas, Philip S. · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Optimizing for the Future in Non-Stationary MDPs
2020/05/17 by Yash Chandak, Chandak, Yash, Georgios Theocharous +9 · 1 citation
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Advanced Multi-Objective Optimization Algorithms
- Towards Safe Policy Improvement for Non-Stationary MDPs
2020/10/23 by Chandak, Yash, Jordan, Scott M., Theocharous, Georgios +2 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Adaptive Rollout Length for Model-Based RL Using Model-Free Deep RL
2022/06/06 by Abhinav Bhatia, Philip S. Thomas, Bhatia, Abhinav +3 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Reinforcement Learning in Robotics #Software Engineering Research
- Enforcing Delayed-Impact Fairness Guarantees
2022/08/24 by Weber, Aline, Metevier, Blossom, Brun, Yuriy +2 · 1 citation
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- From Past to Future: Rethinking Eligibility Traces
2023/12/20 by Gupta, Dhawal, Jordan, Scott M., Chaudhari, Shreyas +3 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- ICU-Sepsis: A Benchmark MDP Built from Real Medical Data
2024/06/09 by Kartik Choudhary, Dhawal Gupta, Choudhary, Kartik +3 · 1 citation
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare
- Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
2025/06/09 by Yaswanth Chittepu, Chittepu, Yaswanth, Blossom Metevier +9 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Applications (stat.AP) #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling