Freedman, Rachel
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 165 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
2024/04/16 by Vincent Conitzer, Conitzer, Vincent, Rachel Freedman +21 · 1 voice · 43 citations
Computer Science · #cs.LG #cs.AI #cs.CL #cs.CY #cs.GT
- Linear Probe Penalties Reduce LLM Sycophancy
2024/12/01 by Henry Papadatos, Papadatos, Henry, Rachel Freedman +1 · 14 citations
Engineering · #Particle accelerators and beam dynamics #Particle Accelerators and Free-Electron Lasers
- Choice Set Misspecification in Reward Inference
2021/01/19 by Rachel Freedman, Freedman, Rachel, Rohin Shah +4 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Reinforcement Learning in Robotics #cs.AI #cs.HC
- The Expertise Problem: Learning from Specialized Feedback
2022/11/12 by Oliver Daniels-Koch, Daniels-Koch, Oliver, Rachel A. Freedman +1 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms