Paul Christiano
- Training language models to follow instructions with human feedback
2022/03/04 by Long Ouyang, Jeff Wu, Ouyang, Long +38 · 9 voices · 2553 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Concrete Problems in AI Safety
2016/06/21 by Dario Amodei, Chris Olah, Amodei, Dario +9 · 7 voices · 213 citations
#cs.AI #cs.LG
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Learning to summarize from human feedback
2020/09/02 by Nisan Stiennon, Stiennon, Nisan, Long Ouyang +16 · 3 voices · 340 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- Deep reinforcement learning from human preferences
2017/06/12 by Paul F. Christiano, Paul Christiano, Christiano, Paul +11 · 1 voice · 580 citations
Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #cs.AI #cs.HC #cs.LG #stat.ML
- AI safety via debate
2018/05/02 by Geoffrey Irving, Paul F. Christiano, Paul Christiano +4 · 3 voices · 73 citations
Computer Science · Mathematics · #Computability, Logic, AI Algorithms #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #cs.LG #stat.ML
- Model evaluation for extreme risks
2023/05/24 by Toby Shevlane, Shevlane, Toby, Sebastian Farquhar +39 · 16 citations
Computer Science · #Software Engineering Research #Software Reliability and Analysis Research #Information and Cyber Security
- Robust Cooperation in the Prisoner's Dilemma: Program Equilibrium via Provability Logic
2014/01/22 by Mihaly Barasz, Barasz, Mihaly, Paul Christiano +9 · 1 voice · 2 citations
Computer Science · #Computer Science and Game Theory (cs.GT) #F.4.1 #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #cs.GT #cs.LO
- A Cryptographic Test of Quantumness and Certifiable Randomness from a\n Single Quantum Device
2018/04/02 by Zvika Brakerski, Brakerski, Zvika, Paul Christiano +7 · 8 citations
Computer Science · #Computational Complexity (cs.CC) #Cryptographic Implementations and Security #Cryptography and Data Security #FOS: Computer and information sciences #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Physics (quant-ph)