vix.ing · top · new · best · stats · spec

Paul Christiano

  1. Training language models to follow instructions with human feedback
    2022/03/04 by Long Ouyang, Jeff Wu, Ouyang, Long +38 · 9 voices · 2553 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
  2. Concrete Problems in AI Safety
    2016/06/21 by Dario Amodei, Chris Olah, Amodei, Dario +9 · 7 voices · 213 citations
    #cs.AI #cs.LG
  3. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  4. Learning to summarize from human feedback
    2020/09/02 by Nisan Stiennon, Stiennon, Nisan, Long Ouyang +16 · 3 voices · 340 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  5. Deep reinforcement learning from human preferences
    2017/06/12 by Paul F. Christiano, Paul Christiano, Christiano, Paul +11 · 1 voice · 580 citations
    Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #cs.AI #cs.HC #cs.LG #stat.ML
  6. AI safety via debate
    2018/05/02 by Geoffrey Irving, Paul F. Christiano, Paul Christiano +4 · 3 voices · 73 citations
    Computer Science · Mathematics · #Computability, Logic, AI Algorithms #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #cs.LG #stat.ML
  7. Model evaluation for extreme risks
    2023/05/24 by Toby Shevlane, Shevlane, Toby, Sebastian Farquhar +39 · 16 citations
    Computer Science · #Software Engineering Research #Software Reliability and Analysis Research #Information and Cyber Security
  8. Robust Cooperation in the Prisoner's Dilemma: Program Equilibrium via Provability Logic
    2014/01/22 by Mihaly Barasz, Barasz, Mihaly, Paul Christiano +9 · 1 voice · 2 citations
    Computer Science · #Computer Science and Game Theory (cs.GT) #F.4.1 #FOS: Computer and information sciences #Logic in Computer Science (cs.LO) #cs.GT #cs.LO
  9. A Cryptographic Test of Quantumness and Certifiable Randomness from a\n Single Quantum Device
    2018/04/02 by Zvika Brakerski, Brakerski, Zvika, Paul Christiano +7 · 8 citations
    Computer Science · #Computational Complexity (cs.CC) #Cryptographic Implementations and Security #Cryptography and Data Security #FOS: Computer and information sciences #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Physics (quant-ph)