vix.ing · top · new · best · stats · spec

Jan Leike

  1. Evaluating Large Language Models Trained on Code
    2021/07/07 by Mark Chen, Chen, Mark, Jerry Tworek +122 · 9 voices · 1320 citations
    Computer Science · #Parallel Computing and Optimization Techniques #Software Engineering Research #Topic Modeling #cs.LG
  2. Training language models to follow instructions with human feedback
    2022/03/04 by Long Ouyang, Jeff Wu, Ouyang, Long +38 · 9 voices · 2616 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
  3. Let's Verify Step by Step
    2023/05/31 by Hunter Lightman, Vineet Kosaraju, Lightman, Hunter +17 · 9 voices · 724 citations
    Computer Science · #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. Deep reinforcement learning from human preferences
    2017/06/12 by Paul F. Christiano, Paul Christiano, Christiano, Paul +11 · 1 voice · 603 citations
    Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #cs.AI #cs.HC #cs.LG #stat.ML
  5. Reasoning Models Don't Always Say What They Think
    2025/05/08 by Yanda Chen, Chen, Yanda, Joe Benton +27 · 6 voices · 100 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  6. Unsupervised Elicitation of Language Models
    2025/06/11 by Jiaxin Wen, Zachary Ankner, Wen, Jiaxin +23 · 15 voices · 4 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Machine Learning and Data Classification
  7. Prover-Verifier Games improve legibility of LLM outputs
    2024/07/18 by Jan Hendrik Kirchner, Jan H. Kirchner, Yining Chen +10 · 1 voice · 16 citations
    Computer Science · #Logic, programming, and type systems #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #cs.CL
  8. Scalable agent alignment via reward modeling: a research direction
    2018/11/19 by Jan Leike, Leike, Jan, David Krueger +9 · 72 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  9. AI Safety Gridworlds
    2017/11/27 by Jan Leike, Leike, Jan, Miljan Martic +13 · 1 voice · 17 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #cs.AI #cs.LG
  10. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
    2025/01/31 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +89 · 8 voices · 40 citations
    Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
  11. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
    2023/12/14 by Collin Burns, Burns, Collin, Pavel Izmailov +21 · 52 citations
    Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  12. Self-critiquing models for assisting human evaluators
    2022/06/12 by William H. Saunders, Catherine Vance Yeh, Saunders, William +11 · 31 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
  13. Recursively Summarizing Books with Human Feedback
    2021/09/22 by Jeff Wu, Wu, Jeff, Long Ouyang +11 · 15 citations
    Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  14. LLM Critics Help Catch LLM Bugs
    2024/06/28 by Nat McAleese, McAleese, Nat, Rai Michael Pokorny +9 · 26 citations
    Computer Science · Agricultural and Biological Sciences · #Law, AI, and Intellectual Property #Phytoplasmas and Hemiptera pathogens #Vector-Borne Animal Diseases
  15. Auditing language models for hidden objectives
    2025/03/14 by Samuel D. Marks, Samuel Marks, Johannes Treutlein +71 · 1 voice · 16 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  16. Active Reinforcement Learning: Observing Rewards at a Cost
    2020/11/13 by David Krueger, Jan Leike, Krueger, David +5 · 5 citations
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
  17. Hidden Incentives for Auto-Induced Distributional Shift
    2020/09/19 by David Krueger, Krueger, David, Tegan Maharaj +3 · 4 citations
    Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Smart Grid Energy Management
  18. Thompson Sampling is Asymptotically Optimal in General Environments
    2016/02/25 by Jan Leike, Leike, Jan, Tor Lattimore +5 · 5 citations
    Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Fractional Differential Equations Solutions #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms
  19. Ranking Templates for Linear Loops
    2014/01/21 by Jan Leike, Matthias Heizmann, Leike, Jan +1 · 1 citation
    Computer Science · #Advanced Software Engineering Methodologies #FOS: Computer and information sciences #Formal Methods in Verification #Logic in Computer Science (cs.LO) #Logic, programming, and type systems
  20. Quantifying Differences in Reward Functions
    2020/06/24 by Adam Gleave, Michael D. Dennis, Gleave, Adam +7 · 1 citation
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  21. Forecasting Rare Language Model Behaviors
    2025/02/24 by Jones, Erik, Meg Tong, Jesse Mu +16 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification
  22. A Formal Solution to the Grain of Truth Problem
    2016/09/16 by Jan Leike, Jessica Taylor, Leike, Jan +3 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computability, Logic, AI Algorithms #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Logic, Reasoning, and Knowledge #Machine Learning (cs.LG) #Machine Learning and Algorithms