Daniel M. Ziegler
- Language Models are Few-Shot Learners
2020/05/28 by T. B. Brown, Tom B. Brown, Benjamin Mann +61 · 16 voices · 3796 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.CL
- Measuring AI Ability to Complete Long Software Tasks
2025/03/18 by Thomas Kwa, Kwa, Thomas, Ben West +48 · 26 voices · 40 citations
#cs.AI #cs.LG
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Learning to summarize from human feedback
2020/09/02 by Nisan Stiennon, Long Ouyang, Stiennon, Nisan +16 · 3 voices · 338 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- Fine-Tuning Language Models from Human Preferences
2019/09/18 by Daniel M. Ziegler, Nisan Stiennon, Ziegler, Daniel M. +13 · 251 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
- Scaling Laws for Autoregressive Generative Modeling
2020/10/28 by Tom Henighan, Jared Kaplan, Henighan, Tom +35 · 75 citations
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
- Recursively Summarizing Books with Human Feedback
2021/09/22 by Jeff Wu, Wu, Jeff, Long Ouyang +11 · 15 citations
Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 16 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Adversarial Training for High-Stakes Reliability
2022/05/03 by Daniel M. Ziegler, Seraphina Nix, Ziegler, Daniel M. +21 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Ethics and Social Impacts of AI