vix.ing · top · new · best · stats · spec

Daniel M. Ziegler

  1. Language Models are Few-Shot Learners
    2020/05/28 by T. B. Brown, Tom B. Brown, Benjamin Mann +61 · 16 voices · 3796 citations
    Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.CL
  2. Measuring AI Ability to Complete Long Software Tasks
    2025/03/18 by Thomas Kwa, Kwa, Thomas, Ben West +48 · 26 voices · 40 citations
    #cs.AI #cs.LG
  3. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  4. Learning to summarize from human feedback
    2020/09/02 by Nisan Stiennon, Long Ouyang, Stiennon, Nisan +16 · 3 voices · 338 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  5. Fine-Tuning Language Models from Human Preferences
    2019/09/18 by Daniel M. Ziegler, Nisan Stiennon, Ziegler, Daniel M. +13 · 251 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  6. Scaling Laws for Autoregressive Generative Modeling
    2020/10/28 by Tom Henighan, Jared Kaplan, Henighan, Tom +35 · 75 citations
    Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
  7. Recursively Summarizing Books with Human Feedback
    2021/09/22 by Jeff Wu, Wu, Jeff, Long Ouyang +11 · 15 citations
    Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  8. Auditing language models for hidden objectives
    2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 16 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  9. Adversarial Training for High-Stakes Reliability
    2022/05/03 by Daniel M. Ziegler, Seraphina Nix, Ziegler, Daniel M. +21 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Ethics and Social Impacts of AI