Christopher Olah
- Affective Coherence Monitoring for Transformer-Based Language Models
2022/12/15 by Yuntao Bai, Bai, Yuntao, Saurav Kadavath +102 · 12 voices · 475 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL
- Conditional Image Synthesis With Auxiliary Classifier GANs
2016/10/30 by Augustus Odena, Christopher Olah, Odena, Augustus +3 · 250 citations
Biochemistry, Genetics and Molecular Biology · Engineering · Computer Science · #Cell Image Analysis Techniques #Image Processing Techniques and Applications #Advanced Image Processing Techniques
- Governance Architecture for Neural Network Superposition: A Structural Solution to Hallucination via Routing and Interference Filtering
2022/09/21 by Nelson Elhage, Tristan Hume, Elhage, Nelson +29 · 2 voices · 147 citations
Computer Science · Physics and Astronomy · #Explainable Artificial Intelligence (XAI) #Model Reduction and Neural Networks #Neural Networks and Applications #cs.LG
- Discovering Language Model Behaviors with Model-Written Evaluations
2022/12/19 by Ethan Perez, Perez, Ethan, Sam Ringer +123 · 129 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- The Capacity for Moral Self-Correction in Large Language Models
2023/02/15 by Deep Ganguli, Ganguli, Deep, Amanda Askell +100 · 1 voice · 11 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling #cs.CL
- Measuring Progress on Scalable Oversight for Large Language Models
2022/11/04 by Samuel R. Bowman, Bowman, Samuel R., Jeeyoon Hyun +88 · 22 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multi-Agent Systems and Negotiation #Speech and dialogue systems #Topic Modeling
- Auditing language models for hidden objectives
2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 16 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Changing Model Behavior at Test-Time Using Reinforcement Learning
2017/02/24 by Augustus Odena, Dieterich Lawson, Odena, Augustus +3 · 1 citation
Computer Science · #Machine Learning and Data Classification #Data Stream Mining Techniques #Machine Learning and Algorithms