Meg Tong
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
2023/10/20 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +36 · 10 voices · 228 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
2023/09/21 by Lukas Berglund, Meg Tong, Berglund, Lukas +11 · 11 voices · 77 citations
Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
- Steering Llama 2 via Contrastive Activation Addition
2023/12/09 by Nick Gabrieli, Panickssery, Nina, Gabrieli, Nick +8 · 160 citations
Computer Science · #Interactive and Immersive Displays
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +89 · 8 voices · 37 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
- Taken out of context: On measuring situational awareness in LLMs
2023/09/01 by Lukas Berglund, Berglund, Lukas, Asa Cooper Stickland +13 · 13 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- Auditing language models for hidden objectives
2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 16 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Forecasting Rare Language Model Behaviors
2025/02/24 by Jones, Erik, Meg Tong, Jesse Mu +16 · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification