vix.ing · top · new · best · stats · spec

Meg Tong

  1. AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
    2023/10/20 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +36 · 10 voices · 228 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
  2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
    2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
  3. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
    2023/09/21 by Lukas Berglund, Meg Tong, Berglund, Lukas +11 · 11 voices · 77 citations
    Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
  4. Steering Llama 2 via Contrastive Activation Addition
    2023/12/09 by Nick Gabrieli, Panickssery, Nina, Gabrieli, Nick +8 · 160 citations
    Computer Science · #Interactive and Immersive Displays
  5. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
    2025/01/31 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +89 · 8 voices · 37 citations
    Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
  6. Taken out of context: On measuring situational awareness in LLMs
    2023/09/01 by Lukas Berglund, Berglund, Lukas, Asa Cooper Stickland +13 · 13 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  7. Auditing language models for hidden objectives
    2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 16 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  8. Forecasting Rare Language Model Behaviors
    2025/02/24 by Jones, Erik, Meg Tong, Jesse Mu +16 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification