Owain Evans
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025/12/10 by Jan Betley, Betley, Jan, Jorio Cocola +12 · 33 voices · 7 citations
Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
- Training large language models on narrow tasks can lead to broad misalignment
2025/02/24 by Jan Betley, Betley, Jan, Daniel Tan +17 · 48 voices · 49 citations
Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
- The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
2018/02/20 by Miles Brundage, Brundage, Miles, Shahar Avin +49 · 6 voices · 19 citations
#cs.AI #cs.CR #cs.CY
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
2023/09/21 by Lukas Berglund, Meg Tong, Berglund, Lukas +11 · 11 voices · 75 citations
Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
#cs.LG #cs.AI
- When Will AI Exceed Human Performance? Evidence from AI Experts
2017/05/24 by Katja Grace, Grace, Katja, John Salvatier +7 · 4 voices · 7 citations
#cs.AI #cs.CY
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 129 citations
#cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
2021/09/08 by Stephanie Lin, Lin, Stephanie, Jacob Hilton +3 · 419 citations
Computer Science · Medicine · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education
- Teaching Models to Express Their Uncertainty in Words
2022/05/28 by Stephanie Lin, Jacob Hilton, Lin, Stephanie +3 · 1 voice · 83 citations
Computer Science · #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
2025/07/29 by Runjin Chen, Andy Arditi, Chen, Runjin +7 · 15 voices · 88 citations
#cs.CL #cs.LG
- Tell me about yourself: LLMs are aware of their learned behaviors
2025/01/19 by Jan Betley, J. Nicholas Betley, Betley, Jan +11 · 6 voices · 31 citations
Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
- Looking Inward: Language Models Can Learn About Themselves by Introspection
2024/10/17 by Felix J Binder, James Chua, Binder, Felix J +15 · 4 voices · 36 citations
Computer Science · #Natural Language Processing Techniques
- Negation Neglect: When models fail to learn negations in training
2026/05/13 by Harry Mayne, Lev McKinney, Jan Dubiński +3 · 8 voices · 1 citation
#cs.CL #cs.AI #cs.LG
- Truthful AI: Developing and governing AI that does not lie
2021/10/13 by Owain Evans, Evans, Owain, Owen Cotton-Barratt +13 · 12 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
2017/07/17 by William S. Saunders, Girish Sastry, Saunders, William +5 · 11 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics
- Taken out of context: On measuring situational awareness in LLMs
2023/09/01 by Lukas Berglund, Asa Cooper Stickland, Berglund, Lukas +13 · 13 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
2023/09/26 by Lorenzo Pacchiardi, Pacchiardi, Lorenzo, Alex Chan +13 · 13 citations
Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Forecasting Future World Events with Neural Networks
2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 9 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
2025/08/24 by Mia Taylor, James Chua, Taylor, Mia +7 · 4 voices · 9 citations
#cs.AI
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
2025/01/14 by James Chua, Chua, James, Owain Evans +1 · 19 citations
Computer Science · #Advanced Database Systems and Queries #Time Series Analysis and Forecasting #Data Visualization and Analytics
- Active Reinforcement Learning: Observing Rewards at a Cost
2020/11/13 by David Krueger, Krueger, David, Jan Leike +5 · 5 citations
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
2024/07/05 by Rudolf Laine, Bilal Chughtai, Laine, Rudolf +15 · 11 citations
Business, Management and Accounting · Health Professions · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research
- Active Reinforcement Learning with Monte-Carlo Tree Search
2018/03/13 by Sebastian Schulze, Owain Evans, Schulze, Sebastian +1 · 4 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Multi-Objective Optimization Algorithms #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Towards evaluations-based safety cases for AI scheming
2024/10/29 by Mikita Balesni, Balesni, Mikita, Marius Hobbhahn +29 · 7 citations
Engineering · Health Professions · Decision Sciences · #Safety Systems Engineering in Autonomy #Occupational Health and Safety Research #Risk and Safety Analysis
- Lessons from Studying Two-Hop Latent Reasoning
2024/11/25 by Mikita Balesni, Tomek Korbak, Balesni, Mikita +3 · 5 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
2025/06/16 by James Chua, J. Nicholas Betley, Chua, James +4 · 15 citations
Social Sciences · Computer Science · #Crime, Illicit Activities, and Governance #Cybercrime and Law Enforcement Studies
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Tell, don't show: Declarative facts influence how LLMs generalize
2023/12/12 by Alexander Meinke, Owain Evans, Meinke, Alexander +1 · 2 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
2025/12/17 by Adam Karvonen, James Chua, Karvonen, Adam +19 · 1 voice · 1 citation
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
2026/07/15 by Jan Betley, Johannes Treutlein, Jan Dubiński +5 · 4 voices
#cs.LG #cs.AI #cs.CR