vix.ing · top · new · best · stats · spec

Owain Evans

  1. Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
    2025/12/10 by Jan Betley, Betley, Jan, Jorio Cocola +12 · 33 voices · 7 citations
    Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
  2. Training large language models on narrow tasks can lead to broad misalignment
    2025/02/24 by Jan Betley, Betley, Jan, Daniel Tan +17 · 48 voices · 49 citations
    Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
  3. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
    2018/02/20 by Miles Brundage, Brundage, Miles, Shahar Avin +49 · 6 voices · 19 citations
    #cs.AI #cs.CR #cs.CY
  4. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
    #cs.AI #cs.LG #stat.ML
  5. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
    2023/09/21 by Lukas Berglund, Meg Tong, Berglund, Lukas +11 · 11 voices · 75 citations
    Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
  6. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
    2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
    #cs.LG #cs.AI
  7. When Will AI Exceed Human Performance? Evidence from AI Experts
    2017/05/24 by Katja Grace, Grace, Katja, John Salvatier +7 · 4 voices · 7 citations
    #cs.AI #cs.CY
  8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
    2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 129 citations
    #cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
  9. TruthfulQA: Measuring How Models Mimic Human Falsehoods
    2021/09/08 by Stephanie Lin, Lin, Stephanie, Jacob Hilton +3 · 419 citations
    Computer Science · Medicine · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education
  10. Teaching Models to Express Their Uncertainty in Words
    2022/05/28 by Stephanie Lin, Jacob Hilton, Lin, Stephanie +3 · 1 voice · 83 citations
    Computer Science · #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  11. Persona Vectors: Monitoring and Controlling Character Traits in Language Models
    2025/07/29 by Runjin Chen, Andy Arditi, Chen, Runjin +7 · 15 voices · 88 citations
    #cs.CL #cs.LG
  12. Tell me about yourself: LLMs are aware of their learned behaviors
    2025/01/19 by Jan Betley, J. Nicholas Betley, Betley, Jan +11 · 6 voices · 31 citations
    Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
  13. Looking Inward: Language Models Can Learn About Themselves by Introspection
    2024/10/17 by Felix J Binder, James Chua, Binder, Felix J +15 · 4 voices · 36 citations
    Computer Science · #Natural Language Processing Techniques
  14. Negation Neglect: When models fail to learn negations in training
    2026/05/13 by Harry Mayne, Lev McKinney, Jan Dubiński +3 · 8 voices · 1 citation
    #cs.CL #cs.AI #cs.LG
  15. Truthful AI: Developing and governing AI that does not lie
    2021/10/13 by Owain Evans, Evans, Owain, Owen Cotton-Barratt +13 · 12 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0
  16. Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
    2017/07/17 by William S. Saunders, Girish Sastry, Saunders, William +5 · 11 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics
  17. Taken out of context: On measuring situational awareness in LLMs
    2023/09/01 by Lukas Berglund, Asa Cooper Stickland, Berglund, Lukas +13 · 13 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  18. How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
    2023/09/26 by Lorenzo Pacchiardi, Pacchiardi, Lorenzo, Alex Chan +13 · 13 citations
    Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  19. Forecasting Future World Events with Neural Networks
    2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 9 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  20. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
    2025/08/24 by Mia Taylor, James Chua, Taylor, Mia +7 · 4 voices · 9 citations
    #cs.AI
  21. Are DeepSeek R1 And Other Reasoning Models More Faithful?
    2025/01/14 by James Chua, Chua, James, Owain Evans +1 · 19 citations
    Computer Science · #Advanced Database Systems and Queries #Time Series Analysis and Forecasting #Data Visualization and Analytics
  22. Active Reinforcement Learning: Observing Rewards at a Cost
    2020/11/13 by David Krueger, Krueger, David, Jan Leike +5 · 5 citations
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
  23. Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
    2024/07/05 by Rudolf Laine, Bilal Chughtai, Laine, Rudolf +15 · 11 citations
    Business, Management and Accounting · Health Professions · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research
  24. Active Reinforcement Learning with Monte-Carlo Tree Search
    2018/03/13 by Sebastian Schulze, Owain Evans, Schulze, Sebastian +1 · 4 citations
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Multi-Objective Optimization Algorithms #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  25. Towards evaluations-based safety cases for AI scheming
    2024/10/29 by Mikita Balesni, Balesni, Mikita, Marius Hobbhahn +29 · 7 citations
    Engineering · Health Professions · Decision Sciences · #Safety Systems Engineering in Autonomy #Occupational Health and Safety Research #Risk and Safety Analysis
  26. Lessons from Studying Two-Hop Latent Reasoning
    2024/11/25 by Mikita Balesni, Tomek Korbak, Balesni, Mikita +3 · 5 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences
  27. Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
    2025/06/16 by James Chua, J. Nicholas Betley, Chua, James +4 · 15 citations
    Social Sciences · Computer Science · #Crime, Illicit Activities, and Governance #Cybercrime and Law Enforcement Studies
  28. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
    2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  29. Tell, don't show: Declarative facts influence how LLMs generalize
    2023/12/12 by Alexander Meinke, Owain Evans, Meinke, Alexander +1 · 2 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling
  30. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
    2025/12/17 by Adam Karvonen, James Chua, Karvonen, Adam +19 · 1 voice · 1 citation
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
  31. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
    2026/07/15 by Jan Betley, Johannes Treutlein, Jan Dubiński +5 · 4 voices
    #cs.LG #cs.AI #cs.CR