vix.ing · top · new · best · stats · spec

Evans, Owain

  1. Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
    2025/12/10 by Jan Betley, Betley, Jan, Jorio Cocola +12 · 33 voices · 7 citations
    Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
  2. Training large language models on narrow tasks can lead to broad misalignment
    2025/02/24 by Jan Betley, Betley, Jan, Daniel Tan +17 · 48 voices · 50 citations
    Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
  3. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
    2018/02/20 by Miles Brundage, Brundage, Miles, Shahar Avin +49 · 6 voices · 20 citations
    #cs.AI #cs.CR #cs.CY
  4. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Korbak, Tomek, Mikita Balesni +79 · 25 voices · 68 citations
    #cs.AI #cs.LG #stat.ML
  5. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
    2023/09/21 by Lukas Berglund, Berglund, Lukas, Meg Tong +11 · 11 voices · 79 citations
    Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
  6. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
    2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 28 citations
    #cs.LG #cs.AI
  7. When Will AI Exceed Human Performance? Evidence from AI Experts
    2017/05/24 by Katja Grace, John Salvatier, Grace, Katja +7 · 4 voices · 7 citations
    #cs.AI #cs.CY
  8. TruthfulQA: Measuring How Models Mimic Human Falsehoods
    2021/09/08 by Stephanie Lin, Lin, Stephanie, Jacob Hilton +3 · 432 citations
    Computer Science · Medicine · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education
  9. Teaching Models to Express Their Uncertainty in Words
    2022/05/28 by Stephanie Lin, Lin, Stephanie, Jacob Hilton +3 · 1 voice · 85 citations
    Computer Science · #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  10. Persona Vectors: Monitoring and Controlling Character Traits in Language Models
    2025/07/29 by Runjin Chen, Chen, Runjin, Andy Arditi +7 · 15 voices · 88 citations
    #cs.CL #cs.LG
  11. Tell me about yourself: LLMs are aware of their learned behaviors
    2025/01/19 by Jan Betley, J. Nicholas Betley, Betley, Jan +11 · 6 voices · 31 citations
    Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
  12. Looking Inward: Language Models Can Learn About Themselves by Introspection
    2024/10/17 by Felix J Binder, Binder, Felix J, James Chua +15 · 4 voices · 37 citations
    Computer Science · #Natural Language Processing Techniques
  13. Truthful AI: Developing and governing AI that does not lie
    2021/10/13 by Owain Evans, Owen Cotton-Barratt, Evans, Owain +13 · 12 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #I.2.0
  14. Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
    2017/07/17 by William S. Saunders, Girish Sastry, Saunders, William +5 · 11 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics
  15. Taken out of context: On measuring situational awareness in LLMs
    2023/09/01 by Lukas Berglund, Asa Cooper Stickland, Berglund, Lukas +13 · 15 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  16. How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions
    2023/09/26 by Lorenzo Pacchiardi, Alex Chan, Pacchiardi, Lorenzo +13 · 14 citations
    Computer Science · Psychology · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Deception detection and forensic psychology #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  17. Forecasting Future World Events with Neural Networks
    2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 10 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  18. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
    2025/08/24 by Mia Taylor, James Chua, Taylor, Mia +7 · 4 voices · 9 citations
    #cs.AI
  19. Are DeepSeek R1 And Other Reasoning Models More Faithful?
    2025/01/14 by James Chua, Chua, James, Owain Evans +1 · 19 citations
    Computer Science · #Advanced Database Systems and Queries #Time Series Analysis and Forecasting #Data Visualization and Analytics
  20. Active Reinforcement Learning: Observing Rewards at a Cost
    2020/11/13 by David Krueger, Krueger, David, Jan Leike +5 · 5 citations
    Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Data Stream Mining Techniques
  21. Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
    2024/07/05 by Rudolf Laine, Bilal Chughtai, Laine, Rudolf +15 · 11 citations
    Business, Management and Accounting · Health Professions · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research
  22. Active Reinforcement Learning with Monte-Carlo Tree Search
    2018/03/13 by Sebastian Schulze, Schulze, Sebastian, Owain Evans +1 · 4 citations
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Advanced Multi-Objective Optimization Algorithms #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  23. Towards evaluations-based safety cases for AI scheming
    2024/10/29 by Mikita Balesni, Balesni, Mikita, Marius Hobbhahn +29 · 7 citations
    Engineering · Health Professions · Decision Sciences · #Safety Systems Engineering in Autonomy #Occupational Health and Safety Research #Risk and Safety Analysis
  24. Lessons from Studying Two-Hop Latent Reasoning
    2024/11/25 by Mikita Balesni, Tomek Korbak, Balesni, Mikita +3 · 5 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences
  25. Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
    2025/06/16 by James Chua, J. Nicholas Betley, Chua, James +4 · 15 citations
    Social Sciences · Computer Science · #Crime, Illicit Activities, and Governance #Cybercrime and Law Enforcement Studies
  26. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
    2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  27. Learning the Preferences of Ignorant, Inconsistent Agents
    2015/12/18 by Owain Evans, Evans, Owain, Andreas Stuhlmueller +3 · 1 citation
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Complex Systems and Decision Making #Decision-Making and Behavioral Economics #FOS: Computer and information sciences
  28. Tell, don't show: Declarative facts influence how LLMs generalize
    2023/12/12 by Alexander Meinke, Owain Evans, Meinke, Alexander +1 · 2 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling
  29. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
    2025/12/17 by Adam Karvonen, James Chua, Karvonen, Adam +19 · 1 voice · 1 citation
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG