Betley, Jan
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025/12/10 by Jan Betley, Jorio Cocola, Betley, Jan +12 · 33 voices · 7 citations
Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
- Training large language models on narrow tasks can lead to broad misalignment
2025/02/24 by Jan Betley, Betley, Jan, Daniel Tan +17 · 48 voices · 49 citations
Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
#cs.LG #cs.AI
- Tell me about yourself: LLMs are aware of their learned behaviors
2025/01/19 by J. Nicholas Betley, Jan Betley, Betley, Jan +11 · 6 voices · 31 citations
Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
- School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
2025/08/24 by Mia Taylor, Taylor, Mia, James Chua +7 · 4 voices · 9 citations
#cs.AI
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
2024/07/05 by Rudolf Laine, Bilal Chughtai, Laine, Rudolf +15 · 11 citations
Business, Management and Accounting · Health Professions · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
2025/06/16 by James Chua, J. Nicholas Betley, Chua, James +4 · 15 citations
Social Sciences · Computer Science · #Crime, Illicit Activities, and Governance #Cybercrime and Law Enforcement Studies
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
2024/06/20 by Johannes Treutlein, Treutlein, Johannes, Dami Choi +12 · 4 voices
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG