Jan Betley
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025/12/10 by Jan Betley, Jorio Cocola, Betley, Jan +12 · 33 voices · 7 citations
Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
- Training large language models on narrow tasks can lead to broad misalignment
2025/02/24 by Jan Betley, Daniel Tan, Betley, Jan +17 · 48 voices · 49 citations
Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
#cs.LG #cs.AI
- Tell me about yourself: LLMs are aware of their learned behaviors
2025/01/19 by J. Nicholas Betley, Jan Betley, Xuchan Bao +11 · 6 voices · 31 citations
Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
- School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
2025/08/24 by Mia Taylor, James Chua, Taylor, Mia +7 · 4 voices · 9 citations
#cs.AI
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
2026/07/15 by Jan Betley, Johannes Treutlein, Jan Dubiński +5 · 4 voices
#cs.LG #cs.AI #cs.CR