vix.ing · top · new · best · stats · spec

Betley, Jan

  1. Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
    2025/12/10 by Jan Betley, Jorio Cocola, Betley, Jan +12 · 33 voices · 7 citations
    Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
  2. Training large language models on narrow tasks can lead to broad misalignment
    2025/02/24 by Jan Betley, Betley, Jan, Daniel Tan +17 · 48 voices · 49 citations
    Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
  3. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
    2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
    #cs.LG #cs.AI
  4. Tell me about yourself: LLMs are aware of their learned behaviors
    2025/01/19 by J. Nicholas Betley, Jan Betley, Betley, Jan +11 · 6 voices · 31 citations
    Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
  5. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
    2025/08/24 by Mia Taylor, Taylor, Mia, James Chua +7 · 4 voices · 9 citations
    #cs.AI
  6. Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
    2024/07/05 by Rudolf Laine, Bilal Chughtai, Laine, Rudolf +15 · 11 citations
    Business, Management and Accounting · Health Professions · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research
  7. Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
    2025/06/16 by James Chua, J. Nicholas Betley, Chua, James +4 · 15 citations
    Social Sciences · Computer Science · #Crime, Illicit Activities, and Governance #Cybercrime and Law Enforcement Studies
  8. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
    2024/06/20 by Johannes Treutlein, Treutlein, Johannes, Dami Choi +12 · 4 voices
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG