vix.ing · top · new · best · stats · spec

Jan Betley

  1. Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
    2025/12/10 by Jan Betley, Jorio Cocola, Betley, Jan +12 · 33 voices · 7 citations
    Computer Science · #Authorship Attribution and Profiling #Topic Modeling #Machine Learning in Healthcare
  2. Training large language models on narrow tasks can lead to broad misalignment
    2025/02/24 by Jan Betley, Daniel Tan, Betley, Jan +17 · 48 voices · 49 citations
    Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
  3. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
    2025/07/20 by Alex Cloud, Cloud, Alex, Minh Le +13 · 36 voices · 26 citations
    #cs.LG #cs.AI
  4. Tell me about yourself: LLMs are aware of their learned behaviors
    2025/01/19 by J. Nicholas Betley, Jan Betley, Xuchan Bao +11 · 6 voices · 31 citations
    Computer Science · #Digital Rights Management and Security #Library Science and Information Systems #Semantic Web and Ontologies #cs.AI #cs.CL #cs.CR #cs.LG
  5. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
    2025/08/24 by Mia Taylor, James Chua, Taylor, Mia +7 · 4 voices · 9 citations
    #cs.AI
  6. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
    2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  7. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
    2026/07/15 by Jan Betley, Johannes Treutlein, Jan Dubiński +5 · 4 voices
    #cs.LG #cs.AI #cs.CR