vix.ing · top · new · best · stats · spec

Samuel Marks

  1. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
    2025/07/20 by Alex Cloud, Minh Le, Cloud, Alex +13 · 36 voices · 28 citations
    #cs.LG #cs.AI
  2. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
    2023/10/10 by Samuel Marks, Max Tegmark, Marks, Samuel +1 · 7 voices · 132 citations
    Computer Science · Social Sciences · #Topic Modeling #Natural Language Processing Techniques #Computational and Text Analysis Methods
  3. Unsupervised Elicitation of Language Models
    2025/06/11 by Jiaxin Wen, Zachary Ankner, Wen, Jiaxin +23 · 15 voices · 4 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Machine Learning and Data Classification
  4. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 93 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  5. Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
    2026/07/03 by Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks · 6 voices
    #cs.CL #cs.AI #cs.LG
  6. Robustly Improving LLM Fairness in Realistic Settings via Interpretability
    2025/06/12 by Adam Karvonen, Samuel Marks, Karvonen, Adam +1 · 3 voices · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  7. The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
    2024/08/02 by Aaron Mueller, Jannik Brinkmann, Mueller, Aaron +25 · 3 voices · 15 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #European and International Law Studies #FOS: Computer and information sciences #Machine Learning (cs.LG) #Qualitative Comparative Analysis Research #cs.AI #cs.LG
  8. Auditing language models for hidden objectives
    2025/03/14 by Samuel Marks, Samuel D. Marks, Marks, Samuel +71 · 1 voice · 17 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
  9. Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
    2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  10. Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
    2025/10/06 by Nevan Wichers, Aram Ebtekar, Wichers, Nevan +19 · 4 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Software Engineering Research #Topic Modeling
  11. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
    2025/12/17 by Adam Karvonen, Karvonen, Adam, James Chua +19 · 1 voice · 1 citation
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG