Lindsey, Jack
- Open Problems in Mechanistic Interpretability
2025/01/27 by Lee Sharkey, Sharkey, Lee, Bilal Chughtai +58 · 7 voices · 52 citations
Computer Science · #Natural Language Processing Techniques #Statistical and Computational Modeling #cs.LG
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
2025/07/29 by Runjin Chen, Chen, Runjin, Andy Arditi +7 · 15 voices · 88 citations
#cs.CL #cs.LG
- Humanity's Last Exam
2025/01/24 by Long Phan, Alice Gatti, Phan, Long +2240 · 9 voices · 130 citations
#cs.LG #cs.AI #cs.CL
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Johannes Treutlein +71 · 1 voice · 22 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Learning to Learn with Feedback and Local Plasticity
2020/06/16 by Jack Lindsey, Lindsey, Jack, Ashok Litwin-Kumar +1 · 2 citations
Computer Science · Engineering · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Advanced Memory and Neural Computing
- Natural Emergent Misalignment from Reward Hacking in Production RL
2025/11/23 by MacDiarmid, Monte, Wright, Benjamin, Uesato, Jonathan +19 · 8 citations
Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Software Engineering Research