Austin Meek
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Marks, Samuel +71 · 1 voice · 15 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
2024/10/09 by Michael Lan, Lan, Michael, Philip Torr +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
2025/10/31 by Austin Meek, Meek, Austin, Eitan Sprejer +7 · 2 citations
Computer Science · Decision Sciences · #Advanced Software Engineering Methodologies #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Personal Information Management and User Behavior #Topic Modeling
- Inducing Human-like Biases in Moral Reasoning Language Models
2024/11/23 by Artem Karpov, Karpov, Artem, Seong Hah Cho +9 · 1 citation
Neuroscience · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Psychology of Moral and Emotional Judgment