Cem Anil
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 100 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Solving Quantitative Reasoning Problems with Language Models
2022/06/29 by Aitor Lewkowycz, Lewkowycz, Aitor, Anders Andreassen +25 · 1 voice · 303 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
2023/08/07 by Roger Grosse, Grosse, Roger, Juhan Bae +31 · 3 voices · 57 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +89 · 8 voices · 41 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
- Sorting out Lipschitz function approximation
2018/11/13 by Cem Anil, Anil, Cem, James Lucas +3 · 18 citations
Computer Science · Engineering · #Advanced Image Processing Techniques #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sparse and Compressive Sensing Techniques
- Exploring Length Generalization in Large Language Models
2022/07/11 by Cem Anil, Yuhuai Wu, Anil, Cem +17 · 23 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
2024/06/20 by Johannes Treutlein, Dami Choi, Treutlein, Johannes +12 · 4 voices
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
2025/12/05 by Igor Shilov, Alex Cloud, Shilov, Igor +13 · 2 voices · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #cs.LG
- Learning to Give Checkable Answers with Prover-Verifier Games
2021/08/27 by Cem Anil, Anil, Cem, Guodong Zhang +5 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Reinforcement Learning in Robotics