Ansh Radhakrishnan
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 100 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Reasoning Models Don't Always Say What They Think
2025/05/08 by Yanda Chen, Chen, Yanda, Joe Benton +27 · 6 voices · 100 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
- Measuring Faithfulness in Chain-of-Thought Reasoning
2023/07/17 by Tamera Lanham, Lanham, Tamera, Anna Chen +59 · 4 voices · 87 citations
Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG
- Debating with More Persuasive LLMs Leads to More Truthful Answers
2024/02/09 by Akbir Khan, John Hughes, Khan, Akbir +17 · 53 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Legal Education and Practice Innovations
- Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
2023/07/17 by Ansh Radhakrishnan, Karina Nguyen, Radhakrishnan, Ansh +45 · 5 citations
Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
2024/11/26 by Jiaxin Wen, Vivek Hebbar, Wen, Jiaxin +20 · 5 citations
Computer Science · #Blockchain Technology Applications and Security