Xander Davies
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Javier Rando, Souly, Alexandra +25 · 34 voices · 18 citations
Medicine · #Medical Imaging and Pathology Studies
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 88 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 56 citations
Computer Science · #Blockchain Technology Applications and Security
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Maxwell Lin, Zou, Andy +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Discovering Variable Binding Circuitry with Desiderata
2023/07/07 by Xander Davies, Max Nadeau, Davies, Xander +7 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
2025/08/08 by Kyle O’Brien, Kyle O'Brien, Stephen Casper +18 · 2 voices · 13 citations
Computer Science · Social Sciences · #cs.LG #cs.AI
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
2025/05/31 by Zhili Feng, Ye Xu, Feng, Zhili +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI