Davies, Xander
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 19 citations
Medicine · #Medical Imaging and Pathology Studies
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 93 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 59 citations
Computer Science · #Blockchain Technology Applications and Security
- SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
2024/10/14 by Nie, Yuzhou, Wang, Zhun, Yang, Yu +7 · 18 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Unifying Grokking and Double Descent
2023/03/10 by Davies, Xander, Langosco, Lauro, Krueger, David · 7 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
2025/08/08 by Kyle O'Brien, Kyle O’Brien, O'Brien, Kyle +18 · 2 voices · 14 citations
Computer Science · Social Sciences · #cs.LG #cs.AI
- Discovering Variable Binding Circuitry with Desiderata
2023/07/07 by Xander Davies, Davies, Xander, Max Nadeau +7 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
2025/02/20 by Davies, Xander, Winsor, Eric, Souly, Alexandra +4 · 4 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
2025/05/31 by Zhili Feng, Ye Xu, Feng, Zhili +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
2025/10/26 by Bazinska, Julia, Mathys, Max, Casucci, Francesco +4 · 2 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- An Example Safety Case for Safeguards Against Misuse
2025/05/23 by Clymer, Joshua, Weinbaum, Jonah, Kirk, Robert +3 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)