vix.ing · top · new · best · stats · spec

Davies, Xander

  1. Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
    2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 19 citations
    Medicine · #Medical Imaging and Pathology Studies
  2. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 93 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  3. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
    2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 59 citations
    Computer Science · #Blockchain Technology Applications and Security
  4. SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
    2024/10/14 by Nie, Yuzhou, Wang, Zhun, Yang, Yu +7 · 18 citations
    #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  5. Unifying Grokking and Double Descent
    2023/03/10 by Davies, Xander, Langosco, Lauro, Krueger, David · 7 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
    2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
    #cs.AI #cs.CL #cs.CY
  7. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
    2025/08/08 by Kyle O'Brien, Kyle O’Brien, O'Brien, Kyle +18 · 2 voices · 14 citations
    Computer Science · Social Sciences · #cs.LG #cs.AI
  8. Discovering Variable Binding Circuitry with Desiderata
    2023/07/07 by Xander Davies, Davies, Xander, Max Nadeau +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  9. Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
    2025/02/20 by Davies, Xander, Winsor, Eric, Souly, Alexandra +4 · 4 citations
    #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  10. Existing Large Language Model Unlearning Evaluations Are Inconclusive
    2025/05/31 by Zhili Feng, Ye Xu, Feng, Zhili +13 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  11. STACK: Adversarial Attacks on LLM Safeguard Pipelines
    2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
  12. Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
    2025/10/26 by Bazinska, Julia, Mathys, Max, Casucci, Francesco +4 · 2 citations
    #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. An Example Safety Case for Safeguards Against Misuse
    2025/05/23 by Clymer, Joshua, Weinbaum, Jonah, Kirk, Robert +3 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)