vix.ing · top · new · best · stats · spec

Xander Davies

  1. Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
    2025/10/08 by Alexandra Souly, Javier Rando, Souly, Alexandra +25 · 34 voices · 18 citations
    Medicine · #Medical Imaging and Pathology Studies
  2. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 88 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  3. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
    2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 56 citations
    Computer Science · #Blockchain Technology Applications and Security
  4. Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
    2025/07/28 by Andy Zou, Maxwell Lin, Zou, Andy +31 · 5 voices · 8 citations
    #cs.AI #cs.CL #cs.CY
  5. Discovering Variable Binding Circuitry with Desiderata
    2023/07/07 by Xander Davies, Max Nadeau, Davies, Xander +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  6. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
    2025/08/08 by Kyle O’Brien, Kyle O'Brien, Stephen Casper +18 · 2 voices · 13 citations
    Computer Science · Social Sciences · #cs.LG #cs.AI
  7. Existing Large Language Model Unlearning Evaluations Are Inconclusive
    2025/05/31 by Zhili Feng, Ye Xu, Feng, Zhili +13 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  8. STACK: Adversarial Attacks on LLM Safeguard Pipelines
    2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI