Barez, Fazl
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Best-of-N Jailbreaking
2024/12/04 by John Hughes, John D. Hughes, Hughes, John +18 · 17 voices · 14 citations
Computer Science · #Digital and Cyber Forensics #cs.AI #cs.CL #cs.LG
- The Larger They Are, the Harder They Fail: Language Models do not Recognize Identifier Swaps in Python
2023/05/24 by Antonio Valerio Miceli-Barone, Miceli-Barone, Antonio Valerio, Fazl Barez +5 · 3 voices · 4 citations
Computer Science · #Software Engineering Research #Topic Modeling #Computational Physics and Python Applications
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
2025/07/03 by Zhu, Yuxuan, Jin, Tengjun, Pruksachatkun, Yada +22 · 5 voices · 12 citations
#A.1 #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.m
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
2024/06/14 by Denison, Carson, MacDiarmid, Monte, Barez, Fazl +11 · 24 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Towards Interpreting Visual Information Processing in Vision-Language Models
2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
2025/02/18 by Adi Simhi, Simhi, Adi, Itay Itzhak +7 · 1 voice · 13 citations
Computer Science · Neuroscience · #Blockchain Technology Applications and Security #Hallucinations in medical conditions #cs.CL
- Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
2024/11/02 by Luke Marks, Alasdair Paren, Marks, Luke +5 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
- Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
2023/05/27 by Jason Hoelscher-Obermaier, Hoelscher-Obermaier, Jason, Julia Persson +7 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
- Near to Mid-term Risks and Opportunities of Open-Source Generative AI
2024/04/25 by Francisco Eiras, Eiras, Francisco, Aleksandar Petrov +45 · 2 voices · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024/10/11 by Tong Fu, Fu, Tingchen, Mrinank Sharma +9 · 7 citations
Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Bengio, Yoshua, Tegan Maharaj +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Understanding Addition in Transformers
2023/10/19 by Philip Quirke, Quirke, Philip, Fazl +1 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
2025/02/19 by Shaona Ghosh, Heather Frase, Ghosh, Shaona +200 · 1 voice · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
2023/11/07 by Michael Lan, Lan, Michael, Torr, Philip +1 · 3 citations
Computer Science · #Text Readability and Simplification #Natural Language Processing Techniques #Topic Modeling
- PMIC: Improving Multi-Agent Reinforcement Learning with Progressive\n Mutual Information Collaboration
2022/03/16 by Pengyi Li, Li, Pengyi, Tang, Hongyao +16 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · Psychology · #Action Observation and Synchronization #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Insect and Arachnid Ecology and Behavior #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics #Robot Manipulation and Learning
- Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
2025/05/30 by Narmeen Oozeer, Oozeer, Narmeen, Luke Marks +4 · 5 citations
Computer Science · #Natural Language Processing Techniques #Multi-Agent Systems and Negotiation #Topic Modeling
- Large Language Models Relearn Removed Concepts
2024/01/03 by Michelle Lo, Shay B. Cohen, Lo, Michelle +3 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
2024/10/09 by Michael Lan, Philip Torr, Lan, Michael +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models
2023/04/22 by Foote, Alex, Nanda, Neel, Kran, Esben +2 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Risks and Opportunities of Open-Source Generative AI
2024/05/14 by Francisco Eiras, Aleksander Petrov, Eiras, Francisco +47 · 1 citation
Computer Science · Decision Sciences · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Scientific Computing and Data Management
- Interpreting Learned Feedback Patterns in Large Language Models
2023/10/12 by Luke Marks, Marks, Luke, Amir Abdullah +9 · 1 citation
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Embodied AI: Emerging Risks and Opportunities for Policy Action
2025/08/28 by Jared Perlo, Alexander Robey, Perlo, Jared +7 · 1 voice · 2 citations
Computer Science · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Human-Automation Interaction and Safety #cs.AI #cs.CY #cs.RO
- In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
2025/04/17 by Ben Bucknall, Bucknall, Ben, Saad Siddiqui +41 · 2 voices · 1 citation
#cs.CY
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
2025/09/28 by Schrodi, Simon, Kempf, Elias, Barez, Fazl +1 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Do Sparse Autoencoders Generalize? A Case Study of Answerability
2025/02/27 by Heindrich, Lovis, Torr, Philip, Barez, Fazl +1 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
2025/05/20 by M. P. Chaudhary, Fazl Barez, Chaudhary, Maheep +1 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
2025/05/28 by Philip Quirke, Narmeen Oozeer, Quirke, Philip +18 · 2 citations
Computer Science · #Topic Modeling
- Scaling sparse feature circuit finding for in-context learning
2025/04/18 by Kharlapenko, Dmitrii, Shabalin, Stepan, Barez, Fazl +2 · 2 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Rethinking AI Cultural Alignment
2025/01/13 by Michal Bravansky, Filip Trhlík, Bravansky, Michal +3 · 1 citation
Social Sciences · #Artificial Intelligence (cs.AI) #Computational and Text Analysis Methods #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Qualitative Comparative Analysis Research
- Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
2025/03/03 by Fu, Tingchen, Barez, Fazl · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Precise In-Parameter Concept Erasure in Large Language Models
2025/05/28 by Yoav Gur-Arieh, Gur-Arieh, Yoav, Clara Haya Suslik +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
2025/09/30 by Oldfield, James, Torr, Philip, Patras, Ioannis +2 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
2024/12/03 by Tony T. Wang, Wang, Tony T., John Hughes +17 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CR #cs.LG