Fazl Barez
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Best-of-N Jailbreaking
2024/12/04 by John D. Hughes, John Hughes, Hughes, John +18 · 17 voices · 14 citations
Computer Science · #Digital and Cyber Forensics #cs.AI #cs.CL #cs.LG
- The Larger They Are, the Harder They Fail: Language Models do not Recognize Identifier Swaps in Python
2023/05/24 by Antonio Valerio Miceli-Barone, Miceli-Barone, Antonio Valerio, Fazl Barez +5 · 3 voices · 4 citations
Computer Science · #Software Engineering Research #Topic Modeling #Computational Physics and Python Applications
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Barez, Fazl, Tingchen Fu +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Towards Interpreting Visual Information Processing in Vision-Language Models
2024/10/09 by Clement Neo, Neo, Clement, C.-H. Luke Ong +9 · 22 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
2025/02/18 by Adi Simhi, Simhi, Adi, Itay Itzhak +7 · 1 voice · 13 citations
Computer Science · Neuroscience · #Blockchain Technology Applications and Security #Hallucinations in medical conditions #cs.CL
- Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
2024/11/02 by Luke Marks, Alasdair Paren, Marks, Luke +5 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Neural Networks and Applications #Explainable Artificial Intelligence (XAI)
- Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
2023/05/27 by Jason Hoelscher-Obermaier, Julia Persson, Hoelscher-Obermaier, Jason +7 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
- Near to Mid-term Risks and Opportunities of Open-Source Generative AI
2024/04/25 by Francisco Eiras, Aleksandar Petrov, Eiras, Francisco +45 · 2 voices · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024/10/11 by Tong Fu, Fu, Tingchen, Mrinank Sharma +9 · 7 citations
Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Bengio, Yoshua, Tegan Maharaj +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
2025/02/19 by Shaona Ghosh, Ghosh, Shaona, Heather Frase +200 · 1 voice · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Large Language Models Relearn Removed Concepts
2024/01/03 by Michelle Lo, Shay B. Cohen, Lo, Michelle +3 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
2024/10/09 by Michael Lan, Lan, Michael, Philip Torr +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Risks and Opportunities of Open-Source Generative AI
2024/05/14 by Francisco Eiras, Aleksander Petrov, Eiras, Francisco +47 · 1 citation
Computer Science · Decision Sciences · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Scientific Computing and Data Management
- Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
2026/02/08 by Adi Simhi, Fazl Barez, Martin Tutek +2 · 4 voices · 1 citation
#cs.CL #cs.AI
- Interpreting Learned Feedback Patterns in Large Language Models
2023/10/12 by Luke Marks, Amir Abdullah, Marks, Luke +9 · 1 citation
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Embodied AI: Emerging Risks and Opportunities for Policy Action
2025/08/28 by Jared Perlo, Alexander Robey, Perlo, Jared +7 · 1 voice · 2 citations
Computer Science · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Human-Automation Interaction and Safety #cs.AI #cs.CY #cs.RO
- In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
2025/04/17 by Ben Bucknall, Bucknall, Ben, Saad Siddiqui +41 · 2 voices · 1 citation
#cs.CY
- Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
2025/05/20 by M. P. Chaudhary, Chaudhary, Maheep, Fazl Barez +1 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
2025/05/28 by Philip Quirke, Quirke, Philip, Narmeen Oozeer +18 · 2 citations
Computer Science · #Topic Modeling
- Rethinking AI Cultural Alignment
2025/01/13 by Michal Bravansky, Filip Trhlík, Bravansky, Michal +3 · 1 citation
Social Sciences · #Artificial Intelligence (cs.AI) #Computational and Text Analysis Methods #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Qualitative Comparative Analysis Research
- Precise In-Parameter Concept Erasure in Large Language Models
2025/05/28 by Yoav Gur-Arieh, Clara Haya Suslik, Gur-Arieh, Yoav +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
- Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
2024/12/03 by Tony T. Wang, Wang, Tony T., John Hughes +17 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CR #cs.LG