Sleight, Henry
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
2024/04/01 by Matthias Gerstgrasser, Gerstgrasser, Matthias, Rylan Schaeffer +26 · 16 voices · 32 citations
Computer Science · #Semantic Web and Ontologies #cs.AI #cs.CL #cs.ET #cs.LG #stat.ML
- Best-of-N Jailbreaking
2024/12/04 by John D. Hughes, John Hughes, Hughes, John +18 · 17 voices · 17 citations
Computer Science · #Digital and Cyber Forensics #cs.AI #cs.CL #cs.LG
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
2025/07/29 by Runjin Chen, Chen, Runjin, Andy Arditi +7 · 15 voices · 88 citations
#cs.CL #cs.LG
- Unsupervised Elicitation of Language Models
2025/06/11 by Jiaxin Wen, Zachary Ankner, Wen, Jiaxin +23 · 15 voices · 4 citations
Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Machine Learning and Data Classification
- Looking Inward: Language Models Can Learn About Themselves by Introspection
2024/10/17 by Felix J Binder, James Chua, Binder, Felix J +15 · 4 voices · 38 citations
Computer Science · #Natural Language Processing Techniques
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
2024/07/22 by Abhay Sheshadri, Aidan Ewart, Sheshadri, Abhay +19 · 31 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Inverse Scaling in Test-Time Compute
2025/07/19 by Aryo Pradipta Gema, Alexander Hägele, Gema, Aryo Pradipta +26 · 4 voices · 26 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Multimodal Machine Learning Applications #Constraint Satisfaction and Optimization
- Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
2024/07/21 by Rylan Schaeffer, Schaeffer, Rylan, Dan Valentine +27 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
2024/11/26 by Jiaxin Wen, Wen, Jiaxin, Vivek Hebbar +20 · 6 citations
Computer Science · #Blockchain Technology Applications and Security
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
2025/06/17 by Kutasov, Jonathan, Sun, Yuqi, Colognese, Paul +9 · 12 citations
Computer Science · Medicine · #Multimodal Machine Learning Applications #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI)
- Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
2024/11/12 by Peng, Alwin, Michael, Julian, Sleight, Henry +2 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
2025/12/05 by Igor Shilov, Alex Cloud, Shilov, Igor +13 · 2 voices · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #cs.LG
- Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
2024/12/03 by Tony T. Wang, John Hughes, Wang, Tony T. +17 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CR #cs.LG
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
2025/10/06 by Nevan Wichers, Wichers, Nevan, Aram Ebtekar +19 · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Software Engineering Research #Topic Modeling