Sharma, Mrinank
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
2023/10/20 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +36 · 10 voices · 226 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Hubinger, Evan, Carson Denison +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
- Best-of-N Jailbreaking
2024/12/04 by John Hughes, John D. Hughes, Hughes, John +18 · 17 voices · 14 citations
Computer Science · #Digital and Cyber Forensics #cs.AI #cs.CL #cs.LG
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
2025/01/31 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +89 · 8 voices · 36 citations
Social Sciences · #Criminal Law and Evidence #Law, Rights, and Freedoms #Legal Systems and Judicial Processes
- Do Bayesian Neural Networks Need To Be Fully Stochastic?
2022/11/11 by Mrinank Sharma, Sharma, Mrinank, Sebastian Farquhar +5 · 1 voice · 7 citations
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.AI #cs.LG #stat.ML
- Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt
2022/06/14 by Mindermann, Sören, Brauner, Jan, Razzak, Muhammed +8 · 21 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
2024/07/21 by Rylan Schaeffer, Dan Valentine, Schaeffer, Rylan +27 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG)
- PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
2024/10/11 by Tong Fu, Fu, Tingchen, Mrinank Sharma +9 · 7 citations
Computer Science · Medicine · Pharmacology, Toxicology and Pharmaceutics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Pharmacovigilance and Adverse Drug Reactions #Poisoning and overdose treatments
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
2024/11/26 by Jiaxin Wen, Wen, Jiaxin, Vivek Hebbar +20 · 4 citations
Computer Science · #Blockchain Technology Applications and Security
- Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
2024/11/12 by Peng, Alwin, Michael, Julian, Sleight, Henry +2 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Understanding and Controlling a Maze-Solving Policy Network
2023/10/12 by Mini, Ulisse, Grietzer, Peli, Sharma, Mrinank +3 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- Forecasting Rare Language Model Behaviors
2025/02/24 by Jones, Erik, Meg Tong, Jesse Mu +16 · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification
- Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
2024/12/03 by Tony T. Wang, John Hughes, Wang, Tony T. +17 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CR #cs.LG