Carlini, Nicholas
- Scalable Extraction of Training Data from (Production) Language Models
2023/11/28 by Milad Nasr, Nicholas Carlini, Nasr, Milad +17 · 21 voices · 79 citations
Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
- Defeating Prompt Injections by Design
2025/03/24 by Edoardo Debenedetti, Debenedetti, Edoardo, Ilia Shumailov +17 · 23 voices · 66 citations
#cs.CR #cs.AI
- Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
2024/06/17 by Robert Hönig, Javier Rando, Hönig, Robert +5 · 38 voices · 9 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.CR
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
2018/02/22 by Nicholas Carlini, Carlini, Nicholas, Chang Liu +7 · 5 voices · 79 citations
#cs.LG #cs.AI #cs.CR
- Audio Adversarial Examples: Targeted Attacks on Speech-to-Text
2018/01/05 by Nicholas Carlini, Carlini, Nicholas, David Wagner +1 · 5 voices · 27 citations
#cs.LG #cs.AI #cs.CR
- Extracting Training Data from Diffusion Models
2023/01/30 by Nicholas Carlini, Carlini, Nicholas, Jamie Hayes +15 · 9 voices · 119 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis
- Poisoning Web-Scale Training Datasets is Practical
2023/02/20 by Nicholas Carlini, Matthew Jagielski, Carlini, Nicholas +16 · 6 voices · 53 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI #Topic Modeling #cs.CR #cs.LG
- Extracting Training Data from Large Language Models
2020/12/14 by Nicholas Carlini, Florian Tramer, Carlini, Nicholas +25 · 8 voices · 288 citations
Computer Science · #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data #Topic Modeling #cs.CL #cs.CR #cs.LG
- Stealing Part of a Production Language Model
2024/03/11 by Nicholas Carlini, Daniel Paleka, Carlini, Nicholas +26 · 11 voices · 33 citations
Computer Science · #Natural Language Processing Techniques #cs.CR
- Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
2023/07/27 by Andy Zou, Zou, Andy, Zi Wang +10 · 5 voices · 538 citations
Computer Science · #Adversarial Robustness in Machine Learning #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG
- Unsolved Problems in ML Safety
2021/09/28 by Dan Hendrycks, Hendrycks, Dan, Nicholas Carlini +5 · 2 voices · 32 citations
#cs.LG #cs.AI #cs.CL #cs.CV
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Javier Rando, Souly, Alexandra +25 · 34 voices · 19 citations
Medicine · #Medical Imaging and Pathology Studies
- MixMatch: A Holistic Approach to Semi-Supervised Learning
2019/05/06 by David Berthelot, Nicholas Carlini, Berthelot, David +9 · 1 voice · 109 citations
Computer Science · #cs.LG #cs.AI #cs.CV #stat.ML
- Towards Evaluating the Robustness of Neural Networks
2016/08/16 by Carlini, Nicholas, Wagner, David · 221 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
2020/01/21 by Kihyuk Sohn, David Berthelot, Sohn, Kihyuk +15 · 210 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Advanced Neural Network Applications
- Persistent Pre-Training Poisoning of LLMs
2024/10/17 by Yiming Zhang, Javier Rando, Zhang, Yiming +13 · 3 voices · 25 citations
#cs.CR #cs.AI
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
2018/02/01 by Anish Athalye, Nicholas Carlini, Athalye, Anish +4 · 120 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Physical Unclonable Functions (PUFs) and Hardware Security
- Membership Inference Attacks From First Principles
2021/12/07 by Nicholas Carlini, Steve Chien, Carlini, Nicholas +9 · 123 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Quantifying Memorization Across Neural Language Models
2022/02/15 by Nicholas Carlini, Daphne Ippolito, Carlini, Nicholas +9 · 90 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Deduplicating Training Data Makes Language Models Better
2021/07/14 by Katherine Lee, Daphne Ippolito, Lee, Katherine +11 · 72 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- On Evaluating Adversarial Robustness
2019/02/18 by Nicholas Carlini, Anish Athalye, Carlini, Nicholas +15 · 83 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Label-Only Membership Inference Attacks
2020/07/28 by Christopher A. Choquette-Choo, Florian Tramèr, Choquette-Choo, Christopher A. +5 · 38 citations
Computer Science · #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data #Machine Learning and Algorithms
- Measuring Robustness to Natural Distribution Shifts in Image Classification
2020/07/01 by Rohan Taori, Achal Dave, Taori, Rohan +9 · 34 citations
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- On Adaptive Attacks to Adversarial Example Defenses
2020/02/19 by Florian Tramèr, Tramer, Florian, Nicholas Carlini +5 · 36 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Security and Verification in Computing
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
2025/10/10 by Milad Nasr, Nicholas Carlini, Nasr, Milad +27 · 4 voices · 18 citations
Computer Science · #Adversarial Robustness in Machine Learning #Network Security and Intrusion Detection #Security and Verification in Computing #cs.CR #cs.LG
- Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy
2022/10/31 by Daphne Ippolito, Florian Tramèr, Ippolito, Daphne +13 · 31 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
2017/05/20 by Carlini, Nicholas, Wagner, David · 19 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Adversary Instantiation: Lower Bounds for Differentially Private Machine\n Learning
2021/01/11 by Milad Nasr, Shuang Song, Nasr, Milad +7 · 23 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- Are aligned neural networks adversarially aligned?
2023/06/26 by Carlini, Nicholas, Nasr, Milad, Choquette-Choo, Christopher A. +8 · 33 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Effective Prompt Extraction from Language Models
2023/07/13 by Yiming Zhang, Zhang, Yiming, Carlini, Nicholas +1 · 29 citations
Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information and Cyber Security #Software Engineering Research
- Counterfactual Memorization in Neural Language Models
2021/12/24 by Chiyuan Zhang, Zhang, Chiyuan, Daphne Ippolito +9 · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- The Privacy Onion Effect: Memorization is Relative
2022/06/21 by Nicholas Carlini, Matthew Jagielski, Carlini, Nicholas +9 · 20 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- High Accuracy and High Fidelity Extraction of Neural Networks
2019/09/03 by Matthew Jagielski, Nicholas Carlini, Jagielski, Matthew +7 · 18 citations
Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Anomaly Detection Techniques and Applications
- Poisoning and Backdooring Contrastive Learning
2021/06/17 by Carlini, Nicholas, Terzis, Andreas · 13 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Measuring Forgetting of Memorized Training Examples
2022/06/30 by Matthew Jagielski, Jagielski, Matthew, Om Thakkar +19 · 15 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Topic Modeling
- ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
2019/11/21 by David Berthelot, Nicholas Carlini, Berthelot, David +11 · 11 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Multimodal Machine Learning Applications
- Provably Minimally-Distorted Adversarial Examples
2017/09/29 by Nicholas Carlini, Carlini, Nicholas, Guy Katz +5 · 19 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
2019/03/22 by Qin, Yao, Carlini, Nicholas, Goodfellow, Ian +2 · 10 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- Backdoor Attacks for In-Context Learning with Language Models
2023/07/27 by Kandpal, Nikhil, Jagielski, Matthew, Tramèr, Florian +1 · 16 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Cryptanalytic Extraction of Neural Network Models
2020/03/10 by Carlini, Nicholas, Jagielski, Matthew, Mironov, Ilya · 9 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Technical Report on the CleverHans v2.1.0 Adversarial Examples Library
2016/10/03 by Papernot, Nicolas, Faghri, Fartash, Carlini, Nicholas +23 · 7 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
2022/12/13 by Florian Tramèr, Gautam Kamath, Tramèr, Florian +3 · 12 citations
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
- AdaMatch: A Unified Approach to Semi-Supervised Learning and Domain Adaptation
2021/06/08 by Berthelot, David, Roelofs, Rebecca, Sohn, Kihyuk +2 · 7 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Handcrafted Backdoors in Deep Neural Networks
2021/06/08 by Sanghyun Hong, Hong, Sanghyun, Nicholas Carlini +3 · 7 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection
- Evading Deepfake-Image Detectors with White- and Black-Box Attacks
2020/04/01 by Carlini, Nicholas, Farid, Hany · 6 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Forcing Diffuse Distributions out of Language Models
2024/04/16 by Yiming Zhang, Zhang, Yiming, Avi Schwarzschild +7 · 1 voice · 10 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
- SoK: Watermarking for AI-Generated Content
2024/11/27 by Xuandong Zhao, Zhao, Xuandong, Sam Gunn +25 · 17 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Artificial Intelligence (cs.AI) #Chaos-based Image/Signal Encryption #Computer Graphics and Visualization Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Defensive Distillation is Not Robust to Adversarial Examples
2016/07/14 by Carlini, Nicholas, Wagner, David · 4 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- On Evaluating the Durability of Safeguards for Open-Weight LLMs
2024/12/10 by Xiangyu Qi, Boyi Wei, Qi, Xiangyu +17 · 2 voices · 12 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Nuclear and radioactivity studies #cs.AI #cs.CR
- Debugging Differential Privacy: A Case Study for Privacy Auditing
2022/02/24 by Florian Tramèr, Andreas Terzis, Tramer, Florian +9 · 5 citations
Computer Science · #Privacy-Preserving Technologies in Data #Cryptography and Data Security #Stochastic Gradient Optimization Techniques
- (Certified!!) Adversarial Robustness for Free!
2022/06/21 by Nicholas Carlini, Florian Tramèr, Carlini, Nicholas +6 · 5 citations
Computer Science · Medicine · Engineering · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI #Advanced X-ray and CT Imaging
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
2021/05/04 by Carlini, Nicholas · 4 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Query-Based Adversarial Prompt Generation
2024/02/19 by Hayase, Jonathan, Borevkovic, Ema, Carlini, Nicholas +2 · 7 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Students Parrot Their Teachers: Membership Inference on Model Distillation
2023/03/06 by Matthew Jagielski, Jagielski, Matthew, Milad Nasr +7 · 1 voice · 3 citations
#cs.CR #cs.LG
- Exploiting Excessive Invariance caused by Norm-Bounded Adversarial\n Robustness
2019/03/25 by Jörn-Henrik Jacobsen, Jacobsen, Jörn-Henrik, Jens Behrmann +8 · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets
2022/03/31 by Tramèr, Florian, Shokri, Reza, Joaquin, Ayrton San +4 · 3 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Privacy Side Channels in Machine Learning Systems
2023/09/11 by Edoardo Debenedetti, Giorgio Severi, Debenedetti, Edoardo +13 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
2017/11/22 by Carlini, Nicholas, Wagner, David · 2 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- No Free Lunch in "Privacy for Free: How does Dataset Condensation Help Privacy"
2022/09/29 by Nicholas Carlini, Vitaly Feldman, Carlini, Nicholas +3 · 3 citations
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- Stateful Detection of Black-Box Adversarial Attacks
2019/07/12 by Chen, Steven, Carlini, Nicholas, Wagner, David · 2 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label Setting
2024/10/08 by Nicholas Carlini, Jorge Chávez-Saab, Carlini, Nicholas +7 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Neural Networks and Applications #Statistical and Computational Modeling
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
2024/04/01 by Yuxin Wen, Wen, Yuxin, Leo Marchyok +9 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
2020/02/11 by Florian Tramèr, Tramèr, Florian, Jens Behrmann +7 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Evading Adversarial Example Detection Defenses with Orthogonal Projected Gradient Descent
2021/06/28 by Oliver Bryniarski, Bryniarski, Oliver, Nabeel Hingun +7 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Bacillus and Francisella bacterial research #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
2024/11/15 by Michael Aerni, Aerni, Michael, Javier Rando +9 · 2 voices · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
- Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
2025/02/04 by Rando, Javier, Zhang, Jie, Carlini, Nicholas +1 · 6 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Report of the 1st Workshop on Generative AI and Law
2023/11/11 by A. Feder Cooper, Cooper, A. Feder, Katherine Lee +67 · 3 citations
Computer Science · Social Sciences · #Artificial Intelligence in Law #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Law, AI, and Intellectual Property
- Initialization Matters for Adversarial Transfer Learning
2023/12/10 by Hua, Andong, Gu, Jindong, Xue, Zhiyu +3 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Remote Timing Attacks on Efficient Language Model Inference
2024/10/22 by Carlini, Nicholas, Nasr, Milad · 4 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
2024/11/22 by Jie Zhang, Zhang, Jie, Christian Schlarmann +11 · 2 voices · 2 citations
#cs.LG #cs.CR
- Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
2025/01/13 by Huang, Yangsibo, Nasr, Milad, Angelopoulos, Anastasios +10 · 4 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
2017/06/15 by He, Warren, Wei, James, Chen, Xinyun +2 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Unrestricted Adversarial Examples
2018/09/22 by Brown, Tom B., Carlini, Nicholas, Zhang, Chiyuan +3 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- A critique of the DeepSec Platform for Security Analysis of Deep Learning Models
2019/05/17 by Carlini, Nicholas · 1 citation
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Distribution Density, Tails, and Outliers in Machine Learning: Metrics and Applications
2019/10/29 by Carlini, Nicholas, Erlingsson, Úlfar, Papernot, Nicolas · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
2025/03/21 by Longpre, Shayne, Klyman, Kevin, Appel, Ruth E. +31 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial Examples
2021/06/18 by Maura Pintor, Pintor, Maura, Luca Demetrio +11 · 2 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Testing and Debugging Techniques
- Defending Against Prompt Injection With a Few DefensiveTokens
2025/07/10 by Chen, Sizhe, Wang, Yizhu, Carlini, Nicholas +2 · 9 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
2025/10/23 by Ziqian Zhong, Aditi Raghunathan, Zhong, Ziqian +3 · 1 voice · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CL #cs.LG
- Is Private Learning Possible with Instance Encoding?
2020/11/10 by Nicholas Carlini, Samuel Deng, Carlini, Nicholas +15 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Privacy-Preserving Technologies in Data
- Data Poisoning Won't Save You From Facial Recognition
2021/06/28 by Evani Radiya-Dixit, Radiya-Dixit, Evani, Hong, Sanghyun +2 · 1 citation
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG)
- AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
2025/03/03 by Nicholas Carlini, Javier Rando, Carlini, Nicholas +7 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Security and Verification in Computing
- Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
2024/05/06 by Nicholas Carlini, Carlini, Nicholas · 1 voice
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CR #cs.LG
- Effective Robustness against Natural Distribution Shifts for Models with Different Training Data
2023/02/02 by Zhouxing Shi, Nicholas Carlini, Shi, Zhouxing +11 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Evading Black-box Classifiers Without Breaking Eggs
2023/06/05 by Debenedetti, Edoardo, Carlini, Nicholas, Tramèr, Florian · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- A LLM Assisted Exploitation of AI-Guardian
2023/07/20 by Nicholas Carlini, Carlini, Nicholas · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Publishing Efficient On-device Models Increases Adversarial Vulnerability
2022/12/28 by Hong, Sanghyun, Carlini, Nicholas, Kurakin, Alexey · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Stealing User Prompts from Mixture of Experts
2024/10/30 by Itay Yona, Yona, Itay, Ilia Shumailov +5 · 1 citation
Computer Science · #Data Mining Algorithms and Applications