vix.ing · top · new · best · stats · spec

Nicholas Carlini

  1. Large-scale online deanonymization with LLMs
    LLMs enable scalable, high-precision deanonymization of pseudonymous online accounts using unstructured text.
    2026/02/18 by Simon Lermen, Daniel Paleka, Joshua Swanson +3 · 85 voices · 1 citation
    #cs.CR #cs.AI #cs.LG
  2. Scalable Extraction of Training Data from (Production) Language Models
    2023/11/28 by Milad Nasr, Nicholas Carlini, Nasr, Milad +17 · 21 voices · 66 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
  3. Defeating Prompt Injections by Design
    2025/03/24 by Edoardo Debenedetti, Ilia Shumailov, Debenedetti, Edoardo +17 · 23 voices · 58 citations
    #cs.CR #cs.AI
  4. Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
    2024/06/17 by Robert Hönig, Hönig, Robert, Javier Rando +5 · 38 voices · 7 citations
    Computer Science · #Adversarial Robustness in Machine Learning #cs.CR
  5. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
    2018/02/22 by Nicholas Carlini, Carlini, Nicholas, Chang Liu +7 · 5 voices · 68 citations
    #cs.LG #cs.AI #cs.CR
  6. Audio Adversarial Examples: Targeted Attacks on Speech-to-Text
    2018/01/05 by Nicholas Carlini, Carlini, Nicholas, David Wagner +1 · 5 voices · 24 citations
    #cs.LG #cs.AI #cs.CR
  7. Extracting Training Data from Diffusion Models
    2023/01/30 by Nicholas Carlini, Carlini, Nicholas, Jamie Hayes +15 · 9 voices · 100 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis
  8. Poisoning Web-Scale Training Datasets is Practical
    2023/02/20 by Nicholas Carlini, Matthew Jagielski, Carlini, Nicholas +16 · 6 voices · 46 citations
    Computer Science · Medicine · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI #Topic Modeling #cs.CR #cs.LG
  9. Extracting Training Data from Large Language Models
    2020/12/14 by Nicholas Carlini, Florian Tramèr, Carlini, Nicholas +25 · 8 voices · 235 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data #Topic Modeling #cs.CL #cs.CR #cs.LG
  10. Stealing Part of a Production Language Model
    2024/03/11 by Nicholas Carlini, Carlini, Nicholas, Daniel Paleka +26 · 11 voices · 26 citations
    Computer Science · #Natural Language Processing Techniques #cs.CR
  11. Universal and Transferable Adversarial Attacks on Aligned Language Models
    2023/07/27 by Andy Zou, Zou, Andy, Zifan Wang +9 · 5 voices · 456 citations
    #cs.CL #cs.AI #cs.CR #cs.LG
  12. Unsolved Problems in ML Safety
    2021/09/28 by Dan Hendrycks, Hendrycks, Dan, Nicholas Carlini +5 · 2 voices · 27 citations
    #cs.LG #cs.AI #cs.CL #cs.CV
  13. Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
    2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 18 citations
    Medicine · #Medical Imaging and Pathology Studies
  14. MixMatch: A Holistic Approach to Semi-Supervised Learning
    2019/05/06 by David Berthelot, Berthelot, David, Nicholas Carlini +9 · 1 voice · 75 citations
    Computer Science · #cs.LG #cs.AI #cs.CV #stat.ML
  15. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
    2020/01/21 by Kihyuk Sohn, Sohn, Kihyuk, David Berthelot +15 · 155 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Advanced Neural Network Applications
  16. Persistent Pre-Training Poisoning of LLMs
    2024/10/17 by Yiming Zhang, Javier Rando, Zhang, Yiming +13 · 3 voices · 20 citations
    #cs.CR #cs.AI
  17. Membership Inference Attacks From First Principles
    2021/12/07 by Nicholas Carlini, Carlini, Nicholas, Steve Chien +9 · 94 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  18. Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
    2018/02/01 by Anish Athalye, Nicholas Carlini, Athalye, Anish +4 · 61 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Physical Unclonable Functions (PUFs) and Hardware Security
  19. Quantifying Memorization Across Neural Language Models
    2022/02/15 by Nicholas Carlini, Daphne Ippolito, Carlini, Nicholas +9 · 71 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  20. Deduplicating Training Data Makes Language Models Better
    2021/07/14 by Katherine Lee, Daphne Ippolito, Lee, Katherine +11 · 58 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  21. On Evaluating Adversarial Robustness
    2019/02/18 by Nicholas Carlini, Anish Athalye, Carlini, Nicholas +15 · 49 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  22. Label-Only Membership Inference Attacks
    2020/07/28 by Christopher A. Choquette-Choo, Florian Tramèr, Choquette-Choo, Christopher A. +5 · 31 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data #Machine Learning and Algorithms
  23. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
    2025/10/10 by Milad Nasr, Nasr, Milad, Nicholas Carlini +27 · 4 voices · 17 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Network Security and Intrusion Detection #Security and Verification in Computing #cs.CR #cs.LG
  24. Adversary Instantiation: Lower Bounds for Differentially Private Machine\n Learning
    2021/01/11 by Milad Nasr, Shuang Song, Nasr, Milad +7 · 16 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
  25. The Privacy Onion Effect: Memorization is Relative
    2022/06/21 by Nicholas Carlini, Matthew Jagielski, Carlini, Nicholas +9 · 16 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
  26. High Accuracy and High Fidelity Extraction of Neural Networks
    2019/09/03 by Matthew Jagielski, Jagielski, Matthew, Nicholas Carlini +7 · 14 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Anomaly Detection Techniques and Applications
  27. Counterfactual Memorization in Neural Language Models
    2021/12/24 by Chiyuan Zhang, Daphne Ippolito, Zhang, Chiyuan +9 · 12 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  28. Measuring Forgetting of Memorized Training Examples
    2022/06/30 by Matthew Jagielski, Jagielski, Matthew, Om Thakkar +19 · 12 citations
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Topic Modeling
  29. ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
    2019/11/21 by David Berthelot, Berthelot, David, Nicholas Carlini +11 · 9 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Multimodal Machine Learning Applications
  30. Provably Minimally-Distorted Adversarial Examples
    2017/09/29 by Nicholas Carlini, Guy Katz, Carlini, Nicholas +5 · 13 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  31. Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
    2022/12/13 by Florian Tramèr, Tramèr, Florian, Gautam Kamath +3 · 8 citations
    Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
  32. ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
    2026/05/11 by Zhun Wang, Nico Schiller, Hongwei Li +13 · 8 voices · 1 citation
    #cs.CR #cs.AI #cs.LG
  33. Forcing Diffuse Distributions out of Language Models
    2024/04/16 by Yiming Zhang, Zhang, Yiming, Avi Schwarzschild +7 · 1 voice · 9 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
  34. Handcrafted Backdoors in Deep Neural Networks
    2021/06/08 by Sanghyun Hong, Nicholas Carlini, Hong, Sanghyun +3 · 5 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection
  35. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
    2026/01/17 by Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82 · 3 voices · 4 citations
    Computer Science · #cs.SE #cs.AI
  36. SoK: Watermarking for AI-Generated Content
    2024/11/27 by Xuandong Zhao, Sam Gunn, Zhao, Xuandong +25 · 13 citations
    Computer Science · #Advanced Steganography and Watermarking Techniques #Artificial Intelligence (cs.AI) #Chaos-based Image/Signal Encryption #Computer Graphics and Visualization Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  37. (Certified!!) Adversarial Robustness for Free!
    2022/06/21 by Nicholas Carlini, Carlini, Nicholas, Florian Tramèr +6 · 4 citations
    Computer Science · Medicine · Engineering · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI #Advanced X-ray and CT Imaging
  38. On Evaluating the Durability of Safeguards for Open-Weight LLMs
    2024/12/10 by Xiangyu Qi, Qi, Xiangyu, Boyi Wei +17 · 2 voices · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #cs.AI #cs.CR
  39. Debugging Differential Privacy: A Case Study for Privacy Auditing
    2022/02/24 by Florian Tramèr, Andreas Terzis, Tramer, Florian +9 · 3 citations
    Computer Science · #Privacy-Preserving Technologies in Data #Cryptography and Data Security #Stochastic Gradient Optimization Techniques
  40. Privacy Side Channels in Machine Learning Systems
    2023/09/11 by Edoardo Debenedetti, Giorgio Severi, Debenedetti, Edoardo +13 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
  41. CryptanalysisBench: Can LLMs do Cryptanalysis?
    2026/07/20 by Lukas Fluri, Avital Shafran, Nicholas Carlini +5 · 3 voices
    Computer Science · #cs.CR
  42. Students Parrot Their Teachers: Membership Inference on Model Distillation
    2023/03/06 by Matthew Jagielski, Jagielski, Matthew, Milad Nasr +7 · 1 voice · 2 citations
    #cs.CR #cs.LG
  43. Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label Setting
    2024/10/08 by Nicholas Carlini, Jorge Chávez-Saab, Carlini, Nicholas +7 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Neural Networks and Applications #Statistical and Computational Modeling
  44. Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
    2020/02/11 by Florian Tramèr, Tramèr, Florian, Jens Behrmann +7 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  45. Exploiting Excessive Invariance caused by Norm-Bounded Adversarial\n Robustness
    2019/03/25 by Jörn-Henrik Jacobsen, Jens Behrmann, Jacobsen, Jörn-Henrik +8 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  46. Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
    2024/11/15 by Michael Aerni, Javier Rando, Aerni, Michael +9 · 2 voices · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
  47. Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial Examples
    2021/06/18 by Maura Pintor, Pintor, Maura, Luca Demetrio +11 · 1 citation
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Testing and Debugging Techniques
  48. Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
    2024/11/22 by Jie Zhang, Zhang, Jie, Christian Schlarmann +11 · 2 voices · 1 citation
    #cs.LG #cs.CR
  49. AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
    2025/03/03 by Nicholas Carlini, Carlini, Nicholas, Javier Rando +7 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Security and Verification in Computing
  50. A LLM Assisted Exploitation of AI-Guardian
    2023/07/20 by Nicholas Carlini, Carlini, Nicholas · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  51. Stealing User Prompts from Mixture of Experts
    2024/10/30 by Itay Yona, Yona, Itay, Ilia Shumailov +5 · 1 citation
    Computer Science · #Data Mining Algorithms and Applications