Tramèr, Florian
- Scalable Extraction of Training Data from (Production) Language Models
2023/11/28 by Milad Nasr, Nasr, Milad, Nicholas Carlini +17 · 21 voices · 68 citations
Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
- On the Opportunities and Risks of Foundation Models
2021/08/16 by Rishi Bommasani, Bommasani, Rishi, Drew A. Hudson +233 · 11 voices · 537 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Topic Modeling
- Defeating Prompt Injections by Design
2025/03/24 by Edoardo Debenedetti, Debenedetti, Edoardo, Ilia Shumailov +17 · 23 voices · 59 citations
#cs.CR #cs.AI
- Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
2024/06/17 by Robert Hönig, Hönig, Robert, Javier Rando +5 · 38 voices · 7 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.CR
- Extracting Training Data from Diffusion Models
2023/01/30 by Nicholas Carlini, Carlini, Nicholas, Jamie Hayes +15 · 9 voices · 105 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis
- Stealing Machine Learning Models via Prediction APIs
2016/09/09 by Florian Tramèr, Tramèr, Florian, Fan Zhang +7 · 3 voices · 49 citations
#cs.CR #cs.LG #stat.ML
- Poisoning Web-Scale Training Datasets is Practical
2023/02/20 by Nicholas Carlini, Carlini, Nicholas, Matthew Jagielski +16 · 6 voices · 46 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI #Topic Modeling #cs.CR #cs.LG
- Stealing Part of a Production Language Model
2024/03/11 by Nicholas Carlini, Daniel Paleka, Carlini, Nicholas +26 · 11 voices · 27 citations
Computer Science · #Natural Language Processing Techniques #cs.CR
- Design Patterns for Securing LLM Agents against Prompt Injections
2025/06/10 by Luca Beurer-Kellner, Beat Buesser, Beurer-Kellner, Luca +25 · 10 voices · 19 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #Topic Modeling
- Advances and Open Problems in Federated Learning
2019/12/10 by Kairouz, Peter, McMahan, H. Brendan, Avent, Brendan +56 · 223 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Persistent Pre-Training Poisoning of LLMs
2024/10/17 by Yiming Zhang, Zhang, Yiming, Javier Rando +13 · 3 voices · 21 citations
#cs.CR #cs.AI
- Ensemble Adversarial Training: Attacks and Defenses
2017/05/19 by Tramèr, Florian, Kurakin, Alexey, Papernot, Nicolas +3 · 38 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Large Language Models Can Be Strong Differentially Private Learners
2021/10/12 by Xuechen Li, Florian Tramèr, Li, Xuechen +5 · 43 citations
Computer Science · #Privacy-Preserving Technologies in Data #Stochastic Gradient Optimization Techniques #Domain Adaptation and Few-Shot Learning
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
2025/10/10 by Milad Nasr, Nasr, Milad, Nicholas Carlini +27 · 4 voices · 17 citations
Computer Science · #Adversarial Robustness in Machine Learning #Network Security and Intrusion Detection #Security and Verification in Computing #cs.CR #cs.LG
- Red-Teaming the Stable Diffusion Safety Filter
2022/10/03 by Javier Rando, Rando, Javier, Daniel Paleka +6 · 28 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection #Adversarial Robustness in Machine Learning
- What Does it Mean for a Language Model to Preserve Privacy?
2022/02/11 by Hannah Brown, Katherine Lee, Brown, Hannah +7 · 22 citations
Computer Science · #Privacy-Preserving Technologies in Data
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
2018/06/08 by Florian Tramèr, Tramèr, Florian, Dan Boneh +1 · 16 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Physical Unclonable Functions (PUFs) and Hardware Security #Security and Verification in Computing
- Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy
2022/10/31 by Daphne Ippolito, Florian Tramèr, Ippolito, Daphne +13 · 23 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Differentially Private Learning Needs Better Features (or Much More Data)
2020/11/23 by Tramèr, Florian, Boneh, Dan · 16 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- The Space of Transferable Adversarial Examples
2017/04/11 by Tramèr, Florian, Papernot, Nicolas, Goodfellow, Ian +2 · 10 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Adversarial Training and Robustness for Multiple Perturbations
2019/04/30 by Florian Tramèr, Dan Boneh, Tramèr, Florian +1 · 12 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research
- SentiNet: Detecting Localized Universal Attacks Against Deep Learning\n Systems
2018/12/01 by Chou Edward, Chou, Edward, Tramer Florian +3 · 10 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- An Adversarial Perspective on Machine Unlearning for AI Safety
2024/09/26 by Jakub Łucki, Boyi Wei, Łucki, Jakub +9 · 3 voices · 19 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.AI #cs.CL #cs.CR #cs.LG
- Measuring Forgetting of Memorized Training Examples
2022/06/30 by Matthew Jagielski, Jagielski, Matthew, Om Thakkar +19 · 13 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Topic Modeling
- Counterfactual Memorization in Neural Language Models
2021/12/24 by Chiyuan Zhang, Zhang, Chiyuan, Daphne Ippolito +9 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Backdoor Attacks for In-Context Learning with Language Models
2023/07/27 by Kandpal, Nikhil, Jagielski, Matthew, Tramèr, Florian +1 · 15 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Universal Jailbreak Backdoors from Poisoned Human Feedback
2023/11/24 by Javier Rando, Florian Tramèr, Rando, Javier +1 · 16 citations
Computer Science · #Adversarial Robustness in Machine Learning #Hate Speech and Cyberbullying Detection
- AgentDojo-PROV: A W3C PROV-O Corpus of LLM Agent Executions
2024/06/19 by Debenedetti, Edoardo, Jie Zhang, Mislav Balunović +8 · 19 citations
Computer Science · #Advanced Malware Detection Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Network Security and Intrusion Detection
- Adversarial Search Engine Optimization for Large Language Models
2024/06/26 by Fredrik Nestaas, Nestaas, Fredrik, Edoardo Debenedetti +3 · 3 voices · 11 citations
#cs.CR #cs.LG
- Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
2022/12/13 by Florian Tramèr, Gautam Kamath, Tramèr, Florian +3 · 8 citations
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
- Evaluating Superhuman Models with Consistency Checks
2023/06/16 by Lukas Fluri, Fluri, Lukas, Daniel Paleka +3 · 2 voices · 2 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #Statistical and Computational Modeling #cs.AI #cs.CR #cs.LG #stat.ML
- Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data
2024/09/29 by Jie Zhang, Debeshee Das, Zhang, Jie +5 · 1 voice · 11 citations
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CR #cs.LG
- Blind Baselines Beat Membership Inference Attacks for Foundation Models
2024/06/23 by Das, Debeshee, Zhang, Jie, Tramèr, Florian · 11 citations
#Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- International AI Safety Report
2025/01/29 by Yoshua Bengio, Bengio, Yoshua, Sören Mindermann +180 · 17 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Query-Based Adversarial Prompt Generation
2024/02/19 by Hayase, Jonathan, Borevkovic, Ema, Carlini, Nicholas +2 · 6 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
2021/07/24 by Tramèr, Florian · 3 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- FairTest: Discovering Unwarranted Associations in Data-Driven Applications
2015/10/08 by Tramèr, Florian, Atlidakis, Vaggelis, Geambasu, Roxana +5 · 2 citations
#Computers and Society (cs.CY) #FOS: Computer and information sciences
- Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
2024/06/12 by Debenedetti, Edoardo, Rando, Javier, Paleka, Daniel +18 · 5 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets
2022/03/31 by Tramèr, Florian, Shokri, Reza, Joaquin, Ayrton San +4 · 3 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Privacy Side Channels in Machine Learning Systems
2023/09/11 by Edoardo Debenedetti, Giorgio Severi, Debenedetti, Edoardo +13 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
2025/05/18 by Zhang, Jie, Petrui, Cezara, Nikolić, Kristina +1 · 11 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
2024/04/22 by Javier Rando, Rando, Javier, Francesco Croce +11 · 4 citations
Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
- Antipodes of Label Differential Privacy: PATE and ALIBI
2021/06/07 by Malek, Mani, Mironov, Ilya, Prasad, Karthik +2 · 2 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- SNAP: Efficient Extraction of Private Properties with Poisoning
2022/08/25 by Harsh Chaudhari, Chaudhari, Harsh, John Abascal +9 · 2 citations
Computer Science · Medicine · #Privacy-Preserving Technologies in Data #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education
- Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
2024/03/30 by Shanglun Feng, Feng, Shanglun, Florian Tramèr +1 · 3 citations
Computer Science · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
2020/02/11 by Florian Tramèr, Tramèr, Florian, Jens Behrmann +7 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning
- The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
2025/04/14 by Kristina Nikolić, Nikolić, Kristina, Jie Zhang +4 · 7 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Exploiting Excessive Invariance caused by Norm-Bounded Adversarial\n Robustness
2019/03/25 by Jörn-Henrik Jacobsen, Jacobsen, Jörn-Henrik, Jens Behrmann +8 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #Bacillus and Francisella bacterial research #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
2025/02/04 by Rando, Javier, Zhang, Jie, Carlini, Nicholas +1 · 4 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Consistency Checks for Language Model Forecasters
2024/12/24 by Daniel Paleka, Paleka, Daniel, Abhimanyu Pallavi Sudhir +9 · 3 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
2024/11/15 by Michael Aerni, Javier Rando, Aerni, Michael +9 · 2 voices · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
- Data Poisoning Won't Save You From Facial Recognition
2021/06/28 by Evani Radiya-Dixit, Radiya-Dixit, Evani, Hong, Sanghyun +2 · 1 citation
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG)
- International Scientific Report on the Safety of Advanced AI (Interim Report)
2024/11/05 by Bengio, Yoshua, Mindermann, Sören, Privitera, Daniel +41 · 2 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
- Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
2024/11/22 by Jie Zhang, Christian Schlarmann, Zhang, Jie +11 · 2 voices · 1 citation
#cs.LG #cs.CR
- Gradient-based Jailbreak Images for Multimodal Fusion Models
2024/10/04 by Javier Rando, Hannah Korevaar, Rando, Javier +7 · 2 citations
Engineering · Medicine · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Medical Imaging Techniques and Applications #Medical Imaging and Analysis
- AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
2025/03/03 by Nicholas Carlini, Javier Rando, Carlini, Nicholas +7 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Security and Verification in Computing
- Evading Black-box Classifiers Without Breaking Eggs
2023/06/05 by Debenedetti, Edoardo, Carlini, Nicholas, Tramèr, Florian · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Pitfalls in Evaluating Language Model Forecasters
2025/05/31 by Daniel Paleka, Paleka, Daniel, Shashwat Goel +5 · 4 citations
Computer Science · #Natural Language Processing Techniques
- Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
2025/09/17 by Apertus, Project, Hernández-Cano, Alejandro, Hägele, Alexander +100 · 9 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)