Stealing Machine Learning Models via Prediction APIs
2016/09/09 by Florian Tramèr, Fan Zhang, Tramèr, Florian +7 · 3 voices · 733 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Adversary #Analytics #Artificial intelligence #Artificial neural network #Computer science #Computer security #Confidentiality #Data mining #Decision tree #Explainable Artificial Intelligence (XAI) #Fidelity #Lasso (programming language) #Machine learning #Network Security and Intrusion Detection #Support vector machine #Threat model #World Wide Web #cs.CR #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1609.02943
published in arXiv (Cornell University), 601-618 (Cornell University) · 19 pages, 7 figures, Proceedings of USENIX Security 2016
openalex publication_date 2016/09/09 · arxiv created 2016/10/03 · arxiv updated 2016/10/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Machine learning (ML) models may be deemed confidential due to their sensitive training data, commercial value, or use in security applications. Increasingly often, confidential ML models are being deployed with publicly accessible query interfaces. ML-as-a-service ("predictive analytics") systems are an example: Some allow users to train models on potentially sensitive data and charge others for access on a pay-per-query basis. The tension between model confidentiality and public access motivates our investigation of model extraction attacks. In such attacks, an adversary with black-box access, but no prior knowledge of an ML model's parameters or training data, aims to duplicate the functionality of (i.e., "steal") the model. Unlike in classical learning theory settings, ML-as-a-service offerings may accept partial feature vectors as inputs and include confidence values with predictions. Given these practices, we show simple, efficient attacks that extract target ML models with near-perfect fidelity for popular model classes including logistic regression, neural networks, and decision trees. We demonstrate these attacks against the online services of BigML and Amazon Machine Learning. We further show that the natural countermeasure of omitting confidence values from model outputs still admits potentially harmful model extraction attacks. Our results highlight the need for careful ML model deployment and new model extraction countermeasures.
Cited by
- CrypTorch: PyTorch-based Auto-tuning Compiler for Machine Learning with Multi-party Computation
- NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs
- Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
- Signal-based Model Access Risk Analysis for AI System Operations Security
- ADS-C: Antidistillation Sampling for Classification
- Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks
- AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
- From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
- Provably Extracting the Features from a General Superposition
- IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol
- Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- Sigil: Server-Enforced Watermarking in U-Shaped Split Federated Learning via Gradient Injection
- RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- Protecting the Neural Networks against FGSM Attack Using Machine Unlearning
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
- On Stealing Graph Neural Network Models
- Black-Box Guardrail Reverse-engineering Attack
- Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations
- Adversarially-Aware Architecture Design for Robust Medical AI Systems
- SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
- SynthID-Image: Image watermarking at internet scale
- Position: Privacy Is Not Just Memorization!
- Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?
- Stealing AI Model Weights Through Covert Communication Channels
- Fingerprinting LLMs via Prompt Injection
- International Intellectual Property Law in the Age of AI
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- Responsible Diffusion: A Comprehensive Survey on Safety, Ethics, and Trust in Diffusion Models
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- MER-Inspector: Assessing model extraction risks from an attack-agnostic perspective
- Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
- Delving into Cryptanalytic Extraction of PReLU Neural Networks
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
- Stabilizing Data-Free Model Extraction
- "Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
- I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks
- Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses
- Surveying the Operational Cybersecurity and Supply Chain Threat Landscape when Developing and Deploying AI Systems
- The AI-Fraud Diamond: A Novel Lens for Auditing Algorithmic Deception
- Decentralized Weather Forecasting via Distributed Machine Learning and Blockchain-Based Model Validation
- AI Security Map: Holistic Organization of AI Security Technologies and Impacts on Stakeholders
- FuSeFL: Fully Secure and Scalable Federated Learning
- Slice or the Whole Pie? Utility Control for AI Models
- GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI
- Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
- Characterizing Linear Alignment Across Language Models
- Striking the Perfect Balance: Preserving Privacy While Boosting Utility in Collaborative Medical Prediction Platforms
- Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
- BarkBeetle: Stealing Decision Tree Models with Fault Injection
- Symbiosis: Multi-Adapter Inference and Fine-Tuning
- Counterfactual Explanation of Shapley Value in Data Coalitions
- Detect & Score: Privacy-Preserving Misbehaviour Detection and Contribution Evaluation in Federated Learning
- GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
- Securing AI Systems: A Guide to Known Attacks and Impacts
- SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
- A Survey on Model Extraction Attacks and Defenses for Large Language Models
- Holmes: Towards Effective and Harmless Model Ownership Verification to Personalized Large Vision Models via Decoupling Common Features
- CEGA: A Cost-Effective Approach for Graph-Based Model Extraction and Acquisition
- AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
- Navigating the Deep: Signature Extraction on Deep Neural Networks
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
- Reassessing Code Authorship Attribution in the Era of Language Models
- Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
- No TPU Left Behind: Retrofitting Side-Channel Protection into Edge TPUs
- Stealix: Model Stealing via Prompt Evolution
- BESA: Boosting Encoder Stealing Attack with Perturbation Recovery
- MISLEADER: Defending against Model Extraction with Ensembles of Distilled Models
- Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing of Graph Models
- CHIP: Chameleon Hash-based Irreversible Passport for Robust Deep Model Ownership Verification and Active Usage Control
- Watermarking Without Standards Is Not AI Governance
- Preventing Adversarial AI Attacks Against Autonomous Situational Awareness: A Maritime Case Study
- DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
- Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated Learning
- IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
- On the Security Risks of ML-based Malware Detection Systems: A Survey
- On the interplay of Explainability, Privacy and Predictive Performance with Explanation-assisted Model Extraction
- Opening the Scope of Openness in AI
- ChainMarks: Securing DNN Watermark with Cryptographic Chain
- LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
- Exploiting PendingIntent Provenance Confusion to Spoof Android SDK Authentication
- PDF: PUF-based DNN Fingerprinting for Knowledge Distillation Traceability
- Enabling Adversarial Robustness in AI Models through Kubeflow MLOps
- Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service
- Robust Privacy: Inference-Stage Privacy through Certified Robustness
- A Prompt-Based Framework for Loop Vulnerability Detection Using Local LLMs
- AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems
- Provably Learning Multi-Head Attention with Queries
- When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
- Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills
- Copycat CNN: Are random non-Labeled data enough to steal knowledge from black-box models?
- Your Image Generator Is Your New Private Dataset
- Towards a robust and trustworthy machine learning system development: An engineering perspective
- Confidential Machine Learning Computation in Untrusted Environments: A Systems Security Perspective
- Algebraic Cryptanalytic Extraction on Hard-Label Neural Networks
Discussions
Related