Accelerating Large Language Model Decoding with Speculative Sampling
2023/02/02 by Charlie Chen, Chen, Charlie, Sebastian Borgeaud +9 · 154 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2302.01318
Abstract
We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the generation of multiple tokens from each transformer call. Our algorithm relies on the observation that the latency of parallel scoring of short continuations, generated by a faster but less powerful draft model, is comparable to that of sampling a single token from the larger target model. This is combined with a novel modified rejection sampling scheme which preserves the distribution of the target model within hardware numerics. We benchmark speculative sampling with Chinchilla, a 70 billion parameter language model, achieving a 2-2.5x decoding speedup in a distributed setup, without compromising the sample quality or making modifications to the model itself.
Cited by
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- WISERouter: LLM Routing with Workload Budget Constraint
- KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
- DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
- Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
- FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
- PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- Parallel Token Prediction for Language Models
- Fast Collaborative Inference via Distributed Speculative Decoding
- LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
- Optimizing Agentic Language Model Inference via Speculative Tool Calls
- DEER: Draft with Diffusion, Verify with Autoregressive Models
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- RADAR: Accelerating Large Language Model Inference With RL-Based Dynamic Draft Trees
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
- AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
- Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
- T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
- RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
- DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
- Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
- DiFR: Inference Verification Despite Nondeterminism
- Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
- F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
- Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits
- Steering Pretrained Drafters during Speculative Decoding
- Parallel Sampling via Autospeculation
- When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
- Next-Latent Prediction Transformers Learn Compact World Models
- Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
- Verifying LLM Inference to Prevent Model Weight Exfiltration
- TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
- Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
- Collaborative Large Language Model Inference via Resource-Aware Parallel Speculative Decoding
- SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
- Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
- SpecAttn: Speculating Sparse Attention
- Kad: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral
- The End of Manual Decoding: Towards Truly End-to-End Language Models
- Polybasic Speculative Decoding Through a Theoretical Perspective
- ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- PSG: Pair-Space Generation for Efficient Generative Reranking
- SSV: Sparse Speculative Verification for Efficient LLM Inference
- What Limits Agentic Systems Efficiency?
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Batch Speculative Decoding Done Right
- Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- Fast Inference via Hierarchical Speculative Decoding
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- Planned Diffusion
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling
- Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- A Survey on Parallel Reasoning
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- 3-Model Speculative Decoding
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
- Faster LLM Inference via Sequential Monte Carlo
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- Efficient Autoregressive Inference for Transformer Probabilistic Models
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
- Speculative Actions: A Lossless Framework for Faster Agentic Systems
- Auditing Pay-Per-Token in Large Language Models
- Self Speculative Decoding for Diffusion Large Language Models
- Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- Drax: Speech Recognition with Discrete Flow Matching
- Self-Speculative Masked Diffusions
- HiSpec: Hierarchical Speculative Decoding for LLMs
- Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
- SpecExit: Accelerating Large Reasoning Model via Speculative Exit
- Learning to Ponder: Adaptive Reasoning in Latent Space
- DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- Infusing Theory of Mind into Socially Intelligent LLM Agents
- Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
- Self-Speculative Biased Decoding for Faster Live Translation
- We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
- FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
- Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
- Divergence Decoding: Training-Free Capability Fusion
- Speculative Safety-Aware Decoding
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
- Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs
- Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
- ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- LATTS: Locally Adaptive Test-Time Scaling
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding
- SpecVLM: Fast Speculative Decoding in Vision-Language Models
- Recurrent State Encoders for Efficient Neural Combinatorial Optimization
- Communication-Efficient Collaborative LLM Inference via Distributed Speculative Decoding
- Set Block Decoding is a Language Model Inference Accelerator
- DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- Confidence-Modulated Speculative Decoding for Large Language Models
- Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
- PC-Sampler: Position-Aware Calibration of Decoding Bias in Masked Diffusion Models
- Cost-Aware Contrastive Routing for LLMs
- Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Camel: Energy-Aware LLM Inference on Resource-Constrained Devices
- CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
- An Efficient and Adaptive Next Edit Suggestion Framework with Zero Human Instructions in IDEs
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
Related