2025/05/30 by Tianqi Chen, Shujian Zhang, Chen, Tianqi +3 · 4 citations
Computer Science · #Benchmark (surveying) #Computation and Language (cs.CL) #Diffusion #Embedding #FOS: Computer and information sciences #Generality #Inference #Language model #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model-Driven Software Engineering Techniques #Natural Language Processing Techniques #Sequence (biology) #Speech and dialogue systems #Speedup
paper · pdf · doi:10.48550/arxiv.2506.00290
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
This paper introduces DLM-One, a score-distillation-based framework for one-step sequence generation with continuous diffusion language models (DLMs). DLM-One eliminates the need for iterative refinement by aligning the scores of a student model's outputs in the continuous token embedding space with the score function of a pretrained teacher DLM. We investigate whether DLM-One can achieve substantial gains in sampling efficiency for language modeling. Through comprehensive experiments on DiffuSeq -- a representative continuous DLM -- we show that DLM-One achieves up to ~500x speedup in inference time while maintaining competitive performance on benchmark text generation tasks used to evaluate the teacher models. We further analyze the method's empirical behavior across multiple datasets, providing initial insights into its generality and practical applicability. Our findings position one-step diffusion as a promising direction for efficient, high-quality language generation and broader adoption of continuous diffusion models operating in embedding space for natural language processing.