vix.ing · top · new · best · stats · spec

Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods

2025/02/03 by Oussama Zekri, Oussama Zékri, Nicolas Boullé +2 · 1 voice · 11 citations
Economics, Econometrics and Finance · #Climate Change Policy and Economics #cs.AI #cs.CL #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2502.01384

openalex publication_date 2025/02/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

Discrete diffusion models have recently gained significant attention due to their ability to process complex discrete structures for language modeling. However, fine-tuning these models with policy gradient methods, as is commonly done in Reinforcement Learning from Human Feedback (RLHF), remains a challenging task. We propose an efficient, broadly applicable, and theoretically justified policy gradient algorithm, called Score Entropy Policy Optimization (\SEPO), for fine-tuning discrete diffusion models over non-differentiable rewards. Our numerical experiments across several discrete generative tasks demonstrate the scalability and efficiency of our method. Our code is available at https://github.com/ozekri/SEPO.

Cited by

Discussions

Related