vix.ing · top · new · best · stats · spec

Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

2024/11/22 by Zhiwei Jia, Yuesong Nan, Jia, Zhiwei +5 · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech Recognition and Synthesis

paper · pdf · doi:10.48550/arxiv.2411.15247

openalex publication_date 2024/11/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Recent research has shown that fine-tuning diffusion models (DMs) with arbitrary rewards, including non-differentiable ones, is feasible with reinforcement learning (RL) techniques, enabling flexible model alignment. However, applying existing RL methods to step-distilled DMs is challenging for ultra-fast (≤2-step) image generation. Our analysis suggests several limitations of policy-based RL methods such as PPO or DPO toward this goal. Based on the insights, we propose fine-tuning DMs with learned differentiable surrogate rewards. Our method, named LaSRO, learns surrogate reward models in the latent space of SDXL to convert arbitrary rewards into differentiable ones for effective reward gradient guidance. LaSRO leverages pre-trained latent DMs for reward modeling and tailors reward optimization for ≤2-step image generation with efficient off-policy exploration. LaSRO is effective and stable for improving ultra-fast image generation with different reward objectives, outperforming popular RL methods including DDPO and Diffusion-DPO. We further show LaSRO's connection to value-based RL, providing theoretical insights. See our webpage \hrefhttps://sites.google.com/view/lasrohere.

Cited by

Related