2024/09/13 by Carles Domingo-Enrich, Michal Drozdzal, Domingo-Enrich, Carles +5 · 2 voices · 77 citations
Computer Science · Economics, Econometrics and Finance · Mathematics · #Applied mathematics #Artificial intelligence #Computer science #Diffusion #Flow (mathematics) #Generative grammar #Generative model #Geometry #Markov Chains and Monte Carlo Methods #Matching (statistics) #Mathematical optimization #Mathematics #Optimal control #Physics #Statistical physics #Statistics #Stochastic control #Stochastic processes and financial applications #Thermodynamics #cs.LG #math.OC #stat.ML
paper · pdf · doi:10.48550/arxiv.2409.08861
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/09/13 · arxiv published 2024/09/13 · openalex created_date 2024/10/23 · arxiv updated 2025/01/07 · openalex updated_date 2026/07/28
Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fine-tuning as stochastic optimal control (SOC). Critically, we prove that a very specific memoryless noise schedule must be enforced during fine-tuning, in order to account for the dependency between the noise variable and the generated samples. We also propose a new algorithm named Adjoint Matching which outperforms existing SOC algorithms, by casting SOC problems as a regression problem. We find that our approach significantly improves over existing methods for reward fine-tuning, achieving better consistency, realism, and generalization to unseen human preference reward models, while retaining sample diversity.