Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
2022/09/07 by Xingchao Liu, Liu, Xingchao, Chengyue Gong +3 · 3 voices · 713 citations
Physics and Astronomy · Computer Science · #cs.LG
paper · pdf · doi:10.48550/arxiv.2209.03003
Abstract
We present rectified flow, a surprisingly simple approach to learning (neural) ordinary differential equation (ODE) models to transport between two empirically observed distributions π0 and π1, hence providing a unified solution to generative modeling and domain transfer, among various other tasks involving distribution transport. The idea of rectified flow is to learn the ODE to follow the straight paths connecting the points drawn from π0 and π1 as much as possible. This is achieved by solving a straightforward nonlinear least squares optimization problem, which can be easily scaled to large models without introducing extra parameters beyond standard supervised learning. The straight paths are special and preferred because they are the shortest paths between two points, and can be simulated exactly without time discretization and hence yield computationally efficient models. We show that the procedure of learning a rectified flow from data, called rectification, turns an arbitrary coupling of π0 and π1 to a new deterministic coupling with provably non-increasing convex transport costs. In addition, recursively applying rectification allows us to obtain a sequence of flows with increasingly straight paths, which can be simulated accurately with coarse time discretization in the inference phase. In empirical studies, we show that rectified flow performs superbly on image generation, image-to-image translation, and domain adaptation. In particular, on image generation and translation, our method yields nearly straight flows that give high quality results even with a single Euler discretization step.
Cited by
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
- High-Performance Self-Supervised Learning by Joint Training of Flow Matching
- FlowLPS: Langevin-Proximal Sampling for Flow-based Inverse Problem Solvers
- Qwen-Music Technical Report
- ThinkGen: Generalized Thinking for Visual Generation
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
- Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- On the Inverse Flow Matching Problem in the One-Dimensional and Gaussian Cases
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
- Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching
- Probing the Geometry of Diffusion Models with the String Method
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
- Energy-Guided Flow Matching Enables Few-Step Conformer Generation and Ground-State Identification
- IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
- Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- SpotEdit: Selective Region Editing in Diffusion Transformers
- LangPrecip: Language-Aware Multimodal Precipitation Nowcasting
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Joint Flow Matching for Generator-Consistent Classification
- Generative Modeling by Value-Driven Transport
- Normalizing Trajectory Models
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- STEER: Steerable Dyadic Head Avatars
- Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation
- Post-FWI Injection of Learned Priors Using a Flow Matching Model
- UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow
- Low-Latency Generative Semantic Communication via Channel-Realization Flow Matching
- Generative Video Compression with Adaptive Score Distillation
- ProEdit: Inversion-based Editing From Prompts Done Right
- ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
- Feedforward 3D Editing Learns from Semantic-Part Transformation
- AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
- Tilt Matching for Scalable Sampling and Fine-Tuning
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
- UPMRI: Unsupervised Parallel MRI Reconstruction via Projected Conditional Flow Matching
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
- AstraNav-World: World Model for Foresight Control and Consistency
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- Generative Spectrum Cartography: Unified Reconstruction and Active Sensing via Diffusion Models
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- Discovering Symmetry Groups with Flow Matching
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- Variational Autoregressive Networks Applied to ϕ4 Field Theory Systems
- StoryMem: Multi-shot Long Video Storytelling with Memory
- OMP: One-step Meanflow Policy with Directional Alignment
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- In-Context Audio Control of Video Diffusion Transformers
- Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
- AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
- Robust and scalable simulation-based inference for gravitational wave signals with gaps
- Loom: Diffusion-Transformer for Interleaved Generation
- Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Generative modeling of conditional probability distributions on the level-sets of collective variables
- SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samples
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- FlowDet: Unifying Object Detection and Generative Transport Flows
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
- Accelerating High-Throughput Catalyst Screening by Direct Generation of Equilibrium Adsorption Structures
- CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
- Native and Compact Structured Latents for 3D Generation
- OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
- Random-Bridges as Stochastic Transports for Generative Models
- FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- Generative Monte Carlo Sampling for Constant-Cost Particle Transport
- MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
- World Models Can Leverage Human Videos for Dexterous Manipulation
- Do-Undo: Generating and Reversing Physical Actions in Vision-Language Models
- Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
- RecTok: Reconstruction Distillation along Rectified Flow
- BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
- Exploring the Design Space of Transition Matching
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment
- PIS: A Generalized Physical Inversion Solver for Arbitrary Sparse Observations via Set Conditioned Flow Matching
- Boosting Monocular Metric Depth Estimation via Bokeh Rendering
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- Bidirectional Normalizing Flow: From Data to Noise and Back
- OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
- Generative Modeling from Black-box Corruptions via Self-Consistent Stochastic Interpolants
- Composing Concepts from Images and Videos via Concept-prompt Binding
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- OmniPSD: Layered PSD Generation with Diffusion Transformer
- TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
- Inferring Compositional 4D Scenes without Ever Seeing One
- One-Step Diffusion Samplers via Self-Distillation and Deterministic Flow
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- FlowSteer: Conditioning Flow Field for Consistent Image Restoration
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Scalable Offline Model-Based RL with Action Chunks
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation
- JoPano: Unified Panorama Generation via Joint Modeling
- IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Pathway to O(√(d)) Complexity bound under Wasserstein metric of flow-based models
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- USV: Unified Sparsification for Accelerating Video Diffusion Models
- InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
- Value Gradient Guidance for Flow Matching Alignment
- NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
- YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models
- UniTS: Unified Time Series Generative Model for Remote Sensing
- ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Plug-and-Play Image Restoration with Flow Matching: A Continuous Viewpoint
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
- Glance: Accelerating Diffusion Models with 1 Sample
- From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Reversible Inversion for Training-Free Exemplar-guided Image Editing
- Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- Forecasting in Offline Reinforcement Learning for Non-stationary Environments
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- High-dimensional Mean-Field Games by Particle-based Flow Matching
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- Accelerating Inference of Masked Image Generators via Reinforcement Learning
- Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
- Flow Matching for Tabular Data Synthesis
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- Self-sufficient Independent Component Analysis via KL Minimizing Flows
- Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- PhysGen: Physically Grounded 3D Shape Generation for Industrial Design
- Vision Bridge Transformer at Scale
- Guiding Visual Autoregressive Models through Spectrum Weakening
- One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Overcoming the Curvature Bottleneck in MeanFlow
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- Test-time scaling of diffusions with flow maps
- Generative models for crystalline materials
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning
- CaFlow: Enhancing Long-Term Action Quality Assessment with Causal Counterfactual Flow
- Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
- Adversarial Flow Models
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- Video Generation Models Are Good Latent Reward Models
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
- MeanFlow Transformers with Representation Autoencoders
- MFM-point: Multi-scale Flow Matching for Point Cloud Generation
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
- Deep Parameter Interpolation for Scalar Conditioning
- Inversion-Free Style Transfer with Dual Rectified Flows
- From Diffusion to One-Step Generation: A Comparative Study of Flow-Based Models with Application to Image Inpainting
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- Segment-Wise Flow Matching for Vision-Aided mmWave V2I Beam Prediction
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- Restora-Flow: Mask-Guided Image Restoration with Flow Matching
- SONIC: Spectral Optimization of Noise for Inpainting with Consistency
- Low-Resolution Editing is All You Need for High-Resolution Editing
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
- Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
- TReFT: Taming Rectified Flow Models For One-Step Image Translation
- ShapeGen: Towards High-Quality 3D Shape Synthesis
- Terminal Velocity Matching
- Flow Map Distillation Without Data
- Efficiency vs. Fidelity: A Comparative Analysis of Diffusion Probabilistic Models and Flow Matching on Low-Resource Hardware
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- EnfoPath: Energy-Informed Analysis of Generative Trajectories in Flow Matching
- Understanding, Accelerating, and Improving MeanFlow Training
- FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
- Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
- VeCoR -- Velocity Contrastive Regularization for Flow Matching
- ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
- Breaking the Likelihood-Quality Trade-off in Diffusion Models by Merging Pretrained Experts
- GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- CoD: A Diffusion Foundation Model for Image Compression
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- FlowPortal: Residual-Corrected Flow for Training-Free Video Relighting and Background Replacement
- Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
- L1 Sample Flow for Efficient Visuomotor Learning
- Score-Regularized Joint Sampling with Importance Weights for Flow Matching
- One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
- Generative Photographic Control for Scene-Consistent Video Cinematic Editing
- Generative design and validation of therapeutic peptides for glioblastoma based on a potential target ATP5A
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching
- Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow Networks
- Towards Stable and Structured Time Series Generation with Perturbation-Aware Flow Matching
- 3D-Guided Scalable Flow Matching for Generating Volumetric Tissue Spatial Transcriptomics from Serial Histology
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving
- Functional Mean Flow in Hilbert Space
- Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
- Lightweight Optimal-Transport Harmonization on Edge Devices
- Learning Straight Flows: Variational Flow Matching for Efficient Generation
- FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- Fast Data Attribution for Text-to-Image Models
- FlowCast: Advancing Precipitation Nowcasting with Conditional Flow Matching
- Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition
- SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
- HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- Rectified Noise: A Generative Model Using Positive-incentive Noise
- Beyond the Training Domain: Robust Generative Transition State Models for Unseen Chemistry
- MotionStream: Real-Time Video Generation with Interactive Motion Controls
- Integrating Reweighted Least Squares with Plug-and-Play Diffusion Priors for Noisy Image Restoration
- Controllable Flow Matching for Online Reinforcement Learning
- Image Restoration via Primal Dual Hybrid Gradient and Flow Generative Model
- A Risk-Neutral Neural Operator for Arbitrage-Free SPX-VIX Term Structures
- Latent Refinement via Flow Matching for Training-free Linear Inverse Problem Solving
- Neodragon: Mobile Video Generation using Diffusion Transformer
- On Flow Matching KL Divergence
- Efficient probabilistic surrogate modeling techniques for partially-observed large-scale dynamical systems
- MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
- Learning Paths for Dynamic Measure Transport: A Control Perspective
- Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
- IllumFlow: Illumination-Adaptive Low-Light Enhancement via Conditional Rectified Flow and Retinex Decomposition
- Lightweight Learning from Actuation-Space Demonstrations via Flow Matching for Whole-Body Soft Robotic Grasping
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Modeling Microenvironment Trajectories on Spatial Transcriptomics with NicheFlow
- RefVTON: person-to-person Try on with Additional Unpaired Visual Reference
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- On the Equivalence of Optimal Transport Problem and Action Matching with Optimal Vector Fields
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- GeneFlow: Translation of Single-cell Gene Expression to Histopathological Images via Rectified Flow
- Curly Flow Matching for Learning Non-gradient Field Dynamics
- Co-Evolving Latent Action World Models
- Efficient Generative AI Boosts Probabilistic Forecasting of Sudden Stratospheric Warmings
- SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
- RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- Flow Map Learning via Nongradient Vector Flow
- RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows
- DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
- Amortized Moment Matching for Visual Generation
- Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data
- Flow Matching for Measure Transport and Feedback Stabilization of Control-Affine Systems
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
- Surflo: Consistent 3D Surface Flow Model with Global State
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
- Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
- Generative Modeling via Drifting
- Distance Marching for Generative Modeling
- Meta Flow Maps enable scalable reward alignment
- LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency
- Balanced conic rectified flow
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Unsupervised Detection of Post-Stroke Brain Abnormalities
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- Residual Diffusion Bridge Model for Image Restoration
- LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
- Nested AutoRegressive Models
- Coupled Flow Matching
- Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling
- FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- ACG: Action Coherence Guidance for Flow-based VLA models
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
- Generalised Flow Maps for Few-Step Generative Modelling on Riemannian Manifolds
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation
- Improved Training Technique for Shortcut Models
- On the flow matching interpretability
- Epipolar Geometry Improves Video Generation Models
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- GuitarFlow: Realistic Electric Guitar Synthesis From Tablatures via Flow Matching and Style Transfer
- Positional Encoding Field
- Target-aware Image Editing via Cycle-consistent Constraints
- Diffusion Bridge Networks Simulate Clinical-grade PET from MRI for Dementia Diagnostics
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- AlphaFlow: Understanding and Improving MeanFlow Models
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- Gradient Variance Reveals Failure Modes in Flow-Based Generative Models
- Latent-Augmented Discrete Diffusion Models
- Demystifying Transition Matching: When and Why It Can Beat Flow Matching
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- Estimating Orbital Parameters of Direct Imaging Exoplanet Using Neural Network
- On Efficiency-Effectiveness Trade-off of Diffusion-based Recommenders
- Adaptive Discretization for Consistency Models
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction
- Latent Diffusion Model without Variational Autoencoder
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport
- Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
- Learning an Image Editing Model without Image Editing Pairs
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Exploring Cross-Modal Flows for Few-Shot Learning
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- Graph Generation with Spectral Geodesic Flow Matching
- VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- DiffLoc: Diffusion Model-Based High-Precision Positioning for 6G Networks
- NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models
- Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- Adapting Noise to Data: Generative Flows from 1D Processes
- WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation
- Diffusion Transformers with Representation Autoencoders
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- PAINT: Parallel-in-time Neural Twins for Dynamical System Reconstruction
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Y-shaped Generative Flows
- Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
- An Eulerian Perspective on Straight-Line Sampling
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- SoundReactor: Frame-level Online Video-to-Audio Generation
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions
- EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
- RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
- Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
- Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Asymmetric Flow Models
- Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
- Conditional Flow Matching for Bayesian Posterior Inference
- Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex Domains
- FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
- Counterfactual Identifiability via Dynamic Optimal Transport
- Expressive Value Learning for Scalable Offline Reinforcement Learning
- IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
- Ctrl-VI: Controllable Video Synthesis via Variational Inference
- Deep Neural Networks Inspired by Differential Equations
- Value Flows
- Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models
- RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
- Optimal Stopping in Latent Diffusion Models
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- SkipSR: Faster Super Resolution with Token Skipping
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- A Denoising Framework for Real-World Ultra-Low-Dose Lung CT Images Based on an Image Purification Strategy
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- DisCo: Reinforcement with Diversity Constraints for Multi-Human Generation
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Three Forms of Stochastic Injection for Improved Distribution-to-Distribution Generative Modeling
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- ESS-Flow: Training-free guidance of flow-based models as inference in source space
- Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Mutual Information Estimation via Score-to-Fisher Bridge for Nonlinear Gaussian Noise Channels
- FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Flow-Matching Based Refiner for Molecular Conformer Generation
- Flow Matching for Conditional MRI-CT and CBCT-CT Image Synthesis
- Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
- Fine-Grained GRPO for Precise Preference Alignment in Flow Models
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Distilled Protein Backbone Generation
- ContextFlow: Context-Aware Flow Matching For Trajectory Inference From Spatial Omics Data
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- LVTINO: LAtent Video consisTency INverse sOlver for High Definition Video Restoration
- Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models
- Multi-Marginal Flow Matching with Adversarially Learnt Interpolants
- Quantifying the noise sensitivity of the Wasserstein metric for images
- Riemannian Consistency Model
- Align Your Tangent: Training Better Consistency Models via Manifold-Aligned Tangents
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- Arbitrary Generative Video Interpolation
- Selective Underfitting in Diffusion Models
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Video Object Segmentation-Aware Audio Generation
- DiffCamera: Arbitrary Refocusing on Images
- Data-to-Energy Stochastic Dynamics
- FLOWER: A Flow-Matching Solver for Inverse Problems
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- EVODiff: Entropy-aware Variance Optimized Diffusion Inference
- Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
- Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image Manipulation
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
- Reweighted Flow Matching via Unbalanced OT for Label-free Long-tailed Generation
- LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
- Swift: An Autoregressive Consistency Model for Efficient Weather Forecasting
- Reframing Generative Models for Physical Systems using Stochastic Interpolants
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Flow Matching with Semidiscrete Couplings
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Chance-constrained Flow Matching for High-Fidelity Constraint-aware Generation
- Score Distillation of Flow Matching Models
- MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation
- MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow Map Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- GLASS Flows: Transition Sampling for Alignment of Flow and Diffusion Models
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Training Agents Inside of Scalable World Models
- OAT-FM: Optimal Acceleration Transport for Improved Flow Matching
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- GeoFunFlow: Geometric Function Flow Matching for Inverse Operator Learning over Complex Geometries
- Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
- Space Group Conditional Flow Matching
- DA-MMP: Learning Coordinated and Accurate Throwing with Dynamics-Aware Motion Manifold Primitives
- FlowLUT: Efficient Image Enhancement via Differentiable LUTs and Iterative Flow Matching
- HunyuanImage 3.0 Technical Report
- Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- RestoRect: Degraded Image Restoration via Latent Rectified Flow & Feature Distillation
- Impute-MACFM: Imputation based on Mask-Aware Flow Matching
- Transport Based Mean Flows for Generative Modeling
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- NIFTY: a Non-Local Image Flow Matching for Texture Synthesis
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- Universal Multi-Domain Translation via Diffusion Routers
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- Taming Flow-based I2V Models for Creative Video Editing
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Error Analysis of Discrete Flow with Generator Matching
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
- DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- Consistency Models as Plug-and-Play Priors for Inverse Problems
- Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
- DistillKac: Few-Step Image Generation via Damped Wave Equations
- Score-based Idempotent Distillation of Diffusion Models
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Energy Guided Geometric Flow Matching
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Federated Flow Matching
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
- A Recovery Theory for Diffusion Priors: Deterministic Analysis of the Implicit Prior Algorithm
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
- RFMSR: Residual Flow Matching for Image Super-Resolution
- RIPPLE: Generating Multi-Channel Phase, Not Recovering It
- Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework
- From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
- Matérn Noise for Triangulation-Agnostic Flow Matching on Meshes
- ELF: Embedded Language Flows
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
- Prompt-Guided Dual Latent Steering for Inversion Problems
- Flow marching for a generative PDE foundation model
- CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- EigenSafe: A Spectral Framework for Learning-Based Stochastic Safety Filtering
- SISMA: Semantic Face Image Synthesis with Mamba
- Discrete-Time Diffusion-Like Models for Speech Synthesis
- STAR: Speech-to-Audio Generation via Representation Learning
- Preference Trajectory Modeling via Flow Matching for Sequential Recommendation
- JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
- ReSeFlow: Rectifying SE(3)-Equivariant Policy Learning Flows
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
- OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Efficient Rectified Flow for Image Fusion
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World Dehazing
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors
- DiffusionNFT: Online Diffusion Reinforcement with Forward Process
- Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement
- Kuramoto Orientation Diffusion Models
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- MeanFlowSE: one-step generative speech enhancement via conditional mean flow
- FlowCast-ODE: Continuous Hourly Weather Forecasting with Dynamic Flow Matching and ODE Solver
- Masked Diffusion Models as Energy Minimization
- RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing
- DiCache: Let Diffusion Model Determine Its Own Cache
- Generative Consistency Models for Estimation of Kinetic Parametric Image Posteriors in Total-Body PET
- Dense-Jump Flow Matching with Non-Uniform Time Scheduling for Robotic Policies: Mitigating Multi-Step Inference Degradation
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
- PINGS: Physics-Informed Neural Network for Fast Generative Sampling
- Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
- Chord: Chain of Rendering Decomposition for PBR Material Estimation from Generated Texture Images
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
- LADB: Latent Aligned Diffusion Bridges for Semi-Supervised Domain Translation
- Physics-Guided Rectified Flow for Low-light RAW Image Enhancement
- ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Reconstruction Alignment Improves Unified Multimodal Models
- Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- CardiacFlow: 3D+t Four-Chamber Cardiac Shape Completion and Generation via Flow Matching
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
- Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
- Transition Models: Rethinking the Generative Learning Objective
- Diffusion Generative Models Meet Compressed Sensing, with Applications to Imaging and Finance
- Scale-Adaptive Generative Flows for Multiscale Scientific Data
- A-FloPS: Accelerating Diffusion Sampling with Adaptive Flow Path Sampler
- Distribution estimation via Flow Matching with Lipschitz guarantees
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom Number
- Latent Space Single-Pixel Imaging Under Low-Sampling Conditions
- Lipschitz-Guided Design of Interpolation Schedules in Generative Models
- Delta Rectified Flow Sampling for Text-to-Image Editing
- ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
- Challenges in Non-Polymeric Crystal Structure Prediction: Why a Geometric, Permutation-Invariant Loss is Needed
- Any-Order Flexible Length Masked Diffusion
- Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
- FNODE: Flow-Matching for data-driven simulation of constrained multibody systems
- PHD: Personalized 3D Human Body Fitting with Point Diffusion
- Efficient Diffusion Model for Image Restoration by Residual Shifting
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation
- Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
- OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
- Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees
- StableIntrinsic: Detail-preserving One-step Diffusion Model for Multi-view Material Estimation
- MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Energy-Based Flow Matching for Generating 3D Molecular Structure
- SAT-SKYLINES: 3D Building Generation from Satellite Imagery and Coarse Geometric Priors
- Provable Mixed-Noise Learning with Flow-Matching
- Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
- CurveFlow: Curvature-Guided Flow Matching for Image Generation
- Source-Guided Flow Matching
- CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
- Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- OmniTry: Virtual Try-On Anything without Masks
- EventTSF: Event-Aware Non-Stationary Time Series Forecasting
- DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- Synthesizing Accurate and Realistic T1-weighted Contrast-Enhanced MR Images using Posterior-Mean Rectified Flow
- FlowMol3: Flow Matching for 3D De Novo Small-Molecule Generation
- FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
- Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models
- Constraint-Aware Flow Matching via Randomized Exploration
- Flow Matching for Efficient and Scalable Data Assimilation
- Next Visual Granularity Generation
- Distribution Matching via Generalized Consistency Models
- LoRAtorio: An intrinsic approach to LoRA Skill Composition
- Noise Matters: Optimizing Matching Noise for Diffusion Classifiers
- 3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- Flow-SLM: Joint Learning of Linguistic and Acoustic Information for Spoken Language Modeling
- Hybrid Long and Short Range Flows for Point Cloud Filtering
- Elucidating Rectified Flow with Deterministic Sampler: Polynomial Discretization Complexity for Multi and One-step Models
- Towards Safe Imitation Learning via Potential Field-Guided Flow Matching
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Score Augmentation for Diffusion Models
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- When and how can inexact generative models still sample from the data manifold?
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- PureSample: Neural Materials Learned by Sampling Microgeometry
- Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- Towards High-Order Mean Flow Generative Models: Feasibility, Expressivity, and Provably Efficient Criteria
- Elastic Diffusion Transformer
- OM2P: Offline Multi-Agent Mean-Flow Policy
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- How and Why: Taming Flow Matching for Unsupervised Anomaly Detection and Localization
- MolSnap: Snap-Fast Molecular Generation with Latent Variational Mean Flow
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Multitask Learning with Stochastic Interpolants
- LayerT2V: Interactive Multi-Object Trajectory Layering for Video Generation
- Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
- SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
- REFLECT: Rectified Flows for Efficient Brain Anomaly Correction Transport
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- Qwen-Image Technical Report
- Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
- Diffusion models for inverse problems
- VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation
- Flow Matching for Probabilistic Learning of Dynamical Systems from Missing or Noisy Data
- PnP-DA: Towards Principled Plug-and-Play Integration of Variational Data Assimilation and Generative Models
- CIF: A Constrained Inversion Framework for Reliable Message Extraction in Diffusion-Based Generative Steganography
- One-Step Flow Policy Mirror Descent
- FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
- Training-free Geometric Image Editing on Diffusion Models
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- Next Tokens Denoising for Speech Synthesis
- Weighted Conditional Flow Matching
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- A Diffusion Model for POI Recommendation
- Why Flow Matching is Particle Swarm Optimization?
- Conditional Diffusion Models for Global Precipitation Map Inpainting
- A New One-Shot Federated Learning Framework for Medical Imaging Classification with Feature-Guided Rectified Flow and Knowledge Distillation
- A Survey of Multimodal Hallucination Evaluation and Detection
- Flow Stochastic Segmentation Networks
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
- TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
- Equivariant Volumetric Grasping
Discussions
Related