Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
2022/09/07 by Xingchao Liu, Liu, Xingchao, Chengyue Gong +3 · 3 voices · 1,196 citations
Physics and Astronomy · Computer Science · #cs.LG
paper · pdf · doi:10.48550/arxiv.2209.03003
Abstract
We present rectified flow, a surprisingly simple approach to learning (neural) ordinary differential equation (ODE) models to transport between two empirically observed distributions π0 and π1, hence providing a unified solution to generative modeling and domain transfer, among various other tasks involving distribution transport. The idea of rectified flow is to learn the ODE to follow the straight paths connecting the points drawn from π0 and π1 as much as possible. This is achieved by solving a straightforward nonlinear least squares optimization problem, which can be easily scaled to large models without introducing extra parameters beyond standard supervised learning. The straight paths are special and preferred because they are the shortest paths between two points, and can be simulated exactly without time discretization and hence yield computationally efficient models. We show that the procedure of learning a rectified flow from data, called rectification, turns an arbitrary coupling of π0 and π1 to a new deterministic coupling with provably non-increasing convex transport costs. In addition, recursively applying rectification allows us to obtain a sequence of flows with increasingly straight paths, which can be simulated accurately with coarse time discretization in the inference phase. In empirical studies, we show that rectified flow performs superbly on image generation, image-to-image translation, and domain adaptation. In particular, on image generation and translation, our method yields nearly straight flows that give high quality results even with a single Euler discretization step.
Cited by
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
- High-Performance Self-Supervised Learning by Joint Training of Flow Matching
- FlowLPS: Langevin-Proximal Sampling for Flow-based Inverse Problem Solvers
- Qwen-Music Technical Report
- ThinkGen: Generalized Thinking for Visual Generation
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
- Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- On the Inverse Flow Matching Problem in the One-Dimensional and Gaussian Cases
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
- Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching
- Probing the Geometry of Diffusion Models with the String Method
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
- Energy-Guided Flow Matching Enables Few-Step Conformer Generation and Ground-State Identification
- IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
- Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- SpotEdit: Selective Region Editing in Diffusion Transformers
- LangPrecip: Language-Aware Multimodal Precipitation Nowcasting
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Joint Flow Matching for Generator-Consistent Classification
- Generative Modeling by Value-Driven Transport
- Normalizing Trajectory Models
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- STEER: Steerable Dyadic Head Avatars
- Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation
- Post-FWI Injection of Learned Priors Using a Flow Matching Model
- UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow
- Low-Latency Generative Semantic Communication via Channel-Realization Flow Matching
- Generative Video Compression with Adaptive Score Distillation
- ProEdit: Inversion-based Editing From Prompts Done Right
- ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
- Feedforward 3D Editing Learns from Semantic-Part Transformation
- AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
- Tilt Matching for Scalable Sampling and Fine-Tuning
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
- UPMRI: Unsupervised Parallel MRI Reconstruction via Projected Conditional Flow Matching
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
- AstraNav-World: World Model for Foresight Control and Consistency
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- Generative Spectrum Cartography: Unified Reconstruction and Active Sensing via Diffusion Models
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- Discovering Symmetry Groups with Flow Matching
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- Variational Autoregressive Networks Applied to ϕ4 Field Theory Systems
- StoryMem: Multi-shot Long Video Storytelling with Memory
- OMP: One-step Meanflow Policy with Directional Alignment
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- In-Context Audio Control of Video Diffusion Transformers
- Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
- AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
- Robust and scalable simulation-based inference for gravitational wave signals with gaps
- Loom: Diffusion-Transformer for Interleaved Generation
- Is There a Better Source Distribution than Gaussian? Exploring Source Distributions for Image Flow Matching
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Generative modeling of conditional probability distributions on the level-sets of collective variables
- SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samples
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- FlowDet: Unifying Object Detection and Generative Transport Flows
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
- Accelerating High-Throughput Catalyst Screening by Direct Generation of Equilibrium Adsorption Structures
- CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
- Native and Compact Structured Latents for 3D Generation
- OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
- Random-Bridges as Stochastic Transports for Generative Models
- FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- Generative Monte Carlo Sampling for Constant-Cost Particle Transport
- MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
- World Models for Learning Dexterous Hand-Object Interactions from Human Videos
- Do-Undo Bench: Reversibility for Action Understanding in Image Generation
- Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
- RecTok: Reconstruction Distillation along Rectified Flow
- BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
- Exploring the Design Space of Transition Matching
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment
- PIS: A Generalized Physical Inversion Solver for Arbitrary Sparse Observations via Set Conditioned Flow Matching
- Boosting Monocular Metric Depth Estimation via Bokeh Rendering
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- Bidirectional Normalizing Flow: From Data to Noise and Back
- OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
- Generative Modeling from Black-box Corruptions via Self-Consistent Stochastic Interpolants
- Composing Concepts from Images and Videos via Concept-prompt Binding
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- OmniPSD: Layered PSD Generation with Diffusion Transformer
- TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
- Inferring Compositional 4D Scenes without Ever Seeing One
- One-Step Diffusion Samplers via Self-Distillation and Deterministic Flow
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- FlowSteer: Conditioning Flow Field for Consistent Image Restoration
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Scalable Offline Model-Based RL with Action Chunks
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation
- JoPano: Unified Panorama Generation via Joint Modeling
- IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Pathway to O(√(d)) Complexity bound under Wasserstein metric of flow-based models
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- USV: Unified Sparsification for Accelerating Video Diffusion Models
- InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
- Value Gradient Guidance for Flow Matching Alignment
- NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
- YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models
- UniTS: Unified Time Series Generative Model for Remote Sensing
- ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Plug-and-Play Image Restoration with Flow Matching: A Continuous Viewpoint
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
- Glance: Accelerating Diffusion Models with 1 Sample
- From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Reversible Inversion for Training-Free Exemplar-guided Image Editing
- Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- Forecasting in Offline Reinforcement Learning for Non-stationary Environments
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- High-dimensional Mean-Field Games by Particle-based Flow Matching
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- Accelerating Inference of Masked Image Generators via Reinforcement Learning
- Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
- Flow Matching for Tabular Data Synthesis
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- Self-sufficient Independent Component Analysis via KL Minimizing Flows
- Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- PhysGen: Physically Grounded 3D Shape Generation for Industrial Design
- Vision Bridge Transformer at Scale
- Guiding Visual Autoregressive Models through Spectrum Weakening
- One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Overcoming the Curvature Bottleneck in MeanFlow
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- Test-time scaling of diffusions with flow maps
- Generative models for crystalline materials
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning
- CaFlow: Enhancing Long-Term Action Quality Assessment with Causal Counterfactual Flow
- Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
- Adversarial Flow Models
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- Video Generation Models Are Good Latent Reward Models
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
- MeanFlow Transformers with Representation Autoencoders
- MFM-point: Multi-scale Flow Matching for Point Cloud Generation
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
- Deep Parameter Interpolation for Scalar Conditioning
- Inversion-Free Style Transfer with Dual Rectified Flows
- From Diffusion to One-Step Generation: A Comparative Study of Flow-Based Models with Application to Image Inpainting
- DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- Segment-Wise Flow Matching for Vision-Aided mmWave V2I Beam Prediction
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- Restora-Flow: Mask-Guided Image Restoration with Flow Matching
- SONIC: Spectral Optimization of Noise for Inpainting with Consistency
- Low-Resolution Editing is All You Need for High-Resolution Editing
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
- Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
- TReFT: Taming Rectified Flow Models For One-Step Image Translation
- ShapeGen: Towards High-Quality 3D Shape Synthesis
- Terminal Velocity Matching
- Flow Map Distillation Without Data
- Efficiency vs. Fidelity: A Comparative Analysis of Diffusion Probabilistic Models and Flow Matching on Low-Resource Hardware
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- EnfoPath: Energy-Informed Analysis of Generative Trajectories in Flow Matching
- Understanding, Accelerating, and Improving MeanFlow Training
- FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
- Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
- VeCoR -- Velocity Contrastive Regularization for Flow Matching
- ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
- Breaking the Likelihood-Quality Trade-off in Diffusion Models by Merging Pretrained Experts
- GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- CoD: A Diffusion Foundation Model for Image Compression
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- FlowPortal: Residual-Corrected Flow for Training-Free Video Relighting and Background Replacement
- Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
- L1 Sample Flow for Efficient Visuomotor Learning
- Score-Regularized Joint Sampling with Importance Weights for Flow Matching
- One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
- Generative Photographic Control for Scene-Consistent Video Cinematic Editing
- Generative design and validation of therapeutic peptides for glioblastoma based on a potential target ATP5A
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow Matching
- Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow Networks
- Towards Stable and Structured Time Series Generation with Perturbation-Aware Flow Matching
- 3D-Guided Scalable Flow Matching for Generating Volumetric Tissue Spatial Transcriptomics from Serial Histology
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- DriveLiDAR4D: Sequential and Controllable LiDAR Scene Generation for Autonomous Driving
- Functional Mean Flow in Hilbert Space
- Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
- Lightweight Optimal-Transport Harmonization on Edge Devices
- Learning Straight Flows: Variational Flow Matching for Efficient Generation
- FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- Fast Data Attribution for Text-to-Image Models
- FlowCast: Advancing Precipitation Nowcasting with Conditional Flow Matching
- Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition
- SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
- HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- Rectified Noise: A Generative Model Using Positive-incentive Noise
- Beyond the Training Domain: Robust Generative Transition State Models for Unseen Chemistry
- MotionStream: Real-Time Video Generation with Interactive Motion Controls
- Integrating Reweighted Least Squares with Plug-and-Play Diffusion Priors for Noisy Image Restoration
- Controllable Flow Matching for Online Reinforcement Learning
- Image Restoration via Primal Dual Hybrid Gradient and Flow Generative Model
- A Risk-Neutral Neural Operator for Arbitrage-Free SPX-VIX Term Structures
- Latent Refinement via Flow Matching for Training-free Linear Inverse Problem Solving
- Neodragon: Mobile Video Generation using Diffusion Transformer
- On Flow Matching KL Divergence
- Efficient probabilistic surrogate modeling techniques for partially-observed large-scale dynamical systems
- MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
- Learning Paths for Dynamic Measure Transport: A Control Perspective
- Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
- IllumFlow: Illumination-Adaptive Low-Light Enhancement via Conditional Rectified Flow and Retinex Decomposition
- Lightweight Learning from Actuation-Space Demonstrations via Flow Matching for Whole-Body Soft Robotic Grasping
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Modeling Microenvironment Trajectories on Spatial Transcriptomics with NicheFlow
- RefTon: Reference person shot assist virtual Try-on
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- On the Equivalence of Optimal Transport Problem and Action Matching with Optimal Vector Fields
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- GeneFlow: Translation of Single-cell Gene Expression to Histopathological Images via Rectified Flow
- Curly Flow Matching for Learning Non-gradient Field Dynamics
- Co-Evolving Latent Action World Models
- Efficient Generative AI Boosts Probabilistic Forecasting of Sudden Stratospheric Warmings
- SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
- RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
- Flow Map Learning via Nongradient Vector Flow
- RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows
- DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
- Amortized Moment Matching for Visual Generation
- Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data
- Flow Matching for Measure Transport and Feedback Stabilization of Control-Affine Systems
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
- Surflo: Consistent 3D Surface Flow Model with Global State
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
- Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
- Generative Modeling via Drifting
- Distance Marching for Generative Modeling
- Meta Flow Maps enable scalable reward alignment
- LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency
- Balanced conic rectified flow
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Unsupervised Detection of Post-Stroke Brain Abnormalities
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- Residual Diffusion Bridge Model for Image Restoration
- LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
- Nested AutoRegressive Models
- Coupled Flow Matching
- Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling
- FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- ACG: Action Coherence Guidance for Flow-based VLA models
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
- Generalised Flow Maps for Few-Step Generative Modelling on Riemannian Manifolds
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation
- Improved Training Technique for Shortcut Models
- On the flow matching interpretability
- Epipolar Geometry Improves Video Generation Models
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- GuitarFlow: Realistic Electric Guitar Synthesis From Tablatures via Flow Matching and Style Transfer
- Positional Encoding Field
- Target-aware Image Editing via Cycle-consistent Constraints
- Diffusion Bridge Networks Simulate Clinical-grade PET from MRI for Dementia Diagnostics
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- AlphaFlow: Understanding and Improving MeanFlow Models
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- Gradient Variance Reveals Failure Modes in Flow-Based Generative Models
- Latent-Augmented Discrete Diffusion Models
- Demystifying Transition Matching: When and Why It Can Beat Flow Matching
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- Estimating Orbital Parameters of Direct Imaging Exoplanet Using Neural Network
- On Efficiency-Effectiveness Trade-off of Diffusion-based Recommenders
- Adaptive Discretization for Consistency Models
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction
- Latent Diffusion Model without Variational Autoencoder
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport
- Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
- Learning an Image Editing Model without Image Editing Pairs
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Exploring Cross-Modal Flows for Few-Shot Learning
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- Graph Generation with Spectral Geodesic Flow Matching
- Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- DiffLoc: Diffusion Model-Based High-Precision Positioning for 6G Networks
- NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models
- Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- Adapting Noise to Data: Generative Flows from 1D Processes
- WaterFlow: Explicit Physics-Prior Rectified Flow for Underwater Saliency Mask Generation
- Diffusion Transformers with Representation Autoencoders
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- PAINT: Parallel-in-time Neural Twins for Dynamical System Reconstruction
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Y-Shaped Generative Flows
- Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
- An Eulerian Perspective on Straight-Line Sampling
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- SoundReactor: Frame-level Online Video-to-Audio Generation
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions
- EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
- RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
- Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
- Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Asymmetric Flow Models
- Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
- Conditional Flow Matching for Bayesian Posterior Inference
- Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex Domains
- FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
- Counterfactual Identifiability via Dynamic Optimal Transport
- Expressive Value Learning for Scalable Offline Reinforcement Learning
- IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
- Ctrl-VI: Controllable Video Synthesis via Variational Inference
- Deep Neural Networks Inspired by Differential Equations
- Value Flows
- Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models
- RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
- Optimal Stopping in Latent Diffusion Models
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- SkipSR: Faster Super Resolution with Token Skipping
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- A Denoising Framework for Real-World Ultra-Low-Dose Lung CT Images Based on an Image Purification Strategy
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- Resolving the Identity Crisis in Text-to-Image Generation
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Three Forms of Stochastic Injection for Improved Distribution-to-Distribution Generative Modeling
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- ESS-Flow: Training-free guidance of flow-based models as inference in source space
- Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Mutual Information Estimation via Score-to-Fisher Bridge for Nonlinear Gaussian Noise Channels
- FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Flow-Matching Based Refiner for Molecular Conformer Generation
- Flow Matching for Conditional MRI-CT and CBCT-CT Image Synthesis
- Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
- Fine-Grained GRPO for Precise Preference Alignment in Flow Models
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Distilled Protein Backbone Generation
- ContextFlow: Context-Aware Flow Matching For Trajectory Inference From Spatial Omics Data
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- LVTINO: LAtent Video consisTency INverse sOlver for High Definition Video Restoration
- Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models
- Multi-Marginal Flow Matching with Adversarially Learnt Interpolants
- Quantifying the noise sensitivity of the Wasserstein metric for images
- Riemannian Consistency Model
- Align Your Tangent: Training Better Consistency Models via Manifold-Aligned Tangents
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- Arbitrary Generative Video Interpolation
- Selective Underfitting in Diffusion Models
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Video Object Segmentation-Aware Audio Generation
- DiffCamera: Arbitrary Refocusing on Images
- Data-to-Energy Stochastic Dynamics
- FLOWER: A Flow-Matching Solver for Inverse Problems
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- EVODiff: Entropy-aware Variance Optimized Diffusion Inference
- Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
- Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image Manipulation
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
- Reweighted Flow Matching via Unbalanced OT for Label-free Long-tailed Generation
- LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
- Swift: An Autoregressive Consistency Model for Efficient Weather Forecasting
- Reframing Generative Models for Physical Systems using Stochastic Interpolants
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Flow Matching with Semidiscrete Couplings
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Chance-constrained Flow Matching for High-Fidelity Constraint-aware Generation
- Score Distillation of Flow Matching Models
- MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation
- MarS-FM: Generative Modeling of Molecular Dynamics via Markov State Models
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow Map Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
- Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- GLASS Flows: Transition Sampling for Alignment of Flow and Diffusion Models
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Training Agents Inside of Scalable World Models
- OAT-FM: Optimal Acceleration Transport for Improved Flow Matching
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- GeoFunFlow: Geometric Function Flow Matching for Inverse Operator Learning over Complex Geometries
- Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
- Space Group Conditional Flow Matching
- DA-MMP: Learning Coordinated and Accurate Throwing with Dynamics-Aware Motion Manifold Primitives
- FlowLUT: Efficient Image Enhancement via Differentiable LUTs and Iterative Flow Matching
- HunyuanImage 3.0 Technical Report
- Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- RestoRect: Degraded Image Restoration via Latent Rectified Flow & Feature Distillation
- Impute-MACFM: Imputation based on Mask-Aware Flow Matching
- Transport Based Mean Flows for Generative Modeling
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- NIFTY: a Non-Local Image Flow Matching for Texture Synthesis
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- Universal Multi-Domain Translation via Diffusion Routers
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- Taming Flow-based I2V Models for Creative Video Editing
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Error Analysis of Discrete Flow with Generator Matching
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
- DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- Consistency Models as Plug-and-Play Priors for Inverse Problems
- Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
- DistillKac: Few-Step Image Generation via Damped Wave Equations
- Score-based Idempotent Distillation of Diffusion Models
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Energy Guided Geometric Flow Matching
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- Federated Flow Matching
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
- A Recovery Theory for Diffusion Priors: Deterministic Analysis of the Implicit Prior Algorithm
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
- RFMSR: Residual Flow Matching for Image Super-Resolution
- RIPPLE: Generating Multi-Channel Phase, Not Recovering It
- Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework
- From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
- Matérn Noise for Triangulation-Agnostic Flow Matching on Meshes
- ELF: Embedded Language Flows
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
- Prompt-Guided Dual Latent Steering for Inversion Problems
- Flow marching for a generative PDE foundation model
- CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- EigenSafe: A Spectral Framework for Learning-Based Probabilistic Safety Assessment
- SISMA: Semantic Face Image Synthesis with Mamba
- Discrete-Time Diffusion-Like Models for Speech Synthesis
- STAR: Speech-to-Audio Generation via Representation Learning
- Preference Trajectory Modeling via Flow Matching for Sequential Recommendation
- JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
- ReSeFlow: Rectifying SE(3)-Equivariant Policy Learning Flows
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
- OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Efficient Rectified Flow for Image Fusion
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World Dehazing
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors
- DiffusionNFT: Online Diffusion Reinforcement with Forward Process
- Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement
- Kuramoto Orientation Diffusion Models
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- MeanFlowSE: one-step generative speech enhancement via conditional mean flow
- FlowCast-ODE: Continuous Hourly Weather Forecasting with Dynamic Flow Matching and ODE Solver
- Masked Diffusion Models as Energy Minimization
- RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing
- DiCache: Let Diffusion Model Determine Its Own Cache
- Generative Consistency Models for Estimation of Kinetic Parametric Image Posteriors in Total-Body PET
- Dense-Jump Flow Matching with Non-Uniform Time Scheduling for Robotic Policies: Mitigating Multi-Step Inference Degradation
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
- PINGS: Physics-Informed Neural Network for Fast Generative Sampling
- Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
- Chord: Chain of Rendering Decomposition for PBR Material Estimation from Generated Texture Images
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
- LADB: Latent Aligned Diffusion Bridges for Semi-Supervised Domain Translation
- Physics-Guided Rectified Flow for Low-light RAW Image Enhancement
- ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Reconstruction Alignment Improves Unified Multimodal Models
- Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- CardiacFlow: 3D+t Four-Chamber Cardiac Shape Completion and Generation via Flow Matching
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
- Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
- Transition Models: Rethinking the Generative Learning Objective
- Diffusion Generative Models Meet Compressed Sensing, with Applications to Imaging and Finance
- Scale-Adaptive Generative Flows for Multiscale Scientific Data
- A-FloPS: Accelerating Diffusion Sampling with Adaptive Flow Path Sampler
- Distribution estimation via Flow Matching with Lipschitz guarantees
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom Number
- Latent Space Single-Pixel Imaging Under Low-Sampling Conditions
- Lipschitz-Guided Design of Interpolation Schedules in Generative Models
- Delta Rectified Flow Sampling for Text-to-Image Editing
- ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
- Challenges in Non-Polymeric Crystal Structure Prediction: Why a Geometric, Permutation-Invariant Loss is Needed
- Any-Order Flexible Length Masked Diffusion
- Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
- FNODE: Flow-Matching for data-driven simulation of constrained multibody systems
- PHD: Personalized 3D Human Body Fitting with Point Diffusion
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation
- Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
- OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
- Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees
- StableIntrinsic: Detail-preserving One-step Diffusion Model for Multi-view Material Estimation
- MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Energy-Based Flow Matching for Generating 3D Molecular Structure
- SAT-SKYLINES: 3D Building Generation from Satellite Imagery and Coarse Geometric Priors
- Provable Mixed-Noise Learning with Flow-Matching
- Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
- CurveFlow: Curvature-Guided Flow Matching for Image Generation
- Source-Guided Flow Matching
- CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
- Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- OmniTry: Virtual Try-On Anything without Masks
- EventTSF: Event-Aware Non-Stationary Time Series Forecasting
- DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- Synthesizing Accurate and Realistic T1-weighted Contrast-Enhanced MR Images using Posterior-Mean Rectified Flow
- FlowMol3: Flow Matching for 3D De Novo Small-Molecule Generation
- FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
- Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models
- Constraint-Aware Flow Matching via Randomized Exploration
- Flow Matching for Efficient and Scalable Data Assimilation
- Next Visual Granularity Generation
- Distribution Matching via Generalized Consistency Models
- LoRAtorio: An intrinsic approach to LoRA Skill Composition
- Noise Matters: Optimizing Matching Noise for Diffusion Classifiers
- 3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- Flow-SLM: Joint Learning of Linguistic and Acoustic Information for Spoken Language Modeling
- Hybrid Long and Short Range Flows for Point Cloud Filtering
- Elucidating Rectified Flow with Deterministic Sampler: Polynomial Discretization Complexity for Multi and One-step Models
- Towards Safe Imitation Learning via Potential Field-Guided Flow Matching
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Score Augmentation for Diffusion Models
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- When and how can inexact generative models still sample from the data manifold?
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- PureSample: Neural Materials Learned by Sampling Microgeometry
- Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- Towards High-Order Mean Flow Generative Models: Feasibility, Expressivity, and Provably Efficient Criteria
- Elastic Diffusion Transformer
- OM2P: Offline Multi-Agent Mean-Flow Policy
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- How and Why: Taming Flow Matching for Unsupervised Anomaly Detection and Localization
- MolSnap: Snap-Fast Molecular Generation with Latent Variational Mean Flow
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
- MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Multitask Learning with Stochastic Interpolants
- LayerT2V: Interactive Multi-Object Trajectory Layering for Video Generation
- Bridging Diffusion Models and 3D Representations: A 3D Consistent Super-Resolution Framework
- SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
- REFLECT: Rectified Flows for Efficient Brain Anomaly Correction Transport
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- Qwen-Image Technical Report
- Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
- Diffusion models for inverse problems
- VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation
- Flow Matching for Probabilistic Learning of Dynamical Systems from Missing or Noisy Data
- PnP-DA: Towards Principled Plug-and-Play Integration of Variational Data Assimilation and Generative Models
- CIF: A Constrained Inversion Framework for Reliable Message Extraction in Diffusion-Based Generative Steganography
- One-Step Flow Policy Mirror Descent
- FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
- Training-free Geometric Image Editing on Diffusion Models
- Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- Next Tokens Denoising for Speech Synthesis
- Weighted Conditional Flow Matching
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- A Diffusion Model for POI Recommendation
- Why Flow Matching is Particle Swarm Optimization?
- Conditional Diffusion Models for Global Precipitation Map Inpainting
- CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers
- FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
- ReDi: Rectified Discrete Flow
- VITA: Vision-to-Action Flow Matching Policy
- A New One-Shot Federated Learning Framework for Medical Imaging Classification with Feature-Guided Rectified Flow and Knowledge Distillation
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Survey of Multimodal Hallucination Evaluation and Detection
- Flow Stochastic Segmentation Networks
- LoViC: Efficient Long Video Generation with Context Compression
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
- SegDT: A Diffusion Transformer-Based Segmentation Model for Medical Imaging
- TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
- Equivariant Volumetric Grasping
- Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models
- SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
- Qwen-Image-Flash: Beyond Objective Design
- Qwen-Image-2.0 Technical Report
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- One-step Latent-free Image Generation with Pixel Mean Flows
- Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
- Flow Matching Meets Biology and Life Science: A Survey
- SADA: Stability-guided Adaptive Diffusion Acceleration
- SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
- TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
- Adaptive Transition State Refinement with Learned Equilibrium Flows
- Hierarchical Rectified Flow Matching with Mini-Batch Couplings
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- Momentum Multi-Marginal Schrödinger Bridge Matching
- Room Impulse Response Generation Conditioned on Acoustic Parameters
- Efficient Diffusion Model for Image Restoration by Residual Shifting
- Protenix-Mini: Efficient Structure Predictor via Compact Architecture, Few-Step Diffusion and Switchable pLM
- SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling
- Minimalist Concept Erasure in Generative Models
- Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
- CharaConsist: Fine-Grained Consistent Character Generation
- Efficient Part-level 3D Object Generation via Dual Volume Packing
- Branched Schrödinger Bridge Matching
- Diffuse and Disperse: Image Generation with Representation Regularization
- Product of Experts for Visual Generation
- Recent Advances in Simulation-based Inference for Gravitational Wave Data Analysis
- Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
- Flow matching for reaction pathway generation
- Straighten Viscous Rectified Flow via Noise Optimization
- Flows and Diffusions on the Neural Manifold
- Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling
- NeuTSFlow: Modeling Continuous Functions Behind Time Series Forecasting
- Spatial Reasoners for Continuous Variables in Any Domain
- How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
- Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow
- I2I-PR: Deep Iterative Refinement for Phase Retrieval using Image-to-Image Diffusion Models
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- Intention-Conditioned Flow Occupancy Models
- BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings
- Demystifying Flux Architecture
- La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching
- Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production
- CoVAE: Consistency Training of Variational Autoencoders
- Geometric Generative Modeling with Noise-Conditioned Graph Networks
- Beyond Scores: Proximal Diffusion Models
- FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency
- When and Where do Data Poisons Attack Textual Inversion?
- Upsample What Matters: Region-Adaptive Latent Sampling for Accelerated Diffusion Transformers
- Reinforcement Learning with Action Chunking
- Dense Temporal Contrast Synthesis via Conditioned Latent Transport
- Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration
- Conditioning Tree-Based Diffusions and Flows for Probabilistic Tabular Regression
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Bridging the Last Mile of Prediction: Enhancing Time Series Forecasting with Conditional Guided Flow Matching
- Exact Evaluation of the Accuracy of Diffusion Models for Inverse Problems with Gaussian Data Distributions
- Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance
- A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
- DreamArt: Generating Interactable Articulated Objects from a Single Image
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
- FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
- MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design
- DICE: Discrete inverse continuity equation for learning population dynamics
- QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation
- StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers
- Meshy T2: Fast Native Mesh Generation with Flow Matching
- Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
- Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
- DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering
- Representations from Pretrained Machine-Learning Interatomic Potentials as Coarse Coordinates for Material Generation and Evaluation
- Flow Matching with Missing Data
- FACM: Flow-Anchored Consistency Models
- MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
- CompoSE: Compositional Synthesis and Editing of 3D Shapes via Part-Aware Control
- RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
- Learning few-step posterior samplers by unfolding and distillation of diffusion models
- P-Flow: Proxy-gradient Flows for Linear Inverse Problems
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
- FMOcc: TPV-Driven Flow Matching for 3D Occupancy Prediction with Selective State Space Model
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation
- Progressive Checkerboards for Autoregressive Multiscale Image Generation
- ActionParty: Multi-Subject Action Binding in Generative Video Games
- IC-Custom: Diverse Image Customization via In-Context Learning
- Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing
- ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
- BoltzNCE: Learning Likelihoods for Boltzmann Generation with Stochastic Interpolants and Noise Contrastive Estimation
- Posterior Inference in Latent Space for Scalable Constrained Black-box Optimization
- Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration
- Epona: Autoregressive Diffusion World Model for Autonomous Driving
- Minimally dissipative multi-bit logical operations
- TurboVSR: Fantastic Video Upscalers and Where to Find Them
- Transition Matching: Scalable and Flexible Generative Modeling
- Towards foundational LiDAR world models with efficient latent flow matching
- JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
- Pyramidal Patchification Flow for Visual Generation
- Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
- VMoBA: Mixture-of-Block Attention for Video Diffusion Models
- FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation
- FunDiff: Diffusion Models over Function Spaces for Physics-Informed Generative Modeling
- Multimodal Atmospheric Super-Resolution With Deep Generative Models
- Unfolding Generative Flows with Koopman Operators: Fast and Interpretable Sampling
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
- TADA: Improved Diffusion Sampling with Training-free Augmented Dynamics
- TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
- Rethinking Oversaturation in Classifier-Free Guidance via Low Frequency
- PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
- The Aging Multiverse: Generating Condition-Aware Facial Aging Tree via Training-Free Diffusion
- Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
- Lightweight Physics-Aware Zero-Shot Ultrasound Plane-Wave Denoising
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- On Convolutions, Intrinsic Dimension, and Diffusion Models
- Telegrapher's Generative Model via Kac Flows
- Ctrl-Z Sampling: Diffusion Sampling with Controlled Random Zigzag Explorations
- Improving Progressive Generation with Decomposable Flow Matching
- ProxelGen: Generating Proteins as 3D Densities
- Real-Time Execution of Action Chunking Flow Policies
- Operator Forces For Coarse-Grained Molecular Dynamics
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Statistical Inference for Optimal Transport Maps: Recent Advances and Perspectives
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
- RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
- Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction
- Efficient many-jet event generation with Flow Matching
- Part2GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
- Schrödinger Bridge Matching for Tree-Structured Costs and Entropic Wasserstein Barycentres
- Reversing Flow for Image Restoration
- Fast and Stable Diffusion Planning through Variational Adaptive Weighting
- RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
- FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
- MoiréXNet: Adaptive Multi-Scale Demoiréing with Linear Attention Test-Time Training and Truncated Flow Matching Prior
- 4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
- Improving Rectified Flow with Boundary Conditions
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- Sampling 3D Molecular Conformers with Diffusion Transformers
- Time-dependent density estimation using binary classifiers
- Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards
- Show-o2: Improved Native Unified Multimodal Models
- Unsupervised Imaging Inverse Problems with Diffusion Distribution Matching
- PairEdit: Learning Semantic Variations for Exemplar-based Image Editing
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Align Your Flow: Scaling Continuous-Time Flow Map Distillation
- Bures-Wasserstein Flow Matching for Graph Generation
- Flexible-length Text Infilling for Discrete Diffusion Models
- Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
- Flow-Based Policy for Online Reinforcement Learning
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- Learning to Integrate
- Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
- Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
- Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
- DiffPR: Diffusion-Based Phase Reconstruction via Frequency-Decoupled Learning
- Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
- Geometric Regularity in Deterministic Sampling of Diffusion-based Generative Models
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- Flow-GRPO: Training Flow Matching Models via Online RL
- BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning
- FontAdapter: Instant Font Adaptation in Visual Text Generation
- Contrastive Flow Matching
- Rectified Point Flow: Generic Point Cloud Pose Estimation
- Aligning Latent Spaces with Flow Priors
- PixCell: A generative foundation model for digital histopathology images
- Diffusion Model Quantization: A Review
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
- SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
- On Fitting Flow Models with Large Sinkhorn Couplings
- FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
- Multi-turn Consistent Image Editing
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation
- Efficient Flow Matching using Latent Variables
- Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
- Solving Inverse Problems via Diffusion-Based Priors: An Approximation-Free Ensemble Sampling Approach
- Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- Large deviations for scaled families of Schrödinger bridges with reflection
- Horizon Reduction Makes RL Scalable
- Physics-Constrained Flow Matching: Sampling Generative Models with Hard Constraints
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
- RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
- Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
- Rectified Flows for Fast Multiscale Fluid Flow Modeling
- Interaction Field Matching: Overcoming Limitations of Electrostatic Models
- LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
- Solving Inverse Problems with FLAIR
- Native-Resolution Image Synthesis
- Grasp2Grasp: Vision-Based Dexterous Grasp Translation via Schrödinger Bridges
- Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
- Chipmunk: Training-Free Acceleration of Diffusion Transformers with Dynamic Column-Sparse Deltas
- UniConFlow: A Unified Constrained Flow-Matching Framework for Certified Motion Planning
- CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
- SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples
- Dual-Process Image Generation
- Synthesis of discrete-continuous quantum circuits with multimodal diffusion models
- Learning of Population Dynamics: Inverse Optimization Meets JKO Scheme
- Playing with Transformer at 30+ FPS via Next-Frame Diffusion
- Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
- Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
- Psi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models
- Latent Stochastic Interpolants
- DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
- Efficient Regression-Based Training of Normalizing Flows for Boltzmann Generators
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
- FourierFlow: Frequency-aware Flow Matching for Generative Turbulence Modeling
- Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space
- Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video Generation
- Graph Flow Matching: Enhancing Image Generation with Neighbor-Aware Flow Fields
- On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning
- Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
- IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
- PDE-Transformer: Efficient and Versatile Transformers for Physics Simulations
- MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
- D-AR: Diffusion via Autoregressive Models
- Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
- FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
- Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
- Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching
- Diffusion Guidance Is a Controllable Policy Improvement Operator
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- Trajectory Generator Matching for Time Series
- Optimization-Free Diffusion Model -- A Perturbation Theory Approach
- ACE-Step: A Step Towards Music Generation Foundation Model
- What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?
- Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
- AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
- FNOPE: Simulation-based inference on function spaces with Fourier Neural Operators
- Scaling Offline RL via Efficient and Expressive Shortcut Models
- ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
- PropMolFlow: Property-Guided Molecule Generation with Geometry-Complete Flow Matching
- Causal Posterior Estimation
- T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models
- Differentiable Solver Search for Fast Diffusion Sampling
- Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
- LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
- Advancing high-fidelity 3D and Texture Generation with 2.5D latents
- Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- Extremum Flow Matching for Offline Goal Conditioned Reinforcement Learning
- Training-Free Multi-Step Audio Source Separation
- Understanding Generalization in Diffusion Distillation via Probability Flow Distance
- Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
- On the Relation between Rectified Flows and Optimal Transport
- Absolute Coordinates Make Motion Generation Easy
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy
- DiSA: Diffusion Step Annealing in Autoregressive Image Generation
- PHI: Bridging Domain Shift in Long-Term Action Quality Assessment via Progressive Hierarchical Instruction
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
- Scalable Generation of Spatial Transcriptomics from Histology Images via Whole-Slide Flow Matching
- Fast Kernel-Space Diffusion for Remote Sensing Pansharpening
- GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
- STFlow: Data-Coupled Flow Matching for Geometric Trajectory Simulation
- MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
- How to build a consistency model: Learning flow maps via self-distillation
- Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
- VORTA: Efficient Video Diffusion via Routing Sparse Attention
- LiveLight: Real-time Streaming Video Relighting with Interactive Control
- ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
- Diffusion Classifiers Understand Compositionality, but Conditions Apply
- CT-OT Flow: Estimating Continuous-Time Dynamics from Discrete Temporal Snapshots
- A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
- T2VUnlearning: A Concept Erasing Method for Text-to-Video Diffusion Models
- Flexible MOF Generation with Torsion-Aware Flow Matching
- Training-Free Efficient Video Generation via Dynamic Token Carving
- Forward-only Diffusion Probabilistic Models
- Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts
- Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
- DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
- Flow Matching based Sequential Recommender Model
- Conditional Panoramic Image Generation via Masked Autoregressive Modeling
- Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow
- Music Restoration via Latent Operator Optimization and Diffusion Model Priors
- Neural Conditional Transport Maps
- Beckmann Transport Models: From Autonomous Flows to One-Step Maps
- A Unified Kullback--Leibler Divergence Analysis of Generative Diffusion Models via Entropy Production Rate
- SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching
- CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
- HaptoFlow: High-Fidelity Real-Time Vibrotactile Generation via Flow Matching for Virtual Reality
- VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
- Hermite Curves as Trajectory Priors for Vision-Language-Action Models
- TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
- Emerging Properties in Unified Multimodal Pretraining
- Riemannian Flow Matching for Brain Connectivity Matrices via Pullback Geometry
- Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
- Learning to Integrate Diffusion ODEs by Averaging the Derivatives
- PhySense: Sensor Placement Optimization for Accurate Physics Sensing
- CacheFlow: Fast Human Motion Prediction by Cached Normalizing Flow
- OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
- Score-Based Training for Energy-Based TTS Models
- Joint Velocity-Growth Flow Matching for Single-Cell Dynamics Modeling
- FlowPure: Continuous Normalizing Flows for Adversarial Purification
- Few-Step Diffusion via Score identity Distillation
- VSA: Faster Video Diffusion with Trainable Sparse Attention
- Mean Flows for One-step Generative Modeling
- Sampling NNLO QCD phase space with normalizing flows
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- Variational Regularized Unbalanced Optimal Transport: Single Network, Least Action
- MiniWorld: Democratizing the Training of Video World Models from Scratch
- ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport
- Conditioning Matters: Training Diffusion Policies is Faster Than You Think
- CN101 - A Digital Thermodynamic Computer for Generative AI
- Where Does Generative Difficulty Reside? An Empirical Study of Target Representations
- Hybrid-Domain Posterior Sampling for Inverse Problems via Latent Flow Matching
- Missing Data Imputation by Reducing Mutual Information with Rectified Flows
- Towards Self-Improvement of Diffusion Models via Group Preference Optimization
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
- Schrödinger Generator for High-Dimensional Integration and Sampling on Quantum Many-Body States
- Path Gradients after Flow Matching
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
- Whitened Score Diffusion: A Structured Prior for Imaging Inverse Problems
- Don't Forget your Inverse DDIM for Image Editing
- Heterogeneous Decentralized Diffusion Models
- LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
- Noise-Robust Conditional Flow Matching: Generating Clean Samples from Noisy Datasets
- Fast Text-to-Audio Generation with Adversarial Post-Training
- Distilling Drifting Transformers with Representation Autoencoders
- 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
- Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
- DanceGRPO: Unleashing GRPO on Visual Generation
- Generative Pre-trained Autoregressive Diffusion Transformer
- Addressing degeneracies in latent interpolation for diffusion models
- Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
- H3DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
- You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling with Gradient Shortcuts
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Climate in a Bottle: Towards a Generative Foundation Model for the Kilometer-Scale Global Atmosphere
- Good Things Come in Pairs: Paired Autoencoders for Inverse Problems
- MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
- ViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- SARe: Structure-Aware Generative 3D Fragment Reassembly
- A Survey on Generative Diffusion Models
- OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction
- Video Generation Models are General-Purpose Vision Learners
- Distilling Two-Timed Flow Models by Separately Matching Initial and Terminal Velocities
- Fast Flow-based Visuomotor Policies via Conditional Optimal Transport Couplings
- Exact Posterior Score Estimation for Solving Linear Inverse Problems
- Discrete Flow Maps
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- Learning Hamiltonian Flow Maps: Mean Flow Consistency for Large-Timestep Molecular Dynamics
- It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
- Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
- CytoSyn: a Foundation Diffusion Model for Histopathology -- Tech Report
- SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
- Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
- The Role of Artificial Intelligence in the SKA Era
- TeLoGraF: Temporal Logic Planning via Graph-encoded Flow Matching
- Adaptive Protein Tokenization
- Fractals made Practical: Denoising Diffusion as Partitioned Iterated Function Systems
- VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
- The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
- AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
- dots.tts Technical Report
- Distribution-Conditioned Transport
- Solaris: Building a Multiplayer Video World Model in Minecraft
- Efficient and robust 3D blind harmonization for large domain gaps
- A Survey of Interactive Generative Video
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
- Unpaired Image-to-Image Translation via a Self-Supervised Semantic Bridge
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- World Action Models are Zero-shot Policies
- Image Generation with a Sphere Encoder
- Demystifying Multimodal Biomolecular Co-design With Intrinsic Geodesic Coupling
- Strong Stochastic Flow Maps
- BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models
- Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
- Zero-Flow Encoders
- Causal World Modeling for Robot Control
- ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
- Follow the Mean: Reference-Guided Flow Matching
- Syn4D: A Multiview Synthetic 4D Dataset
- Categorical Flow Maps
- ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
- ADiff4TPP: Asynchronous Diffusion Models for Temporal Point Processes
- In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
- FastFlow: Accelerating The Generative Flow Matching Models with Bandit Inference
- DanceOPD: On-Policy Generative Field Distillation
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
- Towards Universal Spatial Transcriptomics Super-Resolution: A Generalist Physically Consistent Flow Matching Framework
- Annealed Langevin Monte Carlo for Flow ODE Sampling
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Flow Sampling: Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
- Neural Posterior Estimation of Terrain Parameters from Radar Sounder Data
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method
- MolCrystalFlow: Molecular Crystal Structure Prediction via Flow Matching
- Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps
- Sparsely Supervised Diffusion
- Learning Brenier Potentials with Convex Generative Adversarial Neural Networks
- Integration Flow Models
- Flow Along the K-Amplitude for Generative Modeling
- Learning to Drive from a World Model
- Versatile Framework for Song Generation with Prompt-based Control
- Continuous Adversarial Flow Models
- High-dimensional reliability-based design optimization using stochastic emulators
- DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
- FaithIR: Rethinking Infrared Image Super-Resolution from Perceptual Sharpness to Task Relevant Fidelity
- Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
- RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
- PFM-HR: Pose Flow Matching for Humanoid Robots
- Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
- GraspMeanFlow: SE(3)-Equivariant MeanFlow for Few-Step 6-DoF Grasp Generation
- EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning
- Embedding Empirical Distributions for Computing Optimal Transport Maps
- Dual Prompting Image Restoration with Diffusion Transformers
- Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
- RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models
- Fast Autoregressive Models for Continuous Latent Generation
- When does training on downscaled images yield the same gradients?
- Discretization and Statistical Consistency of Functional Flow Matching
- CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
- OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
- Learning Energy-Based Generative Models via Potential Flow: A Variational Principle Approach to Probability Density Homotopy Matching
- Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
- Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
- HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
- MusFlow: Multimodal Music Generation via Conditional Flow Matching
- U-Shape Mamba: State Space Model for faster diffusion
- Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
- On the minimax optimality of Flow Matching through the connection to kernel density estimation
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
- Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
- A Bidirectional DeepParticle Method for Efficiently Solving Low-dimensional Transport Map Problems
- Energy-Guided Flow Matching
- Flow-Map Distillation on Relation Manifolds for Image Restoration
- LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
- Vorch-Omni: Multi-Task Orchestration of Sight and Sound
- Hierarchical Flow Matching for 3D Point Cloud Generation
- Rethinking Automatic Music Mixing as Sequential Stem Blending
- EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry
- RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- ProtFlow: Fast Protein Sequence Design via Flow Matching on Compressed Protein Language Model Embeddings
- Autoregressive Distillation of Diffusion Transformers
- ADT: Tuning Diffusion Models with Adversarial Supervision
- Omni2: Unifying Omnidirectional Image Generation and Editing in an Omni Model
- On the Contractivity of Stochastic Interpolation Flow
- FLOWR: Flow Matching for Structure-Aware De Novo, Interaction- and Fragment-Based Ligand Generation
- Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers
- Efficient Generative Model Training via Embedded Representation Warmup
- Energy Matching: Unifying Flow Matching and Energy-Based Models for Generative Modeling
- H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
- Flux Already Knows -- Activating Subject-Driven Image Generation without Training
- PixelFlow: Pixel-Space Generative Models with Flow
- VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
- HoloPart: Generative 3D Part Amodal Segmentation
- DiverseFlow: Sample-Efficient Diverse Mode Coverage in Flows
- SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
- EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation
- PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering
- HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
- GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
- Gaussian Mixture Flow Matching Models
- Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
- TabRep: Training Tabular Diffusion Models with a Simple and Effective Continuous Representation
- Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
Discussions
Related