Taming Transformers for High-Resolution Image Synthesis
2020/12/17 by Patrick Esser, Robin Rombach, Esser, Patrick +3 · 2 voices · 716 citations
Computer Science · #cs.CV
paper · pdf · doi:10.48550/arxiv.2012.09841
Changelog can be found in the supplementary
arxiv created 2021/06/23 · arxiv updated 2021/06/24
Abstract
Designed to learn long-range interactions on sequential data, transformers continue to show state-of-the-art results on a wide variety of tasks. In contrast to CNNs, they contain no inductive bias that prioritizes local interactions. This makes them expressive, but also computationally infeasible for long sequences, such as high-resolution images. We demonstrate how combining the effectiveness of the inductive bias of CNNs with the expressivity of transformers enables them to model and thereby synthesize high-resolution images. We show how to (i) use CNNs to learn a context-rich vocabulary of image constituents, and in turn (ii) utilize transformers to efficiently model their composition within high-resolution images. Our approach is readily applied to conditional synthesis tasks, where both non-spatial information, such as object classes, and spatial information, such as segmentations, can control the generated image. In particular, we present the first results on semantically-guided synthesis of megapixel images with transformers and obtain the state of the art among autoregressive models on class-conditional ImageNet. Code and pretrained models can be found at https://github.com/CompVis/taming-transformers .
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
- Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
- GeoDiff-SAR: a geometric prior guided diffusion model for SAR image generation
- Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- DriftXpress: Faster Drifting Models via Projected RKHS Fields
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- TaskTok: Delving into Task Tokens for Task-driven Image Restoration
- Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
- InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
- Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations
- Orbis 2: A Hierarchical World Model for Driving
- Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
- Uni-AdaVD: Universal Concept Erasure for Visual Generation via Orthogonal Value Decomposition
- Introspective Attention Modulation for Safe Text-to-Image Generation
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Allure of Craquelure: A Variational-Generative Approach to Crack Detection in Paintings
- Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
- The Market in the Model: Latent Diffusion as Neural Economy
- One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
- Continuous Autoregressive Language Models
- Towards a Physics Foundation Model
- World Modeling with Probabilistic Structure Integration
- From basic affordances to symbolic thought: A computational phylogenesis of biological intelligence.
- WorldVLA: Towards Autoregressive Action World Model
- OmniSVG: A Unified Scalable Vector Graphics Generation Model
- VGGT: Visual Geometry Grounded Transformer
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
- Distribution Matching Variational AutoEncoder
- ThinkGen: Generalized Thinking for Visual Generation
- RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
- A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
- Generative Latent Coding for Ultra-Low Bitrate Image Compression
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Local Patches Meet Global Context: Scalable 3D Diffusion Priors for Computed Tomography Reconstruction
- PSI3D: Plug-and-Play 3D Stochastic Inference with Slice-wise Latent Diffusion Prior
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- Spherical Leech Quantization for Visual Tokenization and Generation
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- RecTok: Reconstruction Distillation along Rectified Flow
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Multi-temporal Calving Front Segmentation
- Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
- AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- Voxify3D: Pixel Art Meets Volumetric Rendering
- Training-Free Vector Quantization via Gaussian VAEs
- See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
- HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- Efficient Generative Transformer Operators For Million-Point PDEs
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- DeRA: Decoupled Representation Alignment for Video Tokenization
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Rethinking Security in Semantic Communication: Latent Manipulation as a New Threat
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
- PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation
- Visual Generation Tuning
- Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- Bringing Your Portrait to 3D Presence
- Adversarial Flow Models
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- The Collapse of Patches
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Latent Diffusion Inversion Requires Understanding the Latent Space
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Flow Map Distillation Without Data
- Understanding, Accelerating, and Improving MeanFlow Training
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- CoD: A Diffusion Foundation Model for Image Compression
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
- Spanning Tree Autoregressive Visual Generation
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- Flow and Depth Assisted Video Prediction with Latent Transformer
- Decoupling Complexity from Scale in Latent Diffusion Model
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- WiCo-PG: Wireless Channel Foundation Model for Pathloss Map Generation via Synesthesia of Machines
- WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- MixAR: Mixture Autoregressive Image Generation
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- Point Cloud Quantization through Multimodal Prompting for 3D Understanding
- A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space
- FlowCast: Advancing Precipitation Nowcasting with Conditional Flow Matching
- FedeCouple: Fine-Grained Balancing of Global-Generalization and Local-Adaptability in Federated Learning
- TransactionGPT
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- Retrospective motion correction in MRI using disentangled embeddings
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- PADM: A Physics-aware Diffusion Model for Attenuation Correction
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
- PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
- MALeR: Improving Compositional Fidelity in Layout-Guided Generation
- CPO: Condition Preference Optimization for Controllable Image Generation
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
- Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
- DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- BI-DCGAN: A Theoretically Grounded Bayesian Framework for Efficient and Diverse GANs
- Emu3.5: Native Multimodal Models are World Learners
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Product-Quantised Image Representation for High-Quality Image Synthesis
- Amortized Moment Matching for Visual Generation
- Generative Modeling via Drifting
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
- Nested AutoRegressive Models
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- Morphologically Intelligent Perturbation Prediction with FORM
- Pctx: Tokenizing Personalized Context for Generative Recommendation
- Improved Training Technique for Shortcut Models
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Exploring Conditions for Diffusion models in Robotic Control
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Vector Quantization in the Brain: Grid-like Codes in World Models
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- End-to-End Multi-Modal Diffusion Mamba
- NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- Universal Image Restoration Pre-training via Masked Degradation Classification
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- What If : Understanding Motion Through Sparse Interactions
- BIGFix: Bidirectional Image Generation with Token Fixing
- Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Diffusion Transformers with Representation Autoencoders
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- Generative Latent Video Compression
- Variational Secret Common Randomness Extraction
- MelTok: 2D Tokenization for Single-Codebook Audio Compression
- Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Optimal Stopping in Latent Diffusion Models
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- CodeFormer++: Blind Face Restoration Using Deformable Registration and Deep Metric Learning
- Bridging Text and Video Generation: A Survey
- Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Flow Autoencoders are Effective Protein Tokenizers
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
- PUREVQ-GAN: Defending Data Poisoning Attacks through Vector-Quantized Bottlenecks
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Score-based Membership Inference on Diffusion Models
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- Environment-Aware Satellite Image Generation with Diffusion Models
- Real-Aware Residual Model Merging for Deepfake Detection
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Tumor Synthesis conditioned on Radiomics
- Scalable GANs with Transformers
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Object-AVEdit: An Object-level Audio-Visual Editing Model
- Entering the Era of Discrete Diffusion Models: A Benchmark for Schrödinger Bridges and Entropic Optimal Transport
- Stochastic Interpolants via Conditional Dependent Coupling
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
- Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
- CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Codebook-Based Adaptive Feature Compression With Semantic Enhancement for Edge-Cloud Systems
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Learning Dexterous Manipulation with Quantized Hand State
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Efficient Rectified Flow for Image Fusion
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- AToken: A Unified Tokenizer for Vision
- Image Tokenizer Needs Post-Training
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- Reconstruction Alignment Improves Unified Multimodal Models
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- PRIM: Towards Practical In-Image Multilingual Machine Translation
- SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
- Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Acoustic Interference Suppression in Ultrasound images for Real-Time HIFU Monitoring Using an Image-Based Latent Diffusion Model
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
- Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- Taming Transformer for Emotion-Controllable Talking Face Generation
- Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
- 2D Gaussians Meet Visual Tokenizer
- Next Visual Granularity Generation
- Versatile Video Representation via Feed-Forward 2D Gaussian Splatting Tokenization
- Semi-supervised Image Dehazing via Expectation-Maximization and Bidirectional Brownian Bridge Diffusion Models
- Can Synthetic Images Conquer Forgetting? Beyond Unexplored Doubts in Few-Shot Class-Incremental Learning
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Prototype-Guided Diffusion: Visual Conditioning without External Memory
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- Vision Generalist Model: A Survey
- Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UniTalker: Conversational Speech-Visual Synthesis
- Spectral Efficiency-Aware Codebook Design for Task-Oriented Semantic Communications
- Deeper Inside Deep ViT
- Cross-Domain Image Synthesis: Generating H&E from Multiplex Biomarker Imaging
- HPSv3: Towards Wide-Spectrum Human Preference Score
- GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
- CIVQLLIE: Causal Intervention with Vector Quantization for Low-Light Image Enhancement
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
- StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Subtyping Breast Lesions via Generative Augmentation based Long-tailed Recognition in Ultrasound
- GVD: Guiding Video Diffusion Model for Scalable Video Distillation
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- HDR Environment Map Estimation with Latent Diffusion Models
- Kernel Learning for Sample Constrained Black-Box Optimization
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- Local Prompt Adaptation for Style-Consistent Multi-Object Generation in Diffusion Models
- KB-DMGen: Knowledge-Based Global Guidance and Dynamic Pose Masking for Human Image Generation
- RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot Learning
- ReDi: Rectified Discrete Flow
- Latent Denoising Makes Good Tokenizers
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- Reconstruct or Generate: Exploring the Spectrum of Generative Modeling for Cardiac MRI
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Survey of Multimodal Hallucination Evaluation and Detection
- Even Faster Simulations with Flow Matching: A Study of Zero Degree Calorimeter Responses
- Quantizing Text-attributed Graphs for Semantic-Structural Integration
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- DynFaceRestore: Balancing Fidelity and Quality in Diffusion-Guided Blind Face Restoration with Dynamic Blur-Level Mapping and Guidance
- Vec2Face+ for Face Dataset Generation
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Test-time Prompt Refinement for Text-to-Image Models
- HarmonPaint: Harmonized Training-Free Diffusion Inpainting
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
- Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation
- CharaConsist: Fine-Grained Consistent Character Generation
- Implementing Adaptations for Vision AutoRegressive Model
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Diffusion-Based Imaginative Coordination for Bimanual Manipulation
- Latent Space Consistency for Sparse-View CT Reconstruction
- Quantize-then-Rectify: Efficient VQ-VAE Training
- Latent Diffusion Models with Masked AutoEncoders
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
- RefSTAR: Blind Facial Image Restoration with Reference Selection, Transfer, and Reconstruction
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
- RectifiedHR: High-Resolution Diffusion via Energy Profiling and Adaptive Guidance Scheduling
- DALI-PD: Diffusion-based Synthetic Layout Heatmap Generation for ML in Physical Design
- AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
- I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
- CoVAE: Consistency Training of Variational Autoencoders
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- Single-pass Adaptive Image Tokenization for Minimum Program Search
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- Denoising Multi-Beta VAE: Representation Learning for Disentanglement and Generation
- FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents
- LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models
- Text-Guided Token Communication for Wireless Image Transmission
- Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image Restoration
- DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
- ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Sensing
- ICAS: Detecting Training Data from Autoregressive Image Generative Models
- MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
- A Training-Free Style-Personalization via SVD-Based Feature Decomposition
- Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
- Accurate and Efficient World Modeling with Masked Latent Transformers
- RefTok: Reference-Based Tokenization for Video Generation
- UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
- CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation
- LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection
- Progressive Checkerboards for Autoregressive Multiscale Image Generation
- SSL4SAR: Self-Supervised Learning for Glacier Calving Front Extraction from SAR Imagery
- Enhancing Multi-Exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning
- DiffMark: Diffusion-based Robust Watermark Against Deepfakes
- Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
- Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
- BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
- Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration
- Are Large Brainwave Foundation Models Capable Yet? Insights from Fine-tuning
- Towards 3D Semantic Image Synthesis for Medical Imaging
- Epona: Autoregressive Diffusion World Model for Autonomous Driving
- MotionGPT3: Human Motion as a Second Modality
- Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned Diffusion
- Transition Matching: Scalable and Flexible Generative Modeling
- CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
- Learning Counterfactually Decoupled Attention for Open-World Model Attribution
- Riemannian-Geometric Fingerprints of Generative Models
- StableCodec: Taming One-Step Diffusion for Extreme Image Compression
- CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
- Stochastic and Non-local Closure Modeling for Nonlinear Dynamical Systems via Latent Score-based Generative Models
- Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Improving Progressive Generation with Decomposable Flow Matching
- Curating art exhibitions using machine learning
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- Style Transfer: A Decade Survey
- LeVo: High-Quality Song Generation with Multi-Preference Alignment
- AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- Let Your Video Listen to Your Music!
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
- Auto-Regressively Generating Multi-View Consistent Images
- Highly Compressed Tokenizer Can Generate Without Training
- VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
- DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Adapting Vision-Language Models for Evaluating World Models
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
- GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
- Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
- Deep generative models as the probability transformation functions
- Generative Modeling of Weights: Generalization or Memorization?
- Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- Prmpt2Adpt: Prompt-Based Zero-Shot Domain Adaptation for Resource-Constrained Environments
- Single-step Diffusion for Image Compression at Ultra-Low Bitrates
- Watermarking Autoregressive Image Generation
- Audio-Sync Video Generation with Multi-Stream Temporal Control
- Aligning Text, Images, and 3D Structure Token-by-Token
- Privacy-Preserving Chest X-ray Classification in Latent Space with Homomorphically Encrypted Neural Inference
- Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
- Discrete JEPA: Learning Discrete Token Representations without Reconstruction
- VideoMAR: Autoregressive Video Generatio with Continuous Tokens
- Risk Estimation of Knee Osteoarthritis Progression via Predictive Multi-task Modelling from Efficient Diffusion Model using X-ray Images
- Xray2Xray: World Model from Chest X-rays with Volumetric Context
- Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT
- Align Your Flow: Scaling Continuous-Time Flow Map Distillation
- Seismic Acoustic Impedance Inversion Framework Based on Conditional Latent Generative Diffusion Model
- DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization
- Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
- A Watermark for Auto-Regressive Image Generation Models
- Task-Driven Discrete Representation Learning
- ViSAGe: Video-to-Spatial Audio Generation
- SpectralAR: Spectral Autoregressive Visual Generation
- Unsupervised Deformable Image Registration with Structural Nonparametric Smoothing
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
- HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
- RecGPT: A Foundation Model for Sequential Recommendation
- Aligning Latent Spaces with Flow Priors
- Multi-scale Image Super Resolution with a Single Auto-Regressive Model
- Gen-n-Val: Agentic Image Data Generation and Validation
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
- Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- OWT: A Foundational Organ-Wise Tokenization Framework for Medical Imaging
- Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation
- AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data
- PrimitiveAnything: Human-Crafted 3D Primitive Assembly Generation with Auto-Regressive Transformer
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
- One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
- EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
- Hyperspectral Image Generation with Unmixing Guided Diffusion Model
- Playing with Transformer at 30+ FPS via Next-Frame Diffusion
- Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- Data Pruning by Information Maximization
- Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
- Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity
- Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
- Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack
- SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
- Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
- Real-Time Person Image Synthesis Using a Flow Matching Model
- Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining
- Optimizing Sensory Neurons: Nonlinear Attention Mechanisms for Accelerated Convergence in Permutation-Invariant Neural Networks for Reinforcement Learning
- Concept-Centric Token Interpretation for Vector-Quantized Generative Models
- DLM-One: Diffusion Language Models for One-Step Sequence Generation
- D-AR: Diffusion via Autoregressive Models
- DarkDiff: Advancing Low-Light Raw Enhancement by Retasking Diffusion Models for Camera ISP
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- RobSurv: Vector Quantization-Based Multi-Modal Learning for Robust Cancer Survival Prediction
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
- Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
- ReDDiT: Rehashing Noise for Discrete Visual Generation
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning
- MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- DiSA: Diffusion Step Annealing in Autoregressive Image Generation
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- Tokenizing Electron Cloud in Protein-Ligand Interaction Learning
- Plug-and-Play Context Feature Reuse for Efficient Masked Generation
- MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
- Partition Generative Modeling: Masked Modeling Without Masks
- Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
- RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
- PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning
- Deeper Diffusion Models Amplify Bias
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning
- Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- High-Fidelity Functional Ultrasound Reconstruction via A Visual Auto-Regressive Framework
- Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
- Creatively Upscaling Images with Global-Regional Priors
- Training-Free Efficient Video Generation via Dynamic Token Carving
- Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
- Advancing Brainwave Modeling with a Codebook-Based Foundation Model
- MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
- One-Step Diffusion-Based Image Compression with Semantic Distillation
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- ChemMLLM: Chemical Multimodal Large Language Model
- MMaDA: Multimodal Large Diffusion Language Models
- Exploring In-Image Machine Translation with Real-World Background
- Interspatial Attention for Efficient 4D Human Video Generation
- Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
- MSDformer: Multi-scale Discrete Transformer For Time Series Generation
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- RLVR-World: Training World Models with Reinforcement Learning
- Visual Instruction Bottleneck Tuning
- Learning to Integrate Diffusion ODEs by Averaging the Derivatives
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
- Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio
- GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
- FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance
- Mean Flows for One-step Generative Modeling
- MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
- Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy
- Visual Planning: Let's Think Only with Images
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
- Recent Advances in Diffusion Models for Hyperspectral Image Processing and Analysis: A Review
- Where Does Generative Difficulty Reside? An Empirical Study of Target Representations
- Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
- MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation
- ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access
- VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
- Multi-Token Prediction Needs Registers
- Extract the Best, Discard the Rest: CSI Feedback with Offline Large AI Models
- A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
- Continuous Visual Autoregressive Generation via Score Maximization
- H3DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Image Classification Using a Diffusion Model as a Pre-Training Model
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement
- Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
- D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- One Scale at a Time: Scale-Autoregressive Modeling for Fluid Flow Distributions
- Adaptive Protein Tokenization
- Tailor Made Embeddings for Quantum Machine Learning
- SOM-VQ: Topology-Aware Tokenization for Interactive Generative Models
- VLANeXt: Recipes for Building Strong VLA Models
- Cryo-SWAN: the Multi-Scale Wavelet-decomposition-inspired Autoencoder Network for molecular density representation of molecular volumes
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
- The Design Space of Tri-Modal Masked Diffusion Models
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- GPIC: A Giant Permissive Image Corpus for Visual Generation
- Revisiting Diffusion Autoencoder Training for Image Reconstruction Quality
- Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention into Convolutions
- Why Compress What You Can Generate? When GPT-4o Generation Ushers in Image Compression Fields
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
- Large Language Models are Universal Reasoners for Visual Generation
- GarmentX: Autoregressive Parametric Representations for High-Fidelity 3D Garment Generation
- AI-GenBench: A New Ongoing Benchmark for AI-Generated Image Detection
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
- On the Value of Tokeniser Pretraining in Physics Foundation Models
- EarthMapper: Visual Autoregressive Models for Controllable Bidirectional Satellite-Map Translation
- RadioFormer: A Multiple-Granularity Radio Map Estimation Transformer with 1\textpertenthousand Spatial Sampling
- REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
- E-InMeMo: Enhanced Prompting for Visual In-Context Learning
- Latent-Compressed Variational Autoencoder for Video Diffusion Models
- CalM: A Self-Supervised Foundation Model for Population Dynamics in Calcium Imaging Data
- DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
- T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
- Dual Prompting Image Restoration with Diffusion Transformers
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
- MV-Crafter: An Intelligent System for Music-guided Video Generation
- Fast Autoregressive Models for Continuous Latent Generation
- Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution
- Time-adaptive Video Frame Interpolation based on Residual Diffusion
- Latent Denoising Improves Visual Alignment in Large Multimodal Models
- Distilling Specialized Orders for Visual Generation
- Hyper-Transforming Latent Diffusion Models
- FluentLip: A Phonemes-Based Two-stage Approach for Audio-Driven Lip Synthesis with Optical Flow Consistency
- Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
- Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
- TVC: Tokenized Video Compression with Ultra-Low Bit Rate
- MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
- Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- Image Editing with Diffusion Models: A Survey
- KVAE: Family of Tokenizers for Multimodal Generative Models
- Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications
- MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations
- Imagine How To Change: Explicit Procedure Modeling for Change Captioning
- Autoregressive Distillation of Diffusion Transformers
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
- InstructEngine: Instruction-driven Text-to-Image Alignment
- Beyond the Generative Learning Trilemma: Generative Model Assessment in Data Scarcity Domains
- D2iT: Dynamic Diffusion Transformer for Accurate Image Generation
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- Latent Diffusion Autoencoders: Toward Efficient and Meaningful Unsupervised Representation Learning in Medical Imaging
- Diffusion Models for Robotic Manipulation: A Survey
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
- PixelFlow: Pixel-Space Generative Models with Flow
- Scaling Laws for Native Multimodal Models
- Model Discrepancy Learning: Synthetic Faces Detection Based on Multi-Reconstruction
- PRISM-36K: A Benchmark Dataset for AI-Generated Image Attribution
- CamC2V: Context-aware Controllable Video Generation
- Towards Efficient Real-Time Video Motion Transfer via Generative Time Series Modeling
- Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
- Generative Adversarial Networks with Limited Data: A Survey and Benchmarking
- 3D Scene Understanding Through Local Random Access Sequence Modeling
Discussions
Related