Taming Transformers for High-Resolution Image Synthesis
2020/12/17 by Patrick Esser, Robin Rombach, Esser, Patrick +3 · 2 voices · 373 citations
#cs.CV
paper · pdf · doi:10.48550/arxiv.2012.09841
Abstract
Designed to learn long-range interactions on sequential data, transformers continue to show state-of-the-art results on a wide variety of tasks. In contrast to CNNs, they contain no inductive bias that prioritizes local interactions. This makes them expressive, but also computationally infeasible for long sequences, such as high-resolution images. We demonstrate how combining the effectiveness of the inductive bias of CNNs with the expressivity of transformers enables them to model and thereby synthesize high-resolution images. We show how to (i) use CNNs to learn a context-rich vocabulary of image constituents, and in turn (ii) utilize transformers to efficiently model their composition within high-resolution images. Our approach is readily applied to conditional synthesis tasks, where both non-spatial information, such as object classes, and spatial information, such as segmentations, can control the generated image. In particular, we present the first results on semantically-guided synthesis of megapixel images with transformers and obtain the state of the art among autoregressive models on class-conditional ImageNet. Code and pretrained models can be found at https://github.com/CompVis/taming-transformers .
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
- GeoDiff-SAR: A Geometric Prior Guided Diffusion Model for SAR Image Generation
- Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- DriftXpress: Faster Drifting Models via Projected RKHS Fields
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- TaskTok: Delving into Task Tokens for Task-driven Image Restoration
- Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
- InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
- Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations
- Orbis 2: A Hierarchical World Model for Driving
- Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
- Uni-AdaVD: Universal Concept Erasure for Visual Generation via Orthogonal Value Decomposition
- Introspective Attention Modulation for Safe Text-to-Image Generation
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Allure of Craquelure: A Variational-Generative Approach to Crack Detection in Paintings
- Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
- The Market in the Model: Latent Diffusion as Neural Economy
- One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
- Continuous Autoregressive Language Models
- Towards a Physics Foundation Model
- World Modeling with Probabilistic Structure Integration
- From basic affordances to symbolic thought: A computational phylogenesis of biological intelligence.
- WorldVLA: Towards Autoregressive Action World Model
- OmniSVG: A Unified Scalable Vector Graphics Generation Model
- VGGT: Visual Geometry Grounded Transformer
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
- Distribution Matching Variational AutoEncoder
- ThinkGen: Generalized Thinking for Visual Generation
- RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
- A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
- Generative Latent Coding for Ultra-Low Bitrate Image Compression
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Local Patches Meet Global Context: Scalable 3D Diffusion Priors for Computed Tomography Reconstruction
- PSI3D: Plug-and-Play 3D Stochastic Inference with Slice-wise Latent Diffusion Prior
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- Spherical Leech Quantization for Visual Tokenization and Generation
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- RecTok: Reconstruction Distillation along Rectified Flow
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Multi-temporal Calving Front Segmentation
- Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
- AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- Voxify3D: Pixel Art Meets Volumetric Rendering
- Training-Free Vector Quantization via Gaussian VAEs
- See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
- HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- Efficient Generative Transformer Operators For Million-Point PDEs
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- DeRA: Decoupled Representation Alignment for Video Tokenization
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Rethinking Security in Semantic Communication: Latent Manipulation as a New Threat
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
- PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation
- Visual Generation Tuning
- Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image Detection
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- Bringing Your Portrait to 3D Presence
- Adversarial Flow Models
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- The Collapse of Patches
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Latent Diffusion Inversion Requires Understanding the Latent Space
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Flow Map Distillation Without Data
- Understanding, Accelerating, and Improving MeanFlow Training
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- CoD: A Diffusion Foundation Model for Image Compression
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
- Spanning Tree Autoregressive Visual Generation
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- Flow and Depth Assisted Video Prediction with Latent Transformer
- Decoupling Complexity from Scale in Latent Diffusion Model
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- WiCo-PG: Wireless Channel Foundation Model for Pathloss Map Generation via Synesthesia of Machines
- WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- MixAR: Mixture Autoregressive Image Generation
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- Point Cloud Quantization through Multimodal Prompting for 3D Understanding
- A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space
- FlowCast: Advancing Precipitation Nowcasting with Conditional Flow Matching
- FedeCouple: Fine-Grained Balancing of Global-Generalization and Local-Adaptability in Federated Learning
- TransactionGPT
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- Retrospective motion correction in MRI using disentangled embeddings
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- PADM: A Physics-aware Diffusion Model for Attenuation Correction
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
- Seq2Seq Models Reconstruct Visual Jigsaw Puzzles without Seeing Them
- MALeR: Improving Compositional Fidelity in Layout-Guided Generation
- CPO: Condition Preference Optimization for Controllable Image Generation
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
- Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
- DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- BI-DCGAN: A Theoretically Grounded Bayesian Framework for Efficient and Diverse GANs
- Emu3.5: Native Multimodal Models are World Learners
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Product-Quantised Image Representation for High-Quality Image Synthesis
- Amortized Moment Matching for Visual Generation
- Generative Modeling via Drifting
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
- Nested AutoRegressive Models
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- Morphologically Intelligent Perturbation Prediction with FORM
- Pctx: Tokenizing Personalized Context for Generative Recommendation
- Improved Training Technique for Shortcut Models
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Exploring Conditions for Diffusion models in Robotic Control
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- Vector Quantization in the Brain: Grid-like Codes in World Models
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- End-to-End Multi-Modal Diffusion Mamba
- NeuroRVQ: Multi-Scale EEG Tokenization for Generative Large Brainwave Models
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- Universal Image Restoration Pre-training via Masked Degradation Classification
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- What If : Understanding Motion Through Sparse Interactions
- BIGFix: Bidirectional Image Generation with Token Fixing
- Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Diffusion Transformers with Representation Autoencoders
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- Generative Latent Video Compression
- Variational Secret Common Randomness Extraction
- MelTok: 2D Tokenization for Single-Codebook Audio Compression
- Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Optimal Stopping in Latent Diffusion Models
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- CodeFormer++: Blind Face Restoration Using Deformable Registration and Deep Metric Learning
- Bridging Text and Video Generation: A Survey
- Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Flow Autoencoders are Effective Protein Tokenizers
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- MARS: Audio Generation via Multi-Channel Autoregression on Spectrograms
- PUREVQ-GAN: Defending Data Poisoning Attacks through Vector-Quantized Bottlenecks
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Score-based Membership Inference on Diffusion Models
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- Environment-Aware Satellite Image Generation with Diffusion Models
- Real-Aware Residual Model Merging for Deepfake Detection
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Tumor Synthesis conditioned on Radiomics
- Scalable GANs with Transformers
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Object-AVEdit: An Object-level Audio-Visual Editing Model
- Entering the Era of Discrete Diffusion Models: A Benchmark for Schrödinger Bridges and Entropic Optimal Transport
- Stochastic Interpolants via Conditional Dependent Coupling
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
- Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
- CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- Codebook-Based Adaptive Feature Compression With Semantic Enhancement for Edge-Cloud Systems
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Learning Dexterous Manipulation with Quantized Hand State
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Efficient Rectified Flow for Image Fusion
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- AToken: A Unified Tokenizer for Vision
- Image Tokenizer Needs Post-Training
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- Reconstruction Alignment Improves Unified Multimodal Models
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- PRIM: Towards Practical In-Image Multilingual Machine Translation
- Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Acoustic Interference Suppression in Ultrasound images for Real-Time HIFU Monitoring Using an Image-Based Latent Diffusion Model
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
- Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- Taming Transformer for Emotion-Controllable Talking Face Generation
- Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
- 2D Gaussians Meet Visual Tokenizer
- Next Visual Granularity Generation
- Versatile Video Tokenization with Generative 2D Gaussian Splatting
- Semi-supervised Image Dehazing via Expectation-Maximization and Bidirectional Brownian Bridge Diffusion Models
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Prototype-Guided Diffusion: Visual Conditioning without External Memory
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- ThematicPlane: Bridging Tacit User Intent and Latent Spaces for Image Generation
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UniTalker: Conversational Speech-Visual Synthesis
- Spectral Efficiency-Aware Codebook Design for Task-Oriented Semantic Communications
- Deeper Inside Deep ViT
- Cross-Domain Image Synthesis: Generating H&E from Multiplex Biomarker Imaging
- HPSv3: Towards Wide-Spectrum Human Preference Score
- GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
- CIVQLLIE: Causal Intervention with Vector Quantization for Low-Light Image Enhancement
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
- StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Subtyping Breast Lesions via Generative Augmentation based Long-tailed Recognition in Ultrasound
- GVD: Guiding Video Diffusion Model for Scalable Video Distillation
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- HDR Environment Map Estimation with Latent Diffusion Models
- Kernel Learning for Sample Constrained Black-Box Optimization
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
Discussions
Related