Zero-Shot Text-to-Image Generation
2021/02/24 by Aditya Ramesh, Ramesh, Aditya, Mikhail Pavlov +13 · 2 voices · 346 citations
Computer Science · #Advanced Neural Network Applications #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2102.12092
openalex publication_date 2021/02/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.
Citations
Cited by
- ProGuard: Towards Proactive Multimodal Safeguard
- REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
- Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
- Large Vision Model-Enhanced Digital Twin with Deep Reinforcement Learning for User Association and Load Balancing in Dynamic Wireless Networks
- PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
- How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- Generative diffusion models for agricultural AI: plant image generation, indoor-to-outdoor translation, and expert preference alignment
- PSI3D: Plug-and-Play 3D Stochastic Inference with Slice-wise Latent Diffusion Prior
- InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Vision-Language Model Guided Image Restoration
- Next-Embedding Prediction Makes Strong Vision Learners
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Robust and Calibrated Detection of Authentic Multimedia Content
- Where is the Watermark? Interpretable Watermark Detection at the Block Level
- Spherical Leech Quantization for Visual Tokenization and Generation
- Attention-Based Foundation Model for Quantum States
- Directional Textual Inversion for Personalized Text-to-Image Generation
- Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
- JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion
- TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
- AutoMV: An Automatic Multi-Agent System for Music Video Generation
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
- Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
- D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
- Self-Evolving 3D Scene Generation from a Single Image
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- BlurDM: A Blur Diffusion Model for Image Deblurring
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text
- Understanding and Harnessing Sparsity in Unified Multimodal Models
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- Accelerating Inference of Masked Image Generators via Reinforcement Learning
- FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- Semantic Anchoring for Robust Personalization in Text-to-Image Diffusion Models
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
- Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Flow Map Distillation Without Data
- FineXtrol: Controllable Motion Generation via Fine-Grained Text
- Single Image to High-Quality 3D Object via Latent Features
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- EvDiff: High Quality Video with an Event Camera
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- SplitFlux: Learning to Decouple Content and Style from a Single Image
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
- Coffee: Controllable Diffusion Fine-tuning
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
- Point Cloud Quantization through Multimodal Prompting for 3D Understanding
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
- How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- MALeR: Improving Compositional Fidelity in Layout-Guided Generation
- Enhancing Diffusion Model Guidance through Calibration and Regularization
- Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
- Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- Who Made This? Fake Detection and Source Attribution with Diffusion Features
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- FMint-SDE: A Multimodal Foundation Model for Accelerating Numerical Simulation of SDEs via Error Correction
- NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
- From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- Model Inversion Attacks Meet Cryptographic Fuzzy Extractors
- FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models
- LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models
- Product-Quantised Image Representation for High-Quality Image Synthesis
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Synthesis
- Future of AI Models: A Computational perspective on Model collapse
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Uniform Discrete Diffusion with Metric Path for Video Generation
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
- FARMER: Flow AutoRegressive Transformer over Pixels
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
- FAME: Fairness-aware Attention-modulated Video Editing
- Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling
- Human-Centred Evaluation of Text-to-Image Generation Models for Self-expression of Mental Distress: A Dataset Based on GPT-4o
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Exponential Convergence Guarantees for Iterative Markovian Fitting
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
- High-Resolution Complex Scene Synthesis with Transformers
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- MEG-GPT: A transformer-based foundation model for magnetoencephalography data
- OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- Adaptive Discretization for Consistency Models
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Consistent text-to-image generation via scene de-contextualization
- Provenance of AI-Generated Images: A Vector Similarity and Blockchain-based Approach
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- Adaptive Visual Conditioning for Semantic Consistency in Diffusion-Based Story Continuation
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- BIGFix: Bidirectional Image Generation with Token Fixing
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Using neural style transfer to study the evolution of animal signal design: A case study in an ornamented fish
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- PHyCLIP: ℓ1-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Rectified-CFG++ for Flow Based Models
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Graph Conditioned Diffusion for Controllable Histopathology Image Generation
- Expressive and Scalable Quantum Fusion for Multimodal Learning
- Proteome-wide model for human disease genetics
- StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
- Position: Towards Responsible Evaluation for Text-to-Speech
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Signing Right Away
- TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
- Bridging Text and Video Generation: A Survey
- Beyond Static Knowledge Messengers: Towards Adaptive, Fair, and Scalable Federated Learning for Medical AI
- Unlearning in Diffusion models under Data Constraints: A Variational Inference Approach
- Evolutionary Computation as Natural Generative AI
- Unsupervised Transformer Pre-Training for Images: Self-Distillation, Mean Teachers, and Random Crops
- Composite Optimization with Error Feedback: the Dual Averaging Approach
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- SimMIM: A Simple Framework for Masked Image Modeling
- Evaluating Large Language Models Trained on Code
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Importance of localized dilatation and distensibility in identifying determinants of thoracic aortic aneurysm with neural operators
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Discrete Variational Autoencoding via Policy Search
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Zero-shot evaluation reveals limitations of single-cell foundation models
- Scale-Wise VAR is Secretly Discrete Diffusion
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- SlimDiff: Training-Free, Activation-Guided Hands-free Slimming of Diffusion Models
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- LAVA: Explainability for Unsupervised Latent Embeddings
- NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
- ThinkFake: Reasoning in Multimodal Large Language Models for AI-Generated Image Detection
- Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
- Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- BEiT: BERT Pre-Training of Image Transformers
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- From Global to Local: Social Bias Transfer in CLIP
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group Swapping
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- MEF: A Systematic Evaluation Framework for Text-to-Image Models
- Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- Lynx: Towards High-Fidelity Personalized Video Generation
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors
- The Iconicity of the Generated Image
- Computer-Aided Design as Language
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- GenCAD-3D: CAD Program Generation using Multimodal Latent Space Alignment and Synthetic Dataset Balancing
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- GenExam: A Multidisciplinary Text-to-Image Exam
- ShaLa: Multimodal Shared Latent Space Modelling
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image Generation
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Diffusion Models Beat GANs on Image Synthesis
- Paint4Poem: A Dataset for Artistic Visualization of Classical Chinese Poems
- Generative AI in Game Development: A Qualitative Research Synthesis
- Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
- Cross-Modal Retrieval with Cauchy-Schwarz Divergence
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
- PVNet: Point-Voxel Interaction LiDAR Scene Upsampling Via Diffusion Models
- A Novel Local Focusing Mechanism for Deepfake Detection Generalization
- MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Testing chatbots on the creation of encoders for audio conditioned image generation
- TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement
- No Encore: Unlearning as Opt-Out in Music Generation
- Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
- Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- Effectively obtaining acoustic, visual and textual data from videos
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- Pre-Forgettable Models: Prompt Learning as a Native Mechanism for Unlearning
- UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
- From Editor to Dense Geometry Estimator
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
- SOLD: SELFIES-based Objective-driven Latent Diffusion
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-driven Diffusion Models
- TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Zero-Shot Text-Guided Object Generation with Dream Fields
- Diffusion Language Models Know the Answer Before Decoding
- ViTGAN: Training GANs with Vision Transformers
- JVLGS: Joint Vision-Language Gas Leak Segmentation
- Causal Reinforcement Learning using Observational and Interventional Data
- Alias-Free Generative Adversarial Networks
- CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
- CLIPSym: Delving into Symmetry Detection with CLIP
- Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
- Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
- Sketchar: Supporting Character Design and Illustration Prototyping Using Generative AI
- DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Variational latent discrete representation for time series modelling
- SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
- SPG: Style-Prompting Guidance for Style-Specific Content Creation
- CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- NegFaceDiff: The Power of Negative Context in Identity-Conditioned Diffusion for Synthetic Face Generation
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- Understanding Dementia Speech Alignment with Diffusion-Based Image Generation
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Per-Query Visual Concept Learning
- LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering
- Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Multimodal learning with next-token prediction for large multimodal models
- Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
- VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
- Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models
- DreamVE: Unified Instruction-based Image and Video Editing
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- Deeper Inside Deep ViT
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- Diffusion Models with Adaptive Negative Sampling Without External Resources
- When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
- A Closed-Loop Multi-Agent Framework for Aerodynamics-Aware Automotive Styling Design
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
- AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
- Social Media Information Operations
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- StyleSentinel: Reliable Artistic Copyright Verification via Stylistic Fingerprints
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- Agency Among Agents: Designing with Hypertextual Friction in the Algorithmic Web
- Training-free Geometric Image Editing on Diffusion Models
- Next Tokens Denoising for Speech Synthesis
- Improving Generative Ad Text on Facebook using Reinforcement Learning
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration
- Meta CLIP 2: A Worldwide Scaling Recipe
- Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
Discussions
Related