GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
2021/12/20 by Alex Nichol, Prafulla Dhariwal, Nichol, Alex +13 · 1 voice · 288 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Image Retrieval and Classification Techniques
paper · pdf · doi:10.48550/arxiv.2112.10741
Abstract
Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity. We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance. We find that the latter is preferred by human evaluators for both photorealism and caption similarity, and often produces photorealistic samples. Samples from a 3.5 billion parameter text-conditional diffusion model using classifier-free guidance are favored by human evaluators to those from DALL-E, even when the latter uses expensive CLIP reranking. Additionally, we find that our models can be fine-tuned to perform image inpainting, enabling powerful text-driven image editing. We train a smaller model on a filtered dataset and release the code and weights at https://github.com/openai/glide-text2im.
Cited by
- Importance-Aware OBS Pruning for Diffusion Models
- Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
- To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion
- RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
- When 2D Cues Fail: Improving Image Manipulation Localization with Reliable 3D Geometry
- Signed Rectified Flow: Negativity-Controlled Generation
- H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
- ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
- UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation
- Semi-Supervised Conditional Diffusion via Label Augmentation
- Introspective Attention Modulation for Safe Text-to-Image Generation
- Deep models of protein evolution in time generate realistic evolutionary trajectories and functional proteins
- Mechanisms of Projective Composition of Diffusion Models
- Information-Theoretic Proofs for Diffusion Sampling
- Rethinking the Use of Vision Transformers for AI-Generated Image Detection
- LC4-DViT: Land-cover Creation for Land-cover Classification with Deformable Vision Transformer
- M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
- Plug In, Grade Right: Psychology-Inspired AGIQA
- Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction
- PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
- Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design
- Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
- AI-generated Images Challenge Visual Trust in High-risk Scenarios
- FUSE: Unifying Spectral and Semantic Cues for Robust AI-Generated Image Detection
- FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting
- Region-Constraint In-Context Generation for Instructional Video Editing
- Detection of AI Generated Images Using Combined Uncertainty Measures and Particle Swarm Optimised Rejection Mechanism
- InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- DeContext as Defense: Safe Image Editing in Diffusion Transformers
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- 4D-RaDiff: Latent Diffusion for 4D Radar Point Cloud Generation
- FastDDHPose: Towards Unified, Efficient, and Disentangled 3D Human Pose Estimation
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- Directional Textual Inversion for Personalized Text-to-Image Generation
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
- SCAdapter: Content-Style Disentanglement for Diffusion Style Transfer
- TA-KAND: Two-stage Attention Triple Enhancement and U-KAN based Diffusion For Few-shot Knowledge Graph Completion
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
- Food Image Generation on Multi-Noun Categories
- ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment
- BlurDM: A Blur Diffusion Model for Image Deblurring
- GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces
- Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
- SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
- CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
- Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
- SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- LaGen: Towards Autoregressive LiDAR Scene Generation
- Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
- From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained Annotations
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
- DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
- ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
- ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
- Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image Detection
- TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis
- GrOCE:Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models
- SVG360: Editable Multiview Vector Graphics from a Single SVG
- Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
- PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
- How Noise Benefits AI-generated Image Detection
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- CFG-EC: Error Correction Classifier-Free Guidance
- Training-free Detection of AI-generated images via Cropping Robustness
- CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
- Explainable AI-Generated Image Detection RewardBench
- GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
- Fast Data Attribution for Text-to-Image Models
- Top2Ground: A Height-Aware Dual Conditioning Diffusion Model for Robust Aerial-to-Ground View Generation
- CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
- Test-Time Iterative Error Correction for Efficient Diffusion Models
- MALeR: Improving Compositional Fidelity in Layout-Guided Generation
- Seeing What You Say: Expressive Image Generation from Speech
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Evolve to Inspire: Novelty Search for Diverse Image Generation
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Who Made This? Fake Detection and Source Attribution with Diffusion Features
- Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
- CompAgent: An Agentic Framework for Visual Compliance Verification
- Optimal Convergence Analysis of DDPM for General Distributions
- H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models
- A Retrospect to Multi-prompt Learning across Vision and Language
- Group-Equivariant Diffusion Models for Lattice Field Theory
- PEO: Training-Free Aesthetic Quality Enhancement in Pre-Trained Text-to-Image Diffusion Models with Prompt Embedding Optimization
- DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
- Diffusion Models in Vision: A Survey
- EIRES:Training-free AI-Generated Image Detection via Edit-Induced Reconstruction Error Shift
- PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
- UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- Model-free filtering in high dimensions via projection and score-based diffusions
- Optimize Any Topology: A Foundation Model for Shape- and Resolution-Free Structural Topology Optimization
- LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- Unsupervised Deep Generative Models for Anomaly Detection in Neuroimaging: A Systematic Scoping Review
- Spatial Computing Communications for Multi-User Virtual Reality in Distributed Mobile Edge Computing Network
- LOTA: Bit-Planes Guided AI-Generated Image Detection
- Salient Concept-Aware Generative Data Augmentation
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- A Black-Box Debiasing Framework for Conditional Sampling
- Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
- Zero-shot Face Editing via ID-Attribute Decoupled Inversion
- DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing
- Controllable Graph Generation with Diffusion Models via Inference-Time Tree Search Guidance
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Using neural style transfer to study the evolution of animal signal design: A case study in an ornamented fish
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- Cross-Sensor Touch Generation
- MSDM: Generating Task-Specific Pathology Images with a Multimodal Conditioned Diffusion Model for Cell and Nuclei Segmentation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- SummDiff: Generative Modeling of Video Summarization with Diffusion
- Rectified-CFG++ for Flow Based Models
- InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
- Accelerating Diffusion LLM Inference via Local Determinism Propagation
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
- Teleportraits: Training-Free People Insertion into Any Scene
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Teamwork: Collaborative Diffusion with Low-rank Coordination and Adaptation
- Diffusion2: Turning 3D Environments into Radio Frequency Heatmaps
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- Self Speculative Decoding for Diffusion Large Language Models
- Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
- How Diffusion Models Memorize
- Environment-Aware Satellite Image Generation with Diffusion Models
- Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
- UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
- RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
- Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
- DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
- A Unified Framework for Diffusion Model Unlearning with f-Divergence
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- SimDiff: Simulator-constrained Diffusion Model for Physically Plausible Motion Generation
- Efficient Encoder-Free Pose Conditioning and Pose Control for Virtual Try-On
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- MMG: Mutual Information Estimation via the MMSE Gap in Diffusion
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Transformers and genome language models
- CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- Single-Image Depth from Defocus with Coded Aperture and Diffusion Posterior Sampling
- CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
- nDNA -- the Semantic Helix of Artificial Cognition
- Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
- FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
- Lynx: Towards High-Fidelity Personalized Video Generation
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- TrueMoE: Dual-Routing Mixture of Discriminative Experts for Synthetic Image Detection
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- StyleProtect: Safeguarding Artistic Identity in Fine-tuned Diffusion Models
- Double Helix Diffusion for Cross-Domain Anomaly Image Generation
- Data-Efficient Ensemble Weather Forecasting with Diffusion Models
- A Novel Local Focusing Mechanism for Deepfake Detection Generalization
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Feature Space Analysis by Guided Diffusion Model
- Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity
- SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
- From Editor to Dense Geometry Estimator
- Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning
- Enhancing Robustness in Post-Processing Watermarking: An Ensemble Attack Network Using CNNs and Transformers
- A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
- Text-to-Layout: A Generative Workflow for Drafting Architectural Floor Plans Using LLMs
- Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Generative AI in Map-Making: A Technical Exploration and Its Implications for Cartographers
- GSFix3D: Diffusion-Guided Repair of Novel Views in Gaussian Splatting
- MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic Expansion
- Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
- FOCUS: Frequency-Optimized Conditioning of DiffUSion Models for mitigating catastrophic forgetting during Test-Time Adaptation
- Side Effects of Erasing Concepts from Diffusion Models
- SBGD: Improving Graph Diffusion Generative Model via Stochastic Block Diffusion
- AnchorSync: Global Consistency Optimization for Long Video Editing
- Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
- DMS:Diffusion-Based Multi-Baseline Stereo Generation for Improving Self-Supervised Depth Estimation
- Generic Event Boundary Detection via Denoising Diffusion
- SPG: Style-Prompting Guidance for Style-Specific Content Creation
- Object Fidelity Diffusion for Remote Sensing Image Generation
- CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
- Projected Coupled Diffusion for Test-Time Constrained Joint Generation
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
- Learning User Preferences for Image Generation Model
- LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering
- Bridging Semantic Logic Gaps: A Cognition Inspired Multimodal Boundary Preserving Network for Image Manipulation Localization
- CharacterShot: Controllable and Consistent 4D Character Animation
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- DreamVE: Unified Instruction-based Image and Video Editing
- MetaDiT: Enabling Fine-grained Constraints in High-degree-of Freedom Metasurface Design
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection
- DogFit: Domain-guided Fine-tuning for Efficient Transfer Learning of Diffusion Models
- HPSv3: Towards Wide-Spectrum Human Preference Score
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
- Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
- VideoGuard: Protecting Video Content from Unauthorized Editing
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- Can LLMs Generate High-Quality Task-Specific Conversations?
- AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
- Versatile Transition Generation with Image-to-Video Diffusion
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- LeakyCLIP: Extracting Training Data from CLIP
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
- DivControl: Knowledge Diversion for Controllable Image Generation
- Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
- DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent Diffusion
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
- GuidPaint: Class-Guided Image Inpainting with Diffusion Models
Discussions
Related