Analyzing and Improving the Image Quality of StyleGAN
2019/12/03 by Tero Karras, Samuli Laine, Karras, Tero +9 · 2 voices · 173 citations
#cs.CV #cs.LG #cs.NE #eess.IV #stat.ML
paper · pdf · doi:10.48550/arxiv.1912.04958
Abstract
The style-based GAN architecture (StyleGAN) yields state-of-the-art results in data-driven unconditional generative image modeling. We expose and analyze several of its characteristic artifacts, and propose changes in both model architecture and training methods to address them. In particular, we redesign the generator normalization, revisit progressive growing, and regularize the generator to encourage good conditioning in the mapping from latent codes to images. In addition to improving image quality, this path length regularizer yields the additional benefit that the generator becomes significantly easier to invert. This makes it possible to reliably attribute a generated image to a particular network. We furthermore visualize how well the generator utilizes its output resolution, and identify a capacity problem, motivating us to train larger models for additional quality improvements. Overall, our improved model redefines the state of the art in unconditional image modeling, both in terms of existing distribution quality metrics as well as perceived image quality.
Cited by
- Reverse Personalization
- MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
- A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
- TexAvatars : Hybrid Texel-3D Representations for Stable Rigging of Photorealistic Gaussian Head Avatars
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- SkinGenBench: Generative Model and Preprocessing Effects for Synthetic Dermoscopic Augmentation in Melanoma Diagnosis
- Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- Interpretable Similarity of Synthetic Image Utility
- FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- BLANKET: Anonymizing Faces in Infant Video Recordings
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
- VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
- Learning Common and Salient Generative Factors Between Two Image Datasets
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- PersonaLive! Expressive Portrait Image Animation for Live Streaming
- MSN: Multi-directional Similarity Network for Hand-crafted and Deep-synthesized Copy-Move Forgery Detection
- AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment
- EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
- StyleYourSmile: Cross-Domain Face Retargeting Without Paired Multi-Style Data
- Reversible Inversion for Training-Free Exemplar-guided Image Editing
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- Soft Quality-Diversity Optimization
- Vision Bridge Transformer at Scale
- Adversarial Flow Models
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
- Waveform-Based Probabilistic Seismic Hazard Analysis Using Ground-Motion Generative Models
- TSGM: Regular and Irregular Time-series Generation using Score-based Generative Models
- Multiscale Vector-Quantized Variational Autoencoder for Endoscopic Image Synthesis
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image Detection
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- FLUID: Training-Free Face De-identification via Latent Identity Substitution
- Motion Transfer-Enhanced StyleGAN for Generating Diverse Macaque Facial Expressions
- How Noise Benefits AI-generated Image Detection
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection
- SAGA: Source Attribution of Generative AI Videos
- Rethinking Bias in Generative Data Augmentation for Medical AI: a Frequency Recalibration Method
- StyleQoRA: Quality-Aware Low-Rank Adaptation for Few-Shot Multi-Style Editing
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- AvatarTex: High-Fidelity Facial Texture Reconstruction from Single-Image Stylized Avatars
- Knowledge-based anomaly detection for identifying network-induced shape artifacts
- Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
- I Prompt, it Generates, we Negotiate. Exploring Text-Image Intertextuality in Human-AI Co-Creation of Visual Narratives with VLMs
- Generative Hints
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- A New Perspective on Precision and Recall for Generative Models
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
- Instruction-based image editing: a survey on data, models, evaluation, and applications
- Face De-Identification: A Domain-Centric Survey from Capture to Processing
- VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
- LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models
- Fourier-Based GAN Fingerprint Detection using ResNet50
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Uniform Discrete Diffusion with Metric Path for Video Generation
- See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
- Scaling Non-Parametric Sampling with Representation
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- Improved Training Technique for Shortcut Models
- Wavelet-based GAN Fingerprint Detection using ResNet50
- Diffusion Buffer for Online Generative Speech Enhancement
- Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
- Controlling the image generation process with parametric activation functions
- A Multi-domain Image Translative Diffusion StyleGAN for Iris Presentation Attack Detection
- Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation
- CoLoR-GAN: Continual Few-Shot Learning with Low-Rank Adaptation in Generative Adversarial Networks
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- Encoder Decoder Generative Adversarial Network Model for Stock Market Prediction
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations
- FLOWING: Implicit Neural Flows for Structure-Preserving Morphing
- Beyond Real Data: Synthetic Data through the Lens of Regularization
- Towards Real-World Deepfake Detection: A Diverse In-the-wild Dataset of Forgery Faces
- ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
- SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- What Drives Compositional Generalization in Visual Generative Models?
- Secure and reversible face anonymization with diffusion models
- Multi-Domain Brain Vessel Segmentation Through Feature Disentanglement
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Learning Energy-based Variational Latent Prior for VAEs
- Data-to-Energy Stochastic Dynamics
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- Real-Aware Residual Model Merging for Deepfake Detection
- NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
- Scalable GANs with Transformers
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- Enhancing Blind Face Restoration through Online Reinforcement Learning
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- DragGANSpace: Latent Space Exploration and Control for GANs
- LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training
- A Real-Time On-Device Defect Detection Framework for Laser Power-Meter Sensors via Unsupervised Learning
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- Generative Model Inversion Through the Lens of the Manifold Hypothesis
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- StrCGAN: A Generative Framework for Stellar Image Restoration
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- PRNU-Bench: A Novel Benchmark and Model for PRNU-Based Camera Identification
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- Enhancing Reference-based Sketch Colorization via Separating Reference Representations
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- UMind: A Unified Multitask Network for Zero-Shot M/EEG Visual Decoding
- ShaLa: Multimodal Shared Latent Space Modelling
- Defending Deepfake via Texture Feature Perturbation
- DEFT-VTON: Efficient Virtual Try-On with Consistent Generalised H-Transform
- VQT-Light:Lightweight HDR Illumination Map Prediction with Richer Texture.pdf
- Double Helix Diffusion for Cross-Domain Anomaly Image Generation
- Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
- DRAG: Data Reconstruction Attack using Guided Diffusion
- A Controllable 3D Deepfake Generation Framework with Gaussian Splatting
- Beyond Sliders: Mastering the Art of Diffusion-based Image Manipulation
- Styleclone: Face Stylization with Diffusion Based Data Augmentation
- A Discrepancy-Based Perspective on Dataset Condensation
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
- MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
- Reconstruction and Reenactment Separated Method for Realistic Gaussian Head
- Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- From Editor to Dense Geometry Estimator
- DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
- Human Motion Video Generation: A Survey
- Enhancing Robustness in Post-Processing Watermarking: An Ensemble Attack Network Using CNNs and Transformers
- An Investigation of Visual Foundation Models Robustness
- Expandable Residual Approximation for Knowledge Distillation
- CoFE: A Framework Generating Counterfactual ECG for Explainable Cardiac AI-Diagnostics
- Weighted Support Points from Random Measures: An Interpretable Alternative for Generative Modeling
- Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
- Quantum latent distributions in deep generative models
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- PanoHair: Detailed Hair Strand Synthesis on Volumetric Heads
- FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained Poses
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
- subCellSAM: Zero-Shot (Sub-)Cellular Segmentation for Hit Validation in Drug Discovery
- DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
- ID-Card Synthetic Generation: Toward a Simulated Bona fide Dataset
- WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
- TimeMachine: Fine-Grained Facial Age Editing with Identity Preservation
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- Unlocking the Potential of Diffusion Priors in Blind Face Restoration
- Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
- Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Towards Privacy-preserving Photorealistic Self-avatars in Mixed Reality
- HDR Environment Map Estimation with Latent Diffusion Models
Discussions
Related