Progressive Growing of GANs for Improved Quality, Stability, and Variation
2017/10/27 by Tero Karras, Timo Aila, Karras, Tero +5 · 221 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Image Processing Techniques #Digital Media Forensic Detection
paper · pdf · doi:10.48550/arxiv.1710.10196
Abstract
We describe a new training methodology for generative adversarial networks. The key idea is to grow both the generator and discriminator progressively: starting from a low resolution, we add new layers that model increasingly fine details as training progresses. This both speeds the training up and greatly stabilizes it, allowing us to produce images of unprecedented quality, e.g., CelebA images at 10242. We also propose a simple way to increase the variation in generated images, and achieve a record inception score of 8.80 in unsupervised CIFAR10. Additionally, we describe several implementation details that are important for discouraging unhealthy competition between the generator and discriminator. Finally, we suggest a new metric for evaluating GAN results, both in terms of image quality and variation. As an additional contribution, we construct a higher-quality version of the CelebA dataset.
Cited by
- Reverse Personalization
- Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
- Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
- SDUM: A Scalable Deep Unrolled Model for Universal MRI Reconstruction
- Generative Latent Coding for Ultra-Low Bitrate Image Compression
- Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- Preserving Spectral Structure and Statistics in Diffusion Models
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- Disentangled representations via score-based variational autoencoders
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
- Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
- Learning Common and Salient Generative Factors Between Two Image Datasets
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- Unified Control for Inference-Time Guidance of Denoising Diffusion Models
- MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Targeted Data Protection for Diffusion Model by Matching Training Trajectory
- FBA2D: Frequency-based Black-box Attack for AI-generated Image Detection
- FlowSteer: Conditioning Flow Field for Consistent Image Restoration
- MSN: Multi-directional Similarity Network for Hand-crafted and Deep-synthesized Copy-Move Forgery Detection
- SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- Open Set Face Forgery Detection via Dual-Level Evidence Collection
- SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
- Adversarial Flow Models
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
- MFM-point: Multi-scale Flow Matching for Point Cloud Generation
- Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image Detection
- FLUID: Training-Free Face De-identification via Latent Identity Substitution
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Motion Transfer-Enhanced StyleGAN for Generating Diverse Macaque Facial Expressions
- Arbitrary-Resolution and Arbitrary-Scale Face Super-Resolution with Implicit Representation Networks
- How Noise Benefits AI-generated Image Detection
- Training-free Detection of AI-generated images via Cropping Robustness
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection
- Functional Mean Flow in Hilbert Space
- PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
- Diffusion Model Based Signal Recovery Under 1-Bit Quantization
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- Equivariant Sampling for Improving Diffusion Model-based Image Restoration
- Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models
- Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion Models
- HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- CSF-Net: Context-Semantic Fusion Network for Large Mask Inpainting
- Rectified Noise: A Generative Model Using Positive-incentive Noise
- CINEMAE: Leveraging Frozen Masked Autoencoders for Cross-Generator AI Image Detection
- Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
- PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
- Finetuning-Free Personalization of Text to Image Generation via Hypernetworks
- Interpreting the Latent Space of GANs for Semantic Face Editing
- Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- Bayesian model selection and misspecification testing in imaging inverse problems only from noisy and partial measurements
- Face De-Identification: A Domain-Centric Survey from Capture to Processing
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- Determinants of ChatGPT Use and its Impact on Learning Performance: An Integrated Model of BRT and TPB
- High-Resolution Image Synthesis with Latent Diffusion Models
- More Real Than Real: A Study on Human Visual Perception of Synthetic Faces [Applications Corner]
- TextStyleBrush: Transfer of Text Aesthetics From a Single Example
- Fluorescence image deconvolution microscopy via generative adversarial learning (FluoGAN)
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- Optimal 1-Wasserstein Distance for WGANs
- MIND: Monge Inception Distance for Generative Models Evaluation
- Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
- Privacy preserving Generative Adversarial Networks to model Electronic Health Records
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- Generative adversarial networks for speech processing: A review
- Beyond Inference Intervention: Identity-Decoupled Diffusion for Face Anonymization
- Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
- Through the Lens: Benchmarking Deepfake Detectors Against Moiré-Induced Distortions
- Residual Diffusion Bridge Model for Image Restoration
- On the Anisotropy of Score-Based Generative Models
- Self-Attention Decomposition For Training Free Diffusion Editing
- Bridging the gap to real-world language-grounded visual concept learning
- 3DPR: Single Image 3D Portrait Relight using Generative Priors
- BADiff: Bandwidth Adaptive Diffusion Model
- Improved Training Technique for Shortcut Models
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- BrainPuzzle: Hybrid Physics and Data-Driven Reconstruction for Transcranial Ultrasound Tomography
- Efficient Few-shot Identity Preserving Attribute Editing for 3D-aware Deep Generative Models
- Investigating Adversarial Robustness against Preprocessing used in Blackbox Face Recognition
- Latent Diffusion Model without Variational Autoencoder
- WithAnyone: Towards Controllable and ID Consistent Image Generation
- Inference-Time Search using Side Information for Diffusion-based Image Reconstruction
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
- Counting Hallucinations in Diffusion Models
- Learning Latent Energy-Based Models via Interacting Particle Langevin Dynamics
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Zero-shot Face Editing via ID-Attribute Decoupled Inversion
- Encoder Decoder Generative Adversarial Network Model for Stock Market Prediction
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations
- SpecGuard: Spectral Projection-based Advanced Invisible Watermarking
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
- GLVD: Guided Learned Vertex Descent
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- CodeFormer++: Blind Face Restoration Using Deformable Registration and Deep Metric Learning
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- SDAKD: Student Discriminator Assisted Knowledge Distillation for Super-Resolution Generative Adversarial Networks
- Contrastive-SDE: Guiding Stochastic Differential Equations with Contrastive Learning for Unpaired Image-to-Image Translation
- Deep Learning for Image Super-Resolution: A Survey
- Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
- Dale meets Langevin: A Multiplicative Denoising Diffusion Model
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- Secure and reversible face anonymization with diffusion models
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Post-Training Quantization via Residual Truncation and Zero Suppression for Diffusion Models
- Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
- One-shot Conditional Sampling: MMD meets Nearest Neighbors
- VAGUEGAN: Stealthy Poisoning and Backdoor Attacks on Image Generative Pipelines
- Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
- Scalable GANs with Transformers
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- Position-Blind Ptychography: Viability of image reconstruction via data-driven variational inference
- Enhancing Blind Face Restoration through Online Reinforcement Learning
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- Deepfakes: we need to re-think the concept of "real" images
- Consistency Models as Plug-and-Play Priors for Inverse Problems
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- CusEnhancer: A Zero-Shot Scene and Controllability Enhancement Method for Photo Customization via ResInversion
- Sobolev acceleration for neural networks
- Collaborative feature aggregation for face super-resolution and robust re-identification
- Material dynamics analysis with deep generative model
- SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
- CATformer: Contrastive Adversarial Transformer for Image Super-Resolution
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- Improving the color accuracy of lighting estimation models
- \mathttM3VIR: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation
- PMRT: A Training Recipe for Fast, 3D High-Resolution Aerodynamic Prediction
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection
- JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
- TrueMoE: Dual-Routing Mixture of Discriminative Experts for Synthetic Image Detection
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Controllable Localized Face Anonymization Via Diffusion Inpainting
- Bias Amplification in Stable Diffusion's Representation of Stigma Through Skin Tones and Their Homogeneity
- EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
- RaceGAN: A Framework for Preserving Individuality while Converting Racial Information for Image-to-Image Translation
- UTOPY: Unrolling Algorithm Learning via Fidelity Homotopy for Inverse Problems
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- Defending Deepfake via Texture Feature Perturbation
- Adversarial Appearance Learning in Augmented Cityscapes for Pedestrian Recognition in Autonomous Driving
- Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
- ReTrack: Data Unlearning in Diffusion Models through Redirecting the Denoising Trajectory
- Uncovering and Mitigating Destructive Multi-Embedding Attacks in Deepfake Proactive Forensics
- A Novel Local Focusing Mechanism for Deepfake Detection Generalization
- Locality in Image Diffusion Models Emerges from Data Statistics
- Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- DrumGAN: Synthesis of Drum Sounds With Timbral Feature Conditioning Using Generative Adversarial Networks
- Expandable Residual Approximation for Knowledge Distillation
- Face4FairShifts: A Large Image Benchmark for Fairness and Robust Learning across Visual Domains
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- DLGAN : Time Series Synthesis Based on Dual-Layer Generative Adversarial Networks
- Non-expert to Expert Motion Translation Using Generative Adversarial Networks
- Disruptive Attacks on Face Swapping via Low-Frequency Perceptual Perturbations
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- IDF: Iterative Dynamic Filtering Networks for Generalizable Image Denoising
- RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration
- VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
- BrainPath: A Biologically-Informed AI Framework for Individualized Aging Brain Generation
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- Deadline-Aware Bandwidth Allocation for Semantic Generative Communication with Diffusion Models
- TimeMachine: Fine-Grained Facial Age Editing with Identity Preservation
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- Generation of Indian Sign Language Letters, Numbers, and Words
- CLIP-Flow: A Universal Discriminator for AI-Generated Images Inspired by Anomaly Detection
- Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- VQ-VAE Based Digital Semantic Communication with Importance-Aware OFDM Transmission
- Unlocking the Potential of Diffusion Priors in Blind Face Restoration
- CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
- FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
- How Does Bilateral Ear Symmetry Affect Deep Ear Features?
- Anti-Tamper Protection for Unauthorized Individual Image Generation
- EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
- Spatiotemporal wall pressure forecast of a rectangular cylinder with physics-aware DeepU-Fourier neural network
- Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Visual Language Models as Zero-Shot Deepfake Detectors
- Quantum generative modeling for financial time series with temporal correlations
- Evaluating Deepfake Detectors in the Wild
- Learning Kinetic Monte Carlo stochastic dynamics with Deep Generative Adversarial Networks
- Principled Curriculum Learning using Parameter Continuation Methods
- PanoGAN A Deep Generative Model for Panoramic Dental Radiographs
- Harnessing Diffusion-Yielded Score Priors for Image Restoration
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy
Related