A Style-Based Generator Architecture for Generative Adversarial Networks
2019/06/01 by Tero Karras, Samuli Laine, Timo Aila · 59 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Face recognition and analysis #Advanced Image Processing Techniques
paper · doi:10.1109/cvpr.2019.00453
openalex publication_date 2019/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
We propose an alternative generator architecture for generative adversarial networks, borrowing from style transfer literature. The new architecture leads to an automatically learned, unsupervised separation of high-level attributes (e.g., pose and identity when trained on human faces) and stochastic variation in the generated images (e.g., freckles, hair), and it enables intuitive, scale-specific control of the synthesis. The new generator improves the state-of-the-art in terms of traditional distribution quality metrics, leads to demonstrably better interpolation properties, and also better disentangles the latent factors of variation. To quantify interpolation quality and disentanglement, we propose two new, automated methods that are applicable to any generator architecture. Finally, we introduce a new, highly varied and high-quality dataset of human faces.
Citations
Cited by
- Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
- SinGAN: Learning a Generative Model from a Single Natural Image
- Style-aware gloss control for generative non-photorealistic rendering
- InstructAttribute: Fine-grained Object Attributes editing with Instruction
- PixelSmile: Toward Fine-Grained Facial Expression Editing
- Image Generation with a Sphere Encoder
- Revisiting Diffusion Autoencoder Training for Image Reconstruction Quality
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented Reality
- CoCoDiff: Diversifying Skeleton Action Features via Coarse-Fine Text-Co-Guided Latent Diffusion
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- X-Fusion: Introducing New Modality to Frozen Large Language Models
- PixelHacker: Image Inpainting with Structural and Semantic Consistency
- TrueFake: A Real World Case Dataset of Last Generation Fake Images also Shared on Social Networks
- CE-NPBG: Connectivity Enhanced Neural Point-Based Graphics for Novel View Synthesis in Autonomous Driving Scenes
- Multi-axis Analysis of Image Manipulation Localization
- RewardFlow: Generate Images by Optimizing What You Reward
- Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
- Sparsely Supervised Diffusion
- WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
- EarthMapper: Visual Autoregressive Models for Controllable Bidirectional Satellite-Map Translation
- A Langevin sampling algorithm inspired by the Adam optimizer
- DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion
- Equalized Generative Treatment: Matching f-divergences for Fairness in Generative Models
- Designing Digital Humans with Ambient Intelligence
- A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
- DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
- Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion
- Secure Intellicise Wireless Network: Agentic AI for Coverless Semantic Steganography Communication
- FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
- IGAN: Inferent and Generative Adversarial Networks
- Image Generation From Small Datasets via Batch Statistics Adaptation
- PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona
- ScoreField: Neural Inverse Scattering with Score-Based Generative Priors
- Towards Generalized and Training-Free Text-Guided Semantic Manipulation
- Generative Fields: Uncovering Hierarchical Feature Control for StyleGAN via Inverted Receptive Fields
- DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
- IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration
- Can Knowledge Improve Security? A Coding-Enhanced Jamming Approach for Semantic Communication
- Towards Fine-Grained Human Pose Transfer With Detail Replenishing Network
- Learning stochastic object models from medical imaging measurements by use of advanced ambient generative adversarial networks
- Complement Face Forensic Detection and Localization with FacialLandmarks
- StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians
- FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
- A Controllable Appearance Representation for Flexible Transfer and Editing
- Unifying Image Counterfactuals and Feature Attributions with Latent-Space Adversarial Attacks
- Manifold Induced Biases for Zero-shot and Few-shot Detection of Generated Images
- NTIRE 2025 Challenge on Real-World Face Restoration: Methods and Results
- DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
- SEGA: Drivable 3D Gaussian Head Avatar from a Single Image
- Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling
- Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
- PVLM: Parsing-Aware Vision Language Model with Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
- MLEP: Multi-granularity Local Entropy Patterns for Universal AI-generated Image Detection
- Image Editing with Diffusion Models: A Survey
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
- Hadamard product in deep learning: Introduction, Advances and Challenges
- High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
- NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: Methods and Results
- Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
- Beyond Reconstruction: A Physics Based Neural Deferred Shader for Photo-realistic Rendering
- Generative adversarial network [wikipedia]
Related