A Style-Based Generator Architecture for Generative Adversarial Networks
2018/12/12 by Tero Karras, Samuli Laine, Karras, Tero +3 · 2 voices · 474 citations
#cs.NE #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1812.04948
Abstract
We propose an alternative generator architecture for generative adversarial networks, borrowing from style transfer literature. The new architecture leads to an automatically learned, unsupervised separation of high-level attributes (e.g., pose and identity when trained on human faces) and stochastic variation in the generated images (e.g., freckles, hair), and it enables intuitive, scale-specific control of the synthesis. The new generator improves the state-of-the-art in terms of traditional distribution quality metrics, leads to demonstrably better interpolation properties, and also better disentangles the latent factors of variation. To quantify interpolation quality and disentanglement, we propose two new, automated methods that are applicable to any generator architecture. Finally, we introduce a new, highly varied and high-quality dataset of human faces.
Cited by
- Coherent Visualization of 2D Scalar Field Contour Ensembles With Probabilistic Latent Space Modeling
- Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
- Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
- Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
- Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution
- Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
- Show Me Examples: Inferring Visual Concepts from Image Sets
- Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
- To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations
- Fundamental Recovery Bounds for SPAD Signals under Stationary Flux
- Image Editing Models are Numerical Solvers
- GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors
- ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
- Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
- Incomplete Observations Boost Evolutionary Performance in Ocean Modeling
- LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
- Consistent Feature Transport for Image Relighting
- PhantomSeal: Proactive Deepfakes Defense with Identity/Context Protection and Forensic Tracing
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
- STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
- Physics-aware Masked Diffusion-based Flood Simulation for Urban Fisheye Disaster Detection
- Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
- HairPort: In-context 3D-aware Hair Import and Transfer for Images
- NAMESAKES: Probing Identity Memorization in Text-to-Image Models
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- Scaling quantum machine learning without tricks: full-resolution and diverse image generation
- From Faces to Politics: Vision-Language Models (Sometimes) Link Visual Demographic Characteristics to Ideological Labels
- Controllable Probabilistic Forecasting with Stochastic Decomposition Layers
- Fit for Purpose? Deepfake Detection in the Real World
- FakeParts: a New Family of AI-Generated DeepFakes
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Entropic Time Schedulers for Generative Diffusion Models
- Mechanisms of Projective Composition of Diffusion Models
- DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Multi-Agent Adversarial Reinforcement Learning
- Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
- FlowLPS: Langevin-Proximal Sampling for Flow-based Inverse Problem Solvers
- SURE Guided Posterior Sampling: Trajectory Correction for Diffusion-Based Inverse Problems
- Anomaly Detection by Effectively Leveraging Synthetic Images
- Reverse Personalization
- Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
- Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
- DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models
- MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
- ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
- ProEdit: Inversion-based Editing From Prompts Done Right
- AI-generated Images Challenge Visual Trust in High-risk Scenarios
- Real-time Reconstruction of Human Visual Perception from fMRI
- Text-based Tactile Graphics Generation for the Visually Impaired
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-Resolution
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
- Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
- Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
- Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
- FUSE: Unifying Spectral and Semantic Cues for Robust AI-Generated Image Detection
- A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
- Beyond Weight Adaptation: Feature-Space Domain Injection for Cross-Modal Ship Re-Identification
- How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- Generative Latent Coding for Ultra-Low Bitrate Image Compression
- Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
- DVI: Disentangling Semantic and Visual Identity for Training-Free Personalized Generation
- Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
- Efficient Zero-Shot Inpainting with Decoupled Diffusion Guidance
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- An Empirical Study of Sampling Hyperparameters in Diffusion-Based Super-Resolution
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- Instant Expressive Gaussian Head Avatars at Over 100 FPS
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
- Generative Preprocessing for Image Compression with Pre-trained Diffusion Models
- FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake Videos
- Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
- Dual Attention Guided Defense Against Malicious Edits
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution
- Directional Textual Inversion for Personalized Text-to-Image Generation
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
- Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
- Learning Common and Salient Generative Factors Between Two Image Datasets
- Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real
- Progressive Conditioned Scale-Shift Recalibration of Self-Attention for Online Test-time Adaptation
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- PersonaLive! Expressive Portrait Image Animation for Live Streaming
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Inverse problems with diffusion models: MAP estimation via mode-seeking loss
- AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- Food Image Generation on Multi-Noun Categories
- Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps
- Towards Robust Protective Perturbation against DeepFake Face Swapping
- Understanding Diffusion Models via Code Execution
- SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
- AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head Avatars
- A Comparative Study on Synthetic Facial Data Generation Techniques for Face Recognition
- HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-Resolution
- Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity
- PGP-DiffSR: Phase-Guided Progressive Pruning for Efficient Diffusion-based Image Super-Resolution
- Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
- Data-Centric Visual Development for Self-Driving Labs
- StyleYourSmile: Cross-Domain Face Retargeting Without Paired Multi-Style Data
- Weight Space Representation Learning via Neural Field Adaptation
- FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution
- Reversible Inversion for Training-Free Exemplar-guided Image Editing
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Vision Bridge Transformer at Scale
- REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image Detection
- Adapting Neural Audio Codecs to EEG
- Overcoming the Curvature Bottleneck in MeanFlow
- Bringing Your Portrait to 3D Presence
- Adversarial Flow Models
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
- FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision
- Deep Parameter Interpolation for Scalar Conditioning
- Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
- SONIC: Spectral Optimization of Noise for Inpainting with Consistency
- Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
- Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization
- Efficient Transferable Optimal Transport via Min-Sliced Transport Plans
- Solving Diffusion Inverse Problems with Restart Posterior Sampling
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- DiP: Taming Diffusion Models in Pixel Space
- Rethinking Garment Conditioning in Diffusion-based Virtual Try-On
- TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
- Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image Detection
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- UnfoldLDM: Deep Unfolding-based Blind Image Restoration with Latent Diffusion Priors
- Range-Edit: Semantic Mask Guided Outdoor LiDAR Scene Editing
- FLUID: Training-Free Face De-identification via Latent Identity Substitution
- One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing
- Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement
- Motion Transfer-Enhanced StyleGAN for Generating Diverse Macaque Facial Expressions
- How Noise Benefits AI-generated Image Detection
- Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
- APT: Affine Prototype-Timestamp For Time Series Forecasting Under Distribution Shift
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- Birth of a Painting: Differentiable Brushstroke Reconstruction
- MiAD: Mirage Atom Diffusion for De Novo Crystal Generation
- AI-driven Generation of MALDI-TOF MS for Microbial Characterization
- FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
- Training-free Detection of AI-generated images via Cropping Robustness
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection
- Functional Mean Flow in Hilbert Space
- SAGA: Source Attribution of Generative AI Videos
- ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
- SecDiff: Diffusion-Aided Secure Deep Joint Source-Channel Coding Against Adversarial Attacks
- Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis
- PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
- Which Way from B to A: The role of embedding geometry in image interpolation for Stable Diffusion
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- Diffusion Model Based Signal Recovery Under 1-Bit Quantization
- Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
- Model Inversion Attack Against Deep Hashing
- Null-Space Diffusion Distillation Unlocks Speed, Fidelity and Realism in Lensless Imaging
- StyleQoRA: Quality-Aware Low-Rank Adaptation for Few-Shot Multi-Style Editing
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- Detecting Generated Images by Fitting Natural Image Distributions
- Sinkhorn-Drifting Generative Models
- Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
- LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors
- Contact Wasserstein Geodesics for Non-Conservative Schrödinger Bridges
- Integrating Reweighted Least Squares with Plug-and-Play Diffusion Priors for Noisy Image Restoration
- AvatarTex: High-Fidelity Facial Texture Reconstruction from Single-Image Stylized Avatars
- AnoStyler: Text-Driven Localized Anomaly Generation via Lightweight Style Transfer
- Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- Latent Refinement via Flow Matching for Training-free Linear Inverse Problem Solving
- Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
- Diffusion Transformer meets Multi-level Wavelet Spectrum for Single Image Super-Resolution
- Generative Hints
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- A Non-Adversarial Approach to Idempotent Generative Modelling
- A New Perspective on Precision and Recall for Generative Models
- DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
- PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Detecting AI-Generated Images via Diffusion Snap-Back Reconstruction: A Forensic Approach
- Training human super-recognizers’ detection and discrimination of AI-generated faces
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- Bayesian model selection and misspecification testing in imaging inverse problems only from noisy and partial measurements
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Self-Diffusion Driven Blind Imaging
- End-to-End Framework Integrating Generative AI and Deep Reinforcement Learning for Autonomous Ultrasound Scanning
- Blind MIMO Semantic Communication via Parallel Variational Diffusion: A Completely Pilot-Free Approach
- Beyond Data Scarcity Optimizing R3GAN for Medical Image Generation from Small Datasets
- Instruction-based image editing: a survey on data, models, evaluation, and applications
- Face De-Identification: A Domain-Centric Survey from Capture to Processing
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- A Theory of Contrastive Learning with Natural Images
- VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
- Product-Quantised Image Representation for High-Quality Image Synthesis
- Amortized Moment Matching for Visual Generation
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis
- PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
- Beyond Inference Intervention: Identity-Decoupled Diffusion for Face Anonymization
- An efficient probabilistic hardware architecture for diffusion-like models
- Unmasking Facial DeepFakes: A Robust Multiview Detection Framework for Natural Images
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
- Coupled Flow Matching
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling
- Open Multimodal Retrieval-Augmented Factual Image Generation
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- BADiff: Bandwidth Adaptive Diffusion Model
- Noise is All You Need: Solving Linear Inverse Problems by Noise Combination Sampling with Diffusion Models
- On the flow matching interpretability
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- LayerComposer: Multi-Human Personalized Generation via Layered Canvas
- SCEESR: Semantic-Control Edge Enhancement for Diffusion-Based Super-Resolution
- Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
- Efficient Few-shot Identity Preserving Attribute Editing for 3D-aware Deep Generative Models
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- ConsistEdit: Highly Consistent and Precise Training-free Visual Editing
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- Diffusion Models as Dataset Distillation Priors
- Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
- GUIDE: Enhancing Gradient Inversion Attacks in Federated Learning with Denoising Models
- Latent Diffusion Model without Variational Autoencoder
- WithAnyone: Towards Controllable and ID Consistent Image Generation
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Learning Human Motion with Temporally Conditional Mamba
- Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration
- Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
- MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- CoLoR-GAN: Continual Few-Shot Learning with Low-Rank Adaptation in Generative Adversarial Networks
- Zero-shot Face Editing via ID-Attribute Decoupled Inversion
- SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
- Encoder Decoder Generative Adversarial Network Model for Stock Market Prediction
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations
- Decomposer Networks: Deep Component Analysis and Synthesis
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- HeadsUp! High-Fidelity Portrait Image Super-Resolution
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- High-dimensional Analysis of Synthetic Data Selection
- Coupled Data and Measurement Space Dynamics for Enhanced Diffusion Posterior Sampling
- Local MAP Sampling for Diffusion Models
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
- SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- CodeFormer++: Blind Face Restoration Using Deformable Registration and Deep Metric Learning
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- SDAKD: Student Discriminator Assisted Knowledge Distillation for Super-Resolution Generative Adversarial Networks
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- Learning Polynomial Activation Functions for Deep Neural Networks
- PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
- Dale meets Langevin: A Multiplicative Denoising Diffusion Model
- Image Generation Based on Image Style Extraction
- Secure and reversible face anonymization with diffusion models
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Nonparametric Identification of Latent Concepts
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Post-Training Quantization via Residual Truncation and Zero Suppression for Diffusion Models
- FLOWER: A Flow-Matching Solver for Inverse Problems
- Marginal Flow: a flexible and efficient framework for density estimation
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- CharGen: Fast and Fluent Portrait Modification
- SAIP: A Plug-and-Play Scale-adaptive Module in Diffusion-based Inverse Problems
- CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers
- NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
- Tumor Synthesis conditioned on Radiomics
- MAD: Manifold Attracted Diffusion
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Scalable GANs with Transformers
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Calibrated and Resource-Aware Super-Resolution for Reliable Driver Behavior Analysis
- Stochastic Interpolants via Conditional Dependent Coupling
- Enhancing Blind Face Restoration through Online Reinforcement Learning
- Scale-Wise VAR is Secretly Discrete Diffusion
- ConQuER: Modular Architectures for Control and Bias Mitigation in IQP Quantum Generative Models
- Bézier Meets Diffusion: Robust Generation Across Domains for Medical Image Segmentation
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- Overclocking Electrostatic Generative Models
- DragGANSpace: Latent Space Exploration and Control for GANs
- Universal Multi-Domain Translation via Diffusion Routers
- Deepfakes: we need to re-think the concept of "real" images
- DroneFL: Federated Learning for Multi-UAV Visual Target Tracking
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training
- A Real-Time On-Device Defect Detection Framework for Laser Power-Meter Sensors via Unsupervised Learning
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
- Generative Model Inversion Through the Lens of the Manifold Hypothesis
- StrCGAN: A Generative Framework for Stellar Image Restoration
- Achieving Fair Skin Lesion Detection through Skin Tone Normalization and Channel Pruning
- Collaborative feature aggregation for face super-resolution and robust re-identification
- From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
- Diffusion-Based Data Augmentation for Medical Image Segmentation
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- A Gradient Flow Approach to Solving Inverse Problems with Latent Diffusion Models
- Prompt-Guided Dual Latent Steering for Inversion Problems
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem
- PRNU-Bench: A Novel Benchmark and Model for PRNU-Based Camera Identification
- ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
- SISMA: Semantic Face Image Synthesis with Mamba
- ScenGAN: Attention-Intensive Generative Model for Uncertainty-Aware Renewable Scenario Forecasting
- FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection
- Enhancing Reference-based Sketch Colorization via Separating Reference Representations
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Local Mechanisms of Compositional Generalization in Conditional Diffusion
- LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition
- Controllable Localized Face Anonymization Via Diffusion Inpainting
- RaceGAN: A Framework for Preserving Individuality while Converting Racial Information for Image-to-Image Translation
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- A Race Bias Free Face Aging Model for Reliable Kinship Verification
- StyleProtect: Safeguarding Artistic Identity in Fine-tuned Diffusion Models
- Deceptive Beauty: Evaluating the Impact of Beauty Filters on Deepfake and Morphing Attack Detection
- Adaptive Sampling Scheduler
- Image Tokenizer Needs Post-Training
- From Autoencoders to CycleGAN: Robust Unpaired Face Manipulation via Adversarial Learning
- Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
- DRAG: Data Reconstruction Attack using Guided Diffusion
- Reconstructing High-fidelity Plasma Turbulence with Data-driven Tuning of Diffusion in Low Resolution Grids
- Beyond Sliders: Mastering the Art of Diffusion-based Image Manipulation
- Styleclone: Face Stylization with Diffusion Based Data Augmentation
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- A Novel Local Focusing Mechanism for Deepfake Detection Generalization
- A Discrepancy-Based Perspective on Dataset Condensation
- GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
- Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios
- Patch-based Automatic Rosacea Detection Using the ResNet Deep Learning Framework
- Privacy-Preserving Automated Rosacea Detection Based on Medically Inspired Region of Interest Selection
- Integrating Anatomical Priors into a Causal Diffusion Model
- PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed Image
- BIR-Adapter: A Low-Complexity Diffusion Model Adapter for Blind Image Restoration
- CardiacFlow: 3D+t Four-Chamber Cardiac Shape Completion and Generation via Flow Matching
- MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
- Missing Fine Details in Images: Last Seen in High Frequencies
- A Scalable Attention-Based Approach for Image-to-3D Texture Mapping
- From Editor to Dense Geometry Estimator
- Human Motion Video Generation: A Survey
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation
- 2D Gaussian Splatting with Semantic Alignment for Image Inpainting
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- Differentiable Expectation-Maximisation and Applications to Gaussian Mixture Model Optimal Transport
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- Expandable Residual Approximation for Knowledge Distillation
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- Deep Learning for CMB Foreground Removal and Beam Deconvolution: A U-Net GAN Approach
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Weighted Support Points from Random Measures: An Interpretable Alternative for Generative Modeling
- Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
- AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View Images
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- Quantum latent distributions in deep generative models
- Synthetic Image Detection via Spectral Gaps of QC-RBIM Nishimori Bethe-Hessian Operators
- Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
- RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration
- FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained Poses
- BrainPath: A Biologically-Informed AI Framework for Individualized Aging Brain Generation
- Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering
- DiffIER: Optimizing Diffusion Models with Iterative Error Reduction
- TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
- From Classical Probabilistic Latent Variable Models to Modern Generative AI: A Unified Perspective
- REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language
- Leveraging Diffusion Models for Stylization using Multiple Style Images
- Distribution Matching via Generalized Consistency Models
- Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs
- TimeMachine: Fine-Grained Facial Age Editing with Identity Preservation
- StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- NegFaceDiff: The Power of Negative Context in Identity-Conditioned Diffusion for Synthetic Face Generation
- CLIP-Flow: A Universal Discriminator for AI-Generated Images Inspired by Anomaly Detection
- Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection
- Personalized Face Super-Resolution with Identity Decoupling and Fitting
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- Unlocking the Potential of Diffusion Priors in Blind Face Restoration
- Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
- Score Augmentation for Diffusion Models
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
- OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- MotionSwap
- FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
- FLUX-Makeup: High-Fidelity, Identity-Consistent, and Robust Makeup Transfer via Diffusion Transformer
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- UniTalker: Conversational Speech-Visual Synthesis
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection
- Deeper Inside Deep ViT
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- DogFit: Domain-guided Fine-tuning for Efficient Transfer Learning of Diffusion Models
- Anti-Tamper Protection for Unauthorized Individual Image Generation
- Injecting Measurement Information Yields a Fast and Noise-Robust Diffusion-Based Inverse Problem Solver
- FFHQ-Makeup: Paired Synthetic Makeup Dataset with Facial Consistency Across Multiple Styles
- A Closed-Loop Multi-Agent Framework for Aerodynamics-Aware Automotive Styling Design
- Subject or Style: Adaptive and Training-Free Mixture of LoRAs
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- LeakyCLIP: Extracting Training Data from CLIP
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- iLRM: An Iterative Large 3D Reconstruction Model
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Visual Language Models as Zero-Shot Deepfake Detectors
- Trade-offs in Image Generation: How Do Different Dimensions Interact?
- Quantum generative modeling for financial time series with temporal correlations
- Evaluating Deepfake Detectors in the Wild
- Suppressing Gradient Conflict for Generalizable Deepfake Detection
- Locally Controlled Face Aging with Latent Diffusion Models
- VoluMe -- Authentic 3D Video Calls from Live Gaussian Splat Prediction
- History of artificial neural networks [wikipedia]
- Fréchet inception distance [wikipedia]
- StyleGAN [wikipedia]
- List of datasets in computer vision and image processing [wikipedia]
- Generative adversarial network [wikipedia]
Discussions
Related