Improved Techniques for Training GANs
2016/06/10 by Tim Salimans, Salimans, Tim, Ian Goodfellow +9 · 1 voice · 487 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection #Cell Image Analysis Techniques
paper · pdf · doi:10.48550/arxiv.1606.03498
Abstract
We present a variety of new architectural features and training procedures that we apply to the generative adversarial networks (GANs) framework. We focus on two applications of GANs: semi-supervised learning, and the generation of images that humans find visually realistic. Unlike most work on generative models, our primary goal is not to train a model that assigns high likelihood to test data, nor do we require the model to be able to learn well without using any labels. Using our new techniques, we achieve state-of-the-art results in semi-supervised classification on MNIST, CIFAR-10 and SVHN. The generated images are of high quality as confirmed by a visual Turing test: our model generates MNIST samples that humans cannot distinguish from real data, and CIFAR-10 samples that yield a human error rate of 21.3%. We also present ImageNet samples with unprecedented resolution and show that our methods enable the model to learn recognizable features of ImageNet classes.
Cited by
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- dRAE: Representation Autoencoder with Hyper-Spherical Codes
- DriftXpress: Faster Drifting Models via Projected RKHS Fields
- Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion Models
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Conditioning Residuals for Diffusion Models via Representation Feedback
- Statistical Inference for Generative Model Comparison
- VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- Thinking in Video: Can Video Generators Really Reason About the Real World?
- FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting
- Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches
- Signed Rectified Flow: Negativity-Controlled Generation
- DiFA: Inference-Time Forward-Process Alignment for Diffusion Models
- Feature-Guided Diffusion for Non-Differentiable Inverse Rendering
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
- Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
- Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- Training-Free Metrics for Synthetic Object Detection Data: A Proxy for Detector Performance
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- Scaling quantum machine learning without tricks: full-resolution and diverse image generation
- Speedrunning ImageNet Diffusion
- Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
- Plug In, Grade Right: Psychology-Inspired AGIQA
- AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
- Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- SkinGenBench: Generative Model and Preprocessing Effects for Synthetic Dermoscopic Augmentation in Melanoma Diagnosis
- Generative diffusion models for agricultural AI: plant image generation, indoor-to-outdoor translation, and expert preference alignment
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- Interpretable Similarity of Synthetic Image Utility
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- Spherical Leech Quantization for Visual Tokenization and Generation
- Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
- OUSAC: Optimized Guidance Scheduling with Adaptive Caching for DiT Acceleration
- RecTok: Reconstruction Distillation along Rectified Flow
- Differentiable Energy-Based Regularization in GANs: A Simulator-Based Exploration of VQE-Inspired Auxiliary Losses
- TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Bidirectional Normalizing Flow: From Data to Noise and Back
- What matters for Representation Alignment: Global Information or Spatial Structure?
- AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- VABench: A Comprehensive Benchmark for Audio-Video Generation
- A Novel Wasserstein Quaternion Generative Adversarial Network for Color Image Generation
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Training-Free Vector Quantization via Gaussian VAEs
- JoPano: Unified Panorama Generation via Joint Modeling
- TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
- Latent Nonlinear Denoising Score Matching for Enhanced Learning of Structured Distributions
- Evaluating and Preserving High-level Fidelity in Super-Resolution
- HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
- ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
- Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Spatiotemporal Satellite Image Downscaling with Transfer Encoders and Autoregressive Generative Models
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- DPAC: Distribution-Preserving Adversarial Control for Diffusion Sampling
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- The Collapse of Patches
- Temporal Generative Adversarial Nets with Singular Value Clipping
- HVAdam: A Full-Dimension Adaptive Optimizer
- Flow Map Distillation Without Data
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- Q-Save: Towards Scoring and Attribution for Generated Video Evaluation
- DiP: Taming Diffusion Models in Pixel Space
- VeCoR -- Velocity Contrastive Regularization for Flow Matching
- TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis
- Spanning Tree Autoregressive Visual Generation
- Warm Diffusion: Recipe for Blur-Noise Mixture Diffusion Models
- Personalized Reward Modeling for Text-to-Image Generation
- Decoupling Complexity from Scale in Latent Diffusion Model
- Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in vitro
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- Aligning Generative Music AI with Human Preferences: Methods and Challenges
- Photographic Image Synthesis with Cascaded Refinement Networks
- Coffee: Controllable Diffusion Fine-tuning
- Beyond Inferring Class Representatives: User-Level Privacy Leakage From Federated Learning
- SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design
- FoleyBench: A Benchmark For Video-to-Audio Models
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- Distribution Matching Distillation Meets Reinforcement Learning
- Disentangling Pose from Appearance in Monochrome Hand Images
- Training Generative Adversarial Networks with Limited Data
- Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis
- PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
- Hybrid VAE: Improving Deep Generative Models using Partial Observations
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
- Rectified Noise: A Generative Model Using Positive-incentive Noise
- AnoStyler: Text-Driven Localized Anomaly Generation via Lightweight Style Transfer
- Test-Time Iterative Error Correction for Efficient Diffusion Models
- KLASS: KL-Guided Fast Inference in Masked Diffusion Models
- Knowledge-based anomaly detection for identifying network-induced shape artifacts
- A-NICE-MC: Adversarial Training for MCMC
- Scalable Private Learning with PATE
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
- Distributional Evaluation of Generative Models via Relative Density Ratio
- Likelihood Estimation for Generative Adversarial Networks
- Detecting GAN-generated Imagery using Color Cues
- Imitation with Neural Density Models
- GestureGAN for Hand Gesture-to-Gesture Translation in the Wild
- MSG-GAN: Multi-Scale Gradients for Generative Adversarial Networks
- Replay anti-spoofing countermeasure based on data augmentation with post selection
- Generating Multi-label Discrete Patient Records using Generative Adversarial Networks
- Differentially Private Generative Adversarial Network
- Learning to Discover Cross-Domain Relations with Generative Adversarial Networks
- Energy Models for Better Pseudo-Labels: Improving Semi-Supervised Classification with the 1-Laplacian Graph Energy
- An Introduction to Image Synthesis with Generative Adversarial Nets
- Improved Training of Wasserstein GANs
- Continual Learning for Robotics: Definition, Framework, Learning\n Strategies, Opportunities and Challenges
- Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning
- A Novel Framework for Selection of GANs for an Application
- Model Extraction and Defenses on Generative Adversarial Networks
- Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
- Learning to Protect Communications with Adversarial Neural Cryptography
- Sparsely Grouped Multi-task Generative Adversarial Networks for Facial Attribute Manipulation
- Deep Voice 2: Multi-Speaker Neural Text-to-Speech
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Generative OpenMax for Multi-Class Open Set Classification
- Generating Synthetic Multispectral Satellite Imagery from Sentinel-2
- Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning
- Fashion-Gen: The Generative Fashion Dataset and Challenge
- Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
- A Comprehensive Survey of Neural Architecture Search
- Real-Time Adaptive Image Compression
- Mode Collapse and Regularity of Optimal Transportation Maps
- MidiNet: A Convolutional Generative Adversarial Network for Symbolic-domain Music Generation
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data
- On the Origin of Deep Learning
- Channel-Recurrent Autoencoding for Image Modeling
- Invisible Steganography via Generative Adversarial Networks
- A Topology Layer for Machine Learning
- Spectral Normalization for Generative Adversarial Networks
- MIND: Monge Inception Distance for Generative Models Evaluation
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Modeling EEG data distribution with a Wasserstein Generative Adversarial Network to predict RSVP Events
- Unfolding with Generative Adversarial Networks
- On the regularization of Wasserstein GANs
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Generative View Stitching
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- DiffusionX: Efficient Edge-Cloud Collaborative Image Generation with Multi-Round Prompt Evolution
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
- FARMER: Flow AutoRegressive Transformer over Pixels
- Globally and locally consistent image completion
- Nested AutoRegressive Models
- SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
- Generalization and Memorization: The Bias Potential Model
- Semi-Supervised Learning under General Causal Models
- Scaling Non-Parametric Sampling with Representation
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- VISTA: A Test-Time Self-Improving Video Generation Agent
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- MEIcoder: Decoding Visual Stimuli from Neural Activity by Leveraging Most Exciting Inputs
- AccuQuant: Simulating Multiple Denoising Steps for Quantizing Diffusion Models
- IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial Networks
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- The Reality Gap in Robotics: Challenges, Solutions, and Best Practices
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
- GP-GAN: Towards Realistic High-Resolution Image Blending
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Signature Kernel Scoring Rule: A Spatio-Temporal Diagnostic for Probabilistic Weather Forecasting
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Demystifying Transition Matching: When and Why It Can Beat Flow Matching
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- L2P: Unlocking Latent Potential for Pixel Generation
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- One-step Diffusion Models with Bregman Density Ratio Matching
- Region in Context: Text-condition Image editing with Human-like semantic reasoning
- Cost Savings from Automatic Quality Assessment of Generated Images
- Unbiased Auxiliary Classifier GANs with MINE
- Counting Hallucinations in Diffusion Models
- DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
- LayerSync: Self-aligning Intermediate Layers
- Say No to the Discrimination: Learning Fair Graph Neural Networks with Limited Sensitive Attribute Information
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- Diffusion Transformers with Representation Autoencoders
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Robust Learning of Diffusion Models with Extremely Noisy Conditions
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Prediction Under Uncertainty with Error-Encoding Networks
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- VideoVerse: How Far is Your T2V Generator from a World Model?
- Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
- ASBench: Image Anomalies Synthesis Benchmark for Anomaly Detection
- A Distributed Training Algorithm of Generative Adversarial Networks with Quantized Gradients
- Biology-driven assessment of deep learning super-resolution imaging of the porosity network in dentin
- MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- Securing generative artificial intelligence with parallel magnetic tunnel junction true randomness
- Heptapod: Language Modeling on Visual Signals
- SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- TAG:Tangential Amplifying Guidance for Hallucination-Resistant Diffusion Sampling
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- Bridging Text and Video Generation: A Survey
- ObCLIP: Oblivious CLoud-Device Hybrid Image Generation with Privacy Preservation
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- A gradual, semi-discrete approach to generative network training via\n explicit Wasserstein minimization
- Learning Robust Diffusion Models from Imprecise Supervision
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- Visual Self-Refinement for Autoregressive Models
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
- Out-of-Sample Testing for GANs
- Learn to Guide Your Diffusion Model
- Some Theoretical Properties of GANs
- CODED-SMOOTHING: Coding Theory Helps Generalization
- Video Object Segmentation-Aware Audio Generation
- EnScale: Temporally-consistent multivariate generative downscaling via proper scoring rules
- MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
- Reevaluating Convolutional Neural Networks for Spectral Analysis: A Focus on Raman Spectroscopy
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- One-shot Conditional Sampling: MMD meets Nearest Neighbors
- VAGUEGAN: Stealthy Poisoning and Backdoor Attacks on Image Generative Pipelines
- VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning
- Training-Free Multimodal Guidance for Video to Audio Generation
- UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- General-to-Detailed GAN for Infrequent Class Medical Images
- OAT-FM: Optimal Acceleration Transport for Improved Flow Matching
- AudioMoG: Guiding Audio Generation with Mixture-of-Guidance
- BinGAN: Learning Compact Binary Descriptors with a Regularized GAN
- Manifold Adversarial Learning
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Scale-Wise VAR is Secretly Discrete Diffusion
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- What Happens Next? Anticipating Future Motion by Generating Point Trajectories
- No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
- Score-based Idempotent Distillation of Diffusion Models
- Chemical Structure Elucidation from Mass Spectrometry by Matching Substructures
- Deterministic Discrete Denoising
- Generative Adversarial Networks Applied for Privacy Preservation in Biometric-Based Authentication and Identification
- Annealed Generative Adversarial Networks
- BEGAN: Boundary Equilibrium Generative Adversarial Networks
- Improved Training with Curriculum GANs
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- SPG-Net: Segmentation Prediction and Guidance Network for Image Inpainting
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Adversarial Symmetric Variational Autoencoder
- ExtremeWeather: A large-scale climate dataset for semi-supervised\n detection, localization, and understanding of extreme weather events
- StereoFoley: Object-Aware Stereo Audio Generation from Video
- TF-Replicator: Distributed Machine Learning for Researchers
- Echo-Path: Pathology-Conditioned Echo Video Generation
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- A Survey on Evaluation Metrics for Music Generation
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Structured Information for Improving Spatial Relationships in Text-to-Image Generation
- LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition
- Triple Generative Adversarial Nets
- Generative Adversarial Residual Pairwise Networks for One Shot Learning
- Stabilizing GAN Training with Multiple Random Projections
- Noise-Level Diffusion Guidance: Well Begun is Half Done
- Generative Consistency Models for Estimation of Kinetic Parametric Image Posteriors in Total-Body PET
- Image Realness Assessment and Localization with Multimodal Features
- ReTrack: Data Unlearning in Diffusion Models through Redirecting the Denoising Trajectory
- MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling
- Learning Face Age Progression: A Pyramid Architecture of GANs
- Variational Walkback: Learning a Transition Operator as a Stochastic\n Recurrent Net
- Diffusion Models Beat GANs on Image Synthesis
- Double Helix Diffusion for Cross-Domain Anomaly Image Generation
- Adaptive Sampling Scheduler
- Maximum-Likelihood Augmented Discrete Generative Adversarial Networks
- Image Tokenizer Needs Post-Training
- Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis
- Learning Majority-to-Minority Transformations with MMD and Triplet Loss for Imbalanced Classification
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- Synthetic Dataset Evaluation Based on Generalized Cross Validation
- On the Effectiveness of Least Squares Generative Adversarial Networks
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Few-Shot Adversarial Domain Adaptation
- Fréchet ChemNet Distance: A metric for generative models for molecules in drug discovery
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Transparency of medical artificial intelligence systems
- Language Self-Play For Data-Free Training
- DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- Universal Few-Shot Spatial Control for Diffusion Models
- Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
- Breast Cancer Detection in Thermographic Images via Diffusion-Based Augmentation and Nonlinear Feature Fusion
- Autoregressive Quantile Networks for Generative Modeling
- Deep learning-based phase prediction of high-entropy alloys: Optimization, generation, and explanation
- Batch Renormalization: Towards Reducing Minibatch Dependence in Batch-Normalized Models
- Pixel-wise Conditioned Generative Adversarial Networks for Image\n Synthesis and Completion
- Transition Models: Rethinking the Generative Learning Objective
- Least Squares Generative Adversarial Networks
- Sharp Minima Can Generalize For Deep Nets
- Privacy-Utility Trade-off in Data Publication: A Bilevel Optimization Framework with Curvature-Guided Perturbation
- VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
- Implicit Maximum Likelihood Estimation
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Semi-Supervised Bayesian GANs with Log-Signatures for Uncertainty-Aware Credit Card Fraud Detection
- The relativistic discriminator: a key element missing from standard GAN
- Good Semi-supervised Learning that Requires a Bad GAN
- Weighted Support Points from Random Measures: An Interpretable Alternative for Generative Modeling
- Physics Informed Generative Models for Magnetic Field Images
- CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
- UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- Survey on Deep Multi-modal Data Analytics: Collaboration, Rivalry and Fusion
- SenseGen: A Deep Learning Architecture for Synthetic Sensor Data\n Generation
- Quantum latent distributions in deep generative models
- Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
- Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator Correction
- Cross-Domain Few-Shot Learning by Representation Fusion
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
- FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction
- Pretrained Diffusion Models Are Inherently Skipped-Step Samplers
- CurveFlow: Curvature-Guided Flow Matching for Image Generation
- DeepCP: Deep Learning Driven Cascade Prediction Based Autonomous Content Placement in Closed Social Network
- SATURN: Autoregressive Image Generation Guided by Scene Graphs
- Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
- Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
- Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed
- Semantically Consistent Image Completion with Fine-grained Details
- FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- A Tutorial on Deep Latent Variable Models of Natural Language
- Distribution Matching via Generalized Consistency Models
- Invert and Defend: Model-based Approximate Inversion of Generative Adversarial Networks for Secure Inference
- TequilaGAN: How to easily identify GAN samples
- Inverting The Generator Of A Generative Adversarial Network (II)
- Training-Free Anomaly Generation via Dual-Attention Enhancement in Diffusion Model
- Better Supervised Fine-tuning for VQA: Integer-Only Loss
- Every Smile is Unique: Landmark-Guided Diverse Smile Generation
- Toward Diverse Text Generation with Inverse Reinforcement Learning
- Self-Supervised Temporal Super-Resolution of Energy Data using Generative Adversarial Transformer
- Semi-Supervised Learning Enabled by Multiscale Deep Neural Network Inversion
- Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Prototype-Guided Diffusion: Visual Conditioning without External Memory
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- Generation of Indian Sign Language Letters, Numbers, and Words
- Improving GAN Training via Binarized Representation Entropy (BRE)\n Regularization
- Attacks on State-of-the-Art Face Recognition using Attentional Adversarial Attack Generative Network
- Simulation of Charge Stability Diagrams for Automated Tuning Solutions (SimCATS)
- Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach
- Reinforcement Learning for Large Model: A Survey
- Π-nets: Deep Polynomial Neural Networks
- Score Augmentation for Diffusion Models
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- Improved Training for Self-Training by Confidence Assessments
- Composite Functional Gradient Learning of Generative Adversarial Models
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Clinically-guided Data Synthesis for Laryngeal Lesion Detection
- Synthetic Data Generation for Emotional Depth Faces: Optimizing Conditional DCGANs via Genetic Algorithms in the Latent Space and Stabilizing Training with Knowledge Distillation
- Score-Guided Generative Adversarial Networks
- Generative Adversarial Network Architectures For Image Synthesis Using Capsule Networks
- Deep Convolutional Generative Adversarial Network Based Food Recognition\n Using Partially Labeled Data
- Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation
- DISSECT: Disentangled Simultaneous Explanations via Concept Traversals
- Reference-free Adversarial Sex Obfuscation in Speech
- An empirical study on evaluation metrics of generative adversarial networks
- Dataset Condensation with Color Compensation
- Synthesis of High-Quality Visible Faces from Polarimetric Thermal Faces using Generative Adversarial Networks
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
- Towards Deeper Generative Architectures for GANs using Dense connections
- Metric Learning-based Generative Adversarial Network
- PixNerd: Pixel Neural Field Diffusion
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
- Trade-offs in Image Generation: How Do Different Dimensions Interact?
- Model-Agnostic Gender Bias Control for Text-to-Image Generation via Sparse Autoencoder
- Harnessing Diffusion-Yielded Score Priors for Image Restoration
- JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
- DINO: A Conditional Energy-Based GAN for Domain Translation
- ReDi: Rectified Discrete Flow
- Deep learning in ECG diagnosis: A review
- Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis
- Adversarial Distillation of Bayesian Neural Network Posteriors
- Customizing an Adversarial Example Generator with Class-Conditional GANs
- Adversarial and Perceptual Refinement for Compressed Sensing MRI Reconstruction
- Generative Image Inpainting with Contextual Attention
- Feedback GAN (FBGAN) for DNA: a Novel Feedback-Loop Architecture for Optimizing Protein Functions
- G3AN: Disentangling Appearance and Motion for Video Generation
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- Curriculum Learning for Deep Generative Models with Clustering
- T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
- A Comprehensive Review of Diffusion Models in Smart Agriculture: Progress, Applications, and Challenges
- McGan: Mean and Covariance Feature Matching GAN
- Predicting Opioid Relapse Using Social Media Data
- MSGDD-cGAN: Multi-Scale Gradients Dual Discriminator Conditional Generative Adversarial Network
- Generalized Latent Variable Recovery for Generative Adversarial Networks
- Shape Generation using Spatially Partitioned Point Clouds
- Performing Co-Membership Attacks Against Deep Generative Models
- Learning Continuous Face Age Progression: A Pyramid of GANs
- A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis
- Megapixel Size Image Creation using Generative Adversarial Networks
- What Is It Like Down There? Generating Dense Ground-Level Views and Image Features From Overhead Imagery Using Conditional Generative Adversarial Networks
- One-step Latent-free Image Generation with Pixel Mean Flows
- Zero-shot Domain Adaptation without Domain Semantic Descriptors
- Packet-Level Adversarial Network Traffic Crafting using Sequence Generative Adversarial Networks
- Synthetic-to-Real Domain Adaptation for Lane Detection
- Boosting Deep Learning Risk Prediction with Generative Adversarial Networks for Electronic Health Records
- Towards Visible and Thermal Drone Monitoring with Convolutional Neural Networks
- DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
- Attention (as Discrete-Time Markov) Chains
- Bringing Balance to Hand Shape Classification: Mitigating Data Imbalance Through Generative Models
- Maximum Mutation Reinforcement Learning for Scalable Control
- Generative Adversarial Networks and Probabilistic Graph Models for Hyperspectral Image Classification
- Generative Creativity: Adversarial Learning for Bionic Design
- TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
- A Distributed Generative AI Approach for Heterogeneous Multi-Domain Environments under Data Sharing constraints
- Interactive Image Manipulation with Natural Language Instruction Commands
- Novelty Detection with GAN
- GANs for Medical Image Analysis
- A Survey on Generative Adversarial Networks: Variants, Applications, and\n Training
- DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
- Wasserstein Introspective Neural Networks
- Towards Robust Neural Machine Translation
- Adversarial Training Methods for Network Embedding
- Introspective Classification with Convolutional Nets
- Supervised Adversarial Networks for Image Saliency Detection
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- Optimal transport maps for distribution preserving operations on latent\n spaces of Generative Models
- Smooth Neighbors on Teacher Graphs for Semi-supervised Learning
- GILT: Generating Images from Long Text
- Multi-View Image Generation from a Single-View
- Question-Conditioned Counterfactual Image Generation for VQA
- Dual Discriminator Generative Adversarial Nets
- Generative Adversarial Network for Handwritten Text
- CIGLI: Conditional Image Generation from Language & Image
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
- Towards Principled Methods for Training Generative Adversarial Networks
- An Adversarial Regularisation for Semi-Supervised Training of Structured Output Neural Networks
- Learning to Sketch with Shortcut Cycle Consistency
- CATVis: Context-Aware Thought Visualization
- Deep Co-Training for Semi-Supervised Image Recognition
- Text Embedding Knows How to Quantize Text-Guided Diffusion Models
- Frequency Regulation for Exposure Bias Mitigation in Diffusion Models
- Latent Diffusion Models with Masked AutoEncoders
- Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-scarce Scenarios
- Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow
- Learning Private Representations through Entropy-based Adversarial Training
- Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial Inspection
- Learning Diffusion Models with Flexible Representation Guidance
- From source to target and back: symmetric bi-directional adaptive GAN
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- Generative adversarial network [wikipedia]
- Inception score [wikipedia]
- Synthetic media [wikipedia]
Discussions
Related