Generative Adversarial Networks
2014/06/10 by Ian Goodfellow, Ian J. Goodfellow, Goodfellow, Ian J. +14 · 7 voices · 4,605 citations
Computer Science · Mathematics · #Artificial intelligence #Artificial neural network #Computer science #Discriminative model #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Image Processing and 3D Reconstruction #Inference #Machine learning #Mathematical optimization #Mathematics #Minimax #Mistake #Music and Audio Processing #Perceptron #Sample (material) #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1406.2661
published in arXiv (Cornell University) (Cornell University)
arxiv created 2014/06/10 · openalex publication_date 2014/06/10 · arxiv updated 2014/06/11 · openalex created_date 2022/10/01 · openalex updated_date 2026/08/06
Abstract
Large Language Models (LLMS) rely on Key-Value (KV) caches to store attention context during autoregressive decoding. In long-sequence settings, the KV cache can consume large amounts of VRAM and become a practical bottleneck for throughput . We introduce KVHALO, an auxiliary reconstruction model that restores higher-fidelity KV tensors from a compressed cache state when required, reducing persistent memory footprint during inference. In our evaluation, KVHALO achieves up to 91.85% directional cosine alignment at convergence and reduces long-context degradation relative to a low-bit baseline under our stress-test workloads. We used HRM instead of other architectures, which allowed for higher-quality results in only 18,600 steps.
Cited by
- InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation
- Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
- FUSE: Unifying Spectral and Semantic Cues for Robust AI-Generated Image Detection
- Detection of AI Generated Images Using Combined Uncertainty Measures and Particle Swarm Optimised Rejection Mechanism
- Grad: Guided Relation Diffusion Generation for Graph Augmentation in Graph Fraud Detection
- Vision-Language Model Guided Image Restoration
- AI-Augmented Pollen Recognition in Optical and Holographic Microscopy for Veterinary Imaging
- A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
- A Conditional Generative Framework for Synthetic Data Augmentation in Segmenting Thin and Elongated Structures in Biological Images
- Provable Long-Range Benefits of Next-Token Prediction
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- Advanced Unsupervised Learning: A Comprehensive Overview of Multi-View Clustering Techniques
- Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models
- One-Step Generative Channel Estimation via Average Velocity Field
- Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade
- SD-CGAN: Conditional Sinkhorn Divergence GAN for DDoS Anomaly Detection in IoT Networks
- Three-Dimensional Anatomical Data Generation Based on Artificial Neural Networks
- From Healthy Scans to Annotated Tumors: A Tumor Fabrication Framework for 3D Brain MRI Synthesis
- HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
- BrainNormalizer: Anatomy-Informed Pseudo-Healthy Brain Reconstruction from Tumor MRI via Edge-Guided ControlNet
- DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
- Machine learning-driven modeling framework for steam co-gasification applications
- Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
- A Survey of Heterogeneous Graph Neural Networks for Cybersecurity Anomaly Detection
- Reading Radio from Camera: Visually-Grounded, Lightweight, and Interpretable RSSI Prediction
- Manipulate as Human: Learning Task-oriented Manipulation Skills by Adversarial Motion Priors
- DPGLA: Bridging the Gap between Synthetic and Real Data for Unsupervised Domain Adaptation in 3D LiDAR Semantic Segmentation
- DDTR: Diffusion Denoising Trace Recovery
- Collaborative penetration testing suite for emerging generative AI algorithms
- Optimizing DINOv2 with Registers for Face Anti-Spoofing
- Conditional Synthetic Live and Spoof Fingerprint Generation
- A Multi-domain Image Translative Diffusion StyleGAN for Iris Presentation Attack Detection
- Global-focal Adaptation with Information Separation for Noise-robust Transfer Fault Diagnosis
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Using neural style transfer to study the evolution of animal signal design: A case study in an ornamented fish
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- Denoised Diffusion for Object-Focused Image Augmentation
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- cryoTIGER: deep-learning based tilt interpolation generator for enhanced reconstruction in cryo electron tomography
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
- ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
- Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- DiffCamera: Arbitrary Refocusing on Images
- Non-Parametric Simulation of Multivariate Extreme Events via Spectral Bootstrap
- MapGenerator: a framework for learning a diffusion model for text promptable map generation
- Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
- A Biophysical-Model-Informed Source Separation Framework For EMG Decomposition
- Pure Node Selection for Imbalanced Graph Node Classification
- Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
- Modelling and design of transcriptional enhancers
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- Red Teaming Quantum-Resistant Cryptographic Standards: A Penetration Testing Framework Integrating AI and Quantum Security
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
- Latent Conditioned Loco-Manipulation Using Motion Priors
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Generative AI for Misalignment-Resistant Virtual Staining to Accelerate Histopathology Workflows
- Inverse Design of Amorphous Materials with Targeted Properties
- Autonomous Intelligent Monitoring of Photovoltaic Systems: An In‐Depth Multidisciplinary Review
- Self-Evolving LLMs via Continual Instruction Tuning
- Synthetic Dataset Evaluation Based on Generalized Cross Validation
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Privacy-Preserving Automated Rosacea Detection Based on Medically Inspired Region of Interest Selection
- RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- BIR-Adapter: A Low-Complexity Diffusion Model Adapter for Blind Image Restoration
- Systematic Review and Meta-analysis of AI-driven MRI Motion Artifact Detection and Correction
- Set Transformer Architectures and Synthetic Data Generation for Flow-Guided Nanoscale Localization
- From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
- Review on Convolutional Neural Network (CNN) Applied to Plant Leaf Disease Classification
- University students' perceptions on how generative artificial intelligence shape learning and research practices: A case study in Hong Kong
- Semi-Supervised Bayesian GANs with Log-Signatures for Uncertainty-Aware Credit Card Fraud Detection
- Tabular Diffusion Counterfactual Explanations
- Generative AI for Testing of Autonomous Driving Systems: A Survey
- Mining gold from implicit models to improve likelihood-free inference
- From Sound to Sight: Towards AI-authored Music Videos
- DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
- Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
- On the Generalization Limits of Quantum Generative Adversarial Networks with Pure State Generators
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
- Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Deep learning for Alzheimer's disease diagnosis: A survey
- FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
- LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
- Energy-Efficient Federated Learning for Edge Real-Time Vision via Joint Data, Computation, and Communication Design
- Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data
- TopoLiDM: Topology-Aware LiDAR Diffusion Models for Interpretable and Realistic LiDAR Point Cloud Generation
- GeMix: Conditional GAN-Based Mixup for Improved Medical Image Augmentation
- Deep learning in ECG diagnosis: A review
- From South to North: Leveraging Ground-Based LATs for Full-Sky CMB Delensing and Constraints on r
- Quantum-Informed Machine Learning for Predicting Spatiotemporal Chaos
- VAE-GAN Based Price Manipulation in Coordinated Local Energy Markets
- Towards Urban Planing AI Agent in the Age of Agentic AI
- On Inductive Biases for Machine Learning in Data Constrained Settings
- Flow Matching Meets Biology and Life Science: A Survey
- Pyramid Hierarchical Masked Diffusion Model for Imaging Synthesis
- Using Sign Language Production as Data Augmentation to enhance Sign Language Translation
- Local Representative Token Guided Merging for Text-to-Image Generation
- An open dataset of neural networks for hypernetwork research
- Latent Space Consistency for Sparse-View CT Reconstruction
- EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation
- Modeling Partially Observed Nonlinear Dynamical Systems and Efficient Data Assimilation via Discrete-Time Conditional Gaussian Koopman Network
- Physics-Informed Neural Networks with Hard Nonlinear Equality and Inequality Constraints
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- Super-resolution of turbulent velocity fields in two-way coupled particle-laden flows
- Recourse, Repair, Reparation, & Prevention: A Stakeholder Analysis of AI Supply Chains
- Can Artificial Intelligence solve the blockchain oracle problem? Unpacking the Challenges and Possibilities
- Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels
- 3D-Telepathy: Reconstructing 3D Objects from EEG Signals
- Compute Trends Across Three Eras of Machine Learning
- TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
- Efficient Beam Selection for ISAC in Cell-Free Massive MIMO via Digital Twin-Assisted Deep Reinforcement Learning
- Deep learning-based radiointerferometric imaging with GAN-aided training
- Deepfake geography: a novel detection method for identifying manipulated satellite images
- Reversing Flow for Image Restoration
- Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments
- Recursive Variational Autoencoders for 3D Blood Vessel Generative Modeling
- Training-Free Diffusion Framework for Stylized Image Generation with Identity Preservation
- Unleashing Diffusion and State Space Models for Medical Image Segmentation
- Inverse Design of Metamaterials with Manufacturing-Guiding Spectrum-to-Structure Conditional Diffusion Model
- SPC to 3D: Novel View Synthesis from Binary SPC via I2I translation
- Exploring bidirectional bounds for minimax-training of Energy-based models
- AuthGuard: Generalizable Deepfake Detection via Language Guidance
- 3D-StyleGAN: A Style-Based Generative Adversarial Network for Generative Modeling of Three-Dimensional Medical Images
- MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
- Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector
- Enabling Probabilistic Learning on Manifolds through Double Diffusion Maps
- Adversarial Reinforcement Learning: A Duality-Based Approach To Solving Optimal Control Problems
- FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
- Bounding Box-Guided Diffusion for Synthesizing Industrial Images and Segmentation Map
- Semantics-Aware Human Motion Generation from Audio Instructions
- Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer
- A Dataset and Benchmarks for Deep Learning-Based Optical Microrobot Pose and Depth Perception
- Synthesis of Ventilator Dyssynchrony Waveforms using a Hybrid Generative Model and a Lung Model
- Unsupervised anomaly detection in MeV ultrafast electron diffraction
- SynDEc: A Synthetic Data Ecosystem
- Recent Advances in Diffusion Models for Hyperspectral Image Processing and Analysis: A Review
- Learned Lightweight Smartphone ISP with Unpaired Data
- Evaluating the robustness of adversarial defenses in malware detection systems
- AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
- Improving Generalization of Medical Image Registration Foundation Model
- Overcoming Dimensional Factorization Limits in Discrete Diffusion Models through Quantum Joint Distribution Learning
- Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
- Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
- Anomaly detection with Wasserstein GAN
- Advancing Wheat Crop Analysis: A Survey of Deep Learning Approaches Using Hyperspectral Imaging
- Vision Transformers in Precision Agriculture: A Comprehensive Survey
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented Reality
- Yet Another Text Captcha Solver
- Advance Fake Video Detection via Vision Transformers
- Bridging the Generalisation Gap: Synthetic Data Generation for Multi-Site Clinical Model Validation
- TrueFake: A Real World Case Dataset of Last Generation Fake Images also Shared on Social Networks
- Enhancing rice breeding efficiency through semi-supervised detection and segmentation of panicles and leaves
- A comprehensive survey of image synthesis approaches for Deep Learning-based surface defect detection in manufacturing
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
- Deep Sound Change: Deep and Iterative Learning, Convolutional Neural Networks, and Language Change
- Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
- Error Analysis of Deep Ritz Methods for Elliptic Equations
- EDMP: Ensemble-of-costs-guided Diffusion for Motion Planning
- Towards responsible AI for education: Hybrid human-AI to confront the Elephant in the room
- MD-GAN: Multi-Discriminator Generative Adversarial Networks for Distributed Datasets
- Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- Deep Neural Networks and Tabular Data: A Survey
- Prototype-Guided Diffusion for Digital Pathology: Achieving Foundation Model Performance with Minimal Clinical Data
- Efficient Task-specific Conditional Diffusion Policies: Shortcut Model Acceleration and SO(3) Optimization
- Model Discrepancy Learning: Synthetic Faces Detection Based on Multi-Reconstruction
- Interpretable Automatic Rosacea Detection with Whitened Cosine Similarity
- Adversarial Training for Dynamics Matching in Coarse-Grained Models
- QEMesh: Employing A Quadric Error Metrics-Based Representation for Mesh Generation
- Releasing Differentially Private Event Logs Using Generative Models
- Fast Facial Landmark Detection and Applications: A Survey
Discussions
- Generative Adversarial Networks [hn, 21 points, 3 comments]
- This is Ian Goodfellow’s original paper! Let me see if I can find some others!
arxiv.org/abs/1406.2661 [bsky, 2 points, 1 comments]
- Generative Adversarial Nets [pdf] [hn, 1 points, 0 comments]
- arxiv.org/pdf/1406.2661
Digging into GANs today - the idea feels kinda magical to me, want to try to code up my own and see where I go with it!
#NowReading [bsky, 1 points, 0 comments]
- Darüber hinaus bekommt man etwas mehr Kontext zu Schmidhubers Plagiatsvorwürfen gegen das Paper von Goodfellow et al. "Generative Adversarial Networks" (arxiv.org/abs/1406.2661). Zentraler Streitpunkt [bsky, 0 points, 1 comments]
- 수학 이론 위주로 설명하길 정말 잘했다....!! 참고한 논문은 링크로 덧붙일게요! (이미 원본 보셨을 수도 있지만) 2014년 논문인데 딥러닝은 벌써 이걸 교과과정에 반영하고 있군요 몇십 년 전에 논의가 끝난 거 배우는 수학과 입장에서는: 너무무섭습니다 https://arxiv.org/abs/1406.2661 [bsky, 0 points, 1 comments]
- Posit: In the future, generative A.I. will be thought of as the unconscious part of a general A.I.'s mind. [lemmy, -23 points, 24 comments]
Related