Generative Adversarial Networks
2014/06/10 by Ian Goodfellow, Ian J. Goodfellow, Goodfellow, Ian J. +14 · 7 voices · 91 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Image Processing and 3D Reconstruction #Music and Audio Processing #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1406.2661
openalex publication_date 2014/06/10 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/31
Abstract
Large Language Models (LLMS) rely on Key-Value (KV) caches to store attention context during autoregressive decoding. In long-sequence settings, the KV cache can consume large amounts of VRAM and become a practical bottleneck for throughput . We introduce KVHALO, an auxiliary reconstruction model that restores higher-fidelity KV tensors from a compressed cache state when required, reducing persistent memory footprint during inference. In our evaluation, KVHALO achieves up to 91.85% directional cosine alignment at convergence and reduces long-context degradation relative to a low-bit baseline under our stress-test workloads. We used HRM instead of other architectures, which allowed for higher-quality results in only 18,600 steps.
Cited by
- InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation
- Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
- FUSE: Unifying Spectral and Semantic Cues for Robust AI-Generated Image Detection
- Detection of AI Generated Images Using Combined Uncertainty Measures and Particle Swarm Optimised Rejection Mechanism
- Grad: Guided Relation Diffusion Generation for Graph Augmentation in Graph Fraud Detection
- Vision-Language Model Guided Image Restoration
- AI-Augmented Pollen Recognition in Optical and Holographic Microscopy for Veterinary Imaging
- A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
- A Conditional Generative Framework for Synthetic Data Augmentation in Segmenting Thin and Elongated Structures in Biological Images
- Provable Long-Range Benefits of Next-Token Prediction
- General and Domain-Specific Zero-shot Detection of Generated Images via Conditional Likelihood
- Advanced Unsupervised Learning: A Comprehensive Overview of Multi-View Clustering Techniques
- Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models
- One-Step Generative Channel Estimation via Average Velocity Field
- Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade
- SD-CGAN: Conditional Sinkhorn Divergence GAN for DDoS Anomaly Detection in IoT Networks
- Three-Dimensional Anatomical Data Generation Based on Artificial Neural Networks
- From Healthy Scans to Annotated Tumors: A Tumor Fabrication Framework for 3D Brain MRI Synthesis
- HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
- BrainNormalizer: Anatomy-Informed Pseudo-Healthy Brain Reconstruction from Tumor MRI via Edge-Guided ControlNet
- DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
- Machine learning-driven modeling framework for steam co-gasification applications
- Diffusion Models at the Drug Discovery Frontier: A Review on Generating Small Molecules versus Therapeutic Peptides
- A Survey of Heterogeneous Graph Neural Networks for Cybersecurity Anomaly Detection
- Reading Radio from Camera: Visually-Grounded, Lightweight, and Interpretable RSSI Prediction
- Manipulate as Human: Learning Task-oriented Manipulation Skills by Adversarial Motion Priors
- DPGLA: Bridging the Gap between Synthetic and Real Data for Unsupervised Domain Adaptation in 3D LiDAR Semantic Segmentation
- DDTR: Diffusion Denoising Trace Recovery
- Collaborative penetration testing suite for emerging generative AI algorithms
- Optimizing DINOv2 with Registers for Face Anti-Spoofing
- Conditional Synthetic Live and Spoof Fingerprint Generation
- A Multi-domain Image Translative Diffusion StyleGAN for Iris Presentation Attack Detection
- Global-focal Adaptation with Information Separation for Noise-robust Transfer Fault Diagnosis
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Using neural style transfer to study the evolution of animal signal design: A case study in an ornamented fish
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- Denoised Diffusion for Object-Focused Image Augmentation
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- cryoTIGER: deep-learning based tilt interpolation generator for enhanced reconstruction in cryo electron tomography
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
- ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
- Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- DiffCamera: Arbitrary Refocusing on Images
- Non-Parametric Simulation of Multivariate Extreme Events via Spectral Bootstrap
- MapGenerator: a framework for learning a diffusion model for text promptable map generation
- Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
- A Biophysical-Model-Informed Source Separation Framework For EMG Decomposition
- Pure Node Selection for Imbalanced Graph Node Classification
- Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
- Modelling and design of transcriptional enhancers
- Facing & mitigating common challenges when working with real-world data: The Data Learning Paradigm
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- Red Teaming Quantum-Resistant Cryptographic Standards: A Penetration Testing Framework Integrating AI and Quantum Security
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
- Latent Conditioned Loco-Manipulation Using Motion Priors
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Generative AI for Misalignment-Resistant Virtual Staining to Accelerate Histopathology Workflows
- Inverse Design of Amorphous Materials with Targeted Properties
- Autonomous Intelligent Monitoring of Photovoltaic Systems: An In‐Depth Multidisciplinary Review
- Self-Evolving LLMs via Continual Instruction Tuning
- Synthetic Dataset Evaluation Based on Generalized Cross Validation
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Privacy-Preserving Automated Rosacea Detection Based on Medically Inspired Region of Interest Selection
- RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- BIR-Adapter: A Low-Complexity Diffusion Model Adapter for Blind Image Restoration
- Systematic Review and Meta-analysis of AI-driven MRI Motion Artifact Detection and Correction
- Set Transformer Architectures and Synthetic Data Generation for Flow-Guided Nanoscale Localization
- From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
- Review on Convolutional Neural Network (CNN) Applied to Plant Leaf Disease Classification
- University students' perceptions on how generative artificial intelligence shape learning and research practices: A case study in Hong Kong
- Semi-Supervised Bayesian GANs with Log-Signatures for Uncertainty-Aware Credit Card Fraud Detection
- Tabular Diffusion Counterfactual Explanations
- Generative AI for Testing of Autonomous Driving Systems: A Survey
- Mining gold from implicit models to improve likelihood-free inference
- From Sound to Sight: Towards AI-authored Music Videos
- DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
- Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
- On the Generalization Limits of Quantum Generative Adversarial Networks with Pure State Generators
- Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
- Safeguarding Generative AI Applications in Preclinical Imaging through Hybrid Anomaly Detection
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- Deep learning for Alzheimer's disease diagnosis: A survey
- FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
- Energy-Efficient Federated Learning for Edge Real-Time Vision via Joint Data, Computation, and Communication Design
- Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data
Discussions
- Generative Adversarial Networks [hn, 21 points, 3 comments]
- This is Ian Goodfellow’s original paper! Let me see if I can find some others!
arxiv.org/abs/1406.2661 [bsky, 2 points, 1 comments]
- Generative Adversarial Nets [pdf] [hn, 1 points, 0 comments]
- arxiv.org/pdf/1406.2661
Digging into GANs today - the idea feels kinda magical to me, want to try to code up my own and see where I go with it!
#NowReading [bsky, 1 points, 0 comments]
- Darüber hinaus bekommt man etwas mehr Kontext zu Schmidhubers Plagiatsvorwürfen gegen das Paper von Goodfellow et al. "Generative Adversarial Networks" (arxiv.org/abs/1406.2661). Zentraler Streitpunkt [bsky, 0 points, 1 comments]
- 수학 이론 위주로 설명하길 정말 잘했다....!! 참고한 논문은 링크로 덧붙일게요! (이미 원본 보셨을 수도 있지만) 2014년 논문인데 딥러닝은 벌써 이걸 교과과정에 반영하고 있군요 몇십 년 전에 논의가 끝난 거 배우는 수학과 입장에서는: 너무무섭습니다 https://arxiv.org/abs/1406.2661 [bsky, 0 points, 1 comments]
- Posit: In the future, generative A.I. will be thought of as the unconscious part of a general A.I.'s mind. [lemmy, -23 points, 24 comments]
Related