Image-to-Image Translation with Conditional Adversarial Networks
2016/11/21 by Phillip Isola, Isola, Phillip, Jun-Yan Zhu +5 · 3 voices · 492 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CV
paper · pdf · doi:10.48550/arxiv.1611.07004
Website: https://phillipi.github.io/pix2pix/, CVPR 2017
arxiv published 2016/11/21 · arxiv created 2018/11/26 · arxiv updated 2018/11/27
Abstract
We investigate conditional adversarial networks as a general-purpose solution to image-to-image translation problems. These networks not only learn the mapping from input image to output image, but also learn a loss function to train this mapping. This makes it possible to apply the same generic approach to problems that traditionally would require very different loss formulations. We demonstrate that this approach is effective at synthesizing photos from label maps, reconstructing objects from edge maps, and colorizing images, among other tasks. Indeed, since the release of the pix2pix software associated with this paper, a large number of internet users (many of them artists) have posted their own experiments with our system, further demonstrating its wide applicability and ease of adoption without the need for parameter tweaking. As a community, we no longer hand-engineer our mapping functions, and this work suggests we can achieve reasonable results without hand-engineering our loss functions either.
Cited by
- Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators
- PathoSyn: Imaging-Pathology MRI Synthesis via Disentangled Deviation Diffusion
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction
- ISAC and Vision Fusion for Fine-Grained Low-Altitude Target Recognition
- Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Towards Reliable Stain Transfer: An Iterative Data-Model Co-Optimization Framework Based on Multimodal Expert-Guided Assessment
- I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
- Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
- Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
- Efficient Vision Mamba for MRI Super-Resolution via Hybrid Selective Scanning
- GANeXt: A Fully ConvNeXt-Enhanced Generative Adversarial Network for MRI- and CBCT-to-CT Synthesis
- Inverse-Designed Phase Prediction in Digital Lasers Using Deep Learning and Transfer Learning
- Plasticine: A Traceable Diffusion Model for Medical Image Translation
- InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
- EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
- Pixel Super-Resolved Fluorescence Lifetime Imaging Using Deep Learning
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- MFE-GAN: Efficient GAN-based Framework for Document Image Enhancement and Binarization with Multi-scale Feature Extraction
- Towards Physically-Based Sky-Modeling For Image Based Lighting
- Improving the Plausibility of Pressure Distributions Synthesized from Depth Image through Generative Modeling
- Learning Common and Salient Generative Factors Between Two Image Datasets
- Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real
- MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- A Conditional Generative Framework for Synthetic Data Augmentation in Segmenting Thin and Elongated Structures in Biological Images
- WTNN: Weibull-Tailored Neural Networks for survival analysis
- SCU-CGAN: Enhancing Fire Detection through Synthetic Fire Image Generation and Dataset Augmentation
- Improvements to the NSO Farside Mapping Pipeline: Noise Reduction Updates
- Differentially Private Synthetic Data Generation Using Context-Aware GANs
- Voxify3D: Pixel Art Meets Volumetric Rendering
- See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement
- Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control
- Precise Liver Tumor Segmentation in CT Using a Hybrid Deep Learning-Radiomics Framework
- CADE: Continual Weakly-supervised Video Anomaly Detection with Ensembles
- S2WMamba: A Wavelet-Assisted Mamba-Based Dual-Branch Network For Pansharpening
- Synset Signset Germany: a Synthetic Dataset for German Traffic Sign Recognition
- A Comparative Study on Synthetic Facial Data Generation Techniques for Face Recognition
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- Semantic-Guided Two-Stage GAN for Face Inpainting with Hybrid Perceptual Encoding
- Global-Local Aware Scene Text Editing
- Textured Word-As-Image illustration
- FlowEO: Generative Unsupervised Domain Adaptation for Earth Observation
- Spatiotemporal Satellite Image Downscaling with Transfer Encoders and Autoregressive Generative Models
- PIANO: Physics-informed Dual Neural Operator for Precipitation Nowcasting
- Diffusion-Based Synthesis of 3D T1w MPRAGE Images from Multi-Echo GRE with Multi-Parametric MRI Integration
- Digital Elevation Model Estimation from RGB Satellite Imagery using Generative Deep Learning
- Structure-Preserving Unpaired Image Translation to Photometrically Calibrate JunoCam with Hubble Data
- The Age-specific Alzheimer 's Disease Prediction with Characteristic Constraints in Nonuniform Time Span
- Deformation-aware Temporal Generation for Early Prediction of Alzheimers Disease
- Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
- Probabilistic Wildfire Spread Prediction Using an Autoregressive Conditional Generative Adversarial Network
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- TReFT: Taming Rectified Flow Models For One-Step Image Translation
- Experimental insights into data augmentation techniques for deep learning-based multimode fiber imaging: limitations and success
- Leveraging Adversarial Learning for Pathological Fidelity in Virtual Staining
- Facade Segmentation for Solar Photovoltaic Suitability
- From Healthy Scans to Annotated Tumors: A Tumor Fabrication Framework for 3D Brain MRI Synthesis
- Generation of Granular Deposition Interfaces using conditional Generative Adversarial Network (cGAN)
- Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
- Real-Time Cooked Food Image Synthesis and Visual Cooking Progress Monitoring on Edge Devices
- Physics-Informed Machine Learning for Efficient Sim-to-Real Data Augmentation in Micro-Object Pose Estimation
- Denoising weak lensing mass maps with diffusion model and generative adversarial network
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- Adversarial Learning-Based Radio Map Reconstruction for Fingerprinting Localization
- ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
- Example-Based Feature Painting on Textures
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Domain Adaptation for Camera-Specific Image Characteristics using Shallow Discriminators
- GPDM: Generation-Prior Diffusion Model for Accelerated Direct Attenuation and Scatter Correction of Whole-body 18F-FDG PET
- FlowCast: Advancing Precipitation Nowcasting with Conditional Flow Matching
- Augment to Augment: Diverse Augmentations Enable Competitive Ultra-Low-Field MRI Enhancement
- DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
- HAMscope: a snapshot Hyperspectral Autofluorescence Miniscope for real-time molecular imaging
- KPLM-STA: Physically-Accurate Shadow Synthesis for Human Relighting via Keypoint-Based Light Modeling
- Morphing Through Time: Diffusion-Based Bridging of Temporal Gaps for Robust Alignment in Change Detection
- PADM: A Physics-aware Diffusion Model for Attenuation Correction
- AvatarTex: High-Fidelity Facial Texture Reconstruction from Single-Image Stylized Avatars
- K-Stain: Keypoint-Driven Correspondence for H&E-to-IHC Virtual Staining
- Recovering Sub-threshold S-wave Arrivals in Deep Learning Phase Pickers via Shape-Aware Loss
- Identity Card Presentation Attack Detection: A Systematic Review
- TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
- Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
- Text to Sketch Generation with Multi-Styles
- Adversarial and Score-Based CT Denoising: CycleGAN vs Noise2Score
- Generative deep learning for foundational video translation in ultrasound
- Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Progressive Translation of H&E to IHC with Enhanced Structural Fidelity
- Deep Generative Models for Enhanced Vitreous OCT Imaging
- OSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap Data
- A Hierarchical Deep Learning Model for Predicting Pedestrian-Level Urban Winds
- Emu3.5: Native Multimodal Models are World Learners
- Enabling Fast and Accurate Neutral Atom Readout through Image Denoising
- Eddeep: a deep-learning framework for fast eddy-current distortion correction in diffusion MRI
- Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Low-Dose CT Imaging Using a Regularization-Enhanced Efficient Diffusion Probabilistic Model
- Residual Diffusion Bridge Model for Image Restoration
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- TraceTrans: Translation and Spatial Tracing for Surgical Prediction
- Digital Contrast CT Pulmonary Angiography Synthesis from Non-contrast CT for Pulmonary Vascular Disease
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Unsupervised Domain Adaptation via Similarity-based Prototypes for Cross-Modality Segmentation
- NeuralTouch: Neural Descriptors for Precise Sim-to-Real Tactile Robot Control
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Lightweight CycleGAN Models for Cross-Modality Image Transformation and Experimental Quality Assessment in Fluorescence Microscopy
- Diffusion Bridge Networks Simulate Clinical-grade PET from MRI for Dementia Diagnostics
- Curvilinear Structure-preserving Unpaired Cross-domain Medical Image Translation
- Learning and Simulating Building Evacuation Patterns for Enhanced Safety Design Using Generative Models
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- Detecting streaks in smart telescopes images with Deep Learning
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- Cyclic Self-Supervised Diffusion for Ultra Low-field to High-field MRI Synthesis
- Only-Style: Stylistic Consistency in Image Generation without Content Leakage
- VCTR: A Transformer-Based Model for Non-parallel Voice Conversion
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- DISC-GAN: Disentangling Style and Content for Cluster-Specific Synthetic Underwater Image Generation
- SuperEx: Enhancing Indoor Mapping and Exploration using Non-Line-of-Sight Perception
- Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
- Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
- Cross-Sensor Touch Generation
- Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion
- FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
- Biology-driven assessment of deep learning super-resolution imaging of the porosity network in dentin
- WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
- A Probabilistic Basis for Low-Rank Matrix Learning
- Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional Diffusion
- Bridging Text and Video Generation: A Survey
- Real-time Prediction of Urban Sound Propagation with Conditioned Normalizing Flows
- Score-based generative emulation of impact-relevant Earth system model outputs
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- The 1st Solution for CARE Liver Task Challenge 2025: Contrast-Aware Semi-Supervised Segmentation with Domain Generalization and Test-Time Adaptation
- SDAKD: Student Discriminator Assisted Knowledge Distillation for Super-Resolution Generative Adversarial Networks
- Style Brush: Guided Style Transfer for 3D Objects
- Training Variation of Physically-Informed Deep Learning Models
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- OTR: Synthesizing Overlay Text Dataset for Text Removal
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- UrbanGraph: Physics-Informed Spatio-Temporal Dynamic Heterogeneous Graphs for Urban Microclimate Prediction
- MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
- Vector sketch animation generation with differentialable motion trajectories
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- One-shot Conditional Sampling: MMD meets Nearest Neighbors
- Score-based Membership Inference on Diffusion Models
- VAGUEGAN: Stealthy Poisoning and Backdoor Attacks on Image Generative Pipelines
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- Environment-Aware Satellite Image Generation with Diffusion Models
- From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Anatomy-DT: A Cross-Diffusion Digital Twin for Anatomical Evolution
- An Efficient 3D Latent Diffusion Model for T1-contrast Enhanced MRI Generation
- Simulating Post-Neoadjuvant Chemotherapy Breast Cancer MRI via Diffusion Model with Prompt Tuning
- Tumor Synthesis conditioned on Radiomics
- LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- Controllable Generation of Large-Scale 3D Urban Layouts with Semantic and Structural Guidance
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- Comparative Analysis of GAN and Diffusion for MRI-to-CT translation
- Universal Multi-Domain Translation via Diffusion Routers
- StrCGAN: A Generative Framework for Stellar Image Restoration
- RFMSR: Residual Flow Matching for Image Super-Resolution
- ReDiff: Reliability-Guided Diffusion for Trustworthy Ultra-Low-Field to High-Field MRI Synthesis
- Topology Aware Neural Interpolation of Scalar Fields
- Dual-conditioned diffusion model with anatomical guidance for geometric distortion correction in prostate MRI
- One-shot Embroidery Customization via Contrastive LoRA Modulation
- A Kernel Space-based Multidimensional Sparse Model for Dynamic PET Image Denoising
- Self-Alignment Learning to Improve Myocardial Infarction Detection from Single-Lead ECG
- Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid Integration
- Recent Advancements in Microscopy Image Enhancement using Deep Learning: A Survey
- RaceGAN: A Framework for Preserving Individuality while Converting Racial Information for Image-to-Image Translation
- A Race Bias Free Face Aging Model for Reliable Kinship Verification
- Generative AI for Misalignment-Resistant Virtual Staining to Accelerate Histopathology Workflows
- Adversarial Appearance Learning in Augmented Cityscapes for Pedestrian Recognition in Autonomous Driving
- Deep learning approach for flow visualization in background-oriented schlieren
- Radio Propagation Modelling: To Differentiate or To Deep Learn, That Is The Question
- From Autoencoders to CycleGAN: Robust Unpaired Face Manipulation via Adversarial Learning
- BioMetaphor: AI-Generated Biodata Representations for Virtual Co-Present Events
- REGEN: Real-Time Photorealism Enhancement in Games via a Dual-Stage Generative Network Framework
- BREA-Depth: Bronchoscopy Realistic Airway-geometric Depth Estimation
- Geometric Analysis of Magnetic Labyrinthine Stripe Evolution via Deep Learning Segmentation
- Styleclone: Face Stylization with Diffusion Based Data Augmentation
- Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation
- Deep learning-driven adaptive optics for laser wavefront correction
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Mapping of discrete range modulated proton radiograph to water-equivalent path length using machine learning
- Generating Synthetic Contrast-Enhanced Chest CT Images from Non-Contrast Scans Using Slice-Consistent Brownian Bridge Diffusion Network
- VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
- Feature Space Analysis by Guided Diffusion Model
- XOCT: Enhancing OCT to OCTA Translation via Cross-Dimensional Supervised Multi-Scale Feature Learning
- Neural Cone Radiosity for Interactive Global Illumination with Glossy Materials
- Reconstruction and Reenactment Separated Method for Realistic Gaussian Head
- Label-Free Whole Slide Virtual Multi-Staining Using Dual-Excitation Photon Absorption Remote Sensing Microscopy
- Latent space projections and atlases: A cautionary tale in deep neuroimaging using autoencoders
- Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation
- Towards Diagnostic Quality Flat-Panel Detector CT Imaging Using Diffusion Models
- Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
- Improving atomic force microscopy structure discovery via style-translation
- Deep learning-enabled virtual multiplexed immunostaining of label-free tissue for vascular invasion assessment
- Generative AI for Enhanced Wildfire Detection: Bridging the Synthetic-Real Domain Gap
- PRINTER:Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ Fidelity
- A Unified Low-level Foundation Model for Enhancing Pathology Image Quality
- Double-Constraint Diffusion Model with Nuclear Regularization for Ultra-low-dose PET Reconstruction
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Non-expert to Expert Motion Translation Using Generative Adversarial Networks
- Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
- SimShear: Sim-to-Real Shear-based Tactile Servoing
- Self-supervised physics-informed generative networks for phase retrieval from a single X-ray hologram
- MRExtrap: Longitudinal Aging of Brain MRIs using Linear Modeling in Latent Space
- Generative AI in Map-Making: A Technical Exploration and Its Implications for Cartographers
- 2D Ultrasound Elasticity Imaging of Abdominal Aortic Aneurysms Using Deep Neural Networks
- Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
- FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction
- Virtual Multiplex Staining for Histological Images using a Marker-wise Conditioned Diffusion Model
- Taming Transformer for Emotion-Controllable Talking Face Generation
- Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing
- Eliminating Rasterization: Direct Vector Floor Plan Generation with DiffPlanner
- 2D Gaussians Meet Visual Tokenizer
- ID-Card Synthetic Generation: Toward a Simulated Bona fide Dataset
- Self-supervised learning for multiplexing super-resolution confocal microscopy
- Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
- SNNSIR: A Simple Spiking Neural Network for Stereo Image Restoration
- AnatoMaskGAN: GNN-Driven Slice Feature Fusion and Noise Augmentation for Medical Semantic Image Synthesis
- Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension
- ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
- Geospatial Diffusion for Land Cover Imperviousness Change Forecasting
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- Efficient and Scalable Self-Healing Databases Using Meta-Learning and Dependency-Driven Recovery
- Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
- Leveraging GNN to Enhance MEF Method in Predicting ENSO
- GANime: Generating Anime and Manga Character Drawings from Sketches with Deep Learning
- CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
- ContourDiff: Unpaired Medical Image Translation with Structural Consistency
- FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
- MM2CT: MR-to-CT translation for multi-modal image fusion with mamba
- Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
- Automatic Image Colorization with Convolutional Neural Networks and Generative Adversarial Networks
- Deeper Inside Deep ViT
- Cross-Domain Image Synthesis: Generating H&E from Multiplex Biomarker Imaging
- evTransFER: A Transfer Learning Framework for Event-based Facial Expression Recognition
- Learning Latent Representations for Image Translation using Frequency Distributed CycleGAN
- SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis
- CoCoLIT: ControlNet-Conditioned Latent Image Translation for MRI to Amyloid PET Synthesis
- StyDeco: Unsupervised Style Transfer with Distilling Priors and Semantic Decoupling
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
- Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars
- Sample-Aware Test-Time Adaptation for Medical Image-to-Image Translation
- Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
- Data-driven global ocean model resolving ocean-atmosphere coupling dynamics
- LCS: An AI-based Low-Complexity Scaler for Power-Efficient Super-Resolution of Game Content
- DWTGS: Rethinking Frequency Regularization for Sparse-view 3D Gaussian Splatting
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
- Robust Adverse Weather Removal via Spectral-based Spatial Grouping
- Visual Language Models as Zero-Shot Deepfake Detectors
- Evaluating Deepfake Detectors in the Wild
- MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors
- Locally Controlled Face Aging with Latent Diffusion Models
- Conditional Diffusion Models for Global Precipitation Map Inpainting
- JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync
- Physics-Informed Self-Supervised Generative Model for 3D Localization Microscopy
- MoFRR: Mixture of Diffusion Models for Face Retouching Restoration
- Joint Holistic and Lesion Controllable Mammogram Synthesis via Gated Conditional Diffusion Model
- Reconstruct or Generate: Exploring the Spectrum of Generative Modeling for Cardiac MRI
- Facial Demorphing from a Single Morph Using a Latent Conditional GAN
- Benchmarking GANs, Diffusion Models, and Flow Matching for T1w-to-T2w MRI Translation
- MoDyGAN: Combining Molecular Dynamics With GANs to Investigate Protein Conformational Space
- Towards Facilitated Fairness Assessment of AI-based Skin Lesion Classifiers Through GenAI-based Image Synthesis
- Global Modeling Matters: A Fast, Lightweight and Effective Baseline for Efficient Image Restoration
- Pyramid Hierarchical Masked Diffusion Model for Imaging Synthesis
- Novel techniques of imaging interferometry analysis to study gas and plasma density for laser-plasma experiments
- HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
- Pathology-Guided Virtual Staining Metric for Evaluation and Training
- Galaxy image simplification using Generative AI
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Seedance 1.0: Exploring the Boundaries of Video Generation Models
- A Privacy-Preserving Federated Learning Framework for Generalizable CBCT to Synthetic CT Translation in Head and Neck
- Quantize-then-Rectify: Efficient VQ-VAE Training
- Transferring Styles for Reduced Texture Bias and Improved Robustness in Semantic Segmentation Networks
- WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
- Interactive Drawing Guidance for Anime Illustrations with Diffusion Model
- Domain Adaptation-Enabled Realistic Map-Based Channel Estimation for MIMO-OFDM
- Image Translation with Kernel Prediction Networks for Semantic Segmentation
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- A novel attention mechanism for noise-adaptive and robust segmentation of microtubules in microscopy images
- Accelerating Transposed Convolutions on FPGA-based Edge Devices
- Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
- USIGAN: Unbalanced Self-Information Feature Transport for Weakly Paired Image IHC Virtual Staining
- Foreground-aware Virtual Staining for Accurate 3D Cell Morphological Profiling
- Enhancing Underwater Images Using Deep Learning with Subjective Image Quality Integration
- Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors
- MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
- RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- LLM-based Realistic Safety-Critical Driving Video Generation
- UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather Conditions
- SonoGym: High Performance Simulation for Challenging Surgical Tasks with Robotic Ultrasound
- Towards 3D Semantic Image Synthesis for Medical Imaging
- Calligrapher: Freestyle Text Image Customization
- Metadata, Wavelet, and Time Aware Diffusion Models for Satellite Image Super Resolution
- CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
- Advancing Facial Stylization through Semantic Preservation Constraint and Pseudo-Paired Supervision
- FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- CitrusGAN: sparse-view X-ray CT reconstruction for citrus based on generative adversarial networks
- Generative Blocks World: Moving Things Around in Pictures
- Shape2Animal: Creative Animal Generation from Natural Silhouettes
- Photon Absorption Remote Sensing (PARS): Comprehensive Absorption Imaging Enabling Label-Free Biomolecule Characterization and Mapping
- Angio-Diff: Learning a Self-Supervised Adversarial Diffusion Model for Angiographic Geometry Generation
- ReMAR-DS: Recalibrated Feature Learning for Metal Artifact Reduction and CT Domain Transformation
- Style Transfer: A Decade Survey
- GANs vs. Diffusion Models for virtual staining with the HER2match dataset
- Transforming H&E images into IHC: A Variance-Penalized GAN for Precision Oncology
- Staining normalization in histopathology: Method benchmarking using multicenter dataset
- Pix2Geomodel: A Next-Generation Reservoir Geomodeling with Property-to-Property Translation
- MTSIC: Multi-stage Transformer-based GAN for Spectral Infrared Image Colorization
- Reversing Flow for Image Restoration
- Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- A Prior-Guided Joint Diffusion Model in Projection Domain for PET Tracer Conversion
- MoiréXNet: Adaptive Multi-Scale Demoiréing with Linear Attention Test-Time Training and Truncated Flow Matching Prior
- NTIRE 2025 Image Shadow Removal Challenge Report
- Domain Adaptation for Image Classification of Defects in Semiconductor Manufacturing
- One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning
- Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
- You Only Render Once: Enhancing Energy and Computation Efficiency of Mobile Virtual Reality
- D2Diff : A Dual Domain Diffusion Model for Accurate Multi-Contrast MRI Synthesis
- Training-Free Diffusion Framework for Stylized Image Generation with Identity Preservation
- orGAN: A Synthetic Data Augmentation Pipeline for Simultaneous Generation of Surgical Images and Ground Truth Labels
- Deep Diffusion Models and Unsupervised Hyperspectral Unmixing for Realistic Abundance Map Synthesis
- Discovering Hierarchy-Grounded Domains with Adaptive Granularity for Clinical Domain Generalization
- Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments
- Stacked Intelligent Metasurfaces for Multi-Modal Semantic Communications
- Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- SPC to 3D: Novel View Synthesis from Binary SPC via I2I translation
- Tensor-to-Tensor Models with Fast Iterated Sum Features
- Can Foundation Models Generalise the Presentation Attack Detection Capabilities on ID Cards?
- Deep histological synthesis from mass spectrometry imaging for multimodal registration
- InstancePin: Instance-Addressable Layout-to-Image Diffusion via Coordinate Pinning
- MAISY: Motion-Aware Image SYnthesis for Medical Image Motion Correction
- Integrated Image Reconstruction and Target Recognition based on Deep Learning Technique
- Geometry-Aware Texture Generation for 3D Head Modeling with Artist-driven Control
- ConText: Driving In-context Learning for Text Removal and Segmentation
- You Only Train Once
- Facial Appearance Capture at Home with Patch-Level Reflectance Prior
- Improving Post-Processing for Quantitative Precipitation Forecasting Using Deep Learning: Learning Precipitation Physics from High-Resolution Observations
- Res-MoCoDiff: Residual-guided diffusion models for motion artifact correction in brain MRI
- Tactile MNIST: Benchmarking Active Tactile Perception
- One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
- Path and Bone-Contour Regularized Unpaired MRI-to-CT Translation
- A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
- Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- Multi-Platform Methane Plume Detection via Model and Domain Adaptation
- Transport Network, Graph, and Air Pollution
- SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
- Common Inpainted Objects In-N-Out of Context
- On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning
- Segmenting France Across Four Centuries
- Autoregressive regularized score-based diffusion models for multi-scenarios fluid flow prediction
- Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes
- Energy-Embedded Neural Solvers for One-Dimensional Quantum Systems
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- GL-PGENet: A Parameterized Generation Framework for Robust Document Image Enhancement
- Neural Restoration of Greening Defects in Historical Autochrome Photographs Based on Purely Synthetic Data
- Multipath cycleGAN for harmonization of paired and unpaired low-dose lung computed tomography reconstruction kernels
- Unpaired Image-to-Image Translation for Segmentation and Signal Unmixing
- 'Hello, World!': Making GNNs Talk with LLMs
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
- Laparoscopic Image Desmoking Using the U-Net with New Loss Function and Integrated Differentiable Wiener Filter
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- Lesion-Aware Generative Artificial Intelligence for Virtual Contrast-Enhanced Mammography in Breast Cancer
- DeepInverse: A Python package for solving imaging inverse problems with deep learning
- ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
- MAMM: Motion Control via Metric-Aligning Motion Matching
- GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation
- Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos
- ReflectGAN: Modeling Vegetation Effects for Soil Carbon Estimation from Satellite Imagery
- Eye-See-You: Reverse Pass-Through VR and Head Avatars
- F-ANcGAN: An Attention-Enhanced Cycle Consistent Generative Adversarial Architecture for Synthetic Image Generation of Nanoparticles
- RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
- R-Genie: Reasoning-Guided Generative Image Editing
- MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image Recovery
- Learning Shared Representations from Unpaired Data
- Seeing through Satellite Images at Street Views
- Forward-only Diffusion Probabilistic Models
- Paired and Unpaired Image to Image Translation using Generative Adversarial Networks
- Materials Generation in the Era of Artificial Intelligence: A Comprehensive Survey
- ChemMLLM: Chemical Multimodal Large Language Model
- Leveraging the Powerful Attention of a Pre-trained Diffusion Model for Exemplar-based Image Colorization
- Data Augmentation and Resolution Enhancement using GANs and Diffusion Models for Tree Segmentation
- Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning
- A Deep Learning Framework for Two-Dimensional, Multi-Frequency Propagation Factor Estimation
- gen2seg: Generative Models Enable Generalizable Instance Segmentation
- 3D Reconstruction from Sketches
- Handloom Design Generation Using Generative Networks
- Towards Generating Realistic Underwater Images
- RETRO: REthinking Tactile Representation Learning with Material PriOrs
- GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
- Deep Generative Modeling with Spatial and Network Images: An Explainable AI (XAI) Approach
- Generative Adversarial Reconstruction with Adaptive Thresholding for Obstructed Targets in Computational Microwave Imaging
- CHRIS: Clothed Human Reconstruction with Side View Consistency
- From Fibers to Cells: Fourier-Based Registration Enables Virtual Cresyl Violet Staining From 3D Polarized Light Imaging
- Content Generation Models in Computational Pathology: A Comprehensive Survey on Methods, Applications, and Challenges
- X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models
- ROIsGAN: A Region Guided Generative Adversarial Framework for Murine Hippocampal Subregion Segmentation
- How to use score-based diffusion in earth system science: A satellite nowcasting example
- MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models
- Q-space Guided Collaborative Attention Translation Network for Flexible Diffusion-Weighted Images Synthesis
- IMPLICITSTAINER: Resolution Agnostic Data-Efficient Virtual Staining Using Neural Implicit Functions
- FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization
- Beyond Edge Maps: Wavelet-Domain Conditioning for Multi-Adapter Map-to-Satellite Diffusion
- EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
- Few-shot Semantic Encoding and Decoding for Video Surveillance
- Autonomous Robotic Pruning in Orchards and Vineyards: a Review
- HistDiST: Histopathological Diffusion-based Stain Transfer
- Decoupling Multi-Contrast Super-Resolution: Self-Supervised Implicit Re-Representation for Unpaired Cross-Modal Synthesis
- Regression is all you need for medical image translation
- Enhancing AI Face Realism: Cost-Efficient Quality Improvement in Distilled Diffusion Models with a Fully Synthetic Dataset
- MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
- Adversarial Robustness of Deep Learning Models for Inland Water Body Segmentation from SAR Images
- OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction
- Unpaired Cross-Domain Calibration of DMSP to VIIRS Nighttime Light Data Based on CUT Network
- Denoising weak lensing mass maps with diffusion model: systematic comparison with generative adversarial network
- Forget the Learning Rate, Decay Loss
- Tailor Made Embeddings for Quantum Machine Learning
- Scaling Self-Play for End-to-End Driving
- Bubble2Heat: Optical to Thermal Inference in Pool Boiling Using Physics-encoded Generative AI
- Visual Text Processing: A Comprehensive Review and Unified Evaluation
- Text-Conditioned Diffusion Model for High-Fidelity Korean Font Generation
- AI-in-the-Loop Planning for Transportation Electrification: Case Studies from Austin, Texas
- 2D Image Relighting with Image-to-Image Translation
- Axial-UNet: A Neural Weather Model for Precipitation Nowcasting
- What Matters in Practical Learned Image Compression
- EarthMapper: Visual Autoregressive Models for Controllable Bidirectional Satellite-Map Translation
- GAN-SLAM: Real-Time GAN Aided Floor Plan Creation Through SLAM
- A Langevin sampling algorithm inspired by the Adam optimizer
- Exploiting Multiple Representations: 3D Face Biometrics Fusion with Application to Surveillance
- Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking Robustness
- Spectral Collapse in Diffusion Inversion
- A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
- ESDiff: Encoding Strategy-inspired Diffusion Model with Few-shot Learning for Color Image Inpainting
- RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
- DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
- GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
- Virtual Scanning for NSCLC Histology: Investigating the Discriminatory Power of Synthetic PET
- Fried Parameter Estimation from Single Wavefront Sensor Image with Artificial Neural Networks
- Hyper-Transforming Latent Diffusion Models
- GAMBAS: Generalised-Hilbert Mamba for Super-resolution of Paediatric Ultra-Low-Field MRI
- Satellite to GroundScape -- Large-scale Consistent Ground View Generation from Satellite Views
- IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- NTIRE 2025 Challenge on Image Super-Resolution (×4): Methods and Results
- Circular Image Deturbulence using Quasi-conformal Geometry
- PDNNet: PDN-Aware GNN-CNN Heterogeneous Network for Dynamic IR Drop Prediction
- IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion Design
- Image Editing with Diffusion Models: A Survey
- Beyond geometric deformation: High-fidelity orthodontic profile synthesis via ControlNet-guided generative AI
- AdaQual-Diff: Diffusion-Based Image Restoration via Adaptive Quality Prompting
- Mask Image Watermarking
- ADA-Net: Attention-Guided Domain Adaptation Network with Contrastive Learning for Standing Dead Tree Segmentation Using Aerial Imagery
- Enhancing Deterministic Freezing Level Predictions in the Northern Sierra Nevada Through Deep Neural Networks
- Big Brother is Watching: Proactive Deepfake Detection via Learnable Hidden Face
- Digital Staining with Knowledge Distillation: A Unified Framework for Unpaired and Paired-But-Misaligned Data
- DiffuMural: Restoring Dunhuang Murals with Multi-scale Diffusion
- Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance
- seg2med: a bridge from artificial anatomy to multimodal medical images
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- X-DECODE: EXtreme Deblurring with Curriculum Optimization and Domain Equalization
- RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
- ColorizeDiffusion v2: Enhancing Reference-based Sketch Colorization Through Separating Utilities
- Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
- Generative Adversarial Networks with Limited Data: A Survey and Benchmarking
- DA2Diff: Exploring Degradation-aware Adaptive Diffusion Priors for All-in-One Weather Restoration
- AnyArtisticGlyph: Multilingual Controllable Artistic Glyph Generation
- Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal
Discussions
Related