Image quality assessment: from error visibility to structural similarity
2004/04/01 by Zhou Wang, A.C. Bovik, H.R. Sheikh +1 · 1,576 citations
paper · doi:10.1109/tip.2003.819861
Cited by
- ROGR: Relightable 3D Objects using Generative Relighting
- CASIAL: Geometric Distortion Robust Image Watermarking
- 3DGBGS: 3D Granular Ball Gaussian Splatting for Compact Novel View Synthesis
- Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
- SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
- WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models
- Weight and Height Estimation from a Single Human Image Captured in the Wild
- Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
- Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
- Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors
- A Relaxed Gradient Step Denoiser for Splitting Methods in Poisson Inverse Problems
- InkShield: Writing Style Protection Against Unauthorized Handwriting Mimicry
- Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
- Calibrated Pressure-Observable Born and Hessian Actions for Quantum-Assisted Waveform Inversion
- UniForward: Unified 3D Scene and Semantic Field Reconstruction via Feed-Forward Gaussian Splatting from Only Sparse-View Images
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
- Learning Developmental Scaffoldings to Guide Self-Organisation
- SceneExpander: Text-Guided 3D Scene Expansion via Free-Form View Insertion
- Walk through Paintings: Egocentric World Models from Internet Priors
- HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D Avatars
- Seeing Clearly and Deeply: An RGBD Imaging Approach with a Bio-inspired Monocentric Design
- Recover Biological Structure from Sparse-View Diffraction Images with Neural Volumetric Prior
- Larger Hausdorff Dimension in Scanning Pattern Facilitates Mamba-Based Methods in Low-Light Image Enhancement
- Group Relative Attention Guidance for Image Editing
- A Luminance-Aware Multi-Scale Network for Polarization Image Fusion with a Multi-Scene Dataset
- High-Quality and Large-Scale Image Downscaling for Modern Display Devices
- See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
- Variable Projected Augmented Lagrangian Methods for Generalized Lasso Problems
- Differential Privacy: Gradient Leakage Attacks in Federated Learning Environments
- A Survey on Collaborative SLAM with 3D Gaussian Splatting
- Neural USD: An object-centric framework for iterative editing and control
- Adaptive Keyframe Selection for Scalable 3D Scene Reconstruction in Dynamic Environments
- Fast and scalable joint-LORAKS reconstruction and data-driven sampling optimisation of high-dimensional MRI datasets using a GPU-accelerated and learning-free differentiable framework: PyLORAKS
- Super-recognizers sample visual information of superior computational value for facial recognition
- Low-Dose CT Imaging Using a Regularization-Enhanced Efficient Diffusion Probabilistic Model
- Semantic-Guided Cross-Sensor Super Resolution of Remote Sensing Images: A Gated Dual Conditioning Flow Matching Model
- A global inverse-problem approach to quantitative photo-switching optoacoustic mesoscopy
- MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
- Equivariance2Inverse: A Practical Self-Supervised CT Reconstruction Method Benchmarked on Real, Limited-Angle, and Blurred Data
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- Residual Diffusion Bridge Model for Image Restoration
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
- MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- Learning Event-guided Exposure-agnostic Video Frame Interpolation via Adaptive Feature Blending
- Structure Aware Image Downscaling
- Low-Light Image Enhancement Using Gamma Learning And Attention-Enabled Encoder-Decoder Networks
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- DynaPose4D: High-Quality 4D Dynamic Content Generation via Pose Alignment Loss
- TraceTrans: Translation and Spatial Tracing for Surgical Prediction
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
- Frequency-Spatial Interaction Driven Network for Low-Light Image Enhancement
- Scanner-Agnostic MRI Harmonization via SSIM-Guided Disentanglement
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- FlowOpt: Fast Optimization Through Whole Flow Processes for Training-Free Editing
- Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
- LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare Removal
- Depth-Supervised Fusion Network for Seamless-Free Image Stitching
- Cold-Diffusion Driven Downward Continuation of Gravity Data
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- MEIcoder: Decoding Visual Stimuli from Neural Activity by Leveraging Most Exciting Inputs
- Active control the peak value of Hanbury Brown-Twiss effect with classical light by holographic projection
- Positional Encoding Field
- Lightweight CycleGAN Models for Cross-Modality Image Transformation and Experimental Quality Assessment in Fluorescence Microscopy
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- Target-aware Image Editing via Cycle-consistent Constraints
- Inverse Image-Based Rendering for Light Field Generation from Single Images
- From Far and Near: Perceptual Evaluation of Crowd Representations Across Levels of Detail
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Curvilinear Structure-preserving Unpaired Cross-domain Medical Image Translation
- VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction
- BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
- BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
- BrainPuzzle: Hybrid Physics and Data-Driven Reconstruction for Transcranial Ultrasound Tomography
- GRASPLAT: Enabling dexterous grasping through novel view synthesis
- HDR Image Reconstruction using an Unsupervised Fusion Model
- Mono4DGS-HDR: High Dynamic Range 4D Gaussian Splatting from Alternating-exposure Monocular Videos
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- Decoding Dynamic Visual Experience from Calcium Imaging via Cell-Pattern-Aware Pretraining
- DP2O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Adapting Stereo Vision From Objects To 3D Lunar Surface Reconstruction with the StereoLunar Dataset
- ConsistEdit: Highly Consistent and Precise Training-free Visual Editing
- RaindropGS: A Benchmark for 3D Gaussian Splatting under Raindrop Conditions
- Initialize to Generalize: A Stronger Initialization Pipeline for Sparse-View 3DGS
- Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories
- Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
- iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
- CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration
- Rethinking Nighttime Image Deraining via Learnable Color Space Transformation
- Prominence-Aware Artifact Detection and Dataset for Image Super-Resolution
- 2DGS-R: Revisiting the Normal Consistency Regularization in 2D Gaussian Splatting
- A Comprehensive Survey on World Models for Embodied AI
- A GAN-Based Framework for Generating STFT Spectrograms of Rare Acoustic Events in Structural Health Monitoring
- Personalized Image Filter: Mastering Your Photographic Style
- Latent Diffusion Model without Variational Autoencoder
- Compressive Modeling and Visualization of Multivariate Scientific Data using Implicit Neural Representation
- Cost Savings from Automatic Quality Assessment of Generated Images
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- Deep generative priors for 3D brain analysis
- Galaxy Morphology Classification with Counterfactual Explanation
- SaLon3R: Structure-aware Long-term Generalizable 3D Reconstruction from Unposed Images
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- Inpainting the Red Planet: Diffusion Models for the Reconstruction of Martian Environments in Virtual Reality
- LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
- BalanceGS: Algorithm-System Co-design for Efficient 3D Gaussian Splatting Training on GPU
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- Acquisition of interpretable domain information during brain MR image harmonization for content-based image retrieval
- MACE: Mixture-of-Experts Accelerated Coordinate Encoding for Large-Scale Scene Localization and Rendering
- cubic: CUDA-accelerated 3D Bioimage Computing
- ViTacGen: Robotic Pushing with Vision-to-Touch Generation
- NTIRE 2025 Challenge on Low Light Image Enhancement: Methods and Results
- An efficient approach with theoretical guarantees to simultaneously reconstruct activity and attenuation sinogram for TOF-PET
- MRSeqStudio: MRI Sequence Design and Simulation as a Service in a Free and Open-Source Web Platform
- SWIR-LightFusion: Multi-spectral Semantic Fusion of Synthetic SWIR with Thermal IR (LWIR/MWIR) and RGB
- Edit-Your-Interest: Efficient Video Editing via Feature Most-Similar Propagation
- No-Reference Rendered Video Quality Assessment: Dataset and Metrics
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- SimULi: Real-Time LiDAR and Camera Simulation with Unscented Transforms
- Wavefront Coding for Accommodation-Invariant Near-Eye Displays
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- DrivingScene: A Multi-Task Online Feed-Forward 3D Gaussian Splatting Method for Dynamic Driving Scenes
- Readout Representation: Redefining Neural Codes by Input Recovery
- UniGS: Unified Geometry-Aware Gaussian Splatting for Multimodal Rendering
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- MaterialRefGS: Reflective Gaussian Splatting with Multi-view Consistent Material Inference
- Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Perspective-aware 3D Gaussian Inpainting with Multi-view Consistency
- NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes
- Bit Allocation Transfer for Perceptual Quality Enhancement of VVC Intra Coding
- Blade: A Derivative-free Bayesian Inversion Method using Diffusion Priors
- Towards Distribution-Shift Uncertainty Estimation for Inverse Problems with Generative Priors
- Normalization-equivariant Diffusion Models: Learning Posterior Samplers From Noisy And Partial Measurements
- Dynamic Gaussian Splatting from Defocused and Motion-blurred Monocular Videos
- JND-Guided Light-Weight Neural Pre-Filter for Perceptual Image Coding
- Towards Efficient 3D Gaussian Human Avatar Compression: A Prior-Guided Framework
- VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
- Sketch Animation: State-of-the-art Report
- Gesplat: Robust Pose-Free 3D Reconstruction via Geometry-Guided Gaussian Splatting
- Enabling High-Quality In-the-Wild Imaging from Severely Aberrated Metalens Bursts
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- CLoD-GS: Continuous Level-of-Detail via 3D Gaussian Splatting
- BurstDeflicker: A Benchmark Dataset for Flicker Removal in Dynamic Scenes
- FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D Rendering
- Generative Latent Video Compression
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- A Style-Based Profiling Framework for Quantifying the Synthetic-to-Real Gap in Autonomous Driving Datasets
- Simulation-based inference at the theoretical limit for fast, robust microstructural MRI with minimal diffusion data
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- Harnessing Self-Supervised Deep Learning and Geostationary Remote Sensing for Advancing Wildfire and Associated Air Quality Monitoring: Improved Smoke and Fire Front Masking using GOES and TEMPO Radiance Data
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- SynthID-Image: Image watermarking at internet scale
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
- PC-UNet: An Enforcing Poisson Statistics U-Net for Positron Emission Tomography Denoising
- Defense against Unauthorized Distillation in Image Restoration via Feature Space Perturbation
- Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
- Median2Median: Zero-shot Suppression of Structured Noise in Images
- HeadsUp! High-Fidelity Portrait Image Super-Resolution
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense
- X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
- FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
- Spectral Prefiltering of Neural Fields
- Variable-Rate Texture Compression: Real-Time Rendering with JPEG
- SatFusion: A Unified Framework for Enhancing Remote Sensing Images via Multi-Frame and Multi-Source Images Fusion
- FlowLensing: Simulating Gravitational Lensing with Flow Matching
- Interlaced dynamic XCT reconstruction with spatio-temporal implicit neural representations
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction
- PIT-QMM: A Large Multimodal Model For No-Reference Point Cloud Quality Assessment
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Biology-driven assessment of deep learning super-resolution imaging of the porosity network in dentin
- Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection
- ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
- Splat the Net: Radiance Fields with Splattable Neural Primitives
- Low-Compute Watermark Removal via Dual-Domain Natural Projection
- Coupled Data and Measurement Space Dynamics for Enhanced Diffusion Posterior Sampling
- WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
- MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
- MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
- MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics
- Accelerating wave simulations with neural dispersion correctors
- SCas4D: Structural Cascaded Optimization for Boosting Persistent 4D Novel View Synthesis
- RTGS: Real-Time 3D Gaussian Splatting SLAM via Multi-Level Redundancy Reduction
- SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
- How Confident are Video Models? Empowering Video Models to Express their Uncertainty
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- Active Next-Best-View Optimization for Risk-Averse Path Planning
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- Mapping surface height dynamics to subsurface flow physics in free-surface turbulent flow using a shallow recurrent decoder
- Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis
- Life at the extremes: maximally divergent microbes with similar genomic signatures linked to extreme environments
- Learning the detector in optical tomography
- Learning a distance measure from the information-estimation geometry of data
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- CLEAR-IR: Clarity-Enhanced Active Reconstruction of Infrared Imagery
- MACS: Measurement-Aware Consistency Sampling for Inverse Problems
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- No-reference Quality Assessment of Contrast-distorted Images using Contrast-enhanced Pseudo Reference
- QuantDemoire: Quantization with Outlier Aware for Image Demoiréing
- Adaptive double-phase Rudin--Osher--Fatemi denoising model
- The 1st Solution for CARE Liver Task Challenge 2025: Contrast-Aware Semi-Supervised Segmentation with Domain Generalization and Test-Time Adaptation
- Super-resolution image projection over an extended depth of field using a diffractive decoder
- DHQA-4D: Perceptual Quality Assessment of Dynamic 4D Digital Human
- Contrastive-SDE: Guiding Stochastic Differential Equations with Contrastive Learning for Unpaired Image-to-Image Translation
- LoRA Patching: Exposing the Fragility of Proactive Defenses against Deepfakes
- Exploring Instruction Data Quality for Explainable Image Quality Assessment
- AortaDiff: A Unified Multitask Diffusion Framework For Contrast-Free AAA Imaging
- HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
- Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted Perturbations
- PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
- Towards Scalable and Consistent 3D Editing
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- GS-Share: Enabling High-fidelity Map Sharing with Incremental Gaussian Splatting
- Net2Net: When Un-trained Meets Pre-trained Networks for Robust Real-World Denoising
- From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting
- Image Enhancement Based on Pigment Representation
- FSFSplatter: Build Surface and Novel Views with Sparse-Views within 2min
- Longitudinal Flow Matching for Trajectory Modeling
- LVTINO: LAtent Video consisTency INverse sOlver for High Definition Video Restoration
- EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
- Visual Self-Refinement for Autoregressive Models
- Equivariant Splitting: Self-supervised learning from incomplete data
- DIA: The Adversarial Exposure of Deterministic Inversion in Diffusion Models
- SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
- Geometric Spatio-Spectral Total Variation for Hyperspectral Image Denoising and Destriping
- ZQBA: Zero Query Black-box Adversarial Attack
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
- PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
- Universal Beta Splatting
- Observer-Usable Information as a Task-specific Image Quality Metric
- MOLM: Mixture of LoRA Markers
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- HART: Human Aligned Reconstruction Transformer
- PRISM: Progressive Rain removal with Integrated State-space Modeling
- AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image Manipulation
- GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
- Flow Matching with Semidiscrete Couplings
- FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- MANI-Pure: Magnitude-Adaptive Noise Injection for Adversarial Purification
- Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- Environment-Aware Satellite Image Generation with Diffusion Models
- NeuralPVS: Learned Estimation of Potentially Visible Sets
- RIFLE: Removal of Image Flicker-Banding via Latent Diffusion Enhancement
- SAIP: A Plug-and-Play Scale-adaptive Module in Diffusion-based Inverse Problems
- OMeGa: Joint Optimization of Explicit Meshes and Gaussian Splats for Robust Scene-Level Surface Reconstruction
- Asymmetric VAE for One-Step Video Super-Resolution Acceleration
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- Synergizing scientific and local knowledge for ecosystem services assessments: A case study in northern Portugal
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
- Tunable-Generalization Diffusion Powered by Self-Supervised Contextual Sub-Data for Low-Dose CT Reconstruction
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Latent Representation Learning from 3D Brain MRI for Interpretable Prediction in Multiple Sclerosis
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- FlowLUT: Efficient Image Enhancement via Differentiable LUTs and Iterative Flow Matching
- Foundation Model-Based Adaptive Semantic Image Transmission for Dynamic Wireless Environments
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- FM-SIREN & FM-FINER: Nyquist-Informed Frequency Multiplier for Implicit Neural Representation with Periodic Activation
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- Enhanced Quality Aware-Scalable Underwater Image Compression
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
- Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
- Enhancing Blind Face Restoration through Online Reinforcement Learning
- WoW: Towards a World omniscient World model Through Embodied Interaction
- CCNeXt: An Effective Self-Supervised Stereo Depth Estimation Approach
- SSVIF: Self-Supervised Segmentation-Oriented Visible and Infrared Image Fusion
- LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer
- FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
- Clinical Uncertainty Impacts Machine Learning Evaluations
- DragGANSpace: Latent Space Exploration and Control for GANs
- REFINE-CONTROL: A Semi-supervised Distillation Method For Conditional Image Generation
- No-Reference Image Contrast Assessment with Customized EfficientNet-B0
- TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
- SRHand: Super-Resolving Hand Images and 3D Shapes via View/Pose-aware Neural Image Representations and Explicit 3D Meshes
- RED-DiffEq: Regularization by denoising diffusion models for solving inverse PDE problems with application to full waveform inversion
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?
- A Hierarchical Variational Graph Fused Lasso for Recovering Relative Rates in Spatial Compositional Data
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- Efficient Encoder-Free Pose Conditioning and Pose Control for Virtual Try-On
- An Anisotropic Cross-View Texture Transfer with Multi-Reference Non-Local Attention for CT Slice Interpolation
- JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation
- Downscaling climate projections to 1 km with single-image super resolution
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
- SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- Talking Head Generation via AU-Guided Landmark Prediction
- Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
- Don't Trust the AI Ecosystem: Analyzing Privacy Leakage in Compromised Open-Source Components
- CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
- FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
- S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
- RFMSR: Residual Flow Matching for Image Super-Resolution
- FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
- Collaborative feature aggregation for face super-resolution and robust re-identification
- Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
- Convolutional neural shading for high-quality 3D reconstruction from multi-view images
- What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
- Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
- Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI
- Generative Relightable Avatars
- SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
- Habitability Study of Terrestrial Planets: Application to Venus-like Worlds
- StrIPETrack : a real-time, ROI-flexible tracking platform for high-throughput zebrafish behavior
- Dark3R: Learning Structure from Motion in the Dark
- Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured Guidance
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- SeHDR: Single-Exposure HDR Novel View Synthesis via 3D Gaussian Bracketing
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
- Audio-Driven Universal Gaussian Head Avatars
- DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
- FlowCrypt: Flow-Based Lightweight Encryption with Near-Lossless Recovery for Cloud Photo Privacy
- NGD: Neural Gradient Based Deformation for Monocular Garment Reconstruction
- CATformer: Contrastive Adversarial Transformer for Image Super-Resolution
- GeoRemover: Removing Objects and Their Causal Visual Artifacts
- FixingGS: Enhancing 3D Gaussian Splatting via Training-Free Score Distillation
- SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack against No-Reference Image Quality Assessment Models
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
- Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance
- 4D-MoDe: Towards Editable and Scalable Volumetric Streaming via Motion-Decoupled 4D Gaussian Compression
- Audio Super-Resolution with Latent Bridge Models
- Learning Neural Antiderivatives
- RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
- Clothing agnostic Pre-inpainting Virtual Try-ON
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- Modeling Bottom-up Information Quality during Language Processing
- Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
- SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
- PhysHDR: When Lighting Meets Materials and Scene Geometry in HDR Reconstruction
- HyRF: Hybrid Radiance Fields for Memory-efficient and High-quality Novel View Synthesis
- Task-Oriented Communications for 3D Scene Representation: Balancing Timeliness and Fidelity
- MobileUPReg: Identifying User-Perceived Performance Regressions in Mobile OS Versions
- A Multi-Grid Implicit Neural Representation for Multi-View Videos
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging
- Neural Atlas Graphs for Dynamic Scene Decomposition and Editing
- Uniform 2D Target Generation via Inverse-designed Metasurfaces
- Spatio-temporal, multi-field deep learning of shock propagation in meso-structured media
- SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Physics-Informed GCN-LSTM Framework for Long-Term Forecasting of 2D and 3D Microstructure Evolution
- SPATIALGEN: Layout-guided 3D Indoor Scene Generation
- Controllable Localized Face Anonymization Via Diffusion Inpainting
- Dataset Distillation for Super-Resolution without Class Labels and Pre-trained Models
- The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
- ProFusion: 3D Reconstruction of Protein Complex Structures from Multi-view AFM Images
- LamiGauss: Pitching Radiative Gaussian for Sparse-View X-ray Laminography Reconstruction
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- A Hybrid Approach for Unified Image Quality Assessment: Permutation Entropy-Based Features Fused with Random Forest for Natural-Scene and Screen-Content Images for Cross-Content Applications
- Assessing Data Replication in Symbolic Music via Adapted Structural Similarity Index Measure
- Stochastic ion emission perturbation mechanisms in atom probe tomography: Linking simulations to experiment
- UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
- FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
- DEFT-VTON: Efficient Virtual Try-On with Consistent Generalised H-Transform
- Improving 3D Gaussian Splatting Compression by Scene-Adaptive Lattice Vector Quantization
- Temporally Smooth Mesh Extraction for Procedural Scenes with Long-Range Camera Trajectories using Spacetime Octrees
- Deep learning approach for flow visualization in background-oriented schlieren
- Improving Muon Scattering Tomography Performance With A Muon Momentum Measurement Scheme
- Deep Learning Architectures for Medical Image Denoising: A Comparative Study of CNN-DAE, CADTra, and DCMIEDNet
- Evaluation of Objective Image Quality Metrics for High-Fidelity Image Compression
- A Statistical Benchmark for Diffusion Posterior Sampling Algorithms
- Uncovering and Mitigating Destructive Multi-Embedding Attacks in Deepfake Proactive Forensics
- Unrolling Graph-based Douglas-Rachford Algorithm for Image Interpolation with Informed Initialization
- Impact of a Sharpness Based Loss Function for Removing Out-of-Focus Blur
- LaGarNet: Goal-Conditioned Recurrent State-Space Models for Pick-and-Place Garment Flattening
- Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision
- ROSGS: Relightable Outdoor Scenes With Gaussian Splatting
- SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
- Synthetic Dataset Evaluation Based on Generalized Cross Validation
- Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
- Every Camera Effect, Every Time, All at Once: 4D Gaussian Ray Tracing for Physics-based Camera Effect Data Generation
- Mask Consistency Regularization in Object Removal
- A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss
- Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge
- RPD-Diff: Region-Adaptive Physics-Guided Diffusion Model for Visibility Enhancement under Dense and Non-Uniform Haze
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- Data-Driven Discovery of Emergent Dynamics in Reaction-Diffusion Systems from Sparse and Noisy Observations
- VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results
- Objectness Similarity: Capturing Object-Level Fidelity in 3D Scene Evaluation
- MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
- Investigating the Impact of Various Loss Functions and Learnable Wiener Filter for Laparoscopic Image Desmoking
- Images in Motion?: A First Look into Video Leakage in Collaborative Deep Learning
- An U-Net-Based Deep Neural Network for Cloud Shadow and Sun-Glint Correction of Unmanned Aerial System (UAS) Imagery
- GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- Using machine learning to downscale coarse-resolution environmental variables for understanding the spatial frequency of convective storms
- AdsQA: Towards Advertisement Video Understanding
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- First-order State Space Model for Lightweight Image Super-resolution
- Sparse Transformer for Ultra-sparse Sampled Video Compressive Sensing
- LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
- Spectral Bottleneck in Sinusoidal Representation Networks: Noise is All You Need
- SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting
- DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- XOCT: Enhancing OCT to OCTA Translation via Cross-Dimensional Supervised Multi-Scale Feature Learning
- Faster, Self-Supervised Super-Resolution for Anisotropic Multi-View MRI Using a Sparse Coordinate Loss
- GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs
- Privacy Preserving Semantic Communications Using Vision Language Models: A Segmentation and Generation Approach
- Evaluation of Machine Learning Reconstruction Techniques for Accelerated Brain MRI Scans
- Diffusion-Shock PDEs for Deep Learning on Position-Orientation Space
- Perception-oriented Bidirectional Attention Network for Image Super-resolution Quality Assessment
- AIM 2025 Challenge on High FPS Motion Deblurring: Methods and Results
- Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
- Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution
- SpecSwin3D: Generating Hyperspectral Imagery from Multispectral Data via Transformer Networks
- Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance
- Hybrid-illumination multiplexed Fourier ptychographic microscopy with robust aberration correction
- Systematic Review and Meta-analysis of AI-driven MRI Motion Artifact Detection and Correction
- Reverse Browser: Vector-Image-to-Code Generator
- CoRe-GS: Coarse-to-Refined Gaussian Splatting with Semantic Object Focus
- Human Motion Video Generation: A Survey
- A Neural Network Approach to Multi-radionuclide TDCR Beta Spectroscopy
- PromptFlare: Prompt-Generalized Defense via Cross-Attention Decoy in Diffusion-Based Inpainting
- OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
- Deep learning-enabled virtual multiplexed immunostaining of label-free tissue for vascular invasion assessment
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data
- Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
- Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup Tables
- MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space Optimization
- Unsupervised Ultra-High-Resolution UAV Low-Light Image Enhancement: A Benchmark, Metric and Framework
- PRINTER:Deformation-Aware Adversarial Learning for Virtual IHC Staining with In Situ Fidelity
- DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion
- A Unified Low-level Foundation Model for Enhancing Pathology Image Quality
- Seeing through Unclear Glass: Occlusion Removal with One Shot
- Delta Rectified Flow Sampling for Text-to-Image Editing
- GaussianGAN: Real-Time Photorealistic controllable Human Avatars
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
- UPGS: Unified Pose-aware Gaussian Splatting for Dynamic Scene Deblurring
- DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
- DAOVI: Distortion-Aware Omnidirectional Video Inpainting
- LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables
- Low-Rank Regularized Convex-Non-Convex Problems for Image Segmentation or Completion
- Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
- ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
- Diverse Signer Avatars with Manual and Non-Manual Feature Modelling for Sign Language Production
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Unfolding Framework with Complex-Valued Deformable Attention for High-Quality Computer-Generated Hologram Generation
- First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge
- FusionCounting: Robust visible-infrared image fusion guided by crowd counting via multi-task learning
- Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
- CoCoL: A Communication Efficient Decentralized Collaborative Method for Multi-Robot Systems
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- C3-GS: Learning Context-aware, Cross-dimension, Cross-scale Feature for Generalizable Gaussian Splatting
- Disruptive Attacks on Face Swapping via Low-Frequency Perceptual Perturbations
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration
- FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
- Beyond Imaging: Vision Transformer Digital Twin Surrogates for 3D+T Biological Tissue Dynamics
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- StableIntrinsic: Detail-preserving One-step Diffusion Model for Multi-view Material Estimation
- IDF: Iterative Dynamic Filtering Networks for Generalizable Image Denoising
- High-Speed FHD Full-Color Video Computer-Generated Holography
- High-Frequency First: A Two-Stage Approach for Improving Image INR
- Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
- Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion
- VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- Mini-Batch Robustness Verification of Deep Neural Networks
- RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration
- Understanding Benefits and Pitfalls of Current Methods for the Segmentation of Undersampled MRI Data
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- ROSE: Remove Objects with Side Effects in Videos
- Wan-S2V: Audio-Driven Cinematic Video Generation
- Evaluating the visual design of science publications—a quantitative approach comparing legitimate and predatory journal papers
- FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained Poses
- FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction
- Zero-shot CT Super-Resolution using Diffusion-based 2D Projection Priors and Signed 3D Gaussians
- Image-Conditioned 3D Gaussian Splat Quantization
- SPIRiT Regularization: Parallel MRI with a Combination of Sensitivity Encoding and Linear Predictability
- HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
- Systematic Evaluation of Wavelet-Based Denoising for MRI Brain Images: Optimal Configurations and Performance Benchmarks
- Snap-Snap: Taking Two Images to Reconstruct 3D Human Gaussians in Milliseconds
- CUTE-MRI: Conformalized Uncertainty-based framework for Time-adaptivE MRI
- Entanglement-enhanced imaging through scattering media
- Fine-grained Image Quality Assessment for Perceptual Image Restoration
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
- Taming Transformer for Emotion-Controllable Talking Face Generation
- Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning
- Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
- Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting
- GOGS: High-Fidelity Geometry and Relighting for Glossy Objects via Gaussian Surfels
- LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
- Distilled-3DGS:Distilled 3D Gaussian Splatting
- Online 3D Gaussian Splatting Modeling with Novel View Selection
- Self-Supervised Sparse Sensor Fusion for Long Range Perception
- DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-Experts
- Forecasting Smog Events Using ConvLSTM: A Spatio-Temporal Approach for Aerosol Index Prediction in South Asia
- Comparing Conditional Diffusion Models for Synthesizing Contrast-Enhanced Breast MRI from Pre-Contrast Images
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
- Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
- OmniTry: Virtual Try-On Anything without Masks
- MF-LPR2: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
- FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
- A Convergent Primal-Dual Algorithm for Computing Rate-Distortion-Perception Functions
- EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
- Discrete Optimization of Min-Max Violation and its Applications Across Computational Sciences
- Precise Action-to-Video Generation Through Visual Action Prompts
- From Transthoracic to Transesophageal: Cross-Modality Generation using LoRA Diffusion
- Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping
- Frequency-Driven Inverse Kernel Prediction for Single Image Defocus Deblurring
- Neural Rendering for Sensor Adaptation in 3D Object Detection
- DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examples
- Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
- Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- DMS:Diffusion-Based Multi-Baseline Stereo Generation for Improving Self-Supervised Depth Estimation
- Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting
- IntelliCap: Intelligent Guidance for Consistent View Sampling
- ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- Geometry-Aware Video Inpainting for Joint Headset Occlusion Removal and Face Reconstruction in Social XR
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- Demystifying Foreground-Background Memorization in Diffusion Models
- DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
- Guiding WaveMamba with Frequency Maps for Image Debanding
- Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
- Tactile Robotics: An Outlook
- Meta-learning Structure-Preserving Dynamics
- Temporally-Similar Structure-Aware Spatiotemporal Fusion of Satellite Images
- Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2)
- Self-Supervised Stereo Matching with Multi-Baseline Contrastive Learning
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Novel View Synthesis using DDIM Inversion
- Physics-Informed Joint Multi-TE Super-Resolution with Implicit Neural Representation for Robust Fetal T2 Mapping
- HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis
- Physics-Informed Deep Contrast Source Inversion: A Unified Framework for Inverse Scattering Problems
- Multi-Sample Anti-Aliasing and Constrained Optimization for 3D Gaussian Splatting
- TweezeEdit: Consistent and Efficient Image Editing with Path Regularization
- Efficient Image Denoising Using Global and Local Circulant Representation
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Hybrid Quantum-Classical Latent Diffusion Models for Medical Image Generation
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- Hierarchical Graph Attention Network for No-Reference Omnidirectional Image Quality Assessment
- In-place Double Stimulus Methodology for Subjective Assessment of High Quality Images
- MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- DualPhys-GS: Dual Physically-Guided 3D Gaussian Splatting for Underwater Scene Reconstruction
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite Images
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- HyperKD: Distilling Cross-Spectral Knowledge in Masked Autoencoders via Inverse Domain Shift with Spatial-Aware Masking and Specialized Loss
- HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- A Generative Imputation Method for Multimodal Alzheimer's Disease Diagnosis
- Harnessing Input-Adaptive Inference for Efficient VLN
- TaoCache: Structure-Maintained Video Generation Acceleration
- How Does a Virtual Agent Decide Where to Look? -- Symbolic Cognitive Reasoning for Embodied Head Rotation
- DiffPhysCam: Differentiable Physics-Based Camera Simulation for Inverse Rendering and Embodied AI
- A Parametric Bi-Directional Curvature-Based Framework for Image Artifact Classification and Quantification
- Exploring Palette based Color Guidance in Diffusion Models
- Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
- MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion
- SelfHVD: Self-Supervised Handheld Video Deblurring
- Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
- MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
- IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
- Reinforcement Learning for Large Model: A Survey
- Learned Regularization for Microwave Tomography
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Sea-Undistort: A Dataset for Through-Water Image Restoration in High Resolution Airborne Bathymetric Mapping
- QuantEIT: Ultra-Lightweight Quantum-Assisted Inference for Chest Electrical Impedance Tomography
- Multi-view Normal and Distance Guidance Gaussian Splatting for Surface Reconstruction
- Undress to Redress: A Training-Free Framework for Virtual Try-On
- BadPromptFL: A Novel Backdoor Threat to Prompt-based Federated Learning in Multimodal Models
- OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
- Decoupled Functional Evaluation of Autonomous Driving Models via Feature Map Quality Scoring
- Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
- Similarity Matters: A Novel Depth-guided Network for Image Restoration and A New Dataset
- CMAMRNet: A Contextual Mask-Aware Network Enhancing Mural Restoration Through Comprehensive Mask Guidance
- 3DGS-VBench: A Comprehensive Video Quality Evaluation Benchmark for 3DGS Compression
- CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Application of Noise2Inverse and adaptation (Noise2Phase) to single‐mask x‐ray phase contrast micro‐computed tomography
- FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
- FedMeNF: Privacy-Preserving Federated Meta-Learning for Neural Fields
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
- Synthetic Data Generation for Emotional Depth Faces: Optimizing Conditional DCGANs via Genetic Algorithms in the Latent Space and Stabilizing Training with Knowledge Distillation
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- Task complexity shapes internal representations and robustness in neural networks
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
- Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- FLUX-Makeup: High-Fidelity, Identity-Consistent, and Robust Makeup Transfer via Diffusion Transformer
- UGOD: Uncertainty-Guided Differentiable Opacity and Soft Dropout for Enhanced Sparse-View 3DGS
- Laplacian Analysis Meets Dynamics Modelling: Gaussian Splatting for 4D Reconstruction
- Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
- FedMP: Tackling Medical Feature Heterogeneity in Federated Learning from a Manifold Perspective
- MZEN: Multi-Zoom Enhanced NeRF for 3-D Reconstruction with Unknown Camera Poses
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- RetinexDual: Retinex-based Dual Nature Approach for Generalized Ultra-High-Definition Image Restoration
- Combined Image Data Augmentations diminish the benefits of Adaptive Label Smoothing
- UniTalker: Conversational Speech-Visual Synthesis
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- QuantVSR: Low-Bit Post-Training Quantization for Real-World Video Super-Resolution
- From Flat to Round: Redefining Brain Decoding with Surface-Based fMRI and Cortex Structure
- Dyna3DGR: 4D Cardiac Motion Tracking with Dynamic 3D Gaussian Representation
- PIS3R: Very Large Parallax Image Stitching via Deep 3D Reconstruction
- Uncertainty-Aware Spatial Color Correlation for Low-Light Image Enhancement
- DET-GS: Depth- and Edge-Aware Regularization for High-Fidelity 3D Gaussian Splatting
- Age-Diverse Deepfake Dataset: Bridging the Age Gap in Deepfake Detection
- Towards Globally Predictable k-Space Interpolation: A White-box Transformer Approach
- SPJFNet: Self-Mining Prior-Guided Joint Frequency Enhancement for Ultra-Efficient Dark Image Restoration
- CADD: Context aware disease deviations via restoration of brain images using normative conditional diffusion models
- EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware Evaluation
- Diffusion Once and Done: Degradation-Aware LoRA for Efficient All-in-One Image Restoration
- CIVQLLIE: Causal Intervention with Vector Quantization for Low-Light Image Enhancement
- Beyond Illumination: Fine-Grained Detail Preservation in Extreme Dark Image Restoration
- Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems
- BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
- Duplex-GS: Proxy-Guided Weighted Blending for Real-Time Order-Independent Gaussian Splatting
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- RobustGS: Unified Boosting of Feedforward 3D Gaussian Splatting under Low-Quality Conditions
- Low-rankness and Smoothness Meet Subspace: A Unified Tensor Regularization for Hyperspectral Image Super-resolution
- SA-3DGS: A Self-Adaptive Compression Method for 3D Gaussian Splatting
- ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion
- Managing Data for Scalable and Interactive Event Sequence Visualization
- Learning to Incentivize: LLM-Empowered Contract for AIGC Offloading in Teleoperation
- Evaluation of 3D Counterfactual Brain MRI Generation
- DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
- From Pixels to Pathology: Restoration Diffusion for Diagnostic-Consistent Virtual IHC
- GR-Gaussian: Graph-Based Radiative Gaussian Splatting for Sparse-View CT Reconstruction
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- After the Party: Navigating the Mapping From Color to Ambient Lighting
- DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal
- QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots
- Frequency-Domain Denoising-Based in Vivo Fluorescence Imaging
- PMGS: Reconstruction of Projectile Motion Across Large Spatiotemporal Spans via 3D Gaussian Splatting
- Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure
- PoseGuard: Pose-Guided Generation with Safety Guardrails
- A Neural Quality Metric for BRDF Models
- Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
- PRIMU: Uncertainty Estimation for Novel Views in Gaussian Splatting from Primitive-Based Representations of Error and Coverage
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Beyond Vulnerabilities: A Survey of Adversarial Attacks as Both Threats and Defenses in Computer Vision Systems
- Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models
- Physically-based Lighting Generation for Robotic Manipulation
- Integrating Disparity Confidence Estimation into Relative Depth Prior-Guided Unsupervised Stereo Matching
- OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGS
- MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial Geometry
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- Screencast-Based Analysis of User-Perceived GUI Responsiveness
- The Promise of RL for Autoregressive Image Editing
- MASIV: Toward Material-Agnostic System Identification from Videos
- DreamSat-2.0: Towards a General Single-View Asteroid 3D Reconstruction
- LeakyCLIP: Extracting Training Data from CLIP
- FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- Video Color Grading via Look-Up Table Generation
- PIF-Net: Ill-Posed Prior Guided Multispectral and Hyperspectral Image Fusion via Invertible Mamba and Fusion-Aware LoRA
- IN2OUT: Fine-Tuning Video Inpainting Model for Video Outpainting Using Hierarchical Discriminator
- SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth Consistencies
- Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars
- Exploring Fourier Prior and Event Collaboration for Low-Light Image Enhancement
- Sample-Aware Test-Time Adaptation for Medical Image-to-Image Translation
- Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
- Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization
- NeRF Is a Valuable Assistant for 3D Gaussian Splatting
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads
- Towards Measuring and Modeling Geometric Structures in Time Series Forecasting via Image Modality
- EMORe: Motion-Robust 5D MRI Reconstruction via Expectation-Maximization-Guided Binning Correction and Outlier Rejection
- Adversarial-Guided Diffusion for Multimodal LLM Attacks
- LIVE-GS: Online LiDAR-Inertial-Visual State Estimation and Globally Consistent Mapping with 3D Gaussian Splatting
- BS-1-to-N: Diffusion-Based Environment-Aware Cross-BS Channel Knowledge Map Generation for Cell-Free Networks
- UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restoration
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
- Towards High-Resolution Alignment and Super-Resolution of Multi-Sensor Satellite Imagery
- Ray-tracing image simulations of transparent objects with complex shape and inhomogeneous refractive index
- MRpro - open PyTorch-based MR reconstruction and processing package
- DWTGS: Rethinking Frequency Regularization for Sparse-view 3D Gaussian Splatting
- Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
- Gaussian Splatting with Discretized SDF for Relightable Assets
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
- UFV-Splatter: Pose-Free Feed-Forward 3D Gaussian Splatting Adapted to Unfavorable Views
- LAMA-Net: A Convergent Network Architecture for Dual-Domain Reconstruction
- Trade-offs in Image Generation: How Do Different Dimensions Interact?
- SAIGFormer: A Spatially-Adaptive Illumination-Guided Network for Low-Light Image Enhancement
- Neural network enabled wide field-of-view imaging with hyperbolic metalenses
- Text-Aware Image Restoration with Diffusion Models
- Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
- Endoscopic Depth Estimation Based on Deep Learning: A Survey
- Compositional Video Synthesis by Temporal Object-Centric Learning
- Onboard Hyperspectral Super-Resolution with Deep Pushbroom Neural Network
- Benefits of Feature Extraction and Temporal Sequence Analysis for Video Frame Prediction: An Evaluation of Hybrid Deep Learning Models
- Towards trustworthy AI in materials mechanics through domain-guided attention
- Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
- Harnessing Diffusion-Yielded Score Priors for Image Restoration
- Annotation-Free Human Sketch Quality Assessment
- WEEP: A Differentiable Nonconvex Sparse Regularizer via Weakly-Convex Envelope
- JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
- ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control
- Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
- Decomposing Densification in Gaussian Splatting for Faster 3D Scene Reconstruction
- GT-Mean Loss: A Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image Enhancement
- SuperMag: Vision-based Tactile Data Guided High-resolution Tactile Shape Reconstruction for Magnetic Tactile Sensors
- CHADET: Cross-Hierarchical-Attention for Depth-Completion Using Unsupervised Lightweight Transformer
- A Generative Model for Disentangling Galaxy Photometric Parameters
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- Taking Language Embedded 3D Gaussian Splatting into the Wild
- ForCenNet: Foreground-Centric Network for Document Image Rectification
- MoFRR: Mixture of Diffusion Models for Face Retouching Restoration
- Tuning adaptive gamma correction (TAGC) for enhancing images in low ligh
- ChartM3: Benchmarking Chart Editing with Multimodal Instructions
- WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design
- DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering
- Cross-Subject Mind Decoding from Inaccurate Representations
- Fast Learning of Non-Cooperative Spacecraft 3D Models through Primitive Initialization
- Plug and Play Splitting Techniques for Poisson Image Restoration
- Dual Path Learning -- learning from noise and context for medical image denoising
- Learning Efficient and Generalizable Human Representation with Human Gaussian Model
- DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
- Diffusion models for multivariate subsurface generation and efficient probabilistic inversion
- PS-GS: Gaussian Splatting for Multi-View Photometric Stereo
- GEDTM30: global ensemble digital terrain model at 30 m and derived multiscale terrain variables
- G2S-ICP SLAM: Geometry-aware Gaussian Splatting ICP SLAM
- U-Net Based Healthy 3D Brain Tissue Inpainting
- Adapting Large VLMs with Iterative and Manual Instructions for Generative Low-light Enhancement
- Degradation-Consistent Learning via Bidirectional Diffusion for Low-Light Image Enhancement
- OmniVTON: Training-Free Universal Virtual Try-On
- Grounding Degradations in Natural Language for All-In-One Video Restoration
- SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- VisGuard: Securing Visualization Dissemination through Tamper-Resistant Data Retrieval
- PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
- InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling
- Accelerating Parallel Diffusion Model Serving with Residual Compression
- Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
- Perceptual Classifiers: Detecting Generative Images using Perceptual Features
- Efficient Burst Super-Resolution with One-step Diffusion
- UNICE: Training A Universal Image Contrast Enhancer
- Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
- Improving Multislice Electron Ptychography with a Generative Prior
- DFDNet: Dynamic Frequency-Guided De-Flare Network
- Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
- HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
- DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
- FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers
- Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
- Multi-Component VAE with Gaussian Markov Random Field
- SmokeSVD: Smoke Reconstruction from A Single View via Progressive Novel View Synthesis and Refinement with Diffusion Models
- EC-Diff: Fast and High-Quality Edge-Cloud Collaborative Inference for Diffusion Models
- CompressedVQA-HDR: Generalized Full-reference and No-reference Quality Assessment Models for Compressed High Dynamic Range Videos
- Pathology-Guided Virtual Staining Metric for Evaluation and Training
- Wavelet-GS: 3D Gaussian Splatting with Wavelet Decomposition
- Generate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models
- BRUM: Robust 3D Vehicle Reconstruction from 360 Sparse Images
- HPR3D: Hierarchical Proxy Representation for High-Fidelity 3D Reconstruction and Controllable Editing
- Interpretable Prediction of Lymph Node Metastasis in Rectal Cancer MRI Using Variational Autoencoders
- Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
- Deep Equilibrium models for Poisson Imaging Inverse problems via Mirror Descent
- SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields
- SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
- HiSin: A Sinogram-Aware Framework for Efficient High-Resolution Inpainting
- Gaussian2Scene: 3D Scene Representation Learning via Self-supervised Learning with 3D Gaussian Splatting
- MFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection
- A Neural Network Model of Complementary Learning Systems: Pattern Separation and Completion for Continual Learning
- Physically Based Neural LiDAR Resimulation
- WaFusion: A Wavelet-Enhanced Diffusion Framework for Face Morph Generation
- Quantize-then-Rectify: Efficient VQ-VAE Training
- 3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
- VoxelRF: Voxelized Radiance Field for Fast Wireless Channel Modeling
- Ambient Diffusion Omni: Training Good Models with Bad Data
- I2I-PR: Deep Iterative Refinement for Phase Retrieval using Image-to-Image Diffusion Models
- When Schrödinger Bridge Meets Real-World Image Dehazing with Unpaired Training
- RectifiedHR: High-Resolution Diffusion via Energy Profiling and Adaptive Guidance Scheduling
- AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
- Visual Surface Wave Elastography: Revealing Subsurface Physical Properties via Visible Surface Waves
- Physics-Aware Fluid Field Generation from User Sketches Using Helmholtz-Hodge Decomposition
- Generalizable 7T T1-map Synthesis from 1.5T and 3T T1 MRI with an Efficient Transformer Model
- Towards Imperceptible JPEG Image Hiding: Multi-range Representations-driven Adversarial Stego Generation
- M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
- RePaintGS: Reference-Guided Gaussian Splatting for Realistic and View-Consistent 3D Scene Inpainting
- D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping
- Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
- Multigranular Evaluation for Brain Visual Decoding
- Dense Temporal Contrast Synthesis via Conditioned Latent Transport
- D-CNN and VQ-VAE Autoencoders for Compression and Denoising of Industrial X-ray Computed Tomography Images
- Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
- Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
- EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction
- Seg-Wild: Interactive Segmentation based on 3D Gaussian Splatting for Unconstrained Image Collections
- SD-GS: Structured Deformable 3D Gaussians for Efficient Dynamic Scene Reconstruction
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
- QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
- IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation Optimization
- Democratizing High-Fidelity Co-Speech Gesture Video Generation
- Residual Prior-driven Frequency-aware Network for Image Fusion
- FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting
- Capturing Stable HDR Videos Using a Dual-Camera System
- Deep Learning-based Human Gesture Channel Modeling for Integrated Sensing and Communication Scenarios
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement
- DreamArt: Generating Interactable Articulated Objects from a Single Image
- Resolving Extreme Data Scarcity by Explicit Physics Integration: An Application to Groundwater Heat Transport
- Kernel Density Steering: Inference-Time Scaling via Mode Seeking for Image Restoration
- Generative Panoramic Image Stitching
- D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
- TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
- SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model
- Semantic Frame Interpolation
- TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
- SPIDER: Structure-Preferential Implicit Deep Network for Biplanar X-ray Reconstruction
- A Deep Unfolding Framework for Diffractive Snapshot Spectral Imaging
- Backdoors in Conditional Diffusion: Threats to Responsible Synthetic Data Pipelines
- Uncovering Neuroimaging Biomarkers of Brain Tumor Surgery with AI-Driven Methods
- Enhancing Underwater Images Using Deep Learning with Subjective Image Quality Integration
- GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction
- A View-consistent Sampling Method for Regularized Training of Neural Radiance Fields
- Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
- DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
- Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
- Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
- Evaluating Adversarial Protections for Diffusion Personalization: A Comprehensive Study
- ArmGS: Composite Gaussian Appearance Refinement for Modeling Dynamic Urban Environments
- Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
- PhotIQA: A photoacoustic image data set with image quality ratings
- Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
- A Fully Convolutional Approach to Denoising 2D Correlation Spectra
- Learning few-step posterior samplers by unfolding and distillation of diffusion models
- Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning
- VeFIA: An Efficient Inference Auditing Framework for Vertical Federated Collaborative Software
- DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
- MAC-Lookup: Multi-Axis Conditional Lookup Model for Underwater Image Enhancement
- SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
- Sparse-view irradiation processing volumetric additive manufacturing
- LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
- I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation
- ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
- AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
- LLM-based Realistic Safety-Critical Driving Video Generation
- A Multi-Centric Anthropomorphic 3D CT Phantom-Based Benchmark Dataset for Harmonization
- Rapid Salient Object Detection with Difference Convolutional Neural Networks
- Enabling Robust, Real-Time Verification of Vision-Based Navigation through View Synthesis
- EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image Enhancement
- Privacy-Preserving Quantized Federated Learning with Diverse Precision
- UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather Conditions
- LOD-GS: Level-of-Detail-Sensitive 3D Gaussian Splatting for Detail Conserved Anti-Aliasing
- FreNBRDF: A Frequency-Rectified Neural Material Representation
- Parameter-aware high-fidelity microstructure generation using stable diffusion
- Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
- Just Noticeable Difference for Large Multimodal Models
- Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding
- A Systematic Investigation on Deep Learning-Based Omnidirectional Image and Video Super-Resolution
- Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets
- Controllable Reference Guided Diffusion with Local Global Fusion for Real World Remote Sensing Image Super Resolution
- Image Demoiréing Using Dual Camera Fusion on Mobile Phones
- PointSSIM: A novel low dimensional resolution invariant image-to-image comparison metric
- AttentionGS: Towards Initialization-Free 3D Gaussian Splatting via Structural Attention
- OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous Driving
- GViT: Representing Images as Gaussians for Visual Recognition
- On the convergence of iterative regularization method assisted by the graph Laplacian with early stopping
- On the Resilience of Underwater Semantic Wireless Communications
- DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
- PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
- STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic Scene
- CoreMark: Toward Robust and Universal Text Watermarking Technique
- From Coarse to Fine: Learnable Discrete Wavelet Transforms for Efficient 3D Gaussian Splatting
- AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
- Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
- GamerAstra: Supporting 2D Non-Twitch Video Games for Blind and Low-Vision Players through a Multi-Agent Framework
- RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors
- ICME 2025 Generalizable HDR and SDR Video Quality Measurement Grand Challenge
- VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
- Degradation-Modeled Multipath Diffusion for Tunable Metalens Photography
- UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments
- GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
- HazeMatching: Dehazing Light Microscopy Images with Guided Conditional Flow Matching
- Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian Field
- Noise-Inspired Diffusion Model for Generalizable Low-Dose CT Reconstruction
- Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
- ZeroReg3D: A Zero-shot Registration Pipeline for 3D Consecutive Histopathology Image Reconstruction
- OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
- MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation
- Elucidating and Endowing the Diffusion Training Paradigm for General Image Restoration
- Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent Correlations
- Mapping intratumoral heterogeneity through PET-derived washout and deep learning after proton therapy
- Learning to See in the Extremely Dark
- Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
- DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
- Geometry and Perception Guided Gaussians for Multiview-consistent 3D Generation from a Single Image
- Electromagnetic Inverse Scattering from a Single Transmitter
- CitrusGAN: sparse-view X-ray CT reconstruction for citrus based on generative adversarial networks
- Fast ground penetrating radar dual-parameter full waveform inversion method accelerated by hybrid compilation of CUDA kernel function and PyTorch
- Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
- Progressive Alignment Degradation Learning for Pansharpening
- SkinningGS: Editable Dynamic Human Scene Reconstruction Using Gaussian Splatting Based on a Skinning Model
- Photon Absorption Remote Sensing (PARS): Comprehensive Absorption Imaging Enabling Label-Free Biomolecule Characterization and Mapping
- RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
- Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
- NeRF-based CBCT Reconstruction needs Normalization and Initialization
- Video Compression for Spatiotemporal Earth System Data
- Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications
- Angio-Diff: Learning a Self-Supervised Adversarial Diffusion Model for Angiographic Geometry Generation
- Parametric Gaussian Human Model: Generalizable Prior for Efficient and Realistic Human Avatar Modeling
- A Comparative Study of NAFNet Baselines for Image Restoration
- Diffusion-aided Task-oriented Semantic Communications with Model Inversion Attack
- Style Transfer: A Decade Survey
- UniTac-NV: A Unified Tactile Representation For Non-Vision-Based Tactile Sensors
- ICP-3DGS: SfM-free 3D Gaussian Splatting for Large-scale Unbounded Scenes
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- Light of Normals: Unified Feature Representation for Universal Photometric Stereo
- TAMMs: Temporal-Aware Multimodal Model for Satellite Image Change Understanding and Forecasting
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs
- 3D Arena: An Open Platform for Generative 3D Evaluation
- R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
- VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- Transforming H&E images into IHC: A Variance-Penalized GAN for Precision Oncology
- BSMamba: Brightness and Semantic Modeling for Long-Range Interaction in Low-Light Image Enhancement
- A Multi-Scale Spatial Attention-Based Zero-Shot Learning Framework for Low-Light Image Enhancement
- Limitations of NERF with pre-trained Vision Features for Few-Shot 3D Reconstruction
- NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
- BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP
- Adapting Vision-Language Models for Evaluating World Models
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
- StainPIDR: A Pathological Image Decouplingand Reconstruction Method for Stain Normalization Based on Color Vector Quantization and Structure Restaining
- StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets
- Robust Foreground-Background Separation for Severely-Degraded Videos Using Convolutional Sparse Representation Modeling
- 3D Gaussian Splatting for Fine-Detailed Surface Reconstruction in Large-Scale Scene
- MTSIC: Multi-stage Transformer-based GAN for Spectral Infrared Image Colorization
- Reversing Flow for Image Restoration
- Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
- TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion
- RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought
- MetaQAP - A Meta-Learning Approach for Quality-Aware Pretraining in Image Quality Assessment
- Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
- Generative Learning of Differentiable Object Models for Compositional Interpretation of Complex Scenes
- FFINO: Factorized Fourier Improved Neural Operator for Modeling Multiphase Flow in Underground Hydrogen Storage
- MoiréXNet: Adaptive Multi-Scale Demoiréing with Linear Attention Test-Time Training and Truncated Flow Matching Prior
- Active MRI Acquisition with Diffusion Guided Bayesian Experimental Design
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
- NTIRE 2025 Image Shadow Removal Challenge Report
- One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning
- Aligning Text, Images, and 3D Structure Token-by-Token
- Fiber Signal Denoising Algorithm using Hybrid Deep Learning Networks
- You Only Render Once: Enhancing Energy and Computation Efficiency of Mobile Virtual Reality
- A Digital Twin Framework for Adaptive Treatment Planning in Radiotherapy
- ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views
- Causally Steered Diffusion for Automated Video Counterfactual Generation
- GrFormer: A Novel Transformer on Grassmann Manifold for Infrared and Visible Image Fusion
- DepthSeg: Depth prompting in remote sensing semantic segmentation
- Breaking the Multi-Enhancement Bottleneck: Domain-Consistent Quality Enhancement for Compressed Images
- Frequency-Calibrated Membership Inference Attacks on Medical Image Diffusion Models
- 3DGS-IEval-15K: A Large-scale Image Quality Evaluation Database for 3D Gaussian-Splatting
- Exploring Diffusion with Test-Time Training on Efficient Image Restoration
- Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT
- ESRPCB: an Edge guided Super-Resolution model and Ensemble learning for tiny Printed Circuit Board Defect detection
- SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
- TextureSplat: Per-Primitive Texture Mapping for Reflective Gaussian Splatting
- Micro-macro Gaussian Splatting with Enhanced Scalability for Unconstrained Scene Reconstruction
- A Dual-Layer Image Encryption Framework Using Chaotic AES with Dynamic S-Boxes and Steganographic QR Codes
- Efficient multi-view training for 3D Gaussian Splatting
- Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration
- Task-driven real-world super-resolution of document scans
- DM3Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale Matching
- Learning Unpaired Image Dehazing with Physics-based Rehazy Generation
- Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
- Perceptual-GS: Scene-adaptive Perceptual Densification for Gaussian Splatting
- Restoring Gaussian Blurred Face Images for Deanonymization Attacks
- Optimal Transport Driven Asymmetric Image-to-Image Translation for Nuclei Segmentation of Histological Images
- InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning
- Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
- Efficient Star Distillation Attention Network for Lightweight Image Super-Resolution
- Benchmarking Image Similarity Metrics for Novel View Synthesis Applications
- Diffusion-Based Electrocardiography Noise Quantification via Anomaly Detection
- Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
- CGVQM+D: Computer Graphics Video Quality Metric and Dataset
- FCA2: Frame Compression-Aware Autoencoder for Modular and Fast Compressed Video Super-Resolution
- Taming Stable Diffusion for Computed Tomography Blind Super-Resolution
- Multi-Step Guided Diffusion for Image Restoration on Edge Devices: Toward Lightweight Perception in Embodied AI
- SPLATART: Articulated Gaussian Splatting with Estimated Object Structure
- ICME 2025 Grand Challenge on Video Super-Resolution for Video Conferencing
- Voxel-Level Brain States Prediction Using Swin Transformer
- Auditing Data Provenance in Real-world Text-to-Image Diffusion Models for Privacy and Copyright Protection
- SiliCoN: Simultaneous Nuclei Segmentation and Color Normalization of Histological Images
- Gaussian Mapping for Evolving Scenes
- SlotPi: Physics-informed Object-centric Reasoning Models
- NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery
- Edit360: 2D Image Edits to 3D Assets from Any Angle
- PointGS: Point Attention-Aware Sparse View Synthesis with Gaussian Splatting
- DUN-SRE: Deep Unrolling Network with Spatiotemporal Rotation Equivariance for Dynamic MRI Reconstruction
- RealKeyMorph: Keypoints in Real-world Coordinates for Resolution-agnostic Image Registration
- Rethinking Brain Tumor Segmentation from the Frequency Domain Perspective
- Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervals
- FA-INR: Adaptive Implicit Neural Representations for Interpretable Exploration of Simulation Ensembles
- Splat and Replace: 3D Reconstruction with Repetitive Elements
- ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On
- Reliable Evaluation of MRI Motion Correction: Dataset and Insights
- SurGSplat: Progressive Geometry-Constrained Gaussian Splatting for Surgical Scene Reconstruction
- BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading
- Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
- Bidirectional Image-Event Guided Fusion Framework for Low-Light Image Enhancement
- HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
- GS4: Generalizable Sparse Splatting Semantic SLAM
- SSH-Net: A Self-Supervised and Hybrid Network for Noisy Image Watermark Removal
- MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models
- Implicit Neural Representation-Based MRI Reconstruction Method with Sensitivity Map Constraints
- Latent Diffusion Model Based Denoising Receiver for 6G Semantic Communication: From Stochastic Differential Theory to Application
- Controlled Data Rebalancing in Multi-Task Learning for Real-World Image Super-Resolution
- Toward Better SSIM Loss for Unsupervised Monocular Depth Estimation
- FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
- Defurnishing with X-Ray Vision: Joint Removal of Furniture from Panoramas and Mesh
- FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
- Robustness as Architecture: Designing IQA Models to Withstand Adversarial Perturbations
- Enhancing Frequency for Single Image Super-Resolution with Learnable Separable Kernels
- Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
- MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations
- UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
- Splatting Physical Scenes: End-to-End Real-to-Sim from Imperfect Robot Data
- Poisson Informed Retinex Network for Extreme Low-Light Image Enhancement
- GSsplat: Generalizable Semantic Gaussian Splatting for Novel-view Synthesis in 3D Scenes
- WDMamba: When Wavelet Degradation Prior Meets Vision Mamba for Image Dehazing
- RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation
- Active Sampling for MRI-based Sequential Decision Making
- Tensor robust principal component analysis via the tensor nuclear over Frobenius norm
- SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization
- SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting
- Robust Neural Rendering in the Wild with Asymmetric Dual 3D Gaussian Splatting
- Gradient Inversion Attacks on Parameter-Efficient Fine-Tuning
- Point Cloud Quality Assessment Using the Perceptual Clustering Weighted Graph (PCW-Graph) and Attention Fusion Network
- Res-MoCoDiff: Residual-guided diffusion models for motion artifact correction in brain MRI
- HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
- Solving Inverse Problems with FLAIR
- ControlMambaIR: Conditional Controls with State-Space Model for Image Restoration
- ORV: 4D Occupancy-centric Robot Video Generation
- SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
- PhysGaia: A Physics-Aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis
- VTGaussian-SLAM: RGBD SLAM for Large Scale Scenes with Splatting View-Tied 3D Gaussians
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
- A Tree-guided CNN for image super-resolution
- LinkTo-Anime: A 2D Animation Optical Flow Dataset from 3D Model Rendering
- Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
- DRAUN: An Algorithm-Agnostic Data Reconstruction Attack on Federated Unlearning Systems
- Playing with Transformer at 30+ FPS via Next-Frame Diffusion
- Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
- NTIRE 2025 the 2nd Restore Any Image Model (RAIM) in the Wild Challenge
- DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
- Are Pixel-Wise Metrics Reliable for Sparse-View Computed Tomography Reconstruction?
- MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing
- AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
- Globally Consistent RGB-D SLAM with 2D Gaussian Splatting
- Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
- MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
- Video Signature: Implicit Watermarking for Video Diffusion Models
- ViVo: A Dataset for Volumetric Video Reconstruction and Compression
- Real-Time Person Image Synthesis Using a Flow Matching Model
- Revolutionizing Brain Tumor Imaging: Generating Synthetic 3D FA Maps from T1-Weighted MRI using CycleGAN Models
- SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models
- CXR-AD: Component X-ray Image Dataset for Industrial Anomaly Detection
- Beyond Pretty Pictures: Combined Single- and Multi-Image Super-resolution for Sentinel-2 Images
- DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
- Co-designed pre-capture privacy optics for computer vision
- 3D Gaussian Splatting Data Compression with Mixture of Priors
- Digital twins enable full-reference quality assessment of photoacoustic image reconstructions
- IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
- A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
- Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
- VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
- Cora: Correspondence-aware image editing using few step diffusion
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
- LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
- VITON-DRR: Details Retention Virtual Try-on via Non-rigid Registration
- PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- Quality assessment of 3D human animation: Subjective and objective evaluation
- Holistic Large-Scale Scene Reconstruction via Mixed Gaussian Splatting
- Advancing Image Super-resolution Techniques in Remote Sensing: A Comprehensive Survey
- FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
- MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
- Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
- SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model
- HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions
- Super-temporal-resolution Photoacoustic Imaging with Dynamic Reconstruction through Implicit Neural Representation in Sparse-view
- DeepTopoNet: A Framework for Subglacial Topography Estimation on the Greenland Ice Sheets
- Semantics-Guided Generative Image Compression
- ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
- On a novel probabilistic Sampling Kantorovich operators and their application
- UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes
- Anomalies by Synthesis: Anomaly Detection using Generative Diffusion Models for Off-Road Navigation
- PALADIN : Robust Neural Fingerprinting for Text-to-Image Diffusion Models
- Understanding Adversarial Training with Energy-based Models
- GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
- UP-SLAM: Adaptively Structured Gaussian SLAM with Uncertainty Prediction in Dynamic Environments
- FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
- LatentMove: Towards Complex Human Movement Video Generation
- GL-PGENet: A Parameterized Generation Framework for Robust Document Image Enhancement
- VRAG: Learning World Models for Interactive Video Generation
- STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
- Learning Hierarchical Sparse Transform Coding for 3DGS Compression
- How Do Diffusion Models Improve Adversarial Robustness?
- Deep Learning Empowered Sub-Diffraction Terahertz Backpropagation Single-Pixel Imaging
- Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
- VideoMarkBench: Benchmarking Robustness of Video Watermarking
- Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
- Sampling Kantorovich operators for speckle noise reduction using a Down-Up scaling approach and gap filling in remote sensing images
- Instance Data Condensation for Image Super-Resolution
- Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
- Wideband RF Radiance Field Modeling Using Frequency-embedded 3D Gaussian Splatting
- Entropy-Guided Sampling of Flat Modes in Discrete Spaces
- InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
- Video Quality Monitoring for Remote Autonomous Vehicle Control
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
- Structure from Collision
- Plenodium: UnderWater 3D Scene Reconstruction with Plenoptic Medium Representation
- DiMoSR: Feature Modulation via Multi-Branch Dilated Convolutions for Efficient Image Super-Resolution
- MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation
- Dynamic Vision from EEG Brain Recordings, How much does EEG know?
- Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
- Music Source Restoration
- HaloGS: Loose Coupling of Compact Geometry and Gaussian Splats for 3D Scenes
- Long-Context State-Space Video World Models
- Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstruction
- Learning 3D Persistent Embodied World Models
- Structure Disruption: Subverting Malicious Diffusion-Based Inpainting via Self-Attention Query Perturbation
- VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
- Underwater Diffusion Attention Network with Contrastive Language-Image Joint Learning for Underwater Image Enhancement
- HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment
- DeepInverse: A Python package for solving imaging inverse problems with deep learning
- OB3D: A New Dataset for Benchmarking Omnidirectional 3D Reconstruction Using Blender
- GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- Improving Novel view synthesis of 360^∘ Scenes in Extremely Sparse Views by Jointly Training Hemisphere Sampled Synthetic Images
- Deep learning of personalized priors from past MRI scans enables fast, quality-enhanced point-of-care MRI with low-cost systems
- Sparfels: Fast Reconstruction from Sparse Unposed Imagery
- Freqformer: Image-Demoiréing Transformer via Efficient Frequency Decomposition
- MMP-2K: A Benchmark Multi-Labeled Macro Photography Image Quality Assessment Database
- Querying Kernel Methods Suffices for Reconstructing their Training Data
- Geometry-guided Online 3D Video Synthesis with Multi-View Temporal Consistency
- MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation
- Veta-GS: View-dependent deformable 3D Gaussian Splatting for thermal infrared Novel-view Synthesis
- Eye-See-You: Reverse Pass-Through VR and Head Avatars
- CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting
- Beyond Illumination: A Conditional Mutual Information-Guided Network for Low-Light Image Enhancement
- F-ANcGAN: An Attention-Enhanced Cycle Consistent Generative Adversarial Architecture for Synthetic Image Generation of Nanoparticles
- Towards more transferable adversarial attack in black-box manner
- DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
- Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution
- A Unified and Fast-Sampling Diffusion Bridge Framework via Stochastic Optimal Control
- SplatCo: Structure-View Collaborative Gaussian Splatting for Detail-Preserving Rendering of Large-Scale Unbounded Scenes
- DiffusionReward: Enhancing Blind Face Restoration through Reward Feedback Learning
- UniqueSplat: View-Conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction
- Ownership Verification of DNN Models Using White-Box Adversarial Attacks with Specified Probability Manipulation
- Deblending Overlapping Galaxies in DECaLS Using Transformer-Based Algorithm: A Method Combining Multiple Bands and Data Types
- Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- High-Fidelity Functional Ultrasound Reconstruction via A Visual Auto-Regressive Framework
- Render-FM: A Foundation Model for Real-time Photorealistic Volumetric Rendering
- Loop-Mamba: A Loop Mamba with Degradation-Aware and Shared Memory for Old Photo Restoration
- Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
- Seeing through Satellite Images at Street Views
- T2exture: Sparsely Perturbed Thermal-to-Texture Imaging
- Forward-only Diffusion Probabilistic Models
- On the use of Graphs for Satellite Image Time Series
- Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts
- Clear Nights Ahead: Towards Multi-Weather Nighttime Image Restoration
- Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
- SuperPure: Efficient Purification of Localized and Distributed Adversarial Patches via Super-Resolution GAN Models
- DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
- Deep Learning-Driven Ultra-High-Definition Image Restoration: A Survey
- Motion Matters: Compact Gaussian Streaming for Free-Viewpoint Video Reconstruction
- Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
- TRAIL: Transferable Robust Adversarial Images via Latent diffusion
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
- Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
- Masked Conditioning for Deep Generative Models
- Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation
- CP-LLM: Context and Pixel Aware Large Language Model for Video Quality Assessment
- EVA: Expressive Virtual Avatars from Multi-view Videos
- FRN: Fractal-Based Recursive Spectral Reconstruction Network
- R3GS: Gaussian Splatting for Robust Reconstruction and Relocalization in Unconstrained Image Collections
- GT2-GS: Geometry-aware Texture Transfer for Gaussian Splatting
- A Deep Learning Framework for Two-Dimensional, Multi-Frequency Propagation Factor Estimation
- My Face Is Mine, Not Yours: Facial Protection Against Diffusion Model Face Swapping
- MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
- X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
- Bidirectional Variational Autoencoders
- StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
- GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
- Securing Transfer-Learned Networks with Reverse Homomorphic Encryption
- UHD Image Dehazing via anDehazeFormer with Atmospheric-aware KV Cache
- RLVR-World: Training World Models with Reinforcement Learning
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- Enhancing Vision Transformer Explainability Using Artificial Astrocytes
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
- Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning
- Towards Generating Realistic Underwater Images
- Exploring Image Quality Assessment from a New Perspective: Pupil Size
- Memory-Centric Embodied Question Answering
- RETRO: REthinking Tactile Representation Learning with Material PriOrs
- IPENS:Interactive Unsupervised Framework for Rapid Plant Phenotyping Extraction via NeRF-SAM2 Fusion
- Anti-Inpainting: A Proactive Defense Approach against Malicious Diffusion-based Inpainters under Unknown Conditions
- Towards a Universal Image Degradation Model via Content-Degradation Disentanglement
- Single Image Reflection Separation via Dual Prior Interaction Transformer
- Learning Driven Elastic Task Multi-Connectivity Immersive Computing Systems
- Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
- Neural-Enhanced Rate Adaptation and Computation Distribution for Emerging mmWave Multi-User 3D Video Streaming Systems
- FreqSelect: Frequency-Aware fMRI-to-Image Reconstruction
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement
- Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
- CompBench: Benchmarking Complex Instruction-guided Image Editing
- Generative Adversarial Reconstruction with Adaptive Thresholding for Obstructed Targets in Computational Microwave Imaging
- VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
- Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling
- Black-box Adversaries from Latent Space: Unnoticeable Attacks on Human Pose and Shape Estimation
- MiniWorld: Democratizing the Training of Video World Models from Scratch
- NTIRE 2025 Challenge on Efficient Burst HDR and Restoration: Datasets, Methods, and Results
- GTR: Gaussian Splatting Tracking and Reconstruction of Unknown Objects Based on Appearance and Geometric Complexity
- Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment
- Urban Representation Learning for Fine-grained Economic Mapping: A Semi-supervised Graph-based Approach
- From Fibers to Cells: Fourier-Based Registration Enables Virtual Cresyl Violet Staining From 3D Polarized Light Imaging
- Controlling spatial correlation in k-space interpolation networks for MRI reconstruction: denoising versus apparent blurring
- CleanPatrick: A Benchmark for Image Data Cleaning
- Hybrid-Domain Posterior Sampling for Inverse Problems via Latent Flow Matching
- Textured mesh Quality Assessment using Geometry and Color Field Similarity
- X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
- QVGen: Pushing the Limit of Quantized Video Generative Models
- X-Edit: Detecting and Localizing Edits in Images Altered by Text-Guided Diffusion Models
- AeroLLE: Constrained Pseudo-Supervision for Nighttime Aerial Image Enhancement with the AeroNight-1.5K Benchmark
- Equal is Not Always Fair: A New Perspective on Hyperspectral Representation Non-Uniformity
- UGoDIT: Unsupervised Group Deep Image Prior Via Transferable Weights
- What's Inside Your Diffusion Model? A Score-Based Riemannian Metric to Explore the Data Manifold
- A Physics-Informed Spatiotemporal Deep Learning Framework for Turbulent Systems
- ControlGS: Consistent Structural Compression Control for Deployment-Aware Gaussian Splatting
- MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting
- MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs
- From Preimage Search To Source-Grounded Feature Inversion
- Learned Lightweight Smartphone ISP with Unpaired Data
- Large-Scale Gaussian Splatting SLAM
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- RainPro-8: An Efficient Deep Learning Model to Estimate Rainfall Probabilities Over 8 Hours
- Don't Forget your Inverse DDIM for Image Editing
- RPL-UIE: Reliable Prior Learning for Underwater Image Enhancement
- Few-Shot Anomaly-Driven Generation for Anomaly Classification and Segmentation
- An Exploration of Default Images in Text-to-Image Generation
- Meta-learning Slice-to-Volume Reconstruction in Fetal Brain MRI using Implicit Neural Representations
- FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration
- Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
- Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians
- Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
- SymNet: A Multi-Task Network for Joint Radio Map Reconstruction and Transmitter Localization
- Beyond Edge Maps: Wavelet-Domain Conditioning for Multi-Adapter Map-to-Satellite Diffusion
- ACT-R: Adaptive Camera Trajectories for Single View 3D Reconstruction
- ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction
- Highly Undersampled MRI Reconstruction via a Single Posterior Sampling of Diffusion Models
- WaveGuard: Robust Deepfake Detection and Source Tracing via Dual-Tree Complex Wavelet and Graph Neural Networks
- Beyond Consistency: Preserving Temporal Structure in Zero-Shot Video Editing
- PtyRAD: A High-performance and Flexible Ptychographic Reconstruction Framework with Automatic Differentiation
- Enabling Privacy-Aware AI-Based Ergonomic Analysis
- MixBridge: Heterogeneous Image-to-Image Backdoor Attack through Mixture of Schrödinger Bridges
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- TUGS: Physics-based Compact Representation of Underwater Scenes by Tensorized Gaussian
- High-Frequency Prior-Driven Adaptive Masking for Accelerating Image Super-Resolution
- Enhancing Monocular Height Estimation via Sparse LiDAR-Guided Correction
- Fast and low energy approximate full adder based on FELIX logic
- DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- NeuGen: Amplifying the 'Neural' in Neural Radiance Fields for Domain Generalization
- MultiTaskVIF: Segmentation-oriented visible and infrared image fusion via multi-task learning
- PC-SRGAN: Physically Consistent Super-Resolution Generative Adversarial Network for General Transient Simulations
- ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
- Distributionally Robust Contract Theory for Edge AIGC Services in Teleoperation
- A review of advancements in low-light image enhancement using deep learning
- MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
- Document Image Rectification Bases on Self-Adaptive Multitask Fusion
- TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries
- Towards order of magnitude X-ray dose reduction in breast cancer imaging using phase contrast and deep denoising
- A New k-Space Model for Non-Cartesian Fourier Imaging
- SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data Augmentation
- D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
- ItDPDM: Information-Theoretic Discrete Poisson Diffusion Model
- UltraGauss: Ultrafast Gaussian Reconstruction of 3D Ultrasound Volumes
- Learning Heterogeneous Mixture of Scene Experts for Large-scale Neural Radiance Fields
- SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting
- HiLLIE: Human-in-the-Loop Training for Low-Light Image Enhancement
- Hybrid Image Resolution Quality Metric (HIRQM):A Comprehensive Perceptual Image Quality Assessment Framework
- Discrete Spatial Diffusion: Intensity-Preserving Diffusion Modeling
- PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
- Priorconditioned Sparsity-Promoting Projection Methods for Deterministic and Bayesian Linear Inverse Problems
- GauS-SLAM: Dense RGB-D SLAM with Gaussian Surfels
- StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
- FalconWing: An Ultra-Light Indoor Fixed-Wing UAV Platform for Vision-Based Autonomy
- CostFilter-AD: Enhancing Anomaly Detection through Matching Cost Filtering
- High Dynamic Range Novel View Synthesis with Single Exposure
- VRS-UIE: Value-Driven Reordering Scanning for Underwater Image Enhancement
- AdvSplat: Adversarial Attacks on Feed-Forward Gaussian Splatting Models
- Exact Posterior Score Estimation for Solving Linear Inverse Problems
- Generating Synthetic Wildlife Health Data from Camera Trap Imagery: A Pipeline for Alopecia and Body Condition Training Data
- MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers
- The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
- An unsupervised GNN-transformer model for gold geochemical anomaly detection
- Tailor Made Embeddings for Quantum Machine Learning
- Unified Geometry-Guided ML-FTLE for Tracking Transient Chaos from Scalar Time Series
- InstructAttribute: Fine-grained Object Attributes editing with Instruction
- Towards Robust and Generalizable Gerchberg Saxton based Physics Inspired Neural Networks for Computer Generated Holography: A Sensitivity Analysis Framework
- Direct Motion Models for Assessing Generated Videos
- Optimized Lattice-Structured Flexible EIT Sensor for Tactile Reconstruction and Classification
- Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)
- Self-Aware Object Detection via Degradation Manifolds
- Unpaired Image-to-Image Translation via a Self-Supervised Semantic Bridge
- Benchmarking quantum simulation with neutron-scattering experiments
- GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
- Self-Supervised Monocular Visual Drone Model Identification through Improved Occlusion Handling
- Diffusion-based Adversarial Identity Manipulation for Facial Privacy Protection
- Revisiting Diffusion Autoencoder Training for Image Reconstruction Quality
- DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration
- Synergy-CLIP: Extending CLIP with Multi-modal Integration for Robust Representation Learning
- MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric Guidance
- ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
- RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting
- MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
- Unconstrained Large-scale 3D Reconstruction and Rendering across Altitudes
- EfficientHuman: Efficient Training and Reconstruction of Moving Human using Articulated 2D Gaussian
- LoViF 2026 Challenge on Human-oriented Semantic Image Quality Assessment: Methods and Results
- Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
- Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
- Perception-aware Sampling for Scatterplot Visualizations
- Legilimens: Performant Video Analytics on the System-on-Chip Edge
- TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability
- FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
- CompleteMe: Reference-based Human Image Completion
- Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular Video
- Interactive Double Deep Q-network: Integrating Human Interventions and Evaluative Predictions in Reinforcement Learning of Autonomous Driving
- Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition
- Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
- AnimateAnywhere: Rouse the Background in Human Image Animation
- Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
- FusionNet: Multi-model Linear Fusion Framework for Low-light Image Enhancement
- Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
- Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
- RadioFormer: A Multiple-Granularity Radio Map Estimation Transformer with 1\textpertenthousand Spatial Sampling
- REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
- 4DGS-CC: A Contextual Coding Framework for 4D Gaussian Splatting Data Compression
- Augmenting Perceptual Super-Resolution via Image Quality Predictors
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- Depth3DLane: Monocular 3D Lane Detection via Depth Prior Distillation
- Outlier-aware Tensor Robust Principal Component Analysis with Self-guided Data Augmentation
- Adaptive Weight Modified Riesz Mean Filter For High-Density Salt and Pepper Noise Removal
- Physics-Driven Neural Compensation For Electrical Impedance Tomography
- DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion
- HepatoGEN: Generating Hepatobiliary Phase MRI with Perceptual and Adversarial Models
- Latent-Compressed Variational Autoencoder for Video Diffusion Models
- Commodity RF Sensing of Belowground Tuber Growth
- Scene Perceived Image Perceptual Score (SPIPS): combining global and local perception for image quality assessment
- SDiT: Semantic Region-Adaptive for Diffusion Transformers
- Multi-modal MRI-Based Alzheimer's Disease Diagnosis with Transformer-based Image Synthesis and Transfer Learning
- FaithIR: Rethinking Infrared Image Super-Resolution from Perceptual Sharpness to Task Relevant Fidelity
- S3-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images
- Morphology-Aware Implicit Super-Resolution Network for Pathological Images
- XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation
- FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis
- UniNav: A Unified World-Action Diffusion Model for Visual Navigation
- NanoMorph-3D: An End-to-End Physics-Driven Unrolling Framework for Nanomaterial Reconstruction
- I-INR: Iterative Implicit Neural Representations
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
- STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
- LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
- DCT-Shield: A Robust Frequency Domain Defense against Malicious Image Editing
- Deep Reparameterization for Full Waveform Inversion: Architecture Benchmarking, Robust Inversion, and Multiphysics Extension
- ScoreField: Neural Inverse Scattering with Score-Based Generative Priors
- Self-Supervised Noise Adaptive MRI Denoising via Repetition to Repetition (Rep2Rep) Learning
- Unveiling Hidden Vulnerabilities in Digital Human Generation via Adversarial Attacks
- Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
- Casual3DHDR: Deblurring High Dynamic Range 3D Gaussian Splatting from Casually Captured Videos
- High-Quality Cloud-Free Optical Image Synthesis Using Multi-Temporal SAR and Contaminated Optical Data
- UBLLIE: Unified Backlight and Low-Light Image Enhancement
- DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
- Overcoming Statistical Bias in Action-Controllable World Models
- ACA-GS: Adaptive-Capacity Anchored Gaussian Splatting for Compact Dynamic Radiance Fields
- Variable Smoothing for Weakly Convex Problems with Non-Euclidean Directions
- An active-learning framework for real-time depth perception from monocular vision streams
- A geometry-based deep equilibrium model for image restoration under multiplicative Gamma noise
- Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
- Multi-View Face and Gesture Animation with Dynamic Gaussians
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
- OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
- DeclutterNeRF: Generative-Free 3D Scene Recovery for Occlusion Removal
- Subjective Visual Quality Assessment for High-Fidelity Learning-Based Image Compression
- A Machine Learning and Finite Element Framework for Inverse Elliptic PDEs via Dirichlet-to-Neumann Mapping
- MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
- SCAIL-2: Unifying Controlled Character Animation with End-to-End In-Context Conditioning
- Neural electron backscatter diffraction
- IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration
- Image Thresholding: Understanding Bias of Evaluation Metrics Towards Specific Evaluation Functions
- From Understanding to Erasing: Towards Complete and Stable Video Object Removal
- Investigating How the Fractions Skill Score and Brier Divergence Skill Score Reflect Forecast Error
- TactileNet: Bridging the Accessibility Gap with AI-Generated Tactile Graphics for Individuals with Vision Impairment
- Texture2LoD3: Enabling LoD3 Building Reconstruction With Panoramic Images
- ECGDeDRDNet: A deep learning-based method for Electrocardiogram noise removal using a double recurrent dense network
- Targetless LiDAR-Camera Calibration with Neural Gaussian Splatting
- SaENeRF: Suppressing Artifacts in Event-based Neural Radiance Fields
- Iterative Collaboration Network Guided By Reconstruction Prior for Medical Image Super-Resolution
- Self-Guided Diffusion Model for Accelerating Computational Fluid Dynamics
- FluentLip: A Phonemes-Based Two-stage Approach for Audio-Driven Lip Synthesis with Optical Flow Consistency
- DSDNet: Raw Domain Demoiréing via Dual Color-Space Synergy
- StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians
- DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
- MoBGS: Motion Deblurring Dynamic 3D Gaussian Splatting for Blurry Monocular Video
- Structure-guided Diffusion Transformer for Low-Light Image Enhancement
- SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation
- A Controllable Appearance Representation for Flexible Transfer and Editing
- NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study
- Protecting Your Voice: Temporal-aware Robust Watermarking
- IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays
- Frequency-domain Learning with Kernel Prior for Blind Image Deblurring
- NeRFlex: Resource-aware Real-time High-quality Rendering of Complex Scenes on Mobile Devices
- VGNC: Reducing the Overfitting of Sparse-view 3DGS via Validation-guided Gaussian Number Control
- FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
- Metamon-GS: Enhancing Representability with Variance-Guided Densification and Light Encoding
- SEGA: Drivable 3D Gaussian Head Avatar from a Single Image
- RadioDiff-Inverse: Diffusion Enhanced Bayesian Inverse Estimation for ISAC Radio Map Construction
- Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network
- Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
- Fragile Watermarking for Image Certification Using Deep Steganographic Embedding
- SLAM&Render: A Benchmark for the Intersection Between Neural Rendering, Gaussian Splatting and SLAM
- SupResDiffGAN a new approach for the Super-Resolution task
- EG-Gaussian: Epipolar Geometry and Graph Network Enhanced 3D Gaussian Splatting
- Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
- Volume Encoding Gaussians: Transfer Function-Agnostic 3D Gaussians for Volume Rendering
- MGT: Extending Virtual Try-Off to Multi-Garment Scenarios
- CompGS++: Compressed Gaussian Splatting for Static and Dynamic Scene Representation
- Second-order Optimization of Gaussian Splats with Importance Sampling
- DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
- A Survey on Cross-Modal Interaction Between Music and Multimodal Data
- NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models
- AdaQual-Diff: Diffusion-Based Image Restoration via Adaptive Quality Prompting
- High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
- Imaging for All-Day Wearable Smart Glasses
- TTRD3: Texture Transfer Residual Denoising Dual Diffusion Model for Remote Sensing Image Super-Resolution
- Digital Twin Generation from Visual Data: A Survey
- AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
- Event Quality Score (EQS): Assessing the Realism of Simulated Event Camera Streams via Distances in Latent Space
- Cobra: Efficient Line Art COlorization with BRoAder References
- Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
- Wavelet-based Variational Autoencoders for High-Resolution Image Generation
- PCDiff: Proactive Control for Ownership Protection in Diffusion Models with Watermark Compatibility
- 3R-GS: Best Practice in Optimizing Camera Poses Along with 3DGS
- NTIRE 2025 Challenge on Event-Based Image Deblurring: Methods and Results
- VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
- EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
- An Online Adaptation Method for Robust Depth Estimation and Visual Odometry in the Open World
- HyperKING: Quantum-Classical Generative Adversarial Networks for Hyperspectral Image Restoration
- SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
- Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping
- Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
- LiteKD-Net: Lightweight Knowledge-Distilled Network for Mobile Image Denoising
- Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams
- BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models
- RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
- Mapping at First Sense: A Lightweight Neural Network-Based Indoor Structures Prediction Method for Robot Autonomous Exploration
- IPV-Bench: Benchmarking Image Protection Methods under Diverse Image-to-Video Generation Scenarios
- Local-Order Auxiliary Losses Can Improve Autoencoder Reconstruction
- EDGS: Eliminating Densification for Efficient Convergence of 3DGS
- TerraMind: Large-Scale Generative Multimodality for Earth Observation
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
- PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage
- Big Brother is Watching: Proactive Deepfake Detection via Learnable Hidden Face
- WaterFlow: Learning Fast & Robust Watermarks using Stable Diffusion
- Efficient and Robust Remote Sensing Image Denoising Using Randomized Approximation of Geodesics' Gramian on the Manifold Underlying the Patch Space
- Rainy: Unlocking Satellite Calibration for Deep Learning in Precipitation
- H-MoRe: Learning Human-centric Motion Representation for Action Analysis
- PG-DPIR: An efficient plug-and-play method for high-count Poisson-Gaussian inverse problems
- Noise2Ghost: Self-supervised deep convolutional reconstruction for ghost imaging
- A Theory of Universal Rate-Distortion-Classification Representations for Lossy Compression
- EBAD-Gaussian: Event-driven Bundle Adjusted Deblur Gaussian Splatting
- Dual-grid parameter choice method with application to image deblurring
- LL-Gaussian: Low-Light Scene Reconstruction and Enhancement via Gaussian Splatting for Novel View Synthesis
- Investigating the Role of Bilateral Symmetry for Inpainting Brain MRI
- CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
- DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- MedIL: Implicit Latent Spaces for Generating Heterogeneous Medical Images at Arbitrary Resolutions
- Universal Rate-Distortion-Classification Representations for Lossy Compression
- Towards Explainable Partial-AIGC Image Quality Assessment
- Data-Importance-Aware Power Allocation for Adaptive Semantic Communication in Computer Vision Applications
- SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
- RouterKT: Mixture-of-Experts for Knowledge Tracing
- A Knowledge-guided Adversarial Defense for Resisting Malicious Visual Manipulation
- SN-LiDAR: Semantic Neural Fields for Novel Space-time View LiDAR Synthesis
- Geometric Consistency Refinement for Single Image Novel View Synthesis via Test-Time Adaptation of Diffusion Models
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy
- Conformalized Generative Bayesian Imaging: An Uncertainty Quantification Framework for Computational Imaging
- RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
- Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
- Evaluating the robustness of explainable AI in medical image recognition under natural and adversarial data corruption
- Image registration of 2D optical thin sections in a 3D porous medium: Application to a Berea sandstone digital rock image
- SVG-IR: Spatially-Varying Gaussian Splatting for Inverse Rendering
- Distilling Textual Priors from LLM to Efficient Image Fusion
- AstroClearNet: Deep image prior for multi-frame astronomical image restoration
- GIGA: Generalizable Sparse Image-driven Gaussian Humans
- CamC2V: Context-aware Controllable Video Generation
- Micro-splatting: Multistage Isotropy-informed Covariance Regularization Optimization for High-Fidelity 3D Gaussian Splatting
- Time-Aware Auto White Balance in Mobile Photography
- Meta-Continual Learning of Neural Fields
- HiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion Representation
- AI-Driven Reconstruction of Large-Scale Structure from Combined Photometric and Spectroscopic Surveys
- Flash Sculptor: Modular 3D Worlds from Objects
- Towards Efficient Real-Time Video Motion Transfer via Generative Time Series Modeling
- PartStickers: Generating Parts of Objects for Rapid Prototyping
- DA2Diff: Exploring Degradation-aware Adaptive Diffusion Priors for All-in-One Weather Restoration
- TropoDeep: a deep learning-based model for InSAR tropospheric correction on large-scale interferograms using GNSS and WRF outputs
- Feature Importance-Aware Deep Joint Source-Channel Coding for Computationally Efficient and Adjustable Image Transmission
- Exploring Kernel Transformations for Implicit Neural Representations