SGDR: Stochastic Gradient Descent with Warm Restarts
2016/08/13 by Ilya Loshchilov, Frank Hutter, Loshchilov, Ilya +1 · 530 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques
paper · pdf · doi:10.48550/arxiv.1608.03983
openalex publication_date 2016/08/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respectively. We also demonstrate its advantages on a dataset of EEG recordings and on a downsampled version of the ImageNet dataset. Our source code is available at https://github.com/loshchil/SGDR
Citations
Cited by
- Principled Algorithms for Optimizing Generalized Metrics in Binary Classification
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification
- PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation
- AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
- SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion
- Scale Weight Decay and Train Better
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization
- Normalizing Flows to Reconstruct Pseudo-PDFs
- Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors
- Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning
- Fourier Feature Physics-Informed Neural Networks for Elasto-Plastic Analysis of Geomaterials with a Non-Associative Mohr-Coulomb Model
- Learning Association via Track-Detection Matching for Multi-Object Tracking
- Hierarchical Grading in Large Language Models
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time
- A Projected Stochastic Gradient Method for Finite-Sum Problems with Linear Equality Constraints
- Concept-based Visual Counterfactual Explanations with Diffusion Models
- Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex
- UltraLBM-UNet: Ultralight Bidirectional Mamba-based Model for Skin Lesion Segmentation
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- MultiMind at SemEval-2025 Task 7: Crosslingual Fact-Checked Claim Retrieval via Multi-Source Alignment
- Defending against adversarial attacks using mixture of experts
- Optimizer Dynamics at the Edge of Stability with Differential Privacy
- Randomized time stepping of nonlinearly parametrized solutions of evolution problems
- Adaptive Probability Flow Residual Minimization for High-Dimensional Fokker-Planck Equations
- Automated Mosaic Tesserae Segmentation via Deep Learning Techniques
- Robust and scalable simulation-based inference for gravitational wave signals with gaps
- Learning-Based Estimation of Spatially Resolved Scatter Radiation Fields in Interventional Radiology
- Microstructure-based Variational Neural Networks for Robust Uncertainty Quantification in Materials Digital Twins
- InfinityEBSD : Metrics-Guided Infinite-Size EBSD Map Generation With Diffusion Models
- Sharing Knowledge without Sharing Data: Stitches can improve ensembles of disjointly trained models
- Multi-scale Attention-Guided Intrinsic Decomposition and Rendering Pass Prediction for Facial Images
- LaverNet: Lightweight All-in-one Video Restoration via Selective Propagation
- Can Transformers overcome the lack of data in the simulation of history-dependent flows?
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- Hierarchical Neural Surfaces for 3D Mesh Compression
- Residual GRU+MHSA: A Lightweight Hybrid Recurrent Attention Model for Cardiovascular Disease Detection
- Quanvolutional Neural Networks for Spectrum Peak-Finding
- Informing Acquisition Functions via Foundation Models for Molecular Discovery
- Uncertainty Quantification for Machine Learning: One Size Does Not Fit All
- MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater
- Super-Resolved Canopy Height Mapping from Sentinel-2 Time Series Using LiDAR HD Reference Data across Metropolitan France
- A Variable Step Sizes Frequency Offsets-Compensated Least Mean Squares Algorithm
- Decision Feedback-Aided Known-Interference Cancellation
- Extrapolation of Periodic Functions Using Binary Encoding of Continuous Numerical Values
- Sharp Monocular View Synthesis in Less Than a Second
- FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose Estimation
- Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation
- FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- SAQ: Stabilizer-Aware Quantum Error Correction Decoder
- Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
- Leveraging Multispectral Sensors for Color Correction in Mobile Cameras
- Forecasting Dark Matter Subhalo Constraints from Stellar Streams using Implicit Likelihood Inference
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- Learning Relative Gene Expression Trends from Pathology Images in Spatial Transcriptomics
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- Curvature-Regularized Variational Autoencoder for 3D Scene Reconstruction from Sparse Depth
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- Beyond Adam: Disentangling Optimizer Effects in the Fine-Tuning of Atomistic Foundation Models
- Gradient Descent with Provably Tuned Learning-rate Schedules
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
- Parameter efficient hybrid spiking-quantum convolutional neural network with surrogate gradient and quantum data-reupload
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Multi-Scale Visual Prompting for Lightweight Small-Image Classification
- AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
- Network of Theseus (like the ship)
- PGP-DiffSR: Phase-Guided Progressive Pruning for Efficient Diffusion-based Image Super-Resolution
- Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- TPCNet: Triple physical constraints for Low-light Image Enhancement
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- MDiff4STR: Mask Diffusion Model for Scene Text Recognition
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- DAISI: Data Assimilation with Inverse Sampling using Stochastic Interpolants
- Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting
- Adaptive Dataset Quantization: A New Direction for Dataset Pruning
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
- Adam Simplified: Bias Correction Debunked
- Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
- Exploring Surround-View Fisheye Camera 3D Object Detection
- TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception
- Subtract the Corruption: Training-Data-Free Corrective Machine Unlearning using Task Arithmetic
- Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
- Off-Road Navigation via Implicit Neural Representation of Terrain Traversability
- UnfoldLDM: Deep Unfolding-based Blind Image Restoration with Latent Diffusion Priors
- Active Learning with Selective Time-Step Acquisition for PDEs
- Learning Rate Scheduling with Matrix Factorization for Private Training
- Selective Rotary Position Embedding
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- Contrastive vision-language learning with paraphrasing and negation
- From Low-Rank Features to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
- Driving in Spikes: An Entropy-Guided Object Detector for Spike Cameras
- SNAP: Low-Latency Test-Time Adaptation with Sparse Updates
- Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
- CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
- Context Cascade Compression: Exploring the Upper Limits of Text Compression
- nnMIL: A generalizable multiple instance learning framework for computational pathology
- On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- Adaptive Multi-Scale Integration Unlocks Robust Cell Annotation in Histopathology Images
- View-aware Cross-modal Distillation for Multi-view Action Recognition
- Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement Learning
- Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining
- Decoupling Scene Perception and Ego Status: A Multi-Context Fusion Approach for Enhanced Generalization in End-to-End Autonomous Driving
- Monocular 3D Lane Detection via Structure Uncertainty-Aware Network with Curve-Point Queries
- MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation
- Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash Function
- Physically Interpretable Multi-Degradation Image Restoration via Deep Unfolding and Explainable Convolution
- ChemFixer: Correcting Invalid Molecules to Unlock Previously Unseen Chemical Space
- FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
- DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance Segmentation
- STORM: Segment, Track, and Object Re-Localization from a Single Image
- Iterated Population Based Training with Task-Agnostic Restarts
- Ultrafast Pulse Retrieval from Partial FROG Traces Using Implicit Diffusion Models
- One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms
- Mitigating Negative Flips via Margin Preserving Training
- Towards Provably Unlearnable Examples via Bayes Error Optimisation
- Schedulers for Schedule-free: Theoretically inspired hyperparameters
- Enhancing Binary Encoded Crime Linkage Analysis Using Siamese Network
- Recurrent Equivariant Constraint Modulation: Learning Per-Layer Symmetry Relaxation from Data
- Hierarchical Spatial-Frequency Aggregation for Spectral Deconvolution Imaging
- Does TabPFN Understand Causal Structures?
- Countering Multi-modal Representation Collapse through Rank-targeted Fusion
- Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning
- Learning to Restore Multi-Degraded Images via Ingredient Decoupling and Task-Aware Path Adaptation
- DARN: Dynamic Adaptive Regularization Networks for Efficient and Robust Foundation Model Adaptation
- Probabilistic Textual Time Series Depression Detection
- Active Domain Adaptation for mmWave-based HAR via Renyi Entropy-based Uncertainty Estimation
- The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
- Accelerated Sequential Posterior Inference via Reuse for Gravitational-Wave Analyses
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- Disentangling Internal Tides from Balanced Motions with Deep Learning and Surface Field Synergy
- Data-Efficient Realized Volatility Forecasting with Vision Transformers
- Machine Learning the Conformal Manifold of Holographic CFT2s
- Anatomically Constrained Transformers for Echocardiogram Analysis
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
- GeneFlow: Translation of Single-cell Gene Expression to Histopathological Images via Rectified Flow
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- Transcending Sparse Measurement Limits: Operator-Learning-Driven Data Super-Resolution for Inverse Source Problem
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- BasicAVSR: Arbitrary-Scale Video Super-Resolution via Image Priors and Enhanced Motion Compensation
- TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
- Compactly supported radial basis functions as probability density functions
- Paris: A Decentralized Trained Open-Weight Diffusion Model
- Unlimited OCR Works
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- ProbFM: Probabilistic Time Series Foundation Model with Uncertainty Decomposition
- Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- KAN-GCN: Combining Kolmogorov-Arnold Network with Graph Convolution Network for an Accurate Ice Sheet Emulator
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Eigenfunction Extraction for Ordered Representation Learning
- DeshadowMamba: Deshadowing as 1D Sequential Similarity
- Low-Dose CT Imaging Using a Regularization-Enhanced Efficient Diffusion Probabilistic Model
- Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Learning Event-guided Exposure-agnostic Video Frame Interpolation via Adaptive Feature Blending
- Streaming Generation for Music Accompaniment
- EBOP MAVEN: A machine learning model to estimate the input parameters for analytic fitting of detached eclipsing binary light curves
- Weak-to-Strong Generalization under Distribution Shifts
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- 3rd Place Solution to ICCV LargeFineFoodAI Retrieval
- SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- Convergence Analysis of SGD under Expected Smoothness
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- What Does It Take to Build a Performant Selective Classifier?
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
- Mitigating representation bias caused by missing pixels in methane plume detection
- AegisRF: Adversarial Perturbations Guided with Sensitivity for Protecting Intellectual Property of Neural Radiance Fields
- Towards In-Situ Failure Assessment: Deep Learning on DIC Results for Laminated Composites
- ShortcutBreaker: Low-Rank Noisy Bottleneck and Frequency Filtering Block for Multi-Class Unsupervised Anomaly Detection
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- Elastic ViTs from Pretrained Models without Retraining
- AI-Boosted Video Annotation: Assessing the Process Enhancement
- Discovering How Ice Crystals Grow Using Neural ODE's and Symbolic Regression
- DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification under Domain Shift
- DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
- Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Decorrelation Speeds Up Vision Transformers
- EcoScaleNet: A Lightweight Multi Kernel Network for Long Sequence 12 lead ECG Classification
- Ultra Fast Structure-aware Deep Lane Detection
- Scaling Vision Transformers for Functional MRI with Flat Maps
- NTIRE 2025 Challenge on Low Light Image Enhancement: Methods and Results
- Comparative Analysis of Data Augmentation for Clinical ECG Classification with STAR
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
- Chinese ModernBERT with Whole-Word Masking
- Efficient Real-World Deblurring using Single Images: AIM 2025 Challenge Report
- What If : Understanding Motion Through Sparse Interactions
- Self-semi-supervised Learning to Learn from NoisyLabeled Data
- On Feature Decorrelation in Self-Supervised Learning
- Variational Disentanglement for Domain Generalization
- Comparing Symmetrized Determinant Neural Quantum States for the Hubbard Model
- Generalisation of automatic tumour segmentation in histopathological whole-slide images across multiple cancer types
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- Towards fairer public transit: Real-time tensor-based multimodal fare evasion and fraud detection
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- AutoSpeech: Neural Architecture Search for Speaker Recognition
- End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
- Inferring Optical Tissue Properties from Photoplethysmography using Hybrid Amortized Inference
- Stochastic Training is Not Necessary for Generalization
- Label-Aware Distribution Calibration for Long-tailed Classification
- Analysis of Atomistic Representations Using Weighted Skip-Connections
- Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
- FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
- PHyCLIP: ℓ1-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
- Contrastive Representations for Label Noise Require Fine-Tuning
- SatFusion: A Unified Framework for Enhancing Satellite IoT Images via Multi-Temporal and Multi-Source Data Fusion
- TPLinker: Single-stage Joint Extraction of Entities and Relations Through Token Pair Linking
- Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
- Tasks, stability, architecture, and compute: Training more effective\n learned optimizers, and using them to train themselves
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model Training
- Collective is different: Information exchange and speed-accuracy trade-offs in self-organized patterning
- Mid-Training of Large Language Models: A Survey
- DeRainMamba: A Frequency-Aware State Space Model with Detail Enhancement for Image Deraining
- Training Dynamics Impact Post-Training Quantization Robustness
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- QuantDemoire: Quantization with Outlier Aware for Image Demoiréing
- Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised Learning
- Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion
- Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training
- Adaptively Sampling-Reusing-Mixing Decomposed Gradients to Speed Up Sharpness Aware Minimization
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- Not All Memories are Created Equal: Learning to Forget by Expiring
- Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
- Physics-Informed Neural Controlled Differential Equations for Scalable Long Horizon Multi-Agent Motion Forecasting
- OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
- Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty
- Continuous Space-Time Video Super-Resolution with 3D Fourier Fields
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- Intelligent Prediction and Optimization of Open-Hole Wellbore Multiphysics Stability: A Synergistic PINN-DRL Approach
- A Physics-Guided Probabilistic Surrogate Modeling Framework for Digital Twins of Underwater Radiated Noise
- MoReFlow: Motion Retargeting Learning through Unsupervised Flow Matching
- GenVarFormer: Predicting gene expression from long-range mutations in cancer
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Hybrid Layer-Wise ANN-SNN With Surrogate Spike Encoding-Decoding Structure
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots
- Computation Reallocation for Object Detection
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
- Dynamics of Learning: Generative Schedules from Latent ODEs
- Neural Architecture Search in Embedding Space
- Completing <scp>3D</scp> point clouds of individual trees using deep learning
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
- Sharpness-Aware Minimization Can Hallucinate Minimizers
- Reparameterizing 4DVAR with neural fields
- Compute-Optimal Quantization-Aware Training
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
- NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
- Does the Manipulation Process Matter? RITA: Reasoning Composite Image Manipulations via Reversely-Ordered Incremental-Transition Autoregression
- Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
- A Unified Noise-Curvature View of Loss of Trainability
- Anatomically Constrained Transformers for Cardiac Amyloidosis Classification
- Effective Data Fusion with Generalized Vegetation Index: Evidence from Land Cover Segmentation in Agriculture
- Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
- HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
- CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
- Data-free neural PDE solvers based on Graph Neural Networks and weak forms
- DeepViT: Towards Deeper Vision Transformer
- SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
- 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
- Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
- Explaining Data Mixing Scaling Laws
- Natural Image Matting via Guided Contextual Attention
- VINNAS: Variational Inference-based Neural Network Architecture Search
- Use What You Know: Causal Foundation Models with Partial Graphs
- TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning
- On the Bias Against Inductive Biases
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free
- GhostSR: Learning Ghost Features for Efficient Image Super-Resolution
- Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology
- LoCo: Local Contrastive Representation Learning
- Analysing Dropout and Compounding Errors in Neural Language Models
- CorPipe at CRAC 2025: Evaluating Multilingual Encoders for Multilingual Coreference Resolution
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Building effective deep neural network architectures one feature at a time
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- Towards Generalized Synapse Detection Across Invertebrate Species
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
- Design and Development of an Intelligent LLM-based LDAP Honeypot
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World Dehazing
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Domain Consistency Regularization for Unsupervised Multi-source Domain Adaptive Classification
- Global Pre-fixing, Local Adjusting: A Simple yet Effective Contrastive Strategy for Continual Learning
- Unsupervised Object-Level Representation Learning from Scene Images
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- MetricNet: Recovering Metric Scale in Generative Navigation Policies
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- AIM 2025 Low-light RAW Video Denoising Challenge: Dataset, Methods and Results
- FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
- Quickly Tuning Foundation Models for Image Segmentation
- Revisiting ResNets: Improved Training and Scaling Strategies
- Bridging Performance Gaps for ECG Foundation Models: A Post-Training Strategy
- DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval
- ECG-aBcDe: Overcoming Model Dependence, Encoding ECG into a Universal Language for Any LLM
- LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- Self-supervised Pretraining of Visual Features in the Wild
- Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
- SamudrACE: Fast and Accurate Coupled Climate Modeling with 3D Ocean and Atmosphere Emulators
- Stochastic restarting with multiple restart conditions
- NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
- Semantic Fusion with Fuzzy-Membership Features for Controllable Language Modelling
- Human Activity Recognition Based on Electrocardiogram Data Only
- Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
- Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
- Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- Contrastive Learning with Stronger Augmentations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
- MSDANet: A Multiscale Dual-Channel Spatial Attention Network with Depthwise Separable Convolution for Hyperspectral Image Classification
- EfficientIML: Efficient High-Resolution Image Manipulation Localization
- Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
- ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
- Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
- Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
- Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
- LIMFAST. IV. Learning High-Redshift Galaxy Formation from Multiline Intensity Mapping with Implicit Likelihood Inference
- Imitative Membership Inference Attack
- Long-Range Graph Wavelet Networks
- HiPrFlame-An ab initio based real-fluid modeling approach for high-pressure combustion-I. Rationale, methodology, and application to laminar premixed flames
- VQualA 2025 Challenge on Image Super-Resolution Generated Content Quality Assessment: Methods and Results
- Compounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network
- Lookup multivariate Kolmogorov-Arnold Networks
- AIM 2025 Challenge on High FPS Motion Deblurring: Methods and Results
- Self-Adaptive Training: beyond Empirical Risk Minimization
- A Comparison of Surrogate Constitutive Models for Viscoplastic Creep Simulation of HT-9 Steel
- CD-Mamba: Cloud detection with long-range spatial dependency modeling
- Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization
- CPEP: Contrastive Pose-EMG Pre-training Enhances Gesture Generalization on EMG Signals
- ML-PWS: Estimating the Mutual Information Between Experimental Time Series Using Neural Networks
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
- Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
- FoMEMO: Towards Foundation Models for Expensive Multi-objective Optimization
- LatPhon: Lightweight Multilingual G2P for Romance Languages and English
- Pruning Convolutional Neural Networks with Self-Supervision
- Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
- Domain Adaptation via Feature Refinement
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- SC-GIR: Goal-oriented Semantic Communication via Invariant Representation Learning
- ART: Adaptive Resampling-based Training for Imbalanced Classification
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- DarkVRAI: Capture-Condition Conditioning and Burst-Order Selective Scan for Low-light RAW Video Denoising
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- MorphGen: Morphology-Guided Representation Learning for Robust Single-Domain Generalization in Histopathological Cancer Classification
- Domain Attention Consistency for Multi-Source Domain Adaptation
- SIMILAR: Submodular Information Measures Based Active Learning In\n Realistic Scenarios
- Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs
- Representation Learning with Adaptive Superpixel Coding
- IA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Co2L: Contrastive Continual Learning
- Online incremental learning for audio classification using a pretrained audio model
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- Label Uncertainty for Ultrasound Segmentation
- Tune My Adam, Please!
- Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning
- Does simple trump complex? Comparing strategies for adversarial robustness in DNNs
- Image-Conditioned 3D Gaussian Splat Quantization
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- In2x at WMT25 Translation Task
- Busy-Quiet Video Disentangling for Video Classification
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- Neural Predictor for Neural Architecture Search
- Powering Job Search at Scale: LLM-Enhanced Query Understanding in Job Matching Systems
- Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
- Towards Practical Lipreading with Distilled and Efficient Models
- In-Context Decision Making for Optimizing Complex AutoML Pipelines
- Channel DropBlock: An Improved Regularization Method for Fine-Grained Visual Classification
- RICO: Two Realistic Benchmarks and an In-Depth Analysis for Incremental Learning in Object Detection
- DEEP-SEA: Deep-Learning Enhancement for Environmental Perception in Submerged Aquatics
- MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
- Field-level Reconstruction from Foreground-Contaminated 21-cm Maps
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Gaussian Error Linear Units (GELUs)
- Causally-Guided Pairwise Transformer -- Towards Foundational Digital Twins in Process Industry
- Superpixel-informed Continuous Low-Rank Tensor Representation for Multi-Dimensional Data Recovery
- DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
- MBMamba: When Memory Buffer Meets Mamba for Structure-Aware Image Deblurring
- MedFormer: a data-driven model for forecasting the Mediterranean Sea
- Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis
- PVT v2: Improved baselines with pyramid vision transformer
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Using Mode Connectivity for Loss Landscape Analysis
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
- Trajectory-aware Shifted State Space Models for Online Video Super-Resolution
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- Relational Action Forecasting
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud
- Explore Image Deblurring via Blur Kernel Space
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Improving Accuracy of Binary Neural Networks using Unbalanced Activation Distribution
- Arbitrary Marginal Neural Ratio Estimation for Simulation-based\n Inference
- Towards Perfection: Building Inter-component Mutual Correction for Retinex-based Low-light Image Enhancement
- Adaptive Confidence-Wise Loss for Improved Lens Structure Segmentation in AS-OCT
- Point Cloud Registration using Representative Overlapping Points
- AdaS: Adaptive Scheduling of Stochastic Gradients
- Fast and Generalizable parameter-embedded Neural Operators for Lithium-Ion Battery Simulation
- KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
- Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
- CMAMRNet: A Contextual Mask-Aware Network Enhancing Mural Restoration Through Comprehensive Mask Guidance
- D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning
- Semi-Supervised Segmentation of Salt Bodies in Seismic Images using an Ensemble of Convolutional Neural Networks
- Zero-shot self-supervised learning of single breath-hold magnetic resonance cholangiopancreatography (MRCP) reconstruction
- Towards Unified Image Deblurring using a Mixture-of-Experts Decoder
- HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Detecting Model Misspecification in Cosmology with Scale-Dependent Normalizing Flows
- Bootstrap Deep Spectral Clustering with Optimal Transport
- GRASPing Anatomy to Improve Pathology Segmentation
- PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
- Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Beyond Least Squares: Robust Regression Transformer (R2T)
- Optimal Transport for Rectified Flow Image Editing: Unifying Inversion-Based and Direct Methods
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D Prototypes
- VLG-Net: Video-Language Graph Matching Network for Video Grounding
- Tackling Ill-posedness of Reversible Image Conversion with Well-posed Invertible Network
- CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- IAUNet: Instance-Aware U-Net
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- FASTER Recurrent Networks for Efficient Video Classification
- Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
- Rethinking Backbone Design for Lightweight 3D Object Detection in LiDAR
- On Learning Closed-Loop Probabilistic Multi-Agent Simulator
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Modeling turbulent and self-gravitating fluids with Fourier neural operators
- Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
- Simulation-based inference for Precision Neutrino Physics through Neural Monte Carlo tuning
- BS-NAS: Broadening-and-Shrinking One-Shot NAS with Searchable Numbers of Channels
- Robust Adverse Weather Removal via Spectral-based Spatial Grouping
- Adaptive Input Representations for Neural Language Modeling
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Exploiting Diffusion Prior for Task-driven Image Restoration
- Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
- Scaling and Distilling Transformer Models for sEMG
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation
- Cascading and Proxy Membership Inference Attacks
- Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
- FPConv: Learning Local Flattening for Point Convolution
- Annotation-Free Human Sketch Quality Assessment
- Parallel Grid Pooling for Data Augmentation
- STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction
- ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Pic2Diagnosis: A Method for Diagnosis of Cardiovascular Diseases from the Printed ECG Pictures
- Reverse engineering learned optimizers reveals known and novel\n mechanisms
- EnsembleNet: End-to-End Optimization of Multi-headed Models
- MintNet: Building Invertible Neural Networks with Masked Convolutions
- Pre- and Post-Treatment Glioma Segmentation with the Medical Imaging Segmentation Toolkit
- Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
- Modality Agnostic Efficient Long Range Encoder
- A multi-dynamic low-rank deep image prior (ML-DIP) for 3D real-time cardiovascular MRI
- WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design
- Semantics versus Identity: A Divide-and-Conquer Approach towards Adjustable Medical Image De-Identification
- UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis
- Patch Pruning Strategy Based on Robust Statistical Measures of Attention Weight Diversity in Vision Transformers
- On Arbitrary Predictions from Equally Valid Models
- DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
- Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction
- SemiSegECG: A Multi-Dataset Benchmark for Semi-Supervised Semantic Segmentation in ECG Delineation
- GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
- Hyperbolic Deep Learning for Chinese Natural Language Understanding
- Reducing Transformer Depth on Demand with Structured Dropout
- On the Intrinsic Dimensionality of Image Representations
- SpectroscopyNet: Learning to pre-process Spectroscopy Signals without clean data
- Aug3D-RPN: Improving Monocular 3D Object Detection by Synthetic Images with Virtual Depth
- Self-Distilled Self-Supervised Representation Learning
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- All-In-One: Artificial Association Neural Networks
- Adaptive Consistency Regularization for Semi-Supervised Transfer Learning
- RW-Resnet: A Novel Speech Anti-Spoofing Model Using Raw Waveform
- memeBot: Towards Automatic Image Meme Generation
- Learned Step Size Quantization
- Training Sparse Neural Networks using Compressed Sensing
- Learning State-Tracking from Code Using Linear RNNs
- NTIRE 2020 Challenge on Perceptual Extreme Super-Resolution: Methods and Results
- Feature Boosting, Suppression, and Diversification for Fine-Grained Visual Classification
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate\n Schedule
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Consistent Estimators for Learning to Defer to an Expert
Related