ImageNet Large Scale Visual Recognition Challenge
2014/09/01 by Olga Russakovsky, Jia Deng, Russakovsky, Olga +22 · 1 voice · 506 citations
Computer Science · #Image Retrieval and Classification Techniques #Advanced Image and Video Retrieval Techniques #Digital Imaging for Blood Diseases
paper · pdf · doi:10.48550/arxiv.1409.0575
Abstract
The ImageNet Large Scale Visual Recognition Challenge is a benchmark in object category classification and detection on hundreds of object categories and millions of images. The challenge has been run annually from 2010 to present, attracting participation from more than fifty institutions. This paper describes the creation of this benchmark dataset and the advances in object recognition that have been possible as a result. We discuss the challenges of collecting large-scale ground truth annotation, highlight key breakthroughs in categorical object recognition, provide a detailed analysis of the current state of the field of large-scale image classification and object detection, and compare the state-of-the-art computer vision accuracy with human accuracy. We conclude with lessons learned in the five years of the challenge, and propose future directions and improvements.
Cited by
- Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
- Decomposing Task Vectors for Refined Model Editing
- The Quest for Winning Tickets in Low-Rank Adapters
- A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Adaptive Acoustic Monitoring for Endangered Cook Inlet Beluga Whales in Complex Soundscapes
- SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation
- Epistemic Norms for AI Safety and Alignment Research
- DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models
- Multiclass Classification without Labels via Posterior Simplex Geometry
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence
- CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading
- Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge
- Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- DAPA: Distribution Aware Piecewise Activation Functions for On-Device Transformer Inference and Training
- Cognitive Dark Matter: Measuring What AI Misses
- Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
- Progressive Learned Image Compression for Machine Perception
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Event Extraction in Large Language Model
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- RMLer: Synthesizing Novel Objects across Diverse Categories via Reinforcement Mixing Learning
- Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Revisiting the Learning Objectives of Vision-Language Reward Models
- Local Patches Meet Global Context: Scalable 3D Diffusion Priors for Computed Tomography Reconstruction
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- Sharing Knowledge without Sharing Data: Stitches can improve ensembles of disjointly trained models
- Next-Embedding Prediction Makes Strong Vision Learners
- Open Ad-hoc Categorization with Contextualized Feature Learning
- In-Context Semi-Supervised Learning
- The Deleuzian Representation Hypothesis
- Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
- Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory
- An updated efficient galaxy morphology classification model based on ConvNeXt encoding with UMAP dimensionality reduction
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
- Spherical Leech Quantization for Visual Tokenization and Generation
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- Conditional Coverage Diagnostics for Conformal Prediction
- Real-Time Service Subscription and Adaptive Offloading Control in Vehicular Edge Computing
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
- RecTok: Reconstruction Distillation along Rectified Flow
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- Federated Learning with Feedback Alignment
- M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- On the Bayes Inconsistency of Disagreement Discrepancy Surrogates
- Agile Deliberation: Concept Deliberation for Subjective Visual Classification
- Neural Collapse in Test-Time Adaptation
- Inverse problems with diffusion models: MAP estimation via mode-seeking loss
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- Improving Multi-Class Calibration through Normalization-Aware Isotonic Techniques
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- Correction of Decoupled Weight Decay
- Improving the Sensitivity of Backdoor Detectors via Class Subspace Orthogonalization
- Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
- A Geometric Unification of Concept Learning with Concept Cones
- HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
- Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
- Neural expressiveness for beyond importance model compression
- ShadowWolf -- Automatic Labelling, Evaluation and Model Training Optimised for Camera Trap Wildlife Images
- Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
- LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
- On the Theoretical Foundation of Sparse Dictionary Learning in Mechanistic Interpretability
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot
- Diminishing Returns in Self-Supervised Learning
- Studying Various Activation Functions and Non-IID Data for Machine Learning Model Robustness
- Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
- Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptation
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Data-Centric Visual Development for Self-Driving Labs
- On the Unreasonable Effectiveness of Last-layer Retraining
- Directed evolution algorithm drives neural prediction
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
- PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images
- Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
- Saddle-Free Guidance: Improved On-Manifold Sampling without Labels or Additional Training
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher’s Bias
- Adversarial Flow Models
- AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
- Small Object Detection for Birds with Swin Transformer
- Merge and Bound: Direct Manipulations on Weights for Class Incremental Learning
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- Latent Diffusion Inversion Requires Understanding the Latent Space
- Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- XiCAD: Camera Activation Detection in the Da Vinci Xi User Interface
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
- Face, Whole-Person, and Object Classification in a Unified Space Via The Interleaved Multi-Domain Identity Curriculum
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- Flow Map Distillation Without Data
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- Experimental insights into data augmentation techniques for deep learning-based multimode fiber imaging: limitations and success
- Scalable Vision-Guided Crop Yield Estimation
- Robust Nonlinear Transform Coding: A Framework for Generalizable Joint Source-Channel Coding
- CoD: A Diffusion Foundation Model for Image Compression
- Bayesian-based Online Label Shift Estimation with Dynamic Dirichlet Priors
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss Landscapes
- SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
- Formal Abductive Latent Explanations for Prototype-Based Networks
- PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
- Exploiting Inter-Sample Information for Long-tailed Out-of-Distribution Detection
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
- Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
- Artificial intelligence approaches for energy-efficient laser cutting machines
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning
- EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
- Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- Quantifying and Improving Adaptivity in Conformal Prediction through Input Transformations
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
- Dynamic Temperature Scheduler for Knowledge Distillation
- Fast Data Attribution for Text-to-Image Models
- Mined Prompting and Metadata-Guided Generation for Wound Care Visual Question Answering
- Intrinsic Dimensionality as a Model-Free Measure of Class Imbalance
- FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
- When Thinking Pays Off: Incentive Alignment for Human-AI Collaboration
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- Stratified Knowledge-Density Super-Network for Scalable Vision Transformers
- Free-Boundary Quasiconformal Maps via a Least-squares Operator in Diffeomorphism Optimization
- Mitigating Negative Flips via Margin Preserving Training
- Rethinking Explanation Evaluation under the Retraining Scheme
- I2E: Real-Time Image-to-Event Conversion for High-Performance Spiking Neural Networks
- Multi-objective Hyperparameter Optimization in the Age of Deep Learning
- Learning Sparse Label Couplings for Multilabel Chest X-Ray Diagnosis
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
- FreqGRL: Suppressing Low-Frequency Bias and Mining High-Frequency Knowledge for Cross-Domain Few-Shot Learning
- Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- Toward Better Generalization in Few-Shot Learning through the Meta-Component Combination
- Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
- Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances
- Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
- ISC-Perception: A Hybrid Computer Vision Dataset for Object Detection in Novel Steel Assembly
- Benchmarking Attribution Methods with Relative Feature Importance
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- AI-Generated Image Detection: An Empirical Study and Future Research Directions
- CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- A Distributed Plug-and-Play MCMC Algorithm for High-Dimensional Inverse Problems
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Why Federated Optimization Fails to Achieve Perfect Fitting? A Theoretical Perspective on Client-Side Optima
- FedSM: Robust Semantics-Guided Feature Mixup for Bias Reduction in Federated Learning with Long-Tail Data
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- Gaussian Combined Distance: A Generic Metric for Object Detection
- Masked Diffusion Captioning for Visual Feature Learning
- Emu3.5: Native Multimodal Models are World Learners
- Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM
- Active Learning with Task-Driven Representations for Messy Pools
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- Character-level Chinese Writer Identification using Path Signature Feature, DropStroke and Deep CNN
- Product-Quantised Image Representation for High-Quality Image Synthesis
- FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning
- Can neurons speak? Semantic narration of vision at single-cell resolution
- Scene-Centric Unsupervised Video Panoptic Segmentation
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- Is Dimensionality a Barrier for Retrieval Models?
- Confounding factors and biases abound when predicting molecular biomarkers from histological images
- OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Symmetry and Generalisation in Neural Approximations of Renormalisation Transformations
- Neural Tuning for Ordinal Processing: Convergent Patterns in Human Brains and Artificial Networks
- Adversarially Robust Quantum Transfer Learning
- Region-CAM: Towards Accurate Object Regions in Class Activation Maps for Weakly Supervised Learning Tasks
- Resi-VidTok: An Efficient and Decomposed Progressive Tokenization Framework for Ultra-Low-Rate and Lightweight Video Transmission
- Fair Indivisible Payoffs through Shapley Value
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Differentiable, Bit-shifting, and Scalable Quantization without training neural network from scratch
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- A Versatile Framework for Designing Group-Sparse Adversarial Attacks
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Improving the Straight-Through Estimator with Zeroth-Order Information
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- Symmetria: A Synthetic Dataset for Learning in Point Clouds
- The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Adaptive Stochastic Coefficients for Accelerating Diffusion Sampling
- Smart Sensor Placement: A Correlation-Aware Attribution Framework (CAAF) for Real-world Data Modeling
- Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
- GALA: A GlobAl-LocAl Approach for Multi-Source Active Domain Adaptation
- GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
- Caption-Driven Explainability: Probing CNNs for Bias via CLIP
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- Bridging the gap to real-world language-grounded visual concept learning
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease
- Elementary, My Dear Watson: Non-Invasive Neural Keyword Spotting in the LibriBrain Dataset
- More Than Memory Savings: Zeroth-Order Optimization Mitigates Forgetting in Continual Learning
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Hardware-Aware DNN Compression for Homogeneous Edge Devices
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- Transformed Multi-view 3D Shape Features with Contrastive Learning
- Towards Strong Certified Defense with Universal Asymmetric Randomization
- Matrix-Free Least Squares Solvers: Values, Gradients, and What to Do With Them
- A New Type of Adversarial Examples
- Weight Decay may matter more than muP for Learning Rate Transfer in Practice
- Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Elastic ViTs from Pretrained Models without Retraining
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
- Latent Diffusion Model without Variational Autoencoder
- DPTrack:Directional Kernel-Guided Prompt Learning for Robust Nighttime Aerial Tracking
- You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
- Free-Grained Hierarchical Recognition
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- Diffusion Transformers with Representation Autoencoders
- Robust Plant Disease Diagnosis with Few Target-Domain Samples
- Device Placement Optimization with Reinforcement Learning
- Contrastive Dimension Reduction: A Systematic Review
- Neural Weight Compression for Language Models
- Test-Time Adaptation by Causal Trimming
- Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
- Efficient Edge Test-Time Adaptation via Latent Feature Coordinate Correction
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Stroke Locus Net: Occluded Vessel Localization from MRI Modalities
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Unsupervised Dynamic Feature Selection for Robust Latent Spaces in Vision Tasks
- Representational Alignment Across Model Layers and Brain Regions with Hierarchical Optimal Transport
- PHyCLIP: ℓ1-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- Automated Evolutionary Optimization for Resource-Efficient Neural Network Training
- Enhancing Self-Supervised Learning with Semantic Pairs A New Dataset and Empirical Study
- Adaptive Gradient Calibration for Single-Positive Multi-Label Learning in Remote Sensing Image Scene Classification
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Demystifying Deep Learning-based Brain Tumor Segmentation with 3D UNets and Explainable AI (XAI): A Comparative Analysis
- DarkHash: A Data-Free Backdoor Attack Against Deep Hashing
- Long-tailed Recognition with Model Rebalancing
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Benchmarking is Broken -- Don't Let AI be its Own Judge
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Federated Unlearning in the Wild: Rethinking Fairness and Data Discrepancy
- Sharpness-Aware Data Generation for Zero-shot Quantization
- Revisiting Mixout: An Overlooked Path to Robust Finetuning
- Consistent Assistant Domains Transformer for Source-free Domain Adaptation
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- A Novel Technique for Robust Training of Deep Networks With Multisource Weak Labeled Remote Sensing Data
- BuilderBench -- A benchmark for generalist agents
- A Data-Driven Prism: Multi-View Source Separation with Diffusion Model Priors
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- Detection of retinal diseases using an accelerated reused convolutional network
- Sliding Window Attention for Learned Video Compression
- Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL
- SAFA-SNN: Sparsity-Aware On-Device Few-Shot Class-Incremental Learning with Fast-Adaptive Structure of Spiking Neural Network
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
- Robust Context-Aware Object Recognition
- Attack logics, not outputs: Towards efficient robustification of deep neural networks by falsifying concept-based properties
- Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
- Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
- SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
- Cutting the Skip: Training Residual-Free Transformers
- Cat: Post-Training Quantization Error Reduction via Cluster-based Affine Transformation
- From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
- The Impact of Scaling Training Data on Adversarial Robustness
- OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
- Emergent evaluation hubs in a decentralizing large language model ecosystem
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- Echoes of Humanity: Exploring the Perceived Humanness of AI Music
- Score-based Membership Inference on Diffusion Models
- Vehicle Classification under Extreme Imbalance: A Comparative Study of Ensemble Learning and CNNs
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Foveated Retinotopy Improves Classification and Localization in Convolutional Neural Networks
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- Dynamics of Learning: Generative Schedules from Latent ODEs
- Hemorica: A Comprehensive CT Scan Dataset for Automated Brain Hemorrhage Classification, Segmentation, and Detection
- Convolutional Set Transformer
- Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
- CCNeXt: An Effective Self-Supervised Stereo Depth Estimation Approach
- Activation Function Design Sustains Plasticity in Continual Learning
- Category Discovery: An Open-World Perspective
- EfficientDepth: A Fast and Detail-Preserving Monocular Depth Estimation Model
- DeCAF: A Deep Convolutional Activation Feature for Generic Visual\n Recognition
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions in Semi-Supervised Learning
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- CompressAI-Vision: Open-source software to evaluate compression methods for computer vision tasks
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- GraspFactory: A Large Object-Centric Grasping Dataset
- Efficiently Attacking Memorization Scores
- Table Detection with Active Learning
- RFMSR: Residual Flow Matching for Image Super-Resolution
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- Assessing the Alignment of Popular CNNs to the Brain for Valence Appraisal
- Long Story Short: Disentangling Compositionality and Long-Caption Understanding in VLMs
- A Validation Strategy for Deep Learning Models: Evaluating and Enhancing Robustness
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- FlowCrypt: Flow-Based Lightweight Encryption with Near-Lossless Recovery for Cloud Photo Privacy
- Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
- A Single Image Is All You Need: Zero-Shot Anomaly Localization Without Training Data
- Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Training-Free Label Space Alignment for Universal Domain Adaptation
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- Modeling Human Concepts with Subspaces in Deep Vision Models
- From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge
- MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
- Convolutional Neural Network Optimization for Beehive Classification Using Bioacoustic Signals
- PRISM: Precision-Recall Informed Data-Free Knowledge Distillation via Generative Diffusion
- \boldsymbolλ-Orthogonality Regularization for Compatible Representation Learning
- Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
- CoUn: Empowering Machine Unlearning via Contrastive Learning
- Generating Part-Based Global Explanations Via Correspondence
- Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
- [Re] Improving Interpretation Faithfulness for Vision Transformers
- Adversarial Examples Are Not Bugs, They Are Superposition
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- VCBench: Benchmarking LLMs in Venture Capital
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
- SAIL-VL2 Technical Report
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- IS-Diff: Improving Diffusion-Based Inpainting with Better Initial Seed
- Adaptive Spatial Goodness Encoding: Advancing and Scaling Forward-Forward Learning Without Backpropagation
- MAUI: Reconstructing Private Client Data in Federated Transfer Learning
- A Novel Local Focusing Mechanism for Deepfake Detection Generalization
- Some Robustness Properties of Label Cleaning
- SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
- Hierarchical MLANet: Multi-level Attention for 3D Face Reconstruction From Single Images
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Semantic Concentration for Self-Supervised Dense Representations Learning
- Image Recognition with Vision and Language Embeddings of VLMs
- Patch-based Automatic Rosacea Detection Using the ResNet Deep Learning Framework
- Compressing CNN models for resource-constrained systems by channel and layer pruning
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- How Far Are We from True Unlearnability?
- Feature Space Analysis by Guided Diffusion Model
- Object-level Correlation for Few-Shot Segmentation
- Basis Vector Metric: A Method for Robust Open-Ended State Change Detection
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
- Dimensionally Reduced Open-World Clustering: DROWCULA
- Performance of Conformal Prediction in Capturing Aleatoric Uncertainty
- Parameter-Free Logit Distillation via Sorting Mechanism
- SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
- DCV-ROOD Evaluation Framework: Dual Cross-Validation for Robust Out-of-Distribution Detection
- Toward Efficient and Scalable Design of In-Memory Graph-Based Vector Search
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways
- Noisy Label Refinement with Semantically Reliable Synthetic Images
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
- TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
- Isolated Bangla Handwritten Character Classification using Transfer Learning
- Joint Training of Image Generator and Detector for Road Defect Detection
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- Ordinal Adaptive Correction: A Data-Centric Approach to Ordinal Image Classification with Noisy Labels
- Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
- Improving atomic force microscopy structure discovery via style-translation
- Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
- BM-CL: Bias Mitigation through the lens of Continual Learning
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
- Localizing and Mitigating Memorization in Image Autoregressive Models
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations
- Activation Subspaces for Out-of-Distribution Detection
- Multi-Method Ensemble for Out-of-Distribution Detection
- Representation Learning with Adaptive Superpixel Coding
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- Occlusion Robustness of CLIP for Military Vehicle Classification
- Coresets from Trajectories: Selecting Data via Correlation of Loss Differences
- The Role of Teacher Calibration in Knowledge Distillation
- Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
- Efficient Multi-Source Knowledge Transfer by Model Merging
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- CARMA: Collocation-Aware Resource Manager
- A Deep Learning Application for Psoriasis Detection
- Assessing the Noise Robustness of Class Activation Maps: A Framework for Reliable Model Interpretability
- Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- Understanding Data Influence with Differential Approximation
- OASIS: Open-world Adaptive Self-supervised and Imbalanced-aware System
- DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
- A Guide for Manual Annotation of Scientific Imagery: How to Prepare for Large Projects
- GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- SEDEG:Sequential Enhancement of Decoder and Encoder's Generality for Class Incremental Learning with Small Memory
- Hierarchical Conformal Classification
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
- Geometry-Aware Video Inpainting for Joint Headset Occlusion Removal and Face Reconstruction in Social XR
- Infusing fine-grained visual knowledge to Vision-Language Models
- Towards interpretable prediction of recurrence risk in breast cancer using pathology foundation models
- FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
- What Matters for Bioacoustic Encoding
- AIM: Amending Inherent Interpretability via Self-Supervised Masking
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- Semantically Guided Adversarial Testing of Vision Models Using Language Models
- Probing the Representational Power of Sparse Autoencoders in Vision Models
- Versatile Video Tokenization with Generative 2D Gaussian Splatting
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- MPT: Motion Prompt Tuning for Micro-Expression Recognition
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- Separating Knowledge and Perception with Procedural Data
- Frequency-Assisted Adaptive Sharpening Scheme Considering Bitrate and Quality Tradeoff
- IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- Score Augmentation for Diffusion Models
- AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
- ProteoKnight: Convolution-based phage virion protein classification and uncertainty analysis
- BEVANet: Bilateral Efficient Visual Attention Network for Real-Time Semantic Segmentation
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- Towards Robust Red-Green Watermarking for Autoregressive Image Generators
- Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- TopKD: Top-scaled Knowledge Distillation
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Investigating the Impact of Large-Scale Pre-training on Nutritional Content Estimation from 2D Images
- Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
- GaitAdapt: Continual Learning for Evolving Gait Recognition
- Diffusion Models with Adaptive Negative Sampling Without External Resources
- FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
- Towards High Precision: An Adaptive Self-Supervised Learning Framework for Force-Based Verification
- DySTop
- Tackling Ill-posedness of Reversible Image Conversion with Well-posed Invertible Network
- InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
- Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
- Implementing Neural Networks Over-the-Air via Reconfigurable Intelligent Surfaces
- Self-Navigated Residual Mamba for Universal Industrial Anomaly Detection
- Enhancing Diffusion-based Dataset Distillation via Adversary-Guided Curriculum Sampling
- COSTARR: Consolidated Open Set Technique with Attenuation for Robust Recognition
- Structured Spectral Graph Learning for Anomaly Classification in 3D Chest CT Scans
- Reducing the gap between general purpose data and aerial images in concentrated solar power plants
- Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
- VQ-DeepISC: Vector Quantized-Enabled Digital Semantic Communication with Channel Adaptive Image Transmission
- Object-Centric Cropping for Visual Few-Shot Classification
- Continual Learning with Synthetic Boundary Experience Blending
- Causal Identification of Sufficient, Contrastive and Complete Feature Sets in Image Classification
- Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
- PixNerd: Pixel Neural Field Diffusion
Discussions
Related