EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
2019/05/28 by Mingxing Tan, Quoc V. Le · 1 voice · 932 citations
Computer Science · Mathematics · #cs.CV #cs.LG #stat.ML
paper · pdf
published as International Conference on Machine Learning, 2019 · ICML 2019
arxiv created 2020/09/11 · arxiv updated 2020/09/14
Abstract
Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available. In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance. Based on this observation, we propose a new scaling method that uniformly scales all dimensions of depth/width/resolution using a simple yet highly effective compound coefficient. We demonstrate the effectiveness of this method on scaling up MobileNets and ResNet. To go even further, we use neural architecture search to design a new baseline network and scale it up to obtain a family of models, called EfficientNets, which achieve much better accuracy and efficiency than previous ConvNets. In particular, our EfficientNet-B7 achieves state-of-the-art 84.3% top-1 accuracy on ImageNet, while being 8.4x smaller and 6.1x faster on inference than the best existing ConvNet. Our EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 (91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer parameters. Source code is at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet.
Cited by
- A hybrid CNN-machine learning model for anxiety screening with mandala coloring patterns: Feature extraction and classification performance
- Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
- LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models
- ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors
- A Statistical Multi-Objective Framework for Assessing Sensitivity of Radiomic AI Models to Acquisition Parameters
- Traceback Translators Against Forgetting in Continual Fake Speech Detection
- Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
- WiFi Sensing via Reservoir Computing
- Pixel-Space Diffusion Transformers
- SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
- One Round Is All You Need: Analytic Federated Learning for Task-Heterogeneous Multi-Label Medical Image Classification
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- Cross-Dataset Generalization in Breast MRI Tumor Classification via Class-Wise Dataset Mixing
- Privileged Lesion-Context Relational Distillation for Mask-Free Skin Lesion Classification
- CITRUS: Candidate Inference and Temporal-tracking for Reliable, Unobtrusive Sensing of Wearable Heart Rate under Motion
- Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator
- TinyGLASS: Real-Time Self-Supervised In-Sensor Anomaly Detection
- When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection
- DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
- EA-RMENet -- Path Loss Prediction in Urban Environments using Deep Learning
- AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis
- Improving Backward Conformal Prediction via Non-Conformity Score Transformation
- When Bigger is Worse: A Practitioner's Guide to Model Selection Under Data Scarcity
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
- Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study
- DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- OvAi Focus: AI-based Multi-class Segmentation of Functional Ovaries and Adnexal Masses in Gynecological Ultrasound
- Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
- Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- Fit for Purpose? Deepfake Detection in the Real World
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- Deep learning approach to detect and visualize sexual dimorphism in monomorphic species
- Perception Encoder: The best visual embeddings are not at the output of the network
- Accurate phenotyping of luminal A breast cancer in magnetic resonance imaging: A new 3D CNN approach
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- ILIAS: Instance-Level Image retrieval At Scale
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Exploring Syn-to-Real Domain Adaptation for Military Target Detection
- Real-Time In-Cabin Driver Behavior Recognition on Low-Cost Edge Hardware
- Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
- Balanced Soft mixture-of-expert model for Glaucoma Detection
- Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
- Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution
- Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
- Real-time Reconstruction of Human Visual Perception from fMRI
- Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- Visual information extraction from documents via classification-guided large vision-language models
- BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving
- Patch-Discontinuity Mining for Generalized Deepfake Detection
- AI for Mycetoma Diagnosis in Histopathological Images: The MICCAI 2024 Challenge
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- MODE: Multi-Objective Adaptive Coreset Selection
- Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- MedNeXt-v2: Scaling 3D ConvNeXts for Large-Scale Supervised Representation Learning in Medical Image Segmentation
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Generating Risky Samples with Conformity Constraints via Diffusion Models
- MeniMV: A Multi-view Benchmark for Meniscus Injury Severity Grading
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
- Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
- Domain-Aware Quantum Circuit for QML
- Training Together, Diagnosing Better: Federated Learning for Collagen VI-Related Dystrophies
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- Exploring Deep-to-Shallow Transformable Neural Networks for Intelligent Embedded Systems
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
- MFE-GAN: Efficient GAN-based Framework for Document Image Enhancement and Binarization with Multi-scale Feature Extraction
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- DL3M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- ACCOR: Attention-Enhanced Complex-Valued Contrastive Learning for Occluded Object Classification Using mmWave Radar IQ Signals
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Sequence-to-Image Transformation for Sequence Classification Using Rips Complex Construction and Chaos Game Representation
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- NeuroSketch: An Effective Framework for Neural Decoding via Systematic Architectural Optimization
- MelanomaNet: Explainable Deep Learning for Skin Lesion Classification
- FBA2D: Frequency-based Black-box Attack for AI-generated Image Detection
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Data-Efficient Learning of Anomalous Diffusion with Wavelet Representations: Enabling Direct Learning from Experimental Trajectories
- Animal Re-Identification on Microcontrollers
- VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
- UltrasODM: A Dual Stream Optical Flow Mamba Network for 3D Freehand Ultrasound Reconstruction
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- Reevaluating Automated Wildlife Species Detection: A Reproducibility Study on a Custom Image Dataset
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
- Leveraging Pre-trained Neural Network Models for the Classification of Tumor Cells Analyzed by Label-free Phase Holotomographic Microscopy
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Hierarchical Deep Learning for Diatom Image Classification: A Multi-Level Taxonomic Approach
- Phase-OTDR Event Detection Using Image-Based Data Transformation and Deep Learning
- Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
- Invariance Co-training for Robot Visual Generalization
- DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification
- Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World
- Open Set Face Forgery Detection via Dual-Level Evidence Collection
- EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification
- Defense That Attacks: How Robust Models Become Better Attackers
- Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
- Breast Cell Segmentation Under Extreme Data Constraints: Quantum Enhancement Meets Adaptive Loss Stabilization
- Data-Centric Visual Development for Self-Driving Labs
- nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
- M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
- Implementation of a Skin Lesion Detection System for Managing Children with Atopic Dermatitis Based on Ensemble Learning
- ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
- Stacked Ensemble of Fine-Tuned CNNs for Knee Osteoarthritis Severity Grading
- Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
- TeleViT1.0: Teleconnection-aware Vision Transformers for Subseasonal to Seasonal Wildfire Pattern Forecasts
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- Transformer Driven Visual Servoing for Fabric Texture Matching Using Dual-Arm Manipulator
- RefOnce: Distilling References into a Prototype Memory for Referring Camouflaged Object Detection
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- CellFMCount: A Fluorescence Microscopy Dataset, Benchmark, and Methods for Cell Counting
- SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
- Dendritic Convolution for Noise Image Recognition
- Shape-Adapting Gated Experts: Dynamic Expert Routing for Colonoscopic Lesion Segmentation
- LungX: A Hybrid EfficientNet-Vision Transformer Architecture with Multi-Scale Attention for Accurate Pneumonia Detection
- Stro-VIGRU: Defining the Vision Recurrent-Based Baseline Model for Brain Stroke Classification
- Compact neural networks for astronomy with optimal transport bias correction
- A Lightweight, Interpretable Deep Learning System for Automated Detection of Cervical Adenocarcinoma In Situ (AIS)
- Equivariant-Aware Structured Pruning for Efficient Edge Deployment: A Comprehensive Framework with Adaptive Fine-Tuning
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- A Diversity-optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
- Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
- Memory-DD: A Low-Complexity Dendrite-Inspired Neuron for Temporal Prediction Tasks
- Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
- Monocular 3D Lane Detection via Structure Uncertainty-Aware Network with Curve-Point Queries
- H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
- Online Continual Learning on Intel Loihi 2 via a Co-designed Spiking Neural Network
- MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
- SAGE: Saliency-Guided Contrastive Embeddings
- Floor Plan-Guided Visual Navigation Incorporating Depth and Directional Cues
- Lightweight Optimal-Transport Harmonization on Edge Devices
- Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
- TEMPO: Global Temporal Building Density and Height Estimation from Satellite Imagery
- MFI-ResNet: Efficient ResNet Architecture Optimization via MeanFlow Compression and Selective Incubation
- SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
- Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
- D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces
- Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts
- DermAI: Clinical dermatology acquisition through quality-driven image collection for AI classification in mobile
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
- Federated Learning for Pediatric Pneumonia Detection: Enabling Collaborative Diagnosis Without Sharing Patient Data
- REASON: Probability map-guided dual-branch fusion framework for gastric content assessment
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- A vascular code for speed in the spatial navigation system
- Rethinking Explanation Evaluation under the Retraining Scheme
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- LatentPrintFormer: A Hybrid CNN-Transformer with Spatial Attention for Latent Fingerprint identification
- SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
- Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- Physics-Informed Image Restoration via Progressive PDE Integration
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- NeuroFlex: Column-Exact ANN-SNN Co-Execution Accelerator with Cost-Guided Scheduling
- Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
- A Lightweight 3D-CNN for Event-Based Human Action Recognition with Privacy-Preserving Potential
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- FastBoost: Progressive Attention with Dynamic Scaling for Efficient Deep Learning
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Weakly Supervised Pneumonia Localization from Chest X-Rays Using Deep Neural Network and Grad-CAM Explanations
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- How Close Are We? Limitations and Progress of AI Models in Banff Lesion Scoring
- ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding
- Surpassing state of the art on AMD area estimation from RGB fundus images through careful selection of U-Net architectures and loss functions for class imbalance
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
- CNN-Based Reanalysis of Optical Turbulence at the Canary Islands Observatories (OCAN)
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
- HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework
- LLM-to-Phy3D: Physically Conform Online 3D Object Generation with LLMs
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Selective Diabetic Retinopathy Screening with Accuracy-Weighted Deep Ensembles and Entropy-Guided Abstention
- Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- Fair Indivisible Payoffs through Shapley Value
- A Dual-Branch CNN for Robust Detection of AI-Generated Facial Forgeries
- All in one timestep: Enhancing Sparsity and Energy efficiency in Multi-level Spiking Neural Networks
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Training-free Source Attribution of AI-generated Images via Resynthesis
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- PLDC‐Net: A Domain‐Specific Base Model for Plant Leaf Disease Classification Domain Adaptation Tasks
- Does Machine Learning Work? A Comparative Analysis of Strong Gravitational Lens Searches in the Dark Energy Survey
- Explainable Detection of AI-Generated Images with Artifact Localization Using Faster-Than-Lies and Vision-Language Models for Edge Devices
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- Synthetic-to-Real Transfer Learning for Chromatin-Sensitive PWS Microscopy
- DiffusionLane: Diffusion Model for Lane Detection
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- HybridSOMSpikeNet: A Deep Model with Differentiable Soft Self-Organizing Maps and Spiking Dynamics for Waste Classification
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Balanced Multi-Task Attention for Satellite Image Classification: A Systematic Approach to Achieving 97.23% Accuracy on EuroSAT Without Pre-Training
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Knowledge-Informed Neural Network for Complex-Valued SAR Image Recognition
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
- BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- TinyUSFM: Towards Compact and Efficient Ultrasound Foundation Models
- Enhancing Early Alzheimer Disease Detection through Big Data and Ensemble Few-Shot Learning
- Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation
- OmniNWM: Omniscient Driving Navigation World Models
- Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
- ViBED-Net: Video Based Engagement Detection Network Using Face-Aware and Scene-Aware Spatiotemporal Cues
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Fair and Interpretable Deepfake Detection in Videos
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
- Learning to play: A Multimodal Agent for 3D Game-Play
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- Bridging Symmetry and Robustness: On the Role of Equivariance in Enhancing Adversarial Robustness
- A Multi-Task Deep Learning Framework for Skin Lesion Classification, ABCDE Feature Quantification, and Evolution Simulation
- BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection
- PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Removing Cost Volumes from Optical Flow Estimators
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
- Lightweight CNN-Based Wi-Fi Intrusion Detection Using 2D Traffic Representations
- Efficient Edge Test-Time Adaptation via Latent Feature Coordinate Correction
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
- Learning Model Representations Using Publicly Available Model Hubs
- HccePose(BF): Predicting Front & Back Surfaces to Construct Ultra-Dense 2D-3D Correspondences for Pose Estimation
- KTBox: A Modular LaTeX Framework for Semantic Color, Structured Highlighting, and Scholarly Communication
- Exploration of Incremental Synthetic Non-Morphed Images for Single Morphing Attack Detection
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- Structured Output Regularization: a framework for few-shot transfer learning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Weights initialization of neural networks for function approximation
- Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection
- Bridged Clustering: Semi-Supervised Sparse Bridging
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Image-based recognition using advanced neural networks can aid surveillance of Agrilus jewel beetles
- Clinical-grade AI model for molecular subtyping of endometrial cancer: a multi-center cohort study in China
- Mangrove3D: Terrestrial Laser Scanning Dataset for Coastal Mangrove Forests
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- Scalable deep fusion of spaceborne lidar and synthetic aperture radar for global forest structural complexity mapping
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Adaptive Coverage Policies in Conformal Prediction
- Detection of retinal diseases using an accelerated reused convolutional network
- Beyond Static Knowledge Messengers: Towards Adaptive, Fair, and Scalable Federated Learning for Medical AI
- Talking Tennis: Language Feedback from 3D Biomechanical Action Recognition
- CVSM: Contrastive Vocal Similarity Modeling
- ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation Testing
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
- A Fast and Precise Method for Searching Rectangular Tumor Regions in Brain MR Images
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
- Confidence Aware SSD Ensemble with Weighted Boxes Fusion for Weapon Detection
- On The Variability of Concept Activation Vectors
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning
- Evaluating the Impact of Radiographic Noise on Chest X-ray Semantic Segmentation and Disease Classification Using a Scalable Noise Injection Framework
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- Calibrated and Resource-Aware Super-Resolution for Reliable Driver Behavior Analysis
- S3F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- Hemorica: A Comprehensive CT Scan Dataset for Automated Brain Hemorrhage Classification, Segmentation, and Detection
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- Multilingual Vision-Language Models, A Survey
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- No-Reference Image Contrast Assessment with Customized EfficientNet-B0
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions in Semi-Supervised Learning
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
- EnGraf-Net: Multiple Granularity Branch Network with Fine-Coarse Graft Grained for Classification Task
- SD-RetinaNet: Topologically Constrained Semi-Supervised Retinal Lesion and Layer Segmentation in OCT
- RAM-NAS: Resource-aware Multiobjective Neural Architecture Search Method for Robot Vision Tasks
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
- Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
- Affective Computing and Emotional Data: Challenges and Implications in Privacy Regulations, The AI Act, and Ethics in Large Language Models
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
- DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
- DL-QC-fNIRS: a deep learning tool for automated quality control in functional near-infrared spectroscopy signals
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
- Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data
- MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
- TS-P2CL: Plug-and-Play Dual Contrastive Learning for Vision-Guided Medical Time Series Classification
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
- MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
- V-CECE: Visual Counterfactual Explanations via Conceptual Edits
- Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- PM25Vision: A Large-Scale Benchmark Dataset for Visual Estimation of Air Quality
- Benchmarking Class Activation Map Methods for Explainable Brain Hemorrhage Classification on Hemorica Dataset
- Efficient Conformal Prediction for Regression Models under Label Noise
- NeRF-based Visualization of 3D Cues Supporting Data-Driven Spacecraft Pose Estimation
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
- MetricNet: Recovering Metric Scale in Generative Navigation Policies
- Deep Lookup Network
- Real-Time Detection and Tracking of Foreign Object Intrusions in Power Systems via Feature-Based Edge Intelligence
- Performance is not All You Need: Sustainability Considerations for Algorithms
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Human + AI for Accelerating Ad Localization Evaluation
- TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images
- T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
- Automated Landfill Detection Using Deep Learning: A Comparative Study of Lightweight and Custom Architectures with the AerialWaste Dataset
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- The Quest for Universal Master Key Filters in DS-CNNs
- Optimizing Class Distributions for Bias-Aware Multi-Class Learning
- Synthetic vs. Real Training Data for Visual Navigation
- Neural networks in the search for fast radio bursts with RATAN-600
- GraphDerm: Fusing Imaging, Physical Scale, and Metadata in a Population-Graph Classifier for Dermoscopic Lesions
- Hybrid Quantum Neural Networks for Efficient Protein-Ligand Binding Affinity Prediction
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
- An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift
- TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- Tri-Accel: Curvature-Aware Precision-Adaptive and Memory-Elastic Optimization for Efficient GPU Usage
- A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- Compressing CNN models for resource-constrained systems by channel and layer pruning
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- MedicalPatchNet: A Patch-Based Self-Explainable AI Architecture for Chest X-ray Classification
- EfficientNet in Digital Twin-based Cardiac Arrest Prediction and Analysis
- Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
- Improved Classification of Nitrogen Stress Severity in Plants Under Combined Stress Conditions Using Spatio-Temporal Deep Learning Framework
- NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter Editing
- Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
- Realism to Deception: Investigating Deepfake Detectors Against Face Enhancement
- Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
- Khana: A Comprehensive Indian Cuisine Dataset
- Challenges in Deep Learning-Based Small Organ Segmentation: A Benchmarking Perspective for Medical Research with Limited Datasets
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Interpretable Deep Transfer Learning for Breast Ultrasound Cancer Detection: A Multi-Dataset Study
- MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment using UAV Imagery
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
- Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- PracMHBench: Re-evaluating Model-Heterogeneous Federated Learning Based on Practical Edge Device Constraints
- Chest X-ray Pneumothorax Segmentation Using EfficientNet-B4 Transfer Learning in a U-Net Architecture
- Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- MedLiteNet: Lightweight Hybrid Medical Image Segmentation Model
- Prospects for acoustically monitoring ecosystem tipping points
- Vision-Based Embedded System for Noncontact Monitoring of Preterm Infant Behavior in Low-Resource Care Settings
- An Investigation of Visual Foundation Models Robustness
- An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
- Enhancing Fitness Movement Recognition with Attention Mechanism and Pre-Trained Feature Extractors
- Fair Resource Allocation for Fleet Intelligence
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
- AgroSense: An Integrated Deep Learning System for Crop Recommendation via Soil Image Analysis and Nutrient Profiling
- Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
- PrediTree: A Multi-Temporal Sub-meter Dataset of Multi-Spectral Imagery Aligned With Canopy Height Maps
- AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef
- AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
- Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- diveXplore at the Video Browser Showdown 2024
- Unified Acoustic Representations for Screening Neurological and Respiratory Pathologies from Voice
- ATMS-KD: Adaptive Temperature and Mixed Sample Knowledge Distillation for a Lightweight Residual CNN in Agricultural Embedded Systems
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- Paired-Sampling Contrastive Framework for Joint Physical-Digital Face Attack Detection
- Computing-In-Memory Dataflow for Minimal Buffer Traffic
- Formal Algorithms for Model Efficiency
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Unleashing Semantic and Geometric Priors for 3D Scene Completion
- Pixels Under Pressure: Exploring Fine-Tuning Paradigms for Foundation Models in High-Resolution Medical Imaging
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
- Skin Cancer Classification: Hybrid CNN-Transformer Models with KAN-Based Fusion
- OrbitChain: Orchestrating In-orbit Real-time Analytics of Earth Observation Data
- An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Dual-species atomic absorption image reconstruction using deep neural networks
- MultiPark: Multimodal Parking Transformer with Next-Segment Prediction
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- Scalable Geospatial Data Generation Using AlphaEarth Foundations Model
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- Biasing Frontier-Based Exploration with Saliency Areas
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- T-CACE: A Time-Conditioned Autoregressive Contrast Enhancement Multi-Task Framework for Contrast-Free Liver MRI Synthesis, Segmentation, and Diagnosis
- Predictive Uncertainty for Runtime Assurance of a Real-Time Computer Vision-Based Landing System
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- Large-Small Model Collaborative Framework for Federated Continual Learning
- Deep Learning for Automated Identification of Vietnamese Timber Species: A Tool for Ecological Monitoring and Conservation
- Autonomous AI Bird Feeder for Backyard Biodiversity Monitoring
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Low-Regret and Low-Complexity Learning for Hierarchical Inference
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Towards Human-AI Collaboration System for the Detection of Invasive Ductal Carcinoma in Histopathology Images
- MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training
- From Field to Drone: Domain Drift Tolerant Automated Multi-Species and Damage Plant Semantic Segmentation for Herbicide Trials
- Neural Tangent Knowledge Distillation for Optical Convolutional Networks
- DragonFruitQualityNet: A Lightweight Convolutional Neural Network for Real-Time Dragon Fruit Quality Inspection on Mobile Devices
- Position: Ideas Should be the Center of Machine Learning Research
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Don't Reach for the Stars: Rethinking Topology for Resilient Federated Learning
- ULU: A Unified Activation Function
- Perch 2.0: The Bittern Lesson for Bioacoustics
- Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
- Improving Tactile Gesture Recognition with Optical Flow
- From Flat to Round: Redefining Brain Decoding with Surface-Based fMRI and Cortex Structure
- Automated ultrasound doppler angle estimation using deep learning
- Energy-Efficient Real-Time 4-Stage Sleep Classification at 10-Second Resolution: A Comprehensive Study
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- FeDaL: Federated Dataset Learning for Time Series Foundation Models
- Uncertainty-aware Accurate Elevation Modeling for Off-road Navigation via Neural Processes
- FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose Estimation
- DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
- Inductive transfer learning from regression to classification in ECG analysis
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs
- Transfer Learning with EfficientNet for Accurate Leukemia Cell Classification
- Rate-distortion Optimized Point Cloud Preprocessing for Geometry-based Point Cloud Compression
- Benchmarking Adversarial Patch Selection and Location
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- ESM: A Framework for Building Effective Surrogate Models for Hardware-Aware Neural Architecture Search
- Foundation Models for Bioacoustics -- a Comparative Review
- Deep Learning for Pavement Condition Evaluation Using Satellite Imagery
- Classification of Brain Tumors using Hybrid Deep Learning Models
- Rethinking Backbone Design for Lightweight 3D Object Detection in LiDAR
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
- Exploring Superposition and Interference in State-of-the-Art Low-Parameter Vision Models
- Label tree semantic losses for rich multi-class medical image segmentation
- Visual-Language Model Knowledge Distillation Method for Image Quality Assessment
- Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
- Robust Deepfake Detection for Electronic Know Your Customer Systems Using Registered Images
- Cyst-X: A Federated AI System Outperforms Clinical Guidelines to Detect Pancreatic Cancer Precursors and Reduce Unnecessary Surgery
- GeMix: Conditional GAN-Based Mixup for Improved Medical Image Augmentation
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Data Aware Differentiable Neural Architecture Search for Tiny Keyword Spotting Applications
- Evaluating Deepfake Detectors in the Wild
- Suppressing Gradient Conflict for Generalizable Deepfake Detection
- DeSamba: Decoupled Spectral Adaptive Framework for 3D Multi-Sequence MRI Lesion Classification
- Evaluating Deep Learning Models for African Wildlife Image Classification: From DenseNet to Vision Transformers
- Embedding-Aware Quantum-Classical SVMs for Scalable Quantum Machine Learning
- HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
- A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games
- Can Foundation Models Predict Fitness for Duty?
- VLMPlanner: Integrating Visual Language Models with Motion Planning
- AnimalClue: Recognizing Animals by their Traces
- MambaMap: Online Vectorized HD Map Construction using State Space Model
- VAMPIRE: Uncovering Vessel Directional and Morphological Information from OCTA Images for Cardiovascular Disease Risk Factor Prediction
- Taming Domain Shift in Multi-source CT-Scan Classification via Input-Space Standardization
- Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
- Cross-Subject Mind Decoding from Inaccurate Representations
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition
- On Arbitrary Predictions from Equally Valid Models
- Comparative Analysis of Vision Transformers and Convolutional Neural Networks for Medical Image Classification
- VB-Mitigator: An Open-source Framework for Evaluating and Advancing Visual Bias Mitigation
- ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
- Celeb-DF++: A Large-scale Challenging Video DeepFake Benchmark for Generalizable Forensics
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- Improving Bird Classification with Primary Color Additives
- PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
- ACME: Adaptive Customization of Large Models via Distributed Systems
- X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
- BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
- Exp-Graph: How Connections Learn Facial Attributes in Graph-based Expression Recognition
- Comparative evaluation of deep learning architectures for bioclast classification in atomic force microscopy images
- Caching Techniques for Reducing the Communication Cost of Federated Learning in IoT Environments
- Design and Implementation of an Annotation-Driven Drone Autonomy Tool Using YOLOv8–V11 Architectures for Real-Time Object Detection and Distance Estimation
- DUSTrack: Semi-automated point tracking in ultrasound videos
- Rethinking Individual Fairness in Deepfake Detection
- UGPL: Uncertainty-Guided Progressive Learning for Evidence-Based Classification in Computed Tomography
- Interpretable weakly-supervised learning through kernel density matrices: A digital pathology use case
- Computer Vision for Real-Time Monkeypox Diagnosis on Embedded Systems
- Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts
- Improving U-Net Confidence on TEM Image Data with L2-Regularization, Transfer Learning, and Deep Fine-Tuning
- Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
- StarIO: A Lightweight Inertial Odometry for Nonlinear Motion
- Pixel Perfect MegaMed: A Megapixel-Scale Vision-Language Foundation Model for Generating High Resolution Medical Images
- SOD-YOLO: Enhancing YOLO-Based Detection of Small Objects in UAV Imagery
- DASViT: Differentiable Architecture Search for Vision Transformer
- Feature-Enhanced TResNet for Fine-Grained Food Image Classification
- Selective Quantization Tuning for ONNX Models
- Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation
- HomographyAD: Deep Anomaly Detection Using Self Homography Learning
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- 3D Magnetic Inverse Routine for Single-Segment Magnetic Field Images
- SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
- Jellyfish Species Identification: A CNN Based Artificial Neural Network Approach
- DepViT-CAD: Deployable Vision Transformer-Based Cancer Diagnosis in Histopathology
- Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?
- Biologically Inspired Deep Learning Approaches for Fetal Ultrasound Image Classification
- Spatial Lifting for Dense Prediction
- Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
- Recognizing Dementia from Neuropsychological Tests with State Space Models
- Winsor-CAM: Human-Tunable Visual Explanations from Deep Networks via Layer-Wise Winsorization
- Automatic Contouring of Spinal Vertebrae on X-Ray using a Novel Sandwich U-Net Architecture
- DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
- Multigranular Evaluation for Brain Visual Decoding
- Tree-Mamba: A Tree-Aware Mamba for Underwater Monocular Depth Estimation
- Where are we with calibration under dataset shift in image classification?
- Deep Brain Net: An Optimized Deep Learning Model for Brain tumor Detection in MRI Images Using EfficientNetB0 and ResNet50 with Transfer Learning
- A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
- Robust and Safe Traffic Sign Recognition using N-version with Weighted Voting
- GreenHyperSpectra: A multi-source hyperspectral dataset for global vegetation trait prediction
- Mask6D: Masked Pose Priors For 6D Object Pose Estimation
- Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
- Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset
- Zero-Shot Neural Architecture Search with Weighted Response Correlation
- Efficient SAR Vessel Detection for FPGA-Based On-Satellite Sensing
- X-ray transferable polyrepresentation learning
- TinyProto: Communication-Efficient Federated Learning with Sparse Prototypes in Resource-Constrained Environments
- Heterogeneous Federated Learning with Prototype Alignment and Upscaling
- An Explainable Transformer Model for Alzheimer's Disease Detection Using Retinal Imaging
- Siberian radioheliograph image classification using ensemble of CLIP, EfficientNet and CatBoost models
- OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
- Fair Deepfake Detectors Can Generalize
- Detection of Rail Line Track and Human Beings Near the Track to Avoid Accidents
- Age Sensitive Hippocampal Functional Connectivity: New Insights from 3D CNNs and Saliency Mapping
- evMLP: An Efficient Event-Driven MLP Architecture for Vision
- Multi Source COVID-19 Detection via Kernel-Density-based Slice Sampling
- Rapid Salient Object Detection with Difference Convolutional Neural Networks
- Instant Particle Size Distribution Measurement Using CNNs Trained on Synthetic Data
- Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms
- ADAptation: Reconstruction-based Unsupervised Active Learning for Breast Ultrasound Diagnosis
- Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding
- Room Scene Discovery and Grouping in Unstructured Vacation Rental Image Collections
- QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
- SoftStep: Learning Sparse Similarity Powers Deep Neighbor-Based Regression
- Concept-based Adversarial Attack: a Probabilistic Perspective
- From Large-scale Audio Tagging to Real-Time Explainable Emergency Vehicle Sirens Detection
- Low-latency vision transformers via large-scale multi-head attention
- Trident: Detecting Face Forgeries with Adversarial Triplet Learning
- FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
- CaO2: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
- SpikeSMOKE: Spiking Neural Networks for Monocular 3D Object Detection with Cross-Scale Gated Coding
- CAST: Cross-Attentive Spatio-Temporal feature fusion for Deepfake detection
- MedPrompt: LLM-CNN Fusion with Weight Routing for Medical Image Segmentation and Classification
- Tree-based Semantic Losses: Application to Sparsely-supervised Large Multi-class Hyperspectral Segmentation
- Pushing Trade-Off Boundaries: Compact yet Effective Remote Sensing Change Detection
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends
- Multi-model Online Conformal Prediction with Graph-Structured Feedback
- FastRef:Fast Prototype Refinement for Few-Shot Industrial Anomaly Detection
- Feature Hallucination for Self-supervised Action Recognition
- MEL: Multi-level Ensemble Learning for Resource-Constrained Environments
- Causal Representation Learning with Observational Grouping for CXR Classification
- FundaQ-8: A Clinically-Inspired Scoring Framework for Automated Fundus Image Quality Assessment
- IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
- Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
- PhishingHook: Catching Phishing Ethereum Smart Contracts leveraging EVM Opcodes
- Comparative Performance of Finetuned ImageNet Pre-trained Models for Electronic Component Classification
- ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments
- Classification of Tents in Street Bazaars Using CNN
- NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
- Reimagining Parameter Space Exploration with Diffusion Models
- SELFI: Selective Fusion of Identity for Generalizable Deepfake Detection
- EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
- AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
- Transition of AI Models in dependence of noise
- Noise-Informed Diffusion-Generated Image Detection with Anomaly Attention
- LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
- From Lab to Factory: Pitfalls and Guidelines for Self-/Unsupervised Defect Detection on Low-Quality Industrial Images
- TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
- Polyline Path Masked Attention for Vision Transformer
- Enhanced Dermatology Image Quality Assessment via Cross-Domain Training
- Classification of Multi-Parametric Body MRI Series Using Deep Learning
- Clustered Federated Learning via Embedding Distributions
- HiPreNets: High-Precision Neural Networks through Progressive Training
- FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
- DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning
- Vision Transformers for End-to-End Quark-Gluon Jet Classification from Calorimeter Images
- Advancing Image-Based Grapevine Variety Classification with a New Benchmark and Evaluation of Masked Autoencoders
- RBA-FE: A Robust Brain-Inspired Audio Feature Extractor for Depression Diagnosis
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- Active Adversarial Noise Suppression for Image Forgery Localization
- Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models
- GroupNL: Low-Resource and Robust CNN Design over Cloud and Device
- BePo: Dual Representation for 3D Occupancy Prediction
- Evaluating Fairness and Mitigating Bias in Machine Learning: A Novel Technique using Tensor Data and Bayesian Regression
- FAME: A Lightweight Spatio-Temporal Network for Model Attribution of Face-Swap Deepfakes
- TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
- Black-Box Edge AI Model Selection with Conformal Latency and Accuracy Guarantees
- DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
- Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation
- NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery
- DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
- Prompt-Guided Latent Diffusion with Predictive Class Conditioning for 3D Prostate MRI Generation
- HQFNN: A Compact Quantum-Fuzzy Neural Network for Accurate Image Classification
- KNN-Defense: Defense against 3D Adversarial Point Clouds using Nearest-Neighbor Search
- No TPU Left Behind: Retrofitting Side-Channel Protection into Edge TPUs
- An Efficient and Automated Classification System for Rocks Based on Visually Explainable Deep Learning
- DermaCon-IN: A Multi-concept Annotated Dermatological Image Dataset of Indian Skin Disorders for Clinical AI Research
- Tensor-to-Tensor Models with Fast Iterated Sum Features
- Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models
- Can Foundation Models Generalise the Presentation Attack Detection Capabilities on ID Cards?
- Practical Manipulation Model for Robust Deepfake Detection
- Recent Advances in Medical Image Classification
- A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
- Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
- Technology prediction of a 3D model using Neural Network
- FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review
- ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
- GlobalBuildingAtlas: An Open Global and Complete Dataset of Building Polygons, Heights and LoD1 3D Models
- From Pixels to Polygons: A Survey of Deep Learning Approaches for Medical Image-to-Mesh Reconstruction
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- Smartflow: Enabling Scalable Spatiotemporal Geospatial Research
- A Survey of Deep Learning Video Super-Resolution
- ConMamba: Contrastive Vision Mamba for Plant Disease Detection
- OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
- GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
- Enhancing Glass Defect Detection with Diffusion Models: Addressing Imbalanced Datasets in Manufacturing Quality Control
- A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
- Current AI technologies in cancer diagnostics and treatment
- The Promise of Spiking Neural Networks for Ubiquitous Computing: A Survey and New Perspectives
- Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
- Beyond Pixel Agreement: Large Language Models as Clinical Guardrails for Reliable Medical Image Segmentation
- A Large Convolutional Neural Network for Clinical Target and Multi-organ Segmentation in Gynecologic Brachytherapy with Multi-stage Learning
- Advancing automated identification of airborne fungal spores: guidelines for cultivation and reference dataset creation
- CAPAA: Classifier-Agnostic Projector-Based Adversarial Attack
- AWML: An Open-Source ML-based Robotics Perception Framework to Deploy for ROS-based Autonomous Driving Software
- Optimal Weighted Convolution for Classification and Denosing
- Optimal Density Functions for Weighted Convolution in Learning Models
- Benchmarking Foundation Models for Zero-Shot Biometric Tasks
- AnomalyMatch: Discovering Rare Objects of Interest with Semi-supervised and Active Learning
- RCCDA: Adaptive Model Updates in the Presence of Concept Drift under a Constrained Resource Budget
- SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
- Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch
- Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects
- Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
- LeMoRe: Learn More Details for Lightweight Semantic Segmentation
- AquaMonitor: A multimodal multi-view image sequence dataset for real-life aquatic invertebrate biodiversity monitoring
- Improving Respiratory Sound Classification with Architecture-Agnostic Knowledge Distillation from Ensembles
- YH-MINER: Multimodal Intelligent System for Natural Ecological Reef Metric Extraction
- Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
- Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
- Deep Learning-Based BMD Estimation from Radiographs with Conformal Uncertainty Quantification
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
- EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild
- Detecting Informative Channels: ActionFormer
- Intelligent Incident Hypertension Prediction in Obstructive Sleep Apnea
- Towards Cross-Domain Multi-Targeted Adversarial Attacks
- Supervised and Self-Supervised Land-Cover Segmentation & Classification of the Biesbosch Wetlands
- Frequency Composition for Compressed and Domain-Adaptive Neural Networks
- Knowledge Distillation Approach for SOS Fusion Staging: Towards Fully Automated Skeletal Maturity Assessment
- FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
- RoGA: Towards Generalizable Deepfake Detection through Robust Gradient Alignment
- HSplitLoRA: A Heterogeneous Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
- Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
- Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
- Connecting Independently Trained Modes via Layer-Wise Connectivity
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy
- Location-guided lesions representation learning via image generation for assessing plant leaf diseases severity
- Smart Waste Management System for Makkah City using Artificial Intelligence and Internet of Things
- Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
- Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
- A Smart Healthcare System for Monkeypox Skin Lesion Detection and Tracking
- Think Twice before Adaptation: Improving Adaptability of DeepFake Detection via Online Test-Time Adaptation
- Deep Learning for Breast Cancer Detection: Comparative Analysis of ConvNeXT and EfficientNet
- Asymmetric Duos: Sidekicks Improve Uncertainty
- Preserving AUC Fairness in Learning with Noisy Protected Groups
- Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling
- HyperFake: Hyperspectral Reconstruction and Attention-Guided Analysis for Advanced Deepfake Detection
- Clip4Retrofit: Enabling Real-Time Image Labeling on Edge Devices via Cross-Architecture CLIP Distillation
- SecondOpinion: Anatomy-Aware Gated Reasoning for Efficient Medical Image Analysis
- AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models
- VLM Models and Automated Grading of Atopic Dermatitis
- AdaForensics: Learning A Characteristic-aware Adaptive Deepfake Detector
- PawPrint: Whose Footprints Are These? Identifying Animal Individuals by Their Footprints
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- Taming Diffusion for Dataset Distillation with High Representativeness
- A Dataset and Benchmarks for Deep Learning-Based Optical Microrobot Pose and Depth Perception
- Extending Dataset Pruning to Object Detection: A Variance-based Approach
- Detailed Evaluation of Modern Machine Learning Approaches for Optic Plastics Sorting
- AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
- SuperPure: Efficient Purification of Localized and Distributed Adversarial Patches via Super-Resolution GAN Models
- HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification
- TRAIL: Transferable Robust Adversarial Images via Latent diffusion
- Do DeepFake Attribution Models Generalize?
- 15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning
- Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study
- BadSR: Stealthy Label Backdoor Attacks on Image Super-Resolution
- A Deep Learning Framework for Two-Dimensional, Multi-Frequency Propagation Factor Estimation
- CEBSNet: Change-Excited and Background-Suppressed Network with Temporal Dependency Modeling for Bitemporal Change Detection
- Oral Imaging for Malocclusion Issues Assessments: OMNI Dataset, Deep Learning Baselines and Benchmarking
- Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
- Securing Transfer-Learned Networks with Reverse Homomorphic Encryption
- Domain Adaptation for Multi-label Image Classification: a Discriminator-free Approach
- Automated Fetal Biometry Assessment with Deep Ensembles using Sparse-Sampling of 2D Intrapartum Ultrasound Images
- Automated Quality Evaluation of Cervical Cytopathology Whole Slide Images Based on Content Analysis
- AquaSignal: An Integrated Framework for Robust Underwater Acoustic Analysis
- Backward Conformal Prediction
- Deterministic Bounds and Random Estimates of Metric Tensors on Neuromanifolds
- Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
- BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
- Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts
- AGI-Elo: How Far Are We From Mastering A Task?
- FreqSelect: Frequency-Aware fMRI-to-Image Reconstruction
- UCBound-Net: Uncertainty-Guided Boundary-Aware Continual Learning for Domain-Incremental Ultrasound Segmentation
- Energy-Aware Deep Learning on Resource-Constrained Hardware
- Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
- Towards Open-world Generalized Deepfake Detection: General Feature Extraction via Unsupervised Domain Adaptation
- Model alignment using inter-modal bridges
- Equally Critical: Samples, Targets, and Their Mappings in Datasets
- Driver2Map: Imitating Human Driving for Online High-Definition Map Construction
- Semantically-Aware Game Image Quality Assessment
- msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyML
- ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation
- Understanding Nonlinear Implicit Bias via Region Counts in Input Space
- X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
- AW-GATCN: Adaptive Weighted Graph Attention Convolutional Network for Event Camera Data Joint Denoising and Object Recognition
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
- Defect Detection in Photolithographic Patterns Using Deep Learning Models Trained on Synthetic Data
- DeepSeqCoco: A Robust Mobile Friendly Deep Learning Model for Detection of Diseases in Cocos nucifera
- Demystifying AI Agents: The Final Generation of Intelligence
- From Preimage Search To Source-Grounded Feature Inversion
- Ensemble Deep Learning Approaches for AI-Altered Video Detection
- Hierarchical Surgical Robot Transformer (SRT-H): Imitation Learning for Autonomous Surgery
- DCSNet: A Lightweight Knowledge Distillation-Based Model with Explainable AI for Lung Cancer Diagnosis from Histopathological Images
- APR-Transformer: Initial Pose Estimation for Localization in Complex Environments through Absolute Pose Regression
- FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration
- Learning to Detect Multi-class Anomalies with Just One Normal Image Prompt
- Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation (CURVAS) challenge results
- Where the Devil Hides: Deepfake Detectors Can No Longer Be Trusted
- AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
- Validation of Conformal Prediction in Cervical Atypia Classification
- Adaptive Latent-Space Constraints in Personalized Federated Learning
- Security through the Eyes of AI: How Visualization is Shaping Malware Detection
- Differentiable NMS via Sinkhorn Matching for End-to-End Fabric Defect Detection
- Feature Representation Transferring to Lightweight Models via Perception Coherence
- Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification
- Unmasking Deep Fakes: Leveraging Deep Learning for Video Authenticity Detection
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
- Predicting Diabetic Macular Edema Treatment Responses Using OCT: Dataset and Methods of APTOS Competition
- Deep Learning-Based Robust Optical Guidance for Hypersonic Platforms
- Dual-Resolution Attention-Gated Deep Learning with Ordinal Regression for Diabetic Retinopathy Grading: A Quantified Assessment of Cross-Domain Generalization
- DFEN: Dual Feature Equalization Network for Medical Image Segmentation
- V-EfficientNets: Vector-Valued Efficiently Scaled Convolutional Neural Network Models
- Face-D(2)CL: Multi-Domain Synergistic Representation with Dual Continual Learning for Facial DeepFake Detection
- Cross-Branch Orthogonality for Improved Generalization in Face Deepfake Detection
- Altitude-Adaptive Vision-Only Geo-Localization for UAVs in GPS-Denied Environments
- Intelligent Diagnosis Using Dual-Branch Attention Network for Rare Thyroid Carcinoma Recognition with Ultrasound Imaging
- Local Herb Identification Using Transfer Learning: A CNN-Powered Mobile Application for Nepalese Flora
- OODTE: A Differential Testing Engine for the ONNX Optimizer
- Conformal Prediction for Indoor Positioning with Correctness Coverage Guarantees
- Efficient Multi Subject Visual Reconstruction from fMRI Using Aligned Representations
- Adaptively Point-weighting Curriculum Learning
- AI-driven multi-source data fusion for algal bloom severity classification in small inland water bodies: Leveraging Sentinel-2, DEM, and NOAA climate data
- One Search Fits All: Pareto-Optimal Eco-Friendly Model Selection
- Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- AHC: Meta-Learned Adaptive Compression for Continual Object Detection on Memory-Constrained Microcontrollers
- A Framework for Generating Semantically Ambiguous Images to Probe Human and Machine Perception
- MalariAI: A Label-Resilient Decoupled Framework for Annotation-Agnostic Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
- MetaPerch: Learning from metadata for bioacoustics foundation models
- Energy-Efficient Neuromorphic Computing for Edge AI: A Framework with Adaptive Spiking Neural Networks and Hardware-Aware Optimization
- Deep Learning Study of Alkaptonuria Spinal Disease Assesses Global and Regional Severity and Detects Occult Treatment Status
- Automatic detection of fin, operculum and skin deformities in Mediterranean Fish Species
- Combating Falsification of Speech Videos with Live Optical Signatures (Extended Version)
- Resource-Efficient Gesture Recognition through Convexified Attention
- Vision Transformers in Precision Agriculture: A Comprehensive Survey
- A Test Suite for Efficient Robustness Evaluation of Face Recognition Systems
- Efficient Detection and Characterization of Targets of Natural Selection Using Transfer Learning
- A simple and effective approach for body part recognition on CT scans based on projection estimation
- Lightweight and Fast Backdoor Model Detection
- MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
- DSFusionNet: Dynamic Dual-Stream Fusion with Bidirectional Knowledge Distillation for Plant Disease Recognition
- Occlusion-aware Driver Monitoring System using the Driver Monitoring Dataset
- SCOPE-MRI: Bankart Lesion Detection as a Case Study in Data Curation and Deep Learning for Challenging Diagnoses
- SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
- Leveraging Neural Graph Compilers in Machine Learning Research for Edge-Cloud Systems
- AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management
- HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps
- WILD: a new in-the-Wild Image Linkage Dataset for synthetic image attribution
- CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes
- AIBuildAI: An AI Agent for Automatically Building AI Models
- Examining the Impact of Optical Aberrations to Image Classification and Object Detection Models
- Revisiting Data Auditing in Large Vision-Language Models
- DNAD: Differentiable Neural Architecture Distillation
- Classification of SARS-CoV-2 Variants through The Epistatical Circos Plots with Convolutional Neural Networks
- Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion
- Indoor Occupancy Classification using a Compact Hybrid Quantum-Classical Model Enabled by a Physics-Informed Radar Digital Twin
- Integrating APK Image and Text Data for Enhanced Threat Detection: A Multimodal Deep Learning Approach to Android Malware
- A Decade of You Only Look Once (YOLO) for Object Detection: A Review
- OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
- Integrating Multi-Armed Bandit, Active Learning, and Distributed Computing for Scalable Optimization
- RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
- OUI Need to Talk About Weight Decay: A New Perspective on Overfitting Detection
- Coding for Computation: Efficient Compression of Neural Networks for Reconfigurable Hardware
- 4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
- Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
- StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning
- SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets
- Noise-Tolerant Coreset-Based Class Incremental Continual Learning
- Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
- MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
- FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
- An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon
- Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
- MedNNS: Supernet-based Medical Task-Adaptive Neural Network Search
- ForgeBench: A Machine Learning Benchmark Suite and Auto-Generation Framework for Next-Generation HLS Tools
- VeLU: Variance-enhanced Learning Unit for Deep Neural Networks
- ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
- Efficient multiscale feature integration network for lightweight remote sensing images change detection
- Hybrid Knowledge Transfer through Attention and Logit Distillation for On-Device Vision Systems in Agricultural IoT
- A Bayesian Approach to Segmentation with Noisy Labels via Spatially Correlated Distributions
- Application of U-net models in estimating forest canopy closure based on multi-source remote sensing imagery
- RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
- Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- Revisiting CLIP for SF-OSDA: Unleashing Zero-Shot Potential with Adaptive Threshold and Training-Free Feature Filtering
- Segregation and Context Aggregation Network for Real-time Cloud Segmentation
- ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification
- Unreal Robotics Lab: A High-Fidelity Robotics Simulator with Advanced Physics and Rendering
- Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
- Scaling Laws for Data-Efficient Visual Transfer Learning
- Can Masked Autoencoders Also Listen to Birds?
- Artificial intelligence-enabled non-invasive cataract diagnosis and grading system using anterior segment images
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- Early Accessibility: Automating Alt-Text Generation for UI Icons During App Development
- Comparative Evaluation of Radiomics and Deep Learning Models for Disease Detection in Chest Radiography
- BEV-GS: Feed-forward Gaussian Splatting in Bird's-Eye-View for Road Reconstruction
- Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
- ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection
- Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation
- MSA-DCNN: A Data-Efficient Multi-Scale Attention Deformable CNN for Medical Image Classification
- EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection
- Topological shape transform for thymus structures
- Masked Autoencoder Self Pre-Training for Defect Detection in Microelectronics
- GFT: Gradient Focal Transformer
- Dual-Path Enhancements in Event-Based Eye Tracking: Augmented Robustness and Adaptive Temporal Modeling
- Semantic Depth Matters: Explaining Errors of Deep Vision Networks through Perceived Class Similarities
- FROG: Effective Friend Recommendation in Online Games via Modality-aware User Preferences
- FastRSR: Efficient and Accurate Road Surface Reconstruction from Bird's Eye View
- A review on deep learning for vision-based hand detection, hand segmentation and hand gesture recognition in human–robot interaction
- ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps
- Exploring Synergistic Ensemble Learning: Uniting CNNs, MLP-Mixers, and Vision Transformers to Enhance Image Classification
- On Background Bias of Post-Hoc Concept Embeddings in Computer Vision DNNs
- InSPE: Rapid Evaluation of Heterogeneous Multi-Modal Infrastructure Sensor Placement
- DRIP: DRop unImportant data Points -- Enhancing Machine Learning Efficiency with Grad-CAM-Based Real-Time Data Prioritization for On-Device Training
- Benchmarking Image Embeddings for E-Commerce: Evaluating Off-the Shelf Foundation Models, Fine-Tuning Strategies and Practical Trade-offs
- Teaching Humans Subtle Differences with DIFFusion
- S-EO: A Large-Scale Dataset for Geometry-Aware Shadow Detection in Remote Sensing Applications
- Compound and Parallel Modes of Tropical Convolutional Neural Networks
- DefMamba: Deformable Visual State Space Model
- InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
- Memory-Modular Classification: Learning to Generalize with Memory Replacement
- Vibe360+: An extended study on group-level deep emotional understanding in immersive communication
- Find A Winning Sign: Sign Is All We Need to Win the Lottery
- From Specificity to Generality: Revisiting Generalizable Artifacts in Detecting Face Deepfakes
- Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
- Contrastive Language–Image Pre-training [wikipedia]
- EfficientNet [wikipedia]
Discussions
Related