EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
2019/05/28 by Mingxing Tan, Quoc V. Le · 1 voice · 512 citations
#cs.LG #cs.CV #stat.ML
paper · pdf
Abstract
Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available. In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance. Based on this observation, we propose a new scaling method that uniformly scales all dimensions of depth/width/resolution using a simple yet highly effective compound coefficient. We demonstrate the effectiveness of this method on scaling up MobileNets and ResNet. To go even further, we use neural architecture search to design a new baseline network and scale it up to obtain a family of models, called EfficientNets, which achieve much better accuracy and efficiency than previous ConvNets. In particular, our EfficientNet-B7 achieves state-of-the-art 84.3% top-1 accuracy on ImageNet, while being 8.4x smaller and 6.1x faster on inference than the best existing ConvNet. Our EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 (91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer parameters. Source code is at https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet.
Cited by
- A hybrid CNN-machine learning model for anxiety screening with mandala coloring patterns: Feature extraction and classification performance
- Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
- LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models
- ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors
- A Statistical Multi-Objective Framework for Assessing Sensitivity of Radiomic AI Models to Acquisition Parameters
- Traceback Translators Against Forgetting in Continual Fake Speech Detection
- Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
- WiFi Sensing via Reservoir Computing
- Pixel-Space Diffusion Transformers
- SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
- One Round Is All You Need: Analytic Federated Learning for Task-Heterogeneous Multi-Label Medical Image Classification
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- Cross-Dataset Generalization in Breast MRI Tumor Classification via Class-Wise Dataset Mixing
- Privileged Lesion-Context Relational Distillation for Mask-Free Skin Lesion Classification
- CITRUS: Candidate Inference and Temporal-tracking for Reliable, Unobtrusive Sensing of Wearable Heart Rate under Motion
- Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator
- TinyGLASS: Real-Time Self-Supervised In-Sensor Anomaly Detection
- When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection
- DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
- EA-RMENet -- Path Loss Prediction in Urban Environments using Deep Learning
- AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis
- Improving Backward Conformal Prediction via Non-Conformity Score Transformation
- When Bigger is Worse: A Practitioner's Guide to Model Selection Under Data Scarcity
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
- Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study
- DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- OvAi Focus: AI-based Multi-class Segmentation of Functional Ovaries and Adnexal Masses in Gynecological Ultrasound
- Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
- Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- Fit for Purpose? Deepfake Detection in the Real World
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- Deep learning approach to detect and visualize sexual dimorphism in monomorphic species
- Perception Encoder: The best visual embeddings are not at the output of the network
- Accurate phenotyping of luminal A breast cancer in magnetic resonance imaging: A new 3D CNN approach
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- ILIAS: Instance-Level Image retrieval At Scale
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
- Exploring Syn-to-Real Domain Adaptation for Military Target Detection
- Real-Time In-Cabin Driver Behavior Recognition on Low-Cost Edge Hardware
- Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
- Balanced Soft mixture-of-expert model for Glaucoma Detection
- Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
- Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution
- Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
- Real-time Reconstruction of Human Visual Perception from fMRI
- Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- Visual information extraction from documents via classification-guided large vision-language models
- BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving
- Patch-Discontinuity Mining for Generalized Deepfake Detection
- AI for Mycetoma Diagnosis in Histopathological Images: The MICCAI 2024 Challenge
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- MODE: Multi-Objective Adaptive Coreset Selection
- Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- MedNeXt-v2: Scaling 3D ConvNeXts for Large-Scale Supervised Representation Learning in Medical Image Segmentation
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Generating Risky Samples with Conformity Constraints via Diffusion Models
- MeniMV: A Multi-view Benchmark for Meniscus Injury Severity Grading
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
- Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
- Domain-Aware Quantum Circuit for QML
- Training Together, Diagnosing Better: Federated Learning for Collagen VI-Related Dystrophies
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- Exploring Deep-to-Shallow Transformable Neural Networks for Intelligent Embedded Systems
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
- MFE-GAN: Efficient GAN-based Framework for Document Image Enhancement and Binarization with Multi-scale Feature Extraction
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- DL3M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- ACCOR: Attention-Enhanced Complex-Valued Contrastive Learning for Occluded Object Classification Using mmWave Radar IQ Signals
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Sequence-to-Image Transformation for Sequence Classification Using Rips Complex Construction and Chaos Game Representation
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- NeuroSketch: An Effective Framework for Neural Decoding via Systematic Architectural Optimization
- MelanomaNet: Explainable Deep Learning for Skin Lesion Classification
- FBA2D: Frequency-based Black-box Attack for AI-generated Image Detection
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Data-Efficient Learning of Anomalous Diffusion with Wavelet Representations: Enabling Direct Learning from Experimental Trajectories
- Animal Re-Identification on Microcontrollers
- VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
- UltrasODM: A Dual Stream Optical Flow Mamba Network for 3D Freehand Ultrasound Reconstruction
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- Reevaluating Automated Wildlife Species Detection: A Reproducibility Study on a Custom Image Dataset
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
- Leveraging Pre-trained Neural Network Models for the Classification of Tumor Cells Analyzed by Label-free Phase Holotomographic Microscopy
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Hierarchical Deep Learning for Diatom Image Classification: A Multi-Level Taxonomic Approach
- Phase-OTDR Event Detection Using Image-Based Data Transformation and Deep Learning
- Automated Plant Disease and Pest Detection System Using Hybrid Lightweight CNN-MobileViT Models for Diagnosis of Indigenous Crops
- Invariance Co-training for Robot Visual Generalization
- DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification
- Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World
- Open Set Face Forgery Detection via Dual-Level Evidence Collection
- EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification
- Defense That Attacks: How Robust Models Become Better Attackers
- Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
- Breast Cell Segmentation Under Extreme Data Constraints: Quantum Enhancement Meets Adaptive Loss Stabilization
- Data-Centric Visual Development for Self-Driving Labs
- nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis
- M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
- Implementation of a Skin Lesion Detection System for Managing Children with Atopic Dermatitis Based on Ensemble Learning
- ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
- Stacked Ensemble of Fine-Tuned CNNs for Knee Osteoarthritis Severity Grading
- Towards 3D Object-Centric Feature Learning for Semantic Scene Completion
- TeleViT1.0: Teleconnection-aware Vision Transformers for Subseasonal to Seasonal Wildfire Pattern Forecasts
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- Transformer Driven Visual Servoing for Fabric Texture Matching Using Dual-Arm Manipulator
- RefOnce: Distilling References into a Prototype Memory for Referring Camouflaged Object Detection
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- CellFMCount: A Fluorescence Microscopy Dataset, Benchmark, and Methods for Cell Counting
- SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
- Dendritic Convolution for Noise Image Recognition
- Shape-Adapting Gated Experts: Dynamic Expert Routing for Colonoscopic Lesion Segmentation
- LungX: A Hybrid EfficientNet-Vision Transformer Architecture with Multi-Scale Attention for Accurate Pneumonia Detection
- Stro-VIGRU: Defining the Vision Recurrent-Based Baseline Model for Brain Stroke Classification
- Compact neural networks for astronomy with optimal transport bias correction
- A Lightweight, Interpretable Deep Learning System for Automated Detection of Cervical Adenocarcinoma In Situ (AIS)
- Equivariant-Aware Structured Pruning for Efficient Edge Deployment: A Comprehensive Framework with Adaptive Fine-Tuning
- Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
- A Diversity-optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
- Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
- Memory-DD: A Low-Complexity Dendrite-Inspired Neuron for Temporal Prediction Tasks
- Externally Validated Multi-Task Learning via Consistency Regularization Using Differentiable BI-RADS Features for Breast Ultrasound Tumor Segmentation
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
- Monocular 3D Lane Detection via Structure Uncertainty-Aware Network with Curve-Point Queries
- H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
- Online Continual Learning on Intel Loihi 2 via a Co-designed Spiking Neural Network
- MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
- SAGE: Saliency-Guided Contrastive Embeddings
- Floor Plan-Guided Visual Navigation Incorporating Depth and Directional Cues
- Lightweight Optimal-Transport Harmonization on Edge Devices
- Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
- TEMPO: Global Temporal Building Density and Height Estimation from Satellite Imagery
- MFI-ResNet: Efficient ResNet Architecture Optimization via MeanFlow Compression and Selective Incubation
- SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
- Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
- D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces
- Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts
- DermAI: Clinical dermatology acquisition through quality-driven image collection for AI classification in mobile
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
- Federated Learning for Pediatric Pneumonia Detection: Enabling Collaborative Diagnosis Without Sharing Patient Data
- REASON: Probability map-guided dual-branch fusion framework for gastric content assessment
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- A vascular code for speed in the spatial navigation system
- Rethinking Explanation Evaluation under the Retraining Scheme
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- LatentPrintFormer: A Hybrid CNN-Transformer with Spatial Attention for Latent Fingerprint identification
- SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
- Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- Physics-Informed Image Restoration via Progressive PDE Integration
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- NeuroFlex: Column-Exact ANN-SNN Co-Execution Accelerator with Cost-Guided Scheduling
- Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
- A Lightweight 3D-CNN for Event-Based Human Action Recognition with Privacy-Preserving Potential
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- FastBoost: Progressive Attention with Dynamic Scaling for Efficient Deep Learning
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Weakly Supervised Pneumonia Localization from Chest X-Rays Using Deep Neural Network and Grad-CAM Explanations
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- How Close Are We? Limitations and Progress of AI Models in Banff Lesion Scoring
- ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding
- Surpassing state of the art on AMD area estimation from RGB fundus images through careful selection of U-Net architectures and loss functions for class imbalance
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
- CNN-Based Reanalysis of Optical Turbulence at the Canary Islands Observatories (OCAN)
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
- HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Selective Diabetic Retinopathy Screening with Accuracy-Weighted Deep Ensembles and Entropy-Guided Abstention
- Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- Fair Indivisible Payoffs through Shapley Value
- A Dual-Branch CNN for Robust Detection of AI-Generated Facial Forgeries
- All in one timestep: Enhancing Sparsity and Energy efficiency in Multi-level Spiking Neural Networks
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Training-free Source Attribution of AI-generated Images via Resynthesis
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- PLDC‐Net: A Domain‐Specific Base Model for Plant Leaf Disease Classification Domain Adaptation Tasks
- Does Machine Learning Work? A Comparative Analysis of Strong Gravitational Lens Searches in the Dark Energy Survey
- Explainable Detection of AI-Generated Images with Artifact Localization Using Faster-Than-Lies and Vision-Language Models for Edge Devices
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- Synthetic-to-Real Transfer Learning for Chromatin-Sensitive PWS Microscopy
- DiffusionLane: Diffusion Model for Lane Detection
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- HybridSOMSpikeNet: A Deep Model with Differentiable Soft Self-Organizing Maps and Spiking Dynamics for Waste Classification
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Balanced Multi-Task Attention for Satellite Image Classification: A Systematic Approach to Achieving 97.23% Accuracy on EuroSAT Without Pre-Training
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Knowledge-Informed Neural Network for Complex-Valued SAR Image Recognition
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
- BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- TinyUSFM: Towards Compact and Efficient Ultrasound Foundation Models
- Enhancing Early Alzheimer Disease Detection through Big Data and Ensemble Few-Shot Learning
- Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation
- OmniNWM: Omniscient Driving Navigation World Models
- Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
- ViBED-Net: Video Based Engagement Detection Network Using Face-Aware and Scene-Aware Spatiotemporal Cues
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Fair and Interpretable Deepfake Detection in Videos
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- Learning to play: A Multimodal Agent for 3D Game-Play
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- Bridging Symmetry and Robustness: On the Role of Equivariance in Enhancing Adversarial Robustness
- A Multi-Task Deep Learning Framework for Skin Lesion Classification, ABCDE Feature Quantification, and Evolution Simulation
- BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection
- PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Removing Cost Volumes from Optical Flow Estimators
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
- Lightweight CNN-Based Wi-Fi Intrusion Detection Using 2D Traffic Representations
- Efficient Edge Test-Time Adaptation via Latent Feature Coordinate Correction
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
- Learning Model Representations Using Publicly Available Model Hubs
- HccePose(BF): Predicting Front & Back Surfaces to Construct Ultra-Dense 2D-3D Correspondences for Pose Estimation
- KTBox: A Modular LaTeX Framework for Semantic Color, Structured Highlighting, and Scholarly Communication
- Exploration of Incremental Synthetic Non-Morphed Images for Single Morphing Attack Detection
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- Structured Output Regularization: a framework for few-shot transfer learning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Weights initialization of neural networks for function approximation
- Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection
- Bridged Clustering: Semi-Supervised Sparse Bridging
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Image-based recognition using advanced neural networks can aid surveillance of Agrilus jewel beetles
- Clinical-grade AI model for molecular subtyping of endometrial cancer: a multi-center cohort study in China
- Mangrove3D: Terrestrial Laser Scanning Dataset for Coastal Mangrove Forests
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- Scalable deep fusion of spaceborne lidar and synthetic aperture radar for global forest structural complexity mapping
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Adaptive Coverage Policies in Conformal Prediction
- Detection of retinal diseases using an accelerated reused convolutional network
- Beyond Static Knowledge Messengers: Towards Adaptive, Fair, and Scalable Federated Learning for Medical AI
- Talking Tennis: Language Feedback from 3D Biomechanical Action Recognition
- CVSM: Contrastive Vocal Similarity Modeling
- ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation Testing
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
- A Fast and Precise Method for Searching Rectangular Tumor Regions in Brain MR Images
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
- Confidence Aware SSD Ensemble with Weighted Boxes Fusion for Weapon Detection
- On The Variability of Concept Activation Vectors
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning
- Evaluating the Impact of Radiographic Noise on Chest X-ray Semantic Segmentation and Disease Classification Using a Scalable Noise Injection Framework
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- Calibrated and Resource-Aware Super-Resolution for Reliable Driver Behavior Analysis
- S3F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- Hemorica: A Comprehensive CT Scan Dataset for Automated Brain Hemorrhage Classification, Segmentation, and Detection
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- Multilingual Vision-Language Models, A Survey
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- No-Reference Image Contrast Assessment with Customized EfficientNet-B0
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions in Semi-Supervised Learning
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
- EnGraf-Net: Multiple Granularity Branch Network with Fine-Coarse Graft Grained for Classification Task
- SD-RetinaNet: Topologically Constrained Semi-Supervised Retinal Lesion and Layer Segmentation in OCT
- RAM-NAS: Resource-aware Multiobjective Neural Architecture Search Method for Robot Vision Tasks
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
- Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
- Affective Computing and Emotional Data: Challenges and Implications in Privacy Regulations, The AI Act, and Ethics in Large Language Models
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
- DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
- DL-QC-fNIRS: a deep learning tool for automated quality control in functional near-infrared spectroscopy signals
- A typology for visual cues delimiting growth ring boundaries and a deep learning model to detect them in macroscopic images of softwoods
- Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
- Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data
- MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
- TS-P2CL: Plug-and-Play Dual Contrastive Learning for Vision-Guided Medical Time Series Classification
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
- MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
- V-CECE: Visual Counterfactual Explanations via Conceptual Edits
- Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- PM25Vision: A Large-Scale Benchmark Dataset for Visual Estimation of Air Quality
- Benchmarking Class Activation Map Methods for Explainable Brain Hemorrhage Classification on Hemorica Dataset
- Efficient Conformal Prediction for Regression Models under Label Noise
- NeRF-based Visualization of 3D Cues Supporting Data-Driven Spacecraft Pose Estimation
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
- MetricNet: Recovering Metric Scale in Generative Navigation Policies
- Deep Lookup Network
- Real-Time Detection and Tracking of Foreign Object Intrusions in Power Systems via Feature-Based Edge Intelligence
- Performance is not All You Need: Sustainability Considerations for Algorithms
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Human + AI for Accelerating Ad Localization Evaluation
- TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images
- T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
- Automated Landfill Detection Using Deep Learning: A Comparative Study of Lightweight and Custom Architectures with the AerialWaste Dataset
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- The Quest for Universal Master Key Filters in DS-CNNs
- Optimizing Class Distributions for Bias-Aware Multi-Class Learning
- Synthetic vs. Real Training Data for Visual Navigation
- Neural networks in the search for fast radio bursts with RATAN-600
- GraphDerm: Fusing Imaging, Physical Scale, and Metadata in a Population-Graph Classifier for Dermoscopic Lesions
- Hybrid Quantum Neural Networks for Efficient Protein-Ligand Binding Affinity Prediction
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
- An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift
- TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- Tri-Accel: Curvature-Aware Precision-Adaptive and Memory-Elastic Optimization for Efficient GPU Usage
- A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- Compressing CNN models for resource-constrained systems by channel and layer pruning
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- MedicalPatchNet: A Patch-Based Self-Explainable AI Architecture for Chest X-ray Classification
- EfficientNet in Digital Twin-based Cardiac Arrest Prediction and Analysis
- Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
- Improved Classification of Nitrogen Stress Severity in Plants Under Combined Stress Conditions Using Spatio-Temporal Deep Learning Framework
- NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter Editing
- Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
- Realism to Deception: Investigating Deepfake Detectors Against Face Enhancement
- Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
- Khana: A Comprehensive Indian Cuisine Dataset
- Challenges in Deep Learning-Based Small Organ Segmentation: A Benchmarking Perspective for Medical Research with Limited Datasets
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Interpretable Deep Transfer Learning for Breast Ultrasound Cancer Detection: A Multi-Dataset Study
- MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment using UAV Imagery
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
- Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- PracMHBench: Re-evaluating Model-Heterogeneous Federated Learning Based on Practical Edge Device Constraints
- Chest X-ray Pneumothorax Segmentation Using EfficientNet-B4 Transfer Learning in a U-Net Architecture
- Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- MedLiteNet: Lightweight Hybrid Medical Image Segmentation Model
- Prospects for acoustically monitoring ecosystem tipping points
- Vision-Based Embedded System for Noncontact Monitoring of Preterm Infant Behavior in Low-Resource Care Settings
- An Investigation of Visual Foundation Models Robustness
- An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
- Enhancing Fitness Movement Recognition with Attention Mechanism and Pre-Trained Feature Extractors
- Fair Resource Allocation for Fleet Intelligence
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
- AgroSense: An Integrated Deep Learning System for Crop Recommendation via Soil Image Analysis and Nutrient Profiling
- Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
- PrediTree: A Multi-Temporal Sub-meter Dataset of Multi-Spectral Imagery Aligned With Canopy Height Maps
- AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef
- AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
- Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks
- Panoptic Segmentation of Environmental UAV Images : Litter Beach
- ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- diveXplore at the Video Browser Showdown 2024
- Unified Acoustic Representations for Screening Neurological and Respiratory Pathologies from Voice
- ATMS-KD: Adaptive Temperature and Mixed Sample Knowledge Distillation for a Lightweight Residual CNN in Agricultural Embedded Systems
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- Paired-Sampling Contrastive Framework for Joint Physical-Digital Face Attack Detection
- Computing-In-Memory Dataflow for Minimal Buffer Traffic
- Formal Algorithms for Model Efficiency
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Unleashing Semantic and Geometric Priors for 3D Scene Completion
- Pixels Under Pressure: Exploring Fine-Tuning Paradigms for Foundation Models in High-Resolution Medical Imaging
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
- Skin Cancer Classification: Hybrid CNN-Transformer Models with KAN-Based Fusion
- OrbitChain: Orchestrating In-orbit Real-time Analytics of Earth Observation Data
- An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Dual-species atomic absorption image reconstruction using deep neural networks
- MultiPark: Multimodal Parking Transformer with Next-Segment Prediction
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- Scalable Geospatial Data Generation Using AlphaEarth Foundations Model
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- Biasing Frontier-Based Exploration with Saliency Areas
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- T-CACE: A Time-Conditioned Autoregressive Contrast Enhancement Multi-Task Framework for Contrast-Free Liver MRI Synthesis, Segmentation, and Diagnosis
- Predictive Uncertainty for Runtime Assurance of a Real-Time Computer Vision-Based Landing System
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- Large-Small Model Collaborative Framework for Federated Continual Learning
- Deep Learning for Automated Identification of Vietnamese Timber Species: A Tool for Ecological Monitoring and Conservation
- Autonomous AI Bird Feeder for Backyard Biodiversity Monitoring
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Low-Regret and Low-Complexity Learning for Hierarchical Inference
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Towards Human-AI Collaboration System for the Detection of Invasive Ductal Carcinoma in Histopathology Images
- MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training
- From Field to Drone: Domain Drift Tolerant Automated Multi-Species and Damage Plant Semantic Segmentation for Herbicide Trials
- Neural Tangent Knowledge Distillation for Optical Convolutional Networks
- DragonFruitQualityNet: A Lightweight Convolutional Neural Network for Real-Time Dragon Fruit Quality Inspection on Mobile Devices
- Position: Ideas Should be the Center of Machine Learning Research
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Don't Reach for the Stars: Rethinking Topology for Resilient Federated Learning
- ULU: A Unified Activation Function
- Perch 2.0: The Bittern Lesson for Bioacoustics
- Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
- Improving Tactile Gesture Recognition with Optical Flow
- Automated ultrasound doppler angle estimation using deep learning
- Energy-Efficient Real-Time 4-Stage Sleep Classification at 10-Second Resolution: A Comprehensive Study
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- FeDaL: Federated Dataset Learning for Time Series Foundation Models
- Uncertainty-aware Accurate Elevation Modeling for Off-road Navigation via Neural Processes
- FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose Estimation
- DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
- Inductive transfer learning from regression to classification in ECG analysis
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs
- Transfer Learning with EfficientNet for Accurate Leukemia Cell Classification
- Rate-distortion Optimized Point Cloud Preprocessing for Geometry-based Point Cloud Compression
- Benchmarking Adversarial Patch Selection and Location
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- ESM: A Framework for Building Effective Surrogate Models for Hardware-Aware Neural Architecture Search
- Foundation Models for Bioacoustics -- a Comparative Review
- Deep Learning for Pavement Condition Evaluation Using Satellite Imagery
- Classification of Brain Tumors using Hybrid Deep Learning Models
- Rethinking Backbone Design for Lightweight 3D Object Detection in LiDAR
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
- Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings
- Robust Deepfake Detection for Electronic Know Your Customer Systems Using Registered Images
- Cyst-X: A Federated AI System Outperforms Clinical Guidelines to Detect Pancreatic Cancer Precursors and Reduce Unnecessary Surgery
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Evaluating Deepfake Detectors in the Wild
- Contrastive Language–Image Pre-training [wikipedia]
- EfficientNet [wikipedia]
Discussions
Related