Rethinking the Inception Architecture for Computer Vision
2015/12/02 by Christian Szegedy, Vincent Vanhoucke, Szegedy, Christian +7 · 356 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications
paper · pdf · doi:10.48550/arxiv.1512.00567
Abstract
Convolutional networks are at the core of most state-of-the-art computer vision solutions for a wide variety of tasks. Since 2014 very deep convolutional networks started to become mainstream, yielding substantial gains in various benchmarks. Although increased model size and computational cost tend to translate to immediate quality gains for most tasks (as long as enough labeled data is provided for training), computational efficiency and low parameter count are still enabling factors for various use cases such as mobile vision and big-data scenarios. Here we explore ways to scale up networks in ways that aim at utilizing the added computation as efficiently as possible by suitably factorized convolutions and aggressive regularization. We benchmark our methods on the ILSVRC 2012 classification challenge validation set demonstrate substantial gains over the state of the art: 21.2% top-1 and 5.6% top-5 error for single frame evaluation using a network with a computational cost of 5 billion multiply-adds per inference and with using less than 25 million parameters. With an ensemble of 4 models and multi-crop evaluation, we report 3.5% top-5 error on the validation set (3.6% error on the test set) and 17.3% top-1 error on the validation set.
Cited by
- A Neural Network-Based Real-time Casing Collar Recognition System for Downhole Instruments
- Investigating Deep Learning Models for Ejection Fraction Estimation from Echocardiography Videos
- AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
- ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification
- Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- GSpyNetTree-O4: an event validation tool used in the fourth LIGO-Virgo-KAGRA observing run
- OpenPVMapper: A Multi-source, Nationwide Database of Rooftop Photovoltaic Systems in France
- Real-time Reconstruction of Human Visual Perception from fMRI
- Stacked Intelligent Metasurface-Aided Wave-Domain Signal Processing: From Communications to Sensing and Computing
- Benchmarking deep learning models for Raman spectroscopy across open-source datasets
- Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
- Defending against adversarial attacks using mixture of experts
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Localising Shortcut Learning in Pixel Space via Ordinal Scoring Correlations for Attribution Representations (OSCAR)
- Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- A two-stream network with global-local feature fusion for bone age assessment
- A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Interpretable Similarity of Synthetic Image Utility
- Next-Embedding Prediction Makes Strong Vision Learners
- Evaluation of deep learning architectures for wildlife object detection: A comparative study of ResNet and Inception
- Stylized Synthetic Augmentation further improves Corruption Robustness
- Packed Malware Detection Using Grayscale Binary-to-Image Representations
- SemanticBridge - A Dataset for 3D Semantic Segmentation of Bridges and Domain Gap Analysis
- BarcodeMamba+: Advancing State-Space Models for Fungal Biodiversity Research
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- LCMem: A Universal Model for Robust Image Memorization Detection
- Image Diffusion Preview with Consistency Solver
- Learning to Retrieve with Weakened Labels: Robust Training under Label Noise
- Improving the Plausibility of Pressure Distributions Synthesized from Depth Image through Generative Modeling
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- DL3M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- Maritime object classification with SAR imagery using quantum kernel methods
- Bidirectional Normalizing Flow: From Data to Noise and Back
- Stronger Normalization-Free Transformers
- DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
- Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
- CytoDINO: Risk-Aware and Biologically-Informed Adaptation of DINOv3 for Bone Marrow Cytomorphology
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- Identification of Deforestation Areas in the Amazon Rainforest Using Change Detection Models
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- LogicCBMs: Logic-Enhanced Concept-Based Learning
- Synchrony-Gated Plasticity with Dopamine Modulation for Spiking Neural Networks
- Phase-OTDR Event Detection Using Image-Based Data Transformation and Deep Learning
- Wasserstein distance based semi-supervised manifold learning and application to GNSS multi-path detection
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
- Performance Evaluation of Transfer Learning Based Medical Image Classification Techniques for Disease Detection
- Studying Various Activation Functions and Non-IID Data for Machine Learning Model Robustness
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Retrofitting Earth System Models with Cadence-Limited Neural Operator Updates
- Directed evolution algorithm drives neural prediction
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- ForamDeepSlice: A High-Accuracy Deep Learning Framework for Foraminifera Species Classification from 2D Micro-CT Slices
- Adversarial Flow Models
- Semantic-Aware Caching for Efficient Image Generation in Edge Computing
- Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
- Self-Paced Learning for Images of Antinuclear Antibodies
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
- MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
- FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- When +1% Is Not Enough: A Paired Bootstrap Protocol for Evaluating Small Improvements
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- Cross-Domain Generalization of Multimodal LLMs for Global Photovoltaic Assessment
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- Membership Inference Attacks Beyond Overfitting
- Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation
- SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness
- ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model
- Efficiently Training A Flat Neural Network Before It has been Quantizated
- Learning Feature Pyramids for Human Pose Estimation
- Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis
- A comparison of deep machine learning algorithms in COVID-19 disease\n diagnosis
- Multiscale Vision Transformers
- Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- Enhancing Photon Identification with Neural Network Methods
- Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation
- Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
- Query-Efficient Black-box Adversarial Examples (superceded)
- Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning
- An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- A New Perspective on Precision and Recall for Generative Models
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- Sewer pipeline condition assessment and defect detection using computer vision
- Melanoma Classification Through Deep Ensemble Learning and Explainable AI
- Data-Augmented Deep Learning for Downhole Depth Sensing and Field Validation
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- ZEBRA: Towards Zero-Shot Cross-Subject Generalization for Universal Brain Visual Decoding
- Simplex-to-Euclidean Bijections for Categorical Flow Matching
- Merlin L48 Spectrogram Dataset
- EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
- Generative Image Restoration and Super-Resolution using Physics-Informed Synthetic Data for Scanning Tunneling Microscopy
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Face De-Identification: A Domain-Centric Survey from Capture to Processing
- A Hierarchical Framework for Graph Structure Learning in Histopathology Image Classification
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation
- Revisiting Batch Normalization For Practical Domain Adaptation
- Physics-Constrained Inc-GAN for Tunnel Propagation Modeling from Sparse Line Measurements
- Amortized Moment Matching for Visual Generation
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- MIND: Monge Inception Distance for Generative Models Evaluation
- PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report
- An efficient deep learning hashing neural network for mobile visual search
- Copy-Augmented Representation for Structure Invariant Template-Free Retrosynthesis
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series
- Efficient Large-Deformation Medical Image Registration via Recurrent Dynamic Correlation
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- WorldGrow: Generating Infinite 3D World
- Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
- Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
- Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
- BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
- Compressing Biology: Evaluating the Stable Diffusion VAE for Phenotypic Drug Discovery
- Enhancing Graph Neural Networks: A Mutual Learning Approach
- QCFace: Image Quality Control for boosting Face Representation & Recognition
- A New Type of Adversarial Examples
- Image augmentation with invertible networks in interactive satellite image change detection
- Inference-Time Compute Scaling For Flow Matching
- Beam Index Map Prediction in Unseen Environments from Geospatial Data
- Enhancing Cross-Patient Generalization in AI-Based Parkinson s Disease Detection
- Towards a Generalizable Fusion Architecture for Multimodal Object Detection
- A Comprehensive Survey on World Models for Embodied AI
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- FraQAT: Quantization Aware Training with Fractional bits
- AlphaQuanter: An End-to-End Tool-Orchestrated Agentic Reinforcement Learning Framework for Stock Trading
- Approximate Bilevel Graph Structure Learning for Histopathology Image Classification
- Counting Hallucinations in Diffusion Models
- LayerSync: Self-aligning Intermediate Layers
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- SDGraph: Multi-Level Sketch Representation Learning by Sparse-Dense Graph Architecture
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- Diffusion Transformers with Representation Autoencoders
- PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Dynamic Network-Based Two-Stage Time Series Forecasting for Affiliate Marketing
- NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes
- Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
- PENEX: AdaBoost-Inspired Neural Network Regularization
- DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- KTBox: A Modular LaTeX Framework for Semantic Color, Structured Highlighting, and Scholarly Communication
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
- SummDiff: Generative Modeling of Video Summarization with Diffusion
- Biology-driven assessment of deep learning super-resolution imaging of the porosity network in dentin
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Label-frugal satellite image change detection with generative virtual exemplar learning
- How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
- Carré du champ flow matching: better quality-generalisation tradeoff in generative models
- Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional Diffusion
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- Quantization Range Estimation for Convolutional Neural Networks
- HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
- Learning Robust Diffusion Models from Imprecise Supervision
- ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
- Dale meets Langevin: A Multiplicative Denoising Diffusion Model
- PGMEL: Policy Gradient-based Generative Adversarial Network for Multimodal Entity Linking
- Uncertainty-Aware Concept Bottleneck Models with Enhanced Interpretability
- Assessing Foundation Models for Mold Colony Detection with Limited Training Data
- Quantum Probabilistic Label Refining: Enhancing Label Quality for Robust Image Classification
- MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
- EVODiff: Entropy-aware Variance Optimized Diffusion Inference
- Noise-Guided Transport for Imitation Learning
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- DRIFT: Divergent Response in Filtered Transformations for Robust Adversarial Defense
- Fidelity-Aware Data Composition for Robust Robot Generalization
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- S3F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models
- Tag Prediction at Flickr: a View from the Darkroom
- IONext: Unlocking the Next Era of Inertial Odometry
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- SubZeroCore: A Submodular Approach with Zero Training for Coreset Selection
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability
- Deterministic Discrete Denoising
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
- It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
- SAGE:State-Aware Guided End-to-End Policy for Multi-Stage Sequential Tasks via Hidden Markov Decision Process
- Graph-Structured Visual Imitation
- Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI
- Intrinsic Geometric Vulnerability of High-Dimensional Artificial Intelligence
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Neural Networks with Recurrent Generative Feedback
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- Designing Practical Models for Isolated Word Visual Speech Recognition
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Generalization and the Rise of System-level Creativity in Science
- Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem
- Faster CryptoNets: Leveraging Sparsity for Real-World Encrypted Inference
- Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
- Convolutional Neural Network Optimization for Beehive Classification Using Bioacoustic Signals
- A review of Recent Techniques for Person Re-Identification
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- Geometric Mixture Classifier (GMC): A Discriminative Per-Class Mixture of Hyperplanes
- Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image Synthesis
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Autoguided Online Data Curation for Diffusion Model Training
- Interleaved Group Convolutions for Deep Neural Networks
- Statistically Motivated Second Order Pooling
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- Self Identity Mapping
- More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- Diffusion Models Beat GANs on Image Synthesis
- GhostNetV3-Small: A Tailored Architecture and Comparative Study of Distillation Strategies for Tiny Images
- The Quest for Universal Master Key Filters in DS-CNNs
- CSIYOLO: An Intelligent CSI-based Scatter Sensing Framework for Integrated Sensing and Communication Systems
- REGEN: Real-Time Photorealism Enhancement in Games via a Dual-Stage Generative Network Framework
- NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
- Semantic Fusion with Fuzzy-Membership Features for Controllable Language Modelling
- Membership Inference Attacks on Recommender System: A Survey
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- CoAtNeXt:An Attention-Enhanced ConvNeXtV2-Transformer Hybrid Model for Gastric Tissue Classification
- Objectness Similarity: Capturing Object-Level Fidelity in 3D Scene Evaluation
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts
- HyperTTA: Test-Time Adaptation for Hyperspectral Image Classification under Distribution Shifts
- ADHDeepNet From Raw EEG to Diagnosis: Improving ADHD Diagnosis through Temporal-Spatial Processing, Adaptive Attention Mechanisms, and Explainability in Raw EEG Signals
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- Improving Performance, Robustness, and Fairness of Radiographic AI Models with Finely-Controllable Synthetic Data
- Breast Cancer Detection in Thermographic Images via Diffusion-Based Augmentation and Nonlinear Feature Fusion
- Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
- Parameter-Free Logit Distillation via Sorting Mechanism
- Missing Fine Details in Images: Last Seen in High Frequencies
- SL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Layer-wise Analysis for Quality of Multilingual Synthesized Speech
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Scale-interaction transformer: a hybrid cnn-transformer model for facial beauty prediction
- Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion
- An Automated, Scalable Machine Learning Model Inversion Assessment Pipeline
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- High Cursive Complex Character Recognition using GAN External Classifier
- Efficient Pyramidal Analysis of Gigapixel Images on a Decentralized Modest Computer Cluster
- Conditional-t3VAE: Equitable Latent Space Allocation for Fair Generation
- Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints
- An Investigation of Visual Foundation Models Robustness
- Improving atomic force microscopy structure discovery via style-translation
- Systematic Integration of Attention Modules into CNNs for Accurate and Generalizable Medical Image Diagnosis
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
- Adaptive Contrast Adjustment Module: A Clinically-Inspired Plug-and-Play Approach for Enhanced Fetal Plane Classification
- Automatic Identification and Description of Jewelry Through Computer Vision and Neural Networks for Translators and Interpreters
- Multi-Focused Video Group Activities Hashing
- Re-ID done right: towards good practices for person re-identification
- Audio-video Emotion Recognition in the Wild using Deep Hybrid Networks
- E-Stitchup: Data Augmentation for Pre-Trained Embeddings
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- diveXplore 6.0: ITEC's Interactive Video Exploration System at VBS 2022
- EffNetViTLoRA: An Efficient Hybrid Deep Learning Approach for Alzheimer's Disease Diagnosis
- Saddle Hierarchy in Dense Associative Memory
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- A Deep Learning Application for Psoriasis Detection
- Heatmap Regression without Soft-Argmax for Facial Landmark Detection
- A Pursuit of Temporal Accuracy in General Activity Detection
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
- Jamming Identification with Differential Transformer for Low-Altitude Wireless Networks
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
- ModiPick: SLA-aware Accuracy Optimization For Mobile Deep Inference
- Increasing the Utility of Synthetic Images through Chamfer Guidance
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Poaching Hotspot Identification Using Satellite Imagery
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- Generation of Indian Sign Language Letters, Numbers, and Words
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Label Smoothing is a Pragmatic Information Bottleneck
- A Multi-Scale CNN and Curriculum Learning Strategy for Mammogram\n Classification
- Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
- Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- GANime: Generating Anime and Manga Character Drawings from Sketches with Deep Learning
- Representation Understanding via Activation Maximization
- CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
- Practical Block-wise Neural Network Architecture Generation
- ComicGAN: Text-to-Comic Generative Adversarial Network
- Advanced Deep Learning Techniques for Accurate Lung Cancer Detection and Classification
- DSConv: Dynamic Splitting Convolution for Pansharpening
- Recurrent Deep Differentiable Logic Gate Networks
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
- Automated ultrasound doppler angle estimation using deep learning
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Energy-Efficient Real-Time 4-Stage Sleep Classification at 10-Second Resolution: A Comprehensive Study
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- Slice or the Whole Pie? Utility Control for AI Models
- Investigating the Impact of Large-Scale Pre-training on Nutritional Content Estimation from 2D Images
- Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- Rep-GLS: Report-Guided Generalized Label Smoothing for Robust Disease Detection
- Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion
- Multi-vision Attention Networks for On-line Red Jujube Grading
- OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
- Deep Learning for Pavement Condition Evaluation Using Satellite Imagery
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- CIF: A Constrained Inversion Framework for Reliable Message Extraction in Diffusion-Based Generative Steganography
- Calibrated Language Models and How to Find Them with Label Smoothing
- I Am Big, You Are Little; I Am Right, You Are Wrong
- Analysis of Hyperparameter Optimization Effects on Lightweight Deep Models for Real-Time Image Classification
- Training-free Geometric Image Editing on Diffusion Models
- Stress-Aware Resilient Neural Training
Related