ImageNet classification with deep convolutional neural networks
2017/05/24 by Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton · 2648 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Advanced Image Processing Techniques
paper · pdf · doi:10.1145/3065386
Abstract
We trained a large, deep convolutional neural network to classify the 1.2 million high-resolution images in the ImageNet LSVRC-2010 contest into the 1000 different classes. On the test data, we achieved top-1 and top-5 error rates of 37.5% and 17.0%, respectively, which is considerably better than the previous state-of-the-art. The neural network, which has 60 million parameters and 650,000 neurons, consists of five convolutional layers, some of which are followed by max-pooling layers, and three fully connected layers with a final 1000-way softmax. To make training faster, we used non-saturating neurons and a very efficient GPU implementation of the convolution operation. To reduce overfitting in the fully connected layers we employed a recently developed regularization method called "dropout" that proved to be very effective. We also entered a variant of this model in the ILSVRC-2012 competition and achieved a winning top-5 test error rate of 15.3%, compared to 26.2% achieved by the second-best entry.
Cited by
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- Data-Driven Aerospace Engineering: Reframing the Industry with Machine Learning
- YFCC100M
- Convolutional neural network for seismic impedance inversion
- Elastic prestack seismic inversion through discrete cosine transform reparameterization and convolutional neural networks
- Deep Learning Methods for Reynolds-Averaged Navier–Stokes Simulations of Airfoil Flows
- Deep learning with coherent nanophotonic circuits
- Recent Advances and Applications of Machine Learning in Experimental Solid Mechanics: A Review
- Quantum-chemical insights from deep tensor neural networks
- Using deep learning and Google Street View to estimate the demographic makeup of neighborhoods across the United States
- Artificial Intelligence in manufacturing: State of the art, perspectives, and future directions
- The Principles of Deep Learning Theory
- A Comprehensive Survey on Transfer Learning
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing
- Tumour-infiltrating lymphocytes: from prognosis to treatment selection
- Applications and Techniques for Fast Machine Learning in Science
- Multiclass magnetic resonance imaging brain tumor classification using artificial intelligence paradigm
- Prediction of causative genes in inherited retinal disorder from fundus photography and autofluorescence imaging using deep learning techniques
- Toward Causal Representation Learning
- Deep Reinforcement Learning Based Resource Allocation for V2V Communications
- Real-time differentiation of adenomatous and hyperplastic diminutive colorectal polyps during analysis of unaltered videos of standard colonoscopy using a deep learning model
- Artificial intelligence in fetal brain imaging: Advancements, challenges, and multimodal approaches for biometric and structural analysis
- Deep convolutional neural network for the automated detection and diagnosis of seizure using EEG signals
- Obstructive sleep apnea detection from single-lead electrocardiogram signals using one-dimensional squeeze-and-excitation residual group network
- A scoping review of transfer learning research on medical image analysis using ImageNet
- Advancements in automated nuclei segmentation for histopathology using you only look once-driven approaches: A systematic review
- A novel wavelet sequence based on deep bidirectional LSTM network model for ECG signal classification
- Transparency of deep neural networks for medical image analysis: A review of interpretability methods
- Deep learning for denoising
- Regularized elastic full-waveform inversion using deep learning
- Deep learning for forest inventory and planning: a critical review on the remote sensing approaches so far and prospects for further applications
- Deep Neural Network Approximation Theory
- Deep learning and its application in geochemical mapping
- PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM
- Melanoma Recognition in Dermoscopy Images via Aggregated Deep Convolutional Features
- Physics and equality constrained artificial neural networks: Application to forward and inverse problems with multi-fidelity data fusion
- DeepWalk
- Dominant-Current Deep Learning Scheme for Electrical Impedance Tomography
- Efficient Mitchell’s Approximate Log Multipliers for Convolutional Neural Networks
- MTHAEL: Cross-Architecture IoT Malware Detection Based on Neural Network Advanced Ensemble Learning
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- Classification of the Clinical Images for Benign and Malignant Cutaneous Tumors Using a Deep Learning Algorithm
- Deep Learning With Edge Computing: A Review
- Neuro-Inspired Computing With Emerging Nonvolatile Memorys
- FeatureNet: Machining feature recognition based on 3D Convolution Neural Network
- Automated arrhythmia detection with homeomorphically irreducible tree technique using more than 10,000 individual subject ECG records
- Dual non-autonomous deep convolutional neural network for image denoising
- Image-Based Multiresolution Topology Optimization Using Deep Disjunctive Normal Shape Model
- Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations
- Optimal control of PDEs using physics-informed neural networks
- Characterising the neural time-courses of food attribute representations
- Deep Learning for Economists
- LiMuon: Light and Fast Muon Optimizer for Large Models
- AdamNX: An Adam improvement algorithm based on a novel exponential decay mechanism for the second-order moment estimate
- Reclaiming AI as a Theoretical Tool for Cognitive Science
- Hospital Length of Stay Prediction Methods
- Hypergraph convolution and hypergraph attention
- Rasabodha: Understanding Indian classical dance by recognizing emotions using deep learning
- U2-Net: Going deeper with nested U-structure for salient object detection
- Multi-crop Convolutional Neural Networks for lung nodule malignancy suspiciousness classification
- Illumination-aware faster R-CNN for robust multispectral pedestrian detection
- Haar wavelet downsampling: A simple but effective downsampling module for semantic segmentation
- Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
- Optimization Methods for Large-Scale Machine Learning
- Introduction to deep learning methods for multi‐species predictions
- Deep photovoltaic nowcasting
- Wear particle classification considering particle overlapping
- Automatic morphologic classification of Martian craters using imbalanced datasets of Tianwen-1’s MoRIC images with deep neural networks
- Anomalous-diffusion synthesis of non-Gaussian reservoir anomalies for time-lapse seismic inversion
- Exact Neural-Network Representations of the Motzkin States
- Conservative physics-informed neural networks on discrete domains for conservation laws: Applications to forward and inverse problems
- Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed Recognition
- A novel deep learning-based modelling strategy from image of particles to mechanical properties for granular materials with CNN and BiLSTM
- Wireless Networks Design in the Era of Deep Learning: Model-Based, AI-Based, or Both?
- Deep learning for the partially linear Cox model
- Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network
- 4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
- Pointing-Based Object Recognition
- Prediction model of temperature field in dual-mode combustors based on wall pressure
- Characterization of heat release rate by OH* and CH* chemiluminescence
- Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
- Glyce: Glyph-vectors for Chinese Character Representations
- Scaling Synthetic-Image Pre-Training for Federated Fine-Tuning of Large Vision Models
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- LieBN: Batch Normalization over Lie Groups
- Edge-Local and Qubit-Efficient Quantum Graph Learning for the NISQ Era
- Benchmarking NACTI Species Recognition in Long-Tailed Regimes
- The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
- EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database
- UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation
- Provably Lossless Acceleration of DNN Mutation Testing via Memoization
- Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- Feature-Guided Diffusion for Non-Differentiable Inverse Rendering
- VecFontLLM: Anchor-Guided Direct Synthesis of Chinese Vector Fonts
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
- AIMS: An uncertainty-aware AI experimentalist for quantum matter
- Cotton-SF YOLO: Learning Structural and Frequency Cues for Early Cotton Square Detection in Complex Field Environments
- Explicit Over Implicit: Enhancing CNNs Via Complex Structure Tensor Representations for Periocular Recognition
- CNN-Based Surface Temperature Forecasts with Ensemble Numerical Weather Prediction
- Information Theory and Statistical Learning
- Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
- AI solutions for evolutionary genomics of nonmodel species
- BehaveAI enables rapid detection and classification of objects and behavior from motion
- The Acceleration of Artificial Intelligence: Rethinking Organization and Work in an Era of Rapid Technological Change
- The Universal Weight Subspace Hypothesis
- MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
- Artificial intelligence for risk assessment and outcome prediction in malignant haematology
- Impact of Multi-View Fusion and Biomechanical Modeling on Markerless Motion Tracking
- Living Synthetic Benchmarks: A Neutral and Cumulative Framework for Simulation Studies
- FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems
- Video models are zero-shot learners and reasoners
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- The Temporal Scaffolding of Sensory Organization
- scPortrait integrates single-cell images into multimodal modeling
- uGMM-NN: Univariate Gaussian Mixture Model Neural Network
- Power Stabilization for AI Training Datacenters
- Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
- Optimizers Qualitatively Alter Solutions And We Should Leverage This
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Potential role of developmental experience in the emergence of the parvo-magno distinction
- Are Statistical Methods Obsolete in the Era of Deep Learning? A Study of ODE Inverse Problems
- Perception Encoder: The best visual embeddings are not at the output of the network
- Brain-guided convolutional neural networks reveal task-specific representations in scene processing
- Improving Computer Vision Interpretability: Transparent Two-level Classification for Complex Scenes
- What Is Artificial General Intelligence?
- Deep Learning is Not So Mysterious or Different
- Can machines learn density functionals? Past, present, and future of ML in DFT
- ILIAS: Instance-Level Image retrieval At Scale
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Evolution and The Knightian Blindspot of Machine Learning
- Language models and Automated Essay Scoring
- Representations of Sound in Deep Learning of Audio Features from Music
- FedRAD: Federated Robust Adaptive Distillation
- Controllable Data Augmentation Through Deep Relighting
- A Convolutional Attention Network for Extreme Summarization of Source Code
- Modeling Spatial and Temporal Cues for Multi-label Facial Action Unit Detection
- Reducing Data Motion to Accelerate the Training of Deep Neural Networks
- Computer Vision and Abnormal Patient Gait Assessment a Comparison of Machine Learning Models
- TorchQuantumDistributed
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Profiling based Out-of-core Hybrid Method for Large Neural Networks
- How Deep is the Feature Analysis underlying Rapid Visual Categorization?
- Sampled Softmax with Random Fourier Features
- Spatially Correlated Patterns in Adversarial Images
- Learning to Compose Hypercolumns for Visual Correspondence
- Machine Translation: A Literature Review
- Deep Air Quality Forecasting Using Hybrid Deep Learning Framework
- Procrustean Training for Imbalanced Deep Learning
- NullSpaceNet: Nullspace Convoluional Neural Network with Differentiable Loss Function
- A Convolutional Neural Network for gaze preference detection: A\n potential tool for diagnostics of autism spectrum disorder in children
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- RotNet: Fast and Scalable Estimation of Stellar Rotation Periods Using\n Convolutional Neural Networks
- Meta Cross-Modal Hashing on Long-Tailed Data
- Matrix Smoothing: A Regularization for DNN with Transition Matrix under Noisy Labels
- Speeding up convolutional networks pruning with coarse ranking
- Generating Unrestricted 3D Adversarial Point Clouds
- Is Heterophily A Real Nightmare For Graph Neural Networks To Do Node Classification?
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Differentiable Learning-to-Normalize via Switchable Normalization
- NAS-DIP: Learning Deep Image Prior with Neural Architecture Search
- A Simple Saliency Method That Passes the Sanity Checks
- SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)
- Deep Multi-View Spatial-Temporal Network for Taxi Demand Prediction
- Neural Abstractive Text Summarization with Sequence-to-Sequence Models
- MUTE: Data-Similarity Driven Multi-hot Target Encoding for Neural Network Design
- Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection
- Multi-Subspace Neural Network for Image Recognition
- Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control
- A Visual Analytics Framework for Explaining and Diagnosing Transfer Learning Processes
- Colorectal cancer diagnosis from histology images: A comparative study
- Norm-based generalisation bounds for multi-class convolutional neural\n networks
- Universal Approximation with Quadratic Deep Networks
- Unconstrained Road Marking Recognition with Generative Adversarial Networks
- IBM Deep Learning Service
- Fast GPU Linear Algebra via Compile Time Expression Fusion
- Rethinking Intrinsic Dimension Estimation in Neural Representations
- Evolving Deep Convolutional Neural Networks for Image Classification
- Predictive Analysis of COVID-19 Time-series Data from Johns Hopkins University
- Per-pixel Classification Rebar Exposures in Bridge Eye-inspection
- Deep and Shallow Covariance Feature Quantization for 3D Facial Expression Recognition
- Manifold Criterion Guided Transfer Learning via Intermediate Domain Generation
- Gated Feedback Refinement Network for Coarse-to-Fine Dense Semantic Image Labeling
- TopoResNet: A hybrid deep learning architecture and its application to\n skin lesion classification
- PCR-ORB: Enhanced ORB-SLAM3 with Point Cloud Refinement Using Deep Learning-Based Dynamic Object Filtering
- End to End Learning for Self-Driving Cars
- Open Problems in Cooperative AI
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- Deep Learning for Needle Detection in a Cannulation Simulator
- The Dynamics of Gradient Descent for Overparametrized Neural Networks
- Progressive Learning of Low-Precision Networks
- Why do linear SVMs trained on HOG features perform so well?
- Scaling Wide Residual Networks for Panoptic Segmentation
- Learning Hybrid Representation by Robust Dictionary Learning in Factorized Compressed Space
- Deep Learning for Spectrum Sensing
- Energy and Memory-Efficient Federated Learning With Ordered Layer Freezing
- Enhancing Convolutional Neural Networks for Face Recognition with\n Occlusion Maps and Batch Triplet Loss
- Parallel Support Vector Machines in Practice
- A Benchmarking Framework for Interactive 3D Applications in the Cloud
- When science meets geopolitics: global AI research network transformation (2000–2025)
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
- CNN-based Automatic Detection of Bone Conditions via Diagnostic CT Images for Osteoporosis Screening
- Potentials and challenges of polymer informatics: exploiting machine learning for polymer design
- RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network
- Perceiver: General Perception with Iterative Attention
- Alleviating Bottlenecks for DNN Execution on GPUs via Opportunistic Computing
- Learning a Reinforced Agent for Flexible Exposure Bracketing Selection
- Solving Optical Tomography with Deep Learning
- The Challenge of Multi-Operand Adders in CNNs on FPGAs: How not to solve it!
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
- A Simple Method for Commonsense Reasoning
- Alignment of electron optical beam shaping elements using a convolutional neural network
- Partial Graph Reasoning for Neural Network Regularization
- A mountable toilet system for personalized health monitoring via the analysis of excreta
- Immersive exposure to simulated visual hallucinations modulates high-level human cognition
- ChebLieNet: Invariant Spectral Graph NNs Turned Equivariant by\n Riemannian Geometry on Lie Groups
- Semi-supervised Sequence Learning
- MLP-Mixer: An all-MLP Architecture for Vision
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Object Detection from Scratch with Deep Supervision
- Constrained Linear Data-feature Mapping for Image Classification
- DeepVisInterests: CNN-Ontology Prediction of Users Interests from Social Images
- Stochastic Adversarial Gradient Embedding for Active Domain Adaptation
- Vision Transformers Need Registers
- The Pace of Artificial Intelligence Innovations: Speed, Talent, and Trial-and-Error
- Copy-Move Forgery Classification via Unsupervised Domain Adaptation
- Fusion Recurrent Neural Network
- The Golden Ratio of Learning and Momentum
- Spatial Interpolation of Room Impulse Responses based on Deeper Physics-Informed Neural Networks with Residual Connections
- Evolution in Groups: A deeper look at synaptic cluster driven evolution of deep neural networks
- Accelerating CNN Training by Pruning Activation Gradients
- MCMC Guided CNN Training and Segmentation for Pancreas Extraction
- Time-Limited Toeplitz Operators on Abelian Groups: Applications in Information Theory and Subspace Approximation
- Survey of Dropout Methods for Deep Neural Networks
- Deep Selective Combinatorial Embedding and Consistency Regularization for Light Field Super-resolution
- APEX-Net: Automatic Plot Extractor Network
- Semantics, Representations and Grammars for Deep Learning
- Evaluating an Adaptive Multispectral Turret System for Autonomous Tracking Across Variable Illumination Conditions
- A Robust framework for sound event localization and detection on real recordings
- A Simple Framework for Contrastive Learning of Visual Representations
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Optimal Gradient Quantization Condition for Communication-Efficient Distributed Training
- A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification
- Deep Interactive Denoiser (DID) for X-Ray Computed Tomography
- Compression-aware Continual Learning using Singular Value Decomposition
- Segmentation of digital rock images using deep convolutional autoencoder networks
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Recover and Identify: A Generative Dual Model for Cross-Resolution Person Re-Identification
- Improving Image Classification with Location Context
- Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition
- A Sustainable Multi-modal Multi-layer Emotion-aware Service at the Edge
- Reasoning About Human-Object Interactions Through Dual Attention Networks
- Performance tuning for deep learning on a many-core processor (master thesis)
- Associatively Segmenting Instances and Semantics in Point Clouds
- Resolution Switchable Networks for Runtime Efficient Image Recognition
- American Sign Language fingerspelling recognition in the wild
- Transfer Learning Using Classification Layer Features of CNN
- The Loss Surfaces of Multilayer Networks
- TransNFCM: Translation-Based Neural Fashion Compatibility Modeling
- Learning Manipulation under Physics Constraints with Visual Perception
- Multi-Service Mobile Traffic Forecasting via Convolutional Long Short-Term Memories
- ESAI: Efficient Split Artificial Intelligence via Early Exiting Using Neural Architecture Search
- Deep inspection: an electrical distribution pole parts study via deep neural networks
- Improving robustness against common corruptions by covariate shift adaptation
- Node Classification on Graphs with Few-Shot Novel Labels via Meta Transformed Network Embedding
- Tuning Algorithms and Generators for Efficient Edge Inference
- Deep Image Orientation Angle Detection
- Towards High-Level Semantic Intelligence
- The neural architecture of language: Integrative modeling converges on predictive processing
- Data-driven emergence of convolutional structure in neural networks
- Neural model robustness for skill routing in large-scale conversational AI systems: A design choice exploration
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression
- Basic Principles of Clustering Methods
- A predictor-corrector method for the training of deep neural networks
- Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
- WeMix: How to Better Utilize Data Augmentation
- Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous Space
- Accelerating Multi-Model Inference by Merging DNNs of Different Weights
- HABNet: Machine Learning, Remote Sensing Based Detection and Prediction of Harmful Algal Blooms
- A Theoretical Framework for Robustness of (Deep) Classifiers against Adversarial Examples
- Deep Tracking: Visual Tracking Using Deep Convolutional Networks
- Wavelet Integrated CNNs for Noise-Robust Image Classification
- Generalisation in humans and deep neural networks
- Physical world assistive signals for deep neural network classifiers -- neither defense nor attack
- Siamese Anchor Proposal Network for High-Speed Aerial Tracking
- GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation
- Learning from Multi-domain Artistic Images for Arbitrary Style Transfer
- Ablation Studies in Artificial Neural Networks
- Multilingual Image Description with Neural Sequence Models
- A CNN-RNN Architecture for Multi-Label Weather Recognition
- Joint Shape Representation and Classification for Detecting PDAC
- Adversarial Attack Type I: Cheat Classifiers by Significant Changes
- Extraction of Salient Sentences from Labelled Documents
- DeepMRSeg: A convolutional deep neural network for anatomy and abnormality segmentation on MR images
- MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information Maximization
- Iterative temporal differencing with random synaptic feedback weights support error backpropagation for deep learning
- Spectral Unsupervised Domain Adaptation for Visual Recognition
- P-ODN: Prototype based Open Deep Network for Open Set Recognition
- Sparse Vector Transmission: An Idea Whose Time Has Come
- Efficient Multi-Modal Embeddings from Structured Data
- Object Recognition from Short Videos for Robotic Perception
- Framing U-Net via Deep Convolutional Framelets: Application to Sparse-view CT
- Hypernetwork-Based Augmentation
- Post-Earthquake Assessment of Buildings Using Deep Learning
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
- Model-Based Domain Generalization
- Multi-loss ensemble deep learning for chest X-ray classification
- Robust Training in High Dimensions via Block Coordinate Geometric Median Descent
- CLAR: Contrastive Learning of Auditory Representations
- Statistical theory for image classification using deep convolutional neural networks with cross-entropy loss under the hierarchical max-pooling model
- Deep Transfer Learning for Automated Diagnosis of Skin Lesions from Photographs
- Deep Collective Learning: Learning Optimal Inputs and Weights Jointly in Deep Neural Networks
- Lagrangian Neural Networks
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Fully Dynamic Inference with Deep Neural Networks
- An empirical investigation into the properties of standard word embeddings
- Validating predictions of burial mounds with field data: the promise and reality of machine learning
- Cascaded Structure Tensor Framework for Robust Identification of Heavily Occluded Baggage Items from X-ray Scans
- A Survey on Deep Geometry Learning: From a Representation Perspective
- Deep Anchored Convolutional Neural Networks
- DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing
- Single-shot Channel Pruning Based on Alternating Direction Method of Multipliers
- Mutually-aware Sub-Graphs Differentiable Architecture Search
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- AI Empowered Communication and Radar Modulation Recognition: A Survey
- Image-Based Geo-Localization Using Satellite Imagery
- Automated Numerical Stability Analysis of Deep Learning Operators
- Equivariant Q Learning in Spatial Action Spaces
- Guided Evolution for Neural Architecture Search
- Cuttlefish: A Lightweight Primitive for Adaptive Query Processing
- Deep Neural Network for Learning to Rank Query-Text Pairs
- Mind the Missing Split: Resolving Feature Heterogeneity in Swarm Learning with Random Forests
- On the Opportunities and Risks of Foundation Models
- Sparse Symplectically Integrated Neural Networks
- Image2Lego: Customized LEGO Set Generation from Images
- Why is AI hard and Physics simple?
- Deep Kinship Verification via Appearance-shape Joint Prediction and Adaptation-based Approach
- NeuroMAX: A High Throughput, Multi-Threaded, Log-Based Accelerator for Convolutional Neural Networks
- A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions
- Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
- Real-time Reconstruction of Human Visual Perception from fMRI
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- DEAL: Difficulty-aware Active Learning for Semantic Segmentation
- Guided Attention Network for Object Detection and Counting on Drones
- The Causal Loss: Driving Correlation to Imply Causation
- Text-to-image Synthesis via Symmetrical Distillation Networks
- Benchmarking deep learning models for Raman spectroscopy across open-source datasets
- SegSort: Segmentation by Discriminative Sorting of Segments
- Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
- LibContinual: A Comprehensive Library towards Realistic Continual Learning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- An Empirical Study of Machine Learning Robustness and Scalability for Imbalanced Tabular Clinical Data in Emergency and Critical Care
- Incremental Learning Using a Grow-and-Prune Paradigm with Efficient Neural Networks
- Detecting Medical Misinformation on Social Media Using Multimodal Deep Learning
- Granular-ball Guided Masking: Structure-aware Data Augmentation
- Programmable Optical Spectrum Shapers as Computing Primitives for Accelerating Convolutional Neural Networks
- Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
- A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
- Monocular Depth Estimation with Augmented Ordinal Depth Relationships
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- Bloom Filter Encoding for Machine Learning
- Statistically Significant Stopping of Neural Network Training
- Image-specific Convolutional Kernel Modulation for Single Image Super-resolution
- Handling Inter-Annotator Agreement for Automated Skin Lesion Segmentation
- The Seismic Wavefield Common Task Framework
- Intrinsic dimension of data representations in deep neural networks
- Combating Adversarial Misspellings with Robust Word Recognition
- DK-STN: A Domain Knowledge Embedded Spatio-Temporal Network Model for MJO Forecast
- A Convolutional Neural Deferred Shader for Physics Based Rendering
- Deep Kernel Learning
- Predicting Human Trajectories by Learning and Matching Patterns
- Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- Adversarial Robustness in Zero-Shot Learning:An Empirical Study on Class and Concept-Level Vulnerabilities
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- Adversarial Robustness of Vision in Open Foundation Models
- Semi-Supervised Online Learning on the Edge by Transforming Knowledge from Teacher Models
- KOSS: Kalman-Optimal Selective State Spaces for Long-Term Sequence Modeling
- SARMAE: Masked Autoencoder for SAR Representation Learning
- Batch Normalization-Free Fully Integer Quantized Neural Networks via Progressive Tandem Learning
- Hypernetworks That Evolve Themselves
- AI Needs Physics More Than Physics Needs AI
- Higher-Order LaSDI: Reduced Order Modeling with Multiple Time Derivatives
- In Pursuit of Pixel Supervision for Visual Pre-training
- Stylized Synthetic Augmentation further improves Corruption Robustness
- From Risk to Resilience: Towards Assessing and Mitigating the Risk of Data Reconstruction Attacks in Federated Learning
- Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
- An updated efficient galaxy morphology classification model based on ConvNeXt encoding with UMAP dimensionality reduction
- TrajSyn: Privacy-Preserving Dataset Distillation from Federated Model Trajectories for Server-Side Adversarial Training
- Which Coauthor Should I Nominate in My 99 ICLR Submissions? A Mathematical Analysis of the ICLR 2026 Reciprocal Reviewer Nomination Policy
- Prototypical Learning Guided Context-Aware Segmentation Network for Few-Shot Anomaly Detection
- Mimicking Human Visual Development for Learning Robust Image Representations
- DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- On Improving Deep Active Learning with Formal Verification
- Multiple weak biases support adaptive choices without prior experience: a self-supervised strategy
- Unbiased Mean Teacher for Cross-domain Object Detection
- Learning Deep Bilinear Transformation for Fine-grained Image Representation
- Bilinear CNNs for Fine-grained Visual Recognition
- Efficient Dense Modules of Asymmetric Convolution for Real-Time Semantic Segmentation
- Gradient descent aligns the layers of deep linear networks
- Practical Implementation of Memristor-Based Threshold Logic Gates
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- Dropout Neural Network Training Viewed from a Percolation Perspective
- Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory
- Multiclass Graph-Based Large Margin Classifiers: Unified Approach for Support Vectors and Neural Networks
- Explanatory models in neuroscience: Part 1 -- taking mechanistic abstraction seriously
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- Generative Spatiotemporal Data Augmentation
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Machine learning methods for subpixel trajectory reconstruction in discretized position detectors
- Efficient Classification of Very Large Images with Tiny Objects
- Fully Inductive Node Representation Learning via Graph View Transformation
- DREAM-B3P: Dual-Stream Transformer Network Enhanced by Feedback Diffusion Model for Blood-Brain Barrier Penetrating Peptide Prediction
- Where Classification Fails, Interpretation Rises
- Understanding Deep Learning Techniques for Image Segmentation
- FaaT: A Transparent Auto-Scaling Cache for Serverless Applications
- Back to the Baseline: Examining Baseline Effects on Explainability Metrics
- Application of a semantic segmentation convolutional neural network for\n accurate automatic detection and mapping of solar photovoltaic arrays in\n aerial imagery
- A Review of Machine Learning Applications in Fuzzing
- Visual Search Asymmetry: Deep Nets and Humans Share Similar Inherent\n Biases
- Quantum Algorithms for Unsupervised Machine Learning and Neural Networks
- Automated Detection of Equine Facial Action Units
- Poker-CNN: A Pattern Learning Strategy for Making Draws and Bets in Poker Games
- Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
- Attending to Discriminative Certainty for Domain Adaptation
- An Overview and Case Study of the Clinical AI Model Development Life Cycle for Healthcare Systems
- SGNet: A Super-class Guided Network for Image Classification and Object Detection
- Federated Domain Generalization with Latent Space Inversion
- Learning to Optimize Tensor Programs
- ZenSVI: An open-source software for the integrated acquisition, processing and analysis of street view imagery towards scalable urban science
- Dynamic Efficient Adversarial Training Guided by Gradient Magnitude
- Learning Quantized Neural Nets by Coarse Gradient Method for Non-linear Classification
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- DirectSwap: Mask-Free Cross-Identity Training and Benchmarking for Expression-Consistent Video Head Swapping
- Microscopic Vehicle Trajectories from Heterogeneous and Area-Based Traffic
- Causality-inspired Single-source Domain Generalization for Medical Image Segmentation
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Including Image-based Perception in Disturbance Observer for Warehouse Drones
- An Approach for Detection of Entities in Dynamic Media Contents
- Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
- PR-CapsNet: Pseudo-Riemannian Capsule Network with Adaptive Curvature Routing for Graph Learning
- LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks
- On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
- Fully Convolutional Neural Networks for Dynamic Object Detection in Grid Maps (Masters Thesis)
- Atlas: A Dataset and Benchmark for E-commerce Clothing Product\n Categorization
- Natural-Logarithm-Rectified Activation Function in Convolutional Neural Networks
- PCGAN-CHAR: Progressively Trained Classifier Generative Adversarial Networks for Classification of Noisy Handwritten Bangla Characters
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Tackling Graphical NLP problems with Graph Recurrent Networks
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- PolSAR Image Classification Based on Dilated Convolution and Pixel-Refining Parallel Mapping network in the Complex Domain
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- Revisiting Latent-Space Interpolation via a Quantitative Evaluation Framework
- Barrier-Free Large-Scale Sparse Tensor Accelerator (BARISTA) For\n Convolutional Neural Networks
- Explaining Deep Neural Networks
- Life is not black and white -- Combining Semi-Supervised Learning with fuzzy labels
- Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- From Forecast to Action: A Deep Learning Model for Predicting Power Outages During Tropical Cyclones
- Hierarchical Deep Learning for Diatom Image Classification: A Multi-Level Taxonomic Approach
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
- A Perception CNN for Facial Expression Recognition
- CN-CELEB: a challenging Chinese speaker recognition dataset
- Analysis of Social Media Data using Multimodal Deep Learning for Disaster Response
- Cross-Image Region Mining with Region Prototypical Network for Weakly Supervised Segmentation
- Tensor object classification via multilinear discriminant analysis\n network
- Novel Deep Learning Architectures for Classification and Segmentation of Brain Tumors from MRI Images
- Statistical physics for artificial neural networks
- Towards Hardware-Agnostic Gaze-Trackers
- SPOOF: Simple Pixel Operations for Out-of-Distribution Fooling
- A Comparative Study on Synthetic Facial Data Generation Techniques for Face Recognition
- Achieving Approximate Symmetry Is Exponentially Easier than Exact Symmetry
- AI & Human Co-Improvement for Safer Co-Superintelligence
- Affine Non-negative Collaborative Representation Based Pattern Classification
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- TEA: Temporal Excitation and Aggregation for Action Recognition
- On the Effect of Regularization on Nonparametric Mean-Variance Regression
- A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects
- Generative Recursive Reasoning
- ConvSequential-SLAM: A Sequence-based, Training-less Visual Place\n Recognition Technique for Changing Environments
- A Comparative Review of Recent Few-Shot Object Detection Algorithms
- Human-Level Control without Server-Grade Hardware
- TactileSGNet: A Spiking Graph Neural Network for Event-based Tactile Object Recognition
- An all-optical convolutional neural network for image identification
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- A Cascaded Zoom-In Network for Patterned Fabric Defect Detection
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- QoSDiff: An Implicit Topological Embedding Learning Framework Leveraging Denoising Diffusion and Adversarial Attention for Robust QoS Prediction
- CASTLE: Regularization via Auxiliary Causal Graph Discovery
- RRPN++: Guidance Towards More Accurate Scene Text Detection
- Performance Evaluation of Transfer Learning Based Medical Image Classification Techniques for Disease Detection
- Bayes-DIC Net: Estimating Digital Image Correlation Uncertainty with Bayesian Neural Networks
- CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
- Parameter efficient hybrid spiking-quantum convolutional neural network with surrogate gradient and quantum data-reupload
- Deep Unfolding: Recent Developments, Theory, and Design Guidelines
- On the Binding Problem in Artificial Neural Networks
- Multi-Scale Visual Prompting for Lightweight Small-Image Classification
- What does LIME really see in images?
- Compressed Video Action Recognition with Refined Motion Vector
- Using Deep Learning and Machine Learning to Detect Epileptic Seizure with Electroencephalography (EEG) Data
- Group-Structured Adversarial Training
- Detecting Electric Devices in 3D Images of Bags
- Benchmarking machine learning models for multi-class state recognition in double quantum dot data
- Leveraging Large-Scale Pretrained Spatial-Spectral Priors for General Zero-Shot Pansharpening
- OmniPerson: Unified Identity-Preserving Pedestrian Generation
- Associative Memory using Attribute-Specific Neuron Groups-1: Learning between Multiple Cue Balls
- Breast Cell Segmentation Under Extreme Data Constraints: Quantum Enhancement Meets Adaptive Loss Stabilization
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- TGDD: Trajectory Guided Dataset Distillation with Balanced Distribution
- Dependency Aware Filter Pruning
- Classification of Radio Signals and HF Transmission Modes with Deep Learning
- Equilibrium Propagation Without Limits
- TPCNet: Triple physical constraints for Low-light Image Enhancement
- Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning
- Multimodal Mixture-of-Experts for ISAC in Low-Altitude Wireless Networks
- Mid-Level Visual Representations Improve Generalization and Sample Efficiency for Learning Visuomotor Policies
- Neural Networks for Predicting Permeability Tensors of 2D Porous Media: Comparison of Convolution- and Transformer-based Architectures
- Deep Learning for Generic Object Detection: A Survey
- Deep learning in neural networks: An overview
- A systematic study of the class imbalance problem in convolutional neural networks
- Masked Autoencoders Are Scalable Vision Learners
- SRM : A Style-based Recalibration Module for Convolutional Neural Networks
- Compact representations of convolutional neural networks via weight\n pruning and quantization
- CrossedWires: A Dataset of Syntactically Equivalent but Semantically\n Disparate Deep Learning Models
- Non-Parametric Neural Style Transfer
- Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control
- An Introduction to Convolutional Neural Networks
- Directed evolution algorithm drives neural prediction
- Spatiotemporal Satellite Image Downscaling with Transfer Encoders and Autoregressive Generative Models
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Long-term Temporal Convolutions for Action Recognition
- Emotion Recognition from Speech
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- Controllable 3D Object Generation with Single Image Prompt
- Scaling Laws for Neural Language Models
- Novelty Detection in MultiClass Scenarios with Incomplete Set of Class Labels
- All-Weather Object Recognition Using Radar and Infrared Sensing
- Dual-Projection Fusion for Accurate Upright Panorama Generation in Robotic Vision
- Designing a Micro-Benchmark Suite to Evaluate gRPC for TensorFlow: Early Experiences
- ODE guided Neural Data Augmentation Techniques for Time Series Data and its Benefits on Robustness
- CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning
- Structured Context Learning for Generic Event Boundary Detection
- Fixing the train-test resolution discrepancy
- N-ImageNet: Towards Robust, Fine-Grained Object Recognition with Event Cameras
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- "Why the face?": Exploring Robot Error Detection Using Instrumented Bystander Reactions
- SelfVIO: Self-supervised deep monocular Visual–Inertial Odometry and depth estimation
- First Steps towards Machine Learning for Prediction and Pre-Correction in Direct Laser Writing
- GhostNet: More Features from Cheap Operations
- Explaining Deep Learning Models for Structured Data using Layer-Wise Relevance Propagation
- Action Recognition with Kernel-based Graph Convolutional Networks
- A Unified Framework for Multi-View Multi-Class Object Pose Estimation
- Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
- Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
- CyCNN: A Rotation Invariant CNN using Polar Mapping and Cylindrical Convolution Layers
- Learning Deep Structure-Preserving Image-Text Embeddings
- CycleCluster: Modernising Clustering Regularisation for Deep Semi-Supervised Classification
- On-the-Job Learning with Bayesian Decision Theory
- CNN-Based Framework for Pedestrian Age and Gender Classification Using Far-View Surveillance in Mixed-Traffic Intersections
- Hybrid Context-Fusion Attention (CFA) U-Net and Clustering for Robust Seismic Horizon Interpretation
- DAONet-YOLOv8: An Occlusion-Aware Dual-Attention Network for Tea Leaf Pest and Disease Detection
- Adversarial Training Towards Robust Multimedia Recommender System
- Adversarial AutoMixup
- Active Learning for GCN-based Action Recognition
- tfShearlab: The TensorFlow Digital Shearlet Transform for Deep Learning
- A Variational View on Bootstrap Ensembles as Bayesian Inference
- Neurodevelopmental Age Estimation of Infants Using a 3D-Convolutional Neural Network Model based on Fusion MRI Sequences
- Neural Networks, Hypersurfaces, and Radon Transforms
- Photo-Realistic Video Prediction on Natural Videos of Largely Changing Frames
- A multi-task convolutional neural network for mega-city analysis using very high resolution satellite imagery and geospatial data
- Provably Powerful Graph Networks
- Tactile-Based Insertion for Dense Box-Packing
- Decoupled Dynamic Filter Networks
- STEP: Segmenting and Tracking Every Pixel
- Underexposed Image Correction via Hybrid Priors Navigated Deep Propagation
- FOSNet: An End-to-End Trainable Deep Neural Network for Scene Recognition
- Fundamentals of Regression
- Optimal Conversion of Conventional Artificial Neural Networks to Spiking Neural Networks
- BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation
- Mean-Field Limits for Two-Layer Neural Networks Trained with Consensus-Based Optimization
- Differentiable Physics-Neural Models enable Learning of Non-Markovian Closures for Accelerated Coarse-Grained Physics Simulations
- A Physics-Informed U-net-LSTM Network for Data-Driven Seismic Response Modeling of Structures
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and Inference
- Physics-informed neural networks method in high-dimensional integrable systems
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Generalized Out-of-Distribution Detection: A Survey
- Privacy-Preserving Self-Taught Federated Learning for Heterogeneous Data
- Improving the Generalization of End-to-End Driving through Procedural Generation
- WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection
- Comprehensive Graph-conditional Similarity Preserving Network for Unsupervised Cross-modal Hashing
- Deep Learning Assisted Calibrated Beam Training for Millimeter-Wave Communication Systems
- Guaranteed Optimal Compositional Explanations for Neurons
- Open Vocabulary Compositional Explanations for Neuron Alignment
- CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics
- On a Sparse Shortcut Topology of Artificial Neural Networks
- Pre-train to Gain: Robust Learning Without Clean Labels
- NNGPT: Rethinking AutoML with Large Language Models
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Triplet-Based Deep Hashing Network for Cross-Modal Retrieval
- Temporally Distributed Networks for Fast Video Semantic Segmentation
- gradSLAM: Automagically differentiable SLAM
- C3 Framework: An Open-source PyTorch Code for Crowd Counting
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- IDSplat: Instance-Decomposed 3D Gaussian Splatting for Driving Scenes
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Visual Imitation Made Easy
- Experimental insights into data augmentation techniques for deep learning-based multimode fiber imaging: limitations and success
- Prototype-supervised Adversarial Network for Targeted Attack of Deep Hashing
- Combating Ambiguity for Hash-code Learning in Medical Instance Retrieval
- Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
- Deep fusion of multi-view and multimodal representation of ALS point cloud for 3D terrain scene recognition
- Towards Biologically Plausible Convolutional Networks
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- LookSharp: Attention Entropy Minimization for Test-Time Adaptation
- FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework
- Initializing ReLU networks in an expressive subspace of weights
- Cross-individual Recognition of Emotions by a Dynamic Entropy based on Pattern Learning with EEG features
- A Climatology of Quasi-Linear Convective Systems and Their Hazards in the United States
- Enhanced Center Coding for Cell Detection with Convolutional Neural\n Networks
- RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- Using MLIR Transform to Design Sliced Convolution Algorithm
- Decoupled Audio-Visual Dataset Distillation
- APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
- The Rapid Growth of AI Foundation Model Usage in Science
- Responses to Critiques on Machine Learning of Criminality Perceptions (Addendum of arXiv:1611.04135)
- Neural Architecture Search without Training
- Separation of time scales and direct computation of weights in deep\n neural networks
- LassoLayer: Nonlinear Feature Selection by Switching One-to-one Links
- Multiple Code Hashing for Efficient Image Retrieval
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- Spatial-Temporal Dynamic Graph Attention Networks for Ride-hailing Demand Prediction
- Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
- An End-to-End Approach to Automatic Speech Assessment for Cantonese-speaking People with Aphasia
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path
- Feasibility of Embodied Dynamics Based Bayesian Learning for Continuous Pursuit Motion Control of Assistive Mobile Robots in the Built Environment
- MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration
- Using Topological Framework for the Design of Activation Function and Model Pruning in Deep Neural Networks
- Automatic Recognition of Coal and Gangue based on Convolution Neural Network
- A Machine Learning-Driven Solution for Denoising Inertial Confinement Fusion Images
- Evolution Strategies at the Hyperscale
- Compositional Explanations of Neurons
- Learning from Noisy Labels with Distillation
- Spatial-Temporal Transformer Networks for Traffic Flow Forecasting
- Learning without Forgetting
- UniFormer: Unifying Convolution and Self-Attention for Visual Recognition
- FVQA: Fact-Based Visual Question Answering
- Network Pruning for Low-Rank Binary Indexing
- Breast Tumor Classification and Segmentation using Convolutional Neural Networks
- CathAI: Fully Automated Interpretation of Coronary Angiograms Using\n Neural Networks
- GLOBE: Accurate and Generalizable PDE Surrogates using Domain-Inspired Architectures and Equivariances
- A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
- Learning to Navigate Using Mid-Level Visual Priors
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- Med3D: Transfer Learning for 3D Medical Image Analysis
- Task-driven Semantic Coding via Reinforcement Learning
- Incidental Scene Text Understanding: Recent Progresses on ICDAR 2015 Robust Reading Competition Challenge 4
- Syn2Real: Forgery Classification via Unsupervised Domain Adaptation
- A Novel Perspective to Zero-shot Learning: Towards an Alignment of Manifold Structures via Semantic Feature Expansion
- Clustered Object Detection in Aerial Images
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Sigil: Server-Enforced Watermarking in U-Shaped Split Federated Learning via Gradient Injection
- VisEvent: Reliable Object Tracking via Collaboration of Frame and Event Flows
- Deep Discriminative Representation Learning with Attention Map for Scene Classification
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- XNOR-Net++: Improved Binary Neural Networks
- Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
- DCNNs: A Transfer Learning comparison of Full Weapon Family threat detection for Dual-Energy X-Ray Baggage Imagery
- Convolutional Neural Networks for Large-Scale Remote-Sensing Image Classification
- EVAA—Exchange Vanishing Adversarial Attack on LiDAR Point Clouds in Autonomous Vehicles
- Cross-Learning from Scarce Data via Multi-Task Constrained Optimization
- Hamiltonian Neural Networks
- Learning proofs for the classification of nilpotent semigroups
- Region-Point Joint Representation for Effective Trajectory Similarity Learning
- Surgical Visual Domain Adaptation: Results from the MICCAI 2020 SurgVisDom Challenge
- Compositional Generalization for Primitive Substitutions
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Sequence to Sequence Learning with Neural Networks
- IntPhys: A Benchmark for Visual Intuitive Physics Reasoning
- BinaryConnect: Training Deep Neural Networks with binary weights during propagations
- Questions to Guide the Future of Artificial Intelligence Research
- Recommendation or Discrimination?: Quantifying Distribution Parity in Information Retrieval Systems
- Generic decoding of seen and imagined objects using hierarchical visual features
- Information Theory-Guided Heuristic Progressive Multi-View Coding
- To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from Videos
- Learning a Deep Embedding Model for Zero-Shot Learning
- Multi-Interactive Attention Network for Fine-grained Feature Learning in CTR Prediction
- TinyCNN: A Tiny Modular CNN Accelerator for Embedded FPGA
- MineGAN: effective knowledge transfer from GANs to target domains with few images
- A Gaussian Process perspective on Convolutional Neural Networks
- Alpha-Integration Pooling for Convolutional Neural Networks
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
- LaSOT: A High-quality Large-scale Single Object Tracking Benchmark
- Learning to Hash with Graph Neural Networks for Recommender Systems
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- Sparsity-Control Ternary Weight Networks
- Quaternion Convolutional Neural Networks
- A Survey of Deep Reinforcement Learning in Video Games
- 1st Place Solution for Waymo Open Dataset Challenge -- 3D Detection and Domain Adaptation
- Understanding the robustness of deep neural network classifiers for\n breast cancer screening
- LEA-Net: Layer-wise External Attention Network for Efficient Color Anomaly Detection
- Calibrated Adversarial Training
- Separation and Concentration in Deep Networks
- Beyond Dents and Scratches: Logical Constraints in Unsupervised Anomaly Detection and Localization
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation
- Balanced Symmetric Cross Entropy for Large Scale Imbalanced and Noisy Data
- DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars
- Credit scoring using neural networks and SURE posterior probability calibration
- Glance-and-Gaze Vision Transformer
- Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics
- Efficient Transfer Bayesian Optimization with Auxiliary Information
- Discovery and Separation of Features for Invariant Representation Learning
- Accelerating Robustness Verification of Deep Neural Networks Guided by Target Labels
- Batch Normalization Preconditioning for Neural Network Training
- Learning Accurate, Comfortable and Human-like Driving
- End-to-End Wireframe Parsing
- Walsh-Hadamard Variational Inference for Bayesian Deep Learning
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Tackling Over-Smoothing for General Graph Convolutional Networks
- Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
- Multi-level Texture Encoding and Representation (MuLTER) based on Deep Neural Networks
- Dense neural networks as sparse graphs and the lightning initialization
- Curriculum Learning: A Survey
- A Comprehensive Overhaul of Feature Distillation
- Can deep learning help you find the perfect match?
- Estimating or Propagating Gradients Through Stochastic Neurons
- HexCNN: A Framework for Native Hexagonal Convolutional Neural Networks
- Differentiable Rendering: A Survey
- Exemplar Loss for Siamese Network in Visual Tracking
- Extremal Contours: Gradient-driven contours for compact visual attribution
- Effective Regularization Through Loss-Function Metalearning
- Deep Image Prior
- Quantile regression with deep ReLU Networks: Estimators and minimax rates
- A study on using image based machine learning methods to develop the surrogate models of stamp forming simulations
- FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning
- Blow: a single-scale hyperconditioned flow for non-parallel raw-audio\n voice conversion
- Network Moments: Extensions and Sparse-Smooth Attacks
- Global Image Sentiment Transfer
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a\n Survey
- Deep Learning Based Defect Detection for Solder Joints on Industrial X-Ray Circuit Board Images
- A Multicollinearity-Aware Signal-Processing Framework for Cross-β Identification via X-ray Scattering of Alzheimer's Tissue
- Multiscale Vision Transformers
- Machine Learning Framework for Efficient Prediction of Quantum Wasserstein Distance
- Learning Straight Flows: Variational Flow Matching for Efficient Generation
- Deep Spatial Pyramid: The Devil is Once Again in the Details
- Generalizing to unseen domains via distribution matching
- Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
- Compiling to linear neurons
- The modified Physics-Informed Hybrid Parallel Kolmogorov--Arnold and Multilayer Perceptron Architecture with domain decomposition
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Physics-informed Machine Learning for Static Friction Modeling in Robotic Manipulators Based on Kolmogorov-Arnold Networks
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
- Classification of motor faults based on transmission coefficient and reflection coefficient of omni-directional antenna using DCNN
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
- Multi-task Learning with Attention for End-to-end Autonomous Driving
- Learning Memory-guided Normality for Anomaly Detection
- Self-Guided Adaptation: Progressive Representation Alignment for Domain Adaptive Object Detection
- Unsupervised Domain Expansion from Multiple Sources
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
- A Comprehensive Study of Deep Video Action Recognition
- Convolutional neural networks decode visual stimulus positions from local field potentials on the mouse cortex
- C3AE: Exploring the Limits of Compact Model for Age Estimation
- State-Aware Tracker for Real-Time Video Object Segmentation
- PP-LCNet: A Lightweight CPU Convolutional Neural Network
- ProbSelect: Stochastic Client Selection for GPU-Accelerated Compute Devices in the 3D Continuum
- HipKittens: Fast and Furious AMD Kernels
- Advancing credibility and transparency in brain-to-image reconstruction research: Reanalysis of Koide-Majima, Nishimoto, and Majima (Neural Networks, 2024)
- Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
- Quality-Aware Network for Human Parsing
- Integrating Large Circular Kernels into CNNs through Neural Architecture Search
- Under the Skin of Foundation NFT Auctions
- Deep Semantic Hashing with Generative Adversarial Networks
- Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning
- Sample Prior Guided Robust Model Learning to Suppress Noisy Labels
- Edge of chaos as a guiding principle for modern neural network training
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search
- AuxBlocks: Defense Adversarial Example via Auxiliary Blocks
- A Hybrid Autoencoder-Transformer Model for Robust Day-Ahead Electricity Price Forecasting under Extreme Conditions
- Minimum Width of Deep Narrow Networks for Universal Approximation
- Learning Performance Optimization for Edge AI System with Time and Energy Constraints
- Refactoring Neural Networks for Verification
- A Survey of Machine Learning Methods and Challenges for Windows Malware Classification
- Inter-Image Communication for Weakly Supervised Localization
- Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis
- Anti-aliasing Deep Image Classifiers using Novel Depth Adaptive Blurring and Activation Function
- TopoAct: Visually Exploring the Shape of Activations in Deep Learning
- CBIR using Pre-Trained Neural Networks
- The Low-Rank Simplicity Bias in Deep Networks
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Real-Time Face and Landmark Localization for Eyeblink Detection
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- Escaping the Big Data Paradigm with Compact Transformers
- A Visual Perception-Based Tunable Framework and Evaluation Benchmark for H.265/HEVC ROI Encryption
- QDCNN: Quantum Dilated Convolutional Neural Network
- Qu-ANTI-zation: Exploiting Quantization Artifacts for Achieving Adversarial Outcomes
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- Radio AGN feedback sustains quiescence only in a minority of massive galaxies
- Global Multiple Extraction Network for Low-Resolution Facial Expression Recognition
- Adversarial Attacks Beyond the Image Space
- Deep Image: Scaling up Image Recognition
- Toward Better Generalization in Few-Shot Learning through the Meta-Component Combination
- Global Context Networks
- Context-aware Learned Mesh-based Simulation via Trajectory-Level Meta-Learning
- NeuroFlex: Column-Exact ANN-SNN Co-Execution Accelerator with Cost-Guided Scheduling
- Beta Distribution Learning for Reliable Roadway Crash Risk Assessment
- Nowcast3D: Reliable precipitation nowcasting via gray-box learning
- Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
- Accelerating scientific discovery with the common task framework
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- Distribution-Aware Tensor Decomposition for Compression of Convolutional Neural Networks
- Extended Physics Informed Neural Network for Hyperbolic Two-Phase Flow in Porous Media
- Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels
- SAAIPAA: Optimizing aspect-angles-invariant physical adversarial attacks on SAR target recognition models
- Joint Optimization of DNN Model Caching and Request Routing in Mobile Edge Computing
- DKN: Deep Knowledge-Aware Network for News Recommendation
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- A novel method for identifying the deep neural network model with the Serial Number
- On Complex Valued Convolutional Neural Networks
- A survey on human-aware robot navigation
- Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net
- SinReQ: Generalized Sinusoidal Regularization for Low-Bitwidth Deep\n Quantized Training
- Detecting cutaneous basal cell carcinomas in ultra-high resolution and weakly labelled histopathological images
- Deeper and Wider Siamese Networks for Real-Time Visual Tracking
- SPEC2: SPECtral SParsE CNN Accelerator on FPGAs
- An open access repository of images on plant health to enable the development of mobile disease diagnostics
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Efficient Detection and Characterization of Targets of Natural Selection Using Transfer Learning
- MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
- Condition Numbers and Eigenvalue Spectra of Shallow Networks on Spheres
- Neural Network Interoperability Across Platforms
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster
- Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Parameter Interpolation Adversarial Training for Robust Image Classification
- Diluting Restricted Boltzmann Machines
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- Extended dynamic mode decomposition with dictionary learning: A data-driven adaptive spectral decomposition of the Koopman operator
- Image Hashing via Cross-View Code Alignment in the Age of Foundation Models
- Data-Augmented Deep Learning for Downhole Depth Sensing and Field Validation
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- Exploring Landscapes for Better Minima along Valleys
- Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
- Spiking Neural Networks: The Future of Brain-Inspired Computing
- Comparative Analysis of Deep Learning Models for Olive Tree Crown and Shadow Segmentation Towards Biovolume Estimation
- EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- Generative Artificial Intelligence for Air Shower Simulation
- Learning Geometry: A Framework for Building Adaptive Manifold Models through Metric Optimization
- A Review of AI-Driven Approaches for Nanoscale Heat Conduction and Radiation
- DARTS: A Drone-Based AI-Powered Real-Time Traffic Incident Detection System
- Data-driven discovery of thermal illusions through latent-space geometry
- CHIPSIM: A Co-Simulation Framework for Deep Learning on Chiplet-Based Systems
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Feedback Alignment Meets Low-Rank Manifolds: A Structured Recipe for Local Learning
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- Adversarial Domain Randomization
- Beyond Data Scarcity Optimizing R3GAN for Medical Image Generation from Small Datasets
- Deep Multi-Scale Features Learning for Distorted Image Quality Assessment
- Hybrid Models for Open Set Recognition
- Incremental Methods for Weakly Convex Optimization
- Deep Learning Algorithms with Applications to Video Analytics for A Smart City: A Survey
- Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements
- Winner-Take-All Autoencoders
- Channel Pruning via Optimal Thresholding
- Hypergraph Convolution and Hypergraph Attention
- Deep Learning for Anomaly Detection: A Review
- Person Retrieval in Surveillance Video using Height, Color and Gender
- Medical Concept Representation Learning from Electronic Health Records and its Application on Heart Failure Prediction
- Sharpness-aware Quantization for Deep Neural Networks
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- MetaFormer is Actually What You Need for Vision
- Understanding Information Processing in Human Brain by Interpreting Machine Learning Models
- Affinity and Diversity: Quantifying Mechanisms of Data Augmentation
- Deep Learning to Ternary Hash Codes by Continuation
- Camera-aware Proxies for Unsupervised Person Re-Identification
- Joint Object and State Recognition using Language Knowledge
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Deep causal representation learning for unsupervised domain adaptation
- Deep Network Classification by Scattering and Homotopy Dictionary Learning
- Uncertainty-Aware Attention for Reliable Interpretation and Prediction
- A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
- CNN depth analysis with different channel inputs for Acoustic Scene Classification
- Amadeus: Scalable, Privacy-Preserving Live Video Analytics
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- Automatic Target Recognition on Synthetic Aperture Radar Imagery: A Survey
- Malleable 2.5D Convolution: Learning Receptive Fields along the Depth-axis for RGB-D Scene Parsing
- Growing a Brain: Fine-Tuning by Increasing Model Capacity
- Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation
- Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
- Adaptive Gradient for Adversarial Perturbations Generation
- Recent Deep Semi-supervised Learning Approaches and Related Works
- Bridging the Gap Between Spectral and Spatial Domains in Graph Neural Networks
- Deep Learning for the Classification of Lung Nodules
- Deep High-Resolution Representation Learning for Visual Recognition
- Discovering and Explaining the Representation Bottleneck of DNNs
- Face Completion with Semantic Knowledge and Collaborative Adversarial Learning
- MEAL: Multi-Model Ensemble via Adversarial Learning
- Regularizing Explanations in Bayesian Convolutional Neural Networks
- A Spatial-Temporal Decomposition Based Deep Neural Network for Time\n Series Forecasting
- Probabilistic Trust Intervals for Out of Distribution Detection
- Accelerating 3D Deep Learning with PyTorch3D
- Where am I looking at? Joint Location and Orientation Estimation by Cross-View Matching
- Neural Network Architectures for Location Estimation in the Internet of Things
- Multi Receptive Field Network for Semantic Segmentation
- Learning Neural Network Classifiers with Low Model Complexity
- Conditional Deep Learning for Energy-Efficient and Enhanced Pattern\n Recognition
- Tango: A Deep Neural Network Benchmark Suite for Various Accelerators
- Rethinking CNN Models for Audio Classification
- Character-level Chinese Writer Identification using Path Signature Feature, DropStroke and Deep CNN
- Voxceleb: Large-scale speaker verification in the wild
- Replay anti-spoofing countermeasure based on data augmentation with post selection
- Learning the Redundancy-free Features for Generalized Zero-Shot Object Recognition
- Gated Channel Transformation for Visual Recognition
- Variable Selection with Rigorous Uncertainty Quantification using Deep Bayesian Neural Networks: Posterior Concentration and Bernstein-von Mises Phenomenon
- Physical Attribute Prediction Using Deep Residual Neural Networks
- Weakly supervised object detection using pseudo-strong labels
- Universal Perturbation Attack Against Image Retrieval
- N-GCN: Multi-scale Graph Convolution for Semi-supervised Node\n Classification
- Temporal-Clustering Invariance in Irregular Healthcare Time Series
- RL-GAN-Net: A Reinforcement Learning Agent Controlled GAN Network for\n Real-Time Point Cloud Shape Completion
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Deep CTR Prediction in Display Advertising
- Dynamic Curriculum Learning for Imbalanced Data Classification
- Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio Estimation
- Harnessing Deep Neural Networks with Logic Rules
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Reclaiming AI as a theoretical tool for cognitive science
- FINCHES: A Computational Framework for Predicting Intermolecular Interactions in Intrinsically Disordered Proteins
- Hyperparameter Search in Machine Learning
- Fabric Surface Characterization: Assessment of Deep Learning-based Texture Representations Using a Challenging Dataset
- Representation Extraction and Deep Neural Recommendation for Collaborative Filtering
- On hyperparameter optimization of machine learning algorithms: Theory and practice
- NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm
- A New Compensatory Genetic Algorithm-Based Method for Effective Compressed Multi-function Convolutional Neural Network Model Selection with Multi-Objective Optimization
- Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
- Dynamic Graph CNN for Learning on Point Clouds
- Global Filter Networks for Image Classification
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- Rethinking the Hyperparameters for Fine-tuning
- Chargrid-OCR: End-to-end Trainable Optical Character Recognition for Printed Documents using Instance Segmentation
- Places: An Image Database for Deep Scene Understanding
- Real World Robustness from Systematic Noise
- Object Detection in 20 Years: A Survey
- Cnns in land cover mapping with remote sensing imagery: a review and meta-analysis
- Open DNN Box by Power Side-Channel Attack
- A Statistician Teaches Deep Learning
- When Follow is Just One Click Away: Understanding Twitter Follow Behavior in the 2016 U.S. Presidential Election
- Adapting Auxiliary Losses Using Gradient Similarity
- A Deep Multi-task Learning Approach to Skin Lesion Classification
- Adaptive Selection of Deep Learning Models on Embedded Systems
- Dual-level Semantic Transfer Deep Hashing for Efficient Social Image Retrieval
- An interpretable deep hierarchical semantic convolutional neural network for lung nodule malignancy classification
- Continual Learning for Robotics: Definition, Framework, Learning\n Strategies, Opportunities and Challenges
- Spotting insects from satellites: modeling the presence of Culicoides\n imicola through Deep CNNs
- MeshingNet: A New Mesh Generation Method based on Deep Learning
- Ingredient-guided multi-modal interaction and refinement network for RGB-D food nutrition assessment
- Data augmentation using learned transformations for one-shot medical image segmentation
- Automatic Detection of Cerebral Microbleeds From MR Images via 3D Convolutional Neural Networks
- AggNet: Deep Learning From Crowds for Mitosis Detection in Breast Cancer Histology Images
- Fast Convolutional Neural Network Training Using Selective Data Sampling: Application to Hemorrhage Detection in Color Fundus Images
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- For Manifold Learning, Deep Neural Networks can be Locality Sensitive Hash Functions
- Strategies for Pre-training Graph Neural Networks
- Optimal Feature Transport for Cross-View Image Geo-Localization
- An Efficient Multi-Scale Attention two-stream inflated 3D ConvNet network for cattle behavior recognition
- Shuffled Patch-Wise Supervision for Presentation Attack Detection
- Towards All-around Knowledge Transferring: Learning From Task-irrelevant Labels
- Unsupervised Learning of Solutions to Differential Equations with Generative Adversarial Networks
- Deep Long-Tailed Learning: A Survey
- Batch Group Normalization
- Split Slice Training Augmentation and Hyperparameter Tuning of RAKI\n Networks for Simultaneous Multi-Slice Reconstruction
- Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
- FlipReID: Closing the Gap between Training and Inference in Person Re-Identification
- Recent Advances in Neural Question Generation
- Efficient Semantic Scene Completion Network with Spatial Group Convolution
- MUREL: Multimodal Relational Reasoning for Visual Question Answering
- Facial Key Points Detection using Deep Convolutional Neural Network - NaimishNet
- Lifelong Graph Learning
- A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration Framework
- EKT: Exercise-aware Knowledge Tracing for Student Performance Prediction
- Learning Spatiotemporal Features of Ride-sourcing Services with Fusion Convolutional Network
- When Residual Learning Meets Dense Aggregation: Rethinking the Aggregation of Deep Neural Networks
- Global Texture Enhancement for Fake Face Detection in the Wild
- A Simple Semi-Supervised Learning Framework for Object Detection
- FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval
- Deep Residual Learning for Image Recognition
- BigGAN-based Bayesian reconstruction of natural images from human brain activity
- Deep 3D-to-2D Watermarking: Embedding Messages in 3D Meshes and Extracting Them from 2D Renderings
- HYPER-SNN: Towards Energy-efficient Quantized Deep Spiking Neural Networks for Hyperspectral Image Classification
- Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data Set
- Does Data Augmentation Benefit from Split BatchNorms
- Recognition of European mammals and birds in camera trap images using deep neural networks
- Do CNNs Encode Data Augmentations?
- Instance-Aware Predictive Navigation in Multi-Agent Environments
- How Transferable are CNN-based Features for Age and Gender\n Classification?
- DeepIris: Iris Recognition Using A Deep Learning Approach
- A system of different layers of abstraction for artificial intelligence
- Learning Optimal Data Augmentation Policies via Bayesian Optimization for Image Classification Tasks
- TorchIO: A Python library for efficient loading, preprocessing, augmentation and patch-based sampling of medical images in deep learning
- Sentiment Analysis from Images of Natural Disasters
- Distribution Alignment: A Unified Framework for Long-tail Visual Recognition
- Methods for Interpreting and Understanding Deep Neural Networks
- Robust Sparse Linear Discriminant Analysis
- Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
- DiabDeep: Pervasive Diabetes Diagnosis based on Wearable Medical Sensors and Efficient Neural Networks
- The Rise of Radar for Autonomous Vehicles: Signal Processing Solutions and Future Research Directions
- Artificial Intelligence in the Rising Wave of Deep Learning: The Historical Path and Future Outlook [Perspectives]
- Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing
- Skeleton-Based Action Recognition With Gated Convolutional Neural Networks
- Perceptual Image Hashing for Content Authentication Based on Convolutional Neural Network With Multiple Constraints
- Omnidirectional Image Quality Assessment by Distortion Discrimination Assisted Multi-Stream Network
- Tensorizing Neural Networks
- GST: Group-Sparse Training for Accelerating Deep Reinforcement Learning
- Meta-Regularization by Enforcing Mutual-Exclusiveness
- Characterizing machine learning process: A maturity framework
- Circumventing Outliers of AutoAugment with Knowledge Distillation
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Bi-stream Pose Guided Region Ensemble Network for Fingertip Localization from Stereo Images
- Fold bifurcation identification through scientific machine learning
- Spectral Signatures in Backdoor Attacks
- A Dataset and Application for Facial Recognition of Individual Gorillas in Zoo Environments
- What makes visual place recognition easy or hard?
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Deep Residual Dense U-Net for Resolution Enhancement in Accelerated MRI Acquisition
- DANCE: Enhancing saliency maps using decoys
- Competing Ratio Loss for Discriminative Multi-class Image Classification
- How Important is Weight Symmetry in Backpropagation?
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Learning Robust Representations via Multi-View Information Bottleneck
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Dynamic Origin-Destination Matrix Prediction with Line Graph Neural Networks and Kalman Filter
- An Information Theory-inspired Strategy for Automatic Network Pruning
- Detecting COVID-19 from Breathing and Coughing Sounds using Deep Neural\n Networks
- Classification of Pathological and Normal Gait: A Survey
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Domain Adaptive SiamRPN++ for Object Tracking in the Wild
- Quantized Adam with Error Feedback
- A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning
- Machine Learning Etudes in Conformal Field Theories
- Noise-Sampling Cross Entropy Loss: Improving Disparity Regression Via Cost Volume Aware Regularizer
- Interpretable Deep Convolutional Fuzzy Classifier
- SWNet: Small-World Neural Networks and Rapid Convergence
- O sistema tecnológico digital
- LUTNet: Rethinking Inference in FPGA Soft Logic
- Cloud-based or On-device: An Empirical Study of Mobile Deep Inference
- Modulated binary cliquenet
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Training Deep Nets with Sublinear Memory Cost
- Uformer: A General U-Shaped Transformer for Image Restoration
- Controlling Recurrent Neural Networks by Conceptors
- Pseudo-ISP: Learning Pseudo In-camera Signal Processing Pipeline from A Color Image Denoiser
- Random neural networks in the infinite width limit as Gaussian processes
- Protein identification with deep learning: from abc to xyz
- Self-Supervised Multi-View Learning via Auto-Encoding 3D Transformations
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Measuring the Contribution of Multiple Model Representations in\n Detecting Adversarial Instances
- Dual-Sampling Attention Network for Diagnosis of COVID-19 from Community Acquired Pneumonia
- Efficient Reservoir Computing using Field Programmable Gate Array and Electro-optic Modulation
- Deep Reinforcement Learning with Quantum-inspired Experience Replay
- Multimodal Emotion Recognition for One-Minute-Gradual Emotion Challenge
- LiftPool: Bidirectional ConvNet Pooling
- SoftTriple Loss: Deep Metric Learning Without Triplet Sampling
- A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension
- Exclusive: the most-cited papers of the twenty-first century
- Strategies for Conceptual Change in Convolutional Neural Networks
- PIXOR: Real-time 3D Object Detection from Point Clouds
- Hierarchical Neural Representation of Dreamed Objects Revealed by Brain Decoding with Deep Neural Network Features
- Testing Deep Learning Models for Image Analysis Using Object-Relevant Metamorphic Relations
- Pruning at a Glance: Global Neural Pruning for Model Compression
- Transform-Invariant Convolutional Neural Networks for Image\n Classification and Search
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution
- Contrastive Multiview Coding
- Exploiting Deep Features for Remote Sensing Image Retrieval: A Systematic Investigation
- CLCC: Contrastive Learning for Color Constancy
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
- Learning Channel Inter-dependencies at Multiple Scales on Dense Networks for Face Recognition
- VOLO: Vision Outlooker for Visual Recognition
- The Dynamicist Landscape
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- Multimodal mixing convolutional neural network and transformer for Alzheimer’s disease recognition
- Domain Generalization for Semantic Segmentation: A Survey
- Mitigating Spurious Correlation via Distributionally Robust Learning with Hierarchical Ambiguity Sets
- Self-Supervised Convolutional Subspace Clustering Network
- Stochastic Region Pooling: Make Attention More Expressive
- A Comprehensive Survey of Neural Architecture Search
- GPatt: Fast Multidimensional Pattern Extrapolation with Gaussian Processes
- CDGNet: Class Distribution Guided Network for Human Parsing
- Improving the Authentication with Built-in Camera Protocol Using\n Built-in Motion Sensors: A Deep Learning Solution
- Semi-supervised Domain Adaptation via Minimax Entropy
- Instance Adaptive Self-Training for Unsupervised Domain Adaptation
- Revisiting IM2GPS in the Deep Learning Era
- End-to-end Active Object Tracking and Its Real-world Deployment via Reinforcement Learning
- Adaptive Convolutional ELM For Concept Drift Handling in Online Stream Data
- A Survey on Bayesian Deep Learning
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Calibration and Consistency of Adversarial Surrogate Losses
- CvS: Classification via Segmentation For Small Datasets
- A Sentiment-and-Semantics-Based Approach for Emotion Detection in Textual Conversations
- Towards Using Count-level Weak Supervision for Crowd Counting
- Unsupervised learning of clutter-resistant visual representations from natural videos
- Driver Drowsiness Detection Model Using Convolutional Neural Networks\n Techniques for Android Application
- Deep Frequent Spatial Temporal Learning for Face Anti-Spoofing
- Multi-Region Ensemble Convolutional Neural Network for Facial Expression Recognition
- Digital Twin: Values, Challenges and Enablers
- On the Importance of Visual Context for Data Augmentation in Scene\n Understanding
- Graph neural networks: A review of methods and applications
- Rotated Feature Network for multi-orientation object detection
- Defending against substitute model black box adversarial attacks with the 01 loss
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Noise Contrastive Priors for Functional Uncertainty
- Real-Time Steganalysis for Stream Media Based on Multi-channel Convolutional Sliding Windows
- WSOD2: Learning Bottom-up and Top-down Objectness Distillation for Weakly-supervised Object Detection
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- Confusing Image Quality Assessment: Toward Better Augmented Reality Experience
- High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
- Generalizing from a Few Examples
- ROAM: Recurrently Optimizing Tracking Model
- The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)
- Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices: Design Considerations
- Deep Metric Learning for Practical Person Re-Identification
- Deep Lagrangian Networks: Using Physics as Model Prior for Deep Learning
- An Image Patch is a Wave: Phase-Aware Vision MLP
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis
- Any-Precision Deep Neural Networks
- Learning degraded image classification with restoration data fidelity
- Tomato plant disease classification using Multilevel Feature Fusion with adaptive channel spatial and pixel attention mechanism
- High-Performance Large-Scale Image Recognition Without Normalization
- A Semantics-Guided Class Imbalance Learning Model for Zero-Shot\n Classification
- Cross-modal Zero-shot Hashing
- Separating the Effects of Batch Normalization on CNN Training Speed and Stability Using Classical Adaptive Filter Theory
- SurveilEdge: Real-time Video Query based on Collaborative Cloud-Edge Deep Learning
- PAMS: Quantized Super-Resolution via Parameterized Max Scale
- Random VLAD based Deep Hashing for Efficient Image Retrieval
- RGB-based Semantic Segmentation Using Self-Supervised Depth Pre-Training
- Fine-Tuning Models Comparisons on Garbage Classification for Recyclability
- Improving Adversarial Transferability with Gradient Refining
- Generalized Out-of-Distribution Detection: A Survey
- A comprehensive review of object detection with deep learning
- Learning Various Length Dependence by Dual Recurrent Neural Networks
- Universal Lipschitz Approximation in Bounded Depth Neural Networks
- Image Synthesis with a Single (Robust) Classifier
- Symbol Emergence as an Interpersonal Multimodal Categorization
- Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds
- In-domain representation learning for remote sensing
- Deep feature based rice leaf disease identification using support vector machine
- Sequential Graph Convolutional Network for Active Learning
- Performance Analysis and Characterization of Training Deep Learning Models on Mobile Devices
- Vector-quantized Image Modeling with Improved VQGAN
- Adversarial Domain Adaptation with Domain Mixup
- A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities
- Video Modeling with Correlation Networks
- Reversed Active Learning based Atrous DenseNet for Pathological Image Classification
- Usefulness of interpretability methods to explain deep learning based plant stress phenotyping
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- A Kronecker-factored approximate Fisher matrix for convolution layers
- Adaptive Future Frame Prediction with Ensemble Network
- Meta Feature Modulator for Long-tailed Recognition
- Stack-based Buffer Overflow Detection using Recurrent Neural Networks
- Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising
- Pose-Invariant Embedding for Deep Person Re-Identification
- PlantDoc
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- ActBERT: Learning Global-Local Video-Text Representations
- Deep learning in radiology: An overview of the concepts and a survey of the state of the art with focus on MRI
- End-to-End Blind Image Quality Assessment Using Deep Neural Networks
- Deep Learning Markov Random Field for Semantic Segmentation
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
- DDLSTM: Dual-Domain LSTM for Cross-Dataset Action Recognition
- Learning with Group Invariant Features: A Kernel Perspective
- Intelligent Autofocus
- QuaterNet: A Quaternion-based Recurrent Model for Human Motion
- LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search
- A Simple Pooling-Based Design for Real-Time Salient Object Detection
- LightMOT: Lightweight and anchor-free solution for tracking multiple objects in dense populations
- On Learning Over-parameterized Neural Networks: A Functional Approximation Perspective
- Towards Spatial Variability Aware Deep Neural Networks (SVANN): A\n Summary of Results
- A Deep Neural Network for Audio Classification with a Classifier Attention Mechanism
- Deep High-Resolution Representation Learning for Visual Recognition
- PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
- Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments
- Learning to Localize: A 3D CNN Approach to User Positioning in Massive MIMO-OFDM Systems
- In-N-Out: Pre-Training and Self-Training using Auxiliary Information for Out-of-Distribution Robustness
- Material Recognition for Automated Progress Monitoring using Deep Learning Methods
- Data Extraction from Charts via Single Deep Neural Network
- Towards Rapid and Robust Adversarial Training with One-Step Attacks
- Low Precision Floating-point Arithmetic for High Performance FPGA-based CNN Acceleration
- Universal Person Re-Identification
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- ResNet strikes back: An improved training procedure in timm
- Mandarin tone modeling using recurrent neural networks
- Deep Learning in Robotics: A Review of Recent Research
- Searching for A Robust Neural Architecture in Four GPU Hours
- Half and full solar cell efficiency binning by deep learning on electroluminescence images
- Deep Triplet Quantization
- Kernel-Based Smoothness Analysis of Residual Networks
- Toward Intelligent Sensing: Intermediate Deep Feature Compression
- Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks
- Stabilizing DARTS with Amended Gradient Estimation on Architectural Parameters
- Fantastic Four: Differentiable Bounds on Singular Values of Convolution\n Layers
- DeepACEv2: Automated Chromosome Enumeration in Metaphase Cell Images Using Deep Convolutional Neural Networks
- GenURL: A General Framework for Unsupervised Representation Learning
- Artistic Domain Generalisation Methods are Limited by their Deep\n Representations
- Feedback Attention for Cell Image Segmentation
- Dos and Don'ts of Machine Learning in Computer Security
- Attacking and Defending Machine Learning Applications of Public Cloud
- Natural Adversarial Examples
- Decision-based Universal Adversarial Attack
- Branchy-GNN: a Device-Edge Co-Inference Framework for Efficient Point Cloud Processing
- Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- MDCNN: A multimodal dual-CNN recursive model for fake news detection via audio- and text-based speech emotion recognition
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Can Single Neurons Solve MNIST? The Computational Power of Biological Dendritic Trees
- Improving Object Detection from Scratch via Gated Feature Reuse
- Successive Embedding and Classification Loss for Aerial Image Classification
- In-depth Question classification using Convolutional Neural Networks
- Feature Space Transfer for Data Augmentation
- Going Deeper for Multilingual Visual Sentiment Detection
- SAIA: Split Artificial Intelligence Architecture for Mobile Healthcare System
- Training Robust Deep Neural Networks via Adversarial Noise Propagation
- Salient Instance Segmentation via Subitizing and Clustering
- Deep Modulation Recognition with Multiple Receive Antennas: An End-to-end Feature Learning Approach
- Learning Multi-granular Quantized Embeddings for Large-Vocab Categorical Features in Recommender Systems
- Learning Pixel-level Semantic Affinity with Image-level Supervision for Weakly Supervised Semantic Segmentation
- WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition
- ESFNet: Efficient Network for Building Extraction from High-Resolution Aerial Images
- Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey
- A review of vibration-based damage detection in civil structures: From traditional methods to Machine Learning and Deep Learning applications
- DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image Fusion
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval
- Convolutional Neural Network-based Topology Optimization (CNN-TO) By Estimating Sensitivity of Compliance from Material Distribution
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- How to fine-tune deep neural networks in few-shot learning?
- Exploring the Limits of Language Modeling
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Learning Loss for Test-Time Augmentation
- Improving Network Slimming with Nonconvex Regularization
- Efficient Spatialtemporal Context Modeling for Action Recognition
- Universality of deep convolutional neural networks
- Improved Deep Convolutional Neural Network For Online Handwritten Chinese Character Recognition using Domain-Specific Knowledge
- BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation
- Siamese Natural Language Tracker: Tracking by Natural Language Descriptions with Siamese Trackers
- Multiresolution Convolutional Autoencoders
- Deep Convolutional Neural Network for Inverse Problems in Imaging
- SAWNet: A Spatially Aware Deep Neural Network for 3D Point Cloud Processing
- Neuroprosthesis for Decoding Speech in a Paralyzed Person with Anarthria
- Self-training with Noisy Student improves ImageNet classification
- Neural Networks Trained on Natural Scenes Exhibit Gestalt Closure
- Dank Learning: Generating Memes Using Deep Neural Networks
- Improving Prognostic Performance in Resectable Pancreatic Ductal Adenocarcinoma using Radiomics and Deep Learning Features Fusion in CT Images
- BlockDrop: Dynamic Inference Paths in Residual Networks
- When Person Re-identification Meets Changing Clothes
- Deep Facial Expression Recognition: A Survey
- StarGAN v2: Diverse Image Synthesis for Multiple Domains
- Deep Flow Collaborative Network for Online Visual Tracking
- An End-to-End Network for Panoptic Segmentation
- Continual World: A Robotic Benchmark For Continual Reinforcement\n Learning
- Combining Markov Random Fields and Convolutional Neural Networks for Image Synthesis
- Orthogonal Gradient Descent for Continual Learning
- Deep Adaptive Wavelet Network
- Bidirectional Mapping Generative Adversarial Networks for Brain MR to PET Synthesis
- A Survey of FPGA-Based Neural Network Accelerator
- Revisiting Self-Supervised Visual Representation Learning
- Deep Quaternion Features for Privacy Protection
- Unsupervised Learning of Invariant Representations in Hierarchical\n Architectures
- Room Geometry Estimation from Room Impulse Responses using Convolutional Neural Networks
- Learning to Hash for Indexing Big Data - A Survey
- CLDA: Contrastive Learning for Semi-Supervised Domain Adaptation
- Zero-shot World Models Are Developmentally Efficient Learners
- GhostNet: More Features From Cheap Operations
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- DensePoint: Learning Densely Contextual Representation for Efficient Point Cloud Processing
- The Devil is in the Boundary: Exploiting Boundary Representation for Basis-based Instance Segmentation
- Reverse Transfer Learning: Can Word Embeddings Trained for Different NLP\n Tasks Improve Neural Language Models?
- Novel and Effective CNN-Based Binarization for Historically Degraded As-built Drawing Maps
- Multi-label Iterated Learning for Image Classification with Label Ambiguity
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
- HAMBox: Delving into Online High-quality Anchors Mining for Detecting Outer Faces
- Instance-weighted Central Similarity for Multi-label Image Retrieval
- Pairwise Teacher-Student Network for Semi-Supervised Hashing
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for\n Road Construction Safety Applications
- Stanza: Layer Separation for Distributed Training in Deep Learning
- The OoO VLIW JIT Compiler for GPU Inference
- You Only Look & Listen Once: Towards Fast and Accurate Visual Grounding
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent\n Method for Learning Sparse Weight Binarized Activation Neural Networks
- Operational evaluation of data-driven forest fire forecasting models
- Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
- Morphological Network: How Far Can We Go with Morphological Neurons?
- Spatial Pyramid Convolutional Neural Network for Social Event Detection in Static Image
- Data-Efficient Electromagnetic Surrogate Solver Through Dissipative Relaxation Transfer Learning
- Intelligence Primer
- Salient Facial Features from Humans and Deep Neural Networks
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- S2-MLP: Spatial-Shift MLP Architecture for Vision
- Distilling Image Classifiers in Object Detectors
- Content-Aware Convolutional Neural Networks
- Involution: Inverting the Inherence of Convolution for Visual Recognition
- AQD: Towards Accurate Fully-Quantized Object Detection
- Multi-Target Embodied Question Answering
- PUNCH: Positive UNlabelled Classification based information retrieval in\n Hyperspectral images
- Deep Virtual Networks for Memory Efficient Inference of Multiple Tasks
- Wavelet based edge feature enhancement for convolutional neural networks
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian\n Augmentation
- Classification of Motor Imagery EEG Signals by Using a Divergence Based Convolutional Neural Network
- Dropout as a Bayesian Approximation: Appendix
- Adaptive Normalized Risk-Averting Training For Deep Neural Networks
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- EIRES:Training-free AI-Generated Image Detection via Edit-Induced Reconstruction Error Shift
- DRIP: Dynamic patch Reduction via Interpretable Pooling
- Understanding Convolutional Neural Networks with Information Theory: An Initial Exploration
- Region Comparison Network for Interpretable Few-shot Image Classification
- Detection Defense Against Adversarial Attacks with Saliency Map
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- Spatial prediction of apartment rent using regression-based and machine learning-based approaches with a large dataset
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- Representation Learning: A Review and New Perspectives
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Synergizing chemical and AI communities for advancing laboratories of the future
- Derivation and Analysis of Fast Bilinear Algorithms for Convolution
- Audio-visual Representation Learning for Anomaly Events Detection in Crowds
- Taming Visually Guided Sound Generation
- ResNet: Enabling Deep Convolutional Neural Networks through Residual Learning
- Seeking Salient Facial Regions for Cross-Database Micro-Expression Recognition
- Multi-view Low-rank Preserving Embedding: A Novel Method for Multi-view Representation
- A representer theorem for deep neural networks
- LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
- TAda! Temporally-Adaptive Convolutions for Video Understanding
- Ordinal Pooling Networks: For Preserving Information over Shrinking Feature Maps
- Digital video microscopy enhanced by deep learning
- SAND-mask: An Enhanced Gradient Masking Strategy for the Discovery of Invariances in Domain Generalization
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- Deeply Learned Spectral Total Variation Decomposition
- Deep Learning for EEG motor imagery classification based on multi-layer CNNs feature fusion
- Deeply-learned and spatial–temporal feature engineering for human action understanding
- A backdoor attack against LSTM-based text classification systems
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Image Categorization and Search via a GAT Autoencoder and Representative Models
- Hierarchically Robust Representation Learning
- Learned Dual-View Reflection Removal
- MRI-Based Brain Tumor Classification Using Ensemble of Deep Features and Machine Learning Classifiers
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward\n Networks? -- A Neural Tangent Kernel Perspective
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- ATP-Net: An Attention-based Ternary Projection Network For Compressed Sensing
- The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- Privacy-Preserving Semantic Communication over Wiretap Channels with Learnable Differential Privacy
- Training a Convolutional Neural Network for Appearance-Invariant Place Recognition
- Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
- Intrusion Detection: Machine Learning Baseline Calculations for Image\n Classification
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Deep Learning: A Critical Appraisal
- Multiple Convolutional Features in Siamese Networks for Object Tracking
- Multiclass Burn Wound Image Classification Using Deep Convolutional Neural Networks
- Quanvolutional Neural Networks for Pneumonia Detection: An Efficient Quantum-Assisted Feature Extraction Paradigm
- Quantum Machine Learning for Image Classification: A Hybrid Model of Residual Network with Quantum Support Vector Machine
- Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-grained Image Recognition
- Top-Down Semantic Refinement for Image Captioning
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
- An Improved Analysis of Stochastic Gradient Descent with Momentum
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- Image Restoration Using Deep Regulated Convolutional Networks
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- From Black-box to Causal-box: Towards Building More Interpretable Models
- Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
- Task-Adaptive Neural Network Search with Meta-Contrastive Learning
- A Dilated Inception Network for Visual Saliency Prediction
- Widening and Squeezing: Towards Accurate and Efficient QNNs
- EBOP MAVEN: A machine learning model to estimate the input parameters for analytic fitting of detached eclipsing binary light curves
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Bridging the gap to real-world language-grounded visual concept learning
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- 3rd Place Solution to ICCV LargeFineFoodAI Retrieval
- FINE Samples for Learning with Noisy Labels
- Parallelization Techniques for Verifying Neural Networks
- Long-tailed Species Recognition in the NACTI Wildlife Dataset
- Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
- One-pixel Signature: Characterizing CNN Models for Backdoor Detection
- Stand-Alone Self-Attention in Vision Models
- Domain Generalization with MixStyle
- Communication-Efficient Federated Distillation
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- Graph Neural Regularizers for PDE Inverse Problems
- Memory Constrained Dynamic Subnetwork Update for Transfer Learning
- HybridSOMSpikeNet: A Deep Model with Differentiable Soft Self-Organizing Maps and Spiking Dynamics for Waste Classification
- Pose Augmentation: Class-agnostic Object Pose Transformation for Object Recognition
- A Semantic Loss Function for Deep Learning with Symbolic Knowledge
- Anderson-type acceleration method for Deep Neural Network optimization
- SAG-GAN: Semi-Supervised Attention-Guided GANs for Data Augmentation on Medical Images
- Machine learning identification of fractional-order vortex beam diffraction process
- Rethinking Cross-lingual Gaps from a Statistical Viewpoint
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Stealing Neural Networks via Timing Side Channels
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- TransTailor: Pruning the Pre-trained Model for Improved Transfer Learning
- Advances in Inference and Representation for Simultaneous Localization\n and Mapping
- ARC: A Vision-based Automatic Retail Checkout System
- Spatial Attention Point Network for Deep-learning-based Robust Autonomous Robot Motion Generation
- On the Power Saving in High-Speed Ethernet-based Networks for Supercomputers and Data Centers
- Study of Training Dynamics for Memory-Constrained Fine-Tuning
- BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
- Precise classification of low quality G-banded Chromosome Images by reliability metrics and data pruning classifier
- Entity Embeddings of Categorical Variables
- A Semi-Parametric Estimation Method for the Quantile Spectrum with an Application to Earthquake Classification Using Convolutional Neural Network
- Layer-Parallel Training of Deep Residual Neural Networks
- Accelerated WGAN update strategy with loss change rate balancing
- A Unified Perspective on Optimization in Machine Learning and Neuroscience: From Gradient Descent to Neural Adaptation
- Lipschitz regularity of deep neural networks: analysis and efficient estimation
- Data-Driven Short-Term Voltage Stability Assessment Based on Spatial-Temporal Graph Convolutional Network
- Image augmentation with invertible networks in interactive satellite image change detection
- Visual Space Optimization for Zero-shot Learning
- Skin Cancer Recognition using Deep Residual Network
- Data-Driven Analysis of Intersectional Bias in Image Classification: A Framework with Bias-Weighted Augmentation
- Rethinking ResNets: Improved Stacking Strategies With High Order Schemes
- Privacy Inference Attacks and Defenses in Cloud-based Deep Neural Network: A Survey
- End-to-End Spoken Language Translation
- Towards In-Situ Failure Assessment: Deep Learning on DIC Results for Laminated Composites
- Large batch size training of neural networks with adversarial training and second-order information
- Learning from N-Tuple Data with M Positive Instances: Unbiased Risk Estimation and Theoretical Guarantees
- Deep Learning-Based Human Pose Estimation: A Survey
- MCANet: A Coherent Multimodal Collaborative Attention Network for Advanced Modulation Recognition in Adverse Noisy Environments
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Application of Deep Neural Networks to assess corporate Credit Rating
- Decoding Dynamic Visual Experience from Calcium Imaging via Cell-Pattern-Aware Pretraining
- Provable Generalization Bounds for Deep Neural Networks with Momentum-Adaptive Gradient Dropout
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers
- Deep Photovoltaic Nowcasting
- Modeling and Analysis of Energy Harvesting and Smart Grid-Powered Wireless Communication Networks: A Contemporary Survey
- Swin Transformer V2: Scaling Up Capacity and Resolution
- An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity
- Morphology-Aware KOA Classification: Integrating Graph Priors with Vision Models
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- DRAMNet: Authentication based on Physical Unique Features of DRAM Using Deep Convolutional Neural Networks
- UMIche: A UMI-centric analysis platform for enhancing molecular quantification accuracy in bulk and single-cell sequencing
- Becoming episodic: The Development of Objectivity
- Application of CNN to a fine segmented scintillator detector for a single particle and neutrino-nucleon event
- Deep Lesion Tracker: Monitoring Lesions in 4D Longitudinal Imaging Studies
- Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
- Symmetries in PAC-Bayesian Learning
- End-to-end Phoneme Sequence Recognition using Convolutional Neural\n Networks
- Quadratic Autoencoder (Q-AE) for Low-dose CT Denoising
- Themis: Fair and Efficient GPU Cluster Scheduling
- Fine-tuning Flow Matching Generative Models with Intermediate Feedback
- WP-CrackNet: A Collaborative Adversarial Learning Framework for End-to-End Weakly-Supervised Road Crack Detection
- Exploring Structural Degradation in Dense Representations for Self-supervised Learning
- Guiding Query Position and Performing Similar Attention for Transformer-Based Detection Heads
- View Adaptive Neural Networks for High Performance Skeleton-based Human Action Recognition
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- RSG: A Simple but Effective Module for Learning Imbalanced Datasets
- PyKale: Knowledge-Aware Machine Learning from Multiple Sources in Python
- FOX-NAS: Fast, On-device and Explainable Neural Architecture Search
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Enabling Data Diversity: Efficient Automatic Augmentation via Regularized Adversarial Training
- Moment-Based Domain Adaptation: Learning Bounds and Algorithms
- Rebooting Neuromorphic Hardware Design -- A Complexity Engineering Approach
- OpenEDS: Open Eye Dataset
- A deep learning approach to detecting volcano deformation from satellite\n imagery using synthetic datasets
- Video Swin Transformer
- CodeNet: Training Large Scale Neural Networks in Presence of Soft-Errors
- ELASTIC: Improving CNNs with Dynamic Scaling Policies
- Deep Variable-Block Chain with Adaptive Variable Selection
- Video-based Human Action Recognition using Deep Learning: A Review
- VM-BeautyNet: A Synergistic Ensemble of Vision Transformer and Mamba for Facial Beauty Prediction
- Deep generative priors for 3D brain analysis
- SimNets: A Generalization of Convolutional Networks
- A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Tutorial on Variational Autoencoders
- Semantic representations emerge in biologically inspired ensembles of cross-supervising neural networks
- A Multi-domain Image Translative Diffusion StyleGAN for Iris Presentation Attack Detection
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- Efficient Learning of Distributed Linear-Quadratic Controllers
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- On the expressivity of sparse maxout networks
- Composition-based Multi-Relational Graph Convolutional Networks
- Scaling Vision Transformers for Functional MRI with Flat Maps
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- Graph Neural Networks: A Review of Methods and Applications
- Deep learning-based prediction of response to HER2-targeted neoadjuvant chemotherapy from pre-treatment dynamic breast MRI: A multi-institutional validation study
- Accelerating Federated Learning via Momentum Gradient Descent
- O3BNN-R: An Out-of-Order Architecture for High-Performance and Regularized BNN Inference
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Feature Denoising for Improving Adversarial Robustness
- The De-democratization of AI: Deep Learning and the Compute Divide in Artificial Intelligence Research
- Use of covariance matrix images for electroencephalography signal classification for multiclass motor imagery‐based brain computer interface
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- Randomness and Interpolation Improve Gradient Descent
- Plant leaf disease classification using EfficientNet deep learning model
- Stochastic Gradient/Mirror Descent: Minimax Optimality and Implicit Regularization
- Layer Normalization
- Visual7W: Grounded Question Answering in Images
- AnyUp: Universal Feature Upsampling
- GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
- Feature extraction of machine learning and phase transition point of\n Ising model
- Convolution, attention and structure embedding
- A Regularized Convolutional Neural Network for Semantic Image Segmentation
- Cautious Weight Decay
- MetaFormer Is Actually What You Need for Vision
- Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
- Local Background Features Matter in Out-of-Distribution Detection
- Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
- SDGraph: Multi-Level Sketch Representation Learning by Sparse-Dense Graph Architecture
- Faster Meta Update Strategy for Noise-Robust Deep Learning
- Tackling Catastrophic Forgetting and Background Shift in Continual Semantic Segmentation
- Robustness May Be at Odds with Accuracy
- Unsupervised Learning of Object Keypoints for Perception and Control
- Harnessing the Vulnerability of Latent Layers in Adversarially Trained\n Models
- R-Drop: Regularized Dropout for Neural Networks
- Lightweight CNN-Based Wi-Fi Intrusion Detection Using 2D Traffic Representations
- Cross Domain Image Matching in Presence of Outliers
- Fault Localization with Code Coverage Representation Learning
- Hierarchical Qubit-Merging Transformer for Quantum Error Correction
- High-resolution Photo Enhancement in Real-time: A Laplacian Pyramid Network
- Scalable Primitives for Generalized Sensor Fusion in Autonomous Vehicles
- Identification of tea foliar diseases and pest damage under practical field conditions using a convolutional neural network
- An Overview of Attacks and Defences on Intelligent Connected Vehicles
- Generating Binary Tags for Fast Medical Image Retrieval Based on Convolutional Nets and Radon Transform
- Adversarial Invariant Feature Learning with Accuracy Constraint for Domain Generalization
- Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
- Source-Free Object Detection with Detection Transformer
- NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal\n Matrices
- ΔEnergy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
- A compressed code for memory discrimination
- Restricted Receptive Fields for Face Verification
- dN/dx Reconstruction with Deep Learning for High-Granularity TPCs
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- A Strategy of MR Brain Tissue Images' Suggestive Annotation Based on Modified U-Net
- Deep learning methods based on cross-section images for predicting\n effective thermal conductivity of composites
- Learning the mapping x↦ ∑i=1d xi2: the cost of finding the needle in a haystack
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer
- Can a powerful neural network be a teacher for a weaker neural network?
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPU
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Automated machine learning: Review of the state-of-the-art and opportunities for healthcare
- Sketch Animation: State-of-the-art Report
- Understanding Notions of Stationarity in Non-Smooth Optimization
- A Large-Scale Benchmark for Food Image Segmentation
- Deep Learning using Linear Support Vector Machines
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Local-Global Context-Aware and Structure-Preserving Image Super-Resolution
- Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration
- Stochastic Training is Not Necessary for Generalization
- Prismo: A Decision Support System for Privacy-Preserving ML Framework Selection
- Machine learning in ant biology research: A systematic review
- Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear Images
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Artificial intelligence for education: Knowledge and its assessment in AI-enabled learning ecologies
- Computer-aided diagnosis of endobronchial ultrasound images using convolutional neural network
- Self-taught Object Localization with Deep Networks
- Understanding Character Recognition using Visual Explanations Derived from the Human Visual System and Deep Networks
- SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance Recognition
- DeepHash: Getting Regularization, Depth and Fine-Tuning Right
- Piecewise Linear Units Improve Deep Neural Networks
- OR-Net: Pointwise Relational Inference for Data Completion under Partial Observation
- CAMAL: Context-Aware Multi-layer Attention framework for Lightweight Environment Invariant Visual Place Recognition
- Deep Neural Networks for Choice Analysis: Architectural Design with Alternative-Specific Utility Functions
- Architecture Induces Structural Invariant Manifolds of Neural Network Training Dynamics
- Deep prior-based denoising for state-of-the-art scientific imaging and metrology
- Needles in Haystacks: On Classifying Tiny Objects in Large Images
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- Fully Convolutional Networks for Semantic Segmentation
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- Learning Transferable Adversarial Examples via Ghost Networks
- AGDC: Automatic Garbage Detection and Collection
- Provable Watermarking for Data Poisoning Attacks
- Distributionally robust approximation property of neural networks
- Robustness of Object Recognition under Extreme Occlusion in Humans and Computational Models
- Integrating Specialized Classifiers Based on Continuous Time Markov Chain
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- MNIST-NET10: A heterogeneous deep networks fusion based on the degree of certainty to reach 0.1 error rate. Ensembles overview and proposal
- VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
- Self-Training With Noisy Student Improves ImageNet Classification
- RAVEN: A Dataset for Relational and Analogical Visual rEasoNing
- LieTransformer: Equivariant self-attention for Lie Groups
- Exploring Data Aggregation and Transformations to Generalize across Visual Domains
- Image Classification with Classic and Deep Learning Techniques
- The impact of abstract and object tags on image privacy classification
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Deep Neural Networks Inspired by Differential Equations
- C-Net: A Reliable Convolutional Neural Network for Biomedical Image Classification
- Exploring Modality-shared Appearance Features and Modality-invariant Relation Features for Cross-modality Person Re-Identification
- DeepSpectrumLite: A Power-Efficient Transfer Learning Framework for Embedded Speech and Audio Processing from Decentralised Data
- Rectifier Neural Network with a Dual-Pathway Architecture for Image Denoising
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- Joint Architecture and Knowledge Distillation in CNN for Chinese Text Recognition
- Sparse components distinguish visual pathways & their alignment to neural networks
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- SGAS: Sequential Greedy Architecture Search
- Representation Learning from Limited Educational Data with Crowdsourced Labels
- Evaluation of Model Selection for Kernel Fragment Recognition in Corn Silage
- Anomaly Detection in Univariate Time-series: A Survey on the State-of-the-Art
- Time-Frequency Analysis based Blind Modulation Classification for Multiple-Antenna Systems
- AIM 2020 Challenge on Video Extreme Super-Resolution: Methods and Results
- Learning Longterm Representations for Person Re-Identification Using Radio Signals
- A Unified Object Motion and Affinity Model for Online Multi-Object Tracking
- Temporal Accumulative Features for Sign Language Recognition
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- Quick-CapsNet (QCN): A fast alternative to Capsule Networks
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Label-frugal satellite image change detection with generative virtual exemplar learning
- Interactive reconstruction of Monte Carlo image sequences using a recurrent denoising autoencoder
- The dilemma of quantum neural networks
- Dataset Reuse: Toward Translating Principles to Practice
- A review of evolving remote sensing and automated techniques in rock glacier mapping
- ImageNet Large Scale Visual Recognition Challenge
- Random Erasing Data Augmentation
- CoEdge: Cooperative DNN Inference with Adaptive Workload Partitioning over Heterogeneous Edge Devices
- Rapid computation of high-level visual surprise
- SympNets: Intrinsic structure-preserving symplectic networks for identifying Hamiltonian systems
- Verifying Memoryless Sequential Decision-making of Large Language Models
- Associative Memory Model with Neural Networks: Memorizing multiple images with one neuron
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- TreeNet: Layered Decision Ensembles
- Riddled basin geometry sets fundamental limits to predictability and reproducibility in deep learning
- Multi-task Neural Networks for QSAR Predictions
- Learning the detector in optical tomography
- Computing frustration and near-monotonicity in deep neural networks
- Evaluation of fish feeding intensity in aquaculture using a convolutional neural network and machine vision
- Neuroplastic Modular Framework: Cross-Domain Image Classification of Garbage and Industrial Surfaces
- Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
- Visual Representations inside the Language Model
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Improving Large-Scale Recommender Systems with Auxiliary Learning
- Exploring the Efficacy of Modified Transfer Learning in Identifying Parkinson's Disease Through Drawn Image Patterns
- POMO: Policy Optimization with Multiple Optima for Reinforcement Learning
- Improving Transferability of Adversarial Examples with Input Diversity
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Continuous vs. Discrete Optimization of Deep Neural Networks
- FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- A Mathematical Explanation of Transformers for Large Language Models and GPTs
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Attention Branch Network: Learning of Attention Mechanism for Visual Explanation
- Quantization Range Estimation for Convolutional Neural Networks
- Interpolated Convolutional Networks for 3D Point Cloud Understanding
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Evaluation of Transfer Learning for Classification of: (1) Diabetic\n Retinopathy by Digital Fundus Photography and (2) Diabetic Macular Edema,\n Choroidal Neovascularization and Drusen by Optical Coherence Tomography
- End-to-end Active Object Tracking via Reinforcement Learning
- Generation and Comprehension of Unambiguous Object Descriptions
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- Task-Level Contrastiveness for Cross-Domain Few-Shot Learning
- HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
- Accuracy Law for the Future of Deep Time Series Forecasting
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation Testing
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- An Efficient Quality Metric for Video Frame Interpolation Based on Motion-Field Divergence
- Image Generation Based on Image Style Extraction
- Real-time Hand Gesture Detection and Classification Using Convolutional\n Neural Networks
- Fast Object Detection in Compressed Video
- Phase Collaborative Network for Two-Phase Medical Image Segmentation
- TextCAM: Explaining Class Activation Map with Text
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Latency-Aware Differentiable Neural Architecture Search
- Realization of spatial sparseness by deep ReLU nets with massive data
- Indices Matter: Learning to Index for Deep Image Matting
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Max-Pooling Dropout for Regularization of Convolutional Neural Networks
- Hessian-Aware Pruning and Optimal Neural Implant
- Robust Context-Aware Object Recognition
- Signal Classification Recovery Across Domains Using Unsupervised Domain Adaptation
- Assessing Foundation Models for Mold Colony Detection with Limited Training Data
- Quantum Probabilistic Label Refining: Enhancing Label Quality for Robust Image Classification
- FourPhononGPU: A GPU-accelerated framework for calculating phonon scattering rates and thermal conductivity
- Uformer: A General U-Shaped Transformer for Image Restoration
- Digital Passport: A Novel Technological Strategy for Intellectual Property Protection of Convolutional Neural Networks
- Normal-Abnormal Guided Generalist Anomaly Detection
- On-the-Fly Data Augmentation via Gradient-Guided and Sample-Aware Influence Estimation
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- Finding beans in burgers: Deep semantic-visual embedding with\n localization
- PointConv: Deep Convolutional Networks on 3D Point Clouds
- DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
- PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
- On zero-shot recognition of generic objects
- Time-series forecasting with deep learning: a survey
- Deep Neural Networks are Easily Fooled: High Confidence Predictions for\n Unrecognizable Images
- From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
- Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep\n Learning
- FARSA: Fully Automated Roadway Safety Assessment
- Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling
- AIRCHITECT: Learning Custom Architecture Design and Mapping Space
- Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
- From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
- The Impact of Scaling Training Data on Adversarial Robustness
- Sharpness of Minima in Deep Matrix Factorization: Exact Expressions
- Using Images from a Video Game to Improve the Detection of Truck Axles
- Effective Model Pruning
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Human vs. AI Safety Perception? Decoding Human Safety Perception with Eye-Tracking Systems, Street View Images, and Explainable AI
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
- Symmetry-Aware Bayesian Optimization via Max Kernels
- CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
- Automatic recognition of feeding and foraging behaviour in pigs using deep learning
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- Enabling Physical AI through Biological Principles
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- DRIFT: Divergent Response in Filtered Transformations for Robust Adversarial Defense
- PEARL: Performance-Enhanced Aggregated Representation Learning
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
- BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression
- What Are We Automating? On the Need for Vision and Expertise When Deploying AI Systems
- Plant3R: Fusing 3D feature learning with Gaussian splatting to enhance wheat plant 3D reconstruction precision
- MLPerf Training Benchmark
- Learning to Diversify for Single Domain Generalization
- The Devil is in the Margin: Margin-based Label Smoothing for Network Calibration
- Model Watermarking for Image Processing Networks
- A Consolidated Approach to Convolutional Neural Networks and the Kolmogorov Complexity
- Retrieve-Then-Adapt: Example-based Automatic Generation for Proportion-related Infographics
- Road Crack Detection Using Deep Convolutional Neural Network and Adaptive Thresholding
- Exploring Uncertainty in Deep Learning for Construction of Prediction Intervals
- Semantically-Aware Strategies for Stereo-Visual Robotic Obstacle Avoidance
- Tent: Fully Test-time Adaptation by Entropy Minimization
- LifeCLEF Plant Identification Task 2014
- LifeCLEF Plant Identification Task 2014
- Gradient Flow Convergence Guarantee for General Neural Network Architectures
- FairViT-GAN: A Hybrid Vision Transformer with Adversarial Debiasing for Fair and Explainable Facial Beauty Prediction
- Influence-Guided Concolic Testing of Transformer Robustness
- Towards Interpretable Visual Decoding with Attention to Brain Representations
- A Computational Perspective on NeuroAI and Synthetic Biological Intelligence
- Spatially Parallel All-optical Neural Networks
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- Model-agnostic interpretation by visualization of feature perturbations
- Probabilistic Bearing Fault Diagnosis Using Gaussian Process with Tailored Feature Extraction
- Deep Learning Approaches with Explainable AI for Differentiating Alzheimer Disease and Mild Cognitive Impairment
- The unbearable slowness of being: Why do we live at 10 bits/s?
- Machine learning for synthetic gene circuit engineering
- Foundation models in bioinformatics
- Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
- Sparse, Collaborative, or Nonnegative Representation: Which Helps Pattern Classification?
- Leader Stochastic Gradient Descent for Distributed Training of Deep Learning Models: Extension
- Program-Guided Image Manipulators
- Metric-based Regularization and Temporal Ensemble for Multi-task\n Learning using Heterogeneous Unsupervised Tasks
- Stochastic Interpolants via Conditional Dependent Coupling
- HTMA-Net: Towards Multiplication-Avoiding Neural Networks via Hadamard Transform and In-Memory Computing
- DPFNAS: Differential Privacy-Enhanced Federated Neural Architecture Search for 6G Edge Intelligence
- Learning Modulated Loss for Rotated Object Detection
- Learning Generalizable Visual Representations via Interactive Gameplay
- CProp: Adaptive Learning Rate Scaling from Past Gradient Conformity
- On Controlled DeEntanglement for Natural Language Processing
- Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- How Secure is Distributed Convolutional Neural Network on IoT Edge Devices?
- Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models
- A Unified Approximation Framework for Compressing and Accelerating Deep Neural Networks
- MindCraft: How Concept Trees Take Shape In Deep Models
- PAPER: Privacy-Preserving Convolutional Neural Networks using Low-Degree Polynomial Approximations and Structural Optimizations on Leveled FHE
- Survey of Machine Learning Accelerators
- IONext: Unlocking the Next Era of Inertial Odometry
- A Sparse CNN Accelerator for Eliminating Redundant Computations in Intra- and Inter-Convolutional/Pooling Layers
- TSDM: Tracking by SiamRPN++ with a Depth-refiner and a Mask-generator
- Comparison and Benchmarking of AI Models and Frameworks on Mobile Devices
- Generative Image Modeling Using Spatial LSTMs
- Understand Scene Categories by Objects: A Semantic Regularized Scene Classifier Using Convolutional Neural Networks
- Visual Relationship Detection using Scene Graphs: A Survey
- Restructuring Batch Normalization to Accelerate CNN Training
- Mining Domain Knowledge: Improved Framework towards Automatically Standardizing Anatomical Structure Nomenclature in Radiotherapy
- A Learning-from-noise Dilated Wide Activation Network for denoising Arterial Spin Labeling (ASL) Perfusion Images
- On Embeddings in Relational Databases
- deepSELF: An Open Source Deep Self End-to-End Learning Framework
- Reconstructed spatial receptive field structures by reverse correlation technique explains the visual feature selectivity of units in deep convolutional neural networks
- SocialAI 0.1: Towards a Benchmark to Stimulate Research on Socio-Cognitive Abilities in Deep Reinforcement Learning Agents
- Design of Reconfigurable Multi-Operand Adder for Massively Parallel\n Processing
- Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- Light Differentiable Logic Gate Networks
- Prophecy: Inferring Formal Properties from Neuron Activations
- A Data-driven Typology of Vision Models from Integrated Representational Metrics
- LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- AI for Sustainable Future Foods
- Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
- Deep Adversarially-Enhanced k-Nearest Neighbors
- Short-term Load Forecasting with Deep Residual Networks
- Plant identification based on noisy web data: the amazing performance of deep learning (LifeCLEF 2017)
- DeepEMD: Differentiable Earth Mover's Distance for Few-Shot Learning
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- The Taboo Trap: Behavioural Detection of Adversarial Samples
- Explanations can be manipulated and geometry is to blame
- MeshCNN
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks
- Data Augmentation for Skin Lesion using Self-Attention based Progressive\n Generative Adversarial Network
- Audio-Visual Transformer Based Crowd Counting
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
- On the basis of brain: neural-network-inspired changes in general-purpose chips
- Supervised Contrastive Learning
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- Embodied AI: From LLMs to World Models
- GridMask Data Augmentation
- Impact of Loss Weight and Model Complexity on Physics-Informed Neural Networks for Computational Fluid Dynamics
- Sobolev acceleration for neural networks
- Mamba Modulation: On the Length Generalization of Mamba
- Adaptive von Mises-Fisher Likelihood Loss for Supervised Deep Time Series Hashing
- PAC-Bayes Analysis of Sentence Representation
- Learning to see across Domains and Modalities
- Thickened 2D Networks for Efficient 3D Medical Image Segmentation
- Not Only Look But Observe: Variational Observation Model of Scene-Level 3D Multi-Object Understanding for Probabilistic SLAM
- Deep Control - a simple automatic gain control for memory efficient and\n high performance training of deep convolutional neural networks
- Energy-Efficient Processing and Robust Wireless Cooperative Transmission for Edge Inference
- AutoGrow: Automatic Layer Growing in Deep Convolutional Networks
- Seesaw-Net: Convolution Neural Network With Uneven Group Convolution
- Structure-Aware Face Clustering on a Large-Scale Graph with \bf107 Nodes
- Unsupervised Domain Adaptation for Object Detection via Cross-Domain Semi-Supervised Learning
- Multi-task Learning by Leveraging the Semantic Information
- Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
- Deep(er) Learning
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- Hybrid Composition with IdleBlock: More Efficient Networks for Image Recognition
- Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Multi Layer Neural Networks as Replacement for Pooling Operations
- Lets keep it simple, Using simple architectures to outperform deeper and\n more complex architectures
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Large-Scale Long-Tailed Recognition in an Open World
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
- Artificial Intelligence Empowered New Materials: Discovery, Synthesis, Prediction to Validation
- Does the brain's ventral visual pathway compute object shape?
- Application of <scp>MobileNet</scp> and Xception neural networks to identify <scp> <i>Sillago sihama</i> </scp> populations in Vietnam's coastal waters based on otolith morphology
- Deep learning in plant phenotyping: the first ten years
- Quantitative Evaluations on Saliency Methods: An Experimental Study
- Deep learning‐based association analysis of root image data and cucumber yield
- Unifying Relational Sentence Generation and Retrieval for Medical Image Report Composition
- Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
- Ellipse Regression with Predicted Uncertainties for Accurate Multi-View 3D Object Estimation
- DAIL: Dataset-Aware and Invariant Learning for Face Recognition
- When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
- T-BFA: Targeted Bit-Flip Adversarial Weight Attack
- Multi-path Neural Networks for On-device Multi-domain Visual Classification
- Channel Tiling for Improved Performance and Accuracy of Optical Neural Network Accelerators
- Dynamic DNN Decomposition for Lossless Synergistic Inference
- Accuracy and Architecture Studies of Residual Neural Network solving Ordinary Differential Equations
- Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
- Semantic Understanding of Scenes Through the ADE20K Dataset
- PotentialNet for Molecular Property Prediction
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Selecting Data Augmentation for Simulating Interventions
- Gesture Recognition for Initiating Human-to-Robot Handovers
- Composition-Aware Image Aesthetics Assessment
- Unsupervised Intuitive Physics from Past Experiences
- Fuzzy Semantic Segmentation of Breast Ultrasound Image with Breast Anatomy Constraints
- Medical Image Segmentation Using a U-Net type of Architecture
- Thanks for Nothing: Predicting Zero-Valued Activations with Lightweight Convolutional Neural Networks
- A Deep Learning Framework for Classification of in vitro Multi-Electrode Array Recordings
- Inductive Bias of Gradient Descent based Adversarial Training on Separable Data
- Zero-shifting Technique for Deep Neural Network Training on Resistive Cross-point Arrays
- Compact deep neural network models of the visual cortex
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- SEALing Neural Network Models in Secure Deep Learning Accelerators
- Quantum Energy Regression using Scattering Transforms
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- Dense Adaptive Cascade Forest: A Self Adaptive Deep Ensemble for Classification Problems
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- k-Nearest Neighbors by Means of Sequence to Sequence Deep Neural Networks and Memory Networks
- Application of Machine Learning for Aboveground Biomass Modeling in Tropical and Temperate Forests from Airborne Hyperspectral Imagery
- Use What You Know: Causal Foundation Models with Partial Graphs
- The steep cost of capture
- Monitoring urban construction and quarry blasts with low-cost seismic sensors and deep learning tools in the city of Oslo, Norway
- Identifying controlling factors of delta morphology using a convolutional autoencoder
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- EPick: Multi-Class Attention-based U-shaped Neural Network for Earthquake Detection and Seismic Phase Picking
- A Computer-Aided Diagnosis System for Breast Pathology: A Deep Learning Approach with Model Interpretability from Pathological Perspective
- A Performance Comparison of Loss Functions for Deep Face Recognition
- A highly selective response to food in human visual cortex revealed by hypothesis-free voxel decomposition
- Feature Space Augmentation for Long-Tailed Data
- On the Generalization Error Bounds of Neural Networks under Diversity-Inducing Mutual Angular Regularization
- Wireless for Machine Learning
- Auto-Encoding Twin-Bottleneck Hashing
- End-to-End Learning Local Multi-view Descriptors for 3D Point Clouds
- A Geometric Approach to Online Streaming Feature Selection
- Deep Affinity Net: Instance Segmentation via Affinity
- Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition
- Signal-to-Noise Ratio: A Robust Distance Metric for Deep Metric Learning
- GRIPHIN: grids of pharmacophore interaction fields for affinity prediction
- The Brain Abstracted
- Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
- Overcoming Data Sparsity in Group Recommendation
- Learning from Large-scale Noisy Web Data with Ubiquitous Reweighting for Image Classification
- Not All Ops Are Created Equal!
- Handcrafted Backdoors in Deep Neural Networks
- A Survey of Autonomous Driving: <i>Common Practices and Emerging Technologies</i>
- What Do Single-view 3D Reconstruction Networks Learn?
- AutoFlow: Learning a Better Training Set for Optical Flow
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- Identifying and Compensating for Feature Deviation in Imbalanced Deep Learning
- Merchant Category Identification Using Credit Card Transactions
- Shape and Symmetry Induction for 3D Objects
- Detection of Furigana Text in Images
- Game-Theoretic Multiagent Reinforcement Learning
- Can We Faithfully Represent Masked States to Compute Shapley Values on a DNN?
- Robotic grasp detection using a novel two-stage approach
- TRACE: Early Detection of Chronic Kidney Disease Onset with Transformer-Enhanced Feature Embedding
- LoCo: Local Contrastive Representation Learning
- Biologically-inspired Salience Affected Artificial Neural Network (SANN)
- Video Anomaly Detection by Estimating Likelihood of Representations
- A Sheaf and Topology Approach to Generating Local Branch Numbers in Digital Images
- Robustness and Transferability of Universal Attacks on Compressed Models
- Seeing eye-to-eye? A comparison of object recognition performance in\n humans and deep convolutional neural networks under image manipulation
- Uncertainty-driven ensembles of deep architectures for multiclass classification. Application to COVID-19 diagnosis in chest X-ray images
- Why Convolutional Networks Learn Oriented Bandpass Filters: Theory and Empirical Support
- DSRNA: Differentiable Search of Robust Neural Architectures
- ParaNet: Deep Regular Representation for 3D Point Clouds
- KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning
- Meticulous Object Segmentation
- Quantum computing tools for fast detection of gravitational waves in the context of LISA space mission
- Sketch-Specific Data Augmentation for Freehand Sketch Recognition
- Measurement-driven Security Analysis of Imperceptible Impersonation Attacks
- The Sloop System for Individual Animal Identification with Deep Learning
- Assessing the Alignment of Popular CNNs to the Brain for Valence Appraisal
- ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation
- Emergent symbolic language based deep medical image classification
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Visual Recognition Using Directional Distribution Distance
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- Improving Outdoor Multi-cell Fingerprinting-based Positioning via Mobile Data Augmentation
- Hardness-Aware Deep Metric Learning
- Online and Offline Handwritten Chinese Character Recognition: A Comprehensive Study and New Benchmark
- Freehand Sketch Recognition Using Deep Features
- A Tutorial on Quantum Convolutional Neural Networks (QCNN)
- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for\n Autonomous Driving
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
- Density-embedding layers: a general framework for adaptive receptive fields
- A Scalable Lift-and-Project Differentiable Approach For the Maximum Cut Problem
- OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
- Machine learning approach to single-shot multiparameter estimation for the non-linear Schrödinger equation
- Automated Tracking of Primate Behavior
- Bounded PCTL Model Checking of Large Language Model Outputs
- Recognition Of Surface Defects On Steel Sheet Using Transfer Learning
- Detecting Deep Neural Network Defects with Data Flow Analysis
- Self-attention based BiLSTM-CNN classifier for the prediction of ischemic and non-ischemic cardiomyopathy
- Effective and efficient ROI-wise visual encoding using an end-to-end CNN regression model and selective optimization
- Towards Understanding and Modeling Empathy for Use in Motivational Design Thinking
- A Deep Ordinal Distortion Estimation Approach for Distortion Rectification
- SSFN -- Self Size-estimating Feed-forward Network with Low Complexity, Limited Need for Human Intervention, and Consistent Behaviour across Trials
- Six Sigma For Neural Networks: Taguchi-based optimization
- Learning Invariant Representations for Sentiment Analysis: The Missing Material is Datasets
- Effectiveness of Data Augmentation in Cellular-based Localization Using Deep Learning
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge
- Modeling Human Motion with Quaternion-Based Neural Networks
- The evolution of neural network-based chart patterns
- Neural network models and deep learning
- FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Trustworthy Convolutional Neural Networks: A Gradient Penalized-based Approach
- A PMU-Based Machine Learning Application for Fast Detection of Forced Oscillations from Wind Farms
- Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
- nDNA -- the Semantic Helix of Artificial Cognition
- SynergyNet: Fusing Generative Priors and State-Space Models for Facial Beauty Prediction
- SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
- From handcrafted to deep local features
- Deep Clustering for Unsupervised Learning of Visual Features
- Guidelines and Benchmarks for Deployment of Deep Learning Models on\n Smartphones as Real-Time Apps
- A transfer learning method with deep residual network for pediatric pneumonia diagnosis
- Neural Network Based Framework for Passive Intermodulation Cancellation in MIMO Systems
- Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search
- A multiscale neural network based on hierarchical nested bases
- SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- Checking extracted rules in Neural Networks
- Randomized Smoothing Meets Vision-Language Models
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Deep learning algorithms for detection of diabetic retinopathy in retinal fundus photographs: A systematic review and meta-analysis
- Skeletal bone age prediction based on a deep residual network with spatial transformer
- A neural network approach to segment brain blood vessels in digital subtraction angiography
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- A vision transformer based approach for analysis of plasmodium vivax life cycle for malaria prediction using thin blood smear microscopic images
- A deep sift convolutional neural networks for total brain volume estimation from 3D ultrasound images
- Deep Learning and Traffic Classification: Lessons learned from a commercial-grade dataset with hundreds of encrypted and zero-day applications
- Region-Aware Deformable Convolutions
- Recent Advancements in Microscopy Image Enhancement using Deep Learning: A Survey
- Effect of Initial Configuration of Weights on Training and Function of Artificial Neural Networks
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
- Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
- Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling
- Crafting GBD-Net for Object Detection
- Structure-Preserving Margin Distribution Learning for High-Order Tensor Data with Low-Rank Decomposition
- Exploring Self-attention for Image Recognition
- Task-Aware Monocular Depth Estimation for 3D Object Detection
- White-box machine learning for uncovering physically interpretable dimensionless governing equations for granular materials
- Adaptable image quality assessment using meta-reinforcement learning of task amenability
- Domain Generalization via Multidomain Discriminant Analysis
- Pointing Novel Objects in Image Captioning
- Survey on semantic segmentation using deep learning techniques
- Deep Learning For Computer Vision Tasks: A review
- Open-World Class Discovery with Kernel Networks
- Dynamic Spectrum Matching with One-shot Learning
- Unsupervised Object-Level Representation Learning from Scene Images
- MARIC: Multi-Agent Reasoning for Image Classification
- Learning FRAME Models Using CNN Filters
- Incorporating Visual Cortical Lateral Connection Properties into CNN: Recurrent Activation and Excitatory-Inhibitory Separation
- Deep Predictive Coding Network for Object Recognition
- Probabilistic Neural Network with Complex Exponential Activation\n Functions in Image Recognition using Deep Learning Framework
- Techniques for visualizing LSTMs applied to electrocardiograms
- Can you hear me \now? Sensitive comparisons of human and\n machine perception
- DeepPeep: Exploiting Design Ramifications to Decipher the Architecture of Compact DNNs
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification
- Binary Classification of Light and Dark Time Traces of a Transition Edge Sensor Using Convolutional Neural Networks
- Embedding Label Structures for Fine-Grained Feature Representation
- Bridging the Gap between Label- and Reference-based Synthesis in Multi-attribute Image-to-Image Translation
- Rest2Visual: Predicting Visually Evoked fMRI from Resting-State Scans
- A Lightweight ReLU-Based Feature Fusion for Aerial Scene Classification
- TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
- Intelligent Vacuum Thermoforming Process
- RTMobile: Beyond Real-Time Mobile Acceleration of RNNs for Speech Recognition
- Robot-Assisted Feeding: Generalizing Skewering Strategies across Food Items on a Realistic Plate
- On the Structural Sensitivity of Deep Convolutional Networks to the Directions of Fourier Basis Functions
- Deep Learning for Design and Retrieval of Nano-photonic Structures
- Transfer Metric Learning: Algorithms, Applications and Outlooks
- Proof of Federated Learning: A Novel Energy-recycling Consensus Algorithm
- EByFTVeS: Efficient Byzantine Fault Tolerant-based Verifiable Secret-sharing in Distributed Privacy-preserving Machine Learning
- TreeGAN: Syntax-Aware Sequence Generation with Generative Adversarial Networks
- Revisiting Fine-tuning for Few-shot Learning
- BATR-FST: Bi-Level Adaptive Token Refinement for Few-Shot Transformers
- Multi-level Wavelet Convolutional Neural Networks
- High-resolution rainfall-runoff modeling using graph neural network
- DARTS: Deceiving Autonomous Cars with Toxic Signs
- A Robust and Precise ConvNet for small non-coding RNA classification\n (RPC-snRC)
- A Data-Aware Fourier Neural Operator for Modeling Spatiotemporal Electromagnetic Fields
- Ranking Distillation: Learning Compact Ranking Models With High Performance for Recommender System
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- A Survey on Neural Architecture Search
- On the Decision Boundary of Deep Neural Networks
- A Loss Function for Generative Neural Networks Based on Watson's Perceptual Model
- Spherical Convolutional Neural Networks: Stability to Perturbations in SO(3)
- Advancing chest X-ray diagnostics: A novel CycleGAN-based preprocessing approach for enhanced lung disease classification in ChestX-Ray14
- Proposal Learning for Semi-Supervised Object Detection
- A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
- Overlearning Reveals Sensitive Attributes
- Toward Filament Segmentation Using Deep Neural Networks
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Future Frame Prediction for Anomaly Detection -- A New Baseline
- Tubelets: Unsupervised action proposals from spatiotemporal super-voxels
- Scene Labeling with Contextual Hierarchical Models
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
- PhotoApp: Photorealistic Appearance Editing of Head Portraits
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- Beyond Triplet Loss: Meta Prototypical N-tuple Loss for Person Re-identification
- RGP: Neural Network Pruning through Its Regular Graph Structure
- TFD-former: Time-frequency domain fusion decoders for effective and robust fault diagnosis under time-varying speeds
- SCA-PVNet: Self-and-cross attention based aggregation of point cloud and multi-view for 3D object retrieval
- Subject-independent Human Pose Image Construction with Commodity Wi-Fi
- MailLeak: Obfuscation-Robust Character Extraction Using Transfer Learning
- Relating Graph Neural Networks to Structural Causal Models
- Deep Learning in Memristive Nanowire Networks
- Univariate ReLU neural network and its application in nonlinear system identification
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples
- Classifier Crafting: Turn Your ConvNet into a Zero-Shot Learner!
- Learning Deep Representations of Fine-grained Visual Descriptions
- Addressing Failure Prediction by Learning Model Confidence
- Flows Over Periodic Hills of Parameterized Geometries: A Dataset for Data-Driven Turbulence Modeling From Direct Simulations
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Causal-Symbolic Meta-Learning (CSML): Inducing Causal World Models for Few-Shot Generalization
- Synetgy
- Deep learning with asymmetric connections and Hebbian updates
- Sitatapatra: Blocking the Transfer of Adversarial Samples
- A survey on deep learning approaches for breast cancer diagnosis
- Data augmentation and image understanding
- Artificial Neural Networks for Neuroscientists: A Primer
- LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Unrolling Graph-based Douglas-Rachford Algorithm for Image Interpolation with Informed Initialization
- A Multiplexed Network for End-to-End, Multilingual OCR
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
- Efficient Algorithms for Device Placement of DNN Graph Operators
- The Neural Network Approach to Inverse Problems in Differential Equations
- Evaluating State-of-the-Art Classification Models Against Bayes Optimality
- IS-Diff: Improving Diffusion-Based Inpainting with Better Initial Seed
- Determining the boundary of dynamical chaos in the generalized Chirikov map via machine learning
- HyperInverter: Improving StyleGAN Inversion via Hypernetwork
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- Intention-aware Long Horizon Trajectory Prediction of Surrounding Vehicles using Dual LSTM Networks
- A Controllable 3D Deepfake Generation Framework with Gaussian Splatting
- Efficient Byzantine-Robust Privacy-Preserving Federated Learning via Dimension Compression
- DISA at ImageCLEF 2014 Revised: Search-based Image Annotation with DeCAF Features
- Error Control and Loss Functions for the Deep Learning Inversion of Borehole Resistivity Measurements
- Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks
- gen2Out: Detecting and Ranking Generalized Anomalies
- Enhancing Electromagnetic Calorimeter Signal Reconstruction with Machine Learning-Based Noise Discrimination
- AngularGrad: A New Optimization Technique for Angular Convergence of Convolutional Neural Networks
- Straggler-resistant distributed matrix computation via coding theory
- InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
- Neural networks in the search for fast radio bursts with RATAN-600
- Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks
- Self-Attention Capsule Networks for Object Classification
- Squeeze-and-Excitation on Spatial and Temporal Deep Feature Space for Action Recognition
- An Advanced Convolutional Neural Network for Bearing Fault Diagnosis under Limited Data
- Compact Device Models for FinFET and Beyond
- Unlabeled Data Deployment for Classification of Diabetic Retinopathy Images Using Knowledge Transfer
- Weakly Supervised Vulnerability Localization via Multiple Instance Learning
- Sinogram super-resolution and denoising convolutional neural network (SRCN) for limited data photoacoustic tomography
- Review: deep learning on 3D point clouds
- Two-Phase Object-Based Deep Learning for Multi-temporal SAR Image Change Detection
- A survey on Deep Learning based bearing fault diagnosis
- Neural Autoregressive Distribution Estimation
- Learnable Gabor modulated complex-valued networks for orientation robustness
- Deep Learning for Generic Object Detection: A Survey
- Rethinking Channel Dimensions for Efficient Model Design
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- Deep Learning for Insider Threat Detection: Review, Challenges and Opportunities
- A Critical Review of Recurrent Neural Networks for Sequence Learning
- Identification of Crystal Symmetry from Noisy Diffraction Patterns by A Shape Analysis and Deep Learning
- A Sketch Based 3D Shape Retrieval Approach Based on Efficient Deep Point-to-Subspace Metric Learning
- A Capsule-unified Framework of Deep Neural Networks for Graphical Programming
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Deep convolutional neural networks in the face of caricature
- Exploring Weight Symmetry in Deep Neural Networks
- InverSynth: Deep Estimation of Synthesizer Parameter Configurations from\n Audio Signals
- Contrastive Learning with Stronger Augmentations
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- Identifying Pediatric Vascular Anomalies With Deep Learning
- Parameterized Knowledge Transfer for Personalized Federated Learning
- STFCN: Spatio-Temporal FCN for Semantic Video Segmentation
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Expressive Power of Deep Networks on Manifolds: Simultaneous Approximation
- Person Re-identification: Past, Present and Future
- Image Deformation Meta-Networks for One-Shot Learning
- Objectness Similarity: Capturing Object-Level Fidelity in 3D Scene Evaluation
- Illumination-aware Faster R-CNN for Robust Multispectral Pedestrian Detection
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- MSDANet: A Multiscale Dual-Channel Spatial Attention Network with Depthwise Separable Convolution for Hyperspectral Image Classification
- Images in Motion?: A First Look into Video Leakage in Collaborative Deep Learning
- MultiGrain: a unified image embedding for classes and instances
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
- Renovating Parsing R-CNN for Accurate Multiple Human Parsing
- Compressing CNN models for resource-constrained systems by channel and layer pruning
- Modulating human brain responses via optimal natural image selection and synthetic image generation
- CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
- Network Pruning via Transformable Architecture Search
- Deep Learning for Free-Hand Sketch: A Survey
- Line Segment Detection Using Transformers without Edges
- Multispectral CT Denoising via Simulation-Trained Deep Learning: Experimental Results at the ESRF BM18
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Actor-Centric Relation Network
- EAST: An Efficient and Accurate Scene Text Detector
- Trellis Networks for Sequence Modeling
- Distilling the Knowledge in a Neural Network
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- RISC-NN: Use RISC, NOT CISC as Neural Network Hardware Infrastructure
- Rollout-LaSDI: Enhancing the long-term accuracy of Latent Space Dynamics
- Localized PCA-Net Neural Operators for Scalable Solution Reconstruction of Elliptic PDEs
- Impression Space from Deep Template Network
- Residual Dense Network for Image Restoration
- A Practical Deep Learning-Based Acoustic Side Channel Attack on Keyboards
- FusionLane: Multi-Sensor Fusion for Lane Marking Semantic Segmentation Using Deep Neural Networks
- PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy
- Efficient resource management in UAVs for Visual Assistance
- Dataset Distillation
- Language Self-Play For Data-Free Training
- Three Pillars improving Vision Foundation Model Distillation for Lidar
- Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Temporal Image Forensics: A Review and Critical Evaluation
- Enhanced Memory Network: The novel network structure for Symbolic Music Generation
- Revisiting the Calibration of Modern Neural Networks
- Stochastic Sign Descent Methods: New Algorithms and Better Theory
- Inferring brain-computational mechanisms with models of activity\n measurements
- Evaluating the Impact of Adversarial Attacks on Traffic Sign Classification using the LISA Dataset
- Data-driven discovery of dynamical models in biology
- Video Anomaly Detection and Localization via Gaussian Mixture Fully Convolutional Variational Autoencoder
- An Analysis of Scale Invariance in Object Detection - SNIP
- Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- Foldover Features for Dynamic Object Behavior Description in Microscopic Videos
- Deep learning visual analysis in laparoscopic surgery: a systematic review and diagnostic test accuracy meta-analysis
- Comparison Network for One-Shot Conditional Object Detection
- Machine Learning: Algorithms, Real-World Applications and Research Directions
- Unifying Remote Sensing Image Retrieval and Classification with Robust Fine-tuning
- Adaptive, Distribution-Free Prediction Intervals for Deep Networks
- Laguerre-Gauss Preprocessing: Line Profiles as Image Features for Aerial\n Images Classification
- Image Enhanced Rotation Prediction for Self-Supervised Learning
- Dimensionally Reduced Open-World Clustering: DROWCULA
- Lookup multivariate Kolmogorov-Arnold Networks
- AIM 2025 Challenge on High FPS Motion Deblurring: Methods and Results
- Albumentations: Fast and Flexible Image Augmentations
- Data-Efficient Time-Dependent PDE Surrogates: Graph Neural Simulators vs. Neural Operators
- RPC: A Large-Scale Retail Product Checkout Dataset
- Just Jump: Dynamic Neighborhood Aggregation in Graph Neural Networks
- Fooling Computer Vision into Inferring the Wrong Body Mass Index
- Micro-Expression Recognition via Fine-Grained Dynamic Perception
- Agglomerative Attention
- Three-Dimensional Mesh Steganography and Steganalysis: A Review
- A brain-inspired paradigm for scalable quantum vision
- Unity Style Transfer for Person Re-Identification
- DNA: Deeply-supervised Nonlinear Aggregation for Salient Object Detection
- DeepPoison: Feature Transfer Based Stealthy Poisoning Attack
- Explainable AI: A Review of Machine Learning Interpretability Methods
- Dual-Mode Deep Anomaly Detection for Medical Manufacturing: Structural Similarity and Feature Distance
- High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
- Simulation Priors for Data-Efficient Deep Learning
- Rethinking Supervised Pre-training for Better Downstream Transferring
- Prior Distribution and Model Confidence
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Systematic Review and Meta-analysis of AI-driven MRI Motion Artifact Detection and Correction
- Lite-HRNet: A Lightweight High-Resolution Network
- Space-time Mixing Attention for Video Transformer
- TemporalFlowViz: Parameter-Aware Visual Analytics for Interpreting Scramjet Combustion Evolution
- Combating the Elsagate phenomenon: Deep learning architectures for disturbing cartoons
- Synergy: Resource Sensitive DNN Scheduling in Multi-Tenant Clusters
- Graph Unlearning: Efficient Node Removal in Graph Neural Networks
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- Scale-interaction transformer: a hybrid cnn-transformer model for facial beauty prediction
- Deep Learning Approaches for Image Retrieval and Pattern Spotting in\n Ancient Documents
- Multi-Class Lane Semantic Segmentation using Efficient Convolutional Networks
- Convolutional Dictionary Learning in Hierarchical Networks
- Dynamic Sensitivity Filter Pruning using Multi-Agent Reinforcement Learning For DCNN's
- On the Normalization of Confusion Matrices: Methods and Geometric Interpretations
- Bayesian Loss for Crowd Count Estimation with Point Supervision
- SpecNet: Spectral Domain Convolutional Neural Network
- Universality of Gradient Descent Neural Network Training
- Video Analytics with Zero-streaming Cameras
- Towards Open World Detection: A Survey
- Empirical Studies on the Properties of Linear Regions in Deep Neural Networks
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- Delta Activations: A Representation for Finetuned Large Language Models
- Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion
- Light Field Reconstruction via Deep Adaptive Fusion of Hybrid Lenses
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- Transformer-based Multi-Aspect Modeling for Multi-Aspect Multi-Sentiment Analysis
- A Survey on 3D Skeleton-Based Action Recognition Using Learning Method
- Lesion-based Contrastive Learning for Diabetic Retinopathy Grading from Fundus Images
- Improving the Accuracy and Hardware Efficiency of Neural Networks Using Approximate Multipliers
- Interpreting Black-Box Models: A Review on Explainable Artificial Intelligence
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- Understanding Architectures Learnt by Cell-based Neural Architecture Search
- 1D convolutional neural networks and applications: A survey
- From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions
- Spatiotemporal Pyramid Network for Video Action Recognition
- ResiliNet: Failure-Resilient Inference in Distributed Neural Networks
- How Important is the Train-Validation Split in Meta-Learning?
- Watermarking Graph Neural Networks based on Backdoor Attacks
- Semantic Segmentation from Limited Training Data
- Moment Matching for Multi-Source Domain Adaptation
- Robust Training of Social Media Image Classification Models for Rapid Disaster Response
- Learning to See before Learning to Act: Visual Pre-training for Manipulation
- Sparse Autoencoder Neural Operators: Model Recovery in Function Spaces
- LSAM: Asynchronous Distributed Training with Landscape-Smoothed Sharpness-Aware Minimization
- FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
- Isolated Bangla Handwritten Character Classification using Transfer Learning
- Minimizing Perceived Image Quality Loss Through Adversarial Attack\n Scoping
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Understanding deep learning requires rethinking generalization
- End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition
- Directional Bias Amplification
- Long-Term Vehicle Localization by Recursive Knowledge Distillation
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- Human De-occlusion: Invisible Perception and Recovery for Humans
- MIDI-Sandwich2: RNN-based Hierarchical Multi-modal Fusion Generation VAE networks for multi-track symbolic music generation
- Residual Squeeze VGG16
- Stealth by Conformity: Evading Robust Aggregation through Adaptive Poisoning
- Multi-Scale Deep Learning for Colon Histopathology: A Hybrid Graph-Transformer Approach
- A Polynomial-Based Approach for Architectural Design and Learning with\n Deep Neural Networks
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- A Convolutional Hierarchical Deep-learning Neural Network (C-HiDeNN) Framework for Non-linear Finite Element Analysis
- Vision encoders should be image size agnostic and task driven
- MixKD: Towards Efficient Distillation of Large-scale Language Models
- Vision-based deep execution monitoring
- MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices
- Exploring Vicinal Risk Minimization for Lightweight Out-of-Distribution Detection
- The gap between theory and practice in function approximation with deep neural networks
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and Implications
- Open-Domain Conversational Agents: Current Progress, Open Problems, and Future Directions
- FAWA: Fast Adversarial Watermark Attack on Optical Character Recognition (OCR) Systems
- Pruning Convolutional Neural Networks with Self-Supervision
- Robots of the Lost Arc: Self-Supervised Learning to Dynamically Manipulate Fixed-Endpoint Cables
- Robust Small Methane Plume Segmentation in Satellite Imagery
- One Size Does Not Fit All: Multi-Scale, Cascaded RNNs for Radar Classification
- LaSOT: A High-quality Large-scale Single Object Tracking Benchmark
- TensorLib: A Spatial Accelerator Generation Framework for Tensor Algebra
- HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass Lens
- An Investigation of Visual Foundation Models Robustness
- Social gaze fingerprints: identifying social virtual reality users by their eye gaze patterns
- Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- Climate impacts and future trends of hailstorms in China based on millennial records
- Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- Characterizing Types of Convolution in Deep Convolutional Recurrent Neural Networks for Robust Speech Emotion Recognition
- Generalized Zero-Shot Domain Adaptation via Coupled Conditional Variational Autoencoders
- Memory Efficient Class-Incremental Learning for Image Classification
- URLNet: Learning a URL Representation with Deep Learning for Malicious URL Detection
- Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos
- Rotation Invariance Neural Network
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- Diversity inducing Information Bottleneck in Model Ensembles
- A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
- Using Capsule Neural Network to predict Tuberculosis in lens-free microscopic images
- MemeSequencer: Sparse Matching for Embedding Image Macros
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Music Genre Classification Using Machine Learning Techniques
- TransMatch: A Transfer-Learning Framework for Defect Detection in Laser Powder Bed Fusion Additive Manufacturing
- SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities
- Mamba-CNN: A Hybrid Architecture for Efficient and Accurate Facial Beauty Prediction
- Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup Tables
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Contrastive Model Inversion for Data-Free Knowledge Distillation
- MSE Loss with Outlying Label for Imbalanced Classification
- Learning Connectivity of Neural Networks from a Topological Perspective
- Towards Deep Learning Assisted Autonomous UAVs for Manipulation Tasks in GPS-Denied Environments
- Plant Disease Detection and Classification by Deep Learning—A Review
- Review on Convolutional Neural Network (CNN) Applied to Plant Leaf Disease Classification
- Proximal Backpropagation
- SocialGCN: An Efficient Graph Convolutional Network based Model for Social Recommendation
- Performance evaluation of an integrated photonic convolutional neural network based on delay buffering and wavelength division multiplexing
- The Limitations of Deep Learning in Adversarial Settings
- Move Evaluation in Go Using Deep Convolutional Neural Networks
- Probabilistic Permutation Synchronization using the Riemannian Structure\n of the Birkhoff Polytope
- OnlineAugment: Online Data Augmentation with Less Domain Knowledge
- Assessing the (Un)Trustworthiness of Saliency Maps for Localizing\n Abnormalities in Medical Imaging
- Routing Towards Discriminative Power of Class Capsules
- From Synthetic to Real: Unsupervised Domain Adaptation for Animal Pose Estimation
- COVID-19 personal protective equipment detection using real-time deep learning methods
- Modeling the Nonsmoothness of Modern Neural Networks
- Lung cancer identification: a review on detection and classification
- Compact Deep Aggregation for Set Retrieval
- Towards Accurate and Compact Architectures via Neural Architecture Transformer
- Understanding and Enhancing the Use of Context for Machine Translation
- Learning by training: emergent return-point memory from cyclically tuning disordered sphere packings
- Using Distance Estimation and Deep Learning to Simplify Calibration in\n Food Calorie Measurement
- Rethinking FUN: Frequency-Domain Utilization Networks
- Deep Learning on Image Denoising: An overview
- A differential neural network learns stochastic differential equations and the Black-Scholes equation for pricing multi-asset options
- Regularization via Mass Transportation
- Partial success in closing the gap between human and machine vision
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- Incremental Learning with Maximum Entropy Regularization: Rethinking\n Forgetting and Intransigence
- Unsupervised Domain Adaptation Learning Algorithm for RGB-D Staircase Recognition
- Dropout drops double descent
- CNN with large memory layers
- Generative Latent Space Dynamics of Electron Density
- Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability
- Understanding the Behaviour of Contrastive Loss
- Optimization Variance: Exploring Generalization Properties of DNNs
- Artificial Intelligence in Drug Discovery: Applications and Techniques
- AI Compute Architecture and Evolution Trends
- Representation Learning with Adaptive Superpixel Coding
- CuratorNet: Visually-aware Recommendation of Art Images
- Accelerated Training for Massive Classification via Dynamic Class Selection
- DeepOPF: Deep Neural Network for DC Optimal Power Flow
- OmniArt: Multi-task Deep Learning for Artistic Data Analysis
- Accelerating Federated Learning via Momentum Gradient Descent
- Convolutional Neural Fabrics
- Playing Atari with Deep Reinforcement Learning
- Survey of spiking in the mouse visual system reveals functional hierarchy
- EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos
- Cyclone intensity estimate with context-aware cyclegan
- Deep Anomaly Detection by Residual Adaptation
- Estimating Model Uncertainty of Neural Networks in Sparse Information\n Form
- Collective Learning by Ensembles of Altruistic Diversifying Neural Networks
- Efficient Integer-Arithmetic-Only Convolutional Neural Networks
- Learning compact generalizable neural representations supporting perceptual grouping
- IA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers
- GODS: Generalized One-class Discriminative Subspaces for Anomaly Detection
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Microscopic and collective signatures of feature learning in neural networks
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- Data-Free Adversarial Distillation
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- D3PINNs: A Novel Physics-Informed Neural Network Framework for Staged Solving of Time-Dependent Partial Differential Equations
- Network Representation Learning: From Traditional Feature Learning to Deep Learning
- Developing a Multi-Modal Machine Learning Model For Predicting Performance of Automotive Hood Frames
- Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
- 3D-Aided Data Augmentation for Robust Face Understanding
- A Chain Graph Interpretation of Real-World Neural Networks
- Character-Aware Neural Language Models
- Improving Adversarial Robustness via Guided Complement Entropy
- Recovery Guarantees for Compressible Signals with Adversarial Noise
- Computational Emotion Analysis From Images: Recent Advances and Future Directions
- Iterative Low-Rank Approximation for CNN Compression
- What makes instance discrimination good for transfer learning?
- A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
- Exploring Randomly Wired Neural Networks for Image Recognition
- Objective Value Change and Shape-Based Accelerated Optimization for the Neural Network Approximation
- Seam360GS: Seamless 360° Gaussian Splatting from Real-World Omnidirectional Images
- Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
- Improving Calibration for Long-Tailed Recognition
- Radioactive data: tracing through training
- Constructing Geographic and Long-term Temporal Graph for Traffic Forecasting
- Measuring Dataset Granularity
- How Does Batch Normalization Help Optimization?
- Symplectic convolutional neural networks
- Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning
- Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
- Causal Intervention for Weakly-Supervised Semantic Segmentation
- PRADA: Protecting against DNN Model Stealing Attacks
- Regularization With Stochastic Transformations and Perturbations for Deep Semi-Supervised Learning
- Multiscale Deep Equilibrium Models
- Mini-Batch Robustness Verification of Deep Neural Networks
- The GINN framework: a stochastic QED correspondence for stability and chaos in deep neural networks
- DeepOPF: A Deep Neural Network Approach for Security-Constrained DC Optimal Power Flow
- WIDER FACE: A Face Detection Benchmark
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Understanding the Effective Receptive Field in Deep Convolutional Neural Networks
- PanNuke Dataset Extension, Insights and Baselines
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Bladder Cancer Diagnosis with Deep Learning: A Multi-Task Framework and Online Platform
- VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
- Real-Time User-Guided Image Colorization with Learned Deep Priors
- Discretization-Aware Architecture Search
- PIPAL: a Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration
- Systematic evaluation of convolution neural network advances on the Imagenet
- Robust and Efficient Quantum Reservoir Computing with Discrete Time Crystal
- Enhancing Adversarial Example Transferability with an Intermediate Level Attack
- AutoHOOT: Automatic High-Order Optimization for Tensors
- Use of Transfer Learning and Wavelet Transform for Breast Cancer Detection
- Ontology Based Global and Collective Motion Patterns for Event Classification in Basketball Videos
- Deep learning for the semi-classical limit of the Schrödinger equation
- Optimistic and Pessimistic Neural Networks for Scene and Object Recognition
- Attended End-to-end Architecture for Age Estimation from Facial Expression Videos
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- Richer priors for infinitely wide multi-layer perceptrons
- Unpaired Image Translation via Adaptive Convolution-based Normalization
- Learning to Fuse Things and Stuff
- On the approximation of the solution of partial differential equations by artificial neural networks trained by a multilevel Levenberg-Marquardt method
- Frequency-adaptive tensor neural networks for high-dimensional multi-scale problems
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
- Snap-Snap: Taking Two Images to Reconstruct 3D Human Gaussians in Milliseconds
- Study and development of a Computer-Aided Diagnosis system for classification of chest x-ray images using convolutional neural networks pre-trained for ImageNet and data augmentation
- Bias-based Universal Adversarial Patch Attack for Automatic Check-out
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- Real-time Human Detection Model for Edge Devices
- Convolutional Neural Network Pruning with Structural Redundancy Reduction
- Deep Learning for Taxol Exposure Analysis: A New Cell Image Dataset and Attention-Based Baseline Model
- On the Communication Latency of Wireless Decentralized Learning
- Minimizing Task-Oriented Age of Information for Remote Monitoring with Pre-Identification
- Improving OCR using internal document redundancy
- MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
- Towards Precise End-to-end Weakly Supervised Object Detection Network
- Improved Mapping Between Illuminations and Sensors for RAW Images
- CP-mtML: Coupled Projection multi-task Metric Learning for Large Scale\n Face Retrieval
- Perceptual Adversarial Robustness: Defense Against Unseen Threat Models
- Neo
- Discriminative Noise Robust Sparse Orthogonal Label Regression-based Domain Adaptation
- Normalized Cut Loss for Weakly-supervised CNN Segmentation
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- Lottery Jackpots Exist in Pre-trained Models
- Neural Image Beauty Predictor Based on Bradley-Terry Model
- Defeating Catastrophic Forgetting via Enhanced Orthogonal Weights Modification
- Busy-Quiet Video Disentangling for Video Classification
- MOPS-Net: A Matrix Optimization-driven Network forTask-Oriented 3D Point Cloud Downsampling
- Capsule GAN Using Capsule Network for Generator Architecture
- MixHop: Higher-Order Graph Convolutional Architectures via Sparsified\n Neighborhood Mixing
- SVM and ELM: Who Wins? Object Recognition with Deep Convolutional Features from ImageNet
- On the Evaluation Metric for Hashing
- How to Train Your MAML to Excel in Few-Shot Classification
- torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models
- Representative Forgery Mining for Fake Face Detection
- CHAIN: Concept-harmonized Hierarchical Inference Interpretation of Deep Convolutional Neural Networks
- Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
- Open Set Domain Adaptation for Image and Action Recognition
- An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Spherical U-Net on Cortical Surfaces: Methods and Applications
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Towards Fewer Annotations: Active Learning via Region Impurity and Prediction Uncertainty for Domain Adaptive Semantic Segmentation
- FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- A Study of BFLOAT16 for Deep Learning Training
- Genuine multipartite entanglement verification with convolutional neural networks
- SpherePHD: Applying CNNs on a Spherical PolyHeDron Representation of 360\n degree Images
- Training with the Invisibles: Obfuscating Images to Share Safely for Learning Visual Recognition Models
- Deeply Exploit Depth Information for Object Detection
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Unsupervised Domain Adaptation for Semantic Segmentation via Low-level Edge Information Transfer
- NeuroView: Explainable Deep Network Decision Making
- Optimization problems for machine learning: A survey
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Uncertainty Propagation in Deep Neural Network Using Active Subspace
- D2-Mamba: Dual-Scale Fusion and Dual-Path Scanning with SSMs for Shadow Removal
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- SuryaBench: Benchmark Dataset for Advancing Machine Learning in Heliophysics and Space Weather Prediction
- Unsupervised Feature Learning by Cross-Level Instance-Group Discrimination
- The Pitfall of Evaluating Performance on Emerging AI Accelerators
- RED-NET: A Recursive Encoder-Decoder Network for Edge Detection
- NESTA: Hamming Weight Compression-Based Neural Proc. Engine
- Contextual Local Explanation for Black Box Classifiers
- Maximum-Entropy Adversarial Data Augmentation for Improved\n Generalization and Robustness
- Skin Cancer Classification: Hybrid CNN-Transformer Models with KAN-Based Fusion
- Efficient and Verifiable Privacy-Preserving Convolutional Computation for CNN Inference with Untrusted Clouds
- Attentive Deep Regression Networks for Real-Time Visual Face Tracking in Video Surveillance
- Seek and You Will Find: A New Optimized Framework for Efficient Detection of Pedestrian
- An ECC-based Fault Tolerance Approach for DNNs
- Joint Intent Detection and Slot Filling with Wheel-Graph Attention Networks
- Towards Generalizable Human Activity Recognition: A Survey
- One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic Control
- Illusions in Humans and AI: How Visual Perception Aligns and Diverges
- Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- ViPTT-Net: Video pretraining of spatio-temporal model for tuberculosis type classification from chest CT scans
- HALO: Learning to Prune Neural Networks with Shrinkage
- Importance Filtered Cross-Domain Adaptation
- Recent advances and applications of deep learning methods in materials science
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- PCA- and SVM-Grad-CAM for Convolutional Neural Networks: Closed-form Jacobian Expression
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- Itinerary-aware Personalized Deep Matching at Fliggy
- Weakly Supervised Estimation of Shadow Confidence Maps in Fetal Ultrasound Imaging
- DARB: A Density-Aware Regular-Block Pruning for Deep Neural Networks
- Activate Me!: Designing Efficient Activation Functions for Privacy-Preserving Machine Learning with Fully Homomorphic Encryption
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis
- Learning to Predict Trustworthiness with Steep Slope Loss
- CoMoNM: A Cost Modeling Framework for Compute-Near-Memory Systems
- Deciphering the 2016 U.S. Presidential Campaign in the Twitter Sphere: A Comparison of the Trumpists and Clintonists
- Bandicoot: A Templated C++ Library for GPU Linear Algebra
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
- FedNS: Improving Federated Learning for collaborative image classification on mobile clients
- Master Thesis: Neural Sign Language Translation by Learning Tokenization
- Two-Level Residual Distillation based Triple Network for Incremental Object Detection
- Zero-Shot Recognition through Image-Guided Semantic Classification
- Hybrid-Hierarchical Fashion Graph Attention Network for Compatibility-Oriented and Personalized Outfit Recommendation
- health effects of dielectric gases: preliminary report
- Understanding the Error in Evaluating Adversarial Robustness
- Mobile-Friendly Deep Learning for Plant Disease Detection: A Lightweight CNN Benchmark Across 101 Classes of 33 Crops
- Self-Paced Uncertainty Estimation for One-shot Person Re-Identification
- NASA: Neural Articulated Shape Approximation
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Novel View Synthesis using DDIM Inversion
- Predicting the Mumble of Wireless Channel with Sequence-to-Sequence Models
- EventNet: Asynchronous Recursive Event Processing
- A novel data-driven approach for transient stability prediction of power systems considering the operational variability
- ModiPick: SLA-aware Accuracy Optimization For Mobile Deep Inference
- HyperTea: A Hypergraph-based Temporal Enhancement and Alignment Network for Moving Infrared Small Target Detection
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
- Rethink ReLU to Training Better CNNs
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- 3D latent diffusion models for parameterizing and history matching multiscenario facies systems
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- Deep Sparse Subspace Clustering
- Non-uniqueness phenomenon of object representation in modelling IT cortex by deep convolutional neural network (DCNN)
- Convolutional Hough Matching Networks for Robust and Efficient Visual Correspondence
- Improving Automated COVID-19 Grading with Convolutional Neural Networks in Computed Tomography Scans: An Ablation Study
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- A System-Level Solution for Low-Power Object Detection
- Joint Object and Part Segmentation using Deep Learned Potentials
- Hardware-Efficient Structure of the Accelerating Module for Implementation of Convolutional Neural Network Basic Operation
- Explainable AI Technique in Lung Cancer Detection Using Convolutional Neural Networks
- ALTIS: Modernizing GPGPU Benchmarking
- Lifted Relational Neural Networks
- TextureWGAN: Texture Preserving WGAN with MLE Regularizer for Inverse Problems
- Adaptively Denoising Proposal Collection for Weakly Supervised Object Localization
- IPG: Incremental Patch Generation for Generalized Adversarial Patch Training
- Hey Human, If your Facial Emotions are Uncertain, You Should Use Bayesian Neural Networks!
- Temporal Proximity induces Attributes Similarity
- NIRMAL Pooling: An Adaptive Max Pooling Approach with Non-linear Activation for Enhanced Image Classification
- MixSearch: Searching for Domain Generalized Medical Image Segmentation Architectures
- Robust Pollen Imagery Classification with Generative Modeling and Mixup Training
- Rip van Winkle's Razor: A Simple Estimate of Overfit to Test Data
- Spurious Local Minima Are Common for Deep Neural Networks with Piecewise Linear Activations
- Drop-Activation: Implicit Parameter Reduction and Harmonic Regularization
- Source Printer Identification from Document Images Acquired using Smartphone
- How benign is benign overfitting?
- DeepFeatIoT: Unifying Deep Learned, Randomized, and LLM Features for Enhanced IoT Time Series Sensor Data Classification in Smart Industries
- Fast Haar Transforms for Graph Neural Networks
- Modeling the Uncertainty in Electronic Health Records: a Bayesian Deep Learning Approach
- Wide Neural Networks Forget Less Catastrophically
- MPT: Motion Prompt Tuning for Micro-Expression Recognition
- Deep Learning for Automated Identification of Vietnamese Timber Species: A Tool for Ecological Monitoring and Conservation
- Towards Train-Test Consistency for Semi-supervised Temporal Action Localization
- A Finer Calibration Analysis for Adversarial Robustness
- Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Generating Training Data for Denoising Real RGB Images via Camera Pipeline Simulation
- Person Identification with Visual Summary for a Safe Access to a Smart Home
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Efficient motion-based metrics for video frame interpolation
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- Continual Learning in Neural Networks
- ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
- Deep Transform: Cocktail Party Source Separation via Complex Convolution in a Deep Neural Network
- Machine Learning Applications for Precision Agriculture: A Comprehensive Review
- Peer-Assisted Robotic Learning: A Data-Driven Collaborative Learning Approach for Cloud Robotic Systems
- Collegial Ensembles
- Transfer Learning for Melanoma Detection: Participation in ISIC 2017 Skin Lesion Classification Challenge
- Toward Lifelong Learning in Equilibrium Propagation: Sleep-like and Awake Rehearsal for Enhanced Stability
- Hashing as Tie-Aware Learning to Rank
- AI-Skin : Skin Disease Recognition based on Self-learning and Wide Data Collection through a Closed Loop Framework
- PAC-GAN: An Effective Pose Augmentation Scheme for Unsupervised Cross-View Person Re-identification
- Image selective encryption analysis using mutual information in CNN based embedding space
- Forecasting Transportation Network Speed Using Deep Capsule Networks with Nested LSTM Models
- Continuous Perception for Classifying Shapes and Weights of Garmentsfor Robotic Vision Applications
- Automated Decision-based Adversarial Attacks
- Controlling Covariate Shift using Balanced Normalization of Weights
- Flexible Dataset Distillation: Learn Labels Instead of Images
- Implementation of Deep Neural Networks to Classify EEG Signals using Gramian Angular Summation Field for Epilepsy Diagnosis
- Sparse tree-based initialization for neural networks
- Reinforcement learning for batch bioprocess optimization
- Defending Against Image Corruptions Through Adversarial Augmentations
- Interpretable CNNs for Object Classification
- Regularized Adaptation for Stable and Efficient Continuous-Level\n Learning on Image Processing Networks
- An Intelligent Group Event Recommendation System in Social networks
- Understanding the Limitations of Variational Mutual Information Estimators
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
- Self-Supervised Representation Learning for Visual Anomaly Detection
- Active Learning for Breast Cancer Identification
- Cooperative Bi-path Metric for Few-shot Learning
- Lightweight Multi-Scale Feature Extraction with Fully Connected LMF Layer for Salient Object Detection
- Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
- A Bayesian Data Augmentation Approach for Learning Deep Models
- Multi-scale Domain-adversarial Multiple-instance CNN for Cancer Subtype Classification with Unannotated Histopathological Images
- PU-Net: Point Cloud Upsampling Network
- Spatial-Separated Curve Rendering Network for Efficient and High-Resolution Image Harmonization
- Learning to Generate 3D Shapes with Generative Cellular Automata
- Denoising and Verification Cross-Layer Ensemble Against Black-box Adversarial Attacks
- MISS: Multi-Interest Self-Supervised Learning Framework for Click-Through Rate Prediction
- Diffeomorphic Neural Operator Learning
- Confounder Identification-free Causal Visual Feature Learning
- Position: Ideas Should be the Center of Machine Learning Research
- Multimodal learning with next-token prediction for large multimodal models
- Deep Convolutional Decision Jungle for Image Classification
- Lung sounds classification using convolutional neural networks
- Ophthalmic diagnosis using deep learning with fundus images – A critical review
- Medical knowledge embedding based on recursive neural network for multi-disease diagnosis
- Enhancing Recognition and Categorization of Skin Lesions with Tailored Deep Convolutional Networks and Robust Data Augmentation Techniques
- Fast Calculation of Probabilistic Power Flow: A Model-based Deep Learning Approach
- Efficient Cloth Simulation using Miniature Cloth and Upscaling Deep\n Neural Networks
- Reinforced Evolutionary Neural Architecture Search
- Alleviating Mode Collapse in GAN via Diversity Penalty Module
- Affective Image Content Analysis: Two Decades Review and New Perspectives
- Parallel Deep Neural Networks Have Zero Duality Gap
- Facial Information Analysis Technology for Gender and Age Estimation
- Being-ahead: Benchmarking and Exploring Accelerators for Hardware-Efficient AI Deployment
- Learning Representations of Satellite Images with Evaluations on Synoptic Weather Events
- Recurrent Deep Differentiable Logic Gate Networks
- Zero-Shot Learning by Convex Combination of Semantic Embeddings
- Synthetic Data Generation for Emotional Depth Faces: Optimizing Conditional DCGANs via Genetic Algorithms in the Latent Space and Stabilizing Training with Knowledge Distillation
- Multi-view Gaze Target Estimation
- Re-evaluating Evaluation
- Regularized Evolutionary Population-Based Training
- Optimising the Performance of Convolutional Neural Networks across Computing Systems using Transfer Learning
- A geometry-inspired decision-based attack
- Keep It Real: Challenges in Attacking Compression-Based Adversarial Purification
- Input Invex Neural Network
- A Study of Gender Classification Techniques Based on Iris Images: A Deep Survey and Analysis
- Real-time Detection of Practical Universal Adversarial Perturbations
- Multi-Labelled Value Networks for Computer Go
- Rotation Equivariant Arbitrary-scale Image Super-Resolution
- Digital Twin Channel-Aided CSI Prediction: An Environment-Based Subspace Extraction Approach for Achieving Low Overhead and High Robustness
- ULU: A Unified Activation Function
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Self-Error Adjustment: Theory and Practice of Balancing Individual Performance and Diversity in Ensemble Learning
- Smaller Models, Better Generalization
- Toward Errorless Training ImageNet-1k
- Pulmonary embolism identification in computerized tomography pulmonary angiography scans with deep learning technologies in COVID-19 patients
- Robust Processing-In-Memory Neural Networks via Noise-Aware Normalization
- Metric Learning in an RKHS
- PoreFlow-Net: A 3D convolutional neural network to predict fluid flow through porous media
- DeePore: A deep learning workflow for rapid and comprehensive characterization of porous materials
- Adaptive Neuron-wise Discriminant Criterion and Adaptive Center Loss at Hidden Layer for Deep Convolutional Neural Network
- Automated ultrasound doppler angle estimation using deep learning
- Energy-Efficient Real-Time 4-Stage Sleep Classification at 10-Second Resolution: A Comprehensive Study
- Age-Diverse Deepfake Dataset: Bridging the Age Gap in Deepfake Detection
- Slice or the Whole Pie? Utility Control for AI Models
- Deep Cross Residual Learning for Multitask Visual Recognition
- Deep learning framework for crater detection and identification on the Moon and Mars
- Training Deep Neural Networks via Branch-and-Bound
- Automatic Detection of Coronavirus Disease (COVID-19) in X-ray and CT\n Images: A Machine Learning-Based Approach
- CADD: Context aware disease deviations via restoration of brain images using normative conditional diffusion models
- Relationship-Embedded Representation Learning for Grounding Referring Expressions
- Deep Reasoning with Multi-Scale Context for Salient Object Detection
- A Scalable Optimization Mechanism for Pairwise based Discrete Hashing
- Zero Shot Domain Adaptive Semantic Segmentation by Synthetic Data Generation and Progressive Adaptation
- The Power of Many: Synergistic Unification of Diverse Augmentations for Efficient Adversarial Robustness
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- RoIFusion: 3D Object Detection from LiDAR and Vision
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- TrackNet: A Deep Learning Network for Tracking High-speed and Tiny Objects in Sports Applications
- The Geometry of Cortical Computation: Manifold Disentanglement and Predictive Dynamics in VCNet
- Fully-Convolutional Siamese Networks for Object Tracking
- Domain Adaptation with Incomplete Target Domains
- Toward Practical Equilibrium Propagation: Brain-inspired Recurrent Neural Network with Feedback Regulation and Residual Connections
- Infrared Object Detection with Ultra Small ConvNets: Is ImageNet Pretraining Still Useful?
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- Graph Classification Based on Skeleton and Component Features
- Towards Compact and Robust Deep Neural Networks
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- After the Party: Navigating the Mapping From Color to Ambient Lighting
- Deep Learning Face Representation by Joint Identification-Verification
- Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
- Tackling Ill-posedness of Reversible Image Conversion with Well-posed Invertible Network
- Combining Ensembles and Data Augmentation can Harm your Calibration
- Defending Against Adversarial Attacks Using Random Forests
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- Composite Quantization
- Improving Noise Efficiency in Privacy-preserving Dataset Distillation
- RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
- Pulse Shape Discrimination Algorithms: Survey and Benchmark
- Discrete Rotation Equivariance for Point Cloud Recognition
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Deep Built-Structure Counting in Satellite Imagery Using Attention Based\n Re-Weighting
- Reconstructing Trust Embeddings from Siamese Trust Scores: A Direct-Sum Approach with Fixed-Point Semantics
- Intensity augmentation for domain transfer of whole breast segmentation\n in MRI
- Increasing the Generalisation Capacity of Conditional VAEs
- Robust Visual Knowledge Transfer via EDA
- Translating Math Formula Images to LaTeX Sequences Using Deep Neural Networks with Sequence-level Training
- Partial observations and conservation laws: Grey-box modeling in\n biotechnology and optogenetics
- Yelp Food Identification via Image Feature Extraction and Classification
- Fusion Sampling Validation in Data Partitioning for Machine Learning
- A Simple and Effective Method for Uncertainty Quantification and OOD Detection
- Imbalanced Deep Learning by Minority Class Incremental Rectification
- Entropy-Based Uncertainty Calibration for Generalized Zero-Shot Learning
- Learning Competitive and Discriminative Reconstructions for Anomaly Detection
- Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
- Initialization Strategies of Spatio-Temporal Convolutional Neural\n Networks
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- Variational Autoencoder-Based Black-Box Adversarial Attack on Collaborative DNN Inference
- Sinusoidal Approximation Theorem for Kolmogorov-Arnold Networks
- ZipNet-GAN: Inferring Fine-grained Mobile Traffic Patterns via a Generative Adversarial Neural Network
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- L-GTA: Latent Generative Modeling for Time Series Augmentation
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- Particle reconstruction of volumetric particle image velocimetry with strategy of machine learning
- Flow Contrastive Estimation of Energy-Based Models
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Do We Need Fully Connected Output Layers in Convolutional Networks?
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- Initialization Using Perlin Noise for Training Networks with a Limited Amount of Data
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- How deep should be the depth of convolutional neural networks: a backyard dog case study
- CADDA: Class-wise Automatic Differentiable Data Augmentation for EEG Signals
- SemiNLL: A Framework of Noisy-Label Learning by Semi-Supervised Learning
- Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents
- Graph Lineages and Skeletal Graph Products
- An Artificial Intelligence-Based System to Assess Nutrient Intake for Hospitalised Patients
- JQF: Optimal JPEG Quantization Table Fusion by Simulated Annealing on Texture Images and Predicting Textures
- Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
- Tricks and Plug-ins for Gradient Boosting in Image Classification
- Multi-lane Detection Using Instance Segmentation and Attentive Voting
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- Machine Learning for Detecting Data Exfiltration: A Review
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- DACA-Net: A Degradation-Aware Conditional Diffusion Network for Underwater Image Enhancement
- Crowding in humans is unlike that in convolutional neural networks
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Agentic Privacy-Preserving Machine Learning
- RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Knowledge-augmented Column Networks: Guiding Deep Learning with Advice
- StarNet: Targeted Computation for Object Detection in Point Clouds
- FDNAS: Improving Data Privacy and Model Diversity in AutoML
- Synthesizing Optimal Parallelism Placement and Reduction Strategies on Hierarchical Systems for Deep Learning
- TIR-Diffusion: Diffusion-based Thermal Infrared Image Denoising via Latent and Wavelet Domain Optimization
- Moiré Zero: An Efficient and High-Performance Neural Architecture for Moiré Removal
- Object Recognition Datasets and Challenges: A Review
- Distinct contributions of functional and deep neural network features to representational similarity of scenes in human brain and behavior. [europepmc]
- Computer-aided detection in chest radiography based on artificial intelligence: a survey. [europepmc]
- Skin Cancer Classification Using Convolutional Neural Networks: Systematic Review. [europepmc]
- A U-Net Deep Learning Framework for High Performance Vessel Segmentation in Patients With Cerebrovascular Disease. [europepmc]
- Artificial intelligence and machine learning in clinical development: a translational perspective. [europepmc]
- Key Topics in Molecular Docking for Drug Design. [europepmc]
- Artificial Intelligence in Lung Cancer Pathology Image Analysis. [europepmc]
- Evaluation of Combined Artificial Intelligence and Radiologist Assessment to Interpret Screening Mammograms. [europepmc]
- 3D Deep Learning on Medical Images: A Review. [europepmc]
- Deep learning encodes robust discriminative neuroimaging representations to outperform standard machine learning. [europepmc]
- DeepTCR is a deep learning framework for revealing sequence concepts within T-cell repertoires. [europepmc]
- Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. [europepmc]
- The language of proteins: NLP, machine learning & protein sequences. [europepmc]
- Epileptic Seizures Detection Using Deep Learning Techniques: A Review. [europepmc]
- GNINA 1.0: molecular docking with deep learning. [europepmc]
- Combining Machine Learning and Computational Chemistry for Predictive Insights Into Chemical Systems. [europepmc]
- TransMed: Transformers Advance Multi-Modal Medical Image Classification. [europepmc]
- A review on deep learning in medical image analysis. [europepmc]
- Recent Advances in Electrochemical Biosensors: Applications, Challenges, and Future Scope. [europepmc]
- The Role of Artificial Intelligence in Early Cancer Diagnosis. [europepmc]
- Radiology artificial intelligence: a systematic review and evaluation of methods (RAISE). [europepmc]
- A fully automatic AI system for tooth and alveolar bone segmentation from cone-beam CT images. [europepmc]
- Explainable medical imaging AI needs human-centered design: guidelines and evidence from a systematic review. [europepmc]
- Redefining Radiology: A Review of Artificial Intelligence Integration in Medical Imaging. [europepmc]
- How Artificial Intelligence Is Shaping Medical Imaging Technology: A Survey of Innovations and Applications. [europepmc]
Related