Deep Residual Learning for Image Recognition
2015/12/10 by Kaiming He, He, Kaiming, Xiangyu Zhang +5 · 3 voices · 4214 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques
paper · pdf · doi:10.48550/arxiv.1512.03385
Abstract
A curated reading list of foundational machine-learning papers, blog posts, and course notes. Compiled from a Hacker News comment by user cuteboi (news.ycombinator.com/item?id=48822131). This record is a bibliographic list only (titles and links); it does not redistribute the referenced works.
Citations
Cited by
- PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM
- HYCO: Hybrid-Cooperative Learning for Data-Driven PDE Modeling
- Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization
- Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
- Correlating Cross-Iteration Noise for DP-SGD using Model Curvature
- Automatic Stability and Recovery for Neural Network Training
- PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework
- FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
- Low-Altitude Channel Multipath Prediction via Panoramic Perception and Vision-Language Model
- From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning
- Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation
- Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
- Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 Refinement
- On the Convergence of Stochastic Low-Rank Adaptation
- InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
- LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models
- Projection Pursuit CPCANet for Domain Generalization
- Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation
- ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors
- Bowel Obstruction Detection and Localization on Abdominal CT with Deep Learning
- JustDepth: Real-Time Radar-Camera Depth Estimation with Single-Scan LiDAR Supervision
- Toward High-Fidelity 3D Point-Cloud Learning for Brain Folding Morphology Prediction Using Trans-Unet
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- Exact Neural-Network Representations of the Motzkin States
- Pretraining Recurrent Networks without Recurrence
- Deep Convolutional Large-Margin ℓp-SVDD for Visual Anomaly Detection
- A Statistical Multi-Objective Framework for Assessing Sensitivity of Radiomic AI Models to Acquisition Parameters
- Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features
- Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed Recognition
- SocialPulse: On-Device Detection of Social Interactions in Naturalistic Settings Using Smartwatch Sensing
- Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion Models
- DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
- DAUPNet: Domain-Aware Uncertainty Modeling for Reliable Prototype Discrimination in Cross-Domain Few-Shot Semantic Segmentation
- Traceback Translators Against Forgetting in Continual Fake Speech Detection
- Monte Carlo Dropout Uncertainty and Entropy-Thresholded Selective Prediction for Architecture-Agnostic Brain Tumor MRI Triage
- SLT: Robust Quantum Neural Networks for Noisy-Label Medical Image Classification via Supermartingale-based Label Transition
- AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
- PILD: Physics-Informed Learning via Diffusion
- Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
- Auto-adaptive Resonance Equalization using Dilated Residual Networks
- Differentially Private Neural Network Training Under the Hidden State Assumption
- Online Variance Reduction for Domain Adaptation on Streaming Data
- A convergence result of a continuous model of deep learning via a Łojasiewicz--Simon inequality
- Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
- Back to Back with a Copy: A Computational Analysis of AI-Generated Visual Contemporary Art Pastiches
- WiFi Sensing via Reservoir Computing
- Test Case Prioritization for DNNs via Neural Collapse Instability
- Ambiguity-Resolved Micro-Doppler Construction for Asynchronous Bistatic Sensing
- Conjugate Gradient Unrolled Network with PSF Conditioning for Non-Diagonal Data Fidelity in CASSI Reconstruction
- AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing
- IConE: Batch Independent Collapse Prevention for Self-Supervised Representation Learning
- Variance-reduced Domain Adaptation using Paired Sampling
- Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation
- Variational Learning of Physical Intuition from a Few Observations: Charting Manifolds of Variational Physics
- Label-Noise Resistant Learning via Optimal Brain Damage Masking
- Weather Estimation for Integrated Sensing and Communication
- ReRAM-aware Model Finetuning addressing I-V Non-linearity and Retention Errors
- The Well-Tempered Likelihood: Honest Confidence Intervals for Misspecified Models
- End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers
- Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness
- PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
- One Round Is All You Need: Analytic Federated Learning for Task-Heterogeneous Multi-Label Medical Image Classification
- Pipelined Gradient Coding
- Detecting Neural Network Failures through Spectral Analysis of Internal Activations
- Glyce: Glyph-vectors for Chinese Character Representations
- HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
- Great X: A Unified Multi-Modal Simulator Bridging the Sim2Real Gap for 6G
- From Classification to Localization and Clinical Validation: Large-Scale Development of a Deep Learning System for Thoracic Disease Detection on Chest Radiographs in Thailand
- TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
- Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification
- Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification
- LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation
- Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion
- ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning
- End-to-End Differential Privacy in Training Deep Neural Network Classifiers
- A Deep Learning Framework for Predicting Solar EUV Irradiance During Significant Flares
- MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
- Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention
- Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight Aggregation
- TaskTok: Delving into Task Tokens for Task-driven Image Restoration
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- Decafs: Disentangled Conditional adversarial Flows
- Posterior Samplings are Missing Modalities Generators for Medical Image Translation
- Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
- Dual Attention Residuals
- Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making
- Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
- Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention
- GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval
- AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models
- SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
- Finite-Agent Stochastic Differential Games on Large Graphs: II. Graph-Based Architectures
- Bio-SFT: Asymmetric Cortical Guidance and Retinal Adaptation for Robust HDR Reconstruction
- An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers
- LieBN: Batch Normalization over Lie Groups
- InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation
- DECIS: Dual-Evidence Corrective Verification for Interpretable Strabismus Diagnostic Decision-Making
- DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition
- Scalable Model-Assisted Multi-Target Estimation in Large Image Collections
- Benchmarking NACTI Species Recognition in Long-Tailed Regimes
- SGN: A Similarity-based Generative Network for Data Generation under Distribution Shift
- Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation
- (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure
- Early Yield Prediction for Sugar Beet Fields using Satellite Data -- Learnings from Specialized Vision Transformers
- Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator
- COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering
- Efficient Tuning Before Low-Bit Post-Training Quantization for Stochastic Gradient Descent-optimized Models
- mHC-GNN: Manifold-Constrained Hyper-Connections for Graph Neural Networks
- The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Manipulation
- Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning
- The Calibration Channel Determines the Bayes-Error Proxy: An Exact Law for Temperature-Induced Distortion
- EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database
- Dice Loss for Data-imbalanced NLP Tasks
- Rethinking Feature Reliance Evaluation with Semantically Matched Suppression
- Provably Lossless Acceleration of DNN Mutation Testing via Memoization
- Sequential Attention-based Sampling for Histopathological Analysis
- Three-Body Scattering for Generative Modeling
- PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network Encryption
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
- The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
- History-Aware Transformation of ReID Features for Multiple Object Tracking
- AOE: Exhaustive Out-of-Distribution Detection via Recalibrating Outlier Labels
- Cross-Coordinate Correspondence Pruning for Image-to-Point Cloud Registration
- HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
- Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures
- ADEPT: Architecture-Driven Energy-Efficient CNN Fine-Tuning on PIM Accelerators
- The Information Content of Krylov Observables: A Machine Learning Approach
- TinyGLASS: Real-Time Self-Supervised In-Sensor Anomaly Detection
- A Research Prototype for Closed-Loop Generative Design of Customized Foot Orthoses via Semantic-Physics Alignment
- Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
- Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
- When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection
- Do Value Vectors in Deep Layers Need Context from the Residual Stream?
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Unsupervised Incremental Learning Using Confidence-Based Pseudo-Labels
- PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion
- Robust Losses from Univariate Base Functions for Noisy-Label Learning
- Prompt-Guided Foundation Model Tuning for Pathology Image Classification
- MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
- Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery
- Scalable Open-Source Visuotactile Sensor for 6-Axis Contact Wrench Estimation in Tensegrity Robots
- Density-Informed Pseudo-Counts for Calibrated Evidential Deep Learning
- STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- Hybrid Machine Learning for Articulation Angle Estimation of Truck-Semitrailer Combinations
- Online-Score-Aided Federated Learning for Resource-Constrained Wireless Clients with Continual Data Arrival
- Thermal Topology Collapse: Universal Physical Patch Attacks on Infrared Vision Systems
- Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
- Backpropagation-Free Trunk Training via the Split Forward Gradients
- Task-Oriented Communication with Hybrid-Precision Models
- Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training
- GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking
- Dataset Distillation by Influence Matching
- Code-Poisoning Property Inference Attacks
- DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
- A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
- HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection
- From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data
- Toward a mechanistic understanding of inference in visual cortex and diffusion models
- Disentangling Model and Human Data Uncertainty in Apparent Facial Age Estimation
- Measurement of the branching ratio of the K+→π+νν decay
- In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
- Improving Backward Conformal Prediction via Non-Conformity Score Transformation
- Conformal Graph Prediction with Z-Gromov-Wasserstein Distances
- Structure-Induced Information for Rerooting Levin Tree Search
- Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
- Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models
- CoDi -- an exemplar-conditioned diffusion model for low-shot counting
- VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging
- An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches
- Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance
- MuxGel: Simultaneous Dual-Modal Visuo-Tactile Sensing via Spatially Multiplexing and Deep Reconstruction
- Diffusion models recover accurate mixture weights despite score function insensitivity
- xHC: Expanded Hyper-Connections
- The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers
- Scalable Training of Continuous-Time Spiking Neural Networks with Differentiable Spike-Time Discretization
- Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
- Factorized Neural Operators Decompose Dynamic and Persistent Responses
- EdgeFaaS: A Function-based Framework for Edge Computing
- Intentional Electromagnetic Interference Attacks on Facial Recognition
- Knowing You at First Glance: Inferring Apparent Personality from Faces
- From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models
- Estimating Time-Dependent COVID-19 Parameters Using Kolmogorov-Arnold Network and Fourier Series
- GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
- A Step Forward Towards Trustworthy Risk-Aware Facial Retrieval (RA-FR)
- 3D Lane Detection with Odometry for High-Speed Vehicle Racing
- Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks
- Predicting Groundwater Arsenic Concentrations Using Graph Neural Networks
- An Empirical Study of Handcrafted Feature Learning and Convolutional Neural Networks for Facial Expression Recognition
- Representation recycling for streaming video analysis
- CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
- AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis
- Reducing Per-Sample Harm in Stochastic Optimization
- Prediction of the Hubbard <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mi>U</mml:mi> </mml:math> parameter from scanning tunneling microscopy images of moiré systems using image recognition
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Random network structure stabilizes neural manifolds
- What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
- Training-Free Bayesian Filtering with Generative Emulators
- 123D: Unifying Multi-Modal Autonomous Driving Data at Scale
- Deep learning reveals genomic regions introgressed between two recurrently hybridizing lynx species
- The illusory simplicity of the feedforward pass: evidence for the dynamical nature of stimulus encoding along the primate ventral stream
- CURE: Privacy-Preserving Split Learning Done Right
- Information bottleneck for learning the phase space of dynamics from high-dimensional experimental data
- Energy-Efficient Information Representation in MNIST Classification Using Biologically Inspired Learning
- Do Machines Fail Like Humans? A Human-Centred Out-of-Distribution Spectrum for Mapping Error Alignment
- Introducing TropiCam‐AI: A taxonomically flexible automated classifier of Neotropical arboreal mammals and birds from camera‐trap data
- Li-ViP3D++: Query-Gated Deformable Camera–LiDAR Fusion for End-to-End Perception and Trajectory Prediction
- Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
- Benchmarking for practice: Few-shot time-series crop-type classification on the EuroCropsML dataset
- FSL-HDnn: A 40-nm Few-Shot On-Device Learning Accelerator With Integrated Feature Extraction and Hyperdimensional Computing
- Learnable Quantum Efficiency Filters for Urban Hyperspectral Segmentation
- mHC: Manifold-Constrained Hyper-Connections
- SomnoNet: A Lightweight and Interpretable Framework for Sleep Staging Using Single-Channel EEG
- MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
- Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex
- A Primer on Quantum Machine Learning
- ThRIve: Thermally Robust CNN Inference via Low-Rank Adaptation in Heterogeneous PIM Architectures
- Simulating Automotive Radar with Lidar and Camera Inputs
- StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
- A Distributed Emulation Environment for In-Memory Computing Systems
- Model-Guided Microstimulation Steers Primate Visual Behavior
- FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems
- ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
- Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment
- From basic affordances to symbolic thought: A computational phylogenesis of biological intelligence.
- Fast weight programming and linear transformers: from machine learning to neurobiology
- What Neuroscience Can Teach AI About Learning in Continuously Changing Environments
- What does really matter in image goal navigation?
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- HRRRCast: A Data-Driven Emulator for Regional Weather Forecasting at Convection-Allowing Scales
- Adopting a human developmental visual diet yields robust, shape-based AI vision
- Hierarchical Reasoning Model
- WorldVLA: Towards Autoregressive Action World Model
- ShapeEmbed: a self-supervised learning framework for 2D contour quantification
- FLAIR-HUB: Large-scale Multimodal Dataset for Land Cover and Crop Mapping
- NeXtBrain: Combining local and global feature learning for brain tumor classification
- Facial albedo generation for 3D face reconstruction from a single image via a coarse-to-fine approach
- Transformers are Graph Neural Networks
- Are Statistical Methods Obsolete in the Era of Deep Learning? A Study of ODE Inverse Problems
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Perception Encoder: The best visual embeddings are not at the output of the network
- In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
- AssistanceZero: Scalably Solving Assistance Games
- NNN: Next-Generation Neural Networks for Marketing Measurement
- Accurate phenotyping of luminal A breast cancer in magnetic resonance imaging: A new 3D CNN approach
- RANa: Retrieval-Augmented Navigation
- NeuRaLaTeX: A machine learning library written in pure LaTeX
- Transformers without Normalization
- Shared Global and Local Geometry of Language Model Embeddings
- Advancements in automated nuclei segmentation for histopathology using you only look once-driven approaches: A systematic review
- Representational Similarity via Interpretable Visual Concepts
- ILIAS: Instance-Level Image retrieval At Scale
- Traveling Waves Integrate Spatial Information Through Time
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Distributional Diffusion Models with Scoring Rules
- The functional role of oscillatory dynamics in neocortical circuits: A computational perspective
- Rethinking Early Stopping: Refine, Then Calibrate
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- An AI-Powered Autonomous Underwater System for Sea Exploration and Scientific Research
- LC4-DViT: Land-cover Creation for Land-cover Classification with Deformable Vision Transformer
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Distribution Matching Variational AutoEncoder
- The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
- Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
- Image Denoising Using Global and Local Circulant Representation
- Deep learning reveals cross-platform consistency in street view imagery for urban perception mapping
- Multi-label Classification with Panoptic Context Aggregation Networks
- SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel Context
- World Engine: Towards the Era of Post-Training for Autonomous Driving
- Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images
- VGGT-Ω
- Rethinking Intrinsic Dimension Estimation in Neural Representations
- Glacier shrinkage in the Peruvian and Bolivian Andes from a deep learning-based multi-temporal inventory (2016–2024)
- Directly Constructing Low-Dimensional Solution Subspaces in Deep Neural Networks
- Diffusion priors enhanced velocity model building from time-lag images using a neural operator
- A new adaptive two-layer model for opinion spread in hypergraphs: parameter sensitivity and estimation
- Explainable Neural Inverse Kinematics for Obstacle-Aware Robotic Manipulation: A Comparative Analysis of IKNet Variants
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- LIMO: Low-Power In-Memory-Annealer and Matrix-Multiplication Primitive for Edge Computing
- Exploring Syn-to-Real Domain Adaptation for Military Target Detection
- Energy and Memory-Efficient Federated Learning With Ordered Layer Freezing
- EIR: Enhanced Image Representations for Medical Report Generation
- Certifying the Right to Be Forgotten: Primal-Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning
- With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
- JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
- Fusion or Confusion? Multimodal Complexity Is Not All You Need
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Chessformer: A Unified Architecture for Chess Modeling
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- Galaxy Zoo Evo: 1 million human-annotated images of galaxies
- Spatial Interpolation of Room Impulse Responses based on Deeper Physics-Informed Neural Networks with Residual Connections
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples
- Plug In, Grade Right: Psychology-Inspired AGIQA
- Investigating Deep Learning Models for Ejection Fraction Estimation from Echocardiography Videos
- EmoCtrl: Controllable Emotional Image Content Generation
- Bright 4B: Scaling Hyperspherical Learning for Segmentation in 3D Brightfield Microscopy
- DeFloMat: Detection with Flow Matching for Stable and Efficient Generative Object Localization
- CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
- A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
- Frequency Regularization: Unveiling the Spectral Inductive Bias of Deep Neural Networks
- iOS as Acceleration
- Joint Flow Matching for Generator-Consistent Classification
- Attention Residuals
- Greedy dynamical meta-learning
- Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
- Statistics of natural scenes shape contextual modulation in the visual cortex
- To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
- AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
- When Rates Are Geometric: Rate-Certificate Transfer for Contact Splittings in Optimization
- Multi‐angle, cross‐domain fusion strategy enhances automated insect identification and hierarchical categorization: a case study on assassin bugs (Hemiptera: Reduviidae)
- Identifying Dolphin Whistle Producers With Deep Learning: Moving Beyond Signature Whistles
- Adaptive Acoustic Monitoring for Endangered Cook Inlet Beluga Whales in Complex Soundscapes
- End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
- SE-MLP Model for Predicting Prior Acceleration Features in Penetration Signals
- PINNs for Electromagnetic Wave Propagation
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
- Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations
- 6G-Enabled Smart Railways
- How Does a Deep Neural Network Look at Lexical Stress?
- A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
- A Scoping Review of Machine Learning Applications in Power System Protection and Disturbance Management
- Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
- Adaptive Data Admission and Retention for Streaming Federated Learning
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression
- Investigating the Visual Cues of CNNs for Vascular Segmentation: A Case Study in Microscopy and Fundus Imaging
- SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation
- EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability
- LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection
- DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
- Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning
- DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- AI Empowered Communication and Radar Modulation Recognition: A Survey
- Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
- Self-Supervised Learning from Noisy and Incomplete Data
- Automated Numerical Stability Analysis of Deep Learning Operators
- ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
- MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
- Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
- AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
- Balanced Soft mixture-of-expert model for Glaucoma Detection
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- Anti-Backdoor Coreset Selection via Cumulative Entropy
- Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors
- ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
- Multiclass Classification without Labels via Posterior Simplex Geometry
- OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
- PLATO: Pointer Learner for Agent and Task Openness
- Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
- Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix
- Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting
- Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
- Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
- Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification
- A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography
- CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy
- Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems
- Explainable Multimodal Regression via Information Decomposition
- Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer
- Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
- DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification
- Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation
- Lexical discovery in unknown environments orchestrated by Large Language Models
- Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
- Toric code under antiferromagnetic isotropic Heisenberg interactions
- Latent Space Probing for Adult Content Detection in Video Generative Models
- Scalable Neural Decoders for Practical Fault-Tolerant Quantum Computation
- Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers
- Deep Delta Learning
- WiFo-M2: Empower Wireless Communications With Plug-and-Play Environment Sensing via Foundation Model
- Benchmarking deep learning models for Raman spectroscopy across open-source datasets
- AI for Mycetoma Diagnosis in Histopathological Images: The MICCAI 2024 Challenge
- LVLM-Aided Alignment of Task-Specific Vision Models
- Data relativistic uncertainty framework for low-illumination anime scenery image enhancement
- Am I Confused or Is This Confusing?: Deep Ensembles for ENSO Uncertainty Quantification
- LibContinual: A Comprehensive Library towards Realistic Continual Learning
- A Model of Causal Explanation on Neural Networks for Tabular Data
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image Segmentation
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
- First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
- Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- GPF-Net: Gated Progressive Fusion Learning for Polyp Re-Identification
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- GraviBERT: Transformer-based inference for gravitational-wave time series
- Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
- Neural Network-Assisted RIS Weight Optimization for Spatial Nulling in Distorted Reflector Antenna Systems
- UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer
- A Learning Stability Profile for Finite-Dimensional Learning Dynamics
- Hierarchical Modeling Approach to Fast and Accurate Table Recognition
- Understanding Scaling Laws in Deep Neural Networks via Feature Learning Dynamics
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- Granular-ball Guided Masking: Structure-aware Data Augmentation
- X-ray Insights Unleashed: Pioneering the Enhancement of Multi-Label Long-Tail Data
- SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
- ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
- Defending against adversarial attacks using mixture of experts
- NULLBUS: Multimodal Mixed-Supervision for Breast Ultrasound Segmentation via Nullable Global-Local Prompts
- Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
- TrashDet: Iterative Neural Architecture Search for Efficient Waste Detection
- Programmable Optical Spectrum Shapers as Computing Primitives for Accelerating Convolutional Neural Networks
- Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- Neural Scaling Laws for Learning-based Identification of Nonlinear Systems
- CLIP Based Region-Aware Feature Fusion for Automated BBPS Scoring in Colonoscopy Images
- FedDPC : Handling Data Heterogeneity and Partial Client Participation in Federated Learning
- UbiQVision: Quantifying Uncertainty in XAI for Image Recognition
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- Progressive Learned Image Compression for Machine Perception
- Self-motion as a structural prior for coherent and robust formation of cognitive maps
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs
- WSD-MIL: Window Scale Decay Multiple Instance Learning for Whole Slide Image Classification
- TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
- Target Classification for Integrated Sensing and Communication in Industrial Deployments
- LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- Vehicle-centric Perception via Multimodal Structured Pre-training
- Non-Contrast CT Esophageal Varices Grading through Clinical Prior-Enhanced Multi-Organ Analysis
- DK-STN: A Domain Knowledge Embedded Spatio-Temporal Network Model for MJO Forecast
- Deep Legendre Transform
- No Data? No Problem: Robust Vision-Tabular Learning with Missing Values
- Event Extraction in Large Language Model
- Multi-Modal Soccer Scene Analysis with Masked Pre-Training
- Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration
- DSTED: Decoupling Temporal Stabilization and Discriminative Enhancement for Surgical Workflow Recognition
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- From Pixels to Predicates Structuring urban perception with scene graphs
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- AMap: Distilling Future Priors for Ahead-Aware Online HD Map Construction
- D2Stream: Decoupled Dual-Stream Temporal-Speaker Interaction for Audio-Visual Speaker Detection
- Timely Parameter Updating in Over-the-Air Federated Learning
- AI-Driven Subcarrier-Level CQI Feedback
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- DeepQuantum: A PyTorch-based Software Platform for Quantum Machine Learning and Photonic Quantum Computing
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- HyGE-Occ: Hybrid View-Transformation with 3D Gaussian and Edge Priors for 3D Panoptic Occupancy Prediction
- VOIC: Visible-Occluded Integrated Guidance for 3D Semantic Scene Completion
- Localising Shortcut Learning in Pixel Space via Ordinal Scoring Correlations for Attribution Representations (OSCAR)
- VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference
- Misbehavior Forecasting for Focused Autonomous Driving Systems Testing
- Decentralized GNSS at Global Scale via Graph-Aware Diffusion Adaptation
- ML Inference Scheduling with Predictable Latency
- Generating Risky Samples with Conformity Constraints via Diffusion Models
- Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
- WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
- NASTaR: NovaSAR Automated Ship Target Recognition Dataset
- PlantDiseaseNet-RT50: A Fine-tuned ResNet50 Architecture for High-Accuracy Plant Disease Detection Beyond Standard CNNs
- MeniMV: A Multi-view Benchmark for Meniscus Injury Severity Grading
- UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
- On the Convergence Rate of LoRA Gradient Descent
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- FedWiLoc: Federated Learning for Privacy-Preserving WiFi Indoor Localization
- SLIM: Semantic-based Low-bitrate Image compression for Machines by leveraging diffusion
- ALIGN: Advanced Query Initialization with LiDAR-Image Guidance for Occlusion-Robust 3D Object Detection
- Detection of AI Generated Images Using Combined Uncertainty Measures and Particle Swarm Optimised Rejection Mechanism
- A two-stream network with global-local feature fusion for bone age assessment
- FedOAED: Federated On-Device Autoencoder Denoiser for Heterogeneous Data under Limited Client Availability
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients
- InfinityEBSD : Metrics-Guided Infinite-Size EBSD Map Generation With Diffusion Models
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image
- Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
- Self-Supervised Weighted Image Guided Quantitative MRI Super-Resolution
- A Unified Representation of Neural Networks Architectures
- Practical Framework for Privacy-Preserving and Byzantine-robust Federated Learning
- Validation of Diagnostic Artificial Intelligence Models for Prostate Pathology in a Middle Eastern Cohort
- A Systematic Reproducibility Study of BSARec for Sequential Recommendation
- Domain-Aware Quantum Circuit for QML
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
- Cryptanalysis of Pseudorandom Error-Correcting Codes
- Towards Pixel-Wise Anomaly Location for High-Resolution PCBA via Self-Supervised Image Reconstruction
- WDFFU-Mamba: A Wavelet-guided Dual-attention Feature Fusion Mamba for Breast Tumor Segmentation in Ultrasound Images
- A Search for Binary Black Hole Mergers in LIGO O1-O3 Data with Convolutional Neural Networks
- Can Synthetic Images Serve as Effective and Efficient Class Prototypes?
- Enhancing AIGC Service Efficiency with Adaptive Multi-Edge Collaboration in A Distributed System
- Interpretable Similarity of Synthetic Image Utility
- Training Together, Diagnosing Better: Federated Learning for Collagen VI-Related Dystrophies
- Sequencing to Mitigate Catastrophic Forgetting in Continual Learning
- OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
- FlowDet: Unifying Object Detection and Generative Transport Flows
- Few-Shot Fingerprinting Subject Re-Identification in 3D-MRI and 2D-X-Ray
- Exploiting Radio Frequency Fingerprints for Device Identification: Tackling Cross-receiver Challenges in the Source-data-free Scenario
- SARMAE: Masked Autoencoder for SAR Representation Learning
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
- Continuized Nesterov Acceleration for Non-Convex Optimization
- XTC, A Research Platform for Optimizing AI Workload Operators
- Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors
- Hypernetworks That Evolve Themselves
- Graph Neural Networks for Source Detection: A Review and Benchmark Study
- Learning High-Quality Initial Noise for Single-View Synthesis with Diffusion Models
- A Multimodal Approach to Alzheimer's Diagnosis: Geometric Insights from Cube Copying and Cognitive Assessments
- Towards Closing the Domain Gap with Event Cameras
- Dual-View Inference Attack: Machine Unlearning Amplifies Privacy Exposure
- GenAI-enabled Residual Motion Estimation for Energy-Efficient Semantic Video Communication
- Evaluation of deep learning architectures for wildlife object detection: A comparative study of ResNet and Inception
- Privacy Blur: Quantifying Privacy and Utility for Image Data Release
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Open Ad-hoc Categorization with Contextualized Feature Learning
- Higher-Order LaSDI: Reduced Order Modeling with Multiple Time Derivatives
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
- SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks
- In Pursuit of Pixel Supervision for Visual Pre-training
- Artism: AI-Driven Dual-Engine System for Art Generation and Critique
- From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
- Reducing Pilots in Channel Estimation with Predictive Foundation Models
- SCS-SupCon: Sigmoid-based Common and Style Supervised Contrastive Learning with Adaptive Decision Boundaries
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
- From Risk to Resilience: Towards Assessing and Mitigating the Risk of Data Reconstruction Attacks in Federated Learning
- Packed Malware Detection Using Grayscale Binary-to-Image Representations
- Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
- SemanticBridge - A Dataset for 3D Semantic Segmentation of Bridges and Domain Gap Analysis
- Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
- Automated Motion Artifact Check for MRI (AutoMAC-MRI): An Interpretable Framework for Motion Artifact Detection and Severity Assessment
- BBNet: accurate neural network emulator for primordial light element abundances
- Enhancing Alzheimer's Detection through Late Fusion of Multi-Modal EEG Features
- Exploring Deep-to-Shallow Transformable Neural Networks for Intelligent Embedded Systems
- An updated efficient galaxy morphology classification model based on ConvNeXt encoding with UMAP dimensionality reduction
- EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting
- LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
- Which Coauthor Should I Nominate in My 99 ICLR Submissions? A Mathematical Analysis of the ICLR 2026 Reciprocal Reviewer Nomination Policy
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
- Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
- Physics-driven human-like working memory outperforms digital networks in dynamic vision
- Borrowing from anything: A generalizable framework for reference-guided instance editing
- BEV-Patch-PF: Particle Filtering with BEV-Aerial Feature Matching for Off-Road Geo-Localization
- GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
- Isolated Sign Language Recognition with Segmentation and Pose Estimation
- A Roadmap for Applying Graph Neural Networks to Numerical Data: Insights from Cementitious Materials
- Enhancing Visual Sentiment Analysis via Semiotic Isotopy-Guided Dataset Construction
- A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
- FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
- Relaying Signal When Monitoring Traffic: Double Use of Aerial Vehicles Towards Intelligent Low-Altitude Networking
- Enhancing Interpretability for Vision Models via Shapley Value Optimization
- Decoding Orbital Angular Momentum in Turbid Tissue-like Scattering Medium via Fourier-Domain Deep Learning
- Attention-Based Foundation Model for Quantum States
- Transfer learning reveals large discrepancies between air and land surface temperatures in cities
- Seeing the Whole Picture: Distribution-Guided Data-Free Distillation for Semantic Segmentation
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- SELECT: Detecting Label Errors in Real-world Scene Text Data
- Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
- A Single Architecture for Representing Invariance Under Any Space Group
- Ensemble-Guided Distillation for Compact and Robust Acoustic Scene Classification on Edge Devices
- How Does Fourier Analysis Network Work? A Mechanism Analysis and a New Dual-Activation Layer Proposal
- SignIT: A Comprehensive Dataset and Multimodal Analysis for Italian Sign Language Recognition
- Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders
- LitePT: Lighter Yet Stronger Point Transformer
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer for High-Sparsity and Low-Precision Networks
- DP-CSGP: Differentially Private Stochastic Gradient Push with Compressed Communication
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- The Renaissance of Expert Systems: Optical Recognition of Printed Chinese Jianpu Musical Scores with Lyrics
- Face Identity Unlearning for Retrieval via Embedding Dispersion
- Dual-Qubit Hierarchical Fuzzy Neural Network for Image Classification: Enabling Relational Learning via Quantum Entanglement
- WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory
- LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- Time-aware UNet and super-resolution deep residual networks for spatial downscaling
- Reducing Label Dependency in Human Activity Recognition with Wearables: From Supervised Learning to Novel Weakly Self-Supervised Approaches
- TWLR: Text-Guided Weakly-Supervised Lesion Localization and Severity Regression for Explainable Diabetic Retinopathy Grading
- Leveraging Compression to Construct Transferable Bitrate Ladders
- Sharpness-aware Dynamic Anchor Selection for Generalized Category Discovery
- Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- Robust Motion Generation using Part-level Reliable Data from Videos
- Selective Conformal Risk Control
- PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
- GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients
- Federated Learning with Feedback Alignment
- FuXi-γ: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mechanism
- Communication-Efficient Neural Tangent Kernels for Heterogeneous Decentralized Federated Learning
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- TF-MCL: Time-frequency Fusion and Multi-domain Cross-Loss for Self-supervised Depression Detection
- StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain Analysis
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- Generative Spatiotemporal Data Augmentation
- DL3M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- Bench-Push: Benchmarking Pushing-based Navigation and Manipulation Tasks for Mobile Robots
- Optimized Architectures for Kolmogorov-Arnold Networks
- Hybrid algorithm combining matched filtering and convolutional neural networks for searching gravitational waves from binary black hole mergers
- Uncertainty Quantification for Machine Learning: One Size Does Not Fit All
- Active learning potentials for first-principles phase diagrams using replica-exchange nested sampling
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Bridging Data and Physics: A Graph Neural Network-Based Hybrid Twin Framework
- A Hybrid Deep Learning Framework for Emotion Recognition in Children with Autism During NAO Robot-Mediated Interaction
- ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB
- Open Horizons: Evaluating Deep Models in the Wild
- A Benchmark Dataset for Spatially Aligned Road Damage Assessment in Small Uncrewed Aerial Systems Disaster Imagery
- SPDMark: Selective Parameter Displacement for Robust Video Watermarking
- Balancing Accuracy and Speed: A Multi-Fidelity Ensemble Kalman Filter with a Machine Learning Surrogate Model
- Non-Convex Federated Optimization under Cost-Aware Client Selection
- Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
- Fully Inductive Node Representation Learning via Graph View Transformation
- ACCOR: Attention-Enhanced Complex-Valued Contrastive Learning for Occluded Object Classification Using mmWave Radar IQ Signals
- RadarFuseNet: Complex-Valued Attention-Based Fusion of IQ Time- and Frequency-Domain Radar Features for Classification Tasks
- Super-Resolved Canopy Height Mapping from Sentinel-2 Time Series Using LiDAR HD Reference Data across Metropolitan France
- DREAM-B3P: Dual-Stream Transformer Network Enhanced by Feedback Diffusion Model for Blood-Brain Barrier Penetrating Peptide Prediction
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- ARC-AGI Without Pretraining
- Type II and Type III Solar Radio Burst Classification Using Transfer Learning
- On the Bayes Inconsistency of Disagreement Discrepancy Surrogates
- Vision-Based Learning for Cyberattack Detection in Blockchain Smart Contracts and Transactions
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Back to the Baseline: Examining Baseline Effects on Explainability Metrics
- Weak-to-Strong Generalization Enables Fully Automated De Novo Training of Multi-head Mask-RCNN Model for Segmenting Densely Overlapping Cell Nuclei in Multiplex Whole-slice Brain Images
- Bhargava Cube--Inspired Quadratic Regularization for Structured Neural Embeddings
- SATMapTR: Satellite Image Enhanced Online HD Map Construction
- Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
- Task-Specific Distance Correlation Matching for Few-Shot Action Recognition
- CogniSNN: Enabling Neuron-Expandability, Pathway-Reusability, and Dynamic-Configurability with Random Graph Architectures in Spiking Neural Networks
- Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Reliable Detection of Minute Targets in High-Resolution Aerial Imagery across Temporal Shifts
- AMBER: An Adaptive Multimodal Mask Transformer for Beam Prediction with Missing Modalities
- Weakly Supervised Tuberculosis Localization in Chest X-rays through Knowledge Distillation
- Bidirectional Normalizing Flow: From Data to Noise and Back
- PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
- Graph Laplacian Transformer with Progressive Sampling for Prostate Cancer Grading
- Physics-Informed Learning of Microvascular Flow Models using Graph Neural Networks
- Blood Pressure Prediction for Coronary Artery Disease Diagnosis using Coronary Computed Tomography Angiography
- Neural personal sound zones with flexible bright zone control
- Learning to Split: A Reinforcement-Learning-Guided Splitting Heuristic for Neural Network Verification
- Thermal and Size Effects in Ferroelastic Domains by Machine Learning
- Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
- Robust Shape from Focus via Multiscale Directional Dilated Laplacian and Recurrent Network
- Clustered Federated Learning with Hierarchical Knowledge Distillation
- Neural Collapse in Test-Time Adaptation
- Infusing Experimental Reality into Complex Many-Body Hamiltonians: The Observable-Constrained Variational Framework (OCVF)
- Independent Density Estimation
- Sample-wise Adaptive Weighting for Transfer Consistency in Adversarial Distillation
- Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
- Federated Domain Generalization with Latent Space Inversion
- Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- Self-Ensemble Post Learning for Noisy Domain Generalization
- Optimal transport unlocks end-to-end learning for single-molecule localization
- StainNet: Scaling Self-Supervised Foundation Models on Immunohistochemistry and Special Stains for Computational Pathology
- Sequence-to-Image Transformation for Sequence Classification Using Rips Complex Construction and Chaos Game Representation
- Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
- Towards Visual Re-Identification of Fish using Fine-Grained Classification for Electronic Monitoring in Fisheries
- Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
- Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
- CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Masked Registration and Autoencoding of CT Images for Predictive Tibia Reconstruction
- NeuroSketch: An Effective Framework for Neural Decoding via Systematic Architectural Optimization
- QuanvNeXt: An end-to-end quanvolutional neural network for EEG-based detection of major depressive disorder
- Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration
- Benchmarking Real-World Medical Image Classification with Noisy Labels: Challenges, Practice, and Outlook
- Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance
- SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- Contrastive Learning for Semi-Supervised Deep Regression with Generalized Ordinal Rankings from Spectral Seriation
- Accelerating high-order energy-stable discontinous Galerkin solver using auto-differentiation and neural networks
- Natural Geometry of Robust Data Attribution: From Convex Models to Deep Networks
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- Slow dynamics and magnon bound states in the 2D long-range quantum Ising model
- Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Mitigating Individual Skin Tone Bias in Skin Lesion Classification through Distribution-Aware Reweighting
- Fully Decentralized Certified Unlearning
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- When unlearning is free: leveraging low influence points to reduce computational costs
- Bi2MAC: Bimodal Bi-Adaptive Mask-Aware Convolution for Remote Sensing Pansharpening
- RLCNet: An end-to-end deep learning framework for simultaneous online calibration of LiDAR, RADAR, and Camera
- Fast-BEV++: Fast by Algorithm, Deployable by Design
- Spatially-extended Flow Phixer (SpeF-Phixer): A Spatially Extended φ-Mixing Framework for Gene Regulatory Causal Inference in Spatial Gene Field
- Improving the Sensitivity of Backdoor Detectors via Class Subspace Orthogonalization
- Instance-Aware Test-Time Segmentation for Continual Domain Shifts
- Identification of Deforestation Areas in the Amazon Rainforest Using Change Detection Models
- Liver Fibrosis Quantification and Analysis: The LiQA Dataset and Baseline Method
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- CIP-Net: Continual Interpretable Prototype-based Network
- VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
- Relational Visual Similarity
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- Improving action classification with brain-inspired deep networks
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- Bandwidth-Aware Network Topology Optimization for Decentralized Learning
- Dictionary-Based Contrastive Learning for GNSS Jamming Detection
- Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
- How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline
- LogicCBMs: Logic-Enhanced Concept-Based Learning
- A Geometric Unification of Concept Learning with Concept Cones
- DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- START: Spatial and Textual Learning for Chart Understanding
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- UnCageNet: Tracking and Pose Estimation of Caged Animal
- Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification
- R2MF-Net: A Recurrent Residual Multi-Path Fusion Network for Robust Multi-directional Spine X-ray Segmentation
- Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
- Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology
- Balanced Learning for Domain Adaptive Semantic Segmentation
- SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
- Leveraging Pre-trained Neural Network Models for the Classification of Tumor Cells Analyzed by Label-free Phase Holotomographic Microscopy
- Optimal and Diffusion Transports in Machine Learning
- ADAM Optimization with Adaptive Batch Selection
- JOCA: Task-Driven Joint Optimisation of Camera Hardware and Adaptive Camera Control Algorithms
- Rectifying Latent Space for Generative Single-Image Reflection Removal
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- Hierarchical Deep Learning for Diatom Image Classification: A Multi-Level Taxonomic Approach
- On Memory: A comparison of memory mechanisms in world models
- Embodied Referring Expression Comprehension in Human-Robot Interaction
- Programmable and GPU-Accelerated Edge Inference for Real-Time ISAC on NVIDIA Aerial Testbed
- Neural expressiveness for beyond importance model compression
- Automated Deep Learning Estimation of Anthropometric Measurements for Preparticipation Cardiovascular Screening
- When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
- 3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
- A Perception CNN for Facial Expression Recognition
- Phase-OTDR Event Detection Using Image-Based Data Transformation and Deep Learning
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Entropic Confinement and Mode Connectivity in Overparameterized Neural Networks
- Language-driven Fine-grained Retrieval
- Approximate Multiplier Induced Error Propagation in Deep Neural Networks
- Greedy Alignment Principle for Optimizer Selection
- SPOOF: Simple Pixel Operations for Out-of-Distribution Fooling
- DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification
- AQUA-Net: Adaptive Frequency Fusion and Illumination Aware Network for Underwater Image Enhancement
- Measuring the Effect of Background on Classification and Feature Importance in Deep Learning for AV Perception
- Synset Signset Germany: a Synthetic Dataset for German Traffic Sign Recognition
- A Comparative Study on Synthetic Facial Data Generation Techniques for Face Recognition
- Variance Matters: Improving Domain Adaptation via Stratified Sampling
- LDLT L-Lipschitz Network: Generalized Deep End-To-End Lipschitz Network Construction
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- HSCP: A Two-Stage Spectral Clustering Framework for Resource-Constrained UAV Identification
- Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image Segmentation
- Curvature-Regularized Variational Autoencoder for 3D Scene Reconstruction from Sparse Depth
- Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer
- Wasserstein distance based semi-supervised manifold learning and application to GNSS multi-path detection
- VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
- DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
- Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm
- Unleashing Temporal Capacity of Spiking Neural Networks through Spatiotemporal Separation
- University Building Recognition Dataset in Thailand for the mission-oriented IoT sensor system
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- RevoNAD: Reflective Evolutionary Exploration for Neural Architecture Design
- Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
- Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural Networks
- Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting Generation
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Recurrent Neural Networks with Linear Structures for Electricity Price Forecasting
- Rethinking Decoupled Knowledge Distillation: A Predictive Distribution Perspective
- Infrared UAV Target Tracking with Dynamic Feature Refinement and Global Contextual Attention Knowledge Distillation
- Research on Brain Tumor Classification Method Based on Improved ResNet34 Network
- MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
- FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
- Searching for binary black hole mergers with deep learning in Advanced LIGO's third observing run
- Open Set Face Forgery Detection via Dual-Level Evidence Collection
- CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
- Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
- Predicting parameters of a model cuprate superconductor using machine learning
- State Space Models for Bioacoustics: A comparative Evaluation with Transformers
- Parameter efficient hybrid spiking-quantum convolutional neural network with surrogate gradient and quantum data-reupload
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Comparison of neural network training strategies for the simulation of dynamical systems
- HieroGlyphTranslator: Automatic Recognition and Translation of Egyptian Hieroglyphs to English
- EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Studying Various Activation Functions and Non-IID Data for Machine Learning Model Robustness
- Multi-Scale Visual Prompting for Lightweight Small-Image Classification
- MKSNet: Advanced Small Object Detection in Remote Sensing Imagery with Multi-Kernel and Dual Attention Mechanisms
- FeatureLens: A Highly Generalizable and Interpretable Framework for Detecting Adversarial Examples Based on Image Features
- SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
- HBFormer: A Hybrid-Bridge Transformer for Microtumor and Miniature Organ Segmentation
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching
- Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
- NAS-LoRA: Empowering Parameter-Efficient Fine-Tuning for Visual Foundation Models with Searchable Adaptation
- Hierarchical Attention for Sparse Volumetric Anomaly Detection in Subclinical Keratoconus
- Network of Theseus (like the ship)
- Single-Round Scalable Analytic Federated Learning
- DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- PyroFocus: A Deep Learning Approach to Real-Time Wildfire Detection in Multispectral Remote Sensing Imagery
- Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems
- Decision Tree Embedding by Leaf-Means
- Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
- Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks
- GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection
- ProteinPNet: Prototypical Part Networks for Concept Learning in Spatial Proteomics
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
- FastFHE: Packing-Scalable and Depthwise-Separable CNN Inference Over FHE
- Layout Anything: One Transformer for Universal Room Layout Estimation
- Defense That Attacks: How Robust Models Become Better Attackers
- A Framework for Causal Concept-based Model Explanations
- Conformal Correction for Efficiency May be at Odds with Entropy
- A Discrete Neural Operator with Adaptive Sampling for Surrogate Modeling of Parametric Transient Darcy Flows in Porous Media
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
- Leveraging generative adversarial networks with spatially adaptive denormalization for multivariate stochastic seismic data inversion
- Drainage: A Unifying Framework for Addressing Class Uncertainty
- TGDD: Trajectory Guided Dataset Distillation with Balanced Distribution
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- Equilibrium Propagation Without Limits
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Data-Centric Visual Development for Self-Driving Labs
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Physical ID-Transfer Attacks against Multi-Object Tracking via Adversarial Trajectory
- SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- Learned Image Compression for Earth Observation: Implications for Downstream Segmentation Tasks
- Dual Randomized Smoothing: Beyond Global Noise Variance
- On the Unreasonable Effectiveness of Last-layer Retraining
- Multimodal Mixture-of-Experts for ISAC in Low-Altitude Wireless Networks
- A unified framework for geometry-independent operator learning in cardiac electrophysiology simulations
- DB-KAUNet: An Adaptive Dual Branch Kolmogorov-Arnold UNet for Retinal Vessel Segmentation
- ViT3: Unlocking Test-Time Training in Vision
- Toward Content-based Indexing and Retrieval of Head and Neck CT with Abscess Segmentation
- Neural Networks for Predicting Permeability Tensors of 2D Porous Media: Comparison of Convolution- and Transformer-based Architectures
- Walking on the Fiber: A Simple Geometric Approximation for Bayesian Neural Networks
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- Directed evolution algorithm drives neural prediction
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- DPAC: Distribution-Preserving Adversarial Control for Diffusion Sampling
- Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis
- Generalized Medical Phrase Grounding
- PIANO: Physics-informed Dual Neural Operator for Precipitation Nowcasting
- Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- PAGen: Phase-guided Amplitude Generation for Domain-adaptive Object Detection
- Neuroscience-Inspired Memory Replay for Continual Learning: A Comparative Study of Predictive Coding and Backpropagation-Based Strategies
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- Partially Shared Concept Bottleneck Models
- 3D Affordance Keypoint Detection for Robotic Manipulation
- ForamDeepSlice: A High-Accuracy Deep Learning Framework for Foraminifera Species Classification from 2D Micro-CT Slices
- Integrating Skeleton Based Representations for Robust Yoga Pose Classification Using Deep Learning Models
- Structured Context Learning for Generic Event Boundary Detection
- GreenPlanner: Practical Floorplan Layout Generation via an Energy-Aware and Function-Feasible Generative Framework
- SelfAI: Building a Self-Training AI System with LLM Agents
- Time-Series at the Edge: Tiny Separable CNNs for Wearable Gait Detection and Optimal Sensor Placement
- Vision Transformer for Classification of UAV and Helicopters Using Micro-Doppler Spectrograms in Surveillance Radar
- S2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
- An Interpretable Operator-Learning Model for Electric Field Profile Reconstruction in Discharges Based on the EFISH Method
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
- "Why the face?": Exploring Robot Error Detection Using Instrumented Bystander Reactions
- Can we cover navigational perception needs of the visually impaired by panoptic segmentation?
- MANTA: Physics-Informed Generalized Underwater Object Tracking
- SimScale: Learning to Drive via Real-World Simulation at Scale
- PointCNN++: Performant Convolution on Native Points
- GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN acceleration
- Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
- Estimating the Event-Related Potential from Few EEG Trials
- Evaluating the Clinical Impact of Generative Inpainting on Bone Age Estimation
- Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Local and Global Context-and-Object-part-Aware Superpixel-based Data Augmentation for Deep Visual Recognition
- Stable-Drift: A Patient-Aware Latent Drift Replay Method for Stabilizing Representations in Continual Learning
- ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery
- CNN-Based Framework for Pedestrian Age and Gender Classification Using Far-View Surveillance in Mixed-Traffic Intersections
- Adaptive Dataset Quantization: A New Direction for Dataset Pruning
- SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
- A Unified and Stable Risk Minimization Framework for Weakly Supervised Learning with Theoretical Guarantees
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- InstanceV: Instance-Level Video Generation
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- Leveraging Textual Compositional Reasoning for Robust Change Captioning
- Time Series Forecasting via Direct Per-Step Probability Distribution Modeling
- Closing the Generalization Gap in Parameter-efficient Federated Edge Learning
- U Net LSTM with incremental time-stepping for robust long-horizon unsteady flow prediction
- GazeTrack: High-Precision Eye Tracking Based on Regularization and Spatial Computing
- Hard Spatial Gating for Precision-Driven Brain Metastasis Segmentation: Addressing the Over-Segmentation Paradox in Deep Attention Networks
- AnoRefiner: Anomaly-Aware Group-Wise Refinement for Zero-Shot Industrial Anomaly Detection
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
- IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame
- A Spatially Masked Adaptive Gated Network for multimodal post-flood water extent mapping using SAR and incomplete multispectral data
- Rethinking Cross-Generator Image Forgery Detection through DINOv3
- Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation
- DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
- Semi-Supervised Contrastive Learning with Orthonormal Prototypes
- AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
- HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition
- FedRE: A Representation Entanglement Framework for Model-Heterogeneous Federated Learning
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
- Stacked Ensemble of Fine-Tuned CNNs for Knee Osteoarthritis Severity Grading
- HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction
- MRI-Based Brain Age Estimation with Supervised Contrastive Learning of Continuous Representation
- Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving
- The Age-specific Alzheimer 's Disease Prediction with Characteristic Constraints in Nonuniform Time Span
- Self-Paced Learning for Images of Antinuclear Antibodies
- SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
- Biomimetic Metamaterial-based Interface for Decoding Heterogeneous Mechanodermal Activity
- Uni-Hema: Unified Model for Digital Hematopathology
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Lost in Time? A Meta-Learning Framework for Time-Shift-Tolerant Physiological Signal Transformation
- Dynamical Implicit Neural Representations
- PFF-Net: Patch Feature Fitting for Point Cloud Normal Estimation
- Shift-Equivariant Complex-Valued Convolutional Neural Networks
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- DeepRFTv2: Kernel-level Learning for Image Deblurring
- RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression
- Probabilistic Wildfire Spread Prediction Using an Autoregressive Conditional Generative Adversarial Network
- MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
- RefOnce: Distilling References into a Prototype Memory for Referring Camouflaged Object Detection
- Beyond Realism: Learning the Art of Expressive Composition with StickerNet
- AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
- Guaranteed Optimal Compositional Explanations for Neurons
- Open Vocabulary Compositional Explanations for Neuron Alignment
- Pre-train to Gain: Robust Learning Without Clean Labels
- Effects of Initialization Biases on Deep Neural Network Training Dynamics
- Intriguing Properties of Dynamic Sampling Networks
- Adversarial Multi-Task Learning for Liver Tumor Segmentation, Dynamic Enhancement Regression, and Classification
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
- The Driver-Blindness Phenomenon: Why Deep Sequence Models Default to Autocorrelation in Blood Glucose Forecasting
- Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection
- Automated Monitoring of Cultural Heritage Artifacts Using Semantic Segmentation
- Adam Simplified: Bias Correction Debunked
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- HVAdam: A Full-Dimension Adaptive Optimizer
- DRL-Guided Neural Batch Sampling for Semi-Supervised Pixel-Level Anomaly Detection
- Advancing Image Classification with Discrete Diffusion Classification Modeling
- Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization
- XiCAD: Camera Activation Detection in the Da Vinci Xi User Interface
- DiCaP: Distribution-Calibrated Pseudo-labeling for Semi-Supervised Multi-Label Learning
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- Space Alignment Matters: The Missing Piece for Inducing Neural Collapse in Long-Tailed Learning
- RankOOD -- Class Ranking-based Out-of-Distribution Detection
- GazeProphetV2: Head-Movement-Based Gaze Prediction Enabling Efficient Foveated Rendering on Mobile VR
- Stragglers Can Contribute More: Uncertainty-Aware Distillation for Asynchronous Federated Learning
- MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
- Low-Resolution Editing is All You Need for High-Resolution Editing
- Intelligent Image Search Algorithms Fusing Visual Large Models
- LiMT: A Multi-task Liver Image Benchmark Dataset
- Distilling Cross-Modal Knowledge via Feature Disentanglement
- Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
- Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- Annotation-Free Class-Incremental Learning
- AdaCap: An Adaptive Contrastive Approach for Small-Data Neural Networks
- HybriDLA: Hybrid Generation for Document Layout Analysis
- On the Limits of Momentum in Decentralized and Federated Optimization
- Rethinking Message Passing Neural Networks with Diffusion Distance-guided Stress Majorization
- Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
- An Anatomy Aware Hybrid Deep Learning Framework for Lung Cancer Tumor Stage Classification
- Enhancing Conformal Prediction via Class Similarity
- CellFMCount: A Fluorescence Microscopy Dataset, Benchmark, and Methods for Cell Counting
- A CNN-Based Technique to Assist Layout-to-Generator Conversion for Analog Circuits
- DynaMix: Generalizable Person Re-identification via Dynamic Relabeling and Mixed Data Sampling
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- Development of a fully deep learning model to improve the reproducibility of sector classification systems for predicting unerupted maxillary canine likelihood of impaction
- EEG-VLM: A Hierarchical Vision-Language Model with Multi-Level Feature Alignment and Visually Enhanced Language-Guided Reasoning for EEG Image-Based Sleep Stage Prediction
- Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
- When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP
- DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
- Graph-based 3D Human Pose Estimation using WiFi Signals
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
- MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design
- Cross-Domain Generalization of Multimodal LLMs for Global Photovoltaic Assessment
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Leveraging Adversarial Learning for Pathological Fidelity in Virtual Staining
- Scalable Vision-Guided Crop Yield Estimation
- Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
- Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
- Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
- MapRF: Weakly Supervised Online HD Map Construction via NeRF-Guided Self-Training
- AutoMAS: A Generic Multi-Agent System for Algorithm Self-Adaptation in Wireless Networks
- Sampling Control for Imbalanced Calibration in Semi-Supervised Learning
- Re-Key-Free, Risky-Free: Adaptable Model Usage Control
- NI-Tex: Non-isometric Image-based Garment Texture Generation
- OceanForecastBench: A Benchmark Dataset for Data-Driven Global Ocean Forecasting
- Higgs Production Classifier using Weak Supervision
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Single Image to High-Quality 3D Object via Latent Features
- Exploring Surround-View Fisheye Camera 3D Object Detection
- Neural Geometry Image-Based Representations with Optimal Transport (OT)
- Subtract the Corruption: Training-Data-Free Corrective Machine Unlearning using Task Arithmetic
- LMSeg: an end-to-end geometric message-passing network on barycentric dualgraphs for large-scale landscape mesh segmentation
- Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
- Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
- LAA3D: A Benchmark of Detecting and Tracking Low-Altitude Aircraft in 3D Space
- Towards Characterizing Knowledge Distillation of PPG Heart Rate Estimation Models
- CoD: A Diffusion Foundation Model for Image Compression
- Dendritic Convolution for Noise Image Recognition
- AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
- Bayesian-based Online Label Shift Estimation with Dynamic Dirichlet Priors
- CycleSL: Server-Client Cyclical Update Driven Scalable Split Learning
- From Simulations to Surveys: Domain Adaptation for Galaxy Observations
- PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation
- Breaking Forgetting: Training-Free Few-Shot Class-Incremental Learning via Conditional Diffusion
- Shape-Adapting Gated Experts: Dynamic Expert Routing for Colonoscopic Lesion Segmentation
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- AIA-UltraNeRF:Acoustic-Impedance-Aware Neural Radiance Field with Hash Encodings for Robotic Ultrasound Reconstruction and Localization
- Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
- DE-KAN: A Kolmogorov Arnold Network with Dual Encoder for accurate 2D Teeth Segmentation
- UNSEEN: Enhancing Dataset Pruning from a Generalization Perspective
- Hyperspectral Variational Autoencoders for Joint Data Compression and Component Extraction
- Auxiliary Gene Learning: Spatial Gene Expression Estimation by Auxiliary Gene Selection
- Stro-VIGRU: Defining the Vision Recurrent-Based Baseline Model for Brain Stroke Classification
- AFT: Appearance-Based Feature Tracking for Markerless and Training-Free Shape Reconstruction of Soft Robots
- Early Lung Cancer Diagnosis from Virtual Follow-up LDCT Generation via Correlational Autoencoder and Latent Flow Matching
- Nested Unfolding Network for Real-World Concealed Object Segmentation
- Compact neural networks for astronomy with optimal transport bias correction
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
- Learning Rate Scheduling with Matrix Factorization for Private Training
- Adversarial Pseudo-replay for Exemplar-free Class-incremental Learning
- Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
- Understanding Private Learning From Feature Perspective
- Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset Distillation
- L1 Sample Flow for Efficient Visuomotor Learning
- Cost-Sensitive Conformal Training with Provably Controllable Learning Bounds
- State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection
- A multi-view contrastive learning framework for spatial embeddings in risk modelling
- Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
- REXO: Indoor Multi-View Radar Object Detection via 3D Bounding Box Diffusion
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- An Artificial Intelligence Framework for Measuring Human Spine Aging Using MRI
- Radar2Shape: 3D Shape Reconstruction from High-Frequency Radar using Multiresolution Signed Distance Functions
- Self-Supervised Learning by Curvature Alignment
- MemIntelli: A Generic End-to-End Simulation Framework for Memristive Intelligent Computing
- Quantum Masked Autoencoders for Vision Learning
- SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering
- ReBaPL: Repulsive Bayesian Prompt Learning
- QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
- Dual-Path Knowledge-Augmented Contrastive Alignment Network for Spatially Resolved Transcriptomics
- Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
- RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
- Step-E: A Differentiable Data Cleaning Framework for Robust Learning with Noisy Labels
- DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
- A Diversity-optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
- Towards a Safer and Sustainable Manufacturing Process: Material classification in Laser Cutting Using Deep Learning
- Enhancing Adversarial Transferability through Block Stretch and Shrink
- Person Recognition in Aerial Surveillance: A Decade Survey
- MRI Super-Resolution with Deep Learning: A Comprehensive Survey
- Membership Inference Attacks Beyond Overfitting
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- Formal Abductive Latent Explanations for Prototype-Based Networks
- Broad stochastic configuration residual learning system for norm-convergent universal approximation
- Dynamic Participation in Federated Learning: Benchmarks and a Knowledge Pool Plugin
- Hard Samples, Bad Labels: Robust Loss Functions That Know When to Back Off
- ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
- Flow and Depth Assisted Video Prediction with Latent Transformer
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models
- MF-GCN: A Multi-Frequency Graph Convolutional Network for Tri-Modal Depression Detection Using Eye-Tracking, Facial, and Acoustic Features
- Memory-DD: A Low-Complexity Dendrite-Inspired Neuron for Temporal Prediction Tasks
- Green Distributed AI Training: Orchestrating Compute Across Renewable-Powered Micro Datacenters
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Descend or Rewind? Stochastic Gradient Descent Unlearning
- LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
- Exploiting Inter-Sample Information for Long-tailed Out-of-Distribution Detection
- Scenario-Aware Control of Segmented Ladder Bus: Design and FPGA Implementation
- GUIDE: Gaussian Unified Instance Detection for Enhanced Obstacle Perception in Autonomous Driving
- SpectralTrain: A Universal Framework for Hyperspectral Image Classification
- CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- Real-Time 3D Object Detection with Inference-Aligned Learning
- A Spatial Semantics and Continuity Perception Attention for Remote Sensing Water Body Change Detection
- DiffuDepGrasp: Diffusion-based Depth Noise Modeling Empowers Sim2Real Robotic Grasping
- MambaIO: Global-Coordinate Inertial Odometry for Pedestrians via Multi-Scale Frequency-Decoupled Modeling
- Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
- A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
- FunnyNodules: A Customizable Medical Dataset Tailored for Evaluating Explainable AI
- SIGMMA: Hierarchical Graph-Based Multi-Scale Multi-modal Contrastive Alignment of Histopathology Image and Spatial Transcriptome
- A Dataset and Baseline for Deep Learning-Based Visual Quality Inspection in Remanufacturing
- D4C: Data-Free Quantization for Contrastive Language-Image Pre-training Models
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs
- Robust Bayesian Optimisation with Unbounded Corruptions
- Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
- SNAP: Low-Latency Test-Time Adaptation with Sparse Updates
- Enforcing hidden physics in physics-informed neural networks
- Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- TSRE: Channel-Aware Typical Set Refinement for Out-of-Distribution Detection
- Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- Computer Vision Modeling of the Development of Geometric and Numerical Concepts in Humans
- WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines
- Non-Convex Self-Concordant Functions: Practical Algorithms and Complexity Analysis
- Latent space analysis and generalization to out-of-distribution data
- Toward Complete Merger Identification at Cosmic Noon with Deep Learning
- Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
- Efficient Large-Scale Learning of Minimax Risk Classifiers
- Artificial intelligence approaches for energy-efficient laser cutting machines
- GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
- B-Rep Distance Functions (BR-DF): How to Represent a B-Rep Model by Volumetric Distance Functions?
- M2OE2-GL: A Family of Probabilistic Load Forecasters That Scales to Massive Customers
- SLAM-AGS: Slide-Label Aware Multi-Task Pretraining Using Adaptive Gradient Surgery in Computational Cytology
- Parameter Aware Mamba Model for Multi-task Dense Prediction
- Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
- Learning Subglacial Bed Topography from Sparse Radar with Physics-Guided Residuals
- Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
- Sigil: Server-Enforced Watermarking in U-Shaped Split Federated Learning via Gradient Injection
- Cranio-ID: Graph-Based Craniofacial Identification via Automatic Landmark Annotation in 2D Multi-View X-rays
- Stage Aware Diagnosis of Diabetic Retinopathy via Ordinal Regression
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
- HSMix: Hard and Soft Mixing Data Augmentation for Medical Image Segmentation
- Weight Variance Amplifier Improves Accuracy in High-Sparsity One-Shot Pruning
- Free Lunch to Meet the Gap: Intermediate Domain Reconstruction for Cross-Domain Few-Shot Learning
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- Compute-in-Memory Implementation of State Space Models for Event Sequence Processing
- SAE-MCVT: A Real-Time and Scalable Multi-Camera Vehicle Tracking Framework Powered by Edge Computing
- Training-free Detection of AI-generated images via Cropping Robustness
- Robust Client-Server Watermarking for Split Federated Learning
- Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning
- Alpha Divergence Losses for Biometric Verification
- Power Homotopy for Zeroth-Order Non-Convex Optimizations
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Minimax Multi-Target Conformal Prediction with Applications to Imaging Inverse Problems
- Mapping the Vanishing and Transformation of Urban Villages in China
- Hardware optimization on Android for inference of AI models
- TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
- InfoDecom: Decomposing Information for Defending against Privacy Leakage in Split Inference
- Enhancing All-to-X Backdoor Attacks with Optimized Target Class Mapping
- Semi-Supervised Multi-Task Learning for Interpretable Quality As- sessment of Fundus Images
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- Hybrid-Domain Adaptative Representation Learning for Gaze Estimation
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
- Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection
- Birth of a Painting: Differentiable Brushstroke Reconstruction
- Uncertainty-aware Physics-informed Neural Networks for Robust CARS-to-Raman Signal Reconstruction
- Region-Point Joint Representation for Effective Trajectory Similarity Learning
- Synthetic Forgetting without Access: A Few-shot Zero-glance Framework for Machine Unlearning
- MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
- ResAlignNet: A Data-Driven Approach for INS/DVL Alignment
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for Explanations
- Decoupling Scene Perception and Ego Status: A Multi-Context Fusion Approach for Enhanced Generalization in End-to-End Autonomous Driving
- Monocular 3D Lane Detection via Structure Uncertainty-Aware Network with Curve-Point Queries
- H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
- Simple Lines, Big Ideas: Towards Interpretable Assessment of Human Creativity from Drawings
- Learned Adaptive Kernels for High-Fidelity Image Downscaling
- Federated Cyber Defense: Privacy-Preserving Ransomware Detection Across Distributed Systems
- HIT-ROCKET: Hadamard-vector Inner-product Transformer for ROCKET
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
- INC: An Indirect Neural Corrector for Auto-Regressive Hybrid PDE Solvers
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- GeoPl@ntNet: A Platform for Exploring Essential Biodiversity Variables
- Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning
- Efficiently Training A Flat Neural Network Before It has been Quantizated
- Bridging Granularity Gaps: Hierarchical Semantic Learning for Cross-domain Few-shot Segmentation
- Attention-Enhanced Convolutional Autoencoder and Structured Delay Embeddings for Weather Prediction
- FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning
- OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
- Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
- CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- Towards Temporal Fusion Beyond the Field of View for Camera-based Semantic Scene Completion
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection
- Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
- MFI-ResNet: Efficient ResNet Architecture Optimization via MeanFlow Compression and Selective Incubation
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
- HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
- Segmented Exponent Alignment and Dynamic Wordline Activation for Floating-Point Analog CIM Macros
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Leveraging Quantum-Based Architectures for Robust Diagnostics
- AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
- Model Inversion Attack Against Deep Hashing
- Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash Function
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Dynamic Parameter Optimization for Highly Transferable Transformation-Based Attacks
- Sparse by Rule: Probability-Based N:M Pruning for Spiking Neural Networks
- Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
- Supervised Multilabel Image Classification Using Residual Networks with Probabilistic Reasoning
- Intelligent Collaborative Optimization for Rubber Tyre Film Production Based on Multi-path Differentiated Clipping Proximal Policy Optimization
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
- VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation
- Enhancing Road Safety Through Multi-Camera Image Segmentation with Post-Encroachment Time Analysis
- From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
- Evaluation of Attention Mechanisms in U-Net Architectures for Semantic Segmentation of Brazilian Rock Art Petroglyphs
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- A Deep Learning Framework for Thyroid Nodule Segmentation and Malignancy Classification from Ultrasound Images
- CVChess: A Deep Learning Framework for Converting Chessboard Images to Forsyth-Edwards Notation
- Quantifying and Improving Adaptivity in Conformal Prediction through Input Transformations
- Enhancing Photon Identification with Neural Network Methods
- Robust inverse material design with physical guarantees using the Voigt-Reuss Net
- Coordinate Descent for Network Linearization
- D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces
- Virtual Width Networks
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Machine-Learning Based Detection of Coronary Artery Calcification Using Synthetic Chest X-Rays
- MPCGNet: A Multiscale Feature Extraction and Progressive Feature Aggregation Network Using Coupling Gates for Polyp Segmentation
- Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Heterogeneous Complementary Distillation
- YOLO-Drone: An Efficient Object Detection Approach Using the GhostHead Network for Drone Images
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- Dynamic Temperature Scheduler for Knowledge Distillation
- When to Stop Federated Learning: Zero-Shot Generation of Synthetic Validation Data with Generative AI for Early Stopping
- Dynamic LRP-Based Pruning for CNNs in Data-Scarce Transfer Learning: Suppressing Cascading Accuracy Degradation
- GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- Moirai 2.0: When Less Is More for Time Series Forecasting
- On the Detectability of Active Gradient Inversion Attacks in Federated Learning
- Intrinsic Dimensionality as a Model-Free Measure of Class Imbalance
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
- Domain Adaptation for Camera-Specific Image Characteristics using Shallow Discriminators
- RoboBenchMart: Benchmarking Robots in Retail Environment
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
- Soiling detection for Advanced Driver Assistance Systems
- Fairness-Aware Deepfake Detection: Leveraging Dual-Mechanism Optimization
- Physics-informed Machine Learning for Static Friction Modeling in Robotic Manipulators Based on Kolmogorov-Arnold Networks
- 2.5D Transformer: An Efficient 3D Seismic Interpolation Method without Full 3D Training
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- ConSurv: Multimodal Continual Learning for Survival Analysis
- CertMask: Certifiable Defense Against Adversarial Patches via Theoretically Optimal Mask Coverage
- SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data
- Revisiting the Evaluation of Deep Neural Networks for Pedestrian Detection
- Continuum Dropout for Neural Differential Equations
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models
- Data Heterogeneity and Forgotten Labels in Split Federated Learning
- Generalizing PDE Emulation with Equation-Aware Neural Operators
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
- CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
- Federated Learning for Pediatric Pneumonia Detection: Enabling Collaborative Diagnosis Without Sharing Patient Data
- SuperRivolution: Fine-Scale Rivers from Coarse Temporal Satellite Imagery
- AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
- Deep Learning for Metabolic Rate Estimation from Biosignals: A Comparative Study of Architectures and Signal Selection
- Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
- Iterated Population Based Training with Task-Agnostic Restarts
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- Blurred Encoding for Trajectory Representation Learning
- Free-Boundary Quasiconformal Maps via a Least-squares Operator in Diffeomorphism Optimization
- TransactionGPT
- Improving VisNet for Object Recognition
- Weaver: Kronecker Product Approximations of Spatiotemporal Attention for Traffic Network Forecasting
- Improve Contrastive Clustering Performance by Multiple Fusing-Augmenting ViT Blocks
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- FedSDWC: Federated Synergistic Dual-Representation Weak Causal Learning for OOD
- Classifying Phonotrauma Severity from Vocal Fold Images with Soft Ordinal Regression
- A Neural-Operator Preconditioned Newton Method for Accelerated Nonlinear Solvers
- Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
- Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
- Stabilizing Direct Training of Spiking Neural Networks: Membrane Potential Initialization and Threshold-robust Surrogate Gradient
- SENCA-st: Integrating Spatial Transcriptomics and Histopathology with Cross Attention Shared Encoder for Region Identification in Cancer Pathology
- CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
- Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
- RDTE-UNet: A Boundary and Detail Aware UNet for Precise Medical Image Segmentation
- A Generative Adversarial Approach to Adversarial Attacks Guided by Contrastive Language-Image Pre-trained Model
- Measuring the Intrinsic Dimension of Earth Representations
- REASON: Probability map-guided dual-branch fusion framework for gastric content assessment
- Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
- NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
- Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
- RAPTR: Radar-based 3D Pose Estimation using Transformer
- Extreme Model Compression with Structured Sparsity at Low Precision
- Mitigating Negative Flips via Margin Preserving Training
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering
- Rethinking Explanation Evaluation under the Retraining Scheme
- X-IONet: Cross-Platform Inertial Odometry Network with Dual-Stage Attention
- Data-Driven Discovery of Feature Groups in Clinical Time Series
- Fill the gaps: continuous in time interpolation of fluid dynamical simulations
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Towards Provably Unlearnable Examples via Bayes Error Optimisation
- Pixel-level Quality Assessment for Oriented Object Detection
- GAMA: A Neural Neighborhood Search Method with Graph-aware Multi-modal Attention for Vehicle Routing Problem
- WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
- VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- ProbSelect: Stochastic Client Selection for GPU-Accelerated Compute Devices in the 3D Continuum
- BIPPO: Budget-Aware Independent PPO for Energy-Efficient Federated Learning Services
- From Sequential to Recursive: Enhancing Decision-Focused Learning with Bidirectional Feedback
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Invisible Triggers, Visible Threats! Road-Style Adversarial Creation Attack for Visual 3D Detection in Autonomous Driving
- Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition
- Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal Representation
- ReIDMamba: Learning Discriminative Features with Visual State Space Model for Person Re-Identification
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks
- Multi-objective Hyperparameter Optimization in the Age of Deep Learning
- The Impact of Longitudinal Mammogram Alignment on Breast Cancer Risk Assessment
- Data Descriptions from Large Language Models with Influence Estimation
- Visual Bridge: Universal Visual Perception Representations Generating
- MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection
- A General Method for Proving Networks Universal Approximation Property
- VEDA: 3D Molecular Generation via Variance-Exploding Diffusion with Annealing
- Deep Learning Analysis of Prenatal Ultrasound for Identification of Ventriculomegaly
- DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model
- Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation
- Operational machine learning for remote spectroscopic detection of CH4 point sources
- Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement
- CAVER: Curious Audiovisual Exploring Robot
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
- Verifying rich robustness properties for neural networks
- CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
- Revisiting the Neural Tangent Kernel: the role of large width and depth
- Integrating Epigenetic and Phenotypic Features for Biological Age Estimation in Cancer Patients via Multimodal Learning
- BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
- HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving
- Pandar128 dataset for lane line detection
- GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
- Fair Bayesian Data Selection via Generalized Discrepancy Measures
- Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data
- From Attribution to Action: Jointly ALIGNing Predictions and Explanations
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- MI-to-Mid Distilled Compression (M2M-DC): An Hybrid-Information-Guided-Block Pruning with Progressive Inner Slicing Approach to Model Compression
- Minimum Width of Deep Narrow Networks for Universal Approximation
- PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks
- K-Stain: Keypoint-Driven Correspondence for H&E-to-IHC Virtual Staining
- Probably Approximately Global Robustness Certification
- Flexible Concept Bottleneck Model
- REOcc: Camera-Radar Fusion with Radar Feature Enrichment for 3D Occupancy Prediction
- Adaptive Initial Residual Connections for GNNs with Theoretical Guarantees
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
- Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization
- SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation
- Achieving Fairness Without Harm via Selective Demographic Experts
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- Temporal-Guided Visual Foundation Models for Event-Based Vision
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- ReMoD: Rethinking Modality Contribution in Multimodal Stance Detection via Dual Reasoning
- Online Learning of Modular Bayesian Deep Receivers: Single-Step Adaptation with Streaming Data
- Radio AGN feedback sustains quiescence only in a minority of massive galaxies
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
- Adaptive Agent Selection and Interaction Network for Image-to-point cloud Registration
- Global Multiple Extraction Network for Low-Resolution Facial Expression Recognition
- LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
- EndoIR: Degradation-Agnostic All-in-One Endoscopic Image Restoration via Noise-Aware Routing Diffusion
- MACMD: Multi-dilated Contextual Attention and Channel Mixer Decoding for Medical Image Segmentation
- Sign language recognition from skeletal data using graph and recurrent neural networks
- Environment-Aware MIMO Channel Estimation in Pilot-Constrained Upper Mid-Band Systems
- Towards Unified AI-Driven Fracture Mechanics: The Extended Deep Energy Method (XDEM)
- An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
- SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning
- Sharing the Learned Knowledge-base to Estimate Convolutional Filter Parameters for Continual Image Restoration
- Climate Downscaling of Tropical Cyclone Intensity using Deep Learning
- The Causal Round Trip: Generating Authentic Counterfactuals by Eliminating Information Loss
- NeuroFlex: Column-Exact ANN-SNN Co-Execution Accelerator with Cost-Guided Scheduling
- MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification
- From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
- NovisVQ: A Streaming Convolutional Neural Network for No-Reference Opinion-Unaware Frame Quality Assessment
- MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purification for Inference-Time Attacks
- Order-Level Attention Similarity Across Language Models: A Latent Commonality
- Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
- MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- Learning Fourier shapes to probe the geometric world of deep neural networks
- The Future of Fully Homomorphic Encryption System: from a Storage I/O Perspective
- Unveiling the Training Dynamics of ReLU Networks through a Linear Lens
- Beta Distribution Learning for Reliable Roadway Crash Risk Assessment
- What's on Your Plate? Inferring Chinese Cuisine Intake from Wearable IMUs
- Accelerating metamaterial topology optimization using deep super-resolution networks
- Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
- Automatic detection of CMEs using synthetically-trained Mask R-CNN
- Linear Mode Connectivity under Data Shifts for Deep Ensembles of Image Classifiers
- Saliency-Guided Domain Adaptation for Left-Hand Driving in Autonomous Steering
- Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- A MATLAB tutorial on deep feature extraction combined with chemometrics for analytical applications
- Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
- AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
- MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
- Probing the Probes: Methods and Metrics for Concept Alignment
- Proto-LeakNet: Towards Signal-Leak Aware Attribution in Synthetic Human Face Imagery
- An Efficient Algorithm for Learning-Based Visual Localization
- Active Domain Adaptation for mmWave-based HAR via Renyi Entropy-based Uncertainty Estimation
- FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
- Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment
- MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Distribution-Aware Tensor Decomposition for Compression of Convolutional Neural Networks
- Event Reconstruction for Radio-Based In-Ice Neutrino Detectors with Neural Posterior Estimation
- A Transferable Machine Learning Approach to Predict Quantum Circuit Parameters for Electronic Structure Problems
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
- nanoTabPFN: A Lightweight and Educational Reimplementation of TabPFN
- MS2Edge: Towards Energy-Efficient and Crisp Edge Detection with Multi-Scale Residual Learning in SNNs
- Efficient Neural Networks with Discrete Cosine Transform Activations
- Extended Physics Informed Neural Network for Hyperbolic Two-Phase Flow in Porous Media
- Gradient Projection onto Historical Descent Directions for Communication-Efficient Federated Learning
- Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
- Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
- Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels
- Decentralized Federated Learning with Distributed Aggregation Weight Optimization
- Decoupled Entropy Minimization
- Cross-Modal Alignment via Variational Copula Modelling
- SAAIPAA: Optimizing aspect-angles-invariant physical adversarial attacks on SAR target recognition models
- Deploying Rapid Damage Assessments from sUAS Imagery for Disaster Response
- An Augmentation Overlap Theory of Contrastive Learning
- Scalable Autoregressive Deep Surrogates for Dendritic Microstructure Dynamics
- Contextual Role Modulates Object Representational Geometry in the Human Brain
- Conditional Diffusion Model-Enabled Scenario-Specific Neural Receivers for Superimposed Pilot Schemes
- Archaeological Classification of Small Datasets Using Meta- and Transfer Learning Methods: A Case Study on Hittite Stele Fragments
- Learning with less: label-efficient land cover classification at very high spatial resolution using self-supervised deep learning
- Generative Hints
- UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
- Dynamic Priors in Bayesian Optimization for Hyperparameter Optimization
- Improving Unlearning with Model Updates Probably Aligned with Gradients
- Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
- MammoClean: Toward Reproducible and Bias-Aware AI in Mammography through Dataset Harmonization
- Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster
- GAFD-CC: Global-Aware Feature Decoupling with Confidence Calibration for OOD Detection
- Differentiable Hierarchical Visual Tokenization
- Collaborative Attention and Consistent-Guided Fusion of MRI and PET for Alzheimer's Disease Diagnosis
- OmniField: Conditioned Neural Fields for Robust Multimodal Spatiotemporal Learning
- MM-UNet: Morph Mamba U-shaped Convolutional Networks for Retinal Vessel Segmentation
- PrivGNN: High-Performance Secure Inference for Cryptographic Graph Neural Networks
- CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning
- HAGI++: Head-Assisted Gaze Imputation and Generation
- A Proof of Learning Rate Transfer under μP
- TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
- Machine and Deep Learning for Indoor UWB Jammer Localization
- Bridging Lifelong and Multi-Task Representation Learning via Algorithm and Complexity Measure
- Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- Motion-Robust Multimodal Fusion of PPG and Accelerometer Signals for Three-Class Heart Rhythm Classification
- Lightweight ResNet-Based Deep Learning for Photoplethysmography Signal Quality Assessment
- Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion
- Perturbations in the Orthogonal Complement Subspace for Efficient Out-of-Distribution Detection
- Parameter Interpolation Adversarial Training for Robust Image Classification
- Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
- Validating Deep Models for Alzheimer's 18F-FDG PET Diagnosis Across Populations: A Study with Latin American Data
- Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
- Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- Digital Twin of Aerosol Jet Printing
- Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
- Multi-refined Feature Enhanced Sentiment Analysis Using Contextual Instruction
- OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback
- Weakly Supervised Pneumonia Localization from Chest X-Rays Using Deep Neural Network and Grad-CAM Explanations
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
- Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Towards Reliable Pediatric Brain Tumor Segmentation: Task-Specific nnU-Net Enhancements
- A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method
- Towards Automated Petrography
- TRISKELION-1: Unified Descriptive-Predictive-Generative AI
- STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
- Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection
- Melanoma Classification Through Deep Ensemble Learning and Explainable AI
- Machine learning-based cloud resource allocation algorithms: a comprehensive comparative review
- Approximating Young Measures With Deep Neural Networks
- An Efficient and Generalizable Transfer Learning Method for Weather Condition Detection on Ground Terminals
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- Vision Transformer for Robust Occluded Person Reidentification in Complex Surveillance Scenes
- VessShape: Few-shot 2D blood vessel segmentation by leveraging shape priors from synthetic images
- Towards robust quantitative photoacoustic tomography via learned iterative methods
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
- MeisenMeister: A Simple Two Stage Pipeline for Breast Cancer Classification on MRI
- CASR-Net: An Image Processing-focused Deep Learning-based Coronary Artery Segmentation and Refinement Network for X-ray Coronary Angiogram
- Rethinking Robust Adversarial Concept Erasure in Diffusion Models
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
- FedSM: Robust Semantics-Guided Feature Mixup for Bias Reduction in Federated Learning with Long-Tail Data
- Object-IR: Leveraging Object Consistency and Mesh Deformation for Self-Supervised Image Retargeting
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- SpecAware: A Spectral-Content Aware Foundation Model for Unifying Multi-Sensor Learning in Hyperspectral Remote Sensing Mapping
- SilhouetteTell: Practical Video Identification Leveraging Blurred Recordings of Video Subtitles
- AFM-Net: Advanced Fusing Hierarchical CNN Visual Priors with Global Sequence Modeling for Remote Sensing Image Scene Classification
- Exploring Landscapes for Better Minima along Valleys
- Exact Terminal Condition Neural Network for American Option Pricing Based on the Black-Scholes-Merton Equations
- Lightweight CNN Model Hashing with Higher-Order Statistics and Chaotic Mapping for Piracy Detection and Tamper Localization
- Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
- A Retrospect to Multi-prompt Learning across Vision and Language
- Information-Theoretic Greedy Layer-wise Training for Traffic Sign Recognition
- ANCHOR: Integrating Adversarial Training with Hard-mined Supervised Contrastive Learning for Robust Representation Learning
- DP-FedPGN: Finding Global Flat Minima for Differentially Private Federated Learning via Penalizing Gradient Norm
- FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
- End-to-End Framework Integrating Generative AI and Deep Reinforcement Learning for Autonomous Ultrasound Scanning
- Merlin L48 Spectrogram Dataset
- NOMAD -- Navigating Optimal Model Application to Datastreams
- Gaussian Combined Distance: A Generic Metric for Object Detection
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- AD-SAM: Fine-Tuning the Segment Anything Vision Foundation Model for Autonomous Driving Perception
- Incremental Human-Object Interaction Detection with Invariant Relation Representation Learning
- Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget
- Semantic Frame Aggregation-based Transformer for Live Video Comment Generation
- An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
- FlowQ-Net: A Generative Framework for Automated Quantum Circuit Design
- Improving Classification of Occluded Objects through Scene Context
- MSAD: A Deep Dive into Model Selection for Time series Anomaly Detection
- Hybrid Physical-Neural Simulator for Fast Cosmological Hydrodynamics
- SA2Net: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging
- On Measuring Localization of Shortcuts in Deep Networks
- MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
- SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification
- A Hybrid Framework Bridging CNN and ViT based on Theory of Evidence for Diabetic Retinopathy Grading
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- Leveraging Large-Scale Face Datasets for Deep Periocular Recognition via Ocular Cropping
- Exploring Complementarity and Explainability in CNNs for Periocular Verification Across Acquisition Distances
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- Security Risk of Misalignment between Text and Image in Multi-modal Model
- Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
- Comparative Study of UNet-based Architectures for Liver Tumor Segmentation in Multi-Phase Contrast-Enhanced Computed Tomography
- Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
- Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
- CHIPSIM: A Co-Simulation Framework for Deep Learning on Chiplet-Based Systems
- Reading Radio from Camera: Visually-Grounded, Lightweight, and Interpretable RSSI Prediction
- Active Learning with Task-Driven Representations for Messy Pools
- VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
- Feedback Alignment Meets Low-Rank Manifolds: A Structured Recipe for Local Learning
- FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
- Beyond Data Scarcity Optimizing R3GAN for Medical Image Generation from Small Datasets
- Adaptive End-to-End Transceiver Design for NextG Pilot-Free and CP-Free Wireless Systems
- A Deep Learning Framework for Multi-Operator Learning: Architectures and Approximation Theory
- Lightweight Federated Learning in Mobile Edge Computing with Statistical and Device Heterogeneity Awareness
- MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
- When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
- Image Quality Dependent Degradation for AI Systems
- PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter Power
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
- Deep Stacked Networks with Residual Polishing for Image Inpainting
- Binary Stochastic Representations for Large Multi-class Classification
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
- Unsupervised Learning of Solutions to Differential Equations with Generative Adversarial Networks
- Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations
- Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM
- HERMES: A Hybrid Ensemble for Head-and-Neck Tumor Segmentation, TN Staging, and Recurrence-Free Survival on PET/CT
- Calibrating the Digital Twin Channel: Statistics-Consistent Sim-to-Lab Adaptation for W-Band Industrial OFDM Links
- FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
- A Comparison of Data Augmentation Methods for Training Deep Neural Networks on Synthetic Aperture Sonar
- Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution
- Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification
- Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- A Hierarchical Framework for Graph Structure Learning in Histopathology Image Classification
- Lilith: Backdoor Generalization under Training-Inference Trigger Shift
- DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
- HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework
- Shape-Based Inductive Bias for Glioma Grading from Tumor Contours
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- Mixture-of-Depths Attention
- Consistent Accelerated Inference via Confident Adaptive Transformers
- Achieving Human Parity on Automatic Chinese to English News Translation
- Automatic Liver Segmentation from CT Images Using Deep Learning Algorithms: A Comparative Study
- Open Source Face Recognition Performance Evaluation Package
- Fusion++: Volumetric Object-Level SLAM
- Analysing object detectors from the perspective of co-occurring object\n categories
- Maximum Classifier Discrepancy for Unsupervised Domain Adaptation
- Self-Supervision Closes the Gap Between Weak and Strong Supervision in Histology
- Distance-aware Soft Prompt Learning for Multimodal Valence-Arousal Estimation
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Stable Tensor Neural Networks for Rapid Deep Learning
- Efficient Sparse-Winograd Convolutional Neural Networks
- Learning Digital Camera Pipeline for Extreme Low-Light Imaging
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Soft Calibration Objectives for Neural Networks
- Visual Data Augmentation through Learning
- AI and Image
- Dominant Controls on Preferential Flow and Their Implications for Future Soil Water Fluxes
- Understanding Infographics through Textual and Visual Tag Prediction
- Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification
- First-Order Preconditioning via Hypergradient Descent
- Superposition disentanglement of neural representations reveals hidden alignment
- Calibrated Uncertainty Sampling for Active Learning
- Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification
- FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning
- Improve Unsupervised Pretraining for Few-label Transfer
- A one-armed CNN for exoplanet detection from light curves
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- InkShield: Writing Style Protection Against Unauthorized Handwriting Mimicry
- SeGAN: Segmenting and Generating the Invisible
- Deep Feature Flow for Video Recognition
- Knowledge Adaptation for Efficient Semantic Segmentation
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- Learnable Pooling Methods for Video Classification
- Granular Motor State Monitoring of Free Living Parkinson's Disease Patients via Deep Learning
- Mapping small reservoirs across Brazil from 1984 to 2025
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Enforcing Reasoning in Visual Commonsense Reasoning
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- Large-scale empirical tuning and comparison of default optimizers for variational inference
- Scene-Centric Unsupervised Video Panoptic Segmentation
- ShaplEIG: Bayesian Experimental Design for Shapley Value Estimation
- Optimality of Sub-network Laplace Approximations: New Results and Methods
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- LR-to-HR Face Hallucination with an Adversarial Progressive\n Attribute-Induced Network
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Generative Partition Networks for Multi-Person Pose Estimation
- Natural Adversarial Examples
- Hyperspectral image classification using CNN: Application to industrial food packaging
- A Revisit on Deep Hashings for Large-scale Content Based Image Retrieval
- OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI
- A Theory of Generalization in Deep Learning
- Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
- Recurrent 3D Pose Sequence Machines
- GResNet: Graph Residual Network for Reviving Deep GNNs from Suspended Animation
- Dynamic Few-Shot Visual Learning without Forgetting
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Deep Gradient Projection Networks for Pan-sharpening
- Representation range needs for 16-bit neural network training
- SAWNet: A Spatially Aware Deep Neural Network for 3D Point Cloud Processing
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- Deep Spatial Regression Model for Image Crowd Counting
- FHEDN: A based on context modeling Feature Hierarchy Encoder-Decoder Network for face detection
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Detecting abnormal events in video using Narrowed Normality Clusters
- Zero-shot World Models Are Developmentally Efficient Learners
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
- Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
- RaCo: Ranking and Covariance for Practical Learned Keypoints
- Diagnosing Error in Temporal Action Detectors
- Generative Modeling via Drifting
- Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition
- Fourier-Based GAN Fingerprint Detection using ResNet50
- Robust and Generalizable Atrial Fibrillation Detection from ECG Using Time-Frequency Fusion and Supervised Contrastive Learning
- Enhancing Rotated Object Detection via Anisotropic Gaussian Bounding Box and Bhattacharyya Distance
- Generating metamers of human scene understanding
- High-throughput Verticillium wilt detection in cotton: A comparative study of faster R-CNN and YOLOv11
- Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- Where to Focus: Query Adaptive Matching for Instance Retrieval Using Convolutional Feature Maps
- Generalizable and scalable protein stability prediction with rewired protein generative models
- Contrastive learning enhances fairness in pathology artificial intelligence systems
- Deep learning technique for plant disease classification and pest detection and model explainability elevating agricultural sustainability
- EdgeSync: Accelerating Edge-Model Updates for Data Drift through Adaptive Continuous Learning
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- Cost-Sensitive Unbiased Risk Estimation for Multi-Class Positive-Unlabeled Learning
- Energy-Efficient Autonomous Driving with Adaptive Perception and Robust Decision
- Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
- Adversarially Robust Quantum Transfer Learning
- Cosmic Background Removal with Deep Neural Networks in SBND
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- Selective Diabetic Retinopathy Screening with Accuracy-Weighted Deep Ensembles and Entropy-Guided Abstention
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
- Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- A Quadratic Actor Network for Model-Free Reinforcement Learning
- Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
- SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving
- Pyramid Scene Parsing Network
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
- MIC-BEV: Multi-Infrastructure Camera Bird's-Eye-View Transformer with Relation-Aware Fusion for 3D Object Detection
- Eigenfunction Extraction for Ordered Representation Learning
- All in one timestep: Enhancing Sparsity and Energy efficiency in Multi-level Spiking Neural Networks
- Exploring Federated Learning for Thermal Urban Feature Segmentation -- A Comparison of Centralized and Decentralized Approaches
- Time-Embedded Algorithm Unrolling for Computational MRI
- SPLite Hand: Sparsity-Aware Lightweight 3D Hand Pose Estimation
- Deep-Learning-Empowered Programmable Topolectrical Circuits
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- When are radiology reports useful for training medical image classifiers?
- A Domain Adaptive Position Reconstruction Method for Time Projection Chamber based on Deep Neural Network
- Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning
- UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
- Global Chlorophyll-a Retrieval algorithm from Sentinel 2 Using Residual Deep Learning and Novel Machine Learning Water Classification
- Unlocking Out-of-Distribution Generalization in Dynamics through Physics-Guided Augmentation
- CLFSeg: A Fuzzy-Logic based Solution for Boundary Clarity and Uncertainty Reduction in Medical Image Segmentation
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- EddyFormer: Accelerated Neural Simulations of Three-Dimensional Turbulence at Scale
- Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- Differential Privacy: Gradient Leakage Attacks in Federated Learning Environments
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- SLOTH: Lightweight Detection and Localization of On-Chip Fail-Slow Failures for DNN Accelerators
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- Learning from History: A Retrieval-Augmented Framework for Spatiotemporal Prediction
- Mitigating Negative Transfer via Reducing Environmental Disagreement
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- ResNet: Enabling Deep Convolutional Neural Networks through Residual Learning
- An efficient probabilistic hardware architecture for diffusion-like models
- Exploring an image-based b-jet tagging method using convolution neural networks
- Unmasking Facial DeepFakes: A Robust Multiview Detection Framework for Natural Images
- LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
- Improving the Straight-Through Estimator with Zeroth-Order Information
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Incorporating Symmetry into Deep Dynamics Models for Improved Generalization
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- IAEmu: Learning Galaxy Intrinsic Alignment Correlations
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- Lung Cancer Classification from CT Images Using ResNet
- Super-recognizers sample visual information of superior computational value for facial recognition
- Image Categorization and Search via a GAT Autoencoder and Representative Models
- Connectome-Guided Automatic Learning Rates for Deep Networks
- Explaining Robustness to Catastrophic Forgetting Through Incremental Concept Formation
- Chebyshev Moment Regularization (CMR): Condition-Number Control with Moment Shaping
- Robust and Generalizable Background Subtraction on Images of Calorimeter Jets using Unsupervised Generative Learning
- Localising under the drape: proprioception in the era of distributed surgical robotic system
- iPac: Incorporating Intra-image Patch Context into Graph Neural Networks for Medical Image Classification
- CURVETE: Curriculum Learning and Progressive Self-supervised Training for Medical Image Classification
- VIKING: Deep variational inference with stochastic projections
- PrivacyGuard: A Modular Framework for Privacy Auditing in Machine Learning
- Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach
- One-Timestep is Enough: Achieving High-performance ANN-to-SNN Conversion via Scale-and-Fire Neurons
- An Efficient Remote Sensing Super Resolution Method Exploring Diffusion Priors and Multi-Modal Constraints for Crop Type Mapping
- Interpretable Tile-Based Classification of Paclitaxel Exposure
- Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting
- Approaching Domain Generalization with Embeddings for Robust Discrimination and Recognition of RF Communication Signals
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
- Reliable Robotic Task Execution in the Face of Anomalies
- Neural Emulator Superiority: When Machine Learning for PDEs Surpasses its Training Data
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- Softmax is 1/2-Lipschitz: A tight bound across all ℓp norms
- LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image Segmentation
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
- From Uniform to Adaptive: General Skip-Block Mechanisms for Efficient PDE Neural Operators
- Bi-Encoder Contrastive Learning for Fingerprint and Iris Biometrics
- DNN-based Signal Processing for Liquid Argon Time Projection Chambers
- Transforming volcanic monitoring: A dataset and benchmark for onboard volcano activity detection
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- What Do We Understand About Convolutional Networks?
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Error Adjustment Based on Spatiotemporal Correlation Fusion for Traffic Forecasting
- SemiETPicker: Fast and Label-Efficient Particle Picking for CryoET Tomography Using Semi-Supervised Learning
- Estimation of Fireproof Structure Class and Construction Year for Disaster Risk Assessment
- Block Coordinate Descent for Neural Networks Provably Finds Global Minima
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame Projections
- PSScreen V2: Partially Supervised Multiple Retinal Disease Screening
- Prediction-Powered Semi-Supervised Learning with Online Power Tuning
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- An End-to-End Generative Diffusion Model for Heavy-Ion Collisions
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
- Mutual Information guided Visual Contrastive Learning
- C-arm Guidance: A Self-supervised Approach To Automated Positioning During Stroke Thrombectomy
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Stable neural networks and connections to continuous dynamical systems
- DiffusionLane: Diffusion Model for Lane Detection
- Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
- Audio Frequency-Time Dual Domain Evaluation on Depression Diagnosis
- GALA: A GlobAl-LocAl Approach for Multi-Source Active Domain Adaptation
- Simplifying Knowledge Transfer in Pretrained Models
- Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
- LOC: A General Language-Guided Framework for Open-Set 3D Occupancy Prediction
- FermiNets: Learning generative machines to generate efficient neural networks via generative synthesis
- MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification
- Frequentist Validity of Epistemic Uncertainty Estimators
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Caption-Driven Explainability: Probing CNNs for Bias via CLIP
- LiteDiff
- From Black-box to Causal-box: Towards Building More Interpretable Models
- Revisiting Orbital Minimization Method for Neural Operator Decomposition
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
- AURASeg: Attention Guided Upsampling with Residual Boundary-Assistive Refinement for Drivable-Area Segmentation
- FrameShield: Adversarially Robust Video Anomaly Detection
- GRAP-MOT: Unsupervised Graph-based Position Weighted Person Multi-camera Multi-object Tracking in a Highly Congested Space
- ITC-RWKV: Interactive Tissue-Cell Modeling with Recurrent Key-Value Aggregation for Histopathological Subtyping
- Towards Explainable Personalized Recommendations by Learning from Users' Photos
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Large Language Models as Model Organisms for Human Associative Learning
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration
- Cavity Duplexer Tuning with 1d Resnet-like Neural Networks
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
- Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
- Long-tailed Species Recognition in the NACTI Wildlife Dataset
- Confounding Robust Deep Reinforcement Learning: A Causal Approach
- Multimodal Detection of Fake Reviews using BERT and ResNet-50
- Elementary, My Dear Watson: Non-Invasive Neural Keyword Spotting in the LibriBrain Dataset
- Generative Point Tracking with Flow Matching
- Lens Model Accuracy in the Expected LSST Lensed AGN Sample
- HybridSOMSpikeNet: A Deep Model with Differentiable Soft Self-Organizing Maps and Spiking Dynamics for Waste Classification
- H-SPLID: HSIC-based Saliency Preserving Latent Information Decomposition
- Convergence Analysis of SGD under Expected Smoothness
- Hardware-Aware DNN Compression for Homogeneous Edge Devices
- Predicting the 3D microstructure of SOFC anodes from 2D SEM images using stochastic microstructure modeling and CNNs
- Revisiting Model Stitching to Compare Neural Representations
- Multimodal Negative Learning
- PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
- Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Comparison of Deep learning models on time series forecasting : a case study of Dissolved Oxygen Prediction
- What Does It Take to Build a Performant Selective Classifier?
- Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
- HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
- Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process
- GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
- Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
- Deep Discriminative Clustering Analysis
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
- Reliable and Reproducible Demographic Inference for Fairness in Face Analysis
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking
- The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models
- Vision-language models learn the geometry of human perceptual space
- From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
- BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism
- Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
- DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- Exploring "Many in Few" and "Few in Many" Properties in Long-Tailed, Highly-Imbalanced IC Defect Classification
- AutoMT: A Multi-Agent LLM Framework for Automated Metamorphic Testing of Autonomous Driving Systems
- Revisiting the Relation Between Robustness and Universality
- Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
- Modeling Turn-Taking with Semantically Informed Gestures
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- Calibration and Discrimination Optimization Using Clusters of Learned Representation
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
- A Simple Loss Function for Improving the Convergence and Accuracy of Visual Question Answering Models
- Brain-Inspired Perspective on Configurations: Unsupervised Similarity and Early Cognition
- HAMLOCK: HArdware-Model LOgically Combined attacK
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- TinyUSFM: Towards Compact and Efficient Ultrasound Foundation Models
- Enhancing Early Alzheimer Disease Detection through Big Data and Ensemble Few-Shot Learning
- Transformed Multi-view 3D Shape Features with Contrastive Learning
- Online Handwritten Signature Verification Based on Temporal-Spatial Graph Attention Transformer
- FrogDeepSDM: Improving Frog Counting and Occurrence Prediction Using Multimodal Data and Pseudo-Absence Imputation
- Image-Image Domain Adaptation with Preserved Self-Similarity and Domain-Dissimilarity for Person Re-identification
- A Novel Combined Optical Flow Approach for Comprehensive Micro-Expression Recognition
- Towards Strong Certified Defense with Universal Asymmetric Randomization
- RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
- Controllable Machine Unlearning via Gradient Pivoting
- Matrix-Free Least Squares Solvers: Values, Gradients, and What to Do With Them
- QCFace: Image Quality Control for boosting Face Representation & Recognition
- Denoising Complex Covariance Matrices with Hybrid ResNet and Random Matrix Theory: Cryptocurrency Portfolio Applications
- Explainable Deep Learning in Medical Imaging: Brain Tumor and Pneumonia Detection
- Wavelet-based GAN Fingerprint Detection using ResNet50
- Weight Decay may matter more than muP for Learning Rate Transfer in Practice
- Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
- A Unified Perspective on Optimization in Machine Learning and Neuroscience: From Gradient Descent to Neural Adaptation
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Beyond the Pipeline: Analyzing Key Factors in End-to-End Deep Learning for Historical Writer Identification
- Training Keyword Spotters with Limited and Synthesized Speech Data
- Improving Micro-Expression Recognition with Phase-Aware Temporal Augmentation
- Deformable Convolutional Networks
- Towards In-Situ Failure Assessment: Deep Learning on DIC Results for Laminated Composites
- C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Instance-level Human Parsing via Part Grouping Network
- Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel Optimization
- Hybrid Deep Learning Framework for Enhanced Diabetic Retinopathy Detection: Integrating Traditional Features with AI-driven Insights
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- AWSPNet: Attention-based Dual-Tree Wavelet Scattering Prototypical Network for MIMO Radar Target Recognition and Jamming Suppression
- S2AP: Score-space Sharpness Minimization for Adversarial Pruning
- Learning Human-Object Interaction as Groups
- Robust High-Resolution Multi-Organ Diffusion MRI Using Synthetic-Data-Tuned Prompt Learning
- AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
- ShortcutBreaker: Low-Rank Noisy Bottleneck with Global Perturbation Attention for Multi-Class Unsupervised Anomaly Detection
- MCANet: A Coherent Multimodal Collaborative Attention Network for Advanced Modulation Recognition in Adverse Noisy Environments
- Uncertainty Estimation by Flexible Evidential Deep Learning
- Beyond Frequency: Scoring-Driven Debiasing for Object Detection via Blueprint-Prompted Image Synthesis
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning
- POLAR: Policy-based Layerwise Reinforcement Learning Method for Stealthy Backdoor Attacks in Federated Learning
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers
- Adaptive transfer learning for surgical tool presence detection in laparoscopic videos through gradual freezing fine-tuning
- MARS-M: When Variance Reduction Meets Matrices
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- Enhancing Cross-Patient Generalization in AI-Based Parkinson s Disease Detection
- Towards 3D Objectness Learning in an Open World
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning
- How human–AI feedback loops alter human perceptual, emotional and social judgements
- RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries
- MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning
- Exploration via Feature Perturbation in Contextual Bandits
- Facial Expression-based Parkinson's Disease Severity Diagnosis via Feature Fusion and Adaptive Class Balancing
- Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
- Deep Neural Network extraction of Unpolarized Transverse Momentum Distributions
- Pano2CAD: Room Layout From A Single Panorama Image
- 2D3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
- GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection
- Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training
- Boosting Fidelity for Pre-Trained-Diffusion-Based Low-Light Image Enhancement via Condition Refinement
- Benchmarking Out-of-Distribution Detection for Plankton Recognition: A Systematic Evaluation of Advanced Methods in Marine Ecological Monitoring
- Fair and Interpretable Deepfake Detection in Videos
- Mapping Hidden Heritage: Self-supervised Pre-training on High-Resolution LiDAR DEM Derivatives for Archaeological Stone Wall Detection
- GUIDE: Enhancing Gradient Inversion Attacks in Federated Learning with Denoising Models
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- Annotation-Efficient Universal Honesty Alignment
- EndoCIL: A Class-Incremental Learning Framework for Endoscopic Image Classification
- NeuCo-Bench: A Novel Benchmark Framework for Neural Embeddings in Earth Observation
- Conditional Synthetic Live and Spoof Fingerprint Generation
- STARK: Strategic Team of Agents for Refining Kernels
- State estimation in homogeneous isotropic turbulence using super-resolution with a 4DVar training algorithm
- Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision
- SAMOSA: Sharpness Aware Minimization for Open Set Active learning
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Symmetric Entropy-Constrained Video Coding for Machines
- Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
- An Efficient Semantic Segmentation Decoder for In-Car or Distributed Applications
- Class-N-Diff: Classification-Induced Diffusion Model Can Make Fair Skin Cancer Diagnosis
- The Sherpa.ai Blind Vertical Federated Learning Paradigm to Minimize the Number of Communications
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Geospatial Machine Learning Libraries
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Residual Correction Models for AC Optimal Power Flow Using DC Optimal Power Flow Solutions
- Learning a Generalized Model for Substation Level Voltage Estimation in Distribution Networks
- EdgeNavMamba: Mamba Optimized Object Detection for Energy Efficient Edge Devices
- GENESIS: A Generative Model of Episodic-Semantic Interaction
- Deep Neural ODE Operator Networks for PDEs
- Compressive Modeling and Visualization of Multivariate Scientific Data using Implicit Neural Representation
- Dynamic Recalibration in LiDAR SLAM: Integrating AI and Geometric Methods with Real-Time Feedback Using INAF Fusion
- OpenEDS: Open Eye Dataset
- Semantic segmentation with coarse annotations
- Rethinking Convergence in Deep Learning: The Predictive-Corrective Paradigm for Anatomy-Informed Brain MRI Segmentation
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
- Automated C-Arm Positioning via Conformal Landmark Localization
- RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation
- Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
- VM-BeautyNet: A Synergistic Ensemble of Vision Transformer and Mamba for Facial Beauty Prediction
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- Hyperparameter Optimization and Reproducibility in Deep Learning Model Training
- Fourier Transform Multiple Instance Learning for Whole Slide Image Classification
- Extending Temporal Disturbance Estimations For Magnetic Anomaly Navigation and Mapping
- A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects
- Adiabatic transport of neural network quantum states
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
- C4D: 4D Made from 3D through Dual Correspondences
- OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- A Multi-Task Deep Learning Framework for Skin Lesion Classification, ABCDE Feature Quantification, and Evolution Simulation
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
- Geometric Moment Alignment for Domain Adaptation via Siegel Embeddings
- EuroMineNet: A Multitemporal Sentinel-2 Benchmark for Spatiotemporal Mining Footprint Analysis in the European Union (2015-2024)
- Causality Enhancement for Cross-Domain Recommendation
- SteeringTTA: Guiding Diffusion Trajectories for Robust Test-Time-Adaptation
- Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Selective Labeling with False Discovery Rate Control
- Boosted Attention: Leveraging Human Attention for Image Captioning
- EcoScaleNet: A Lightweight Multi Kernel Network for Long Sequence 12 lead ECG Classification
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Vision Mamba for Permeability Prediction of Porous Media
- Unsupervised Deep Generative Models for Anomaly Detection in Neuroimaging: A Systematic Scoping Review
- Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
- BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection
- A Multi-domain Image Translative Diffusion StyleGAN for Iris Presentation Attack Detection
- TED++: Submanifold-Aware Backdoor Detection via Layerwise Tubular-Neighbourhood Screening
- CLEAR: Causal Learning Framework For Robust Histopathology Tumor Detection Under Out-Of-Distribution Shifts
- PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
- Spiking Neural Network Architecture Search: A Survey
- When Flatness Does (Not) Guarantee Adversarial Robustness
- LOTA: Bit-Planes Guided AI-Generated Image Detection
- Global-focal Adaptation with Information Separation for Noise-robust Transfer Fault Diagnosis
- Beat Tracking as Object Detection
- ViTacGen: Robotic Pushing with Vision-to-Touch Generation
- Nondeterminism-Aware Optimistic Verification for Floating-Point Neural Networks
- Conditional Clifford-Steerable CNNs with Complete Kernel Basis for PDE Modeling
- Revisiting Video Saliency: A Large-scale Benchmark and a New Model
- Axial Neural Networks for Dimension-Free Foundation Models
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
- Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
- Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
- Removing Cost Volumes from Optical Flow Estimators
- UniVector: Unified Vector Extraction via Instance-Geometry Interaction
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- Approximate Bilevel Graph Structure Learning for Histopathology Image Classification
- DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Counting Hallucinations in Diffusion Models
- Transformer-based Scalable Beamforming Optimization via Deep Residual Learning
- Beyond Pixels: A Differentiable Pipeline for Probing Neuronal Selectivity in 3D
- DP-TTA: Test-time Adaptation for Transient Electromagnetic Signal Denoising via Dictionary-driven Prior Regularization
- BlendFL: Blended Federated Learning for Handling Multimodal Data Heterogeneity
- Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
- XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation
- FlashWorld: High-quality 3D Scene Generation within Seconds
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- DocVQA: A Dataset for VQA on Document Images
- Computationally Efficient Neural Receivers via Axial Self-Attention
- CuMPerLay: Learning Cubical Multiparameter Persistence Vectorizations
- AnyUp: Universal Feature Upsampling
- KoALA: KL-L0 Adversarial Detector via Label Agreement
- Assessing the Potential for Catastrophic Failure in Dynamic Post-Training Quantization
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- Layer-Aware Influence for Online Data Valuation Estimation
- Learning Human Motion with Temporally Conditional Mamba
- PubSub-VFL: Towards Efficient Two-Party Split Learning in Heterogeneous Environments via Publisher/Subscriber Architecture
- Fast Visuomotor Policy for Robotic Manipulation
- MS-GAGA: Metric-Selective Guided Adversarial Generation Attack
- A Function Centric Perspective On Flat and Sharp Minima
- Cautious Weight Decay
- Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking
- Simple Projection Variants Improve ColBERT Performance
- Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
- Local Background Features Matter in Out-of-Distribution Detection
- Faster State Preparation with Randomization
- A Gradient Guided Diffusion Framework for Chance Constrained Programming
- Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
- SDGraph: Multi-Level Sketch Representation Learning by Sparse-Dense Graph Architecture
- Beyond the Brightest: A Deep Learning Approach to Identifying Major and Minor Galaxy Mergers in CANDELS at z ∼ 1
- nuGPR: GPU-Accelerated Gaussian Process Regression with Iterative Algorithms and Low-Rank Approximations
- MultiFoodhat: A potential new paradigm for intelligent food quality inspection
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- DIANet: A Phase-Aware Dual-Stream Network for Micro-Expression Recognition via Dynamic Images
- EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
- VCTR: A Transformer-Based Model for Non-parallel Voice Conversion
- Pyramid Stereo Matching Network
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- A deep learning theory for neural networks grounded in physics
- Snapshot renormalization group for quantum matter
- Post-surgical Endometriosis Segmentation in Laparoscopic Videos
- PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
- Lightweight CNN-Based Wi-Fi Intrusion Detection Using 2D Traffic Representations
- Image Classification of Melanoma, Nevus and Seborrheic Keratosis by Deep Neural Network Ensemble
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Adversarial Attacks Leverage Interference Between Features in Superposition
- Hierarchical Qubit-Merging Transformer for Quantum Error Correction
- FACE: Faithful Automatic Concept Extraction
- How many samples to label for an application given a foundation model? Chest X-ray classification study
- Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
- Phase Aware Ear-Conditioned Learning for Multi-Channel Binaural Speaker Separation
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- Exploring and Leveraging Class Vectors for Classifier Editing
- DTEA: Dynamic Topology Weaving and Instability-Driven Entropic Attenuation for Medical Image Segmentation
- Building machines that learn and think like people
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer
- Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
- Metropolis-Scale Road Network Datasets for Fine-Grained Urban Traffic Modeling
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- Source-Free Object Detection with Detection Transformer
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Efficient Edge Test-Time Adaptation via Latent Feature Coordinate Correction
- Benchmarking Deep Learning Models for Laryngeal Cancer Staging Using the LaryngealCT Dataset
- XGrasp: Gripper-Aware Grasp Detection with Multi-Gripper Data Generation
- Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
- Perspective-aware 3D Gaussian Inpainting with Multi-view Consistency
- Adversarial Robustness in One-Stage Learning-to-Defer
- Mixup Helps Understanding Multimodal Video Better
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
- SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
- Predicting the single-site and multi-site event discrimination power of dual-phase time projection chambers
- Nonlinearly Preconditioned Gradient Methods: Momentum and Stochastic Analysis
- Audio-Guided Visual Perception for Audio-Visual Navigation
- Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment
- Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Optimally Deep Networks -- Adapting Model Depth to Datasets for Superior Efficiency
- Restricted Receptive Fields for Face Verification
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping
- Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
- Understanding Self-supervised Contrastive Learning through Supervised Objectives
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- SoundReactor: Frame-level Online Video-to-Audio Generation
- SpurBreast: A Curated Dataset for Investigating Spurious Correlations in Real-world Breast MRI Classification
- Unlocking Symbol-Level Precoding Efficiency Through Tensor Equivariant Neural Network
- PENEX: AdaBoost-Inspired Neural Network Regularization
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- Learning Model Representations Using Publicly Available Model Hubs
- Bridging Perspectives: Foundation Model Guided BEV Maps for 3D Object Detection and Tracking
- MRI Brain Tumor Detection with Computer Vision
- Performance of heavy-flavour jet identification in Lorentz-boosted topologies in proton-proton collisions at √(s) = 13 TeV
- Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
- HccePose(BF): Predicting Front & Back Surfaces to Construct Ultra-Dense 2D-3D Correspondences for Pose Estimation
- YOLOv11-Litchi: Efficient Litchi Fruit Detection based on UAV-Captured Agricultural Imagery in Complex Orchard Environments
- LocalViT: Analyzing Locality in Vision Transformers
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Collaborative Learning of Semantic-Aware Feature Learning and Label Recovery for Multi-Label Image Recognition with Incomplete Labels
- DREAM: A Benchmark Study for Deepfake REalism AssessMent
- ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications
- Gradient-based Model Shortcut Detection for Time Series Classification
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration
- Mathematical Modeling and Convergence Analysis of Deep Neural Networks with Dense Layer Connectivities in Deep Learning
- Emergence of Spatial Representation in an Actor-Critic Agent with Hippocampus-Inspired Sequence Generator
- A Style-Based Profiling Framework for Quantifying the Synthetic-to-Real Gap in Autonomous Driving Datasets
- BiFuseNet: A Multimodal Network for Estimating Blood Alcohol Concentration via Bidirectional Hierarchical Fusion
- A Multicentric Dataset for Training and Benchmarking Breast Cancer Segmentation in H&E Slides
- TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
- A Methodology for Transparent Logic-Based Classification Using a Multi-Task Convolutional Tsetlin Machine
- Three-Stream Convolutional Networks for Video-based Person Re-Identification
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- Rethinking the shape convention of an MLP
- High-resolution velocity model estimation with neural operator and the time-shift imaging condition
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- Automated Segmentation of Brain Gray Matter Nuclei on Quantitative Susceptibility Mapping Using Deep Convolutional Neural Network
- Guided Proofreading of Automatic Segmentations for Connectomics
- GANs for Urban Design
- Stability of Transformers under Layer Normalization
- Holistic Order Prediction in Natural Scenes
- Cell Instance Segmentation: The Devil Is in the Boundaries
- Cross-Sensor Touch Generation
- HeSRN: Representation Learning On Heterogeneous Graphs via Slot-Aware Retentive Network
- Architecture Induces Structural Invariant Manifolds of Neural Network Training Dynamics
- Failure Prediction at Runtime for Generative Robot Policies
- A deep architecture for unified aesthetic prediction
- Uncovering Overconfident Failures in CXR Models via Augmentation-Sensitivity Risk Scoring
- A Novel Multi-branch ConvNeXt Architecture for Identifying Subtle Pathological Features in CT Scans
- MemLoss: Enhancing Adversarial Training with Recycling Adversarial Examples
- Provable Watermarking for Data Poisoning Attacks
- TARO: Toward Semantically Rich Open-World Object Detection
- FLOWING: Implicit Neural Flows for Structure-Preserving Morphing
- Denoised Diffusion for Object-Focused Image Augmentation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- SegTrans: Transferable Adversarial Examples for Segmentation Models
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- Training Feature Attribution for Vision Models
- Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning
- Cross-Receiver Generalization for RF Fingerprint Identification via Feature Disentanglement and Adversarial Training
- Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
- Non-Rigid Structure-from-Motion via Differential Geometry with Recoverable Conformal Scale
- VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
- Vision Language Models: A Survey of 26K Papers
- Towards Information-Optimized Multi-Agent Path Finding: A Hybrid Framework with Reduced Inter-Agent Information Sharing
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference
- SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense
- Theoretical guarantees for change localization using conformal p-values
- ReSplat: Learning Recurrent Gaussian Splats
- ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
- DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model
- AI-Driven Radiology Report Generation for Traumatic Brain Injuries
- Detecting and Recognizing Human-Object Interactions
- A Multimodal Depth-Aware Method For Embodied Reference Understanding
- Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
- A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
- Multi-Condition Conformal Selection
- A class-driven hierarchical ResNet for classification of multispectral remote sensing images
- DM1: MeanFlow with Dispersive Regularization for 1-Step Robotic Manipulation
- On the Occurence of Critical Learning Periods in Neural Networks
- FlowLensing: Simulating Gravitational Lensing with Flow Matching
- Learning to Navigate Socially Through Proactive Risk Perception
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- Self-Supervised Learning Strategies for a Platform to Test the Toxicity of New Chemicals and Materials
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Demystifying Deep Learning-based Brain Tumor Segmentation with 3D UNets and Explainable AI (XAI): A Comparative Analysis
- Deep Neural Networks Inspired by Differential Equations
- FedLAM: Low-latency Wireless Federated Learning via Layer-wise Adaptive Modulation
- Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method
- Sparse components distinguish visual pathways & their alignment to neural networks
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- Video Surveillance for Road Traffic Monitoring
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- RayFusion: Ray Fusion Enhanced Collaborative Visual Perception
- Mutual Learning for Hashing: Unlocking Strong Hash Functions from Weak Supervision
- MONKEY: Masking ON KEY-Value Activation Adapter for Personalization
- Curriculum Learning with Synthetic Data for Enhanced Pulmonary Nodule Detection in Chest Radiographs
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Weights initialization of neural networks for function approximation
- Geometry-aware Policy Imitation
- Robust Canonicalization through Bootstrapped Data Re-Alignment
- High-dimensional Analysis of Synthetic Data Selection
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- Rényi Sharpness: A Novel Sharpness that Strongly Correlates with Generalization
- Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
- Robust Source-Free Domain Adaptation for Medical Image Segmentation based on Curriculum Learning
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- EMPalm: Exfiltrating Palm Biometric Data via Electromagnetic Side-Channels
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
- SpecGuard: Spectral Projection-based Advanced Invisible Watermarking
- Evaluating Fundus-Specific Foundation Models for Diabetic Macular Edema Detection
- AppleCiDEr II: SpectraNet -- A Deep Learning Network for Spectroscopic Data
- Accelerating Inference for Multilayer Neural Networks with Quantum Computers
- Bridged Clustering: Semi-Supervised Sparse Bridging
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Sharpness-Aware Data Generation for Zero-shot Quantization
- Revealing the Temporally Stable Bimodal Energy Distribution of FRB 20121102A with a Tripled Burst Set from AI Detections
- Learning Global Representation from Queries for Vectorized HD Map Construction
- High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization
- Interactive reconstruction of Monte Carlo image sequences using a recurrent denoising autoencoder
- HARP-NeXt: High-Speed and Accurate Range-Point Fusion Network for 3D LiDAR Semantic Segmentation
- Multi-hop Deep Joint Source-Channel Coding with Deep Hash Distillation for Semantically Aligned Image Retrieval
- Online Generic Event Boundary Detection
- Consistent Assistant Domains Transformer for Source-free Domain Adaptation
- CardioRAG: A Retrieval-Augmented Generation Framework for Multimodal Chagas Disease Detection
- Clinical-grade AI model for molecular subtyping of endometrial cancer: a multi-center cohort study in China
- Recurrence-Complete Frame-based Action Models
- Synaptic facilitation and learning of multiplexed neural signals
- Getting the Numbers Right\unicodex2014Modelling Multi-Class Object Counting in Dense and Varied Scenes
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
- Adaptive Stain Normalization for Cross-Domain Medical Histology
- Mangrove3D: Terrestrial Laser Scanning Dataset for Coastal Mangrove Forests
- metabeta -- A fast neural model for Bayesian mixed-effects regression
- DPA-Net: A Dual-Path Attention Neural Network for Inferring Glycemic Control Metrics from Self-Monitored Blood Glucose Data
- The Framework That Survives Bad Models: Human-AI Collaboration For Clinical Trials
- Angular Constraint Embedding via SpherePair Loss for Constrained Clustering
- Improving Artifact Robustness for CT Deep Learning Models Without Labeled Artifact Images via Domain Adaptation
- Lung Infection Severity Prediction Using Transformers with Conditional TransMix Augmentation and Cross-Attention
- Learning from Limited Multi-Phase CT: Dual-Branch Prototype-Guided Framework for Early Recurrence Prediction in HCC
- How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Kaputt: A Large-Scale Dataset for Visual Defect Detection
- Enhancing Automotive Security with a Hybrid Approach towards Universal Intrusion Detection System
- A Novel Technique for Robust Training of Deep Networks With Multisource Weak Labeled Remote Sensing Data
- Empirical Comparison of Membership Inference Attacks in Deep Transfer Learning
- Data Factory with Minimal Human Effort Using VLMs
- Membership Inference Attacks on Tokenizers of Large Language Models
- Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
- VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
- TreeNet: Layered Decision Ensembles
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
- Critical attention scaling in long-context transformers
- Human Action Recognition from Point Clouds over Time
- Machine Learning Detection of Road Surface Conditions: A Generalizable Model using Traffic Cameras and Weather Data
- Emergent AI Surveillance: Overlearned Person Re-Identification and Its Mitigation in Law Enforcement Context
- Quantifying the Accuracy-Interpretability Trade-Off in Concept-Based Sidechannel Models
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- Deep Generative Model for Human Mobility Behavior
- Towards Data-Efficient Medical Imaging: A Generative and Semi-Supervised Framework
- Large Language Model-Based Uncertainty-Adjusted Label Extraction for Artificial Intelligence Model Development in Upper Extremity Radiography
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- Computing frustration and near-monotonicity in deep neural networks
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional Diffusion
- Unsupervised Active Learning via Natural Feature Progressive Framework
- ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures
- Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Forecasting-based Biomedical Time-series Data Synthesis for Open Data and Robust AI
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Improving Large-Scale Recommender Systems with Auxiliary Learning
- Knowledge Distillation Detection for Open-weights Models
- MultiModal Action Conditioned Video Generation
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Trade-off in Estimating the Number of Byzantine Clients in Federated Learning
- Fusion-Based Neural Generalization for Predicting Temperature Fields in Industrial PET Preform Heating
- Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction
- Bridging Text and Video Generation: A Survey
- Closed-Form Last Layer Optimization
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
- Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
- BLADE: Bias-Linked Adaptive DEbiasing
- Efficient Training of Spiking Neural Networks by Spike-aware Data Pruning
- Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
- Atomistic Machine Learning with Irreducible Cartesian Natural Tensors
- RAP: 3D Rasterization Augmented End-to-End Planning
- RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
- Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
- Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
- Quantization Range Estimation for Convolutional Neural Networks
- Deep Defense: Training DNNs with Improved Adversarial Robustness
- Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis
- FHEON: A Configurable Framework for Developing Privacy-Preserving Neural Networks Using Homomorphic Encryption
- VBM-NET: Visual Base Pose Learning for Mobile Manipulation using Equivariant TransporterNet and GNNs
- Privacy Enhancement in Over-the-Air Federated Learning via Adaptive Receive Scaling
- Zero-shot Recognition via Semantic Embeddings and Knowledge Graphs
- No bad local minima: Data independent training error guarantees for\n multilayer neural networks
- Deep Learning for Physical Processes: Incorporating Prior Scientific\n Knowledge
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- MambaCAFU: Hybrid Multi-Scale and Multi-Attention Model with Mamba-Based Fusion for Medical Image Segmentation
- A Benchmark Study of Deep Learning Methods for Multi-Label Pediatric Electrocardiogram-Based Cardiovascular Disease Classification
- Adaptively Sampling-Reusing-Mixing Decomposed Gradients to Speed Up Sharpness Aware Minimization
- LoRA Patching: Exposing the Fragility of Proactive Defenses against Deepfakes
- Click Here: Human-Localized Keypoints as Guidance for Viewpoint\n Estimation
- What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis
- MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- Achieving Universal Approximation and Universal Interpolation via Nonlinearity of Control Families
- On Provable Benefits of Muon in Federated Learning
- Personalized federated prototype learning in mixed heterogeneous data scenarios
- A Novel Cloud-Based Diffusion-Guided Hybrid Model for High-Accuracy Accident Detection in Intelligent Transportation Systems
- UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization
- AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networks
- CVSM: Contrastive Vocal Similarity Modeling
- Not every day is a sunny day: Synthetic cloud injection for deep land cover segmentation robustness evaluation across data sources
- Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
- Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
- A physics-informed neural network approach to the point defect model for electrochemical oxide film growth
- Representing Beauty: Towards a Participatory but Objective Latent Aesthetics
- Forensic Similarity for Speech Deepfakes
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- Distributed Low-Communication Training with Decoupled Momentum Optimization
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Net2Net: When Un-trained Meets Pre-trained Networks for Robust Real-World Denoising
- Accuracy Law for the Future of Deep Time Series Forecasting
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation Testing
- Image Enhancement Based on Pigment Representation
- A Statistical Method for Attack-Agnostic Adversarial Attack Detection with Compressive Sensing Comparison
- Uncertainty Quantification In Surface Landmines and UXO Classification Using MC Dropout
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- On residual network depth
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Self-Supervised Representation Learning as Mutual Information Maximization
- Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
- Towards Adversarial Training under Hyperspectral Images
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- TextCAM: Explaining Class Activation Map with Text
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
- Looking Alike From Far to Near: Enhancing Cross-Resolution Re-Identification via Feature Vector Panning
- Feature Identification for Hierarchical Contrastive Learning
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Uncertainty-Aware Concept Bottleneck Models with Enhanced Interpretability
- Deep learning motion correction of quantitative stress perfusion cardiovascular magnetic resonance
- From 2D to 3D, Deep Learning-based Shape Reconstruction in Magnetic Resonance Imaging: A Review
- Batch-CAM: Introduction to better reasoning in convolutional deep learning models
- U-DFA: A Unified DINOv2-Unet with Dual Fusion Attention for Multi-Dataset Medical Segmentation
- Assessing Foundation Models for Mold Colony Detection with Limited Training Data
- Sentry: Authenticating Machine Learning Artifacts on the Fly
- Vicinity-Guided Discriminative Latent Diffusion for Privacy-Preserving Domain Adaptation
- Diagnosing Shortcut-Induced Rigidity in Continual Learning: The Einstellung Rigidity Index (ERI)
- SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
- On-the-Fly Data Augmentation via Gradient-Guided and Sample-Aware Influence Estimation
- Automated Structured Radiology Report Generation with Rich Clinical Context
- Generative AI for subgrid turbulence in large-eddy simulations
- End-to-End Training of High-Dimensional Optimal Control with Implicit Hamiltonians via Jacobian-Free Backpropagation
- Cascaded Diffusion Framework for Probabilistic Coarse-to-Fine Hand Pose Estimation
- VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
- Teaching Categories to Human Learners with Visual Explanations
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- Robust Federated Inference
- F-scheduler: illuminating the free-lunch design space for fast sampling of diffusion models
- Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
- Cutting the Skip: Training Residual-Free Transformers
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- CODED-SMOOTHING: Coding Theory Helps Generalization
- Board Gender Diversity and Carbon Emissions Performance: Insights from Panel Regressions, Machine Learning and Explainable AI
- Creative synthesis of kinematic mechanisms
- Object-Centric Case-Based Reasoning via Argumentation
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Approximately Unimodal Likelihood Models for Ordinal Regression
- Uncertainty Quantification for Regression using Proper Scoring Rules
- Bayesian Influence Functions for Hessian-Free Data Attribution
- IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks
- FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
- Point2RBox-v3: Self-Bootstrapping from Point Annotations via Integrated Pseudo-Label Refinement and Utilization
- Cat: Post-Training Quantization Error Reduction via Cluster-based Affine Transformation
- Marginal Flow: a flexible and efficient framework for density estimation
- Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
- Architecturally Constrained Solutions to Ill-Conditioned Problems in QUBIC
- Machine Learning and Control: Foundations, Advances, and Perspectives
- Scaling Equilibrium Propagation to Deeper Neural Network Architectures
- The Impact of Scaling Training Data on Adversarial Robustness
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
- MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
- Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
- Enhancing Certifiable Semantic Robustness via Robust Pruning of Deep Neural Networks
- Using Images from a Video Game to Improve the Detection of Truck Axles
- Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
- Annotation-Efficient Active Test-Time Adaptation with Conformal Prediction
- Growing Winning Subnetworks, Not Pruning Them: A Paradigm for Density Discovery in Sparse Neural Networks
- DescribeEarth: Describe Anything for Remote Sensing Images
- Effective Model Pruning
- LAPIS: A Performance Portable, High Productivity Compiler Framework
- Ascent Fails to Forget
- Neural Hamilton--Jacobi Characteristic Flows for Optimal Transport
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- MetaChest: Generalized few-shot learning of pathologies from chest X-rays
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Online Mapping for Autonomous Driving: Addressing Sensor Generalization and Dynamic Map Updates in Campus Environments
- Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
- TDHook: A Lightweight Framework for Interpretability
- Multi-patch isogeometric neural solver for partial differential equations on computer-aided design domains
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
- LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model
- Physics-Informed Inductive Biases for Voltage Prediction in Distribution Grids
- Crop Spirals: Re-thinking the field layout for future robotic agriculture
- MANI-Pure: Magnitude-Adaptive Noise Injection for Adversarial Purification
- GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
- DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
- Vehicle Classification under Extreme Imbalance: A Comparative Study of Ensemble Learning and CNNs
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- ClustRecNet: A Novel End-to-End Deep Learning Framework for Clustering Algorithm Recommendation
- Vision Function Layer in Multimodal LLMs
- Quantifying Generalisation in Imitation Learning
- Spatial-Functional awareness Transformer-based graph archetype contrastive learning for Decoding Visual Neural Representations from EEG
- Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm
- Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
- VNODE: A Piecewise Continuous Volterra Neural Network
- RIFLE: Removal of Image Flicker-Banding via Latent Diffusion Enhancement
- Foggy Crowd Counting: Combining Physical Priors and KAN-Graph
- Enabling Physical AI through Biological Principles
- Performance-Efficiency Trade-off for Fashion Image Retrieval
- Hybrid Layer-Wise ANN-SNN With Surrogate Spike Encoding-Decoding Structure
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- DRIFT: Divergent Response in Filtered Transformations for Robust Adversarial Defense
- H+: An Efficient Similarity-Aware Aggregation for Byzantine Resilient Federated Learning
- Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
- Adaptive Source-Channel Coding for Multi-User Semantic and Data Communications
- Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
- High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
- Neural Visibility of Point Sets
- Analysis of Bias in Deep Learning Facial Beauty Regressors
- EYE-DEX: Eye Disease Detection and EXplanation System
- BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression
- GLASS Flows: Transition Sampling for Alignment of Flow and Diffusion Models
- BERT Bi-modal self-supervised learning for crop classification using Sentinel-2 and Planetscope
- Improving Simple Models with Confidence Profiles
- Foveated Retinotopy Improves Classification and Localization in Convolutional Neural Networks
- Characterizing and Improving Stability in Neural Style Transfer
- A Pixel-Based Framework for Data-Driven Clothing
- Relation Networks for Optic Disc and Fovea Localization in Retinal Images
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
- Deep Learning Reconstruction on Crab LST-1 Data
- A Systematic Review of Digital Twin-Driven Predictive Maintenance in Industrial Engineering: Taxonomy, Architectural Elements, and Future Research Directions
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- On The Variability of Concept Activation Vectors
- End-to-end Topographic Auditory Models Replicate Signatures of Human Auditory Cortex
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
- Enhancing hyperspectral image prediction with contrastive learning in low-label regimes
- Learning-Based Testing for Deep Learning: Enhancing Model Robustness with Adversarial Input Prioritization
- Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- Gradient Flow Convergence Guarantee for General Neural Network Architectures
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
- FairViT-GAN: A Hybrid Vision Transformer with Adversarial Debiasing for Fair and Explainable Facial Beauty Prediction
- CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- Transparent Visual Reasoning via Object-Centric Agent Collaboration
- PVTAdpNet: Polyp Segmentation using Pyramid vision transformer with a novel Adapter block
- Merge Now, Regret Later: The Hidden Cost of Model Merging is Adversarial Transferability
- Decentralized Dynamic Cooperation of Personalized Models for Federated Continual Learning
- Virtual Nodes based Heterogeneous Graph Convolutional Neural Network for Efficient Long-Range Information Aggregation
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- From Past To Path: Masked History Learning for Next-Item Prediction in Generative Recommendation
- Evaluating the Impact of Radiographic Noise on Chest X-ray Semantic Segmentation and Disease Classification Using a Scalable Noise Injection Framework
- A Recall-First CNN for Sleep Apnea Screening from Snoring Audio
- EfficientMIL: Efficient Linear-Complexity MIL Method for WSI Classification
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Deep Taxonomic Networks for Unsupervised Hierarchical Prototype Discovery
- Multi-Level Heterogeneous Knowledge Transfer Network on Forward Scattering Center Model for Limited Samples SAR ATR
- Channel, Trend and Periodic-Wise Representation Learning for Multivariate Long-term Time Series Forecasting
- Disentanglement of Variations with Multimodal Generative Modeling
- HunyuanImage 3.0 Technical Report
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
- Imaging-Based Mortality Prediction in Patients with Systemic Sclerosis
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- S3F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- Neural Language Priors
- Generative Modeling of Shape-Dependent Self-Contact Human Poses
- UniPose: Unified Cross-modality Pose Prior Propagation towards RGB-D data for Weakly Supervised 3D Human Pose Estimation
- Graph Your Own Prompt
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Adversarial Scheduling
- Seeing Through the Blur: Unlocking Defocus Maps for Deepfake Detection
- Learning Regional Monsoon Patterns with a Multimodal Attention U-Net
- PAL-AI reveals genetic determinants that control poly(A)-tail length during oocyte maturation, with relevance to human fertility
- More Data or Better Algorithms: Latent Diffusion Augmentation for Deep Imbalanced Regression
- Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
- UltraUNet: Real-Time Ultrasound Tongue Segmentation for Diverse Linguistic and Imaging Conditions
- Leave No Observation Behind: Real-time Correction for VLA Action Chunks
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- Effective Quantization of Muon Optimizer States
- HTMA-Net: Towards Multiplication-Avoiding Neural Networks via Hadamard Transform and In-Memory Computing
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- Dynamics of Learning: Generative Schedules from Latent ODEs
- Activation Matching for Explanation Generation
- DPFNAS: Differential Privacy-Enhanced Federated Neural Architecture Search for 6G Edge Intelligence
- Learning without Global Backpropagation via Synergistic Information Distillation
- ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction
- Single-Shot Refinement Neural Network for Object Detection
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer
- URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization
- Beyond Aggregation: Guiding Clients in Heterogeneous Federated Learning
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- Beyond Gaussian Initializations: Signal Preserving Weight Initialization for Odd-Sigmoid Activations
- Training Deep Normalization-Free Spiking Neural Networks with Lateral Inhibition
- Multi-Miner: Object-Adaptive Region Mining for Weakly-Supervised Semantic Segmentation
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- Targeted perturbations reveal brain-like local coding axes in robustified, but not standard, ANN-based brain models
- Hemorica: A Comprehensive CT Scan Dataset for Automated Brain Hemorrhage Classification, Segmentation, and Detection
- T-TAMER: Provably Taming Trade-offs in ML Serving
- MindCraft: How Concept Trees Take Shape In Deep Models
- FedCF: Fair Federated Conformal Prediction
- Convolutional Set Transformer
- AntiFLipper: A Secure and Efficient Defense Against Label-Flipping Attacks in Federated Learning
- ControlEvents: Controllable Synthesis of Event Camera Datawith Foundational Prior from Image Diffusion Models
- Seeing Isn't Believing: Context-Aware Adversarial Patch Synthesis via Conditional GAN
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- Toward a Physics of Deep Learning and Brains
- Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- IONext: Unlocking the Next Era of Inertial Odometry
- Transformers for single-cell RNA sequencing: a survey
- Category Discovery: An Open-World Perspective
- γ-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition
- U-MAN: U-Net with Multi-scale Adaptive KAN Network for Medical Image Segmentation
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
- A model of errors in transformers
- GPT-4 for Occlusion Order Recovery
- Multidimensional Uncertainty Quantification via Optimal Transport
- Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
- Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- Clinical Uncertainty Impacts Machine Learning Evaluations
- A Law of Data Reconstruction for Random Features (and Beyond)
- Red Teaming Quantum-Resistant Cryptographic Standards: A Penetration Testing Framework Integrating AI and Quantum Security
- Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
- Multilingual Vision-Language Models, A Survey
- Concept activation vectors: a unifying view and adversarial attacks
- Self-driving cars: Are we there yet?
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- Concept-SAE: Active Causal Probing of Visual Model Behavior
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- Closing the Oracle Gap: Increment Vector Transformation for Class Incremental Learning
- LG-CD: Enhancing Language-Guided Change Detection through SAM2 Adaptation
- Zubov-Net: Adaptive Stability for Neural ODEs Reconciling Accuracy with Robustness
- Enhancing Low-Rank Adaptation with Structured Nonlinear Transformations
- Deepfakes: we need to re-think the concept of "real" images
- SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
- Sharpness-Aware Minimization Can Hallucinate Minimizers
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- SubZeroCore: A Submodular Approach with Zero Training for Coreset Selection
- HyperCore: Coreset Selection under Noise via Hypersphere Models
- On the Status of Foundation Models for SAR Imagery
- Machine Learning for Quantum State Tomography: Robust Covariance Matrix Estimation for Squeezed Vacuum States with Thermal Noise
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- Light Differentiable Logic Gate Networks
- Learning Admissible Heuristics for A*: Theory and Practice
- POEM: Explore Unexplored Reliable Samples to Enhance Test-Time Adaptation
- Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
- Task-Agnostic Federated Continual Learning via Replay-Free Gradient Projection
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
- Functional Encryption in Secure Neural Network Training: Data Leakage and Practical Mitigations
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- VISION: Prompting Ocean Vertical Velocity Reconstruction from Incomplete Observations
- Model reduction of parametric ordinary differential equations via autoencoders: structure-preserving latent dynamics and convergence analysis
- Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability
- MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
- Learning Conformal Explainers for Image Classifiers
- Differential-Integral Neural Operator for Long-Term Turbulence Forecasting
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
- LayerNorm Induces Recency Bias in Transformer Decoders
- Shapley Features for Robust Signal Prediction in Tactile Internet
- SiNGER: A Clearer Voice Distills Vision Transformers Further
- Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement
- Plant identification based on noisy web data: the amazing performance of deep learning (LifeCLEF 2017)
- Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- FERD: Fairness-Enhanced Data-Free Robustness Distillation
- LiLAW: Lightweight Learnable Adaptive Weighting to Learn Sample Difficulty & Improve Noisy Training
- Dual-supervised Asymmetric Co-training for Semi-supervised Medical Domain Generalization
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Enhancing Cross-View Geo-Localization Generalization via Global-Local Consistency and Geometric Equivariance
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
- DATS: Distance-Aware Temperature Scaling for Calibrated Class-Incremental Learning
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Normalizing Flows are Capable Models for Bi-manual Visuomotor Policy
- Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
- SemSight: Probabilistic Bird's-Eye-View Prediction of Multi-Level Scene Semantics for Navigation
- Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
- How fine can fine-tuning be? Learning efficient language models
- Plant identification in an open-world (LifeCLEF 2016)
- Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition
- Neural Integrated Sensing and Communication for the MIMO-OFDM Downlink
- The Unwinnable Arms Race of AI Image Detection
- An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image Clustering
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
- Iterative Visual Reasoning Beyond Convolutions
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
- PerFace: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
- Velocity model building from seismic images using a Convolutional Neural Operator
- Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
- Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
- Generative Model Inversion Through the Lens of the Manifold Hypothesis
- Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- mmHSense: Multi-Modal and Distributed mmWave ISAC Datasets for Human Sensing
- Embodied AI: From LLMs to World Models
- Does the Manipulation Process Matter? RITA: Reasoning Composite Image Manipulations via Reversely-Ordered Incremental-Transition Autoregression
- MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
- Anomaly Detection by Clustering DINO Embeddings using a Dirichlet Process Mixture
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach
- When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- ExpFace: Exponential Angular Margin Loss for Deep Face Recognition
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Adaptive Model Ensemble for Continual Learning
- Sobolev acceleration for neural networks
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- MoTiC: Momentum Tightness and Contrast for Few-Shot Class-Incremental Learning
- IBiT: Utilizing Inductive Biases to Create a More Data Efficient Attention Mechanism
- Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
- Unsupervised Domain Adaptation for Binary Classification with an Unobservable Source Subpopulation
- Efficiently Attacking Memorization Scores
- Are Foundation Models Ready for Industrial Defect Recognition? A Reality Check on Real-World Data
- ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
- Region-of-Interest Augmentation for Mammography Classification under Patient-Level Cross-Validation
- Achieving Fair Skin Lesion Detection through Skin Tone Normalization and Channel Pruning
- The Impact of 2D Segmentation Backbones on Point Cloud Predictions Using 4D Radar
- High Clockrate Free-space Optical In-Memory Computing
- Adaptive von Mises-Fisher Likelihood Loss for Supervised Deep Time Series Hashing
- CURE: Centroid-guided Unsupervised Representation Erasure for Facial Recognition Systems
- Confidence Calibration in Large Language Model-Based Entity Matching
- Benign on Label, Malicious by Design: Clean-Label Dormant-to-Activated Backdoor via Machine Unlearning with Removable Camouflage
- ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
- Don't Trust the AI Ecosystem: Analyzing Privacy Leakage in Compromised Open-Source Components
- CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- Simplifying Neural Networks During Training
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- Looped Transformers with Source-Centered State Evolution
- Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
- MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging
- Efficient Training of Convolutional Neural Nets on Large Distributed\n Systems
- Deep learning-based hierarchical insect classification using camera trap imagery
- Meteosat Third Generation imagery improves CNN-based SSI retrieval
- ATF6 activation alters colonic lipid metabolism causing tumour-associated microbial adaptation
- Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging
- Collaborative feature aggregation for face super-resolution and robust re-identification
- What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- CompoVista: A Composition-Graph-Based Visual Analytics System for Compositional Analysis of Traditional Chinese Paintings
- Lets keep it simple, Using simple architectures to outperform deeper and\n more complex architectures
- Multi-Head Attention Residuals
- Shortcut to Nowhere: Demystifying Deep Spurious Regression
- Toward a More Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data
- Successive convex optimization for transformer encoder model predictive control
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Deep learning to achieve clinically applicable segmentation of head and neck anatomy for radiotherapy
- A transformer-based multi-task deep learning model for urban livability evaluation by fusing remote sensing and textual geospatial data
- Learning Convolutional Networks for Content-weighted Image Compression
- Orexin population activity precisely reflects net body movement across behavioral and metabolic states
- Temporal Straightening for Latent Planning
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural\n Networks
- Face-like holistic Processing in Non-Face Stimuli
- Topology Aware Neural Interpolation of Scalar Fields
- DL-QC-fNIRS: a deep learning tool for automated quality control in functional near-infrared spectroscopy signals
- Stiffness: A New Perspective on Generalization in Neural Networks
- Effect of Demographic Bias on Skin Lesion Classification
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
- Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
- Vocoder-Projected Feature Discriminator
- Face Detection through Scale-Friendly Deep Convolutional Networks
- Outplaying elite table tennis players with an autonomous robot
- XNOR-Net: ImageNet Classification Using Binary Convolutional Neural\n Networks
- Stuck on Suggestions: Automation Bias, the Anchoring Effect, and the Factors That Shape Them in Computational Pathology
- GAN-based Virtual Re-Staining: A Promising Solution for Whole Slide Image Analysis
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- Multi-style Generative Network for Real-time Transfer
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Learning to Segment Every Thing
- Art of singular vectors and universal adversarial perturbations
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- DeepDeblur: Fast one-step blurry face images restoration
- Train, Diagnose and Fix: Interpretable Approach for Fine-grained Action Recognition
- Artificial intelligence-enabled electrocardiography from scientific research to clinical application
- Box-Level Class-Balanced Sampling for Active Object Detection
- LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
- Fast Neural Architecture Construction using EnvelopeNets
- Assessing the Alignment of Popular CNNs to the Brain for Valence Appraisal
- Self-evolved Imitation Learning in Simulated World
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Defending against Stegomalware in Deep Neural Networks with Permutation Symmetry
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- Probabilistic Runtime Verification, Evaluation and Risk Assessment of Visual Deep Learning Systems
- World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- A Kernel Space-based Multidimensional Sparse Model for Dynamic PET Image Denoising
- Rethinking Federated Learning Over the Air: The Blessing of Scaling Up
- Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- Deep Learning At Scale and At Ease
- Source-Free Domain Adaptive Semantic Segmentation of Remote Sensing Images with Diffusion-Guided Label Enrichment
- MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
- Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation
- MER-Inspector: Assessing model extraction risks from an attack-agnostic perspective
- SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack against No-Reference Image Quality Assessment Models
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
- Towards Application Aligned Synthetic Surgical Image Synthesis
- MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
- Chiplet-Based RISC-V SoC with Modular AI Acceleration
- Learning to Condition: A Neural Heuristic for Scalable MPE Inference
- Recognition Of Surface Defects On Steel Sheet Using Transfer Learning
- Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem
- Lipschitz-Based Robustness Certification for Recurrent Neural Networks via Convex Relaxation
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient Noise
- Why Do Deep Neural Networks Still Not Recognize These Images?: A Qualitative Analysis on Failure Cases of ImageNet Classification
- Degradation-Aware All-in-One Image Restoration via Latent Prior Encoding
- Automated Labeling of Intracranial Arteries with Uncertainty Quantification Using Deep Learning
- Accelerated characterization of two-level systems in superconducting qubits via machine learning
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
- Tailored Transformation Invariance for Industrial Anomaly Detection
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- A2M2-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition
- PRNU-Bench: A Novel Benchmark and Model for PRNU-Based Camera Identification
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Multimodal Radio and Vision Fusion for Robust Localization in Urban V2I Communications
- MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
- Training-Free Label Space Alignment for Universal Domain Adaptation
- Exploring Machine Learning Models for Physical Dose Calculation in Carbon Ion Therapy Using Heterogeneous Imaging Data -- A Proof of Concept Study
- Optimizing Split Federated Learning with Unstable Client Participation
- Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
- Attention-based Mixture of Experts for Robust Speech Deepfake Detection
- Faster CryptoNets: Leveraging Sparsity for Real-World Encrypted Inference
- Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion
- Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge
- The evolution of neural network-based chart patterns
- Domain Adaptive Object Detection for Space Applications with Real-Time Constraints
- The unreasonable effectiveness of the forget gate
- Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
- Convolutional Neural Network Optimization for Beehive Classification Using Bioacoustic Signals
- Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- Intra-Cluster Mixup: An Effective Data Augmentation Technique for Complementary-Label Learning
- Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
- Computational Scaffolding of Composition, Value, and Color for Disciplined Drawing
- SynergyNet: Fusing Generative Priors and State-Space Models for Facial Beauty Prediction
- Flow-Induced Diagonal Gaussian Processes
- CoBEVMoE: Heterogeneity-aware Feature Fusion with Dynamic Mixture-of-Experts for Collaborative Perception
- MARS: A Malignity-Aware Backdoor Defense in Federated Learning
- Data-Driven Reconstruction of Significant Wave Heights from Sparse Observations
- SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
- A review of Recent Techniques for Person Re-Identification
- SVDNet for Pedestrian Retrieval
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
- DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality Assessment
- LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object Detection
- 3D Reconstruction in Canonical Co-ordinate Space from Arbitrarily Oriented 2D Images
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- PRISM: Precision-Recall Informed Data-Free Knowledge Distillation via Generative Diffusion
- SegFlow: Joint Learning for Video Object Segmentation and Optical Flow
- Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
- Graph Coloring for Multi-Task Learning
- Long-Tailed Out-of-Distribution Detection with Refined Separate Class Learning
- FedEL: Federated Elastic Learning for Heterogeneous Devices
- ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis
- Automatic Classification of Magnetic Chirality of Solar Filaments from H-Alpha Observations
- Optimized Learned Image Compression for Facial Expression Recognition
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
- Looking in the mirror: A faithful counterfactual explanation method for interpreting deep image classification models
- When Confidence Fails: Revisiting Pseudo-Label Selection in Semi-supervised Semantic Segmentation
- Towards a Transparent and Interpretable AI Model for Medical Image Classifications
- IPF-RDA: An Information-Preserving Framework for Robust Data Augmentation
- Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
- Learned Digital Codes for Over-the-Air Federated Learning
- Single Path One-Shot Neural Architecture Search with Uniform Sampling
- ST-GS: Vision-Based 3D Semantic Occupancy Prediction with Spatial-Temporal Gaussian Splatting
- No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- TranTac: Leveraging Transient Tactile Signals for Contact-Rich Robotic Manipulation
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
- Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- DISCO: Disentangled Communication Steering for Large Language Models
- SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
- PM25Vision: A Large-Scale Benchmark Dataset for Visual Estimation of Air Quality
- CAMBench-QR : A Structure-Aware Benchmark for Post-Hoc Explanations with QR Understanding
- Factorizing Diffusion Policies for Observation Modality Prioritization
- CoUn: Empowering Machine Unlearning via Contrastive Learning
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- Learning Safety for Obstacle Avoidance via Control Barrier Functions
- DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
- Detecting Photoshopped Faces by Scripting Photoshop
- A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
- PAN: Pillars-Attention-Based Network for 3D Object Detection
- Enriched Feature Representation and Motion Prediction Module for MOSEv2 Track of 7th LSVOS Challenge: 3rd Place Solution
- RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
- Layout Stroke Imitation: A Layout Guided Handwriting Stroke Generation for Style Imitation with Diffusion Model
- Toward Efficient Influence Function: Dropout as a Compression Tool
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Contrastive Learning with Spectrum Information Augmentation in Abnormal Sound Detection
- Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
- MEC-Quant: Maximum Entropy Coding for Extremely Low Bit Quantization-Aware Training
- Saccadic Vision for Fine-Grained Visual Classification
- Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
- Distributed Multi-Task Learning for Joint Wireless Signal Enhancement and Recognition
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- Filter-and-Attend: Wireless Channel Foundation Model with Noise-Plus-Interference Suppression Structure
- MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
- Boosting Active Learning with Knowledge Transfer
- On the Convergence of Muon and Beyond
- SolarCrossFormer: Improving day-ahead Solar Irradiance Forecasting by Integrating Satellite Imagery and Ground Sensors
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
- MoE-CE: Enhancing Generalization for Deep Learning based Channel Estimation via a Mixture-of-Experts Framework
- How many classes do we need to see for novel class discovery?
- Efficient Multimodal Dataset Distillation via Generative Models
- CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan Reconstruction
- Causal Fingerprints of AI Generative Models
- Generating Part-Based Global Explanations Via Correspondence
- Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
- CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization
- TITAN: A Trajectory-Informed Technique for Adaptive Parameter Freezing in Large-Scale VQE
- Who to Trust? Aggregating Client Knowledge in Logit-Based Federated Learning
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Loss Barcode: A Topological Measure of Escapability in Loss Landscapes
- Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Physics-Informed GCN-LSTM Framework for Long-Term Forecasting of 2D and 3D Microstructure Evolution
- UCorr: Wire Detection and Depth Estimation for Autonomous Drones
- Brain-HGCN: A Hyperbolic Graph Convolutional Network for Brain Functional Network Analysis
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- FlowCast-ODE: Continuous Hourly Weather Forecasting with Dynamic Flow Matching and ODE Solver
- Designing Latent Safety Filters using Pre-Trained Vision Models
- Pre-training Autoencoder for Acoustic Event Classification via Blinky
- Multiscale super-resolution reconstruction of fluid flows with deep neural networks
- SpeechMLC: Speech Multi-label Classification
- How Does Instrumental Music Help SingFake Detection?
- Emotion-Aware Speech Generation with Character-Specific Voices for Comics
- An Attention Free Transformer
- Efficient 3D Perception on Embedded Systems via Interpolation-Free Tri-Plane Lifting and Volume Fusion
- CUFG: Curriculum Unlearning Guided by the Forgetting Gradient
- Enhancing Feature Fusion of U-like Networks with Dynamic Skip Connections
- Towards Privacy-Preserving and Heterogeneity-aware Split Federated Learning via Probabilistic Masking
- Domain Adaptation for Ulcerative Colitis Severity Estimation Using Patient-Level Diagnoses
- DiffVL: Diffusion-Based Visual Localization on 2D Maps via BEV-Conditioned GPS Denoising
- TriSPrompt: A Hierarchical Soft Prompt Model for Multimodal Rumor Detection with Incomplete Modalities
- Learning to Pick: A Visuomotor Policy for Clustered Strawberry Picking
- Object Recognition and Force Estimation with the GelSight Baby Fin Ray
- Survey on Deep Learning-based Kuzushiji Recognition
- BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
- URNet: Uncertainty-aware Refinement Network for Event-based Stereo Depth Estimation
- Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
- Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers
- VRScout: Towards Real-Time, Autonomous Testing of Virtual Reality Games
- Semi-Supervised 3D Medical Segmentation from 2D Natural Images Pretrained Model
- FedKLPR: Personalized Federated Learning for Person Re-Identification with Adaptive Pruning
- Spectral/Spatial Tensor Atomic Cluster Expansion with Universal Embeddings in Cartesian Space
- Incorporating Visual Cortical Lateral Connection Properties into CNN: Recurrent Activation and Excitatory-Inhibitory Separation
- Interleaved Group Convolutions for Deep Neural Networks
- Statistically Motivated Second Order Pooling
- Generative Adversarial Residual Pairwise Networks for One Shot Learning
- Class-Invariant Test-Time Augmentation for Domain Generalization
- BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
- An Exploratory Study on Abstract Images and Visual Representations Learned from Them
- Data Leakage in Visual Datasets
- Neural Proteomics Fields for Super-resolved Spatial Proteomics Prediction
- FedERL: Federated Efficient and Robust Learning for Common Corruptions
- VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
- SHaRe-RL: Structured, Interactive Reinforcement Learning for Contact-Rich Industrial Assembly Tasks
- Noise-Level Diffusion Guidance: Well Begun is Half Done
- Masked Feature Modeling Enhances Adaptive Segmentation
- Efficient Quantization-Aware Neural Receivers: Beyond Post-Training Quantization
- Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
- Self Identity Mapping
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
- ParaAegis: Parallel Protection for Flexible Privacy-preserved Federated Learning
- Deep Lookup Network
- Improved Segmentation of Polyps and Visual Explainability Analysis
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
- WatchAnxiety: A Transfer Learning Approach for State Anxiety Prediction from Smartwatch Data
- UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
- Integrated diffractive full-Stokes spectro-polarimetric imaging
- FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
- Federated Learning for Deforestation Detection: A Distributed Approach with Satellite Imagery
- EvHand-FPV: Efficient Event-Based 3D Hand Tracking from First-Person View
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Defending Deepfake via Texture Feature Perturbation
- Behavior Foundation Model for Humanoid Robots
- FlowDrive: Energy Flow Field for End-to-End Autonomous Driving
- Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
- Learning Progression-Guided AI Evaluation of Scientific Models To Support Diverse Multi-Modal Understanding in NGSS Classroom
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
- Image Realness Assessment and Localization with Multimodal Features
- ResidualViT for Efficient Temporally Dense Video Encoding
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- Cosine Normalization: Using Cosine Similarity Instead of Dot Product in Neural Networks
- UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
- CvT: Introducing Convolutions to Vision Transformers
- Bridging Performance Gaps for ECG Foundation Models: A Post-Training Strategy
- Reversible Deep Equilibrium Models
- EByFTVeS: Efficient Byzantine Fault Tolerant-based Verifiable Secret-sharing in Distributed Privacy-preserving Machine Learning
- HQCNN: A Hybrid Quantum-Classical Neural Network for Medical Image Classification
- Performance is not All You Need: Sustainability Considerations for Algorithms
- AdaGAT: Adaptive Guidance Adversarial Training for the Robustness of Deep Neural Networks
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- Modelling and analysis of the 8 filters from the "master key filters hypothesis" for depthwise-separable deep networks in relation to idealized receptive fields based on scale-space theory
- Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
- CIARD: Cyclic Iterative Adversarial Robustness Distillation
- Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew
- High-Energy Concentration for Federated Learning in Frequency Domain
- 4DRadar-GS: Self-Supervised Dynamic Driving Scene Reconstruction with 4D Radar
- FunKAN: Functional Kolmogorov-Arnold Network for Medical Image Enhancement and Segmentation
- SegStereo: Exploiting Semantic Information for Disparity Estimation
- Curvature Learning for Generalization of Hyperbolic Neural Networks
- Detecting Cyberattacks in Industrial Control Systems Using Convolutional Neural Networks
- SAUNet: Shape Attentive U-Net for Interpretable Medical Image Segmentation
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- Multi-scale Scanning Network for Machine Anomalous Sound Detection
- FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
- Conflect: Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications
- VQT-Light:Lightweight HDR Illumination Map Prediction with Richer Texture.pdf
- TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images
- T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
- Introduce the Result Into Self-Attention
- End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
- Exploring Training Data Attribution under Limited Access Constraints
- Contextualized Representation Learning for Effective Human-Object Interaction Detection
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- SRN: Side-output Residual Network for Object Symmetry Detection in the Wild
- AON: Towards Arbitrarily-Oriented Text Recognition
- Rethinking Radiology: An Analysis of Different Approaches to BraTS
- A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
- Beyond Gaze Overlap: Analyzing Joint Visual Attention Dynamics Using Egocentric Data
- GhostNetV3-Small: A Tailored Architecture and Comparative Study of Distillation Strategies for Tiny Images
- Uncertainty-Aware Hourly Air Temperature Mapping at 2 km Resolution via Physics-Guided Deep Learning
- U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
- FedDAF: Federated Domain Adaptation Using Model Functional Distance
- Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
- Automatically detecting pig position and posture by 2D camera imaging and deep learning
- Reconciling Communication Compression and Byzantine-Robustness in Distributed Learning
- LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- A Geometric Graph-Based Deep Learning Model for Drug-Target Affinity Prediction
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Take it in your stride: Do we need striding in CNNs?
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- PD-Loss: Proxy-Decidability for Efficient Metric Learning
- MSMA: Multi-Scale Feature Fusion For Multi-Attribute 3D Face Reconstruction From Unconstrained Images
- DRAG: Data Reconstruction Attack using Guided Diffusion
- The Quest for Universal Master Key Filters in DS-CNNs
- DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition
- Determining the boundary of dynamical chaos in the generalized Chirikov map via machine learning
- Optimizing Class Distributions for Bias-Aware Multi-Class Learning
- A Fine-Grained 3D Radio Map Construction Paradigm with Ultra-Low Sampling Rates by Large Generative Models
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
- DARD: Dice Adversarial Robustness Distillation against Adversarial Attacks
- Synthetic vs. Real Training Data for Visual Navigation
- Adaptive Spatial Goodness Encoding: Advancing and Scaling Forward-Forward Learning Without Backpropagation
- Error Control and Loss Functions for the Deep Learning Inversion of Borehole Resistivity Measurements
- Advanced Layout Analysis Models for Docling
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- EMeRALDS: Electronic Medical Record Driven Automated Lung Nodule Detection and Classification in Thoracic CT Images
- Reduced Order Modeling of Energetic Materials Using Physics-Aware Recurrent Convolutional Neural Networks in a Latent Space (LatentPARC)
- Two-Stage Decoupling Framework for Variable-Length Glaucoma Prognosis
- Reconstruction of IACT events using deep learning techniques with CTLearn
- NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks
- Logit Mixture Outlier Exposure for Fine-grained Out-of-Distribution Detection
- MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation
- Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation
- Promoting Shape Bias in CNNs: Frequency-Based and Contrastive Regularization for Corruption Robustness
- Neural networks in the search for fast radio bursts with RATAN-600
- FEWT: Improving Humanoid Robot Perception with Frequency-Enhanced Wavelet-based Transformers
- SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
- Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain Adaptation
- End-to-End Visual Autonomous Parking via Control-Aided Attention
- Hybrid Quantum Neural Networks for Efficient Protein-Ligand Binding Affinity Prediction
- SPHERE: Semantic-PHysical Engaged REpresentation for 3D Semantic Scene Completion
- Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
- CIFNet: An Analytic Neural Learning Framework for Efficient and Calibrated Class-Incremental Learning
- Make Identity Unextractable yet Perceptible: Synthesis-Based Privacy Protection for Subject Faces in Photos
- UnLoc: Leveraging Depth Uncertainties for Floorplan Localization
- Policy-Driven Transfer Learning in Resource-Limited Animal Monitoring
- Point-Plane Projections for Accurate LiDAR Semantic Segmentation in Small Data Scenarios
- A Modern Look at Simplicity Bias in Image Classification Tasks
- Group Evidence Matters: Tiling-based Semantic Gating for Dense Object Detection
- ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation
- Building a General SimCLR Self-Supervised Foundation Model Across Neurological Diseases to Advance 3D Brain MRI Diagnoses
- SSL-AD: Spatiotemporal Self-Supervised Learning for Generalizability and Adaptability Across Alzheimer's Prediction Tasks and Datasets
- EfficientNet-Based Multi-Class Detection of Real, Deepfake, and Plastic Surgery Faces
- A Symmetry-Integrated Approach to Surface Code Decoding
- FedBiF: Communication-Efficient Federated Learning via Bits Freezing
- Towards Understanding the Data Dependency of Mixup-style Training
- Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- Preserving Domain Generalization in Fine-Tuning via Joint Parameter Selection
- Balanced Sharpness-Aware Minimization for Imbalanced Regression
- Local Information Matters: A Rethink of Crowd Counting
- Neural Scaling Laws for Deep Regression
- Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
- Service Function Chaining Architecture for Multi-hop Split Inference and Learning
- FedRP: A Communication-Efficient Approach for Differentially Private Federated Learning Using Random Projection
- The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagram
- Hierarchical MLANet: Multi-level Attention for 3D Face Reconstruction From Single Images
- Disentangling Polysemantic Neurons with a Null-Calibrated Polysemanticity Index and Causal Patch Interventions
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
- Why and How Auxiliary Tasks Improve JEPA Representations
- TUNI: Unifying Pre-training and Fine-tuning with Modality-Aware Mutual Learning and Rectification for RGB-T Semantic Segmentation
- Imagined Autocurricula
- HGEN: Heterogeneous Graph Ensemble Networks
- DVQA: Understanding Data Visualizations via Question Answering
- Deep Learning with Apache SystemML
- From the Gradient-Step Denoiser to the Proximal Denoiser and their associated convergent Plug-and-Play algorithms
- ZORRO: Zero-Knowledge Robustness and Privacy for Split Learning (Full Version)
- Uncertainty Propagation Networks for Neural Ordinary Differential Equations
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- Invisible Attributes, Visible Biases: Exploring Demographic Shortcuts in MRI-based Alzheimer's Disease Classification
- Cough Classification using Few-Shot Learning
- Composable Score-based Graph Diffusion Model for Multi-Conditional Molecular Generation
- MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
- CoAtNeXt:An Attention-Enhanced ConvNeXtV2-Transformer Hybrid Model for Gastric Tissue Classification
- Modular, On-Site Solutions with Lightweight Anomaly Detection for Sustainable Nutrient Management in Agriculture
- Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection
- Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing
- OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
- FPI-Det: a face--phone Interaction Dataset for phone-use detection and understanding
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Scalable extensions to given-data Sobol' index estimators
- Improvement of Human-Object Interaction Action Recognition Using Scene Information and Multi-Task Learning Approach
- Tri-Accel: Curvature-Aware Precision-Adaptive and Memory-Elastic Optimization for Efficient GPU Usage
- AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
- MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
- MSDANet: A Multiscale Dual-Channel Spatial Attention Network with Depthwise Separable Convolution for Hyperspectral Image Classification
- OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
- Patch-based Automatic Rosacea Detection Using the ResNet Deep Learning Framework
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- Images in Motion?: A First Look into Video Leakage in Collaborative Deep Learning
- Privacy-Preserving Automated Rosacea Detection Based on Medically Inspired Region of Interest Selection
- Model compression via distillation and quantization
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Similarity-based Outlier Detection for Noisy Object Re-Identification Using Beta Mixtures
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
- ArgoTweak: Towards Self-Updating HD Maps through Structured Priors
- Compressing CNN models for resource-constrained systems by channel and layer pruning
- RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts
- CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
- Audio Deepfake Verification
- Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
- Rethinking the Backbone in Class Imbalanced Federated Source Free Domain Adaptation: The Utility of Vision Foundation Models
- Chordless cycle filtrations for dimensionality detection in complex networks via topological data analysis
- Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Advancing Few-Shot Pediatric Arrhythmia Classification with a Novel Contrastive Loss and Multimodal Learning
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
- Value bounds and Convergence Analysis for Averages of LRP attributions
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- Lightweight Deep Unfolding Networks with Enhanced Robustness for Infrared Small Target Detection
- CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded Scenes
- Sparsity in Deep Neural Networks - An Empirical Investigation with\n TensorQuant
- Rollout-LaSDI: Enhancing the long-term accuracy of Latent Space Dynamics
- ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis
- Localized PCA-Net Neural Operators for Scalable Solution Reconstruction of Elliptic PDEs
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- Hammer and Anvil: A Principled Defense Against Backdoors in Federated Learning
- How Far Are We from True Unlearnability?
- Neuromorphic Simulation of Drosophila Melanogaster Brain Connectome on Loihi 2
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Feature Space Analysis by Guided Diffusion Model
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
- Object-level Correlation for Few-Shot Segmentation
- Active Membership Inference Test (aMINT): Enhancing Model Auditability with Multi-Task Learning
- Two Stage Context Learning with Large Language Models for Multimodal Stance Detection on Climate Change
- Efficient resource management in UAVs for Visual Assistance
- EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration
- EHWGesture -- A dataset for multimodal understanding of clinical gestures
- Generating Transferrable Adversarial Examples via Local Mixing and Logits Optimization for Remote Sensing Object Recognition
- Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
- Deep Image Retrieval: Learning global representations for image search
- MixGAN: A Hybrid Semi-Supervised and Generative Approach for DDoS Detection in Cloud-Integrated IoT Networks
- Using Optimal Transport Aligned Latent Embeddings for Separated Flow Analysis
- Basis Vector Metric: A Method for Robust Open-Ended State Change Detection
- Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
- XSRD-Net: EXplainable Stroke Relapse Detection
- Bias-Aware Machine Unlearning: Towards Fairer Vision Models via Controllable Forgetting
- A machine learning assistant for detecting fraudulent activities in synchronous online programming exams
- ACE and Diverse Generalization via Selective Disagreement
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Sketch-R2CNN: An Attentive Network for Vector Sketch Recognition
- SBS: Enhancing Parameter-Efficiency of Neural Representations for Neural Networks via Spectral Bias Suppression
- G3CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition
- NeuroFabric: Identifying Ideal Topologies for Training A Priori Sparse Networks
- Representation Transfer by Optimal Transport
- XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
- H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
- BatStation: Toward In-Situ Radar Sensing on 5G Base Stations with Zero-Shot Template Generation
- Are Targeted Data Poisoning Attacks as Effective as We Think?
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
- Imitative Membership Inference Attack
- MRI-Based Brain Tumor Detection through an Explainable EfficientNetV2 and MLP-Mixer-Attention Architecture
- MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis
- Food Ingredients Recognition through Multi-label Learning
- Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- Integrated Detection and Tracking Based on Radar Range-Doppler Feature
- QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients
- FSG-Net: Frequency-Spatial Synergistic Gated Network for High-Resolution Remote Sensing Change Detection
- An Analysis of Scale Invariance in Object Detection - SNIP
- Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- Multi-Modal Camera-Based Detection of Vulnerable Road Users
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- Perception-oriented Bidirectional Attention Network for Image Super-resolution Quality Assessment
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
- NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables
- Pothole Detection and Recognition based on Transfer Learning
- A Multi-Modal Deep Learning Framework for Colorectal Pathology Diagnosis: Integrating Histological and Colonoscopy Data in a Pilot Study
- Signal-Based Malware Classification Using 1D CNNs
- Breaking SafetyCore: Exploring the Risks of On-Device AI Deployment
- IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter Editing
- Dimensionally Reduced Open-World Clustering: DROWCULA
- Predicting Brain Tumor Response to Therapy using a Hybrid Deep Learning and Radiomics Approach
- Analysis of Transferability Estimation Metrics for Surgical Phase Recognition
- Lookup multivariate Kolmogorov-Arnold Networks
- Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
- Exploring Light-Weight Object Recognition for Real-Time Document Detection
- Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- A Surrogate model for High Temperature Superconducting Magnets to Predict Current Distribution with Neural Network
- Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Micro-Expression Recognition via Fine-Grained Dynamic Perception
- Khana: A Comprehensive Indian Cuisine Dataset
- Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
- AttriPrompt: Dynamic Prompt Composition Learning for CLIP
- Challenges in Deep Learning-Based Small Organ Segmentation: A Benchmarking Perspective for Medical Research with Limited Datasets
- DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
- Parameter-Free Logit Distillation via Sorting Mechanism
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement\n Learning
- CRAB: Camera-Radar Fusion for Reducing Depth Ambiguity in Backward Projection based View Transformation
- Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics
- SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
- Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning
- High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
- JRN-Geo: A Joint Perception Network based on RGB and Normal images for Cross-view Geo-localization
- Transformer-based Topology Optimization
- Toward Efficient and Scalable Design of In-Memory Graph-Based Vector Search
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- Causal Multi-fidelity Surrogate Forward and Inverse Models for ICF Implosions
- Prior Distribution and Model Confidence
- TripOptimizer: Generative 3D Shape Optimization and Drag Prediction using Triplane VAE Networks
- Stochastic Analysis of Overlapping Generations Models Under Incomplete Markets
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
- On Evaluating the Poisoning Robustness of Federated Learning under Local Differential Privacy
- NSML: Meet the MLaaS platform with a real-world case study
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- On Hyperparameters and Backdoor-Resistance in Horizontal Federated Learning
- Semi-supervised Deep Transfer for Regression without Domain Alignment
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- TemporalFlowViz: Parameter-Aware Visual Analytics for Interpreting Scramjet Combustion Evolution
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
- Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
- Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
- Scale-interaction transformer: a hybrid cnn-transformer model for facial beauty prediction
- An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- Adapt in the Wild: Test-Time Entropy Minimization with Sharpness and Feature Regularization
- Dynamic Sensitivity Filter Pruning using Multi-Agent Reinforcement Learning For DCNN's
- Towards Open World Detection: A Survey
- Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
- CD-Mamba: Cloud detection with long-range spatial dependency modeling
- Exploiting Unlabeled Structures through Task Consistency Training for Versatile Medical Image Segmentation
- MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment using UAV Imagery
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
- Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
- Beyond Output Faithfulness: Learning Attributions that Preserve Computational Pathways
- MICACL: Multi-Instance Category-Aware Contrastive Learning for Long-Tailed Dynamic Facial Expression Recognition
- Noisy Label Refinement with Semantically Reliable Synthetic Images
- Suppressing the Unusual: towards Robust CNNs using Symmetric Activation Functions
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression
- An Automated, Scalable Machine Learning Model Inversion Assessment Pipeline
- DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
- YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- Revisiting Simple Baselines for In-The-Wild Deepfake Detection
- A Synthetic-to-Real Dehazing Method based on Domain Unification
- TriLiteNet: Lightweight Model for Multi-Task Visual Perception
- Using Deep Learning to Identify Artificial Satellite Trails in Multi-band Photometric Astronomical Images
- Learning from Majority Label: A Novel Problem in Multi-class Multiple-Instance Learning
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation
- PracMHBench: Re-evaluating Model-Heterogeneous Federated Learning Based on Practical Edge Device Constraints
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions
- Learning Neural Decoding with Parallelism and Self-Coordination for Quantum Error Correction
- Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
- Map as a By-product: Collective Landmark Mapping from IMU Data and User-provided Texts in Situated Tasks
- 2018 Robotic Scene Segmentation Challenge
- Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial
- Insights from Gradient Dynamics: Gradient Autoscaled Normalization
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- An Empirical Evaluation of Factors Affecting SHAP Explanation of Time Series Classification
- Teacher-Student Model for Detecting and Classifying Mitosis in the MIDOG 2025 Challenge
- The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
- Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
- Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
- ANNIE: Be Careful of Your Robots
- Prototype-Guided Robust Learning against Backdoor Attacks
- InWaveSR: Topography-Aware Super-Resolution Network for Internal Solitary Waves
- Temporal social network modeling of mobile connectivity data with graph neural networks
- SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather
- RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion
- A Plug-and-play Model-agnostic Embedding Enhancement Approach for Explainable Recommendation
- LSAM: Asynchronous Distributed Training with Landscape-Smoothed Sharpness-Aware Minimization
- FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
- High Cursive Complex Character Recognition using GAN External Classifier
- Isolated Bangla Handwritten Character Classification using Transfer Learning
- Putting An End to End-to-End: Gradient-Isolated Learning of Representations
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- Enhancing Robustness in Post-Processing Watermarking: An Ensemble Attack Network Using CNNs and Transformers
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- Unsupervised Instance Segmentation with Superpixels
- SPENet: Self-guided Prototype Enhancement Network for Few-shot Medical Image Segmentation
- S2M2ECG: Spatio-temporal bi-directional State Space Model Enabled Multi-branch Mamba for ECG
- Single Domain Generalization in Diabetic Retinopathy: A Neuro-Symbolic Learning Approach
- Artificial Intelligence-derived Cardiotocography Age as a Digital Biomarker for Predicting Future Adverse Pregnancy Outcomes
- LGBP-OrgaNet: Learnable Gaussian Band Pass Fusion of CNN and Transformer Features for Robust Organoid Segmentation and Tracking
- Invariant Features for Global Crop Type Classification
- Gradient Estimation Methods of Approximate Multipliers for High-Accuracy Retraining of Deep Learning Models
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity
- Continuous Saudi Sign Language Recognition: A Vision Transformer Approach
- Application of Decision Rules for Handling Class Imbalance in Semantic Segmentation
- Information transmission: Inferring change area from change moment in time series remote sensing images
- DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
- Backdoor Poisoning Attack Against Face Spoofing Attack Detection Methods
- Heatmap Guided Query Transformers for Robust Astrocyte Detection across Immunostains and Resolutions
- Lattice Annotated Temporal (LAT) Logic for Non-Markovian Reasoning
- Perturbing the Derivative: Wild Refitting for Model-Free Evaluation of Machine Learning Models under Bregman Losses
- RiverScope: High-Resolution River Masking Dataset
- A Convolutional Hierarchical Deep-learning Neural Network (C-HiDeNN) Framework for Non-linear Finite Element Analysis
- Fisher information flow in artificial neural networks
- Vision encoders should be image size agnostic and task driven
- Earthquake Source Depth Determination using Single Station Waveforms and Deep Learning
- SynthGenNet: a self-supervised approach for test-time generalization using synthetic multi-source domain mixing of street view images
- DeepGeo: Photo Localization with Deep Neural Network
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
- Superhuman Accuracy on the SNEMI3D Connectomics Challenge
- A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
- Prospects for acoustically monitoring ecosystem tipping points
- Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models
- Robust Small Methane Plume Segmentation in Satellite Imagery
- SALAD -- Semantics-Aware Logical Anomaly Detection
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- A flexible FPGA accelerator for convolutional neural networks
- Exploring Diffusion Models for Generative Forecasting of Financial Charts
- An Investigation of Visual Foundation Models Robustness
- Physics-Informed Machine Learning with Adaptive Grids for Optical Microrobot Depth Estimation
- A Continuous Encoding-Based Representation for Efficient Multi-Fidelity Multi-Objective Neural Architecture Search
- Energy-Efficient Split Learning for Resource-Constrained Environments: A Smart Farming Solution
- HydroVision: Predicting Optically Active Parameters in Surface Water Using Computer Vision
- Enhancing Fitness Movement Recognition with Attention Mechanism and Pre-Trained Feature Extractors
- Scalable, End-to-End, Deep-Learning-Based Data Reconstruction Chain for Particle Imaging Detectors
- Deep learning-enabled virtual multiplexed immunostaining of label-free tissue for vascular invasion assessment
- Augmented KRnet for density estimation and approximation
- VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
- Systematic Integration of Attention Modules into CNNs for Accurate and Generalizable Medical Image Diagnosis
- Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
- CLINN: Conservation Law Informed Neural Network for Approximating Discontinuous Solutions
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- Synesthesia of Machines (SoM)-Based Task-Driven MIMO System for Image Transmission
- A Multimodal Cross-View Model for Predicting Postoperative Neck Pain in Cervical Spondylosis Patients
- Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
- Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
- Towards deep-learning based detection and quantification of intestinal metaplasia on digitized gastric biopsies: a multi-expert comparative study
- ADVMEM: Adversarial Memory Initialization for Realistic Test-Time Adaptation via Tracklet-Based Benchmarking
- Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks
- BM-CL: Bias Mitigation through the lens of Continual Learning
- Handling imbalance and few-sample size in ML based Onion disease classification
- Examination of PCA Utilisation for Multilabel Classifier of Multispectral Images
- One-Shot Clustering for Federated Learning Under Clustering-Agnostic Assumption
- Structured AI Decision-Making in Disaster Management
- Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
- LiFeChain: Lightweight Blockchain for Secure and Efficient Federated Lifelong Learning in IoT
- Mamba-CNN: A Hybrid Architecture for Efficient and Accurate Facial Beauty Prediction
- Domain Adaptation via Feature Refinement
- Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup Tables
- Uirapuru: Timely Video Analytics for High-Resolution Steerable Cameras on Edge Devices
- CbLDM: A Diffusion Model for recovering nanostructure from atomic pair distribution function
- AgroSense: An Integrated Deep Learning System for Crop Recommendation via Soil Image Analysis and Nutrient Profiling
- Non-local RoIs for Instance Segmentation
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- RT-DETRv2 Explained in 8 Illustrations
- Enhanced Fingerprint-based Positioning With Practical Imperfections: Deep learning-based approaches
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
- SC-GIR: Goal-oriented Semantic Communication via Invariant Representation Learning
- Calibration-sample free distortion correction of electron diffraction patterns using deep learning
- Seeing through Unclear Glass: Occlusion Removal with One Shot
- Node-Equivariant Message Passing for Efficient and Accurate Machine Learning Interatomic Potentials
- Person Search with Natural Language Description
- Expandable Residual Approximation for Knowledge Distillation
- Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
- Controllable Generation of Implied Volatility Surfaces with Variational Autoencoders
- Deep Learning-Based Rock Particulate Classification Using Attention-Enhanced ConvNeXt
- Investigating Transfer Learning Capabilities of Vision Transformers and CNNs by Fine-Tuning a Single Trainable Block
- ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
- PrediTree: A Multi-Temporal Sub-meter Dataset of Multi-Spectral Imagery Aligned With Canopy Height Maps
- Wavelet-Enhanced PaDiM for Industrial Anomaly Detection
- Practical and Private Hybrid ML Inference with Fully Homomorphic Encryption
- Cross-Domain Few-Shot Segmentation via Ordinary Differential Equations over Time Intervals
- CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
- Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
- Sequential Difference Maximization: Generating Adversarial Examples via Multi-Stage Optimization
- A computer vision-based approach to enhance seismic catalogues
- IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects
- Decomposing and Revising What Language Models Generate
- MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
- Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
- A Closer Look at Local Aggregation Operators in Point Cloud Analysis
- Face4FairShifts: A Large Image Benchmark for Fairness and Robust Learning across Visual Domains
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
- Learning Frequency-aware Dynamic Network for Efficient Super-Resolution
- AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- Data Distillation: Towards Omni-Supervised Learning
- Cascade Residual Learning: A Two-stage Convolutional Neural Network for Stereo Matching
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- Multi-Focused Video Group Activities Hashing
- MLLMRec: Exploring the Potential of Multimodal Large Language Models in Recommender Systems
- NoiseCutMix: A Novel Data Augmentation Approach by Mixing Estimated Noise in Diffusion Models
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
- Target-Oriented Single Domain Generalization
- Scene Graph Generation via Conditional Random Fields
- Towards Zero-shot Cross-lingual Image Retrieval and Tagging
- Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks
- CryptoFace: End-to-End Encrypted Face Recognition
- Re-ID done right: towards good practices for person re-identification
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Theory Foundation of Physics-Enhanced Residual Learning
- Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
- Estimating Parameter Fields in Multi-Physics PDEs from Scarce Measurements
- Generative Latent Space Dynamics of Electron Density
- Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering
- Learning from Silence and Noise for Visual Sound Source Localization
- Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability
- FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- Activation Subspaces for Out-of-Distribution Detection
- Mapping like a Skeptic: Probabilistic BEV Projection for Online HD Mapping
- Self-supervised large-scale kidney abnormality detection in drug safety assessment studies
- Failure Prediction Is a Better Performance Proxy for Early-Exit Networks Than Calibration
- Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks
- Trees as Gaussians: Large-Scale Individual Tree Mapping
- Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs
- Automated Multi-label Classification of Eleven Retinal Diseases: A Benchmark of Modern Architectures and a Meta-Ensemble on a Large Synthetic Dataset
- Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609
- Mini Autonomous Car Driving based on 3D Convolutional Neural Networks
- Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
- Representation Learning with Adaptive Superpixel Coding
- Weakly-Supervised 3D Pose Estimation from a Single Image using Multi-View Consistency
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
- A Simple Cache Model for Image Recognition
- Learning Longer-term Dependencies in RNNs with Auxiliary Losses
- MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems
- Deep Active Learning for Lung Disease Severity Classification from Chest X-rays: Learning with Less Data in the Presence of Class Imbalance
- Class Incremental Continual Learning with Self-Organizing Maps and Variational Autoencoders Using Synthetic Replay
- Unitary Synthesis with AlphaZero via Dynamic Circuits
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
- Network of Experts for Large-Scale Image Categorization
- Fast Convergence Rates for Subsampled Natural Gradient Algorithms on Quadratic Model Problems
- Scaling Neuro-symbolic Problem Solving: Solver-Free Learning of Constraints and Objectives
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- SphereFace: Deep Hypersphere Embedding for Face Recognition
- Fast, Exact and Multi-Scale Inference for Semantic Image Segmentation\n with Deep Gaussian CRFs
- Multidimensional Distributional Neural Network Output Demonstrated in Super-Resolution of Surface Wind Speed
- Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
- Deep Learning Framework for Early Detection of Pancreatic Cancer Using Multi-Modal Medical Imaging Analysis
- Adapting Foundation Model for Dental Caries Detection with Dual-View Co-Training
- SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
- Self-Composing Neural Operators with Depth and Accuracy Scaling via Adaptive Train-and-Unroll Approach
- GLaRE: A Graph-based Landmark Region Embedding Network for Emotion Recognition
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- Domain Adaptation Techniques for Natural and Medical Image Classification
- MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
- IAENet: An Importance-Aware Ensemble Model for 3D Point Cloud-Based Anomaly Detection
- CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
- Prediction of mortality and resource utilization in critical care: a deep learning approach using multimodal electronic health records with natural language processing techniques
- MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
- More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
- Developing a Multi-Modal Machine Learning Model For Predicting Performance of Automotive Hood Frames
- Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
- Probability Density from Latent Diffusion Models for Out-of-Distribution Detection
- Unified Acoustic Representations for Screening Neurological and Respiratory Pathologies from Voice
- End-to-End Analysis of Charge Stability Diagrams with Transformers
- RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration
- A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
- Turning Tabular Foundation Models into Graph Foundation Models
- Exploring Machine Learning and Language Models for Multimodal Depression Detection
- Coresets from Trajectories: Selecting Data via Correlation of Loss Differences
- Split-Merge Pooling
- HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
- Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
- Dual Attention Networks for Multimodal Reasoning and Matching
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
- NM-Hebb: Coupling Local Hebbian Plasticity with Metric Learning for More Accurate and Interpretable CNNs
- Self-supervised structured object representation learning
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- From Research to Reality: Feasibility of Gradient Inversion Attacks in Federated Learning
- AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
- The Art of Hide and Seek: Making Pickle-Based Model Supply Chain Poisoning Stealthy Again
- SCAR: A Characterization Scheme for Multi-Modal Dataset
- ALSA: Anchors in Logit Space for Out-of-Distribution Accuracy Estimation
- Intriguing Properties of Contrastive Losses
- Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
- Advanced Deep Learning Techniques for Classifying Dental Conditions Using Panoramic X-Ray Images
- Weed Detection in Challenging Field Conditions: A Semi-Supervised Framework for Overcoming Shadow Bias and Data Scarcity
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Fast 3D Diffusion for Scalable Granular Media Synthesis
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
- EffNetViTLoRA: An Efficient Hybrid Deep Learning Approach for Alzheimer's Disease Diagnosis
- Development and Evaluation of an AI-Driven Telemedicine System for Prenatal Healthcare
- DoSReMC: Domain Shift Resilient Mammography Classification using Batch Normalization Adaptation
- Get Global Guarantees: On the Probabilistic Nature of Perturbation Robustness
- Few-Shot Connectivity-Aware Text Line Segmentation in Historical Documents
- AT-CXR: Uncertainty-Aware Agentic Triage for Chest X-rays
- Tackling Federated Unlearning as a Parameter Estimation Problem
- No Label Left Behind: A Unified Surface Defect Detection Model for all Supervision Regimes
- DeeperGCN: All You Need to Train Deeper GCNs
- ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval
- Enhancing Document VQA Models via Retrieval-Augmented Generation
- TaiBai: A fully programmable brain-inspired processor with topology-aware efficiency
- The GINN framework: a stochastic QED correspondence for stability and chaos in deep neural networks
- HierCVAE: Hierarchical Attention-Driven Conditional Variational Autoencoders for Multi-Scale Temporal Modeling
- SegReConcat: A Data Augmentation Method for Voice Anonymization Attack
- Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
- DQEN: Dual Query Enhancement Network for DETR-based HOI Detection
- Stabilizing Open-Set Test-Time Adaptation via Primary-Auxiliary Filtering and Knowledge-Integrated Prediction
- T-MLP: Tailed Multi-Layer Perceptron for Level-of-Detail Signal Representation
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- EMind: A Foundation Model for Multi-task Electromagnetic Signals Understanding
- Infant Cry Detection In Noisy Environment Using Blueprint Separable Convolutions and Time-Frequency Recurrent Neural Network
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
- Flatness-aware Curriculum Learning via Adversarial Difficulty
- Class-wise Flooding Regularization for Imbalanced Image Classification
- Enhancing Video-Based Robot Failure Detection Using Task Knowledge
- Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
- Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
- Clustering-based Feature Representation Learning for Oracle Bone Inscriptions Detection
- eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
- Zero-Shot Visual Recognition using Semantics-Preserving Adversarial Embedding Networks
- UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
- DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability
- A Deep Learning Application for Psoriasis Detection
- Bladder Cancer Diagnosis with Deep Learning: A Multi-Task Framework and Online Platform
- PneuGelSight: Soft Robotic Vision-Based Proprioception and Tactile Sensing
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- CellINR: Implicitly Overcoming Photo-induced Artifacts in 4D Live Fluorescence Microscopy
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- ANO : Faster is Better in Noisy Landscape
- Practical GPU Choices for Earth Observation: ResNet-50 Training Throughput on Integrated, Laptop, and Cloud Accelerators
- BirdRecorder's AI on Sky: Safeguarding birds of prey by detection and classification of tiny objects around wind turbines
- FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction
- Robust and Efficient Quantum Reservoir Computing with Discrete Time Crystal
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- Towards Reliable and Generalizable Differentially Private Machine Learning (Extended Version)
- A Cascaded Residual UNET for Fully Automated Segmentation of Prostate and Peripheral Zone in T2-weighted 3D Fast Spin Echo Images
- Improved Recurrent Neural Networks for Session-based Recommendations
- RCDINO: Enhancing Radar-Camera 3D Object Detection with DINOv2 Semantic Features
- GEN2: A Generative Prediction-Correction Framework for Long-time Emulations of Spatially-Resolved Climate Extremes
- Transfer learning optimization based on evolutionary selective fine tuning
- BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
- Are Virtual DES Images a Valid Alternative to the Real Ones?
- Integrated Sensing, Communication, and Computation for Over-the-Air Federated Edge Learning
- Lift-the-flap: what, where and when for context reasoning
- BasketLiDAR: The First LiDAR-Camera Multimodal Dataset for Professional Basketball MOT
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
- SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
- HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
- Towards Source-Free Machine Unlearning
- "What happens if..." Learning to Predict the Effect of Forces in Images
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- You Only Pose Once: A Minimalist's Detection Transformer for Monocular RGB Category-level 9D Multi-Object Pose Estimation
- Snap-Snap: Taking Two Images to Reconstruct 3D Human Gaussians in Milliseconds
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- Adversarial Hospital-Invariant Feature Learning for WSI Patch Classification
- Real-time Human Detection Model for Edge Devices
- Deep Learning for Taxol Exposure Analysis: A New Cell Image Dataset and Attention-Based Baseline Model
- Understanding Data Influence with Differential Approximation
- Reliable Smoke Detection via Optical Flow-Guided Feature Fusion and Transformer-Based Uncertainty Modeling
- Minimizing Task-Oriented Age of Information for Remote Monitoring with Pre-Identification
- Locality-aware Concept Bottleneck Model
- Artificial Intelligence-Based Multiscale Temporal Modeling for Anomaly Detection in Cloud Services
- Global-Distribution Aware Scenario-Specific Variational Representation Learning Framework
- WeedSense: Multi-Task Learning for Weed Segmentation, Height Estimation, and Growth Stage Classification
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Computing-In-Memory Dataflow for Minimal Buffer Traffic
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
- AFABench: A Generic Framework for Benchmarking Active Feature Acquisition
- Quantization Meets Spikes: Nearly Lossless Conversion at the First Timestep via Polarity Multi-Spike Mapping
- DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
- Adaptive Interpolating Quantum Transform: A Quantum-Native Framework for Efficient Transform Learning
- DOPA: Stealthy and Generalizable Backdoor Attacks from a Single Client under Challenging Federated Constraints
- AnchorSync: Global Consistency Optimization for Long Video Editing
- Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning
- Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
- Rethinking the Potential of Layer Freezing for Efficient DNN Training
- Maintaining Discrimination and Fairness in Class Incremental Learning
- Effect of Data Augmentation on Conformal Prediction for Diabetic Retinopathy
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- CLIPSym: Delving into Symmetry Detection with CLIP
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- In-hoc Concept Representations to Regularise Deep Learning in Medical Imaging
- Priority-based Parameter Propagation for Distributed DNN Training
- GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks
- OmViD: Omni-supervised active learning for video action detection
- Prospects for Deep-Learning-Based Mass Reconstruction of Ultra-High-Energy Cosmic Rays using Simulated Air-Shower Profiles
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Communication-Efficient Federated Learning with Adaptive Number of Participants
- Refining Contrastive Learning and Homography Relations for Multi-Modal Recommendation
- MUFFIN: Mixture of User-Adaptive Frequency Filtering for Sequential Recommendation
- DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
- Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
- AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN Training
- CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter
- Fracture Detection and Localisation in Wrist and Hand Radiographs using Detection Transformer Variants
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- Revisiting Deep Learning Models for Tabular Data
- Airborne acoustic emission enables sub-scanline keyhole porosity quantification and effective process characterization for metallic laser powder bed fusion
- Few Shot Speaker Recognition using Deep Neural Networks
- Unsupervised Open Domain Recognition by Semantic Discrepancy Minimization
- FAMNet: Integrating 2D and 3D Features for Micro-expression Recognition via Multi-task Learning and Hierarchical Attention
- Electromagnetic Signal Modulation Recognition based on Subgraph Embedding Learning
- Genuine multipartite entanglement verification with convolutional neural networks
- Training with the Invisibles: Obfuscating Images to Share Safely for Learning Visual Recognition Models
- Hotels-50K: A Global Hotel Recognition Dataset
- Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Graph Concept Bottleneck Models
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Zero-Resource Neural Machine Translation with Multi-Agent Communication Game
- A Dual-Attention Graph Network for fMRI Data Classification
- GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
- SlimComm: Doppler-Guided Sparse Queries for Bandwidth-Efficient Cooperative 3-D Perception
- Learning local and global prototypes with optimal transport for unsupervised anomaly detection and localization
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
- Exploration of Deep Learning Based Recognition for Urdu Text
- An Empirical Model of Large-Batch Training
- SocialTrack: Multi-Object Tracking in Complex Urban Traffic Scenes Inspired by Social Behavior
- CLAIRE-DSA: Fluoroscopic Image Classification for Quality Assurance of Computer Vision Pipelines in Acute Ischemic Stroke
- Deep Semantic Inference over the Air: An Efficient Task-Oriented Communication System
- Range-Angle Likelihood Maps for Indoor Positioning Using Deep Neural Networks
- Frequency-Driven Inverse Kernel Prediction for Single Image Defocus Deblurring
- Unlearning Comparator: A Visual Analytics System for Comparative Evaluation of Machine Unlearning Methods
- Neural Rendering for Sensor Adaptation in 3D Object Detection
- Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
- Refine-and-Contrast: Adaptive Instance-Aware BEV Representations for Multi-UAV Collaborative Object Detection
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- SuryaBench: Benchmark Dataset for Advancing Machine Learning in Heliophysics and Space Weather Prediction
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
- DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examples
- A Self-Ensemble Inspired Approach for Effective Training of Binary-Weight Spiking Neural Networks
- Towards High-Resolution Industrial Image Anomaly Detection
- Towards SISO Bistatic Sensing for ISAC
- EgoTwin: Dreaming Body and View in First Person
- Deploying Models to Non-participating Clients in Federated Learning without Fine-tuning: A Hypernetwork-based Approach
- Morphological classification of eclipsing binary stars using computer vision methods
- Field-level Reconstruction from Foreground-Contaminated 21-cm Maps
- Constraint-Aware Flow Matching via Randomized Exploration
- TCUQ: Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML
- Skin Cancer Classification: Hybrid CNN-Transformer Models with KAN-Based Fusion
- FedUNet: A Lightweight Additive U-Net Module for Federated Learning with Heterogeneous Models
- GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations
- Efficient and Verifiable Privacy-Preserving Convolutional Computation for CNN Inference with Untrusted Clouds
- Omni Survey for Multimodality Analysis in Visual Object Tracking
- Visual-Neural-Inspired Image Inpainting for Specific Objects-of-Interest Imaging
- Dextr: Zero-Shot Neural Architecture Search with Singular Value Decomposition and Extrinsic Curvature
- FLARE: Fast Low-rank Attention Routing Engine
- Anatomic Feature Fusion Model for Diagnosing Calcified Pulmonary Nodules on Chest X-Ray
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- Hierarchical Conformal Classification
- MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
- Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems
- HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
- Jamming Identification with Differential Transformer for Low-Altitude Wireless Networks
- WXSOD: A Benchmark for Robust Salient Object Detection in Adverse Weather Conditions
- Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
- Attention Pooling Enhances NCA-based Classification of Microscopy Images
- Splat Feature Solver
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis
- AICRN: Attention-Integrated Convolutional Residual Network for Interpretable Electrocardiogram Analysis
- TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- Generic Event Boundary Detection via Denoising Diffusion
- Towards interpretable prediction of recurrence risk in breast cancer using pathology foundation models
- WiseLVAM: A Novel Framework For Left Ventricle Automatic Measurements
- FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- Consensus-based Sequence Training for Video Captioning
- Large Kernel Modulation Network for Efficient Image Super-Resolution
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- DynamicPose: Real-time and Robust 6D Object Pose Tracking for Fast-Moving Cameras and Objects
- Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
- BConformeR: A Conformer Based on Mutual Sampling for Unified Prediction of Continuous and Discontinuous Antibody Binding Sites
- Dual-species atomic absorption image reconstruction using deep neural networks
- Regularized Ensembles and Transferability in Adversarial Learning
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Machine Learning-Based AES Key Recovery via Side-Channel Analysis on the ASCAD Dataset
- Mitigating Adversarial Effects Through Randomization
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- Grab-n-Go: On-the-Go Microgesture Recognition with Objects in Hand
- Recurrent Residual Module for Fast Inference in Videos
- Activate Me!: Designing Efficient Activation Functions for Privacy-Preserving Machine Learning with Fully Homomorphic Encryption
- TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
- A Real-time Concrete Crack Detection and Segmentation Model Based on YOLOv11
- Semi-Supervised Learning with Online Knowledge Distillation for Skin Lesion Classification
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- Repetitive TMS-based Identification of Methamphetamine-Dependent Individuals Using EEG Spectra
- CoMoNM: A Cost Modeling Framework for Compute-Near-Memory Systems
- LKFMixer: Exploring Large Kernel Feature For Efficient Image Super-Resolution
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
- Harmonized Gradient Descent for Class Imbalanced Data Stream Online Learning
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Boosting the Robustness-Accuracy Trade-off of SNNs by Robust Temporal Self-Ensemble
- Probing the Representational Power of Sparse Autoencoders in Vision Models
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Unsupervised Domain Adaptation for Spatio-Temporal Action Localization
- A Coarse-to-Fine Human Pose Estimation Method based on Two-stage Distillation and Progressive Graph Neural Network
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels
- BeeNet: Reconstructing Flower Shapes from Electric Fields using Deep Learning
- Hybrid-Hierarchical Fashion Graph Attention Network for Compatibility-Oriented and Personalized Outfit Recommendation
- EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
- Generalizable Federated Learning using Client Adaptive Focal Modulation
- A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots
- Axis-level Symmetry Detection with Group-Equivariant Representation
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- Symmetry-Constrained Multi-Scale Physics-Informed Neural Networks for Graphene Electronic Band Structure Prediction
- Lightweight CNNs for Embedded SAR Ship Target Detection and Classification
- Understanding the Disharmony between Dropout and Batch Normalization by Variance Shift
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- HyperTea: A Hypergraph-based Temporal Enhancement and Alignment Network for Moving Infrared Small Target Detection
- Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- FIND-Net -- Fourier-Integrated Network with Dictionary Kernels for Metal Artifact Reduction
- Fourier-Guided Attention Upsampling for Image Super-Resolution
- On Spectral Properties of Gradient-based Explanation Methods
- Adapting SAM via Cross-Entropy Masking for Class Imbalance in Remote Sensing Change Detection
- PSScreen: Partially Supervised Multiple Retinal Disease Screening
- EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
- DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- CTAP: Complementary Temporal Action Proposal Generation
- Super LiDAR Reflectance for Robotic Perception
- PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
- Improving OCR for Historical Texts of Multiple Languages
- Glo-UMF: A Unified Multi-model Framework for Automated Morphometry of Glomerular Ultrastructural Characterization
- SPATL: Salient Parameter Aggregation and Transfer Learning for Heterogeneous Clients in Federated Learning
- 3D latent diffusion models for parameterizing and history matching multiscenario facies systems
- Efficient Image Denoising Using Global and Local Circulant Representation
- BERT-VQA: Visual Question Answering on Plots
- Y-GAN: A Generative Adversarial Network for Depthmap Estimation from Multi-camera Stereo Images
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- Hierarchical Representations for Efficient Architecture Search
- Explainable AI Technique in Lung Cancer Detection Using Convolutional Neural Networks
- Out-of-Distribution Detection using Counterfactual Distance
- Deep Learning Enables Large-Scale Shape and Appearance Modeling in Total-Body DXA Imaging
- Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
- Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
- January Food Benchmark (JFB): A Public Benchmark Dataset and Evaluation Suite for Multimodal Food Analysis
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
- SHREC'25 Track on Multiple Relief Patterns: Report and Analysis
- Perceptual Reality Transformer: Neural Architectures for Simulating Neurological Perception Conditions
- Reverse Convolution and Its Applications to Image Restoration
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental Learning
- Analysis of the Compaction Behavior of Textile Reinforcements in Low-Resolution In-Situ CT Scans via Machine-Learning and Descriptor-Based Methods
- TriForecaster: A Mixture of Experts Framework for Multi-Region Electric Load Forecasting with Tri-dimensional Specialization
- The Role of Radiographic Knee Alignment in Total Knee Replacement Outcomes and Opportunities for Artificial Intelligence-Driven Assessment
- Predictive Uncertainty for Runtime Assurance of a Real-Time Computer Vision-Based Landing System
- Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction
- NegFaceDiff: The Power of Negative Context in Identity-Conditioned Diffusion for Synthetic Face Generation
- Artificial Intelligence, Domain AI Readiness, and Firm Productivity
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
- TFPose: Direct Human Pose Estimation with Transformers
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- Scalable Person Re-identification on Supervised Smoothed Manifold
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- Dynamical Alignment: A Principle for Adaptive Neural Computation
- Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
- Large-Small Model Collaborative Framework for Federated Continual Learning
- Deep Learning for Automated Identification of Vietnamese Timber Species: A Tool for Ecological Monitoring and Conservation
- DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
- What-Meets-Where: Unified Learning of Action and Contact Localization in a New Dataset
- Implicit Hypergraph Neural Networks: A Stable Framework for Higher-Order Relational Learning with Provable Guarantees
- HKT: A Biologically Inspired Framework for Modular Hereditary Knowledge Transfer in Neural Networks
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- Predicting protein variant properties with electrostatic representations
- Inter-expert reliability in multi-field-of-view automatic drusen segmentation analysis using optical coherence tomography
- Harnessing Input-Adaptive Inference for Efficient VLN
- Automated Charge Transition Detection in Quantum Dot Charge Stability Diagrams
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Low-Regret and Low-Complexity Learning for Hierarchical Inference
- Label Smoothing is a Pragmatic Information Bottleneck
- Exploring Cross-Stage Adversarial Transferability in Class-Incremental Continual Learning
- Frequency-Assisted Adaptive Sharpening Scheme Considering Bitrate and Quality Tradeoff
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
- SafeFix: Targeted Model Repair via Controlled Image Generation
- Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
- P-CAFE: Personalized Cost-Aware Incremental Feature Selection For Electronic Health Records
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
- QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
- Feature Enhancement Network: A Refined Scene Text Detector
- Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
- Resource-Aware Aggregation and Sparsification in Heterogeneous Ensemble Federated Learning
- Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Toward Lifelong Learning in Equilibrium Propagation: Sleep-like and Awake Rehearsal for Enhanced Stability
- Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent Space
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- Adaptive High-Frequency Preprocessing for Video Coding
- Image selective encryption analysis using mutual information in CNN based embedding space
- Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking
- Pep2Prob Benchmark: Predicting Fragment Ion Probability for MS2-based Proteomics
- Hierarchical Adaptive networks with Task vectors for Test-Time Adaptation
- Towards Efficient and Practical GPU Multitasking in the Era of LLM
- IPBA: Imperceptible Perturbation Backdoor Attack in Federated Self-Supervised Learning
- Identifying nonequilibrium degrees of freedom in high-dimensional stochastic systems
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- RedDino: A foundation model for red blood cell analysis
- TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation
- Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
- Deep Generative Models for Discrete Genotype Simulation
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- EFU: Enforcing Federated Unlearning via Functional Encryption
- Morphological Analysis of Semiconductor Microstructures using Skeleton Graphs
- Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
- AgentWorld: An Interactive Simulation Platform for Scene Construction and Mobile Robotic Manipulation
- Collaborative Learning of Scattering and Deep Features for SAR Target Recognition with Noisy Labels
- GAPNet: A Lightweight Framework for Image and Video Salient Object Detection via Granularity-Aware Paradigm
- ImageDDI: Image-enhanced Molecular Motif Sequence Representation for Drug-Drug Interaction Prediction
- Domain Generalization of Pathological Image Segmentation by Patch-Level and WSI-Level Contrastive Learning
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Fast and Generalizable parameter-embedded Neural Operators for Lithium-Ion Battery Simulation
- CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
- Enhanced Generative Structure Prior for Chinese Text Image Super-resolution
- Vision-Based Localization and LLM-based Navigation for Indoor Environments
- Prototype-Guided Curriculum Learning for Zero-Shot Learning
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- Enhancing Small-Scale Dataset Expansion with Triplet-Connection-based Sample Re-Weighting
- Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformer
- Adaptive Source-Channel Coding for Semantic Communications
- Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation
- Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- When Is Prior Knowledge Helpful? Exploring the Evaluation and Selection of Unsupervised Pretext Tasks from a Neuro-Symbolic Perspective
- Revisiting Data Attribution for Influence Functions
- SUIT: Spatial-Spectral Union-Intersection Interaction Network for Hyperspectral Object Tracking
- Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Lightweight Multi-Scale Feature Extraction with Fully Connected LMF Layer for Salient Object Detection
- Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications
- SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI
- ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
- Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be Forgotten
- Representation Understanding via Activation Maximization
- DexFruit: Dexterous Manipulation and Gaussian Splatting Inspection of Fruit
- ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting
- Membership Inference Attacks with False Discovery Rate Control
- QuiZSF: An efficient data-model interaction framework for zero-shot time-series forecasting
- eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
- Data-Efficient Neural Training with Dynamic Connectomes
- Label Inference Attacks against Federated Unlearning
- TADoc: Robust Time-Aware Document Image Dewarping
- Can Multitask Learning Enhance Model Explainability?
- BoRA: Towards More Expressive Low-Rank Adaptation with Block Diversity
- Sensory robustness through top-down feedback and neural stochasticity in recurrent vision models
- Robust-Sub-Gaussian Model Predictive Control for Safe Ultrasound-Image-Guided Robotic Spinal Surgery
- Rethinking Key-frame-based Micro-expression Recognition: A Robust and Accurate Framework Against Key-frame Errors
- CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
- Position: Ideas Should be the Center of Machine Learning Research
- Practical Block-wise Neural Network Architecture Generation
- Learnable pooling with Context Gating for video classification
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Feature-Space Oversampling for Addressing Class Imbalance in SAR Ship Classification
- A Classification-Aware Super-Resolution Framework for Ship Targets in SAR Imagery
- Are you In or Out (of gallery)? Wisdom from the Same-Identity Crowd
- Zero-shot self-supervised learning of single breath-hold magnetic resonance cholangiopancreatography (MRCP) reconstruction
- Vertex reconstruction in the TAO experiment
- Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
- Learning Representations of Satellite Images with Evaluations on Synoptic Weather Events
- Ensemble-Based Graph Representation of fMRI Data for Cognitive Brain State Classification
- MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
- Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis
- Efficient Bayer-Domain Video Computer Vision with Fast Motion Estimation and Learned Perception Residual
- Hybrid(Transformer+CNN)-based Polyp Segmentation
- Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements
- ASAudio: A Survey of Advanced Spatial Audio Research
- GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
- Accelerating Quantum Monte Carlo Calculations with Set-Equivariant Architectures and Transfer Learning
- Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball Rotation
- MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment
- FedX: Explanation-Guided Pruning for Communication-Efficient Federated Learning in Remote Sensing
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Interpretable Rheumatoid Arthritis Scoring via Anatomy-aware Multiple Instance Learning
- Multi-view Gaze Target Estimation
- When Deepfake Detection Meets Graph Neural Network:a Unified and Lightweight Learning Framework
- SMOL-MapSeg: Show Me One Label as prompt
- Keep It Real: Challenges in Attacking Compression-Based Adversarial Purification
- Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
- Long-Term Client Selection for Federated Learning with Non-IID Data: A Truthful Auction Approach
- Divide-and-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-supervised Continual Learning
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
- Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Neural Estimation of Information Leakage for Secure Communication System Design
- Human Activity Recognition from Smartphone Sensor Data for Clinical Trials
- Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
- X-MoGen: Unified Motion Generation across Humans and Animals
- PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning
- Digital Twin Channel-Aided CSI Prediction: An Environment-Based Subspace Extraction Approach for Achieving Low Overhead and High Robustness
- HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
- Learning from Similarity-Confidence and Confidence-Difference
- Multimodal Fact Checking with Unified Visual, Textual, and Contextual Representations
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer Arithmetic
- ULU: A Unified Activation Function
- Automatic Image Colorization with Convolutional Neural Networks and Generative Adversarial Networks
- A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
- Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
- MedMambaLite: Hardware-Aware Mamba for Medical Image Classification
- Attribute Guidance With Inherent Pseudo-label For Occluded Person Re-identification
- Unified modality separation: A vision-language framework for unsupervised domain adaptation
- FedMP: Tackling Medical Feature Heterogeneity in Federated Learning from a Manifold Perspective
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
- Functional Connectivity Graph Neural Networks
- Few-Shot Deployment of Pretrained MRI Transformers in Brain Imaging Tasks
- Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
- Gaussian mixture layers for neural networks
- Occupancy Learning with Spatiotemporal Memory
- BEVCon: Advancing Bird's Eye View Perception with Contrastive Learning
- PixCuboid: Room Layout Estimation from Multi-view Featuremetric Alignment
- Robust Processing-In-Memory Neural Networks via Noise-Aware Normalization
- How Does Bilateral Ear Symmetry Affect Deep Ear Features?
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- Drone Detection with Event Cameras
- TopKD: Top-scaled Knowledge Distillation
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection
- Learning Robust Intervention Representations with Delta Embeddings
- Benchmarking Foundation Models for Mitotic Figure Classification
- InfoQ: Mixed-Precision Quantization via Global Information Flow
- WSS-CL: Weight Saliency Soft-Guided Contrastive Learning for Efficient Machine Unlearning Image Classification
- S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- FLAT: Latent-Driven Arbitrary-Target Backdoor Attacks in Federated Learning
- Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
- Automated ultrasound doppler angle estimation using deep learning
- DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
- DP-DocLDM: Differentially Private Document Image Generation using Latent Diffusion Models
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Bootstrap Deep Spectral Clustering with Optimal Transport
- Multilingual Source Tracing of Speech Deepfakes: A First Benchmark
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- Energy-Efficient Real-Time 4-Stage Sleep Classification at 10-Second Resolution: A Comprehensive Study
- Excavate the potential of Single-Scale Features: A Decomposition Network for Water-Related Optical Image Enhancement
- CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
- SenseCrypt: Sensitivity-guided Selective Homomorphic Encryption for Joint Federated Learning in Cross-Device Scenarios
- Hybrid Quantum--Classical Machine Learning Potential with Variational Quantum Circuits
- Convolutional autoencoders for the reconstruction of three-dimensional interfacial multiphase flows
- Slice or the Whole Pie? Utility Control for AI Models
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- FeDaL: Federated Dataset Learning for Time Series Foundation Models
- CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
- Decoupled Contrastive Learning for Federated Learning
- Investigating the Impact of Large-Scale Pre-training on Nutritional Content Estimation from 2D Images
- CityPersons: A Diverse Dataset for Pedestrian Detection
- Deep learning framework for crater detection and identification on the Moon and Mars
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
- AttZoom: Attention Zoom for Better Visual Features
- Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
- DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- SoilNet: A Multimodal Multitask Model for Hierarchical Classification of Soil Horizons
- MURAUER: Mapping Unlabeled Real Data for Label AUstERity
- 4D-PreNet: A Unified Preprocessing Framework for 4D-STEM Data Analysis
- GaitAdapt: Continual Learning for Evolving Gait Recognition
- A Multimodal Late Fusion Model for E-Commerce Product Classification
- Evaluating the Predictive Value of Preoperative MRI for Erectile Dysfunction Following Radical Prostatectomy
- MedCAL-Bench: A Comprehensive Benchmark on Cold-Start Active Learning with Foundation Models for Medical Image Analysis
- Variety Is the Spice of Life: Detecting Misinformation with Dynamic Environmental Representations
- DeepAP: Deep Learning-based Aperture Photometry Feasibility Assessment and Aperture Size Prediction
- Zero Shot Domain Adaptive Semantic Segmentation by Synthetic Data Generation and Progressive Adaptation
- Ultralight Polarity-Split Neuromorphic SNN for Event-Stream Super-Resolution
- BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
- The Power of Many: Synergistic Unification of Diverse Augmentations for Efficient Adversarial Robustness
- PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
- COFFEE: A Shadow-Resilient Real-Time Pose Estimator for Unknown Tumbling Asteroids using Sparse Neural Networks
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Pseudo-label Induced Subspace Representation Learning for Robust Out-of-Distribution Detection
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- T2UE: Generating Unlearnable Examples from Text Descriptions
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- On the Fast Adaptation of Delayed Clients in Decentralized Federated Learning: A Centroid-Aligned Distillation Approach
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- Toward Practical Equilibrium Propagation: Brain-inspired Recurrent Neural Network with Feedback Regulation and Residual Connections
- Architectural Insights into Knowledge Distillation for Object Detection: A Comprehensive Review
- MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
- Neural Networks with Orthogonal Jacobian
- Communication and Computation Efficient Split Federated Learning in O-RAN
- Rethinking Transparent Object Grasping: Depth Completion with Monocular Depth Estimation and Instance Mask
- ASMR: Angular Support for Malfunctioning Client Resilience in Federated Learning
- Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion
- TRUDI and TITUS: A Multi-Perspective Dataset and A Three-Stage Recognition System for Transportation Unit Identification
- Assessing the Impact of Image Super Resolution on White Blood Cell Classification Accuracy
- Is Uncertainty Quantification a Viable Alternative to Learned Deferral?
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
- Failure Cases Are Better Learned But Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial Training
- Test-Time Model Adaptation for Quantized Neural Networks
- TrackletGait: A Robust Framework for Gait Recognition in the Wild
- AdvGAN++ : Harnessing latent layers for adversary generation
- DySTop
- Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure
- Semi-Supervised Dual-Threshold Contrastive Learning for Ultrasound Image Classification and Segmentation
- Tackling Ill-posedness of Reversible Image Conversion with Well-posed Invertible Network
- Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
- mCardiacDx: Radar-Driven Contactless Monitoring and Diagnosis of Arrhythmia
- YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
- Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model
- PRIME: Plasticity-Robust Incremental Model for Encrypted Traffic Classification in Dynamic Network Environments
- Multi-contrast machine learning improves schistosomiasis diagnostic performance
- NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
- Neuromorphic Computing with Multi-Frequency Oscillations: A Bio-Inspired Approach to Artificial Intelligence
- SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis
- Large-Scale Model Enabled Semantic Communication Based on Robust Knowledge Distillation
- Model Recycling Framework for Multi-Source Data-Free Supervised Transfer Learning
- AID4AD: Aerial Image Data for Automated Driving Perception
- Quantum Machine Learning-based Test Oracle for Autonomous Mobile Robots
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- Enhancing Object Discovery for Unsupervised Instance Segmentation and Object Detection
- Benchmarking Deep Learning-Based Object Detection Models on Feature Deficient Astrophotography Imagery Dataset
- Forecasting West Nile virus with deep graph encoders
- SitEmb-v1.5: Improved Context-Aware Dense Retrieval for Semantic Association and Long Story Comprehension
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- InspectVLM: Unified in Theory, Unreliable in Practice
- M3AD: Multi-task Multi-gate Mixture of Experts for Alzheimer's Disease Diagnosis with Conversion Pattern Modeling
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Boosting Generalization Performance in Model-Heterogeneous Federated Learning Using Variational Transposed Convolution
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Minimal High-Resolution Patches Are Sufficient for Whole Slide Image Representation via Cascaded Dual-Scale Reconstruction
- IMU: Influence-guided Machine Unlearning
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- CLIMD: A Curriculum Learning Framework for Imbalanced Multimodal Diagnosis
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
- Neural Predictive Control to Coordinate Discrete- and Continuous-Time Models for Time-Series Analysis with Control-Theoretical Improvements
- Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- Adverse Weather-Independent Framework Towards Autonomous Driving Perception through Temporal Correlation and Unfolded Regularization
- Glass Surface Segmentation with an RGB-D Camera via Weighted Feature Fusion for Service Robots
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- VAGPO: Vision-augmented Asymmetric Group Preference Optimization for Graph Routing Problems
- Beyond Vulnerabilities: A Survey of Adversarial Attacks as Both Threats and Defenses in Computer Vision Systems
- SPARTA: Advancing Sparse Attention in Spiking Neural Networks via Spike-Timing-Based Prioritization
- Energy-Efficient Federated Learning for Edge Real-Time Vision via Joint Data, Computation, and Communication Design
- Imbalance-Robust and Sampling-Efficient Continuous Conditional GANs via Adaptive Vicinal Learning and Auxiliary Regularization
- MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
- OmniEvent: Unified Event Representation Learning
- Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models
- The Vanishing Gradient Problem for Stiff Neural Differential Equations
- PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
- Physically-based Lighting Generation for Robotic Manipulation
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening
- Enhancing Diffusion-based Dataset Distillation via Adversary-Guided Curriculum Sampling
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- Calibrated Prediction Set in Fault Detection with Risk Guarantees via Significance Tests
- Deep Learning for Pavement Condition Evaluation Using Satellite Imagery
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Classification of Brain Tumors using Hybrid Deep Learning Models
- Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
- Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
- Large Language Models Facilitate Vision Reflection in Image Classification
- COSTARR: Consolidated Open Set Technique with Attenuation for Robust Recognition
- Evading Data Provenance in Deep Neural Networks
- Structured Spectral Graph Learning for Anomaly Classification in 3D Chest CT Scans
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- A Simple and Effective Method for Uncertainty Quantification and OOD Detection
- Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles
- Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
- Using Deep Learning for Segmentation and Counting within Microscopy Data
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
- DBLP: Noise Bridge Consistency Distillation For Efficient And Reliable Adversarial Purification
- HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning
- Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
- Representation Shift: Unifying Token Compression with FlashAttention
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- ECGTwin: Personalized ECG Generation Using Controllable Diffusion Model
- Multimodal Referring Segmentation: A Survey
- Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
- Backdoor Attacks on Deep Learning Face Detection
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- Rethinking Backbone Design for Lightweight 3D Object Detection in LiDAR
- BOOD: Boundary-based Out-Of-Distribution Data Generation
- MoSSDA: A Semi-Supervised Domain Adaptation Framework for Multivariate Time-Series Classification using Momentum Encoder
- Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- Deep Joint Source-Channel Coding for Small Satellite Applications
- Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition
- Segmenting proto-halos with vision transformers
- Explainable Image Classification with Reduced Overconfidence for Tissue Characterisation
- Max-Mahalanobis Linear Discriminant Analysis Networks
- Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
- Mamba-based Efficient Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
- DA-Occ: Direction-Aware 2D Convolution for Efficient and Geometry-Preserving 3D Occupancy Prediction
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- Beyond topography: Topographic regularization improves robustness and reshapes representations in convolutional neural networks
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Deep Contextual Recurrent Residual Networks for Scene Labeling
- I Am Big, You Are Little; I Am Right, You Are Wrong
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
Discussions
Related