Cross-Entropy Loss Functions: Theoretical Analysis and Applications
2023/04/14 by Anqi Mao, Mehryar Mohri, Mao, Anqi +3 · 92 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Integrated Circuits and Semiconductor Failure Analysis #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · pdf · doi:10.48550/arxiv.2304.07288
openalex publication_date 2023/04/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Cross-entropy is a widely used loss function in applications. It coincides with the logistic loss applied to the outputs of a neural network, when the softmax is used. But, what guarantees can we rely on when using cross-entropy as a surrogate loss? We present a theoretical analysis of a broad family of loss functions, comp-sum losses, that includes cross-entropy (or logistic loss), generalized cross-entropy, the mean absolute error and other cross-entropy-like loss functions. We give the first H-consistency bounds for these loss functions. These are non-asymptotic guarantees that upper bound the zero-one loss estimation error in terms of the estimation error of a surrogate loss, for the specific hypothesis set H used. We further show that our bounds are tight. These bounds depend on quantities called minimizability gaps. To make them more explicit, we give a specific analysis of these gaps for comp-sum losses. We also introduce a new family of loss functions, smooth adversarial comp-sum losses, that are derived from their comp-sum counterparts by adding in a related smooth term. We show that these loss functions are beneficial in the adversarial setting by proving that they admit H-consistency bounds. This leads to new adversarial robustness algorithms that consist of minimizing a regularized smooth adversarial comp-sum loss. While our main purpose is a theoretical analysis, we also present an extensive empirical analysis comparing comp-sum losses. We further report the results of a series of experiments demonstrating that our adversarial robustness algorithms outperform the current state-of-the-art, while also achieving a superior non-adversarial accuracy.
Cited by
- Principled Algorithms for Optimizing Generalized Metrics in Binary Classification
- Benchmarking LLMs for Predictive Applications in the Intensive Care Units
- Orthogonal Activation with Implicit Group-Aware Bias Learning for Class Imbalance
- Interpretable Plant Leaf Disease Detection Using Attention-Enhanced CNN
- Model Agnostic Preference Optimization for Medical Image Segmentation
- Machine learning discovers new champion codes
- ID-PaS+ : Identity-Aware Predict-and-Search for General Mixed-Integer Linear Programs
- Partitioning the Sample Space for a More Precise Shannon Entropy Estimation
- MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection
- Learning Visual Affordance from Audio
- An Efficient Privacy-preserving Intrusion Detection Scheme for UAV Swarm Networks
- Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- SpectralTrain: A Universal Framework for Hyperspectral Image Classification
- Blind Quality Enhancement of Compressed Video via Fine-Grained Degradation-Guided Sequential Inference
- X-IONet: Cross-Platform Inertial Odometry Network with Dual-Stage Attention
- Quantum optical neural networks using atom-cavity interactions to provide all-optical nonlinearity
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- Budgeted Multiple-Expert Deferral
- MIN-Merging: Merge the Important Neurons for Model Merging
- Readability-Robust Code Summarization via Meta Curriculum Learning
- Energy-Efficient Autonomous Driving with Adaptive Perception and Robust Decision
- RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Relative-Based Scaling Law for Neural Language Models
- Enabling Fine-Grained Operating Points for Black-Box LLMs
- ArmFormer: Lightweight Transformer Architecture for Real-Time Multi-Class Weapon Segmentation and Classification
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- Adversarial Robustness in One-Stage Learning-to-Defer
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- On the Rademacher Complexity of Graph Neural Networks: Unifying Expressivity and Geometry
- To Ask or Not to Ask: Learning to Require Human Feedback
- Mangrove3D: Terrestrial Laser Scanning Dataset for Coastal Mangrove Forests
- TFM Dataset: A Novel Multi-task Dataset and Integrated Pipeline for Automated Tear Film Break-Up Segmentation
- What Scales in Cross-Entropy Scaling Law?
- A Greedy PDE Router for Blending Neural Operators and Classical Methods
- PEARL: Performance-Enhanced Aggregated Representation Learning
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- EmbeddedML: A New Optimized and Fast Machine Learning Library
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Some Robustness Properties of Label Cleaning
- FASL-Seg: Anatomy and Tool Segmentation of Surgical Scenes
- Adaptive Contrast Adjustment Module: A Clinically-Inspired Plug-and-Play Approach for Enhanced Fetal Plane Classification
- Distributed Gossip-GAN for Low-overhead CSI Feedback Training in FDD mMIMO-OFDM Systems
- Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs
- FusionSort: Enhanced Cluttered Waste Segmentation with Advanced Decoding and Comprehensive Modality Optimization
- DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- Mamba-FCS: Joint Spatio- Frequency Feature Fusion, Change-Guided Attention, and SeK Loss for Enhanced Semantic Change Detection in Remote Sensing
- Introducing Fractional Classification Loss for Robust Learning with Noisy Labels
- SSFMamba: Symmetry-driven Spatial-Frequency Feature Fusion for 3D Medical Image Segmentation
- TensorHyper-VQC: A Tensor-Train-Guided Hypernetwork for Robust and Scalable Variational Quantum Computing
- Probabilistic Consistency in Machine Learning and Its Connection to Uncertainty Quantification
- Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications
- Modeling Insider Filing Delays in Financial Markets with an Interpretable XGBoost Framework
- Rethinking Memorization Measures and their Implications in Large Language Models
- HEIMDALL: a grapH-based sEIsMic Detector And Locator for microseismicity
- Capsule-ConvKAN: A Hybrid Neural Approach to Medical Image Classification
- Surg-SegFormer: A Dual Transformer-Based Model for Holistic Surgical Scene Segmentation
- Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
- Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
- CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical Vision
- Mastering Multiple-Expert Routing: Realizable H-Consistency and Strong Guarantees for Learning to Defer
- Multi-Modal Beamforming with Model Compression and Modality Generation for V2X Networks
- Quality over Quantity: An Effective Large-Scale Data Reduction Strategy Based on Pointwise V-Information
- ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge Processors
- VQC-MLPNet: An Unconventional Hybrid Quantum-Classical Architecture for Scalable and Robust Quantum Machine Learning
- An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry
- StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
- Integrated Image Reconstruction and Target Recognition based on Deep Learning Technique
- Active Sampling for MRI-based Sequential Decision Making
- DSAGL: Dual-Stream Attention-Guided Learning for Weakly Supervised Whole Slide Image Classification
- Universal Visuo-Tactile Video Understanding for Embodied Interaction
- A Unified Foundation Model for Wireless Technology Recognition and Localization
- 15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning
- Federated prediction for scalable and privacy-preserved knowledge-based planning in radiotherapy
- Field Matters: A lightweight LLM-enhanced Method for CTR Prediction
- Real-World fNIRS-Based Brain-Computer Interfaces: Benchmarking Deep Learning and Classical Models in Interactive Gaming
- ZENN: A Thermodynamics-Inspired Computational Framework for Heterogeneous Data-Driven Modeling
- Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young Losses
- EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
- No Query, No Access
- Diffusion-driven SpatioTemporal Graph KANsformer for Medical Examination Recommendation
- SAMSEM – A Generic and Scalable Approach for IC Metal Line Segmentation
- Multi-Hierarchical Fine-Grained Feature Mapping Driven by Feature Contribution for Molecular Odor Prediction
- Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements
- SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space
- Dynamic Graph-Like Learning with Contrastive Clustering on Temporally-Factored Ship Motion Data for Imbalanced Sea State Estimation in Autonomous Vessel
- Cross-Asset Risk Management: Integrating LLMs for Real-Time Monitoring of Equity, Fixed Income, and Currency Markets
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k Experts
- DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation
- Comorbidity-Informed Transfer Learning for Neuro-developmental Disorder Diagnosis
Related