vix.ing · top · new · best · stats

Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks

2015/11/19 by Stefan Lee, Senthil Purushwalkam, Lee, Stefan +7 · 1 voice · 204 citations
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Artificial intelligence #Artificial neural network #Class (philosophy) #Computer science #Convolutional neural network #Deep neural networks #Domain Adaptation and Few-Shot Learning #Engineering #Ensemble forecasting #Ensemble learning #Initialization #Machine Learning and Data Classification #Machine learning #Oracle #Range (aeronautics) #Variation (astronomy) #cs.CV #cs.LG #cs.NE

paper · pdf · doi:10.48550/arxiv.1511.06314

published in arXiv (Cornell University) (Cornell University)

arxiv created 2015/11/19 · openalex publication_date 2015/11/19 · arxiv published 2015/11/19 · arxiv updated 2015/11/20 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

Convolutional Neural Networks have achieved state-of-the-art performance on a wide range of tasks. Most benchmarks are led by ensembles of these powerful learners, but ensembling is typically treated as a post-hoc procedure implemented by averaging independently trained models with model variation induced by bagging or random initialization. In this paper, we rigorously treat ensembling as a first-class problem to explicitly address the question: what are the best strategies to create an ensemble? We first compare a large number of ensembling strategies, and then propose and evaluate novel strategies, such as parameter sharing (through a new family of models we call TreeNets) as well as training under ensemble-aware and diversity-encouraging losses. We demonstrate that TreeNets can improve ensemble performance and that diverse ensembles can be trained end-to-end under a unified loss, achieving significantly higher "oracle" accuracies than classical ensembles.

Cited by

Discussions

Related