vix.ing · top · new · best · stats · spec

A streaming ensemble algorithm (SEA) for large-scale classification

2001/08/26 by W. Nick Street, YongSeog Kim · 5 citations
Computer Science · #Data Stream Mining Techniques #Machine Learning and Data Classification #Anomaly Detection Techniques and Applications #Computer science #Boosting (machine learning) #Resampling #Machine learning #Decision tree #Artificial intelligence #Concept drift #Data mining #Ensemble learning #Scale (ratio) #Heuristic #Streaming data #Context (archaeology) #Statistical classification #Data stream mining

paper · doi:10.1145/502512.502568

openalex publication_date 2001/08/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04

Abstract

Ensemble methods have recently garnered a great deal of attention in the machine learning community. Techniques such as Boosting and Bagging have proven to be highly effective but require repeated resampling of the training data, making them inappropriate in a data mining context. The methods presented in this paper take advantage of plentiful data, building separate classifiers on sequential chunks of training points. These classifiers are combined into a fixed-size ensemble using a heuristic replacement strategy. The result is a fast algorithm for large-scale or streaming data that classifies as well as a single decision tree built on all the data, requires approximately constant memory, and adjusts quickly to concept drift.

Citations

Cited by