1999/01/01 by H. Drucker, Harris Drucker, Donghui Wu +2 · 7 citations
Computer Science · #Face and Expression Recognition #Spam and Phishing Detection #Text and Document Classification Technologies
paper · doi:10.1109/72.788645
openalex publication_date 1999/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study the use of support vector machines (SVM's) in classifying e-mail as spam or nonspam by comparing it to three other classification algorithms: Ripper, Rocchio, and boosting decision trees. These four algorithms were tested on two different data sets: one data set where the number of features were constrained to the 1000 best features and another data set where the dimensionality was over 7000. SVM's performed best when using binary features. For both data sets, boosting trees and SVM's had acceptable test performance in terms of accuracy and speed. However, SVM's had significantly less training time.