2019/12/17 by Yashaswi Verma, Verma, Yashaswi
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Text and Document Classification Technologies
paper · pdf · doi:10.48550/arxiv.1912.08140
openalex publication_date 2019/12/17 · openalex created_date 2024/04/10 · openalex updated_date 2026/07/28
The goal of eXtreme Multi-label Learning (XML) is to automatically annotate a\ngiven data point with the most relevant subset of labels from an extremely\nlarge vocabulary of labels (e.g., a million labels). Lately, many attempts have\nbeen made to address this problem that achieve reasonable performance on\nbenchmark datasets. In this paper, rather than coming-up with an altogether new\nmethod, our objective is to present and validate a simple baseline for this\ntask. Precisely, we investigate an on-the-fly global and structure preserving\nfeature embedding technique using random projections whose learning phase is\nindependent of training samples and label vocabulary. Further, we show how an\nensemble of multiple such learners can be used to achieve further boost in\nprediction accuracy with only linear increase in training and prediction time.\nExperiments on three public XML benchmarks show that the proposed approach\nobtains competitive accuracy compared with many existing methods. Additionally,\nit also provides around 6572x speed-up ratio in terms of training time and\naround 14.7x reduction in model-size compared to the closest competitors on the\nlargest publicly available dataset.\n