2010/09/19 by David R. Hardoon, Hardoon, David R., Kristiaan Pelcksman +1
Mathematics · #Applications (stat.AP) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Statistics Theory (math.ST) #math.ST #stat.AP #stat.ML #stat.TH
paper · pdf · doi:10.48550/arxiv.1009.3601
arxiv created 2010/09/19 · arxiv updated 2010/09/21
This paper studies the problem of learning clusters which are consistently present in different (continuously valued) representations of observed data. Our setup differs slightly from the standard approach of (co-) clustering as we use the fact that some form of `labeling' becomes available in this setup: a cluster is only interesting if it has a counterpart in the alternative representation. The contribution of this paper is twofold: (i) the problem setting is explored and an analysis in terms of the PAC-Bayesian theorem is presented, (ii) a practical kernel-based algorithm is derived exploiting the inherent relation to Canonical Correlation Analysis (CCA), as well as its extension to multiple views. A content based information retrieval (CBIR) case study is presented on the multi-lingual aligned Europal document dataset which supports the above findings.