2018/03/01 by Alain Rakotomamonjy, Abraham Traoré, Rakotomamonjy, Alain +8 · 4 citations
Computer Science · Mathematics · #Artificial intelligence #Computer science #Data mining #Discrete mathematics #Discriminative model #Distance measures #Divergence (linguistics) #Domain Adaptation and Few-Shot Learning #Embedding #Empirical probability #FOS: Computer and information sciences #Face and Expression Recognition #Kernel (algebra) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Mathematics #Measure (data warehouse) #Pairwise comparison #Posterior probability #Probability distribution #Similarity (geometry) #Statistics #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1803.00250
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2018/03/01 · arxiv created 2018/11/15 · arxiv updated 2018/11/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework extends the theory of similarity of Balcan et al. (2008) to the population distribution case and we show that, for some learning problems, some dissimilarity on distribution achieves low-error linear decision functions with high probability. Our key result is to prove that the theory also holds for empirical distributions. Algorithmically, the proposed approach consists in computing a mapping based on pairwise dissimilarity where learning a linear decision function is amenable. Our experimental results show that the Wasserstein distance embedding performs better than kernel mean embeddings and computing Wasserstein distance is far more tractable than estimating pairwise Kullback-Leibler divergence of empirical distributions.