2020/03/19 by Malik Boudiaf, Jérôme Rony, Boudiaf, Malik +11 · 5 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face and Expression Recognition #Face recognition and analysis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2003.08983
openalex publication_date 2020/03/19 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Recently, substantial research efforts in Deep Metric Learning (DML) focused\non designing complex pairwise-distance losses, which require convoluted schemes\nto ease optimization, such as sample mining or pair weighting. The standard\ncross-entropy loss for classification has been largely overlooked in DML. On\nthe surface, the cross-entropy may seem unrelated and irrelevant to metric\nlearning as it does not explicitly involve pairwise distances. However, we\nprovide a theoretical analysis that links the cross-entropy to several\nwell-known and recent pairwise losses. Our connections are drawn from two\ndifferent perspectives: one based on an explicit optimization insight; the\nother on discriminative and generative views of the mutual information between\nthe labels and the learned features. First, we explicitly demonstrate that the\ncross-entropy is an upper bound on a new pairwise loss, which has a structure\nsimilar to various pairwise losses: it minimizes intra-class distances while\nmaximizing inter-class distances. As a result, minimizing the cross-entropy can\nbe seen as an approximate bound-optimization (or Majorize-Minimize) algorithm\nfor minimizing this pairwise loss. Second, we show that, more generally,\nminimizing the cross-entropy is actually equivalent to maximizing the mutual\ninformation, to which we connect several well-known pairwise losses.\nFurthermore, we show that various standard pairwise losses can be explicitly\nrelated to one another via bound relationships. Our findings indicate that the\ncross-entropy represents a proxy for maximizing the mutual information -- as\npairwise losses do -- without the need for convoluted sample-mining heuristics.\nOur experiments over four standard DML benchmarks strongly support our\nfindings. We obtain state-of-the-art results, outperforming recent and complex\nDML methods.\n