vix.ing · top · new · best · stats · spec

Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders

2020/10/18 by Masha Itkina, Itkina, Masha, Boris Ivanovic +7 · 1 citation
Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #I.2.10 #I.2.6 #I.2.9 #Machine Learning (cs.LG) #Machine Learning in Healthcare #Robotics (cs.RO) #cs.AI #cs.CV #cs.LG #cs.RO

paper · pdf · doi:10.48550/arxiv.2010.09164

21 pages, 15 figures, 34th Conference on Neural Information Processing Systems (NeurIPS 2020)

openalex publication_date 2020/10/19 · arxiv created 2021/01/18 · arxiv updated 2021/01/19 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Discrete latent spaces in variational autoencoders have been shown to effectively capture the data distribution for many real-world problems such as natural language understanding, human intent prediction, and visual scene representation. However, discrete latent spaces need to be sufficiently large to capture the complexities of real-world data, rendering downstream tasks computationally challenging. For instance, performing motion planning in a high-dimensional latent representation of the environment could be intractable. We consider the problem of sparsifying the discrete latent space of a trained conditional variational autoencoder, while preserving its learned multimodality. As a post hoc latent space reduction technique, we use evidential theory to identify the latent classes that receive direct evidence from a particular input condition and filter out those that do not. Experiments on diverse tasks, such as image generation and human behavior prediction, demonstrate the effectiveness of our proposed technique at reducing the discrete latent sample space size of a model while maintaining its learned multimodality.

Cited by

Related