vix.ing · top · new · best · stats · spec

Disentangling by Partitioning: A Representation Learning Framework for\n Multimodal Sensory Data

2018/05/29 by Wei-Ning Hsu, James Glass, Hsu, Wei-Ning +1 · 2 citations
Computer Science · #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.1805.11264

Abstract

Multimodal sensory data resembles the form of information perceived by humans\nfor learning, and are easy to obtain in large quantities. Compared to unimodal\ndata, synchronization of concepts between modalities in such data provides\nsupervision for disentangling the underlying explanatory factors of each\nmodality. Previous work leveraging multimodal data has mainly focused on\nretaining only the modality-invariant factors while discarding the rest. In\nthis paper, we present a partitioned variational autoencoder (PVAE) and several\ntraining objectives to learn disentangled representations, which encode not\nonly the shared factors, but also modality-dependent ones, into separate latent\nvariables. Specifically, PVAE integrates a variational inference framework and\na multimodal generative model that partitions the explanatory factors and\nconditions only on the relevant subset of them for generation. We evaluate our\nmodel on two parallel speech/image datasets, and demonstrate its ability to\nlearn disentangled representations by qualitatively exploring within-modality\nand cross-modality conditional generation with semantics and styles specified\nby examples. For quantitative analysis, we evaluate the classification accuracy\nof automatically discovered semantic units. Our PVAE can achieve over 99%\naccuracy on both modalities.\n

Citations

Cited by

Related