2019/09/25 by Tristan Bepler, Ellen D. Zhong, Bepler, Tristan +7 · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Advanced Electron Microscopy Techniques and Applications #Advanced Image Processing Techniques #Cell Image Analysis Techniques #Computational Physics and Python Applications #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Biological sciences #FOS: Computer and information sciences #Image Processing Techniques and Applications #Machine Learning (cs.LG) #Quantitative Methods (q-bio.QM)
paper · pdf · doi:10.48550/arxiv.1909.11663
openalex publication_date 2019/09/25 · openalex created_date 2019/10/03 · openalex updated_date 2026/07/28
Given an image dataset, we are often interested in finding data generative\nfactors that encode semantic content independently from pose variables such as\nrotation and translation. However, current disentanglement approaches do not\nimpose any specific structure on the learned latent representations. We propose\na method for explicitly disentangling image rotation and translation from other\nunstructured latent factors in a variational autoencoder (VAE) framework. By\nformulating the generative model as a function of the spatial coordinate, we\nmake the reconstruction error differentiable with respect to latent translation\nand rotation parameters. This formulation allows us to train a neural network\nto perform approximate inference on these latent variables while explicitly\nconstraining them to only represent rotation and translation. We demonstrate\nthat this framework, termed spatial-VAE, effectively learns latent\nrepresentations that disentangle image rotation and translation from content\nand improves reconstruction over standard VAEs on several benchmark datasets,\nincluding applications to modeling continuous 2-D views of proteins from single\nparticle electron microscopy and galaxies in astronomical images.\n