2019/08/12 by R. Annunziata, Annunziata, Roberto, Christos Sagonas +3 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
paper · pdf · doi:10.48550/arxiv.1908.04130
openalex publication_date 2019/08/12 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28
Extrapolating fine-grained pixel-level correspondences in a fully\nunsupervised manner from a large set of misaligned images can benefit several\ncomputer vision and graphics problems, e.g. co-segmentation, super-resolution,\nimage edit propagation, structure-from-motion, and 3D reconstruction. Several\njoint image alignment and congealing techniques have been proposed to tackle\nthis problem, but robustness to initialisation, ability to scale to large\ndatasets, and alignment accuracy seem to hamper their wide applicability. To\novercome these limitations, we propose an unsupervised joint alignment method\nleveraging a densely fused spatial transformer network to estimate the warping\nparameters for each image and a low-capacity auto-encoder whose reconstruction\nerror is used as an auxiliary measure of joint alignment. Experimental results\non digits from multiple versions of MNIST (i.e., original, perturbed, affNIST\nand infiMNIST) and faces from LFW, show that our approach is capable of\naligning millions of images with high accuracy and robustness to different\nlevels and types of perturbation. Moreover, qualitative and quantitative\nresults suggest that the proposed method outperforms state-of-the-art\napproaches both in terms of alignment quality and robustness to initialisation.\n