2021/01/12 by Md Abul Hayat, Abul Hayat, Peter Harrington +9 · 1 voice
Computer Science · Physics and Astronomy · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Artificial intelligence #Artificial neural network #Astrophysics #Computer science #Cosmology and Nongalactic Astrophysics (astro-ph.CO) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Physical sciences #Galaxy #Image (mathematics) #Instrumentation and Methods for Astrophysics (astro-ph.IM) #Machine learning #Pattern recognition (psychology) #Physics #Representation (politics) #Sample (material) #Similarity (geometry) #Sky #Supervised learning #Task (project management) #Video Surveillance and Tracking Methods #astro-ph.CO #astro-ph.IM #cs.AI
paper · pdf · doi:10.48550/arxiv.2101.04293
arxiv created 2021/01/12 · openalex publication_date 2021/01/12 · arxiv published 2021/01/12 · arxiv updated 2021/01/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We use a contrastive self-supervised learning framework to estimate distances to galaxies from their photometric images. We incorporate data augmentations from computer vision as well as an application-specific augmentation accounting for galactic dust. We find that the resulting visual representations of galaxy images are semantically useful and allow for fast similarity searches, and can be successfully fine-tuned for the task of redshift estimation. We show that (1) pretraining on a large corpus of unlabeled data followed by fine-tuning on some labels can attain the accuracy of a fully-supervised model which requires 2-4x more labeled data, and (2) that by fine-tuning our self-supervised representations using all available data labels in the Main Galaxy Sample of the Sloan Digital Sky Survey (SDSS), we outperform the state-of-the-art supervised learning method.