vix.ing · top · new · best · stats · spec

Convolutional Neural Network (CNN) vs Vision Transformer (ViT) for Digital Holography

2021/08/20 by Stéphane Cuenat, Cuenat, Stéphane, Raphaël Couturier +1
Biochemistry, Genetics and Molecular Biology · Engineering · Physics and Astronomy · #Cell Image Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #Digital Holography and Microscopy #FOS: Computer and information sciences #FOS: Electrical engineering #Image Processing Techniques and Applications #Image and Video Processing (eess.IV) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2108.09147

openalex publication_date 2021/08/20 · openalex created_date 2022/07/20 · openalex updated_date 2026/07/28

Abstract

In Digital Holography (DH), it is crucial to extract the object distance from a hologram in order to reconstruct its amplitude and phase. This step is called auto-focusing and it is conventionally solved by first reconstructing a stack of images and then by sharpening each reconstructed image using a focus metric such as entropy or variance. The distance corresponding to the sharpest image is considered the focal position. This approach, while effective, is computationally demanding and time-consuming. In this paper, the determination of the distance is performed by Deep Learning (DL). Two deep learning (DL) architectures are compared: Convolutional Neural Network (CNN) and Vision Transformer (ViT). ViT and CNN are used to cope with the problem of auto-focusing as a classification problem. Compared to a first attempt [11] in which the distance between two consecutive classes was 100μm, our proposal allows us to drastically reduce this distance to 1μm. Moreover, ViT reaches similar accuracy and is more robust than CNN.

Related