2024/10/20 by R. Docherty, Docherty, Ronan, Antonis Vamvakeros +3 · 1 citation
Computer Science · Engineering · #Advanced Neural Network Applications #Advanced X-ray and CT Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #FOS: Physical sciences #Image and Video Processing (eess.IV) #Industrial Vision Systems and Defect Detection #Materials Science (cond-mat.mtrl-sci) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2410.19836
openalex publication_date 2024/10/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The features of self-supervised vision transformers (ViTs) contain strong semantic and positional information relevant to downstream tasks like object localization and segmentation. Recent works combine these features with traditional methods like clustering, graph partitioning or region correlations to achieve impressive baselines without finetuning or training additional networks. We leverage upsampled features from ViT networks (e.g DINOv2) in two workflows: in a clustering based approach for object localization and segmentation, and paired with standard classifiers in weakly supervised materials segmentation. Both show strong performance on benchmarks, especially in weakly supervised segmentation where the ViT features capture complex relationships inaccessible to classical approaches. We expect the flexibility and generalizability of these features will both speed up and strengthen materials characterization, from segmentation to property-prediction.