vix.ing · top · new · best · stats · spec

Multi-scale hybrid vision transformer and Sinkhorn tokenizer for sewer defect classification

2022/10/20 by Joakim Bruslund Haurum, Meysam Madadi, Sérgio Escalera +2 · 2 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Digital Media Forensic Detection #Infrastructure Maintenance and Monitoring

paper · pdf · doi:10.1016/j.autcon.2022.104614

openalex publication_date 2022/10/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02

Abstract

A crucial part of image classification consists of capturing non-local spatial semantics of image content. This paper describes the multi-scale hybrid vision transformer (MSHViT), an extension of the classical convolutional neural network (CNN) backbone, for multi-label sewer defect classification. To better model spatial semantics in the images, features are aggregated at different scales non-locally through the use of a lightweight vision transformer, and a smaller set of tokens was produced through a novel Sinkhorn clustering-based tokenizer using distinct cluster centers. The proposed MSHViT and Sinkhorn tokenizer were evaluated on the Sewer-ML multi-label sewer defect classification dataset, showing consistent performance improvements of up to 2.53 percentage points.

Citations

Cited by