2021/12/05 by Tal Shaharabany, Shaharabany, Tal, Lior Wolf +1
Computer Science · Environmental Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Remote Sensing and LiDAR Applications #Visual Attention and Saliency Detection #cs.CV
paper · pdf · doi:10.48550/arxiv.2112.02535
arxiv created 2021/12/05 · openalex publication_date 2021/12/05 · arxiv updated 2021/12/07 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
The leading segmentation methods represent the output map as a pixel grid. We study an alternative representation in which the object edges are modeled, per image patch, as a polygon with k vertices that is coupled with per-patch label probabilities. The vertices are optimized by employing a differentiable neural renderer to create a raster image. The delineated region is then compared with the ground truth segmentation. Our method obtains multiple state-of-the-art results: 76.26% mIoU on the Cityscapes validation, 90.92% IoU on the Vaihingen building segmentation benchmark, 66.82% IoU for the MoNU microscopy dataset, and 90.91% for the bird benchmark CUB. Our code for training and reproducing these results is attached as supplementary.