2021/12/01 by Tomer Amit, Tal Shaharbany, Amit, Tomer +5 · 44 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Artificial intelligence #Benchmark (surveying) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Encoder #Encoding (memory) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Geography #Image (mathematics) #Image segmentation #Machine Learning (cs.LG) #Pattern recognition (psychology) #Probabilistic logic #Scale-space segmentation #Segmentation #Segmentation-based object categorization #Statistical model #Video Surveillance and Tracking Methods #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2112.00390
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2021/12/01 · arxiv created 2022/09/07 · arxiv updated 2022/09/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08
Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.