2019/06/04 by Junkyung Kim, Kim, Junkyung, Drew Linsley +6 · 10 citations
Computer Science · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Artificial intelligence #Cognitive neuroscience of visual object recognition #Cognitive psychology #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Face Recognition and Perception #Gestalt psychology #Horizontal and vertical #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Neuroscience #Object (grammar) #Perception #Psychology #Task (project management) #Top-down and bottom-up design #Visual Attention and Saliency Detection #Visual perception #Visual perception and processing mechanisms #cs.AI #cs.CV #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1906.01558
published in arXiv (Cornell University) (Cornell University) · Published in ICLR 2020
openalex publication_date 2019/06/04 · arxiv created 2020/10/28 · arxiv updated 2020/10/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Forming perceptual groups and individuating objects in visual scenes is an essential step towards visual intelligence. This ability is thought to arise in the brain from computations implemented by bottom-up, horizontal, and top-down connections between neurons. However, the relative contributions of these connections to perceptual grouping are poorly understood. We address this question by systematically evaluating neural network architectures featuring combinations bottom-up, horizontal, and top-down connections on two synthetic visual tasks, which stress low-level "Gestalt" vs. high-level object cues for perceptual grouping. We show that increasing the difficulty of either task strains learning for networks that rely solely on bottom-up connections. Horizontal connections resolve straining on tasks with Gestalt cues by supporting incremental grouping, whereas top-down connections rescue learning on tasks with high-level object cues by modifying coarse predictions about the position of the target object. Our findings dissociate the computational roles of bottom-up, horizontal and top-down connectivity, and demonstrate how a model featuring all of these interactions can more flexibly learn to form perceptual groups.