2026/07/01 by Niklas Müller, Cees G. M. Snoek, Iris I.A. Groen +1 · 1 voice
Neuroscience · Computer Science · Psychology · #Face Recognition and Perception #Face recognition and analysis #Child and Animal Learning Development
paper · pdf · doi:10.1080/09540091.2026.2694256
Convolutional Neural Networks (CNNs) surpass human-level performance on visual object recognition, yet their behavior differs from humans in important ways. One prominent example is that CNNs trained on ImageNet exhibit a texture bias, while humans show a strong shape bias. Although CNN shape bias can be increased through e.g., data augmentation or additional training, the source of this discrepancy remains unclear. Developmental research suggests that one factor driving human shape bias is that during early childhood, toddlers tend to fill their field-of-view with close-up objects. We operationalize this close-up as a zoom-in on objects during CNN training, which increases shape bias without additional training or data augmentation. Systematic manipulation of background-object ratios during training reveals a strong inverse correlation with shape bias. Notably, zooming in on objects, thereby more closely emulating child vision, aligns classification accuracy and shape bias between humans and CNNs. Finally, we achieve a near human-like shape bias when using a developmentally-inspired background-object ratio for training and shape bias assessment. These findings demonstrate that a simple adjustment to image datasets — zooming in on objects — can produce human-like shape bias. This suggests that human learning strategies offer a promising avenue for developing human-aligned, efficient, and robust vision CNNs.