2020/07/31 by A. Emin Orhan, Orhan, A. Emin, Vaibhav V. Gupta +3 · 5 citations
Psychology · Social Sciences · #Child and Animal Learning Development #Computer Vision and Pattern Recognition (cs.CV) #Education and Critical Thinking Development #FOS: Computer and information sciences #Innovative Teaching and Learning Methods #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
paper · pdf · doi:10.48550/arxiv.2007.16189
openalex publication_date 2020/07/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Within months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives.