2016/03/03 by Adam Lerer, Sam Gross, Lerer, Adam +3 · 1 voice · 81 citations
Computer Science · Mathematics · #Advanced Vision and Imaging #Artificial intelligence #Block (permutation group theory) #Cognitive science #Computer science #Geometry #Human Pose and Action Recognition #Intuition #Machine learning #Mathematics #Music Technology and Sound Studies #Physics engine #Simulation #cs.AI
paper · pdf · doi:10.48550/arxiv.1603.01312
published in arXiv (Cornell University), 430-438 (Cornell University)
arxiv created 2016/03/03 · openalex publication_date 2016/03/03 · arxiv updated 2016/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Wooden blocks are a common toy for infants, allowing them to develop motor skills and gain intuition about the physical behavior of the world. In this paper, we explore the ability of deep feed-forward models to learn such intuitive physics. Using a 3D game engine, we create small towers of wooden blocks whose stability is randomized and render them collapsing (or remaining upright). This data allows us to train large convolutional network models which can accurately predict the outcome, as well as estimating the block trajectories. The models are also able to generalize in two important ways: (i) to new physical scenarios, e.g. towers with an additional block and (ii) to images of real wooden blocks, where it obtains a performance comparable to human subjects.