2019/02/23 by Lea Frermann, Mirella Lapata, Frermann, Lea +1
Computer Science · Psychology · Social Sciences · #Categorization, perception, and language #Child and Animal Learning Development #Computation and Language (cs.CL) #FOS: Computer and information sciences #Language and cultural evolution #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.1902.08830
openalex publication_date 2019/02/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Categories such as animal or furniture are acquired at an early age and play\nan important role in processing, organizing, and communicating world knowledge.\nCategories exist across cultures: they allow to efficiently represent the\ncomplexity of the world, and members of a community strongly agree on their\nnature, revealing a shared mental representation. Models of category learning\nand representation, however, are typically tested on data from small-scale\nexperiments involving small sets of concepts with artificially restricted\nfeatures; and experiments predominantly involve participants of selected\ncultural and socio-economical groups (very often involving western native\nspeakers of English such as U.S. college students) . This work investigates\nwhether models of categorization generalize (a) to rich and noisy data\napproximating the environment humans live in; and (b) across languages and\ncultures. We present a Bayesian cognitive model designed to jointly learn\ncategories and their structured representation from natural language text which\nallows us to (a) evaluate performance on a large scale, and (b) apply our model\nto a diverse set of languages. We show that meaningful categories comprising\nhundreds of concepts and richly structured featural representations emerge\nacross languages. Our work illustrates the potential of recent advances in\ncomputational modeling and large scale naturalistic datasets for cognitive\nscience research.\n