2017/11/22 by Nikhil Garg, Londa Schiebinger, Dan Jurafsky +1 · 1 voice · 3 citations
Computer Science · Physics and Astronomy · Social Sciences · #Anthropology #Artificial intelligence #Cartography #Census #Computational and Text Analysis Methods #Computer science #Demography #Dynamics (music) #Embedding #Ethnic group #Geography #Immigration #Intersection (aeronautics) #Language and cultural evolution #Linguistics #Opinion Dynamics and Social Influence #Population #Sociology #Word (group theory) #Word embedding #cs.CL #cs.CY
paper · pdf · doi:10.1073/pnas.1720347115
arxiv created 2017/11/22 · arxiv published 2017/11/22 · openalex publication_date 2018/04/03 · arxiv updated 2018/06/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Word embeddings use vectors to represent words such that the geometry between vectors captures semantic relationship between the words. In this paper, we develop a framework to demonstrate how the temporal dynamics of the embedding can be leveraged to quantify changes in stereotypes and attitudes toward women and ethnic minorities in the 20th and 21st centuries in the United States. We integrate word embeddings trained on 100 years of text data with the U.S. Census to show that changes in the embedding track closely with demographic and occupation shifts over time. The embedding captures global social shifts -- e.g., the women's movement in the 1960s and Asian immigration into the U.S -- and also illuminates how specific adjectives and occupations became more closely associated with certain populations over time. Our framework for temporal analysis of word embedding opens up a powerful new intersection between machine learning and quantitative social science.