2021/05/31 by Jussi Karlgren, Karlgren, Jussi
Computer Science · Mathematics · #Advanced Text Analysis Techniques #Analogy #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Data science #Digital humanities #FOS: Computer and information sciences #Linguistics #Mathematics #Natural Language Processing Techniques #Point (geometry) #Political science #Relevance (law) #Scholarship #Similarity (geometry) #Topic Modeling #World Wide Web #cs.CL
paper · pdf · doi:10.48550/arxiv.2105.14921
arxiv created 2021/05/31 · openalex publication_date 2021/05/31 · arxiv updated 2021/06/01 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
This paper describes how the current lexical similarity and analogy gold standards are built to conform to certain ideas about what the models they are designed to evaluate are used for. Topical relevance has always been the most important target notion for information access tools and related language technology technologies, and while this has proven a useful starting point for much of what information technology is used for, it does not always align well with other uses to which technologies are being put, most notably use cases from digital scholarship in the humanities or social sciences. This paper argues for more systematic formulation of requirements from the digital humanities and social sciences and more explicit description of the assumptions underlying model design.