2021/08/20 by David W. Schindler, Schindler, David, Felix Bensmann +5 · 5 citations
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Scientific Computing and Data Management #Software Engineering Research #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2108.09070
openalex publication_date 2021/08/20 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Knowledge about software used in scientific investigations is important for\nseveral reasons, for instance, to enable an understanding of provenance and\nmethods involved in data handling. However, software is usually not formally\ncited, but rather mentioned informally within the scholarly description of the\ninvestigation, raising the need for automatic information extraction and\ndisambiguation. Given the lack of reliable ground truth data, we present\nSoMeSci (Software Mentions in Science) a gold standard knowledge graph of\nsoftware mentions in scientific articles. It contains high quality annotations\n(IRR: \κ=.82) of 3756 software mentions in 1367 PubMed Central\narticles. Besides the plain mention of the software, we also provide relation\nlabels for additional information, such as the version, the developer, a URL or\ncitations. Moreover, we distinguish between different types, such as\napplication, plugin or programming environment, as well as different types of\nmentions, such as usage or creation. To the best of our knowledge, SoMeSci is\nthe most comprehensive corpus about software mentions in scientific articles,\nproviding training samples for Named Entity Recognition, Relation Extraction,\nEntity Disambiguation, and Entity Linking. Finally, we sketch potential use\ncases and provide baseline results.\n