vix.ing · top · new · best · stats · spec

Data-to-Value: An Evaluation-First Methodology for Natural Language Projects

2022/01/19 by Jochen L. Leidner, Leidner, Jochen L. · 1 voice
Computer Science · Mathematics · #62H99 #68T50 #68U15 #91B02 #Computation and Language (cs.CL) #D.2.9 #FOS: Computer and information sciences #H.0 #I.2.7 #I.7.m #Methodology (stat.ME) #cs.CL #stat.ME

paper · pdf · doi:10.48550/arxiv.2201.07725

Abstract

Big data, i.e. collecting, storing and processing of data at scale, has recently been possible due to the arrival of clusters of commodity computers powered by application-level distributed parallel operating systems like HDFS/Hadoop/Spark, and such infrastructures have revolutionized data mining at scale. For data mining project to succeed more consistently, some methodologies were developed (e.g. CRISP-DM, SEMMA, KDD), but these do not account for (1) very large scales of processing, (2) dealing with textual (unstructured) data (i.e. Natural Language Processing (NLP, "text analytics"), and (3) non-technical considerations (e.g. legal, ethical, project managerial aspects). To address these shortcomings, a new methodology, called "Data to Value" (D2V), is introduced, which is guided by a detailed catalog of questions in order to avoid a disconnect of big data text analytics project team with the topic when facing rather abstract box-and-arrow diagrams commonly associated with methodologies.

Discussions

Related