2017/01/18 by Hui Guan, Thanos Gentimis, Guan, Hui +5
Business, Management and Accounting · Decision Sciences · Social Sciences · #Big Data and Business Intelligence #Data Quality and Management #FOS: Computer and information sciences #General Literature (cs.GL) #Information Retrieval (cs.IR) #Misinformation and Its Impacts
paper · pdf · doi:10.48550/arxiv.1702.02107
openalex publication_date 2017/01/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We introduce the idea of Data Readiness Level (DRL) to measure the relative richness of data to answer specific questions often encountered by data scientists. We first approach the problem in its full generality explaining its desired mathematical properties and applications and then we propose and study two DRL metrics. Specifically, we define DRL as a function of at least four properties of data: Noisiness, Believability, Relevance, and Coherence. The information-theoretic based metrics, Cosine Similarity and Document Disparity, are proposed as indicators of Relevance and Coherence for a piece of data. The proposed metrics are validated through a text-based experiment using Twitter data.