vix.ing · top · new · best · stats · spec

Are the discretised lognormal and hooked power law distributions\n plausible for citation data?

2016/03/16 by Mike Thelwall, Thelwall, Mike
Computer Science · Decision Sciences · #Advanced Text Analysis Techniques #Data Analysis with R #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Forecasting Techniques and Applications #Meta-analysis and systematic reviews #Statistical and Computational Modeling #scientometrics and bibliometrics research

paper · pdf · doi:10.48550/arxiv.1603.05078

openalex publication_date 2016/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

There is no agreement over which statistical distribution is most appropriate\nfor modelling citation count data. This is important because if one\ndistribution is accepted then the relative merits of different citation-based\nindicators, such as percentiles, arithmetic means and geometric means, can be\nmore fully assessed. In response, this article investigates the plausibility of\nthe discretised lognormal and hooked power law distributions for modelling the\nfull range of citation counts, with an offset of 1. The citation counts from 23\nScopus subcategories were fitted to hooked power law and discretised lognormal\ndistributions but both distributions failed a Kolmogorov-Smirnov goodness of\nfit test in over three quarters of cases. The discretised lognormal\ndistribution also seems to have the wrong shape for citation distributions,\nwith too few zeros and not enough medium values for all subjects. The cause of\npoor fits could be the impurity of the subject subcategories or the presence of\ninterdisciplinary research. Although it is possible to test for subject\nsubcategory purity indirectly through a goodness of fit test in theory with\nlarge enough sample sizes, it is probably not possible in practice. Hence it\nseems difficult to get conclusive evidence about the theoretically most\nappropriate statistical distribution.\n

Related