2005/01/31 by L.L. Gonçalves, L. L. Goncalves, L. B. Goncalves +1
Biochemistry, Genetics and Molecular Biology · Economics, Econometrics and Finance · Physics and Astronomy · #Complex Systems and Time Series Analysis #Fractal and DNA sequence analysis #cond-mat.other #physics.soc-ph
paper · pdf · doi:10.1016/j.physa.2005.06.049
27 pages, 10 tables,15 figures. Revised version accepted in Physica A
arxiv created 2005/06/03 · openalex publication_date 2005/07/19 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
We present in this paper a numerical investigation of literary texts by various well-known English writers, covering the first half of the twentieth century, based upon the results obtained through corpus analysis of the texts. A fractal power law is obtained for the lexical wealth defined as the ratio between the number of different words and the total number of words of a given text. By considering as a signature of each author the exponent and the amplitude of the power law, and the standard deviation of the lexical wealth, it is possible to discriminate works of different genres and writers and show that each writer has a very distinct signature, either considered among other literary writers or compared with writers of non-literary texts. It is also shown that, for a given author, the signature is able to discriminate between short stories and novels.