1997/12/24 by Franco Orsucci, F. Orsucci, K. Walter +12 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Physics and Astronomy · #Chaos control and synchronization #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fractal and DNA sequence analysis #Neural Networks and Applications #cmp-lg #cs.CL
paper · pdf · doi:10.48550/arxiv.cmp-lg/9712010
8 pages, 7 figures, submitted to Intl. J. Chaos Theory Applications
arxiv created 1997/12/24 · openalex publication_date 1997/12/24 · arxiv updated 2012/08/27 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
A methodology based upon recurrence quantification analysis is proposed for the study of orthographic structure of written texts. Five different orthographic data sets (20th century Italian poems, 20th century American poems, contemporary Swedish poems with their corresponding Italian translations, Italian speech samples, and American speech samples) were subjected to recurrence quantification analysis, a procedure which has been found to be diagnostically useful in the quantitative assessment of ordered series in fields such as physics, molecular dynamics, physiology, and general signal processing. Recurrence quantification was developed from recurrence plots as applied to the analysis of nonlinear, complex systems in the physical sciences, and is based on the computation of a distance matrix of the elements of an ordered series (in this case the letters consituting selected speech and poetic texts). From a strictly mathematical view, the results show the possibility of demonstrating invariance between different language exemplars despite the apparent low-level of coding (orthography). Comparison with the actual texts confirms the ability of the method to reveal recurrent structures, and their complexity. Using poems as a reference standard for judging speech complexity, the technique exhibits language independence, order dependence and freedom from pure statistical characteristics of studied sequences, as well as consistency with easily identifiable texts. Such studies may provide phenomenological markers of hidden structure as coded by the purely orthographic level.