vix.ing · top · new · best · stats · spec

B(eo)W(u)LF: Facilitating recurrence analysis on multi-level language

2013/08/12 by Alexandra Paxton, A. Paxton, Rick Dale +3
Computer Science · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1308.2696

3 pages plus 6 appendices (including code and sample data)

arxiv created 2013/08/12 · openalex publication_date 2013/08/12 · arxiv updated 2013/08/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Discourse analysis may seek to characterize not only the overall composition of a given text but also the dynamic patterns within the data. This technical report introduces a data format intended to facilitate multi-level investigations, which we call the by-word long-form or B(eo)W(u)LF. Inspired by the long-form data format required for mixed-effects modeling, B(eo)W(u)LF structures linguistic data into an expanded matrix encoding any number of researchers-specified markers, making it ideal for recurrence-based analyses. While we do not necessarily claim to be the first to use methods along these lines, we have created a series of tools utilizing Python and MATLAB to enable such discourse analyses and demonstrate them using 319 lines of the Old English epic poem, Beowulf, translated into modern English.

Related