2010/06/01 by V. A. Yatsko, M. S. Starikov, A. V. Butakov · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Authorship Attribution and Profiling #Automatic summarization #Natural language processing #Computer science #Information retrieval #Artificial intelligence #Speech recognition
paper · doi:10.3103/s0005105510030027
openalex publication_date 2010/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/06/11
This paper describes an experimental method for automatic text genre recognition based on 45 statistical, lexical, syntactic, positional, and discursive parameters. The suggested method includes: (1) the development of software permitting heterogeneous parameters to be normalized and clustered using the k-means algorithm; (2) the verification of parameters; (3) the selection of the parameters that are the most significant for scientific, newspaper, and artistic texts using two-factor analysis algorithms. Adaptive summarization algorithms have been developed based on these parameters.