vix.ing · top · new · best · stats · spec

Dictionary-based methods for information extraction

2004/02/29 by Andrea Baronchelli, A. Baronchelli, E. Caglioti +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Physics and Astronomy · #Algorithms and Data Compression #Fractal and DNA sequence analysis #cond-mat.other #cond-mat.stat-mech #cs.IR #q-bio.GN #q-bio.OT #semigroups and automata theory

paper · pdf · doi:10.1016/j.physa.2004.01.072

published as Physica A - Vol 342/1-2 pp 294-300 (2004) · 7 pages, Latex, elsart style

openalex publication_date 2004/05/31 · arxiv created 2004/09/14 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper we present a general method for information extraction that exploits the features of data compression techniques. We first define and focus our attention on the so-called "dictionary" of a sequence. Dictionaries are intrinsically interesting and a study of their features can be of great usefulness to investigate the properties of the sequences they have been extracted from (e.g. DNA strings). We then describe a procedure of string comparison between dictionary-created sequences (or "artificial texts") that gives very good results in several contexts. We finally present some results on self-consistent classification problems.

Citations