1999/01/01 by John Bentley, D. L. McIlroy · 67 citations
Computer Science · #Algorithms and Data Compression #semigroups and automata theory #Network Packet Processing and Optimization #Computer science #Data compression #Compression (physics) #Algorithm #Data structure #Theoretical computer science #Data mining #Programming language #Physics
paper · doi:10.1109/dcc.1999.755678
openalex publication_date 1999/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
We describe a precompression algorithm that effectively represents any long common strings that appear in a file. The algorithm interacts well with standard compression algorithms that represent shorter strings that are near in the input text. Our experiments show that some real data sets do indeed contain many long common strings. We extend the fingerprint mechanisms of our algorithm to a program that identifies long common strings in an input file. This program gives interesting insights into the structure of real data files that contain long common strings.