vix.ing · top · new · best · stats · spec

Nonapproximablity of the Normalized Information Distance

2009/10/22 by Sebastiaan A. Terwijn, Terwijn, Sebastiaan A., Leen Torenvliet +3
Biochemistry, Genetics and Molecular Biology · Computer Science · #Algorithms and Data Compression #Computability, Logic, AI Algorithms #Computational Complexity (cs.CC) #FOS: Computer and information sciences #Fractal and DNA sequence analysis #Information Theory (cs.IT)

paper · pdf · doi:10.48550/arxiv.0910.4353

openalex publication_date 2009/10/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Normalized information distance (NID) uses the theoretical notion of Kolmogorov complexity, which for practical purposes is approximated by the length of the compressed version of the file involved, using a real-world compression program. This practical application is called `normalized compression distance' and it is trivially computable. It is a parameter-free similarity measure based on compression, and is used in pattern recognition, data mining, phylogeny, clustering, and classification. The complexity properties of its theoretical precursor, the NID, have been open. We show that the NID is neither upper semicomputable nor lower semicomputable up to any reasonable precision.

Related