2006/06/21 by A. Asensio Ramos · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · Physics and Astronomy · #Algorithm #Artificial intelligence #Artificial neural network #Blind Source Separation Techniques #Cellular Automata and Applications #Computer science #Data set #ENCODE #Fractal and DNA sequence analysis #Inversion (geology) #Mathematics #Minimum description length #Model selection #Overfitting #Selection (genetic algorithm) #Set (abstract data type) #Simple (philosophy) #astro-ph
paper · pdf · doi:10.1086/505136
published as Astrophys.J.646:1445-1451,2006 · Accepted for publication in the Astrophysical Journal
arxiv created 2006/06/21 · openalex publication_date 2006/07/28 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
It is shown that the two-part minimum description length principle can be used to discriminate among different models that can explain a given observed data set. The description length is chosen to be the sum of the lengths of the message needed to encode the model plus the message needed to encode the data when the model is applied to the data set. It is verified that the proposed principle can efficiently distinguish the model that correctly fits the observations while avoiding overfitting. The capabilities of this criterion are shown in two simple problems for the analysis of observed spectropolarimetric signals. The first is the denoising of observations with the aid of the PCA technique. The second is the selection of the optimal number of parameters in LTE inversions. We propose this criterion as a quantitative approach for distinguishing the most plausible model among a set of proposed models. This quantity is very easy to implement as an additional output on the existing inversion codes.