2005/08/27 by Fabrice Rossi, Amaury Lendasse, Damien François +3 · 4 citations
Chemistry · Computer Science · Engineering · Environmental Science · Mathematics · #Advanced Chemical Sensor Technologies #Spectroscopy and Chemometric Analyses #Water Quality Monitoring and Analysis #cs.LG #cs.NE #stat.AP
paper · pdf · doi:10.1016/j.chemolab.2005.06.010
published as Chemometrics and Intelligent Laboratory Systems / I Mathematical Background Chemometrics Intell Lab Syst 80, 2 (2006) 215-226
openalex publication_date 2005/08/27 · arxiv created 2007/09/21 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Data from spectrophotometers form vectors of a large number of exploitable variables. Building quantitative models using these variables most often requires using a smaller set of variables than the initial one. Indeed, a too large number of input variables to a model results in a too large number of parameters, leading to overfitting and poor generalization abilities. In this paper, we suggest the use of the mutual information measure to select variables from the initial set. The mutual information measures the information content in input variables with respect to the model output, without making any assumption on the model that will be used; it is thus suitable for nonlinear modelling. In addition, it leads to the selection of variables among the initial set, and not to linear or nonlinear combinations of them. Without decreasing the model performances compared to other variable projection methods, it allows therefore a greater interpretability of the results.