2014/03/19 by Tommi Suvitaival, Suvitaival, Tommi, Simon Rogers +3
Biochemistry, Genetics and Molecular Biology · Chemistry · Mathematics · #Analytical Chemistry and Chromatography #Applications (stat.AP) #Biomolecules (q-bio.BM) #FOS: Biological sciences #FOS: Computer and information sciences #Metabolomics and Mass Spectrometry Studies #Quantitative Methods (q-bio.QM) #Spectroscopy and Chemometric Analyses #q-bio.BM #q-bio.QM #stat.AP
paper · pdf · doi:10.48550/arxiv.1403.4732
arxiv created 2014/03/19 · openalex publication_date 2014/03/19 · arxiv updated 2014/03/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Mass spectrometry-based metabolomic analysis depends upon the identification of spectral peaks by their mass and retention time. Statistical analysis that follows the identification currently relies on one main peak of each compound. However, a compound present in the sample typically produces several spectral peaks due to its isotopic properties and the ionization process of the mass spectrometer device. In this work, we investigate the extent to which these additional peaks can be used to increase the statistical strength of differential analysis. We present a Bayesian approach for integrating data of multiple detected peaks that come from one compound. We demonstrate the approach through a simulated experiment and validate it on ultra performance liquid chromatography-mass spectrometry (UPLC-MS) experiments for metabolomics and lipidomics. Peaks that are likely to be associated with one compound can be clustered by the similarity of their chromatographic shape. Changes of concentration between sample groups can be inferred more accurately when multiple peaks are available. When the sample-size is limited, the proposed multi-peak approach improves the accuracy at inferring covariate effects. An R implementation, data and the supplementary material are available at http://research.ics.aalto.fi/mi/software/peakANOVA/ .