2021/09/21 by Yiwen Liu, Liu, Yiwen, Xiaoxiao Sun +5
Biochemistry, Genetics and Molecular Biology · Chemistry · #Applications (stat.AP) #FOS: Computer and information sciences #Gene expression and cancer classification #Metabolomics and Mass Spectrometry Studies #Methodology (stat.ME) #Spectroscopy and Chemometric Analyses
paper · pdf · doi:10.48550/arxiv.2109.09940
openalex publication_date 2021/09/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Very often for the same scientific question, there may exist different techniques or experiments that measure the same numerical quantity. Historically, various methods have been developed to exploit the information within each type of data independently. However, statistical data fusion methods that could effectively integrate multi-source data under a unified framework are lacking. In this paper, we propose a novel data fusion method, called B-scaling, for integrating multi-source data. Consider K measurements that are generated from different sources but measure the same latent variable through some linear or nonlinear ways. We seek to find a representation of the latent variable, named B-mean, which captures the common information contained in the K measurements while takes into account the nonlinear mappings between them and the latent variable. We also establish the asymptotic property of the B-mean and apply the proposed method to integrate multiple histone modifications and DNA methylation levels for characterizing epigenomic landscape. Both numerical and empirical studies show that B-scaling is a powerful data fusion method with broad applications.