vix.ing · top · new · best · stats · spec

Fusing heterogeneous data sets

2019/08/23 by Yipeng Song, Song, Yipeng
Biochemistry, Genetics and Molecular Biology · Mathematics · #FOS: Biological sciences #FOS: Computer and information sciences #Gene expression and cancer classification #Genomics (q-bio.GN) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Metabolomics and Mass Spectrometry Studies #Methodology (stat.ME) #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.1908.09653

openalex publication_date 2019/08/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

In systems biology, it is common to measure biochemical entities at different levels of the same biological system. One of the central problems for the data fusion of such data sets is the heterogeneity of the data. This thesis discusses two types of heterogeneity. The first one is the type of data, such as metabolomics, proteomics and RNAseq data in genomics. These different omics data reflect the properties of the studied biological system from different perspectives. The second one is the type of scale, which indicates the measurements obtained at different scales, such as binary, ordinal, interval and ratio-scaled variables. In this thesis, we developed several statistical methods capable to fuse data sets of these two types of heterogeneity. The advantages of the proposed methods in comparison with other approaches are assessed using comprehensive simulations as well as the analysis of real biological data sets.

Citations

Related