2013/03/27 by Christopher Minas, Minas, Christopher, Edward Curry +3
Biochemistry, Genetics and Molecular Biology · #Applications (stat.AP) #FOS: Computer and information sciences #Gene expression and cancer classification #Genetic Associations and Epidemiology #Genomic variations and chromosomal abnormalities #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.1303.7002
openalex publication_date 2013/03/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Due to rapid technological advances, a wide range of different measurements\ncan be obtained from a given biological sample including single nucleotide\npolymorphisms, copy number variation, gene expression levels, DNA methylation\nand proteomic profiles. Each of these distinct measurements provides the means\nto characterize a certain aspect of biological diversity, and a fundamental\nproblem of broad interest concerns the discovery of shared patterns of\nvariation across different data types. Such data types are heterogeneous in the\nsense that they represent measurements taken at very different scales or\ndescribed by very different data structures. We propose a distance-based\nstatistical test, the generalized RV (GRV) test, to assess whether there is a\ncommon and non-random pattern of variability between paired biological\nmeasurements obtained from the same random sample. The measurements enter the\ntest through distance measures which can be chosen to capture particular\naspects of the data. An approximate null distribution is proposed to compute\np-values in closed-form and without the need to perform costly Monte Carlo\npermutation procedures. Compared to the classical Mantel test for association\nbetween distance matrices, the GRV test has been found to be more powerful in a\nnumber of simulation settings. We also report on an application of the GRV test\nto detect biological pathways in which genetic variability is associated to\nvariation in gene expression levels in ovarian cancer samples, and present\nresults obtained from two independent cohorts.\n