vix.ing · top · new · best · stats · spec

Identifying Outliers using Influence Function of Multiple Kernel Canonical Correlation Analysis

2016/06/01 by Md Ashad Alam, Alam, Md Ashad, Yu‐Ping Wang +1
Biochemistry, Genetics and Molecular Biology · Mathematics · #Advanced Statistical Methods and Models #FOS: Computer and information sciences #Genetic and phenotypic traits in livestock #Machine Learning (stat.ML) #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.1606.00113

openalex publication_date 2016/06/01 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

Imaging genetic research has essentially focused on discovering unique and co-association effects, but typically ignoring to identify outliers or atypical objects in genetic as well as non-genetics variables. Identifying significant outliers is an essential and challenging issue for imaging genetics and multiple sources data analysis. Therefore, we need to examine for transcription errors of identified outliers. First, we address the influence function (IF) of kernel mean element, kernel covariance operator, kernel cross-covariance operator, kernel canonical correlation analysis (kernel CCA) and multiple kernel CCA. Second, we propose an IF of multiple kernel CCA, which can be applied for more than two datasets. Third, we propose a visualization method to detect influential observations of multiple sources of data based on the IF of kernel CCA and multiple kernel CCA. Finally, the proposed methods are capable of analyzing outliers of subjects usually found in biomedical applications, in which the number of dimension is large. To examine the outliers, we use the stem-and-leaf display. Experiments on both synthesized and imaging genetics data (e.g., SNP, fMRI, and DNA methylation) demonstrate that the proposed visualization can be applied effectively.

Citations

Related