2012/06/27 by Hanhuai Shan, Shan, Hanhuai, Jens Kattge +9
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · #Applications (stat.AP) #Computational Engineering #FOS: Computer and information sciences #Finance #Gene expression and cancer classification #Genetic Mapping and Diversity in Plants and Animals #Genomics and Phylogenetic Studies #Machine Learning (cs.LG) #Plant and animal studies #and Science (cs.CE)
paper · pdf · doi:10.48550/arxiv.1206.6439
openalex publication_date 2012/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Plant traits are a key to understanding and predicting the adaptation of\necosystems to environmental changes, which motivates the TRY project aiming at\nconstructing a global database for plant traits and becoming a standard\nresource for the ecological community. Despite its unprecedented coverage, a\nlarge percentage of missing data substantially constrains joint trait analysis.\nMeanwhile, the trait data is characterized by the hierarchical phylogenetic\nstructure of the plant kingdom. While factorization based matrix completion\ntechniques have been widely used to address the missing data problem,\ntraditional matrix factorization methods are unable to leverage the\nphylogenetic structure. We propose hierarchical probabilistic matrix\nfactorization (HPMF), which effectively uses hierarchical phylogenetic\ninformation for trait prediction. We demonstrate HPMF's high accuracy,\neffectiveness of incorporating hierarchical structure and ability to capture\ntrait correlation through experiments.\n