vix.ing · top · new · best · stats · spec

Multidimensional mutual information methods for the analysis of\n covariation in multiple sequence alignments

2014/04/26 by Greg Clark, Sharon H. Ackerman, Clark, Greg W. +5
Biochemistry, Genetics and Molecular Biology · Materials Science · #Enzyme Structure and Function #FOS: Biological sciences #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #Protein Structure and Dynamics #Quantitative Methods (q-bio.QM)

paper · pdf · doi:10.48550/arxiv.1404.6684

openalex publication_date 2014/04/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Several methods are available for the detection of covarying positions from a\nmultiple sequence alignment (MSA). If the MSA contains a large number of\nsequences, information about the proximities between residues derived from\ncovariation maps can be sufficient to predict a protein fold. If the structure\nis already known, information on the covarying positions can be valuable to\nunderstand the protein mechanism.\n In this study we have sought to determine whether a multivariate extension of\ntraditional mutual information (MI) can be an additional tool to study\ncovariation. The performance of two multidimensional MI (mdMI) methods,\ndesigned to remove the effect of ternary/quaternary interdependencies, was\ntested with a set of 9 MSAs each containing <400 sequences, and was shown to be\ncomparable to that of methods based on maximum entropy/pseudolikelyhood\nstatistical models of protein sequences. However, while all the methods tested\ndetected a similar number of covarying pairs among the residues separated by <\n8 AA in the reference X-ray structures, there was on average less than 65%\noverlap between the top scoring pairs detected by methods that are based on\ndifferent principles.\n We have also attempted to identify whether the difference in performance\namong methods is due to different efficiency in removing covariation\noriginating from chains of structural contacts. We found that the reason why\nmethods that derive partial correlation between the columns of a MSA provide a\nbetter recognition of close contacts is not because they remove chaining\neffects, but because they filter out the correlation between distant residues\nthat originates from general fitness constraints. In contrast we found that\ntrue chaining effects are expression of real physical perturbations that\npropagate inside proteins, and therefore are not removed by the derivation of\npartial correlation between variables.\n

Citations

Related