2018/07/31 by Anne‐Florence Bitbol, Anne-Florence Bitbol · 2 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · Physics and Astronomy · #Artificial intelligence #Bioinformatics and Genomic Networks #Biology #Coevolution #Computational biology #Computer science #Entropy (arrow of time) #Entropy maximization #Evolutionary biology #Gene #Genetics #Inference #Interaction information #Mathematics #Microbial Metabolic Engineering and Bioproduction #Multiple sequence alignment #Mutual information #Pairwise comparison #Peptide sequence #Physics #Principle of maximum entropy #Protein Structure and Dynamics #Protein family #Protein sequencing #Protein structure #Protein–protein interaction #Sequence (biology) #Sequence alignment #Statistics #physics.bio-ph #q-bio.BM
paper · pdf · doi:10.1371/journal.pcbi.1006401
published as PLoS Comput. Biol. 14(11): e1006401 (2018) · 26 pages, 11 figures, published version
arxiv created 2018/11/13 · openalex publication_date 2018/11/13 · arxiv updated 2018/11/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Functional protein-protein interactions are crucial in most cellular processes. They enable multi-protein complexes to assemble and to remain stable, and they allow signal transduction in various pathways. Functional interactions between proteins result in coevolution between the interacting partners, and thus in correlations between their sequences. Pairwise maximum-entropy based models have enabled successful inference of pairs of amino-acid residues that are in contact in the three-dimensional structure of multi-protein complexes, starting from the correlations in the sequence data of known interaction partners. Recently, algorithms inspired by these methods have been developed to identify which proteins are functional interaction partners among the paralogous proteins of two families, starting from sequence data alone. Here, we demonstrate that a slightly higher performance for partner identification can be reached by an approximate maximization of the mutual information between the sequence alignments of the two protein families. Our mutual information-based method also provides signatures of the existence of interactions between protein families. These results stand in contrast with structure prediction of proteins and of multi-protein complexes from sequence data, where pairwise maximum-entropy based global statistical models substantially improve performance compared to mutual information. Our findings entail that the statistical dependences allowing interaction partner prediction from sequence data are not restricted to the residue pairs that are in direct contact at the interface between the partner proteins.