vix.ing · top · new · best · stats · spec

Fast pseudolikelihood maximization for direct-coupling analysis of protein structure from many homologous amino-acid sequences

2014/01/20 by Magnus Ekeberg, Tuomo Hartonen, Erik Aurell · 3 citations
Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #q-bio.QM #physics.comp-ph #physics.data-an

paper · pdf · doi:10.1016/j.jcp.2014.07.024

published as Journal of Computational Physics 276 (2014) 341-356 · 33 pages, 4 figures; M. Ekeberg and T. Hartonen are joint first authors; code and supplementary information on http://plmdca.csc.kth.se/

arxiv created 2014/01/20 · arxiv updated 2014/09/16

Abstract

Direct-Coupling Analysis is a group of methods to harvest information about coevolving residues in a protein family by learning a generative model in an exponential family from data. In protein families of realistic size, this learning can only be done approximately, and there is a trade-off between inference precision and computational speed. We here show that an earlier introduced l2-regularized pseudolikelihood maximization method called plmDCA can be modified as to be easily parallelizable, as well as inherently faster on a single processor, at negligible difference in accuracy. We test the new incarnation of the method on 148 protein families from the Protein Families database (PFAM), one of the largest tests of this class of algorithms to date.

Cited by