TY - JOUR
T1 - HMMSTR
T2 - A hidden Markov model for local sequence-structure correlations in proteins
AU - Bystroff, Christopher
AU - Thorsson, Vesteinn
AU - Baker, David
N1 - Funding Information:
We wish to thank Phil Green, Richard M. Karp, Anders Krogh, Chip Lawrence, Ingo Ruczinski, and Ed Thayer for helpful discussions, and a referee for pointing out an error in an earlier version of this manuscript. This work was supported by a University of Washington Training Grant in Interdisciplinary Genome Sciences (V.T.), by a Sloan Foundation/Department of Energy Fellowship in Computational Molecular Biology(V.T.), by the NSF (STC cooperative agreement BIR-9214821, D.B.), and by a Packard Fellowship in Science and Engineering (D.B.) and a grant from the Howard Hughes Medical Institute to the Rensselaer Bioinformatics Program (C.B.).
PY - 2000/8/4
Y1 - 2000/8/4
N2 - We describe a hidden Markov model, HMMSTR, for general protein sequence based on the I-sites library of sequence-structure motifs. Unlike the linear hidden Markov models used to model individual protein families, HMMSTR has a highly branched topology and captures recurrent local features of protein sequences and structures that transcend protein family boundaries. The model extends the I-sites library by describing the adjacencies of different sequence-structure motifs as observed in the protein database and, by representing overlapping motifs in a much more compact form, achieves a great reduction in parameters. The HMM attributes a considerably higher probability to coding sequence than does an equivalent dipeptide model, predicts secondary structure with an accuracy of 74.3 %, backbone torsion angles better than any previously reported method and the structural context of β strands and turns with an accuracy that should be useful for tertiary structure prediction. (C) 2000 Academic Press.
AB - We describe a hidden Markov model, HMMSTR, for general protein sequence based on the I-sites library of sequence-structure motifs. Unlike the linear hidden Markov models used to model individual protein families, HMMSTR has a highly branched topology and captures recurrent local features of protein sequences and structures that transcend protein family boundaries. The model extends the I-sites library by describing the adjacencies of different sequence-structure motifs as observed in the protein database and, by representing overlapping motifs in a much more compact form, achieves a great reduction in parameters. The HMM attributes a considerably higher probability to coding sequence than does an equivalent dipeptide model, predicts secondary structure with an accuracy of 74.3 %, backbone torsion angles better than any previously reported method and the structural context of β strands and turns with an accuracy that should be useful for tertiary structure prediction. (C) 2000 Academic Press.
KW - Clustering
KW - Hidden Markov models
KW - I-sites library
KW - Motifs
KW - Sequence patterns
UR - https://www.scopus.com/pages/publications/0034604368
U2 - 10.1006/jmbi.2000.3837
DO - 10.1006/jmbi.2000.3837
M3 - Article
C2 - 10926500
AN - SCOPUS:0034604368
SN - 0022-2836
VL - 301
SP - 173
EP - 190
JO - Journal of Molecular Biology
JF - Journal of Molecular Biology
IS - 1
ER -