Metapredict enables accurate disorder prediction across the Tree of Life
Lotthammer, J. M.; Hernandez-Garcia, J.; Griffith, D.; Weijers, D.; Holehouse, A. S.; Emenecker, R. J.
Show abstract
Intrinsically disordered proteins and protein regions (collectively IDRs) are critical in numerous cellular processes. To understand how IDRs facilitate function, we need tools to accurately and rapidly identify them from sequence. While many methods for disorder prediction exist, we are currently limited by throughput and accuracy for evolutionary scale analyses. To bridge this gap, we developed metapredict V3, an updated version of our disorder predictor that enables evolutionary-scale disorder prediction. Metapredict V3 enables proteome-scale prediction with state-of-the-art accuracy in seconds and was developed with a focus on usability. It is distributed as a web server, Python software package, command-line interface, and Google Colab notebook. Here, we leverage the accuracy and throughput of metapredict V3 to predict disorder for over 20,000 proteomes to evaluate the prevalence of disorder across the kingdoms of life.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 95%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 95%
Similar papers in this journal
Similar papers in this journal
- The amino acid sequence determines protein abundance through its conformational stability and reduced synthesis cost. 96%
- Learning deep representations of enzyme thermal adaptation 95%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.