Protein Sequence Domain Annotation using Language Models
Sarkar, A.; Krishnan, K.; Eddy, S. R.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWProtein function inference relies on annotating protein domains via sequence similarity, often modeled through profile Hidden Markov Models (profile HMMs), which capture evolutionary diversity within related domains. However, profile HMMs make strong simplifying independence assumptions when modeling residues in a sequence. Here, we introduce PSALM (Protein Sequence Annotation using Language Models), a hierarchical approach that relaxes these assumptions and uses representations of protein sequences learned by protein language models to enable high-sensitivity, high-specificity residue-level protein sequence annotation. We also develop the Multi-Domain Protein Homology Benchmark (MDPH-Bench), a benchmark for protein sequence domain annotation, where training and test sequences have been rigorously split to share no similarity between any of their domains at a given threshold of sequence identity. Prior benchmarks, which split one domain family at a time, do not support methods for annotating multi-domain proteins, where training and test sequences need to have multiple domains from different families. We validate PSALMs performance on MDPH-Bench and highlight PSALM as a promising alternative to HMMER, a state-of-the-art profile HMM-based method, for protein sequence annotation.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Prop3D: A Flexible, Python-based Platform for Machine Learning with Protein Structural Properties and Biophysical Data 95%
- DPCfam: a new method for unsupervised protein family classification 94%
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 93%
Similar papers in this journal
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 99%
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 92%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.