Back

Beyond profiles: supervised repeat annotation using protein embeddings

Qiu, K.; Ludwiczak, J.; Lupas, A. N.; Dunin-Horkawicz, S.

2026-05-20 bioinformatics
10.64898/2026.05.19.725729 bioRxiv
Show abstract

Repeated sequence motifs give rise to diverse protein structures and functions, yet their detection is often challenging due to weak sequence similarity between repeat units. Most sensitive approaches rely on homology information directly represented by alignments or derived profiles, thereby limiting flexibility and scalability. Here, we introduce TREAD (Transfer learning-based REpeat Annotation using Protein EmbeDdings), a supervised framework that reformulates repeat detection as a residue-level annotation problem and operates directly on embeddings from protein language models. Instead of constructing explicit probabilistic profiles, TREAD learns repeat-specific features implicitly, enabling residue-resolved scoring and flexible repeat segment extraction. Across complementary benchmarks based on RepeatsDB and Pfam, TREAD consistently matches or outperforms the widely used profile-based tool HMMER, particularly in low-data and high-divergence settings. The model exhibits robustness to score thresholding and demonstrates better generalization to independent test sets, including those containing remote homologs. To illustrate the practical utility of TREAD, we applied it to survey {beta}-propeller proteins across the AlphaFold Database and representative proteomes to generate a comprehensive census of this fold. This analysis highlights extensive propeller diversity, identifies lineage-specific expansion patterns across the tree of life, and suggests previously unrecognized relationships between propellers and other repeat folds. Together, TREAD provides a flexible and scalable alternative to profile-based repeat annotation and establishes a general motif-centric framework for protein sequence annotation.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.