Back

Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval

Urokov, R.; Khan, A.; Eshboyev, F.; Asadov, D.; Rahman, S.; Kushokova, D.

2026-08-22 bioinformatics
10.64898/2026.08.21.746127 bioRxiv
Show abstract

Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired inference. Of 6,953 atlas BGCs from 182 deduplicated Streptomyces griseus genome accessions, 5,325 silver-labeled BGCs are split into 98 training, 21 validation, and 21 test reference groups. Of the test reference groups, 16 are eligible for retrieval diagnostics. Pfam Jaccard scored Recall@50 of 0.8788, while Pfam-augmented BGC-SetNet scored 0.8472. The combination of ESM and Pfam-augmented BGC-SetNet scored 0.8769. A weighted Pfam Jaccard obtained a slightly higher score of 0.8789, which has a negligible difference compared to unweighted Pfam accard. Our results do not support the claim that sequence-derived representations can recover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the major signal for this objective. Our results de[fi]ne the curation and pathway-level validation processes that are necessary for a more robust biological test.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.