AstraROLE2 & AstraSUIT2: Multi-Task Annotation Models for Functional Profiling of Proteins
Bozkurt, C.; Vasilyeva, A.; Goteti, A.
Show abstract
AbstractMost in-silico protein characterisation tools focus on only one aspect of protein function, forcing researchers to use multiple models or to bypass computational checks. Here we introduce AO_SCPLOWSTRAC_SCPLOWROLE2 and AO_SCPLOWSTRAC_SCPLOWSUIT2, two transformer-based, multi-task annotators that deliver an integrated functional profile in a single pass. A 1,351-dimensional input (ESM-2 CLS embeddings plus physicochemical Orbion enrichments) is mapped by a 512-unit encoder and task-specific linear heads: four in AO_SCPLOWSTRAC_SCPLOWROLE2 (EC class, GO term, molecular pathway, protein category) and nine in AO_SCPLOWSTRAC_SCPLOWSUIT2 (cofactor group, specific cofactor, domain, host, membrane association type, transmembrane helix number, subcellular localization, quaternary category, quaternary stoichiometry). Models were trained on 730k UniProt proteins with stratified 70/15/15 splits; class-weighted BCE and Optuna hyper-parameter search countered imbalance. On hold-out sets the heads reached macro F1=0.84-0.98 and MCC=0.85-0.98. Highest scores were seen for cofactor binding (0.98), membrane association type (F1=0.97) and top-level EC number (0.96); GO term classification was hardest (0.85). Against recent comparators (incl. DeepGOPlus and TargetP 2.0), the Astra models matched or exceeded performance, especially on metal-ion binding and cofactor binding. Additional tests on three novel proteins not included in initial dataset showed good predictions for most labels, underscoring the potential for hypothesis generation. Overall, AO_SCPLOWSTRAC_SCPLOWROLE2 and AO_SCPLOWSTRAC_SCPLOWSUIT2 supplied fast, state-of-the-art multi-label protein annotation within one unified model network.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LMPred: Predicting Antimicrobial Peptides Using Pre-Trained Language Models and Deep Learning 95%
- ProteinPrompt: a webserver for predicting protein-protein interactions 94%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 94%
Similar papers in this journal
- Protein language models can capture protein quaternary state 95%
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.