Human-in-the-loop approach to identify functionally important residues of proteins from literature
Vollmar, M.; Tirunagari, S.; Harrus, D.; Armstrong, D.; Gaborova, R.; Gupta, D.; Afonso, M. Q. L.; Evans, G. L.; Velankar, S.
Show abstract
We present a novel system that leverages curators in the loop to develop a dataset and model for detecting residue-level functional annotations and other protein structure features from standard publication text. Our approach involves the integration of data from multiple resources, including PDBe, EuropePMC, PubMedCentral, and PubMed, combined with annotation guidelines from UniProt, while employing LitSuggest and Huggingface models as tools in the annotation process. A team of seven annotators manually curated ten articles for named entities, which we utilized to train a starting PubmedBert model from Huggingface. Using a human-in-the-loop annotation system, we developed the best model with commendable performance metrics of 0.90 for precision, 0.92 for recall, and 0.91 for F1-measure. Our proposed system showcases a successful synergy of machine learning techniques and human expertise in curating a dataset for residue-level functional annotations and protein structure features. The results demonstrate the potential for broader applications in protein research, bridging the gap between advanced machine learning models and the indispensable insights of domain experts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 94%
- APICURON: a database to credit and acknowledge the work of biocurators 93%
- PDB NextGen Archive: Centralising Access to Integrated Annotations and Enriched Structural Information by the Worldwide Protein Data Bank 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.