Fine-tuning BERT models to extract transcriptional regulatory interactions of bacteria from biomedical literature
Varela-Vega, A.; Posada-Reyes, A.-B.; Mendez-Cruz, C.-F.
Show abstract
AbstractCuration of biomedical literature has been the traditional approach to extract relevant biological knowledge; however, this is time-consuming and demanding. Recently, Large language models (LLMs) based on pre-trained transformers have addressed biomedical relation extraction tasks outperforming classical machine learning approaches. Nevertheless, LLMs have not been used for the extraction of transcriptional regulatory interactions between transcription factors and regulated elements (genes or operons) of bacteria, a first step to reconstruct a transcriptional regulatory network (TRN). These networks are incomplete or missing for many bacteria. We compared six state-of-the-art BERT architectures (BERT, BioBERT, BioLinkBERT, BioMegatron, BioRoBERTa, LUKE) for extracting this type of regulatory interactions. We fine-tuned 72 models to classify sentences in four categories: activator, repressor, regulator, and no relation. A dataset of 1562 sentences manually curated from literature of Escherichia coli was utilized. The best model of LUKE architecture obtained a relevant performance in the evaluation dataset (Precision: 0.8601, Recall: 0.8788, F1-Score Macro: 0.8685, MCC: 0.8163). An examination of model predictions revealed that the model learned different ways to express the regulatory effect. The model was applied to reconstruct a TRN of Salmonella Typhimurium using 264 complete articles. We were able to accurately reconstruct 82% of the network. A network analysis confirmed that the transcription factor PhoP regulated many genes (uppermost degree), some of them responsible for antimicrobial resistance. Our work is a starting point to address the limitations of curating regulatory interactions, especially for the reconstruction of TRNs of bacteria or diseases of biological interest.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 95%
- Designing of thermostable proteins with a desired melting temperature 94%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 94%
Similar papers in this journal
- Relation extraction between bacteria and biotopes from biomedical texts with attention mechanisms and domain-specific contextual representations 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.