Instruction-tuned extraction of virus-host interactions from integrated scientific evidence
Zhang, Z.; Li, W.; Ren, J.; Chen, Y.; He, B.; Sun, Q.; Wang, H.
Show abstract
MotivationViral infectious diseases continue to pose a major threat to global health. Understanding protein-protein interactions (PPIs) and RNA-protein interactions (RPIs) between viruses and hosts is essential for elucidating infection mechanisms. However, manual curation of these interactions from biomedical literature is inefficient, creating a pressing need for automated and scalable extraction methods. Large language models (LLMs), such as the generative pre-trained transformer (GPT) and bidirectional encoder representations from transformers (BERT), offer promising solutions. Yet, most existing datasets focus on abstracts, overlooking other information-rich sections. We aim to develop a data-efficient approach to extract virus-host interaction (VHI) entities from full-text biomedical articles, including Results, Methods and tables. To our knowledge, this is the first study to apply instruction tuning to full-text VHI extraction ResultsWe curated a dataset containing 3, 395 PPI and 674 RPI entities from the Results, Materials and Methods sections, along with 566 PPIs and 793 RPIs from tables. Under low-resource conditions (<500 training examples), our instruction-tuned ChatMed-VHI model achieves the best overall performance (F1: 89.7%, Precision: 95.3%), outperforming PubMedBERT (F1: 74.6%, Precision: 75.1%). When scaled to the full dataset (>4, 000 training examples), ChatMed-VHI maintained the highest overall performance, while PubMedBERT achieved slightly higher precision (92.3% vs. 91.3%). Notably, ChatMed-VHI improved F1 and recall by 2.79% but precision dropped by 4.20% with more training data, whereas PubMedBERT improved consistently across all metrics. These results demonstrate the effectiveness of instruction-tuned LLMs for full-text biomedical extraction tasks, and position ChatMed-VHI as a scalable, domain-adaptable solution for VHI mining.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GRU-SCANET: Unleashing the Power of GRU-based Sinusoidal CApture Network for Precision-driven Named Entity Recognition 95%
- Gilda: biomedical entity text normalization with machine-learned disambiguation as a service 94%
- CoNECo: A Corpus for Named Entity recognition and normalization of protein Complexes 94%
Similar papers in this journal
- An Analysis of Protein Language Model Embeddings for Fold Prediction 94%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 93%
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 96%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 95%
- Discovering nuclear localization signal universe through a novel deep learning model with interpretable attention units 93%
Similar papers in this journal
- Learning universal knowledge graph embedding for predicting biomedical pairwise interactions 94%
- DeepGene: An Efficient Foundation Model for Genomics based on Pan-genome Graph Transformer 93%
- GAT-HiC: Efficient Reconstruction of 3D Chromosome Structure via Residual Graph Attention Neural Networks 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.