Database-Augmented Transformer-Based Large Language Models Achieve High Accuracy in Mapping Gene-Phenotype Relationships
Suhardi, V.; Suhardi, N.; Oktarina, A.; Bostrom, M.; Yang, X.
Show abstract
Transformer-based large language models (LLMs) have demonstrated significant potential in the biological and medical fields due to their ability to effectively learn from large-scale, diverse datasets and perform a wide range of downstream tasks. However, LLMs are limited by issues such as information processing inaccuracies and data confabulation, which hinder their utility for literature searches and other tasks requiring accurate and comprehensive extraction of information from extensive scientific literature. In this study, we evaluated the performance of various LLMs in accurately retrieving peer-reviewed literature and mapping correlations between 102 genes and four phenotypes: bone formation, cartilage formation, fibrosis, and cell proliferation. Our analysis included standard transformer-based LLMs (ChatGPT4o and Gemini1.5 Pro), fine-tuned LLMs with dedicated custom databases containing peer-reviewed articles (SciSpace and ScholarAI), and fine-tuned LLMs without dedicated databases (PubMedGPT and ScholarGPT). Using human-curated gene-to-phenotype mappings as the ground truth, we found that fine-tuned LLMs with dedicated databases (SciSpace and ScholarAI) achieved high accuracy (>80%) in gene-to-phenotype mapping. Additionally, these models were able to provide relevant peer-reviewed publications supporting each gene-to-phenotype correlation. These findings underscore the importance of database augmentation and finetuning in enhancing the reliability and utility of LLMs for biomedical research applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural language processing of gene descriptions for overrepresentation analysis with GeneTEA 92%
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 91%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 91%
Similar papers in this journal
- Citation needed?Wikipedia bibliometrics during the first waveof the COVID-19 pandemic. 93%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 92%
- TooManyCellsInteractive: a visualization tool for dynamic exploration of single-cell data 91%
Similar papers in this journal
- A natural language processing system for the efficient extraction of cell markers 94%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 92%
- Network and pathway expansion of genetic disease associations identifies successful drug targets 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.