Back

PGxRAG: A Retrieval Augmented Generation supported Pharmacogenomics Assistant

Borishetty, D.; Banda, P.; Andhi, N.; Desai, A.; Uppili, B.; Samudrala, N. M.; Palanki, S. S.; Bobbili, D. R.; Iyer, G. R.

2025-09-29 health informatics
10.1101/2025.09.24.25336524 medRxiv
Show abstract

Pharmacogenomics enables personalized medicine by predicting individual drug responses based on genetic makeup, but complex guideline retrieval remains challenging for clinicians, particularly in resource-limited settings. While large language models (LLMs) show promise for numerous healthcare applications, their performance on domain-specific pharmacogenomics queries without expert knowledge integration remains limited. We evaluated whether Retrieval-Augmented Generation (RAG) enhancement improves LLM accuracy for pharmacogenomics applications compared to native model performance. We conducted comparative evaluation of four LLMs with and without RAG enhancement, constructing a knowledge base from PharmGKB (now ClinPGx), CPIC, Dutch Pharmacogenetics Working Group (KNMP), and FDA guidelines containing 2,617 embedded document chunks. We developed 225 multiple-choice questions representing patient and healthcare provider perspectives, then systematically evaluated hyperparameter combinations testing different temperatures, embedding dimensions, retrieval methods, and k-values with different sample sizes. RAG-enhanced models consistently outperformed native LLMs, with optimal configuration (GPT-4o) achieving 95.1% accuracy compared to 89.8% for the same native model. The RAG approach significantly enhances LLM performance in pharmacogenomics applications, providing a scalable solution for making complex pharmacogenomic guidelines accessible to healthcare providers while maintaining high clinical decision support accuracy.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.