Back

Leveraging Pretrained Large Language Model for Prognosis of Type 2 Diabetes Using Longitudinal Medical Records

Nguyen, P. B. H.; Menden, M. P.; Holl, R. W.; Hungele, A.

2025-02-05 health informatics
10.1101/2025.02.04.24313200 medRxiv
Show abstract

Timely prognosis of type 2 diabetes (T2D) is critical for effective interventions and reducing economic burden. Longitudinal medical records offer potential for extracting clinical insights but face challenges due to the sparse, high-dimensional data, data privacy, domain compatibility and interpretability issues. This study introduced PRIME-LLM, a framework that leverages the prediction power of pretrained large language models for disease prognosis. PRIME-LLM overcomes the challenges by synthetic data generation, missingness modeling, and a learnable embedding layer prepended to a pretrained LLM backbone. We finetuned and evaluated the model performance using a large real world dataset of 449,185 T2D patients. The PRIME-LLM-fine-tuned model outperformed baselines in forecasting HbA1c, LDL, blood pressure, improving MSE up to 12.8%. The model also demonstrated robust long-term prediction over 578.8 days (95% CI: [180,1155]). Integrated gradient analysis identified significant clinical features and visits, revealing potential biomarkers for early intervention. Overall, the results showed the possibility to leverage the prediction power of LLM in T2D prognosis using sparse medical time series, assisting clinical prognosis and biomarker discovery, ultimately advancing precision medicine.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.