Leveraging Pretrained Large Language Model for Prognosis of Type 2 Diabetes Using Longitudinal Medical Records
Nguyen, P. B. H.; Menden, M. P.; Holl, R. W.; Hungele, A.
Show abstract
Timely prognosis of type 2 diabetes (T2D) is critical for effective interventions and reducing economic burden. Longitudinal medical records offer potential for extracting clinical insights but face challenges due to the sparse, high-dimensional data, data privacy, domain compatibility and interpretability issues. This study introduced PRIME-LLM, a framework that leverages the prediction power of pretrained large language models for disease prognosis. PRIME-LLM overcomes the challenges by synthetic data generation, missingness modeling, and a learnable embedding layer prepended to a pretrained LLM backbone. We finetuned and evaluated the model performance using a large real world dataset of 449,185 T2D patients. The PRIME-LLM-fine-tuned model outperformed baselines in forecasting HbA1c, LDL, blood pressure, improving MSE up to 12.8%. The model also demonstrated robust long-term prediction over 578.8 days (95% CI: [180,1155]). Integrated gradient analysis identified significant clinical features and visits, revealing potential biomarkers for early intervention. Overall, the results showed the possibility to leverage the prediction power of LLM in T2D prognosis using sparse medical time series, assisting clinical prognosis and biomarker discovery, ultimately advancing precision medicine.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning for Real-Time Aggregated Prediction of Hospital Admission for Emergency Patients 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 94%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 93%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 93%
Similar papers in this journal
- Precision Diagnostic Approach to Predict 5-year Risk for Microvascular Complications in Type 1 Diabetes 92%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 92%
- 1 H-NMR metabolomics-based surrogates to impute common clinical risk factors and endpoints 91%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 95%
- Explainable AI-based analysis of human pancreas sections identifies traits of type 2 diabetes 93%
- Deep transfer learning for reducing health care disparities arising from biomedical data inequality 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.