Large language models improve transferability of electronic health record-based predictions across countries and coding systems
Kirchler, M.; Ferro, M.; Lorenzini, V.; FinnGen, ; Lippert, C.; ganna, a.
Show abstract
Variation in medical practices and reporting standards across healthcare systems limits the transferability of prediction models based on structured electronic health record (EHR) data. We introduce GRASP, a novel transformer-based architecture that enhances the generalizability of EHR-based prediction by embedding medical codes into a unified semantic space using a large language model. We applied GRASP to predict the onset of 21 diseases and all-cause mortality in over one million individuals from UK Biobank (UK), FinnGen (Finland) and Mount Sinai (USA), all harmonized to OMOP common data model. Trained on the UK Biobank and evaluated in FinnGen and Mount Sinai, GRASP achieved an average {Delta}C-index that was 83% and 35% higher than language-unaware models, respectively. GRASP also showed significantly higher correlations with polygenic risk scores for 62% of diseases. Notably, GRASP mantained robust performance even when datasets were not harmonized to the same data model, accurately predicting disease risk from ICD-10-CM codes without direct mappings to OMOP. GRASP enables accurate and transferable disease predictions across heterogeneous healthcare systems with minimal resource requirements.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 94%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 93%
- Multi-ancestry omic Mendelian randomization revealing putative drug targets of COVID-19 severity 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.