Back

Domain Adaptation Strategies for Transformer-Based Disease Prediction using Electronic Health Records

Driever, J.; Lentzen, M.; Madan, S.; Fröhlich, H.

2025-06-02 health informatics
10.1101/2025.06.02.25328621 medRxiv
Show abstract

Electronic Health Records (EHRs) offer rich data for machine learning, but model generalizability across institutions is hindered by statistical and coding biases. This study investigates domain adaptation (DA) techniques to improve model transfer, focusing on the Ex-Med-BERT transformer architecture for structured EHR data. We compare supervised and un-supervised DA methods in transferring predictive capabilities from the large-scale IBM Explorys database (U.S.) to the UK Biobank. Results across six clinical endpoints show that DA methods outperform fine-tuning, especially with limited target data. These findings emphasize selecting DA strategies based on target data availability and the benefit of incorporating source domain data for robust adaptation.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.