Enhancing patient representation learning from electronic health records through predicted family relations
Huang, X.; Arora, J.; Erzurumluoglu, A. M.; Lam, D.; Boehringer Ingelheim Global Computational Biology and Digital Sciences, ; Zhao, H.; Ding, Z.; Wang, Z.; de Jong, J.
Show abstract
Artificial intelligence and machine learning are powerful tools in analyzing electronic health records (EHRs) for healthcare research. Despite the recognized importance of family health history, in healthcare research individual patients are often treated as independent samples, overlooking family relations. To address this gap, we present ALIGATEHR, which models predicted family relations in a graph attention network and integrates this information with a medical ontology representation. Taking disease risk prediction as a use case, we demonstrate that explicitly modeling family relations significantly improves predictions across the disease spectrum. We then show how ALIGATEHRs attention mechanism, which links patients disease risk to their relatives clinical profiles, successfully captures genetic aspects of diseases using only EHR diagnosis data. Finally, we use ALIGATHER to successfully distinguish the two main inflammatory bowel disease subtypes (Crohns disease and ulcerative colitis), illustrating its great potential for improving patient representation learning for predictive and descriptive modeling of EHRs.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 96%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 94%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.