Transport-based transfer learning on Electronic Health Records: Application to detection of treatment disparities
Li, W.; Park, Y.; Dao Duc, K.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWElectronic Health Records (EHRs) sampled from different populations can introduce unwanted bi-ases, limit individual-level data sharing, and make the data and fitted model hardly transferable across different population groups. In this context, our main goal is to design an effective method to transfer knowledge between population groups, with computable guarantees for suitability, and that can be applied to quantify treatment disparities. For a model trained in an embedded feature space of one subgroup, our proposed framework Optimal Transport-based Transfer Learning for EHRs (OT-TEHR) combines feature embedding of the data and unbalanced optimal transport (OT) for domain adaptation to another population group. To test our method, we processed and divided the MIMIC-III and MIMIC-IV databases into multiple population groups using ICD codes and multiple labels. We derive a theoretical bound for the generalization error of our method, and interpret it in terms of the Wasserstein distance, unbalancedness between the source and target domains, and labeling divergence, which can be used as a guide for assessing the suitability of binary classification and regression tasks. In general, our method achieves better accuracy and computational efficiency compared to standard and machine learning transfer learning methods on various tasks. Upon testing our method for populations with different insurance plans, we detect various levels of disparities in hospital duration stay between groups. By leveraging tools from OT theory, our proposed frame-work allows to compare statistical models on EHR data between different population groups. As a potential application for clinical decision making, we quantify treatment disparities between different population groups. Future directions include applying OTTEHR to broader regression and classification tasks and extending the method to semi-supervised learning. Data and Code AvailabilityThis paper uses the MIMIC-III dataset [Johnson et al., 2016], which is available on the PhysioNet repository [Moody et al., 2001]. The anonymized code repository is available at this link. Institutional Review Board (IRB)This research does not require IRB approval.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 96%
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 96%
- AITL: Adversarial Inductive Transfer Learning with input and output space adaptation for pharmacogenomics 96%
Similar papers in this journal
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 95%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 95%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.