A Pharmacogenomic-Informed Representation Improves Multimodal EHR Survival Prediction
Lee, M. H.; Xiao, Y.; Li, X.; Klee, E.; Yang, P.; Sio, T.; Wang, L.; Cerhan, J. R.; Zong, N.
Show abstract
BackgroundElectronic health record (EHR)-based prognostic modeling is increasingly used in oncology, yet incorporating pharmacogenomic (PGx) knowledge derived from experimental systems into clinical prediction frameworks remains challenging. This gap is driven by fundamental mismatches between controlled drug-mutation assays and heterogeneous, incomplete real-world clinical data. MethodsWe propose a representation transfer framework that integrates PGx embeddings learned from large-scale in vitro pharmacogenomic screens into patient-level EHR models. A frozen pharmacogenomic encoder is used to generate interaction-aware embeddings from patient mutation profiles and administered therapies, which are aggregated into a fixed-length PGx Complementarity Representation. These representations are incorporated into multimodal survival prediction models alongside standard clinical features. Performance was evaluated using systematic modality ablation analyses, attribution analyses, and exploratory unsupervised representation analyses. ResultsIntegrating PGx embeddings yielded consistent performance improvements across all evaluated modality combinations. Relative gains were largest in modality-sparse settings, where baseline EHR features encode limited biological context, and were attenuated--but remained significant--in biologically enriched configurations. Attribution analyses indicated that PGx embeddings contributed non-redundant predictive signal beyond standard clinical features. Exploratory unsupervised analyses further demonstrated that the learned representations exhibit interpretable association patterns aligned with known therapeutic exposures and pathway-level associations. ConclusionThese findings suggest that externally learned pharmacogenomic representations can be transferred into real-world EHR models as a context-dependent, non-redundant augmentation. By framing PGx knowledge as an interaction-aware representation rather than a mechanistic model, this work provides an informatics framework for integrating experimental pharmacogenomic data into clinical prediction tasks in a reproducible and interpretable manner.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
- Biologically-informed deep neural networks provide quantitative assessment of intratumoral heterogeneity in post-treatment glioblastoma 94%
- FedWeight: Mitigating Covariate Shift of Federated Learning on Electronic Health Records Data through Patients Re-weighting 94%
Similar papers in this journal
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 93%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 92%
Similar papers in this journal
- Deep Learning identifies new morphological patterns of Homologous Recombination Deficiency in luminal breast cancers from whole slide images. 93%
- Pancreatic cancer risk prediction using deep sequential modeling of longitudinal diagnostic and medication records 92%
- Estimating the Treatment Effects of Multiple Drug Combinations on Multiple Outcomes in Hypertension 92%
Similar papers in this journal
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 93%
- Obtaining Spatially Resolved Tumor Purity Maps Using Deep Multiple Instance Learning In A Pan-cancer Study 93%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.