Leveraging multi-source to resolve inconsistency across pharmacogenomic datasets in drug sensitivity prediction
Das, T.; Bhattarai, K.; Rajaganapathy, S.; Wang, L.; Cerhan, J. R.; Zong, N.
Show abstract
Pharmacogenomics datasets have been generated for various purposes, such as investigating different biomarkers. However, when studying the same cell line with the same drugs, differences in drug responses exist between studies. These variations arise from factors such as inter-tumoral heterogeneity, experimental standardization, and the complexity of cell subtypes. Consequently, drug response prediction suffers from limited generalizability. To address these challenges, we propose a computational model based on Federated Learning (FL) for drug response prediction. By leveraging three pharmacogenomics datasets (CCLE, GDSC2, and gCSI), we evaluate the performance of our model across diverse cell line-based databases. Our results demonstrate superior predictive performance compared to baseline methods and traditional FL approaches through various experimental tests. This study underscores the potential of employing FL to leverage multiple data sources, enabling the development of generalized models that account for inconsistencies among pharmacogenomics datasets. By addressing the limitations of low generalizability, our approach contributes to advancing drug response prediction in precision oncology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance 95%
- Creating a computer assisted ICD coding system: performance metric choice and use of the ICD hierarchy 94%
- ARCH: Large-scale Knowledge Graph via Aggregated Narrative Codified Health Records Analysis 94%
Similar papers in this journal
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 94%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.