Enhancing Gene Expression Representation and Drug Response Prediction with Data Augmentation and Gene Emphasis
Lu, D.; Pamar, D. P. S.; Ohnmacht, A. J.; Kutkaite, G.; Menden, M. P.
Show abstract
Representation learning for tumor gene expression (GEx) data with deep neural networks is limited by the large gene feature space and the scarcity of available clinical and preclinical data. The translation of the learned representation between these data sources is further hindered by inherent molecular differences. To address these challenges, we propose GExMix (Gene Expression Mixup), a data augmentation method, which extends the Mixup concept to generate training samples accounting for the imbalance in both data classes and data sources. We leverage the GExMix-augmented training set in encoder-decoder models to learn a GEx latent representation. Subsequently, we combine the learned representation with drug chemical features in a dual-objective enhanced gene-centric drug response prediction, i.e., reconstruction of GEx latent embeddings and drug response classification. This dual-objective design strategically prioritizes gene-centric information to enhance the final drug response prediction. We demonstrate that augmenting training samples improves the GEx representation, benefiting the gene-centric drug response prediction model. Our findings underscore the effectiveness of our proposed GExMix in enriching GEx data for deep neural networks. Moreover, our proposed gene-centricity further improves drug response prediction when translating preclinical to clinical datasets. This highlights the untapped potential of the proposed framework for GEx data analysis, paving the way toward precision medicine.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Looking at the BiG picture: Incorporating bipartite graphs in drug response prediction 97%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 97%
- AITL: Adversarial Inductive Transfer Learning with input and output space adaptation for pharmacogenomics 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.