Benchmarking Chemical, Genetic, and Cell Line Encodings for Cancer Perturbation Response Prediction
Zinchenko, V.; Schlicker, A.; Kurilov, R.; Pouplin, A.; Horlacher, M.
Show abstract
Estimating the response of tumor cells to specific perturbations is crucial for identifying effective treatments that selectively target cancer cells while sparing healthy ones, enabling personalized medicine approaches. Large-scale initiatives, such as DepMap, have profiled cancer cell line responses to various drug treatments and gene knockouts, facilitating the development of computational models that predict sensitivity of cancer cells to different perturbations. Existing models utilize diverse methods for encoding perturbations, including various chemical fingerprints and types of gene-gene relationships. They also rely on different architectures and are often trained on distinct datasets. This variability makes it unclear which chemical, genetic, or cell line encoding is most informative for predicting cancer cell viability following perturbation treatment. To address this gap, we systematically evaluated various approaches to encode chemical and genetic perturbations and cell lines on the tasks of predicting cell viability and gene dependency. We found that for genetic perturbations, STRING-based encodings yield the highest performance, considerably outperforming GO-term and protein language model based encodings, which showed promising results in previous perturbation prediction studies. For chemical perturbations, while most encoders showed comparable performance, those pre-trained on other bio-assay data yielded the highest performance. Finally, we found that for cell line encodings, raw gene expression features outperformed more sophis-ticated approaches, such as transcriptomics foundation model embeddings, as well as genotype-based encodings. Together, our results identify promising approaches for encoding chemical and genetic perturbations and enable virtual screening for perturbations with selective toxicity.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Discovering Governing Equations of Biological Systems through Representation Learning and Sparse Model Discovery 96%
- Prediction of G4 formation in live cells with epigenetic data: a deep learning approach 94%
- A computational method for direct imputation of cell type-specific expression profiles and cellular compositions from bulk-tissue RNA-Seq in brain disorders 94%
Similar papers in this journal
- Discovering differential genome sequence activity with interpretable and efficient deep learning 95%
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 94%
- Prioritizing and characterizing functionally relevant genes across human tissues 94%
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 96%
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.