Back

Residual Multi-Modal Learning for Pan-Breast-Cancer Drug Response Prediction

Huang, B.; Tasaka, L.; Li, J.; Islam, T.; Zhang, S.

2026-07-08 bioinformatics
10.64898/2026.07.03.736239 bioRxiv
Show abstract

Predicting drug sensitivity across diverse cancer cell lines remains a fundamental challenge in precision oncology, particularly for data-scarce cell lines where per-cell-line models overfit and lookup-table approaches cannot generalise to unseen biological contexts. We present DL4DR, a Two Tower Residual Late Fusion deep learning model that addresses this challenge through content-based, identity-free genomic conditioning. The Cell Line Tower encodes each cell line as a 3 x 139 x 139 genomic image - encoding gene expression, mutation severity, and copy-number variation as RGB channels - using a convolutional encoder that maps directly from biological content, never from a cell line ID. The Compound Tower combines three complementary molecular representations: D-MPNN graph message passing, ORNN octave convolutional image features, and an ECFP hard-memorization head that preserves activity-cliff resolution. Predictions are composed as a residual sum: f = fhard + {lambda}(zc). fresidual, where the learned gate $\lambda$ modulates how much interaction signal supplements the memorization baseline. Evaluated across 51 breast cancer cell lines(136,342 records), Residual Fusion outperforms the ECFP-Only baseline in 48/51 cell lines (94.1%), with {Delta}R2 > 0.02 in 26/51 (51.0%). On the leave-cell-line-out split - the decisive test of genomic generalisation - the mean {Delta} R2 = 0.016 across all 51 lines demonstrates that the genomic encoder learns transferable biological signal beyond cell line identity. External validation on 601 cell lines across 27 cancer tissue types (CellTiter-Glo dataset; 0 cell line overlap with training) achieves median R2 = 0.627, within the range of the internal random-split performance (R2 = 0.61--0.69), confirming pan-cancer generalisation. GradCAM interpretability on the Cell Line Tower recovers TP53 among the top-five cross-cell-line genomic activators (5/51 cell lines) alongside several uncharacterised candidate genes (e.g.FSIP2, 6/51) - without any prior pathway annotation - providing partial biological validation of the learned representation, while also indicating that a substantial share of the encoder's top-ranked signal corresponds to genes with no current annotation as breast cancer drivers. Code and data are available at https://github.com/bayjuan5/DL4DR.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 18%
9.6%
2
Bioinformatics
1204 papers in training set
Top 4%
6.6%
3
Nature Machine Intelligence
70 papers in training set
Top 0.4%
6.2%
4
Cell Systems
201 papers in training set
Top 0.7%
6.2%
5
Genome Medicine
183 papers in training set
Top 0.6%
5.4%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.3%
7
Nature Methods
385 papers in training set
Top 2%
4.3%
8
Nature Genetics
286 papers in training set
Top 2%
4.2%
9
PLOS ONE
5266 papers in training set
Top 39%
3.1%
10
Nucleic Acids Research
1281 papers in training set
Top 6%
2.7%
50% of probability mass above
11
Genome Biology
637 papers in training set
Top 4%
2.7%
12
Nature Biotechnology
172 papers in training set
Top 2%
2.4%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.3%
14
Scientific Reports
3612 papers in training set
Top 51%
1.9%
15
Cancer Research
130 papers in training set
Top 2%
1.7%
16
Cell Reports Medicine
153 papers in training set
Top 2%
1.7%
17
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
18
Nature
645 papers in training set
Top 7%
1.5%
19
npj Precision Oncology
53 papers in training set
Top 1.0%
1.5%
20
Communications Biology
993 papers in training set
Top 19%
1.3%
21
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 36%
1.1%
22
npj Systems Biology and Applications
125 papers in training set
Top 2%
1.1%
23
Molecular Systems Biology
162 papers in training set
Top 2%
1.1%
24
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.6%
1.1%
25
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
26
npj Digital Medicine
118 papers in training set
Top 3%
1.0%
27
Science Advances
1243 papers in training set
Top 28%
1.0%
28
Communications Medicine
113 papers in training set
Top 4%
1.0%
29
iScience
1154 papers in training set
Top 30%
1.0%
30
Nature Biomedical Engineering
47 papers in training set
Top 1%
1.0%