Back

MiT4SL: multi-omics triplet representation learning for cancer cell line-adapted prediction of synthetic lethality

Tao, S.; Feng, Y.; Yang, Y.; Wu, M.; Zheng, J.

2025-04-25 bioinformatics
10.1101/2025.04.20.649694 bioRxiv
Show abstract

Synthetic lethality (SL) offers a promising approach for targeted cancer therapies. Current SL prediction models heavily rely on extensive labeled data for specific cell lines to accurately identify SL pairs. However, a major limitation is the scarcity of SL labels across most cell lines, which makes it challenging to predict SL pairs for target cell lines with limited or even no available labels in real-world scenarios. Furthermore, gene interactions could be opposite between training and test cell lines, i.e. SL vs. non-SL, which further aggravates the challenge of generalization among cell lines. A promising strategy is to transfer knowledge learned from cell lines with relatively abundant SL labels to those with limited SL labels for the discovery of novel SL pairs, i.e., cell line-adapted SL prediction. Here, we propose MiT4SL, a multi-omics triplet representation learning model for cell line-adapted SL prediction. The core idea of MiT4SL is to model cell lineage information as embeddings, which are generated by combining a protein-protein interaction network representation tailored to each cell line with the corresponding protein sequence embeddings. We then combine these cell line embeddings with gene pair representations derived from a biomedical knowledge graph and protein sequences. This triplet representation learning strategy enables MiT4SL to capture both shared biological mechanisms across cell lines and those unique to each cell line, effectively mitigating distribution shift and improving generalization to target cell lines. Additionally, explicit cell line embeddings provide the necessary signals for MiT4SL to effectively differentiate between cell line contexts, enabling it to adjust predictions and mitigate possible label conflicts for the same gene pair across different cell lines. Experimental results across various cell line-adapted scenarios show that MiT4SL outperforms six state-of-the-art models. To the best of our knowledge, MiT4SL is the first deep learning model designed specifically for cancer cell line-adapted SL prediction. AvailabilityThe code of our work is available at https://github.com/JieZheng-ShanghaiTech/MiT4SL.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.1%
14.4%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.5%
11.4%
3
Bioinformatics
1204 papers in training set
Top 2%
11.4%
4
Nature Communications
5641 papers in training set
Top 22%
7.6%
5
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.2%
4.6%
6
Nature Machine Intelligence
70 papers in training set
Top 0.7%
4.1%
50% of probability mass above
7
Patterns
78 papers in training set
Top 0.4%
3.9%
8
Bioinformatics Advances
203 papers in training set
Top 2%
3.9%
9
Genome Biology
637 papers in training set
Top 4%
2.6%
10
Advanced Science
286 papers in training set
Top 3%
2.3%
11
Journal of Computational Biology
48 papers in training set
Top 0.4%
2.3%
12
Nucleic Acids Research
1281 papers in training set
Top 8%
2.0%
13
Cell Systems
201 papers in training set
Top 2%
2.0%
14
Genome Research
468 papers in training set
Top 3%
2.0%
15
PLOS Computational Biology
1863 papers in training set
Top 14%
1.8%
16
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.6%
1.7%
17
BMC Bioinformatics
457 papers in training set
Top 4%
1.7%
18
Genomics, Proteomics & Bioinformatics
172 papers in training set
Top 1%
1.7%
19
iScience
1154 papers in training set
Top 21%
1.4%
20
Scientific Reports
3612 papers in training set
Top 61%
1.4%
21
npj Digital Medicine
118 papers in training set
Top 3%
1.1%
22
Nature Methods
385 papers in training set
Top 7%
0.8%
23
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.8%
24
PNAS Nexus
159 papers in training set
Top 5%
0.6%
25
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
26
Computational and Structural Biotechnology Journal
242 papers in training set
Top 9%
0.6%