Back

Computing tumor specificity of cancer antigen targets by k-mer indexing of healthy tissue transcriptomes

Hausmann, J.; Lang, F.; Muslu, O.; Kress, L.; Landry, J.; Suchan, M.; Nubbemeyer, A.; Kuner, R.; Weber, D.; Schrörs, B.; Schulz, M. H.; Gaida, M. M.; Sahin, U.; Ibn-Salem, J.

2026-07-03 bioinformatics
10.64898/2026.06.29.734488 bioRxiv
Show abstract

Individualized cancer immunotherapies rely on tumor-specific T-cell antigens, often predicted from somatic mutations as neoantigens. For tumors with low mutational burden, mRNA transcript variants, including gene fusions and novel splice junctions, can serve as important alternative targets. A main challenge in their identification from tumor RNA-seq is to confirm that their expression is tumor-restricted. Although large public collections of healthy-tissue RNA-seq exist, verifying tumor-specific expression requires computationally expensive re-analysis of these data for every novel candi-date. To address this, we benchmarked nine k-mer indexing algorithms and devel-oped k4neo, which leverages k-mer indexing of raw RNA-seq reads to compute the tumor specificity of any transcript variant. This mapping-free and transcript-class ag-nostic approach screens any candidate sequence against 18,960 samples across 51 healthy tissue types. We confirmed k4neo's detection accuracy with qRT-PCR and showed that k4neo accurately classifies somatic and germline variants, gene fusions, and isoforms by tumor specificity. Applied to nine tumor cohorts, it nominated a medi-an of 4-80 tumor-specific splice junctions per patient, including recurrent, long-read-validated novel antigen candidates. Together, k4neo enables efficient access to large-scale sequencing cohorts and accurately computes tumor specificity for any in-put transcript sequence, thereby expanding the repertoire of individual and shared cancer antigen targets.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Methods
385 papers in training set
Top 0.4%
18.3%
2
Nature Communications
5641 papers in training set
Top 12%
14.9%
3
Nature Biotechnology
172 papers in training set
Top 0.1%
11.8%
4
Cell Systems
201 papers in training set
Top 0.2%
11.8%
50% of probability mass above
5
Genome Biology
637 papers in training set
Top 2%
5.4%
6
Genome Medicine
183 papers in training set
Top 2%
2.8%
7
Nucleic Acids Research
1281 papers in training set
Top 6%
2.7%
8
Cell
431 papers in training set
Top 5%
2.1%
9
Genome Research
468 papers in training set
Top 3%
2.1%
10
Bioinformatics
1204 papers in training set
Top 6%
1.9%
11
Cell Reports Methods
165 papers in training set
Top 2%
1.7%
12
Cancer Cell
42 papers in training set
Top 0.9%
1.5%
13
Cell Reports
1498 papers in training set
Top 21%
1.4%
14
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
15
Cell Reports Medicine
153 papers in training set
Top 3%
1.1%
16
Nature Genetics
286 papers in training set
Top 4%
1.1%
17
Nature Biomedical Engineering
47 papers in training set
Top 1%
1.1%
18
Cell Genomics
172 papers in training set
Top 3%
1.1%
19
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
20
Science
477 papers in training set
Top 8%
1.0%
21
PLOS Computational Biology
1863 papers in training set
Top 19%
1.0%
22
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 39%
1.0%
23
Science Immunology
88 papers in training set
Top 2%
0.8%
24
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
25
Genomics, Proteomics & Bioinformatics
172 papers in training set
Top 2%
0.6%
26
Communications Medicine
113 papers in training set
Top 6%
0.6%