Back

Transformer models of mutation risk at base-pair resolution identify non-coding hotspot cancer driver mutations

Galvan-Femenia, I.; Veiner, M.; Naro, D.; Supek, F.

2026-07-10 genomics
10.64898/2026.07.07.736824 bioRxiv
Show abstract

Recurrent somatic mutations reveal cancer drivers, but in whole genomes many non-coding hotspots are passengers generated by localized mutational processes. We developed MutFormer, a transformer/convolutional neural net model that predicts base-pair-resolution somatic mutation risk from DNA sequence alone, separately for COSMIC signatures. Trained on >90 million high-confidence mutational signature-assigned SNVs from cancer genomes, MutFormer learns extended sequence determinants beyond trinucleotide context, often spanning up to ~20 nucleotides, and recovers APOBEC, UV, POLE and SBS17 sequence preferences as well as various additional mutation risk-prone motifs. We integrated MutFormer predictions with mutation burden, signature exposures and epigenetic covariates to model neutral recurrence of individual hotspots in >18,000 tumor whole genomes. Coding-region analyses calibrated the framework against known driver genes and AlphaMissense scores, supporting conservative false-discovery estimates. In non-coding regions, most recurrent hotspots were explained by passenger mutability, whereas selected outliers were enriched near cancer genes and supported by SpliceAI, PromoterAI, AlphaGenome and expression data. Prioritized candidates include splice-region or deep-intronic hotspots in BCL6, PTEN, TCF7L2, PBRM1, PTPRT and VHL, and promoter hotspots in SHKBP1, PRSS3 and BCL2.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.