Back

Long deletion signatures in repetitive genomic regions track somatic evolution and enable sensitive detection of microsatellite instability

Guo, Q.; Househam, J.; Lakatos, E.; Nowinski, S.; Al Bakir, I.; Grant, H.; Balarajah, V.; Hughes, C. S.; Zapata, L.; Kocher, H.; Sottoriva, A.; Baker, A.-M.; Mustonen, V.; Graham, T.

2024-10-04 bioinformatics
10.1101/2024.10.03.616572 bioRxiv
Show abstract

Deficiency in the mismatch repair system (MMRd) causes microsatellite instability (MSI) in cancers and determines eligibility for immunotherapy. Here, we show that MMRd tumours harbour long-deletion signatures ([≥]2-5+ base pairs deleted in repetitive regions), which provide new insights into MSI evolution and enable sensitive MSI detection particularly in challenging clinical samples. Long deletions, accumulated through stepwise DNA slippage errors, are significantly more prevalent in metastatic MMRd tumours compared to primary tumours. Importantly, we show that long-deletion signatures harbour features that are distinct from background noise, making them robustly detectable even in shallow whole genome sequencing (sWGS, [~]0.1X coverage) of formalin-fixed samples. We constructed a machine learning classifier that uses these distinct features to detect Microsatellite Instability in LOw-quality (MILO) samples. MILO achieved 100% accuracy in detecting MSI in sWGS data with only 2%-15% tumour purity and demonstrated promise in identifying MMRd clones in precancerous intestinal lesions. We propose that MILO could be clinically used for the sensitive monitoring of MMRd cancer evolution from early to late stages, using minimal sequencing data from both archival and fresh-frozen samples with low tumour content. SignificanceMutational signatures characterised by long deletions in repetitive genomic regions provide a sensitive route to detect and track MMRd clone evolution, even with low purity shallow whole genome sequencing data.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.