Back

Reference-free microsatellite instability detection from tumor sequencing using intrasample variability modeling

Vlachos, G.; Moser, T.; Patel, M.; White, J. R.; Heitzer, E.; Diaz, L. A.

2026-01-14 bioinformatics
10.64898/2026.01.13.699258 bioRxiv
Show abstract

Microsatellite instability (MSI) is a key predictive biomarker across multiple tumor types, but current next-generation sequencing (NGS)-based callers oftendependon matched normals, large reference panels, or pretrained machine-learning models, which limits portability across assays and sequencing centers. We developed PROMIS (PROfiling of Microsatellite InStability), a tumor-only, reference-free pipeline that detects MSI by leveraging intrasample variability at predefined microsatellite loci. PROMIS models repeat-length distributions with a discrete mixture framework to distinguish stable germline configurations from unstable loci with additional allele populations. Locus-level classifications are aggregated into a continuous MSI score, defined as the fraction of unstable loci, which can then be used to classify samples as MSI-high or microsatellite-stable. We benchmarked PROMIS in colorectal (CRC), endometrial (UCEC), and gastric (STAD) cancers from The Cancer Genome Atlas (TCGA). PROMIS achieved an overall AUC of 0.995 and cohort-specific AUCs of 1.00 in CRC and STAD and 0.999 in UCEC, comparable to establi shed tools despite not using matched normals or pretrained models. Subsampling and in silico dilutions showed robust performance with substantially fewer loci and down to 3% tumor fraction. Finally, in prostate and CRC cell-free DNA (cfDNA) cohorts, including Illumina TSO500 data and an 18-gene panel, PROMIS yielded assay-agnostic MSI scores concordant with orthogonal tissue- and panel-based classifications.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.