MOSAIC: Methylation-Oriented Site Analysis and Information Classifier for Robust Epigenomic Classification of Acute Leukemia in Clinical Cohorts with Variable Tumor Purity
Shah, A.; Green, D.; Wainmann, L.; Karrs, J.; Shah, P.
Show abstract
DNA methylation-based classification offers a rapid diagnostic complement to conventional molecular workflows in acute leukemia. Existing classifiers are trained on array-derived reference cohorts whose construction favors specimens with adequate tumor content, leaving clinically relevant low-purity specimens underrepresented and classifier robustness in this regime uncharacterized. On held-out low-purity specimens, existing classifiers were concordant with expert pathology in only 7 of 10 (MARLIN) and 5 of 10 (ALMA) cases, motivating a classifier built to maintain accuracy at low tumor purity. We developed MOSAIC (Methylation-Oriented Site Analysis and Information Classifier), a neural network classifier built to maintain accuracy across the full range of tumor purities encountered in clinical practice. MOSAIC is a neural network trained on publicly available array-based methylation data augmented with native methylation calls from Oxford Nanopore sequencing. MOSAIC was evaluated on low-purity specimens held out entirely from training. On these held-out low-blast leukemia specimens, all below 25% blasts and including a case at 1.4%, MOSAIC was concordant with expert pathology in every case, recovering the correct subtype where diluted disease signal would otherwise be mistaken for normal or unrelated tissue. Gradient-based saliency analysis showed that the network relies on a partially distinct set of discriminative CpG probes when classifying low-blast specimens. MOSAIC demonstrates that augmenting training with clinically representative clinical specimens yields methylation-based leukemia classification that maintains effectiveness under the variable tumor purity of real clinical cohorts.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mapping AML heterogeneity – multi-cohort transcriptomic analysis identifies novel clusters and divergent ex-vivo drug responses 94%
- Continuous Indexing of Fibrosis (CIF): Improving the Assessment and Classification of MPN Patients 93%
- Clinical Impact of Panel Based Error Corrected Next Generation Sequencing versus Flow Cytometry to Detect Measurable Residual Disease (MRD) in Acute Myeloid Leukemia (AML) 92%
Similar papers in this journal
- Acute myeloid leukemia expresses a specific group of olfactory receptors 93%
- Development of a Notch pathway assay and quantification of functional Notch pathway activity in T-cell acute lymphoblastic leukemia 92%
- Loss of FBXO9 enhances proteasome activity and promotes aggressiveness in acute myeloid leukemia 92%
Similar papers in this journal
- Methylome-based cell-of-origin modeling (Methyl-COOM) identifies aberrant expression of immune regulatory molecules in CLL 93%
- CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells 93%
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 92%
Similar papers in this journal
- Cas9-directed long-read sequencing to resolve optical genome mapping findings in leukemia diagnostics. 94%
- Mebendazole for Differentiation Therapy of Acute Myeloid Leukemia Identified by a Lineage Maturation Index 93%
- Refined detection and phasing of structural aberrations in pediatric acute lymphoblastic leukemia by linked-read whole genome sequencing 93%
Similar papers in this journal
- Single-Cell Atlas of AML Reveals Age-Related Gene Regulatory Networks in t(8;21) AML 95%
- COMPARE, an ultra-fast and robust suite for multiparametric screening, identifies phenotypic drug responses in acute myeloid leukemia 94%
- Unsupervised machine learning reveals key immune cell subsets in COVID-19, rhinovirus infection, and cancer therapy 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.