The gene expression classifier ALLCatchR identifies B-precursor ALL subtypes and underlying developmental trajectories across age
Beder, T.; Hansen, B.-T.; Hartmann, A. M.; Zimmermann, J.; Amelunxen, E.; Wolgast, N.; Walter, W.; Zaliova, M.; Antic, Z.; Chouvarine, P.; Bartsch, L.; Barz, M.; Bultmann, M.; Horns, J.; Bendig, S.; Kaessens, J.; Kaleta, C.; Cario, G.; Schrappe, M.; Neumann, M.; Goekbuget, N.; Bergmann, A. K.; Trka, J.; Haferlach, C.; Brueggemann, M.; Baldus, C. D.; Bastian, L.
Show abstract
Current classifications (WHO-HAEM5 / ICC) define up to 26 molecular B-cell precursor acute lymphoblastic leukemia (BCP-ALL) disease subtypes, which are defined by genomic driver aberrations and corresponding gene expression signatures. Identification of driver aberrations by RNA-Seq is well established, while systematic approaches for gene expression analysis are less advanced. Therefore, we developed ALLCatchR, a machine learning based classifier using RNA-Seq expression data to allocate BCP-ALL samples to 21 defined molecular subtypes. Trained on n=1,869 transcriptome profiles with established subtype definitions (4 cohorts; 55% pediatric / 45% adult), ALLCatchR allowed subtype allocation in 3 independent hold-out cohorts (n=1,018; 75% pediatric / 25% adult) with 95.7% accuracy (averaged sensitivity across subtypes: 91.1% / specificity: 99.8%). High confidence predictions were achieved in 84.6% of samples with 99.7% accuracy. Only 1.2% of samples remained unclassified. ALLCatchR outperformed existing tools and identified novel candidates in previously unassigned samples. We established a novel RNA-Seq reference of human B-lymphopoiesis. Implementation in ALLCatchR enabled projection of BCP-ALL samples to this trajectory, which identified shared patterns of proximity of BCP-ALL subtypes to normal lymphopoiesis stages. ALLCatchR sustains RNA-Seq routine application in BCP-ALL diagnostics with systematic gene expression analysis for accurate subtype allocations and novel insights into underlying developmental trajectories.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- IntegrateALL: an end-to-end RNA-seq analysis pipeline for multilevel data extraction and interpretable subtype classification in B-precursor ALL 96%
- Bone marrow lymphocyte dynamics during chemotherapy in pediatric acute myeloid leukemia 95%
- Ex Vivo Drug Responses and Molecular Profiles of 597 Pediatric Acute Lymphoblastic Leukemia Patients 95%
Similar papers in this journal
Similar papers in this journal
- Mapping AML heterogeneity – multi-cohort transcriptomic analysis identifies novel clusters and divergent ex-vivo drug responses 96%
- AML/T cell interactomics uncover correlates of patient outcomes and the key role of ICAM1 in T cell killing of AML 94%
- A transcriptomic continuum of differentiation arrest identifies myeloid interface acute leukemias with poor prognosis 94%
Similar papers in this journal
- Genome-wide CRISPR Screens Identify Ferroptosis as a Novel Therapeutic Vulnerability in Acute Lymphoblastic Leukemia 97%
- Quantification of measurable residual disease using duplex sequencing in adults with acute myeloid leukemia 95%
- First-born twin has a higher risk of acute leukemia in a population-based assessment of cancer in twins in California, and lower than anticipated rate of twin concordance 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.