IntegrateALL: an end-to-end RNA-seq analysis pipeline for multilevel data extraction and interpretable subtype classification in B-precursor ALL
Wolgast, N.; Beder, T.; Mondal, M.; Walter, W.; Hutter, S.; Bendig, S.; Kaessens, J. C.; Hansen, B.-T.; Iben, K.; Wolf, S.; Cremer, A.; Barz, M.; Neumann, M.; Goekbuget, N.; Haferlach, C.; Brueggemann, M.; Baldus, C. D.; Hartmann, A. M.; Bastian, L.
Show abstract
Transcriptome sequencing (RNA-seq) is emerging as a diagnostic standard for B-cell precursor acute lymphoblastic leukemia (B-ALL). Expression-based classifiers reach [~]95% accuracy, but reproducible end-to-end solutions that also integrate transcript-derived genomic drivers and quantitative virtual karyotyping are lacking. We developed IntegrateALL, a Snakemake pipeline that standardizes RNA-seq analysis from FASTQ to rule-based subtype assignment across 26 WHO-HAEM5/ICC entities by integrating expression-based subtype prediction, gene fusion- / hotspot SNV calling and virtual karyotyping. We introduce KaryALL, a machine-learning classifier that uses normalized expression and minor-allele-frequency features (RNASeqCNV) to distinguish near haploid, hypodiploid and high hyperdiploid B-ALL and chromosome-21 gains/iAMP21 (accuracy: 0.98 / F1-score: 0.96 on 615 independent test samples). SNP-array concordance supported RNA-based karyotyping. Applied to 774 unselected B-ALL cases, IntegrateALL yielded unambiguous subtype assignments in 81.5%, based on concordance of gene expression class with a defining driver (75.3% of all cases) or, in selected cases, high-confidence expression-based classification alone (6.2%); the remainder (18.5%) were flagged for manual curation. Independent validation (3 cohorts; n=436, including pediatric cases) reproduced these distributions. Across all patients (n=1,210), 2.6% harbored two subtype defining drivers, including hyperdiploidy in fusion-driven subtypes where it was not expected or subtype-defining SNVs (e.g., PAX5 P80R / IKZF1 N159Y) co-occurring with BCR::ABL1-positive/-like, KMT2A- or DUX4-fusions. In most dual-driver cases, one subtype gene expression signature predominated, indicating a hierarchy of oncogenic control and the value of systematic driver screening alongside expression-based calls. IntegrateALL provides an adaptable fully reproducible workflow for molecular B-ALL characterization by systematically integrating genomic drivers and downstream gene regulation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mutational and transcriptional landscape of pediatric B-cell precursor lymphoblastic lymphoma 96%
- Enhancer heterogeneity in acute lymphoblastic leukemia drives differential gene expression between patients 95%
- Leukemia escapes immunity by imposing a Type-1 regulatory program on neoantigen-specific CD4+ T cells. 95%
Similar papers in this journal
- Modeling IKZF1 lesions in B-ALL reveals distinct chemosensitivity patterns and potential therapeutic vulnerabilities 96%
- Genomics of PDGFR-rearranged hypereosinophilic syndrome 95%
- Single cell long-read genotyping of transcriptomes reveals discrete mechanisms of clonal evolution in post-myeloproliferative neoplasm acute myeloid leukemia. 94%
Similar papers in this journal
- Copy number signatures predict chromothripsis and associate with poor clinical outcomes in patients with newly diagnosed multiple myeloma 95%
- Role of Stem-Like Cells in Chemotherapy Resistance and Relapse in pediatric T Cell Acute Lymphoblastic Leukemia 95%
- Joint profiling of DNA and proteins in single cells to dissect genotype-phenotype associations in leukemia 95%
Similar papers in this journal
- Mapping AML heterogeneity – multi-cohort transcriptomic analysis identifies novel clusters and divergent ex-vivo drug responses 96%
- AML/T cell interactomics uncover correlates of patient outcomes and the key role of ICAM1 in T cell killing of AML 94%
- Identification of leukemia stem cell subsets with distinct transcriptional, epigenetic and functional properties 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.