External Validation and Calibration Assessment of Explainable Machine Learning Models for GVHD Prediction After Allogeneic HSCT
Syed, N.; Ahmed, N.; Abuhaleeqa, M.; Al Kaabi, F. M.; Raza, A.; Al Zaki, A.; Sammour, F.; Alkhatib, Y.; Gopalakrishnan, D.; Afrooz, I.; Damlaj, M.; Abu Jazar, H.; Abdel-Razeq, H.; Halahleh, K.; Yaqub, M.; Hashmi, S.
Show abstract
Background Graft versus host disease (GVHD) remains a major determinant of morbidity and mortality following allogeneic hematopoietic stem cell transplantation (allo HSCT). Existing GVHD prediction models demonstrate modest discrimination and limited generalizability, and calibration drift across external populations is rarely characterized despite its essential role in the clinical interpretability of predicted probabilities. Objectives To develop and externally validate an explainable machine learning framework for predicting acute and chronic GVHD and associated overall survival in patients with acute myeloid leukemia (AML), acute lymphoblastic leukemia (ALL), and myelodysplastic syndromes (MDS) undergoing allo HSCT, and to systematically characterize calibration across heterogeneous external validation cohorts to inform deployment requirements. Study Design The model was developed on three publicly available registry-derived datasets (N = 2,509) and externally validated across six independent cohorts (N = 14,788) comprising adult and pediatric allo HSCT recipients, including a regional Middle Eastern cohort (UAE and Jordan). A standardized preprocessing pipeline harmonized heterogeneous datasets. Gradient boosting models (CatBoost) were used for binary GVHD prediction; exploratory overall survival analysis used a Cox proportional hazards model with predicted acute GVHD risk as a covariate. Discrimination (AUROC with bootstrap 95% CI), calibration (logistic recalibration intercept and slope with analytical 95% CI), and feature importance (SHapley Additive exPlanations, SHAP) were assessed in training out-of-fold and all external cohorts. Results In internal validation, AUROC was 0.63 (95% CI 0.61-0.65) for acute GVHD and 0.72 (95% CI 0.70-0.74) for chronic GVHD. External validation demonstrated AUROC ranges of 0.51-0.57 (acute) and 0.54-0.64 (chronic), with consistent performance across disease subgroups despite substantial heterogeneity in transplant practices and feature availability. In exploratory survival analysis, the acute-GVHD-informed Cox model achieved a training-cohort C-index of 0.679 (95% CI 0.658-0.697); external C-indices ranged from 0.47-0.53. Calibration analysis identified systematic external risk overestimation (negative calibration intercept in 10 of 11 evaluable external cohort-target combinations) with heterogeneous slope drift requiring cohort-specific recalibration. Key predictors included recipient age, graft source, conditioning intensity, GVHD prophylaxis, and HLA match ratio. Conclusions An explainable, externally validated GVHD prediction framework was developed using heterogeneous registry-derived datasets, with systematic characterization of calibration drift across multiple external cohorts, an analysis rarely reported in prior GVHD prediction literature. Predictive performance was modest for acute GVHD and moderate for chronic GVHD, constrained by missing immunobiological variables and incomplete HLA characterization. Per-cohort recalibration is required before clinical deployment, with prospective validation and benchmarking against established GVHD risk scores identified as priority next steps.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dental biofilm microbiota dysbiosis is associated with the risk of acute graft-versus-host disease after allogeneic hematopoietic stem cell transplantation 94%
- Development and Validation of Multivariable Prediction Models of Serological Response to SARS-CoV-2 Vaccination in Kidney Transplant Recipients 92%
- Dampened inflammatory signalling and myeloid-derived suppressor-like cell accumulation reduces circulating monocytic HLA-DR density and associates with malignancy risk in long-term renal transplant recipients 92%
Similar papers in this journal
- A Combined HLA Molecular Mismatch and Expression Analysis for Evaluating HLA-DPB1 de novo Donor-Specific Antibody Risk in Pediatric Solid Organ Transplantation: Implications for Solid Organ & Hematopoietic Stem Cell Transplantation 93%
- Tixagevimab-cilgavimab as an early treatment for COVID-19 in kidney transplant recipients 90%
- FOXP3 mRNA profile prognostic of T cell mediated rejection and human kidney allograft survival 90%
Similar papers in this journal
- Pretransplant targeting of TNFRSF25 and CD25 stimulates recipient Tregs in target tissues ameliorating GVHD post-HSCT 93%
- Clinicopathologic Correlates and Natural History of Atypical Chronic Myeloid Leukemia 92%
- Inhibitory KIR-Ligand Interactions and Relapse Protection Following HLA Matched Allogeneic Hematopoietic Cell Transplantation for Acute Myelogenous Leukemia 91%
Similar papers in this journal
- Exploring the role of Large Language Models (LLMs) in hematology: a systematic review of applications, benefits, and limitations 90%
- Validation of the IMPEDE VTE Score for Prediction of Venous Thromboembolism in Multiple Myeloma: A Retrospective Cohort Study 90%
- Immune response to COVID-19 vaccination is attenuated by poor disease control and antimyeloma therapy with vaccine driven divergent T cell response 88%
Similar papers in this journal
- Stratified computational meta-analysis of 2213 acute myeloid leukemia patients reveals age- and sex-dependent gene expression signatures 91%
- Serum-free differentiation platform for the generation of B lymphocytes and natural killer cells from human CD34+ cord blood progenitors 91%
- Complex genotype-phenotype relationships shape the response to treatment of Down Syndrome Childhood Acute Lymphoblastic Leukaemia 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.