Back

A Supervised Machine Learning Model to Predict Therapy Response and Mortality at 90 days After Acute Myeloid Leukemia Diagnosis

Delgado Sanchis, J. A.; Pons-Suner, P.; Alvarez, N.; Sargas, C.; Dorado, S.; Gil Orti, J. V.; Signol, F.; Llop, M.; Arnal, L.; Llobet, R.; Perez-Cortes, J.-C.; Ayala, R.; Barragan, E.

2023-06-29 health informatics
10.1101/2023.06.26.23291731 medRxiv
Show abstract

Background and ObjectiveThe main objective in this paper is to validate a machine-learning model trained to predict the 90-day risk of complications for patients with Acute Myeloid Leukemia using variables available at diagnosis. This is a first fundamental step towards the development of a tool that could help physicians in their therapeutic decisions. Methods266 patients and 36 variables form the training dataset collected by Hospital 12 de Octubre (Madrid, Spain). The external test cohort provided by Instituto de Investigacion Sanitaria La Fe (Valencia, Spain) contains 162 observations. An XGBoost model was trained with one dataset and validated with the other. Additionally, the features were ranked by permutation importance and compared with the ELN 2022 risk classification by genetics at initial diagnosis. ResultsThe model was evaluated with the training cohort using leave-one-out cross-validation, reaching a ROC-AUC of 0.85. By setting the functioning point that maximises Youdens index, 3 out of 4 patients with complications and 84 out of 100 in remission are correctly classified. The model was validated with external data collected in a different hospital, achieving 0.7 ROC-AUC. At the best functioning point, almost 6 out of 10 patients with complications and 8 out of 10 patients in remission are correctly classified. Ranking the variables by descending importance, the top four are, in order: age, white-blood-cells count, Gender, and TP53. The list exhibits good coherence with the ELN 2022 risk classification. ConclusionsThe model achieves performances that suggest it could be used as a therapeutical decision support tool. Important variables are coherent with ELN 2022 risk classification. Further work is needed to understand the reasons for the drop in test performance. The 90-day model should be supplemented by others that predict the risk of complications at six months or one year.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.