Back

Two-Stage Machine Learning Based Prediction of Thrombophilia Management

McRae, H. L.; Kahl, F.; Kapsecker, M.; Ruehl, H.; Jonas, S. M.; Poetzsch, B.

2025-10-23 health informatics
10.1101/2025.10.22.25338547 medRxiv
Show abstract

Thrombophilia diagnosis and management rely on the nuanced interpretation of clinical history, risk factors, and laboratory data, yet significant variability exists in clinical practice due to subjective assessment, institutional differences, and limited consensus guidelines. In this retrospective study, we evaluated the use of a two-stage machine learning (ML) approach to predict thrombophilia diagnosis and subsequent anticoagulation treatment. The study included data (14 clinical and 26 laboratory features) from 496 patients evaluated at the University Hospital Bonn between 2019 and 2024. The hybrid nature of the diagnostic categories combining ordinal (including thrombophilia diagnosis severity and treatment recommendations) and categorical (e.g. antiphospholipid syndrome) classes necessitated a two-stage approach. XGBClassifier was used to distinguish categorical from ordinal classes, followed by an ordinal-specific XGBOrdinalV2 model. A sliding window approach improved classification performance across all ordinal categories reaching sensitivities above 89% across all ordinal classes. Feature importance analysis revealed that age at first thrombosis and antiphospholipid antibody status were key predictive variables. Of the 496 patients, 362 (73%) experienced no discrepancy between the ML-based predictions and the practitioner diagnosis and subsequent treatment recommendations. Re-evaluation of the remaining 134 patients revealed that, while the ML models correctly classified 36 patients, it underestimated or overestimated the thrombophilia severity in 71 and 27 patients, respectively. This study highlights the potential of interpretable ML models to support standardized thrombophilia management and improve diagnostic accuracy. Future prospective studies and external validation are needed to assess generalizability and clinical impact.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.