Back

Machine Learning-Based Identification of Sickle Cell Disease Subphenotypes in Clinical Trial Data

Xiao, W.; Oneal, P.; Wang, M.; Mehta, N. J.; Liu, Q.; Zhang, R.; Perrine, S.; Ryan, Q.

2025-06-02 hematology
10.1101/2025.06.01.25328537 medRxiv
Show abstract

Sickle Cell Disease (SCD) is a rare autosomal recessive disorder caused by a point mutation producing abnormal hemoglobin S, leading to deformed red blood cells and a wide range of clinical manifestations, including pain crises, organ damage, and an increased risk of infection. These devastating complications often result in significant morbidity and early mortality, presenting significant therapeutic challenges. Currently, there is a lack of clinically validated predictive tools to assess individual SCD patients prognoses and therapeutic responses. This is largely due to the complexity and variability of the clinical manifestations, which vary widely among patients. As a result, there remains an unmet need for a systematic approach to SCD disease subphenotype classification that can guide and tailor therapeutic strategies, predict outcomes, and improve patients lives. Over a decade ago, two clinical subphenotypes in SCD were proposed based on literature and clinical observations. However, this concept has not been applied or explored in the design of clinical trials (CT). Recent advances in machine learning (ML) applications in medicine, and growing availability of SCD clinical trial data evaluating therapeutics which target different pathophysiologic aspects of the disease, provides opportunity to enhance understanding of therapeutic responses within SCD populations. Applying ML techniques to a large CT database could support development of robust disease models capable of identifying and validating disease subphenotypes, with potential to predict outcomes to specific therapies based on mechanism of action and to optimize care in SCD. In this study, we constructed a comprehensive database comprising 3,551 patients with SCD from 16 clinical trials that supported therapeutic approvals for SCD. Using this database, we applied a machine learning pipeline to develop a rule-based classification method, which identified two distinct clinical subphenotypes of SCD: the Vaso-occlusive Primary (VP) subphenotype, primarily characterized by a higher frequency of vaso-occlusive pain crises, and the Hemolytic Dominant (HD) subphenotype, characterized by chronic hemolysis and its associated complications. Biomarker comparisons demonstrated that the VP subphenotype was associated with a significantly higher annual rate of vasoocclusive crisis events, significantly higher levels of total and fetal hemoglobin, and leukocytosis, while the HD subphenotype exhibited significantly higher levels of hemolysis-related biomarkers of indirect bilirubin. The biomarker profiles were validated using an independent clinical trial dataset, which confirmed these two subphenotypes in SCD. Our study demonstrated that the integration of ML with disease pathophysiology enables robust identification of clinically meaningful subphenotypes of SCD from an international clinical trial database. This approach provides a basis for developing predictive disease models, which may optimize treatment strategies and improve patients outcomes. Further, our methodological framework offers a scalable model for application to identify subsets in other rare genetic diseases.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Blood
74 papers in training set
Top 0.1%
22.1%
2
Haematologica
25 papers in training set
Top 0.1%
15.2%
3
Journal of Clinical Investigation
179 papers in training set
Top 0.2%
9.7%
4
British Journal of Haematology
15 papers in training set
Top 0.1%
9.7%
50% of probability mass above
5
JCI Insight
277 papers in training set
Top 0.5%
7.9%
6
eLife
5828 papers in training set
Top 29%
4.1%
7
Blood Advances
62 papers in training set
Top 0.5%
3.2%
8
Molecular Therapy
81 papers in training set
Top 0.8%
1.9%
9
Frontiers in Immunology
638 papers in training set
Top 6%
1.7%
10
Journal of Thrombosis and Haemostasis
32 papers in training set
Top 0.2%
1.7%
11
PLOS Computational Biology
1863 papers in training set
Top 16%
1.3%
12
Molecular Therapy - Methods & Clinical Development
38 papers in training set
Top 0.4%
1.3%
13
Leukemia
42 papers in training set
Top 0.5%
1.3%
14
Nature Communications
5641 papers in training set
Top 51%
1.1%
15
PLOS ONE
5266 papers in training set
Top 54%
1.1%
16
Science Advances
1243 papers in training set
Top 27%
1.1%
17
PLOS Genetics
862 papers in training set
Top 11%
1.0%
18
Circulation: Genomic and Precision Medicine
48 papers in training set
Top 0.8%
0.8%
19
iScience
1154 papers in training set
Top 34%
0.8%
20
Cell Reports Medicine
153 papers in training set
Top 6%
0.6%
21
Molecular Therapy Nucleic Acids
39 papers in training set
Top 1%
0.6%
22
Molecular Therapy - Nucleic Acids
25 papers in training set
Top 0.8%
0.6%
23
EMBO Reports
263 papers in training set
Top 8%
0.6%
24
npj Digital Medicine
118 papers in training set
Top 4%
0.6%
25
Scientific Reports
3612 papers in training set
Top 78%
0.6%
26
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 44%
0.6%