Back

Machine Learning Investigation Of Gene Expression Datasets Reveals TP53 Mutant-like AML With Wild Type TP53 And Poor Prognosis

Lee, Y.; Baughn, L. B.; Sachs, Z.; Myers, C. L.

2023-02-23 cancer biology
10.1101/2023.02.22.529592 bioRxiv
Show abstract

Acute myeloid leukemia (AML) with TP53 mutations (TP53Mut) has poor clinical outcomes with 1-year survival rates of less than 10%. We investigated whether this AML subtype harbors a distinct gene expression profiling (GEP), what this GEP reveals about TP53Mut AML pathophysiology, and whether this GEP is prognostic in TP53 wild type (TP53WT) AML. We applied a supervised machine-learning approach to assess whether a unique TP53Mut GEP could be detected. Using the BEAT-AML dataset, we randomly divided the samples into training and testing datasets, while the TCGA dataset was reserved as a validation dataset. We trained a ridge regression machine learning model to classify TP53Mut and TP53WT cases. This model was highly accurate in distinguishing TP53Mut versus TP53WT cases in both the test and validation data sets. Additionally, we noted a cohort of TP53WT samples with high ridge regression scores and poor overall survival, suggesting share clinical and GEP features with TP53Mut AML. We defined these TP53WT samples as TP53 mutant-like (TP53Mut-like) AMLs. We trained a second ridge regression model to specifically detect TP53Mut-like samples in the BEAT AML dataset and found that TCGA data also harbors TP53Mut-like samples. The TP53Mut-like samples in the TCGA also have a worse OS rate than TP53WT cases. Using drug sensitivity data from 122 small molecules in the BEAT AML dataset, we found TP53Mut-like AMLs have distinct drug sensitivity patterns compared to TP53WT. Finally, we identified a 25 gene signature that can identify TP53Mut-like cases. This signature could be used clinically to identify this novel subset of poor-prognosis AML.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.