Effective tree-based classification for automated flow cytometry data analysis on samples with suspected haematological malignancy
Rothwell, A.; Carter, A.; Green, P. L.; Jones, A. R.
Show abstract
Flow cytometry is a commonly used diagnostic technique for haematological malignancies. The gold standard method for analysis of flow cytometry data is manual gating, which is time consuming and requires a highly skilled operator, generating a bottleneck in the workflow and potentially increasing time to diagnose malignancy. For nearly 20 years attempts have been made at replacing manual analysis with automated algorithms, however these are not deemed accurate enough for clinical practice. Clustering methods have been the focus of previous automated attempts, though supervised methods have been shown to be more accurate and require less manual intervention. Tree-based classification algorithms make decisions using an analogous process to manual gating. One hundred and fifty-two flow cytometry files were generated from peripheral blood samples of patients with suspected haematological malignancies. A trained operator labelled events in these files as one of nine cell types. CART, Random Forest and XGBoost were trained on the labelled dataset and the performance was evaluated against previously published clustering methods. Classification algorithms showed higher mean F1 scores than clustering methods. There was no significant difference between CART, Random Forest and XGBoost mean F1 scores, and all three algorithms showed mean prediction times per sample of less than 25 seconds. Tree-based methods struggled to differentiate B cell subtypes, which show similar phenotypic signatures and present an area for future improvement. This work demonstrates the effectiveness of tree-based classification algorithms for flow cytometry analysis. Overall, CART may offer a solution to automated flow cytometry analysis for the purpose of haematological malignancies due to showing high agreement with manual analysis, and short prediction and training times.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 93%
- A distinct four-value blood signature of pyrexia under combination therapy of malignant melanoma with BRAF/MEK-inhibitors evidenced by an algorithm-defined pyrexia score 92%
- FaDA: A Shiny web application to accelerate common lab data analyses 92%
Similar papers in this journal
- Classification of human white blood cells using machine learning for stain-free imaging flow cytometry 95%
- Hematologist-level classification of mature B-cell neoplasm using deep learning on multiparameter flow cytometry data 93%
- Integration, exploration, and analysis of high-dimensional single-cell cytometry data using Spectre 92%
Similar papers in this journal
- Hilab system, a new point-of-care hematology analyzer supported by the Internet of Things and Artificial Intelligence 92%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 92%
- Distinct SARS-CoV-2 Antibody Reactivity Patterns in Coronavirus Convalescent Plasma Revealed by a Coronavirus Antigen Microarray 91%
Similar papers in this journal
- Topological Structures in the Space of Treatment-Naive Patients With Chronic Lymphocytic Leukemia 94%
- Use of high-plex data reveals novel insights into the tumour microenvironment of clear cell renal cell carcinoma 90%
- piNET: An Automated Proliferation Index Calculator Framework for Ki67 Breast Cancer Images 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.