Comparison of machine learning methods for clinical data imputation among a real-world lung cancer cohort
Yang, D. X.; Hui, Y.; Wang, B.; Avesta, A.; Zuber, C.; Park, H. S.; Aneja, S.
Show abstract
PURPOSECancer registries are important sources of real-world data (RWD) that reveal insights into practice patterns and cancer patient outcomes, but the prevalence of missing data can be high. Machine learning (ML) imputation methods can be applied to large RWD sets, but the performance of these approaches within cancer registries is unclear. METHODSWe identified non-small cell lung cancer (NSCLC) patients within the National Cancer Database diagnosed in 2014 with complete data in 19 variables of known clinical and prognostic significance. We generated synthetic missing data for each variable, then performed imputation using substitution (control) and five different ML approaches. Imputation efficacy was measured by normalized root-mean-square error (RMSE) for continuous variables and proportion of falsely classified entries (PFC) for categorical variables. We also measured algorithm runtimes and the impact of incorporating imputed values on survival modeling. RESULTS50,790 NSCLC patients were included for this study, with 81 features for each patient after data preprocessing. Among the tested ML methods, SoftImpute had the lowest RMSE (best performance) for continuous variables ranging from 0.071 to 0.080 for 10% to 50% missing data, and MissForest had the lowest PFC (best performance) for categorical variables ranging from 0.251 to 0.311 for 10 to 50% missing data. SoftImpute had a runtime of 3.28x10-4 seconds per patient record, and MissForest averaged 2.96x10-3 seconds per patient record. Deep learning imputation using a denoising autoencoder did not achieve improved performance despite higher algorithm runtimes. Cox models incorporating ML imputed data achieved similar C-index ranging from 0.787 to 0.801 for all ML methods tested. CONCLUSIONML imputation achieved promising performance for NSCLC patients within a large national cancer registry.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 97%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 92%
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 92%
Similar papers in this journal
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 93%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 92%
Similar papers in this journal
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 92%
- Predicting mortality in SARS-COV-2 (COVID-19) positive patients in the inpatient setting using a Novel Deep Neural Network 92%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.