Predicting cause of death from free-text health summaries: development of an interpretable machine learning tool
McWilliams, C.; Walsh, E.; Huxor, A.; Turner, E. L.; Santos-Rodriguez, R.
Show abstract
Structured AbstractO_ST_ABSPurposeC_ST_ABSAccurately assigning cause of death is vital to understanding health outcomes in the population and improving health care provision. Cancer-specific cause of death is a key outcome in clinical trials, but assignment of cause of death from death certification is prone to misattribution, therefore can have an impact on cancer-specific trial mortality outcome measures. MethodsWe developed an interpretable machine learning classifier to predict prostate cancer death from free-text summaries of medical history for prostate cancer patients (CAP). We developed visualisations to highlight the predictive elements of the free-text summaries. These were used by the project analysts to gain an insight of how the predictions were made. ResultsCompared to independent human expert assignment, the classifier showed >90% accuracy in predicting prostate cancer death in test subset of the CAP dataset. Informal feedback suggested that these visualisations would require adaptation to be useful to clinical experts when assessing the appropriateness of these ML predictions in a clinical setting. Notably, key features used by the classifier to predict prostate cancer death and emphasised in the visualisations, were considered to be clinically important signs of progressing prostate cancer based on prior knowledge of the dataset. ConclusionThe results suggest that our interpretability approach improve analyst confidence in the tool, and reveal how the approach could be developed to produce a decision-support tool that would be useful to health care reviewers. As such, we have published the code on GitHub to allow others to apply our methodology to their data (https://zenodo.org/badge/latestdoi/294910364).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 93%
- Modeling physician variability to prioritize relevant medical record information 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 92%
Similar papers in this journal
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 90%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 92%
- LinkR: an open source, low-code and collaborative data science platform for healthcare data analysis and visualization 92%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.