Classifying progression status statements from radiology exams among non-small cell lung cancer patients using natural language processing
Davoudi, A.; Yu, S.; Doucette, A.; Gabriel, P.; Miller, M.; Williams, H.; Desai, H.; Le, A.; Stoeckert, C.; Maxwell, K.; Mowery, D.
Show abstract
Although NLP has been used to support cancer research more broadly, the development of NLP algorithms to extract evidence of progression from clinical notes to support lung cancer research is still in its infancy. In this study, we trained supervised machine learning classifiers using rich semantic features to detect and classify statements of progression status from radiology exams. Our progression status classifier achieves high F1-scores for detecting and discerning progression (0.80), stable (0.82), and not relevant (0.92) sentences, demonstrating promising performance. We are actively integrating these extractions with structured electronic health record data using ontologies to instantiate a longitudinal model of progression among non-small cell lung cancer patients.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural language inference for clinical registry curation 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 93%
Similar papers in this journal
Similar papers in this journal
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 94%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 93%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 91%
Similar papers in this journal
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 94%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.