Systematic Evaluation of Common Natural Language Processing Techniques to Codify Clinical Notes
Tavabi, N.; Singh, M.; Pruneski, J.; Kiapour, A.
Show abstract
Proper codification of medical diagnoses and procedures is essential for optimized health care management, quality improvement, research, and reimbursement tasks within large healthcare systems. Assignment of diagnostic or procedure codes is a tedious manual process, often prone to human error. Natural Language Processing (NLP) have been suggested to facilitate these manual codification process. Yet, little is known on best practices to utilize NLP for such applications. Here we comprehensively assessed the performance of common NLP techniques to predict current procedural terminology (CPT) from operative notes. CPT codes are commonly used to track surgical procedures and interventions and are the primary means for reimbursement. The direct links between operative notes and CPT codes makes them a perfect vehicle to test the feasibility and performance of NLP for clinical codification. Our analysis of 100 most common musculoskeletal CPT codes suggest that traditional approaches (i.e., TF-IDF) can outperform resource intensive approaches like BERT, in addition to providing interpretability which can be very helpful and even crucial in the clinical domain. We also proposed a complexity measure to quantify the complexity of a classification task and how this measure could influence the effect of dataset size on models performance. Finally, we provide preliminary evidence that NLP can help minimize the codification error, including mislabeling due to human error.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 95%
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 94%
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 96%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 95%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 95%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 99%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 95%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.