Data Quality in Clinical Coding: A Critical Analysis and Preliminary Study
Khadka, S.; Jiang, X.; Palade, V.
Show abstract
Clinical coding is a vital yet complex part of healthcare practice. Most automated coding research relies on imperfect training data, which has an unavoidable negative impact on the quality of code prediction. A key contributing issue, vastly overlooked in current research, is the ubiquitous presence of undercoding and various types of coding errors in widely used coding datasets. From another angle, coding audit is as challenging as coding itself due to the lack of assistive tools, which is also under-studied compared to automated coding research. In this work, we uncover substantial undercoding and errors in commonly used datasets and present the first empirical study on their impacts on the performances of automated coding algorithms. To enable this, we develop an interpretable coding pipeline, using large language models (LLMs) for evidence extraction and code verification, and a multiclass classifier trained on a large-scale dataset of silver-standard evidence-code pairs for code prediction. Assisted by the pipeline, three professional coders systematically identify, categorise, and correct errors in two widely-used coding datasets that have human-annotated evidence texts for assigned codes. As an AI-assisted coding audit tool, the current study uncovers significant data quality issues, including a 76.3% undercoding rate in MDACE and a 29.7% error rate in CodiEsp. Re-evaluating existing models on error-corrected datasets results in consistent performance improvements. Additionally, the pipeline is a highly potential coding framework, which achieves superior or comparative performances to state-of-the-art LLM-based methods. The results underscore the necessity of shifting research focus from model-centric to data-centric solutions in clinical AI.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 95%
Similar papers in this journal
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 95%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.