Back

Data Quality in Clinical Coding: A Critical Analysis and Preliminary Study

Khadka, S.; Jiang, X.; Palade, V.

2025-08-26 health informatics
10.1101/2025.08.24.25334321 medRxiv
Show abstract

Clinical coding is a vital yet complex part of healthcare practice. Most automated coding research relies on imperfect training data, which has an unavoidable negative impact on the quality of code prediction. A key contributing issue, vastly overlooked in current research, is the ubiquitous presence of undercoding and various types of coding errors in widely used coding datasets. From another angle, coding audit is as challenging as coding itself due to the lack of assistive tools, which is also under-studied compared to automated coding research. In this work, we uncover substantial undercoding and errors in commonly used datasets and present the first empirical study on their impacts on the performances of automated coding algorithms. To enable this, we develop an interpretable coding pipeline, using large language models (LLMs) for evidence extraction and code verification, and a multiclass classifier trained on a large-scale dataset of silver-standard evidence-code pairs for code prediction. Assisted by the pipeline, three professional coders systematically identify, categorise, and correct errors in two widely-used coding datasets that have human-annotated evidence texts for assigned codes. As an AI-assisted coding audit tool, the current study uncovers significant data quality issues, including a 76.3% undercoding rate in MDACE and a 29.7% error rate in CodiEsp. Re-evaluating existing models on error-corrected datasets results in consistent performance improvements. Additionally, the pipeline is a highly potential coding framework, which achieves superior or comparative performances to state-of-the-art LLM-based methods. The results underscore the necessity of shifting research focus from model-centric to data-centric solutions in clinical AI.

Published in International Journal of Medical Informatics (predicted rank #11) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.