Validation of 13,102 ICD-10-CM Codes Using a Large Language Model-Based System
Wang, Y.; Song, Y.; Siu, R.; Nimma, I. R.; Yan, Y.; Savage, T. R.; Wang, Y.; Li, Z.; Ramai, D.; Wang, J.; Badurdeen, D.; Tao, C.; Kumbhari, V.; Huang, Y.
Show abstract
ObjectiveTo comprehensively evaluate the validity of ICD-10-CM codes for both prevalent diagnoses and less common diseases, and to assess the performance of a large language model (LLM)-based system in validating these codes. Materials and MethodsThis retrospective study analyzed hospital admissions from the Medical Information Mart for Intensive Care (MIMIC-IV) database. We developed a validated LLM-based system using GPT-4o, refined through iterative prompt engineering, to assess ICD-10-CM code validity. We measured the PPV of ICD-10-CM codes, PPV of principal and secondary diagnoses, and the performance of an LLM-based system in code validation. ResultsAmong 865,079 assigned codes, the PPV was 84.6% (95% CI, 84.5%-84.6%). Principal diagnoses had a PPV of 93.9% (95% CI, 93.7%-94.1%), while secondary diagnoses had a PPV of 83.8% (95% CI, 83.7%-83.9%). The LLM system demonstrated high performance in validating ICD codes, achieving 93.6% accuracy, 95.4% sensitivity and 85.2% specificity. Among correctly assigned secondary diagnoses, the majority (67.9%) represented historical or baseline conditions, while 32.1% reflected active conditions that deviated from baseline status; 22.3% of these emerged after hospital admission. PPV decreases with later diagnosis positions, with the largest decline occurring between principal and secondary diagnoses. Discussion and ConclusionIn this large-scale evaluation, ICD-10-CM codes exhibited generally high accuracy, though variability existed by position and condition type. A validated LLM system performed comparably to physician review and offers a scalable means to improve coding accuracy. These findings support the potential for integrating LLM-based auditing into routine workflows to strengthen the quality of administrative and research data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 96%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 96%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 93%
Similar papers in this journal
- Developing and Evaluating Mappings of ICD-10 and ICD-10-CM Codes to PheCodes 94%
- Can we trust the prediction model? Demonstrating the importance of external validation by investigating the COVID-19 Vulnerability (C-19) Index across an international network of observational healthcare datasets 93%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 92%
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
Similar papers in this journal
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 92%
- Development of a COVID-19 Application Ontology for the ACT Network 92%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.