Benchmarking Large Language Models for Extraction of International Classification of Diseases Codes from Clinical Documentation
Simmons, A.; Takkavatakarn, K.; McDougal, M.; Dilcher, B.; Pincavitch, J.; Meadows, L.; Kauffman, J.; Klang, E.; Wig, R.; Smith, G.; Soroush, A.; Freeman, R.; Apakama, D. J.; Charney, A.; Kohli-Seth, R.; Nadkarni, G.; Sakhuja, A.
Show abstract
BackgroundHealthcare reimbursement and coding is dependent on accurate extraction of International Classification of Diseases-tenth revision - clinical modification (ICD-10-CM) codes from clinical documentation. Attempts to automate this task have had limited success. This study aimed to evaluate the performance of large language models (LLMs) in extracting ICD-10-CM codes from unstructured inpatient notes and benchmark them against human coder. MethodsThis study compared performance of GPT-3.5, GPT4, Claude 2.1, Claude 3, Gemini Advanced, and Llama 2-70b in extracting ICD-10-CM codes from unstructured inpatient notes against a human coder. We presented deidentified inpatient notes from American Health Information Management Association Vlab authentic patient cases to LLMs and human coder for extraction of ICD-10-CM codes. We used a standard prompt for extracting ICD-10-CM codes. The human coder analyzed the same notes using 3M Encoder, adhering to the 2022-ICD-10-CM Coding Guidelines. ResultsIn this study, we analyzed 50 inpatient notes, comprising of 23 history and physicals and 27 progress notes. The human coder identified 165 unique codes with a median of 4 codes per note. The LLMs extracted varying numbers of median codes per note: GPT 3.5: 7, GPT4: 6, Claude 2.1: 6, Claude 3: 8, Gemini Advanced: 5, and Llama 2-70b:11. GPT 4 had the best performance though the agreement with human coder was poor at 15.2% for overall extraction of ICD-10-CM codes and 26.4% for extraction of category ICD-10-CM codes. ConclusionCurrent LLMs have poor performance in extraction of ICD-10-CM codes from inpatient notes when compared against a human coder.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 94%
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 94%
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- Developing and Evaluating Mappings of ICD-10 and ICD-10-CM Codes to PheCodes 94%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 92%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 92%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.