Automatic ICD coding using LLMs: a systematic review
Gershon, A.; Soffer, S.; Klang, E.; Nadkarni, G.
Show abstract
BackgroundManual assignment of International Classification of Diseases (ICD) codes is error-prone. Transformer-based large language models (LLMs) have been proposed to automate coding, but their accuracy and generalizability remain uncertain. MethodsWe performed a systematic review registered with PROSPERO (CRD42024576236) and reported according to PRISMA guidelines. PubMed, Embase, and Google Scholar were searched through January 2025 for peer-reviewed studies that evaluated an LLM (e.g., BERT, GPT) for ICD coding and reported at least one performance metric. Two reviewers independently screened articles, extracted data, and assessed methodological quality with the Joanna Briggs Institute Critical Appraisal Checklist for Analytical Cross-Sectional Studies. Outcomes included micro-F1, macro-F1, accuracy, precision, recall, and AUC, capturing both overall predictive performance and sensitivity to rare ICD codes. ResultsOf 590 records screened, 35 studies met the inclusion criteria. 24 assessed general-purpose coding across broad clinical text, 10 focused on specific clinical contexts, and 11 addressed multilingual interoperability; some studies belonged to more than one theme. Median micro-F1 for frequent codes was 0.79 (range, 0.73-0.94), exceeding that of legacy machine-learning baselines in all comparative studies. Performance for infrequent codes was lower (median macro-F1, 0.42) but improved modestly with data augmentation, contrastive retrieval, or graph-based decoders. Only 1 study used federated learning across institutions, and 3 conducted external validation. The risk-of-bias assessment rated 18 studies (51%) as moderate, primarily due to unclear blinding of assessors and selective reporting. ConclusionsLLM-based systems reliably automate common ICD codes and frequently match or surpass professional coders, but accuracy declines for rare diagnoses, and external validation is scant. Prospective, multicenter trials and transparent reporting of prompts and post-processing rules are required before clinical deployment.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 93%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 92%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 92%
Similar papers in this journal
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 93%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 93%
- Medication information extraction using local large language models 93%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 89%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.