Large Language Models Improve Coding Accuracy and Reimbursement in a Neonatal Intensive Care Unit
Holmes, E.; Massarelli, C.; Richter, F.; Bernard, S.; Freeman, R.; Gavin, N.; Juliano, C.; Gelb, B. D.; Glicksberg, B. S.; Nadkarni, G. N.; Klang, E.
Show abstract
ImportanceDiagnosis coding is essential for clinical care, research validity, and hospital reimbursement. In neonatal settings, manual coding is frequently error-prone, contributing to misclassification and financial losses. Large language models (LLMs) offer a scalable approach to improve diagnostic consistency and optimize revenue. ObjectiveTo compare the diagnostic accuracy of LLMs with human coders in identifying common neonatal diagnoses and assess the potential impact on revenue from Diagnosis-Related Group (DRG) assignment. DesignThis was a retrospective cross-sectional study conducted using data from 2022 to 2023. LLMs were prompted with all physician notes from the admission. Two neonatologists independently and blindly adjudicated diagnoses from three sources: human coders, GPT-4o, and GPT-o3-mini. SettingA single academic referral centers neonatal intensive care unit (NICU). ParticipantsThe study included a consecutive sample of 100 infants admitted to the NICU who did not require respiratory support. All available physician notes from the hospital stay were included. ExposureTwo HIPAA-compliant LLMs (GPT-4o and GPT-o3-mini) were prompted to assign diagnoses from a standardized list based on physician notes. Three prompt iterations were developed and reviewed for optimization prior to final evaluation. Main Outcomes and MeasuresThe primary outcome was diagnostic accuracy compared with physician adjudication. Secondary outcomes included changes in expected DRG assignment and projected annual revenue. ResultsAmong 100 infants (median gestational age 35.6 weeks, 52% male), GPT-o3-mini achieved 79.1% diagnostic accuracy (95% CI, 74.0%-84.2%), comparable to human coders at 76.3% (95% CI, 70.9%-81.7%; P = .52). GPT-4o underperformed at 58.6% (95% CI, 52.5%-64.7%; P < .001 vs both). Accuracy of GPT-o3-mini did not differ by DRG impact. Extrapolated to one year, correct GPT-o3-mini diagnoses yielded projected revenue of $5.71 million, compared to $4.82 million from human coders, an 18% increase. Conclusions and RelevanceA HIPAA-compliant LLM demonstrated diagnostic accuracy comparable to human coders in neonatal billing while identifying higher-acuity diagnoses that improved projected reimbursement. LLMs may serve as effective adjuncts to manual coding in neonatal care, with potential clinical and financial benefit. Key PointsO_ST_ABSQuestionC_ST_ABSCan a large language model support accurate diagnosis generation for neonatal billing? FindingsIn this retrospective study of 100 neonates hospitalized in the Neonatal Intensive Care Unit, GPT-o3-mini demonstrated diagnostic accuracy comparable to human coders, as confirmed by physician review. Its implementation could yield an estimated 18% increase in revenue. MeaningLarge language models may serve as effective adjuncts in neonatal coding, offering both diagnostic precision and financial benefit.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 94%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- A comprehensive digital phenotype for postpartum hemorrhage 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 92%
- Regional Variation in Antenatal Late Preterm Steroid Use following the ALPS Trial 92%
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 92%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 92%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 92%
Similar papers in this journal
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 93%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 93%
- Ranked severe maternal morbidity index for population-level surveillance at delivery hospitalization based on hospital discharge data 93%
Similar papers in this journal
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 92%
- Identification of an ANCA-Associated Vasculitis Cohort Using Deep Learning and Electronic Health Records 90%
- Completion of electronic nursing documentation of inpatient admission assessment: insights from Australian metropolitan hospitals 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.