Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using synthetic controls to estimate the population-level effects of Ontario's recently implemented overdose prevention sites and consumption and treatment services 90%
- Diagnosis and treatment of opioid-related disorders in a South African private sector medical insurance scheme: a cohort study 90%
- Characterizing Declines in US Overdose Deaths Compared to Exponential Predictions 89%
Similar papers in this journal
- The impact of later trading hours for bars and clubs on alcohol-related ambulance call-outs and crimes in Scotland: a controlled interrupted time series study 92%
- Hospital admissions among adolescents with local authority care experience or special educational needs in England: a population-based cohort study using linked administrative data from health, education and social care services 86%
- COVID-19 mass testing in adult social care in England 85%
Similar papers in this journal
- UK medical students’ self-reported knowledge and harm assessment of psychedelics and their application in clinical research: a cross-sectional study 92%
- Prevalence and characteristics of hazardous and harmful drinkers receiving general practitioners’ brief advice on and support with alcohol consumption in Germany: results of a population survey 91%
- Neutrophil-lymphocyte ratio across psychiatric diagnoses: An electronic health record investigation 91%
Similar papers in this journal
Similar papers in this journal
- The impact of the COVID-19 pandemic on health service utilisation following self-harm: a systematic review 92%
- Ethnic inequalities in compulsory psychiatric hospital detentions during UK COVID-19 'lockdowns': A Regression Discontinuity Design in time study 92%
- A thematic analysis of Prevention of Future Death Reports for Children who died by suicide in England and Wales: January 2015 to November 2023 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.