Can a General-Purpose Coding Agent Analyze a Production Hospital Data Warehouse?
Marshall, N. P.; Haberkorn, W.; Faulkenberry, J.; Palakkattil, G.; Nateghi Haredasht, F.; Schwenk, H. T.; Chen, J. H.; Morse, K.
Show abstract
Background. Health systems answer most questions by having expert analysts hand-write queries against a complex electronic health record data warehouse, a slow, resource-intensive process. Whether an autonomous coding agent can do this accurately is unknown. Methods. In a single-center quality-improvement evaluation, we posed ten questions about a common pediatric infection, acute otitis media. The questions went to analysts, whose adjudicated answers were the reference, and to an autonomous coding agent (OpenAI Codex), which wrote and ran read-only queries on a full copy of the production data warehouse (Epic Caboodle). We ran the agent under four conditions: an autonomous baseline (each question answered three times at two reasoning-effort settings), a variant in which it listed its assumptions, interactive analyst feedback, and reuse of a corrected definition across related questions. Outcomes were accuracy, patient-level agreement (F1), reproducibility, and cost. Results. Working autonomously, the agent wrote valid queries and never fabricated data, but rarely produced the exact answer. At medium effort, it came within 5% of the reference on 27 of 30 runs but exact on only 10. Higher effort produced no improvement. Reproducibility was the greater weakness, with the three runs returning an identical answer on only 3 of 10 questions. When a run matched the reference, it had found the same patients (F1 = 1.00); the exception was an over-counted procedure (F1 = 0.72). With analyst feedback on four questions, it answered two exactly, and reusing the corrected definition on related questions restored reproducibility and accuracy. Conclusions. A general-purpose coding agent matched an adjudicated analyst reference on most routine questions but not reproducibly, because key definitions depended on warehouse knowledge that the data dictionary omits. Supplying that knowledge, by prompt or feedback, restored reproducibility. Under expert analyst supervision, the agent is already a capable drafting aid and a promising step toward broader hospital analytic support.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 93%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 91%
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 91%
Similar papers in this journal
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 93%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.