Orchestrated multi agents sustain accuracy under clinical scale workloads compared to a single agent
Omar, M.; Klang, E.; Ganesh, J.; Agbareia, R.; Tsimina, P.; Freeman, R.; Gavin, N.; Stump, L.; Charney, A. W.; Glicksberg, B. S.; Nadkarni, G.
Show abstract
We tested state-of-the-art large language models (LLMs) in two configurations for clinical-scale workloads: a single agent handling heterogeneous tasks versus an orchestrated multi-agent system assigning each task to a dedicated worker. Across retrieval, extraction, and dosing calculations, we varied batch sizes from 5 to 80 to simulate clinical traffic. Multi-agent runs maintained high accuracy under load (pooled accuracy 90.6% at 5 tasks, 65.3% at 80) while single-agent accuracy fell sharply (73.1% to 16.6%), with significant differences beyond 10 tasks (FDR-adjusted p < 0.01). Multi-agent execution reduced token usage up to 65-fold and limited latency growth compared with single-agent runs. The designs isolation of tasks prevented context interference and preserved performance across four diverse LLM checkpoints. This is the first evaluation of LLM agent architectures under sustained, mixed-task clinical workloads, showing that lightweight orchestration can deliver accuracy, efficiency, and auditability at operational scale.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero Shot Health Trajectory Prediction Using Transformer 93%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 92%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 92%
Similar papers in this journal
- Application of Generative Artificial Intelligence to Utilise Unstructured Clinical Data for Acceleration of Inflammatory Bowel Disease Research 90%
- Multimodal surveillance of SARS-CoV-2 at a university enables development of a robust outbreak response framework 88%
- The Medical Action Ontology: A Tool for Annotating and Analyzing Treatments and Clinical Management of Human Disease 86%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.