General-Purpose vs. Domain-Specific Large Language Models in Antibiotic Clinical Decision-Making: A Double-Blind Evaluation with a 2X2 Factorial Design
Liu, Y.; Zhang, C.; Wang, F.; Xu, W.; Zhang, Y.; Ma, S.; zhang, H.
Show abstract
Background: Antimicrobial resistance poses a major threat to global public health. Large language models (LLMs) offer new possibilities for optimizing antibiotic prescribing decisions, but the capabilities of general-purpose versus domain-specific medical LLMs under different prompting strategies remain to be clarified. Methods: This double-blind, randomized-sequence evaluation used a 2X2 factorial design comparing four AI conditions-the domain-specific model MedGo and the general-purpose model DeepSeek V3.5, each under standard direct prompting and chain-of-thought (CoT) prompting-alongside real physician prescriptions across 59 complex inpatient infection cases. Five parallel regimens were generated per case and independently evaluated by three senior clinicians (1-5 comprehensive score and five domain sub-scores). ChatGPT 5.2 was additionally assessed as an automated evaluation tool. Results: Score ranking: real physicians > MedGo-CoT > DeepSeek-CoT > MedGo> DeepSeek (Friedman test, p<0.001). In base mode, MedGo significantly outperformed DeepSeek (Holm-adjusted p=0.040). CoT improved both models (Holm-adjusted p<0.001 for DeepSeek; p=0.024 for MedGo) and reduced score dispersion. MedGo-CoT significantly outperformed DeepSeek-CoT in individualized adjustment (adjusted p<0.001) and dosing precision (adjusted p=0.005). ChatGPT-expert correlation was negligible (overall Kendall {tau}=0.153, p=0.003; subgroup {tau}=0.06-0.20, all p>0.05). Conclusions: Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision. However, notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
Similar papers in this journal
Similar papers in this journal
- Informing antimicrobial stewardship with explainable AI 95%
- Community-acquired pneumonia identification from electronic health records in the absence of a gold standard: a Bayesian latent class analysis 92%
- ePOCT+ and the medAL-suite: Development of an electronic clinical decision support algorithm and digital platform for pediatric outpatients in low- and middle-income countries 92%
Similar papers in this journal
- Serum Concentration of Continuously administered Vancomycin influences Efficacy and Safety in Critically Ill Adults: A Systematic Review 91%
- No one-size-fits-all approach: Retrospective analysis of efficacy and safety of serum concentrations of continuously administered vancomycin in critically ill adults reveals different target serum concentrations depending on disease severity 91%
- Analysis of time-to-positivity data in tuberculosis treatment studies: Identifying a new limit of quantification 91%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 93%
- A Sepsis Treatment Algorithm to Improve Early Antibiotic De-escalation While Maintaining Adequacy of Coverage (Early-IDEAS): A Prospective Observational Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.