Large language model scoring of medical student reflection essays: Accuracy and reproducibility of prompt-model variations
Cook, D. A.; Laack, T. A.; Pankratz, V. S.
Show abstract
Purpose: Evaluate large language models (LLMs) for scoring medical student essays, and compare various prompting techniques and models. Methods: OpenAI GPT scored 51 medical student reflection essays (15 real, 36 fabricated) using a previously-reported 6-point rubric (April-May 2025). We compared 29 prompt-model conditions by systematically varying the LLM prompts (including the persona, scoring rubric, few-shot learning [exemplars], chain-of-thought reasoning, and temperature), fine-tuning, and model (including GPT-4.1, GPT-4.1-mini, GPT-o4-mini, and GPT-4-Turbo). Outcomes were accuracy (compared with human raters, measured using single-score intraclass correlation coefficient [ICC] and mean absolute difference [MAD; zero indicates perfect agreement]), within-condition reproducibility, and cost. Results: Across all conditions, it took mean (SD) 3.73 (3.12) seconds to score 1 essay. The cost to score 100 essays was USD $0.04 for GPT-4.1-mini, $0.21 for GPT-4.1, $0.57 for GPT-4.1 with 3 exemplars, and $2.00 for fine-tuned GPT-4.1. When the one-time cost of fine-tuning was amortized across 10,000 essays, the cost for fine-tuned GPT-4.1 was $0.20 per 100. Accuracy was "almost perfect" (ICC >0.80) for 28/29 conditions (97%). Fine-tuned models were more accurate than non-fine-tuned models (MAD difference -0.24 [95% CI, -0.34, -0.14]). Conditions with exemplars were more accurate than those without (MAD difference -0.44 [CI, -0.57, -0.31]). Accuracy progressively decreased as 6, 3, 1, and 0 rubric levels were explicitly defined in the prompt (P<.001). Contrary to hypotheses, accuracies for chain-of-thought prompts and variations in temperature and persona were not significantly different from the baseline prompt. Reproducibility ICC was >0.80 for 28/29 conditions (97%). Discussion: Automated LLM essay scoring demonstrated near-perfect accuracy and reproducibility for most prompt-model conditions. Fine-tuned models and prompts with exemplars had higher accuracy but higher cost. Fine-tuned models had lower per-essay costs for larger essay volumes. For smaller volumes, non-fine-tuned GPT-4.1 provided excellent results at moderate cost. GPT-4.1-mini provided very good results at low cost.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 92%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
Similar papers in this journal
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 93%
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 91%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 91%
Similar papers in this journal
- Quality of Human Expert vs. Large Language Model Generated Multiple Choice Questions in the Field of Mechanical Ventilation 92%
- Development and Prospective Validation of a Transparent Deep Learning Algorithm for Predicting Need for Mechanical Ventilation 86%
- Effect of Ventilator Mode on Ventilator-Free Days in Critically Ill Adults: A Randomized Trial 83%
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 92%
- Fitness tracking reveals task-specific associations between memory, mental health, and physical activity 91%
- Solitary Silence and Social Sounds: Music influences mental imagery, inducing thoughts of social interactions 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.