Development and Internal Validation of a Large Language Model Pipeline for Multi-Label Classification of Patient Portal Messages
Steitz, B. D.; Ogunsan, O. O.; Ancker, J. S.; Carlson, B. R.; Gaynor, L. S.; Higashi, R. T.; Morrow, E. L.; Reese, T. J.; Romano, R. R.; Stern, S.; Turer, R. W.; Rosenbloom, S. T.; Wright, A.
Show abstract
Objectives: Characterizing patient portal message content at scale can help target efforts to manage administrative work. We developed and validated a large language model (LLM) pipeline for multi-label classification of messages using an expert-derived topic taxonomy, then characterized topic distribution across a two-year corpus. Materials and Methods: We studied all medical advice request messages sent to ambulatory clinicians at an academic medical center from 2024-2025. We convened an expert panel that derived an 11-category taxonomy through a modified Delphi process. Two annotators labeled 750 randomly selected messages (Cohen kappa 0.80), holding out 500 for evaluation. The pipeline used GPT-4o-mini in a zero-shot prompt. On the held-out set, we measured micro- and macro-averaged precision, recall, and F1, and label stability across runs. We then characterized topic distribution and co-occurrence across the corpus. Results: The pipeline achieved micro- and macro-averaged F1 of 0.89 and 0.86. Labels were identical across runs for 93.6% of messages. Across 2.4 million messages, content concentrated on a few topics. The two most common topics, Problems & Management and Medications & Prescriptions, were present in 67.9% of messages, and the four most common in 93.9%. 51.7% of messages addressed multiple topics. Discussion and Conclusion: The pipeline classified patient message topics accurately and stably across millions of messages. Message content was concentrated within a small number of topics, highlighting opportunities for targeted interventions and enabling more efficient triage, routing, and patient-facing support.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 94%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 93%
Similar papers in this journal
- Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.