Unleashing the Power of Language Models in Clinical Settings: A Trailblazing Evaluation Unveiling Novel Test Design
Li, Q.; Min, X.
Show abstract
The realm of clinical medicine stands on the brink of a revolutionary breakthrough as large language models (LLMs) emerge as formidable allies, propelled by the prowess of deep learning and a wealth of clinical data. Yet, amidst the disquieting specter of misdiagnoses haunting the halls of medical treatment, LLMs offer a glimmer of hope, poised to reshape the landscape. However, their mettle and medical acumen, particularly in the crucible of real-world professional scenarios replete with intricate logical interconnections, re-main shrouded in uncertainty. To illuminate this uncharted territory, we present an audacious quantitative evaluation method, harnessing the ingenuity of tho-racic surgery questions as the litmus test for LLMs medical prowess. These clinical questions covering various diseases were collected, and a test format consisting of multi-choice questions and case analysis was designed based on the Chinese National Senior Health Professional Technical Qualification Examination. Five LLMs of different scales and sources were utilized to answer these questions, and evaluation and feedback were provided by professional thoracic surgeons. Among these models, GPT-4 demonstrated the highest performance with a score of 48.67 out of 100, achieving accuracies of 0.62, 0.27, and 0.63 in single-choice, multi-choice, and case-analysis questions, respectively. However, further improvement is still necessary to meet the passing threshold of the examination. Additionally, this paper analyzes the performance, advantages, disadvantages, and risks of LLMs, and proposes suggestions for improvement, providing valuable insights into the capabilities and limitations of LLMs in the specialized medical domain.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 98%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
Similar papers in this journal
- Large language models (GPT-5, Grok-4, Claude Opus 4.1, Gemini 2.5 Pro) achieved textbook-level accuracy on the Japanese medical licensing examination by 2025: A comparative study 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 93%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
Similar papers in this journal
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 93%
- A fast, accurate, and generalisable heuristic-based negation detection algorithm for clinical text 93%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 93%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 95%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 94%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.