Large language models to differentiate vasospastic angina using patient information
Kiyohara, Y.; Kodera, S.; Sato, M.; Ninomiya, K.; Sato, M.; Shinohara, H.; Takeda, N.; Akazawa, H.; Morita, H.; Komuro, I.
Show abstract
BackgroundVasospastic angina is sometimes suspected from patients medical history. It is essential to appropriately distinguish vasospastic angina from acute coronary syndrome because its standard treatment is pharmacotherapy, not catheter intervention. Large language models have recently been developed and are currently widely accessible. In this study, we aimed to use large language models to distinguish between vasospastic angina and acute coronary syndrome from patient information and compare the accuracies of these models. MethodWe searched for cases of vasospastic angina and acute coronary syndrome which were written in Japanese and published in online-accessible abstracts and journals, and randomly selected 66 cases as a test dataset. In addition, we selected another ten cases as data for few-shot learning. We used generative pre-trained transformer-3.5 and 4, and Bard, with zero- and few-shot learning. We evaluated the accuracies of the models using the test dataset. ResultsGenerative pre-trained transformer-3.5 with zero-shot learning achieved an accuracy of 52%, sensitivity of 68%, and specificity of 29%; with few-shot learning, it achieved an accuracy of 52%, sensitivity of 26%, and specificity of 86%. Generative pre-trained transformer-4 with zero-shot learning achieved an accuracy of 58%, sensitivity of 29%, and specificity of 96%; with few-shot learning, it achieved an accuracy of 61%, sensitivity of 63%, and specificity of 57%. Bard with zero-shot learning achieved an accuracy of 47%, sensitivity of 16%, and specificity of 89%; with few-shot learning, this model could not be assessed because it failed to produce output. ConclusionGenerative pre-trained transformer-4 with few-shot learning was the best of all the models. The accuracies of models with zero- and few-shot learning were almost the same. In the future, models could be made more accurate by combining text data with other modalities.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 93%
- Detecting Heart Failure using novel bio-signals and a Knowledge Enhanced Neural Network 92%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 92%
Similar papers in this journal
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
- Electrocardiogram Analysis of Post-Stroke Elderly People Using One-dimensional Convolutional Neural Network Model with Gradient-weighted Class Activation Mapping 93%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 92%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
Similar papers in this journal
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 93%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.