I answer to clinical scenarios as a real doctor because I "Think": "ChatGPTo1" new reasoning ability.
Mondillo, G.; Colosimo, S.; Perrotta, A.; Frattolillo, V.; Masino, M.; Marzuillo, P.
Show abstract
The adoption of the ChatGPT o1 model represents a significant advancement in the management of clinical cases due to the introduction of a new structured reasoning capability, the "chain-of-thought reasoning" (CoT). In this study, 350 general medicine clinical cases were tested using ChatGPT o1 and ChatGPT o1 mini, and their performance was compared with ChatGPT 4o and ChatGPT 4o mini to evaluate diagnostic accuracy. The results showed that ChatGPT o1 achieved a correct answer rate of 93.4%, outperforming both ChatGPT 4o (82.2%) and the mini versions (ChatGPT o1 mini: 70.2% and ChatGPT 4o mini: 66.2%). The CoT technique enabled the model to provide more coherent and transparent responses, reducing the occurrence of so-called "hallucinations." This study highlights how the ChatGPT o1 model can be a valuable tool in clinical practice, although its use requires supervision to ensure patient safety, especially in critical settings.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 95%
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 95%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
Similar papers in this journal
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 94%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 94%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 93%
Similar papers in this journal
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 96%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 94%
Similar papers in this journal
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 96%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 95%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.