Artificial Intelligence in Diabetes Care: Evaluating GPT-4's Competency in Reviewing Diabetic Patient Management Plan in Comparison to Expert Review
Mondal, A.; Naskar, A.
Show abstract
BackgroundThe escalating global burden of diabetes necessitates innovative management strategies. Artificial intelligence, particularly large language models like GPT-4, presents a promising avenue for improving guideline adherence in diabetes care. Such technologies could revolutionize patient management by offering personalized, evidence-based treatment recommendations. MethodsA comparative, blinded design was employed, involving 50 hypothetical diabetes mellitus case summaries, emphasizing varied aspects of diabetes management. GPT-4 evaluated each summary for guideline adherence, classifying them as compliant or non-compliant, based on the ADA guidelines. A medical expert, blinded to GPT-4s assessments, independently reviewed the summaries. Concordance between GPT-4 and the experts evaluations was statistically analyzed, including calculating Cohens kappa for agreement. ResultsGPT-4 labelled 30 summaries as compliant and 20 as non-compliant, while the expert identified 28 as compliant and 22 as non-compliant. Agreement was reached on 46 of the 50 cases, yielding a Cohens kappa of 0.84, indicating near-perfect agreement. GPT-4 demonstrated a 92% accuracy, with a sensitivity of 86.4% and a specificity of 96.4%. Discrepancies in four cases highlighted challenges in AIs understanding of complex clinical judgments related to medication adjustments and treatment modifications. ConclusionGPT-4 exhibits promising potential to support health-care professionals in reviewing diabetes management plans for guideline adherence. Despite high concordance with expert assessments, instances of non-agreement underscore the need for AI refinement in complex clinical scenarios. Future research should aim at enhancing AIs clinical reasoning capabilities and exploring its integration with other technologies for improved healthcare delivery.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 94%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 93%
- Using convolutional neural network to predict remission of diabetes after gastric bypass surgery: a machine learning study from the Scandinavian Obesity Surgery Register 93%
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 92%
Similar papers in this journal
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 92%
- Genetically Guided Precision Medicine Clinical Decision Support Tools: A Systematic Review 92%
Similar papers in this journal
- SoK: Intelligent Detection for Polycystic Ovary Syndrome(PCOS) 89%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 88%
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.