Assessing the performance of GPT-4 in the filed of osteoarthritis and orthopaedic case consultation
li, j.; Gao, X.; Dou, T.; Gao, Y.; Zhu, W.
Show abstract
BackgroundLarge Language Models (LLMs) like GPT-4 demonstrate potential applications in diverse areas, including healthcare and patient education. This study evaluates GPT-4s competency against osteoarthritis (OA) treatment guidelines from the United States and China and assesses its ability in diagnosing and treating orthopedic diseases. MethodsData sources included OA management guidelines and orthopedic examination case questions. Queries were directed to GPT-4 based on these resources, and its responses were compared with the established guidelines and cases. The accuracy and completeness of GPT-4s responses were evaluated using Likert scales, while case inquiries were stratified into four tiers of correctness and completeness. ResultsGPT-4 exhibited strong performance in providing accurate and complete responses to OA management recommendations from both the American and Chinese guidelines, with high Likert scale scores for accuracy and completeness. It demonstrated proficiency in handling clinical cases, making accurate diagnoses, suggesting appropriate tests, and proposing treatment plans. Few errors were noted in specific complex cases. ConclusionsGPT-4 exhibits potential as an auxiliary tool in orthopedic clinical practice and patient education, demonstrating high accuracy and completeness in interpreting OA treatment guidelines and analyzing clinical cases. Further validation of its capabilities in real-world clinical scenarios is needed.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Investigation of locomotive syndrome improvement by total hip arthroplasty in patients with hip osteoarthritis: a before-after comparative study focusing on 25-question geriatric locomotive function scale 96%
- Intra- and inter-rater reliability, agreement, and minimal detectable change of the handheld dynamometer in individuals with symptomatic hip osteoarthritis 95%
- Action observation intervention using three - dimensional movies improves the usability of hands with distal radius fractures in daily life: a nonrandomized controlled trial in women 95%
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 93%
- Natural Language Processing Algorithms Outperform ICD Codes in the Development of Fall Injuries Registry 92%
- Model-based reasoning methods for diagnosis in integrative medicine based on electronic medical records and natural language processing 91%
Similar papers in this journal
- A multiscale modeling approach to study the role of mechanics and inflammation in pathophysiology of articular cartilage 91%
- SymScore: Machine Learning Accuracy Meets Transparency in a Symbolic Regression-Based Clinical Score Generator 91%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 91%
Similar papers in this journal
- Are Electromyography data a fingerprint for patients with cerebral palsy (CP)? 93%
- Angle-angle diagrams in the assessment of locomotion in persons with multiple sclerosis: A preliminary study 91%
- Genetic deletion of interleukin-15 is not associated with major structural changes following experimental post-traumatic knee osteoarthritis in rats 91%
Similar papers in this journal
- Standardizing Phenotypic Algorithms for the Classification of Degenerative Rotator Cuff Tear from Electronic Health Record Systems 91%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 90%
- Modeling physician variability to prioritize relevant medical record information 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.