Clinical Med students' validation of Arkangel AI: Are their responses any better when supported by the AI?
Castano-Villegas, N.; Llano, I.; Villa, C.; Zea, J.; Velasquez, L.
Show abstract
IntroductionLarge Language Models (LLMs) in healthcare practice and education have been evaluated using medical question-answering (QA) datasets, with excellent performance. However, multiple-choice questions fall short when assessing more complex language interactions. ObjectiveTo evaluate the time invested and validity of medical students responses to clinical questions using ArkangelAI, compared to traditional search methods. MethodsRandomized, double-blind trial with clinical medical students assigned to two groups. Each group answered four clinical questions from each of four clinical cases, one using ArkangelAI, the other using traditional research methods: Google, PubMed, etc. Field specialists evaluated the responses using six pre-established criteria to define the answers validity. Total average validity (the mean of individual scores) was compared by groups with hypothesis testing and 95% CI. The time to respond was also compared. ResultsEighty-three medical students were randomized to groups A (43) and B (40). Average differences responded in half the time (three minutes faster) than the control group, with 98% fewer searches needed. The models answers were valid (accurate, non-biased, aligned with consensus, and safe) with a total validity score of 2.84 (group A) and 2.69 (group B). Most Arkangel AI users found it helpful for daily practice and would recommend it to colleagues. Conclusion: LLM-supported methods appear to have a positive influence on effective clinical search without sacrificing, and even augmenting, the quality of answers. This is applicable at the clinical medical student level and for non-critical clinical reasoning. Validations, including graduated physicians and specialists, are needed to further understand the effect of LLMs in education and clinical practice.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 96%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 93%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 94%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 94%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 95%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 93%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 97%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 96%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.