Real-world Validation of MedSearch: a conversational agent for real-time, evidence-based medical question-answering.
Castano-Villegas, N.; Llano, I.; Villa, M. C.; Zea, J.
Show abstract
IntroductionApplication of Large Language Models (LLMs) powered Conversation Agents (CAs) in healthcare has been evaluated using medical question-answering (QA) datasets, with excellent performance in international medical licensing exams [1, 2, 3]. However, multiple-choice questions fall short when the intention is to assess more complex language interactions and open-ended responses. Objectiveto evaluate the time invested and the validity of health care personnels (HCP) responses to clinical questions using Med-Search compared to traditional search methods without AI. MethodsThis was a randomized, double-blind trial with 100 physicians assigned to two groups. Each group answered four clinical cases with four questions, one group using MedSearch and the other using traditional research methods such as Google and PubMed, except AI. Field specialists evaluated responses in six aspects established to define the validity of answers. Time to respond was also recorded, described, and compared between the two groups. ResultsMore than 70% of the sample were medical students. Differences in results between groups were statistically significant in all evaluated aspects (p <0.01): the intervention (MedSearch) group arrived at a final answer in half the time (three minutes faster) of the control group (traditional research methods), with approximately 66% fewer searches per case. The models answers were valid (accurate, current, aligned with consensus, and safe) with an average score of 2.8 on a scale from 1 to 3. Most MedSearch users found it useful for daily practice and would recommend it to colleagues. Conclusionthe present results suggest a positive impact of LLM-supported methods for a more effective clinical search, without sacrificing, and even augmenting the quality of answers. More clinical validations are needed to understand further the effect of LLMs use in education and clinical practice, using broader sample sizes and across professionals from different fields.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 95%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 95%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 94%
Similar papers in this journal
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 95%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 94%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 97%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 96%
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.