VITRUVIUS: A conversational agent for real-time, evidence-based medical question-answering
Villa, M. C.; Llano, I.; Castano-Villegas, N.; Martinez, J.; Guevara, M. F.; Zea, J.; Velasquez, L.
Show abstract
BackgroundThe application of Large Language Models (LLMs) to create conversational agents (CAs) that can aid health professionals in their daily practice is increasingly popular, mainly due to their ability to understand and communicate in natural language. Conversational agents can manage enormous amounts of information, comprehend and reason with clinical questions, extract information from reliable sources and produce accurate answers to queries. This presents an opportunity for better access to updated and trustworthy clinical information in response to medical queries. ObjectiveWe present the design and initial evaluation of Vitruvius, an agent specialized in answering queries in healthcare knowledge and evidence-based medical research. MethodologyThe model is based on a system containing 5 LLMs; each is instructed with precise tasks that allow the algorithms to automatically determine the best search strategy to provide an evidence-based answer. We assessed our systems comprehension, reasoning, and retrieval capabilities using the public clinical question-answer dataset MedQA-USMLE. The model was improved accordingly, and three versions were manufactured. ResultsWe present the performance assessment for the three versions of Vitruvius, using a subset of 288 QA (Accuracy V1 86%, V2 90%, V3 93%) and the complete dataset of 1273 QA (Accuracy V2 85%, V3 90.3%). We also evaluate intra-inter-class variability and agreement. The final version of Vitruvius (V3) obtained a Cohens kappa of 87% and a state-of-the-art (SoTA) performance of 90.26%, surpassing current SoTAs for other LLMs using the same database. ConclusionsVitruvius demonstrates excellent performance in medical QA compared to standard database responses and other popular LLMs. Future investigations will focus on testing the model in a real-world clinical environment. While it enhances productivity and aids healthcare professionals, it should not be utilized by individuals unqualified to reason with medical data to ensure that critical decision-making remains in the hands of trained professionals.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 96%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 97%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
Similar papers in this journal
Similar papers in this journal
- Medication information extraction using local large language models 95%
- Detecting Goals of Care Conversations in Clinical Notes with Active Learning 95%
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.