ENTAgents: AI Agents for Complex Knowledge Otolaryngology
Dinh, N.-D.; Chan, T. K.
Show abstract
Various healthcare applications based on large language models (LLMs) have emerged as LLMs show improved efficiency and error reduction. Recently, retrieval augmented generation (RAG) has been adopted frequently for LLM applications to solve the problem of hallucinations. Despite the success of RAG, it has its drawbacks, including incomplete semantic meanings, and large-scale dataset requirements. AI Agents have shown great potential in medicine and healthcare applications by leveraging their rich background knowledge and reasoning capabilities. In this paper, we introduce ENTAgents, a framework that utilizes both the RAG and multi-agent systems. To achieve better decision-making and enhanced explainability, we use the verbal reinforcement learning framework, Reflexion, as a reference for the task planning of ENTAgents. With the verbal feedback received from the last agent node, the agentic system can decide the best agent to utilize to improve the response of ENTAgents. We tested ENTAgents with various types of questions, including short questions, essay questions, and multiple-choice questions. ENTAgents improve accuracy by over 11.3% in handling short questions, with 2.78 folds in the length of the text and explaining clearly. ENTAgents shows higher accuracy and is more comprehensive in its answer compared to other LLM models. ENTAgents also demonstrates its ability to refine its answer with additional information to make it more thorough and educational. It also presents the capability to change its response to the correct answer in multiple-choice questions according to its self-reflection. Overall, ENTAgents is easy to use, can self-correct, and can provide detailed information in complex knowledge to users. We look forward to integrating ENTAgents into various healthcare scenarios, such as medical education and clinical decision support, particularly otolaryngology.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 94%
- Regional medical inter-institutional cooperation in medical provider network constructed using patient claims data from Japan 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
Similar papers in this journal
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 92%
Similar papers in this journal
- Improving Tuberculosis Detection in Chest X-ray Images through Transfer Learning and Deep Learning: A Comparative Study of CNN Architectures 90%
- Learning from the resilience of hospitals and their staff to the COVID-19 pandemic: a scoping review 90%
- Google Trends as a predictive tool for COVID-19 vaccinations in Italy: a retrospective infodemiological analysis 89%
Similar papers in this journal
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 96%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.