MedRAGent: An Automatic Literature Retrieval and Screening System Utilizing Large Language Models with Retrieval-Augmented Generation
Chen, Z.; Liu, T.; Mo, Y.; Fu, Q.; Lei, S.; Tong, T.; Tang, X.
Show abstract
BackgroundSystematic reviews play a critical role in synthesizing evidence across numerous studies, providing a foundation for informed decision-making in medical practice. However, the process is resource-intensive, requiring proficiency in constructing Boolean queries and screening extensive literature, which are time-consuming and susceptible to inconsistencies, especially for non-expert researchers. While large language models (LLMs) offer a potential solution, their tendency to generate inaccurate or hallucinated content restricts their direct application in systematic reviews. ObjectiveThis study introduces and evaluates MedRAGent, a novel system that integrates LLMs with retrieval-augmented generation (RAG), designed to automate and enhance the efficiency and accuracy of Boolean query formulation and title/abstract screening in systematic reviews. MethodsMedRAGent employs DeepSeek-V3-0324 and Kimi-K2-0711-preview LLMs within an RAG framework tailored for PubMed. The system utilizes the official Medical Subject Headings (MeSH) database to construct precise Boolean queries. For screening, it employs the LLMs with a structured prompt to automatically evaluate the relevance of retrieved articles based on predefined inclusion and exclusion criteria. Its performance was assessed using 53,054 articles from 6 research topics. ResultsOur results showed that MedRAGent achieved an overall precision of 0.0271, recall of 0.8308, and F1-score of 0.0525 in Boolean query construction. For automated literature screening, the system attained an overall sensitivity of 0.8131, specificity of 0.9891, and G-mean of 0.8968 when using DeepSeek-V3-0324 as the underlying LLM. Performance improved when using Kimi-K2-0711-preview, with sensitivity of 0.8582, specificity of 0.9919, and G-mean of 0.9226. It efficiently processed 4,000-7,000 articles per day at low operational cost. ConclusionsMedRAGent demonstrates strong potential for automating Boolean query construction and abstract-level screening in systematic reviews. It effectively accelerates literature processing, supporting researchers in conducting efficient and evidence-based medical reviews.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 95%
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 94%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 94%
Similar papers in this journal
- Catchii: empowering literature review screening in healthcare 96%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 95%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 95%
Similar papers in this journal
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
- Predicting Emerging Themes in Rapidly Expanding COVID-19 Literature with Dynamic Word Embedding Networks and Machine Learning 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.