Evaluating Loon Lens Pro, an AI-Driven Tool for Full-Text Screening in Systematic Reviews: A Validation Study
Janoudi, G.; Uzun, M.; Jurdana, M.; Hutton, B.
Show abstract
BackgroundSystematic literature reviews (SLRs) are essential for evidence synthesis but are hampered by the resource-intensive full-text screening phase. Loon Lens Pro, a publicly available agentic AI tool, automates full-text screening without prior training by using user-defined inclusion/exclusion criteria and multiple specialized AI agents. This study validated Loon Lens Pro against human reviewers to assess its accuracy, efficiency, and confidence scoring in screening. MethodsIn this comparative validation study, 84 full-text articles from eight SLRs were screened by both Loon Lens Pro and human reviewers (gold standard). The AI provided binary inclusion/exclusion decisions along with a transparent rationale and confidence ratings (low, medium, high). Performance metrics-- including accuracy, sensitivity, specificity, negative predictive value, precision, and F1 score--were derived from a confusion matrix. Logistic regression with bootstrap resampling (1,000 iterations) evaluated the association between confidence scores and screening errors. ResultsLoon Lens Pro correctly classified 70 of 84 full texts, achieving an accuracy of 83.3% (95% CI: 75.0- 90.5%), sensitivity of 94.7% (95% CI: 82.4-100%), and specificity of 80.0% (95% CI: 70.1-89.2%). The negative predictive value was 98.1% (95% CI: 93.8-100%), with a precision of 58.1% (95% CI: 41.4- 76.0%) and an F1 score of 0.72. Logistic regression revealed a strong inverse relationship between confidence level and error probability: low, medium, and high confidence decisions were associated with predicted error probabilities of 46.9%, 30.9%, and 3.5%, respectively (C-index = 0.87). ConclusionOur study provides evidence that Loon Lens Pro is a viable and effective tool for automating the full-text screening phase of systematic reviews. Its high sensitivity, robust confidence scoring mechanism, and transparent rationale generation collectively support its potential to alleviate the burden of manual screening without compromising the quality of study selection.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Catchii: empowering literature review screening in healthcare 96%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 96%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 95%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 94%
Similar papers in this journal
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 95%
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 94%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.