Loon Lens 1.0 Validation: Agentic AI for Title and Abstract Screening in Systematic Literature Reviews
Janoudi, G.; Rada (Uzun), M.; Jurdana, M.; Fuzul, E.; Ivkovic, J.
Show abstract
IntroductionSystematic literature reviews (SLRs) are critical for informing clinical research and practice, but they are time-consuming and resource-intensive, particularly during Title and Abstract (TiAb) screening. Loon Lens, an autonomous, agentic AI platform, streamlines TiAb screening without the need for human reviewers to conduct any screening. MethodsThis study validates Loon Lens against human reviewer decisions across eight SLRs conducted by Canadas Drug Agency, covering a range of drugs and eligibility criteria. A total of 3,796 citations were retrieved, with human reviewers identifying 287 (7.6%) for inclusion. Loon Lens autonomously screened the same citations based on the provided inclusion and exclusion criteria. Metrics such as accuracy, recall, precision, F1 score, specificity, and negative predictive value (NPV) were calculated. Bootstrapping was applied to compute 95% confidence intervals. ResultsLoon Lens achieved an accuracy of 95.5% (95% CI: 94.8-96.1), with recall at 98.95% (95% CI: 97.57-100%) and specificity at 95.24% (95% CI: 94.54-95.89%). Precision was lower at 62.97% (95% CI: 58.39-67.27%), suggesting that Loon Lens included more citations for full-text screening compared to human reviewers. The F1 score was 0.770 (95% CI: 0.734-0.802), indicating a strong balance between precision and recall. ConclusionLoon Lens demonstrates the ability to autonomously conduct TiAb screening with a substantial potential for reducing the time and cost associated with manual or semi-autonomous TiAb screening in SLRs. While improvements in precision are needed, the platform offers a scalable, autonomous solution for systematic reviews. Access to Loon Lens is available upon request at https://loonlens.com/.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 93%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 93%
Similar papers in this journal
Similar papers in this journal
- Analysis of clinical trial registry entry histories using the novel R package cthist 94%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
- COVID-19 L·OVE repository is highly comprehensive and can be used as a single source for COVID-19 studies 93%
- Diagnostic test accuracy in longitudinal study settings: Theoretical approaches with use cases from clinical practice 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.