Validation of Synthesa AI, a Large Language Model-Based Screening Tool for Systematic Reviews: Results from Nine Studies
Teperikidis, L.; Polymenakis, K.; Trampoukis, C.
Show abstract
Systematic review screening is often burdensome, prone to human error, and requires significant manual effort. Synthesa AI, a large language model (LLM)-based tool, was developed to address these challenges by offering a transparent and prompt-driven approach to abstract screening. In this validation study, Synthesa AI was evaluated across 17 benchmark meta-analyses encompassing nine clinical domains. Using user-defined PICOS criteria, the tool screened a total of 270,626 abstracts retrieved from PubMed and Scopus. Synthesa AI accurately identified all 163 benchmark-included studies, yielding a sensitivity of 100% and a pooled specificity of 99.4%. Remarkably, it reduced reviewer workload by 91.7%, flagging only 1,797 abstracts for manual review. Furthermore, the tool identified 32 relevant studies that had been missed in the original reviews, representing a 19.6% increase in evidence yield. These findings demonstrate that Synthesa AI delivers high precision, efficiency, and reproducibility in systematic review workflows. Its auditable and deterministic architecture adheres to Good Machine Learning Practice (GMLP) guidelines, making it suitable for both academic and regulatory applications. Synthesa AI represents a promising solution for living systematic reviews and large-scale evidence synthesis initiatives, offering a transformative alternative to traditional human-led screening.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 97%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
Similar papers in this journal
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 95%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 95%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 95%
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 91%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.