Dual-Model LLM Ensemble via Web Chat Interfaces Reaches Near-Perfect Sensitivity for Systematic-Review Screening: A Multi-Domain Validation with Equivalence to API Access
Fagerberg, P.; Sallander, O.; Vikhe Patil, K.; Berg, A.; Nyman, A.; Borg, N.; Linden, T.
Show abstract
BackgroundPrior work showed that state-of-the-art (mid-2025) large language models (LLMs) prompted with varying batch sizes can perform well on systematic review (SR) abstract screening via public APIs within a single medical domain. Whether comparable performance holds when using no-code web interfaces (GUIs) and whether results generalize across medical domains remain unclear. ObjectiveTo evaluate the screening performance of a zero-shot, large-batch, two-model LLM ensemble (OpenAI GPT-5 Thinking; Google Gemini 2.5 Pro) operated via public chat GUIs across a diverse range of medical topics, and to compare its performance with an equivalent API-based workflow. MethodsWe conducted a retrospective evaluation using 736 titles and abstracts from 16 Cochrane reviews (330 included, 406 excluded), all published in May-June 2025. The primary outcome was the sensitivity of a pre-specified "OR" ensemble rule designed to maximize sensitivity, benchmarked against final full-text inclusion decisions (reference standard). Secondary outcomes were specificity, single-model performance, and duplicate-run reliability (Cohens {kappa}). Because models saw only titles/abstracts while the reference standard reflected full-text decisions, specificity estimates are conservative for abstract-level screening. ResultsThe GUI-based ensemble achieved 99.7% sensitivity (95% CI, 98.3%-100.0%) and 49.3% specificity (95% CI, 44.3%-54.2%). The API-based workflow yielded comparable performance, with 99.1% sensitivity (95% CI, 97.4%-99.8%) and 49.3% specificity (95% CI, 44.3%-54.2%). The difference in sensitivity was not statistically significant (McNemar p=0.625) and met equivalence within a {+/-}2-percentage-point margin (TOST<0.05). Duplicate-run reliability was substantial to almost perfect (Cohens {kappa}: 0.78-0.93). The two models showed complementary strengths: Gemini 2.5 Pro consistently achieved higher sensitivity (94.5%-98.2% across single runs), whereas GPT-5 Thinking yielded higher specificity (62.3%-67.0%). ConclusionsA zero-code, browser-based workflow using a dual-LLM ensemble achieves near-perfect sensitivity for abstract screening across multiple medical domains, with performance equivalent to API-based methods. Ensemble approaches spanning two model families may mitigate model-specific blind spots. Prospective studies should quantify workload, cost, and operational feasibility in end-to-end systematic review pipelines.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 95%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 94%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 93%
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 96%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 95%
- Catchii: empowering literature review screening in healthcare 95%
Similar papers in this journal
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 95%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 93%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 93%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 94%
- Characteristics and completeness of reporting of systematic reviews of prevalence studies in adult populations: a meta-epidemiological study 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.