Back

Research Synthesis Methods

Wiley

All preprints, ranked by how well they match Research Synthesis Methods's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Amount and certainty of evidence in Cochrane systematic reviews of interventions: a large-scale meta-research study

Starck, T.; Ravaud, P.; Boutron, I.

2025-12-21 public and global health 10.64898/2025.12.19.25342674 medRxiv
Top 0.1%
55.9%
Show abstract

ObjectivesTo quantify the amount and certainty of evidence in Cochrane systematic reviews of interventions, and to describe how this evidence has evolved over time. DesignLarge-scale meta-research study Data sourceCochrane Database of Systematic Reviews (search date April 8, 2025) Eligibility criteriaCochrane systematic reviews assessing interventions reporting "Summary of findings" tables. Data extractionData were automatically extracted using web scraping and a large language model, with quality control performed by humans on a random sample. AnalysisWe describe the certainty of evidence for each population-intervention-comparison-outcome (PICO) question reported in all Cochrane "Summary of findings" tables. When available, we compared the certainty of evidence between the initial version and the latest update. ResultsWe identified 5,116 reviews that reported a "Summary of findings" table, containing 64,849 PICO questions. Overall, 24% (n = 15,768) of PICOs had no study included, 31% (n = 20,390) included only 1 study, 14% (n = 8,796) 2 studies, and 31% (n = 19,895) more than 2 studies. Nearly all PICOs (97%) only included randomized trials. The median [Q1-Q3] number of included participants was 123 [0-557]. The certainty of evidence was rated as high for 4% (n = 2,852), moderate for 16% (n = 10,574), low for 27% (n = 17,409), very low for 26% (n = 17,012), and not assessed for 26% (n = 17,002). Of the 7,461 PICO questions with an update (median time to update of 4.3 years [Q1-Q3: 2.6-6.4]), the number of included studies in the latest update remained the same for 63%; the certainty of evidence was unchanged for 71%; upgraded for 13% and downgraded for 15%. ConclusionThe amount and certainty of evidence is low and has not improved over time with review updates. These results question the efficiency of the research ecosystem. SummaryO_ST_ABSWhat is already known on this topicC_ST_ABSO_LIHigh quality up-to-date evidence synthesis is essential for decision-makers. C_LIO_LIConfidence in the evidence informing decision-making can be limited by the amount and quality of primary research on a specific research question C_LI What this study addsO_LIThis large-scale meta-research study analyzed all Cochrane "Summary of findings" tables (i.e., 64,849 population-intervention-comparison-outcome - PICO - questions), and found that about two thirds of the PICO questions were informed by two or fewer studies, with a median [Q1-Q3] of 123 [0-557] participants per PICO; the associated certainty of evidence was rated as high in only 4% of the cases. C_LIO_LIAfter an update of the review (i.e., 7,461 PICOs), 63% PICOs did not include additional studies, and 71% showed no change in certainty of evidence; upgrades and downgrades of certainty occurred at similar frequencies. C_LIO_LIThese results question the efficiency of the research ecosystem. C_LI

2
NMA: Network meta-analysis based on multivariate meta-analysis and meta-regression models in R

Noma, H.; Maruo, K.; Tanaka, S.; Furukawa, T. A.

2025-09-18 epidemiology 10.1101/2025.09.15.25335823 medRxiv
Top 0.1%
53.4%
Show abstract

Network meta-analysis has become an established methodology within systematic reviews for comparing the effectiveness of multiple treatments, and it has been now a standard approach in comparative effectiveness research. However, the underlying statistical methods are often highly technical for non-statisticians in practice, and no freely available software package has been developed that can handle a general framework based on the multivariate meta-analysis and meta-regression models. To address these issues, we developed NMA, a comprehensive and user-friendly R package that covers extensive analysis and graphical tools of network meta-analysis with simple commands. The NMA package provides generic functional tools for evidence synthesis based on the multivariate meta-analysis models, network meta-regression, assessment of heterogeneity and inconsistency, comparative effectiveness analyses, and a range of graphical tools. In addition, NMA includes data-handling functions that facilitate the integration of both arm-level data and summary effect measure statistics easily. In this article, we provide a gentle introduction to the NMA package and illustrate its application through a case study of a network meta-analysis of antihypertensive drugs. HighlightsWhat is already known? O_LISeveral freely computational packages are available for network meta-analysis, but no general frequentist tool based on the multivariate meta-analysis and meta-regression models, introduced by White et al. 13, has been developed. C_LI What is new? O_LIWe developed NMA, a comprehensive R package for network meta-analysis based on multivariate meta-analysis and meta-regression models with frequentist approach. C_LIO_LIThe NMA package provides a broad range of functions for evidence synthesis, heterogeneity and inconsistency assessment, comparative effectiveness analysis, and graphical visualization. C_LIO_LIKey analytical tools--such as Higgins global inconsistency test 12, network meta-regression, and advanced inferential and prediction methods to address invalidity issues of the ordinary approaches 17,21 --are fully implemented. C_LIO_LIGeneric data-handling tools can now integrate arm-level data with summary statistics. This provides greater flexibility in network meta-analysis, making it especially useful for studies of survival outcomes. C_LI Potential impact for RSM readers O_LIThe package facilitates the practical use of network meta-analysis for a broad range of researchers, including non-statisticians, thereby enhancing the accessibility of systematic reviews on important clinical and public health questions. C_LIO_LIBy comprehensively covering standard analyses and graphical tools, the NMA package is also valuable for educational purposes, serving as a practical resource for students, researchers, and clinicians to learn the research methods through real-world case studies. C_LI

3
Ranked (In)direct Citation Searching in Systematic Reviews: A methodological case study

Woelfle, T.; Fucile, G.; Hirt, J.; Pena, R. C. G.; Vogt, M.; Nordhausen, T.; Ewald, H.; Appenzeller-Herzog, C.

2026-05-27 medical education 10.64898/2026.05.26.26354093 medRxiv
Top 0.1%
53.0%
Show abstract

Systematic Review (SR) is a prosperous study type in modern medicine and beyond. Many SR authors complement their primary database searches by supplementary techniques. Among these, citation-based techniques known as citation searching (CS) are widespread. Unranked Direct CS (UDCS) to identify directly cited and citing literature of seed references is currently most prevalent. Ranked (In)direct CS (RICS) additionally collects co-cited and co-citing literature combined with a ranking and cut-off procedure. However, RICS workflows remain non-standardized and tedious, and associated benefits unclear. This work aims to create a framework for the prospective international comparison of supplementary UDCS and RICS. To prime RICS research, we developed the open-source Co*Citation Network application and assessed parallel supplementary UDCS and RICS retrospectively in three completed SRs and prospectively in one case study. Automated RICS collected and ranked cited, citing, co-cited, and co-citing literature of seed references from OpenAlex database and applied an empirical rank cut-off to approximate the volume of UDCS results. In RICS compared to UDCS, we consistently noted higher overlap with primary database search results. Title/abstract screening in the case study showed a precision (number needed to read) of 1.8% (57) for UDCS and 2.1% (48) for RICS results. After full text screening, two additional articles were included for review, one of which was identified by UDCS and RICS, and one exclusively by UDCS. The present study indicates potential benefits of RICS for SR authors and will enable the formation of a research consortium to compare supplementary UDCS and RICS on larger scale.

4
Sensitivity, specificity and avoidable workload of using a large language models for title and abstract screening in systematic reviews and meta-analyses

Tran, V.-T.; Gartlehner, G.; Yaacoub, S.; Boutron, I.; Schwingshackl, L.; Stadelmaier, J.; Sommer, I.; Aboulayeh, F.; Afach, S.; Meerpohl, J.; Ravaud, P.

2023-12-17 epidemiology 10.1101/2023.12.15.23300018 medRxiv
Top 0.1%
52.7%
Show abstract

ImportanceSystematic reviews are time-consuming and are still performed predominately manually by researchers despite the exponential growth of scientific literature. ObjectiveTo investigate the sensitivity, specificity and estimate the avoidable workload when using an AI-based large language model (LLM) (Generative Pre-trained Transformer [GPT] version 3.5-Turbo from OpenAI) to perform title and abstract screening in systematic reviews. Data SourcesUnannotated bibliographic databases from five systematic reviews conducted by researchers from Cochrane Austria, Germany and France, all published after January 2022 and hence not in the training data set from GPT 3.5-Turbo. DesignWe developed a set of prompts for GPT models aimed at mimicking the process of title and abstract screening by human researchers. We compared recommendations from LLM to rule out citations based on title and abstract with decisions from authors, with a systematic reappraisal of all discrepancies between LLM and their original decisions. We used bivariate models for meta-analyses of diagnostic accuracy to estimate pooled estimates of sensitivity and specificity. We performed a simulation to assess the avoidable workload from limiting human screening on title and abstract to citations which were not "ruled out" by the LLM in a random sample of 100 systematic reviews published between 01/07/2022 and 31/12/2022. We extrapolated estimates of avoidable workload for health-related systematic reviews assessing therapeutic interventions in humans published per year. ResultsPerformance of GPT models was tested across 22,666 citations. Pooled estimates of sensitivity and specificity were 97.1% (95%CI 89.6% to 99.2%) and 37.7%, (95%CI 18.4% to 61.9%), respectively. In 2022, we estimated the workload of title and abstract screening for systematic reviews to range from 211,013 to 422,025 person-hours. Limiting human screening to citations which were not "ruled out" by GPT models could reduce workload by 65% and save up from 106,268 to 276,053-person work hours (i.e.,66 to 172-person years of work), every year. Conclusions and RelevanceAI systems based on large language models provide highly sensitive and moderately specific recommendations to rule out citations during title and abstract screening in systematic reviews. Their use to "triage" citations before human assessment could reduce the workload of evidence synthesis.

5
boutliers: R package of outlier detection and influence diagnostics for meta-analysis

Noma, H.; Maruo, K.; Gosho, M.

2025-09-19 health informatics 10.1101/2025.09.18.25336125 medRxiv
Top 0.1%
52.0%
Show abstract

Meta-analysis is an established methodology for evidence synthesis. In practice, substantial heterogeneity often arises among studies, and random-effects models are widely employed as standard tools. However, in many cases of data synthesis, some studies exhibit markedly different characteristics from others, beyond the degree expected from statistical error, and may become influential outliers that affect the overall conclusions. Although outlier detection and influence diagnostic methods have been discussed in the context of meta-analysis, there has been a lack of user-friendly statistical packages that quantify the statistical uncertainty of diagnostic measures. We developed the R package boutliers, which implements influence diagnostics based on bootstrap methods using simple commands. The package provides three leave-one-out diagnostics: (1) studentized residuals, (2) relative change measures for the variance of the grand mean parameter, and (3) relative change measures for heterogeneity variance. In addition, a model-based approach using a likelihood ratio statistic under a mean-shifted outlier detection model is also available. This article offers a practical tutorial for the boutliers package, illustrated with an application to a meta-analysis of chronic low back pain.

6
Enough evidence and other endings: a descriptive study of stable Cochrane systematic reviews in 2019.

Bastian, H.; Hemkens, L. G.

2019-12-09 epidemiology 10.1101/19013912 medRxiv
Top 0.1%
52.0%
Show abstract

BackgroundFrom 2006 to 2019, Cochrane reviews could be designated "stable" if they were not being updated but highly likely to be current. This provides an opportunity to observe practice in ending systematic reviewing and what is regarded as enough evidence. MethodsWe identified Cochrane reviews designated stable in 2013 and 2019 and reasons for this designation. For those with conclusions stated to be so firm that new evidence is unlikely to change them, we assessed conclusions, strength of evidence ratings, and recommendations for further research. We assessed the fate of the 2013 stable reviews. We also estimated usage of formal analytic methods to determine when there is enough evidence in protocols for Cochrane reviews. ResultsCochrane reviews were rarely designated stable. In 2019, there were 507 stable Cochrane reviews (6.6% of 7,645 non-withdrawn reviews). The most common reasons related to no, little, or infrequent research activity expected (331 of 505; 65.5%). Only 39 reviews were stable because of firm conclusions unlikely to be changed by new evidence (7.7%), but that declaration was mostly not supported by judgments made in the review about strength of evidence and implications for research. Among the 180 reviews stable in 2013, 16 reverted to normal status (8.9%), with 2 of those changing conclusions because of new studies. Few Cochrane protocols specified an analytic method for determining when there was enough evidence to stop updating the review (116 of 2,415; 4.8%). ConclusionCochrane reviews were more likely to end because important future primary research activity was believed to be unlikely, than because there was enough evidence. Judgments about the strength of evidence and need for research were often inconsistent with the declaration that conclusions were unlikely to change. The inconsistencies underscore the need for reliable analytic methods to support decision-making about the conclusiveness of evidence.

7
The epidemiology of systematic review updates: a longitudinal study of updating of Cochrane reviews, 2003 to 2018.

Bastian, H.; Doust, J.; Clarke, M.; Glasziou, P.

2019-12-11 epidemiology 10.1101/19014134 medRxiv
Top 0.1%
51.7%
Show abstract

BackgroundThe Cochrane Collaboration has been publishing systematic reviews in the Cochrane Database of Systematic Reviews (CDSR) since 1995, with the intention that these be updated periodically. ObjectivesTo chart the long-term updating history of a cohort of Cochrane reviews and the impact on the number of included studies. MethodsThe status of a cohort of Cochrane reviews updated in 2003 was assessed at three time points: 2003, 2011, and 2018. We assessed their subject scope, compiled their publication history using PubMed and CDSR, and compared them to all Cochrane reviews available in 2002 and 2017/18. ResultsOf the 1,532 Cochrane reviews available in 2002, 11.3% were updated in 2003, with 16.6% not updated between 2003 and 2011. The reviews updated in 2003 were not markedly different to other reviews available in 2002, but more were retracted or declared stable by 2011 (13.3% versus 6.3%). The 2003 update led to a major change of the conclusions of 2.8% of updated reviews (n = 177). The cohort had a median time since publication of the first full version of the review of 18 years and a median of three updates by 2018 (range 1-11). The median time to update was three years (range 0-14 years). By the end of 2018, the median time since the last update was seven years (range 0-15). The median number of included studies rose from eight in the version of the review before the 2003 update, to 10 in that update and 14 in 2018 (range 0-347). ConclusionsMost Cochrane reviews get updated, however they are becoming more out-of-date over time. Updates have resulted in an overall rise in the number of included studies, although they only rarely lead to major changes in conclusion.

8
Evaluating Loon Lens Pro, an AI-Driven Tool for Full-Text Screening in Systematic Reviews: A Validation Study

Janoudi, G.; Uzun, M.; Jurdana, M.; Hutton, B.

2025-02-14 epidemiology 10.1101/2025.02.11.25322087 medRxiv
Top 0.1%
46.4%
Show abstract

BackgroundSystematic literature reviews (SLRs) are essential for evidence synthesis but are hampered by the resource-intensive full-text screening phase. Loon Lens Pro, a publicly available agentic AI tool, automates full-text screening without prior training by using user-defined inclusion/exclusion criteria and multiple specialized AI agents. This study validated Loon Lens Pro against human reviewers to assess its accuracy, efficiency, and confidence scoring in screening. MethodsIn this comparative validation study, 84 full-text articles from eight SLRs were screened by both Loon Lens Pro and human reviewers (gold standard). The AI provided binary inclusion/exclusion decisions along with a transparent rationale and confidence ratings (low, medium, high). Performance metrics-- including accuracy, sensitivity, specificity, negative predictive value, precision, and F1 score--were derived from a confusion matrix. Logistic regression with bootstrap resampling (1,000 iterations) evaluated the association between confidence scores and screening errors. ResultsLoon Lens Pro correctly classified 70 of 84 full texts, achieving an accuracy of 83.3% (95% CI: 75.0- 90.5%), sensitivity of 94.7% (95% CI: 82.4-100%), and specificity of 80.0% (95% CI: 70.1-89.2%). The negative predictive value was 98.1% (95% CI: 93.8-100%), with a precision of 58.1% (95% CI: 41.4- 76.0%) and an F1 score of 0.72. Logistic regression revealed a strong inverse relationship between confidence level and error probability: low, medium, and high confidence decisions were associated with predicted error probabilities of 46.9%, 30.9%, and 3.5%, respectively (C-index = 0.87). ConclusionOur study provides evidence that Loon Lens Pro is a viable and effective tool for automating the full-text screening phase of systematic reviews. Its high sensitivity, robust confidence scoring mechanism, and transparent rationale generation collectively support its potential to alleviate the burden of manual screening without compromising the quality of study selection.

9
Estimation of the COVID-19 Average Incubation Time:Systematic Review, Meta-analysis and SensitivityAnalyses

Weng, Y.; Yi, G. Y.

2022-01-18 epidemiology 10.1101/2022.01.17.22269421 medRxiv
Top 0.1%
45.7%
Show abstract

ObjectivesWe aim to provide sensible estimates of the average incubation time of COVID-19 by capitalizing available estimates reported in the literature and explore different ways to accommodate heterogeneity involved with the reported studies. MethodsWe search through online databases to collect the studies about estimates of the average incubation time and conduct meta-analyses to accommodate heterogeneity of the studies and the publication bias. Cochrans heterogeneity statistic Q and Higgins & Thompsons I2 statistic are employed. Subgroup analyses are conducted using mixed effects models and publication bias is assessed using the funnel plot and Eggers test. ResultsUsing all those reported mean incubation estimates, the average incubation time is estimated to be 6.43 days with a 95% confidence interval (CI) (5.90, 6.96), and using all those reported mean incubation estimates together with those transformed median incubation estimates, the estimated average incubation time is 6.07 days with a 95% CI (5.70,6.45). ConclusionsProviding sensible estimates of the average incubation time for COVID-19 is important yet complex, and the available results vary considerably due to many factors including heterogeneity and publication bias. We take different angles to estimate the mean incubation time, and our analyses provide estimates to range from 5.68 days to 8.30 days.

10
Applying Machine Learning to Increase Efficiency and Accuracy of Meta-Analytic Review

Gorelik, A. J.; Gorelik, M. G.; Ridout, K. K.; Nimarko, A. F.; Peisch, V.; Kuramkote, S. R.; Low, M.; Pan, T.; Singh, S.; Nrusimha, A.; Singh, M. K.

2020-10-08 neuroscience 10.1101/2020.10.06.314245 medRxiv
Top 0.1%
41.5%
Show abstract

The rapidly burgeoning quantity and complexity of publications makes curating and synthesizing information for meta-analyses ever more challenging. Meta-analyses require manual review of abstracts for study inclusion, which is time consuming, and variation among reviewer interpretation of inclusion/exclusion criteria for selecting a paper to be included in a review can impact a studys outcome. To address these challenges in efficiency and accuracy, we propose and evaluate a machine learning approach to capture the definition of inclusion/exclusion criteria using a machine learning model to automate the selection process. We trained machine learning models on a manually reviewed dataset from a meta-analysis of resilience factors influencing psychopathology development. Then, the trained models were applied to an oncology dataset and evaluated for efficiency and accuracy against trained human reviewers. The results suggest that machine learning models can be used to automate the paper selection process and reduce the abstract review time while maintaining accuracy comparable to trained human reviewers. We propose a novel approach which uses model confidence to propose a subset of abstracts for manual review, thereby increasing the accuracy of the automated review while reducing the total number of abstracts requiring manual review. Furthermore, we delineate how leveraging these models more broadly may facilitate the sharing and synthesis of research expertise across disciplines.

11
Clinical study type classification, validation, and PubMed filter comparison with natural language processing and active learning

van IJzendoorn, D. G. P.; Habets, P. C.; Vinkers, C. H.; Otte, W. M.

2022-11-03 epidemiology 10.1101/2022.11.01.22281685 medRxiv
Top 0.1%
40.9%
Show abstract

Each day, many thousands of new studies are published. Identifying specific study types with high sensitivity and specificity may improve searchability and accelerate updating systematic reviews and meta-analyses. Machine learning transformer models could facilitate this identification process if sufficient training data is available. We used an active learning strategy to construct a large training set (n=50,000) and fine-tuned the PubMedBERT language model to classify PubMed abstracts as randomized controlled trials, human studies, systematic reviews with and without meta-analyses, protocols, and rodent studies. In an external dataset (n=5,000), the average sensitivity and specificity across study types were 0.94 and 0.96, respectively. PubMeds internal filters had a low sensitivity for both systematic reviews with meta-analysis (0.175, CI: 0.057-0.293) and randomized controlled trials (0.256, CI: 0.119-0.393). We applied this labeling to all 34 million PubMed abstracts currently available and provide the results within an online meta-information platform (EvidenceHunt). In conclusion, we show that study type classification in PubMed is opportune, given the available language models. The high accuracy in this study invites extending these models to more elaborate and hierarchical identification schemes.

12
Determining fragility and robustness to missing data in binary outcome meta-analyses, illustrated with conflicting associations between vitamin D and cancer mortality

Grimes, D. R.

2025-08-19 epidemiology 10.1101/2025.08.15.25333793 medRxiv
Top 0.1%
40.8%
Show abstract

Meta-analysis is a vital component in clinical decision making, but previous work found binary event meta-analytic results can be fragile, affected by only a small number of patients in specific trials. Meta-analyses can also miss literature, and a method for estimating how much additional unseen data would flip results would be a useful tool. This works establishes a complementary and generalisable definition of meta-analytic fragility, based on Ellipse of Insignificance (EOI) and Region of Attainable Redaction (ROAR) methods originally developed for dichotomous outcome trials. This method does not require trial-specific alterations to estimate fragility and yields a general method to estimate robustness of a meta-analysis to data redaction or addition of hypothetical trial outcomes. This method is applied to 3 meta-analyses with conflicting findings on the association of vitamin D supplementation and cancer mortality. A full meta-analysis of all trials cited in the 3 meta-analyses yielded no association between vitamin D supplementation and cancer mortality. Using the method outlined here, it was determined that meta-analytic fragility was high in all cases, with recoding of just 5 patients in the full cohort of 133,262 patients was enough to cross the significance threshold. Small amounts of redacted or non-included data also had substantial impact on each meta-analysis, with addition of just 3 hypothetical patients to an ostensibly significant meta-analyses (N = 38,538) enough to yield a null result. This method for analytical fragility is complementary to previous investigations that suggested meta-analyses are frequently fragile. It further shows that merely increasing the sample size is not an assurance against fragility. Caution should be advised when interpreting the results of meta-analyses and conflicting results may stem from inherent fragility and should be carefully employed.

13
What we do in the shadows: Methodologists use a range of synthesis methods when meta-analysis of all results is not possible but describe challenges in planning and selecting methods

Cumpston, M. S.; Brennan, S. E.; Ryan, R.; Thomas, J.; McKenzie, J. E.

2026-07-18 epidemiology 10.64898/2026.07.15.26358140 medRxiv
Top 0.1%
40.5%
Show abstract

Introduction Systematic review authors commonly encounter situations where the data required for meta-analysis are incompletely reported (e.g. when effect estimates are reported without a measure of precision). In this circumstance, many systematic review authors use a method other than meta-analysis (e.g. vote counting), but rarely describe those methods or the rationale for selecting them. We aimed to investigate what methods authors consider when meta-analysis of all study results is not possible, and what factors influence their decisions. Methods We interviewed 12 experienced systematic review authors, editors and methodologists, presenting four scenarios in which it was not possible to combine all results using meta-analysis. Scenarios varied in the number and size of included studies, available data, and risk of bias. Participants discussed the methods they considered to summarise, synthesise and present the results; whether they would synthesise available results; and how they would draw overall conclusions. Results Factors that informed decisions included participants' overall purpose in conducting synthesis, existing beliefs about study results and synthesis methods, trust in the available data, and the decision-making needs of end users. Participants differed in which synthesis methods to use, whether they would use multiple synthesis methods, and which studies they would analyse with each method. Conclusions We identified several synthesis methods considered when meta-analysis of all results is not possible, and factors that influence the selection of methods, neither of which are routinely reported. More complete reporting of these methods and the factors informing decisions would allow readers to better understand the decisions made.

14
Accelerating the pace and accuracy of systematic reviews using AI: a validation study

Zhan, J.; Suvada, K.; Xu, M.; Tian, W.; Cara, K. C.; Wallace, T. C.; Ali, M. K.

2024-12-11 epidemiology 10.1101/2024.12.10.24318803 medRxiv
Top 0.1%
39.6%
Show abstract

BackgroundArtificial intelligence (AI) can greatly enhance efficiency in systematic literature reviews and meta-analyses, but its accuracy in screening titles/abstracts and full-text articles is uncertain. ObjectivesThis study evaluated the performance metrics (sensitivity, specificity) of a GPT-4 AI program, Review Copilot, against human decisions (gold standard) in screening titles/abstracts and full-text articles from four published systematic reviews/meta-analyses. Research DesignParticipant data from four already-published systematic literature reviews were used for this validation study. This was a study comparing Review Copilot to human decision-making (gold standard) in screening titles/abstracts and full-text articles for systematic reviews/meta-analyses. The four studies that were used in this study included observational studies and randomized control trials. Review Copilot operates on the OpenAI, GPT-4 server. We examined the performance metrics of Review Copilot to include and exclude titles/abstracts and full-text articles as compared to human decisions in four systematic reviews/meta-analyses. Sensitivity, specificity, and balanced accuracy of title/abstract and full-text screening were compared between Review Copilot and human decisions. ResultsReview Copilots sensitivity and specificity for title/abstract screening were 99.2% and 83.6%, respectively, and 97.6% and 47.4% for full-text screening. The average agreement between two runs was 95.4%, with a kappa statistic of 0.83. Review Copilot screened in one-quarter of the time compared to humans. ConclusionsAI use in systematic reviews and meta-analyses is inevitable. Health researchers must understand these technologies strengths and limitations to ethically leverage them for research efficiency and evidence-based decision-making in health.

15
Large language models for abstract screening in systematic- and scoping reviews: A diagnostic test accuracy study

Krag, C. H.; Balschmidt, T.; Bruun, F.; Brejnebol, M. W.; Xu, J. J.; Boesen, M.; Andersen, M. B.; Müller, F. C.

2024-10-02 radiology and imaging 10.1101/2024.10.01.24314702 medRxiv
Top 0.1%
39.5%
Show abstract

IntroductionWe investigated if large language models (LLMs) can be used for abstract screening in systematic- and scoping reviews. MethodsTwo broad reviews were designed: a systematic review structured according to the PRISMA guideline with abstract inclusion based on PICO criteria; and a scoping review, where we defined abstract characteristics and features of interest to look for. For both reviews 500 abstracts were sampled. Two readers independently screened abstracts with disagreements handled with arbitrations or consensus, which served as the reference standard. The abstracts were analysed by six LLMs (GPT-4o, GPT-4T, GPT-3.5, Claude3-Opus, Claude3-Sonnet, and Claude3-Haiku). Primary outcomes were diagnostic test accuracy measures for abstract inclusion, abstract characterisation and feature of interest detection. Secondary outcome was the degree of automation using LLMs as a function of the error rate. ResultsIn the systematic review 12 studies were marked as include by the human consensus. GPT-4o, GPT-4T, and Claude3-Opus achieved the highest accuracies (97% to 98%) comparable to the human readers (96% and 98%), although sensitivity was low (33% to 50%). In the scoping review 130 features of interest were present and the LLMs achieved sensitivities between 74-84%, comparable to the human readers (73% and 86%). The specificity of GPT-4o (98%) and GPT-4T (>99%) greatly surpassed the other LLMs (between 33% and 93%). For abstract characterization all LLMs achieved above 95% accuracy for language, manuscript type and study participant characterisation. For characterisation of disease-specific features only GPT-4T and GPT-4o showed very high accuracy. For abstract inclusion the highest automation rate (91%) at the lowest error rate (8%) was achieved by use of two LLMs with disagreement solved by human arbitration. An LLM pre screening before human abstract screening achieved an automation rate of 55% with no missed abstracts. ConclusionAbstract characterisation and specific feature of interest detection with LLMs is feasible and accurate with GPT-4o and GPT-4T. The majority of abstract screenings for systematic reviews can be automated with use of LLMs, at low error rates.

16
ChatGPT for assessing risk of bias of randomized trials using the RoB 2.0 tool: A methods study

Pitre, T.; Jassal, T.; Talukdar, J. R.; Shahab, M.; Ling, M.; Zeraatkar, D.

2023-11-22 epidemiology 10.1101/2023.11.19.23298727 medRxiv
Top 0.1%
39.4%
Show abstract

BackgroundInternationally accepted standards for systematic reviews necessitate assessment of the risk of bias of primary studies. Assessing risk of bias, however, can be time- and resource-intensive. AI-based solutions may increase efficiency and reduce burden. ObjectiveTo evaluate the reliability of ChatGPT for performing risk of bias assessments of randomized trials using the revised risk of bias tool for randomized trials (RoB 2.0). MethodsWe sampled recently published Cochrane systematic reviews of medical interventions (up to October 2023) that included randomized controlled trials and assessed risk of bias using the Cochrane-endorsed revised risk of bias tool for randomized trials (RoB 2.0). From each eligible review, we collected data on the risk of bias assessments for the first three reported outcomes. Using ChatGPT-4, we assessed the risk of bias for the same outcomes using three different prompts: a minimal prompt including limited instructions, a maximal prompt with extensive instructions, and an optimized prompt that was designed to yield the best risk of bias judgements. The agreement between ChatGPTs assessments and those of Cochrane systematic reviewers was quantified using weighted kappa statistics. ResultsWe included 34 systematic reviews with 157 unique trials. We found the agreement between ChatGPT and systematic review authors for assessment of overall risk of bias to be 0.16 (95% CI: 0.01 to 0.3) for the maximal ChatGPT prompt, 0.17 (95% CI: 0.02 to 0.32) for the optimized prompt, and 0.11 (95% CI: -0.04 to 0.27) for the minimal prompt. For the optimized prompt, agreement ranged between 0.11 (95% CI: -0.11 to 0.33) to 0.29 (95% CI: 0.14 to 0.44) across risk of bias domains, with the lowest agreement for the deviations from the intended intervention domain and the highest agreement for the missing outcome data domain. ConclusionOur results suggest that ChatGPT and systematic reviewers only have "slight" to "fair" agreement in risk of bias judgements for randomized trials. ChatGPT is currently unable to reliably assess risk of bias of randomized trials. We advise against using ChatGPT to perform risk of bias assessments. There may be opportunities to use ChatGPT to streamline other aspects of systematic reviews, such as screening of search records or collection of data.

17
Stage-wise algorithmic bias, its reporting, and relation to classical systematic review biases in AI-based automated screening in health sciences: A structured literature review

Pardal-Refoyo, J. L.; Pardal-Pelaez, B.

2025-05-16 health informatics 10.1101/2025.05.16.25327774 medRxiv
Top 0.1%
37.8%
Show abstract

IntroductionAlgorithmic bias in systematic reviews that use automatic screening is a major challenge in the application of AI in health sciences. This article presents preliminary findings from the project titled "Identification, Reporting, and Mitigation of Algorithmic Bias in Systematic Reviews with AI-Assisted Screening: Systematic Review and Development of a Checklist for its Evaluation" registered in PROSPERO with the registration number CRD420251036600 (https://www.crd.york.ac.uk/PROSPERO/view/CRD420251036600). The results presented here are preliminary and part of ongoing work. ObjectiveTo synthesize knowledge about the taxonomies of algorithmic bias, reporting, relationships with classical biases, and use of visualizations in AI-supported systematic reviews in health sciences. MethodsA specific literature review was conducted, focusing on systematic reviews, conceptual frameworks, and reporting standards for bias in AI in healthcare, as well as studies cataloguing detection and mitigation strategies, with an emphasis on taxonomies, transparency practices, and visual/illustrative tools. ResultsA mature body of work describes stage-based taxonomies and mitigation methods for algorithmic bias in general clinical AI. Common improvements in reporting and transparency (e.g. CONSORT-AI, SPIRIT-AI) are described. However, there is a notable absence of direct application to AI-automated screening of systematic reviews or empirical analyses of the interactions of biases with classical biases at the review level. Visualization techniques, such as bias heatmaps and pipe diagrams, are available, but have not been adapted to review workflows. ConclusionsThere are fundamental methodologies to identify and mitigate algorithmic bias in AI in health, but significant gaps remain in the understanding and operationalization of these frameworks within AI-assisted systematic reviews. Future research should address this translational gap to ensure transparency, fairness, and methodological rigor in the synthesis of evidence.

18
Agentic AI for Streamlining Title and Abstract Screening: Addressing Precision and evaluating calibration of AI guardrails

Disher, T.; Janoudi, G.; Uzun, M.

2024-11-15 epidemiology 10.1101/2024.11.15.24317267 medRxiv
Top 0.1%
35.2%
Show abstract

1.BackgroundTitle and abstract (TiAb) screening in systematic literature reviews (SLRs) is labor-intensive. While agentic artificial intelligence (AI) platforms like Loon Lens 1.0 offer automation, lower precision can necessitate increased full-text review. This study evaluated the calibration of Loon Lens 1.0s confidence ratings to prioritize citations for human review. MethodsWe conducted a post-hoc analysis of citations included in a previous validation of Loon Lens 1.0. The data set consists of records screened by both Loon Lens 1.0 and human reviewers (gold standard). A logistic regression model predicted the probability of discrepancy between Loon Lens and human decisions, using Loon Lens confidence ratings (Low, Medium, High, Very High) as predictors. Model performance was assessed using bootstrapping with 1000 resamples, calculating optimism-corrected calibration, discrimination (C-index), and diagnostic metrics. ResultsLow and Medium confidence citations comprised 5.1% of the sample but accounted for 60.6% of errors. The logistic regression model demonstrated excellent discrimination (C-index = 0.86) and calibration, accurately reflecting observed error rates. "Low" confidence citations had a predicted probability of error of 0.65 (95% CI: 0.56-0.74), decreasing substantially with higher confidence: 0.38 (95% CI 0.28-0.49) for "Medium", 0.05 (95% CI 0.04-0.07) for "High", and 0.01 (95% CI 0.007-0.01) for "Very High". Human review of "Low" and "Medium" confidence abstracts would lead to improved overall precision from 62.97% to 81.4% while maintaining high sensitivity (99.3%) and specificity (98.1%). ConclusionsLoon Lens 1.0s confidence ratings show good calibration used as the basis for a model predicting the probability of making an error. Targeted human review significantly improves precision while preserving recall and specificity. This calibrated model offers a practical strategy for optimizing human-AI collaboration in TiAb screening, addressing the challenge of lower precision in automated approaches. Further research is needed to assess generalizability across diverse review contexts.

19
Research waste from poor reporting of core methods and results and redundancy in studies of reporting guideline adherence: a meta-research review

Dal Santo, T.; Rice, D. B.; Amiri, L. S.; Tasleem, A.; Li, K.; Boruff, J. T.; Geoffroy, M.-C. B.; Benedetti, A.; Thombs, B.

2022-12-20 epidemiology 10.1101/2022.12.19.22283669 medRxiv
Top 0.1%
35.2%
Show abstract

ObjectivesWe investigated meta-research studies that evaluated adherence to prominent reporting guidelines (CONSORT, PRISMA, STARD, STROBE) in health research studies to determine the proportion that (1) provided an explanation for how complex guideline items were rated for adherence and (2) provided results from individual studies reviewed in addition to aggregate results. We also examined the conclusions of each meta-research study to assess redundancy of findings across studies. DesignCross-sectional meta-research review. Data sourcesMEDLINE (Ovid) searched on July 5, 2022. Eligibility criteria for selecting studiesStudies in any language were eligible if they used any version of the CONSORT, PRISMA, STARD, or STROBE reporting guidelines or their extensions to evaluate reporting in at least 10 human health research studies. We excluded studies that modified a reporting guideline or its items or evaluated fewer than half of reporting guideline items. Main outcomes were (1) the proportion of meta-research studies that provided a coding explanation that could be used to replicate the study or verify its results and (2) the proportion that provided individual-level study results in the main text, supplemental materials, or via an internet link. ResultsOf 148 included meta-research studies, 14 (10%, 95% confidence interval [CI] 6% to 15%) provided a fully replicable coding explanation, and 49 (33%, 95% CI 26% to 41%) completely reported individual study results. Of 90 studies that classified reporting as adequate or inadequate in the study abstract, 6 (7%, 95% CI 3% to 14%) concluded that reporting was adequate but none of those 6 studies provided information on how items were coded or provided item-level results for included studies. ConclusionsMuch of published meta-research on reporting in health research is likely wasteful. Few studies report enough information for verification or replication, and almost all find that reporting in health research studies is suboptimal. These findings highlight the importance of shifting the focus from assessing reporting adequacy to developing, testing, and implementing strategies to improve reporting. FundingThere was no specific funding for this study. ProtocolPosted on the Open Science Framework June 29, 2022 (https://osf.io/gtm4z/).

20
Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies

Henry, M.; O'Connell, N.; Riley, R.; Moons, K.; Shea, B.; Hooft, L.; Wallwork, S.; Damen, J.; Skoetz, N.; Appiah, R.; Berryman, C.; Crouch, S.; Ferencz, G.; Grant, A.; Henry, K.; Herman, A.; Karran, E.; Koralegedera, I.; Leake, H.; MacIntyre, E.; Mouatt, B.; Phuentsho, K.; Van Der Laan, D.; Welsby, E.; Wiles, L.; Wilkinson, E.; Wilson, M.; Wilson, M.; Moseley, L.

2025-04-14 epidemiology 10.1101/2025.04.10.25325555 medRxiv
Top 0.1%
35.1%
Show abstract

BackgroundThis paper details initial testing of the agreeability and usability of a novel quality appraisal tool for systematic reviews of prognostic factor studies: AMSTAR-PF. MethodsFourteen appraisers each assessed eight systematic reviews using AMSTAR-PF. Their ratings for each question and each article were compared, with interrater, inter-pair and intrapair agreeability calculated using Gwets agreement coefficient. Time of use and time to reach consensus were also recorded. ResultsInterrater agreement averaged 0.59 (range, 0.21-0.90), inter-pair 0.61 (range 0.24-0.91) and intrapair 0.75 (range 0.45-0.95) across the domains, with agreement for the overall rating 0.46 (95%CI 0.30-0.62) for interrater, 0.46 (95%CI 0.17-0.74) for inter-pair, and 0.68 (range of averages 0.22-1.00) for intrapair agreement. The majority (60.7%) of intrapair ratings were identical, with 94.6% of final ratings either identical or only one category different for the overall appraisal. The time taken to appraise a study with AMSTAR-PF improved with use and averaged around 34 minutes after the first two appraisals. ConclusionsDespite some variance in agreeability for different domains and between different appraisers, the testing results suggest that AMSTAR-PF has clear utility for appraising the quality of systematic reviews of prognostic factor studies.