Performance of Google NotebookLM for AI-assisted data extraction and consensus statement generation in a heterogenous systematic review on inflammatory bowel disease, obesity, and cardiometabolic comorbidities: A Methodological Report
Samaan, S.; Devi, J.; Vincent, M.; Coombs, S.; Sehgal, P.; Mouhamed, M.; Rai, V.; Johnson, A. M.; Yarur, A. J.; Barnes, E. L.; Deepak, P.
Show abstract
Background: Large language models (LLMs) offer promise for systematic review data extraction, but performance in complex multidisciplinary domains and utility for clinical statement generation remain insufficiently described. Objectives: To evaluate Google NotebookLM for AI-assisted data extraction and RAND/UCLA consensus statement generation in a systematic review of IBD, obesity, and cardiometabolic comorbidities. Methods: Studies were organized into domain-specific notebooks; structured prompts generated standardized evidence tables. Two independent reviewers validated outputs against full-text articles using a four-category error classification. Cell-level accuracy and critical accuracy (cells free of major factual errors) were the primary metrics; workflow time was compared against a published conventional extraction benchmark. Concordance between AI-generated and expert-finalized statements was assessed. Results: Across 57 articles, 1,710 data cells were extracted; 151 (8.83%) were flagged, yielding 91.17% cell-level accuracy. Major factual errors occurred in only 4 cells (0.23%), for a critical accuracy of 99.77%. Most errors were minor omissions (59.6%) or incomplete extractions (30.5%); domain error rates ranged from 7.08% to 11.33%. The pipeline required 17.7 versus a projected 165.1 person-hours (89.3% reduction). PICO-structured prompting generated 70 candidate statements; 58 of 112 finalized panel statements (51.8%) were AI-derived, and 85.7% were retained in the finalized set. Conclusion: Google NotebookLM demonstrates feasibility as a primary extraction and synthesis tool in a multidisciplinary systematic review, with extractive incompleteness as the principal limitation and substantial time savings over conventional approaches. Its novel application to RAND/UCLA consensus statement generation extends AI-assisted evidence synthesis to clinical consensus generation workflow.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Erroneous risks in pharmaceutical versus non-pharmaceutical interventions during data extraction in evidence synthesis practice: Study protocol for a randomized controlled trial 91%
- A Retrospective Analysis of Serious Adverse Events and Deaths in US-Based Lifestyle Clinical Trials for Cognitive Health 91%
- The IGNITE Trial: Participant Recruitment Lessons Prior to SARS-CoV-2 89%
Similar papers in this journal
- Methods used to select results to include in meta-analyses of nutrition research: a meta-research study 95%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 93%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 93%
Similar papers in this journal
Similar papers in this journal
- Transparency and reporting characteristics of COVID-19 randomized controlled trials 93%
- Tool to assess risk of bias due to missing evidence in network meta-analysis (ROB-MEN): elaboration and examples 92%
- Evidence of Unreliable Data and Poor Data Provenance in Clinical Prediction Model Research and Clinical Practice 91%
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 94%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 94%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.