End-to-End Reliability of Automated Systems for Diagnostic Data Extraction: A Benchmark Study in Uro-Oncologic Evidence Synthesis
May, M.; Garzaro, J.; Kravchuk, A.; Gilfrich, C.; Donabauer, G.; Ateia, S.; Kruschwitz, U.; Haas, M.; Eckl, C.; Burger, M.
Show abstract
BackgroundAutomated systems, including large language models, are increasingly used to support data extraction in diagnostic systematic reviews. However, their reliability, safety, and repeatability under realistic extraction conditions remain insufficiently characterized. ObjectiveTo benchmark the end-to-end reliability of automated systems for extracting diagnostic accuracy data from published uro-oncologic studies, with a focus on correctness, abstention behavior in non-derivable scenarios, repeatability across repeated runs, and operational efficiency. MethodsThis prospective, protocol-driven benchmarking study evaluates a purpose-built extraction system (MedNuggetizer) and three contemporary large language models. Systems are applied to a fixed corpus of published full-text PDFs and publicly available supplementary material reporting on Uromonitor and urine cytology for bladder cancer detection. A locked, uniform extraction prompt is used across all systems. The primary endpoint is dataset-run correctness, defined as either exact extraction of the complete 2-by-2 diagnostic table or correct declaration of non-derivability. A non-inferiority design with an exact one-sided binomial test is employed. Secondary endpoints include hallucination behavior on pre-specified sentinel datasets, repeatability across repeated runs, fidelity of derived diagnostic metrics, and execution time compared with human extraction. ResultsThe study is powered for a non-inferiority margin of 5 percentage points relative to a predefined correctness threshold of 95 percent. Twenty independent runs per system are performed, yielding 320 dataset-run observations. Primary inference is conducted at the run level, with consensus-level results reported as supportive robustness analyses. ConclusionsThis protocol establishes a conservative and reproducible framework for evaluating automated systems used in diagnostic evidence synthesis. By integrating correctness, abstention, repeatability, and safety into a single end-to-end evaluation, the study addresses key methodological gaps in the clinical assessment of generative AI tools.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 93%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 92%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 92%
Similar papers in this journal
- Automating the cancer registry: An Autonomous, Resource-Efficient AI for Multi-Cancer Data Abstraction from Pathology Reports 91%
- Commercially-available heart rate monitor repurposed into a 9-gram standalone device for automatic arrhythmia detection with snapshot electrocardiographic capability: a pilot validation 87%
- Transparent and robust Artificial intelligence-driven Electrocardiogram model for Left Ventricular Systolic Dysfunction 87%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.