A blinded, counterbalanced rater design for evaluating AI-assisted summarisation of tertiary clinical genomics reports: methodology of the QNOMX-VHIR-CPSP-001 Phase 1 study
Creeden, J.; Olivecrona, M.; Soriano, A.
Show abstract
Background. Tertiary clinical genomics reports condense layered molecular findings into documents that treating oncologists must read, translate, and act upon; manual summarisation of these reports is time-consuming and variable. Tools that assist summarisation and translation into local languages are emerging, yet the field lacks an agreed methodology for evaluating such tools before any downstream clinical use. The appropriate first endpoint is fidelity of the generated summary to its source report, assessed by qualified human raters under blinded scoring, not downstream variant classification. Methods. QNOMX-VHIR-CPSP-001 Phase 1 is a single-site, non-interventional clinical performance study conducted at Vall d'Hebron Institut de Recerca (VHIR) under ISO 20916:2019 as a Clinical Performance Study Protocol. De-identified tertiary cancer genomics reports from pediatric oncology cases are summarised by the AI-assisted summarisation system under evaluation and, in parallel, by the standard manual workflow. Qualified raters score both summary types against the source genomics report using the Quality Summary Index (QSI), a six-dimension, five-point rubric adapted from the Provider Documentation Summarization Quality Instrument, under a blinded, counterbalanced, two-period crossover with a minimum fourteen-day washout. Two co-primary composite endpoints, content and presentation, are analysed for non-inferiority under a Bayesian hierarchical model, with a frequentist linear mixed model as the convergence check. Inter-rater reliability is reported as Krippendorff's ; a Monte-Carlo power analysis of the fixed clustered design is pre-specified. Discussion. The design isolates summarisation quality from clinical decision-making by scoring both summary types against the same source report under blinding, counterbalancing, and a fourteen-day washout. Conclusion. The QSI rubric, the counterbalanced crossover, and the pre-specified Bayesian primary with frequentist convergence check define a replicable protocol for early-stage evaluation of AI-assisted summarisation in tertiary genomics reporting; observed variance components will inform sample-size determination for Phase 2.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Supporting Reanalysis and Reuse of Clinical Trial Data: A Case Study 94%
- The Melanoma Genomics Managing Your Risk Study randomised controlled trial: Statistical Analysis Plan 92%
- Controlled evaLuation of Angiotensin Receptor Blockers for COVID-19 respIraTorY disease (CLARITY): Statistical analysis plan for a randomised controlled Bayesian adaptive sample size trial 92%
Similar papers in this journal
- Analysis of clinical trial registry entry histories using the novel R package cthist 93%
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 93%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 93%
Similar papers in this journal
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 94%
- Reproducibility and transparency characteristics of oncology research evidence 93%
- Surgical Resection, Radiotherapy, And Percutaneous Thermal Ablation for Treatment of Stage 1 Non-Small Cell Lung Cancer: A Systematic Review and Network Meta-Analysis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.