BMJ
● BMJ
Preprints posted in the last 90 days, ranked by how well they match BMJ's content profile, based on 51 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Mismetti, P.; Bertoletti, L.; Elias, A.; Assante, C.; Sanchez, O.; Schmidt, J.; PRESLES, E.; Chapelle, C.; Accassat, S.; Couturaud, F.; Mahe, I.; Laporte, S.
Show abstract
BACKGROUND. In patients with venous thromboembolism (VTE), renal impairment increases the risks for both recurrence and bleeding. Because these patients are underrepresented in clinical trials, we assessed whether standard direct factor Xa inhibitor (DXI) lead-in followed by early dose reduction was noninferior to standard anticoagulation in patients with acute VTE and moderate-to-severe renal impairment. METHODS. The VERDICT trial was a randomized, prospective, multicenter, open-label, blinded-endpoint, noninferiority trial. Consecutive patients with acute proximal deep-vein thrombosis or pulmonary embolism and chronic renal impairment (creatinine clearance 15?50 mL/min) were randomized 1:1 to an early DXI dose reduction strategy or standard therapy (heparin plus a vitamin K antagonist). Patients allocated to the DXI strategy underwent a second 1:1 randomization to apixaban or rivaroxaban, each administered at an initial standard lead-in dose followed by early dose reduction. The primary outcome was net clinical benefit at 3 months, defined as the composite of major bleeding and symptomatic recurrent VTE. RESULTS. Due to slow recruitment, the trial was prematurely terminated after enrolling 200 of the planned 800 patients (DXI: n=104; standard: n=96). The median age was 85.9 years, 31% were male, and 29.0% had severe renal impairment. The primary outcome occurred in 8 patients (7.7%) in the DXI group and 9 patients (9.3%) in the standard therapy group (adjusted subhazard ratio [sHR], 0.87; 95% CI, 0.26 to 2.87; P = 0.19 for noninferiority; noninferiority margin 1.30). Major bleeding occurred in 6.7% and 6.2% of patients and recurrent VTE occurred in 1.0% and 3.1% of patients, respectively. CONCLUSION. In patients with VTE and moderate-to-severe renal impairment, noninferiority of an early DXI dose-reduction strategy versus standard therapy could not be demonstrated for net clinical benefit. Although no major differences in efficacy or safety outcomes were observed between groups, the reduced sample size precludes definitive conclusions.
Kumar, R. S. P.; Ye, J.
Show abstract
Background: Major soccer tournaments may temporarily change recreational soccer activity, community gatherings, and injury-prevention needs, but evidence for population-level emergency department (ED) injury patterns during these events is limited. Understanding whether ED-treated soccer injury burden changes during Men's FIFA World Cup periods may help inform surveillance readiness and prevention planning for future tournaments. Objective: To evaluate whether Men's FIFA World Cup tournament periods temporally coincided with changes in ED-treated soccer-coded injury burden in the United States and to assess the implications for public health surveillance and injury-prevention preparedness. Methods: We conducted a retrospective, repeated cross-sectional calendar-period analysis of publicly available national ED injury surveillance records from 1999 through 2025. Soccer-coded injuries were identified using product code 1267 in any available product field. The primary exposure was the set of official Men's FIFA World Cup tournament dates from 2002, 2006, 2010, 2014, 2018, and 2022. Tournament dates were compared with matched same-calendar dates in adjacent years, excluding dates that overlapped other FIFA World Cup tournament windows. The primary estimands were the mean daily difference and ratio in weighted national ED-treated soccer-coded injury estimates between tournament and matched-control periods. Results: The analytic cohort included 170,679 soccer-coded ED cases, corresponding to an estimated 5,366,681 ED-treated soccer-coded injuries nationally. Mean daily weighted estimates were 453.1 during Men's World Cup tournament dates and 384.1 during matched control dates. The absolute mean daily difference was 68.9 injuries per day (95% CI, -0.5 to 138.3), and the mean daily ratio was 1.18 (95% CI, 1.00 to 1.39). Tournament-specific estimates were heterogeneous, with a near-null estimate for the 2022 winter tournament and higher estimates for prior summer tournaments. Conclusions: Men's FIFA World Cup periods were associated with a modest, imprecise increase in mean daily ED-treated soccer-coded injury estimates, but the findings were heterogeneous and compatible with no difference to a moderate increase. These results should be interpreted as ecological and hypothesis-generating rather than causal. The primary implication is not that World Cup tournaments directly cause injuries, but that major soccer events provide a practical opportunity for real-time ED injury surveillance, targeted recreational soccer injury-prevention messaging, concussion awareness, and coordinated preparedness for community and fan-event injury patterns during future tournaments.
Moe-Byrne, T.; Knapp, P.; Golder, S.
Show abstract
Background People with lower levels of literacy or health literacy may struggle to understand conventional health information. Video animations show promise as information tools, yet it is unclear whether video animations help reduce these inequalities in understanding. This study examined whether the effectiveness of video animations in health settings differs according to level of literacy or health literacy. Methods We drew on trials from a recent systematic review of video animations about healthcare or public health topics for patients or the public. We extracted available data on literacy, health literacy, or proxy indicators. One reviewer extracted data and a second checked all entries. Where possible, we conducted subgroup analyses of low and high literacy levels or interaction meta-analyses comparing low versus high literacy groups; otherwise, results were summarised narratively. Results From 88 eligible trials, we extracted health literacy data for 12. Across nine trials reporting knowledge, animations mostly improved knowledge compared with controls in both lower and higher health literacy groups. Effects on attitudes and behaviours were mixed and often small, with few studies reporting results by health literacy level. Across the subgroup analyses available, there was no consistent evidence of a pooled interaction effect of animations according to low and high literacy groups, but both statistical heterogeneity and small subgroup sizes limited precision of estimates. Across 88 trials, 54 (61%) reported education level, 22 (25%) did not, and 12 (14%) involved children or adolescents likely to have similar education levels. Conclusions Overall, the available data suggest that video animations can improve knowledge outcomes in both lower and higher health literacy groups, but their impact on attitudes and behaviour is less clear. Because literacy was rarely reported or analysed in the trials, it remains uncertain whether animations help to reduce literacy-related inequalities in access to, and use of health information.
Li, Z.; Yagis, E.; Riad, A.; Windrath-Carr, O.; Arribas, M.; Sodiq, T.; Goldsmith, K.; Glampson, B.; Flott, K.; Haji, G.; Khan, Z.; Baker, C.; Mayer, E. K.
Show abstract
Venous thromboembolism (VTE) is a leading cause of preventable inpatient mortality, while the real-world performance of mandated risk assessment and the potential for automating using electronic health record (EHR) data remain unclear. We analysed 577,904 admissions and 726,896 VTE assessment forms across five NHS hospitals between 2015 and 2025 to evaluate assessment completion, concordance with structured EHR data, clinical validity, and feasibility of EHR-based automation assisted by machine learning. Overall completion was high (96.7%), and timely completion improved from 47.4% in 2015 to 90.5% in 2024. Agreement between forms and EHR data was good for common risk factors, but low-prevalence variables were often under-documented in the forms. Despite these discrepancies, form-derived thrombosis risk was associated with increased VTE incidence (OR 3.31, 95% CI 2.81-3.90). Machine learning models using first-14-hour EHR data achieved discrimination comparable to clinician-recorded variables (AUROC 0.709 vs 0.704), supporting real-time EHR-integrated assessment pre-population and decision support.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Agbalalah, T.; Rowaiye, A.
Show abstract
Abstract Background:Sickle cell disease (SCD) is concentrated in sub-Saharan Africa, where delivery of guideline-referenced care remains challenging. Current evaluation approaches rely largely on access indicators and clinical outcomes, which do not directly measure care delivery. We developed the Care Delivery Gap (CDG) framework, a patient-reported approach for identifying care-process omissions, and conducted a proof-of-concept study to assess feasibility and explore variation across income strata. Methods: We conducted a cross-sectional framework-development study involving a proof-of-concept sample of 52 individuals with SCD or caregivers recruited through clinics and moderated SCD communities across Africa, North America, and Europe between June 2025 and March 2026. The CDG framework assessed patient-reported omissions in specialist involvement, follow-up continuity, cardiovascular screening, and biochemical surveillance. Analyses were descriptive. Results: Substantial multi-domain care-process omissions were identified despite high reported healthcare engagement. Across geographic income strata, cardiovascular screening was reported by 4/35 (11%) LMIC versus 16/17 (94%) HIC participants, and regular follow-up within the preceding 12 months by 14/35 (40%) versus 16/17 (94%), respectively. High CDG scores, representing 1 omissions across three or four domains, occurred in 20/35 (57%) LMIC compared with 1/17 (6%) HIC participants. Similar disparities were observed across specialist review and vitamin B12 surveillance domains. Conclusion: A structured patient-reported framework identified multi-domain omissions in guideline-referenced SCD care, including among individuals reporting healthcare access. The divergence between access indicators and reported care delivery suggests that service contact alone may not reflect care quality. The framework provides a feasible foundation for future process-level quality measurement in high-burden settings.
Mitchell, B.; White, N. M.; Cheng, A.; Russo, P.; Brain, D.; Tehan, P.; Matterson, G.; King, J.; Havers, S.; Browne, K.
Show abstract
The CATION study is a parallel two-arm randomised controlled trial on the prevention of catheter-associated urinary tract infections (CAUTI) in hospitalised patients. The intervention is the use of a sterile wipe containing 0.1% chlorhexidine solution for meatal cleaning prior to urinary catheter insertion as part of usual care; the intervention will be compared with the use of a sterile wipe containing 0.9% normal saline as the control. This document is the Statistical Analysis Plan for evaluating primary, secondary and tertiary effectiveness outcomes. The trial was preregistered on the Australian and New Zealand Clinical Trials registry (ACTRN12625000278437). A copy of the study protocol and a signed version of this Statistical Analysis Plan are available on request from the corresponding author (BM).
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Bauer, N.; Binnie, A.; Lad, V.; Marticorena, M.; Tsang, J.; Poirier Zytaruk, N.; Heels-Ansdell, D.; Cook, D. J.
Show abstract
Background: In Canada, there is a lack of data relating sociodemographic characteristics to the likelihood of consent and clinical trial participation. Objective: The overall objective of this study is to examine the association of hospital-level sociodemographic variables with a priori informed consent rates for participation in the REVISE trial. Design: This study is a retrospective observational analysis of Canadian sites participating in the international REVISE trial. Methods: Sociodemographic characteristics for 42 hospitals participating in the REVISE trial will be supplemented by national data from the 2021 Canadian Census of Population Profile at the census tract level corresponding to the hospital's location. Hospital level information for Ontario sites will be derived from the Institute for Clinical Evaluate Sciences (ICES) database. Site clustering will be performed using latent class analysis, a flexible clustering technique that identifies meaningful subgroups based on sociodemographic variables purposively selected from data available through the Statistics Canada 2021 census profile, ICES, and hospital-reported data. Clustering analysis will be performed for all Ontario hospitals with available ICES data, followed by a separate analysis for all Canadian REVISE sites using Statistics Canada data. Concordance in the clustering of REVISE sites will be examined by comparing the assignment of hospitals to the latent classes separately identified using ICES and Statistics Canada data. If there is a high degree of agreement between the two datasets, sociodemographic predictors will be analyzed using the clusters identified through ICES for Ontario sites with the concordant classes based on Statistics Canada data for Canadian sites outsite Ontario. If there is disagreement in cluster assignment between the two datasets, separate analyses of sociodemographic factors will be conducted for Ontario sites using ICES data and for all Canadian sites using the 2021 Census Profile. Multivariate linear regression models will be used to analyze the association between hospital-level characteristics and the likelihood of a priori and deferred consent. Results: Results of this study will generate information about the relationship between informed consent to participate in a low-risk critical care clinical trial using different consent models, and socioeconomic patient characteristics at the hospital site level (e.g., educational attainment, knowledge of official languages, citizenship rates, family income, poverty, rurality and immigration patterns). Conclusions: This study will fill an evidence gap by generating information on the relationship between sociodemographic variables and the likelihood of informed consent to participate in a critical care clinical trial in Canada.
Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.
Show abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Fabian-Therond, C.; Ahuja, S.; Papachristou Nadal, I.; Holt, R. I.; Watson, S. I.; Hussain, S.; Choudhary, P.; Ajjan, R.; Harris, R.; Peck, M.; Mohammadi, J.; Sims, S.; Fiorentino, F.; Due-Christensen, M.; Huber, J.; Fisher, L.; Hardenberg, K.; Stadler, M.; Jin, H.; Halliday, J. A.; Sturt, J.; on behalf of the D-stress study collaborators,
Show abstract
Introduction Diabetes distress describes the psychological and emotional burden of living with diabetes and is associated with reduced self-management and adverse diabetes outcomes. Clinical guidelines recommend routine assessment and management of diabetes distress, but this is not always implemented. Therefore, there is a need to develop approaches to deliver emotional health support in routine clinical care more effectively. We describe here the protocol for a study to I) assess the feasibility of implementation of the D-stress Pathway, comprising Enhanced Usual Care (EUC) and an online, group-based, psychological diabetes distress reduction intervention called REDUCE, ii) evaluate the feasibility of the study protocol iii) detect an effect signal of diabetes distress score and Interstitial Glucose Time in Range and iv) refine initial programme theories of how both interventions (EUC and REDUCE) work, for whom, and under what circumstances. Methods This feasibility study includes a multicentre trial within a cohort design (TWICs) where sites have a staggered exposure to the interventions alongside a realist process evaluation. Four UK NHS diabetes services will recruit 80 adults with type 1 diabetes ([≥]1 year) using continuous glucose monitoring (CGM) ([≥]3 months). All participants will receive EUC and provide monthly data over 7 months on diabetes distress (measured by the Type 1 Diabetes Distress Assessment System (T1DDAS) and interstitial glucose measured by using continuous glucose monitoring. Participants with elevated diabetes distress, will be offered the six-week, group-based, online REDUCE intervention plus EUC, compared to EUC alone. Up to twenty participants with type 1 diabetes, ten family members/friends, sixteen healthcare professionals delivering EUC and five REDUCE facilitators will be interviewed to explore their experience of receiving training and delivering the D-stress Pathway. Up to 20 EUC consultations and REDUCE sessions will be observed. Analysis Feasibility will be assessed against pre-specified progression criteria and analysed descriptively using summary statistics. Primary outcomes include baseline level of diabetes distress, recruitment rate, intervention uptake, and data completeness, which will be analysed descriptively. Qualitative data will be analysed using framework analysis guided by realist programme theories developed for this study. Ethics Ethics approval has been granted by NHS Research Ethics Committee (REC) (Bromley REC: 25/LO/0469) and Health Research Authority obtained. All participants will provide informed consent. Trial registration no: Registered at ClinicalTrials.gov number NCT07193446 on 26/11/2025. Protocol and statistical analysis plan: The trial protocol and statistical analysis plan can be accessed at ClinicalTrials.gov.
Cornett, C.; Tilston, G.; Martin, G.; Palin, V.
Show abstract
Background: Maternal postpartum checks with a general practitioner (GP) are recognised as an essential service in England and vital for recovery after pregnancy and reducing risk of long-term morbidity. Despite this, its reported fewer than of women have a record of the examination in the recommended 6-8 weeks, with observed disparities in uptake nationally. The impact of the COVID-19 pandemic disrupted delivery of these checks nationally, but there is limited data on the impact of the pandemic and its recovery for regional populations representing diversity and areas of dense poverty and ethnic minority populations. This study utilised region level data to assess the impact of COVID-19 on postnatal care. Methods: Anonymised electronic health records with clinical coded birth events for females, aged 16-49 years, were analysed for patients registered with a GP using the Greater Manchester Care Record (GMCR) between January 2018 and August 2023. Unique delivery episodes were defined and monthly rates calculated separately for women with a postnatal-related code within 4-, 6-, 8-, or 12-weeks or 1 year follow-up. Rates were also generated by key maternal demographics to assess any differences in postpartum care. Interrupted time series, modelling the onset of the pandemic estimated the IRR of 0.49 (95% CI 0.40-0.58). To assess the impact of maternal characteristics on the odds of non-attendance at examination, a logistic regression adjusting for various maternal characteristics was fitted. Results: There were 114,874 unique delivery episodes, relating to 85,076 women in the 12-week follow up cohort; 72,595 episodes to 55,784 women in 8-weeks and 28,846 episodes to 24,018 women in 6-weeks. The rate of postpartum checks was greater the longer the follow-up period. For checks within 8 weeks the first lockdown reduced from ~325 per 1000 delivery episodes in 2019 to 225 per 1000 by April 2020 (30.8%), which remained low, before returning to pre-pandemic rates by rates by October 2022. Rates remained lower overall for Black, or Asian women compared to White. Conclusion: The COVID-19 pandemic reduced postnatal follow-up in primary care across Greater Manchester, with rates frequently falling outside the recommended 6-8 week window. Significant disparities exist in the provision and uptake of these services. Improved integration of data across care sites, combined with enhanced risk management, could increase equity in access and support the timely delivery of care for those at greatest risk of postnatal complications and longer-term health issues.
Sidebotham, D.; Barlow, J.
Show abstract
Background Multicentre trials in anaesthesia and critical care report low rates of statistically significant differences. This finding may partly reflect conventional sample size methods, which assume a fixed treatment effect. Assurance methods use a design prior to represent uncertainty in the expected treatment effect, which may provide a more realistic way of estimating sample sizes. Methods We calculated power curves across a range of effect sizes, design priors, and sample sizes using frequentist and Bayesian assurance methods and compared the sample sizes required to achieve 80% and 90% power to the conventional method. We standardised the design priors across effect sizes using the coefficient of variation. We derived a theoretical limit for achievable power. We validated a normal approximation to the Bayesian posterior distribution. Results Frequentist and Bayesian assurance methods produced similar power curves across all scenarios. At a coefficient of variation of 0.5 - reflecting realistic prior uncertainty in the expected effect size - both methods required sample sizes that were approximately 1.5 to 3.5 times larger than the conventional method. The theoretical power limit depends only on the coefficient of variation of the design prior and holds true across all effect sizes. The normal approximation to the Bayesian posterior distribution matched the results obtained from Markov chain Monte Carlo sampling. Conclusions Incorporating clinical uncertainty in the expected effect size substantially increases the sample size required to achieve adequate power, which has important implications for the feasibility of randomised trials in anaesthesia and critical care.
Pryymachenko, Y.; Wilson, R.; Abbott, J. H.
Show abstract
Objectives To analyse the long-term effects of a cruciate ligament (CL) injury on health and socioeconomic outcomes. Methods We used a comprehensive national injury insurance database to identify CL injuries occurring in New Zealand between 2009 and 2022, and employed a doubly robust staggered difference-in-differences research design to identify the effects of these injuries on outcomes up to 10 years after injury. The outcomes of interest were healthcare use (hospitalisations, emergency department visits, medications, knee replacement surgery for osteoarthritis), associated healthcare costs, and labour market outcomes (employment rates, income, and government benefit payments). Results We identified 61 344 CL injuries for inclusion in the analysis. Over 10-year follow-up, a CL injury resulted in increased healthcare use (0.6 more hospitalizations [95%CI 0.4 to 0.7], 1.7 more days spent in hospital [95%CI 1.3 to 2.1], 0.4 more emergency department visits [95%CI 0.3 to 0.6], 2.5 more outpatient visits [95%CI 1.8 to 3.2], and 4.7 more medications dispensed [95%CI -1.8 to 11.2]) and public healthcare costs ($7 537; 95%CI 5 888 to 9 186), reduced income (-$6 060; 95%CI -11 644 to -475), and increased benefit payments ($1 152; 95%CI 542 to 1 761). Conclusion CL injuries have long-term impacts on healthcare use and socioeconomic outcomes. Strategies to reduce the incidence of CL injuries have the potential to realise large health and economic benefits.
YOSHIHIRO, S.; KATAOKA, Y.; NISHIKIMI, M.; SHIME, N.; MATSUO, H.
Show abstract
Purpose To estimate the per-protocol effect of red blood cell (RBC) transfusion strategies on ICU-acquired infection in critically ill adults with sepsis using a target trial emulation framework. We evaluated whether restrictive strategy and liberal strategy, defined by hemoglobin (Hgb) thresholds, differ in their effect on ICU-acquired infection during ICU stay. Methods We conducted a target trial emulation using the MIMIC-IV database and included adults who met Sepsis criteria at ICU admission. Clones were assigned to restrictive or liberal transfusion strategies. Under the restrictive strategy, RBC transfusion was permitted only when Hgb was [≤]7.0 g/dL, whereas under the liberal strategy, transfusion was permitted when Hgb was >7.0 g/dL. The primary outcome was the first ICU-acquired infection occurring at least 72 hours after ICU admission. Per-protocol effects were estimated using a clone-censor-weight approach with a marginal structural model. A parametric g-formula was used as a complementary analysis that jointly modeled ICU discharge and ICU mortality as competing events to derive strategy-specific 28-day cumulative incidences and risk differences. Results 8 Among 4,013 eligible ICU stays, the liberal-versus-restrictive comparison provided little evidence of a difference in the risk of ICU-acquired infection (adjusted conditional OR, 0.954; 95% CI, 0.797 to 1.142). In the complementary g-formula analysis, the 28-day risk difference for the liberal versus restrictive comparison was -0.02 percentage points (95% CI, -0.15 to 0.11), consistent with the primary analysis. Findings were generally robust across prespecified subgroup and sensitivity analyses. Conclusion In this target trial emulation of adults with sepsis, we observed no clinically meaningful difference in ICU-acquired infection between RBC transfusion strategies defined by hemoglobin thresholds.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Baughman, D. J.; Liu, S.; Jee, S.; Young, C.; Knight, A. M.; Davis, A.; Yegnasubramanian, S.; Najjar, P.; Whitbread, J. J.; Ahumada, L.; Chused, A.; Haut, E. R.; Lau, B. D.; Sridharan, A.; Streiff, M.; Aziz, K. B.
Show abstract
Background: Guiding risk-appropriate inpatient thromboprophylaxis requires venous thromboembolism (VTE) risk stratification; however, reliable risk determination remains inconsistent in routine care. Health systems increasingly pilot artificial intelligence (AI) tools, yet few studies demonstrate rigorous evaluation in the context of a learning health system (LHS). We evaluated the performance of a pilot electronic health record (EHR)-integrated generative AI (GenAI) system, inHealth General Reasoner (iHGR), for VTE risk stratification versus clinician order set classifications and physician-adjudicated chart review. Methods: This multisite retrospective validation study included adult inpatient admissions at Johns Hopkins Medicine between June 21, 2025, and Dec 18, 2025 (checklist-based order set from June 21, 2025 - November 19, 2025, and clinician judgement-based order set from November 29 - December 18, 2025). From 758 eligible admissions, we randomly sampled 500 balanced by site and order set periods. iHGR and clinician-selected order set classifications were compared with the reference standard (RS). Primary outcomes were iHGR sensitivity and specificity. Secondary analyses compared the order sets with the same RS to evaluate workflow comparators and error patterns. Results: iHGR achieved 81.8% sensitivity (95% CI 77.3-85.6) and 70.9% specificity (63.6-77.3). The checklist-based order set had 61.3% sensitivity (53.7-68.5) and 86.2% specificity (77.4-91.9). The clinician judgement-based order set had 78.1% sensitivity (71.3-83.7) and 65.4% specificity (54.3-75.0). False-negative iHGR classifications were associated with missed narrative risk factors. Conclusion: iHGR showed higher sensitivity for VTE risk than checklist-based order sets and clinician judgement without introducing systematic bias. In silico evaluation of pilot AI systems within LHSs can identify clinically important performance trade-offs and implementation targets before operational scale-up. Narrative clinical data abstraction remained a key limitation, supporting the use of GenAI to support rather than supplant clinician judgement.
Young, K. G.; Banerjee, A.; Dayan, C.; Denaxas, S.; Eastwood, S. V.; Jeffery, A.; Rutter, M. K.; Sattar, N.; Valabhji, J.; Horswood, R.; Humphreys, R.; Molete, M.; Murray, K.; Rogers, P.; Veiro, D.; Ireland, H.; Walker, C.; Shields, B. M.; Pearson, E. R.; McGovern, A. P.; Dennis, J. M.
Show abstract
Aims To develop a 'core' dataset of diabetes related variables to support reproducible research using UK routinely collected health data. Methods A workshop was conducted bringing together diabetes healthcare professionals, researchers, and patient and public representatives to discuss and prioritise variables for inclusion in the Diabetes Core Dataset. Core variables were those considered to be highest priority for diabetes research and available at high quality in NHS data routinely used for research (primary care [GP] and Hospital Episode Statistics [HES] data). Candidate variables for inclusion in the Diabetes Core Dataset were from a review of existing core datasets and expert opinion. Participants scored variables anonymously based on priority for diabetes research. Results 25 variables from existing diabetes core datasets and 87 other candidate variables were considered for inclusion in the Diabetes Core Dataset. All 25 of those from existing diabetes core datasets and 5 of the 87 candidate variables met the core requirements for inclusion. In addition, 7 variables were identified as high priority but not included in the core dataset as they are not currently available in GP/HES data; these were labelled as 'future high priority' variables for diabetes research. Conclusions A new diabetes core dataset for UK EHR research has been developed using a consensus-based process. The core dataset is openly available and can be flexibly applied in UK EHR (https://healthdatagateway.org/en/tool/426), including in new NHS Research Secure Data Environment platforms, to enhance reproducible research to improve the clinical care of people with diabetes and associated conditions.
Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.
Show abstract
Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [≥]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.