Back

BMJ

BMJ

Preprints posted in the last 30 days, ranked by how well they match BMJ's content profile, based on 51 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Do video animations reduce literacy-related inequalities in understanding of health information? A secondary analysis of a systematic review of intervention trials

Moe-Byrne, T.; Knapp, P.; Golder, S.

2026-08-17 health informatics 10.64898/2026.08.14.26360327 medRxiv
Top 0.1%
14.9%
Show abstract

Background People with lower levels of literacy or health literacy may struggle to understand conventional health information. Video animations show promise as information tools, yet it is unclear whether video animations help reduce these inequalities in understanding. This study examined whether the effectiveness of video animations in health settings differs according to level of literacy or health literacy. Methods We drew on trials from a recent systematic review of video animations about healthcare or public health topics for patients or the public. We extracted available data on literacy, health literacy, or proxy indicators. One reviewer extracted data and a second checked all entries. Where possible, we conducted subgroup analyses of low and high literacy levels or interaction meta-analyses comparing low versus high literacy groups; otherwise, results were summarised narratively. Results From 88 eligible trials, we extracted health literacy data for 12. Across nine trials reporting knowledge, animations mostly improved knowledge compared with controls in both lower and higher health literacy groups. Effects on attitudes and behaviours were mixed and often small, with few studies reporting results by health literacy level. Across the subgroup analyses available, there was no consistent evidence of a pooled interaction effect of animations according to low and high literacy groups, but both statistical heterogeneity and small subgroup sizes limited precision of estimates. Across 88 trials, 54 (61%) reported education level, 22 (25%) did not, and 12 (14%) involved children or adolescents likely to have similar education levels. Conclusions Overall, the available data suggest that video animations can improve knowledge outcomes in both lower and higher health literacy groups, but their impact on attitudes and behaviour is less clear. Because literacy was rarely reported or analysed in the trials, it remains uncertain whether animations help to reduce literacy-related inequalities in access to, and use of health information.

2
Machine Learning-Supported Efficient VTE Risk Assessment using Routinely Collected Electronic Health Record Data

Li, Z.; Yagis, E.; Riad, A.; Windrath-Carr, O.; Arribas, M.; Sodiq, T.; Goldsmith, K.; Glampson, B.; Flott, K.; Haji, G.; Khan, Z.; Baker, C.; Mayer, E. K.

2026-08-19 health informatics 10.64898/2026.08.18.26360687 medRxiv
Top 0.1%
10.5%
Show abstract

Venous thromboembolism (VTE) is a leading cause of preventable inpatient mortality, while the real-world performance of mandated risk assessment and the potential for automating using electronic health record (EHR) data remain unclear. We analysed 577,904 admissions and 726,896 VTE assessment forms across five NHS hospitals between 2015 and 2025 to evaluate assessment completion, concordance with structured EHR data, clinical validity, and feasibility of EHR-based automation assisted by machine learning. Overall completion was high (96.7%), and timely completion improved from 47.4% in 2015 to 90.5% in 2024. Agreement between forms and EHR data was good for common risk factors, but low-prevalence variables were often under-documented in the forms. Despite these discrepancies, form-derived thrombosis risk was associated with increased VTE incidence (OR 3.31, 95% CI 2.81-3.90). Machine learning models using first-14-hour EHR data achieved discrimination comparable to clinician-recorded variables (AUROC 0.709 vs 0.704), supporting real-time EHR-integrated assessment pre-population and decision support.

3
Effects of collaborative clinical visit agenda-setting interventions: A systematic review and meta-analysis

Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.

2026-09-03 medical education 10.64898/2026.08.30.26361729 medRxiv
Top 0.1%
10.1%
Show abstract

Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.

4
Algorithmic Ascertainment of Cause of Death from Longitudinal Real-World Medical Claims Data: Development and Validation

McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360606 medRxiv
Top 0.1%
9.7%
Show abstract

This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.

5
People living with multiple long-term conditions have different pathways of unscheduled care in hospital: findings from an analysis of routinely-collected clinical data

Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.

2026-09-01 health informatics 10.64898/2026.08.28.26361696 medRxiv
Top 0.1%
6.9%
Show abstract

Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.

6
Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.

2026-08-21 health informatics 10.64898/2026.08.18.26360701 medRxiv
Top 0.1%
6.7%
Show abstract

Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [&ge;]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.

7
Optimising the prevention, detection and management of diabetes distress in adults with type 1 diabetes: Feasibility study protocol (D-stress feasibility study).

Fabian-Therond, C.; Ahuja, S.; Papachristou Nadal, I.; Holt, R. I.; Watson, S. I.; Hussain, S.; Choudhary, P.; Ajjan, R.; Harris, R.; Peck, M.; Mohammadi, J.; Sims, S.; Fiorentino, F.; Due-Christensen, M.; Huber, J.; Fisher, L.; Hardenberg, K.; Stadler, M.; Jin, H.; Halliday, J. A.; Sturt, J.; on behalf of the D-stress study collaborators,

2026-08-18 endocrinology 10.64898/2026.08.17.26360546 medRxiv
Top 0.1%
5.5%
Show abstract

Introduction Diabetes distress describes the psychological and emotional burden of living with diabetes and is associated with reduced self-management and adverse diabetes outcomes. Clinical guidelines recommend routine assessment and management of diabetes distress, but this is not always implemented. Therefore, there is a need to develop approaches to deliver emotional health support in routine clinical care more effectively. We describe here the protocol for a study to I) assess the feasibility of implementation of the D-stress Pathway, comprising Enhanced Usual Care (EUC) and an online, group-based, psychological diabetes distress reduction intervention called REDUCE, ii) evaluate the feasibility of the study protocol iii) detect an effect signal of diabetes distress score and Interstitial Glucose Time in Range and iv) refine initial programme theories of how both interventions (EUC and REDUCE) work, for whom, and under what circumstances. Methods This feasibility study includes a multicentre trial within a cohort design (TWICs) where sites have a staggered exposure to the interventions alongside a realist process evaluation. Four UK NHS diabetes services will recruit 80 adults with type 1 diabetes ([&ge;]1 year) using continuous glucose monitoring (CGM) ([&ge;]3 months). All participants will receive EUC and provide monthly data over 7 months on diabetes distress (measured by the Type 1 Diabetes Distress Assessment System (T1DDAS) and interstitial glucose measured by using continuous glucose monitoring. Participants with elevated diabetes distress, will be offered the six-week, group-based, online REDUCE intervention plus EUC, compared to EUC alone. Up to twenty participants with type 1 diabetes, ten family members/friends, sixteen healthcare professionals delivering EUC and five REDUCE facilitators will be interviewed to explore their experience of receiving training and delivering the D-stress Pathway. Up to 20 EUC consultations and REDUCE sessions will be observed. Analysis Feasibility will be assessed against pre-specified progression criteria and analysed descriptively using summary statistics. Primary outcomes include baseline level of diabetes distress, recruitment rate, intervention uptake, and data completeness, which will be analysed descriptively. Qualitative data will be analysed using framework analysis guided by realist programme theories developed for this study. Ethics Ethics approval has been granted by NHS Research Ethics Committee (REC) (Bromley REC: 25/LO/0469) and Health Research Authority obtained. All participants will provide informed consent. Trial registration no: Registered at ClinicalTrials.gov number NCT07193446 on 26/11/2025. Protocol and statistical analysis plan: The trial protocol and statistical analysis plan can be accessed at ClinicalTrials.gov.

8
Reduced maternal healthcare interactions with general practice in the postnatal period during the COVID-19 pandemic, a cohort study of Greater Manchester residents.

Cornett, C.; Tilston, G.; Martin, G.; Palin, V.

2026-08-22 health informatics 10.64898/2026.08.18.26360757 medRxiv
Top 0.1%
4.9%
Show abstract

Background: Maternal postpartum checks with a general practitioner (GP) are recognised as an essential service in England and vital for recovery after pregnancy and reducing risk of long-term morbidity. Despite this, its reported fewer than of women have a record of the examination in the recommended 6-8 weeks, with observed disparities in uptake nationally. The impact of the COVID-19 pandemic disrupted delivery of these checks nationally, but there is limited data on the impact of the pandemic and its recovery for regional populations representing diversity and areas of dense poverty and ethnic minority populations. This study utilised region level data to assess the impact of COVID-19 on postnatal care. Methods: Anonymised electronic health records with clinical coded birth events for females, aged 16-49 years, were analysed for patients registered with a GP using the Greater Manchester Care Record (GMCR) between January 2018 and August 2023. Unique delivery episodes were defined and monthly rates calculated separately for women with a postnatal-related code within 4-, 6-, 8-, or 12-weeks or 1 year follow-up. Rates were also generated by key maternal demographics to assess any differences in postpartum care. Interrupted time series, modelling the onset of the pandemic estimated the IRR of 0.49 (95% CI 0.40-0.58). To assess the impact of maternal characteristics on the odds of non-attendance at examination, a logistic regression adjusting for various maternal characteristics was fitted. Results: There were 114,874 unique delivery episodes, relating to 85,076 women in the 12-week follow up cohort; 72,595 episodes to 55,784 women in 8-weeks and 28,846 episodes to 24,018 women in 6-weeks. The rate of postpartum checks was greater the longer the follow-up period. For checks within 8 weeks the first lockdown reduced from ~325 per 1000 delivery episodes in 2019 to 225 per 1000 by April 2020 (30.8%), which remained low, before returning to pre-pandemic rates by rates by October 2022. Rates remained lower overall for Black, or Asian women compared to White. Conclusion: The COVID-19 pandemic reduced postnatal follow-up in primary care across Greater Manchester, with rates frequently falling outside the recommended 6-8 week window. Significant disparities exist in the provision and uptake of these services. Improved integration of data across care sites, combined with enhanced risk management, could increase equity in access and support the timely delivery of care for those at greatest risk of postnatal complications and longer-term health issues.

9
Long-term outcomes of cruciate ligament injury: evidence from New Zealand linked register data

Pryymachenko, Y.; Wilson, R.; Abbott, J. H.

2026-09-01 epidemiology 10.64898/2026.08.27.26361565 medRxiv
Top 0.1%
4.5%
Show abstract

Objectives To analyse the long-term effects of a cruciate ligament (CL) injury on health and socioeconomic outcomes. Methods We used a comprehensive national injury insurance database to identify CL injuries occurring in New Zealand between 2009 and 2022, and employed a doubly robust staggered difference-in-differences research design to identify the effects of these injuries on outcomes up to 10 years after injury. The outcomes of interest were healthcare use (hospitalisations, emergency department visits, medications, knee replacement surgery for osteoarthritis), associated healthcare costs, and labour market outcomes (employment rates, income, and government benefit payments). Results We identified 61 344 CL injuries for inclusion in the analysis. Over 10-year follow-up, a CL injury resulted in increased healthcare use (0.6 more hospitalizations [95%CI 0.4 to 0.7], 1.7 more days spent in hospital [95%CI 1.3 to 2.1], 0.4 more emergency department visits [95%CI 0.3 to 0.6], 2.5 more outpatient visits [95%CI 1.8 to 3.2], and 4.7 more medications dispensed [95%CI -1.8 to 11.2]) and public healthcare costs ($7 537; 95%CI 5 888 to 9 186), reduced income (-$6 060; 95%CI -11 644 to -475), and increased benefit payments ($1 152; 95%CI 542 to 1 761). Conclusion CL injuries have long-term impacts on healthcare use and socioeconomic outcomes. Strategies to reduce the incidence of CL injuries have the potential to realise large health and economic benefits.

10
Effect of red blood cell transfusion strategies on ICU-acquired infection in patients with sepsis: A target trial emulation using the Medical Information Mart for Intensive Care IV database

YOSHIHIRO, S.; KATAOKA, Y.; NISHIKIMI, M.; SHIME, N.; MATSUO, H.

2026-08-17 intensive care and critical care medicine 10.64898/2026.08.14.26360449 medRxiv
Top 0.1%
4.1%
Show abstract

Purpose To estimate the per-protocol effect of red blood cell (RBC) transfusion strategies on ICU-acquired infection in critically ill adults with sepsis using a target trial emulation framework. We evaluated whether restrictive strategy and liberal strategy, defined by hemoglobin (Hgb) thresholds, differ in their effect on ICU-acquired infection during ICU stay. Methods We conducted a target trial emulation using the MIMIC-IV database and included adults who met Sepsis criteria at ICU admission. Clones were assigned to restrictive or liberal transfusion strategies. Under the restrictive strategy, RBC transfusion was permitted only when Hgb was [&le;]7.0 g/dL, whereas under the liberal strategy, transfusion was permitted when Hgb was >7.0 g/dL. The primary outcome was the first ICU-acquired infection occurring at least 72 hours after ICU admission. Per-protocol effects were estimated using a clone-censor-weight approach with a marginal structural model. A parametric g-formula was used as a complementary analysis that jointly modeled ICU discharge and ICU mortality as competing events to derive strategy-specific 28-day cumulative incidences and risk differences. Results 8 Among 4,013 eligible ICU stays, the liberal-versus-restrictive comparison provided little evidence of a difference in the risk of ICU-acquired infection (adjusted conditional OR, 0.954; 95% CI, 0.797 to 1.142). In the complementary g-formula analysis, the 28-day risk difference for the liberal versus restrictive comparison was -0.02 percentage points (95% CI, -0.15 to 0.11), consistent with the primary analysis. Findings were generally robust across prespecified subgroup and sensitivity analyses. Conclusion In this target trial emulation of adults with sepsis, we observed no clinically meaningful difference in ICU-acquired infection between RBC transfusion strategies defined by hemoglobin thresholds.

11
Multi-model LLM assessment of Quality Control Circlemethodological quality: a designed-anchor reliabilitystudy

LIn, H.; Lyu, J.

2026-08-13 health informatics 10.64898/2026.08.12.26360276 medRxiv
Top 0.1%
4.1%
Show abstract

BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.

12
Acceptability, feasibility, quality of life and diabetes distress score outcomes: A pragmatic randomised clinical trial on continuous glucose monitoring for people with type 1 diabetes

Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.

2026-08-31 endocrinology 10.64898/2026.08.26.26361479 medRxiv
Top 0.2%
4.0%
Show abstract

Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [&ge;]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.

13
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.2%
3.5%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

14
Measuring implementation of clinical guidelines through the COVID-19 pandemic, using linked national health records: a national study of type 2 diabetes in England

Biglarbeigi, P.; Dale, C.; Lambarth, A.; Mason, A.; Takher, R.; Ballabio, G.; Minshull, J.; Mamas, M. A.; Tomlinson, C.; Rowark, S.; Rayman, G.; Pearson, E. R.; Khunti, K.; Sattar, N.; Sofat, R.

2026-08-10 health policy 10.64898/2026.08.07.26359950 medRxiv
Top 0.2%
3.4%
Show abstract

Objectives: To examine the conformance to type 2 diabetes NICE guidelines across cardiovascular risk strata; and to quantify geographical variation in treatment pathways following the COVID-19 pandemic, encompassing guideline changes. Design: We carried out a retrospective observational study using linked electronic health records across England. Process mining, a data driven method that can reconstruct clinical treatment pathways, was applied to map 12-month treatment trajectories after treatment initiation. Conformance with NICE NG28 (2022) was quantified using a structural similarity index. Further, behavioural and entropy-based similarity (capturing treatment variability and complexity) measures were used to assess sequencing and heterogeneity of treatment. Setting: Primary and secondary care in England datasets within the National Health Service England Secure Data Environment (NHSE SDE), analysed first at national level and then across 42 Integrated Care Boards (ICBs) which are the devolved health care geographical delivery regions in England. Participants: 822,650 individuals with newly diagnosed T2DM between 1-February-2022 and 1-November-2025, stratified into low cardiovascular risk (LR-C; QRISK3<10), high risk (HR-C; QRISK3>=10 or receiving statins/blood pressure lowering treatment), and established cardiovascular disease (eCVD-C). Participants were followed for 12 months after first dispensed glucose lowering therapy. Main outcome measure: First line therapy, treatment intensification and switching within 12 months; change in glycated haemoglobin (HbA1c); quantified conformance to NICE recommended pathways; and regional variation in broader similarity measures. Results: Metformin monotherapy was the dominant initiation strategy in LR-C and HR-C cohorts (92.4% and 90.2%, respectively), whereas eCVD-C showed lower uptake of metformin (68.9%) and higher uptake of SGLT2 inhibitors (26.3%). Intensification from metformin to combination therapy was infrequent across all cohorts (<1%), although HR-C demonstrated the highest treatment transitions and switching behaviour. Dispensed SGLT2 inhibitor use was nearly threefold higher in eCVD-C (26.7%) than in LR-C (9.0%) or HR-C (10.7%). Overall, conformance to NICE-recommended pathways remained modest nationally, particularly in LR-C and HR-C. Across 42 ICBs, substantial regional heterogeneity in treatment pathways and guideline conformance was observed, with conformance ranging from 0.29 to 1.00 in LR-C pathways, 0.40 to 1.00 in HR-C pathways, and 0.54 to 0.92 in eCVD pathways. Conclusion: National T2DM treatment pathways for post-pandemic showed higher alignment to NICE guidelines in eCVD-C compared to the other risk groups, with ongoing gaps and large regional variations in other risk groups. Process mining offers a scalable approach to monitor implementation of guideline recommended care that could support learning health systems. Using T2DM during and post COVID-19 pandemic as a case study, this work demonstrates how these methods can assess the use of existing and innovative therapies, identify gaps and guide future adoption to ensure recommended treatments reach the right patient groups.

15
Prognostic Language and Subsequent Code-Status Limitation After Acute Brain Injury: A Multidatabase Observational Study

Gorenshtein, A.; Adiniaev, Y.; Srour, A.; Klang, E.; Daniel, O.

2026-08-31 intensive care and critical care medicine 10.64898/2026.08.27.26361534 medRxiv
Top 0.2%
2.9%
Show abstract

Purpose. Prognostic assessments after acute brain injury are largely narrative, and how prognostic language relates to subsequent care has not been measured at scale. We quantified where it is written and its association with a subsequent code-status limitation. Materials and Methods. Multidatabase observational study of adults with acute brain injury or a related neurologic emergency, using MIMIC-IV (2008-2019; discharge summaries and radiology reports) and a timestamped MIMIC-III cohort (notes and code-status orders). The exposure was documented prognostic language; outcomes were its association with a subsequent full-code-to-limitation transition, note-stream location, and completeness of documented command-following relative to structured Glasgow Coma Scale (GCS) motor scores. Results. Among 31,993 admissions (27,054 patients; median age, 69 years; 54.9% male), prognostic language in the timestamped cohort (MIMIC-III) was associated with a subsequent code-status limitation after multivariable adjustment (adjusted hazard ratio, 4.3; 95% CI, 2.9-6.5; unadjusted 14-day cumulative incidence, 40% vs 8.5%), including the comfort-measures component (3.9), a higher-risk subgroup (4.4), and after acute-physiology adjustment (4.1); the association was concentrated in the first 3 days. Non-prognostic severity language showed no comparable association (hazard ratios, 1.1-1.3). Prognostic language localized almost entirely to the narrative (4.9% of discharge summaries vs 0.015% of radiology reports); command-following was undocumented in 55.7% of summaries, and no final-24-hour GCS motor score was charted in 72.8%. Conclusions. Documented prognostic language after acute brain injury was written in the narrative, not structured fields, and was associated with a subsequent code-status limitation after multivariable adjustment. This observational association cannot establish causation but warrants prospective study.

16
Efficacy and safety of PCSK9 inhibitors for children and adolescents with heterozygous familial hypercholesterolaemia: Systematic review and meta-analysis of randomised controlled trials

Llewellyn, A.; Simmonds, M.; Marshall, D.; Harden, M.; Humphries, S. E.; Woods, B.; Gomes, M.; Priestley-Barnham, L.; Ramaswami, U.; Fisher, M.; Qureshi, N.; Tata, L. J.

2026-08-06 cardiovascular medicine 10.64898/2026.08.04.26359681 medRxiv
Top 0.2%
2.8%
Show abstract

Background Statins and ezetimibe are the preferred lipid-lowering therapies (LLTs) for children with heterozygous familial hypercholesterolaemia (HeFH). Proprotein convertase subtilisin/kexin type 9 inhibitors (PCSK9i) are newer add-on therapies for individuals not achieving low-density lipoprotein-cholesterol (LDL-C) targets. We evaluated the efficacy and safety of PCSK9i in children aged <18 years with HeFH. Methods Systematic review and pairwise meta-analyses of randomised-controlled trials (RCTs) of evolocumab, alirocumab and inclisiran. Comprehensive bibliographic searches were conducted in February 2026. Risk of bias was assessed with Cochrane RoB 2. Results Of 2798 unique records screened, three RCTs were included (n=451, mean age 13 years, follow-up 24 to 47 weeks). Each trial evaluated either evolocumab, alirocumab or inclisiran against placebo as add-on to baseline LLT. Participants had elevated LDL-C (>3.4 mmol/L [130 mg/dL]) despite stable LLT. Overall risk of bias was low. PCSK9i reduced LDL-C by an average of 35.44% (95% CI -41.74 to -29.14, I2=50.8%) and by 1.63 mmol/L [62.93 mg/dL] (95% CI -1.86 to -1.39, I2=18.6%) compared with placebo. There was no evidence of differences between PCSK9i and placebo in tolerability, growth and maturation, and overall incidence of adverse events. Conclusions PCSK9i add-on therapy leads to substantial reductions in LDL-C in paediatric patients with HeFH failing to achieve LDL-C targets with standard LLT. While the findings of this review support the use of PCSK9i in a subset of children and young people with HeFH, limited trial numbers and short follow-up periods underscore the need for future high-quality studies evaluating long-term safety, effectiveness and cost-effectiveness.

17
Radiographically identified vertebral fractures in haemochromatosis-associated HFE C282Y homozygotes in the UK Biobank

Banfield, L. R.; Pilling, L. C.; Melzer, D.; Shearman, J.; Knapp, K.; Atkins, J. L.

2026-08-22 epidemiology 10.64898/2026.08.19.26360796 medRxiv
Top 0.2%
2.7%
Show abstract

Abstract Purpose: Haemochromatosis due to HFE-C282Y homozygosity can lead to excess iron absorption and is typically associated with liver malignancy, plus widespread arthritis. Recent evidence suggests that limb fractures are more common, but little is known about vertebral effects. This study investigated the association of vertebral compression fractures, assessed with intelligent dual-energy X-ray absorptiometry (iDXA), and HFE genotype in a large community cohort. Methods: UK Biobank data from 227 European genetic ancestry C282Y homozygotes (mean 64.6 years) and 234 age, sex, and BMI-matched controls without common HFE haemochromatosis variants were included. Lateral vertebral assessment scans (iDXA, GE-Lunar) were acquired at imaging reassessment (2014-2020) and reviewed, blind to genotype, for radiological evidence of vertebral fracture. Matched logistic regression models assessed associations between C282Y homozygosity and vertebral fractures. Results: 78 vertebral fractures (16.9%) were identified within 461 participants. Male C282Y homozygotes had increased odds of vertebral fracture (n=22/89, 24.7%) compared to participants without HFE alleles (n=9/90, 10.0%); Odds Ratio [OR]: 2.95, 95%CI: 1.28-6.85, p=0.01. The association persisted after excluding individuals with a diagnosis of haemochromatosis (OR: 3.37, 95% CI: 1.41-8.10, p=0.007). No excess fracture risk was observed in female C282Y homozygotes (n=23/138, 16.7%) vs those without HFE alleles (n=24/144, 16.7%); OR: 0.99, 95%CI: 0.53-1.87, p=1.00. Conclusion: In this community-based imaging study, male HFE C282Y homozygotes had a markedly higher likelihood of vertebral fractures than those without HFE variants. These findings support further evaluation of vertebral fracture assessment in C282Y homozygous men to ensure prompt treatment to prevent future fracture if appropriate.

18
Real-World Performance of the 2026 AHA/ACC Pulmonary Embolism Framework in a Multi-System CTPA Cohort

Alwakeel, M.; Zaveri, S.; Buck, E.; Rajagopal, S.; Verma, D.; Loriaux, D.; Henao, R.; Tapson, V. F.; Ortel, T. L.; Jones, W. S.; Martin, J. G.; Haines, K. L.; Freeman, N. L.; Wong, A.-K. I.

2026-08-10 health informatics 10.64898/2026.08.06.26359865 medRxiv
Top 0.3%
2.7%
Show abstract

Background: The 2026 American Heart Association/American College of Cardiology (AHA/ACC) guidelines replaced the 2019 European Society of Cardiology (ESC) four-tier pulmonary embolism (PE) risk scheme with five clinical categories (A-E) and subcategories. These categories were set by expert consensus and have not been validated against outcomes. How patients are reclassified relative to ESC, or how the two systems compare prognostically, is unknown. Methods: We utilized three cohorts of patients with confirmed PE using structured electronic health record data, laboratory biomarkers, and large-language-model abstraction of radiology reports: Duke University Health System (n=12,992, drawn from 95,760 consecutive inpatient CT pulmonary angiography studies, 2014-2025, with no referral or registry enrollment step between imaging and cohort entry), INSPECT (Stanford; n=3,870), and MIMIC-IV (Beth Israel Deaconess; n=361). Patients were assigned AHA/ACC categories B through E, subcategorized where data allowed, and mapped to 2019 ESC risk strata. The primary outcome was 30-day mortality; discrimination was assessed with Harrell C-index. Results: Among 17,223 patients with confirmed PE, pooled 30-day mortality rose monotonically across categories: 1.5% (B), 8.9% (C), 15.5% (D), and 31.9% (E), with the ordering preserved in all three cohorts despite differing baseline mortality. Subcategory-level discrimination was reliable only at the high-acuity extreme (D2-E2); across subcategories C1 through D1, mortality did not order monotonically (9.2%, 10.8%, 8.1%, 10.9%), and adding subcategories to category C did not improve discrimination at Duke (C-index 0.699 vs 0.699). Category C patients lacking both echocardiography and biomarker testing (12.7% of category C) had mortality (10.4%) equal to or exceeding classified peers. Relative to ESC, the frameworks were concordant at the extremes, but 5.7%of ESC intermediate-risk patients were reclassified to category D, with modestly higher but non-significant 30-day mortality than those remaining in category C (10.8% versus 8.9%). Conclusions: Across a three-health-system cohort, the 2026 AHA/ACC framework produced a reproducible mortality gradient at the category level, with added subcategory granularity refining risk chiefly at the highest-acuity tiers. Discrimination across the broad intermediate band was limited, and reclassification from ESC fell almost entirely within this range.

19
Twelve-Year Real-World Evaluation of a Regulated Guideline-Based Warfarin Dosing and Care Automation System

Tiihonen, M.

2026-08-12 health informatics 10.64898/2026.08.10.26360059 medRxiv
Top 0.3%
2.6%
Show abstract

Background: Warfarin therapy requires repetitive dose adjustments based on INR (International Normalised Ratio) monitoring. We evaluated the long-term real-world performance of Forsante Warfarin Advisor (WA), a CE-marked class IIb guideline-based decision support and care automation medical device used in anticoagulation management. Methods: Retrospective real-world data from routine clinical use between 2016 and 2026 were analysed. Treatment quality was assessed using Time in Therapeutic Range (TTR). Recommendation performance was evaluated by comparing achievement of target INR after clinician acceptance or modification of Warfarin Advisor recommendations. Results: Among 1348 patients in March 2026 median TTR was 83%, compared with 70% in March 2016. Dosages congruent with Warfarin Advisor recommendations were strongly associated with achieving target INR at follow-up in INR target ranges of 2.0-3.0 and 2.5-3.5. Treatment quality remained consistently high across years of deployment. No serious device-attributable safety incidents, regulatory incident reports, or CAPA cases were identified during 12 calendar years and 82,709 patient years of routine use. Conclusions: The findings provide real-world long-term evidence that a guideline-based warfarin dosing and care automation system can support sustained high-quality anticoagulation control in routine clinical practice. The findings support the feasibility of deploying workflow-integrated execution of selected guideline-driven clinical processes, while the causal effects on clinical outcomes require prospective confirmation. Keywords: Clinical decision support systems, Guideline execution, Real-world evidence, Warfarin, Anticoagulation

20
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 0.3%
2.4%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.