The Lancet Digital Health
○ Elsevier BV
All preprints, ranked by how well they match The Lancet Digital Health's content profile, based on 25 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Tamm, A.; Shine, B.; James, T.; Withers, J.; Salih, H.; East, J. E.; Oke, J.; Davies, J.; Morris, E. J.; Nicholson, B. D.
Show abstract
Background The faecal immunochemical test (FIT) is central to triaging symptomatic patients with suspected colorectal cancer (CRC) in UK primary care, yet only about one in eleven patients above the NICE 10 ug/g threshold have CRC. Existing prediction models attempting to improve on FIT have relied on conventional statistics and limited predictors. Methods GP-requested FITs with linked data (Jan 2017 - May 2025) were extracted from the Oxford University Hospitals (OUH) datawarehouse. Patients aged [≥]18 with core bloods and 180-day CRC follow-up were included. Machine learning (ML) models were trained on up to 1,025 predictors: FIT, age, sex, blood tests and their time series slopes, diagnoses/procedures/prescriptions, deprivation, BMI, and ethnicity. Models comprised penalised logistic regression, generalised additive models (EBM, NAM, SNAM, NODE-GAM), decision tree ensembles (random forests, XGBoost), and a multilayer perceptron. Referral reduction versus FIT [≥]10 ug/g was evaluated at model risk score thresholds capturing the same cancers (conservative) or same proportion of cancers (less conservative) as FIT. Potential to prioritise referred patients was assessed by examining whether positive predictive value (PPV) is very high (>30%) at any substantial sensitivity (>10%). Nested twice-repeated five-fold cross-validation provided unbiased estimates. An existing COLOFIT model was evaluated alongside. Findings 62,219 individuals (746 CRC) were analysed; 30,862 patients (315 CRC) with high/low risk symptoms and buffered FITs formed the primary subset. At [≥]10 ug/g, FIT had 91.4% sensitivity, 84.2% specificity, 5.6% PPV, and 99.9% NPV. No model reduced referrals when required to capture the same cancers as in the FIT [≥]10 ug/g cohort. Generalised additive models achieved up to 18.5% referral reduction when detecting the same proportion but some different cancers as FIT [≥]10 ug/g (EBM: 18.5%, NODE-GAM: 17.5%, SNAM: 17.4%, COLOFIT: 16.7%). At 30% sensitivity, EBM, NAM and NODE-GAM had average PPVs between 34.6%-35.0%, while FIT had a PPV of 14.6%. Interpretation Generalised additive models (GAMs) reduced referrals on average by 19% if a small proportion of the FIT-positive CRCs were substituted with originally FIT-negative CRCs by the models. No model, including COLOFIT, reduced referrals while capturing all FIT-positive cancers. Generalised additive models could detect about a third of CRCs faster, as one in three patients flagged by the models had CRC at 30% sensitivity. Funding EPSRC Centre for Doctoral Training in Health Data Science; National Institute for Health Research (NIHR) Oxford Biomedical Research Centre; Cancer Research UK. Keywords Colorectal cancer, faecal immunochemical test, machine learning, positive predictive value
Rao, S.; Li, Y.; Mamouei, M.; Salimi-Khorshidi, G.; Wamil, M.; Nazarzadeh, M.; Yau, C.; Collins, G.; Jackson, R.; Vickers, A.; Danaei, G.; Rahimi, K.
Show abstract
ObjectiveTo develop and validate the Transformer-based Risk assessment survival model (TRisk), a novel deep learning model, for prediction of 10-year risk of cardiovascular disease (CVD) in both the general population and individuals with diabetes. DesignProspective open cohort study design. SettingPrimary and secondary care in England as provided by Clinical Practice Research Datalink (CPRD) Gold ParticipantsAn open cohort of 3 million adults aged 25 to 84 years was identified using linked primary and secondary electronic health records from 291 and 98 general practices in England and were used for model development and validation, respectively (i.e., general population cohort). Additionally, a second cohort of patients with diabetes was extracted. At study entry, patients in both cohorts were free of CVD and not prescribed statins. MethodsTRisk utilised all diagnosis, medication, procedure, and clinical test data up to study entry in linked longitudinal primary and secondary care electronic health records for prediction of 10-year risk of CVD. Discrimination, calibration, and decision curve analyses were conducted to investigate predictive performance. The proposed model was also compared against QRISK3 and a deep learning derivation model of QRISK3 (DeepSurv). Additional analyses compared discriminatory performance in other age groups, by sex, and across categories of socioeconomic status. Main outcome measuresIncident cardiovascular disease recorded in either linked general practice or hospital admission datasets provided by CPRD Gold. ResultsTRisk demonstrated superior discrimination (C-index in the general population: 0.910; 95% confidence interval [CI]: 0.906 to 0.913). TRisks performance was found to be less sensitive to population age range than the benchmark models and outperformed other models also in analyses stratified by age, sex or socioeconomic status. All models were overall well-calibrated. In decision curve analyses, TRisk demonstrated greater net benefit than benchmark models across the range of relevant thresholds. At both the recommended 10% risk threshold and the 15% risk threshold, TRisk reduced both the total number of patients classified at high risk (by 22% and 35% respectively) and the number of false negatives as compared with currently recommended strategies. TRisk similarly outperformed other models in patients with diabetes. Compared with the widely recommended treat-all policy approach for patients with diabetes, TRisk at a 10% risk threshold would lead to deselection of 24% of individuals with a small fraction of false negatives (0.2% of cohort). ConclusionTRisk enabled a more targeted selection of individuals at risk of CVD compared to benchmark statistical and deep learning models, in both the general population and patients with diabetes. Incorporation of TRisk into routine clinical care would allow a reduction in the number of treatment-eligible patients by approximately one-third while preventing at least as many events as with currently adopted approaches.
van Dokkum, E. D.; Kraaijenbrink, N.; le Cessie, S.; Sijbom, M.; van der Schoor, A. S.; Visser, L. G.; van Nieuwkoop, C.; Borgdorff, H.
Show abstract
While socioeconomic status (SES) and migration background have been linked to complicated lower respiratory tract infections (LRTIs) in population-based studies, their predictive value in primary care remains unclear. Using routine care data from Dutch general practices (Leiden-The Hague-Zoetermeer region, n {approx} 750,000 adult patients, 2014 to 2023, excluding COVID-19 years), linked to sociodemographic and hospital claims data, we developed a multivariable logistic regression model to predict 30-day hospitalisation or death following LRTI. Among 186,094 LRTI episodes, 2.19% were classified as complicated. After adjusting for established clinical factors, SES was a strong predictor, whereas migration background was not. Patients in the lowest SES category had an adjusted odds ratio of 1.46 (95%CI: 1.31 - 1.62) for a complicated course compared to the highest. The incorporation of SES into clinical decision tools and guidelines has the potential to enhance risk-stratification of patients with LRTI in daily practice of primary care, thereby supporting more equitable care.
Gupta, R. K.; Harrison, E. M.; Ho, A.; Docherty, A. B.; Knight, S. R.; van Smeden, M.; Abubakar, I.; Lipman, M.; Quartagno, M.; Pius, R. B.; Buchan, I.; Carson, G.; Drake, T. M.; Dunning, J.; Fairfield, C. J.; Gamble, C.; Green, C. A.; Halpin, S.; Hardwick, H.; Holden, K.; Horby, P.; Jackson, C.; McLean, K.; Merson, L.; Nguyen-Van-Tam, J. S.; Norman, L.; Olliaro, P. L.; Pritchard, M. G.; Russell, C. D.; Scott-Brown, J.; Shaw, C. A.; Sheikh, A.; Solomon, T.; Sudlow, C. L.; Swann, O. V.; Turtle, L.; Openshaw, P. J.; Baillie, J. K.; Semple, M. G.; Noursadeghi, M.
Show abstract
Prognostic models to predict the risk of clinical deterioration in acute COVID-19 are required to inform clinical management decisions. Among 75,016 consecutive adults across England, Scotland and Wales prospectively recruited to the ISARIC Coronavirus Clinical Characterisation Consortium (ISARIC4C) study, we developed and validated a multivariable logistic regression model for in-hospital clinical deterioration (defined as any requirement of ventilatory support or critical care, or death) using 11 routinely measured variables. We used internal-external cross-validation to show consistent measures of discrimination, calibration and clinical utility across eight geographical regions. We further validated the final model in held-out data from 8,252 individuals in London, with similarly consistent performance (C-statistic 0.77 (95% CI 0.75 to 0.78); calibration-in-the-large 0.01 (-0.04 to 0.06); calibration slope 0.96 (0.90 to 1.02)). Importantly, this model demonstrated higher net benefit than using other candidate scores to inform decision-making. Our 4C Deterioration model thus demonstrates unprecedented clinical utility and generalisability to predict clinical deterioration among adults hospitalised with COVID-19.
Wang, J.; Tang, W.; Ma, X.; Yan, H. m.; Yuan, Y.
Show abstract
Large language models (LLMs) are increasingly used for automated quality control (QC) of radiology reports. However, the reliability of LLMs on reports in Mandarin, and the relative performance of domestic versus international flagship models, remain unknown. We benchmarked 14 LLM configurations, seven Chinese-developed ("domestic") and seven international models, on 1,000 whole-body 18F-FDG PET/CT reports split into an error-injected "junior-docto" arm and a low-residual "finalised" arm (500 each), using a controlled error-injection gold standard. Under each blinded zero-shot prompt, each model flagged six error types and assigned a 1-5 overall score. Two distinct abilities: error-detection macro-F1 (0.356-0.667) and overall-score calibration (ICC[2,1] 0.099-0.627), were weakly and not significantly correlated across models (Spearman {rho} = 0.38, p = 0.18); the dissociation was instead evident in sharp rank reversals, the strongest detector (Claude-Opus-4.8 0.667) calibrating poorly (0.491), while the three best-calibrated models were all domestic (MiMo 0.627, GLM-5 0.612, DeepSeek 0.609). Once the access channel was controlled, domestic and international error detection were statistically indistinguishable ({Delta}macro-F1= -0.011, P = 0.84); domestic models showed consistent but not significant advantages in calibration ({Delta}ICC = +0.142) and Chinese-character-error detection ({Delta}F1 = +0.109), accompanied with large reductions in cost (US$0.09-2.71 vs $0.26-14.5 per 1,000 reports) and on-premise deployability. Re-running two flagships through both agent channels and clean APIs showed that agent channel inflated both detection and calibration (GPT-5.5 {Delta}ICC = +0.098, 95% CI 0.070-0.128), confirming that uncontrolled benchmarks over-credit agent-channel models. Missed-diagnosis detection was the universal weakness (best 0.467) and the one category where the human physicians outperformed every model. Raw detection ability does not guarantee a trustworthy score, and domestic and international models differ by deployment-relevant profile rather than by overall performance rank; both essential distinctions for performing clinical nuclear-medicine QC.
Tan, J.; Tang, P. H.
Show abstract
BackgroundPa4ediatric pneumonia is a major cause of childhood morbidity and mortality. Chest X-rays (CXR) are central to diagnosis, but shortages of specialist radiologists can delay reporting. Multimodal large language models (MLLMs) may assist clinical workflows by analysing images and communicating findings, although their diagnostic performance remains below state-of-the-art classifiers. ObjectiveTo evaluate whether ensemble strategies improve MLLM diagnostic performance for paediatric radiological pneumonia detection on CXRs. MethodsIn this retrospective study, paediatric CXRs from two datasets (balanced and real-world) at KK Womens and Childrens Hospital were analysed. Images were independently reviewed by two board-certified radiologists, with pneumonia severity assigned to three classes using a predefined consensus algorithm. Fifteen MedGemma-4B-it agents classified each CXR into five likelihood categories, which were mapped to the three severity classes for evaluation. Majority voting, soft voting and GPTOSS-20B aggregation were compared with baseline average agent performance. The primary outcome was One-vs-Rest (OvR) AUROC. Secondary metrics included accuracy, sensitivity, specificity, F1-score, Cohens {kappa} and One-vs-One (OvO) AUROC. ResultsThe balanced dataset contained 900 CXRs and the real-world dataset 1300 CXRs. Soft voting significantly improved OvR-AUROC compared with baseline in both datasets (Balanced: 0.829>0.764; 95%CI=0.752-0.779; P=0.0002. Real-world: 0.728>0.655; 95%CI=0.638-0.679; P=0.0003). Soft voting also improved accuracy, Cohens {kappa}, OvO-AUROC in both datasets and F1-score in the balanced dataset. ConclusionSoft voting enhances MedGemmas diagnostic discriminatory performance for paediatric radiological pneumonia detection. Our system enables privacy-preserving, near real-time clinical decision support with explainable outputs, having potential for integration into emergency departments. Our systems high specificity supports triage by flagging high-risk radiological pneumonia cases. Clinical ImpactO_LIPaediatric CXRs often face reporting delays exceeding 24 hours due to radiologist shortages. C_LIO_LIOur proposed MLLM ensemble framework achieves better than average MLLM diagnostic discrimination for radiological pneumonia without requiring cloud-based systems. C_LIO_LISoft-voting aggregation enhances diagnostic discriminatory effectiveness for paediatric pneumonia severity, while preserving explainable outputs. C_LIO_LIOur system acts as a decision support tool that identifies higher-risk pneumonia cases for urgent review, supporting safer triage. C_LI
ADETUNJI, S. A.
Show abstract
BackgroundAcute coronary syndrome at first contact in primary care must be recognised rapidly and safely--especially in women and in adults with diabetes, who more often present without classic chest pain and are at risk of under-triage. We synthesised evidence on symptom constellations and triage strategies--rapid electrocardiogram, high-sensitivity cardiac troponin pathways, and primary-care risk tools--that reduce missed or late recognition of acute coronary syndrome. MethodsTargeted review (Jan 1, 2007-May 31, 2025) of prospective cohorts, diagnostic or implementation studies, risk-tool derivations and validations, telephone-triage studies, and contemporary guidelines relevant to first-contact primary care (office, urgent or emergency primary care, and out-of-hours services). Prespecified outcomes were missed acute coronary syndrome, 30-day major adverse cardiac events, time to electrocardiogram, time to troponin and decision, emergency-department transfer or admission, length of stay, and diagnostic performance (sensitivity, specificity, negative and positive predictive value). Risk of bias was assessed with Quality Assessment of Diagnostic Accuracy Studies-2, Newcastle-Ottawa Scale, and Risk of Bias 2, and certainty of evidence with Grading of Recommendations, Assessment, Development and Evaluation. FindingsAcross 18 sources, symptom clusters alone were insufficient to safely rule out acute coronary syndrome; history-based rules showed heterogeneous sensitivity and are best used to structure history and trigger testing. Rapid electrocardiogram (with repeats when concern persists) and high-sensitivity cardiac troponin algorithms provided the greatest diagnostic safety. Ambulatory implementations of assay-specific strategies achieved very high negative predictive values (approximately 99-100%) for rule-out in low-risk populations and accelerated disposition; an emergency primary-care deployment of the European Society of Cardiology 0/1-hour high-sensitivity troponin pathway ruled out about 64% at 1 hour, yielded a conclusive decision for about 77% by 4 hours, and reduced length of stay by roughly 2.2 hours. Safety margins were lower in patients with known coronary artery disease (negative predictive value about 96-98%) and the rule-out fraction was smaller, supporting more conservative thresholds, brief observation, or adjunctive clinical risk scoring. Telephone and out-of-hours case-control data linked non-retrosternal descriptors and system factors to missed acute coronary syndrome, arguing for up-triage to in-person electrocardiogram and troponin testing or emergency medical services transfer. InterpretationFor adults first assessed in primary care--particularly women and people with diabetes--a protocolised pathway (electrocardiogram within minutes, serial high-sensitivity troponin using assay-specific delta thresholds, and clear escalation) offers high rule-out safety with faster, more equitable care. Known coronary artery disease is a caution zone, warranting stricter discharge thresholds or extended observation. Health systems should invest in clinic electrocardiogram access, point-of-care high-sensitivity troponin where laboratory turnaround is slow, emergency-medical-services-first scripts, and routine audit (time to electrocardiogram and troponin, rule-out at 1 and 4 hours, emergency-department transfer, and 30-day major adverse cardiac events) stratified by sex, diabetes status, and coronary artery disease. RESEARCH IN CONTEXTO_ST_ABSEvidence before this studyC_ST_ABSWe searched MEDLINE (via PubMed), Embase, and Web of Science from Jan 1, 2007, to May 31, 2025, using terms for acute coronary syndrome (ACS)/chest pain, primary care/general practice/out-of-hours services, women/sex differences, diabetes, high-sensitivity cardiac troponin (hs-cTn; 0/1-hour and very-low strategies), telephone triage, and primary-care risk scores (Marburg Heart Score, INTERCHEST, HEART/HEAR). We hand-searched guideline repositories and reference lists. Three signals consistently emerged: (1) guidelines endorse rapid electrocardiogram (ECG) and hs-cTn-based pathways; (2) symptom clusters alone--particularly in women and people with diabetes--are insufficient to rule out ACS; and (3) primary-care data are sparse and heterogeneous, with safety concerns in known coronary artery disease (CAD) when accelerated rule-out algorithms are used. Few studies linked findings to operational metrics important at the primary-care front door (time-to-ECG/hs-cTn, rule-out proportions, length of stay [LOS]). Added value of this studyWe integrate first-contact primary-care evidence across 18 sources into a single, implementable pathway centered on rapid ECG, assay-specific hs-cTn strategies (0/1-hour or single very-low when symptom onset [≥]3 h), and explicit escalation thresholds. We foreground women and adults with diabetes, quantify rule-out safety including the known-CAD caveat, and pair clinical accuracy with operational outcomes (time-to-test, rule-out at 1/4 h, emergency department transfers, LOS). We provide a concise summary table and figure plus an audit set that clinicians can deploy immediately, alongside formal risk-of-bias and GRADE certainty judgments. Implications of all the available evidencePrimary care can safely accelerate ACS evaluation by pairing ECG within 10 minutes and hs-cTn algorithms with conservative thresholds in known CAD and proactive escalation for non-chest presentations common in women and people with diabetes (e.g., dyspnea, epigastric discomfort, unusual fatigue). Health systems should invest in clinic ECG access, point-of-care hs-cTn where laboratory turnaround is slow, emergency medical services (EMS)-first triage scripts, and routine audit of time-to-ECG/hs-cTn, rule-out proportions, emergency department transfers, and 30-day major adverse cardiac events (MACE) stratified by sex, diabetes, and CAD. Research priorities include prospective, consecutive primary-care cohorts with adjudicated 30-day outcomes, pragmatic trials of point-of-care hs-cTn, decision support tailored to known CAD, and cluster-randomized improvements to telephone/virtual triage.
Amiruddin, N.; Mellor, S.; Crisp, R.; Nair, A.; Patel, M.
Show abstract
Background Ventilator-associated pneumonia (VAP) is the most frequent nosocomial infection in critical care, affecting 20-36% of mechanically ventilated patients. Early prediction is hampered by the absence of a reliable, objective diagnostic standard. We developed ADVISE (Automated Dudley Ventilation Infection Series Evaluation), a machine learning model to predict physiological deterioration consistent with developing VAP using routinely collected electronic health record data from a UK NHS intensive care unit. Methods Retrospective observational study of admissions at Russell's Hall Hospital ICU (2008-2026). Following National Data Opt-Out exclusion (158 admissions, 4.2%), 3,566 admissions generated 33,208 candidate 48-hour observation blocks. Six temporal variables - FiO2, ventilator mode, P:F ratio, procalcitonin (PCT), secretion amount, and secretion description - were extracted across the baseline window (hours 1-24). A composite VAP-surrogate outcome required concurrent P:F ratio decline (>=5%) and PCT rise (>=0.5 ng/mL) across the outcome window (hours 25-48). After sequential quality filters, 2,134 blocks (18 positive, 0.84% prevalence) were retained. An XGBoost classifier was trained using nested 5-fold cross-validation with scale_pos_weight=114.0 and ROC-based hyperparameter optimisation on 1,495 training blocks, evaluated on 639 held-out test blocks. Performance was assessed via AUROC, AUPRC, and calibration (Brier score). Bootstrap resampling (1,000 iterations) generated 95% confidence intervals. Results On the held-out test set (n=639, 5 positive outcomes), ADVISE achieved AUROC 0.874 [95% CI: 0.771-0.939] and AUPRC 0.031 [0.008-0.069], representing a 4.0-fold improvement over the no-skill baseline. Nested cross-validation mean AUROC was 0.844 +/- 0.078 (range 0.716-0.915). At the Youden-optimal threshold, sensitivity was 0% with specificity 97.8%, reflecting extreme class imbalance (0.78% test prevalence). A threshold targeting 80% sensitivity achieved sensitivity 80.0% [33.3-100.0%], specificity 87.4% [84.8-89.9%], positive predictive value 4.8% [1.1-9.9%], and negative predictive value 99.8% [99.4-100.0%], detecting 4 of 5 VAP cases with approximately 80 false alarms (12.6% false positive rate). Brier score was 0.0078. Feature importance identified baseline P:F ratio as the dominant predictor (41.3% total gain), followed by ventilator mode (26.1%), secretion amount (13.2%), secretion description (9.1%), procalcitonin (5.9%), and FiO2; (4.5%). Conclusions ADVISE demonstrates that baseline oxygenation trajectory and ventilatory support patterns - derived exclusively from routinely charted ICCA variables - can identify admissions at risk of VAP-related physiological deterioration with meaningful discrimination (AUROC 0.874) despite severe class imbalance. The 80% sensitivity operating point offers a clinically actionable alert rate (12.6% FPR), supporting integration into existing ICU workflows. This proof-of-concept study establishes feasibility; multi-site prospective validation is required before clinical deployment.
Aghlmandi, S.; Shafiezadeh, S.; Huber, C.; Godet, P.; Bucher, H. C.; Bielicki, J. A.
Show abstract
ObjectivesTo evaluate whether machine learning (ML) applied to comprehensive claims data without diagnostic codes can distinguish a high proportion of antibiotic treatment episodes as urinary tract infection (UTI) or non-UTI cases. Such approaches may be valuable for antimicrobial stewardship when diagnosis-linked datasets are unavailable. MethodsOutpatient antibiotic prescription claims from three major Swiss insurers (2017-2020; [~]40% of the Swiss population) were analyzed. Based on clinical input, specific constellations of claims codes (e.g. positive urine culture plus typical antibiotic) were a priori assigned as indicating UTI episodes, providing the reference classification. Predictors included sex, age group, comorbidity, and diagnostic tests ordered during the episode. Four ML classifiers were tested; performance and interpretability were evaluated, with XGBoost prioritized. ResultsAfter cleaning and balancing, 38,982 records (19,491 UTI; 19,491 non-UTI) were included. XGBoost achieved an AUC of 0.94, accuracy of 87.6%, sensitivity of 79.2%, and specificity of 96.1%. Misclassification was asymmetric: 11% of non-UTI cases were labeled UTI, while 2% of UTI cases were misclassified as non-UTI. Diagnostics ordered were the strongest predictors, followed by female sex and older age. ConclusionsEven in the absence of diagnosis codes, ML applied to claims data can reliably identify UTI-related prescriptions. This supports the feasibility of claims-based surveillance tools for stewardship, while in parallel highlighting the need for scalable, low-burden approaches to improve direct diagnostic coding in routine data.
Xu, Y.; Teutsch, B.; Zeng, W.; Hu, Y.; Rastogi, S.; Hu, E. Y.; DeGregorio, I. M.; Fung, C. W.; Richter, B. I.; Cummings, R.; Goldberg, J. E.; Mathieu, E.; Appiah Asare, B.; Hegedus, P.; Gurza, K.-B.; Szabo, I. V.; Tarjan, H.; Szentesi, A.; Borbely, R.; Molnar, D.; Faluhelyi, N.; Vincze, A.; Marta, K.; Hegyi, P.; Lei, Q.; Gonda, T.; Huang, C.; Shen, Y.
Show abstract
Background and aimsAcute pancreatitis (AP) is a common gastrointestinal disease with rising global incidence. While most cases are mild, severe AP (SAP) carries high mortality. Early and accurate severity prediction is crucial for optimal management. However, existing severity prediction models, such as BISAP and mCTSI, have modest accuracy and often rely on data unavailable at admission. This study proposes a deep learning (DL) model to predict AP severity using abdominal contrast-enhanced CT (CECT) scans acquired within 24 hours of admission. MethodsWe collected 10,130 studies from 8,335 patients across a multi-site U.S. health system. The model was trained in two stages: (1) self-supervised pretraining on large-scale unlabeled CT studies and (2) fine-tuning on 550 labeled studies. Performance was evaluated against mCTSI and BISAP on a hold-out internal test set (n=100 patients) and externally validated on a Hungarian AP registry (n=518 patients). ResultsOn the internal test set, the model achieved AUROCs of 0.888 (95% CI: 0.800-0.960) for SAP and 0.888 (95% CI: 0.819-0.946) for mild AP (MAP), outperforming mCTSI (p = 0.002). External validation showed robust AUROCs of 0.887 (95% CI: 0.825-0.941) for SAP and 0.858 (95% CI: 0.826-0.888) for MAP, surpassing mCTSI (p = 0.024) and BISAP (p = 0.002). Retrospective simulation suggested the models potential to support admission triage and serve as a second reader during CECT interpretation. ConclusionsThe proposed DL model outperformed standard scoring systems for AP severity prediction, generalized well to external data, and shows promise for providing early clinical decision support and improving resource allocation.
Reitsam, N. G.; Gustav, M.; Jesinghaus, M.; Maerkl, B.; Foersch, S.; Kather, J. N.
Show abstract
Large language models (LLMs) are evolving into diagnostic co-pilots, yet current benchmarks fail to test the integrated, stepwise reasoning required in diagnostic pathology. Here, we present Pathologys Last Exam (PLE), a curated, highly detailed, text-based benchmark of 100 complex cases spanning organ systems, enriched for rare/challenging entities, plus 20 adversarial cases designed to stress-test model safety. Each case provides structured blocks (Primary, Clinical, Histopathology, IHC/Special Stains, Molecular Pathology) with stepwise information release mirroring real sign-out. We evaluated five LLMs (one proprietary, four open-source) across different stages. While the best model (GPT-5) achieved 70% accuracy on full evidence, performance on safety tests was alarming. Models frequently failed to detect biological contradictions, confidently diagnosing nonsensical "mix-up" cases rather than refusing them. This reveals a critical safety gap: high diagnostic capability is currently coupled with a dangerous inability to recognize impossible clinical scenarios. PLE provides a framework to measure and mitigate these risks before clinical deployment, as well as a foundation for developing multimodal evaluation protocols that can be extended to vision-language models and autonomous diagnostic agents in the future.
Sparnon, E.; Stevens, K.; Song, E.; Harris, R. J.; Strong, B. W.; Bruno, M. A.; Baird, G. L.
Show abstract
The present study evaluates the real-world clinical predictive performance of FDA-authorized artificial intelligence (AI) devices used in radiology, focusing on the false positive paradox (FPP) and its implications for clinical practice. To do this, we analyzed publicly available FDA data on AI radiology devices from 2024 and 2025 from 510(k) summaries, demonstrating how diagnostic accuracy metrics like sensitivity and specificity do not necessarily translate into high positive predictive value (PPV) due to the influence of target disease prevalence. We show the importance of disclosing the false discovery (FDR) and false omission rates (FOR) and argue that this transparency enables clinicians to select AI systems that balance false positive and false negative costs in a clinically, ethically, and financially appropriate manner. Finally, we provide recommendations for what data should be provided to best serve practices and radiologists.
Zhang, H.; Shi, T.; Wu, X.; Zhang, X.; Wang, K.; Bean, D.; Dobson, R.; Teo, J. T.; Sun, J.; Zhao, P.; Li, C.; Dhaliwal, K.; Wu, H.; Li, Q.; Guthrie, B.
Show abstract
BackgroundAccurate risk prediction of clinical outcome would usefully inform clinical decisions and intervention targeting in COVID-19. The aim of this study was to derive and validate risk prediction models for poor outcome and death in adult inpatients with COVID-19. MethodsModel derivation using data from Wuhan, China used logistic regression with death and poor outcome (death or severe disease) as outcomes. Predictors were demographic, comorbidity, symptom and laboratory test variables. The best performing models were externally validated in data from London, UK. Findings4.3% of the derivation cohort (n=775) died and 9.7% had a poor outcome, compared to 34.1% and 42.9% of the validation cohort (n=226). In derivation, prediction models based on age, sex, neutrophil count, lymphocyte count, platelet count, C-reactive protein and creatinine had excellent discrimination (death c-index=0.91, poor outcome c-index=0.88), with good-to-excellent calibration. Using two cut-offs to define low, high and very-high risk groups, derivation patients were stratified in groups with observed death rates of 0.34%, 15.0% and 28.3% and poor outcome rates 0.63%, 8.9% and 58.5%. External validation discrimination was good (c-index death=0.74, poor outcome=0.72) as was calibration. However, observed rates of death were 16.5%, 42.9% and 58.4% and poor outcome 26.3%, 28.4% and 64.8% in predicted low, high and very-high risk groups. InterpretationOur prediction model using demography and routinely-available laboratory tests performed very well in internal validation in the lower-risk derivation population, but less well in the much higher-risk external validation population. Further external validation is needed. Collaboration to create larger derivation datasets, and to rapidly externally validate all proposed prediction models in a range of populations is needed, before routine implementation of any risk prediction tool in clinical care. FundingMRC, Wellcome Trust, HDR-UK, LifeArc, participating hospitals, NNSFC, National Key R&D Program, Pudong Health and Family Planning Commission Research in contextO_ST_ABSEvidence before this studyC_ST_ABSSeveral prognostic models for predicting mortality risk, progression to severe disease, or length of hospital stay in COVID-19 have been published.1 Commonly reported predictors of severe prognosis in patients with COVID-19 include age, sex, computed tomography scan features, C-reactive protein (CRP), lactic dehydrogenase, and lymphocyte count. Symptoms (notably dyspnoea) and comorbidities (e.g. chronic lung disease, cardiovascular disease and hypertension) are also reported to have associations with poor prognosis.2 However, most studies have not described the study population or intended use of prediction models, and external validation is rare and to date done using datasets originating from different Wuhan hospitals.3 Given different patterns of testing and organisation of healthcare pathways, external validation in datasets from other countries is required. Added value of this studyThis study used data from Wuhan, China to derive and internally validate multivariable models to predict poor outcome and death in COVID-19 patients after hospital admission, with external validation using data from Kings College Hospital, London, UK. Mortality and poor outcome occurred in 4.3% and 9.7% of patients in Wuhan, compared to 34.1% and 42.9% of patients in London. Models based on age, sex and simple routinely available laboratory tests (lymphocyte count, neutrophil count, platelet count, CRP and creatinine) had good discrimination and calibration in internal validation, but performed only moderately well in external validation. Models based on age, sex, symptoms and comorbidity were adequate in internal validation for poor outcome (ICU admission or death) but had poor performance for death alone. Implications of all the available evidenceThis study and others find that relatively simple risk prediction models using demographic, clinical and laboratory data perform well in internal validation but at best moderately in external validation, either because derivation and external validation populations are small (Xie et al3) and/or because they vary greatly in casemix and severity (our study). There are three decision points where risk prediction may be most useful: (1) deciding who to test; (2) deciding which patients in the community are at high-risk of poor outcomes; and (3) identifying patients at high-risk at the point of hospital admission. Larger studies focusing on particular decision points, with rapid external validation in multiple datasets are needed. A key gap is risk prediction tools for use in community triage (decisions to admit, or to keep at home with varying intensities of follow-up including telemonitoring) or in low income settings where laboratory tests may not be routinely available at the point of decision-making. This requires systematic data collection in community and low-income settings to derive and evaluate appropriate models.
Jia, E.; Omar, M.; Barash, Y.; Brook, O. R.; Ahmed, M.; Kruskal, J. B.; Gorenshtein, A.; Klang, E.
Show abstract
Ramaswamy et al. recently reported in Nature Medicine that ChatGPT Health, a consumer-facing health AI tool, undertriaged 51.6% of true emergencies. It was also susceptible to social anchoring in a structured stress test of triage recommendations. We applied the same vignette-based benchmark to OpenEvidence, a widely used physician-facing AI platform for clinical decision support. The benchmark included 960 prompts across 21 clinical domains (Supplementary Table S3). OpenEvidence undertriaged 12.5% of emergencies, a four-fold reduction relative to ChatGPT Health. It also showed no anchoring effect. Its errors skewed in a safer direction, including 68.0% overtriage of Home presentations. In 65 of 960 responses (6.8%), it declined to assign a triage level. These refusals occurred only in symptom-only prompts and never in urgent or emergency cases. Performance improved when objective clinical data were provided. Under the same benchmark, a widely used physician-facing system showed a different safety profile from a consumer-facing one. This suggests that who a health AI is built for can shape how it fails.
Isaksen, A. A.; Schaarup, J. R.; Bjerg, L.; Hulman, A.
Show abstract
BackgroundArtificial intelligence (AI) is expected to become an integral part of healthcare services, and the widespread adoption of AI tools in all areas of life is making AI accessible to the general public. Public perception of the benefits and risks of AI in healthcare is key to large-scale acceptance and implementation, and is increasingly influenced by first-hand experiences of AI. The aim of this study was to assess how exposure to ChatGPT changed public perception of AI in healthcare. MethodsWe used baseline and follow-up data from 5,899 survey participants, who reported their perception of AI in 2022 and 2024, and ChatGPT use in 2024. Administrative and healthcare data from nationwide Danish registers was used for weighting and adjustment. Multinomial multivariate logistic regression was used to model how exposure to ChatGPT use affected changes in perception of AI. ResultsAt baseline (before ChatGPTs launch) 2,236 individuals (37%) were unsure of the benefits and risks of AI in healthcare, 2,384 (40%) perceived net benefits, 1,083 (18%) perceived benefits and risks as equal, and 196 (3.3%) perceived net risks. At follow-up, 1,195 individuals (20%) had been exposed to ChatGPT use, which was associated with higher odds of changing perception of AI to benefits (OR 3.21 [95% CI: 2.34-4.40]) among individuals who were unsure at baseline, and lower odds of changing to uncertainty from more defined baseline perceptions (from benefits (OR 0.32 [0.24-0.42]), equal (OR 0.47 [0.32-0.69]) and risks (OR 0.27 [0.08-0.98])). ConclusionExposure to ChatGPT was associated with a change towards positive perception of benefits and risks of AI in healthcare among individuals who were uncertain prior to exposure, and individuals with more defined perceptions of AI were less likely to become uncertain after exposure to ChatGPT.
Samuels, T. H.; Forrest-Hammond, R.; Stockford, C.; Harris, S. K.; Eyre, D. W.; Gupta, R. K.; Noursadeghi, M.
Show abstract
Background: Bacteraemia is associated with poor outcomes but the diagnostic gold standard, peripheral blood culture, takes up to 24 hours to become clinically actionable, hampering early management decisions in suspected infection. Single predictors and existing sepsis risk scores discriminate poorly, and few multivariable bacteraemia models have been adequately validated in UK populations. Methods: We developed a logistic regression model, using backwards AIC based selection of predefined candidate predictors routinely available within hours of hospital attendance, in a retrospective cohort of 33,874 hospital encounters at University College London Hospitals (UCLH) between 2019 and 2024. Continuous predictors were modelled using restricted cubic splines and missing data handled using multiple imputation. Model performance was assessed via internal external cross validation and prediction instability analysis, before temporal validation in held-out 2024 UCLH data and external validation in 53,669 hospital encounters from the Infections in Oxfordshire Research Database (IORD). Results: Bacteraemia occurred in 5.2% of UCLH and 8.9% of IORD encounters, respectively. Twenty predictors were retained, spanning demographics, comorbidities, vital signs and blood tests. Discrimination was stable across development time periods (pooled c-statistic 0.82, 95%CI 0.81 to 0.84) and was maintained in temporal (0.83, 0.79 to 0.87) and external validation (0.83, 0.82 to 0.83), with excellent calibration in external validation (calibration slope 1.08 (1.05 to 1.11); calibration-in-the-large 0.01 (-0.02 to 0.04)). The model outperformed single predictors, established risk scores, and a reconstructed comparator model, and showed superior net benefit in decision curve analysis. Performance was consistent across age, sex, ethnicity and socioeconomic subgroups but degraded when blood cultures were sampled more than six hours after attendance and varied by likely infection site. Conclusions: This model accurately predicts bacteraemia using routinely collected data available within hours of hospital attendance, with performance maintained in a large, independent external validation cohort. It offers a generalisable, clinically interpretable tool to support early decision-making in suspected infection, pending further work to establish optimal implementation thresholds.
Graham, S.; Minhas, F.; Bilal, M.; Ali, M.; Tsang, Y. W.; Eastwood, M.; Wahab, N.; Jahanifar, M.; Hero, E.; Dodd, K.; Sahota, H.; Wu, S.; Lu, W.; Azam, A.; Benes, K.; Nimir, M.; Hewitt, K.; Bhalerao, A.; Robinson, A.; Eldaly, H.; E Ahmed Raza, S.; Gopalakrishnan, K.; Snead, D.; Rajpoot, N.
Show abstract
ObjectivesDevelop an interpretable AI algorithm to rule out normal large bowel endoscopic biopsies saving pathologist resources. DesignRetrospective study. SettingOne UK NHS site was used for model training and internal validation. External validation conducted on data from two other NHS sites and one site in Portugal. Participants6,591 whole-slides images of endoscopic large bowel biopsies from 3,291 patients (54% Female, 46% Male). Main outcome measuresArea under the receiver operating characteristic and precision recall curves (AUC-ROC and AUC-PR), measuring agreement between consensus pathologist diagnosis and AI generated classification of normal versus abnormal biopsies. ResultsA graph neural network was developed incorporating pathologist domain knowledge to classify the biopsies as normal or abnormal using clinically driven interpretable features. Model training and internal validation were performed on 5,054 whole slide images of 2,080 patients from a single NHS site resulting in an AUC-ROC of 0.98 (SD=0.004) and AUC-PR of 0.98 (SD=0.003). The predictive performance of the model was consistent in testing over 1,537 whole slide images of 1,211 patients from three independent external datasets with mean AUC-ROC = 0.97 (SD=0.007) and AUC-PR = 0.97 (SD=0.005). Our analysis shows that at a high sensitivity threshold of 99%, the proposed model can, on average, reduce the number of normal slides to be reviewed by a pathologist by 55%. A key advantage of IGUANA is its ability to provide an explainable output highlighting potential abnormalities in a whole slide image as a heatmap overlay in addition to numerical values associating model prediction with various histological features. Example results with can be viewed online at https://iguana.dcs.warwick.ac.uk/. ConclusionsAn interpretable AI model was developed to screen abnormal cases for review by pathologists. The model achieved consistently high predictive accuracy on independent cohorts showing its potential in optimising increasingly scarce pathologist resources and for achieving faster time to diagnosis. Explainable predictions of IGUANA can guide pathologists in their diagnostic decision making and help boost their confidence in the algorithm, paving the way for future clinical adoption. What is already known on this topicO_LIIncreasing screening rates for early detection of colon cancer are placing significant pressure on already understaffed and overloaded histopathology resources worldwide and especially in the United Kingdom1. C_LIO_LIApproximately a third of endoscopic colon biopsies are reported as normal and therefore require minimal intervention, yet the biopsy results can take up to 2-3 weeks2. C_LIO_LIAI models hold great promise for reducing the burden of diagnostics for cancer screening but require incorporation of pathologist domain knowledge and explainability. C_LI What this study addsO_LIThis study presents the first AI algorithm for rule out of normal from abnormal large bowel endoscopic biopsies with high accuracy across different patient populations. C_LIO_LIFor colon biopsies predicted as abnormal, the model can highlight diagnostically important biopsy regions and provide a list of clinically meaningful features of those regions such as glandular architecture, inflammatory cell density and spatial relationships between inflammatory cells, glandular structures and the epithelium. C_LIO_LIThe proposed tool can both screen out normal biopsies and act as a decision support tool for abnormal biopsies, therefore offering a significant reduction in the pathologist workload and faster turnaround times. C_LI
Hjärtström, M.; Didriksson, I.; Spangfors, M.; Friberg, H.; Jakobsson, A.; Frigyesi, A.
Show abstract
Mortality among patients admitted to intensive care with coronavirus disease 2019 (COVID-19) remains substantial despite advances in management. The contribution of pre-admission medication profiles to long-term survival is poorly defined. We analysed 497 adults with confirmed COVID-19 admitted to six intensive care units in southern Sweden between May 2020 and May 2021. Clinical and laboratory data were combined with prescription information from the national drug registry; drugs dispensed at least twice within eight months before admission were classified by Anatomical Therapeutic Chemical code. Polypharmacy was defined as the use of five or more medications. An XGBoost survival model with a Cox partial-likelihood objective was trained to predict one-year mortality and interpreted using SHapley Additive exPlanations (SHAP). The model achieved a concordance index of 0.74. Age was the strongest predictor of mortality, followed by the number of medications per patient, which ranked above the Charlson Comorbidity Index and Clinical Frailty Scale. Proton pump inhibitors were the only individual drug class among the top predictors, showing a modest positive association with mortality, whereas angiotensin-converting enzyme inhibitors and angiotensin II receptor blockers had negligible contributions. These findings identify cumulative medication burden as an independent and clinically relevant marker of vulnerability in critical COVID-19.
Callender, T.; Imrie, F.; Cebere, B.; Pashayan, N.; Navani, N.; van der Schaar, M.; Janes, S. M.
Show abstract
BackgroundEnsemble machine learning could support the development of highly parsimonious prediction models that maintain the performance of more complex models whilst maximising simplicity and generalisability, supporting the widespread adoption of personalised screening. In this work, we aimed to develop and validate ensemble machine learning models to determine eligibility for risk-based lung cancer screening. MethodsFor model development, we used data from 216,714 ever-smokers in the UK Biobank prospective cohort and 26,616 high-risk ever-smokers in the control arm of the US National Lung Screening randomised controlled trial. We externally validated our models amongst the 49,593 participants in the chest radiography arm and amongst all 80,659 ever-smoking participants in the US Prostate, Lung, Colorectal and Ovarian Screening Trial (PLCO). Models were developed to predict the risk of two outcomes within five years from baseline: diagnosis of lung cancer, and death from lung cancer. We assessed model discrimination (area under the receiver operating curve, AUC), calibration (calibration curves and expected/observed ratio), overall performance (Brier scores), and net benefit with decision curve analysis. ResultsModels predicting lung cancer death (UCL-D) and incidence (UCL-I) using three variables - age, smoking duration, and pack-years - achieved or exceeded parity in discrimination, overall performance, and net benefit with comparators currently in use, despite requiring only one-quarter of the predictors. In external validation in the PLCO trial, UCL-D had an AUC of 0.803 (95% CI: 0.783-0.824) and was well calibrated with an expected/observed (E/O) ratio of 1.05 (95% CI: 0.95-1.19). UCL-I had an AUC of 0.787 (95% CI: 0.771-0.802), an E/O ratio of 1.0 (0.92-1.07). The sensitivity of UCL-D was 85.5% and UCL-I was 83.9%, at 5-year risk thresholds of 0.68% and 1.17%, respectively 7.9% and 6.2% higher than the USPSTF-2021 criteria at the same specificity. ConclusionsWe present parsimonious ensemble machine learning models to predict the risk of lung cancer in ever-smokers, demonstrating a novel approach that could simplify the implementation of risk-based lung cancer screening in multiple settings.
Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.
Show abstract
Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.