Healthcare
○ MDPI AG
Preprints posted in the last 7 days, ranked by how well they match Healthcare's content profile, based on 17 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Barzideh, A.; Devasahayam, A. J.; Marzolini, S.; Munce, S.; Sibley, K. M.; Inness, E. L.; Mansfield, A.
Show abstract
Background: Aerobic exercise is recommended during stroke rehabilitation to improve cardiorespiratory fitness and support recovery; however, participation rates remain low. While institutional and system-level barriers have been widely examined, less is known about how individual patient factors influence engagement in aerobic exercise during rehabilitation. Objectives: We aimed to determine whether depressive symptoms, apathy, self-efficacy and outcome expectations for exercise, perceived barriers, or past exercise history were associated with aerobic exercise participation in stroke rehabilitation. Methods: In this prospective cohort sub-study, adults admitted to in- or out-patient stroke rehabilitation at three urban hospitals completed validated questionnaires assessing depressive symptoms, apathy, exercise self-efficacy, outcome expectations for exercise, perceived barriers to being active, and premorbid exercise history. Participants were separated into two groups for analysis: those who completed aerobic exercise during rehabilitation and those who did not. Equivalence testing and between-group comparisons were performed. Results: Sixty-two participants were enrolled; 16 participated in aerobic exercise and 46 did not. Groups were not equivalent on any individual-level factors. Compared to non-participants, those who performed aerobic exercise had significantly higher depressive symptom scores (p=0.0025) and lower self-efficacy for exercise (p=0.0087). Non-participants demonstrated significantly higher apathy (p=0.0007). No significant differences were found for outcome expectations, perceived barriers, or exercise history. Conclusion: Depressive symptoms and lower self-efficacy did not impede aerobic exercise participation during rehabilitation. Increased apathy, however, was associated with non-participation. Findings highlight the need for individually tailored aerobic exercise prescriptions that consider motivational and affective factors to optimize engagement during stroke rehabilitation.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Masters, N. B.; Farrar, K. G.; Holler, E.; Lancaster, J. M.
Show abstract
Background: Vitamin K prophylaxis is universally recommended for newborns to prevent life threatening vitamin K deficiency bleeding. Although not on the immunization schedule, vitamin K prophylaxis is often coadministered with hepatitis B birth dose and erythromycin ophthalmic ointment, and rising hesitancy around vaccines/preventive care may spill over into vitamin K administration. Methods: We conducted a retrospective cohort study using Truveta electronic health record data with linked mother-child dyads. Live births to mothers aged 15-49 from January 1, 2019 through June 30, 2026 were included. Vitamin K administration was defined as documentation on the birth date or following day. Logistic regression assessed sociodemographic predictors of non-receipt, and interrupted time series analysis evaluated changes after January 2026. Results: Among 1,026,375 infants, 995,628 (96.97%) had documented vitamin K administration. Non-receipt increased from an average of 2.1% during 2019-2022 to 4.3% in 2025 and 6.1% in 2026, reaching 8.10% in June 2026. Older maternal age, non-Hispanic or Latino ethnicity, Medicaid or unknown insurance, and year of delivery were associated with greater odds of non-receipt. After January 2026, there was no immediate step change, but the odds of vitamin K receipt declined an additional 10% per month (OR: 0.90; 95% CI, 0.88-0.91). Conclusions: Vitamin K non-receipt increased over the study period and accelerated after January 2026. Because vitamin K recommendations were not changed by the January vaccine schedule, this association may reflect broader impacts to confidence in newborn preventive care. Future studies should examine causal mechanisms, parental decision-making, and associated clinical outcomes.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Tang, P.; Lu, M. W.-H.; Yeung, K.-T.; Guo, B. J.; Wei, K.-F. N.
Show abstract
Background Global labor migration from LMIC to higher-income destinations has expanded rapidly, placing increasing pressure on destination-country health. Existing research on cross-border migrant workers has focused largely on occupational health, general healthcare utilization, and disease-specific risks, while there is considerably less evidence on their sexual and reproductive health. This study contributes to this understudied field by examining the policy and health-system factors that shape the sexual and reproductive health services for migrant workers in Taiwan. Methods A qualitative study was conducted in Taiwan between November 2025 and August 2026. 22 stakeholders were purposively recruited from academia, healthcare, nongovernmental organizations, government, labor brokerage, and employers. Data were collected through semi-structured interviews and small focus groups. Interviews were conducted in Mandarin Chinese, transcribed verbatim, and translated into English. Data were analyzed using framework analysis combining deductive coding based on the AAAQ framework with inductive coding of implementation and contextual themes. Results Gaps were identified across all four AAAQ dimensions. Participants described limited migrant-responsive SRH programming; physical, financial, administrative, social, and information barriers; shortcomings in linguistic and cultural responsiveness; and weaknesses in interpretation, coordination, and continuity of care, despite generally favorable views of Taiwan's clinical quality. Conclusions Our findings show that broad insurance coverage and strong clinical capacity do not by themselves ensure the realization of migrant workers' SRHR. In Taiwan, rights were mediated through labor brokerage, gendered live-in work arrangements, and fragmented governance across health, labor, immigration, and social-welfare systems. Improving migrant SRHR therefore requires stronger implementation of existing protections, reduced dependence on informal intermediaries, and more integrated institutional responsibility for cross-sector migrant health needs.
Gao, C.; Zhang, Y.; He, X.; Yuan, M.; Mou, F.; Zhou, J.; Chen, H.; Wang, H.; Guo, W.; Wei, Y.; Zhang, Z.; Yin, T.; Zhang, C.; Lian, Z.; Zhu, B.; Liu, J.; Zhang, R.; Fu, G.; Onuma, Y.; Wang, D.; Serruys, P. W.; Yi, F.; Tao, L.
Show abstract
BACKGROUND The optimal antiplatelet regimen in patients with acute coronary syndrome (ACS) and multivessel disease undergoing drug-coated balloon (DCB) angioplasty remains unclear. METHODS This was a prespecified subgroup analysis of the REC-CAGEFREE II trial, which was conducted at 41 sites in China and randomized 1948 exclusively DCB-treated participants with ACS to stepwise dual antiplatelet therapy (DAPT) de-escalation or standard DAPT. The primary endpoint was net adverse clinical events (NACE; including all-cause death, stroke, myocardial infarction, revascularization, and BARC type 3 or 5 bleeding) at 12 months. Participants were stratified into multivessel and single-vessel subgroups according to angiographic characteristics. RESULTS Overall, 720/1948 (37.0%) patients had multivessel disease. The multivessel subgroup was associated with a significantly higher risk of NACE compared with the single-vessel subgroup (12.5% versus 6.7%, HR IPTW:1.84, 95%CI:1.35-2.51, P<0.001). No significant interaction was observed between vessel status (multivessel or single-vessel) and treatment allocation with respect to NACE (Pinteraction=0.542). In the multivessel subgroup, NACE occurred in 44/368 (12.1%) and 45/352 (12.9%) in the stepwise de-escalation and standard DAPT groups (HR IPTW:0.95, 95%CI:0.62-1.75, P=0.818), respectively. In the single-vessel subgroup, NACE occurred in 43/607 (7.1%) and 39/621 (6.3%) in the stepwise de-escalation and standard groups (HR IPTW:1.12, 95%CI:0.72-1.70, P=0.611), respectively. For the prespecified hierarchical secondary endpoint, win ratio analyses yielded more wins for stepwise de-escalation in both subgroups. CONCLUSIONS Among patients with ACS undergoing DCB-only angioplasty, those with multivessel disease were associated with a higher risk of NACE than those with single-vessel disease. Stepwise DAPT de-escalation and standard DAPT exhibited similar risk-benefit profiles in both subgroups.
Mannava, S.; Ramkumar, V.; Murthy, G.
Show abstract
Introduction Hearing loss (HL) affects over 1{middle dot}5 billion people globally and India shares a disproportionately high burden including Disabling Hearing Loss (DHL). HL affects an Individual socio-economically, but there are limited studies on the broader societal economic consequences of HL in India.Methods Using Cost-of-Illness (COI) approach, we studied the societal economic burden of HL in India. This study uses epidemiological and macroeconomic data and modelling to estimate the loss of Gross National Income (GNI) due to HL and DHL across three economic pathways. Uncertainty is evaluated using deterministic and Probabilistic Sensitivity Analyses (PSA).Results The model estimates that there are in India, 289 million and 85{middle dot}9 million people with HL and DHL respectively. Direct Loss of GNI and Indirect Loss of GNI (Caregiver burden) are estimated as INR 4,648{middle dot}4 billion (USD 55{middle dot}6 billion) and INR 3,268 billion (USD 39 billion) respectively. The Loss of GNI due to Low Education amongst those with HL is estimated as INR 1,041{middle dot}9 billion (USD 12{middle dot}45 billion).Discussion Economic burden of HL is presented across three pathways with Direct Loss of GNI due to DHL being the greatest. It also presents age stratified caregiver economic burden. The findings of the study help in estimating similar cost pathways, advocacy, and policy decisions towards reducing HL prevalence in India and LMICs. This study also highlights the need for India specific estimations related to the HL attributable low education, state-wise disaggregates, and prevalence studies. Funding This study has not received any funding.
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.
Show abstract
In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.
Leuenberger, L. M.; Shoman, Y.; Romero, F.; Sasaki, M.; Deligianni, X.; Goebel, N.; Mozun, R.; Bielicki, J. A.; Burckhardt, M.-A.; Saner, C.; Schwitzgebel, V.; Hauschild, M.; Righini Grunder, F.; Mueller, P.; Schlapbach, L. J.; Jenni, O.; Spycher, B. D.; Kuehni, C. E.; Belle, F. N.; SwissPedHealth consotrium,
Show abstract
BACKGROUND: We used anthropometric data from electronic health records (EHRs) of Swiss childrens hospitals to evaluate growth references and estimate centile curves. METHODS: We received EHRs extracted from seven Swiss childrens hospitals and analysed two samples: all children with a height, weight, body mass index (BMI), or head circumference recording, and a subsample restricted to children without diseases potentially affecting growth, weighted to represent the general population. We calculated mean z-scores based on the World Health Organization growth references adopted for Switzerland in 2011 (CH-WHO 2011) and current Swiss growth references (Swiss 2026). We estimated sex-specific centile curves in the subsample using generalised additive models for location, scale, and shape. RESULTS: We included 213,868 children with height, 448,002 with weight, 209,244 with BMI, and 67,397 with head circumference recordings. Mean z-scores in the all children sample were (CH-WHO 2011; Swiss 2026): height (0.10; -0.19), weight (0.16; -0.09), BMI (0.04; -0.07), head circumference (-0.28, -0.28); and in the subsample: height (0.34; 0.00), weight (0.27; 0.01), BMI (0.18; 0.05), and head circumference (0.04; 0.01). The 50th height, weight, BMI, and head circumference centiles of girls and boys in the subsample closely followed those of Swiss 2026, with slightly wider 3rd and 97th centiles in infancy and adolescence. CONCLUSION: Height, weight, BMI, and head circumference centiles aligned well with the Swiss 2026 growth references in Switzerland, demonstrating that hospital EHRs could contribute to future growth references.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
Chen, Y.; Puckett, H.; Clarot, G.; Hawkins, B.; Sharp, K.; Todd, D. A.; Lopez, A.; Bertollo, J. R.; Behar, H. E.; Zeithamova, D.; Xie, H.; Verbalis, A.; VanMeter, A. S.; Gaillard, W. D.; Kenworthy, L.; Vaidya, C. J.
Show abstract
Generalization is a key cognitive process that allows humans to flexibly apply prior knowledge to guide new behaviors. Difficulties with generalization and flexibility are observed across neurodevelopmental disorders, especially autism, limiting adaptive function and quality of life. Cognitive-behavioral treatment benefits some but not all autistic individuals. As treatment requires application of learned skills to everyday life, variability in generalization ability may limit intervention success in autism. While cognitive substrates of learning and generalization are well established, their potential for explaining clinical outcomes is not known. Here, we combined a category learning task with computational modelling to distinguish two learning strategies underlying generalization -- prototype abstraction vs. exemplar memorization -- and tested whether individual differences in these learning strategies predicted real-world intervention outcomes in autistic youth. Fifty-four participants completed the category learning task at two pre-intervention timepoints, and then completed Unstuck and On Target:14-22 intervention targeting flexible problem solving, goal setting, and planning. We found that participants who consistently relied on prototype abstraction (N=26) were subsequently more likely to benefit from the intervention, showing improvement in parent- and self-reported flexibility. These findings identify prototype abstraction as a clinically relevant cognitive capacity that may help explain individual differences in intervention response and support the tailoring of interventions. More broadly, they demonstrate the value of linking basic cognitive mechanisms to clinical outcomes and may inform strategies to enhance the effectiveness of cognitive-behavioral interventions for youth with developmental disabilities.
Sahputri, V.; Angeline, A.; Tenggono, E.
Show abstract
Perioperative safety checklists standardize critical actions, but reliable completion depends on the surrounding work system and team behavior. We conducted a prospective observational analytic study from April to May 2026 in the central surgical unit of a high-volume public teaching referral hospital in Indonesia to examine whether patient safety culture and teamwork were associated with directly observed perioperative safety compliance and whether teamwork mediated the culture-compliance relationship. Patient safety culture was measured with the Hospital Survey on Patient Safety Culture 2.0, teamwork with a 35-item TeamSTEPPS Teamwork Perceptions Questionnaire research adaptation, and compliance by direct role-based observation using a 45-item checklist derived from the AORN Comprehensive Surgical Checklist. Eighty of 92 recruited professionals contributed 240 person-operation observations across 50 operations. Overall compliance was 74.75%, with sign-out lowest at 70.68%. Patient safety culture was associated with teamwork ({beta} = 0.590; 95% CI 0.510-0.770) and directly with compliance ({beta} = 0.407; 95% CI 0.187-0.712). The teamwork-compliance coefficient was positive ({beta} = 0.285; p = 0.046), but the prespecified percentile 95% CI included zero (-0.045 to 0.517). The indirect effect through teamwork was not supported ({beta} = 0.168; p = 0.079). These findings support a system-level interpretation of perioperative safety and identify learning-oriented responses to error, situation monitoring, and sign-out fidelity as measurable targets for future improvement efforts.
Yang, T.; Wei, S.; Wang, Y.; Bai, D.
Show abstract
Background Mirror therapy (MT)-specifically paradigms using mirror visual feedback (MVF)-is widely used in neurorehabilitation; however, mechanistic implementations vary substantially in movement content, rhythmicity and attentional demands. This protocol describes an acute mechanistic, within-participant fNIRS screening study designed to compare three prespecified upper-limb mirror-therapy task paradigms and to quantify associated subjective experience after each condition in healthy adults during a single visit. Methods and analysis This is a single-centre, within-participant, randomised crossover study conducted at Wuhan Wuchang Hospital (Wuhan, China). Healthy adults aged 18-35 years will complete three task conditions once each in a counterbalanced order using a 3*3 Latin-square scheme: UMT1 (task-oriented rhythmic functional movement), UMT2 (open-ended free movement with auditory control), and UMT3 (non-functional rhythmic movement). fNIRS will be acquired using the NirSmart-6000A system during a standardised block design. The primary outcome is ROI-level HbO activation quantified as GLM-derived {beta} estimates within the prespecified primary ROIs (bilateral SM1/M1 and bilateral PMC). Secondary outcomes include ROI-level windowed {Delta}HbO (5-20 s post-onset relative to the immediately preceding rest; descriptive only), ROI-level {Delta}HbR, and post-condition subjective ratings (illusion, immersion, confusion and fatigue; 1-7 Likert). Condition effects will be analysed using linear mixed-effects models with fixed effects for condition and period and prespecified multiplicity-adjusted pairwise contrasts. Ethics and dissemination Ethics approval was obtained from the Ethics Committee of Wuchang Hospital Affiliated to Wuhan University of Science and Technology (Approval No.: 2025-112-01; approved on 2025-08-21). The study is expected to be minimal risk. Findings will be disseminated through publication of this protocol manuscript and subsequent results manuscripts and conference presentations. Trial registration number Chinese Clinical Trial Registry (ChiCTR2600116634). This study is conducted as a prespecified mechanistic sub-study under the overarching registered project.