CHEST
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match CHEST's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
van Eijk, J.; Schober, P.; van Schuppen, H.; ter Schure, J.
Show abstract
We present our Stage-1 Registered Report as a full clinical trial article with all methods in past tense and including mock results, table and figures for the primary analysis. To remind the reader that this Stage-1 article is written before data collection, we highlight in color that these mock results are only for illustrative purposes and will be replaced by the actual results in the Stage-2 Registered Report. Background In patients experiencing out-of-hospital cardiac arrest, optimization of oxygen delivery during cardiopulmonary resuscitation is a critical. Although both positive end-expiratory pressure (PEEP) and zero end-expiratory pressure (ZEEP) are employed during CPR, their respective impacts on clinically relevant outcomes is yet to be clearly established. Methods This investigator-initiated, pragmatic, registry-based, multicenter, triple-blind randomized controlled superiority trial evaluates whether applying 8 cm H2O PEEP during cardiopulmonary resuscitation improves outcomes compared with ZEEP in adults with non-traumatic, non-drowning out-of-hospital cardiac arrest. Pre-randomized CPR kits (1:1 PEEP vs. sham) were used by ambulance sites during manual ventilation throughout the resuscitation process. The primary analysis was conducted in the principal stratum of patients who received either a supraglottic airway or endotracheal tube. The primary outcome was neurological status at hospital discharge measured by a utility-weighted score on the modified Rankin Scale. Secondary outcomes included prehospital return of spontaneous circulation, 30-day survival, and 6-month quality of life. The primary safety outcome was clinically significant pneumothorax.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Alwakeel, M.; Zaveri, S.; Buck, E.; Rajagopal, S.; Verma, D.; Loriaux, D.; Henao, R.; Tapson, V. F.; Ortel, T. L.; Jones, W. S.; Martin, J. G.; Haines, K. L.; Freeman, N. L.; Wong, A.-K. I.
Show abstract
Background: The 2026 American Heart Association/American College of Cardiology (AHA/ACC) guidelines replaced the 2019 European Society of Cardiology (ESC) four-tier pulmonary embolism (PE) risk scheme with five clinical categories (A-E) and subcategories. These categories were set by expert consensus and have not been validated against outcomes. How patients are reclassified relative to ESC, or how the two systems compare prognostically, is unknown. Methods: We utilized three cohorts of patients with confirmed PE using structured electronic health record data, laboratory biomarkers, and large-language-model abstraction of radiology reports: Duke University Health System (n=12,992, drawn from 95,760 consecutive inpatient CT pulmonary angiography studies, 2014-2025, with no referral or registry enrollment step between imaging and cohort entry), INSPECT (Stanford; n=3,870), and MIMIC-IV (Beth Israel Deaconess; n=361). Patients were assigned AHA/ACC categories B through E, subcategorized where data allowed, and mapped to 2019 ESC risk strata. The primary outcome was 30-day mortality; discrimination was assessed with Harrell C-index. Results: Among 17,223 patients with confirmed PE, pooled 30-day mortality rose monotonically across categories: 1.5% (B), 8.9% (C), 15.5% (D), and 31.9% (E), with the ordering preserved in all three cohorts despite differing baseline mortality. Subcategory-level discrimination was reliable only at the high-acuity extreme (D2-E2); across subcategories C1 through D1, mortality did not order monotonically (9.2%, 10.8%, 8.1%, 10.9%), and adding subcategories to category C did not improve discrimination at Duke (C-index 0.699 vs 0.699). Category C patients lacking both echocardiography and biomarker testing (12.7% of category C) had mortality (10.4%) equal to or exceeding classified peers. Relative to ESC, the frameworks were concordant at the extremes, but 5.7%of ESC intermediate-risk patients were reclassified to category D, with modestly higher but non-significant 30-day mortality than those remaining in category C (10.8% versus 8.9%). Conclusions: Across a three-health-system cohort, the 2026 AHA/ACC framework produced a reproducible mortality gradient at the category level, with added subcategory granularity refining risk chiefly at the highest-acuity tiers. Discrimination across the broad intermediate band was limited, and reclassification from ESC fell almost entirely within this range.
Dyer, B. P.; Deery, M.; Heyman, R.; Robinson, P.; Wainwright, C.; Sly, P.; Ware, R.; Blake, T.
Show abstract
Background Elexacaftor-tezacaftor-ivacaftor (ETI) has been demonstrated to improve lung function in clinical trials; however, evidence describing effects on trajectories and whether long-term improvements are sustained (>1-year) is lacking. We estimated within-person lung clearance index (LCI) trajectories before and after ETI initiation, assessing changes in level and rate of change, alongside acute LCI change, up to three years after ETI initiation. Methods Prospective observational study of children at a tertiary hospital. Children aged 3-17 years with [≥]2 LCI testing occasions (i) before and (ii) after starting ETI were used to describe lung function trajectories. Children with [≥]1 pre-ETI and [≥]1 post-ETI LCI occasion(s) were used to describe acute LCI change after ETI initiation. Age-adjusted LCI trajectories for time periods (i) before and (ii) after ETI initiation were estimated using linear mixed-effects models, and pre- and post-ETI LCIs were compared using paired Wilcoxon tests. Results Mean pre-ETI and post-ETI longitudinal changes in LCI were -0.007 (95% CI: -0.28, 0.27; n=35) and 0.12 (95% CI: -0.17, 0.41; n=20) turnovers per year, respectively. Before ETI initiation, 57% (30/53) of patients had an LCI[≥]7.1 turnovers (indicating impaired lung function), compared to 26% (14/53) post-ETI, with a median LCI difference of -0.70 (95% CI -0.84, -0.46; p<0.001) turnovers. Within-individual variability in LCI decreased post-ETI. Conclusions Our real-world data within a unique longitudinal study provide a comprehensive picture of ETI benefit by outlining not only acute improvement in LCI but maintained stability in LCI trajectories and improved LCI stability sustained up to three years post-initiation.
Otte, J. H.; Cartagena, A.
Show abstract
Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.
Pritz, S.; Bordag, N.; Foris, V.; Biasin, V.; Billensteiner, H.; Habisch, H.; Madl, T.; Marsche, G.; Nagaraj, C.; Suessner, S.; Kovacs, G.; Heresi, G.; Bodenhofer, U.; Olschewski, H.; Olschewski, A.
Show abstract
Rationale: Pulmonary hypertension is defined by pulmonary hemodynamics, but diagnostic and prognostic biomarkers remain limited. Nuclear magnetic resonance (NMR) spectroscopy provides detailed insights, particularly in the lipid metabolism. Objectives: To explore circulating NMR-derived metabolites and lipoprotein-related parameters for their association with pulmonary hemodynamics and to analyse their prognostic properties in pulmonary arterial hypertension (PAH). Methods: Retrospective analysis of a PAH cohort with complete diagnostic workup including right heart catheterization and baseline serum samples, from the prospective GRaz Pulmonary Hypertension-Metabolism (GRAPH-M) registry. Measurements: NMR-derived metabolites and lipoprotein-related parameters were analyzed for their association with clinically relevant parameters of PAH. We defined PHIHDL, a score derived from high-density lipoprotein (HDL) related measures based on their strong association with pulmonary hemodynamics, and evaluated its prognostic value. Results: We included 100 patients with PAH treated at the PH clinic of LKH University Hospital, Medical University of Graz, between 2011 and 2021. Age was 61{+/-}15 years, female/male ratio 2.5, BMI 26 {+/-}7 kg/m2, mPAP 41{+/-}16 mmHg, PAWP 8.8{+/-}3.2 mmHg, PVR 8.0{+/-}4.9 WU, and median survival was 8.0 years. During follow-up, 46 patients died. We identified a cluster of 12 HDL-related measures that showed significant inverse association to pulmonary hemodynamics and derived PHIHDL from the reversed scaled average of these particles. PHIHDL was associated with all-cause mortality after adjustment for age and sex (HR 2.96, 95% CI 1.52-5.70), independent of the clinical risk scores COMPERA 2.0 and REVEAL Lite2. Conclusion: PHIHDL, a pulmonary hemodynamics-based metabolomic score, provides independent prognostic information beyond established risk scores in PAH.
Matsubayashi, S.; Ito, S.; Hosaka, Y.; Yoshida, M.; Kadota, T.; Hashimoto, M.; Hatano, S.; Maruyama, T.; Fujimoto, S.; Nishioka, S.; Inukai, S.; Fujita, Y.; Minagawa, S.; Hara, H.; Nakada, T.; Nakayama, K.; Ohtuska, T.; Kuwano, K.; Araya, J.
Show abstract
Inadequate autophagy promotes smoking-induced cellular senescence involved in chronic obstructive pulmonary disease (COPD) pathogenesis. Transcription factor EB (TFEB) is a master regulator of the autophagy-lysosome axis. For the first time, we investigated the therapeutic potential of pemafibrate, a putative TFEB inducer. COPD lung epithelial cells showed reduced TFEB expression. Pemafibrate enhanced autophagy/mitophagy flux and restored lysosomal acidification observed during cigarette smoke (CS) extract exposure in human bronchial epithelial cells, resulting in reduced cellular senescence. TFEB knockdown demonstrated involvement of pemafibrate-induced TFEB in these effects. Pemafibrate induced TFEB expression, mitigated alveolar enlargement and airflow obstruction, and attenuated the CS-induced increase in static lung compliance in a long-term CS-exposed mouse model. It reduced the CS exposure-induced cellular senescence, possibly through autophagy/mitophagy, as suggested by bulk RNA sequencing of mouse lungs. A retrospective cohort study showed that patients given pemafibrate displayed attenuated FEV1.0 decline compared with those given bezafibrate or fenofibrate. In conclusion, pemafibrate is a promising therapeutic agent for COPD, potentially exerting its effects through the regulation of the TFEB-autophagy/mitophagy-lysosome axis.
Joseph-Delaffon, K.; Desgrouas, M.; Catanese, S.; Lejeune, J.; Nait-Kaci, J.; Piver, E.; Breteau, I.; Leducq, S.; Gatault, P.; Khanna, R. K.; Angoulvant, D.; Vallet, N.
Show abstract
Background. Designing high-quality Objective Structured Clinical Examination (OSCE) stations is a time-consuming process. Generative artificial intelligence (AI) represents a promising path to accelerate content creation by automating the generation of scenarios. A growing number of AI tools is now available for this purpose. Objective. To assess the variability between generative AI models in their ability to produce OSCE stations in the field of paediatrics. Methods. A structured prompt was developed based on the French national OSCE guidelines for medical education. Five distinct AI models were provided with this prompt, alongside the neonatal jaundice chapter from the French pediatric reference textbook, to generate 6 complete OSCE stations. Results. Prompt compliance was high for ChatGPT 5.1, ChatGPT 5.2, Gemini 3.0 Pro, and Claude Opus 4.5, while it was lower for Grok 4.1. Expert-rated quality was generally high, with few factual errors or missing information across models. However usability differed significantly between models. This was also true for several quality dimensions such as checklist clarity, embedding of checklist answers within vignettes, and ease of standardized patient formation. ChatGPT 5.1 required the most revisions and Gemini most often rated usable as is. Significant inter-model differences were observed in diagnostics, only with ChatGPT 5.1 sampling all three neonatal jaundice categories. Contextual variables showed systematic narrowing across models. Clinical grid density was consistent (10-12 items per station), but thematic distribution differed markedly. Soft skills coverage varied significantly across models (p=0.002), none of them consistently representing all communication competency domains. Conclusion. Large language models can generate structurally compliant OSCE stations, but surface compliance conceals substantive inter-model differences in diagnostic coverage, contextual diversity, and soft skills representation, that compromise content validity. No model currently meets the criteria for unsupervised deployment in a summative assessment bank. The choice of model carries pedagogical implications and expert curation remains essential before integration into high-stakes assessment workflows.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Weibel, S.; Duengfelder, H.; Pscheidl, T.; Krone, M.; Meybohm, P.
Show abstract
Background Despite numerous randomized controlled trials (RCTs) and systematic reviews (SRs), current sepsis guidelines continue to issue only weak recommendations for corticosteroids. We examined the clinical scope, underlying study pools, and mortality conclusions of SRs evaluating corticosteroids for sepsis. Methods We conducted a meta-research study of SRs on corticosteroids in sepsis (2015 to 2025), extracting SR characteristics, mortality results, and included RCTs. Study-pool overlap was assessed using an SRxRCT inclusion matrix, Jaccard similarity (J), and hierarchical clustering. SRs and RCTs were classified according to standardized Population, Intervention, Comparison, Outcome (PICO) profiles. We explored discordance in short-term mortality conclusions among clinically comparable SRs and potential associations with study-pool composition, target populations, and methodological characteristics. Results Forty-two SRs including 121 unique RCTs were identified. More than half of pairwise SR comparisons shared no RCTs, and only three pairs showed high overlap (J>0.8). SRs addressing similar intervention and target population profiles frequently relied on different study pools. Among 38 SRs with short-term mortality meta-analyses, 15 (39%) reported benefit and 23 (61%) no evidence of effect. Discordance occurred exclusively among SRs evaluating broad, non-specific corticosteroid strategies; conclusions were consistent for hydrocortisone plus fludrocortisone (benefit) and hydrocortisone, ascorbic acid, and thiamine (no evidence of effect). SRs including sepsis +/- shock populations more frequently reported benefit than those restricted to septic shock (62% vs 22%), although estimates were imprecise. No single methodological or clinical factor consistently explained discordance. Conclusions SRs addressing apparently similar clinical questions frequently synthesized different underlying evidence bases and reported discordant conclusions. Guideline developers should therefore consider not only methodological quality and reported PICO, but also whether the RCTs included in an SR adequately represent the intended clinical question. Clinically coherent evidence syntheses may improve the interpretability of pooled treatment effects and support more targeted corticosteroid therapy in sepsis.
Khodi Babaroudi, E.; Pham, M. H. X.; Lenz, I. T.; melgaard, e. l. r.; Grand, J.; Hove, J. D.; Seven, E.
Show abstract
Introduction: Nicotine Pouches are increasingly used as a smokeless alternative to cigarettes and other nicotine products, yet their acute cardiovascular effects remain poorly documented. While nicotine's impact on heart rate and electrocardiogram (ECG) parameters is well-documented in smoking, no trials have evaluated these effects specifically for nicotine pouches. Methods: This study is a single-center, double-blind, placebo-controlled, crossover trial which will include 20 healthy adult nicotine users. Participants will undergo three sessions, receiving either a placebo, 6 mg, or 14 mg nicotine pouch in random order. Heart rate obtained by an ECG and various other ECG parameters, vital signs, and subjective symptoms will be measured at baseline, and multiple time points over 30 minutes. Conclusions: This study aims to determine whether nicotine pouches cause acute changes in heart rate, ECG parameters, vital signs, and self-reported symptoms. We hypothesize that higher nicotine pouch does will lead to measurable increases in heart rate and other autonomic effects compared to placebo.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
McKinnon, G.; Tsai, W. H.; Ip-Buting, A.; Duff, N.; Fabreau, G. E.; McBrien, K.; David, O.; Donald, M.; Pendharkar, S. R.
Show abstract
Abstract Importance: Socially vulnerable patients have a high burden of obstructive sleep apnea, but the stage of the referral pathway at which access barriers arise is uncertain. Objective: To determine whether area-level social deprivation was associated with appointment scheduling, wait time, cancellations, or no-shows among adults referred for specialist obstructive sleep apnea care. Design: This was a cross-sectional study evaluating patients referred from December 1, 2016 through November 30, 2019. Data were analyzed from January 30, 2026 to May 16, 2026. Setting: Foothills Medical Centre Sleep Centre in Calgary, Canada. Participants: Adults referred to a tertiary academic sleep centre in Calgary, Alberta, Canada. Eligible patients had valid provincial health insurance and either a scheduled clinic appointment or home sleep apnea test data available. Exposures: Quintiles of the four Canadian Index of Multiple Deprivation domains: residential instability, economic dependency, ethnocultural composition, and situational vulnerability. Main Outcomes and Measures: The primary outcome was receipt of a scheduled specialist appointment. Secondary outcomes were time from referral to the first attended appointment and number of appointment cancellations or no-shows. Results: Among 3111 patients (mean [SD] age, 53.7 [14.3] years; 40.7% female), 1766 (56.7%) were scheduled and 1647 (52.9%) attended an appointment. Each quintile increase in situational vulnerability was associated with lower odds of scheduling (adjusted odds ratio [95% confidence interval] 0.86 [0.81-0.92]), whereas each quintile increase in ethnocultural composition was associated with higher odds (adjusted odds ratio [95% confidence interval] 1.32 [1.22-1.43]). Residential instability and economic dependency were not associated with scheduling. No deprivation domain was associated with time to the first attended appointment, cancellations, or no-shows. Conclusions and Relevance: In this cohort, area-level deprivation was associated with whether patients were scheduled for an appointment but not with wait time or missed visits after scheduling. These findings suggest that equity interventions should focus on completion of referral and scheduling processes.
Lenihan, S.; Barr, M.; Coates, K.; Kedroff, L.; Battle, C.; Sorice, V.; Faghy, M. A.; Edwards, J.; Papaioannou, D.; Young, T.; Rombach, I.; Carlton, E.; Goodacre, S.; Mani, N.
Show abstract
Background: Pain management post-rib fractures is often difficult. High pain levels can lead to altered respiratory mechanics and delayed complications such as poor mobility. Whilst as-needed Opioids are the mainstay of treatment, the potential negative side effects have led to research into alternatives such as kinesiotaping, single-shot chest wall regional anaesthesia, and incentive spirometry. Methods: A systematic review was undertaken using Medline (Ovid), Emcare, CINAHL, and the Cochrane Library. Article review and selection were undertaken by two independent reviewers using Covidence. Quality was assessed through the Mixed Methods Appraisal Tool (MMAT). Where appropriate, meta-analysis was undertaken using R studio with a REML random effects model. Forest plots were completed, and Higgins I2 and Chi2 were calculated. Results: Kinesiotaping demonstrates a reduced pain score than medication alone (SMD: -1.87, 95% CI [-2.65, -1.08]), as did single-shot chest wall regional anaesthesia (SMD: -0.79 [-1.15, -0.01]). The regional anaesthesia group had lower opioid consumption (SMD -0.84 [-2.18, 0.50]) and reduced length of hospital stay (SMD: -0.18 [-0.39, 0.03]) but no change in the risk of complications (RR: 0.92 [0.36, 2.36]). The incentive spirometry group had an increased risk of complications (RR: 3.35 [0.68, 16.44]); however, the causative effect could not be inferred due to significant confounding variables. Conclusions: Low-to-moderate certainty evidence suggests that kinesiotaping and single-shot chest wall regional anaesthesia may reduce pain in adult emergency department patients with rib fractures. However, evidence is insufficient to show a clear benefit for opioid reduction, length of stay, or complications. The current evidence does not support routine use of incentive spirometry in this setting, but the evidence is severely confounded by baseline injury severity in the current published studies.
Krasnova, T.; Zarkovic, M.; Nigg, C.; Sasaki, M.; Ganbat, M.; Casaulta, C.; Moeller, A.; Kuehni, C. E.
Show abstract
Background Exposure to environmental tobacco smoke (ETS) negatively affects children`s health, but few studies examined parental smoking behaviour in families of children with respiratory diseases. We studied parental smoking prevalence, characteristics, and changes over one year among families in the Swiss Paediatric Airway Cohort (SPAC). Methods We included children aged 0-17 years referred to paediatric respiratory outpatient clinics in Switzerland from 2017 to 2024. Parents answered a questionnaire at the initial clinic visit and again after one year. We used multivariable logistic regression to explore the characteristics of mothers and fathers who smoked and assessed changes in smoking behavior over one year. Results Among 4,199 children (median age 9 years [IQR 5-12]), 31% were exposed to parental smoking at baseline (paternal smoking: 16%; maternal smoking: 6%; both parents smoking: 9%). Mothers were more likely to smoke if they had a lower education level (OR 2.0, 95%CI 1.6-2.5 for compulsory education vs university education), did not have Swiss nationality (OR 1.3, 1.0-1.6) and lived in a socially disadvantaged neighborhood (OR 1.3, 1.0-1.7). Similar associations were observed for fathers. In addition, fathers were more likely to smoke if they were unemployed (OR 2.0, 1.3-3.2 vs having a full-time job. The strongest predictor of smoking was having a partner who smoked, with ORs above 6 for both mothers and fathers. Parents of 2,338 children completed the one-year follow-up questionnaire. Data from 2226 mothers and 1895 fathers showed that among baseline smokers with follow-up data, 225 (78%) mothers and 382 (81%) of fathers continued smoking, and only 63 (22%) of mothers and 90 (19%) of fathers quit. Among baseline non-smokers, 47 (2%) mothers and 54 (3%) fathers started smoking. Conclusions One-third of children consulting respiratory specialists in Switzerland are exposed to parental smoking. ETS exposure was strongly associated with socio-economic factors. Even after visiting a specialized clinic, most parents continued to smoke. This highlights the urgent need for stronger national smoking policies and targeted support to help these parents quit and stay smoke-free.
Patchigolla, V.; Jhand, A. S.; Lee, H. J.; Benjamins, L. J.
Show abstract
Evidence-based medicine (EBM) concepts are difficult for medical students to grasp. We developed a Python-Streamlit web application providing interactive visualizations to enhance EBM education. Preliminary use with first year medical students demonstrated high engagement and improved conceptual understanding, supporting the feasibility of integrating interactive, web-based tools into EBM curricula.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Wang, N.; Huang, H.; Chu, J.; Hsu, J.
Show abstract
Objectives: Healthcare data can reveal actionable opportunities to prevent asthma hospitalizations. Limited national-level data exist regarding social determinants of health (SDOH) and asthma hospitalizations. We examined SDOH-related International Classification of Diseases, Tenth Revision (ICD-10) Z-codes in national administrative data on asthma hospitalizations and described patient- and hospital-level characteristics associated with documented SDOH Z-codes. Methods: Pooled cross-sectional analysis of 2016-2022 Nationwide Inpatient Sample for 200,452 U.S. hospitalizations (all ages) with a primary diagnosis of asthma. Presence of SDOH Z-codes (codes Z55-Z65) assessed by descriptive statistics and multivariable logistic regression to calculate odds ratios (ORs) and 95% confidence intervals (95% CIs) for associations between SDOH Z-codes and patient- and hospital-level characteristics. Results: In unweighted analyses, 3,149 asthma hospitalizations had SDOH Z-codes (1.57%). The most common SDOH Z-codes were homelessness (Z59.0; n=942) and unemployment (Z56.0; n=349). Weighted chi-square analyses found all selected variables were associated with asthma hospitalization SDOH Z-code documentation. Logistic regression results varied; adjusted odds for SDOH Z-code documentation were higher for asthma hospitalizations involving male patients (aOR=1.51; 95% CI, 1.39-1.63; P < .001) compared to female patients. Asthma hospitalizations involving rural hospitals had lower odds of SDOH Z-codes documentation (aOR=0.57; 95% CI, 0.47-0.70; P < .001) compared to urban teaching hospitals. Conclusions: National 2016-2022 data indicate housing- and employment-related Z-codes were the most commonly documented SDOH within asthma hospitalizations. Future analyses could consider establishing causality and exploring how relationships between these SDOH may be used by public health practitioners and others to improve program interventions.
Huapaya, J.; Burbelo, P.; Robbins, E. W.; Tian, X.; Gao, S.; Turan, S.; Gairhe, S.; Ward, J.; Redekar, N.; Li, J.; Pastor, G.; Gupta, N.; Noroozi Farhadi, P.; Sarkar, K.; Casal-Dominguez, M.; Pinal-Fernandez, I.; Christopher-Stine, L.; Schiffenbauer, A.; Rider, L.; Mammen, A. L.; Danoff, S. K.; Suffredini, A. F.
Show abstract
Introduction: Idiopathic inflammatory myopathy-associated interstitial lung disease (IIM-ILD) is a major cause of morbidity and mortality. We tested whether quantitative myositis-specific autoantibodies and proteomic profiling capture biological heterogeneity and prognosis beyond categorical serology. Methods: Myositis-specific autoantibodies were quantified using the luciferase immunoprecipitation systems assay, and 184 serum proteins were measured in 226 IIM patients; 199 with higher-ILD-risk autoantibodies (Jo-1/MDA5/PL-7/PL-12/EJ), 27 with lower-ILD-risk autoantibodies (Mi-2/NXP2/TIF1{gamma}) and 35 healthy controls. We identified shared and subgroup-specific differences by comparing each subgroup with controls, then correlated quantitative autoantibody and protein levels within higher-risk subgroups. Additional analyses included pathway enrichment, unsupervised clustering, longitudinal lung-function change, and mortality. Results: Higher-ILD-risk subgroups shared interferon-responsive CXCR3 chemokine, IL-6/JAK/STAT3, and apoptosis signaling. Dominant autoantibody subgroup profiles differed: interferon/CXCR3 chemokine signaling with T-cell activation and monocyte recruitment in anti-Jo-1; proteostasis/antigen-processing and vascular/cellular stress signals in anti-MDA5; IL-6/macrophage and profibrotic signals in anti-PL-12; and apoptotic and innate immune activation with metabolic/redox-stress signals in anti-PL-7. Within higher-ILD-risk subgroups, autoantibody levels correlated with interferon-response, profibrotic, and metabolic/vascular proteins (r=0.40-0.74; nominal p<0.05). Unsupervised clustering identified four proteomic endotypes beyond autoantibody type, including an injury-stress endotype associated with worse lung function and poorer survival, and a chemokine/checkpoint-high endotype with relatively preserved lung function. Across 203 participants with 38 deaths, a weighted 10-protein score was associated with all-cause mortality (HR, 3.28; 95% CI, 2.12-5.08; p<0.001). Conclusions: Integrated quantitative autoantibodies and proteomic profiling revealed shared inflammatory biology, autoantibody-associated signatures, and an injury-stress endotype associated with poor survival in IIM-ILD, supporting risk stratification beyond categorical serology.