Back

CHEST

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match CHEST's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Positive end-expiratory pressure versus sham valve/zero end-expiratory pressure in cardiopulmonary resuscitation during manual ventilation toimprove neurological outcomes in adult patients suffering an out-of-hospital cardiac arrest - an investigator-initiated, pragmatic, registry-based, multicenter, parallel-group, triple-blind randomized controlled superiority clinical trial in the ARREST registry (REVIVE-PEEP protocol Stage-1 Registered Report)

van Eijk, J.; Schober, P.; van Schuppen, H.; ter Schure, J.

2026-08-31 emergency medicine 10.64898/2026.08.27.26361533 medRxiv
Top 0.1%
11.5%
Show abstract

We present our Stage-1 Registered Report as a full clinical trial article with all methods in past tense and including mock results, table and figures for the primary analysis. To remind the reader that this Stage-1 article is written before data collection, we highlight in color that these mock results are only for illustrative purposes and will be replaced by the actual results in the Stage-2 Registered Report. Background In patients experiencing out-of-hospital cardiac arrest, optimization of oxygen delivery during cardiopulmonary resuscitation is a critical. Although both positive end-expiratory pressure (PEEP) and zero end-expiratory pressure (ZEEP) are employed during CPR, their respective impacts on clinically relevant outcomes is yet to be clearly established. Methods This investigator-initiated, pragmatic, registry-based, multicenter, triple-blind randomized controlled superiority trial evaluates whether applying 8 cm H2O PEEP during cardiopulmonary resuscitation improves outcomes compared with ZEEP in adults with non-traumatic, non-drowning out-of-hospital cardiac arrest. Pre-randomized CPR kits (1:1 PEEP vs. sham) were used by ambulance sites during manual ventilation throughout the resuscitation process. The primary analysis was conducted in the principal stratum of patients who received either a supraglottic airway or endotracheal tube. The primary outcome was neurological status at hospital discharge measured by a utility-weighted score on the modified Rankin Scale. Secondary outcomes included prehospital return of spontaneous circulation, 30-day survival, and 6-month quality of life. The primary safety outcome was clinically significant pneumothorax.

2
Effects of collaborative clinical visit agenda-setting interventions: A systematic review and meta-analysis

Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.

2026-09-03 medical education 10.64898/2026.08.30.26361729 medRxiv
Top 0.1%
3.6%
Show abstract

Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.

3
Lung function trajectories in children with cystic fibrosis aged 3-17 years: impact of elexacaftor-tezacaftor-ivacaftor on lung function

Dyer, B. P.; Deery, M.; Heyman, R.; Robinson, P.; Wainwright, C.; Sly, P.; Ware, R.; Blake, T.

2026-09-02 respiratory medicine 10.64898/2026.08.31.26361791 medRxiv
Top 0.1%
3.2%
Show abstract

Background Elexacaftor-tezacaftor-ivacaftor (ETI) has been demonstrated to improve lung function in clinical trials; however, evidence describing effects on trajectories and whether long-term improvements are sustained (>1-year) is lacking. We estimated within-person lung clearance index (LCI) trajectories before and after ETI initiation, assessing changes in level and rate of change, alongside acute LCI change, up to three years after ETI initiation. Methods Prospective observational study of children at a tertiary hospital. Children aged 3-17 years with [&ge;]2 LCI testing occasions (i) before and (ii) after starting ETI were used to describe lung function trajectories. Children with [&ge;]1 pre-ETI and [&ge;]1 post-ETI LCI occasion(s) were used to describe acute LCI change after ETI initiation. Age-adjusted LCI trajectories for time periods (i) before and (ii) after ETI initiation were estimated using linear mixed-effects models, and pre- and post-ETI LCIs were compared using paired Wilcoxon tests. Results Mean pre-ETI and post-ETI longitudinal changes in LCI were -0.007 (95% CI: -0.28, 0.27; n=35) and 0.12 (95% CI: -0.17, 0.41; n=20) turnovers per year, respectively. Before ETI initiation, 57% (30/53) of patients had an LCI[&ge;]7.1 turnovers (indicating impaired lung function), compared to 26% (14/53) post-ETI, with a median LCI difference of -0.70 (95% CI -0.84, -0.46; p<0.001) turnovers. Within-individual variability in LCI decreased post-ETI. Conclusions Our real-world data within a unique longitudinal study provide a comprehensive picture of ETI benefit by outlining not only acute improvement in LCI but maintained stability in LCI trajectories and improved LCI stability sustained up to three years post-initiation.

4
AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Otte, J. H.; Cartagena, A.

2026-08-31 medical education 10.64898/2026.08.26.26361437 medRxiv
Top 0.1%
2.8%
Show abstract

Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.

5
PHIHDL: A Novel HDL Index Predicting Baseline Pulmonary Hemodynamics and Long-Term Survival in PAH

Pritz, S.; Bordag, N.; Foris, V.; Biasin, V.; Billensteiner, H.; Habisch, H.; Madl, T.; Marsche, G.; Nagaraj, C.; Suessner, S.; Kovacs, G.; Heresi, G.; Bodenhofer, U.; Olschewski, H.; Olschewski, A.

2026-09-02 respiratory medicine 10.64898/2026.08.31.26361587 medRxiv
Top 0.1%
2.1%
Show abstract

Rationale: Pulmonary hypertension is defined by pulmonary hemodynamics, but diagnostic and prognostic biomarkers remain limited. Nuclear magnetic resonance (NMR) spectroscopy provides detailed insights, particularly in the lipid metabolism. Objectives: To explore circulating NMR-derived metabolites and lipoprotein-related parameters for their association with pulmonary hemodynamics and to analyse their prognostic properties in pulmonary arterial hypertension (PAH). Methods: Retrospective analysis of a PAH cohort with complete diagnostic workup including right heart catheterization and baseline serum samples, from the prospective GRaz Pulmonary Hypertension-Metabolism (GRAPH-M) registry. Measurements: NMR-derived metabolites and lipoprotein-related parameters were analyzed for their association with clinically relevant parameters of PAH. We defined PHIHDL, a score derived from high-density lipoprotein (HDL) related measures based on their strong association with pulmonary hemodynamics, and evaluated its prognostic value. Results: We included 100 patients with PAH treated at the PH clinic of LKH University Hospital, Medical University of Graz, between 2011 and 2021. Age was 61{+/-}15 years, female/male ratio 2.5, BMI 26 {+/-}7 kg/m2, mPAP 41{+/-}16 mmHg, PAWP 8.8{+/-}3.2 mmHg, PVR 8.0{+/-}4.9 WU, and median survival was 8.0 years. During follow-up, 46 patients died. We identified a cluster of 12 HDL-related measures that showed significant inverse association to pulmonary hemodynamics and derived PHIHDL from the reversed scaled average of these particles. PHIHDL was associated with all-cause mortality after adjustment for age and sex (HR 2.96, 95% CI 1.52-5.70), independent of the clinical risk scores COMPERA 2.0 and REVEAL Lite2. Conclusion: PHIHDL, a pulmonary hemodynamics-based metabolomic score, provides independent prognostic information beyond established risk scores in PAH.

6
Senotherapeutic role of pemafibrate through autophagy/mitophagy regulation in chronic obstructive pulmonary disease

Matsubayashi, S.; Ito, S.; Hosaka, Y.; Yoshida, M.; Kadota, T.; Hashimoto, M.; Hatano, S.; Maruyama, T.; Fujimoto, S.; Nishioka, S.; Inukai, S.; Fujita, Y.; Minagawa, S.; Hara, H.; Nakada, T.; Nakayama, K.; Ohtuska, T.; Kuwano, K.; Araya, J.

2026-09-02 respiratory medicine 10.64898/2026.08.31.26361865 medRxiv
Top 0.2%
1.9%
Show abstract

Inadequate autophagy promotes smoking-induced cellular senescence involved in chronic obstructive pulmonary disease (COPD) pathogenesis. Transcription factor EB (TFEB) is a master regulator of the autophagy-lysosome axis. For the first time, we investigated the therapeutic potential of pemafibrate, a putative TFEB inducer. COPD lung epithelial cells showed reduced TFEB expression. Pemafibrate enhanced autophagy/mitophagy flux and restored lysosomal acidification observed during cigarette smoke (CS) extract exposure in human bronchial epithelial cells, resulting in reduced cellular senescence. TFEB knockdown demonstrated involvement of pemafibrate-induced TFEB in these effects. Pemafibrate induced TFEB expression, mitigated alveolar enlargement and airflow obstruction, and attenuated the CS-induced increase in static lung compliance in a long-term CS-exposed mouse model. It reduced the CS exposure-induced cellular senescence, possibly through autophagy/mitophagy, as suggested by bulk RNA sequencing of mouse lungs. A retrospective cohort study showed that patients given pemafibrate displayed attenuated FEV1.0 decline compared with those given bezafibrate or fenofibrate. In conclusion, pemafibrate is a promising therapeutic agent for COPD, potentially exerting its effects through the regulation of the TFEB-autophagy/mitophagy-lysosome axis.

7
Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry

Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.

2026-09-03 dentistry and oral medicine 10.64898/2026.09.01.26361980 medRxiv
Top 0.2%
1.5%
Show abstract

Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.

8
An interpretable, formally verified point-of-care ultrasound risk equation for difficult videolaryngoscopy: development and internal validation

Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.

2026-09-02 anesthesia 10.64898/2026.08.28.26361621 medRxiv
Top 0.2%
1.5%
Show abstract

Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.

9
Acute Cardiovascular and Electrocardiographic Effects of Nicotine Pouches: A study protocle for a Randomized, Double-Blind, Placebo-Controlled Crossover Trial (NICOTUNE STUDY)

Khodi Babaroudi, E.; Pham, M. H. X.; Lenz, I. T.; melgaard, e. l. r.; Grand, J.; Hove, J. D.; Seven, E.

2026-08-31 public and global health 10.64898/2026.08.29.26361726 medRxiv
Top 0.3%
1.1%
Show abstract

Introduction: Nicotine Pouches are increasingly used as a smokeless alternative to cigarettes and other nicotine products, yet their acute cardiovascular effects remain poorly documented. While nicotine's impact on heart rate and electrocardiogram (ECG) parameters is well-documented in smoking, no trials have evaluated these effects specifically for nicotine pouches. Methods: This study is a single-center, double-blind, placebo-controlled, crossover trial which will include 20 healthy adult nicotine users. Participants will undergo three sessions, receiving either a placebo, 6 mg, or 14 mg nicotine pouch in random order. Heart rate obtained by an ECG and various other ECG parameters, vital signs, and subjective symptoms will be measured at baseline, and multiple time points over 30 minutes. Conclusions: This study aims to determine whether nicotine pouches cause acute changes in heart rate, ECG parameters, vital signs, and self-reported symptoms. We hypothesize that higher nicotine pouch does will lead to measurable increases in heart rate and other autonomic effects compared to placebo.

10
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 0.3%
0.9%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

11
AURORA: Analysing and understanding responses to oncological regimens with artificial intelligence

Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.

2026-09-02 health informatics 10.64898/2026.08.30.26361778 medRxiv
Top 0.4%
0.8%
Show abstract

Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.

12
No Overall Survival Benefit with Adding Chemotherapy to Immunotherapy in PD-L1 TPS >= 50% NSCLC: An Agent-Stratified Reassessment

Han, F.; Wang, J.; Shi, S.; Jin, M.; Ren, C.

2026-09-03 oncology 10.64898/2026.09.01.26361919 medRxiv
Top 0.7%
0.4%
Show abstract

IMPORTANCE: A recent meta-analysis showed that chemoimmunotherapy was associated with improved overall survival (OS) compared with immune checkpoint inhibitor (ICI) monotherapy for programmed death-ligand 1 (PD-L1) tumor proportion score (TPS) [&ge;] 50% advanced non-small-cell lung cancer (NSCLC). However, whether this benefit reflects chemotherapy effect or ICI heterogeneity remains unclear. OBJECTIVE: To reassess the survival benefit of adding chemotherapy to ICI monotherapy using agent-stratified comparisons anchored to chemotherapy. DATA SOURCES: The 24 phase 3 randomized clinical trials included in the original meta-analysis (search date, August 3, 2025). DATA EXTRACTION AND SYNTHESIS: Hazard ratios (HRs) for OS and progression-free survival (PFS) were extracted from each trial in the original meta-analysis. Two analytic frameworks were used: within-agent comparisons (same ICI in both chemoimmunotherapy and monotherapy) and across-agent comparisons (ICI in one treatment strategy only). For within-agent comparisons, a two-stage random-effects meta-analysis was conducted. In stage 1, ICI-specific HRs for chemoimmunotherapy and ICI monotherapy versus chemotherapy were pooled and their ratio was calculated (RHR = HRchemoimmuno/HRmono; RHR < 1 favors chemoimmunotherapy). The RHRs were pooled in stage 2. For across-agent comparisons, RHR was derived from pooled HRs by treatment strategy. MAIN OUTCOMES AND MEASURES: Endpoints were OS and PFS. RESULTS: In within-agent comparisons (4 ICIs; 13 trials; N = 3252), pooled RHR was 0.94 (95% CI, 0.78-1.13; P = .48; I2 = 0.0%) for OS and 0.85 (95% CI, 0.68-1.06; P = .14; I2 = 0.0%) for PFS. In across-agent comparisons (7 ICIs; 11 trials; N = 2231), RHR favored chemoimmunotherapy for OS (0.68; 95% CI, 0.50-0.92; P = .01) and PFS (0.46; 95% CI, 0.37-0.58; P < .001). In a sensitivity analysis restricted to trials of NCCN-recommended regimens, pooled RHR was 1.02 (95% CI, 0.81-1.28; P = .87) for OS. CONCLUSIONS AND RELEVANCE: In the within-agent comparisons, adding chemotherapy to ICI monotherapy did not improve OS or PFS in patients with PD-L1 TPS [&ge;] 50% advanced NSCLC. The benefit in the original meta-analysis appears driven by across-ICI heterogeneity. These findings are consistent with ICI monotherapy as a standard first-line option and underscore the need for agent-level stratification in across-trial comparisons.

13
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 0.7%
0.3%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

14
Postoperative analgesia and recovery after minimally invasive cardiac surgery

Note, H.; Kajiura, T.; Muramatsu, A.; Inagaki, Y.; Takahashi, T.; Sato, K.; Nakamura, K.; Sadatoshi, T.; Sakurai, Y.; Tochii, M.; Watanuki, H.; Matsuyama, K.; Okamoto, S.

2026-08-31 intensive care and critical care medicine 10.64898/2026.08.27.26361580 medRxiv
Top 0.7%
0.3%
Show abstract

Introduction Postoperative analgesic management after minimally invasive cardiac surgery (MICS) should facilitate early recovery while providing adequate pain control. However, direct evidence comparing postoperative remifentanil- and fentanyl-based analgesic strategies after MICS remains limited. We compared these strategies and explored their associations with postoperative recovery, postoperative nausea and vomiting (PONV), and pain management. Methods This retrospective single-center observational cohort study included patients who underwent MICS via a right mini-thoracotomy between January 2023 and June 2026. Patients were categorized according to postoperative remifentanil- or fentanyl-based analgesia in the intensive care unit. Outcomes included time to extubation, PONV, postoperative pain assessed using the numerical rating scale (NRS), additional analgesic use, and intensive care unit length of stay. Multivariable logistic regression examined the association between postoperative opioid strategy and PONV, adjusting for age, sex, and smoking history. Results PONV occurred less frequently in the remifentanil group than in the fentanyl group (20.6% vs 45.0%, P = 0.004), and this association remained significant after adjustment (adjusted odds ratio, 0.23; 95% confidence interval, 0.10-0.56; P = 0.001). Time to extubation was shorter with remifentanil (median, 179 [interquartile range, 134-240.5] vs 247 [190.2-276.5] min; P < 0.001). In contrast, NRS pain scores on postoperative day 0 were higher with remifentanil (3 [1-6] vs 1 [0-2]; P < 0.001), and additional analgesics were used more frequently (80.6% vs 33.3%; P < 0.001). Pain scores on postoperative day 1 did not differ significantly between groups. Conclusion Postoperative remifentanil-based analgesia after MICS was associated with less PONV and earlier extubation but also with greater early postoperative pain and more frequent additional analgesic use than fentanyl-based analgesia. Appropriate transition to longer-acting analgesics with multimodal analgesia may help preserve the potential benefits of remifentanil while maintaining adequate postoperative pain control.

15
Comparative effectiveness of preventive strategies against medically-attended respiratory syncytial virus in U.S. infants during the first six months of life, 2023-2025

Kim, S. S.; Zissette, S. Z.; Van Meter, C.; Shiiba, M.; Bruck, M.; Tippett, A.; Kamidani, S.; Benkeser, D.; McQuade, E. R.

2026-08-31 epidemiology 10.64898/2026.08.25.26361361 medRxiv
Top 0.9%
0.3%
Show abstract

Importance: Maternal vaccination and long-acting monoclonal antibodies are now available in the U.S. to prevent RSV. Long-acting monoclonal antibody administration in the U.S. commonly occurs after hospital discharge in outpatient settings, leaving some infants unprotected early in life when severe RSV risk is highest. Comparative effectiveness between the two interventions and whether delays affect effectiveness estimates have not been quantified. Objective: To evaluate the effectiveness of infant long-acting monoclonal antibody strategies and a maternal vaccination strategy, each compared to no intervention, and the comparative effectiveness of intervention strategies when accounting for real-world delays in monoclonal antibody receipt. Design: Cohort study using target trial emulation to compare four strategies for prevention of RSV-related outcomes. Setting: The U.S. between 2023 and 2025 using a nationwide database of employer-sponsored commercial insurance claims. Participants: 120,586 commercially insured mother-infants, whose infants were born in the U.S. during the 2023-2024 or 2024-2025 RSV season. Infants who could not be paired with their mother's record, did not enroll in commercial insurance within 75 days from birth, received palivizumab, and had an implausible birth date were excluded. Interventions: Comparison of four RSV prevention strategies: (i) maternal RSVpreF; (ii) long-acting monoclonal antibody given within the first week of life (mAb as intended); (iii) long-acting monoclonal antibody given within a six-month grace period from birth (mAb within grace period); and (iv) a control. Main outcomes and measures: Effectiveness against first RSV-associated hospitalization and medically-attended RSV illness was summarized using adjusted hazard ratios (aHR) and estimated using an inverse propensity weighting approach, with weights accounting for maternal age, maternal comorbidities affecting pregnancy, obstetric and newborn complications, season, region, and birth timing relative to October 1. A weighted Kaplan Meier estimator was used to estimate strategy-specific cumulative incidence of RSV outcomes over time. Results: In the first five weeks of life, the mAb within grace period strategy doubled the hazard of RSV hospitalization (aHR: 2.0 [95% CI: 1.0-4.9]) and increased the hazard of medically-attended RSV (aHR: 1.6 [95% CI: 1.0-2.7]) compared to the maternal RSVpreF strategy. The hazard for RSV hospitalization was similar for the mAb as intended strategy compared to the maternal RSVpreF strategy (aHR = 0.9 [95% CI: 0.3-1.9]). Conclusions and relevance: RSVpreF and monoclonal antibodies were similarly effective when monoclonal antibodies were administered close to birth, but when accounting for real-world delays in monoclonal antibody receipt, the maternal RSVpreF strategy was more effective than the mAb within grace period strategy.

16
Predicting COVID-19 hospitalisation and common disease risk from comorbid diagnoses in 13 million individuals

Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,

2026-09-01 health informatics 10.64898/2026.08.27.26361302 medRxiv
Top 0.9%
0.3%
Show abstract

Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.

17
Dynamic Clinical States and Transitions During the First 72 Hours of Intensive Care After Acute Stroke

LEI, P.; XU, Y.; ZHANG, Y.

2026-09-01 intensive care and critical care medicine 10.64898/2026.08.30.26361738 medRxiv
Top 0.9%
0.3%
Show abstract

Background: The condition of a patient with acute stroke often changes within hours of ICU admission. Prognostic work here targets fixed endpoints predicted from admission data, and trajectory phenotyping assigns one label per patient. We used longitudinal ICU data to identify interpretable dynamic clinical states, characterize transitions between them, and relate the current state to later events. Methods: Retrospective cohort study of 6368 adults with acute stroke in MIMIC IV v3.1. The first 72 h were divided into twelve 6-hour windows, and a hidden Markov model was fitted to 21 neurological, physiological and organ support variables. State number was chosen against criteria fixed before fitting: statistical fit, restart stability, state occupancy and clinical interpretability. Generalized estimating equations related the current state to new mechanical ventilation and vasopressor use within 12 h, and to ICU death within 72 h. Eleven sensitivity analyses assessed the robustness of the state solution. Results: Four states were selected: neurologically preserved-low support, neurological impairment low support, impairment renal dysfunction and impairment-respiratory support (63.3%, 7.8%, 11.8% and 17.1% of windows). Within 72 h, 40.3% of patients changed state at least once, and transitions ran in both directions rather than along a single severity gradient. States were identified without outcome data, yet ICU mortality by last state ranged from 2.9% to 43.9%. Adjusted for age, sex, subtype and Charlson index, the current state remained associated with organ-support escalation and death. State prevalence differed by at most 1.1 percentage points between training and test sets, and 10 of 11 sensitivity analyses gave a stable four-state solution (ARI 0.754 0.955). Conclusions: The early ICU course of acute stroke can be represented as movement among a small number of clinically interpretable states. The representation was reproducible in a held out set and across admission eras, but requires validation in an independent database before any clinical use.

18
Acute Renal, Hepatic, Thromboembolic and Functional Complications after Community-Acquired Acute Lower Respiratory Tract Infection: A Prospective Cohort Study in Bristol, UK, 2022-2024

Chatzilena, A.; Hyams, C.; Challen, R.; Lahuerta, M.; McGuinness, S.; Clout, M.; Begier, E.; King, J.; Morales-Aza, B.; Duale, K.; Rodriguez Pereira, A.; Healy, W.; Southern, J.; Wells, P.; Lihou, K.; Grimes, C.; Campling, J. A.; Maskell, N.; Oliver, J.; Vyse, A.; Gessner, B.; Finn, A.; Danon, L.; The AvonCAP Research Group,

2026-09-02 respiratory medicine 10.64898/2026.08.28.26361617 medRxiv
Top 1%
0.2%
Show abstract

Introduction Acute lower respiratory tract disease (aLRTD) is a leading cause of hospitalisation and death, particularly in older adults and adults with comorbidities, with acute lower respiratory tract infection (aLRTI; pneumonia and non-pneumonic LRTI) being a major component. Non-pulmonary complications and functional decline after aLRTI are recognised, but their pathogen-specific burden is poorly described. We aimed to quantify renal, hepatic, thromboembolic and functional complications, and mortality, after aLRTI hospitalisation, by clinical phenotype and pathogen. Methods We conducted a cohort study of adults (>18 years) admitted with aLRTD to two hospitals in Bristol, UK (01 August 2022-31 July 2024). aLRTD was classified as pneumonia, non-pneumonic LRTI (NP-LRTI) or no diagnosis of aLRTI. Pathogens were identified from standard-of-care and research microbiology. Outcomes were acute kidney injury (AKI), acute liver dysfunction, venous thromboembolism (VTE), in-hospital falls, reduced mobility at discharge, increased care requirements, and 30-day and 1-year mortality. Analyses were descriptive. Results Among 246,797 adult admissions, 21,456 aLRTD hospitalisations were included: 10,239 (47.7%) pneumonia, 7,742 (36.1%) NP-LRTI and 3,475 (16.2%) with no evidence of aLRTI. Of 19,152 tested aLRTD admissions, 8,503 (44.4%) had a positive microbiological/virological test, yielding 9,204 pathogen detections; 1,194 (6.2%) had co-infections, and SARS-CoV-2 was most frequent, with influenza the second most common in pneumonia and NP-LRTI. Pneumonia had greater severity than NP-LRTI and no diagnosis of aLRTI (median length of stay 6 vs 4 vs 4 days; ICU admission 3.4% vs 0.7% vs 0.5%, respectively). Overall, 22.2% developed AKI, 6.1% acute liver dysfunction, 0.6% DVT and 2.4% PE; 1.8% had a fall, 11.5% reduced mobility, and 16.6% required increased care at discharge. 30-day and 1-year mortality were highest for pneumonia (14.0% and 32.0%, respectively). Pathogen-specific analyses showed longer stays and higher complications and mortality rates for SARS-CoV-2 and Streptococcus pneumoniae, and shorter stays with lower complication and mortality rates for influenza and Haemophilus influenzae. Conclusions Non-cardiovascular complications and functional decline after aLRTI were common, particularly in pneumonic and SARS-CoV-2 or pneumococcal disease. These findings support routine surveillance for renal, hepatic, thromboembolic events, early mobilisation and rehabilitation, and consideration of multi-system outcomes when evaluating public health and economic value of vaccines and therapies.

19
Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model

Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.

2026-08-31 health informatics 10.64898/2026.08.26.26361443 medRxiv
Top 1%
0.2%
Show abstract

Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [&ge;]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.

20
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 1%
0.2%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.