JAMA
● American Medical Association (AMA)
Preprints posted in the last 7 days, ranked by how well they match JAMA's content profile, based on 18 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Boden-Albala, B.; Wing, J.; Landry, M. J.; Castro, M.; Gutierrez, D.; Cardenas, C.; Rousseau, J.; Rahmani, A. M.; Chavez, A.; Ding, X.; Kurzman, A.; Albala, B.
Show abstract
Background: Cardiovascular disease (CVD) disproportionately burdens underserved communities, where social determinants of health (SDOH) perpetuate persistent disparities. Family-based interventions leveraging social support represent a promising yet understudied approach. We describe the rationale, design, and methods of the Skills-based Educational strategies for the Reduction of Vascular Events in Orange County (SERVE OC) RCT and present baseline characteristics of enrolled families. Methods: SERVE OC is a 2-arm RCT of 190 Latino and Vietnamese families (486 individuals) randomized to the family-based intervention or individual self-management. The intervention was grounded in social network theory while employing community engaged strategies. Primary outcomes include achieving ideal cardiovascular health (CVH) defined by AHA Life's Essential 8 (LE8) and systolic blood pressure reduction at 12, 24, and 36 months. Baseline assessments include demographics, LE8, psychosocial factors, food security, and SDOH. Descriptive statistics and regression analyses examined cohort characteristics and associations between SDOH, food security, and LE8. Results: Over 83% of participants had suboptimal LE8 scores. Average adult total LE8 scores were 66.61 {plus minus}11.96, with physical activity as the weakest domain, compared to an average of 76.52{plus minus}10.15 in children. Greater SDOH burden and food security were associated with significantly lower odds of ideal CVH and lower LE8 scores respectively. Conclusions: SERVE OC demonstrates the feasibility of enrolling families in community-engaged RCT targeting CVD disparities in underserved population. Baseline findings confirm substantial CVD risk and SDOH burden underscoring the need for multi-level, culturally tailored interventions. Trials results will inform scalable, family-focused strategies for CVD prevention across the life course. Clinical Trial Registration: URL: https://www.clinicaltrials.gov/; Unique Identifier: NCT05641519.
Pichkar, Y.; Manolakos, S.; Phillips, K. M.; Schabath, M. B.; Chaudhary, A.
Show abstract
Background: Low-dose computed tomography (LDCT) screening reduces lung cancer mortality but is limited by low uptake and associated with high rates of false-positives and indeterminate-nodules. Breath volatile organic compound (VOC) analysis is a non-invasive candidate biomarker approach that could complement LDCT, but prior work has relied on laboratory-based high-resolution mass spectrometry (HRMS), limiting point-of-care deployment. Methods: In this pilot study, breath samples were collected from 40 patients with treatment-naive, pathologically confirmed non-small cell lung cancer (NSCLC) and 25 lung-cancer-screening-eligible healthy controls. Paired samples were analyzed via a compact point-of-care GC-MS platform (CLARION) and a laboratory HRMS reference. Diagnostic classification models were built independently for each platform using elastic net logistic regression with leave-one-out cross-validation, and performance was evaluated by area under the receiver operating characteristic curve (AUC). Results: CLARION identified 103 VOCs across breath specimens, compared to over 900 identified by HRMS. Despite this difference in panel size, CLARION achieved diagnostic performance nearly identical to HRMS for distinguishing NSCLC cases from controls (AUC 0.864 vs. 0.863). Compared to controls, performance statistics were similar for early-stage NSCLC (AUC 0.854 vs. 0.841) and adenocarcinoma (AUC 0.770 vs. 0.787). VOCs of interest include p-cymene, phenol, propylbenzene, tetradecane, {beta}-ocimene, 2,3-dihydro-indole, and 1-methylthio-(Z)-1-propene. Conclusion: A compact, point-of-care breath GC-MS platform achieved diagnostic performance for NSCLC detection comparable to a laboratory HRMS reference despite a substantially smaller detected VOC panel. These findings support continued development of point-of-care breath VOC testing as a non-invasive, field-deployable complement to LDCT-based lung cancer screening.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Jafree, D. J.; Sun, M.; Stewart, G. W.; Gishen, F.; Swanton, C.; Motallebzadeh, R.; UCL MB-PhD Outcomes Study Group,
Show abstract
Background: Clinician-scientists translate clinical observation into discovery, trials, and policy, yet this workforce is shrinking across health systems worldwide. Integrated MB-PhD training, pausing medical training to complete a PhD before clinical exposure or specialisation, is one route into this career. We aimed to evaluate the long-term value of MB-PhD training and the barriers to clinical-academic careers these face after graduation. Methods: We evaluated all 131 graduates (29.8% female) who entered the University College London (UCL) MB-PhD programme over a 25-year period (1994-2018). Bibliometric outputs were collated via an inter-linked information system. Concurrently, all 131 graduates were invited to respond to open-ended questions on career benefits and structural barriers; 99 (75.6%) responded, and responses were independently coded into themes, which were then reviewed and confirmed by a Study Group of 107 individuals, including the 91 respondents who agreed to participate further. Results: Graduates produced 5,877 publications (1,141 first-author, 819 corresponding-author), attracting 350,754 citations, with a mean relative citation ratio of 3.30 {+/-} 0.47, approximately three times the field average and sustained across three decades of programme entry. Graduates secured an estimated $157.55 million across 99 grants, released 465 public datasets, and were named investigators on 31 clinical trials across five continents. Among the 99 survey respondents, 49.5% held consultant-grade posts, 72.7% remained research-active, and 25.3% had reached senior academic grade. Open-ended responses were coded into five recurring structural barriers, subsequently confirmed by the Study Group: insufficient protected research time (72.2% of responses), unsupportive training structures and limited career opportunities (36.7%, 24.4% of responses), funding and pay barriers (22.2% of responses), and lack of mentorship or geographical/family constraints (14.4%, 13.3% of responses). Conclusions: Integrated MB-PhD training generates sustained academic productivity and leadership, but structural barriers threaten retention of graduates within clinical-academic careers. Protecting research time, stabilising funding and pay, and reducing geographic instability are needed to retain the clinician-scientists that health systems have already invested in training.
Satorres-Perez, E.; Castillo-Marco, N.; Igual, M.; Cordero, T.; Munoz-Blat, I.; Monfort-Ortiz, R.; Marcos-Puig, B.; Simon, C.; Garrido-Gomez, T.; Perales-Marin, A.
Show abstract
Background. In Europe, first-trimester combined screening with the Fetal Medicine Foundation (FMF) algorithm identifies women at increased risk of preeclampsia who may benefit from personalized aspirin prophylaxis. However, a substantial proportion of early-onset preeclampsia (EOPE) remains undetected at clinically acceptable specificity. Objective. To evaluate the first-trimester performance of MaiRa for early-onset preeclampsia (EOPE) risk stratification by benchmarking it against FMF screening in the same women, characterizing discordant patient-level classification profiles and exploring potential implementation strategies. Study Design. This secondary case-control analysis was nested within the prospective, multicentre PREMOM cohort [NCT04990141], which enrolled women with singleton pregnancies across 14 tertiary hospitals in Spain. First-trimester MaiRa and FMF risk estimates were evaluated in the same 126 pregnant women, comprising 99 uncomplicated controls and 27 EOPE cases, defined by disease onset before 34 weeks. Discrimination was compared using a stratified paired bootstrap analysis of the areas under the receiver-operating-characteristic curves. Performance was assessed at prespecified clinical thresholds, and detection rates were evaluated at fixed false-positive rates. Universal and contingent MaiRa implementation strategies were also evaluated. Results. MaiRa showed greater first-trimester discrimination for EOPE than FMF combined screening (AUC, 0.974 vs 0.900; P=.040) and consistently achieved higher detection rates across fixed false-positive rates. At false-positive rates of 5% and 10%, MaiRa detected 85.2% and 92.6% of EOPE cases, compared with 44.4% and 70.4% for FMF, respectively. Patient-level analysis demonstrated that MaiRa identified 12 of 27 EOPE cases (44.4%) classified as low risk by FMF; these pregnancies generally exhibited less abnormal conventional first-trimester profiles, including fewer maternal risk factors, lower mean arterial pressure and lower uterine artery pulsatility index, yet 8 of 12 (66.7%) subsequently developed severe EOPE. Exploratory implementation analyses showed that universal MaiRa screening achieved the highest EOPE detection, whereas a contingent strategy using FMF for triage and reflex MaiRa testing reduced molecular testing to 35.7% of pregnancies while maintaining 77.8% sensitivity and 97.0% specificity. Conclusion. MaiRa provided greater first-trimester discrimination for EOPE than conventional combined screening and detected additional pregnancies that later developed severe disease despite less abnormal conventional screening profiles. The findings suggest that maternal plasma cfRNA profiling captures biological alterations not fully reflected by combined first-trimester screening and support further prospective evaluation in an independent, unselected obstetric population. Key words: early-onset preeclampsia; first-trimester screening; cell-free RNA; liquid biopsy; Fetal Medicine Foundation algorithm; combined screening; risk stratification; aspirin prophylaxis.
Han, F.; Wang, J.; Shi, S.; Jin, M.; Ren, C.
Show abstract
IMPORTANCE: A recent meta-analysis showed that chemoimmunotherapy was associated with improved overall survival (OS) compared with immune checkpoint inhibitor (ICI) monotherapy for programmed death-ligand 1 (PD-L1) tumor proportion score (TPS) [≥] 50% advanced non-small-cell lung cancer (NSCLC). However, whether this benefit reflects chemotherapy effect or ICI heterogeneity remains unclear. OBJECTIVE: To reassess the survival benefit of adding chemotherapy to ICI monotherapy using agent-stratified comparisons anchored to chemotherapy. DATA SOURCES: The 24 phase 3 randomized clinical trials included in the original meta-analysis (search date, August 3, 2025). DATA EXTRACTION AND SYNTHESIS: Hazard ratios (HRs) for OS and progression-free survival (PFS) were extracted from each trial in the original meta-analysis. Two analytic frameworks were used: within-agent comparisons (same ICI in both chemoimmunotherapy and monotherapy) and across-agent comparisons (ICI in one treatment strategy only). For within-agent comparisons, a two-stage random-effects meta-analysis was conducted. In stage 1, ICI-specific HRs for chemoimmunotherapy and ICI monotherapy versus chemotherapy were pooled and their ratio was calculated (RHR = HRchemoimmuno/HRmono; RHR < 1 favors chemoimmunotherapy). The RHRs were pooled in stage 2. For across-agent comparisons, RHR was derived from pooled HRs by treatment strategy. MAIN OUTCOMES AND MEASURES: Endpoints were OS and PFS. RESULTS: In within-agent comparisons (4 ICIs; 13 trials; N = 3252), pooled RHR was 0.94 (95% CI, 0.78-1.13; P = .48; I2 = 0.0%) for OS and 0.85 (95% CI, 0.68-1.06; P = .14; I2 = 0.0%) for PFS. In across-agent comparisons (7 ICIs; 11 trials; N = 2231), RHR favored chemoimmunotherapy for OS (0.68; 95% CI, 0.50-0.92; P = .01) and PFS (0.46; 95% CI, 0.37-0.58; P < .001). In a sensitivity analysis restricted to trials of NCCN-recommended regimens, pooled RHR was 1.02 (95% CI, 0.81-1.28; P = .87) for OS. CONCLUSIONS AND RELEVANCE: In the within-agent comparisons, adding chemotherapy to ICI monotherapy did not improve OS or PFS in patients with PD-L1 TPS [≥] 50% advanced NSCLC. The benefit in the original meta-analysis appears driven by across-ICI heterogeneity. These findings are consistent with ICI monotherapy as a standard first-line option and underscore the need for agent-level stratification in across-trial comparisons.
Beukema, M.; Vermeulen, E.; de Vries-Idema, J.; Huckriede, A.; Joshi, M.
Show abstract
The increasing incidence of H5N1 influenza virus transmission from animal species to humans has heightened concerns about an imminent H5N1 pandemic. Prior studies using recombinant hemagglutinin and neuraminidase proteins have reported age-dependent cross-reactivity to H5N1, attributed to immune imprinting from an individual's first influenza virus exposure. However, whether this pattern holds when using whole inactivated virus (WIV), capturing antibodies against diverse viral proteins, and is stable over time remains unknown. We therefore aimed to determine whether H5N1 cross-reactivity of pre-existing antibodies to whole virus follows an age-dependent or imprinting-specific pattern, and whether this pattern is stable over a five-year period. To this end, we measured serum antibody levels in adolescents, adults and seniors by ELISA using whole inactivated H5N1 virus as antigen rather than purified proteins. Detectable, albeit generally low, levels of H5N1-reactive antibodies were present in most individuals, irrespective of age. Comparison of antibody levels against H5N1 with those to five historical influenza virus strains revealed a consistent positive correlation between H5N1-reactive antibodies and responses to the H1N1pdm09 strain A/California/7/2009 (CA), across all age groups. Using unbiased clustering of antibody titers against H5N1, CA, and the H3N2 strain A/Perth/16/2009 (PE), we identified seven distinct age-transcending antibody profiles. These profiles covered individuals with varying titers to all three included viruses but also identified individuals with high anti-CA levels, yet low anti-H5N1 levels and vice versa. Moreover, despite stable antibody levels over a five-year interval in the study population, individual antibody levels and profiles fluctuated considerably over this period. Taken together, our results confirm the presence of H5N1-reactive antibodies in human sera and their association with previously circulating strains. However, they also caution against inferring antibody levels against a new strain based solely on responses to antigenically related strains and highlight the limitations of extrapolating immune status from single timepoint measurements.
Montanez-Valverde, R. A.; Kim, V.; Duran-Luciano, P.; Yuan, Y.; Sofer, T.; Kaplan, R. C.; Gallo, L. C.; Talavera, G. A.; Perreira, K. M.; Daviglus, M. L.; Rosas, S. E.; Llabre, M. M.; Elfassy, T.; Li, X.; Isasi, C. R.; Rodriguez, C. J.
Show abstract
Background. The imprecision of current metrics to capture the complex genetic admixture and racial identity among Hispanic/Latino individuals in the United States [US] is a concern. We examined the relationship of self-reported race and genetic ancestry with hypertension [HTN] among Hispanics/Latinos. Methods. Cross-sectional study of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), including 10,586 Hispanic/Latino unrelated adults. Genetic ancestry: West African [AA], Amerindian [AI], and European [EA]. Self-reported race: White, Black, Native American, or Multiple/Missing (More than one race or Unknown/Not reported/Refused). HTN: systolic (SBP) [≥]130 mmHg, diastolic blood pressure (DBP) [≥]80 mmHg, and/or use of HTN medications. Age- and sex adjusted models were used. Results. Self-reported race was White (38{middle dot}6%), Black (3{middle dot}6%), Native American (4{middle dot}1%), and Multiple/Missing (53{middle dot}7%), with Unknown/Not reported/Refused representing 32{middle dot}7%. Black and White Hispanics/Latinos had the greatest AA (55{middle dot}7%) and EA (69{middle dot}3%) ancestries, respectively. Each 10% AA increase was associated with OR 1{middle dot}15, SBP beta +0{middle dot}9 mmHg, and DBP beta +0{middle dot}7 mmHg. Conversely, each 10% AI increase was associated with OR 0{middle dot}83, SBP beta -0{middle dot}4 mmHg, and DBP beta -0{middle dot}6 mmHg. HTN prevalence was highest among those with Black race or in the highest AA quantile (45{middle dot}6% and 48{middle dot}0%, respectively), and lowest among those with Native American race or in the highest AI quantile (37{middle dot}6% and 26{middle dot}7%, respectively). Conclusion. One-third of Hispanics/Latinos did not self-report race. Black or White self-reporting race did somewhat relate to AA or EA ancestry, respectively. HTN profiles were related to self-reported race and genetic ancestry in this admixed population.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Rakhimov, B.; Choi, J.; Kim, K.; Tuychiev, L.; Shadmanov, A.; Mamatkulov, B.
Show abstract
Background. The clinical course of coronavirus disease 2019 (COVID-19), and the ability to anticipate which patients will require intensive care, were poorly characterized in Central Asia during the first pandemic wave. We aimed to describe the clinical features of hospitalized COVID-19 patients at the Tashkent State Medical University, Uzbekistan, and to identify risk factors for intensive care unit (ICU) admission. Methods. In this single-centre cross-sectional study, we reviewed the records of 2500 consecutive patients hospitalized between 11 April and 8 August 2020. Patients were grouped as asymptomatic or symptomatic, and symptomatic patients were compared by ICU versus non-ICU status. Groups were compared with chi-square or Fisher's exact and Mann-Whitney U tests. Univariable and multivariable logistic regression identified risk factors for ICU admission. Results. Of 2500 patients (median age 36 years; 60.9% male), 989 (39.6%) were asymptomatic and 1511 (60.4%) symptomatic. In total, 129 (5.2%) were admitted to the ICU and 38 (1.5%) died. ICU patients were older (median 56 vs 40.5 years) and more often had bilateral pneumonia, oxygen desaturation and cardiometabolic comorbidity. In the multivariable model (AUC 0.82), the independent predictors of ICU admission were ischemic heart disease (aOR 4.20), shortness of breath (aOR 3.22), hypertensive heart disease (aOR 2.93) and male sex (aOR 2.00). Conclusions. Older age, cardiometabolic comorbidity and respiratory compromise identified patients at high ICU risk. As one of the first clinical COVID-19 descriptions from Uzbekistan, these data provide a baseline for preparedness in Central Asia.
Kim, S. S.; Zissette, S. Z.; Van Meter, C.; Shiiba, M.; Bruck, M.; Tippett, A.; Kamidani, S.; Benkeser, D.; McQuade, E. R.
Show abstract
Importance: Maternal vaccination and long-acting monoclonal antibodies are now available in the U.S. to prevent RSV. Long-acting monoclonal antibody administration in the U.S. commonly occurs after hospital discharge in outpatient settings, leaving some infants unprotected early in life when severe RSV risk is highest. Comparative effectiveness between the two interventions and whether delays affect effectiveness estimates have not been quantified. Objective: To evaluate the effectiveness of infant long-acting monoclonal antibody strategies and a maternal vaccination strategy, each compared to no intervention, and the comparative effectiveness of intervention strategies when accounting for real-world delays in monoclonal antibody receipt. Design: Cohort study using target trial emulation to compare four strategies for prevention of RSV-related outcomes. Setting: The U.S. between 2023 and 2025 using a nationwide database of employer-sponsored commercial insurance claims. Participants: 120,586 commercially insured mother-infants, whose infants were born in the U.S. during the 2023-2024 or 2024-2025 RSV season. Infants who could not be paired with their mother's record, did not enroll in commercial insurance within 75 days from birth, received palivizumab, and had an implausible birth date were excluded. Interventions: Comparison of four RSV prevention strategies: (i) maternal RSVpreF; (ii) long-acting monoclonal antibody given within the first week of life (mAb as intended); (iii) long-acting monoclonal antibody given within a six-month grace period from birth (mAb within grace period); and (iv) a control. Main outcomes and measures: Effectiveness against first RSV-associated hospitalization and medically-attended RSV illness was summarized using adjusted hazard ratios (aHR) and estimated using an inverse propensity weighting approach, with weights accounting for maternal age, maternal comorbidities affecting pregnancy, obstetric and newborn complications, season, region, and birth timing relative to October 1. A weighted Kaplan Meier estimator was used to estimate strategy-specific cumulative incidence of RSV outcomes over time. Results: In the first five weeks of life, the mAb within grace period strategy doubled the hazard of RSV hospitalization (aHR: 2.0 [95% CI: 1.0-4.9]) and increased the hazard of medically-attended RSV (aHR: 1.6 [95% CI: 1.0-2.7]) compared to the maternal RSVpreF strategy. The hazard for RSV hospitalization was similar for the mAb as intended strategy compared to the maternal RSVpreF strategy (aHR = 0.9 [95% CI: 0.3-1.9]). Conclusions and relevance: RSVpreF and monoclonal antibodies were similarly effective when monoclonal antibodies were administered close to birth, but when accounting for real-world delays in monoclonal antibody receipt, the maternal RSVpreF strategy was more effective than the mAb within grace period strategy.
Takeuchi, J. S.; Kurokawa, M.; Yamamoto, K.; Yamanaka, J.; Morino, E.; Takayanagi-Nishisako, S.; Ohmagari, N.; Sugiura, W.; Kimura, M.
Show abstract
Background The COVID-19 pandemic substantially altered respiratory pathogen circulation worldwide. However, longitudinal analyses of changes in respiratory pathogen ecology across the pandemic and post-pandemic periods remain limited. Methods We analyzed 19,968 respiratory samples tested with the BioFire(R) FilmArray(R) Respiratory Panel at a hospital in Tokyo, Japan, between January 2020 and March 2026. We evaluated temporal changes in pathogen circulation, age-specific epidemiology, co-detection patterns, pairwise pathogen associations, and clinical parameters. Results At least one respiratory pathogen was detected in 27.8% of tests. Respiratory pathogens resurged asynchronously following the relaxation of COVID-19-related public health measures. Influenza virus circulation remained markedly suppressed until late 2022 before re-emerging in successive large seasonal epidemics, whereas other pathogens, including RSV, human metapneumovirus, and Mycoplasma pneumoniae, exhibited distinct resurgence patterns. Pathogen distributions also varied by age. Human rhinovirus/enterovirus remained predominant among young children, whereas SARS-CoV-2 predominated among older adults. Co-detection occurred in 14.0% of positive specimens and was significantly more frequent in younger patients. Pairwise analysis identified both positive and negative pathogen associations; however, the patterns varied across age groups and study periods. Conclusions Respiratory pathogen circulation changed substantially during the transition from the COVID-19 pandemic to the post-pandemic period, with pathogen-specific, age- and period-dependent patterns. Continued surveillance is warranted to determine how respiratory pathogen circulation will evolve and to inform infection control strategies in the post-pandemic era.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Bandini, V.; Whitaker, L. H.; Vincent, K.; Salmeri, N.; Mawson, R.; Vercellini, P.; Horne, A. W.
Show abstract
Background: Endometriosis is a chronic pain condition in which hormonal therapies form the cornerstone of long-term management. Treatment tolerability is critical for adherence and therapeutic success, but most comparative studies and reviews have focused on their ability to reduce menstrual pain, while their impact on non-menstrual pelvic pain (NMPP), bleeding patterns, adverse events (AEs), treatment discontinuation and quality of life (QoL) remain poorly characterised. This systematic review and meta-analysis evaluate these outcomes across currently available hormonal therapies, providing practical evidence for clinical decision-making. Methods: PubMed/MEDLINE, Scopus, and Embase were searched up to November 2025 for randomised controlled trials comparing at least two active first- or second-line hormonal treatments for endometriosis. Studies without confirmed endometriosis, treatment duration less than three months and comparing therapies to placebo only were excluded. Data were extracted by two reviewers from reports. Pain outcomes were pooled as mean differences (MD, 95% CI), with bleeding patterns, AEs, and discontinuations as proportions. Analyses were performed in R. PROSPERO: CRD420251137785. Findings: Of 1892 records screened, 48 trials (5583 women) met our inclusion criteria. Overall pelvic pain (0-10 scale) was significantly reduced across all treatment categories (p<0.001): combined oral contraceptives (COCs) (MD 3.17), oral and long-acting progestogens (MD 3.83; MD 4.29), and GnRH-analogues (MD 3.81). Sensitivity analyses restricted to studies reporting NMPP yielded comparable results. GnRH-agonists showed the most favourable bleeding profile, followed by continuous COCs. However, all regimens reported class-specific AEs, including mood changes, nausea, headache, weight gain, and decreased libido (pooled proportions >10%). Overall discontinuation due to AEs was 7.7%, and vaginal bleeding was the leading cause. Heterogeneity across meta-analyses was high. Risk of bias (RoB2) was moderate to high. Interpretation: Given similar reductions in overall pelvic pain across hormonal therapies, treatment decisions should prioritise differences in bleeding profiles, therapy-specific AEs, and QoL. Funding: None.