Back

Nature Medicine

Springer Science and Business Media LLC

All preprints, ranked by how well they match Nature Medicine's content profile, based on 125 papers previously published here. The average preprint has a 0.13% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Learning the shared structure of human health across diseases, modalities, and time

Hager, P.; Roth, B.; Buehler, N.; Lu, D.; Burks, J. H.; Shilova, L.; Kfuri-Rubens, R.; Roellin, E.; Pan, J.; Di Folco, M.; Chan, E.; Schnabel, J. A.; Adams, L.; Rueckert, D.; Theis, F. J.; Casale, F. P.

2026-07-09 health informatics 10.64898/2026.07.07.26357373 medRxiv
Top 0.1%
33.0%
Show abstract

Human disease risk emerges from the shared influences of genetics, environment, lifestyle, and concurrent diseases over time, resulting in recurring patterns of susceptibility across conditions. However, most risk prediction models treat diseases as independent outcomes or rely on limited input variables, restricting their ability to capture these shared patterns. Here we present RisQ, a framework that learns a unified representation of human health across diseases, modalities, and time. This representation is queried with natural language to estimate disease risk for arbitrary diseases and prediction horizons. Generalization to unseen disease groups and prediction horizons indicates that information is shared across diseases and time, revealing a common structure of disease risk that is learnable. Trained and validated in 488,170 participants from the UK Biobank and evaluated without retraining in 257,538 participants from the independent All of Us cohort, RisQ leverages this shared structure to outperform disease-specific models, multi-disease frameworks, and tabular foundation models in risk prediction. We show that jointly modeling increasing numbers of diseases, input modalities, and prediction horizons improves performance, indicating that scaling these axes increases information transfer and enriches the learned structure. We then show this structure is multi-scale: it captures demographic determinants of disease susceptibility, while also organizing individuals into reproducible cross-disease risk clusters within demographically restricted subgroups. Genetic analyses further support the biological grounding of the structure by linking gene-level loss of function to cross-disease risk profiles. This surfaces known relationships of HBB, SLC22A12, CASR, and LDLR, while also highlighting less characterized associations. Together, these results indicate that human disease risk exhibits a shared structure that can be learned from multimodal data to improve risk prediction, stratify individuals by cross-disease susceptibility, and support the discovery of relationships across diseases.

2
A deep learning transformer model predicts high rates of undiagnosed rare disease in large electronic health systems

Jordan, D. M.; Vy, H. M. T.; Do, R.

2023-12-24 health informatics 10.1101/2023.12.21.23300393 medRxiv
Top 0.1%
31.3%
Show abstract

It is estimated that as many as 1 in 16 people worldwide suffer from rare diseases. Rare disease patients face difficulty finding diagnosis and treatment for their conditions, including long diagnostic odysseys, multiple incorrect diagnoses, and unavailable or prohibitively expensive treatments. As a result, it is likely that large electronic health record (EHR) systems include high numbers of participants suffering from undiagnosed rare disease. While this has been shown in detail for specific diseases, these studies are expensive and time consuming and have only been feasible to perform for a handful of the thousands of known rare diseases. The bulk of these undiagnosed cases are effectively hidden, with no straightforward way to differentiate them from healthy controls. The ability to access them at scale would enormously expand our capacity to study and develop drugs for rare diseases, adding to tools aimed at increasing availability of study cohorts for rare disease. In this study, we train a deep learning transformer algorithm, RarePT (Rare-Phenotype Prediction Transformer), to impute undiagnosed rare disease from EHR diagnosis codes in 436,407 participants in the UK Biobank and validated on an independent cohort from 3,333,560 individuals from the Mount Sinai Health System. We applied our model to 155 rare diagnosis codes with fewer than 250 cases each in the UK Biobank and predicted participants with elevated risk for each diagnosis, with the number of participants predicted to be at risk ranging from 85 to 22,000 for different diagnoses. These risk predictions are significantly associated with increased mortality for 65% of diagnoses, with disease burden expressed as disability-adjusted life years (DALY) for 73% of diagnoses, and with 72% of available disease-specific diagnostic tests. They are also highly enriched for known rare diagnoses in patients not included in the training set, with an odds ratio (OR) of 48.0 in cross-validation cohorts of the UK Biobank and an OR of 30.6 in the independent Mount Sinai Health System cohort. Most importantly, RarePT successfully screens for undiagnosed patients in 32 rare diseases with available diagnostic tests in the UK Biobank. Using the trained model to estimate the prevalence of undiagnosed disease in the UK Biobank for these 32 rare phenotypes, we find that at least 50% of patients remain undiagnosed for 20 of 32 diseases. These estimates provide empirical evidence of a high prevalence of undiagnosed rare disease, as well as demonstrating the enormous potential benefit of using RarePT to screen for undiagnosed rare disease patients in large electronic health systems.

3
A Multimodal Framework for Organ- and Cell-Resolved Biological Aging and Longevity Intervention Discovery

Al Dajani, S. A.; Williams, J. R.; Fuentealba, M.; Zhai, T.; Furman, D.; Snyder, M.; Abudayyeh, O. O.; Gootenberg, J. S.; Gladyshev, V. N.

2026-05-12 geriatric medicine 10.64898/2026.05.08.26352759 medRxiv
Top 0.1%
31.2%
Show abstract

Aging is the primary driver of chronic disease and mortality, requiring comprehensive frameworks for quantification of aging and nomination of longevity interventions. We developed mAge (multimodal age), a biological aging framework that integrates plasma proteomics, wearables, and mortality hazard to predict biological age, intrinsic capacity, and mortality risk. By combining proteomic and wearable data in UK Biobank samples, mAge exceeds unimodal baseline age prediction to 0.87 test R{superscript 2} and 2.3 years mean error, and reduces unimodal baseline mortality prediction error by 21%. We further constructed organ-and cell type-specific biological clocks that quantify aging across 49 distinct subsystems, revealing that cardiac, immune, and intracellular protein signatures benefit most from wearable integration. By mapping data to FDA-approved drug targets, we identified interventions, such as GLP-1 receptor agonists, gabapentin, and ACE inhibitors, that are associated with lower overall and subsystem-specific proteomic age and mortality risk or are associated with longer time-to-death and later age-at-death in longitudinal and deceased cohorts. mAge establishes a scalable framework for nominating and validating personalized longevity interventions, bridging continuous digital monitoring with molecular aging diagnostics.

4
Combined values alignment and epistemic verification prevent delusional reinforcement in conversational AI agents

Carrano, A.; Patel, M. S.; Hartono, S.; Ekker, S. C.

2026-06-02 health informatics 10.64898/2026.05.29.26354389 medRxiv
Top 0.1%
31.1%
Show abstract

Conversational AI is being deployed into medical decision support, mental-health triage, and social companionship, where reinforcement of a user's false or delusional belief can cause direct harm. Most deployed safety techniques are evaluated for factual accuracy in isolation; the question of whether they protect against belief-level harm, and whether layered architectures behave additively or synergistically, has not been answered empirically. We compared four configurations of the same underlying model: a bare language model (condition A); an explicit values constraint we call the First Law architecture (condition B); a real-time epistemic verification layer called Aletheia (condition C); and the complete architecture combining all components together (condition D). Across 156 scored responses spanning 39 probe items in four belief-harm domains, condition A only passed 3 of 36 main-battery probes (8.3%; 95% CI 1.8 to 22.5%) under triple-blind human consensus rating demonstrating the core limitations of unmodified LLM deployments. In contrast, the three safety architectures (B-D) passed at least 97% of items (Fisher's exact, P < 0.001 versus A). On a synergy battery designed to test items at the intersection of value- and epistemic-domain failures (16 scored items, AI-rated), only the complete architecture passed every item; single-layer conditions failed on 7 of 16 items (43.8%) where neither values constraint nor verification was individually sufficient. Linear mixed-effects modelling of three-turn emotional escalation gave a slope of -1.00 points per turn for the values-only condition (t = -6.20) and -0.75 points per turn for the verification-only condition (t = -4.65); the complete architecture was flat at {beta} = 0.00. We describe a mechanistic failure of single-layer verification we call bot-validates-kernel-endorses-inference, in which accurate confirmation of a true factual element embedded in a delusional claim transfers epistemic authority to the surrounding false inference. Values alignment and factual verification address different failure modes, and the combined VaaS-Aletheia architecture is what produces stable protection across emotional escalation in conversational settings. The complete architecture evaluated here represents evidence-based specification for safer deployment of AI in high-stakes advisory contexts and serves as a benchmark against which future safety architectures can be compared.

5
Increased burden of influenza A/H1N1pdm09 in older adults following the COVID-19 pandemic

de Jong, S. P. J.; Russell, C. A.

2026-05-28 infectious diseases 10.64898/2026.05.20.26353664 medRxiv
Top 0.1%
28.3%
Show abstract

Of the two influenza A virus (IAV) subtypes circulating endemically in humans, A/H3N2 and A/H1N1pdm09, A/H3N2 has historically been the dominant driver of disease burden in older adults. Based on an analysis of publicly available global surveillance data from 2015 to 2025 (>300,000 subtyped, age-stratified infections), we report a substantially increased contribution of A/H1N1pdm09 to influenza morbidity in older adults since approximately 2022. Birth cohort-stratified analyses suggest elevated A/H1N1pdm09 burden among individuals born before 1955-1959, consistent with erosion of pre-existing immunity originally generated by exposure to historical A/H1N1 strains. Pooled estimates across datasets and analytical approaches indicate the increase in A/H1N1pdm09 burden rises with earlier birth year, ranging from 1.22-fold (95% CI 1.08-1.37) for the 1955-1959 birth cohort to 3.10-fold (95% CI 2.58-3.72) for the 1930-1934 cohort. These findings point to a substantial rise in the overall influenza burden among the most vulnerable age groups, with implications for vaccine policy, clinical management, and public health planning.

6
Whole-blood transcriptomic traces of organ pathology and their causal triage

Sinha, R. K.; Alvarez, K.; Sinha, S.

2026-08-03 health informatics 10.64898/2026.07.31.26359425 medRxiv
Top 0.1%
26.9%
Show abstract

Whole blood offers a non-invasive window into organ health, yet it remains unclear which organ pathologies leave a detectable trace in the blood transcriptome, and whether such traces are causal or reactive. Progress has been limited because paired whole-blood profiles and pathologist-graded organ pathology are rarely available together. Here we assembled a pathology-linked benchmark of 59 pathologies across 20 organs in 803 GTEx v10 donors with matched whole-blood RNA-seq and postmortem histology. Because a postmortem cohort's blood is strongly shaped by age, sex, and the circumstances of death, our framework, TRACE, counts a signal only when blood expression predicts a pathology beyond these donor factors. Four pathologies passed, with liver cirrhosis by far the strongest (AUC 0.79). The cirrhosis signature replicated in an independent cohort of living patients (AUC 0.80) and remained specific against severe systemic illness. To separate candidate drivers from reactive markers, we used human genetics: Mendelian randomization linking plasma proteins to liver disease, which recovered established fibrosis drivers, including PAI-1 (SERPINE1), tenascin-C, thrombospondin-2 and nominated further candidates. Together, TRACE provides a resource and a confounder-aware framework for learning which organ pathologies the blood transcriptome can, and cannot, detect, and which of those signals are likely causal.

7
FeverIQ - A Privacy-Preserving COVID-19 SymptomTracker with 3.6 Million Reports

Ranjan, A.; Li, S.; Chen, B.; Chiu, A.; Jagadeesh, K.; Liphardt, J.

2020-09-25 public and global health 10.1101/2020.09.23.20200006 medRxiv
Top 0.1%
26.9%
Show abstract

Population-scale COVID-19 management benefits from timely and honest information from billions of people. Here, we provide a first report on the FeverIQ symptom tracker, a global effort to collect symptom and test data which has received more than 3.6 million submissions. Unlike other trackers, FeverIQ uses secure multiparty computation (SMC) to cryptographically guarantee user privacy while providing insights to scientists and public health efforts. We performed basic integrity checks of the FeverIQ dataset, such as by comparing it to other publicly released data. We then trained a linear classifier on diagnosis scores which were computed securely, without unprotected symptom data ever leaving a users phone or computer. FeverIQ is currently the worlds largest application of SMC in a health context, demonstrating the practicality of privacy-preserving analytics for population-scale digital health interventions.

8
Cross-database validation reveals distinct layers of transportability in ICU delirium prediction

Ni, S.; Sato, K.

2026-07-21 health informatics 10.64898/2026.07.19.26358409 medRxiv
Top 0.1%
26.9%
Show abstract

External validation of clinical AI emphasizes discrimination, although deployment requires the endpoint, probability estimates and operating policy to transport. Here we show that these layers diverged in retrospective bidirectional evaluation of five model families across eICU and MIMIC-IV. Coarse-label AUROC fell from 0.87-0.92 internally to 0.66-0.83 during source-only transfer. For assessment-conditioned repeated monitoring of persistence or recurrence, external AUROC reached 0.76-0.94, but removing assessment history reduced it by 0.16-0.32; broader features did not help consistently. Transported scores concentrated future-positive ICU stays 2.4-6.9-fold in the top risk decile. Development-selected cutoffs alerted 0.3-2.0% of prediction rows and captured 9.2-11.0% of future-positive rows; after deduplication, 4.9-12.2% of stays were alerted, capturing 43.9-49.4% of future-positive stays. Thus, ranking can persist while probability and policy transport remain site dependent. Layered validation is a prerequisite for prospective evaluation, not evidence of clinical benefit.

9
Distance of covariance (DISCO), a novel measure of network homeostatic dysregulation, reveals organ system interconnections underlying mortality and disease risk

Hao, M.; Zhang, H.; Li, Y.; Huang, Y.; Wu, J.; Zhang, S.; Hu, Z.; Li, X.; Jiang, S.; Tanner, K. T.; Chen, J.; Bao, Z.; Wang, J.; Cohen, A. A.; Huang, Y.; Jin, L.; Wang, X.

2025-05-07 geriatric medicine 10.1101/2025.05.06.25327108 medRxiv
Top 0.1%
26.4%
Show abstract

Aging manifests as the progressive declines of homeostatic resilience and repair mechanisms, marked by dysregulations across systems and increasing individual heterogeneity. However, the breadth of measures of homeostatic dysregulation remains underexplored. Here, we introduce DISCO as a novel measure of homeostatic dysregulation, integrating clinical, proteomics, metabolomics, and microbiomes data. DISCO demonstrated moderate correlation with chronological age but robustly predicted mortality, frailty, and chronic disease risk, outperforming Mahalanobis distance in health outcome prediction, comparable to the best epigenetic clocks. Organ/tissue-specific DISCO analysis revealed limited organ-disease specificity, suggesting systemic rather than localized dysregulation drives health decline. Network analysis identified aging-associated proteins as central hubs strongly linked to DISCO scores; further, organ-level DISCO metrics most predictive of age and outcomes were also central within biological networks. Collectively, DISCO emerges as a validated measure of whole-body homeostatic dysregulation, providing a tool for aging risk stratification and insights into systemic aging mechanisms.

10
Dual foundation models for accelerometry predict future health

Dige, M.; Lorenzen, N. R.; Kjer, M. R.; Burns, A. C.; Jennum, P. J.; During, E. H.; Zou, J.; Mignot, E.; Brink-Kjaer, A.

2026-07-27 health informatics 10.64898/2026.07.24.26358894 medRxiv
Top 0.1%
26.2%
Show abstract

Wrist accelerometers are ubiquitous and capture activity, sleep, and cardiorespiratory motion, but how this relates to future disease across the phenome is unclear. We encoded one week of UK Biobank accelerometry from 97,696 participants using two frozen self-supervised models and trained a multilabel survival model on these embeddings with participant age and sex for 390 outcomes. In 5,253 held-out participants, mean concordance was 0.688. A single component explained 76% of predicted risk variance and was associated with future disease burden and mortality. Nevertheless, disease-specific scores added discrimination beyond this shared axis for 85 of 101 well-powered outcomes. A higher-powered full-cohort out-of-fold analysis identified prodromal neurodegenerative signatures, strongest for Parkinson's disease (five-year time-dependent AUROC 0.90; 428 cases), with limited attenuation after lead-time washout. Daytime movement contributed most, whereas sleep-related and genetic information contributed selectively. These findings establish one week of wrist movement as a scalable, low-cost representation of future health for wearable-based risk assessment.

11
Conserved neuroectodermal aging encodes primate health and longevity

Yang, S.; Xin, Z.; Wang, W.

2026-05-07 geriatric medicine 10.64898/2026.05.05.26352498 medRxiv
Top 0.1%
26.1%
Show abstract

Neuroectoderm-derived tissues are highly metabolically active and exhibit minimal regenerative turnover, rendering them uniquely vulnerable to age-related stress while preserving undiluted degenerative signals. Yet aging dynamics in these tissues remain elusive in living primates. Here, we introduce an in vivo neuroectodermal aging clock and trace its trajectory in 66,602 human adults and six rhesus macaques across nine health and disease cohorts using an in situ optical biopsy. Through a digital histology atlas integrated with artificial intelligence, we resolve tissue representations of neuroectodermal aging within the human retina, predominantly localized to the metabolically active ganglion and bipolar cell populations and the photoreceptor complex, while demonstrating their evolutionary conservation across primate species. Neuroectodermal aging predicts health and longevity, scales across space and time, and captures preclinical aging signals within and beyond the neuroectodermal compartment. This framework is further validated in a diabetic population, where robust prognostic and dynamic sensitivity are preserved across physiological and perturbed states. Our work establishes a scalable framework for resolving neuroectodermal aging in living primates and linking tissue-level vulnerability to systemic health trajectories.

12
Time-to-event modeling with multimodal clinical and genetic features improves risk stratification of liver complications in chronic hepatitis C

Islam, H.; Arian, A.; Franses, J. W.; Ahsan, H.

2026-03-09 health informatics 10.64898/2026.03.06.26347819 medRxiv
Top 0.1%
25.8%
Show abstract

Chronic hepatitis C (CHC) remains a leading cause of cirrhosis, hepatocellular carcinoma (HCC), and premature mortality despite effective antiviral therapy, underscoring the need for individualized risk stratification beyond fibrosis stage alone. Using harmonized data from the All of Us Research Program, we developed and internally validated an interpretable multimodal survival framework to predict incident cirrhosis, HCC, and all-cause mortality, explicitly accounting for competing death. Baseline predictors within a {+/-}180-day window around CHC diagnosis included demographics, comorbidities, medications, laboratory biomarkers, socioeconomic context, and selected germline variants. Penalized Cox, ensemble, gradient-boosted, and neural survival models were compared under a consistent training and held-out testing strategy. Best-performing models achieved test C-indices of 0.67 for cirrhosis (Coxnet-LASSO), 0.71 for HCC, and 0.75 for mortality (Random Survival Forest), with stable time-dependent AUROC up to 0.81. Substantial feature compression preserved discrimination: restricting to the top 50% or 25% of predictors resulted in minimal absolute change in test performance (3.5%). Reduced models were anchored in clinically interpretable domains, including age, liver injury markers, hepatic reserve, cardiometabolic burden, deprivation index, and chromosome 19/22 loci. Feature importance reinforces existing known clinical and biological risk factors for liver complications: liver injury markers were most influential for cirrhosis and HCC, whereas hepatic reserve and cardiometabolic burden were more predictive of mortality, with age serving as a central baseline determinant across outcomes. Together, these results support a scalable and parsimonious framework for individualized CHC risk stratification that integrates multimodal determinants.

13
Immune Correlates Analysis of a Single Ad26.COV2.S Dose in the ENSEMBLE COVID-19 Vaccine Efficacy Clinical Trial

Fong, Y.; McDermott, A. B.; Benkeser, D.; Roels, S.; Stieh, D. J.; Vandebosch, A.; Le Gars, M.; Van Roey, G. A.; Houchens, C. R.; Martins, K.; Jayashankar, L.; Castellino, F.; Amoa-Awua, O.; Basappa, M.; Flach, B.; Lin, B. C.; Moore, C.; Naisan, M.; Naqvi, M.; Narpala, S.; O'Connell, S.; Mueller, A.; Serebryannyy, L.; Castro, M.; Wang, J.; Petropoulos, C. J.; Luedtke, A.; Hyrien, O.; Lu, Y.; Yu, C.; Borate, B.; van der Laan, L. W. P.; Hejazi, N. S.; Kenny, A.; Carone, M.; Wolfe, D. N.; Sadoff, J.; Gray, G. E.; Grinsztejn, B.; Goepfert, P. A.; Little, S. J.; Paiva de Sousa, L.; Maboa, R.; Randh

2022-04-12 infectious diseases 10.1101/2022.04.06.22272763 medRxiv
Top 0.1%
22.8%
Show abstract

Anti-spike IgG binding antibody, anti-receptor binding domain IgG antibody, and pseudovirus neutralizing antibody measurements four weeks post-vaccination were assessed as correlates of risk of moderate to severe-critical COVID-19 outcomes through 83 days post-vaccination and as correlates of protection following a single dose of Ad26.COV2.S COVID-19 vaccine in the placebo-controlled phase of ENSEMBLE, an international, randomized efficacy trial. Each marker had evidence as a correlate of risk and of protection, with strongest evidence for 50% inhibitory dilution (ID50) neutralizing antibody titer. The outcome hazard ratio was 0.49 (95% confidence interval 0.29, 0.81; p=0.006) per 10-fold increase in ID50; vaccine efficacy was 60% (43, 72%) at nonquantifiable ID50 (< 2.7 IU50/ml) and rose to 89% (78, 96%) at ID50 = 96.3 IU50/ml. Comparison of the vaccine efficacy by ID50 titer curves for ENSEMBLE-US, the COVE trial of the mRNA-1273 vaccine, and the COV002-UK trial of the AZD1222 vaccine supported consistency of the ID50 titer correlate of protection across trials and vaccine types.

14
Clinical trajectories and genetic architecture across the neurological-psychiatric boundary

Kopal, J.; Smeland, O. B.; Hagen, E.; Amanzadi, A.; Erdos, B.; Fuhrer, J.; Shadrin, A. A.; Frei, O.; van der Meer, D.; O'Connell, K. S.; Dale, A. M.; Andreassen, O. A.

2026-08-10 health informatics 10.64898/2026.08.06.26359855 medRxiv
Top 0.1%
22.7%
Show abstract

Foundation models trained on health records are increasingly used to represent human disease, but whether their embeddings reflect biology is hard to establish. We validate disease trajectory embeddings from an attention-based transformer against an external signal: genome-wide genetic architecture. Across 19 neurological and psychiatric disorders, clinical trajectory similarity mirrors genetic similarity, and the model recovers the same neurological-psychiatric boundary that emerges from genetic data, including which disorders cross it. A model with no access to diagnostic labels or genetic data thus recovers biological structure it was never trained on.

15
Therapeutic potential of IL6R blockade for the treatment of sepsis and sepsis-related death: Findings from a Mendelian randomisation study

Hamilton, F. W.; Thomas, M.; Arnold, D. T.; Palmer, T. M.; Moran, E.; Mentzer, A. J.; Maskell, N. A.; Baillie, J. K.; Summers, C.; Hingorani, A. D.; MacGowan, A. P.; Khandaker, G. M.; Mitchell, R. E.; Davey Smith, G.; Ghazal, P.; Timpson, N. J.

2022-07-15 infectious diseases 10.1101/2022.07.14.22277638 medRxiv
Top 0.1%
22.4%
Show abstract

IntroductionSepsis is characterised by dysregulated, life-threatening immune responses, which are thought to be driven by cytokines such as interleukin-6 (IL-6). Genetic variants in IL6R known to downregulate IL-6 signalling are associated with improved COVID-19 outcomes, a finding later confirmed in randomised trials of IL-6 receptor antagonists (IL6RA). We hypothesised that blockade of IL6R could also improve outcomes in sepsis. MethodsWe performed a Mendelian randomisation analysis using single nucleotide polymorphisms (SNPs) in and near IL6R to evaluate the likely causal effects of IL6R blockade on sepsis, sepsis severity, other infections, and COVID-19. We weighted SNPs by their effect on CRP and combined results across them in inverse variance weighted meta-analysis, proxying the effect of IL6RA. Our outcomes were measured in UK Biobank, FinnGen, the COVID-19 Host Genetics Initiative (HGI), and the GenOSept and GainS consortium. We performed several sensitivity analyses to test assumptions of our methods, including utilising variants around CRP in a similar analysis. ResultsIn the UK Biobank cohort (N=485,825, including 11,643 with sepsis), IL6R blockade was associated with a decreased risk of sepsis (OR=0.80; 95% CI 0.66-0.96, per unit of natural log transformed CRP decrease). The size of this effect increased with severity, with larger effects on 28-day sepsis mortality (OR=0.74; 95% CI 0.38-0.70); critical care admission with sepsis (OR=0.48, 95% CI 0.30-0.78) and critical care death with sepsis (OR=0.37, 95% CI 0.14 - 0.98) Similar associations were seen with severe respiratory infection: OR for pneumonia in critical care 0.69 (95% CI 0.49 - 0.97) and for sepsis survival in critical care (OR=0.22; 95% CI 0.04- 1.31) in the GainS and GenOSept consortium. We also confirm the previously reported protective effect of IL6R blockade on severe COVID-19 (OR=0.69, 95% 0.57 - 0.84) in the COVID-19 HGI, which was of similar magnitude to that seen in sepsis. Sensitivity analyses did not alter our primary results. ConclusionsIL6R blockade is causally associated with reduced incidence of sepsis, sepsis related critical care admission, and sepsis related mortality. These effects are comparable in size to the effect seen in severe COVID-19, where IL-6 receptor antagonists were shown to improve survival. This data suggests a randomised trial of IL-6 receptor antagonists in sepsis should be considered.

16
A Longitudinal Clinical Foundation Model on Nationwide Veteran Health Trajectories

Zamora-Resendiz, R.; Yin, J.; Kimbrel, N. A.; Beckham, J. C.; Crivelli, S.

2026-05-17 health informatics 10.64898/2026.05.13.26353133 medRxiv
Top 0.1%
22.0%
Show abstract

We present VA-LLM, a 1.62-billion-parameter autoregressive transformer pre-trained from scratch on 1.74 trillion tokens of clinical text spanning 22 years of care for 13.8 million patients in the Veterans Health Administration, with mortality outcomes confirmed through the National Death Index for 7.8 million patients. In a retrospective-prospective evaluation on 107,555 withheld patients, VA-LLM achieved higher 5-year AUPRC than Llama-2 (7 billion parameters), BioGPT _large (1.57 billion parameters), and GatorTron (3.91 billion parameters), matching GatorTron's 100,000-patient performance with only 10,000 labeled patients. In a clinical validation against the VA's operational Care Assessment Need (CAN) score on 5.5 million patients one year beyond the pre-training corpus, VA-LLM achieved a 90-day mortality AUROC of 90.00% versus 87.74% (p < 0.001) and a 45% relative improvement in AUPRC; post-hoc recalibration recovered calibration comparable to CAN (Brier 0.0091 versus 0.0093) without sacrificing discrimination. Across 21 pre-training checkpoints, discriminative performance correlated more strongly with cumulative mortality experience (CME), the total person-years contributed by patients with confirmed deaths, than with token count ({Delta}R2 = 0.15; Williams p < 10-6). Performance plateaued once marginal cohorts added fewer confirmed deaths, even as pre-training loss continued to decrease. These findings suggest that the clinical composition of pre-training data, particularly the completeness of documented patient trajectories, correlates with predictive performance more closely than corpus size alone.

17
Large Language Models for Psychiatric Phenotype Extraction from Electronic Health Records

Frydman-Gani, C.; Arias, A.; Perez Vallejo, M.; Londono Martinez, J. D.; Valencia-Echeverry, J.; Castano, M.; Bui, A. A. T.; Freimer, N. B.; Lopez-Jaramillo, C.; Olde Loohuis, L. M.

2025-08-12 psychiatry and clinical psychology 10.1101/2025.08.07.25333172 medRxiv
Top 0.1%
22.0%
Show abstract

The accurate detection of clinical phenotypes from electronic health records (EHRs) is pivotal for advancing large-scale genetic and longitudinal studies in psychiatry. Free-text clinical notes are an essential source of symptom-level information, particularly in psychiatry. However, the automated extraction of symptoms from clinical text remains challenging. Here, we tested 11 open-source generative large language models (LLMs) for their ability to detect 109 psychiatric phenotypes from clinical text, using annotated EHR notes from a psychiatric clinic in Colombia. The LLMs were evaluated both "out-of-the-box" and after fine-tuning, and compared against a traditional natural language processing (tNLP) method developed from the same data. We show that while base LLM performance was poor to moderate (0.2-0.6 macro-F1 for zero-shot; 0.2-0.74 macro-F1 for few shot), it improved significantly after fine-tuning (0.75-0.86 macro-F1), with several fine-tuned LLMs outperforming the tNLP method. In total, 100 phenotypes could be reliably detected (F1>0.8) using either a fine-tuned LLM or tNLP. To generate a fine-tuned LLM that can be shared with the scientific and medical community, we created a fully synthetic dataset free of patient information but based on original annotations. We fine-tuned a top-performing LLM on this data, creating "Mistral-small-psych", an LLM that can detect psychiatric phenotypes from Spanish text with performance comparable to that of LLMs trained on real EHR data (macro-F1=0.79). Finally, the fine-tuned LLMs underwent an external validation using data from a large psychiatric hospital in Colombia, the Hospital Mental de Antioquia, highlighting that most LLMs generalized well (0.02-0.16 point loss in macro-F1). Our study underscores the value of domain-specific adaptation of LLMs and introduces a new model for accurate psychiatric phenotyping in Spanish text, paving the way for global precision psychiatry.

18
Case-level artificial intelligence for multi-photo teledermatology submissions: development and internal validation using patient-submitted dermatology images

Patel, V. P.; Sheth, N.; Patel, A.; Patel, Y.

2026-06-01 dermatology 10.64898/2026.05.21.26353816 medRxiv
Top 0.1%
21.9%
Show abstract

Background: Store-and-forward teledermatology commonly relies on several patient-submitted photographs of the same concern, but most dermatology artificial intelligence models classify single images independently. Objective: To develop and internally validate a case-level diagnostic-support model that aggregates multiple patient-submitted photographs for common dermatologic conditions. Methods: We conducted a retrospective diagnostic-modeling study using the Skin Condition Image Network, a public dataset of deidentified self-taken dermatology images from US adults. We curated 2,336 cases comprising 5,041 images across 10 common inflammatory, allergic, and infectious conditions. Cases were split at the submission level into training, validation, and held-out test sets. Frozen general-purpose and dermatology-specific encoders were compared with image-level classifiers and a gated-attention multiple instance learning model that generated one case-level output from 1-3 images. Results: The strongest image-level baseline, dermatology-specific embeddings with random forest classification, achieved macro/micro ROC-AUCs of 0.797/0.854. Case-level aggregation improved discrimination, with dermatology-specific embeddings plus multiple instance learning achieving mean macro/micro ROC-AUCs of 0.819/0.863 across repeated stratified experiments. The locked final model achieved macro/micro ROC-AUCs of 0.800/0.849 on the held-out test set. Balanced-threshold sensitivity/specificity examples were 0.702/0.688 for eczema and 0.818/0.826 for urticaria. Limitations: Internal validation used a 10-condition subset from a US volunteer dataset; external validation, calibration, subgroup performance analysis, and prospective workflow studies are required. Conclusion: Modeling the teledermatology submission as a multi-image case better reflects asynchronous dermatology workflow than single-image classification. The model is preliminary clinician-facing support for structured review and triage, not autonomous diagnosis.

19
Circulating protein profiling identifies prognostic biomarkers in amyotrophic lateral Sclerosis

Klimovski, H.; Weinreich, M.; Strange, A.; Magen, I.; Alhathli, E.; Cohen, Y.; Melamed-Kadosh, D.; Lester, D. G.; Taylor, A.; Zhou, Y.; Ziv, T.; Abramovich, B.; Subramanian, A.; Perlson, E.; Drory, V.; Admon, A.; Shaw, P.; Malaspina, A.; Turner, M. R.; Talbot, K.; Cooper-Knock, J.; Thompson, A. G.; Hornstein, E.

2026-07-23 neurology 10.64898/2026.07.20.26354798 medRxiv
Top 0.1%
21.8%
Show abstract

In the pathologically and clinically heterogeneous neurodegenerative disorder amyotrophic lateral sclerosis (ALS), objective biochemical predictors of survival are essential to handle complexity in clinical trials, enrich clinical decision-making and interrogate the biology of disease progression. In this longitudinal study, we performed high-depth proximity extension assay proteomics using 1,095 samples of serum (N=851) and CSF (N=244) from 426 people with ALS, with orthogonal replication in an external cohort of 349 people with ALS. Age- and sex-adjusted Cox analysis identified 57 proteins in serum, including neurofilament light chain (NEFL) and peripherin, as well as five proteins in CSF, including tropomyosin 3 (TPM3) that were associated with survival (FDR-adjusted p<0.05). Penalised Cox regression identified a panel of 9 serum proteins - including NEFL, peripherin, TNF receptor superfamily member 27 (EDA2R) and calcitonin - that reflect the extent of disease as well as the progression rate, improving survival prediction compared with models using clinical parameters and NEFL. Joint modelling identified associations between the longitudinal trajectories of serum EDA2R and calcitonin with survival, highlighting their potential role in measuring disease progression. This work indicates the utility of multiple proteins reflecting diverse biological pathways in refining survival stratification and highlights systemic factors in ALS progression.

20
Diversification without convergence: national childhood respiratory pathogen spectra diverge as they diversify, 1990-2023

Li, D.; Feng, Q.; Zhang, Y.; Chen, H.; Wang, X.; Shen, C.

2026-09-03 pediatrics 10.64898/2026.09.01.26361890 medRxiv
Top 0.1%
21.7%
Show abstract

Background National childhood respiratory pathogen spectra are diversifying nearly everywhere - within-country diversity rose in 203 of 204 countries between 1990 and 2023 - yet whether countries are diversifying toward a common spectrum or along divergent paths is unknown. We quantified between-country compositional distance of national pathogen spectra over the same period. Methods We built national pathogen share vectors from Global Burden of Disease Study 2023 lower respiratory infection etiologic attributions (26 pathogens, 204 countries, ages 0-19 years) at five timepoints spanning 1990-2023. Between-country distance was measured as all pairwise Jensen-Shannon divergences (JSD; primary) and Bray-Curtis dissimilarities, with Baselga and Jaccard decompositions; robustness was assessed across metrics, pathogen panels, low-count thresholds and a balanced panel of 107 countries. Results Mean pairwise JSD rose from 0.0084 in 1990 to 0.0283 in 2023 (+238%; trend p = 0.030), peaking in 2021 (+283%) with a partial 2023 pullback. Bray-Curtis dissimilarity rose +120% and the balanced panel +423%. Divergence was entirely balanced variation (share reallocation), with spectrum richness rising from 18.5 to 21.1 of 26 pathogens. Dispersion rose fastest for influenza (coefficient of variation 0.03 to 0.55) and respiratory syncytial virus (0.08 to 0.48). Within-region distance rose in every computable GBD super-region (five of seven): divergence occurs within regions, not between blocs. Conclusions National spectra are re-sorting along country-specific axes as vaccine-preventable dominance recedes at different speeds. Diversification is universal, but convergence is absent: the transition at the etiologic-spectrum level is asynchronous and path-dependent, with implications for empirical treatment policy and pathogen surveillance.