Med
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Med's content profile, based on 39 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Yang, S.; Chen, V. L.; Ng, W. H.; Zhang, S.; Qiu, S.; Zhu, J.; Hsieh, T. Y.-J.; Ji, F.; Yeo, Y. H.
Show abstract
Background Clinical data analysis typically requires statistical programming skills, whereas cloud-based artificial intelligence (AI) agents risk exposing sensitive patient records. We developed and functionally validated a privacy-preserving, zero-code conversational statistical analysis framework that translates natural-language clinical research requests into executable R workflows while strictly retaining raw patient data within local computing environments. Methods Orchestrated by the n8n engine, the system integrates the DeepSeek-Reasoner model with a Pinecone vector database for retrieval-augmented generation (RAG), grounding statistical selection in curated biostatistical guidance and R templates. Core functionalities include data schema perception, interactive data cleaning, requirements refinement, and local R code execution via a controlled command-line interface. System performance was evaluated by replicating a published prognostic model study on metabolic dysfunction-associated steatotic liver disease (MASLD). Findings All core analytical workflows, including data cleaning, multivariable Cox proportional hazards modeling, model diagnostics, and publication-ready tables and figures (e.g., baseline characteristics, Schoenfeld residuals, receiver operating characteristic curves, and forest plots), were executed solely through natural-language dialogues without manual coding. The external large language model actively clarified analytical prompts while receiving zero row-level patient data. Interpretation Decoupling remote cloud reasoning from local code execution lowers the technical threshold for clinicians conducting data-driven research while safeguarding data privacy. This architecture provides a practical, scalable, and reproducible framework for converting natural-language clinical questions into executable statistical workflows. Funding National Natural Science Foundation of China (82473291), Shaanxi Province "Three Qin Scholars" Innovation Team Project (2023001), and Fundamental Research Funds for the Central Universities (xtr062023003).
Xiao, H.; Jiang, N.; Zhang, T.; Yin, Z.; Gui, T.; Zhang, Z.; Shao, K.; Ge, J.; Wei, R.; Pan, J.; Ma, J.; Yang, L.; Zhao, Z.; Zhou, J.; Fan, J.; Jiang, Y.; Torr, P.; Zheng, S.; Wu, Y. C.; Gao, Q.
Show abstract
Clinical research advances slowly because its core tasks, from evidence synthesis to mechanistic validation, remain fragmented. We present MedGenesis, a clinical artificial intelligence (AI) scientist built on a world-model reasoning loop that jointly updates a Latent Hypothesis Space and a Latent Action Space under expected information gain (EIG), uncertainty reduction (UR), and a safety prior P(safe), and integrates longitudinal electronic health records (EHRs) via the Virtual Clinical Trajectory and Observation Representation (ViCTOR) for cohort retrieval, trajectory stratification, and time-to-event analysis. On two benchmarks - ClinicalResBench (1,697 expert-curated questions) and ClinicalRepBench (40 paper-reproduction tasks) - MedGenesis outperformed frontier language models and biomedical AI systems while reducing hallucination. Across 1 million patient observations spanning five clinical evidence formats, it generated traceable outputs across meta-analysis, randomized controlled trials, real-world trajectories, case-control studies, and case reports, with one wet-lab-coupled run nominating a 3-hydroxybutyrate - neutrophil axis modulating antitumor immunity. These results compress hypothesis-to-evidence cycles from years to hours, creating a continuous clinical discovery process.
Yang, S.; Xin, Z.; Wang, W.
Show abstract
Environmental exposures are major modifiable determinants of human aging, yet the evidence remains fragmented across organ-agnostic summaries and rarely confronts population inequity. Here we present an exposomic atlas of pan-organ aging in ~300,000 UK Biobank adults, mapping 164 environmental and behavioural exposures onto biological aging of the whole body and nine organ subsystems. Comprising 1,476 systematically tested exposure-subsystem associations, the atlas reveals that environmental effects on human aging are pervasively organ-specific, with 65.9% of exposures acting divergently across organ subsystems. This landscape resolves into nine navigable modules that preserve organ selectivity, predict 23 major age-related diseases, and expose distinct dimensions of health inequity. In-silico analyses further show that priorities for ameliorating aging are target-dependent rather than universal, diverge markedly from the whole-body ranking (Kendall's {tau} = 0.52 to 0.39), reorder substantially across population strata, with findings externally validated in an ethnically distinct cohort. The atlas establishes an organ-resolved and target-aware foundation for precision environmental health.
Li, J.; Pan, Y.; Han, Y.; Zhou, C.; Zhao, L.; He, Y.
Show abstract
Abstract The association between COVID-19 vaccination and Guillain-Barre syndrome (GBS) has been previously investigated with inconsistent results, largely due to limited data and lack of concurrent controls. To address this problem, a large longitudinal cohort study was conducted using National COVID Cohort Collaborative (N3C) data. While COVID-19 infection was associated with increased GBS occurrence, COVID-19 vaccination was associated with significantly reduced GBS risk relative to unexposed (unvaccinated and uninfected) control, corresponding to a 61% lower 30-day risk (incidence risk ratio: IRR = 0.39, P < 0.01), consistent with multivariable Cox regression showing a similar reduction (adjusted hazard ratio: aHR = 0.41, P < 0.01). This protective association was observed only among recipients of mRNA vaccines (BNT162b2: IRR = 0.38, P < 0.01; mRNA-1273: IRR = 0.24, P < 0.01), but not among recipients of adenoviral-vector vaccines (IRR = 1.38, P > 0.05). Prior COVID-19 vaccination also reduced infection-associated GBS risk. Additional factors associated with GBS risk included sex, vaccine dose, and pre-existing comorbidities such as stroke, neurological disorders, and autoimmune diseases. Overall, our N3C large-scale study provides evidence that COVID-19 mRNA vaccination reduces GBS risk, supporting the safety profile of mRNA vaccines and warranting further mechanistic investigation.
Fu, B.; DeSchepper, L. B.; Sun, J.; McKeithen-Mead, S. A.; Kapili, B.; Ochoa-Andersen, P.; Spencer, S. P.; Fardeen, T.; Ricardo, M.; El Kamari, V.; Sinha, S.; Relman, D. A.; Grembi, J. A.; Shalon, D.; Estrela, S.; Huang, K. C.
Show abstract
The human small intestine (SI) plays a central role in nutrient processing, host-microbe interactions, and immune regulation, yet remains poorly characterized due to the lack of minimally disruptive sampling methods. Here, we present a protocol for deploying, recovering, and analyzing samples collected using an ingestible device that enables multi-region, lumen-targeted SI sampling during normal digestion. The device incorporates a ~30-cm collapsible tube wound into pH- or time-responsive layers that sequentially unfurl in situ, typically capturing three spatially ordered samples with high yield and reliable retrieval. This protocol outlines study design, participant handling, device recovery, contamination control, and standardized workflows for analyses, including cell quantification, culturomics, sequencing, and metabolomics. We further describe benchmarking approaches for evaluating spatial resolution and strategies for assay prioritization when sample volume is limiting. By reducing participant burden and facilitating integration with stool, saliva, and clinical metadata, this approach enables longitudinal and large-cohort studies linking SI microbial ecology and host physiology to human health.
Zhang, M.; Zhao, J.; Tang, W.; Xing, J.; Li, J.; Zhang, H.; Qiu, J.; Zhang, Y.
Show abstract
In primary care and outpatient settings, clinically important patient information is often embedded in fragmented, ambiguous, repetitive, and noisy communication between physicians and patients. This limits physicians ability to obtain a clear preconsultation overview of symptoms, history of present illness, and visit intent, while also preventing real world clinical dialogues from being reused in hospital information systems and medical artificial intelligence applications. To address this challenge, we developed PCRAgent, a centrally coordinated multi agent framework for preconsultation clinical information organization. Guided by physician inquiry logic, PCRAgent identifies, extracts, corrects, and standardizes patient-reported information from noisy consultations. Its coordinated modules including error detection, semantic editing, output control, contextual memory, and intent recognition enable robust parallel handling of spelling errors, repetitions, grammatical inconsistencies, medical ambiguities, and non-medical interference. A traceable edit list records intermediate corrections and context, allowing iterative refinement without redundant modifications. PCRAgent generates two complementary outputs. One is a PreConsultation Clinical Report for rapid physician review. The other is a Structured Clinical Conversation Dataset for hospital data construction and downstream AI applications. In evaluations using 220000 strongly perturbed consultations, PCRAgent maintained high robustness, achieving a clinical information accuracy of 4.99 out of 5 and key element completeness of 5 out of 5, outperforming GPT4o. Expert review of Chinese and English dialogues confirmed high clinical accuracy of 4.85 out of 5 and high safety of 4.79 out of 5. Multicenter validation in real-world outpatient workflows further demonstrated practical utility. These findings indicate that PCRAgent can efficiently transform noisy and unstructured consultations into physician ready reports and AI ready structured data, improving outpatient efficiency, reducing cognitive burden, ensuring information completeness, supporting precise decision-making, and enabling high-quality reuse of clinical data.
Fan, H.; Mugoya, R.; Finnegan, A.; Thate, J.; Jia, H.; Rossetti, S. C.; Yen, P.-Y.
Show abstract
Despite contributing substantially to clinician burnout, nursing documentation lacks empirical evidence distinguishing clinically essential from administratively driven documentation. Exploiting a COVID-19 documentation relaxation policy as a natural experiment, we analyzed 520,357 patient shifts from 36,321 patients in 54 inpatient units (2019 - 2022) using large language model-assisted flowsheet classification and structural equation modeling. When permitted, front-line nurses reliably distinguished two types of documentation: in acute care units, primary nurses reduced compliance-driven Cares & Safety documentation by 19% (106.4 to 86.2 entries, r = -0.19), while maintaining or increasing documentation directly relevant to respiratory management, with no impact on patient respiratory outcomes. Documentation intensity also co-varied with real-time patient deterioration, consistently across unit types (|{beta}| = 0.13 - 0.14). Together, these findings provide the first large-scale quantitative evidence distinguishing clinically essential documentation from compliance-driven documentation and demonstrate that targeted reduction of the latter is a viable strategy for alleviating documentation burden without compromising care quality for respiratory care management.
Liu, C.; Geltzeiler, A.; Afyouni, A.; Nie, M.; Ravi, K.; French, C.; Chung, W.
Show abstract
Background: Rare diseases affect a significant portion of the global population, yet patients often endure a lengthy diagnostic odyssey, frequently missing the opportunity for timely whole-exome or whole-genome sequencing (WES/WGS). Existing informatics tools often rely on pre-identified patients or rigid, institution-specific rule sets, failing to address the broader operational question of clinical necessity and feasibility. Method: We introduce RESCUE, an end-to-end, multi-agent LLM-powered workflow designed for proactive rare-disease screening across the entire electronic health record (EHR). RESCUE utilizes a team of specialized agents including Ontology, Modeling, Screening, and Review, to automate the screening process. The Ontology Agent classifies clinical data into a four-tier genetic-evidence taxonomy; the Modeling Agent builds a positive-unlabeled (PU) XGBoost classifier to identify potential cases; the Screening Agent applies these models across the EHR population; and the Review Agent evaluates candidates by sampling clinical notes to ensure medical necessity and operational feasibility for sequencing. Results: Our retrospective evaluation on a holdout set (n=12,591) demonstrated strong discrimination (AUC 0.808). Among 175,842 eligible patients from an institutional base of ~494,577, RESCUE-flagged candidates were 7.4-fold more likely to receive subsequent genetic workups compared to controls. Blinded manual chart reviews confirmed that RESCUE identifies previously missed, medically necessary patients with 80% precision, while simultaneously accounting for prior testing history. Conclusion: By decoupling expert roles into modular agents, RESCUE offers a flexible, scalable, and adaptable framework for rare-disease screening. This approach overcomes the limitations of traditional rule-based methods and provides a reproducible, agentic pathway to reduce diagnostic delays and improve patient care at an institutional scale.
Kim, S.; Yoo, H.; Yoo, S.-K.; Lee, J.; Min, Y. W.; Lee, H.
Show abstract
Background and Aims: Endoscopic artificial intelligence is commonly validated on selected single images, whereas gastric cancer interpretation requires integrating whole examinations. We developed GutCore and evaluated whether whole-case endoscopic images could be used for patient-level assessment of gastric cancer depth, biomarkers, and prognosis. Methods: GutCore was pretrained on 5.6 million de-identified endoscopic images from more than ten hospitals. We compared it with five general, medical, and endoscopy-specific foundation models using open image-level datasets and an internal tertiary-center cohort of 11,035 de-identified endoscopic examinations (2019-2023): 8,049 with early or advanced gastric cancer and 2,986 with benign gastritis or intestinal metaplasia. All examination images were aggregated for patient-level assessment of cancer status, invasion depth, molecular biomarkers, and overall survival. Results: Aggregating all stored images from each examination enabled patient-level gastric cancer assessment without selecting representative frames. GutCore achieved AUCs of 0.995 for cancer detection, 0.960 for muscularis propria invasion, and 0.804 for SM2-or-deeper invasion. Prediction of tissue-defined biomarker status was strongest for Epstein-Barr virus status and MLH1 loss (AUC, 0.831 and 0.854), with lower HER2 performance (AUC, 0.673). In the held-out advanced gastric cancer test set, GutCore-derived risk groups showed marked survival separation (log-rank P < .0001; high-risk vs low-risk hazard ratio, 13.18; 95% CI, 6.06-28.66), with stratification persisting within pathological stage II and III disease. External frame-level benchmarks showed strong performance for anatomical landmark recognition, disease grading, and segmentation. Conclusions: GutCore supported whole-case patient-level gastric cancer assessment using routinely stored endoscopic images. Further validation in independent clinical cohorts is needed to establish generalizability and clinical utility.
Justen, L. J.; Zulli, A.; Kantor, R. S.; Linfield, R. Y.; Moskatel, L. S.; Cunningham-Bryant, D.; Kaufman, J.; Johnson, M. C.; McLaren, M. R.; Sabeti, P.
Show abstract
Wastewater metagenomic sequencing (WW-MGS) enables simultaneous detection of hundreds of pathogens, but its use for quantitative pathogen tracking has not been robustly validated. Like wastewater PCR (WW-PCR), WW-MGS is affected by biases from variable fecal dilution and sample processing, but must additionally contend with the compositional structure of sequencing data, where a taxon's apparent abundance depends on the abundance of every other taxon in the sample. Simple summaries such as a pathogen's fraction of total reads may therefore be poorly suited to quantitative use. We retrospectively evaluated seven normalization approaches that attempt to control for these sources of bias against a baseline of total read relative abundance, using 1,425 samples from the CASPER consortium spanning 25 U.S. sites. Each approach was compared against WW-PCR and clinical data across eight total pathogens. Among the normalization strategies we evaluated, tobamovirus markers, diet-derived plant viruses abundant in human stool, performed best. Normalizing WW-MGS data by tobamovirus-genus counts improved median site concordance for 18 of 19 pathogen and comparison-source combinations. Gains were largest for year-round-circulating SARS-CoV-2 and norovirus and smaller for sharply seasonal pathogens such as influenza and respiratory syncytial virus, where baseline concordance was already high. Tobamovirus normalization rarely degraded concordance, with median gains roughly five times larger than median losses. Tobamovirus-normalized WW-MGS reached clinical concordance comparable to targeted WW-PCR, supporting its use as a quantitative trend-monitoring tool alongside pathogen-agnostic detection.
Nie, M.; Chung, W.; Waxler, J.; Lee, M.; Weng, C.; Lewis, R.; Ahimaz, P.; Wang, K.; Liu, C.
Show abstract
Purpose: Prior authorization (PA) for exome or genome sequencing is a time-consuming process that impedes timely rare disease diagnosis. Large language model-based browser agents offer potential for automating these workflows, but their clinical reliability remain uncharacterized. Methods: We developed a sandbox compromising a simulated ES/GS PA submission payer portal and a synthetic EHR containing 836 patient records spanning compliant profiles and deficient profiles with different types of issues. Gemini 3 Pro, Gemini 3 Flash, and Claude Opus 4.5 were evaluated on task completion rate, form completion accuracy, and appropriate withholding for deficient profiles. Results: Larger models achieved much higher task completion rates (Gemini 3 Pro 95.45%, Claude Opus 4.5 93.67%) compared to Gemini 3 Flash (56.05%), but nearly universally failed to withhold submission for deficient profiles whereas Gemini 3 Flash ironically demonstrated superior withholding performance (17.33%). In a non-agentic setting, Gemini 3 Pro correctly identified 91% of the issues in deficient profiles, indicating that withholding failure is attributable to the browser interaction rather than the model's reasoning limitations. Conclusion: Current LLM-based browser agents exhibit a systematic bias towards form submission that poses risks in PA workflows. A modular, multi-agent architecture with human supervision is necessary for a safe clinical deployment.
Winnett, A. V.; Tabachnikova, A.; Chen, J.; Greene, K.; Romano, A. E.; Pei, X. P.; Cooper, M. M.; Silva, J.; Carter, A. M.; Jiang, J.; Kong, Y.; Roos, M.; Middle, C.; Zhang, H.; Thomson, M.; Booher, K.; Kuersten, S.; Iwasaki, A.; Ismagilov, R. F.
Show abstract
COVID-19 vaccines markedly reduce disease severity, but their ability to block infection and transmission remains limited and variable.1 A better understanding of early mucosal immune programs that constrain viral replication at susceptible upper respiratory sites is needed to develop more effective antiviral strategies. However, temporal and anatomical antiviral dynamics are difficult to resolve without longitudinal, paired-site sampling beginning before infection onset. We quantified longitudinal viral load by RT-qPCR and human gene expression by mRNA sequencing in 1,237 samples prospectively collected daily from the nasal cavity, oral cavity, and oropharynx of 16 individuals starting from the onset of naturally acquired SARS-CoV-2 infection, and 16 age-, sex-, and vaccination-matched uninfected individuals. Here we show that Type I interferon (IFN) responses are initiated concurrently across these upper respiratory sites, even before local viral detection. In contrast, Type II IFN initiation is more spatially variable, and earlier nasal Type II IFN initiation is associated with reduced viral replication, prior COVID-19 vaccination, and higher tissue-resident memory T cell (TRM) signature expression. These findings demonstrate that in addition to humoral immunity, prior vaccination primes rapid, inducible mucosal Type II IFN responses, likely mediated by rare TRM upon viral encounter, to limit viral replication and spread.
Villafuerte-Galvez, J. A.; Noriega, M. A.; Cakir Colak, S.; Crawford, C. V.
Show abstract
Background. Clostridioides difficile infection (CDI) imposes a burden that extends well beyond the gastrointestinal tract, yet existing outcome measures only partially capture the patient experience. We used frontier large language models (LLMs) on patient and caregiver narratives at scale to describe how burden shifts with disease course. Methods. We analyzed 189 testimonials from the Peggy Lillis Foundation corpus, sorted into four cohorts with recurrence (r) and fulminant (f) severity as axes (rfCDI, fCDI, rCDI, non-rfCDI). Two independent LLMs coded eight thematic domains, four fulminant flags, thirteen emerging semantic fields, the dominant dimension, and narrative arcs. Two clinicians independently coded a subset for inter-rater reliability (PABAK, Gwet's AC1). Results. Treatment trajectory was the dominant theme in recurrent disease, whereas death and near-death dominated non-recurrent fulminant narratives. Psychological burden was near-universal in fulminant disease (98.0% in rfCDI, 97.2% in fCDI). Caregiver and bereavement content concentrated in fCDI (66.7%). Diagnostic failure was frequent across recurrent cohorts (47.6 - 56.1%). Bacteriotherapy tracked recurrence (60.2% rfCDI versus 5.6% fCDI). Financial, mental-health, and caregiver burdens were prominent and are currently unaddressed by guidelines. Human-human reliability was substantial (PABAK 0.79 for semantic fields, 0.76 for domains); arc coding was least reliable. Conclusions. Patient narratives reveal a course-dependent, multidimensional burden in CDI. Concrete gaps exist between what patients prioritize, what guidelines recommend, and what therapy access provides. Frontier-LLM coding, validated against clinicians, offers a reproducible route to translate these priorities into research, care, and policy.
Podkowik, M.; Welling, A. R.; Dey, S.; Tillman, A.; Putzel, G.; Takats, C.; McWilliams, J.; Bartlett, S.; Samhadaneh, N.; Ulrich, R. J.; Rabii, K. B.; Olusanya, O.; Otto, C.; Drlica, K.; Ortigoza, M. B.; Renson, A.; Pironti, A.; Hochman, S.; Shopsin, B.
Show abstract
Background Mupirocin, a widely used topical agent for decolonization of Staphylococcus aureus, is increasingly compromised by resistance. Although plasmid-mediated mupirocin resistance is a recognized cause of decolonization failure, its role in facilitating hospital-wide transmission is unknown. Methods We conducted genomic surveillance of S. aureus at two interconnected urban hospitals where mupirocin decolonization is routine. Genome sequencing of >10,000 isolates was integrated with patient data to identify transmission and resistance determinants. Bacterial phenotypes and fitness were evaluated in vitro and in murine colonization models. Findings Genome sequencing identified 475 hospital transmission events; none were detected by conventional surveillance. The mupA (ileS2) resistance determinant, carried on conjugative plasmids, was enriched eightfold in methicillin-resistant S. aureus (MRSA) relative to methicillin-susceptible strains. mupA was associated with nearly a threefold greater chance of hospital transmission, especially within endemic healthcare-associated MRSA lineages, and was enriched twofold in hospital-onset infections compared with admission colonizing isolates. Multiple independently evolved inactivating mutations in the essential chromosomal gene ileS1 co-occurred with mupA, creating plasmid addiction in which mupA became indispensable for bacterial survival. Addiction arose most frequently within the dominant community-acquired MRSA lineage, where plasmid carriage reduced colonization fitness in mice. Plasmid-containing strains exhibited stringent-response activation, explaining the fitness costs and collateral tolerance to disinfectants, such as ethanol and peroxide. Although addiction reduced S. aureus fitness, it increased plasmid transfer, and addicted variants spread across hosts, demonstrating adaptation that mitigates these costs. Unexpectedly, we identified a mupirocin-dependent vulnerability to isoleucine limitation, revealing a potential strategy to target mupA-mediated resistance. Interpretation Plasmids promote hospital transmission of mupirocin-resistant S. aureus and create an evolutionary trap in which antibiotic use selects for bacterial dependence on otherwise costly resistance elements. This dependence revealed a collateral bacterial vulnerability that could be exploited to target resistant strains and preserve the effectiveness of mupirocin.
Li, Q.; Xu, L.; Wang, J.; Li, C.; Wen, W.; Shu, X.; Yang, Y.; Shu, X.-o.; Cai, Q.; Long, J.; Singh, B.; Lau, K. S.; Yin, Z.; Casey, G.; Song, M.; Peters, U.; Zheng, W.; Guo, X.
Show abstract
Bulk tissue-based DNA methylation-wide (MWAS) and transcriptome-wide association studies (TWAS) have identified CpG sites and genes associated with colorectal cancer (CRC) risk, but do not account for cellular heterogeneity. To address this, we developed a deconvolution-informed framework to infer cell-type specific DNA methylation and gene expression profiles from bulk normal colon tissues using reference single-cell epigenomic and transcriptomic datasets. We performed cell-type specific MWAS (ctMWAS) using deconvoluted DNA methylation data from 293 normal colon samples and conducted cell-type specific TWAS (ctTWAS) using deconvoluted gene expression data from 707 normal colon samples. Genetically predicted methylation and expression models were integrated with CRC GWAS summary statistics (78,473 cases and 107,143 controls) to identify risk-associated CpG sites and genes. Through ctMWAS, ctTWAS, and colocalization analyses, we identified 178 significant cell-type-specific CpG sites in 106 loci and 68 risk genes in 40 loci, including 26 previously unreported loci. Through additional integrative methylation-gene analysis, we prioritized 132 candidate risk genes, the majority of which were supported by multi-omics evidence and stage-specific dysregulation across the adenoma-carcinoma and serrated-carcinoma progression pathways. Pathway enrichment analyses implicated pathways involved in DNA double-strand break repair, TP53 regulation, TGF-{beta} signaling, and innate immune responses. Among prioritized genes, 14 were identified as putative druggable targets linked to 90 FDA-approved or clinical-stage drugs. Experimental validation supports an oncogenic role for SF3A3. These findings demonstrate that deconvolution-informed integrative analyses enable cell-type-resolved identification of epigenetic and transcriptional mechanisms underlying CRC susceptibility and provide insights into disease biology, prevention, and therapeutic target discovery.
Wang, L.; Poenaru, D.
Show abstract
Background: Medical-AI publications do more than report technical performance; they also frame AI as beneficial, uncertain, or risky. How this evaluative stance has changed across the medical literature is not well characterized. Objective: To characterize evaluative stance in published medical-AI discourse abstracts from January 2021 through April 2026 and examine variation over time, concern themes, failure mechanisms, specialties, first-author geography, and publication format. Methods: We conducted an LLM-assisted computational content analysis of medical-AI abstracts from first- and second-quartile medical journals. Of 97,492 post-cutoff Q1/Q2 records entering the prefilter, 16,759 were retained as discourse or evaluative. Claude Sonnet 4.6 assigned 16,749 valid stance classifications using Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Annual analyses used 16,747 records dated 2021-2026. Critical stance was defined as Alarm plus Caution and indexed evaluative scrutiny rather than opposition or author psychology. Each LLM step was validated against blinded human coding by one author: prefilter Cohen kappa = 0.51, stance quadratic-weighted kappa = 0.79 (95% CI 0.72-0.84) for codable, in-scope records, specialty kappa = 0.75, and mechanism-axis kappa = 0.84 for model type and 0.57 for failure mode. Results: Advocacy declined from 2.9% in 2021 to 0.6% in partial 2026, while Cautious Optimism remained the majority stance. Among 16,749 valid classifications, 30.8% were critical. Critical share increased from 25.4% to 32.6%, a 7.25-percentage-point increase based on unrounded estimates. Among critical records, patient safety remained the most prevalent concern. Hallucination/errors increased by 30.9 percentage points. Regulation declined by 22.0 percentage points and ethics/bias by 8.1 percentage points in prevalence share; these declines do not necessarily indicate lower publication counts. Within the hallucination/error theme, factual error was more common than fabrication. Fabrication estimates should be treated as an upper bound because failure-mode agreement was moderate. Specialty patterns were heterogeneous. Critical rate was inversely associated with FDA-cleared device availability (Spearman rho = -0.65, two-sided p = 0.004), which does not measure adoption, deployment, maturity, or clinical use. First-author geography described publication metadata and discourse, not national attitudes or research quality. Reviews were the least critical and most favourable format. In exploratory forward validation, 2 of 78 early Advocacy predictions were fully borne out, although the analysis was single-rater and retrieval-dependent. Conclusions: Published medical-AI abstracts became modestly less promotional and more focused on specific errors and safety concerns. Unqualified promotion declined, but qualified favourable framing remained dominant, and the rise in critical stance was modest. Concern moved toward errors and patient safety, with factual error discussed more often than fabrication. These findings describe published discourse, not AI capability or whether the evaluations were correct.
Shelley, J. P.; Lake, A. M.; Sealock, J. M.; Ueland, T. E.; Peterson, J. F.; Davis, L. K.; Mosley, J. D.
Show abstract
Objective: Genomic research using electronic health record (EHR)-linked biobanks is influenced by heterogeneity in the clinical settings (care sites) where encounters occur. We developed two methods leveraging care site data: ClinicScan identifies where phenotype documentation occurs, and ClinicWAS identifies specialty utilization patterns associated with a risk factor. Materials and Methods: We extracted care sites for each clinical encounter at an academic medical center and mapped each to a clinical specialty. ClinicScan summarizes the specialty distribution of a user-specified diagnosis; ClinicWAS fits a logistic regression for each care site to identify specialty encounters associated with a user-specified risk factor. We applied ClinicScan to depression to test whether requiring a psychiatry encounter strengthened the association between a polygenic risk score (PRS) and a depression phenotype, and ClinicWAS to a coronary heart disease (CHD) PRS to identify sites enriched for high-risk patients. Results: Across 64,983,257 encounters, 2,544 care sites mapped to 57 specialties. Most depression diagnoses occurred in primary care (30.3%) and psychiatry (19.8%). Requiring a psychiatry encounter strengthened the PRS-phenotype association (OR=1.30, 95% CI 1.26-1.35) versus two or more diagnosis codes alone (OR=1.21, 95% CI 1.19-1.24). CHD ClinicWAS identified 19 associated care sites, including 5 catheterization labs. Men and women with high genetic risk (PRS[≥]95th percentile) underwent catheterization for CHD 3.1 (1.5-4.6) and 4.6 (2.5-6.7) years earlier than normal-risk participants, respectively. Discussion: Care site data capture phenotype heterogeneity that otherwise distorts EHR-based phenotypes and obscures high-risk subpopulations. Conclusion: Clinical care site data are an under-utilized resource in EHR-linked biobanks.
Diament, A.; Sapir, G.; Gorodetski, M.; Wolf, A.; Rice, A.; Azouri, D.; Etzion-Fuchs, A.; Gelbard Solodkin, D.; Talmor-Barkan, Y.; Lutsker, G.; Segal, E.; Rossman, H.
Show abstract
General-purpose language models generate fluent health reports that can fabricate derived clinical metrics. In an illustrative comparison on identical two-week CGM and meal data, leading foundation models produced reports with invented MAGE values, inflated meal counts, and unreferenced complication-risk projections: failures invisible to non-expert readers and plausible enough to mislead clinicians. We describe the HPP Personal Health Agent (PHA), a metabolic health agent that grounds generation in four layers: the Human Phenotype Project (HPP), a deep-phenotyped cohort of 13,000+ participants supplying population references and trained predictive models; 21 domain-expert tools and trained-model wrappers that compute clinical metrics and risk predictions; declarative behavioural skills that constrain what the model may claim; and 21 automated evals across 8 categories developed via a test-driven cycle in which each eval encodes a failure mode discovered during iterative development. In a 210-report matrix (14 participants x 3 prompts x 5 system conditions), the gains are largest on the system's primary use case (meal-grounded metabolic reports, the report it was designed for), where the full system raises a deterministic form/provenance score from 0.37 (the same foundation model with no tools or skills) to 0.91; this score measures structural completeness, numerical accuracy, tool grounding, and clinical-language compliance: a necessary condition for trustworthy health reporting, with clinical quality as a complementary axis examined qualitatively. A skills-vs-tools decomposition shows the two layers act on different axes: tools drive numerical accuracy (from about 14% to 90% of reported metrics correct), while the declarative skills add most of the remaining gain in citations, completeness, and structure (tools alone recover only part of the gap, 0.49 from the same 0.37 baseline). The lift generalises beyond the primary use case: to a second metabolic prompt (0.72) and a cardiovascular extension (0.70), each from a 0.37-0.39 baseline. The architecture extends across clinical domains: adding a SCORE2 cardiovascular risk tool and a corresponding skill (with no changes to orchestration, eval harness, or existing tools) produced a cardiovascular risk report from the same system. Trustworthy domain-specialised health AI is a systems design problem: deep-phenotyped cohort data, domain-expert tools and models, and eval-driven development together form a replicable pattern.
Jampani, Y.; McClay, J.; Rana, M. K. Z.; Mandhadi, V.; Mcmahon, T.; Niu, X.; Islam, M. S.; Hossain, S.
Show abstract
Academic medical centers face growing demand for clinical data access that traditional informatics infrastructure cannot support at scale. We describe HABITAT, a governed, tiered research data ecosystem at the University of Missouri integrating internal and external sources, including unstructured clinical data, into a layered architecture supporting PCORnet and OMOP common data models. A three-tier model provides self-service feasibility tools, secure cloud analytics, and curated honest broker services under a Five Safes governance framework. From mid- 2022 through February 2026, HABITAT received 1,150 requests with 689 fulfilled and supported 29 externally funded projects. Median fulfillment for internal data requests improved from 30 to 20 days, and an early survey captured 54 scholarly outputs, including peer-reviewed publications in high-impact journals. HABITAT supported coursework to multi-site research without proportional staff increases, while cost r
Gorenshtein, A.; Omar, M.; Barash, Y.; Kruskal, J. B.; Ahmed, M.; Brook, O. R.; Klang, E.
Show abstract
Clinical AI agents may be assigned to individual patients, but hospital resources are shared across many patients. We tested what agents do when helping their assigned patient would violate the hospital's rule for a scarce resource. We analyzed 22,916 simulated cases comprising 274,992 logged agent actions across 20 AI models. In each scenario, the agent could claim a scarce resource for its patient even though the hospital rule gave another patient priority. We varied only the agent's assigned role, from responsibility for the whole ward to strong advocacy for one patient. Violations of the hospital rule rose from 32.5% under whole-ward responsibility to 69.4% under strong patient advocacy, a 36.9-point increase (95% CI, 25.7-48.0). Agents correctly identified which patient should receive the resource in 95.7% of tests, yet still took it for their own patient in 65.9% of those episodes. Asking the agent to apply its own allocation judgment immediately before acting reduced violations to 0-2% in a three-model follow-up experiment. Assigned roles can shape how clinical AI agents use shared hospital resources, even when they identify the correct priority patient. Patient-focused agents should not independently control shared resources without an allocation check.