Med
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match Med's content profile, based on 39 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Diament, A.; Sapir, G.; Gorodetski, M.; Wolf, A.; Rice, A.; Azouri, D.; Etzion-Fuchs, A.; Gelbard Solodkin, D.; Talmor-Barkan, Y.; Lutsker, G.; Segal, E.; Rossman, H.
Show abstract
General-purpose language models generate fluent health reports that can fabricate derived clinical metrics. In an illustrative comparison on identical two-week CGM and meal data, leading foundation models produced reports with invented MAGE values, inflated meal counts, and unreferenced complication-risk projections: failures invisible to non-expert readers and plausible enough to mislead clinicians. We describe the HPP Personal Health Agent (PHA), a metabolic health agent that grounds generation in four layers: the Human Phenotype Project (HPP), a deep-phenotyped cohort of 13,000+ participants supplying population references and trained predictive models; 21 domain-expert tools and trained-model wrappers that compute clinical metrics and risk predictions; declarative behavioural skills that constrain what the model may claim; and 21 automated evals across 8 categories developed via a test-driven cycle in which each eval encodes a failure mode discovered during iterative development. In a 210-report matrix (14 participants x 3 prompts x 5 system conditions), the gains are largest on the system's primary use case (meal-grounded metabolic reports, the report it was designed for), where the full system raises a deterministic form/provenance score from 0.37 (the same foundation model with no tools or skills) to 0.91; this score measures structural completeness, numerical accuracy, tool grounding, and clinical-language compliance: a necessary condition for trustworthy health reporting, with clinical quality as a complementary axis examined qualitatively. A skills-vs-tools decomposition shows the two layers act on different axes: tools drive numerical accuracy (from about 14% to 90% of reported metrics correct), while the declarative skills add most of the remaining gain in citations, completeness, and structure (tools alone recover only part of the gap, 0.49 from the same 0.37 baseline). The lift generalises beyond the primary use case: to a second metabolic prompt (0.72) and a cardiovascular extension (0.70), each from a 0.37-0.39 baseline. The architecture extends across clinical domains: adding a SCORE2 cardiovascular risk tool and a corresponding skill (with no changes to orchestration, eval harness, or existing tools) produced a cardiovascular risk report from the same system. Trustworthy domain-specialised health AI is a systems design problem: deep-phenotyped cohort data, domain-expert tools and models, and eval-driven development together form a replicable pattern.
Gorenshtein, A.; Omar, M.; Barash, Y.; Kruskal, J. B.; Ahmed, M.; Brook, O. R.; Klang, E.
Show abstract
Clinical AI agents may be assigned to individual patients, but hospital resources are shared across many patients. We tested what agents do when helping their assigned patient would violate the hospital's rule for a scarce resource. We analyzed 22,916 simulated cases comprising 274,992 logged agent actions across 20 AI models. In each scenario, the agent could claim a scarce resource for its patient even though the hospital rule gave another patient priority. We varied only the agent's assigned role, from responsibility for the whole ward to strong advocacy for one patient. Violations of the hospital rule rose from 32.5% under whole-ward responsibility to 69.4% under strong patient advocacy, a 36.9-point increase (95% CI, 25.7-48.0). Agents correctly identified which patient should receive the resource in 95.7% of tests, yet still took it for their own patient in 65.9% of those episodes. Asking the agent to apply its own allocation judgment immediately before acting reduced violations to 0-2% in a three-model follow-up experiment. Assigned roles can shape how clinical AI agents use shared hospital resources, even when they identify the correct priority patient. Patient-focused agents should not independently control shared resources without an allocation check.
Tyagi, S.; Ramakrishnaiah, Y.; Hawkey, J.; Wisniewski, J.; Blakeway, L.; Christian, T.; Sikric, V.; Librata, W.; Song, J.; Webb, G. I.; Ashok, A.; Bain, C.; Macesic, N.; Peleg, A. Y.
Show abstract
Artificial intelligence (AI) has the potential to transform healthcare, with advanced multimodal approaches showing great promise in leveraging diverse health-related data. Here, we applied multimodal AI to entire electronic health record (EHR) and complete pathogen genome data to predict patient outcomes from life-threatening infection. An automated, scalable pipeline was developed for EHR data preprocessing, quality control, and standardisation. A deep learning fusion model was trained to predict in-hospital mortality, need for ICU admission, prolonged length of stay and 30-day unplanned readmission. We then developed a novel genomic large language model (gLLM) architecture to incorporate bacterial genomic features into the multimodal fusion model. The cohort comprised 2,656 bloodstream infection hospitalisations involving 2,535 patients. Deep learning fusion models using entire structured and unstructured EHR data outperformed traditional APACHE II score mortality prediction (AUROC [95% confidence intervals] 0.93 [0.92-0.94] versus 0.77 [0.77-0.78]). The model also showed strong performance for predicting the need for ICU admission (AUROC 0.978 [0.966 - 0.986]), prolonged hospital length of stay (AUROC 0.803 [0.790 - 0.812]) and unplanned readmission (AUROC 0.696 [0.690 - 0.701]). As proof of principle, incorporating entire microbial genomic features from the causative pathogen further enhanced prediction and enabled identification of key bacterial virulence pathways relevant for human disease. Multimodal AI integrating harmonised EHR and genomic data can accurately identify hospitalised patients at risk of poor outcomes. These approaches are scalable to other subspecialities of medicine.
Otieno, C. O.; Seagle, H. M.; Akerele, A. T.; Jaworski, J.; Guare, L.; Setia-Verma, S.; Velez Edwards, D. R.; Edwards, T. L.
Show abstract
Transcriptome-wide association studies (TWAS) can identify genes where genetically predicted gene expression is associated with disease risk, but translating those signals into therapeutic opportunities remains time-consuming, manual, and difficult to reproduce. We developed TRACE (TWAS-driven Repurposing through AI-assisted Curation of Evidence), a gene- and phenotype-agnostic computational pipeline that accepts a TWAS gene and effect-size direction, normalizes the gene symbol, retrieves FDA-approved drug-gene candidates from four online resources, collects related peer-reviewed literature from PubMed, and uses a fine-tuned biomedical language model to classify whether the literature supports a direct drug-gene relationship, the mechanism of action, and the direction of effect. The pipeline then compares the drug-derived direction with the direction implied by the TWAS effect estimate to rank candidate therapeutic pairs and flag potential drug safety concerns. The local classifier, built on BiomedBERT, was trained using pipeline-derived labels, BioCreative VI ChemProt gold-standard chemical-protein relation examples, and author-reviewed active-learning cases, reaching a held-out macro F1 of 0.809 across three simultaneous classification tasks. We validated the pipeline against a manually curated endometriosis gold standard of 43 drug-gene pairs spanning six TWAS-identified genes, developed through S-PrediXcan analysis of endometriosis GWAS summary statistics, manual querying of four drug-gene interaction databases for each gene, literature review of drug-gene mechanistic evidence, and Mendelian randomization validation of candidate pairs. External validation used two independently published genetically informed drug-repurposing studies in metabolic dysfunction-associated steatotic liver disease (MASLD) and type 2 diabetes (T2D). The pipeline recovered 90.7% of endometriosis pairs, 88.2% of MASLD pairs, and 92.9% of T2D pairs that were present in at least one queried database. Applied to 99 endometriosis-associated TWAS genes, the pipeline identified 1,089 FDA-approved drug-gene pairs, 32 candidate therapeutic pairs, and 77 potential safety concerns, including independent recovery of leuprolide acetate, an established endometriosis therapy. This framework provides a scalable, literature-grounded bridge from TWAS discovery to prioritized therapeutic hypotheses, while preserving uncertainty through manual-review flags and requiring downstream Mendelian randomization, electronic health record-based validation, and experimental follow-up before clinical interpretation.
Gorenshtein, A.; Jia, E. L.; Omar, M.; Brook, O. R.; Ahmed, M.; Kruskel, J. B.; Barash, Y.; Klang, E.
Show abstract
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.
Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O.; Klang, E.; Barash, Y.
Show abstract
Background Large language models are increasingly proposed to post-edit decoded text in communication brain-computer interfaces and augmentative communication. A fluent model can substitute a different intent than attempted (intent drift). Whether meaning survives or confidence flags failure is unmeasured. Methods In-silico benchmark of 20 open-weight models post-editing text (4,252,326 labeled generations) corrupted with an empirical P300 confusion matrix at five levels (0-40% character error rate, CER) across the ALS message-banking vocabulary (AUTH), a message-critical probe set, and matched controls. Outputs were scored faithful, degraded, or drift by an ensemble benchmarked against physicians. A substudy re-ran 562 messages under six interface policies (seven-model panel). Findings Detected drift rose steeply with corruption in all three corpora, from 2.2% to 60.3% at 0-40% target CER in AUTH, a stress-test upper bound, not an expected clinical rate (odds ratio 2.30 per 10-percentage-point rise in target CER). Stated confidence discriminated faithful outputs reasonably well (AUROC 0.83, 0.80-0.85) but was poorly calibrated (expected calibration error 0.32, 0.27-0.37): 28.4% of outputs at confidence 90 or higher were not faithful. Message-critical content carried a small excess after matching, surviving detector removal (rule-free OR 1.10). The ratio of faithful rescues to fluent errors exceeded 1 at low corruption but fell below 1 at 20-30% target CER. No interface policy removed drift: conservative editing and abstention lowered it, alternatives and expansion raised it; the best drifted on 18.0 per 100. A 2,281-item panel (16 of 20 models) gave moderate ensemble-versus-consensus agreement (kappa 0.41); correction lowered pooled drift 31.4% to 28.3%, and a CER-stratified physician-corrected re-analysis confirmed the dose-response at each level. Interpretation Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed. This does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed. Funding: A.G. and E.K. were supported in part by the Clinical and Translational Science Awards (CTSA) grant UL1TR002541 from the National Center for Advancing Translational Sciences, through the Harvard Catalyst | The Harvard Clinical and Translational Science Center Pilot Award Program. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Competing interests: The authors declare that they have no competing interests.
Lane, T.; Green, T. E.; Garza, D.; Brown, N. J.; de Silva, M. G.; Bennett, M. F.; Tubb, C.; Macdonald, S. M. W.; Gascoigne, A.; Phillips, R. J.; Slavin, J.; D'Arcy, C.; MacGregor, D.; Clifford, A.; Pathmanathan, L.; Robertson, S. J.; Bekhor, P.; Simpson, J.; Gooley, S.; Scheffer, I. E.; Berkovic, S. F.; Penington, A. J.; Hildebrand, M.
Show abstract
Targeted precision therapies are increasingly used in the treatment of individuals with vascular anomalies (VAs). This increases the need for rapid, accurate and inexpensive genetic diagnosis. Droplet digital polymerase chain reaction (ddPCR) is an alternative to next-generation sequencing (NGS), permitting rapid, highly sensitive interrogation of recurrent pathogenic mosaic variants. We examined the feasibility of ddPCR as a primary diagnostic tool in a large cohort of individuals with VAs. Lesional tissue was collected for ddPCR of up to 46 recurrent pathogenic variants across 16 genes associated with VAs. Specimens were assessed on a subset of assays for each individual based on clinical phenotype. Most individuals who had negative ddPCR results went on to high-depth gene panel or deep exome NGS, or Sanger sequencing. Here we report the phenotypic and molecular findings for 78 newly recruited and tested individuals in addition to the 60 individuals already reported from our cohort. The overall diagnostic yield for our cohort when combined with individuals previously reported was 104/138 (75%). Of 138 individuals tested, recurrent pathogenic variants were detected in 71 (51%) on ddPCR. Variants were most frequently identified in PIK3CA (n=28), TEK (n=18), GNAQ (n=12), or MAP2K1 (n=7). In a further 33 individuals, pathogenic variants were identified on NGS or Sanger sequencing. Our findings indicate that ddPCR is an efficient method achieving a high diagnostic yield in our cohort when used prior to sequencing.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Pavia, M. J.; Amaro, I. F.; Xu, D.; Gonzalez-Hernandez, G.; Scotch, M.
Show abstract
Influenza vaccine effectiveness (VE) is estimated from a limited number of clinics using a test-negative design. These standard estimates face geographic, temporal, and operational constraints. Using Twitter/X data, we applied few-shot chain-of-thought prompting to identify self-reported vaccination status and influenza test results, then implemented a test-negative-like design to estimate VE. Our estimates fell within the range of interim reports and could complement current systems, improving feasibility, timeliness, and scalability.
Hayder, N. S.; Bukhari, S. A. C.
Show abstract
Synthetic clinical data are increasingly used for healthcare machine-learning development, model validation, data sharing, and predeployment testing, yet such data often claim to be trustworthy after passing a limited collection of realism tests. A synthetic dataset may indeed claim statistical similarity while leaking training membership, erasing rare subgroups, failing on held-out real patients, or lacking sufficient artifacts for reproduction. We introduce SynTrustBench, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization. Its Evidence Assessment component audits published reports and produces a five-element Evidence Maturity Profile (EMP) together with a separate evaluability gate. Its executable structured-tabular protocol accepts frozen real training data, held-out real test data, a synthetic table, and a declarative configuration; computes dimension-specific metrics and uncertainty; and produces subgroup results, failure flags, benchmark cards, and provenance manifests. In a frozen pilot audit of 30 reports, 17 of 30 quantitatively evaluated privacy, 2 of 30 documented a formal privacy guarantee to the audit threshold, 2 of 30 evaluated equity, 12 of 30 evaluated robustness, and only 4 of 30 passed the evaluability gate. The executable implementation operationalizes the same dimensions through distribution and dependency checks, frozen train-on-real/test-on-real (TRTR) and train-on-synthetic/test-on-real (TSTR) utility, empirical privacy attacks, subgroup analysis, perturbation testing, and a controlled failure-injection harness. SynTrustBench does not certify clinical safety or collapse trustworthiness into a single score. Instead, it provides an inspectable predeployment contract for identifying what was evaluated, what failed, what remains unknown, and whether evidence is sufficiently complete and reproducible for comparison or downstream healthcare AI use.
Myers, M.; Robson, F.; Baig, S.; Kular, S.; Aziz, M.; Burchi, E.; Battacharyya, D.; Li, S.; Majid, A.; Ali, A. N.
Show abstract
Background: Aneurysmal subarachnoid haemorrhage (aSAH) is frequently complicated by delayed cerebral ischaemia (DCI), for which current therapies incompletely target the underlying multifactorial pathophysiology. Transauricular vagus nerve stimulation (taVNS) modulates inflammatory, vasoactive and autonomic pathways and may attenuate secondary brain injury after aSAH. Methods: We conducted a prospective, single-centre, single-blind, randomised, sham-controlled pilot trial in adults within 5 days of aneurysm securing for non-traumatic aSAH. Participants were allocated 1:1 to active taVNS (left tragus) or sham (left earlobe) using a portable device delivered for 45 minutes twice daily over 5 days. Primary outcomes were safety (taVNS-related serious adverse events), acceptability, and compliance; secondary outcomes included inflammatory biomarkers, DCI, in-hospital complications, and functional outcomes to 1 month. Results: Thirty patients were randomised (16 taVNS, 14 sham), with numerically more severe aSAH at baseline in the taVNS arm. No taVNS-related serious adverse events occurred; side effects were generally mild and transient, and over 80% of planned sessions were completed. TaVNS produced greater reductions in serum tumour necrosis factor- and trends towards reductions in interleukin-1{beta} and interleukin-10, with numerically fewer DCI events (6.6% vs 35.7%) and neurological impairments (16.7% vs 53.8%), although functional outcomes were not statistically different at 1 month. Conclusions: Early taVNS after aSAH is safe, acceptable, and feasible in the neurocritical care setting and shows biologically plausible signals warranting evaluation in larger multi-centre trials.
Ma, Y.; Weissenbacher, D.; Patock, J.; Gonzalez-Hernandez, G.
Show abstract
Adverse drug event (ADE) evidence is produced across patient-generated, clinical, and scientific settings that differ in language, documentation purpose, terminology, and degree of standardization. These differences shape both which adverse experiences become visible to pharmacovigilance systems and how readily they can be linked to curated drug-safety knowledge. We examine these relationships across five corpora representing distinct data-production settings: ADE Corpus V2 (medical case reports), SMM4H-2026 Task 1 (multi-lingual user-generated health content), CADEC V2 (patient-forum narratives), the Dutch ADE Corpus (EHR clinical notes), and TwiMed-PubMed (biomedical literature). A shared BERTopic analysis of ADE-positive texts concerning antidepressants and antihypertensives across the four English-language corpora identified nine interpretable topics. CADEC V2 contained a more differentiated distribution of symptom-specific themes, including sexual effects, suicidal or panic-related thoughts, vivid dreams, and memory difficulties, whereas SMM4H-2026, TwiMed-PubMed, and ADE Corpus V2 were dominated by a broader medication, sleep, tiredness, and pain theme. These patterns indicate that data-production context shapes what adverse experiences are expressed and standardized, with patient-generated narratives surfacing subjective, symptom-specific experience largely absent from clinical and scientific sources. We further show that this context shapes how readily real-world drug mentions can be linked to curated pharmacovigilance knowledge. Using SIDER 4.1 as a retrieval resource, we find substantial cross-corpus mismatches between real-world drug mentions and SIDER's predominantly English, generic-name vocabulary: CADEC V2 achieved only 9.5% exact-match coverage, with unmatched mentions frequently involving brand names, misspellings, and language-specific variants, compared to 91.0% coverage in TwiMed-PubMed's formally standardized biomedical literature. To probe how these representational differences interact with automated detection, we compare corpus-specific QLoRA fine-tuning of Llama-3.2-3B with retrieval-augmented inference using Llama-3.1-70B and Llama-3.1-405B grounded in SIDER-retrieved evidence. QLoRA-Llama-3B achieved the highest micro-averaged F1 scores on ADE Corpus V2 (0.91), CADEC V2 (0.88), and SMM4H-2026 (0.80), whereas SIDER-grounded inference with Llama-3.1-405B achieved the highest scores on Dutch ADE (0.95) and TwiMed-PubMed (0.91); these corpus-dependent patterns should not be interpreted as a controlled comparison of adaptation strategies, since model scale, task formulation, and available supervision differ across datasets. Together, our findings indicate that data-production context influences what adverse experiences are expressed, how they are standardized, and how readily they can be retrieved and computationally detected. Pharmacovigilance systems should therefore combine source-sensitive supervision with external knowledge grounding while explicitly monitoring gaps between real-world language and curated drug-safety resources.
Nielsen, M.; Castelo, A.; Altaie, M.; Bennett, J.; Anthony, A.; Siddiqi, N. S.; Gupta, A. C.; Brock, K. K.; Woodland, M.
Show abstract
Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.
Lyman, K.; Thinzar, L. P.; Vargas, D.; Falcone, G. J.; Gilmore, E.; Kim, J. A.; Magid-Bernstein, J.; de Havenon, A.; Matouk, C. C.; Hebert, R.; Sheth, K. N.; Ortega-Gutierrez, S.; Petersen, N. H.
Show abstract
Optimal blood pressure management after thrombectomy remains uncertain, and individualized autoregulation-based targets typically require continuous neuromonitoring. We developed an angiography-derived autoregulatory metric using intraprocedural data and applied it retrospectively to a single-center cohort of patients who underwent thrombectomy for acute stroke. From 62 patients with 3-month functional outcomes, greater time within the predicted autoregulatory range during the first 24 hours after thrombectomy was independently associated with improved outcome after adjustment for covariates (odds ratio per 10% increase, 1.86; 95% CI, 1.31-2.66; P = .0006). These findings support routine angiography as a potential source of early, patient-specific hemodynamic targets after thrombectomy.
Fu, Z.; Fastiggi, V. A.; Phelan, A.; Bell, K.; Lucarelli, S.; Wilson, S. S.; Lindner, J. M.; Cutler, A. A.
Show abstract
Chronic inflammation drives persistent systemic cytokine signaling that contributes to vascular dysfunction and secondary vasculitis, yet mechanistic studies are limited by models that fail to capture the multicellular architecture and dynamics of human arteries. In contrast, perfusing intact vessels ex vivo has limited tractability because of material availability and difficulty of genetic or biochemical manipulation. We developed a modular, perfused artery-on-a-chip platform by tri-axially bioprinting primary human vascular cells to recapitulate the concentric organization of the intimal, medial, and adventitial layers. The engineered vessels are viable longer than 21 days, with functional endothelial barriers, contractile smooth muscle behavior, and actively remodeled extracellular matrices bearing hallmarks of native vascular tissue. Addition of tumor necrosis factor alpha (TNF) induces altered transcript levels of proinflammatory mediators and secretion of cytokines and matrix-remodeling enzymes without compromising vessel viability. Importantly, this secretory response is effectively attenuated by both a small-molecule JAK1 inhibitor (ABT-317) and anti-TNF antibody (Infliximab), demonstrating the models utility for therapeutic evaluation.
Kocot, J.; Pradhan, S. H.; Maric, D.; Kosa, P.; Winkler, C.; Oguz, C.; Myers, T. G.; Wigerblad, G.; Lack, J.; Haigh, C.; Peterson, K.; Bielekova, B.
Show abstract
Modeling neural-immune interactions in neurodegenerative and immune-mediated central nervous system (CNS) diseases requires human 3D models that capture cellular diversity and long-term tissue maturation. Here, we present an enhanced human induced pluripotent stem cell (hiPSC)-derived cerebral organoid (CO) platform optimized to mitigate core hypoxia for over 200 days. Timed pro-myelinating cues established organized neuronal layering and progressive axonal myelination through day 140, while vascular fusion yielded assembloids incorporating endothelial structures and microglia. Extended culture (>500-750 days) spontaneously reproduced hallmark features of human CNS aging, including cellular senescence signatures, neuroaxonal loss, hypomyelination, and the autonomous emergence of a neurotoxic astrocyte transcriptional profile in the complete absence of microglia or immune cells. Co-culture with autologous activated peripheral blood mononuclear cells (PBMC) resulted in transient immune infiltration and a pronounced type II interferon response across CNS lineages. High-plex spatial transcriptomics revealed that immune cell infiltration was associated with oligodendrocyte loss and in aged organoids also with downregulated oligodendrocyte myelin gene transcription. While not fully reproducing adult tissue stoichiometry, this platform enables longitudinal modeling of neural-immune crosstalk in age-related and neuroinflammatory CNS disorders.
George, C. A.; Brown, M. E.; Rana, P.; Killebrew, D. A.; Wilson, R. C.
Show abstract
SummaryA catch-all intronic guide RNA pair excises the KIAA1549--BRAF oncofusion across its major variants, with productive junction excision confirmed by gain-of-function PCR in patient-derived glioma cells. An allele-specific guide selectively disrupts BRAF V600E, in patient-derived pediatric low-grade glioma cells. Pediatric low-grade glioma (pLGG) is the most common brain tumor of childhood, accounting for 30--50% of all pediatric central nervous system malignancies1. The disease is almost universally driven by activating mutations in the BRAF serine/threonine kinase: a chromosomal tandem duplication generating the KIAA1549--BRAF oncofusion in approximately 70% of cases, or the BRAF V600E gain-of-function point mutation in approximately 15%2. Current targeted pharmacotherapies, including the RAF inhibitor tovorafenib, require continuous dosing, are not allele-specific, and carry risks of long-term toxicity in children. A one-time genomic intervention that permanently disables the oncogenic BRAF alteration while preserving wild-type BRAF signaling represents a compelling therapeutic alternative. In this study, we describe the design and experimental validation of allele-specific CRISPR guide RNAs targeting both the KIAA1549--BRAF oncofusion and the BRAF V600E point mutation. For the oncofusion, we developed a double-cut intronic excision strategy in which a guide RNA targeting KIAA1549 intron 14 is paired with a guide RNA targeting BRAF intron 11. Because the genomic breakpoints of all four major fusion variants (KB 16:9, 15:9, 16:11, and 15:11) fall within these introns, a single guide pair can address the full landscape of fusion heterogeneity in a single intervention. For BRAF V600E, we exploited a unique PAM sequence created by the pathogenic TBA transversion at codon 600, enabling allele-specific SpCas9 and AsCas12a guide designs that distinguish the mutant from the wild-type allele at single-nucleotide resolution. We screened guide RNA candidates by ribonucleoprotein (RNP) nucleofection in A375 human melanoma cells (BRAF V600E homozygous) and in patient-derived 3635 PXA glioma cells (BRAF V600E heterozygous). The top KIAA1549 intron 14 guide, K9_i14_A_Cas9, achieved 66% indel frequency in A375 cells. The top BRAF intron 11 guides, B_i11_A_Cas9 and B_i11_D_Cas9, achieved 84% and 85% indel frequency, respectively. For BRAF V600E, the best allele-specific SpCas9 guide achieved l57% editing in A375 cells and l74% editing in 3635 PXA patient-derived glioma cells. Dual-cut excision of the KIAA1549--BRAF junction was confirmed by a gain-of-function PCR assay designed to detect the excision junction amplicon ([~]191 bp) produced by NHEJ-mediated rejoining of the KIAA1549 intron 14 and BRAF intron 11 cut ends.
NING, Z.; Wu, G.; Luo, J.; Li, Y.; Li, Y.; Shi, J.; Fang, W.; To, W. L. W.; Ruan, S.; Zhou, Y.; Chow, S.; Zhang, J.; Jiang, X.; Wang, T.; Gao, H.; Xu, S.; Li, B.; Zhuang, M.; Zheng, P.; Zhu, L.; Lin, C.; Liu, Q.; Yuan, C.-S.; Lam, Y. Y.; Zhai, L.; Zhao, L.; Bian, Z.
Show abstract
How ecological architectures within the gut microbiome convert complex inputs into specific host physiological outcomes remains poorly understood. We used CDD-2101, a multi-component botanical drug operating under an FDA (U.S. Food and Drug Administration) Investigational New Drug program, as a defined ecological perturbation in functional constipation (FC). Integrating a randomized, double-blind, placebo-controlled clinical trial with genome-resolved metagenomics, targeted metabolomics, staged prediction modeling, and receptor-level validation, we show that clinical efficacy of CDD-2101 depends on remodeling a function-specific substructure of the stable Two Competing Guilds (TCG) architecture. We term this substructure the FC-TCG, demonstrate its role along the gut-motility axis, and confirm its effect in three independent gut hypomotility cohorts. The two guilds responded asymmetrically: the intervention selectively suppressed the C1B guild (the pathobiont guild) while largely sparing the C1A guild, the foundation guild that anchors the core gut community, restoring its ecological dominance, producing a coordinated metabolic shift that elevates lithocholic acid and propionic acid. Through gnotobiotic transplantation and receptor antagonism, we demonstrate that lithocholic acid and propionic acid restore gut motility via concurrent engagement of Takeda G protein-coupled receptor 5 (TGR5) and G-protein coupled receptor 43 (GPR43). These findings identify microbial guild architecture as a function-resolved signal-transducing layer that converts multi-component botanical intervention into multi-receptor-mediated gut motility restoration, reframing the gut microbiome from a compositional system into a structural transducer between complex environmental inputs and host physiology.
Markovits, H.; Cohen, Y. J.; Grupel, D.; Goldstein, R.; Goldenstein, H.; Katz Hanein, N.; Razi, T.; Schonmann, Y.; Arbel, R.; Netzer, D.; Tsanani, S. E.; Yamin, D.
Show abstract
Pneumococcal vaccination of older adults is primarily guided by age and clinical eligibility, despite substantial variation in individual risk of severe pneumonia. Here, we used longitudinal electronic health records from 787,538 adults aged [≥]65 years to evaluate the real-world effectiveness of the 20-valent pneumococcal conjugate vaccine (PCV20) and quantify clinical benefit according to baseline risk of pneumonia hospitalization. We developed and validated a machine-learning model using pre-PCV20 data to estimate individual 12-month hospitalization risk and integrated these predictions into a propensity score matching framework. Overall vaccine effectiveness against pneumonia hospitalization was 16.5% (95% CI, 10.6-22.1), but this population-level estimate masked substantial heterogeneity in clinical benefit. The 60% at lowest predicted risk, characterized by younger age and fewer pulmonary and other chronic conditions, showed no measurable reduction in hospitalization (VE, 3.1%; 95% CI, -14.4 to 18.0) and had an estimated 1-year number needed to vaccinate (NNV) of 7,423, compared with 184 and 115 in the intermediate- and high-risk groups, respectively. These findings suggest that incorporating baseline risk into adult pneumococcal vaccination strategies could enable more targeted and potentially better-timed vaccination.
Chen, X.; Jiang, X.; Shan, C.; Wang, Z.; Li, D.; Zhao, C.
Show abstract
Medical AI models have made a great impact on biomedical research and real-world clinical applications, but conducting interdisciplinary medical AI research remains challenging, requiring close collaboration between clinicians and AI experts. Recent advances in large language models (LLMs) and autonomous code agents present an opportunity for low cost medical AI development, where clinicians can build AI tailored to their own research questions, even without continuous support from dedicated AI experts. However, enabling code agents to autonomously tackle complex multimodal medical AI development tasks requires clinicians to construct and supervise an AI research loop with detailed technical specifics, demanding substantial expertise in AI and computer science that they often lack. To address this challenge, we introduce the Medical AI Research Loop Agent (MARLA), an agentic framework that completely abstracts the construction and supervision of medical AI research loops from clinicians. Given a clinician-defined research intent, MARLA automatically translates high-level research goals into executable hierarchical research loops, decomposes them into verifiable sub-loops, and specifies the models, datasets, tools, and evaluation protocols required for each task. During execution, MARLA coordinates specialized code agents, monitors progress, diagnoses failures, and iteratively refines research strategies based on experimental feedback to drive the research process toward optimal outcomes. We evaluate MARLA on multimodal medical AI tasks that require closed-loop conversion from high-level clinical study objectives to trained and validated AI models. The results demonstrate MARLA's ability to autonomously conduct complex medical AI research while substantially reducing the need for AI expertise.