Back

Patterns

Elsevier BV

All preprints, ranked by how well they match Patterns's content profile, based on 78 papers previously published here. The average preprint has a 0.09% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Diagnosis of Blood Diseases and Disorders with Topological Deep Learning

Koung, P.; Fatema, S.; Pakasticali, N.; Luu, H.; Coskunuzer, B.

2025-01-23 hematology 10.1101/2025.01.21.25320908 medRxiv
Top 0.1%
39.7%
Show abstract

Blood diseases and disorders, including leukemia and infectious diseases of red blood cells, pose significant diagnostic challenges due to their complex presentations and reliance on time-consuming cytomorphological analysis to detect subtle morphologic features. While microscopic examination remains the gold standard in their diagnosis, its dependence on expert interpretation high-lights the need for advanced methods to enhance the diagnostic workflow. Recently, deep learning (DL) methods have shown promise in medical imaging by automating and improving accuracy. However, these approaches often require large, annotated datasets, which are scarce for rare diseases, and they often face interpretability issues, limiting their integration into clinical practice. In this study, we present a novel framework that integrates topological features with state-of-the-art DL techniques to enhance the analysis of cytomorphological images for diagnosing blood disorders. By combining global topological information with the localized patterns captured by DL models, our approach improves diagnostic accuracy while ensuring robustness, interpretability, and reproducibility. Experimental results demonstrate that the inclusion of topological features not only enhances model performance but also proves particularly effective in limited data settings. This methodology addresses critical limitations of existing techniques, advancing the classification and diagnosis of blood disorders while improving efficiency and reliability in medical imaging workflows.

2
MyeGPT: an AI agent for Multiple Myeloma

Chang, J. G.; Gout, A. M.; Rodiger, J.; Chung, T.-H.; Mulligan, G.; Chng, W. J.

2026-05-20 hematology 10.64898/2026.05.14.26353252 medRxiv
Top 0.1%
30.6%
Show abstract

Today, advancements in our understanding of cancer biology are increasingly attributed to large-scale clinical-molecular datasets. The case in point for multiple myeloma, the second-most prevalent haematological malignancy, is the CoMMpass study, a dataset with the paired clinical and sequencing data of 1,143 patients. Given its complexity, the multi-omics data of CoMMpass demands programming skills which imposes a hurdle for experimental myeloma researchers who want to validate their hypotheses on population data. The rise of agentic AI over the past few years presents unparalleled opportunities to bridge this technical gap. We propose MyeGPT (Myeloma Generative Pretrained Transformer), an AI bioinformatician for multiple myeloma that relies on the CoMMpass dataset as its ground truth. MyeGPT converts natural language queries such as 'What are the characteristics of patients who relapse after induction therapy' or 'Compare the overall survival of high vs normal NSD2 expression' into de novo analyses backed on real data, then pro-actively generates plots to visualize the results. We develop a set of evaluation questions based on CoMMpass, complete with scoring criteria, and ran benchmarks to identify the best choice for LLMs and text-embedding models. We package MyeGPT as a ready-to-use browser application, enabling CoMMpass-grounded hypothesis validation from a smartphone.

3
Evaluation of AI-Generated Synthetic Data for Clinical Research in Secondary Cardiovascular Prevention among Dyslipidemia Patients

Bonomi, A.; Werba, J. P.; Saccani, S.; Lu, L. L.; Coser, A.; Franchi, M.; Valsecchi, C.; Teruzzi, E.; Terragni, A.; Centenaro, C.; Scatigna, M.; Pompilio, G.

2026-06-15 cardiovascular medicine 10.64898/2026.06.12.26355456 medRxiv
Top 0.1%
27.9%
Show abstract

Background: Access to high-quality clinical data is essential for advancing medical research and developing effective medical statistical and Artificial Intelligence models. However, privacy regulations and logistical barriers often hinder timely access to real-world data. Synthetic data offer a promising solution, preserving the statistical characteristics of original datasets while protecting patient privacy. Objectives: This study investigates the use of synthetic data for secondary cardiovascular prevention in patients with dyslipidemia, using two real-world datasets from Centro Cardiologico Monzino. Methods: Given the high dimensionality and limited sample size of the datasets, we employed a custom generative framework based on Large Language Models (LLMs). Pre-trained LLMs were fine-tuned on original clinical records to synthesize tabular data replicating source-data distributions. Fine-tuning was performed within the Centro Cardiologico Monzino's secure infrastructure to ensure data sovereignty. We evaluate clinical utility and privacy using fidelity and privacy metrics, identifying the optimal generative model and benchmarking against traditional anonymization methods. Results: Synthetic data achieved a superior trade-off than classically anonymized datasets. Real and synthetic datasets showed strong agreement, with significant distributional differences limited to few variables. Models trained on synthetic data replicated key associations from the original dataset, including therapy modification and creatine phosphokinase as predictors of SAMS, and pharmacological intensity as the main driver of LDL-C reduction. Conclusions: Results support the feasibility of using synthetic data as a proxy for real-world datasets in exploratory analyses and model development. Despite slight attenuation of some effect sizes, preserved clinical relationships reinforce the validity of synthetic data in medical research.

4
DeepAtlas: a tool for effective manifold learning

Hughes, S.; Hamilton, T.; Kolokotrones, T.; Deeds, E. J.

2025-08-31 systems biology 10.1101/2025.08.26.672474 medRxiv
Top 0.1%
26.5%
Show abstract

Manifold learning builds on the "manifold hypothesis," which posits that data in high-dimensional datasets are drawn from lower-dimensional manifolds. Current tools generate global embeddings of data, rather than the local maps used to define manifolds mathematically. These tools also cannot assess whether the manifold hypothesis holds true for a dataset. Here, we describe DeepAtlas, an algorithm that generates lower-dimensional representations of the datas local neighborhoods, then trains deep neural networks that map between these local embeddings and the original data. Topological distortion is used to determine whether a dataset is drawn from a manifold and, if so, its dimensionality. Application to test datasets indicates that DeepAtlas can successfully learn manifold structures. Interestingly, many real datasets, including single-cell RNA-sequencing, do not conform to the manifold hypothesis. In cases where data is drawn from a manifold, DeepAtlas builds a model that can be used generatively and promises to allow the application of powerful tools from differential geometry to a variety of datasets.

5
Domain specific models outperform large vision language models on cytomorphology tasks

Kukuljan, I.; Dasdelen, M. F.; Schaefer, J.; Buck, M.; Goetze, K.; Marr, C.

2025-05-06 hematology 10.1101/2025.05.05.25326989 medRxiv
Top 0.1%
26.5%
Show abstract

Large vision-language models (LVLMs) show impressive capabilities in image understanding across domains. However, their suitability for high-risk medical diagnostics remains unclear. We systematically evaluate four state-of-the-art LVLMs and three domain-specific models on key cytomorphological benchmarks: peripheral blood cell classification, morphology assessment, bone marrow cell classification, and cervical smear malignancy detection. Performance is assessed under zero-shot, few-shot, and fine-tuned conditions. LVLMs underperform significantly: the best LVLM achieves a zero-shot F1 score of 0.057 {+/-} 0.008 for malignancy detection--near random (0.039)--and only 0.15 {+/-} 0.01 in few-shot. In contrast, domain-specific models reach up to 0.83 in accuracy. Even after fine-tuning, a dedicated hematology model outperforms GPT-4o. While LVLMs offer explainability via text, we find the visual-language grounding unreliable, and the morphological features mention by the model often do not match the single cell properties. Our findings suggest that LVLMs require substantial improvements before use in high-stakes diagnostic settings. Key findingsO_LILVLMs perform poorly on cytomorphology tasks, often near chance level and far below domain-specific models. C_LIO_LIEven after fine-tuning, LVLMs lag behind domain-specific models. C_LIO_LIWhile LVLMs provide textual justifications, these often reflect generic descriptions rather than image-specific morphological features. C_LI

6
Transformer-based artificial intelligence on single-cell clinical data for homeostatic mechanism inference and rational biomarker discovery

Tozzo, V.; Zhang, L. H.; Ranganath, R.; Higgins, J. M.

2025-03-25 hematology 10.1101/2025.03.24.25324556 medRxiv
Top 0.1%
25.9%
Show abstract

Artificial intelligence (AI) applied to single-cell data has the potential to transform our understanding of biological systems by revealing patterns and mechanisms that simpler traditional methods miss. Here, we develop a general-purpose, interpretable AI pipeline consisting of two deep learning models: the Multi- Input Set Transformer++ (MIST) model for prediction and the single-cell FastShap model for interpretability. We apply this pipeline to a large set of routine clinical data containing single-cell measurements of circulating red blood cells (RBC), white blood cells (WBC), and platelets (PLT) to study population fluxes and homeostatic hematological mechanisms. We find that MIST can use these single-cell measurements to explain 70-82% of the variation in blood cell population sizes among patients (RBC count, PLT count, WBC count), compared to 5-20% explained with current approaches. MISTs accuracy implies that substantial information on cellular production and clearance is present in the single-cell measurements. MIST identified substantial crosstalk among RBC, WBC, and PLT populations, suggesting co-regulatory relationships that we validated and investigated using interpretability maps generated by single-cell FastShap. The maps identify granular single-cell subgroups most important for each populations size, enabling generation of evidence-based hypotheses for co-regulatory mechanisms. The interpretability maps also enable rational discovery of a single-WBC biomarker, "Down Shift", that complements an existing marker of inflammation and strengthens diagnostic associations with diseases including sepsis, heart disease, and diabetes. This study illustrates how single-cell data can be leveraged for mechanistic inference with potential clinical relevance and how this AI pipeline can be applied to power scientific discovery.

7
OpenScientist: evaluating an open agentic AI co-scientist to accelerate biomedical discovery

Roberts, K. F.; Abrams, Z. B.; Cappelletti, L.; Moqri, M.; Heugel, N.; Caufield, J. H.; Bourdenx, M.; Li, Y.; Banerjee, J.; Foschini, L.; Galeano, D.; Harris, N. L.; Li, M.; Ying, K.; Melendez, J. A.; Barthelemy, N. R.; Bollinger, J. G.; He, Y.; Ovod, V.; Benzinger, T. L. S.; Flores, S.; Gordon, B.; Ojewole, A. A.; Phatak, M.; Elbert, D. L.; Biber, S.; Landsness, E. C.; Mungall, C. J.; Bateman, R. J.; Reese, J.

2026-03-18 health informatics 10.64898/2026.03.15.26348338 medRxiv
Top 0.1%
23.1%
Show abstract

BackgroundAdvances in medicine depend on analyzing large and complex data sources, but discovery is partly constrained by the limited time and domain expertise of human researchers. Agentic artificial intelligence (agentic AI) can accelerate discovery by automating components of the scientific workflow, including information retrieval, data analysis, and knowledge synthesis. AimOpenScientist, an open-source agentic AI co-scientist, aims to accelerate biomedical discovery by semi-autonomously investigating scientist-defined queries and generating clinically relevant, verifiable scientific insights. MethodsDomain experts evaluated OpenScientist for novel discoveries in four clinical case studies: (1) a prespecified analysis in a community-based Alzheimers disease biomarker cohort, (2) unsupervised modeling for plasma proteomic survival prediction, (3) hypothesis investigation in single-cell transcriptomic data from neurons with neurofibrillary tangles, and (4) hypothesis generation with validation in a multiple myeloma dataset with a randomized negative control. ResultsOpenScientist completed analyses in minutes that otherwise would take weeks to months of human time and expertise. It identified %ptau217 as the best predictor of amyloid PET status, generated a plasma proteomic survival model with performance comparable to published models, proposed a mechanism linking tau pathology to altered lysosomal acidification, and generated multiple myeloma hypotheses that were validated in an external cohort while distinguishing true signal from randomized controls. ConclusionOpenScientist demonstrates that open, auditable, agentic AI can support real-world clinical research by generating hypotheses, executing analyses, and discovering insights from complex datasets.

8
Differential Privacy Protection Against Membership Inference Attack on Machine Learning for Genomic Data

Chen, J.; Wang, W. H.; Shi, X.

2020-08-04 bioinformatics 10.1101/2020.08.03.235416 medRxiv
Top 0.1%
22.3%
Show abstract

Machine learning is powerful to model massive genomic data while genome privacy is a growing concern. Studies have shown that not only the raw data but also the trained model can potentially infringe genome privacy. An example is the membership inference attack (MIA), by which the adversary, who only queries a given target model without knowing its internal parameters, can determine whether a specific record was included in the training dataset of the target model. Differential privacy (DP) has been used to defend against MIA with rigorous privacy guarantee. In this paper, we investigate the vulnerability of machine learning against MIA on genomic data, and evaluate the effectiveness of using DP as a defense mechanism. We consider two widely-used machine learning models, namely Lasso and convolutional neural network (CNN), as the target model. We study the trade-off between the defense power against MIA and the prediction accuracy of the target model under various privacy settings of DP. Our results show that the relationship between the privacy budget and target model accuracy can be modeled as a log-like curve, thus a smaller privacy budget provides stronger privacy guarantee with the cost of losing more model accuracy. We also investigate the effect of model sparsity on model vulnerability against MIA. Our results demonstrate that in addition to prevent overfitting, model sparsity can work together with DP to significantly mitigate the risk of MIA.

9
Attention Guided Mechanism Interpretable Drug-Gene Interaction (MIDI) Modeling for Cancer Drug Response Prediction and Target Effect Explanation

Wanyan, T.; Cai, L.; Nijhawan, D.; Posner, B.; Kim, J.; Yuan, R.; Wang, R.; Huang, J.; Yang, S.; Mallipeddi, P.; Wang, T.; Zhan, X.; Xiao, G.; Xie, Y.

2025-04-02 cancer biology 10.1101/2025.03.31.646490 medRxiv
Top 0.1%
22.3%
Show abstract

Cancer drug discovery using genetic information is still poorly developed. Precisely locating drug atoms and explaining the targeting effect is crucial in precision medicine since it helps understand the drugs mechanism of action. Much data has been collected regarding drug response against cancer cell lines, and many models predict the drug response based on genomic information. However, to our knowledge, none of the data-driven techniques propose to detect the targeting mechanism of small drug molecules against genetic targets. In this work, we propose MIDI (Mechanism Interpretable Drug-Gene Interaction) model to delve deep into the targeting relation between drug molecules against genetic patterns. We show that purely based on a data-driven approach, the attention mechanism in our model could capture the important binding effect of small molecules towards gene targets. We provide both theoretical derivation and experiment results to show the information flow regarding the attention mechanism. In the meantime, we demonstrate that our model presents much higher prediction performance with the interpretation mechanism than the other state-of-the-art drug response prediction models.

10
Emerging Concern of Scientific Fraud: Deep Learning and Image Manipulation

Qi, C.; Zhang, J.; Luo, P.

2021-01-17 scientific communication and education 10.1101/2020.11.24.395319 medRxiv
Top 0.1%
22.1%
Show abstract

Scientific fraud by image duplications and manipulations within western blot images is a rising problem. Currently, problematic western blot images are mainly detected by checking repeated bands or through visual observation. However, the completeness of the above methods in detecting problematic images has not been demonstrated. Here we show that Generative Adversarial Nets (GANs) can generate realistic western blot images that indistinguishable from real western blots. The overall accuracy of researchers for identifying synthetic western blot images is 0.52, which almost equal to blind guess (0.5). We found that GANs can generate western blot images with bands of the expected lengths, widths, and angles in desired positions that can fool researchers. For the case study, we find that the accuracy of detecting the synthetic western blot images is related to years of researchers performed studies relevant to western blots, but there was no apparent difference in accuracy among researchers with different academic degrees. Our results demonstrate that GANs can generate fake western blot images to fool existing problematic image detection methods. Therefore, more information is needed to ensure that the western blots appearing in scientific articles are real. We argue to require every western blot image to be uploaded along with a unique identifier generated by the laboratory machine and to peer review these images along with the corresponding submitted articles, which may reduce the incidence of scientific fraud.

11
Cracking the code of co-authorship networks geo-temporally using interpretable machine learning

Keshari, S.; Rarani, Z. H.; Kishore, A.; Das, J.

2025-03-07 scientific communication and education 10.1101/2025.03.05.641725 medRxiv
Top 0.1%
21.8%
Show abstract

An exponential growth in the scientific literature necessitates the development of highly scalable computational tools that can effectively analyze and distill insights from complex, interconnected research landscapes. We introduce Distributed, Interpretable, and Scalable computing for Co-authorship Networks (DISCo-Net), a robust and scalable tool engineered to curate and examine large-scale co-authorship networks by harnessing the power of distributed computing and advanced relational database queries. We use DISCo-Net to analyze co-authorship networks derived from millions of papers in the life sciences and physical sciences over more than two decades. Using a range of deep learning approaches, we surprisingly found that pre-trained zero-shot embeddings from a sentence transformer better captured global co-authorship relationships than a complex graphical attention transformer. Even more surprisingly, a simple interpretable Term Frequency-Inverse Document Frequency (TF-IDF) model performed as well as the Bidirectional Encoder Representations from Transformers (BERT) model. Through topic modeling on TF-IDF document descriptors, we identified nine major research areas prevalent globally over the past 24 years and captured topic-specific shifting trends in scientific output. Our study draws an innovative parallel between collaborative research networks and genomic regulatory structures, applying genomics data analysis methodologies to uncover patterns in global scientific collaboration. This approach reveals interpretable alignments between research interests and human developmental stages, while also identifying emerging influential players in the global research landscape. The findings highlight potential far-reaching consequences of current funding challenges, particularly in the U.S., and offer actionable insights for optimizing resource allocation and fostering innovation in an interconnected global scientific community.

12
Machine Learning the Phenomenology of COVID-19 From Early Infection Dynamics

Magdon-Ismail, M.

2020-03-20 health informatics 10.1101/2020.03.17.20037309 medRxiv
Top 0.1%
19.3%
Show abstract

We present a robust data-driven machine learning analysis of the COVID-19 pandemic from its early infection dynamics, specifically infection counts over time. The goal is to extract actionable public health insights. These insights include the infectious force, the rate of a mild infection becoming serious, estimates for asymtomatic infections and predictions of new infections over time. We focus on USA data starting from the first confirmed infection on January 20 2020. Our methods reveal significant asymptomatic (hidden) infection, a lag of about 10 days, and we quantitatively confirm that the infectious force is strong with about a 0.14% transition from mild to serious infection. Our methods are efficient, robust and general, being agnostic to the specific virus and applicable to different populations or cohorts.

13
Use of Federated Learning for validating and updating privacy-preserving decentralized multi-study prognostic models in Traumatic Brain Injury

Torres-Espin, A.; Wong, J. C.; Hinson, H. E.; Kuipers, T. B.; Hoekstra, B. P. T.; Jain, S.; Sun, X.; Yue, J. K.; Pisi?, D.; Mikolic, A.; Lingsma, H. F.; Markowitz, A. J.; Ferguson, A. R.; Menon, D. K.; Maas, A. I. R.; Steyerberg, E. W.; Manley, G. T.; Belton, P. J.

2026-08-02 health informatics 10.64898/2026.07.30.26359337 medRxiv
Top 0.1%
19.0%
Show abstract

Developing modern clinical prediction models (CPMs) and advanced analytics requires large datasets, often necessitating data from different studies. Privacy regulations may hinder data sharing, especially across countries. Decentralized federated data infrastructures, where data remain in their original location and analyses are run only in a shared, secure environment, may address these challenges. We implemented a privacy-preserving federated learning (FL) infrastructure and evaluated and updated the IMPACT prognostic models for traumatic brain injury (TBI) using 2 studies. A multi-continental federated infrastructure was established between 2 large-scale studies (TRACK-TBI from the United States and CENTER-TBI from Europe and Israel). Three IMPACT prognostic models for post-TBI 6-month mortality and unfavorable outcomes were evaluated, followed by model updates through 2 FL approaches trained across the TRACK-TBI and CENTER-TBI studies. Internal validation, external cross-validation, and sub-study validations were performed. CPMs were evaluated for discrimination and calibration. The federated cohort included 1616 participants (TRACK-TBI: n=441, CENTER-TBI: n=1175). Both FL performed well, with comparable coefficient estimates, AUCs (area under the receiver operating characteristics curve) between 0.77-0.88, and calibrated probabilities. Compared to the original IMPACT and single-study models, both federated models presented similar discrimination (AUC), were well-calibrated, were more efficient (higher precision), and reduced the impact of missing data in model estimation. FL is feasible for privacy-preserving development and evaluation of CPMs, and can enable validation and updating across large, virtually analyzed datasets while overcoming regulatory constraints on data combination. Federated infrastructures can facilitate global collaboration to advance data-hungry analytical methods, such as artificial intelligence.

14
The results of Transcriptome-wide Mendelian Randomization (TWMR) in large-scale populations can directly validate, across scales, the results of causal inference from deep learning combined with double machine learning on single-cell transcriptomes of human samples.

ye, w.; Jiang, X.; Shen, F.

2026-03-19 rheumatology 10.64898/2026.03.16.26348532 medRxiv
Top 0.1%
18.7%
Show abstract

ObjectiveAiming at the core problems prevalent in biomedical research, including the "translational distance", the difficulty in aligning cross-scale studies, and the lack of direct validation of single-cell systems biology models in human samples, this study aims to verify whether the results of transcriptome-wide Mendelian randomization (TWMR) based on large-scale populations are consistent with the causal inference results of deep learning combined with double machine learning (DML) using single-cell transcriptome data from human samples, to clarify whether statistical biology and systems biology can converge to the same biological truth, and provide methodological support for mechanism dissection and precision medicine research of complex diseases such as rheumatoid arthritis (RA). MethodsThis study integrated multi-omics data to conduct a two-stage causal inference and cross-scale validation analysis. In the first stage, based on the summary statistics of RA genome-wide association study (GWAS) from 456,348 individuals of European ancestry in the UK Biobank (UKB), and cis-expression quantitative trait locus (cis-eQTL) data from 31,684 individuals in the eQTLGen Consortium, a two-sample Mendelian randomization approach was adopted. Transcriptome-wide causal effect analysis was performed using the inverse-variance weighted (IVW) method, MR Egger regression, and weighted median method, and gene-level causal effect values were obtained after strict quality control and multiple testing correction. In the second stage, based on single-cell RNA sequencing (scRNA-seq) data from RA patients and healthy controls (RA group: 11 samples, 211,867 cells; Healthy control group: 38 samples, 456,631 cells), after preprocessing via the Seurat pipeline, batch effect correction, and cell type annotation, a hierarchical deep neural network was constructed to complete feature compression of high-dimensional expression data, and the DML framework was used to estimate the causal effects of genes on RA disease status. Finally, Pearson correlation analysis was performed to conduct cell type-specific cross-scale validation of gene-level causal effect values obtained by the two methods, and the validated model was used to quantify the causal effects of 16 RA-related pathways from the Reactome database. ResultsThis study confirmed that the gene causal effect values obtained from large-scale population TWMR analysis were significantly correlated with those calculated by the deep learning combined with DML model based on single-cell transcriptome data. Among them, the correlation was extremely significant (p<0.001) in core naive B cells (r=0.202, p=3.2e-05, n=414) and core naive CD4 T cells (r=0.102, p=0.037, n=412). The validated DML model successfully quantified the cell type-specific causal effect values of 16 RA-related signaling pathways. ConclusionStatistical biology and systems biology can converge to the same biological truth. The cross-scale consistency between the two can significantly shorten the "translational distance" in biomedical research, and realizes the direct validation of the single-cell systems biology causal model of human samples based on large-scale population genetic data, getting rid of the excessive dependence on animal/cell experimental models in traditional research. This research paradigm not only provides a new path for mechanism dissection and therapeutic target screening of complex diseases such as RA, but also provides a feasible solution for rare disease research to break through the limitation of GWAS sample size, and lays an important theoretical and methodological foundation for constructing standardized systems biology models of human complex diseases and promoting the development of precision medicine.

15
Reward-Guided Generation Improves the Scientific Utility of Synthetic Biomedical Data

Jackson, N. J.; Espinosa-Dice, N.; Yan, C.; Malin, B. A.

2026-03-16 health informatics 10.64898/2026.03.11.26348077 medRxiv
Top 0.1%
18.7%
Show abstract

Synthetic data generation is a promising approach for biomedical data sharing and dataset augmentation, yet existing methods lack mechanisms to preserve statistical properties necessary for scientific analysis. To address this, we introduce RLSYN+REG, a reinforcement learning-driven generative model, which encourages that regression models trained on synthetic data reproduce the coefficients and predictions of their real-data counterparts. We evaluate RL-SO_SCPLOWYNC_SCPLOW+RO_SCPLOWEGC_SCPLOW on MIMIC-III and the American Community Survey (ACS) across regression model reproduction, fidelity to real data, and privacy. Synthetic data from RLSO_SCPLOWYNC_SCPLOW+RO_SCPLOWEGC_SCPLOW substantially improves upon that of RLSO_SCPLOWYNC_SCPLOW, raising correlations between real and synthetic regression coefficients from 0.054 to 0.600 on MIMIC-III and from 0.160 to 0.376 on ACS. Predictive performance also improves, reducing the gap between real-data baselines by 81.4% and 97.6% on MIMIC-III and ACS, respectively. These improvements come with negligible cost to fidelity or privacy and are robust to reductions in training data.

16
LORE: A Literature Semantics Framework for Evidenced Disease-Gene Pathogenicity Prediction at Scale

Li, P.-H.; Sun, Y.-Y.; Juan, H.-F.; Chen, C.-Y.; Tsai, H.-K.; Huang, J.-H.

2024-08-11 health informatics 10.1101/2024.08.10.24311801 medRxiv
Top 0.1%
18.5%
Show abstract

Effective utilization of academic literature is crucial for Machine Reading Comprehension to generate actionable scientific knowledge for wide real-world applications. Recently, Large Language Models (LLMs) have emerged as a powerful tool for distilling knowledge from scientific articles, but they struggle with the issues of reliability and verifiability. Here, we propose LORE, a novel unsupervised two-stage reading methodology with LLM that models literature as a knowledge graph of verifiable factual statements and, in turn, as semantic embeddings in Euclidean space. Applied to PubMed abstracts for large-scale understanding of disease-gene relationships, LORE captures essential information of gene pathogenicity. Furthermore, we demonstrate that modeling a latent pathogenic flow in the semantic embedding with supervision from the ClinVar database leads to a 90% mean average precision in identifying relevant genes across 2,097 diseases. Finally, we have created a disease-gene relation knowledge graph with predicted pathogenicity scores, 200 times larger than the ClinVar database.

17
Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula

Bahmani, A.; Cha, K.; Alavi, A.; Dixit, A.; Ross, A.; Park, R.; Goncalves, F.; Ma, S.; Saxman, P.; Nair, R.; Sarraf, R. A.; Zhou, X.; Wang, M.; Contrepois, K.; Than, J. L. P.; Monte, E.; Rodriguez, D. J. F.; Lai, J.; Babu, M.; Tondar, A.; Rose, S. M. S.-F.; Akbari, I.; Zhang, X.; Yegnashankaran, K.; Yracheta, J.; Dale, K.; Miller, A. D.; Edmiston, S.; McGhee, E. M.; Nebeker, C.; Wu, J. C.; Kundaje, A.; Snyder, M.

2024-08-01 health informatics 10.1101/2024.07.31.24311182 medRxiv
Top 0.1%
18.5%
Show abstract

BackgroundPrecision medicine promises significant health benefits but faces challenges such as complex data management and analytics, interdisciplinary collaboration, and education of researchers, healthcare professionals, and participants. Addressing these needs requires the integration of computational experts, engineers, designers, and healthcare professionals to develop user-friendly systems and shared terminologies. The widespread adoption of large language models (LLMs) such as Generative Pretrained Transformer (GPT) and Claude highlights the importance of making complex data accessible to non-specialists. MethodsWe evaluated the Stanford Data Ocean (SDO) precision medicine training programs learning outcomes, AI Tutor performance, and learner satisfaction by assessing self-rated competency on key learning objectives through pre- and post-learning surveys, along with formative and summative assessment completion rates. We also analyzed AI Tutor accuracy and learners self-reported satisfaction, and post-program academic and career impacts. Additionally, we demonstrated the capabilities of the AI Data Visualization tool. ResultsSDO demonstrates the ability to improve learning outcomes for learners from broad educational and socioeconomic backgrounds with the support of the AI Tutor. The AI Data Visualization tool enables learners to interpret multi-omics and wearable data and replicate research findings. ConclusionsSDO strives to mitigate challenges in precision medicine through a scalable, cloud-based platform that supports data management for various data types, advanced research, and personalized learning. SDO provides AI tutors and AI-powered data visualization tools to enhance educational and research outcomes and make data analysis accessible to users from broad educational backgrounds. By extending engagement and cutting-edge research capabilities globally, SDO particularly benefits economically disadvantaged and historically marginalized communities, fostering interdisciplinary biomedical research and bridging the gap between education and practical application in the biomedical field. Plain Language SummaryPrecision medicine is the use of various types of health data specific to an individual to improve disease prevention, diagnosis, or treatment. We used artificial intelligence to build a precision medicine learning platform for clinicians and researchers in training. Students in 95 countries accessed the platform and found it helpful. It could be particularly helpful for training students in low- and middle-income countries.

18
A Privacy-Preserving Zero-Code Conversational Statistical Analysis System for Clinical Research Using Agentic AI and Local R Execution

Yang, S.; Chen, V. L.; Ng, W. H.; Zhang, S.; Qiu, S.; Zhu, J.; Hsieh, T. Y.-J.; Ji, F.; Yeo, Y. H.

2026-08-02 health informatics 10.64898/2026.07.30.26359367 medRxiv
Top 0.1%
17.9%
Show abstract

Background Clinical data analysis typically requires statistical programming skills, whereas cloud-based artificial intelligence (AI) agents risk exposing sensitive patient records. We developed and functionally validated a privacy-preserving, zero-code conversational statistical analysis framework that translates natural-language clinical research requests into executable R workflows while strictly retaining raw patient data within local computing environments. Methods Orchestrated by the n8n engine, the system integrates the DeepSeek-Reasoner model with a Pinecone vector database for retrieval-augmented generation (RAG), grounding statistical selection in curated biostatistical guidance and R templates. Core functionalities include data schema perception, interactive data cleaning, requirements refinement, and local R code execution via a controlled command-line interface. System performance was evaluated by replicating a published prognostic model study on metabolic dysfunction-associated steatotic liver disease (MASLD). Findings All core analytical workflows, including data cleaning, multivariable Cox proportional hazards modeling, model diagnostics, and publication-ready tables and figures (e.g., baseline characteristics, Schoenfeld residuals, receiver operating characteristic curves, and forest plots), were executed solely through natural-language dialogues without manual coding. The external large language model actively clarified analytical prompts while receiving zero row-level patient data. Interpretation Decoupling remote cloud reasoning from local code execution lowers the technical threshold for clinicians conducting data-driven research while safeguarding data privacy. This architecture provides a practical, scalable, and reproducible framework for converting natural-language clinical questions into executable statistical workflows. Funding National Natural Science Foundation of China (82473291), Shaanxi Province "Three Qin Scholars" Innovation Team Project (2023001), and Fundamental Research Funds for the Central Universities (xtr062023003).

19
Differentiating central nervous system demyelinating disorders using Graph Attention Networks

Choudalakis, S.; Rapti, A.; Karathanasis, D.; Kastis, G. A.; Evangelopoulos, M.-E.; Mavragani, C. P.; Dikaios, N.

2026-01-02 rheumatology 10.64898/2025.12.29.25343068 medRxiv
Top 0.1%
15.4%
Show abstract

Autoimmune demyelinating central nervous system (CNS) disorders encompass a wide array of clinical entities ranging from the organ specific Multiple Sclerosis (MS) to systemic autoimmune diseases (SADs) such as Systemic Lupus Erythematosus (SLE) and Sjogrens syndrome (SS). Despite international research efforts, distinction of these entities at clinical, imaging and laboratory level remains challenging, with almost 20% of patients being misdiag-nosed with MS, out of which more than 50% carry the misdiagnosis for at least 3 years, while 5% are misdiagnosed for 20 years or more. This work aims to identify early biomarkers that can discriminate MS from other clinical entities that are MS mimickers. For this reason, we have meticulously curated a unique biobank with serum, DNA, RNA, and cerebrospinal fluid (CSF) samples from 396 treatment-naive patients who presented with a demyelinating episode and for which we have recorded over 300 clinical, serological, and imaging parameters. These patients have undergone follow up at 6 months and 12 months from their first demyelinating episodes and have been categorised as follows: MS, SAD with CNS involvement, and demyelination with autoimmune features (DAF). A hybrid model for semi-supervised tabular classification is proposed that integrates a graph attention network with dynamic graph learning via Random Forest proximity and k-NN graphs with tabular prior-data fitted network for direct classification. The model also performed feature selection, self-supervised contrastive learning, self-training and data augmentation.

20
Exploratory electronic health record analysis with ehrapy

Heumos, L.; Ehmele, P.; Treis, T.; Upmeier zu Belzen, J.; Namsaraeva, A.; Horlava, N.; Shitov, V. A.; Zhang, X.; Zappia, L.; Knoll, R.; Lang, N. J.; Hetzel, L.; Virshup, I.; Sikkema, L.; Roellin, E.; Curion, F.; Eils, R.; Schiller, H. B.; Hilgendorff, A.; Theis, F.

2023-12-11 health informatics 10.1101/2023.12.11.23299816 medRxiv
Top 0.1%
15.3%
Show abstract

With progressive digitalization of healthcare systems worldwide, large-scale collection of electronic health records (EHRs) has become commonplace. However, an extensible framework for comprehensive exploratory analysis that accounts for data heterogeneity is missing. Here, we introduce ehrapy, a modular open-source Python framework designed for exploratory end-to-end analysis of heterogeneous epidemiology and electronic health record data. Ehrapy incorporates a series of analytical steps, from data extraction and quality control to the generation of low-dimensional representations. Complemented by rich statistical modules, ehrapy facilitates associating patients with disease states, differential comparison between patient clusters, survival analysis, trajectory inference, causal inference, and more. Leveraging ontologies, ehrapy further enables data sharing and training EHR deep learning models paving the way for foundational models in biomedical research. We demonstrated ehrapys features in five distinct examples: We first applied ehrapy to stratify patients affected by unspecified pneumonia into finer-grained phenotypes. Furthermore, we revealed biomarkers for significant differences in survival among these groups. Additionally, we quantify medication-class effects of pneumonia medications on length of stay. We further leveraged ehrapy to analyze cardiovascular risks across different data modalities. Finally, we reconstructed disease state trajectories in SARS-CoV-2 patients based on imaging data. Ehrapy thus provides a framework that we envision will standardize analysis pipelines on EHR data and serve as a cornerstone for the community.