Back

Journal of Biomedical Informatics

Elsevier BV

All preprints, ranked by how well they match Journal of Biomedical Informatics's content profile, based on 47 papers previously published here. The average preprint has a 0.07% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Synonym Augmentation for Rare Disease Identification in Unstructured Data

Valinejad, J.; Moon, S.; Xu, Y.; Zhu, Q.

2026-05-13 health informatics 10.64898/2026.05.11.26352910 medRxiv
Top 0.1%
61.4%
Show abstract

The significant challenges associated with rare diseases in the medical and research domains include the scarcity of information, which is often confined to unstructured formats. Although existing approaches provide valuable insights, there is a need to develop effective methods to identify information pertinent to rare diseases for advancing rare disease research. We identified mentions of rare diseases in relevant texts and assessed their relevance using derived scores, the confidence score and semantic similarity from a fine-tuned BioMedBERT encoder. This encoder was fine-tuned using rare disease related text from Online Mendelian Inheritance in Man (OMIM), Orphanet, a manually validated dataset, and STS benchmark datasets. The process of identifying meaningful rare disease mentioned was presented through two case studies that retrieved relevant NIH-funded projects, utilizing a generated knowledge graph in Neo4j to host data on 2,067 GARD diseases with over 320,000 NIH funded projects. Through various case studies with NIH-funded projects related to rare diseases, we demonstrated the effectiveness of our approach in systematically providing rare disease related data to enhance our understanding of rare diseases for future investigations.

2
Hierarchical Representation of Complex Intervention Sequences for Automated Subgroup Analysis in Critical Care Settings

Trivedi, A.; Ogallo, W.; Abebe, G.

2021-10-18 intensive care and critical care medicine 10.1101/2021.10.16.21265074 medRxiv
Top 0.1%
51.6%
Show abstract

Our understanding of the impact of interventions in critical care is limited by the lack of techniques that represent and analyze complex intervention spaces applied across heterogeneous patient populations. Existing work has mainly focused on selecting a few interventions and representing them as binary variables, resulting in oversimplification of intervention representation. The goal of this study is to find effective representations of sequential interventions to support intervention effect analysis. To this end, we have developed Hi-RISE (Hierarchical Representation of Intervention Sequences), an approach that transforms and clusters sequential interventions into a latent space, with the resulting clusters used for heterogenous treatment effect analysis. We apply this approach to the MIMIC III dataset and identified intervention clusters and corresponding subpopulations with peculiar odds of 28-day mortality. Our approach may lead to a better understanding of the subgrouplevel effects of sequential interventions and improve targeted intervention planning in critical care settings.

3
Summarizing Clinical Notes using LLMs for ICU Bounceback and Length-of-Stay Prediction

Choudhuri, A.; Polgreen, P.; Segre, A.; Adhikari, B.

2025-01-20 health informatics 10.1101/2025.01.19.25320797 medRxiv
Top 0.1%
51.5%
Show abstract

Recent advances in the Large Language Models (LLMs) provide a promising avenue for retrieving relevant information from clinical notes for accurate risk estimation of adverse patient outcomes. In this empirical study, we quantify the gain in predictive performance obtained by prompting LLMs to study the clinical notes and summarize potential risks for downstream tasks. Specifically, we prompt LLMs to generate a summary of progress notes and state potential complications that may arise. We then learn representations of the generated notes in sequential order and estimate the risks of patients in the ICU getting readmitted in ICU after discharge (ICU bouncebacks) and predict the overall length of stay in the ICU. Our analysis in the real-world MIMIC III dataset shows performance gains of 7.17% in terms of AUC-ROC and 14.16% in terms of AUPRC for the ICU bounceback task and 2.84% in terms of F-1 score and 7.12% in terms of AUPRC for the ICU LOS Prediction task. This demonstrates that the LLM-infused models outperform the approaches that only directly rely on clinical notes and other EHR data.

4
Using Large Language Models to Explore Mechanisms of Life Course Exposure-Outcome Associations

Wang, S.; Gao, Y.; Zhang, Y.; DU, J.

2024-10-17 health informatics 10.1101/2024.10.17.24315648 medRxiv
Top 0.1%
50.3%
Show abstract

Large language models (LLMs) enhanced with Graph Retrieval-Augmented Generation (GRAG) are promising for life-course epidemiology, which typically depends on costly and incomplete cohort data. Inspired by the epidemiological pathway model, we introduce EpiPathAI, which combines literature-derived causal knowledge graphs with LLMs to mine bridging variables and synthesize potential mechanisms between gestational diabetes and dementia. We test four GRAG strategies on GPT-4 and evaluate the identified mediators with clinical experts and three other LLM reviewers. The knowledge graph identifies 118 bridging variables, including coronary heart disease and chronic kidney disease, previously validated in our data-driven approach through the UK Biobank. EpiPathAI has identified additional clinically meaningful mediators, including high-level low-density lipoprotein (9.8% of effect, 95% CI: 3.7%-23.2%), and depression, which is a reasonable but statistically non-significant mediator in UK Biobank. EpiPathAI serves as a knowledge-driven mechanism mining agent that complements the data-driven approach, providing a compelling foundation for investigating other mediating pathways in future longitudinal cohort studies.

5
A Survey on Optimization and Machine -learning-based Fair Decision Making in Healthcare

Chen, Z.; Marrero, W. J.

2024-03-18 health systems and quality improvement 10.1101/2024.03.16.24304403 medRxiv
Top 0.1%
45.2%
Show abstract

The unintended biases introduced by optimization and machine learning (ML) models are a topic of great interest to medical researchers and professionals. Bias in healthcare decisions can cause patients from vulnerable populations (e.g., racially minoritized, low-income, or living in rural areas) to have lower access to resources and inferior outcomes, thus exacerbating societal unfairness. In this systematic literature review, we present a structured overview of the literature regarding fair decision making in healthcare until April 2024. After screening 801 unique references, we identified 114 articles within the scope of our review. In our review, we comprehensively examine fair decision-making methodologies in healthcare by systematically identifying and categorizing biases within both data and models. Initially, we elucidate existing bias within healthcare decision making. Then, we present a range of fairness metrics drawn from different use cases, followed by analyzing and classifying bias mitigation strategies into pre-processing, in-processing, and post-processing techniques. We provide a broad conceptual overview and practical illustrations of each approach. Additionally, we examine emerging bias mitigation technologies that, though not yet applied in healthcare, show substantial promise for future integration. Our review aims to increase awareness of fairness in healthcare decision making and facilitate the selection of appropriate approaches under varying scenarios.

6
Faithful AI in Healthcare and Medicine

Xie, Q.; Wang, F.

2023-04-19 health informatics 10.1101/2023.04.18.23288752 medRxiv
Top 0.1%
42.0%
Show abstract

Artificial intelligence (AI), especially the most recent large language models (LLMs), holds great promise in healthcare and medicine, with applications spanning from biological scientific discovery and clinical patient care to public health policymaking. However, AI methods have the critical concern for generating factually incorrect or unfaithful information, posing potential long-term risks, ethical issues, and other serious consequences. This review aims to provide a comprehensive overview of the faithfulness problem in existing research on AI in healthcare and medicine, with a focus on the analysis of the causes of unfaithful results, evaluation metrics, and mitigation methods. We systematically reviewed the recent progress in optimizing the factuality across various generative medical AI methods, including knowledge-grounded LLMs, text-to-text generation, multimodality-to-text generation, and automatic medical fact-checking tasks. We further discussed the challenges and opportunities of ensuring the faithfulness of AI-generated information in these applications. We expect that this review will assist researchers and practitioners in understanding the faithfulness problem in AI-generated information in healthcare and medicine, as well as the recent progress and challenges in related research. Our review can also serve as a guide for researchers and practitioners who are interested in applying AI in medicine and healthcare.

7
Deep Learning Approach to Parse Eligibility Criteria in Dietary Supplements Clinical Trials Following OMOP Common Data Model

Bompelli, A.; Li, J.; Xu, Y.; Wang, N.; Wang, Y.; Adam, T.; He, Z.; Zhang, R.

2020-09-18 health informatics 10.1101/2020.09.16.20196022 medRxiv
Top 0.1%
39.5%
Show abstract

Dietary supplements (DSs) have been widely used in the U.S. and evaluated in clinical trials as potential interventions for various diseases. However, many clinical trials face challenges in recruiting enough eligible patients in a timely fashion, causing delays or even early termination. Using electronic health records to find eligible patients who meet clinical trial eligibility criteria has been shown as a promising way to assess recruitment feasibility and accelerate the recruitment process. In this study, we analyzed the eligibility criteria of 100 randomly selected DS clinical trials and identified both computable and non-computable criteria. We mapped annotated entities to OMOP Common Data Model (CDM) with novel entities (e.g., DS). We also evaluated a deep learning model (Bi-LSTM-CRF) for extracting these entities on CLAMP platform, with an average F1 measure of 0.601. This study shows the feasibility of automatic parsing of the eligibility criteria following OMOP CDM for future cohort identification.

8
Early Detection of Rare Disease Using Hierarchical Set-to-Sequence Modeling of Structured Electronic Health Records

Ma, Y.; Chinthala, L.; Mohammed, A.; Davis, R. L.; Colonna, V.

2026-05-06 health informatics 10.64898/2026.05.04.26352393 medRxiv
Top 0.1%
39.0%
Show abstract

Rare diseases are characterized by heterogeneous, weak, and sparse phenotypic signals that emerge gradually across longitudinal clinical visits, making early detection a persistent challenge. In this study, we propose a hierarchical set-to-sequence (HSS) framework for prospective rare disease detection using structured EHR data. HSS decomposes the problem into two levels: (1) intra-visit encoding via Multi-Query Attention (MQA), which treats heterogeneous clinical events within a single clinical visit as an unordered set to generate unified visit-level representations, and (2) inter-visit temporal modeling with transformer encoders conditioned on patient visit age and inter-visit time gaps to capture the disease progression and the irregular intervals between clinical visits. We construct a real-world cohort of 40,223 patients comprising 708,422 visits from a single academic medical center (2005-2025), with 3,032 rare disease cases identified by curated rule-based phenotyping including severe neuro-developmental, congenital, or genetic conditions. We formulate the task as multi-horizon prospective binary classification with five prediction horizons of 7, 30, 90, 180, and 365 days prior to first diagnosis. Experimental results show that the proposed HSS model consistently outperforms linear logistic regression, tree-based XGBoost, and Transformer-based baselines at every prediction horizon, ranging from AUROC = 0.893 and AUPRC = 0.601 at 7 days with 5.17% prevalence to AUROC = 0.829 and AUPRC = 0.228 at 365 days with at 3.98% prevalence. Notably, the performance gap between HSS and the strongest competing baseline is largest at the 365 days horizon, indicating stronger advantages for long-horizon prediction where phenotypic signals for rare diseases are weak and sparse. Additional analyses further clarify the contribution of the hierarchical components and confirm the importance of hierarchical modeling. This work contributes to the ongoing development of AI methodologies tailored to rare diseases by introducing a hierarchical framework for early detection using structured longitudinal clinical data.

9
Distilling the Knowledge from Large-language Model for Health Event Prediction

Ding, S.; Ye, J.; Hu, X.; Zou, N.

2024-06-24 health informatics 10.1101/2024.06.23.24309365 medRxiv
Top 0.1%
38.7%
Show abstract

Health event prediction is empowered by the rapid and wide application of electronic health records (EHR). In the Intensive Care Unit (ICU), precisely predicting the health related events in advance is essential for providing treatment and intervention to improve the patients outcomes. EHR is a kind of multi-modal data containing clinical text, time series, structured data, etc. Most health event prediction works focus on a single modality, e.g., text or tabular EHR. How to effectively learn from the multi-modal EHR for health event prediction remains a challenge. Inspired by the strong capability in text processing of large language model (LLM), we propose the framework CKLE for health event prediction by distilling the knowledge from LLM and learning from multi-modal EHR. There are two challenges of applying LLM in the health event prediction, the first one is most LLM can only handle text data rather than other modalities, e.g., structured data. The second challenge is the privacy issue of health applications requires the LLM to be locally deployed, which may be limited by the computational resource. CKLE solves the challenges of LLM scalability and portability in the healthcare domain by distilling the cross-modality knowledge from LLM into the health event predictive model. To fully take advantage of the strong power of LLM, the raw clinical text is refined and augmented with prompt learning. The embedding of clinical text are generated by LLM. To effectively distill the knowledge of LLM into the predictive model, we design a cross-modality knowledge distillation (KD) method. A specially designed training objective will be used for the KD process with the consideration of multiple modality and patient similarity. The KD loss function consists of two parts. The first one is cross-modality contrastive loss function, which models the correlation of different modalities from the same patient. The second one is patient similarity learning loss function to model the correlations between similar patients. The cross-modality knowledge distillation can distill the rich information in clinical text and the knowledge of LLM into the predictive model on structured EHR data. To demonstrate the effectiveness of CKLE, we evaluate CKLE on two health event prediction tasks in the field of cardiology, heart failure prediction and hypertension prediction. We select the 7125 patients from MIMIC-III dataset and split them into train/validation/test sets. We can achieve a maximum 4.48% improvement in accuracy compared to state-of-the-art predictive model designed for health event prediction. The results demonstrate CKLE can surpass the baseline prediction models significantly on both normal and limited label settings. We also conduct the case study on cardiology disease analysis in the heart failure and hypertension prediction. Through the feature importance calculation, we analyse the salient features related to the cardiology disease which corresponds to the medical domain knowledge. The superior performance and interpretability of CKLE pave a promising way to leverage the power and knowledge of LLM in the health event prediction in real-world clinical settings.

10
Improving Clinical Decision Making with a Two-Stage Recommender System Based on Language Models: A Case Study on MIMIC-III Dataset

Raza, S.

2023-02-24 health informatics 10.1101/2023.02.21.23286247 medRxiv
Top 0.1%
38.7%
Show abstract

Clinical decision-making is a challenging and time-consuming task that involves integrating a vast amount of patient data, including medical history, test results, and notes from clinicians. To assist this process, clinical recommender systems have been developed to provide personalized recommendations to healthcare practitioners. However, creating effective clinical recommender systems is complex due to the diversity and intricacy of clinical data and the need for customized recommendations. In this paper, we propose a two-stage recommender framework for clinical decision-making basedon the publicly available MIMIC dataset of electronic health records. The first stage of the framework employs a deep neural networkbased model to retrieve a set of candidate items, such as diagnosis, medication, and prescriptions, from the patients electronic health records. The model is trained to extract relevant information from clinical notes using a pre-trained language model. The second stage of the framework utilizes a deep learning model to rank and recommend the most pertinent items to healthcare providers. The model considers the patients medical history and the context of the current visit to offer personalized recommendations. To evaluate the proposed model, we compared it to various baseline models using multiple evaluation metrics. The findings indicate that the proposed model achieved a precision of 89% and a macro-average F1 score of approximately 84%, indicating its potential to improve clinical decision-making and reduce information overload for healthcare providers. The paper also discusses challenges, such as data availability, privacy, and bias, and suggests areas for future research in this field.

11
A Novel Explainable AI Method to Assess Associations between Temporal Patterns in Patient Trajectories and Adverse Outcome Risks: Analyzing Fitness as a Risk Factor of ADRD

Shao, Y.; Zamrini, E. Y.; Ahmed, A.; Cheng, Y.; Nelson, S. J.; Kokkinos, P.; Zeng-Treitler, Q.

2024-05-17 health informatics 10.1101/2024.05.17.24307541 medRxiv
Top 0.1%
38.4%
Show abstract

We present a novel explainable artificial intelligence (XAI) method to assess the associations between the temporal patterns in the patient trajectories recorded in longitudinal clinical data and the adverse outcome risks, through explanations for a type of deep neural network model called Hybrid Value-Aware Transformer (HVAT) model. The HVAT models can learn jointly from longitudinal and non-longitudinal clinical data, and in particular can leverage the time-varying numerical values associated with the clinical codes or concepts within the longitudinal data for outcome prediction. The key component of the XAI method is the definitions of two derived variables, the temporal mean and the temporal slope, which are defined for the clinical concepts with associated time-varying numerical values. The two variables represent the overall level and the rate of change over time, respectively, in the trajectory formed by the values associated with the clinical concept. Two operations on the original values are designed for changing the values of the two derived variables separately. The effects of the two variables on the outcome risks learned by the HVAT model are calculated in terms of impact scores and impacts. Interpretations of the impact scores and impacts as being similar to those of odds ratios are also provided. We applied the XAI method to the study of cardiorespiratory fitness (CRF) as a risk factor of Alzheimers disease and related dementias (ADRD). Using a retrospective case-control study design, we found that each one-unit increase in the overall CRF level is associated with a 5% reduction in ADRD risk, while each one-unit increase in the changing rate of CRF over time is associated with a 1% reduction. A closer investigation revealed that the association between the changing rate of CRF level and the ADRD risk is nonlinear, or more specifically, approximately piecewise linear along the axis of the changing rate on two pieces: the piece of negative changing rates and the piece of positive changing rates.

12
Enhanced Adverse-Event Detection and Drug-Event Relation Extraction from Clinical Notes

Alharbi, O.; Wu, C. H.; Chen, C.; Shanker, V.

2026-05-08 health informatics 10.64898/2026.05.06.26352616 medRxiv
Top 0.1%
34.7%
Show abstract

Adverse drug events (ADEs) are a significant source of preventable patient harm, yet many ADE signals remain buried in free-text clinical notes. Clinical notes often describe adverse events (AEs) in relation to drugs in two ways: whether a drug causes the AE (the AE is an ADE) or a drug is given to treat an AE (it is considered the Reason for drug treatment). In the N2C2 2018 benchmark, ADEs and Reasons are annotated as separate entity types, despite often being similar in both wording and clinical meaning. This shared similarity makes them difficult to distinguish during entity extraction, leading to errors in relation classification. Therefore, we propose a two-stage framework that first detects AEs as a unified event category and then classifies drug-event pairs into Drug-ADE, Drug-Reason, or No-Relation. In the end-to-end evaluation on the N2C2 2018 benchmark, our system achieves F1 scores of 0.93 for Drug-ADE and 0.94 for Drug-Reason, improving over previously reported end-to-end benchmarks of 0.48 for Drug-ADE and 0.59 for Drug-Reason. Overall, these results support a more precise task formulation in which AEs are detected broadly first, and the ADE vs Reason distinction is resolved at the relation layer. Furthermore, they motivate the development of AE-focused datasets annotated independently of drug linkage to enable more reliable end-to-end pharmacovigilance systems.

13
Enhancing automated indexing of publication types and study designs in biomedical literature using full-text features

Menke, J. D.; Ming, S.; Radhakrishna, S.; Kilicoglu, H.; Smalheiser, N. R.

2025-04-25 health informatics 10.1101/2025.04.23.25326300 medRxiv
Top 0.1%
34.6%
Show abstract

ObjectiveSearching for biomedical articles by publication type or study design is essential for tasks like evidence synthesis. Prior work has relied solely on PubMed information or addressed a limited set of types (e.g., randomized controlled trials). In this study, we build on previous work by lever-aging full-text features, enriched text representations, and advanced optimization techniques for comprehensive indexing. MethodsUsing a dataset of PubMed articles published between 1987 and 2023 with human-annotated indexing terms, we fine-tuned BERT-based encoders (PubMedBERT, BioLinkBERT, SPECTER, SPECTER2-Base, SPECTER2-Clf) to investigate whether text representations based on different pre-training objectives could benefit the task. We incorporated textual and verbalized metadata features, full-text extraction (rule-based, extractive, and abstractive summarization), and additional topical information about the articles. To mitigate potential label noise and improve calibration, we used asymmetric loss and label smoothing. We also explored contrastive learning approaches (SimCSE, ADNCE, HeroCon, WeighCon). Models were evaluated using precision, recall, F1 score (both micro- and macro-), and area under ROC curve (AUC). ResultsFine-tuning SPECTER2-Base with asymmetric loss, label smoothing and contrastive learning (ADNCE and HeroCon) improved performance significantly over the previous best model (micro-F1: 0.658 [-&gt;] 0.670 [+1.8%]; macro-F1: 0.643 [-&gt;] 0.677 [+5.3%]; p < 0.001). Asymmetric loss and using SPECTER2-Base instead of PubMedBERT contributed most to this gain, while contrastive learning provided more moderate gains. Full-text features boosted performance by 2.4% (micro-F1) and 0.8% (macro-F1) over the baseline (micro-F1: 0.656 [-&gt;] 0.672; macro-F1: 0.595 [-&gt;] 0.600; p < 0.001). ConclusionFull-text features, citation-aware encoders, and fine-tuning optimizations significantly improve publication type and study design indexing. Future work should refine label accuracy, better distill relevant full-text information, and expand label sets to meet needs of the research community. Data, code, and models are available at https://github.com/ScienceNLP-Lab/MultiTagger-v2. HighlightsO_LIWe trained and validated Transformer-based models for automatic indexing of publication types and study designs in biomedical articles, using a dataset with 61 labels derived primarily from expert-assigned PubMed indexing terms. C_LIO_LIWe investigated whether enriched article representations, advanced optimization techniques, and fine-grained labels could enhance model performance. C_LIO_LIThe largest performance improvement came from using citation-aware article representations and asymmetric loss. C_LIO_LIModels trained using full-text features outperformed models trained using PubMed-only features, demonstrating the utility of full-text content for this task. C_LI Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=146 SRC="FIGDIR/small/25326300v3_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@1716286org.highwire.dtl.DTLVardef@fb65ccorg.highwire.dtl.DTLVardef@d861b7org.highwire.dtl.DTLVardef@1f7670d_HPS_FORMAT_FIGEXP M_FIG C_FIG

14
Fine-tuning large language models for effective nutrition support in residential aged care: a domain expertise approach

Alkhalaf, M.; Deng, C.; Shen, J.; Chang, H.-C.; Yu, P.

2024-07-21 health informatics 10.1101/2024.07.21.24310775 medRxiv
Top 0.1%
34.3%
Show abstract

PurposeMalnutrition is a serious health concern, particularly among the older people living in residential aged care facilities. An automated and efficient method is required to identify the individuals afflicted with malnutrition in this setting. The recent advancements in transformer-based large language models (LLMs) equipped with sophisticated context-aware embeddings, such as RoBERTa, have significantly improved machine learning performance, particularly in predictive modelling. Enhancing the embeddings of these models on domain-specific corpora, such as clinical notes, is essential for elevating their performance in clinical tasks. Therefore, our study introduces a novel approach that trains a foundational RoBERTa model on nursing progress notes to develop a RAC domain-specific LLM. The model is further fine-tuned on nursing progress notes to enhance malnutrition identification and prediction in residential aged care setting. MethodsWe develop our domain-specific model by training the RoBERTa LLM on 500,000 nursing progress notes from residential aged care electronic health records (EHRs). The models embeddings were used for two downstream tasks: malnutrition note identification and malnutrition prediction. Its performance was compared against baseline RoBERTa and BioClinicalBERT. Furthermore, we truncated long sequence text to fit into RoBERTas 512-token sequence length limitation, enabling our model to handle sequences up to1536 tokens. ResultsUtilizing 5-fold cross-validation for both tasks, our RAC domain-specific LLM demonstrated significantly better performance over other models. In malnutrition note identification, it achieved a slightly higher F1-score of 0.966 compared to other LLMs. In prediction, it achieved significantly higher F1-score of 0.655. We enhanced our models predictive capability by integrating the risk factors extracted from each clients notes, creating a combined data layer of structured risk factors and free-text notes. This integration improved the prediction performance, evidenced by an increased F1-score of 0.687. ConclusionOur findings suggest that further fine-tuning a large language model on a domain-specific clinical corpus can improve the foundational models performance in clinical tasks. This specialized adaptation significantly improves our domain-specific models performance in tasks such as malnutrition risk identification and malnutrition prediction, making it useful for identifying and predicting malnutrition among older people living in residential aged care or long-term care facilities.

15
Utilizing LLMs to Evaluate the Argument Quality of Triples in SemMedDB for Enhanced Understanding of Disease Mechanisms

Wang, S.; Zhang, Y.; Du, J.

2024-03-22 health informatics 10.1101/2024.03.20.24304652 medRxiv
Top 0.1%
33.9%
Show abstract

Current semantic extraction tools have limited performance in identifying causal relations, neglecting variations in argument quality, especially persuasive strength across different sentences. The present study proposes a five-element based (evidence cogency, concept, relation stance, claim-context relevance, conditional information) causal knowledge mining framework and automatically implements it using large language models (LLMs) to improve the understanding of disease causal mechanisms. As a result, regarding cogency evaluation, the accuracy (0.84) of the fine-tuned Llama2-7b largely exceeds the accuracy of GPT-3.5 turbo with few-shot. Regarding causal extraction, by combining PubTator and ChatGLM, the entity first-relation later extraction (recall, 0.85) outperforms the relation first-entity later means (recall, 0.76), performing great in three outer validation sets (a gestational diabetes-relevant dataset and two general biomedical datasets), aligning entities for further causal graph construction. LLMs-enabled scientific causality mining is promising in delineating the causal argument structure and understanding the underlying mechanisms of a given exposure-outcome pair.

16
CODE - XAI: Construing and Deciphering Treatment Effects via Explainable AI using Real-world Data.

Lu, M.; Covert, I.; White, N. J.; Lee, S.-I.

2024-10-11 health informatics 10.1101/2024.09.04.24312866 medRxiv
Top 0.1%
33.7%
Show abstract

Clinicians rely on evidence from randomized controlled trials (RCTs) to decide on medical treatments for patients. However, RCTs often lack the granularity needed to inform decisions for individual patients or specific clinical scenarios. Recent advances in machine learning, particularly conditional average treatment effect (CATE) modeling, offer a promising approach for estimating patient-specific treatment effects. Yet, their adoption in clinical practice remains limited because these models often function as "black boxes", making it difficult to understand which features drive treatment effect heterogeneity among individuals. To overcome these barriers, we introduce LIFT-XAI, a robust and principled framework that interprets ensemble CATE models with Shapley values, thereby accurately identifying unique features and patient subgroups driving the treatment effects. We validate LIFT-XAI on real-world clinical data from four RCTs comprising over 42,000 patients. We demonstrate that LIFT-XAI uncovered fasting glucose as the primary factor explaining conflicting outcomes between two major blood pressure trials (SPRINT and ACCORD) and discovered age as a critical treatment modifier in a new clinical setting-a finding later validated by a subsequent RCT. By enabling such powerful cross-cohort analysis and the discovery of novel patient subgroups, LIFT-XAI advances precision medicine across diverse settings, from chronic disease management to acute care.

17
Developing a Health Score and Predicting Disease Risks Using DKAbio-Clusters

Cheng, K.-F.; Yang, Y.-H.; Su, C.-H.; Tsai, M.-C.

2024-06-18 health informatics 10.1101/2024.06.16.24308995 medRxiv
Top 0.1%
33.7%
Show abstract

In our research, accurately estimating the morbidity of individuals with specific conditions, plays a pivotal role in enhancing healthcare delivery systems. Introducing DKABio-clusters, we delve into their distinct characteristics, showcasing their profound implications for healthcare management. A primary focus of DKABio-clusters lies in developing a unique health assessment tool, termed DKABio-HS, alongside predictive risk analysis. DKABio-HS facilitates the computation of a comprehensive "disease-related" score, condensing an individuals health status into a singular numerical value. Our investigation reveals the remarkable consistency of this health score, with minimal variations observed between training and validation datasets (mean absolute percentage errors within 0 to 10 years remaining below 0.1%, with all mean absolute percentage errors ranging between 1.2-1.6%). A higher health score denotes better health or reduced disease risk, diminishing with age or the presence of multiple diseases. Utilizing this health score, we establish a classification framework termed the "disease map," enabling precise differentiation of individuals across various health states. Through this framework, individuals without diseases can be categorized as either healthy or sub-healthy, facilitating tailored health management strategies for preventive interventions. Our analysis indicates that individuals classified as sub-healthy exhibit significantly elevated disease risks compared to those deemed healthy (Female (male) 5-year risks of developing at least one disease are 29% vs. 15% (29% vs. 16.5%)). Furthermore, leveraging a carefully selected set of health variables, we can delineate the distribution of DKABio-clusters and concurrently predict the 10-year risks associated with 15 diseases/conditions. Validating the predictive capabilities of our model, we compare predicted risks with true risks derived from extensive datasets, demonstrating non-statistically significant differences in the majority of cases. All analyses are grounded in data sourced from the National Health Insurance Research Database (more than 2 million participants) released by the National Health Research Institute, Taiwan and the Mei Jau Health Management Institution database (more than 0.75 million participants), spanning the years 2000 to 2016 in Taiwan.

18
Large Language Models Struggle to Encode Medical Concepts - A Multilingual Benchmarking and Comparative Analysis

Rouhizadeh, H.; Yazdani, A.; Zhang, B.; Vicente Alvarez, D.; Hueser, M.; Vanobberghen, A.; Yang, R.; Li, I.; Walter, A.; Teodoro, D.

2025-01-15 health informatics 10.1101/2025.01.15.25320579 medRxiv
Top 0.1%
33.6%
Show abstract

Interoperability in health information systems is crucial for accurate data exchange across environments such as electronic health records, clinical notes, and medical research. The main challenge arises from the wide variation in biomedical concepts, their representation across different systems and languages, and the limited context, complicating data integration and standardization. Inspired by recent advances in large language models (LLMs), this study explores their potential role as biomedical knowledge engineers to (semi-)automate multilingual biomedical concept normalization, a key task for semantic interoperability of medical concepts. We developed a novel multilingual dataset comprising 59104 unique terms mapped to 27280 distinct biomedical concepts, designed to assess language model performance across this task within five European languages: English, French, German, Spanish, and Turkish. We then proposed a multi-stage pipeline based on a retrieve-then-rerank approach using sparse and dense retrievers, rerankers, and fusion approaches, leveraging discriminative and generative LLMs, with a predefined primary knowledge organization system. Our experiments show that the best discriminative model, e5, achieves an accuracy of 71%, surpassing the best generative model, Mistral, by 2% (p-value < 0.001). For semi-automated workflows, e5 maintained superior performance with 82% recall@10 versus Mistrals 78%. Our findings demonstrate a pathway to how LLM-based approaches can advance the normalization of multilingual biomedical terms as well as the limitations of LLMs in encoding biomedical concepts.

19
Large Language Models for Social Determinants of Health Information Extraction from Clinical Notes - A Generalizable Approach across Institutions

Keloth, V. K.; Selek, S.; Chen, Q.; Gilman, C.; Fu, S.; Dang, Y.; Chen, X.; Hu, X.; Zhou, Y.; He, H.; Fan, J. W.; Wang, K.; Brandt, C.; Tao, C.; Liu, H.; Xu, H.

2024-05-22 health informatics 10.1101/2024.05.21.24307726 medRxiv
Top 0.1%
33.4%
Show abstract

The consistent and persuasive evidence illustrating the influence of social determinants on health has prompted a growing realization throughout the health care sector that enhancing health and health equity will likely depend, at least to some extent, on addressing detrimental social determinants. However, detailed social determinants of health (SDoH) information is often buried within clinical narrative text in electronic health records (EHRs), necessitating natural language processing (NLP) methods to automatically extract these details. Most current NLP efforts for SDoH extraction have been limited, investigating on limited types of SDoH elements, deriving data from a single institution, focusing on specific patient cohorts or note types, with reduced focus on generalizability. This study aims to address these issues by creating cross-institutional corpora spanning different note types and healthcare systems, and developing and evaluating the generalizability of classification models, including novel large language models (LLMs), for detecting SDoH factors from diverse types of notes from four institutions: Harris County Psychiatric Center, University of Texas Physician Practice, Beth Israel Deaconess Medical Center, and Mayo Clinic. Four corpora of deidentified clinical notes were annotated with 21 SDoH factors at two levels: level 1 with SDoH factor types only and level 2 with SDoH factors along with associated values. Three traditional classification algorithms (XGBoost, TextCNN, Sentence BERT) and an instruction tuned LLM-based approach (LLaMA) were developed to identify multiple SDoH factors. Substantial variation was noted in SDoH documentation practices and label distributions based on patient cohorts, note types, and hospitals. The LLM achieved top performance with micro-averaged F1 scores over 0.9 on level 1 annotated corpora and an F1 over 0.84 on level 2 annotated corpora. While models performed well when trained and tested on individual datasets, cross-dataset generalization highlighted remaining obstacles. To foster collaboration, access to partial annotated corpora and models trained by merging all annotated datasets will be made available on the PhysioNet repository.

20
Advancements in Multilingual Biomedical Natural Language Processing: exploring Large Language Models for Named Entity Recognition and Linking

Mazzucato, S.; Seinen, T. M.; Moccia, S.; Micera, S.; Bandini, A.; van Mulligen, E. M.

2026-01-23 health informatics 10.64898/2026.01.22.26344605 medRxiv
Top 0.1%
33.3%
Show abstract

ObjectiveNamed Entity Recognition (NER) and Biomedical Entity Linking (BEL) are essential for transforming unstructured Electronic Health Records (EHRs) into structured information. However, tools for these tasks are limited in non-English biomedical texts such as Dutch and Italian. This study investigates the use of prompt-based learning with Large Language Models (LLMs) to perform multilingual NER and BEL using minimal domainspecific data, while addressing annotation preservation during corpus translation. MethodsAn English-annotated corpus from the ShARe/CLEF dataset was translated into Dutch and Italian using a strategy that embeds annotations directly into the text prior to translation and retrieves them afterwards. GPT-4o was applied in zero-shot and few-shot settings to extract biomedical entities, which were then mapped to Unified Medical Language System Concept Unique Identifiers using contextual word embeddings. Performance was evaluated with precision, recall, and F1-score, and compared with goldstandard clinician annotations. ResultsThe multilingual NER pipeline achieved strong performance, with an overall F1-score of 0.98 across languages. BEL experiments showed reliable entity normalization, with an overall accuracy of 0.91 and a mean reciprocal rank of 0.95. The combined performance of the NER and BEL achieved 0.90 supporting the utility of LLMs in standardizing biomedical concepts across languages. ConclusionPrompt-based LLMs can effectively perform NER and BEL in languages with less annotated resources, even with limited annotated training data. The proposed annotation-preserving translation method, combined with generative and discriminative LLM capabilities, provides a scalable approach to multilingual clinical information extraction. These findings highlight the potential for broader adoption of LLM-based natural language processing systems to support multilingual healthcare data harmonization. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=133 SRC="FIGDIR/small/26344605v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@100c385org.highwire.dtl.DTLVardef@12467d8org.highwire.dtl.DTLVardef@11d9ca5org.highwire.dtl.DTLVardef@1173e3d_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIThis study shows the feasibility of using prompt-based learning with large language models (LLMs) to perform multilingual named entity recognition (NER) and biomedical entity linking (BEL) in Dutch and Italian, two languages with less annotated resources. C_LIO_LIAn annotation-preserving translation strategy was proposed to adapt the ShARe/CLEF eHealth corpus, enabling consistent evaluation across English, Dutch, and Italian without loss of gold-standard annotations. C_LIO_LIThe multilingual NER pipeline achieved strong overall performance (F1-score: 0.89), while BEL experiments showed reliable entity normalization (F1-score: 0.64, MRR: 0.68) to standardized clinical concepts. C_LIO_LIThe approach highlights the potential of generative and discriminative LLM capabilities for scalable multilingual clinical information extraction, supporting broader European initiatives for cross-lingual health data harmonization. C_LI