Literature-derived, context-aware gene regulatory networks improve biological predictions and mathematical modeling
Tsutsui, M.; Arakane, K.; Okada, M.
Show abstract
MotivationComplex gene regulatory networks (GRNs) underlie most disease processes, and understanding disease-specific network structures and dynamics is crucial for developing effective treatments. Yet, literature-based analyses of GRNs often treat gene regulations as context-independent interactions, overlooking how their biological relevance can differ depending on the disease type, cell lineage, or experimental condition. ResultsIn an attempt to improve on existing methods for leveraging knowledge present in the scientific literature, we developed a framework to assign quantitative, context-dependent weights to gene regulations extracted from literature. We demonstrate that the context-specific GRNs reconstructed with our method can effectively capture disease biology, showing strong correlation with transcriptomics across a wide range of diseases. Furthermore, we show that utilizing contextual information improves accuracy in drug-target prediction tasks. Finally, we showcase the utility of the contextualized GRNs through the automated construction of an ordinary differential equation model of a breast cancer-specific signaling network. The large language model-based framework allows the integration of literature- and experimentally derived information and streamlines the process of assembling a biologically relevant and functional mathematical model. Our findings indicate the importance of considering the context when making biological predictions, and we demonstrate the use of natural language processing tools to effectively mine associations between gene regulations and biological contexts. Availability and implementationAll reproducibility code is available at https://github.com/okadalabipr/context-dependent-GRNs, along with the automated mathematical model construction package at https://github.com/okadalabipr/BioMathForge. The dataset used in this study is available at Zenodo, DOI: 10.5281/zenodo.16416117.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 95%
- Mining hidden knowledge: Embedding models of cause-effect relationships curated from the biomedical literature 95%
- Beyond synthetic lethality in large-scale metabolic and regulatory network models via genetic minimal intervention sets 95%
Similar papers in this journal
- RIDDEN: Data-driven inference of receptor activity from transcriptomic data 96%
- Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery 96%
- Joint representation of molecular networks from multiple species improves gene classification 95%
Similar papers in this journal
- Deciphering the Signaling Network Landscape of Breast Cancer Improves Drug Sensitivity Prediction 95%
- Interpretable Machine Learning for Perturbation Biology 94%
- TRIAGE: A web-based iterative analysis platform integrating pathway and network approaches optimizes hit selection from high- throughput assays. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.