Interpreting Omics Data Analysis with Large Language Models for Disease Target and Drug Discovery
XU, Z.; Chen, W.; Ren, W.; Xu, T.; Amaechin, S.; Khan, R.; Chen, Y.; Province, M.; Payne, P.; Li, F.
Show abstract
In biomedical scientific discovery, synthesizing prior knowledge from the literature is an essential component of interpreting numerical omics data analyses for disease target identification and drug discovery. Large language models (LLMs) alone can rapidly retrieve disease mechanisms from biomedical text, but text-only outputs are general and unreliable for target and drug prioritization without cohort-specific quantitative evidence. Herein, we propose a provenance-aware Text-to-Target framework that couples schema-constrained multi-model LLM retrieval with numeric omics data analysis. The key design is a modality-aware fusion step: candidates are partitioned into overlap-supported anchors, retrieval-only hidden hubs, and network-emergent novelty nodes, then propagated into staged hypothesis and strategy generation under topology constraints. We evaluate the model in Alzheimers disease (AD) and pancreatic ductal adenocarcinoma (PDAC). In PDAC, the workflow produced a balanced 75-gene candidate universe and a 23-strategy portfolio, with significant DepMap support at both target level and strategy level. In AD, stricter candidate controls yielded a compact 34-gene universe and 14 strategies; under an expanded CRISPRbrain registry, both target-level axes were significant, with strong strategy-level enrichment. Across both diseases, final strategies preserved full provenance closure to the candidate pool, enabling end-to-end auditability from retrieval artifacts to validation outputs. These results support a transferable discovery architecture in which omics evidence constrains biological activity, LLM retrieval expands mechanistic search space, and network-aware fusion preserves interpretability. The framework provides a reproducible basis for dual-disease target prioritization and motivates continuous literature-mechanism concordance with agentic evidence-refresh loops.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Addressing biases in gene-set enrichment analysis: a case study of Alzheimer's Disease 92%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 92%
- Impact of between-tissue differences on pan-cancer predictions of drug sensitivity 91%
Similar papers in this journal
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 92%
- SpaPheno: Linking Spatial Transcriptomics to Clinical Phenotypes with Interpretable Machine Learning 92%
- Personalized Cancer Therapy Prioritization Based on Driver Alteration Co-occurrence Patterns 91%
Similar papers in this journal
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 93%
- DrDimont: Explainable drug response predictionfrom differential analysis of multi-omics networks 92%
- CLEP: A Hybrid Data- and Knowledge- Driven Framework for Generating Patient Representations 92%
Similar papers in this journal
- Single-Cell Trajectory Inference for Detecting Transient Events in Biological Processes 92%
- DeepSpaceDB: a spatial transcriptomics atlas for interactive in-depth analysis of tissues and tissue microenvironments 91%
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.