OriGene: A Self-Evolving Virtual Disease Biologist Automating Therapeutic Target Discovery
Zhang, Z.; Qiu, Z.; Wu, Y.; Li, S.; Wang, D.; Zhou, Z.; An, D.; Chen, Y.; Li, Y.; Wang, Y.; Ou, C.; Wang, Z.; Chen, J. X.; Zhang, B.; Hu, Y.; Zhang, W.; Wei, Z.; Ma, R.; Liu, Q.; Dong, B.; He, Y.; Feng, Q.; Bai, L.; Gao, Q.; Sun, S.; Zheng, S.
Show abstract
Here, we present OriGene, a self-evolving multi-agent system that functions as a virtual disease biologist, systematically identifying original and mechanistically grounded therapeutic targets at scale. OriGenes architecture integrates over 600 specialized tools through a Model Context Protocol (MCP), enabling it to reason across diverse data modalities including genomics, protein networks, pharmacology, clinical records and literature evidence, to generate and prioritize target discovery hypotheses. We implemented a strategy combining a knowledge graph-based Tool RAG with an advanced agent selection mechanism to enable dynamic, context-aware tool deployment. Through a self-evolving framework, OriGene continuously integrates human and experimental feedback to iteratively refine its core thinking templates, tool composition, and analytical protocols, thereby enhancing both accuracy and adaptability over time. To comprehensively evaluate its performance, we established TRQA, an original benchmark comprising over 1,900 expert-level question-answer pairs spanning a wide range of diseases and target classes. OriGene consistently outperforms human experts, leading research agents, and state-of-the-art large language models in accuracy, recall, and robustness, particularly under conditions of data sparsity or noise. Critically, OriGene nominated previously underexplored therapeutic targets for liver (GPR160) and colorectal cancer (ARG2), which demonstrated significant anti-tumor activity in patient-derived organoid and tumor fragment models mirroring human clinical exposures. These findings demonstrate OriGenes potential as a scalable and adaptive platform for AI-driven discovery of mechanistically grounded therapeutic targets, offering a new paradigm to accelerate drug development.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 94%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 94%
- Community assessment of cancer drug combination screens identifies strategies for synergy prediction 94%
Similar papers in this journal
Similar papers in this journal
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 94%
- KG-COVID-19: a framework to produce customized knowledge graphs for COVID-19 response 93%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.