Harnessing Data-Intelligence-Intensive Multi-Agent System for Life Science Research
Liu, Y.; Shen, R.; Zhou, L.; Xiao, Q.; Yuan, J.; Li, Y.
Show abstract
Advancements in high-throughput sequencing technologies and artificial intelligence offer unprecedented opportunities for groundbreaking discoveries in bioinformatics research. However, the challenges of exponential growth of omics data and the rapid development of artificial intelligence technologies require automated big biological data analysis capability and interdisciplinary knowledge-driven scientific insight. Here we propose a data-intelligence-intensive bioinformatics copilot (Bio-Copilot) system that synergizes AI capabilities with human expertise to facilitate hypothesis-free exploratory research and inspire novel scientific insights in large-scale omics studies. Bio-Copilot forms high-quality intensive intelligence through close collaboration between multiple agents, driven by large language models (LLMs), and human experts. To augment the capabilities of Bio-Copilot, this study devises an agent group management strategy, an effective human-agent interaction mechanism, a shared interdisciplinary knowledge database, and continuous learning strategies for the agents. We comprehensively compare Bio-Copilot against GPT-4o and several leading AI agents across diverse bioinformatics tasks, using a broad range of evaluation metrics. Bio-Copilot achieves the overall state-of-the-art performance across all tasks, while showcases exceptional task completeness. Furthermore, in the application of constructing a large-scale human lung cell atlas, Bio-Copilot not only reproduces the intricate data integration process detailed in a seminal study but also introduces a hierarchical annotation strategy to capture the continuous nature of cellular states and uncovers the characteristics of rare cell types, highlighting its potential to unravel hidden complexities in biological systems. Beyond the technical achievements, this study also underscores the profound implications of integrating AI capabilities with expert knowledge in accelerating impactful biological discoveries and exploring uncharted territories in life sciences.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Robust and Scalable Graph Neural Network for Accurate Single Cell Classification 95%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 94%
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 95%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- TooManyCellsInteractive: a visualization tool for dynamic exploration of single-cell data 94%
Similar papers in this journal
- Discovering nuclear localization signal universe through a novel deep learning model with interpretable attention units 95%
- Federated Learning for multi-omics: a performance evaluation in Parkinson's disease 95%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.