Back

CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis

Xiao, Y.; Liu, J.; Zheng, Y.; Jiao, S.; Hao, J.; Xie, X.; Li, M.; Wang, R.; Ni, F.; Li, Y.; Wang, Z.; Shang, X.; Bao, Z.; Yang, C.; Peng, J.

2025-05-05 bioinformatics
10.1101/2024.05.13.593861 bioRxiv
Show abstract

Single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) data analysis are pivotal for advancing biological research, enabling precise characterization of cellular heterogeneity. However, existing analysis approaches require extensive manual programming and tool manipulation, posing significant challenges for researchers. To address this, we introduce CellAgent, an autonomous, LLM-driven approach that performs end-to-end scRNA-seq and spatial transcriptomics data analysis through natural language interactions. CellAgent employs a multi-agent hierarchical decision-making framework, simulating a "deep-thinking" workflow to ensure that each analytical step remains consistent with the overall task objective. To further enhance its capabilities, we developed sc-Omni, a high-performance, expert-curated toolkit that consolidates essential tools for scRNA-seq and spatial transcriptomics analysis. Additionally, we introduce a self-reflective optimization mechanism, enabling automated, iterative refinement of results through specialized evaluation methods, effectively replacing traditional manual assessments. Benchmarking against human experts demonstrates that CellAgent achieves approximately 60% improvement in efficiency across multiple downstream applications. In terms of accuracy, it maintains performance comparable to existing approaches while preserving natural language interactions. By translating natural language interactions into optimized analytical workflows, CellAgent establishes a scalable paradigm for LLM-driven scientific discovery, bridging the gap between experimental biologists and complex data analytics. This framework minimizes reliance on manual coding and exhaustive deliberation, ushering in the era of the "AI Agent for Science."

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.