Multi-agent AI enables evidence-based cell annotation in single-cell transcriptomics
Ahuja, G.; Antill, A.; Su, Y.; Dall'Olio, G. M.; Basnayake, S.; Karlsson, G.; Dhapola, P.
Show abstract
Cell type annotation remains a critical bottleneck, with current methods often inaccurate and requiring extensive manual validation, particularly in disease contexts. While large language models (LLMs) show promise, they can be unreliable due to hallucinations. We developed CyteType, a multi-agent framework that generates competing hypotheses grounded in full expression data and study context, validates against external databases, and iteratively self-evaluates. Comprehensive benchmarking demonstrates that CyteType substantially outperforms reference-based and LLM-based methods, with self-generated confidence scores reliably identifying trustworthy annotations. CyteType transforms cell type annotation from label assignment into evidence-grounded biological discovery. Python (AnnData compatible): https://github.com/NygenAnalytics/CyteType R (Seurat compatible): https://github.com/NygenAnalytics/CyteTypeR
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 94%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 94%
Similar papers in this journal
- SELINA: Single-cell Assignment using Multiple-Adversarial Domain Adaptation Network with Large-scale References 94%
- Clustering-independent estimation of cell abundances in bulk tissues using single-cell RNA-seq data 94%
- Quantifying tumor specificity using Bayesian probabilistic modeling for drug target discovery and prioritization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.