Back

Large Language Model Consensus Substantially Improves the Cell Type Annotation Accuracy for scRNA-seq Data

Yang, C.; Zhang, X.; Chen, J.

2025-04-17 bioinformatics
10.1101/2025.04.10.647852 bioRxiv
Show abstract

Different large language models (LLMs) have the potential to complement one another. We introduce an iterative multi-LLM consensus framework for annotating single-cell RNA sequencing data. This framework outperforms the best state-of-the-art method by nearly 15% in mean accuracy (77.3% vs 61.3%) across 50 diverse datasets from 26 tissues, encompassing over 8 million cells. By leveraging cross-model deliberation, our framework quantifies uncertainty, identifies ambiguous clusters for expert review, provides transparent reasoning chains, and minimizes the effort and expertise needed for cell type annotation in large-scale studies. Additionally, our framework enables users to seamlessly integrate new LLMs.

Published in Communications Biology (predicted rank #21) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.