Benchmarking Cell Type Annotation by Large Language Models with AnnDictionary
Crowley, G.; Tabula Sapiens Consortium, ; Quake, S. R.
Show abstract
We developed an open-source package called AnnDictionary (https://github.com/ggit12/anndictionary/) to facilitate the parallel, independent analysis of multiple anndata. AnnDictionary is built on top of LangChain and AnnData and supports all common large language model (LLM) providers. AnnDictionary only requires 1 line of code to configure or switch the LLM backend and it contains numerous multithreading optimizations to support the analysis of many anndata and large anndata. We used AnnDictionary to perform the first benchmarking study of all major LLMs at de novo cell-type annotation. LLMs varied greatly in absolute agreement with manual annotation based on model size. Inter-LLM agreement also varied with model size. We find that LLM annotation of most major cell types to be more than 80-90% accurate, and will maintain a leaderboard of LLM cell type annotation at https://singlecellgpt.com/celltype-annotation-leaderboard. Furthermore, we benchmarked these LLMs at functional annotation of gene sets, and found that Claude 3.5 Sonnet recovers close matches of functional gene set annotations in over 80% of test sets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A comparison of marker gene selection methods for single-cell RNA sequencing data 96%
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 96%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
Similar papers in this journal
Similar papers in this journal
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 95%
- CoSTA: Unsupervised Convolutional Neural Network Learning for Spatial Transcriptomics Analysis 94%
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.