Accelerating scRNA-seq analysis: automated cell type annotation using representation learning and vector search
Williams, S. R.; Grab, F.; Kamath, G. M.; Ordabayev, Y.; Mellen, J.; Roelli, P.; Cibulskis, K.; Lehnert, E.; Xie, F.; Covarrubias, M.; Rahman, N.-T.; Tickle, T.; Erhan, E.; Malfroy-Camine, N.; Lydon, K.; Babadi, M.; Delaney, N. F.
Show abstract
Cell type annotation in single-cell RNA sequencing (scRNA-seq) experiments is the fundamental step of assigning cell types to individual cells or clusters of cells based on their gene expression profiles. This process is crucial for developing biological insights from scRNA-seq experiments. We present a service that automates cell type annotation for 10x Genomics single-cell gene expression samples, enabling researchers to rapidly and accurately categorize cells within a sample. This service operates on the basis of reverse search: it compares each cells gene expression profile against the Chan Zuckerberg CELL by GENE (CZ CELLxGENE) Census, a comprehensive repository of published scRNA-seq datasets enriched with community-annotated cell types, and yields cell type annotations through summarizing the labels associated with similar cells. The annotation algorithm employed in this service avoids reliance on predefined marker genes or tissue-specific references, providing both fine-grained and coarse annotations. These initial annotations can be further refined by investigators to suit their specific research needs.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Single-cell RNA-seq Training and Analysis Suite using the Galaxy Framework 97%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 97%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.