Back

Contrastive Learning for Robust Cell Annotation and Representation from Single-Cell Transcriptomics

Andrekson, L.; Mercado, R.

2025-02-28 bioinformatics
10.1101/2024.06.20.599868 bioRxiv
Show abstract

Batch effects are a significant concern in single-cell RNA sequencing (scRNA-Seq) data analysis, where variations in the data can be attributed to factors unrelated to cell types. This can make downstream analysis a challenging task. In this study, we present a novel DL approach using contrastive learning and a carefully designed loss function for learning a generalizable embedding space from scRNA-Seq data. We call this model CELLULAR: CELLUlar contrastive Learning for Annotation and Representation. When benchmarked against multiple established methods for scRNA-Seq integration, CELLULAR outperforms existing methods in learning a generalizable embedding space on multiple datasets. Cell annotation was also explored as a downstream application for the learned embedding space. When compared against multiple well-established methods, CELLULAR demonstrates competitive performance with top cell classification methods in terms of accuracy, balanced accuracy, and F1 score. CELLULAR is also capable of performing novel cell type detection. These findings assess the biological relevance of the models learned embedding space by demonstrating the robustness of its cell representations across various applications. The model has been structured into an opensource Python package, specifically designed to simplify and streamline its usage for bioinformaticians and other scientists interested in cell representation learning.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.