CZ CELLxGENE Discover: A single-cell data platform for scalable exploration, analysis and modeling of aggregated data
CZI Single-Cell Biology, ; Abdulla, S.; Aevermann, B.; Assis, P.; Badajoz, S.; Bell, S. M.; Bezzi, E.; Cakir, B.; Chaffer, J.; Chambers, S.; Cherry, J. M.; Chi, T.; Chien, J.; Dorman, L.; Gloria, N.; Garcia-Nieto, P.; Hastie, M.; Hegeman, D.; Hilton, J.; Huang, T.; Infeld, A.; Istrate, A.-M.; Jelic, I.; Katsuya, K.; Kim, Y. J.; Liang, K.; Lin, M.; Lombardo, M.; Marshall, B.; Martin, B.; McDade, F.; Megill, C.; Patel, N.; Predeus, A.; Raymor, B.; Robatmili, B.; Rogers, D.; Rutherford, E.; Sadgat, D.; Shin, A.; Small, C.; Smith, T.; Sridharan, P.; Tarashansky, A.; Tavares, N.; Thomas, H.; Tolop
Show abstract
Hundreds of millions of single cells have been analyzed to date using high throughput transcriptomic methods, thanks to technological advances driving the increasingly rapid generation of single-cell data. This provides an exciting opportunity for unlocking new insights into health and disease, made possible by meta-analysis that span diverse datasets building on recent advances in large language models and other machine learning approaches. Despite the promise of these and emerging analytical tools for analyzing large amounts of data, a major challenge remains the sheer number of datasets and inconsistent format, data models and accessibility. Many datasets are available via unique portals platforms that often lack interoperability. Here, we present CZ CellxGene Discover (cellxgene.cziscience.com), a data platform that provides curated and interoperable data. This single-cell data resource, available via a free-to-use online data portal, hosts a growing corpus of community contributed data that spans more than 50 million unique cells. Curated, standardized, and associated with consistent cell-level metadata, this collection of interoperable single-cell transcriptomic data is the largest of its kind. A suite of tools and features enables accessibility and reusability of the data via both computational and visual interfaces to allow researchers to rapidly explore individual datasets and perform cross-corpus analysis. This functionality is enabling meta-analyses of tens of millions of cells across studies and tissues and providing global views of human cells at the resolution of single cells.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CellContrast: Reconstructing Spatial Relationships in Single-Cell RNA Sequencing Data via Deep Contrastive Learning 94%
- EmptyNN: A neural network based on positive-unlabeled learning to remove cell-free droplets and recover lost cells in single-cell RNA sequencing data 94%
- MANGEM - a web app for Multimodal Analysis of Neuronal Gene expression, Electrophysiology and Morphology 94%
Similar papers in this journal
- The 4D Nucleome Data Portal: a resource for searching and visualizing curated nucleomics data 96%
- Scarf: A toolkit for memory efficient analysis of large-scale single-cell genomics data 96%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 96%
Similar papers in this journal
- Flexible comparison of batch correction methods for single-cell RNA-seq using BatchBench 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 95%
- DeepSpaceDB: a spatial transcriptomics atlas for interactive in-depth analysis of tissues and tissue microenvironments 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.