Determining gene specificity from multivariate single-cell RNA sequencing data
Swarna, N. P.; Booeshaghi, A. S.; Rebboah, E.; Gordon, M. G.; Kathail, P.; Li, T.; Alvarez, M.; Ye, C. J.; Wold, B. J.; Mortazavi, A.; Pachter, L.
Show abstract
An important application of single-cell genomics experiments is to identify genes specific to biological categories or experimental conditions. Although numerous approaches have been proposed to identify such genes, we consider an axiomatic approach based on defining properties that a specificity measure should have. This leads us to develop ember (Entropy Metrics for Biological ExploRation), which we show is the only method satisfying four key desired properties for a specificity measure. Applying ember to eight tissues from eight founder mouse strains, we find that gene specificity is often unintuitive: canonical markers can be supplanted, housekeeping genes are context-dependent, and mouse strain can drive unexpected cell type switching. Unsupervised learning on entropy metrics uncovers shared genes specialized to male gonads and kidney, as well as genes specific to non-consecutive developmental stages in the kidney. To facilitate further exploration of gene specificity in mice, we have also developed a comprehensive specificity database, along with a web interface, API and MCP server. Extending ember to a human PBMC dataset collected from 255 diverse individuals, we find that variation in PBMCs is largely localized to classical monocytes. We also find genes with unique specificity by sex, age and ancestral background. Together, these applications establish ember as a powerful tool and provide a roadmap for elucidating the impact of human genetic variation using the murine model.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A NMF-based approach to discover overlooked differentially expressed gene regions from single-cell RNA-seq data 93%
- ConsHMM Atlas: conservation state annotations for major genomes and human genetic variation 93%
- Sequence-based chromatin activity modeling and regulatory impact prediction of genetic variants in farmed animals using deep learning 92%
Similar papers in this journal
- Identifying gene function and module connections by the integration of multi-species expression compendia 94%
- PRAM: a novel pooling approach for discovering intergenic transcripts from large-scale RNA sequencing experiments 94%
- Alignment of single-cell RNA-seq samples without over-correction using kernel density matching 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.