Single-cell transcriptomics for the 99.9% of species without reference genomes
Botvinnik, O. B.; Vemuri, P.; Pierce Ward, N. T.; Logan, P. A.; Nafees, S.; Karanam, L.; Travaglini, K. J.; Ezran, C. S.; Ren, L.; Wang, J.; Huang, Y.; Wang, J.; Brown, C. T.
Show abstract
Single-cell RNA-seq (scRNA-seq) is a powerful tool for cell type identification but is not readily applicable to organisms without well-annotated reference genomes. Of the approximately 10 million animal species predicted to exist on Earth, >99.9% do not have any submitted genome assembly. To enable scRNA-seq for the vast majority of animals on the planet, here we introduce the concept of "k-mer homology," combining biochemical synonyms in degenerate protein alphabets with uniform data subsampling via MinHash into a pipeline called Kmermaid. Implementing this pipeline enables direct detection of similar cell types across species from transcriptomic data without the need for a reference genome. Underpinning Kmermaid is the tool Orpheum, a memory-efficient method for extracting high-confidence protein-coding sequences from RNA-seq data. After validating Kmermaid using datasets from human and mouse lung, we applied Kmermaid to the Chinese horseshoe bat (Rhinolophus sinicus), where we propagated cellular compartment labels at high fidelity. Our pipeline provides a high-throughput tool that enables analyses of transcriptomic data across divergent species transcriptomes in a genome- and gene annotation-agnostic manner. Thus, the combination of Kmermaid and Orpheum identifies cell type-specific sequences that may be missing from genome annotations and empowers molecular cellular phenotyping for novel model organisms and species.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Enhanced recovery of single-cell RNA-sequencing reads for missing gene expression data 95%
- Towards Universal Cell Embeddings: Integrating Single-cell RNA-seq Datasets across Species with SATURN 94%
- Helixer--de novo Prediction of Primary Eukaryotic Gene Models Combining Deep Learning and a Hidden Markov Model 94%
Similar papers in this journal
Similar papers in this journal
- Highly accurate reference and method selection for universal cross-dataset cell type annotation with CAMUS 95%
- Automated quality control and cell identification of droplet-based single-cell data using dropkick 95%
- An atlas of fish genome evolution reveals delayed rediploidization following the teleost whole-genome duplication 95%
Similar papers in this journal
- Comprehensive prediction of robust synthetic lethality between paralog pairs in cancer cell lines 94%
- Geometric Sketching Compactly Summarizes the Single-Cell Transcriptomic Landscape 93%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.