A wealth of novel cell-specific expressed SNVs from tumor and normal scRNA-seq datasets
Dillard, C.; Ulianova, E.; Prashant, N.; Liu, H.; Edwards, N. J.; Horvath, A.
Show abstract
We demonstrate a novel variant calling strategy using barcode-stratified alignments on 25 tumor and normal 10XGenomics scRNA-seq datasets (>200,000 cells). Our approach identified 24,528 exonic non-dbSNP single cell expressed (sce)SNVs, a third of which are shared across multiple samples. The novel sceSNVs include unreported somatic and germline variants, as well as RNA-originating variants; some are expressed in up to 17% of the cells, and many are found in known cancer genes. Our findings suggest that there is an unacknowledged repertoire of expressed genetic variants, possibly recurrent and common across samples, in the normal and cancer transcriptome.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detecting haplotype-specific transcript variation in long reads with FLAIR2 95%
- JAFFAL: Detecting fusion genes with long read transcriptome sequencing 94%
- CHESS 3: an improved, comprehensive catalog of human genes and transcripts based on large-scale expression data, phylogenetic analysis, and protein structure 94%
Similar papers in this journal
- T Cell Receptor Beta Germline Variability is Revealed by Inference From Repertoire Data 94%
- Retroelement co-option disrupts the cancer transcriptional programme 93%
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 92%
Similar papers in this journal
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 93%
- Enhancing the annotation of small ORF-altering variants using MORFEE: introducing MORFEEdb, a comprehensive catalog of SNVs affecting upstream ORFs in human 5'UTRs 93%
- Covering all your bases: incorporating intron signal from RNA-seq data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.