Scywalker: scalable end-to-end data analysis workflow for nanopore single-cell transcriptome sequencing
De Rijk, P.; Watzeels, T.; Kucukali, F.; Van Dongen, J.; Faura, J.; Willems, P.; De Deyn, L.; Duchateau, L.; Grones, C.; Eekhout, T.; De Pooter, T.; Joris, G.; Rombauts, S.; De Rybel, B.; Rademakers, R.; Van Breusegem, F.; Strazisar, M.; Sleegers, K.; De Coster, W.
Show abstract
We introduce scywalker, an innovative and scalable package developed to comprehensively analyze long-read nanopore sequencing data of full-length single-cell or single-nuclei cDNA. Existing nanopore single-cell data analysis tools showed severe limitations in handling current data sizes. We developed novel scalable methods for cell barcode demultiplexing and single-cell isoform calling and quantification and incorporated these in an easily deployable package. Scywalker streamlines the entire analysis process, from sequenced fragments in FASTQ format to demultiplexed pseudobulk isoform counts, into a single command suitable for execution on either server or cluster. Scywalker includes data quality control, cell type identification, and an interactive report. Assessment of datasets from the human brain, Arabidopsis leaves, and previously benchmarked data from mixed cell lines, demonstrate excellent correlation with short-read analyses at both the cell-barcoding and gene quantification levels. At the isoform level, we show that scywalker facilitates the direct identification of cell-type-specific expression of novel isoforms.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- QClus: A droplet-filtering algorithm for enhanced snRNA-seq data quality in challenging samples 96%
- Disentangling single-cell omics representation with a power spectral density-based feature extraction 96%
- LINE-1 Retrotransposon expression in cancerous, epithelial and neuronal cells revealed by 5'-single cell RNA-Seq 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.