Single cell and spatial alternative splicing analysis with long read sequencing
Fu, Y.; Kim, H.; Adams, J. I.; Grimes, S. M.; Huang, S.; Lau, B.; Sathe, A.; Hess, P.; Ji, H.; Zhang, N.
Show abstract
Long-read sequencing has become a powerful tool for alternative splicing analysis. However, technical and computational challenges have limited our ability to explore alternative splicing at single cell and spatial resolution. The higher sequencing error of long reads, especially high indel rates, have limited the accuracy of cell barcode and unique molecular identifier (UMI) recovery. Read truncation and mapping errors, the latter exacerbated by the higher sequencing error rates, can cause the false detection of spurious new isoforms. Downstream, there is yet no rigorous statistical framework to quantify splicing variation within and between cells/spots. In light of these challenges, we developed Longcell, a statistical framework and computational pipeline for accurate isoform quantification for single cell and spatial spot barcoded long read sequencing data. Longcell performs computationally efficient cell/spot barcode extraction, UMI recovery, and UMI-based truncation- and mapping-error correction. Through a statistical model that accounts for varying read coverage across cells/spots, Longcell rigorously quantifies the level of inter-cell/spot versus intra-cell/ spot diversity in exon-usage and detects changes in splicing distributions between cell populations. Applying Longcell to single cell long-read data from multiple contexts, we found that intra-cell splicing heterogeneity, where multiple isoforms co-exist within the same cell, is ubiquitous for highly expressed genes. On matched single cell and Visium long read sequencing for a tissue of colorectal cancer metastasis to the liver, Longcell found concordant signals between the two data modalities. Finally, on a perturbation experiment for 9 splicing factors, Longcell identified regulatory targets that are validated by targeted sequencing.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Specific splice junction detection in single cells with SICILIAN 97%
- Comprehensive characterization of single cell full-length isoforms in human and mouse with long-read sequencing 97%
- ZetaSuite, A Computational Method for Analyzing Multi-dimensional High-throughput Data, Reveals Genes with Opposite Roles in Cancer Dependency 97%
Similar papers in this journal
Similar papers in this journal
- Automated quality control and cell identification of droplet-based single-cell data using dropkick 96%
- scTIE: data integration and inference of gene regulation using single-cell temporal multimodal data 96%
- Highly accurate reference and method selection for universal cross-dataset cell type annotation with CAMUS 95%
Similar papers in this journal
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 97%
- The SpliZ generalizes "Percent Spliced In" to reveal regulated splicing at single-cell resolution 97%
- Cancer subclone detection based on DNA copy number in single cell and spatial omic sequencing data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.