Robust Integration of Sparse Single-Cell Alternative Splicing and Gene Expression Data with SpliceVI
Vaidyanathan, S.; Isaev, K.; Zweig, A.; Knowles, D. A.
Show abstract
Alternative splicing (AS) and gene expression (GE) are tightly related regulatory processes, critical for defining cell types and states, yet are rarely modeled together in single-cell analyses. This hinders a comprehensive understanding of cellular identity. We address this by introducing SpliceVI, adapted from MultiVI (Multi-modal Variational Inference) to specifically handle AS. Applied to a large multisample mouse Smart-seq2 dataset (n = 142, 315 cells/nuclei), SpliceVI jointly learns from both AS and GE using a partial variational autoencoder that effectively handles the sparsity and missingness of splicing data. We show that SpliceVIs joint embeddings are more expressive and informative of biological correlates like age than a GE-only approach (scVI). SpliceVI also uncovers splicingbased differences between neuronal subclusters. This approach reveals the distinct yet synergistic relationship between AS and GE in shaping cellular diversity in mouse.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Single-Cell Omics for Transcriptome CHaracterization (SCOTCH): isoform-level characterization of gene expression through long-read single-cell RNA sequencing 96%
- multiDGD: A versatile deep generative model for multi-omics data 95%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 95%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
- Liam tackles complex multimodal single-cell data integration challenges 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.