Pan-cell-type prediction of splicing patterns from sequence and splicing factor expression
Vetsigian, K.; Lancaster, J.; Ieremie, I.; Radens, C. M.; Smyth, P.; Young, S.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWAlternative splicing is a core determinant of cell-type-specific gene expression in humans, and its dysregulation contributes to many diseases including neurodegeneration, autoimmunity, and cancer. However, current deep learning models for predicting RNA expression from sequence are limited in how they handle cellular context dependence. These models typically achieve cell-type specificity by training separate models or heads for each tissue or cell-type, which assumes discrete, predefined cell types. This design not only prevents learning from pathological or experimentally perturbed transcriptomes, but also prevents generalization to new cellular contexts. Here we introduce PanExonNet, a deep learning framework that integrates cis- and trans-regulation by conditioning sequence-based splicing predictions on an inferred "splicing state" derived from the expression of RNA-binding proteins (RBPs) and spliceosome components. PanExonNet is trained on diploid, individual-specific gene sequences containing indels to predict splicing profiles derived from short-read RNA-seq, and it outputs multiple tracks at single-nucleotide resolution as well as donor-acceptor junction usage. Compared to multi-headed baselines such as Borzoi and Pangolin, PanExonNet exhibits substantially higher context specificity--as quantified by a {Delta}PSI correlation metric designed to isolate context-specific variation--and, crucially, generalizes to unseen cell types. Adding perturbational knockdown followed by RNA sequencing (KD-RNA-seq) data to the training set further improves generalization. This performance is enabled in part by our introduction of contextualizable convolutions, a modular layer that may broadly benefit genomic sequence modeling. This framework provides a scalable foundation for future DNA-to-RNA models that could improve variant effect prediction, the design of oligonucleotide therapeutics, and biomarker discovery across diverse cellular contexts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- Machine learning-optimized targeted detection of alternative splicing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.