CLADES - Contrastive Learning Augmented DifferEntial Splicing with Orthologous Positive Pairs
Talukder, A.; Keung, N.; Pe'er, I.; Knowles, D. A.
Show abstract
Alternative splicing (AS) reshapes transcript and protein repertoires across biological, e.g. cellular, contexts. However, learning sequence [->] content-specific splicing mappings is challenging due to limited labels across tissues and cell types and variability introduced by experimental protocols. We propose a contrastive representation learning pre-training approach grounded in evolutionary conservation. Orthologous exon-intron junction sequences are treated as semantically consistent views of the same regulatory program: evolutionary orthologs are positive pairs, non-homologous junctions are negatives. This discriminative objective aligns embeddings of regulatory equivalents while separating functionally unrelated sequences, inducing invariances to unconstrained sequence and emphasizing conserved motif/RBP and positional signals. We show that this pre-training strategy provides representations that help predict{Delta}{psi} , the change in exon inclusion between conditions, which encodes both direction and magnitude of splicing shifts. Specifically, we finetune a lightweight supervised head on available labels to predict{Delta}{psi} . To make these predictions biologically meaningful, we further introduce an interpretable, splice-motif-aware classification framework grounded in known regulatory signals. On benchmarks spanning tissue- and cell-type differential splicing, the learned representations yield strong{Delta}{psi} classification performance (AUPRC/AUROC for increased/decreased inclusion) and competitive results for regression (RMSE, Spearman). These findings indicate that evolution-as-augmentation, instantiated via contrastive learning, is an effective and biologically principled route to context-resolved splicing prediction.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 96%
- Single-Cell Omics for Transcriptome CHaracterization (SCOTCH): isoform-level characterization of gene expression through long-read single-cell RNA sequencing 95%
- MOCCASIN: A method for correcting for known and unknown confounders in RNA splicing analysis 95%
Similar papers in this journal
- Using single-cell perturbation screens to decode the regulatory architecture of splicing factor programs 95%
- Shiba: A versatile computational method for systematic identification of differential RNA splicing across platforms 95%
- Inference of cell state transitions and cell fate plasticity from single-cell with MARGARET 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.