Back

CLADES - Contrastive Learning Augmented DifferEntial Splicing with Orthologous Positive Pairs

Talukder, A.; Keung, N.; Pe'er, I.; Knowles, D. A.

2026-02-21 genomics
10.64898/2026.02.20.707118 bioRxiv
Show abstract

Alternative splicing (AS) reshapes transcript and protein repertoires across biological, e.g. cellular, contexts. However, learning sequence [->] content-specific splicing mappings is challenging due to limited labels across tissues and cell types and variability introduced by experimental protocols. We propose a contrastive representation learning pre-training approach grounded in evolutionary conservation. Orthologous exon-intron junction sequences are treated as semantically consistent views of the same regulatory program: evolutionary orthologs are positive pairs, non-homologous junctions are negatives. This discriminative objective aligns embeddings of regulatory equivalents while separating functionally unrelated sequences, inducing invariances to unconstrained sequence and emphasizing conserved motif/RBP and positional signals. We show that this pre-training strategy provides representations that help predict{Delta}{psi} , the change in exon inclusion between conditions, which encodes both direction and magnitude of splicing shifts. Specifically, we finetune a lightweight supervised head on available labels to predict{Delta}{psi} . To make these predictions biologically meaningful, we further introduce an interpretable, splice-motif-aware classification framework grounded in known regulatory signals. On benchmarks spanning tissue- and cell-type differential splicing, the learned representations yield strong{Delta}{psi} classification performance (AUPRC/AUROC for increased/decreased inclusion) and competitive results for regression (RMSE, Spearman). These findings indicate that evolution-as-augmentation, instantiated via contrastive learning, is an effective and biologically principled route to context-resolved splicing prediction.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.