Probabilistic Modelling of Prime Editing Variant CorrectionEfficiency
Ozden, F.; Lu, P.; Minary, P.
Show abstract
Prime editing has emerged as a versatile genome editing, technology capable of installing precise genetic modifications without requiring double-strand breaks or donor templates. However, designing pegRNAs with high editing efficiency remains a challenge due to the complex interplay of sequence features that affect editing outcomes. Current approaches predominantly provide point predictions without capturing the inherent uncertainty in editing efficiency, limiting risk assessment, and decision-making in pegRNA design. Here, we present crispAIPE, a transformer-based probabilistic framework for predicting prime editing variant correction efficiency with uncertainty quantification. Our approach models editing outcomes in a 3D simplex space, enabling comprehensive uncertainty estimation while achieving superior predictive performance compared to existing models. crispAIPE leverages transformer encoders to capture long-range sequence dependencies and contextual relationships, surpassing existing models in the point estimate prediction task. The model also predicts efficiency distributions for all edit types, including single nucleotide replacements, insertions, and deletions. Trained on 73, 939 pegRNAs in multiple cell lines, for all outcome types, on overall, crispAIPE achieves a Spearman correlation of 0.881 and a Pearson correlation of 0.894 while providing calibrated uncertainty estimates that allow the selection of risk-sensitive pegRNA. Additionally, we identify key sequence motifs and positional features that drive editing efficiency, providing interpretable insights into the sequence determinants of prime editing. We demonstrate that uncertainty-aware predictions could significantly improve pegRNA design outcomes, with high-confidence predictions showing higher success rates compared to low-confidence designs. crispAIPE represents the first probabilistic deep learning framework for prime editing, bridging the gap between predictive accuracy and uncertainty quantification to enable more reliable and interpretable pegRNA design. The source code and example data are available at https://github.com/furkanozdenn/pe-uncert.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards In-Silico CLIP-seq: Predicting Protein-RNA Interaction via Sequence-to-Signal Learning 95%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Pair consensus decoding improves accuracy of neural network basecallers for nanopore sequencing 95%
Similar papers in this journal
- DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data 95%
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 95%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 94%
Similar papers in this journal
- Genome-wide functional screens enable the prediction of high activity CRISPR-Cas9 and -Cas12a guides in Yarrowia lipolytica 95%
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 95%
- Clair3-RNA: A deep learning-based small variant caller for long-read RNA sequencing data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.