Accurate prediction of CRISPR editing outcomes in somatic cell lines and zygote with few-shot learning
Zheng, W.; Yu, L.; Wang, G.; Cao, S.; Ho, K.; Song, J.; Cheng, C.; Ho, J. W. K.; Liu, X.; Wu, M.; Liu, Z.; Wang, H.; Liu, P.; Lan, G.; Huang, Y.
Show abstract
The CRISPR-Cas system has revolutionized gene editing, while its outcome prediction remains unsatisfactory, especially in new cell states due to their distinct DNA repair preferences. In this study, we introduce inDecay, a flexible system for predicting CRISPR editing outcomes from target sequence, returning probabilities of nearly the full spectrum of indel events. Uniquely, inDecay utilizes informative and parameter-efficient features for each indel event and incorporates cell-type-specific repair preferences through a multi-stage design. While both inDecay and existing methods achieve accurate results for prediction within cell lines, only inDecay with transfer learning can retain the high performance for cross-cell line prediction. We then applied inDecay to mouse embryo editing using our newly generated data and observed remarkable accuracy by including as few as 30 fine-tuning embryonic samples. Notably, inDecay is the first software to predict embryonic editing. Therefore, our few-shot learning-supported system may accelerate guide RNA prioritization in mouse model generation, mammal embryonic gene editing, and cellular therapeutics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Linking genotypes with multiple phenotypes in single-cell CRISPR screens 95%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 95%
- Genome-wide CRISPR guide RNA design and specificity analysis with GuideScan2 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.