Large Serine Integrase Off-Target Discovery with Deep Learning for Genome Wide Prediction
Bakalar, M. H.; Biondi, T.; Liang, X.; Santesmasses, D.; Barra, A. M.; Mehta, J. B.; Wang, J.; Hazelbaker, D. Z.; Finn, J. D.; O'Connell, D. J.
Show abstract
Large Serine Integrases (LSIs) hold significant therapeutic promise due to their ability to efficiently incorporate gene-sized DNA into the human genome, offering a method to integrate healthy genes in patients with monogenic disorders or to insert gene circuits for the development of advanced cell therapies. To advance the application of LSIs for human therapeutic applications, new technologies and analytical methods for predicting and characterizing off-target recombination by LSIs are required. It is not experimentally tractable to validate off-target editing at all potential off-target sites in therapeutically relevant cell types because of sample limitations and genetic variation in the human population. To address this gap, we constructed a deep learning model named IntQuery that can predict LSI activity genome-wide. For Bxb1 integrase, IntQuery was trained on quantitative off-target data from 410,776 cryptic attB sequences discovered by Cryptic-seq, an unbiased in vitro discovery technology for LSI off-target recombination. We show that IntQuery can accurately predict in vitro LSI activity, providing a tool for in silico off-target prediction of large serine integrases to advance therapeutic applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing DSB repair to promote efficient homology-dependent and -independent prime editing 95%
- Improved prime editors enable pathogenic allele correction and cancer modelling in adult mice 95%
- A generalizable Cas9/sgRNA prediction model using machine transfer learning with small high-quality datasets 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.