Applications of machine learning to solve genetics problems
Sowunmi, K.; Soyebo, T. A.; Okosesi, E. A.; Adesiyan, A. L.; Oladimeji, K. A.; Ajibola, O. A.; Ogunlana, Y. O.; Agboola, O. W.; Kaur, G.; Atoromola, H.; Oladipupo, T. A.
Show abstract
The development of precise DNA editing nucleases that induce double-strand breaks (DSBs) - including zinc finger nucleases, TALENs, and CRISPR/Cas systems - has revolutionized gene editing and genome engineering. Endogenous DNA DSB repair mechanisms are often leveraged to enhance editing efficiency and precision. While the non-homologous end joining (NHEJ) and homologous recombination (HR) DNA DSB repair pathways have already been the topic of an excellent deal of investigation, an alternate pathway, microhomology-mediated end joining (MMEJ), remains relatively unexplored. However, the MMEJ pathways ability to supply reproducible and efficient deletions within the course of repair makes it a perfect pathway to be used in gene knockouts. (Microhomology Evoked Deletion Judication EluciDation) may be a random forest machine learning-based method for predicting the extent to which the location of a targeted DNA DSB are going to be repaired using the MMEJ repair pathway. On an independent test set of 24 HeLa cell DSB sites, MEDJED achieved a Pearson coefficient of correlation (PCC) of 81.36%, Mean Absolute Error (MAE) of 10.96%, and Root Mean Square Error (RMSE) of13.09%. This performance demonstrates MEDJEDs value as a tool for researchers who wish to leverage MMEJ to supply efficient and precise gene knock outs.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 92%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 92%
- VarSCAT: A computational tool for sequence context annotations of genomic variants 92%
Similar papers in this journal
Similar papers in this journal
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 92%
- ENNGene: an Easy Neural Network model building tool for Genomics 92%
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.