Identification of mobile element insertion from whole genome sequencing data using deep neural network model
Bu, F.; Xu, X.; Huang, Y.; Wang, X.; Cheng, J.; Yuan, H.
Show abstract
Mobile element insertions (MEIs) are a major contributor to genome evolution and play an essential role in the regulation of gene expression, as well as being implicated in various human diseases. This study introduces DeepMEI, a tool based on a convolutional neural network model that transforms the MEI identification process into an image recognition problem and automatically learns complex and abstract representations of MEI features in whole genome sequencing data. DeepMEI outperformed existing tools in the benchmark dataset from the Genome in a Bottle consortium, with a precision of 0.90 and recall of 0.70. Moreover, factors such as sequencing depth, ME integrity, and genome mappability can affect MEI identification accuracy. Using DeepMEI, we reanalyzed 3,202 high-coverage whole-genome sequencing samples from the 1000 Genome Project (1kGP) phase 4 release, discovering 1.71-fold more non-reference MEIs, totaling 6,218,088, with 92.2% of the increase coming from rare MEIs (allele frequency <1%). This enhances our understanding of MEIs role in human disease and evolution. The DeepMEI tool and the updated 1kGP MEI dataset can be accessed at https://github.com/xuxif/DeepMEI.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pre-processing of paleogenomes: Mitigating reference bias and postmortem damage in ancient genome data 95%
- Personalized and graph genomes reveal missing signal in epigenomic data 95%
- Quartet DNA reference materials and datasets for comprehensively evaluating germline variants calling performance 94%
Similar papers in this journal
- Detection of simple and complex de novo mutations without, with, or with multiple reference sequences 95%
- Predicting unrecognized enhancer-mediated genome topology by an ensemble machine learning model 94%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 94%
Similar papers in this journal
- PEPATAC: An optimized pipeline for ATAC-seq data analysis with serial alignments 95%
- Detection of homozygous and hemizygous partial exon deletions by whole-exome sequencing 95%
- svCapture: Efficient and specific detection of very low frequency structural variant junctions by error-minimized capture sequencing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.