Systematic analysis of insertions signature in gnomAD revealed large set of novel processed pseudogenes
Podvalnyi, A.; Kucherenko, V.; Doroschuk, N.; Sarygina, E.; Sagaydak, O.; Mityaeva, O.; Bogdanov, V.; Krupinova, J.; Woroncow, M.; Volchkov, P.; Albert, E.
Show abstract
Pseudogenes are non-functional copies of protein-coding genes that arise through genomic duplication or retrotransposition. Processed pseudogenes (PPs) is the most abundant class of pseudogenes, which is generated via mRNA reverse transcription and subsequent cDNA integration. Presence of PPs complicates the analysis of short read sequencing data due to high similarity with parental gene and frequent absence from reference genome. Here we demonstrate that the presence of non-reference (absent from reference genome) PPs leads to the very distinctive artefact of germline variant calling - long insertions on exon-intron boundaries, which sequences could be mapped to other exons of the same gene. We showed that by detecting these artifacts it is possible to identify non-reference PPs existence based on the cohort summary statistics without analysing sample-level data. We used identified signature of PPs presence to systematically mine the gnomAD database which currently contains over 70,000 whole-genome and over 700,000 exome samples to describe novel non-reference PPs. Our approach uncovered 1498 non-reference PPs of which 1268 were novel and absent in the latest GENCODE release. This resource enhances the accuracy of variant interpretation and contributes to a deeper understanding of pseudogenes diversity across human populations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discovery of an unusual high number of de novo mutations in sperm of older men using duplex sequencing 94%
- Long-read genome sequencing and variant reanalysis increase diagnostic yield in neurodevelopmental disorders 94%
- Discordant calls across genotype discovery approaches elucidate variants with systematic errors 94%
Similar papers in this journal
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 94%
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 93%
- Mutational Constraint Analysis Workflow for Overlapping Short Open Reading Frames and Genomic Neighbours 93%
Similar papers in this journal
- DeNovoCNN: A deep learning approach to de novo variant calling in next generation sequencing data 93%
- Long-read whole-genome sequencing-based concurrent haplotyping and aneuploidy profiling of single cells 93%
- TypeTE: a tool to genotype mobile element insertions from whole genome resequencing data 92%
Similar papers in this journal
- Establishment of an eHAP1 Human Haploid Cell Line Hybrid Reference Genome Assembled from Short and Long Reads 91%
- Guiding eQTL mapping and genomic prediction of gene expression in three pig breeds with tissue-specific epigenetic annotations from early development 90%
- Comparative analysis of capture methods for genomic profiling of circulating tumor cells in colorectal cancer 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.