VNTR prediction on sequence characteristics using long-read annotation and validation by short-read pileup
Cames van Batenburg, D. F. J.; Linthorst, J.; Holstege, H.; Reinders, M. J. T.
Show abstract
Tandem repeats (TRs) are contiguously repetitive sequences with a high mutation rate. Several human diseases have been associated with an expansion of TR, a mutation which constitutes a change in their number of repetitions. Nevertheless, these Variable Number Tandem Repeats (VNTRs) have not been included in many genome-wide studies. The reason is that VNTR genotyping is inaccurate using short-read sequencing while new technology like long-read sequencing is expensive and lacks throughput. Here, we propose a sequence based random forest classifier that is able to predict variable expansion of TR regions, given by incomplete VNTR annotation from long-read sequencing of 5 haplotypes. The classifier mainly predicted VNTRs using the features TR length. The second most used feature is a novel finding: the Mfold predicted likelihood of self-folding for which more stable foldings are correlated with VNTRs. We validated VNTR candidates predicted by this classifier by clustering short-read pileup patterns compared across 17 genomes. TRs labeled VNTR by the classifier showed similar local variance in their pileup profiles. Contactdiederik.cvb@gmail.com Supplementary informationSupplementary data are available at bioRxiv
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tailored machine learning models for functional RNA detection in genome-wide screens 95%
- Fast and memory-efficient mapping of short bisulfite sequencing reads using a two-letter alphabet 94%
- NGSTroubleFinder: A tool for detection and quantification of contamination and kinship across human NGS data 94%
Similar papers in this journal
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 95%
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 95%
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 94%
Similar papers in this journal
Similar papers in this journal
- Distinct sequencing success at non-B-DNA motifs 95%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 94%
- Haplocheck: Phylogeny-based Contamination Detection in Mitochondrial and Whole-Genome Sequencing Studies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.