A comprehensive tandem repeat catalog of the human genome
Chiu, R.; Rajan Babu, I. S.; Friedman, J. M.; Birol, I.
Show abstract
With the increasing availability of long-read sequencing data, high-quality human genome assemblies, and software for fully characterizing tandem repeats, genome-wide genotyping of tandem repeat loci on a population scale becomes more feasible. Such efforts not only expand our knowledge of the tandem repeat landscape in the human genome but also enhance our ability to differentiate pathogenic tandem repeat mutations from benign polymorphisms. To this end, we analyzed 272 genomes assembled using datasets from three public initiatives that employed different long-read sequencing technologies. Here, we report a catalog of over 18 million tandem repeat loci, many of which were previously unannotated. Some of these loci are highly polymorphic, and many of them reside within coding sequences.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- STRling: a k-mer counting approach that detects short tandem repeat expansions at known and novel loci 96%
- Quartet DNA reference materials and datasets for comprehensively evaluating germline variants calling performance 96%
- Paragraph: A graph-based structural variant genotyper for short-read sequence data 95%
Similar papers in this journal
- Multi-modal investigation of the schizophrenia-associated 3q29 genomic interval reveals global genetic diversity with unique haplotypes and segments that increase the risk for non-allelic homologous recombination 94%
- Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches 93%
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.