Back

Benchmarking of small and large variants across tandem repeats

English, A.; Dolzhenko, E.; Ziaei-Jam, H.; Olson, N. D.; Mckenzie, S.; De Coster, W.; Park, J.; Gu, B.; Wagner, J.; Eberle, M.; Gymrek, M.; Chaisson, M.; Zook, J. M.; Sedlazeck, F. J.

2023-11-01 bioinformatics
10.1101/2023.10.29.564632 bioRxiv
Show abstract

Tandem repeats (TRs) are highly polymorphic in the human genome, have thousands of associated molecular traits, and are linked to over 60 disease phenotypes. However, their complexity often excludes them from at-scale studies due to challenges with variant calling, representation, and lack of a genome-wide standard. To promote TR methods development, we create a comprehensive catalog of TR regions and explore its properties across 86 samples. We then curate variants from the GIAB HG002 individual to create a tandem repeat benchmark. We also present a variant comparison method that handles small and large alleles and varying allelic representation. The 8.1% of the genome covered by the TR catalog holds [~]24.9% of variants per individual, including 124,728 small and 17,988 large variants for the GIAB HG002 TR benchmark. We work with the GIAB community to demonstrate the utility of this benchmark across short and long read technologies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.