Back

Aardvark: Sifting through differences in a mound of variants

Holt, J. M.; Kronenberg, Z.; Saunders, C. T.; Dolzhenko, E.; Krusche, P.; Olson, N. D.; Zook, J. M.; Eberle, M. A.

2025-10-06 bioinformatics
10.1101/2025.10.03.680257 bioRxiv
Show abstract

Variant benchmarking is critical in assessing the accuracy of genomic secondary pipelines. However, traditional benchmarking tools that require exact genotype matches inject biases from variant representation and are ill-suited for tandem repeat or structural variation. We describe Aardvark, a variant benchmarking tool that introduces the basepair score to directly compare haplotype sequences, reducing representation biases while allowing for partial credit scoring. The tool also includes a traditional genotype score and supports separate or joint benchmarking of small variants, tandem repeats, and structural variants (<10 kb). Aardvark accepts standard inputs, runs {approx}16x faster than hap.py, and is freely available and open source (https://github.com/PacificBiosciences/aardvark).

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.