Back

Quality control of large genome datasets using genome fingerprints

Robinson, M.; Joshi, A.; Vidyarthi, A.; Maccoun, M.; Rangavajjhala, S.; Glusman, G.

2021-12-06 genomics
10.1101/600254 bioRxiv
Show abstract

The 1000 Genomes Project (TGP) is a foundational resource which serves the biomedical community as a standard reference cohort for human genetic variation. There are now seven public versions of these genomes. The TGP Consortium produced the first by mapping its final data release against human reference sequence GRCh37, then "lifted over these genomes to the improved reference sequence (GRCh38) when it was released, and remapped the original data to GRCh38 with two similar pipelines. As best practice quality validation, the pipelines that generated these versions were benchmarked against the Genome In A Bottle Consortiums platinum quality genome (NA12878). The New York Genome Center recently released the results of independently resequencing the cohort at greater depth (30X), a phased version informed by the inclusion of related individuals, and independently remapped the original variant calls to GRCh38. We evaluated all seven versions using genome fingerprinting, which supports ultrafast genome comparison even across reference versions. We noted multiple issues including discrepancies in cohort membership, disagreement on the overall level of variation, evidence of substandard pipeline performance on specific genomes and in specific regions of the genome, cryptic relationships between individuals, inconsistent phasing, and annotation distortions caused by the history of the reference genome itself. We therefore recommend global quality assessment by rapid genome comparisons, using genome fingerprints and other metrics, alongside benchmarking as part of best practice quality assessment of large genome datasets. Our observations also help inform the decision of which version to use, to support analyses by individual researchers.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.