Quality control of large genome datasets using genome fingerprints
Robinson, M.; Joshi, A.; Vidyarthi, A.; Maccoun, M.; Rangavajjhala, S.; Glusman, G.
Show abstract
The 1000 Genomes Project (TGP) is a foundational resource which serves the biomedical community as a standard reference cohort for human genetic variation. There are now seven public versions of these genomes. The TGP Consortium produced the first by mapping its final data release against human reference sequence GRCh37, then "lifted over these genomes to the improved reference sequence (GRCh38) when it was released, and remapped the original data to GRCh38 with two similar pipelines. As best practice quality validation, the pipelines that generated these versions were benchmarked against the Genome In A Bottle Consortiums platinum quality genome (NA12878). The New York Genome Center recently released the results of independently resequencing the cohort at greater depth (30X), a phased version informed by the inclusion of related individuals, and independently remapped the original variant calls to GRCh38. We evaluated all seven versions using genome fingerprinting, which supports ultrafast genome comparison even across reference versions. We noted multiple issues including discrepancies in cohort membership, disagreement on the overall level of variation, evidence of substandard pipeline performance on specific genomes and in specific regions of the genome, cryptic relationships between individuals, inconsistent phasing, and annotation distortions caused by the history of the reference genome itself. We therefore recommend global quality assessment by rapid genome comparisons, using genome fingerprints and other metrics, alongside benchmarking as part of best practice quality assessment of large genome datasets. Our observations also help inform the decision of which version to use, to support analyses by individual researchers.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 94%
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 94%
- A Complete Pedigree-Based Graph Workflow for Rare Candidate Variant Analysis 93%
Similar papers in this journal
Similar papers in this journal
- A reference-quality, fully annotated genome from a Puerto Rican individual 94%
- Robust, flexible, and scalable tests for Hardy-Weinberg Equilibrium across diverse ancestries 92%
- Transmission distortion and genetic incompatibilities between alleles in a multigenerational mouse advanced intercross line 92%
Similar papers in this journal
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 95%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 94%
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 92%
Similar papers in this journal
- The FORCE panel: An all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications 93%
- VarGenius-HZD allows accurate detection of rare homozygous or hemizygous deletions in targeted sequencing leveraging breadth of coverage 92%
- Optical genome mapping as a next-generation cytogenomic tool for detection of structural and copy number variations for prenatal genomic analyses 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.