WxS-QC - a quality control pipeline for human Whole-Genome and Whole Exome sequencing cohorts
Zakharov, G.; Eberhardt, R.; Frolova, A.; Horyslavets, D.; Hidari, E.; Popov, I.; Dennehy, J.; Koreshkov, M.; Antoniou, P.; Kantsypa, V.; Ovchinnikov, V.; Iyer, V.
Show abstract
MotivationWhole-exome (WES) and whole-genome (WGS) sequencing are rapidly becoming preferred methods for population-scale analysis of the human genetic landscape. However, there are currently no standardized quality control (QC) pipelines for human WES and WGS datasets. Moveover, there are no open datasets that can be used to test QC pipelines, because most projects (like 1000 genomes and gnomAD) publish only post-QC results. ResultsWe present WxS-QC, a powerful, scalable, and convenient pipeline for the QC of human WGS and WES cohorts, developed at the Wellcome Sanger Institute (WSI). Our pipeline is based on a deep refactoring of the gnomAD quality control pipeline code and is aligned with current best practices in WGS/WES cohort data QC. It offers a set of novel QC techniques, automatic export of resulting graphs and summary tables, excellent performance and scalability, and incorporates comprehensive documentation. To test our pipeline and similar solutions, we also assembled an open dataset with all required metadata. Availability and implementationThe pipeline code is written in Python using the Hail library and is freely available under the BSD-3 license here: https://github.com/wtsi-hgi/wxs-qc. It can run in any UNIX-like environment and has been able to efficiently process cohorts of up to 200,000 whole-exome samples with the potential to handle bigger datasets, depending on the available hardware. The detailed description of the pipeline, resources and test data is available in the pipeline documentation: https://github.com/wtsi-hgi/wxs-qc/blob/main/README.md The open test dataset is available to download from https://wxs-qc-data.cog.sanger.ac.uk/wxs-qc_public_dataset_v3.tar. An example of test dataset analysis is available in the supplementary materials.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SVIM-asm: Structural variant detection from haploid and diploid genome assemblies 96%
- Generative Haplotype Prediction Outperforms Statistical Methods for Small Variant Detection in NGS Data 95%
- MosaiCatcher v2: a single-cell structural variations detection and analysis reference framework based on Strand-seq 95%
Similar papers in this journal
Similar papers in this journal
- GGoutlieR: an R package to identify and visualize unusual geo-genetic patterns of biological samples 92%
- BREADR: An R Package for the Bayesian Estimation of Genetic Relatedness from Low-coverage Genotype Data 91%
- alignparse: A Python package for parsing complex features from high-throughput long-read sequencing 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.