COPILOT: a Containerised wOrkflow for Processing ILlumina genOtyping daTa
Patel, H.; Lee, S.-H.; Breen, G.; Menzel, S.; Ojewunmi, O.; Dobson, R.
Show abstract
BackgroundThe Illumina genotyping microarrays generate data in image format, which is processed by the platform-specific software GenomeStudio, followed by an array of complex bioinformatics analyses. This process can be time-consuming, lead to reproducibility errors, and be a daunting task for novice bioinformaticians. ResultsHere we introduce the COPILOT (Containerised wOrkflow for Processing ILlumina genOtyping daTa) protocol, which provides an in-depth and clear guide to process raw Illumina genotype data in GenomeStudio, followed by a containerised workflow to automate an array of complex bioinformatics analyses involved in a GWAS quality control (QC). The COPILOT protocol was applied to two independent cohorts consisting of 2791 and 479 samples genotyped on the Infinium Global Screening (GSA) array with Multi-disease (MD) drop-in (~750,000 markers) and the Infinium H3Africa consortium array (~2,200,000 markers) respectively. Following the COPILOT protocol, an average sample quality improvement of 1.24% was observed across sample call rates, with notable improvement for low-quality samples. For example, from the 3270 samples processed, 141 samples had an initial sample call rate below 98%, averaging 96.6% (95% CI 95.6-97.7%), which is considered below the acceptable sample call rate threshold for a typical GWAS analysis. However, following the COPILOT protocol, all 141 samples had a call rate above 98% after QC and averaged 99.6% (95% CI 99.5-99.7%). In addition, the COPILOT pipeline automatically identified potential data issues, including gender discrepancies, heterozygosity outliers, related individuals, and population outliers through ancestry estimation. ConclusionsThe COPILOT protocol makes processing Illumina genotyping data transparent, effortless and reproducible. The container is deployable on multiple platforms, improves data quality, and the end product is analysis-ready PLINK formatted data, with a comprehensive and interactive summary report to guide the user for further data analyses.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 94%
- H3AGWAS : A portable workflow for Genome Wide Association Studies 94%
- ILIAD: A suite of automated Snakemake workflows for processing genomic data for downstream applications 94%
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 92%
- genepanel.iobio - an easy to use web tool for generating disease- and phenotype-associated gene lists 92%
- Accuracy and Reproducibility of Somatic Point Mutation Calling in Clinical-Type Targeted Sequencing Data 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.