Streamlining Large-Scale Genomic Data Management: Insights from the UK Biobank Whole-Genome Sequencing Data
Li, X.; Wood, A. R.; Yuan, Y.; Zhang, M.; Huang, Y.; Hawkes, G.; Beaumont, R. N.; Weedon, M. N.; Li, W.; Li, X.; Lin, X.; Li, Z.
Show abstract
Biobank-scale Whole-Genome Sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets sheer volume and complexity presents significant challenges. We propose vcf2agds, an all-in-one toolkit that efficiently converts WGS data from Variant Call Format (VCF) format to the annotated Genomic Data Structure (aGDS) format, significantly reducing data size while supporting seamless genomic and functional data integration for comprehensive genetic analyses. The toolkit was applied to the UK Biobank 500k WGS data, resulting in twenty-three aGDS files, one for each chromosome, which collectively compressed 1,473.85 Tebibytes of pVCF data into 1.10 Tebibytes. Utilizing these aGDS files, we conducted a functionally informed rare variant association analysis of total cholesterol employing the STAARpipeline and detected 480 genome-wide significant coding and noncoding associations. Overall, vcf2agds offers a streamlined approach facilitating the efficient management and analysis of biobank-scale WGS data across hundreds of thousands of samples.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 93%
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 93%
- ColocQuiaL: A QTL-GWAS colocalization pipeline 93%
Similar papers in this journal
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 95%
- Efficient phasing and imputation of low-coverage sequencing data using large reference panels 95%
- Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.