Back

Streamlining Large-Scale Genomic Data Management: Insights from the UK Biobank Whole-Genome Sequencing Data

Li, X.; Wood, A. R.; Yuan, Y.; Zhang, M.; Huang, Y.; Hawkes, G.; Beaumont, R. N.; Weedon, M. N.; Li, W.; Li, X.; Lin, X.; Li, Z.

2025-01-28 genetic and genomic medicine
10.1101/2025.01.27.25321225 medRxiv
Show abstract

Biobank-scale Whole-Genome Sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets sheer volume and complexity presents significant challenges. We propose vcf2agds, an all-in-one toolkit that efficiently converts WGS data from Variant Call Format (VCF) format to the annotated Genomic Data Structure (aGDS) format, significantly reducing data size while supporting seamless genomic and functional data integration for comprehensive genetic analyses. The toolkit was applied to the UK Biobank 500k WGS data, resulting in twenty-three aGDS files, one for each chromosome, which collectively compressed 1,473.85 Tebibytes of pVCF data into 1.10 Tebibytes. Utilizing these aGDS files, we conducted a functionally informed rare variant association analysis of total cholesterol employing the STAARpipeline and detected 480 genome-wide significant coding and noncoding associations. Overall, vcf2agds offers a streamlined approach facilitating the efficient management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Published in Cell Genomics (predicted rank #4) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.