Back

Data Descriptor: Human whole exome genotype data for Alzheimer's Disease

Leung, Y. Y.; Naj, A. C.; Chou, Y.-F.; Valladares, O.; Wheeler, N.; Lin, H.; Gangadharan, P.; Qu, L.; Clark, K.; Cantwell, L.; Nicaretta, H.; the Alzheimer's Disease Sequencing Project, ; Seshadri, S.; Brkanac, Z.; Cruchaga, C.; Pericak-Vance, M. A.; Mayeux, R.; Kuzma, A. B.; Lee, W.-P.; Bush, W. S.; DeStefano, A. L.; Martin, E.; Schellenberg, G. D.; Wang, L.-S.

2022-10-13 genetics
10.1101/2022.10.11.511653 bioRxiv
Show abstract

Bigger sample size can help to identify new genetic variants contributing to an increased risk of developing Alzheimers disease. However, the heterogeneity of the whole-exome sequencing (WES) data generation methods presents a challenge to a joint analysis. Here we present a bioinformatics strategy for joint calling 20,504 WES samples collected across nine studies and sequenced using ten different capture kits in fourteen sequencing centers in the Alzheimers Disease Sequencing Project. gVCFs of samples were joint-called by the Genome Center for Alzheimers Disease into a single VCF, containing only positions within the union of capture kits. The VCF was then processed using specific strategies to account for the batch effects arising from the use of different capture kits from different studies. We identified 8.2 million autosomal variants. 96.82% of the variants are high-quality, and are located in 28,579 Ensembl transcripts. 41% of the variants are intronic and 15% are missense variants. 1.8% of the variants are with CADD>30. Our new strategy for processing these diversely generated WES samples has shown to generate high-quality data. The improved ability to combine data sequenced in different batches benefits the whole genomics research community. The WES data are accessible to the scientific community via https://dss.niagads.org/.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.