An efficient and accurate frailty model approach for genome-wide survival association analysis controlling for population structure and relatedness in large-scale biobanks
Dey, R.; Zhou, W.; Kiiskinen, T.; Havulinna, A.; Elliott, A.; Karjalainen, J.; Kurki, M.; Qin, A.; FinnGen, ; Lee, S.; Palotie, A.; Neale, B. M.; Daly, M. J.; Lin, X.
Show abstract
With decades of electronic health records linked to genetic data, large biobanks provide unprecedented opportunities for systematically understanding the genetics of the natural history of complex diseases. Genome-wide survival association analysis can identify genetic variants associated with ages of onset, disease progression and lifespan. We developed an efficient and accurate frailty (random effects) model approach for genome-wide survival association analysis of censored time-to-event (TTE) phenotypes in large biobanks by accounting for both population structure and relatedness. Our method utilizes state-of-the-art optimization strategies to reduce the computational cost. The saddlepoint approximation is used to allow for analysis of heavily censored phenotypes (>90%) and low frequency variants (down to minor allele count 20). We demonstrated the performance of our method through extensive simulation studies and analysis of five TTE phenotypes, including lifespan, with heavy censoring rates (90.9% to 99.8%) on ~400,000 UK Biobank participants with white British ancestry and ~180,000 samples in FinnGen, respectively. We further performed genome-wide association analysis for 871 TTE phenotypes in UK Biobank and presented the genome-wide scale phenome-wide association (PheWAS) results with the PheWeb browser.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The impact of non-additive genetic associations on age-related complex diseases. 98%
- OTTERS: A powerful TWAS framework leveraging summary-level reference data 97%
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 96%
Similar papers in this journal
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 97%
- A resource-efficient tool for mixed model association analysis of large-scale data 97%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 97%
Similar papers in this journal
- Capturing additional genetic risk from family history for improved polygenic risk prediction 96%
- Multivariate analysis reveals shared genetic architecture of brain morphology and human behavior 95%
- Accuracy of haplotype estimation and whole genome imputation affects complex trait analyses in complex biobanks 95%
Similar papers in this journal
- Controlling for background genetic effects using polygenic scores improves the power of genome-wide association studies 96%
- Polygenic prediction of human longevity on the supposition of pervasive pleiotropy 95%
- Synteny: a high throughput web tool to streamline causal gene prioritisation and provide insight into protein function 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.