Case-base-control designs
Elhezzani, N. S.; Bergsma, W.; Weale, M.
Show abstract
Most genome-wide association studies (GWASs) use randomly selected samples from the population (hereafter bases) as the control set. This approach is successful when the trait of interest is rare; otherwise, a loss in the statistical power to detect disease-associated variants is expected. To address this, a proposal to combine the three sample types, cases, controls and bases is introduced, for instances when the disease under study is prevalent. This is done by modelling the bases as a mixture of multinomial logistic functions of cases and controls, according to the disease prevalence. The maximum likelihood method is used to estimate the underlying parameters using the EM algorithm. Three classical tests of association; score, Walds, and likelihood ratio tests are derived and their power of detecting genetic associations under different designs is compared. Simulations show that combining the three samples can increase the power to detect disease-associated variants, though a very large base sample set can compensate for the lack of controls.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Testing gene-environment interactions for rare and/or common variants in sequencing association studies 97%
- The PPLD has advantages over conventional regression methods in application to moderately sized genome-wide association studies 95%
- HCLC-FC: a novel statistical method for phenome-wide association studies 95%
Similar papers in this journal
- Two-phase sample selection strategies for design and analysis in post-genome wide association fine-mapping studies 97%
- Using a supervised principal components analysis for variable selection in high-dimensional datasets reduces false discovery rates 95%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.