Back

Population Stratification at the Phenotypic Variance level and Implication for the Analysis of Whole Genome Sequencing Data from Multiple Studies

Sofer, T.; Zheng, X.; Laurie, C. A.; Gogarten, S. M.; Brody, J. A.; Conomos, M. P.; Bis, J. C.; Thornton, T. A.; Szpiro, A.; O'Connell, J. R.; Lange, E. M.; Gao, Y.; Cupples, L. A.; Psaty, B. M.; Trans-Omics for Precision Medicine (TOPMed) Consortium, ; Rice, K. M.

2020-03-05 genetics
10.1101/2020.03.03.973420 bioRxiv
Show abstract

In modern Whole Genome Sequencing (WGS) epidemiological studies, participant-level data from multiple studies are often pooled and results are obtained from a single analysis. We consider the impact of differential phenotype variances by study, which we term variance stratification. Unaccounted for, variance stratification can lead to both decreased statistical power, and increased false positives rates, depending on how allele frequencies, sample sizes, and phenotypic variances vary across the studies that are pooled. We describe a WGS-appropriate analysis approach, implemented in freely-available software, which allows study-specific variances and thereby improves performance in practice. We also illustrate the variance stratification problem, its solutions, and a corresponding diagnostic procedure in data from the Trans-Omics for Precision Medicine Whole Genome Sequencing Program (TOPMed), used in association tests for hemoglobin concentrations and BMI.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.