Polygenic Prediction of Molecular Traits using Large-Scale Meta-analysis Summary Statistics
Pain, O.; Gerring, Z. F.; Derks, E. F.; Wray, N. R.; Gusev, A.; Al-Chalabi, A.
Show abstract
IntroductionTranscriptome-wide association study (TWAS) integrates expression quantitative trait loci (eQTL) data with genome-wide association study (GWAS) results to infer differential expression. TWAS uses multi-variant models trained using individual-level genotype-expression datasets, but methodological development is required for TWAS to utilise larger eQTL summary statistics. MethodsTWAS models predicting gene expression were derived using blood-based eQTL summary statistics from eQTLGen, the Young Finns Study (YFS), and MetaBrain. Summary statistic polygenic scoring methods were used to derive TWAS models, evaluating their predictive utility in GTEx v8. We investigated gene inclusion criteria and omnibus tests for aggregating TWAS associations for a given gene. We performed a schizophrenia TWAS using summary statistic-based TWAS models, comparing results to existing resources and methods. ResultsTWAS models derived using eQTL summary statistics performed comparably to models derived using individual-level data. Multi-variant TWAS models significantly improved prediction over single variant models for 8.6% of genes. TWAS models derived using eQTLGen summary statistics significantly improved prediction over models derived using a smaller individual-level dataset. The eQTLGen-based schizophrenia TWAS, using the ACAT omnibus test to aggregate associations for each gene, identified novel significant and colocalised associations compared to summary-based mendelian randomisation (SMR) and SMR-multi. ConclusionsUsing multi-variant TWAS models and larger eQTL summary statistic datasets can improve power to detect differential expression associations. We provide TWAS models based on eQTLGen and MetaBrain summary statistics, and software to easily derive and apply summary statistic-based TWAS models based on eQTL and other molecular QTL datasets released in the future.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integration of multidimensional splicing data and GWAS summary statistics for risk gene discovery 94%
- Combinations of genes at the 16p11.2 and 22q11.2 CNVs contribute to neurobehavioral traits 94%
- Tissue specificity-aware TWAS (TSA-TWAS) framework identifies novel associations with metabolic, immunologic, and virologic traits in HIV-positive adults 94%
Similar papers in this journal
- BinomiRare: A carriers-only test for association of rare genetic variants with a binary outcome for mixed models and any case-control proportion 94%
- Evaluation of imputation performance of multiple reference panels in a Pakistani population 94%
- Evaluating Genomic Polygenic Risk Scores for Childhood Acute Lymphoblastic Leukemia in Latinos 92%
Similar papers in this journal
- RNAseqCovarImpute: a multiple imputation procedure that outperforms complete case and single imputation differential expression analysis 93%
- GBAT: a gene-based association method for robust trans-gene regulation detection 92%
- Modelling group heteroscedasticity in single-cellRNA-seq pseudo-bulk data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.