Impact of sample size and tissue relevance on T2D gene identification
Davtian, D.; Dupuis, T.; Mansour Aly, D.; Atabaki Pasdar, N.; Walker, M.; Franks, P. W.; Rutters, F.; Im, H. K. W.; Pearson, E. R.; van de Bunt, M.; Vinuela, A.; Brown, A.
Show abstract
Identification of genes and proteins mediating the activity of GWAS variants requires molecular data from disease relevant tissues, but these may be difficult to collect. Using multiple gene expression reference datasets and GWAS summary statistics for T2D we identified 1,818 unique genes associated with T2D. Comparing the performance of different reference datasets, we found that sample size, and not the relevance of the tissue to the disease, was the critical factor in identifying relevant genes. Genes implicated using a well powered expression dataset were also more likely to have multiple lines of genetic evidence. A targeted proteomics reference dataset from plasma samples showed similar power to identify T2D related proteins as gene expression with the same sample size. Accounting for BMI reduces power across all tissues and phenotypes by [~]30%, suggesting that many GWAS links to T2D are mediated by BMI, potentially implicating insulin resistance related effects. Finally, using data from smaller GWAS studies with precisely defined T2D subtypes uncovers genes directly relevant to that subtype, such as LST1, an immune response gene for Severe Autoimmune Diabetes and TRMT2A, involved in beta-cell apoptosis, for Severe Insulin Deficient Diabetes. Our work demonstrates the benefits of well powered reference datasets in accessible tissues and well-defined disease subtypes when studying complex diseases involving multiple tissues.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrative proteogenomic analyses provide novel interpretations of type 1 diabetes risk loci through circulating proteins 96%
- Characterizing common and rare variations in non-traditional glycemic biomarkers using multivariate approaches on multi-ancestry ARIC study 96%
- Genome-Wide Association Meta-Analysis Using a Recessive Model Illuminates Genetic Architecture of Type 2 Diabetes 93%
Similar papers in this journal
- The power of TOPMed imputation for the discovery of Latino enriched rare variants associated with type 2 diabetes 95%
- Epigenome-wide association study of incident type 2 diabetes in Black and White participants from the Atherosclerosis Risk in Communities Study 95%
- High-throughput Genetic Clustering of Type 2 Diabetes Loci Reveals Heterogeneous Mechanistic Pathways of Metabolic Disease 94%
Similar papers in this journal
Similar papers in this journal
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 92%
- Metabolic reaction fluxes as amplifiers and buffers of risk alleles for coronary artery disease 92%
- Machine learning-guided deconvolution of plasma protein levels 91%
Similar papers in this journal
- Genome-wide study on 72,298 Korean individuals in Korean biobank data for 76 traits identifies hundreds of novel loci 95%
- Genetic associations with ratios between protein levels detect new pQTLs and reveal protein-protein interactions 94%
- Type 1 diabetes risk genes mediate pancreatic beta cell survival in response to proinflammatory cytokines 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.