Simulating genetic risk scores from summary statistics
Squires, S.; Weedon, M. N.; Oram, R. A.
Show abstract
MotivationGenetic risk scores (GRS) summarise genetic data into a single number and allow for discrimination between cases and controls. Many applications of GRSs would benefit from comparisons with multiple datasets to assess quality of the GRS across different groups. However, genetic data is often unavailable. If summary statistics of the genetic data could be used to simulate GRSs more comparisons could be made, potentially leading to improved research. ResultsWe present a methodology that utilises only summary statistics of genetic data to simulate GRSs with an example of a type 1 diabetes (T1D) GRS. An example on European populations of the mean T1D GRS for real and simulated data are 10.31 (10.12-10.48) and 10.38 (10.24-10.53) respectively. An example of a case-control set for T1D has a area under the receiver operating characteristic curve of 0.917 (0.903-0.93) for real data and 0.914 (0.898-0.929) for simulated data. AvailabilityThe code is available at https://github.com/stevensquires/simulating_genetic_risk_scores. Contacts.squires@exeter.ac.uk
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The PPLD has advantages over conventional regression methods in application to moderately sized genome-wide association studies 94%
- An R Package for Divergence Analysis of Omics Data 93%
- Estimating multiplicity of infection, haplotype frequencies, and linkage disequilibria from multi-allelic markers for molecular disease surveillance 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.