Phenotype prediction from genome-wide genotyping data: a crowdsourcing experiment
Naret, O.; Baranger, D.; Greshake Tzovaras, B.; Mohanty, S. P.; Salathe, M.; Fellay, J.
Show abstract
BackgroundThe increasing statistical power of genome-wide association studies is fostering the development of precision medicine through genomic predictions of complex traits. Nevertheless, it has been shown that the results remain relatively modest. A reason might be the nature of the methods typically used to construct genomic predictions. Recent machine learning techniques have properties that could help to capture the architecture of complex traits better and improve genomic prediction accuracy. MethodsWe relied on crowd-sourcing to efficiently compare multiple genomic prediction methods. This represents an innovative approach in the genomic field because of the privacy concerns linked to human genetic data. There are two crowd-sourcing elements building our study. First, we constructed a dataset from openSNP (opensnp.org), an open repository where people voluntarily share their genotyping data and phenotypic information in an effort to participate in open science. To leverage this resource we release the openSNP Cohort Maker, a tool that builds a homogeneous and up-to-date cohort based on the data available on opensnp.org. Second, we organized an open online challenge on the CrowdAI platform (crowdai.org) aiming at predicting height from genome-wide genotyping data. ResultsThe openSNP Height Prediction challenge lasted for three months. A total of 138 challengers contributed to 1275 submissions. The winner computed a polygenic risk score using the publicly available summary statistics of the GIANT study to achieve the best result (r2 = 0.53 versus r2 = 0.49 for the second-best). ConclusionWe report here the first crowd-sourced challenge on publicly available genome-wide genotyping data. We also deliver the openSNP Cohort Maker that will allow people to make use of the data available on opensnp.org.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Interpretable network-guided epistasis detection 91%
- The Great Genotyper: A Graph-Based Method for Population Genotyping of Small and Structural Variants 90%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 90%
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 92%
- genepanel.iobio - an easy to use web tool for generating disease- and phenotype-associated gene lists 90%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 90%
Similar papers in this journal
- Single-molecule optical mapping enables quantitative measurement of D4Z4 repeats in facioscapulohumeral muscular dystrophy (FSHD) 91%
- How do clinician and parent reported data differ? An analysis of similarity and difference in the datasets from a cross-syndrome genetics cohort study(GenROC) 89%
- Assessing performance of pathogenicity predictors using clinically-relevant variant datasets 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.