Novel discoveries and enhanced genomic prediction from modelling genetic risk of cancer age-at-onset
Ojavee, S. E.; Maksimova, E. S.; Läll, K.; Sadler, M. C.; Mägi, R.; Kutalik, Z.; Robinson, M. R.
Show abstract
Genome-wide association studies seek to attribute disease risk to DNA regions and facilitate subject-specific prediction and patient stratification. For later-life diseases, inference from case-control studies is hampered by the uncertainty that control group subjects might later be diagnosed. Time-to-event analysis treats controls as right-censored, making no additional assumptions about future disease occurrence and represents a more sound conceptual alternative for more accurate inference. Here, using data on 11 common cancers from the UK and Estonian Biobank studies, we provide empirical evidence that discovery and genomic prediction are greatly improved by analysing age-at-diagnosis, compared to a case-control model of association. We replicate previous findings from large-scale case-control studies and find an additional 7 previously unreported independent genomic regions, out of which 3 replicated in independent data. Our novel discoveries provide new insights into underlying cancer pathways, and our model yields a better understanding of the polygenicity and genetic architecture of the 11 tumours. We find that heritable germline genetic variation plays a vital role in cancer occurrence, with risk attributable to many thousands of underlying genomic regions. Finally, we show that Bayesian modelling strategies utilising time-to-event data increase prediction accuracy by an average of 20% compared to a recent summary statistic approach (LDpred-funct). As sample sizes increase, incorporating time-to-event data should be commonplace, improving case-control studies by using richer information about the disease process.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cancer PRSweb - an Online Repository with Polygenic Risk Scores (PRS) for Major Cancer Traits and Their Phenome-wide Exploration in Two Independent Biobanks 97%
- Assessing digital phenotyping to enhance genetic studies of human diseases 95%
- The contribution of coding variants to the heritability of multiple cancer types using UK Biobank whole-exome sequencing data 95%
Similar papers in this journal
- Constructing germline research cohorts from the discarded reads of clinical tumor sequences 96%
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 94%
- Sequence dependencies and mutation rates of localized mutational processes in cancer 94%
Similar papers in this journal
- Set-based rare variant association tests for biobank scale sequencing data sets 94%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 94%
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 94%
Similar papers in this journal
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 95%
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 95%
- Cross-Cancer Evaluation of Polygenic Risk Scores for 17 Cancer Types in Two Large Cohorts 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.