Leveraging Large-Scale Biobanks for Therapeutic Target Discovery
Ferolito, B. R.; Dashti, H.; Giambartolomei, C.; Peloso, G. M.; Golden, D. J.; Gravel-Pucillo, K.; Rasooly, D.; Horimoto, A. R. V. R.; Matty, R.; Gaziano, L.; Liu, Y.; Smit, I. A.; Zdrazil, B.; Tsepilov, Y.; Costa, L.; Kosik, N. M.; Huffman, J. E.; Tartaglia, G. G.; Bini, G.; Proietti, G.; Ioannidis, H.; Hunter, F.; Hemani, G.; Butterworth, A. S.; Di Angelantonio, E.; Langenberg, C.; Ghoussaini, M.; Leach, A. R.; Liao, K. P.; Damrauer, S.; Selva, L. E.; Whitbourne, S.; Tsao, P. S.; Moser, J.; Gaunt, T.; Cai, T.; Whittaker, J. C.; Program, M. V.; Casas, J. P.; Muralidhar, S.; Gaziano, J. M.; Ch
Show abstract
Large biobanks, including the Million Veteran Program (MVP), the UK Biobank, and FinnGen, provide genetic association results for more than 1,000,000 individuals for hundreds of phenotypes. To select targets for pharmaceutical development, as well as to improve the understanding of existing targets, we harmonized these studies, and performed two-sample Mendelian Randomization (MR) on 2,003 phenotypes using genetic variants associated with gene expression (derived from GTEx and eQTLGen) and plasma protein levels (derived from ARIC, Fenland, and DeCODE) as proxies of target modulation. We found 69,669 gene-trait pairs with evidence (p [≤] 1.6 x 10-9) for causal effects. From the selected gene-trait pairs, we observed 6,447 genes with strong causal evidence for at least one of 2,003 investigated traits. As expected, being identified as a gene-trait pair in our approach was significantly associated with higher odds of being an approved drug target and indication. We were able to rediscover 9% of approved drug targets in ChEMBL 34. Moreover, identified gene-traits were significantly associated with higher odds of being previously described as a gene-trait pair in OMIM, ClinVar, mouse knock-out data, and rare variant burden studies. To enhance the translational potential of the resource, we developed a predictive ranking model trained using approved drug targets described in ChEMBL 34 as well as several different biological annotations. This model was able to accurately predict the odds of a particular significant MR result being developed into an approved drug and its clinical indication (precision-recall AUC 0.79). We make our results publicly available in CIPHER.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 95%
- A novel preclinical secondary pharmacology resource illuminates target-adverse drug reaction associations of marketed drugs 95%
- Machine Learning Identifies Novel Candidates for DrugRepurposing in Alzheimer's Disease 94%
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
- scDrugPrio: A framework for the analysis of single-cell transcriptomics to address multiple problems in precision medicine in immune-mediated inflammatory diseases 93%
- Mendelian gene identification through mouse embryo viability screening 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.