Identification of disease mechanisms and novel disease genes using clinical concept embeddings learned from massive amounts of biomedical data.
Bugrim, A.
Show abstract
MotivationKnowledge of relationship and similarity among human diseases can be leveraged in many biomedical applications such as drug repositioning, biomarker discovery, differential diagnostics, and understanding of disease mechanisms. Recently developed cui2vec resource provides embeddings of approximately 109 thousand biomedical terms and allows computing novel measures of disease similarity, directly related to patterns in real-world data. We investigate whether disease embeddings from cui2vec can be utilized to identify functional relations among diseases, to uncover their molecular mechanisms, and to generate hypotheses about novel gene-disease associations and potential drug targets. Methods and resultsWe focus on a subset of 3,568 cui2vec terms corresponding to human diseases annotated in DisGeNET database. Disease-disease distance matrix is computed for this set of diseases based on their embedding vectors. Clustering of this matrix reveals a well-defined structure with good correspondence between disease clusters and the top MeSH disease categories. Using pulmonary embolism as an example we show how disease clustering is related to known mechanistic relations among diseases. Next, we combine disease embeddings with annotated gene-disease associations from DisGeNET to generate joint gene-disease co-embeddings. From these we identify molecular pathways most characteristic for each disease group and show that they are highly relevant to known disease physiology. Finally, we leverage disease similarity to generate and rank hypothesis for gene-disease associations and demonstrate that this method generates highly accurate results and can suggest relevant drug targets. ConclusionsWe show that combination of disease embeddings learned from massive amounts of biomedical records with curated data on gene-disease associations can reliably reveal groups of functionally related diseases and their molecular mechanisms and predict novel gene-disease associations. Importantly, our analysis does not require knowledge of associated genes for every disease to identify patterns in the embedding space, therefore it can be used to suggest mechanisms for conditions that have not been functionally understood. In this respect our analysis can be applied to identify potential markers and drug targets for poorly characterized orphan and rare diseases. It can also reveal unexpected novel connections among diseases and between diseases and molecular pathways.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Predicting Gene Disease Associations With Knowledge Graph Embeddings For Diseases With Curtailed Information 96%
- Discovering Governing Equations of Biological Systems through Representation Learning and Sparse Model Discovery 93%
- Prognostic importance of splicing-triggered aberrations of protein complex interfaces in cancer 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.