Transformers Enhance the Predictive Power of Network Medicine
Spector, J.; Aldana, A.; Sebek, M.; Ehlert, J.; De Frondeville, C.; Ghiassian, S. D.; Barabasi, A.-L.
Show abstract
BackgroundSelf-attention mechanisms and token embeddings behind transformers allow the extraction of complex patterns from large datasets, and enhance the predictive power over traditional machine learning models. Yet, being trained to make predictions about individual cells or genes, it is not clear if transformers can learn the inherent interaction patterns between genes, ultimately responsible for their mechanism of action. We use Geneformer, pretrained on single-cell transcriptomes, to ask if transformers can implicitly capture molecular dependencies, including protein-protein interactions (PPIs), allowing us to explore the use of transformers to improve network medicine tasks such as disease gene identification and drug repurposing. MethodsWe extracted the cosine similarity of gene embeddings and the attention weights contained in Geneformer, allowing us to test if these weights capture experimentally validated protein interactions. Using dilated cardiomyopathy as a case study, we evaluated the effectiveness of the resulting weighted networks in disease module detection and drug repurposing. ResultsWe found that Geneformer displays awareness of experimentally documented Protein-Protein Interactions, exhibiting higher cosine similarity and attention weights for gene pairs with physical interactions. Weighting PPI networks with the cosine similarity and attention weights improved the detection of disease-associated genes and the accuracy of drug repurposing predictions for dilated cardiomyopathy, surpassing the accuracy of unweighted networks. Finally, we find that combining attention weights and cosine similarities with ranking methods enhances drug candidate prioritization for drug repurposing. ConclusionsWe find that transformers, by implicitly learning the interactions between genes, offer a promising pathway for advancing medicine and drug discovery when integrated with the graph theoretic algorithms used in network medicine.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepDRIM: a deep neural network to reconstruct cell-type-specific gene regulatory network using single-cell RNA-seq data 95%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 94%
- KGETCDA: an efficient representation learning framework based on knowledge graph encoder from transformer for predicting circRNA-disease associations 94%
Similar papers in this journal
Similar papers in this journal
- Inferring latent temporal progression and regulatory networks from cross-sectional transcriptomic data of cancer samples 95%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 94%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 94%
Similar papers in this journal
- Genome-wide discovery of hidden genes mediating known drug-disease association using KDDANet 93%
- Methylation risk scores are associated with a collection of phenotypes within electronic health record systems 91%
- Geometric Network Analysis Provides Prognostic Information in Patients with High Grade Serous Carcinoma of the Ovary Treated with Immune Checkpoint Inhibitors 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.