Domain-specific embeddings uncover latent genetics knowledge
Ho, S. S.; Mills, R. E.
Show abstract
The inundating rate of scientific publishing means every researcher will miss new discoveries from overwhelming saturation. To address this limitation, we employ natural language processing to overcome human limitations in reading, curation, and knowledge synthesis, with domain-specific applications to genetics and genomics. We construct a corpus of 3.5 million normalized genetics and genomics abstracts and implement both semantic and network-based embedding models. Our methods not only capture broad biological concepts and relationships but also predict complex phenomena such as gene expression. Through a rigorous temporal validation framework, we demonstrate that our embeddings successfully predict gene-disease associations, cancer driver genes, and experimentally-verified protein interactions years before their formal documentation in literature. Additionally, our embeddings successfully predict experimentally verified gene-gene interactions absent from the literature. These findings demonstrate that substantial undiscovered knowledge exists within the collective scientific literature and that computational approaches can accelerate biological discovery by identifying hidden connections across the fragmented landscape of scientific publishing.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 96%
- Context-aware deconvolution of cell-cell communication with Tensor-cell2cell 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
Similar papers in this journal
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 94%
- Conserved epigenetic regulatory logic infers genes governing cell identity 94%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.