Back

LORE: A Literature Semantics Framework for Evidenced Disease-Gene Pathogenicity Prediction at Scale

Li, P.-H.; Sun, Y.-Y.; Juan, H.-F.; Chen, C.-Y.; Tsai, H.-K.; Huang, J.-H.

2024-08-11 health informatics
10.1101/2024.08.10.24311801 medRxiv
Show abstract

Effective utilization of academic literature is crucial for Machine Reading Comprehension to generate actionable scientific knowledge for wide real-world applications. Recently, Large Language Models (LLMs) have emerged as a powerful tool for distilling knowledge from scientific articles, but they struggle with the issues of reliability and verifiability. Here, we propose LORE, a novel unsupervised two-stage reading methodology with LLM that models literature as a knowledge graph of verifiable factual statements and, in turn, as semantic embeddings in Euclidean space. Applied to PubMed abstracts for large-scale understanding of disease-gene relationships, LORE captures essential information of gene pathogenicity. Furthermore, we demonstrate that modeling a latent pathogenic flow in the semantic embedding with supervision from the ClinVar database leads to a 90% mean average precision in identifying relevant genes across 2,097 diseases. Finally, we have created a disease-gene relation knowledge graph with predicted pathogenicity scores, 200 times larger than the ClinVar database.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.