Back

GENPHIRE: Enhancing Disease Risk Prediction Using Large Language Model

Yao, D.; Liu, C.; Yan, S.; Zhang, J.; Sun, Y. V.; Qin, Z. S.

2025-12-04 genetic and genomic medicine
10.64898/2025.12.03.25341576 medRxiv
Show abstract

BackgroundEstimating an individuals liability to a disease is a fundamental problem in genome research. By exploiting findings from genome-wide association studies (GWASs), many powerful polygenic risk scores (PRSs) have been developed to predict disease risk based on genetic profile. Despite much success, the performance of PRS models is hindered by its inability to capture complex, nonlinear effects and interactions among variants. ResultsIn this study, we introduce GENPHIRE or Genetic-Phenotypic Representation, a novel machine learning framework designed for disease risk prediction. The central idea in GENPHIRE is to translate an individuals genotype profile to a "sentence" consist of basic clinical information together with an ordered list of top phenotypes for which the individual is found to have elevated number of risk alleles. After translation, the sentence is converted to an embedded vector by an pre-trained large language model (LLM) to assess its disease risk. We have tested GENPHIRE using UK Biobank data across a broad range of diseases and found it outperforms state-of-the-art PRS models more than 80% of the time. ConclusionsOur results demonstrated that LLM-derived embeddings can be leveraged for disease risk prediction when an individuals genotype profile is effectively represented. Our findings highlight a promising alternative strategy that complements existing PRS approaches.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.