GENPHIRE: Enhancing Disease Risk Prediction Using Large Language Model
Yao, D.; Liu, C.; Yan, S.; Zhang, J.; Sun, Y. V.; Qin, Z. S.
Show abstract
BackgroundEstimating an individuals liability to a disease is a fundamental problem in genome research. By exploiting findings from genome-wide association studies (GWASs), many powerful polygenic risk scores (PRSs) have been developed to predict disease risk based on genetic profile. Despite much success, the performance of PRS models is hindered by its inability to capture complex, nonlinear effects and interactions among variants. ResultsIn this study, we introduce GENPHIRE or Genetic-Phenotypic Representation, a novel machine learning framework designed for disease risk prediction. The central idea in GENPHIRE is to translate an individuals genotype profile to a "sentence" consist of basic clinical information together with an ordered list of top phenotypes for which the individual is found to have elevated number of risk alleles. After translation, the sentence is converted to an embedded vector by an pre-trained large language model (LLM) to assess its disease risk. We have tested GENPHIRE using UK Biobank data across a broad range of diseases and found it outperforms state-of-the-art PRS models more than 80% of the time. ConclusionsOur results demonstrated that LLM-derived embeddings can be leveraged for disease risk prediction when an individuals genotype profile is effectively represented. Our findings highlight a promising alternative strategy that complements existing PRS approaches.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep transfer learning provides a Pareto improvement for multi-ancestral clinico-genomic prediction of diseases 94%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
- scGRNom: a computational pipeline of integrative multi-omics analyses for predicting cell-type disease genes and regulatory networks 93%
Similar papers in this journal
- A Scalable Framework for Identifying Allelic Series from Summary Statistics 96%
- The Construction of Multi-ethnic Polygenic Risk Score using Transfer Learning 95%
- Cancer PRSweb - an Online Repository with Polygenic Risk Scores (PRS) for Major Cancer Traits and Their Phenome-wide Exploration in Two Independent Biobanks 95%
Similar papers in this journal
- SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification 96%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 95%
- Deep representation learning for clustering longitudinal survival data from electronic health records 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.