Large language models identify causal genes in complex trait GWAS
Shringarpure, S. S.; Wang, W.; Karagounis, S.; Wang, X.; Reisetter, A. C.; Auton, A.; Khan, A. A.
Show abstract
Identifying causal genes at genome-wide association study (GWAS) loci remains a major challenge. Literature evidence for disease-gene co-occurrence, whether through automated approaches or human expert annotation, is one way of nominating causal genes at GWAS loci. However, current automated approaches are limited in accuracy and generalizability, and expert annotation is not scalable to hundreds of thousands of significant findings. Here, we demonstrate that large language models (LLMs) can accurately prioritize likely causal genes at GWAS loci. We rigorously evaluated several widely available general-purpose LLMs using a benchmark of high-confidence causal gene annotations, including a novel set of 26 previously unpublished GWAS. Our results show that LLMs outperform current state-of-the-art methods and substantially augment their performance. These findings establish LLMs as a powerful, efficient, and scalable approach to causal gene discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups 95%
- Best practices for multi-ancestry, meta-analytic transcriptome-wide association studies: lessons from the Global Biobank Meta-analysis Initiative 95%
- A Unifying Statistical Framework to Discover Disease Genes from GWAS 94%
Similar papers in this journal
Similar papers in this journal
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 95%
- Optimizing and benchmarking polygenic risk scores with GWAS summary statistics 94%
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 94%
Similar papers in this journal
- Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores 96%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 95%
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.