Assessing the Reliability of LLM-Generated Phenotype-Genotype Associations Through External Validation
Sun, C.; Xin, Y.; Zeng, S.; Sunthankar, S. D.; Su, W.-C.; Lynn, J.; Mundo, S.; Babanejad, M.; Feng, Q.; Wei, W.-Q.
Show abstract
BackgroundPhenotype-genotype associations underpin precision medicine by enabling disease prevention, early diagnosis, risk stratification, therapeutic target discovery, and personalized treatment. However, the rapid growth of scientific evidence has made manual curation of these associations increasingly labor-intensive, time-consuming, and incomplete. Large Language Models (LLMs) offer a potential path to scalable genomic generation and synthesis of this knowledge, but their ability to accurately identify phenotype-genotype associations and the extent to which these outputs are supported by established genomic knowledge bases remain unclear. Materials and MethodsFour LLMs, Claude Sonnet 4.6, DeepSeek V4 Flash, Gemini 3 Flash Preview, and GPT-5.5, were benchmarked on six zero-shot task categories covering forward and reverse phenotype-gene and phenotype-SNP generation. A total of 4,196 associations were identified from curated inputs and evaluated through a multistage external verification pipeline comprising phenotype normalization, ontology mapping, genomic identifier validation against Ensembl, and evidence verification using both the GWAS Catalog and OMIM. Associations were assigned a fused evidence level of strong, moderate, weak, or none. ResultsOverall, 74.19% of generated associations were matched to at least one external genomic knowledge base; 9.15% received strong support and 54.46% moderate support. Phenotype-gene associations were more verifiable than phenotype-SNP associations (strong or moderate: 67.19% vs 54.06%). Among existing associations, Claude Sonnet 4.6 achieved the highest overall strong or moderate rate (69.2%), followed by GPT-5.5 (65.1%), DeepSeek V4 Flash (61.7%), and Gemini 3 Flash Preview (56.9%). ConclusionLLMs can support scalable generation of candidate phenotype-genotype associations. Performance varied substantially by relation type and was lower for SNP-level and rare disease associations, highlighting both the limitations of current genomic resources and the need for rigorous validation pipelines.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 95%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 92%
- COBT: A gene-based rare variant burden test for case-only study designs using aggregated genotypes from public reference cohorts. 92%
Similar papers in this journal
Similar papers in this journal
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 95%
- MARRVEL-MCP enables natural language variant interpretation through autonomous workflow construction 94%
- The Phenotype-Genotype Reference Map: Improving biobank data science through replication. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.