Phenotype-first patient matching with SimPheny identifies diagnostic candidates beyond curated gene associations
Cooperstein, I. B.; Ward, A.; Kobren, S. N.; Lebleu, E.; Moore, B.; Spillmann, R. C.; Shashi, V.; Undiagnosed Diseases Network, ; Marth, G. T.
Show abstract
Diagnostic tools for rare diseases typically rely on curated gene-phenotype associations and static disease models, limiting their effectiveness in cases with atypical presentations or previously uncharacterized disorders. To address these limitations, we present SimPheny, a phenotype-first algorithm for gene prioritization that operates independently of documented gene-phenotype associations. SimPheny identifies phenotypically similar diagnosed patients by comparing an undiagnosed patients disease presentation to a reference cohort of diagnosed cases, and returns gene hypotheses by matching the undiagnosed patients candidate gene list to the causative genes of similar patients using a statistical scoring model. Evaluated in diagnosed probands from the Undiagnosed Diseases Network (UDN) with the true diagnostic gene blinded, SimPheny consistently ranked the diagnostic gene among the top five candidates, outperforming existing tools, particularly for genes with limited gene-phenotype association data. When applied to previously unsolved UDN cases, clinical review confirmed that SimPhenys high-confidence causative gene predictions were diagnostic in nearly half of the analyzed cases. As the size of the diagnosed reference cohort increases, SimPhenys diagnostic reach expands without sacrificing ranking performance. By leveraging real patient data rather than curated models, SimPheny provides a generalizable, scalable framework for improving diagnostic yield in rare disease cohorts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COBT: A gene-based rare variant burden test for case-only study designs using aggregated genotypes from public reference cohorts. 94%
- Defining and Reducing Variant Classification Disparities 94%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
Similar papers in this journal
- Discovering Monogenic Patients with a Confirmed Molecular Diagnosis in Millions of Clinical Notes with MonoMiner 96%
- A Framework for Automated Gene Selection in Genomic Screening 94%
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 94%
Similar papers in this journal
Similar papers in this journal
- Inferring compound heterozygosity from large-scale exome sequencing data 95%
- Set-based rare variant association tests for biobank scale sequencing data sets 94%
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.