Back

Phenotype-first patient matching with SimPheny identifies diagnostic candidates beyond curated gene associations

Cooperstein, I. B.; Ward, A.; Kobren, S. N.; Lebleu, E.; Moore, B.; Spillmann, R. C.; Shashi, V.; Undiagnosed Diseases Network, ; Marth, G. T.

2026-01-17 genetic and genomic medicine
10.64898/2026.01.15.26344236 medRxiv
Show abstract

Diagnostic tools for rare diseases typically rely on curated gene-phenotype associations and static disease models, limiting their effectiveness in cases with atypical presentations or previously uncharacterized disorders. To address these limitations, we present SimPheny, a phenotype-first algorithm for gene prioritization that operates independently of documented gene-phenotype associations. SimPheny identifies phenotypically similar diagnosed patients by comparing an undiagnosed patients disease presentation to a reference cohort of diagnosed cases, and returns gene hypotheses by matching the undiagnosed patients candidate gene list to the causative genes of similar patients using a statistical scoring model. Evaluated in diagnosed probands from the Undiagnosed Diseases Network (UDN) with the true diagnostic gene blinded, SimPheny consistently ranked the diagnostic gene among the top five candidates, outperforming existing tools, particularly for genes with limited gene-phenotype association data. When applied to previously unsolved UDN cases, clinical review confirmed that SimPhenys high-confidence causative gene predictions were diagnostic in nearly half of the analyzed cases. As the size of the diagnosed reference cohort increases, SimPhenys diagnostic reach expands without sacrificing ranking performance. By leveraging real patient data rather than curated models, SimPheny provides a generalizable, scalable framework for improving diagnostic yield in rare disease cohorts.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.