Performance of Information Theory Derived Semantic Similarity Algorithms for Differential Diagnosis and Clustering
COLEMAN, B. D.; Danis, D.; Reese, J.; Robinson, P. N.
Show abstract
Semantic similarity analysis with Human Phenotype Ontology (HPO) enables fuzzy, specificity weighted comparisons of clinical manifestations of individuals and diseases and can be used to support differential diagnostics or to stratify cohorts. Many methods have been proposed to calculate semantic similarity for various applications, including the Phenomizer, which calculates the average best match over all terms in the query and disease, and set-based methods ranging from the Jaccard Intersection to methods that leverage the conditional information content to calculate similarity. However, these methods have not been described under a single mathematical model or robustly compared using a comprehensive data set. Here, we describe several semantic similarity algorithms using derivations based on information theory, propose three of our own variations to these models, and compare the performance of each approach for differential diagnostic ranking and phenotypic clustering. We find that Phenomizer performs better when diseases are ranked by similarity alone, without generating p-values. Additionally, non-normalized algorithms that use conditional information perform similarly to Phenomizer for differential diagnosis. In contrast, normalized algorithms perform best when clustering cohorts. AvailabilityData is available through the Phenopacket-Store (https://github.com/monarch-initiative/phenopacket-store). Algorithms are implemented in the Python package SetSim (https://github.com/P2GX/setsim).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Consensus clustering applied to multi-omic disease subtyping 95%
- Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework 95%
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.