PhenoExam: an R package and Web application for the examination of phenotypes linked to genes and gene sets
Cisterna Garcia, A.; Gonzalez-Vidal, A.; Ruiz Villa, D.; Ortiz Murillo, J.; Gomez-Pascual, A.; Chen, Z.; Nalls, M. A.; Faghri, F.; Hardy, J.; Diez, I.; Maietta, P.; Alvarez, S.; Ryten, M.; Botia, J. A.
Show abstract
Gene set based phenotype enrichment analysis (detecting phenotypic terms that emerge as significant in a set of genes) can improve the rate of genetic diagnoses amongst other research purposes. To facilitate diverse phenotype analysis, we developed PhenoExam, a freely available R package for tool developers and a web interface for users, which performs: (1) phenotype and disease enrichment analysis on a gene set; (2) measures statistically significant phenotype similarities between gene sets and (3) detects significant differential phenotypes or disease terms across different databases. PhenoExam achieves these tasks by integrating databases or resources such as the HPO, MGD, CRISPRbrain, CTD, ClinGen, CGI, OrphaNET, UniProt, PsyGeNET, and Genomics England Panel App. PhenoExam accepts both human and mouse genes as input. We developed PhenoExam to assist a variety of users, including clinicians, computational biologists and geneticists. It can be used to support the validation of new gene-to-disease discoveries, and in the detection of differential phenotypes between two gene sets (a phenotype linked to one of the gene set but no to the other) that are useful for differential diagnosis and to improve genetic panels. We validated PhenoExam performance through simulations and its application to real cases. We demonstrate that PhenoExam is effective in distinguishing gene sets or Mendelian diseases with very similar phenotypes through projecting the disease-causing genes into their annotation-based phenotypic spaces. We also tested the tool with early onset Parkinsons disease and dystonia genes, to show phenotype-level similarities but also potentially interesting differences. More specifically, we used PhenoExam to validate computationally predicted new genes potentially associated with epilepsy. Therefore, PhenoExam effectively discovers links between phenotypic terms across annotation databases through effective integration. The R package is available at https://github.com/alexcis95/PhenoExam and the Web tool is accessible at https://snca.atica.um.es/PhenoExamWeb/.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ranking of cell clusters in a single-cell RNA-sequencing analysis framework using prior knowledge 95%
- GenOtoScope: Towards automating ACMG classification of variants associated with congenital hearing loss 95%
- GeneCOCOA: Detecting context-specific functions of individual genes using co-expression data 95%
Similar papers in this journal
- Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework 95%
- Random Walk with Restart on multilayer networks: from node prioritisation to supervised link prediction and beyond 95%
- GenEpi: Gene-based Epistasis Discovery Using Machine Learning 94%
Similar papers in this journal
Similar papers in this journal
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 95%
- Assessing the performance of genome-wide association studies for predicting disease risk 94%
- Phenogenon: Gene to Phenotype Associations for Rare Genetic Diseases 94%
Similar papers in this journal
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 96%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
- Integrative network analysis interweaves the missing links in cardiomyopathy diseasome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.