Back

AyurPhenoClusters define common molecular roots for rare diseases and uncover ciliary dysfunctions in syndromic conditions

Joshi, A.; Jangir, D.; Sharma, A.; Anand, T.; Verma, H.; Yadav, M.; Rangani, N.; Joshi, P.; Singh, R. P.; Kumar, S.; Girdhar, S.; Sharma, R.; Kumar, A.; Dey, L.; Mukerji, M.

2024-09-19 genetics
10.1101/2024.09.13.612844 bioRxiv
Show abstract

Managing rare genetic diseases with an organ-centric approach poses challenges in linking genotypes to phenotypes. Ayurveda, however, takes a multisystem perspective, assessing diseases through kinetic (Vata:V), metabolic (Pitta:P), and structural (Kapha:K) dimensions, each with distinct phenotypic and molecular correlates. This study explores rare diseases from a systems perspective by integrating Ayurveda and unifying terminologies from both disciplines using Human Phenotype Ontology (HPO). Experts categorized 10,610 HPO terms into Ayurvedic phenotypic groups (V/P/K) and applied the Expectation Maximization (EM) algorithm to cluster 12,678 diseases. This yielded six distinct clusters, termed "AyurPhenoClusters," with 2,814 diseases uniquely classified and enriched for V/P/K phenotypes. Functional annotation of these, highlighted key biological processes: (i) embryogenesis and skeletal morphogenesis, (ii) endocrine and ciliary functions, (iii) DNA damage response and cell cycle regulation, (iv) inflammation and immune response, (v) immune function, hemopoiesis, and telomere aging, and (vi) small molecule metabolism and transport. Notably, the K predominant cluster had the highest ciliary gene enrichment (43%), followed by the V predominant cluster (16%), suggesting potential ciliopathies in V cluster. This systems-based approach enhances the understanding of rare diseases by bridging Ayurveda with modern medicine, for improved diagnostics and therapeutics.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.