Back

Unsupervised Variant Clustering Identifies Genetic Subtypes of Disease

Mizikovsky, D.; Wray, N. R.; Shah, S.; Palpant, N.

2025-07-04 genomics
10.1101/2025.07.01.662666 bioRxiv
Show abstract

Current approaches to identify disease subtypes rely on prior knowledge or external phenotypic references that limit their generalisability across traits. This study presents an unsupervised, phenotype- free framework that analyses the genetic co-occurrence of variants across individuals to identify variant clusters underpinning disease heterogeneity. The approach is developed and validated using simulated combinations of real phenotypes then applied to dissect disease subtypes of type 2 diabetes and asthma. The approach captures clinically established mechanistic profiles in type 2 diabetes and asthma without requiring any reference phenotype data and enables identification of plasma proteins for novel subtype- specific biomarkers and drug targets. Lastly, we demonstrate that variants with conflicting effects on disease-relevant traits are resolved into distinct clusters, identifying subtypes with opposing phenotypic profiles that are lost by the aggregated genetic risk. Co-occurrence-based clustering provides a scalable strategy for understanding disease heterogeneity across diverse populations by identifying biologically meaningful subtype structure from genotype data alone.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.