Back

MULTICLUST - Fast multinomial clustering of multiallelic genotypes to infer genetic population structure

Sethuraman, A.; Chen, W.-C.; Wanjiku, M.; Dorman, K.

2025-08-07 evolutionary biology
10.1101/2025.08.06.668969 bioRxiv
Show abstract

Identifying population structure from multilocus genotype data is key to down-stream genetic analyses, including analysis of genome-wide association (GWAS), genetic genealogy and phylogenetics, in the fields of conservation, forensics, evolution and more. While inference of population structure has been dominated by Bayesian methods, maximum likelihood methods have some benefits in terms of reliability, consistency and efficiency. Here we extend the methods of Tang et al. [53] and Alexander et al. [2] to handle multi-allelic (e.g., SNP, STR, allozyme) and polyploid loci with missing data to infer genetic admixture proportions and subpopulation allele frequencies. Comparative analyses of our method, MULTICLUST, and STRUCTURE [42] on both simulated and empirical data indicate comparable, fast, reproducible, and accurate estimates of population admixture proportions and allele frequencies using MULTICLUST.MULTICLUST is implemented in the C programming language and is publicly accessible via www.github.com/arunsethuraman/multiclust.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.