New machine learning method identifies subtle fine-scale genetic stratification in diverse populations
Qin, X.; Jia, P.
Show abstract
Fine-scale genetic structure impacts genetic risk predictions and furthers the understanding of the demography of populations. Current approaches (e.g., PCA, DAPC, t-SNE, and UMAP) either produce coarse and ambiguous cluster divisions or fail to preserve the correct genetic distance between populations. We proposed a new machine learning algorithm named ALFDA. ALFDA considers both local and global genetic affinity between individuals and also preserves the multimodal structure within populations. ALFDA outperformed the existing approaches in identifying fine-scale genetic structure and in retaining population geogenetic distance, providing a valuable tool for geographic ancestry inference as well as correction for spatial stratification in population health studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Kernel Local Fisher Discriminant Analysis of Principal Components (KLFDAPC) significantly improves the accuracy of predicting geographic origin of individuals 96%
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 93%
- SSMD: A semi-supervised approach for a robust cell type identification and deconvolution of mouse transcriptomics data 93%
Similar papers in this journal
Similar papers in this journal
- A spectral framework to map QTLs affecting joint differential networks of gene co-expression 94%
- HiCImpute: A Bayesian Hierarchical Model for Identifying Structural Zeros and Enhancing Single Cell Hi-C Data. 94%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 94%
Similar papers in this journal
- CNETML: Maximum likelihood inference of phylogeny from copy number profiles of spatio-temporal samples 93%
- A comprehensive benchmark of graph-based genetic variant genotyping algorithms on plant genomes for creating an accurate ensemble pipeline 93%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.