Why most Principal Component Analyses (PCA) in population genetic studies are wrong
Elhaik, E.
Show abstract
Principal Component Analysis (PCA) is a multivariate analysis that allows reduction of the complexity of datasets while preserving data covariance and visualizing the information on colorful scatterplots, ideally with only a minimal loss of information. PCA applications are extensively used as the foremost analyses in population genetics and related fields (e.g., animal and plant or medical genetics), implemented in well-cited packages like EIGENSOFT and PLINK. PCA outcomes are used to shape study design, identify, and characterize individuals and populations, and draw historical and ethnobiological conclusions on origins, evolution, dispersion, and relatedness. The replicability crisis in science has prompted us to evaluate whether PCA results are reliable, robust, and replicable. We employed an intuitive color-based model alongside human population data for eleven common test cases. We demonstrate that PCA results are artifacts of the data and that they can be easily manipulated to generate desired outcomes. PCA results may not be reliable, robust, or replicable as the field assumes. Our findings raise concerns about the validity of results reported in the literature of population genetics and related fields that place a disproportionate reliance upon PCA outcomes and the insights derived from them. We conclude that PCA may have a biasing role in genetic investigations. An alternative mixed-admixture population genetic model is discussed.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FSTruct: An Fst-based tool for measuring ancestry variation in inference of population structure 95%
- Benchmarking Imputed Low Coverage Genomes in a Human Population Genetics Context 95%
- Commonly used Hardy-Weinberg equilibrium filtering schemes impact population structure inferences using RADseq data 94%
Similar papers in this journal
- Estimating allele frequencies, ancestry proportions and genotype likelihoods in the presence of mapping bias 94%
- Limited population structure but signals of recent selection in introduced African Fig Fly (Zaprionus indianus) in North America 92%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.