GraphPop: graph-native computation decouples population genomics complexity from sample count
Estaji, E.; Zhao, S.-W.; Chen, Z.-Y.; Nie, S.; Mao, J.-F.
Show abstract
Matrix-based population genomics tools scale as O(V x N ), re-reading the full genotype matrix for every analysis. Here we present GraphPop, a graph database engine that reduces summary statistic complexity to O(V x K) where K is population count--independent of sample count--by computing on pre-aggregated allele-count arrays stored as graph node properties. The same architecture enables annotation-conditioned queries via edge traversal, persistent analytical records, and multi-statistic composition. Applied to rice 3K (29.6M SNPs, 3,024 accessions) and human 1000 Genomes (3,202 samples, 22 autosomes), GraphPop reveals that all 12 rice subpopulations show{pi} N /{pi}S > 1.0, uncovers opposite consequence-level Fst regimes between species, and identifies KCNE1 as a candidate pre-Out-of-Africa sweep via convergence of five stored statistics. GraphPop achieves 146-327x query-time speedup for pre-aggregated statistics and 63-179x for bit-packed haplotype computation (iHS, XP-EHH, nSL), at constant[~] 160 MB procedure working memory. This complexity reduction makes systematic, annotation-integrated population genomics practical at scale.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Efficient detection and characterization of targets of natural selection using transfer learning 95%
- soibean: High-resolution Taxonomic Identification of Ancient Environmental DNA Using Mitochondrial Pangenome Graphs 94%
- Tensor decomposition based feature extraction and classification to detect natural selection from genomic data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.