What limits local ancestry inference at low divergence: a feasibility threshold, a metric that conceals failure, and a deficit of input more than architecture
Tian, Q.
Show abstract
Local ancestry inference assigns each position along an admixed chromosome to a source population, underpinning admixture mapping, ancestry-specific association testing and admixture dating. Validation is almost exclusively on continentally divergent sources (Hudsons FST{approx} 0.1) and coalescent simulations; we examine both restrictions. Across FST from 0.0022 to 0.243 we compare five methods -- two likelihood baselines, RFMix, FLARE and a dilated convolutional network -- on identical sites with exact ground truth, and on 11 real 1000 Genomes pairs. Three findings follow. First, a feasibility floor: at FST= 0.0022 no method exceeds 0.575, and CHB/CHS at FST= 0.00042 yields at best 0.551. Pairs motivating fine-scale analysis, such as northern versus southern Han, fall below it. Second, per-site accuracy conceals a failure of tract structure: the most accurate method per site produces 78.8x too many tracts, implying an admixture time 61.2x too old, which Viterbi decoding removes at no cost to accuracy (+0.0002). Third, the simulated lead does not survive real data, and the deficit is one of input more than of the architectures we varied: attention, state-space layers, capacity, objective and self-supervised pretraining each move accuracy by at most 0.006, while supplying the haplotype information the released tools receive recovers +0.031 on 8 of 8 pairs below FST = 0.04 and nothing above it -- necessary but not sufficient, since the network still trails on 10 of 11 pairs. Two quantities usually held fixed matter more than architecture: the statistic summarising reference matching, and reference panel size, which no method is near saturating.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaGLIMPSE: Meta Imputation of Low Coverage Sequencing Data for Modern and Ancient Genomes 95%
- Simultaneous inference of parental admixture proportions and admixture times from unphased local ancestry calls 94%
- A genealogy-based approach for revealing ancestry-specific structures in admixed populations 93%
Similar papers in this journal
- Genotyping common, large structural variations in 5,202 genomes using pangenomes, the Giraffe mapper, and the vg toolkit 91%
- A lineage-resolved molecular atlas of C. elegans embryogenesis at single cell resolution 91%
- Lineage tracing on transcriptional landscapes links state to fate during differentiation 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.