Back

What limits local ancestry inference at low divergence: a feasibility threshold, a metric that conceals failure, and a deficit of input more than architecture

Tian, Q.

2026-08-03 bioinformatics
10.64898/2026.07.30.741148 bioRxiv
Show abstract

Local ancestry inference assigns each position along an admixed chromosome to a source population, underpinning admixture mapping, ancestry-specific association testing and admixture dating. Validation is almost exclusively on continentally divergent sources (Hudsons FST{approx} 0.1) and coalescent simulations; we examine both restrictions. Across FST from 0.0022 to 0.243 we compare five methods -- two likelihood baselines, RFMix, FLARE and a dilated convolutional network -- on identical sites with exact ground truth, and on 11 real 1000 Genomes pairs. Three findings follow. First, a feasibility floor: at FST= 0.0022 no method exceeds 0.575, and CHB/CHS at FST= 0.00042 yields at best 0.551. Pairs motivating fine-scale analysis, such as northern versus southern Han, fall below it. Second, per-site accuracy conceals a failure of tract structure: the most accurate method per site produces 78.8x too many tracts, implying an admixture time 61.2x too old, which Viterbi decoding removes at no cost to accuracy (+0.0002). Third, the simulated lead does not survive real data, and the deficit is one of input more than of the architectures we varied: attention, state-space layers, capacity, objective and self-supervised pretraining each move accuracy by at most 0.006, while supplying the haplotype information the released tools receive recovers +0.031 on 8 of 8 pairs below FST = 0.04 and nothing above it -- necessary but not sufficient, since the network still trails on 10 of 11 pairs. Two quantities usually held fixed matter more than architecture: the statistic summarising reference matching, and reference panel size, which no method is near saturating.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.