Back

A Unified Deep Learning-Based Framework for Reference-Based and Reference-Free Local Ancestry Inference

Diem, T.

2026-08-06 genomics
10.64898/2026.08.01.742246 bioRxiv
Show abstract

Local ancestry inference (LAI) identifies the ancestral origin of genomic segments within admixed individuals and is an important tool for population genetics and disease association studies. Existing LAI methods rely on reference panels composed of individuals from ancestral populations, limiting their applicability when such panels are unavailable or poorly characterized. We present Optional Reference Inference Ancestry Network (ORIAN), a software package containing two complementary algorithms for local ancestry inference. The first is a reference-based approach that combines neural network predictions with a hidden Markov model to produce probabilistic ancestry assignments. The second is a reference-free method that introduces an iterative framework in which ad-mixed individuals are used as probabilistic references for one another, enabling local ancestry inference without labeled ancestral reference panels. Both methods are trained on a diverse set of simulated admixture scenarios to promote generalization across populations. We evaluate ORIAN on human, Drosophila melanogaster, and fully simulated datasets, comparing its performance against RFMix and LOTER across a range of admixture times and proportions. In the reference-based setting, ORIAN achieves the highest median diploid accuracy for recent admixture while remaining competitive across a broad range of scenarios. In the reference-free setting, ORIAN produces competitive local ancestry estimates using only admixed individuals, extending local ancestry inference to settings where ancestral reference panels are unavailable. These results demonstrate that ORIAN provides an accurate and flexible framework for both conventional and reference-free local ancestry inference.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.