LoGicAl: Local Ancestry and Genotype Calling Uncertainty-aware Ancestry-specific Allele Frequency Estimation from Admixed Samples
Wang, J.; Zöllner, S.
Show abstract
In admixed groups, it is of interest to estimate allele frequencies of their ancestral contributing populations. These ancestry-specific allele frequencies inform genetic drivers of disease etiologies, facilitate genome-wide association study interpretations, enhance polygenic risk prediction and portability, and provide insights into demographic history. Their estimation leverages inferred locus-specific ancestry background, i.e., local ancestry, generated by tools like RFMix. However, existing estimation methods lose accuracy and incur biases by failing to model uncertainty from upstream local ancestry inference and genotype calling. Here, we introduce LoGicAl, a novel likelihood-based method for estimating ancestry-specific allele frequencies from admixed samples, simultaneously accommodating uncertainty from ancestry calling, genotyping, and statistical phasing. We demonstrate that modeling these uncertainties substantially reduces estimation errors, resulting in superior accuracy for both sequence-based and array-based genotyping with different levels of local ancestry inference quality. By integrating an accelerated fixed-point algorithm, LoGicAl achieves high scalability and enhanced computational efficiency compared with existing approaches. Applying LoGicAl to admixed cohorts in the 1000 Genomes Project, we illustrate the benefits of local-ancestry-based allele frequency estimates. Together, LoGicAl contributes to genomic analyses of admixed samples by providing precise and rapid ancestry-specific allele frequency estimates, and to constructing the spatial landscape and dynamics of genetic variations in admixed populations at a finer scale.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 96%
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 94%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 94%
Similar papers in this journal
- Scalable probabilistic PCA for large-scale genetic variation data 96%
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 96%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 95%
Similar papers in this journal
- Reducing reference bias using multiple population reference genomes 95%
- ContamLD: Estimation of Ancient Nuclear DNA Contamination Using Breakdown of Linkage Disequilibrium 94%
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.