An Atlas of Linkage Disequilibrium Across Species
Zhu, T.-N.; Huang, X.; Yang, M.-y.; Qi, G.-A.; Zhang, Q.-X.; Lin, F.; Zhang, W.; Zhang, Z.; Jin, X.; Zheng, H.-F.; Xu, H.; Yu, S.; Chen, G.-B.
Show abstract
Linkage disequilibrium (LD) is a key metric that characterizes populations in flux. To reach a genomic scale LD illustration, which has a substantial computational cost of[O] (nm2), we introduce a framework with two novel algorithms for LD estimation: X-LD, with a time complexity of[O] (n2m) suitable for small sample sizes (n < 104); X-LDR, a stochastic algorithm with a time complexity of[O] (nmB) for biobank-scale data (B iterations); n the sample size, and m the number of SNPs. These methods can refine the entire genome into high-resolution LD grids, such as more than 9 million grids for UK Biobank samples ([~]4.2 million SNPs). The efficient resolution for genome-wide LD leads to intriguing biological discoveries. I) High-resolution LD illustrations revealed how the pericentromeric regions and the HLA region lead to intense and extended LD patterns. II) Two universal LD patterns, identified as Norm I and Norm II patterns, provide insights on the evolutionary history of populations and can also highlight genomic regions of deviation, such as chromosomes 6 and 11 or ncRNA regions. III) The results of our innovative LD decay method aligned with the LD decay scores of 59.5 for Europeans, 60.2 for East Asians, and 33.2 for Africans; correspondingly, the length of the LD was approximately 2.85 Mb, 2.18 Mb, and 1.58 Mb for these three ethnicities. Rare or imputed variants universally increased LD. IV) An unprecedented LD atlas for 25 reference populations contoured interspecies diversity in terms of their Norm I and Norm II LD patterns, highlighting the impact of refined population structure, quality of reference genomes, and uncovered a profound status quo of these populations. The algorithms have been implemented in C++ and are freely available (https://github.com/gc5k/gear2).
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Characterization of Structural Variation in Tibetans Reveals New Evidence of High-altitude Adaptation and Introgression 95%
- Mapping genetic variants for nonsense-mediated mRNA decay regulation across human tissues 95%
- HOPS: a quantitative score reveals pervasive horizontal pleiotropy in human genetic variation is driven by extreme polygenicity of human traits and diseases 95%
Similar papers in this journal
- sgcocaller and comapr: personalised haplotype assembly and comparative crossover map analysis using single-gamete sequencing data 96%
- Organisation of gene programs revealed by unsupervised analysis of diverse gene-trait associations 96%
- scIGANs: single-cell RNA-seq imputation using generative adversarial networks 95%
Similar papers in this journal
- Human and rat skeletal muscle single-nuclei multi-omic integrative analyses nominate causal cell types, regulatory elements, and SNPs for complex traits 96%
- Predicting unrecognized enhancer-mediated genome topology by an ensemble machine learning model 96%
- A limited set of transcriptional programs define major cell types 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.