Pangenome-based human genome analysis improves trait association and genomic prediction
Lu, S.; Liao, W.-W.; DeGorter, M. K.; Goddard, P. C.; Ebler, J.; Lu, T.-Y.; Chaisson, M. J. P.; Marschall, T.; Montgomery, S. B.; Stitziel, N. O.; Hall, I. M.
Show abstract
The Human Pangenome Reference Consortium has generated 462 open-access reference genomes and a variation graph that represents differences among them, providing a substrate for pangenome-based analysis methods that overcome the longstanding limitation of comparing all genomic data to a single linear reference. A key unresolved question is the extent to which these approaches can improve trait mapping. We investigate this using the genetics of gene expression variation as a model. We developed a graph-based method (EdgeDepth) for associating sequence variation with traits using short-read genome sequencing data, and show that it captures complex forms of genetic variation missed by other methods. We evaluated trait mapping performance using 430 samples with deep RNA-seq data, and found that pangenomic methods enable the detection of expression quantitative trait loci involving multiallelic indels and structural variants, leading to increased power at a subset of genes. These include 812 genes (7.9% of total) with [≥]20% improvement in statistical significance relative to the 1000 Genomes Project callset, and 185 (1.8%) with a 50% improvement, 10 of which are candidates to explain prior GWAS results. Notably, these analyses implicate GBAP1 pseudogene copy number as a causal factor in Crohn's disease, likely via miRNA-mediated regulation of GBA1, which explains prior GWAS results based on flanking SNPs. The inclusion of pangenome-specific variation also improved the performance of gene expression prediction models, with median variance explained increasing from 10.1% to 12.5%, and 14.6% of genes showing significant improvement ({Delta}r2>0.05). Taken together, these results suggest that integration of pangenomic methods into human genetic studies will improve trait association and genomic prediction at a meaningful subset of genes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 98%
- Systematic assessment of regulatory effects of human disease variants in pluripotent cells 97%
- Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank 97%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.