Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
Ercelen, D.; Caggiano, C.; Border, R.; Sankararaman, S.; Mangul, S.; Zaitlen, N.; Thompson, M.
Show abstract
Upwards of 40% of reads in sequencing datasets may be unmapped and discarded by standard protocols. Recent work has shown the utility of re-analyzing these unmapped reads to construct meaningful features, such as immune diversity repertoires or copy number variation in mtDNA and rDNA. While previous analyses of these features have produced significant correlations with diverse traits, they have generally been limited to analyses of RNA-sequencing data in phenotype-specific cohorts. Here, we explore whether associations can be identified using population-scale, whole-exome sequencing data in the UK BioBank. Using recently developed tools, we constructed multiple features including T-cell receptor diversity metrics, microbial load, and mtDNA and rDNA copy numbers for nearly 50,000 individuals in the UK BioBank. We first verify the validity of our method by showing that GWAS on these constructed traits results in replication of associations from studies in which the phenotypes were explicitly measured. Next, across several GWAS, we identified 21 novel independent significant loci in 11 genes, most of them in genes implicated in the innate immune response. Finally, we further analyzed the read-constructed features by establishing correlations to other population-level biobank traits such as immune disorders, metabolic disorders, neuropsychiatric disorders, and blood cell counts. Our results suggest that existing tools for feature construction from unmapped reads can offer novel information at the population level, and that these features can be used to establish novel genetic associations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A functional genomics approach to understand host genetic regulation of COVID-19 severity 96%
- Human leukocyte antigen class II gene diversity tunes antibody repertoires to common pathogens 94%
- Differential haplotype expression in class I MHC genes during SARS-CoV-2 infection of human lung cell lines 94%
Similar papers in this journal
Similar papers in this journal
- Common variation in a long non-coding RNA gene modulates variation of circulating TGF- β 2 levels in metastatic colorectal cancer patients (Alliance) 93%
- Expert Curation of the Human and Mouse Olfactory Receptor Gene Repertoires Identifies Conserved Coding Regions Split Across Two Exons 93%
- Characterization of a strain-specific CD-1 reference genome reveals potential inter- and intra-strain functional variability 93%
Similar papers in this journal
- ImmuneMirror: a Machine Learning-based Integrative Pipeline and Web Server for Neoantigen Prediction 92%
- RIBOSS detects novel translational events by combining long- and short-read transcriptome and translatome profiling 91%
- xQTLbiolinks: a comprehensive and scalable tool for integrative analysis of molecular QTLs 91%
Similar papers in this journal
- Genome-wide Copy Number Variations in a Large Cohort of Bantu African Children 93%
- Exome-wide analysis of copy number variation shows association of the human leukocyte antigen region with asthma in UK Biobank 93%
- Co-expression in tissue-specific gene networks links genes in cancer-susceptibility loci to known somatic driver genes 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.