Structured Pooling Improves Detection of Rare Regulatory Mutations in Population-Scale Reporter Assays
Dura, K.; Siklenka, K.; Strouse, K. P.; Morrow, S.; Zhang, C.; Barrera, A.; Allen, A. S.; Reddy, T. E.; Majoros, W. H.
Show abstract
Identifying genetic variants in noncoding DNA that impact gene expression and thereby contribute to disease risk remains a difficult but important challenge in genomic medicine. Modern reporter assays such as STARR-seq and MPRA provide an efficient and effective means of testing, in very high throughput, millions of variants captured directly from patient genomes. While these assays have previously been scaled to whole genomes and, separately, to populations, we report findings from the first whole-genome population-scale STARR-seq experiment performed on 100 individuals. In order to achieve that scale we devised a novel experimental design that partitions samples into pools so as to increase allele frequencies within pools and thereby reduce expected dropout and increase signal-to-noise ratio in experimental readouts. We show that this design produces more accurate estimates of variant effect sizes, and we provide a Bayesian model for robust estimation of those effect sizes that also reports full posterior distributions for assessment of confidence in estimates. Together, these methodological innovations facilitate the detection of functional regulatory variants, particularly rare variants, with much higher accuracy and at greater scale than previously possible. We demonstrate the utility of this approach on the task of functional annotation of quantitative trait loci such as eQTLs and caQTLs, and show concordance with patterns of constraint in transcription factor binding profiles.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Conditional resampling improves calibration and sensitivity in single cell CRISPR screen analysis 96%
- Deep-learning prediction of gene expression from personal genomes 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
Similar papers in this journal
- Quantitative occupancy of myriad transcription factors from one DNase experiment enables efficient comparisons across conditions 96%
- Allo: Accurate allocation of multi-mapped reads enables regulatory element analysis at repeats 96%
- DeepArk: modeling cis-regulatory codes of model species with deep learning 96%
Similar papers in this journal
Similar papers in this journal
- Machine learning-optimized targeted detection of alternative splicing 95%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 94%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.