Characterizing uncertainty in predictions of genomic sequence-to-activity models
Bajwa, A.; Rastogi, R.; Kathail, P.; Shuai, R. W.; Ioannidis, N. M.
Show abstract
Genomic sequence-to-activity models are increasingly utilized to understand gene regulatory syntax and probe the functional consequences of regulatory variation. Current models make accurate predictions of relative activity levels across the human reference genome, but their performance is more limited for predicting the effects of genetic variants, such as explaining gene expression variation across individuals. To better understand the causes of these shortcomings, we examine the uncertainty in predictions of genomic sequence-to-activity models using an ensemble of Basenji2 model replicates. We characterize prediction consistency on four types of sequences: reference genome sequences, reference genome sequences perturbed with TF motifs, eQTLs, and personal genome sequences. We observe that models tend to make high-confidence predictions on reference sequences, even when incorrect, and low-confidence predictions on sequences with variants. For eQTLs and personal genome sequences, we find that model replicates make inconsistent predictions in >50% of cases. Our findings suggest strategies to improve performance of these models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A base-resolution panorama of the in vivo impact of cytosine methylation on transcription factor binding 96%
- A unified encyclopedia of human functional DNA elements through fully automated annotation of 164 human cell types 96%
- HATCHet2: clone- and haplotype-specific copy number inference from bulk tumor sequencing data 96%
Similar papers in this journal
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 96%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.