Predicting cell type-specific coverage profiles from DNA sequence
Linder, J.; Yuan, H.; Kelley, D. R.
Show abstract
Predicting expression profiles from RNA-seq experiments provides a powerful approach for universal sequence-based variant effect prediction, enabling researchers to score variants that affect total gene expression and relative isoform abundances. These models can be repurposed for new prediction tasks through transfer learning. However, current base models train primarily on bulk RNA-seq profiles derived from tissues and cell lines, overlooking the wealth of single-cell 3-seq data that captures cell type-specific gene regulation. Here, we extend the capabilities of our recently developed Borzoi model by training on single-cell 3-seq expression profiles from the Tabula Sapiens, Tabula Muris, and the Adult Brain Atlas, aggregated by cell type. This new model, Borzoi Prime, enables accurate variant interpretation across diverse cells, spanning erythrocytes to microglia. Training on 3-seq profiles improves the models ability to predict cell- and tissue-specific alternative polyadenylation, even in the original bulk RNA-seq data. Through UTR-wide mutagenesis experiments of alternatively polyadenylated genes, we highlight determinants of cell type-specific 3 UTR regulation learned by the model. This cell type-resolved approach opens new possibilities for understanding genetic variant effects via multiple layers of regulation in specific cellular contexts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 96%
- Leveraging supervised learning for functionally-informed fine-mapping of cis-eQTLs identifies an additional 20,913 putative causal eQTLs 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.