Single-cell gene expression prediction from DNA sequence at large contexts
Schwessinger, R.; Deasy, J.; Woodruff, R. T.; Young, S.; Branson, K. M.
Show abstract
Human genetic variants impacting traits such as disease susceptibility frequently act through modulation of gene expression in a highly cell-type-specific manner. Computational models capable of predicting gene expression directly from DNA sequence can assist in the interpretation of expression-modulating variants, and machine learning models now operate at the large sequence contexts required for capturing long-range human transcriptional regulation. However, existing predictors have focused on bulk transcriptional measurements where gene expression heterogeneity can be drowned out in broadly defined cell types. Here, we use a transfer learning framework, seq2cells, leveraging a pre-trained epigenome model for gene expression prediction from large sequence contexts at single-cell resolution. We show that seq2cells captures cell-specific gene expression beyond the resolution of pseudo-bulked data. Using seq2cells for variant effect prediction reveals heterogeneity within annotated cell types and enables in silico transfer of variant effects between cell populations. We demonstrate the challenges and value of gene expression and variant effect prediction at single-cell resolution, and offer a path to the interpretation of genomic variation at uncompromising resolution and scale.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 97%
- Towards Universal Cell Embeddings: Integrating Single-cell RNA-seq Datasets across Species with SATURN 96%
- Joint probabilistic modeling of paired transcriptome and proteome measurements in single cells 96%
Similar papers in this journal
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 96%
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 96%
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.