Deconvolution-derived cell-type expression targets for personal genome sequence-to-expression prediction
Sim, S.; Shen, L.
Show abstract
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue-cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122-0.142 for sequence-derived approaches and 0.081-0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution-pseudobulk agreement for sequence-derived models (r = 0.35-0.43 across tissue-cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- omnideconv: a unifying framework for using and benchmarking single-cell-informed deconvolution of bulk RNA-seq data 95%
- Multi-cell type deconvolution using a probabilistic model for single-molecule DNA methylation haplotypes 95%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 95%
Similar papers in this journal
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 94%
- SpaTM: Topic Models for Inferring Spatially Informed Transcriptional Programs 94%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
Similar papers in this journal
Similar papers in this journal
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 95%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 94%
- On the discovery of population-specific state transitions from multi-sample multi-condition single-cell RNA sequencing data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.