A large-scale foundation model for bulk transcriptomes
kang, b.; Fan, R.; Yi, M.; Cui, C.; Cui, Q.
Show abstract
Large language models (LLMs) have emerged as powerful foundation models leading to breakthroughs in transcriptome analysis. However, current RNA-seq foundation models are exclusively pretrained on sparse single-cell RNA-seq (scRNA-seq) data, which typically detects only [~]3000 genes per cell. This thus creates a critical gap in models specifically designed for bulk transcriptomes, a fundamentally different modality capable of profiling [~]16,000 genes per sample. Here we propose BulkFormer, a large-scale foundation model for bulk transcriptome analysis. With 150 million parameters covering about 20,000 protein-coding genes, BulkFormer is pretrained on over 500,000 human bulk transcriptomic profiles. BulkFormer incorporates a hybrid encoder architecture, combining a graph neural network to capture explicit gene-gene interactions and a performer module to model global expression dependencies. As a result, despite incurring much lower training costs than scRNA-seq foundation models, BulkFormer consistently outperforms them in all six downstream tasks: transcriptome imputation, disease annotation, prognosis modeling, drug response prediction, compound perturbation simulation, and gene essentiality scoring. Notably, BulkFormer not only enhances the discovery of novel clinical biomarkers but also uncovers latent disease mechanisms by imputing biologically meaningful gene expression. Collectively, these results demonstrate BulkFormers power as a versatile and robust framework for bulk transcriptome modeling and biomedical discovery, bridging a critical gap in the current foundation model landscape.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- scTrace+: enhance the cell fate inference by integrating the lineage-tracing and multi-faceted transcriptomic similarity information 96%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 96%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 96%
Similar papers in this journal
- stGCL: A versatile cross-modality fusion method based on multi-modal graph contrastive learning for spatial transcriptomics 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- CREaTor: zero-shot cis-regulatory pattern modeling with attention mechanisms 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.