TifBERT: a self-supervised foundation model for normalization-robust bulk RNA-seq representation learning
Hosseini, S.; Sharma, D.
Show abstract
Bulk RNA sequencing remains central to translational genomics, yet foundation-model development has largely focused on single-cell data. Existing transformer approaches for bulk RNA-seq often rely on expression discretization, numerical reconstruction, external gene embeddings, or restricted gene sets, limiting robustness across normalization schemes and cohorts. Here, we introduce TifBERT, a self-supervised framework for full-transcriptome bulk RNA-seq representation learning. TifBERT converts each unordered expression profile into a sample-specific gene sequence using term frequency-inverse document frequency (TF-IDF) ordering, prioritizing genes that are both highly expressed within a sample and selectively expressed across the cohort. It is then pretrained using masked gene modeling, predicting gene identities from transcriptomic context rather than reconstructing expression values. Pretrained on harmonized TCGA Pan-Cancer data spanning five RNA-seq normalization schemes, TifBERT learns contextual representations across approximately 10,000 genes without expression binning, landmark-gene restriction, or external biological embeddings. Across 33 TCGA cancer types, TifBERT achieved 90.83% accuracy, 0.996 macro AUC-ROC, and 0.903 MCC. It also captured pathway-level biology, achieving mean sample-wise and pathway-wise Pearson correlations of 0.754 and 0.762 across 1,387 PARADIGM pathway activities. Independent evaluation on GTEx healthy tissues showed preservation of tissue-level transcriptomic structure without retraining. In comparison with existing models, TifBERT achieves competitive subtype discrimination with substantially greater stability and produces markedly richer embedding geometry (effective rank 95.6 versus 6.3), without requiring expression discretization or in-distribution pretraining exposure. Together, TifBERT provides a scalable, normalization-independent foundation model for reusable bulk transcriptomic representation learning.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning tissue representation by identification of persistent local patterns in spatial omics data 95%
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 95%
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 95%
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 95%
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
Similar papers in this journal
- SCOPE: a normalization and copy number estimation method for single-cell DNA sequencing 94%
- Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck 93%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.