Back

An atlas-scale generative model for unified representation learning of bulk RNA-seq data

Pande, A.; Uyar, B.; Akalin, A.

2026-06-24 bioinformatics
10.64898/2026.06.18.733198 bioRxiv
Show abstract

Public bulk RNA-seq repositories contain hundreds of thousands of samples, creating opportunities for large-scale representation learning, but integration across studies remains challenging because of heterogeneous annotations, experimental protocols, and technical variation. While pre-trained foundation models are now widely available for single-cell RNA-seq, comparable resources for bulk RNA-seq remain scarce, motivating a model that learns a unified, tissue-aware representation directly from bulk data. We trained a supervised variational autoencoder (VAE) on a compendium of 118,263 bulk RNA-seq samples that we assembled from TCGA, GTEx, and ARCHS4 and mapped to 42 tissue categories. The model classifies tissue of origin at 94.9% balanced accuracy (weighted F1 96.2%) and compresses 16,115 genes into a 121-dimensional latent space. Tissue identity is the primary organizing axis of the latent space, while source effects remain secondary. To assess the impact of data volume, we constructed training sets at three different scales (38K, 75K, and 118K samples). Our results demonstrated that reconstruction fidelity improved incrementally with each expansion of the dataset, but with diminishing returns. We validated the model on an independent cohort of 734 paediatric tumour samples from TARGET, achieving 84.6% agreement with the expected tissue of origin. The trained model and code are available at GitHub (https://github.com/BIMSBbioinfo/flexynesis_tissue_vae_manuscript) with an interactive web application.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 2%
33.8%
2
Nature Machine Intelligence
70 papers in training set
Top 0.1%
11.7%
3
Nature Methods
385 papers in training set
Top 1%
7.2%
50% of probability mass above
4
Genome Biology
637 papers in training set
Top 2%
5.4%
5
Nature Biotechnology
172 papers in training set
Top 0.7%
5.4%
6
Nature Genetics
286 papers in training set
Top 1%
4.8%
7
Nucleic Acids Research
1281 papers in training set
Top 5%
4.0%
8
Nature
645 papers in training set
Top 5%
2.4%
9
PLOS Computational Biology
1863 papers in training set
Top 14%
1.9%
10
Bioinformatics
1204 papers in training set
Top 7%
1.7%
11
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.7%
12
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.3%
13
Science Advances
1243 papers in training set
Top 26%
1.1%
14
eLife
5828 papers in training set
Top 58%
1.1%
15
Genome Research
468 papers in training set
Top 5%
1.0%
16
Nature Biomedical Engineering
47 papers in training set
Top 1%
1.0%
17
Communications Biology
993 papers in training set
Top 26%
1.0%
18
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 40%
0.9%
19
Cell Systems
201 papers in training set
Top 5%
0.8%
20
Cell Genomics
172 papers in training set
Top 4%
0.8%
21
Scientific Reports
3612 papers in training set
Top 79%
0.6%
22
Genome Medicine
183 papers in training set
Top 6%
0.6%