Back

bMAE: Masked Autoencoder Latent Representations for Bulk RNA-seq Tissues

Wan, Z.; Untalan, M. Z. G.; Vasconcellos Vargas, D.

2026-03-05 bioinformatics
10.64898/2026.03.03.709470 bioRxiv
Show abstract

Bulk tissue RNA-sequencing data from large-scale consortia such as GTEx provide comprehensive gene expression profiles across diverse human tissues. However, the high-dimensional nature of bulk RNA-seq data, combined with technical noise and batch effects, poses challenges for downstream analyses. While dimensionality reduction methods are routinely applied, standard approaches such as PCA often fail to optimally preserve tissue-discriminative information and exhibit poor generalization to unseen tissue types. We developed a masked autoencoder for bulk tissue RNA-seq that learns compressed latent representations through self-supervised learning with variable masking schedules. Evaluated on GTEx data (31 tissue types, 19,788 samples, 19,308 genes), our method substantially outperformed all baselines in leave-one-tissue-out (LOTO) cross-validation. We achieved mean silhouette 0.20, ARI 0.58, and NMI 0.84 versus best baseline UMAP (0.007, 0.25, 0.47), representing 28.6-fold, 2.3-fold, and 1.8-fold improvements. Remarkably, within held-out tissue categories containing subtissues, our latent space revealed hierarchical structure with enhanced subtissue separation (silhouette 0.16, ARI 0.35, NMI 0.36) versus baselines (silhouette 0.15, ARI 0.20, NMI 0.20), despite never observing these distinctions during training. The method compressed 19,308 genes to 128 dimensions while preserving multi-scale structure.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.