Back

EES-Transformer: A Dual-Path Transformer for Tissue Classification and Gene Representation Learning from Extreme Expression Sets

Park, J.-S.; Lee, Y.; Kang, Y. J.

2026-02-23 bioinformatics
10.64898/2026.02.22.707241 bioRxiv
Show abstract

Accurate tissue classification from gene expression data is fundamental to transcriptomic analysis. Here we introduce EES-Transformer V2, a dual-path transformer architecture that learns from Extreme Expression Sets (EES)--sequences of genes at expression extremes (above 95th or below 5th percentile). The architecture separates tissue classification from masked language modeling through independent branches: the classification branch operates without tissue label information, while the generative branch receives tissue conditioning. This design enables fair evaluation of classification performance while learning tissue-specific gene relationships. Applied to 12,212 Arabidopsis thaliana RNA-seq samples spanning 47 tissue types, EES-Transformer achieves 91-92% classification accuracy (varying across evaluation runs due to stochastic input masking)--substantially above the 2.1% random baseline. Attention-based analysis identifies 3,524 gene-tissue high-attention associations whose importance patterns reflect known biology. Critically, while individual high-attention genes appear broadly across tissues, gene pairs from attention-derived regulatory networks show higher tissue specificity: pollen gene pairs show a 2.7-fold enrichment over single-gene rates, and root and leaf pairs each show 1.5-fold enrichment. This finding reveals that tissue identity is encoded in combinatorial gene expression patterns rather than individual genes. Attention-derived gene regulatory networks exhibit scale-free topology and biologically coherent hub gene programs, with pollen networks consisting entirely of DOWN-DOWN interactions among silenced vegetative genes. EES-Transformer provides accurate tissue classification, interpretable gene importance scores, and attention-derived regulatory networks for biological discovery.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.