Back

Assessing large-scale genomic language models in predicting personal gene expression: promises and limitations

Li, S.; Luo, R.; Huang, Y.

2025-07-14 bioinformatics
10.1101/2025.07.09.664024 bioRxiv
Show abstract

Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to predict personal gene expression remains largely unexplored. We developed a framework, gLM2X-Tower, to benchmark gLMs and sequence-to-function (S2F) models on this task with paired personal genome-transcriptome data. With individual-level training, we found that similar to S2F models (e.g., AlphaGenome), gLMs (e.g., Evo2) remain incapable of predicting the inter-person variability on held-out genes. However, such training improves prediction for seen genes in new individuals, particularly by gLMs, highlighting the potential applications in few-shot settings like for rare variants

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.