Back

Exploring molecular signatures of senescence with markeR, an R toolkit for evaluating gene sets as phenotypic markers

Martins-Silva, R.; Kaizeler, A.; Barbosa-Morais, N. L.

2025-12-09 bioinformatics
10.64898/2025.12.05.692517 bioRxiv
Show abstract

Many biological processes, including cellular senescence, manifest as diverse phenotypes that vary across cell types and conditions. In the absence of single, definitive markers, researchers often rely on the expression of sets of genes to identify these complex states. However, there are multiple ways to summarise gene set expression into quantitative metrics (i.e., signatures), each with its own strengths and limitations, and we know of no consensual framework to systematically evaluate their performance across datasets. We therefore developed markeR (https://bioconductor.org/packages/markeR), an open-source, modular R package that evaluates gene sets as phenotypic markeRs using various scoring and enrichment-based approaches. markeR generates interpretable metrics and intuitive visualisations that enable benchmarking of gene signatures and exploration of their associations with chosen study variables. As a case study, we applied markeR to 9 published senescence-related gene sets across 25 RNA-seq datasets, covering 6 human cell types and 12 senescence-inducing conditions. There was wide variability in gene set performance, as some signatures (e.g., SenMayo) were robust senescence markers across contexts, while others (e.g., those from MSigDB), performed poorly as such. We also used markeR to analyse gene expression in 49 GTEx tissues, revealing tissue- and age-related differences in senescence-associated signals. Together, these findings emphasise the difficulty of characterising molecular phenotypes and demonstrate the potential of markeR in facilitating the systematic evaluation of gene sets in various biological contexts.

Published in NAR Genomics and Bioinformatics (predicted rank #13) · training set

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.