Back

Uncertainty-Aware Gene Rankings Reveal Key Players in Coexpression Networks

Mahapatra, S.; Subramanian, N. A.; Narayanan, M.

2025-12-16 bioinformatics
10.64898/2025.12.13.694156 bioRxiv
Show abstract

Key genes of a complex biological system are often identified by inferring a coexpression network from transcriptomic data and analyzing it using network science measures such as centrality. However, relatively modest sample sizes of transcriptomic data, along with heterogeneity in the human population, can lead to uncertainty about the "true" coexpression network. Earlier studies have quantified this uncertainty using bootstrap resampling or similar analysis, but fewer have investigated the extent to which this uncertainty affects downstream network analyses like centrality. This work presents a systematic workflow that estimates and propagates uncertainty about the coexpression network to downstream measures such as degree/PageRank centrality, with the goal of producing a robust (uncertainty-aware) gene ranking and comparing its performance to traditional centrality based rankings. Specifically, we propose bootstrap-based node scores (BOONS) of the form {micro}-c{sigma} (for different values of c) that combines the expectation ({micro}) and variance ({sigma}2) of a centrality measure computed across boot-strapped coexpression networks to prioritize stable central genes. We assess their efficacy with respect to reproducibility, replicability, and tissue specificity on simulated and real-world (GTEx) transcriptomic datasets, and provide sample size recommendations as well. In several of these benchmarks, our BOONS show significantly better performance than the baseline measure that does not incorporate uncertainty. For instance, across five tested GTEx tissues each subsampled to a modest 237 sample size, our proposed ({micro} -{sigma} )-based ranking shows an average improvement of 48.5% over the baseline for recovering a set of 100 reference genes. Based on results from these benchmarks, especially one pertaining to the recovery of tissue-specific reference genes, we recommend the usage of ({micro} -{sigma} ) scoring. All our results taken together quantify the negative impact of uncertainty on vanilla centrality based gene rankings, and underscore how adoption of uncertainty-aware analysis such as BOONS proposed in this work is long overdue to obtain stable biological network measures. Code availabilityThe code for bootstrap-based node scores generation and all other associated analyses is at https://github.com/BIRDSgroup/Uncertainty-Propagation-using-Bootstrap-Analysis. Data availabilityThe computed bootstrap-based node scores and all other associated results, both for simulated and real-world datasets, are publicly available at https://drive.google.com/drive/folders/1Ru2zukLijqTYofn8lNLri2s7z9CTOS7i?usp=sharing.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.