Back

BfBio: a graph-based tool for the prediction of Angiogenic Stalk Cell genes using a Personalized PageRank algorithm

Bettoni, L.; Dmitrieva, J.; Mousa, M.; Alsafar, H.; Saeys, Y.; Zakeri, P.; Carmeliet, P.

2026-08-24 cancer biology
10.64898/2026.08.23.746493 bioRxiv
Show abstract

Although most human protein coding genes have functional annotations in databases, such as GeneCards, many remain poorly characterized. To address this gap, computational tools can be leveraged to predict the functional roles of under-annotated genes by extracting patterns from complex biological networks. Here we introduce Brain-for-Biotech (BfBio), a framework designed to identify genes important for vascular endothelial cells (EC), which are crucial cells for vessel formation (angiogenesis), vascular homeostasis, hemostasis and blood/tissue barrier function but also critical mediators of immunity and cancer progression. BfBio utilizes a Personalized PageRank (PPR) algorithm on an integrated network of different omics datasets and publicly available gene-gene/protein-protein interaction databases. In this study, we apply the predictive capabilities of BfBio to infer angiogenic stalk cell phenotype function in genes for which this function was not known before. By leveraging a set of genes characterizing the stalk cell cluster in lung tumor EC models previously identified, we have achieved a high Area Under Receiver Operative Characteristic (AUC-ROC) performance (0.837). Enrichment analysis, coupled with a text mining application, further confirmed that among the 49 predicted genes four of them were poorly characterized yet possessed biologically relevant properties and were linked to cancer, thereby validating BfBio as a robust tool for prioritizing novel therapeutic targets in vascular biology.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.