Back

Open resource of bacterial biofilm for the prediction of biofilm-associated proteins

Zhang, Z.; Wajid, H.; Li, E.

2023-09-13 bioinformatics
10.1101/2023.09.10.556831 bioRxiv
Show abstract

SummaryBacterial biofilms are organized heterogeneous assemblages of microbial cells encased within a self-produced matrix of exopolysaccharides, extracellular DNA and proteins. Over the last decade, more and more biofilm-associated proteins have been discovered and investigated. Furthermore, omics techniques such as transcriptomes, proteomes also play important roles in identifying new biofilm-associated genes or proteins. However, those important data have been uploaded separately to various databases, which creates obstacles for biofilm researchers to have a comprehensive access to these data. In this work, we constructed BBSdb, a state-of-the-art open resource of bacterial biofilm-associated protein. It includes 48 different bacteria species, 105 transcriptome datasets, 21 proteome datasets, 1205 experimental samples, 57,823 differentially expressed genes (DEGs), 13,605 differentially expressed proteins (DEPs), 1,930 Top 5% differentially expressed genes and a predictor for prediction of biofilm-associated protein. In addition, 1,781 biofilm-associated proteins, including annotation and sequences, were extracted from 942 articles and public databases via text-mining analysis. We used E. coli as an example to represent how to explore potential biofilm-associated proteins in bacteria. We believe that this study will be of broad interest to researchers in field of bacteria, especially biofilms, which are involved in bacterial growth, pathogenicity, and drug resistance. Availability and implementationThe BBSdb is freely available at http://124.222.145.44/#!/.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.