Back

Predicting Rhizosphere Competence Related Catabolic Gene Clusters in plant-associated bacteria with RhizoSMASH

Li, Y.; Sun, M.; Raaijmakers, J. M.; Mommer, L.; Zhang, F.; Song, C.; Medema, M. H.

2025-04-02 bioinformatics
10.1101/2025.03.29.646099 bioRxiv
Show abstract

Plants release a substantial fraction of their photosynthesized carbon into the rhizosphere as root exudates, a mix of chemically diverse compounds that drive microbiome assembly. Deciphering how plants modulate the composition and activities of rhizosphere microbiota through root exudates is challenging, as no dedicated computational methods exist to systematically identify microbial root exudate catabolic pathways. Here we used and integrated published information on catabolic genes in bacterial taxa that contribute to their rhizosphere competence. We developed the RhizoSMASH algorithm for genome-synteny-based annotation of rhizosphere-competence-related catabolic gene clusters (rCGCs) in bacteria by means of a set of 58 knowledge-based logic detection rules carefully curated through sequence similarity network analysis. Our analysis revealed large heterogeneity of rCGC prevalence both across and within plant-associated bacterial taxa, indicating extensive niche specialization. Furthermore, we validated that the presence or absence of rCGCs in bacterial genomes reflects their catabolic capacity and is predictive for their rhizosphere competence by aligning rhizoSMASH results with paired genome/metabolome datasets of rhizobacterial taxa. RhizoSMASH provides an extensible framework for studying rhizosphere bacterial catabolism, allowing targeted selection of beneficial bacterial taxa for microbiome-assisted breeding approaches for sustainable agriculture.

Published in Nature Communications (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.