Back

BATTER: Accurate Prediction of Rho-dependent and Rho-independent Transcription Terminators in Metagenomes

Jin, Y.; Ma, H.; Xu, Z. Z.; Lu, Z. J.

2023-10-03 bioinformatics
10.1101/2023.10.02.560326 bioRxiv
Show abstract

Bacterial transcription termination is a critical yet underexplored mechanism of gene regulation in microbial ecosystems. Existing computational tools, however, primarily focus on predicting transcript 3 ends generated by Rho-independent terminators (RITs) in model species, leaving significant gaps in understanding those generated by Rho-dependent terminators (RDTs), especially in non-model species. To address these limitations, we developed BATTER (BActeria Transcript Three Prime End Recognizer), a comprehensive computational tool for bacterial transcript 3 termini prediction. BATTER builds on the observation that conserved stem-loop structures are frequently associated with 3 ends of primary transcripts generated by both RIT and RDT mechanisms across distantly related bacterial species. BATTER demonstrated its advantage compared to existing tools. It enabled a comprehensive analysis of 42,905 representative bacterial genomes, uncovering that stem-loop structures exhibit clade-specific properties with greater variations between species than between gene families. Notably, BATTER uncovered that certain Cyanobacteria lineages, despite lacking Rho homologs, harbor Rho utilization (RUT) site-like sequences near 3 ends, with preliminary experimental validation in E. coli suggesting their partial functionality in transcription termination. Additionally, BATTER systematically identified pervasive premature termination events in antimicrobial resistance (AMR) genes, highlighting their regulatory roles in translation protection and drug efflux. This study advances our understanding of transcription termination across diverse bacterial lineages and provides a robust computational approach for exploring transcription regulation in complex microbial ecosystems.

Published in Microbiome (predicted rank #12) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.