Back

UniOP: a universal operon prediction for high-throughput prokaryotic (meta-)genomic data using intergenic distance

Su, H.; Soeding, J.; Zhang, R. S.

2024-11-11 bioinformatics
10.1101/2024.11.11.623000 bioRxiv
Show abstract

The study of the deluge of metagenomic and genomic sequences is challenging due to the severe lack of function information. Predicting operons, groups of functionally related genes in prokaryotic genomes, is critical for bridging this gap. However, existing methods for operon prediction heavily rely on experimental data, functional annotations, or extensive characterization of homologous genes, making it difficult to accurately predict operons in newly sequenced or poorly characterized genomes. Here, we introduce UniOP, an unsupervised approach that uses a statistical model to predict operons from intergenic distances directly derived from the target genomic sequence. UniOP not only outperforms alternative approaches on ten complete genomes but also shows superior results on 3269 metagenome-assembled genomes across 13 bacterial and 2 archaeal phyla. Furthermore, we explored enhancing UniOP by incorporating the conservation of gene neighborhood and strandedness in respective genomes and examined the influence of Pfam annotations and motif searching on its performance.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.