Integrating suboptimal secondary structures, AI-assisted genomic synteny, and evolutionary conservation to identify bacterial ncRNA homologs beyond sequence similarity
Panek, J.
Show abstract
A bioinformatic approach for genome-wide identification of homologs of bacterial non-coding RNAs (ncRNAs) integrating structural similarity, genomic synteny, and evolutionary conservation is presented. The structural similarity is detected using an algorithm for genome-wide identification of loci in genomic intergenic regions (IGRs) containing sequences capable of adopting secondary structures similar to that of the query ncRNA. The algorithm scans IGR sequences using a sliding window with a predefined step. For each window, suboptimal secondary structures are predicted and compared with the template structure to compute structural similarity scores. These scores are evaluated statistically on a genome-wide scale to infer homology of the RNAs represented by the predicted structures. Loci encoding statistically significant structures are further filtered using genomic synteny of the query ncRNAs inferred from genomic annotations. ChatGPT was used to assist in identifying literature-supported biological relationships between genes with distinct functional annotations. Syntenic loci with the structures are then examined for homologs in related species, as evolutionary conservation among related species is a common feature of ncRNAs Using this approach, we predicted novel homologs of the spot42 RNA-encoding spf gene in Glaciecola and Pseudoalteromonas genomes, and ms1 RNA genes in Frankia and Bifidobacterium genomes, where previous homology searches had failed. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=44 SRC="FIGDIR/small/737749v1_ufig1.gif" ALT="Figure 1"> View larger version (13K): org.highwire.dtl.DTLVardef@dd18baorg.highwire.dtl.DTLVardef@18261eforg.highwire.dtl.DTLVardef@eb95c3org.highwire.dtl.DTLVardef@b535ad_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Identifying Genomic Islands with Deep Neural Networks 94%
- Comprehensive genome-wide identification of angiosperm upstream ORFs with peptide sequences conserved in various taxonomic ranges using a novel pipeline, ESUCA 94%
- Complete genome sequence and annotation of the laboratory reference strain Shigella flexneri serovar 5a M90T and genome-wide transcriptional start site determination 93%
Similar papers in this journal
- Platon: identification and characterization of bacterial plasmid contigs in short-read draft assembliesexploiting protein-sequence-based replicon distribution scores 94%
- cazy_webscraper: local compilation and interrogation of comprehensive CAZyme datasets 93%
- Whokaryote: distinguishing eukaryotic and prokaryotic contigs in metagenomes based on gene structure 93%
Similar papers in this journal
- Identification of RNA 3' ends and termination sites in Haloferax volcanii 94%
- PresRAT: A server for identification of bacterial small-RNA sequences and their targets with probable binding region. 94%
- From reporters to endogenous genes: the impact of the first five codons on translation efficiency in Escherichia coli 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.