Back

Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

Li, S.; Sprague, D. A.; Eberhard, Q. E.; Boyson, S. P.; Laederach, A.; Calabrese, J. M.

2026-07-08 bioinformatics
10.64898/2026.07.03.736302 bioRxiv
Show abstract

Long noncoding RNAs (lncRNAs) play roles in gene regulation across kingdoms of life. However, lncRNAs with related functions often lack linear sequence similarity, making it difficult to leverage studies of one lncRNA to inform the understanding of others. We describe a k-mer-based hidden Markov model, hmSEEKR, that enables the scanning of transcriptomes for regions of non-linear sequence similarity to a query domain, without prior knowledge of where within the transcriptome the similarities may be located. When individual lncRNA domains were used as search features, hmSEEKR successfully identified regions in other RNAs that harbor non-linear sequence similarity and bind similar sets of proteins. Applying hmSEEKR to transcriptome-wide searches, we found that certain domains within the lncRNAs XIST, NEAT1, and MALAT1 exhibited widespread regional similarity to both lncRNA and protein-coding genes, while others were more unique, exhibiting similarity to ~100 genes or fewer. Combinatorial searches uncovered RNAs containing sequential matches to core functional domains of XIST and NEAT1, and eCLIP-inferred protein-interaction networks within these RNAs more closely resembled those of XIST and NEAT1, respectively, than would be expected by chance, suggesting the searches recovered RNAs with similar biological properties. Finally, within annotated sets of cis-activating and cis-repressive lncRNAs, we observed opposing enrichments for similarity to domains associated with transcription-promoting complexes and heterogeneous nuclear ribonucleoprotein (hnRNP) binding, respectively, suggesting the enriched sequences may contribute to regulatory functions. hmSEEKR can be applied with minimal training data and enables the a priori discovery of RNA domains that share nonlinear similarity, offering a sequence-informed approach to discover functional elements within noncoding transcriptomes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.1%
11.9%
2
Nucleic Acids Research
1281 papers in training set
Top 2%
11.6%
3
Genome Biology
637 papers in training set
Top 0.9%
10.7%
4
Nature Communications
5641 papers in training set
Top 26%
6.1%
5
PLOS Computational Biology
1863 papers in training set
Top 6%
6.1%
6
RNA
189 papers in training set
Top 0.4%
5.4%
50% of probability mass above
7
Genome Research
468 papers in training set
Top 1%
5.3%
8
Bioinformatics
1204 papers in training set
Top 4%
4.7%
9
Scientific Reports
3612 papers in training set
Top 39%
2.7%
10
Nature Methods
385 papers in training set
Top 3%
2.3%
11
BMC Genomics
406 papers in training set
Top 4%
2.1%
12
PLOS ONE
5266 papers in training set
Top 46%
2.1%
13
Cell Systems
201 papers in training set
Top 2%
2.1%
14
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
15
eLife
5828 papers in training set
Top 51%
1.6%
16
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.1%
1.5%
17
Cell Reports
1498 papers in training set
Top 22%
1.3%
18
PeerJ
308 papers in training set
Top 9%
1.1%
19
RNA Biology
78 papers in training set
Top 1.0%
1.1%
20
Molecular Cell
350 papers in training set
Top 4%
1.1%
21
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
22
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
23
Nature Biotechnology
172 papers in training set
Top 4%
0.8%
24
Frontiers in Genetics
230 papers in training set
Top 6%
0.8%
25
iScience
1154 papers in training set
Top 37%
0.8%
26
Molecular Biology and Evolution
542 papers in training set
Top 6%
0.6%
27
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%