Back

CircSSNN: circRNA-binding site prediction via sequence self-attention neural networks with pre-normalization

Cao, C.; Yang, S.; Li, M.; Li, C.

2023-02-07 bioinformatics
10.1101/2023.02.07.527436 bioRxiv
Show abstract

Circular RNAs (circRNAs) play a significant role in some diseases by acting as transcription templates. Therefore, analyzing the interaction mechanism between circRNA and RNA-binding proteins (RBPs) has far-reaching implications for the prevention and treatment of diseases. Existing models for circRNA-RBP identification most adopt CNN, RNN, or their variants as feature extractors. Most of them have drawbacks such as poor parallelism, insufficient stability, and inability to capture long-term dependence. To address these issues, we designed a Seq_transformer module to extract deep semantic features and then propose a CircRNA-RBP identification model based on Sequence Self-attention with Pre-normalization. We test it on 37 circRNA datasets and 31 linear RNA datasets using the same set of hyperparameters, and the overall performance of the proposed model is highly competitive and, in some cases, significantly out-performs state-of-the-art methods. The experimental results indicate that the proposed model is scalable, transformable, and can be applied to a wide range of applications without the need for task-oriented fine-tuning of parameters. The code is available at https://github.com/cc646201081/CircSSNN. Author summaryIn this paper, we propose a new method completely using the self-attention mechanism to capture deep semantic features of RNA sequences. On this basis, we construct a CircSSNN model for the cirRNA-RBP identification. The proposed model constructs a feature scheme by fusing circRNA sequence representations with statistical distributions, static local context, and dynamic global context. With a stable and efficient network architecture, the distance between any two positions in a sequence is reduced to a constant, so CircSSNN can quickly capture the long-term dependence and extract the deep semantic features. Experiments on 37 circRNA datasets show that the proposed model has overall advantages in stability, parallelism, and prediction performance. Keeping the network structure and hyperparameters unchanged, we directly apply CircSSNN to linRNA datasets. The favorable results show that CircSSNN can be transformed simply and efficiently without task-oriented tuning. In conclusion, CircSSNN can serve as an appealing circRNA-RBP identification tool with good identification performance, excellent scalability, and wide application scope, which is expected to reduce the professional threshold required for hyperparameter tuning in bioinformatics analysis.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Briefings in Bioinformatics
354 papers in training set
Top 0.1%
27.3%
2
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.1%
10.9%
3
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.1%
8.1%
4
BMC Bioinformatics
457 papers in training set
Top 2%
4.1%
50% of probability mass above
5
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.2%
3.5%
6
Bioinformatics
1204 papers in training set
Top 5%
3.3%
7
Genomics, Proteomics & Bioinformatics
172 papers in training set
Top 0.6%
3.3%
8
Frontiers in Genetics
230 papers in training set
Top 1%
2.8%
9
Nature Machine Intelligence
70 papers in training set
Top 1%
2.7%
10
PLOS Computational Biology
1863 papers in training set
Top 12%
2.5%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.0%
12
PLOS ONE
5266 papers in training set
Top 47%
1.8%
13
Computers in Biology and Medicine
128 papers in training set
Top 2%
1.7%
14
IEEE Access
35 papers in training set
Top 0.7%
1.7%
15
Scientific Reports
3612 papers in training set
Top 59%
1.5%
16
Journal of Computational Biology
48 papers in training set
Top 0.8%
1.2%
17
Bioinformatics Advances
203 papers in training set
Top 4%
1.2%
18
Nature Communications
5641 papers in training set
Top 50%
1.2%
19
Science China Life Sciences
29 papers in training set
Top 0.4%
1.1%
20
Computational Biology and Chemistry
28 papers in training set
Top 0.8%
1.0%
21
Neurocomputing
13 papers in training set
Top 0.3%
0.9%
22
Advanced Science
286 papers in training set
Top 10%
0.6%
23
Science Bulletin
21 papers in training set
Top 0.5%
0.6%
24
Patterns
78 papers in training set
Top 3%
0.6%