Back

Meta-PseU: A Meta-Classifier for Robust Prediction of RNA Pseudouridine Modification Sites from Long Sequences

Sutou, T.; HARUN-OR-ROSHID, M.; Kurata, H.

2025-12-11 bioinformatics
10.64898/2025.12.08.693080 bioRxiv
Show abstract

Pseudouridine ({Psi}) represents one of the most abundant and evolutionarily conserved RNA modifications. {Psi} provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of {Psi} sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, current machine-learning and deep-learning predictors suffer from limitations such as small datasets and limited generalizability. To overcome these issues, we have constructed new long-sequence datasets derived from RMBase 3.0 and developed Meta-PseU, a logistic regression-based meta-classifier that stacks multiple single-feature or baseline classifiers across three species of human, mouse, and yeast. Meta-PseU substantially reduced the performance gaps between training and independent test datasets, presenting superior generalization. Meta-PseU substantially outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. This work offers a new framework for robust {Psi}-site identification by using long sequences. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU.

Published in Computer Methods and Programs in Biomedicine · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.