miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA Frequency Class Bias
Sammut, S.; Gresova, K.; Tzimotoudis, D.; Marsalkova, E.; Cechak, D.; Alexiou, P.
Show abstract
MotivationMicroRNAs (miRNAs) are crucial regulators of gene expression, but the precise mechanisms governing their binding to target sites remain unclear. A major contributing factor to this is the lack of unbiased experimental datasets for training accurate prediction models. While recent experimental advances have provided numerous miRNA-target interactions, these are solely positive interactions. Generating negative examples in silico is challenging and prone to introducing biases, such as the miRNA frequency class bias identified in this work. Biases within datasets can compromise model generalization, leading models to learn dataset-specific artifacts rather than true biological patterns. ResultsWe introduce a novel methodology for negative sample generation that effectively mitigates the miRNA frequency class bias. Using this methodology, we curate several new, extensive datasets and benchmark several state-of-the-art methods on them. We find that a simple convolutional neural network model, retrained on some of these datasets, is able to outperform state-of-the-art methods. This highlights the potential for leveraging unbiased datasets to achieve improved performance in miRNA binding site prediction. To facilitate further research and lower the barrier to entry for machine learning researchers, we provide an easily accessible Python package, miRBench, for dataset retrieval, sequence encoding, and the execution of state-of-the-art models. AvailabilityThe miRBench Python Package is accessible at https://github.com/katarinagresova/miRBench/releases/tag/v1.0.0
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Advancing mRNA subcellular localization prediction with graph neural network and RNA structure 93%
- A Pseudo-Temporal Causality Approach to Identifying miRNA-mRNA Interactions During Biological Processes 93%
- DataRemix: a universal data transformation for optimal inference from gene expression datasets 93%
Similar papers in this journal
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 93%
- Identification of Monotonically Classifying Pairs of Genes for Ordinal Disease Outcomes 92%
- GeneSNAKE: a Python package for benchmarking and simulation of gene regulatory networks and perturbation-induced expression data 92%
Similar papers in this journal
- Identifying promoter sequence architectures via a chunking-based algorithm using non-negative matrix factorisation 92%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 92%
- Predicting Mean Ribosome Load for 5'UTR of any length using Deep Learning 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.