Back

Asymmetric trichotomous data partitioning enables development of predictive machine learning models using limited siRNA efficacy datasets

Monopoli, K. R.; Korkin, D.; Khvorova, A.

2022-07-10 bioinformatics
10.1101/2022.07.08.499317 bioRxiv
Show abstract

Chemically modified small interfering RNAs (siRNAs) are promising therapeutics guiding sequence-specific silencing of disease genes. However, identifying chemically modified siRNA sequences that effectively silence target genes is a challenge. Such determinations necessitate computational algorithms. Machine Learning (ML) is a powerful predictive approach for tackling biological problems, but typically requires datasets significantly larger than most available siRNA datasets. Here, we describe a framework for applying ML to a small dataset (356 modified sequences) for siRNA efficacy prediction. To overcome noise and biological limitations in siRNA datasets, we apply a trichotomous (using two thresholds) partitioning approach, producing several combinations of classification threshold pairs. We then test the effects of different thresholds on random forest (RF) ML model performance using a novel evaluation metric accounting for class imbalances. We identify thresholds yielding a model with high predictive power outperforming a simple linear classification model generated from the same data. Using a novel method to extract model features, we observe target site base preferences consistent with current understanding of the siRNA-mediated silencing mechanism, with RF providing higher resolution than the linear model. This framework applies to any classification challenge involving small biological datasets, providing an opportunity to develop high-performing design algorithms for oligonucleotide therapies.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.