Back

LinearCapR: Linear-time computation of per-nucleotide structural-context probabilities of RNA without base-pair span limits

Otagaki, T.; Hosokawa, H.; Fukunaga, T.; Iwakiri, J.; Terai, G.; Asai, K.

2025-12-29 bioinformatics
10.64898/2025.12.26.696559 bioRxiv
Show abstract

MotivationRNA molecules adopt dynamic ensembles of secondary structures, where the local structural context of each nucleotide-such as whether it resides in a stem or a specific type of loop-strongly shapes molecular interactions and regulatory function. Structural-context probabilities therefore provide a more functionally informative view of RNA folding than the minimum free energy structures or base-pairing probabilities. However, existing tools either require O (N3) time or employ span-restricted approximations that omit long-range base-pairs, limiting their applicability to large and biologically important RNAs. ResultsWe introduce LinearCapR, enabling linear-time, span-unrestricted computation of structural-context marginalized probabilities, using beam-pruned Stochastic Context Free Grammar-based computation. LinearCapR retains global ensemble features lost by span-limited methods and yields superior predictive power on bpRNA-1m(90) dataset, especially for multiloops and exterior regions, as well as long-distance stems. LinearCapR supports analysis of long RNAs, demonstrated on the full genome of SARS-CoV-2. ConclusionsLinearCapR provides the first base-pair-span-unrestricted, linear-time framework for RNA structural-context analysis, retaining key thermodynamic ensemble features essential for functional interpretation. It enables large-scale studies of viral genomes, long non-coding RNAs, and downstream analyses such as RNA-binding protein site prediction. AvailabilityThe source code of LinearCapR is available at https://github.com/hoget157/LinearCapR.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.