Back

Population-scale interpretation of RNA isoform diversity enabled by Isopedia

Zheng, X.; Kronenberg, Z.; Garcia-Ruiz, S.; Layer, R. M.; Gustavsson, E. K.; Ryten, M.; Sedlazeck, F. J.

2026-03-25 bioinformatics
10.64898/2026.03.23.713667 bioRxiv
Show abstract

Alternative splicing generates extensive transcriptomic complexity, yet "novelty" is often inflated because of incomplete reference annotations, with 20-70% of transcripts in RNA-Seq studies labeled as novel. Isopedia provides an expandable data structure for reference-agnostic isoform annotation, which we demonstrate here through a population-scale catalog of 1,007 long-read datasets spanning 37 diverse biological contexts. By transitioning from reference-dependent to evidence-weighted annotation, Isopedia provides the frequency-based context necessary to distinguish stochastic noise from biologically active isoforms. In HG002 benchmarks, Isopedia reduced apparent isoform novelty by up to 26-fold, achieving a >95% annotation rate even for low-abundance isoforms typically missed by standard catalogs. The framework further supports systematic exploration of challenging loci such as pseudogenes and gene fusions. Isopedia transforms isoform discovery into a systematic interpretation of the human transcriptome, providing a critical foundation for clinical and functional RNA research. Isopedia is open source and freely available: https://github.com/zhengxinchang/isopedia.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.