Back

CORAL: Accurate annotation of compact genomes using long-read RNA-seq, demonstrated in Oikopleura dioica

Torres-Aguila, N. P.; Cassa, B.; Canestro, C.

2025-12-08 bioinformatics
10.64898/2025.12.04.692336 bioRxiv
Show abstract

The burst of high-quality genome assemblies has created an urgent need to improving genome annotation tools, especially for problematic species such as those with compact, fast-evolving genomes in which standard annotation tools underperform when facing short intergenic regions, overlapping UTRs, non-canonical splicing and widespread operons. We present CORAL (Compact-genome Oriented RNA-based Annotation using Long-reads), a long-read RNA-seq-based workflow for de novo annotation of compact eukaryotic genomes, together with GAMBA (Gene Aggregation tool for Multicistronic Block Annotation), a Rust-based module that detects polycistronic transcriptional units directly from GTF annotations and is integrated within CORAL. Using the exceptionally compact chordate genome of Oikopleura dioica as a case study, CORAL increases the number of annotated genes, improves gene-model completeness and reduces chimeric predictions relative to existing annotations, while GAMBA recovers thousands of operons, most of which are independently supported by splice-leader data. Together, CORAL and GAMBA provide an accurate, scalable framework for annotating compact, structurally atypical genomes using long-read transcriptomic data alone. Availability and ImplementationCORAL and GAMBA are open source and available at https://github.com/EvoDevoGenomics-UB/CORAL and https://github.com/nurie05/gamba-tool.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.