Back

TATAT: a containerized software for generating annotated coding transcriptomes from raw RNA-seq data

Brown, A. J.; Seifert, S. N.

2025-07-14 bioinformatics
10.1101/2025.07.09.663867 bioRxiv
Show abstract

MotivationMany transcriptome creation workflows are not standardized, are difficult to install or share, prone to breaking as dependencies update or cease to be maintained and are resource intensive. Due to a lack of authoritative literature, many also overlook potentially important steps, such as thinning contig over-assembly or identifying transcript consensus across samples, which reduce resource demands during annotation and increase the accuracy of final transcripts. ResultsWe developed TATAT, a modular, Dockerized software that contains all the tools necessary to generate an annotated coding transcriptome from raw RNA-seq data. The tools remain in a static state and can be coordinated with bash and python scripts provided therein, making TATAT a standardized, reproducible workflow that can easily be shared and installed. We preferentially incorporate tools that are not only accurate, but are fast and require less RAM, and subsequently show TATAT can generate a comprehensive transcriptome for a non-model organism, the Egyptian rousette bat (Rousettus aegyptiacus), in [~]8 hours in a high-performance computing (HPC) environment. Availability and implementationThe TATAT code, instructions, and tutorial are available at https://github.com/viralemergence/tatat.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.