TATAT: a containerized software for generating annotated coding transcriptomes from raw RNA-seq data
Brown, A. J.; Seifert, S. N.
Show abstract
MotivationMany transcriptome creation workflows are not standardized, are difficult to install or share, prone to breaking as dependencies update or cease to be maintained and are resource intensive. Due to a lack of authoritative literature, many also overlook potentially important steps, such as thinning contig over-assembly or identifying transcript consensus across samples, which reduce resource demands during annotation and increase the accuracy of final transcripts. ResultsWe developed TATAT, a modular, Dockerized software that contains all the tools necessary to generate an annotated coding transcriptome from raw RNA-seq data. The tools remain in a static state and can be coordinated with bash and python scripts provided therein, making TATAT a standardized, reproducible workflow that can easily be shared and installed. We preferentially incorporate tools that are not only accurate, but are fast and require less RAM, and subsequently show TATAT can generate a comprehensive transcriptome for a non-model organism, the Egyptian rousette bat (Rousettus aegyptiacus), in [~]8 hours in a high-performance computing (HPC) environment. Availability and implementationThe TATAT code, instructions, and tutorial are available at https://github.com/viralemergence/tatat.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparative evaluation of full-length isoform quantification from RNA-Seq 96%
- SPEAQeasy: a Scalable Pipeline for Expression Analysis and Quantification for R/Bioconductor-powered RNA-seq analyses 95%
- PIPETS: A statistically informed, gene-annotation agnostic analysis method to study bacterial termination using 3'-end sequencing. 95%
Similar papers in this journal
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 96%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
- tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction 95%
Similar papers in this journal
- Differential Expression Gene Explorer (DrEdGE): A tool for generating interactive online data visualizations for exploration of quantitative transcript abundance datasets 94%
- cloudrnaSPAdes: Isoform assembly using bulk barcoded RNA sequencing data 94%
- SplicingFactory - Splicing diversity analysis for transcriptome data 93%
Similar papers in this journal
Similar papers in this journal
- REVERSE: A user-friendly web server for analyzing next-generation sequencing data from in vitro selection/evolution experiments 95%
- TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes 94%
- PerturbAtlas: A Comprehensive Atlas of Public Genetic Perturbation Bulk RNA-seq Datasets 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.