Back

Restrander: rapid orientation and QC of long-read cDNA data

Schuster, J.; Ritchie, M. E.; Gouil, Q.

2023-05-03 bioinformatics
10.1101/2023.05.02.539165 bioRxiv
Show abstract

In transcriptomic analyses, it is helpful to keep track of the strand of the RNA molecules. However, the Oxford Nanopore long-read cDNA sequencing protocols generate reads that correspond to either the first or second-strand cDNA, therefore the strandedness of the initial transcript has to be inferred bioinformatically. Reverse transcription and PCR can also introduce artefacts which should be flagged in data pre-processing. Here we introduce Restrander, a lightning-fast and highly accurate tool for restranding and quality checking long-read cDNA sequencing data. Thanks to its C++ implementation, Restrander was faster than Oxford Nanopore Technologies existing tool Pychopper, and correctly restranded more reads due to its strategy of searching for polyA/T tails in addition to primer sequences from the reverse transcription and template-switch steps. We found that restranding improved the process of visualising and exploring data, and increased the number of novel isoforms discovered by bambu, particularly in regions where sense and antisense transcripts co-occur. The artefact detection implemented in Restrander quantifies reads which do not have the correct 5 and 3 ends, a feature which is useful in quality control for library preparation. Restrander is pre-configured for all major cDNA protocols, and can be customised with user-defined primers. Restrander is available at https://github.com/jakob-schuster/restrander

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.