Back

A short read de novo transcriptome construction pipeline optimized by long reads reveals novel developmentally regulated gene isoforms and disease targets in hundreds of eye samples

Swamy, V. S.; Fufa, T. D.; Hufnagel, R. B.; McGaughey, D. M.

2020-09-22 genomics
10.1101/2020.08.21.261644 bioRxiv
Show abstract

De novo transcriptome construction from short-read RNA-seq is a common method for reconstructing mRNA transcripts within a given sample. However, the precision of this process is unclear as it is difficult to obtain a ground-truth measure of transcript expression. With advances in third generation sequencing, full length transcripts of whole transcriptomes can be accurately sequenced to generate a ground-truth transcriptome. We generated long-read PacBio and short-read Illumina RNA-seq data from a human induced pluripotent stem cell- derived retinal pigmented epithelium (iPSC-RPE) cell line. We use long-read data to identify simple metrics for assessing de novo transcriptome construction and optimize a short-read based de novo transcriptome construction pipeline. We apply this this pipeline to construct transcriptomes for 340 short-read RNA-seq samples originating from healthy adult and fetal human retina, cornea, and RPE. We identify hundreds of novel gene isoforms and examine their significance in the context of ocular development and disease.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.