The ENCODE4 long-read RNA-seq collection reveals distinct classes of transcript structure diversity
Reese, F.; Williams, B.; Balderrama-Gutierrez, G.; Wyman, D.; Celik, M. H.; Rebboah, E.; Rezaie, N.; Trout, D.; Razavi-Mohseni, M.; Jiang, Y.; Borsari, B.; Morabito, S.; Liang, H. Y.; McGill, C. J.; Rahmanian, S.; Sakr, J.; Jiang, S.; Zeng, W.; Carvalho, K.; Weimer, A. K.; Dionne, L. A.; McShane, A.; Bedi, K.; Elhajjajy, S. I.; Upchurch, S.; Jou, J.; Youngworth, I.; Gabdank, I.; Sud, P.; Jolanki, O.; Strattan, J. S.; Kagda, M. S.; Snyder, M. P.; Hitz, B. C.; Moore, J. E.; Weng, Z.; Bennett, D.; Reinholdt, L.; Ljungman, M.; Beer, M. A.; Gerstein, M. B.; Pachter, L.; Guigo, R.; Wold, B. J.; Mort
Show abstract
The majority of mammalian genes encode multiple transcript isoforms that result from differential promoter use, changes in exonic splicing, and alternative 3 end choice. Detecting and quantifying transcript isoforms across tissues, cell types, and species has been extremely challenging because transcripts are much longer than the short reads normally used for RNA-seq. By contrast, long-read RNA-seq (LR-RNA-seq) gives the complete structure of most transcripts. We sequenced 264 LR-RNA-seq PacBio libraries totaling over 1 billion circular consensus reads (CCS) for 81 unique human and mouse samples. We detect at least one full-length transcript from 87.7% of annotated human protein coding genes and a total of 200,000 full-length transcripts, 40% of which have novel exon junction chains. To capture and compute on the three sources of transcript structure diversity, we introduce a gene and transcript annotation framework that uses triplets representing the transcript start site, exon junction chain, and transcript end site of each transcript. Using triplets in a simplex representation demonstrates how promoter selection, splice pattern, and 3 processing are deployed across human tissues, with nearly half of multitranscript protein coding genes showing a clear bias toward one of the three diversity mechanisms. Evaluated across samples, the predominantly expressed transcript changes for 74% of protein coding genes. In evolution, the human and mouse transcriptomes are globally similar in types of transcript structure diversity, yet among individual orthologous gene pairs, more than half (57.8%) show substantial differences in mechanism of diversification in matching tissues. This initial large-scale survey of human and mouse long-read transcriptomes provides a foundation for further analyses of alternative transcript usage, and is complemented by short-read and microRNA data on the same samples and by epigenome data elsewhere in the ENCODE4 collection.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Long-read transcriptomics of a diverse human cohort reveals widespread ancestry bias in gene annotation 97%
- Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals 97%
- Widespread naturally variable human exons aid genetic interpretation 97%
Similar papers in this journal
- BamQuery: a proteogenomic tool for the genome-wide exploration of the immunopeptidome 96%
- scDALI: Modelling allelic heterogeneity of DNA accessibility in single-cells reveals context-specific genetic regulation 96%
- A read count-based method to detect multiplets and their cellular origins from snATAC-seq data 96%
Similar papers in this journal
Similar papers in this journal
- Transcriptional kinetics and molecular functions of long non-coding RNAs 97%
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 96%
- Dynamic network-guided CRISPRi screen reveals CTCF loop-constrained nonlinear enhancer-gene regulatory activity in cell state transitions 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.