CLN3 transcript complexity revealed by long-read RNA sequencing analysis
Zhang, H.-Y.; Minnis, C.; Gustavsson, E. K.; Ryten, M.; Mole, S. E.
Show abstract
BackgroundBatten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common mutation shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the "1-kb" deletion: the "major" and "minor" transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. MethodsWe leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. ResultsWe found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated "major" transcripts are detected. Together, they have median usage of 1.51% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. ConclusionOverall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the "1-kb" deletion and rare mutations on CLN3 transcription and disease pathogenesis.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Familial ALS/FTD-associated RNA-Binding deficient TDP-43 mutants cause neuronal and synaptic transcript dysregulation in vitro 94%
- A KLHL40 3’ UTR splice-altering variant causes milder NEM8, an under-appreciated disease mechanism 94%
- Redefining the PTEN Promoter: Identification of Two Upstream Transcription Start Regions 94%
Similar papers in this journal
- Expert Curation of the Human and Mouse Olfactory Receptor Gene Repertoires Identifies Conserved Coding Regions Split Across Two Exons 95%
- Mutational Constraint Analysis Workflow for Overlapping Short Open Reading Frames and Genomic Neighbours 93%
- Cataloging the potential functional diversity of Cacna1e splice variants using long-read sequencing 93%
Similar papers in this journal
- Enhancing the annotation of small ORF-altering variants using MORFEE: introducing MORFEEdb, a comprehensive catalog of SNVs affecting upstream ORFs in human 5'UTRs 95%
- A comprehensive atlas of fetal splicing patterns in the brain of adult myotonic dystrophy type 1 patients 95%
- Covering all your bases: incorporating intron signal from RNA-seq data 94%
Similar papers in this journal
- Common tissue-specific expressions and regulatory mechanisms of c-KIT isoforms with and without GNNK and GNSK sequences across five mammals 95%
- Single cell RNA sequencing of nc886, a non-coding RNA transcribed by RNA polymerase III, with a primer spike-in strategy 93%
- Formation of human long intergenic non-coding RNA genes and pseudogenes: ancestral sequences are key players 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.