Evidence of distal regulations orchestrated by RNAs initiating at short tandem repeats
Grapotte, M.; Vroland, C.; Garrido-Martin, D.; Vignoli, A.; Calero, L.; Bouvier, Q.; Robin, M.; Yip, C. W.; Carninci, P.; Chatelain, C.; Brehelin, L.; Notredame, C.; Guigo, R.; Lecellier, C. H.
Show abstract
Short Tandem Repeats (STRs), also called microsatellites, correspond to tandemly repeated short DNA motifs (1 to 6 bp) and are one of the most polymorphic and abundant repetitive elements in the human genome. Variations of their length (i.e. number of consecutive repeats) have been implicated in gene expression regulation (termed expression(e)STRs). Using captrapping followed by long read sequencing, we discovered that STRs can host transcription start sites, the presence of which depends mainly on STR flanking sequences. Here, we investigate the effect of SNPs located in these sequences and ask whether STR-initiating RNAs have regulatory potential. First, we develop fully interpretable deep learning-based models, called Modular Neural Networks, able to predict, for each STR class, the level of RNAs using 101bp-long sequences encompassing STRs. Analysis of MNN filters allows us to identify multiple regulatory elements and candidate transcription factors. Second, leveraging genome sequencing and gene expression data from the Genotype-Tissue Expression project, we use the output of our models to link the levels of STR-initiating RNAs to the expression of nearby genes. We identify 14,340 significant associations (coined RNA(r)STRs) and illustrate how this novel resource can help interpret non-coding variants associated with complex traits and diseases. Third, we unveil an intricate transcriptional interplay between STR-initiating RNAs and Alu repeats that may couple their regulatory actions, extending both the nature and the functional importance of non-coding transcription and shedding new light on the complexity of distal regulations orchestrated by repeated sequences.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification and analysis of splicing quantitative trait loci across multiple tissues in the human genome 97%
- Landscape of allele-specific transcription factor binding in the human genome 97%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 97%
Similar papers in this journal
- Integrated annotation and analysis of genomic features reveal new types of functional elements and large-scale epigenetic phenomena in the developing zebrafish 97%
- Linking candidate causal autoimmune variants to T cell networks using genetic and epigenetic screens in primary human T cells. 96%
- DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of enhancers 96%
Similar papers in this journal
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 96%
- Enhancer regulatory networks globally connect non-coding breast cancer loci to cancer genes 96%
Similar papers in this journal
- Recruitment of Homodimeric Proneural Factors by Conserved CAT-CAT E-Boxes Drives Major Epigenetic Reconfiguration in Cortical Neurogenesis 97%
- A high-resolution map of functional miR-181 response elements in the thymus reveals the role of coding sequence targeting and an alternative seed match 97%
- Extensive long-range polycomb interactions and weak compartmentalization are hallmarks of human neuronal 3D genome 97%
Similar papers in this journal
- Variant-resolved prediction of context-specific isoform variation with a graph-based attention model 97%
- Normal and cancer tissues are accurately characterised by intergenic transcription at RNA polymerase 2 binding sites 96%
- Comprehensive locus-specific L1 DNA methylation profiling reveals the epigenetic and transcriptional interplay between L1s and their integration sites. 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.