Mobile DNA
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match Mobile DNA's content profile, based on 31 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Halabian, R.; Storer, J. M.; Hoyt, S. J.; Hartley, G. A.; Brosius, J.; O'Neill, R. J.; Makalowski, W.
Show abstract
Long terminal repeats (LTRs) and non-LTRs retrotransposons, aka retroelements, collectively occupy a substantial part of the human genome. Certain non-LTR retroelements, such as L1 and SVA, have the potential for DNA transduction, which involves the concurrent mobilization of flanking non-transposon DNA during retrotransposition. These events can be detected by computational approaches. Despite being the most abundant short interspersed sequences (SINEs) that are still active within the genomes of humans and other primates, the transduction rate caused by Alu sequences remains unexplored. Therefore, we conducted an analysis to address this research gap and utilized an in-house program to probe for the presence of Alu-related transductions in the human genome. We analyzed 118,489 full-length AluY subfamilies annotated within the first complete human reference genome, T2T-CHM13. For comparative insights, we extended our exploration to two non-human primate genomes, the chimpanzee and the rhesus monkey. After manual curation, our findings did not confirm any Alu-mediated transductions, whose source genes are, unlike L1 or SVA, transcribed by RNA polymerase III, implying that they are infrequent or possibly absent not only in the human but also in chimpanzee and rhesus monkey genomes. Although we identified loci in which the 3 Target Site Duplication (TSD) was located distantly from the retrotransposed AluYs, a transduction hallmark, our study could not find further support for such events. The observation of these instances can be explained by the incorporation of other nucleotides into the poly(A) tails in conjunction with polymerase slippage.
Maiwald, S.; Maiwald, F.; Heitkam, T.
Show abstract
Plant genomes are filled with retrotransposons and their derivatives, subject to constant sequence turnover. As short, non-autonomous retrotransposons do not encode a protein product, they experience reduced selective constraints on their DNA sequence, leading to diversification into multiple families, usually limited to only a few species. This absence of any coding capacity and their tendency to form subfamilies are the reasons for the incomplete description of non-autonomous LTR retrotransposons in most to all genomic repeat annotations. Here, we focus on non-autonomous LTR retrotransposon identification. Are all of these sequences derivatives of easier-to-identify full-length elements? Or is there more variability, which is currently overlooked? For this, we capitalize on our comprehensive understanding of the TE landscape in sugar beet to assess the extent of the blind spot on non-autonomous LTR retrotransposons Here, we present a workflow to identify non-autonomous LTR retrotransposons without prior sequence information, retrieving more than 100 families within the sugar beet genome. We only include TEs without the ability for complete self mobilization. Spanning up to 15,000 bp, these non-autonomous families are often longer than expected and characterized by reshuffling and modular evolution. Most strikingly, only a few of these families are directly derived from autonomous partners, showing that there is a large, undiscovered TE variety in the non-autonomous TE fraction. We highlight that a large fraction of non-autonomous TEs wont be retrieved with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study is the first insight into the non-autonomous LTR retrotransposon landscape within a single genome and serves as an example to estimate the error in non-autonomous TE detection.
Ricci, M.; Peona, V.; Taccioli, C.
Show abstract
The presence in nature of closely related species showing drastic differences in lifespan and cancer incidence has recently increased the interest of the scientific community on these topics. In particular, the adaptations and genomic characteristics underlying the evolution of cancer-resistant and long-lived species have recently focused on the presence of alterations in the number of non-coding RNAs, on epigenetic regulation and, finally, on the activity of transposable elements (TEs). In this study, we compared the content and dynamics of TE activity in the genomes of four rodent and six bat species exhibiting different lifespans and cancer susceptibility. Mouse, rat and guinea pig (short-lived and cancer-prone organisms) were compared with the naked mole rat (Heterocephalus glaber) which is the rodent with the longest lifespan. The long-lived and cancer-resistant bats of the genera Myotis, Rhinolophus, Pteropus and Rousettus were instead compared with the Molossus, which is instead a short-lived and cancer-resistant organism. Analyzing the patterns of recent accumulations of TEs in the genome in these species, we found a strong suppression or negative selection to accumulation, of non-LTR retrotransposons in long-lived and cancer-resistant organisms. On the other hand, all short-lived and cancer-prone species have shown recent accumulation of this class of TEs. Among bats, the Molossus molossus turned out to be a very particular species and, at the same time, an important model because, despite being susceptible to rapid ageing, it is resistant to cancer. In particular, we found that its genome has the highest density of SINE (non-LTR retrotransposons), but, on the other hand, a total lack of active LINE retrotransposons. Our hypothesis is that the lack of LINEs presumably makes the Molossus cancer resistant due to lack of retrotransposition but, at the same time, the high presence of SINE, may be related to their short life span due to "sterile inflammation" and high mutation load. We suggest that research on ageing and cancer evolution should put particular attention to the involvement of non-LTR retrotransposons in these phenomena.
Hassan, N. T.; Adelson, D. L.; Galbraith, J. D.
Show abstract
Horizontal transfer of transposable elements (HTT) has been reported across many species and the impact of such events on genome structure and function has been well described. However, few studies have focused on reptilian genomes, especially HTT events in Testudines (turtles). Here, we investigated the repetitive content of Malaclemys terrapin terrapin (Diamondback turtle) and found a high similarity hAT-6 DNA transposon shared between other turtle species, ray-finned fishes, and a frog. hAT-6 was notably absent in taxa closely related to turtles, such as crocodiles and birds. Successful invasion of DNA transposons into new genomes requires the conservation of specific residues in the encoded transposase, and through structural analysis, these residues were identified indicating retention of functional transposition activity. We document a rare and recent HTT event of a DNA transposon between turtles which are known to have a low genomic evolutionary rate and ancient repeats.
Packiaraj, J.; Thakur, J.
Show abstract
Centromeres are essential for faithful chromosome segregation during mitosis and meiosis. However, the organization of satellite DNA and chromatin at mouse centromeres and pericentromeres is poorly understood due to the challenges of sequencing and assembling repetitive genomic regions. Using recently available PacBio long-read sequencing data from the C57BL/6 strain and chromatin profiling, we found that contrary to the previous reports of their highly homogeneous nature, centromeric and pericentromeric satellites display varied sequences and organization. We find that both centromeric minor satellites and pericentromeric major satellites exhibited sequence variations within and between arrays. While most arrays are continuous, a significant fraction is interspersed with non-satellite sequences, including transposable elements. Additionally, we investigated CENP-A and H3K9me3 chromatin organization at centromeres and pericentromeres using Chromatin immunoprecipitation sequencing (ChIP-seq). We found that the occupancy of CENP-A and H3K9me3 chromatin at centromeric and pericentric regions, respectively, is associated with increased sequence abundance and homogeneity at these regions. Furthermore, the transposable elements at centromeric regions are not part of functional centromeres as they lack CENP-A enrichment. Finally, we found that while H3K9me3 nucleosomes display a well-phased organization on major satellite arrays, CENP-A nucleosomes on minor satellite arrays lack phased organization. Interestingly, the homogeneous class of major satellites phase CENP-A and H3K27me3 nucleosomes as well, indicating that the nucleosome phasing is an inherent property of homogeneous major satellites. Overall, our findings reveal that house mouse centromeres and pericentromeres, which were previously thought to be highly homogenous, display significant diversity in satellite sequence, organization, and chromatin structure.
Li, Z.; Clavereau, I.; Pollet, N.
Show abstract
Hammerhead ribozymes (HHRs) are small catalytic RNAs found across diverse life forms. In animal genomes, they can be encoded by genes organised in dispersed copies or in tandem genomic arrangements. These tandemly organised forms, known as Non-LTR retrozymes, were recently identified as a distinct group of non-autonomous retrotransposons, likely mobilised via a rolling-circle transposition mechanism and potentially involved in host transcriptome regulation. However, their evolutionary origins remain poorly understood. Here, we investigate the presence, genomic distribution, and possible origins of Non-LTR retrozymes across a broad range of vertebrate species. We find that these elements display a patchy phylogenetic distribution, notably absent from the Aves and Mammalia lineages. In species where they are present, retrozyme copy number, consensus length, and monomer proportion vary widely across species and retrozyme families, suggesting diverse amplification dynamics. Genomic mapping reveals a significant enrichment of Non-LTR retrozymes in intergenic regions and their exclusion from introns and exons, indicating selective pressure against genic insertion. Strikingly, the phylogenetic distribution of Non-LTR retrozymes coincides with that of Penelope-like elements (PLEs). Phylogenetic analysis further shows that the pLTR region of PLEs is closely related to Non-LTR retrozymes, supporting the hypothesis that Non-LTR retrozymes are non-autonomous derivatives of PLEs. Together, our findings shed new light on the evolutionary origin and genomic behaviour of Non-LTR retrozymes and underscore their potential regulatory roles in vertebrate genomes.
Choi, J. D.; Del Pinto, L. A.; Sutter, N. B.
Show abstract
BackgroundMessenger RNA 3 untranslated regions (3UTRs) control many aspects of gene expression and determine where the transcript will terminate. The polyadenylation signal (PAS) AAUAAA is a key regulator of transcript termination and this hexamer, or a similar sequence, is very frequently found within 30 bp of 3UTR ends. Short interspersed element (SINE) retrotransposons are found throughout genomes in high copy number. When inserted into genes they can disrupt expression, alter splicing, or cause nuclear retention of mRNAs. The genomes of the domestic dog and other carnivores carry hundreds of thousands Can-SINEs, a tRNA-related SINE with transcription termination potential. Because of this we asked whether Can-SINEs may help terminate transcript in some dog genes. ResultsDog 3UTRs have several peaks of AATAAA PAS frequency within 40 bp of the 3UTR end, including four bp-interval peaks at 28, 32, and 36 bp from the end. The periodicity is partly explained by TAAA(n) repeats within Can-SINE AT-rich tails. While density of antisense-oriented Can-SINEs in 3UTRs is fairly constant with distances from 3end, sense-oriented Can-SINEs are common at the 3end but nearly absent farther upstream. There are nine Can-SINE sub-types in the dog genome and the consensus sequence sense strands (head to tail) all carry at least three PASs while antisense strands usually have none. We annotated all repeat-masked Can-SINE copies in the Boxer reference genome and found that the young SINEC_Cf type has a mode of 15 bp for target site duplications (TSDs). We find that all Can-SINE types favor integration at TSDs beginning with A(4). The count of AATAAA PASs differs significantly between sense and antisense-oriented retrotransposons in transcripts. Can-SINEs near 3UTR ends are very likely to carry AATAAA on the mRNA sense strand while those farther upstream are not. We also identified loci where Can-SINE insertion has truncated or altered a dog 3UTR compared to the human ortholog. ConclusionDog Can-SINE activity has imported AATAAA PASs into gene transcripts and led to alteration of 3UTRs. AATAAA sequences are selectively removed from Can-SINEs in introns and upstream 3UTR regions but are retained at the far downstream end of 3UTRs, which we infer reflects their role as termination sequences for these transcripts.
Mercuri, R. L. V.; Miller, T. L. A.; dos Santos, F. F.; de Lima, M. F.; Rangel-Pozzo, A.; Galante, P. A. F.
Show abstract
BackgroundTransposable elements (TEs) constitute a significant portion of mammalian genomes, accounting for about 50% of the total DNA. Intragenic TEs are of particular interest as they are co-transcribed with their host genes in pre-mRNA, potentially leading to the formation of novel chimeric transcripts and the exonization of TEs. The abundance of RNA sequencing data currently available offers a unique opportunity to explore transcriptomic variations. However, a significant limitation is the capability of existing computational tools. Here, we introduce FREDDIE, an innovative algorithm designed to detect the exonization of retrotransposable elements using RNA-seq data. FREDDIE can process short and long RNA sequencing data, assemble and quantify transcripts, evaluate coding potential, and identify protein domains in chimeric transcripts involving exonized TEs and retrocopies. ResultsTo demonstrate the efficacy of FREDDIE, we analyzed and validated TE exonization in two human cancer cell lines, K562 and U251. We have identified 322 chimeric transcripts, of which 126 were from K562, and 196 were from U251. Among these chimeric transcripts, there were 35 that showed similar exonization patterns and host genes. These transcripts involve protein-coding genes of the host and exonization of LINE-1 (L1), Alu elements, and retrocopies of coding genes. We have selected some candidates and validated them experimentally through RT-PCR. The validation rate for these candidates was 70%, later confirmed by long-read sequencing. Additionally, we applied FREDDIE to analyze TE exonization across 157 glioblastoma samples, identifying 1,010 chimeric transcripts. The majority of these transcripts involved the exonization of Alu elements (69.8%), followed by L1 (20.6%) and retrocopies (9.6%). Notably, we discovered a highly expressed L1 exonization within the ROS gene, resulting in a truncated open reading frame (ORF) with the deletion of two protein domains. ConclusionsFREDDIE is an efficient and user-friendly tool for identifying chimeric transcripts that involve exonization of intragenic TEs. Overall, FREDDIE enables comprehensive investigations into the contributions of TEs to transcriptome evolution, variation, and disease-associated abnormalities, and it operates effectively on standard computing systems. FREDDIE is publicly available: https://github.com/galantelab/freddie
Moreira Mombach, D.; Mendez-Dorantes, C.; Mercuri, R. L. V.; Schofield, P.; Soares Baal, S. C.; Poersch, M. A.; Burns, K. H.; Carvalho de Oliveira, J.; Loreto, E. L. S.; Galante, P. A. F.
Show abstract
BackgroundTriple-negative breast cancer (TNBC) is an aggressive subtype with limited therapeutic options. While PARP inhibitors, such as olaparib, show promise in BRCA1-deficient TNBC through synthetic lethality, up to 50% of patients fail to respond, highlighting the need to understand the molecular mechanisms underlying PARP inhibitors efficacy. Transposable elements (TEs), particularly LINE-1 elements, are increasingly recognized as modulators of genomic instability associated with DNA repair processes and potential key players in synthetic lethality. Here, we investigate the functional relationship between TE activity and olaparib treatment in TNBC with distinct BRCA1 functional status. MethodsWe performed comprehensive multi-OMICs analysis of four TNBC cell lines (two BRCA1-deficient: SUM1315 and MDA-MB-436; two BRCA1-proficient: MDA-MB-468 and BT549) treated with olaparib. We analyzed expression and differential expression of protein-coding genes, TEs, and gene-TE chimeric transcripts. Long-read whole-genome sequencing was employed to detect de novo TE insertions, complemented by a functional assay to quantify LINE-1 retrotransposition activity in olaparib-treated cells. ResultsOlaparib treatment induces extensive transcriptomic and genomic disorganization mediated by TEs, especially LINE-1, exclusively in BRCA1-deficient cells. We observed aberrant overexpression of both genes and TEs, including gene-TE chimeric transcripts harboring poison exons within tumorigenic genes and multi-exonic TE-TE chimeras capable of forming immunostimulatory double-stranded RNA (dsRNA) structures. Functional enrichment analyses revealed activation of antiviral immune pathways linked to LINE-1 activity. Consistently, orthogonal assays confirmed LINE-1 retrotransposition in BRCA1-deficient cells following olaparib exposure. ConclusionsOur findings demonstrate that olaparib treatment induces TE activation especially in BRCA1-deficient cells, a novel mechanism that may underlie synthetic lethality in TNBC. This TE activation triggers immune responses and genomic instability, providing new therapeutic opportunities through immunotherapy combinations and suggesting that TE activity may serve as a potential biomarker for treatment stratification of TNBC. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=119 SRC="FIGDIR/small/738694v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@e8b202org.highwire.dtl.DTLVardef@fea1c5org.highwire.dtl.DTLVardef@12ea397org.highwire.dtl.DTLVardef@f60f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Ikemoto, K.; Fujimoto, H.; Fujimoto, A.
Show abstract
BackgroundLong-read sequencing technologies have the potential to overcome the limitations of short reads and provide a comprehensive picture of the human genome. However, it remains hard to characterize repetitive sequences by reconstructing genomic structures at high resolution solely from long reads. Here, we developed a localized assembly method (LoMA) that constructs highly accurate consensus sequences (CSs) from long reads. MethodsWe first developed LoMA, by combining minimap2, MAFFT, and our algorithm, which classifies diploid haplotypes based on structural variants and constructs CSs. Using this tool, we analyzed two human samples (NA18943 and NA19240) sequenced with the Oxford Nanopore sequencer. We defined target regions in each genome based on mapping patterns and then constructed a high-quality catalog of the human insertion solely from the long-read data. ResultsThe assessment of LoMA showed high accuracy of CSs (error rate < 0.3%) compared with raw data (error rate > 8%) and superiority to the previous study. The genome-wide analysis of NA18943 and NA19240 identified 5,516 and 6,542 insertions ({zeta} 100 bp) respectively. Most insertions ([~]80%) were derived from the tandem repeat and transposable elements. We also detected processed pseudogenes, insertions in transposable elements, and long insertions (> 10 kbp). Further, our analysis suggested that short tandem duplications were association with gene expression and transposons. ConclusionsOur analysis showed that LoMA constructs high-quality sequences from long reads with substantial errors. This study revealed the true structures of insertions with high accuracy and inferred mechanisms for the insertions. Our approach contributes to the future human genome studies. LoMA is available at our GitHub page: https://github.com/kolikem/loma.
Signor, S.; Vedanayagam, J.; Kim, B. Y.; Wierzbicki, F.; Kofler, R.; Lai, E. C.
Show abstract
Effective suppression of transposable elements (TEs) is paramount to maintain genomic integrity and organismal fitness. In D. melanogaster, flamenco is a master suppressor of TEs, preventing their movement from somatic ovarian support cells to the germline. It is transcribed by Pol II as a long (100s of kb), single-stranded, primary transcript, that is metabolized into Piwi-interacting RNAs (piRNAs) that target active TEs via antisense complementarity. flamenco is thought to operate as a trap, owing to its high content of recent horizontally transferred TEs that are enriched in antisense orientation. Using newly-generated long read genome data, which is critical for accurate assembly of repetitive sequences, we find that flamenco has undergone radical transformations in sequence content and even copy number across simulans clade Drosophilid species. D. simulans flamenco has duplicated and diverged, and neither copy exhibits synteny with D. melanogaster beyond the core promoter. Moreover, flamenco organization is highly variable across D. simulans individuals. Next, we find that D. simulans and D. mauritiana flamenco display signatures of a dual-stranded cluster, with ping-pong signals in the testis and/or embryo. This is accompanied by increased copy numbers of germline TEs, consistent with these regions operating as functional dual stranded clusters. Overall, the physical and functional diversity of flamenco orthologs is testament to the extremely dynamic consequences of TE arms races on genome organization, not only amongst highly related species, but even amongst individuals.
Montserrat-Ayuso, T.; Pujol, A.; Esteve-Codina, A.
Show abstract
Human endogenous retroviruses (HERVs), remnants of ancient retroviral infections, account for over 8% of the human genome and remain an underexplored source of cis-regulatory elements and protein-coding remnants. Here we present HERVarium, an integrated and interactive database that provides access to systematic annotations of the protein-domain architecture of internal HERV regions and the regulatory landscape of their LTRs across the human genome. Building on our previous catalog of conserved retroviral domains, we classified over 400,000 LTRs by structure (solo, 5', or 3' LTR), proximity to transcription start sites (TSS), potential transcription-factor binding-motif (TFBM) burden, and computationally reconstructed their canonical U3-R-U5 substructure, revealing that two-thirds retain recognizable segmentation, particularly those adjacent to conserved internal domains. LTRs flanking internal regions with conserved Gag, Pol, or Env domains tend to be longer and richer in motifs, suggesting coordinated maintenance of coding and regulatory potential. Conversely, solo LTRs positioned at gene TSSs exhibit promoter-like architectures enriched in motifs for developmental and proliferative regulators, while selectively lacking motifs associated with neuronal differentiation. Some 5' LTRs at lncRNA TSSs also flank conserved retroviral domains, indicating that certain lncRNAs may derive from transcriptionally active, structurally intact HERV loci. HERVarium provides a comprehensive, interactive resource to explore and download these annotations at single-locus resolution.
Tang, W.; Liang, P.
Show abstract
Mobile elements (MEs) can be divided into two major classes based on their transposition mechanisms as retrotransposons and DNA transposons. DNA transposons move in the genomes directly in the form of DNA in a cut-and-paste style, while retrotransposons utilize an RNA-intermediate to transpose in a "copy-and-paste" fashion. In addition to the target site duplications (TSDs), a hallmark of transposition shared by both classes, the DNA transposons also carry terminal inverted repeats (TIRs). DNA transposons constitute ~3% of primate genomes and they are thought to be inactive in the recent primate genomes since ~37My ago despite their success during early primate evolution. Retrotransposons can be further divided into Long Terminal Repeat retrotransposons (LTRs), which are characterized by the presence of LTRs at the two ends, and non-LTRs, which lack LTRs. In the primate genomes, LTRs constitute ~9% of genomes and have a low level of ongoing activity, while non-LTR retrotransposons represent the major types of MEs, contributing to ~37% of the genomes with some members being very young and currently active in retrotransposition. The four known types of non-LTR retrotransposons include LINEs, SINEs, SVAs, and processed pseudogenes, all characterized by the presence of a polyA tail and TSDs, which mostly range from 8 to 15 bp in length. All non-LTR retrotransposons are known to utilize the L1-based target-primed reverse transcription (TPRT) machineries for retrotransposition. In this study, we report a new type of non-LTR retrotransposon, which we named as retro-DNAs, to represent DNA transposons by sequence but non-LTR retrotransposons by the transposition mechanism in the recent primate genomes. By using a bioinformatics comparative genomics approach, we identified a total of 1,750 retro-DNAs, which represent 748 unique insertion events in the human genome and nine non-human primate genomes from the ape and monkey groups. These retro-DNAs, mostly as fragments of full-length DNA transposons, carry no TIRs but longer TSDs with ~23.5% also carrying a polyA tail and with their insertion site motifs and TSD length pattern characteristic of non-LTR retrotransposons. These features suggest that these retro-DNAs are DNA transposon sequences likely mobilized by the TPRT mechanism. Further, at least 40% of these retro-DNAs locate to genic regions, presenting significant potentials for impacting gene function. More interestingly, some retro-DNAs, as well as their parent sites, show certain levels of current transcriptional expression, suggesting that they have the potential to create more retro-DNAs in the current primate genomes. The identification of retro-DNAs, despite small in number, reveals a new mechanism in propagating the DNA transposons sequences in the primate genomes with the absence of canonical DNA transposon activity. It also suggests that the L1 TPRT machinery may have the ability to retrotranspose a wider variety of DNA sequences than what we currently know.
Gozashti, L.
Show abstract
The genomic landscape of transposable elements (TEs) varies dramatically across species, with some TEs demonstrating greater success in colonizing particular lineages than others. In mammals, LINE retrotransposons typically occupy more of the genome than any other TE and most LINE content is represented by a single family: L1. Here, we report an unusual genomic landscape of TEs in the deer mouse, Peromyscus maniculatus, a model for studying the genomic basis of adaptation. In contrast to other previously examined mammalian species, LTR elements occupy more of the deer mouse genome than LINEs (11% and 10% respectively). This pattern reflects a combination of relatively low LINE activity in addition to a massive invasion of lineage-specific endogenous retroviruses (ERVs). Deer mouse ERVs exhibit diverse origins spanning the retroviral phylogeny suggesting that these rodents have been host to a wide range of exogenous retroviruses. Notably, we were able to trace the origin of one ERV lineage, which arose within the last [~]11-18 million years, to a close relative of feline leukemia virus, revealing inter-ordinal horizontal transmission of these zoonotic viruses. Several lineage-specific ERV subfamilies have attained very high copy numbers, with the top five most abundant accounting for [~]2% of the genome. Concomitant to the expansive diversification of ERVs, we also observe a massive expansion of Kruppel-associated box domain-containing zinc finger genes (KZNFs), which likely control ERV activity and whose expansion may have been partially facilitated by ectopic recombination between ERVs. We also find evidence that ERVs directly impacted the evolutionary trajectory of LINEs by outcompeting them for genomic sites and frequently disrupting autonomous LINE copies. Together, our results illuminate the genomic ecology that shaped the deer mouse genomes TE landscape, opening up a range of opportunities to investigate the evolutionary processes that give rise to variation in mammalian genome structure. SummaryTransposable elements (TEs) are a highly diverse collection of genetic elements capable of mobilizing in genomes and function as important drivers of genome evolution. The landscape of TEs in a genome have been compared to a genomic ecosystem, with interactions between TEs and each other as well as TEs and their host, dictating the evolutionary success of TE lineages. While TE diversity and copy numbers can vary dramatically across taxa, the evolutionary reasons for this variation remain poorly understood. In mammals, long interspersed nuclear elements (LINEs) typically dominate, occupying more of the genome than any other TE. Here, we report a unique case in the deer mouse (Peromyscus maniculatus) in which long terminal repeat (LTR) retrotransposons occupy more of the genome than LINEs. We investigate the evolutionary origins and implications of the deer mouses distinct genomic landscape, revealing ecological processes that helped shape its evolution. Together, our results provide much-needed insight into the evolutionary processes that give rise to variation in mammalian genome structure.
Sahlstrom, H. M.; Datsomor, A. K.; Monsen, O.; Hvidsten, T. R.; Sandve, S. R.
Show abstract
BackgroundTransposable elements (TEs) are hypothesized to play important roles in shaping genome evolution following whole genome duplications (WGD), including rewiring of gene regulation. In a recent analysis, duplicate gene copies that had evolved higher expression in liver following the salmonid WGD ~100 million years ago were associated with higher numbers of predicted TE-derived cis-regulatory elements (TE-CREs). Yet, the ability of these TE-CREs to recruit transcription factors (TFs) in vivo and impact gene expression remains unknown. ResultsHere, we evaluated the gene regulatory functions of 11 TEs using luciferase promoter reporter assays in Atlantic salmon (Salmo salar) primary liver cells. Canonical Tc1-Mariner elements from intronic regions showed no or small repressive effects on transcription. However, other TE-derived cis-regulatory elements upstream of transcriptional start sites increased expression significantly. ConclusionOur results question the hypothesis that TEs in the Tc1-Mariner superfamily, which were extremely active following WGD in salmonids, had a major impact on regulatory rewiring of gene duplicates, but highlights the potential of other TEs in post-WGD rewiring of gene regulation in the Atlantic salmon genome.
Hill, T.
Show abstract
BackgroundThe evolutionary dynamics of transposable elements (TEs) vary across the tree of life and even between closely related species with similar ecologies. In Drosophila, most of the focus on TE dynamics has been completed in Drosophila melanogaster and the overall pattern indicates that TEs show an excess of low frequency insertions, consistent with their frequent turn over and high fitness cost in the genome. Outside of D. melanogaster, insertions in the species Drosophila algonquin, suggests that this situation may not be universal, even within Drosophila. Here we test whether the pattern observed in D. melanogaster is similar across five Drosophila species that share a common ancestor more than fifty million years ago.\n\nResultsFor the most part, TE family and order insertion frequency patterns are broadly conserved between species, supporting the idea that TEs have invaded species recently, are mostly costly and dynamics are conserved in orthologous regions of the host genome\n\nConclusionsMost TEs retain similar activities and fitness costs across the Drosophila phylogeny, suggesting little evidence of drift in the dynamics of TEs across the phylogeny, and that most TEs have invaded species recently.
Hector Rosche-Flores, H.; Fischer, S.; Picard, C. J.
Show abstract
BackgroundThe black soldier fly (Hermetia illucens) is an emerging model for bioconversion and industrial rearing. Its genome is highly repetitive, yet the contribution of transposable elements (TEs) to population divergence and demographic processes. The sampled populations represent a gradient of demographic histories, including wild and near-wild North American populations, and domesticated European strains with shared industrial origins. Difference in TE composition may influence genome structure, regulatory variation, and evolutionary responses to captive environments. ResultsA comparative analysis of the repetitive landscape was done for four H. illucens genomes, one of which is a wild-caught specimen. Total repeat content was high across all assemblies (67.6% to 70.8%) and dominated by LINE elements. Class-level TE diversity was nearly identical among genomes, but multiple DNA transposon families showed distinct lineage-specific differences. Large families including Maverick and Academ were generally depleted relative to the wild sample. Divergence profiles revealed patterns consistent with recent turnover in several families. Family level turnover, rather than class level change, accounted for the most difference among the genomes. TE-associated structural variants (TESVs) were also not uniformly distributed. Most chromosomes showed mid-chromosome enrichment, and a pronounced TESV peak on chromosome 5 overlapped a histone rich region containing many unclassified repeats. Use of a repeat library derived from multiple genomes increased the number of detected TESVs and improved classification within complex regions, demonstrating that multi-genome libraries enhance annotation accuracy compared to single reference-based models. ConclusionsMultiple DNA transposon families show evidence of recent or lineage-specific amplification in H. illucens, suggesting that TE amplification contributes to genome variation during demography-associated TE turnover. The multi-genome-based library improved TE detection and classification, providing a proof of concept that even a small lineage-inclusive repeat library enhances annotation accuracy and capture TE diversity missed by single-reference approaches. Together, these findings demonstrate that TE family turnover plays a significant role in shaping genome architecture and adaptation in this species.
Saylor, B.; Kremer, S. C.; Gregory, T. R.; Cottenie, K.
Show abstract
BackgroundDespite decades of research the factors that cause differences in transposable element (TE) distribution and abundance within and between genomes are still unclear. Transposon Ecology is a new field of TE research that promises to aid our understanding of this often-large part of the genome by treating TEs as species within their genomic environment, allowing the use of methods from ecology on genomic TE data. Community ecology methods are particularly well suited for application to TEs as they commonly ask questions about how diversity and abundance of a community of species is determined by the local environment of that community.\n\nResultsUsing a redundancy analysis, we found that ~ 50% of the TEs within a diverse set of genomes are distributed in a predictable pattern along the chromosome, and the specific TE superfamilies that show these patterns are relate to the phylogeny of the host taxa. In a more focused analysis, we found that ~60% of the variation in the TE community within the human genome is explained by its location along the chromosome, and of that variation two thirds (~40% total) was explained by the 3D location of that TE community within the genome (i.e. what other strands of DNA physically close in the nucleus). Of the variation explained by 3D location half (20% total) was explained by the type of regulatory environment (sub compartment) that TE community was located in. Using an analysis to find indicator species, we found that some TEs could be used as predictors of the environment (sub compartment type) in which they were found; however, this relationship did not hold across different chromosomes.\n\nConclusionsThese analyses demonstrated that TEs are non-randomly distributed across many diverse genomes and were able to identify the specific TE superfamilies that were non-randomly distributed in each genome. Furthermore, going beyond the one-dimensional representation of the genome as a linear sequence was important to understand TE patterns within the genome. Additionally, we extended the utility of traditional community ecology methods in analyzing patterns of TE diversity.
Volakhava, A.; Pavlova, S.; Radova, L.; Tausova, K.; Svozilova, H.; Zenatova, M.; Doubek, M.; Mamedov, I.; Pospisilova, S.; Plevova, K.
Show abstract
Retroelements (RE), particularly autonomous Long Interspersed Element-1 (LINE-1), function as potent drivers of genomic instability in various malignancies. While normally silenced by epigenetic mechanisms, their reactivation in cancer cells can drive tumor evolution. TP53 is known to repress LINE-1 transcription; however, the consequences of TP53 dysfunction on retrotransposition and transcriptional activity in chronic lymphocytic leukemia (CLL) remain unknown. To investigate the relationship between TP53 status, LINE-1 retrotranspositional potential, and transcriptomic alterations, we utilized a multi-omics approach, combining a highly sensitive NGS protocol for detecting novel LINE-1 insertions with transcriptomic profiling of transposable elements and protein-coding genes. We applied these methods to a cohort of CLL patients stratified by the presence or absence of TP53 clonal evolution and to CLL-derived cell lines MEC1 and HG3, including CRISPR/Cas9-engineered TP53 mutants. Genomic analysis revealed no evidence of widespread somatic retrotransposition, suggesting that CLL exhibits resistance against de novo LINE-1 insertions. Conversely, transcriptomic profiling uncovered distinct transposon expression signatures aligned with patterns of TP53 mutation status evolution. Notably, differentially expressed genes were significantly enriched in the RNA splicing pathway, indicating that while LINE-1 elements remain largely constrained at the genomic level, their transcriptomic activity may influence cellular rewiring, affecting patterns of TP53 mutation clonal evolution. Based on these results, we argue that the pathogenic contribution of retroelements in CLL lies in transcriptomic dysregulation and splicing alterations, rather than in direct DNA damage caused by LINE-1 insertions.
Rogers, S. O.; Bendich, A. J.
Show abstract
Introns and transposons exhibit many similar features, but the connections between them have yet to be firmly established. Group I introns have commonalities with DNA transposons, while group II introns share many features with retrotransposons. Here, we report the results of an analysis of 214 introns (including group I, group II, group III, twintrons, spliceosomal, and archaeal introns) from members of seven major taxa (within Eukarya, Bacteria, and Archaea) that all have direct repeats at or near both exon/intron borders, indicating that they were inserted via transposition events. Border sequence analysis indicates that after splicing, most mature transcripts would be functionally compromised because they do not restore the DNA sequence information before intron insertion. Transposons and introns thus appear to be members of a diverse assemblage of parasitic mobile genetic elements that secondarily may benefit their host cell and have expanded greatly in eukaryotes from their presumed prokaryotic ancestors. Author SummaryIntrons are found in all domains of life. While they are limited in prokaryotes, they have greatly expanded in number and diversity in eukaryotes. We found direct repeat sequences at or near both exon/intron borders for all 214 introns analyzed among eukaryotes, bacteria, and archaea. We infer that all introns were inserted into genes via transposon-like mechanisms and are members of a large family of mobile genetic elements.