Genome Biology and Evolution
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match Genome Biology and Evolution's content profile, based on 338 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Stuart, A. J.; Du, Z.; Hassan, N. T.; Adelson, D. L.
Show abstract
Retrotransposons are mobile, repetitive DNA sequences that are ubiquitous across eukaryotes and widely recognised as key drivers of both gene and genome evolution. The CR1 group of retrotransposons is thought to have been present in the most recent common ancestor of chordates [~]560 mya, and is the dominant retrotransposon in the majority of chordate species. The advent of long-read sequencing technologies has enabled the assembly of high-quality genomes from representatives of almost all major chordate orders, enabling comparative analysis with deeply divergent species. To better understand the composition of CR1-group elements (CGEs) in chordates, we systematically characterised full-length, recently active transposable elements across representative species from every available extant order of chordate. Our analysis uncovered previously unknown phylogenetic relationships of CGEs within and between species and has pushed back the origin of certain CR1-group subclades by tens of millions of years. Additionally, entirely novel elements with no close relatives in existing databases were uncovered within several of the species analysed. We also detected numerous putative horizontal transfer events, many of which had not been previously documented. Overall, this investigation has provided the first chordate-wide analysis of an element that is historically understudied yet plays a pivotal role in genome biology and evolution.
Roy, S.; Chu, D.; Singh, S.
Show abstract
Histone variants are paralogs that replace canonical histones in nucleosomes, often imparting novel functions. Despite their importance, how histone variants arise and evolve is poorly understood. Reconstruction of histone protein evolution is challenging due to high amino acid conservation and large differences in evolutionary rates across gene lineages and sites. Here we combined amino acid sequences and intron position data from 108 nematode genomes to trace the evolutionary histories of the three H2A variants found in Caenorhabditis elegans: the ancient H2A.ZHTZ-1, the sperm-specific HTAS-1, and HIS-35, which differs from canonical H2A by a single glycine-to-alanine C-terminal change. We find disparate evolutionary histories. Although the H2A.ZHTZ-1 protein is highly conserved, its gene exhibits recurrent intron gain and loss. This pattern suggests that it is intron presence, rather than specific intron sequences or positions, that may be important to H2A.Z functionality. In contrast, for HTAS-1 and HIS-35, we find variant-specific intron positions that are conserved across species. HIS-35 arose in the ancestor of Caenorhabditis and its sister group, including the genus Diploscapter, while the sperm-specific variant HTAS-1 arose more recently in the ancestor of a subset of Caenorhabditis species. HIS-35 exhibits gene retention in some descendent lineages but also recurrent gene loss in others, suggesting that histone variant use or functionality is highly flexible in this case. We also find that the single amino acid differentiating HIS-35 from core H2A is ancestral and common across canonical Caenorhabditis H2A sequences and identify one nematode species that bear identical HIS-35 and canonical H2A proteins, findings that are not predicted from the hypothesis that HIS-35 has a distinct function. Instead, we speculate that HIS-35 enables H2A expression across the cell cycle or in distinct tissues; genes encoding such partially-redundant functions may be advantageous yet relatively replaceable over evolutionary times, consistent with the patchwork pattern of retention and loss of both genes. Our study shows the evolutionary trajectory for histone H2A variants with distinct functions and the utility of intron positions for reconstructing the evolutionary history of gene families, particularly those undergoing idiosyncratic sequence evolution.
Eyre-Walker, Y. C.; Conradsen, C.; Vos, M.; Eyre-Walker, A.
Show abstract
Bacterial genomes often contain many genes that are only present in a subset of strains, the so-called accessory genes. Whether these genes are adaptive, neutral or deleterious remains contentious. Here we introduce a simple test to differentiate between these possibilities. If an accessory gene is adaptive then the sequence of the gene should be conserved, and the ratio of non-synonymous to synonymous diversity,{pi} n/{pi}s, should be less than one. In contrast, if the gene is neutral or deleterious, selection should not conserve the gene sequence, and{pi} n/{pi}s should equal one. We apply this test to accessory genes in Escherichia coli and Staphylococcus aureus; two highly divergent bacterial species with a large and a small pangenome respectively. We find{pi} n/{pi}s<1 for genes at all frequencies in both species demonstrating that many are adaptive. We estimate that at least 75% of all the accessory genes are maintained by selection in the two samples of 500 genomes that we have analysed, equating to thousands of adaptive accessory genes in both species, a substantial increase on previous estimates.
Schultz, D. T.; Heath-Heckman, E. A. C.; Winchell, C. J.; Kuo, D.-H. T.; Yu, Y.-s.; Oberauer, F.; Kocot, K.; Cho, S.-J.; Simakov, O.; Weisblat, D. A.
Show abstract
Comparisons of multiple metazoan genomes have revealed the existence of ancestral linkage groups (ALGs), genomic scaffolds sharing sets of orthologous genes that have been inherited from ancestral animals for hundreds of millions of years (Simakov et al. 2022; Schultz et al. 2023) These ALGs have persisted across major animal taxa including Cnidaria, Deuterostomia, Ecdysozoa and Spiralia. Notwithstanding this general trend of chromosome-scale conservation, ALGs have been obliterated by extensive genome rearrangements in certain groups, most notably including Clitellata (oligochaetes and leeches), a group of easily overlooked invertebrates that is of tremendous ecological, agricultural and economic importance (Charles 2019; Barrett 2016). To further investigate these rearrangements, we have undertaken a comparison of 12 clitellate genomes (including four newly sequenced species) and 11 outgroup representatives. We show that these rearrangements began at the base of the Clitellata (rather than progressing gradually throughout polychaete annelids), that the inter-chromosomal rearrangements continue in several clitellate lineages and that these events have substantially shaped the evolution of the otherwise highly conserved Hox cluster.
Matsuo, L.; Novak Vanclova, A. M. G.; Pomiankowski, A.; Lane, N.; Dacks, J. B.
Show abstract
The origin of meiotic sex was a key milestone in the evolution of the eukaryotic cell. The DNA recombinases Rad51 and DMC1 have been used previously to trace the timing and origins of the meiotic machinery, and warrant revisiting in the face of the recent increased diversity of reported asgardarchaeal taxa. Here we perform comparative genomics and phylogenetic analyses of RadA protein sequences from a broad sampling of eukaryotic and archaeal taxa. We show that even with increased and new sampling, the eukaryotic Rad51 and DMC1 proteins still resolve separately from any archaeal RadA sequences. Taking into account recent evolutionary cell biological discoveries, our data are most consistent with a scenario whereby the asgardarchaeal host cell was evolving cytoskeletal and membrane-protein machinery that was later incorporated into eukaryotic endomembrane systems following the acquisition of mitochondria and the evolution of sex.
Sharma, A.; Khushi, K.; Ravindran, F.; Kadandale, J. S.; Choudhary, B.; Srinivasan, S.; Nanjundiah, V.
Show abstract
The life cycle of the Dictyostelid amoebae is unusual in that it alternates between a free-living solitary phase and an aggregative social phase. We used six previously collected Dictyostelium giganteum strains from distinct ecological niches in the Mudumalai Nature Reserve, India. From them, we generated short Illumina reads and assembled a consensus genome, comprising nuclear and mitochondrial genomes, representative of all six strains. The nuclear assembly has an AT content of 75.76%, accounting for 38.52 Mb, and resolves into five chromosome-scale scaffolds that are consistent with the published karyotype. Its N50 is 3.01 Mb and L50 is 5. The BUSCO analysis shows 3.9% fragmentation and 92.1% completeness. Genome assembly and completeness are also validated using [~]5,500 genes from a publicly available transcriptome dataset (PRJNA48443) derived from the post-aggregation stage. 13,251 predicted proteins are encoded by the genome, including ABC transporters, polyketide synthases, Ras/Rho GTPases, and expanded families of protein kinases. Comparative analysis demonstrates extensive conservation of syntenic blocks related to the dictyostelids D. discoideum and D. firmibasis as well as lineage-specific rearrangements. About 18% of the nuclear genome is made up of repetitive DNA, mostly in the form of simple repeats. Major transposable element classes, including piggyBac-like fragments, were found by homology searches. Long poly-asparagine/glutamine tracts are less common than in D. discoideum, but low-complexity sequences are common due to strong AT-driven codon bias. Comparative proteome-level orthology analysis across Dictyostelium species and Entamoeba identified a conserved Amoebozoan core together with a substantial Dictyostelium-specific gene repertoire. Domain-level comparisons further revealed widespread conservation of intracellular signalling and cytoskeletal modules shared with animals, whereas canonical metazoan extracellular adhesion domains were absent, highlighting the deep evolutionary roots of regulatory complexity underlying aggregative multicellularity.
Koludarov, I.; Jackson, T. N.; Suranse, V.; Pozzi, A.; Sunagar, K.; Mikheyev, A. S.
Show abstract
Gene duplication is associated with the evolution of many novel biological functions at the molecular level. The dominant view, often referred to as "neofunctionalization", states that duplications precede many novel gene functions by creating functionally redundant copies which are less constrained than singletons. However, numerous alternative models have been formulated, including some in which novel functions emerge prior to duplication. Unfortunately, few studies have reconstructed the evolutionary history of a functionally diverse gene family sufficiently well to differentiate between these models. Here we examined the evolution of the g2 family of phospholipase A2 (EC 3.1.1.4) in the genomes of 93 species from all major lineages of Vertebrata. This family is evolutionarily important and has been co-opted for a diverse range of functions, including innate immunity and venom. The genomic region in which this family is located is remarkably syntenic. This allowed us to reconstruct all duplication events over hundreds of millions of years of evolutionary history using manual annotation of gene clusters, which enabled the discovery of a large number of previously un-annotated genes. Intriguingly, we found that the same ancestral gene in the phospholipase gene cluster independently acquired novel molecular functions in birds, mammals and snake, and all subsequent expansion of the cluster originates from this locus. This suggests that the locus has a deep ancestral propensity for multiplication, likely conferred by a structural arrangement of genomic material (i.e. the "genomic context" of the locus) that dates back at least the amniote MRCA. These results highlight the underlying complexity of gene family evolution, as well as the historical- and context-dependence of gene family evolution. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=137 SRC="FIGDIR/small/583344v3_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@1c3228forg.highwire.dtl.DTLVardef@12053forg.highwire.dtl.DTLVardef@11683c0org.highwire.dtl.DTLVardef@123eb3a_HPS_FORMAT_FIGEXP M_FIG C_FIG
Maddamsetti, R.
Show abstract
Bacteria, Archaea, and Eukarya all share a common set of metabolic reactions. This implies that the function and topology of central metabolism has been evolving under purifying selection over deep time. Central metabolism may similarly evolve under purifying selection during longterm evolution experiments, although it is unclear how long such experiments would have to run (decades, centuries, millennia) before signs of purifying selection on metabolism appear. I hypothesized that central and superessential metabolic enzymes would show evidence of purifying selection in the long-term evolution experiment with Escherichia coli (LTEE). I also hypothesized that enzymes that specialize on single substrates would show stronger evidence of purifying selection in the LTEE than generalist enzymes that catalyze multiple reactions. I tested these hypotheses by analyzing metagenomic time series covering 62,750 generations of the LTEE. I find mixed support for these hypotheses, because the observed patterns of purifying selection are idiosyncratic and population-specific. To explain this finding, I propose the Jenga hypothesis, named after a childrens game in which blocks are removed from a tower until it falls. The Jenga hypothesis postulates that loss-of-function mutations degrade costly, redundant, and nonessential metabolic functions. Replicate populations can therefore follow idiosyncratic trajectories of lost redundancies, despite purifying selection on overall function. I tested the Jenga hypothesis by simulating the evolution of 1,000 minimal genomes under strong purifying selection. As predicted, the minimal genomes converge to different metabolic networks. Strikingly, the core genes common to all 1,000 minimal genomes show consistent signatures of purifying selection in the LTEE. Significance StatementPurifying selection conserves organismal function over evolutionary time. However, few studies have examined the role of purifying selection during adaptation to novel environments. I tested metabolic enzymes for purifying selection in an ongoing long-term evolution experiment with Escherichia coli. While some populations show signs of purifying selection, the overall pattern is inconsistent. To explain these findings, I propose the Jenga hypothesis, in which loss-of-function mutations first degrade costly, redundant, and nonessential metabolic functions, after which purifying selection begins to dominate. I then tested several predictions of the Jenga hypothesis using computational simulations. On balance, the simulations confirm that we should find evidence of purifying selection on the metabolic pathways that sustain growth in a novel environment.
Descorps-Declere, S.; Richard, G.-F.
Show abstract
Since the formation of the first proto-eukaryotes, more than 1.5 billion years ago, eukaryotic gene repertoire as well as genome complexity has significantly increased. Among genetic elements that are responsible for this increase in genome coding capacity and plasticity are tandem repeats such as microsatellites, minisatellites and their bigger brothers, megasatellites. Although microsatellites have been thoroughly studied in many organisms for the last 20 years, little is known about the distribution and evolution of mini- and megasatellites. Here, we describe the first genome-wide analysis of megasatellites in 58 vertebrate genomes, belonging to 12 monophyletic groups. We show that two bursts of megasatellite formation occurred, one after the radiation between agnatha et gnathostomata fishes and the second one later, in therian mammals. Megasatellites are frequently encoded in genes involved in transcription regulation (zinc-finger proteins) and intracellular trafficking, but also in cell membrane metabolism, reminiscent of what was observed in fungi genomes. The presence of many introns within young megasatellites suggests a model in which an exon-intron DNA segment is first duplicated and amplified before the accumulation of mutations in intronic parts partially erase the tandem repeat in such a way that it becomes detectable only in exonic regions. In addition, evidence for the genetic transfer of megasatellites between unrelated genes suggests that megasatellite formation and evolution is a very dynamic and still ongoing process in vertebrate genomes.
Nishida, A.; Ochman, H.
Show abstract
Bacterial strains evolve in response to the gut environment of their hosts, with genomic changes that influence their interactions with hosts as well as with other members of the gut community. Great apes in captivity have acquired strains of Bacteroides xylanisolvens, which are common within gut microbiome of humans but not typically found other apes, thereby enabling characterization of strain evolution following colonization. Here, we isolate, sequence and reconstruct the history of gene gain and loss events in numerous captive-ape-associated strains since their divergence from their closest human-associated strains. We show that multiple captive-ape-associated B. xylanisolvens lineages have independently acquired gene complexes that encode functions related to host mucin metabolism. Our results support the finding of high genome fluidity in Bacteroides, in that several strains, in moving from humans to captive apes, have rapidly gained large genomic regions that augment metabolic properties not previously present in their relatives. Significance statementChronicling the changes that occur in bacterial genomes after a host-switch event is normally difficult due to age of most bacteria-host associations, which renders uncertainties about the bacterial ancestor (and ancestral genome) prior to colonization of the new host. However, the gut microbiomes of great apes in captivity contain bacterial strains that are unique to humans, allowing fine-scale assessment and reconstruction of the genomic changes that follow colonization. By sequencing and comparing closely related strains of Bacteroides that are restricted both to human and to captive great apes, we found that multiple bacterial lineages convergently acquired sets of genes involved in the metabolism of dietary polysaccharides. These results show that over relatively short timescales, the incorporation of strains into microbiomes involves large-scale genomic events that correspond to characteristics of the new host environment.
Chang, E. S.; Connelly, M. T.; Travert, M.; Barreira, S. N.; Rivera, A. M.; Katzer, A. M.; Yu, R.; Cartwright, P.; Baxevanis, A. D.
Show abstract
Cnidarians are important models for the studying the evolution of animal development, regeneration, cell type differentiation, and allorecognition. The marine hydrozoan Podocoryna americana is related to the well-established model species Hydractinia symbiolongicarpus. Although both species possess a sessile polyp stage, P. americana differs in that it also has a free-swimming medusa (jellyfish) stage in its life cycle. We used a combination of PacBio CLR long-read and Illumina Hi-C short-read genome sequencing to produce a chromosome-level genome assembly for P. americana. The final assembly is 327 Mbp in total length with 17 chromosome-scale scaffolds representing 98% of the assembly. Comprehensive functional annotation with BRAKER3 generated a total of 19,085 predicted protein-coding genes in this assembly, covering 91.2% of the metazoan BUCSO gene set. Comparison of the P. americana genome to other chromosome-level cnidarian genome assemblies revealed a high degree of macrosynteny conservation, and ortholog identification and gene family evolution analysis identified 522 expanded and 1,026 contracted gene families in P. americana. This high-quality, chromosome-level genome assembly of P. americana will be an invaluable resource for researchers studying the evolution of development, regeneration, and allorecognition in cnidarians and other metazoans.
Fierst, J. L.; Milwood, J. D.; Adams, P. E.; Sutton, J. M.; Pienaar, J.
Show abstract
Genome size has been measurable since the 1940s but we still do not understand the basis of genome size variation. Caenorhabditis nematodes show strong conservation of chromosome number but vary in genome size between closely related species. Androdioecy, where populations are composed of males and self-fertile hermaphrodites, has evolved from outcrossing, female-male dioecy, three times in this group. Androdioecious genomes are 10-30% smaller than dioecious species but large phylogenetic distances and rapid protein evolution have made it difficult to pinpoint the basis of these changes. Here, we analyze the genome sequences of Caenorhabditis and and test three hypotheses explaining genome evolution: 1) genomes evolve through deletions and genome shrinkage in androdioecious species; 2) genome size is determined by transposable element (TE) expansion and DNA loss through large deletions (the accordion model); and 3) TE dynamics differ in androdioecious and dioecious species. We find no evidence for these hypotheses in Caenorhabditis. Across both short and long evolutionary distances Caenorhabditis genomes evolve through small structural variant (SV) mutations including frequent duplications and insertions, predominantly in genic regions. Caenorhabditis have rapid rates of gene family expansion and contraction and we identify 71 protein families with significant, parallel decreases across self-fertile Caenorhabditis. These include genes involved in the sensory system, regulatory proteins and membrane-associated immune responses, reflecting the shifting selection pressures that result from self-fertility. Our results suggest that the rules governing genome evolution differ between organisms based on ecology, life style and reproductive system.
Farrell, A. A.; Nesbo, C. L.; Zhaxybayeva, O.
Show abstract
The placement of a non-hyperthermophilic order Mesoaciditogales at the base of Thermotogota tree challenges the prevailing hypothesis that the last common ancestor of Thermotogota was a hyperthermophile. Yet, given the long branch leading to the only two Mesoaciditogales described to-date, the phylogenetic position of the order may be due to the long branch attraction artifact. By testing various models and applying data recoding in phylogenetic reconstructions, we observed that Mesoaciditogales basal placement is strongly supported by the conserved marker genes assumed to be vertically inherited. However, based on the taxonomic content of 1,181 gene families and a phylogenetic analysis of 721 gene family trees, we also found that a substantial number of Mesoaciditogales genes are more closely related to species from the order Petrotogales. These genes contribute to coenzyme transport and metabolism, fatty acid biosynthesis, genes known to respond to heat and cold stressors, and include many genes of unknown functions. The Petrotogales comprise moderately thermophilic and mesophilic species with similar temperature tolerances to that of Mesoaciditogales. Our findings hint at extensive horizontal gene transfer between, or parallel independent gene gains by, the two ecologically similar lineages, and suggest that the exchanged genes may be important for adaptation to comparable temperature niches. SignificanceThe high-temperature phenotype is often referenced when conjecturing about characteristics of the last common ancestor of all present-day organisms. Such inferences rely on accuracy of phylogenetic trees, especially with respect to lineages that branch closest to the last common ancestor. Here, we examined evolutionary history of Mesoaciditogales, an early-branching lineage within Thermotogota phylum, which is one of the early-diverging groups of bacteria. Thermotogota is composed of thermophiles, hyperthermophiles and mesophiles, who collectively can grow between 20 to 90 degrees Celsius, making it challenging to infer the growth temperature of their common ancestor. Our analysis revealed a complex evolutionary history of Mesoaciditogales genome content impacted by horizontal gene transfer, highlighting the challenges of ancestral phenotype inferences using present-day genomes.
Wang, J.; Zhang, G.; Sun, C.; Chang, L.; Wang, Y.; Yang, X.; Chen, G.; Itgen, M. W.; Haley, A.; Tang, J.; Mueller, R. L.
Show abstract
Size evolution among gigantic genomes involves gain and loss of many gigabases of transposable elements (TEs), sequences that parasitize host genomes. Animals suppress TEs using piRNA and KRAB-ZFP pathways. TEs and hosts coevolve in an arms race, where suppression strength reflects TE fitness costs. In enormous genomes, additional TE costs become miniscule. How, then, do TEs and host suppression invoke further addition of massive DNA amounts? We analyzed TE proliferation histories, deletion rates, and community diversities in six salamander genomes (21.3 - 49.9 Gb), alongside gonadal expression of TEs and suppression pathways. TE activity is higher in testes than ovaries, attributable to lower KRAB-ZFP suppression. Unexpectedly, genome size/expansion is uncorrelated with TE deletion rate, proliferation history, expression, and host suppression. Also, TE community diversity increases with genome size, contrasting theoretical predictions. TE/host antagonism in gigantic genomes likely produces stochastic TE accumulation, determined by noisy intermolecular interactions in huge genomes/cells.
Reinar, W. B.; Torresen, O. K.; Nederbragt, A. J.; Matschiner, M.; Jentoft, S.; Jakobsen, K. S.
Show abstract
Repetitive DNA make up a considerable fraction of most eukaryotic genomes. In fish, transposable element (TE) activity have coincided with rapid species diversification. Here, we annotated the repetitive content in 100 genome assemblies, covering the major branches of the diverse lineage of teleost fish. We investigated if TE content correlates with family level net diversification rates and found support for a weak negative correlation. Further, we found that TE content, the degree of parental care and short tandem repeat (STR) content contributed to genome size variability. In contrast to TEs, STR content showed a negative relationship with genome size. STR content did not correlate with TE content, which implies independent evolutionary paths. Last, marine and freshwater fish have large differences in STR content. The most extreme propagation was found in the genomes of codfish species and Atlantic herring. Such a high density of STRs is likely to increase the mutational load, which we propose could be counterbalanced by high fecundity as seen in codfishes and herring.
Rivera-Rivera, C. J.; Feuda, R.; Pisani, D.
Show abstract
The evolution of the nervous system has been shrouded in controversy since the onset of the genomics era. A large part of this controversy stems from the lack of phylogenetic consensus for the main branches of the animal tree, where often animals with nervous systems do not form a monophyletic group. However, this question can be informed from other non-phylogenetic perspectives, such as comparative genomics. Here we ask how similar are genes differentially expressed in neurons across a representative set of eight animals with a nervous system and tally the presence or absence of their homologs in 10 other animals and two choanoflagellates. We show that proteins from 39 families are differentially expressed in all neurons, regardless of the phylogenetic placement of their lineage, and that the majority of these gene families are present in the (unicellular) closest relatives of animals, choanoflagellates. We found that the members of these 39 gene families are enriched in domains for ion transport and juxtacrine signalling, and that there is one gene family of zinc-dependent extracellular matrix-remodelling proteins which is only found in animals bearing a nervous system. Our results show that common genetic toolkits are in place for the function of nervous systems. We identify a large number of potential new genomic markers linked to the nervous system and hope they can complement ongoing research efforts to better understand this quintessentially animal system.
McGowan, J.; Richards, T. A.; Hall, N.; Swarbreck, D.
Show abstract
The translation of nucleotide sequences into amino acid sequences, governed by the genetic code, is one of the most conserved features of molecular biology. The standard genetic code, which uses 61 sense codons to encode one of the 20 standard amino acids and 3 stop codons (UAA, UAG, and UGA) to terminate translation, is used by most extant organisms. The protistan phyla Ciliophora (the ciliates) are an unusual exception to this norm, exhibiting the greatest diversity of non-canonical nuclear genetic code variants and evidence of repeated changes in code. In this study, we report the discovery of multiple independent genetic code changes within the Phyllopharyngea class of ciliates. By mining publicly available ciliate genome datasets, we discovered that three ciliate species from the TARA Oceans eukaryotic metagenome dataset use the UAG codon to putatively encode leucine. We identified novel suppressor tRNA genes in two of these genomes. Phylogenomics analysis revealed that these three uncultivated taxa form a monophyletic lineage within the Phyllopharyngea class. Expanding our analysis by reassembling published phyllopharyngean genome datasets led to the discovery that the UAG codon had also been reassigned to putatively code for glutamine in Hartmannula sinica and Trochilia petrani. Phylogenomics analysis suggests that this occurred via two independent genetic code change events. These data demonstrate that the reassigned UAG codons have widespread usage as sense codons within the phyllopharyngean ciliates. Furthermore, we show that the function of UAA is firmly fixed as the preferred stop codon. These findings shed light on the evolvability of the genetic code in understudied microbial eukaryotes.
Naser-Khdour, S.; Scheuber, F.; Fields, P. D.; Ebert, D.
Show abstract
Genomic regions that play a role in parasite defense are often found to be highly variable, with the MHC serving as an iconic example. Single nucleotide polymorphisms may represent only a small portion of this variability, with Indel polymorphisms and copy number variation further contributing. In extreme cases, haplotypes may no longer be recognized as homologs. Understanding the evolution of such highly divergent regions is challenging because the most extreme variation is not visible using reference-assisted genomic approaches. Here we analyze the case of the Pasteuria Resistance Complex (PRC) in the crustacean Daphnia magna, a defense complex in the host against the common and virulent bacterium Pasteuria ramosa. Two haplotypes of this region have been previously described, with parts of it being non-homologous, and the region has been shown to be under balancing selection. Using pan-genome analysis and tree reconciliation methods to explore the evolution of the PRC and its characteristics within and between species of Daphnia and other Cladoceran species, our analysis revealed a remarkable diversity in this region even among host species, with many non-homologous hyper-divergent-haplotypes. The PRC is characterized by extensive duplication and losses of Fucosyltransferase (FuT) and Galactosyltransferase (GalT) genes that are believed to play a role in parasite defense. The PRC region can be traced back to common ancestors over 250 million years. The unique combination of an ancient resistance complex and a dynamic, hyper-divergent genomic environment presents a fascinating opportunity to investigate the role of such regions in the evolution and long-term maintenance of resistance polymorphisms. Our findings offer valuable insights into the evolutionary forces shaping disease resistance and adaptation, not only in the genus Daphnia, but potentially across the entire Cladocera class. SignificanceUnderstanding how organisms adapt to their environment requires insights into the evolution of genetic defenses against their parasites. While the Major Histocompatibility Complex (MHC) is a well-known example of a highly variable immune-related gene region, much remains unknown about the evolution of other such regions. Our study investigates the Pasteuria Resistance Complex (PRC) in water fleas, a genomic region crucial for defense against a parasitic bacterium. We discovered that the PRC is exceptionally diverse, with a history spanning hundreds of millions of years. This research provides new insights into the mechanisms underlying the maintenance of genetic diversity in the face of persistent parasite pressure. Our findings contribute to a broader understanding of how organisms evolve robust defenses against infectious diseases.
Tico, M.; Mariotti, M.
Show abstract
Selenoproteins incorporate the rare selenium-containing amino acid selenocysteine (Sec) and play crucial roles for redox homeostasis, stress response, and hormone regulation. Sec is inserted by co-translational recoding of the UGA codon, normally a stop. As a consequence, selenoproteins are often misannotated in public databases and require specialized bioinformatic methods and resources. Here, we present a refined characterization of the composition and evolution of the vertebrate selenoproteome. Based on analyses of 19 gene families across hundreds of genomes, we show that extant selenoproteomes were shaped by extensive gene duplications (56 selenoproteins), losses (50), and Sec-to-cysteine (Cys) conversions (21). Tetrapods including mammals encode 24-25 selenoproteins, with variations in 6 families. Notably, the same genes underwent convergent evolutionary events in multiple tetrapods, namely Sec-to-Cys substitutions (SELENOU1, GPX6) and gene losses (SELENOV). In contrast, ray-finned fish exhibit larger and more dynamic selenoproteomes, reinforcing the hypothesis that the selective advantage of Sec is stronger in aquatic environments. We detected selenoprotein duplications spread across the actinopterygian clade involving 13 families, mainly involved in antioxidant defense. The richest selenoproteomes were found in Salmonidae and Cyprinoidei fish with 56 and 44 selenoproteins, respectively, owing to whole genome duplications. Among our findings, the SELENOP family stands out in lampreys, carrying up to an unprecedented 162 UGAs putatively recoded to Sec. Our study presents the most comprehensive evolutionary map of vertebrate selenoproteins to date and delineates the specific selenoproteome of each lineage, establishing a foundational framework for selenium biology research in the era of biodiversity genomics.
Sahm, A.; Cherkasov, A.; Liu, H.; Voronov, D.; Siniuk, K.; Schwarz, R.; Ohlenschlaeger, O.; Foerste, S.; Bens, M.; Groth, M.; Goerlich, I.; Paturej, S.; Klages, S.; Braendl, B.; Olsen, J.; Bushnell, P.; Poulsen, A. B.; Ferrando, S.; Garibaldi, F.; Drago, D. L.; Tozzini, E. T.; Mueller, F.-J.; Fischer, M.; Kretzmer, H.; Domenici, P.; Steffensen, J. F.; Cellerino, A.; Hoffmann, S.
Show abstract
The Greenland shark (Somniosus microcephalus) is the longest-lived vertebrate known, with an estimated lifespan of [~] 400 years. Here, we present a chromosome-level assembly of the 6.45 Gb Greenland shark, rendering it one of the largest non-tetrapod genomes sequenced so far. Expansion of the genome is mostly accounted for by a substantial expansion of transposable elements. Using public shark genomes as a comparison, we found that genes specifically duplicated in the Greenland shark form a functionally connected network enriched for DNA repair function. Furthermore, we identified a unique insertion in the conserved C-terminal region of the key tumor suppressor p53. We also provide a public browser to explore its genome.