Back

Genomics

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Genomics's content profile, based on 64 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.

1
Functional Characterization of Transcriptome-Wide Isoform Switching in Hürthle Cell Carcinoma (HCC)

Butt, R. S.; Amir, A.; Paracha, R. Z.

2026-07-27 bioinformatics 10.64898/2026.07.23.740299 medRxiv
Top 0.1%
13.1%
Show abstract

Hurthle cell carcinoma (HCC) is an aggressive form of thyroid cancer. While mitochondrial DNA mutations and chromosomal losses have been identified in HCC, isoform switching, and its functional consequences remain uncharacterized. This study reanalyzed NCBI GEO dataset GSE228870 (n = 32), using Salmon and IsoformSwitchAnalyzeR() to identify isoform switching. The analysis resulted in 371 switches across 335 genes showing functional consequences including loss of protein domains, shorter open reading frames (ORFs), loss of signal peptides and novel sub-cellular localizations. Most significant isoform switches (q-value < 0.05, |dIF| > 0.1) were observed in LAMA2, LSP1, MAD2L2, FBLN2 and CXCL12, implicating extracellular matrix dysregulation, DNA damage response, immune signaling and cytoskeleton regulation. These genes are expressed in normal thyroid (median TPM 20.69, 11.66, 14.79, 134.1 & 80.76). However, specific isoforms of LAMA2 and MAD2L2 are not expressed in normal thyroid, explaining tumor-specific expression in HCC. Alternative transcription termination site (ATTS) gain was significant, suggesting altered 3 end in HCC transcripts. TCGA SpliceSeq showed LSP1, FBLN2 and CXCL12 undergo alternative promoter (LSP1 exon1 PSI=94.5%, FBLN2 exon2 PSI=99.0%) and alternative termination (CXCL12 exon3.3 PSI=53.9%) in thyroid cancer, suggesting ATTS and alternative transcription start site (ATSS) as shared splicing dysregulation mechanisms. This is the first systematic characterization of isoform-level dysregulation in HCC.

2
Evolutionary Stratification of Codon Usage Bias In Plants Arises from GC3 Composition and Translational Optimization

Mohanta, T. K.

2026-07-01 genomics 10.64898/2026.06.26.734692 medRxiv
Top 0.1%
6.7%
Show abstract

Codon usage bias is a fundamental genomic characteristic that prefers non-random preferential use of synonymous codons. It is a major determinant of translational efficiency, gene regulation, and molecular evolution. However, the evolutionary bias and functional relevance of codon usage bias across the plant lineage is poorly defined and yet to understand what are the major factors responsible for relative synonymous codon usage (RSCU) in genomes and how codon usage bias influences the gene regulation, molecular evolution genomes. A genome-wide codon usage bias study of coding DNA sequences of 262 plant genome was conducted. It encompassed more than 4.6 billion codons from > 11 million coding sequences. Relative synonymous codon usage, codon adaptation index, codon-anticodon mapping, effective number of codon (ENC)-GC3, GC1,2-GC3, parity rule 2 (PR2-bias), molecular economy, and machine learning approaches were used for the study. It was found that codon usage bias was strongly non-random and exhibited a clear phylogenetic structuring. The higher plants favoured A/T-ending, whereas early-diverging lineages were enriched in G/C-ending codons. Analysis of RSCU, codon adaptation index, and codon-anticodon pairing indicated that translational selection is mediated by tRNA availability, contributing sustainability to these molecular patterns. Machine-learning approaches identified a small subset of codons having outsized influence on genome-wide codon usage landscapes. Further studies revealed the presence of robust inverse relationships between the effective number of codons and GC content at synonymous third positions. Neutrality analysis revealed approximately 61% of variation was driven by mutational pressure, tempered by selective constraints. Phylogenetic reconstruction showed a progressive relaxation of codon bias from algae to angiosperms while maintaining a conserved molecular economy cost of ~ 30 ATP per codon across the lineages. The study revealed codon usage bias is lineage-specific evolutionary conserved trait governed by mutation, selection, and translational optimization.

3
A single-nucleus atlas of the adult laying hen liver reveals metabolic specialization and improved cellular resolution through enhanced genome annotation

Lagoutte, L.; Allain, C.; Lebez, B.; Cossard, G.; Lecerf, F.; Blum, Y.; lagarrigue, S.; Degalez, F.

2026-07-29 genomics 10.64898/2026.07.27.740890 medRxiv
Top 0.1%
5.4%
Show abstract

The liver of laying hens plays a central role in metabolism and reproduction, supporting the synthesis of egg yolk precursors under strong hormonal regulation. Despite its physiological importance, a high-resolution cellular reference of the adult chicken liver is still lacking. Here, we generated a single-nucleus RNA sequencing atlas of the adult laying hen liver from eight individuals, providing a comprehensive view of its cellular composition and transcriptional landscape. Using this framework, we identified major hepatic cell populations, including hepatocytes, endothelial cells, cholangiocytes, hepatic stellate cells, and diverse immune cell types, revealing a broadly conserved vertebrate liver architecture. However, hepatocyte zonation, a key feature of mammalian liver organization, was not observed, consistent with the absence of hepatocyte zonation reported in birds. Importantly, we demonstrate that the use of an enriched genome annotation, incorporating additional protein-coding and long non-coding RNA models, substantially improves transcript detection and enhances cell-type resolution in single-nucleus datasets. This improved resolution allows more accurate marker-based assignment of hepatocyte subpopulations and refines the interpretation of hepatic cellular heterogeneity. Within hepatocytes, we uncovered transcriptionally distinct subpopulations associated with lipid metabolism and reproductive function, including estrogen-responsive programs involving cytochrome P450 genes such as CYP2C23A and CYP2C23B. In parallel, we characterized a complex immune compartment composed of resident macrophages and adaptive immune cells, highlighting the dual metabolic and immunological roles of the avian liver. Overall, this atlas provides a high-resolution reference for avian liver biology and demonstrates that improved genome annotation enhances the resolution and interpretation of cellular heterogeneity in single-cell transcriptomic studies.

4
Germline genomic and methylomic dynamics following three generations of early-life metabolic challenges

de Anca Prado, V.; Pertille, F.; Andersson, D.; Mourin-Fernandez, M.; Godia, M.; Jimenez-Chillaron, J. C.; Ruegg, J.; Guerrero-Bosagna, C.

2026-07-24 genomics 10.64898/2026.07.21.739755 medRxiv
Top 0.1%
5.4%
Show abstract

Environmental and dietary factors can exert multigenerational effects on health and development. In this study, we investigated whether early-life metabolic challenge affects the germline genome and epigenome across three generations. Using a murine model of early life obesity via litter size reduction (overnutrition group, ON) and a control group (CT), we followed the paternal lineage focusing on germline genomic and methylation changes employing Genotyping-by-Sequencing (GBS) coupled with methyl-immunoprecipitation (GBS-MeDIP). We found that unrelated ON families clustered together based on identified Single-Nucleotide Polymorphism (SNP), suggesting that the treatment may have genomic impact. Copy number variations (CNVs) events were identified in ON individuals, being enriched in Long Interspersed Nuclear Elements (LINEs) and Long Terminal Repeats (LTRs). While Principal Component Analysis (PCA) of the methylome showed no clear treatment effect, pathway enrichment and regional analyses revealed methylation changes associated with transposable elements and developmental genes. Notably, the ON group exhibited a disruption in the methylation of Repetitive Elements (RE), which was significant in the same type of RE that were also enriched in the observed CNVs. The ON also showed reduced emergence of novel SNPs in offspring compared to the CT group. These findings suggest that multigenerational metabolic challenge can constrain genetic variability and induce genome instability, potentially mediated by transposable element activity rather than widespread changes in DNA methylation. This work highlights the importance of studying both genome and epigenome dynamics under realistic, multigenerational exposure scenarios and suggests that early metabolic challenges can have long-lasting impacts on genomic architecture and evolutionary potential.

5
Title: A de novo transcriptomic atlas of early embryo development in the Arabian killifish

Akinmusola, R. Y.; Minhas, R.; O'Neill, P.; Kon-Nanjo, K.; Kon, T.; Shimada, Y.; Ramsdale, M.; Kudoh, T.

2026-07-21 genomics 10.64898/2026.07.17.738637 medRxiv
Top 0.1%
5.0%
Show abstract

The Arabian killifish (Aphaniops dispar) is new tractable vertebrate model system for developmental, ecological and biomedical research, including drug screening, pharmacological and infection biology studies. It is a relatively small euryhaline teleost with broad thermal tolerance and adaptability across a wide range of salinities from freshwater to hypersaline habitats. The embryos and early larvae are tolerant to environmental stressors and exhibit a delayed period of nutritional independence before hatching. This advantage offers an extended window for experimenting on the early developmental processes. Here, we describe time-course gene expression profiling of Arabian killifish embryos across nine developmental time points, from the 1-cell stage to the larval pre-hatching stage. Clustering of dynamic expression profiles for 27,564 Trinity genes revealed coordinated transcriptional modules corresponding to the maternal, blastula, maternal-to-zygotic transition (MZT)-related, gastrulation, organogenesis and larval maturation stages. The maternal stage displayed a highly distinct expression profile, dominated by maternal-specific transcripts that are rapidly degraded during the MZT. The later stages, from 48 hpf onward, revealed a shift from early regulatory mechanisms to the expression of organogenesis-related genes. The ZGA stage showed the conserved up-regulation of many zinc finger-associated genes, consistent with zebrafish and other teleost genomes. Overall, embryo development in A. dispar is slower than in zebrafish, with equivalent stages occurring several hours later. We propose a delayed onset of zygotic genome activation (ZGA) in the blastula stage, corresponding to 6 hpf in the Arabian killifish. Taken together, this study provides a transcriptomic resource for mining embryo development-related genes in the Arabian killifish.

6
Systematic benchmarking of low-input whole exome sequencing workflows for longitudinal ctDNA profiling in pancreatic ductal adenocarcinoma

James, L. G.; Thorn, G. J.; Morel, C.; PCRFTB, ; Kocher, H. M.; Ross-Adams, H. E.; Chelala, C.

2026-07-03 genomics 10.64898/2026.06.29.734743 medRxiv
Top 0.1%
4.9%
Show abstract

Whole exome sequencing (WES) of circulating tumour DNA (ctDNA) enables longitudinal monitoring of tumour dynamics, evolution and treatment response but remains technically challenging in low-input, low-shedding settings such as pancreatic ductal adenocarcinoma (PDAC). Here, we systematically compared three commercially available low-input WES workflows incorporating Agilent (V6, V8) and Qiagen exome capture designs using ultra-low input cfDNAs extracted from multiple matched longitudinal plasma samples from PDAC patients. Using predefined performance metrics including coverage, duplication rate and variant detection and additional metrics relevant for clinical genomic profiling in patient care, we show that all three workflows produced high-quality sequencing data, even from very low input cfDNA. Within the conditions tested here, the Agilent V8 workflow provided the most favourable balance of coverage uniformity, sequencing efficiency and hotspot coverage for low input, low tumour fraction cfDNA WES. These findings demonstrate that workflow design, including capture footprint, substantially influences ctDNA WES performance in low-input clinical contexts. These findings are particularly relevant in early stage and/or minimal residual disease settings, where tumour fractions are low and recovery of genomic information from limited-input samples is critical.

7
A gapless Landrace pig genome resolves centromeres and telomeres and highlights telomere repeat structures in different pig breeds

Grove, H.; Stenlokk, K. S. R.; Lien, S.; Gjuvsland, A. B.; Arnyasi, M.; van Son, M.; Kent, M.

2026-06-30 genomics 10.64898/2026.06.25.734473 medRxiv
Top 0.2%
4.2%
Show abstract

Abstract The Duroc-derived reference genome Sscrofa11.1 has provided a critical foundation for pig genomics, providing a high-quality reference genome for accurate variant detection and comparative genomics but does not capture breed-specific variation. Here, we present a near-complete, gap-free genome assembly for the Landrace pig (Landrace_v1, GCA_963921485.1), spanning all 20 chromosomes and totaling 2.6 Gb, including 176 Mb of sequence absent from Sscrofa11.1. Comparative analyses with recently published high-quality pig genomes reveal a conserved centromere organization across breeds, accompanied by substantial variation in repeat composition and length, and identify a pig specific pattern of telomere variant repeats across eight pig breeds. The improved resolution of repetitive regions in Landrace_v1 enables more complete reconstruction of complex gene families, including olfactory receptors, and uncovers structural variation at the KIT proto-oncogene receptor tyrosine kinase locus not represented in the Duroc reference. Together, these findings highlight the limitations of single-reference genomes and demonstrate the value of breed-specific assemblies for capturing genomic diversity and improving downstream analyses.

8
A gapless telomere-to-telomere reference genome of Ostreococcus tauri RCC4221 with expanded annotation of medium-sized ncRNAs

Liu, G.; Bousquet, L.; Mayeur, H.; Manirakiza, E.; Daric, V.; Klopp, C.; Noirot, C.; Lopez-Escardo, D.; Grimsley, N. H.; Yau, S.; Krasovec, M.; Echeverria, M.; PIGANEAU, G.

2026-07-14 genomics 10.64898/2026.07.10.737489 medRxiv
Top 0.2%
4.0%
Show abstract

Marine photosynthetic microbes contribute substantially to global primary production, yet many algal lineages still lack reference genomes with the continuity and annotation quality required for fine-scale structural, regulatory and comparative analyses. Ostreococcus tauri, one of the smallest known free-living photosynthetic eukaryotes, has been a model marine picoeukaryote for over two decades. Despite successive improvements to its historical reference genome, previous assemblies retained hundreds of gaps and incomplete genes, hampering high-resolution genomic analyses. Here, we present O. tauri RCC4221 genome version 2026, a telomere-to-telomere assembly of all 20 chromosomes spanning 13.34 Mb with no gaps. This assembly combines PacBio long-read sequencing, Illumina short-read polishing, correction of unresolved regions guided by independent Nanopore-based assemblies. The updated reference supports a curated annotation comprising 7,683 protein-coding genes, 48 tRNA genes, 3 rRNA operons, 116 medium-sized noncoding RNAs, one signal recognition particle RNA and 138 small nucleolar RNAs. It also improves gene-model integrity and recovers candidate coding loci absent from the 2014 reference. Structural analyses resolved the organization of the two atypical low-GC chromosome 2 and 19 that contain duplicated regions that were collapsed or misrepresented in previous assemblies. Finally, bisulfite sequencing and PacBio SMRT sequencing revealed a dual DNA methylation landscape, with CG-context cytosine methylation concentrated in gene bodies and N6-methyladenosine (m6A) enriched at the start codon. The updated O. tauri 2026 assembly provides a complete and curated reference resource for chromosome biology, comparative genomics, epigenomics and RNA biology in a model marine picoeukaryote.

9
Nitrogen use efficiency in pigs is associated with transcriptomic signatures related to amino acid metabolism, immune activity, and nutrient partitioning

Monney, B.; Ewaoluwagbemiga, E. O.; Kasper, C.

2026-07-01 genomics 10.64898/2026.06.26.733976 medRxiv
Top 0.2%
4.0%
Show abstract

Dietary protein restriction challenges the allocation of amino acids to growth and other physiological functions and therefore requires coordinated metabolic adaptation. Domestic pigs provide an informative system in which to study such responses, because nitrogen retention directly affects lean growth and can be quantified accurately under controlled feeding and housing conditions. Under reduced-protein diets, pigs differ in how effectively they retain nitrogen, and this variation has a genetic basis, making them well suited to investigate the molecular regulation of nitrogen use efficiency (NUE). Here, we characterise differential gene expression and enriched pathways in liver and skeletal muscle of more than 80 pigs with two divergent NUE phenotypes (high and low) maintained under the same protein-reduced, ad libitum dietary conditions. The two NUE phenotypes were clearly distinct at the transcriptomic level, with 177 differentially expressed genes in the liver and 133 in the muscle. In the liver, differential expression and enrichment analyses indicate reduced amino acid catabolism, lower inflammatory and detoxification activity, and a metabolic state that favours lipid processing and insulin-related regulation over the use of amino acids as energy sources. In skeletal muscle, they point to reduced lipid uptake, lower reliance on amino acid oxidation, and a greater emphasis on protein synthesis, translational regulation, mitochondrial energy metabolism, and growth-related processes. These gene-level patterns were supported and extended by pathway and gene-set enrichment analyses. Together, the results suggest that high and low-NUE pigs differ through coordinated, tissue-specific molecular adaptations. Overall, variation in NUE appears to reflect coordinated, tissue-specific differences in how nutrients are allocated between energy use, storage, and lean tissue growth.

10
H3K4me3 exhibits length-dependent deposition patterns at transcription initiation regions in Trypanosoma cruzi and correlates with transcriptional activity

Lopez, M. d. R.; Gitman, I. F. B.; Prego, A. F.; Lavignolle-Heguy, R.; Zambrano-Siri, R. T.; Carena, S.; Arguello, R. J.; Vilchez-Larrea, S. C.; Alonso, G. D.; Ocampo, J.

2026-06-29 genomics 10.64898/2026.06.26.734760 medRxiv
Top 0.2%
4.0%
Show abstract

In trypanosmatids genes, transcribed by RNA polymerase II do not have canonical promoters and are organized into directional gene clusters that mature into monocistronic transcripts by a co-transcriptional process known as trans-splicing. Even though gene expression is regulated mainly post-transcriptionally, it is currently understood that chromatin and epigenetics are also involved in this regulation. In eukaryotes, specific signals are normally required for the occurrence of an appropriate transcription initiation. Among them, trimethylation of histone H3 in lysine 4 is the most conserved signal normally detected at transcription start sites of actively transcribed genes. Unlike many model organisms, trypanosomes do not have defined promoters. Instead, transcription initiates in a bidirectional manner from dispersed regions coincident with divergent strand switch regions located between directional gene clusters (DGCs). In T. cruzi, H3K4me3 was observed at the origins of transcription coincident with divergent strand switch regions (dSSRs) in epimastigotes, but it has not been mapped throughout the whole genome at base-pair resolution or in other life stages so far. Here, we set up the CUT&RUN technique for T. cruzi epimastigotes and trypomastigotes. Consistent with a predominant post-transcriptional regulation along the life cycle, we did not find significant differences between life stages. We corroborated that H3K4me3 is enriched at dSSR adjacent to actively expressed DGCs. Moreover, we noticed that this histone mark exhibits different patterns that correlate with the genomic span of the transcription initiation regions and with transcriptional activity. Furthermore, we unveiled that the most actively transcribed DGCs are associated with shorter dSSRs and are located within the core compartment of the genome displaying a more accessible chromatin.

11
Genetic introgression and transcriptomic plasticity are associated with enhanced Leishmania infantum pathogenicity causing human cutaneous leishmaniasis in Tunisia

SANTI, A. M. M.; LI, B.; PIEL, L.; PIPOLI DA FONSECA, J.; BACQ-DAIAN, D.; OLASO, R.; DELEUZE, J. F.; COKELAER, T.; AOUN, K.; BOURATBINE, A.; SPÄTH, G. F.

2026-06-22 genomics 10.64898/2026.06.20.733521 medRxiv
Top 0.3%
3.3%
Show abstract

The protozoan parasite Leishmania infantum exhibits significant genetic variability among isolates, influencing disease manifestation and treatment response. Although L. infantum is classically described as the causative agent of Visceral Leishmaniasis (VL) - often associated with immune deficiency, cases of Cutaneous Leishmaniasis (CL) caused by this species in immunocompetent individuals have been reported in different countries. To investigate the molecular basis of this unusual shift in tissue tropism and pathogenicity, we applied comparative genomic and transcriptomic approaches on two canine isolates (CanL) and two human isolates associated with Cutaneous Leishmaniasis (CL) in Tunisia. While the CanL isolates showed close genetic similarity to the L. infantum reference strain (JPCM5), the CL isolates formed a separate, highly divergent cluster based on SNP localization and frequency, differing not only from JPCM5 but also from each other. Utilizing the metagenomics sequence classification tool Kraken, we revealed a complex hybrid nature of the CL isolates, showing introgression from L. donovani and L. tropica, suggesting that hybridization has played a key role in generating novel phenotypic traits. Integration of RNA-seq and DNA-seq data demonstrated that only a minority of gene expression variation within and in-between the CanL or CL groups reflected gene dosage effects due to copy number variation, while the majority of expression differences were independent of gene dosage, implying post-transcriptional regulatory mechanisms contributing to parasite adaptation. In conclusion, our study identifies hybridization, genome instability, and transcriptomic adaptation as interconnected drivers of the L. infantum evolutionary potential. These mechanisms can collectively enhance parasite fitness gain, potentially explaining the emergence of cutaneous disease forms in a species traditionally linked to visceral infection. Author SummaryThis research reveals that hybridization between distinct Leishmania parasite species could be a key mechanism driving the evolution of new disease forms. By demonstrating that cutaneous leishmaniasis (CL) cases are caused by hybrid L. infantum parasites whose genomes show introgression with DNA from L. donovani and L. tropica, this study reveals a molecular mechanism potentially linked to the emergence of tegumentary disease from a species traditionally known to cause visceral infection. These findings contribute to our understanding of Leishmania evolution, the emergence of atypical forms of leishmaniasis linked to hybridization, and the impact of genome instability and transcriptomic adaptation as potent forces for generating phenotypic diversity and enhancing parasite fitness.

12
A cis-regulatory variant in ASIP causes gray coat color in the donkey

li, y.; Liu, Y.; wu, j.; liu, s.; lin, x.; guo, k.; yang, t.; feng, m.; zhang, h.; wang, x.; xing, w.; qian, s.; yang, r.; zhao, c.

2026-06-28 genomics 10.64898/2026.06.23.733953 medRxiv
Top 0.3%
3.2%
Show abstract

BackgroundGray is one of the relatively rare coat colors in donkeys. The Hetian Gray donkey is a distinctive indigenous breed from the Xinjiang Uygur Autonomous Region of Northwestern China, characterized by progressive hair depigmentation with aging while retaining dark skin pigmentation. However, the genetic basis underlying this unique gray coat color phenotype remains unclear. ResultsTo elucidate the genetic basis, we conducted whole-genome resequencing of Gray and non-Gray donkeys. Genome-wide selection signature analyses identified a candidate region on chromosome 15. Subsequent fine-mapping using mass spectrometry-based genotyping of 42 loci refined the candidate interval and revealed a SNP within intron 2 of the ASIP gene, located in a genomic fragment with highly similar sequences, showing complete association with the gray coat color. Association analysis in an expanded population further confirmed a strong correlation between this variant and the gray phenotype. Gene expression analyses also supported the role of ASIP in regulating pigmentation in donkeys. ConclusionsThese findings identify a genetic determinant of gray coat color in donkeys and provide new insights into the molecular mechanisms underlying age-related depigmentation in domestic animals.

13
Blood-based transcriptomic classification of lung cancer: a leakage-free nested cross-validation framework with LASSO

Bakim, S.; UrluOzalan, N.; Gulbahce Mutlu, E.; Demir, V.; Gulbahce, E.

2026-07-13 oncology 10.64898/2026.07.11.26357823 medRxiv
Top 0.3%
3.2%
Show abstract

Peripheral whole-blood gene expression profiling offers a minimally invasive route to lung cancer detection, but high-dimensional transcriptomic data are prone to optimistic bias when preprocessing and model selection are not properly separated from performance evaluation. We applied L1-penalised (LASSO) logistic regression to 303 peripheral whole-blood microarray profiles (123 lung cancer cases and 180 healthy controls; Gene Expression Omnibus accession GSE252168; Illumina HumanHT-12 v4) within a leakage-free nested cross-validation framework (5 outer and 3 inner folds), in which all data-dependent steps (imputation, univariate feature screening by ANOVA F-test with k = 500, and standardisation) were confined strictly to training partitions. Statistical significance was assessed by permutation testing (B = 100), and feature selection stability was quantified across outer folds. LASSO was compared with ridge logistic regression, linear support vector machines, and random forest under the same framework. The LASSO model identified a sparse 29-probe signature with a pooled out-of-fold area under the ROC curve (AUC) of 0.990 (nested estimate 0.989 +/- 0.015), accuracy 97.4%, sensitivity 94.3%, and specificity 99.4% at a 0.50 threshold; permutation testing confirmed significance (p = 0.0099). Six probes, including CDC42, U2AF1, and RPS15A, were selected in all five outer folds, forming a stable core, and all classifiers exceeded AUC 0.987, indicating a strong, algorithm-independent signal. A leakage-free nested cross-validation framework enables unbiased performance estimation and reproducible feature selection in blood-based lung cancer classification. The 29-probe panel is an internally validated candidate requiring prospective, multicentre external validation before clinical use.

14
DNA 6mA marks transcriptionally active chromatin in malaria parasites

Seshan, D.; Lauer, W.; Sarkar, G.; Govindasamy, M.; Murray, C. S.; Greer, E. L.; Smith, M. L.; Vembar, S. S.

2026-06-13 genomics 10.64898/2026.06.12.732001 medRxiv
Top 0.3%
3.1%
Show abstract

DNA N6-methyladenine (6mA) has emerged as a significant epigenetic modification across a broad range of eukaryotes, from unicellular protists to metazoa. However, its role in unicellular eukaryotic parasites with highly AT-rich genomes, such as malaria-causing Plasmodium falciparum, remains unclear. Using mass spectrometry, South-western blotting, and Single Molecule Real-Time sequencing (Pacific Biosciences) across four stages of P. falciparum intra-erythrocytic development (IED), we confirmed that 0.02-0.04% of genomic adenines are modified to 6mA, with over 60% of the sites being stably maintained during the IED cycle. Notably, 6mA is enriched at transcription start sites, with genes bearing 6mA marks within their 5 and 3 untranslated regions exhibiting significantly elevated steady-state transcript levels. Consistent with this, 6mA loci show a strong positive correlation with activating histone post-translational modifications, while showing no significant association with repressive histone marks. Furthermore, in contrast to unicellular ciliates such as Oxytricha and Tetrahymena - organisms that share ancestry with Plasmodium - 6mA-marked genomic regions do not occlude nucleosomes. Lastly, we identified a putative 6mA methyltransferase belonging to the METTL4 family in P. falciparum, PfN6AMT encoded by the PF3D7_1303100 gene, and demonstrate that recombinant PfN6AMT exhibits robust methyltransferase activity in vitro, with mutation of its active site residues abolishing catalytic activity. Collectively, our findings demonstrate that 6mA is a low-abundance, yet reproducible, feature of the P. falciparum epigenome that is associated with transcriptionally active chromatin, and that the molecular mechanisms governing DNA adenine methylation may have undergone substantial evolutionary divergence, even among closely related eukaryotic lineages.

15
Synonymous codon usage is biased for and against m⁶A RRACH motifs in mammals

Creasey, L. D.; Tauber, E.

2026-07-14 genomics 10.64898/2026.07.09.737571 medRxiv
Top 0.3%
3.1%
Show abstract

RNA methylation at N6-adenosine (m6A) predominantly occurs within RRACH motifs, yet the forces shaping these motifs in coding regions remain unclear. Here we show that the synonymous codon combinations able to form or disrupt RRACH sites are used non-randomly across mammals. Using 13,491 protein-coding genes from 261 species, we identified genes significantly enriched or depleted in RRACH motifs, a pattern consistent with gene-specific selection for or against m6A potential. Genes enriched in RRACH sites were linked to ubiquitin-like conjugation and cell cycle regulation, whereas transmembrane and HOX genes were RRACH-poor, likely reflecting sequence incompatibility with CpG dinucleotides. Cross-species comparison with Caenorhabditis elegans, which lacks mRNA m6A methylation, revealed reciprocal RRACH frequencies, as expected if these motifs are under selection in m6A-competent genomes but evolve without this constraint otherwise. At the codon level, specific amino acid pairs, particularly threonine-ending dyads, were biased toward RRACH-forming codons while others were depleted, indicating that synonymous codon choice is skewed for and against motif formation. RRACH motifs were also non-randomly distributed along coding sequences, depleted near start codons and enriched toward the 3' end, consistent with known m6A profiles. Finally, analysis of cancer mutations revealed tissue-specific gain and loss of RRACH sites, reflecting context-dependent remodeling of methylation potential. Together, these results show that synonymous codon usage is systematically biased for and against m6A RRACH motifs, pointing to an evolutionary coupling between the genetic code and the epitranscriptomic landscape.

16
The role of long-range transcriptional regulation in interpretation of non-coding variants associated with human disease

Mandic, K.; Hrsak, D.; Uljanic, F.; Lenhard, B.; Baresic, A.

2026-06-17 genomics 10.64898/2026.06.15.731051 medRxiv
Top 0.3%
3.1%
Show abstract

Genome-wide association studies (GWAS) are the key tools for the discovery of associations between single nucleotide polymorphisms (SNPs) and phenotypic traits and have been successfully applied to many diseases and disorders. However, a great challenge is to find the gene affected by the non-coding fraction of SNPs, especially if the gene is distal in terms of genomic distance. In this study, we present a novel approach, named targPred, which utilises genomic regulatory blocks (GRBs) for inference of a connection between a certain SNP/locus and the target gene located in the same GRB, in a more robust and generalisable manner. We identified that many disease traits such as cancer and psychiatric disease have a propensity for long-range regulation. Furthermore, we showcased a childhood obesity locus which is connected to the distal BDNF gene. Finally, we propose a new web-based service based on enhancer-promoter association, to facilitate finding the causal genes for a wide array of traits and conditions.

17
The epigenomic landscape of deep lineage divergence: The case of the European sea bass

Longo, A.; Babbucci, M.; Jiao, Z.; Ferraresso, S.; Franch, R.; Bortoletti, M.; Bertotto, D.; Faggion, S.; Ilsley, G. R.; Papadogiannis, V.; Manousaki, T.; Kristoffersen, J.; Tsigenopoulos, C. S.; Macqueen, D. J.; Bargelloni, L.

2026-07-28 evolutionary biology 10.64898/2026.07.27.738391 medRxiv
Top 0.3%
3.0%
Show abstract

BackgroundUnderstanding the role of non-coding genomic variation in speciation remains a major challenge in evolutionary biology. Here, we investigated whether regulatory elements contribute to this process between Atlantic and Mediterranean lineages of European sea bass (Dicentrarchus labrax), a well-characterized case-study near speciation where barriers to introgression exist in the presence of connectivity between diverging populations. ResultsWe generated a novel, highly contiguous genome assembly, which was annotated at the epigenomic level using ATAC-seq and ChIP-seq with six embryonic developmental stages and five tissue types in adult fish, identifying thousands of promoters, enhancers, and open chromatin regions. Integrating this annotation with whole-genome sequence data from 65 individuals across three geographically distinct populations, we identified 57,505 outlier SNPs and 332 structural variants (SVs) showing elevated differentiation between Atlantic and East Mediterranean lineages. Outlier SVs affected key regulatory elements and coding genes, while outlier SNPs were enriched in regulatory elements, particularly enhancers active in adult tissues. Local genomic divergence correlated positively with regulatory element density, especially on chromosomes 1, 9, and 18, which are enriched in genes related to osmoregulation, immune response, and oxidative stress -- processes relevant to adaptation across contrasting marine environments. ConclusionsThese findings support a major role for regulatory variation in driving deep lineage divergence through local adaptation.

18
Phasing of the 'Wonderful' Pomegranate Genome Using Haploid DNA Extracted from Pollen Grains

Lana, G.; Traband, R.; Ferrante, S. P.; Resendiz, M.; Yu, L.; Qu, H.; Eurmsirilerd, E.; Deng, Z.; Roose, M.; Merhaut, D.; Beaulieu, T.; Seymour, D.; Gmitter, F.; Jia, Z.; Chater, J.

2026-07-18 genomics 10.64898/2026.07.13.738218 medRxiv
Top 0.4%
2.7%
Show abstract

The scientific and commercial interest in pomegranate (Punica granatum L.) cultivation has increased noticeably during the last two decades. Because of the high concentration of bioactive compounds and its promising nutraceutical properties, pomegranate has been defined as a functional food. In order to develop advanced genomic tools to improve pomegranate breeding program efficiency, we present the chromosome-scale and haplotype-resolved genome assembly of Wonderful, a pomegranate cultivar widely grown around the world. DNA isolated from diploid leaf tissues was sequenced using long read sequencing technology (PacBio and Nanopore), while DNA extracted from haploid pollen grains was sequenced using a short-reads platform (Illumina). Genomic data from 11 single haploid gamete cells were analyzed using the R package called Hapi to phase the genome. The final genome assembly size was of 372.51 Mbp anchored to eight pairs of homologous chromosomes. The present study provides an insight on the adoption of an innovative and efficient approach for the assembly of haplotype-resolved genomes, which enables a higher resolution of DNA variant detection and offers the opportunity to investigate crossover events in single gamete cells during meiosis.

19
Genomic insights into the karyotypic radiation of a narrow endemic holocentric plant Carex helodes

Gomez-Ramos, I.; Sanchez-Villegas, R.; Mohan, A. V.; Cornet, C.; Marques, A.; Maguilla, E.; Martin-Bravo, S.; Lucek, K.; Escudero, M.

2026-07-18 genomics 10.64898/2026.07.14.738159 medRxiv
Top 0.4%
2.7%
Show abstract

Holocentric chromosomes allow rapid genome changes through chromosomal rearrangements such as fissions, fusions, inversions or translocations. The plant genus Carex shows one of the highest rates of karyotypic evolution among holocentric organisms. We studied the genomic patterns underlying chromosomal rearrangements in the karyotypic radiation of the narrow endemic species Carex helodes (2n = 68-75). Comparing genome assemblies of C. helodes from the two karyologically distinct extremes of its European distribution, revealed a striking number of eight chromosomal rearrangements including fusions, translocations and inversions. Genomic breakpoints are gene-poor and TE-rich, corroborating findings in other species and suggesting common genomic characteristics that facilitate the evolution and establishment of chromosomal rearrangements. We identified a chromosomal inversion exhibiting patterns of purifying selection and enrichment in functional genes that potentially mediate rearrangement tolerance. Conversely, another inversion displayed elevated sequence divergence and enrichment in response to temperature stress and phosphate limitation, matching key environmental variables that differ between the study localities. The establishment of chromosomal rearrangements along Carex helodes European populations was likely driven by demographic bottlenecks and distinct genomic features at breakpoints. Our findings provide preliminary evidence on the rearrangement role in population differentiation either as reproductive barriers or as genomic islands of differentiation.

20
Genome-wide association study of susceptibility to pneumococcal carriage amongst children

Kandasamy, R.; Gurung, M.; Shrestha, S.; Bibi, S.; Thorson, S.; Carter, M.; O'Connor, D.; Murdoch, D. R.; Kelly, D. F.; Shrestha, S.; Levin, M.; Pollard, A. J.

2026-07-16 genetic and genomic medicine 10.64898/2026.07.13.26356474 medRxiv
Top 0.4%
2.6%
Show abstract

Background Pneumococcal disease is a leading cause of paediatric pneumonia and meningitis. Pneumococcal colonisation is the fundamental step to pneumococcal disease causation. We aimed to identify genetic loci associated with pneumococcal colonisation amongst children. Methods We conducted a genome-wide association study on 2111 Nepalese children, comprising 1346 cases carrying pneumococcus and 765 controls. We tested 8.1 million imputed variants using logistic regression and ten principal components as covariates. Fine mapping and functional evidence were used to identify suspected causal variants and related genes of interest. Findings A cluster of 22 variants of genome-wide significance (p<5x10-8) were identified on chromosome 12q21.31, eight of which were within PPFIA2. Fine mapping of this region identified 5 variants within 0.1 Mb of the 5-prime region of PPFIA2 all of which are significant eQTLs for PPFIA2. We further describe three loci (10q23.31, 12q23.1, and 20p11.21) which had variants with highly suggestive associations (p<5x10-7)with pneumococcal carriage. Interpretation Our study demonstrate human susceptibility to pneumococcal carriage to be polygenic with genetic variations which regulate PPFIA2 expression playing a key role in the ability for pneumococcus to colonise children. Targeting these genetic factors and the associated pathways are a means for preventing pneumococcal disease. Funding This study was supported by funding from Gavi - the vaccine alliance, the European Unions Horizon 2020 research and innovation program under grant agreement number 668303 (PERFORM), and a Robert Austrian Research Award.