Back

Genomics

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Genomics's content profile, based on 64 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.

1
Evolutionary Stratification of Codon Usage Bias In Plants Arises from GC3 Composition and Translational Optimization

Mohanta, T. K.

2026-07-01 genomics 10.64898/2026.06.26.734692 medRxiv
Top 0.2%
4.3%
Show abstract

Codon usage bias is a fundamental genomic characteristic that prefers non-random preferential use of synonymous codons. It is a major determinant of translational efficiency, gene regulation, and molecular evolution. However, the evolutionary bias and functional relevance of codon usage bias across the plant lineage is poorly defined and yet to understand what are the major factors responsible for relative synonymous codon usage (RSCU) in genomes and how codon usage bias influences the gene regulation, molecular evolution genomes. A genome-wide codon usage bias study of coding DNA sequences of 262 plant genome was conducted. It encompassed more than 4.6 billion codons from > 11 million coding sequences. Relative synonymous codon usage, codon adaptation index, codon-anticodon mapping, effective number of codon (ENC)-GC3, GC1,2-GC3, parity rule 2 (PR2-bias), molecular economy, and machine learning approaches were used for the study. It was found that codon usage bias was strongly non-random and exhibited a clear phylogenetic structuring. The higher plants favoured A/T-ending, whereas early-diverging lineages were enriched in G/C-ending codons. Analysis of RSCU, codon adaptation index, and codon-anticodon pairing indicated that translational selection is mediated by tRNA availability, contributing sustainability to these molecular patterns. Machine-learning approaches identified a small subset of codons having outsized influence on genome-wide codon usage landscapes. Further studies revealed the presence of robust inverse relationships between the effective number of codons and GC content at synonymous third positions. Neutrality analysis revealed approximately 61% of variation was driven by mutational pressure, tempered by selective constraints. Phylogenetic reconstruction showed a progressive relaxation of codon bias from algae to angiosperms while maintaining a conserved molecular economy cost of ~ 30 ATP per codon across the lineages. The study revealed codon usage bias is lineage-specific evolutionary conserved trait governed by mutation, selection, and translational optimization.

2
Mapping non-coding functional elements in allotetraploid Cyprinus carpio embryo development reveals subgenome variation of transcription regulation

Jimenez-Gonzalez, A.; Madrero Pardo, A.; Hadzhiev, Y.; Blasweiler, A.; Zunar, B.; Csenki-Bakos, Z.; Muller, T.; Megens, H.-J.; Wiegertjes, G. F.; Lenhard, B.; Baranasic, D.; Mueller, F.

2026-06-26 genomics 10.64898/2026.06.22.733679 medRxiv
Top 0.3%
3.2%
Show abstract

Common carp (Cyprinus carpio) is an important freshwater species for ornamental and aquaculture purposes, and a key cyprinid model for studying allotetraploidy. Its two chromosomally-separated subgenomes show distinct gene expression profiles, but how their regulatory landscapes control gene expression dynamics during development remains unknown. We generated a regulatory atlas by combining transcriptomes across 12 developmental stages with chromatin accessibility maps, transcription start sites and gene regulation-associated histone post-translational modifications. Subgenome-specific annotation and comparison of 254,276 developmental regulatory elements (PADREs) revealed that regulatory subgenome divergence is most prominent during early development, converging toward the phylotypic period, mirroring expression convergence between subgenomes at the same stages. This dynamic was driven by enhancers, while promoters maintained a more stable subgenome bias, extending the hourglass model of developmental constraint to allotetraploid subgenome regulation. Subgenome-specific enhancers were preferentially retained in subgenome B, whereas subgenome A shifted toward homeologous enhancer activity near the phylotypic stage, indicating directional regulatory divergence between subgenomes. Comparison with zebrafish revealed high concordance with sequence conservation and that subgenome B retained more ancestral cyprinid regulatory elements than subgenome A. This developmental regulatory atlas provides a foundational resource for investigating cis-regulatory evolution following the fourth round of vertebrate genome duplication.

3
A gapless Landrace pig genome resolves centromeres and telomeres and highlights telomere repeat structures in different pig breeds

Grove, H.; Stenlokk, K. S. R.; Lien, S.; Gjuvsland, A. B.; Arnyasi, M.; van Son, M.; Kent, M.

2026-06-30 genomics 10.64898/2026.06.25.734473 medRxiv
Top 0.4%
2.7%
Show abstract

Abstract The Duroc-derived reference genome Sscrofa11.1 has provided a critical foundation for pig genomics, providing a high-quality reference genome for accurate variant detection and comparative genomics but does not capture breed-specific variation. Here, we present a near-complete, gap-free genome assembly for the Landrace pig (Landrace_v1, GCA_963921485.1), spanning all 20 chromosomes and totaling 2.6 Gb, including 176 Mb of sequence absent from Sscrofa11.1. Comparative analyses with recently published high-quality pig genomes reveal a conserved centromere organization across breeds, accompanied by substantial variation in repeat composition and length, and identify a pig specific pattern of telomere variant repeats across eight pig breeds. The improved resolution of repetitive regions in Landrace_v1 enables more complete reconstruction of complex gene families, including olfactory receptors, and uncovers structural variation at the KIT proto-oncogene receptor tyrosine kinase locus not represented in the Duroc reference. Together, these findings highlight the limitations of single-reference genomes and demonstrate the value of breed-specific assemblies for capturing genomic diversity and improving downstream analyses.

4
Comparative analysis of Illumina and Ultima-Genomics sequencing for plasma cell-free small RNA profiling in pancreatic cancer

Levon, A.; Volkov, H.; Shlayem, R.; Shomron, N.

2026-06-25 genomics 10.64898/2026.06.21.733585 medRxiv
Top 0.5%
2.2%
Show abstract

Plasma-derived cell-free small non-coding RNAs are promising non-invasive biomarkers for cancer detection and monitoring. However, variability in sequencing output limits standardization, and cross-platform performance for plasma small RNA profiling has not been systematically evaluated. Illumina short-read sequencing is the current standard, whereas the newcomer, Ultima-Genomics platform, has been less extensively studied for circulating small RNA in plasma. To directly compare platform performance, we sequenced plasma cell-free RNA from 39 patients with pancreatic cancer and 39 matched controls on both platforms. After filtering, Ultima-Genomics retained more mature microRNA reads, whereas Illumina achieved slightly higher enrichment efficiency and mapping rates. Despite these technical differences, both platforms produced concordant expression profiles, with strong cross-platform correlations for shared microRNAs and clear separation of cases and controls within each dataset. Differential expression analysis identified 14 significant microRNAs on both platforms with concordant directions of change, most of which are supported by pancreatic cancer databases. Pathway enrichment analysis highlighted signaling pathways implicated in pancreatic cancer, supporting the biological relevance of both shared and platform-specific signatures. These findings indicate that both Illumina and Ultima Genomics platforms are suitable for plasma small RNA profiling and capture biologically relevant signals in pancreatic cancer.

5
Linking geography and mutation profiles across goat species

Bionda, A.; Crepaldi, P.; Prendergast, J. G. D.; Neupane, M.; Amills, M.; Rosen, B. D.; Tosser-Klopp, G.; Milanesi, M.; Talenti, A.; The VarGoats Consortium,

2026-07-14 genomics 10.64898/2026.07.09.737241 medRxiv
Top 0.7%
1.9%
Show abstract

Recent studies have characterised the mutational profile across multiple mammalian species, highlighting substantial differences across lineages. However, none of these studies investigated whether mutation profiles and geography are significantly correlated. In this study, we present a multi-genome alignment spanning several Capra taxa, reconstruct the ancestral genome of Capra hircus and use it to characterize the mutational profiles across multiple Capra species by using the 1000 genomes VarGoats dataset. Results confirmed that the scale of differences among Capra species largely reflects their phylogenetic relationships, in particular with the Bezoar being genetically closer to domestic goats than to other wild species. Subsequently, we correlated the mutational profile and the geographical origin of the different individuals. In particular, ACG>ATG changes have the strongest correlation with longitude (r = -0.79, P-value = 3.02*10-204), while TCA>TGA are strongly correlated with latitude (r = -0.51, P-value = 4.30*10-63). We highlight how sequential dinucleotide mutations (SDMs) place cosmopolitan breeds closer to the sampling location, rather than the country of origin, showing how the recent relocation of cosmopolitan breeds to new continents is reshaping the genome of these animals. Finally, we used the mutational profile to predict the coordinate of origin of each animal in the dataset. In conclusion, we show the important role that geography had in shaping the genomes of domestic goats.

6
Genomic and Functional Insights into the Cluster V Mycobacteriophage ‘EniyanLRS’ and its therapeutically relevant LysB

Nadar, K.;Eniyan, K.;Bajpai, U.

2026-06-27 Molecular Biology 10.64898/2026.06.26.734815 medRxiv
Top 0.7%
1.8%
Show abstract

Drug-resistant tuberculosis and the rising incidence of nontuberculous mycobacterial (NTM) infections are a growing concern that demands innovative therapeutic strategies. Despite advances in diagnostics, drug discovery, and vaccine strategies, significant gaps remain. Mycobacteriophages and their lytic enzymes offer a promising solution due to their natural abundance and diversity, host specificity and ability to disrupt complex cell envelopes and biofilms. In this study, we report the genomic and functional characterization of a V-Cluster mycobacteriophage, EniyanLRS, isolated near a hospital in Delhi and the encoded endolysins LysA and LysB. EniyanLRS features a 78.53 kbp genome with a notably low GC content (56.9%) as compared to other mycobacteriophages, and an exceptionally long Tape Measuring Protein (TMP) gene (5.97 kbp). Its genome lacks genes related to lysogeny and harbours 24 tRNAs, suggesting high translational efficiency. Phenotypically, EniyanLRS exhibits a siphovirus morphology, lytic lifecycle and infects Mycobacterium smegmatis and drug-resistant Mycobacterium fortuitum. LysA, with its lysozyme-chitinase-amidase domain architecture, did not demonstrate significant antibacterial or antibiofilm activity. Conversely, LysB, an /{beta}-hydrolase, exhibited superior in vitro esterase activity compared to previously reported LysB enzymes and showed pronounced cell wall disruption of M. smegmatis and M. fortuitum, along with considerable antibiofilm efficacy (62.77% and 41.91% inhibition, respectively). Collectively, these findings highlight the potential of EniyanLRS and its LysB enzyme as potent biocontrol agents against pathogenic mycobacteria, which can be explored to treat planktonic cells and biofilm-associated infections.

7
Nitrogen use efficiency in pigs is associated with transcriptomic signatures related to amino acid metabolism, immune activity, and nutrient partitioning

Monney, B.; Ewaoluwagbemiga, E. O.; Kasper, C.

2026-07-01 genomics 10.64898/2026.06.26.733976 medRxiv
Top 0.8%
1.7%
Show abstract

Dietary protein restriction challenges the allocation of amino acids to growth and other physiological functions and therefore requires coordinated metabolic adaptation. Domestic pigs provide an informative system in which to study such responses, because nitrogen retention directly affects lean growth and can be quantified accurately under controlled feeding and housing conditions. Under reduced-protein diets, pigs differ in how effectively they retain nitrogen, and this variation has a genetic basis, making them well suited to investigate the molecular regulation of nitrogen use efficiency (NUE). Here, we characterise differential gene expression and enriched pathways in liver and skeletal muscle of more than 80 pigs with two divergent NUE phenotypes (high and low) maintained under the same protein-reduced, ad libitum dietary conditions. The two NUE phenotypes were clearly distinct at the transcriptomic level, with 177 differentially expressed genes in the liver and 133 in the muscle. In the liver, differential expression and enrichment analyses indicate reduced amino acid catabolism, lower inflammatory and detoxification activity, and a metabolic state that favours lipid processing and insulin-related regulation over the use of amino acids as energy sources. In skeletal muscle, they point to reduced lipid uptake, lower reliance on amino acid oxidation, and a greater emphasis on protein synthesis, translational regulation, mitochondrial energy metabolism, and growth-related processes. These gene-level patterns were supported and extended by pathway and gene-set enrichment analyses. Together, the results suggest that high and low-NUE pigs differ through coordinated, tissue-specific molecular adaptations. Overall, variation in NUE appears to reflect coordinated, tissue-specific differences in how nutrients are allocated between energy use, storage, and lean tissue growth.

8
CNSigs: An R Package for the Identification of Copy Number Mutational Signatures

Tallman, D.; Striker, S.; Byappanahalli, A. M.; Stockard, S.; Jenison, J.; Collier, K. A.; Blige, E.; Vater, M.; Stover, D. G.

2026-06-25 bioinformatics 10.64898/2026.06.21.733646 medRxiv
Top 1%
1.2%
Show abstract

BackgroundCopy number aberrations (CNAs) are gains and losses of large genomic segments present across most cancer types and are a hallmark of cancer genomic alterations. However, the processes underlying CNAs and characteristic patterns of CNAs are poorly understood. Bioinformatic advances have identified underlying single nucleotide variant (SNV) mutational signatures resulting from distinct mutational processes, yet development of algorithms able to uncover similar signatures for CNAs remains less advanced. MethodsUsing segmented data files from DNA sequencing, six copy number features are extracted for signature determination: segment size, breakpoints per 10 megabases, copy number oscillation events, average changepoint size, average copy number, and breakpoints per chromosome arm, along with ploidy. Mixed model approaches and non-negative matrix factorization (NMF) are utilized to derive CNA signatures across cancer types. The full methodology was packaged in a robust R package, termed CNSigs that is publicly available. ResultsTo verify the reproducibility of the signatures, we derived five signatures from two independent breast cancer datasets (total n>3000), demonstrating high accuracy (average cosine similarity = 0.89). Pan-cancer application of CNSigs in the TCGA dataset resulted in derivation of 13 pan-cancer signatures which were significantly associated with disease-specific survival. Benchmarking CNSigs to two other CNA signature approaches within TCGA demonstrated non-overlapping signatures and favorable compute speed for CNSigs. We evaluated n=24 pairs of tumor and circulating tumor DNA (ctDNA) acquired at the same time and demonstrated that CNSigs are detectable and reproducible via ctDNA, with significant association of CNSig11 with metastatic triple-negative breast cancer progression-free survival for taxane but not platinum or capecitabine chemotherapy. CNSigs association with immunophenotype was evaluated in low-grade glioma (LGG) and CNSig 3 was found to be highly prognostic for LGG yet complementary to immune features. ConclusionsThe CNSigs R package allows researchers to easily analyze their own samples to derive copy number signatures and evaluate clinical associations. We demonstrate potential application in ctDNA and association with treatment response. The development of this package allows further investigation of underlying processes that may be responsible for these CNA fingerprints.

9
Chromosome organization of Entamoeba histolytica and Entamoeba dispar

Kawano-Sugaya, T.; Kobayashi, S.; Kawashima, A.; Saito-Nakano, Y.; Izumiyama, S.; Nozaki, T.; Nakada-Tsukui, K.

2026-07-09 genomics 10.64898/2026.07.06.736064 medRxiv
Top 1%
1.1%
Show abstract

Entamoeba histolytica is a clinically important pathogenic eukaryote and the causative agent of amoebic dysentery. Entamoeba dispar, a nonpathogenic commensal species that resides in the human colon, is the closest sibling species, and serves as an appropriate comparator for genome-wide analysis. Although the genome of E. histolytica is approximately 26.9 Mb, and the largest known genome within the genus, that of E. invadens, is approximately 40.9 Mb, obtaining high-quality assemblies in this genus has remained challenging due to extensive repetitive regions, tRNA gene arrays, and aneuploidy. Here, we used PacBio HiFi sequencing to assemble the genomes of the pathogenic E. histolytica and the nonpathogenic E. dispar. We reconstructed all 36 chromosomes of E. histolytica and 35 chromosomes of E. dispar, assembling each as a single continuous DNA sequence (contig). The two species exhibited high genome-wide nucleotide similarity and conserved synteny at the amino acid level. At one end of each chromosome, we identified tRNA arrays, whereas the opposite end lacked such arrays, resulting in an asymmetric chromosomal architecture. Analysis of unique-read depth revealed widespread aneuploidy in both species: E. histolytica is predominantly tetraploid, whereas E. dispar is diploid, a conclusion further supported by SNP allele-frequency distributions. These assemblies provide a robust foundation for comparative genomics in Entamoeba and offer detailed insights into chromosome-end structure and ploidy.

10
Systematic benchmarking of low-input whole exome sequencing workflows for longitudinal ctDNA profiling in pancreatic ductal adenocarcinoma

James, L. G.; Thorn, G. J.; Morel, C.; PCRFTB, ; Kocher, H. M.; Ross-Adams, H. E.; Chelala, C.

2026-07-03 genomics 10.64898/2026.06.29.734743 medRxiv
Top 1%
1.1%
Show abstract

Whole exome sequencing (WES) of circulating tumour DNA (ctDNA) enables longitudinal monitoring of tumour dynamics, evolution and treatment response but remains technically challenging in low-input, low-shedding settings such as pancreatic ductal adenocarcinoma (PDAC). Here, we systematically compared three commercially available low-input WES workflows incorporating Agilent (V6, V8) and Qiagen exome capture designs using ultra-low input cfDNAs extracted from multiple matched longitudinal plasma samples from PDAC patients. Using predefined performance metrics including coverage, duplication rate and variant detection and additional metrics relevant for clinical genomic profiling in patient care, we show that all three workflows produced high-quality sequencing data, even from very low input cfDNA. Within the conditions tested here, the Agilent V8 workflow provided the most favourable balance of coverage uniformity, sequencing efficiency and hotspot coverage for low input, low tumour fraction cfDNA WES. These findings demonstrate that workflow design, including capture footprint, substantially influences ctDNA WES performance in low-input clinical contexts. These findings are particularly relevant in early stage and/or minimal residual disease settings, where tumour fractions are low and recovery of genomic information from limited-input samples is critical.

11
Identification of implications of m6A regulators and autophagy-associated genes for prognosis in ovarian cancer

Chen, Y.; Yu, X.; Chu, W.; Shang, S.; He, N.; guo, l.

2026-06-29 obstetrics and gynecology 10.64898/2026.06.25.26356535 medRxiv
Top 1%
1.1%
Show abstract

The most prevalent RNA alteration in the mammalian genome is N-6-methylenediosine (m6A). There is mounting evidence linking dysregulation of m6A regulatory factors and alterations in m6A levels to the development, course, or prognosis of ovarian cancer. Genes having prognostic value were screened using the univariate, multifactorial, and Least Absolute Shrinkage Selection Operator (LASSO) Cox regression analyses. Important genes' m6A expression in clinical material was verified by real-time fluorescent quantitative polymerase chain reaction (RT-qPCR). In present study, all 23 regulators were significantly differentially expressed in ovarian cancer tissues. LASSO regression analysis screened for 10 key genes associ-ated with both autophagy and m6A. A risk score was constructed and nomogram was developed to forecast the prognosis of ovarian cancer patients. Additionally, individuals with ovarian cancer were classified as high-risk or low-risk; and the low-risk group might be more likely to benefit from im-munotherapy. RT-qPCR was used for the bioinformatics study of human ovarian cancer and normal tissues. Lastly, PLK2 and LEPR were confirmed to be associated with tumorigenesis in scRNA-seq. The risk score established by m6A and autophagy can be used to predict prognosis and susceptibility to anticancer drugs in patients with ovarian cancer.

12
H3K4me3 exhibits length-dependent deposition patterns at transcription initiation regions in Trypanosoma cruzi and correlates with transcriptional activity

Lopez, M. d. R.; Gitman, I. F. B.; Prego, A. F.; Lavignolle-Heguy, R.; Zambrano-Siri, R. T.; Carena, S.; Arguello, R. J.; Vilchez-Larrea, S. C.; Alonso, G. D.; Ocampo, J.

2026-06-29 genomics 10.64898/2026.06.26.734760 medRxiv
Top 1%
1.1%
Show abstract

In trypanosmatids genes, transcribed by RNA polymerase II do not have canonical promoters and are organized into directional gene clusters that mature into monocistronic transcripts by a co-transcriptional process known as trans-splicing. Even though gene expression is regulated mainly post-transcriptionally, it is currently understood that chromatin and epigenetics are also involved in this regulation. In eukaryotes, specific signals are normally required for the occurrence of an appropriate transcription initiation. Among them, trimethylation of histone H3 in lysine 4 is the most conserved signal normally detected at transcription start sites of actively transcribed genes. Unlike many model organisms, trypanosomes do not have defined promoters. Instead, transcription initiates in a bidirectional manner from dispersed regions coincident with divergent strand switch regions located between directional gene clusters (DGCs). In T. cruzi, H3K4me3 was observed at the origins of transcription coincident with divergent strand switch regions (dSSRs) in epimastigotes, but it has not been mapped throughout the whole genome at base-pair resolution or in other life stages so far. Here, we set up the CUT&RUN technique for T. cruzi epimastigotes and trypomastigotes. Consistent with a predominant post-transcriptional regulation along the life cycle, we did not find significant differences between life stages. We corroborated that H3K4me3 is enriched at dSSR adjacent to actively expressed DGCs. Moreover, we noticed that this histone mark exhibits different patterns that correlate with the genomic span of the transcription initiation regions and with transcriptional activity. Furthermore, we unveiled that the most actively transcribed DGCs are associated with shorter dSSRs and are located within the core compartment of the genome displaying a more accessible chromatin.

13
Genome-wide meQTL mapping in cattle blood reveals cis and trans regulation of DNA methylation

Fouere, C.; Costes, V.; Besnard, F.; Le Danvic, C.; Patry, C.; Fritz, S.; Boussaha, M.; Jouin, M.; Boichard, D.; Kiefer, H.; Costa Monteiro Moreira, G.; Sanchez, M.-P.

2026-07-08 genetics 10.64898/2026.07.07.736355 medRxiv
Top 1%
1.0%
Show abstract

Background Complex traits are influenced by numerous variants, most of which have regulatory effects on gene expression that can be mediated by DNA methylation. Molecular QTL mapping is an approach that aims to dissect these effects. However, obtaining molecular phenotypes on a large scale is challenging, particularly in livestock species. In cattle, an epigenotyping array called EpiChip has recently been developed in the European RUMIGEN project. The EpiChip, which contains 43,317 CpG sites distributed all over the bovine genome, enables large-scale measurement of DNA methylation. This study aims to characterize the genetic determinism of blood DNA methylation in cows by estimating heritability and mapping cis- and trans-methylation QTLs (meQTLs). Results Whole blood samples from 4,457 genotyped Holstein cows were epigenotyped. Across all CpG sites, the heritability estimates averaged 24.6%. The local meQTL mapping at sequence-level for variable CpG sites (SD > 2.5%; n = 28,806) detected cis-meQTLs for 80.1% of the CpG sites, with sentinel SNPs located close to their associated CpGs. A two-step analysis was also conducted to identify long-range associations, with a particular focus on trans-meQTL hotspots. First, we identified CpG-SNP trans-associations using medium-density genotypes (50k SNPs) that revealed 31,846 SNPs with significant effects on 1 to 530 trans-CpG sites. Then, regions associated with at least 34 independent trans-CpGs were retained defining 31 hotpots. For each hotspot, a local sequence-level GWAS was conducted using the first principal component derived from the associated trans-CpGs. Out of the 31 detected hotspots, three were located close to transcription factor genes (RUNX1, NFIC and FOXA3) for which the associated trans-CpGs were enriched for the corresponding binding motif. Two other hotspots were located within KDM5A and KDM5B, and their corresponding trans-CpGs were strongly overrepresented in H3K4me3 narrow peaks in blood as well as in other tissues. Conclusions By identifying functional candidate genes associated with blood DNA methylation in cattle, these findings provide new insights into the regulatory architecture of DNA methylation in mammals, highlighting the value of large-scale molecular data from livestock populations.

14
Genetic introgression and transcriptomic plasticity are associated with enhanced Leishmania infantum pathogenicity causing human cutaneous leishmaniasis in Tunisia

SANTI, A. M. M.; LI, B.; PIEL, L.; PIPOLI DA FONSECA, J.; BACQ-DAIAN, D.; OLASO, R.; DELEUZE, J. F.; COKELAER, T.; AOUN, K.; BOURATBINE, A.; SPÄTH, G. F.

2026-06-22 genomics 10.64898/2026.06.20.733521 medRxiv
Top 1%
1.0%
Show abstract

The protozoan parasite Leishmania infantum exhibits significant genetic variability among isolates, influencing disease manifestation and treatment response. Although L. infantum is classically described as the causative agent of Visceral Leishmaniasis (VL) - often associated with immune deficiency, cases of Cutaneous Leishmaniasis (CL) caused by this species in immunocompetent individuals have been reported in different countries. To investigate the molecular basis of this unusual shift in tissue tropism and pathogenicity, we applied comparative genomic and transcriptomic approaches on two canine isolates (CanL) and two human isolates associated with Cutaneous Leishmaniasis (CL) in Tunisia. While the CanL isolates showed close genetic similarity to the L. infantum reference strain (JPCM5), the CL isolates formed a separate, highly divergent cluster based on SNP localization and frequency, differing not only from JPCM5 but also from each other. Utilizing the metagenomics sequence classification tool Kraken, we revealed a complex hybrid nature of the CL isolates, showing introgression from L. donovani and L. tropica, suggesting that hybridization has played a key role in generating novel phenotypic traits. Integration of RNA-seq and DNA-seq data demonstrated that only a minority of gene expression variation within and in-between the CanL or CL groups reflected gene dosage effects due to copy number variation, while the majority of expression differences were independent of gene dosage, implying post-transcriptional regulatory mechanisms contributing to parasite adaptation. In conclusion, our study identifies hybridization, genome instability, and transcriptomic adaptation as interconnected drivers of the L. infantum evolutionary potential. These mechanisms can collectively enhance parasite fitness gain, potentially explaining the emergence of cutaneous disease forms in a species traditionally linked to visceral infection. Author SummaryThis research reveals that hybridization between distinct Leishmania parasite species could be a key mechanism driving the evolution of new disease forms. By demonstrating that cutaneous leishmaniasis (CL) cases are caused by hybrid L. infantum parasites whose genomes show introgression with DNA from L. donovani and L. tropica, this study reveals a molecular mechanism potentially linked to the emergence of tegumentary disease from a species traditionally known to cause visceral infection. These findings contribute to our understanding of Leishmania evolution, the emergence of atypical forms of leishmaniasis linked to hybridization, and the impact of genome instability and transcriptomic adaptation as potent forces for generating phenotypic diversity and enhancing parasite fitness.

15
Promoter Structural Variants are Drivers of Genome-Wide Differential Expression in Maize

Munasinghe, M.; Read, A.; Schulz, A. J.; Brandvain, Y. J.; Springer, N. M.; Hirsch, C.

2026-06-28 genomics 10.64898/2026.06.23.734061 medRxiv
Top 2%
1.0%
Show abstract

BackgroundStructural variants (SVs) are large insertions or deletions of DNA sequences. While less numerous than single nucleotide polymorphisms, SVs often account for a greater proportion of nucleotide differences between genomes. Their size and frequent association with repetitive sequences has historically hindered their detection, which has limited the ability to associate this variation with molecular and phenotypic trait variation. While some SVs have been linked to observable traits, it remains unclear whether such effects are rare or broadly distributed across the genome. ResultsTo test for genome-wide relationships between SVs and gene expression, we analyzed genome assemblies and transcriptomic data from 10 tissues across 26 diverse maize inbred lines. We identified SVs amongst these lines and examined variants located within the 1kb promoter region upstream of genes. Thousands of genes showed expression differences associated with promoter SVs, often in a tissue-specific manner. One common feature of these SVs was the presence of transposable element sequences. LTR retrotransposons were enriched amongst promoter SVs associated with differential expression and often reduced expression of the nearby gene. Despite widespread expression changes, we found no enrichment for specific biological functions or pathways among affected genes. ConclusionsOur findings indicate that extant TE-mediated promoter SVs play a significant role in shaping gene expression patterns across the maize genome. However, their phenotypic effects appear limited or context-dependent, suggesting that many variants may have minimal impact outside specific developmental stages or environmental conditions.

16
Long-range regulatory target prediction reveals shared genetic background across ulcerative colitis, Crohn's disease, primary sclerosing cholangitis and ankylosing spondylitis

Dulcic, D.; Mandic, K.; Hrsak, D.; Baresic, A.

2026-07-03 genomics 10.64898/2026.06.29.735270 medRxiv
Top 2%
0.8%
Show abstract

Common variants detected by the genome-wide association studies (GWAS) create a wealth of knowledge on genetic component of individual traits and diseases. Elucidating the molecular mechanism behind the vast majority of these variants that are found to be non-coding remains a largely unsolved task, especially when distal and pleiotropic interactions between regulatory elements where these variants occur and gene promoters are taken into account. Focusing on four diseases with immune-mediated mechanisms namely ulcerative colitis, Crohn's disease, primary sclerosing cholangitis and ankylosing spondylitis, we demonstrate the utility of the targPred tool, providing prediction of genes targeted by the regulatory variants. We demonstrate that taking into account evolutionary and comparative genomic data, previously unobserved mechanistic trends (the platelet, vascular and sterol clusters) can be detected in terms of implicated genes targeted by the regulatory elements containing common variants, shared between all four diseases, as well as specific trends for subsets of diseases, e.g. two IBD phenotypes. We also elucidate a clinically-relevant target COG6 shared between IBD and PSC, as well as a whole range of other target genes missed by the conventional SNP-to-gene assignments methods.

17
Three Plasmid Strategies, One Intermediate Convergence State: Lineage-Specific Resistance, Virulence Architecture in Dominant Indian Carbapenem-Resistant Klebsiella pneumoniae Clones

Kulkarni, S. M.; Jacob, J. J.; Rajendra, S.; S, P.; T, M. P.; Velmurugan, A.; Nelson, R.; Neeravi, A.; Balaji, L.; Gunasekaran, K.; Manesh, A.; Rajni, E.; Walia, K.; Veeraraghavan, B.

2026-06-24 microbiology 10.64898/2026.06.23.734134 medRxiv
Top 2%
0.8%
Show abstract

Carbapenem-resistant Klebsiella pneumoniae (CRKp) is a critical global healthcare threat driven by high-risk multidrug-resistant (MDR) clones that acquire hypervirulence genes. Although resistance-virulence co-occurrence is extensively documented, the plasmid-level mechanisms facilitating this convergence remain unclear. In this study, we utilized hybrid short- and long-read whole-genome sequencing of 376 clinical CRKp strains to define the evolutionary trajectories and structural plasmid dynamics of three predominant high-risk clones: ST147 (n=157), ST231 (n=108), and ST2096 (n=111). Carbapenemase genes were present in 90% of isolates, predominantly blaOXA-48-like and blaNDM-5 co-harbored with blaCTX-M-15. Virulence profiling indicated high aerobactin (iuc) prevalence (62.7%), while salmochelin and colibactin were undetected. Hypermucoviscosity occurred infrequently (6.6%) and was independent of rmpA/rmpA2, confirming a clear genotype-phenotype discordance. Comparative plasmid mapping revealed three distinct, lineage-specific plasmid configurations underlying this intermediate convergent pathotype: ST147 exhibited dynamic, mosaic hybrid IncFIB-IncHI1B plasmids; ST2096 showed structurally stabilized hybrids; and ST231 retained virulence and resistance determinants on separate, segregated plasmids. These findings show that convergence is regulated by multiple, clone-specific evolutionary routes rather than a single path, highlighting the critical need for more in-depth genomic surveillance capable of identifying convergent plasmids along with high-risk lineages

18
Head-to-head organized segmental paralogs AtOFP2 and AtOFP17 exhibit differential, spatio-temporal partitioning of function, and negative regulation of multiple developmental traits including seed-yield and root architecture

Chahar, N.; Pokhriyal, E.; Yadav, S.; Ren, B.; Dangwal, M.; Das, S.

2026-07-09 plant biology 10.64898/2026.06.30.735610 medRxiv
Top 2%
0.8%
Show abstract

Ovate Family Proteins (OFPs) are a class of plant-specific, negative nuclear transcriptional regulators characterized by conserved C-terminal OVATE domain. This study on comparative functional characterization of two head-to-head arranged OFPs - AtOFP2 (Ovate-OFP with full ovate domain) and AtOFP17 (Ovate-Like OFP with partial ovate domain) provides critical insight into how structural variations in ovate domain leads to functional divergence. Detailed phenotypic analysis of 28 physical and physiological traits of loss- and gain-of-function mutants revealed that both genes act as broad, pleotropic repressors of plant growth and development. Removal of repression in knock-down mutants of both genes exhibited reduced duration of seed dormancy, faster rate of germination and growth, bigger plants and significantly higher seed yield. In contrast, constitutive over-expression showed a generalized repressive nature of both genes, with nuanced differences for fine tuning of specific traits. For example, both genes showed antagonistic behaviours on root hair architecture. AtOFP2 act as a strong repressor of root hair development whereas AtOFP17 is a stronger repressor of hypocotyl and root cell architecture. AtOFP17 owing to partial ovate domain exerts a mild level of repression throughout life span as indicated by smaller plants and lesser yield in knock-down AtOFP17 mutants. On the contrary, AtOFP2 exerted a much stronger repressor effect in which > 90% over-expression mutants died at the juvenile stage ; the survival of remaining 10% is probably owing to activation of dosage-dependent feedback loop mechanism as indicated by normal growth of mature plants, and is also evident by transcriptome data. Transcriptome analysis of roots of 7-day old seedling of knock-down and over-expression mutants of AtOFP2 showed downregulation of OFP2 in over-expressed mutants. However, severely stunted phenotype indicated presence of stable OFP2 protein to exert effects. Analysis of DEGs in OFP2 mutants revealed that it acts as an important regulator working at intersection of hormonal signalling affecting critical genes required for auxin, cytokinin, GA, BR and ABA functioning. Perturbations across hormonal signalling pathways affects cell wall remodelling factors such as EXPANSINS, Xyloglucan hydrolases (XTHs) and cellulose synthases (CSLs) causing overall stunted growth; and epidermal patterning genes such as WER, GL1, EGL3, TTG1 leading to severely reduced root length and root hairs. Significantly, functional analysis of this master regulator highlighted a significant economic potential. Knockdown of both these genes relieves their natural repression on reproductive traits, leading to longer siliques, bigger and heavier seeds, and substantially increased overall seed yield, positioning AtOFP2 and AtOFP17 as highly valuable targets for agricultural crop improvement.

19
Studying the regulons of OmrA and OmrB paralogous small RNAs reveals targets involved in central carbon metabolism and lipogenesis

Korepanov, A.;Jagodnik, J.;Quenette, F.;LAM, T.;HAMON, M.;Fromont, J.;Sismeiro, O.;Gherdol-Nouvion, V.;Maes, A.;Guillier, M.

2026-06-27 Molecular Biology 10.64898/2026.06.26.734639 medRxiv
Top 2%
0.8%
Show abstract

Small regulatory RNAs (sRNAs) are key players in bacterial adaptation to stress. They often occupy central positions in regulatory networks and control the expression of multiple targets. In a striking example of this, the enterobacterial OmrA and OmrB paralogous sRNAs are known to regulate about ten different targets, with extensive data suggesting the regulon is in fact much larger. Here we performed transcriptome and proteome analyses and identified more than fifteen new targets of Escherichia coli OmrA and OmrB. We validated several, including genes involved in central carbon metabolism and fatty acid synthesis, among which ppc, actP and fabA. Consistent with a role in carbon metabolism, overproducing OmrA or OmrB inhibited growth on glucose minimal medium. The analysis of suppressor mutants shows that this is due to a decreased carbon flux through the TCA cycle. Incorporating other datasets such as RIL-seq, we generated a multi-omics-based prediction of target candidates. Together, our results show that OmrA/B base-pair to various regions of their mRNA targets, and therefore likely act through diverse regulatory mechanisms. Hence, this work extends the OmrA and OmrB regulons, establishes an unsuspected connection with carbon usage, and shows the benefits of combining global analyses to investigate sRNA regulons.

20
A Draft Male Genome Assembly of the Slipper Lobster (Thenus australiensis) Reveals an XY System and a Validated Diagnostic Marker for Monosex Aquaculture.

Tran Nguyen, A. H.; Ha, G.-H.; Tran, D.-P.; Le, N. T.; Glendining, S.; Fitzgibbon, Q.; Herzig, V.; Luu, P.-L.; Ventura, T.

2026-06-29 genomics 10.64898/2026.06.24.734161 medRxiv
Top 2%
0.8%
Show abstract

The slipper lobster (Thenus australiensis) is rapidly emerging as a high-potential species for commercial aquaculture. Because females exhibit superior growth characteristics due to less frequent moulting after sexual maturity, developing monosex breeding strategies is highly desirable for industry profitability. However, the lack of genomic resources and early sex-identification tools has hindered this development. Here, we report the first draft male genome assembly for T. australiensis, generated using a combination of whole-genome shotgun sequencing, DArT-seq, and multi-tissue transcriptomics. The curated assembly spans 0.913 Gbp with high functional completeness (93.0% BUSCO), providing a robust repertoire of 30,100 protein-coding genes. Through k-mer subtraction and population-level DArT-seq genotyping, we provide definitive evidence that T. australiensis utilizes an XX/XY sex-determination system. Crucially, by identifying male-specific structural variations within a neo-Y locus, we developed a diagnostic PCR assay targeting a male-exclusive sequence. This 171 bp marker achieved 100% accuracy in phenotypic sex identification across wild-caught populations. Ultimately, these foundational genomic resources, combined with a highly reliable molecular sexing tool, provide the critical framework necessary for early sex sorting, broodstock management, and the commercial advancement of monosex slipper lobster farming.