Back

Genomics

Elsevier BV

Preprints posted in the last 7 days, ranked by how well they match Genomics's content profile, based on 64 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.

1
Construction of a risk prediction model for postoperative bleeding in patients with thyroid cancer based on clinical data

zhang, y.; chen, w.; li, x.; shen, w.

2026-07-18 oncology 10.64898/2026.07.16.26358297 medRxiv
Top 2%
0.5%
Show abstract

Objective To develop and validate a risk model for predicting postoperative bleeding in patients with thyroid cancer. Methods A total of 2800 consecutive patients diagnosed with thyroid cancer in the Department of Thyroid and Breast Surgery of the Affiliated Hospital of Xuzhou Medical University between January 2020 and December 2023 were retrospectively analyzed. Patients were categorized into two groups based on postoperative bleeding occurrence: bleeding and non-bleeding groups. Univariate and multivariate logistic regression analyses were utilized to screen independent risk factors. Meanwhile, risk prediction models were developed and nomogram . Subgroup analysis was performed to identify independent risk factors. The predictive effects of the models were assessed using the Hosmer-Lemeshow test and receiver operating characteristic (ROC) curves. Results Of the 2800 recruited patients, 50 had postoperative bleeding, with an incidence rate of 1.7%. Multivariate logistic regression analysis showed that age, hypertension, total thyroidectomy, tumor size [&ge;]4 cm, and operation time [&ge;]90 min were the risk factors for postoperative bleeding in thyroid cancer patients (P<0.05). A risk prediction model was established based on the above factors, and the area under the ROC curve was 0.881, with a sensitivity of 94.0%, a specificity of 67.3%, and an accuracy of 74.0%. Decision curve analysis revealed that the model had good predictive ability. Conclusions The constructed risk prediction model has good predictive power and can provide a reference for healthcare professionals to predict the risk of bleeding in patients after thyroid cancer surgery.

2
Genome-wide association study of susceptibility to pneumococcal carriage amongst children

Kandasamy, R.; Gurung, M.; Shrestha, S.; Bibi, S.; Thorson, S.; Carter, M.; O'Connor, D.; Murdoch, D. R.; Kelly, D. F.; Shrestha, S.; Levin, M.; Pollard, A. J.

2026-07-16 genetic and genomic medicine 10.64898/2026.07.13.26356474 medRxiv
Top 3%
0.5%
Show abstract

Background Pneumococcal disease is a leading cause of paediatric pneumonia and meningitis. Pneumococcal colonisation is the fundamental step to pneumococcal disease causation. We aimed to identify genetic loci associated with pneumococcal colonisation amongst children. Methods We conducted a genome-wide association study on 2111 Nepalese children, comprising 1346 cases carrying pneumococcus and 765 controls. We tested 8.1 million imputed variants using logistic regression and ten principal components as covariates. Fine mapping and functional evidence were used to identify suspected causal variants and related genes of interest. Findings A cluster of 22 variants of genome-wide significance (p<5x10-8) were identified on chromosome 12q21.31, eight of which were within PPFIA2. Fine mapping of this region identified 5 variants within 0.1 Mb of the 5-prime region of PPFIA2 all of which are significant eQTLs for PPFIA2. We further describe three loci (10q23.31, 12q23.1, and 20p11.21) which had variants with highly suggestive associations (p<5x10-7)with pneumococcal carriage. Interpretation Our study demonstrate human susceptibility to pneumococcal carriage to be polygenic with genetic variations which regulate PPFIA2 expression playing a key role in the ability for pneumococcus to colonise children. Targeting these genetic factors and the associated pathways are a means for preventing pneumococcal disease. Funding This study was supported by funding from Gavi - the vaccine alliance, the European Unions Horizon 2020 research and innovation program under grant agreement number 668303 (PERFORM), and a Robert Austrian Research Award.

3
European-derived coronary artery disease polygenic scores over-flag genetic risk in Vietnamese and Southeast Asian populations: a multi-score analysis in 1000 Genomes

Hoang, Q. P.; Le, T. X.; Doan, D. D.

2026-07-15 genetic and genomic medicine 10.64898/2026.07.10.26357796 medRxiv
Top 3%
0.4%
Show abstract

Background. Polygenic scores (PRS) for coronary artery disease (CAD) are derived almost entirely from European-ancestry data. Their portability to Southeast Asian populations, including the Vietnamese, is largely uncharacterised and clinically consequential when scores are used with risk thresholds. Methods. We evaluated four independent European-derived CAD scores from the PGS Catalog (PGS000058, PGS000349, PGS002809, PGS004198; 70 - 5,723 variants) in 2,504 individuals from the 1000 Genomes Project, focusing on the Vietnamese Kinh (KHV) and Dai (CDX) samples. Per-individual scores were computed with PLINK2 and standardised. We assessed (i) the cross-ancestry distribution (calibration) and (ii) a clinically-relevant consequence: the proportion of each population flagged high genetic risk when the European top-20% threshold is applied (20% if perfectly calibrated). Results. For the primary score (PGS000058) the standardised PRS differed across super-populations (ANOVA F(4, 2499) = 121.1, p < 0.001); the Vietnamese Kinh mean was +0.47 SD above the European mean (Welch t = 7.77, p = 2.0 x 10^ -14). Applying the European top-20% high-risk threshold, the fraction of Vietnamese Kinh flagged ranged from 22.2% to 57.6% across the four scores, and of Dai from 21.5% to 43.0%, versus the intended 20%. Three of the four scores over-flagged Vietnamese (25-58%); the largest score (PGS004198) was approximately calibrated for East/Southeast Asians ([~]22%) but markedly over-flagged Africans (69.3%). Conclusions. European-derived CAD polygenic scores are inconsistently calibrated in Vietnamese and other Southeast Asian samples, and most substantially over-flag high genetic risk when a European threshold is applied. The magnitude and even the direction of miscalibration depend on the specific score, so no such score can be assumed transferable without local validation and recalibration. Distribution shift bounds, but does not by itself quantify, loss of predictive accuracy, which requires phenotyped data.

4
Single-cell gene programs define subtype identity and metastatic trajectories in renal cell carcinoma

Madrigal, A.; Kim, M.; Mehrjoo, Z.; Nishimura, T.; Saatci, O.; Osakwe, A.; Zavacky, E.; Moslemi, E.; Glennon, K. I.; Dankner, M.; Maritan, S. M.; Kuasne, H.; Pilon, V.; Monast, A.; Soytas, M.; Arseneault, M.; Oikonomopoulos, S.; Harutyunyan, A.; Lu, T.; Rayes, R.; Soto, L. M.; Hernandez-Corchado, A.; Spicer, J. D.; Petrecca, K.; Siegel, P.; Park, M.; Ragoussis, J.; Sahin, O.; Brimo, F.; Tanguay, S.; Riazalhosseini, Y.; Najafabadi, H. S.

2026-07-16 genetic and genomic medicine 10.64898/2026.07.14.26357682 medRxiv
Top 5%
0.3%
Show abstract

While extensive cellular heterogeneity in renal cell carcinomas (RCC) is linked to diverse clinical outcomes, our understanding of this diversity is limited to those driven by clonal patterns or activity of canonical pathways. Here, we present a compendium of over 85,000 single-cell gene expression profiles from primary and metastatic tumors as well as patient-derived models across four RCC subtypes, including the rare clear cell papillary renal cell tumors, which we show are often misclassified and for which we identify CASP14 as a highly sensitive and specific biomarker. We dissect malignant cell variation within and across tumors using a generative modeling framework that accounts for clonal and copy number-driven expression shifts, defining 59 gene expression programs that deconstruct canonical pathways into functional submodules with divergent activity patterns, distinct regulators, and differential association with clinical outcomes. Despite the canonical view that VHL-deficient clear cell RCC exists in a constitutive pseudohypoxic state, we show strong intra-tumor variability of a hypoxia inducible factor 2 (HIF2)-driven program linked to poor outcome. We also identify early, spatially organized activation of a complete epithelial-to-mesenchymal transition (EMT) program, loss of epithelial identity, and upregulation of protein translation programs as key characteristics of metastatic progression. Finally, a metastatic signature capturing cellular de-differentiation and translational activity identifies primary tumors associated with adverse clinical outcomes. Together, this resource establishes a framework for dissecting malignant cell heterogeneity, refines RCC subtype classification, and defines transcriptional programs underlying metastasis progression.

5
PRANA: A Deep Learning Method for Adapting Polygenic Risk Scores to Diverse Ethnic Groups

Levi, H.; The Breast Cancer Association Consortium, ; Michailidou, K.; Elkon, R.; Shamir, R.

2026-07-15 genetic and genomic medicine 10.64898/2026.07.12.26357860 medRxiv
Top 5%
0.3%
Show abstract

Polygenic risk scores (PRSs), which quantify inherited susceptibility to complex traits and diseases, have emerged as valuable tools for risk stratification and precision medicine. Despite their promise, PRS developed on European cohorts often demonstrate substantially reduced predictive accuracy in non-European populations, due to differences in genetic architecture. The disproportionate representation of European ancestry cohorts in genome-wide association studies (GWAS) leads to inequitable deployment of PRS technologies across diverse populations. Here, we introduce PRANA (Polygenic Risk Adaptation via Neural-network Architecture), a deep learning framework that adapts an existing PRS developed on one population to other ancestries. Unlike methods that require large-scale GWAS in the target population, PRANA leverages pre-trained PRS models derived from European cohorts and adapts them using modestly sized cohorts from the target population. We evaluated PRANA on seven complex traits in South Asian, East Asian and Ashkenazi Jewish populations, as well as in selected smaller East Asian subpopulations where the scarcity of training data poses a particular challenge. PRANA mostly improved predictive performance of the baseline PRS models by 5%-20% in terms of effect size and Nagelkerke's R^2, and, in most cases, outperformed existing cross-ancestry multi-PRS approaches. These results highlight PRANA as a scalable and practical strategy to reduce disparities in genomic risk prediction and advance the equitable application of PRS in diverse populations.

6
PARIS (Pneumonia: Acute Respiratory Infection +/- Sepsis): a prospective single-centre observational cohort study of hospitalised patients with pneumonia

Nasser, S. T.; Piercy, C. R.; Falinska, A.; O'Sullivan, D. M.; Devonshire, A.; Martinez-Estrada, F.; Huggett, J.; Creagh-Brown, B. C.

2026-07-17 respiratory medicine 10.64898/2026.07.15.26357955 medRxiv
Top 6%
0.2%
Show abstract

Introduction Hospitalised community-acquired pneumonia (CAP) is heterogeneous in aetiology, severity, and outcome. Phenotyping and endotyping approaches offer potential to stratify patients biologically and guide targeted therapy, but require well-characterised cohorts with linked biosamples. We describe the PARIS (Pneumonia: Acute Respiratory Infection +/- Sepsis) study: a prospective observational cohort of hospitalised patients with pneumonia, designed to characterise functional outcomes and to provide a biobank for translational immunological research. Methods Adults admitted with CAP to a single NHS district general hospital were enrolled within 24 hours of admission between December 2020 and March 2022. Clinical, functional, and physiological data were collected at enrolment, hospital discharge, and 6-8 week follow-up. Serial blood samples were collected for flow cytometry, transcriptomics, pathogen DNA detection, and plasma biobanking. Results Forty-seven patients were enrolled (15 without and 32 with sepsis [SOFA >=2] at enrolment); 87% met sepsis criteria by 24 hours post enrolment. Most patients (30/47, 64%) were managed as COVID-19, microbiologically confirmed in 27. Mean age was 57 years (SD 16), 70% were male, and baseline comorbidity burden was low. Severity was moderate (median NEWS2 4 at enrolment, rising to 6 by 24 hours post enrolment; p<0.001). Mortality was 4/47 (8.5%), with 44/47 (94%) alive at hospital discharge. Median length of stay was 8 days (IQR 5.5-11). Translational samples were collected from the majority: fresh flow cytometry (44/47, 94%), transcriptomics from the sepsis subgroup (31/32, 97%), pathogen DNA sampling (35 samples received across study timepoints; see Table 5), and stored plasma (29/47, 62%). The primary outcome of functional decline (Barthel score decrease >=1.85) occurred in only 1/29 patients with paired assessments (3.4%). Persistent CRP elevation (>3 mg/L) at 6-8 week follow-up was present in 16/31 (52%) survivors with available data. Conclusions The PARIS cohort provides a well-characterised clinical platform and linked biobank to support translational studies of pneumonia and sepsis. The low rate of functional decline reflects the younger, lower-comorbidity, COVID-predominant population recruited. Primary protocol endpoints were not achieved owing to pandemic-related disruption. Data and samples underpin a programme of linked translational studies.

7
Evaluation of polygenic risk scores and ambient air pollutants for lung cancer risk stratification in a lung cancer screening cohort

Trap, L.; Buyukcelik, R.; Antonissen, N.; Sidorenkov, G. A.; Ruiter, R.; Van Heemst, J.; Sedaghati-Khayat, B.; Stikker, B. S.; Dumoulin, D. W.; Gietema, H. A.; Heuvelmans, M. A.; Mohamed Hoesein, F. A. A.; De Jong, P. A.; Uitterlinden, A. G.; Brusselle, G.; Jacobs, C.; Aerts, J. G. J. V.; Vermeulen, R. C. H.; De Bock, G. H.; Groen, H. J. M.; Vliegenthart, R.; Downward, G. S.; Stadhouders, R.; Van Rooij, J.; NELSON-POP consortium,

2026-07-16 respiratory medicine 10.64898/2026.07.14.26358054 medRxiv
Top 6%
0.2%
Show abstract

Background: Randomized controlled trials have shown that computed tomographic (CT) screening reduces lung cancer mortality. Improved identification of at-risk groups, by leveraging non-smoking risk factors, could help refine screening selection. Aim: To evaluate polygenic risk scores (PRSs) and ambient air pollution (AAP) exposure for risk stratification in the NELSON lung cancer screening cohort. Methods: Two PRSs (PRS-McKay/PRS-Byun) and several AAPs (including nitrogen dioxide, ozone, and particulate matter [PM]) were assessed in the NELSON lung cancer screening trial (N=7,364). PRSs were validated in the Rotterdam Study (N=11,493). Associations with lung cancer, mortality, screening results, and discriminative ability to distinguish lung cancer were evaluated. Results: PRS-McKay and PRS-Byun were associated with lung cancer (odds ratio [OR] per SD [95%CI]: 1.22 [1.08-1.37] and 1.28 [1.13-1.44], respectively) and lung cancer-specific mortality (OR [95%CI]: 1.24 [1.05-1.47], for both), but not with non-lung cancer mortality (OR [95%CI]: 1.01 [0.94-1.10] and 1.03 [0.95-1.12], respectively). Exposure to PM2.5 was associated with lung cancer (OR [95%CI]: 1.11 [1.01-1.22]). PM constituents were associated with adenocarcinoma, particularly PM10 (OR [95%CI]: 1.16 [1.01-1.32]) and ultra-fine particles (OR [95%CI]: 1.16 [1.04-1.30]). PRS and AAP added modestly to the discriminative ability for lung cancer on top of pack-years, age, and sex (area under the curve [95%CI]: 0.659 [0.624-0.695] vs. 0.643 [0.608-0.679]). Conclusions: PRSs and exposure to PM were associated with lung cancer in a high-risk screening population. The primary potential of PRSs may reside in refining lung cancer screening selection toward individuals at higher risk of dying from lung cancer specifically.

8
Sex-different phenotypic correlations: Due to genes or environment?

Fritz, A.; Darrous, L.; Bonnelykke, K.; Pedersen, A. G.; Kutalik, Z.

2026-07-15 genetic and genomic medicine 10.64898/2026.07.13.26357694 medRxiv
Top 6%
0.2%
Show abstract

Differences in physical features and disease prevalence between men and women are examples of sexual dimorphisms. However, sex differences can manifest not only in trait means but also in how strongly risk factors are linked to diseases (e. g. BMI to cardiovascular disease), a question heavily under-researched. To fill this gap, we set out to identify sex differences in phenotype correlations (rP) and decompose them into genetic (rG) and environmental (rE) contributions. Our analysis revealed 250 trait pairs with significant sex-different phenotypic correlations in the UK Biobank. Overall, we observed a predominance of environmental contributions to sex-different effects: 182 trait pairs (73%) exhibited exclusively sex-different rE, while 68 (27%) showed sex differences in both rE and rG, and no trait pair was affected solely by sex-specific rG. For example, we detected sex-different environmental correlation between C-reactive protein and BMI (rE(men) = 0.07 vs rE(women) = 0.25), but no sex-difference in genetic correlation. On the contrary, glycated haemoglobin and LDL cholesterol showed genetic correlation only in women (rG(women) = 0.17; 95% CI = [0.1, 0.23]), but environmental correlation only in men (rE(men) = -0.18; 95% CI = [-0.19, -0.16]). Some of the observed sex differences - including those involving testosterone, SHBG, urate, waist-hip ratio, and triglycerides - may reflect underlying sex-specific genetic architectures, as evidenced by low between-sex genetic correlations. In conclusion, environmental factors are the predominant contributors to sex differences in phenotypic correlations between complex traits, with modest detectable contributions from sex-specific genetic architectures. Recognising these patterns can inform the development of more effective, sex-informed interventions.

9
Multi-tissue analyses of allele-specific chromatin accessibility nominate likely functional variants for type 2 diabetes

Narisu, N.; Li, H. X.; Rathbun, C. J. M.; Varshney, A.; Swift, A. J.; Yan, T.; Sinha, N.; Currin, K. W.; Xue, D.; Robertson, C. C.; Taylor, D. L.; Taylor, H. J.; Beck, A.; Lee, B. N.; Wang, L.; Broadaway, K. A.; Wilson, E. P.; Stringham, H.; Saramies, J.; Lakka, T. A.; Spracklen, C. N.; Scott, L. J.; Stitzel, M. L.; Tuomilehto, J.; Laakso, M.; Koistinen, H. A.; Boehnke, M.; Arda, H. E.; Chen, S.; Biesecker, L. G.; Bonnycastle, L. L.; Erdos, M. R.; Mohlke, K. L.; Parker, S. C. J.; Collins, F. S.

2026-07-15 health informatics 10.64898/2026.07.14.26358094 medRxiv
Top 7%
0.1%
Show abstract

Genome-wide association studies (GWAS) have identified >1,200 signals associated with type 2 diabetes (T2D), yet identifying functional variants remains challenging because the majority of them lie in noncoding regions of the genome and are in areas of high linkage disequilibrium (LD). While chromatin accessibility QTL (caQTL) and expression QTL (eQTL) analyses are useful for nominating regulatory mechanisms underlying GWAS signals, limitations still exist in pinpointing functional variants within regions of high LD. A complementary approach that has been less frequently applied is to focus on the allele-specific effect on chromatin accessibility at heterozygous single-nucleotide polymorphisms (SNPs), hereafter referred to as allelic imbalance. We analyzed the allelic imbalance of reads generated from an assay for transposase-accessible chromatin with sequencing (ATAC-seq) across genotyped samples from 490 donors in T2D-relevant tissues: skeletal muscle, liver, pancreatic islets, adipose tissue, and relevant cell types. We identified 119,949 allelically imbalanced SNPs (FDR<0.05) across the genome. The allelic imbalance was often most prominent in one tissue and showed an enrichment overlapping with tissue-specific transcription factor (TF) binding footprints. Focusing on the 8,581 SNPs in previously published 99% credible sets from 338 T2D GWAS signals, we identified 256 imbalanced SNPs across 123 (36.4% of) signals, each showing allelic imbalance in at least one tissue or cell type. Of these, 71 signals contained only a single imbalanced SNP, representing excellent candidate causative variants. As a proof-of-concept, we showed that 23 of the 256 imbalanced SNPs were supported by allelic assays from previous studies. Further, we experimentally validated two imbalanced SNPs as likely functional variants: rs34584161 among a seven-SNP T2D credible set at the RNF6 signal in islets and rs849134 among a 13-SNP credible set at the JAZF1 signal in liver. This study demonstrates the power of integrating ATAC-seq allelic imbalance (ASAI) with GWAS statistical fine-mapping to identify candidate functional regulatory variants from among tightly linked GWAS variants in disease-relevant tissues. While applied here in T2D, this approach represents a widely applicable high-throughput framework for refining the genetic architecture of complex traits.

10
Evidence for a Nod-like signalling system in cyanobacterial symbiosis with O. sativa

Sanchez del Solar, C.; Jimenez-Rios, L.; Jurado-Flores, A.; Frias, J. E.; Mariscal, V.; Alvarez, C.

2026-07-15 microbiology 10.64898/2026.07.13.738138 medRxiv
Top 9%
0.1%
Show abstract

Symbiotic interactions between plants and nitrogen-fixing microorganisms are essential for sustainable agriculture, yet the molecular mechanisms underlying plant-cyanobacterium symbiosis remain poorly understood. In particular, the nature of the signalling mechanisms mediating partner recognition in associations involving Nostoc species is largely unknown. Recent proteomic analyses have identified proteins homologous to rhizobial Nod factors biosynthetic enzymes in Nostoc punctiforme, suggesting the existence of a Nod-like signalling system. However, the functional role of these components has not been experimentally validated. Here, we investigate the contribution of nod-like biosynthetic and regulatory genes to symbiosis by analysing mutants of N. punctiforme affected in genes with homology to nodB and nodD. Phenotypic characterization revealed that disruption of nodB-like genes does not impair free-living growth but affects early stages of plant association and colonization. Specifically, the nodB1 mutant is impaired in plant association and shows a mild defect in colonization, whereas the nodB3 mutant exhibits a severe defect in colonization. In contrast, nodD-like mutants exhibited altered symbiotic phenotypes, with specific regulators differentially affecting interaction and colonization efficiency in rice (Oryza sativa). In particular, mutation of nodD2 and nodD3 reduced plant association and severely compromised colonization in Oryza sativa, with a more pronounced phenotype in nodD3 mutant. Altogether, our results provide genetic evidence supporting the involvement of Nod-like components in cyanobacterial symbiosis and suggest the existence of a regulatory and biosynthetic module contributing to plant colonization. These findings shed new light on the evolution and diversity of symbiotic signalling mechanisms across plant-microbe interactions.

11
Computational design of a multi-epitope vaccine against M. tuberculosis

Buhari, A.; Okutu, P.; Oyeleke, U. A.; Sivakumar, A.; Hameed, S. A.

2026-07-15 bioinformatics 10.64898/2026.07.09.737463 medRxiv
Top 9%
0.1%
Show abstract

BackgroundTuberculosis remains a leading global infectious killer, with BCG offering inconsistent adult protection and rising drug-resistant strains demanding novel vaccine strategies. We report the first multi-epitope vaccine construct simultaneously targeting three previously unexplored Mycobacterium tuberculosis virulence proteins; EccB3, MycP, and polyketide synthase which collectively govern nutrient acquisition, ESX secretion integrity, and innate immune evasion. MethodsUsing a reverse vaccinology pipeline, B-cell, CTL, and HTL epitopes were predicted, filtered for allergenicity, toxicity, and IFN-{gamma} induction, then assembled into an 823-residue chimeric construct incorporating beta-defensin and PADRE adjuvants with AAY/GPGPG linkers, covering [~]90% global HLA diversity. The construct underwent AlphaFold structure prediction, 3DRefine refinement, disulfide engineering, PROCHECK/ProSA validation, ClusPro 2.0 docking against TLR1/TLR2, and C-IMMSIM immune simulation. ResultsThe construct (82.3 kDa, instability index 32.48) showed strong structural quality (94.7% favoured Ramachandran residues), stable TLR1/TLR2 binding (weighted energy: -1,371.0 kcal/mol), and robust in silico immune responses and durable memory cell formation following booster simulation. ConclusionThis computationally validated construct represents a promising multi-target TB vaccine candidate warranting experimental advancement.

12
Municipal wastewater surveillance reveals socioeconomic and immigration gradients in antimicrobial resistance across Alberta, Canada

Lee, J.; Gonzalez, C.; Au, E.; Acosta, N.; Waddell, B. J.; Xu, Z. S.; Clark, R. G.; Weyant, R. B.; Dalton, B.; Zaheer, R.; McAllister, T. A.; Barkema, H.; Nobrega, D.; Bhatnagar, S.; Lee, B. E.; Pang, X.; O'Grady, C.; Frankowski, K.; Bertazzon, S.; Conly, J. M.; Hubert, C. R. J.; Parkins, M. D.

2026-07-21 infectious diseases 10.64898/2026.07.19.26358431 medRxiv
Top 10%
0.1%
Show abstract

Antimicrobial resistance (AMR) is an ever-increasing threat to population health. Industrial, environmental and societal factors are increasingly recognized as important contributors to AMR within communities. Here, we investigated the spatial distribution of AMR genes (ARGs) across Alberta, Canada and their association with socio-economic, immigration-related, and agro-industrial characteristics using municipal wastewater-based surveillance. We analyzed monthly wastewater metagenomes collected between March 2022 and March 2023 across eleven municipalities, representing 39% of Alberta's population. Integration with census data enabled multivariate analysis, revealing that municipal resistome profiles were strongly structured along income and immigration-related population gradients. ARGs spanning 14 resistance classes exhibited distinct distributional patterns across income and immigration gradients, including contrasting associations among beta-lactam, aminoglycoside, and macrolide-lincosamide-streptogramin ARGs, consistent with heterogeneous selection pressures across sub-populations. These findings demonstrate the capacity of longitudinal wastewater surveillance to identify persistent population-level resistome patterns and highlight the importance of incorporating sociodemographic context into AMR surveillance and mitigation strategies.

13
CuGen: A GPU-accelerated framework for large-scale genomics

Kiiskinen, T.; Richland, J.; Wang, W.; Lu, W. S.; Balasubramanian, N.; Hastie, T.; Tibshirani, R.; Rivas, M. A.

2026-07-17 genetic and genomic medicine 10.64898/2026.07.15.26358178 medRxiv
Top 10%
0.1%
Show abstract

Biobank-scale genomic analyses remain computationally expensive, CPU-bound workflows, particularly when adjusting for confounding. Here, we present CuGen, a GPU-accelerated framework for large-scale genomics. CuGen uses UltraLasso, a novel hierarchical application of univariate-guided sparse regression (uniLasso), to select a compact, phenotype-informed active set of fewer than 30,000 variants. This achieves robust leave-one-chromosome-out (LOCO) confounding control, enabling both downstream GWAS and in-sample fine-mapping. Additionally, we introduce the .cugen file format, a genotype representation designed for memory-optimized, high-throughput streaming and random access on GPU hardware. Building on this substrate, we provide a general GPU-accelerated genomics toolkit handling polygenic prediction, data manipulation, quality control, analysis, and visualization. We demonstrate CuGen's efficacy in the UK Biobank with up to 408,624 individuals, where the full GWAS pipeline and fine-mapping against 6.8 million imputed variants completes in approximately 10 minutes on a single high-throughput GPU with 80 GB of memory. The pipeline scales efficiently to massive phenome-wide analyses with sublinear resource consumption.

14
Natural genetic variation reveals divergent transcriptomic responses to hyperoxia in two Chlamydomonas reinhardtii ecotypes

Temple, J. A.; Neofotis, P. G.; Lucker, B. F.; Bibik, J. D.; Kramer, D. M.; Strenkert, D.

2026-07-15 genomics 10.64898/2026.07.09.737578 medRxiv
Top 10%
0.1%
Show abstract

Green algae must continuously balance resource availability to maintain photosynthetic performance. The O2:CO2 ratio is a key determinant of their metabolic mode. Under hyperoxia or low CO2, many algae induce a carbon concentrating mechanism (CCM). In the model green alga Chlamydomonas reinhardtii, the CCM relies on a pyrenoid, a specialized microcompartment that elevates CO2 around rubisco. While ambient CO2 acclimation is well-studied, responses to hyperoxia remain poorly understood, despite its frequent occurrence in nature under high light. Using controlled bioreactors, we exposed two diverse Chlamydomonas ecotypes, CC1009 and CC2343, to 95% oxygen to analyze time-dependent, genome-wide transcriptomic and phenotypic changes. Both ecotypes induced CCM genes, but they exhibited distinct molecular and physiological phenotypes. The tolerant ecotype (CC1009) successfully adapted, developing a functional CCM with a structured starch sheath. Conversely, the sensitive ecotype (CC2343) suffered growth arrest and formed malformed pyrenoids. Transcriptomics revealed that CC1009 initiated a rapid initial response, upregulating chloroplast proteostasis and downregulating nucleotide metabolism. CC2343 showed a massive, delayed transcriptional response, downregulating genes coding for photosystems and tetrapyrrole biosynthesis. This unbiased transcriptomic approach identifies key candidate genes driving algal acclimation to hyperoxic stress in natural, high-light environments.

15
Storing >1 byte of information in 16S ribosomal RNA using orthogonal trans-splicing ribozymes

Dysart, M. J.; Fang, L.; Karinje, L. K.; Chappell, J.; Stadler, L. B.; Silberg, J. J.

2026-07-15 synthetic biology 10.64898/2026.07.14.738544 medRxiv
Top 10%
0.1%
Show abstract

TEXT ABSTRACTCatalytic-RNA (cat-RNA) expressed from mobile DNA can record cellular events, such as the uptake of plasmids via horizontal gene transfer, by splicing a barcode onto 16S ribosomal RNA (rRNA) - a system termed RNA addressable modification (RAM). However, scaling RAM to record multiple simultaneous biological events requires large numbers of orthogonal cat-RNA whose signals reflect the biological features under investigation rather than variability arising from the barcode sequence. Here, we explore how to design orthogonal cat-RNA to record information about multiple plasmid-encoded traits in parallel. We show that cat-RNA having tRNA-derived barcodes with sequence variation in the anticodon stem-loop present greater signal consistency within Escherichia coli than mRNA-derived barcodes. When orthogonal cat-RNA designs harboring tRNA-derived barcodes were evaluated in Vibrio natriegens and Pseudomonas putida, increased variance was observed compared with Escherichia coli. Nevertheless, the signal consistency was sufficient to use these orthogonal cat-RNAs to report on the relative activities of four promoters and two origins of replication by sequencing barcoded-rRNA derived from the three organisms. These results show how RAM can be multiplexed to report on mobile DNA features in microbial communities and illustrate the importance of accounting for variability in RNA outputs when designing and interpreting multiplexed RNA barcoding data. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=88 SRC="FIGDIR/small/738544v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@406ebaorg.highwire.dtl.DTLVardef@259751org.highwire.dtl.DTLVardef@1f1512corg.highwire.dtl.DTLVardef@8384b_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
Association between serum CEA levels and ctDNA-detected Epidermal Growth Factor Receptor mutations in lung adenocarcinoma

Roy, S.; Soroar, M. K. I.; Ara, H.; Nur, S. A.; Akanda, R. A.; Saha, S.; Alam, M. M.

2026-07-17 oncology 10.64898/2026.07.14.26358115 medRxiv
Top 11%
0.1%
Show abstract

Background with objective: Detecting EGFR mutations is critical for treating lung adenocarcinoma with highly effective targeted therapies. However, standard genetic testing is expensive, complex, and often unavailable in resource-limited settings like Bangladesh. Because elevated serum CEA has been linked to these genetic alterations, it could serve as an accessible screening tool. This study aims to evaluate the association between serum CEA levels and EGFR mutation status to determine if routine CEA testing can reliably predict these mutations and guide treatment. Methodology: In this cross-sectional analytical study, we recruited 58 patients with histologically confirmed treatment naive lung adenocarcinoma. The presence of EGFR mutations in the ctDNA was determined via ARMS (Amplification Refractory Mutation System) PCR. Patient data was statistically analyzed to assess the diagnostic correlation between serum CEA levels and the presence of EGFR mutations. Result: The overall EGFR mutation rate was 43.1% with exon 19 deletion (48%) and exon 21 mutations (44%) were the predominant types. Median serum CEA levels were significantly higher in patients with EGFR mutations compared to wild-type cases (14.6 ng/ml vs 2.8 ng/ml, p<0.001). A multivariate analysis revealed a 14% increased likelihood of an EGFR mutation for 1 ng/ml rise in serum CEA. Furthermore, serum CEA showed strong diagnostic accuracy for ctDNA samples at a 6.39 ng/ml cut-off (AUC 0.82, sensitivity 68.0%, specificity 84.8%). Conclusion: Serum CEA is a valuable, cost-effective, and non-invasive biomarker demonstrating significantly higher levels and strong diagnostic accuracy in EGFR-mutated lung adenocarcinoma compared to wild-type cases.

17
Prevalence and factors associated with multidrug resistant Mycobacterium tuberculosis infection in Cameroon: a systematic review and meta-analysis

Cheuyem, F. Z. L.; Achangwa, C.; Mbarga, P. E.; Tchamani, R.; Dabou, S.; Mutarambirwa, H. D.; Temgoua, M. N.

2026-07-15 infectious diseases 10.64898/2026.07.13.26357969 medRxiv
Top 11%
0.1%
Show abstract

Background: Multidrug-resistant tuberculosis (MDR-TB) remains a significant public health threat in low- and middle-income countries, including Cameroon. This systematic review and meta-analysis aimed to determine the pooled prevalence of MDR-TB and other specific anti-tuberculosis drug resistance patterns, as well as to identify factors associated with drug-resistant tuberculosis in Cameroon. Methods: A comprehensive literature search was conducted in PubMed, Scopus, Web of Science, Embase, Cochrane Library, and African Journals Online. Additional studies were identified through Google Scholar and reference list screening. Observational studies (cross-sectional, cohort, and case-control) reporting drug resistance among bacteriologically confirmed tuberculosis patients in Cameroon were eligible. Joanna Briggs Institute critical appraisal tools were used to critically assessed the study quality. Pooled prevalence estimates were calculated using random-effects meta-analysis. Subgroup analyses and meta-regression explored sources of heterogeneity. A p-value 0.05 was considered statistically significant. Results: Twenty-eight studies conducted between 1995 and 2022 were included. The pooled prevalence of MDR-TB was 5.2% (95% CI: 2.7-9.6; 21 studies; n = 7,515), with significantly higher acquired resistance (11.6%; 95% CI: 6.3-20.3) than initial resistance (2.0%; 95% CI: 1.1-3.5). The highest pooled MDR-TB prevalence was observed in the most recent studies (38.8%; 95% CI: 33.7-44.2), and the lowest in 2015-2019 (2.7%; 95% CI: 0.4-15.2). Any resistance to anti-tuberculosis drugs was 16.0% (95% CI: 10.3-23.9; 28 studies; n = 9,931), and rifampicin resistance was 4.6% (95% CI: 2.4-8.6; 25 studies; n = 8,728). Monoresistance was highest for streptomycin (6.4%; 95% CI: 3.7-10.8) and isoniazid (4.7%; 95% CI: 3.0-7.4). Previous tuberculosis infection was the strongest predictor of drug resistance (OR = 3.9; 95% CI: 1.8-8.4), followed by alcohol consumption (OR = 1.8; 95% CI: 1.2-2.7) and history of incarceration (OR = 1.7; 95% CI: 1.1-2.6). High heterogeneity was observed across most of the pooled estimates. Conclusions: Drug-resistant tuberculosis, particularly MDR-TB, poses a substantial burden in Cameroon, with acquired resistance significantly exceeding initial resistance. Previous tuberculosis infection, alcohol use, and incarceration are key modifiable risk factors. These findings underscore the urgent need to strengthen routine drug susceptibility testing, scale up rapid molecular diagnostics, enhance treatment adherence strategies, and implement targeted interventions for high-risk populations.

18
Prescribing Trends of Antimicrobials in Obstetric and Gynaecological Inpatients: A Prospective Drug Utilization Study with Concurrent Antimicrobial Stewardship Audit from a Tertiary Care Hospital in Karachi, Pakistan

Ansari, T.; Zehra, A.; Jabbar, S.; Fatima, M.; Syed, B.; Shah, S. S. A. M.; Ahmed, A. S.; Hamid, A.; Ashafaq, H.

2026-07-17 obstetrics and gynecology 10.64898/2026.07.16.26358229 medRxiv
Top 11%
0.1%
Show abstract

Background: Antimicrobial resistance (AMR) disproportionately affects low- and middle-income countries (LMICs) such as Pakistan, where obstetric and gynaecological (OBGYN) patients carry high antibiotic exposure. Specialty-specific drug utilization data with concurrent stewardship audit remain scarce. This study evaluated antibiotic prescribing patterns, consumption metrics, and antimicrobial stewardship program (AMS) compliance in OBGYN inpatients at a public sector tertiary care hospital. Methods: A prospective cross-sectional study was conducted in OBGYN wards of Dow University Hospital, Karachi, from 1 September to 31 October 2025. Women receiving [&ge;]1 systemic antibiotic were included. Daily AMS rounds were conducted by an Infectious Diseases physician and pharmacist. Antibiotic consumption was measured as Defined Daily Doses (DDD) and Days of Therapy (DOT) per 1,000 patient-days (total = 821). Antibiotics were classified by WHO AWaRe (2023) framework. Results: Of 812 total admissions, 278 patients (34.2%) received [&ge;]1 antibiotic and were enrolled (205 obstetric, 73 gynaecological), generating 636 prescriptions (mean 2.29/patient). Surgical prophylaxis was the predominant documented indication (213, 33.5%); 65.1% carried no documented indication. By AWaRe classification, 53.6% were Access-group and 46.1% Watch-group. Ceftriaxone (38.4%) and metronidazole (36.8%) together represented 75.2% of prescriptions. Combined DDD/1,000 patient-days was 1,758.6 and DOT/1,000 patient-days was 1,852.7. AMS compliance was 0%. Conclusions: This study documents high antibiotic prescribing burden, near-universal documentation failure, and zero AMS compliance in OBGYN inpatients at a Pakistani public sector hospital. The predominance of Watch-group antibiotics and undocumented surgical prophylaxis highlights structural stewardship gaps. Findings support urgent need for institutional OBGYN antibiotic guidelines and structured pharmacist-led AMS programs.

19
Knowledge and misconceptions of the French population regarding medical genetics: a survey of 3,000 respondents

MERCIER, S.; PETIT, F.; MISRAHI, M.; BERTA, P.; CAMBON-THOMSEN, A.; CHAUMETTE, B.; CHNEIWEISS, H.; CRETOLLE, C.; EDERY, P.; HEARD, D.; KONYUKH, M.; LAENG, C.; MAHLAOUI, N.; PASQUIER, L.; PLUTINO, M.; ODENT, S.; STOPPA-LYONNET, D.; "Genetics and the General Public" FFGH Ethics Working Group,

2026-07-19 genetic and genomic medicine 10.64898/2026.07.17.26358259 medRxiv
Top 13%
0.1%
Show abstract

Advances in high-throughput sequencing and genetic research have expanded the role of genetics in medicine and society. Population-based screening programs, including neonatal and preconception testing, are increasingly implemented globally, alongside the rise of direct-to-consumer (DTC) genetic testing. The "Genetics and the General Public" Ethics Working Group of the French Federation of Human Genetics (FFGH) assessed knowledge and awareness of genetics within the French population through a nationally representative survey (n=3,013) conducted by the polling firm Ipsos bva. Results indicated that 69% of respondents report an interest in genetics, although their level of knowledge remains limited. Most respondents expressed positive attitudes toward genetics, perceiving it as a major source of hope in healthcare. While a majority indicated willingness to undergo genetic testing for medical purposes, they also reported legitimate concerns regarding the potential results. Despite legal restrictions, 12% reported having ordered a DTC genetic test (5% for genealogical; 5% for medical and 2% for both purposes), and 45% of non-users expressed strong interest in this type of test. Notably, there is a substantial lack of awareness regarding the limitations of these tests and the French legal framework governing their use. These findings highlight critical gaps in public knowledge, emphasizing the need for improved genetic education, including incorporating genetics into school curricula and launching targeted awareness campaigns. These initiatives should help clarify the distinctions between clinically validated genetic tests and DTC genetic testing services, addressing both their benefits and their ethical, legal, and scientific limitations, in order to promote informed decision-making.

20
Comparative Efficacy of Vancomycin and Fidaxomicin Regimens for the Prevention of Recurrent Clostridioides difficile Infection: A Systematic Review and Network Meta-Analysis of Randomized Controlled Trials

Prosty, C.; Butler-Laporte, G.; Brophy, J.; Frenette, C.; Loo, V.; Coburn, B.; Hota, S.; Longtin, Y.; Kong, L.; Muller, M.; Steiner, T.; Valiquette, L.; Daneman, N.; Daley, P.; Nott, C.; MacFadden, D. R.; Kandel, C.; Chen, Y.; Perez- Patrigeon, S.; Lee, T. C.; McDonald, E.

2026-07-17 infectious diseases 10.64898/2026.07.14.26358112 medRxiv
Top 14%
0.0%
Show abstract

Background and Aims The optimal treatment for first episodes and first recurrences of Clostridioides difficile infections (CDI) is unknown and there is emerging evidence for pulse and taper (P-T) regimens. Therefore, we sought to estimate the relative efficacy of treatment options. Methods MEDLINE and CENTRAL were searched from database inception to May 21, 2025 and unpublished conference abstracts were searched from recent infectious disease conferences. RCTs on the treatment of first episodes or first recurrences of CDI comparing fixed-dose or P-T regimens of fidaxomicin or vancomycin were included. The primary and secondary outcomes were 40- and 56-day CDI recurrence, respectively. A random-effects network meta-analysis on the risk ratio (RR) scale was conducted using a standard regimen (10-14 days) of vancomycin as the comparator. Treatments were ranked using the surface under the cumulative ranking curve (SUCRA). Results 8 RCTs were included comprising a total of 2181 patients. For 40-day recurrence, fidaxomicin P-T had the highest probability of ranking best (RR=0.10, 95%Confidence Interval [95%CI]=0.10-0.49, SUCRA=1.00), followed by vancomycin P-T (RR=0.49, 95%CI=0.32-0.76, SUCRA=0.61), fixed-dose fidaxomicin (RR=0.61, 95%CI=0.49-0.76, SUCRA=0.39), and, finally, fixed-dose of vancomycin (SUCRA=0.00). The treatments ranked in the same order for 56-day recurrence, though only 3 RCTs reported on this timepoint. Conclusion Vancomycin P-T, fidaxomicin P-T, and fixed-dose fidaxomicin were all superior to a fixed-dose vancomycin. Head-to-head comparative effectiveness RCTs are needed to quantify their relative effect sizes of and impact on long-term prevention of recurrent CDI.