iMeta
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match iMeta's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Zhang, Z.; Feng, Y.; Ge, X.; Meng, X.; Peng, Y.
Show abstract
Viral phenotypes such as host and tissue tropism are critical determinants of viral infection and transmission. Inferring viral phenotypes presents unique challenges compared to cellular organisms, as viruses rely entirely on host machinery for replication and survival. Current methods for predicting viral phenotypes mainly rely on viral genomic data, often overlooking host-related information. Here, we evaluated the utility of predicted virus-human protein-protein interactions (PPIs) in inferring diverse viral phenotypes using machine-learning algorithms. For predicting human infectivity, a PPI-based machine learning model outperformed both virus genomic and protein sequence-based models that used large language model embeddings. It also surpassed previous methods that incorporated both viral and host genomic data. The human proteins identified by the model were significantly enriched in functions related to viral infection and immune response. In predicting various phenotypes of human RNA viruses, PPI-based models performed better than virus sequence-based models in forecasting virulence, human transmissibility and transmission routes, while showing comparable performance to genomic sequence-based models in predicting tissue tropism. Finally, we demonstrated that a PPI-based model could distinguish high-risk HPV genotypes from low-risk ones. Proteins associated with high-risk HPV were involved in apoptosis and immune regulation, whereas those linked to low-risk HPV were enriched in telomere maintenance and DNA repair. Collectively, this study is the first to demonstrate the value of predicted virus-human PPIs in inferring viral phenotypes, thereby enhancing our understanding of the molecular mechanisms underlying these phenotypes. It also provides effective tools for risk assessment of emerging viruses, contributing to improved pandemic preparedness.
Zierep, P.; Hojat Ansari, M.; Buehler, P.; Faack, S.; Fosso, B.; Defazio, G.; Beracochea, M.; Sanchez, S.; Hottmann, A.; Batut, B.
Show abstract
Advances in whole-genome sequencing (WGS) technologies have enabled large-scale recovery of metagenome-assembled genomes (MAGs), providing unprecedented insights into microbial diversity across diverse environments. However, the reconstruction of MAGs remains computationally demanding and methodologically complex, requiring the integration of multiple tools for quality control, assembly, binning, refinement, and annotation. Existing workflows often rely on scripting-based implementations, constrain user-driven modification and stepwise execution, and require advanced expertise in high-performance computing (HPC) system administration, thereby limiting accessibility, reproducibility, and adaptability. Here, we present FAIRyMAGs, a Findable, Accessible, Interoperable, and Reusable (FAIR)-compliant, modular pipeline implemented within the Galaxy platform for the generation and analysis of MAGs. FAIRyMAGs consists of six interconnected workflows covering all major steps of MAG reconstruction, including read preprocessing, host and contaminant removal, assembly, binning, dereplication, and downstream taxonomic and functional annotation. The workflows are accompanied by extensive training material, including tutorials, a learning pathway, FAQs, test datasets and video walk-throughs by domain experts, supporting community adaptation. By leveraging Galaxys graphical interface and federated infrastructure, FAIRyMAGs enables users to execute complex analyses on public or private compute resources without requiring local installation or workflow programming expertise. The modular design further supports flexible adaptation, iterative optimization, and seamless integration of new tools contributed by the community. To demonstrate applicability, FAIRyMAGs was applied to four real-world microbiome datasets spanning various host-associated and environmental systems. These analyses revealed substantial variability in MAG recovery, community complexity, and clustering structure, underscoring the importance of flexible workflows adaptable to dataset-specific characteristics. Overall, FAIRyMAGs provides an accessible, extensible, and reproducible framework for genome-resolved metagenomics, reducing technical barriers and enabling methodological innovation through community-driven development within the adaptable Galaxy ecosystem.
Mainguy, J.; Lemane, T.; Bazin, A.; Arnoux, J.; Gautreau, G.; Medigue, C.; Calteau, A.; Vallenet, D.
Show abstract
PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=63 SRC="FIGDIR/small/742796v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@186b88dorg.highwire.dtl.DTLVardef@1be33d0org.highwire.dtl.DTLVardef@3bc596org.highwire.dtl.DTLVardef@292401_HPS_FORMAT_FIGEXP M_FIG C_FIG
Andrade, R. L.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.
Show abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearsons r improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTXGUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
Kananen, K.; Tran, N.; Bradley, P. H.
Show abstract
In microbiome studies, associations between microbial functions and the environment are often confounded by phylogeny. While some methods explicitly account for this confounder, they require information about genome content, limiting their use in biomes where few genomes have been available. To make these methods more universally accessible, we have developed Phylogenize2, a redesigned phylogeny-aware tool for linking microbial gene families to abundance phenotypes. Phylogenize2 integrates large metagenome-assembled genome collections, including both biome-specific collections from MGnify and a broadly sampled general purpose database, GlobDB, to substantially expand species coverage, allowing its application in environments like the mouse gut and ocean. In addition, by default, Phylogenize2 uses a new robust phylogenetic testing framework that has been optimized for microbial abundance data, while also allowing the use of other comparative methods such as POMS. In an experimental mouse study, Phylogenize2 identifies that Muribaculaceae with higher abundance on a high-fat diet are enriched for proteins in the thioredoxin family, with likely roles in oxidative stress. When we apply Phylogenize2 to a polar ocean study, we find that a molybdenum-dependent PaoABC/YagTSR-like aldehyde oxidoreductase system differentiates mesopelagic from surface-dwelling Flavobacteriaceae, suggesting that aldehyde detoxification may be important for organisms that degrade marine snow. Together, these results show that Phylogenize2 expands phylogeny-aware microbiome analysis beyond the human gut and can provide insight into the genetic basis of microbiome-encoded traits in diverse environments. ImportanceMicrobiome studies often set out to identify which microbes are more or less abundant across environments, but these patterns can be difficult to interpret. Phylogenize2 is an open-source software package that allows researchers to ask whether individual microbial gene families are associated with the environment across independent branches of the microbial tree of life. By incorporating large collections of genomes from uncultivated microbes, as well as modern statistical methods designed for microbial abundance data, Phylogenize2 makes this approach practical for microbiomes beyond the human gut, including in model organisms like lab mice and free-living environments like the ocean. We also provide a pipeline that allows the use of new genome collections. In two case studies, we demonstrate that Phylogenize2 effectively prioritizes specific genes and pathways from metagenomic data, thereby leading researchers from changes in microbial abundance to more biologically interpretable explanations.
Zhou, Y.; Huang, F.; Zhao, Y.
Show abstract
Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.
Loureiro, C.; Schorn, M. A.; Alanjary, M.; Kuipers, B.; Louwen, J. J. R.; van der Oost, J.; Medema, M. H.; Sipkema, D.
Show abstract
Marine sponges are known sources of bioactive natural products (NPs), many of which are produced by associated bacterial symbionts via encoded biosynthetic gene clusters (BGCs). A particularly interesting subclass of sponge-derived NPs is comprised of small, brominated alkaloids, which are recovered from diverse habitats and host sponge taxonomies. Despite having been described decades ago, most of these NPs do not have an elucidated biosynthetic origin. We queried metagenomes of several sponge species by making use of a minimal set of core enzymes that we postulate to be necessary to produce these small peptidic NPs: an FADH2-dependent halogenase and an AMP-binding adenylation enzyme. This revealed a variety of novel BGC architectures, many of which showed conservation among sponge host phylogenies and were encoded in the genomes of diverse sponge-associated bacteria. Furthermore, we identified a BGC in the sponge G. barretti that is potentially linked to the production of the iconic barettins, given its enzymatic machinery and specific acidobacterial origin. The present work contributes to the challenging quest to link orphan brominated NPs to their parent BGCs in the sponge holobiont and beyond.
Fritscher, J.; Duncan, A.; Hildebrand, F.
Show abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal -- profiling through alignment -- an ultra-fast alignment-based method for species- and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom benchmarks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler). Because protal retains read alignments, it can reconstruct intraspecific phylogenetic relationships among detected bacteria with accuracy comparable to StrainPhlAn 4. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
Paliyal, S.; Kaur, B.; Rao, L.; Chakrabortty, A.; Singh, L.; Sehgal, I.; Sharma, M.; Singh, D.; Chaudhry, V.; Mantri, S. S.
Show abstract
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strains and NPs have been reported harboring this rare bis-heterocyclic moiety, underscoring a largely unexplored chemical space. Here, we performed large-scale genome mining and identified 277 putative biosynthetic gene clusters (BGCs) across diverse bacterial hosts, including previously unreported bacterial genera and strains. The BGCs were grouped into three compound classes: benzoxazolinate, benzobactin, and ashimides based on sequence similarity network clustering. Bioactivity predictions of the identified BGCs revealed the predominance of antibacterial and cytotoxic potential, highlighting promising candidates for future experimental validation and functional studies. This study also presents a neural network-based bioprospecting model that efficiently detects rare BGCs encoding benzoxazolinate-containing molecules from genomic sequences. Overall, our findings expand the known repertoire of bacterial hosts with the potential to produce benzoxazolinate-containing NPs and provide a comprehensive framework for the discovery and identification of candidate BGCs.
Goemann, H.; Jiraska, L.; Perry, M.; Hillary, L. S.; Emerson, J. B.
Show abstract
We present an update to the PIGEON (Phages and Integrated Genomes Encapsidated or Not) reference database of DNA viral sequences (vOTUs, mostly dsDNA bacteriophages) from global ecosystems. To reflect the inclusion of only virus size-fractionated metagenome- (virome-)derived vOTUs, we reintroduce the database as PEONY (Phages Encapsidated ONlY) and present new data summarizing the utility of this database.
Chasapi, I. N.; Aplakidou, E.; Chasapi, M. N.; Lamari, E.; Galaras, A.; Diplari, S.; Iliopoulos, I.; Emiris, I. Z.; Georgakopoulos-Soares, I.; Patalano, S.; Stravopodis, D. J.; Karatzas, E.; Baltoumas, F. A.; Kyrpides, N.; Pavlopoulos, G. A.
Show abstract
Metagenomic studies of arthropod-associated microbiomes have generated vast amounts of sequence data, yet the functional and structural organization of these proteins remains largely unexplored. Here, we present ArthroVerse, the first comprehensive database of protein families derived from arthropod-associated metagenomes. Non-redundant protein families were generated after rigorous filtering, deduplication, and clustering. The protein families were further annotated with microbial taxonomy, host associations, protein structural information, and Carbohydrate-active enzymes (CAZyme) predictions. The resulting dataset integrates both metagenomic and reference genome-derived proteins, enabling systematic exploration of functional diversity, evolutionary relationships, and host-microbe interactions in insect microbiomes. ArthroVerse provides a valuable resource for the study of microbial ecology and arthropod physiology, offering unprecedented insight into the protein landscape of insect-associated microbial communities.
De Santiago, A.; Bik, H.
Show abstract
Microbes closely interact with every living organism, including meiofauna (i.e., microbial eukaryotes 38 m - 1 mm in length), and influence the development, life cycle, and evolution of diverse metazoans. Together, meiofauna and their microbiomes, collectively referred to as the holobiont, underpin biogeochemical cycles and drive decomposition of organic matter. However, our understanding of the ecological and evolutionary dynamics of meiofauna microbiomes are limited, typically owed to low-resolution 16S rRNA surveys, which cannot accurately delineate bacterial taxa. Single-specimen holobiont sequencing can help overcome the limitations of metabarcoding approaches by 1) generating metagenome-assembled genomes (MAGs) of the host microbiome and 2) recovering host single-copy genes (SCGs) to phylogenetically confirm the identity of the host organism. However, most bioinformatics pipelines for the assembly of metagenomic datasets have been developed for the assembly of high-complexity microbial communities of bulk sediment or soil samples (and cannot be used for the assembly of host genomes), rely on co-assembly approaches (which collapses strain-level genomic information of bacterial taxa), and focus on binning either prokaryotic or eukaryotic taxa. Therefore, there is a tremendous need for a computational workflow for the dual analysis of host genomes and their microbiomes. Here, we developed MeioBIOME, a modular Snakemake pipeline for the reproducible analysis of holobiont metagenomes obtained from individually sequenced microbial metazoa. We analyze publicly available single-specimen metagenomics datasets to show the utility of MeioBIOME and recover host-associated symbiont MAGs and host SCGs. Additionally, we integrate state-of-the-art binning algorithms which generate more MAGs than the DOE Joint Genome Institute metagenomic pipeline. We anticipate that MeioBIOME will facilitate studies of phylosymbiosis by generating high-quality host genome skims (to build well-supported host phylogenetic trees) and host-associated prokaryotic MAGs obtained from single specimens.
Zeng, Z.; Wang, Y.
Show abstract
FigTree is a long-standing phylogenetic tree viewer, but its GUI-centered workflow does not itself provide a versioned, batch-replayable record of styling operations. We present FigTreeKit, a Python package that serializes a supported subset of FigTree 1.4.4 annotations (!hilight, !color, and !font), audits taxonomy mappings before topology-gated clade collapse, retains selected BEAST-style metadata in the tested fixtures, and invokes a patched FigTree renderer for headless PNG, PDF, and SVG output. Across 60 independently generated balanced trees with 50-10,000 taxa (10 trees per size, each timed 10 times as technical replicates), the tree-level log-log slope of export time was 0.96 (95% confidence interval [CI], 0.91-1.01), which is compatible with approximately linear scaling over the tested range but does not prove it. The 189,801-taxon GTDB R232 bacterial reference tree was parsed and exported as a large-data scalability demonstration. On the 10,122-taxon GTDB R232 archaeal reference tree, the scripted workflow assessed 179 order-level groups; 142 multi-tip groups produced non-trivial collapses, whereas 37 singleton groups did not alter the display. The software is accompanied by 796 passing tests, a golden conformance corpus that includes acceptance tests against the bundled FigTree JAR, deterministic scenario-based topology checks, and an overall statement coverage of 81%, reported as a descriptive engineering metric. FigTreeKit is released under the GPL-2.0-or-later license as the figtreekit package on PyPI, with source code, documentation, and benchmark data archived on Zenodo.
Dang, T.; Lysenko, A.; Tsunoda, T.
Show abstract
The microbiome plays a significant role in the development and progression of many diseases, yet extracting interpretable insights from multi-omics data remains challenging. Existing approaches face a recurring practical trade-off: deep learning methods achieve high predictive performance but lack uncertainty quantification, whereas probabilistic methods provide interpretable results but require data-type-specific likelihood functions that limit generalization across diverse omics modalities. Here, we introduce DBayesCM (Deep Bayesian Clustering for Multi-omics), which combines deep learning modeling with Bayesian nonparametric methods. DBayesCM employs separate encoders to project microbiome and host omics data into a shared latent space, where an infinite mixture model with a Dirichlet process prior determines the number of clusters automatically while quantifying the uncertainty of each samples assignment. Spike-and-slab priors identify discriminative features, and a Bayesian neural network estimates probabilistic co-occurrence between microbial species and host omics features. To isolate the effect of latent geometry, we evaluate two variants that are identical except for their latent space: DBayesCM-vMF constrains the latent to the unit hypersphere and applies a von Mises-Fisher mixture, while DBayesCM-GMM uses a Euclidean latent space and a Gaussian mixture. On simulated data, the hyperspherical variant recovered the correct number of clusters, whereas the Euclidean variant over-segmented, demonstrating that the latent geometry affects cluster recovery. Applied to colon, breast, and kidney cancer cohorts spanning metagenomics, host metabolomics, RNA-seq, and miRNA data, and to an obstructive sleep apnea model, DBayesCM ranked consistently among the existing methods while uniquely combining data-driven cluster-number determination, sample-level uncertainty, and interpretable feature selection within a single framework. DBayesCM reveals conditional probabilistic co-occurrence between core microbial species and host omics features, enabling uncertainty-aware exploration of microbiome-host relationships across diverse diseases.
Dyball, X.; Ponsero, A. J.; Docherty, J. A. D.; Telatin, A.; Crost, E. H.; Juge, N.; Cook, R.; Adriaenssens, E. M.
Show abstract
Prophages are major drivers of bacterial evolution, mediating horizontal gene transfer and lysogenic conversion to alter host phenotypes. Nevertheless, identifying prophages within bacterial genomes remains challenging due to their heterogeneity and similarity to other mobile genetic elements. Here we present PHORAGER (Prophage Hunting, vOtu Retrieval, Annotation and Genomic ExploRation), a scalable Nextflow pipeline for the standardised identification and quality assessment of prophages from bacterial genomes. PHORAGER incorporates bacterial genome pre-processing, consolidation of predictions from multiple mining tools, annotation-based filtering to reduce false positives, and generation of ready-to-analyse summary tables. We validated PHORAGER using 30,824 publicly available ESKAPE pathogen genomes. PHORAGER recovered more high-quality prophages than individual mining tools alone, and through extensive quality assessments removed a substantial number of false-positive predictions. In total 23,132 putative prophages were identified, the majority belonging to the class Caudoviricetes, and exhibiting a high degree of host-specificity. Putative antimicrobial resistance genes were detected in 0.48% of prophages, whereas virulence factors were most abundant in S. aureus prophages. ESKAPE prophages also frequently encoded anti-phage defence systems. PHORAGER is freely available as open-source software and the ESKAPE prophage collection generated in this study provides a reusable resource for further investigations. GRAPHICAL ABTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/742953v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@1ce582org.highwire.dtl.DTLVardef@11fc004org.highwire.dtl.DTLVardef@17765f5org.highwire.dtl.DTLVardef@1c6f0ff_HPS_FORMAT_FIGEXP M_FIG C_FIG
SHEKARRIZ, E.; VIJENDRAN, E.; Ho, J. W. K.
Show abstract
Viral identification for terabyte-scale metagenomic data is limited by scalability and computational resources. We present vOMIX-MEGA, an end-to-end viral metagenomic framework that overcomes performance bottlenecks by significantly improving parallelization and memory usage in critical steps. Benchmarked on empirical datasets, it completes processing in up to 50 minutes with 24 GB of RAM, bypassing four other state-of-the-art pipelines that require 7 hours (383 GB) to 14 days (32 GB). vOMIX-MEGA is on average 21% and 13% more accurate when benchmarked on mock and experimental data and is available via https://github.com/holab-hku/vOMIX-MEGA.
McLaughlin, R. J.; Chen, S.; Nag, A.; Noonan, A. J. C.; Bartolomeu, C.; Borden, S. A.; Lam, S.; Myers, R.; Hallam, S. J.
Show abstract
Microbial communities inhabiting the respiratory tract contribute to health status through interactions with host physiology, immune function, and local environmental conditions. Advances in small subunit ribosomal RNA (SSU or 16S rRNA) gene amplicon sequencing enable culture-independent profiling of microbial communities as amplicon sequence variants (ASVs), revealing links between microbial dysbiosis and respiratory diseases, and the use of mass spectrometry to measure volatile organic compounds (VOCs) in exhaled breath shows emerging promise for biomarker discovery. Here we present ASPIRE, the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems, an accessible Nextflow workflow for processing, analyzing, and interpreting linked ASV-VOC data from respiratory microbiome studies. ASPIRE is designed to support scalable comparative analysis across respiratory sample types while preserving intermediate file outputs for inspection and reuse within a standardized file structure. Availability and implementationThe ASPIRE workflow, quick start guide, and usage are freely available on the GitHub platform https://github.com/hallamlab/ASPIRE. ASPIRE is implemented in Nextflow with a modular design to allow for reproducibility, extensibility and scalability.
Skudnov, A.; Badamshin, E.; Efimenko, B.; Popadin, K.; Gunbin, K.; Denisov, S.
Show abstract
The mutational spectrum is an increasingly important molecular phenotype that quantitatively describes mutagenesis in a given gene and species, enabling future comparative analyses to reveal differences in underlying mutagenic processes, whether internal, such as DNA repair processes, or external, such as ecological niches and conditions. Mutation accumulation experiments, although time-consuming and costly, remain the standard approach for reconstructing bacterial neutral mutation spectra. Here, we present BacNeMu, a phylogenetically informed pipeline that reconstructs neutral mutational spectra of bacterial genomes using open databases GTDB, AnnoTree and KEGG Orthology, building on previously developed NeMu pipeline. BacNeMu reconstructs mutation spectra that closely match mutation accumulation experiments results while requiring substantially less time, enabling comparative analyses across diverse bacterial taxa. Applied to obligate aerobes and anaerobes, BacNeMu recovered the expected excess of T:A>C:G transitions, consistent with oxidative-damage-associated mutational patterns previously described in mitochondrial genomes and yeast single-strand. We further asked if any other ecologic factors influence a mutational spectrum. As a pilot we compared three species living under different temperatures: one strong thermophile - Thermotoga maritima, one psychrophile - Clostridium algidicarnis, and one with intermediate temperature tolerance - Psychrobacter sanguinis. In the thermophile, the relative frequency of T:A>C:G substitutions was higher than in the psychrophile, consistent with the hypothesis that GC-biased mutagenesis contributes to thermal adaptation, although C:G>T:A transitions predominate across all three species. BacNeMu provides a rapid, phylogenetically informed framework for generating biologically meaningful mutation spectra from open databases.
Zeng, Z.; Wang, Y.
Show abstract
Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).
Gueguen, L.-M.; Mathieu, A.; Perin, O.; Droit, A.
Show abstract
Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.