iMeta
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match iMeta's content profile, based on 10 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Mainguy, J.; Lemane, T.; Bazin, A.; Arnoux, J.; Gautreau, G.; Medigue, C.; Calteau, A.; Vallenet, D.
Show abstract
PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=63 SRC="FIGDIR/small/742796v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@186b88dorg.highwire.dtl.DTLVardef@1be33d0org.highwire.dtl.DTLVardef@3bc596org.highwire.dtl.DTLVardef@292401_HPS_FORMAT_FIGEXP M_FIG C_FIG
Andrade, R. L.; Fiuza, T. d. S.; Ferraz, R. S.; Kroll, J. E.; Barbosa Araujo, P. V.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; Alves Sobrinho, P. d. A.; de Souza, S. J.
Show abstract
The human gut microbiome plays a central role in host physiology and disease, yet metagenomic analysis pipelines remain fragmented across sample preparation, taxonomic classification, and clinical interpretation stages, complicating reproducibility and translational use. Here we present GTX-GUT, a fully automated, containerized Snakemake pipeline for 16S rRNA gut microbiome profiling that integrates quality control, taxonomic classification (QIIME2/DADA2 against Greengenes 13.8), diversity and compositional metrics benchmarked against a curated healthy reference population, enterotype classification, a clinical association module spanning 11 disease categories, and automated natural-language report generation. We validated the pipeline using the ZymoBIOMICS mock community, showing that BBDuk preprocessing substantially reduced genus-level quantification error (Mean Absolute Error reduced from 7.34 to 1.58 percentage points; Pearsons r improved from 0.576 to 0.833). Application to a human sample from a patient with type 2 Diabetes Mellitus recovered a dysbiotic signature consistent with the literature, including reduced Firmicutes abundance, elevated Bacteroidetes and Proteobacteria, and a predominance of clinical associations within metabolic and gastrointestinal categories. These results demonstrate that GTXGUT provides a reproducible, end-to-end framework linking raw sequencing data to clinically interpretable output, with direct applicability to research and translational microbiome studies.
Zhou, Y.; Huang, F.; Zhao, Y.
Show abstract
Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.
Loureiro, C.; Schorn, M. A.; Alanjary, M.; Kuipers, B.; Louwen, J. J. R.; van der Oost, J.; Medema, M. H.; Sipkema, D.
Show abstract
Marine sponges are known sources of bioactive natural products (NPs), many of which are produced by associated bacterial symbionts via encoded biosynthetic gene clusters (BGCs). A particularly interesting subclass of sponge-derived NPs is comprised of small, brominated alkaloids, which are recovered from diverse habitats and host sponge taxonomies. Despite having been described decades ago, most of these NPs do not have an elucidated biosynthetic origin. We queried metagenomes of several sponge species by making use of a minimal set of core enzymes that we postulate to be necessary to produce these small peptidic NPs: an FADH2-dependent halogenase and an AMP-binding adenylation enzyme. This revealed a variety of novel BGC architectures, many of which showed conservation among sponge host phylogenies and were encoded in the genomes of diverse sponge-associated bacteria. Furthermore, we identified a BGC in the sponge G. barretti that is potentially linked to the production of the iconic barettins, given its enzymatic machinery and specific acidobacterial origin. The present work contributes to the challenging quest to link orphan brominated NPs to their parent BGCs in the sponge holobiont and beyond.
Fritscher, J.; Duncan, A.; Hildebrand, F.
Show abstract
Large-scale metagenomic studies increasingly require taxonomic profiles that are sensitive, precise, strain-resolved and computationally tractable. Existing profilers typically trade taxonomic breadth, sensitivity, precision and speed against one another, limiting their utility for high-resolution microbiome analyses. Here we present protal -- profiling through alignment -- an ultra-fast alignment-based method for species- and strain-resolved profiling of metagenomes. Protal combines a newly developed alignment algorithm, machine-learning-based classification and conserved bacterial marker genes to profile species represented in the standardized and regularly updated GTDB taxonomy. Protal reliably profiles all 143,614 bacterial and archaeal species in GTDB r226 and achieved higher precision on CAMI2 benchmarks than all tested contemporary profilers, including MetaPhlAn 4, mOTUs4, sylph and Kraken2+Bracken (mean species-level precision 98.3% versus 97.5% for the next-best profiler, sylph). In custom benchmarks, protal showed particularly strong gains for rare species represented by a single reference genome and for highly complex communities containing 10,000 species (F1-score 9% and 14% higher than the second best profiler). Because protal retains read alignments, it can reconstruct intraspecific phylogenetic relationships among detected bacteria with accuracy comparable to StrainPhlAn 4. Unlike dedicated strain-profiling workflows, however, protal performs strain analysis concurrently with species-level profiling, making it up to 40-fold faster without requiring additional steps. Together, these features make strain-resolved profiling of thousands of metagenomes feasible on commodity hardware. The software, databases and tutorials are available at https://github.com/4less/protal and http://protal.earlham.ac.uk.
Paliyal, S.; Kaur, B.; Rao, L.; Chakrabortty, A.; Singh, L.; Sehgal, I.; Sharma, M.; Singh, D.; Chaudhry, V.; Mantri, S. S.
Show abstract
The benzoxazolinate moiety is a key functional group found in a few natural products (NPs), exhibiting diverse bioactivities, including antitumor, antibacterial, and cytotoxic activities. Despite their clinical importance, only a few bacterial strains and NPs have been reported harboring this rare bis-heterocyclic moiety, underscoring a largely unexplored chemical space. Here, we performed large-scale genome mining and identified 277 putative biosynthetic gene clusters (BGCs) across diverse bacterial hosts, including previously unreported bacterial genera and strains. The BGCs were grouped into three compound classes: benzoxazolinate, benzobactin, and ashimides based on sequence similarity network clustering. Bioactivity predictions of the identified BGCs revealed the predominance of antibacterial and cytotoxic potential, highlighting promising candidates for future experimental validation and functional studies. This study also presents a neural network-based bioprospecting model that efficiently detects rare BGCs encoding benzoxazolinate-containing molecules from genomic sequences. Overall, our findings expand the known repertoire of bacterial hosts with the potential to produce benzoxazolinate-containing NPs and provide a comprehensive framework for the discovery and identification of candidate BGCs.
Zeng, Z.; Wang, Y.
Show abstract
FigTree is a long-standing phylogenetic tree viewer, but its GUI-centered workflow does not itself provide a versioned, batch-replayable record of styling operations. We present FigTreeKit, a Python package that serializes a supported subset of FigTree 1.4.4 annotations (!hilight, !color, and !font), audits taxonomy mappings before topology-gated clade collapse, retains selected BEAST-style metadata in the tested fixtures, and invokes a patched FigTree renderer for headless PNG, PDF, and SVG output. Across 60 independently generated balanced trees with 50-10,000 taxa (10 trees per size, each timed 10 times as technical replicates), the tree-level log-log slope of export time was 0.96 (95% confidence interval [CI], 0.91-1.01), which is compatible with approximately linear scaling over the tested range but does not prove it. The 189,801-taxon GTDB R232 bacterial reference tree was parsed and exported as a large-data scalability demonstration. On the 10,122-taxon GTDB R232 archaeal reference tree, the scripted workflow assessed 179 order-level groups; 142 multi-tip groups produced non-trivial collapses, whereas 37 singleton groups did not alter the display. The software is accompanied by 796 passing tests, a golden conformance corpus that includes acceptance tests against the bundled FigTree JAR, deterministic scenario-based topology checks, and an overall statement coverage of 81%, reported as a descriptive engineering metric. FigTreeKit is released under the GPL-2.0-or-later license as the figtreekit package on PyPI, with source code, documentation, and benchmark data archived on Zenodo.
McLaughlin, R. J.; Chen, S.; Nag, A.; Noonan, A. J. C.; Bartolomeu, C.; Borden, S. A.; Lam, S.; Myers, R.; Hallam, S. J.
Show abstract
Microbial communities inhabiting the respiratory tract contribute to health status through interactions with host physiology, immune function, and local environmental conditions. Advances in small subunit ribosomal RNA (SSU or 16S rRNA) gene amplicon sequencing enable culture-independent profiling of microbial communities as amplicon sequence variants (ASVs), revealing links between microbial dysbiosis and respiratory diseases, and the use of mass spectrometry to measure volatile organic compounds (VOCs) in exhaled breath shows emerging promise for biomarker discovery. Here we present ASPIRE, the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems, an accessible Nextflow workflow for processing, analyzing, and interpreting linked ASV-VOC data from respiratory microbiome studies. ASPIRE is designed to support scalable comparative analysis across respiratory sample types while preserving intermediate file outputs for inspection and reuse within a standardized file structure. Availability and implementationThe ASPIRE workflow, quick start guide, and usage are freely available on the GitHub platform https://github.com/hallamlab/ASPIRE. ASPIRE is implemented in Nextflow with a modular design to allow for reproducibility, extensibility and scalability.
Zeng, Z.; Wang, Y.
Show abstract
Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).
Gueguen, L.-M.; Mathieu, A.; Perin, O.; Droit, A.
Show abstract
Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.
Farrell, M. V.; Rix, L.; O'brien, P. A.; Dunbar, T. L.; Mahesh, S.; Kuek, F.; Shikuma, N. J.
Show abstract
A major barrier to scaling marine restoration and aquaculture is the lack of reliable tools to induce invertebrate larvae to settle and metamorphose when and where needed. Although microbial cues are known to induce metamorphosis in many invertebrates, existing methods rely on natural biofilms that are variable, difficult to standardize, and unsuitable for large-scale deployment. Here we introduce ReefTiles, a non-living bacterial coating that preserves inductive activity from metamorphosis-stimulating marine bacteria in a stable, reproducible format. Using both tubeworm and coral larvae, we show that dried and inactivated bacterial films retain full settlement-inducing capacity, matching or exceeding live biofilms while eliminating concerns associated with releasing viable microbes into the environment. Viability assays confirm inactivation, and the coating adheres reliably to common substrate materials. Because ReefTiles can be manufactured and stored at scale and tailored to different inductive strains, they provide a practical microbe-based biotechnology for enhancing larval settlement in reef restoration, sustainable aquaculture, and engineered marine infrastructure.
Meng, L.; Zhang, R.; De Castro, C.; Uchiyama, I.; Kanehisa, M.; Ogata, H.
Show abstract
Carbohydrate-active enzymes (CAZymes) shape virus-host interactions by modifying virion structures, host surfaces and extracellular glycans. However, the diversity and evolutionary origins of viral carbohydrate-active enzymes remain poorly understood, partly due to limited viral protein annotations. To address this, we present VirGenes, a database of viral orthologous groups constructed from the KEGG viral gene dataset. VirGenes uses a hierarchical framework that integrates sequence similarity, remote homology, and structural similarity to support evolutionary and functional analyses of viral proteins. By screening the sequence space of VirGenes, we identified 558 CAZyme-associated gene clusters spanning 102 CAZyme families, revealing particularly enriched repertoires in dsDNA viral lineages. Two bacteriophage families, Kleczkowskaviridae and Pootjesviridae, encoded more than 10 CAZymes per genome, followed by Mimiviridae, a representative family of eukaryotic giant viruses. Phylogenetic analyses systematically revealed divergent evolutionary histories of viral carbohydrate-active genes, including frequent horizontal transfer of endolysin genes from bacteria, which likely represents a viral strategy in the ongoing evolutionary arms race with their cellular hosts. Within the structural space of VirGenes, a large number of viral genes were found to contain CAZyme-like folds despite more than 85% of them lacking detectable sequence similarity to annotated CAZyme sequences. Notably, numerous hypothetical sequences from giant viruses exhibited glycoside hydrolase-like five-bladed {beta}-propeller folds. Overall, by integrating sequence, structural and functional evidence, we show that viral carbohydrate-active systems exemplify how distributed innovations, constrained by ancient folds, collectively build the functional complexity of the global virosphere. VirGenes is publicly accessible at https://www.genome.jp/vogdb/.
Li, X.; Zhou, C.; Xiao, K.; Xu, J.; Jia, X.; Zhao, D.; Chen, L.; Li, Y.; Peng, J.; Zhu, J.; Liu, Y.; Shang, X.; Kong, H.
Show abstract
The continuous accumulation of genetic mutations in influenza A viruses (IAVs) drives antigenic drift, necessitating precise antigenic prediction for optimal vaccine strain selection. While sequence-based methods have advanced antigenic surveillance, they neglect the three-dimensional structural context that fundamentally dictates viral antigenicity. Here, we introduce Vir3D, which leverages ESMFold-derived structural information from amino acid sequences to precisely predict viral antigenicity and guide vaccine strain selection. Across both human H3 and highly pathogenic avian H5 subtypes, Vir3D not only accurately discriminates antigenic variants and infers pairwise antigenic distances, but also mechanistically delineates key structural residues driving viral immune evasion. In a decade-long retrospective analysis, Vir3D-prioritized vaccines consistently achieve broader antigenic coverage of circulating strains than World Health Organization (WHO) recommendations. Crucially, Vir3D successfully predicts that the emerging U.S. dairy cattle H5N1 virus (TX/24) remains antigenically stable relative to clade 2.3.4.4b vaccine strains, and subsequent wet-laboratory validation of hemagglutination inhibition (HI) assays definitively corroborates this finding. Overall, Vir3D establishes a powerful, structure-driven framework for proactive influenza surveillance and pandemic preparedness.
Shedleur-Bourguignon, F.; Theriault, W. P.; Thibodeau, A.
Show abstract
Full-length 16S rRNA gene sequencing using Oxford Nanopore Technologies has emerged as a promising approach to improve species-level resolution in microbiota studies. However, the accuracy of taxonomic assignment remains highly dependent on the bioinformatics s used to process Nanopore long-read data. Therefore, the only way to ensure a good level of certainty in obtained results is to use positive controls in the form of mock communities in the experimental designs. In this study, we compared the performance of Epi2Me 16S (using Minimap2 or Kraken2) workflows provided by Oxford Nanopore Technologies and an EMU workflow for full-length 16S rRNA gene analysis. Using a commercial mock community sequenced across multiple Nanopore runs, taxonomic assignment accuracy and reproducibility was evaluated. Epi2Me-Kraken2 exhibited 18 % of incorrect genus-level assignments and failed to identify 3 species present in the mock community. While Epi2Me-Minimap2 achieved an excellent genus-level classification, reporting 9 % of sequences assigned to a genus not in the mock community, species-level assignments were inconsistent for several community members such as Listeria. In contrast, EMU provided accurate and consistent species-level taxonomic profiles, with all species correctly identified while keeping the number of genus absent from the mock community at 1.2%. ImportanceThese results highlight that Epi2Me integrated workflows are not the best option for specie-level taxonomic assignation. More importantly, this paper underscores the importance of routine inclusion of positive controls for microbiota studies, in the form of mock communities, as a critical safeguard for accurate data interpretation. Without the use of a mock community, a paper published would be at risk of reporting wrong observations and inaccurate conclusions.
Demircioglu, E.; Bole, M.; da Rocha, U. N.; Kallies, R.
Show abstract
MotivationAnalysing and concatenating phage annotation is time-consuming. Further, the output of phage annotation tools cannot be directly submitted to public repositories. To deal with these issues, we developed PhaGAMeToo. This command-line workflow for Linux integrates the functional annotations of two major viral annotation tools (Pharokka and VIBRANT), enabling faster and more accurate functional annotation. Furthermore, the workflow provides merged annotations as submission-ready GenBank files. ResultsPhaGAMeToo uses three steps to generate submission-ready GenBank files. The user uses the reoriented viral genomes as inputs for Pharokka and VIBRANT. Pharokka and VIBRANT-generated files are parsed through the PhaGAMeToo workflow to produce a merged GenBank file. Further, PhaGAMeToo also enables the use of BLASTP to annotate hypothetical proteins not identified by Pharokka and VIBRANT. It then merges the results into a submission-ready GenBank file(s). We tested PhaGAMeToo in three different Use Cases. We analysed reference and uncultivated viral genomes manually curated or directly recovered using MuDoGeR in our Use Cases. In the Use Case 1, we analysed four different NCBI reference genomes. In the Use Cases 2 and 3, we analysed seven recently described huge phage genomes and 56 uncultivated viral genomes recovered from 30 soil metagenomes, respectively. Availability and implementationThe source code, documentation, and installation instructions for PhaGAMeToo are available at https://github.com/NFDI4Microbiota/PhaGAMeToo ContactRene.Kallies@uba.de; ebrardemircioglu25@hacettepe.edu.tr Supplementary informationSupplementary data will be made available upon publication.
Dogbegah, W. A.; Opiyo, S. O.; Proscovia Aber, P.; Tiambo, C. K.
Show abstract
Antimicrobial resistance (AMR) continues to pose a major public health threat across Africa, yet available surveillance data remain highly fragmented across private, public, and academic sources. This study analysed continent-wide AMR surveillance patterns by integrating datasets from multiple independent repositories and visualising them through interactive Geographic Information System (GIS) dashboards. The objective was to generate an integrated evidence base that highlights resistance patterns, surveillance disparities, reporting gaps, and opportunities for improved AMR monitoring across Africa. Data were compiled from major private AMR surveillance programmes including Pfizers ATLAS, GSKs SOAR, Johnson & Johnsons DREAM, Venatorxs GEARS, and Shionogis SIDERO-WT covering the period 2004-2022. Public datasets from the WHO Global Antimicrobial Resistance and Use Surveillance System (GLASS) and the Fleming Funds Mapping Antimicrobial Resistance and Antimicrobial Use Partnership (MAAP) were incorporated for 2016-2020, together with published AMR studies conducted between 2010 and 2024. Datasets were harmonised to align key variables including bacterial species, isolate identifiers, antibiotics tested, surveillance source, geographical location, and categorical AMR outcomes while preserving the original structure of the contributing datasets. Interactive dashboards were developed using R Shiny to support spatial visualisation and dynamic analytical exploration of resistance patterns, temporal trends, species distribution, and country-level surveillance coverage. Descriptive analyses including means, standard deviations, medians, interquartile ranges (IQR), frequency distributions, Gini coefficients, Shannon entropy, Herfindahl-Hirschman Index (HHI), and Lorenz curves were used to assess inequality and concentration in country-level AMR reporting across surveillance systems. The integrated analyses revealed substantial heterogeneity and concentration in AMR surveillance reporting across Africa, reflecting major differences in surveillance intensity, laboratory infrastructure, reporting systems, and diagnostic capacity across countries. Private datasets demonstrated broader antibiotic panels and longer temporal coverage, whereas public datasets exhibited substantial gaps in country participation and pathogen-antibiotic representation. Published AMR studies additionally highlighted important surveillance information absent from formal surveillance databases. By integrating multiple streams of AMR evidence, this study demonstrates the value of interactive GIS dashboards as exploratory and updateable surveillance-support tools for improving visibility of fragmented AMR datasets, identifying surveillance disparities, supporting geographically informed interpretation of resistance trends, and strengthening future AMR surveillance harmonisation efforts across Africa.
Nakamura, J.; Miyazaki, K.; Torii, S.; Kitajima, M.; Mikamo, K.; Kimihira, T.; Morimoto, L.; Ashayqa, H.; Ito, J.; Takeshita, K.; Kosugi, S.; Minegishi, Y.; Ito, M.; Hirano, R.; Ishida, S.; Yoshimi, K.; Halfmann, P. J.; Kawaoka, Y.; Mashimo, T.
Show abstract
Rapidly converting viral genome information into deployable molecular tests remains a major challenge in outbreak preparedness. We developed CONAN-SWIFT (Simple Workflow for Isothermal Field Testing), a sequence-to-test platform that integrates computational assay design, reverse-transcription loop-mediated isothermal amplification, CRISPR-Cas3 detection, reagent lyophilization and lateral-flow readout. Sequence-guided assays for Andes virus and Bundibugyo virus were established within approximately three weeks and extended to four additional filoviruses. A web-based designer supported crRNA selection, and systematic RT-LAMP primer optimization improved amplification performance. Recombinant Escherichia coli-expressed Cascade enabled standardized preparation of lyophilized Cas3-detection reagents, which were combined with a battery-operated isothermal device. The portable system detected as few as 10 input RNA copies per reaction within approximately 40 min. It also detected viral RNA and biologically contained, replication-incompetent Ebola virus in spiked human blood and concentrated wastewater. These findings establish the analytical feasibility of a rapidly adaptable CRISPR-Cas3 engineering framework for decentralized detection of emerging RNA viruses.
Salgado, A.; Tomaz, C. R.; Freitas, A. T.; Almeida, A. S.
Show abstract
Microbiome-based biomarkers have been proposed for colorectal cancer (CRC), yet candidate taxa are often interpreted without knowing whether taxonomic profiling workflows can reliably detect and quantify them in human samples. Existing ground-truth studies commonly rely on simplified communities that do not preserve the biological and technical complexity of clinical stool metagenomes. We hypothesized that weak CRC-associated signals, particularly those relevant to early-stage disease, may be missed through analytical non-recovery rather than biological absence. We developed an in silico spike-in framework that embeds CRC-associated signals into clinical stool metagenomes. Ten taxa were introduced individually at six fractions (0.01-5%) or as an equally weighted community at seven total fractions (0.01-10%; effective per-taxon fractions, 0.001-1%), generating 5,770 spike-in metagenomes from 310 samples. The resulting metagenomes were profiled with Kraken2/Bracken and MetaPhlAn 4 to quantify detection, abundance accuracy, false-positive signals, biomarker recovery, and calibration against a known ground truth. Recovery depended strongly on workflow, taxon, abundance, and clinical background. At 0.01%, four taxa- F. nucleatum, P. micra, P. stomatis, and P. intermedia-showed good recovery in 85-90% of samples under Kraken2/Bracken, whereas none achieved good recovery in at least 50% under MetaPhlAn 4. Greater low-abundance recovery was accompanied by a broader artefact-prone background (54.9% versus 0.5% of non-target taxa). Artefact-prone taxa accounted for 96.3% and 100% of enriched off-target differential-abundance calls, respectively. Spike-in-derived artefact exclusion substantially reduced off-target detections where present, while abundance-response modelling provided proof-of-principle correction of systematic abundance distortions in both evaluated configurations. Overall, the evaluated profiling configurations demonstrate that known low-abundance CRC-associated signals can be missed or distorted across a complete biomarker-discovery pipeline. Analytical non-recovery may cause early-detection biomarkers to be missed rather than indicate biological absence. Importantly, artefact-aware filtering and abundance calibration show that these limitations can be partially overcome. Improvements in taxonomic profiling may help bring reliable microbiome-based CRC diagnostics closer to clinical application.
Palmer, C. M.; Thompson, J.; Hwang, J. H.; Ranger, W.; Ane, J.-M.; Venturelli, O. S.
Show abstract
Nitrogen fixation performed by rhizosphere bacteria has the potential to improve the sustainability of cereal crop cultivation. Deciphering the role of interspecies interactions on nitrogen fixation is crucial for devising strategies to enhance this process. To unravel the contributions of interspecies interactions, we constructed synthetic microbial communities from the bottom-up that contain diazotrophic bacteria that fix nitrogen and maize rhizosphere bacteria that do not have this capability. Interactions that impacted nitrogenase activity via growth-independent mechanisms were prevalent in the system. Nitrogenase activity increased and eventually saturated as a function of the number of inoculated diazotrophs. Using a tailored machine learning model for microbiome dynamics and explainable artificial intelligence, we deciphered species contributions on nitrogenase activity and diazotroph growth. We identified a community containing Klebsiella variicola, Herbaspirillum seropedicae, and Stutzerimonas stutzeri as a starting point for developing microbial inoculants for cereal crops. Taken together, these results provide insights into the role of interspecies interactions on nitrogenase activity.
Verma, S.; Arora, N.; Ajay, C. P.; Singh, P.; Mallick, H.; Ghosh, T. S.
Show abstract
Deciphering gut microbiome to host metabolome interaction is critical for understanding how microbial communities generate bioactive signals that shape host physiology and disease. Progress, however, has been hindered by inconsistent metabolite annotations, poor interoperability across studies, and the absence of integrated resources placing microbiome-derived metabolites within their functional, microbial, physiological, and clinical context. Here we present HuMMANet (Human Microbiome Metabolome Annotation Network), a harmonized resource integrating 46 paired gut microbiome metabolome studies (59 study-units; 14,405 samples; 13 disease categories plus a healthy/control reference category) with a scalable metabolite-harmonization framework. HuMMANet resolves heterogeneous annotations through a multi-stage workflow spanning RefMet, HMDB, PubChem, Metabolomics Workbench, SMPDB, MiMeDB 2.0, GNPS/microbeMASST, DrugBank, and DrugCentral, yielding a reference atlas of 54,914 unique metabolites, annotated with standardized chemical identifiers, biochemical pathways, microbial producer associations, physiological distributions, disease links, and structural relationships to approved therapeutics, a unified reference framework for microbiome metabolome research. Applying HuMMANet to a multi-cohort integration of adult serum and fecal metabolomes, we identified 519 serum and 322 fecal metabolites reproducibly associated with gut microbial community composition (PERMANOVA, P < 0.05 in at least 50% of studies in which detected), enriched for specific biomolecular classes and pathways. Cross-referencing these against Health Associated Core Keystone (HACK) taxa revealed 58 serum and 25 fecal metabolites (HACK positive) whose taxon-level associations tracked positively with the taxon specific HACK indices. These reproducible metabolomic signatures of microbiome health included indole3propionic acid, a gut barrier-protective microbial tryptophan metabolite, and 3phenylpropionate. Drug similarity annotation within HuMMANet linked 16 of this serum and 13 fecal HACK positive metabolites to therapeutics used in neurological, inflammatory, and vascular disease. Conversely, 38 serum and 65 fecal metabolites, including imidazole propionate and long-chain acylcarnitines such as ACar 18:0, showed HACK negative signatures previously associated with dysbiosis-linked disease. GNPS/microbeMASST and MiMeDB 2.0 annotations further traced subsets of these metabolites to putative bacterial producers. HuMMANet thus provides a standardized framework for reproducible microbiome metabolome integration, enabling cross study discovery and translational prioritization of conserved microbiome derived metabolic signatures across human populations and disease states.