Back

EcoKMER: One-stop shop for spatio-temporal metagenomic exploration using DataFed

George, A.; Leach, D.; Philips, C.; Brown, J.; Duncan, J. M.; Nedved, B.; Leiser, O. P.; Boise, N. R.; Wu, R.; Anderson, L. N.; Kuchar, O.; Cheung, M. S.; Pollock, D. D.; Widener, P.; Johnson, C. G. M.

2025-12-17 bioinformatics
10.64898/2025.12.15.694490 bioRxiv
Show abstract

Spatially distributed environmental sampling generates highly complex and multidimensional datasets illuminating key insights into microbial diversity, evolutionary-coevolutionary processes, and host-pathogen interactions. While these sampling methods generate high value datasets, dataset size, the dataset integration, visualization, analysis, and provenance tracking present significant bottlenecks to scientific discovery. To address this bottleneck, we developed EcoKMER, an R-Shiny front-end application designed to streamline metagenomic data accessibility and provide geospatial context to data in support of hypothesis-driven investigations into environmental sampling, supported by DataFed as its back-end data management platform. EcoKMER enables interactive visualization and filtering of harmonized metagenomic data and metadata using an interoperable approach, allowing users to extract spatially distributed sample-based metadata on top of environmental parameters such as geolocation, temperature, pH, for investigating ecological changes across time and space enhancing sample processing methods. As an example, we deployed this tool to track analysis of metagenomes from the organisms in the Salish Sea Estuary, consistent with existing community-accepted standards. Built on top of DataFed, a flexible and robust scientific data management system built for data lakehouse architectures, EcoKMER is positioned as a powerful tool to improve sampling strategy decision making, accelerate new insights for collaborative biological and environmental research, and fostering AI-ready analyses designed to enhance discovery and guidance for bioeconomic engineering.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.