Back

MetaClaw: an auditable AI agent for end-to-end, multi-directional metagenomic and multi-omics analysis

Zhang, H.; Li, Z.; Lagniton, P. N. P.; Wang, Z.; Zhao, L.; Li, W.; Duan, p.; Jiang, X.; Ning, K.

2026-07-24 bioinformatics
10.64898/2026.07.21.739769 bioRxiv
Show abstract

Omics studies increasingly depend on long, multi-directional workflows, making auditability as important as individual analytical tools. Existing LLM-driven bioinformatics agents automate parts of this work, but few have been tested for conclusion-level reproduction with traceable execution. MetaClaw splits analysis into a deterministic FlowHub upstream tool flow and a customizable OpenClaw downstream skill container, coupled through one YAML pipeline registry. Per-job bundles archive FlowHub specifications, skill scripts, pinned Dockerfiles, parameters and outputs, while downstream containers run without network access. We benchmarked MetaClaw on downsampled cohorts from four published studies, comprising 94 samples in total. Although downsampling likely limited sensitivity to low-prevalence markers, MetaClaw recovered 2/4 sorghum markers, 3/3 RRMS features, 3/5 CRC markers and 5/5 permafrost marker groups. Across 45 model-by-prompt runs, all backends completed upstream processing, whereas downstream validity depended on model and instruction detail; three decoy-tested endpoints showed no significant differences. In 35 ablation sessions, removing the registry, reference scripts, planning loop or manifest caused distinct losses. Relative to registry removal, the intact scaffold reduced time by 32%, tool calls by 37% and token cost by 47%. MetaClaw therefore links modular architecture, secure rerunnable provenance, biological reproduction, perturbation robustness and measurable resource efficiency in an auditable agentic framework.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.