Back

BioMaster: Multi-agent System for Automated Bioinformatics Analysis Workflow

Su, H.; Long, W.; Zhang, Y.

2025-01-26 bioinformatics
10.1101/2025.01.23.634608 bioRxiv
Show abstract

MotivationThe rapid expansion of biological data has significantly increased the complexity of bioinformatics workflows, which often involve intricate, multi-step processes. These tasks demand considerable manual effort from bioinformaticians, creating inefficiencies and limiting scalability. Recent advancements in large language model (LLM)-powered agents offer promising solutions to streamline and automate these workflows. However, existing automated systems, while effective for short, well-defined tasks, often struggle with long, multi-step workflows due to challenges such as error propagation, limited adaptability to emerging tools, and the inability of LLMs to generalize to niche bioinformatics tasks. Achieving effective workflow automation requires robust task coordination, dynamic knowledge retrieval, and mechanisms to ensure errors are identified and resolved before they impact downstream processes. ResultsWe present BioMaster, a multi-agent framework designed to automate and streamline complex bioinformatics workflows. BioMaster incorporates specialized agents with role-based responsibilities, enabling precise task decomposition, execution, and validation. It leverages Retrieval-Augmented Generation (RAG) to dynamically retrieve domain-specific knowledge, improving adaptability to new tools and niche analyses. BioMaster also introduces enhanced control over input and output validation to ensure pipeline consistency and employs a memory management strategy optimized for handling long workflows. Experiments across diverse bioinformatics tasks, including RNA-seq, ChIP-seq, single-cell analysis, and Hi-C processing, demonstrate that BioMaster significantly outperforms existing methods in accuracy, efficiency, and scalability. By addressing key limitations in workflow automation, BioMaster offers a robust solution for modern bioinformatics challenges. Availabilityhttps://github.com/ai4nucleome/BioMaster Contactyanlinzhang@hkust-gz.edu.cn

Published in Patterns · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.