BioInformatics Agent (BIA): Unleashing the Power of Large Language Models to Reshape Bioinformatics Workflow
Xin, Q.; Kong, Q.; Ji, H.; Shen, Y.; Liu, Y.; Sun, Y.; Zhang, Z.; Li, Z.; Xia, X.; Deng, B.; Bai, Y.
Show abstract
Bioinformatics plays a crucial role in understanding biological phenomena, yet the exponential growth of biological data and rapid technological advancements have heightened the barriers to in-depth exploration of this domain. Thereby, we propose Bio-Informatics Agent (BIA), an intelligent agent leveraging Large Language Models (LLMs) technology, to facilitate autonomous bioinformatic analysis through natural language. The primary functionalities of BIA encompass extraction and processing of raw data and metadata, querying both locally deployed and public databases for information. It further undertakes the formulation of workflow designs, generates executable code, and delivers comprehensive reports. Focused on the single-cell RNA sequencing (scRNA-seq) data, this paper demonstrates BIAs remarkable proficiency in information processing and analysis, as well as executing sophisticated tasks and interactions. Additionally, we analyzed failed executions from the agent and demonstrate prospective enhancement strategies including selfrefinement and domain adaptation. The future outlook includes expanding BIAs practical implementations across multi-omics data, to alleviating the workload burden for the bioinformatics community and empowering more profound investigations into the mysteries of life sciences. BIA is available at: https://github.com/biagent-dev/biagent.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 96%
- A workflow reproducibility scale for automatic validation of biological interpretation results. 96%
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 95%
Similar papers in this journal
- polars-bio - fast, scalable and out-of-core operations on large genomic interval datasets 96%
- simpleaf: A simple, flexible, and scalable framework for single-cell transcriptomics data processing using alevin-fry 95%
- ExpOmics: a comprehensive web platform empowering biologists with robust multi-omics data analysis capabilities 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.