BioMANIA: Simplifying bioinformatics data analysis through conversation
Dong, Z.; Zhong, V.; Lu, Y.
Show abstract
The rapid advancements in high-throughput sequencing technologies have produced a wealth of omics data, facilitating significant biological insights but presenting immense computational challenges. Traditional bioinformatics tools require substantial programming expertise, limiting accessibility for experimental researchers. Despite efforts to develop user-friendly platforms, the complexity of these tools continues to hinder efficient biological data analysis. In this paper, we introduce BioMANIA- an AI-driven, natural language-oriented bioinformatics pipeline that addresses these challenges by enabling the automatic and codeless execution of biological analyses. BioMANIA leverages large language models (LLMs) to interpret user instructions and execute sophisticated bioinformatics work-flows, integrating API knowledge from existing Python tools. By streamlining the analysis process, BioMANIA simplifies complex omics data exploration and accelerates bioinformatics research. Compared to relying on general-purpose LLMs to conduct analysis from scratch, BioMANIA, informed by domain-specific biological tools, helps mitigate hallucinations and significantly reduces the likelihood of confusion and errors. Through comprehensive benchmarking and application to diverse biological data, ranging from single-cell omics to electronic health records, we demonstrate BioMANIAs ability to lower technical barriers, enabling more accurate and comprehensive biological discoveries.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- kmtricks: Efficient and flexible construction of Bloom filters for large sequencing data collections 95%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 95%
- ScaleSC: A superfast and scalable single cell RNA-seq data analysis pipeline powered by GPU. 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.