Pipette: Encoding scientific literature into an executable Skill Graph for multi-agent bioinformatics
Gupta, C.; Sharma, A.
Show abstract
The cost of genomic sequencing has fallen by several orders of magnitude, yet data analysis remains a bottleneck concentrated among researchers with specialized computational expertise. While Large Language Models can generate bioinformatics code, they frequently produce incoherent multi-step workflows due to the absence of domain-specific analytical constraints. Here, we present Pipette, a multi-agent AI framework that orchestrates end-to-end bioinformatics workflows through natural language interaction, guided by a literature-derived Skill Graph. This directed, edge-weighted knowledge graph, extracted from over 20,000 peer-reviewed publications, constrains workflow generation to biologically valid analytical transitions, preventing incomplete or incoherent workflows. We benchmarked Pipette across four biological domains using published datasets: single-cell RNA-seq analysis of peripheral blood mononuclear cells and a human pancreas atlas, bulk RNA-seq differential expression in rice under environmental stress, and two structure-based computational drug design workflows. In ablation against two LLMs operating without Skill Graph constraints, Pipette matched or exceeded both baselines across all quantitative metrics while uniquely completing multi-step cross-domain transitions. We further evaluated Pipette on a clinical genomics task, where it executed an ACMG/AMP-compliant variant classification on a reference human genome. In all cases, Pipette recapitulated established biological and clinical findings while generating a fully reproducible, machine-readable provenance record. By reducing the computational expertise required to execute standard genomic analyses, Pipette lowers the barrier between sequencing data and biological insight for bench scientists. Pipette is available at https://pipette.bio.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pan-cancer classification of single cells in the tumour microenvironment 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 96%
Similar papers in this journal
- Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening 95%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 95%
- Conserved epigenetic regulatory logic infers genes governing cell identity 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.