PathoBench: an open community-driven benchmark registry for pathogen bioinformatics tools
Dong, Y.; Li, N.; Chiribau, C. B.; Mitchell, M.; Liu, X.; Perkins, A.
Show abstract
SummaryThe proliferation of pathogen bioinformatics pipelines has outpaced the communitys ability to compare them on common ground. Self-reported performance numbers, ad-hoc evaluation datasets, and inconsistent metrics make pipeline selection difficult for clinical and public-health researchers. We present PathoBench, an open web platform that addresses this gap through three coordinated mechanisms: (i) a curated registry of 26 standard benchmark datasets across 10 human pathogens, each with persistent identifiers and direct download links; (ii) pathogen-specific evaluation metrics that submissions must report, allowing direct head-to-head comparison only on the same dataset; and (iii) a credibility framework combining mandatory dataset attestation, ORCID-linked attribution, public peer comments, and administrator verification. As a case study, four published Mycobacterium tuberculosis drug-resistance pipelines were evaluated against the WHO TB mutation catalogue, demonstrating the frameworks discriminating power. PathoBench is open for community contributions across all ten supported pathogens. Availability and implementationPathoBench is freely available at https://pathobench.vercel.app. Source code is released under the MIT license at https://github.com/BPHL-Molecular/pathobench. The platform requires no installation for end users; programmatic access is available via a Supabase REST API. Contactyibo.dong@flhealth.gov Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.