Back

novelBGC: An interactive dual-score framework for biosynthetic gene cluster novelty assessment and candidate prioritisation

Shukla, G.; Merugu, B.; Sharma, G.

2026-06-18 bioinformatics
10.64898/2026.06.15.732227 bioRxiv
Show abstract

Genome mining now yields tens of thousands of putative biosynthetic gene clusters (BGCs) per project, yet, separating genuinely novel candidates from rediscoveries of known compounds remains the rate-limiting step before experimental validation. Single-axis prioritisation tools, antiSMASH similarity, BiG-FAM GCF distance, and self-resistance-enzyme (SRE) filters such as ARTS, each surface a different facet of evidence, yet their isolated use systematically over-ranks rediscovery-prone BGCs and overlooks genuinely orphan clusters. We present novelBGC, a web-hosted framework that converts these disparate outputs into two deliberately non-inverse continuous metrics per BGC, a Novelty (N) and a Reference Similarity (RS) score which together define a 2D decision plane that resolves rediscoveries, divergent family members, contig-edge artefacts, and uncharted chemistry with interactive visualisations, with all component weights user-tuneable at submission. Retrospective validation across three independent experimental datasets demonstrates the utility of the framework for candidate prioritization. Within the first 186-BGC SRE-guided cloning study, every confirmed bioactive product fell within the low-to-mid N band whereas 55 high-N (N [≥] 0.50) BGCs were never selected. Moreover, in the other two studies, it correctly prioritised the fully orphan lariocidin BGC of Paenibacillus sp. M2 and the divergent within-family indanopyrrole-A idp BGC of Streptomyces sp. CNX-425. Together, these case studies demonstrate that the joint (N, RS) space facilitates prioritization decisions that are difficult to achieve using any single criterion alone. from identical input data. novelBGC requires no command-line expertise, no local tool installation, and no manual integration of intermediate output formats, addressing a well-documented accessibility barrier for wet-laboratory researchers engaging with genome-mining workflows. novelBGC is freely available at https://project.iith.ac.in/sharmaglab/novelbgc/.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.1%
8.8%
2
Bioinformatics
1204 papers in training set
Top 3%
6.7%
3
Bioinformatics Advances
203 papers in training set
Top 0.8%
6.2%
4
Nucleic Acids Research
1281 papers in training set
Top 3%
6.2%
5
PLOS Computational Biology
1863 papers in training set
Top 6%
5.5%
6
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
4.3%
7
PLOS ONE
5266 papers in training set
Top 35%
4.0%
8
Cell Reports Methods
165 papers in training set
Top 0.5%
4.0%
9
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.5%
10
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.4%
50% of probability mass above
11
Nature Communications
5641 papers in training set
Top 34%
3.4%
12
Communications Chemistry
48 papers in training set
Top 0.3%
2.6%
13
Metabolites
53 papers in training set
Top 0.3%
2.6%
14
Nature Protocols
33 papers in training set
Top 0.2%
2.4%
15
ACS Synthetic Biology
287 papers in training set
Top 1%
2.1%
16
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.1%
17
Scientific Reports
3612 papers in training set
Top 57%
1.7%
18
PeerJ
308 papers in training set
Top 7%
1.3%
19
eLife
5828 papers in training set
Top 58%
1.1%
20
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
21
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
22
Frontiers in Bioinformatics
49 papers in training set
Top 0.9%
1.1%
23
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.2%
1.1%
24
Metabolic Engineering
75 papers in training set
Top 0.6%
1.0%
25
Protein Science
246 papers in training set
Top 3%
1.0%
26
Microbial Genomics
225 papers in training set
Top 2%
1.0%
27
iScience
1154 papers in training set
Top 31%
1.0%
28
BMC Genomics
406 papers in training set
Top 8%
0.8%
29
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.8%
30
GigaScience
212 papers in training set
Top 5%
0.6%