Back

BGC-QDR: A Quantum-Assisted Pipeline for Biosynthetic Gene Cluster Discovery and Ranking from Environmental DNA

Mishra, A.; Rai, A.

2026-06-25 bioinformatics
10.64898/2026.06.21.733574 bioRxiv
Show abstract

Biosynthetic gene clusters (BGCs) encode enzymatic pathways for natural products with pharmaceutical potential, yet prioritizing candidates from fragmented environmental DNA (eDNA) assemblies remains computationally challenging. We present BGC-QDR (Biosynthetic Gene Cluster Quantum Discovery and Ranking), an open-source pipeline that integrates input quality control, Prodigal ORF prediction, Pfam HMM domain annotation, rule-based BGC classification, MiBIG 4.0 novelty assessment, and variational quantum classifier (VQC) ranking via PennyLane. BGC-QDR is designed as a quantum-assisted ranking framework for biologically informed BGC prioritization, not as a claim of quantum computational advantage over classical machine learning. We evaluate the pipeline on MiBIG 4.0 (2,636 annotated BGCs) using a 20-dimensional biosynthetic feature vector and stratified 10-fold cross-validation. The integrated VQC (6 qubits x 3 layers, 54 parameters) achieves accuracy of 0.789 {+/-} 0.076 and ROC-AUC of 0.835 {+/-} 0.057. Random Forest achieves the highest ROC-AUC (0.898 {+/-} 0.032), followed by Logistic Regression (0.874 {+/-} 0.020) and MLP (0.872 {+/-} 0.024). Wilcoxon signed-rank tests on per-fold AUC scores show that VQC ROC-AUC is significantly lower than Random Forest (p = 0.0098) and Logistic Regression (p = 0.037) at = 0.05, with no significant difference versus MLP (p = 0.064). Architecture ablation identifies 4 qubits x 3 layers as the best VQC configuration on hold-out validation (AUC = 0.737). Feature importance analysis highlights peptidyl carrier protein domains, cluster length, and module count as dominant predictors. BGC-QDR provides a reproducible, end-to-end workflow for eDNA-derived BGC discovery with integrated novelty scoring and quantum-assisted candidate ranking. The complete BGC-QDR source code, benchmark datasets, and reproduction instructions are publicly available at: Abhishekmishra2808/BGC-PIPELINE

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
14.8%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.6%
9.6%
3
Nucleic Acids Research
1281 papers in training set
Top 2%
9.6%
4
Briefings in Bioinformatics
354 papers in training set
Top 1%
6.2%
5
Nature Communications
5641 papers in training set
Top 30%
4.8%
6
Nature Methods
385 papers in training set
Top 2%
4.8%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 1.0%
4.2%
50% of probability mass above
8
Cell Reports Methods
165 papers in training set
Top 0.6%
3.4%
9
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
10
BMC Bioinformatics
457 papers in training set
Top 3%
3.2%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.7%
12
Communications Chemistry
48 papers in training set
Top 0.3%
2.4%
13
Nature Biotechnology
172 papers in training set
Top 2%
2.1%
14
Genome Biology
637 papers in training set
Top 5%
1.9%
15
Nature Protocols
33 papers in training set
Top 0.2%
1.9%
16
Advanced Science
286 papers in training set
Top 4%
1.9%
17
Scientific Reports
3612 papers in training set
Top 55%
1.7%
18
Patterns
78 papers in training set
Top 2%
1.3%
19
PLOS Computational Biology
1863 papers in training set
Top 16%
1.3%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 36%
1.1%
21
Journal of Cheminformatics
29 papers in training set
Top 0.6%
1.1%
22
PLOS ONE
5266 papers in training set
Top 57%
1.0%
23
eLife
5828 papers in training set
Top 62%
1.0%
24
Nature Machine Intelligence
70 papers in training set
Top 2%
1.0%
25
Nature Computational Science
55 papers in training set
Top 1%
0.9%
26
Protein Science
246 papers in training set
Top 3%
0.9%
27
Chemical Science
73 papers in training set
Top 2%
0.8%
28
mAbs
32 papers in training set
Top 0.6%
0.6%
29
GigaScience
212 papers in training set
Top 5%
0.6%
30
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.4%
0.6%