Back

Multi-feature Classification to Improve Colorimetric Loop-Mediated Isothermal Amplification Fidelity

Melton, G.; Negron, D. A.; Hauser, K.; Jagannathan, S.; Tolli, N.; Jennings, K.; Necciai, B.; Sozhamannan, S.; Abramson, B.

2026-06-08 bioinformatics
10.64898/2026.06.03.728514 bioRxiv
Show abstract

Loop-mediated isothermal amplification (LAMP) is a cost-effective and portable assay technique for performing nucleic acid-based diagnostics in the field whose adoption is hindered by design and reproducibility issues. This is due to a complex primer design process that fine-tunes parameters across 6-8 binding regions. The likelihood of assay success depends on satisfying thermodynamic and secondary structure constraints while maintaining target specificity and avoiding overlaps between multiple primers. Software such as the NEB(R) LAMP Primer Design Tool, PREMIER Biosoft LAMP Designer, Primer3, PCR Signature Erosion Tool (PSET), and PrimerExplorer enable automation of this task for researchers. However, in our experience, these programs can sometimes yield inconsistent results in laboratory testing. Here, we approached the issue by comparing and training multiple machine learning (ML) models on primer sets targeting various organisms from working assays and failing ones to determine significant features and improve predictions prior to ordering primer sets. A literature review produced an initial list of primer sets (n=116), which were then filtered down based on reference template availability to discern their FIP/BIP components (F2/F1c and B1c/B2). The final training set (n=109) included sequence and thermodynamic features derived from primers collected from the review (n=74) and those designed in-house with PSET (n=35). Failing assays were difficult to obtain from the publications, so we provided our own (n=23). Using WEKA Experimenter, models were created based on decision tree and Bayesian learning algorithms using an experimental scheme that performed a parameter grid search, seeded replicates, feature selection, and cross-validation while avoiding data-leakage and outputting logs for model comparison, feature analysis, and overfit assessment. Notably, thermodynamic features associated with the F1c and B1c primers consistently appeared in the top ranks according to consensus between information gain, class-correlation, and model-based feature ranking. For classification, the NaiveBayes algorithm had a TP and TN rate of 0.90 ({+/-} 0.02) and 0.73 ({+/-} 0.05) while achieving Cohens kappa coefficient and F-score values of 0.61 ({+/-} 0.06) and 0.91 ({+/-} 0.01). This work highlights how a practical model was built from a small, imbalanced training set incorporating negative research results, of which more are needed to improve generalization and refine parameters critical to assay success.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 15%
12.8%
2
Microbiology Spectrum
469 papers in training set
Top 2%
5.6%
3
Analytical Chemistry
218 papers in training set
Top 0.6%
5.6%
4
Briefings in Bioinformatics
354 papers in training set
Top 1%
5.6%
5
Frontiers in Microbiology
427 papers in training set
Top 3%
4.1%
6
Journal of Clinical Microbiology
130 papers in training set
Top 0.4%
4.1%
7
Scientific Reports
3612 papers in training set
Top 29%
3.5%
8
ACS Synthetic Biology
287 papers in training set
Top 1.0%
2.9%
9
Applied and Environmental Microbiology
339 papers in training set
Top 3%
2.7%
10
BMC Microbiology
49 papers in training set
Top 0.4%
2.5%
11
BMC Genomics
406 papers in training set
Top 3%
2.5%
50% of probability mass above
12
The Journal of Molecular Diagnostics
39 papers in training set
Top 0.3%
2.2%
13
BMC Bioinformatics
457 papers in training set
Top 3%
2.2%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.0%
15
BioTechniques
25 papers in training set
Top 0.2%
2.0%
16
Genomics
64 papers in training set
Top 0.7%
1.8%
17
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
18
Genes
144 papers in training set
Top 2%
1.7%
19
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 2%
1.2%
20
PeerJ
308 papers in training set
Top 8%
1.2%
21
Biology
45 papers in training set
Top 0.5%
1.1%
22
BioMed Research International
28 papers in training set
Top 2%
1.0%
23
Clinical Chemistry
22 papers in training set
Top 0.3%
0.9%
24
SLAS Technology
14 papers in training set
Top 0.2%
0.9%
25
mSystems
394 papers in training set
Top 6%
0.9%
26
Diagnostics
50 papers in training set
Top 2%
0.9%
27
PLOS Computational Biology
1863 papers in training set
Top 19%
0.9%
28
Viruses
332 papers in training set
Top 5%
0.6%
29
SLAS Discovery
25 papers in training set
Top 0.3%
0.6%
30
Microbial Genomics
225 papers in training set
Top 3%
0.6%