Back

GeneFior: A back to basics and transparent multi-tool approach tosequence detection

Dimonaco, N. J.; Lawther, K.

2026-05-18 bioinformatics
10.64898/2026.05.15.724838 bioRxiv
Show abstract

The detection of sequences of interest, such as antimicrobial resistance genes, directly from genomic and metagenomic sequencing data has become routine, enabled by curated reference databases and rapid in silico sequence search tools. Yet most workflows depend on prior assembly, an inherently lossy process in which a substantial proportion of reads fail to assemble or are collapsed into consensus sequences, causing low-abundance variants and nucleotide-level diversity to be systematically obscured. The tools used to interrogate the resulting assemblies compound this further, clustering reference sequences at arbitrary identity thresholds, imposing hidden parameter defaults, and reducing intermediate alignment evidence to summarised outputs that cannot be critically evaluated or reproduced. Here we present GeneFior, a transparent, multitool workflow integrating BLAST, DIAMOND, Bowtie2, BWA, and Minimap2 to search both DNA and protein sequences against any user-supplied reference database. By enforcing genecentric identity and coverage thresholds at both the read and gene level, GeneFior reduces false positives while retaining sensitivity to genuine, low-abundance variants, including those differing at single-nucleotide resolution. Crucially, by exposing all alignment parameters, preserving intermediate outputs, and generating cross-tool consensus detection matrices, GeneFior makes the influence of tool choice, database selection, and parameter configuration on reported gene profiles directly observable and reproducible.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Bioinformatics Advances
203 papers in training set
Top 0.1%
14.6%
2
Bioinformatics
1204 papers in training set
Top 2%
10.7%
3
Genome Biology
637 papers in training set
Top 0.9%
10.3%
4
Nature Methods
385 papers in training set
Top 1%
7.0%
5
Nucleic Acids Research
1281 papers in training set
Top 4%
4.7%
6
GigaScience
212 papers in training set
Top 0.7%
4.7%
50% of probability mass above
7
Nature Biotechnology
172 papers in training set
Top 1.0%
4.2%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
3.9%
9
Nature Communications
5641 papers in training set
Top 33%
3.9%
10
PLOS ONE
5266 papers in training set
Top 36%
3.4%
11
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
3.1%
12
PeerJ
308 papers in training set
Top 3%
3.1%
13
BMC Bioinformatics
457 papers in training set
Top 3%
2.7%
14
Molecular Biology and Evolution
542 papers in training set
Top 3%
2.3%
15
Microbial Genomics
225 papers in training set
Top 1%
2.3%
16
Genome Research
468 papers in training set
Top 3%
2.0%
17
Cell Systems
201 papers in training set
Top 2%
2.0%
18
Scientific Reports
3612 papers in training set
Top 57%
1.6%
19
Cell Reports Methods
165 papers in training set
Top 3%
1.1%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
21
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
22
BMC Genomics
406 papers in training set
Top 10%
0.6%
23
Journal of Molecular Biology
232 papers in training set
Top 5%
0.6%