Back

Revisiting use of DNA characters in taxonomy with MolD - a tree independent algorithm to retrieve diagnostic nucleotide characters from monolocus datasets

Fedosov, A.; Puillandre, N.; Achaz, G.

2019-11-11 bioinformatics
10.1101/838151 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWWhile DNA characters are increasingly used for phylogenetic inference, taxa delimitation and identification, their use for formal description of taxa (i.e. providing either a formal description or a diagnosis) remains scarce and inconsistent. The impediments are neither nomenclatural, nor conceptual, but rather methodological issues: lack of agreement of what DNA character should be provided, and lack of a suitable operational algorithm to identify such characters. Furthermore, the reluctance of using DNA data in taxonomy may also be due to the concerns of insufficient reliability of DNA characters as robustness of the DNA based diagnoses has never been thoroughly assessed. Removing these impediments will enhance integrity of systematics, and will enable efficient treatment of traditionally problematic cases, such as for example, cryptic species. We have developed a novel versatile and scalable algorithm MolD to recover diagnostic combinations of nucleotides (DNCs) for pre-defined groups of DNA sequences, corresponding to taxa. We applied MolD to four published monolocus datasets to examine 1) which type of DNA characters compilation allows for more robust diagnosis, and 2) how the robustness of DNA based diagnosis changes depending on the sampled fraction of taxons diversity. We demonstrate that the redundant DNCs, termed herein sDNCs, allow for higher robustness. Furthermore, we show that a reliable DNA-based diagnosis may be obtained when a rather small fraction of the entire data set is available. Based on our results we propose improvements to the existing practices of handling DNA data in taxonomic descriptions, and discuss a workflow of contemporary systematic study, where the integrative taxonomy part precedes the proposition of a DNA based diagnosis and the diagnosis itself can be efficiently used as a DNA barcode. Our analysis fills existing methodological gaps, thus setting stage for a wider use of the DNA data in taxa description.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Systematic Entomology
14 papers in training set
Top 0.1%
18.4%
2
Molecular Phylogenetics and Evolution
69 papers in training set
Top 0.1%
11.8%
3
PeerJ
308 papers in training set
Top 0.1%
10.6%
4
Peer Community Journal
281 papers in training set
Top 0.5%
7.2%
5
Molecular Ecology Resources
171 papers in training set
Top 0.4%
5.5%
50% of probability mass above
6
Systematic Biology
144 papers in training set
Top 0.3%
5.1%
7
PLOS ONE
5266 papers in training set
Top 30%
5.1%
8
Ecology and Evolution
267 papers in training set
Top 2%
3.2%
9
Scientific Reports
3612 papers in training set
Top 34%
3.2%
10
Metabarcoding and Metagenomics
14 papers in training set
Top 0.1%
2.4%
11
Applications in Plant Sciences
23 papers in training set
Top 0.2%
2.4%
12
Methods in Ecology and Evolution
176 papers in training set
Top 0.9%
2.0%
13
BMC Bioinformatics
457 papers in training set
Top 4%
1.7%
14
Royal Society Open Science
214 papers in training set
Top 5%
1.1%
15
BMC Genomics
406 papers in training set
Top 6%
1.1%
16
Insect Science
11 papers in training set
Top 0.2%
0.9%
17
Zoological Journal of the Linnean Society
18 papers in training set
Top 0.4%
0.8%
18
Bioinformatics
1204 papers in training set
Top 9%
0.8%
19
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.8%
20
Frontiers in Plant Science
256 papers in training set
Top 4%
0.8%
21
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.6%
22
Plants
43 papers in training set
Top 2%
0.6%
23
Insects
42 papers in training set
Top 0.8%
0.6%