The challenge of delimiting cryptic species, and a supervised machine learning solution
Derkarabetian, S.; Starrett, J.; Hedin, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe diversity of biological and ecological characteristics of organisms, and the underlying genetic patterns and processes of speciation, makes the development of universally applicable genetic species delimitation methods challenging. Many approaches, like those incorporating the multispecies coalescent, sometimes delimit populations and overestimate species numbers. This issue is exacerbated in taxa with inherently high population structure due to low dispersal ability, and in cryptic species resulting from nonecological speciation. These taxa present a conundrum when delimiting species: analyses rely heavily, if not entirely, on genetic data which over split species, while other lines of evidence lump. We showcase this conundrum in the harvester Theromaster brunneus, a low dispersal taxon with a wide geographic distribution and high potential for cryptic species. Integrating morphology, mitochondrial, and sub-genomic (double-digest RADSeq and ultraconserved elements) data, we find high discordance across analyses and data types in the number of inferred species, with further evidence that multispecies coalescent approaches over split. We demonstrate the power of a supervised machine learning approach in effectively delimiting cryptic species by creating a "custom" training dataset derived from a well-studied lineage with similar biological characteristics as Theromaster. This novel approach uses known taxa with particular biological characteristics to inform unknown taxa with similar characteristics, and uses modern computational tools ideally suited for species delimitation while also considering the biology and natural history of organisms to make more biologically informed species delimitation decisions. In principle, this approach is universally applicable for species delimitation of any taxon with genetic data, particularly for cryptic species.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systematics of a radiation of Neotropical suboscines (Aves: Thamnophilidae: Epinecrophylla) 97%
- Molecular species delimitation in the primitively segmented spider genus Heptathela endemic to Japanese islands 97%
- Cytonuclear discordance in the crowned-sparrows, Zonotrichia atricapilla and Zonotrichia leucophrys: a mitochondrial selective sweep? 97%
Similar papers in this journal
Similar papers in this journal
- Whole-genomes illuminate the drivers of gene tree discordance and the tempo of tinamou diversification (Aves: Tinamidae) 97%
- Paralogs and off-target sequences improve phylogenetic resolution in a densely-sampled study of the breadfruit genus (Artocarpus, Moraceae) 96%
- Sexual signals persist over deep time: ancient co-option of bioluminescence for courtship displays in cypridinid ostracods 96%
Similar papers in this journal
- Parallel evolution and cryptic diversification in the common and widespread Amazonian tree, Protium subserratum 97%
- Divergence, gene flow and the origin of leapfrog geographic distributions: The history of color pattern variation in Phyllobates poison-dart frogs 96%
- Genomic, phenotypic and environmental correlates of speciation in the midwife toads (Alytes) 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.