naturaList: a package to classify occurrence records in levels of confidence in species identification
Rodrigues, A. V.; Nakamura, G.; Duarte, L.
Show abstract
O_LIThere is a big volume of occurrence records available in biodiversity databases, but researchers should guarantee its quality before use it in scientific studies. A problem that might compromise the quality of occurrence data is species misidentification. We address this issue by presenting naturaList, a R package designed to classify species occurrence data according to identification reliability. C_LIO_LInaturaList allows to classify species occurrences up to six levels of confidence in species identification, and to filter occurrence data accordingly. The highest level of confidence is assigned to records identified by a specialist, whose name must be provided by the user. The other five levels of confidence are derived from the occurrence data. We demonstrate naturaList functions using occurrences of Alsophila setosa, a tree fern species from Atlantic Forest, as example. We classified and filtered data in grid cells in order to maintain only the highest-level records in each cell. Then we selected only those records classified in the two highest levels of confidence. C_LIO_LIFrom 323 occurrences of Alsophila setosa displaying geographic coordinates, 69 (21%) were identified by a specialist. After filtering the highest-level records inside grid cells, 102 records remained. From these grid cell filtered data, 38 occurrences (37%) were classified into the highest confidence level. Three records were removed using an interactive map module, due to falling in sea sites or outside the native range size of the species. Since we selected only records classified in the two highest levels of confidence, the final dataset contained 94 occurrence records. C_LIO_LInaturaList guarantees the reproducibility of occurrence data processing and cleaning. Macroecologists, biogeographers and taxonomists might benefit from using naturaList package to evaluate the quality of species identification in occurrence data and by identify sites that need evaluation of taxonomic classification of species. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Evaluating the data quality of iNaturalist termite records 94%
- Reference Sequence Browser: An R application with a User-Friendly GUI to rapidly query sequence databases 94%
- SecSel, a new software tool for conservation prioritization that is applicable to ordinal-scale data for multiple biodiversity features 94%
Similar papers in this journal
- Transects, quadrats, or points? What is the best combination to get a precise estimation of a coral community? 93%
- NGScloud2: optimized bioinformatic analysis using Amazon Web Services 92%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.