species2vec: A novel method for species representation
Angelov, B.
Show abstract
Word embeddings are omnipresent in Natural Language Processing (NLP) tasks. The same technology which defines words by their context can also define biological species. This study showcases this new method - species embedding (species2vec). By proximity sorting of 6761594 mammal observations from the whole world (2862 different species), we are able to create a training corpus for the skip-gram model. The resulting species embeddings are tested in an environmental classification task. The classifier performance confirms the utility of those embeddings in preserving the relationships between species, and also being representative of species consortia in an environment.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The Potential of the Primitive: a Network Analysis of Early Arthropod Evolution 92%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 91%
- OxCOVID19 Database: a multimodal data repository for better understanding the global impact of COVID-19 91%
Similar papers in this journal
- Genomic style: yet another deep-learning approach to characterize bacterial genome sequences 91%
- Predicting Phenotypes From Novel Genomic Markers Using Deep Learning 90%
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 90%
Similar papers in this journal
- QuoVidi : a open-source web application for the organisation of large scale biological treasure hunts 92%
- Historical demography and species distribution models shed light on past speciation in primates of northeast India 90%
- occAssess: An R package for assessing potential biases in species occurrence data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.