Back

Biotext: Exploiting Biological-like format for text mining

Machado, D. J. S.; De Pierri, C. R.; Santos, L. G. C.; Pedrosa, F. O.; Raittz, R. T.

2021-04-11 bioinformatics
10.1101/2021.04.08.439078 bioRxiv
Show abstract

The large amount of existing textual data justifies the development of new text mining tools. Bioinformatics tools can be brought to Text Mining, increasing the arsenal of resources. Here, we present BIOTEXT, a package of strategies for converting natural language text into biological-like information data, providing a general protocol with standardized functions, allowing to share, encode and decode textual data for amino acid and DNA. The package was used to encode the arbitrary information present in the headings of the biological sequences found in a BLAST survey. The protocol implemented in this study consists of 12 steps, which can be easily executed and/ or changed by the user, depending on the study area. BIOTEXT empowers users to perform text mining using bioinformatics tools. BIOTEXT is freely available at https://pypi.org/project/BIOTEXT/ (Python package) and https://sourceforge.net/projects/BIOTEXTtools/files/AMINOcode_GUI/ (Standalone tool).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.