Biotext: Exploiting Biological-like format for text mining
Machado, D. J. S.; De Pierri, C. R.; Santos, L. G. C.; Pedrosa, F. O.; Raittz, R. T.
Show abstract
The large amount of existing textual data justifies the development of new text mining tools. Bioinformatics tools can be brought to Text Mining, increasing the arsenal of resources. Here, we present BIOTEXT, a package of strategies for converting natural language text into biological-like information data, providing a general protocol with standardized functions, allowing to share, encode and decode textual data for amino acid and DNA. The package was used to encode the arbitrary information present in the headings of the biological sequences found in a BLAST survey. The protocol implemented in this study consists of 12 steps, which can be easily executed and/ or changed by the user, depending on the study area. BIOTEXT empowers users to perform text mining using bioinformatics tools. BIOTEXT is freely available at https://pypi.org/project/BIOTEXT/ (Python package) and https://sourceforge.net/projects/BIOTEXTtools/files/AMINOcode_GUI/ (Standalone tool).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation 95%
- Sequence alignment using machine learning for accurate template-based protein structure prediction 95%
- MLDSP-GUI: An alignment-free standalone tool with an interactive graphical user interface for DNA sequence comparison and analysis 94%
Similar papers in this journal
- NGScloud2: optimized bioinformatic analysis using Amazon Web Services 94%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 93%
- Non-synonymous to synonymous substitutions suggest that orthologs tend to keep their functions, while paralogs are a source of functional novelty 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.