Back

From Text-based Genome, Population Variations, and Transcriptome Datafiles to SQLite Database and Web Application: A Bioinformatical Study on Alfalfa

Qiao, S.; Shen, C.

2021-08-27 bioinformatics
10.1101/2021.08.25.457609 bioRxiv
Show abstract

In this study, a web database application with the Flask framework was developed to implement three types of queries and visualize the results over a bioinformatical dataset from Alfalfa (Medicago sativa). A backend SQLite database was constructed from genome FASTA, population variations, transcriptome, and annotation files with extensions ".fasta", ".gff", "vcf", ".annotate", etc. Further, a supplementary command-line-based Java application was also developed for faster access to the database without direct SQL programming. Overall, Python, Java, and HTML were the main programming languages used in this application. Those scripts and the development procedures are valuable for bioinformaticians to build online databases from similar raw datasets of other species.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.