An Evolutionary Statistics Toolkit for Simplified Sequence Analysis on Web with Client-Side Processing
Karagol, A.; Karagol, T.
Show abstract
We present the Evolutionary Statistics Toolkit, a user-friendly web-based platform designed for specialized analysis of genetic sequences, which integrates multiple evolutionary statistics. The toolkit focuses on a selection of specialized tools, including Tajimas D calculator with Site Frequency Spectrum (SFS), Shannons Entropy (H), alignment re-formatting, HGSV to FASTA conversion, pair-wise frequency analysis, FASTA to SEQRES, RNA 2D structure alignment, Kyte-Doolittle hydrophilicity plot tool, Chou-Fasman tool, and kurtosis coefficient calculator. Tajimas D is calculated using the reference formula: D = ({pi} - {theta}W) / sqrt(VD), where {pi} corresponds to the average number of differences, {theta}W is Wattersons estimator of {theta}, and VD is the variance of {pi} - {theta}W. Shannons Entropy is defined as H = -{sum} pi * log2(pi), where pi is the probability of occurrence of each unique character (nucleotide or amino acid) in the sequence. The toolkit facilitates streamlined workflows for early researchers in evolutionary biology, genomics, and related fields. With comparing with existing codes, we propose it also emerges as an educational interactive website for beginners in evolutionary statistics. The source code for each tool in the toolkit is available through GitHub links provided on the website. This open-source approach allows users to inspect the code, suggest improvements, or further adapt the tools for their specific usage and research needs. This article describes the functionalities, and validation of each tool within the platform, along with comparison with accessible existing statistical utilities. The toolkit is freely accessible on: https://www.alperkaragol.com/toolkit
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.