VIVIDHA: Variant Prediction and Visualization Interface for Dynamic High-throughput Analysis
Barasiya, T.; Bharti, N.; Gadhari, R.; Bayaskar, A.; Jain, S.; Saxena, A.; Kasibhatla, S. M.; Joshi, R.
Show abstract
Large scale genome sequencing projects have produced huge datasets that pose challenges of high processing times especially for variant calling, a significant downstream analysis step. Efficient utilization of computational resources for accurate variant prediction in a timely manner is possible using Hadoop MapReduce framework. We have developed VIVIDHA (Variant Prediction and Visualization Interface for Dynamic High-throughput Analysis), a high throughput methodology for prediction of variants based on splitting the alignment file using overlapping regions using Hadoop MapReduce framework. The size of overlap region is user-defined. Three variant callers viz. GATK, VarScan2 and BCFTools have been included to predict variants using a consensus approach. Speed-up observed provides the rationale of better performance as number of compute cores and file size are increased. VIVIDHA is available in both GUI as well as command-line modes and can be downloaded from URL: https://github.com/bioinformatics-cdac/VividhaInstaller
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Academic Tracker: Software for Tracking and Reporting Publications Associated with Authors and Grants 94%
- Artificial intelligence tool for the study of COVID-19 microdroplet spread across the human diameter and airborne space 94%
- BioDiscViz : a visualization support and consensus signature selector for BioDiscML results 94%
Similar papers in this journal
- Machado: open source genomics data integration framework 95%
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 95%
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.