HAT: de novo variant calling for highly accurate short-read and long-read sequencing data
Ng, J. K.; Turner, T. N.
Show abstract
Motivationde novo variant (DNV) calling is challenging from parent-child sequenced trio data. We developed Hare And Tortoise (HAT) to work as an automated workflow to detect DNVs in highly accurate short-read and long-read sequencing data. Reliable detection of DNVs is important for human genetics studies (e.g., autism, epilepsy). ResultsHAT is a workflow to detect DNVs from short-read and long read sequencing data. This workflow begins with aligned read data (i.e., CRAM or BAM) from a parent-child sequenced trio and outputs DNVs. HAT detects high-quality DNVs from short-read whole-exome sequencing, short-read wholegenome sequencing, and highly accurate long-read sequencing data. Availabilityhttps://github.com/TNTurnerLab/HAT Contacttychele@wustl.edu Supplementary informationSupplementary data are available at bioRxiv.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Algorithmic improvements for discovery of germline copy number variants in next-generation sequencing data 94%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 94%
- NucBreak: Location of structural errors in a genome assembly by using paired-end Illumina reads 94%
Similar papers in this journal
- A graph clustering algorithm for detection and genotyping of structural variants from long reads 96%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
- Identifying, understanding, and correcting technical biases on the sex chromosomes in next-generation sequencing data 95%
Similar papers in this journal
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 95%
- WeavePop: A bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations 94%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.