Haplomatic: A Deep-Learning Tool for Adaptively Scaling Resolution in Genetic Mapping Studies
Douglas, T.; Tarvin, R.; Long, A. D.
Show abstract
Genomic mapping studies face a fundamental trade-off between accuracy and resolution: increasing resolution improves localization of genetic signals but typically reduces the accuracy of frequency estimates due to increased statistical noise. In pooled sequencing studies, this trade-off impacts the accuracy of haplotype frequency estimates, a primary statistic used to identify genetic associations.To mitigate this trade-off we introduce Haplomatic, a novel deep-learning-based tool that adaptively adjusts genomic resolution by predicting haplotype frequency estimation error. Haplomatic generates simulated population data from known recombinant inbred line populations, predicts error through a transformer-based neural network, and adjusts resolution until a target error is achieved. Haplomatic achieves significant resolution gains (15% on average) over previous methods without sacrificing accuracy across multiple evaluated sequencing depths (10x, 50x, and 100x). To our knowledge, this is the first instance of applying deep learning to directly predict estimation error and dynamically scale resolution in trait mapping studies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome wide detection of somatic mosaicism at short tandem repeats 96%
- Generative Haplotype Prediction Outperforms Statistical Methods for Small Variant Detection in NGS Data 95%
- Polaris: Polarization of ancestral and derived polymorphic alleles for inferences of extended haplotype homozygosity in human populations. 95%
Similar papers in this journal
- Robust, flexible, and scalable tests for Hardy-Weinberg Equilibrium across diverse ancestries 96%
- Distinct error rates for reference and non-reference genotypes estimated by pedigree analysis 96%
- Powerful, efficient QTL mapping in Drosophila melanogaster using bulked phenotyping and pooled sequencing 95%
Similar papers in this journal
- Simulation with RADinitio Improves RADseq Experimental Design and Sheds Light on Sources of Missing Data 96%
- pixy: Unbiased estimation of nucleotide diversity and divergence in the presence of missing data 95%
- Versatile simulations of admixture and accurate local ancestry inference with mixnmatch and ancestryinfer 95%
Similar papers in this journal
- SVCollector: Optimized sample selection for cost-efficient long-read population sequencing 95%
- Variation in mutation, recombination, and transposition rates in Drosophila melanogaster and Drosophila simulans 95%
- Assessing and mitigating privacy risk of sparse, noisy genotypes by local alignment to haplotype databases 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.