Population-scalable genotyping from low-coverage sequencing data using pangenome graphs
Bolognini, D.; Guarracino, A.; Paleni, C.; Dudley, T. S.; Iacoviello, L.; Raveane, A.; Sudmant, P. H.; Garrison, E.; Soranzo, N.
Show abstract
Pangenome-based genotyping of structurally complex loci remains challenging at low sequencing depths, particularly when samples are of low quality, such as in ancient DNA. We present COSIGT (COsine SImilarity-based GenoTyper), a method that infers structural genotypes by matching short-read coverage patterns to pangenome haplotypes using cosine similarity. COSIGT maintains robust accuracy at low coverage (1-2X), outperforming existing methods where depth-sensitive approaches degrade. We demonstrate scalability to thousands of modern and ancient genomes, enabling population-scale analyses of complex variation from low-coverage data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 96%
- Inferring allele-specific copy number aberrations and tumor phylogeography from spatially resolved transcriptomics 95%
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 95%
Similar papers in this journal
- ProSolo: Accurate Variant Calling from Single Cell DNA Sequencing Data 95%
- Long-read transcriptomics of a diverse human cohort reveals widespread ancestry bias in gene annotation 94%
- Identity-by-descent detection across 487,409 British samples reveals fine-scale population structure, evolutionary history, and trait associations 94%
Similar papers in this journal
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 95%
- A read count-based method to detect multiplets and their cellular origins from snATAC-seq data 95%
- Epiphany: predicting Hi-C contact maps from 1D epigenomic signals 95%
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 96%
- Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation 95%
- Imputation of low-coverage sequencing data from 150,119 UK Biobank genomes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.