A novel workflow to improve multi-locus genotyping of wildlife species: an experimental set-up with a known model system
Gillingham, M.; Montero, B. K.; Wilhelm, K.; Grudzus, K.; Sommer, S.; Santos, P.
Show abstract
AO_SCPCAPBSTRACTC_SCPCAPGenotyping novel complex multigene systems is particularly challenging in non-model organisms. Target primers frequently amplify simultaneously multiple loci leading to high PCR and sequencing artefacts such as chimeras and allele amplification bias. Most next-generation sequencing genotyping pipelines have been validated in non-model systems whereby the real genotype is unknown and the generation of artefacts may be highly repeatable. Further hindering accurate genotyping, the relationship between artefacts and copy number variation (CNV) within a PCR remains poorly described. Here we investigate the latter by experimentally combining multiple known major histocompatibility complex (MHC) haplotypes of a model organism (chicken, Gallus gallus, 43 artificial genotypes with 2-13 alleles per amplicon). In addition to well defined "optimal" primers, we simulated a non-model species situation by designing "naive" primers, with sequence data from closely related Galliform species. We applied a novel open-source genotyping pipeline (ACACIA) to the data, and compared its performance with another, previously published, pipeline. ACACIA yielded very high allele calling accuracy (>98%). Non-chimeric artefacts increased linearly with increasing CNV but chimeric artefacts leveled when amplifying more than 4-6 alleles. As expected, we found heterogeneous amplification efficiency of allelic variants when co-amplifying multiple loci. Using our validated ACACIA pipeline and the example data of this study, we discuss in detail the pitfalls researchers should avoid in order to reliably genotype complex multigene systems. ACACIA and the datasets used in this study are publicly available at GitLab and FigShare (https://gitlab.com/psc_santos/ACACIA and https://figshare.com/projects/ACACIA/66485).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Robust and efficient software for reference-free genomic diversity analysis of GBS data on diploid and polyploid species 95%
- A de novo chromosome-level genome assembly of Coregonus sp. \"Balchen\": one representative of the Swiss Alpine whitefish radiation 95%
- Characterization of a Y-specific duplication/insertion of the anti-Mullerian hormone type II receptor gene based on a chromosome-scale genome assembly of yellow perch, Perca flavescens 94%
Similar papers in this journal
- GenErode: a bioinformatics pipeline to investigate genome erosion in endangered and extinct species 96%
- Performance analysis of conventional and AI-based variant callers using short and long reads 95%
- HISS: Snakemake-based workflows for performing SMRT-RenSeq assembly, AgRenSeq and dRenSeq for the discovery of novel plant disease resistance genes. 95%
Similar papers in this journal
Similar papers in this journal
- Comparative analysis of novel MGISEQ-2000 sequencing platform vs Illumina HiSeq 2500 for whole-genome sequencing 94%
- Variant calling and genotyping accuracy of ddRAD-seq: comparison with 20X WGS in layers 94%
- On taming the effect of transcript level intra-condition count variation during differential expression analysis: a story of dogs, foxes and wolves 94%
Similar papers in this journal
- Significantly improving the quality of genome assemblies through curation 95%
- Draft genome assemblies using sequencing reads from Oxford Nanopore Technology and Illumina platforms for four species of North American killifish from the Fundulus genus 94%
- A high-throughput multiplexing and selection strategy to complete bacterial genomes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.