Analytic optimization of Plasmodium falciparum marker gene haplotype recovery from amplicon deep sequencing of complex mixtures
Lapp, Z.; Freedman, E.; Huang, K.; Markwalter, C. F.; Obala, A. A.; Prudhomme-O'Meara, W.; Taylor, S. M.
Show abstract
Molecular epidemiologic studies of malaria parasites commonly employ amplicon deep sequencing (AmpSeq) of marker genes derived from dried blood spots (DBS) to answer public health questions related to topics such as transmission and drug resistance. As these methods are increasingly employed to inform direct public health action, it is important to rigorously evaluate the risk of false positive and false negative haplotypes derived from clinically-relevant sample types. We performed a control experiment evaluating haplotype recovery from AmpSeq of 5 marker genes (ama1, csp, msp7, sera2, and trap) from DBS containing mixtures of DNA from 1 to 10 known P. falciparum reference strains across 3 parasite densities in triplicate (n=270 samples). While false positive haplotypes were present across all parasite densities and mixtures, we optimized censoring criteria to remove 83% (148/179) of false positives while removing only 8% (67/859) of true positives. Post-censoring, the median pairwise Jaccard distance between replicates was 0.83. We failed to recover 35% (477/1365) of haplotypes expected to be present in the sample. Haplotypes were more likely to be missed in low-density samples with <1.5 genomes/{micro}L (OR: 3.88, CI: 1.82-8.27, vs. high-density samples with [≥]75 genomes/{micro}L) and in samples with lower read depth (OR per 10,000 reads: 0.61, CI: 0.54-0.69). Furthermore, minority haplotypes within a sample were more likely to be missed than dominant haplotypes (OR per 0.01 increase in proportion: 0.96, CI: 0.96-0.97). Finally, in clinical samples the percent concordance across markers for multiplicity of infection ranged from 40%-80%. Taken together, our observations indicate that, with sufficient read depth, haplotypes can be successfully recovered from DBS while limiting the false positive rate.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic and transcriptomic evidence for descent from Plasmodium and loss of blood schizogony in Hepatocystis parasites from naturally infected red colobus monkeys 94%
- Emergence of artemisinin-resistant Plasmodium falciparum with kelch13 C580Y mutations on the island of New Guinea 94%
- Oxamniquine resistance alleles are widespread in Old World Schistosoma mansoni and predate drug deployment 94%
Similar papers in this journal
- Impact of sickle cell trait hemoglobin on the intraerythrocytic transcriptional program of Plasmodium falciparum 94%
- Selective whole genome amplification as a tool to enrich specimens with low Treponema pallidum genomic DNA copies for whole genome sequencing 92%
- Metabolic Adaptability and Nutrient Scavenging in Toxoplasma gondii: Insights from Ingestion Pathway-Deficient Mutants 91%
Similar papers in this journal
- Sensitive, highly multiplexed sequencing of microhaplotypes from the Plasmodium falciparum heterozygome 97%
- Amplicon sequencing reveals complex infection in infants congenitally infected with Trypanosoma cruzi and informs the dynamics of parasite transmission 95%
- Fitness costs of pfhrp2 and pfhrp3 deletions underlying diagnostic evasion in malaria parasites 94%
Similar papers in this journal
Similar papers in this journal
- Measuring Growth, Resistance and Recovery after Artemisinin Treatment of Plasmodium falciparum in a semi-high-throughput Assay 94%
- Development of copy number assays for detection and surveillance of piperaquine resistance associated plasmepsin 2/3 copy number variation in Plasmodium falciparum 94%
- Analysis of nucleic acids extracted from rapid diagnostic tests reveals a significant proportion of false positive test results associated with recent malaria treatment 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.