Back

eProbe: a capture probe design toolkit for genetic diversity reconstructions from ancient environmental DNA

Huang, Z.; Gu, Z.; Cai, Y.; Macleod, R.; Xue, Z.; Dong, H.; Overballe-Petersen, S.; Liu, S.; Gao, Y.; Li, H.; Tang, S.; Diao, X.; Joergensen, M. E.; Dockter, C.; Vinner, L.; Willerslev, E.; Chen, F.; Wang, H.; Wang, Y.

2024-09-03 bioinformatics
10.1101/2024.09.02.610737 bioRxiv
Show abstract

Ancient environmental DNA (aeDNA) is now commonly used in paleoecology and evolutionary ecology, yet due to difficulties in gaining sufficient genome coverage on individual species from metagenome data, its genetic perspectives remain largely uninvestigated. Hybridization capture has proven as an effective approach for enriching the DNA of target species, thus increasing the genome coverage of sequencing data and enabling population and evolutionary genetics analysis. However, to date there is no tool available for designing capture probe sets tailored for aeDNA based population genetics. Here we present eProbe, an efficient, flexible and easy-to-use program toolkit that provides a complete workflow for capture probe design, assessment and validation. By benchmarking a probe set for foxtail millet, an annual grass, made by the eProbe workflow, we demonstrate a remarkable increase of capturing efficiency, with the target taxa recovery rate improved by 577-fold, and the genome coverage achieved by soil capture-sequencing data even higher than data directly shotgun sequenced from the plant tissues. Probes that underwent our filtering panels show notably higher efficiency. The capture sequencing data enabled accurate population and evolutionary genetic analysis, by effectively inferring the fine-scale genetic structures and patterns, as well as the genotypes on functional genes.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.