Back

DeepKin: Predicting relatedness from low-coverage genomes and paleogenomes with convolutional neural networks

Guler, M. N.; Yilmaz, A.; Katircioglu, B.; Kantar, S.; Unver, T. E.; Vural, K. B.; Altinisik, N. E.; Akbas, E.; Somel, M.

2024-08-09 bioinformatics
10.1101/2024.08.08.607159 bioRxiv
Show abstract

DeepKin is a novel tool designed to predict relatedness from genomic data using convolutional neural networks (CNNs). Traditional methods for estimating relatedness often struggle when genomic data is limited, as with paleogenomes and degraded forensic samples. DeepKin addresses this challenge by leveraging two CNN models trained on simulated genomic data to classify relatedness up to the third-degree and to identify parent-offspring and sibling pairs. Our benchmarking shows DeepKin performs comparably or better than the widely used tool READv2. We validated DeepKin on empirical paleogenomes from two paleological sites, demonstrating its robustness and adaptability across different genetic backgrounds, with accuracy >90% above 10K shared SNPs. By capturing information across genomic segments, DeepKin offers a new methodological path for relatedness estimation in settings with highly degraded samples, with applications in ancient DNA, as well as forensic and conservation genetics.

Published in Molecular Ecology Resources (predicted rank #7) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.