DeepKin: Predicting relatedness from low-coverage genomes and paleogenomes with convolutional neural networks
Guler, M. N.; Yilmaz, A.; Katircioglu, B.; Kantar, S.; Unver, T. E.; Vural, K. B.; Altinisik, N. E.; Akbas, E.; Somel, M.
Show abstract
DeepKin is a novel tool designed to predict relatedness from genomic data using convolutional neural networks (CNNs). Traditional methods for estimating relatedness often struggle when genomic data is limited, as with paleogenomes and degraded forensic samples. DeepKin addresses this challenge by leveraging two CNN models trained on simulated genomic data to classify relatedness up to the third-degree and to identify parent-offspring and sibling pairs. Our benchmarking shows DeepKin performs comparably or better than the widely used tool READv2. We validated DeepKin on empirical paleogenomes from two paleological sites, demonstrating its robustness and adaptability across different genetic backgrounds, with accuracy >90% above 10K shared SNPs. By capturing information across genomic segments, DeepKin offers a new methodological path for relatedness estimation in settings with highly degraded samples, with applications in ancient DNA, as well as forensic and conservation genetics.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pre-processing of paleogenomes: Mitigating reference bias and postmortem damage in ancient genome data 96%
- ContamLD: Estimation of Ancient Nuclear DNA Contamination Using Breakdown of Linkage Disequilibrium 96%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 95%
Similar papers in this journal
- CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data 95%
- ForestQC: quality control on genetic variants from next-generation sequencing data using random forest 95%
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 95%
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 93%
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 93%
- Deciphering signatures of natural selection via deep learning 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.