ConfuseNN: Interpreting convolutional neural network inferences in population genomics with data shuffling
Tran, L. N.; Castellano, D.; Gutenkunst, R. N.
Show abstract
Convolutional neural networks (CNNs) have become powerful tools for population genomic inference, yet understanding which genomic features drive their performance remains challenging. We introduce ConfuseNN, a method that systematically shuffles input haplotype matrices to disrupt specific population genetic features and evaluate their contribution to CNN performance. By sequentially removing signals from linkage disequilibrium, allele frequency, and other population genetic patterns in test data, we evaluate how each feature contributes to CNN performance. We applied ConfuseNN to three published CNNs for demographic history and selection inference, confirming the importance of specific data features and identifying limitations of network architecture and of simulated training and testing data design. ConfuseNN provides an accessible biologically motivated framework for interpreting CNN behavior across different tasks in population genetics, helping bridge the gap between powerful deep learning approaches and traditional population genetic theory.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimating allele frequencies, ancestry proportions and genotype likelihoods in the presence of mapping bias 93%
- Phase-free local ancestry inference mitigates the impact of switch errors on phase-based methods 92%
- Which mouse multiparental population is right for your study? The Collaborative Cross inbred strains, their F1 hybrids, or the Diversity Outbred population 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.