On convolutional neural networks for selection inference: revealing the lurking role of preprocessing, and the surprising effectiveness of summary statistics
Cecil, R. M.; Sugden, L. A.
Show abstract
A central challenge in population genetics is the detection of genomic footprints of selection. As machine learning tools including convolutional neural networks (CNNs) have become more sophisticated and applied more broadly, these provide a logical next step for increasing our power to learn and detect such patterns; indeed, CNNs trained on simulated genome sequences have recently been shown to be highly effective at this task. Unlike previous approaches, which rely upon human-crafted summary statistics, these methods are able to be applied directly to raw genomic data, allowing them to potentially learn new signatures that, if well-understood, could improve the current theory surrounding selective sweeps. Towards this end, we examine a representative CNN from the literature, paring it down to the minimal complexity needed to maintain comparable performance; this low-complexity CNN allows us to directly interpret the learned evolutionary signatures. We then validate these patterns in more complex models using metrics that evaluate feature importance. Our findings reveal that common preprocessing steps play a central role in the learned prediction method, most commonly resulting in models that mimic a previously-defined summary statistic, which itself achieves similarly high accuracy. In other cases, preprocessing steps introduce artifacts that can lead to "shortcut learning". We conclude that human decisions still wield significant influence on these methods, hindering their potential to learn novel signatures. To gain new insights into the workings of evolutionary processes through the use of machine learning, we propose that the field focus on methods that avoid human-dependent preprocessing. Author summaryThe ever-increasing power and complexity of machine learning tools presents the scientific community with both unique opportunities and unique challenges. On the one hand, these data-driven approaches have led to state-of-the-art advances on a variety of research problems spanning many fields. On the other, these apparent performance improvements come at the cost of interpretability: it is difficult to know how the model makes its predictions. This is compounded by the computational sophistication of machine learning models which can lend a deceptive air of objectivity, often masking ways in which human bias may be baked into the modeling decisions or the data itself. We present here a case study, examining these issues in the context of a central problem in population genetics: detecting patterns of selection from genome data. Through this application, we show how human decision-making can influence model predictions behind the scenes, sometimes encouraging the model to see what we want it to see, and at other times, presenting the model with signals that allow it to circumvent the learning process. By understanding how these models work, and how they fail, we have a chance of creating new frameworks that are more robust to human biases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Domain-adaptive neural networks improve supervised machine learning based on simulated population genetic data 98%
- An approximate full-likelihood method for inferring selection and allele frequency trajectories from DNA sequence data 96%
- Mapping gene flow between ancient hominins through demography-aware inference of the ancestral recombination graph 96%
Similar papers in this journal
Similar papers in this journal
- Efficient detection and characterization of targets of natural selection using transfer learning 97%
- Tensor decomposition based feature extraction and classification to detect natural selection from genomic data 97%
- Discovery of ongoing selective sweeps within Anopheles mosquito populations using deep learning 96%
Similar papers in this journal
Similar papers in this journal
- MEDICC2: whole-genome doubling aware copy-number phylogenies for cancer evolution 95%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 95%
- CNETML: Maximum likelihood inference of phylogeny from copy number profiles of spatio-temporal samples 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.