Optimizing Strain Selection for Association Studies Under Hard Cost Constraints
Rau, C. D.; Bradley, P. H.
Show abstract
Quantitative genetics methods that link genotype to phenotype can be especially powerful tools in model organisms and non-human populations, where researchers can apply well-controlled perturbations to commercially-available collections of strains. However, purchasing and phenotyping large collections can be cost-prohibitive. We evaluate several approaches to select subsets of a panel of strains with the aim of maximizing power under fixed budgetary constraints. Some approaches focus solely on costs, others on genetic diversity, and some on both simultaneously. To assess how these results generalize, we simulate data across two biological and statistical settings: phylogenetic regression across bacterial isolates, and linear-mixed model regression across mouse strains. Surprisingly, we find that simply selecting the cheapest strains until the budget is exhausted (MinCost) is usually among the top-performing methods, and that algorithms that consider genetic diversity without weighting by cost tend to show much lower power overall. Methods that consider both cost and diversity are typically, though not universally, the top performers. A comparison of diversity-maximizing objectives reveals that the most commonly studied objective (MaxMin) tends to be the among the least effective at retaining power. Instead, two alternative objectives that have not been previously studied in this context -- maximizing the total minor allele frequency (MaxMAF) and maximizing the total sum of pairwise genetic distances (MaxSum) -- typically perform best. Finally, in addition to fully simulated data, we also test these approaches by subsampling real data on cardiovascular phenotypes from the Hybrid Mouse Diversity Panel (HMDP). The three best methods on these data are cost-aware (p-Median, MaxSum, and MinCost). Based on these results, we recommend investigators consider approaches that maximize total MAF or pairwise genetic distance weighted by strain cost, but we also conclude that a surprisingly robust approach is to simply pick the most strains one can afford.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phase-free local ancestry inference mitigates the impact of switch errors on phase-based methods 93%
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 92%
- Which mouse multiparental population is right for your study? The Collaborative Cross inbred strains, their F1 hybrids, or the Diversity Outbred population 92%
Similar papers in this journal
Similar papers in this journal
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 93%
- Deep convolutional and conditional neural networks for large-scale genomic data generation 93%
- Discovering functional sequences with RELICS, an analysis method for tiling CRISPR screens 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.