Back

Combining multiple genetic estimates of Ne

Waples, R. S.

2025-10-24 evolutionary biology
10.1101/2025.06.26.661823 bioRxiv
Show abstract

Researchers often use multiple genetic methods to estimate contemporary effective population size (Ne), but few formally combine estimates despite potential benefits for increasing precision. Maximizing benefits requires an optimal, inverse-variance weighting scheme. Methods should be estimating the same parameter, which can be appropriate either for estimates using the same method applied to different time periods, or estimates using different methods applied to the same time period. Previous approaches focused on [Formula] for weighting, but that is problematical because [Formula] is highly skewed and can be infinitely large. A new approach is described using weights inversely proportional to [Formula], which is the drift signal that [Formula] estimation methods respond to. The distribution of [Formula] is close to normal even when [Formula] assumes extreme values. Benefits are maximized under three general conditions: estimators have approximately equal variances; they are uncorrelated or have weak positive correlations; individual estimates have low precision (i.e., if data are limited and/or true Ne is large). Analytical and numerical results demonstrate that: (1) existing theory allows robust estimates of [Formula] for the temporal and LD methods, which provide independent information about Ne - both of which facilitate optimally combining those methods; (2) estimates for the LD and sibship methods are essentially uncorrelated when data are limited but can be strongly positively correlated in genomics-scale datasets. General theory predicting [Formula] for the sibship method is lacking, but values for specific scenarios have been published. New software (CO_SCPLOWOMBOC_SCPLOWNO_SCPLOWEC_SCPLOW) is introduced to calculate combined estimates.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.