Back

DivQuant: Estimation of Species Richness and Entropy from Small Samples

Schmitz, J. E.; Rahmann, S.

2026-06-11 bioinformatics
10.64898/2026.06.08.730836 bioRxiv
Show abstract

Estimating diversity properties of discrete distributions from a small observed sample is a fundamental problem in algorithmic statistics that has applications in many fields, in particular bioinformatics, but also in ecology or linguistics. The two most common diversity measures are the number of distinct elements in a multiset, also referred to as "species richness" in ecology or "alpha diversity" in microbial analysis, and the Shannon entropy, also referred to as "evenness". Estimating these properties from a small sample is particularly challenging for distributions with many rare elements. Thus, many estimators have been proposed in the past that, in practice, work well for different types of distributions. We present DivQuant, an optimization-based, extrapolating richness and entropy estimator with three contributions. First, we formulate the upsampling problem as a convex quadratic program with a Neyman{chi} 2 objective. Unlike the linear program of its predecessor RichnEst, DivQuant admits confidence intervals via{chi} 2 test inversion that are empirically well-calibrated. Second, we replace RichnEsts fixed-threshold fingerprint truncation with the rare/abundant fingerprint split of Valiant and Valiant, which strongly reduces problem size and preserves enough degrees of freedom for the confidence-interval program to remain valid and feasible. Third, we plug the optimal population fingerprint returned by the program into Shannons entropy formula to obtain an entropy estimate. DivQuant attains close-to-nominal 95% confidence intervals in essentially all tested regimes, including six simulated distribution families, Tara Oceans microbiome data, and 10X Genomics scRNA-seq data, while competing state-of-the-art methods (RichnEst, iNext, PreSeq) miss the true richness in up to 80% of instances, well above the nominal 5%. In addition, DivQuant outperforms classical asymptotic entropy estimators (Miller-Madow, CAE) and the extrapolating iNext estimator. Running times remain competitive, with DivQuant typically completing in seconds. DivQuant is available as a command-line tool at https://gitlab.com/rahmannlab/divquant. 2012 ACM Subject ClassificationMathematics of computing[->] Probability and statistics; Mathematics of computing[->] Linear programming; Mathematics of computing[->] Quadratic programming; Applied computing[->] Bioinformatics

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.1%
62.2%
50% of probability mass above
2
Journal of Computational Biology
48 papers in training set
Top 0.1%
5.7%
3
BMC Bioinformatics
457 papers in training set
Top 2%
3.4%
4
PLOS Computational Biology
1863 papers in training set
Top 11%
2.5%
5
Biometrics
23 papers in training set
Top 0.1%
2.2%
6
Nature Biotechnology
172 papers in training set
Top 2%
2.1%
7
BMC Genomics
406 papers in training set
Top 4%
1.8%
8
Nature Communications
5641 papers in training set
Top 47%
1.6%
9
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.4%
10
Biostatistics
24 papers in training set
Top 0.2%
1.4%
11
Genome Biology
637 papers in training set
Top 7%
1.1%
12
Statistics in Medicine
40 papers in training set
Top 0.4%
1.1%
13
Methods in Ecology and Evolution
176 papers in training set
Top 2%
0.9%
14
Bioinformatics Advances
203 papers in training set
Top 4%
0.9%
15
Frontiers in Genetics
230 papers in training set
Top 5%
0.9%
16
Scientific Reports
3612 papers in training set
Top 77%
0.6%
17
The Annals of Applied Statistics
19 papers in training set
Top 0.2%
0.6%
18
GigaScience
212 papers in training set
Top 5%
0.6%
19
PLOS ONE
5266 papers in training set
Top 63%
0.6%
20
Nature Methods
385 papers in training set
Top 7%
0.6%
21
Systematic Biology
144 papers in training set
Top 0.9%
0.5%
22
Cell Systems
201 papers in training set
Top 6%
0.5%
23
BioData Mining
22 papers in training set
Top 1%
0.5%
24
Nature Computational Science
55 papers in training set
Top 2%
0.5%