Back

Distillation enables scalable high-fidelity virtual screening across ultra-large chemical libraries

Dai, J.; Wang, Y.; Shan, N. L.; Mariani, M.; Yu, Z.; Yan, Q.; Golani, L. K.; Surovtseva, Y. V.; Lee, W. H.; Pusztai, L.

2026-07-03 bioinformatics
10.64898/2026.06.29.735361 bioRxiv
Show abstract

Accurate virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on lower-fidelity scoring functions or sampling-based strategies that can limit predictive accuracy and bias the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of the structure-based model Boltz-2 into an efficient sequence-based surrogate. Trained on ~1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. Applied to histone deacetylase 11 (HDAC11), FastBindRank substantially enriched high-confidence binders relative to the background chemical space. The lightweight model captured structural patterns associated with predicted binding, revealing structural determinants of binding. Under a comparable computational budget, FastBindRank achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. Experimental validation confirmed the activity of two novel compounds. These results establish distillation as a practical strategy for scalable, high-fidelity virtual screening of ultra-large chemical libraries.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.1%
30.6%
2
Nature Communications
5641 papers in training set
Top 21%
7.8%
3
Communications Chemistry
48 papers in training set
Top 0.1%
6.6%
4
Advanced Science
286 papers in training set
Top 1%
4.8%
5
Nature Methods
385 papers in training set
Top 2%
4.3%
50% of probability mass above
6
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 14%
4.0%
7
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
8
Cell Systems
201 papers in training set
Top 1%
3.2%
9
Nature Machine Intelligence
70 papers in training set
Top 1%
1.9%
10
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
11
Nature Biotechnology
172 papers in training set
Top 2%
1.7%
12
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
13
Journal of Cheminformatics
29 papers in training set
Top 0.4%
1.7%
14
Patterns
78 papers in training set
Top 2%
1.4%
15
Scientific Reports
3612 papers in training set
Top 62%
1.3%
16
Bioinformatics
1204 papers in training set
Top 7%
1.3%
17
Chemical Science
73 papers in training set
Top 1%
1.1%
18
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
19
Nature Structural & Molecular Biology
18 papers in training set
Top 0.3%
1.1%
20
Journal of Medicinal Chemistry
77 papers in training set
Top 0.7%
1.1%
21
ACS Central Science
71 papers in training set
Top 1%
1.1%
22
Communications Biology
993 papers in training set
Top 25%
1.0%
23
eLife
5828 papers in training set
Top 66%
0.8%
24
PLOS ONE
5266 papers in training set
Top 62%
0.8%
25
Nature Chemical Biology
119 papers in training set
Top 3%
0.6%
26
Cell Chemical Biology
94 papers in training set
Top 2%
0.6%
27
iScience
1154 papers in training set
Top 41%
0.6%