Back

COFFEE-PRESC: A fast pre-screening method using compound retrieval by pairwise positional relationship of representative fragments

Shimizu, M.; Yoneyama, S.; Yanagisawa, K.; Akiyama, Y.

2025-12-10 bioinformatics
10.64898/2025.12.07.692874 bioRxiv
Show abstract

Protein-ligand docking is one of the most widely used methods in structure-based virtual screening in the early stages of drug discovery. Its calculations require approximately one minute per compound, making exhaustive evaluation of ultra-large libraries containing billions of molecules computationally impractical. In this study, we propose COFFEE-PRESC (COmpound Filtering by Fragment pair-based Efficient Evaluation for PRE-SCreening), a fast, fragment-based pre-screening method. COFFEE-PRESC first docks fragments in a pre-constructed fragment set to the target protein and enumerates multiple favorable protein-fragment docking poses and then pairs them to consider pairwise positional relationship. The fragment set is composed of a small number of representative fragments that exhibit high similarity to many other fragments, enabling coverage of a large and diverse chemical space. Compounds that contain stuructures similar to fragment pairs are then retrieved through similarity-based searches. This retrieval methodology guarantees that the mutual positional relationship of the two matched fragments does not spatially collide. Finally, the retrieved compounds are evaluated using docking scores of the representative fragments and similarity values between the representative and individual fragments matched in the compound retrieval process. COFFEE-PRESC was 32-fold faster while higher accuracy than Spresso, a existing pre-screening tool, highlighting its potential for application to ultra-large compound library screening. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/692874v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@c9a5eborg.highwire.dtl.DTLVardef@abeddcorg.highwire.dtl.DTLVardef@18d0bfdorg.highwire.dtl.DTLVardef@10e1e33_HPS_FORMAT_FIGEXP M_FIG TOC Graphic C_FIG

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.