Back

Extensive exploration of structure activity relationships for the SARS-CoV-2 macrodomain from shape-based fragment merging and active learning

Correy, G. J.; Rachman, M. M.; Togo, T.; Gahbauer, S.; Doruk, Y. U.; Stevens, M. G. V.; Jaishankar, P.; Kelley, B.; Goldman, B.; Schmidt, M.; Kramer, T.; Ashworth, A.; Riley, P.; Shoichet, B. K.; Renslo, A. R.; Walters, W. P.; Fraser, J. S.

2024-08-26 biophysics
10.1101/2024.08.25.609621 bioRxiv
Show abstract

The macrodomain contained in the SARS-CoV-2 non-structural protein 3 (NSP3) is required for viral pathogenesis and lethality. Inhibitors that block the macrodomain could be a new therapeutic strategy for viral suppression. We previously performed a large-scale X-ray crystallography-based fragment screen and discovered a sub-micromolar inhibitor by fragment linking. However, this carboxylic acid-containing lead had poor membrane permeability and other liabilities that made optimization difficult. Here, we developed a shape- based virtual screening pipeline - FrankenROCS - to identify new macrodomain inhibitors using fragment X-ray crystal structures. We used FrankenROCS to exhaustively screen the Enamine high-throughput screening (HTS) collection of 2.1 million compounds and selected 39 compounds for testing, with the most potent compound having an IC50 value equal to 130 M. We then paired FrankenROCS with an active learning algorithm (Thompson sampling) to efficiently search the Enamine REAL database of 22 billion molecules, testing 32 compounds with the most potent having an IC50 equal to 220 M. Further optimization led to analogs with IC50 values better than 10 M, with X-ray crystal structures revealing diverse binding modes despite conserved chemical features. These analogs represent a new lead series with improved membrane permeability that is poised for optimization. In addition, the collection of 137 X-ray crystal structures with associated binding data will serve as a resource for the development of structure-based drug discovery methods. FrankenROCS may be a scalable method for fragment linking to exploit ever-growing synthesis-on- demand libraries.

Published in Science Advances (predicted rank #10) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.