Structure-free, site-resolved contrastive learningextends small-molecule discovery beyond the reachof structure-based modeling
Fondrie, W. E.; Canzani, D.; Tatka, L.; Paez, J. S.; Prymolenna, A.; Gutierrez, A.; Robbins, J.; McEllin, B.; Hubbard, E.; Siebenthall, K.; Pino, L. K.; Federation, A. J.
Show abstract
Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer this question by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket. However, the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer no such pocket to explore. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Freed from the requirement for protein structures, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well- folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. Ptarmigan-1 localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a nearest-neighbor query in a rich latent space shared by protein residues and the compounds that bind them. SummaryPtarmigan-1 predicts which small molecules engage a protein directly from sequence and chemistry, with no three-dimensional pose, and screens a 3.4 billion compound library against the entire human proteome in under a day. It matches leading structure-based methods on standard, well-folded targets, and beats them on cryptic and poorly structured sites where those methods have no pocket to work from. It localizes engagement to the correct residues from sequence alone, generalizing to proteins and chemotypes withheld from training and discovering cryptic pockets. Precomputed embeddings make it thousands of times faster per target than a co-folding model, and the same embedding space answers proteome-wide queries as readily as single- target ones. A patent-only STAT6 inhibitor series, unseen in training, provides a retrospective test of the approach. Ptarmigan-1 recovers the actives and correctly localizes them, where a co- folding baseline mislocalizes them.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Development and validation of a potent and specific inhibitor for the CLC-2 chloride channel 95%
- Machine learning classification can reduce false positives in structure-based virtual screening 95%
- Predicting clinical drug response from model systems by non-linear subspace-based transfer learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.