Back

Nevermore: Target-Conditioned Protein-Ligand Representation Learning for Multi-Objective Lead Optimization with Database-Grounded Retrieval

Refahi, M. S.; Toutounchian, M.; Sokhansanj, B.; Yoo, H.; Brown, J. R.; Ji, H.-F.; Rosen, G.

2026-01-21 bioinformatics
10.64898/2026.01.20.700610 bioRxiv
Show abstract

Target-conditioned molecular design requires optimizing binding affinity to a proposed therapeutic protein target while balancing competing developability constraints (e.g., absorption, distribution, metabolism, excretion, and toxicity; ADMET). Yet many computational pipelines either optimize a single objective or rely on fully de novo generation that can be difficult to control and interpret. We present Nevermore, a target-conditioned, database-grounded framework that combines a geometry-aware protein-ligand affinity oracle with Pareto-aware multi-objective search over an explicit molecular feature space. A central design choice is to optimize in count-based Morgan fingerprint space, where each feature corresponds to a chemically meaningful substructure count, enabling discrete, interpretable "bucket-level" edits. Nevermore learns target-conditioned scores by aligning protein and ligand representations under contrastive objectives and using a similarity-based prediction head; the resulting affinity oracle improves over previously reported benchmark baselines, providing a stronger scoring signal for downstream optimization. Nevermore then steers candidate selection by proposing sparse fingerprint edits, re-ranking candidates under multiple objectives, and projecting edited fingerprints back to valid molecules via nearest-neighbor retrieval from a large compound library. This yields efficient screening without exhaustive enumeration and provides transparent attributions that connect optimization steps to concrete chemical motifs. We evaluate Nevermore on two target case studies (Menin and SARS-CoV-2 Mpro). Across targets, the closed-loop search consistently retrieves candidate sets with improved affinity-property trade-offs compared with random sampling and similarity-only retrieval baselines, while maintaining explicit control and interpretability through discrete feature-space edits. These results support database-grounded, feature-space steering as a practical route to target-conditioned multi-objective lead refinement without relying on fully de novo generation.

Published in Biology · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.