// NATURE NEWS — SPAZIO & SCIENZA
Development of a random background to understand ligand optimization
Nature
(2026) Cite this article
Ligand optimization is central to drug discovery, with hundreds of analogues often designed and synthesized between an initial hit and a therapeutic candidate1,2. The efficiency of this process is unclear, partly because there is no random background for optimization to compare against. Such a random background might emerge from systematic random small substitutions across starting ligands, measuring the likelihood of achieving a substantial improvement in affinity or potency, or other property by any single perturbation. Recent literature has suggested that perhaps 10% of analogues with minor modifications improve upon the potency of a parent by tenfold or more3,4, but this number is clouded by reporting bias, intentional improvement and inter-group variability. To begin to establish a background expectation for ligand optimization, here we systematically modified 18 lead molecules across six targets with single-atom changes; 257 compounds were synthesized. Unexpectedly, 11.3% of these random small perturbation analogues improved potency by tenfold or more. Conversely, they typically had worse in vitro pharmacokinetics. Although it was possible to find analogues where the potency increase compensated for inferior exposure and half-life, resulting in more potent compounds in vivo, overall, a frustrated landscape for ligand optimization is revealed. This study begins to establish a background expectation for ligand potency optimization and offers a simple strategy to do so. It also begins to quantify the challenges confronting the field in moving beyond in vitro potency.
Ligand optimization is central to chemical probe and drug discovery, as initial active molecules rarely have the potency or pharmacokinetic (PK) properties to be viable in vivo1,2. Between the discovery of an initial hit and a clinical candidate many hundreds, occasionally thousands of optimized analogues might be synthesized1,5. Several strategies6,7,8,9,10 have been developed to improve this process, ranging from early empirical approaches such as Topliss Trees11 and quantitative structure–activity relationships12 to contemporary artificial intelligence (AI)-guided PK13 prediction and free energy calculations14,15. The efficiency of these strategies remains uncertain as there is no random background against which to compare them. Using tenfold improvement between analogue and parent as a benchmark for substantial impact, 10% of analogues meet this standard in the ChEMBL database16 (Extended Data Fig. 1a). As these ChEMBL results suffer from success bias, sample multiple types of perturbation and suffer from inter-group irreproducibility17, the public domain offers no sure guidance on what level of improvement one might expect in ligand optimization from random conservative changes.
Random backgrounds have long been used in biology to quantify significance. In genomics, they distinguish between artificial and natural selection18. In epidemiology, random incidence distinguishes genetic diseases from those driven by environmental factors19,20. In protein engineering, random backgrounds help to evaluate improvements in enzyme function21, and alanine scanning introduces a minimal perturbation to find hot spots for binding and function22. A random background for ligand optimization might quantify progress across chemotypes and targets and compare different optimization strategies.
Like alanine scanning, an unbiased background for ligand optimization might involve conservative perturbations, systematically and comprehensively applied across multiple parents with different targets. This resembles a ‘positional analogue scanning’ approach that substitutes ligand non-polar hydrogens with groups such as CH3, OH, Cl, F and Br, and aromatic carbon atoms with nitrogen23. In a retrospective analysis of 110,000 matched molecular pairs in the ChEMBL24 database3,4, about 30% of analogues improved over threefold i