PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion¶
Conference: NeurIPS 2026 (task-list assignment; the version read is arXiv v3)
arXiv: 2609.08101
Area: Computational Biology / Structure-Based Drug Design
Keywords: pocket-conditioned generation, variance-exploding diffusion, geometric stability, classifier-free guidance, pocket perturbation
TL;DR¶
PocketVE combines VE/EDM denoising that preserves the shared protein–ligand coordinate scale, training-time pocket perturbation, and inference-time multi-property CFG, raising 3D validity from 58.6% to 80.6% and reducing strain energy from 457.4 to 127.9 relative to guided TAGMol on CrossDocked2020, without attributing these gains to VE alone or claiming superiority on every docking metric.
Background & Motivation¶
Structure-based drug design requires more than a valid molecular graph: a ligand must occupy a specified protein pocket with both plausible internal conformation and appropriate external spatial relationships. Non-autoregressive methods such as TargetDiff jointly denoise all atoms, while TAGMol steers generation with gradients from external property predictors. However, a lower docking score does not guarantee a credible generated conformation: a model can obtain favorable proxy scores by distorting bond lengths, local conformations, or placement within the pocket. Redocking can also obscure problems in the original generated pose, making separate checks of molecular properties, 3D conformation, and protein–ligand clashes necessary.
Consequently, guidance and the denoising backbone cannot be treated as unrelated plug-ins. Property guidance changes Euclidean coordinates that lack pixel-like value bounds, while the protein pocket is an explicit spatial condition. If ligand inputs receive standard EDM scaling but protein coordinates remain unchanged, intermolecular distances no longer reflect their original physical scale. Treating one crystallographic structure as an exact condition can additionally make the denoiser overly dependent on brittle local coordinate details.
PocketVE therefore studies the property–geometry operating curve of an integrated system rather than proving the intrinsic superiority of a diffusion family. It retains TAGMol's equivariant architecture while changing coordinate parameterization, scaling, the sampler, the optimizer, and guidance, and uses several diagnostics to examine the gains. Core idea: first build a hybrid denoiser with consistent pocket coordinates and smooth responses to small perturbations, then apply CFG between its property-conditioned and property-null branches so inference-time control strength tunes the property–geometry trade-off.
Method¶
Overall Architecture¶
The inputs are protein-pocket atom coordinates and features; the outputs are ligand atom coordinates and types. Atom count is sampled from a pocket-size-conditioned prior rather than being a principal contribution. Ligand positions are referenced to the pocket center, not normalized independently for each ligand. Coordinates use variance-exploding diffusion (VE), and atom types use discrete flow matching (DFM); both branches share TAGMol's equivariant features.
Three designs cooperate: pocket perturbation regularization during training, a pocket-scale VE backbone used in training and sampling, and multi-property CFG that amplifies property preferences only at inference time. During training, clean ligands are corrupted, pockets receive small perturbations, and the model learns to recover the original coordinates and atom types. At inference time, sampling begins from noise near the pocket center, progressively reduces noise, and optionally applies CFG to the predictions. The main setting does not add training-style perturbations to test pockets.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
P["Protein pocket"] -->|Training only| A["Pocket perturbation<br/>regularization"]
P -->|Original pocket at inference| B["Pocket-scale<br/>VE backbone"]
A --> B
X["Ligand state"] -->|Training corruption / inference noise initialization| B
B -->|Training: coordinate and type supervision| L["Joint denoising loss"]
B -->|Inference: conditioned and property-null predictions| C["Multi-property CFG"]
Q["Requested property bins"] --> C
C -->|Euler coordinate update + DFM type update| O["Ligand coordinates<br/>and atom types"]
Key Designs¶
1. Pocket perturbation regularization: avoid reliance on one exact crystallographic coordinate realization
During training, small Gaussian noise is applied only to pocket atom coordinates. Atom features remain unchanged, and the ligand supervision target is still the original clean ligand. The main pocket-noise standard deviation is \(\sigma_p=\min(0.1\sigma,0.5)\), where \(\sigma\) is the current ligand-noise scale. At high ligand noise, conditioning can be coarser; at low noise, perturbations decrease to preserve the spatial constraints needed for final refinement. The cap limits augmentation magnitude rather than making the protein itself a jointly generated object.
The rationale is that input noise encourages smoother predictions under nearby pocket-coordinate changes. Appendix E gives a second-order expansion under smooth-loss and small-perturbation assumptions: the extra term involves the loss Hessian in pocket directions. For squared loss with small residuals, it contains a nonnegative Jacobian-sensitivity term. This is a local regularization interpretation, not evidence that the model learns a realistic protein conformational ensemble or remains invariant to arbitrary pocket corruption. Stronger perturbations can impair target specificity, and the experiments do not show simultaneous improvement on every metric.
2. Pocket-scale VE backbone: stable denoising must preserve intermolecular spatial scale
The forward coordinate process directly adds Gaussian noise with standard deviation \(\sigma\) to clean ligand positions without progressively attenuating the clean coordinates. Atom types undergo a mixture of retaining the clean type and uniform categorical corruption at the same noise level; higher noise leaves less type information. The shared equivariant network predicts both clean coordinates and clean-type logits, coupling geometry recovery to atom-identity recovery rather than treating types as continuous coordinate regression.
Standard EDM input normalization creates a pocket-specific problem: shrinking only the ligand at low noise without similarly scaling the protein changes protein–ligand distances. PocketVE retains EDM output preconditioning but uses the following scale-preserving input (Section 3.2 and Appendix B):
As noise approaches zero, the input multiplier approaches one and restores a ligand scale consistent with the unscaled pocket; at high noise, it still attenuates input magnitude. Fixed \(\sigma_{\mathrm{data}}=10.0\) covers both internal ligand extent and displacement from the pocket center, rather than asserting that every molecule has a measured standard deviation of 10. Appendix B further explains that the implementation predicts positions in the scaled coordinate frame: output reconstruction subtracts the scaled input before using the EDM residual and skip terms to return to the original frame. This should be distinguished from the main text's abstract residual-predictor notation; absolute network positions must not be used directly as residuals.
During sampling, the coordinate branch constructs an update direction from the difference between the current state and the clean prediction, then uses first-order Euler updates along a decreasing noise schedule. The type branch performs DFM discrete updates using the predicted distribution. The main configuration uses a 100-step generalized arcsin schedule, noise range \([10^{-3},800]\), and schedule parameters \(\beta=2.2\), \(p=0.57\). Appendix H additionally reports substantial deterioration when sampling-time noise treatment is removed, so the system should not be described as interchangeable with arbitrary deterministic Euler sampling. The sampling details do not fully specify implementation parameters for this noise injection; no additional update equation is reconstructed here.
3. Multi-property CFG: control properties with one model rather than additional gradient heads
Vina, QED, and normalized SA are independently divided into five training-set percentile bins, and the three bin indices form an auxiliary condition. During training, these property conditions are jointly dropped with probability 0.5 while the protein pocket is always retained. The null branch therefore means “no requested properties,” not “no target.” At inference time, the model predicts coordinates under both requested and null properties, then combines them as follows; \(D_{\mathrm{cond}}\) and \(D_{\mathrm{null}}\) are predictions for the same pocket and noisy state.
\(s=0\) selects the property-null branch, \(s=1\) selects the ordinary property-conditioned model, and \(s>1\) performs extrapolative guidance. The type branch uses the same linear combination in logit space followed by softmax, rather than extrapolating normalized probabilities. This eliminates separate noise-dependent property predictors and reduces head-specific calibration of numeric scales and maximization/minimization directions. However, conditioning errors are still amplified: strong guidance does not automatically enforce molecular geometry.
The main experiment requests favorable low-Vina, high-QED, and high-SA bins and uses \(s=5\) as a representative operating point. The main text describes the “lowest Vina bin” according to the original property direction, whereas Appendix O states that larger indices are more favorable and uses \((4,4,4)\) for the favorable joint request. These descriptions do not justify guessing the raw index direction in code. Appendix O also states that training Vina boundaries use Vina Dock while distribution diagnostics use Vina Score. Consequently, the results support directional steering, not strict Vina-bin attainment or independently calibrated control of three properties.
Loss & Training¶
The objective combines noise-weighted squared coordinate reconstruction error and atom-type cross-entropy. Coordinate weighting follows the inverse squared EDM output-preconditioning coefficient. Pocket perturbations and property-condition dropout enter through training inputs; no additional true-binding-energy supervision loss is required. The small-perturbation regularization interpretation also does not mean that the implementation explicitly computes a Jacobian penalty.
Training runs for 400,000 steps with batch size 4, validation every 2,000 steps, and selection of the checkpoint with the lowest validation loss. Hidden two-dimensional weight matrices use Muon at learning rate \(10^{-3}\); other parameters use AdamW at \(10^{-4}\). Learning rates are multiplied by 0.6 after 20,000 steps without validation improvement. Training takes approximately 24 hours on one A100. Changing CFG strength reuses one checkpoint, whereas changing training-time pocket perturbation requires the corresponding training configuration.
Key Experimental Results¶
Main Results¶
The experiments use the standard CrossDocked2020 split filtered for pose quality and sequence similarity, evaluating 100 test pockets with 100 generated ligands each. For baselines with publicly released generated molecules, the authors re-evaluate them using the same split, docking protocol, and GenBench3D pipeline; this does not mean that all methods are retrained with a matched training budget.
The metrics need separate interpretation. Vina Score, Vina Min, and Vina Dock are distinct evaluation variants, with lower values preferred. QED measures drug-likeness, and the normalized SA used here is higher for greater synthetic accessibility. High Affinity is the per-pocket fraction with Vina Dock no worse than the reference ligand. Joint-Ref additionally requires QED and SA individually no worse than the reference ligand in that pocket, then aggregates across pockets. Joint-Ref is neither the fixed-threshold Hit Rate nor the 3D-validity rate.
For geometry, Valid\(_{3\mathrm{D}}\) is the fraction passing GenBench3D conformation checks, including reference-distribution checks of bond lengths and valence angles. Strain energy measures the internal energetic cost of generated conformations, with lower values preferred. GenBench3D clash-free is the clash-free fraction under that protocol; centroid distance checks displacement and is not atom-level pose RMSD. JSD compares generated and reference distance or atom-type distributions: it measures distribution fidelity, not proof of per-molecule physical validity.
The following selection is from Table 1. Vina Score, Vina Dock, QED, and SA are medians; Joint-Ref is the pocket-average percentage; strain follows the source table's scale.
| Method | Vina Score | Vina Dock | QED | SA | Joint-Ref (%) | Valid\(_{3\mathrm{D}}\) (%) | Strain energy | GenBench3D clash-free (%) |
|---|---|---|---|---|---|---|---|---|
| TargetDiff | -6.30 | -7.91 | 0.48 | 0.58 | 5.7 | 78.4 | 306.0 | 86.41 |
| TAGMol (guided) | -7.77 | -8.69 | 0.56 | 0.56 | 7.1 | 58.6 | 457.4 | 93.21 |
| PocketXMol | -6.14 | -7.59 | 0.51 | 0.80 | Unavailable | 67.0 | 80.0 | 93.46 |
| PAFlow | -8.92 | -9.49 | 0.50 | 0.57 | 8.9 | 47.4 | 1834.6 | 96.46 |
| PocketVE (\(s=5\)) | -7.89 | -8.89 | 0.64 | 0.71 | 26.4 | 80.6 | 127.9 | 94.02 |
Relative to TargetDiff, paired bootstrap over complete pockets gives a 3D-validity difference of +2.25 percentage points with 95% CI [-0.75, 5.13], which does not establish significantly higher validity. The strain reduction is 178.10 with CI [150.38, 203.44], a clearer gain. Against guided TAGMol, the mean Vina Score advantage is 0.43 with CI [-0.02, 1.07], also insufficient for strict docking superiority. PocketXMol has better strain and SA, while PAFlow has better Vina scores, further showing that PocketVE does not dominate every method.
Ablation Study¶
The following selection from Appendix Table 5 reuses the same checkpoint and sampling settings, making it suitable for examining inference-time guidance changes. Properties and Vina are medians.
| CFG scale \(s\) | QED | SA | Vina Score | Valid\(_{3\mathrm{D}}\) (%) | Strain energy | GenBench3D clash-free (%) |
|---|---|---|---|---|---|---|
| 0 | 0.48 | 0.63 | -6.70 | 78.64 | 164.9 | 90.59 |
| 1 | 0.60 | 0.70 | -7.20 | 77.99 | 127.4 | 90.74 |
| 5 | 0.64 | 0.71 | -7.89 | 80.60 | 127.9 | 94.02 |
| 10 | 0.64 | 0.72 | -8.12 | 78.54 | 137.7 | 95.03 |
| 20 | 0.60 | 0.69 | -7.61 | 72.04 | 187.5 | 94.51 |
In cumulative Table 2, moving from unguided TAGMol to the PocketVE backbone without perturbation or CFG changes validity from 65.56% to 69.62% and strain from 401.3 to 274.4. Adding the generalized arcsin schedule yields 77.75% / 179.1. This combines parameterization, scaling, sampling, and optimization changes rather than isolating VE causally. Finally, adding pocket perturbation to the no-perturbation \(s=5\) configuration raises validity from 77.33% to 80.60%, but changes strain from 125.9 to 127.9, so perturbation does not improve every metric.
The pocket-scale diagnostic in Appendix Table 17 more directly separates internal and external geometry: removing scale preservation changes mean Vina from -7.589 to +92.237 and clash-free from 94.23% to 3.39%, while 3D validity actually rises from 78.72% to 82.39%. This separate diagnostic uses 100 pockets × 10 ligands and must not be merged into the main table. A plausible internal conformation can still be entirely incompatible with the protein pocket.
Source reporting differences need to be retained. The cumulative \(s=1\) row in Table 2 reports Vina -7.41, validity 78.55%, and strain 137.8, versus -7.20, 77.99%, and 127.4 in Table 5; the rows are not interchangeable. Appendix Table 7 separately summarizes three runs as Vina -7.734 ± 0.0150, validity 79.27 ± 0.68, and strain 124.7 ± 1.29, stating that Table 1 is one of those runs. Its Vina dispersion is difficult to reconcile with Table 1's -7.89 and requires author clarification rather than numerical correction here.
Key Findings¶
- Stronger guidance is not always better: \(s=10\) improves Vina relative to \(s=5\) but reduces validity and increases strain; at \(s=20\), properties and geometry both deteriorate. JSD diagnostics likewise indicate greater reference-distribution drift under stronger guidance.
- Original-pose PoseCheck clash-free rates are only 4.54% for PocketVE and 1.65% for TargetDiff (Appendix Table 10). These differ from the main table's GenBench3D rates of 94.02% and 86.41% and must not be conflated as a single “protein-clash-free rate.” More total contacts do not necessarily indicate better binding.
- Sampling 100 ligands for one pocket takes 234.1 seconds at \(s=5\), 127.8 seconds at \(s=1\), and 1382.5 seconds for TargetDiff. Conditional extrapolation adds two-branch computation. Timings use one A100, batch size 10, and no mixed precision, excluding docking, MMFF, and post-processing.
Highlights & Insights¶
- Internal geometry and pocket compatibility are different axes. The scale diagnostic exposes valid molecules in the wrong spatial context, showing why 3D validity alone is insufficient.
- The property-null branch still sees the pocket, so CFG changes property preferences for the same target. This separation of conditions better reflects the task than dropping every condition together.
- An operating curve over guidance strength is more informative than one best score. Other scientific coordinate-generation tasks can likewise report preference metrics, internal geometry, and external-condition compatibility together.
Limitations & Future Work¶
- The principal evidence comes from one filtered 100-pocket benchmark, property control relies on docking proxies, and there is no experimental binding validation. \(s=5\) is a descriptive choice from the test-benchmark sweep, not a choice from an independent validation set.
- VE, scaling, noise treatment, scheduling, and the optimizer change together, so the experiments do not establish VE-only causal attribution. Future controls should match backbones, optimizers, atom counts, and sampling budgets.
- Gaussian pocket perturbation is local augmentation rather than a protein-flexibility dynamics model. Severe test-pocket noise still damages Vina and clash metrics.
- Property bins are coupled, and favorable joint requests in Appendix O still attain desired property bins at limited rates. Training boundaries based on Vina Dock differ from Vina Score diagnostics, preventing strict calibrated three-objective control claims.
- Appendix A uses \((1+w)\) as the external-guidance coefficient in Eq. (15), but \(w\) in Eq. (18); main-text Eq. (7) also uses \(w\). The derivation has inconsistent coefficient conventions. This note uses the explicit CFG implementation and does not repair the external-guidance derivation.
Related Work & Insights¶
- vs TAGMol: The equivariant denoising architecture is shared, but TAGMol uses external gradient guidance, whereas PocketVE jointly modifies the coordinate backbone, sampling, and property conditioning. The parent comparison is informative at the system level but does not isolate VE's contribution.
- vs TargetDiff / PAFlow: TargetDiff provides a strong geometry reference and PAFlow a strong docking reference. PocketVE offers a favorable multi-metric compromise rather than treating one docking or geometry metric as sufficient.
- vs VEDA / EDM: Standard preconditioning is transferred to a setting with explicit protein coordinates, where the new issue is shared coordinate scale. Whenever samples and conditioning objects enter one geometric graph, normalization should be checked for distorted cross-object distances.
- Relationship to GenBench3D / PoseCheck: They complement conformation-quality and original protein–ligand pose diagnostics, respectively. Pocket permutation is a post-generation specificity check, not regeneration under mismatched conditions or a causal network-component control.
Rating¶
- Novelty: 4/5; the integrated mechanism addresses pocket geometry explicitly, although most core tools come from existing diffusion and CFG methods.
- Experimental Thoroughness: 4/5; multi-axis ablations, distribution checks, and original-pose diagnostics are included, but fully matched component controls and independent validation-based selection are missing.
- Writing Quality: 3/5; claims are generally restrained, but cross-table values, run variability, and the guidance derivation need clarification.
- Value: 4/5; a practical framework for studying coupled property control and geometric stability, not evidence of validated drug-discovery outcomes.