Skip to content

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Conference: NeurIPS 2026 Evaluations and Datasets Track (source: arXiv comment; not the regular main track)
arXiv: 2606.30170
Code: https://github.com/blaschma/TheNanotechnologyMolecularOptimizationBenchmark
Area: Computational Biology (molecular optimization; applications in nanophysics)
Keywords: molecular optimization, quantum simulations, electrode anchors, synthetic pretraining, generative flow networks

TL;DR

NMO tests molecular optimization beyond pharmaceutical priors through three nanophysics tasks constrained by electrode binding, and combines explicit-anchor GGS representations, random-molecule pretraining, and genetic-guided GFNs to discover scientifically promising candidates, although the full model does not maximize AUC on every task.

Background & Motivation

Molecular generation is commonly evaluated on pharmaceutical benchmarks such as PMO, with objectives including similarity, drug-related properties, and their combinations. Their oracles are usually inexpensive, while large pharmaceutical datasets such as ZINC provide strong pretraining priors. Consequently, leaderboard performance can reflect search ability, proximity between the pretraining distribution and the solution, and task-specific vocabularies; it does not establish that a model can discover molecules in an unfamiliar scientific domain. NMO uses molecular design for nanodevices to expose this distinction: the object being optimized is not an isolated molecule but a molecule connected to gold electrodes, where attachment positions alter electronic and vibrational transport.

The barrier to a general-purpose benchmark is more than missing data. A molecular string must be converted into a device geometry, relaxed, and evaluated through quantum transport or spectroscopy. A small structural change can substantially shift interference features, producing a rugged reward landscape with expensive feedback. The paper packages these calculations into scalar oracles using semi-empirical xTB methods, enabling participation without expertise in quantum physics. Synthetic accessibility, conformational complexity, device geometry, and stability constraints prevent optimization from relying solely on extreme structures with attractive scores.

The authors also find that pharmaceutical priors help inconsistently across tasks, while implicit anchor placement particularly constrains two-sided junction tasks. The contribution therefore extends beyond three scoring functions: it includes a representation with optimizable attachment sites and a training route independent of historical pharmaceutical data. Core Idea: under a shared configuration and evaluation budget, optimize molecular structure together with electrode anchors, using physical oracles to distinguish cross-domain optimization ability from the benefits of pharmaceutical distributional priors.

Method

Overall Architecture

NMO contains thermoelectrics (TE), phonon transport (PH), and molecular optomechanics (MO). Given a molecular string and attachment information, the backend parses the structure, adds gold–thiol anchors, relaxes and aligns the geometry, constructs the physical system, and runs simulations to return fitness and scientific metadata. TE and PH require two-sided electrode binding. MO attaches only one end to the substrate, but the other endpoint still determines orientation and device geometry.

GGS is an optional interface, not a participation requirement: ordinary SMILES with explicit atom indices can also specify anchors. Many SMILES baselines in the paper instead use the first and last non-hydrogen atoms as implicit anchors; these experiments do not prove that SMILES cannot express explicit anchor positions. The proposed baseline first pretrains on random GGS graphs and then alternates policy sampling, graph-level genetic refinement, oracle evaluation, and replay training. DCD, DEX, and descriptor supervision affect training and search; they are not additional physical evaluation metrics.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Physical tasks and<br/>shared protocol"] -.->|defines oracle and budget| C
    A --> B["GGS and synthetic<br/>pretraining"]
    B -.->|pretrained weights| C["Genetic-guided GFN"]
    C -->|policy and genetic candidates| O["Filters and physical oracle"]
    O -->|scores and candidates| R["Replay buffer and<br/>final candidates"]
    R -.->|training samples| D["Adaptive stability and<br/>descriptor supervision"]
    D -.->|training update| C

Solid arrows represent candidate or data flow; dashed arrows represent protocol, initialization, or training signals. Final candidates come from the search history across the evaluation budget, rather than from treating the training loop as a purely feed-forward inference network.

Key Designs

1. Physical tasks and shared protocol: connect scores to interpretable device properties

Each task multiplies its target physical property by synthetic-accessibility and rotatable-bond penalties, along with hard constraints. Violating a hard constraint yields zero fitness: examples include unreasonable junction geometry, unstable-substructure filters, and an excessively small HOMO–LUMO gap. The PH implementation specifies a gap threshold of 0.2 eV. The shared rotatable-bond penalty discourages large, experimentally difficult conformer ensembles, so the objective is not an unconstrained physical extreme.

\[ f_{\text{task}}(x)=c_{\text{task}}\mathcal{O}_{\text{task}}(x)\mathcal{P}_{\mathrm{SA,task}}(x)\mathcal{P}_{\mathrm{rot}}(x)\Theta(x),\qquad \mathcal{P}_{\mathrm{rot}}(x)=\frac{1}{1+e^{N_{\mathrm{rot}}-3.5}}. \]

PH minimizes room-temperature phononic thermal conductance: the physical objective is its reciprocal, the SA penalty is reciprocal SA, and the scaling constant is 0.5. The simulation derives vibrational information from the relaxed structure's mass-weighted Hessian, calculates phonon transmission through Green functions describing electrode coupling, and integrates it to obtain thermal conductance at 300 K. This explains why attachment positions, substituent masses, and inter-ring torsion matter: they change transport channels rather than simply improving a pharmaceutical descriptor.

TE maximizes thermoelectric conversion efficiency by retaining electronic transport and the Seebeck coefficient while suppressing electronic and phononic heat transport. It uses Z as the objective, a scaling constant of 0.01, and an SA penalty decreasing linearly from 1 at SA=1 to 0 at SA=10. The relationship below explains both the connection to PH and the difference: PH targets only the phononic contribution, whereas TE must avoid suppressing useful electric current together with heat transport. Physical candidates are generally reported using dimensionless ZT, not fitness; these quantities are not interchangeable.

\[ Z=\frac{GS^{2}}{\kappa_{\mathrm{ph}}+\kappa_{\mathrm{el}}},\qquad ZT=Z\,T. \]

TE explicitly includes gold electrode clusters and combines the electronic Hamiltonian, overlap matrix, and electrode surface Green functions to calculate electronic transmission, conductance, the Seebeck coefficient, and electronic thermal conductance. Electrode level alignment matters: even with the same internal molecular structure, contact geometry changes transmission near the Fermi energy and therefore thermoelectric performance. NMO's structural constraints arise from this device-level physics, not arbitrary string-format requirements.

MO searches for molecules that upconvert THz/MIR radiation into visible or near-infrared signals. Useful vibrational modes must support both infrared absorption and Raman scattering. The paper sums upconversion intensities over modes in the 30–1000 cm⁻¹ range, applies a logarithm and standardization, and obtains P. The geometric contribution F penalizes larger cavity spacing g and molecular footprint S: the former reduces field enhancement, while the latter reduces the number of molecules in the cavity. S below denotes an area, not the Seebeck coefficient in the preceding equation; the standardization parameters come from an existing reference molecular set.

\[ P=\frac{1}{\sigma}\left(\log_{10}\left(\sum_{m\in M}I_m^c\right)-\mu\right),\qquad F=-\frac{2}{\sigma}\log_{10}(g)-\frac{1}{\sigma}\log_{10}(S). \]

MO uses P+F as its physical objective, a scaling constant of 1/4, and the square of TE's linear SA penalty. The backend estimates spectral intensities using GFN2-xTB vibrational frequencies and PTB dipole and polarizability derivatives. These are computationally tractable approximations, not experimental measurements or a guarantee of DFT validation for every candidate.

The protocol allows at most 10,000 fitness evaluations per task and seed, averaged over five consecutive, non-cherry-picked seeds. Expensive oracle calls before optimization are prohibited, as are task-specific datasets and separate hyperparameter tuning; all three tasks must share a configuration and fragment library. Duplicates and candidates rejected by cheap filters need not consume the evaluation budget. PH permits length bounds to avoid a degenerate regime where arbitrarily long molecules artificially suppress thermal conductance.

Evaluation includes Top-10 AUC, final Mean Top-10 Fitness, Mean Top-10 SA, and the Relevance Indicator (RI). AUC tracks the quality of the ten best molecules as oracle calls accumulate, capturing both sample efficiency and final quality. RI checks whether a candidate exceeds a domain threshold while passing the SA criterion. The physical thresholds are ZT>3 for TE, thermal conductance<0.25 pW/K for PH, and P>7.88 at the xTB level for MO. The practical SA boundary is 4.5, but the paper defines RI strictly using SA<4.5; a candidate rounded to 4.50 cannot automatically be counted as passing. An n/5 entry counts seeds finding a relevant candidate, not whether every Top-10 molecule qualifies.

2. GGS and synthetic pretraining: encode optimizable anchors and chemical assembly rules

Group SELFIES represents molecules using predefined fragments, incoming and outgoing coupling sites, and branch-return tokens. However, its post-hoc parser can truncate a sequence when attachment sites are unavailable. Different action sequences may therefore map to the same shortened molecule, disconnecting reward from the intended generation trajectory. GGS instead maintains a fragment-level directed acyclic graph, tracks available coupling sites and valence during construction, and permits only compatible connections. Source and sink fragments, together with their selected coupling sites, naturally specify electrode anchors rather than leaving attachment information to an arbitrary scoring-time choice.

“Valid by construction” has a precise boundary: a successfully constructed GGS graph corresponds to a chemically valid molecule, but an autoregressive agent can still emit strings that cannot be converted into graphs. Consequently, the full method retains a nonzero invalid rate. Chemical validity also does not establish stability, synthesizability, or feasible device geometry; filters, SA assessment, and geometry checks remain necessary. The current implementation limits branch-return depth to one and does not explicitly encode chirality or stereochemistry. A fragment-level DAG does not preclude aromatic rings inside fragments.

The same graph constraints support targeted genetic operations. Crossover exchanges subgraphs around legal cut points. Mutations can replace fragments, move internal bonds, modify electrode anchor positions, insert or delete fragments, add side branches, and change terminal fragments, while checking attachment-site requirements. Anchor-position mutation specifically searches how the same backbone should contact an electrode. Chemically valid graph-level offspring can still receive zero fitness because of physical hard constraints.

The authors randomly assemble 300,000 molecular graphs and convert them into GGS sequences for pretraining, also providing corresponding SMILES data to help distinguish data and representation effects. Random assembly uses the fragment library, connection rules, and length/branch restrictions; cheap stability filtering can be applied without calculating target quantum properties. The model learns a validity and assembly-syntax prior rather than the frequency of structures in pharmaceutical databases. Independence from domain-specific datasets does not mean absence of priors: fragment selection and the random assembly distribution remain explicit biases.

The fragment vocabulary can become a “Chemist's Shop” based on available laboratory reagents. Across the NMO physics tasks, sulfur-containing backbone fragments are excluded so that sulfur is reserved for anchors, reducing alternative gold-binding sites. The library includes acetylene and other units known to matter for molecular transport. This provides a controllable, discussable search space, but neither proves that every fragment combination has a feasible synthesis route nor removes the physical knowledge embedded in the library.

3. Genetic-guided GFN: use expensive scores to improve both candidates and the sampling policy

Genetic GFN aims to sample proportionally to reward rather than training only a greedy generator. At each iteration, the agent autoregressively samples candidates, which reach the oracle only after deduplication and filtering. High-reward candidates enter a replay buffer. A genetic algorithm selects parents by rank, applies crossover and mutation, filters and evaluates new offspring, and returns sufficiently strong candidates to the buffer. Policy updates use buffered trajectories, allowing structures discovered by genetic search to influence later generation rather than remaining isolated local improvements.

This cooperation matters when physical evaluations are expensive and feasible structures are sparse. Policy sampling alone may rarely find valid junction geometries, while genetic search alone may be limited by its initial population. In the representative TE run shown in the appendix, policy sampling supplies early candidates and genetic refinement finds stronger structures after approximately 3,000 oracle calls. In PH, crossover discovers a strong candidate near call 8,000, after which the policy continues exploring that region. These are analyses of selected runs, not fixed behavior promised across seeds.

The authors replace the original GRU with a causal decoder-only Transformer using RoPE and separate next-action policy and molecular-descriptor regression heads. GGS actions include fragments, incoming/outgoing coupling sites, pop, and end; removing sulfur-containing fragments leaves 91 actions for the physics tasks. The mechanism is not simply that a larger Transformer performs better: the policy learns sequences with explicit attachment semantics and reuses expensive evaluations through replay.

4. Adaptive stability and descriptor supervision: address syntax forgetting, mode collapse, and representation degradation

Exceptionally high physical rewards can pull the policy too aggressively toward a few trajectories, producing two distinct failures. DCD activates when invalid sequences exceed 35%, temporarily switching replay training from rank-based to uniform sampling and quadrupling batch size. Exposure to a broader set of valid sequences helps recover syntax. DEX activates when unique molecules fall below 30% of a generated batch, increasing the rank-sampling coefficient to flatten the sampling distribution and counter duplication and mode collapse. It changes sampling, not the softmax temperature directly.

The appendix's trajectory analysis also specifies deactivation hysteresis: DCD stops once the invalid rate falls below 0.1, and DEX stops once the duplicate rate falls below 0.3. The method removes the original Genetic GFN's KL anchoring because its random prior teaches chemical syntax; there is no reason to keep the optimized policy close to a random-molecule distribution. Adding stability mechanisms and removing KL is a joint intervention, so the ablation cannot isolate DCD or DEX individually.

Descriptor supervision predicts 17 inexpensive RDKit properties from the Transformer's latent representation, including molecular weight, ring counts, rotatable bonds, polar surface area, and LogP. Targets are standardized using the random pretraining dataset's mean and standard deviation to prevent large-magnitude properties from dominating gradients. This supervision remains active during both pretraining and task optimization without additional quantum oracle calls. Its purpose is to preserve interpretable structural information, not to treat the descriptors as task objectives or extra input labels. The benefit is task-dependent: TE improves, whereas MO AUC and final fitness decline.

Loss & Training

Pretraining minimizes sequence negative log-likelihood to learn valid GGS syntax. With descriptors, the sum of sequence and descriptor losses is multiplied by 1/2. During task optimization, physical fitness is exponentiated into a positive reward, and a simplified trajectory-balance objective trains the policy and a learnable normalization constant. The following equations reflect the paper rather than adding other terms from general GFN formulations.

\[ R(x)=\exp(\beta f(x)),\qquad \mathcal{L}_{\mathrm{TB}}=\left(\log Z_{\mathrm p}+\sum_{t=0}^{T-1}\log P_{\theta}(s_{t+1}\mid s_t)-\log R(x)\right)^2. \]
\[ \mathcal{L}_{\mathrm{desc}}=\frac{1}{K}\sum_{k=1}^{K}\left(\hat y_{\mathrm d,k}-\frac{y_{\mathrm d,k}-\mu_k}{\sigma_k}\right)^2,\qquad K=17,\qquad \mathcal{L}_{\mathrm{opt}}=\mathcal{L}_{\mathrm{TB}}+0.1\mathcal{L}_{\mathrm{desc}}. \]

The reward coefficient is 30. GGS pretraining runs for 3 epochs and SMILES pretraining for 5, with learning rate 0.001 and batch size 128. During optimization, the policy learning rate is 0.0005, the log-partition learning rate is 0.1, gradient norm is clipped to 10, base batch size is 64, and replay capacity is 1,024. GGS token limits are 30 for TE/MO and 18 for PH. The PH exception is a protocol-permitted length bound; these fragment-level token limits cannot be directly compared with the 140-token atomic SMILES limit as measures of molecular size.

Key Experimental Results

Main Results

The table selects Top-10 AUC (mean±standard deviation) and RI seed counts from Table 1. Each task uses five seeds and 10,000 evaluations per seed. Task-specific scoring functions and scaling constants make AUC unsuitable for comparing absolute physical difficulty across columns. SMILES methods use their respective reported settings, while the GGS molGA variant also changes genetic operators.

Method TE AUC TE RI PH AUC PH RI MO AUC MO RI
f-RAG 0.00±0.00 0/5 0.02±0.01 0/5 0.10±0.06 0/5
GenMol 0.00±0.00 0/5 0.02±0.00 0/5 0.00±0.00 0/5
molGA, SMILES 0.13±0.10 1/5 0.07±0.01 0/5 0.57±0.14 5/5
molGA, GGS+graph operators 0.68±0.05 5/5 0.35±0.13 0/5 0.36±0.04 5/5
Genetic GFN, SMILES+ZINC 0.14±0.06 0/5 0.08±0.03 1/5 1.29±0.32 5/5
Full method 0.78±0.23 5/5 0.33±0.16 4/5 0.42±0.08 5/5

The full method's Mean Top-10 Fitness is 1.19±0.42, 0.75±0.48, and 0.59±0.16 for TE/PH/MO; Mean Top-10 SA is 4.06±0.30, 3.37±0.42, and 4.40±0.37. Original Genetic GFN achieves higher MO AUC but has mean SA 4.55±0.39. A mean exceeding the boundary is compatible with RI=5/5 because RI requires a qualifying candidate, not universal qualification among the ten best molecules. Values displayed as 0.00 are rounded, not evidence that no positive score was obtained.

Ablation Study

The table follows the cumulative configurations in Appendix Table 4. TE invalid rate measures invalid samples at the final step, not the failure fraction across all oracle calls. Selected metrics avoid reproducing every reported number.

Cumulative configuration TE AUC TE invalid rate PH AUC MO AUC
Original Genetic GFN, SMILES+ZINC 0.14±0.06 0.37±0.26 0.08±0.03 1.29±0.32
+Synthetic data, still SMILES 0.22±0.07 0.32±0.17 0.10±0.04 0.62±0.31
+GGS 0.63±0.22 0.46±0.33 0.52±0.22 0.55±0.10
+Transformer 0.65±0.13 0.53±0.36 0.31±0.15 0.69±0.29
+DCD/DEX, remove KL 0.60±0.11 0.19±0.07 0.32±0.18 0.74±0.13
+Descriptors, full method 0.78±0.23 0.21±0.13 0.33±0.16 0.42±0.08

Replacing the pretraining data alone does not solve TE/PH and lowers MO AUC from 1.29 to 0.62. GGS substantially improves TE/PH, but this stage also changes representation, anchor semantics, and genetic operations, so the gain cannot be assigned to a single isolated factor. After the Transformer change, stability mechanisms together with KL removal reduce TE invalid rate from 0.53 to 0.19 without improving AUC. Descriptors subsequently raise TE AUC to 0.78 but lower MO AUC from 0.74 to 0.42.

Key Findings

  • The full method finds relevant candidates on all three tasks, but not in every seed: PH succeeds in 4/5. It also does not achieve the highest PH/MO AUC among the configurations.
  • Removing acetylene from the library lowers PH AUC from 0.33±0.16 to 0.11±0.01 and changes RI from 4/5 to 0/5; TE and MO remain at 5/5. This demonstrates the contribution of a physically important building block, which should not be obscured by claims of “unbiased” random pretraining.
  • The illustrated TE candidate reaches ZT=8.50 at 300 K with SA=3.9; the PH candidate has thermal conductance 0.0990 pW/K and SA=3.18; the MO candidate has PTB/DFT P values 9.91/8.31 and SA=4.35. These are selected physical case studies, not Mean Top-10 metrics.
  • DFT validation covers 180 high-performing MO candidates across runs, of which 78 have P>7.88. Some structures predicted at 26–29 by PTB yield only 5–12 under DFT, showing that the oracle's high-score tail cannot be accepted at face value.
  • Source ambiguities remain explicit: Table 4 marks some RI entries as “cross (1/5)” or “cross (3/5),” inconsistent with Table 1's checkmark convention; this note uses counts. Likewise, the PH claim of exceeding literature performance is not an absolute thermal-conductance record: the appendix also reports a prior candidate at 0.07 pW/K with SA=4.4. The new candidate improves synthetic accessibility and overall practicality, not the raw minimum below 0.07.

Highlights & Insights

  • Electrode anchors are optimization variables rather than auxiliary metadata. This principle transfers to inverse-design problems dominated by adsorption, attachment sites, or contact geometry without requiring the specific GGS encoding.
  • Random-molecule pretraining partly separates historical pharmaceutical distributions from optimization performance while retaining chemical assembly rules. Its value is converting hidden distributional biases into auditable fragment and assembly biases, not eliminating all priors.
  • RI separates scientific candidate discovery from leaderboard integration. High AUC, favorable mean SA, and finding a structure worth validating are different objectives; reporting them together makes the trade-offs visible.

Limitations & Future Work

  • xTB/PTB are high-throughput approximations rather than guarantees of experimental accuracy. MO frequently overestimates the P>15 tail, while TE is sensitive to level alignment and electrode modeling. Final candidates require higher-level calculations and experimental screening.
  • PH's long-molecule degeneracy requires length bounds. SA and filters cannot fully establish stability, synthesis routes, or device formation. Available building blocks do not establish that the complete molecule is synthesizable.
  • GGS branching depth, stereochemistry support, and the fixed fragment library constrain exploration. Acetylene ablations show that library knowledge affects conclusions. Explicit-anchor SMILES and open-vocabulary approaches deserve further comparison.
  • Five seeds and substantial standard deviations limit statistical confidence. Cumulative ablations combine interventions rather than fully separating representation, anchors, operators, and KL. Reported per-seed runtimes are approximately 24 hours for TE, 50 for PH, and 40 for MO on the specified 36-core CPU setting; these timings should not be transferred directly across hardware.
  • vs PMO: PMO uses inexpensive pharmaceutical proxy objectives; NMO introduces device-level physical simulations and restricts task-specific tuning. The full method's total AUC over 23 PMO tasks is 14.70±0.18, below original Genetic GFN's 15.81±0.34, supporting continued applicability rather than universal improvement.
  • vs TARTARUS: Both move beyond simple proxy oracles. NMO emphasizes independence from task-specific data and starting points while incorporating electrode structures and a shared three-task protocol.
  • vs Group SELFIES / Genetic GFN: Group SELFIES supplies the fragment-string foundation; GGS moves validity constraints into graph construction and adds native anchors. Genetic GFN supplies the hybrid policy/genetic-search framework; the paper changes representation, priors, and training stabilization rather than introducing an entirely new GFN algorithm.

Rating

  • Novelty: 4/5 — The combination of physical tasks, electrode anchors, and a strict transfer protocol is distinctive.
  • Experimental Thoroughness: 4/5 — Multiple methods, cumulative ablations, and partial DFT validation are informative, but orthogonal controls and statistical scale remain limited.
  • Writing Quality: 4/5 — The benchmark and appendix explain mechanisms clearly; RI symbols and some scientific-superiority wording require careful interpretation.
  • Value: 4/5 — A reproducible entry point into unfamiliar molecular spaces, with device-level value still awaiting downstream verification.