Behavioral Foundation Models for Quality Diversity¶
Conference: NeurIPS2026
arXiv: 2609.35615
Area: Robotics & Embodied AI
Keywords: behavioral foundation models, quality-diversity, latent-space search, backward inference, offline pretraining
TL;DR¶
BFM-QD freezes an offline-pretrained behavioral foundation model, builds a repertoire of high-quality, diverse behaviors by searching latent codes rather than policy weights, and uses closed-form backward inference for small directed mutations, substantially improving sparse navigation and contact-rich manipulation while still requiring environment rollouts and pretraining investment.
Background & Motivation¶
Quality-Diversity (QD) seeks not just the highest-return policy but a collection of behaviorally distinct policies that each perform well. Forward locomotion can use different foot-contact ratios, and cube transport can use different grasp angles. MAP-Elites discretizes behavioral descriptors into cells and retains the best policy in each cell, producing a repertoire for selection, composition, or adaptation. Traditional methods, however, perturb all neural-network weights, and most of this search space does not produce coherent behavior. Random weight perturbations rarely generate useful trajectories in tasks requiring sustained navigation or an approach–grasp–transport sequence.
CMA-ME adapts mutation directions through covariance estimation, while PGA-ME and DCRL-ME learn online critics to supply policy gradients, but they still need informative experience in difficult environments. Sparse rewards deprive the critic of positive examples; DDE-Elites, which trains a policy-weight autoencoder from existing elites, is constrained by the quality of that archive. Rather than modifying QD selection, this paper changes the object of search: a Behavioral Foundation Model (BFM), pretrained on diverse offline interactions, supplies learned dynamics and behavioral structure so that candidates start near executable skills rather than random weights.
BFMs usually infer one task code from a reward function and deploy the corresponding policy, leaving their other behaviors underused. QD can explore those behaviors, but must not move the entire archive to the same reward-optimal point. Core Idea: use the behavioral latent space of a frozen BFM as the QD backbone and mix a closed-form global task direction into each parent only in small steps, preserving diversity through behavioral-cell competition and random mutation.
Method¶
Overall Architecture¶
Inputs are offline transitions from the same environment, a downstream fitness function, behavioral descriptors, and a search budget; the output is an archive of latent codes indexed by behavioral cells, each callable through the same frozen actor. The offline phase trains the BFM. The QD phase changes only latent codes, not actor weights, repeating parent selection, variation, normalization, environment evaluation, and within-cell replacement.
The framework allows different BFM–optimizer combinations: FB-ME pairs Forward-Backward (FB) representations with Gaussian mutation, FB-CMA-ME pairs FB with CMA-ME, and TD-JEPA-CMA-ME changes the behavioral backbone. Backward Inference (BI) is an optional directed variation operator, not a separate training module required by every BFM-QD variant.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
D["Offline transitions"] -->|Offline training only| A["Frozen behavioral backbone"]
A --> B["Latent archive and<br/>spherical constraint"]
B --> C["Small-step backward inference"]
R["Past evaluation trajectories"] -->|States and rewards; closed-form inference| C
C -->|Half BI; half Gaussian| E["Environment rollout<br/>Fitness and descriptor"]
A -.->|Frozen actor executes candidates| E
E --> F["Behavioral-cell competition"]
E -->|Collect interactions| R
F -->|Retain elites; select next parents| B
Offline training and downstream environment evaluation incur different costs. BI reuses past evaluation data, but each new candidate still needs a complete environment episode. Avoiding critic training and back-propagation for BI does not eliminate offline learning or environment interaction.
Key Designs¶
1. Frozen behavioral backbone: learn executable behaviors before searching their combinations
The BFM takes the current state and a latent code and outputs an action; holding a code fixed instantiates a conditional policy that can execute over time. It does not decode a code into another independently optimized set of network weights: all candidates share an actor, and their behaviors differ through conditioning. The main experiments use 50-dimensional codes rather than the roughly 20,000-dimensional parameter-space baseline. This reduces the search and storage objects, but does not make the actor itself a network with only 50 parameters.
FB's forward representation describes future state visitation under a conditional policy, while its backward representation maps states into features that express tasks. Temporal-difference consistency of successor measures trains these representations to capture long-term dynamics, and the actor learns to act for different codes. TD-JEPA uses long-term, policy-conditioned latent prediction while retaining the interface in which changing a task code changes behavior. The contribution is primarily to reuse these interfaces for repertoire search, not to introduce a new BFM pretraining loss.
Freezing allows pretraining in one environment to be amortized over multiple reward tasks and seeds, and avoids training every candidate network during search. The trade-off is that the archive can express only behaviors supported by the frozen model. If the backbone has not learned grasping, a larger search budget alone does not guarantee a new grasping skill. Zero-shot refers to downstream reward-specific policy extraction without additional model training from an already pretrained model, not learning without pretraining.
2. Latent archive and spherical constraint: change the mutation object, but evaluate actual behavior
Each cell stores a latent code, fitness, and behavioral descriptor rather than policy parameters. Initialization samples random codes; subsequent generations select parents from occupied cells, generate offspring through Gaussian mutation or CMA-ME, and execute the frozen actor in the environment. A rollout's behavioral descriptor determines the cell index, not two coordinates of the latent code. The low-dimensional search space and the behavioral archive space are therefore distinct.
Because BFM training constrains code magnitude, mutated codes must return to the same sphere. For a nonzero candidate, normalization is:
Gaussian perturbation and linear mixing therefore cannot merely grow code magnitude into regions outside the training constraint. The sphere itself is not the source of behavioral richness. Random-code comparisons between trained and untrained FB show that architecture and geometry alone are insufficient; the useful structure comes from offline learning.
Reducing dimension is likewise insufficient. DDE-Elites also uses a 50-dimensional VAE latent space, but mainly encodes policy weights already in its archive. An archive initially lacking grasping or navigation behaviors does not automatically acquire the dynamics-informed behavioral backbone available through BFM pretraining.
3. Small-step backward inference: improve each parent with a global reward direction instead of replacing the archive
Gaussian mutation does not know which direction improves the current reward, while online policy gradients require critic training. BI samples state–reward pairs from a replay buffer containing past evaluation trajectories and projects reward onto learned features. Equation (2) and Algorithm 3 give the closed-form task code:
Here \(\rho\) is the replay-data state distribution and \(\phi\) is the backbone's learned reward-feature representation. This fits a task vector that best explains observed rewards linearly; it does not solve for a separate optimal code in each behavioral cell. The source uses a matrix inverse. Finite-sample invertibility and numerical stability need implementation treatment, but unspecified regularization is not presented here as part of the authors' algorithm.
The inferred \(z^*\) supplies a global direction, while every parent retains its own starting point:
The mixed code is projected back onto the sphere. Replacing every candidate with \(z^*\) would erase differences between parents; small-step mixing instead attempts to move distinct skills toward higher fitness. Half of each generation's offspring use BI, and half use Gaussian mutation with standard deviation 1, preventing search from contracting solely along one direction. Each inference samples 10,000 state–reward pairs from existing evaluations rather than requiring dedicated inference rollouts.
Appendix A interprets mixing as a damped Newton step on a local quadratic surrogate: reward must be approximately representable in the feature space, ideal conditional policies must correspond to task vectors, and the local Hessian is assumed strictly positive definite. The actual and surrogate objectives differ through reward projection residuals, and the Taylor expansion is local. BI should therefore not be described as globally identical to arbitrary true policy gradients or as guaranteeing improvement for every parent. After spherical projection, actual evaluation remains necessary rather than being replaced by the derivation.
4. Behavioral-cell competition: reward improvements become QD improvements only when behavioral differences remain
Evaluation identifies each offspring's behavioral cell. An empty cell accepts the candidate; an occupied cell replaces its elite only if fitness increases. Many globally high-return candidates entering the same cell still contribute only one elite. Conversely, a lower-return candidate occupying a new cell can expand the repertoire, distinguishing QD from single-objective CMA-ES.
The three evaluation metrics answer different questions. Coverage is the fraction of all cells that are occupied; QD-score sums elite fitness over occupied cells; maximum fitness measures only the best elite. QD-score combines coverage and within-cell quality, whereas maximum fitness does not measure diversity. A low QD-score does not automatically mean that no success ever occurred: interpretation requires the task, coverage, and raw values.
Reward scales differ between tasks. The paper applies per-task min-max normalization using observed minima and maxima across all methods and seeds, then averages tasks within each environment. Maximum-fitness curves receive the same normalization. A value between 0 and 1 is relative to the compared methods, not a success rate or coverage fraction, and its scale can change when methods are added.
A Worked Example¶
Consider Cube-single with cube-height descriptors and energy-efficiency fitness. The backbone first learns manipulation behaviors from offline noisy data; QD then searches for policies that maintain different planar positions and heights while reducing action expenditure. The archive stores codes whose descriptors occupy different cells, rather than training one robot policy per cell.
With the default budget of 400 candidates per generation, the BI variant allocates half to small-step backward inference and half to Gaussian mutation. One parent mixes 0.02 of the global \(z^*\) into its own code and returns to the sphere of radius \(\sqrt{50}\). It must still execute a 1,000-step episode to establish whether cube height is maintained, energy efficiency improves, and which cell should receive the candidate.
Even if the offspring saves energy, returning the cube to the table and entering another cell does not improve the original elevated cell. The replay buffer retains its trajectory for the next generation's inference, while the archive updates only from actual descriptors and fitness. Reward direction, behavioral diversity, and environment feedback cannot substitute for one another. This is a procedural illustration, not an additional reported experimental trajectory.
Loss & Training¶
Offline BFM training and QD search are separate. FB jointly learns forward/backward representations and conditional policies through successor-measure consistency; TD-JEPA learns representations and policies through temporal-difference long-term latent prediction. No new actor-training loss is introduced during QD, and BI reward fitting does not retrain a neural critic.
Default FB pretraining uses 2,000,000 optimization steps, batches of 1,024, learning rate \(10^{-4}\), discount \(\gamma=0.99\), and EMA coefficient 0.01. Walker and HalfCheetah data are collected with RND; AntMaze uses antmaze-medium-explore, and Cube uses cube-single-noisy. Reward-free pretraining does not require expert demonstrations, but does require sufficient behavioral coverage.
Search uses the stated \(50\times50\) archive, 500 generations, and 400 complete episodes per generation, totaling 200,000 evaluations. Walker/HalfCheetah use 500 steps per episode, totaling 100M transitions; AntMaze/Cube use 1,000 steps, totaling 200M transitions. The source also lists a four-dimensional Ant foot-contact descriptor, but the implementation details read here do not explain its relationship to the two-dimensional grid configuration. This remains a reproducibility question rather than a confirmed four-dimensional discretization scheme.
On one H100, Appendix Table 16 reports 2.32 GPU-hours for one-time FB pretraining and per-search times of 1.63 for FB-ME, 1.83 for FB-CMA-ME, and 1.28 for ME. A first FB-ME workflow includes pretraining plus search; reporting search time alone cannot establish lower end-to-end cost. The table also does not include the complete offline data-collection cost. Reuse amortizes training, but environment evaluations remain necessary in every QD run.
Key Experimental Results¶
Main Results¶
The four environments contain 18 tasks. The table selects raw, unnormalized QD-scores from Appendix Tables 17–20, reported as mean and standard deviation over 5 seeds. Columns have different scales: comparisons are valid within a column, not by adding values across tasks or treating them as success rates.
| Method | Walker forward velocity | HalfCheetah backward walk | AntMaze sparse Top Wall | Cube grasp style |
|---|---|---|---|---|
| MAP-Elites | 15253 ± 538 | 11874 ± 534 | 0 ± 0 | 6 ± 0 |
| PGA-ME | 13081 ± 535 | 14326 ± 473 | 0 ± 0 | 8 ± 1 |
| DCRL-ME | 14089 ± 456 | 13103 ± 479 | 0 ± 0 | 6 ± 0 |
| DDE-Elites | 13723 ± 132 | 8311 ± 125 | 14 ± 3 | 0 ± 0 |
| FB-ME | 16471 ± 175 | 11246 ± 228 | 61 ± 45 | 229 ± 103 |
| FB-CMA-ME | 15986 ± 169 | 13617 ± 201 | 611 ± 50 | 1073 ± 111 |
| FB-BI | 18908 ± 180 | 12514 ± 221 | 70 ± 48 | 255 ± 93 |
| FB-PGA-ME | 17906 ± 186 | 12520 ± 219 | 63 ± 43 | 255 ± 101 |
These values come from Appendix E.12. For AntMaze sparse Top Wall, FB-CMA-ME's 611 clearly exceeds ME's 0, but DDE-Elites' 14 is not literally zero. For Cube grasp style, FB-CMA-ME outperforms FB-BI, showing that closed-form direction does not replace every benefit of covariance-based search.
The main text's claim that BFM-QD variants consistently lead is stronger than the numerical tables support. On HalfCheetah backward walk, PGA-ME scores 14326, FB-CMA-ME 13617, and FB-ME 11246. The note follows the tables and limits the conclusion to pronounced advantages on difficult tasks, rather than claiming that every variant wins every task.
Ablation Study¶
Appendix E.8 Table 14 reports normalized QD-score over 3 seeds, comparing sampling around one inferred optimum with parent-wise BI. The alternative computes \(z^*\) from the pretraining dataset to reduce early-search data-quality effects; this differs from BI's replay-based updates each generation. The comparison is not a strict single-factor ablation changing only the geometric operator.
| Config | Walker | Cube | Note |
|---|---|---|---|
| Sample around \(z^*\), \(\sigma=0.5\) | 0.16 ± 0.03 | 0.09 ± 0.01 | Candidates concentrate near the global inferred point |
| Sample around \(z^*\), \(\sigma=1.0\) | 0.18 ± 0.02 | 0.09 ± 0.02 | More local noise still provides limited behavioral coverage |
| FB-BI | 0.92 ± 0.04 | 0.96 ± 0.05 | Each parent takes a small step while retaining its own starting point |
These normalized values are not directly comparable to the main table's raw scores, nor should they be merged with normalized aggregates from other appendix tables into one ranking. The gap supports preserving parent-specific behavioral structure, but does not mean that BI has a 96% actual success rate.
Key Findings¶
- BI and latent-space gradient updates are task-dependent: both score 255 on Cube grasp style, while Walker forward velocity gives 18908 and 17906 respectively. Overall similarity is not exact equality on every metric.
- The checkpoint study finds rapid QD-score gains in the first 40k steps and early saturation while zero-shot performance continues improving. This evidence comes from two energy-efficiency tasks, not a demonstration that all environments work without pretraining.
- A 3-seed Walker step-size sweep is stable for \(\alpha\in[0.005,0.1]\), with larger steps harming diversity. Random-code, untrained-FB, and interpolation figures support the role of the behavioral backbone, but no numerical values are inferred from curves here.
Highlights & Insights¶
- Changing the search object matters more than further repairing weight mutation. BFM offers coherent behaviors in a shared actor rather than dimension reduction alone, enabling useful candidates earlier in difficult tasks.
- A global task direction can coexist with local behavioral starting points. BI fits reward direction in shared features instead of training a separate actor for every cell, while the archive preserves distinct parents.
- Behavioral richness needed by QD differs from value precision needed by zero-shot extraction. Search can compensate for imperfect task extraction when behavioral structure emerges first, but still depends on sufficient behaviors already being present.
Limitations & Future Work¶
- The frozen backbone and offline data constrain expressivity; narrow expert trajectories or QD-collected data may be poor pretraining sources. Online backbone expansion is a possible direction, but could disrupt the stable relationship between learned representations and archive codes.
- Each BFM is environment-specific, and experiments cover simulated continuous control only. Environment generality, real-robot transfer, and behavioral novelty beyond the pretraining data remain unverified.
- Rewards and descriptors are supplied externally, and BI does not infer tasks conditioned directly on a target descriptor. Larger budgets, data-collection costs, matrix-inverse stability, and the four-dimensional descriptor's archive implementation need further clarification.
- Absolute claims of leadership in the main text exceed some table results, and normalized appendix comparisons should not be merged across tables. The checklist provides no verifiable public code link, so reproduction cannot rely on an unspecified repository.
Related Work & Insights¶
- vs MAP-Elites / CMA-ME: behavioral-cell competition and variation remain, but frozen BFM codes replace policy weights. CMA-ME can still serve as the optimizer rather than being discarded entirely.
- vs PGA-ME / DCRL-ME: BI avoids online critic training and infers reward direction through existing representations. DCRL-ME's descriptor-conditioned guidance is not a capability already supplied by BI.
- vs DDE-Elites / PoMS: these methods learn low-dimensional representations from archive policy parameters; BFM-QD first learns a behavioral backbone from offline dynamics and then searches conditional policies. The representation's training target and source of behavioral coverage matter more than latent dimension alone.
Rating¶
- Novelty: 4/5 — Extends single-task BFM extraction to repertoire search, with a meaningful closed-form small-step operator.
- Experimental Thoroughness: 4/5 — Covers 18 tasks and multiple representations, datasets, and operators, but some controls are not strictly single-factor and real-robot validation is absent.
- Writing Quality: 3/5 — The method is clear, but absolute claims in the main text do not fully match the appendix tables.
- Value: 4/5 — Provides an actionable route to diverse skill libraries through offline control-model reuse, contingent on pretraining coverage and cost amortization.