Skip to content

Soft Geometric Inductive Bias for Object Centric Dynamics

Conference: NeurIPS2026
arXiv: 2512.15493
Code: https://github.com/hlinander/soft-geometric-inductive-bias
Area: Physics / Scientific Computing, Object-Centric Dynamics
Keywords: soft geometric inductive bias, Clifford algebra, geometric attention, block-causal modeling, dynamics prediction

TL;DR

The paper encodes provided object states as Clifford multivectors and learns next-timestep states with geometric-product layers and a block-causal Transformer that do not enforce exact equivariance, improving prediction and autoregressive rollouts in symmetry-breaking settings such as wall collisions, anisotropic confinement, and real driving trajectories.

Background & Motivation

Object-centric world models decompose a scene into objects and their interactions, providing a suitable representation for collisions, motion, and multi-object composition. Object positions, velocities, and orientations have natural geometric structure, so strictly equivariant networks can reduce the burden of relearning rotation or translation relationships. This advantage, however, depends on the input and task satisfying the assumed symmetry: fixed boundaries, gravity directions, material differences, and traffic rules do not always transform together with the objects.

The central example is rigid-body motion inside a box: dynamic object states are inputs, but static walls are not encoded as objects. Rotating or translating only the dynamic states does not rotate or translate the box, so the observed transition is not a strictly geometrically equivariant mapping. Enforcing equivariance limits the model's ability to learn special responses near walls, while a generic MLP or Transformer may need more data to learn basic motion and contact behavior. The paper does not discover objects from images; it asks how to retain useful geometric parameterization while accommodating non-equivariant environmental dynamics when object states are already available.

Clifford algebra offers one route: states and transformations can both be represented as multivectors, and the geometric product structures feature interactions. Instead of restricting every layer to strictly equivariant operators, the authors allow learnable multivector weights to change relationships among geometric components. Core idea: use geometric algebra to structure how a model expresses dynamics, rather than exact equivariance to restrict which dynamics it can express, and learn sequential state transitions through object attention organized into temporal blocks.

Method

Overall Architecture

The input is a sequence of states with corresponding object identities, not raw images. Each object at each timestep forms one token, with position, velocity, orientation, and other quantities embedded in separate multivector channels. After channel expansion and temporal positional encoding, the model flattens time and object dimensions into a token sequence and applies geometric-product parameterization, geometric attention, and block-causal temporal modeling to predict the next frame for all objects.

The 2D model uses \(Cl(2,0,1)\), with 8 real components per multivector; the 3D model uses \(Cl(3,0,1)\), with 16 components. A 2D sequence is organized as \((B,S,K,C,8)\), with distinct batch, time, object, and channel dimensions, and is flattened to \((B,SK,C,8)\). The 8 algebra components of an object are not 8 separate tokens: together, they form the internal representation of one object token.

Training uses ground-truth history as input and ground-truth next-frame multivectors as supervision. At inference time, predictions are fed back to generate an autoregressive rollout. Continuous-variable training error is computed in multivector space; outputs are decoded into object coordinates when displaying or evaluating trajectories. Waymo also includes activity flags, with binary supervision handled separately from continuous-state supervision.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Provided object states<br/>Training: true history"] --> B["Multivector embedding<br/>Channel expansion + time encoding"]
    B --> C["Geometric-Product<br/>Parameterization"]
    C --> D["Geometric Attention"]
    D --> E["Block-Causal<br/>Temporal Modeling"]
    E --> F["Next-frame multivectors<br/>Decoded object states"]
    F -->|Inference: rollout feedback| A
    F -.->|Training: continuous variables| L["Multivector L2<br/>Activity-flag BCE"]
    G["True next frame<br/>Training supervision only"] -.-> L

Key Designs

1. Geometric-Product Parameterization: retain structured interactions without excluding non-equivariant dynamics

A standard linear layer mixes real-valued features directly; here, weights are themselves multivectors and act on inputs through the geometric product. In 2D, position uses projective-related bivector components, velocity uses vector components, and orientation and angular velocity use rotors. These quantities occupy 4 separate multivector channels for a circular object, and shape-related information can also be included. Different physical quantities need not be collapsed into one scalar feature vector, yet can interact within a shared algebra.

The main soft model, S-CliffordTransformer, uses a one-sided geometric product, while S-Ad-CliffordTransformer uses a sandwich product. The two core mappings are retained below, with repeated channel indices summed and the star denoting the geometric product:

\[ \mathrm{S}(x)^i=W^{ij}\star x_j,\qquad \mathrm{S}_{\mathrm{Ad}}(x)^i=W^{ij}\star x_j\star(W^{ij})^r. \]

The sandwich form borrows geometric-transformation parameterization more directly, but the superscript \(r\) denotes the Clifford reverse, not the inverse; learnable weights are also not restricted to valid unit rotors. The layer therefore cannot be interpreted as always executing an exact rigid transformation. Gated sigmoid nonlinearities are used between layers, and attention and MLP sublayers retain residual connections.

โ€œSoftโ€ refers to architectural parameterization, not an additional equivariance regularization loss or a mathematical guarantee of small approximate-equivariance error. The appendix explains this through the noncommutativity of the geometric product: fixed weights generally cannot commute with the group transformation acting on the input, so neither the one-sided nor the sandwich mapping guarantees exact equivariance. The comparison model, E-CliffordTransformer, instead uses equivariant linear combinations organized by algebraic grade, imposing stronger restrictions.

This design specifically addresses fixed environments that are not explicitly modeled. The model can still learn motion relationships through geometric products, while learnable weights permit special responses to an absolute direction or wall region. The cost is losing strict group-symmetry guarantees; effectiveness must be established experimentally rather than inferred from the use of Clifford algebra alone.

The original embedding conventions contain a difference that must be preserved: the method section writes 2D orientation as \(\cos(\theta)+\sin(\theta)e_{12}\), whereas the Waymo description and algebra appendix write \(\cos(\theta/2)+\sin(\theta/2)e_{12}\). These use different full-angle and half-angle conventions without a unified explanation. Reproduction should inspect the code rather than treating either expression as the uniquely established implementation.

2. Geometric Attention: compare objects within the algebra and transmit state information through learnable projections

Multivector queries, keys, and values are produced by their corresponding linear mappings. Attention similarity does not arbitrarily take a standard dot product over all components: it uses an inner product induced by projective geometric algebra. In 2D, only the scalar, the two Euclidean vector components, and their planar bivector component contribute to similarity; projective-basis components do not contribute directly. This introduces algebraic structure into object matching rather than merely replacing the input encoding of a standard Transformer.

Absence from the inner product does not mean that position information is discarded by the entire network. The preceding learnable multivector projections and geometric products can change component combinations, while the value branch continues to carry multivector information. Because these projections can be non-equivariant, a geometric inner product alone does not make the full attention layer or network exactly equivariant.

The scaling factor in the paper's Equation (4) appears outside softmax; its placement is preserved here:

\[ \mathrm{att}(x)_i=\frac{1}{\sqrt{dC}}\mathrm{softmax}_j\!\left(Q(x)_i\cdot K(x)_j\right)V(x)_j. \]

Here, \(d\) is the algebra dimension and \(C\) is the number of multivector channels. This differs from the usual scaling of logits before softmax: the two placements affect output magnitude and the attention distribution, respectively, and are not equivalent. The cached paper does not explain the difference. This note neither moves the factor inside softmax nor guesses the implementation.

3. Block-Causal Temporal Modeling: jointly reason over same-frame objects while keeping future frames hidden

Applying a tokenwise triangular mask to flattened object sequences would prevent earlier objects within a frame from seeing later ones, making prediction depend on an arbitrary object ordering. The paper defines causal blocks by timestep rather than individual object: all objects at one timestep can attend to each other and to earlier timesteps, but not to future timesteps. Each object's output is supervised with its next-frame state, so this does not expose next-frame labels to the input.

For example, two rigid bodies in the current frame can each read the other's current position and velocity when predicting the next frame; collision reasoning need not wait for the other object to be generated sequentially. All objects in the next frame are output in parallel and then enter the next prediction together. The next-token objective therefore does not imply object-by-object autoregression. Sinusoidal temporal positional encoding identifies history order, while the block-causal mask controls visibility.

This organization also accommodates real data and synthetic physics within the same framework. A Waymo frame contains up to 32 actor tokens, 64 lane-centerline segments, and 24 miscellaneous map tokens, totaling 120 tokens. Maps provide static environmental information, while actors carry dynamic states. The 51 categorical features are placed in the scalar components of additional multivector channels, avoiding the assumption that all traffic information is a geometric vector.

The two tasks differ in environmental visibility: Kinetix walls are absent from the dynamic-object input, whereas Waymo explicitly includes lanes and maps. The paper does not ignore static environments in every dataset, and Waymo gains cannot be attributed entirely to learning hidden walls.

A Worked Example

Consider 10 polygons with walls and gravity and a 2-frame context, yielding 20 flattened object tokens. An object in the later frame can access all 20 tokens; an object in the earlier frame can access only that frame's 10 tokens. The 10 outputs associated with the later frame jointly predict the third frame, rather than predicting one object and using its prediction to generate the others.

During training, the true third-frame state is used only for supervision, and subsequent history still comes from the true trajectory. During rollout, the predicted third frame enters the context to predict the fourth frame, so small wall-response errors can accumulate. No explicit wall token is provided: the soft parameterization permits learning transitions near fixed boundaries from states and the training distribution, but does not guarantee generalization to a different box.

Loss & Training

Continuous states use an L2 loss between predicted and target multivectors, with binary cross entropy added for activity flags on relevant Waymo tokens. Waymo regression specifically targets actor position deltas and heading rotor deltas; it should not be described as predicting motion for every map token. Training uses teacher forcing and does not add noise to mitigate rollout drift, helping isolate the contribution of the geometric bias.

The main configuration has 10 blocks, 8 heads, and 24 multivector channels, with approximately 1.5M parameters. The standard Transformer retains the same attention-block and block-causal organization, adjusting embedding width to match parameter count. Comparisons thus test geometric parameterization against ordinary attention, not just Transformers against MLPs.

All models use batch size 128, AdamW, and learning rate \(5\times10^{-4}\). Weight decay is \(10^{-7}\) for Transformers and CliffordMLP, and \(3\times10^{-6}\) for the standard MLP. The authors report convergence in approximately 20 hours on one A100 for the Kinetix 10k configuration; this is a specific training setting, not a universal runtime across dataset sizes.

Key Experimental Results

Main Results

Experiments cover boxed 2D rigid bodies, 3D charged particles, and Waymo driving trajectories. Kinetix dataset generation specifies 100 or 1000 episodes of 128 frames each, giving 12,800 or 128,000 frames. Some experimental figures use the label 10k; this naming difference should be retained rather than treating every figure as having the same exact frame count. The 3D experiments use 5 particles and 3000 training trajectories of 50 frames each.

The following table comes from appendix Table 2 and reports Waymo next-step XY position RMSE with 4096 training scenarios. It is not 1.6-second rollout ADE or a unified error across datasets. Lower is better.

Agent Type S-CliffordTransformer E-CliffordTransformer Transformer
SDC (self-driving car) 0.0334 0.0708 0.1023
Non-SDC (other agents) 4.5321 4.6867 4.8162

The soft model's advantage is substantial for SDC agents, but smaller for non-SDC agents, whose absolute errors are much higher. Differences in difficulty across actor subsets cannot be ignored, and SDC results should not be generalized to every road participant.

For long-horizon prediction, the Kinetix main text evaluates up to 35 frames. With a 1-frame context, S-Ad-CliffordTransformer performs best at the longest horizon; with a 2-frame context, the standard Transformer can approach Clifford models over short rollouts, but a gap persists at longer horizons. โ€œThe soft model ranks first under every conditionโ€ is therefore not an accurate summary.

The Waymo main-text paragraph calls Figure 7 position RMSE, whereas the metrics section and caption identify it as 1.6-second ADE. This note preserves that conflict rather than merging the metrics. Figure 7 compares 4k and 16k scenarios, with the 16k points representing single-seed runs; it does not establish multi-seed validation at both scales. The cache contains no exact vertical-axis values for Figure 7, so no ADE numbers are invented.

Ablation Study

The per-variable analysis in appendix Tables 3 and 4 shows that the soft model does not win on every output. The following table uses the same 4096-scenario setting and is an output-variable analysis, not a module-removal ablation. Only entries that most clearly test the boundaries of the claim are retained.

Agent Type and Metric S-CliffordTransformer E-CliffordTransformer Transformer Lowest Error
SDC: x position RMSE 0.0163 0.0369 0.0703 S-CliffordTransformer
SDC: y velocity RMSE 0.0682 0.0651 0.0775 E-CliffordTransformer
SDC: combined velocity RMSE 0.0791 0.0777 0.1028 E-CliffordTransformer
SDC: heading RMSE (rad) 0.0083 0.0084 0.0099 S-CliffordTransformer
Non-SDC: combined velocity RMSE 0.6514 0.6603 1.0993 S-CliffordTransformer
Non-SDC: heading RMSE (rad) 0.2337 0.2270 0.2261 Transformer

A more direct mechanism test is controlled symmetry breaking in 3D. The authors apply constant gravity or position-dependent harmonic confinement along z. As confinement strength \(k_z\in\{0,0.5,1,3,10\}\) increases, E-CliffordTransformer deteriorates relative to the soft model; with constant gravity alone, they perform comparably. This evidence comes from Figure 5 and its description, without reliably extractable per-point numerical values in the cache.

Architecture ablations compare 5, 10, and 20 blocks; 4, 8, and 16 heads; and 12, 24, and 48 soft-model channels. The authors select 10 blocks, 8 heads, and width 64, settings favorable to the standard Transformer, then match the Clifford model with 24 multivector channels. This reduces concern about deliberately weak baseline settings, although matching parameters does not match FLOPs.

Key Findings

  • Wall collisions test the soft bias more strongly than free motion. The strictly equivariant model can achieve the lowest error in free motion, while soft models have the advantage at walls. The lesson is to choose constraints according to whether symmetry holds, not to reject equivariance universally.
  • Rollout RMSE compares a predicted trajectory with the true trajectory from the same initial state. Euler RMSE instead takes one true simulator step from the model's current predicted state and compares it with the next prediction. The latter separates local dynamics errors from accumulated drift; it is not another long-horizon trajectory ADE.
  • Qualitative appendix examples still exhibit excessive energy dissipation and post-collision trajectory deviations. Lower error does not imply exact conservation or indefinitely drift-free rollouts.

Highlights & Insights

  • Geometric representations can be separated from strict equivariance constraints. Multivectors and geometric products provide structure, while non-equivariant weights permit environmental exceptions; this is relevant to prediction tasks with explicit geometric states but omitted environmental variables.
  • Using frames as causal units fits interacting systems. Objects share same-frame information while retaining temporal direction, avoiding the conversion of arbitrary token order into a physical dependency.
  • Controlled symmetry breaking is more explanatory than simply adding another dataset. Varying a position-dependent confinement strength links the point at which hard constraints become a liability to observed error changes, rather than inferring the mechanism solely from driving-data performance.

Limitations & Future Work

  • Inputs assume available object states and identity correspondences; the paper does not validate object extraction from images, severe occlusion, or joint perception-and-dynamics learning. Multivector embedding advantages may not transfer directly when upstream slots lack clear geometric meaning.
  • The soft geometric bias provides no explicit equivariance-error or conservation guarantee, and long rollouts still drift. Future work could retain non-equivariant expressivity while checking contact validity or energy behavior, but these are not modules already present in the paper.
  • Comparisons mainly cover parameter-matched MLPs, Transformers, and an equivariant Clifford variant, not every physics simulator or driving-prediction method. Benefits for non-autoregressive tasks are also unproven.
  • The larger Waymo setting has a single-seed limitation, and the main text and caption disagree on the metric name. Half-angle embedding and Equation (4) scaling placement need reproduction checks. Another appendix discrepancy is that the mixed-object caption specifies 10 frames while the following text specifies 20, so that horizon should not be treated as uniformly confirmed.
  • Geometric products add computational overhead; improved data efficiency does not directly establish better runtime or fixed-FLOPs efficiency. References to fixed-budget studies of related architectures do not replace a complete computational-efficiency evaluation of this paper.
  • Relation to GATr: both use geometric-algebra representations and attention, but GATr's equivariant mappings emphasize strict group structure, while this paper relaxes projection and MLP restrictions. It also learns a reusable next-step transition model rather than only predicting a short-integration terminal state from initial conditions.
  • Relation to Residual Pathway Priors: that approach accommodates symmetry deviations through residual pathways and priors; this paper chooses structural parameterization with non-equivariant Clifford layers. Neither should be reduced to the claim that both simply add a soft-equivariance regularizer to the loss.
  • Relation to object-centric models such as SlotFormer: those approaches obtain object representations from visual content and predict dynamics, whereas this paper assumes known object states to isolate geometric dynamics. Connecting perception models is a potential next step, but stable, decodable geometric quantities in the slots would need validation.

Rating

  • Novelty: 4/5. Combines non-equivariant Clifford parameterization with object-wise temporal blocks through a clear mechanism, while reusing established components such as geometric-algebra attention.
  • Experimental Thoroughness: 4/5. Covers 2D, 3D, and real trajectories with controlled symmetry breaking; computational-budget and large-scale multi-seed evidence remain incomplete.
  • Writing Quality: 3/5. The problem and mechanism are understandable, but rotor angles, attention scaling, and some metric and horizon descriptions require checking.
  • Value: 4/5. Offers a reusable structure for approximately symmetric dynamics, primarily with available geometric object states and autoregressive prediction.