Skip to content

SafeSAE-VLA: Interpreting OpenVLA Progress Dynamics with Sparse Feature Analysis

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/socratesosorio/safesae-vla-eccv2026
Area: Multimodal VLM
Keywords: Vision-Language-Action Models, Sparse Autoencoders, Mechanistic Interpretability, Task Progress Representation, Embodied AI Monitoring

TL;DR

SafeSAE-VLA presents an interpretable sparse representation analysis pipeline for Vision-Language-Action (VLA) models, demonstrating that decomposing OpenVLA action-token residual streams with a sparse autoencoder enables top-20 sparse features alone to achieve an AUROC of 0.894 in linear progress discrimination, proving that robotic policies internally maintain compact, inspectable, and writable task-progress subspaces.

Background & Motivation

Vision-Language-Action (VLA) models have emerged as standard foundation policies for robotic manipulation, mapping visual observations and natural language instructions directly into low-level continuous action sequences. As these 7B-parameter Transformer-based architectures transition into real closed-loop deployment, robotic safety and interpretability demand answering a central mechanistic question: does the underlying policy backbone internally track temporal task progress, or does it merely generate reactive control commands based on instantaneous observations?

In embodied trajectory monitoring, conventional binary safe/unsafe or failure labels suffer from severe statistical degeneration. Across complex multi-stage manipulation environments such as LIBERO, physical collisions or minor trajectory violations occur almost universally, causing binary violation labels to collapse across rollouts and leaving minimal discriminative variance for representation probes. Furthermore, existing mechanistic interpretability efforts have predominantly investigated pure language models or static vision encoders, leaving open the question of how embodied policies encode continuous physical dynamics and operational task completion in their hidden states.

To overcome the limitations of collapsed binary classifications, this paper re-frames the analysis around relative geometric progress derived from telemetry data. Core Idea: Decompose the action-token residual stream activations of OpenVLA into high-dimensional monosemantic features via Sparse Autoencoders (SAEs), revealing that a compact subset of just 20 readable and intervention-compatible sparse directions linearly decodes task progress with near-dense predictive fidelity.

Method

Overall Architecture

SafeSAE-VLA consists of a five-stage interpretable analysis pipeline: rollout trajectory collection with activation caching, suite-normalized quartile progress relabeling, non-parametric differential feature analysis with FDR control, lightweight runtime progress monitor evaluation, and direct causal feature-setting interventions. Taking OpenVLA rollouts across LIBERO manipulation tasks as input, the pipeline produces isolated, interpretable progress directions and demonstrates their causal sensitivity in action generation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Trajectory Caching & Residual Stream Extraction<br/>Capture OpenVLA action-token hidden states"] --> B["Quartile Progress Relabeling<br/>Within-suite normalization and extreme-group split"]
    B --> C["High-Dimensional Sparse Feature Encoding<br/>Expand to 16,384 dims via layer-20 SAE"]
    C --> D["Non-Parametric Differential Feature Analysis<br/>Mann-Whitney U test and BH-FDR filtering"]
    D --> E["Lightweight Progress Monitoring & Causal Intervention<br/>Top-20 sparse probe and feature-mean injection"]

Key Designs

1. Quartile Progress Relabeling: Mitigating Scale Variance with Extreme-Group Partitioning Addressing the variance collapse observed in binary safety monitoring under continuous physical interaction, this design extracts relative end-effector displacement from simulator telemetry and calculates progress scores normalized within each task suite (goal, object, spatial, long). To prevent cross-suite distribution shift stemming from disparate physical task scales, Min-Max normalization is performed per suite. The pipeline discards the ambiguous middle 50% of episodes, designating the top 25% as high-progress (\(y=1\)) and the bottom 25% as low-progress (\(y=0\)). This clean extreme-contrast partition yields balanced, high-variance datasets for rigorous statistical representation testing.

2. High-Dimensional Sparse Decomposition of Action Tokens: Recovering Monosemantic Directions Addressing the superposition and polysemantic nature of dense Transformer hidden states where individual neuron activations couple multiple disparate concepts, this design attaches forward hooks to mid-to-late layers of OpenVLA (hidden dimension \(d=4096\)) to cache residual stream states \(h_t^{(20)}\) specifically at action token positions. A pre-trained BatchTopK sparse autoencoder (\(d_{in}=4096\), \(d_{sae}=16384\), \(k=32\)) then projects these states into sparse latent space. By constraining exactly 32 latent units to activate at any timestep, the model decomposes dense action representations into inspectable, individually attributable directional axes.

3. Dual-Filtered Non-Parametric Differential Significance Testing: Isolating Core Progress Axes To isolate genuine progress-encoding features from spurious correlations within the 16,384-dimensional space, the pipeline conducts feature-wise non-parametric Mann-Whitney U tests across high-progress and low-progress cohorts without assuming Gaussianity. To strictly control the false positive rate under mass parallel testing, the Benjamini-Hochberg false discovery rate (BH-FDR) procedure is enforced at \(\alpha=0.05\). For each significant candidate feature, the rank-biserial correlation \(r_f\) is computed alongside a composite importance score: $\(score(f) = |r_f| \cdot -\log_{10}(p_{\text{adj}, f} + \epsilon)\)$ where \(\epsilon = 10^{-300}\) prevents numerical underflow. This score balances effect magnitude and statistical confidence to rank the most critical progress-bearing directions.

4. Sparse Linear Probes and Feature-Setting Interventions: Verifying Causal Sensitivity To establish that identified sparse directions retain virtually all predictive capacity rather than acting as weak correlates, the design constructs lightweight \(L_2\)-regularized Logistic Regression monitors using only the top-20 ranked features. Crucially, to move beyond observational correlations, the system performs direct feature-setting interventions by clamping the top-20 features to their high-progress class means before projecting back into the policy residual stream. Changes are evaluated against independent dense probes and raw action distribution shifts, ruling out internal metric circularity or probe leakage.

Loss & Training

The layer-20 SAE was optimized on cached OpenVLA residual stream activations using mean squared reconstruction loss under explicit top-\(k\) sparsity: $\(\mathcal{L}_{\text{SAE}} = \|h_t^{(20)} - \hat{h}_t^{(20)}\|_2^2\)$ where the forward pass preserves only the \(k=32\) largest latent activations and sets the remaining coordinates to zero. Downstream linear progress probes were trained with \(L_2\)-penalized binary logistic regression and evaluated via stratified cross-validation.

Key Experimental Results

Main Results

The primary evaluation utilizes 374 quartile-filtered episodes (balanced 187 low-progress and 187 high-progress episodes comprising 112,200 timesteps) across 750 LIBERO rollouts to evaluate progress discrimination.

Method / Representation AUROCโ†‘ F1โ†‘ Precisionโ†‘ Recallโ†‘ PR-AUCโ†‘
Random Baseline 0.554 0.550 0.566 0.536 0.563
Action Magnitude Only 0.572 - - - -
End-Effector Velocity Only 0.711 - - - -
Magnitude + Velocity 0.572 - - - -
Top-20 Feature LR (Top-20 SAE) 0.894 0.752 0.844 0.679 0.885
Full SAE LR (16,384-d) 0.918 0.817 0.797 0.839 0.913

Suite-stratified evaluations demonstrate the stability of progress representations across varied manipulation tasks:

Task Suite Progress AUROC (Layer 20) Episodes Empirical Characteristics
goal 0.985 174 Near-perfect linear progress separability
object 0.947 142 Highly robust and stable separation
long 0.700 44 Substantially above chance despite long-horizon variance
spatial 0.667 14 Low-sample regime showing consistent positive trend

Ablation Study

Ablation experiments explore comparative probe capacities across representation spaces (pooled all-episode encoding with 2000-sample bootstrap 95% confidence intervals) and quantify direct causal intervention effects.

Probe Architecture (Layer 20) AUROC (95% CI) Interpretability & Mechanics
Raw Activation LR 0.975 (0.960โ€“0.987) Dense black-box baseline
PCA-20 0.966 (0.947โ€“0.982) Dense orthogonal compression, uninterpretable
Raw Activation MLP 0.957 (0.936โ€“0.975) Non-linear dense representation
Full SAE LR 0.947 (0.920โ€“0.970) 16,384-d inspectable sparse basis
Top-20 SAE LR 0.926 (0.900โ€“0.951) Compact, monosemantic, writable directions
Random-Projection-20 0.916 (0.883โ€“0.945) Arbitrary dense projection
Motion Telemetry 0.902 (0.870โ€“0.933) Pure external physical displacement proxy
Random-20 SAE 0.783 (0.736โ€“0.828) Marked collapse, confirming feature selectivity

In causal interventions, clamping top-20 sparse features to their high-progress class means shifts an independent raw-activation progress probe by +0.763 (95% CI [0.65, 0.87]), compared to only +0.078 for matched-random controls (permutation \(p=0.005\)). In the pre-registered goal-task-1 evaluation (\(n=64\)), intervention on 13 active features reduces the episode violation incidence from 1.00 to 0.50 (\(\Delta = -0.50\), \(p=1.5\times 10^{-8}\)).

Key Findings

  • Progress Information is Sparsely Concentrated: Although 1,117 active SAE features pass BH-FDR significance thresholds (\(\alpha=0.05\)) out of 1,881 tested, the top-20 ranked features alone preserve 0.894 (or 0.926 under pooled cross-validation) AUROC. This demonstrates that internal policy progress dynamics reside in a compact, low-dimensional linear subspace.
  • Independence from Gross Motion Magnitude: Control probes trained purely on action magnitude or end-effector velocity achieve at most 0.711 AUROC, substantially below the 0.918 AUROC of SAE features, proving that the learned representations capture structural geometric progression rather than trivial movement speeds.
  • Geometric Progress Dissociates from Semantic Completion: Auditing episode completion demonstrates that the progress-tuned top-20 features score at chance on final success classification (0.471 AUROC), while the complete SAE basis decodes success at 0.968 AUROC. This confirms that top-20 directions track relative geometric progression rather than memorizing goal completion states.

Highlights & Insights

  • Interpretability and Interventions over Predictive Dominance: While dense raw-activation classifiers reach 0.975 AUROC, the primary value of SafeSAE-VLA is decomposing this black-box capacity into inspectable, named directions that support targeted causal manipulation without retraining the underlying policy.
  • Methodological Remedy for Degenerate Safety Labels: Discarding collapsed binary violation metrics in favor of suite-normalized geometric progress quartiles provides a principled blueprint for analyzing internal representations in continuous contact-rich manipulation tasks.
  • Consistent Cross-Layer and Cross-Model Signatures: Progress signals linearly decode across layers 16, 20, and 24 (AUROC 0.871โ€“0.896). Parallel interventions on Octo-small-1.5 demonstrate directional XYZ steering (e.g., full sign inversion across 3 axes on F1383), showing shared architectural control traits across the VLA model family.

Limitations & Future Work

  • Reliance on Relative Geometric Telemetry: Progress labels currently derive from end-effector spatial displacements, which do not fully capture non-geometric milestones such as gripper state transitions or articulated object state changes. Future extensions should incorporate explicit semantic goal distances.
  • Task-Local Interventions vs. Closed-Loop Repair: Direct feature clamping alters local action statistics and lowers violation rates but does not yet achieve full closed-loop recovery of stalled manipulation policies.
  • Broader Benchmark and Hardware Validation: Experiments are confined to simulated LIBERO environments; testing representation stability on real-world bimanual systems and mobile manipulators remains a crucial next step.
  • vs. Black-Box Robot Safety Monitors (e.g., Dalal et al., Thananjeyan et al.): Traditional external safety monitors learn black-box classifiers predicting catastrophic failure. SafeSAE-VLA exposes internal policy representations, allowing mechanistic feature-level attribution and lightweight monitoring with minimal compute.
  • vs. LLM Mechanistic Interpretability (e.g., Bricken et al., Templeton et al.): Prior SAE investigations focused on linguistic or multimodal static vision concepts. SafeSAE-VLA extends dictionary learning into continuous closed-loop robotic action trajectories, demonstrating temporal progress decomposition.

Rating

  • Novelty: โญโญโญโญโญ Pioneering application of Sparse Autoencoders to analyze internal progress dynamics in Vision-Language-Action models.
  • Experimental Thoroughness: โญโญโญโญโญ Rigorous validation encompassing stratified probing, dense baseline bounds, motion controls, layer sweeps, and causal feature interventions.
  • Writing Quality: โญโญโญโญโญ Lucid technical narrative with explicit scoping distinguishing geometric progress from semantic success.
  • Value: โญโญโญโญโญ Establishes a foundational framework for white-box monitoring, neural debugging, and mechanistic steering in embodied AI policies.