Skip to content

PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation

Conference: ECCV 2026
Paper: ECCV 2026
Project Page: https://www.shinjeongkim.com/pixvod/
Area: 3D Vision
Keywords: Visual Odometry / Pixel Processor Array / Gaussian Belief Propagation / Depth Estimation / Direct Tracking

TL;DR

Addressing the on-sensor computing paradigm (Pixel Processor Array), PixVOD introduces the first fully pixel-distributed direct visual odometry and depth estimation framework using Gaussian Belief Propagation (GBP) and keyframe anchoring, achieving global 6DoF consensus and metric depth recovery solely via local inter-pixel message passing.

Background & Motivation

In high-frame-rate, low-latency, and power-constrained mobile edge domains such as robotics and wearable edge systems, standard computer vision architectures adhere to the traditional separation of sensing and processing: capturing raw image pixels on an image sensor, streaming massive raw video streams across off-sensor interconnects, and executing state estimation centrally on an off-chip CPU or GPU. When scaling to high-rate streams of hundreds or thousands of frames per second, this conventional split incurs excessive interconnect bandwidth, thermal dissipation, and transmission latency. Focal-plane sensor-processors and Pixel Processor Arrays (PPAs, such as the SCAMP vision chip series) integrate local digital/analog memory registers and programmable compute logic directly within each photo-sensitive pixel, offering a transformative paradigm by computing directly on focal-plane visual signals.

However, realizing global 3D geometric state estimation directly on a distributed pixel array presents severe structural hurdles. Prior PPA-based VO/SLAM pipelines perform only low-level feature extraction or binary descriptor matching on-sensor, still requiring continuous frame-rate transmission of features to an external CPU to solve the nonlinear 6DoF bundle adjustment. Furthermore, recent in-pixel geometric estimation prototypes (e.g., PixRO and BP-SF) are either confined to 3DoF rotational tracking or depend on dedicated active depth sensors. In monocular vision, global 6DoF camera motion is an intrinsically global latent state across the entire visual field, while metric scene depth is tightly coupled with camera motion and fundamentally unscaled. Under the severe hardware constraints where each pixel possesses only local memory and nearest-neighbor communication lines, achieving camera motion consensus while reliably recovering metric scene geometry constitutes the central tension of distributed visual odometry.

This paper identifies that dense direct visual odometry is actually better suited to focal-plane array mapping than sparse feature-based approaches, because natural spatial continuity in 3D scenes provides well-behaved local priors that directly map onto local grid graph factors. Core idea: formulate monocular 6DoF tracking and dense depth estimation as an in-pixel Gaussian factor graph, synchronize global camera motion via sharded hierarchical identity factors, regularize dense depth via normal integration factors, and introduce a keyframe anchoring mechanism to maintain sufficient baseline, all solved distributively via Gaussian Belief Propagation without global memory access.

Method

Overall Architecture

PixVOD is tailored to pixel processor array hardware, completely bypassing off-sensor video transmission and centralized global memory structures. The system receives high-frame-rate monocular image frames and surface normal maps predicted locally by a lightweight in-pixel convolutional network, outputting a globally converged 6DoF relative pose estimate alongside a dense log-depth map. Each pixel processor stores an independent composite manifold state variable, while the distributed factor graph incorporates local photometric factors, motion priors, sharded identity consensus factors, and bilateral normal integration factors, updated iteratively through local Gaussian Belief Propagation message passing and belief updates.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input: Consecutive Frame Pair (Is, It) & Surface Normals"] --> B["Composite Manifold State Initialization<br/>Per-pixel Local 6DoF Pose SE(3) & Log-Depth R"]
    B --> C["Local Photometric Factor Formulation<br/>Pixel-wise Local Warping Residual Evaluation"]
    C --> D["Sharded Hierarchical Identity Propagation<br/>Quadtree Message Passing for Global 6DoF Consensus"]
    D --> E["Normal Integration Geometric Factors<br/>Bilateral Surface Normal Constraints across Neighbors"]
    E --> F["Local Keyframe Anchoring Strategy<br/>Stabilizing Effective Baseline across Multi-Target Tracking"]
    F --> G["Output: Consensual 6DoF Camera Pose & Dense Depth Map"]

Key Designs

1. Composite Manifold State and Local Photometric Factors: Fully In-Pixel Direct Tracking

Conventional direct visual odometry solves a centralized nonlinear least-squares problem \(\arg\min_\mu \sum_{p} \rho(\| I_s[p] - I_t[\mathcal{W}(p; \mu)] \|)\), which requires loading the full image into a shared memory pool. PixVOD decentralizes this formulation by allocating a stacked variable \(v_i = \mu_i \in \langle SE(3), \mathbb{R} \rangle\) locally at each pixel \(p_i\). The first six components \(\mu_{i,[0:6]} \in SE(3)\) represent the local estimate of the global camera motion, and the scalar component \(\mu_{i,[6]} \in \mathbb{R}\) represents the local log-depth (metric depth equals \(\exp(\mu_{i,[6]})\)). The direct photometric objective is decomposed into decoupled single-pixel factor terms: $\(E_D^i(\mu_i) = \frac{1}{2} \rho \left( \left\| I_s[p_i] - I_t[\mathcal{W}(p_i; \mu_i)] \right\|_{\Lambda_D}^2 \right)\)$ where the warping function \(\mathcal{W}(p_i; \mu_i) = \pi(K \exp(\mu_{i,[0:6]}) K^{-1} \pi^{-1}(p_i, \exp(\mu_{i,[6]})))\) maps coordinates into the target image frame. Each pixel processor only accesses interpolated intensities at the warped target location via short-range local inter-pixel routing, eliminating centralized image caching.

2. Sharded Hierarchical Identity Factors: Rapid 6DoF Motion Consensus

Optimizing photometric factors independently per pixel causes severe divergence in textureless or repetitive scene regions where visual flow is ill-posed. To constrain all pixel processors to agree upon a single rigid camera trajectory, PixVOD introduces pairwise identity regularisation factors: $\(E_R^{i,j}(\mu_{i,[0:6]}, \mu_{j,[0:6]}) = \frac{1}{2} \left\| \mu_{i,[0:6]} \boxminus \mu_{j,[0:6]} \right\|_{\Lambda_R}^2\)$ While purely local 4-neighbor grid connectivity matches SCAMP-like physical wires, spatial diffusion of motion estimates across an entire array requires hundreds of iterations. PixVOD adopts a sharded quadtree graph topology inspired by PixRO. Virtual aggregation nodes connect local clusters across hierarchical levels. At each higher level, noise parameters \(\sigma_R, \sigma_P\) are halved (quadrupling the information matrix precision), enabling rapid long-range consensus in logarithmically bounded local communication hops without long-distance physical interconnects.

3. Normal Integration Geometric Factors: Photometrically Guided Dense Depth

Monocular depth cannot be stably optimized from photometric residuals alone in low-texture regions. Unlike metric depth, surface normals are local geometric differentials that can be accurately estimated by compact focal-plane neural networks. Drawing upon Bilateral Normal Integration (BINI), PixVOD introduces pairwise depth-difference factors enforcing local gradient alignment with surface normals: $\(E_{N_x}^{i, j}(\mu_{i,[6]}, \mu_{j,[6]}) = \frac{1}{2} \left\| (\mu_{j,[6]} - \mu_{i,[6]}) + \frac{n_x}{\tilde{n}_z} \right\|_{\Lambda_N}^2, \quad E_{N_y}^{i, j}(\mu_{i,[6]}, \mu_{j,[6]}) = \frac{1}{2} \left\| (\mu_{j,[6]} - \mu_{i,[6]}) + \frac{n_y}{\tilde{n}_z} \right\|_{\Lambda_N}^2\)$ where \(\tilde{n}_z\) is the scaled z-component of the surface normal. These factors bridge flat surfaces with smooth curvature. Crucially, the joint coupling between photometric factors and normal factors resolves normal integration's unobservable scale and boundary-drift issues: photometric correspondence anchors metric scale across multi-view baselines and separates disconnected object boundaries.

4. Local Keyframe Anchoring Strategy: Regulating Effective Baseline against Degeneracy

On-sensor processing naturally operates at ultra-high frame rates, producing microscopic inter-frame displacements. However, naively running frame-to-frame visual odometry and integrating poses causes catastrophic geometric degeneracy: when the baseline approaches zero, translation and rotation become indistinguishable, degrading metric trajectory recovery. PixVOD introduces an anchored local keyframing protocol. A reference frame \(I_s\) is fixed as an anchor keyframe along with its normal priors, while subsequent target frames \(I_t\) update continuously across 100~300 frames. Every pixel estimates motion relative to this stable anchor, maintaining a healthy triangulation baseline (\(0.05 \sim 0.15\) m) that broadens the basin of convergence while warm-starting states across sequential arrivals.

Loss & Training

The entire nonlinear factor graph is optimized using synchronous Gaussian Belief Propagation on the composite manifold. For a state variable \(v = \bar{\mu} \diamondplus \tau\) with perturbation \(\tau \in \langle \mathfrak{se}(3), \mathbb{R} \rangle\), factors are approximated via first-order Taylor expansion using right Jacobians \(r(\bar{\mu} \diamondplus \tau) \approx \bar{r} + \bar{J}\tau\).

A Huber loss with threshold \(t_{\text{huber}} = 400\) provides robustness against photometric outliers. Leaf hyper-parameters are configured as: photometric precision \(\sigma_D = 5 \times 10^{-3}\), pose prior \(\sigma_P = 1.0\), identity factor \(\sigma_R = 4 \times 10^{-4}\), and normal factor \(\sigma_N = 10^{-3}\). Higher quadtree levels iteratively halve \(\sigma_P\) and \(\sigma_R\). The parallel algorithm is implemented in JAX on an NVIDIA RTX 4090 GPU, executing 100 synchronized message passing and belief update iterations per frame, with the final camera pose obtained by averaging all per-pixel pose means.

Key Experimental Results

Main Results

Evaluation is conducted on sequences office_{0, 1, 2, 3} from the Replica dataset at \(128 \times 128\) resolution, upsampled \(10\times\) via ScLERP trajectory interpolation to model PPA frame-rates. The table compares PixVOD against a centralized baseline (Gauss-Newton optimization solved via Iteratively Reweighted Least Squares, IRLS) across multiple frame-rate regimes:

Optimization / Frame-rate Config Communication Paradigm Memory Access Model Convergence Iterations (Per Frame) Relative Translation Error (Median) Relative Rotation Error (Median)
Centralized (IRLS + Gauss-Newton) Centralized global aggregation Shared global memory ~1,500 - 2,500 0.06 - 0.08 0.04 - 0.06
PixVOD (Framerate 10) Pixel-distributed local message passing Distributed per-pixel private memory 100 (cumul. ~4,500) 0.14 - 0.18 0.10 - 0.15
PixVOD (Framerate 5) Pixel-distributed local message passing Distributed per-pixel private memory 200 0.16 - 0.20 0.12 - 0.17
PixVOD (Framerate 0.2) Pixel-distributed local message passing Distributed per-pixel private memory 5,000 0.85 - 1.15 0.75 - 1.05

Note: While the centralized baseline converges faster with superior asymptotic accuracy, it strictly requires centralized memory access off-sensor. PixVOD achieves competitive tracking without any global bus, while framerate 0.2 diverges due to large inter-frame displacement exceeding the direct tracking convergence basin.

Ablation Study

The ablation investigates the benefits of photometric guidance over standalone normal integration (BINI) across camera baselines, and evaluates pose estimation robustness under different structural priors:

Experiment Group Geometric Prior / Configuration Translational Baseline (m) / Test Condition Absolute Depth Error / Relative Pose Error (Median) Performance Summary
Depth Reconstruction Standalone Normal Integration (BINI) [10] Baseline: 0.015 - 0.285 m Depth Abs Error: ~0.26 - 0.38 m Severe scale drift without photometric anchors
Depth Reconstruction PixVOD (BINI init + Photometric Guidance) Baseline: 0.075 - 0.165 m Depth Abs Error: ~0.06 - 0.12 m Over 60% depth error reduction, sharp boundary separation
Prior Ablation Ground-Truth Depth GT Geodesic: 0.02 - 0.38 Relative Pose Error: ~0.08 - 0.12 Upper-bound benchmark
Prior Ablation Ground-Truth Normals GT Geodesic: 0.02 - 0.38 Relative Pose Error: ~0.15 - 0.24 Primary baseline setup, robust convergence across whole range
Prior Ablation DSINE Predicted Normals [4] GT Geodesic: 0.02 - 0.20 Relative Pose Error: ~0.18 - 0.28 Practical CNN normals match GT convergence closely

Key Findings

  • Photometry anchors normal integration: Standalone BINI produces scale-unconstrained reconstructions that collapse across discontinuous scene boundaries; incorporating pixel-wise photometric factors reduces absolute depth error from \(\sim 0.32\) m to \(< 0.10\) m.
  • The optimal baseline sweet spot: High frame-rate tracking requires keyframe anchoring; when baseline is excessively narrow (\(< 0.02\) m), rotation and translation are geometrically ambiguous. Conversely, large baselines (\(> 0.25\) m) fall outside direct photometric basins. Maintaining an anchored baseline of \(0.05 \sim 0.15\) m achieves the optimal convergence basin.
  • Robustness to practical surface normal predictors: Replacing ground-truth surface normals with predictions from a real-world neural network (DSINE) causes only an \(\sim 18\%\) error increase within moderate convergence basins, demonstrating strong feasibility for future in-pixel neural inference.

Highlights & Insights

  • Shifting geometric estimation into the sensor focal plane: PixVOD demonstrates that complex 6DoF visual odometry and metric mapping do not inherently mandate centralized computing, proving that local Gaussian message passing can achieve global rigid body consensus on distributed 2D pixel grids.
  • Complementary synergy of normals and direct tracking: Surface normal prediction is naturally local and amenable to focal-plane CNN execution, while direct photometric alignment supplies multi-view triangulation and scale constraints.
  • Keyframe anchoring resolves high-frame-rate ambiguity: Rather than naively performing consecutive step-by-step differentiation at ultra-high FPS, anchoring an explicit local keyframe provides the essential baseline geometry needed to decouple translation from rotation.

Limitations & Future Work

  • GPU-based emulation rather than physical PPA tape-out: Currently validated via GPU simulation in JAX on an RTX 4090; existing SCAMP-6 chips feature limited analog/fixed-point registers incapable of directly running double-precision Lie algebra and matrix inversion.
  • Hierarchical tree wiring overhead: Sharded quadtrees accelerate information diffusion compared to flat 2D lattices, but physical routing of tree hierarchies on silicon poses hardware layout challenges.
  • Generalization to non-rigid scene flow: By relaxing the rigidity constraints enforced by the camera motion identity factors, the framework can be naturally extended to per-pixel 3D non-rigid scene flow and dynamic object tracking.
  • vs SCAMP-VO / Bit-VO [8, 32]: SCAMP-VO and Bit-VO extract FAST corners and binary descriptors on-sensor, but offload pose optimization to an external CPU; PixVOD performs fully decentralized 6DoF tracking and dense depth estimation entirely in-pixel.
  • vs PixRO [2]: PixRO pioneered GBP-based pixel-distributed tracking for 3DoF camera rotation; PixVOD expands the state representation onto \(\langle SE(3), \mathbb{R} \rangle\), resolving the tightly coupled monocular scale and translation challenge.
  • vs BINI [10]: BINI is a centralized, offline normal integration solver; PixVOD exploits its spatial sparsity within a distributed factor graph and couples it with multi-view photometry for metric scale recovery.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering the first fully pixel-distributed 6DoF direct visual odometry and dense depth framework for focal-plane sensor-processors]
  • Experimental Thoroughness: โญโญโญโญโ˜† [Rigorous validation against centralized baselines and comprehensive ablation on baselines and geometric priors, though physical chip deployment remains future work]
  • Writing Quality: โญโญโญโญโญ [Clear mathematical formulation on composite Lie groups with intuitive factor graph illustrations]
  • Value: โญโญโญโญโญ [Provides a foundational algorithmic blueprint for next-generation on-sensor intelligent perception]