Skip to content

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

Conference: NeurIPS2026 (task archive; reviewed version: arXiv v3, 2026-09-29)
arXiv: 2605.16981
Paper: https://arxiv.org/abs/2605.16981
Area: 3D Vision
Keywords: streaming reconstruction, recurrent state, frame-level gating, long-term drift, memory retention

TL;DR

AFG uses changes in adjacent frames' internal features to generate a positive frame-level write weight in frozen CUT3R/TTT3R inference, multiplying it with the token gate to reduce redundant overwriting and lower mean KITTI ATE from 68.54 m to 43.96 m without discarding frames, retraining, or growing memory, although not every sequence or error metric improves.

Background & Motivation

Streaming 3D reconstruction must receive RGB images incrementally while producing camera poses, intrinsics, and pixel-aligned geometry. Offline global-attention models can exploit multiple frames jointly but struggle with online memory constraints on long videos. CUT3R instead compresses history into a fixed number of state tokens, using dual-stream cross-attention to read history and generate a write-back residual for each frame. Memory independent of sequence length does not guarantee stable retention of early geometry: repeated views can keep perturbing the state, eventually producing trajectory drift and fragmented point clouds.

TTT3R assigns each state token a gate derived from cross-attention logits, regulating where the current residual is written. Across five datasets, 39 sequences, and 18.15M observations, this paper finds a median gate value of 0.31 and a maximum observed value of 0.558; per-sequence means remain within 0.27โ€“0.34. On a representative ScanNet sequence, the standard deviation over all frame-token values is 0.055, compared with 0.024 for per-frame means. The gate distinguishes state slots within a frame more than it distinguishes the information value of different frames. Statistics and mechanism analysis support this structural bottleneck, but the observed values below 0.6 are not a mathematical upper bound of the sigmoid.

Classical SLAM selects keyframes to limit map growth, whereas a fixed-size recurrent state does not need frame rejection to bound its map size. What matters is how strongly each frame overwrites history. Redundant frames should still produce outputs and weak writes, with stronger writes restored at substantial scene or viewpoint changes. Core Idea: retain TTT3R's token selectivity and add frame-level write strength from changes in existing adjacent-frame features, allowing a limited state to forget redundant observations more slowly instead of updating equally strongly on every frame.

Method

Overall Architecture

The input remains an RGB video processed frame by frame, and the outputs remain the backbone's camera poses, intrinsics, and pointmaps. AFG replaces neither the encoder nor the decoder. The current frame undergoes the usual forward computation; novelty is extracted from an already available encoder global feature or decoder pose token, converted into a scalar gate, and used to scale the TTT3R residual at the final state write-back site. Beyond the fixed-size state, only the previous frame's selected feature must be retained; no frame list or KV cache grows with the video.

The signals define two separate variants: AFG-Img uses image features, and AFG-Pose uses pose-related representations. They are not an automatically switching dual-branch system, nor is there a validated fusion strategy. Variant selection is an external deployment configuration. The training node below only indicates checkpoint provenance; its dashed edge does not mean this paper performs training or uses ground-truth supervision.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    W["Previously pretrained<br/>CUT3R checkpoint"] -.->|Frozen; no retraining here| D["Current frame and old state<br/>Backbone forward computation"]
    I["Frame-by-frame RGB video"] --> D
    D --> N["Internal-feature novelty<br/>Img or Pose"]
    P["Previous selected feature"] --> N
    N --> G["Positive bounded gating"]
    G --> U["Factorized state write-back"]
    D -->|Residual and token gate| U
    D --> O["Per-frame pose and geometry"]
    U -->|Fixed state for next frame| D

Key Designs

1. Internal-feature novelty: no additional forward pass or ground-truth motion

AFG-Img spatially averages the current encoder's patch tokens into a global feature, then measures its Euclidean distance from the previous frame's global feature. CUT3R already uses this feature as its LocalMemory query, so AFG reuses existing activations rather than adding optical flow, feature matching, or a novelty predictor. The distance captures changes in visual content. It does not read the recurrent state and therefore has no direct path through which state errors change the gate.

AFG-Pose instead uses the first-row state representation in the final decoder layer that feeds pose prediction, measuring the Euclidean distance between adjacent pose tokens. This token is neither the output camera translation nor a rotation angle computed from ground truth. It is the pose-related hidden representation generated in the current backbone forward pass, available before applying the gate at write-back. Thus, although state updates are subsequently interpreted as test-time learning, there is no extra optimization loop.

The Pose signal depends jointly on the current image and old state, so it could reflect both geometry changes and model instability. The appendix checks this specifically on TUM walking_xyz: the gate's Spearman correlation is +0.50 with ground-truth adjacent-frame rotation and โˆ’0.12 with aligned pose error. This supports motion-driven gating on that sequence, but does not establish the absence of state coupling everywhere. Img avoids that coupling, while Pose may respond more sensitively to motion. This motivates complementary variants, not a guarantee of automatic motion-regime recognition.

2. Positive bounded gating: continuous write strength rather than a keyframe switch

Both variants pass the feature distance through the same sigmoid form, differing only in scale constants. Let the selected internal feature \(f_t\) denote the Img global feature or the Pose token. The source mechanism can be written in unified form as:

\[ \alpha_t=\sigma\!\left(\kappa\left(\|f_t-f_{t-1}\|_2-\tau\right)\right),\qquad (\kappa,\tau)= \begin{cases} (1,1.5),&\text{AFG-Img},\\ (3,1.0),&\text{AFG-Pose}. \end{cases} \]

\(\tau\) is the sigmoid center: a distance equal to it yields a gate of 0.5. \(\kappa\) controls response sharpness. These are hand-set constants, not learned parameters; the experiments use the same defaults without per-scene adjustment. As the distance grows, the gate approaches 1, restoring the original TTT3R write strength rather than amplifying a novel frame beyond the baseline.

Since the Euclidean distance is nonnegative, the gate at zero change remains at least \(\sigma(-\kappa\tau)\): approximately 0.18 for Img and 0.047 for Pose. For finite inputs, the sigmoid is strictly below 1; the paper uses \((0,1]\) to describe its bounded range. The positive lower bound prevents rejection through a zero frame-level gate, but does not imply a nonzero residual at every token on every frame.

This distinction makes AFG different from a frame-dropping algorithm. Every image is still encoded, decoded, and evaluated; only the state subsequently stored for the next frame changes. If adjacent-frame novelty is underestimated, the current frame still writes weakly, and later observations may continue accumulating evidence. The source does not specify how the first frame is handled when no predecessor feature exists, so no undocumented initialization rule is assumed here.

Boundedness also limits update magnitudes on the frozen checkpoint. Raising the novel-frame cap to 2 or 5 in the appendix increases mean KITTI ATE to 55.88 and 97.71 m, respectively, compared with 43.96 m at the default cap of 1. Some indoor depth and rotation metrics improve instead, making this a robust deployment choice rather than a universal claim that amplification is harmful. A single write no larger than TTT3R's does not imply a final trajectory error no larger than TTT3R's.

3. Factorized state write-back: frame strength and token-level spatial selection

TTT3R generates the token gate \(\beta_t\) by averaging cross-attention logits across layers, heads, and image tokens, scaling the residual separately for each state slot. AFG does not construct another token selector. It multiplies every slot's update from the same frame by the scalar \(\alpha_t\). Its central update is:

\[ S_t=S_{t-1}+\alpha_t\,\beta_t\odot\Delta S_t. \]

\(\Delta S_t\) is the candidate state from the current backbone forward computation minus the old state; \(\odot\) denotes elementwise scaling across state tokens. A small frame gate weakens the whole frame's influence on persistent state. A large frame gate still leaves the token gate to choose where to write. Gate computation only needs adjacent feature vectors and the fixed-dimensional state, so added storage does not grow with sequence length. Zero extra forward passes does not mean zero arithmetic overhead: averaging, distance computation, the sigmoid, and multiplication still occur, and the paper provides no separate latency measurement.

The authors rewrite the update as a convex combination of old and candidate states, interpreting memory through an exponential moving average (EMA). If candidate writes are treated as independent inputs and the effective gate is approximately constant, the explicit old-state coefficient after \(k\) steps is \((1-\alpha\beta)^k\). Under this scalarized interpretation, the lag at which it reaches \(1/e\) and its small-gate approximation are:

\[ H_{1/e}=-\frac{1}{\log(1-\alpha\beta)} \;\approx\;\frac{1}{\alpha\beta}. \]

The exact lag above is a conditional expansion of the paper's geometric-decay expression, not a memory theorem established by the authors for the complete network. Actual candidate states and gates depend on earlier states, so information may reenter the candidate through nonlinear paths. The explicit EMA coefficient cannot therefore be equated directly with the disappearance rate of information influence in the full reconstructor. In particular, \(\beta\approx0.31\) is not very small; the paper's approximately three-frame claim uses \(1/\beta\), not the exact lag.

On redundant segments, a Pose gate around 0.048 substantially reduces the approximate effective step size. The appendix measures \(\bar\beta=0.352\), reports a horizon of roughly 60 frames, and uses an independently measured drift-decay time of about 64 frames as mechanism evidence; substituting those values into the reciprocal approximation gives about 59.2 frames. This result concerns artificially repeated input, not a fixed 60-frame memory for every natural video, and certainly not indefinite retention of complete history.

A Worked Example

Consider the input in Appendix A1: the first 500 frames of TUM walking_xyz, followed by 100 copies of frame 499 with the ground-truth camera held stationary. Each repeated observation still undergoes forward computation and produces a pose; computation and evaluation points are not skipped.

TTT3R's mean token gate remains around 0.352 on this segment. Natural residual reduction alone leaves a state-update norm of roughly 0.31. Identical adjacent encoder features bring the Img frame gate to approximately 0.18. Pose representations can still vary slightly with recurrent state, but the measured gate approaches 0.048 and the update norm falls to 0.043. The protected object is the state entering the redundant segment, not a newly selected map keyframe.

The source reports approximately 9 cm of TTT3R position drift on the repeated segment and less than 0.5 cm for AFG-Pose, describing an approximately 18-fold reduction. This is a controlled experiment on one sequence, not a benchmark-wide average gain. The upper bound โ€œless than 0.5 cmโ€ also cannot by itself reproduce an exactly equal 18-fold ratio.

Loss & Training

AFG introduces no loss, training examples, learnable parameters, or fine-tuning. All experiments add TTT3R and AFG to the same frozen CUT3R checkpoint, using one NVIDIA A100 80 GB. Test-time learning here is an interpretation of state write-back, not inference-stage retraining of checkpoint weights.

Key Experimental Results

Main Results

Camera trajectories use Sim(3) Umeyama alignment. ATE is aligned absolute trajectory error, with lower values better. The main Bonn experiments use metric alignment without per-frame scale-and-shift compensation; AbsRel is absolute relative depth error, also lower-is-better. The table selects results from source Table 1 and appendix Tables 11, 12, and 14. Values from different datasets should not be ranked directly against one another.

Evaluation condition / metric TTT3R TTSA3R MeMix AFG-Img AFG-Pose
TUM, 1000 frames, ATE (m) 0.109 0.093 0.101 0.084 0.054
ScanNet, 100 frames, ATE (m) 0.072 0.063 0.075 0.071 0.084
ScanNet, 500 frames, ATE (m) 0.278 0.259 0.261 0.212 0.250
ScanNet, 1000 frames, ATE (m) 0.405 0.388 0.367 0.282 0.281
Bonn, 500 frames, AbsRel 0.0997 0.0942 0.1009 0.0970 0.0863
KITTI, mean of 11 sequences, ATE (m) 68.54 Not reported Not reported 56.60 43.96

LongStream and Keyframe-VO have mean KITTI ATE values of 51.90 and 87.00 m, respectively. Baseline values come from their original papers rather than a complete rerun in one experiment. AFG-Pose reduces the TTT3R mean by 35.9%, but sequence 04 worsens from 4.51 to 6.49 m, 05 from 31.62 to 35.16 m, 07 from 12.00 to 17.44 m, and 10 from 29.82 to 31.09 m.

Reconstruction uses sparse input sampling, taking one frame every two frames. The following table retains only the 500-frame means from appendix Table 18: Acc is reconstruction accuracy error and Comp is completeness error, both lower-is-better; NC is normal consistency, higher-is-better. Sparse sampling is an evaluation input protocol, not AFG's frame-rejection mechanism.

Dataset, 500 frames Method Acc Comp NC
7-Scenes-S TTT3R 0.065 0.031 0.550
7-Scenes-S TTSA3R 0.044 0.024 0.557
7-Scenes-S AFG-Pose 0.031 0.022 0.558
7-Scenes-S AFG-Img 0.024 0.021 0.554
NRGBD-S TTT3R 0.166 0.087 0.588
NRGBD-S TTSA3R 0.121 0.049 0.604
NRGBD-S AFG-Pose 0.101 0.033 0.611
NRGBD-S AFG-Img 0.092 0.027 0.614

Img is stronger on reconstruction errors but not optimal on every metric: its 7-Scenes NC of 0.554 is below TTSA3R's 0.557 and Pose's 0.558. VGGT runs out of memory at 300/400/500 frames under this protocol; this does not imply that every global model fails on all hardware and settings.

Ablation Study

Source Table 2 uses 1000-frame TUM pose sequences and 500-frame Bonn depth sequences. RPEr is local relative rotation error, lower-is-better. Depth \(\delta<1.25\) is the percentage of pixels whose maximum predicted-to-ground-truth depth ratio or inverse ratio is below 1.25, higher-is-better. โ€œNot measuredโ€ does not mean zero.

Config TUM ATE TUM RPEr Bonn AbsRel Bonn \(\delta<1.25\) (%)
TTT3R, no frame gate 0.109 0.443 0.100 92.1
Fixed frame gate 0.1 0.082 1.849 0.089 Not measured
Fixed frame gate 0.3 0.066 0.854 0.095 93.2
Fixed frame gate 0.7 0.093 0.429 0.100 92.3
Write once every 3 frames; evaluate every frame 0.066 0.648 0.098 92.7
Adaptive frame gate only, no token gate 0.063 0.376 0.085 95.1
AFG-Pose, frame and token gates 0.054 0.520 0.086 94.7

The adaptive gate lowers ATE from the best constant gate's 0.066 to 0.054, but does not improve every metric simultaneously. Adding the token gate further improves ATE while slightly worsening rotation and depth relative to frame-only gating. In the appendix threshold sweep, Pose with \(\tau=0.5\) achieves ATE 0.064 and RPEr 0.376, showing that the default operating point favors low long-term drift over the best local rotation accuracy.

Key Findings

  • The reported average ATE reduction on TUM at lengths of at least 600 frames is 51%; the mean Bonn AbsRel reduction across all tested lengths is 13.0%. Neither is a uniform reduction for every sequence or length.
  • Fixed reductions in writing also help, but content-dependent writing is stronger. In Appendix A4, a binary gate with a low value of 0 gives TUM ATE/RPEr of 0.061/0.880; a low value of 0.05 gives 0.052/0.647; the graded gate gives 0.054/0.520. The graded gate does not beat binary gating on every individual metric.
  • After stride-3 input subsampling, AFG-Pose becomes worse than TTT3R on TUM: ATE 0.0650 versus 0.0604, and RPEr 1.137 versus 0.934. Reducing input frames and reducing state writes are different interventions.
  • The source claims Img is best on ScanNet at \(L\geq300\), but Table 12 lists Pose at 0.281 and Img at 0.282 for 1000 frames. This small ranking conflict is retained rather than resolved by changing table values.
  • The source says increasing the threshold progressively improves ATE, but Table 4 rises from 0.051 at \(\tau=1.25\) to 0.058 at \(\tau=1.5\). The overall benefit is not strictly monotonic.

Highlights & Insights

  • Separate spatial selection from temporal strength. The token gate determines which slots receive information, while the frame gate determines how strongly the frame should write. This decomposition targets repeated observations more directly than adding further masks along the same token axis.
  • Weak writes offer another use for keyframe novelty. Novelty need not be used to reduce computation; it can protect fixed memory instead. A positive graded gate preserves some subsequent influence of underestimated observations, without guaranteeing lossless recovery.
  • Controlled repeated inputs are more explanatory than long-video curves alone. Drift under stationary ground truth shows that camera-motion difficulty is not the only cause. The probe supports the redundant-write mechanism but does not replace causal validation across scenes.

Limitations & Future Work

  • Sustained novel content can still saturate the limited state. AFG adds neither capacity nor loop closure or global map optimization. Constant memory is a storage-complexity property, not a guarantee of retaining complete historical information.
  • Default Pose gating lowers ATE but raises TUM RPEr: 0.520 versus 0.443. Several KITTI sequences also worsen, exposing costs for fine short-term rotation or particular motion regimes.
  • Adjacent feature distance is not explicit geometric novelty. Lighting changes, dynamic objects, and state errors can interfere with it. Evidence about Pose coupling comes from one sequence, motivating broader failure-case testing.
  • There is no validated automatic Img/Pose selection policy. OR, product, and weighted fusion do not strictly exceed the best individual gate, so the variants should not be described as one adaptively switching model.
  • The EMA horizon is a conditional diagnostic. The source calls it an exact property of the update rule, but full nonlinear state dependence is not covered by that approximation. The roughly 60/64-frame result in Appendix A1 does not establish a universal memory duration.
  • The source's claim that all threshold settings outperform TTT3R holds for ATE/AbsRel, not RPEr at the default and higher thresholds. This metric boundary is retained, and a single-step update cap is not treated as a final-error safety bound.
  • vs CUT3R / TTT3R: CUT3R supplies the fixed-state backbone, while TTT3R regulates residuals with a token gate. AFG uses the same frozen checkpoint and adds a frame scalar at write-back; it is not a retrained reconstruction network.
  • vs TTSA3R / MeMix: These comparisons primarily regulate token-level writing. AFG adds the frame axis, but other methods can remain better on short ScanNet, short Bonn, or some NC metrics. MeMix on TTT3R and MeMix on TTSA3R must also be distinguished.
  • vs Keyframe-VO / classical SLAM: The former learns a discrete admission/skipping policy; the latter selects keyframes for the map. AFG computes continuous weights in closed form from existing activations and processes every input. It shares a novelty signal, not map-management or computational-savings functionality.
  • vs LongStream: The source categorizes LongStream as a retrained long-sequence model. AFG has lower training cost and retains a fixed state, but its KITTI advantage concerns aligned mean ATE, not a universal win in raw scale drift or every sequence.
  • Research direction: Soft frame-level writing could be explored in other fixed-memory online models, separately measuring responses to redundancy, sustained novelty, and noisy changes. This is a proposed transfer direction, not a cross-task result established by the paper.

Rating

  • Novelty: 4/5. The frame/token factorization is a clear contribution, although the gate itself is simple and draws on keyframe ideas.
  • Experimental Thoroughness: 4/5. Six benchmarks, multiple ablations, and a controlled repetition probe provide broad coverage; a full-network memory proof and separate latency measurements are absent.
  • Writing Quality: 3/5. The main argument is clear, but exact-horizon wording, ScanNet ranking, and threshold monotonicity require explicit boundaries or conflict notes.
  • Value: 4/5. Low-cost adaptation of a frozen backbone is useful for long streams, but default rotation degradation and losses on some trajectories limit universal deployment.