Skip to content

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

Conference: NeurIPS2026 (task-list assignment; the cached source is arXiv v1)
arXiv: 2609.38519
Paper: https://arxiv.org/abs/2609.38519
Area: Human Understanding / Egocentric Gaze Prediction
Keywords: gaze trajectories, conditional flow matching, task conditioning, temporal dependencies, egocentric video

TL;DR

GazeFlow generates whole gaze trajectories through flow matching conditioned on local video features and global task tokens, reaching F1 0.491 and average angular error 9.01 on EGTEA Gaze+ while improving displacement statistics, but its default model uses future frames and is not directly an online eye tracker.

Background & Motivation

An egocentric camera records the scene visible to its wearer but does not directly record where the eyes are looking. Inferring gaze from video could reduce eye-tracking hardware and calibration costs, supporting gaze-driven video compression, assistive understanding, and imitation learning. The challenge is not simply to locate the most salient pixel in each frame: people can choose different targets in the same scene, but their choices are constrained by task intent and the temporal structure of alternating fixations and saccades.

Framewise regression can average several plausible targets into an implausible location. Heatmap models such as GLC preserve uncertainty within a frame, but points sampled independently from successive heatmaps need not form a natural trajectory. AT and EgoM2P incorporate historical dependencies through recurrence or autoregression, although sequential generation can accumulate errors. Local saliency also cannot fully explain which object a person intends to handle next: subsequent actions can help explain the current fixation. The paper therefore adopts distribution modeling conditioned on a complete window for offline prediction.

An annotated trajectory is treated as one realization of a video-conditioned distribution rather than the only correct answer. Core idea: a temporally interacting velocity network reads both framewise visual evidence and task context extracted from the video, transporting a whole Gaussian noise trajectory into a coherent gaze trajectory rather than predicting independent heatmaps and stitching them together.

Method

Overall Architecture

The input is an egocentric video window, and the output is a complete trajectory of two-dimensional gaze coordinates, one per frame. Dual-Path Video Conditioning supplies evidence about the current scene and the activity as a whole; the Conditional Velocity Network repeatedly updates the entire noisy trajectory; Trajectory Flow Transport generates different plausible behaviors from different noise initializations. Overlapping-Window Blending supports long videos, while evaluation heatmaps approximate framewise marginals from multiple trajectory samples.

Two distinct time axes must be separated: video frame indices describe the progression of behavior, whereas flow time describes generation from noise to a trajectory. One velocity-network call handles all frames in the window; it does not predict just the next video frame. Training uses 64-frame windows, so joint modeling applies within a window rather than through a single global ODE over arbitrarily long videos.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    V["Egocentric video window"] --> C["Dual-Path Video<br/>Conditioning"]
    C -->|Framewise visual and global task tokens| N["Conditional Velocity<br/>Network"]
    X["Whole noisy trajectory<br/>and flow time"] --> N
    N -->|Whole-trajectory coordinate velocities| F["Trajectory Flow<br/>Transport"]
    F -->|Next integration step| X
    G["Annotated trajectory and random noise"] -.->|Training only: velocity regression supervision| N
    F -->|Trajectory samples from multiple windows| W["Overlapping-Window<br/>Blending"]
    W --> O["Long trajectories or sampled heatmaps"]

Key Designs

1. Dual-Path Video Conditioning: separating local visual evidence from activity context

The visual pathway uses V-JEPA 2 ViT-L. Frames resized to 256ร—256 yield a 16ร—16 grid of spatial tokens. Features of width 1024 are projected to 256, augmented with two-dimensional sinusoidal positional encoding, and processed by a two-layer framewise Transformer that organizes spatial information. This additional spatial encoder does not mix frames, but V-JEPA 2 is itself a video backbone: frame-local access to conditioning tokens does not imply a backbone without temporal context.

The task pathway spatially averages tokens into one summary per frame and then averages over time to form a global token. Because averaging can dilute a few important actions, 4 learnable queries cross-attend to all frame summaries. The resulting task bank contains the global token and 4 query outputs. It extracts selective, low-bandwidth activity cues from the same video rather than taking annotated action classes or language instructions as inputs.

This separation is useful for manipulation: spatial features locate currently visible objects, while global task cues help identify which visible object deserves attention. Action, verb, and noun labels are used only for post-training linear probes, not as required supervision for the gaze model; the Ego4D subset lacks these labels but still supports training. Probe results indicate semantic information in task tokens, not a demonstration that those tokens directly encode interpretable human psychological intent.

2. Conditional Velocity Network: combining trajectory dependencies, the current scene, and the task at every update

Each frame's noisy two-dimensional coordinates and flow time are embedded into a gaze token and processed by 6 DiT-style blocks. Each block applies gaze self-attention with one-dimensional RoPE, same-frame visual cross-attention, task cross-attention, and a feed-forward network, in that order. A final linear head outputs two-dimensional velocity per frame. Attention uses 8 heads with model width 256.

Gaze self-attention allows candidate positions across the window to constrain one another. It is the velocity network's channel for direct cross-frame interaction between gaze tokens. RoPE represents relative frame distances rather than a frame's arbitrary absolute position within a cropped window. The default network has no causal mask, so a position can consult candidate trajectory states on both sides in time.

Visual cross-attention restricts each gaze token to spatial tokens from its own frame, avoiding direct mixing of object coordinates from different frames. Task cross-attention allows every frame to access the shared task bank. Both conditions are used in every DiT block rather than concatenated only at the input, enabling repeated task queries as the candidate trajectory evolves.

The task residual has a layer-specific, channelwise sigmoid gate initialized with parameters -2, giving an initial weight of approximately 0.12; the visual residual has no such gate. Randomly initialized task queries therefore do not strongly disrupt visual prediction at the start of training. Appendix analyses find self-attention heads favoring past or future frames and visual-attention heads with different hand/object preferences. These are behavioral analyses of the model, not additional hand, object, or saccade supervision branches.

3. Trajectory Flow Transport: modeling randomness of complete paths instead of independent framewise sampling

Generation starts with standard Gaussian noise in the space of entire trajectories. The velocity network specifies where each coordinate should move at the current generation stage; numerical integration updates all frames, and its endpoint is one trajectory sample. A fresh noise draw produces another trajectory. The default configuration generates 50 trajectories in parallel, each with 50 explicit Euler steps. Video conditions are encoded once and reused throughout integration rather than recomputed by the backbone at each step.

\[ \mathbf{x}_0\sim\mathcal{N}(\mathbf{0},\mathbf{I}_{2T}),\qquad \frac{d\mathbf{x}_s}{ds}=\mathbf{v}_\theta(\mathbf{x}_s,s,\mathbf{c}),\qquad \hat{\mathbf{g}}=\mathbf{x}_1. \]

Here \(T\) is the number of video frames, \(s\) is flow time, and \(\mathbf{c}\) contains both video-conditioning pathways. Shared transport and cross-frame self-attention can represent maintaining a fixation before moving to another target, rather than independently sampling each frame and smoothing afterward. Fixation durations and physiological saccade parameters are not explicitly prescribed; the model must learn such regularities from trajectory data.

Appendix B explains that even perfectly estimated conditional framewise marginals lose temporal dependencies when multiplied together; the residual KL error is conditional total correlation. At the population level, ideal CFM velocity regression can recover the target distribution. This motivates joint modeling but does not certify recovery of the true human distribution with finite data, a finite model, or finitely many Euler steps.

Implementation also uses self-conditioning: the previous estimate of the clean trajectory is supplied as an additional input, starting from zero. During training, this estimate is provided with probability one-half, otherwise replaced by zero; it is enabled at every inference step. The network can therefore refine an existing whole-path hypothesis instead of inferring anew solely from its current noisy state.

4. Overlapping-Window Blending: extending video length without overstating global joint consistency

Long videos are split into windows of length 64 and stride 32. Each window is encoded and sampled independently, and predictions in the overlap are linearly interpolated by frame index, progressively shifting weight from the earlier window to the later one. Short clips repeat the final frame for padding, outputs are truncated to their original length, and padding positions are excluded from training loss. This addresses seams and memory use, but adjacent windows are not sampled jointly by one complete long-video model, and blending does not establish global probabilistic consistency.

When a heatmap is needed, a Gaussian kernel with standard deviation 25 pixels is placed at each of the 50 sampled points per frame, and the kernels are averaged. This converts joint generation into a marginal framewise heatmap for comparison with heatmap models such as GLC. Improved heatmap evaluation does not automatically imply the same accuracy for a single sampled trajectory. Training coordinates are normalized to \([-1,1]\) and restored to the evaluation pixel space before heatmap construction.

A Worked Example

Consider a 64-frame window showing tomato cutting. This is an illustrative walkthrough, not an additional reported experiment. Spatial tokens provide local evidence about the hand, knife, and tomato; the task bank extracts activity context from the complete sequence of frame summaries without receiving the class label โ€œcut tomato.โ€ Initially, the Gaussian trajectory has no meaningful fixation interpretation. Repeated DiT updates combine same-frame object evidence, global task information, and cross-frame position relationships to form fixation targets and transitions.

50 different noise draws can produce different plausible paths, whose averaged framewise heatmap represents their marginals. If the video extends beyond one window, the next window is generated independently and blended over a 32-frame overlap. This does not imply recovery of the true joint distribution at arbitrary lengths. If the wearer suddenly looks at an off-camera target, the input lacks evidence about that target and the model may remain focused on visible objects.

Loss & Training

The main text and Algorithm 1 use linear interpolation: an intermediate state between an annotated trajectory and random noise is sampled, and the network regresses the velocity from noise to annotation. A validity mask excludes EGTEA blink frames and padding. Valid fixation and saccade frames remain trajectory targets, so training is not restricted to fixations.

\[ \mathbf{x}_s=(1-s)\mathbf{x}_0+s\mathbf{g},\qquad \mathcal{L}_{\mathrm{CFM}}= \mathbb{E}_{\mathcal{V},s,\mathbf{x}_0,\mathbf{g}} \left[\left\|\mathbf{m}\odot\left(\mathbf{v}_\theta(\mathbf{x}_s,s,\mathbf{c})-(\mathbf{g}-\mathbf{x}_0)\right)\right\|_2^2\right]. \]

Appendix E.4 separately specifies an OT-form interpolant, \(\mathbf{x}_s=(1-(1-\sigma_{\min})s)\mathbf{x}_0+s\mathbf{g}\), with \(\sigma_{\min}=10^{-3}\), describing it as a small correction to the main text's linear interpolant. These are not literally identical paths, particularly because the endpoint retains a small noise component. They are recorded separately here: the main-text equation should not be labeled the implementation's exact complete objective. Self-conditioning is another implementation detail omitted from the simplified main algorithm.

The backbone trains only LoRA adapters on query/value projections, with rank 16, scaling coefficient 32, and dropout 0.05; other backbone weights are frozen. The projection and velocity network are trainable. AdamW uses a learning rate cosine-decayed from \(10^{-4}\) to \(10^{-6}\) over at most 80 epochs, effective batch size 16, weight decay \(10^{-4}\), and gradient clipping norm 1.0. EMA decay is 0.995 and evaluation uses EMA weights; early-stopping patience is 15.

Training uses BF16 on 4 H100 GPUs and takes approximately 3 days for EGTEA and 5 days for Ego4D. The appendix estimates approximately 12.2M trainable parameters, a 10.4M velocity network, and a 327.8M video backbone. These are approximate source counts; ignoring the frozen backbone would incorrectly suggest that the entire model has only around a dozen million parameters.

Key Experimental Results

Main Results

EGTEA Gaze+ contains cooking videos, with 8,299/2,022 training/validation clips at native resolution 640ร—480 and 24 fps. The Ego4D Aria gaze subset contains 12,178/5,202 clips at 1088ร—1080 and 30 fps. Main comparisons use the LoRA model on complete validation splits. Metrics average frames within clips and then average clips, so longer clips do not automatically receive greater weight.

The table selects core localization metrics from source Table 2. F1 is Adaptive F1 obtained by threshold sweeping, with P/R measured at the corresponding threshold. AAE is average angular error, where lower is better. The appendix first mentions a fixed threshold but later specifies sweeping; this note follows the more specific latter definition rather than describing a uniformly preset threshold.

Dataset Method AUC โ†‘ F1 โ†‘ Recall โ†‘ AAE โ†“
EGTEA Gaze+ GLC 0.953 0.421 0.531 10.21
EGTEA Gaze+ AT 0.956 0.419 0.547 9.88
EGTEA Gaze+ EgoM2P LoRA 0.921 0.293 0.515 12.81
EGTEA Gaze+ GazeFlow 0.964 0.491 0.603 9.01
Ego4D GLC 0.947 0.355 0.530 11.52
Ego4D AT 0.954 0.346 0.487 12.01
Ego4D Random Walk 0.763 0.109 0.662 14.38
Ego4D GazeFlow 0.958 0.392 0.540 10.46

Relative to GLC, the strongest F1 baseline on each dataset, F1 improves by 16.6% and 10.4%. EGTEA AAE falls from AT's 9.88 to 9.01. These results should not be summarized as winning every metric: Random Walk's Ego4D Recall of 0.662 exceeds the model's 0.540, and Appendix Table 9 reports better Ego4D KL for AT, 2.156 versus 2.213.

The next table comes from Table 3 and evaluates complete EGTEA paths and inter-frame movement. ADE averages framewise Euclidean distance; DTW measures length-normalized distance under monotone temporal alignment. Both use units of 100 pixels. Best uses ground truth to select the best of 50 samples and is not a result that can be selected at deployment without ground truth. Displacement mean/median are ratios to human reference values of 12.5/4.0 pixels; JSD compares 100-bin displacement histograms using natural logarithms.

Method ADE Mean โ†“ ADE Best โ†“ DTW Mean โ†“ DTW Best โ†“ Displacement mean ratio Displacement median ratio JSD โ†“
GLC 1.13 1.01 1.03 0.87 5.77 11.75 0.352
AT 1.32 1.12 1.22 0.99 9.70 24.55 0.476
EgoM2P LoRA 1.50 0.96 1.37 0.81 2.69 3.03 0.058
GazeFlow 1.04 0.60 0.95 0.48 0.82 1.08 0.002

Ablation Study

These results come from Table 6, using a frozen encoder on a fixed EGTEA validation subset rather than the LoRA model on the complete validation set above. ADE/DTW columns are omitted because this table does not explicitly label Mean/Best as Table 3 does, avoiding insertion of protocol-ambiguous values into the main trajectory comparison.

Config AUC โ†‘ F1 โ†‘ AAE โ†“ Displacement median ratio JSD โ†“
full model 0.970 0.493 8.81 1.08 0.0011
Without framewise visual conditioning 0.961 0.431 9.22 1.17 0.0012
Without task conditioning 0.965 0.476 9.14 1.09 0.0013
Without temporal RoPE 0.968 0.477 8.94 1.40 0.0027
Deterministic regression replaces joint generation 0.903 0.451 9.12 2.01 0.0729

Removing visual conditioning produces the largest F1 drop. Removing RoPE affects displacement distributions more clearly than AUC. The regression replacement changes both stochastic distribution modeling and trajectory coupling, so it does not isolate a claim that CFM outperforms every other joint generator. Appendix Table 15 provides a more direct objective comparison with the same network: CFM/diffusion F1 is 0.493/0.485 at 50 steps and 0.506/0.359 at 5 steps, favoring CFM for low-step sampling in this setting.

Key Findings

  • The task bank contributes beyond global averaging: Table 5 reports F1 falling from 0.493 to 0.484 without learnable queries and to 0.469 with conditioning concatenation. In linear-probe Table 4, the global token gives action Top-1 49.1%, while the union of two probe predictions reaches 54.6%. This union counts success when at least one probe is correct; it is not the accuracy of one jointly trained classifier.
  • On one H100, the default 64-frame, 50-sample, 50-step configuration takes 3.37 seconds, comprising approximately 450 milliseconds of encoding and 2.92 seconds of sampling, with peak memory 9.2 GB. Table 10's 30-sample, 5-step setting takes 0.73 seconds with F1 0.493. This improves offline throughput but still requires the complete input window.
  • In Table 12's sample-count analysis, single-sample ADE/DTW is 1.05/0.96; means over 50 samples are 1.04/0.95, while best values improve to 0.60/0.48. More samples primarily improve candidate coverage, not each individual path. Single-sample F1 differs between Tables 10 and 12, at 0.405 and 0.398. The source does not explain this discrepancy, so the sweeps are not merged into one value.
  • In causal-variant Table 14, F1 falls from 0.493 to 0.437, but AAE improves from 8.81 to 8.71; not every metric deteriorates. The filtered subset contains 1,536 paired frame-level observations, and the authors use Wilcoxon tests. Significance across correlated video frames is not evidence of stability across training random seeds.

Highlights & Insights

  • Separating spatial accuracy from human-like movement is an important evaluation choice. A model with strong framewise scores may still produce excessively abrupt paths, so gaze applications should not be assessed by heatmap F1 alone.
  • A global mean token plus a few learnable queries extracts task conditioning without action labels. This is transferable to video-conditioned behavioral trajectory generation, but semantic claims should be tested through downstream tasks and interventions rather than attention maps alone.
  • Fewer Euler steps do not necessarily reduce heatmap accuracy in this setting. Engineering evaluation should sweep both sample count and step count rather than treating the conservative main-experiment budget as the only deployment option.

Limitations & Future Work

  • The default model reads future frames and is suitable for offline estimation or scenarios allowing a complete buffered segment; the causal variant has a separate accuracy trade-off. Window throughput of 88 fps does not establish low-latency causal prediction for newly arriving frames.
  • Video conditioning fixes the observed head motion, leaving the model to estimate residual eye movement under that video rather than complete head-eye behavior across freely moving observers. Joint head and eye generation is a natural extension.
  • Off-camera saccades lack target evidence, and recorded gaze can be clamped to the image boundary. Wider fields of view, off-camera state detection, and saccade-event modeling address this more directly than merely strengthening visible-object saliency.
  • Each video has only one human trajectory. Many samples or low displacement JSD do not demonstrate a calibrated conditional distribution. Repeated observations, multiple participants, coverage, and probability-calibration evaluation are needed; Best-of-50 should also be reported alongside ground-truth-free selection or mean-sample performance.
  • Baseline protocols are not completely equivalent: AT retrains only its SP stage on Ego4D because fixation/saccade event labels are unavailable; Random Walk uses the preceding frame's ground truth, making it a privileged-history reference rather than an ordinary video-only predictor. The appendix describes some baselines being uniformly upscaled to 480ร—640 but describes heatmap conversion at native evaluation resolution elsewhere. The exact cross-resolution implementation requires code inspection rather than reconstruction from the text.
  • vs GLC: GLC predicts framewise heatmaps through global-local correlation; GazeFlow generates joint trajectories first and then estimates framewise heatmaps. The distinction is the modeled output, not a claim that GLC never accesses global video context.
  • vs AT / EgoM2P: AT expresses historical dependencies through recurrent attention transitions, while EgoM2P autoregressively emits discrete gaze tokens. GazeFlow updates a whole continuous trajectory instead of rolling out frame by frame, at the cost of repeated velocity-network computation and complete-window conditioning.
  • vs diffusion gaze generation: DiffEye generates eye movements conditioned on natural images, whereas this paper conditions trajectories on egocentric video. Its same-architecture diffusion replacement supports a low-step advantage in this setting, not a universal superiority of flow matching for all image- or video-conditioned gaze generation.
  • An extension question: treat window boundaries as an explicit probabilistic connection problem, sample across windows with overlapping-trajectory conditioning, and evaluate long-range consistency and calibration using repeated human-viewing data. This is a direction inferred from the paper's boundaries, not an experiment already performed by the authors.

Rating

  • Novelty: 4/5 โ€” Connects behavioral mechanisms, dual-path video conditioning, and joint trajectory flow matching into a clear approach; the contribution primarily concerns task modeling and architecture integration.
  • Experimental Thoroughness: 4/5 โ€” Two datasets, trajectory metrics, ablations, and cost analyses provide substantial evidence, but repeated-observation distribution validation and cross-seed error bars are absent.
  • Writing Quality: 4/5 โ€” The method has a clear narrative and detailed training appendix; interpolants, thresholds, and some table references/protocol statements require careful distinction.
  • Value: 4/5 โ€” A useful baseline for offline egocentric gaze estimation, with explicit boundaries for real-time and off-camera behavior.