Skip to content

ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding

Conference: ECCV 2026
Paper: ECCV 2026
Area: Object Detection / Video Understanding
Keywords: Reynolds Transport Theorem, Helmholtz-Hodge Decomposition, Physics-Inspired Vision, Optical Flow, Spatiotemporal Representation

TL;DR

Addressing the brightness constancy bottleneck and hefty computational costs of traditional optical flow and deep spatiotemporal models, ReynoldsFlow introduces an unsupervised, training-free representation based on the Reynolds transport theorem and Helmholtz-Hodge decomposition that decouples motion into curl-free and divergence-free components while fusing motion magnitude with appearance cues, delivering substantial accuracy gains and real-time efficiency across pose estimation, action recognition, and tiny object detection.

Background & Motivation

Video understanding serves as the core cornerstone of modern computer vision, powering vital tasks including human pose estimation, action recognition, video object detection, and multi-target tracking. For years, the community has predominantly relied on data-driven spatiotemporal deep neural architectures, ranging from 3D convolutional neural networks and recurrent neural networks to modern spatiotemporal Vision Transformers and state-space Mamba models. While demonstrating impressive empirical results on canonical benchmarks, these purely data-driven frameworks incur prohibitive computational footprints, depend heavily on heuristic spatiotemporal designs, require extensive supervised training, and provide minimal interpretability as their representations remain ungrounded in fundamental physical principles.

Conversely, classical optical flow and motion estimation methods offer mathematically grounded formulations, yet virtually all depend on the restrictive brightness constancy assumption. In realistic, unconstrained open-world videos, physical factors such as camera zoom-induced geometric divergence, non-rigid structural deformations, cast shadows, and abrupt ambient illumination variations routinely violate this assumption, inducing catastrophic tracking drift and spurious motion artifacts. Furthermore, modern deep learning optical flow networks (such as RAFT and SEA-RAFT) necessitate intensive task-specific fine-tuning and cross-domain adaptation, while the ubiquitous HSV color-coding visualization scheme causes non-linear perceptual distortion and completely obliterates high-frequency spatial appearance textures, rendering it unsuitable as direct input for fine-grained downstream perception.

This work departs from both empirical heuristic architectures and the brittle brightness constancy constraint by treating continuous video frames as fluid flows governed by continuum mechanics. Core idea: Grounded in the Reynolds transport theorem (RTT) and Helmholtz-Hodge decomposition (HHD), ReynoldsFlow decouples motion into an irrotational (curl-free) component that isolates zoom/illumination variations and a solenoidal (divergence-free) component that resolves rigid motion, constructing an unsupervised, training-free representation that couples motion magnitudes with image appearance.

Method

Overall Architecture

The ReynoldsFlow pipeline formulates spatiotemporal motion through continuum mechanics, orthogonal velocity field decomposition, divergence-compensated residual estimation, and appearance-preserving feature construction. Taking consecutive grayscale video frames as input, it derives the area Jacobian under the Reynolds transport theorem, analytically solves the curl-free (CF) velocity field via variational boundary integration on local patches, compensates the temporal divergence residual, and executes Gaussian-weighted least-squares regression to resolve the divergence-free (DF) velocity field. Finally, it couples the decomposed motion magnitudes with current frame intensity into a three-channel dynamic tensor that directly feeds off-the-shelf vision backbones.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Consecutive Video Frame Input<br/>Intensity sequence f(p, t)"] --> B["1. Spatiotemporal Modeling via Reynolds Transport Theorem<br/>Control volume integration & area Jacobian transform"]
    B --> C["2. Variational Solution of Curl-Free Field<br/>CF component isolates geometric divergence & photometric flux"]
    C --> D["3. Divergence Compensation & Divergence-Free Field Estimation<br/>Residual map construction & weighted least-squares regression"]
    D --> E["4. Motion Magnitude and Appearance Coupling<br/>Construct three-channel dynamics-aware feature tensor"]
    E --> F["Downstream Vision Applications<br/>Pose estimation / Action recognition / UAV detection"]

Key Designs

1. Spatiotemporal Modeling via Reynolds Transport Theorem: Breaking the Brightness Constancy Constraint Conventional optical flow hinges upon the total time derivative \(\frac{df}{dt} = 0\), an idealized assumption that fails under scale variations and illumination transitions. The authors model a local image patch as a time-evolving two-dimensional control volume \(\omega(t)\) where scalar function \(f(\bm{p}, t)\) designates pixel intensity. Applying the Reynolds transport theorem (RTT), the rate of change of total intensity within the moving domain expands into local temporal change plus net convective flux across the boundary: $\(\frac{d}{dt} \int_{\omega(t)} f \, dA = \int_{\omega(t)} \left( \frac{\partial f}{\partial t} + \nabla \cdot (f \bm{v}) \right) dA = \int_{\omega(t)} \left( \frac{\partial f}{\partial t} + \nabla f \cdot \bm{v} + f \nabla \cdot \bm{v} \right) dA\)$ Analyzing the differential area transformation from \(\Omega^n\) to \(\Omega^{n+1}\) under explicit Euler stepping \(\bm{p}^{n+1} \approx \bm{p}^n + \bm{v}^n \Delta t\), the differential wedge product yields \(dx^{n+1} \wedge dy^{n+1} \approx (1 + \nabla \cdot \bm{v}^n \Delta t) dx^n \wedge dy^n\). This establishes that \((1 + \nabla \cdot \bm{v}^n \Delta t)\) represents the local Jacobian determinant of the spatial domain transformation, explicitly formalizing camera zoom-in/zoom-out and object scaling dynamics.

2. Variational Solution of Curl-Free Field: Physically Isolating Dilatational and Photometric Residuals According to the Helmholtz-Hodge decomposition (HHD), any smooth velocity field decomposes uniquely into a curl-free (CF, irrotational) component \(\bm{v}_c\) and a divergence-free (DF, solenoidal) component \(\bm{v}_d\), such that \(\bm{v} = \bm{v}_c + \bm{v}_d\), where \(\nabla \times \bm{v}_c = 0\) and \(\nabla \cdot \bm{v}_d = 0\). Substituting this decomposition into the transport theorem and equating discrete Euler LHS approximation with continuous RHS formulation cancels out matching terms, isolating the condition \(\int_{\omega^n} \delta f^n \nabla \cdot \bm{v}_c^n \, dA^n = 0\), where \(\delta f^n = f^{n+1} - f^n\). Integration by parts transforms this constraint into a variational boundary-domain equilibrium: $\(\int_{\partial \omega^n} \delta f^n (\bm{v}_c^n \cdot \bm{n}) \, dS^n - \int_{\omega^n} \nabla \delta f^n \cdot \bm{v}_c^n \, dA^n = 0\)$ Assuming \(\bm{v}_c^n\) is locally invariant on a \(3 \times 3\) window \(\omega^n_{3 \times 3}\), Simpson's rule computes the boundary line integral via discrete convolutional stencils, and central difference operators discretize the domain integral. Convolving with a 2D Gaussian kernel \(G\) yields the regularized CF velocity field \(\bm{v}_c^n\). This irrotational field encapsulates non-rigid expansions, perspective zooming, and photometric flux, isolating them from pure rigid translation.

3. Divergence Compensation & Divergence-Free Field Estimation: Closed-Form Solenoidal Flow Recovery With the CF velocity field \(\bm{v}_c^n\) established, the authors equate discrete temporal differencing with Taylor expansion up to second order. Since second-order temporal acceleration is highly susceptible to high-frequency image noise, the authors apply the local intensity preservation hypothesis \(\int_{\omega^n} f^{n+1} dA^n \approx \int_{\omega^n} f^n dA^n\) and physical localization, defining the divergence-compensated residual map \(D\): $\(D \equiv - \delta f^n - \nabla f^n \cdot \bm{v}_c^n - f^n \nabla \cdot \bm{v}_c^n = \nabla f^n \cdot \bm{v}_d^n\)$ Notice that when \(\bm{v}_c^n = 0\), this formulation strictly reduces to classical optical flow \(\nabla f^n \cdot \bm{v}_d^n = -\delta f^n\). To solve for local translation \(\bm{v}_d^n = (u_d, v_d)^\top\), a Gaussian-weighted spatial least-squares problem is established: \(\min_{\bm{v}_d^n} \sum_{\bm{x} \in \omega} W(\bm{x}) (\nabla f^n \cdot \bm{v}_d^n - D)^2\), yielding the symmetric positive-definite \(2 \times 2\) closed-form system: $\(\begin{bmatrix} \sum W (f_x^n)^2 & \sum W f_x^n f_y^n \\ \sum W f_x^n f_y^n & \sum W (f_y^n)^2 \end{bmatrix} \begin{bmatrix} u_d \\ v_d \end{bmatrix} = \begin{bmatrix} \sum W f_x^n D \\ \sum W f_y^n D \end{bmatrix}\)$ Direct analytical matrix inversion yields the divergence-free vector field \(\bm{v}_d^n\), completing the total velocity field \(\bm{v}_R^n = \bm{v}_c^n + \bm{v}_d^n\).

4. Motion Magnitude and Appearance Coupling: Building Dynamics-Aware Feature Tensors Conventional optical flow visualizations map direction and magnitude non-linearly to Hue and Saturation in HSV space. This projection introduces strong chromatic discontinuity and discards high-frequency spatial appearance cues, causing tiny targets to vanish into background noise. ReynoldsFlow constructs a three-channel tensor representation: $\(\bm{F}_R^n = \left[ |\bm{v}_d^n|, \, |\bm{v}_c^n|, \, f^n \right]\)$ Here, the Red channel encodes solenoidal motion magnitude \(|\bm{v}_d^n|\), the Green channel encodes irrotational divergence/illumination magnitude \(|\bm{v}_c^n|\), and the Blue channel preserves original frame intensity \(f^n\). Assigning the CF component to the Green channel mirrors the RGGB Bayer pattern where green sensors possess double sampling density and maximal luminance sensitivity in human and machine perception. This formulation provides rich physical dynamics while preserving clear structural boundaries, functioning as an immediate drop-in replacement for downstream models.

Loss & Training

ReynoldsFlow operates in a fully unsupervised, training-free manner derived entirely from continuum mechanics and variational calculus. It requires zero backpropagation, no pre-trained weights, and no loss functions. The entire pipeline comprises local 2D convolutions, spatial derivatives, and analytical \(2 \times 2\) matrix inversions, ensuring strict \(O(HW)\) computational complexity with deterministic, real-time runtime efficiency.

Key Experimental Results

Main Results

The authors evaluated ReynoldsFlow against representative classical and deep learning optical flow methods on a motorized slider and turntable testbed across four controlled scenarios: Geometric Divergence, Varying Illumination, Horizontal Translation, and Pure Rotation.

Scenario / Method Radial Corr. \(R_{sp} \uparrow\) Ang. Cos. Sim. \(\text{ACS} \uparrow\) End. Error \(\text{AEPE} \downarrow\) Ang. Error \(\text{AAE} (^\circ) \downarrow\)
Geometric Divergence
Horn-Schunck \(0.9491 \pm 0.08\) \(0.9695 \pm 0.01\) \(0.1885 \pm 0.04\) \(10.3510 \pm 1.02\)
Farneback \(0.9452 \pm 0.11\) \(0.9965 \pm 0.00\) \(0.0635 \pm 0.01\) \(3.3412 \pm 0.29\)
SEA-RAFT (M) \(0.7625 \pm 0.25\) \(0.9825 \pm 0.01\) \(0.1652 \pm 0.02\) \(9.1250 \pm 0.68\)
ReynoldsFlow (Ours) \(\mathbf{0.9654 \pm 0.06}\) \(\mathbf{0.9975 \pm 0.00}\) \(\mathbf{0.0582 \pm 0.01}\) \(\mathbf{3.1045 \pm 0.21}\)
Varying Illumination
Horn-Schunck \(0.6251 \pm 0.12\) \(0.9999 \pm 0.00\) \(0.0026 \pm 0.00\) \(0.1385 \pm 0.05\)
Lucas-Kanade \(0.1245 \pm 0.02\) \(0.9988 \pm 0.00\) \(0.0105 \pm 0.00\) \(0.4512 \pm 0.06\)
SEA-RAFT (M) \(-0.0712 \pm 0.01\) \(0.9991 \pm 0.00\) \(0.0258 \pm 0.01\) \(1.4150 \pm 0.28\)
ReynoldsFlow (Ours) \(\mathbf{0.6582 \pm 0.14}\) \(\mathbf{0.9999 \pm 0.00}\) \(\mathbf{0.0020 \pm 0.00}\) \(\mathbf{0.1165 \pm 0.02}\)
Pure Rotation
Horn-Schunck \(0.7021 \pm 0.09\) \(0.7305 \pm 0.01\) \(1.0852 \pm 0.01\) \(42.0120 \pm 4.15\)
Lucas-Kanade \(0.4589 \pm 0.14\) \(0.9964 \pm 0.00\) \(0.1235 \pm 0.01\) \(3.7145 \pm 0.08\)
Brox \(0.8582 \pm 0.08\) \(0.9975 \pm 0.00\) \(0.1271 \pm 0.01\) \(3.4102 \pm 0.07\)
SEA-RAFT (M) \(0.0415 \pm 0.09\) \(0.9915 \pm 0.01\) \(0.1925 \pm 0.02\) \(6.1050 \pm 0.52\)
ReynoldsFlow (Ours) \(\mathbf{0.9015 \pm 0.03}\) \(\mathbf{0.9993 \pm 0.00}\) \(\mathbf{0.0615 \pm 0.01}\) \(\mathbf{1.9840 \pm 0.07}\)

Across downstream vision applications, ReynoldsFlow seamlessly integrated into SwingNet (GolfDB), C3D (HMDB51, UCF101), and YOLOv11n (Anti-UAV, ARD100, UAVDB), outperforming RGB and flow baselines.

Method / Input GolfDB (PCE \(\uparrow\)) HMDB51 (Acc \(\uparrow\)) UCF101 (Acc \(\uparrow\)) Anti-UAV (\(\text{AP}_{50}\)) Anti-UAV (\(\text{AP}_{50-95}\)) ARD100 (\(\text{AP}_{50}\)) ARD100 (\(\text{AP}_{50-95}\)) UAVDB (\(\text{AP}_{50}\)) UAVDB (\(\text{AP}_{50-95}\))
RGB 0.705 0.372 0.698 โ€” โ€” 0.554 0.304 0.811 0.518
Grayscale / Infrared 0.698 0.328 0.614 0.781 0.418 0.376 0.167 0.660 0.281
Farneback (HSV) 0.717 0.248 0.383 0.500 0.246 0.182 0.103 0.258 0.145
TV-L1 (HSV) 0.810 0.284 0.537 0.600 0.278 0.227 0.127 0.779 0.409
SEA-RAFT (M) (HSV) 0.782 0.419 0.722 0.357 0.188 0.089 0.048 0.486 0.243
DPFlow (HSV) 0.786 0.411 0.705 0.427 0.265 0.072 0.016 0.270 0.101
ReynoldsFlow (Ours) \(\mathbf{0.812}\) 0.402 0.714 \(\mathbf{0.792}\) \(\mathbf{0.446}\) \(\mathbf{0.602}\) \(\mathbf{0.326}\) \(\mathbf{0.895}\) \(\mathbf{0.547}\)

Ablation Study

The ablation investigates the impact of the visualization scheme (HSV vs magnitude-appearance coupling) and the contribution of individual decomposed components across all six benchmarks.

Representation Config GolfDB (PCE) HMDB51 (Acc) UCF101 (Acc) Anti-UAV (\(\text{AP}_{50}\)) Anti-UAV (\(\text{AP}_{50-95}\)) ARD100 (\(\text{AP}_{50}\)) ARD100 (\(\text{AP}_{50-95}\)) UAVDB (\(\text{AP}_{50}\)) UAVDB (\(\text{AP}_{50-95}\)) Note
HSV [30] 0.804 0.382 0.597 0.646 0.320 0.417 0.211 0.500 0.288 Directional chromatic coding; loses texture
$[ \bm{v}_d^n , f^n]$ 0.788 0.375 0.684 0.765 0.407 0.478 0.262 0.803
$[ \bm{v}_c^n , \bm{v}_d^n , f^n]$ 0.791 0.367 0.699 0.784 0.386 0.509
**$[ \bm{v}_d^n , \bm{v}_c^n , f^n]$ (Full Model)** \(\mathbf{0.812}\) \(\mathbf{0.402}\) \(\mathbf{0.714}\) \(\mathbf{0.792}\) \(\mathbf{0.446}\) \(\mathbf{0.602}\)

Key Findings

  • Crucial Role of Curl-Free Flux: Removing the curl-free magnitude \(|\bm{v}_c^n|\) severely hurts detection performance, causing ARD100 \(\text{AP}_{50}\) to drop by 12.4% (from 0.602 to 0.478) and UAVDB \(\text{AP}_{50}\) to degrade by 9.2%. This underscores that physically accounting for scale expansion and illumination flux is indispensable in dynamic real-world videos.
  • Sensor-Aligned Channel Assignment: Placing the CF magnitude in the Green channel yields consistent accuracy gains over placing it in the Red channel across all downstream tasks. This aligns with sensor Bayer patterns where the green channel captures double spatial samples and carries primary luminance perception.
  • Directional Coding vs. Magnitude in Tiny Object Detection: Deep optical flow models underperform on tiny UAV detection when represented in HSV (e.g. SEA-RAFT achieves only 0.089 \(\text{AP}_{50}\) on ARD100 compared to 0.554 for RGB) because directional noise obliterates minute targets. ReynoldsFlow circumvents directional artifacts and reinforces spatial texture with motion magnitude, delivering up to +14.7% \(\text{AP}_{50}\) gains.

Highlights & Insights

  • First-Principles Continuum Mechanics in Vision: By formulating frame sequences as fluid flows and harnessing the Reynolds transport theorem alongside Helmholtz-Hodge decomposition, the work delivers an elegant mathematical bridge between physical transport theory and computer vision representation learning.
  • Lightweight, Zero-Shot Plug-and-Play Utility: Requiring zero model training, backpropagation, or GPU parameter storage, ReynoldsFlow executes fast analytical matrix inversions with minimal computational overhead, acting as an off-the-shelf enhancement for existing 2D/3D networks.
  • Broad Transferability for Physics-Informed Perception: The decoupling of rigid translation from volumetric expansion and photometric changes offers immense potential for tasks beyond UAV detection, including bio-cellular tracking, robotic manipulation in specular environments, and autonomous driving in fog and rain.

Limitations & Future Work

  • Admitted Limitations: The current formulation assumes two-dimensional planar fluid projections and does not extend to 3D volumetric velocity fields; large cross-frame displacements remain constrained by local window sizes without coarse-to-fine multi-scale pyramids.
  • Underlying Assumptions & Edge Cases: Under severe motion blur, low SNR night vision, or semi-transparent reflections, spatial gradient degradation can introduce numerical instability in boundary integrals; the local spatial constancy assumption is coarse near sharp motion boundaries.
  • Potential Improvement Directions: Future research could integrate this analytical transport formulation as explicit physical loss constraints or neural operator layers (e.g., Fourier Neural Operators) inside foundational video self-supervised models.
  • vs Classical Optical Flow (Horn-Schunck / Lucas-Kanade / Farneback): Classical approaches rely on strict brightness constancy and zero divergence, collapsing under camera zoom and variable lighting; ReynoldsFlow analytically isolates dilation and illumination flux, lowering angular error to 3.10ยฐ on zoom tests.
  • vs Deep Learning Flow Models (RAFT / SEA-RAFT / DPFlow): Deep flow networks incur heavy iterative refinement costs and suffer severe accuracy drops on tiny object detection due to HSV visualization artifacts; ReynoldsFlow is training-free, computationally negligible, and doubles tiny drone detection AP.

Rating

  • Novelty: โญโญโญโญโญ Elegant integration of the Reynolds transport theorem and Helmholtz-Hodge decomposition into spatiotemporal motion representation.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive physical benchmark verification on an optical rail testbed complemented by six diverse downstream datasets.
  • Writing Quality: โญโญโญโญโญ Rigorous mathematical derivation paired with lucid physical intuition and well-structured comparative analysis.
  • Value: โญโญโญโญโญ High theoretical elegance combined with immediate practical plug-and-play engineering utility for real-time video understanding.