Skip to content

GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning

Conference: NeurIPS 2026 (task-list assignment; the cached source is arXiv v1)
arXiv: 2609.36056
Code: https://github.com/DUAL-Xiao/GeoWind2Plan/
Area: Robotics / Urban UAV Energy Planning and 3D Wind Prediction
Keywords: urban mean wind, Fourier neural operator, wind-direction covariance, local corridor, UAV trajectory optimization

TL;DR

GeoWind2Plan converts building geometry and background wind into a mission-corridor 3D mean wind field and jointly optimizes UAV path and speed; on block A at a background speed of 4 m/s, CFD-evaluated energy decreases from \(13.3\pm1.8\) for wind-agnostic planning to \(12.5\pm1.5\) Wh/km, while wind inference for a 20% corridor takes about 3 seconds.

Background & Motivation

Path length alone does not determine the energy of low-altitude urban flight. Buildings transform the same background wind into street-canyon speed-up, sheltering, wakes, and local turning; the UAV experiences air-relative velocity rather than ground speed. Consequently, the shortest collision-free route may cross high-drag regions, while a detour or altitude change can save energy. Existing wind-aware planners often assume that a 3D wind field is already available, or replace local flow with one wind vector or a height-only profile. These inputs cannot identify which side of a building offers more favorable exposure.

High-fidelity computational fluid dynamics (CFD) can resolve building-scale wind, but cannot conveniently be run after a mission arrives. A 1.2 km block case in this paper takes about 8 hours, and each CFD solution is tied to specific inlet speed, direction, and boundary conditions; changing the wind creates a new boundary-value problem. Directly training a model over combinations of city geometry, continuous wind direction, and wind speed would also require an impractical number of 3D CFD labels. Meanwhile, one flight queries wind only near a limited set of routes, making full-city inference unnecessary for each mission.

The paper connects these difficulties: physical coordinate transformations convert direction changes into geometry changes, high-Reynolds-number mean-flow similarity approximately handles speed changes, and prediction is restricted to the region relevant to the current endpoints. Evaluation shifts from whether every grid cell matches CFD to whether predicted wind supports better flight decisions. Core Idea: rapidly construct a mission-relevant mean wind field with a local geometry-conditioned neural operator in a reference-wind frame, use a shared physical energy model for continuous 3D path and speed optimization, and evaluate the resulting route's energy benefit under offline CFD wind.

Method

Overall Architecture

Online inputs are only 3D building geometry, a horizontal background wind vector, and a start-goal pairโ€”not the current mission's CFD field. The system determines a route corridor, rotates geometry into the reference inflow direction, and predicts relevant 3D patches using a local Fourier neural operator (FNO). After overlapping predictions are blended, velocity vectors are rotated back into city coordinates and rescaled for background wind speed. Multiple rapidly-exploring random tree (RRT) paths then initialize nonlinear optimization, which produces a continuous 3D route, speed profile, and timing.

The four key designs are reference-wind adaptation, corridor-local prediction, air-relative energy, and multi-start continuous planning. The first two make wind information available within the mission budget; the latter two turn it into flight decisions. CFD supplies offline supervision and evaluation only, not an online planning input or an inference-time corrector.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Building geometry and wind<br/>Start and goal"] --> A["Reference-wind adaptation<br/>Geometry rotation"]
    A --> B["Corridor-local prediction<br/>Geometry encoding and FNO"]
    B -->|Blend then rotate vectors back and rescale| C["Air-relative energy<br/>Wind to drag and power"]
    C --> D["Multi-start continuous planning<br/>RRT and nonlinear optimization"]
    D --> O["3D route and speed<br/>Offline CFD evaluation"]
    T["Offline CFD labels"] -.->|Training supervision| B

Key Designs

1. Reference-wind adaptation: convert continuous direction changes into geometry changes in a reference frame

Training a separate model for every inflow angle would exhaust the available CFD data. The paper learns geometry-to-mean-wind prediction only at a reference direction and speed. When mission direction changes, it rotates building geometry and query locations to align inflow with the reference direction. A scalar occupancy field requires only transformed sampling coordinates, but wind is a vector field: all three velocity components must also be rotated back after prediction. Merely returning the prediction grid to its original location is insufficient.

Source Eq. (4) states yaw covariance for the ideal mean-flow operator:

\[ \mathcal{S}(\mathcal{G},U\bm{d}_{\theta})(\bm{x})=Q_{\theta}^{\top}\mathcal{S}(\mathcal{R}_{\theta}\mathcal{G},U\bm{e}_{1})(Q_{\theta}\bm{x}). \]

Here, \(\mathcal{G}\) is building occupancy, \(\mathcal{S}\) is the physical mean-flow solution operator, and \(Q_{\theta}\) aligns the mission inflow unit vector \(\bm{d}_{\theta}\) with reference direction \(\bm{e}_{1}\). Rotated geometry at \(\bm{y}\) samples the original geometry at \(Q_{\theta}^{\top}\bm{y}\). The relation requires neutral incompressible flow, a unique steady or ensemble-mean solution, and boundary conditions that rotate consistently with geometry and inflow or are yaw symmetric. It does not mean that wind stays unchanged when a city rotates, nor that the learned model automatically achieves error-free equivariance. Finite-domain boundaries and voxel interpolation can still introduce errors.

Speed adaptation has a different physical status. Source Eq. (5) assumes that, for fixed geometry, direction, and inlet-profile shape in the tested neutral high-Reynolds-number regime, mean flow normalized by inlet speed changes relatively little:

\[ \mathcal{S}(\mathcal{G},U\bm{e}_{1})(\bm{x})\approx\frac{U}{U_{\mathrm{ref}}}\mathcal{S}(\mathcal{G},U_{\mathrm{ref}}\bm{e}_{1})(\bm{x}). \]

After prediction and vector rotation back, wind velocity is multiplied by \(U/U_{\mathrm{ref}}\). This is an empirically tested mean-flow similarity approximation, not a claim that the Navierโ€“Stokes equations are linear in velocity. It also does not imply that gusts, thermal stratification, or turbulence statistics scale proportionally.

2. Corridor-local prediction: compute only relevant wind while retaining directional building information

The local FNO receives more than building presence. Occupancy identifies solids and free space, the signed distance function (SDF) describes distance to the nearest building surface, and multi-directional distance features (MDDFs) use horizontal and wind-aligned vertical rays to encode obstacle bearing, canyon openings, and height-dependent blockage. Spatial coordinates are also supplied. MDDFs are precomputed over the full domain before cropping, allowing buildings outside a patch to affect features inside it through ray distances. This is more informative for wakes and sheltering than relying entirely on within-patch occupancy, but does not establish exact recovery of all long-range fluid interactions.

Each prediction patch contains \(45\times45\times45\) grid cells and uses the same shared FNO. The network lifts feature dimensions, mixes spatial information through Fourier layers and channels through MLPs, and projects to three time-averaged velocity components. Inference uses overlapping patches and smooth blending to reduce boundary discontinuities, followed by interpolation for wind queries at continuous planning locations. The target is mean wind, not instantaneous turbulence or gusts.

The horizontal start-goal segment defines a 3D corridor extruded through the flight-altitude band. Width parameter \(r\) means total width equals \(r\) times horizontal endpoint distance; it is neither a one-sided half-width nor a strict statement that only fraction \(r\) of the city's volume is computed. After geometry rotation, only patches intersecting the rotated corridor are evaluated. Corridor experiments restrict RRT and continuous optimization to that same corridor, so computational savings also change the searchable route set. They cannot be interpreted as reducing wind computation without affecting planning.

3. Air-relative energy: let local wind affect the objective through drag, thrust, and flight duration

The planner cannot simply assign low edge costs to favorable-wind regions; it must compare integrated power for specific route-speed combinations. Source Eqs. (7)โ€“(9) compute quadratic parasite drag from nodal air-relative velocity and required thrust from force balance. Nodal power combines useful mechanical power, induced power, and blade profile power, with trajectory energy obtained by trapezoidal integration:

\[ \begin{aligned} \bm{v}_{a,i}&=\bm{v}_{i}-\bm{w}(\bm{x}_{i}),\\ \bm{D}_{i}&=-\frac{1}{2}\rho C_{d}A_{f}\|\bm{v}_{a,i}\|\bm{v}_{a,i},\qquad \bm{T}_{i}=m\bm{a}_{i}-m\bm{g}-\bm{D}_{i},\\ P_{i}&=\bm{T}_{i}^{\top}\bm{v}_{a,i}+P_{\mathrm{ind}}(\bm{T}_{i},\bm{v}_{a,i})+P_{\mathrm{prof}}(\bm{T}_{i}),\\ E(\tau;\bm{w})&=\sum_{i=0}^{N-2}\frac{1}{2}(P_{i}+P_{i+1})\Delta t_{i}. \end{aligned} \]

Here, \(\bm{v}_{i}\) is ground-relative velocity, \(\bm{g}\) points downward, and \(\rho\), \(C_d\), \(A_f\), and \(m\) are air density, drag coefficient, frontal area, and mass. Appendix E approximates induced velocity using momentum theory and scales blade profile power by the \(3/2\) power of thrust relative to hover thrust. The model includes acceleration and altitude changes but remains quasi-steady, without full attitude dynamics, motor-controller transients, or an electrochemical battery model.

This explains why minimum drag is not minimum total energy: a multirotor continues spending lift-related power throughout a longer flight. In tailwind, higher ground speed can increase average drag yet reduce total energy by shortening duration; in headwind, lower speed can exchange longer duration for lower air-relative velocity. The source permits negative useful mechanical power in some descending or strongly wind-assisted states. This modeling convention must not be interpreted as a real battery recovering the same amount of energy.

4. Multi-start continuous planning: find obstacle-avoiding routes before jointly adjusting position and speed

Buildings make energy optimization highly nonconvex. RRT generates different routes in 3D free space, using occupancy and SDF for obstacle checks, with larger steps in open areas and smaller steps near buildings. These paths are collision-free at the checking resolution, but are neither energy optimal nor guaranteed to satisfy full dynamic feasibility in advance. Each mission uses 10 RRT initializations; each path is resampled into nodes and refined by an interior-point optimizer over positions, ground-relative velocities, accelerations, and a uniform time step.

Source Eq. (11) uses trapezoidal collocation to connect position to velocity and velocity to acceleration:

\[ \bm{x}_{i+1}=\bm{x}_{i}+\frac{1}{2}(\bm{v}_{i}+\bm{v}_{i+1})\Delta t,\qquad \bm{v}_{i+1}=\bm{v}_{i}+\frac{1}{2}(\bm{a}_{i}+\bm{a}_{i+1})\Delta t. \]

Optimization fixes endpoint positions, bounds altitude, map extent, airspeed, acceleration, thrust, and time step, and imposes SDF clearance at nodes and selected intermediate segment samples. Endpoint velocities and accelerations are not fixed to zero by default. Continuous variables allow smoothing, lateral shifts, climbing, descending, and timing changes, but finite sampling is not a safety certificate for the entire continuous trajectory. Ten starts followed by local nonlinear optimization also provide no global-optimality guarantee.

The candidate with lowest predicted energy under its planning wind is selected, then wind queries are replaced with high-fidelity CFD wind for evaluation using the same discretized trajectory and time step. All compared methods share the energy model, constraints, RRT initializations, and optimizer; only planning wind changes. The paper's Ground truth setting is reference planning with CFD wind, not a unique true optimal trajectory or a proven minimum of physical battery consumption.

A Worked Example

Consider a headwind mission with default endpoint altitude 75 m and background speed 4 m/s. The endpoints define a corridor spanning the 30โ€“120 m flight band, and both corridor and building geometry are transformed into reference-wind coordinates. The network predicts only intersecting patches; blending, vector rotation back, and speed scaling create a field that the planner can query around sheltered building regions, canyons, and different heights.

Ten RRT routes supply different obstacle-avoiding starts. Each route adjusts lateral position, altitude, and speed under the same wind field and physical cost. In headwind, optimization may favor lower, slower flight with local building shelter instead of fast travel along the shortest horizontal line through an exposed canyon. This is an explanatory mechanism example, not a claim that the source published numerical trajectory results for this particular mission. Actual candidate selection and energy evaluation still follow the shared protocol above.

Loss & Training

Supervised labels are generated offline with CityFFD CFD/LES and averaged in time after convergence. Training contains approximately 20 \(1.2\times1.2\) km urban cases and one \(3\times3\) km case; evaluation blocks share no CFD patches with training. Inflow follows a power-law profile with exponent 0.15 and reference height 10 m, and simulation cells measure \(4\times4\times1.5\) m. Reference training speed is 4 m/s.

The network is supervised with velocity-field RMSE, not trained end-to-end by back-propagating planning energy. It uses 4 Fourier layers, hidden width 60, and 12 Fourier modes per spatial direction; AdamW weight decay is \(10^{-4}\), batch size 16, and initial learning rate \(10^{-2}\), halved every 20 epochs. Angular MDDF samples are compressed in the frequency domain, retaining a small number of low-frequency modes. Training takes about 7 hours on one 32 GB V100. Local sampling increases patch examples from existing CFD cases, not the number of independent city-weather simulations.

Key Experimental Results

Main Results

Evaluation includes four \(1.2\times1.2\) km blocks Aโ€“D and one \(3\times3\) km block E. Each block has 256 sampled feasible endpoint missions, with very short missions removed. Default background speed is 4 m/s, with 2 and 8 m/s additionally tested on A; the flight band is 30โ€“120 m and default endpoints are at 75 m. Relative mission-wind angle \(\psi\) defines tailwind by \(|\psi|\le45^{\circ}\), headwind by \(|\psi-180^{\circ}|\le45^{\circ}\), and crosswind otherwise.

The following subset of source Table 1 reports CFD-evaluated unit-distance energy as mean ยฑ standard deviation, in Wh/km. Parentheses retain the source's energy overhead relative to CFD-reference planning; they are not recomputed from rounded means. The profile baseline is precomputed by averaging urban-interior training CFD fields and retains only height dependence.

Block / wind speed / missions CFD reference Wind-agnostic Height-only profile GeoWind2Plan
A / 4 m/s / All \(12.3\pm1.4\) \(13.3\pm1.8\) (7.5%) \(12.8\pm1.6\) (4.1%) \(12.5\pm1.5\) (1.4%)
A / 4 m/s / Tailwind \(10.2\pm0.6\) \(11.0\pm0.4\) (8.4%) \(10.6\pm0.6\) (3.8%) \(10.2\pm0.5\) (0.6%)
A / 4 m/s / Headwind \(13.7\pm0.5\) \(15.4\pm1.1\) (12.1%) \(14.6\pm0.7\) (6.4%) \(14.0\pm0.5\) (2.1%)
A / 8 m/s / Headwind \(15.2\pm0.7\) \(21.5\pm3.2\) (41.0%) \(18.1\pm1.9\) (18.9%) \(16.4\pm1.0\) (7.6%)
E / 4 m/s / All \(12.2\pm2.1\) \(14.0\pm3.2\) (14.6%) \(12.7\pm2.3\) (3.8%) \(12.5\pm2.3\) (2.5%)

Strong headwind most clearly amplifies the value of local wind information: in A's 8 m/s headwind missions, wind-agnostic planning incurs 41.0% overhead versus 7.6% for GeoWind2Plan. The larger domain E also stays close to reference planning, but establishes scaling on that tested block rather than a guarantee at arbitrary city scales.

Ablation Study

Source Table 2 analyzes corridor width rather than removing FNO or MDDF modules. The table below preserves its ordering and reports Wh/km over all relative-angle regimes on block A. Latency is included only where the text explicitly provides wind-inference timing.

Method Corridor width CFD-evaluated energy Wind inference / analysis
CFD reference Full field \(12.3\pm1.4\) CFD case takes about 8 hours
GeoWind2Plan Full field \(12.5\pm1.5\) Neural prediction takes about 15 seconds
GeoWind2Plan 50% \(12.4\pm1.5\) Latency not separately reported
GeoWind2Plan 20% \(12.4\pm1.6\) Neural prediction takes about 3 seconds
GeoWind2Plan 10% \(12.4\pm1.9\) Failures increase with blocking tall buildings
Wind-agnostic Full field \(13.3\pm1.8\) No online wind prediction required

Similar energies do not establish equal success rates. A building taller than 120 m can span a 10% corridor and eliminate feasible routes. Table 2 gives no success rate for each width, so it cannot support a claim that narrow corridors preserve feasibility without loss.

Appendix H.4, Table 5 independently tests speed rescaling: CFD fields at different speeds are divided by their respective inlet speeds while geometry, direction, and inlet-profile shape are fixed. Relative normalized vector difference is the global \(L_2\) norm of the field difference divided by the global \(L_2\) norm of the target-speed field. Mean angular difference averages pointwise vector angles over the free-space flight envelope.

Speed comparison (4 m/s reference) Relative normalized vector difference Mean angular difference
2 vs. 4 m/s 0.0669 0.1115 rad (6.39ยฐ)
8 vs. 4 m/s 0.1124 0.1365 rad (7.82ยฐ)

Key Findings

  • 3D wind prediction supports lateral bending around buildings as well as altitude changes. A height-only profile represents vertical shear but misses horizontal differences in wake and shelter exposure.
  • Speed similarity is imperfect: normalized vector difference is greater at 8 m/s than at 2 m/s. Cross-speed results should therefore consider both wind-field discrepancy and final trajectory energy under corresponding CFD wind.
  • Appendix G separates computation stages: 20% corridor wind inference takes about 3 seconds; one RRT start plus optimization takes about 2 CPU seconds. Ten starts in parallel on 10 cores give approximately 2 seconds of planning wall time, versus about 20 CPU seconds serially. Planning time must not be presented as a 2-second complete pipeline, and wind inference cannot be omitted.

Highlights & Insights

  • Physical structure reduces label demand: converting inflow direction into reference-frame geometry avoids learning an independent mapping per angle. Transfer requires consistent coordinates, vector components, and boundariesโ€”not simply rotated input images.
  • The task region sets the compute budget: corridor inference preserves high-resolution structures near routes instead of downsampling the entire city. It also contracts the planning space, motivating adaptive widening or a fallback during deployment.
  • Decision evaluation complements field error: planning under predicted wind and scoring under CFD wind tests whether the predictor misleads the planner. This is transferable to control with physical surrogates, but the paper still uses supervised field training rather than a decision-level training loss.

Limitations & Future Work

  • Prediction uses time-averaged wind with fixed background conditions during a mission. There is no empirical support for gust handling, turbulence-risk planning, or replanning under time-varying inflow; turbulent kinetic energy, wind-input uncertainty, and online updates are concrete extensions.
  • Yaw covariance assumes a unique mean solution and consistent boundaries. Speed rescaling is limited to tested neutral, mechanically driven conditions with fixed profile shape; thermal stratification, complex terrain, and sensor errors need further validation.
  • Narrow corridors reduce computation but can exclude necessary detours; an energy table does not replace a success-rate table. The paper also provides neither a continuous safety certificate beyond finite-node checks nor a global-optimality proof.
  • CFD is a high-fidelity numerical reference, not a flight measurement. No real-flight battery-saving experiment is reported, and the quasi-steady convention permitting negative mechanical power needs consistent treatment for a non-regenerative battery cost.
  • Source ยง3.2 describes aggregate percentages as averages over โ€œfour blocks,โ€ while the overall evaluation lists five blocks Aโ€“E. That sentence does not fully clarify the aggregate scope; this note prioritizes per-block table values and does not present the aggregate as a five-block result.
  • vs Flume-FNO: the paper builds on local FNOs and directional building-distance encoding. Its emphasis is reference-wind adaptation, mission-region inference, and the CFD-evaluated UAV planning loop; the entire wind-prediction architecture should not be credited as newly introduced.
  • vs Ebert et al.'s Gappy POD wind-estimation planning: that approach uses a 2D wind slice and an empirical power curve. GeoWind2Plan uses 3D mean wind, force-based thrust modeling, and continuous optimization with variable speed and altitude.
  • vs Gu et al.'s urban-logistics planning: that work includes a detailed multirotor wind-resistance model but mainly addresses 2D steady cruise. This paper explicitly optimizes 3D nodes, acceleration, and time step jointly.
  • vs risk-oriented CAE-DNN / A* methods: those predict wind speed and turbulent kinetic energy for risk avoidance. This paper targets energy and joint route-speed optimization but currently omits turbulence safety; the objectives are complementary, and energy improvement does not establish a safety advantage.

Rating

  • Novelty: 4/5 โ€” The main contribution is the physical-adaptation, local-inference, and continuous-planning loop rather than a new FNO architecture.
  • Experimental Thoroughness: 4/5 โ€” Held-out blocks, speeds, angles, scale, and corridor analyses are included, but real flight and detailed failure rates are missing.
  • Writing Quality: 4/5 โ€” Method and appendix assumptions are clear; aggregate block counts and latency scope require careful interpretation.
  • Value: 4/5 โ€” Useful for exploiting building-resolved wind within mission deadlines, with uncertainty and safety validation still needed for deployment.