CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/CubicSplat/repo
Area: Others (Vector Graphics & Differentiable Rendering / 3D Vision & Graphics)
Keywords: Differentiable rendering, Vector graphics, Gaussian splatting, Gradient conditioning, Error-bounded relaxation
TL;DR¶
CubicSplat resolves the gradient seesaw between forward geometric exactness and backward gradient conditioning in differentiable vector graphics through error-bounded forward relaxation, replacing Bézier closest-point root solvers with an \(\mathcal{O}(S^{-2})\) polyline surrogate to enable closed-form distance evaluation, achieving up to 4× faster training and over 2 dB PSNR gains on closed-fill benchmarks.
Background & Motivation¶
Vector graphics represent images as resolution-independent, compact, and directly editable parametric primitives such as Bézier curves and filled polygons. Fitting these primitives via gradient-based differentiable rasterization offers an elegant pipeline for image abstraction and generative vector design. However, classical rasterization exhibits intrinsic discontinuities at geometric boundaries, which leads to gradients that are either zero almost everywhere across pixel interiors or explosive at primitive boundaries. Early remedies like DiffVG mitigate this via area-averaged prefiltering but must differentiate through iterative cubic Bézier closest-point solvers, whereas more recent methods like Bézier Splatting place discrete 2D Gaussian splats along curves.
Despite compelling visual reconstructions, existing approaches require fragile optimization scaffolding as scene complexity scales, including ad-hoc opacity heuristics, curvature regularization losses, and complex learning rate annealing schedules. Diagnostic experiments reveal that monotonically refining forward geometric fidelity (such as tighter subdivision or denser point sampling) quickly reaches a performance plateau while dramatically inflating computational cost and destabilizing training. This counter-intuitive behavior stems from the dual role of the forward rasterizer: it must not only evaluate the objective accurately, but more importantly, serve as a well-conditioned gradient oracle capable of guiding parameter optimization through a non-convex landscape.
Current vector rasterizers suffer from a fundamental gradient seesaw: (1) Exact forward passes induce ill-conditioned backward passes (such as DiffVG encountering gradient explosions at degenerate control-point configurations due to iterative root finding); (2) Smooth gradients induce structurally distorted forward models (such as Bézier Splatting failing to commute with spatial rescaling, leading to phase-dependent artifacts and high-curvature geometric distortion under extreme magnification); (3) Adaptive subdivision triggers discontinuous computation graphs (such as de Casteljau splitting introducing discrete structural mutations that yield heteroscedastic gradient variance across primitives). The core idea is: rather than pursuing maximal forward geometric exactness at the expense of backward conditioning, CubicSplat embraces Error-Bounded Forward Relaxation, substituting iterative Bézier solvers with a uniform polyline surrogate whose geometric bias decays at \(\mathcal{O}(S^{-2})\) to establish a static, closed-form computation graph with well-conditioned gradients, while leveraging compositing-derived transmittance Gini pruning to eliminate degenerate primitives without auxiliary regularization.
Method¶
Overall Architecture¶
CubicSplat frames differentiable vector graphics optimization as principled gradient-oracle design. Given \(N\) parametric primitives \(\theta\) with fixed front-to-back layer ordering—encompassing Bézier control points, stroke half-width \(r_p\), color \(c_p\), and opacity—the forward relaxation operator \(R_S\) maps parameters to pixel values through three sequential stages: First, a uniform polyline surrogate \(\Pi_S\) discretizes each Bézier primitive into \(S\) linear segments, converting complex curve distance queries into closed-form point-to-segment Euclidean distances. Second, a signed residual \(\delta_p(x)\) unifies geometry and winding-number topology into a continuous scalar, which a smooth monotone coverage kernel \(\phi\) maps to anti-aliased per-pixel soft opacity \(\alpha_p(x)\). Finally, standard front-to-back alpha compositing synthesizes the output image while intermediate pixel transmittance values supply per-primitive visibility statistics, enabling Gini-coefficient-driven pruning and densification of occluded or degenerate primitives without external geometric penalties.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Differentiable Vector Parameters<br/>Control points, half-width, color, opacity"] --> B["Uniform Polyline Surrogate & Closed-Form Distance<br/>S-segment discretization and closed-form Euclidean distance"]
B --> C["Signed Residual & Smooth Coverage Kernel<br/>Winding-number topology and anti-aliasing transition"]
C --> D["Transmittance Visibility Statistics & Gini Pruning<br/>Alpha compositing and dynamic capacity reallocation"]
D --> E["Output: Rasterized Image & Backward Gradient Flow<br/>Static computation graph and well-conditioned oracle"]
Key Designs¶
1. Uniform Polyline Surrogate & Closed-Form Distance: Eliminating Iterative Solvers under an \(\mathcal{O}(S^{-2})\) Geometric Error Bound
To eradicate the severe gradient ill-conditioning caused by differentiating through iterative cubic Bézier closest-point solvers at degenerate geometries in analytical rasterizers, CubicSplat constructs a uniform polyline surrogate \(\mathcal{P}_S = \Pi_S(\mathcal{B})\). Each Bézier segment over \(t \in [0, 1]\) is discretized into \(S\) uniform line segments, reducing the Euclidean distance query from any pixel center \(x\) to the curve to a closed-form minimum distance computation:
This formulation eliminates numerical loops and dynamic branching, preserving a static computation graph throughout back-propagation. Drawing on non-smooth analysis and spline approximation theory, the Hausdorff distance discrepancy between the true Bézier curve \(\mathcal{B}\) and the surrogate \(\Pi_S(\mathcal{B})\) is strictly bounded by \(\mathcal{O}(S^{-2})\) under mild smoothness conditions. Crucially, empirical optimization exhibits a broad "solution-equivariance" regime: once the sampling density surpasses a modest geometric threshold \(S_0\) (\(S=24\) for open strokes, \(S=12\) for closed fills), further increasing \(S\) produces negligible variation in the converged parameters. Thus, \(S\) acts as a forward-relaxation knob for gradient conditioning rather than an approximation bottleneck.
2. Signed Residual & Smooth Coverage Kernel: Decoupling Geometric Distance from Winding-Number Derivatives
To provide unified, differentiable support for both open strokes and closed fills without allowing discrete topological decisions to corrupt the gradient flow, CubicSplat formulates a unified continuous signed residual:
where \(r_p \ge 0\) is the learnable stroke half-width. For closed fills, \(s_p(x) \in \{-1, +1\}\) is determined by a discrete nonzero winding-number test on the polyline boundary. A smooth, monotone coverage kernel \(\phi(\delta_p(x) / \tau_p)\) converts the normalized residual into soft opacity \(\alpha_p(x)\), where \(\tau_p > 0\) defines the anti-aliasing transition width. The decisive design mechanism is gradient routing decoupling: the winding sign \(s_p(x)\) is treated as constant with respect to parameter updates (\(\nabla_\theta s_p(x) = 0\)), directing all spatial gradients through the continuous distance field \(d(x, \mathcal{P}_p)\). Because changes in \(s_p(x)\) correspond to codimension-1 boundary crossing events in parameter space, the smooth kernel \(\phi\) already supplies well-conditioned boundary motion gradients, preventing topological boundary jitter during closed-shape optimization.
3. Transmittance Visibility Statistics & Gini Pruning: Internal Capacity Allocation Replacing Explicit Geometric Regularization
As the number of vector primitives scales, severe occlusion often produces degenerate or redundant curves (e.g., collapsed micro-loops or fully occluded paths) that corrupt the optimization landscape. Prior works enforce ad-hoc convexity, non-self-intersection, or curvature penalties, which severely restrict the expressive representation of complex natural contours. CubicSplat avoids explicit geometric regularization entirely, harvesting intrinsic physical signals exposed during front-to-back alpha compositing (with initial transmittance \(T_0(x) = 1\)):
The effective primitive contribution \(T_{i-1}(x)\alpha_i(x)\) is aggregated across all pixels into an importance weight \(w_p\) at zero additional compute cost. When the Gini coefficient computed over the importance distribution \(\{w_p\}\) exceeds 0.35, the system identifies severe capacity concentration and automatically prunes primitives with persistently low visibility. The freed parameter budget is immediately redensified in regions exhibiting high pixel residual error. This converts external geometric regularization into renderer-intrinsic capacity allocation, maintaining gradient conditioning while preserving full geometric flexibility.
Loss & Training¶
The framework minimizes a standard pixel-space \(\ell_2\) reconstruction objective \(\mathcal{L}(\theta) = \| R_S(\theta) - I^* \|_2^2\) without auxiliary geometric penalty terms. The parameters are optimized via the Adan optimizer with learning rates set to \(0.1\) for color and opacity, \(10^{-3}\) for control points, and \(10^{-4}\) for stroke width. Training runs for 12,000 steps in open mode and 10,000 steps in closed mode. Gini pruning and densification are executed every 500 steps and frozen during the final 2,000 steps to ensure fine boundary convergence. Evaluation renders at \(S^* = \max(S, 24)\) or using an exact Bézier rasterizer, guaranteeing zero inference quality loss.
Key Experimental Results¶
Main Results¶
CubicSplat is evaluated against leading differentiable vector rasterizers on the 200-image DIV2K validation subset and the 24-image Kodak dataset at full native resolutions on an NVIDIA RTX 4090 GPU.
Table 1: Single training-step execution time on a 2040×1344 image with 2048 curves on RTX 4090
| Topology Mode | Step Phase | DiffVG | Bézier Splatting | CubicSplat (Ours) | Relative Speedup over Baselines |
|---|---|---|---|---|---|
| Open Curves | Forward Pass | 141.3 ms | 4.5 ms | 2.70 ms | 52.3× faster than DiffVG, 1.67× faster than Bézier Splatting |
| Backward Pass | 701.3 ms | 4.7 ms | 4.07 ms | 172.3× faster than DiffVG, 1.15× faster than Bézier Splatting | |
| Total Step Time | 842.6 ms | 9.2 ms | 6.77 ms | 124.5× faster than DiffVG, 1.36× faster than Bézier Splatting | |
| Closed Curves | Forward Pass | 85.2 ms | 14.1 ms | 10.68 ms | 7.98× faster than DiffVG, 1.32× faster than Bézier Splatting |
| Backward Pass | 448.3 ms | 24.58 ms | 8.96 ms | 50.0× faster than DiffVG, 2.74× faster than Bézier Splatting | |
| Total Step Time | 533.5 ms | 38.68 ms | 19.64 ms | 27.2× faster than DiffVG, 1.97× faster than Bézier Splatting |
Table 2: Quantitative reconstruction quality and training runtime on the DIV2K benchmark (200 images)
| Topology Mode | Method | 256 Curves (SSIM / PSNR / Time) | 512 Curves (SSIM / PSNR / Time) | 1024 Curves (SSIM / PSNR / Time) |
|---|---|---|---|---|
| Open Mode | DiffVG | 0.552 / 19.83 dB / 18.9 min | 0.587 / 21.47 dB / 22.0 min | 0.616 / 22.62 dB / 30.6 min |
| Bézier Splatting | 0.600 / 22.17 dB / 3.4 min | 0.646 / 23.79 dB / 3.3 min | 0.699 / 25.45 dB / 3.2 min | |
| CubicSplat (\(S=24\)) | 0.624 / 23.07 dB / 1.8 min | 0.662 / 24.43 dB / 1.8 min | 0.708 / 25.93 dB / 1.8 min | |
| Closed Mode | DiffVG | 0.578 / 20.69 dB / 16.4 min | 0.601 / 21.82 dB / 18.5 min | 0.631 / 22.95 dB / 25.1 min |
| LIVE | 0.576 / 20.09 dB / 2.6 h | 0.611 / 21.70 dB / 4.2 h | 0.648 / 23.11 dB / 5.1 h | |
| LIVSS | 0.586 / 17.71 dB / 39.2 min | 0.630 / 18.71 dB / 54.3 min | 0.678 / 19.83 dB / 1.4 h | |
| Bézier Splatting | 0.580 / 20.74 dB / 7.8 min | 0.607 / 22.11 dB / 8.3 min | 0.639 / 23.45 dB / 8.6 min | |
| CubicSplat (\(S=12\)) | 0.629 / 23.00 dB / 1.8 min | 0.664 / 24.31 dB / 1.9 min | 0.705 / 25.78 dB / 2.1 min |
Ablation Study¶
Ablations on the uniformly sampled DIV2K dataset examine sampling density \(S\), geometric regularization, Gini pruning, and adaptive subdivision, reporting throughput in Tasks Per Minute (TPM).
Table 3: Ablation study across curve budgets, topologies, and architectural variations on DIV2K
| Topology Mode | Configuration Variant | 256 Curves (PSNR / TPM) | 512 Curves (PSNR / TPM) | 1024 Curves (PSNR / TPM) | Core Finding |
|---|---|---|---|---|---|
| Open Mode | Default (\(S=24\)) | 23.07 dB / 0.861 | 24.43 dB / 0.883 | 25.93 dB / 0.873 | Optimal quality and throughput trade-off |
| + Geometric Regularization | 22.60 dB / 0.894 | 23.89 dB / 0.938 | 25.28 dB / 0.917 | Restricts non-convex primitives, dropping PSNR by 0.65 dB | |
| - Without Gini Pruning | 22.38 dB / 0.971 | 23.56 dB / 0.982 | 24.83 dB / 0.952 | Occluded redundant curves accumulate, dropping 1.10 dB | |
| Denser Sampling \(S=32\) | 23.06 dB / 0.859 | 24.42 dB / 0.868 | 25.94 dB / 0.830 | Demonstrates solution-equivariance; no quality gain | |
| Coarse Sampling \(S=8\) | 22.43 dB / 0.753 | 23.64 dB / 0.775 | 24.97 dB / 0.832 | Forward bias exceeds well-conditioned threshold | |
| de Casteljau Subdivision | 23.01 dB / 0.645 | 24.37 dB / 0.631 | 25.87 dB / 0.633 | Dynamic graph splits induce gradient variance and slow runtime | |
| Closed Mode | Default (\(S=12\)) | 22.99 dB / 0.879 | 24.32 dB / 0.786 | 25.78 dB / 0.654 | Winding number ensures exact interior labeling at small \(S\) |
| Denser Sampling \(S=24\) | 23.01 dB / 0.572 | 24.33 dB / 0.481 | 25.80 dB / 0.413 | Marginal 0.02 dB gain while throughput drops over 36% | |
| + Geometric Regularization | 22.63 dB / 0.650 | 23.95 dB / 0.578 | 25.38 dB / 0.503 | Impairs complex boundary fitting, dropping 0.40 dB | |
| - Without Gini Pruning | 22.21 dB / 0.546 | 23.22 dB / 0.456 | 24.21 dB / 0.390 | Capacity collapse under large budgets (-1.57 dB drop) | |
| de Casteljau Subdivision | 22.97 dB / 0.517 | 24.29 dB / 0.531 | 25.73 dB / 0.481 | Trails static uniform sampling in both PSNR and throughput |
Key Findings¶
- High Parameter Efficiency & Generational Scaling: CubicSplat demonstrates remarkable parameter compactness. With only 256 closed curves on DIV2K (23.00 dB), it matches or exceeds the fidelity of DiffVG with 1024 curves (22.95 dB). At 1024 closed curves, it outperforms Bézier Splatting by 2.33 dB in PSNR while slashing per-image training time from 8.6 minutes to 2.1 minutes (a 4× wall-clock reduction).
- Solution-Equivariance Across Sampling Densities: Beyond minimal geometric fidelity thresholds (\(S=12\) for closed shapes, \(S=24\) for open strokes), refining polyline resolution up to \(S=48\) alters converged PSNR by less than 0.03 dB while inflating computation by 4×. This confirms that \(S\) serves as an optimization-oracle conditioning parameter rather than a precision bottleneck.
- Amplified Pruning Utility under Dense Primitive Regimes: Omitting transmittance Gini pruning causes minor degradation at 256 curves (~0.7 dB) but results in a massive 1.10–1.57 dB collapse at 1024 curves, proving that pruning occluded primitives is essential to prevent gradient interference in dense scenes.
- Extreme Resolution Scaling & High-Primitive Capacity: Under extreme 24K magnification (144× pixel count), downsampled reconstruction PSNR remains within 0.5 dB of native 2K baselines, completely avoiding the high-frequency phase artifacts of point-based splatting. Furthermore, scaling closed curves to 32,768 (32K) primitives achieves 38.09 dB PSNR on complex illustrations while consuming only 2,118 MB of VRAM, validating tile-parallel static graph scalability.
Highlights & Insights¶
- Rethinking Differentiable Rendering Philosophy: Instead of chasing maximal forward geometric exactness, the paper reframes forward rasterization as an optimization oracle. An \(\mathcal{O}(S^{-2})\) controlled relaxation eliminates nonlinear root finding and dynamic branching, establishing a well-conditioned static computation graph.
- Zero-Cost Repurposing of Transmittance Byproducts: The method extracts per-primitive visibility weights directly from forward alpha compositing, substituting external geometric regularizers with an internal Gini-driven capacity allocation mechanism.
- Generalizable Paradigm for Physics-Based Optimization: The core insight—relaxing the forward pass to stabilize backward conditioning while strictly bounding forward error—offers an actionable framework for other gradient-sensitive inverse problems such as NeRFs, differentiable physics, and implicit geometry fitting.
Limitations & Future Work¶
- Perceptual Micro-Texture Discrepancy under Sparse Budgets: At low curve budgets on open strokes, CubicSplat achieves higher PSNR and SSIM but occasionally exhibits higher LPIPS than Bézier Splatting. Discretized Gaussian splats accidentally mimic high-frequency micro-textures, whereas smooth parametric curves appear stylized when budget-constrained.
- Uniform Global Discretization: A single global sampling density \(S\) is shared across all primitives, which can lead to slight over-sampling on tiny details or under-sampling on extremely long strokes.
- Future Directions: The author highlights integrating CubicSplat's fast static-graph pipeline into text-to-vector diffusion frameworks (such as VectorFusion/DiffSketcher) to enable real-time interactive design and temporally consistent video vectorization.
Related Work & Insights¶
- vs DiffVG (ACM TOG 2020): DiffVG evaluates exact analytical coverage by differentiating through iterative cubic root solvers, leading to severe gradient ill-conditioning and slow runtime; CubicSplat substitutes closed-form polyline distances, accelerating step times by 27–124× and eliminating gradient blowups.
- vs Bézier Splatting (arXiv 2025): Bézier Splatting relies on discrete Gaussian point splats, requiring millions of primitives for closed fills and introducing scale-dependent distortion under magnification; CubicSplat maintains true vector primitives with signed residuals, speeding up closed fills by 2–4× while preserving scale consistency up to 24K.
- vs LIVE / LIVSS (CVPR 2022 / 2025): Layer-wise greedy vectorization demands hours of sequential optimization per image; CubicSplat optimizes all primitives concurrently in an end-to-end global pass, taking roughly 2 minutes per image on DIV2K.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ (Formulates the fundamental "gradient seesaw" paradox and resolves it through error-bounded forward relaxation and oracle conditioning theory)
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Comprehensive validation across DIV2K, Kodak, 24K magnification, 32K primitives, throughput analysis, and exhaustive ablations)
- Writing Quality: ⭐⭐⭐⭐⭐ (Rigorous mathematical derivation, elegant presentation, and cohesive exposition from oracle theory to empirical findings)
- Value: ⭐⭐⭐⭐⭐ (Significant performance and efficiency breakthrough for differentiable vector graphics, open-source code, and foundational implications for inverse graphics)