Skip to content

Synthetic Sub-Aperture Phase Augmentation for Demosaicing 2ร—2 Shared Microlens Sensors

Conference: ECCV 2026
Paper: ECCV Original
Area: Image Restoration
Keywords: Demosaicing, Shared Microlens, Phase Modeling, Synthetic Data, Phase-Detection Autofocus

TL;DR

Addressing grid-aligned checkerboard and chromatic artifacts caused by sub-aperture phase shifts under defocus in 2ร—2 shared microlens sensors, this paper proposes Synthetic Sub-Aperture Phase Augmentation (SPA)โ€”a calibration-free physical data synthesis pipelineโ€”which, paired with a lightweight network, suppresses periodic artifacts while preserving natural optical defocus blur.

Background & Motivation

Modern high-resolution mobile CMOS image sensors exceeding 200 megapixels frequently adopt clustered Color Filter Array (CFA) layouts, such as Quad (2ร—2) and Hexadeca (4ร—4) Bayer patterns, to boost low-light signal-to-noise ratio via on-sensor pixel binning within tiny pixel pitches. Simultaneously, premium mobile sensors incorporate quad-pixel phase-detection autofocus (PDAF) architectures utilizing 2ร—2 on-chip-lens (OCL) configurations with shared microlenses (termed Q-cell lenses). Under this design, four adjacent same-color photodiodes share a single microlens, capturing top-left (TL), top-right (TR), bottom-left (BL), and bottom-right (BR) directional sub-aperture rays in a single exposure. While providing omnidirectional, dense autofocus coverage, this optical layout introduces an underexplored, severe physical artifact during full-resolution demosaicing.

Traditional deep demosaicing architectures (e.g., DPN, NAFNet, and KLAP) rely fundamentally on the assumption of intra-cluster photometric consistency across adjacent pixels of the same color. However, under optical defocus, the four sub-apertures beneath a shared microlens experience perspective parallax displacements. This induces severe polarity alternation and intensity imbalance across sub-pixel locations within each 2ร—2 unit cell. Furthermore, oblique ray incidence toward sensor peripheriesโ€”governed by the chief ray angle (CRA)โ€”imposes directional asymmetric attenuation, creating high-frequency periodic sampling biases directly in the RAW domain. Because conventional demosaicers misinterpret this structured optical bias as true high-frequency image edges and chrominance textures, the reconstruction severely amplifies them into stubborn checkerboard artifacts, zippering patterns, and color misalignment. Calibrating real point spread functions (PSFs) across every sensor model, lens stack, field position, and focus distance is prohibitively expensive and unscalable for mobile deployment.

This paper's angle of attack is that demosaicing networks fail on these periodic artifacts primarily due to training distribution deficiency: standard RAW synthesis pipelines (such as unprocessing) only simulate noise and inverse ISP curves while completely overlooking sub-aperture phase displacements. Core idea: propose Synthetic Sub-Aperture Phase Augmentation (SPA), a physics-inspired data generation framework that models four-directional sub-aperture kernels and chief-ray-angle tilt perturbations with a symmetric blur training target, guiding lightweight networks to disentangle phase polarity biases from natural textures without requiring measured PSFs or sensor-specific optical calibration.

Method

Overall Architecture

The SPA framework simulates the four-directional phase sampling of shared microlens sensors during data synthesis and performs full-resolution RGB restoration via an ultra-compact demosaicing network. Given an uncompressed scene image, it is first mapped to scene-referred linear RGB via inverse ISP operations (undoing tone mapping, gamma, color correction, and white balance). A base circle-of-confusion (CoC) kernel is synthesized using a thin-lens model and Butterworth radial profiling, from which four anisotropic sub-aperture kernels (TL, TR, BL, BR) are derived through orthogonal linear attenuation fields and randomized local CRA tilt factors. After filtering the linear image into four parallax views, they are interleaved into a 2ร—2 clustered RAW mosaic, masked by CFA layouts, and perturbed with Poissonโ€“Gaussian noise. The network, PhaseUNet, is supervised end-to-end by the symmetric blur target, learning to suppress phase-induced grid artifacts.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Scene-Referred Linear RGB Image"] --> B["Physics-Inspired Anisotropic Sub-Aperture Kernel Sampling<br/>Thin-lens CoC + directional linear attenuation fields"]
    B --> C["Chief-Ray-Angle Tilt Perturbation Modeling<br/>Local tilt sampling to introduce controlled energy imbalance"]
    C --> D["Symmetric Defocus Target Constraint & RAW Interleaving<br/>4-view interleaving + CFA masking + noise injection"]
    D --> E["Lightweight Phase-Aware Demosaicing Network<br/>PhaseUNet full-resolution RGB reconstruction"]

Key Designs

1. Physics-Inspired Anisotropic Sub-Aperture Kernel Sampling: Decoupling Four-Directional Parallax

To realistically emulate the energy imbalance of Q-cell architectures under defocus without sensor-specific calibration, the method builds upon thin-lens optics and a radially increasing Butterworth profile to generate a symmetric base kernel \(K_0\). Given object distance \(d\) and focus distance \(s\), the signed circle-of-confusion (CoC) radius \(r\) determines front- or back-focus polarity as well as kernel footprint. A central-dip disk mask is smoothed to create \(K_0\). Horizontal and vertical linear attenuation ramps \(D_x(i, j) = \frac{j}{k-1}\) and \(D_y(i, j) = \frac{i}{k-1}\) are combined with spatial flipping to construct cardinal sub-aperture kernels:

\[K_L = K_0 \odot D_x, \quad K_R = \text{flip}_x(K_L), \quad K_T = K_0 \odot D_y, \quad K_B = \text{flip}_y(K_T)\]

Diagonal kernels representing four microlens quadrants are generated via pairwise averaging (e.g., \(K_{TL} = \frac{1}{2}(K_T + K_L)\)). Front- and back-focus polarity reversal is seamlessly modeled by flipping attenuation fields \((D_x, D_y) \to (1-D_x, 1-D_y)\), accurately reproducing the phase-induced polarity alternation observed in physical sensors.

2. Chief-Ray-Angle Tilt Perturbation Modeling: Covering Peripheral Oblique Incidence

In mobile optical stacks, light incident on the sensor periphery forms steep chief ray angles, causing asymmetric vignetting and differential optical crosstalk across the four photodiodes. To impart robust artifact suppression across variable field heights without explicit coordinate conditioning, SPA introduces a stochastic patch-level tilt parameter \(\delta \sim \text{Uniform}(-0.15, 0.15)\) to induce anti-correlated scaling across orthogonal attenuation axes:

\[\alpha_x = 1 + \delta, \quad \alpha_y = 1 - \delta; \quad D_x' = \alpha_x D_x, \quad D_y' = \alpha_y D_y\]

Applying these tilted attenuation fields \(D_x', D_y'\) injects physically plausible energy asymmetry into the synthesized sub-aperture kernels. This prevents the network from overfitting to perfectly balanced laboratory conditions and builds high tolerance against real-world lens shading and microlens manufacturing misalignments.

3. Symmetric Defocus Target Constraint & RAW Interleaving: Constructing Unbiased Supervision

The linear RGB image is convolved with the four directional kernels to produce four sub-aperture views \(I_{lin}^{(v)}\) (\(v \in \{TL, TR, BL, BR\}\)). These views are spatially multiplexed into a single sensor RAW frame according to the parity of spatial coordinates \((i, j)\) in a \(2\times 2\) grid, followed by CFA spectral subsampling (Hexadeca or Quad) and Poissonโ€“Gaussian noise addition. Crucially, the training target is defined as \(I_{tgt} = K_0 * I_{lin}\), filtered solely by the symmetric base kernel. Rather than forcing the network to hallucinate sharp edges in out-of-focus areas (which destroys aesthetic bokeh), this target maintains genuine optical defocus blur while entirely stripping away directional phase disparity, compelling the network to suppress periodic polarity alternation.

4. Lightweight Phase-Aware Demosaicing Network: Validating Data-Driven Superiority with Minimal Compute

To verify that performance gains stem from physics-grounded phase modeling rather than architectural scaling, the authors deploy PhaseUNet, an ultra-compact U-Net variant with base width \(C=32\) and only 0.90M parameters. The network comprises a 3-stage downsampling encoder, a 2-stage upsampling decoder, additive skip connections, and a residual refinement head. Optimized end-to-end using an \(\ell_1\) reconstruction loss \(\mathcal{L}_1 = \|I_{pred} - I_{tgt}\|_1\), this lightweight backbone outperforms models with nearly twenty times more parameters, confirming that physical phase augmentation resolves the root cause of Q-cell demosaicing artifacts.

Key Experimental Results

Main Results

Evaluated on synthetic Q-cell RAW benchmarks (averaged over Kodak, McM, Set5, Set14, and Urban100 with realistic defocus distributions: 50% on-focus, 30% slight, 10% moderate, and 10% severe defocus), PhaseUNet establishes new state-of-the-art results across both Hexadeca (4ร—4) and Quad (2ร—2) CFA patterns:

Model Params (M) Compute (MACs, M) Hexadeca PSNR (dB) โ†‘ Hexadeca SSIM โ†‘ Hexadeca LPIPS โ†“ Quad PSNR (dB) โ†‘ Quad SSIM โ†‘ Quad LPIPS โ†“
DPN 5.58 10775 35.09 0.942 0.0747 35.42 0.945 0.0698
NAFNet 17.11 3985 35.31 0.946 0.0719 35.68 0.948 0.0673
KLAP 17.80 4016 33.64 0.922 0.0965 34.12 0.928 0.0892
PyNet-Qร—Q 1.07 4408 28.35 0.836 0.2057 โ€” โ€” โ€”
PhaseUNet (Ours) 0.90 3010 36.60 0.954 0.0434 37.15 0.958 0.0401

Ablation Study

On the Hexadeca CFA setting, evaluating the impact of SPA across progressive defocus levels using the identical PhaseUNet-H backbone:

Defocus Severity (CoC radius) Configuration PSNR (dB) โ†‘ SSIM โ†‘ LPIPS โ†“ Relative Gain
Slight (\(0 < \|r\| \le 2\)) w/o SPA 32.91 0.941 0.083 Baseline
w/ SPA (Ours) 35.44 0.954 0.043 PSNR +2.53 dB / LPIPS -48%
Moderate (\(2 < \|r\| \le 5\)) w/o SPA 34.76 0.932 0.192 Baseline
w/ SPA (Ours) 41.62 0.985 0.028 PSNR +6.86 dB / LPIPS -85%
Severe (\(\|r\| > 5\)) w/o SPA 33.46 0.827 0.332 Baseline
w/ SPA (Ours) 47.33 0.994 0.015 PSNR +13.87 dB / LPIPS -95%

Quantitative evaluation on real prototype Q-cell RAW captures (116 Hexadeca handheld natural scenes and 52 Quad chart captures) using no-reference CLIPIQA alongside Spatial Lattice Imbalance (SLI) and Nyquist Energy Ratio (NER):

Sensor Layout Evaluation Metric DPN NAFNet KLAP PyNet-Qร—Q PhaseUNet (Ours)
Real Hexadeca CLIPIQA โ†‘ 0.376 0.368 0.356 0.257 0.390
SLI (ร—10โปยณ) โ†“ 4.61 3.97 5.94 10.25 2.97
NER (ร—10โปยฒ) โ†“ 5.067 7.274 3.957 1.204 0.012 (\(1.2\times 10^{-4}\))
Real Quad CLIPIQA โ†‘ 0.409 0.401 0.394 โ€” 0.467
SLI (ร—10โปยณ) โ†“ 2.05 3.26 1.39 โ€” 0.56
NER โ†“ 2.929 2.557 1.909 โ€” 0.00007 (\(7.0\times 10^{-5}\))

Key Findings

  • Phase misalignment becomes the dominant error source under increasing defocus. Under severe blur, omitting SPA caps PSNR at 33.46 dB due to severe checkerboard false structures; incorporating SPA propels PSNR to 47.33 dB (+13.87 dB) and slashes LPIPS by 95%, proving that structured phase imbalance is the primary performance bottleneck.
  • Spectral and spatial periodic artifacts are virtually eradicated. On real Quad sensor captures, NER (measuring periodic energy concentrated at 2ร—2 Nyquist frequencies) collapses from ~2.9 in baselines to \(7\times 10^{-5}\), shifting the Fourier spectrum from sharp discrete artifact spikes to smooth natural image decay.
  • Outstanding zero-shot generalization across CFA patterns. With only 0.90M parameters and 3010M MACs (24% lower compute and 19ร— fewer parameters than NAFNet), the phase-aware model establishes superior perceptual quality and artifact suppression across both 4ร—4 Hexadeca and 2ร—2 Quad sensors.

Highlights & Insights

  • Physics-grounded data augmentation outclasses brute-force model scaling: The breakdown of deep demosaicers on shared microlens sensors is not due to restricted network capacity, but rather a fundamental mismatch in training data physics. Modeling sub-aperture disparity in synthetic RAW enables a 0.9M parameter network to easily surpass 17M+ parameter backbones.
  • Defocus-preserving symmetric blur supervision: Rather than naively targeting sharp ground truth (which produces unnatural ringing and destroys depth-of-field realism), supervising with a symmetric blur kernel isolates the phase disparity error, forcing the model to solely remove periodic grid alternations.
  • Purpose-built physical metrics (SLI & NER): Standard full-reference metrics often overlook localized high-frequency grid ripples. SLI and NER directly quantify spatial sub-lattice variance and Fourier Nyquist peaks, providing invaluable diagnostic tools for computational photography and sensor pipeline design.

Limitations & Future Work

  • First-order optical approximation: The thin-lens and linear attenuation assumptions do not explicitly simulate higher-order optical aberrations (e.g., coma, astigmatism, or field curvature). Extremely wide-angle lenses with severe peripheral distortion may exhibit residual unmodeled phase patterns.
  • Scope and sensor architectures: The primary design centers on 2ร—2 shared microlens (Q-cell) layouts. Adapting to dual-photodiode (1ร—2 2PD) or irregular sparse PDAF layouts requires reconfiguring directional attenuation ramps.
  • Future improvements: Exploring differentiable, field-dependent aberration modeling conditioned on optical metadata, and validating across a broader diversity of real-world multi-camera mobile platforms.
  • vs DPN / NAFNet / KLAP (Clustered CFA Demosaicing): Existing clustered CFA demosaicers focus entirely on spatial interpolation across sparse color clusters under an assumption of intra-cluster smoothness. This work identifies the critical vulnerability of Q-cell phase disparity and resolves it at the data level.
  • vs Dual-Pixel Deblurring: Dual-pixel methods leverage two-view disparity for depth estimation and sharp image restoration in the image domain. In contrast, this paper tackles four-directional phase shifts coupled with Bayer mosaic sampling in the RAW domain to resolve color-domain demosaicing artifacts.

Rating

  • Novelty: โญโญโญโญโญ [Pioneering identification of 2ร—2 shared microlens phase artifacts in clustered CFAs with an elegant, calibration-free synthetic pipeline]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive synthetic benchmarks, real prototype sensor RAW evaluations, and innovative spatial/spectral diagnostic metrics]
  • Writing Quality: โญโญโญโญโญ [Clear mathematical derivation, intuitive physical insights, and compelling visual/quantitative validation]
  • Value: โญโญโญโญโญ [Highly practical, compute-efficient solution for next-generation 200MP mobile sensor ISP pipelines]