The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation¶
Conference: NeurIPS2026
arXiv: 2609.30478
Code: https://github.com/snskysk/event2rgb-distillation
Area: Event-based Vision / Interpretability
Keywords: event camera, cross-domain knowledge distillation, shape bias, edge structure, spectral trade-off
TL;DR¶
The paper uses standard knowledge distillation as an analytical tool to transfer an event teacher's predictions into an RGB student, finding joint improvements in tolerance to color changes, shape preference, and high-frequency-band noise, but with lower clean accuracy and sensitivity to disrupted geometric continuity rather than comprehensive robustness.
Background & Motivation¶
ImageNet-trained CNNs often classify using local texture, but such cues can fail under changes in color, imaging conditions, and distribution. Event cameras record temporal changes in pixel brightness rather than complete RGB appearance: static uniform regions typically emit no events, while motion converts spatial brightness gradients into temporal changes that preserve edges and contours more reliably. Texture can also generate events through motion, however, so event input is not simply a texture-free shape map. The substantive question is how this sensor selectivity changes learned representations.
Direct analysis is constrained by the event domain's limited diagnostic toolkit. RGB vision already has ImageNet-C, shape–texture cue-conflict images, and shape-preserving evaluations, whereas the event domain lacks equally mature tools. The authors therefore reverse the usual RGB-to-event distillation direction: an event teacher guides an RGB student, whose behavior is then examined in the RGB domain. The objective is neither a smaller network nor primarily higher ImageNet accuracy, but observing what heterogeneous supervision transfers within a familiar evaluation space.
This transfer also permits mechanistic controls: if grayscale, binary, RGB-edge, or still-image-difference teachers reproduce every phenomenon, event supervision may add nothing distinctive. The paper investigates the joint profile across several diagnostic axes rather than claiming that color invariance or high-frequency tolerance can only arise from event data. Core idea: use standard logit distillation on paired event–RGB instances to transfer event-domain brightness-gradient selectivity into RGB representations, then use complementary diagnostics to characterize the edge-structure prior and its costs.
Method¶
Overall Architecture¶
Training uses instance-paired ImageNet RGB images and N-ImageNet event samples. A frozen DiST ResNet-34 teacher reads the event representation, while an RGB ResNet-34 student learns both class labels and the teacher's class probabilities. After training, the student requires only RGB input, with neither the event camera nor the teacher needed at inference time.
The analytical pipeline comprises Paired Event Supervision, Dual-Branch Distillation, and Complementary Diagnostics. The first two produce comparable RGB models; the last examines representation changes through color, shape, frequency, and transfer behavior rather than introducing another training loss. Solid edges below indicate training data and supervision, while dashed edges indicate post-training RGB evaluation.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
I["Paired events and RGB<br/>with class labels"] --> P["Paired Event Supervision<br/>Frozen DiST teacher"]
I --> D["Dual-Branch Distillation<br/>Weak RGB KL / strong RGB CE"]
P -->|Event soft labels, training only| D
D --> S["Trained RGB student"]
S -.->|RGB-only inference, not training supervision| A["Complementary Diagnostics<br/>Color / shape / frequency / transfer"]
Key Designs¶
1. Paired Event Supervision: turn sensor selectivity into transferable class distributions
N-ImageNet is generated by recapturing ImageNet images displayed on a monitor with an event camera, retaining the same 1000 classes and instance correspondence. During distillation, the teacher reads the event sample paired with that RGB instance—not the student's current RGB view or a fixed class template. Instance pairing gives each KL comparison a clear semantic meaning: both domains describe the same object, and the student must reproduce the teacher's relative judgments among classes given that event evidence.
The principal experiment uses publicly released DiST ResNet-34 weights without updating the teacher, whose N-ImageNet validation top-1 accuracy is 48.43%. Matching teacher and student backbones helps concentrate the comparison on supervision rather than model size. Nevertheless, teacher accuracy and the RGB baseline's 73.90% concern different input domains and cannot directly rank teacher quality. Teacher weakness and output entropy remain potential influences; the authors offer arguments based on coherent behavior but no entropy-matched causal control.
The relevant event property here is brightness-gradient selectivity, not the full set of event-camera hardware advantages. Monitor recapture does not test real-scene HDR or depth-dependent motion parallax. The experiment therefore supports the transfer of a structural prior from this event-training distribution, not a universal advantage of native dynamic event data.
2. Dual-Branch Distillation: retain RGB classification supervision while accepting event constraints
The same RGB instance passes through strong and weak augmentation pipelines. The strong view receives ground-truth cross-entropy supervision, whereas the weak view matches the temperature-softened event-teacher output. Strong augmentation supports ordinary RGB classification, and the more stable weak view carries cross-domain class-distribution matching. Both losses update one student; this is neither two independently trained students nor a two-branch prediction ensemble at inference time.
ED_Student uses \(\alpha=0.2\), ED_SoftLabelOnly uses \(\alpha=1.0\), and the RGB baseline uses \(\alpha=0\). Removing RGB ground-truth supervision in the soft-only setting makes event-teacher signatures more prominent but yields only 53.50% clean accuracy. Mixed supervision recovers 69.70%, still below the baseline's 73.90%. ED_SoftLabelOnly is thus primarily a mechanism probe, whereas ED_Student tests whether the structural prior survives alongside more practical classification capability.
The weak RGB view uses deterministic Resize and CenterCrop, while the event side independently samples time reversal, horizontal flipping, and spatial shifts. Their spatial views are not synchronized per iteration. Class-probability matching does not require pixel alignment, but class preservation does not guarantee identical prediction distributions. The authors interpret the asymmetry as stochastic regularization and acknowledge that they did not experimentally isolate this confound; it should not be described as ruled out.
3. Complementary Diagnostics: constrain the mechanism using both behavioral changes and failure boundaries
The color axis uses hue rotation and grayscale inputs to test dependence on color. The shape axis combines shape–texture cue conflict, ImageNet-Sketch, binary images, Canny edges, and silhouette masks. Under the 16-class restricted-argmax protocol, shape bias measures the proportion of shape choices among valid shape or texture choices. A higher fraction indicates a relative shift toward shape, not overall classification accuracy or proof of exclusive shape reliance. ED_Student's 0.375 also remains far below the human reference of 0.959.
Patch shuffling and rotation preserve some within-patch local cues while disrupting geometric continuity across patches. Event-distilled students degrade more sharply under these conditions, complementing the preference shift in cue-conflict images. Such controlled information-suppression probes add mechanistic evidence, but are not interchangeable with the shape-bias metric or with every other geometric transformation.
The frequency axis distinguishes removing unused texture from contaminating relied-upon bands. JPEG compression can remove high-frequency detail while preserving contours, making it relatively favorable to the student. Severe pixelation, blur, or deformation can disrupt edge continuity and eliminate or reverse the advantage. Broadband noise also contaminates low–mid-frequency structure, so tolerance to high-frequency-band noise does not imply uniformly better performance on Gaussian, shot, or impulse noise. This note reports only defensive measurement conclusions, not perturbation construction or optimization procedures.
Internal evidence includes early-layer filters, correlation statistics, and cross-model layerwise similarity. ED_SoftLabelOnly has visibly smoother first-layer filters and higher correlations, but ED_Student's mean first-layer correlations remain close to baseline. The soft-only result cannot simply be assigned to the main model. Appendix C.4.2 further reports a mean maximum filter cosine similarity of 0.161 between the event students at layer1, above the baseline–ED_Student value of 0.142. The authors therefore locate the change mainly in early representations rather than explaining all behavior with one first-layer scalar.
Input statistics supply upstream evidence: radial PSD power-law exponents fitted over the same intermediate band are 2.46 for ImageNet and 1.48 for N-ImageNet, indicating a flatter event spectrum. This does not contradict reduced reliance on unstructured high-frequency texture in the RGB student. Relatively emphasized edges are not confined to extreme frequencies; low–mid bands also carry broad contours and continuous gradients. Spectral statistics, filters, and behavior form a coherent hypothesis, not a complete causal demonstration.
A Worked Example¶
Consider one N-ImageNet event sample paired with its original ImageNet RGB image. The frozen DiST teacher produces a class distribution from the events; the student matches it on the weak RGB view and computes ground-truth CE on the strong RGB view. This describes one training instance's data flow without assigning unreported probabilities or local edges to a particular image.
After training, inference uses RGB alone. In Table 14, ED_Student's clean accuracy is 69.70%, below baseline's 73.90%, whereas its Canny-input accuracy is 13.27%, above 10.83%. These aggregate test results illustrate how the transferred structural preference can help under a specific information-removal condition; they neither establish that this particular instance is correct nor introduce the event teacher into RGB inference.
Loss & Training¶
The paper uses a standard knowledge-distillation objective rather than a new loss. Its Eq. (1) is retained below; both distributions in the KL term come from temperature-softened logits, while CE supplies ordinary ground-truth classification supervision:
Here \(y\) is the ground-truth label; \(x_{\text{strong}}\) and \(x_{\text{weak}}\) derive from the same RGB instance; \(x_{\text{event}}\) is its paired event sample; and \(p_t\) and \(p_s\) denote teacher and student outputs. All KD students use \(T=3.0\).
“Soft:hard = 2:8” describes only \(\alpha\) and \(1-\alpha\) before multiplication by \(T^2\). ED_Student's actual KL coefficient is \(0.2\times3^2=1.8\), while the CE coefficient is 0.8. Neither effective loss contributions nor gradient contributions are therefore established as 2:8. Relative gradients were not measured, so this shorthand cannot prove that CE necessarily dominates a particular layer.
The principal baseline, ED_Student, and ED_SoftLabelOnly runs use AdamW, batch size 512, learning rate \(2\times10^{-3}\), weight decay \(5\times10^{-2}\), and 120 epochs with 5 epochs of linear warmup followed by cosine decay. They share seed 1. The main ResNet-34 setting does not use MixUp or CutMix; the 300-epoch stronger-recipe check is a separate experiment.
Downstream evaluation is strict linear probing: the entire backbone is frozen and only a new linear classification layer is trained. Datasets share a 50-epoch, batch-size-1024 SGD recipe. Each model–dataset pair has two seeds, however, and Table 1 reports the higher of their best validation accuracies rather than a mean and standard deviation. This measures frozen-feature transferability, not full fine-tuning capacity.
Key Experimental Results¶
Main Results¶
The following selection from original Table 14 reports top-1 accuracy (%). Sketch is the independent ImageNet-Sketch dataset, not a transformation of ImageNet validation images; the other information-removal conditions use processed validation images.
| Model | Clean RGB | Grayscale | Sketch | B&W | Canny | Silhouette |
|---|---|---|---|---|---|---|
| baseline | 73.90 | 62.45 | 23.70 | 41.39 | 10.83 | 3.30 |
| ED_Student | 69.70 | 64.64 | 28.08 | 49.01 | 13.27 | 5.43 |
| ED_SoftLabelOnly | 53.50 | 52.58 | 24.47 | 34.87 | 11.45 | 5.42 |
| Gray Distilled | 74.60 | 69.51 | 25.60 | 45.56 | 7.87 | 3.63 |
| SIN | 72.18 | 62.53 | 30.84 | 51.62 | 25.77 | 5.38 |
ED_Student loses 4.20 percentage points on clean RGB but exceeds baseline on grayscale, Sketch, and the three further information-removal inputs. Gray Distilled is better on grayscale, and SIN is stronger on Sketch, B&W, and Canny, so event distillation is not optimal on every individual metric. Silhouette accuracy is also only 5.43%, hardly a solution to contour-only recognition.
Across the 19 downstream tasks in Table 1, mean linear-probe accuracy rises from 73.5% to 75.1%, a 1.6-point gain. CIFAR-100, CIFAR-10, Chest X-ray, and MNIST gain 5.2, 4.8, 4.3, and 3.4 points, respectively; Oxford-IIIT Pet and Food101 lose 0.1 and 0.4 points. The mean excludes pretraining ImageNet and is not a sample-weighted pooled accuracy across tasks.
Ablation Study¶
The following selection is from original Table 2. Clean is absolute accuracy; all four remaining columns are percentage-point differences from baseline, signed so that positive means improvement. The high-frequency column measures reduction in the accuracy drop under \(0.95\pi\) band noise, not absolute corrupted-input accuracy.
| Supervision | Clean (%) | Mean hue gain | Sketch gain | High-frequency drop reduction | Mean linear-probe gain |
|---|---|---|---|---|---|
| ED_Student | 69.7 | +3.6 | +4.4 | +17.4 | +1.6 |
| Canny_Distilled | 71.1 | +4.3 | +4.2 | −5.9 | +0.2 |
| Synth_Distilled | 71.8 | −3.4 | −3.2 | +18.9 | −5.4 |
| Adversarial Training | 66.0 | −11.2 | −2.4 | +28.1 | −1.8 |
| Noise Training | 74.4 | +0.8 | +1.7 | +13.4 | −0.8 |
| SIN | 72.2 | +0.5 | +7.1 | +19.3 | −0.9 |
Baseline references on these four axes are mean hue accuracy 61.8%, Sketch accuracy 23.7%, a high-frequency accuracy drop of 29.6 points, and mean linear-probe accuracy 73.5%. ED_Student drops only 12.2 points under the high-frequency condition, yielding the 17.4-point improvement above. Robustification baselines retain their own training recipes rather than differing only in supervision source.
The Canny teacher uses RGB edge maps and reproduces color and Sketch benefits, but not high-frequency tolerance or the mean transfer gain. The Synth teacher uses an ON/OFF proxy from grayscale differences between augmented still-RGB views, not an event sensor or event simulator. It tolerates high-frequency noise but regresses on color, Sketch, and mean transfer. The conclusion is that the tested non-event supervisions fail to reproduce the joint four-axis profile, not that event data are indispensable for any individual property.
The following mechanism analysis is from original Table 15, not an independent performance ablation. RGB Correlation describes cross-RGB-channel filter correlation, and Spatial Autocorrelation concerns neighboring pixels in first-layer filters. The full text does not provide a more detailed aggregation formula, so an exact implementation is not reconstructed here.
| Model | RGB Correlation | Spatial Autocorrelation |
|---|---|---|
| baseline | 0.549 | 0.391 |
| ED_Student | 0.544 | 0.398 |
| ED_SoftLabelOnly | 0.747 | 0.575 |
Key Findings¶
- Shape bias increases from baseline's 0.212 to ED_Student's 0.375; the soft-only model reaches 0.487. This supports preference transfer, not strict reliance or human-level performance at the 0.959 reference.
- ImageNet-C covers 19 corruptions, but gains are uneven. JPEG and fog are favorable; severe pixelate, contrast, elastic_transform, and some blur conditions reverse the advantage. Broadband noise may also match or underperform baseline.
- Across five architectures in the limited defensive measurements, ED_Student exceeds baseline in only 19 of 30 architecture–evaluation–budget combinations. Most of the 11 exceptions involve ConvNeXt-Base and RepLKNet31B; uniformly improved cross-architecture robustness is not established.
- Defensive adversarial conclusions are restricted to \(0<\epsilon\leq1/255\) and the paper's limited evaluation, without larger budgets or AutoAttack. The stronger training recipe removes the advantage on that axis while retaining gains in color, Sketch, high-frequency-band tolerance, and transfer.
- Under the stronger recipe in Table 12, clean top-1 is 75.92% versus 70.80%, Sketch is 26.98% versus 28.94%, and mean downstream accuracy is 74.19% versus 75.45%. Retained benefits still accompany lower clean accuracy; stronger training does not remove the cost.
Highlights & Insights¶
- Distillation as a measurement bridge. A domain lacking diagnostics need not first build every tool before being analyzed: transferring supervision into an RGB student enables existing controlled evaluations. The novelty lies in the experimental question and cross-domain setup, not a more elaborate loss.
- Joint behavior is more informative than a single win. Canny, Synth, and SIN reproduce different subsets of the benefits but do not improve all four axes together. This confines event-specific claims to the joint profile among current controls rather than over-attributing one metric.
- Failure boundaries also inform the mechanism. Greater fragility to disrupted geometry is an expected cost if the model emphasizes edge continuity. Separating texture removal from structural contamination is more explanatory than treating all corruptions as generic noise.
Limitations & Future Work¶
- Limited data coverage. Monitor recapture isolates brightness-gradient selectivity, not HDR or real three-dimensional motion parallax. Native paired event–frame data are needed to test these extrapolations.
- Remaining mechanistic confounds. Teacher confidence, cross-domain view mismatch, and the soft-only model's lack of strong CE augmentation are not fully isolated. Entropy-matched teachers and augmentation-matched controls are needed; explanatory arguments are not experimental elimination of a confound.
- Capacity and task boundaries. Validation focuses on CNN image classification, with weaker gains for large backbones. ViTs, detection, and segmentation remain untested, and logit distillation is not comprehensively compared with feature-level distillation.
- Incomplete statistical reporting. Main training uses one shared seed, and downstream results select the better of two validation runs without variance reporting. Repeated experiments are needed to assess stability and selection bias.
- Real accuracy costs. Soft-only supervision is not a general deployable replacement, and mixed supervision still loses clean accuracy. Selection should depend on expected distribution shifts and the relevance of structure, not shape bias alone.
Related Work & Insights¶
- vs standard KD and cross-modal distillation: The method retains Hinton-style soft-label distillation but uses the event-to-RGB direction to analyze heterogeneous supervision, rather than compression or new-task accuracy. DeiT and Patient Teacher motivate dual-view training but do not validate view consistency in this experiment.
- vs Stylized-ImageNet / SIN: SIN reduces texture shortcuts through stylization and is substantially stronger on some shape inputs. The distinction here is the joint trade-off across color, frequency, and frozen-feature transfer—not the highest possible shape bias.
- vs Burgert et al.'s controlled-suppression analysis: Cue conflict measures choice preference; strict reliance requires complementary information-suppression evidence. Patch diagnostics and shape-input tests add evidence, but their measurements remain distinct.
- Research direction: Comparing native event supervision, event representations, and entropy-matched RGB teachers while tracking early-layer representations and frozen transfer could separate sensor statistics from distillation regularization. This remains a proposed investigation, not a result already established by the paper.
Rating¶
- Novelty: 4/5 — Event-to-RGB distillation provides a clear representation-analysis contribution, while the objective remains standard.
- Experimental Thoroughness: 4/5 — Multi-axis controls, non-event supervision, and a stronger-recipe check are extensive, but causal confounds and statistical uncertainty remain.
- Writing Quality: 4/5 — The narrative is coherent and appendices clarify filter correlations and robustness boundaries; some strong mechanistic wording requires its qualifications.
- Value: 4/5 — A useful route to analyzing event supervision and transferable structural priors, not comprehensive robustification or a cost-free accuracy improvement.