Skip to content

Cumulative-Goodness Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant

Conference: NeurIPS2026
arXiv: 2605.06240
Code: https://github.com/amirhossein-yousefi/ff-free-riding
Area: Optimization & Theory
Keywords: Forward-Forward, cumulative goodness, gradient attenuation, local learning, separation–accuracy dissociation

TL;DR

The paper proves that cumulative goodness exactly attenuates deep blocks' local discrimination gradients on examples already separated upstream, and repairs block health through history removal, hardness gating, and gradient compensation, but finds no accuracy-dominant benefit from these repairs; MGC even reduces CIFAR-100 Stage-1 single-crop accuracy by 1.05 percentage points.

Background & Motivation

Forward-Forward (FF) trains layers or blocks using local goodness discrimination objectives on positive and negative examples, avoiding end-to-end cross-block back-propagation; gradient descent remains available within a block. Cumulative-goodness variants add earlier blocks' scores to the current block's training objective to share existing discrimination evidence. The problem is that a low current-block loss may reflect inherited positive margin rather than the block's own ability to distinguish classes: examples already solved upstream can present an almost saturated softplus objective downstream.

Deep-block free-riding is therefore a plausible optimization diagnosis, but it does not establish that this pathology is the main reason FF trails back-propagation (BP). If a repair only reallocates evidence while the classifier still sums scalar goodness across blocks, healthier deep blocks need not improve the total score's classification ability. Three questions must be separated: whether local gradients are suppressed, whether each block's own separation recovers, and whether the deployed readout becomes more accurate. Examining only one can mistake mechanism evidence for performance evidence.

The paper follows this causal chain with analysis and controlled experiments rather than presenting a new accuracy champion. Core idea: first characterize exactly how cumulative margin suppresses local discrimination gradients, then intervene with different repairs and separately measure block health, the native FF readout, and frozen-feature probes to test whether repairing deep blocks narrows the accuracy gap.

Method

Overall Architecture

The inputs are images and candidate classes. Positive examples use the correct class; negatives include wrong labels (NL) and wrong images (NI). Stage 1 trains multiple blocks simultaneously, but each block applies parameter updates generated only by its own loss: upstream tokens and goodness are detached before entering the current block's objective. At test time, the classifier evaluates candidate classes, sums current goodness across blocks without weighting, and selects the highest-scoring class. The training history weight does not enter this inference sum.

The default backbone comprises a convolutional stem, patch embedding, and 4 FF Hybrid Blocks, each containing self-attention with RoPE, a feed-forward network, attention pooling, and a multi-aspect goodness head. The large CIFAR-10 model uses 32 experts with top-4 routing; the CIFAR-100 model carrying the main causal result is instead a dense 4-block network of width 256, with neither MoE nor SAM. The goodness head has four slots: prototype alignment, activation energy, attention sharpness, and a learned score. With memory disabled by default, the attention-sharpness slot is identically zero, leaving only three active aspects.

Within the same block-local training framework, the paper compares history rules and local auxiliary terms, using the exact gradient ratio to explain their effects. Stage 2 freezes the Stage-1 backbone and trains a separate attentive readout, typically for no more than 20 epochs. It is a representation-quality probe, not native FF or backbone retraining. In the diagram, dashed edges denote mechanism choices or training supervision, while solid edges denote model products and readout relationships; the repairs are alternatives, not components that must be stacked sequentially.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Images and positive/negative classes"] --> B["Local Gradient Diagnosis"]
    B -. "Optional history rule or auxiliary" .-> C["Heuristic Local Repairs"]
    B -. "Replace depth-scaled auxiliary" .-> D["Missing-Gradient Compensation"]
    C -. "Own-loss supervision" .-> E["Stage 1 block-local training<br/>Upstream stop-gradient"]
    D -. "Own-loss supervision" .-> E
    E --> F["Dual-Readout Test<br/>S1 goodness sum<br/>S2 frozen-feature attentive probe"]
    F --> G["Report block health and accuracy separately"]

Key Designs

1. Local Gradient Diagnosis: identify how inherited success weakens the current block's learning pressure

For a positive–negative pair, the current block's margin is its own positive goodness minus its negative goodness, denoted by \(m\); summing all upstream block margins gives \(P\). The cumulative objective uses \(M=m+\gamma P\), applying the same history weight to all previous blocks rather than recursive geometric weighting by depth. The softplus discrimination loss is \(\ell_\beta(u)=\log(1+e^{-\beta u})\). Because \(P\) is detached from the current block's parameters \(\theta_d\), the chain rule yields an exact identity for each example and negative stream:

\[ R_\gamma(m,P)=\frac{1+e^{\beta m}}{1+e^{\beta(m+\gamma P)}},\qquad \nabla_{\theta_d}\ell_\beta(m+\gamma P)=R_\gamma(m,P)\nabla_{\theta_d}\ell_\beta(m). \]

When \(m\geq0\), \(P\geq0\), and \(\gamma\geq0\), the ratio lies between \(e^{-\beta\gamma P}\) and \(\min\{1,2e^{-\beta\gamma P}\}\). A larger positive upstream margin therefore weakens the gradient the current block receives from cumulative discrimination. If \(P<0\), the ratio can exceed 1, so upstream failure can amplify the gradient. The theorem does not state that all deep gradients attenuate: it covers only this per-example discrimination term, not the total update containing reconstruction, contrastive, and other terms. It also does not exclude gradient cancellation within a batch or establish global convergence.

To avoid interpreting a saturated loss as success, the authors inspect both gradients and separation. The free-riding index is \(\mathcal{F}_d=\mathbb{E}[1-\min\{1,R_d\}]\), clipping amplification caused by negative upstream margins; values near 1 indicate severe suppression. In the CIFAR-10 mechanism check, the cumulative setting has indices of 0.930/0.952/0.955 at blocks 1–3. These values are not classification error rates.

Two easily confused quantities must also be distinguished: \(\mathrm{sep}^{\mathrm{cur}}_{\mathrm{nl}}(d)=\mathbb{E}[m^{(d)}]\) is the current block's own wrong-label margin, measured in unbounded goodness-margin units; \(\mathrm{sep}_{\mathrm{nl}}(d)=\Pr[M_\gamma^{(d)}>0]\) is the fraction of examples separated by the cumulative objective, ranging over \([0,1]\). Wrong labels come from hard-negative mining, and these separation values are logged on Stage-1 training batches, not measured as all-class test accuracy. DS is accuracy using blocks 0 through the current depth divided by full-model accuracy. LC is the average barrier loss at the end of Stage 1, computed from the inference margin without the history weight. Low LC supports a free-riding diagnosis only when accompanied by low own-block separation.

2. Heuristic Local Repairs: reduce inheritance or make deep blocks directly address unresolved examples

The simplest intervention sets \(\gamma=0\): the training objective no longer includes historical margin, making the ratio above identically 1. This removes the specific cumulative-attenuation mechanism, but does not prevent other depth pathologies or require abandoning multi-block summation at inference. Hardness gating retains selective history sharing by applying \(\gamma_0\sigma(\tau(\kappa-\sum_{j<d}g^{(j)}))\) to the upstream goodness for the same input: high upstream scores reduce inheritance, while low scores permit up to \(\gamma_0\). The gate uses goodness, not the positive–negative margin \(P\) in the theorem; these are not interchangeable difficulty variables.

A third approach adds a current-block-only auxiliary loss to the cumulative objective, increasing its coefficient with depth. In the four-block setting, \(\lambda_0=0.25\) and \(\rho=3.0\) yield coefficients of 0.25/0.50/0.75/1.00, giving deeper blocks a stronger direct discrimination term. Upstream residual margins also reweight examples to emphasize those not yet solved by earlier blocks:

\[ \mathcal{L}^{(d)}_{\mathrm{curr}}=\frac{1}{B}\sum_i w_i^{(d)}\ell_\beta(m_i^{(d)}),\qquad \lambda_{\mathrm{curr}}(d)=\lambda_0\left(1+\rho\frac{d}{L-1}\right),\qquad w_i^{(d)}=\frac{\sigma(-\beta P_i^{(d-1)})}{B^{-1}\sum_k\sigma(-\beta P_k^{(d-1)})+\epsilon}. \]

Weights are computed separately for the NL and NI streams, detached, and not clipped; block 0 has no residual weighting. Their batch mean is exactly 1 when \(\epsilon=0\), and approximately 1 with the implemented \(\epsilon=10^{-6}\). This preserves the average weight scale, not equality between the numerical values of the reweighted and unweighted losses.

For an unresolved example with \(m\leq0\), the discrimination subobjective's margin-gradient magnitude is at least \(\lambda_{\mathrm{curr}}(d)w_i^{(d)}\beta/2\). This floor still depends on the individual residual weight, which can become very small, so there is no uniform positive batch-wide floor. A parameter-gradient floor additionally requires a non-degenerate margin Jacobian and the absence of cancellation by other terms. More importantly, later controls show that this auxiliary can participate in deep-block separation collapse in some recipes: a gradient floor is not a guarantee of healthier training.

3. Missing-Gradient Compensation: replenish local discrimination without claiming full training-gradient recovery

Missing-gradient compensation (MGC) starts from the ratio \(R\): if cumulative discrimination supplies only part of the local gradient, inject the missing part rather than simply increasing the deep-block loss coefficient. With \(s(u)=\sigma(-\beta u)\), the implementation uses the following objective and stops gradients through the compensation coefficient:

\[ \widetilde{R}_i^{(d)}=\frac{s(M_i^{(d)})}{s(m_i^{(d)})+\epsilon},\qquad \lambda_i^{(d)}=[1-\widetilde{R}_i^{(d)}]_+,\qquad \widetilde{\mathcal{L}}^{(d)}_{\mathrm{MGC}}=\frac{1}{B}\sum_i\left[\ell_\beta(M_i^{(d)})+\lambda_i^{(d)}\ell_\beta(m_i^{(d)})\right]. \]

Under the ideal conditions \(\gamma P\geq0\) and \(\epsilon=0\), \(R\leq1\), and adding the cumulative and compensating margin gradients exactly recovers the local objective's magnitude \(\beta s(m)\); the batch-mean loss also carries \(1/B\). Detaching the coefficient is essential, because differentiating through \(R\) would introduce an additional path and invalidate this missing-component addition. MGC omits residual weighting from its local compensator to avoid mixing emphasis on unresolved examples with attenuation cancellation.

Actual training uses \(\epsilon=10^{-6}\), slightly underestimating the ratio and potentially over-compensating. The positive-part clamp prevents a negative compensator when history already amplifies the gradient. The paper therefore claims approximate compensation of the gradient envelope, not exact recovery of the entire block gradient or reproduction of a \(\gamma=0\) network's training trajectory. MGC replaces the depth-scaled auxiliary rather than merely adding another term. Block 0 has no history and is skipped by the compensator, so it also loses its original 0.25 auxiliary.

Consequently, MGC differs from the cumulative baseline through both auxiliary removal and compensation addition. A CIFAR-10 appendix control separates these changes: cumulative training without the auxiliary or compensation already achieves about 92% of the deepest-block separation gain, leaving about 8% attributable to the added compensator. A valid theoretical construction therefore does not establish that all empirical improvements arise from its proposed mechanism.

4. Dual-Readout Test: separate block health, evidence allocation, and final classification

Native Stage-1 prediction always uses \(\hat y=\arg\max_y\sum_d g^{(d)}(x,y)\). Adding to block \(a\) an arbitrary \(q(x,y)\) and subtracting the same quantity from block \(b\) leaves every class's total score, and hence every prediction, exactly unchanged, while potentially changing both blocks' own margins substantially. This is arithmetic invariance under redistribution of fixed scores, not a theorem that retrained networks are equivalent.

Actual retraining changes many predictions: the three CIFAR-100 history rules disagree on approximately 13% of predictions under S2-TTA, while MGC and cumulative training disagree on 22.7–26.2% across protocols. Similar accuracy means that corrected and newly introduced errors approximately balance, not that the models learn identical representations. The Stage-2 attentive head and fitted structured goodness readouts check whether this similarity is peculiar to simple summation, but their outcomes remain empirical findings.

The readout boundary matters: in the appendix, a block-3-only goodness logistic probe spans 8.40 percentage points across the three history rules, while 10-NN on raw goodness spans 8.58 percentage points. The accuracy-dominance conclusion is therefore limited to the deployed rule and tested structured readouts, not every classifier capable of extracting information from the features.

Loss & Training

The total block loss also contains per-aspect goodness objectives, cumulative discrimination, the current-block auxiliary, depth ordering, supervised contrastive learning, token reconstruction, and MoE router regularization. The attenuation theorem analyzes only cumulative discrimination. The implemented depth-order term compares the current block's cumulative loss score with the preceding block's own goodness; the appendix's idealized hard-constraint proposition instead compares consecutive cumulative scores. These are different constraints, so the latter cannot guarantee monotonic margin growth for the implemented training rule.

Hard-negative mining uses an EMA teacher to select the highest-goodness candidate among sampled incorrect classes. Sampling is with replacement: the CIFAR-10 schedule 8→16 counts draws, not distinct classes. NL and NI discrimination streams are blended with equal default weights of 0.5; neither repair updates upstream parameters through history.

The main CIFAR-100 control uses dense L4/D256, batch size 256, 362+20 epochs, and 3 seeds. The three history rules share the depth-scaled auxiliary, which MGC replaces. The CIFAR-10 L4/D128 MGC control uses 180+10 epochs, SAM enabled, and 3 seeds per arm. The larger CIFAR-10 accuracy calibration uses an MoE recipe; results at different scales must not be combined into one controlled experiment.

The locality audit reveals an implementation detail: in the reported runs, the depth-order term could generate a nonzero gradient on the preceding block's mixing logits, but its optimizer cleared that gradient before applying it, leaving actual updates block-local. The released implementation also places this path under stop-gradient; the authors' small-scale check obtained bit-identical parameters. The precise claim is applied-update locality, not that every historical implementation's computational graph was free of cross-block gradients.

Key Experimental Results

Main Results

The following table comes from source Table 3, the central CIFAR-100 control. sep is the deepest block's own wrong-label margin at the last Stage-1 epoch; accuracy is evaluated at validation-selected EMA checkpoints. The separation measurement must not be described as coming from the selected checkpoint. Accuracy is in %, and \(\pm\) denotes sample standard deviation over 3 seeds; TTA means horizontal-flip test-time augmentation.

History rule / objective sep (L3) S1 single crop S1-TTA S2 single crop (probe) S2-TTA (probe)
Cumulative, \(\gamma=0.7\) \(0.96\pm0.01\) \(66.78\pm0.28\) \(67.41\pm0.35\) \(68.51\pm0.51\) \(68.96\pm0.43\)
History-free, \(\gamma=0\) \(4.78\pm0.02\) \(66.37\pm0.20\) \(66.93\pm0.14\) \(68.54\pm0.25\) \(69.23\pm0.34\)
Hardness-gated, \(\kappa=0\) \(1.71\pm0.01\) \(66.35\pm0.41\) \(67.35\pm0.66\) \(68.44\pm0.51\) \(69.07\pm0.61\)
MGC \(5.61\pm0.04\) \(65.73\pm0.35\) \(66.58\pm0.25\) \(68.24\pm0.51\) \(69.09\pm0.27\)

All four-protocol paired 95% CIs for comparisons among the three history-rule arms lie inside \(\pm1\) percentage point, but this does not mean no difference: history-free minus cumulative has an S1 paired difference of −0.42, with CI [−0.76, −0.07], detecting a small negative effect. Differences between displayed rounded means can differ by 0.01 from the source's unrounded paired differences; this note preserves the reported paired results.

MGC raises deepest-block sep from 0.96 to 5.61, approximately 5.9-fold, but its S1 paired difference is −1.05 percentage points, with 95% CI [−1.46, −0.62]; the S1-TTA difference is −0.83, [−1.23, −0.42]. Neither CI lies entirely inside \(\pm1\), so “all repairs stay within 1 percentage point” is not an accurate blanket statement. S2 and S2-TTA differences are −0.27, [−0.68, +0.14], and +0.13, [−0.27, +0.53], respectively. These support recovery of the accuracy cost by a frozen-feature probe, not recovery by pure FF.

The bootstrap CIs use matched-seed test predictions and are conditional on the trained checkpoints; they are not generalization guarantees over all possible training runs. In the history-rule trio, 5/9 runs were resumed, as was one MGC seed. The authors disclose the associated log coverage and continuation limitations.

Ablation Study

The next table comes from appendix Table 6 and tests attribution of MGC's gain. All four arms retain \(\gamma=0.7\), CIFAR-10 L4/D128, SAM enabled, 180+10 epochs, and 3 seeds. Accuracy is in %, and sep remains the last-epoch own-block margin.

Config sep (L3) S1 single crop S2-TTA (probe) Note
Cumulative + depth-scaled auxiliary \(0.75\pm0.01\) \(85.74\pm0.37\) \(87.04\pm0.51\) SAM-matched baseline for MGC
Cumulative, no auxiliary or compensation \(4.88\pm0.05\) \(85.31\pm0.10\) \(86.90\pm0.22\) Auxiliary removal supplies most of the gain
MGC \(5.25\pm0.05\) \(85.45\pm0.37\) \(87.27\pm0.29\) Compensation replaces the auxiliary
MGC + retained block-0 auxiliary \(5.28\pm0.03\) \(85.32\pm0.19\) \(86.93\pm0.09\) No evidence that the block-0 change masks an accuracy gain

The separation gain from baseline to MGC is 4.50: auxiliary removal contributes 4.13 (about 92%), and the compensator adds 0.37 (about 8%). The authors' preregistered prediction that the auxiliary-free cumulative arm would remain below 4.72 failed. These results do not validate a compensation-dominant explanation.

The auxiliary-free and auxiliary-bearing cumulative arms ran on different GPU/software stacks, fully confounding the auxiliary-removal contrast with those stacks. The 92%/8% split is therefore the authors' empirical decomposition of these results, not a universal causal fraction free of confounding. CIFAR-100 lacks the corresponding auxiliary-removal control, so the same split cannot be applied to its MGC results.

For effect-size calibration, source Table 2 also provides the following CIFAR-10 comparisons. The first four rows are 3-seed means, while the last uses only seed 42. The strong-augmentation recipes differ between FF and BP-strong, so this is not an augmentation-matched comparison of the FF and BP algorithms.

Backbone Training / augmentation Single-crop test accuracy (%)
Plain CNN, 1.15 M parameters BP / strong \(90.08\pm0.28\)
Plain CNN, 1.15 M parameters FF / FF-pipeline \(29.85\pm0.50\)
Stripped FF backbone, 5.54 M parameters BP / weak \(87.43\pm0.62\)
Stripped FF backbone, 5.54 M parameters BP / BP-strong \(93.85\pm0.18\)
Stripped FF backbone, 6.74 M parameters FF / FF-pipeline 89.03

Key Findings

  • A real mechanism, not a universal explanation: in the squared-hinge replication, history removal raises deepest-block sep from 1.67 to 4.44, while S2-TTA changes from 87.08% to 86.93%. This extends the barrier evidence, not the claim to every goodness statistic.
  • Depth amplifies health differences, not automatically accuracy gains: for matched seed 42 at CIFAR-10 L8/D128, gated sep is 6.94 versus LCFF's 0.15, while S2-TTA is only 86.84% versus 86.79%. The gated 3-seed mean of 87.12% must not replace this matched single-seed comparison.
  • Other depth problems exist beyond historical attenuation: on Tiny ImageNet, history-free own-block sep peaks at block 1 and then falls, while accuracy rises with depth. Three-seed S1 single-crop accuracy is \(48.54\pm0.19\%\), and S2-TTA is \(52.32\pm0.34\%\); that comparison changes both readout and TTA.
  • Recipe calibration moves accuracy more than local repair: the plain CNN's FF–BP gap is approximately 60.2 percentage points, and strong versus weak augmentation changes BP on the stripped backbone by 6.42 percentage points. These are pipeline-specific effect sizes, not independently additive architecture and augmentation contributions.

Highlights & Insights

  • Separate diagnosis from the target metric: low cumulative loss can hide weak own-block discrimination, while the exact ratio and own margin provide complementary views. Accuracy under a clearly specified readout still determines whether a repair helps, preventing a health proxy from replacing task performance.
  • A negative result with a testable mechanism: the paper proves a local gradient relationship and intervenes on it rather than merely reporting no gain. MGC and auxiliary-removal controls also force a narrower attribution, which is more informative than crediting every separation improvement to compensation.
  • Treat readout as a research object: similar accuracy can coexist with substantial prediction churn, indicating changes in representation and error distribution. How to extract complementary information deserves further study, but the existing fitted goodness-weight experiments do not create large between-arm accuracy differences.

Limitations & Future Work

  • The theorem characterizes a local surrogate under upstream stop-gradient, without global convergence, sample-complexity, or generalization bounds. Actual total gradients also depend on Jacobians, auxiliary terms, and cancellation across examples.
  • The negative result covers softplus, one reduced-scale squared-hinge replication, and the paper's goodness head, backbones, and datasets. ImageNet scale, classic FF architectures, and substantially different local objectives remain untested.
  • Stage 2 is a back-propagation-trained frozen-feature probe, not pure FF accuracy. Broader probes already show differences above 1 percentage point, ruling out extrapolation to invariant accuracy under all readouts.
  • CIFAR-100's MGC benefits and costs have not been decomposed into auxiliary removal and compensation, while the CIFAR-10 control has hardware/software-stack confounding. Same-hardware, fully factorial multi-seed controls and trajectory-level gradient/margin measurements should precede stronger attribution.
  • The depth-scaled auxiliary has the best single-seed result in the 8-block, 32-expert configuration, leaving open whether it becomes useful for accuracy at greater depth. Tiny ImageNet's history-free depth deterioration also lies outside the present theorem.
  • vs Hinton's FF: the paper retains positive/negative goodness and block-local learning, but uses attention backbones, multi-aspect goodness, and class injection at every block. Its analysis concerns cumulative training objectives, not a proof of free-riding in every FF variant.
  • vs Layer Collaboration: that work emphasizes insufficient information sharing across layers; this paper identifies the opposite failure direction, where excessive inherited positive margin saturates the local objective. Collaboration is not a binary choice: sharing representations differs from sharing an already-satisfied objective.
  • vs ASGE, DeeperForward, and SCFF: those methods improve performance through goodness, architecture, or training recipes, whereas this paper separates block-health repair from accuracy gain. Cross-work accuracy is only calibration, not a controlled SOTA ranking, because supervision, parameter counts, augmentation, and readouts differ.
  • A follow-up question: under fixed parameter budgets and a shared hardware stack, jointly varying residual weights, auxiliary strength, and history rules, then testing readouts that use complementary cross-block information, could explain why health improvements fail to become accuracy gains. This is an untested direction, not a demonstrated benefit of the paper.

Rating

  • Novelty: 4/5, the combination of exact local-gradient analysis and mechanism–accuracy separation is distinctive.
  • Experimental Thoroughness: 4/5, the central 3-seed control, second barrier, and counterfactual controls are substantial, but scale and hardware confounding limit extrapolation.
  • Writing Quality: 4/5, probe boundaries and failed predictions are disclosed clearly; the long appendix and protocol differences require careful checking.
  • Value: 4/5, it offers reusable local-objective debugging principles and constrains over-attribution of FF's accuracy bottleneck.