Skip to content

Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning

Conference: ECCV 2026
Paper: ECCV Official Page
Code: https://github.com/asimukaye/celm
Area: AI Safety
Keywords: Federated Learning, Contribution Estimation, Logit Maximization, Label Skew, Collaborative Fairness

TL;DR

Addressing the challenge of evaluating client contributions under severe non-IID and label skew conditions without accessing raw data, validation sets, or self-reported metadata, this paper proposes CELM (Contribution Estimation from Logit Maximization), which probes client models via class-wise logit maximization, de-biases evidence against a global baseline, and normalizes class shares to fairly reward informative rare-class holders (Mavericks) while suppressing non-contributing free-riders.

Background & Motivation

In cross-silo federated learning (FL) systems across distributed institutions, preserving private data privacy and adhering to regulatory compliance are paramount requirements. However, practical deployments frequently exhibit severe statistical heterogeneity across clients, characterized by disparate dataset volumes, partial class coverage, and pathological label skew. Standard aggregation protocols such as FedAvg rely heavily on client self-reported sample quantities, implicitly assuming that all samples contribute equally to global generalization. Under extreme non-IID partitions, this causes the global model to overfit dominant clients holding large majority-class subsets, severely degrading performance on underrepresented minority classes.

Accurately and fairly estimating each client's actual contribution without relying on raw local data is crucial for robust aggregation and collaborative fairness, yet existing paradigms face structural limitations. Client-reported metadata (e.g., sample counts or label histograms) is susceptible to strategic manipulation when incentives or influence are at stake. Server-side validation datasets (as employed in CFFL or FedCE) are often unavailable in privacy-restricted domains, and any inherent distributional mismatch in the auxiliary dataset directly distorts contribution scores. Conversely, recent data-free methods based on gradient similarity (such as CGSV) operate under a "similarity-to-average" premise; when confronted with informative clients that possess rare or unique classes ("Mavericks"), their atypical gradient trajectories are actively penalized, collapsing rare-class accuracy down to near-zero.

This paper addresses this fundamental tension by circumventing parameter-space average alignment and instead directly probing the class-discriminative representations synthesized inside local model parameters. Core idea: repurpose activation maximization into a server-side black-box probe that performs class-wise logit maximization on client models, subtracts a global baseline to de-bias general confidence shifts, normalizes per-class shares across clients to fairly evaluate rare-class competence, and freezes contribution weights after an early warm-up window to guarantee aggregation stability and minimal computational overhead.

Method

Overall Architecture

CELM operates within standard federated learning communication rounds and decouples the training lifecycle into an early warm-up phase (\(t \le T_w\)) and a subsequent freeze phase (\(t > T_w\)). During warm-up, upon receiving client models, the server performs class-wise gradient ascent in the synthetic input space to maximize class-specific logits under \(\ell_2\) regularization, extracting raw evidence scores. It then computes a de-biasing baseline from the prior global model, constructs a de-biased client-class evidence matrix via ReLU suppression, and normalizes evidence class-wise across clients to capture relative class shares. The resulting scores are averaged across classes, projected onto the probability simplex, and smoothed via an exponential moving average (EMA). Once the warm-up threshold is passed, the contribution weights are frozen and reused for all remaining rounds.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Client uploads local model $w_i^{(t)}$"] --> B["Class-wise Logit Maximization Probing"]
    B --> C["Global Baseline De-biasing & Relative Share Normalization"]
    C --> D["Simplex Projection & EMA Dynamic Smoothing"]
    D --> E["Warm-up Decoupling & Weight Freezing"]
    E --> F["Weighted Aggregation into Global Model $w_g^{(t)}$"]

Key Designs

1. Class-wise Logit Maximization Probing: Extracting Local Discriminative Evidence

To circumvent the absence of validation data and the vulnerability of self-reported metadata, CELM repurposes input-space activation maximization into a quantitative model competence probe. If client \(i\) has thoroughly trained on class \(c\), its model parameters encode discriminative features that trigger high logit responses for synthetic inputs optimized from standard Gaussian noise. For each client model \(w_i^{(t)}\) and class \(c \in \{1, \dots, K\}\), the server optimizes an input \(x\) via Adam gradient ascent with \(\ell_2\) regularization: $\(x_{i,c}^{\star,(t)} = \arg\max_{x \in \mathcal{X}} s_c(x; w_i^{(t)}) - \lambda \|x\|_2^2\)$ where \(s_c(x; w)\) denotes the pre-softmax logit for class \(c\), and \(\lambda\) suppresses high-frequency noise and unbounded pixel growth. The resulting raw score is \(\tilde{q}_{i,c}^{(t)} = s_c(x_{i,c}^{\star,(t)}; w_i^{(t)})\). All class probes run embarrassingly parallel on the server and yield purely synthetic, non-semantic visual patterns that do not compromise client training data privacy.

2. Global Baseline De-biasing & Relative Share Normalization: Eliminating Calibration Drift and Protecting Rare Classes

Raw logit magnitudes cannot be compared naively across clients due to global calibration shifts and differing optimization dynamics. CELM introduces a reference baseline derived from the previous global model \(w_g^{(t-1)}\) averaged over all classes: $\(b^{(t-1)} = \frac{1}{K} \sum_{c=1}^K s_c(x_{g,c}^{\star,(t-1)}; w_g^{(t-1)})\)$ Subtracting this generic confidence level followed by ReLU truncation yields the de-biased evidence score \(q_{i,c}^{(t)} = \max(0, \tilde{q}_{i,c}^{(t)} - b^{(t-1)})\), forming the evidence matrix \(Q^{(t)} \in \mathbb{R}^{N \times K}\). Crucially, to prevent clients dominating frequent classes from drowning out those specializing in rare classes, CELM computes relative class share across clients along each class column: $\(r_{i,c}^{(t)} = \frac{q_{i,c}^{(t)}}{\sum_{j=1}^N q_{j,c}^{(t)} + \epsilon}\)$ By evaluating competence on a per-class basis, an informative Maverick client that uniquely holds a rare class achieves \(r_{i,c}^{(t)} \approx 1.0\) for that class, safeguarding its contribution weight from majority dilution.

3. Simplex Projection & EMA Dynamic Smoothing: Dampening Round-to-Round Volatility

A client's instantaneous raw contribution score is computed as the unweighted mean over all class shares: \(\hat{c}_i^{(t)} = \frac{1}{K} \sum_{c=1}^K r_{i,c}^{(t)}\). To maintain the convex combination property for global model aggregation, the vector is normalized over the probability simplex: \(\bar{c}_i^{(t)} = \hat{c}_i^{(t)} / \sum_{j=1}^N \hat{c}_j^{(t)}\). Because single-round SGD updates and input-space optimization introduce high-frequency fluctuations, CELM incorporates exponential moving average (EMA) smoothing with momentum parameter \(\beta \in [0, 1)\) (set to 0.5): $\(c_i^{(t)} = \beta c_i^{(t-1)} + (1 - \beta)\bar{c}_i^{(t)}\)$ This dampens gradient noise and guarantees a stable trajectory during global parameter aggregation.

4. Warm-up Decoupling & Weight Freezing: Preserving Specificity and Limiting Overhead

As federated aggregation progresses across rounds, client models assimilate global parameters, causing local class-specific logit variations to diminish while server probing costs accumulate. CELM establishes a warm-up horizon spanning only the initial \(5\%\) of rounds (\(T_w = 0.05 T\)). During warm-up, the server broadcasts only the aggregated backbone while clients retain their local classification heads, preserving local label sensitivity during probing. For all subsequent rounds (\(t > T_w\)), the contribution weights are frozen as \(c_i^{(t)} = c_i^{(T_w)}\), and the full global model is broadcast. This bounds additional server compute strictly to early training and reverts post-warm-up overhead to that of standard FedAvg.

Key Experimental Results

Main Results

Evaluations were conducted across FashionMNIST (4-layer MLP), CIFAR-10 (5-layer CNN), and the real-world clinical skin lesion benchmark FedISIC (ViT-B/16), spanning Pure Label Skew (PLS), Step Label Skew (SLS), and Dirichlet splits (\(\alpha \in \{0.01, 0.05, 0.10\}\)).

Dataset Split FedAvg CFFL CGSV ShapFed CELM (Ours) Advantage / Gain
FashionMNIST Dir. (\(\alpha=0.01\)) \(80.64 \pm 2.25\) \(78.89 \pm 2.93\) \(47.40 \pm 7.82\) \(81.15 \pm 1.83\) 81.76 ± 1.85 +0.61% over ShapFed
Dir. (\(\alpha=0.05\)) \(81.63 \pm 2.69\) \(81.89 \pm 0.45\) \(46.30 \pm 3.34\) \(81.97 \pm 2.61\) 83.64 ± 0.42 +1.67% over 2nd best with lower variance
Dir. (\(\alpha=0.10\)) \(84.83 \pm 0.62\) \(84.58 \pm 0.64\) \(63.37 \pm 3.71\) \(85.14 \pm 0.65\) 85.34 ± 0.09 Consistent top performer
PLS \(80.62 \pm 0.52\) \(80.95 \pm 1.81\) \(49.41 \pm 2.17\) \(81.80 \pm 0.56\) 83.70 ± 0.19 +1.90% over ShapFed
SLS \(83.13 \pm 2.61\) 88.13 ± 0.35 \(55.39 \pm 5.06\) \(84.17 \pm 2.07\) 87.00 ± 1.14 Highest among all data-free methods
CIFAR-10 Dir. (\(\alpha=0.01\)) \(63.41 \pm 1.42\) \(53.59 \pm 6.74\) \(18.25 \pm 6.09\) \(63.12 \pm 1.39\) 64.37 ± 0.98 Strongest robustness under extreme skew
Dir. (\(\alpha=0.05\)) \(65.12 \pm 1.89\) \(55.92 \pm 9.31\) \(24.48 \pm 1.57\) \(65.83 \pm 1.71\) 67.16 ± 1.29 +2.04% over FedAvg
Dir. (\(\alpha=0.10\)) \(68.98 \pm 0.67\) \(60.41 \pm 4.85\) \(29.03 \pm 3.16\) \(68.98 \pm 0.42\) 69.21 ± 0.76 Consistent top accuracy
PLS \(56.82 \pm 2.96\) \(49.93 \pm 6.28\) \(28.75 \pm 4.92\) \(56.80 \pm 2.80\) 59.11 ± 1.96 +2.29% over FedAvg
SLS \(67.40 \pm 0.95\) \(70.47 \pm 2.10\) \(33.58 \pm 1.76\) \(69.91 \pm 0.43\) 71.96 ± 0.08 Beats validation-based CFFL
FedISIC Natural clinical split \(61.25 \pm 0.04\) \(62.60 \pm 6.34\) \(26.17 \pm 0.59\) \(62.31 \pm 0.16\) 70.18 ± 0.58 Balanced Acc. +7.58% over CFFL

Maverick Performance & Ablation Analysis

To verify whether CELM correctly values rare-class holders, the authors evaluated Maverick splits (where specific clients exclusively possess rare classes) alongside warm-up schedule ablations.

Predictive Performance under Maverick Split:

Dataset Metric FedAvg CFFL CGSV ShapFed CELM (Ours)
FashionMNIST Balanced Acc. \(84.79 \pm 0.24\) \(84.39 \pm 2.43\) \(50.74 \pm 0.23\) \(84.87 \pm 0.32\) 87.32 ± 0.10
Rare Class Acc. \(81.76 \pm 0.68\) \(82.99 \pm 8.56\) \(0.00 \pm 0.00\) \(82.02 \pm 0.81\) 90.77 ± 0.24
CIFAR-10 Balanced Acc. \(64.75 \pm 0.29\) \(62.10 \pm 1.15\) \(42.12 \pm 4.17\) \(64.14 \pm 0.27\) 68.60 ± 0.53
Rare Class Acc. \(39.95 \pm 1.06\) \(44.69 \pm 18.06\) \(0.00 \pm 0.00\) \(38.16 \pm 1.13\) 61.21 ± 0.59

Warm-up Horizon \(T_w / T\) Ablation (Test Accuracy %):

Dataset Split \(T_w=5\%\) (Default) \(T_w=10\%\) \(T_w=15\%\) \(T_w=20\%\) \(T_w=30\%\)
FashionMNIST Dir. (\(\alpha=0.05\)) 83.64 ± 0.42 \(83.31 \pm 0.99\) \(83.07 \pm 0.95\) \(83.06 \pm 0.75\) \(82.70 \pm 0.91\)
PLS \(83.70 \pm 0.19\) \(83.57 \pm 0.45\) \(83.73 \pm 0.34\) \(83.99 \pm 0.10\) 84.14 ± 0.51
CIFAR-10 Dir. (\(\alpha=0.05\)) 67.16 ± 1.29 \(66.88 \pm 1.54\) \(66.68 \pm 1.49\) \(66.54 \pm 1.43\) \(66.43 \pm 1.49\)
PLS \(59.11 \pm 1.96\) \(59.03 \pm 2.32\) \(59.36 \pm 1.42\) \(59.44 \pm 0.95\) 59.86 ± 0.57

Key Findings

  • Dramatic Rare-Class Gain: Under Maverick splits, CGSV collapses completely (rare-class accuracy drops to \(0.00\%\)), while CELM elevates rare-class accuracy on FashionMNIST and CIFAR-10 to \(90.77\%\) and \(61.21\%\) respectively (outperforming FedAvg by \(+9.01\%\) and \(+21.26\%\)), demonstrating the effectiveness of class-relative share estimation.
  • Superior Clinical Generalization: On the real-world FedISIC clinical dataset exhibiting intrinsic class imbalance, CELM achieves \(70.18\%\) balanced accuracy, outpacing the validation-based CFFL by \(+7.58\%\).
  • Fast and Frugal Warm-up Horizon: A minimal \(5\%\) warm-up window (\(T_w = 0.05 T\)) suffices to establish reliable client contribution vectors. Longer warm-ups yield diminishing returns while incurring unnecessary server compute.

Highlights & Insights

  • Repurposing Feature Inversion as a Zero-Data Auditor: Instead of using activation maximization purely for feature interpretability, CELM adapts it as a server-side probe that reverse-engineers class-discriminative evidence without inspecting raw training images.
  • De-biasing via Global Reference Model: Subtracting average logit activations from the preceding global model effectively cancels out systemic model overconfidence and cross-client calibration discrepancies.
  • Orthogonal to Local Client Optimization: Because all modifications reside strictly on the aggregation server, CELM seamlessly integrates with client-side loss modifications such as FedRS (Restricted Softmax), yielding further compounding improvements (CELM(RS)).

Limitations & Future Work

  • Server Compute Complexity Scales with Class Count: The warm-up phase requires \(K \times (N+1)\) backward passes per round; scaling to datasets with hundreds or thousands of classes will require dynamic class sub-sampling or hierarchical probing.
  • Sensitivity to Miscalibrated Architectures: The fidelity of evidence estimation relies on reasonable model calibration; severely overconfident or temperature-distorted local networks could skew relative evidence scores.
  • vs FedAvg: FedAvg relies on self-reported sample quantities, making it vulnerable to false reporting and unable to account for qualitative label skew; CELM eliminates metadata dependence by evaluating true parameter capability.
  • vs CGSV & ShapFed: CGSV rewards alignment with the cohort average gradient, severely penalizing informative outlier clients (Mavericks); ShapFed attempts class-wise alignment in gradient space but remains susceptible to noise. CELM operates directly on class output logits, providing robust rare-class preservation.

Rating

  • Novelty: ⭐⭐⭐⭐☆ (Creative adaptation of activation maximization to data-free federated client valuation)
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ (Extensive evaluation across synthetic splits, clinical benchmarks, Maverick setups, and free-rider detection)
  • Writing Quality: ⭐⭐⭐⭐⭐ (Crisp mathematical formulation and coherent narrative flow)
  • Value: ⭐⭐⭐⭐☆ (Offers an actionable, validation-free mechanism for collaborative fairness and robust federated aggregation)