Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/asimukaye/celm
Area: AI Safety / Federated Learning
Keywords: Federated learning, contribution estimation, logit maximization, label skew, collaborative fairness
TL;DR¶
Addressing the challenge of evaluating client contributions under severe non-IID label skew without raw data, auxiliary validation sets, or self-reported metadata, CELM probes client models via server-side class-level logit maximization, builds a debiased evidence matrix with relative share normalization, and accurately rewards rare-class holders (Mavericks) while penalizing free-riders.
Background & Motivation¶
In cross-silo federated visual learning across institutions, protecting private data and maintaining regulatory compliance are non-negotiable requirements. However, client datasets in real-world deployments exhibit extreme statistical heterogeneity, characterized by unequal sample sizes, incomplete category coverage, and severe label distribution skew. Standard federated aggregation protocols, such as classic FedAvg weighting models strictly by local sample counts, implicitly assume that every sample contributes equally to global optimization. Under heavy heterogeneity, updates from clients possessing large majority-class subsets drown out minority-class signals, causing global model generalization to collapse on tail and rare categories.
Fairly estimating client contributions to guide aggregation weighting is critical, yet conventional approaches encounter major practical hurdles. Relying on self-reported client metadata (such as sample counts or label histograms) invites strategic manipulation and adversarial falsification. Introducing server-side validation sets (e.g., CFFL, FedCE) is infeasible in privacy-critical domains and injects the validation set's own distribution bias. While recent data-free gradient alignment methods (e.g., CGSV) attempt to circumvent these limitations, they rely on a flawed "similarity-to-average" premise: updates aligned with the population mean gradient are deemed beneficial, whereas high-value clients holding unique or rare classes ("Mavericks") are penalized for diverging from the majority direction, often driving rare-class accuracy to zero.
To overcome this dilemma, the model inspection must sidestep parameter-space averaging and probe class-discriminative knowledge encoded directly within the trained neural parameters. Core idea: adapt activation maximization as a server-side black-box diagnostic probe, optimizing synthetic inputs to maximize class-specific logits as evidence of local learning, normalizing relative class shares across clients to reward rare-class ownership, and freezing weights after an initial warm-up window to guarantee aggregation stability and minimal computation.
Method¶
Overall Architecture¶
CELM (Contribution Estimation from Logit Maximization) operates directly within standard federated communication rounds, partitioned into an early warm-up phase (\(t \le T_w\)) and a subsequent frozen phase (\(t > T_w\)). During warm-up, upon receiving client weights, the server synthesizes inputs via constrained gradient ascent to maximize unnormalized logits for each class, extracting raw evidence scores. These scores are debiased against the previous global model's baseline response and assembled into a client-by-class evidence matrix. Next, relative evidence shares are normalized along the class dimension, averaged across classes, and projected onto the probability simplex with Exponential Moving Average (EMA) smoothing. Once warm-up concludes, the resulting contribution weights are frozen for all subsequent communication rounds.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Clients upload local model weights $w_i^{(t)}$"] --> B["Class-level Logit Maximization Probing"]
B --> C["Global Baseline Debiasing & Relative Share Normalization"]
C --> D["Simplex Projection & EMA Smoothing"]
D --> E["Warm-up Decoupling & Weight Freezing"]
E --> F["Weighted Aggregation yields Global Model $w_g^{(t)}$"]
Key Designs¶
1. Class-Level Logit Maximization Probing: Extracting Class Discrimination Without Data Lacking access to private data and untrusting client metadata, CELM transforms feature visualization into an evidence probe. A client that has thoroughly trained on class \(c\) encodes strong discriminative pathways, allowing a synthetic input initialized from Gaussian noise to be optimized rapidly to a high logit value. The server performs \(\ell_2\)-regularized gradient ascent for each client model \(w_i^{(t)}\) and class \(c \in \{1,\dots,K\}\): $\(x_{i,c}^{\star,(t)} = \arg\max_{x \in \mathcal{X}} s_c(x; w_i^{(t)}) - \lambda \|x\|_2^2\)$ where \(s_c(x; w)\) is the unnormalized logit for class \(c\), and \(\lambda\) penalizes excessive pixel magnitude and high-frequency artifacts. The resulting response forms raw evidence \(\tilde{q}_{i,c}^{(t)} = s_c(x_{i,c}^{\star,(t)}; w_i^{(t)})\). All probes execute in parallel on the server and yield abstract synthetic patterns that reveal zero private image semantics.
2. Global Baseline Debiasing and Relative Share Normalization: Protecting Rare-Class Mavericks Raw logits cannot be compared directly across clients due to confidence calibration drift across architectures and optimization states. CELM establishes a global background baseline from the previous global model \(w_g^{(t-1)}\): $\(b^{(t-1)} = \frac{1}{K} \sum_{c=1}^K s_c(x_{g,c}^{\star,(t-1)}; w_g^{(t-1)})\)$ Subtracting this baseline and passing through ReLU yields debiased evidence \(q_{i,c}^{(t)} = \max(0, \tilde{q}_{i,c}^{(t)} - b^{(t-1)})\), forming evidence matrix \(Q^{(t)} \in \mathbb{R}^{N \times K}\). To prevent majority-class clients from dominating total numerical volume, CELM normalizes relative shares within each class column: $\(r_{i,c}^{(t)} = \frac{q_{i,c}^{(t)}}{\sum_{j=1}^N q_{j,c}^{(t)} + \epsilon}\)$ This formulation guarantees that rare classes receive equal weight to majority classes: even if a rare class is held by only one Maverick client, its relative share \(r_{i,c}^{(t)} \approx 1.0\), preventing majority-rule gradient suppression.
3. Simplex Projection and Exponential Moving Average: Suppressing Optimization Variance A client's instantaneous score is computed by averaging relative shares across all classes: \(\hat{c}_i^{(t)} = \frac{1}{K} \sum_{c=1}^K r_{i,c}^{(t)}\), followed by probability simplex projection \(\bar{c}_i^{(t)} = \hat{c}_i^{(t)} / \sum_{j=1}^N \hat{c}_j^{(t)}\). To counteract transient perturbations from local mini-batch sampling and logit optimization noise, CELM updates weights via exponential moving average with momentum factor \(\beta \in [0, 1)\) (default 0.5): $\(c_i^{(t)} = \beta c_i^{(t-1)} + (1 - \beta)\bar{c}_i^{(t)}\)$ This smoothing mechanism stabilizes weight trajectories across rounds.
4. Warm-Up Decoupling and Weight Freezing: Preserving Discrimination at Minimal Cost As global aggregation rounds progress, client parameters homogenize and global priors attenuate local discriminative signatures, while continual probing incurs recurring computation. CELM restricts logit probing strictly to the initial \(5\%\) of communication rounds (\(T_w = 0.05 T\)). During warm-up, the server broadcasts the global backbone while allowing clients to retain private local heads, accentuating head sensitivity to local distributions. Once \(t > T_w\), weights are permanently frozen (\(c_i^{(t)} = c_i^{(T_w)}\)) and full models including heads are synchronized, reducing post-warmup compute cost to that of standard FedAvg.
Key Experimental Results¶
Main Benchmarks¶
Evaluated across FashionMNIST (4-layer MLP), CIFAR-10 (5-layer CNN), and real-world clinical dermoscopy benchmark FedISIC (ViT-B/16) under Pure Label Skew (PLS), Step Label Skew (SLS), and Dirichlet distributions (\(\alpha \in \{0.01, 0.05, 0.10\}\)).
Table 1: Benchmark test accuracy comparison across heterogeneous label skew partitions | Dataset | Partition | FedAvg | CFFL | CGSV | ShapFed | CELM (Ours) | Relative Gain / Observation | |---|---|---|---|---|---|---|---| | FashionMNIST | Dir. (\(\alpha=0.01\)) | \(80.64 \pm 2.25\) | \(78.89 \pm 2.93\) | \(47.40 \pm 7.82\) | \(81.15 \pm 1.83\) | 81.76 ± 1.85 | +0.61% over ShapFed | | | Dir. (\(\alpha=0.05\)) | \(81.63 \pm 2.69\) | \(81.89 \pm 0.45\) | \(46.30 \pm 3.34\) | \(81.97 \pm 2.61\) | 83.64 ± 0.42 | +1.67% over runner-up, lower variance | | | Dir. (\(\alpha=0.10\)) | \(84.83 \pm 0.62\) | \(84.58 \pm 0.64\) | \(63.37 \pm 3.71\) | \(85.14 \pm 0.65\) | 85.34 ± 0.09 | Consistently top performance | | | PLS | \(80.62 \pm 0.52\) | \(80.95 \pm 1.81\) | \(49.41 \pm 2.17\) | \(81.80 \pm 0.56\) | 83.70 ± 0.19 | +1.90% over ShapFed | | | SLS | \(83.13 \pm 2.61\) | 88.13 ± 0.35 | \(55.39 \pm 5.06\) | \(84.17 \pm 2.07\) | 87.00 ± 1.14 | Outperforms all data-free baselines | | CIFAR-10 | Dir. (\(\alpha=0.01\)) | \(63.41 \pm 1.42\) | \(53.59 \pm 6.74\) | \(18.25 \pm 6.09\) | \(63.12 \pm 1.39\) | 64.37 ± 0.98 | Substantially outperforms all baselines | | | Dir. (\(\alpha=0.05\)) | \(65.12 \pm 1.89\) | \(55.92 \pm 9.31\) | \(24.48 \pm 1.57\) | \(65.83 \pm 1.71\) | 67.16 ± 1.29 | +2.04% over FedAvg | | | Dir. (\(\alpha=0.10\)) | \(68.98 \pm 0.67\) | \(60.41 \pm 4.85\) | \(29.03 \pm 3.16\) | \(68.98 \pm 0.42\) | 69.21 ± 0.76 | Consistent competitive lead | | | PLS | \(56.82 \pm 2.96\) | \(49.93 \pm 6.28\) | \(28.75 \pm 4.92\) | \(56.80 \pm 2.80\) | 59.11 ± 1.96 | +2.29% over FedAvg | | | SLS | \(67.40 \pm 0.95\) | \(70.47 \pm 2.10\) | \(33.58 \pm 1.76\) | \(69.91 \pm 0.43\) | 71.96 ± 0.08 | Beats validation-based CFFL | | FedISIC | Clinical Non-IID | \(61.25 \pm 0.04\) | \(62.60 \pm 6.34\) | \(26.17 \pm 0.59\) | \(62.31 \pm 0.16\) | 70.18 ± 0.58 | +7.58% balanced accuracy gain |
Maverick Performance and Ablation Analysis¶
Table 2: Balanced accuracy and rare-class accuracy under Maverick settings (%) | Dataset | Metric | FedAvg | CFFL | CGSV | ShapFed | CELM (Ours) | |---|---|---|---|---|---|---| | FashionMNIST | Balanced Acc. | \(84.79 \pm 0.24\) | \(84.39 \pm 2.43\) | \(50.74 \pm 0.23\) | \(84.87 \pm 0.32\) | 87.32 ± 0.10 | | | Rare Class Acc. | \(81.76 \pm 0.68\) | \(82.99 \pm 8.56\) | \(0.00 \pm 0.00\) | \(82.02 \pm 0.81\) | 90.77 ± 0.24 | | CIFAR-10 | Balanced Acc. | \(64.75 \pm 0.29\) | \(62.10 \pm 1.15\) | \(42.12 \pm 4.17\) | \(64.14 \pm 0.27\) | 68.60 ± 0.53 | | | Rare Class Acc. | \(39.95 \pm 1.06\) | \(44.69 \pm 18.06\) | \(0.00 \pm 0.00\) | \(38.16 \pm 1.13\) | 61.21 ± 0.59 |
Table 3: Sensitivity to warm-up ratio \(T_w/T\) on test accuracy (%) | Dataset | Partition | \(T_w=5\%\) (Default) | \(T_w=10\%\) | \(T_w=15\%\) | \(T_w=20\%\) | \(T_w=30\%\) | |---|---|---|---|---|---|---| | FashionMNIST | Dir. (\(\alpha=0.05\)) | 83.64 ± 0.42 | \(83.31 \pm 0.99\) | \(83.07 \pm 0.95\) | \(83.06 \pm 0.75\) | \(82.70 \pm 0.91\) | | | PLS | \(83.70 \pm 0.19\) | \(83.57 \pm 0.45\) | \(83.73 \pm 0.34\) | \(83.99 \pm 0.10\) | 84.14 ± 0.51 | | CIFAR-10 | Dir. (\(\alpha=0.05\)) | 67.16 ± 1.29 | \(66.88 \pm 1.54\) | \(66.68 \pm 1.49\) | \(66.54 \pm 1.43\) | \(66.43 \pm 1.49\) | | | PLS | \(59.11 \pm 1.96\) | \(59.03 \pm 2.32\) | \(59.36 \pm 1.42\) | \(59.44 \pm 0.95\) | 59.86 ± 0.57 |
Highlights & Insights¶
- Feature Inversion as an Honest Diagnostic Probe: Repurposes activation maximization from interpretability and adversarial generation into an active server-side probe that reverse-engineers class support directly from weights.
- Global Baseline Subtracts Systemic Overconfidence: By subtracting the global model's mean logit probe, CELM removes intrinsic architectural biases toward easily separable classes.
- Orthogonal Integration with Local Solvers: Operating entirely on the server, CELM can be combined seamlessly with client-side regularization such as FedRS (Restricted Softmax) to yield enhanced gains.
Limitations & Future Work¶
- Linear Scaling with Class Cardinality: Probing requires \(K \times (N+1)\) gradient ascent passes per warm-up round, creating computational overhead on thousand-class datasets that calls for dynamic class sub-sampling.
- Calibration Dependence: Overfitted or poorly calibrated local models may exhibit logit magnitudes disconnected from genuine generalization, motivating joint temperature scaling.
Related Work & Insights¶
- vs FedAvg: Eliminates reliance on self-reported sample quantities that misrepresent heterogeneous contributions.
- vs CGSV & ShapFed: Avoids average-gradient similarity metrics that suppress rare-class holders, ensuring fair treatment of Maverick participants.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Novel use of class-level activation maximization for federated contribution auditing]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across multiple vision benchmarks, real clinical datasets, and Maverick edge cases]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical derivations, thorough conceptual motivation, and well-structured experiments]
- Practical Value: ⭐⭐⭐⭐☆ [Highly practical and privacy-compliant framework for cross-silo federated systems]