FedDO: Dynamic Client Optimization for Adaptive Federated Learning¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/leafuan/FedDO_code
Area: Reinforcement Learning
Keywords: Federated Learning, Deep Reinforcement Learning, Continuous Control, Scale Heterogeneity, DSAC-T
TL;DR¶
FedDO formulates client coordination in federated learning as a continuous control problem, leveraging a Distributional Soft Actor-Critic with Three Refinements (DSAC-T) agent to dynamically allocate per-client data-usage ratios and integrating low-rank parameterization to scale to thousands of clients while mitigating gradient drift.
Background & Motivation¶
Federated learning (FL) enables distributed edge devices to collaboratively train machine learning models without exposing their raw local data, thereby addressing stringent data privacy and compliance mandates. However, real-world federated deployments are hindered by the compounded challenges of statistical heterogeneity and scale heterogeneity. On the one hand, non-independent and identically distributed (non-IID) client data distributions cause local optimization paths to diverge sharply, resulting in severe gradient drift during global server aggregation. On the other hand, the substantial disparity in dataset sizes across devices allows data-heavy clients to dominate the global optimization trajectory under standard weighted averaging, drowning out the critical, distinct knowledge held by scarce-data nodes and inducing training instability and slow convergence.
To alleviate the adverse impacts of heterogeneity, recent adaptive control research has explored reinforcement learning (RL) to automate client selection and scheduling. Nevertheless, existing RL-FL frameworks typically treat client scheduling as a binary decision—deciding whether to select or reject a client in each round—without considering the continuous tunability of each client's internal data utilization. This binary granularity prevents such algorithms from fine-tuning contribution strength and managing local computation overhead under severe scale imbalances. Furthermore, many prior frameworks rely on on-policy RL algorithms (such as PPO), which suffer from low sample efficiency and demand frequent online environment interactions, frequently degenerating into training instability or premature policy collapse under stochastic and non-stationary reward signals.
The angle of attack in this work is to break away from the discrete participation paradigm by formulating client data-usage ratios as a continuous action space, driven by an off-policy distributional reinforcement learning agent that is both sample-efficient and robust against noise. Core idea: formulate federated client coordination as a continuous control problem where a DSAC-T agent adaptively allocates per-client data-usage ratios based on parameter drift and global model state, paired with a low-rank parameterization technique to prevent action space explosion across massive client cohorts.
Method¶
Overall Architecture¶
The overall architecture of FedDO frames the federated learning server as an autonomous agent that continuously observes system dynamics and emits continuous control actions. At the beginning of round \(t\), the server inspects the global model state and evaluates the parameter drift of each client relative to the global model, constructing a compact, normalized state vector that is fed into the DSAC-T policy network. The policy network outputs a continuous data-usage ratio vector \(a_t = \langle d_t^1, d_t^2, \ldots, d_t^K \rangle\), where \(d_t^k \in [0, 1]\). Each selected client unbiasedly samples a subset of its local dataset corresponding to the assigned ratio, performs local stochastic gradient descent updates, and uploads the refined model parameters back to the server. The server aggregates the received models via sample-weighted averaging to produce the updated global model, evaluates the accuracy gain on a held-out validation set to construct a scalar reward, and refines the reinforcement learning agent using off-policy replay transitions.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Global Model and Client Updates<br/>$w_t$ and $\{w_t^k\}_{k=1}^K$"] --> B["State Construction & Low-Rank Mapping<br/>Parameter drift $o_t^k = \|w_t^k - w_t\|$"]
B --> C["Continuous Data-Usage Ratio Allocation<br/>Dynamic sampling ratios $d_t^k \in [0, 1]$"]
C --> D["Client Subset Sampling & Local Training<br/>$|\mathcal{S}_t^k| = d_t^k |\mathcal{D}_k|$"]
D --> E["Model Aggregation & Validation Evaluation<br/>Update $w_{t+1}$ and compute accuracy delta"]
E --> F["Distributional Soft Actor-Critic Agent<br/>DSAC-T value distribution learning & policy update"]
F -.->|guide next action generation| B
Key Designs¶
1. Continuous Data-Usage Ratio Allocation: balancing statistical utility and gradient drift
To overcome the inability of discrete client selection to modulate individual client contribution intensity and alleviate data-scale imbalance, FedDO defines each client \(k\)'s sampling ratio \(d_t^k \in [0, 1]\) in round \(t\) as a continuous control variable. For a client possessing a private dataset \(\mathcal{D}_k\), the active training subset \(\mathcal{S}_t^k \subseteq \mathcal{D}_k\) is drawn via unbiased random sampling in exact proportion to the assigned ratio. Only clients with \(d_t^k > 0\) form the active participant set \(\mathcal{K}_t\). The server aggregation objective and model parameter update are formalized as:
This continuous action space introduces a flexible trade-off: under severe non-IID conditions, the policy automatically curtails the number of active clients while assigning them larger data-usage budgets to preserve local statistical consistency; conversely, in homogeneous distributions, the agent engages a broader set of clients with lower individual data fractions, achieving extensive sample coverage with minimal computation and communication overhead.
2. Distributional Soft Actor-Critic Agent: resolving off-policy bias and non-stationary noise
To resolve the sample inefficiency and instability of conventional RL in communication-constrained federated environments, FedDO incorporates the Distributional Soft Actor-Critic with Three Refinements (DSAC-T) algorithm. Instead of estimating a scalar expected Q-value, twin critic networks model the full return distribution \(Z_{\theta_i}(s, a) \sim \mathcal{N}(Q_{\theta_i}(s, a), \sigma^2_{\theta_i}(s, a))\), explicitly capturing uncertainty caused by client stochasticity and non-stationary dynamics. When evaluating target values, the critic with the smaller expected value determines the target, while expected-value substitution and variance-aware clipping boundary suppress over-optimistic value propagation:
where \(\bar{i} = \arg\min_{i=1,2} Q_{\bar{\theta}_i}(s_{t+1}, a_{t+1})\). Furthermore, the critic loss is reweighted by a variance factor \(\omega_i = \mathbb{E}_{\mathcal{B}}[\sigma_{\theta_i}^2]\), aligning the update magnitude with the return variance and rendering the training process inherently invariant to arbitrary reward scale bases.
3. Low-Rank Parameterization: mitigating the curse of dimensionality in large-scale cohorts
In practical enterprise federated learning systems, the number of clients \(K\) frequently scales into hundreds or thousands. Because the actor's output dimension and the critic's input dimension grow linearly with \(K\), fully connected layers experience explosive parameter growth, inducing prohibitive computation overhead and gradient instability. FedDO addresses this scalability bottleneck via a low-rank parameterized restructuring strategy. Without altering the underlying DSAC-T algorithmic structure, low-rank projection is applied to compress high-dimensional input states, and low-rank matrix factorization is introduced into the actor's output mapping. This factorization reduces the agent's parameter footprint by 20% to over 52%, ensuring smooth scalability from 100 to 2000 clients while preserving policy expressiveness.
Loss & Training¶
The DSAC-T agent is trained under the maximum-entropy reinforcement learning paradigm to encourage active exploration:
The critic networks are optimized using the variance-scaled distributional Kullback-Leibler (KL) divergence loss:
with the variance-aware clipping boundary \(C(y_z^{\min}; b) = \text{clip}(y_z^{\min}, Q_{\theta_i} - b, Q_{\theta_i} + b)\) guided by the moving uncertainty threshold \(b = \xi \mathbb{E}_{\mathcal{B}}[\sigma_{\theta_i}]\). The actor parameters are updated by maximizing the pessimistic soft value surrogate:
The reward function measures the delta of global accuracy on the validation set using an asymmetric exponential formulation (\(\Xi = 64\)):
Key Experimental Results¶
Main Results¶
All methods are evaluated on standard 100-client federated benchmarks against seven baseline approaches: FedAvg (averaging), MOON (contrastive), Oort (heuristic selection), FedLC (logit calibration), FedNTD (not-true distillation), FedCross (cross-model collaboration), and FedAA (discrete RL-based aggregation). Data heterogeneity is governed by Dirichlet distributions \(\text{Dir}(\alpha)\), where \(\alpha=0.1\) denotes extreme non-IID skew, \(\alpha=0.5\) denotes moderate skew, and IID represents uniform distribution.
Table 1: Top-1 test accuracy of baseline methods and FedDO across non-IID and IID partitions (%)
| Model | Dataset | Heterogeneity | FedAvg | MOON | Oort | FedLC | FedNTD | FedCross | FedAA | FedDO (Ours) |
|---|---|---|---|---|---|---|---|---|---|---|
| CNN | FashionMNIST | \(\alpha=0.1\) | 83.66 ± 1.49 | 74.00 ± 10.10 | 84.29 ± 0.60 | 82.96 ± 1.83 | 83.78 ± 1.32 | 84.75 ± 0.31 | 75.53 ± 1.86 | 87.49 ± 0.07 |
| FashionMNIST | \(\alpha=0.5\) | 83.52 ± 2.76 | 80.48 ± 1.95 | 82.99 ± 1.67 | 83.48 ± 2.31 | 84.53 ± 1.22 | 84.57 ± 0.69 | 76.58 ± 1.48 | 97.34 ± 0.01 | |
| FashionMNIST | IID | 90.80 ± 0.20 | 90.77 ± 0.14 | 91.06 ± 0.09 | 90.16 ± 0.21 | 90.74 ± 0.23 | 89.35 ± 0.03 | 89.88 ± 0.03 | 90.99 ± 0.02 | |
| CIFAR-10 | \(\alpha=0.1\) | 66.56 ± 0.72 | 46.19 ± 3.81 | 70.78 ± 0.27 | 67.83 ± 0.08 | 59.51 ± 0.43 | 55.68 ± 1.27 | 50.76 ± 0.10 | 71.36 ± 0.10 | |
| CIFAR-10 | \(\alpha=0.5\) | 56.88 ± 4.60 | 46.89 ± 4.04 | 34.08 ± 3.26 | 56.23 ± 1.93 | 59.52 ± 0.69 | 63.28 ± 0.67 | 43.39 ± 2.11 | 89.70 ± 0.05 | |
| CIFAR-10 | IID | 64.31 ± 0.06 | 62.76 ± 0.24 | 63.98 ± 0.09 | 64.51 ± 0.15 | 66.13 ± 0.07 | 59.25 ± 0.07 | 57.10 ± 0.05 | 67.03 ± 0.08 | |
| CIFAR-100 | \(\alpha=0.1\) | 28.92 ± 0.06 | 26.41 ± 0.35 | 25.71 ± 0.32 | 27.98 ± 0.18 | 28.93 ± 0.12 | 49.81 ± 0.05 | 20.22 ± 0.21 | 50.14 ± 0.15 | |
| CIFAR-100 | \(\alpha=0.5\) | 26.65 ± 0.18 | 24.82 ± 0.47 | 25.17 ± 0.53 | 27.02 ± 0.32 | 27.86 ± 0.19 | 50.92 ± 0.06 | 19.12 ± 0.23 | 53.44 ± 0.21 | |
| CIFAR-100 | IID | 23.44 ± 0.16 | 22.50 ± 0.15 | 22.88 ± 0.16 | 24.29 ± 0.03 | 25.91 ± 0.06 | 31.44 ± 0.03 | 19.61 ± 0.06 | 32.39 ± 0.14 | |
| TinyImageNet | \(\alpha=0.1\) | 12.01 ± 0.18 | 10.69 ± 0.43 | 12.49 ± 0.05 | 11.96 ± 0.15 | 14.44 ± 0.06 | 12.98 ± 0.02 | 9.47 ± 0.10 | 32.45 ± 0.04 | |
| TinyImageNet | \(\alpha=0.5\) | 11.81 ± 0.13 | 10.11 ± 0.39 | 11.94 ± 0.18 | 12.98 ± 0.03 | 14.57 ± 0.05 | 13.54 ± 0.05 | 9.42 ± 0.13 | 33.17 ± 0.16 | |
| TinyImageNet | IID | 10.81 ± 0.08 | 8.85 ± 0.09 | 10.22 ± 0.13 | 10.52 ± 0.03 | 6.50 ± 0.08 | 8.21 ± 0.05 | 6.54 ± 0.05 | 12.59 ± 0.04 | |
| ResNet-18 | FashionMNIST | \(\alpha=0.1\) | 95.30 ± 0.32 | 80.87 ± 3.79 | 65.34 ± 2.01 | 95.79 ± 0.22 | 83.51 ± 0.77 | 94.90 ± 0.02 | 67.87 ± 0.61 | 97.68 ± 0.03 |
| FashionMNIST | \(\alpha=0.5\) | 95.83 ± 0.07 | 82.35 ± 4.39 | 63.85 ± 0.42 | 96.26 ± 0.16 | 87.28 ± 1.21 | 95.34 ± 0.04 | 72.21 ± 0.58 | 97.69 ± 0.01 | |
| FashionMNIST | IID | 90.69 ± 0.09 | 90.61 ± 0.06 | 92.11 ± 0.03 | 90.77 ± 0.05 | 91.10 ± 0.10 | 88.23 ± 0.04 | 91.00 ± 0.06 | 92.11 ± 0.03 | |
| CIFAR-10 | \(\alpha=0.1\) | 55.75 ± 4.16 | 54.18 ± 1.27 | 59.25 ± 0.25 | 66.42 ± 0.04 | 53.54 ± 1.19 | 82.50 ± 0.17 | 42.89 ± 1.92 | 83.24 ± 0.10 | |
| CIFAR-10 | \(\alpha=0.5\) | 87.08 ± 0.36 | 65.92 ± 4.89 | 22.75 ± 4.48 | 69.99 ± 3.44 | 72.79 ± 2.86 | 85.27 ± 0.46 | 54.09 ± 0.61 | 88.26 ± 0.10 | |
| CIFAR-10 | IID | 66.07 ± 0.13 | 67.03 ± 0.08 | 66.24 ± 0.21 | 68.19 ± 0.06 | 68.38 ± 0.13 | 75.06 ± 0.03 | 57.70 ± 0.33 | 77.03 ± 0.08 | |
| CIFAR-100 | \(\alpha=0.1\) | 23.89 ± 0.29 | 24.22 ± 0.20 | 23.48 ± 0.26 | 27.62 ± 0.25 | 27.30 ± 0.14 | 26.25 ± 0.09 | 19.20 ± 0.16 | 31.65 ± 0.04 | |
| CIFAR-100 | \(\alpha=0.5\) | 19.74 ± 0.59 | 22.09 ± 0.45 | 24.09 ± 0.34 | 19.83 ± 0.88 | 20.65 ± 0.50 | 17.15 ± 0.11 | 19.16 ± 0.15 | 26.61 ± 0.20 | |
| CIFAR-100 | IID | 32.37 ± 0.32 | 31.32 ± 0.09 | 33.21 ± 0.06 | 33.08 ± 0.05 | 36.03 ± 0.10 | 32.65 ± 0.06 | 27.65 ± 0.14 | 37.87 ± 0.01 | |
| TinyImageNet | \(\alpha=0.1\) | 11.11 ± 0.15 | 11.90 ± 0.11 | 11.34 ± 0.10 | 10.84 ± 0.27 | 18.49 ± 0.13 | 13.89 ± 0.06 | 8.30 ± 0.07 | 22.02 ± 0.23 | |
| TinyImageNet | \(\alpha=0.5\) | 10.42 ± 0.05 | 12.36 ± 0.09 | 11.01 ± 0.16 | 20.47 ± 0.02 | 15.17 ± 0.49 | 30.48 ± 0.10 | 8.12 ± 0.07 | 31.22 ± 0.11 | |
| TinyImageNet | IID | 17.15 ± 0.07 | 15.54 ± 0.10 | 15.82 ± 0.12 | 18.29 ± 0.05 | 19.73 ± 0.17 | 16.65 ± 0.10 | 8.95 ± 0.12 | 20.01 ± 0.06 | |
| VGG-11 | FashionMNIST | \(\alpha=0.1\) | 77.96 ± 1.38 | 40.52 ± 4.52 | 38.05 ± 1.57 | 74.84 ± 4.29 | 76.11 ± 1.60 | 74.06 ± 0.09 | 60.17 ± 2.16 | 90.65 ± 0.02 |
| FashionMNIST | \(\alpha=0.5\) | 72.29 ± 2.11 | 55.08 ± 1.94 | 40.20 ± 2.24 | 81.81 ± 2.47 | 72.53 ± 3.94 | 74.13 ± 0.36 | 51.92 ± 1.56 | 94.63 ± 0.01 | |
| FashionMNIST | IID | 91.11 ± 0.06 | 90.79 ± 0.10 | 92.06 ± 0.17 | 91.00 ± 0.05 | 90.60 ± 0.04 | 89.67 ± 0.03 | 89.40 ± 0.21 | 92.49 ± 0.05 | |
| CIFAR-10 | \(\alpha=0.1\) | 75.11 ± 1.77 | 69.58 ± 2.66 | 64.44 ± 0.59 | 88.55 ± 0.03 | 76.72 ± 0.54 | 71.90 ± 0.36 | 46.05 ± 1.95 | 81.76 ± 0.08 | |
| CIFAR-10 | \(\alpha=0.5\) | 78.67 ± 1.30 | 35.74 ± 3.24 | 14.66 ± 2.96 | 76.28 ± 2.15 | 65.05 ± 1.54 | 71.75 ± 0.11 | 29.38 ± 1.53 | 90.35 ± 0.14 | |
| CIFAR-10 | IID | 78.54 ± 0.20 | 79.69 ± 0.19 | 80.27 ± 0.29 | 78.73 ± 0.30 | 81.80 ± 0.10 | 73.06 ± 0.04 | 69.35 ± 0.25 | 82.37 ± 0.07 | |
| CIFAR-100 | \(\alpha=0.1\) | 31.37 ± 0.20 | 22.07 ± 0.52 | 12.98 ± 0.61 | 27.27 ± 0.35 | 22.55 ± 0.30 | 47.15 ± 0.09 | 11.29 ± 0.59 | 61.92 ± 0.15 | |
| CIFAR-100 | \(\alpha=0.5\) | 26.86 ± 0.07 | 24.66 ± 0.33 | 12.89 ± 0.42 | 26.78 ± 0.15 | 27.87 ± 0.19 | 46.12 ± 0.10 | 9.19 ± 1.48 | 61.34 ± 0.09 | |
| CIFAR-100 | IID | 44.18 ± 0.12 | 35.18 ± 0.24 | 37.30 ± 0.38 | 37.71 ± 0.66 | 40.91 ± 0.55 | 38.85 ± 0.15 | 27.98 ± 0.27 | 44.84 ± 0.06 | |
| TinyImageNet | \(\alpha=0.1\) | 6.92 ± 0.54 | 10.41 ± 0.12 | 3.75 ± 0.23 | 6.62 ± 0.34 | 5.97 ± 0.41 | 11.61 ± 0.17 | 5.91 ± 0.05 | 14.73 ± 0.09 | |
| TinyImageNet | \(\alpha=0.5\) | 6.55 ± 0.36 | 8.35 ± 0.22 | 4.14 ± 0.22 | 7.02 ± 0.21 | 6.62 ± 0.24 | 13.91 ± 0.06 | 4.13 ± 0.22 | 15.10 ± 0.29 | |
| TinyImageNet | IID | 21.85 ± 0.19 | 20.28 ± 0.64 | 13.27 ± 0.34 | 18.93 ± 0.55 | 19.24 ± 0.46 | 18.71 ± 0.12 | 7.62 ± 0.50 | 24.14 ± 0.07 |
Ablation Study & Scalability¶
Table 2: Communication rounds \(R\) and accuracy (Acc, %) required to reach target accuracy on CIFAR-10 with ResNet-18
| Method | \(\text{Dir}(0.1)\) Rounds & Acc (Target 71.4%) | \(\text{Dir}(0.5)\) Rounds & Acc (Target 87.6%) | \(\text{IID}\) Rounds & Acc (Target 66.6%) | Note |
|---|---|---|---|---|
| FedAvg | (>1000, 55.7) | (784, 87.7) | (>1000, 66.2) | Severe drift under non-IID; requires 784 rounds under Dir(0.5) |
| MOON | (>1000, 14.5) | (>1000, 37.3) | (773, 67.1) | Model-contrastive regularization converges slowly under skew |
| Oort | (>1000, 59.35) | (>1000, 24.1) | (>1000, 66.4) | Discrete exploration-exploitation fails to arrest gradient drift |
| FedLC | (>1000, 66.42) | (>1000, 71.7) | (639, 68.1) | Calibrating logits improves stability but misses high target |
| FedNTD | (>1000, 54.6) | (>1000, 73.9) | (625, 68.0) | High distillation overhead, fails target under high skew |
| FedCross | (729, 82.4) | (>1000, 85.7) | (821, 75.0) | Cross-collaboration stabilizes updates but demands more rounds |
| FedAA | (>1000, 42.1) | (>1000, 54.7) | (>1000, 57.1) | Discrete RL action space struggles in high dimensions |
| FedDO (Ours) | (344, 71.5) | (429, 87.8) | (476, 67.4) | Halves communication rounds; hits target across all splits |
Note: > denotes failure to reach target accuracy within 1000 rounds; the accompanying number indicates the best accuracy reached.
Table 3: Impact of parameterized restructuring on agent parameter size and accuracy on FashionMNIST (VGG-11, IID)
| Client Cohort \(K\) | Original Acc (%) | Original Parameters (Para) | Low-Rank Acc (%) | Low-Rank Parameters (Para) | Parameter Reduction |
|---|---|---|---|---|---|
| 100 | 90.62 | 0.93M | 90.29 | 0.74M | ↓ 20.4% |
| 500 | 91.03 | 2.26M | 89.27 | 1.28M | ↓ 43.4% |
| 1000 | 89.78 | 3.96M | 89.55 | 1.95M | ↓ 50.8% |
| 2000 | 88.51 | 7.25M | 85.96 | 3.47M | ↓ 52.1% |
Key Findings¶
- Substantial convergence acceleration: In the challenging non-IID \(\text{Dir}(0.1)\) setting, nearly all baselines fail to achieve the target accuracy within 1000 communication rounds. FedDO converges in only 344 rounds, achieving a \(2.1\times\) speedup over the only converging baseline (FedCross) while securing superior accuracy.
- Adaptive participation-intensity dynamics: The learned policy displays distinct behavioral patterns: under heavy distribution skew (\(\text{Dir}(0.1)\)), the agent stabilizes on a lean cohort of 20–25 active clients with high data usage (0.8–0.9) to preserve gradient consistency; under IID distributions, it engages over 50 clients with modest per-client data allocations (0.4–0.5) to maximize computation and communication efficiency.
- High-dimensional scalability: As detailed in Table 3, scaling to 2000 clients causes the original policy network to expand to 7.25M parameters, whereas low-rank parameterization compresses it to 3.47M (\(\downarrow 52.1\%\)) with only a 2.55% accuracy trade-off, breaking the dimensionality barrier for reinforcement learning in massive federated systems.
Highlights & Insights¶
- From binary selection to continuous control: While conventional methods rely on on/off client switches, FedDO identifies data-usage ratios as a continuous control lever that smooths gradient variance and eliminates the oscillations inherent in discrete client sampling.
- Natural alignment of distributional RL with FL: Federated systems inherently feature stochastic network latency and non-stationary distribution shifts. DSAC-T's twin distributional critics explicitly model value variance, and variance-aware clipping suppresses optimistic drift, making off-policy learning stable in volatile environments.
- Decoupled complexity via low-rank factorizations: Low-rank matrix factorization decouples the agent's parameter scaling from the number of clients, offering a viable blueprint for applying continuous DRL agents to large-scale distributed edge fleets.
Limitations & Future Work¶
- Offline policy pretraining overhead: The DSAC-T policy requires 200 complete training cycles during the offline preparation phase. Although this overhead is amortized across subsequent deployments, zero-shot generalization across unseen data distributions remains an open research avenue.
- Synchronous communication assumption: The framework currently operates in synchronous communication rounds without fully addressing device stragglers and unexpected disconnections. Future work should extend continuous data-usage optimization to asynchronous federated learning and incorporate differential privacy mechanisms.
Related Work & Insights¶
- vs FedAvg & FedProx: FedAvg applies unweighted or sample-proportional averaging, suffering from severe gradient drift under non-IID data; FedProx regularizes local update distances but cannot reallocate data usage. FedDO dynamically controls data ingestion at the source, preventing divergent gradients from contaminating aggregation.
- vs Oort: Oort selects clients via discrete exploration-exploitation based on loss variance and device capabilities. FedDO refines this into continuous data-usage assignment, consistently outperforming static heuristic rules under complex heterogeneity.
- vs FedAA: FedAA leverages RL for client selection and aggregation weights, but uses discrete action spaces and basic value estimation. FedDO employs continuous distributional control via DSAC-T and low-rank scaling, achieving higher accuracy, faster convergence, and superior scalability.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Pioneers continuous client data-usage control paired with distributional RL and low-rank scaling]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive verification across 4 vision datasets, 3 architectures, 3 heterogeneity levels, and 2000-client cohorts]
- Writing Quality: ⭐⭐⭐⭐⭐ [Rigorous mathematical formulation, clear narrative structure, and compelling empirical analysis]
- Value: ⭐⭐⭐⭐☆ [Offers a practical, robust, and scalable optimization paradigm for heterogeneous federated systems]