Skip to content

H-SFP: Hierarchical Federated Learning with Decoupled Split-Model Prototyping

Conference: ECCV 2026
Paper: ECCV Official
Full Cache: /Users/zy/workspace/paper_cache/ECCV2026/eccv-5748.txt
Area: Optimization & Theory
Keywords: Federated Learning, Split Federated Learning, Feature Prototyping, Communication Efficiency, Non-IID Data

TL;DR

Addressing the bandwidth bottlenecks and synchronous locks of transmitting high-dimensional activations and backward gradients in split federated learning, H-SFP decouples cross-tier backpropagation by exchanging compact class-wise first- and second-order feature statistics, synthesizing Gaussian feature distributions at upper tiers to cut communication overhead by over two orders of magnitude while substantially curbing client drift.

Background & Motivation

Federated learning (FL) enables distributed edge clients to collaboratively train deep learning models without centralizing raw training data, serving as a cornerstone paradigm for privacy-sensitive visual learning. However, practical deployments over heterogeneous networks must navigate an inherent trilemma: severe statistical heterogeneity (Non-IID distributions) across edge devices, stringent compute and memory budgets on resource-constrained clients, and crippling communication bottlenecks over bandwidth-limited wireless links. Conventional FL algorithms such as FedAvg and FedProx inherently presume that edge clients possess the computational capability to train and evaluate the full model locally. When deployed on lightweight edge devices with limited RAM, loading the entire network triggers out-of-memory errors and excludes weak nodes from participation.

To alleviate client-side hardware constraints, split federated learning (SFL) and hierarchical split federated learning (HSFL) partition the global network across client, edge, and cloud tiers. The client executes only a shallow sub-network, offloading intermediate processing and heavy prediction layers to upper tiers. Nonetheless, existing split-learning paradigms preserve a tightly coupled end-to-end optimization graph: clients must transmit sample-level intermediate activations (smashed data) during the forward pass, and servers must transmit per-sample gradients back during backpropagation. This bidirectional sample-dependent dependency incurs an \(O(B \cdot d)\) communication burden per batch, enforces strict inter-tier synchronization locks, and exposes intermediate activations to feature-inversion privacy attacks.

While prototype-based federated frameworks like FedProto exchange class prototypes to compress communication, they primarily employ prototypes as client-side regularization penalties rather than using them to synthesize training data for upper tiers, frequently suffering from severe prototype drift under extreme heterogeneity. This paper revisits the communication interface of hierarchical split learning: can compact class-level feature statistics preserve sufficient geometric structure to train upper-tier model components independently without sample-level interactions? Core idea: sever the cross-tier end-to-end backpropagation graph by having clients transmit only compact class-wise first- and second-order moments, enabling edge and cloud tiers to synthesize representative Gaussian feature distributions for local decoupled training, complemented by dual-timescale model averaging to ensure measure-theoretic asymptotic convergence.

Method

Overall Architecture

H-SFP is tailored for a three-tier cloud-edge-client topology consisting of a central cloud server, \(M\) intermediate edge servers, and \(K\) distributed clients. The global network \(W\) is split into three decoupled components: a lightweight client-side feature extractor \(W_c\) (e.g., the initial residual block of a ResNet), an intermediate edge-side representation model \(W_e\), and a cloud-side task prediction head \(W_g\). Clients train \(W_c\) locally on private data using a self-supervised contrastive objective and package each class's extracted representations into compact mean and variance moments. Edge servers aggregate these moments within their clusters, synthesize intermediate feature samples via a multivariate Gaussian distribution, and train \(W_e\) with a contrastive loss. The cloud tier subsequently aggregates the edge-level distributions, sampling global synthetic features to train \(W_g\) via supervised classification. The entire pipeline operates on a dual-timescale schedule, where statistics are communicated frequently while shallow client models undergo low-frequency model averaging.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Private Image Inputs"] --> B["Multi-Tier Decoupled Self-Learning<br/>Local NT-Xent Contrastive Optimization of Wc"]
    B --> C["Compact Moment Prototype Packaging & Generative Synthesis<br/>Compute and upload class statistics (μ, σ) to Edge"]
    C --> D["Edge Server Feature Aggregation & Synthetic Sampling<br/>Gaussian Feature Resampling & Contrastive Training on We"]
    D --> E["Cloud Global Distribution Aggregation & Task Training<br/>Synthesize Global Features & Train Task Model Wg"]
    E -->|Low-Frequency Interval Ic| F["Dual-Timescale Transmission & Model Averaging<br/>Aggregate and Update Shallow Client Parameters Wc"]

Key Designs

1. Multi-Tier Decoupled Self-Learning: Severing Cross-Tier Backpropagation via Independent Representation Optimization
Conventional hierarchical split learning requires a continuous end-to-end gradient backpropagation chain linking cloud, edge, and client tiers, causing straggler latency and cascading communication bottlenecks. H-SFP eliminates cross-tier backpropagation entirely, allowing each tier to conduct closed-loop optimization within its physical boundary. Client \(k\) computes shallow representations \(z_i = W_{c,k}(x_i)\) on its private dataset \(\mathcal{D}_k\) using the supervised NT-Xent contrastive loss, pulling representations of the same class together while repelling distinct classes to enforce tight, well-separated feature clusters: $\(\mathcal{L}_{\text{con},k}(W_{c,k}) = \sum_{i \in \mathcal{B}} \frac{-1}{|P(i)|} \sum_{p \in P(i)} \log \frac{\exp(z_i \cdot z_p / \tau)}{\sum_{a \in A(i)} \exp(z_i \cdot z_a / \tau)}\)$ where \(P(i)\) denotes positive samples in batch \(\mathcal{B}\) sharing class \(i\)'s label, \(A(i)\) denotes all other candidates, and \(\tau\) is the temperature hyperparameter. Crucially, when edge server \(m\) trains intermediate model \(W_{e,m}\) on synthesized features, it applies the exact same contrastive loss \(\mathcal{L}_{m}(W_{e,m}) = \mathcal{L}_{\text{con},m}(W_e(\mathcal{D}_{\text{syn}}^{\text{edge}}))\). This objective forces \(W_e\) to learn relative topological relationships and structural manifold geometries rather than overfitting to synthetic noise. Finally, the cloud optimizes standard supervised cross-entropy loss \(\mathcal{L}_{\text{task}}(W_g)\) over globally synthesized representations.

2. Compact Moment Prototype Packaging & Generative Synthesis: Lightweight Gaussian Manifold Reconstruction with Elliptical Boundaries
To eradicate the massive bandwidth consumption of transmitting per-sample activations (smashed data), H-SFP computes class-conditional statistical moments after local training. For each class \(j\) present on client \(k\), the client calculates the empirical feature mean \(\mu_k^{(j)}\) and coordinate-wise standard deviation \(\sigma_k^{(j)}\): $\(\mu_k^{(j)} = \frac{1}{|\mathcal{D}_k^{(j)}|} \sum_{x_i \in \mathcal{D}_k^{(j)}} W_{c,k}(x_i), \quad \sigma_k^{(j)} = \sqrt{\frac{1}{|\mathcal{D}_k^{(j)}|} \sum_{x_i \in \mathcal{D}_k^{(j)}} \left(W_{c,k}(x_i) - \mu_k^{(j)}\right)^2}\)$ Communication complexity drops from sample-dependent \(O(B \cdot d)\) to class-dependent \(O(J_k \cdot d_c)\). The authors explicitly employ a diagonal covariance matrix: transmitting a full covariance matrix would escalate bandwidth to \(O(d_c^2)\), recreating the communication bottleneck; conversely, collapsing to a mean-only prototype fails to capture the oriented elliptical boundaries of feature clusters. Edge and cloud servers synthesize training features via multivariate normal sampling: $\(z_{\text{syn}} \sim \mathcal{N}\left(\mu, \text{diag}(\sigma^2) + \epsilon I\right)\)$ Under feature-space Lipschitz continuity, the paper proves that combining first-order means with second-order diagonal variances strictly bounds the Maximum Mean Discrepancy (MMD) between true and synthesized distributions, reconstructing feature geometry with high fidelity while naturally thwarting activation-inversion privacy attacks.

3. Dual-Timescale Transmission & Model Averaging: Measure-Theoretic Stabilization Against Long-Term Client Drift
In an asynchronous generative setup without cross-tier gradient feedback, client-side shallow extractors would eventually suffer from representation collapse or severe client drift if left unconstrained. H-SFP establishes a dual-timescale operational cadence: on the fast timescale (every round), clients upload only lightweight statistical moments \((\mu_k, \sigma_k)\) to their respective edge servers, requiring minimal bandwidth; on the slow timescale (every \(I_c\) communication rounds), edge servers and the cloud execute infrequent parameter averaging over client shallow models \(W_{c,k}\). From a measure-theoretic standpoint, upper-tier models optimize against a dynamic probability measure \(Q_e^t\) parameterized by the lower tier. Under Wasserstein-2 Lipschitz continuity and bounded gradient variance assumptions, the stabilization of client models guarantees that the cross-tier distribution drift \(\delta_m^t = L_m \cdot \mathcal{W}_2(Q_e^{t+1}, Q_e^t) \to 0\), theoretically proving that decoupled global parameters asymptotically converge to a bounded stationary neighborhood.

Key Experimental Results

Main Results

H-SFP was evaluated across 200 distributed clients and 10 edge servers, benchmarked against standard FL (FedAvg, FedProx), prototype and distillation methods (FedProto, FedGen, FedDF), and split-learning frameworks (SplitFed, HeteroSFL, HSFL). Table 1 presents classification accuracy on CIFAR-10 and CIFAR-100 under IID and Dirichlet Non-IID (\(\alpha=0.3, 0.7\)) partitions. Table 2 reports results on large-scale ImageNet-1K and dermatoscopic HAM10000.

Dataset Distribution Setting Ours: H-SFP (A) Strong Baseline: HSFL (A) Prototype Baseline: FedProto Standard: FedAvg Key Takeaway / Gain
CIFAR-10 IID 67.44 ± 0.85% 51.35 ± 1.10% 55.10 ± 1.05% 57.25 ± 1.24% Outperforms HSFL by +16.09%
CIFAR-10 Dirichlet(0.3) 40.15 ± 1.25% 17.80 ± 1.75% 38.50 ± 1.45% 12.15 ± 2.10% Massive gain under extreme Non-IID
CIFAR-10 Dirichlet(0.7) 48.85 ± 1.10% 26.54 ± 1.50% 45.20 ± 1.30% 23.60 ± 1.85% Outperforms FedProto by +3.65%
CIFAR-100 IID 55.10 ± 0.95% 35.47 ± 1.35% 46.50 ± 1.15% 48.85 ± 1.40% Tops all federated baselines
CIFAR-100 Dirichlet(0.3) 7.15 ± 0.65% 5.10 ± 0.80% 6.50 ± 0.95% 2.10 ± 0.85% Surpasses FedGen (6.70%)
CIFAR-100 Dirichlet(0.7) 10.48 ± 0.85% 8.25 ± 1.15% 8.80 ± 1.05% 5.40 ± 1.15% Robust in fine-grained multi-class
HAM10000 ResNet-50 (IID) 80.86 ± 0.75% 68.82 ± 1.25% 67.20 ± 1.25% 68.28 ± 1.45% Nears Centralized Oracle (86.40%)
HAM10000 ResNet-50 (dir(0.7)) 37.50 ± 1.05% 30.45 ± 1.40% 35.00 ± 1.40% 13.45 ± 1.65% +7.05% absolute gain over HSFL
ImageNet-1K ResNet-50 (IID) 27.9 ± 0.3% 24.6 ± 0.9% 25.2 ± 1.2% 23.1 ± 1.4% Leads under strict 200-round budget
ImageNet-1K ResNet-50 (dir(0.7)) 22.4 ± 0.6% 14.8 ± 1.1% 18.5 ± 1.4% 11.2 ± 1.7% Leads HeteroSFL (16.2%) by +6.2%

On the dense medical image segmentation benchmark ISIC-2018 (ResNet50-UNet), H-SFP reaches 68.3 ± 0.5% IoU and 79.5 ± 0.4% DICE with a wall-clock training time of only 6.8 ± 0.2 hours. In contrast, SplitFed requires 72.4 hours to reach 53.8% IoU, and HSFL takes 14.5 hours for 62.1% IoU, demonstrating a 10× training speedup alongside a +14.5% IoU gain.

Ablation Study

To validate each algorithmic component and assess the role of second-order variance in feature distribution synthesis, the authors conducted comprehensive ablations on CIFAR-100 and ImageNet-1K (200 clients, \(I_c=5, I_e=10\), Table 5), alongside temperature hyperparameter sensitivity on \(\tau\) (Table 6).

Configuration / Variant CIFAR-100 (IID) CIFAR-100 (dir(0.7)) ImageNet-1K (IID) ImageNet-1K (dir(0.7)) Mechanism Analysis & Findings
HSFL Baseline (No Prototyping) 35.47 ± 1.35% 8.25 ± 1.15% 24.6 ± 0.9% 14.8 ± 1.1% Sample-level gradient propagation collapses under Non-IID
+ Contrastive Loss (CL) 41.20 ± 1.25% 8.95 ± 1.05% 25.4 ± 0.8% 16.3 ± 1.0% Normalizes client feature geometry for steady gains
+ Synthetic Data (SD) 47.60 ± 1.15% 9.45 ± 0.95% 26.1 ± 0.7% 18.2 ± 0.9% Generative feature resampling enhances generalization
Full (μ-only, No Variance) 52.35 ± 1.05% 9.85 ± 0.90% 26.8 ± 0.5% 20.5 ± 0.8% Discards elliptical bounds, losing 2.75% accuracy
Purely Generative (No Model Avg.) 48.20 ± 1.10% 9.15 ± 1.00% 25.8 ± 0.7% 19.4 ± 0.9% Shallow feature drift impairs upper-tier training
Full H-SFP (μ + σ + Model Avg.) 55.10 ± 0.95% 10.48 ± 0.85% 27.9 ± 0.3% 22.4 ± 0.6% Second-order moments + dual-timescale achieve peak SOTA

In the ablation on contrastive temperature \(\tau\), values that are too small (\(\tau=0.2\)) cause sharp, unstable gradients (48.15% on CIFAR-100), whereas values that are too large (\(\tau=2.0\)) over-smooth representations (48.65%); \(\tau = 0.5\) consistently yields the optimal balance across all benchmarks (55.10%).

Key Findings

  • Second-Order Variance Is Critical for Manifold Preservation: The Full (μ-only) ablation suffers a 2.75% accuracy drop on CIFAR-100 compared to full H-SFP. Centroid-only prototypes fail to express intra-class variance and elliptical boundaries, whereas diagonal variance maintains geometric MMD bounds at negligible extra communication cost.
  • Order-of-Magnitude Reductions in Communication and RAM: Across 200 rounds of CIFAR-100 training, FedAvg consumes 3,865 GB, and HSFL consumes 116 GB (including 43.47 GB in gradients and 54.93 GB in activations). H-SFP (A) requires only 13.8 GB, and H-SFP (C) consumes a mere 10.94 GB—yielding a 353× reduction over FedAvg and a 10.4× reduction over HSFL. Concurrently, client RAM usage plummets from ~600 MB (full model) to sub-100 MB.
  • Scalability Under Extreme Data Sparsity: As client cohort size increases from 20 to 200 (reducing local sample count), FedAvg drops sharply from 60.50% to 48.85% (-11.65%), whereas H-SFP demonstrates robust resilience, maintaining 55.10% on CIFAR-100 and consistently leading all baselines.

Highlights & Insights

  • Decoupled Architecture Without Backpropagation: By eliminating inter-tier backpropagation, H-SFP transforms upper tiers into consumers of synthesized distributions rather than consumers of raw activations, removing distributed synchronization stragglers.
  • Cost-Effective Diagonal Gaussian Synthesis: Rather than training heavy generative adversarial networks or diffusion models on the server, H-SFP leverages contrastive pre-structuring to make diagonal Gaussian sampling sufficient for high-fidelity representation modeling at \(O(C \cdot d)\) complexity.
  • Measure-Theoretic Convergence Foundation: Coupling fast-timescale prototype exchange with slow-timescale shallow model averaging allows the paper to bound cross-tier Wasserstein-2 drift, rigorously proving convergence for decoupled split training.

Limitations & Future Work

  • Independent Coordinate Assumption: Restricting feature synthesis to a diagonal covariance matrix assumes feature dimensions are uncorrelated, which may struggle with complex multimodal or highly curved manifold distributions.
  • Unoptimized Random Tier Topology: The current benchmark randomly assigns clients to edge servers without accounting for geographic proximity, physical network ping, or client data distribution affinities; topology-aware hierarchical clustering remains a promising avenue.
  • vs FedAvg / FedProx: Standard FL requires clients to maintain and optimize the entire model locally, incurring high memory footprints (~600 MB); H-SFP offloads deep layers (<100 MB RAM) and avoids full-parameter exchanges.
  • vs HSFL / SplitFed: Conventional split learning transmits sample-level activations and backward gradients, incurring \(O(B \cdot d)\) communication and synchronization stalls; H-SFP transmits only class moments at \(O(C \cdot d)\) complexity, eliminating backward gradients and reducing bandwidth by 10× to 350×.
  • vs FedProto / FedGen: FedProto confines prototypes to client-side regularization penalties without enabling split execution; FedGen requires training compute-heavy server-side generative models. H-SFP utilizes parameter-free Gaussian sampling conditioned on empirical moments to train split upper layers directly.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Decoupling split federated learning via statistical moment prototyping and eliminating cross-tier backpropagation is highly original.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation across 4 classification datasets, medical image segmentation, varying Non-IID Dirichlet shifts, client scaling (20-200), and detailed ablation studies.
  • Writing Quality: ⭐⭐⭐⭐⭐ Mathematically rigorous, well-structured, with clear communication complexity comparisons and intuitive visualizations.
  • Value: ⭐⭐⭐⭐⭐ Highly practical for deploying deep neural networks across bandwidth-constrained and compute-limited edge and IoT ecosystems.