HER-Count: Learning Hyper-Exemplar Representation for Generalized Zero-Shot Object Counting¶
Conference: ECCV 2026
Paper: ECCV 2026
Area: Object Detection
Keywords: zero-shot object counting, hyper-exemplar representation, multimodal large language model, hierarchical feature fusion, global discrimination enhancement
TL;DR¶
Addressing the issues where patch-based exemplars fail to cover intra-class diversity and accumulate detector errors while text embeddings lack image-specific instance cues, HER-Count leverages a multimodal large language model (MLLM) to synthesize holistic hyper-exemplar representations (HER), enabling high-accuracy, noise-robust zero-shot counting and localization via hierarchical fusion and global discrimination enhancement.
Background & Motivation¶
Object counting aims to determine the total number of objects of a specified category within an image, playing an indispensable role in practical scenarios like surveillance, crowd monitoring, traffic flow tracking, and biological cell analysis. Early counting techniques were predominantly class-specific, training dedicated regressors on dense dot annotations for closed categories. However, these models cannot generalize to unseen classes encountered in open environments. To eliminate category restrictions, class-agnostic counting has emerged as a major focus, encompassing few-shot and zero-shot setups. Few-shot methods require human annotators to crop several representative exemplar bounding boxes at test time to match similar instances, but this manual intervention hinders autonomous end-to-end deployment. Consequently, zero-shot counting—relying purely on textual category names—has gained significant traction. Existing methods typically follow two paradigms: either deploying off-the-shelf object detectors (such as Grounding DINO) to identify and select 1–3 visual patches as exemplars, or directly mapping category names to generic text embeddings for cross-modal matching with image regions.
Nonetheless, both paradigms encounter fundamental bottlenecks in diverse open-world scenarios. For patch selection methods, objects of the same class within an image frequently undergo substantial appearance variations due to arbitrary scale differences, viewing angle shifts, partial occlusions, and lighting changes. A handful of sparse patches cannot capture such comprehensive intra-class diversity. More detrimentally, any missing, truncated, or false-positive bounding boxes produced by automated detectors propagate irreversibly through the multi-stage pipeline, compounding matching errors. On the other hand, methods reliant solely on text embeddings remain trapped in generic category semantics; they are entirely blind to the concrete visual appearance and instance-level details present in the specific target image. Furthermore, prevailing approaches derive counts by summing across predicted continuous density maps, where unavoidable low-level diffuse background noise accumulates into substantial count overestimations while completely discarding spatial localization information.
The crux of the tension lies between the fragility of localized physical patch crops and the excessive semantic abstraction of text-only embeddings. This paper argues that exemplars need not be restricted to discrete pixel crops cropped from the image, nor should they be isolated text vectors detached from visual context. Instead, a multimodal large language model (MLLM) can jointly ingest the holistic image along with category instructions to synthesize a continuous latent representation that encapsulates the visual commonalities of all target instances in that image. Core idea: exploit the cross-modal synthesis capacity of an MLLM to construct a multimodal Hyper-Exemplar Representation (HER) from the full image and textual category, progressively injecting it across visual backbone layers via hierarchical fusion and regularizing it with sample- and class-level global discrimination constraints for accurate, noise-resistant zero-shot counting and localization.
Method¶
Overall Architecture¶
HER-Count adopts a compact end-to-end architecture structured into four cooperative stages: hyper-exemplar synthesis, multi-scale hierarchical feature fusion, density map decoding with peak localization, and dual-level discrimination optimization. Given an input image and a textual category label, an MLLM jointly processes the multimodal sequence, aggregating instance appearances into the hidden states of a trailing boundary token to generate a multi-tier set of hyper-exemplars spanning shallow texture to high-level semantics. Next, a Hierarchical Fusion Strategy (HFS) integrates these multi-level hyper-exemplars into intermediate stages of the visual encoder, guiding visual representations to concentrate dynamically on category-relevant features. Finally, a convolutional decoder converts the conditioned multi-scale features into an object density map, from which Local-Maxima Detection (LMD) extracts discrete peak coordinates to filter background clutter and produce exact integer counts.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input: Image $I$ and Category Text $T$"] --> B["Hyper-Exemplar Representation Generation<br/>MLLM jointly encodes image and text to extract HERs from last $n$ layers at </s>"]
B --> C["Hierarchical Fusion Strategy<br/>HFS injects multi-level HERs progressively into vision encoder intermediate layers"]
C --> D["Local-Maxima Detection Decoding<br/>Decoder predicts density map; LMD filters background noise to yield coordinates and count"]
D --> E["Dual Optimization Objective<br/>Joint supervision via density MSE loss and GDE sample/class-level contrastive constraints"]
Key Designs¶
1. Hyper-Exemplar Representation Generation: Unifying Visual Nuances and Textual Semantics via MLLM Latent Space Conventional paradigms suffer either from detector inaccuracies when cropping sparse local patches or from the visual detachment of generic class-name embeddings. To resolve this, HER-Count uses an MLLM (specifically Qwen3-VL-2B) consisting of an image encoder \(\Phi_I\) and a multimodal large language encoder \(\Phi_T\) as the exemplar synthesizer. Given input image \(I\) and target class name \(T\), \(\Phi_I\) transforms \(I\) into visual tokens, while a special termination token \(\texttt{</s>}\) is appended to \(T\) to form \([T, \texttt{</s>}]\). The concatenated visual and textual tokens are fed into \(\Phi_T\) of depth \(L\). Through multi-head cross-modal attention, the hidden state at \(\texttt{</s>}\) aggregates holistic visual cues across all target instances guided by the textual prompt. Taking the hidden states \(\mathbf{e}_i[\texttt{</s>}]\) from the final \(n\) layers (\(i = L-n+1, \dots, L\)) and projecting them via a lightweight MLP yields a set of hierarchical hyper-exemplar representations: $\(\mathbf{F} = \{\mathbf{f}_i \mid i = L-n+1, \dots, L\}\)$ This design completely circumvents bounding box misalignments and exemplar selection errors inherent in multi-stage detectors, while capitalizing on rich MLLM pretraining priors to guarantee strong generalization on unseen test categories.
2. Hierarchical Fusion Strategy (HFS): Layer-Wise Multi-Scale Guidance Injection into Vision Encoder After obtaining the hyper-exemplar set \(\mathbf{F}\), naively modulating the final output feature map of the vision backbone provides guidance only at the tail end of processing, leaving the preceding feature extraction layers blind to category-specific visual attributes. To establish end-to-end task awareness throughout the backbone, HER-Count introduces the Hierarchical Fusion Strategy (HFS). Because the components in \(\mathbf{F}\) originate from varying depths of the MLLM, shallower embeddings capture localized textures while deeper ones encode high-level category semantics. In the visual encoder (instantiated as HRNet-W32), the feature map at the \(i\)-th stage \(P_i\) is dynamically conditioned on its corresponding hyper-exemplar \(\mathbf{f}_i\): $\(\mathbf{h}_i = P_i(\mathbf{h}_{i-1}) \odot \mathbf{f}_i\)$ where \(\odot\) denotes channel-wise feature modulation. By infusing multi-scale hyper-exemplars hierarchically from early to late layers, the visual encoder progressively suppresses irrelevant background clutter and enriches features aligned with the target class, passing a purified representation \(\mathbf{h}_L\) to the downstream convolutional decoder.
3. Global Discrimination Enhancement (GDE): Dual Sample- and Class-Level Adaptive Hard Contrastive Learning Relying solely on the final density map Mean Squared Error (MSE) loss provides only diluted, indirect supervision to the MLLM latent tokens, making them vulnerable to intra-class appearance drift or inter-class ambiguity in cluttered scenes. To impose direct metric regularization on hyper-exemplars, the Global Discrimination Enhancement (GDE) constraint optimizes the top-level exemplar \(\mathbf{f}_L\). GDE establishes dual-level contrastive anchors: individual sample exemplars serve as image-level anchors \(A^I\), while class-wise exemplar centroids define class-level anchors \(A^C\). The objective optimizes relative cosine similarities, drawing \(\mathbf{f}_L\) toward its positive anchors while repelling all negative anchors across the batch: $\(\mathcal{L}_{gde} = \overline{\log\left(1 + \frac{S_n^I}{\text{sim}(\mathbf{f}_L, A^I[p^I])}\right)} + \overline{\log\left(1 + \frac{S_n^C}{\text{sim}(\mathbf{f}_L, A^C[p^C])}\right)}\)$ where \(\text{sim}(\mathbf{u}, \mathbf{v}) = \exp(\mathbf{u}^\top \mathbf{v} / \tau)\) with temperature \(\tau = 0.05\), and \(S_n^I, S_n^C\) represent sums of negative anchor similarities. Due to the logarithmic ratio structure, the gradient magnitude naturally amplifies on hard samples with low positive similarity or close negative neighbors, dynamically enforcing tight intra-class clustering and sharp inter-class margin separation without complex sampling heuristics.
4. Local-Maxima Detection Decoding (LMD): Suppressing Background Diffuse Noise and Yielding Exact Peak Coordinates Standard density regression frameworks integrate continuous pixel predictions to calculate total object counts: \(N = \sum D\). However, widespread low-intensity background noise frequently accumulates into substantial positive count deviations, and spatial point locations remain completely obscured. Because HER-guided density maps exhibit high signal-to-noise ratios and sharp Gaussian peaks at instance centers, HER-Count incorporates Local-Maxima Detection (LMD) using a \(3 \times 3\) max-pooling filter and a response confidence threshold (set to 0.05). LMD extracts discrete spatial coordinates \([x_k, y_k]\) corresponding to genuine instance centers, determining the total count directly from the cardinality of valid local peaks. This eliminates background accumulation error and provides spatial point annotations at minimal computational cost.
Loss & Training¶
The overall training objective combines density regression and representation regularization: $\(\mathcal{L} = \mathcal{L}_{mse} + \lambda_{gde} \mathcal{L}_{gde}\)$ where \(\mathcal{L}_{mse}\) represents the pixel-wise mean squared error between predicted density map \(D\) and ground-truth Gaussian density map \(D^*\): $\(\mathcal{L}_{mse} = \frac{1}{HW} \sum_{x, y} (D(x, y) - D^*(x, y))^2\)$ The balancing weight is set to \(\lambda_{gde} = 0.1\). The visual backbone is HRNet-W32, and the decoder comprises 4 convolutional layers with 256 channels. Training inputs are resized to \(384 \times 384\). The framework is optimized with AdamW for 100 epochs on FSC147 using a batch size of 16 and weight decay of 0.05. The learning rate begins at 0.03 and follows a cosine decay schedule down to 0, completing training in roughly 5 hours on 4 NVIDIA RTX 4090 GPUs.
Key Experimental Results¶
Main Results¶
Quantitative evaluations on the standard FSC147 zero-shot benchmark against state-of-the-art methods are detailed below. Metrics include Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE):
| Method | Source | Input Size | Val MAE | Val RMSE | Test MAE | Test RMSE |
|---|---|---|---|---|---|---|
| ZSC | CVPR 2023 | 384 | 26.93 | 88.63 | 22.09 | 115.17 |
| CLIP-Count | ACM MM 2023 | 224 | 18.79 | 61.18 | 17.78 | 106.62 |
| PseCo | CVPR 2024 | 1024 | 23.90 | 100.33 | 16.58 | 129.77 |
| VA-Count | ECCV 2024 | 384 | 17.87 | 73.22 | 17.88 | 129.31 |
| DAVE (prompt) | CVPR 2024 | 512 | 15.48 | 52.57 | 14.90 | 103.42 |
| GeCo | NeurIPS 2024 | 1024 | 14.81 | 64.95 | 13.30 | 108.72 |
| YOLO-Count | ICCV 2025 | 640 | 15.43 | 58.36 | 14.80 | 96.14 |
| T2ICount | CVPR 2025 | 384 | 13.78 | 58.78 | 11.76 | 97.86 |
| HER-Count (Ours) | ECCV 2026 | 384 | 13.56 | 55.52 | 14.23 | 95.31 |
Cross-dataset generalization evaluated on the CARPK car-parking dataset demonstrates significant domain transfer robustness:
| Method | Source | #Shot | CARPK \(\to\) CARPK MAE | CARPK \(\to\) CARPK RMSE | FSC147 \(\to\) CARPK MAE | FSC147 \(\to\) CARPK RMSE |
|---|---|---|---|---|---|---|
| BMNet+ | CVPR 2022 | 3-shot | 5.76 | 7.83 | 10.44 | 13.77 |
| CounTR | BMVC 2022 | 3-shot | 5.75 | 7.45 | - | - |
| RCC | CVPR 2023 | 0-shot | 9.21 | 11.33 | 21.38 | 26.61 |
| CLIP-Count | ACM MM 2023 | 0-shot | - | - | 11.96 | 16.61 |
| VA-Count | ECCV 2024 | 0-shot | 8.75 | 10.30 | 10.63 | 13.20 |
| HER-Count (Ours) | ECCV 2026 | 0-shot | 5.19 | 6.77 | 5.56 | 7.14 |
Ablation Study¶
Ablation analysis on FSC147 assessing individual components and input modalities:
| Config | Modification | Val MAE | Val RMSE | Test MAE | Test RMSE | Note |
|---|---|---|---|---|---|---|
| Baseline HER | Top-level injection only, w/o GDE | 16.85 | 67.96 | 17.17 | 107.65 | Baseline hyper-exemplar |
| HER + HFS | Add Hierarchical Fusion Strategy | 15.53 | 59.71 | 15.71 | 96.12 | Multi-scale layer-wise alignment |
| HER + \(\text{GDE}^I\) | Add image-level contrastive loss | 15.18 | 60.35 | 16.25 | 97.97 | Strengthens intra-image invariance |
| HER + GDE | Image- + class-level complete GDE | 14.33 | 58.11 | 15.56 | 96.79 | Tight intra-class, separated inter-class |
| HER-Count (Full) | Complete model (HFS + full GDE) | 13.56 | 55.52 | 14.23 | 95.31 | Optimal synergy across all designs |
Comparison of input modalities for hyper-exemplar synthesis:
| Input Source | Description | Val MAE | Val RMSE | Test MAE | Test RMSE | Note |
|---|---|---|---|---|---|---|
| Image-only | Synthesize HER from image alone | 16.55 | 60.56 | 16.83 | 101.76 | Lacks class intent; confuses categories |
| Text-only | Synthesize HER from class name alone | 15.77 | 58.73 | 15.13 | 99.56 | Misses concrete image-specific cues |
| Multi-modal | Joint image + category text input | 13.56 | 55.52 | 14.23 | 95.31 | Customized representation, best performance |
Inference efficiency and lightweight variants (FSC147-Val, single RTX 4090 GPU, Batch Size = 1):
| Method | Parameters | Speed (FPS) | Val MAE | Val RMSE | Note |
|---|---|---|---|---|---|
| VA-Count | 0.95B | 6.4 | 17.87 | 73.22 | Multi-stage detector pipeline |
| GeCo | 1.2B | 1.1 | 14.81 | 64.95 | Heavy multi-stage resampling |
| T2ICount | 2.0B | 1.4 | 13.78 | 58.78 | Iterative diffusion sampling steps |
| HER-Count (CLIP) | 0.44B | 27.1 | 14.73 | 59.64 | Lightweight variant, fast and competitive |
| HER-Count (Qwen3-VL) | 2.1B | 14.1 | 13.56 | 55.52 | Single-pass forward, 10x faster than SOTA |
Key Findings¶
- Hierarchical fusion (HFS) is the primary driver of variance reduction: Incorporating HFS onto the baseline HER reduces Test RMSE from 107.65 to 96.12 (a reduction of 11.53 points). This proves that intermediate backbone features benefit substantially from layer-aligned multimodal cues, whereas late-stage modulation fails to steer early visual feature extraction.
- GDE strictly suppresses latent intra-class variance: Metric distance distributions confirm that GDE significantly contracts cosine distances among intra-class instances while expanding the margin between visually confusable classes, effectively resolving false-negative omissions under varied colors and lighting.
- Single-pass forward architecture shatters zero-shot counting latency bottlenecks: Whereas diffusion-based methods like T2ICount stall at 1.4 FPS, HER-Count achieves 14.1 FPS at 2.1B parameters and 27.1 FPS in its 0.44B CLIP variant, proving that performance gains stem from the hyper-exemplar formulation rather than raw parameter scaling.
Highlights & Insights¶
- Shifting from physical patch crops to continuous latent exemplars: Rather than wrestling with error-prone detectors to extract immaculate physical patches that inevitably miss intra-image variance, HER-Count innovatively abstracts exemplars into continuous latent representations synthesized by an MLLM.
- Direct adaptive contrastive supervision on MLLM tokens: Instead of treating MLLMs purely as frozen feature extractors with downstream loss backpropagation, GDE applies adaptive contrastive gradients directly to specific MLLM tokens at sample and class levels, offering a strong paradigm for other zero-shot multimodal vision tasks.
- Replacing unconstrained summation with peak detection: By generating sharp, high-SNR density maps and decoding counts via local maxima detection, HER-Count eliminates diffuse background integration drift while recovering instance center coordinates at near-zero overhead.
Limitations & Future Work¶
- Peak overlap in extremely congested clusters: In ultra-dense microscopic scenes with heavy instance occlusion, neighboring density peaks may merge after downsampling, causing undercounting in LMD filtering that may necessitate higher-resolution decoders or adaptive non-maximum suppression.
- Sensitivity to fine-grained semantic distractors: As hyper-exemplar synthesis inherits priors from MLLM pretraining, scenes with visually similar distractors (e.g., fabric polka dots vs. small round buttons) can occasionally induce subtle prompt-confusion errors, which could be mitigated via interactive negative text prompting.
Related Work & Insights¶
- vs VA-Count (ECCV 2024): VA-Count relies on explicit detectors to select 1–3 local patches, struggling when instances display variable appearances (such as bottle caps of differing colors), while running at only 6.4 FPS; HER-Count captures holistic instance variance in latent space in a single pass, delivering higher accuracy at over double the throughput.
- vs T2ICount (CVPR 2025): T2ICount relies solely on text prompts refined through multi-step diffusion decoding, lacking image-specific instance context and resulting in contiguous, blurry responses at a sluggish 1.4 FPS; HER-Count conditions hyper-exemplars on both the target image and text, producing clean, sharp density maps 10x faster.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneers multimodal hyper-exemplar synthesis via MLLMs, overcoming the physical patch versus generic text dichotomy]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive benchmarking across FSC147, tough CARPK cross-dataset zero-shot transfer, comprehensive ablations, and speed/scale audits]
- Writing Quality: ⭐⭐⭐⭐⭐ [Cohesive prose narrative, sound mathematical formulations, insightful comparative motivations, and clear visual evidence]
- Value: ⭐⭐⭐⭐⭐ [Provides a practical, noise-resistant, high-throughput (14–27 FPS) end-to-end framework for zero-shot counting and localization]