Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/jinyanglii/Ceptor
Area: Multimodal VLM
Keywords: open-set object detection, vision-language model, visual prompt, Pyramid RoI Extractor, harmonization strategies
TL;DR¶
Ceptor draws inspiration from human visual search by employing a pre-trained vision-language model (VLM) as both the image backbone and multi-prompt generator, integrating a spatial Pyramid RoI Extractor and Multi-Route FFN to achieve unified, universal open-set object detection across interactive visual, generic visual, textual, and hybrid prompts.
Background & Motivation¶
Open-set object detection (OSOD) aims to transcend fixed predefined category vocabularies, serving as an indispensable foundation for autonomous driving, industrial anomaly inspection, medical diagnostics, and vision-centric modules in multimodal large language models. Mainstream text-guided open-vocabulary detectors heavily rely on cross-modal alignment within the pre-trained vision-language embedding space. However, when encountering long-tail categories, domain-specific objects, or physical instances that lack semantic abstraction and defy precise verbal description, the textual feature space turns sparse due to language discreteness, causing detector responsiveness to text guidance to deteriorate dramatically.
To bypass the bottleneck of text prompts on poorly abstracted categories, recent endeavors have investigated coordinate-based visual prompting frameworks (such as T-Rex2 and YOLOE). Nonetheless, existing visual prompt methods exhibit critical limitations: first, both the perception backbone and prompt generation lack intrinsic generalization, relying on naive region crops or raw coordinates that lose contextual cues and inflate computational overhead; second, text and visual prompt pathways remain isolated without end-to-end coordinated optimization, lacking intermediate prompting mechanisms to navigate divergent category distributions; third, cross-modal representation disparities introduce optimization conflicts and gradient interference when fused inside deep decoder layers.
In contrast, human visual search integrates heterogeneous perceptual signals flexibly: when targeting familiar concepts, the brain activates high-level semantic labels; when searching for unfamiliar or complex long-tail items, humans leverage visual exemplars and dense sensory templates to perform robust comparative matching. The core idea is to employ a pre-trained vision-language model (VLM) simultaneously as the detector backbone and diverse prompt generator, building a shared, generalizable feature space where visual and textual prompts are inherently aligned or derived from the same source, orchestrated by a Pyramid RoI Extractor (PRE) and Multi-Route expert harmonization network for unified multi-prompt open-set detection.
Method¶
Overall Architecture¶
Ceptor is built upon the Deformable DETR architecture, composed of three collaborative core components: the VLM Infuser, the Prompt Zoo, and the Prompt-Oriented Decoder (POD). The input image is processed by the VLM Infuser, wherein the VLM's Image Encoder (IE) and a lightweight convolutional stream form a Spatial Enhancer to inject multi-scale spatial inductive biases. Diverse prompt candidates are synthesized within the Prompt Zoo directly from the VLM, engaging in Early Interaction with image features via multi-head cross-attention (MHCA). Subsequently, object queries in the POD interact through self-attention, image cross-attention, and dedicated prompt-cross-attention, before being dispatched to a Multi-Route Feed-Forward Network (Multi-Route FFN) to generate final box coordinates and classification logits.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image & Multi-Modal Cues"] --> B["VLM Infuser & Spatial Enhancer<br/>IE extraction + lightweight Conv stream + SFP multi-scale pyramid"]
A --> C["Prompt Zoo Generation<br/>TE text / PRE interactive visual / offline generic visual / hybrid prompt"]
B & C --> D["Early Cross-Modal Interaction<br/>Enhanced image features and prompt features interact via MHCA"]
D --> E["Efficient Encoder Fusion<br/>Top-level feature MHSA + CCFM cross-scale fusion"]
E & C --> F["Prompt-Oriented Decoder<br/>Query self-attention + image cross-attention + prompt-cross-attention"]
F --> G["Multi-Route Expert FFN & Harmonized Output<br/>Dispatch via instruction embeddings for box regression & classification"]
G --> H["Cross-Modal Instance-Level Feature Alignment<br/>L_align distillation between predicted features F_align and GT visual embeddings V_G'"]
Key Designs¶
1. VLM Infuser & Spatial Enhancer: Inductive transfer from plain ViT to multi-scale detection representations Pre-trained plain ViT backbones maintain single-scale features without explicit spatial inductive biases, hindering dense multi-scale localization in Deformable DETR. Ceptor introduces a Spatial Enhancer pairing a Simple Feature Pyramid (SFP) with a lightweight convolutional network \(C\). The input image \(I\) is encoded into features \(F_I\) by the Image Encoder (IE) to construct SFP, which are concatenated at each level \(l\) with convolutional features \(C(I)^l\), followed by a \(1\times 1\) convolution and normalization to yield enhanced multi-scale features \(F_{\text{enh}}^l\). In the Efficient Encoder stage, multi-head self-attention and FFN are applied solely to the top-level feature \(F_{\text{fused}}^L\), supplemented by a cross-scale feature fusion module (CCFM) to reinforce semantic consistency across levels at minimal computational overhead.
2. Pyramid RoI Extractor (PRE): Context-preserving interactive visual prompt extraction Traditional visual prompting relies on raw bounding box coordinates or cropped patches passed through extra encoders, which either misses context or drastically inflates latency. For interactive visual prompts \(V_I\), Ceptor designs the Pyramid RoI Extractor. PRE first deploys a lightweight CCFM to smooth ViT patch-level features into a continuous pyramid \(\{F_{\text{pre}}^m\}_{m=1}^M\), allowing arbitrary bounding box scales to map precisely to optimal resolution levels. Multi-scale RoI Align is executed across levels, naturally inheriting the wide contextual awareness encoded by ViT's global self-attention. The multi-scale region tokens are then concatenated along the channel dimension, blended via a \(1\times 1\) convolution, and downsampled through dedicated convolutional layers to preserve spatial structure, yielding compact, discriminative interactive visual embeddings \(V_I'\).
3. Diverse Guidance Ecosystem: From single-shot exemplar queries to generic cross-image retrieval The Prompt Zoo accommodates four synergistic prompt modalities to cover arbitrary detection requirements: - Text Prompts (\(T\)): Extracted from 80 category templates using the frozen Text Encoder (TE) and normalized, providing global, high-level abstract conceptual guidance; - Interactive Visual Prompts (\(V_I\)): Generated on-the-fly from user-specified bounding boxes on the current image using PRE, tailored for intra-image interactive segmentation, object counting, and user-in-the-loop annotation; - Generic Visual Prompts (\(V_G\)): Extracted offline from multiple cropped positive exemplars across images via the frozen IE and clustered into class centroid vectors, designed for open-set cross-image search of long-tail entities; - Hybrid Prompts (\(H\)): Formed by averaging and normalizing generic visual embeddings \(V_G\) and text embeddings \(T\), uniting visual specificity with lexical abstraction for heightened robustness under ambiguous category boundaries.
4. Multi-Route FFN Experts & Harmonization Strategies: Mitigating multi-prompt gradient conflicts Diverse prompt modalities inhabit distinct manifold distributions; forcing all prompt streams through shared feed-forward layers causes parameter contention and optimization instability. Ceptor incorporates a Multi-Route FFN in the decoder, allocating four dedicated feed-forward networks acting as specialized experts for each prompt modality. Concurrently, four independent learnable instruction embeddings are introduced to condition the network explicitly on the active prompt pathway. Furthermore, an alignment mechanism matches decoder instance features \(F_{\text{align}}\) against frozen ground-truth IE embeddings \(V_G'\) via bipartite matching: $\(\mathcal{L}_{\text{align}} = \frac{1}{n \cdot d} \sum_{i=1}^n \sum_{j=1}^d \left( F_{i,j,\text{align}} - V'_{i,j,\text{align}} \right)^2\)$ This objective distills VLM open-set generalization into the output heads and aligns latent spaces across different prompt routes from the prediction boundary.
Loss & Training¶
Ceptor optimizes bounding box coordinates and classification via Hungarian bipartite matching. The overall objective balances classification contrastive loss, box regression losses, and feature alignment: $\(\mathcal{L} = \lambda_{\text{cls}} \mathcal{L}_{\text{cls}} + \lambda_{\text{L1}} \mathcal{L}_{\text{L1}} + \lambda_{\text{GIoU}} \mathcal{L}_{\text{GIoU}} + \lambda_{\text{align}} \mathcal{L}_{\text{align}}\)$ The classification loss \(\mathcal{L}_{\text{cls}}\) adopts contrastive focal loss computed via the dot product of output query embeddings and prompt vectors. Training is conducted in two progressive stages: Phase 1 conducts text-guided pre-training on Objects365, V3Det, and GoldG (1.56M labeled images), fine-tuning the IE backbone with a 0.05 learning rate factor at \(1024\times 1024\) resolution; Phase 2 executes full multi-prompt joint training on Objects365, OpenImages, V3Det, and HierText (2.5M images), alternating generic visual and text iterations while periodically interleaving interactive visual and hybrid batches every 5 cycles.
Key Experimental Results¶
Main Results¶
Ceptor was evaluated on COCO-val and the challenging LVIS benchmark (minival and full val set) across interactive visual (Visual-I), generic visual (Visual-G), text (Text), and hybrid (Hybrid) modes, consistently outperforming state-of-the-art open-set detectors.
| Method | Open Source | Backbone | Data Scale | Mode | COCO AP | LVIS-minival AP_all | LVIS-minival AP_r | LVIS-val AP_all | LVIS-val AP_r |
|---|---|---|---|---|---|---|---|---|---|
| T-Rex2 | ✗ | Swin-T | 7.3M | Visual-I | 56.6 | 59.3 | 64.4 | 62.6 | 71.9 |
| Ceptor | ✓ | PEcoreS | 3.2M | Visual-I | 65.1 (+8.5) | 66.9 (+7.6) | 77.0 (+12.6) | 63.0 (+0.4) | 74.6 (+2.7) |
| T-Rex2 | ✗ | Swin-L | 7.3M | Visual-I | 58.5 | 62.5 | 70.1 | 65.8 | 72.6 |
| Ceptor | ✓ | PEcoreL | 3.5M | Visual-I | 71.9 (+13.4) | 72.8 (+10.3) | 81.3 (+11.2) | 68.8 (+3.0) | 80.2 (+7.6) |
| T-Rex2 | ✗ | Swin-T | 7.3M | Visual-G | 38.8 | 37.4 | 29.9 | 34.9 | 32.4 |
| YOLOE-11-L | ✓ | - | 1.4M | Visual-G | - | 33.7 | 28.1 | - | - |
| Ceptor | ✓ | PEcoreS | 3.2M | Visual-G | 47.3 (+8.5) | 39.1 (+1.7) | 37.0 (+7.1) | 30.9 | 28.8 |
| Ceptor | ✓ | PEcoreL | 3.5M | Visual-G | 54.0 (+7.5) | 47.8 (+0.2) | 47.2 (+1.8) | 40.5 | 40.5 |
| Grounding DINO | ✗ | Swin-T | 5.4M | Text | 48.4 | 27.4 | 18.1 | - | - |
| T-Rex2 | ✗ | Swin-T | 7.3M | Text | 45.8 | 42.8 | 37.4 | 34.8 | 29.0 |
| Ceptor | ✓ | PEcoreS | 3.2M | Text | 51.7 (+5.9) | 43.1 (+0.3) | 40.6 (+3.2) | 34.3 | 28.2 |
| Ceptor | ✓ | PEcoreL | 3.5M | Text | 54.8 (+2.6) | 54.0 | 54.7 (+5.5) | 45.5 | 42.8 |
| T-Rex2 | ✗ | Swin-T | 7.3M | Hybrid | 42.5 | - | - | 37.0 | 34.3 |
| Ceptor | ✓ | PEcoreS | 3.2M | Hybrid | 50.5 (+8.0) | 44.1 | 40.9 | 35.9 | 31.4 |
Ablation Study¶
Ablation experiments conducted on Objects365 pre-training using the PEcoreS backbone demonstrate the individual efficacy of the core architectural contributions:
| Mode | Model Config | COCO-val AP | LVIS-minival AP_all | LVIS-minival AP_r | LVIS-minival AP_c | LVIS-minival AP_f | Note |
|---|---|---|---|---|---|---|---|
| Visual-I | Ceptor Full Model | 62.3 | 61.1 | 72.4 | 68.5 | 52.5 | full model |
| Visual-I | w/o Infusion Strategies | 61.0 (-1.3) | 57.5 (-3.6) | 66.7 (-5.7) | 64.0 (-4.5) | 50.1 (-2.4) | Missing SFP and early alignment drops rare categories steeply |
| Visual-I | w/o Harmonization Strategies | 61.3 (-1.0) | 59.2 (-1.9) | 69.5 (-2.9) | 66.0 (-2.5) | 51.3 (-1.2) | Omitting Multi-Route FFN and instruction tags harms multi-tasking |
| Visual-I | w/o VLM (standard ViT) | 61.5 (-0.8) | 53.6 (-7.5) | 59.7 (-12.7) | 59.9 (-8.6) | 47.0 (-5.5) | Lacking pre-trained VLM prior drops rare classes by 12.7 AP |
| Visual-I | w/o PRE (w/ Cross-Attention) | 59.3 (-3.0) | 59.3 (-1.8) | 70.3 (-2.1) | 65.4 (-3.1) | 51.9 (-0.6) | Validates PRE in retaining instance context and spatial precision |
| Visual-G | Ceptor Full Model | 48.1 | 31.5 | 25.6 | 30.2 | 33.7 | full model |
| Visual-G | w/o Infusion Strategies | 47.9 (-0.2) | 30.8 (-0.7) | 24.9 (-0.7) | 29.5 (-0.7) | 33.0 (-0.7) | Backbone feature representation degraded |
| Visual-G | w/o Harmonization Strategies | 46.5 (-1.6) | 29.6 (-1.9) | 24.1 (-1.5) | 27.8 (-2.4) | 32.1 (-1.6) | Severe cross-modal interference |
| Visual-G | w/o VLM (standard ViT) | 47.6 (-0.5) | 30.4 (-1.1) | 21.4 (-4.2) | 29.2 (-1.0) | 33.0 (-0.7) | Generic long-tail generalization compromised |
| Text | Ceptor Full Model | 50.2 | 33.7 | 32.2 | 31.9 | 35.7 | full model |
| Text | w/o Infusion Strategies | 48.9 (-1.3) | 32.8 (-0.9) | 30.4 (-1.8) | 31.0 (-0.9) | 34.9 (-0.8) | Weaker text-to-vision grounding |
| Text | w/o Harmonization Strategies | 48.1 (-2.1) | 32.3 (-1.4) | 31.1 (-1.1) | 30.8 (-1.1) | 33.9 (-1.8) | Multi-Route FFN essential for text queries |
| Text | w/o VLM (standard ViT) | 49.1 (-1.1) | 33.1 (-0.6) | 30.5 (-1.7) | 31.6 (-0.3) | 34.9 (-0.8) | Loss of intrinsic vision-language alignment |
Key Findings¶
- Same-Source Visual Representations Excel at the Long Tail: On LVIS rare classes (\(AP_r\)), interactive visual prompts with PEcoreS reached 77.0 AP, dramatically outpacing pure text (40.6 AP) and generic visual prompts (37.0 AP). Replacing the VLM with standard ViT caused a 12.7 AP plummet on rare categories under Visual-I, proving that dense, continuous feature representations from VLM image encoders are vital for capturing long-tail visual similarity.
- Synergistic Complementarity in Hybrid Prompts: On LVIS-minival, the hybrid prompt achieved 44.1 \(AP_{\text{all}}\), exceeding both pure text prompts (43.1 AP) and generic visual prompts (39.1 AP). This confirms that pairing high-level semantic abstractions with instance-level visual cues resolves category ambiguities effectively.
- Superiority of the PRE Pipeline: Replacing PRE with raw cross-attention over full image features caused a 3.0 AP drop on COCO-val, demonstrating the necessity of multi-scale RoI pooling combined with structured convolutional downsampling.
Highlights & Insights¶
- Unified VLM Origin: Rather than treating backbones and prompt generators as separate disconnected sub-networks, Ceptor deploys a single VLM to serve dual roles, guaranteeing inherent latent alignment and feature homogeneity.
- Crop-Free Pyramid RoI Extractor: By harmonizing feature smoothing CCFM and multi-level RoI Align, PRE eliminates costly image cropping and re-encoding while preserving global contextual cues via ViT's self-attention.
- Modular Multi-Prompt Scalability: Ceptor proves that multi-route expert FFNs and instruction embeddings decouple multi-modal training dynamics, allowing hybrid prompt mixtures to converge smoothly and offering an ideal foundation for automated prompt-agent workflows.
Limitations & Future Work¶
- Limitations Admitted by Authors: On the complete LVIS-val benchmark under generic visual guidance (Visual-G), Ceptor's 30.9 AP trails T-Rex2's 34.9 AP, which benefited from over 7.3M training samples including proprietary high-quality annotations. This indicates that generic visual cluster prototypes remain sensitive to training set scale and label purity.
- Future Directions: Beyond static centroid averaging for generic visual prompts, future iterations could leverage MLLMs to generate adaptive category visual descriptors; additionally, extending PRE to support point clicks, scribbles, and mask exemplars would broaden practical deployment scenarios.
Related Work & Insights¶
- vs Grounding DINO / OV-DINO: These detectors rely exclusively on text prompts, performing well on common vocabulary but failing when encountering undescribed long-tail entities; Ceptor supplements textual guidance with interactive and generic visual prompts, achieving dominant gains on rare classes.
- vs T-Rex2: While T-Rex2 pioneered multi-modal prompt synergies, it relies on proprietary code, decouples modalities during inference, and lacks multi-scale contextual RoI extraction; Ceptor delivers an end-to-end trained open-source framework with superior data efficiency.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ Elegantly repurposes pre-trained VLM representations for dual backbone-prompt roles; PRE and multi-route harmonization are soundly designed.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive benchmarks across COCO and LVIS, covering four guidance routes with extensive component ablations.
- Writing Quality: ⭐⭐⭐⭐⭐ Rigorous logical progression from human visual cognition to architectural realization, supported by transparent visualizations.
- Value: ⭐⭐⭐⭐⭐ Provides an open-source, versatile "detect-anything" foundation with strong zero-shot and interactive capabilities for robotics and vision agents.