Skip to content

OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/StarLxc3/OBBSeg
Area: Medical Imaging
Keywords: Medical Image Segmentation, Weakly Supervised Learning, Oriented Bounding Box, Shape Bias, Prompt Guidance

TL;DR

To tackle the lack of orientation constraints and severe rectangular shape bias in traditional weakly supervised medical image segmentation, OBBSeg introduces an oriented bounding box (OBB) supervision framework equipped with a differentiable Mask-to-OBB (M2O) loss that eliminates geometric contour bias, alongside encoder-level prompt-assisted foreground enhancement and differential contrast modules, delivering segmentation accuracy on par with fully supervised methods at minimal annotation cost.

Background & Motivation

In computer-aided diagnosis and treatment planning, the precise delineation of anatomical structures and pathological lesion boundaries across medical images serves as the cornerstone for disease screening and quantitative prognosis. However, mainstream fully supervised segmentation networks, such as U-Net and its variants, rely heavily on dense pixel-level annotations. For lesions characterized by irregular contours, ambiguous boundaries, and high aspect ratios, manual mask delineation is notoriously labor-intensive—averaging over 90 seconds per image—presenting a fundamental bottleneck to scaling deep segmentation models in clinical practice.

To alleviate this prohibitive annotation burden, weakly supervised segmentation paradigms using points, scribbles, and axis-aligned bounding boxes (AABBs) have emerged. While these coarse annotations significantly reduce labeling effort, they provide substantially weaker geometric constraints. In particular, conventional horizontal bounding boxes fail to capture the principal axis orientation and spatial tilt of targets. When lesions are elongated or anisotropic, horizontal boxes inevitably incorporate extensive healthy background tissues; in regions with clustered or adjacent lesions, axis-aligned boxes frequently suffer from extensive spatial overlap, obscuring true instance boundaries.

Incorporating Oriented Bounding Boxes (OBBs) introduces rotation-sensitive supervision with negligible extra annotation effort (requiring only 4 corner points, taking ~17.8 s per image). Nevertheless, directly supervising soft predictions with rigid rectangular OBB masks injects severe rectangular shape bias into the network, encouraging unnatural blocky outputs. Core idea: establish OBBSeg as an intermediate supervision paradigm that leverages oriented bounding boxes, in which a differentiable Mask-to-OBB projection eliminates rectangular contour bias while preserving spatial extent and orientation, combined with encoder-stage prompt-guided foreground enhancement and differential contrast to achieve fully supervised-level segmentation under weak supervision.

Method

Overall Architecture

The overall architecture of OBBSeg is built upon a hierarchical feature backbone (such as a ViT / SAM2 encoder or Res2Net50) organized into four feature extraction stages. At each stage, the model integrates sparse prompt guidance into the encoder feature stream to suppress healthy tissue false alarms and amplify foreground-background contrast via two complementary modules. At the supervision stage, rather than imposing naive binary mask matching, OBBSeg employs a geometric projection pipeline (Mask-to-OBB loss) alongside hierarchical prompt supervision, strictly bounding the lesion's spatial span and orientation while liberating boundary delineation to data-driven representations.

The forward flow passes through the Prompt-assisted Foreground Enhancer (PAFE) and Differential-based Foreground Enhancer (DBFE), generates soft mask predictions, and is supervised end-to-end via the Mask-to-OBB transformation pipeline combined with scale-consistency constraints.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image & Multi-stage Features"] --> B["Prompt-assisted Foreground Enhancer PAFE<br/>Prompt mask filtering suppresses background"]
    B --> C["Differential-based Foreground Enhancer DBFE<br/>Foreground-background contrast amplification"]
    C --> D["Feature Decoding & Soft Mask Prediction"]
    D --> E["Mask-to-OBB Transformation<br/>Affine rotation → Axis projection → Min fusion"]
    E --> F["Joint Optimization Objectives<br/>M2O geometric loss + Hierarchical prompt loss + Scale consistency"]

Key Designs

1. Prompt-assisted Foreground Enhancer: blocking background noise in early encoder stages

Traditional promptable segmentation architectures (e.g., SAM / SAM2) typically inject geometric prompt embeddings solely into the mask decoder, leaving the deep image encoder unguided during feature extraction and prone to wasting capacity on complex background textures. OBBSeg introduces the Prompt-assisted Foreground Enhancer (PAFE), which takes user-provided prompts (points, boxes, scribbles, or circles) to generate intermediate binary prompt masks \(p_i\) dynamically resized to match feature map \(x_i\) across stages. By performing element-wise multiplication between \(p_i\) and \(x_i\), non-lesion background noise is filtered out. A residual connection preserves unmasked context to prevent accidental truncation of lesion boundaries: $\(x_e = x_i + p_i \odot x_i\)$

2. Differential-based Foreground Enhancer: explicitly magnifying weak foreground-background contrast

Medical imaging frequently exhibits low visual contrast and infiltrative margins where pathological lesions blend into adjacent healthy parenchyma. The Differential-based Foreground Enhancer (DBFE) sharpens discriminability by explicitly isolating and contrasting foreground against background representations. DBFE extracts background features by applying the inverted prompt mask \((1 - p_{i+1})\) to feature \(x_e\), followed by spatial average pooling to produce a global background prototype \(x_e'\). It then subtracts \(x_e'\) from \(x_e\) to derive a differential representation, which is gated by the foreground mask \(p_{i+1}\) and added back residually: $\(x_{i+1} = x_e + (x_e - x_e') \odot p_{i+1}\)$ This operation acts as a local feature-level background subtraction, expanding the margin between lesion signatures and healthy tissues.

3. Mask-to-OBB (M2O) Loss: eliminating rectangular shape bias via geometric projection

Directly penalizing segmentation predictions against box-shaped OBB ground truth forces models to memorize rectangular contours. The Mask-to-OBB (M2O) loss breaks this coupling by projecting predicted instance masks into a coordinate-aligned 1D manifold before computing the loss: - Instance Decomposition & Affine Rectification: Given predicted soft mask \(s\) and target OBB masks \(\{b_m\}_{m=1}^M\), instance sub-masks \(s_m = s \odot b_m\) are extracted. Using the orientation angle \(\theta_m\) of \(b_m\), an affine transformation matrix \(R_m\) rotates \(s_m\) clockwise to align its principal axes horizontally and vertically: \(s_a = \text{Warp}(s_m, R_m)\). - Axis-wise Max Projection: Axis-aligned mask \(s_a\) is max-projected onto horizontal and vertical axes, yielding 1D maximum probability vectors \(s_h = \max(s_a, \text{axis}=1)\) and \(s_v = \max(s_a, \text{axis}=0)\). This step entirely removes inner contours and boundary irregularities, retaining only the true spatial extent and center offset. - Minimum Back-projection & Inverse Affine Transformation: Vectors \(s_h\) and \(s_v\) are replicated across spatial dimensions and merged via element-wise minimum: \(s_{hw} = \min(s_h, s_w)\). Applying the inverse affine transformation \(R_m^{-1}\) rotates \(s_{hw}\) counterclockwise by \(\theta_m\) back to the original orientation, yielding reconstructed instance mask \(s_o\). - Aggregation & Loss Computation: All instance predictions are combined via element-wise maximum: \(s_t = \max(s_o^1, \dots, s_o^M)\). Because \(s_t\) is transformed into a clean bounding envelope, computing Binary Cross-Entropy and Dice loss against ground-truth OBB \(y\) enforces geometric orientation and bounds without inducing rectangular boundary artifacts.

4. Hierarchical Prompt Supervision: multi-stage alignment across diverse prompt modalities

To supplement sparse OBB supervision and stabilize multi-scale feature hierarchies, OBBSeg applies a hierarchical prompt supervision loss (\(\mathcal{L}_{\text{PS}}\)). The loss aligns intermediate stage predictions \(p_i\) as well as final prediction \(s\) with user prompt \(p\): $\(\mathcal{L}_{\text{PS}} = \mathcal{L}(s, p) + \sum_{i=1}^4 \mathcal{L}(p_i, p)\)$ The supervision operator \(\mathcal{L}\) adaptively accommodates point, scribble, box, and circular prompts, computing gradients exclusively on verified prompt locations to deliver noise-free guidance throughout network depths.

Loss & Training

The overall training objective combines the M2O loss \(\mathcal{L}_{\text{M2O}}\), hierarchical prompt loss \(\mathcal{L}_{\text{PS}}\), and scale consistency loss \(\mathcal{L}_{\text{SC}}\): $\(\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{M2O}} + \mathcal{L}_{\text{PS}} + \lambda \mathcal{L}_{\text{SC}}\)$ where \(\mathcal{L}_{\text{M2O}} = \mathcal{L}_{\text{BCE}}(s_t, y) + \mathcal{L}_{\text{Dice}}(s_t, y)\), and \(\mathcal{L}_{\text{SC}}\) enforces multi-scale structural agreement following WeakPolyp.

Training was conducted on two NVIDIA RTX 4090 GPUs using AdamW (initial learning rate \(1\times 10^{-4}\), weight decay \(5\times 10^{-4}\), batch size 16) for 16 epochs. When utilizing the SAM2 backbone, the pre-trained encoder weights remain frozen, and only OBBSeg-specific enhancement and decoder modules are trained, keeping computational overhead minimal.

Key Experimental Results

Main Results

On five standard polyp segmentation benchmarks (ClinicDB, ColonDB, ETIS, Kvasir, Endo) and the SUN-SEG video polyp dataset, OBBSeg was benchmarked against fully supervised and weakly supervised baselines (evaluated by Dice %):

Method Year/Venue Supervision ClinicDB ColonDB ETIS Kvasir Endo FSPD Avg. SUN-SEG Avg.
U-Net MICCAI 2015 Full 82.3 51.2 39.8 81.8 71.0 56.1 -
PraNet MICCAI 2020 Full 89.9 70.9 62.8 89.8 87.1 74.0 67.7
CASCADE WACV 2023 Full 94.3 82.5 80.1 92.6 90.5 84.7 -
EMCAD CVPR 2024 Full 95.2 92.3 92.3 92.8 - 92.6 -
MedSAM2 (Box) 2025 Full + Prompt 94.6 91.7 93.2 94.4 93.7 92.8 -
WeakPolyp MICCAI 2023 Weak 85.2 74.5 71.1 86.1 84.8 76.7 79.8
WeakPolyp + M2O Ours (Plug-in) Weak (OBB) 89.2 75.9 72.4 90.0 90.0 79.0 81.0
SAM2 + OBBSeg (Point) Ours Weak + Prompt 91.9 87.3 86.9 91.2 92.1 88.5 90.0
SAM2 + OBBSeg (Scribble) Ours Weak + Prompt 92.4 89.2 90.4 94.2 92.6 90.6 92.1
SAM2 + OBBSeg (Box) Ours Weak + Prompt 95.1 93.1 92.7 95.7 94.3 93.6 94.7
SAM2 + OBBSeg (Circle) Ours Weak + Prompt 95.1 93.8 92.9 95.7 94.8 94.0 95.0

Across cross-modality datasets, OBBSeg demonstrated consistent generalizability: achieving 94.4% average Dice on skin lesion datasets (ISIC2017 / ISIC2018), 92.0% average Dice on ultrasound nodule and tumor benchmarks (TN3K / DDTI / BUSI), and 88.2% average Dice on the multi-organ CT Synapse dataset.

Ablation Study

Component-wise ablation on the five standard polyp datasets shows the individual and compounding benefits of each proposed design (Dice %):

Config Base (SAM2+OBB) \(\mathcal{L}_{\text{M2O}}\) \(\mathcal{L}_{\text{PS}}\) DBFE PAFE Prompt-free +Point +Scribble +Box +Circle
1 ✓ 77.2 - - - -
2 ✓ ✓ 80.5 - - - -
3 ✓ ✓ ✓ 83.3 78.2 78.5 82.4 82.5
4 ✓ ✓ ✓ 81.3 - - - -
5 ✓ ✓ ✓ ✓ 83.6 81.7 86.4 91.5 91.8
6 (Full model) ✓ ✓ ✓ ✓ ✓ - 88.5 90.6 93.6 94.0

Additionally, for elongated anisotropic lesions (aspect ratio > 2), adding the M2O loss yielded an impressive +5.3% Dice improvement (FSPD: 65.9% \(\to\) 71.2%; SUN-SEG: 76.2% \(\to\) 81.5%), confirming the distinct geometric superiority of orientation-aware supervision.

Key Findings

  • De-biasing Impact of M2O Loss: In a purely prompt-free setting, substituting conventional bounding box supervision with OBB + M2O loss boosted the base model from 77.2% to 80.5% (+3.3%). Directly integrating M2O into WeakPolyp improved average Dice from 76.7% to 79.0% while achieving a 5.3% gain on elongated lesions, demonstrating that M2O successfully eliminates rectangular bias.
  • Synergistic Gain of PAFE & DBFE: DBFE provides modest gains when unguided (80.5% \(\to\) 81.3%), but once prompt supervision \(\mathcal{L}_{\text{PS}}\) provides reliable foreground seeds, DBFE and PAFE jointly elevate Box-prompted performance from 82.4% to 93.6%, matching top fully supervised models.
  • Robustness to Annotation Noise: Angular perturbations within \(\pm 30^\circ\) resulted in less than 2.3% Dice variation, and padding expansions up to 15 pixels caused negligible degradation. In real human annotation evaluations (21.20% deviation), WeakPolyp plummeted by 12.2%, whereas WeakPolyp+M2O degraded by only 1.5%, exhibiting outstanding clinical noise tolerance.

Highlights & Insights

  • Differentiable Affine Projection Pipeline: By combining rotation, axis-wise maximum projection, and inverse mapping, the M2O loss effectively scrubs internal and edge contour information while retaining spatial span and orientation, formulating an elegant, general-purpose supervision mechanism free of boundary shape bias.
  • Early-stage Encoder Prompt Integration: Deviating from the convention of injecting prompts solely into the mask decoder, OBBSeg introduces prompt-driven filtering and differential background subtraction within early encoder stages, preventing feature corruption in low-contrast medical backgrounds.
  • High Practical Transferability: Because M2O relies entirely on standard matrix transformations and incurs zero inference compute, it can be seamlessly ported as a drop-in loss to other CNN/Transformer architectures and extended to remote sensing or defect inspection where anisotropic targets abound.

Limitations & Future Work

  • Representational Bottleneck on Complex Topologies: Because M2O relies on 1D continuous maximum projection along orthogonal axes, it can oversimplify complex branching structures (e.g., retinal vascular trees or bronchial networks) where topological cavities and multi-path connectivity are essential.
  • Lack of Native 3D Volumetric and Temporal Modeling: The current framework is optimized for static 2D slices without native support for multi-phase dynamic enhancement or 3D volumetric OBB bounding surfaces.
  • Promising Next Steps: Developing adaptive contour evolution strategies under hybrid weak labels (e.g., sparse points + OBB) and extending oriented bounding geometry to 3D tensor fields for volumetric clinical imaging.
  • vs WeakPolyp: WeakPolyp relies on horizontal bounding boxes (AABBs) with projection-backprojection, which inevitably captures substantial background when lesions are angled or elongated. OBBSeg introduces orientation-aware OBBs with affine rectification, outperforming WeakPolyp by 5.3% on slender lesions.
  • vs MedSAM / SAM2: Native SAM2 treats prompts as decoding conditions, lacking encoder-level background pruning and automated prompt-free capability. OBBSeg embeds prompt filtering directly into the backbone (PAFE + DBFE), maintaining state-of-the-art segmentation across both interactive and automated clinical inference modes.

Rating

  • Novelty: ⭐⭐⭐⭐ [Systematically introduces oriented bounding boxes into weakly supervised medical image segmentation with a mathematically elegant, differentiable M2O projection loss]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across 5 imaging modalities, 13 datasets, 20 competing methods, complemented by human annotation timing and perturbation studies]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, self-contained mathematical formulation, and exceptionally well-structured empirical analyses]
  • Value: ⭐⭐⭐⭐⭐ [Drastically reduces clinical annotation overhead while achieving segmentation performance matching fully supervised baselines]