iMED: A Multi-Endoscope Dataset for Surgical 3D Perception¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/surgical-vision/imed
Area: Medical Imaging
Keywords: Surgical 3D Perception / Multi-Endoscope Dataset / Novel View Synthesis / Robot-Assisted Minimally Invasive Surgery / Geometric Generalization
TL;DR¶
Addressing the fundamental limitation of surgical vision benchmarks confined to narrow single-trajectory forward motion, iMED introduces the first synchronized dual-stereo-endoscope dataset (340 sequences, ~170K timepoints) and a cross-trajectory evaluation protocol across 23 SOTA algorithms, revealing photometric collapse under held-out viewpoints and the remarkable transferability of multi-domain foundation models.
Background & Motivation¶
Non-rigid scenes observed under severely constrained camera motion constitute one of the most hostile operating regimes for visual 3D perception, including camera pose estimation, feature matching, optical flow, depth estimation, and novel view synthesis (NVS). In robot-assisted minimally invasive surgery (RAMIS), these challenges reach an extreme: endoscopes operate inside enclosed anatomical cavities via narrow trocar ports, following constrained, forward-facing trajectories through deforming tissue plagued by specular reflections, smoke, bleeding, and topological tearing. For years, the 3D scene representation community has relied on photometric metrics like PSNR and SSIM evaluated on held-out frames along the same trajectory. However, such metrics act as proxies for local image reconstruction rather than measures of true geometric consistency: high photometric fidelity frequently reflects dense temporal interpolation between adjacent frames without demonstrating whether the model has uncovered the underlying 3D physical surface.
In general computer vision and dynamic human capture, measuring true geometric generalization is achieved via multi-camera rigs where models train on a subset of viewpoints and evaluate on physically independent, held-out cameras. Unfortunately, physiological space limitations and surgical robot hardware constraints have made multi-camera endoscopic setups exceptionally difficult to deploy in clinical environments. Consequently, established surgical datasets (e.g., SCARED, SERV-CT, STIR, StereoMIS) remain restricted to single-endoscope recordings along a single forward-facing path, leaving surgical vision without any benchmark capable of cross-trajectory geometric validation.
To bridge this fundamental gap between photometric interpolation artifacts and rigorous 3D geometric generalization, the core idea is to build a hardware-synchronized dual-stereo-endoscope platform with miniature ArUco fiducials across ex vivo, postmortem, and live surgical tissues, introducing a "train-on-one-endoscope, test-on-another" held-out evaluation protocol to genuinely benchmark surgical 3D perception.
Method¶
Overall Architecture¶
The iMED methodology encompasses hardware-synchronized dual-stereo capture yielding four concurrent HDMI streams, spatial-temporal geometric calibration via custom miniature ArUco cubes and a five-stage trajectory estimation pipeline, an isolated cross-endoscope benchmarking protocol, and comprehensive multi-task 3D perception evaluation across 23 state-of-the-art methods. The pipeline proceeds from synchronized acquisition to sub-millimeter reference pose estimation, enabling rigorous held-out geometric testing.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Synchronized Dual-Stereo Acquisition<br/>4 Concurrent HDMI Streams (60 FPS)"] --> B["Dual-Endoscope Acquisition Platform<br/>Ex vivo / Postmortem / In vivo Sessions"]
B --> C["Custom ArUco Markers & 5-Stage Trajectory Pipeline<br/>PnP + Stereo Fusion + Filtering & Uncertainty"]
C --> D["Cross-Endoscope Generalization Protocol<br/>Train on Endoscope 2 / Test on Endoscope 1"]
D --> E["Multi-Task 3D Perception Benchmarking<br/>Evaluate 23 SOTA Methods Across NVS & Geometry"]
Key Designs¶
1. Dual-Endoscope Acquisition Platform: Wide-Baseline Multi-View Geometry in Clinical Settings
To overcome the lack of viewpoint diversity in single-endoscope setups, the authors constructed a multi-endoscope capture system in a clinical lab utilizing three surgical systems with three 8mm 0° and four 8mm 30° stereo endoscopes. Endoscope 1 is mounted on an articulated passive arm and kept stationary during acquisition, while Endoscope 2 is mounted on the active robotic patient cart with integrated illumination. To preserve authentic clinical lighting distributions, only Endoscope 2's illumination is activated, and photometric differences between endoscopes are normalized using sequence-level statistics. The two hardware-synchronized stereo endoscopes output four HDMI streams (60 FPS) separated by an average baseline of \(2.6 \pm 1.2\text{ cm}\) with an orientation difference of \(18.7^\circ \pm 10.0^\circ\). This inter-camera separation is an order of magnitude larger than the internal 4mm stereo baseline, spanning 20% to 50% of the typical 6–13 cm laparoscopic surgical field of view. The dataset spans 10 ex vivo organ specimens (chicken, porcine, bovine; ~7 sequences each), 3 human cadaver subjects (~6 anatomical sites each), and 1 live porcine subject (10 anatomical sites), totaling 340 sequences (~170K synchronized timepoints) with an exceptional mean inter-frame synchronization error of only \(3.36\text{ ms}\) (\(\text{SD} = 3.90\text{ ms}\)).
2. Custom ArUco Markers & 5-Stage Trajectory Pipeline: Sub-Millimeter Reference Pose Tracking
Because standard surgical trocars are constrained to diameters \(\le 12\text{ mm}\) and abdominal fluids readily cause reflection and staining, standard calibration targets cannot be used. The authors developed custom 8mm hard-wooden calibration cubes laser-etched with 5.7mm ArUco markers (ARUCO_MIP_36h12 dictionary), easily manipulated by robotic instruments and adhering steadily to tissue surfaces via surface tension. The 6-DOF reference camera trajectories are computed via a five-stage pipeline: - RANSAC Initialization & Refinement: Detected marker corners are solved via Perspective-n-Point (PnP); inlier hypotheses filtered by RANSAC are refined using Levenberg-Marquardt optimization minimizing bidirectional reprojection error across all corners; - Stereo Fusion: Left-to-left and right-to-right transformations between endoscopes are independently solved and fused via spherical quaternion interpolation for global consistency; - Outlier Detection: Operating under known surgical kinematics where expert laparoscopic tool speeds stay below \(8.5\text{ cm/s}\) (\(< 0.14\text{ cm/frame}\) at 60 FPS), physically impossible pose jumps (\(> 0.7\text{ cm/frame}\) translation or \(> 3.0^\circ/\text{frame}\) rotation) are rejected; - Temporal Smoothing: Translation axes are filtered using a Savitzky-Golay filter (\(W = 15, p = 1\)), while rotations are smoothed in quaternion space with sign-continuity enforcement and SLERP interpolation for missing frames; - Uncertainty Quantification: Geometric consistency between independent stereo pairs provides frame-wise translation and rotation confidence bounds. At a 5cm working distance, the pipeline achieves 0.58 mm physical accuracy (11.87 px reprojection error) on live tissue.
3. Cross-Endoscope Generalization Protocol: Disentangling True 3D Geometry from Photometric Interpolation
Conventional novel view synthesis benchmarks evaluate on held-out frames sparsely sampled along the training trajectory, allowing neural networks to achieve high PSNR/SSIM purely via temporal radiance interpolation. To eliminate this shortcut, iMED enforces a "train-on-Endoscope-2, test-on-Endoscope-1" protocol. Because Endoscope 1 occupies a physically separate port with an ~18.7° viewpoint offset and non-overlapping optical centers, models cannot rely on memorized ray proximities. This benchmark forces algorithms to extrapolate consistent 3D geometry rather than overfitting view-dependent photometric cues, serving as the first strict spatial generalization benchmark in surgical computer vision.
4. Multi-Task 3D Perception Benchmarking: Geometric Regularization & Multi-Domain Foundation Models
Across rigid NVS, deformable NVS, relative pose/feature matching, and monocular depth estimation, iMED benchmarks 23 state-of-the-art algorithms under this held-out-view protocol. The findings expose critical architectural trade-offs: for coordinate MLPs (NeRF), monocular depth supervision induces semi-transparent "ghost" volumetric artifacts, causing non-monotonic SSIM degradation during camera interpolation; conversely, explicit 3D Gaussian Splatting benefits from depth priors. In deformable dynamics, purely photometric models (D-3DGS) collapse on held-out views (PSNR 14.50 dB), whereas dense geometric constraints (DynOmo's local rigidity, isometry, and foundation feature grouping) achieve photorealistic novel view synthesis (PSNR 25.29 dB). Crucially, general-purpose vision foundation models (RoMa, MASt3R, UniDepth-V2, MoGe-V2) consistently outperform domain-specific surgical models, demonstrating that large-scale multi-domain visual priors transfer remarkably well to surgical environments compared to narrow in-domain fine-tuning.
Key Experimental Results¶
Main Results¶
Under the cross-endoscope evaluation protocol (trained on Endoscope 2, evaluated on Endoscope 1), the performance of rigid and deformable novel view synthesis methods is detailed below, alongside execution speeds on NVIDIA RTX 6000 Ada GPUs.
| Task Category | Method & Supervision | PSNR (dB)↑ | SSIM↑ | Render (FPS)↑ | Train (FPS)↑ |
|---|---|---|---|---|---|
| Rigid NVS | NeRF (w/o depth prior) | 17.24 | 0.65 | 0.11 | 0.21 |
| Rigid NVS | NeRF + Depth Anything [70] | 16.29 | 0.55 | 0.11 | 0.19 |
| Rigid NVS | NeRF + FoundationStereo [66] | 15.93 | 0.55 | 0.11 | 0.18 |
| Rigid NVS | 3DGS (w/o depth prior) | 14.82 | 0.59 | 487.62 | 0.59 |
| Rigid NVS | 3DGS + Depth Anything [70] | 14.77 | 0.63 | 509.99 | 0.36 |
| Rigid NVS | 3DGS + FoundationStereo [66] | 15.33 | 0.63 | 497.43 | 0.31 |
| Deformable NVS | 4DGS (Yang et al.) [71] | 13.74 | 0.57 | 578.46 | 0.38 |
| Deformable NVS | D-3DGS [72] | 14.50 | 0.51 | 31.41 | 0.14 |
| Deformable NVS | 4DGS (Wu et al.) [68] | 15.65 | 0.33 | 38.99 | 0.76 |
| Deformable NVS | Endo-4DGS [24] + Depth Anything | 16.26 | 0.59 | 309.37 | 0.66 |
| Online Deformable NVS | Hayoz et al. [23] + RAFT | 15.66 | 0.40 | 224.15 | 0.15 |
| Online Deformable NVS | Hayoz et al. [23] + FoundationStereo | 15.92 | 0.40 | 229.39 | 0.14 |
| Online Deformable NVS | DynOmo [55] + Depth Anything | 24.16 | 0.87 | 65.18 | 0.07 |
| Online Deformable NVS | DynOmo [55] + FoundationStereo | 25.29 | 0.90 | 63.58 | 0.06 |
The table below reports monocular depth estimation evaluated against high-accuracy stereo reference depth (> \(6 \times 10^8\) valid pixels across 30 sequences).
| Method | Med. Scaled | Abs Rel↓ | RMSE↓ | RMSElog↓ | \(\delta_1 (< 1.25)\uparrow\) | \(\delta_2 (< 1.25^2)\uparrow\) |
|---|---|---|---|---|---|---|
| UniDepth-V2 [47] (Foundation Model) | Yes | 0.081 | 0.012 | 0.111 | 0.936 | 0.992 |
| MoGe-V2 [65] (Foundation Model) | Yes | 0.109 | 0.014 | 0.140 | 0.889 | 0.989 |
| EndoDAC [11] (Surgical Fine-tuned) | Yes | 0.119 | 0.017 | 0.164 | 0.854 | 0.968 |
| Endo3R [21] (Surgical Only) | Yes | 0.187 | 0.023 | 0.236 | 0.704 | 0.927 |
| UniDepth-V2 [47] (Metric Direct) | No | 11.573 | 1.098 | 2.465 | 0.000 | 0.000 |
| MoGe-V2 [65] (Metric Direct) | No | 14.556 | 1.329 | 2.683 | 0.000 | 0.000 |
Ablation Study¶
For camera pose estimation and feature matching across 30 calibrated sequences, the table below compares rotation error, temporal drift on 23 static-camera dynamic-tissue sequences, and inlier statistics.
| Method Family | Feature Matcher | Rot. Err. (°)↓ | Temporal Drift (°)↓ | Inlier Rate (%)↑ | Avg Inliers↑ |
|---|---|---|---|---|---|
| Dense Foundation Matcher | RoMa [15] | 1.79 ± 1.41 | 1.41 | 78.1 | 3905 |
| Dense Foundation Matcher | MASt3R [31] | 1.81 ± 1.38 | 1.38 | 70.3 | 1676 |
| Learned Sparse Matcher | GIM + LightGlue [36,56] | 1.87 ± 1.48 | 0.41 | 43.1 | 264 |
| Learned Sparse Matcher | ALIKE + LightGlue [36,75] | 1.91 ± 1.52 | 0.64 | 68.2 | 211 |
| Detector-free Matcher | LoFTR [59] | 1.91 ± 1.40 | 0.43 | 47.2 | 513 |
| Learned Sparse Matcher | DISK + LightGlue [36,63] | 1.93 ± 1.52 | 0.59 | 55.2 | 446 |
| Lightweight Accelerated | XFeat [48] | 2.13 ± 1.47 | 0.77 | 25.9 | 854 |
| Learned Sparse Detector | SuperPoint + LightGlue [13,36] | 2.15 ± 1.41 | 0.70 | 34.2 | 74 |
| Surgical Domain Matcher | EndoMatcher [69] | 3.40 ± 1.91 | 2.11 | 39.9 | 1255 |
| Classical Handcrafted | SIFT [38] | 4.04 ± 14.45 | 5.44 | 52.9 | 91 |
Furthermore, trajectory accuracy validation across session types (Table 3) shows reprojection errors of 38.33 px (ex vivo), 17.00 px (in vivo postmortem), and 11.87 px (in vivo live), corresponding to physical accuracy errors at a 5cm working distance of 1.88 mm, 0.83 mm, and 0.58 mm, respectively.
Key Findings¶
- Catastrophic Collapse of Photometry-Only Deformable NVS: Models lacking spatial geometric regularization (e.g., D-3DGS) suffer severe visual degradation on held-out viewpoints (14.50 dB PSNR, 0.51 SSIM), proving that unconstrained temporal deformations overfit view-dependent radiance rather than reconstructing physical surfaces.
- The Heavy Computational Toll of Explicit 3D Regularization: While DynOmo achieves state-of-the-art NVS rendering (25.29 dB PSNR) by enforcing dense local rigidity, isometry, and foundation feature grouping, it requires 16 seconds per frame to optimize (0.06 FPS). This creates a 16× latency bottleneck compared to Endo-4DGS (0.66 FPS), highlighting real-time geometric regularization as a major open challenge.
- Surprising Dominance of Foundation Models over Domain-Specific Fine-Tuning: Pre-trained foundation models (UniDepth-V2, MoGe-V2, RoMa, MASt3R) consistently outperform surgical-specific architectures (Endo3R, EndoDAC, EndoMatcher) in relative depth accuracy and correspondence density. However, metric depth models fail without median scaling (Abs Rel > 11), emphasizing that scale ambiguity remains critical in surgical spaces.
- Accuracy versus Temporal Stability Trade-off in Feature Matching: Dense foundation matchers offer superior absolute accuracy and thousands of inliers, whereas video-trained sparse matchers like GIM+LightGlue achieve the lowest temporal drift (0.41°), proving highly advantageous for smooth online visual odometry and SLAM.
Highlights & Insights¶
- First True Geometric Generalization Benchmark for Surgical Vision: By collecting multi-endoscope synchronized video, iMED breaks the decade-long reliance on single-trajectory interpolation benchmarks, forcing algorithms to demonstrate physical 3D fidelity across independent viewpoints.
- Ingenious Low-Cost Miniature Wooden ArUco Cubes: Laser-etched 8mm wooden cubes solve specular reflection, trocar dimension constraints (≤ 12mm), and tissue adhesion challenges, enabling sub-millimeter trajectory calibration in wet operating fields.
- Challenging the Domain-Specific Fine-Tuning Dogma: The findings demonstrate that aggressive fine-tuning on limited surgical data can discard valuable geometric invariances learned by large vision foundation models, offering a fresh direction for medical visual foundation models.
Limitations & Future Work¶
- Hand-Guided Motion in Ex Vivo Sessions: Ex vivo sequences utilized hand-guided endoscope manipulation, which introduces slight mechanical irregularities compared to robotic wrist motion.
- Dependence on In-Scene Calibration Cubes for Dynamic Tracking: Dynamic sequences rely on visible fiducial cubes requiring mask-based inpainting removal, as markerless multi-endoscope tracking in non-rigid cavities remains an open research frontier.
- Limited In Vivo Subject Cohort: Although spanning 14 specimens and ~170K frames, the live in vivo subset originates from a single porcine subject; broadening the benchmark across diverse human clinical procedures will be an essential next step.
Related Work & Insights¶
- vs SCARED / SERV-CT / StereoMIS: Previous surgical benchmarks either provided static structured light/CT ground truth on phantoms/cadavers or relied on single-endoscope trajectories; iMED provides the first synchronized dual-endoscope, four-view benchmark with dynamic live tissue deformations.
- vs Panoptic Studio / Immersive Video: Traditional multi-view benchmarks deploy dozens of external cameras around isolated subjects; iMED solves the inverse challenge of placing synchronized multi-camera systems inside confined, fully deforming anatomical cavities.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ Pioneering multi-endoscope surgical benchmark that establishes a new paradigm for geometric generalization evaluation.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluation of 23 SOTA algorithms across 4 tasks on 14 specimens and ~170K frames.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear exposition, thorough calibration derivations, and deep insights into photometric vs. geometric trade-offs.
- Value: ⭐⭐⭐⭐⭐ Foundational infrastructure for future robot-assisted minimally invasive surgery, surgical digital twins, and embodied surgical AI.