At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections¶
Conference: NeurIPS2026
arXiv: 2604.10766
Area: Computational Biology
Keywords: cryogenic electron tomography, open-set macromolecule detection, tilt-series, visual prompting, multi-view geometry
TL;DR¶
FullTilt feeds aligned 2D tilt-series directly into a visually prompted multiclass 3D detector, replacing volumetric sliding-window detection with cross-tilt row attention, near-zero-tilt query initialization, and training-time geometric augmentation to achieve subsecond zero-shot detection on three real-world cryo-ET datasets, while still requiring simulated-data pretraining and input alignment.
Background & Motivation¶
Cryogenic electron tomography (cryo-ET) records 2D projections while rotating a specimen, then reconstructs the aligned tilt-series into a 3D tomogram. Conventional workflows locate macromolecules in this volume before extracting subtomograms for averaging and structural refinement. Closed-set detectors require new annotations and training for a new protein, whereas open-set methods such as TomoTwin and ProPicker locate targets from visual reference particles. However, removing target-specific training does not remove volumetric traversal: an entire tomogram does not fit in GPU memory, so models repeatedly read 3D subvolumes or, as in CryoSAM, traverse 2D slices along multiple axes.
The bottleneck is therefore not merely a slow detection network, but an expensive representation chosen before detection. The paper illustrates this with a tilt-series containing 41 projections versus a corresponding tomogram with 256 depth slices; a reconstruction convenient for inspection need not imply that detection must traverse every reconstructed voxel. Reading projections directly can bypass volumetric detection, but loses the depth separation provided by reconstruction. Particles may overlap at particular angles, while the signal-to-noise ratio in a single image may be too low to recognize them, so independent 2D detection followed by back-projection converts unreliable image evidence into 3D localization errors.
The paper therefore asks whether fusing cross-tilt evidence before detection can preserve the geometry needed for localization without constructing and traversing dense 3D features. Alignment provides a useful constraint: rotation around the same axis preserves a particle's vertical image coordinate, while its horizontal position changes with tilt angle and depth. Core Idea: exploit this row correspondence to fuse the complete projection series, identify requested classes through visual prompts, initialize 3D queries from a near-zero tilt, and refine them across views to predict particle coordinates without tomogram sliding-window detection.
Method¶
Overall Architecture¶
The inputs are a 2D tilt-series already aligned along the y-axis, the angle of each image, and class-labeled 2D particle boxes; the outputs are particle centers in 3D, diameters, classes, and confidence scores. A Swin Transformer extracts multiscale features independently from each image, followed by “Tilt-Series Encoding,” “Multiclass Visual Prompt Encoding,” and “Tilt-Aware 3D Decoding.” The resulting 3D-aware features remain organized as 2D views rather than a reconstructed dense 3D volume.
“Auxiliary Geometric Primitives” provide additional supervision only during training and are not inference inputs. The model requires aligned images and known angles, so preprocessing between microscope acquisition and alignment is not replaced. The standard 3D prompting workflow still uses a reconstructed tomogram, whereas direct 2D prompting on the zero-tilt image can avoid that step.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Aligned projections<br/>and tilt angles"] --> B["Per-image 2D backbone"]
B --> C["Tilt-Series Encoding"]
C --> D["Multiclass Visual<br/>Prompt Encoding"]
P["Class-labeled 2D prompts"] --> D
D --> E["Tilt-Aware 3D Decoding"]
C -->|Cross-view features| E
E --> F["3D centers, diameters<br/>classes and confidence"]
G["Auxiliary Geometric<br/>Primitives"] -.->|Training only: synthetic projections| A
G -.->|Training only: 3D ground-truth supervision| E
Key Designs¶
1. Tilt-Series Encoding: turn low-SNR views into shared geometric evidence
Each projection's features receive 2D positional, feature-level, and tilt embeddings. The tilt embedding is produced by an MLP over multifrequency sine and cosine encodings, informing the network that visually similar images may come from different viewing directions. The encoder alternates local attention within a view with global row attention across all views, progressively updating cross-view features. The local step organizes spatial context within each image; the cross-view step connects candidate responses along the same row, allowing a particle obscured by noise to borrow evidence from other angles instead of requiring a reliable single-image detection first.
Row attention does not simply reduce the number of views: it restricts attention using acquisition and alignment geometry. Rotation changes a particle's horizontal projection coordinate but not its row, so cross-tilt search can extend horizontally along that row without interactions between every pair of 2D positions. This removes irrelevant computation while retaining horizontal trajectories that vary with depth. It does not yield an exact depth estimate: the encoder supplies fused evidence for 3D localization, while the subsequent query decoder determines the centers. If the input violates the alignment assumption, row correspondence itself may fail.
2. Multiclass Visual Prompt Encoding: aggregate only the reference particles of each class
Users provide visual examples rather than fixed training-class identifiers: each prompt contains a particle center, a box defined by its diameter, and a class label. Images may contain different numbers of prompts and multiple requested classes. To accommodate this irregular input, the model pads local prompts within each image, then appends a full-image global box and a learnable class token for every requested class. Local prompts also receive content tokens and encoded box positions. Boxes direct cross-attention to features near the target, while global class tokens collect evidence from same-class prompts.
Prompt tokens first read fused image features through deformable cross-attention, then undergo class-masked self-attention. The mask permits interaction only between valid tokens with identical labels, preventing a ribosome prompt from directly mixing with another protein class and excluding padded positions as valid evidence. After each layer, a class token is averaged only across views containing prompts for that class, producing a final class prototype. Prototypes support both query initialization and class prediction during decoding; their number follows the user's request rather than a fixed list of training proteins.
This explains why reducing prompt projections differs from reducing input projections. Even with a prompt on only one tilt, the encoder still observes the complete tilt-series, and features near the prompt already incorporate other views. Removing input images instead removes cross-view evidence, so the same robustness cannot be assumed. Prototypes can also be extracted from one instance and reused for other instances in the same dataset, although whether target appearance remains consistent across instances requires empirical evaluation.
3. Tilt-Aware 3D Decoding: start from a reliable plane and let all views correct depth
The initializer selects the image closest to zero tilt, computes similarity between each feature location and all class prototypes, and takes the maximum over classes to obtain candidate responses. The 900 highest-response locations and their associated 2D boxes provide initial centers and diameters. Depth starts at normalized coordinate 0.5, the middle plane, while content queries are learnable parameters. Near-zero tilt is chosen because the horizontal and vertical coordinates in the zero-degree projection directly correspond to 3D x and y, avoiding simultaneous exploration of image-plane position and depth by random 3D queries under low-SNR conditions.
This initialization does not measure the true z coordinate: it supplies a more reliable starting point for refinement and assumes zero-mean depth bias. The nearest-to-zero image need not be exactly at zero degrees, so the coordinate correspondence motivates the strategy rather than guaranteeing exact recovery of x and y at every angle. The final depth must be corrected through consistency across multiple projections instead of remaining fixed at the middle plane.
Each decoder layer applies self-attention among queries, cross-attention between queries and class prototypes, and deformable cross-attention between queries and image features. The third step projects the current 3D anchor into every tilt and samples nearby features. Rather than detecting every image independently before back-projection, one shared 3D hypothesis gathers corresponding evidence from all views. Under the pixel-coordinate convention in the paper, the central projection relation is:
Here W is image width, D is the depth of the adopted 3D coordinate space, and \(\theta\) is the tilt angle. Candidates sharing x and y but differing in z project to different horizontal locations at nonzero angles, allowing multi-view features to distinguish particles that overlap in one projection. Although a tomogram is not reconstructed, the 3D coordinate scale and depth range must still be defined: avoiding 3D volume inputs does not eliminate the need for 3D geometry.
Query updates from all views are averaged, and anchors are iteratively refined using DINO's “look forward twice” approach. The final anchor supplies the 3D center and a single diameter, while inner products between queries and class prototypes provide class predictions and confidence scores. This separates responsibilities: 2D encoding exposes particles within noisy data, 3D queries seek locations that explain multiple projections, and visual prototypes determine whether the particles match the requested classes.
4. Auxiliary Geometric Primitives: train cross-view localization with inexpensive exact geometry
Physics-based cryo-ET simulation alone is expensive, and low SNR, occlusion, and non-uniform illumination further complicate learning geometry. During training, the auxiliary module generates 2D projections of geometric primitives, such as differently sized circles, with exact 3D ground-truth coordinates. It varies object distributions, including dense clustering near the middle depth plane, and simulates physical occlusion and non-uniform illumination. The network therefore learns both where one 3D position should appear at different angles and how to exploit other views when some evidence is corrupted.
These data are generated on the fly rather than stored as large collections of files, and training alternates between them and simulated cryo-ET data each epoch. Simple circles do not replace all protein training, and the module is not a geometry-reconstruction stage required before real-data detection. Instead, it is an auxiliary training source with exact spatial labels, providing inexpensive, controlled supervision for cross-view correspondence while physics-based simulations of protein shapes support visual recognition. The text provides this high-level mechanism but not the full primitive distributions or augmentation parameters, so unspecified implementation details should not be presented as published settings.
A Worked Example¶
Consider detecting one particle class within a projection series. A user supplies one class-labeled box on the zero-tilt image. The backbone encodes every input image, row attention gathers evidence from multiple tilts along corresponding rows, and the prompt encoder compresses fused features around the box into a class prototype. Even though there is only one 2D prompt, this process uses the full image series rather than independent single-image detection.
The initializer uses the prototype to select up to 900 candidates from the near-zero-tilt image and places them at the middle depth initially. If one candidate has an incorrect depth, its anchor projects away from the particle response at nonzero tilts; the decoder uses multi-view samples to refine its position and diameter, then assigns a class and confidence through prototype similarity. This example illustrates the mechanism, not a numerical refinement trajectory reported for a specific particle, and it does not imply that all 900 candidates become valid final detections.
Loss & Training¶
Model outputs are paired with 3D ground truth through bipartite matching, after which an L1 loss supervises positions and sizes and a focal loss supervises classification. Auxiliary losses at decoder layers and a contrastive denoising strategy accelerate convergence; the text does not provide complete loss weights, so none are reconstructed here. Although detection classes are determined by visual prototypes, training still requires spatially labeled simulated particles and is not unsupervised.
Pretraining uses the same 119 proteins as prior open-set methods to generate 1,138 simulated tomograms and corresponding noisy and clean aligned tilt-series. Geometric-primitive training alternates with these data each epoch. Evaluation on real data does not retrain the model for the test targets, so “zero-shot” means no target-specific retraining, not an untrained model or detection without visual prompts.
The tilt-series encoder, prompt encoder, and decoder each have 6 layers, with hidden dimension 256 and 900 queries. Training uses AdamW with learning rate \(1\times10^{-4}\) for 200 epochs on 4 NVIDIA RTX 6000 Ada GPUs, a per-GPU batch size of 1, and PyTorch 2.5.1. Inference removes the geometric-data generation path but retains the trained detector.
Key Experimental Results¶
Main Results¶
The table selects intra-instance results with 1 prompt per class from Tables 1–3; means and variation terms are preserved as reported. [email protected] and mAP@1r use center-distance thresholds of 0.5 and 1 times the target radius, respectively, while F1 uses a 1-radius threshold; these are not 3D-box IoU metrics. Measurements use one NVIDIA RTX 6000 Ada GPU and a 48-core CPU, with prompt selection repeated for 10 trials. The text does not explicitly identify the variation term as a standard deviation or standard error.
| Dataset | Method | [email protected] | mAP@1r | F1 | Runtime (s) | Peak VRAM (MB) |
|---|---|---|---|---|---|---|
| CZII | TomoTwin | 0.089 ± 0.021 | 0.132 ± 0.024 | 0.078 ± 0.012 | 2960 | 20800 |
| CZII | FullTilt | 0.118 ± 0.013 | 0.148 ± 0.013 | 0.193 ± 0.008 | 0.499 | 3430 |
| EMPIAR-10304 | TomoTwin | 0.189 ± 0.057 | 0.339 ± 0.063 | 0.532 ± 0.061 | 1730 | 20800 |
| EMPIAR-10304 | FullTilt | 0.199 ± 0.004 | 0.437 ± 0.002 | 0.593 ± 0.002 | 0.229 | 2410 |
| EMPIAR-10499 | TomoTwin | 0.035 ± 0.008 | 0.113 ± 0.020 | 0.249 ± 0.027 | 2330 | 20800 |
| EMPIAR-10499 | FullTilt | 0.159 ± 0.004 | 0.288 ± 0.004 | 0.418 ± 0.005 | 0.389 | 3410 |
These results demonstrate acceleration of the detection stage, not end-to-end processing from microscope acquisition. FullTilt substantially reduces runtime and memory relative to volumetric open-set methods, but is not faster or more memory-efficient than every multi-view baseline. For example, on EMPIAR-10304, DETR3D uses 2290 MB and 0.136 seconds versus FullTilt's 2410 MB and 0.229 seconds, with FullTilt delivering better detection quality.
Ablation Study¶
The following CZII results from Table 6 distinguish removing input images from reducing the views containing a prompt, using a baseline of 1 prompt per class.
| Config | [email protected] | mAP@1r | Runtime (s) | Peak VRAM (MB) |
|---|---|---|---|---|
| Full input and prompt projections | 0.118 ± 0.013 | 0.148 ± 0.013 | 0.499 | 3430 |
| Retain 50% of input tilt images | 0.100 ± 0.013 | 0.133 ± 0.011 | 0.246 | 1860 |
| Retain 25% of input tilt images | 0.085 ± 0.010 | 0.120 ± 0.011 | 0.153 | 1280 |
| Retain 12.5% of input tilt images | 0.038 ± 0.007 | 0.073 ± 0.014 | 0.087 | 700 |
| Full input, prompt in only one image | 0.106 ± 0.014 | 0.134 ± 0.016 | 0.513 | 3430 |
Key Findings¶
- Prompts can be sparse, but measurements cannot be arbitrarily sparse: single-image prompting changes [email protected] from 0.118 to 0.106, whereas retaining only 12.5% of input images reduces it to 0.038. The full projection evidence is the critical resource.
- In Table 7, removing tilt-series encoding, initialization, and geometric augmentation together gives CZII [email protected] of only 0.001 ± 0.001, compared with 0.118 ± 0.013 for the full model. The text states that initialization first enables viable detection among single-module configurations, followed by joint improvements from fusion and geometric supervision; the gains cannot be attributed merely to replacing the backbone.
- Multiclass requests share computation: on CZII with 1 prompt per class, increasing the requested classes from 1 to 6 changes runtime from 0.499 to 0.508 seconds and VRAM from 3430 to 3440 MB, while [email protected] changes from 0.118 ± 0.013 to 0.113 ± 0.008.
- FullTilt does not lead in every setting: with 4 intra-instance prompts on EMPIAR-10304, TomoTwin achieves [email protected] of 0.241 ± 0.022 versus FullTilt's 0.201 ± 0.004, although FullTilt still has higher mAP@1r and F1. Table 4 shows similar exceptions for the strict distance metric with 2 and 4 cross-instance prompts on this dataset.
Highlights & Insights¶
- Optimizing the input representation is more fundamental than merely compressing a 3D detector. Predicting 3D coordinates does not require constructing a dense intermediate 3D feature volume.
- Row attention converts an explicit rotation and alignment constraint into computational sparsity. It retains horizontal particle trajectories while avoiding impossible correspondences in low-SNR projections.
- Visual prompts read cross-view fused features before producing class prototypes. A small number of examples can therefore exploit the whole series, while multiple requested classes share the main encoding cost.
Limitations & Future Work¶
- The authors acknowledge difficulty with particularly small macromolecules, and a single-diameter representation is unsuitable for highly non-globular particles. Center localization should be distinguished from recovering shape, orientation, or high-resolution structure.
- The geometry assumes aligned inputs, known tilt angles, and a defined coordinate scale. Preprocessing, reconstruction for the standard 3D prompting workflow, and downstream subtomogram averaging have not all disappeared; the real-time claim primarily concerns the measured detection inference.
- The 900 queries impose a candidate budget, and absolute real-data mAP values indicate that the task remains difficult. Stability under denser specimens, alignment errors, and different acquisition conditions warrants further evaluation rather than inferring universal reliability from three datasets.
- The cached text contains no additional appendix and leaves details such as complete geometric-primitive parameters and inclusion of all prompting interactions in timing insufficiently specified. It only promises public code and data upon publication, without a verifiable repository link.
Related Work & Insights¶
- vs TomoTwin / ProPicker: all use visual references for open-set localization, but prior methods primarily scan reconstructed volumes while FullTilt directly encodes projections. Its main efficiency gain is reduced volumetric traversal, not elimination of training.
- vs CryoSAM / Zeng et al.: CryoSAM segments tomogram slices, whereas Zeng's method detects individual projections before back-projection. FullTilt fuses the entire tilt-series before localization and lets shared 3D hypotheses gather evidence across views, reducing error propagation from unrecognizable single images.
- vs DETR3D / PETR: FullTilt inherits multi-view 3D queries and positional-encoding ideas but adds cryo-ET-specific cross-tilt fusion, visual class prototypes, and initialization. The results show that natural-image multi-view frameworks do not directly solve detection in extremely low-SNR projections.
Rating¶
- Novelty: 5/5; moves open-set cryo-ET detection to aligned projections and exploits row geometry for fusion.
- Experimental Thoroughness: 4/5; covers three real datasets, prompt counts, cross-instance transfer, multiclass detection, and component ablations, but needs further error and preprocessing-boundary analyses.
- Writing Quality: 4/5; component responsibilities are clear, although geometric-data details and some efficiency claims need tighter qualification.
- Value: 5/5; substantially reduces large-scale particle-localization costs while retaining interaction without target-specific retraining.