PISCO: Precise Video Instance Insertion with Sparse Control¶
Conference: NeurIPS2026
arXiv: 2602.08277
Paper: Project page
Area: Video Generation
Keywords: video instance insertion, sparse keyframe control, temporal masking, depth conditioning, video diffusion model
TL;DR¶
PISCO propagates a few instance keyframes into existing footage through variable-density conditioning, pre-encoding frame completion and post-encoding masking, and depth and appearance augmentation; on PISCO-Bench, first-and-last-frame control with its 14B model reduces whole-video FVD from VACE's 371 to 204, although the compared methods receive different inputs.
Background & Motivation¶
Video instance insertion is not about generating an approximately matching video from scratch. It adds a specified object to existing footage while preserving camera motion, background actors, and scene dynamics. Users also want precise placement and timing, natural subsequent motion, and environmental effects such as shadows and reflections. Conventional video inpainting uses per-frame spatial masks to define the editable region, but an object that has not yet been inserted has no existing trajectory to segment. Asking users to draw its entire mask sequence turns generation into labor-intensive manual animation.
Another approach edits an image first and uses an image-to-video model to generate subsequent frames. It can preserve the first-frame appearance but tends to reimagine the original background dynamics. Reference-guided video editing can use the original video, yet typically cannot constrain position, pose, and identity through a few instance cutouts at arbitrary timestamps. PISCO addresses an asymmetry in its conditions: the background video is fully observed, whereas the inserted instance is observed only at a few times. Missing instance observations neither mean that the object should disappear nor justify treating fabricated dense conditions as genuine observations.
The difficulty also begins in the pretrained video encoder. Zeroing all non-keyframes gives a temporal VAE a sequence that differs substantially from natural video. Fine-tuning the generator alone does not explain why the conditions themselves cause flickering and miscoloring. Core idea: separate providing the encoder with a more distribution-compatible continuous input from telling the generator which information is actually available, then learn propagation across conditioning densities and use depth and appearance augmentation to handle occlusion and illumination.
Method¶
Overall Architecture¶
The inputs are a complete background video, a few spatially placed instance RGB cutouts and masks, and background and instance depth conditions; the output is the entire video with the instance inserted. Training additionally uses a target video containing the instance as reconstruction supervision, not as an inference input. The backbone combines Wan and a VACE context adapter rather than first reconstructing a complete 4D scene.
Variable-Information Guidance first determines which instance frames are available during training; at inference time, these timestamps come directly from the user. Distribution-Preserving Temporal Masking completes the conditioning sequence in pixel space before suppressing unavailable information in latent space. Geometry- and Appearance-Robust Conditioning supplies relative depth and teaches occlusion of complete objects and illumination adaptation through training augmentation. Multi-Channel Adaptation and Staged Training then integrates these conditions into video denoising.
The spatial mask identifies where the object lies within a frame; the availability mask identifies which frames contain user information; latent-space masking identifies which encoded conditions can serve as valid hints. These have different roles: โno condition was supplied at this timeโ does not mean โthe object must be absent at this time.โ The background video and background depth are not dropped with instance conditions, so the model can still use the full shot to infer motion and occlusion.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
T["Training instance sequence"] -.-> A["Variable-Information<br/>Guidance"]
A -.->|Sample available timestamps| B["Distribution-Preserving<br/>Temporal Masking"]
I["Inference: user keyframes"] -->|Specify available timestamps| B
B --> C["Geometry- and Appearance-<br/>Robust Conditioning"]
V["Complete background and depth"] --> C
C --> D["Multi-Channel Adaptation<br/>and Staged Training"]
Y["Training target video"] -.->|Denoising supervision| D
D --> O["Video with inserted instance"]
Solid arrows denote inference conditioning flow; dashed arrows denote training-time condition sampling and target supervision. Amodal completion and relighting are training augmentations, not a mandatory post-processing chain at every inference run.
Key Designs¶
1. Variable-Information Guidance: adapt one model to control ranging from a single frame to dense observations
Variable-Information Guidance (VIG) is dynamic contextual dropout, not an additional classifier or a paper-specified inference guidance-scale formula. During training, a binary availability mask samples combinations spanning a single retained frame, varying sparsity, and complete conditioning. The model repeatedly encounters the task of generating a complete target from partial instance information. Sparse training encourages inference of intermediate motion from the background and video prior; dense training constrains identity and detail so that propagation does not come at the expense of appearance fidelity.
With \(A\) denoting frame availability, \(I\) instance RGB, \(M\) the spatial mask, and \(D_I\) instance depth, the condition-masking boundary is:
Background depth \(D_V\) is not masked by \(A\), and the complete background video remains an input. Availability describes the reliability of the condition source, not an output presence constraint. A first-frame input cannot specify every subsequent trajectory precisely or guarantee a unique motion at every unobserved timestamp. The paper describes hybrid-density training but does not provide branch probabilities or a precise density distribution; a reproducible sampling recipe should not be invented from this description.
2. Distribution-Preserving Temporal Masking: completion serves encoding, while masking preserves the information boundary
Distribution-Preserving Temporal Masking (DPTM) addresses missing conditioning frames, not missing output frames: extreme conditioning sparsity disrupts the input distribution of the pretrained temporal VAE. It first fills missing frames with the temporally nearest available instance frame, propagating keyframe cutouts forward and backward. The encoder therefore no longer receives abrupt alternation between normal images and all-zero frames. This is neither optical-flow estimation nor generation of intermediate poses; the completed content is a placeholder that supports encoding.
After VAE encoding, condition tokens associated with originally unobserved timestamps are masked so that the generator does not treat repeated cutouts as a user-specified trajectory. Temporal compression places several original frames within one compressed token, so a single availability bit per token would lose within-group information. The authors move the frame-wise availability pattern of a local temporal window into the channel dimension and provide it to the adapter alongside other conditions.
The paper specifies an availability tensor of shape \(C_t\times T'\times H'\times W'\), where \(C_t\) is the temporal compression factor and \(T'=\lceil T/C_t\rceil\). This preserves the pattern of observed original frames within each compressed position. However, the cache does not specify the complete token-masking implementation for groups containing both observed and unobserved frames. It should not be described as proven frame-wise information isolation, nor as eliminating all cross-frame mixing in the temporal VAE.
3. Geometry- and Appearance-Robust Conditioning: learn to occlude objects while allowing scene-compatible illumination changes
Depth conditions come from two sources: Depth Anything V3 estimates depth for the complete background video, and instance depth is extracted by masking the depth of the target video containing the object. Background depth remains fully available; instance depth is available only at retained keyframes. These conditions teach relative ordering rather than pasting a complete object over every foreground occluder. This is learned conditional generation, not a geometric renderer with calibrated cameras. The paper also does not fully explain how depth for practical user-supplied instances is placed on a common scale with the target scene.
Training instances are usually visible regions obtained through segmentation, whereas users typically provide complete objects. Training exclusively on fragmented cutouts risks treating missing regions as intrinsic appearance rather than scene-induced occlusion. The authors build examples from fully visible instances, randomly occlude them with other cutouts, and fine-tune an image editing model with LoRA to recover complete appearance. Frame-wise completion produces pseudo-amodal instances; masks are obtained by thresholding, and missing depth is interpolated from available values. Pairing these complete conditions with the original occluded target teaches the relationship โcomplete input, scene-occluded output.โ
Relighting augmentation uses IC-Light to change instance conditions under random background lighting, while the target still requires compatibility with the original scene. This discourages mechanical copying of input colors and illumination and permits adaptation to environmental light. It does not justify arbitrary identity changes: the aim is to correct incompatible illumination while preserving identity and keyframe control. Shadows and reflections can occur outside the spatial instance mask, so the method is not a copy-and-paste system that hard-locks every background pixel.
4. Multi-Channel Adaptation and Staged Training: establish the condition interface before modifying the video prior
The VACE context adapter receives background, instance RGB, depth, spatial masks, availability, and related signals. Its first linear projection is modified to accept the new input dimensionality, while the other layers initially reuse pretrained weights. The principal adaptation therefore concerns how task conditions enter the backbone rather than training a video generator from scratch. Condition tokens influence denoising through the adapter, and the model generates the entire video with the instance and necessary environmental effects.
Training first updates only the new input projection, then the complete adapter, and then both the adapter and diffusion backbone. Completion and relighting augmentation are subsequently introduced with continued joint adaptation using LoRA, followed by a spatial and temporal extension. This sequence lets the randomly initialized condition interface learn useful representations before the expensive generator is unfrozen, reducing the risk of immediately disrupting pretrained video priors. PISCO-14B uses Wan2.2's separate high-noise and low-noise denoisers, each trained with the same schedule; these are not an additional instance-propagation network.
A Worked Example¶
For a 49-frame clip under the paper's first-and-last-frame protocol, the user supplies spatially placed cutouts, masks, and corresponding depth for the same object at frames 1 and 49; all 49 background frames remain available. The endpoint keyframes anchor identity, position, and pose. The intermediate 47 frames lack instance observations, rather than prohibiting the object from appearing.
DPTM fills intermediate conditioning frames with the nearest endpoint cutout for temporal VAE encoding, then masks conditioning information for unsupplied timestamps while retaining the encoded availability of the two genuine endpoints. The model generates intermediate object states using complete background motion and depth. A pillar in the background should occlude the object, and changing environmental light can alter its illumination. Endpoint conditions reduce path ambiguity but do not specify an exact trajectory for every intermediate frame.
Loss & Training¶
Supervision comes from the target video containing the instance and a denoising prediction error. Keyframe sparsification affects the input conditions, while the target covers the entire sequence. Cached equation (2) displays the error's second norm without a square; this note does not silently reconstruct it as the authors' exact mean-squared-error formula or assume additional occlusion, background-copying, or physical-constraint losses.
Training videos come from ROSE, VPData, MOSE, and DAVIS. Filtering for removable instances yields 16,642 clips with at least 49 frames each. A side-effect-aware removal model trained on ROSE produces instance-free backgrounds for target/background pairs. This differs from simply blacking out the instance region and also means that the training pairs inherit removal-model errors.
The backbones are Wan2.1-VACE-1.3B and Wan2.2-VACE-Fun-A14B. The first four stages use \(832\times480\) and 49 frames; the fifth extends to \(1280\times720\) and 120 frames. Optimization uses AdamW on NVIDIA H100 GPUs. The cache does not provide stage-specific step counts, learning rates, batch sizes, or sampling probabilities; a five-stage schedule is not a complete reproduction specification.
Key Experimental Results¶
Main Results¶
PISCO-Bench selects 100 BURST videos without training overlap, manually checks and corrects instance masks, and uses ROSE to generate clean backgrounds. All comparisons use \(832\times480\), 49 frames, and 50 denoising steps. These results should not be treated as direct validation of the extended 720p, 120-frame configuration.
Whole-video metrics compare generated footage with the original video containing the instance. Foreground metrics multiply both generated and target videos by the target instance mask before evaluation. Foreground LPIPS is therefore evaluated on masked videos, not independent object crops, and does not assess all shadows and reflections outside the mask.
The following selection comes from the paper's Table 2; lower FVD and LPIPS and higher PSNR are better.
| Method and control | Whole-video FVD | Whole-video LPIPS | Whole-video PSNR | Foreground FVD | Foreground LPIPS |
|---|---|---|---|---|---|
| Image editing + I2V, first and last | 624 | 0.392 | 16.44 | 250 | 0.030 |
| CoCoCo, dense masks | 590 | 0.191 | 23.62 | 398 | 0.031 |
| VideoPainter, dense masks | 524 | 0.154 | 23.11 | 384 | 0.035 |
| VACE-14B, dense masks | 371 | 0.103 | 25.55 | 273 | 0.028 |
| UniVideo, no mask conditioning | 485 | 0.211 | 19.22 | 310 | 0.031 |
| PISCO-1.3B, first and last | 269 | 0.103 | 27.01 | 171 | 0.024 |
| PISCO-14B, first and last | 204 | 0.097 | 26.58 | 138 | 0.022 |
Inputs are not equivalent. The I2V pipeline edits endpoints with Nano-banana-Pro and generates with Wan2.2-Fun-A14B-InP, without using the full background video as a frame-wise constraint. CoCoCo and VideoPainter receive dense masks and descriptions generated by Qwen3-VL-32B-Instruct, but no direct reference instance image. VACE receives the background video, reference image, text, and full masks; UniVideo receives the background video, reference image, and text, without mask support. PISCO receives the complete background plus sparse instance images, masks, and depth. The comparison evaluates different workflows rather than an identical information budget.
Ablation Study¶
The cache contains no quantitative component-removal table for VIG, DPTM, or depth conditioning. Figures 3, 4, and 6 provide qualitative analyses of depth, DPTM, and relighting, respectively; they cannot justify invented numerical ablations. The verifiable quantitative analysis varies control density, as shown below from Tables 2 and 3.
| Model and control | Whole-video FVD | Foreground FVD | VBench subject consistency | VBench eight-metric average |
|---|---|---|---|---|
| PISCO-1.3B, first only | 398 | 243 | 87.16 | 64.43 |
| PISCO-1.3B, first and last | 269 | 171 | 91.33 | 65.45 |
| PISCO-1.3B, five frames | 172 | 104 | 91.45 | 65.89 |
| PISCO-14B, first only | 337 | 222 | 87.26 | 64.56 |
| PISCO-14B, first and last | 204 | 138 | 91.57 | 65.64 |
| PISCO-14B, five frames | 136 | 75 | 91.98 | 66.04 |
Five-frame control uses the first and last frames plus three random intermediate frames. It is an additional reference setting, excluded from standard-protocol rankings. VBench background and subject consistency isolate their respective regions with masks and use CLIP/DINO features; other metrics follow the official implementation. The average is therefore neither a user-control success rate nor a physical-correctness rate.
Key Findings¶
- First-and-last-frame PISCO-14B reduces whole-video FVD by 167 and foreground FVD by 135 relative to VACE. However, it receives instance appearance at two timestamps, whereas VACE receives dense placement masks; the improvement cannot be attributed to a single module.
- Moving from first-only to first-and-last to five-frame control consistently improves FVD and average VBench scores for both models, supporting the value of additional anchors in reducing propagation ambiguity. This is not a complete sensitivity curve over arbitrary control counts and placements.
- Larger models do not improve every metric: first-and-last whole-video PSNR is 27.01 for 1.3B versus 26.58 for 14B. First-only 14B whole-video LPIPS, 0.116, is also worse than VACE's 0.103. The paper's claim of consistently outperforming every baseline should be restricted to supported metrics and protocols.
- โMonotonic improvementโ does not hold metric by metric: 14B VBench Temporal Flickering changes from 97.26 with first-only control to 97.21 with first-and-last control, then to 97.34 with five frames. Aggregate gains should not be described as strict monotonic improvement in every metric.
Highlights & Insights¶
- Missing conditions can disrupt the encoder input distribution before becoming a generator problem. Completing before encoding and masking afterward separately address encoding stability and observational validity more directly than zero filling.
- Complete background and sparse instance observations are distinct information sources and should not share indiscriminate dropout. Retaining background geometry provides motion and occlusion references when instance observations are absent.
- Pairing a complete object condition with an occluded target better matches practical insertion than training exclusively on fragmented segmentation cutouts. The transferable insight is the visibility relationship between conditions and targets, not a requirement to copy the condition into every output frame.
Limitations & Future Work¶
- Single-frame conditions leave intermediate trajectories ambiguous, and the qualitative comparison acknowledges drift. Additional keyframes help, but explicit frame-wise placement errors and control-failure rates are not reported.
- Removal-model-generated backgrounds may retain object clues or introduce inpainting errors. A self-constructed benchmark of 100 videos does not fully represent arbitrary new-object insertion in open-world scenes.
- Depth-scale alignment, masking rules for mixed-availability temporal tokens, and training hyperparameters remain underspecified. Future evaluation should provide explicit implementations and sensitivity tests for depth errors and keyframe placement.
- Quantitative evaluation concentrates on 49-frame clips, without systematic long-video or 720p evaluation, latency and memory tables, or complete numerical component ablations.
- Background changes, repositioning, rescaling, speed changes, and dynamics simulation are application examples, not physical simulation capabilities validated by independent benchmarks. Plausible occlusions and shadows do not guarantee physical laws.
Related Work & Insights¶
- vs VACE / UniVideo: General reference editing supplies generation backbones and control interfaces, whereas PISCO specializes in instance conditions at arbitrary sparse timestamps. It improves instance-level control, but dense masks, text, and appearance anchors carry different information, so this is not a same-interface substitution.
- vs CoCoCo / VideoPainter: Inpainting relies more heavily on per-frame edit regions; PISCO infers propagation from a few instance cutouts and allows effects outside those regions. Background copying helps consistency but does not itself solve accurate identity or natural illumination adaptation.
- vs InsertAnywhere and other 4D methods: Explicit geometry prioritizes spatial structure, while PISCO combines depth conditioning with generative priors to avoid the reconstruction pipeline. Occlusion and motion consequently rely mainly on learned generation, without geometric or physical hard guarantees.
- Research direction: Confidence-aware availability could distinguish genuine user anchors, estimated depth, and pseudo-completed information, with trajectory error and background side effects evaluated separately. This is an extension, not a module implemented in this paper.
Rating¶
- Novelty: 4/5, a targeted combination addressing encoder distribution shift under sparse instance control.
- Experimental Thoroughness: 3/5, paired evaluation and control-density analysis, but no quantitative component ablations or fully matched input protocols.
- Writing Quality: 3/5, a clear task definition, with incomplete masking and hyperparameter details and some claims exceeding table evidence.
- Value: 4/5, useful for low-interaction video editing, while physical consistency and long-video deployment require further validation.