Skip to content

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Conference: NeurIPS 2026 Creative AI Track (not the main conference)
arXiv: 2609.27489
Paper: Artwork page
Area: Audio & Speech (interactive audiovisual installation)
Keywords: video-to-audio synthesis, spacetime reconstruction, real-time sound synthesis, viewer presence, distributed creative agency

TL;DR

Passing resamples a monorail-window recording into a nonlinear journey with branching playback, lets viewer presence indirectly change the visual path, and uses an installation-adapted SpecMaskFoley to generate sound in real time, presenting an exhibition case study and creative observations rather than a controlled experiment in acoustic reconstruction accuracy.

Background & Motivation

Video-to-audio synthesis typically uses images to generate sounds that match scene semantics and event timing. The synchronization capabilities of methods such as MMAudio and SpecMaskFoley allow the technology to move beyond offline soundtrack production into continuously running audiovisual installations. Passing, however, does not present ordinary video: the artist stacks monorail-window frames into a spatiotemporal volume and changes the orientation and trajectory of a slicing plane. Locations originally belonging to the same moment can be reorganized, while motion traces from different moments can appear together. A unique, physically correct original soundtrack therefore cannot be assigned to each reconstructed image.

The artwork does not attempt to repair this inconsistency. It asks how a machine might interpret the world it โ€œhearsโ€ when familiar space and linear time are disrupted. Audiovisual synchronization remains important because sound should respond to the current visual flow, but synchronization is not faithful reconstruction. The creative setup also imposes concrete engineering requirements: finite material must support sustained playback, audience participation should not become direct volume or pitch control, and the generator must keep up with visual changes driving a high-frame-rate presentation.

Passing consequently connects the artist's rules, viewer presence, and the model's sonic interpretation rather than treating the model as an isolated soundtrack tool. Core Idea: use viewer presence to select paths through pre-rendered spacetime slices, then let a real-time video-to-audio model anchored in a train-interior acoustic context interpret the current visual flow, distributing creative agency among artist, machine, and audience.

Method

Overall Architecture

The inputs are a continuous monorail-window recording, ambient sounds recorded inside the train, and a camera-based estimate of viewer presence. The visual pipeline performs spacetime slice reconstruction, followed by presence-mediated sequence reordering; displayโ€“conditioning separation feeds the screen and audio model independently, and audio-semantics-guided synchronized synthesis produces stereo sound.

Solid arrows represent exhibition-time control or data flow, not training supervision. Viewer presence enters only the video controller, while train-interior sounds enter the audio-conditioning pathway through CLAP. Model training and adaptation are described separately below; viewer behavior is not used as a learning label.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Continuous monorail footage"] --> A["Spacetime slice reconstruction"]
    A --> B["Presence-mediated<br/>sequence reordering"]
    P["Camera-based presence estimate"] --> B
    B --> C["Displayโ€“conditioning separation"]
    C -->|4K HDR / 120 fps| V["Screen"]
    C -->|224 ร— 224 / 25 fps| D["Audio-semantics-guided<br/>synchronized synthesis"]
    S["Train-interior sounds<br/>CLAP conditioning"] --> D
    D --> O["32 kHz stereo audio"]

Key Designs

1. Spacetime slice reconstruction: extract different temporal organizations from finite footage

The artwork uses a Tama Monorail window recording captured on March 27, 2025, along the approximately 20 min route from Tachikawa to Tama Center. The source footage is 4K, 480 fps, and HDR. Ordinary playback advances frame by frame along time; Passing instead treats each frame as a slice and stacks the slices along the temporal axis into a virtual spatiotemporal volume. Changing the slicing plane's position and angle extracts a new two-dimensional image from the same footage, while continuously translating or rotating the plane creates a new moving sequence.

Unlike a generative model inventing scenery, this process keeps the content of the recorded scene and changes how space and time are sampled. The paper uses a 90ยฐ rotation to illustrate exchanging the temporal axis with the horizontal image dimension; this is a schematic example, not evidence that all displayed clips use one angle. As the plane tilts, train movement, relative motion of near and distant objects, and local vehicle or pedestrian activity can appear suspended, stretched, reversed, or displaced.

To preserve the embodied impression of scenery passing a window, the artist constrains the overall landscape flow to remain right to left, combining horizontal mirroring and reverse playback where necessary. Trajectories with different departure timings and rotational speeds are pre-rendered into clips. The live system therefore selects among prepared spacetime trajectories rather than recomputing high-resolution slices. The endless journey reorganizes finite material; it does not generate an unlimited supply of new landscapes.

2. Presence-mediated sequence reordering: let viewers change the world presented to the model

Video clips are organized into Groups A, B, C, and D. A contains standard forward sequences; B combines reverse playback with leftโ€“right mirroring. The default path plays A1 through A131, then B131 through B1, before returning to A1. Forward and reversed temporal organization alternate, while mirroring helps maintain a consistent horizontal landscape flow. A and B each contain 131 clips of approximately 10 s.

When a viewer is present, the controller probabilistically enables C or D branches; detection does not guarantee an immediate transition. C contains 58 clips connecting forward material to reverse-playback and mirrored material, while D contains 52 clips for the inverse transition. After C, playback automatically returns to the B-series; after D, it returns to the A-series. If no branch is selected, or no viewer is detected, the default path continues. Recently played C/D clips are temporarily excluded from the candidate pool for approximately 60 min to reduce immediate repetition.

The camera sends only a presence state to the controller and does not record or store camera imagery. When a transition is triggered, the controller sends an Open Sound Control (OSC) cue to the playback engine. Interaction thus affects the path rather than sound parameters: viewers change visual conditions such as motion speed, event density, and transition rhythm, and the model changes its sound in response. Viewers do not receive an interface for arbitrarily choosing clips or directly manipulating timbre.

3. Displayโ€“conditioning separation: high-quality presentation does not require processing every frame as model input

The screen plays pre-rendered 4K HDR clips at 120 fps, whereas the audio model receives corresponding 224 ร— 224 conditioning images at 25 fps. These image sequences are converted and cached before the exhibition, then loaded according to the selected clip. The pipeline has three distinct sampling specifications: 480 fps for capture, 120 fps for screen playback, and 25 fps for visual conditioning. They cannot be substituted for one another or interpreted as a single inference speed.

Separating presentation from conditioning avoids runtime processing of every displayed 4K, 120 fps frame. High-frame-rate video serves the viewing experience, while the lower-resolution, lower-frame-rate pathway supplies information for sound synchronization. This is a caching and input-bandwidth design, not evidence of frame-by-frame inference over the full display stream. The paper explicitly states that conditioning images are cached, but does not provide enough detail to claim that all visual features are extracted offline.

The model runs on an RTX 4080 workstation, generates 32 kHz stereo audio in real time, and routes it through an audio interface to loudspeakers. Video and conditioning inputs are prepared in advance, not the final soundtrack; sound is still generated continuously for the selected visual path. The paper does not report end-to-end latency, a real-time factor, or long-duration failure rates. โ€œReal timeโ€ therefore denotes the authors' reported installation capability rather than a quantified performance metric.

4. Audio-semantics-guided synchronized synthesis: anchor interpretations of unusual motion in a train-interior soundscape

SpecMaskFoley uses a pre-trained SpecMaskGIT audio backbone and introduces SynchFormer-based temporal video features through ControlNet and an FT-Aligner. The FT-Aligner adapts those features to the time-frequency ControlNet pathway, which supports audiovisual temporal alignment for sustained motion and instantaneous events. Passing adopts this mechanism rather than introducing a new audio-generation backbone.

The installation removes the CLIP visual encoder from the original configuration and delegates semantic guidance to CLAP audio conditioning. A pre-trained CLAP encoder processes ambient sounds recorded inside the train, providing SpecMaskGIT with a stable train-interior acoustic context. The visual pathway supplies the temporal structure of current motion, while the audio pathway constrains the broad semantic space of the sound. This permits responses to distorted visual motion without unrestricted drift toward unrelated or excessively artificial textures; it is not sequential replay of the original soundtrack.

Removing CLIP has both artistic and engineering motivations. Artistically, visual semantics should not explicitly determine the sonic interpretation. Technically, the authors identify CLIP as a major real-time inference bottleneck. The final configuration uses 32 kHz SpecVQGAN, SpecMaskGIT, and Vocos to produce playable audio. The paper provides component specifications and removal rationale, but no before-and-after latency or quality measurements, so this choice should not be presented as a quantified ablation result.

A Worked Example

Suppose the system is showing forward window scenery along the A-series when a viewer enters the viewing zone. The camera sends a presence state. If no branch is selected, the A-series continues and the audio model still receives the current clip's conditioning images. Viewer arrival does not directly change volume or immediately select an arbitrary sound.

If the controller selects a C clip, the playback engine shows its spacetime-distorted imagery and loads its 25 fps conditioning sequence. The audio model retains the CLAP semantic condition supplied by train-interior ambient sound, but synthesizes a new soundscape in response to the altered motion structure. After C, playback returns to the B-series, and the transition clip is temporarily excluded from the candidate pool for approximately 60 min. This example explains the system rules; it is not a user trial recorded in the paper.

Loss & Training

SpecVQGAN, SpecMaskGIT, and Vocos are trained on AudioSet and VGGSound. Source separation removes human speech components during preprocessing to reduce the risk of unintentionally generating human-like voices. This is a risk-mitigation measure, not a guarantee against voice-like outputs.

SpecMaskFoley fine-tunes pre-trained SpecMaskGIT using VGGSound plus approximately 35 min of custom Tama Monorail audiovisual data. To expand the limited train material, augmentation includes spatial video crops, audio-channel manipulations, and synchronized temporal reversal of audiovisual pairs. Preserving pair synchronization during reversal is relevant to the installation's nonlinear visual material, but the paper does not separately measure this augmentation's benefit.

Appendix Table 1 lists V2A adaptation data as VGGSound 500 h plus the monorail recordings. Approximately 20 min describes the window-recording route, while approximately 35 min describes the custom adaptation data; these are different quantities and should not be rewritten as one duration. The paper does not provide a new loss function, learning rate, or complete optimization configuration, and those unreported details are not reconstructed here.

Key Experimental Results

Main Results

This is an artwork paper. Its evidence consists of deployment specifications, a comparison between the original model and the installation configuration, and on-site stability tests and artist listening observations. The table summarizes representative changes from Appendix Table 1 without treating configuration differences as quality improvements.

Item Original SpecMaskFoley Passing configuration Supported conclusion
V2A adaptation data VGGSound 500 h VGGSound 500 h + custom monorail data Installation-specific material added; the main text reports approximately 35 min
Audio sampling rate 22.05 kHz 32 kHz Output specification changes, not a subjective quality score
SpecVQGAN parameters 72M 72M Tokenizer parameter count is unchanged
Mel representation 80 ร— 848 128 ร— 800 Mel-bin and temporal-frame configurations differ
Token grid 265 (5 ร— 53) 400 (8 ร— 50) Discrete audio representation length differs
Compression ratio 820ร— 800ร— Source values retained; no speedup inferred
SpecMaskFoley parameters 300M 300M The table reports the same total count
Main network / ControlNet 24 / 12 Transformer blocks 24 / 12 Transformer blocks Both columns report 768 dimensions and 8 attention heads
Vocoder HiFi-GAN Vocos Waveform-generation component changes

Appendix Figure 5 cites the original SpecMaskFoley evaluation of FAD and DeSync on the VGGSound test set; lower is better for both. It contextualizes model selection rather than independently evaluating Passing's final soundscape. The text cache does not supply reliably transcribable plot values, so no numerical ranking or improvement is reported here.

Ablation Study

The paper contains no controlled ablation study. The following deployment and observation analysis distinguishes specifications, interaction rules, and non-quantitative on-site observations.

Aspect Reported configuration or observation Evidence boundary
Source recording 4K, 480 fps, HDR; approximately 20 min route Source specification, not model throughput
Display 4K HDR, 120 fps Pre-rendered video playback specification
Visual conditioning 224 ร— 224, 25 fps; cached beforehand Not every displayed frame enters the model
Runtime and output RTX 4080; 32 kHz stereo No latency, real-time factor, or failure statistics reported
Default A/B path 131 clips each, approximately 10 s per clip Alternating playback preserves right-to-left landscape flow
C/D branches C: 58 clips; D: 52 clips; recent clips excluded for approximately 60 min Probabilistically enabled by presence, not deterministic user commands
On-site A/B observation Relatively stable train-interior atmosphere Artist listening and author descriptions, without listener statistics
On-site C/D observation Distorted visuals accompany stronger sonic variation Not a quantified causal effect or quality improvement

Key Findings

  • The installation combines pre-rendered visuals with real-time sound, rather than generating every modality online. Cached conditioning inputs also do not imply playback of prerecorded audio.
  • The A/B versus C/D listening difference informs creative decisions, but no participant count, rating distribution, or significance test supports a general improvement in immersion.
  • CLIP removal, Vocos replacement, and scene adaptation jointly define the final setup. Their contributions are not isolated, so the component producing the largest benefit cannot be identified.

Highlights & Insights

  • Separate synchronization from truth: the visual world has no unique correct soundtrack, but sound can still respond to its motion timing. Video-to-audio synthesis becomes an interpretive mechanism while retaining perceptible audiovisual relations.
  • Give indirect interaction a concrete causal path: presence changes clips, clips change model conditioning, and conditioning changes generated sound. This prevents audience participation from being conflated with direct control and makes the boundaries of creative authority explicit.
  • Source semantic guidance from audio: train-interior sounds maintain the listening context while visual features supply temporal variation. This division serves installations with a stable acoustic setting and continuously transformed imagery better than simply following visual categories.

Limitations & Future Work

  • The authors treat mismatches, ambiguity, and unexpected audiovisual correspondences as creative material, but this does not demonstrate that every mismatch has artistic value. The interpretation mainly comes from the authors and artist.
  • Branch probabilities, presence-detection accuracy, complete training hyperparameters, and end-to-end latency are unreported, limiting engineering reproducibility. Future reporting could include latency distributions across paths, audio interruption rates, and long-duration runtime logs.
  • One route and acoustic setting do not establish cross-scene generalization. Audience studies would need explicit participant counts, comparison conditions, and evaluation questions, separating artistic interpretation from generation quality.
  • Speech separation and the absence of stored camera imagery show risk awareness, but no audit of residual voice-like generation or presence-detection errors is provided. Both would benefit from targeted checks.
  • vs SpecMaskFoley: the earlier work introduces synchronized generation using SpecMaskGIT, ControlNet, and FT-Aligner. Passing contributes installation adaptation, visual-path organization, and a case study of creative agency; its backbone architecture should not be credited as a new method here.
  • vs MMAudio: MMAudio is cited as a high-quality synchronized video-to-audio method. Passing selects SpecMaskFoley to support continuous on-site operation, not to demonstrate superiority over MMAudio. No matched installation comparison is conducted.
  • vs Khronos Projector / Liquid Time: these works incorporate viewers' present actions into recorded video's temporal structure. Passing uses pre-rendered trajectories and presence-enabled branches for high-resolution presentation, then adds real-time generated sound; interaction is not arbitrary per-pixel spacetime manipulation.
  • vs Studies for: the earlier project explores collaboration through an artist's sound archive and text/audio conditioning. Passing adds dynamically reordered visual conditioning and viewer presence, mediating sonic interpretation through a variable visual path.

Rating

These ratings assess an artwork case study, not benchmark standing for a main-conference algorithm paper.

  • Novelty: 4/5 โ€” Nonlinear spacetime imagery, indirect audience interaction, and real-time sonic interpretation form a clear creative position.
  • Experimental Thoroughness: 2/5 โ€” Specifications and on-site observations are documented, but controlled ablations, system latency statistics, and formal audience studies are absent.
  • Writing Quality: 4/5 โ€” Actor roles and appendix rules are clear, with an explicit boundary between model-performance context and artwork evaluation.
  • Value: 4/5 โ€” A reusable design for continuous audiovisual installations, valuable primarily for creative practice and system integration rather than new algorithmic performance.