Skip to content

๐ŸŽต Audio & Speech

๐Ÿง  NeurIPS2026 ยท 5 paper notes

๐Ÿ“Œ Same area in other venues: ๐ŸŽž๏ธ ECCV2026 (23) ยท ๐Ÿ“ท CVPR2026 (22) ยท ๐Ÿ”ฌ ICLR2026 (80) ยท ๐Ÿ’ฌ ACL2026 (70) ยท ๐Ÿงช ICML2026 (36) ยท ๐Ÿค– AAAI2026 (30)

Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Audible World Models connects panorama generation, semantic source parsing, dry-audio synthesis, and geometric acoustics to store sound as persistent 3D world state; it produces more accurate motion-dependent directional cues while retaining strong semantic alignment, but currently remains an offline construction system for static worlds.

OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations

OpenWhistle organizes years of underwater recordings from one dolphin group into a pretraining corpus containing approximately 180,000 whistles and an expert-annotated subset of 8,354 whistles, establishing session-split classification and detection benchmarks with frozen-representation linear probes; corpus-pretrained Wav2Vec2.0 achieves 81.1% classification accuracy and 75.6% detection mAP, without decoding the meaning of dolphin communication.

Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound

Passing resamples a monorail-window recording into a nonlinear journey with branching playback, lets viewer presence indirectly change the visual path, and uses an installation-adapted SpecMaskFoley to generate sound in real time, presenting an exhibition case study and creative observations rather than a controlled experiment in acoustic reconstruction accuracy.

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

The paper models the vocal tract as a differentiable dynamic acoustic tube and combines frequency-domain solving, smooth turbulence gating, and neural geometry parameterization to infer area functions and MRI visualizations from speech; its default forward synthesis runs at 71.5 times real time, but the reconstructed geometry is an acoustically compatible explanation, not unique anatomical ground truth.

SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding

SENSE combines electrode-geometry graph encoding with EEG semantic conditioning learned only from congruent trials to improve waveform reconstruction on the passive-listening N400 dataset; unseen-subject WER falls from FE-Phoneme's 1.1231 to 0.9948, but remains far from reliable language recovery.