๐ต Audio & Speech¶
๐ง NeurIPS2026 ยท 5 paper notes
๐ Same area in other venues: ๐๏ธ ECCV2026 (23) ยท ๐ท CVPR2026 (22) ยท ๐ฌ ICLR2026 (80) ยท ๐ฌ ACL2026 (70) ยท ๐งช ICML2026 (36) ยท ๐ค AAAI2026 (30)
- Audible World Models: Spatially Aware Sound Generation for 3D Worlds
-
Audible World Models connects panorama generation, semantic source parsing, dry-audio synthesis, and geometric acoustics to store sound as persistent 3D world state; it produces more accurate motion-dependent directional cues while retaining strong semantic alignment, but currently remains an offline construction system for static worlds.
- OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
-
OpenWhistle organizes years of underwater recordings from one dolphin group into a pretraining corpus containing approximately 180,000 whistles and an expert-annotated subset of 8,354 whistles, establishing session-split classification and detection benchmarks with frozen-representation linear probes; corpus-pretrained Wav2Vec2.0 achieves 81.1% classification accuracy and 75.6% detection mAP, without decoding the meaning of dolphin communication.
- Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound
-
Passing resamples a monorail-window recording into a nonlinear journey with branching playback, lets viewer presence indirectly change the visual path, and uses an installation-adapted SpecMaskFoley to generate sound in real time, presenting an exhibition case study and creative observations rather than a controlled experiment in acoustic reconstruction accuracy.
- Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
-
The paper models the vocal tract as a differentiable dynamic acoustic tube and combines frequency-domain solving, smooth turbulence gating, and neural geometry parameterization to infer area functions and MRI visualizations from speech; its default forward synthesis runs at 71.5 times real time, but the reconstructed geometry is an acoustically compatible explanation, not unique anatomical ground truth.
- SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding
-
SENSE combines electrode-geometry graph encoding with EEG semantic conditioning learned only from congruent trials to improve waveform reconstruction on the passive-listening N400 dataset; unseen-subject WER falls from FE-Phoneme's 1.1231 to 0.9948, but remains far from reliable language recovery.