Reconstructing the Vocal Tract with Differentiable Acoustic Simulation¶
Conference: NeurIPS2026
arXiv: 2609.36737
Area: Audio & Speech
Keywords: differentiable acoustic simulation, articulatory inversion, frequency-domain synthesis, neural fields, vocal tract visualization
TL;DR¶
The paper models the vocal tract as a differentiable dynamic acoustic tube and combines frequency-domain solving, smooth turbulence gating, and neural geometry parameterization to infer area functions and MRI visualizations from speech; its default forward synthesis runs at 71.5 times real time, but the reconstructed geometry is an acoustically compatible explanation, not unique anatomical ground truth.
Background & Motivation¶
Speech production can be understood as a source filtered by the vocal tract: vocal fold vibration supplies periodic excitation, while the tongue, lips, and jaw change the tube geometry and move its formants; turbulence at constrictions produces consonants. Articulatory synthesizers such as Kelly–Lochbaum, Maeda, and VocalTractLab describe these mechanisms, but sequential time-domain solving is expensive, and separate vowel and consonant branches introduce differentiation discontinuities. Neural vocoders produce realistic waveforms but do not directly explain the articulatory geometry behind the sound.
Recovering articulation from audio requires more than a regression network. Paired speech–MRI or electromagnetic articulography (EMA) data are difficult to collect, and the audio-to-geometry mapping is inherently ill-posed and non-convex: multiple tract shapes can produce similar sounds. Even with a differentiable simulator, independently updating each section's area can create competing gradients and trap optimization in poor local solutions. The paper therefore addresses computational efficiency, consonant differentiability, and the geometry optimization space together, rather than simply attaching automatic differentiation to a classical synthesizer.
The method retains acoustic tube physics but reformulates short-window propagation using parallelizable frequency-domain transfer matrices and couples areas through neural networks. This supports per-utterance optimization, an encoder that predicts areas directly from speech, and back-propagation of acoustic errors into MRI generator latents. Core Idea: use speech itself to supervise vocal tract geometry through differentiable physical synthesis, with neural parameterization improving inversion without paired speech–articulatory measurement data.
Method¶
Overall Architecture¶
The input is target speech; the outputs are a time-varying area function with 32 spatial sections and resynthesized audio, plus a generated vocal tract video in the MRI application. Swift-F0 first estimates pitch, and a Liljencrants–Fant (LF) model constructs a glottal pulse. Neural geometry parameterization supplies the areas to frequency-domain synthesis and differentiable turbulence gating, which produce lip pressure. A multiscale log Mel spectrogram error supervises geometry without ground-truth area labels.
Neural geometry parameterization has three alternative uses, not three networks connected in series: a neural field fits each utterance, a Wav2Vec 2.0 encoder learns speech-to-area mapping, and StyleGAN2 supplies geometry through an MRI image prior. Solid arrows indicate forward synthesis with given parameters; dashed arrows indicate loss evaluation and updates during training or per-utterance inversion. A trained encoder predicts areas directly, whereas neural field fitting and GAN latent inversion still require optimization for new audio.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Target speech"] --> B["Neural geometry parameterization<br/>Three alternative paths"]
A --> S["Pitch estimation and<br/>LF glottal pulse"]
B --> C["Frequency-domain synthesis"]
S --> C
C --> D["Differentiable turbulence gating"]
D --> O["Synthesized audio"]
A -.-> L["Multiscale spectrogram loss"]
O -.-> L
L -.->|Training or per-utterance inversion| B
Key Designs¶
1. Neural geometry parameterization: share optimization structure instead of letting sections compete independently
The most direct representation is a spatial-section-by-time-window area array. However, the paper finds that optimization from a uniform tube can oscillate across sections and fail to recover basic vowels. A neural field takes position and time as inputs to a shared network and outputs the area function, so parameter updates affect multiple locations and implicitly couple them. The authors compare a fully connected network with random Fourier features (RFF) against a multiplicative filter network (MFN), attributing improved coarse-to-fine recovery to shared weights. This is an empirical optimization benefit, not proof that the non-convex problem is globally solved.
The RFF network has 4 layers, a 64-dimensional positional encoding, and a hidden dimension of 256. The MFN has 3 layers, an input scale of 256, and the same hidden dimension of 256. The appendix states that outputs use softmax and minimum areas are clipped at \(10^{-2}\); this note preserves that description rather than changing it to softplus. Positive-area constraints do not amount to full anatomical constraints. Likewise, the main text's claim of requiring no regularizers does not mean that network representation and clipping introduce no priors.
Refitting a neural field for every utterance remains expensive, so the second path amortizes area prediction into an encoder. A raw waveform passes through a CNN and a 12-layer Transformer to obtain areas, which the physical decoder converts back into audio. Training requires speech reconstruction but no area labels, enabling real speech data instead of only synthetic consonant–vowel combinations. Here, self-supervision means the absence of paired articulatory measurement supervision, not unique identification of the true vocal tract.
The third path uses an MRI image prior to restrict searchable geometries. StyleGAN2 is trained on MRI images with dataset identity labels as conditions; inversion fixes the generator and optimizes a 512-dimensional latent per frame. Images undergo sigmoid soft binarization, followed by summation along a fixed guide from the glottis to the lips to approximate cross-sectional areas. The guide combines a quarter-circle arc and a line segment. Acoustic losses can therefore back-propagate through area extraction into the latents, but these two-dimensional pixel sections are acoustic tube proxies rather than directly measured three-dimensional areas.
The MRI model uses \(64\times64\) images and optimizes 30 latents for each 3-second utterance. It does not train an inverse mapping on paired speech–MRI examples, but the visual prior still requires real MRI data and identity labels. The optimized frame count and the 9,900 evaluation video frames use different accounting; the paper does not fully explain the resampling between them.
2. Frequency-domain synthesis: replace sample-by-sample temporal recursion with independent frequency solves
The vocal tract is approximated as a one-dimensional acoustic tube whose cross-sectional area varies along its length. Within a short window, areas are fixed, and the linearized Euler equations define a linear time-invariant filter; formants are peaks in its frequency response. The assumptions are small perturbations and approximately planar waves across each section. The appendix explains why the plane-wave approximation is reasonable below 4 kHz; this is not a full three-dimensional fluid solve.
In the frequency domain, pressure and volume velocity across each section are related by a \(2\times2\) transfer matrix. Multiplying consecutive matrices gives the total transfer matrix, and glottal excitation plus lip radiation impedance determine the tract filter. Using the main text's matrix-element notation, the key relationship is:
Here, \(\mathbf{A},\mathbf{C}\) are total transfer matrix elements, not cross-sectional areas; \(Z_{\text{lips}}\) represents the radiation load at the lip opening. The glottal excitation spectrum passes through this filter, and an inverse Fourier transform produces the output. The appendix incorporates viscous boundary-layer losses and yielding tract walls through equivalent circuit impedances, so this is not merely a freely learned all-pole filter.
The benefit is not that frequency-domain simulation is inherently more realistic. Frequencies can be computed in parallel on a GPU, avoiding gradient propagation through extensive temporal recursion. The vowel interpolation loss plots show smoother landscapes for the implemented frequency-domain synthesizer, followed by improved audio fitting. However, the time-domain baseline replaces frequency-dependent wall-friction resistance with DC resistance. The implementations differ in more than temporal recursion, so their entire quality gap cannot be attributed solely to changing domains.
Dynamic speech uses consecutive windows: a Tukey analysis window with \(\alpha=0.25\) is applied to the source, a Hann synthesis window to the output, and the windows are combined by overlap-add. The default simulation window is 25 ms with a hop ratio of \(1/4\), or 6.25 ms. This is a within-window quasi-static approximation with between-window assembly, not an exact frequency-domain solution for a continuously and rapidly moving tract. Errors during rapid closures or motion require further validation.
3. Differentiable turbulence gating: smoothly vary consonant noise with constriction location and airflow
Linear acoustic propagation synthesizes vowels but does not create nonlinear turbulence at constrictions. Classical auxiliary models inject noise at the narrowest section and activate it using a critical Reynolds number; hard location selection and thresholds interrupt gradients. This method uses softmax over negative areas to distribute noise according to constriction strength, and softplus to smooth the activation gate. The mechanism from Equations (5)–(6) is:
\(u_n\) is the section's volume velocity, \(A_n\) its area, and \(\rho\) and \(\mu\) are air density and dynamic viscosity. Here, \(\alpha\) is the noise gain, not the window parameter used in the previous design. The source \(z_n\) is fractal Perlin noise, and the critical Reynolds number is 3500. Smaller areas increase gating strength, while soft location allocation lets the loss update multiple candidate constrictions rather than only one hard-selected section.
Turbulence is computed in the time domain and transformed into frequency-domain sources. Each source propagates through the transfer function from its injection position to the lips, and its pressure contribution is added to that from glottal excitation. Vowels and consonants can thus transition within one differentiable pipeline, with noise filtered by tract geometry rather than simply appended to the output waveform. This is an optimizable auxiliary turbulence model, not a full Navier–Stokes solve; smooth softplus activation is also not an exact physical threshold.
A Worked Example¶
Consider generating an MRI visualization from a 3-second speech clip. Pitch and RMS energy first determine the LF source. The method initializes 30 latents and generates 30 identity-conditioned MRI images. Soft binarization and the fixed guide convert each image into 32 section areas, forming a coarse temporal geometry sequence.
The simulator computes propagation within 25 ms quasi-static windows and performs overlap-add with a 6.25 ms hop. At constrictions, smoothly gated turbulence noise is filtered by the tube as well. Spectrogram error sends mismatched formants or consonant energy back to image latents, progressively changing the available tongue positions and lip shapes. This example illustrates the reported components without inventing the interpolation rule between the latent sequence and simulation windows, which the paper does not fully specify.
Loss & Training¶
The three inversion paths share a multiscale log Mel spectrogram objective rather than pointwise waveform alignment. Appendix Equation (56) averages \(L_2\) norms across resolutions, whereas the main text calls the objective a squared distance. The displayed formula is retained without adding an unreported square:
Each resolution uses 128 Mel bins. The FFT size, window length, and hop length in \(\mathcal{C}\) are \((512,160,40)\), \((1024,400,80)\), and \((2048,800,160)\), in samples. These loss STFT configurations must not be confused with the 25 ms acoustic simulation window. Multiple resolutions expose resonant structure at different time–frequency scales and reduce dependence on a single spectrogram resolution.
Neural field fitting uses Adam at a learning rate of \(10^{-2}\) for 200 steps; the appendix reports approximately 1 minute to fit 5 seconds of speech. Encoder training uses Adam at \(10^{-4}\), initially training on 3-second English segments with a batch size of 64 for 15,000 iterations, followed by 10 fine-tuning epochs for each other language. The authors describe 800 hours of training audio, not explicitly 800 hours of unique data.
The encoder Transformer has 12 attention heads and a hidden dimension of 768. Training uses LibriTTS-R for English, Multilingual LibriSpeech for other European languages, and KSS, JVS, and AISHELL-3 for Korean, Japanese, and Chinese. MRI reconstruction targets undergo Adobe Express speech enhancement, so visual experiments do not use unprocessed MRI-synchronized recordings.
Key Experimental Results¶
Main Results¶
Frequency- versus time-domain fitting uses 25 random LibriTTS-R utterances, an MFN area representation, and 200 Adam steps. The standard setup uses 16 kHz and 32 spatial sections. SI-SDR, STOI, and PESQ are estimated with TorchAudio-Squim and reflect distortion, intelligibility, and perceptual quality, respectively; higher is better. They are not separately conducted subjective listening evaluations.
| Synthesis | Window / hop ratio | SI-SDR | STOI | PESQ | Real-time factor |
|---|---|---|---|---|---|
| TDS | Not applicable | 7.32 ± 2.42 | 0.77 ± 0.03 | 1.55 ± 0.14 | 1.08× |
| FDS | 25 ms / 1/8 | 17.83 ± 2.19 | 0.93 ± 0.02 | 2.02 ± 0.31 | 35.0× |
| FDS | 25 ms / 1/4 | 16.39 ± 2.42 | 0.93 ± 0.02 | 1.98 ± 0.26 | 71.5× |
| FDS | 50 ms / 1/8 | 17.47 ± 1.94 | 0.93 ± 0.02 | 1.97 ± 0.26 | 36.0× |
| FDS | 50 ms / 1/4 | 15.82 ± 2.47 | 0.93 ± 0.02 | 1.98 ± 0.25 | 70.0× |
These values come from Table 1. Here, real-time factor means synthesized duration divided by forward runtime, not speedup relative to TDS or inversion optimization speed. Figure 3 and the text separately report roughly 70× average speedup over TDS. The text also gives 1.08 seconds versus 14.2 ms to synthesize 1 second, but its TDS runtime of 1.08 seconds conflicts with Table 1's claim of 1.08× faster than real time; the two statements should not be silently reconciled.
The self-supervised encoder is tested on 11 languages, with 200 synthesized test utterances per language. The following excerpt is from Table 3. Lower WER/CER is better. Speaker ID is cosine similarity between ECAPA-TDNN identity embeddings, retained on the table's numeric scale rather than interpreted as identification accuracy; error ranges are 95% confidence intervals.
| Language | Original audio WER / CER | TensorTract2 WER / CER | Ours WER / CER | TensorTract2 ID | Ours ID |
|---|---|---|---|---|---|
| English | 3.91 ± 0.84 / 2.05 ± 0.70 | 13.00 ± 1.82 / 7.37 ± 1.11 | 5.30 ± 0.94 / 2.82 ± 0.75 | 16.86 ± 1.01 | 52.0 ± 1.39 |
| Spanish | 3.35 ± 0.69 / 1.37 ± 0.44 | 9.71 ± 0.93 / 5.13 ± 0.44 | 9.27 ± 1.25 / 4.43 ± 0.59 | 14.74 ± 1.49 | 52.9 ± 1.81 |
| Korean | 8.00 ± 2.42 / 2.10 ± 0.79 | 32.32 ± 4.27 / 14.58 ± 2.42 | 19.45 ± 3.49 / 6.05 ± 1.25 | 23.51 ± 0.77 | 63.4 ± 0.94 |
| Chinese | — / 13.54 ± 3.63 | — / 53.27 ± 42.97 | — / 30.11 ± 13.0 | 17.06 ± 1.12 | 47.9 ± 1.52 |
| Japanese | — / 6.89 ± 1.19 | — / 15.01 ± 2.38 | — / 11.88 ± 1.57 | 23.76 ± 1.41 | 53.9 ± 1.44 |
Whisper computes WER/CER, and original audio also has nonzero recognition error, so results depend on the recognizer and language as well. TensorTract2 trains on synthetic consonant–vowel data, whereas this method trains on real speech with language-specific fine-tuning. The comparison reflects the full training approach, not an isolated simulator effect.
MRI evaluation compares 40 clips of 3 seconds, totaling 9,900 frames. Ours versus Speech2rtMRI yields FVD of 623 versus 2949, SSIM of 0.352 versus 0.317, and LPIPS of 0.159 versus 0.355. Ours also synthesizes audio, with SI-SDR 13.08, STOI 0.92, and PESQ 2.09; the competitor has no audio metrics. The ratio 2949/623 is approximately 4.73, not the order-of-magnitude difference claimed in the text. Low SSIM also indicates that visual matching remains far from exact.
Ablation Study¶
Appendix Table 5 compares area parameterizations on the same 25-utterance fitting setup, showing that representation affects optimization beyond simulator differentiability.
| Area parameterization | SI-SDR | STOI | PESQ |
|---|---|---|---|
| Discrete | 10.78 ± 2.22 | 0.81 ± 0.02 | 1.57 ± 0.12 |
| RFF | 14.51 ± 3.78 | 0.84 ± 0.05 | 1.76 ± 0.26 |
| MFN | 16.39 ± 2.42 | 0.93 ± 0.02 | 1.98 ± 0.26 |
MFN outperforms RFF on all three real-speech metrics, but area reconstruction in Table 2 has no universal winner. For /e/, area MSE in \(\mathrm{cm}^{4}\) is 26.22 ± 1.36 for Discrete, 1.29 ± 0.90 for RFF, and 15.00 ± 7.06 for MFN. For /u/, the respective values are 28.19 ± 2.54, 27.26 ± 22.30, and 19.98 ± 5.04. RFF remains close to the discrete baseline with large variation on /u/, so neural fields do not always recover anatomy accurately.
Key Findings¶
- Default FDS reaches SI-SDR 16.39 versus TDS 7.32, a difference of 9.07. The denser 1/8 hop improves SI-SDR but reduces the real-time factor to 35.0×.
- Improved audio fitting and geometry fitting provide complementary evidence, but correct formants do not guarantee correct areas or a unique MRI tongue position.
- From 1,000 randomly generated MRIs, the authors extract areas and compute the first two formants, observing a space resembling the IPA vowel trapezoid. This supports meaningful acoustic structure, not quantitative validation of an individual's anatomy.
- No numerical ablation separately disables smooth turbulence gating, so the tables cannot assign gains to each of the three contributions.
Highlights & Insights¶
- The method turns physical equations into a frequency-domain circuit suitable for GPU execution and back-propagation, rather than adding automatic differentiation to an old temporal loop. Numerical representation affects optimizability, an issue relevant to other wave inverse problems.
- Shared network weights act as an implicit geometry prior. MFN's real-speech advantage also shows that static geometry error and dynamic acoustic quality need not select the same representation.
- A physical decoder turns audio reconstruction error into supervision without articulatory labels. Transferring this idea requires identifying which latent variables observations constrain; low reconstruction error does not establish identifiability of every hidden variable.
Limitations & Future Work¶
- The authors acknowledge ill-posed inversion: different geometries can produce the same sound. Current MRI results are acoustically grounded visualizations, not diagnostic-grade reconstructions.
- The MRI frames omit the nasal cavity, so nasal geometry is not explicitly reconstructed. Bifurcated transmission lines and equivalent areas are discussed in the appendix, but do not replace validation of actual nasal anatomy and anti-formants.
- One-dimensional plane waves, within-window quasi-static geometry, and auxiliary noise are approximations. Errors should be analyzed by phoneme, rapid closure, and high-frequency range, with higher-fidelity simulators as references.
- Per-utterance fitting uses only 25 speech samples. MRI results rely on a low-resolution prior and enhanced target audio; robustness to unseen speakers, noise, and atypical articulation needs dedicated testing.
- The authors report that DPS struggles with large tongue deformations and performs worse than GAN inversion, but provide no quantitative comparison table. Future work could develop priors better suited to deformation and acoustic constraints, rather than generalizing this observation to all diffusion models.
- The appendix requires clear labeling of simulated samples. Educational or clinical applications should also expose uncertainty and multiple solutions instead of presenting one acoustically compatible result as a real internal observation.
Related Work & Insights¶
- vs VocalTractLab / Maeda: Classical articulatory synthesis retains physical interpretation; this paper emphasizes GPU parallelism and continuous gradients. Its advantage is integration with learning systems, not replacement of all high-fidelity articulatory models.
- vs DDSP / neural source-filter: All use differentiable signal-processing chains, but this method determines filters through areas, transfer impedances, and lip radiation, linking latent variables to geometry rather than only acoustic filter parameters.
- vs Südholt et al.'s gradient-based area inversion: The earlier work primarily demonstrates static vowels and consonants with explicit areas. This method improves optimization through neural fields and extends to moving words, encoders, and MRI priors.
- vs TensorTract2 / Speech2rtMRI: Paired-data models learn articulatory encoding or video generation. Here, acoustic consistency replaces paired inverse supervision, but real speech, visual priors, and independent anatomical validation remain necessary.
Rating¶
- Novelty: 4/5 — Combines differentiable frequency-domain acoustics, consonant gating, and neural geometry inversion into reusable infrastructure.
- Experimental Thoroughness: 3/5 — Covers multiple languages and MRI, but per-utterance fitting is limited and independent gating ablation and ambiguity evaluation are missing.
- Writing Quality: 3/5 — Mechanisms and appendices are substantial, but table references, squared-loss wording, and runtime conventions are inconsistent.
- Value: 4/5 — Useful for interpretable articulation modeling and self-supervised acoustic inversion; medical use requires additional validation.