Audio Samples · IEEE/ACM TASLP Submission

Articulatory Source-Filter TTS:
Physically Grounded Control through Vocal Tract Kinematics

Supplementary listening page. Synthesised speech is decomposed into an articulation-conditioned filter and a flow-matched glottal source, enabling explicit, physically grounded control over the vocal tract response.

Jesuraj Bandekar/Shinji Watanabe/Prasanta Kumar Ghosh

01

Baseline Comparison

Ten utterances per test set, synthesised by each system under evaluation. Use the toggle to switch between LibriTTS test-clean and test-other. The proposed model column is shaded.

Utterance Matcha-TTS StyleTTS2 YourTTS XTTS HierSpeech++ Source-Filter Proposed

Place files at audio/baselines/{set}/{system}_{NN}.wav — e.g. audio/baselines/clean/proposed_01.wav

02

Source & Filter Decomposition

Ten utterances from the proposed model, each shown as five signals with spectrograms: the source excitation before and after OT-CFM refinement, the articulation-conditioned filter, and the combined output both before and after refinement — isolating what the OT-CFM module contributes to the final audio.

source stream filter stream

Audio: audio/components/{type}_{NN}.wav  ·  Spectrogram: spectrograms/{type}_{NN}.png
where {type} ∈ src_pre, src_post, filter, combined_pre, combined_post

03

Cross-Speaker Source / Filter Recombination

The source (pitch, energy, temporal) is routed from Speaker A while the filter (vocal-tract kinematics) is routed from Speaker B. All clips are model-generated; the matched conditions (A–A, B–B), where source and filter share one speaker, serve as self-reconstruction baselines for the cross-speaker A–B routing.

Audio: audio/crossspeaker/{combo}_{NN}.wav  ·  Spectrogram: spectrograms/crossspeaker_{combo}_{NN}.png
where {combo} ∈ AA, BB, AB

04

Articulatory & Accent Control

Audio corresponding to the kinematic analyses in the paper: the minimal pair “pat”/“cat” (Fig. 2), the American rhotic in “cart” (Fig. 3), and the rhotic→non-rhotic edit of “bird” (Fig. 4).

Minimal pair — “pat” vs “cat” · Fig. 2

Lower-lip activity for the bilabial /p/ versus tongue-dorsum elevation for the velar /k/, with a shared /-æt/ rhyme. See Fig. 2 in the paper for the kinematic trajectories.

“pat”
pat spectrogram
“cat”
cat spectrogram

American rhotic — “cart” · Fig. 3

F3 lowering toward F2 with the simultaneous retroflex tongue-tip elevation and retraction characteristic of the post-vocalic /r/. See Fig. 3 in the paper for the kinematic trajectories.

“cart” — synthesised
cart spectrogram

Accent edit — “bird” rhotic → non-rhotic · Fig. 4

Flattening the tongue-tip and tongue-body elevation suppresses the F3 drop, converting the rhotic American realisation into a non-rhotic one — with no acoustic post-processing. See Fig. 4 in the paper for the kinematic trajectories.

Original — rhotic
bird rhotic spectrogram
Modified — non-rhotic
bird non-rhotic spectrogram

Audio: audio/figures/{name}.wav  ·  Spectrogram: spectrograms/{name}.png
where {name} ∈ pat, cat, cart, bird_rhotic, bird_nonrhotic