Audio Samples · IEEE/ACM TASLP Submission
Supplementary listening page. Synthesised speech is decomposed into an articulation-conditioned filter and a flow-matched glottal source, enabling explicit, physically grounded control over the vocal tract response.
Ten utterances per test set, synthesised by each system under evaluation. Use the toggle to switch between LibriTTS test-clean and test-other. The proposed model column is shaded.
| Utterance | Matcha-TTS | StyleTTS2 | YourTTS | XTTS | HierSpeech++ | Source-Filter | Proposed |
|---|
Place files at audio/baselines/{set}/{system}_{NN}.wav — e.g. audio/baselines/clean/proposed_01.wav
Ten utterances from the proposed model, each shown as five signals with spectrograms: the source excitation before and after OT-CFM refinement, the articulation-conditioned filter, and the combined output both before and after refinement — isolating what the OT-CFM module contributes to the final audio.
Audio: audio/components/{type}_{NN}.wav · Spectrogram: spectrograms/{type}_{NN}.png
where {type} ∈ src_pre, src_post, filter, combined_pre, combined_post
The source (pitch, energy, temporal) is routed from Speaker A while the filter (vocal-tract kinematics) is routed from Speaker B. All clips are model-generated; the matched conditions (A–A, B–B), where source and filter share one speaker, serve as self-reconstruction baselines for the cross-speaker A–B routing.
Audio: audio/crossspeaker/{combo}_{NN}.wav · Spectrogram: spectrograms/crossspeaker_{combo}_{NN}.png
where {combo} ∈ AA, BB, AB
Audio corresponding to the kinematic analyses in the paper: the minimal pair “pat”/“cat” (Fig. 2), the American rhotic in “cart” (Fig. 3), and the rhotic→non-rhotic edit of “bird” (Fig. 4).
Lower-lip activity for the bilabial /p/ versus tongue-dorsum elevation for the velar /k/, with a shared /-æt/ rhyme. See Fig. 2 in the paper for the kinematic trajectories.
F3 lowering toward F2 with the simultaneous retroflex tongue-tip elevation and retraction characteristic of the post-vocalic /r/. See Fig. 3 in the paper for the kinematic trajectories.
Flattening the tongue-tip and tongue-body elevation suppresses the F3 drop, converting the rhotic American realisation into a non-rhotic one — with no acoustic post-processing. See Fig. 4 in the paper for the kinematic trajectories.
Audio: audio/figures/{name}.wav · Spectrogram: spectrograms/{name}.png
where {name} ∈ pat, cat, cart, bird_rhotic, bird_nonrhotic