Speech Synthesis
-
- KVRist
- 35 posts since 5 Feb, 2007
Can anyone tell me how/with which tools this was synthesized? Sounds too fluid to be patched together from a phoneme bank (unless it's one friggin huge one), and unlike anything I've heard before, except maybe speak'n'spell and similar LPC speech synths...
http://noouch.de/audio/voxcollection.mp3
http://noouch.de/audio/voxcollection.mp3
-
- KVRian
- 604 posts since 7 Jul, 2004 from Somewhere between the 2nd and 3rd dimensions.
Sounds like the Speech Synth in FL Studio on its "Munchkin" setting, sped up and fed through a phaser.
Well, it does!
Well, it does!

Analogue or digital – which is better? There's only one way to find out... FI-I-IGHT!!!
-
- KVRian
- 604 posts since 7 Jul, 2004 from Somewhere between the 2nd and 3rd dimensions.
Doesn't sound vocoded to me, the pitch varies too much (i.e. the carrier is too natural). Maybe it's just human speech that's been timestretched.gsoto wrote:Doesn't sound synthesized to me. And too clean to be a vocoder, but I never got into vocoders, so I can't tell.
Anyway, I am the Gorfian consciousness.
EDIT: Reminds me of Software Speech - used on old Epyx / US Gold games on the C64. It was human speech, digitised at an extraordinarily low bitrate and compressed like hell. It wasn't synthesised but did sound very cool.

Analogue or digital – which is better? There's only one way to find out... FI-I-IGHT!!!
-
- KVRian
- 604 posts since 7 Jul, 2004 from Somewhere between the 2nd and 3rd dimensions.
I'll try again [wipes rabid foam from chin]...
I reckon this is synthesised, but there is a lot of distortion (intentional?). This is my quick and dirty attempt at a similar effect, could probably get it somewhere near with some tweaking. Amazed at how clearly it says "light bulb" and "wank".
Don't know what system it is, imagine it's some mid-80s Votrax hardware that pre-dates the system Stephen Hawking uses. Would be interesting to know though.
I reckon this is synthesised, but there is a lot of distortion (intentional?). This is my quick and dirty attempt at a similar effect, could probably get it somewhere near with some tweaking. Amazed at how clearly it says "light bulb" and "wank".
Don't know what system it is, imagine it's some mid-80s Votrax hardware that pre-dates the system Stephen Hawking uses. Would be interesting to know though.
Last edited by Modeler on Mon Jun 18, 2007 10:59 pm, edited 1 time in total.

Analogue or digital – which is better? There's only one way to find out... FI-I-IGHT!!!
-
- KVRAF
- 3222 posts since 23 Dec, 2002
-
- KVRian
- 673 posts since 15 Nov, 2004 from Montevideo, Uruguay
lol, Modeler. Nice attempt.
But yours is clearly text to speech. Without being an expert I can hear the common artifacts of synthesized voice. In the original mp3 the articulation is too perfect, I can't spot a single strange thing in the entire audio. Listen to the s's, they are too natural.
But yours is clearly text to speech. Without being an expert I can hear the common artifacts of synthesized voice. In the original mp3 the articulation is too perfect, I can't spot a single strange thing in the entire audio. Listen to the s's, they are too natural.
-
- KVRian
- 604 posts since 7 Jul, 2004 from Somewhere between the 2nd and 3rd dimensions.
Agreed, but it pronounces "reservoir" as "reserv-oyer".gsoto wrote:In the original mp3 the articulation is too perfect, I can't spot a single strange thing in the entire audio. Listen to the s's, they are too natural.

Analogue or digital – which is better? There's only one way to find out... FI-I-IGHT!!!
-
- KVRist
- Topic Starter
- 35 posts since 5 Feb, 2007
One thing I noticed is that all the S words start almost exactly the same. Modeler's example comes really close. Less distortion and breath and it's pretty much spot-on. Care to share the settings for this?
-
- KVRian
- 604 posts since 7 Jul, 2004 from Somewhere between the 2nd and 3rd dimensions.
Sure, but I didn't spend much time refining it. Speech synth is "Large Male" set at pitch level C2, channel volume ~50%. Put it through Delay Bank with "Time" set to 1:00 and volume set at ~20%. I put the mix of that through the JCM900 simulator that comes with Guitar Suite (free) and put it on channel B - that's it. Childish swearing is optional.noouch wrote:Modeler's example comes really close. Less distortion and breath and it's pretty much spot-on. Care to share the settings for this?
Have fun.

Analogue or digital – which is better? There's only one way to find out... FI-I-IGHT!!!
