2001: A Space Odyssey
How To Create Text-To-Speech Audio Files:
A Practical Step-By-Step Course For Beginners
Introduction To Text-To-Speech
Background & Brief History Of Text-To-Speech.
Popular Text-To-Speech Engines.
Text-To-Speech Terms.
© [Link]
A Brief History Of Text-To-Speech
© [Link]
A Brief History Of Text-To-Speech
[Link]
© [Link]
A Brief History Of Text-To-Speech
Kratzenstein's resonators
VODER speech synthesizer
von Kempelen's speaking machine [Link]
© [Link]
A Brief History Of Text-To-Speech
© [Link]
A Brief History Of Text-To-Speech
[Link]
© [Link]
A Brief History Of Text-To-Speech
Milestones In Speech Synthesis
Amiga SoftVoice speech synthesis
[Link]
© [Link]
A Brief History Of Text-To-Speech
© [Link]
A Brief History Of Text-To-Speech
[Link]
© [Link]
A Brief History Of Text-To-Speech
[Link]
© [Link]
A Brief History Of Text-To-Speech
© [Link]
A Brief History Of Text-To-Speech
(The Office - Season 3, Episode 13 Travelling Salesman)
© [Link]
A Brief History Of Text-To-Speech
• Learning
• Teaching
• Sales
• News
• Information
• Entertainment
• Recipes
• Home
• Office
© [Link]
A Brief History Of Text-To-Speech
© [Link]
Text-To-Speech (TTS) Technologies
© [Link]
Text-To-Speech (TTS) Technologies
Overview Of A Typical Text-To-Speech System
[Link]
© [Link]
Text-To-Speech (TTS) Technologies
Qualities Of A Great Speech Synthesis System
• NATURALNESS - how closely the synthetic generated voice sounds like human speech.
• INTELLIGIBILITY - how easily the speech can be understood.
The ideal speech synthesizer aims to sound as natural and
intelligible as possible!
© [Link]
Text-To-Speech (TTS) Technologies
Concatenative Speech Synthesis
• A very large database of short speech fragments (units) are recorded from
a single speaker and recombined to form complete utterances.
• Stringing segments of recorded speech together.
• PROS: Produces natural-sounding synthesized speech.
• CONS: It’s difficult to modify the voice (for example switching to a
different speaker or altering the emphasis or emotion of their speech)
without recording a whole new database.
Many Text-To-Speech Applications Use Concatenative
Synthesis
© [Link]
Text-To-Speech (TTS) Technologies
Parametric Speech Synthesis
[Link]
© [Link]
Text-To-Speech (TTS) Technologies
WaveNet Speech Synthesis
[Link]
© [Link]
Text-To-Speech (TTS) Technologies
WaveNet Speech Synthesis
• Same technology used to create speech for Google Assistant, Google Search, and Google Translate.
• Sounds more natural than other text-to-speech systems.
• Most people prefer WaveNet speech audio over other text-to-speech technologies.
[Link]
© [Link]
Text-To-Speech (TTS) Technologies
© [Link]
Text-To-Speech Engines
© [Link]
Text-To-Speech Engines
TTS engines provide users with access to text-to-speech functionality.
• Microsoft Speak (Word, Outlook, PowerPoint, Etc.)
• Amazon Polly
• Google Text-To-Speech
© [Link]
Text-To-Speech Terms
Here are some common terms we’ll be using in this course:
• TTS (Text-To-Speech)
• Speech Synthesis (e.g. Concatenative, Parametric, WaveNet)
• Neural Networks
• Machine Learning AI (Artificial Intelligence) Voices
• SSML & Markup Tags
• Prosody (speech volume, pitch & speed)
• Phonemes & Phonetic Pronunciations
© [Link]
How To Create Text-To-Speech Audio Files:
A Practical Step-By-Step Course For Beginners
End Of Lesson
© [Link]
How To Create Text-To-Speech Audio Files:
A Practical Step-By-Step Course For Beginners
© [Link]