Chapter one
Introduction
1.0Background Concept/ Study
Text to speech (TTS) is a natural language modelling process that requires
changing units of text into units of speech for audio presentation. This is the
opposite of speech to text, where a technology takes in spoken words and tries
to accurately record them as text. Text to speech is now common in
technologies that seek to render audio output from digital text to assist those
who are unable to read, or for other kinds of uses.
Text-to-speech synthesis -TTS – is the automatic conversion of a text into
speech that resembles, as closely as possible, a native speaker of the language
reading that text. Text-to-speech/ Audio system is the technology which lets
computer speak to you. The TTS system gets the text as the input and then a
computer algorithm which called TTS engine analyses the text, pre-processes
the text and synthesizes the speech with some mathematical models. The TTS
engine usually generates sound data in an audio format as the output.
Speech synthesis can be described as artificial production of human speech. A
computer system used for this purpose is called a speech synthesizer, and can be
implemented in software or hardware. A text-to-speech (TTS) system converts
normal language text into speech. Synthesized speech can be created by
concatenating pieces of recorded speech that are stored in a database. Systems
differ in the size of the stored speech units; a system that stores phones or
diaphones provides the largest output range, but may lack clarity. For specific
usage domains, the storage of entire words or sentences allows for high-quality
output.
1.1 Statement of the problem
The problem area in speech synthesis is very wide. There are several problems
in text pre-processing, such as numerals, abbreviations, and acronyms. This
system will help solve the problems of those with learning disabilities, some
people have basic literary levels. They often get frustrated trying to browse the
internet because so much of it is in text form. People with visual impairment –
Text to speech can be a very useful tool for the mild or moderately visually
impaired.
Even for people with the visual capability to read, the process can often cause
too much strain to be of any use or enjoyment. With text to speech, people with
visual impairment can take in all manner of content in comfort instead of strain.
1.2 Aim And Objective Of The Study
The main objective of this research is to design and implement a Text-to-
Speech/Audio System. The Speech/Audio system focuses precisely on the
following objectives:
I. To Design and Implement a Speech synthesizer that converts text to
audio.
II. To Design and Implement a System that can read out text in any
frequency that user specifies.
III. To design and implement a speech synthesizer that can read out text in
both female and male voices.
1.3 Justification/Significance of Study
The significance of this study is:
The application will build a platform to aid people with disabilities especially
on reading and also help get information easily without any stress.
The project could also help children learn how to pronounce words and how to
read.
The study will serve as a foundation and guide to other research students
interested in researching on Text-to-Speech systems.
1.4 Scope of the Study
The scope of the research is focused on implementing a text to speech (TTS)
system to improve the usage of text documents and in order to achieve a more
flexible speech system.
1.5 Limitations of the Study
During the development of the research, this were some drawbacks
encountered.
Limited research material available at the school library and on the internet.
High cost of setting up the system as it requires a high programming language.
Combining school work and carrying out the research.