Speech Emotion Recognition with Python
Speech Emotion Recognition with Python
Abstract
Over the last few decades, there has been a tremendous amount of research on the use of
However, in recent years, research has concentrated on using deep learning for
speechrelated applications. This new machine area Learning has produced far superior
results in a variety of applications, including speech, when compared to others, and has
thus become a very appealing area of research. For this we use neural networks
methods, our proposed classifier makes easy for the speaker to get their emotions, our
classifier predicts the speaker emotions entered manually. The result showed algorithms
i
Acknowledgement
First and foremost, we are grateful for the prayers of our mothers, fathers, and our
classmates unwavering love and unwavering support throughout our lives as well as our
studies Without you, this thesis project would not have been accomplished, and you
have truly molded us into the people we are today We would like to thank the Almighty
ALLAH for the gift of life, intelligence, and understanding that he has bestowed upon
us, as well as the cause for our being. And to our families for the love and support they
had provided throughout our life. Also, we would like to express our sincere
regard as our mentor and supervisor, we thank him for the expertise and intelligence he
has displayed while supervising this project. This excellent work, we believe, is a result
of his excellent supervision and cooperation. Finally, we'd want to thank our Faculty's
instructors on their outstanding work during the course of our four-year curriculum.
iii
2.3 Types of machine learning ......................................................................................
8
15 2.7
conclusion ............................................................................................................. 26
iv
3.4.1 Hardware Component ....................................................................................
30 3.4.2. Software
component ..................................................................................... 30
v
5.4 Data pre-processing ...............................................................................................
41
44 5.7.2 Dashboard
page .............................................................................................. 45
47 6.3
limitations .............................................................................................................. 48
work ........................................................................ 50
References ...................................................................................
52
vi
LIST OF FIGURES
FIGURES PAGE
1Fig 2.1: List of Development in Technologies used for Speech Recognition .............
17
vii
11Figure 5.6 Login Page .................................................................................................
44
List of Tables
viii
CHAPTER I: INTRODUCTION
1.0 Introduction
Speech is one of the primary means of communication among human beings. It is not
only used to convey linguistic information but also to express paralinguistic cues such
as emotions, attitudes, and psychological states. Human interactions heavily rely on the
tone, rhythm, pitch, and intensity of speech to decode the speaker’s emotional condition.
Effective communication, therefore, is not only based on what is said, but also how it is
said.
expressions help individuals connect socially, offer empathy, and enhance mutual
human-computer interaction more natural and user-friendly. Devices or systems that can
detect and respond to emotional cues can significantly improve applications such as
As we move toward developing intelligent systems, the ability to analyze and process
(SER) is an emerging research area that focuses on identifying and classifying emotions
from speech signals using computational methods. The integration of SER in real-time
applications aims to bridge the gap between human and machine communication by
making systems more aware, responsive, and adaptive to the emotional context of the
user.
1
1.1 Background of the study
Over the last few decades, there has been a tremendous amount of research on the use of
However, in recent years, research has concentrated on using deep learning for
speechrelated applications. This new machine area Learning has produced far superior
results in a variety of applications, including speech, when compared to others, and has
In earlier days, people used speech as a means of communication or the way a listener is
conveyed by voice or expression. But the idea of machine learning and various methods
are necessary for the recognition of speech in the matter of interaction with machines.
With a voice as a bio-metric through use and significance, speech has become an
Sreyas
2
Automatic SER aids virtual assistants and smart speakers with better understanding its
users, particularly when they detect phrases with ambiguous meaning. For instance, the
word "truly" can be used both positively and negatively to underline and stress out a
statement or to cast doubt on a reality. Read the following phrases several times: I like
having that gadget a lot. The same program can translate between languages, which is
especially useful given that different languages have distinct conventions for expressing
understand the feelings of others by conveying our feelings and giving feedback to
others. Research has revealed the powerful role that emotion plays in shaping human
social interaction. Emotional displays convey considerable information about the mental
state of an individual. This has opened up a new research field called automatic emotion
recognition, having basic goals to understand and retrieve desired emotions. In prior
studies, several modalities have been explored to recognize the emotional states such as
Speech analysis has become a crucial tool in closing the gap between the real and digital
There are several uses for speech emotion recognition. The main goal of this essay is to
identify spoken emotions and categorize them into seven different emotion output
classes: anger, boredom, disgust, anxiety, happiness, sorrow, and neutral. Research
studies have provided evidence that human emotions in fluency the decision making
It is known that some physiological changes occur in the body due to people's emotional
state. Some variables such as pulse, blood pressure, facial expressions, body
movements, brain waves, and acoustic properties vary Often in the interest to increase
the acceptability of speech technology for human users. The speech signal
information about speaker’s emotions personality’s attitude feelings levels of stress and
current mental states. depending on the emotional state. Pulse, blood pressure, brain
waves, and so forth. Although changes cannot be detected without a portable medical
device, facial expressions and voice signals can be received directly without connecting
Words are not enough to correctly understand the mood and intention of speaker and
paramount importance. This can be achieved by the researching and creating methods of
speech modeling and analysis that embrace to signal linguistic and emotional aspects of
communication. In this study we are going to develop a speech recognition system that
•To collect and prepare a dataset of speech recordings labelled with different
emotions.
•To train and optimize the chosen machine learning algorithm(s) on the dataset,
• How to train and optimize the chosen machine learning algorithm(s) on the
Emotions recognition has wide scope in many areas such as human computer
This study will define the purpose of the study speech recognition through emotions. So
it provides insight into artificial intelligence or machine intelligence that uses various
supervised and unsupervised machine learning algorithms to simulate the human brain
This project’s geographical scope or context is in Mogadishu, Somalia. This study will
The following systems can be cited as an example of the areas in which these studies are
Education: a course system for distance education can detect bored users so that they
can change the style or level of material provided in addition, provide emotional
incentives or compromises.
Automobile: driving performance and the emotional state of the driver are often linked
internally. Therefore, these systems can be used to promote the driving experience and
5
Security: They can be used as support systems in public spaces by detecting extreme
integrated with the interactive voice response system, it can help improve customer
service.
Health: It can be beneficial for people with autism who can use portable devices to
understand their own feelings and emotions and possibly adjust their social behaviour
accordingly.
2.0 introduction
This chapter will go over several speech emotion recognition systems and how they
relate to the speech emotion recognition system. Which include, machine learning voice
recognition, emotion recognition from speech and other types of speech recognition will
be discussed.
machine learning to process and classify speech signals in order to detect emotions.
6
2.1 History of machine learning
The term machine learning was first coined in the 1950s when Artificial Intelligence
pioneer Arthur Samuel built the first self-learning system for playing checkers. He
noticed that the more the system played, the better it performed.
Fuelled by advances in statistics and computer science, as well as better datasets and the
growth of neural networks, machine learning has truly taken off in recent years.
translation, image recognition, voice search technology, self-driving cars, and beyond.
Machine learning (ML) is a branch of artificial intelligence (AI) that enables computers
to “self-learn” from training data and improve over time, without being explicitly
programmed. Machine learning algorithms are able to detect patterns in data and learn
from them, in order to make their own predictions. In short, machine learning
instruct a computer how to transform input data into a desired output. Instructions are
mostly based on an IF-THEN structure: when certain conditions are met, the program
Machine learning, on the other hand, is an automated process that enables machines to
solve problems with little or no human input, and take actions based on past
observations.
7
2.3 Types of machine learning
Gartner, a business consulting firm, predicts that supervised learning will remain the
2022. This type of machine learning feeds historical input and output data in machine
learning algorithms, with processing in between each input/output pair that allows the
algorithm to shift the model to create outputs as closely aligned with the desired result
This machine learning type got its name because the machine is “supervised” while it's
learning, which means that you’re feeding the algorithm information to help it learn.
The outcome you provide the machine is labelled data, and the rest of the information
For example, if you were trying to learn about the relationships between loan defaults
and borrower information, you might provide the machine with 500 cases of customers
who defaulted on their loans and another 500 who didn't. The labelled data “supervises”
the machine to figure out the information you're looking for. Supervised learning is
8
2.3.2 Unsupervised learning
While supervised learning requires users to help the machine learn, unsupervised
learning doesn't use the same labelled training sets and data. Instead, the machine looks
for less obvious patterns in the data. This machine learning type is very helpful when
you need to identify patterns and use data to make decisions. Common algorithms used
clustering, and Gaussian mixture models. Using the example from supervised learning,
let's say you didn't know which customers did or didn't default on loans. Instead, you'd
provide the machine with borrower information and it would look for patterns between
This type of machine learning is widely used to create predictive models. Common
applications also include clustering, which creates a model that groups objects together
based on specific properties, and association, which identifies the rules existing between
Pinpointing associations in customer data (for example, customers who buy a specific
Reinforcement learning is the closest machine learning type to how humans learn. The
algorithm or agent used learns by interacting with its environment and getting a positive
networks, and Q-learning. Going back to the bank loan customer example, you might
classifies them as high-risk and they default, the algorithm gets a positive reward. If
9
they don't default, the algorithm gets a negative reward. In the end, both instances help
the machine learn by understanding both the problem and environment better. Gartner
notes that most ML platforms don't have reinforcement learning capabilities because it
requires higher computing power than most organizations have. Reinforcement learning
is applicable in areas capable of being fully simulated that are either stationary or have
large volumes of relevant data. Because this type of machine learning requires less
management than supervised learning, it’s viewed as easier to work with dealing with
unlabelled data sets. Practical applications for this type of machine learning are still
Training robots to learn policies using raw video images as input that they can use to
has gained prominence and use with the rise of artificial intelligence and intelligent
Voice recognition systems let consumers interact with technology simply by speaking to
Voice recognition can identify and distinguish voices using automatic speech
recognition (ASR) software programs. Some ASR programs require users first train the
program to recognize their voice for a more accurate speech-to-text conversion. Voice
10
Although voice recognition and speech recognition are referred to interchangeably, they
aren't the same, and a critical distinction must be made. Voice recognition identifies the
a signal, it must have a digital database of words or syllables as well as a quick process
for comparing this data to signals. The speech patterns are stored on the hard drive and
loaded into memory when the program is run. A comparator checks these stored patterns
against the output of the A/D converter -- an action called pattern recognition. Speech is
the basic way of interaction between the listener to the speaker by voice or expression.
Humans can easily understand the speakers' message, but machines can't understand the
Speech recognition gives the usual impression to believe the subject on recurrent neural
To correctly recognize the emotion from speech data, it is very important to extract the
features which accurately represent the emotional aspect of speech signals. One of the
biggest challenges in this field is to extract efficient features for the best classification of
emotions. Some notable works in this area include analysis and synthesis of emotional
Existing studies have found that MFCCs are a far preferable way of analysing emotions
compared to other commonly used speech features (e.g., loudness, formants, linear
11
Speech emotion analysis is the process of identifying vocal cues in speech that serve as
possible to forecast a speaker's emotional state using objectively quantifiable clues. This
idea is acceptable given that emotional emotions trigger physiological responses that
have an impact on how we speak. For instance, the emotional state of dread typically
causes muscle tension, a quick heartbeat, rapid breathing, and perspiration. The
vibration of the vocal folds and the structure of the vocal tract alter as a result of these
physiological actions. All of this has an impact on the vocal qualities of speech, which
enable the listener to discern the emotional state that the speaker is experiencing
(Gjoreski,Gjoresk, 2014).
There are various technologies and methods which are used to recognize and processing
of speech which have their own methods and way of functioning to recognize the
following articles. In these, the trend-based and application-based publications are about
vector machines for speech-based emotional recognition, age, gender, and speech-based
Recently, the combination of different kinds of features has been widely used for
emotion recognition in speech (Bhavan et al., 2019) Showed that the combination of
MFCCs, Melenergy spectrum dynamic coefficients (MEDCs) and energy with a SVM
difference between MEDCs and MFCCs is that MEDCs are calculated as the
logarithmic average of energies after the filter bank, while MFCCs are calculated as the
logarithmic after the filter bank (Chen et al., 2012) used a three-level speech emotion
recognition model to solve the speaker independent emotion recognition problem and
12
extracted the energy, zero crossing rate (ZCR), pitch, the first to third formants,
and five Mel frequency bands energy. The three levels classify the six emotions
pairwise, with each level providing finer classification than the last Schuler (Chen et al.,
2012).
Proposed the usage of a multiple-stage classifier with a support vector machine over 7
emotional classes, with the aim of employing both acoustic and linguistic features for
emotion classification. A deep belief network was used for spotting emotional
Networks, Nearest Neighbours) were used for training on the acoustic features, and then
combined with the belief network using a neural network and their performances
multiple models in other to give a final, (usually) stronger learner. Such methods have
been used in a wide range of application areas — including credit scoring, medical
diagnosis(Wang et al., 2011). Ensemble methods have been applied to audio data as well
accuracy on data scraped from movie content. Morrison (Morrison et al., 2007).
Particularly, bagged ensembles of support vector machines have been analyzed in a few
works (Kim et al., 2002) used such an ensemble for the problem of fault detection in
rotating machinery. However, such works are few and far in between, and we hope to
further analyze this model and apply it for emotion recognition from speech (Hu et al.,
2007).
13
2.5 Sensory modalities for emotion expression
There is vigorous debate about what exactly individual can express nonverbally.
Humans can express their emotions through many different types of nonverbal
physiological signals of the human body. In this section, we discuss each of these
The human face is extremely expressive, able to express countless emotions without
saying a word. And unlike some forms of nonverbal communication, facial expressions
are universal. The facial expressions for happiness, sadness, anger, surprise, fear, and
2.5.2 Speech
In addition to faces, voices are an important modality for emotional expression. Speech
is a relevant communicational channel enriched with emotions: the voice in speech not
only conveys a semantic message but also the information about the emotional state of
the speaker. Some important voice feature vectors that have been chosen for research
social media and Machine Learning electrocardiogram (ECG), respiration (RSP), blood
pressure (BP), electromyogram (EMG), skin conductance (SC), blood volume pulse
(BVP), and skin temperature (ST). Using physiological signals to recognize emotions is
also helpful to those people who suffer from physical or mental illness thus exhibit
14
2.6 Related Work
In the last few years, new method is introduced where static feature vectors are obtained
functionals, by using this approach a big number of large feature vectors is obtained.
The downside is that not all of the feature vectors are of good value, especially not for
used(Gjoreski,Gjoresk, 2014).
that results in physical and psychological changes. These changes influence thought and
behaviour. According to other theories, emotions are not causal forces but simply
is based on three measures that create a three-dimensional space that describes all the
(Abbaschian et al., 2021) A combination of these qualities will create a vector that will
be in one of the defined emotion territories, and based on that, we can report the most
relevant emotion (Mehrabian, 1996) Using pleasure, arousal, and dominance, we can
describe almost any emotion, but such a deterministic system will be very complex to
use statistical models and cluster samples into one of the named qualitative emotions
such as anger, happiness, sadness, and so forth. To be able to classify and cluster any of
the mentioned emotions, we need to model them using features extracted from the
15
speech; this is usually done by extracting different categories of prosody, voice quality,
Any of these categories have benefits in classifying some emotions and weaknesses in
detecting others. The prosody features usually focused on fundamental frequency (F0),
speaking rate, duration, and intensity, are not able to confidently differentiate angry and
The immediate advantage that they have compared to prosody features is that they can
confidently distinguish angry from happy. However, an area of concern is that the
magnitude and shift of the formants for the same emotions vary across different vowels,
and this would add more complexity to an emotion recognition system, and it needs to
speech algorithms, the main areas of speech enhancements include noise reduction and
16
1Fig 2.1: List of Development in Technologies used for Speech Recognition
Teleconference, recognition of speech and audio equipment. Three simple classes can
group speech enhancement algorithms for noise reduction: spectral restoration, methods
based on the model and filtering techniques (Professor, Sreyas Institute of Engineering
In 2020, published a brief review on the importance of speech emotion datasets and
classification approaches, including SVM and HMM. The strength of the research is the
weakness is the leak of more modern methods’ investigation and briefly mentions
17
The main question is: Are there any objective feature profiles of the voice that can be
used for speaker emotion recognition? A lot of studies are done for the sake of providing
such feature profiles that can be used for representation of the emotions, but results are
not always consistent. For some basic problems like distinguishing normal speech from
angry speech or distinguishing normal speech from bored speech the experimental
results converge, For example such converging results are showing that compared to
normal speech, when expressing fear or happiness human speak with higher pitch
This simple analysis is just an example of how we can compare speech signals by using
their physical characteristics. This simple approach cannot be used for speech emotion
recognition. The problem arises when we have to distinguish emotional states like anger
from happiness or fear from happiness. By using the basic speech audio features for
describing these emotional states, the feature profiles are quite similar so distinguishing
Recognition of emotions in audio signals has been a field of extensive study in the past.
Previous work in this area included use of various classifiers like SVM, Neural
Networks, Bayes Classifier etc. The number of emotions classified varied from study to
study, they play an important aspect in evaluating the accuracy of the different
classifiers. Reduction in the number of emotions used for recognition has generated
more accurate results as depicted below. The following table summarises the previous
(%)
18
[Kamran Soltani, Two layer Neural 6 77.1
Ainon,2007]
[Li Wern Chew, Kah PCA, LDA and RBF 6 (divided into
three
Phooi Seng, Li-Minn independent 81.67
classes)
Ang, Vish
Ramakonar, 2011]
2008]
2012]
emotion, and it is still an open problem in psychology. According to Plutchik, more than
ninety definitions of emotion were proposed in the twentieth century. Emotions are
these definitions, two models have become common in speech emotion recognition:
discrete emotional model, and dimensional emotional model(Akçay & Oğuz, 2020).
19
Automatic SER helps smart speakers and virtual assistants to understand their users
better, especially when they recognize dubious meaning words. For example, the term
“really” can be used to question a fact or emphasize and stress out a statement in both
positive and negative ways. Read the following sentences in different ways: “I really
liked having that tool.” The same application can help translate from one language to
through speech. SER is also beneficial in online interactive tutorials and courses.
Understanding the student’s emotional state will help the machine decide how to present
the rest of the course contents. Speech emotion recognition can also be very
instrumental in vehicles’ safety features. . It can recognize the driver’s state of mind and
help prevent accidents and disasters. Another related application is in therapy sessions;
by employing SER, therapists will understand their patients’ state and possibly
underlying hidden emotions as well. It has been proven that in stressful and noisy
environments like aircraft cockpits, the application of SER can significantly help to
increase the performance of automatic speech recognition systems. The service industry
and e-commerce can utilize speech emotion recognition in call centre’s to give early
alerts to customer service and supervisors of the caller’s state of mind. In addition,
to understand viewers’ emotions. The interactive film could then go along different
datasets. For SER tasks, there are generally three types of training datasets, natural, semi
natural, and simulated. The natural datasets are extracted from available videos and
audios, either broadcasted on TV or online. There are also databases from call centers
and similar environments. Semi-natural datasets are made by defining a scenario for
professional voice actors and asking them to play them. The third and most widely used
20
type, the simulated datasets, are similar to semi-naturals. The difference is that the voice
actors are acting the same sentences with different emotions. Traditionally SER used to
follow the steps of automatic speech recognition (ASR), and methods based on HMMS,
GMMs, and SVMs were widespread. Those approaches needed lots of feature
engineering and any changes in the features usually required restructuring the entire
architecture of the method. However, lately, by the development of deep learning tools
and processes, solutions for SER can be changed as well. There is a lot of effort and
neural networks and the use of long short-term memory (LSTM) networks, auto
encoders, and generative adversarial models, there has been a wave of studies on SER
using these techniques to solve the problem. The rest of the paper is organized as
follows: In Section 2, we define SER, and in Section 3, we present some related studies.
where we review several traditional and deep learning methods used in SER. Finally, in
the last chapter, we discuss and conclude our work while proposing direction for future
There are several emotional speech databases that are extensively used in the literature
German, English, Japanese, Spanish, Chinese, Russian, Dutch etc. One main
the speech. Whether they are simulated or they are extracted from real life situations.
The advantage of having a simulated speech is that the researcher has a complete control
over the emotion that it is expressed and complete control over the quality of the audio.
However, the disadvantage is that there is loss in the level of naturalness and
21
speech that is extracted from real life scenarios like call-centers, interviews, meetings,
movies, short videos and similar situations where the naturalness and spontaneity is
kept. The disadvantage is that in these databases there is not a complete control over the
expressed emotions. Also the low quality of the audio can be problem(Gjoreski,Gjoresk,
2014).
In general, there are two approaches in human emotions analysis. In the first approach
the emotions are represented as discrete and distinct recognition classes. The other
emotional distance, level of activeness, level of dominance and level of pleasure are be
observed(Gjoreski,Gjoresk, 2014).
speech. Our approach uses the discrete type of approach; therefore the emotional states
are represented by seven classes: anger, fear, sadness, happiness, boredom, disgust and
neutral. Even though ML approaches have been proposed in the literature, our approach
algorithm parameters optimization. With this analysis, we try to find the optimal ML
configuration of: features, algorithms and parameters, for the task of emotion
22
Figure 2.2: ML approach for emotion recognition.
Speech emotion recognition is a difficult problem for machine learning. The analysis of
Speech is digitized using signal processing methods and then sound characteristics are
obtained through acoustic analysis. However, the overall success rate changes as the
changes in these characteristics differ according to the emotions (sadness, fear, anger,
happiness, neutral, displeasure, etc.). Although different methods are utilized in both
feature extraction and emotion recognition, the success rate varies according to
Sound cannot spread in space, because there are no particles, such as atoms or
molecules, which could lead to the contraction and expansion of a substance. Sound
travel also depends on the temperature of the environment. Sound also creates energy.
The type of energy emerging from a substance's oscillation and vibration is called sound
energy. Sound has physical qualities, such as frequency, wavelength, period, speed of
travel, intensity, loudness, timbre, echo, pressure, amplitude, and resonance. Wavelength
is the distance travelled by sound for the formation of the sound wave. The number of
vibrations of sound in a unit of time (usually seconds) is called “frequency.” Its unit of
measurement is Hertz (Hz). Unlike frequency, the period is the amount of time required
Sound waves are grouped into three categories, based on their frequencies: the human
ear can hear sounds with a frequency of between 20 Hz and 20000 Hz (20 kHz). The
sound waves within the sensitivity thresholds of the human ear are called audible sound
waves. These sounds can be produced in many ways such as using musical instruments
or vocal cords. Sound waves less than 20 Hz are called infrasonic sound waves. For
example, earthquake waves. Sound waves of more than 20000 Hz are called ultrasonic
23
sound waves. Dogs and bats can hear sounds at this frequency. Sound waves are used in
different ways in science and technology. These are Ultrasound devices are used when
imaging internal organs, using ultrasonic (high frequency) waves. Sonars and
echolocation are used in maritime for scanning fish, submarines, and geographical
formations under the sea. Sound waves are also used to break up stones in the kidneys.
Sound intensity is related to how loud or soft a sound is. Sound intensity depends upon
the frequency of the sound wave. We sense sound waves with a low [Link] as soft,
and sound waves with a high frequency, loud. Our perception of the highness of sound
crying sound of babies is loud, which means its frequency is high. On the other hand, an
adult person's voice is soft, which means its frequency is low. The intensity of sound
depends directly on the mechanical pressure on the eardrum sound depends on, and is
directly proportional to, the energy and amplitude of sound waves. As the amplitude of
sound waves increases, the energy and intensity of the sound increase. On the contrary,
as the amplitude of sound waves decrease, the energy and intensity of the sound decline.
The measure of sound intensity is called the sound level, and the unit of measurement of
the sound level is a decibel (dB). Our ear can pick up sounds in between 0 and 140
dB(Sönmez,arol, 2019).
improving. Technology for speech awareness can interact with other disabled persons.
This allows the control of the digital system. Great opportunity in the future to extend
the spoken network of engineering. Enhancing speech recognition can provide improved
services to people with disabilities and provide our system with a secure environment
with voice authentication. We are still well on the way before us because of the high
level of competition on the market between this tech giant and the growing prevalence
that are related to speech, procurement, processing and generation of output and
application of based on the function of the Speech extracted. Here, as vocal cords are
used to produce speech from the mouth, sensors and recording techniques can detect it.
The voice will then be processed using expression techniques and processes. Speech
synthesis then takes place to rearrange the voice bits to a certain frequency to produce a
coherent word that is further used by the product based on the use of speech technology
But when it comes to Speech Emotion recognition, the process is explained in five main
steps and these steps are involved in each and every application related to SER which is
based on machine learning techniques. They are classified as Speech input, Feature
are:
A typical set of emotion contains 300 Emotional States. While the primary emotions are
Anger, Disgust, Fear, Joy, Sadness, and Surprise. In those, Anger, Joy, Sadness and fear
are primary, these emotions are differentiated based on the corresponding changes that
occur in speech rate, pitch, energy and spectrum one of the main speech features that
indicate emotions is energy and study of energy. It depends on the short-term energy and
25
short-term average amplitude. The speech input is taken based on the required format
and then feature extraction is done by MFCC and LPCC methods and selection process
based
Let us see how the classification of Speech Processing is done base on each and every
for many workplaces that integrate application-based computing and telephony through
different platforms. The present device becomes a terminal for the personal remote
workstation with added speech input and output capabilities, allowing access to its
functionality from anywhere depending on its result and outcome. The required speech
is also minimized by the advent of effective and spectral algorithms, very large circuits
with integrated and processors of digital signals and detection of the audio signal based
translate the spoken word to a specific speech message response, while speech
verification is used to check the voice features of the clients. The aim of speech
recognition systems is simply to understand the speaker's spoken word and develop the
In contributions to the other published surveys, this research provides a thorough study
of significant databases and deep learning discrete approaches in SER. The reason for
not focusing on the other older techniques is recent progress in neural networks and,
more specifically, deep learning. Based on the best of our knowledge, this study is the
26
first survey in SER focusing on deep learning along with unified experimental results
2021).
2.7 conclusion
Speech is the basic way of interaction between the listener to the speaker by voice or
expression. Humans can easily understand the speakers' message, but machines can't
Initially, the trained RNN layer-based feature extraction is done to get the speech
3.0 Introduction
This chapter includes the study's presentation as well as the research methodologies
used to complete it. There's also a system description, a system overview, system
First an emotional speech database is used, which consists of simulated and annotated
Then, feature selection method is used for decreasing the number of features and
selecting only the most relevant ones. Finally, the emotion recognition is performed by a
classification algorithm.
The performance of an SER system using machine learning is evaluated using various
metrics such as accuracy, precision, recall, and F1 score. This system is typically trained
27
and tested using speech data, where each speech sample is labeled with the
powerful tool for recognizing and analyzing the emotional content of speech. With the
increasing availability of speech data and advances in machine learning techniques, SER
systems are expected to become more accurate and robust in the future.
The system architecture of a speech emotion recognition system using machine learning
selection, classification, and system evaluation. The success of the system depends on
the choice of features, the machine learning algorithm used, and the quality of the TESS
This system predicts emotion of speech. Our system will have a variety of features,
including:
28
Speech signal sampling is the reduction of a continuous-time signal to a
sequence of "samples".
Is a program that processes its input data to produce output that is used as input
features that can be processed while preserving the information in the original
data set. , many common features are extracted, such as energy, pitch, formant,
Feature Selection is the process of selecting a subset of relevant features for use
Redundant features are those which provide no more information than the
in any context to deal with this issue, we used a method for feature selection.
Feature selection (FS) aims to choose a subset of the relevant features from the
29
Recursive feature elimination (RFE) uses a model (e.g., linear regression or
SVM) to select either the best- or worst-performing feature and then excludes
this feature.
Classification Many machine learning algorithms have been used for discrete
emotion classification. The goal of these algorithms is to learn from the training
samples and then use this learning to classify new observation. In fact, there is
no definitive answer to the choice of the learning algorithm; every technique has
For this reason, here we chose to compare the performance of three different
classifiers:
✓ Training a classifier is where you determine the optimal values (based on your
training content set) to make the trade-offs that make sense to your application.
✓ Test Train Split When machine learning algorithms are used to generate
predictions on data that was not used to train the model, this approach is used to
30
3.4 System requirement
System Requirement Specification Our system will have a lot of integrated requirements,
including hardware and software. The software requirement of our system python &
JupyterNotebook.
components:
✓ Computer Hp
✓ Cori5
✓ Ram8
✓ SSD
✓ Speaker
✓ Microphone
The following components will make up the software component of our system:
✓ python
✓ JupyterNotebook
✓ Python flask
dynamic typing and dynamic binding, make it very attractive for Rapid
31
emphasizes readability and 37 therefore reduces the cost of program
program modularity and code reuse. The Python interpreter and the extensive
standard library are available in source or binary form without charge for all
major platforms, and can be freely distributed. Often, programmers fall in love
programs is easy: a bug or bad input will never cause a segmentation fault.
Instead, when the interpreter discovers an error, it raises an exception. When the
program doesn't catch the exception, the interpreter prints a stack trace. A source
time, and so on. The debugger is written in Python itself, testifying to Python's
introspective power. On the other hand, often the quickest way to debug a
program is to add a few print statements to the source: the fast edit-test-debug
cycle makes this simple approach very effective. Python flask is Flask is a
for developing web applications. Also is a small and lightweight Python web
framework that provides useful tools and features making creating web
wrapper around Werkzeug and Jinja and has become one of the most popular
scientists to create and share documents that include live code, equations, and
32
other multimedia resources. Jupyter notebooks are used for all sorts of data
science tasks such as exploratory data analysis (EDA), data cleaning and
deep learning. Jupyter notebooks are especially useful for "showing the work"
that your data team has done through a combination of code, markdown, links,
and images. They are easy to use and can be run cell by cell to better understand
what the code does. Visual Studio Code comes includes an incredibly quick
source code editor that is perfect for daily use. With support for hundreds of
4.1 Introduction
This chapter focuses on the analysis and design of our research on emotion-based
speech recognition using machine learning, which is one of the key forms of
communication among humans. In the system analysis, we will address the existing
method of speech emotion, which is the manual method, the disadvantages of the
current method, and the necessity for this system, as well as the requirements of our
system. The system requirements can be divided into two categories. Functional and
33
our system that are unique to the system; non-functional requirements discuss the
System analysis involves the study of machine learning methods of speech emotion
prediction and the most important face facts is to collect dataset after teaching the
models by removing any Data cleaning as mentioned Chapter three after those models
have been tested and any data cleaning will be done, the interface will be streamlined
various emotions conveyed in speech, some current systems have limited integration
conveyed in speech.
different extensions or may have restrictions on the types of audio files they
can process.
predictions.
➢ Inadequate Training Data: Many existing systems may not have access to a
model.
34
4.4 The Need of This Work
An SER is an important for various reasons. This system can increase and improve the
communication. SER can interact with other disabled persons. also It can recognize the
driver’s state of mind and help prevent accidents and disasters. Another related
To deal with the problem, we developed automatic emotion prediction using machine
learning techniques, that takes all voice extensions like mp3 or .wav formats and returns
a result that shows by emoji’s with text of the voice. So, the users can analyze and
understand our system. This system is designed to prevent the approval of emotion
prediction. Our Speech Emotion Recognition system aims to provide an improved and
reliable solution for accurately recognizing and analysing emotions in speech. Our
system is built using Flask, a web framework, allowing seamless integration into
In this section, researchers discuss requirements for the speech emotion System and
classify the requirement into two parts Functional and Non-functional requirements
system. The following are the requirements that are proposed system should
functionally attain:
❖ Audio Input: The system should be able to accept audio inputs of various
35
❖ Emotion Prediction: The system should accurately predict one of the seven
emotions (Angry, Happy, Sad, Surprise, Disgust, Fear, and Normal) based on
learning model using the provided dataset and evaluating its performance. It
should be able to update and retrain the model as new data becomes available.
❖ Real-Time Processing: The system should process the input audio and provide
❖ User Interface: The system should have a user-friendly interface where users
can interact with the system, upload audio files, and receive emotion
predictions as output.
The following are the requirements that the proposed system should non-functionally
achieve:
in emotion recognition.
36
❖ Performance: The system should be able to handle multiple concurrent
prediction.
System design is the process of creating the architecture, modules, and components of a
system, as well as the various interfaces between those components and the data that
passes through it. The aim of system design is to make a new system as good as the old
one. Use case diagram, Flow chart diagram or UML are used and analyzed to display
37
Figure 4.1 follow chart diagram
For SER tasks, there are generally three types of training datasets, natural, semi-natural,
and simulated. The natural datasets are extracted from available videos and audios,
either broadcasted on TV or online. There are also databases from call centers and
similar environments. The performance and robustness of the recognition systems will
essential to have sufficient and suitable phrases in the database to train the emotion
recognition system and subsequently evaluate its performance. There are three main
emotions. In this work, we used an acted emotion databases because they contain strong
emotional expressions. The literature on speech emotion recognition shows that the
majority of studies have been conducted with emotional acted speech. In this section we
detailed the emotional speech database used for classifying discrete emotions in our
experiments on emotion recognition with raw audio, we built a dataset from Toronto
Emotional Speech Set (TESS). Is an acted dataset primarily developed for analysing the
effect of age on the ability to recognize emotions this dataset is all comprised of two
female actors, about 60 and 20 years old? Each actor has simulated seven emotions for
200 neutral sentences. Emotions in this dataset are: angry, pleasantly surprised,
disgusted, happy, sad, fearful, and neutral. To label the dataset, 56 undergraduate
students were asked to identify emotions from the sentences. After the identification
task, the sentences with over 66% confidence have been selected to be in the dataset.
38
CHAPTER V: IMPLEMENTATION AND TESTING
5.1 introduction
In this chapter, we discuss the details of our proposed SER system's implementation and
testing. This chapter also includes Snapshots of the project's key components as well as
descriptions of how the system operates. This project explains how to build a SER using
a neural networking (NN) model that continuously predicts emotions. We also present
the results of our experiments and analyze the performance of the proposed system.
To implement our proposed SER system, we have chosen to use two different
environments Kaggle and jupyter notebooks and python flask. These platforms provide
a web-based interface for running Jupyter notebooks and executing code, with access to
hardware resources such as GPUs and TPUs, which are essential for training deep
39
Kaggle is a platform for data science competitions and collaborative data science
notebooks, with access to GPUs and TPUs. In addition to the computing resources,
Kaggle also provides access to a large community of data scientists and machine
hours at a time, making it an excellent choice for experimenting with deep learning
models without the need for expensive hardware. Kaggle allows users to collaborate
In our implementation of the SER system, we used Kaggle to train and evaluate our
model. We leveraged the GPU and TPU resources available on this platform to
We collecting Audio data from publicly available and combining it with self-made audio
samples. The publicly available dataset is taken from kaggle included 2800 audios
appraisal record; this data is in .wav file format. System will read this file using
Important libraries
40
Figure 5.1 libraries
the model.
To pre-process the audio data for our SER system, we converted the audio recordings to
Wave show and spectrograms using a Librosa library .we then extracted the data and
used Mel-frequency cepstral coefficients (MFCC) It extract 2800 audio files and 40
To train our SER model, we split the pre-processed dataset into batches for training and
validation. We implemented batch training and validation sets to improve the training
We trained the model for 50 epochs with a learning rate of 0.2 and a batch size of 64.
We split the data using 70/30 ration meaning 70% of the data is used to train the model,
41
and the remaining part, 30%, is used to test the model. We trained the model, using
model selection, with the data samples from the training set.
42
10Figure 5.5 train loss and Val loss
After we have trained our SER model, we will deploy it in a production environment
where it will be used for real-time speech recognition. To accomplish this, we used
python Flask, an open-source platform for deploying and managing machine learning
models.
used for developing web applications. We used Flask to create a RESTful API that can
receive audio files and return the corresponding emotions using our SER model.
43
To connect the Flask server to a React client, we created a frontend web application
using html/CSS, and python for backend the client sends HTTP requests to the Flask
server through the REST API, and displays the transcriptions on the user interface.
We took snapshots some of the most significant parts of the Administrators &
Merchants, and these Snapshots will help system users to understand it effectively.
The first page that shows when the portal loads is the Login Page, where the user must
enter a valid username and password to determine, the login form is used to validate the
44
5.7.2 Dashboard page
After a user uploads or selects an audio file, they will be redirected to the page shown in
Figure 5.8. This page provides the user with several options, including the ability to
listen to the audio, to see audio as a wave, to download the audio, and to change
playback speed.
Additionally, the page also includes a section where the user can make a prediction to
Show the result of prediction page where the users can view emotions of the audio
45
13 Figure 5.8 result page
6.1 Introduction
In this chapter, we will review several studies that have addressed the challenge of
developing SER systems. We will discuss the findings of these studies, including the
46
data augmentation, and deep learning methods in improving SER performance for these
languages. We will discuss the findings of these studies, including the importance of
language modelling, data augmentation, and deep learning methods in improving SER
performance for these languages. Finally, we will discuss the challenges that remain in
developing SER systems for complex emotions and the need for further research to
6.2 Results
The results of our study indicate that the development of an SER system is feasible and
holds great potential for improving the accessibility of emotion recognition technology
and it is recreational with good accuracy. Through our analysis of an audio data in the
form of numerical, we were able to identify the emotion of different audio samples and
that are essential for the accurate recognition of emotions from the speech.
(TESS). Is an acted dataset primarily developed for analysing the effect of age on the
ability to recognize emotions this dataset is all comprised of two female actors, about 60
and 20 years old? Each actor has simulated seven emotions for 200 neutral sentences.
Emotions in this dataset are: angry, pleasantly surprised, disgusted, happy, sad, fearful,
and neutral. In this section, experimentation results are presented and discussed. We
structure used is a simple LSTM. It consists of two consecutive LSTM with 256 layers
with dense 128 activation that gets 40 features as input, followed by two classification
dense layers. Features from data are scaled to [40, 1] before applying classifiers. Scaling
47
unscaled data, it is possible for large inputs to slow down the learning and convergence
and in some cases prevent the used classifier from effectively learning for the
classification problem.
As a baseline SER system, we used a RNNS classifier with long short-term memory
(LSTM) its relative insensitivity to gap length is its advantage over other RNNs, hidden
Markov models and other sequence learning methods. It aims to provide a short-term
memory for RNN It is applicable to classification, processing and predicting data based on
Our findings suggest that the proposed deep learning model with RNNs, holds great
6.3 limitations
Our study has several limitations that should be considered when interpreting the results,
there are many advancements on speech emotion recognition systems, and there are still
several obstacles that need to be removed for successful recognition. One of the most
important limitations is the generation of the Somali dataset that is used for the learning
process. Most of the data sets used for SER are acted or elicited that are recorded in
special silent rooms. However, the real-life data is noisy and has far more different
characteristics than the others. Although natural data sets are also available, they are
fewer in numbers. There are legal and ethical problems to record and use natural
emotions. Most of the utterances in natural data sets are taken from talk-shows, call-
center recordings, and similar cases where the involved parties are informed of the
recording. Somali data sets do not contain all emotions and may not reflect the emotions
that are felt. In addition, there are problems during the labelling of the utterances. There
are human annotators labelling the speech data after the utterances are recorded, and
there is no annotators in Somalia. The actual emotion felt by the speaker and emotions
48
perceived by human annotators may show differences. Even the recognition rates of
human annotator are not over 90%. In Favor of humans, however, we believe that we
also depend on the content and the context of the speech as we are evaluating. There are
also cultural and language effects on SER. There are several studies available working
on cross-language SER.
However, the results show that current systems and features used are not sufficient for it.
The intonation of emotions on speech among various languages may show differences
for example. An overlooked challenge is the case of multiple speech signals, where the
SER system has to decide which signal to focus on. Although it can be handled via a
speech separation algorithm in the preprocessing stage, current systems fail to notice
this problem.
In most modern SER systems, semi-natural and simulated datasets are utilized that are
49
7.1 introduction
The goal of this chapter is to provide a clear and succinct summary of the research's
findings, which were reached after six months of in-depth inquiry and analysis. The
research's conclusion and an analysis of the extent to which the research objectives
outlined in chapter one were met. Furthermore, the chapter offers useful guidelines and
the future, as well as suggestions for areas of further analysis and exploration.
7.2 Conclusion
Our study, we presented speech emotion recognition (SER) system using machine
learning algorithms (RNN) to classify seven emotions. Thus, two types of features
(MFCC and MS) were extracted from acted databases (TESS databases), and a
combination of these features was presented. In fact, we study how classifiers and
information is not always good in machine learning applications. The machine learning
models were trained and evaluated to recognize emotional states from these features.
SER reported the best recognition accuracy rate of 98% is achieved by RNN classifier
without SN and with FS, on the TESS database using RNN classifier. Overall, our
research project highlights the potential of deep learning-based SER models, providing
It is recommended to support the used dataset with a greater number of algorithms to get
high accuracy for the predictive model the study speech emotion recognition to improve
50
the current algorithms it can be concluded that. There are many factors that need to be
SER such as MLR, SVM, means algorithm can be applied to emotion prediction.
REFERENCES
Abbaschian, B. J., Sierra-Sosa, D., & Elmaghraby, A. S. (2021). Deep Learning Techniques for
Akçay, M. B., & Oğuz, K. (2020). Speech emotion recognition: Emotional models,
51
Basu, S., Chakraborty, J., Bag, A., & Aftabuddin, M. (2017). A review on emotion
Bhavan, A., Chauhan, P., Hitkul, & Shah, R. R. (2019). Bagged support vector machines
[Link]
Chen, L., Mao, X., Xue, Y., & Cheng, L. L. (2012). Speech emotion recognition:
Gjoreski,Gjoresk, M., Hristijan i. (2014). Machine Learning Approach for Emotion Recognition
in Speec.
Hasan, M. R., Jamil, M., & Rahman, M. (2004). Speaker identification using mel
Hu, Q., He, Z., Zhang, Z., & Zi, Y. (2007). Fault diagnosis of rotating machinery based
on improved wavelet package transform and SVMs ensemble. Mechanical Systems and
Hwan , Yung-Hwan, P., Kim,Oh. (2009). Feature vector classification based speech emotion
Kerkeni, L., Serrestou, Y., Mbarki, M., Raoof, K., Ali Mahjoub, M., & Cleder, C.
(2020).
[Link]
52
Kim, H.-C., Pang, S., Je, H.-M., Kim, D., & Bang, S.-Y. (2002). Support Vector
Machine
Ensemble with Bagging. In S.-W. Lee & A. Verri (Eds.), Pattern Recognition with
Support Vector Machines (Vol. 2388, pp. 397–408). Springer Berlin Heidelberg.
[Link]
Kumar, Dr. T., Villalba-Condori, K. O., Arias-Chavez, D., K., R., M, K. C., & S., Dr. S.
Liu, Z.-T., Wu, M., Cao, W.-H., Mao, J.-W., Xu, J.-P., & Tan, G.-Z. (2018). Speech
emotion recognition based on feature selection and extreme learning machine decision
292.
Morrison, D., Wang, R., & De Silva, L. C. (2007). Ensemble methods for spoken
S., Jason, C. A., & [Link] Student, Sreyas Institute of Engineering and Technology,
53
Vlasenko, B., Philippou-Hübner, D., Prylipko, D., Böck, R., Siegert, I., & Wendemuth,
emotions. 1–6.
Wang, G., Hao, J., Ma, J., & Jiang, H. (2011). A comparative assessment of ensemble
learning for credit scoring. Expert Systems with Applications, 38(1), 223–230.
[Link]
Zareapoor, M., & Shamsolmoali, P. (2015). Application of Credit Card Fraud Detection:
[Link]
Zhou, N.-R., Li, J.-F., Yu, Z.-B., Gong, L.-H., & Farouk, A. (2016). New quantum
Bhavan, A., Chauhan,P. ,Shah,R. R. (2019). Bagged support vector machines for emotion
Recognition in Speech.
Kerkeni, L., Serrestou,Y. ,Mbarki,M. ,A. ,Cleder,C. (2019). Automatic Speech Emotion
54
import pandas as pd import numpy as np from tensorflow import
Model model =
[Link].load_model('./model/speech_model.h5')
mfcc
write_to_file(username, password):
{password}\n')
55
# to check if user already exist def
check_credentials(username, password):
return True
return False
@[Link]("/home") def
home():
return render_template("[Link]")
@[Link]("/about")
def about():
return render_template("[Link]") #
username = [Link]['username']
password = [Link]['password']
56
write_to_file(username, password)
render_template('[Link]') # Login
page rendering
login():
if [Link] == 'POST':
username = [Link]['username']
password = [Link]['password'] if
check_credentials(username, password):
return redirect(url_for('home'))
else:
return render_template('[Link]')
[Link]({'speech':paths}) pred_X_mfcc =
57
pred_df['speech'].apply(lambda x: extract_mfcc(x)) pred_X = [x for x in
[Link](debug=True,port=5001)
Appendix C: Prediction
filenames: [Link]([Link](dirname,
label = [Link]('.')[0]
[Link]([Link]()) if len(paths) ==
2800:
break
print('Dataset is Loaded')
## Create a dataframe df =
[Link]() df['speech']
58
= paths df['label'] = labels
[Link]() df = df[df['label'] !
= 'disgust']
df['label'].value_counts()
labels =
[Link](df['label'].value_cou
path = [Link](df['speech']
[df['label']==emotion])[0]
data, sampling_rate =
[Link](path) #display
[Link](da
of audio file
spectogram(data,
sampling_rate, emotion)
Audio(path)
Audio(path)
extract_mfcc(filename):
mfcc
X = [x for x in X_mfcc]
X = [Link](X)
[Link]
X = np.expand_dims(X, -1)
[Link]
60
# Mapped label into numbers example (sad -> 0 , happy -> 1)
OneHotEncoder()
y = enc.fit_transform(df[['label']]) y
= [Link]()
[Link]
# show the label mapped values label_index_mapping = {i: label for i, label
= [Link](y_test, axis=1)
accuracy_score(y_test_labels, y_pred_labels)
print(accuracy*100,"% accuracy")
61
pred_X_mfcc = pred_df['speech'].apply(lambda x: extract_mfcc(x))
62
Machine learning enables the development of SER systems by allowing algorithms to self-learn from training data, improving emotion recognition over time without explicit programming. This approach helps in processing and classifying complex speech signals accurately .
Feature extraction in SER systems improves performance by isolating critical aspects of speech signals that are indicative of emotional states. This step often uses advanced algorithms to emphasize features that contribute most to accurate emotion classification .
In healthcare, SER systems can assist people with autism by helping them understand and adjust their emotional expressions and social behaviors. They can also support therapists in assessing patients' emotional states, providing insights into hidden emotions .
Deep learning methods offer superior results due to their ability to learn complex patterns from large datasets, which are particularly beneficial in speech-related applications. These methods outperform earlier techniques by providing more accurate and reliable emotion classification from speech signals .
Current SER systems face challenges like limited emotion classification accuracy, lack of flexibility in handling various audio formats, latency in real-time processing, and inadequate training data diversity, which all impact their performance and accessibility .
Developers must ensure SER systems avoid biases and ensure fairness in emotion recognition. This involves careful dataset selection and model training processes to prevent discrimination and inaccuracies .
SER systems improve driving safety by monitoring the emotional state of drivers, recognizing states like stress or anger, and providing feedback to promote safer driving behaviors and reduce accidents .
Jupyter Notebooks are used for tasks such as exploratory data analysis, data cleaning, transformation, and for executing and testing machine learning models during the development of SER systems. This enables easy visualization and iterative testing .
SER systems enhance human-computer interaction by making systems more aware, responsive, and adaptive to users' emotional contexts. This is crucial in applications like virtual assistants, customer service, and therapy bots, where emotional awareness can improve user experience and interaction effectiveness .
A typical SER system architecture involves stages like signal pre-processing, feature extraction, feature selection, and classification. These components work together to process speech data, extract relevant features, and classify emotions accurately .