Baby Chillanto Database for Asphyxia Cry
Baby Chillanto Database for Asphyxia Cry
Classification of asphyxia infant cry using hybrid speech features and deep
learning models
Hua-Nong Ting a, b, c, *, Yao-Mun Choo d, Azanna Ahmad Kamar d
a
Department of Biomedical Engineering, Faculty of Engineering, Universiti Malaya, Jalan Pantai Baharu, 50603 Kuala Lumpur, Malaysia
b
Centre for Image and Signal Processing, Faculty of Engineering, Universiti Malaya, Jalan Pantai Baharu, 50603 Kuala Lumpur, Malaysia
c
Faculty of Medical Engineering, Jining Medical University, University Park, National High-tech Zone, 272067 Jining City, Shandong Province, China
d
Neonatal Intensive Care Unit, Department of Paediatrics, Faculty of Medicine, Universiti Malaya, Jalan Pantai Baharu, 50603 Kuala Lumpur, Malaysia
A R T I C L E I N F O A B S T R A C T
Keywords: Single speech feature such as Mel-Frequency Cepstral Coefficient (MFCC) has been used in most of the studies to
Asphyxia classify asphyxia cry among infants. Other speech features such as Chromagram, Mel-scaled Spectrogram,
Infant cry Spectral Contrast and Tonnetz have not been reported in any study related to the classification of asphyxia cry.
Hybrid features
The study investigated the use of hybrid features of MFCC, Chromagram, Mel-scaled Spectrogram, Spectral
Deep Neural Network
Convolutional Neural Network
Contrast and Tonnetz and deep learning models in classifying asphyxia cry. Deep learning models such as Deep
Neural Network (DNN) and Convolutional Neural Network (CNN) were used to classify infant cry between
normal/non-asphyxia and asphyxia. The performance of the deep learning models was compared using
concatenated hybrid features and single feature of MFCC. The Baby Chillanto Database was used in this study.
CNN model performed better than DNN models when MFCC was used. DNN models performed better with hybrid
features compared to that with single feature of MFCC. DNN with multiple hidden layers achieved an accuracy of
100% in classifying normal and asphyxia cry, and 99.96% for non-asphyxia and asphyxia cry when the hybrid
features were used.
* Corresponding author at: Department of Biomedical Engineering, Faculty of Engineering, Universiti Malaya, Jalan Pantai Baharu, 50603 Kuala Lumpur, Malaysia.
E-mail address: tinghn@[Link] (H.-N. Ting).
[Link]
Received 6 July 2021; Received in revised form 23 May 2022; Accepted 3 July 2022
Available online 7 July 2022
0957-4174/© 2022 Elsevier Ltd. All rights reserved.
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
MFCC. presents and discusses the results. Finally, the study is concluded in
Several classification models have been proposed for classifying the Section 5.
cries of babies. The Artificial Neural Network (ANN) classifier is the
most popular model. Among the ANN models used by the researchers are 2. Related works
TDNN (Hariharan, Saraswathy, et al., 2012; Reyes-Galaviz et al., 2004),
Multilayer Perceptron (Zabidi et al., 2010), and the Probabilitic Neural Recent studies on the classification of asphyxia cry using the Baby
Network (PNN)(Hariharan, Saraswathy, et al., 2012; Hariharan et al., Chillanto Database are summarized in Table 1. The Chillanto database
2011). Machine learning model such as Support Vector Machine (SVM) consists of 507 normal cry samples and 340 asphyxia cry samples. Most
is used by some researchers (Badreldine et al., 2018; Ji et al., 2021; of the studies use parts of the database (Badreldine et al., 2018; Har
Sahak et al., 2018; Sahak et al., 2010; Salehian Matikolaie & Tadj, iharan, Saraswathy, et al., 2012; Hariharan et al., 2011; Sahak et al.,
2020). Recently, deep learning models are used in the classification of 2018; Sahak et al., 2010; Zabidi et al., 2010; Zabidi et al., 2018), and
asphyxia cry. Examples are Deep Neural Network (DNN)(Ji et al., 2019; only five studies include all the cry samples of Baby Chillanto Database
Lahmiri et al., 2022; Sachin et al., 2017), GoogleNet and AlexNet in their experiments (Hariharan et al., 2018; Ji et al., 2019; Moharir
(Moharir et al., 2017), and Convolutional Neural Network (CNN) (Ji et al., 2017; Rosales-Pérez et al., 2015; Sachin et al., 2017). Ji et al.
et al., 2021; Lahmiri et al., 2022; Sahak et al., 2018). However, the (2019) classified the infant cry sounds into asphyxiated and non-
performance of the deep learning models is lower than that of the pre asphyxiated cry. The non-asphyxiated cry group included the cry sam
vious studies using ANN models. Although the Recurrent Neural ples of normal, deaf, hungry, and pain, and the total number of non-
Network (RNN) and Long Short-Term Memory (LSTM) are good at asphyxiated cry is 1927. Hariharan et al. (2018) added an extra of 45
classifying time series data or sequence data, their performance is lower self-collected normal cry samples to make the total number of the
than CNN for automatic cry event detection and for the diagnosis of normal cry samples 552.
infant cry (Cohen et al., 2020; Cohen, 2020; Lahmiri et al., 2022). MFCC is the most commonly used feature extraction technique for
Our contributions are as follows: (1) We propose the use of hybrid classification and the most used numbers of MFCC coefficients which are
speech features, comprising of concatenated features of MFCC, Chro 12 and 16 (Badreldine et al., 2018; Hariharan et al., 2018; Ji et al., 2019;
magram, Mel-scaled Spectrogram, Spectral Contrast, and Tonnetz in Rosales-Pérez et al., 2015; Sahak et al., 2018; Sahak et al., 2010; Zabidi
classifying asphyxia cry. These features have not been used in any study et al., 2010; Zabidi et al., 2018). Other feature extraction techniques
related to the classification of asphyxia cry. The hybrid speech features include Wavelet Packet Transform (Hariharan et al., 2018; Hariharan
contain rich and extra information, which MFCC does not have. (2) et al., 2011), LPC (Hariharan et al., 2018; Reyes-Galaviz et al., 2004;
Performance is compared between different deep learning models with Rosales-Pérez et al., 2015), Short Time Fourier Transform (STFT)
MFCC and hybrid speech features for asphyxia cry recognition. (3) Our (Hariharan, Sindhu, et al., 2012), and prosodic features such pitch, in
proposed method achieves the state-of-the-art accuracy and outperforms tensity, fundamental frequency and formants (Ji et al., 2019). Most of
various existing methods. Our proposed method improves the accuracy the studies use a single feature such as MFCC for classification. Some
of asphyxia cry recognition without using a feature selection algorithm. studies explored the use of multiple speech features to improve the ac
Asphyxia infant cry classification has the potential for use as a surrogate curacy of classification (Hariharan et al., 2018; Ji et al., 2019). Some
clinical parameter in the evaluation of intrauterine hypoxia. The iden studies use feature selection techniques to select the optimal features for
tification of asphyxia infant cry may possibly be used to monitor infants classification. Sahak et al. (2018) used Principal Component Analysis
at risk, for example, infants with mild hypoxia who do not require (PCA) and Orthogonal Least Square (OLS) in selecting the optimal fea
extensive resuscitation, or those with perinatal risk factors such as tures. Hariharan et al. (2018) selected the most salient speech features
growth restriction and placental anomalies. using particle swarm optimization, binary version of dragonfly optimi
The rest of this paper is organized as follows: Section 2 explains zation algorithm, and improved binary dragonfly optimization algo
current related studies in infant crying analysis of asphyxia cry. Section rithm. Hariharan et al. (2011) used PCA to reduce the high dimension of
3 presents the materials and methods used in the study. Section 4 the features and select the non-redundant features based on the Eigen
Table 1
Summary of the works on normal-asphyxia classification using Baby Chillanto Database.
Reference Data (samples) Feature Classifier Accuracy (%) Cross-Validation
(Ji et al., 2019) Normal (1927) 12 MFCC with weighted prosodics DNN 96.74 5-fold
Asphyxia (340)
(Badreldine et al., 2018) Normal (340) DWT, 6–12 MFCC RBR Kernel SVM 98.5 10-fold
Asphyxia (340)
(Zabidi et al., 2018) Normal (316) 12 MFCC CNN 92.78 None
Asphyxia (284) (70%/30%)
(Sahak et al., 2018) Normal (316) 10–20 MFCC SVM 94.84 10-fold
Asphyxia (284)
(Hariharan et al., 2018) Normal (552) 16 MFCC, 16 LPC, Wavelet Packet Transform ELM 100 None/NA
Asphyxia (340)
(Moharir et al., 2017) Normal (507) Waveform GoogleNet, AlexNet 94 None
Asphyxia (340) (75%/25%)
(Sachin et al., 2017) Normal (1049) Waveform DNN 92 None
Asphyxia (340) (75%/25%)
(Rosales-Pérez et al., 2015) Normal (507) 16 MFCC, LPC Fuzzy model 90.68 10-fold
Asphyxia (340)
(Hariharan, Saraswathy, et al., 2012) Normal (340) STFT PNN, GRNN, TDNN, MLP 99.19 10-fold
Asphyxia (340)
(Hariharan et al., 2011) Normal (340) Wavelet Packet Transform PNN 99.41 None
Asphyxia (340) (60%/40%)
(Sahak et al., 2010) Normal (169) 16 MFCC SVM 95.86 5-fold
Asphyxia (169)
(Zabidi et al., 2010) Normal (316) 20–40 MFCC MLP 93.38 None
Asphyxia (284) (60%/20%/20%
2
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
3
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
MCMCT coefficients per cry sample. The mean coefficients were ob up to 7 layers. ReLU activation was used at each of the outputs of the
tained by averaging the coefficient over all the signal frames, which was hidden or dense layer. The output layer has two neurons, which corre
the entire 1-second of the cry samples. 49, 61, and 73 coefficients of spond to the normal/non-asphyxia and asphyxia cry. Softmax activation
MCMCT were extracted per cry sample. 49 MCMCT coefficients con was applied at the output layer. Kernel regularizer (L2) of 0.001was
sisted of 12 MFCCs, 12 Chromagrams coefficients, 12 Mel-Spectrogram applied at all the hidden (dense) layers. Adaptive Moment Estimation
coefficients, 7 Spectral Contrast coefficients, and 6 Tonnetz coefficients. (Adam) was used as the optimizer. Mean squared error was used as the
61 MCMCT coefficients contained 16 MFCCs, 16 Chromagrams co loss function. The number of training epochs was set at 200. The ar
efficients, 16 Mel-Spectrogram coefficients, 7 Spectral Contrast co chitecture of the DNN2 model is shown in Fig. 2. The DNN2 model was
efficients, and 6 Tonnetz coefficients. 73 MCMCT coefficients were trained at different hidden 1 neuron numbers, which were set at 50, 100,
comprised of 20 MFCCs, 20 Chromagrams coefficients, 20 Mel- 200, 500, 1000, and 2000 neurons. The hidden neuron number for the
Spectrogram coefficients, 7 Spectral Contrast coefficients, and 6 Ton 2nd layer until the 7th layers were set at 256, 128, 64, 32, 16, and 8
netz coefficients. The description of the spectral representation is sum neurons, respectively. The DNN2 model was tested on Dataset 1 and
marized in Table 3. Dataset 2 using Feature 1, Feature 2, and Feature 3. The training was
conducted using 5-fold cross-validation. For each fold, accuracy was
3.3. Classifier computed after every epoch. The maximum accuracy for each fold was
used to calculate the overall accuracy.
3.3.1. Baseline model
Deep Neural Network with one hidden (dense) layer (DNN1) was 3.3.3. Convolutional Neural Network
chosen as the baseline model due to its simplicity. The DNN1 model’s Convolutional Neural Network (CNN) was used to classify the cry
architecture is shown in Fig. 1. The DNN1 model was trained at different samples. CNN was chosen instead of RNN and LSTM because it had a
hidden neuron numbers (50, 100, 200, 500, 1000, and 2000) with higher recognition rate in the recognizing cry sounds. Since the MFCCs
different signal frame lengths (20 ms, 30 ms, 40 ms, and 50 ms) and were extracted frame by frame, N frames of MFCCs were extracted over
MFCCs (12, 16, and 20). ReLU activation was used at the output of the the entire 1-second of the cry samples. N depends on the signal frame
hidden or dense layer. The output layer has two neurons, which corre length. A total of 101, 67, 51, and 41 frames of MFCCs were extracted for
spond to the normal/non-asphyxia and asphyxia cry. Softmax activation signal frame lengths of 20 ms, 30 ms, 40 ms, and 50 ms respectively with
was applied at the output layer. Kernel regularizer (L2) of 0.001was a hop length of 50%. These N frames of MFCCs formed a 2-D image,
applied at all the hidden (dense) layers. Adaptive Moment Estimation which was fed into the CNN. Four convolutional layers were used. Filter
(Adam) was used as the optimizer. Mean squared error was used as the size of 32 was used at each convolutional layer. The kernel size of 3 × 3
loss function. The number of training epochs was set at 200. The DNN1 and stride of 1 were used at each convolutional layer. ReLU activation
model was tested on Dataset 1 and Dataset 2 using Feature 1, Feature 2, was used at each output of the convolutional layer and the two fully
and Feature 3. In Feature 1, the MFCCs of the N frames were concate connected dense layers. No Max-pooling and dropout were applied be
nated in series to form one dimensional MFCCs before they were fed into tween the convolution layers. The output of the 4th convolutional layer
the DNN1 model. In Feature 2, MFCCs were averaged over all the signal was flattened before feeding into two fully connected dense layers with
frames to form the Mean MFCC. Feature 3 was the concatenated mean neurons of 256 and 128 respectively. The output layer has two neurons,
coefficients of MFCCs, Chromagram, Mel-scaled Spectrogram, Spectral which correspond to normal/non-asphyxia and asphyxia cry. Softmax
Contrast, and Tonnetz in one dimension. Since the data size is small, 5- activation was applied at the output layer. Kernel regularizer (L2) of
fold cross-validation was implemented, where the database (total cry 0.001was applied at the convolutional layers and fully connected layers.
samples) was split randomly into 5 equal sets. The testing was conducted Adaptive Moment Estimation (Adam) was used as the optimizer. Mean
on one set of cry samples while the training was carried out using the squared error was used as the loss function. The CNN architecture is
remaining 4 sets of cry samples. The process was repeated 5 times. By shown in Fig. 3. The number of training epochs was set at 200. The CNN
implementing 5-fold cross-validation, the number of the testing data is was tested on Dataset 1 and Dataset 2 using Feature 1 only. The training
the total number of cry samples in the database. For each fold, accuracy was conducted using 5-fold cross-validation. For each fold, accuracy was
was computed after every epoch. The maximum accuracy for each fold computed after every epoch. The maximum accuracy for each fold was
was used to calculate the overall accuracy. used to calculate the overall accuracy.
3.3.2. Deep Neural Network with two and more hidden layers 3.4. Experiments and performance criteria
Once the training and testing of the DNN1 model were completed,
the optimal signal frame lengths and MFCC numbers were used to train The algorithms for speech feature extraction and deep learning
and test the Deep Neural Network with two and more hidden layers models (DNN1, DNN2, and CNN) were implemented under Jupyter
(DNN2 model) based on the two highest accuracies achieved by the Notebook (Version 6.1.6) using Python (Version 3.8.5). Librosa audio
DNN1 model. The DNN2 models were trained with hidden (dense) layers library was used to extract the speech features (Librosa, 2021; McFee
Table 3
Spectral representation in terms of MFCC, Chromagram, mel-Spectrogram, Spectral Contrast and Tonnetz.
Feature Item Number of coefficients per signal frame Signal frame length (Total frame number per cry sample) Number of coefficients
per cry sample
4
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
Fig. 1. DNN1 architecture and spectral features, where MFCC = 12, 16, 20; N = 101, 67, 51, 41; MCMCT = 49, 61, 73.
Fig. 2. DNN2 architecture and spectral features, where MFCC = 12, 16, 20; N = 101, 67, 51; 41; MCMCT = 49, 61, 73.
Fig. 3. CNN architecture and spectral features where MFCC = 12, 16, 20; N = 101, 67, 51, 41.
et al., 2015). Tensorflow Keras 2 (Keras, 2021) was used to build the length of 50 ms was used. DNN2 model achieved the highest accuracy of
DNN1, DNN2, and CNN models. The experiments were run using a 96.69% at a signal frame length of 50 ms. As for the CNN model, the
desktop computer with Intel Core i3-10100 CPU@3.60 GHz and 16 GB highest accuracy was achieved at 98.34% when a signal frame length of
RAM. Experiments were conducted to determine the optimal perfor 40 ms was used. CNN model outperformed the DNN1 model and DNN2
mance of DNN1, DNN2, and CNN using Dataset 1 and Dataset 2. Three model while the DNN2 model performed better than the DNN1 model.
different features were used: Feature 1 (MFCC × N Frames), Mean The DNN1 model performed best with 12 and 20 MFCCs while the DNN2
MFCC, and MCMCT. The models were trained with different signal model worked best with 12 MFCCs. CNN model performed best with 16
frame lengths (20 ms, 30 ms, 40 ms, and 50 ms), MFCCs (12, 16, and 20) MFCCs. DNN2 model performed the best when 5 hidden layers were
and MCMCT (49, 61, and 73), hidden neuron numbers (50, 100, 200, used.
500, 1000, and 2000) and hidden layer numbers (up to 7 layers). All the The performance of the classifiers using Feature 2 is shown in
models were trained and tested with 5-fold cross-validation. Accuracy Table 5. DNN1 model achieved the highest accuracy of 96.69% when a
was used as the performance metric. For each fold, accuracy was signal frame length of 50 ms was used. DNN2 model performed better
computed after every epoch. The maximum accuracy for each fold was than the DNN1 model with the highest accuracy achieved at 98.11%.
used to calculate the overall accuracy. The workflow of the proposed DNN2 model was able to achieve the highest accuracy when 4 to 6
method is shown in Fig. 4. hidden layers were used.
The performance of classifiers using Feature 3 is shown in Table 6.
4. Results and discussion DNN1 model achieved the highest accuracy of 99.76% when a signal
frame length of 50 ms was used. DNN2 model outperformed the DNN1
4.1. Performance of deep learning models using Dataset 1 model with the highest accuracy achieved at 100%. DNN2 model was
able to achieve the highest accuracy when 4 to 6 hidden layers were
The performance of the deep learning models using Dataset 1 and used.
Feature 1 is shown in Table 4. Generally, a longer signal frame length In the classification of Dataset 1, the CNN model performed better
was needed to achieve optimal performance of the classifiers. DNN1 than the DNN1 model and DNN2 model while the DNN2 model per
model achieved the highest accuracy of 96.45% when a signal frame formed better than the DNN1 model when Feature 1, Feature 2, and
5
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
Feature 3 were used. However, in terms of computing complexity, CNN Database. Studies by Rosales-Pérez et al. (2015) and Moharir et al.
required a much longer time for training and validation compared to (2017) used the same number of normal and asphyxia cry samples for
DNN1 and DNN2. Training of DNN2 was longer than DNN1 since DNN2 the classification and the highest accuracy was achieved at 94%. Har
had more hidden layers. Classifiers performed better with Feature 3 iharan et al. (2011) used parts of the Baby Chillanto Database and
compared to Feature 1 and Feature 2. DNN2 model was able to fully achieved an accuracy of 99.41%. However, the study did not implement
classify all the cry samples correctly using Feature 3. No feature selec cross-validation and the database was simply split into 60% of the
tion was implemented in this study. The speech features were extracted training set and 40% of the testing set. Hariharan et al. (2018) added an
and fed directly into the classifiers. In terms of computing speed, DNN2 extra of 45 self-collected normal cry samples and the study achieved an
with Feature 3 was faster than DNN2 with Feature 1 but slower than accuracy of 100%. However, cross-validation method was not imple
DNN2 with Feature 2. This was due to the number of the coefficients at mented, and no information was available for the percentage of the
the input layer. DNN with longer signal frame length and more hidden database used for training data and testing data.
layers performed better using Feature 3. Shorter signal frame length and
DNN with one hidden layer degraded the classification performance.
4.2. Performance of deep learning models using Dataset 2
In this study of using Dataset 1, an accuracy of 100% was achieved by
using hybrid features of MCMCT and DNN2 model. The experiment was
The size of Dataset 2 (2268 samples) is bigger than Dataset 1 (847
conducted using a 5-fold cross-validation involving 507 normal cry
samples). The non-asphyxia group includes cry samples of normal, deaf,
samples and 340 asphyxia cry samples from the Baby Chillanto
hunger and pain. The performance of the classifiers using Dataset 2 and
6
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
Table 4 Table 6
Highest accuracy achieved by DNN1, DNN2 and CNN when tested on Dataset 1 Highest accuracy achieved by DNN1 and DNN2 when tested on Dataset 1 using
using Feature 1 with different signal frame lengths, MFCC numbers, hidden Feature 3 with different signal frame lengths, MFCC numbers, hidden neuron
neuron numbers and hidden layer numbers. numbers and hidden layer numbers.
Signal 12 MFCC 16 MFCC 20 MFCC Classifier Signal Frame 49 MCMCT 61 MCMCT 73 MCMCT Classifier
Frame (Hidden 1 (Hidden 1 (Hidden 1 Length (ms) (Hidden 1 (Hidden 1 (Hidden 1
Length (ms) neuron no) neuron no) neuron no) neuron no) neuron no) neuron no)
20 95.86% (2000) 95.62% (500) 95.38% DNN1 20 98.58% 98.58% 98.22% DNN1
(500) (2000) (2000) (2000)
97.75% (32) 97.87% (32) 97.16% CNN 30 98.82% 98.82% 99.53% DNN1
(32) (2000) (2000) (2000)
30 95.86% (500) 95.62% (200) 95.74% DNN1 40 98.34% 97.87% 97.87% DNN1
(2000) (2000) (2000) (2000)
97.51% (32) 98.22% (32) 97.87% CNN 50 99.76% 99.53% 99.64% DNN1
(32) (2000) (2000) (1000)
40 96.33% (2000) 96.21% (1000) 96.09% DNN1 50 99.88% (100) – 99.88% (100) DNN2 with 2
(2000) hidden layers
98.11% (32) 98.34% (32) 97.87% CNN 99.88% (50) – 99.88% (50) DNN2 with 3
(32) hidden layers
50 96.45% 96.09% (1000) 96.45% DNN1 99.88% (50) – 100.00% DNN2 with 4
(2000) (2000) (200) hidden layers
97.87% (32) 97.75% (32) 97.51% CNN 99.88% (50) – 100.00% DNN2 with 5
(32) (1000) hidden layers
50 96.57% (1000) – 96.09% DNN2 with 2 100.00% – 99.88% (50) DNN2 with 6
(200) hidden layers (50) hidden layers
96.09% (500) – 96.45% DNN2 with 3 99.88% (500) – 99.88% (500) DNN2 with 7
(500) hidden layers hidden layers
96.33% (1000) – 96.21% DNN2 with 4
(2000) hidden layers
96.69% – 96.45% DNN2 with 5
(1000) (1000) hidden layers Table 7
96.57% (500) – 96.45% DNN2 with 6 Highest accuracy achieved by DNN1, DNN2 and CNN when tested on Dataset 2
(1000) hidden layers using Feature 1 with different signal frame lengths, MFCC numbers, hidden
96.57% (500) – 96.45% DNN2 with 7 neuron numbers and hidden layer numbers.
(2000) hidden layers
Signal 12 MFCC 16 MFCC 20 MFCC Classifier
Frame (Hidden 1 (Hidden 1 (Hidden 1
Length (ms) neuron no) neuron no) neuron no)
Table 5 20 97.84% (50) 97.88% (200) 97.97% DNN1
Highest accuracy achieved by DNN1 and DNN2 when tested on Dataset 1 using (100)
Feature 2 with different signal frame lengths, MFCC numbers, hidden neuron 99.03% (32) 98.98% (32) 98.90% CNN
numbers and hidden layer numbers. (32)
30 97.92% (500) 98.06% (200) 98.01% DNN1
Signal Frame 12 Mean 16 Mean 20 Mean Classifier
(2000)
Length (ms) MFCC MFCC MFCC
98.63% (32) 98.85% (32) 98.85% CNN
(Hidden 1 (Hidden 1 (Hidden 1
(32)
neuron no) neuron no) neuron no)
40 98.01% (1000) 98.15% (1000) 98.23% DNN1
20 95.50% 95.62% 95.74% DNN1 (2000)
(2000) (2000) (2000) 98.76% (32) 98.85% (32) 98.94% CNN
30 95.38% 96.33% 95.86% DNN1 (32)
(2000) (2000) (2000) 50 98.15% (2000) 98.15% (200) 98.23% DNN1
40 95.62% 96.35% 96.45% DNN1 (500)
(2000) (2000) (2000) 98.94% (32) 98.76% (32) 98.85% CNN
50 95.74% 96.69% 96.57% DNN1 (32)
(2000) (2000) (1000) 40 – – 97.92% DNN2 with 2
50 – 97.04% 97.04% (50) DNN2 with 2 (1000) hidden layers
(100) hidden layers – – 98.01% DNN2 with 3
– 97.99% 97.87% DNN2 with 3 (1000) hidden layers
(500) (2000) hidden layers – – 98.01% DNN2 with 4
– 97.99% 98.11% DNN2 with 4 (200) hidden layers
(200) (2000) hidden layers – – 97.97% DNN2 with 5
– 97.99% 98.11% DNN2 with 5 (200) hidden layers
(100) (100) hidden layers – – 97.92% DNN2 with 6
– 97.75% 98.11% DNN2 with 6 (500) hidden layers
(2000) (200) hidden layers – – 97.97% DNN2 with 7
– 95.98% 96.45% DNN2 with 7 (1000) hidden layers
(1000) (1000) hidden layers 50 – – 98.10% DNN2 with 2
(200) hidden layers
– – 98.06% DNN2 with 3
Feature 1 is shown in Table 7. DNN1 model achieved the highest ac (100) hidden layers
98.28% DNN2 with 4
curacy of 98.23% when signal frame lengths of 50 ms and 40 ms were – –
(200) hidden layers
used. DNN2 model achieved the highest accuracy of 98.28% at a signal – – 98.19% DNN2 with 5
frame length of 50 ms. CNN model outperformed the DNN1 model and (2000) hidden layers
DNN2 model, with the highest accuracy achieved at 99.03% when a – – 98.10% DNN2 with 6
signal frame length of 20 ms was used. DNN1 model and DNN2 model (500) hidden layers
98.10% DNN2 with 7
performed the best with 20 MFCCs. CNN model performed the best with
– –
(50) hidden layers
7
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
Table 9 Table 10
Highest accuracy achieved by DNN1 and DNN2 when tested on Dataset 2 using Comparison between our proposed method and existing studies.
Feature 3 with different signal frame lengths, MFCC numbers, hidden neuron
numbers and hidden layer numbers. Study Data Accuracy Accuracy Improvement
(samples) (%) achieved by (%)
Signal Frame 49 MCMCT 61 MCMCT 73 MCMCT Classifier our method
Length (ms) (Hidden 1 (Hidden 1 (Hidden 1 (%)
neuron no) neuron no) neuron no)
(Ji et al., 2019) Normal 96.74 99.96 3.22
20 98.81% 98.98% 98.94% DNN1 (1927)
(1000) (2000) (500) Asphyxia
30 98.90% 98.94% 99.03% DNN1 (340)
(2000) (2000) (2000) (Badreldine Normal 98.5 100.00 1.5
40 98.59% 98.90% 98.76% DNN1 et al., 2018) (340)
(2000) (2000) (2000) Asphyxia
50 99.38% 99.16% 99.43% DNN1 (340)
(2000) (2000) (2000) (Zabidi et al., Normal 92.78 100.00 7.22
50 99.74% – 99.69% DNN2 with 2 2018) (316)
(2000) (500) hidden layers Asphyxia
99.87% – 99.91% DNN2 with 3 (284)
(500) (200) hidden layers (Sahak et al., Normal 94.84 100.00 5.16
99.96% – 99.96% DNN2 with 4 2018) (316)
(500) (500) hidden layers Asphyxia
99.96% – 99.91% (50) DNN2 with 5 (284)
(100) hidden layers (Hariharan Normal 100 100.00 0
99.96% – 99.91% (50) DNN2 with 6 et al., 2018) (552)
(200) hidden layers Asphyxia
99.96% (50) – 99.96% DNN2 with 7 (340)
(2000) hidden layers (Moharir et al., Normal 94 100.00 6
2017) (507)
Asphyxia
12 MFCCs. DNN2 model performed the best when 4 hidden layers were (340)
used. (Sachin et al., Normal 92 100.00 8
2017) (1049)
The performance of the classifiers using Feature 2 is shown in
Asphyxia
Table 8. DNN1 model achieved the highest accuracy of 98.28% when a (340)
signal frame length of 50 ms was used. The DNN2 model performed (Rosales-Pérez Normal 90.68 100.00 9.32
better than the DNN1 model with the highest accuracy achieved at et al., 2015) (507)
98.54%. DNN2 model was able to achieve the highest accuracy when 5 Asphyxia
(340)
and 6 hidden layers were used. (Hariharan, Normal 99.19 100.00 0.81
The performance of classifiers using Feature 3 is shown in Table 9. Saraswathy, (340)
DNN1 model achieved the highest accuracy of 99.43% when a signal et al., 2012) Asphyxia
frame length of 50 ms was used. DNN2 model outperformed the DNN1 (340)
(Hariharan Normal 99.41 100.00 0.59
model with the highest accuracy achieved at 99.96%. DNN2 model was
et al., 2011) (340)
able to achieve the highest accuracy when 4 to 7 hidden layers were Asphyxia
used. (340)
8
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
study used 1927 non-asphyxia cry samples and 340 asphyxia cry References
samples.
The proposed method is compared with other studies in Table 10. Badreldine, O. M., Elbeheiry, N. A., Haroon, A. N. M., Elshehaby, S., & Marzook, E. M.
(2018). Automatic diagnosis of asphyxia infant cry signals using wavelet based mel
Our proposed method shows improvement in terms of accuracy, which frequency cepstrum features. ICENCO 2018 - 14th international computer engineering
varies between 0.59% and 9.32%. conference: Secure smart societies, Cairo, Egypt.
Baucas, M. J., & Spachos, P. (2020). Using cloud and fog computing for large scale IoT-
based urban sound classification. Simulation Modelling Practice and Theory, 101,
5. Conclusion Article 102013. [Link]
Cohen, R., Ruinskiy, D., Zickfeld, J., IJzerman, H., & Lavner, Y. (2020). Baby cry
The study investigated the use of hybrid features of MFCC, Chro detection: deep learning and classical approaches. In W. Pedrycz, & S.-M. Chen
(Eds.), Development and analysis of deep learning architectures (pp. 171–196). Springer
magram, Mel-scaled Spectrogram, Spectral Contrast, and Tonnetz in International Publishing, 10.1007/978-3-030-31764-5_7.
classifying asphyxia cry. These concatenated hybrid speech features Cohen, R. a. R. D. a. Z. J. a. I. H. a. L. Y. (2020). Baby Cry Detection: Deep Learning and
have not been used in any study to classify the asphyxia cry. This study Classical Approaches. 10.1007/978-3-030-31764-5_7.
Espinosa, R., Ponce, H., & Gutiérrez, S. (2021). Click-event sound detection in
demonstrated that its use may increase the accuracy of the classification
automotive industry using machine/deep learning. Applied Soft Computing, 108,
of infant asphyxia cry. Three different speech features were studied. Article 107465. [Link]
Feature 1 is the extraction of MFCC frame by frame and the total number Groenendaal, F., & van Bel, F. (2021). Perinatal asphyxia in term and late preterm infants.
of features is obtained by multiplying the number of MFCCs with the Retrieved 30th October 2021 from [Link]
asphyxia-in-term-and-late-preterm-infants.
total frame number. Feature 2 is the mean MFCC, which is calculated Hariharan, M., Saraswathy, J., Sindhu, R., Khairunizam, W., & Yaacob, S. (2012). Infant
over the total number of frames. Feature 3 is the concatenated hybrid cry classification to identify asphyxia using time-frequency analysis and radial basis
features of MFCC, Chromagram, Mel-scaled Spectrogram, Spectral neural networks. Expert Systems with Applications, 39(10), 9515–9523. [Link]
org/10.1016/[Link].2012.02.102
Contrast, and Tonnetz. Three different deep learning models were Hariharan, M., Sindhu, R., Vijean, V., Yazid, H., Nadarajaw, T., Yaacob, S., & Polat, K.
investigated in this study. The baseline model was the DNN1, which had (2018). Improved binary dragonfly optimization algorithm and wavelet packet
one hidden layer. DNN2 had 2 to 7 hidden layers. CNN was used to based non-linear features for infant cry classification. Computer Methods and
Programs in Biomedicine, 155, 39–51. [Link]
classify Feature 1 only as the speech features were fed into CNN in 2 Hariharan, M., Sindhu, R., & Yaacob, S. (2012). Normal and hypoacoustic infant cry
dimensions, i.e. total number of frames × MFCC number. Experiments signal classification using time–frequency analysis and general regression neural
were conducted to determine the optimal signal frame length, MFCC network. Computer Methods and Programs in Biomedicine, 108(2), 559–569. https://
[Link]/10.1016/[Link].2011.07.010
number, and hidden neuron number. The Baby Chillanto Database was Hariharan, M., Yaacob, S., & Awang, S. A. (2011). Pathological infant cry analysis using
used in this study where it was reorganized into two datasets. Dataset 1 wavelet packet transform and probabilistic neural network. Expert Systems with
consisted of 540 normal cry samples and 340 asphyxia cry samples. Applications, 38(12), 15377–15382. [Link]
Harte, C., Sandler, M., & Gasser, M. (2006). Detecting harmonic change in musical audio.
Dataset 2 comprised of 1928 non-asphyxia cry samples and 340
Proceedings of the 1st ACM workshop on audio and music computing multimedia, Santa
asphyxia cry samples. CNN model performed better than the DNN1 and Barbara, California, USA.
the DNN2 models when Feature 1 was used. DNN1 and DNN2 performed Issa, D., Fatih Demirci, M., & Yazici, A. (2020). Speech emotion recognition with deep
better with Feature 3 compared with Feature 1 and Feature 2. DNN2 was convolutional neural networks. Biomedical Signal Processing and Control, 59, Article
101894. [Link]
able to achieve an accuracy of 100% and 99.96% for Dataset 1 and Ji, C., Chen, M., Li, B., & Pan, Y. (2021). infant cry classification with graph
Dataset 2 respectively, when the hybrid features were used. The high convolutional networks. In 2021 IEEE 6th international conference on computer and
accuracy obtained may therefore be of potential clinical application for communication systems (ICCCS) (pp. 322–327).
Ji, C., Xiao, X., Basodi, S., & Pan, Y. (2019). Deep learning for asphyxiated infant cry
reliable detection and monitoring of infants with birth asphyxia once classification based on acoustic features and weighted prosodic features. Proceedings
clinically validated. of the 2019 international conference on internet of things (iThings) and IEEE green
computing and communications (GreenCom) and IEEE cyber, physical and social
computing (CPSCom) and IEEE smart data (SmartData), Atlanta, GA, USA.
CRediT authorship contribution statement Jiang, D. N., Lu, L., Zhang, H. J., Tao, J. H., & Cai, L. H. (2002). August). Music type
classification by spectral contrast feature. Proceedings of the IEEE international
Hua-Nong Ting: Conceptualization, Methodology, Software, Data conference on multimedia and expo.
Keras. (2021). Retrieved 1st July 2021 from [Link]
curation, Writing – original draft, Visualization, Investigation, Project Lahmiri, S., Tadj, C., Gargour, C., & Bekiros, S. (2022). Deep learning systems for
administration, Funding acquisition. Yao-Mun Choo: Conceptualiza automatic diagnosis of infant cry signals. Chaos, Solitons & Fractals, 154, Article
tion, Writing – review & editing. Azanna Ahmad Kamar: Conceptual 111700. [Link]
Librosa. (2021). Retrieved 1st July 2021 from [Link]
ization, Writing – review & editing.
McFee, B., Raffel, C., Liang, D., Ellis, D., McVicar, M., Battenberg, E., & Nieto, O. (2015).
Librosa: Audio and music signal analysis in python. Proceedings of the 14th Python in
Declaration of Competing Interest Science Conference, Austion, Texas, USA.
Moharir, M., Sachin, M., Nagaraj, R., Samiksha, M., & Rao, S. (2017). Identification of
asphyxia in newborns using gpu for deep learning. Proceedings of the 2017 2nd
The authors declare that they have no known competing financial International Conference for Convergence in Technology (I2CT), Mumbai, India.
interests or personal relationships that could have appeared to influence Moshiro, R., Mdoe, P., & Perlman, J. M. (2019). A global view of neonatal asphyxia and
the work reported in this paper. resuscitation. Frontiers in Pediatrics, 7, 489. [Link]
fped.2019.00489
Reyes-Galaviz, O. F., Cano-Ortiz, S. D., & Reyes-García, C. A. (2008). Evolutionary-neural
Acknowledgements system to classify infant cry units for pathologies identification in recently born
babies. Proceedings of the 2008 seventh Mexican international conference on artificial
intelligence, Atizapan de Zaragoza, Mexico.
We would like to thank Dr. Carlos A. Reyes-Garcia, Dr. Emilio Arch- Reyes-Galaviz, O. F., Reyes-García, C., Óptica, A., & Erro, L. (2004). A system for the
Tirado and his INR-Mexico group, and Dr. Edgar M. Garcia-Tamayo for processing of infant cry to recognize pathologies in recently born babies with neural
their dedication to the collection of the infant cry database. We would networks. Proceedings of the 9th Conference on Speech and Computer (SPECOM’2004),
St. Petersburg, Russia.
like to thank Boon-Fei Yong for partially coding the Python programme Rosales-Pérez, A., Reyes-García, C. A., Gonzalez, J. A., Reyes-Galaviz, O. F.,
for hybrid features and deep learning models. We would like to thank Escalante, H. J., & Orlandi, S. (2015). Classifying infant cry patterns by the Genetic
the Universiti Malaya for supporting this research under the Faculty Selection of a Fuzzy Model. Biomedical Signal Processing and Control, 17, 38–46.
[Link]
Research Grant (Grant Number: GPF074A-2018). Sachin, M. U., Nagaraj, R., Samiksha, M., Rao, S., & Moharir, M. (2017). GPU based Deep
learning to detect asphyxia in neonates. Indian Journal of Science and Technology, 10
Appendix A. Supplementary data (3), 1–5. [Link]
Sahak, R., Mansor, W., Lee, K. Y., & Zabidi, A. (2018). Support vector machine
performance with optimal parameters identification in recognising asphyxiated
Supplementary data to this article can be found online at [Link] infant cry. International Journal of Engineering and Technology(UAE), 7(3), 114–119.
org/10.1016/[Link].2022.118064. [Link]
9
H.-N. Ting et al. Expert Systems With Applications 208 (2022) 118064
Sahak, R., Mansor, W., Lee, Y. K., Yassin, A. I. M., & Zabidi, A. (2010). Performance of Zabidi, A., Mansor, W., Khuan, L. Y., Yassin, I. M., & Sahak, R. (2010). The effect of F-
combined support vector machine and principal component analysis in recognizing ratio in the classification of asphyxiated infant cries using multilayer perceptron
infant cry with asphyxia. Proceedings of the 2010 Annual International Conference of Neural Network. Proceedings of the 2010 IEEE EMBS Conference on Biomedical
the IEEE Engineering in Medicine and Biology, Buenos Aires, Argentina. Engineering and Sciences (IECBES), Kuala Lumpur, Malaysia.
Salehian Matikolaie, F., & Tadj, C. (2020). On the use of long-term features in a newborn Zabidi, A., Yassin, I. M., Hassan, H. A., Ismail, N., Hamzah, M. M. A. M., Rizman, Z. I., &
cry diagnostic system. Biomedical Signal Processing and Control, 59, Article 101889. Abidin, H. Z. (2018). Detection of asphyxia in infants using deep learning
[Link] Convolutional Neural Network (CNN) trained on Mel Frequency Cepstrum
Su, Y., Zhang, K., Wang, J., Zhou, D., & Madani, K. (2020). Performance analysis of Coefficient (MFCC) features extracted from cry sounds. Journal of Fundamental and
multiple aggregated acoustic features for environment sound classification. Applied Applied Sciences, 9(3S), 768–778. [Link]
Acoustics, 158, 107050. [Link]
10
Deep learning models face challenges such as lower performance compared to traditional neural networks when applied to the relatively small Baby Chillanto Database. Researchers have addressed these challenges by integrating hybrid speech features, which provide richer data for training, thereby improving model accuracy. Despite these efforts, the complexity and computational demands of deep learning remain significant issues .
Artificial neural networks (ANNs) were preferred in past studies because they achieved higher classification accuracies than deep learning models for asphyxia cries. ANNs like TDNN, MLP, and PNN demonstrated strong performance, with accuracies up to 99.41%, possibly due to their effective handling of smaller datasets typical of cry classification studies .
Different neural network architectures, such as DNN1, DNN2, and CNN, have varied impacts on the accuracy of cry classification models. For instance, CNNs perform best when using hybrid features but require longer training times. DNN2 achieves high accuracy with multiple hidden layers, especially using Feature 3. The architecture choice impacts computation time, accuracy, and the handling of different feature sets .
Computational challenges associated with using CNNs for infant cry classification include their demand for extensive computational resources and longer processing times than simpler models like DNNs. These challenges stem from the model's complexity, which involves numerous layers and parameters, thus requiring significant computational power and time for training and validation .
MFCC is popular in infant cry classification studies because it effectively represents the power spectrum of sound, capturing phonetic nuances crucial for distinguishing cry types. Its continued use is supported by its proven accuracy, simplicity, and efficiency in processing acoustic signals .
The Baby Chillanto Database is used extensively in studies to classify asphyxia cries. It provides a standard set of cry samples to train and test classification models. However, its limitation lies in its size, which is relatively small, prompting the use of techniques like K-fold cross-validation. The database comprises 507 normal and 340 asphyxia cry samples .
Findings suggest that asphyxia cry classification offers potential clinical applications as a surrogate parameter in evaluating intrauterine hypoxia. It could help monitor at-risk infants, such as those with mild hypoxia or perinatal risk factors, by identifying asphyxiated cries .
Feature selection enhances the performance of asphyxia cry classification by reducing redundant and non-discriminative features, thus improving classifier efficiency. Techniques like genetic algorithms have been used to optimize this process, leading to significant performance gains in some studies .
The main techniques used for speech feature extraction in classifying asphyxia in infant cries are Linear Predictive Coding (LPC), Mel Frequency Cepstral Coefficients (MFCC), Wavelet Packet Transform, and spectrograms, with MFCC being the most commonly used technique .
Hybrid speech features, which include MFCC, Chromagram, Mel-scaled Spectrogram, Spectral Contrast, and Tonnetz, achieve higher classification accuracy than single features like MFCC alone. The hybrid features contain richer information, leading to improved accuracy in asphyxia cry recognition .