0% found this document useful (0 votes)
19 views14 pages

Deep Learning for Parkinson's Detection

This document presents a study on a computer-aided diagnostic system for Parkinson's disease using spectrogram-based deep features. The research proposes three methods for classification, with the second method achieving the highest accuracy of 99.7% using deep features from speech recordings. The results indicate that the deep feature-based approach outperforms traditional acoustic feature methods and transfer learning techniques in detecting Parkinson's disease.

Uploaded by

m699599499
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views14 pages

Deep Learning for Parkinson's Detection

This document presents a study on a computer-aided diagnostic system for Parkinson's disease using spectrogram-based deep features. The research proposes three methods for classification, with the second method achieving the highest accuracy of 99.7% using deep features from speech recordings. The results indicate that the deep feature-based approach outperforms traditional acoustic feature methods and transfer learning techniques in detecting Parkinson's disease.

Uploaded by

m699599499
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Received January 17, 2020, accepted February 6, 2020, date of publication February 14, 2020, date of current version

February 28, 2020.


Digital Object Identifier 10.1109/ACCESS.2020.2974008

A Spectrogram-Based Deep Feature Assisted


Computer-Aided Diagnostic System for
Parkinson’s Disease
LAIBA ZAHID 1 , MUAZZAM MAQSOOD 1 , MEHR YAHYA DURRANI1 , MAHEEN BAKHTYAR2 ,
JUNAID BABER2 , HABIBULLAH JAMAL3 , IRFAN MEHMOOD4 , AND OH-YOUNG SONG 5
1 Department of Computer Science, COMSATS University Islamabad, Attock Campus, Attock 43600, Pakistan
2 Department of Computer Science and Information Technology, University of Balochistan, Quetta 87300, Pakistan
3 Facultyof Engineering Sciences, Ghulam Ishaq Khan Institute, Topi 23460, Pakistan
4 Department of Media Design and Technology, Faculty of Engineering & Informatics, University of Bradford, Bradford BD7 1AZ, U.K.
5 Department of Software, Sejong University, Seoul 05006, South Korea

Corresponding authors: Muazzam Maqsood ([Link]@[Link]) and Oh-young Song (oysong@[Link])


This work was supported in part by the MSIT (Ministry of Science and ICT), South Korea, through the Information Technology Research
Center (ITRC) Support Program supervised by the Institute for Information & Communications Technology Planning & Evaluation (IITP)
under Grant IITP-2019-2016-0-00312, and in part by the Faculty Research Fund of Sejong University in 2019.

ABSTRACT Parkinson’s disease is a neural degenerative disease. It slowly progresses from mild to severe
stage, resulting in the degeneration of dopamine cells of neurons. Due to the deficiency of dopamine cells in
the brain, it leads to a motor (tremor, slowness, impaired posture) and non-motor (speech, olfactory) defects
in the body. Early detection of Parkinson’s disease is a difficult chore as the symptoms of disease appear
overtime. However, different diagnostic systems have contributed towards disease detection by considering
gait, tremor and speech characteristics. Recent work has shown that speech impairments can be considered
as a possible predictor for Parkinson’s disease classification and remains an open research area. The speech
signals show major differences and variations for Parkinson patients as compared to normal human beings.
Therefore, variation in speech should be modeled using acoustic features to identify these variations.
In this research, we propose three methods- the first method employs a transfer learning-based approach
using spectrograms of speech recordings, the second method evaluates deep features extracted from speech
spectrograms using machine learning classifiers and the third method evaluates simple acoustic feature of
recordings using machine learning classifiers. The proposed frameworks are evaluated on a Spanish dataset
pc-Gita. The results show that the second framework shows promising results with deep features. The highest
99.7% accuracy on vowel \o\ and read text is observed using a multilayer perceptron. Whereas 99.1%
accuracy observed on vowel \i\ deep features using random forest. The deep feature-based method performs
better as compared to simple acoustic features and transfer learning approaches. The proposed methodology
outperforms the existing techniques on the pc-Gita dataset for Parkinson’s disease detection.

INDEX TERMS Parkinson disease, classification, deep features, speech signals, transfer learning.

I. INTRODUCTION disease [2]. The disease is common among people age 50 or


Parkinson’s disease (PD) is a slowly progressing neural above. However, early symptoms have also been observed at
degenerative disease. The main origin of the disease is still the age of 30-50. The disease primarily affects the central
unknown [1]. However, researchers have found that certain nervous system by causing the degeneration of dopamine-
hereditary and environmental factors aid in the cause of producing neuron cells. Dopamine is a chemical produced
Parkinson’s disease. Research has shown that nearly 100- by substantial nigra (basal ganglia) which is responsible
250 persons out of 100,000 are suffering from Parkinson’s for transferring signals within the brain. Loss of dopamine-
producing cells results in movement disorder in PD patients.
The associate editor coordinating the review of this manuscript and The symptoms of Parkinson’s disease are classified into
approving it for publication was Ahmed Farouk. the motor and non-motor symptoms. Motor symptoms are

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
35482 VOLUME 8, 2020
L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

associated with movements and are more perceptible as 2. We propose an acoustic-phonetic based approach for
compared to non-motor symptoms [3]. In motor symptoms, the detection of PD disease
the patient suffers from slowness of movement referred 3. We conducted a comparison between a proposed deep
to as bradykinesia, rigidity, postural instability and tremor. feature-based approach with simple acoustic features
Non-motor symptoms are evident at a particular interval; and transfer learning-based methods
include sleep disorder, speech, and swallowing problem and The rest of the research paper is organized as follows.
olfactory disorder (loss of sense of smell) [1], [3]. The
Section 2 explains the literature review, section 3 explain the
effect of Parkinson’s disease on speech is characterized as
proposed methodology, section 4 shows experimental setup
phonation, articulation, and prosody. Phonation is the use and dataset details, section 5 gives detail view of results and
of vocal folds for speech and articulation refers to the use simulations, followed by the conclusion and future work.
of special tissues in speech production. While prosody is
related to amplitude, loudness and pitch to produce sound.
Most of the work in Parkinson disease detection consid- II. LITERATURE REVIEW
ered phonation that includes the pronunciation of vowels This section explains the existing techniques on Parkinson’s
\a\, \e\, \i\, \o\, \u\ [4]. disease detection using the Spanish speech dataset. The tech-
Speech signals are usually considered as one of the main niques presented by different researchers are grouped into
methods to diagnose Parkinson’s disease. In [5] authors per- two categories that are explained in detail in the following
formed experimentation on speech signal recordings of three subsections.
languages identifying speech pronunciation problems in it.
It is observed that speech pronunciation like vowels, sen- A. MACHINE LEARNING BASED METHODS
tences, words are affected by this disease. Therefore, speech In the past few years, machine learning-based disease classi-
is considered as a major predictor of PD disease. In [6] fication is widely used in the medical field and has acquired
authors accounted that articulation, intelligibility, prosody remarkable significance [16], [17]. A Gaussian based den-
features of speech signal shows promising results in the detec- sity was considered by Moro-Velazquez et al. [18] from
tion of PD. In [7] the author presented that somehow the age four to five different PD corpora. They exploited phonetic
factor also contributes to disease. They assessed that the text-dependent utterances that require vocal tract features
speech recordings of young speakers show significant defects of speech signals. They assessed words, sentences, mono-
in speech pronunciation tasks. The researchers also observed logues and vowels in three corpora with more male patients
the monitoring of skype calls using normal sentences shows (Czech dataset). They showed better results as compared
significant errors in the pronunciation of PD patients [8]. to other (Spanish) datasets and reported 81% accuracy.
Traditionally, acoustic features are considered in most of the Rueda et al. [6] used a wrapper feature selection method
recent works along with SVM for PD detection. Most of for vowel \a\ and words \pa-ta-ka\ (Articulation, phona-
the recent work has performed disease detection using gait tion, Diadochokinetic features) from pc-Gita recordings and
[9], handwriting [10], [11] and speech datasets. Furthermore, achieved 70% accuracy. Pérez-Toro et al. [19] considered
the literature study shows Gaussian based model, several classical features like term frequency and bag of words of
machine learning techniques, convolution neural networks monologues only. They used pc-Gita recordings and reported
[12] are contributing to PD diagnosis. Parkinson’s disease that language pronunciation contains enough information
detection using most suitable speech impairment features is for PD classification and achieved 72% accuracy results.
imperative and still is an open research area. Karan et al. [20] assessed inherent and decomposition-based
This research work contemplates speech recordings using features from vowel \a\, \o\ only from two datasets pc-Gita
spectrograms and acoustic features. In our work, all record- and Saarbrucken dataset. They have reported 96% accuracy
ings are transformed into short-time Fourier transform (spec- on the Spanish dataset using random forest and support vector
trograms) that are used in the transfer learning method. machine. Kacha et al. [21] showed that use of PCA for artic-
We proposed a simple acoustic features based method and ulation features from sentences of Spanish pc-Gita speech
also considered a pre-trained convolution neural network spectrograms (STFT) is efficient for the detection of Parkin-
architecture [13], [14] Alexnet model for deep feature extrac- son’s disease. Parra-Gallego et al. [4] considered articulation
tion and detection of PD. For fair comparison, we have and intelligibility features from the words like \pa-ta-ka\
also used transfer learning-based classification. To evalu- from Spanish pc-Gita recordings. The authors reported 88%
ate the performance of the proposed methods, Parkinson’s accuracy in their work when classifiers were trained using
disease speech recordings from PC-GITA [15] dataset are intelligibility and articulation-based features.
used. The results show that the deep features based tech- Vasquez-Correa et al. [22] proposed a novel approach by
nique produced better results. The main contributions of considering the on-off state of vocal folds (i.e. on when
our research for PD detection using speech signals are the the candidate starts speaking and off when they stop speak-
following: ing). It includes words, vowels, monologues, and sentences
1. We propose a spectrogram based approach to extract from pc-Gita recordings. They reported 94.9% accuracy
deep feature to distinguish PD patients from healthy in speech signals for PD classification. Garcia et al. [23]

VOLUME 8, 2020 35483


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

assessed the appropriateness of i-vectors to classify Parkin- 91% accuracy using SVM. El Maachi et al. [30] evaluated the
son’s disease using articulation, prosody, phonation features gait physionet dataset for PD diagnosis. The author assessed
for words, vowels and 10 sentences from pc-Gita dataset. Parkinson disease from gait using a deep 1-D neural network.
They computed cosine difference that confers for disease They achieved 98.7% accuracy using deep neural networks
detection contributing 78% accuracy using articulation fea- and 85.3% accuracy in finding the severity of PD. Turker
tures. Klumpp et al. [24] developed a model that considered and Dogan [31] proposed an octopus based multiple pool-
the voice signal during phone calls and also considered syl- ing method (comprising of eight poling method) for feature
lable \pa-ta-ka\. They evaluated the severity and onset of extraction. They evaluated the vowels dataset and achieved
PD disease in their work. Arias-Vergara et al. [7] assessed 99.2% accuracy using SVM for PD and gender classification.
the aging factor in their research work. They considered Diogo et al. [32] presented an approach for early diagno-
articulation, prosody features, and age factors from pc-Gita sis of PD using three distinct database consisting of vowel
Spanish speech vowels. pronunciation in different language (Portuguese, uci data
The authors also accounted for gender factors and reported etc). They evaluated acoustic and phonetic characteristics of
that the age factor plays a very important role in classify- speech. Their work achieved 99.94% of highest accuracy
ing Parkinson’s disease. Their research evaluated that young using random forest.
speaker signals are more contributing towards PD classifica-
tion process than the old age speaker signals. They modeled B. CONVOLUTION NEURAL NETWORK-BASED METHODS
binary and multi-class support vector machines and compared Recent studies show that neural networks immensely con-
their results with neural networks and reported 95% accuracy. tribute to speech classification [13], [33]. Trinh and Darragh
Moro-Velázquez et al. [25] presented a new approach for [12] proposed a convolution neural network-based approach
classifying a speech signal into Parkinson’s patient or healthy for two PD datasets- Saarbrucken voice databases and pc-Gita
patient. They proposed a phonological feature-based method dataset and achieved 96.7% accuracy for a pc-Gita dataset.
in which the speech signal of words, monologues and read Naranjo et al. [29] proposed a convolution neural net-
text from the Spanish pc-Gita dataset. The authors reported work model for feature extraction and classification pro-
that this approach is quite useful in the assessment of Parkin- cess. They considered articulation from /pa-ta-ka/ words,
son’s disease patients in clinics. Moro-Velázquez et al. [25] sentences and read text from pc-Gita and observed start,
used traditional machine learning methods considering artic- stop signal of speech. They extracted features from spectro-
ulation and phonological features for Parkinson’s disease grams and achieved 89% accuracy using Gaussian mixture
detection. They assessed kinetic features for /pa-ta-ka/, two modeling and i vector. Teixeira et al. [34]. concentrated on
read sentences and a sustained vowel /a/ from Spanish pc-Gita discretizing a neural network and its information sources by
dataset from Parkinson disease patient (PDP) speech sig- utilizing quantization and weight scaling. They applied linear
nals. They performed classification tasks using Gaussian homomorphic encryption batching technique on the Spanish
mixture modeling and i-vectors and reported 87% accuracy. pc-Gita dataset. They observed the time of 1.4ms rather than
Orozco- Arroyave et al. [26] proposed an open-source soft- the original approach where prediction took 4.5s to com-
ware for Parkinson’s disease. The authors have used phona- pute results. Arias-Vergara et al. [35] considered phonation,
tion, articulation, prosody, and intelligibility dimensions of articulation, and prosody data from monologue recordings
speech signal from pc-Gita dataset for vowels using conven- of extended version of the Spanish language dataset. They
tional machine learning to identify Parkinson. They designed used support vector machine and convolution neural network-
a system that can be easily adopted by clinicians to assess based model to extract the most suitable features and support
different voice diseases. Arias-Vergara et al. [8] proposed a vector machine for classification. The achieved 84% accuracy
model for assessment of Parkinson’s disease using individual and showed prosody features are effective than the others.
speaker speech signal analysis. They assessed phonation,
articulation, and prosody to model recordings of spontaneous III. PROPOSED METHODOLOGY
speech and a read text from the Spanish pc-Gita dataset In this work, we proposed spectrogram and acoustic feature-
from different channels (mobile phone calls, online calls like based frameworks for Parkinson’s disease classification.
skype). In this work, authors observed that skype speech The first framework employs a transfer learning approach
signals were effective in distant observation of Parkinson’s for speech spectrograms. In our second proposed method,
disease patients. They performed evaluation using Gaussian we evaluated deep learning-based feature extraction from
mixture modeling and i-vectors and obtained a 0.77% cor- speech spectrograms while the third method considers the
relation. Vásquez-Correa et al. [27] presented an improved simple acoustic feature method for Parkinson disease detec-
version of m-FDA for PD detection. They considered phona- tion. Deep learning has seen tremendous results in many
tion, articulation, prosody, and intelligibility features from fields such as computer vision, image processing and speech
Spanish vowel /a/, sentences and words. Orhan et al. [28] signal recognition [14]. In our work, we used pre-trained
considered vowel recordings of freely available datasets [29]. convolution neural network architecture Alexnet for fea-
They incorporated statistical pooling for increasing features ture extraction from speech signal and speech spectrograms.
and used ReliefF for selecting the best features and achieved Fig 1 shows PD classification using the transfer learning

35484 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

FIGURE 1. Spectrogram based transfer learning model for PD classification.

FIGURE 2. Waveform of monologue of healthy candidate. FIGURE 3. Waveform of monologue of parkinson patient.

method. In both models, the first step is signal preprocessing TABLE 1. Signal processing parameters.
followed by deep feature extraction using Alexnet and hand-
crafted feature extraction. We have performed classification
using transfer learning and machine learning classifiers. All
steps of the methodology section are explained in detail in
following subsections.

A. SIGNAL PREPROCESSING
To input data into a classifier, the speech signals are first
converted to spectrograms. A spectrogram is a visual repre-
sentation of the signal spectrum that changes over time [36].
In our work, we converted Parkinson disease speech signal
into a spectrogram. In the time domain, digitally sampled data the pc-Gita dataset. Our dataset consists of spectrograms of
is divided into segments that overlap and form Fourier trans- vowels \a\, \e\, \i\, \o\, \u\, monologues and read text.
form that calculate the spectral amplitude of each segment. Alexnet is a deep neural network-based model that consists
Each segment corresponds to vertical line in image. Table 1. of 8 layers. The first five layers form the convolution layer
Shows signal processing parameters, Fig 2. shows waveform whereas the last three layers combine to form fully connected
of monologue pronounced by healthy candidate and Fig 3. layers. Fig 4. shows feature extraction process using Alexnet.
shows waveform of Parkinson disease. To input data into Alexnet model, we scaled spectrograms
to fit into the model. Spectrograms obtained from signal
B. FEATURE EXTRACTION: DEEP FEATURES USING preprocessing are of size 224 × 224 whereas Alexnet accepts
ALEXNET MODEL input images of size 227 × 227.
In our work, we used deep learning and machine learning Thus, all images are scaled accordingly. Alexnet model has
methods for the classification process. In order to input data a convolution layer, hidden layers, and classification layer
into our classifiers, deep features were extracted [37], [38] at the end. First five convolution layers of the network is
from our speech signal dataset. We used deep learning con- trained on ImageNet data and last three fully connected layer
volution model Alexnet for extracting deep features from are replaced with target Parkinson disease data. In this work,

VOLUME 8, 2020 35485


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

FIGURE 4. Transfer learning model using Alexnet.

TABLE 2. Deep features extracted using Alexnet. b: PRE-TRAINING USING ALEXNET


Initially, Alexnet model is trained on ImageNet data that
consist of 1.2 million images. This model accepts images of
size 227 × 227 so all RGB images are resized from 224 × 224
size to 227 × 227 to fit into the model. The first five layers of
Alexnet is trained using ImageNet data. The fully connected
layers fc1, fc2, fc3 are shown in Fig 4, are replaced by the
pc-Gita Spanish speech dataset.

c: CONVOLUTION LAYER
This layer performs all computation and transfer it to the
we extracted features from the first five convolution layers next fully connected layers. This layer generates convolu-
thus fully connected classification layers are not used in this tion feature map c1, c2, c3, c4, c5 of input images and
architecture. The model extracts deep shallow features from refers each feature map of the previous layer to the next
individual data vowels, monologues and reads text. Table 2. layer.
shows a total number of features extracted from each set
of Parkinson’s Spanish speech data. The generic transfer
d: POOLING LAYER
learning architecture is presented here:
This layer reduces all computations and parameters to reduce
network complexity. It reduces the dimensionality of input
a: INPUT LAYER
data by mixing the output of the previous layer with the input
In the transfer learning model, the first layer accepts input
of the next layer.
that is basically images of size 227 × 227. We input RGB
spectrograms (vowels, monologues, read the text) and each
e: FULLY CONNECTED LAYER
of them is given as a separate input.
A fully connected layer connects neurons in a layer to each
Input = Ni × Wi × Hi × D (eq.1)
neuron in another. In principle, it is identical to the traditional
In eq.1, Ni represents a number of images, Wi is with for input multilayer perceptron. The dashed matrix passes through
image i, Hi is height D is the depth. fully connected layers to classify data.

35486 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

TABLE 3. Handcrafted features extracted using Mir toolbox.

TABLE 4. Varying parameter values in deep learning-based transfer and healthy patient for each set of data monologues, vowels,
learning approach.
read text and words. We extract simple acoustic features [39]
of each recording. These features include spectral features
and statistical features. Table 3 shows the details of features
extracted for each recording and derivative of each feature.

D. PARKINSON DISEASE CLASSIFICATION


After signal preprocessing and feature extraction, we per-
formed classification. To classify PD patients, we utilized
deep learning convolution neural network architecture and
machine learning models; support vector machine, random
forest, and multilayer perceptron. The following subsection
explains methods for classification in detail.

f: REPLACEMENT OF LAST LAYERS 1) TRANSFER LEARNING APPROACH.


To perform classification, we initially trained our Alexnet The transfer learning method is the most widely used tech-
model on ImageNet data images embedded in it. To test our nique in the deep learning model. We trained model on base
model’s accuracy we replaced layers with our target data data and utilized it to learn features and transfer it to target
pc-Gita Spanish speech recordings’ spectrograms. data [40], [41] using Alexnet model [42], [43].

g: NETWORK TRAINING 2) MACHINE LEARNING BASED APPROACH


We trained our network for each set of dataset i-e vowels, the In this work, we utilized machine learning classifiers for
weighted learn rate is also varied at different points from 30 to Parkinson disease detection. We used simple and deep fea-
70 to assess accuracy at different learn rates. The bias factor tures that were extracted using feature extraction shown
is varied from 40 to 70 and batch sizes 5 to 10. The initial in Fig 5 using Alexnet. We performed 5 cross-validation for
learn rate is set to le-4. Table 4 show complete details of the each set of our data. The following sections explain classifiers
parameters. in detail with varying parameters used to train and test data.

C. FEATURE EXTRACTION: HANDCRAFTED a: SUPPORT VECTOR MACHINE


FEATURE-BASED MODEL The support vector machine has been widely used in dif-
In this feature extraction model, we extracted simple acoustic ferent fields like computer-aided diagnostic system, recog-
features from Spanish speech recordings. Separate sets of nition and vision system [37], [38]. It is the most popular
features were extracted for both Parkinson’s disease patient used machine learning model for binary classification due

VOLUME 8, 2020 35487


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

FIGURE 5. Machine learning-based approach for PD classification using deep features and handcrafted features.

to its generalizability. It generates different support vectors system. It has been observed from recent research that a
of the given input. It identifies the linear and non-linear sur- multilayer perceptron is extensively used in the medical diag-
faces in the input support vectors by constructing hyperplane, nostic field [43], [44]. It is a feedforward neural network that
which later classifies the data. The complexity parameters in is made up of an input layer, an output layer and in between,
this model build hyperplane from the class label. The hyper- them is a hidden layer. The input layer accepts the input value
plane that computes the largest distance from the training whereas the hidden layer sends information from the input
data has the highest classification result. The parameters like layer to the output layer as shown in Fig 6. A hidden layer
gamma rate are set to 0.01, complexity value is 1.0, the degree consists of a number of neurons, each of the hidden layer
value is 3 and the coefficient coef0 is 1. We have used the neuron has information for its input influencing it by growing
linear in our research. In the linear kernel, the projection for them by their connection weights. The output of each neuron
input task is considered by dot product for the input value y is defined as
and the support vector yi is calculated. X
wi = f ( xii zi ) (eq.3)
f (y) = B(0) + sum (zi (y, yi )) (eq.2)
In eq.2, coefficient B (0) and zi is estimated for each input In eq.3, f is defined as an activation function, which is
value y and evaluated by learning a data learning algorithm. proportional to input weights. It is mostly some thresh-
old value a simple sigmoid or a hyperbolic tangent
b: RANDOM FOREST function. The learning rate for the multilayer percep-
An ensemble method widely used in the classification pro- tron ranges from 0 to 1 where 0.3 is set as a default
cesses that make use of different decision trees for classifying value.
data [44]. It builds the bootstrap templates from the random
forest original data and grows a raw classification or regres- IV. EXPERIMENTAL SETUP
sion tree for each bootstrap template. It considers each node A. DATASET
instead of choosing the only best disclosure from all pre- We utilized PC-GITA [15] Spanish language dataset. The
dictors. It performs a random selection of predictors and dataset consists of a Spanish speech signal recording
chooses the best split between them as shown in Fig 6. of 50 people that are PD patients and 50 HC people as shown
In our research, we utilized default parameters for random in Table 5. The dataset includes recordings of 25 male and
forest. 25 female persons. The dataset belongs to Spanish language
Table 5 depicts dataset details. The dataset consists of the
c: MULTILAYER PERCEPTRON recording of vowels, monologues and read the text. Each
A multilayer perceptron is an artificial neural network that recording consists of different voice features that are dis-
are broadly utilized in speech, image, and vision recognition cussed below.

35488 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

FIGURE 6. Random forest and multilayer perceptron model.

1) PHONATION 2) SENSITIVITY
The phonation analysis in continuous speech is performed by Sensitivity is the ability of a test to correctly identify those
extracting voiced segments from the utterance. The feature with the disease (true positive rate)
set includes seven descriptors such as jitter and shimmer. TP
The first and second derivatives of F0, long term perturbation Sensitivity = (eq.5)
TP + TN
features such as the amplitude perturbation quotient, the pitch
perturbation quotient, and the energy. 3) SPECIFICITY
is the ability of the test to correctly identify those without the
2) PROSODY disease (true negative rate).
The prosody features are based on duration, the F0 contour, TN
and the energy contour. We computed 13 features per utter- Specificity = (eq.6)
TN + TP
ance including the average, standard deviation, and maximum
value of F0. 4) F1 SCORE
It is defined as a ratio or numerical average of precision and
3) ARTICULATION recall values from the classification result.
The articulatory capability of the patients is evaluated precision × recall
f 1 score = 2 × (eq.7)
with information from the onset/offset transitions to model precision + recall
the difficulties of patients to start/stop the movement
V. RESULTS
of the vocal folds. The set of features extracted from
the onset and offset includes 12 Mel-Frequency Cep- This section explains in detail the results obtained transfer
stral Coefficients (MFCCs) with their first and second learning, deep feature-based, and machine learning approach.
derivatives.
A. TRANSFER LEARNING BASED APPROACH RESULTS
In our first framework, the initial step comprises the conver-
B. EVALUATION METRICS
sion of Spanish speech recordings into spectrograms.
The results obtained after classification from deep and
All speech recordings are transformed into their rela-
machine learning models are evaluated using the following
tive spectrograms by using parameters described in signal
evaluation metrics explained in the below subsection.
preprocessing. In transfer, the learning approach model is
trained using a source dataset, which is replaced by our target
1) ACCURACY
dataset [6]. Once our model is trained on the source dataset,
It is defined as the total number of samples that are truly we replaced it with our target dataset speech recordings spec-
classified and the total number of negatively classified results. trograms. The spectrograms of each set of data i.e. vowels,

(TP + TN )
 monologues, read text and words are individually considered
Accuracy = (eq.4) by varying different parameters. The model is trained by
Total
varying parameters discussed in the network training section
In this TP is true positive and TN is true negative, and total discussed earlier. The corresponding results obtained by vary-
represents total number of class predictions. ing parameters are shown in Table 6. The results depicted in

VOLUME 8, 2020 35489


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

TABLE 5. Dataset table consists of details of data with total number of recordings and age range of Parkinson disease patient recordings and healthy
candidate for each gender (M/F).

the vowel dataset that vowel \a\ shows the highest accuracy
of 73.7% while vowel \e\ shows the highest accuracy of 75%.
In the case of vowels, the accuracy achieved for \i\ is 81.4%
and vowel \o\ achieved maximum accuracy of 72.6%. Vowel
\u\ attained 82.3% on epoch size 9. The read text dataset
showed the highest accuracy of 91% on epoch size 8 where
the weighted learn rate is 30. For monologues, 86.36% high-
est accuracy observed on weighted learn rate 30 and epoch
size 7.
For words, /apto/ we achieved 77.2 % accuracy on epoch
size 7 and weighted learn rate 60, whereas for \atelta\ we
achieved 73.7% accuracy on epoch size 9 and weighted learn
rate 40. It is observed from the above results that the highest
accuracy is observed in reading text 91% that is a major con- FIGURE 7. Accuracy observed during training process.
tribution using transfer learning in this research. Fig 7 shows
the training process for reading text data. Figure 7 depicts
training done using epoch size 6, blue lines depict the train-
ing process that is smoothed, light blue lines show the
actual training process and black lines show the accuracy
achieved on validation data. The accuracy achieved in this
figure is 77.27% on epoch size 7 and weighted learn rate 40.
Fig 8 shows the loss observed in read text data on epoch size 6.

B. MACHINE LEARNING BASED APPROACH RESULTS


1) DEEP FEATURES BASED RESULTS
In our second framework, we evaluated our dataset Span-
ish speech signal dataset using a machine learning model.
The initial step comprises of deep feature extraction. The
Spanish speech recordings spectrograms obtained after signal
preprocessing steps are used to extract the speech features. FIGURE 8. Loss observed during training process.
In this approach, spectrograms are used as input into the
feature extraction Alexnet model. The model extracts deep
features from spectrograms as shown in the transfer learning separately input into different machine learning classifiers.
architecture model in the above sections. The total number The support vector machine, random forest and multilayer
of features extracted for the vowel recordings dataset is 150. perceptron are separately validated on our Spanish speech
The 50 features are extracted for monologues, read text and recordings dataset. All of these classifiers used five-fold
word dataset. The extracted features for each of the data cross-validation. The support vector machine results show

35490 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

TABLE 6. Accuracy obtained using transfer learning model by varying parameters, bias learn rate 50, initial learn rate is le-4, epoch size vary from 6-10.

that the highest accuracy obtained for vowel \o\ is 93% acoustic features. Fig 10. depicts the results obtained for each
whereas the least accuracy result is 83% on the monologue set of data. It is observed that the highest accuracy is observed
dataset. The highest accuracy in this model obtained is on on vowel \e\ that is 84.6% using random forest, vowel
vowel \ e\ 99.4% whereas other vowel results also show an \o\ showed 83.6% accuracy and vowel \a\ presents 83%
average accuracy of 99.1%. accuracy while vowel \I\ show 82.8% accuracy using ran-
The multilayer perceptron showed the highest accuracy dom forest. Least accuracy in the vowel dataset is observed
of 99.7% on vowel \e\, \i\, \o\ whereas on the other data using vowel \u\. In the case of the reading text dataset, the
this model outperforms the other classifiers results. The highest accuracy 73% is observed using the random forest.
results of the machine learning model clearly depict that vow- Monologues dataset showed a bad accuracy of 37% using a
els are sufficient in the classification of Parkinson’s disease. multilayer perceptron and 15% using random forest. Thus,
Fig 9. depicts the accuracy obtained for each classifier. Mini- it is observed that vowels are efficient in Parkinson disease
mum accuracies obtained using a machine learning approach detection when handcrafted features are utilized. However,
is using support vector machine 83%. In the monologues in comparison to deep features, this accuracy is far less as
dataset, the highest 99.7% accuracy is obtained using a mul- the highest 99.7% accuracy is achieved. Thus, deep features
tilayer perceptron. based methods outperform other methods.

2) HANDCRAFTED FEATURE-BASED RESULTS C. COMPARATIVE ANALYSIS


Different acoustic features are used for Parkinson’s disease This section performs a brief comparison of recent existing
detection. In this work, we consider handcrafted acoustic techniques with the proposed technique in this research work
features from the Spanish speech dataset. This part of our shown in Table 7. In this research work, we evaluated deep
research work performs a comparison with our deep feature- learning and machine learning-based approaches. Each of
based machine learning model and transfer learning classi- these approaches considered a pc-Gita dataset that consists
fication. Fig 10. depicts the results obtained using simple of Spanish speech recordings which include pronunciation of

VOLUME 8, 2020 35491


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

FIGURE 9. Accuracy observed using deep features.

FIGURE 10. Accuracy observed using handcrafted features.

vowels, monologues, words and read the text. All of them are In comparison to our transfer learning approach
separately used as an input in both our model and their corre- Trinh and Darragh [12] used a convolution neural network
sponding accuracies are recorded. The results clearly depict in their work. Tuncer and Dogan [31] assessed the extended
that deep feature-based method outperformed by presenting version of Spanish recordings corpus using a convolution
98.3% of average accuracy for random forest and 99.3% for neural network. Trinh and Darragh [12] used Gaussian
multilayer perceptron. In comparison to Karan et al. [20] mixture modeling and i vector in their proposed work.
incorporated inherent based features from two different sets The comparison shows that our proposed machine learn-
of data observed 96% accuracy when classified using random ing method outperformed existing techniques, however, the
forest and support vector machine. These results are better transfer learning approach compared with existing tech-
than the already published work for the same datasets which niques does not prove to be as much as accurate than the
are presented in Table 7. convolution neural network research that has been already

35492 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

TABLE 7. Comparative analysis of proposed method with existing techniques.

presented [45], [46]. Parkinson disease shows several motors text, the same classifiers achieved 97.8%, 99%, 91% whereas
and non-motor symptoms. Gait, tremor and handwriting for monologues 97%, 99.3%, 86.36% accuracy is achieved.
problem appear overtime. However, changes in the speech Thus, our work shows that speech analysis using deep fea-
and handwriting start appearing early and are more evident. tures are efficient in distinguishing between healthy and
Changes in speech show a clear distinction between a healthy Parkinson patient with high accuracy. Further, we also con-
and PD patient. They suffer pause in pronouncing as well clude that pronunciation recordings of vowels are enough
as jarring sound. Classification of PD is crucial; however, in distinguishing patients from healthy. In future, the gait,
speech impairments are considered as an important biomarker tremor and other symptoms data can be assessed together
for PD detection. However, it is essential to perform the using a deep feature method to identify that to what extent this
speech collection of PD in a noise-free environment. Speech approach is suitable with other scenarios. Furthermore, the
analysis using deep learning-based methods proved as a feature selection process can be used for handcrafted feature
very good method because the variation in speaking can be method.
modeled effectively. However, the symptoms of Parkinson
can be observed in handwriting and gait. Therefore, the PD REFERENCES
identification can be done in a more efficient way by com- [1] D. K. Simon, C. M. Tanner, and P. Brundin, ‘‘Parkinson disease epidemi-
bining all the above-mentioned biomarkers. ology, pathology, genetics, and pathophysiology,’’ Clinics Geriatric Med.,
vol. 36, no. 1, pp. 1–12, Feb. 2020.
[2] A. E. Lang and A. M. Lozano, ‘‘Parkinson’s disease,’’ New England
D. CONCLUSION J. Med., vol. 339, pp. 1130–1143, 1998.
[3] D. Georgiev, M. Domellof, K. Hamberg, L. Forsgren, and G. M. Hariz,
Parkinson’s disease is one of the common diseases among ‘‘Sex differences, quality of life and non-motor symptoms in Parkinson’s
people worldwide. Early diagnosis of the disease is an open disease,’’ Neurology, vol. 85, no. 1, Jul. 2015.
research and many researchers have shown significant work [4] L. F. Parra-Gallego, T. Arias-Vergara, J. C. Vásquez-Correa,
N. Garcia-Ospina, J. R. Orozco-Arroyave, and E. Nöth, ‘‘Automatic
in achieving the highest accuracy for its detection and diag- intelligibility assessment of Parkinson’s disease with diadochokinetic
nostic. In our work, we used Alexnet model for deep fea- exercises,’’ in Proc. Workshop Eng. Appl., 2018, pp. 223–230.
ture extraction and handcrafted feature extraction from the [5] N. Hosseini-Kivanani, J. C. Vásquez-Correa, M. Stede, and E. Nöth,
‘‘Automated cross-language intelligibility analysis of Parkinson’s disease
Spanish speech recordings dataset. For classification, we used patients using speech recognition technologies,’’ in Proc. 57th Annu. Meet-
transfer learning, deep feature and acoustic-phonetic feature- ing Assoc. Comput. Linguistics, Student Res. Workshop, 2019, pp. 74–80.
based methods. We proposed that deep features extracted [6] A. Rueda, J. C. Vásquez-Correa, C. D. Rios-Urrego, J. R. Orozco-
Arroyave, S. Krishnan, and E. Nöth, ‘‘Feature representation of patho-
using Alexnet are efficient to distinguish Parkinson patients physiology of parkinsonian dysarthria,’’ in Proc. Interspeech, Sep. 2019,
from healthy patients. In our proposed models, random pp. 3048–3052.
forest, multilayer perceptron, transfer learning achieved high- [7] T. Arias-Vergara, J. C. Vásquez-Correa, and J. R. Orozco-Arroyave,
‘‘Parkinson’s disease and aging: Analysis of their effect in phonation
est accuracy of 99%, 99.7%, 72% respectively. For vowels and articulation of speech,’’ Cogn. Comput., vol. 9, no. 6, pp. 731–748,
\o\ the same classifiers achieved 99%, 99.6, 76%. For read Dec. 2017.

VOLUME 8, 2020 35493


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

[8] T. Arias-Vergara, J. C. Vásquez-Correa, J. R. Orozco-Arroyave, and [26] J. R. Orozco-Arroyave, J. C. Vásquez-Correa, J. F. Vargas-Bonilla,


E. Nöth, ‘‘Speaker models for monitoring Parkinson’s disease progression R. Arora, N. Dehak, P. S. Nidadavolu, H. Christensen, F. Rudzicz,
considering different communication channels and acoustic conditions,’’ M. Yancheva, H. Chinaei, A. Vann, N. Vogler, T. Bocklet, M. Cernak,
Speech Commun., vol. 101, pp. 11–25, Jul. 2018. J. Hannink, and E. Nöth, ‘‘NeuroSpeech: An open-source software for
[9] E. Rastegari, S. Azizian, and H. Ali, ‘‘Machine learning and Parkinson’s speech analysis,’’ Digit. Signal Process., vol. 77, pp. 207–221,
similarity network approaches to support automatic classifi- Jun. 2018.
cation of parkinson’s diseases using accelerometer-based gait [27] J. C. Vásquez-Correa, J. R. Orozco-Arroyave, T. Bocklet, and E. Nöth,
analysis,’’ in Proc. 52nd Hawaii Int. Conf. Syst. Sci., 2019, ‘‘Towards an automatic evaluation of the dysarthria level of patients
pp. 4231–4242. with Parkinson’s disease,’’ J. Commun. Disorders, vol. 76, pp. 21–36,
[10] C. D. Rios-Urrego, J. C. Vásquez-Correa, J. F. Vargas-Bonilla, E. Nöth, Nov. 2018.
F. Lopera, and J. R. Orozco-Arroyave, ‘‘Analysis and evaluation of hand- [28] O. Yaman, F. Ertam, and T. Tuncer, ‘‘Automated Parkinson’s disease
writing in patients with Parkinson’s disease using kinematic, geometrical, recognition based on statistical pooling method using acoustic features,’’
and non-linear features,’’ Comput. Methods Programs Biomed., vol. 173, Med. Hypotheses, vol. 135, Feb. 2020, Art. no. 109483.
pp. 43–52, May 2019. [29] L. Naranjo, C. J. Pérez, J. Martín, and Y. Campos-Roca, ‘‘A two-stage vari-
[11] R. Castrillon, A. Acien, J. R. Orozco-Arroyave, A. Morales, J. F. Vargas, able selection and classification approach for Parkinson’s disease detec-
R. Vera-Rodrıguez, J. Fierrez, J. Ortega-Garcia, and A. Villegas, tion by using voice recording replications,’’ Comput. Methods Programs
‘‘Characterization of the handwriting skills as a biomarker for Biomed., vol. 142, pp. 147–156, Apr. 2017.
parkinson disease,’’ 2019, arXiv:1903.08226. [Online]. Available: [30] I. El Maachi, G.-A. Bilodeau, and W. Bouachir, ‘‘Deep 1D-convnet for
[Link] accurate Parkinson disease detection and severity prediction from gait,’’
[12] N. Trinh and O. B. Darragh, ‘‘Pathological speech classification using Expert Syst. Appl., vol. 143, Apr. 2020, Art. no. 113075.
a convolutional neural network,’’ in Proc. IMVIP, Ireland, 2019, [31] T. Tuncer and S. Dogan, ‘‘A novel octopus based Parkinson’s disease
pp. 72–75. and gender recognition method using vowels,’’ Appl. Acoust., vol. 155,
[13] R. Miikkulainen, J. Liang, E. Meyerson, A. Rawal, D. Fink, and pp. 75–83, Dec. 2019.
O. Francon, ‘‘Evolving deep neural networks,’’ in Artificial Intelligence [32] D. Braga, A. M. Madureira, L. Coelho, and R. Ajith, ‘‘Automatic detection
in the Age of Neural Networks and Brain Computing. Amsterdam, of Parkinson’s disease based on acoustic analysis of speech,’’ Eng. Appl.
The Netherlands: Elsevier, 2019, pp. 293–312. Artif. Intell., vol. 77, pp. 148–158, Jan. 2019.
[14] A. Graves, A.-R. Mohamed, and G. Hinton, ‘‘Speech recognition with deep [33] J. C. Vásquez-Correa, J. R. Orozco-Arroyave, and E. Nöth, ‘‘Convolu-
recurrent neural networks,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal tional neural network to model articulation impairments in patients with
Process., May 2013, pp. 6645–6649. Parkinson’s disease,’’ in Proc. Interspeech, Aug. 2017, pp. 314–318.
[15] J. R. Orozco-Arroyave, J. D. Arias-Londoño, J. F. Vargas-Bonilla, [34] F. Teixeira, A. Abad, and I. Trancoso, ‘‘Privacy-preserving paralinguistic
M. C. Gonzalez-Rátiva, and E. Nöth, ‘‘New Spanish speech corpus tasks,’’ in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP),
database for the analysis of people suffering from Parkinson’s disease,’’ May 2019, pp. 6575–6579.
in Proc. LREC, 2014, pp. 342–347. [35] T. Arias-Vergara, J. C. Vasquez-Correa, J. R. Orozco-Arroyave, P. Klumpp,
[16] J. A. Nichols, H. W. Herbert Chan, and M. A. B. Baker, ‘‘Machine learning: and E. Noth, ‘‘Unobtrusive monitoring of speech impairments of Parkin-
Applications of artificial intelligence to imaging and diagnosis,’’ Biophys. son’s disease patients through mobile devices,’’ in Proc. IEEE Int. Conf.
Rev., vol. 11, no. 1, pp. 111–118, Feb. 2019. Acoust., Speech Signal Process. (ICASSP), Apr. 2018, pp. 6004–6008.
[17] H. R. Pereira and H. A. Ferreira, ‘‘Classification of patients with [36] M. Stolar, M. Lech, R. S. Bolia, and M. Skinner, ‘‘Acoustic characteristics
Parkinson’s disease using medical imaging and artificial intelligence of emotional speech using spectrogram image classification,’’ in Proc. 12th
algorithms,’’ in Proc. Medit. Conf. Med. Biol. Eng. Comput., 2019, Int. Conf. Signal Process. Commun. Syst. (ICSPCS), Dec. 2018, pp. 1–5.
pp. 2043–2056. [37] J. I. Forcén, M. Pagola, E. Barrenechea, and H. Bustince, ‘‘Aggregation
[18] L. Moro-Velazquez, J. A. Gomez-Garcia, J. I. Godino-Llorente, J. Villalba, of deep features for image retrieval based on object detection,’’ in Proc.
J. Rusz, S. Shattuck-Hufnagel, and N. Dehak, ‘‘A forced Gaussians based Iberian Conf. Pattern Recognit. Image Anal., 2019, pp. 553–564.
methodology for the differential evaluation of Parkinson’s disease by [38] F. Nazir, M. N. Majeed, M. A. Ghazanfar, and M. Maqsood, ‘‘Mispro-
means of speech processing,’’ Biomed. Signal Process. Control, vol. 48, nunciation detection using deep convolutional neural network features and
pp. 205–220, Feb. 2019. transfer learning-based model for Arabic phonemes,’’ IEEE Access, vol. 7,
[19] P. Pérez-Toro, J. Vásquez-Correa, M. Strauss, J. Orozco-Arroyave, pp. 52589–52608, 2019.
and E. Nöth, ‘‘Natural language analysis to detect Parkinson’s [39] C. Castro, E. Vargas-Viveros, A. Sánchez, E. Gutiérrez-López, and
disease,’’ in Proc. Int. Conf. Text, Speech, Dialogue, 2019, D.-L. Flores, ‘‘Parkinson’s disease classification using artificial neural
pp. 82–90. networks,’’ in Proc. Latin Amer. Conf. Biomed. Eng., 2019, pp. 1060–1065.
[20] B. Karan, S. S. Sahu, and K. Mahto, ‘‘Parkinson disease pre- [40] M. Shu, ‘‘Deep learning for image classification on very small datasets
diction using intrinsic mode function based features from speech using transfer learning,’’ M.S. thesis, Iowa State Univ., Ames, IA, USA,
signal,’’ Biocybern. Biomed. Eng., vol. 40, no. 1, pp. 249–264, 2019.
Jan. 2020. [41] S. Afzal, M. Javed, M. Maqsood, F. Aadil, S. Rho, and I. Mehmood,
[21] A. Kacha, F. Grenez, J. R. Orozco-Arroyave, and J. Schoentgen, ‘‘Principal ‘‘A segmentation-less efficient alzheimer detection approach using hybrid
component analysis of the spectrogram of the speech signal: Interpretation image features,’’ in Handbook of Multimedia Information Security: Tech-
and application to dysarthric speech,’’ Comput. Speech Lang., vol. 59, niques and Applications. Springer, 2019, pp. 421–429.
pp. 114–122, Jan. 2020. [42] M. Z. Alom, T. M. Taha, C. Yakopcic, S. Westberg, P. Sidike, M. S. Nasrin,
[22] J. C. Vasquez-Correa, T. Arias-Vergara, J. R. Orozco-Arroyave, B. C Van Esesn, A. A S. Awwal, and V. K. Asari, ‘‘The history began from
B. Eskofier, J. Klucken, and E. Noth, ‘‘Multimodal assessment of AlexNet: A comprehensive survey on deep learning approaches,’’ 2018,
Parkinson’s disease: A deep learning approach,’’ IEEE J. Biomed. Health arXiv:1803.01164. [Online]. Available: [Link]
Inform., vol. 23, no. 4, pp. 1618–1630, Jul. 2019. [43] M. Maqsood, F. Nazir, U. Khan, F. Aadil, H. Jamal, I. Mehmood, and
[23] N. Garcia, J. R. Orozco-Arroyave, L. F. D’Haro, N. Dehak, and E. O.-Y. Song, ‘‘Transfer learning assisted classification and detection of
Nöth, ‘‘Evaluation of the neurological state of people with Parkin- Alzheimer’s disease stages using 3D MRI scans,’’ Sensors, vol. 19, no. 11,
son’s disease using i-vectors,’’ in Proc. Interspeech, Aug. 2017, p. 2645, Jun. 2019.
pp. 299–303. [44] C. M. Travieso, J. B. Alonso, J. R. Orozco-Arroyave, J. F. Vargas-Bonilla,
[24] P. Klumpp, T. Janu, T. Arias-Vergara, J. C. Vásquez-Correa, E. Nöth, and A. G. Ravelo-García, ‘‘Detection of different voice diseases
J. R. Orozco-Arroyave, and E. Nöth, ‘‘Apkinson—A mobile monitoring based on the nonlinear characterization of speech signals,’’ Expert Syst.
solution for Parkinson’s disease,’’ in Proc. Interspeech, Aug. 2017, Appl., vol. 82, pp. 184–195, Oct. 2017.
pp. 1839–1843. [45] K. Seppi, K. Ray Chaudhuri, M. Coelho, S. H. Fox, R. Katzenschlager,
[25] L. Moro-Velázquez, J. A. Gómez-García, J. I. Godino-Llorente, J. Villalba, and S. Perez Lloret, ‘‘Update on treatments for nonmotor symptoms of
J. R. Orozco-Arroyave, and N. Dehak, ‘‘Analysis of speaker recogni- Parkinson’s disease—An evidence-based medicine review,’’ Movement
tion methodologies and the influence of kinetic changes to automatically Disorders, vol. 34, pp. 180–198, Feb. 2019.
detect Parkinson’s disease,’’ Appl. Soft Comput., vol. 62, pp. 649–666, [46] H. Gunduz, ‘‘Deep learning-based Parkinson’s disease classification using
Jan. 2018. vocal feature sets,’’ IEEE Access, vol. 7, pp. 115540–115551, 2019.

35494 VOLUME 8, 2020


L. Zahid et al.: Spectrogram-Based Deep Feature Assisted Computer-Aided Diagnostic System for PD

LAIBA ZAHID received the B.S. degree (Hons.) HABIBULLAH JAMAL received the [Link]. degree
from the Comsats University Islamabad–Attock, in EE from the University of Engineering and
in 2017, where she is currently pursuing the M.S. Technology, Lahore, Pakistan, in 1974, and the
degree in computer science. Her research interests [Link]. and Ph.D. degrees in electrical engi-
include machine learning and signal processing. neering from the University of Toronto, Canada,
in 1979 and 1982, respectively.
He is currently a Professor of engineering sci-
ences with the Ghulam Ishaq Khan Institute,
Topi, Pakistan. He is the author of two textbooks
and 132 research articles. His research interests
include (but not limited to) signal processing, the design of microelectronic
MUAZZAM MAQSOOD received the Ph.D. circuits and development of novel computer architectures for telecommuni-
degree from UET Taxila, in 2017. He is currently cation, and national defense and other applications.
an Assistant Professor with COMSATS University Dr. Jamal was a recipient of prestigious national level awards 8th
Islamabad–Attock, Pakistan. His research interests TERADATA National IT Excellence Awards for Excellence in IT Education,
include medical imaging, machine learning’s, rec- in 12 April 2008; Performance Excellence in Engineering Awarded by
ommender systems, and image processing. The Institution of Engineers, Pakistan, in the Engineers Day 29 May 2007;
9th Pakistan Education Forum, National Education Award in 2003, and the
National Book Council of Pakistan Award in 1991.

MEHR YAHYA DURRANI graduated in informa-


tion technology. He received the master’s degree
in computer science. He is currently pursuing the
Ph.D. degree in pattern recognition. He is also
an Assistant Professor with COMSATS University IRFAN MEHMOOD is currently a Senior Lec-
Islamabad–Attock, Pakistan. turer with the University of Bradford, U.K. His
sustained contribution at various research and
industry-collaborative projects gives him an extra
edge to meet the current challenges faced in the
field of multimedia analytics. Specifically, he has
made significant contribution in the areas of video
summarization, medical image analysis, visual
MAHEEN BAKHTYAR received the master’s and surveillance, information mining, deep learning in
Ph.D. degrees from the Asian Institute of Tech- industrial applications, and data encryption.
nology (AIT), Thailand, with one year of research
experience from the National Institute of Infor-
matics, Tokyo, Japan. She is currently working
as an Assistant Professor with the Department of
Computer Science and Information Technology,
University of Balochistan, Pakistan. Her research
interests mainly include information/knowledge
management and retrieval, natural language pro-
cessing, sentiment analysis, text processing, language understanding, ques- OH-YOUNG SONG received the B.S., M.S., and
tion answering systems, and ontology processing. Ph.D. degrees from the School of Electrical Engi-
neering and Computer Science, Seoul National
University, South Korea, in 1998, 2000, and 2004,
JUNAID BABER received the M.S. and Ph.D. respectively. He was a Postdoctoral Fellow at the
degrees in computer science from the Asian Insti- School of Electrical Engineering and Computer
tute of Technology, Thailand. He has spent one Science, Seoul National University, from 2004 to
year as a Research Scientist with the National 2006. He is currently an Associate Professor with
Institute of Informatics, Tokyo. He is currently the Department of Software, Sejong University,
working as a faculty member with the University South Korea. His research interests include com-
of Balochistan, Quetta. His research interests lie puter graphics, simulation, and machine learning. Especially, he has made
in machine learning, high performance computing, a contribution in the areas of physics-based animation, human motion,
and data analytics. numerical algorithms, VR/AR, medical image analysis, and deep learning.

VOLUME 8, 2020 35495

You might also like