0% found this document useful (0 votes)
17 views9 pages

EEG-Based CNN for Driver Fatigue Detection

This article presents a novel EEG-based spatial-temporal convolutional neural network (ESTCNN) designed for evaluating driver fatigue, achieving a classification accuracy of 97.37%. The framework effectively extracts temporal dependencies and spatial features from EEG signals, outperforming traditional machine learning methods. The study highlights the importance of EEG signals for fatigue detection due to their rich physiological information and proposes practical applications in brain-computer interface systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views9 pages

EEG-Based CNN for Driver Fatigue Detection

This article presents a novel EEG-based spatial-temporal convolutional neural network (ESTCNN) designed for evaluating driver fatigue, achieving a classification accuracy of 97.37%. The framework effectively extracts temporal dependencies and spatial features from EEG signals, outperforming traditional machine learning methods. The study highlights the importance of EEG signals for fatigue detection due to their rich physiological information and proposes practical applications in brain-computer interface systems.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

This article has been accepted for inclusion in a future issue of this journal.

Content is final as presented, with the exception of pagination.

IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS 1

EEG-Based Spatio–Temporal Convolutional Neural


Network for Driver Fatigue Evaluation
Zhongke Gao , Xinmin Wang, Yuxuan Yang, Chaoxu Mu , Qing Cai, Weidong Dang , and Siyang Zuo

Abstract— Driver fatigue evaluation is of great importance for These observations indicate that fatigue working has caused
traffic safety and many intricate factors would exacerbate the great troubles in the transportation industry, demanding for
difficulty. In this paper, based on the spatial–temporal structure detecting techniques to precisely recognize fatigue states [3].
of multichannel electroencephalogram (EEG) signals, we develop
a novel EEG-based spatial–temporal convolutional neural net- Up until now, several types of human clues have been
work (ESTCNN) to detect driver fatigue. First, we introduce used for fatigue detection, including facial expressions [4],
the core block to extract temporal dependencies from EEG speech signals, and some physiological indexes such as elec-
signals. Then, we employ dense layers to fuse spatial fea- troencephalogram (EEG) data and dermal resistance [5]. How-
tures and realize classification. The developed network could ever, the nonphysiological sources such as face expressions
automatically learn valid features from EEG signals, which
outperforms the classical two-step machine learning algorithms. vary from different living habits and cultural backgrounds,
Importantly, we carry out fatigue driving experiments to collect thus may not reliable to a certain extent. In addition, some
EEG signals from eight subjects being alert and fatigue states. physiological indexes such as dermal resistance could be
Using 2800 samples under within-subject splitting, we compare affected by temperature and humidity. On consideration, EEG
the effectiveness of ESTCNN with eight competitive methods. signal is the chief source among these, which contains a
The results indicate that ESTCNN fulfills a better classification
accuracy of 97.37% than these compared methods. Furthermore, great deal of physiological information of a working brain.
the spatial–temporal structure of this framework advantages in It can truly reflect the pathological and mental states of
computational efficiency and reference time, which allows further a human body due to its good temporal resolution and
implementations in the brain–computer interface online systems. information richness [6]. Moreover, with the rapid devel-
Index Terms— Brain–computer interface (BCI), convolutional opment of wearable EEG devices and dry electrode tech-
neural network (CNN), deep learning (DL), electroencephalo- niques [7], various tasks can be easily implemented with EEG
gram (EEG), fatigue driving, spatio–temporal data. online systems and are promising for practical applications.
In view of these foundations, we focus on driver fatigue
I. I NTRODUCTION evaluation based on continuously recording multichannel
EEG signals.

M ENTAL fatigue refers to a complicated physiolog-


ical and psychological condition accompanied with
lessened alertness and decremental integrated performance.
It is noticed that the EEG signal is extremely weak with
low signal-to-noise ratios, which makes it fairly difficult
to develop computational algorithms for fatigue detection.
According to individual motivation, it is defined as a sensation
Against these odds, considerable studies have been conducted
of weariness with one’s unwillingness to carry on performing
on fatigue evaluation tasks. Subjects may be suffered from an
in the task and may lead to reduced work efficiency and
aggravated fatigue reaction in a monotonous and cumulative
increased accident possibility [1]. As reported by W orld
process [8], which has an effect on physiological indicators
H ealth Organi zati on [2], about 127 000 people were killed
in turn. A suitable fatigue threshold could be established
by road accidents annually, and about one-third of the victims
according to individual feedbacks and derivable indicators,
were aged from 15 to 29 years. Among the causes of the
and thus fatigue detection can be turned into a classification
accidents, fatigue driving devotes to a conservative estimate
exploration. Many novel methods have been proposed to
of above 10 000 deaths, which is a great threat to road safety.
extract valuable information from EEG signals, e.g., time–
Manuscript received November 30, 2017; revised July 11, 2018 and frequency analysis [9], complex network methods [10]–[12],
October 26, 2018; accepted December 2, 2018. This work was supported and nonlinear analysis [13]. Particularly, several studies have
in part by the National Natural Science Foundation of China under Grant given great ideas to distinguish subjects’ alert and fatigue
61873181, Grant 61473203, and Grant 61773284 and in part by the Nat-
ural Science Foundation of Tianjin, China under Grant 16JCYBJC18200. states [14]. In [15], four frequency features were extracted
(Corresponding author: Zhongke Gao.) across four principal frequency bands from EEG signals and
Z. Gao, X. Wang, Y. Yang, C. Mu, Q. Cai, and W. Dang are with the School then fed into the support vector machine (SVM), which
of Electrical and Information Engineering, Tianjin University, Tianjin 300072,
China (e-mail: zhongkegao@[Link]). achieved an excellent classification effect on fatigue detection.
S. Zuo is with the Key Laboratory of Mechanism Theory and Equipment In [16], an EEG-based system was presented to effectively
Design, Ministry of Education, Tianjin University, Tianjin 300072, China. detect drivers’ fatigue states by calculating four entropies as
Color versions of one or more of the figures in this paper are available
online at [Link] features and analyzing the effects of multiple entropy fusion.
Digital Object Identifier 10.1109/TNNLS.2018.2886414 To enhance the performance, the feature extraction processes
2162-237X © 2019 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission.
See [Link] for more information.
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

2 IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

cannot ignore the interference of electrooculogram (EOG) In another work [30], by using fast Fourier transform
and other physiological signal noises. In [17], a multimodal (FFT), EEG sequences were converted into 2-D meshlike
approach was proposed to estimate driver fatigue by combin- hierarchies and then fed into the convolutional recurrent neural
ing EEG and forehead EOG with adopted electrode place- network for human intended movement identification. The
ment, which targeted at integrating the temporal dependencies results confirmed the effectiveness of FFT-based DL methods
of both and received an improved performance. In [18], in EEG recognition tasks. Recently, by modifying the filter-
an independent component was used for source separation bank CSPs methods, EEG signals were turned into new tem-
from EEG signals, then coupled with autoregressive modeling poral representations and a CNN architecture was introduced
and Bayesian neural network, which worked effectively for for motor imagery EEG signal classification. The framework
fatigue identification. Overall, it is challenging to design a outperformed the existing results on the motor imagery data
single framework to overcome these problems and be available set [31]. These two studies converted EEG signals into new
for wide varieties of users. characteristics by using feature extraction methods, which can
According to international standards [19], EEG signals are be fed into DL frameworks with more concrete information.
recorded from multiple active electrodes attached to cerebral During the analysis procedure, they both focused on the
cortex with a fixed spatial arrangement. Among the EEG treatments of spatial relations and temporal dependencies from
classification methods mentioned above, the feature extraction EEG signals.
process of time–frequency analysis may neglect the valuable In this paper, we develop a novel framework, namely
information of electrode correlations from EEG signals. It can as EEG-based spatial–temporal CNN (ESTCNN), for driver
be concretely interpreted as a spatial dimension, and some fatigue detection. It attaches importance to temporal depen-
methods based on the common spatial patterns (CSPs) method dencies learning for each electrode and also strengthens spa-
were proposed to manage with spatial information. The CSP tial information extraction. The developed framework has an
method aims to find a set of spatial filters that can maximize effective performance on EEG signal classification tasks. First,
the between-class distance. Then, the computed relative energy we introduce the core block which has some advantages on
of the filtered channels is fed into linear classifiers for classifi- temporal dependencies extraction and then combine it with
cation. A regularized CSPs algorithm was presented as feature dense layers to meet the spatial–temporal information of EEG
extraction in [20], which significantly outperformed tradi- signals. Second, it constantly reduces the data dimension in
tional CSPs in motor imagery classification. Several successful the inference procedure, which gives rise to computational
attempts have been conducted to improve the CSP methods. efficiency and reference response. The proposed framework
The filter bank CSP method in [21] was extended to help has achieved the best performance on the fatigue data set
improve the performance of signal decomposition. In applied compared with eight competitive methods.
contexts, summarily, spatial–temporal filtering methods have To gain insight into fatigue detection, this paper is organized
an outstanding capacity to deal with EEG signal analysis. as follows. First, we introduce the core block and present the
Even though some CSP-based methods can gradually increase ESTCNN framework with the learning process and implemen-
the performances, they still have room for improvement in tation details. Second, we present a systematic introduction of
mapping relationships, especially the concern on temporal the fatigue experiment and the acquisition and preprocessing of
dependencies. EEG signals. Third, we present the overall performance of the
Recently, deep learning (DL) techniques have shown dis- model and compare it with eight competitive methods. Finally,
tinctive capabilities in object detection, speech recognition, we present conclusions with an outlook of future applications
and time series classification [22], [23]. These models possess on broader brain–computer interface (BCI) systems.
high computational efficiency and low model complexity.
In addition, many advanced methods based on DL techniques II. M ETHODS
have been explored and successfully converted into practical To build a proper network, it is important to consider
applications. Among these previous studies, DL methods have spatial relations and temporal dependencies using EEG sig-
also shown effective performances in the analysis of EEG nals. Due to its inherent distribution, the rational information
classification tasks. The DL framework should first reduce the along different dimensions could greatly enhance the model
dimension of EEG signals and then transform them into new performance and make the model more explanatory. We first
representations without any significant information loss. Some introduce the core block of the ESTCNN framework and
studies have applied convolutional neural networks (CNNs) then present the concrete model architecture and learning
into many kinds of EEG recognition tasks, such as emotion procedures. Finally, we give the implementation details.
recognition [24], P300 detection [25], effective radiated power
classification [26], and memory performance prediction [27].
In [28], a novel channel-wise CNN was proposed to evaluate A. Core Block
fatigue states on raw EEG data and ICA-transformed data, Several DL methods, such as convolutional layer and recur-
and it received an improved performance. These studies con- rent layer, have been proposed to deal with the dynamic
vey solid and meaningful design solutions for EEG signals. properties of EEG signals for excavating effective information
Furthermore, a detailed survey was presented in [29], which of partial snippets. Temporal convolution can be used as a
reviewed how to design and train CNNs without handcrafted feature extraction module to process time series, and when
features for EEG-based brain mapping. compared with recurrent models, it has better computational
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

GAO et al.: ESTCNN FOR DRIVER FATIGUE EVALUATION 3

Fig. 1. High-level pipeline of the ESTCNN framework.

efficiency on time series classification tasks [32]. We get TABLE I


inspired and propose this core block for two reasons in the D ETAILS OF THE ESTCNN F RAMEWORK . T HE S YMBOL [] C ONTAINS THE
K ERNEL S IZE , THE N UMBER OF F EATURE M APS , AND THE T YPE OF
task. First, convolution layers in the core block could process L AYER , R ESPECTIVELY
information in a hierarchical form, so be easier to capture
high-level features from EEG signals. Second, we also add
pooling layer to balance train accuracy and generalization.
The core block consists of three convolutional blocks and a
pooling layer. Each convolutional block orderly consists of
a 1 × 3 convolution, a rectified linear activation [33], and
a batch normalization (BN) [34], referring to [35]. We use
valid padding for convolution so that the output through
three convolutions has six fewer dimensions than the input.
In addition, three convolutions in each core block share the
same hyperparameters.
From the view of feature level, the core block can be
explained as follows. Each convolution layer in the core block
performs weighted feature combiners with nonlinear activa- are initially performed to handle input samples. Each pooling
tion. Then, the extracted features are cross channel convoluted layer goes only in temporal dimension without overlap, where
repeatedly in the next convolution layer. It contributes to the first two are set as max pooling with kernel size 2, and
extract the high-level information across the temporal dimen- the third is set as average pooling with kernel size 7. Two
sion. In addition, a pooling layer is followed to ease overfitting. dense layers, behind the third core block, are followed with
Here, we take this convolution structure as core block to make a Softmax classifier. In addition, no normalization is adopted
better predictions. This change makes it straight to understand in the data preprocessing process. The node numbers of two
the process on temporal dimension. dense layers are 50 and 2, respectively. In the convolutional
layers, kernel size is set as 3 with default stride and 16x filters,
B. Model Architecture where x is initialized to 1 and doubled every core block.
Fig. 1 shows the high-level pipeline of the proposed method
for EEG-based fatigue evaluation. The framework could learn C. Model Learning
effective information through the combination of the core A unit in the CNN is denoted by xl,k,(m,n) , where l is the
blocks and dense layers. layer, k is the feature map, and (m, n) is the position of the unit
Let matrix X ∈ R E·T serve as input samples (for 30 × in this feature map. Likewise, σl,k,(m,n) is denoted as the scalar
100 volumes) with paired labels Y ∈ R, where E denotes product between a group of input neurons. Then, xl,k,(m,n) can
recorded electrodes, T denotes sampling points, and Y varies be obtained as
from 0 to 1 as the associated label prediction. The proposed
xl,k,(m,n) = f (σl,k,(m,n) ) (1)
framework contains convolutional layers, pooling layers, and
dense layers. where f is the rectified linear units function [33] used for
In the experiment, we reach a network with depth selected whole network layers.
to 14 layers, which consists of three core blocks, two dense Notably, each neuron of feature maps in the convolutional
layers, and a softmax layer. The details of network structure layer shares the same set of weights, which aims to decrease
are presented in Table I. In this framework, three core blocks the number of weight parameters. They are attached to a subset
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

4 IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

of the neurons of the former layer, which depends on the where ω14,0,n is a threshold and N13 denotes the neuron
exact position of this neuron. Or rather, the neuron weights are number in layer L 13 . L 14 is fully connected to L 13 .
trained independently to their corresponding receptive fields. The layer aims to select the valid spatial information
In ESTCNN structure, the computation process of three core for classification.
blocks is quite similar, so we introduce the first block and other
reasoning parts. Let layer n denote L n , then the information D. Implementation
transmission process could be described as follows.
We take the robust normalization strategy in the fancy
1) For Layer L 1
14-layer network, adopting the CNN architecture from elec-

A trocardiogram arrhythmias detection in [36]. Hence, BN layers
σ1,k,(m,n) = ω1,k,0 + Im+i−1,n · ω1,k,i (2) are repeatedly employed to optimize such a deep network
i=1 manageably, and the stacking structure of the network is
where ω1,k,0 is a threshold and ω1,k, j denotes a set of used to enlarge the receptive field of high-level convolution
weights with 1 ≤ i ≤ A ( A = 3). Here, m corresponds kernels as well as the nonlinearity of model fitting. We use
to the convolutional kernel used in this framework, the classical backpropagation as a learning algorithm to tune
which amounts for temporal filters. With C + 1 weights up the thresholds and weights of the network [37], which is
for each map, this layer is designed to extract more valid reflected by the promotion of model accuracy on the validation
temporal features among all electrodes. set. The cross-entropy objective function is employed as a loss
2) For Layer L 2 function to estimate the model performance.

N1 
B During the process of training the models, the weight
σ2,k,(m,n) = ω2,k,0 + X 1,i,(m+ j −1,n) · ω2,k,i (3) of the convolutional layers is initialized with glorot_normal
i=1 j =1 initializer mentioned in [38]. We adopt the stochastic gradient
descent [39] optimizer with the learning rate 0.001 and other
where N1 denotes the number of feature maps in layer
default parameters. During the optimization for backpropaga-
L 1 and B denotes the kernel size of L 2 . This layer is
tion, we save the optimal model evaluated on the validation
employed to further process the temporal information
data set. The model is implemented on a workstation with
extract from layer L 1 by cross channel.
an Intel CPU (i7-6850k, 3.6 GHz) and an NVIDIA GPU
3) For Layer L 3
(Titan Xp) using the Keras library,1 which is extended from

N2 
C
Google Tensorflow.2 The typical duration for model training
σ3,k,(m,n) = ω3,k,0 + X 2,i,(m+ j −1,n) · ω3,k,i (4) is about 5 min.
i=1 j =1
where N2 denotes the number of feature maps in layer III. E XPERIMENT
L 2 and C denotes the kernel size of L 3 . This layer For the purpose of investigating the effective representations
is employed to extract the high-level information of of fatigue states, we elaborately designed a fatigue driving
temporal dependencies based on the feature maps from experiment for fatigue evaluation to collect EEG data, which
layer L 2 . could produce valuable original data sets. In our experiments,
4) For Layer L 4 we took full consideration about the time factor and subjects’
σ4,k,(m,n) = max(x 3,k,(i,n) , x 3,k,(i+1,n) ). (5) cooperative attitudes. In addition, we arranged the driving test
environment as real life as possible, which aimed to elicit
No parameter is occupied in this layer and k is fixed.
the strong physiological changes for subjects. In this section,
The pooling layer reduces the dimension of feature maps
we introduce the subject information, experiment protocol, and
by half, which aims to ease overfitting. The second and
data acquisition and preprocessing, respectively.
third core blocks (L 5 –L 8 and L 9 –L 12 ) follow the same
rules of the first core block (L 1 –L 4 ) and can be deduced
from it. A. Subjects
5) For Layer L 13 In the study, eight right-handed undergraduates (five males
and three females; mean: 22.73, standard deviation (STD):

N12 
D
σ13,n = ω13,0,n + X 12,i, j · ω13,n (6) 1.69) aged from 19 to 26 voluntarily performed the exper-
i=1 j =1 iment. None of them had any disorders related to psy-
chiatric. Subjects were required to refrain from antifatigue
where ω13,0,n is a threshold, N12 denotes the number
drinks or drowsiness-causing medications for 2 days before the
of feature maps and each has D neurons in layer L 12 .
experiment. Concurrently, they needed to keep reasonable rest
In addition, L 13 has n neurons and is fully connected
with sleep durations of more than 7 h per night. They should
to the flattened L 12 . This layer plays a role of making
comply with these regulations so as to join the study. As all
channel feature combinations.
subjects had no exposure to driving simulators, they were
6) For Layer L 14
asked to practice driving until they get skilled. Subjects were

N13
σ14,n = ω14,0,n + X 13,i · ω14,n (7) 1 [Link]
2 [Link]
i=1
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

GAO et al.: ESTCNN FOR DRIVER FATIGUE EVALUATION 5

transition, the subject went through another 30-min driving,


which was used as a fatigue state. After this stage, each subject
was asked to answer some questions and report postexperiment
comments. The recorded duration for each subject was nearly
90 min, which varied with individual differences. The timeline
of the experiment is shown in Fig. 3. Due to that, the exper-
iment procedure was a little tiresome, repetitive but mental
engaged; the fatigue level of subjects became increasingly
stronger during the experiment. As the vital physiological
measurements of fatigue states, their subjective assessments
and actual behaviors revealed the increasingly suffered fatigue.

C. Data Acquisition and Preprocessing


As a measurement of driver fatigue states, the EEG record-
ing device was equipped with the Neuroscan system with
40 electrodes at a sampling frequency of 1000 Hz, which
were arranged in accordance of the standard international
10/20 system. Before the acquisition, the skin impedance of
EEG electrodes was adjusted below 5 k by injecting conduc-
tive gel. Among these 40 electrodes, except for 4 electrodes as
the interior structure, 2 were defined as reference electrodes
and 4 (placed across horizontal and vertical directions) were
Fig. 2. Experimental settings. (a) Experimental scene from a researcher’s
used for monitoring eye movements. All subjects were advised
perspective. (b) Simulated driving system. (c) Forehead EOG placement from to possibly restrict unnecessary body movements, maintain a
Neuroscan system. constant speed, and avoid car collision during data collection.
Raw EEG signals were preprocessed with EEGLAB tool-
box [41]. We reduced the sampling rate to 100 Hz and
advised to stop driving at any moment during the experiment performed a bandpass filter of 1–50 Hz on the EEG signals
when any discomfort appeared. to remove artifacts since the power frequency interference is
above 50 Hz and some useless physiological noises are below
B. Experiment Protocol 1 Hz. After that, there were 30 channels for the preprocessed
EEG signals. We removed the 10-min transition period and
Driving on a real highway and simultaneously conducting obtained two pieces of about 30-min signals, labeled with
another task was highly perilous for subjects and other drivers. alert and fatigue categories. Then, we divided the signals
Therefore, the study was conducted in the Laboratory of Com- into samples of 1-s epochs without overlap and intentionally
plex Networks and Intelligent Systems, Tianjin University. selected larger class samples to make classes balanced. There
We used the driving simulator, PG F D001, equipped with a are 2800 samples per subject, of which the number of each
pedal, a steering wheel, and a clutch. In the virtual driving soft- category is 1400.
ware 3D I nstr uctor 2, we used the general car Phaeton2.0L
with automatic shifting by default. The simulated environment IV. R ESULT
was a monotonous expressway with few bends, sunny day,
and bare roadside scenery. Furthermore, we added a webcam In this section, to validate the performance of the proposed
360D618, a projector, and a stereo cabinet for better feel. The framework, we first give the overall performance evaluated
experimental settings are shown in Fig. 2. with 10-fold cross validation. Then, we make a comparison
To heighten the fatigue sensations of subjects, the trials with other common structures and competitive works to prove
began during 14:00–15:30, which was proven as an easy- the temporal superiority of our proposed structure. Finally,
trapped period for fatigue. Each trial lasted about 90 min. we analyze the impact of the ESTCNN framework on the
Full-course scalp EEG signals of subjects were collected in an spatial–temporal information with some brief discussions.
isolated and silent room. In addition, we also monitored sub-
jects’ facial states via a front-facing camera to verify fatigue A. Overall Performance
levels. According to Karolinska Sleepiness Scale that assesses The proposed ESTCNN is trained to detect fatigue states
from 1 (extremely alert) to 9 (very sleepy) [40], drivers’ fatigue for each subject. The individual performances are obtained on
states are divided into alert, mild fatigue, and fatigue for eight subjects by 10-fold cross validation. For each subject,
evaluation. Before the experiment, there were 10 min for scene 90% of the samples are randomly selected for training and
setup and about 20 min for driving practice. First, after a 3-min the remaining 10% are reserved for validation. We show the
survey, the subject would keep driving until he/she report recognition effects with classification accuracy and STD on
his/her mild fatigue, which generally lasted about 30 min as the validation sets. Fig. 4 presents the performance of the
the alert state. Then, after a 10 min continuously driving for ESTCNN framework on the fatigue data set.
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

6 IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

Fig. 3. Timeline of the experiment.

Actually, convolutional and recurrent layers are often


used to extract the temporal dependencies in deep net-
works [31], [42]. Therefore, we give three baselines to ver-
ify the spatial–temporal extraction capabilities of ESTCNN.
We also select other advanced EEG-based signal classification
methods, including some feature-based methods, to check the
performance of the proposed model. Here, we briefly introduce
some of the details for these compared models.
1) PSD-SVM: It extracted the EEG power spectrum den-
sity features and used SVM classifier to determine the
fatigue level.
2) CSP-SVM: It fed the relative energy of the filtered chan-
nels from CSPs methods into SVM classifiers, which
was used to test the results in [28].
3) LSTM: It used the deep long short-term memory
(LSTM) architecture for binary classification in [42],
Fig. 4. Overall performance of the ESTCNN framework. which consisted of two LSTM layers and a sigmoid
activation function.
4) CNN-A: It replaced each core block with a 1 × 7
From Fig. 4, we find that the ESTCNN model is stably convolution layer and a max-pooling layer, then added
effective on the whole data set and all the accuracies sur- a 1 × 3 convolutional layer and a global average pooling
pass 92%. The average accuracy achieves 97.37% with STD layer at the end.
3.30%. Among them, the accuracy of five subjects is over the 5) CNN-B: It replaced each 1 × 7 convolution layer with
mean accuracy. Except for some indeed uncontrollable factors, three 1 × 3 convolution layers in CNN-A model. Other
the accuracy difference is possibly caused by subjects’ physi- hyperparameters remained consistent.
cal conditions. By examining the recorded experiment videos, 6) CNN-C: It developed a four-layer CNN for spatial
we found that S4 showed slight fatigue at the beginning of the feature fusion and temporal feature extraction on steady-
experiment, but reported it after 30 min of the experiment. Just state visual evoked potential classification in [43].
the reverse, S7 almost stayed alert throughout the experiment 7) CNN-D: It proposed a novel channel-wise CNN with
but reported slight fatigue half an hour before the end of the raw EEG data on driver’s cognitive performance predic-
experiment. These two subjects did not well report the degree tion tasks in [28].
of subjective fatigue, which accounts for the lower accuracies 8) FFT-CNN: It transformed EEG signals into multispectral
of the two subjects. It suggests that, during carrying out the images and trained with a combined CNN and LSTM
experiments, we should inform the subjects of reporting actual framework for classification in [30].
individual situations, which reflects that the proposed method
For an effective comparison, we select the eight above-
has the robust ability to learn effective information from two
mentioned methods, which are some baselines and typi-
categories of samples for recognition.
cal works from different perspectives. There are some con-
ventional methods for EEG analysis, such as PSD-SVM
in [44]. It extracted power spectrum density features to
B. Method Comparison
decode time series while neglecting the spatial informa-
From the overall performance, we find that our proposed tion of EEG signals. CSP-SVM considered the spatial pat-
method has effective performances, which should be due to terns to discover discriminative information. DL methods
the spatial–temporal structure of the framework. In order to brought some successful attempts to improve recognition
deeply explore the spatial–temporal ability of the method performance taking full consideration for spatial–temporal
on recognition tasks, we make a comparison with several properties. Therefore, we give three baselines to verify
commonly used structures and other existing works. the superiority of our proposed framework with details
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

GAO et al.: ESTCNN FOR DRIVER FATIGUE EVALUATION 7

TABLE II
S TRUCTURES OF T HREE BASELINE M ETHODS . T HE U SAGE
OF S YMBOL [] I S S AME AS A BOVE

Fig. 5. Recognition accuracies of the compared methods.

CNN-B by over 4%. It suggests that the overall structure is


shown in Table II, including LSTM, CNN-A, and CNN-B. beneficial for EEG-based fatigue evaluation, of which the core
We also choose five competitive works for further com- block could enhance temporal information extraction.
parison, including PSD-SVM, CSP-SVM, CNN-C, CNN-D, Some DL and non-DL methods have been conducted to
and FFT-CNN. From the above, we select two feature-based explore different kinds of EEG analysis tasks. We employ
methods and six DL frameworks to test and evaluate our the 10-fold cross validation to estimate the classification
developed algorithm. effects among ESTCNN and other competitive methods.
According to the structure parameters presented in the Frequency features are often used in the above-mentioned
original papers, we reproduce these methods to analyze our studies, including PSD-SVM and FFT-CNN. Meanwhile,
EEG data set. Note that we use the open-source code released CSP method is also used to investigate stable patterns, so we
from the FFT-CNN method. These above-mentioned are all take CSP-SVM into comparison. PSD-SVM gets the mean
designed for EEG signal analysis in the original papers. accuracy of 70.45% while CSP-SVM has a better performance
We apply these methods to all eight subjects with 10-fold cross of 73.29%. However, these two methods are worse than other
validation and receive the mean accuracy of each method. DL-related methods, which reflects the robust capacity of
Fig. 5 shows the recognition accuracies of the compared DL methods on learning representations. The classification
methods. accuracy of these DL methods varies between 87% and
In order to investigate the importance of extracting temporal 97%. Among the methods using the EEG signals as input,
information, we take three baseline methods LSTM, CNN-A, CNN-B and ESTCNN work well in the whole data set.
and CNN-B for comparison. It is proven that the receptive Moreover, some prior knowledge-based DL methods also
field of a 5 × 5 kernel is equal to the receptive field of two benefit a lot from prior information. FFT-CNN combined
stacked 3 × 3 kernels [45]. We replace the large 7 × 7 kernel with FFT has a performance of over 94%. Note that in terms
with three small 3×3 kernels. The small kernel has advantages of STD, the STD of these eight methods varies from 9.84%
in many aspects, such as increasing the nonlinearity of model to 3.28%. CNN-D receives the maximum variance of 9.84%
fitting and reducing the occupied memory of model. Thus, and ESTCNN has the STD of 3.3%.
we build the third baseline method CNN-B. From Fig. 5, Among these eight methods, ESTCNN shows the best
LSTM and CNN-A perform well and achieve 91.34% and performance for driving fatigue evaluation tasks with consider-
87.34%, respectively. It indicates that convolutional and recur- able advantages compared with other methods. This suggests
rent layers are capable of extracting temporal information from that our ESTCNN framework can robustly capture effective
EEG signals. By replacing large kernels with small kernels, information using EEG signals from eight subjects, due to its
CNN-B contributes to a slight increase in its performance, good abilities for temporal dependencies extraction and spatial
which reaches 92.73%. features fusion. Overall, the ESTCNN framework delivers an
The classification accuracy of all three baseline methods excellent performance on the fatigue data set.
exceeds 87% with some differences, in which LSTM outper-
forms CNN-A by 4%. After increasing the model nonlinearity, C. Discussion
CNN-B has a better performance than LSTM. Note that we The above-mentioned comparison analysis shows that our
employ global average pooling layer as the final layer to ESTCNN framework performs best than other methods under
average every channel into a value in the CNN baselines. In the the same validation scheme. As the ESTCNN can be taken
model modification, we reduce the degree of signal compres- as an integration of the core blocks and dense layers, it is
sion by using two values to represent each channel and receive novel to investigate the importance of the core block for
a considerable improvement. The results indicate that two extracting temporal dependencies and dense layer for spatial
representative values are more robust to retain effective infor- feature fusion. Comparing LSTM and CNN-A with ESTCNN,
mation across spatial dimensions. As the adjusted model, our the core block in the ESTCNN framework performs better
ESTCNN provides an accuracy of 97.37%, which surpasses than conventional DL layers, while these layers or blocks
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

8 IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS

are targeted for extracting temporal dependencies. Meanwhile, [6] J. A. Horne and S. D. Baulk, “Awareness of sleepiness when driving,”
comparing CNN-B with ESTCNN, the accuracy of ESTCNN PsychoPhysiology, vol. 41, no. 1, pp. 161–165, Jan. 2004.
[7] Y. M. Chi, Y.-T. Wang, Y. J. Wang, C. Maier, T.-P. Jung, and
is about 5% higher than CNN-B with lower deviations. These G. Cauwenberghs, “Dry and noncontact eeg sensors for mobile brain–
improvements of performance suggest the effectiveness of our computer interfaces,” IEEE Trans. Neural Syst. Rehabil. Eng., vol. 20,
proposed framework which learns both spatial and temporal no. 2, pp. 228–235, Mar. 2012.
[8] S. K. L. Lal and A. Craig, “A critical review of the psychophysiology of
information using EEG signals. driver fatigue,” Biol. Psychol., vol. 55, no. 3, pp. 173–194, Feb. 2001.
In the model evaluation, two typical indicators are the [9] C.-T. Lin et al., “Adaptive EEG-based alertness estimation system by
validate accuracy and standard variance. In comparison, our using ICA-based fuzzy neural networks,” IEEE Trans. Circuits Syst.
I, Reg. Papers, vol. 53, no. 11, pp. 2469–2476, Nov. 2006.
proposed ESTCNN method has considerable advantages in [10] Z.-K. Gao, Q. Cai, Y.-X. Yang, N. Dong, and S.-S. Zhang, “Visibility
accuracy and variance. By applying the core blocks and dense graph from adaptive optimal kernel time-frequency representation for
layers to reduce temporal dimensions, the network is designed classification of epileptiform EEG,” Int. J. Neural Syst., vol. 27, no. 4,
p. 1750005, Jun. 2017.
to learn representations from multichannel EEG signals. Due [11] Z. K. Gao, M. Small, and J. Kurths, “Complex network analysis of time
to the data-driven nature of the model, the developed ESTCNN series,” Europhys. Lett., vol. 116, no. 5, p. 50001, Jan. 2017.
model keeps enough flexibility to address many kinds of EEG- [12] Z. K. Gao et al., “An adaptive optimal-Kernel time-frequency
based recognition tasks. Compared to the above-mentioned representation-based complex network method for characterizing
fatigued behavior using the SSVEP-based BCI system,” Knowl.-Based
approaches, ESTCNN model requires less preprocessing on Syst., vol. 152, pp. 163–171, Jul. 2018.
the multichannel data and is more convenient to be imple- [13] L.-L. Chen, Y. Zhao, J. Zhang, and J.-Z. Zou, “Automatic detection
mented in the BCI online system. of alertness/drowsiness from physiological signals using wavelet-based
nonlinear features and machine learning,” Expert Syst. Appl., vol. 42,
no. 21, pp. 7344–7355, Nov. 2015.
V. C ONCLUSION [14] G. Borghini, L. Astolfi, G. Vecchiato, D. Mattia, and F. Babiloni, “Mea-
suring neurophysiological signals in aircraft pilots and car drivers for
Fatigued driving is a social problem that requires more the assessment of mental workload, fatigue and drowsiness,” Neurosci.
attentions, and DL methods could promote the development of Biobehav. Rev., vol. 44, pp. 58–75, Jul. 2014.
[15] V. M. Y. Mervyn, X. Li, K. Shen, and E. P. V. Wilder-Smith, “Can SVM
computational models for detecting task. In summary, we have be used for automatic EEG detection of drowsiness during car driving?”
developed a spatial–temporal CNN to detect driver fatigue Saf. Sci., vol. 47, no. 1, pp. 115–124, Jan. 2009.
from EEG signals. The proposed method shows significant [16] J. Min, P. Wang, and J. Hu, “Driver fatigue detection through multiple
improvements in the model performance and can learn more entropy fusion analysis in an EEG-based system,” PLoS ONE, vol. 12,
no. 12, p. e0188756, Dec. 2017.
robust representations from EEG signals. The vital procedures [17] W. L. Zheng and B.-L. Lu, “A multimodal approach to estimating
are in two parts: first, we introduce the core block to deal vigilance using EEG and forehead EOG,” J. Neural Eng., vol. 14, no. 2,
with the information on the temporal dimension. Second, p. 026017, Feb. 2017.
[18] R. Chai et al., “Driver fatigue classification with independent component
we utilize the dense layer to fuse the spatial features among by entropy rate bound minimization analysis in an EEG-based system,”
the electrodes. We investigate the importance of spatial infor- IEEE J. Biomed. Health Inform., vol. 21, no. 3, pp. 715–724, May 2017.
mation and temporal dependencies by giving three baseline [19] G. H. Klem, H. O. Lüeders, H. H. Jasper, and C. Elger, “The ten-twenty
electrode system of the international federation,” Electroencephalogr.
methods and five competitive studies. The results show that Clin. Neurophysiol., vol. 52, no. 3, pp. 3–6, 1999.
the performance is greatly improved due to the involvement [20] F. Lotte and C. Guan, “Regularizing common spatial patterns to improve
of spatial–temporal information in EEG-based classification BCI designs: Unified theory and new algorithms,” IEEE Trans. Biomed.
Eng., vol. 58, no. 2, pp. 355–362, Feb. 2010.
tasks. [21] K. K. Ang, Z. Y. Chin, H. Zhang, and C. Guan, “Filter bank common
It would be a great potential to extend the proposed method spatial pattern (FBCSP) in brain-computer interface,” in Proc. IEEE Int.
to numerous areas, such as multisource information fusion Joint Conf. Neural Netw., Hong Kong, Jun. 2008, pp. 2390–2397.
[22] Y. Bengio, A. Courville, and P. Vincent, “Representation learning:
tasks. Further works would focus on the combination with A review and new perspectives,” IEEE Trans. Pattern Anal. Mach. Intell.,
feature-based methods to improve the model performance. vol. 35, no. 8, pp. 1798–1828, Aug. 2013.
Look forward, with the effectiveness and generality of the [23] S. Lin and G. C. Runger, “GCRNN: Group-constrained convolutional
ESTCNN model, we expect it to be useful for broader appli- recurrent neural network,” IEEE Trans. Neural Netw. Learn. Syst.,
vol. 29, no. 10, pp. 4709–4718, Oct. 2018.
cations in BCI systems. [24] Y. X. Yang et al., “A recurrence quantification analysis-based channel-
frequency convolutional neural network for emotion recognition from
EEG,” Chaos, vol. 28, no. 8, p. 085724, Aug. 2018.
R EFERENCES [25] H. Cecotti and A. Graser, “Convolutional neural networks for P300
[1] M. Tops and M. A. S. Boksem, “Absorbed in the task: Personality detection with application to brain-computer interfaces,” IEEE Trans.
measures predict engagement during task performance as tracked by Pattern Anal. Mach. Intell., vol. 33, no. 3, pp. 433–445, Mar. 2011.
error negativity and asymmetrical frontal activity,” Cogn. Affect. Behav. [26] H. Cecotti, M. P. Eckstein, and B. Giesbrecht, “Single-trial classification
Neurosci., vol. 10, no. 4, pp. 441–453, Dec. 2010. of event-related potentials in rapid serial visual presentation tasks using
[2] F. Racioppi, L. Eriksson, C. Tingvall, and A. Villaveces, Preventing supervised spatial filtering,” IEEE Trans. Neural Netw. Learn. Syst.,
Road Traffic Injury: A Public Health Perspective for Europe. Geneva, vol. 25, no. 11, pp. 2030–2042, Nov. 2014.
Switzerland: World Health Organ, 2004. [27] X. Sun, C. Qian, Z. Chen, Z. Wu, B. Luo, and G. Pan, “Remembered or
[3] D. F. Dinges, “An overview of sleepiness and accidents,” J. Sleep Res., forgotten?—An EEG-based computational prediction approach,” PLoS
vol. 4, no. 2, pp. 4–14, Dec. 1995. ONE, vol. 11, no. 12, p. e0167497, Dec. 2016.
[4] M. Karchani et al., “Presenting a model for dynamic facial expression [28] M. Hajinoroozi, Z. Mao, T.-P. Jung, C.-T. Lin, and Y. Huang, “EEG-
changes in detecting drivers’ drowsiness,” Electron. Phys., vol. 7, no. 2, based prediction of driver’s cognitive performance by deep convolutional
pp. 1073–1077, Apr./Jun. 2015. neural network,” Signal Process., Image Commun., vol. 47, pp. 549–555,
[5] M. Fallahi, M. Motamedzade, R. Heidarimoghadam, A. R. Soltanian, Sep. 2016.
and S. Miyake, “Effects of mental workload on physiological and [29] R. T. Schirrmeister et al., “Deep learning with convolutional neural
subjective responses during traffic density monitoring: A field study,” networks for EEG decoding and visualization,” Hum. Brain Mapping,
Appl. Ergonom., vol. 52, pp. 95–103, Jan. 2016. vol. 38, no. 11, pp. 5391–5420, Nov. 2017.
This article has been accepted for inclusion in a future issue of this journal. Content is final as presented, with the exception of pagination.

GAO et al.: ESTCNN FOR DRIVER FATIGUE EVALUATION 9

[30] P. Bashivan, I. Rish, M. Yeasin, and N. Codella, “Learning representa- Yuxuan Yang received the bachelor’s degree in
tions from EEG with deep recurrent-convolutional neural networks,” in automation from Anhui University, Hefei, China,
Proc. (ICLR), San Juan, Puerto Rico, Feb. 2016, pp. 1–15. in 2014, the master’s degree in automation from the
[31] S. Sakhavi, C. Guan, and S. Yan, “Learning temporal information for School of Electrical and Information Engineering,
brain-computer interface using convolutional neural networks,” IEEE Tianjin University, Tianjin, China, in 2017, where
Trans. Neural Netw. Learn. Syst., vol. 29, no. 11, pp. 5619–5629, she is currently pursuing the Ph.D. degree with the
Nov. 2018. School of Electrical and Information Engineering.
[32] C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional Her current research interests include brain–
networks: A unified approach to action segmentation,” in Proc. Eur. computer interface, machine learning, and complex
Conf. Comput. Vis. (ECCV), Amsterdam, The Netherlands, Oct. 2016, networks.
pp. 47–54.
[33] V. Nair and G. E. Hinton, “Rectified linear units improve restricted
boltzmann machines,” in Proc. 27th Int. Conf. Mach. Learn., Haifa,
Israel, 2000, pp. 807–814.
[34] S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep
network training by reducing internal covariate shift,” in Proc. Int. Conf.
Mach. Learn., Lille, France, Mar. 2015, pp. 448–456.
[35] K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep Chaoxu Mu received the B.S. degree from the
residual networks,” in Proc. Eur. Conf. Comput. Vis. (ECCV), Zürich, Harbin Institute of Technology, Harbin, China,
Switzerland, Oct. 2016, pp. 630–645. in 2006, the M.S. degree from Hohai University,
[36] P. Rajpurkar, A. Y. Hannun, M. Haghpanahi, C. Bourn, and A. Y. Ng. Nanjing, China, in 2009, and the Ph.D. degree
(2017). “Cardiologist-level arrhythmia detection with convolutional in control science and engineering from Southeast
neural networks.” [Online]. Available: [Link] University, Nanjing, in 2012.
[37] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning repre- Her current research interests include computa-
sentations by back-propagating errors,” Nature, vol. 323, pp. 533–536, tional intelligence, nonlinear system control and
Oct. 1986. optimization, adaptive and learning systems, smart
[38] X. Glorot and Y. Bengio, “Understanding the difficulty of training deep grid, and machine learning.
feedforward neural networks,” in Proc. 13th Int. Conf. Artif. Intell.
Statist., Sardinia, Italy, Mar. 2010, pp. 249–256.
[39] L. Bottou, “Large-scale machine learning with stochastic gradient
descent,” in Proc. COMPSTAT, Paris, France, 2010, pp. 177–186.
[40] T. Åkerstedt and M. Gillberg, “Subjective and objective sleepiness in
the active individual,” Int. J. Neurosci., vol. 52, nos. 1–2, pp. 29–37,
Jul. 1990.
[41] A. Delorme and S. Makeig, “EEGLAB: An open source toolbox for Qing Cai received the [Link]. degree in automation
analysis of single-trial EEG dynamics including independent component from Liaoning Shihua University, Liaoning, China,
analysis,” J. Neurosci. Methods, vol. 134, no. 1, pp. 9–21, Mar. 2004. in 2014. She is currently pursuing the Ph.D. degree
[42] R. G. Hefron, B. J. Borghetti, J. C. Christensen, and C. M. S. Kabban, with the School of Electrical and Information Engi-
“Deep long short-term memory structures model temporal dependen- neering, Tianjin University, Tianjin, China.
cies improving cognitive workload estimation,” Pattern Recognit. Lett., Her current research interests include EEG, fMRI,
vol. 94, pp. 96–104, Jul. 2017. brain network, and complex networks.
[43] N.-S. Kwak, K.-R. Müller, and S.-W. Lee, “A convolutional neural
network for steady state visual evoked potential classification under
ambulatory environment,” PLoS ONE, vol. 12, no. 2, p. e0172578,
Feb. 2017.
[44] X. Zhang et al., “Design of a fatigue detection system for high-
speed trains based on driver vigilance using a wireless wearable EEG,”
Sensors, vol. 17, no. 3, p. 486, Mar. 2017.
[45] C. Szegedy et al., “Going deeper with convolutions,” in Proc. IEEE
Int. Conf. Comput. Vis. Pattern Recognit. (CVPR), Boston, MA, USA,
Jun. 2015, pp. 1–9. Weidong Dang received the [Link]. degree in automa-
tion from Tianjin University, Tianjin, China, in 2016,
where he is currently pursuing the Ph.D. degree
Zhongke Gao received the [Link]. and Ph.D. degrees with the School of Electrical and Information
from Tianjin University, Tianjin, China, in 2007 and Engineering.
2010, respectively. His current research interests include sensor
Since 2016, he has been a Full Professor with the design, multisource information fusion, measure-
School of Electrical and Information Engineering, ment science and technology, multiphase flow, and
Tianjin University, where he is currently the Director complex networks.
of the Laboratory of Complex Networks and Intelli-
gent Systems. His current research interests include
deep learning, EEG analysis, complex networks,
brain–computer interface, and wearable intelligent
devices.

Xinmin Wang received the bachelor’s degree in Siyang Zuo received the [Link]. and Ph.D. degrees
automation from the School of Electrical and Infor- in information science and technology from the
mation Engineering, Tianjin University, Tianjin, University of Tokyo, Tokyo, Japan, in 2009 and
China, in 2017, where he is currently pursuing the 2013, respectively.
master’s degree in control science and engineering He is currently a Professor with the Key Labora-
with the School of Electrical and Information Engi- tory of Mechanism Theory and Equipment Design,
neering. Ministry of Education, School of Mechanical Engi-
His current research interests include brain– neering, Tianjin University, Tianjin, China. His cur-
computer interface, EEG analysis, and machine rent research interests include medical robotics and
learning. imaging techniques.

You might also like