0% found this document useful (0 votes)
7 views13 pages

Schizophrenia Classification Using FMRI Data Based

Uploaded by

rishavchand02
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views13 pages

Schizophrenia Classification Using FMRI Data Based

Uploaded by

rishavchand02
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

This article has been accepted for publication in a future issue of this journal, but has not been

fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Date of publication xxxx 00, 0000, date of current version xxxx 00, 0000.
Digital Object Identifier 10.1109/[Link]

Schizophrenia Classification using fMRI


Data Based on a Multiple Feature Image
Capsule Network Ensemble
∗ ∗ ∗
BO YANG1, ,YUAN CHEN1, ,QUAN-MING SHAO1, ,RUI YU1 ,WEN-BIN LI1 ,GUAN-QI
† †
GUO1 ,JUN-QIANG JIANG1, ,LI PAN1,
1
School of Information Science and Engineering, Hunan Institute of Science and Technology, Yueyang, 414006, China

Corresponding author: JUN-QIANG JIANG(jjq@[Link]) and LI PAN([Link]@[Link])
This work was partially funded by the Natural Science Foundation of Hunan Province (Grant No. 2019JJ40105, 2018JJ2153, 2017JJ2016,
2018JJ2152), the Science and Technology Program of Hunan Province (Grant No. 2016TP1021), the Scientific Research Fund of Hunan
Provincial Education Department (Grant No. 15A079, 17A089), and the Scientific Research Innovation Project for Postgraduate of Hunan
Province(Grant No. CX2018B777).

BO YANG, YUAN CHEN and QUAN-MING SHAO contributed equally to this work.

ABSTRACT Automatic diagnosis and classification of schizophrenia based on functional magnetic


resonance imaging (fMRI) data have attracted increasing attention in recent years. Most previous studies
abstracted highly compressed functional features from the view of brain science and fed them into shallow
classifiers for this purpose. However, their classification performance in practical applications is unstable
and unsatisfactory. As an acute psychotic disorder, schizophrenia shows functional complexity in fMRI
data. Therefore, additional features and deep classification methods are needed to improve classification
performance. In this study, we propose a multiple feature image capsule network ensemble approach for
schizophrenia classification. The proposed approach proceeds in three steps: 1) extracting multiple image
features from the perspective of linear sparse representation, nonlinear multiple kernel representation, and
function connection of brain areas respectively; 2) feeding these image features into three specially designed
independent capsule networks for classification; 3) obtaining the final results by fusing the outputs of these
three deep capsule network using a ensemble approach. To further improve the classification performance,
we design a optimization model of maximizing the square of correlation coefficients and propose a weighted
ensemble technology based on this model, which is mathematically proved to be solved as a eigenvalue
decomposition problem in certain case. Finally, the proposed approach is implemented and evaluated on
the schizophrenia fMRI dataset from COBRE, UCLA and WUSTL. From the experimental results, we
conclude that the proposed method outperforms some current methods and further improves the accuracy
of schizophrenia classification.

INDEX TERMS Schizophrenia Classification, Multiple Features Extraction, Deep Capsule Network,
Classifier Ensemble

I. INTRODUCTION greatly to the improved analysis and diagnosis of schizophre-


A. BACKGROUND nia in recent years [7]–[11].
Chizophrenia is a devastating mental disease with ex- As a complex psychiatric disorder, schizophrenia shows
S traordinary complexity. Diagnosis of schizophrenia with
high confidence is important in neurosciences and medical
local abnormalities of brain activity and functional connec-
tivity networks in the schizophrenic brain feature disrupted
science [1], [2]. High-resolution brain imaging techniques, topological properties [12]. As a gold-standard functional
such as functional magnetic resonance imaging (fMRI) [3], imaging technique in neurosciences, fMRI has become the
structural magnetic resonance imaging(sMRI) [4], diffu- most widely used imaging technique among all the above
sion tensor imaging(DTI) [5], positron emission tomogra- mentioned imaging technologies for the analysis and diagno-
phy(PET) [6], facilitate understanding of the structure and sis of schizophrenia [13]. To free imaging specialists from the
function of human brain. These techiques have contributed heavy task of interpreting fMRI images, methods to diagnose

VOLUME 4, 2016 1

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

schizophrenia automatically with high reliability need to be suitable for schizophrenia classification with small sample
developed to simplify analysis of fMRI images. Different ma- size should also be determined.
chine learning methods, including classification and feature
extraction, have been introduced in recent years to realize the C. OUR CONTRIBUTIONS
automatic analysis and diagnosis of schizophrenia [14]. In view of these problems, this paper proposes a deep learn-
From the perspective of statistical machine learning, ing approach for schizophrenia classification from several
schizophrenia analysis and diagnosis can be regarded as a perspectives, such as multiple features extraction, deep net-
typical statistical classification task. Similar to other com- work architecture design and classifier ensemble, to improve
mon statistical classification approaches, most of existing the accuracy and stability of classification. The details of our
approaches for schizophrenia classification based on fMRI contributions are listed as below:
have the same processing steps. After some data preprocess- (1)We present a novel whole schizophrenia classification
ing operations, these existing approaches can be divided into approach which contains 3 steps: multiple features extrac-
three main steps: 1) extracting features from original fMRI tion, deep capsule network training and multiple classifiers
data based on functional region of interests(ROI) informa- weighed ensemble. In our experiments, the approach outper-
tion; 2) transforming ROI features into compressed features forms some current methods.
using linear or nonlinear feature transformation techniques; (2)We present multiple fMRI features extraction method
3) optimizing a classifier by training it on these compressed and extract linear sparse feature image, nonlinear multiple k-
features. Recently proposed approaches using deep learning ernel feature image and functional connectivity feature image
[19]–[21] involve concise operations in which the last two from three different perspectives: statistical linear analysis,
steps are merged. And automatic feature extraction is realized statistical nonlinear analysis and brain regions analysis. This
by end-to-end learning. multiple fMRI features extraction method can extract more
useful information from the view of data science and brain
science separately.
B. MOTIVATION
(3)We introduce capsule network into schizophrenia di-
Despite the continuous advances in schizophrenia classifica- agnosis and design new capsule network architecture more
tion approaches based on fMRI data, existing approaches are suitable for schizophrenia diagnosis. At the basis of original
still in the infancy stage, and their diagnosis levels are con- capsule network, we add more convolution layers to enlarge
siderably lower than those of human experts. The research the local receptive field and cancel the RELU nonlinear map-
of schizophrenia classification based on fMRI data still faces ping layers to control network capacity, which are effective
several inevitable problems and challenges. for solving over-fitting.
First, little is presently known about the structures and (4)We present a weighted ensemble method to complete
functions of human brain. In specific, the etiology of the final classifier ensemble of our schizophrenia classifica-
schizophrenia and the abnormal modes of the brains of pa- tion approach. This weighted ensemble method based on an
tients with this condition remain unclear [14]. As mentioned optimization model maximizes the square of correlation co-
above, existing approaches obtain ROI features from current efficients. We discuss and prove in theory that in certain case
knowledge about brain regions. The only available materi- the presented optimization model is essentially a eigenvalue
als in schizophrenia classification are fMRI data and prior decomposition problem, which means we can obtain the best
knowledge about this condition. Therefore, a method to ex- weights rapidly by the common eigenvalue decomposition
tract deep information from fMRI data by using information algorithms.
procesing technologies has become increasingly important. The rest of the paper is organized as follows: In section
Second, most existing schizophrenia classification ap- II, we review the related literature. In section III, we present
proaches can be classified as shallow learning. Intuitionally, multiple feature image capsule networks ensemble approach
deep learning approaches [15]–[18] that are suitable for com- and describe it in details. In Section IV, the experimental
plex tasks should be used for highly complex schizophrenia results on a multi-site schizophrenia fMRI dataset from the
classification tasks. Besides, recently proposed approaches center for biomedical research excellence(COBRE, available
using deep learning [19]–[21] involve concise operations at [Link] the university of California, Los An-
and automatic feature extraction is realized by end-to-end gles(UCLA, available at [Link] and the conte
learning. However, mainstream deep learning methods are center for the neuroscience of mental disorders at Washington
data hungry. In addition, these methods are easily over-fitting university school of medicine in St. Louis(WUSTL, available
and show poor performance in tasks with a small sample at [Link] are presented and discussed by com-
size. Unfortunately, the number of samples in a schizophrenia paring with some other representative methods. In Section V,
classification research is usually small. As listed in the paper we conclude the paper.
[48], the mean and median sample size of 21 researches on
schizophrenia classification in the recent 5 years(2014-2018) II. RELATED WORK
are 208 and 147 respectively. Hence, how to choose appropri- The researches of schizophrenia classification based on fMRI
ate deep learning methods and design network architecture data have made great progress in recent years. AS a clas-
2 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

FIGURE 1: Illustration of multiple feature image capsule networks ensemble

sification task, one of the most important things is choos- [47]. Because of the powerful learning ability and automatic
ing a good classifier. So far many linear and nonlinear features extraction ability of deep learning, these approaches
classifiers have been applied in schizophrenia classification, are hopeful to improve classification performance of brain
which contain bayesian classifier [22], c-means classifier disease. However, due to the relatively small number of sam-
[23], k-nearest neighbor algorithm [24], artificial neural net- ples [48], so far the deep neural networks widely used in clas-
work [25], least square classifier [26], support vector ma- sical computer vision tasks still have not shown distinguished
chine(SVM) [27] and extreme learning machine(ELM) [28], classification power. As mentioned in the paper [48], there
[29]. are still only 18 papers published in recent 5 years(2014–
The works mentioned above are generally based on ROIs 2018) using deep neural network for classification of various
information extracted from the original fMRI data. As point- disorders(less than 4%, 18/209 papers) and SVM, a famous
ed out in the paper [7], the dimension of functional con- shallow learning method, is still the most popular method-
nectivity is large even if ones only evaluate connectivity s(more than 55%, 117/209 papers). Hence, how to improve
between defined ROIs. To avoid overfitting problem, many the generalization of deep learning in classification of brain
feature transformation technologies were also applied, which disease becomes a key technology needed to be solved now.
contain principle components analysis(PCA) [30], kernel P-
CA(KPCA) [31], fisher linear discriminant analysis (FLDA), III. OUR SCHIZOPHRENIA CLASSIFICATION APPROACH
kernel FLDA(KFLDA) [32], independent component analy- A. ILLUSTRATION OF OUR APPROACH
sis (ICA) [33], locally linear embedding(LLE) [23], canon- Suppose we have a labeled fMRI image data set
ical correlation analysis(CCA) [34], [35]. From the view of S = {(x1 , y1 ), . . . , (xi , yi ), . . . , (xn , yn )}, where xi (xi ∈
machine learning, all the above approaches can be classified Rl1 ×l2 ×l3 ×l4 , 1 ≤ i ≤ n)is the ith 4D fMRI image data,
as shallow machine learning. However, because there exists yi (1 ≤ i ≤ n) is the category label of xi and yi ∈
a gap between the highly complexity of brain function and {1, . . . , C} if there are C categories. Automatic diagnosis of
the relatively poor representation power of shallow machine brain diseases can be treated as a classification task in which
learning, further improving classification performance using unlabeled fMRI image are classified into certain categories
shallow approaches becomes very difficult. by a classifier y = f (x). In order to ensure the classifi-
Recently, deep learning approaches were introduced in cation accuracy, the classifier y = f (x) should be trained
classification of brain disease [48], [49], which contain deep and
Pn optimized based on the minimization of empirical loss
belief network [21], [36]–[38], Auto Encoder networks [19], i=1 L(f (xi ), yi ).
[20], [39], [40], Convolution Neural Networks(CNN) [41]– This paper proposes a multiple feature image capsule
VOLUME 4, 2016 3

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

network ensemble approach for fMRI images classification. Algorithm 1 Optimization Algorithm for solving model(2)
the proposed approach is illustrated in Fig.1. Different from Input: the sample matrix R, the initialized Dictionary D(0), the
the existing approaches, which mainly use functional connec- regularization coefficient λ, stop threshold ε, t = 0
Output: the optimized Dictionary D ∗ , the optimized sparse pre-
tivity feature set from ROI features to make decision, here
sentation matrix S ∗
we use classifier ensemble on multiple features to improve 1: Fix dictionary D(t) and optimize presentation matrix S(t)
the accuracy and stability of brain disease classification. In according to Eq.(3).
our ensemble model, we first extract different 2D feature 2: Fix sparse presentation matrix S(t) and optimize dictionary
images based on the original high-dimensional 4D fMRI D(t + 1) according to Eq.(4).
3: If ||S(t − 1) − S(t)||2F > ε or ||D(t − 1) − D(t)||2F > ε,
image data from three different perspectives: linear sparse
t = t + 1, goto 1.
representation, nonlinear kernel mapping and prior knowl- 4: return the optimized Dictionary D ∗ = D(t + 1), the optimized
edge from brain science. Then the three 2D feature image sparse presentation matrix S ∗ = S(t).
are input into three capsule networks for individual decision-
making. Finally, the classifier ensemble technology is used to
integrate the three individual classification results and output operator and Least Angle Regression, while Eq.(4) can be
the final ensemble decision-making results. solved by K Singular Value Decomposition algorithm [50].
Using dictionary D, a l-dimensional sample vector ri is
B. LINEAR SPARSE REPRESENTATION IMAGE represented as a m-dimensional feature vector si . Noting that
Linear sparse representation images are generated by sparse the regularization coefficient λ has the effect of adjusting the
dictionary learning [50]. Because there is no need for tem- sparsity of representation S, here we set d different λ values
poral and spatial information here, first the 4D fMRI image λ1 , . . . , λd and obtain d different feature vectors of different
xi (xi ∈ Rl1 ×l2 ×l3 ×l4 ) is vectorized to a 1D vector ri (ri ∈ degrees of sparsity. Finally, we combine these d different
Rl×1 , l = l1 × l2 × l3 × l4 ). feature vectors into a sparsity representation matrix Si which
Suppose there exists a Dictionary matrix D(D ∈ Rl×m ) is called as linear sparsity representation image here
to be optimized, which contains m atoms {di }1≤i≤m and
D = (d1 , . . . , di , . . . , dm ). The optimization model of Si = (si (λ1 ), . . . , si (λd )). (5)
sparse dictionary learning is defined as
C. MULTIPLE KERNEL REPRESENTATION IMAGE
min T
T r((R − DS) (R − DS)) + λΩ(S), (1) Multiple kernel representation images are generated by using
D,S kernel mapping method [51]. For a l-dimensional 1D vector
where matrix R is the sample matrix and R ∈ ri in sample set {ri }1≤i≤n , its nonlinear mapping can be
Rl×n , R = (r1 , . . . , ri , . . . , rn ), matrix S(S ∈ Rm×n , S = described as
(s1 , . . . , si , . . . , sn )) is the sparse presentation matrix using
dictionary D, Ω(S) is the sparsity regularization term for S, ri → φ(ri ), (6)
λ(λ ≥ 0) is the regularization coefficient.
where φ(.) is a nonlinear mapping function.
At present, the commonly used sparsity regularization
Considering that φ(.) can hardly be defined in general, we
terms are mainly based on L0 norm and L1 norm. Because
cannot analyse it directly in this implicit nonlinear space. As
L1 norm is of convexity in mathematics and brings more
an alternative, a well-defined kernel inner product function
effective algorithms, here sparse dictionary learning based
can be used. For two samples ri , rj , the kernel inner product
on L1 norm regularization is adopted. So model(1) can be
about them is defined as
rewritten as
k(ri , rj ) = φ(ri )T φ(rj ), (7)
T
min T r((R − DS) (R − DS)) + λ||S||1 . (2)
D,S where k(., .) is the kernel inner product function.
Eq.(2) can be optimized by alternating optimization of the Furthermore, the sample ri can be mapped into n-
below 2 sub problems dimensional kernel sampling space using kernel inner prod-
uct function as a kernel sample ϕ(ri ), which can be described
min T r((R − DS)T (R − DS)) + λ||S||1 , (3) as
S

min T r((R − DS)T (R − DS)). (4) ri → ϕ(ri ) = (k(r1 , ri ), . . . , k(rn , ri ))T . (8)
D

Details of the optimization algorithm is described in Algo- Kernel sample ϕ(ri ) is related not only to the types
rithm 1. of kernel inner product function but also to the values of
Eq.(3) and (4) in Algorithm 1 are both convex optimization super parameters. Here we set s different super parameter
models. Eq.(3) can be solved quickly by some rapid algo- values σ1 , . . . , σs and obtain d different kernel samples for
rithms such as the Least Absolute Shrinkage and Selection ri . Finally, we combine these s different kernel samples
4 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

into a multiple kernel representation matrix Ki denoted as (1) The single convolution layer of the original capsule
nonlinear multiple kernel representation image network is extended to multiple convolution layers to
  reduce the number of network parameters and expand
kσ1 (r1 , ri ) · · · kσs (r1 , ri ) the local receptive field.
Ki =  .. .. ..
. (9)
 
. . . (2) All convolution layers of our capsule network are
kσ1 (rn , ri ) · · · kσs (rn , ri ) designed not containing nonlinear activation, such as
RELU, to avoid over fitting.
D. FUNCTIONAL CONNECTIVITY IMAGE
The above two presentation feature images are generated F. WEIGHTED CLASSIFIER ENSEMBLE
from the perspective of [Link] them, functional Suppose the score outputs of n training samples from
connectivity feature images are based on brain science. As m different classifiers are {oi }1≤i≤m , where oi (oi =
we know, the first three dimensions of a 4D fMRI image (o1i , . . . , oji , . . . , oni )) is the output vector from the ith classi-
xi (xi ∈ Rl1 ×l2 ×l3 ×l4 ) compose the spatial dimensions of fier and oji is the output of the jth training sample from the ith
brain voxels and the last one dimension is the temporal classifier. To find the best score fusion results, here we pro-
dimension. The time series of a brain voxel describes the vari- pose a multiple classifier weighted ensemble method based
ation in blood oxygen concentration with time at this brain on the maximization of the square of correlation coefficients.
voxel and evaluates the variation in activation and repression The optimization model is formulated as below
of this brain voxel. From the perspective of brain science, we m
can obtain the functional connectivity matrix about all ROIs
X αT O T oi oTi Oα
max J(α) =
by calculating the correlation coefficients of time series about i=1
||Oα||22 ||oi ||22 , (12)
every two ROIs. The calculation of functional connectivity s.t. α ≥ 0, αT e = 1
matrix is described as follows:
where O(O ∈ Rn×m ) is the output matrix and O =
(1) Select ROIs and obtain time series v1i , . . . , vm
i
in fMRI (o1 , . . . , om ), α(α ∈ Rm×1 ) is the weight vector, and
image xi accrording to brain science; e(e ∈ Rm×1 ) is a vector whose elements are equal to 1.
i
(2) Calculate Pearson correlation coefficient wpq of time In Eq.(12), we find the best weight vector to maximize the
i i
series vp and vq . The Pearson correlation coefficient square of correlation coefficients between the fusion output
i
wpq is defined as vector Oα and all outputs from the m single classifiers oi .
We use the square of correlation coefficients instead of cor-
relation coefficientsbecause the former is more convenient
i
(vpi − mvpi )T (vqi − mvqi ) to calculate. Considering the mathematical significance of
wpq = , (10)
||vpi − mvpi ||2 ||vqi − mvqi ||2 weights, we introduce the constraints α ≥ 0, αT e = 1 into
our model.
where mvpi , mvqi are the mean values of time series vpi , After defining inner product matrix of outputs K = O T O,
vqi . Finally, we obtain a functional connectivity matrix Wi Eq.(12) can be further reformulated as
denoted as functional connectivity image here
αT KSKα
max J(α) =
αT Kα ,
 i i (13)

w11 · · · w1m
T
Wi =  ...
 .. .. . (11) s.t. α ≥ 0, α e = 1
. . 
i
wm1 ··· i
wmm where S(S ∈ Rm×m ) is a diagonal matrix and Sii = ||o1i ||2 .
2
Observing the main optimization item of Eq.(13) separate-
E. CAPSULE NETWORK DESIGN ly, we can find that it is essentially a generalized eigenvalue
Capsule network is a neural network by Hinton in 2017 decomposition model. Without considering the constraints,
[52]. Compared with CNN, the capsule network discovers the best weight vector is exactly the first eigenvector cor-
the equivariance of features by introducing several capsule responding to the first eigenvalue of the main optimization
layers. It has been found experimentally that capsule network item. Unfortunately, after introducing the constraints, the
can solve some problems with a relatively small sample size conclusion is usually invalid.
effectively [52]. After further analysis, we find that Eq.(13) can be treated
To improve the generalization performance of the capsule as an unconstrained optimization model under certain condi-
network for schizophrenia classification, here we adjusted tions, which is described in Theorem 1.
the architecture of the original capsule network, which is Theorem 1: If outputs set {oi }m≥i≥1 are linearly indepen-
illustrated in Fig.2. dent set and every 2 outputs oi , oj satisfy the condition
In addition to the basic layers of the original capsule net- oTi oj > 0, then Eq.(13) is equivalent to the generalized
T
work, such as Convolution layer, PrimaryCaps layer, Class- eigenvalue decomposition model max α αKSKα T Kα .
Capes layer and L2 output layer, we make the following Proof 3.1: Suppose the best solution of weight vector for
adjustments: Eq.(13) is α∗ , then we can deduce that using β = cα∗ (c ∈
VOLUME 4, 2016 5

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

FIGURE 2: Basic framework of our capsule network

R, c 6= 0) can also obtain the best objective value and satisfy the condition oTi oj > 0, the best solution α∗ ≥ 0
T
α∗T KSKα∗ always holds, which means that the constraint α∗ ≥ 0 can
α∗T Kα∗
= β βKSKβ
T Kβ . It implies that the constraint
α e = 1 is unnecessary in Eq.(13). Let β ∗ be the best
∗T also be removed from Eq.(13). 
solution of the similar model not containing the constraint In application, the condition oTi oj > 0 in Theorem 1
αT e = 1. We can obtain the best vector α∗ for Eq (13) by can be achieved naturally. Taking a frequently-used classifier,
normalizing β ∗ softmax classifier, as an example, the outputs oi of the ith
softmax classifier are the probabilities of samples belonging
β∗
α∗ = ∗T . (14) to all categories. Clearly, oi ≥ 0 always holds, which means
β e that oTi oj > 0 also holds for any 2 outputs oi , oj . SVM
In this case, J(α∗ ) = J(β ∗ ). Hence, the constraint classifier, which is another frequently used classifier, is also
αT e = 1 can be removed from Eq.(13). The main optimiza- used in this paper. The original outputs oi (oi = {oij }, oij ∈
tion item of Eq.(13) is a generalized eigenvalue decomposi- R) of the ith SVM classifier can also be transformed into
tion model. According to Lagrangian Multiplier method, the 1
the probability outputs 1+exp(Ao ij +B)
(Refer to paper [53] for
best solution satisfies more details). Hence, here we find the best weight vector by
KSKα = λKα, (15) solving eigenvalue decomposition (16) directly.

where λ(λ > 0) is the Lagrangian multiplier. Because


{oi }m≥i≥1 are linearly independent, matrix K is always IV. EXPERIMENTS
invertible. Eq.(15) can be reformulated as A. DATASET
The dataset includes 385 subjects that are composed of 153
SKα = λα. (16)
patients with schizophrenia and 232 healthy controls from
Eq.(16) shows that the generalized eigenvalue decomposi- three imaging resources.
tion problem can be treated as a eigenvalue decomposition The first sub-dataset is COBRE, which includes 72 pa-
problem for matrix SK when matrix K is invertible. tients with schizophrenia and 75 healthy controls. The o-
Let α∗ be the first eigenvector corresponding to the first riginal fMRI data were obtained from a 3 Tesla SIEMENS
eigenvalue of matrix P (P = SK). In this case, the best TIM scanner with the following parameters:time of repetition
∗T ∗
objective value is αα∗TPαα∗ . In general, α∗ ≥ 0 does not hold. TR=2000ms, echo time TE=29ms, flip angle FA=75◦ , field
However, when the condition, every 2 outputs oi , oj satisfy of view FOV=192mm, 4mm thickness and 0mm gap, matrix
the condition oTi oj > 0, holds, we find it always holds. size 64×64 and the number of axial slices is 32.
In this case, K ≥ 0 and P ≥ 0 also hold, and we have The second sub-dataset is UCLA, which includes 58
Pm Pm ∗ ∗ patients with schizophrenia and 134 healthy controls. The
α∗T P α∗ i=1 j=1 αi αj Pij
= original fMRI data were obtained from a 3 Tesla SIEMENS
α∗T α∗ P α∗T α∗
Pm Pm ∗ ∗ m Pm ∗ ∗ TIM scanner with the following parameters:time of repetition
|αi ||αj |P ij j=1 |αi ||αj |Pij

i=1 j=1
=
i=1
. TR=2000ms, echo time TE=30ms, flip angle FA=90◦ , field
α∗T α∗ |α∗ |T |α∗ | of view FOV=192mm, 4mm thickness and 0mm gap, matrix
The above inequality shows that |α∗ | is a better solution size 64×64 and the number of axial slices is 34.
than the best vector α∗ . Thus, when every two outputs oi , oj The last sub-dataset is WUSTL, which includes 23 patients
6 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

with schizophrenia and 41 healthy controls. All sub-datasets kernel representation feature image with size 100×50.
are available at the web site ([Link] The F unctional connectivity image. All voxels are
original fMRI data were obtained from a 3 Tesla SIEMENS
first divided into 116 brain regions according to the
TIM scanner with the following parameters:time of repetition
automated anatomical atlas(AAL). Then, we calculated
TR=2500ms, echo time TE=27ms, flip angle FA=90◦ , field the mean blood oxygen concentration of all brain regions,
of view FOV=256mm, 4mm thickness and 0mm gap, matrix which are treated as the time series of all ROIs. We
size 64×64 and the number of axial slices is 33. further calculate all the correlation coefficients between
every 2 ROIs. Finally, we obtain a functional connectivity
feature image with size 116×116.
B. PREPROCESSING AND FEATURE GENERATION
All fMRI data were preprocessed as previously described
[54], [55] by using a statistical parametric mapping software C. APPROACHES AND EXPERIMENTAL SETTINGS
package(SPM8, which can be downloaded freely from the In our experiments, the 10 fold cross validation method is
web site: [Link] For each subject, used to evaluate the proposed approach. Then, SVM, ELM,
the first 5 frames of the scanned data were discarded for DAN, CNN, the original capsule network(capsule network-
magnetic saturation. The following preprocessing steps then 1) and our modified capsule network(capsule network-2) are
proceeded in turn: a) slice timing correction; b) motion cor- employed as the final comparative classifiers for both single
rection; c) normalization with an EPI template in the Montre- and multiple [Link] details of settings are as follows:
al Neurological Institute atlas space (3mm isotropic voxels); SVM. We use the libSVM toolbox which can be down-
d) spatial smoothing with a 6mm fullwidth half-maximum loaded freely from the web site([Link]/
Gaussian kernel; e) linear detrend and band-pass tempo- cjlin/libsvm/). In our experiments, we use linear SVM
ral filtering (frequency range:0.01-0.08Hz); f) regression of and the penalty parameter C of SVM is selected from
nuisance variables, including the 6 parameters obtained by {10−5 , 10−4 , . . . , 104 , 105 } by grid search method.
rigid body head motion correction, ventricular and white
matter signals, and their first temporal derivatives, quadratic ELM. The ELM is programmed on MATLAB platfor-
terms, and squares of derivatives (32P); and g) if frame-wise m. Our ELM is of single hidden layer and the num-
displacement at any point in time exceeded 0.3mm, then that ber of nodes in the hidden layer is selected from
time point was scrubbed. {50, 100, . . . , 450, 500} by grid search method. Here
After the above processing steps, data cleaning operations the activation function of our ELM is sigmoid function
1
such as control of motion artifact, balancing for age and sigmoid(x) = 1+exp(−x) .
gender between the patient and control groups were then DAN. The DAN is also programmed on MATLAB platfor-
performed according to paper [19], [20]. Finally, we ob- m. The number of hidden layers is chosen in {2, 3, 4, 5}.
tained 222 image samples composed of 102 patients with For simplicity, the numbers of nodes in hidden layers are
schizophrenia and 120 healthy controls, which were well all same. The number of nodes in a hidden layer is se-
matched in gender (patients vs. controls: 47/55 vs. 70/51 lected from {50, 100, 150}. Here the activation function
males/females) and age (patients vs. controls: 33.41±9.47 is tanh function tanh(x) = exp(x)−exp(−x)
exp(x)+exp(−x) .
vs. 31.99±10.08 years). All samples were then adjusted to
4D images with the same size 53×63×52×140. The linear CNN. In our experiment, the CNN is designed and com-
sparse representation image, nonlinear multiple kernel rep- pleted on the TensorFlow platform. The designed CNN
resentation image, and functional connectivity image are all contains K composite convolution layers and one fully
from the above preprocessed 4D images. connected classification layer. The output size of the fully
connected classification layer is 2, which is the same
Linear sparse representation image. The dic- as the number of classes in our experiment. For sim-
tionary contains 80 atoms. All the atoms are selected plicity, every composite convolution layer is composed
initially and randomly from another fMRI image dataset, of a convolution layer and a max pooling layer. The
Human Connectome Project and 20 values of the regu- RELU nonlinearly mapping layers are excluded from our
larization coefficient λ : {2−10 , 2−9 , . . . , 28 , 29 } are set. composite convolution layers. In our CNN, the window
Finally, we obtain a linear sparse representation feature size and stride size are 2 for all max pooling layers. In
image with size 80×20. all convolution layers, the kernel size is 3, the stride size
N onlinear multiple kernel representation is 1, and padding is equal to "same". The number of
image. The Gaussian kernel function k(ri , rj ) = output channels is chosen in {8,16,32}, and the number
||r −r ||2 of composite convolution layers k is chosen in {1,2,3,4}.
exp( i−σ2j 2 ) is used and 20 values of the kernel param-
When training our CNN, batch size is 5, flop size is 200,
eter σ : {σs ×1.20 , σs ×1.21 , . . . , σs ×1.248 , σs ×1.249 }
are set. σs is calculated as previously described [56]. In initial learning rate is 0.01 and Adam optimizer is used
every training stage, all 100 samples are used as repre- for learning rate adjustment.
sentation basis. Finally, we obtain a nonlinear multiple Capsule network-1. Here we use a publicly accessible edi-
VOLUME 4, 2016 7

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

tion of capsule network([Link]


CapsNetTensorflow) on the TensorFlow platform. The
capsule network-1 contains 1 RELU convolution layers,
1 primary capsule layer, 1 class capsule layer, and 1 L2
output layer. In the RELU convolution layer, the kernel
size is 3, the stride size is 1, padding is equal to "valid".
The number of output channels is chosen in {8,16,32}.In
the primary capsule layer, the length of capsule is 4, the
kernel size is 2, and the stride size is 2. In the class
capsule layer, the length of capsule is 30 and the number (a) linear sparse representation feature image
of capsules is 2, which is equal to the number of classes
in our experiment. When training capsule network-1, the
batch size, flop size, initial learning rate, and optimizer
are the same as those used in our CNN.
Capsule network-2. The capsule network-2 is designed
and completed based on the above capsule network-1.
Different from only 1 RELU convolution layer in the cap-
sule network-1, the capsule network-2 contains k linear
convolution layers. The number of convolution layers k
is chosen in {3,4,5,6}, the number of output channels is (b) nonlinear multiple kernel representation
also chosen in {8,16,32}, and the values of the other super feature image
parameters in the capsule network-2 are the same as those
in the capsule network-1.

D. EXPERIMENTAL RESULTS AND ANALYSIS


1) CLASSIFICATION ACCURACY AND ROC CURVES
Tables 1-4 show the average specificity(SC), sensitivi-
ty(SS), classification accuracy(ACC), positive predictive val-
ue(PPV), negative predictive value(NPV) and F1-Score of all
classifiers when using linear sparse representation feature im-
age, nonlinear multiple kernel representation feature image, (c) functional connectivity feature image
functional connectivity feature image and our score weighted
fusion on the multi-site data set.
Tables 1-3 show that the capsule network-2 has better clas-
sification performance than the others when using each single
features and achieves the best average classification accuracy
of 81.82% when using functional connectivity feature image,
with 2.53% higher than the second-best average classification
accuracy of 79.29% from DAN. By comparison, the capsule
network-1 shows no obvious improvement of classification
performance. It proves that the proposed architecture ad- (d) our weighted ensemble
justment for capsule network works indeed. Tables 4 shows
that our weighted ensemble technology further improves the FIGURE 3: ROC curves
classification performance and that the final classification
accuracy when using the capsule network-2 is up to 82.83%,
with 2.02% higher than the second-best average classification resentation feature image, finally nonlinear multiple kernel
accuracy of 80.81% from DAN. For the other classifiers representation feature image. Using our weighted ensemble
except SVM and ELM, our weighted ensemble technology technology, the best AUC value is up to 0.9141.
also improves the classification performance to some extent.
Furthermore, we calculate the AUC values and draw the 2) GRID SEARCH FOR PARAMETERS OF CNN AND
ROC curves of all classifiers and illustrate them in Fig.3. CAPSULE NETWORK
Fig.3 shows that the capsule network-2 obtains the best The above experimental results are obtained by using the
AUC values, followed by the DAN, finally the other classi- best parameters, which are determined through grid search
fiers. From the view of AUC values, functional connectivity method on validation samples. Figs.4, 5, and 6 illustrate the
feature image is the best, followed by linear sparse rep- average validation correct rates of SVM, ELM and capsule
8 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

TABLE 1: classification results using linear sparse representation feature image(%)


Classif iers SC SS ACC PPV NP V F 1 − Score
SV M 85.12±14.40 55.84±14.92 73.74±11.94 76.12±24.38 75.07±7.58 62.55±15.28
ELM 75.21±8.22 74.03±14.02 74.75±7.19 66.04±8.99 82.62±8.70 69.08±9.06
DAN 74.38±20.65 71.43±14.74 73.23±12.06 70.44±17.59 73.97±18.46 68.48±13.36
CN N 80.17±13.37 58.44±20.66 71.72±8.03 67.88±15.87 76.07±7.56 59.9±15.95
capsule network − 1 74.38±19.27 72.73±23.12 73.74±15.48 66.14±20.61 81.12±13.93 68.25±19.79
capsule network − 2 80.99±12.50 68.83±15.41 76.26±8.99 71.84±14.39 80.85±8.04 69.23±11.52

TABLE 2: classification results using nonlinear multiple kernel representation feature image(%)
Classif iers SC SS ACC PPV NP V F 1 − Score
SV M 90.91±8.13 42.86±18.07 72.22±9.94 75.00±20.22 71.77±7.60 53.33±18.64
ELM 85.95±8.49 48.05±21.45 71.21±11.34 67.14±19.76 72.84±9.22 54.99±20.53
DAN 76.86±17.08 68.83±14.59 73.74±10.35 71.00±15.94 78.55±15.46 67.64±12.66
CN N 87.60±10.18 44.16±25.12 70.71±11.14 70.38±22.34 72.19±10.07 51.15±22.84
capsule network − 1 85.95±15.41 50.65±22.46 72.22±13.38 75.49±26.42 73.75±10.79 57.54±19.94
capsule network − 2 80.17±17.52 72.73±14.29 77.27±10.10 75.89±16.31 80.40±16.59 71.93±12.64

TABLE 3: classification results using functional connectivity feature image(%)


Classif iers SC SS ACC PPV NP V F 1 − Score
SV M 85.95±13.09 58.44±22.55 75.25±11.75 73.15±18.17 77.26±10.21 63.25±19.92
ELM 88.43±23.10 61.04±14.67 77.77±17.74 79.19±21.80 78.22±16.32 68.08±16.77
DAN 85.12±9.34 70.13±21.62 79.29±8.10 76.99±11.59 83.17±10.39 71.00±13.84
CN N 78.51±13.65 71.43±20.20 75.76±12.49 69.03±17.64 82.07±11.84 69.25±16.61
capsule network − 1 82.64±13.15 71.43±23.04 78.28±9.45 76.40±16.24 83.47±10.89 70.43±14.88
capsule network − 2 88.43±10.03 71.43±19.17 81.82±7.89 82.77±14.86 83.92±8.76 74.28±13.69

TABLE 4: classification results using our weighted ensemble(%)


Classif iers SC SS ACC PPV NP V F 1 − Score
SV M 85.12±13.72 57.14±18.07 74.24±11.19 73.58±20.80 76.07±8.34 62.81±16.79
ELM 78.51±10.83 76.62±12.92 77.77±8.96 73.93±18.93 82.21±6.01 74.27±13.31
DAN 87.60±10.18 70.13±17.44 80.81±7.60 81.19±14.57 82.91±7.50 73.11±13.04
CN N 80.99±12.50 68.83±15.41 76.26±8.99 71.84±14.39 80.85±8.04 69.08±11.52
capsule network − 1 80.99±14.56 75.32±13.48 78.79±9.33 74.74±10.08 84.74±7.95 73.17±9.49
capsule network − 2 85.95±9.42 77.92±17.34 82.83±7.64 79.05±10.55 86.90±8.84 77.30±11.36

FIGURE 4: Average validation correct rates using SVM FIGURE 5: Average validation correct rates using ELM
under different parameter settings under different parameter settings

network-1 when using different single parameter settings. number of hidden layer nodes for ELM is 400, the best
Figs.7, 8, and 9 illustrate the average validation correct rates number of channels for capsule network-1 is 16, the best
of DAN, CNN and capsule network-2 when using different number of layers and the best number of nodes for DAN
combinations of 2 parameters. is 3 and 50 respectively, the best number of layers and the
Using linear sparse representation feature image, we find best number of channels for capsule network-2 are 5 and 16
that the best penalty parameter C for SVM is 10−3 , the best respectively, the best number of layers and the best number of
VOLUME 4, 2016 9

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

(a) DAN
FIGURE 6: Average validation correct rates using capsule
network-1 under different parameter settings

(b) CNN

(a) DAN

(c) Capsule Network-2

FIGURE 8: Average validation correct rates for nonlinear


(b) CNN multiple kernel representation feature image using DAN, C-
NN and capsule network-2 under different parameter settings

channels for CNN are 2 and 16 respectively. Using nonlinear


multiple kernel representation feature image, we find that the
best penalty parameter C for SVM is 10−1 , the best number
of hidden layer nodes for ELM is 100, the best number
of channels for capsule network-1 is 8, the best number of
layers and the best number of nodes for DAN is 2 and 50
respectively, the best number of layers and the best number
of channels for our capsule network are 6 and 16 respectively,
(c) Capsule Network-2 the best number of layers and the best number of channels for
CNN are 2 and 8 respectively. Using functional connectivity
FIGURE 7: Average validation correct rates for linear sparse feature image, we find that the best penalty parameter C for
representation feature image using DAN, CNN and capsule SVM is 100 , the best number of hidden layer nodes for ELM
network-2 under different parameter settings is 300, the best number of channels for capsule network-1 is
10 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

depth of network can avoid over-fitting effectively while im-


proving the ability of covariant features extraction. Besides,
the Figs.5,6,7,8 and 9 illustrate that the average validation
correct rates of all neural networks including ELM, DAN,
CNN, capsule network-1 and capsule network-2 are not very
sensitive to the width of network.

V. CONCLUSION
To improve the effectiveness of schizophrenia classifica-
tion, we propose a multiple feature image capsule networks
ensemble method. In the proposed method, we introduce
(a) DAN multiple features extraction, deep capsule network design,
and a novel weighted classifier ensemble to increase the com-
plementarity of features and improve the generality of clas-
sification. Finally, we conduct the comparative experiments
on the schizophrenia fMRI dataset from COBRE, UCLA
and WUSTL. Experimental results show that the proposed
method performs better than the other comparative methods
and and that the average correct rate of schizophrenia classi-
fication increases to 82.83%.

ACKNOWLEDGEMENT
BO YANG, YUAN CHEN, and QUAN-MING SHAO con-
(b) CNN
tributed equally to this work. Correspondence to JUN-
QIANG JIANG and LI PAN. The authors would like to thank
the anonymous reviewers for their valuable comments on
improving the paper.

REFERENCES
[1] J. U. Duncombe, “Diagnosis and classification of schizophrenia,”
Schizophr Bull, vol. 19, no. 2, pp. 199–214, Feb. 1993.
[2] N. C. Andreasen, “Diagnosis of schizophrenia,” Schizophr Bull, vol. 13,
no. 1, pp. 9–22, Jan. 1987.
[3] N. K. Logothetis, J. Pauls, M. Augath, T. Trinath and A. Oeltermann,
“Neurophysiological investigation of the basis of the fmri signal,” Nature,
vol. 412, no. 6843, pp. 150–157, Jul. 2001.
[4] Y. Zhang, Z. Dong, L. Wu and S. Wang, “A hybrid method for mri brain
image classification,” Expert Systems with Applications, vol. 38, no. 8, pp.
(c) Capsule Network-2 10049–10053, Aug. 2011.
[5] Y. Assaf and O. Pasternak, “Diffusion tensor imaging (DTI)-based white
FIGURE 9: Average validation correct rates for functional matter mapping in brain research: a review,” Journal of Molecular Neuro-
science, vol. 34, no. 1, pp. 51–61, Jan. 2008.
connectivity feature image using DAN, CNN and capsule
[6] K. Vunckx et al, “Evaluation of three mri-based anatomical priors for
network-2 under different parameter settings quantitative pet brain imaging,” IEEE Transactions on Medical Imaging,
vol. 31, no. 3, pp. 599–612, Mar. 2012.
[7] P. Mcguire, O. D. Howes, J. Stone and P. Fusar-Poli, “Functional neu-
roimaging in schizophrenia: diagnosis and drug discovery,” Trends in
16, the best number of layers and the best number of nodes Pharmacological Sciences, vol. 29, no. 2, pp. 91–98, Feb. 2008.
[8] S. Ehrlich et al, “Associations of cortical thickness and cognition in
for DAN is 2 and 100 respectively, the best number of layers patients with schizophrenia and healthy controls,” Schizophrenia Bulletin,
and the best number of channels for our capsule network are vol. 38, no. 5, pp. 1050–1062, Sep. 2012.
6 and 16 respectively, the best number of layers and the best [9] J. Sui et al, “Combination of resting state fmri, dti, and smri data to
discriminate schizophrenia by n-way MCCA+JICA,” Frontiers in Human
number of channels for CNN are 2 and 16 respectively. Neuroscience, vol. 7, pp. 235, May. 2013.
Figs.7, 8, and 9 illustrate that the average validation correct [10] D. F. Wong et al, “Quantification of cerebral cannabinoid receptors sub-
rates of DAN and CNN are reduced obviously as the number type 1 (CB1) in healthy subjects and schizophrenia by the novel PET
of layers is increased, which implies that over-fitting be- radioligand [11C]OMAR,” NeuroImage, vol. 52, no. 4, pp. 1505–1513,
Oct. 2010.
comes more and more obvious with the increase of the num- [11] P. Orban et al, “Multisite generalizability of schizophrenia diagnosis clas-
ber of nonlinear layers. By contrast, the variation of average sification based on functional brain connectivity,” Schizophrenia Research,
validation correct rates of capsule network-2 with the number vol. 192, pp.167-171, Aug. 2017.
[12] R. Irina et al, “Schizophrenia as a Network Disease: Disruption of Emer-
of layers are not the case, which shows that our introducing gent Brain Function in Patients with Auditory Hallucinations,” PLoS ONE,
multiple linear layers into capsule network to increase the vol. 8, pp. e50625, Jan. 2013.

VOLUME 4, 2016 11

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

[13] G. S. Malhi et al, “Bipolaroids: functional imaging in bipolar disorder,” [35] P. Guo, G. Q. Xie and R. F. Li, “Object Detection Using Multiview
Acta Psychiatr Scand Suppl, vol. 110, no. s422, pp. 46–54, Feb. 2010. CCA-Based Graph Spectral Learning,” Journal of Circuits, Systems and
[14] V. Elisa et al, “Machine Learning Approaches: From Theory to Appli- Computers, vol. 28, no. 4, Feb. 2019.
cation in Schizophrenia,” Computational and Mathematical Methods in [36] H. Suk, S. W. Lee and D. Shen, “Hierarchical feature representation and
Medicine, vol. 2013, pp. 1–12, Dec. 2013. multimodal fusion with deep learning for AD/MCI diagnosis,” Neuroim-
[15] Y. Lecun, Y. Bengio and G. Hinton, “Deep learning,” Nature, vol. 521, no. age, vol. 101, pp. 569–582, Nov, 2014.
7553, pp. 436–444, May. 2015. [37] Y. Yoo et al, “Deep learning of joint myelin and T1w MRI features in
[16] A. Voulodimos, N. Doulamis, A. Doulamis and E. Protopapadakis, “Deep normal-appearing brain tissue to distinguish between multiple sclerosis
learning for computer vision: A brief review,” Computational Intelligence patients and healthy controls,” Neuroimage Clinical, vol. 17, pp. 169–178,
and Neuroscience, vol. 2018, pp. 1–13, Feb. 2018. Oct, 2018.
[38] M. Akhavan Aghdam, A. Sharifi and M.M. Pedram, “Combination of
[17] K. Noda et al, “Audio-visual speech recognition using deep learning,”
rsfMRI and sMRI data to discriminate autism spectrum disorders in young
Applied Intelligence, vol. 42, no. 4, pp. 722–737, Jun. 2015.
children using deep belief network,” Journal of Digital Imaging, vol. 31,
[18] S. P. Singh et al, “Machine translation using deep learning: An overview,”
no. 6, pp. 895–903, Dec, 2018.
in International Conference on Computer, Communications and Electron-
[39] H. Suk, C.Y. Wee, S.W. Lee and D. Shen, “State-space model with
ics, New York, USA, 2017, pp. 162–167.
deep learning for functional dynamics estimation in resting-state fMRI,”
[19] L. L. Zeng et al, “Multi-site diagnostic classification of schizophrenia Neuroimageg, vol. 129, pp. 292–307, Apr, 2016.
using discriminant deep learning with functional connectivity MRI,” E- [40] A.S. Heinsfeld, A.R. Franco, R.C. Craddock, A. Buchweitz and F.
BioMedicne, vol. 30, pp. 74–85, Apr. 2018. Meneguzzi, “Identification of autism spectrum disorder using deep learn-
[20] J. Kim, V. Calhoun, E. Shim and J. Lee, “Deep neural network with ing and the ABIDE dataset,” Neuroimageg Clinical, vol. 17, pp. 16–23,
weight sparsity control and pretraining extracts hierarchical features and Apr, 2018.
enhances classification performance: evidence from wholebrain resting- [41] A. Karim, “Classification of sMRI for AD Diagnosis with Convolutional
state functional connectivity patterns of schizophrenia,” NeuroImage, vol. Neuronal Networks: A Pilot 2-D+ Study on ADNI,” MultiMedia Model-
124, pp. 127–146, Jan. 2016. ing 23rd International Conference, MMM 2017, , 2017, pp. 690–701.
[21] W.H. Pinaya et al, “Using deep belief network modelling to characterize [42] H. Choi and K.H. Jin, “Predicting cognitive decline with deep learning of
differences in brain morphometry in schizophrenia,” Scientific Reports, brain metabolism and amyloid imaging,” Behavioural Brain Research, vol.
vol. 6, pp. 38897, Dec, 2016. 344, pp. 103–109, May, 2018.
[22] J. I. Arribas, V. D. Calhoun and T. Adali, “Automatic bayesian classifica- [43] H. Suk, S.W. Lee and D. Shen, “Alzheimer’s Disease Neuroimaging I.
tion of healthy controls,bipolar disorder and schizophrenia using intrinsic Deep ensemble learning of sparse regression models for brain disease
connectivity maps from fmri data,” IEEE transactions on biomedical diagnosis,” Medical Image Analysis, vol. 37, pp. 101–113, Apr, 2017.
engineering, vol. 57, no. 12, pp. 2850–2860, Dec. 2010. [44] M. Liu, D. Cheng, K. Wang and Y. Wang, “Alzheimer’s Disease Neu-
[23] H. Shen, Y. Liu and D. Hu, “Discriminative analysis of resting-state roimaging I. Multi-modality cascaded convolutional neural networks for
functional connectivity patterns of schizophrenia using low dimensional Alzheimer’s disease diagnosis,” Neuroinformatics, vol. 16, no. (3–4), pp.
embedding of fmri,” Neuroimage, vol. 49, no. 4, pp. 3110–3121, Feb. 295–308, Oct, 2018.
2010. [45] S.H. Wang, P. Phillips, Y. Sui, B. Liu, M. Yang and H. Cheng, “Classifi-
[24] M. R. Arbabshirani, G. D. Pearlson and V. D. Calhoun, “Classification of cation of Alzheimer’s disease based on eight-layer convolutional neural
schizophrenia patients based on resting-state functional network connec- network with leaky rectified linear unit and max pooling,” Journal of
tivity,” Frontiers in Neuroscience, vol. 7, pp. 133, Jul. 2013. Medical Systems, vol. 42, no. 5, pp. 295–308, May, 2018.
[25] M. J. Jafri and V. D. Calhoun, “Functional classification of schizophrenia [46] B. Jie, M. Liu, J. Liu, D. Zhang and D. Shen, “Temporally constrained
using feed forward neural networks,” in International Conference of the group sparse learning for longitudinal data analysis in Alzheimer’s dis-
IEEE Engineering in Medicine and Biology Society, New York, USA, ease,” IEEE transactions on biomedical engineering, vol. 64, no. 1, pp.
2006, pp. 6631–6642. 238–249, Jan, 2017.
[47] Z. Akkus et al, “Predicting deletion of chromosomal arms 1p/19q in low-
[26] B. Yang et al, “A Study on Regularized Weighted Least Square Support
grade gliomas from MR images using machine intelligence,” Journal of
Vector Classifier,” Pattern Recognition Letters, vol. 108, pp. 48–55, Jun.
Digital Imaging, vol. 30, no. 4, pp. 469–476, Aug, 2017.
2018.
[48] K. Sakai and K. Yamada, “Machine learning studies on major brain
[27] S. Laconte, S. Strother, V. Cherkassky, J. Anderson and X. Hu, “ Support
diseases: 5-year trends of 2014–2018,” Japanese Journal of Radiology, vol.
vector machines for temporal classification of block design fmri data,”
37, pp. 34–72, Nov. 2018.
Neuroimage, vol. 26, no. 2, pp. 317–329, Jun. 2005.
[49] K. Justin et al, “Deep learning applications in medical image analysis,”
[28] D. Chyzhyk, A. Savio and M. Grana, “ Computer aided diagnosis of IEEE Access, vol. 6, pp. 9375–9389, Dec. 2017.
schizophrenia on resting state fMRI data by ensembles of ELM,” Neural [50] K. K. Delgado et al, “Dictionary learning algorithms for sparse represen-
Network, vol. 68, pp. 23–33, Aug. 2015. tation,” Neural Computation, vol. 15, no. 4, pp. 349–396, Feb. 2014.
[29] M.N.I. Qureshi et al, “ Multimodal discrimination of schizophrenia using [51] J. S. Taylor and N. Cristianini, Kernel Methods for Pattern Analysis. New
hybrid weighted feature concatenation of brain functional connectivity York, USA: Cambridge University Press, 2004, pp. 123–135.
and anatomical features with an extreme learning machine,” Frontiers in [52] S. Sabour, N. Frosst and G. E. Hinton, “Dynamic routing between cap-
Neuroinformatics, pp. 11–59, Sep. 2017. sules,” in Advances in Neural Information Processing Systems, pp. 1–11,
[30] J. Ford et al, “A combined structural-functional classification of 2017.
schizophrenia using hippocampal volume plus fmri activation,” in the 2nd [53] H. Lin, C. Lin and R. Weng, “A note on platt’s probabilistic outputs for
Joint Conference and the Fall Meeting of the Biomedical Engineering support vector machine,” Machine Learning, vol. 68, no. 3, pp. 267–276,
Society, New York, USA, 2002, pp. 48–49. Oct. 2007.
[31] P. Khurd, S. Baloch, R. Gur and C. Davatzikos, “Manifold learning tech- [54] L. L. Zeng, H. Shen, L. Liu and D. Hu, “Unsupervised classification
niques in image analysis of high-dimensional diffusion tensor magnetic of major depression using functional connectivity MRI,” Human Brain
resonance images,” in IEEE Conference on Computer Vision and Pattern Mapping, vol. 35, no. 4, pp. 1630–1641, Apr. 2014.
Recognition, New York, USA, 2007, pp. 1–7. [55] L. L. Zeng et al, “Neurobiological basis of headmotion in brain imaging,”
[32] F. M. Vos et al, “Linear and kernel fisher discriminant analysis for studying Proceedings of the National Academy of Sciences, vol. 111, no. 16, pp.
diffusion tensor images in schizophrenia,” in the 4th IEEE International 6058–6062, Apr. 2014.
Symposium on Biomedical Imaging: From Nano to Macro, New York, [56] I. W. Tsang, A. Kocsor and J. T. Kwok, “Efficient kernel feature extraction
USA, 2007, pp. 764–767. for massive data sets,” in Proceedings of the 12th ACM SIGKDD inter-
[33] S. A. Meda et al, “Evidence for anomalous network connectivity during national conference on Knowledge discovery and data mining, New York,
working memory encoding in schizophrenia: an ICA based analysis,” Plos USA, 2006, pp. 724–729.
One, vol. 4, no. 11, pp. e7911, Nov. 2009.
[34] J. Sui et al, “A CCA+ICA based model for multi-task brain imaging data
fusion and its application to schizophrenia,” Neuroimage, vol. 51, no. 1,
pp. 123–134, May. 2010.

12 VOLUME 4, 2016

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]
This article has been accepted for publication in a future issue of this journal, but has not been fully edited. Content may change prior to final publication. Citation information: DOI
10.1109/ACCESS.2019.2933550, IEEE Access

Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS

BO YANG received the [Link]. degree in me- GUAN-QI GUO received the [Link]. degree in elec-
chanical engineering from Zhengzhou University, tronics from the Huazhong Institute of Science and
China, in 1996, the [Link]. degree in computer Technology, Wuhan, China, in 1983, and the Ph.D.
application technology from Xiangtan University, degree in control theory and control engineering
China, in 2004, and the Ph.D. degrees in mechani- from Central South University, Changsha, China
cal and electronic engineering from Central South in 2003. He currently works as a Professor in
University, China, in 2010. Since 2012, he has college of information science and engineering,
been an associate professor with the College of Hunan Institute of Science and Technology. His
Information Science and Technology, Hunan Insti- main research interests include evolutionary com-
tute of Science and Technology. His main research putation, multi-objective optimization, and ma-
interests include MR brain image analysis, statistical pattern recognition, and chine learning.
machine learning.

YUAN CHEN received the [Link]. degree in infor-


mation engineering from Hunan Institute of Sci-
ence and Technology, China, in 2017. He is cur-
rently pursuing the [Link]. degree with the College
of Information Science and Technology, Hunan
Institute of Science and Technology. His research
interests include deep learning, and MR brain im-
age analysis. JUN-QIANG JIANG received the Ph.D. degree
in software engineering from Hunan University,
Changsha, China, in 2017. He is currently a master
supervisor in the School of Information Science
and Engineering, Hunan Institute of Science and
Technology, China. His main research interests
include cloud computing, parallel computing and
QUAN-MING SHAO received the [Link]. degree workflow scheduling, and machine learning. He is
in information engineering from Hunan Institute a member of China Computer Federation.
of Science and Technology, China, in 2017. His
research interests include deep learning, MR brain
image analysis, and computer vision.

RUI YU received the [Link]. degree in information


engineering from Hunan Institute of Science and
Technology, China, in 2018. He is currently pur- LI PAN received the M.S. degree in computer
suing the [Link]. degree with the College of Infor- software and theory from Sun Yat-Sen University,
mation Science and Technology, Hunan Institute Guangzhou, China, in 2004, and the Ph.D. degree
of Science and Technology. His research interests in computer applied technology from Tongji Uni-
include deep learning, and multi-objective opti- versity, Shanghai, China, in 2009. He is currently a
mization. Professor with the Department of Information Sci-
ence and Engineering, Hunan Institute of Science
and Technology, China. He has published more
than 20 papers in refereed journals and conference
proceedings. His main research interests include
graph networks, Petri nets, and computational intelligence.
WEN-BIN LI received the B.S. degree in com-
puter science and technology from Hunan Normal
University, Changsha, China, in 2003, and the
M.S. degree in Computer Applications Technol-
ogy from Changsha University of Science and
Technology, Changsha, Chinačňin 2006. He is
currently pursuing the Ph.D. degree with Central
South University, Changsha, China. His research
interests include deep learning, evolutionary com-
putation, and multi-objective optimization.

VOLUME 4, 2016 13

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see [Link]

You might also like