0% found this document useful (0 votes)
11 views10 pages

Multi-Resolution CNN for Lung Nodule Detection

The paper presents a deep learning-based method for lung nodule detection in chest X-ray radiographs using multi-resolution patch-based convolutional networks. The proposed approach demonstrates superior performance, achieving over 99% detection of lung nodules with a low false positive rate when evaluated on the JSRT database. This method has significant potential for clinical application, addressing the shortage of radiologists and improving early lung cancer detection.

Uploaded by

Omar Abdullah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views10 pages

Multi-Resolution CNN for Lung Nodule Detection

The paper presents a deep learning-based method for lung nodule detection in chest X-ray radiographs using multi-resolution patch-based convolutional networks. The proposed approach demonstrates superior performance, achieving over 99% detection of lung nodules with a low false positive rate when evaluated on the JSRT database. This method has significant potential for clinical application, addressing the shortage of radiologists and improving early lung cancer detection.

Uploaded by

Omar Abdullah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Artificial Intelligence In Medicine 103 (2020) 101744

Contents lists available at ScienceDirect

Artificial Intelligence In Medicine


journal homepage: [Link]/locate/artmed

Multi-resolution convolutional networks for chest X-ray radiograph based T


lung nodule detection
Xuechen Lia,b,c, Linlin Shena,b,c,*, Xinpeng Xiea, Shiyun Huangd, Zhien Xiee, Xian Honge, Juan Yuf
a
College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, Guangdong province, PR China
b
Shenzhen Institute of Artificial Intelligence and Robotics for Society, PR China
c
Guangdong Key Laboratory of Itelligent Information Processing, Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen University, PR China
d
Sun Yat-Sen University Public Health Insititue, Guangzhou, Guangdong province, PR China
e
GuangzhHou Thoracic Hospital, Guangzhou, Guangdong province, PR China
f
Imaging Department of Shenzhen University Health Science Center, Shenzhen University School of Medicine, Shenzhen Second People's Hospital, First Affiliated Hospital of
Shenzhen University, Shenzhen, Guangdong, PR China

ARTICLE INFO ABSTRACT

Keywords: Lung cancer is the leading cause of cancer death worldwide. Early detection of lung cancer is helpful to provide
Computer-aided detection the best possible clinical treatment for patients. Due to the limited number of radiologist and the huge number of
x-ray radiograph chest x-ray radiographs (CXR) available for observation, a computer-aided detection scheme should be devel-
Lung nodule detection oped to assist radiologists in decision-making. While deep learning showed state-of-the-art performance in
Multi-resolution patch-based convolutional
several computer vision applications, it has not been used for lung nodule detection on CXR. In this paper, a deep
neural network
learning-based lung nodule detection method was proposed. We employed patch-based multi-resolution con-
volutional networks to extract the features and employed four different fusion methods for classification. The
proposed method shows much better performance and is much more robust than those previously reported
researches. For publicly available Japanese Society of Radiological Technology (JSRT) database, more than 99%
of lung nodules can be detected when the false positives per image (FPs/image) was 0.2. The FAUC and R-CPM
of the proposed method were 0.982 and 0.987, respectively. The proposed approach has the potential of ap-
plications in clinical practice.

1. Introduction of radiologists is 4.1% per year, that of medical image is about 30% per
year. As a result, in the near future, hospitals will suffer from a lack of
Lung cancer is the leading cause of cancer death worldwide. Early experienced radiologists. On the other hand, in countries with in-
detection of lung cancer is helpful to provide the best possible clinical sufficient medical and health resource, the workload of radiologist is
treatment for patients. As usually there are no symptoms associated heavy. As a large number of CXR images need examination every day,
with pulmonary nodules, most pulmonary nodules are discovered when the diagnosis process has to be complete in a short time and a large
a chest x-ray (CXR) is done for some other reasons [1]. Although low- number of inconspicuous nodules could be missed [6,7]. In this case, an
dose computed tomography (CT) showed higher sensitivity to lung automatic nodule detection system with stable performance shall be
cancer at early stage, and has been used for lung cancer screening [2–4] used to help radiologists’ diagnosis [8,9]. A medical research [10]
for about 20 years, a recent research showed that, computer-aided di- suggested that, the diagnosis accuracy of doctors can be massively
agnosis (CAD) software may broaden the usage of CXR and finally be- boosted when a CADe system with low false positive (FP) rate and high
come the low cost solution in countries where low-dose CT is not sensitivity (above 80%) was applied.
popular [5]. As lung cancer appears as nodules in CXR images, nodule Early researches about lung nodule detection used appearance and
detection plays an important role in the early diagnosis of lung cancer. morphological characteristics to distinguish nodules from possible
In real clinical practice, CXR images are evaluated by radiologists. candidates [11,12]. The performances were not satisfactory since only
However, the number of radiologists is far less than the one demanded the intensity and the shape features were considered. Later researches
by the huge number of CXR images. In China, while the increasing rate [13–18] added gradient and texture features for performance

Corresponding author.

E-mail addresses: llshen@[Link], xxp_angle@[Link] (L. Shen), 551759864@[Link] (S. Huang), wuzhuanghong@[Link] (Z. Xie),
yujuan72@[Link] (J. Yu).

[Link]
Received 17 October 2018; Received in revised form 23 October 2019; Accepted 23 October 2019
0933-3657/ © 2019 Elsevier B.V. All rights reserved.
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

improvement. For example, in the work of Wei et al. [13], about two
hundred texture features were extracted for lung nodule description.
The reported sensitivity was 80% when the FPs/image was 5.4. How-
ever, the dimension of features was larger than the number of training
lung nodules, which may lead to over-fitting problem. In another study
[14], about a hundred features including texture and location were
employed to represent lung nodule. The reported sensitivities were 51%
(two FPs/image) and 67% (four FPs/image). In later studies, more
features such as shape, surface, gray-level intensity, and gradient fea-
tures were employed and the reported sensitivities increased to 78%
(four FPs/image) [15] and 79% (five FPs/image) [16]. According to
[19–22], the lung nodules appears more distinct in the soft-tissue
images generated by dual-energy radiographs, where the bones were
suppressed. Chen et al. [17] proposed a CADe method based on rib
suppressed radiographs. The reported sensitivity was 85% when the Fig. 1. The nodule size distributions for JSRT, GDH and SZH databases.
FPs/image was five. Our previous work [18] proposed a solitary fea-
ture-based approach for rib suppressed CXR images. 93% of lung no-
dules can be detected with five FPs/image. Although a large number of SZH dataset, the number of nodules with length 0–10 mm, 10–20 mm,
works have been proposed, the lung nodule detection systems in lit- 20–30 mm and larger than 30 mm are 358, 55, 25 and 19, respectively.
erature still produced a relatively large number of FPs to achieve good The average nodule size is 7.1 mm and 7.3 mm for GDH and SZH da-
sensitivity. tabase, respectively. The size distributions of nodules in different da-
In recent years, thanks to the large amount of available data and tasets were given in Fig. 1.
computational power of modern computers, convolutional neural net- While the X-ray images in JSRT and SZH database have a size of
works (CNNs) have shown state-of-the-art performance in a number of 2048 × 2048 pixels, the x-ray images in GDH database have a size of
computer vision applications [23–28]. Because CNNs can be trained 1024 × 1024 pixels. All images have a gray-scale depth of 12 bits. The
end-to-end in a supervised framework to learn highly discriminative JSRT dataset is publicly available and provided clinical information
features, they are well suited to lung nodule detection in CXR. Algo- about the position and radius of lung nodules. The GDH and SZH da-
rithms such as network fusion [26], Hough-CNN [29] and auto-context tasets provided clinical reports where the locations of lung nodules
[28] were developed to improve the performance of CNN. There have were recorded. The nodule areas were annotated by three radiologists
been a lot of applications of CNN on lung nodule detection using CT from Guangzhou Thoracic Hospital and Shenzhen Second People's
images [27,30,31], either 2D or 3D CNN was employed in these ap- Hospital. Each CXR was annotated by one of the three radiologists.
plications and promising results have been reported. For CXR images, While the ground truth of JSRT database was confirmed using CT, the
CNN has also been applied for disease diagnosis [32–35], image view annotation of lung nodules in GDH and SZH databases were based on
classification [36] and body parts segmentation [37–39]. However, to CXR only.
the best of our knowledge, this is the first CNNs-based lung nodule Four databases were generated to evaluate the proposed nodule
detection study on CXR images by far. In this work, we developed a detection network. Previous studies [15–17] divided JSRT database
multi-resolution patch-based CNN approach for lung nodule detection into two databases. The first database consists of 140 nodule cases of
on CXR. We also evaluated four different fusion methods for the final JSRT where the lung nodules are located in lung field (expressed as
decision. JSRT-A). The second database includes all cases in the JSRT database
The paper is organized as follows. Section 2 gives the introduction (expressed as JSRT-B). The 14 images that were not selected in JSRT-A
of databases employed for training and testing the networks. Section 3 were included in JSRT-B as normal cases. We employed the same di-
proposes multi-resolution patch-based CNN and fusion methods. Sec- vision in this study for performance comparison. Table 1 gives the de-
tion 4 presents the experimental results. Section 5 gives the discussion. tails of the databases employed in this study.
Finally, Section 6 draws conclusion and describes the future work.
2. Method
1.1. Material
The proposed lung nodule detection method includes two main
The proposed method was evaluated using databases from three parts, i.e. training and testing, as shown in Fig. 2. For training, the
different sources. The first database consists of frontal CXR images images was preprocessed using lung field segmentation and rib sup-
provided by Japanese Society of Radiological Technology (JSRT: pression [40,41] to get the region of interest and enhance the visibility
[Link] [40], which were collected from 13 of lung nodules. In the second step, the images were enhanced using
medical centers in Japan and one institution in the United States. The histogram operation. In the third step, multi-resolution patch-based
JSRT database contains 247 digitized CXR images, among which, 154 CNNs were trained for lung nodule detection. Patches centred at all
contain lung nodules and the remaining 93 are normal. Only one no- pixels in lung field were extracted to train CNNs for feature extraction.
dule was available for nodule cases, which was confirmed by CT ex- Fusion strategy was employed for patch classification. Four different
amination. The number of nodules with length 0–10 mm, 11–15 mm, fusion methods were evaluated, and the method with the best perfor-
16–20 mm, 21–25 mm, 26–30 mm and 31–60 mm are 31, 52, 36, 14, 17 mance was selected.
and 4, respectively. The average size of all nodules included in the
database is 17.3 mm. Other two databases were obtained from hospitals Table 1
in Guangzhou (GDH) and Shenzhen (SZH), respectively. The images of The CXR databases.
GDH and SZH database are digitally acquired chest radiographs (DR). Database Nodule cases Number of nodules Normal cases
The GDH database contains 238 DRs, 168 of them contain nodules and
70 are normal. The SZH database contains 240 CXR images and all of JSRT-A 140 140 0
JSRT-B 140 140 107
them contain nodules. For GDH datasets, most of the nodules are small.
GDH 168 577 70
The number of nodules with length 0–10 mm, 10–20 mm, 20–30 mm SZH 240 457 0
and larger than 30 mm are 455, 62, 48 and 12, respectively. For the

2
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Fig. 2. The block diagram of the proposed method.

For testing part, the lung field was firstly segmented and the ribs to pull the gray-level intensity of the second brightest pixel in the lung
were suppressed. Then the image was enhanced to increase the visibi- field area to one and depress the second darkest pixels to zero. Other
lity of lung nodules. After that the image was processed by the pre- pixels were linearly stretched.
trained CNNs. Finally, morphological processing was employed as After gray-level intensity stretching, the images still showed large
postprocessing to remove small false positives. difference across various databases. A histogram matching [42] method
was employed to make all images similar in different databases. Fifty
images with clear visibility of lung nodules were manually selected and
2.1. Lung field segmentation and rib suppression
the histogram of their average image was used as the standard histogram.
Histograms of all images were matched to the standard histogram.
The lung field segmentation and rib suppression methods were pre-
sented in our previous studies [40,41]. The active shape model (ASM) was
employed for lung field segmentation. The principal component analysis 2.3. Network structure
(PCA) was employed to model the rib. The suppressed image was obtained
by subtracting the modelled ribs from the original image. The suppression As shown in Fig. 3, our CNN consists of three dense blocks [43]. For
method could substantially enhance the visibility of nodules. each dense block, shown as the right image in Fig. 3, the convolutional
block consists of three convolutional layers, with size 3 × 3, 1 × 1 and
2.2. Image enhancement 3 × 3, respectively. The depth and growth rate of the dense block are set
as 13 and 12. Each convolutional block in the dense block is followed by
Since the brightness of images is different in the database, nor- rectified linear units (ReLUs) and drop-out layers. Average-pooling of size
malization of gray-level intensity is necessary by stretching the gray- 2 × 2 was applied at the end of dense block. The final layer is a global
level intensities. However, most images in the JSRT database were average-pooling layer, which output the extracted deep features for nodule
edited to protect the information of patients. For example, the areas of detection. All network parameters are randomly initialized according to a
patients’ information are covered by black rectangles. This leads to normal distribution with variance 0.01. The network was trained using
some problems for the stretching operations (the darkest pixel is always stochastic gradient descent with learning rate 0.001, which dropped
in the black rectangles). In addition, the images are digitized from the 0.00001 every iteration. The cost function was defined as follows:
chest films. There are occasionally white borders around the image, B
which would be the brightest pixel. To overcome these problems, we C (l , s ) = li log si + (1 li ) log (1 si )
implemented an efficient gray-level stretching operation. The idea was i=1 (1)

3
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Fig. 3. The architecture of proposed CNN. DenseNet

where s is the assigned pixel probability score, l is the reference pixel label 2.5. CNNs fusion
and B is the size of mini-batch.
As shown in Fig. 4, we designed the following four strategies to fuse
2.4. Network training the three CNNs for final decision.

2.4.1. Pixel-based patch extraction and augmentation 1) Voting-Fusion: In this scheme, each CNN is considered as the voter
As the size of lung nodules is different among various images, with same weight. If majority of the CNNs classify a patch as lung
therefore, we extracted patches with different resolutions, i.e. nodule, the patch should be predicted as lung nodule. Otherwise, the
112 × 112 , 168 × 168 and 224 × 224 , centred at each pixel in lung field patch should be classified as non-nodule.
and assign them the label of the central pixel. To address the bordering 2) Committee-Fusion: In this scheme, we connected the output of the
effect, we expand the image for 112 pixels at each border and filled the fully connected layer of each CNN to a classification layer that
expanded border with zero. Patches of all resolutions were resized into consists of an additional fully connected layer with sigmoid function
224 × 224 as the input of deep network. [44,45]. The final prediction of CNNs is the combination of the
Since the databases were small, additional training samples were output of each CNN, as shown in Fig. 4(a).
generated by random rotations (0–360 degrees) and mirroring. This is 3) Late-Fusion: The late-fusion method [26,46,47] concatenates the
especially important as there are extremely few lung nodule examples outputs of the Global average pooling layers and connects the con-
in the training set, which leads to unbalanced number of nodule and catenated outputs directly to the classification layer (see Fig. 4(a)).
non-nodule patches. Since the number of normal patches is much larger The classification layer can thus learn the multi-resolution char-
than that of the nodule patches, we only perform data augmentation on acteristics by combining the outputs of CNNs. In this configuration,
nodule patches and randomly selected the same number of normal the parameters of the convolutional layers for different CNNs are
patches for training and validation. shared.
Five-fold cross-validation was employed to evaluate the proposed 4) Full-Fusion: the full fusion method [48] was proposed for the in-
network. The images in the database were divided into five folds. When tegration of handcrafted and deep learning features, and it can also
one fold was used for testing, three and one of the remaining folds were be employed for the fusion of multi-resolution inputs. As shown in
used for training and validation, respectively. The numbers of nodules Fig. 4(b), the information of different resolutions interchanges
and normal patches are roughly the same. across every dense block. After global average pooling, the multi-
resolution features are concatenated and connected to the classifi-
2.4.2. Training cation layer. Compared to the previously mentioned three fusion
The training processing was given below methods, the full-fusion method not only combined the multi-re-
solution information at the end of network, but shared the para-
meters between every dense block.
for each epoch
for each training data
propagate error through the network; 2.6. Postprocessing
adjust the weights;
calculate the accuracy over training data;
For each pixel in the image, a patch was extracted and a classifi-
for each validation data
calculate the accuracy over the validation data;
cation result was given by the CNNs. Therefore, given a CXR image, a
if validation accuracy continuously decrease for ten epochs lung nodule possibility map could be obtained. While solitary white
exit training; pixels (false alarms) could appear in the lung nodule possibility map, a
else morphological operation was employed as postprocessing. A 5 × 5
continue training
close operation followed by a 5 × 5 open operation was employed in

4
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Global pooling

Fig. 4. The proposed four CNNs fusion methods (a. voting, committee and late fusion; b. full fusion).

this paper. The close operation was used to link white pixels near each radiologist shows nearly a hundred percent sensitivity with one FPs/
other. The open operation was used to remove small solitary white image. FAUC shows the performance of CNNs on classifying candidates
pixels. as nodules or non-nodules, while R-CPM shows the performance of CAD
at operating points likely used in practice. The definitions of FROC
curve were given as follows.
2.7. Evaluation metrics FROC curve is a tool for characterizing the performance of a free-
response system at all decision thresholds simultaneously. It mainly
Two performance metrics, i.e. area under the Free-Response evaluates the cost (the number of FPs) of a method to achieve proper
Receiver Operating Characteristic (FROC) curve (FAUC) and Refined sensitivities. The horizontal axis of FROC curve represents the number
Competition Performance Metric (R-CPM) were applied in this paper. of FPs and the vertical axis represents the sensitivity of the detection
Since the FROC points are discrete, we used interpolation to calculate method. The sensitivity here means the percentage of the nodules that
FAUC. While the original CPM measures the average sensitivity at seven can be detected by the proposed method.
operation points of FROC curve: 1/8, 1/4, 1/2, 1, 2, 4 and 8 FPs per When measuring the detection performance, a detection is decided
image, R-CPM only measures the average sensitivity at four operating to be correct when the overlap of the detected nodule and the ground
points of the FROC curve: 1/8, 1/4, 1/2 and 1 FPs per image, since the truth was larger than 0.2. Similar strategies were employed in [17,18].

5
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Table 2
The performance of three fusion methods on JSRT-B.
Methods FAUC R-CPM

Best single CNN 0.9501 0.9555


Voting-fusion 0.9717 0.9735
Committee-fusion 0.9763 0.9810
Late-fusion 0.9792 0.9825
Full-fusion 0.9823 0.9875

Table 3
The performance of the proposed method, state-of-the-art methods and radi-
ologist on JSRT-A and B.
CADe methods FAUC R-CPM

JSRT-A Handcrafted Chen et al. [17] 0.4082 0.3825


feature Li et al. [18] 0.5426 0.3354
Deep feature U-net [39] 0.9328 0.9312
Fig. 5. The FROC curves of the proposed three CNNs at different resolution.
Best single CNN (DenseNet) 0.9035 0.9050
[43]
Full-fusion method 0.9568 0.9625
3. Experimental results
JSRT-B Handcrafted Chen et al. [17] 0.5301 0.5175
feature Li et al. [18] 0.3593 0.3300
3.1. Performance of single and fused CNN Deep feature U-net [39] 0.9547 0.9475
Best single CNN (DenseNet) 0.9501 0.9555
As JSRT-B consist of images from both normal and nodule cases, it is [43]
Full-fusion method 0.9823 0.9875
firstly used to evaluate the performance of three single CNNs. The FROC
Radiologist [49] 0.8246 0.8175
curves were given in Fig. 5. One can observe from the figure that the
performance of each single CNN was similar. All three CNNs can detect
over 90% of lung nodules in JSRT-B when the FPs/image was one. The 3.2. Performance comparison with state of the art
CNN at resolution 168 × 168 showed better performance, especially at
low FPs/image. We evaluated the proposed method on both JSRT-A and B using
The FROC curves of all three fusion methods and the best single five-fold cross-validation. Table 3 gives the details of FAUC and R-CPM
CNN on JSRT-B was shown in Fig. 6. The orange, yellow, purple, green of the proposed method and previous works. While our previous work
and blue curves represent voting-fusion, committee-fusion, late-fusion, [18] achieved the best performance in literature on JSRT-A, the pro-
full-fusion and the best single CNN (at resolution 168 × 168), respec- posed method achieved substantial improvements. The state-of-the-art
tively. The details of FAUC and R-CPM were given in Table 2. All fusion method - U-net [39], which often achieves competitive performance,
methods achieved better performance than single CNN (FAUC = 0.950, was also included for comparison. In this work, considering the limited
R-CPM = 0.956). The full-fusion method shows the best performance number of images, we used patch-based U-net (patch size 224 × 224 ) for
(FAUC = 0.9823, R-CPM = 0.9875), which is method shows the best nodule detection. Fig. 7 shows the performance of the mentioned
performance, which is slightly better than other three fusion methods methods and radiologist in the obvious (solid curves), intermediate
(about 1% higher for FAUC and R-CPM). Therefore, we chose full-fu- (dashed-curves) and subtle (dot-curves) cases in JSRT-A. The black,
sion as the fusion method for the following experiments. blue, red and green curves are the FROC curves of the full-fusion
method, U-net, state-of-the-art CADe method [18] and radiologists

Fig. 7. The FROC curves of the proposed method, the state-of-the-art method
Fig. 6. The FROC curves of the proposed three CNNs fusion methods. (For in- and average performance of radiologists in obvious, intermediate and subtle
terpretation of the references to colour in the text, the reader is referred to the cases in JSRT-A. (For interpretation of the references to colour in the text, the
web version of this article.) reader is referred to the web version of this article.)

6
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Fig. 8. The FROC curves of the proposed methods and the state-of-the-art
methods and average performance of radiologists. (For interpretation of the Fig. 10. The cross-databases FROC curves of the proposed method.
references to colour in the text, the reader is referred to the web version of this
article.)
Table 5
The cross-database performance of the proposed method.
Database Model FAUC R-CPM

JRST-A Trained on JRST-A 0.9568 0.9625


JRST-B Trained on JRST-B 0.9823 0.9875
GZH Trained on JRST-B, fine-tuned on GZH 0.7169 0.6900
SZH Trained on JRST-B, fine-tuned on SZH 0.7466 0.7575

Fig. 9. The FROC curves (without fine-tuning) on GDH and SZH.

Table 4
Comparison of the performances on GDH and SZH without fine-tuning.
Method Database FAUC R-CPM

Full-fusion GDH 0.3847 0.3775


SZH 0.4213 0.4025 Fig. 11. The cross-databases performance for different nodule sizes.
Best single CNN GDH 0.2135 0.2025
SZH 0.2892 0.2450
Li et al. [18] GDH 0.0610 0.0950 U-net [39], state-of-the-art [17] and [18], respectively, while the green
curve represent the performance of radiologist. [17,18] achieved very
competitive performance when using handcrafted features. U-net [39]
[49], respectively. According to [49], 20 radiologists participated the is one of the most wildly used deep network in medical image seg-
test, nine of them were chest radiologists, 11 of them were general mentation. The radiologist’s performance was directly referred from
radiologists. The FROC curve represents the average performance of [49]. The details of FAUC and R-CPM were given in Table 3. It can be
these radiologists. For radiologists, the performance decreased with the seen that, except the deep learning-based methods, the state-of-the-art
subtle degree of lung nodules obviously. For our previous method, the CADe approaches cannot match the radiologist’s performance. They
performances were worse than the radiologist, but showed similar were far from the requirement of real clinical practice. The proposed
pattern. For the proposed method and U-net method, the performances deep learning-based method shows much better performance than
were much better than the state-of-the-art method and radiologist, the previous works. Compared to DenseNet and U-net, the proposed
performance decrease is small that we can say the performances are method also showed better performance, while patch-based DenseNet
similar for obvious and subtle cases. In contrast, the proposed method and U-net showed similar performance. Compared to the average per-
showed better performance than U-net. formance of radiologists (FAUC = 0.82, R-CPM = 0.82), the proposed
The FROC curves obtained on JSRT-B were shown in Fig. 8. The method also shows higher sensitivity at low FPs per image
blue, orange, yellow and purple curves represent the proposed method, (FAUC = 0.98, R-CPM = 0.99). When FPs/image is larger than 0.2, the

7
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Fig. 12. The examples of lung nodule detection (first row: JSRT-A; second row: GDH and SZH. The red circles annotated the lung nodules. The blue circles were the
detection of the proposed method). (For interpretation of the references to colour in this figure legend, the reader is referred to the web version of this article.)

sensitivity of the proposed method was nearly one (0.99). images of these two new databases to fine-tune the model (10 epochs).
We employed five-fold cross-validation to test the model (three folds for
fine-tuning, one for validation and one for testing). The FROC curves of
3.3. Cross databases performance
the fine-tuned model on GDH and SZH were given in Fig. 10, where the
FROC curves on JSRT-A and B were also listed. The detail of FAUC and
The proposed method achieved state-of-the-art performance on both
R-CPM were given in Table 5. The performances of the fine-tuned
JSRT-A and B. In this section, we evaluated the proposed method with
models have a significant lift over the model trained on JSRT-B only.
GDH and SZH. We firstly tested the model trained on JSRT-B (80% for
However, it seems that the model performances on GDH and SZH were
training, 20% for validation) and its FROC curves were presented in
still worse than that on JSRT-A and B.
Fig. 9 (the dot and solid curves represent the best single CNN and full-
To find out why the performances of the proposed network on GDH
fusion model, respectively). We also gave the performance of [18] when
and SZH were worse, we evaluated the performance of the network
the images in GDH were not used for training and validation. The de-
trained on JSRA-B as a function of nodule sizes. The performance
tailed performances were given in Table 4. It can be seen that when
variances related to nodule size was shown in Fig. 11. It can be seen
using multi-resolution and full-fusion method, the model performance
that 1) the model performed worse on small nodules than large ones. 2)
were much better than single CNN (0.16 and 0.13 higher FAUC, and
The fine-tuned models showed similar FAUC variances across different
0.17 and 0.16 higher R-CPM for database GDH and SZH, respectively)
databases, i.e. different nodule sizes. When the size of nodule is smaller
and previous handcrafted feature method [18] (0.36 and 0.31 higher
than 10 mm and 10–20 mm, the FAUC on JSRT-A is about 0.15 and 0.05
for FAUC and R-CPM, respectively).
higher than that on GDH and SZH, respectively. For the nodules larger
However, the model trained on JSRT-B still did not achieve sa-
than 20 mm, the model showed no difference across all databases.
tisfying performance on these two databases. Therefore, we used the

8
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

Fig. 13. The examples of false positives (marked by blue curve only). (For interpretation of the references to colour in this figure legend, the reader is referred to the
web version of this article.)

According to the size distributions shown in Fig. 1, the percentage of performance on nodules larger than 20 mm. For each CNN, we only
nodules smaller than 10 mm for JSRT-A, GDH and SZH is 20%, 80% and used hundreds cases for training and validation. If more data is avail-
80%, respectively. Therefore, the model showed worse overall perfor- able to train the network, the detection accuracy and robustness should
mance on GDH and SZH. 3) While the multi-resolution fusion methods be further improved.
achieved similar detection performance to single CNN (DenseNet) on There were still some false positives detected by the proposed
nodules larger than 20 mm, its performance on nodules smaller than method. As shown in Fig. 13, the areas marked by blue curve only were
20 mm was significantly better than DenseNet, i.e. the improvement of falsely detected lung nodules. These false positive areas are similar with
FAUC was 4%, 13% and 12% for JSRT-A, GDH and SZH, respectively. real lung nodules in terms of intensity, texture and shape, and are thus
difficult to discriminate.

4. Discussion
5. Conclusion and future work
One primary contribution of the proposed method is the employ-
In this paper, a CNN based lung nodule detection approach for CXR
ment of CNNs for lung nodule detection on CXR. The experimental
is proposed. Firstly, the CXR images were preprocessed by lung filed
results showed that using CNNs for features extraction and pixel clas-
segmentation and rib suppression. Then patches were extracted for
sification achieved a significant improvement to the sensitivity of lung
every pixel in lung field, three CNNs were trained at different image
nodule detection. Compared to handcrafted feature-based machine
resolutions. Finally, feature fusion method was used to combine all
learning methods [17,18], CNN based method can learn the most sui-
information extracted at different resolutions. Four fusion methods
table features via training. Therefore, the features extracted by CNN are
were tested, the full-fusion method showed the best performance. The
often better than handcrafted features. The proposed method employed
proposed method can detect 99% lung nodules on JSRT database when
multi-resolution technology, i.e. the image was analysed at three dif-
FPs/image was under 0.2. The performance was better than the average
ferent resolutions, and three CNNs were trained at each resolution.
radiologist. Compared with previous studies, the proposed method
When single CNN achieves comparable result with radiologist, the
achieved much higher sensitivity at much lower FPs/image. The pro-
method combining all three CNNs achieved even better performance
posed method was accurate, robust and has the potential to be used in
(The examples of detection result are shown in Fig. 12).
real clinical practice.
Another contribution of the proposed method is that we compared
Our future work will mainly try to employ large CXR database
four different fusion methods. The result showed that the full-fusion
collected from different hospitals and machines for network training, to
method achieves the best performance. The reason is that full-fusion
improve the detection performance and robustness of the network. The
method interchanges features extracted from patches of different re-
architecture of CNN should also be optimized for larger-scale database.
solutions at middle layers, and combined all information in the end,
Furthermore, the cases that the lung nodules are located outside lung
which make the classifier has more comprehensive information for lung
field should be considered as well.
nodule detection. The late- and committee-fusion showed slightly lower
performance than full-fusion. Compared to voting-fusion, in which each
classifier has equal contribution, the late- and committee-fusion method 6. Declaration of Competing Interest
is more reasonable.
The experimental result showed that, on the rib suppressed CXR of We declare that all authors have no conflicts of interest in the au-
JSRT database, the method showed better performance than radi- thorship or publication of this contribution.
ologist. On other two different databases, the method also showed sa-
tisfactory performance after fine-tuning, which indicated the robustness Acknowledgements
of the proposed method. When considering nodule size, the model
showed worse performance on nodules smaller than 20 mm, and similar The work is supported by National Natural Science Foundation of

9
X. Li, et al. Artificial Intelligence In Medicine 103 (2020) 101744

China (Grant No. 61702337, 61672357 and U1713214) and the Science [24] Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolu-
and Technology Funding of Guangdong Province (2018A050501014). tional neural networks. International Conference on Neural Information Processing
Systems 2012:1097–105.
We thank Prof. Qingmao Hu, Shenzhen Institutes of Advanced [25] Li Y, Shen L, Yu S. HEp-2 specimen image segmentation and classification using
Technology, for providing us the CXR images. very deep fully convolutional network. IEEE Trans Med Imaging 2017;3:1561–72.
[26] Setio AAA, Ciompi F, Litjens G, Gerke P, Jacobs C, van Riel SJ, et al. Pulmonary
nodule detection in ct images: false positive reduction using multi-view convolu-
References tional networks. IEEE Trans Med Imaging 2016;35:1160–9.
[27] Jiang H, He M, Wei Q, Gao M, Yan L. An automatic detection system of lung nodule
[1] Chunhua X, Keke H, Yong S, Like Y, Zhibo H, Ping Z. Early diagnosis of solitary based on multi-group patch-based deep learning network. IEEE J Biomed Health
pulmonary nodules. J Thorac Dis 2013;5:830–40. Inform 2017;22:1227–37.
[2] Pedersen JH, Ashraf H, Dirksen A, Bach K, Hansen H, Toennesen P, et al. The Danish [28] Ssm S, Erdogmus D, Gholipour A. Auto-context convolutional neural network
randomized lung cancer CT screening trial—overall design and results of the pre- (Auto-Net) for brain extraction in magnetic resonance imaging. IEEE Trans Med
valence round. J Thorac Oncol 2009;4:608–14. Imaging 2017;36:2319–30.
[3] Pastorino U, Rossi M, Rosato V, Marchianò A, Sverzellati N, Morosi C, et al. Annual [29] Milletari F, Ahmadi SA, Kroll C, Plate A, Rozanski V, Maiostre J, et al. Hough-CNN:
or biennial CT screening versus observation in heavy smokers: 5-year results of the deep learning for segmentation of deep brain regions in MRI and ultrasound.
MILD trial. Eur J Cancer Prev 2012;21:308–15. Comput Vis Image Underst 2017;164:92–102.
[4] Infante M, Cavuto S, Lutman FR, Brambilla G, Chiesa G, Ceresoli G, et al. A ran- [30] Qi D, Hao C, Yu L, Jing Q, Heng PA. Multilevel contextual 3-D CNNs for false
domized study of lung cancer screening with spiral computed tomography: three- positive reduction in pulmonary nodule detection. IEEE Trans Biomed Eng
year results from the DANTE trial. Am J Respir Crit Care Med 2009;180:445–53. 2017;64:1558–67.
[5] van Beek EJ, Mirsadraee S, Murchison JT. Lung cancer screening: computed to- [31] Ding J, Li A, Hu Z, Wang L. Accurate pulmonary nodule detection in computed
mography or chest radiographs? World J Radiol 2015;7:189–93. tomography images using deep convolutional neural networks arXiv preprint
[6] Austin JH, Romney BM, Goldsmith LS. Missed bronchogenic carcinoma: radio- arXiv:1706.04303 2017.
graphic findings in 27 patients with a potentially resectable lesion evident in ret- [32] Avni U, Greenspan H, Konen E, Sharon M, Goldberger J. X-ray categorization and
rospect. Radiology 1992;182:115–22. retrieval on the organ and pathology level, using patch-based visual words. IEEE
[7] Shah PK, Austin JH, White CS, Patel P, Haramati LB, Pearson GD, et al. Missed non- Trans Med Imaging 2011;30:733.
small cell lung cancer: radiographic findings of potentially resectable lesions evi- [33] Jaeger S, Karargyris A, Candemir S, Folio L, Siegelman J, Callaghan F, et al.
dent only in retrospect. Radiology 2003;226:235–41. Automatic tuberculosis screening using chest radiographs. IEEE Trans Med Imaging
[8] Kobayashi T, Xu XW, Macmahon H, Metz CE, Doi K. Effect of a computer-aided 2014;33:233–45.
diagnosis scheme on radiologists’ performance in detection of lung nodules on [34] Shin H-C, Roberts K, Lu L, Demner-Fushman D, Yao J, Summers RM. Learning to
radiographs. Radiology 1996;199:843–8. read chest x-rays: recurrent neural cascade model for automated image annotation.
[9] Macmahon H, Engelmann R, Behlen FM, Hoffmann KR, Ishida T, Roe C, et al. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Computer-aided diagnosis of pulmonary nodules: results of a large-scale observer 2016:2497–506.
Test1. Radiology 1999;213:723–6. [35] X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers, Chestx-ray8:
[10] Matsumoto T, Doi K, Kano A, Nakamura H, Nakanishi T. [Evaluation of the po- Hospital-scale chest x-ray database and benchmarks on weakly-supervised classi-
tential benefit of computer-aided diagnosis (CAD) for lung cancer screenings using fication and localization of common thorax diseases, presented at the 2017 IEEE
photofluorography: analysis of an observer study]. Nihon Igaku Hoshasen Gakkai Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
Zasshi 1993;53:1195–207. [36] Xue Z, You D, Candemir S, Jaeger S, Antani S, Long LR, et al. Chest x-ray image
[11] Xu X-W, Doi K, Kobayashi T, MacMahon H, Giger ML. Development of an improved view classification. Proceedings of the IEEE Symposium on Computer-Based
CAD scheme for automated detection of lung nodules in digital chest images. Med. Medical Systems, vol. 2015 2015:66–71.
Phys 1997;24:1395–403. [37] Boussaid H, Kokkinos I. Fast and exact: ADMM-based discriminative shape seg-
[12] Giger ML, Ahn N, Doi K, MacMahon H. Computerized detection of pulmonary no- mentation with loopy part models. 2014 IEEE Conference on Computer Vision and
dules in digital chest images: use of morphological filters in reducing false-positive Pattern Recognition (CVPR) 2015:4058–65.
detections. Med Phys 1990;17:861–5. [38] Hermann S. Evaluation of scan-line optimization for 3D medical image registration.
[13] Wei J, Hagihara Y, Shimizu A, Kobatake H. Optimal image feature set for detecting IEEE Conference on Computer Vision and Pattern Recognition 2014:3073–80.
lung nodules on chest X-ray images. Workshop on Computer-Aided Diagnosis Proc. [39] Ronneberger O, Fischer P, Brox T. U-net: convolutional networks for biomedical
of SPIE 2002:1–6. image segmentation. International Conference on Medical Image Computing and
[14] Schilham AM, van Ginneken B, Loog M. A computer-aided diagnosis system for Computer-Assisted Intervention 2015:234–41.
detection of lung nodules in chest radiographs with an evaluation on a public da- [40] Li X, Luo S, Hu Q, Li J, Wang D. Rib suppression in chest radiographs for lung
tabase. Med Image Anal 2006;10(Apr):247–58. nodule enhancement. Proceeding of the 2015 IEEE International Conference on
[15] Hardie RC, Rogers SK, Wilson T, Rogers A. Performance analysis of a new computer Information and Automation, Lijiang, China 2015:50–5.
aided detection system for identifying lung nodules on chest radiographs. Med [41] Li X, Luo S, Hu Q, Li J, Wang D, Chiong F. Automatic lung field segmentation in X-
Image Anal 2008;12(Jun):240–58. ray radiographs using statistical shape and appearance models. J Med Imaging
[16] Chen S, Suzuki K, MacMahon H. Development and evaluation of a computer-aided Health Inform 2016;6:338–48.
diagnostic scheme for lung nodule detection in chest radiographs by means of two- [42] Morovic J, Shaw J, Sun PL. A fast, non-iterative and exact histogram matching
stage nodule enhancement with support vector classification. Med Phys algorithm. Pattern Recognit Lett 2002;23:127–35.
2011;38:1844. [43] Huang G, Liu Z, Weinberger KQ, van der Maaten L. Densely connected convolu-
[17] Chen S, Suzuki K. Computerized detection of lung nodules by means of “Virtual tional networks arXiv preprint arXiv:1608.06993 2016.
dual-energy” radiography. IEEE Trans Biomed Eng 2013;60:369–78. [44] Cireşan D, Meier U, Masci J, Schmidhuber J. Multi-column deep neural network for
[18] Li X, Shen L, Luo S. A solitary feature-based lung nodule detection approach for traffic sign classification. Neural Netw 2012;32:333.
chest X-Ray radiographs. IEEE J Biomed Health Inform 2018;22:516–27. [45] Ginneken BV, Setio AAA, Jacobs C, Ciompi F. Off-the-shelf convolutional neural
[19] Oda S, Awai K, Suzuki K, Yanaga Y, Funama Y, MacMahon H, et al. Performance of network features for pulmonary nodule detection in computed tomography scans.
radiologists in detection of small pulmonary nodules on chest radiographs: effect of IEEE International Symposium on Biomedical Imaging 2015:286–9.
rib suppression with a massive-training artificial neural network. AJR Am J [46] Prasoon A, Petersen K, Igel C, Lauze F, Dam E, Nielsen M. Deep feature learning for
Roentgenol 2009;193(Nov):W397–402. knee cartilage segmentation using a triplanar convolutional neural network.
[20] Li F, Engelmann R, Pesce LL, Doi K, Metz CE, MacMahon H. Small lung cancers: Medical Image Computing & Computer-assisted Intervention: Miccai International
improved detection by use of bone suppression imaging-comparison with dual-en- Conference on Medical Image Computing & Computer-assisted Intervention
ergy subtraction chest radiography. Radiology 2011;261:937–49. 2013:246–53.
[21] Li F, Hara T, Shiraishi J, Engelmann R, MacMahon H, Doi K. Improved detection of [47] Karpathy A, Toderici G, Shetty S, Leung T, Sukthankar R, Fei-Fei L. Large-scale
subtle lung nodules by use of chest radiographs with bone suppression imaging: video classification with convolutional neural networks. IEEE Conference on
receiver operating characteristic analysis with and without localization. AJR Am J Computer Vision and Pattern Recognition 2014:1725–32.
Roentgenol 2011;196(May):W535–41. [48] Li X, Shen L, Shen M, Qiu CS. Integrating handcrafted and deep features for optical
[22] Novak RD, Novak NJ, Gilkeson R, Mansoori B, Aandal GE. A comparison of com- coherence tomography based retinal disease classification. IEEE Access 2019;PP:1.
puter-aided detection (CAD) effectiveness in pulmonary nodule identification using [49] Shiraishi J, Katsuragawa S, Ikezoe J, Matsumoto T, Kobayashi T, Komatsu K, et al.
different methods of bone suppression in chest radiographs. J Digit Imaging Development of a digital image database for chest radiographs with and without a
2013;26(August):651–6. lung nodule: receiver operating characteristic analysis of radiologists’ detection of
[23] An D, Meier U, Masci J, Schmidhuber J, Rgen. 2012 Special Issue: multi-column pulmonary nodules. Am J Roentgenol 2000;174:71–4.
deep neural network for traffic sign classification. Neural Netw 2012;32:333–8.

10

You might also like