Enhanced YOLO-V3 for Remote Sensing Detection
Enhanced YOLO-V3 for Remote Sensing Detection
Article
Improved YOLO-V3 with DenseNet for Multi-Scale
Remote Sensing Target Detection
Danqing Xu and Yiquan Wu *
College of Electronic and Information Engineering, Nanjing University of Aeronautics and Astronautics,
Nanjing 211106, China; xudanqing@[Link]
* Correspondence: imagestrong@[Link]; Tel.: +86-137-7666-7415
Received: 18 June 2020; Accepted: 28 July 2020; Published: 31 July 2020
Abstract: Remote sensing targets have different dimensions, and they have the characteristics of
dense distribution and a complex background. This makes remote sensing target detection difficult.
With the aim at detecting remote sensing targets at different scales, a new You Only Look Once
(YOLO)-V3-based model was proposed. YOLO-V3 is a new version of YOLO. Aiming at the defect of
poor performance of YOLO-V3 in detecting remote sensing targets, we adopted DenseNet (Densely
Connected Network) to enhance feature extraction capability. Moreover, the detection scales were
increased to four based on the original YOLO-V3. The experiment on RSOD (Remote Sensing
Object Detection) dataset and UCS-AOD (Dataset of Object Detection in Aerial Images) dataset
showed that our approach performed better than Faster-RCNN, SSD (Single Shot Multibox Detector),
YOLO-V3, and YOLO-V3 tiny in terms of accuracy. Compared with original YOLO-V3, the mAP
(mean Average Precision) of our approach increased from 77.10% to 88.73% in the RSOD dataset.
In particular, the mAP of detecting targets like aircrafts, which are mainly made up of small targets
increased by 12.12%. In addition, the detection speed was not significantly reduced. Generally
speaking, our approach achieved higher accuracy and gave considerations to real-time performance
simultaneously for remote sensing target detection.
Keywords: remote sensing image; target detection; multi-scale; YOLO-V3; convolutional neural
network; DenseNet
1. Introduction
Recently, remote sensing images [1–4] have attracted more research in the field of computer
version (CV) with the rapid development of satellite and imaging technology. There is a significant
value on information extraction of remote sensing images. Remote sensing target detection [5–7] has
important and extensive applications in military, navigation, salvage, and other aspects, which requires
high speed and accuracy for target detection algorithms.
The rapid development of computer technology makes it possible for the applications of the
convolutional neural network (CNN) [8–11], which requires high computing power. Compared with
traditional target detection algorithms like HOG-SVM (Histogram of Oriented Gradients-Support Vector
Machine) [12,13], DPM (Deformable Parts Model) [14,15], and HOG-Cascade [16,17], CNN-based
target detection algorithms have great advantages in many aspects such as speed and accuracy.
Convolutional neural network (CNN) is a kind of feed forward neural network with convolutional
computing and it usually has a deep structure. It is one of the most important components of deep
learning [18–20]. Recently, the research of deep learning in target detection has become a hot spot.
The CNN-based target detection models can mainly be divided into two categories, which include the
two-stage ones and the one-stage ones.
Currently, the two-stage ones are represented by R-CNN [21], and then Fast R-CNN [22,23],
Faster R-CNN [24,25], and Mask R-CNN [26,27], which have been developed on the basis of it. As the
name implies, the two-stage target detection algorithms divide the detection processes into two steps.
First, the Region Proposed Network (RPN) [28–30] is used to extract the information of the targets
and then the detection layers predict location and category information of the targets. The other
ones are one-stage target detection algorithms including SSD (Single Shot Multibox Detector) [31–33],
DSSD (Deconvolution Single Shot Multibox Detector) [34], FSSD (Feature Fusion Single Shot Multibox
Detector) [35], YOLO [36], YOLO-V2 [37], and YOLO-V3 [38]. Instead of using the region proposed
network (RPN), the one-stage algorithms obtain the predictive information of location and category
directly. Therefore, they are also called the regression-based algorithms and they can usually achieve
higher detection speed than the two-stage ones. At present, numerous state-of-the-art target detection
models with higher speed are proposed based on YOLO such as YOLO-V3 tiny [39] and TF-YOLO [40].
Therefore, the accuracy of them is not satisfactory.
From the current research, remote sensing target detection usually faces the following challenges:
one is that remote sensing targets are usually small and take up fewer pixels, which makes it difficult to
extract features. The second challenge is that remote sensing images are usually disturbed by shadow,
light, and other external factors. In addition, the scales of the remote sensing targets are usually
different. To solve these problems, researchers have made unremitting efforts.
In order to realize remote sensing target detection, the inchoate research is mainly based on
template matching, which is to match the target with a specific template for detection. For example,
Weber et al. [41] proposed a method of making use of image analysis to extract coastline templates and
adopted this method to detect oil tanks. It achieved good results. However, although the method of
template matching is simple and effective, its overall robustness is poor and it is sensitive to the shapes
of the targets and geometric deformation. The algorithms based on image analysis are to judge whether
each region of the remote sensing image has a target by segmentation and classification. For example,
Feng et al. [42] proposed the algorithm named multi-resolution segmentation, which segmented remote
sensing images into multiple regions for detection by three parameters: shape, scale, and density.
Compared with the method of template matching, this method is more flexible and can combine
contextual semantic information, which has achieved good results in some tasks. However, this kind
of algorithm still needs to be designed manually for segmentation, which is not universal.
Compared with the previous two algorithms, the remote sensing target detection algorithms
based on deep learning have better accuracy and robustness because they no longer use the features of
manual design. Sun et al. [43] extracted the region of interest with sliding windows, and then used the
features of Bag-of-Words to detect the targets. Zhang et al. [44] and Yu et al. [45] combined the prior
characteristics of the airports and coasts, respectively, with deep learning to conduct remote sensing
target detection. More commonly, researchers use existing target detection algorithms such as Faster
R-CNN in remote sensing target detection tasks. However, when these models are applied to remote
sensing target detection tasks, their performance is poor due to the factors such as illumination, cloud
cover, and complex background interference.
As an advanced target detection model, YOLO-V3 adopts a feature pyramid network (FPN) [46,47],
ResNet (Residual Network) [48], and achieves good performance in speed and accuracy. YOLO-V3
predicts targets at three different scales. Compared with the previous two versions, YOLO-V3 enhances
the capability of detecting multi-scale targets, especially small targets. Abundant improved algorithms
have been proposed since YOLO-V3 came out. References [49,50] adopted four detection layers to
enhance the performance of detecting small targets. Reference [51] adopted circular ground truth
to realize tomato detection. Reference [52] increased another shortcut connection to concatenate 2
CBLs (Convolution-Batch Normalization-Leak ReLU) between two ‘residual units’ to enhance the
performance for the feature extraction network of information transfer. Reference [40] simplified the
feature extraction network to obtain faster detection speed and adopted multiple layers concatenation
to enhance the performance of feature extraction. The above improved algorithms achieved a good
Sensors 2020, 20, 4276 3 of 23
detection effect. However, the resolution of remote sensing images is large. The scales of remote
sensing targets are small and the backgrounds are complex. These algorithms, which have excellent
performance on routine datasets, are not suitable for remote sensing target detection. Therefore, we need
to design a feature extraction network and detection networks for our proposed algorithm elaborately.
According to the characteristics of remote sensing targets, the proposed method was improved
based on the YOLO-V3 model. The main contributions in this paper include the following. (1) In order
to reduce reliance on ResNet and enhance the ability of feature information extraction, which is
inspired by DenseNet, improved densely connected units proposed to replace some of the residual
units of Darknet53. (2) To further improve the ability of detecting multi-scale remote sensing targets,
we extended the original three output layers of YOLO-V3 to 4. (3) In order to avoid gradient vanishing,
instead of five convolutional layers in each detection layers, three residual units were adopted.
The experimental results on remote sensing images show that the proposed method not only has
good performance in accuracy, but also gives attention to real-time performance for remote sensing
target detection.
The rest of this paper is as follows. In Section 2, we introduced the theory of YOLO and the
framework of YOLO-V3. In Section 3, we described the improved method of our approach in details.
Section 4 gives the experiments of the proposed algorithm on the RSOD dataset and compared the
performance of our approach with other classical algorithms. Lastly, the conclusion is shown in
Section 5.
YOLO has made greater achievements than Faster R-CNN in terms of speed, but it also brings
the low accuracy of detection. On the basis of YOLO-V1, YOLO-V2 introduces the concept of the
anchor box and runs k-means on the dataset to generate appropriate prediction boxes at the beginning.
Instead of full connected layers (FC), YOLO-V2 introduces convolutional layers in the output end.
In addition, YOLO-V2 also adopts Batch Normalized, New feature extraction network (Darknet19),
which greatly improves the performance compared with YOLO-V1.
YOLO-V3 is a further improved version based on YOLO-V2 by upgrading the original Darknet19
to Darknet53 and adopts multi-scale detection layers (three scales) to detect the targets. This allows
YOLO-V3 to detect small targets more effectively.
Sensors 2020, 20, 4276 4 of 23
UP
CBL concat CBL*5 CBL CONV
Sample
26*26*2
55
UP
CBL concat CBL*5 CBL CONV
Sample
52*52*2
55
Leaky Res
CBL CONV BN CBL CBL add
relu unit
RES Zero
CBL
Res
Res
Res
Res
n padding unit
unit
unit
unit*n
Table 1. The feature extraction network of You Only Look Once (YOLO)-V3.
Table 1. The feature extraction network of You Only Look Once (YOLO)-V3.
Layer Filter Size Output
Layer Filter Size Output
Convolutional 32 3×3 416 × 416 × 32
Convolutional 32 3×3 416 × 416 × 32
Convolutional 64 3 × 3/2 208 × 208 × 64
Convolutional 64 3 × 3/2 208 × 208 × 64
Convolutional 32 1×1
Convolutional 32 1×1
1× Convolutional 64 3×3
1× Convolutional 64 3×3
Residual
Residual 208
208××208
208××64
64
Convolutional
Convolutional
128
128
3 × 3/2
3 × 3/2
104 × 104 × 128
104 × 104 × 128
Convolutional 64 1×1
Convolutional 64 1×1
2×
2×
Convolutional
Convolutional 128
128 3 ××33
3
Residual
Residual 104
104××104104××128
128
Convolutional
Convolutional 256
256 33××3/2
3/2 5252××5252××256
256
Convolutional 128 1×1
Convolutional 128 1×1
8×
8× Convolutional
Convolutional 256
256 3 ××33
3
Residual
Residual 5252××5252××256
256
Convolutional
Convolutional 512
512 33××3/2
3/2 2626××2626××512
512
Convolutional
Convolutional 256
256 11××11
8×
8× Convolutional
Convolutional 512
512 33××33
Residual
Residual 2626××2626××512
512
Convolutional 1024
Convolutional 1024 33××3/2
3/2 1313××1313××1024
1024
Convolutional
Convolutional 512
512 11××11
4× Convolutional 1024
Convolutional 1024 33××33
Residual
Residual 1313××1313××1024
1024
[Link]
RelatedWork
Work
[Link]
3.1. ImprovedDensely
DenselyConnected
ConnectedNetwork
Network
The improvement
The improvement of ofYou
YouOnly
OnlyLook Once
Look Once(YOLO)-V3
(YOLO)-V3is mainly basedbased
is mainly on theon concept of a residual
the concept of a
network. Darknet53 uses several residual units, and the ResNet made up of these
residual network. Darknet53 uses several residual units, and the ResNet made up of these residual residual units
contains
units a large
contains number
a large of parameters
number andand
of parameters it isitresponsible for for
is responsible thethe
main calculations
main calculationsforfor
YOLO-V3
YOLO-
network.
V3 [Link] ResNet,
Unlike which
ResNet, addsadds
which the values of theof
the values subsequent layers by
the subsequent constructing
layers an identity
by constructing an
map, DenseNet [53] connects all the layers for channel merging to achieve feature
identity map, DenseNet [53] connects all the layers for channel merging to achieve feature reuse. reuse. Compared
with ResNet,
Compared theResNet,
with back propagation of the gradient
the back propagation of theisgradient
enhanced, which canwhich
is enhanced, make canbetter usebetter
make of feature
use
information
of and improveand
feature information the transmittance
improve theoftransmittance
the informationofbetween [Link] structure
the information between [Link]
DenseNet
is shown of
structure in DenseNet
Figure 2. is shown in Figure 2.
H1 H2 H3 H4
X0 X1 X2 X3 X4
Thestructure
[Link]
Figure structureof
ofDenseNet.
DenseNet.
In Figure 2, x1 , x2 , x3 , and x4 represent the feature maps of the output layers, while H1 , H2 , H3 ,
x1 nonlinear
x x transformations.
x4 H
H4Figure
and In refers to
2, the , 2 , 3 , and representThe
the network contains
feature maps l(l +
of the output layers, whilewith1 ,l
1)/2 connections
H2 H3 H4
, , and refers to the nonlinear transformations. The network contains l ( l+1) / 2
Sensors 2020, 20, x FOR PEER REVIEW 6 of 24
Sensors 2020, 20, 4276 6 of 23
connections with l layers. Each layer is connected to all the other layers. Thus, each layer can receive
all the feature
layers. maps
Each layer is connected to all the( lother
of the preceding − 1) layers. The feature
Thus, map can
each layer of each layer
receive allcan
the be expressed
feature maps
of the preceding
in Equation (2). ( l − 1 ) layers. The feature map of each layer can be expressed in Equation (2).
xl =HH
xl = l [xl [0x
, 0x,1x, 1. ,...,
. . , xxl−1
l −1
]] (2)(2)
The proposed
proposed densely
denselyconnected
connectednetwork
networkininthis
thispaper
paperborrows
borrowsfrom
from the
the idea
idea of of residual
residual units
units in
in Figure
Figure 1. The
1. The convolution,
convolution, Batch
Batch Normalization,
Normalization, andand Leaky-ReLU
Leaky-ReLU make make up CBL
up the the CBL module,
module, whilewhile
two
two
CBL CBL
modulesmodules are cascaded
are cascaded into a into a Double-CBL
Double-CBL (DCBL)(DCBL)
[Link]. We DCBL
We use the use the DCBLasmodule
module as
transport
transport layer H
layer Hi : Conv ( 1i :×Conv ( 1 × 1 × 32)-BN-ReLU-Conv
1 × 32)-BN-ReLU-Conv (3 × 3 × 64)-BN-ReLU
(3 × 3 × 64)-BN-ReLU and Conv and (1Conv
× 1 ×(1 × 1 × 64)-BN-
64)-BN- ReLU-
ReLU- × 3 ×(3128)-BN-ReLU.
Conv (3Conv × 3 × 128)-BN-ReLU. Thus, too many
Thus, toolayers
many of DenseNet
layers will lead
of DenseNet thelead
will feature maps getting
the feature maps
redundant
getting and decrease
redundant the speedthe
and decrease of detection, we set fourwe
speed of detection, layers
set for
foureach module.
layers Themodule.
for each incrementTheof
the featureof
increment maps for eachmaps
the feature layerfor
in module
each layer’DENSE 1st’ is’DENSE
in module 64 while theisincrement
1st’ 64 while the of the feature of
increment maps
the
for eachmaps
feature layerfor
in module
each layer ’DENSE 2nd’ ’DENSE
in module is 128. 2nd’ is 128.
With the
theaimaimofof reducing
reducing the the
network’s dependence
network’s on residual
dependence units, aunits,
on residual part ofathe lower
part resolution
of the lower
layers of the
resolution feature
layers of theextraction network isnetwork
feature extraction replacedis by the improved
replaced densely connected
by the improved network.
densely connected
The structure
network. The diagram
structureof the proposed
diagram feature extraction
of the proposed network isnetwork
feature extraction shown inis Figure
shown 3. in Figure 3.
X0 X1 X2 X3 X4
X0 X1 X2 X3 X4
The structure
Figure 3. The structure diagram of the feature extraction network.
To show
To show the
the structure
structure of
of our
our approach
approach in
in detail, Table 22 gives
detail, Table gives the
the feature
feature extraction
extraction network of
network of
our approach.
our approach.
Sensors 2020, 20, 4276 7 of 23
4 16 * 4 16 *
3
CBL
RES
1
RES
2
UP RES
CBL concat CBL CONV
Sample 4th 1 04 * 1 04 *
N Scale 4
RES Residual
4 units×3
1 × 1 × 64
DENSE 3 × 3 × 128
1st
UP RES
CBL concat CBL CONV
Sample 3rd 52 * 52 * N Scale 3
Residual
units×3
RES 1 × 1 × 128
4 3 × 3 × 256
DENSE
2nd
UP RES
CBL concat CBL CONV
Sample 2nd 26 * 26 * N
Scale 2
Residual
units×3
1 × 1 × 256
RES RES 3 × 3 × 512
CBL CONV
4 1st 13 * 13 * N
Scale 1
Residual
units×3
1 × 1 × 512
3 × 3 × 1024
Inspired by Faster-RCNN, YOLO-V2 and YOLO-V3 introduced the ideal of the anchor box to
predict the bounding boxes more accurately. In our approach, we ran K-means to generate the anchor
Sensors
boxes.2020,
The20,function
4276 of the K-means algorithm is conducting latitude clustering to make anchor9 boxes
of 23
and adjacent ground truth have larger IOU values, which is not directly related to the size of anchor
boxes.
3.3. K-Means for Anchor Boxes
d ( box ,centroid ) = 1 − IOU ( box ,centroid )
Inspired by Faster-RCNN, YOLO-V2 and YOLO-V3 introduced the ideal of the anchor box (3) to
predictIOU
the bounding boxes
refers to the more accurately.
intersection ratio andInitour approach,
is defined we ran K-means
in Equation (4). to generate the anchor
boxes. The function of the K-means algorithm is conducting latitude clustering to make anchor
boxes and adjacent ground truth have larger IOU Soverlapwhich is not directly related to the size of
IOUvalues,
= (4)
anchor boxes. Sunion
d(box, centroid) = 1 − IOU(box, centroid) (3)
Soverlap refers to the overlap area between the predicted box and the ground truth and Sunion
IOU refers to the intersection ratio and it is defined in Equation (4).
refers to the union area between them. The pseudocode of K-means in this paper is shown in
Algorithm 1. Soverlap
IOU = (4)
Algorithm 1: The pseudocode of K-means Sunion
1: SGiven
overlap refers to the
K cluster overlap
center area (between
points: ∈{1,predicted
Wi , Hi ), ithe 2,..., k} , W
box
i
, Hand the ground truth and Sunion refers
i refer to the width and height of
to theeach
union area between
anchor box. them. The pseudocode of K-means in this paper is shown in Algorithm 1.
2: Calculate the distance between each ground truth and each cluster center:
Algorithm 1: The pseudocode of K-means
d (box, centroid ) = 1 − IOU(box, centroid ) . Since the position of the anchor box is not fixed, the
1: Given K cluster
center pointcenter points:
of each (Wi ,truth
ground Hi ), i ∈is{1, 2, . . . , k},Wwith
coincident i , Hi refer
the to the widthcenter.
clustering and height of each anchor
box.
' 1 ' 1
2: 3:
Calculate the distance
Recalculate between
the cluster each
center ground
for truth andW
each cluster: each
i wcenter:
=cluster i
,H i = hi
d(box, centroid) = 1 − IOU(box, centroid). Since the position of the Ni anchor box isNnot
i fixed, the center point of
each ground truth is coincident with the clustering
4: Repeat step 2 and step 3 until the clusters converge. center.
3: Recalculate the cluster center for each cluster: W 0 i = N1i wi , H0 i = N1i hi
P P
Figure 5. The relationship between the number of clusters and average IOU by K-means clustering.
Figure 5. The relationship between the number of clusters and average IOU by K-means clustering.
values
[Link] to the GridcCell
the network. x and c y represent the offset of the gird relative to the upper left. The
values ofWhen
bounding boxes
detecting thecan be represented
targets, as:the values of bounding boxes based on the predicted
we need to get
values. The process is shown in Figure 6. In Figure 6, tx , t y , tw , and th represent the predicted values of
b = σ (t ) + c x
the network. cx and c y represent the offset of xthe girdx relative to the upper left. The values of bounding
boxes can be represented as: b y = σ (t y ) + c y
bx = σ(tx ) t+ cx
b y b= = p y e) +
w σ(tw
w
cy (5)
bw = pw e th t w (5)
bh = ph e
bh = ph eth
−x
σ(xσ)( x=) =1/
1(/1(1++ee−x ) )
Figure6.
Figure The final
[Link] final prediction.
prediction.
In Equation (7), λcoord refers to the weight of the coordinate error and we selected λcoord = 5 in
our model. S2 refers to the number of the grids (S × S). B refers to the number of bounding boxes
obj
per grid. Iij refers to whether there is an object that falls in the jth bounding box of the ith grid cell.
(xi , yi , wi , hi ) and (xi , yi , wi , hi ) refer to the center coordinate, height, and width of the predicted box
and the ground truth, respectively.
Sensors 2020, 20, 4276 12 of 23
λnoobj refers to the confidence penalty when there is no object and we selected λnoobj = 0.5 in our
model. Ci And Ci refer to the true and predicted confidence, respectively.
Errorcls refers to the classification error and it is defined as:
Xs2 XB X 2
obj
Errorcls = Iij (pi (c) − p̂i (c)) (9)
i=1 j=1 c∈classes
TP
Precision = (10)
TP + FP
TP
Recall = (11)
TP + FN
Mean average precision (mAP) is a performance metric for predicting target locations and
categories. The accuracy and recall are mutually restricted in practice, and there will be ambiguity
when compared separately. Therefore, in our experiment, we introduced mAP, which is one of the
most important metrics to evaluate the performance of target detection algorithms.
1. Scale diversity. Remote sensing images can be taken from hundreds of meters to nearly 10,000 meters
in height, and ground targets may be of different sizes even if they are of the same kind.
For example, ships in ports may be only tens of meters to more than 300 meters in size.
2. Perspective particularity. The perspective of remote sensing images is basically overhead, but most
of the conventional datasets are still ground level, so the mode of the same target is usually
different. The detector trained well on the conventional datasets, which may have a poor effect
on the remote sensing images.
3. Problem of small targets. Most of the remote sensing targets are small in size. As a result, the
target information is limited. The information of the targets has been lost due to the down
sampling layers of the Convolutional Neural Network (CNN). After four times of down sampling,
the feature map of the target with 24 × 24 pixels may take up only 1 pixel.
4. Problem of multi-directions. The viewing angle of remote sensing images are usually overhead,
while the directions of the targets are uncertain while there is a degree of certainty in
conventional datasets.
5. The high complexity of the background. The fields of remote sensing images are relatively large
(usually covering several square kilometers). The fields of vision may contain various backgrounds,
which will produce strong interference to the target detection.
Sensors 2020, 20, 4276 13 of 23
Based on the above reasons, it is often difficult to train an ideal target detector from conventional
datasets for target detection tasks of remote sensing images. A special remote sensing image database
is needed.
(a)
(b)
Figure
Figure 7. The
7. The samples
samples of of
thethedatasets:
datasets: (a)
(a) the
the samples
samplesofofremote
remote sensing object
sensing detection
object (RSOD)
detection (RSOD)
dataset and (b) the samples of dataset of object detection in aerial images (UCS-AOD) dataset.
dataset and (b) the samples of dataset of object detection in aerial images (UCS-AOD) dataset.
Target Amount
Dataset Class Image Instances
Small Medium Large
Aircraft 446 4993 3714 833 446
Training Oil tank 165 1586 724 713 149
Set Overpass 176 180 0 0 180
Playground 189 191 0 12 179
Aircraft 176 1257 741 359 157
Test Oil tank 63 567 257 213 97
Set Overpass 36 41 0 0 41
Playground 49 52 0 0 52
Metric (%)
Method Backbone mAP FPS
Aircraft Oil Tank Overpass Playground
(IOU = 0.5)
Faster RCNN VGG-16 85.85 86.67 88.15 90.35 87.76 6.7
SSD VGG-16 69.17 71.20 70.23 81.26 72.97 62.2
DSSD ResNet-101 72.12 72.49 72.10 83.56 75.07 6.1
ESSD VGG-16 73.08 72.94 73.61 84.27 75.98 37.3
YOLO-V2 DarkNet19 62.35 67.74 68.38 78.51 69.25 35.6
YOLO-V3 DarkNet53 74.30 73.85 75.08 85.16 77.10 29.7
YOLO-V3 tiny DarkNet19 54.14 56.21 59.28 64.20 58.46 69.8
UAV-YOLO [52] Figure 1 in [52] 74.68 74.20 76.32 85.96 77.79 30.1
DC-SPP-YOLO [54] Figure 5 in [54] 73.16 73.52 74.82 84.82 76.58 33.5
ours (Figure 3) 86.42 87.57 89.37 91.56 88.73 25.8
Table 7 shows that our approach is superior to other classical algorithms in the index of mAP.
The detection speed is not significantly reduced relative to YOLO-V3. For aircrafts and oil tanks,
which are mainly small and medium-sized targets, our approach has a clear improvement in detection
accuracy compared to YOLO-V3. The experimental results show that our improved YOLO-V3 can
effectively detect the remote sensing targets under the complex background in the condition of real-time
detection. In Table 8, we divide target categories by size. We can see that our approach has more
advantages than YOLO-V3 in detecting small-sized targets.
For the universality of our algorithm, we ran the experiment on the UCS-AOD dataset.
The comparison results are shown in Table 9. In addition, from Tables 8 and 9, we can see that
the leak detection rate is significantly lower than YOLO-V3 and other state-of-the-art algorithms.
Sensors 2020, 20, x FOR PEER REVIEW 16 of 24
Table 9. Experimental comparisons of accuracy and speed in the UCS-AOD dataset.
Table 9. Experimental comparisons of accuracy and speed in the UCS-AOD dataset.
Metric (%)
Method Backbone Metric (%) FPS
Method Backbone Leak Detection mAP mAPFPS
Aircraft CarCar Leak Detection
Aircraft
RateRate 0.5) = 0.5)
(%) (%) (IOU =(IOU
Faster RCNN
Faster RCNN VGG-16
VGG-16 87.31 86.48
87.31 86.48 13.8 13.8 86.90 86.90 6.1 6.1
SSD SSD VGG-16
VGG-16 70.24
70.24 72.6172.61 23.7 23.7 71.43 71.4361.5 61.5
DSSD DSSD ResNet-101
ResNet-101 73.17
73.17 74.1974.19 16.1 16.1 73.68 73.68 5.2 5.2
ESSD ESSD VGG-16
VGG-16 73.62
73.62 75.06
75.06 15.9 15.9 74.34 74.3433.2 33.2
YOLO-V2 DarkNet19 63.17 68.42 23.0 65.80 34.3
YOLO-V2 DarkNet19 63.17 68.42 23.0 65.80 34.3
YOLO-V3 YOLO-V3 DarkNet53
DarkNet53 75.71
75.71 75.62
75.62 18.5 18.5 75.67 75.6727.6 27.6
YOLO-V3 YOLO-V3
tiny tiny DarkNet19
DarkNet19 57.58
57.58 56.3556.35 35.2 35.2 56.97 56.9765.3 65.3
UAV-YOLO UAV-YOLO
[52] Figure 1 inFigure 1 in [52]
Reference 75.12 75.6075.60
75.12 16.5 16.5 75.36 75.3628.4 28.4
DC-SPP-YOLO [54][52] Figure 5 Reference
in Reference[52][54] 76.52 74.61 17.4 75.57 30.4
DC-SPP-YOLO
Ours Figure3)5 in
(Figure 89.31 88.24 9.3 88.78 24.9
76.52 74.61 17.4 75.57 30.4
[54] Reference [54]
Ours (Figure 3) 89.31 88.24 9.3 88.78 24.9
Under different backgrounds, partial detection results of our approach in RSOD and UCS-AOD
Under different
dataset are shown in Figure backgrounds,
8. In thepartial detection of
conditions results of our approach
different in RSODdifferent
illumination, and UCS-AOD distributions,
dataset are shown in Figure 8. In the conditions of different illumination, different distributions, and
and different target sizes, our approach can detect the target accurately, which proves excellent detection
different target sizes, our approach can detect the target accurately, which proves excellent detection
performance for multi-scale
performance remote
for multi-scale sensing
remote sensingtargets.
targets.
(a) Small aircraft targets under a (b) Small aircraft targets under a (c) Densely distributed small
strong light condition. weak light condition. aircraft targets.
Figure 8. Cont.
Sensors 2020, 20, 4276 16 of 23
Sensors 2020, 20, x FOR PEER REVIEW 17 of 24
(h) Densely distributed small oil (i) Oil tank targets under a bad
(g) Oil tank targets.
tank targets. weather condition.
(o) Densely distributed car targets. (p) Densely distributed car targets.
Figure 8. The detection results of the improved You Only Look Once (YOLO)-V3.
Figure 8. The detection results of the improved You Only Look Once (YOLO)-V3.
4.3.3. Ablation Experiments
4.3.3. AblationInExperiments
this section, we need to verify the effectiveness of each improved module we proposed. In
order to analyze the influence of module ‘DENSE 1st’ and module ‘DENSE 2nd’ (Figure 3) on the
In this section, we need to verify the effectiveness of each improved module we proposed. In order
detection accuracy, different combination modes were set up in the experiment under the condition
to analyze the influence of module ‘DENSE 1st’ and module ‘DENSE 2nd’ (Figure 3) on the detection
accuracy, different combination modes were set up in the experiment under the condition of three
detection scales. The experimental results of each combination in the RSOD dataset are shown
in Table 10.
Sensors 2020, 20, 4276 17 of 23
Table 10. Experimental comparisons of each combination in the feature extraction network.
Metric (%)
DENSE DENSE mAP
Aircraft Oil Tank Overpass Playground FPS
1st 2nd (IOU = 0.5)
1 74.30 73.85 75.08 85.16 77.10 29.7
2 X 76.81 75.38 77.21 85.37 78.69 30.9
3 X 77.28 76.39 79.65 85.92 79.81 31.4
4 X X 82.16 83.52 85.12 86.73 84.38 32.3
It can be seen from the first experiment and the fourth experiment that the feature extraction
network of the fourth experiment introduced dense connection modules based on Darknet53. mAP
of its model improved from 77.10% to 84.38%. On the other hand, the detection speed of the fourth
experiment increased from 29.7 FPS to 32.3 FPS compared to the first experiment. The experimental
results show that our proposed feature extraction network can improve the performance of remote
sensing target detection and also has advantages in detection speed.
In addition, Table 11 compared the experimental results of each module at the detection end
based on an improved feature extraction network. It can be found in the comparison of the first
experiment and the second experiment, and the comparison between the third experiment and the
fourth experiment, that the fourth detection scale increased and improved mAP up to 5.95% and
5.78%, respectively. Among them, for the small-sized targets like aircraft, the accuracy is improved by
8.72% and 7.04%, respectively. This shows that the increased detection scale can effectively improve
the detection accuracy of small targets. Compared with six convolutional layers, the ‘Res 3’ module
can avoid gradient fading and reduce the number of parameters. The comparison of experiment 1
and experiment 3, and the comparison between experiment 2 and experiment 4 show that the ‘Res 3’
module can slightly increase the detection speed.
Metric (%)
4th Scale Res 3 FPS
mAP
Aircraft Oil Tank Overpass Playground
(IOU = 0.5)
1 77.25 76.38 84.36 86.12 81.03 29.7
2 X 85.97 85.18 87.15 89.61 86.98 24.8
3 X 79.38 78.85 85.29 88.28 82.95 30.1
4 X X 86.42 87.57 89.37 91.56 88.73 25.8
The ablation experiments result in Tables 9 and 10, which proved that the improved feature
extraction network and detection end we proposed can improve the feature extraction ability of
the network and enhanced the detection accuracy of multi-scale remote sensing targets, especially
small-sized targets. In addition, the detection speed of our approach is not significantly reduced when
compared to YOLO-V3 and meets the real-time requirements.
Figure 9. Cont.
SensorsSensors
2020, 20, 4276
2020, 20, x FOR PEER REVIEW 19 of 23
20 of 24
Figure 9. The comparison results of YOLO-V3 and our approach: (a1)–(a8) The detection results of
Figure 9. The comparison
YOLO-V3; (b1)–(b8) Theresults of YOLO-V3
detection and our
results of Faster approach:
RCNN; (c1)–(c8)(a1–a8) The detection
The detection results ofresults
our of
YOLO-V3; (b1–b8) The detection results of Faster RCNN; (c1–c8) The detection results of our approach.
approach.
In Figure 9, a9,total
In Figure a totalofof24
24detection resultsofof
detection results eight
eight groups
groups werewere
chosenchosen in thedataset
in the RSOD RSODand dataset
and UCS-AOD dataset
UCS-AOD dataset to to prove
prove the the superiority
superiority of theof the improved
improved YOLO-V3. YOLO-V3.
The pictures The pictures
in the in the
first list
are are
first list the detection resultsresults
the detection of the YOLO-V3 [Link].
of the YOLO-V3 The picturesThein pictures
the second inlist
theare the detection
second list are the
results
detection of Faster
results of RCNN
Faster andRCNNthe pictures
and theinpictures
the thirdin listthe
arethird
the detection
list are results of our approach.
the detection results of It our
can be clearly seen that there are several small targets missed and detected by YOLO-V3.
approach. It can be clearly seen that there are several small targets missed and detected by YOLO-V3. Although
Faster RCNN performed better than YOLO-V3, leak detection still exists. On the other hand, all the
Although Faster RCNN performed better than YOLO-V3, leak detection still exists. On the other
targets were detected by our approach. The contrast experiments of eight groups and the ablation
hand, all the targets were detected by our approach. The contrast experiments of eight groups and
experiments showed that, by improving the feature extraction network and increasing the fourth
the ablation
detectionexperiments showedenhanced
scale, our approach that, by improving
the performance the feature extraction
of detecting network
small targets and
with increasing
complex
the fourth detection scale, our approach enhanced
background conditions in remote sensing images. the performance of detecting small targets with
complex background conditions in remote sensing images.
5. Conclusions
5. Conclusions
In practical engineering applications, we need to consider both accuracy and speed of detection.
The
In existing engineering
practical remote sensing target detection
applications, algorithms
we need often fail
to consider to accuracy
both consider both
andof [Link].
speed this
paper, we proposed an improved YOLO-V3-based model for multi-scale remote sensing
The existing remote sensing target detection algorithms often fail to consider both of them. In this paper, target
[Link]
we proposed Onimproved
account ofYOLO-V3-based
the complexity ofmodel
the background of remote
for multi-scale sensing
remote targets,
sensing this detection.
target puts
forward a higher requirement for the ability of the network to extract features. In this paper, we
On account of the complexity of the background of remote sensing targets, this puts forward a
focused on improving the original feature extraction network. Several improvements have been
higher requirement for the ability of the network to extract features. In this paper, we focused on
introduced to the original YOLO-V3 network. First, in order to extract feature information more
improving the original
effectively, feature extraction
a dense connection network. was
network (DenseNet) Several improvements
introduced have
in the feature been introduced
extraction network. to
the original
Second,YOLO-V3
to enhance network. First, in
the performance order to small-sized
of detecting extract feature information
targets, we extendedmore effectively,
the detection a dense
scales
connection network (DenseNet) was introduced in the feature extraction network. Second,
from 3 to 4. Third, we replaced three residual units with five convolutional layers, which is in each to enhance
the performance
detection layerof detecting small-sized
to avoid gradient targets,
fading. We canwe extended
see from thethe detection scales fromthat
ablation experiments 3 toeach
4. Third,
we replaced three residual units with five convolutional layers, which is in each detection layer to avoid
gradient fading. We can see from the ablation experiments that each improved module we proposed
is effective in improving the detection accuracy. Experiments on RSOD and UCS-AOD datasets
show that our approach achieves better performance on multi-scale remote sensing target detection.
Sensors 2020, 20, 4276 20 of 23
The improvement on the feature extraction network greatly improved the ability of extracting the
features of the targets. The additional fourth detection scale strengthens the performance of detecting
small targets. In the case of losing a portion of detection speed, the accuracy is greatly improved,
especially for small remote sensing targets compared with YOLO-V3. Although numerous improved
networks based on YOLO-V3 have been proposed, they usually detected targets in routine images.
When facing complex remote sensing images, they did not do well. On the contrast, with the above
measures adopted, our proposed algorithm is more suitable for remote sensing target detection than
other state-of-the-art target detection algorithms. In further work, multi-receptive fields for the feature
extraction of the network will be researched to boost the performance of remote sensing target detection.
In addition, the latest version of YOLO: YOLO-V4 [55] has been proposed and this will be researched
in further work.
Author Contributions: D.X. provided the original ideal, finished the experiment and this paper, and collected the
dataset. Y.W. contributed the modifications and suggestions to the paper. All authors have read and agreed to the
published version of the manuscript.
Funding: The National Nature Science Founding of China under Grant 61573183 and Open Project Program of
the National Laboratory of Pattern Recognition (NLPR) under Grant 201900029 funded this research.
Acknowledgments: The authors wish to thank the editor and reviewers for their suggestions and thank Yiquan
Wu for his guidance.
Conflicts of Interest: The authors declare no conflicts of interest.
Abbreviations:
The abbreviations in this paper are as follows:
YOLO You Only Look Once
CV Computer Version
SVM Support Vector Machine
HOG Histograms of Oriented Gradients
DPM Deformable Parts Model
IOU Intersection over Union
FC Full Connected Layer
FCN Full Convolutional Network
CNN Convolutional Neural Network
GT Ground Truth
RPN Region Proposal Network
FPN Feature Pyramid Network
ResNet Residual Network
DenseNet Densely Connected Network
NMS Non-Maximum Suppression
TP True Positive
FP False Positive
FN False Negative
AP Average Precision
mAP Mean Average Precision
FPS Frames Per Second
References
1. Shi, W.; Jiang, J.; Bao, S.; Tan, D. CISPNet: Automatic Detection of Remote Sensing Images from Google Earth
in Complex Scenes Based on Context Information Scene Perception. Appl. Sci. 2019, 9, 4836. [CrossRef]
2. Zhong, Y.; Weng, W.; Li, J.; Zhu, S. Collaborative Cross-Domain $k$ NN Search for Remote Sensing Image
Processing. IEEE Geosci. Remote Sens. Lett. 2019, 16, 1801–1805. [CrossRef]
3. Zhu, H.; Zhang, P.; Wang, L.; Zhang, X.; Jiao, L. A multiscale object detection approach for remote sensing
images based on MSE-DenseNet and the dynamic anchor assignment. Remote Sens. Lett. 2019, 10, 959–967.
[CrossRef]
Sensors 2020, 20, 4276 21 of 23
4. Zhang, Z.; Chen, J.; Liu, Z. SLIC segmentation method for full-polarised remote-sensing image. J. Eng. 2019,
2019, 6404–6407. [CrossRef]
5. Shi, Y.; Wang, W.; Gong, Q.; Li, D. Superpixel segmentation and machine learning classification algorithm for
cloud detection in remote-sensing images. J. Eng. 2019, 2019, 6675–6679. [CrossRef]
6. Li, Y.; Xu, J.; Xia, R.; Wang, X.; Xie, W. A two-stage framework of target detection in high-resolution
hyperspectral images. Signal Image Video Process. 2019, 13, 1339–1346. [CrossRef]
7. Li, S.; Xu, Y.; Zhu, M.; Ma, S.; Tang, H. Remote Sensing Airport Detection Based on End-to-End Deep
Transferable Convolutional Neural Networks. IEEE Geosci. Remote Sens. Lett. 2019, 16, 1640–1644. [CrossRef]
8. Kujawa, S.; Mazurkiewicz, J.; Czekala, W. Using convolutional neural networks to classify the maturity of
compost based on sewage sludge and rapeseed straw. J. Clean. Prod. 2020, 258, 120814. [CrossRef]
9. Xiao, B.; Xu, Y.; Bi, X.; Zhang, J.; Ma, X. Heart sounds classification using a novel 1-D convolutional neural
network with extremely low parameter consumption. Neurocomputing 2020, 392, 153–159. [CrossRef]
10. Hashimoto, R.; Requa, J.; Dao, T.; Ninh, A.; Tran, E.; Mai, D.; Lugo, M.; El-Hage Chehade, N.; Chang, K.J.;
Karnes, W.E.; et al. Artificial intelligence using convolutional neural networks for real-time detection of
early esophageal neoplasia in Barrett’s esophagus (with video). Gastrointest. Endosc. 2020, 91, 1264–1271.
[CrossRef]
11. Chen, R.-C. Automatic License Plate Recognition via sliding-window darknet-YOLO deep learning.
Image Vis. Comput. 2019, 87, 47–56. [CrossRef]
12. Bilal, M.; Hanif, M.S. Benchmark Revision for HOG-SVM Pedestrian Detector Through Reinvigorated
Training and Evaluation Methodologies. IEEE Trans. Intell. Transp. Syst. 2020, 21, 1277–1287. [CrossRef]
13. Wang, L.; Wu, J.; Wu, D. Research on vehicle parts defect detection based on deep learning. J. Phys. Conf. Ser.
2020, 1437, 012004. [CrossRef]
14. Zhang, D. Vehicle target detection methods based on color fusion deformable part model. EURASIP J. Wirel.
Commun. Netw. 2018, 2018, 1–6. [CrossRef]
15. Shen, J.; Pan, L.; Hu, X. Building Detection from High Resolution Remote Sensing Imagery Based on a
Deformable Part Model. Geomat. Inf. Sci. Wuhan Univ. 2017, 42, 1285–1291. (In Chinese) [CrossRef]
16. Chen, J.; Takiguchi, T.; Ariki, Y. Rotation-reversal invariant HOG cascade for facial expression recognition.
Signal Image Video Process. 2017, 11, 1485–1492. [CrossRef]
17. Jin, M.; Jeong, K.; Yoon, S.; Park, D.S. Real-time Pedestrian Detection based on GMM and HOG Cascade.
In Sixth International Conference on Machine Vision; Verikas, A., Vuksanovic, B., Zhou, J., Eds.; SPIE: Bellingham,
WA, USA, 2013; Volume 9067.
18. Xu, Z.; Huo, Y.; Liu, K.; Liu, S. Detection of ship targets in photoelectric images based on an improved
recurrent attention convolutional neural network. Int. J. Distrib. Sens. Netw. 2020, 16. [CrossRef]
19. Liu, Z.; Zhang, G.; Zhao, J.; Yu, L.; Sheng, J.; Zhang, N.; Yuan, H. Second-Generation Sequencing with Deep
Reinforcement Learning for Lung Infection Detection. J. Healthc. Eng. 2020, 2020. [CrossRef]
20. Xue, D.; Sun, J.; Hu, Y.; Zheng, Y.; Zhu, Y.; Zhang, Y. Dim small target detection based on convolutinal neural
network in star image. Multimed. Tools Appl. 2020, 79, 4681–4698. [CrossRef]
21. Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. IEEE. Rich Feature Hierarchies for Accurate Object Detection
and Semantic Segmentation. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern
Recognition, Columbus, OH, USA, 24–27 June 2014; pp. 580–587. [CrossRef]
22. Li, X.; Shang, M.; Qin, H.; Chen, L. Fast Accurate Fish Detection and Recognition of Underwater Images with Fast
R-CNN; IEEE: Piscataway, NJ, USA, 2015; pp. 921–925.
23. Girshick, R. IEEE. Fast R-CNN. In Proceedings of the 2015 IEEE International Conference on Computer
Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [CrossRef]
24. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal
Networks. In Advances in Neural Information Processing Systems 28; Cortes, C., Lawrence, N.D., Lee, D.D.,
Sugiyama, M., Garnett, R., Eds.; IEEE Computer Society: Los Alamitos, CA, USA, 2015; Volume 28.
25. Sun, N.; Zhu, Y.; Hu, X. Faster R-CNN Based Table Detection Combining Corner Locating; IEEE Computer Society:
Los Alamitos, CA, USA, 2019; pp. 1314–1319. [CrossRef]
26. Kaiming, H.; Gkioxari, G.; Dollar, P.; Girshick, R. Mask R-CNN. In Proceedings of the 2017 IEEE International
Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [CrossRef]
Sensors 2020, 20, 4276 22 of 23
27. Huang, Z.; Zhong, Z.; Sun, L.; Huo, Q. Mask R-CNN with Pyramid Attention Network for Scene Text
Detection. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV),
Waikoloa Village, HI, USA, 7–11 January 2019; pp. 1550–5790.
28. Shih, K.-H.; Chiu, C.-T.; Pu, Y.-Y. IEEE. Real-Time Object Detection via Pruning and a Concatenated
Multi-Feature Assisted Region Proposal Network. In Proceedings of the 2019 IEEE International Conference
on Acoustics, Speech and Signal Processing, Brighton, UK, 12–17 May 2019; pp. 1398–1402.
29. Shree, C.; Kaur, R.; Upadhyay, S.; Joshi, J. Multi-Feature Based Automated Flower Harvesting Techniques in
Deep Convolutional Neural Networking. In Proceedings of the 2019 4th International Conference on Internet
of Things: Smart Innovation and Usages (IoT-SIU), Ghaziabad, India, 18–19 April 2019; p. 6. [CrossRef]
30. Yuan, J.; Xue, B.; Zhang, W.; Xu, L.; Sun, H.; Zhou, J. RPN-FCN Based Rust Detection on Power Equipment.
In 2018 International Conference on Identification, Information and Knowledge in the Internet of Things; Bie, R.,
Sun, Y., Yu, J., Eds.; Elsevier Science Bv: Amsterdam, The Netherlands, 2019; Volume 147, pp. 349–353.
31. Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot MultiBox
Detector. In Computer Vision—ECCV 2016, Pt I; Leibe, B., Matas, J., Sebe, N., Welling, M., Eds.; Springer
International Publishing Ag: Cham, Switzerland, 2016; Volume 9905, pp. 21–37.
32. Lin, M.; Bing, L.; Zhiyu, Z.; Aravinda, C.V.; Kamitoku, N.; Yamazaki, K. Oracle Bone Inscription Detector Based
on SSD; Springer International Publishing: Cham, Switzerland, 2019; pp. 126–136. [CrossRef]
33. Tang, J.; Yao, X.; Kang, X.; Shun, N.; Ren, F. Position-Free Hand Gesture Recognition Using Single Shot
Multibox Detector Based Neural Network. In Proceedings of the 2019 IEEE International Conference on
Mechatronics and Automation (ICMA), Tianjin, China, 4–7 August 2019; pp. 2251–2256. [CrossRef]
34. Cui, L.; Ma, R.; Lv, P.; Jiang, X.; Gao, Z.; Zhou, B.; Xu, M. MDSSD: Multi-scale deconvolutional single shot
detector for small objects. Sci. China Inf. Sci. 2020, 63, 120113. [CrossRef]
35. Haque, M.F.; Dae-Seong, K. Multi Scale Object Detection Based on Single Shot Multibox Detector with
Feature Fusion and Inception Network. J. Korean Inst. Inf. Technol. 2018, 16, 93–100. [CrossRef]
36. Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. IEEE. You Only Look Once: Unified, Real-Time Object
Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition,
Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [CrossRef]
37. Zhang, X.; Qiu, Z.; Huang, P.; Hu, J.; Luo, J. IEEE. Application Research of YOLO v2 Combined with Color
Identification. In Proceedings of the 2018 International Conference on Cyber-Enabled Distributed Computing
and Knowledge Discovery, Zhengzhou, China, 18–20 October 2018; pp. 138–141. [CrossRef]
38. Redmon, J.; Farhadi, A. YOLOv3: An Incremental Improvement. 2018. Available online: [Link]
com/media/files/papers/[Link] (accessed on 30 July 2020).
39. Adarsh, P.; Rathi, P.; Kumar, M. YOLO v3-Tiny: Object Detection and Recognition Using one Stage Improved
Model. In Proceedings of the 2020 6th International Conference on Advanced Computing and Communication
Systems (ICACCS), Coimbatore, India, 6–7 March 2020; pp. 687–694. [CrossRef]
40. He, W.; Huang, Z.; Wei, Z.; Li, C.; Guo, B. TF-YOLO: An Improved Incremental Network for Real-Time
Object Detection. Appl. Sci. 2019, 9, 3225. [CrossRef]
41. Weber, J.; Lefevre, S. A multivariate Hit-or-Miss Transform for Conjoint Spatial and Spectral Template
Matching. In Image and Signal Processing; Elmoataz, A., Lezoray, O., Nouboud, F., Mammass, D., Eds.;
Springer-Verlag Berlin: Berlin, Germany, 2008; Volume 5099, pp. 226–235.
42. Feng, T.; Ma, H.; Cheng, X.; Zhang, H. Calculation of the optimal segmentation scale in object-based
multiresolution segmentation based on the scene complexity of high-resolution remote sensing images.
J. Appl. Remote Sens. 2018, 12, 025006. [CrossRef]
43. Sun, H.; Sun, X.; Wang, H.; Li, Y.; Li, X. Automatic Target Detection in High-Resolution Remote Sensing
Images Using Spatial Sparse Coding Bag-of-Words Model. IEEE Geosci. Remote Sens. Lett. 2012, 9, 109–113.
[CrossRef]
44. Zhang, P.; Niu, X.; Dou, Y.; Xia, F. Airport Detection on Optical Satellite Images Using Deep Convolutional
Neural Networks. IEEE Geosci. Remote Sens. Lett. 2017, 14, 1183–1187. [CrossRef]
45. Yu, Y.; Yang, X.; Xiao, S.; Lin, J. Automated Ship Detection from Optical Remote Sensing Images. In Advanced
Materials in Microwaves and Optics; Wang, D., Ed.; Trans Tech Publications Ltd.: Zurich, Switzerland, 2012;
Volume 500, pp. 785–791.
46. Guo, C.; Fan, B.; Zhang, Q.; Xiang, S.; Pan, C. AugFPN: Improving Multi-scale Feature Learning for Object
Detection. arXiv 2019, arXiv:1912.05384.
Sensors 2020, 20, 4276 23 of 23
47. Wong, F.; Hu, H. Adaptive learning feature pyramid for object detection. IET Comput. Vis. 2019, 13, 742–748.
[CrossRef]
48. Zeng, Y.; Ritz, C.; Zhao, J.; Lan, J. Attention-Based Residual Network with Scattering Transform Features for
Hyperspectral Unmixing with Limited Training Samples. Remote Sens. 2020, 12, 400. [CrossRef]
49. Li, J.; Gu, J.; Huang, Z.; Wen, J. Application Research of Improved YOLO V3 Algorithm in PCB Electronic
Component Detection. Appl. Sci. 2019, 9, 3750. [CrossRef]
50. Ju, M.; Luo, H.; Wang, Z.; Hui, B.; Chang, Z. The Application of Improved YOLO V3 in Multi-Scale Target
Detection. Appl. Sci. 2019, 9, 3775. [CrossRef]
51. Liu, G.; Nouaze, J.C.; Mbouembe, P.L.T.; Kim, J.H. YOLO-Tomato: A Robust Algorithm for Tomato Detection
Based on YOLOv3. Sensors 2020, 20, 2145. [CrossRef]
52. Liu, M.; Wang, X.; Zhou, A.; Fu, X.; Ma, Y.; Piao, C. UAV-YOLO: Small Object Detection on Unmanned Aerial
Vehicle Perspective. Sensors 2020, 20, 2238. [CrossRef]
53. Zhu, Y.; Newsam, S. IEEE. Densenet for Dense Flow. In Proceedings of the 2017 24th IEEE International
Conference on Image Processing, Beijing, China, 17–20 September 2017; pp. 790–794.
54. Huang, Z.; Wang, J. DC-SPP-YOLO: Dense connection and spatial pyramid pooling based YOLO for object
detection. Inf. Sci. 2020, 522, 241–258. [CrossRef]
55. Bochkovskiy, A.; Chien-Yao, W.; Liao, H.Y.M. YOLOv4: Optimal speed and accuracy of object detection.
arXiv 2020, arXiv:2004.10934.
© 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access
article distributed under the terms and conditions of the Creative Commons Attribution
(CC BY) license ([Link]