0% found this document useful (0 votes)
18 views10 pages

Robust Vehicle Detection with Faster R-CNN

This research presents a robust vehicle detection algorithm based on Faster R-CNN, addressing challenges such as occlusion and scale variations in real-time vehicle identification. The proposed framework utilizes multiscale feature maps and a modified VGG16 architecture to enhance detection accuracy and processing speed, outperforming earlier Faster R-CNN models. Results from a custom dataset indicate significant improvements in detection efficiency, achieving a mean average accuracy of 83.92%.

Uploaded by

Ganga
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views10 pages

Robust Vehicle Detection with Faster R-CNN

This research presents a robust vehicle detection algorithm based on Faster R-CNN, addressing challenges such as occlusion and scale variations in real-time vehicle identification. The proposed framework utilizes multiscale feature maps and a modified VGG16 architecture to enhance detection accuracy and processing speed, outperforming earlier Faster R-CNN models. Results from a custom dataset indicate significant improvements in detection efficiency, achieving a mean average accuracy of 83.92%.

Uploaded by

Ganga
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Journal of Real-Time Image Processing (2023) 20:93

[Link]

RESEARCH

Faster RCNN based robust vehicle detection algorithm for identifying


and classifying vehicles
Md Khorshed Alam1 · Asif Ahmed2 · Rania Salih3 · Abdullah Faiz Saeed Al Asmari4 · Mohammad Arsalan Khan5,6 ·
Noman Mustafa7 · Mohammad Mursaleen8 · Saiful Islam4

Received: 12 April 2023 / Accepted: 7 July 2023 / Published online: 30 July 2023
© The Author(s) 2023

Abstract
Deep convolutional neural networks (CNNs) have shown tremendous success in the detection of objects and vehicles in
recent years. However, when using CNNs to identify real-time vehicle detection in a moving context remains difficult. Many
obscured and truncated cars, as well as huge vehicle scale fluctuations in traffic photos, provide these issues. To improve the
performance of detection findings, we used multiscale feature maps from CNN or input pictures with numerous resolutions
to adapt the base network to match different scales. This research presents an enhanced framework depending on Faster
R-CNN for rapid vehicle recognition which presents better accuracy and fast processing time. Research results on our custom
dataset indicate that our recommended methodology performed better in terms of detection efficiency and processing time,
especially in comparison to the earlier age of Faster R-CNN models.

Keywords Classification · Deep learning · Modified vgg16 · Vehicle detection

1 Introduction [1], license plate identification [2], incident detection [3],


driver facial emotion identification [4–7],and internet of
Technologies of vehicle detection have existed under things (IOT) source location and identification [8]. On
development in industry and academics in recent years. the other hand, vision-based approaches can fully use the
Many state-of-art image identification algorithms have abundance of visual patterns to distinguish target objects
not been able to compete in the field of vehicle detection in a human-like manner. For example, radar sensor-based
standards. The primary obstacles in automobile identifi- methods can only identify cars in a relatively small area,
cation include big differences in object sizes, substantial but vision-based systems may use a camera to discover all
occlusion, and considerable fluctuations in light. Sensor- the vehicles in a vast viewable region and describe addi-
based algorithms may be used to solve some surveillance tional aspects of each detected vehicle at the same time. As
tasks in urban traffic systems, such as vehicle counting a result, numerous computer vision and machine learning

4
* Rania Salih Civil Engineering Department, College of Engineering, King
rania.salih2@[Link] Khalid University, Abha 61421, Saudi Arabia
5
* Mohammad Arsalan Khan Primary Affiliation: Geomechanics and Geotechnics Group,
[Link]@[Link] Kiel University, 24118 Kiel, Germany
6
* Mohammad Mursaleen Department of Civil Engineering, Z. H. College
mursaleenm@[Link] of Engineering and Technology, Aligarh Muslim University,
Uttar Pradesh, Aligarh, India
1
School of Automotive Studies, Tongji University, Shanghai, 7
School of Electronic Information and Electrical Engineering,
China
Shanghai Jiao Tong University, 800 Dong Chuan Road,
2
Department of Geotechnical Engineering, College of Civil Shanghai 200240, China
Engineering, Tongji University, 1239 Siping Road, 8
China Medical University Hospital, China Medical
Shanghai 200092, China
University (Taiwan), Taichung 40402, Taiwan
3
Department of Civil Engineering, Red Sea University,
Port‑Sudan, Sudan

13
Vol.:(0123456789)
93 Page 2 of 10 Journal of Real-Time Image Processing (2023) 20:93

models are thoroughly investigated in this study in order performed well, with a mean average accuracy of 83.92
to solve a range of fascinating challenges in intelligent percent on the KITTI [31] automobile detection standard.
transportation systems. Researchers have suggested vari- Another recent work the use of a faster R-CNN with domain
ous classic vehicle detection algorithms from the earliest adaption in road vehicles detection research shows about
stages of the field to current days [9–15]. The performance accuracy 83.1% and 0.56 s testing time on COCO dataset
of techniques is determined by handcrafted characteris- [32].
tics. The most often utilized features are the Haar-like [16] Faster R-CNN's competitive performance on the KITTI
and Histogram of Oriented Gradient (HOG) [17]. The cas- vehicle identification benchmark may be explained by one
caded detector [18], exhibiting a commendable level of primary factor that’s the wide range of vehicle scales. The
precision, stands out as one of the pioneering real-time RPN is fed convolutional feature maps and produces possi-
detection systems. Two well-known methods of the part- ble ROI. RPN ignores tiny objects and vehicle overlooking
based design method are Support Vector Machines (SVM) due to the wide range of vehicle scales. However, we think
[19] and deformable part-based models (DPM) [20]. The there is a scope to further improve the faster R-CNN per-
researchers focus on three major practical issues in vehicle formance, therefore we propose a model to solve the issue
detection such as huge variations in light, heavy occlusion, of wide-scale variance in vehicle detection. Not only for
and big variations in sizes. To solve the issue of high vari- adequate transportation management or administration but
ance in light, Saini [21] presented a strong CNN model for also for efficient damages detection in insurance solutions,
traffic management light recognition for autonomous cars. accurate classification of automobiles into distinct kinds is
As input, the framework uses raw picture data, identifies critical. Therefore, a work towards automatic damage assess-
candidate regions, and later detects and recognizes traffic ment procedure on vehicles is essential to prevent work acci-
lights. Heavy occlusion makes distinguishing occluded dents that may be caused by individuals while assessing the
vehicles hard to detect. Phan [22] presented a strategy damage.
for dealing with thick occlusion caused by surveillance
cameras that are fixed. The approach includes background
removal, occlusion detection, and automobile detection, 2 Custom dataset creation
which extracts occluded cars separately based on exterior
attributes. For big variations of size, Lu [23] has suggested 2.1 Preparing dataset
a scale-aware Region Proposal Network (RPN) to handle
the challenge of identifying vehicles of various sizes. The This study has created a unique and customized dataset as
scale-aware RPN is composed consisting of two particular well as KITTI Vision Benchmark Suite for training and
sub-networks: one that detects big proposals and the other evaluation purposes. Every subcategory has at least 66,000
detects little proposals, which are then fed into two differ- illustrations in our collection. Figure 1 show some samples
ent XGBoost [24] classifiers to create the final prediction. of our custom dataset. Our dataset was gathered from the
The first appropriate technique, known as region-based different roadside and upper sides of a different roads. We
convolutional neural network (R-CNN) [25], performed well installed the camera at a specific place and recorded videos
in vehicle identification. The region-based convolution neu- at 60 frames per second at a different place for three days.
ral network has a region proposal network with the CNN to The images were then retrieved from videos and duplicate
outperform HOG [17] features with an SVM classifier. Raw pictures were eliminated. We categorized data following
picture data is fed into a region-based convolutional neu- gathering it according to its classifications. The five cat-
ral network, which generates region recommendations. The egories in our database are car, bus, truck, motorcycle, and
region suggestions are then put into the CNN to extract the cycle. We offered images of every category from several
features map, and the support vector machine [26] is used perspectives, including panoramic, front base, and lateral
to forecast. In the Pascal VOC 2010 competition, the basic views. There are 4000 samples capture in this study and
RCNN obtained a mean average accuracy of 53% and spatial more than 2000 were taken from the KITTI dataset. alto-
Pyramid Pooling (SPP) [27] employs a convolution layer gether we got 6000 images and for our training purpose, we
on the whole picture and extracts the features map using used 4000 images and for validation purposes 2000 images
SPP-net, avoiding the high cost of computation of R-CNN. were used. The dataset is graded on three difficulty levels:
In 2015, He et al. suggested a faster R-CNN [28] for easy, moderate, and hard. The easy type objects are made up
object recognition. Faster R-CNN was the primary to use of anchor boxes with the least height of 40 pixels and higher,
the Region Proposal Network (RPN) as a candidate genera- the moderate type objects are made up of anchor boxes with
tor for Regions of Interest (Roi). For the COCO [29] and the height of 25 pixels to 40 pixels, and the hard type images
Pascal VOC [30] two-dimensional object detection stand- are made up of anchor boxes with a Smaller than 25 pixels
ards, the faster R-CNN performs well. One recent work has considered as hard.

13
Journal of Real-Time Image Processing (2023) 20:93 Page 3 of 10 93

Fig. 1  Few examples of our custom dataset

2.2 Annotating dataset wide variety of labeling techniques. To date, the tool has
been used to annotate over 6000 photos that are only in the
The key difficulties in machine learning involve object training dataset which has over 66,000 automobile illustra-
detection and categorization. The detection and classifica- tions. The labeling method that we used is the bounding box
tion methods help to identify numerous items on the street, method. The below Figure shows some samples from our
including automobiles, humans, and fixed things like traffic annotation image and annotation XML file information. Our
lights, street signs, and lamp posts. Real-time training data- annotation XML files contain the information as different
sets are required for the creation of identification and classi- vehicle types like the car, bus/truck, cycle/motorbike, and
fication methods. However, we created our own dataset with the difficulty level are easy, moderate, and hard.
in different streets with a different expression. The images in
these datasets are typically manually labeled by us. we build
bounding boundaries around the recognized items and save 3 Methodology
the features of such objects through an annotation stage like
most of the researchers use some open-source software to 3.1 The implemented architecture
do the annotation process manually[33].
Utilizing the tools, we generate shapes for image segmen- At the beginning of our method, we take our custom
tation, build anchor boxes in object detection and recogni- annotated dataset as import Then the imported images go
tion, and add captions to the selected regions. through the base network. In the base network we used dif-
The annotation information is saved in a variety of forms, ferent kinds of architecture, those are our modified Vgg16,
including text, JSON, YOLO, XML, ILSVRC, and others. ResNet50, ResNet101, and [Link] output from
In our case, we saved the annotation information as XML. the base network (base network feature map) fed to RPN,
The hand annotation process is very costly and also time- soft NMS, as a result, we got our proposal layer (with anchor
consuming. For example, the YOLO object detection data- box). Roi pooling layer takes the input from the base net-
base requires approximately 35 s to build an anchor box over work and proposal to perform max polling with transposed
an object [28]. Experts use two alternative ways to make the convolution and gives the output as the refined proposal.
operation of bounding box annotations affordable and effi- Finally, the output of the refined proposal goes through the
cient. There are three types of annotation processes: manual, classification and regression layer to show the final detection
semi-automated, and completely automatic that we can use. results with the regression box and classification box. The
A manual annotation tool had been employed for manually suggested approach's general architecture in our research is
labeling the images. This tool allows us to label using a illustrated in Fig. 2.

13
93 Page 4 of 10 Journal of Real-Time Image Processing (2023) 20:93

Fig. 2  Proposed conventional faster-RCNN model

3.2 Modified VGG16 the network achieves loss convergence in lesser time. At


the place of last max pooling, we added a Global Aver-
As the backbone framework, we have used the modified age Pooling (GAP) layer before fully connecting FC-12
VGG 16. It has the large-sized Kernel Filtering, with 11 and the SoftMax layer. Traditional VGG16 doesn’t have
in the first Convolutional Layer, each with 3 × 3 numerous a dropout and batch normalization layer in it. With our
kernel-sized filters. It is preferable to transmit 224 × 224 × 3 modified VGG-16, the weight that has been pre-trained
to Conv1 layers, after which the images go through numer- will be benefited. After adding dropout, batch normaliza-
ous convolutional layers with very narrow filtering of size tion layers and using global average pooling, the modified
3 × 3. A 1 × 1 convolutional filter is also employed in some VGG16 networks exhibit excellent classification perfor-
circumstances. After every block data goes to the dropout mance and shorter testing time. On the other hand, the
layer and then batch normalization. Very small size kernels other three base network (MobileNetV3, ResNet-101and.
are layered together around the perception to retain the spa- ResNet-50) that we have also tested, those were unchanged
tially numerously small shaped kernels layered together on and fine-tuned for our model which didn’t perform well
the receptive field since it assists to understand complicated like modified Vgg16.
features at low cost as numerously nonlinear-layers boost the From Fig. 3, we can see that the RPN creates a col-
depth of the network. The padding is set appropriately (for lection of anchor boxes from the base network's convolu-
the 3 × 3 Convolutional layer, padding is set to 1 × 1 pixels), tion feature map. Those anchors frequently overlap, and
however, the strides are set to 1. The spatial Pooling is per- proposals frequently overlap over the identical object. To
formed by the five Max Pooling stages; this comes after a solve the problem of overwriting proposals, the soft non-
few convolutional layers but is not complete of them. Max maximum suppression (SNMS) algorithm is used. The
Pooling makes use of the 2 × 2 Pixels Kernel or the window NMS algorithm is often used to erase duplicate proposals
with the 2 strides. At the last instant of max pooling, we for many state-of-the-art object recognition techniques,
added a Global Average Pooling (GAP) layer before fully including Faster R-CNN. Classical NMS eliminates any
connecting FC-12 and the SoftMax layer. other proposal that overlaps a winning proposal by over
We mainly have done modifications on VGG-16 and a predetermined threshold. The classical NMS algorithm
named it as modified VGG-16. After every max-pooling could remove beneficial proposals surprisingly caused by
stage of Traditional VGG16, we add a dropout and a batch heavy automotive occlusion in traffic (Fig. 4).
normalization layer. The goal of Batch Normalization is to This study used a soft-NMS algorithm to solve the
achieve a stable distribution of activation values through- NMS problem with overlapped vehicles. The neighboring
out training, and in our experiments, we apply it before proposals of successful proposals really aren't completely
the activation layer. BN layer performs scaling operation effectively suppressed with soft-NMS. Rather, those are
on the outputs of the layer before it. This process brings suppressed relying on the neighboring proposals' updated
stability to the weights updating during the training of the objectiveness scores, which are calculated depending on
model. This has the effect of stabilizing and speeding-up the level of overlap between the neighboring proposals or
the training process of deep neural networks. In result, the winning proposal.
the network weights optimization becomes simplified and

13
Journal of Real-Time Image Processing (2023) 20:93 Page 5 of 10 93

Fig. 3  Our the proposed VGG16 illustrate

Fig. 4  Illustration result of SNMS

3.3 Refined proposal an error buildup in backward propagation throughout


the training phase. The detection of tiny cars will be
The Roi pooling layer is often used in various two-stage decreased. To reduce the size of suggestions to a fixed
object recognition methods, including Faster R-CNN, and size while preserving the original features of tiny cars and
Fast R-CNN [23], to regulate reducing the size of proposals improving the efficiency of the suggested technique for
to a certain size. The Roi pooling layer employs the max identifying small vehicles by the refined proposal.
pooling. that turn the features within each acceptable region Figure 5 depicts the refined proposal idea. When the
of interest from the proposal layer into a compact feature size of the proposal is bigger than the specified feature size
map with such a specific geographic area H × W. The h × w map output in the refined proposal process, max-pooling is
Roi proposal is divided into an H × W matrix of sub-win- employed to decrease the proposal's size to a specific size.
dows of similar sizes of (h/H) × (w/W), and the elements When the proposal size is lower than the output fea-
in every sub-window are max-pooled into the appropriate ture map's predetermined size, Transposed Convolution
output bounding box. is used to extend the proposal size to the fixed size. The
If a proposal size is smaller compared to H × W, that proportion between the refined proposal size and the input
will be extended to accommodate the extra space by adding proposal size determines the kernel size. Moreover, while
repeated values. Because Roi pooling avoids processing the the wideness of a proposal is greater than the stable out-
convolutional layers again, it may drastically reduce both put feature map's size and the proposal height is lower
training and testing time. Adding repeated values to tiny than the specified height of the outputs feature, transposed
proposals, on the other hand, is not acceptable, particularly convolution is used to increase the proposal height while
with little vehicles, since it may ruin the actual shapes of the max-pooling used to decrease the width of that proposal.
small automobiles. The proposal size has been regulated to a fixed size with
Furthermore, adding duplicated values for minor pro- improved proposal.
posals would result in erroneous forward propagating

13
93 Page 6 of 10 Journal of Real-Time Image Processing (2023) 20:93

Fig. 5  Refined proposal

4 Result and analysis 1.2


Training
Validation
In this research, we primarily apply the modified Vgg- 1.0

16, MobileNetV3, ResNet-101, and ResNet-50 model to


our custom-made dataset, which is then fine-tuned on the 0.8

KITTI dataset for the base network. The training environ-


ment utilized for our experiments involved the utilization 0.6
Loss

of the Nvidia RTX3080 GPU. This high-performance GPU


from Nvidia played a crucial role in accelerating the train- 0.4

ing process and enabling efficient model optimization. Its


advanced capabilities provided the necessary computational 0.2
power to handle the complex training tasks and achieve opti-
mal results. Each batch normalization layer's weights and 0.0
dropout in the pre-trained model were increased to speed
-5 0 5 10 15 20 25 30 35
up training and reduce overfitting. Turn after turn, the clas-
Epoch
sifier and the RPN are trained. A mini-batch is used to train
the RPN initially, with the base network and RPN variables
Fig. 6  Training and validation losses
changed just once. The RPN's negative and positive propos-
als are then used to update and train the classification. The
classifier parameters are adjusted once, then the characteris- In Fig. 7 The training accuracy and validation accuracy
tics of the basic convolutional layers are adjusted once more. both stabilize at a specific point which means we got a well-
RPN and the modified Vgg-16, MobileNetV3, ResNet-101, trained network with modified VGG16 as a base network.
and ResNet-50-based classifier both use the same underly- Table 1 shows some of the outcomes of our modified
ing convolutional layers. The loss function for bounding box model. We also discovered that the traditional Faster R-CNN
regression and coordinate parameterization are similar to the fails to detect little objects (less than 64 pixels). As a result,
traditional Faster R-CNN work. In the loss function, the bal- we proposed Modified VGG16 with soft NMS and a refined
ance parameter is set to 1. The loss functions are optimized proposal to accommodate tiny objects. Table 1 reports the
using the SGD with momentum. With the learning rate per analysis of the traditional Faster R-CNN and our improved
mini-batch set at 0.0001, the RPN and the classifier's starting model comparison.
learning rates are predetermined to 0.0001 and we used 200 Table 1 displays the results that our proposed model gives
epochs for training. better MAP and processing time performance than the older
We can see in Fig. 6 that at some time, the training and version of Faster R-CNN. Research results indicate that our
validation losses both diminish and stabilize. This demon- recommended methodology performed better in terms of
strates that our model's ideal fit does not underfeed or over- detection efficiency and processing time, especially in com-
feed the data. parison to the traditional Faster R-CNN models.

13
Journal of Real-Time Image Processing (2023) 20:93 Page 7 of 10 93

Table 2  Results of the different base networks on our proposed model


1.0
Base Network Learning Rate mAP Car Truck/ Bus Motor-
bike/
Cyclist
0.9
Vgg16 0.0001 88.35 90.25 87.43 87.37
Accuracy

Modified 0.0001 91.78 93.67 89.15 92.52


0.8 Vgg16
MobileNetV3 0.0001 89.52 92.41 90.66 85.49
ResNet50 0.0001 72.61 75.42 72.25 70.14
0.7
ResNet101 0.0001 75.99 78.14 75.39 74.45

Train
0.6
Valid indicating that it performs well on different hardness level
0 5 10 15 20 25 30 35
by pixel of the bounding box.
Epoch Our model obtains 91.78% AP on a different degree of
difficulty with a duration of 0.11 s per picture by utilizing a
GPU with 11 GB of RAM.
Fig. 7  Training and validation accuracy
Figure 9 can show some prediction accuracy in dataset
detection results. Each layer is divided into proposed areas
As we can see from Table 1, our modified Vgg16 and by our modified VGG16, which also predicts the locations
MobileNetV3 provide a better MAP accuracy and our of many anchor boxes of various scales and sizes for each
model can detect different categories of vehicle (car as object. The predicted box placement is adjusted using a
medium size vehicle, bus/truck big size vehicle, and cycle global optimization, and we can see that our model predicted
Motorbike as small size vehicle). Therefore, we decided accurately although some images have poor lighting condi-
to choose modified Vgg16 as a base network for our final tion and vehicle is located in the shadow area of the road.
comparison with those recent publication techniques using We also discovered that the traditional Faster R-CNN fails
the KITTI dataset. We can see the difference of improve- to detect little objects (less than 64 pixels). As a result, we
ment in Table 2. proposed Modified VGG16 with soft NMS and a refined
As we can see from Table 2 our modified Vgg16 and proposal to accommodate tiny objects.
MobileNetV3 give a better map and our model can detect We tried four different kinds of base networks as feature
different categories of vehicle (car as medium size vehicle, extractors (Modified Vgg16, MobilenetV3, ResNet50, and
bus/truck big size vehicle, and cycle/ Motorbike as small ResNet101). Using our modified model, we were able to
size vehicle) So we decided to choose modified Vgg16 as recognize the automobile category in our custom detec-
a base network for our final comparison with those recent tion dataset. In traditional faster R-CNN they used the
publication techniques using the KITTI dataset. Later, classical VGG16 as a feature extractor but in our model
Fig. 8 has shown the results of Precision-Recall for three we used modified VGG16 which give better accuracy
different categories (car, bus/truck, and cycle/ Motorbike) and faster testing time. the soft-NMS method replaces
average precision (AP) measures given by our model with the NMS (non-maximum suppression) method after the
modified VGG16 in terms of easy, medium, and hard level. RPN (region proposal network) in the traditional Faster
On our custom dataset, the suggested model with R-CNN to tackle the issue of duplicated proposals and it
modified VGG16 has 93.67% for cars, 89.15% for truck/ also slightly improve the AP performance. The proposals
buses, and 92.52% for motorbikes/cycle Global Accuracy, are then adjusted to the appropriate size using a refined

Table 1  Our proposed model Method Easy Object mAP Moderate Object hard Object mAP Process-
result improvement after mAP ing time (s/
implementing different steps Image)

Faster R-CNN [1] 86.71 81.84 71.12 2


Soft NMS 88.43 86.78 76.31 2
Refined proposal 89.59 91.39 81.24 2
Modified Vgg16 87.27 81.84 73.21 0.11
MobilenetV3 85.42 78.04 69.64 0.13

13
93 Page 8 of 10 Journal of Real-Time Image Processing (2023) 20:93

Fig. 8  Precision-recall categories with modified VGG16 model on easy, medium and hard levels

proposal layer without compromising vital contextual Our proposed model gives better mAP and processing
information which gives better performance to detect tiny time performance than the older version of Faster R-CNN.
sized vehicle than the traditional faster R-CNN.
We evaluated our model with the custom dataset for
state-of-the-art detection on the KITTI testing dataset. We 5 Conclusion and future work
chose to select a slightly unique training dataset that we
were using for assessment on our custom dataset since The purpose of this research is also to use deep learning
the KITTI testing dataset had comparable scenarios to the to get a better understanding of real-time road vehicles,
training set. Since we often encounter situations in the including preparing our own dataset with image annotation
testing set where cars appear stranded on the street, we and vehicle recognition. Tuning the number and density of
decided to include heavily occluded vehicles in our analy- the network's convolutional layers demonstrates the neu-
sis. As a result, a dataset containing all occluded labels ral network and data flexibility. We chose our modified
were utilized to train the network that was used to sub- VGG16 as the core base network model for the feature
mit to the scoreboard. Aside from that, all of the training extractor after assessing all of the evaluation indicators in
settings were identical. Table 3 presents the performance general. In future work, the major component that needs to
of our recommended approach on the custom dataset, as be focused on is a range of photograph collections, such as
well as the results of other approaches on the KITTI test lighting settings and background surroundings. CNN mod-
dataset. It provides a comprehensive overview of our posi- els can readily notice patterns and output with a greater
tion in the KITTI benchmark, showcasing the effectiveness accuracy rate when given various input components. In
and competitiveness of our proposed method compared to addition, the volume of the dataset can play a role in learn-
existing approaches. ing algorithms.

13
Journal of Real-Time Image Processing (2023) 20:93 Page 9 of 10 93

Fig. 9  Visual representation of detection results on test datasets samples

Table 3  Performance Method Easy Object Moderate hard Object mAP Process-
comparison on KITTI mAP Object mAP ing time (s/
benchmark Image)

Faster R-CNN [10] 86.71 81.84 71.12 2


Faster R-CNN [52] 89.20 87.86 74.72 0.15
Complexer-YOLO[53] 79.43 71.97 67.62 0.06
IA-SSD (single)[54] 83.98 76.37 71.73 0.013
Cascade MS-CNN[55] 94.26 91.60 78.84 0.25
Proposed Method (with original VGG16) 88.35 86.87 77.68 0.14
Proposed Method (with modified Vgg16) 91.78 89.54 79.54 0.11

Author contributions Alam and Ahmed wrote the main manuscript, Declarations
Alam and Asmari visualised, Salih did data curation, Khan did the
re-writing, Mustafa, Mursaleen and Islam did the formal analyses. All Conflict of interest The authors declare that there is no conflict of in-
authors reviewed the manuscript. terest regarding the publication of this paper.

Funding Open Access funding enabled and organized by Projekt Data Availability Statement The data is available as per request to cor-
DEAL. responding author.

13
93 Page 10 of 10 Journal of Real-Time Image Processing (2023) 20:93

Open Access This article is licensed under a Creative Commons Attri- 15. Zaman, K., et al., A novel driver emotion recognition system based
bution 4.0 International License, which permits use, sharing, adapta- on deep ensemble classification. Complex & Intelligent Systems,
tion, distribution and reproduction in any medium or format, as long 2023: p. 1–26.
as you give appropriate credit to the original author(s) and the source, 16. Wen, X., et al.: Efficient feature selection and classification for
provide a link to the Creative Commons licence, and indicate if changes vehicle detection. IEEE Trans. Circuits Syst. Video Technol.
were made. The images or other third party material in this article are 25(3), 508–517 (2014)
included in the article's Creative Commons licence, unless indicated 17. Tomasi, C., Histograms of oriented gradients. Computer Vision
otherwise in a credit line to the material. If material is not included in Sampler, 2012: p. 1–6.
the article's Creative Commons licence and your intended use is not 18. Saipullah, K., et al., COMPARISON OF FEATURE EXTRAC-
permitted by statutory regulation or exceeds the permitted use, you will TORS FOR REAL-TIME OBJECT DETECTION ON ANDROID
need to obtain permission directly from the copyright holder. To view a SMARTPHONE. Journal of Theoretical & Applied Information
copy of this licence, visit [Link] Technology, 2013. 47(1).
19. Suykens, J., Vandewalle, J.: Neural Process. Lett 9, 293 (1999)
20. Hsiao, E., et al., A discriminatively trained, multiscale, deformable
part model. 2009.
References 21. Saini, S., et al. An efficient vision-based traffic light detection and
state recognition for autonomous vehicles. in 2017 IEEE Intel-
1. Bas, E., A.M. Tekalp, and F.S. Salman. Automatic vehicle count- ligent Vehicles Symposium (IV). 2017. IEEE.
ing from video for traffic flow analysis. in 2007 IEEE intelligent 22. Phan, H.N., et al. Occlusion vehicle detection algorithm in
vehicles symposium. 2007. Ieee. crowded scene for traffic surveillance system. in 2017 Interna-
2. Chen, R.-C.: Automatic License Plate Recognition via sliding- tional Conference on System Science and Engineering (ICSSE).
window darknet-YOLO deep learning. Image Vis. Comput. 87, 2017. IEEE.
47–56 (2019) 23. Ding, L., et al. Scale-aware RPN for vehicle detection. in Advances
3. Hussain, T., et al.: Real time violence detection in surveillance in Visual Computing: 13th International Symposium, ISVC 2018,
videos using Convolutional Neural Networks. Multimedia Tools Las Vegas, NV, USA, November 19–21, 2018, Proceedings 13.
and Applications 81(26), 38151–38173 (2022) 2018. Springer.
4. Zaman, K., et al.: Driver Emotions Recognition Based on 24. Ramraj, S., et al.: Experimenting XGBoost algorithm for predic-
Improved Faster R-CNN and Neural Architectural Search Net- tion and classification of different datasets. International Journal
work. Symmetry 14(4), 687 (2022) of Control Theory and Applications 9(40), 651–662 (2016)
5. Shah, S.M., et al.: A driver gaze estimation method based on deep 25. Girshick, R., et al. Rich feature hierarchies for accurate object
learning. Sensors 22(10), 3959 (2022) detection and semantic segmentation. in Proceedings of the IEEE
6. Ullah, R., et al.: Auction Mechanism-Based Sectored Fractional conference on computer vision and pattern recognition. 2014.
Frequency Reuse for Irregular Geometry Multicellular Networks. 26. Suykens, J.A., Vandewalle, J.: Least squares support vector
Electronics 11(15), 2281 (2022) machine classifiers. Neural Process. Lett. 9, 293–300 (1999)
7. Zaman, K., et al.: EEDLABA: Energy-Efficient Distance-and 27. He, K., et al.: Spatial pyramid pooling in deep convolutional net-
Link-Aware Body Area Routing Protocol Based on Clustering works for visual recognition. IEEE Trans. Pattern Anal. Mach.
Mechanism for Wireless Body Sensor Network. Appl. Sci. 13(4), Intell. 37(9), 1904–1916 (2015)
2190 (2023) 28. Ren, S., et al., Faster r-cnn: Towards real-time object detection
8. Hussain, T., et al.: Improving Source location privacy in social with region proposal networks. Advances in neural information
Internet of Things using a hybrid phantom routing technique. processing systems, 2015. 28.
Comput. Secur. 123, 102917 (2022) 29. Lin, T.-Y., et al. Microsoft coco: Common objects in context.
9. Ojha, A., S.P. Sahu, and D.K. Dewangan. VDNet: vehicle detection in Computer Vision–ECCV 2014: 13th European Conference,
network using computer vision and deep learning mechanism for Zurich, Switzerland, September 6–12, 2014, Proceedings, Part V
intelligent vehicle system. in Proceedings of Emerging Trends and 13. 2014. Springer.
Technologies on Intelligent Systems: ETTIS 2021. 2022. Springer. 30. Everingham, M., et al.: The pascal visual object classes (voc)
10. Dewangan, D.K. and S.P. Sahu. Predictive control strategy for challenge. Int. J. Comput. Vision 88, 303–338 (2010)
driving of intelligent vehicle system against the parking slots. in 31. Nguyen, H.: Improving faster R-CNN framework for fast vehicle
2021 5th international conference on intelligent computing and detection. Math. Probl. Eng. 2019, 1–11 (2019)
control systems (ICICCS). 2021. IEEE. 32. Yin, G., et al.: Research on highway vehicle detection based on
11. Dewangan, D.K. and S.P. Sahu. Real time object tracking for intel- faster R-CNN and domain adaptation. Appl. Intell. 52(4), 3483–
ligent vehicle. in 2020 first international conference on power, 3498 (2022)
control and computing technologies (ICPC2T). 2020. IEEE. 33. Torralba, A., Russell, B.C., Yuen, J.: Labelme: Online image
12. Ottakath, N., Al-Maadeed, S.: Vehicle instance segmentation annotation and applications. Proc. IEEE 98(8), 1467–1484 (2010)
polygonal dataset for a private surveillance system. Sensors 23(7),
3642 (2023) Publisher's Note Springer Nature remains neutral with regard to
13. Dewangan, D.K. and S.P. Sahu, Lane detection for intelligent jurisdictional claims in published maps and institutional affiliations.
vehicle system using image processing techniques. Data Science:
Theory, Algorithms, and Applications, 2021: p. 329–348.
14. Farid, A., et al.: A Fast and Accurate Real-Time Vehicle Detection
Method Using Deep Learning for Unconstrained Environments.
Appl. Sci. 13(5), 3059 (2023)

13

You might also like