Robust Vehicle Detection with Faster R-CNN
Robust Vehicle Detection with Faster R-CNN
[Link]
RESEARCH
Received: 12 April 2023 / Accepted: 7 July 2023 / Published online: 30 July 2023
© The Author(s) 2023
Abstract
Deep convolutional neural networks (CNNs) have shown tremendous success in the detection of objects and vehicles in
recent years. However, when using CNNs to identify real-time vehicle detection in a moving context remains difficult. Many
obscured and truncated cars, as well as huge vehicle scale fluctuations in traffic photos, provide these issues. To improve the
performance of detection findings, we used multiscale feature maps from CNN or input pictures with numerous resolutions
to adapt the base network to match different scales. This research presents an enhanced framework depending on Faster
R-CNN for rapid vehicle recognition which presents better accuracy and fast processing time. Research results on our custom
dataset indicate that our recommended methodology performed better in terms of detection efficiency and processing time,
especially in comparison to the earlier age of Faster R-CNN models.
4
* Rania Salih Civil Engineering Department, College of Engineering, King
rania.salih2@[Link] Khalid University, Abha 61421, Saudi Arabia
5
* Mohammad Arsalan Khan Primary Affiliation: Geomechanics and Geotechnics Group,
[Link]@[Link] Kiel University, 24118 Kiel, Germany
6
* Mohammad Mursaleen Department of Civil Engineering, Z. H. College
mursaleenm@[Link] of Engineering and Technology, Aligarh Muslim University,
Uttar Pradesh, Aligarh, India
1
School of Automotive Studies, Tongji University, Shanghai, 7
School of Electronic Information and Electrical Engineering,
China
Shanghai Jiao Tong University, 800 Dong Chuan Road,
2
Department of Geotechnical Engineering, College of Civil Shanghai 200240, China
Engineering, Tongji University, 1239 Siping Road, 8
China Medical University Hospital, China Medical
Shanghai 200092, China
University (Taiwan), Taichung 40402, Taiwan
3
Department of Civil Engineering, Red Sea University,
Port‑Sudan, Sudan
13
Vol.:(0123456789)
93 Page 2 of 10 Journal of Real-Time Image Processing (2023) 20:93
models are thoroughly investigated in this study in order performed well, with a mean average accuracy of 83.92
to solve a range of fascinating challenges in intelligent percent on the KITTI [31] automobile detection standard.
transportation systems. Researchers have suggested vari- Another recent work the use of a faster R-CNN with domain
ous classic vehicle detection algorithms from the earliest adaption in road vehicles detection research shows about
stages of the field to current days [9–15]. The performance accuracy 83.1% and 0.56 s testing time on COCO dataset
of techniques is determined by handcrafted characteris- [32].
tics. The most often utilized features are the Haar-like [16] Faster R-CNN's competitive performance on the KITTI
and Histogram of Oriented Gradient (HOG) [17]. The cas- vehicle identification benchmark may be explained by one
caded detector [18], exhibiting a commendable level of primary factor that’s the wide range of vehicle scales. The
precision, stands out as one of the pioneering real-time RPN is fed convolutional feature maps and produces possi-
detection systems. Two well-known methods of the part- ble ROI. RPN ignores tiny objects and vehicle overlooking
based design method are Support Vector Machines (SVM) due to the wide range of vehicle scales. However, we think
[19] and deformable part-based models (DPM) [20]. The there is a scope to further improve the faster R-CNN per-
researchers focus on three major practical issues in vehicle formance, therefore we propose a model to solve the issue
detection such as huge variations in light, heavy occlusion, of wide-scale variance in vehicle detection. Not only for
and big variations in sizes. To solve the issue of high vari- adequate transportation management or administration but
ance in light, Saini [21] presented a strong CNN model for also for efficient damages detection in insurance solutions,
traffic management light recognition for autonomous cars. accurate classification of automobiles into distinct kinds is
As input, the framework uses raw picture data, identifies critical. Therefore, a work towards automatic damage assess-
candidate regions, and later detects and recognizes traffic ment procedure on vehicles is essential to prevent work acci-
lights. Heavy occlusion makes distinguishing occluded dents that may be caused by individuals while assessing the
vehicles hard to detect. Phan [22] presented a strategy damage.
for dealing with thick occlusion caused by surveillance
cameras that are fixed. The approach includes background
removal, occlusion detection, and automobile detection, 2 Custom dataset creation
which extracts occluded cars separately based on exterior
attributes. For big variations of size, Lu [23] has suggested 2.1 Preparing dataset
a scale-aware Region Proposal Network (RPN) to handle
the challenge of identifying vehicles of various sizes. The This study has created a unique and customized dataset as
scale-aware RPN is composed consisting of two particular well as KITTI Vision Benchmark Suite for training and
sub-networks: one that detects big proposals and the other evaluation purposes. Every subcategory has at least 66,000
detects little proposals, which are then fed into two differ- illustrations in our collection. Figure 1 show some samples
ent XGBoost [24] classifiers to create the final prediction. of our custom dataset. Our dataset was gathered from the
The first appropriate technique, known as region-based different roadside and upper sides of a different roads. We
convolutional neural network (R-CNN) [25], performed well installed the camera at a specific place and recorded videos
in vehicle identification. The region-based convolution neu- at 60 frames per second at a different place for three days.
ral network has a region proposal network with the CNN to The images were then retrieved from videos and duplicate
outperform HOG [17] features with an SVM classifier. Raw pictures were eliminated. We categorized data following
picture data is fed into a region-based convolutional neu- gathering it according to its classifications. The five cat-
ral network, which generates region recommendations. The egories in our database are car, bus, truck, motorcycle, and
region suggestions are then put into the CNN to extract the cycle. We offered images of every category from several
features map, and the support vector machine [26] is used perspectives, including panoramic, front base, and lateral
to forecast. In the Pascal VOC 2010 competition, the basic views. There are 4000 samples capture in this study and
RCNN obtained a mean average accuracy of 53% and spatial more than 2000 were taken from the KITTI dataset. alto-
Pyramid Pooling (SPP) [27] employs a convolution layer gether we got 6000 images and for our training purpose, we
on the whole picture and extracts the features map using used 4000 images and for validation purposes 2000 images
SPP-net, avoiding the high cost of computation of R-CNN. were used. The dataset is graded on three difficulty levels:
In 2015, He et al. suggested a faster R-CNN [28] for easy, moderate, and hard. The easy type objects are made up
object recognition. Faster R-CNN was the primary to use of anchor boxes with the least height of 40 pixels and higher,
the Region Proposal Network (RPN) as a candidate genera- the moderate type objects are made up of anchor boxes with
tor for Regions of Interest (Roi). For the COCO [29] and the height of 25 pixels to 40 pixels, and the hard type images
Pascal VOC [30] two-dimensional object detection stand- are made up of anchor boxes with a Smaller than 25 pixels
ards, the faster R-CNN performs well. One recent work has considered as hard.
13
Journal of Real-Time Image Processing (2023) 20:93 Page 3 of 10 93
2.2 Annotating dataset wide variety of labeling techniques. To date, the tool has
been used to annotate over 6000 photos that are only in the
The key difficulties in machine learning involve object training dataset which has over 66,000 automobile illustra-
detection and categorization. The detection and classifica- tions. The labeling method that we used is the bounding box
tion methods help to identify numerous items on the street, method. The below Figure shows some samples from our
including automobiles, humans, and fixed things like traffic annotation image and annotation XML file information. Our
lights, street signs, and lamp posts. Real-time training data- annotation XML files contain the information as different
sets are required for the creation of identification and classi- vehicle types like the car, bus/truck, cycle/motorbike, and
fication methods. However, we created our own dataset with the difficulty level are easy, moderate, and hard.
in different streets with a different expression. The images in
these datasets are typically manually labeled by us. we build
bounding boundaries around the recognized items and save 3 Methodology
the features of such objects through an annotation stage like
most of the researchers use some open-source software to 3.1 The implemented architecture
do the annotation process manually[33].
Utilizing the tools, we generate shapes for image segmen- At the beginning of our method, we take our custom
tation, build anchor boxes in object detection and recogni- annotated dataset as import Then the imported images go
tion, and add captions to the selected regions. through the base network. In the base network we used dif-
The annotation information is saved in a variety of forms, ferent kinds of architecture, those are our modified Vgg16,
including text, JSON, YOLO, XML, ILSVRC, and others. ResNet50, ResNet101, and [Link] output from
In our case, we saved the annotation information as XML. the base network (base network feature map) fed to RPN,
The hand annotation process is very costly and also time- soft NMS, as a result, we got our proposal layer (with anchor
consuming. For example, the YOLO object detection data- box). Roi pooling layer takes the input from the base net-
base requires approximately 35 s to build an anchor box over work and proposal to perform max polling with transposed
an object [28]. Experts use two alternative ways to make the convolution and gives the output as the refined proposal.
operation of bounding box annotations affordable and effi- Finally, the output of the refined proposal goes through the
cient. There are three types of annotation processes: manual, classification and regression layer to show the final detection
semi-automated, and completely automatic that we can use. results with the regression box and classification box. The
A manual annotation tool had been employed for manually suggested approach's general architecture in our research is
labeling the images. This tool allows us to label using a illustrated in Fig. 2.
13
93 Page 4 of 10 Journal of Real-Time Image Processing (2023) 20:93
13
Journal of Real-Time Image Processing (2023) 20:93 Page 5 of 10 93
13
93 Page 6 of 10 Journal of Real-Time Image Processing (2023) 20:93
13
Journal of Real-Time Image Processing (2023) 20:93 Page 7 of 10 93
Train
0.6
Valid indicating that it performs well on different hardness level
0 5 10 15 20 25 30 35
by pixel of the bounding box.
Epoch Our model obtains 91.78% AP on a different degree of
difficulty with a duration of 0.11 s per picture by utilizing a
GPU with 11 GB of RAM.
Fig. 7 Training and validation accuracy
Figure 9 can show some prediction accuracy in dataset
detection results. Each layer is divided into proposed areas
As we can see from Table 1, our modified Vgg16 and by our modified VGG16, which also predicts the locations
MobileNetV3 provide a better MAP accuracy and our of many anchor boxes of various scales and sizes for each
model can detect different categories of vehicle (car as object. The predicted box placement is adjusted using a
medium size vehicle, bus/truck big size vehicle, and cycle global optimization, and we can see that our model predicted
Motorbike as small size vehicle). Therefore, we decided accurately although some images have poor lighting condi-
to choose modified Vgg16 as a base network for our final tion and vehicle is located in the shadow area of the road.
comparison with those recent publication techniques using We also discovered that the traditional Faster R-CNN fails
the KITTI dataset. We can see the difference of improve- to detect little objects (less than 64 pixels). As a result, we
ment in Table 2. proposed Modified VGG16 with soft NMS and a refined
As we can see from Table 2 our modified Vgg16 and proposal to accommodate tiny objects.
MobileNetV3 give a better map and our model can detect We tried four different kinds of base networks as feature
different categories of vehicle (car as medium size vehicle, extractors (Modified Vgg16, MobilenetV3, ResNet50, and
bus/truck big size vehicle, and cycle/ Motorbike as small ResNet101). Using our modified model, we were able to
size vehicle) So we decided to choose modified Vgg16 as recognize the automobile category in our custom detec-
a base network for our final comparison with those recent tion dataset. In traditional faster R-CNN they used the
publication techniques using the KITTI dataset. Later, classical VGG16 as a feature extractor but in our model
Fig. 8 has shown the results of Precision-Recall for three we used modified VGG16 which give better accuracy
different categories (car, bus/truck, and cycle/ Motorbike) and faster testing time. the soft-NMS method replaces
average precision (AP) measures given by our model with the NMS (non-maximum suppression) method after the
modified VGG16 in terms of easy, medium, and hard level. RPN (region proposal network) in the traditional Faster
On our custom dataset, the suggested model with R-CNN to tackle the issue of duplicated proposals and it
modified VGG16 has 93.67% for cars, 89.15% for truck/ also slightly improve the AP performance. The proposals
buses, and 92.52% for motorbikes/cycle Global Accuracy, are then adjusted to the appropriate size using a refined
Table 1 Our proposed model Method Easy Object mAP Moderate Object hard Object mAP Process-
result improvement after mAP ing time (s/
implementing different steps Image)
13
93 Page 8 of 10 Journal of Real-Time Image Processing (2023) 20:93
Fig. 8 Precision-recall categories with modified VGG16 model on easy, medium and hard levels
proposal layer without compromising vital contextual Our proposed model gives better mAP and processing
information which gives better performance to detect tiny time performance than the older version of Faster R-CNN.
sized vehicle than the traditional faster R-CNN.
We evaluated our model with the custom dataset for
state-of-the-art detection on the KITTI testing dataset. We 5 Conclusion and future work
chose to select a slightly unique training dataset that we
were using for assessment on our custom dataset since The purpose of this research is also to use deep learning
the KITTI testing dataset had comparable scenarios to the to get a better understanding of real-time road vehicles,
training set. Since we often encounter situations in the including preparing our own dataset with image annotation
testing set where cars appear stranded on the street, we and vehicle recognition. Tuning the number and density of
decided to include heavily occluded vehicles in our analy- the network's convolutional layers demonstrates the neu-
sis. As a result, a dataset containing all occluded labels ral network and data flexibility. We chose our modified
were utilized to train the network that was used to sub- VGG16 as the core base network model for the feature
mit to the scoreboard. Aside from that, all of the training extractor after assessing all of the evaluation indicators in
settings were identical. Table 3 presents the performance general. In future work, the major component that needs to
of our recommended approach on the custom dataset, as be focused on is a range of photograph collections, such as
well as the results of other approaches on the KITTI test lighting settings and background surroundings. CNN mod-
dataset. It provides a comprehensive overview of our posi- els can readily notice patterns and output with a greater
tion in the KITTI benchmark, showcasing the effectiveness accuracy rate when given various input components. In
and competitiveness of our proposed method compared to addition, the volume of the dataset can play a role in learn-
existing approaches. ing algorithms.
13
Journal of Real-Time Image Processing (2023) 20:93 Page 9 of 10 93
Table 3 Performance Method Easy Object Moderate hard Object mAP Process-
comparison on KITTI mAP Object mAP ing time (s/
benchmark Image)
Author contributions Alam and Ahmed wrote the main manuscript, Declarations
Alam and Asmari visualised, Salih did data curation, Khan did the
re-writing, Mustafa, Mursaleen and Islam did the formal analyses. All Conflict of interest The authors declare that there is no conflict of in-
authors reviewed the manuscript. terest regarding the publication of this paper.
Funding Open Access funding enabled and organized by Projekt Data Availability Statement The data is available as per request to cor-
DEAL. responding author.
13
93 Page 10 of 10 Journal of Real-Time Image Processing (2023) 20:93
Open Access This article is licensed under a Creative Commons Attri- 15. Zaman, K., et al., A novel driver emotion recognition system based
bution 4.0 International License, which permits use, sharing, adapta- on deep ensemble classification. Complex & Intelligent Systems,
tion, distribution and reproduction in any medium or format, as long 2023: p. 1–26.
as you give appropriate credit to the original author(s) and the source, 16. Wen, X., et al.: Efficient feature selection and classification for
provide a link to the Creative Commons licence, and indicate if changes vehicle detection. IEEE Trans. Circuits Syst. Video Technol.
were made. The images or other third party material in this article are 25(3), 508–517 (2014)
included in the article's Creative Commons licence, unless indicated 17. Tomasi, C., Histograms of oriented gradients. Computer Vision
otherwise in a credit line to the material. If material is not included in Sampler, 2012: p. 1–6.
the article's Creative Commons licence and your intended use is not 18. Saipullah, K., et al., COMPARISON OF FEATURE EXTRAC-
permitted by statutory regulation or exceeds the permitted use, you will TORS FOR REAL-TIME OBJECT DETECTION ON ANDROID
need to obtain permission directly from the copyright holder. To view a SMARTPHONE. Journal of Theoretical & Applied Information
copy of this licence, visit [Link] Technology, 2013. 47(1).
19. Suykens, J., Vandewalle, J.: Neural Process. Lett 9, 293 (1999)
20. Hsiao, E., et al., A discriminatively trained, multiscale, deformable
part model. 2009.
References 21. Saini, S., et al. An efficient vision-based traffic light detection and
state recognition for autonomous vehicles. in 2017 IEEE Intel-
1. Bas, E., A.M. Tekalp, and F.S. Salman. Automatic vehicle count- ligent Vehicles Symposium (IV). 2017. IEEE.
ing from video for traffic flow analysis. in 2007 IEEE intelligent 22. Phan, H.N., et al. Occlusion vehicle detection algorithm in
vehicles symposium. 2007. Ieee. crowded scene for traffic surveillance system. in 2017 Interna-
2. Chen, R.-C.: Automatic License Plate Recognition via sliding- tional Conference on System Science and Engineering (ICSSE).
window darknet-YOLO deep learning. Image Vis. Comput. 87, 2017. IEEE.
47–56 (2019) 23. Ding, L., et al. Scale-aware RPN for vehicle detection. in Advances
3. Hussain, T., et al.: Real time violence detection in surveillance in Visual Computing: 13th International Symposium, ISVC 2018,
videos using Convolutional Neural Networks. Multimedia Tools Las Vegas, NV, USA, November 19–21, 2018, Proceedings 13.
and Applications 81(26), 38151–38173 (2022) 2018. Springer.
4. Zaman, K., et al.: Driver Emotions Recognition Based on 24. Ramraj, S., et al.: Experimenting XGBoost algorithm for predic-
Improved Faster R-CNN and Neural Architectural Search Net- tion and classification of different datasets. International Journal
work. Symmetry 14(4), 687 (2022) of Control Theory and Applications 9(40), 651–662 (2016)
5. Shah, S.M., et al.: A driver gaze estimation method based on deep 25. Girshick, R., et al. Rich feature hierarchies for accurate object
learning. Sensors 22(10), 3959 (2022) detection and semantic segmentation. in Proceedings of the IEEE
6. Ullah, R., et al.: Auction Mechanism-Based Sectored Fractional conference on computer vision and pattern recognition. 2014.
Frequency Reuse for Irregular Geometry Multicellular Networks. 26. Suykens, J.A., Vandewalle, J.: Least squares support vector
Electronics 11(15), 2281 (2022) machine classifiers. Neural Process. Lett. 9, 293–300 (1999)
7. Zaman, K., et al.: EEDLABA: Energy-Efficient Distance-and 27. He, K., et al.: Spatial pyramid pooling in deep convolutional net-
Link-Aware Body Area Routing Protocol Based on Clustering works for visual recognition. IEEE Trans. Pattern Anal. Mach.
Mechanism for Wireless Body Sensor Network. Appl. Sci. 13(4), Intell. 37(9), 1904–1916 (2015)
2190 (2023) 28. Ren, S., et al., Faster r-cnn: Towards real-time object detection
8. Hussain, T., et al.: Improving Source location privacy in social with region proposal networks. Advances in neural information
Internet of Things using a hybrid phantom routing technique. processing systems, 2015. 28.
Comput. Secur. 123, 102917 (2022) 29. Lin, T.-Y., et al. Microsoft coco: Common objects in context.
9. Ojha, A., S.P. Sahu, and D.K. Dewangan. VDNet: vehicle detection in Computer Vision–ECCV 2014: 13th European Conference,
network using computer vision and deep learning mechanism for Zurich, Switzerland, September 6–12, 2014, Proceedings, Part V
intelligent vehicle system. in Proceedings of Emerging Trends and 13. 2014. Springer.
Technologies on Intelligent Systems: ETTIS 2021. 2022. Springer. 30. Everingham, M., et al.: The pascal visual object classes (voc)
10. Dewangan, D.K. and S.P. Sahu. Predictive control strategy for challenge. Int. J. Comput. Vision 88, 303–338 (2010)
driving of intelligent vehicle system against the parking slots. in 31. Nguyen, H.: Improving faster R-CNN framework for fast vehicle
2021 5th international conference on intelligent computing and detection. Math. Probl. Eng. 2019, 1–11 (2019)
control systems (ICICCS). 2021. IEEE. 32. Yin, G., et al.: Research on highway vehicle detection based on
11. Dewangan, D.K. and S.P. Sahu. Real time object tracking for intel- faster R-CNN and domain adaptation. Appl. Intell. 52(4), 3483–
ligent vehicle. in 2020 first international conference on power, 3498 (2022)
control and computing technologies (ICPC2T). 2020. IEEE. 33. Torralba, A., Russell, B.C., Yuen, J.: Labelme: Online image
12. Ottakath, N., Al-Maadeed, S.: Vehicle instance segmentation annotation and applications. Proc. IEEE 98(8), 1467–1484 (2010)
polygonal dataset for a private surveillance system. Sensors 23(7),
3642 (2023) Publisher's Note Springer Nature remains neutral with regard to
13. Dewangan, D.K. and S.P. Sahu, Lane detection for intelligent jurisdictional claims in published maps and institutional affiliations.
vehicle system using image processing techniques. Data Science:
Theory, Algorithms, and Applications, 2021: p. 329–348.
14. Farid, A., et al.: A Fast and Accurate Real-Time Vehicle Detection
Method Using Deep Learning for Unconstrained Environments.
Appl. Sci. 13(5), 3059 (2023)
13