Computer Vision: A Comprehensive Guide
Computer Vision: A Comprehensive Guide
started with
Computer
Vision
A guide to the knowledge and application
of visual systems
[Link]
The goal of computer vision is to extract
meaning from pixels and perform visual
tasks similar to the human visual system. Interest in
how machines ‘see’ and how computer vision
can be used to build products for consumers and
businesses is growing rapidly.
Rea
Identification
ognition
l-t
Tracking
Capabilities of
ime analys
computer vision
Rec
is
If you are reading the printed version of this brochure, you can download a hyperlinked pdf at [Link]/brochures
1
Contents
1 An introduction to computer vision 3
a. FAQs 3
b. A brief history of computer vision 4
c. The evolution of computer vision 5
d. Deep learning breakthrough 5
2 Application examples 7
a. Smart homes 7
b. Smart cities 7
c. Industry 7
d. Healthcare 8
e. Agriculture 8
f. Security 8
g. Autonomous vehicles 8
h. AR/VR & immersive technologies 8
3 How is computer vision used in business? 9
a. Benefits for business, industry and society 9
b. Technical challenges 9
c. Privacy 9
4 How to set up a computer vision system 10
a. Basic components 10
b. Hardware platforms 10
c. Software tools 10
d. Digital imaging system stack 10
5 How to process and interpret images 11
a. Image as an array 11
b. Image processing 12
c. Machine learning 12
d. Deep learning 13
e. Choosing machine learning or deep learning 13
f. Image processing libraries 13
g. Machine learning frameworks 14
6 Embedded vision 15
a. Embedded vision platforms 15
b. Camera modules 16
c. Interfaces 16
2
1 An introduction to
computer vision
a. FAQs Relationship between AI,
machine learning and deep learning
Artificial intelligence:
What is computer vision? AI is the theory and
al intellige development of computer
Of the five human senses, vision is the one that provides most tifici nc systems to perform tasks
of the data we receive and is considered our dominant sense. Ar normally requiring human
e
intelligence.
It provides us with a detailed description of the surrounding
Machine learning: is an
world which is constantly changing. Although vision involves hine learnin application of AI based
a huge amount of information and complex processing, the
ac around the idea of giving
M
machines access to data
g
human visual system can interpret this information easily. and letting them learn for
themselves.
The ability to see, process and then act on visual input is
Deep Deep learning: is a special
something that most humans take for granted.
learning type of machine learning
algorithm, multiple layers of
Computer vision engineering is the practice of using neural networks that mimic
technology and machines to replicate, and even improve upon, the connectivity of the
human brain in processing
human vision. The technology captures and stores images data and creating patterns
before transforming them into information that can be further for use in decision making.
acted upon. As a minimum an AI system must be able to reproduce aspects
of human intelligence
This requires expertise across a range of fields, including sensor
technology, image and signal processing, computer graphics,
Image processing takes an image as an input and provides a
computer architecture, algorithms and machine learning.
processed image as an output. The purpose of the processing
is usually to improve the quality of the image. Typical methods
What are the fundamental computer vision used are filtering, noise removal, sharpening and edge detection.
techniques?
Computer vision broadens the purpose of image processing
Image classification gives a computer the ability to interpret to include quantitative and qualitative information from visual
the input from an image sensor and categorise what it ‘sees’. data. Similar to the process of human visual reasoning, computer
vision can distinguish between objects, classify them and sort
Object Detection detecting instances of a certain class (such
them according to their attributes. Computer vision, like image
as vehicles, humans, buildings) in images or videos.
processing, takes an image as an input. However, it returns an
Object Tracking detecting and recognizing a defined item output with additional information interpreted from the image
in each frame of a video to distinguish it from other objects in such as size, colour, number, location or orientation.
the scene.
This can be extended beyond the extraction of meaningful
3D Image Reconstruction the process of capturing the shape information from a single image to multiple images or video,
and appearance of real objects. for example, to count the number of cars passing by a point on
the street as they are recorded by a video camera. Temporal
Semantic Image Segmentation when specific regions of an information therefore plays a role in computer vision, much as
image are labelled according to what the object is. it does with our own understanding of the world.
3
What is artificial intelligence (AI)? The computer vision
Artificial intelligence is intelligence demonstrated by machines,
where any device can perceive its environment and mimic
market is expected
human functions such as ‘learning’ and ‘problem solving’.
Artificial intelligence, or AI, is the broad concept of machines to reach close to $22 billion
being able to carry out tasks in a way that is considered ‘smart’.
by 2026
What are neural networks? [Link]
Neural networks are a means of machine learning, where computer-vision-market-size-and-forecast-to-2025/
a computer learns to perform a task by analysing training
examples or datasets. Usually, the dataset examples have been
manually labelled in advance. An object recognition system
might be fed thousands of labelled images of cars, houses, b. A brief history of
cups and would find visual patterns in the images that correlate
consistently with the particular label.
computer vision
What is deep learning? Computer vision has a long history in commercial and
Deep learning is the use of neural network methods to perform government use where light wave sensors in various spectrum
image analysis, moving away from statistical methods to neural ranges have been deployed in many applications such as:
network algorithms which are developed to mimic the neurons
of the human brain. • Remote sensing for environmental observation
and management
What applications can computer vision • High resolution cameras that collect intelligence over
be used for? battlefields
Applications of computer vision are many and varied. • Thermal imagers to detect people during police operations
Common applications you may be familiar with include
augmented reality, facial recognition, gesture and handwriting • X-ray sensors for airport security.
recognition, machine vision, remote sensing, robotics,
autonomous vehicles, people counting and iris recognition. The sensors can be stationary or attached to moving objects,
such as satellites, drones and vehicles. When combined with
What business sectors use computer vision? connectivity technologies such as Wi-Fi, Bluetooth or 3G/4G/5G,
they create a new set of applications that were not possible before.
Computer vision has numerous applications such as
remote sensing, healthcare (particularly around medical Computer vision, coupled with connectivity, advanced data
imaging such as MRI scans or ultrasound imaging), security, analytics and artificial intelligence, are catalysts for each other,
manufacturing, automotive, transport, robotics, sports, giving rise to revolutionary leaps in IoT innovations
gaming and many others. and applications.
4
c. The evolution of computer vision
1960s
Computer vision technology started in the early 1960s with
the aim of trying to mimic human vision systems and to ask
computers to tell us what they see.
5
1950:
computer vision
emerges 1957:
pixel invented, first digital
image
1990s:
computer graphics
& computer vision
1966: (image morphing,
MIT Artificial
view interpolation,
intelligence lab
panoramic image
stitching)
1969:
CCD invented
2001:
real-time face
detection
1970s:
first commercial
computer vision 1990s:
application (OCR) projective 3D reconstructions,
stereo imaging, statistical
learning techniques for facial
1975: recognition
first commercial digital
camera 2010s:
GPUs/neural networks
1980s:
mathematical and
quantitative analysis
developments
2012
AlexNet, deep neural
network for image
recognition
6
2 Application examples
a.
Person detection
Computer can be used to adjust
vision-based user data Facial recognition Indoor security lighting and temperature
will increasingly become will be used to cameras will to the number of people
a feature of the home. send an alert to in any room to ensure a
Smart When systems can
unlock the door, or
a smartphone if comfortable
to remain locked if
homes detect and recognise an unfamiliar person an elderly family environment and save
objects, they can deliver member falls, or if a on electricity and a TV
approaches box that recognises
smart actions according toddler is climbing
to what they were up stairs individuals can turn on a
programmed tailored interface for
entertainment
to do
d. Computer vision
applications in
Healthcare Robots will need
healthcare have been These can be used
robotics can to be able to
developed to aid to detect if elderly
Healthcare healthcare professionals people have fallen help with assisting navigate the world
nurses to clean around them
with medical imaging or require other
hospitals through 3D computer
diagnosis, surgery and forms of assistance
vision
health monitoring
7
The agriculture The quality of food
e.
This can help with Properties such as
industry is better productivity, products can be colour, shape, size,
increasingly using crop monitoring, assessed and graded surface defects and Agriculture
computer vision precision agriculture into specific grades, contamination can
technology for and locating weeds while detecting also be estimated
applications and pests defects
Examples include
Cameras can be
The challenge f.
Intelligent scene Automatic Number is to identify the
placed in offices,
monitoring systems Plate Recognition scene and context,
hospitals, banks,
are playing an (ANPR), people and
ports, car parks,
understanding what Security
increasingly vehicle tracking, demands immediate
stadiums, shopping
significant role in crowd analysis and attention, what is
centres, airports
society zone detection for valuable and what
and more
health & safety can be ignored
Computer vision
Self-driving vehicles
must be able to
capture visual data in
High quality
images and videos
must be obtained
g.
Self-driving vehicles technology is
can be made real time to create 3D in low light
being applied maps to understand Autonomous
intelligent, self-reliant conditions as well
to autonomous the surroundings,
and reliable using vehicles to make it
as daylight, using vehicles
computer vision while detecting and LiDAR sensors and
safe for passengers classifying objects
technology thermal cameras
and pedestrians in their path such alongside visible
as traffic lights and camera sensors
pedestrians
8
3 How is computer vision
used in business?
Facial recognition Financial institutions
b. Technical challenges
There is a high level of technical understanding team of professionals with technical expertise.
required to create software that collects and interprets Companies may also need to have a dedicated team
visual data. To train a computer vision system for regular monitoring and evaluation of the vision
powered by machine learning, companies need to have a system performance.
c. Privacy
Privacy is the biggest social threat that computer vision poses. their behaviour or monitoring their habits. Everyone’s
The capabilities of computer vision – identification, recognition, information is stored on a cloud.
tracking and real-time analysis – impact directly with individual It is important to understand the potential negative effects of
rights for privacy. With computers learning from thousands and computer vision applications on society. This is crucial to ensure
thousands of images and videos, computers are getting better that computer vision applications make our lives more comfortable
at recognising individuals by their facial features, by identifying and efficient and not for purposes of constrain and control.
9
4 How to set up a computer
vision system
Almost everyone has experienced computer vision and machine learning, often without even knowing.
This section explains how to set up a computer vision system.
a. Basic components
The components of a standard computer vision system are:
b. Hardware platforms
CPU and can run in to the thousands. The greater number of
The central processing unit of a computer used to cores allows multiple calculations to be worked on at the
perform arithmetic computations. Most modern CPUs same time which allows image processing to be
have 2 to 256 cores. performed efficiently.
GPU FPGA
The graphics processing unit of a computer used to Field programmable gate arrays have parallel processing
process graphics. GPUs start at a couple of hundred cores capabilities which make them suitable for image processing.
c. Software tools
There are many software tools with the necessary techniques to perform image and video processing tasks as well as machine
learning algorithms.
CPU GPU
• OpenCV • Scilab • Octave • R • Matlab • Tensorflow • PyTorch • Keras • Caffe
processing
5 Image post-processor Image data optimisation
Numeric
4 Image storage Formatting and storing image data presentation
10
5 How to process and
interpret images
a. Image as an array
A digital image is an array of pixels where each pixel is a A grayscale image refers to the number of different shades, or
combination of numerical values representing the colours depth, of a particular colour. A grayscale image can be created
and intensities at a particular point on the image. from any single channel or colour of the image.
A pixel, or picture element, is the smallest visual element of an The image resolution gives the number of pixels and the aspect
image and typically contains three component intensities, or ratio gives the width:height pixel ratio.
channels, such as red, green and blue. Colour digital images The channel contains the number of samples per point, which for
are created by combining the channels to reproduce the broad grayscale images is a single sample per pixel, whereas for colour
range of colours seen by the human eye. images is three samples per pixel (red, green, blue).
Pixel
Smallest visual element
10 10 16 28
Digital Image
65 70 56 43
9 90 96 67
A multidimensional array 99 70 56 78
32 90 96 67
of numbers 15 85 43 92
60 90 96 67
Aspect Ratio 23 85 43 92
32 65 87 99
Width:Height 85 85 43 92
Channel: No. of samples per point 54 65 87 99
Resolution
32 65 87 99
Width x Height Single plane: Grayscale / B&W images
©maxEmbedded.com2012
11
b. Image processing
The main purpose of image processing is to improve the etc.) and treat those features as a ‘definition’ of the object.
quality of the image by sharpening and restoration; extract These ‘definitions’ are then searched for in other images. If
the features of an image to help discriminate objects and/or a significant number of features from one type of object are
classes of objects; classify objects, locate their position and found in another image, the image can then be classified as
get an overall understanding of the scene. containing that specific object (bicycle, horse, etc.).
Standard methods and algorithms include edge detection, When the number of classes go up or the image clarity goes
corner detection, blobs, correlation and thresholding. down, traditional computer vision algorithms find it harder
These techniques are used to extract as many features from to cope and machine learning techniques become more
images of a specific class of object (e.g., bicycles, horses, suitable.
c. Machine learning
• Labelled data
Machine learning uses patterns in large data to perform tasks
• Direct feedback
without being explicitly told what to do.
• Predict outcome/future
There are mainly three different ways machines can learn:
which label the new inputs will be classified as, based on prior
Un
learning
en
m
up
or
Unsupervised learning
inf
ise
12
d. Deep learning
Deep learning is a special subset of machine learning and has Deep learning introduced the concept of end-to-end learning
revolutionised computer vision. Many problems that once where the machine is just given a dataset of images which
seemed improbable to be solved are solved to the point have been annotated with what class of object is present in
where machines are getting better results than humans. each image.
13
g. Machine learning frameworks
Computers learn by viewing thousands of labelled images to understand the traits of what’s being visualised. They learn to
associate characteristics they detect in the images with each label. This method of machine learning means that the same
principle can be applied to diverse areas such as:
• Evaluating the quality of packages in a factory • Identifying trends in the stock market
• Diagnosing organ function from an MRI scan • Locating traffic signs and many more.
There are a great variety of free open-source tools to help to get started with machine learning tasks.
TensorFlow The image processing algorithms can be used for tasks such
An open-source platform for machine learning created by as face recognition, image joining, or tracking moving objects.
Google. It has a comprehensive, flexible ecosystem of tools, Accord also include libraries that provide a more traditional
range of machine learning functions starting from neural
libraries and community resources that lets researchers push the
networks and ending with decision tree systems.
state-of-the-art in ML and developers easily build and deploy
ML-powered applications. Tensorflow works best for image Caffe
classification, image recognition, image segmentation, image
Convolutional Architecture for Fast Feature Embedding (Caffe)
to image translation. Tensorflow includes a set of libraries for
is an open-source framework that can be used for creating and
creating and training custom deep learning models and neural
training popular types of deep learning architectures. Caffe is
networks. Tensorflow supports several popular programming good for tasks such as image classification, segmentation and
languages, including C++, Python, and Java. recognition. Caffe is written in C++ but it also has a Python
PyTorch interface.
A Python based scientific computing package that uses the Google Colab
power of GPUs, currently one of the preferred deep learning Google Colaboratory, or simply Colab, is one of the top
research platforms built to provide maximum flexibility and image processing services. While it’s a cloud service rather
speed. than a framework, it can still be used for building custom
deep learning applications from scratch. Tasks such as image
Keras classification, segmentation and object detection can be
Keras is an open-source Python library for creating deep performed. Google Colab offers free usage of both CPU- and
learning models. It’s a great solution for those who only begin GPU-based acceleration.
to use machine learning algorithms in their projects as it
simplifies the creation of a deep learning model from scratch. NVIDIA DeepStream SDK
To build and deploy AI-powered Intelligent Video Analytics
[Link] apps and services. DeepStream offers a multi-platform
Accord also includes a .Net machine learning framework scalable framework with TLS security to deploy on the edge
combined with audio and image processing libraries written in and connect to any cloud. [Link]
C#. It is a good framework for both creative and general tasks. deepstream-sdk
Computer vision developments are evolving very quickly with new frameworks being written, new networks and datasets being
released and new chips being designed at increasing pace. There are many more frameworks and platforms available, both open-
source or subscription. Picking the right framework for the machine learning application is an important step of project development.
14
6 Embedded vision
Computer vision systems have traditionally relied on a PC due to the processing power required to perform image analysis.
A frame grabber or interface card sends image data from the camera to the computer which then analyses the images and
relays information to another part of the system. These systems can be bulky or complex, however they offer good
performance specifications.
The industry now is using more and more single-board Driven by the need to integrate small cameras into mobile
computers and camera electronics have also become smaller. phones, embedded vision technology advances are now at
New camera and computer systems for applications the stage where it is practical to incorporate computer vision
are now capabilities almost anywhere.
• Highly compact Embedded vision systems are usually easier to use and
integrate than PC-based systems. They often only include a
• Powerful
small camera without a housing connected to a processing
• Low-cost board (embedded board/module) via a connector. The
components are combined into one device and images
• Large memory
sent from the camera are processed directly on the system’s
• Energy-efficient processing board.
Raspberry Pi Zero / Zero W 1GHz, Single Core - 2,4 or 8GB $5 and $10
NVIDIA Jetson Nano Quad-core ARM A57 128-core 4GB 64-bit $99
Maxwell GPU
Seeed Studio Rock Pi N10 Dual Cortex-A72, Mali T860MP4 4/6/8GB $99 - $169
1.8GHz, quad
Cortex-A53
Cheapest: Raspberry Pi Zero / Zero W Best flexibility: NVIDIA Jetson Nano Dev Kit
Best for beginners: Raspberry Pi 4 Best for machine learning with Tensorflow: Google Coral Dev Board
15
b. Camera modules
As image sensor components become smaller, cheaper and more efficient, the range of applications they can be
applied to increases.
• Image quality, with true colours, clear contrast and resolution as important factors
• Easy operation and prototyping capabilities, often with a development kit and plug and play interfaces
c. Interfaces
Choosing the right interface is crucial for any imaging application. Understanding the applications requirements in terms of
resolution, frame rates, transfer speed requirements, among others, will determine the best interface to use. A comparison of
popular digital camera interfaces is shown in the table below.
Interface FireWire 1394.b Camera Link® USB 2.0 USB 3.0 GigE
Data Transfer Rate 800 Mb/s 3.6 Gb/s 480 Mb/s 5Gb/s 1000 Mb/s
16
7 Computer vision & IoT
Connecting computer vision systems to the Internet of Things learning or used to train deep learning models. However,
(IoT) creates a powerful network capability. Being able to machine learning inference and training require substantial
identify objects from cameras allows the local node to be computational and memory resources to run quickly.
more intelligent and have greater autonomy, thus reducing
the processing load on central servers and allowing a more Edge computing, where computer nodes are placed close to
distributed control architecture. end devices, is a viable way to meet the high computation and
low-latency requirements of deep learning on edge devices
Devices such as smartphones and IoT sensors are generating and also provides additional benefits in terms of privacy,
data that needs to be analysed in real time using machine bandwidth efficiency and scalability.
Latency
There is a time lag between the collection and processing of data in the cloud, which is unnoticeable in many use cases.
However, in time-sensitive applications, this time lag, which may only be milliseconds, becomes essential. Real-time inference
is critical to many applications such as autonomous vehicles or voice-based assistance solutions. Sending data to the cloud for
inference or training may incur delays from the network.
Scalability
Sending data from the sources to the cloud consumes significant bandwidth, which in turn increases data processing and
transfer times, introducing scalability issues, as network access to the cloud can become a bottleneck as the number of
connected devices increases.
Privacy
Sending data to the cloud risks privacy concerns from users who own the data or whose behaviours are captured in the data.
Users may be wary of uploading sensitive information to the cloud and how an application may use that data.
17
Vision processing at the edge
Cloud processing is not ideal for real time and mission-critical applications. Once sensors detect an anomaly in a high volume
continuous manufacturing process, for example, the system must take corrective action immediately, otherwise the defect will
propagate. The time from detection to correction must be in seconds.
In the case of a self-driving car, the response time must be in milliseconds. For these applications, the round trip from device to
gateway to the cloud and back takes too long. A different architecture is needed where the data collection and processing are
closer to the devices (or edge).
Edge processing moves the computer, storage and networking closer to the source of the data, significantly reducing travel time
and latency. Embedded smart devices enable more sophisticated processing at the sensor level.
Key to this has been the introduction of lower-cost, compact embedded boards with processing power required for real-time
image analysis. Placing the processing at the edge of the network allows for real-time results, low power consumption, strong
privacy and is a viable solution to meet the challenges introduced by cloud processing. Embedded smart devices are ideal for
repeated and automated robotic processes, such as edge detection in a pick-and-place system.
18
8 Implementing
computer vision
As the demand for intelligent vision solutions grows, tools must integrate computer vision, processing, analytics,
machine learning and connectivity into applications to help translate visual data into meaningful insights.
Cameras Processing
• Angle of illumination
19
b. How CENSIS can help
CENSIS launched the Vision Lab, a dedicated facility to help
businesses adopt or deliver innovative computer vision or
imaging solutions.
20
9 Incubators & learning
resources
Incubators
• NVIDIA Inception
[Link]
Learning Resources
Links to useful online courses,videos and resources:
21
10 The computer vision
community in Scotland
Computer vision research and development in Scotland has a long history
going back to the 1960s with the Department of Machine Intelligence
and Perception at the University of Edinburgh. Research robot Freddy,
built in the 1960s, was one of the earliest systems to integrate perception
and action. Freddy utilised a heavy robot arm fixed to an overhead gantry
with adaptive grippers. A binocular vision system was also mounted to
the gantry. Freddy was able to recognise a variety of objects and could be
instructed to assemble simple artefacts, such as a toy car, from a random
heap of components.
[Link]
a. Companies in Scotland
Advances in computer vision and machine learning are making it possible to build exciting new solutions for a range of industrial
applications. Scotland has a strong base of computer vision companies - a selection is listed below.
22
b. Research in Scotland
There is an important and growing imaging and computer vision research community in universities throughout Scotland.
A selection of research areas and universities are listed in the table below.
SINAPSE Scottish Imaging Medical Imaging – MRI, PET, SPECT, EEG, [Link]
Network deep learning in medical imaging
A consortium of
Aberdeen, Dundee,
Edinburgh, Glasgow,
St Andrews, Stirling,
Strathclyde
23
Glossary
TERM MEANING
AI Artificial Intelligence
Join our
community at
[Link]
24
CENSIS is the centre of excellence for sensor and imaging
systems (SIS) and Internet of Things (IoT) technologies.
Interest in how
machines ‘see’
and how computer
vision can be used
is growing.
Contact details:
CENSIS
The Inovo Building
121 George Street
Glasgow
G1 1RD
Contact details:
Contact
CENSIS details:
Tel: 0141Building
The Inovo 330 3876
CENSIS
Email:
121 Georgeinfo@[Link]
Street
Glasgow
The Inovo Building
G1 1RD
121 George Street
Glasgow
Tel: 0141 330 3876
Email: info@[Link]
G1 1RD
Join theCENSIS
Join the CENSIS mailing
mailing at [Link]
[Link]
list at:
Join the CENSIS mailing list at [Link]
Follow
Follow
Follow usus us
onon on Twitter
Twitter
Twitter
@@CENSIS121
CENSIS121
@CENSIS121
[Link]