0% found this document useful (0 votes)
3 views26 pages

Computer Vision: A Comprehensive Guide

The document is a comprehensive guide to computer vision, detailing its goals, applications, and the technology behind it. It covers the evolution of computer vision, its integration with AI and machine learning, and various sectors where it is applied, such as healthcare, smart cities, and industry. Additionally, it outlines the fundamental techniques, challenges, and resources for implementing computer vision systems.

Uploaded by

ktswiercz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views26 pages

Computer Vision: A Comprehensive Guide

The document is a comprehensive guide to computer vision, detailing its goals, applications, and the technology behind it. It covers the evolution of computer vision, its integration with AI and machine learning, and various sectors where it is applied, such as healthcare, smart cities, and industry. Additionally, it outlines the fundamental techniques, challenges, and resources for implementing computer vision systems.

Uploaded by

ktswiercz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Getting

started with
Computer
Vision
A guide to the knowledge and application
of visual systems

[Link]
The goal of computer vision is to extract
meaning from pixels and perform visual
tasks similar to the human visual system. Interest in
how machines ‘see’ and how computer vision
can be used to build products for consumers and
businesses is growing rapidly.

Rea
Identification

ognition

l-t
Tracking

Capabilities of

ime analys
computer vision
Rec

is

If you are reading the printed version of this brochure, you can download a hyperlinked pdf at [Link]/brochures

1
Contents
1 An introduction to computer vision 3
a. FAQs 3
b. A brief history of computer vision 4
c. The evolution of computer vision 5
d. Deep learning breakthrough 5
2 Application examples 7
a. Smart homes 7
b. Smart cities 7
c. Industry 7
d. Healthcare 8
e. Agriculture 8
f. Security 8
g. Autonomous vehicles 8
h. AR/VR & immersive technologies 8
3 How is computer vision used in business? 9
a. Benefits for business, industry and society 9
b. Technical challenges 9
c. Privacy 9
4 How to set up a computer vision system 10
a. Basic components 10
b. Hardware platforms 10
c. Software tools 10
d. Digital imaging system stack 10
5 How to process and interpret images 11
a. Image as an array 11
b. Image processing 12
c. Machine learning 12
d. Deep learning 13
e. Choosing machine learning or deep learning 13
f. Image processing libraries 13
g. Machine learning frameworks 14

6 Embedded vision 15
a. Embedded vision platforms 15
b. Camera modules 16
c. Interfaces 16

7 Computer vision & IoT 17


a. Cloud vs edge processing 17
b. Cloud platform and machine learning vendors 18
c. Machine vision and IoT 18

8 Implementing computer vision 19


a. Your first prototype 19
b. How CENSIS can help 20
c. IoT2Go Vision kit 20

9 Incubators and learning resources 21

10 The computer vision community in Scotland 22


a. Companies in Scotland 22
b. Research in Scotland 23
Glossary 24

2
1 An introduction to
computer vision
a. FAQs Relationship between AI,
machine learning and deep learning
Artificial intelligence:
What is computer vision? AI is the theory and
al intellige development of computer
Of the five human senses, vision is the one that provides most tifici nc systems to perform tasks
of the data we receive and is considered our dominant sense. Ar normally requiring human

e
intelligence.
It provides us with a detailed description of the surrounding
Machine learning: is an
world which is constantly changing. Although vision involves hine learnin application of AI based
a huge amount of information and complex processing, the
ac around the idea of giving

M
machines access to data

g
human visual system can interpret this information easily. and letting them learn for
themselves.
The ability to see, process and then act on visual input is
Deep Deep learning: is a special
something that most humans take for granted.
learning type of machine learning
algorithm, multiple layers of
Computer vision engineering is the practice of using neural networks that mimic
technology and machines to replicate, and even improve upon, the connectivity of the
human brain in processing
human vision. The technology captures and stores images data and creating patterns
before transforming them into information that can be further for use in decision making.
acted upon. As a minimum an AI system must be able to reproduce aspects
of human intelligence
This requires expertise across a range of fields, including sensor
technology, image and signal processing, computer graphics,
Image processing takes an image as an input and provides a
computer architecture, algorithms and machine learning.
processed image as an output. The purpose of the processing
is usually to improve the quality of the image. Typical methods
What are the fundamental computer vision used are filtering, noise removal, sharpening and edge detection.
techniques?
Computer vision broadens the purpose of image processing
Image classification gives a computer the ability to interpret to include quantitative and qualitative information from visual
the input from an image sensor and categorise what it ‘sees’. data. Similar to the process of human visual reasoning, computer
vision can distinguish between objects, classify them and sort
Object Detection detecting instances of a certain class (such
them according to their attributes. Computer vision, like image
as vehicles, humans, buildings) in images or videos.
processing, takes an image as an input. However, it returns an
Object Tracking detecting and recognizing a defined item output with additional information interpreted from the image
in each frame of a video to distinguish it from other objects in such as size, colour, number, location or orientation.
the scene.
This can be extended beyond the extraction of meaningful
3D Image Reconstruction the process of capturing the shape information from a single image to multiple images or video,
and appearance of real objects. for example, to count the number of cars passing by a point on
the street as they are recorded by a video camera. Temporal
Semantic Image Segmentation when specific regions of an information therefore plays a role in computer vision, much as
image are labelled according to what the object is. it does with our own understanding of the world.

Machine learning is the application of intelligence that


What’s the difference between image provides the computer system with the ability to automatically
processing, computer vision and machine learn and improve from experience without having to be
learning? programmed. In computer vision terms, this means ‘training’ a
Each of these fields is based on the input of an image. They system. Algorithms and statistical models are used to perform
process the pixels and give us an altered output in return. While image analysis using patterns and inference trained on data sets
their names imply their goals and methodologies, these fields of many thousands of images for automatic learning, rather
depend substantially on one another. than using explicit instructions as image processing would.

3
What is artificial intelligence (AI)? The computer vision
Artificial intelligence is intelligence demonstrated by machines,
where any device can perceive its environment and mimic
market is expected
human functions such as ‘learning’ and ‘problem solving’.
Artificial intelligence, or AI, is the broad concept of machines to reach close to $22 billion
being able to carry out tasks in a way that is considered ‘smart’.
by 2026
What are neural networks? [Link]
Neural networks are a means of machine learning, where computer-vision-market-size-and-forecast-to-2025/
a computer learns to perform a task by analysing training
examples or datasets. Usually, the dataset examples have been
manually labelled in advance. An object recognition system
might be fed thousands of labelled images of cars, houses, b. A brief history of
cups and would find visual patterns in the images that correlate
consistently with the particular label.
computer vision
What is deep learning? Computer vision has a long history in commercial and
Deep learning is the use of neural network methods to perform government use where light wave sensors in various spectrum
image analysis, moving away from statistical methods to neural ranges have been deployed in many applications such as:
network algorithms which are developed to mimic the neurons
of the human brain. • Remote sensing for environmental observation
and management

What applications can computer vision • High resolution cameras that collect intelligence over
be used for? battlefields
Applications of computer vision are many and varied. • Thermal imagers to detect people during police operations
Common applications you may be familiar with include
augmented reality, facial recognition, gesture and handwriting • X-ray sensors for airport security.
recognition, machine vision, remote sensing, robotics,
autonomous vehicles, people counting and iris recognition. The sensors can be stationary or attached to moving objects,
such as satellites, drones and vehicles. When combined with
What business sectors use computer vision? connectivity technologies such as Wi-Fi, Bluetooth or 3G/4G/5G,
they create a new set of applications that were not possible before.
Computer vision has numerous applications such as
remote sensing, healthcare (particularly around medical Computer vision, coupled with connectivity, advanced data
imaging such as MRI scans or ultrasound imaging), security, analytics and artificial intelligence, are catalysts for each other,
manufacturing, automotive, transport, robotics, sports, giving rise to revolutionary leaps in IoT innovations
gaming and many others. and applications.

4
c. The evolution of computer vision

1960s
Computer vision technology started in the early 1960s with
the aim of trying to mimic human vision systems and to ask
computers to tell us what they see.

Computers ‘see’ the world differently from humans


Robot cars
• They capture an image as an array of pixels

• Borders between objects are discerned by measuring


tested on
shades of colour roads by Google
• Spatial relations between objects can be estimated.
in 2010
3D models and representations of the environment from
2D images began to be developed. Research continued by
developing ways to analyse real world images which led to 2010s
techniques such as edge detection and segmentation. These
were the foundations for low-level scene understanding and • Hardware technology evolution
steps towards automating the process of image analysis.
Throughout the 2010s, single board computers with
increasingly powerful GPUs, FPGAs and mobile hardware
1970s platforms have been designed, built and adapted to accelerate
machine learning based computer vision algorithms.
The 1970s saw the first commercial application of computer
Increased power and efficiency at lower costs have allowed
vision technology, which was an optical character recognition
breakthroughs in using machine learning for computer vision
program. Combined with text-to-speech technology it provided
and deployment is increasing at an exponential rate.
the first print-to-speech reading machine for the blind.

• Sensor technology developments


1980s
Advancements are also happening rapidly in many areas
In 1980 the precursor of modern convolutional neural beyond conventional camera sensors. For example, infrared
networks was developed. As neural networks evolved sensors and lasers combine to sense depth and distance,
throughout the 1980s, algorithms started to be programmed which are one of the critical enablers of self-driving cars and
to solve individual challenges 3D mapping applications.

2000s • Data generation


Face detection in real-time was first developed in 2001, by One of the driving factors behind the growth of computer
Viola & Jones and was the first object detection framework to vision is the amount of data generated which can be used to
successfully perform in real time. create datasets to train and make computer vision better.

d. Deep learning breakthrough


Although computer vision techniques started in the late 1950s
and many of the machine learning algorithms were developed
in the 1980s, computer vision has grown exponentially
in the last decade due to the increased computational Computer vision has
power offered by processing chips, cloud technologies and
other advancements. Alongside the dedicated hardware grown exponentially
developments, in recent years, the emergence of deep learning
algorithms has reinvigorated computer vision. Throughout
the 2010s, computer performance, accelerated by graphics
in the last decade
processing units (GPUs), have grown powerful enough for us to
realise the capabilities of neural network algorithms.

5
1950:
computer vision
emerges 1957:
pixel invented, first digital
image
1990s:
computer graphics
& computer vision
1966: (image morphing,
MIT Artificial
view interpolation,
intelligence lab
panoramic image
stitching)
1969:
CCD invented
2001:
real-time face
detection

1970s:
first commercial
computer vision 1990s:
application (OCR) projective 3D reconstructions,
stereo imaging, statistical
learning techniques for facial
1975: recognition
first commercial digital
camera 2010s:
GPUs/neural networks

1980s:
mathematical and
quantitative analysis
developments
2012
AlexNet, deep neural
network for image
recognition

6
2 Application examples
a.
Person detection
Computer can be used to adjust
vision-based user data Facial recognition Indoor security lighting and temperature
will increasingly become will be used to cameras will to the number of people
a feature of the home. send an alert to in any room to ensure a
Smart When systems can
unlock the door, or
a smartphone if comfortable
to remain locked if
homes detect and recognise an unfamiliar person an elderly family environment and save
objects, they can deliver member falls, or if a on electricity and a TV
approaches box that recognises
smart actions according toddler is climbing
to what they were up stairs individuals can turn on a
programmed tailored interface for
entertainment
to do

b. Smart cities employ Computer vision


Smart city
applications include
a combination of and related
low power sensors, monitoring of traffic Smart parking
Smart cameras and
technologies
and pedestrian systems could also
can play a significant
cities machine learning role in managing
flows using energy direct motorists to
software to monitor efficient, intelligent a free parking spot
smart cities as they
the efficient working street lighting
serve as the ‘eyes’ of
of the city with ambient light
the city
sensors

c. Computer vision can be Automated Machine vision


Predictive
combined with methods applications such as combined with
and technologies to maintenance and
package inspection, robotics provides defect reduction
Industry provide applications barcode reading, applications also typically use
in industry. Computer 3D inspection, such as product machine vision
vision used in this field is track and trace are and component technology
often referred to as commonly used assembly
‘machine vision’

d. Computer vision
applications in
Healthcare Robots will need
healthcare have been These can be used
robotics can to be able to
developed to aid to detect if elderly
Healthcare healthcare professionals people have fallen help with assisting navigate the world
nurses to clean around them
with medical imaging or require other
hospitals through 3D computer
diagnosis, surgery and forms of assistance
vision
health monitoring

7
The agriculture The quality of food
e.
This can help with Properties such as
industry is better productivity, products can be colour, shape, size,
increasingly using crop monitoring, assessed and graded surface defects and Agriculture
computer vision precision agriculture into specific grades, contamination can
technology for and locating weeds while detecting also be estimated
applications and pests defects

Examples include
Cameras can be
The challenge f.
Intelligent scene Automatic Number is to identify the
placed in offices,
monitoring systems Plate Recognition scene and context,
hospitals, banks,
are playing an (ANPR), people and
ports, car parks,
understanding what Security
increasingly vehicle tracking, demands immediate
stadiums, shopping
significant role in crowd analysis and attention, what is
centres, airports
society zone detection for valuable and what
and more
health & safety can be ignored

Computer vision
Self-driving vehicles
must be able to
capture visual data in
High quality
images and videos
must be obtained
g.
Self-driving vehicles technology is
can be made real time to create 3D in low light
being applied maps to understand Autonomous
intelligent, self-reliant conditions as well
to autonomous the surroundings,
and reliable using vehicles to make it
as daylight, using vehicles
computer vision while detecting and LiDAR sensors and
safe for passengers classifying objects
technology thermal cameras
and pedestrians in their path such alongside visible
as traffic lights and camera sensors
pedestrians

Computer vision aids


AR/VR applications in AR/VR applications
h.
virtual reality with Computer vision-
vision capabilities like based AR overlays
e-commerce allow in the healthcare AR/VR and
the user to visualise industry empower
SLAM (Simultaneous imagery or audio products within their professionals to immersive
localisation and onto existing real-
mapping), user body world scenery
homes or virtually try provide better technologies
tracking and gaze on clothes to find the diagnosis and make
tracking perfect fit surgery safer

8
3 How is computer vision
used in business?
Facial recognition Financial institutions

Autonomous There are a huge


vehicles Manufacturing
range of applications
where the ability to
extract meaning from
Medicine ‘seeing’ visual data Agriculture
is useful

Digital marketing Handwriting extraction


and analysis

a. Benefits for business, industry and society


Computer vision has the potential to revolutionise many • Accurate outcomes – computer vision systems can
everyday aspects of our lives. Having the ability to see and provide high quality image processing capabilities
interpret a scene reliably and without tiring, computer vision
systems automate tasks without needing human intervention. • Cost-reductions – errors and therefore faulty products
As a result, business users can have benefits such as or services can be minimised, so companies can save
a lot of money that would otherwise be spent on
• Faster and simpler processes – computer vision systems fixing flawed processes and products
can carry out monotonous, repetitive tasks at a faster rate,
making the entire process simpler

b. Technical challenges
There is a high level of technical understanding team of professionals with technical expertise.
required to create software that collects and interprets Companies may also need to have a dedicated team
visual data. To train a computer vision system for regular monitoring and evaluation of the vision
powered by machine learning, companies need to have a system performance.

c. Privacy
Privacy is the biggest social threat that computer vision poses. their behaviour or monitoring their habits. Everyone’s
The capabilities of computer vision – identification, recognition, information is stored on a cloud.
tracking and real-time analysis – impact directly with individual It is important to understand the potential negative effects of
rights for privacy. With computers learning from thousands and computer vision applications on society. This is crucial to ensure
thousands of images and videos, computers are getting better that computer vision applications make our lives more comfortable
at recognising individuals by their facial features, by identifying and efficient and not for purposes of constrain and control.

9
4 How to set up a computer
vision system
Almost everyone has experienced computer vision and machine learning, often without even knowing.
This section explains how to set up a computer vision system.

a. Basic components
The components of a standard computer vision system are:

• Digital camera/image sensor • Lens


At the heart of any camera is the sensor. Modern sensors To focus or enhance the scene
are solid-state electronic devices containing up to millions • Frame grabber
of discrete photodetector sites called pixels. To capture individual frames
• Lighting devices • Image processing software
Many computer vision systems are optimised by illuminating To analyse the captured scene
the scene to be captured, and may require filters to enhance • Machine learning algorithms
the sensor characteristics. For pattern recognition

b. Hardware platforms
CPU and can run in to the thousands. The greater number of
The central processing unit of a computer used to cores allows multiple calculations to be worked on at the
perform arithmetic computations. Most modern CPUs same time which allows image processing to be
have 2 to 256 cores. performed efficiently.

GPU FPGA
The graphics processing unit of a computer used to Field programmable gate arrays have parallel processing
process graphics. GPUs start at a couple of hundred cores capabilities which make them suitable for image processing.

c. Software tools
There are many software tools with the necessary techniques to perform image and video processing tasks as well as machine
learning algorithms.

CPU GPU
• OpenCV • Scilab • Octave • R • Matlab • Tensorflow • PyTorch • Keras • Caffe

d. Digital imaging system stack


Presentation
Software 6 Visualisation and reproduction Viewing image in visual format

processing
5 Image post-processor Image data optimisation

Numeric
4 Image storage Formatting and storing image data presentation

Hardware 3 Digital signal processor Manipulation of digital signal


processing
2 Sensor Converting light to electrical signal

1 Optics Gathering Light


Light

10
5 How to process and
interpret images
a. Image as an array
A digital image is an array of pixels where each pixel is a A grayscale image refers to the number of different shades, or
combination of numerical values representing the colours depth, of a particular colour. A grayscale image can be created
and intensities at a particular point on the image. from any single channel or colour of the image.

A pixel, or picture element, is the smallest visual element of an The image resolution gives the number of pixels and the aspect
image and typically contains three component intensities, or ratio gives the width:height pixel ratio.
channels, such as red, green and blue. Colour digital images The channel contains the number of samples per point, which for
are created by combining the channels to reproduce the broad grayscale images is a single sample per pixel, whereas for colour
range of colours seen by the human eye. images is three samples per pixel (red, green, blue).

Pixel
Smallest visual element

10 10 16 28
Digital Image
65 70 56 43
9 90 96 67
A multidimensional array 99 70 56 78
32 90 96 67
of numbers 15 85 43 92
60 90 96 67
Aspect Ratio 23 85 43 92
32 65 87 99
Width:Height 85 85 43 92
Channel: No. of samples per point 54 65 87 99
Resolution
32 65 87 99
Width x Height Single plane: Grayscale / B&W images

Three Planes: Colour images

©maxEmbedded.com2012

11
b. Image processing
The main purpose of image processing is to improve the etc.) and treat those features as a ‘definition’ of the object.
quality of the image by sharpening and restoration; extract These ‘definitions’ are then searched for in other images. If
the features of an image to help discriminate objects and/or a significant number of features from one type of object are
classes of objects; classify objects, locate their position and found in another image, the image can then be classified as
get an overall understanding of the scene. containing that specific object (bicycle, horse, etc.).
Standard methods and algorithms include edge detection, When the number of classes go up or the image clarity goes
corner detection, blobs, correlation and thresholding. down, traditional computer vision algorithms find it harder
These techniques are used to extract as many features from to cope and machine learning techniques become more
images of a specific class of object (e.g., bicycles, horses, suitable.

The main steps for image processing are:


Image acquisition Image compression and decompression
Captures the image with a sensor or camera and convert it into To allow for changes in image resolution and size, to
a manageable format reduce or restore images depending on the requirement
Image enhancement Morphological processing
The input image quality is enhanced and important Defines the object structure and shape in the image
details extracted
Feature extraction
Image restoration
For a particular object, the specific features are identified
Any corruption such as blur, noise, or camera misfocus
is removed to get a cleaner image in the image and techniques like object detection are used
Colour image processing Representation and description
The coloured images are processed with RGB or other Store and visualise the processed data with a suitable file
colour space methods format and output

c. Machine learning
• Labelled data
Machine learning uses patterns in large data to perform tasks
• Direct feedback
without being explicitly told what to do.
• Predict outcome/future
There are mainly three different ways machines can learn:

Supervised learning algorithms


These are designed to learn by example. When training a
supervised learning algorithm, the training data will consist
of inputs paired with the correct outputs. During training, the Supervised
algorithm will search for patterns in the data that correlate
with the desired outputs. After training, a supervised learning
algorithm will take in new unseen inputs and will determine Machine
t

which label the new inputs will be classified as, based on prior
Un

learning
en

training data. The objective of a supervised learning model is


s

m
up

to predict the correct label for newly presented input data.


ce
erv

or

Unsupervised learning
inf
ise

Give the machine unlabelled data and it will find patterns in


Re
d

the data. The algorithm will pick up the difference between


objects as they find logical patterns.
• No labels • Decision process
Reinforcement learning
• No feedback • Reward system
The algorithm is trained in a reward and punishment • ‘Find hidden structure’ • Learn series of actions
mechanism. The agent is rewarded for correct moves and
punished for the wrong ones. In doing so, the algorithm tries
to minimize wrong moves and maximize the right ones.

12
d. Deep learning
Deep learning is a special subset of machine learning and has Deep learning introduced the concept of end-to-end learning
revolutionised computer vision. Many problems that once where the machine is just given a dataset of images which
seemed improbable to be solved are solved to the point have been annotated with what class of object is present in
where machines are getting better results than humans. each image.

e. Choosing machine learning or deep learning


Classic computer vision analysis excels at measurements, Deep learning has unlocked a myriad of sophisticated new AI
finding defects or matching patterns. It is the ideal solution applications.
for repeatable dimension measurements of an object in a
The performance of deep learning algorithms with complex
controlled environment, such as examining machined parts or
tasks have made it particularly appealing as a solution.
printed circuit boards. Traditional techniques work very well in
However, it is not always the best approach to computer
constrained environments. However they don’t handle novel
vision and machine learning related problems. Deep learning
situations very well.
methods are ideal for replacing human eyes for object
In comparison, machine learning is trainable, and as it gains classification problems, or to emulate expertise by interpreting
access to a wider data set, it’s able to locate, identify and images such as medical x-rays. Deep learning algorithms take
segment a wider number of objects or faults with more a long time to train however, requiring a lot of code compared
variable appearance or perspective, such as identifying and to relatively few lines of classic computer vision code.
counting foods such as broccoli on a conveyor belt.
While each use case is unique and will depend on business
Breakthroughs in the field of artificial neural networks in recent objectives, AI maturity, timescale, data and resources, among
years have driven companies across industries to implement other things, are general considerations to take into account
deep learning solutions, from chatbots in customer service before deciding whether or not to use deep learning to solve
to image and object recognition in retail, and many more. a given problem.

f. Image processing libraries


OpenCV Scilab An open source software similar to MATLAB,
with a computer vision and image processing module.
The Open Source Computer Vision Library (OpenCV)
is one of the most popular computer vision libraries that
provides many algorithms and functions. It includes Octave An open source software also similar to MATLAB,
modules such as image processing, object detection and with a computer vision and image processing module.
deep learning to name just a few.
The library is written in C++ and supports C++, Java, Python R An open source data analysis library with packages for
and MATLAB interfaces. image processing.

13
g. Machine learning frameworks
Computers learn by viewing thousands of labelled images to understand the traits of what’s being visualised. They learn to
associate characteristics they detect in the images with each label. This method of machine learning means that the same
principle can be applied to diverse areas such as:

• Evaluating the quality of packages in a factory • Identifying trends in the stock market

• Diagnosing organ function from an MRI scan • Locating traffic signs and many more.

There are a great variety of free open-source tools to help to get started with machine learning tasks.

TensorFlow The image processing algorithms can be used for tasks such
An open-source platform for machine learning created by as face recognition, image joining, or tracking moving objects.
Google. It has a comprehensive, flexible ecosystem of tools, Accord also include libraries that provide a more traditional
range of machine learning functions starting from neural
libraries and community resources that lets researchers push the
networks and ending with decision tree systems.
state-of-the-art in ML and developers easily build and deploy
ML-powered applications. Tensorflow works best for image Caffe
classification, image recognition, image segmentation, image
Convolutional Architecture for Fast Feature Embedding (Caffe)
to image translation. Tensorflow includes a set of libraries for
is an open-source framework that can be used for creating and
creating and training custom deep learning models and neural
training popular types of deep learning architectures. Caffe is
networks. Tensorflow supports several popular programming good for tasks such as image classification, segmentation and
languages, including C++, Python, and Java. recognition. Caffe is written in C++ but it also has a Python
PyTorch interface.
A Python based scientific computing package that uses the Google Colab
power of GPUs, currently one of the preferred deep learning Google Colaboratory, or simply Colab, is one of the top
research platforms built to provide maximum flexibility and image processing services. While it’s a cloud service rather
speed. than a framework, it can still be used for building custom
deep learning applications from scratch. Tasks such as image
Keras classification, segmentation and object detection can be
Keras is an open-source Python library for creating deep performed. Google Colab offers free usage of both CPU- and
learning models. It’s a great solution for those who only begin GPU-based acceleration.
to use machine learning algorithms in their projects as it
simplifies the creation of a deep learning model from scratch. NVIDIA DeepStream SDK
To build and deploy AI-powered Intelligent Video Analytics
[Link] apps and services. DeepStream offers a multi-platform
Accord also includes a .Net machine learning framework scalable framework with TLS security to deploy on the edge
combined with audio and image processing libraries written in and connect to any cloud. [Link]
C#. It is a good framework for both creative and general tasks. deepstream-sdk

Computer vision developments are evolving very quickly with new frameworks being written, new networks and datasets being
released and new chips being designed at increasing pace. There are many more frameworks and platforms available, both open-
source or subscription. Picking the right framework for the machine learning application is an important step of project development.

14
6 Embedded vision
Computer vision systems have traditionally relied on a PC due to the processing power required to perform image analysis.
A frame grabber or interface card sends image data from the camera to the computer which then analyses the images and
relays information to another part of the system. These systems can be bulky or complex, however they offer good
performance specifications.

The industry now is using more and more single-board Driven by the need to integrate small cameras into mobile
computers and camera electronics have also become smaller. phones, embedded vision technology advances are now at
New camera and computer systems for applications the stage where it is practical to incorporate computer vision
are now capabilities almost anywhere.

• Highly compact Embedded vision systems are usually easier to use and
integrate than PC-based systems. They often only include a
• Powerful
small camera without a housing connected to a processing
• Low-cost board (embedded board/module) via a connector. The
components are combined into one device and images
• Large memory
sent from the camera are processed directly on the system’s
• Energy-efficient processing board.

a. Embedded vision platforms


There are many popular devices that are commonly used for running computer vision algorithms

Provider Board CPU GPU RAM Price

Raspberry Pi Zero / Zero W 1GHz, Single Core - 2,4 or 8GB $5 and $10

NVIDIA Jetson Nano Quad-core ARM A57 128-core 4GB 64-bit $99
Maxwell GPU

Raspberry Pi RPi 4 Quad core Broadcom 1, 2 or 4GB $35 - $55


Cortex-A72 VideoCore VI

Google Coral dev board NXP [Link] 8M SOC Integrated GC7000


(quad Cortex-A53, Lite Graphics
Cortex-M4F)

Seeed Studio Rock Pi N10 Dual Cortex-A72, Mali T860MP4 4/6/8GB $99 - $169
1.8GHz, quad
Cortex-A53

Cheapest: Raspberry Pi Zero / Zero W Best flexibility: NVIDIA Jetson Nano Dev Kit

Best for beginners: Raspberry Pi 4 Best for machine learning with Tensorflow: Google Coral Dev Board

15
b. Camera modules
As image sensor components become smaller, cheaper and more efficient, the range of applications they can be
applied to increases.

• Image quality, with true colours, clear contrast and resolution as important factors

• Easy operation and prototyping capabilities, often with a development kit and plug and play interfaces

• Easy system integration with well-defined interfaces and software protocols

c. Interfaces
Choosing the right interface is crucial for any imaging application. Understanding the applications requirements in terms of
resolution, frame rates, transfer speed requirements, among others, will determine the best interface to use. A comparison of
popular digital camera interfaces is shown in the table below.

Comparison of popular digital camera interfaces

Interface FireWire 1394.b Camera Link® USB 2.0 USB 3.0 GigE

Data Transfer Rate 800 Mb/s 3.6 Gb/s 480 Mb/s 5Gb/s 1000 Mb/s

Max Cable Length 100m 10m 5m 3m 100m

No. devices Up to 63 1 Up to 127 Up to 127 Unlimited

Connector 9pin-9pin 26pin USB USB Rj45/Cat53 or 6

Capture board Optional Required Optional Optional Not required

Power Optional Required Optional Optional Required

Source: Edmund Optics: [Link]

16
7 Computer vision & IoT
Connecting computer vision systems to the Internet of Things learning or used to train deep learning models. However,
(IoT) creates a powerful network capability. Being able to machine learning inference and training require substantial
identify objects from cameras allows the local node to be computational and memory resources to run quickly.
more intelligent and have greater autonomy, thus reducing
the processing load on central servers and allowing a more Edge computing, where computer nodes are placed close to
distributed control architecture. end devices, is a viable way to meet the high computation and
low-latency requirements of deep learning on edge devices
Devices such as smartphones and IoT sensors are generating and also provides additional benefits in terms of privacy,
data that needs to be analysed in real time using machine bandwidth efficiency and scalability.

a. Cloud vs edge processing


Computer vision tasks typically require fast processing capabilities – particularly for real-time image and scene understanding.

Vision processing in the cloud


To use cloud resources, data must be moved from the data source location on the network edge (i.e. camera modules,
smartphones) to a centralised location on the cloud using a remote server or data centre. Moving data from the source to the
cloud can introduce several challenges

Latency
There is a time lag between the collection and processing of data in the cloud, which is unnoticeable in many use cases.
However, in time-sensitive applications, this time lag, which may only be milliseconds, becomes essential. Real-time inference
is critical to many applications such as autonomous vehicles or voice-based assistance solutions. Sending data to the cloud for
inference or training may incur delays from the network.

Scalability
Sending data from the sources to the cloud consumes significant bandwidth, which in turn increases data processing and
transfer times, introducing scalability issues, as network access to the cloud can become a bottleneck as the number of
connected devices increases.

Privacy
Sending data to the cloud risks privacy concerns from users who own the data or whose behaviours are captured in the data.
Users may be wary of uploading sensitive information to the cloud and how an application may use that data.

17
Vision processing at the edge
Cloud processing is not ideal for real time and mission-critical applications. Once sensors detect an anomaly in a high volume
continuous manufacturing process, for example, the system must take corrective action immediately, otherwise the defect will
propagate. The time from detection to correction must be in seconds.

In the case of a self-driving car, the response time must be in milliseconds. For these applications, the round trip from device to
gateway to the cloud and back takes too long. A different architecture is needed where the data collection and processing are
closer to the devices (or edge).

Edge processing moves the computer, storage and networking closer to the source of the data, significantly reducing travel time
and latency. Embedded smart devices enable more sophisticated processing at the sensor level.

Key to this has been the introduction of lower-cost, compact embedded boards with processing power required for real-time
image analysis. Placing the processing at the edge of the network allows for real-time results, low power consumption, strong
privacy and is a viable solution to meet the challenges introduced by cloud processing. Embedded smart devices are ideal for
repeated and automated robotic processes, such as edge detection in a pick-and-place system.

b. Cloud platform and machine learning vendors


Cloud capabilities and resources for machine learning are increasingly significantly. Today, cloud computing providers increasingly
offer GPU and FPGA co-processors to accelerate processing workloads, including:

Amazon Rekognition: [Link] IBM Watson Visual Recognition: [Link]


cloud/watson-visual-recognition
Amazon Sagemaker: [Link]
Microsoft Azure Computer Vision API: [Link]
Google Cloud Vision API: [Link] com/en-gb/services/cognitive-services/computer-vision/

c. Machine vision and IoT


Machine vision systems, which is a general term for computer operations and open up a wide range of applications,
vision used for industrial applications, connected to the IoT can offering valuable insights into the operation of industrial
create a powerful network capability. Allowing the local node systems. This in turn is opening up new ways of monitoring
to be more intelligent and have greater autonomy, reducing equipment and connecting autonomous robotic systems
the processing load on central servers, can provide efficient to the IoT infrastructure.

18
8 Implementing
computer vision
As the demand for intelligent vision solutions grows, tools must integrate computer vision, processing, analytics,
machine learning and connectivity into applications to help translate visual data into meaningful insights.

a. Your first prototype


To develop a computer vision prototype, there are many important considerations to be made regarding choices
of hardware and software to suit the application requirements. A brief summary of the main areas are:

Cameras Processing

• Image sensor performance • Suitable development environment

• Camera features and characteristics • Software development kits

• Data rate/transfer • Language flexibility

• Camera/PC interfaces • Single processor/multi-core support

• Image processing tools


Optics
• GPU utilisation
• Focal length
• Cloud-based analytics
• Field of view
• Memory requirements
• Magnification

• Image quality Output/Display

• Processed image or video


Illumination
• Data analytics of object count or location
• How to enhance features of interest
• System information or monitoring
• Monochrome or colour image

• Angle of illumination

19
b. How CENSIS can help
CENSIS launched the Vision Lab, a dedicated facility to help
businesses adopt or deliver innovative computer vision or
imaging solutions.

CENSIS is uniquely positioned to help kickstart or accelerate


businesses’ use of computer vision due to our connections
with academia and industry and the funding we can bring to
innovative technology projects. Our in-house technical and
business development teams can also provide engineering
support and consultancy.

The hardware and software we have in the Vision Lab greatly


develops our technical capabilities in computer vision and
related fields and includes:

• Development kits for image sensing, machine learning

• 3D time-of-flight sensor for industrial machine


vision applications

• Machine vision cameras


c. IoT2Go Vision Kit
• Camera modules for embedded vision
CENSIS has created IoT2Go, a series of plug and play IoT
• MVTec Halcon, a comprehensive standard software development kits for organisations to try out an IoT solution in
package for machine vision industries, with capabilities their own premises. IoT2Go was developed as part of a Scottish
in areas such as analysis, matching, measuring, Government programme to raise awareness of IoT technologies.
identification, 3D vision and deep learning algorithms
The kits are quick and easy to set up and can be used by people
with no technical or coding experience.
The CENSIS Vision Lab can help SMEs with product
development or product enhancement around new The IoT2Go Vision kit can capture a real-time count of
technology, through funding, collaboration, consultancy people and objects and has image classification and object
and access to equipment and expertise. detection demos.

20
9 Incubators & learning
resources
Incubators
• NVIDIA Inception
[Link]

• NVIDIA Deep Learning Institute


[Link]

• Intel Edge AI Incubator


[Link]

• Imagimob AI Early Access Program


[Link]

Learning Resources
Links to useful online courses,videos and resources:

• Introduction to Computer Vision on Udacity, free course


[Link]

• Awesome Computer Vision, a list of resources on Github


[Link]

• Computer Vision course by Subhransu Maji


[Link]

• Video Tutorial by Alberto Romay


[Link]

21
10 The computer vision
community in Scotland
Computer vision research and development in Scotland has a long history
going back to the 1960s with the Department of Machine Intelligence
and Perception at the University of Edinburgh. Research robot Freddy,
built in the 1960s, was one of the earliest systems to integrate perception
and action. Freddy utilised a heavy robot arm fixed to an overhead gantry
with adaptive grippers. A binocular vision system was also mounted to
the gantry. Freddy was able to recognise a variety of objects and could be
instructed to assemble simple artefacts, such as a toy car, from a random
heap of components.

[Link]

The Department of Machine Intelligence and Perception, later the


Department of Artificial Intelligence, was the forerunner to both the Turing
Institute in Glasgow, formed in 1983 and developed to combine research
in AI with technology transfer to industry, and also to the current School of
Informatics at Edinburgh which has leading research expertise in computer
vision and machine learning.

a. Companies in Scotland
Advances in computer vision and machine learning are making it possible to build exciting new solutions for a range of industrial
applications. Scotland has a strong base of computer vision companies - a selection is listed below.

Company City Specialist Areas Link

Odos Imaging Edinburgh 3D sensing and imaging solutions [Link]

Peacock Stirling Robotics, automation, image processing, [Link]


Technology machine vision

Five AI Edinburgh Autonomous vehicles [Link]

Machines Edinburgh Highly accurate train positioning system [Link]


with Vision for continuously monitoring track condition

Optos Dunfermline Retina imaging devices and development [Link]

STMicroelectronics Edinburgh CMOS image sensor development, imaging systems, [Link]


Design Centre optical engineering, semiconductor solutions for
autonomous driving and IoT

NCTech Edinburgh High-resolution 360deg imagery and LiDAR [Link]

Sense Edinburgh 3D perception systems for mobility, [Link]


Photonics industrial and robotics autonomy

22
b. Research in Scotland
There is an important and growing imaging and computer vision research community in universities throughout Scotland.
A selection of research areas and universities are listed in the table below.

University Research Group Areas of Research Link

Heriot-Watt Visionlab Robotics, [Link]


University Automotive driver assistance,
Surveillance,
Human behaviour inference,
Detection & tracking,
Analysis of shape in 2D,
Range and LiDAR analysis
Signal & Image MRI, Ultrasound imaging, Novel imaging [Link]
Processing modalities, Imaging techniques in radio engineering-physical-sciences/institutes/
Laboratory and optical astronomy sensors-signals-systems/[Link]

Dundee Computer Healthcare and biomedical imaging, [Link]


University Vision & Image Visual perception of people and places.
Processing

The University Edinburgh Centre Virtual reality environments [Link]


of Edinburgh for Robotics
Machine Iconic vision in 2D [Link]
Vision Unit
Institute of Statistical machine learning, computer [Link]
Perception, Action vison, mobile and humanoid robotics,
and Behaviour motor control, graphics and visualisation

University Computer Vision 3D vision systems [Link]


of Glasgow & Autonomous research/researchsections/ida-section/
Systems computervisionandautonomoussystems/
Computer Human body modelling in 3D [Link]
Vision & Graphics

University Vision & Image Text and language processing, [Link]


of Stirling Processing Special visual perception
Interest Group

Glasgow School of Neural networks applied to condition [Link]


Caledonian Computing, monitoring
University Engineering &
Built Environment

Glasgow School of Real-time 3D visualisation, interaction [Link]


School of Art Simulation and technologies centres/school-of-simulation-and-
Visualisation visualisation/

SINAPSE Scottish Imaging Medical Imaging – MRI, PET, SPECT, EEG, [Link]
Network deep learning in medical imaging
A consortium of
Aberdeen, Dundee,
Edinburgh, Glasgow,
St Andrews, Stirling,
Strathclyde

23
Glossary
TERM MEANING

AI Artificial Intelligence

CMOS Complementary Metal-Oxide Semiconductor

CPU Central Processing Unit

FPGA Field Programmable Gate Array

GPU Graphics Processing Unit

IoT Internet of Things

IoT2Go CENSIS IoT starter kit

LiDAR Light Detection And Ranging

SDK Software Development Kit

SME Small and Medium Enterprises

Join our
community at
[Link]

24
CENSIS is the centre of excellence for sensor and imaging
systems (SIS) and Internet of Things (IoT) technologies.

We help organisations of all sizes explore innovation


and overcome technology barriers to achieve business
transformation.

As one of Scotland’s Innovation Centres, our focus is not


only creating sustainable economic value in the Scottish
economy, but also generating social benefit. Our industry-
experienced engineering and project management teams
work with companies or in collaborative teams with university
research experts.
We act as independent trusted advisers, allowing organisations
to implement quality, efficiency and performance
improvements and fast-track the development of new
products and services for global markets.

Interest in how
machines ‘see’
and how computer
vision can be used
is growing.
Contact details:
CENSIS
The Inovo Building
121 George Street
Glasgow
G1 1RD
Contact details:
Contact
CENSIS details:
Tel: 0141Building
The Inovo 330 3876
CENSIS
Email:
121 Georgeinfo@[Link]
Street
Glasgow
The Inovo Building
G1 1RD
121 George Street
Glasgow
Tel: 0141 330 3876
Email: info@[Link]
G1 1RD

Tel: 0141 330 3876


Email: info @[Link]

Join theCENSIS
Join the CENSIS mailing
mailing at [Link]
[Link]
list at:
Join the CENSIS mailing list at [Link]
Follow
Follow
Follow usus us
onon on Twitter
Twitter
Twitter
@@CENSIS121
CENSIS121
@CENSIS121

[Link]

You might also like