Deep Learning
Deep Learning
History
• In 1995, Yann LeCun
and Yoshua Bengio
introduced the concept
of convolutional neural
networks.
“Deep Learning doesn’t do different things,
it does things differently”
Performance vs Sample Size
Performance
Traditional ML algorithms
Size of Data
Outline
Supervised Learning
Convolutional Neural Network
Sequence Modelling: RNN and its extensions
Unsupervised Learning
Autoencoder
Stacked Denoising Autoencoder
• Unsupervised Learning (+Supervised)
Generative Adversarial Networks
Reinforcement Learning
Deep Reinforcement Learning
Outline
GAN
Output
Generator
Generated
Input
Network
Shakespeare Fake
Poetry Discriminator
Network
Real Real
Shakespeare
Poetry
Supervised Learning
Traditional pattern recognition models work with hand crafted
features and relatively simple trainable classifiers.
Trainable
Extract Hand Output
Classifier
Crafted (e.g. Outdoor
(e.g. SVM,
Features Yes or No)
Random
Limitations Forrest)
Limitations
Image
Pixel Edge Texture Motif Part Object
Text
Character Word Word-group Clause Sentence Story
x1 a1(1)
x2 a2(1)
a1(2) Y
x3 a3(1)
x4 a4(1)
x4 a4(1)
𝑎1(1) = 𝑓 𝑤1 ∗ 𝑥1 + 𝑤2 ∗ 𝑥2 + 𝑤3 ∗ 𝑥3 + 𝑤4 ∗ 𝑥4
𝑅𝑒𝑙𝑢: max(0, 𝑥)
𝑎1(1) = 𝑚𝑎𝑥 0, 𝑤1 ∗ 𝑥1 + 𝑤2 ∗ 𝑥2 + 𝑤3 ∗ 𝑥3 + 𝑤4 ∗ 𝑥4
Number of Parameters
Softmax
x1 a1(1)
x2 a2(1)
a1(2) Y
x3 a3(1)
x4 a4(1)
4*4 + 4 +1
If the input is an Image?
x1 a1(1)
x2 a2(1)
a1(2) Y
x3 a3(1)
400 X 400 X 3
a480000(1)
x480000
Number of Parameters
480000*480000 + 480000 +1 = approximately 230 Billion !!!
480000*1000 + 1000 +1 = approximately 480 million !!!
Let us see how convolutional layers
help.
Convolutional Layers
0 1 0
Filter 1 -4 1
0 1 0
1 1 1 1 1 1 0.015686 0.015686 0.011765 0.015686 0.015686 0.015686 0.015686 0.964706 0.988235 0.964706 0.866667 0.031373 0.023529 0.007843
0.007843 0.741176 1 1 0.984314 0.023529 0.019608 0.015686 0.015686 0.015686 0.011765 0.101961 0.972549 1 1 0.996078 0.996078 0.996078 0.058824 0.015686
0.019608 0.513726 1 1 1 0.019608 0.015686 0.015686 0.015686 0.007843 0.011765 1 1 1 0.996078 0.031373 0.015686 0.019608 1 0.011765
0.015686 0.733333 1 1 0.996078 0.019608 0.019608 0.015686 0.015686 0.011765 0.984314 1 1 0.988235 0.027451 0.015686 0.007843 0.007843 1 0.352941
0.015686 0.823529 1 1 0.988235 0.019608 0.019608 0.015686 0.015686 0.019608 1 1 0.980392 0.015686 0.015686 0.015686 0.015686 0.996078 1 0.996078
0.015686 0.913726 1 1 0.996078 0.019608 0.019608 0.019608 0.019608 1 1 0.984314 0.015686 0.015686 0.015686 0.015686 0.952941 1 1 0.992157
0.019608 0.913726 1 1 0.988235 0.019608 0.019608 0.019608 0.039216 0.996078 1 0.015686 0.015686 0.015686 0.015686 0.996078 1 1 1 0.007843
0.019608 0.898039 1 1 0.988235 0.019608 0.015686 0.019608 0.968628 0.996078 0.980392 0.027451 0.015686 0.019608 0.980392 0.972549 1 1 1 0.019608
0.043137 0.905882 1 1 1 0.015686 0.035294 0.968628 1 1 0.023529 1 0.792157 0.996078 1 1 0.980392 0.992157 0.039216 0.023529
1 1 1 1 1 0.992157 0.992157 1 1 0.984314 0.015686 0.015686 0.858824 0.996078 1 0.992157 0.501961 0.019608 0.019608 0.023529
0.996078 0.992157 1 1 1 0.933333 0.003922 0.996078 1 0.988235 1 0.992157 1 1 1 0.988235 1 1 1 1
0.015686 0.74902 1 1 0.984314 0.019608 0.019608 0.031373 0.984314 0.023529 0.015686 0.015686 1 1 1 0 0.003922 0.027451 0.980392 1
0.019608 0.023529 1 1 1 0.019608 0.019608 0.564706 0.894118 0.019608 0.015686 0.015686 1 1 1 0.015686 0.015686 0.015686 0.05098 1
0.015686 0.015686 1 1 1 0.047059 0.019608 0.992157 0.007843 0.011765 0.011765 0.015686 1 1 1 0.015686 0.019608 0.996078 0.023529 0.996078
0.019608 0.015686 0.243137 1 1 0.976471 0.035294 1 0.003922 0.011765 0.011765 0.015686 1 1 1 0.988235 0.988235 1 0.003922 0.015686
0.019608 0.019608 0.027451 1 1 0.992157 0.223529 0.662745 0.011765 0.011765 0.011765 0.015686 1 1 1 0.015686 0.023529 0.996078 0.011765 0.011765
0.015686 0.015686 0.011765 1 1 1 1 0.035294 0.011765 0.011765 0.011765 0.015686 1 1 1 0.015686 0.015686 0.964706 0.003922 0.996078
0.007843 0.019608 0.011765 0.054902 1 1 0.988235 0.007843 0.011765 0.011765 0.015686 0.011765 1 1 1 0.015686 0.015686 0.015686 0.023529 1
0.007843 0.007843 0.015686 0.015686 0.960784 1 0.490196 0.015686 0.015686 0.015686 0.007843 0.027451 1 1 1 0.011765 0.011765 0.043137 1 1
0.023529 0.003922 0.007843 0.023529 0.980392 0.976471 0.039216 0.019608 0.007843 0.019608 0.015686 1 1 1 1 1 1 1 1 1
a b c d w1 w2
h1 h2
e f g h w3 w4
i j k l
m n o p
ℎ2 = 𝑓 𝑏 ∗ 𝑤1 + 𝑐 ∗ 𝑤2 + 𝑓 ∗ 𝑤3 + 𝑔 ∗ 𝑤4
w1 w2
w3 w4
w5 w6
w7 w8
Filter 1
Filter 2
Input Image
Layer 1 Layer 2
Feature Map Feature Map
In Convolutional neural networks, hidden units are only connected to local receptive field.
Pooling
Max pooling: reports the maximum output within a rectangular
neighborhood.
Average pooling: reports the average output of a rectangular
neighborhood.
Living Room
Bed Room
128
256
256
512
512
512
512
256
128
512
512
Kitchen
64
64
Bathroom
Outdoor
Max Pool
Filter
Fully Connected
Layers
Convolutional Neural Networks
Output: Binary, Multinomial, Continuous, Count
Input: fixed size, can use padding to make all images same
size.
Architecture: Choice is ad hoc
requires experimentation.
Optimization: Backward propagation
hyper parameters for very deep model can be estimated properly only if you
have billions of images.
Use an architecture and trained hyper parameters from other papers
(Imagenet or Microsoft/Google APIs etc)
Computing Power: Buy a GPU!!
Automatic Colorization of Black and White Images
Optimizing Images
Kernel
f * g (x) f ( )g(x )d
N 1 Output is
f ( )g(x ) sometimes called
0 Feature map
2D (continuous, discrete) :
f * g (x, y) f ( , )g(x , y )dd
N 1 N 1
f ( , )g(x , y )
0 0
Convolution Properties
• Commutative:
f*g = g*f
• Associative:
(f*g)*h = f*(g*h)
• Homogeneous:
f*(g)= f*g
• Additive (Distributive):
f*(g+h)= f*g+f*h
• Shift-Invariant
f*g(x-x0,y-yo)= (f*g) (x-x0,y-yo)
ConvNet
• ConvNet architectures for images:
– fully-connected structure does not scale to large
images
– the explicit assumption that the inputs are images
– allows us to encode certain properties into the
architecture.
– These then make the forward function more efficient
to implement
– Vastly reduce the amount of parameters in the
network.
• 3D volumes: neurons arranged in 3 dimensions:
width, height, depth.
Convnets
translated
image image
32x32x3 image
32 height
32 width
3 depth
Convolutions: More detail
32x32x3 image
5x5x3 filter
32
32
3
Convolutions: More detail
Convolution Layer
32x32x3 image
5x5x3 filter
32
1 number:
the result of taking a dot product between the
filter and a small 5x5x3 chunk of the image
32 (i.e. 5*5*3 = 75-dimensional dot product + bias)
3
Convolutions: More detail
Convolution Layer
activation map
32x32x3 image
5x5x3 filter
32
28
32 28
3 1
Convolutions: More detail
consider a second, green filter
Convolution Layer
32x32x3 image activation maps
5x5x3 filter
32
28
32 28
3 1
Convolutions: More detail
For example, if we had 6 5x5 filters, we’ll get 6 separate activation maps:
activation maps
32
28
Convolution Layer
32 28
3 6
32 28
CONV,
ReLU
e.g. 6
5x5x3
32 filters 28
3 6
Convolutions: More detail
Preview: ConvNet is a sequence of Convolutional Layers, interspersed with activation
functions
32 28 24
….
CONV, CONV, CONV,
ReLU ReLU ReLU
e.g. 6 e.g. 10
5x5x3 5x5x6
32 filters 28 filters 24
3 6 10
Convolutions: More detail
[From recentYann
Preview LeCun slides]
Convolutions: More detail
one filter =>
one activation map example 5x5 filters
(32 total)
28
32 28
3 1
Convolutions: More detail
A closer look at spatial dimensions:
• 7
• 7x7 input
(spatially)
assume 3x3
filter
• 7
Convolutions: More detail
A closer look at spatial dimensions:
• 7
• 7x7 input
(spatially)
assume 3x3
filter
• 7
Convolutions: More detail
A closer look at spatial dimensions:
• 7
• 7x7 input
(spatially)
assume 3x3
filter
• 7
Convolutions: More detail
A closer look at spatial dimensions:
• 7
• 7x7 input
(spatially)
assume 3x3
filter
• 7
Convolutions: More detail
A closer look at spatial dimensions:
• 7
• 7x7 input (spatially)
assume 3x3 filter
7 => 5x5 output
Convolutions: More detail
A closer look at spatial dimensions:
7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2
7
Convolutions: More detail
A closer look at spatial dimensions:
7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2
7
Convolutions: More detail
A closer look at spatial dimensions:
7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2
=> 3x3 output!
7
Convolutions: More detail
A closer look at spatial dimensions:
7
7x7 input (spatially)
assume 3x3 filter
applied with stride 3?
7
Convolutions: More detail
A closer look at spatial dimensions:
7
7x7 input (spatially)
assume 3x3 filter
applied with stride 3?
7 doesn’t fit!
cannot apply 3x3 filter on
7x7 input with stride 3.
Convolutions: More detail
N
Output size:
(N - F) / stride + 1
F
e.g. N = 7, F = 3:
F N
stride 1 => (7 - 3)/1 + 1 = 5
stride 2 => (7 - 3)/2 + 1 = 3
stride 3 => (7 - 3)/3 + 1 = 2.33 :\
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0
(recall:)
(N - F) / stride + 1
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0
0
7x7 output!
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0
0
7x7 output!
in general, common to see CONV layers with
stride 1, filters of size FxF, and zero-padding with
(F-1)/2. (will preserve size spatially)
e.g. F = 3 => zero pad with 1
F = 5 => zero pad with 2
F = 7 => zero pad with 3
(N + 2*padding - F) / stride + 1
Convolutions: More detail
Examples time:
Max
Sum
3. Spatial Pooling
• Sum or max over non-overlapping / overlapping regions
• Role of pooling:
• Invariance to small transformations
• Larger receptive fields (neurons see more of input)
Pooling Layer
• Insertion of pooling layer:
– reduce the spatial size of the representation
reduce the amount of parameters and computation in the network, and
hence also control overfitting.
• The Pooling Layer operates independently on every depth slice of
the input and resizes it spatially, using the MAX operation.
• The most common form is a pooling layer with filters of size 2x2
applied with a stride of 2 -- downsamples every depth slice in the
input by 2 along both width and height,
• MAX operation would in take a max over 4 numbers (little 2x2
region in some depth slice).
• The depth dimension remains unchanged.
General pooling layer
• Accepts a volume of size W1×H1×D1
• Requires two hyperparameters:
– their spatial extent F
– the stride S
• Produces a volume of size W2×H2×D2 where:
– W2=(W1−F)/S+1
– H2=(H1−F)/S+1
– D2=D1
• Introduces zero parameters
• Other pooling functions: Average pooling, L2-
norm pooling
General pooling
Simonyan, Karen, and Andrew Zisserman. "Very deep convolutional networks for large-scale
image recognition." arXiv preprint arXiv:1409.1556 (2014).
Fully-connected layer
• Neurons in a fully connected layer have full connections to all
activations in the previous layer
• Their activations can hence be computed with a matrix
multiplication followed by a bias offset.
• Converting FC layers to CONV layers
• the only difference between FC and CONV layers is that the
neurons in the CONV layer are connected only to a local
region in the input, and that many of the neurons in a CONV
volume share parameters.
• However, the neurons in both layers still compute dot
products, so their functional form is identical.
Converting FC layers to CONV layers
• For any CONV layer there is an FC layer that implements the same forward
function.
• The weight matrix would be a large matrix that is mostly zero except for at
certain blocks (due to local connectivity) where the weights in many of the
blocks are equal (due to parameter sharing).
• Conversely, any FC layer can be converted to a CONV layer.
• For example, an FC layer with K=4096 that is looking at some input volume
of size 7×7×512
• can be equivalently expressed as a CONV layer with F=7,P=0,S=1,K=4096.
• In other words, we are setting the filter size to be exactly the size of the
input volume, and hence the output will simply be 1×1×4096 since only a
single depth column “fits” across the input volume, giving identical result
as the initial FC layer.
ConvNet Architectures
Layer Patterns
• The most common architecture
• stacks a few CONV-RELU layers,
• follows them with POOL layers,
• and repeats this pattern until the image has been merged spatially
to a small size.
• At some point, it is common to transition to fully-connected layers.
The last fully-connected layer holds the output, such as the class
scores. In other words, the most common ConvNet architecture
follows the pattern:
INPUT -> [[CONV -> RELU]*N -> POOL?]*M ->[FC -> RELU]*K -> FC
• N >= 0 (and usually N <= 3), M >= 0, K >= 0
Prefer a stack of small filter CONV to one large receptive field CONV layer.
three layers of 3x3 CONV vs a single CONV layer with 7x7
receptive fields.
• The receptive field size is identical in spatial extent (7x7), but
with several disadvantages.
1. The neurons would be computing a linear function over the input,
while the three stacks of CONV layers contain non-linearities that
make their features more expressive.
2. If we suppose that all the volumes have C channels, the single 7x7
CONV layer would contain C×(7×7×C)=49C2 parameters, while the
three 3x3 CONV layers would contain 3×(C×(3×3×C))=27C2
parameters.
• Intuitively, stacking CONV layers with tiny filters as opposed to
having one CONV layer with big filters allows us to express
more powerful features of the input, and with fewer
parameters.
Recent Departures
• The conventional paradigm of a linear list of layers
has recently been challenged, in
1. Google’s Inception architectures
2. current (state of the art) Residual Networks from
Microsoft Research Asia.
• Both of these feature more intricate and different
connectivity structures.
DIABETIC RETINOPATHY
LEARNING OBJECTIVES
• Recognize the importance of diabetic retinopathy as a public
health problem
• Discuss diabetic retinopathy as a leading cause of blindness in
developed countries
• Identify the risk factors for diabetic retinopathy
• Describe and distinguish between the stages of diabetic
retinopathy
• Understand the role of risk factor control and annual dilated eye
exams in the prevention of vision loss
DIABETES MELLITUS
Diabetes Mellitus is a group of diseases characterized by high blood glucose
levels. Diabetes results from defects in the body's ability to produce and/or use
insulin.
• Type 1 diabetes is usually diagnosed in children and young adults, and was
previously known as juvenile diabetes. In type 1 diabetes, the body does not
produce insulin. 5% of people with diabetes have this form of the disease.
• In Type 2 diabetes, either the body does not produce enough insulin or the
cells ignore the insulin. This is the most common form of diabetes.
DIABETIC RETINOPATHY (DR)
DEFINITION
• Progressive dysfunction of the retinal blood vessels
caused by chronic hyperglycemia.
• DR can be a complication of diabetes type 1 or
diabetes type 2.
• Initially, DR is asymptomatic, if not treated though it
can cause low vision and blindness.
[Link]
WHAT IS THE RETINA?
• The retina is a multilayered, light sensitive neural tissue
lining the inner eye ball. Light is focused onto the retina
and then transmitted to the brain through the optic
nerve.
• The macula is a highly sensitive area in the center of
the retina, responsible for central vision. The macula is
needed for reading, recognizing faces and executing
other activities that require fine, sharp vision.
RETINA
Healthy Retina Diabetic Retinopathy
DIABETIC RETINOPATHY
EPIDEMIOLOGY
• Duration of diabetes
• Poor Blood Sugar control
• HTN
• Hyperlipidemia
• Barriers to care
[Link]
The Effect of Intensive Diabetes Treatment
On the Progression of Diabetic Retinopathy
In Insulin-Dependent Diabetes Mellitus
[Link]
HOW DIABETES CAUSES VISION LOSS
How diabetes cause vision loss
Macular Clinical
significant
edema
macular edema
Vitreous hemorrhage
Preproliferative Proliferative and/or Retinal
DR DR detachment and/or
neovascular glaucoma
PATHOPHYSIOLOGY
Diabetic Retinopathy is a microvasculopathy that
causes:
• Retinal capillary occlusion
• Retinal capillary leakage
MICROVASCULAR OCCLUSION
Microvascular occlusion is caused by:
• Thickening of capillary basement membranes
• Abnormal proliferation of capillary endothelium
• Increased platelet adhesion
• Increased blood viscosity
• Defective fibrinolysis
Ischemia
Infarction
Increased VEFG
Neovascularization
Vitreous Neovascular
Fibrovascular bands
hemorrhage glaucoma
Tractional retinal
detachment Retina in systemic disease : a color manual of
ophthalmoscopy / Homayoun Tabandeh, Morton F.
Goldberg 2009
MICROVASCULAR LEAKAGE
Microvascular leakage is caused by:
• Impairment of endothelial tight junctions
• Loss of pericytes
• Weakening of capillary walls
• Elevated levels of vascular endothelial growth factor (VEGF)
Retinal
Edema Hard exudates
hemorrhage
.
RECOMMENDED
DiabeticEYE EXAMINATION
Eye Disease
SCHEDULE Key Points
Diabetes Type Recommended Time of Recommended Follow-
First Examination up*
Microaneurysms
MODERATE NONPROLIFERATIVE DIABETIC
RETINOPATHY (NPDR)
Characteristics
• More than just microaneurysms but less than severe NPDR but
less than severe NPD
MODERATE NONPROLIFERATIVE DIABETIC
RETINOPATHY (NPDR)
Microaneurysm
Hard exudates
Flamed shaped
hemorrhage
MODERATE NONPROLIFERATIVE
DIABETIC RETINOPATHY (NPDR)
Hard exudates
microaneurysm
SEVERE NONPROLIFERATIVE
DIABETIC RETINOPATHY (NPDR)
Any of the following:
• More than 20 intraretinal hemorrhages in each of four
quadrants
• Definite venous beading in two or more quadrants
• Prominent Intraretinal Microvascular Abnormalities
(IRMA) in one or more quadrants
• And no signs of proliferative retinopathy
Severe Nonproliferative Diabetic Retinopathy
(NPDR)
Venous beading
Proliferative Diabetic Retinopathy (PDR)
Characteristics
• Neovascularization
• Vitreous/preretinal
hemorrhage
PROLIFERATIVE
DIABETIC
RETINOPATHY Cotton-wool
spot
Neovascularization
Neovascularization
Hard exudate
Blot hemorrhage
HIGH-RISK PROLIFERATIVE DIABETIC
RETINOPATHY
Basic and Clinical Science Course, Section 12: Retina and Vitreous AAO
DIABETIC MACULAR EDEMA
• Diabetic macular edema is the leading cause of legal
blindness in diabetics.
• Diabetic macular edema can be present at any stage of
the disease, but is more common in patients with
proliferative diabetic retinopathy.
Meta analysis and review on the effect on bevacizumab id diabetic macular edema
Graefes Arch Clin Exp Ophthalmol(2011) 249:15-27
Why is Diabetic macular edema so important?
• The macula is responsible for central vision.
• Diabetic macular edema may be asymptomatic at
first. As the edema moves in to the fovea (the center
of the macula) the patient will notice blurry central
vision. The ability to read and recognize faces will be
compromised.
Macula
Fovea
Normal Macular Edema
CLINICALLY SIGNIFICANT MACULAR EDEMA
(CSME)
• Thickening of the retina at or within 500 µm of the
center of the macula.
• Hard exudates at or within 500 µm of the center of the
macula, if associated with thickening of the adjacent
retina.
• Area of retinal thickening 1 disc area or larger, within 1
disc diameter of the center of the macula.
ETDRS
INTERNATIONAL CLINICAL DIABETIC MACULAR EDEMA
DISEASE SEVERITY SCALE
Secondary prevention
Annual eye exams
Tertiary prevention
Retinal Laser photocoagulation
Vitrectomy
DIABETIC RETINOPATHY TREATMENT
Diabetic Retinopathy is
preventable through strict
glycemic control and annual
dilated eye exams by an
ophthalmologist.
The Guerrilla Eye Service of the UPMC Eye Center is dedicated
to eliminating barriers to eye care for patients in the Western
Pennsylvania area.
Self-driving cars
Question
How would you define a self-driving car?
Definition: What is an autonomous car?
● Autonomous Car: A driverless vehicle capable of fulfilling the main
transportation capabilities of a traditional car.
Classifications of Autonomy according to the NHTSA.
● Level 0: The driver completely controls the vehicle at all times.
● Level 1: Individual vehicle controls are automated, such as electronic stability
control or automatic braking.
● Level 2: At least two controls can be automated in unison, such as adaptive
cruise control in combination with lane keeping.
● Level 3: The driver can fully cede control of all safety-critical functions in certain
conditions and the car provides a "sufficiently comfortable transition time" for the
driver to do so.
● Level 4: The vehicle performs all safety-critical functions for the entire trip, with
the driver not expected to control the vehicle at any time.
Purpose
What kinds of things does a self-driving car need to be able to do?
Purpose
● navigate to a given destination based on passenger-provided instructions
● The range finder mounted on the top is a Velodyne 64-beam laser. This laser
allows the vehicle to generate a detailed 3D map of its environment.
● The car uses data collected from these mechanisms to drive itself.
Google’s Technology
How it works: Lidar system
● Laser + radar
● The system detects obstacles and tells the car when to avoid them to
navigate safely.
● It uses a 3D point cloud output provide the necessary data for robot software
to determine where potential obstacles exist in the environment and where
the car is is located relative to those obstacles.
How it works: Velodyne
● Company started experimenting with laser distance in 2005 with the DARPA
Grand Challenge
● Since then, they have vastly reduced the size of the sensor and weight while
improving its performance.
● It is a premier lidar system
How does communication among driverless cars
work?
● vehicles and roadside units as the communicating nodes
○ DSRC devices- 5.9 GHz band with bandwith of 75 MHz- range of 1000m
Communication among driverless cars cont.
● Smart intersections
Bosch Peugeot
Nissan Uber
Renault Google
Toyota Tesla
Mercedes Benz
Audi
Tesla’s Current Auto Pilot
Potential advantages
● being able to get things done while in traffic or on the road
● fewer traffic collisions. Experts estimate 300,000 lives can be saved per
decade
● Software reliability
● Loss of privacy
Legislation
In the United States, state vehicle codes generally do not envisage — but do not
necessarily prohibit — highly automated vehicles.
Public Opinion
What do you think?
● By 2018, Elon Musk expects Tesla Motors to have developed mature serial
production version of fully self-driving cars, where the driver can fall asleep
behind the wheel.
Predictions: Possible Developments
● By 2018, Nissan anticipates to have a feature that can allow the vehicle
maneuver its way on multi-lane highways.
● By 2020, Google autonomous car project head's goal to have all outstanding
problems with the autonomous car be resolved.
SMART SPEAKER
CONSUMERADOPTION
REPORT
MARCH2019
U.S.
G I V I N G V OICE T O A RE V O L U TI O N
Table of Contents AboutVoicebot AboutVoicify
Introduction // 3 Voicebot produces the leading online publication, Voicify is the market leader in voice experience
newsletter and podcast focused on the voice and AI management software that combines voice
Smart Speaker Ownership // 6 industries. Thousands of entrepreneurs, developers, optimized content management, cross-platform
investors, analysts and other industry leaders look deployment, and voice-specific customer insights.
Smart Speaker Use Cases //15 to Voicebot each week for the latest news, data, The Voicify Voice Experience Platform™ enables
analysis and insights defining the trajectory of the marketers to connect with their customers by
Voice Assistants on Smart Phones // 22
next great computing platform. At Voicebot, we give creating highly engaging and personalized voice
Voice App Discovery // 25 voice to a revolution. experiences that are automatically deployed to
a broad array of voice platforms such a s voice
Consumer Sentiment about Smart Speakers // 29 assistants (Amazon Alexa, Google Assistant and
Microsoft Cortana), chatbots and other services.
Conclusion // 32 Methodology The platform enables non-technical users to
deploy feature-rich voice applications quickly and
Additional Resources // 33 The survey was conducted online during the first
efficiently while offering the flexibility of unlimited
week of January 2019 and was completed by 1,038
customization.
U.S. adults age 18 or older that were representative
of U.S. Census demographic averages. Because we
reached only online adults which represent 89% of [Link]
the population according to Pew Research Center,
some totals are adjusted downward to provide
device and usage numbers relevant to the entire
adult population. Other findings are relative to device
ownership and do not require adjustment.
SMART SPEAKER CONSUMER ADOPTIONREPORT
Regardless, when more than one-in-four consumers are using a device and its voice assistant, the
media, brands, game makers, service providers, independent developers, and even governments are
sure to take notice. This recognition is playing out with more voice apps published. The number of
Alexa skills rose by 2.2 times to nearly 60,000 in the U.S. alone in 2018. During the s ame period Google
Actions grew at a slightly faster rate of 2.5 times to over 4,000.
A Different Smart Speaker Ecosystem, but the Same Leaders Smart Speakersare Solidly in the Early Majority Market
Voicebot reported in the fall of 2018 that Phase 1 of smart speaker adoption One way we can put the current state of smart speaker adoption in perspective is
was over and we were entering Phase 2. The second phase is characterized by to consider a standard technology adoption lifecycle first developed in the 1950’s
the influx of more casual users but also by the introduction of new product form at Iowa State University and popularized in the 1990’s by Geoffrey Moore.
factors and new manufacturers.
The model posits that about 16% of the user population will be “innovators” and
The most significant of these changes h as been the emergence of smart “early adopters” followed by 34% that will be among the “early majority.” With more
displays. When Amazon was the only manufacturers of these voice-first devices than 26% population adoption, smart speakers are securely in the “early majority”
with display screens, adoption was minimal. However, the introduction of Google segment today.
Assistant enabled smart displays has helped drive sales, including Amazon, a s it
An interesting aspect of moving along the adoption curve is that later adopters
brought more attention to the product category.
have different preferences than early adopters. Two areas of difference are
There are also many more manufacturers today than in 2017. Big names in audio typically placing higher value in broader feature sets and integrations with other
such a s Bose, Bang & Olufsen, and Klipsch all entered the smart speaker segment devices. You should expect to see smart speaker makers emphasize features,
in 2018 offering more consumer choice. However, the most significant new smart convenience of access, and third-party integrations more in the coming year.
speaker launch in 2018 was Apple HomePod. That appears to have captured
a significant number of new sales in Q1 and Q2, but seems to have tapered off
in Q3 and Q4. Although Apple was threatening to break up the smart speaker 2019
duopoly, it appears that Amazon and Google enter 2019 nearly a s strong a s they
did in 2018 by maintaining 85% in total installed base market share.
Amazon continued to have the leading installed base of smart speakers in 2018 despite its market
share shrinking from about 72% to 61%. Google was a big mover shifting from 18.4% to nearly 24%,
accounting for precisely half of Amazon’s market share decline. U.S. Smart Speaker Market Share by Brand
January 2018 &2019
The “Other” category was led by Apple and Sonos, and overall the non-Amazon, non-Google device
market share rose by 50% over 2017. More than half of this growth is attributed to Apple HomePod
2019
which had a strong debut in the first half of 2018, but then tapered off in sales a s the year went on.
There were several new smart speakers introduced in 2018 and many focused on sound quality. It
61.1% 23.9% 15.0%
appears consumers are open to adding these higher end smart speakers to their device collection a s
Amazon Google Other
over three-quarters of “Other” category smart speaker owners also report having either an Amazon
Echo or Google Home device.
2018
Sonos went public in 2018 and was clear in its investor documents that voice assistant integration
was critical to the company’s future competitiveness. However, the inability to launch a Google
Assistant enabled speaker may have hurt its appeal with consumers a s the company’s overall smart 71.9% 18.4% 9.7%
Amazon Google Other
speaker market share fell during the year. We can surmise that most of the Sonos fans that wanted
an Alexa-based speaker already bought their device in 2017. As the overall market expanded in 2018,
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
few additional Sonos One devices were purchased and the company’s relative market share fell.
Adding Google Assistant support in 2019 may help reverse this market share slide.
Amazon Echo Dot is the most widely adopted smart speaker by a significant
margin. The sub $50 list price device is frequently available for less than $30 and
U.S. Smart Speaker Market Share by Device - January 2019
refurbished models can be acquired for under $20. This device has proven more
popular than Amazon’s higher priced offerings such a s the Echo, Echo Plus, Echo
31.4% 11.2% 10.0%
Spot, and Echo Show.
Amazon Echo Dot Google Home Other
In the Google portfolio, the Home and Home Mini appear to be equally popular
with 11.2% share each. There are likely to be more Home Minis in use today in
terms of total devices a s this analysis reflects the number of users with access to
a device. If you have one Home and three Minis, you are counted a s one in each
category. And, this may be common a s 87% of Google smart speaker owners
report having both devices. 11.2%
23.2% Home Mini
Apple HomePod and Sonos One lead with smart speaker market share in the Echo orPlus
“Other” category. It appears that smart displays with Google Assistant along with
2.7%
the introduction of Apple HomePod in February 2018 were the key drivers leading Apple
to a 50% growth in this category during the year. Keep in mind that aside from HomePo
HomePod, the “Other” category devices all have Alexa or Google Assistant on d
board, so the dominance of Amazon and Google voice assistants extends beyond 3.5% // EchoSpot
2.2%
1.2% Sonos One
their own products. 3.0% Voicebot
Source: // Amazon
Smart EchoShow
Speaker Consumer Adoption Report Jan 2019 0.2% ////Home
HomeHub
Max
U.S. Smart Speaker Frequency of Use 2018 New Smart Speaker Owners are Less
63.6%
Likely to be Daily Users
Maybe the biggest change in the composition of smart speaker owners is the
47.4% influx of more casual users of the devices. Nearly 64% of device owners in
January 2018 reported being daily users. In January 2019, that number fell to only
about 47%. Monthly users were fairly similar with the offsetting difference being
the infrequent users which rose from 13% to over 26%.
26.5% 26.1%
23.5%
This seems like a natural progression. Early
12.9% adopters of technology are more likely to
incorporate them quickly into their daily habits than
consumers that tend to adopt later. However, this
2018 2019 2018 2019 2018 2019 will be a metric to monitor going forward. Three
NEVER ORRARELY MONTHLY DAILY
out of four smart speaker owners still report being
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 monthly active users. As long
a s we see that type of consistent usage along
with continued growth, smart speakers will
continue to grow in importance a s a voice
assistant channel for consumer engagement.
The data indicate that the industry sold about 48 million smart speakers in the U.S. in 2018 bringing the total
in use to about 133 million up from about 85 million at the end of 2017. Of the 19 million new smart speaker
owners, 31% have purchased multiple devices. That compares to 49% of U.S. adults that have owned smart
speakers for more than a year and have multiple devices.
8.0% 14.4%
3 devices
3 devices
65.7% 58.1%
19.3% 1 device 1 device
2 devices
23.2%
2 devices
2018 2019
© [Link] - All Rights Reserved 2019 Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 PAGE 11
SMART SPEAKER CONSUMER ADOPTIONREPORT
For the second straight year, the living room was the Where Consumers Have Smart Speakers
most common location for smart speakers. At just
under 45% it was ahead of the bedroom at 37.6%
2.3% // Garage
which had about the same percentage a s last year 37.6% 14.4%
but moved up from third to the second most popular Bedroom Home Office 32.7%
spot. Third place went to the kitchen. At right around
Kitchen
33%, the kitchen seems to have fallen from favor a
bit among smart speaker owners. It’s still popular, 44.4%
but down from 41% in January 2018. Living Room
Most of the other locations were fairly similar to last
year with the exception of the home office which
grew by about one-third. As consumers have been
adding more smart speakers to their collection,
the home office seems to be a common second
location. 2.0%
6.2% // Bathroom 6.5% // DiningRoom Work
Office
Note: Multiple responses accepted, numbers total more than 100%
© [Link] - All Rights Reserved 2019 Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 PAGE 12
SMART SPEAKER CONSUMER ADOPTIONREPORT
Amazon Prime and Gmail Users More Likely to be Smart Speaker Owners
AMAZON PRIME GMAIL USERS
It will surprise few people that Amazon Prime members are Gmail users are also more likely than all users to own a smart
50% more likely to own a smart speaker and more likely to speaker, in this case by about 31%. However, Gmail users are
own an Echo branded device. Amazon Echo smart speakers no more likely than all users to own a Google Home device.
command a 70% market share among Prime members, but In fact, they are almost exactly representative of all users
surprisingly also adopt Google Home products at almost a when it comes to Amazon, Google, and third-party branded
22% rate. Non-Prime members are more likely to adopt third smart speakers. Whereas a Prime membership and Gmail use
party smart speakers made by manufacturers other than suggests a bias toward early technology adoption, only the
Amazon or Google. Prime membership seems to materially influence consumer
choice of smart speakers.
always. When you look at monthly and daily use, it provides a far more accurate
For the second straight year, asking general questions is the top indication of why consumers are using the devices and in many cases why they
use case most commonly tried by smart speaker owners. may be buying a second or third smart speaker for the home. For example, nearly
However, it is not the top use case employed on a monthly or one-in-four smart speaker owners say they set an alarm to play on their smart
daily basis. That distinction goes to listening to streaming music speaker daily. That would suggest a device location for the bedroom may become
services a s it did in 2018. Third place both years was asking increasingly important.
about the weather which is followed by Timers and Alarms in the
fourth and fifth positions. Number six in 2019 was listening to the The biggest variance is smart home control which is ninth in terms of “ever tried”
radio. and fourth for “daily active use.” You must have a smart home device to use
this feature so that automatically eliminates some people from trial. However,
You may have noticed that four of the top five use cases are what controlling lights or thermostats are already daily functions and if you have smart
are considered first-party services. That means they are provided home devices for these features, then switching your habits from smartphone
by the voice assistant natively. Two of the top six use cases app control to voice interaction is a relatively easy change. What smart speaker
involve music which are third-party entertainment services. and voice assistant developers want to see is consumers using these devices
Positions 7-9 all go to the more traditional third-party services, frequently and incorporating them into daily routines. This not only leads to a
many of which were made by independent developers of Alexa higher perception of value by consumers but also leads to stickiness which
skills and Google Actions. So, the order of use frequency at a means the devices are less likely to be removed or swapped out by consumers for
category level are first-party utilities, third-party entertainment, and a competing product.
third-party apps and services.
Frequency Sometimes
© [Link] - All Rights Reserved 2019 More Important Than Trial PAGE 15
You will notice that many analyses of smart speaker use only
focus on what users have tried. This offers a pretty solid
guide to what users value, but not
SMART SPEAKER CONSUMERADOPTIONREPORT
VOICE COMMERCE
Voice commerce was a mover for a different reason. It had
Monthly Active U.S. Smart Speaker Voice Commerce Users
the lowest frequency of use cases tracked this year. However,
it also showed relative growth in monthly active users during
2018. Monthly active users rose 10.5% from 13.6% to 15.0%.
This is still a relatively new use case that consumers are
becoming accustomed to, but the growth is indicative of the 13.6% 15.0%
utility of shopping by voice. And, this isn’t just users searching 2017 2018
for products. The responses were specific to making purchases.
When it comes to product search, over 40% of users have
attempted this use case on smart speakers and 28% do so
monthly. These are figures that are increasingly difficult for
consumer brands toignore.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019 PAGE 17
SMART SPEAKER CONSUMER ADOPTIONREPORT
Smart home devices are popular among Smart Home Devices Used by U.S. Smart Speaker Owners
smart speaker owners. More than 55%
of smart speaker owners say they have
at least one smart home device that they
control by voice. Of course, being able to
interact by voice with your smart home
devices doesn’t mean you are going to
use it a s about one-in-five consumers 33.3% 21.2% 14.4% 12.4%
with smart home devices have never tried Smart media controller,
SmartTV SmartLights SmartThermostat
controlling them with their smart speaker. game console or cable box
Smart speaker owners are 10% more likely to have used a voice assistant on a Voice Assistant Use Frequency on Smartphones by Smart Speaker Ownership
smartphone. They are also more likely to be daily users. Of all voice assistant
users on smartphones 27.6% report being daily users. Among smart speaker
owners that figure rises to 39.8%. 49.3%
RARELY
One-third of smart speaker owners say after purchasing the device they are using 26.8%
voice assistants on their smartphones more frequently. J u s t over 50% say they
are using smartphone-based voice assistants about the same and only 14% say
A 31.5%
they are using them less. There is a growing consensus that smart speakers are
T
displacing time normally spent on smartphones and many people posit that this L
will help reduce screen time. An accelerant for reducing screen time may be using E
voice assistants on smartphones a s well. This reduces the touch, swipe, and look A
S 1
for many use cases.
T
39.8%
M 33.3%
Smart speaker owners are about a s likely a s non-owners to be monthly voice
O
assistant users on smartphones. However, they are twice a s likely to be daily N Smartphone Voice Assistant Smartphone Voice Assistant
T Use Frequency of Non Smart Use Frequency of Smart
users. Data is consistently showing that usage of voice on one platform increases H Speaker Owners Speaker Owners
usage of voice on others platforms. L
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
Y
A 9.1%
T
L
E
A
S
T
Voice App Discovery
SMART SPEAKER CONSUMER ADOPTIONREPORT
For the second year in a row, just under half of smart speaker owners say they How Smart Speaker Owners Discover Voice Apps
don’t actually discover new voice apps. One-in-four device owners rely on friends I don’t
to introduce them to new voice apps followed by just 15% that note social media 49.7%
a s a discovery channel. Friends
A smaller group of users between 11-14% cite Amazon and Google’s primary 26.8%
promotion channels a s key sources of discovery, such a s their in-app and online Social media
stores and weekly emails. Not far behind these channels is advertising which was 15.4%
less visible in previous years, but now is a source of voice app discovery for one- Alexa skill store / Google Assistant discover section
in-ten smart speaker owners. 13.7%
Email newsletter from Alexa or GoogleAssistant
Discovery is the top issue facing third-party voice app publishers today. The voice
assistant user base is growing quickly, but about half of these users are only 11.1%
discovering first-party solutions provided by the voice assistants themselves Ads / commercials
such a s Alexa and Google Assistant. Many third-parties are having more trouble 10.5%
capturing new users. Word-of-mouth appears to be the most effective channel, Newsmedia
but is the hardest to tap into. So, most voice app publishers should focus on a 7.2%
variety of techniques ranging from news media coverage and social media to Other
advertising to drive discovery today.
2.9%
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019 PAGE 25
SMART SPEAKER CONSUMER ADOPTIONREPORT
Only 14% of smart speaker owners have ever left a review of a third-party voice U.S. Smart Speaker Owners That Have Left a Voice App Review
app. This is up from just 11% at the end of 2017 and is a reflection of the fact that
most reviews must be submitted through a visual interface for a user experience Once
designed for no visual interaction. 6.9% More than once
7.2%
Amazon introduced voice ratings in late 2018 which enabled users to offer a star 85.9%
rating for an Alexa skill by voice after using it. However, this was limited to a few Never
skills and not available for skill publishers to set a s a feature on their own. By
contrast, Google Assistant users are more likely to use the voice assistant both on
smartphones and smart speakers. That multimodal use profile might explain why
Google Home owners are about 11% more likely to have left a review than those
with Amazon Echo devices.
With all of that said, 14% seems like a small number of smart speaker owners
leaving reviews until you consider the fact that only 48.7% say they have even
used a third-party voice app. That means 29% of device owners that have tried a
third-party voice app have left a review. This is a promising figure given the friction
involved in actually leaving a review provided voice app publishers can increase
the proportion of smart speaker owners that try third-party apps.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019
Consumer Sentiment
About Smart Speakers
SMART SPEAKER CONSUMER ADOPTIONREPORT
Amazon, Apple, and Google executives have spoken many times How fast itresponds 45.1%
about their focus on adding personality to voice assistants despite
the fact that it is considered important by only 15.4% of smart Its personality 15.4%
speaker owners. That lower rating m a y be influenced by the fact
that personality is offered by all of the leading voice assistant Whether it has my
favorite mediaentertainment 14.4%
providers, but it is clearly not something having an impact today.
Whether the voice assistant is
the same as my mobile device 10.1%
It is also notable that smartphone ownership only influenced
smart speaker selection for about one-in-ten consumers. Apple I am not interested
in a smart speaker 9.8%
and Google would like that linkage to be higher given their
dominance of smartphone-based voice assistants worldwide.
Whether it has goodgames 4.6%
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019