0% found this document useful (0 votes)
16 views208 pages

Deep Learning

The document provides an overview of Convolutional Neural Networks (CNNs), detailing their history, structure, and functionality in deep learning. It discusses the evolution from traditional machine learning to deep learning, emphasizing the automatic feature extraction capabilities of CNNs and their applications in various tasks such as image processing. Additionally, it covers convolutional layers, pooling methods, and the architecture of CNNs, highlighting the importance of optimization and computational power in training these models.

Uploaded by

mokaifasif
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views208 pages

Deep Learning

The document provides an overview of Convolutional Neural Networks (CNNs), detailing their history, structure, and functionality in deep learning. It discusses the evolution from traditional machine learning to deep learning, emphasizing the automatic feature extraction capabilities of CNNs and their applications in various tasks such as image processing. Additionally, it covers convolutional layers, pooling methods, and the architecture of CNNs, highlighting the importance of optimization and computational power in training these models.

Uploaded by

mokaifasif
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Convolutional Neural Network

History
• In 1995, Yann LeCun
and Yoshua Bengio
introduced the concept
of convolutional neural
networks.
“Deep Learning doesn’t do different things,
it does things differently”
Performance vs Sample Size

Performance

Traditional ML algorithms

Size of Data
Outline
Supervised Learning
Convolutional Neural Network
Sequence Modelling: RNN and its extensions
Unsupervised Learning
Autoencoder
Stacked Denoising Autoencoder
• Unsupervised Learning (+Supervised)
Generative Adversarial Networks
Reinforcement Learning
Deep Reinforcement Learning
Outline

GAN

Produce Poetry like Shakespeare

Output
Generator
Generated
Input

Network
Shakespeare Fake
Poetry Discriminator
Network
Real Real
Shakespeare
Poetry
Supervised Learning
Traditional pattern recognition models work with hand crafted
features and relatively simple trainable classifiers.

Trainable
Extract Hand Output
Classifier
Crafted (e.g. Outdoor
(e.g. SVM,
Features Yes or No)
Random
Limitations Forrest)

Limitations

Very tedious and costly to develop hand crafted features.


The hand-crafted features are usually highly dependents on one
application.
Deep Learning
Deep learning has an inbuilt automatic multi stage feature learning
process that learns rich hierarchical representations (i.e. features).

Low Level Mid Level High Output


Trainable (e.g. outdoor,
Features Features Level
Classifier indoor)
Features
Deep Learning

Low Level Mid Level High


Trainable Output
Input Features Features Level
Features
Classifier

Image
Pixel Edge Texture Motif Part Object
Text
Character Word Word-group Clause Sentence Story

•Each module in Deep Learning transforms its input


representation into a higher-level one, in a way similar to
human cortex.
Let us see how it all
works!
A Simple Neural Network
An Artificial Neural Network is an information processing paradigm
that is inspired by the biological nervous systems, such as the
human brain’s information processing mechanism.

x1 a1(1)

x2 a2(1)
a1(2) Y
x3 a3(1)

x4 a4(1)

Input Hidden Layers Output


A Simple Neural Network
Softmax 1
w1
x1 a1(1) (2)
1 + 𝑒 −𝑤∗𝑎1
w2
x2 a2(1)
w3
a1(2) Y
x3 a3(1)
w4

x4 a4(1)

𝑎1(1) = 𝑓 𝑤1 ∗ 𝑥1 + 𝑤2 ∗ 𝑥2 + 𝑤3 ∗ 𝑥3 + 𝑤4 ∗ 𝑥4

f( ) is activation function: Relu or sigmoid

𝑅𝑒𝑙𝑢: max(0, 𝑥)
𝑎1(1) = 𝑚𝑎𝑥 0, 𝑤1 ∗ 𝑥1 + 𝑤2 ∗ 𝑥2 + 𝑤3 ∗ 𝑥3 + 𝑤4 ∗ 𝑥4
Number of Parameters
Softmax
x1 a1(1)

x2 a2(1)
a1(2) Y
x3 a3(1)

x4 a4(1)

Input Hidden Layers Output

4*4 + 4 +1
If the input is an Image?
x1 a1(1)

x2 a2(1)
a1(2) Y
x3 a3(1)

400 X 400 X 3
a480000(1)

x480000

Input Hidden Layers Output

Number of Parameters
480000*480000 + 480000 +1 = approximately 230 Billion !!!
480000*1000 + 1000 +1 = approximately 480 million !!!
Let us see how convolutional layers
help.
Convolutional Layers
0 1 0
Filter 1 -4 1
0 1 0
1 1 1 1 1 1 0.015686 0.015686 0.011765 0.015686 0.015686 0.015686 0.015686 0.964706 0.988235 0.964706 0.866667 0.031373 0.023529 0.007843
0.007843 0.741176 1 1 0.984314 0.023529 0.019608 0.015686 0.015686 0.015686 0.011765 0.101961 0.972549 1 1 0.996078 0.996078 0.996078 0.058824 0.015686
0.019608 0.513726 1 1 1 0.019608 0.015686 0.015686 0.015686 0.007843 0.011765 1 1 1 0.996078 0.031373 0.015686 0.019608 1 0.011765
0.015686 0.733333 1 1 0.996078 0.019608 0.019608 0.015686 0.015686 0.011765 0.984314 1 1 0.988235 0.027451 0.015686 0.007843 0.007843 1 0.352941
0.015686 0.823529 1 1 0.988235 0.019608 0.019608 0.015686 0.015686 0.019608 1 1 0.980392 0.015686 0.015686 0.015686 0.015686 0.996078 1 0.996078
0.015686 0.913726 1 1 0.996078 0.019608 0.019608 0.019608 0.019608 1 1 0.984314 0.015686 0.015686 0.015686 0.015686 0.952941 1 1 0.992157
0.019608 0.913726 1 1 0.988235 0.019608 0.019608 0.019608 0.039216 0.996078 1 0.015686 0.015686 0.015686 0.015686 0.996078 1 1 1 0.007843
0.019608 0.898039 1 1 0.988235 0.019608 0.015686 0.019608 0.968628 0.996078 0.980392 0.027451 0.015686 0.019608 0.980392 0.972549 1 1 1 0.019608
0.043137 0.905882 1 1 1 0.015686 0.035294 0.968628 1 1 0.023529 1 0.792157 0.996078 1 1 0.980392 0.992157 0.039216 0.023529
1 1 1 1 1 0.992157 0.992157 1 1 0.984314 0.015686 0.015686 0.858824 0.996078 1 0.992157 0.501961 0.019608 0.019608 0.023529
0.996078 0.992157 1 1 1 0.933333 0.003922 0.996078 1 0.988235 1 0.992157 1 1 1 0.988235 1 1 1 1
0.015686 0.74902 1 1 0.984314 0.019608 0.019608 0.031373 0.984314 0.023529 0.015686 0.015686 1 1 1 0 0.003922 0.027451 0.980392 1
0.019608 0.023529 1 1 1 0.019608 0.019608 0.564706 0.894118 0.019608 0.015686 0.015686 1 1 1 0.015686 0.015686 0.015686 0.05098 1
0.015686 0.015686 1 1 1 0.047059 0.019608 0.992157 0.007843 0.011765 0.011765 0.015686 1 1 1 0.015686 0.019608 0.996078 0.023529 0.996078
0.019608 0.015686 0.243137 1 1 0.976471 0.035294 1 0.003922 0.011765 0.011765 0.015686 1 1 1 0.988235 0.988235 1 0.003922 0.015686
0.019608 0.019608 0.027451 1 1 0.992157 0.223529 0.662745 0.011765 0.011765 0.011765 0.015686 1 1 1 0.015686 0.023529 0.996078 0.011765 0.011765
0.015686 0.015686 0.011765 1 1 1 1 0.035294 0.011765 0.011765 0.011765 0.015686 1 1 1 0.015686 0.015686 0.964706 0.003922 0.996078
0.007843 0.019608 0.011765 0.054902 1 1 0.988235 0.007843 0.011765 0.011765 0.015686 0.011765 1 1 1 0.015686 0.015686 0.015686 0.023529 1
0.007843 0.007843 0.015686 0.015686 0.960784 1 0.490196 0.015686 0.015686 0.015686 0.007843 0.027451 1 1 1 0.011765 0.011765 0.043137 1 1
0.023529 0.003922 0.007843 0.023529 0.980392 0.976471 0.039216 0.019608 0.007843 0.019608 0.015686 1 1 1 1 1 1 1 1 1

Input Image Convoluted Image

 Inspired by the neurophysiological experiments conducted by Hubel and Wiesel 1962.


Convolutional Layers
What is Convolution?
ℎ1 = 𝑓 𝑎 ∗ 𝑤1 + 𝑏 ∗ 𝑤2 + 𝑒 ∗ 𝑤3 + 𝑓 ∗ 𝑤4

a b c d w1 w2
h1 h2
e f g h w3 w4
i j k l
m n o p

Input Image Filter Convolved Image


(Feature Map)

ℎ2 = 𝑓 𝑏 ∗ 𝑤1 + 𝑐 ∗ 𝑤2 + 𝑓 ∗ 𝑤3 + 𝑔 ∗ 𝑤4

Number of Parameters for one feature map = 4


Number of Parameters for 100 feature map = 4*100
Lower Level to More Complex Features

w1 w2

w3 w4
w5 w6

w7 w8
Filter 1
Filter 2

Input Image
Layer 1 Layer 2
Feature Map Feature Map
 In Convolutional neural networks, hidden units are only connected to local receptive field.
Pooling
Max pooling: reports the maximum output within a rectangular
neighborhood.
Average pooling: reports the average output of a rectangular
neighborhood.

MaxPool with 2X2 filter with


1 3 5 3 stride of 2
4 2 3 1
4 5
3 1 1 3
3 4
0 1 0 4

Input Matrix Output Matrix


Convolutional Neural Network
Maxpool
Output
Feature Extraction Architecture Vector

Living Room

Bed Room
128

256
256

512
512

512
512
256
128

512

512
Kitchen
64
64

Bathroom

Outdoor
Max Pool
Filter

Fully Connected
Layers
Convolutional Neural Networks
Output: Binary, Multinomial, Continuous, Count
Input: fixed size, can use padding to make all images same
size.
Architecture: Choice is ad hoc
requires experimentation.
Optimization: Backward propagation
hyper parameters for very deep model can be estimated properly only if you
have billions of images.
Use an architecture and trained hyper parameters from other papers
(Imagenet or Microsoft/Google APIs etc)
Computing Power: Buy a GPU!!
Automatic Colorization of Black and White Images
Optimizing Images

Post Processing Feature Optimization


(Color Curves and Details)

Post Processing Feature Optimization Post Processing Feature Optimization


(Illumination) (Color Tone: Warmness)
Convolution
1D (continuous, discrete) : Input

 Kernel
f * g (x)   f ( )g(x  )d
 
N 1 Output is
  f ( )g(x   ) sometimes called
 0 Feature map
2D (continuous, discrete) :
 
f * g (x, y)    f ( ,  )g(x   , y   )dd
   
N 1 N 1
   f ( ,  )g(x  , y   )
 0  0
Convolution Properties
• Commutative:
f*g = g*f
• Associative:
(f*g)*h = f*(g*h)
• Homogeneous:
f*(g)=  f*g
• Additive (Distributive):
f*(g+h)= f*g+f*h
• Shift-Invariant
f*g(x-x0,y-yo)= (f*g) (x-x0,y-yo)
ConvNet
• ConvNet architectures for images:
– fully-connected structure does not scale to large
images
– the explicit assumption that the inputs are images
– allows us to encode certain properties into the
architecture.
– These then make the forward function more efficient
to implement
– Vastly reduce the amount of parameters in the
network.
• 3D volumes: neurons arranged in 3 dimensions:
width, height, depth.
Convnets

Layers used to build ConvNets:


• a stacked sequence of
layers. 3 main types
• Convolutional Layer,
Pooling Layer, and Fully-
Connected Layer • every layer of a ConvNet
transforms one volume of
activations to another through a
differentiable function.
The replicated feature approach
• Use many different copies of the
same feature detector with The red connections all
different positions. have the same weight.
– Could also replicate across scale and
orientation (tricky and expensive)
– Replication greatly reduces the
number of free parameters to be
learned.
• Use several different feature types,
each with its own map of
replicated detectors.
– Allows each patch of image to be
represented in several ways.
Backpropagation with weight constraints

• It’s easy to modify the To constrain : w1  w2


backpropagation algorithm to we need : w1  w2
incorporate linear constraints
between the weights. E E
compute: and
• We compute the gradients as w1 w2
usual, and then modify the
gradients so that they satisfy E E
use  for w1and w2
the constraints. w1 w2
– So if the weights started off
satisfying the constraints, they
will continue to satisfy them.
What does replicating the feature detectors achieve?
• Equivariant activities: Replicated features do not make the
neural activities invariant to translation. The activities are
equivariant.
representation translated
by active representation
neurons

translated
image image

• Invariant knowledge: If a feature is useful in some locations


during training, detectors for that feature will be available in
all locations during testing.
Pooling the outputs of replicated feature
detectors
• Get a small amount of translational invariance at each level
by averaging four neighboring replicated detectors to give a
single output to the next level.
– This reduces the number of inputs to the next layer of feature
extraction, thus allowing us to have many more different feature
maps.
– Taking the maximum of the four works slightly better.
• Problem: After several levels of pooling, we have lost
information about the precise positions of things.
– This makes it impossible to use the precise spatial relationships
between high-level parts for recognition.
Example Architecture for CIFAR-10
• [INPUT - CONV - RELU - POOL - FC]
• INPUT [32x32x3] : the raw pixel values of the image
• CONV will compute the output of neurons that are connected to
local regions in the input. With 12 filters, the output volume is
[32x32x12]
• RELU : apply an elementwise activation function, such as the
max(0,x)
• POOL will perform a downsampling operation along the spatial
dimensions (width, height), resulting in volume such as [16x16x12].
• FC layer will compute the class scores, resulting in volume of size
[1x1x10], where each of the 10 numbers correspond to a class
score, such as among the 10 categories of CIFAR-10
Convolution Layer
• The Conv layer is the core building block of a CNN
• The parameters consist of a set of learnable filters.
• Every filter is small spatially (width and height), but extends through the
full depth of the input volume, eg, 5x5x3
• During the forward pass, we slide (convolve) each filter across the width
and height of the input volume and compute dot products between the
entries of the filter and the input at any position.
• produce a 2-dimensional activation map that gives the responses of that
filter at every spatial position.
• Intuitively, the network will learn filters that activate when they see some
type of visual feature
• A set of filters in each CONV layer
– each of them will produce a separate 2-dimensional activation map
– We will stack these activation maps along the depth dimension and produce
the output volume.
Convolutional Neural Network 2
Convolution
Convolutions: More detail

32x32x3 image

32 height

32 width
3 depth
Convolutions: More detail
32x32x3 image

5x5x3 filter
32

Convolve the filter with the image


i.e. “slide over the image spatially,
computing dot products”

32
3
Convolutions: More detail
Convolution Layer
32x32x3 image
5x5x3 filter
32

1 number:
the result of taking a dot product between the
filter and a small 5x5x3 chunk of the image
32 (i.e. 5*5*3 = 75-dimensional dot product + bias)
3
Convolutions: More detail
Convolution Layer
activation map
32x32x3 image
5x5x3 filter
32

28

convolve (slide) over all


spatial locations

32 28
3 1
Convolutions: More detail
consider a second, green filter
Convolution Layer
32x32x3 image activation maps
5x5x3 filter
32

28

convolve (slide) over all


spatial locations

32 28
3 1
Convolutions: More detail
For example, if we had 6 5x5 filters, we’ll get 6 separate activation maps:
activation maps

32

28

Convolution Layer

32 28
3 6

We stack these up to get a “new image” of size 28x28x6!


Convolutions: More detail
Preview: ConvNet is a sequence of Convolution Layers, interspersed with
activation functions

32 28

CONV,
ReLU
e.g. 6
5x5x3
32 filters 28
3 6
Convolutions: More detail
Preview: ConvNet is a sequence of Convolutional Layers, interspersed with activation
functions

32 28 24

….
CONV, CONV, CONV,
ReLU ReLU ReLU
e.g. 6 e.g. 10
5x5x3 5x5x6
32 filters 28 filters 24
3 6 10
Convolutions: More detail
[From recentYann
Preview LeCun slides]
Convolutions: More detail
one filter =>
one activation map example 5x5 filters
(32 total)

We call the layer convolutional


because it is related to convolution
of two signals:

elementwise multiplication and sum of


a filter and the signal (image)
Convolutions: More detail
A closer look at spatial dimensions:
activation map
32x32x3 image
5x5x3 filter
32

28

convolve (slide) over all


spatial locations

32 28
3 1
Convolutions: More detail
A closer look at spatial dimensions:

• 7
• 7x7 input
(spatially)
assume 3x3
filter

• 7
Convolutions: More detail
A closer look at spatial dimensions:

• 7
• 7x7 input
(spatially)
assume 3x3
filter

• 7
Convolutions: More detail
A closer look at spatial dimensions:

• 7
• 7x7 input
(spatially)
assume 3x3
filter

• 7
Convolutions: More detail
A closer look at spatial dimensions:

• 7
• 7x7 input
(spatially)
assume 3x3
filter

• 7
Convolutions: More detail
A closer look at spatial dimensions:

• 7
• 7x7 input (spatially)
assume 3x3 filter
7 => 5x5 output
Convolutions: More detail
A closer look at spatial dimensions:

7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2

7
Convolutions: More detail
A closer look at spatial dimensions:

7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2

7
Convolutions: More detail
A closer look at spatial dimensions:

7
7x7 input (spatially)
assume 3x3 filter
applied with stride 2
=> 3x3 output!
7
Convolutions: More detail
A closer look at spatial dimensions:

7
7x7 input (spatially)
assume 3x3 filter
applied with stride 3?

7
Convolutions: More detail
A closer look at spatial dimensions:

7
7x7 input (spatially)
assume 3x3 filter
applied with stride 3?

7 doesn’t fit!
cannot apply 3x3 filter on
7x7 input with stride 3.
Convolutions: More detail
N
Output size:
(N - F) / stride + 1
F
e.g. N = 7, F = 3:
F N
stride 1 => (7 - 3)/1 + 1 = 5
stride 2 => (7 - 3)/2 + 1 = 3
stride 3 => (7 - 3)/3 + 1 = 2.33 :\
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0

(recall:)
(N - F) / stride + 1
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0

0
7x7 output!
Convolutions: More detail
In practice: Common to zero pad the border
0 0 0 0 0 0
e.g. input 7x7
0
3x3 filter, applied with stride 1
0 pad with 1 pixel border => what is the output?
0

0
7x7 output!
in general, common to see CONV layers with
stride 1, filters of size FxF, and zero-padding with
(F-1)/2. (will preserve size spatially)
e.g. F = 3 => zero pad with 1
F = 5 => zero pad with 2
F = 7 => zero pad with 3

(N + 2*padding - F) / stride + 1
Convolutions: More detail
Examples time:

Input volume: 32x32x3


10 5x5 filters with stride 1, pad 2

Output volume size: ?


Convolutions: More detail
Examples time:

Input volume: 32x32x3


10 5x5 filters with stride 1, pad 2

Output volume size:


(32+2*2-5)/1+1 = 32 spatially, so
32x32x10
Convolutions: More detail
Examples time:

Input volume: 32x32x3


10 5x5 filters with stride 1, pad 2

Number of parameters in this layer?


Convolutions: More detail
Examples time:

Input volume: 32x32x3


10 5x5 filters with stride 1, pad 2

Number of parameters in this layer?


each filter has 5*5*3 + 1 = 76 params (+1 for bias)
=> 76*10 = 760
Convolutions: More detail
Spatial arrangement
• Three hyperparameters control the size of the
output volume
– Depth: no of filters, each learning to look for
something different in the input.
– the stride with which we slide the filter.
– pad the input volume with zeros around the
border.
Spatial arrangement
• We compute the spatial size of the output
volume as a function of
– the input volume size (W)
– the receptive field size of the Conv Layer neurons (F)
– the stride with which they are applied (S)
– the amount of zero padding used (P) on the border.
• The number of neurons that “fit” is given by
(W−F+2P)/(S+1)
– For a 7x7 input and a 3x3 filter with stride 1 and pad 0
we would get a 5x5 output.
– With stride 2 we would get a 3x3 output.
• one spatial dimension (x-axis), one neuron with a receptive field
size of F = 3, the input size is W = 5, and zero padding of P = 1
• Stride = 1, 2

• The Krizhevsky et al. architecture that won the ImageNet 2012


• images of size [227x227x3].
• the first Convolutional Layer, used neurons with receptive field size F=11,
stride S=4, no zero padding P=0
• Since (227 - 11)/4 + 1 = 55, the Conv layer had a depth of K=96,
• the Conv layer output volume had size [55x55x96].
• Each of the 55*55*96 neurons in this volume was connected to a region of
size [11x11x3] in the input volume.
• Moreover, all 96 neurons in each depth column are connected to the same
[11x11x3] region of the input,
Parameter Sharing
• Parameter sharing controls the number of parameters.
• If there are 55*55*96 = 290,400 neurons in the first Conv Layer, and
each has 11*11*3 = 363 weights and 1 bias. Together, this adds up
to 290400 * 364 = 105,705,600 parameters on the first layer of the
ConvNet alone.
• Reduce by parameter sharing
• now have only 96 unique set of weights (one for each depth slice),
for a total of 96*11*11*3 = 34,848 unique weights, or 34,944
parameters (+96 biases)
• During backpropagation, every neuron in the volume will compute
the gradient for its weights, but these gradients will be added up
across each depth slice and only update a single set of weights per
slice.
• Example filters learned by Krizhevsky.
• 96 filters each of size [11x11x3], each is
shared by the 55*55 neurons in one depth
slice.
Summary of Conv Layer
• Accepts a volume of size W1×H1×D1
• Requires four hyperparameters:
– Number of filters K
– their spatial extent F
– the stride S
– the amount of zero padding P
• Produces a volume of size W2×H2×D2
– W2=(W1−F+2P)/S+1
– H2=(H1−F+2P)/S+1
– D2=K
• With parameter sharing, it introduces F⋅F⋅D1 weights per filter, for a total
of (F⋅F⋅D1)⋅K weights and K biases.
• In the output volume, the d-th depth slice (of size W2×H2) is the result of
performing a valid convolution of the d-th filter over the input volume
with a stride of S, and then offset by d-th bias.
Spatial Pooling
• Sum or max over non-overlapping / overlapping regions
• Role of pooling:
• Invariance to small transformations
• Larger receptive fields (neurons see more of input)

Max

Sum
3. Spatial Pooling
• Sum or max over non-overlapping / overlapping regions
• Role of pooling:
• Invariance to small transformations
• Larger receptive fields (neurons see more of input)
Pooling Layer
• Insertion of pooling layer:
– reduce the spatial size of the representation
reduce the amount of parameters and computation in the network, and
hence also control overfitting.
• The Pooling Layer operates independently on every depth slice of
the input and resizes it spatially, using the MAX operation.
• The most common form is a pooling layer with filters of size 2x2
applied with a stride of 2 -- downsamples every depth slice in the
input by 2 along both width and height,
• MAX operation would in take a max over 4 numbers (little 2x2
region in some depth slice).
• The depth dimension remains unchanged.
General pooling layer
• Accepts a volume of size W1×H1×D1
• Requires two hyperparameters:
– their spatial extent F
– the stride S
• Produces a volume of size W2×H2×D2 where:
– W2=(W1−F)/S+1
– H2=(H1−F)/S+1
– D2=D1
• Introduces zero parameters
• Other pooling functions: Average pooling, L2-
norm pooling
General pooling

• Backpropagation. the backward pass for a max(x, y) operation


routes the gradient to the input that had the highest value in
the forward pass.
• Hence, during the forward pass of a pooling layer you may
keep track of the index of the max activation (sometimes also
called the switches) so that gradient routing is efficient during
backpropagation.
Getting rid of pooling
1. Striving for Simplicity: The All Convolutional Net proposes to
discard the pooling layer and have an architecture that only
consists of repeated CONV layers.
• To reduce the size of the representation they suggest using larger
stride in CONV layer once in a while.
• Argument:
– The purpose of pooling layers is to perform dimensionality reduction to
widen subsequent convolutional layers' receptive fields.
– The same effect can be achieved by using a convolutional layer: using a
stride of 2 also reduces the dimensionality of the output and widens the
receptive field of higher layers.
• The resulting operation differs from a max-pooling layer in that
– it cannot perform a true max operation
– it allows pooling across input channels.

Springenberg, Jost Tobias, et al. "Striving for simplicity: The all


convolutional net." arXiv preprint arXiv:1412.6806 (2014).
Getting Rod of Pooling 2
2. Very Deep Convolutional Networks for Large-Scale Image
Recognition.
• The core idea here is that hand-tuning layer kernel sizes to achieve
optimal receptive fields (say, 5×5 or 7×7) can be replaced by simply
stacking homogenous 3×3 layers.
• The same effect of widening the receptive field is then achieved by
layer composition rather than increasing the kernel size
– three stacked 3×3 have a 7×7 receptive field.
– At the same time, the number of parameters is reduced:
– a 7×7 layer has 81% more parameters than three stacked 3×3 layers.

Simonyan, Karen, and Andrew Zisserman. "Very deep convolutional networks for large-scale
image recognition." arXiv preprint arXiv:1409.1556 (2014).
Fully-connected layer
• Neurons in a fully connected layer have full connections to all
activations in the previous layer
• Their activations can hence be computed with a matrix
multiplication followed by a bias offset.
• Converting FC layers to CONV layers
• the only difference between FC and CONV layers is that the
neurons in the CONV layer are connected only to a local
region in the input, and that many of the neurons in a CONV
volume share parameters.
• However, the neurons in both layers still compute dot
products, so their functional form is identical.
Converting FC layers to CONV layers
• For any CONV layer there is an FC layer that implements the same forward
function.
• The weight matrix would be a large matrix that is mostly zero except for at
certain blocks (due to local connectivity) where the weights in many of the
blocks are equal (due to parameter sharing).
• Conversely, any FC layer can be converted to a CONV layer.
• For example, an FC layer with K=4096 that is looking at some input volume
of size 7×7×512
• can be equivalently expressed as a CONV layer with F=7,P=0,S=1,K=4096.
• In other words, we are setting the filter size to be exactly the size of the
input volume, and hence the output will simply be 1×1×4096 since only a
single depth column “fits” across the input volume, giving identical result
as the initial FC layer.
ConvNet Architectures
Layer Patterns
• The most common architecture
• stacks a few CONV-RELU layers,
• follows them with POOL layers,
• and repeats this pattern until the image has been merged spatially
to a small size.
• At some point, it is common to transition to fully-connected layers.
The last fully-connected layer holds the output, such as the class
scores. In other words, the most common ConvNet architecture
follows the pattern:
INPUT -> [[CONV -> RELU]*N -> POOL?]*M ->[FC -> RELU]*K -> FC
• N >= 0 (and usually N <= 3), M >= 0, K >= 0
Prefer a stack of small filter CONV to one large receptive field CONV layer.
three layers of 3x3 CONV vs a single CONV layer with 7x7
receptive fields.
• The receptive field size is identical in spatial extent (7x7), but
with several disadvantages.
1. The neurons would be computing a linear function over the input,
while the three stacks of CONV layers contain non-linearities that
make their features more expressive.
2. If we suppose that all the volumes have C channels, the single 7x7
CONV layer would contain C×(7×7×C)=49C2 parameters, while the
three 3x3 CONV layers would contain 3×(C×(3×3×C))=27C2
parameters.
• Intuitively, stacking CONV layers with tiny filters as opposed to
having one CONV layer with big filters allows us to express
more powerful features of the input, and with fewer
parameters.
Recent Departures
• The conventional paradigm of a linear list of layers
has recently been challenged, in
1. Google’s Inception architectures
2. current (state of the art) Residual Networks from
Microsoft Research Asia.
• Both of these feature more intricate and different
connectivity structures.
DIABETIC RETINOPATHY
LEARNING OBJECTIVES
• Recognize the importance of diabetic retinopathy as a public
health problem
• Discuss diabetic retinopathy as a leading cause of blindness in
developed countries
• Identify the risk factors for diabetic retinopathy
• Describe and distinguish between the stages of diabetic
retinopathy
• Understand the role of risk factor control and annual dilated eye
exams in the prevention of vision loss
DIABETES MELLITUS
Diabetes Mellitus is a group of diseases characterized by high blood glucose
levels. Diabetes results from defects in the body's ability to produce and/or use
insulin.
• Type 1 diabetes is usually diagnosed in children and young adults, and was
previously known as juvenile diabetes. In type 1 diabetes, the body does not
produce insulin. 5% of people with diabetes have this form of the disease.
• In Type 2 diabetes, either the body does not produce enough insulin or the
cells ignore the insulin. This is the most common form of diabetes.
DIABETIC RETINOPATHY (DR)
DEFINITION
• Progressive dysfunction of the retinal blood vessels
caused by chronic hyperglycemia.
• DR can be a complication of diabetes type 1 or
diabetes type 2.
• Initially, DR is asymptomatic, if not treated though it
can cause low vision and blindness.

[Link]
WHAT IS THE RETINA?
• The retina is a multilayered, light sensitive neural tissue
lining the inner eye ball. Light is focused onto the retina
and then transmitted to the brain through the optic
nerve.
• The macula is a highly sensitive area in the center of
the retina, responsible for central vision. The macula is
needed for reading, recognizing faces and executing
other activities that require fine, sharp vision.
RETINA
Healthy Retina Diabetic Retinopathy
DIABETIC RETINOPATHY
EPIDEMIOLOGY

• The total number of people with diabetes


is projected to rise from 285 million in
2010 to 439 million in 2030.
• Diabetic retinopathy is responsible for
1.8 million of the 37 million cases of
blindness throughout the world .
• Diabetic retinopathy (DR) is the leading
cause of blindness in people of working
age in industrialized countries.
Causes of global blindness in millions of people
(WHO 2002)
20
18
16
14
12
10
8
6
4
2
0
DIABETIC RETINOPATHY
EPIDEMIOLOGY

• The best predictor of diabetic retinopathy is the duration


of the disease
• After 20 years of diabetes, nearly 99% of patients with
type 1 diabetes and 60% with type 2 have some degree
on diabetic retinopathy
• 33% of patients with diabetes have signs of diabetic
retinopathy
• People with diabetes are 25 times more likely to become
blind than the general population.
PREVALENCE OF DIABETIC RETINOPATHY AFTER
20 YEARS OF DIAGNOSIS
[Link]
DIABETIC RETINOPATHY SYMPTOMS
Diabetic retinopathy is asymptomatic in early stages of the disease
As the disease progresses symptoms may include
• Blurred vision
• Floaters
• Fluctuating vision
• Distorted vision
• Dark areas in the vision
• Poor night vision
• Impaired color vision
• Partial or total loss of vision
Risk factors

• Duration of diabetes
• Poor Blood Sugar control
• HTN
• Hyperlipidemia
• Barriers to care

[Link]
The Effect of Intensive Diabetes Treatment
On the Progression of Diabetic Retinopathy
In Insulin-Dependent Diabetes Mellitus

The Diabetes Control and Complications Trial

The Diabetes Control and Complications Trial Research Group

Intensive control reduced the risk of developing retinopathy by 76%


and slowed progression of retinopathy by 54%; intensive control
also reduced the risk of clinical neuropathy by 60% and albuminuria
by 54%.
RISK FACTORS DIABETIC RETINOPATHY

Duration of diabetes is a major risk


factor associated with the development
of diabetic retinopathy

The severity of hyperglycemia is the


key alterable risk factor associated with
the development of diabetic retinopathy

[Link]
HOW DIABETES CAUSES VISION LOSS
How diabetes cause vision loss

Macular Clinical
significant
edema
macular edema

Preclinical Background Vision


Diabetes changes DR loss

Vitreous hemorrhage
Preproliferative Proliferative and/or Retinal
DR DR detachment and/or
neovascular glaucoma
PATHOPHYSIOLOGY
Diabetic Retinopathy is a microvasculopathy that
causes:
• Retinal capillary occlusion
• Retinal capillary leakage
MICROVASCULAR OCCLUSION
Microvascular occlusion is caused by:
• Thickening of capillary basement membranes
• Abnormal proliferation of capillary endothelium
• Increased platelet adhesion
• Increased blood viscosity
• Defective fibrinolysis

Retina in systemic disease : a color manual of ophthalmoscopy / Homayoun


Tabandeh, Morton F. Goldberg 2009
Microvascular
Occlusion

Ischemia

Infarction

Increased VEFG

Cotton – wool spot

Neovascularization

Vitreous Neovascular
Fibrovascular bands
hemorrhage glaucoma

Tractional retinal
detachment Retina in systemic disease : a color manual of
ophthalmoscopy / Homayoun Tabandeh, Morton F.
Goldberg 2009
MICROVASCULAR LEAKAGE
Microvascular leakage is caused by:
• Impairment of endothelial tight junctions
• Loss of pericytes
• Weakening of capillary walls
• Elevated levels of vascular endothelial growth factor (VEGF)

Retina in systemic disease : a color manual of ophthalmoscopy / Homayoun Tabandeh,


Morton F. Goldberg 2009
Microvascular Leakage

Retinal
Edema Hard exudates
hemorrhage

.
RECOMMENDED
DiabeticEYE EXAMINATION
Eye Disease
SCHEDULE Key Points
Diabetes Type Recommended Time of Recommended Follow-
First Examination up*

Type 1 3-5 years after Yearly


diagnosis

Type 2 At time of diagnosis Yearly

Prior to pregnancy • Treatments


Prior exist but workNobest
to conception retinopathy to mild
(type 1 or type 2) and early in the first moderate NPDR every
before vision
trimesteris lost 3-12 months
Severe NPDR or worse
every 1-3 months.

*Abnormal findings may dictate more frequent follow-up examinations


h ttp://[Link]/CE/PracticeGuidelines/PPP_Content.aspx?cid=d0c853d3-219f-487b-a524-326ab3cecd9a
International Clinical Diabetic Retinopathy Disease Severity
Scale
Findings Observable upon Dilated
Proposed Disease Severity Level Ophthalmoscopy
Findings Obsd
No apparent retinopathy No abnormalities

Mild nonproliferative diabetic retinopathy Microaneurysms only

More than just microaneurysms but less than severe NPDR


Moderate nonproliferative diabetic retinopathy

Any of the following:


Severe nonproliferative diabetic retinopathy More than 20 intraretinal hemorrhages in each of four
quadrants
Definite venous beading in two or more quadrants
Prominent IRMA in one or more quadrants
and no signs of proliferative retinopathy.

One or both of the following:


Proliferative diabetic retinopathy Neovascularization
Vitreous/preretinal hemorrhage
No retinopathy
MILD NONPROLIFERATIVE
DIABETIC RETINOPATHY
Characteristics
• Microaneurysms only
MILD NONPROLIFERATIVE DIABETIC
RETINOPATHY

Microaneurysms
MODERATE NONPROLIFERATIVE DIABETIC
RETINOPATHY (NPDR)

Characteristics
• More than just microaneurysms but less than severe NPDR but
less than severe NPD
MODERATE NONPROLIFERATIVE DIABETIC
RETINOPATHY (NPDR)

Microaneurysm

Hard exudates

Flamed shaped
hemorrhage
MODERATE NONPROLIFERATIVE
DIABETIC RETINOPATHY (NPDR)

Hard exudates

microaneurysm
SEVERE NONPROLIFERATIVE
DIABETIC RETINOPATHY (NPDR)
Any of the following:
• More than 20 intraretinal hemorrhages in each of four
quadrants
• Definite venous beading in two or more quadrants
• Prominent Intraretinal Microvascular Abnormalities
(IRMA) in one or more quadrants
• And no signs of proliferative retinopathy
Severe Nonproliferative Diabetic Retinopathy
(NPDR)

Venous beading
Proliferative Diabetic Retinopathy (PDR)

Characteristics
• Neovascularization
• Vitreous/preretinal
hemorrhage
PROLIFERATIVE
DIABETIC
RETINOPATHY Cotton-wool
spot

Neovascularization

Neovascularization
Hard exudate
Blot hemorrhage
HIGH-RISK PROLIFERATIVE DIABETIC
RETINOPATHY

At risk for serious vision loss


Any combination of three of the following four findings
• Presence of vitreous or preretinal hemorrhage.
• Presence of new vessels (neovascularization, NV)
• Location of NV on or near the optic disc.
• Moderate to severe extent of new vessels.

Basic and Clinical Science Course, Section 12: Retina and Vitreous AAO
DIABETIC MACULAR EDEMA
• Diabetic macular edema is the leading cause of legal
blindness in diabetics.
• Diabetic macular edema can be present at any stage of
the disease, but is more common in patients with
proliferative diabetic retinopathy.
Meta analysis and review on the effect on bevacizumab id diabetic macular edema
Graefes Arch Clin Exp Ophthalmol(2011) 249:15-27
Why is Diabetic macular edema so important?
• The macula is responsible for central vision.
• Diabetic macular edema may be asymptomatic at
first. As the edema moves in to the fovea (the center
of the macula) the patient will notice blurry central
vision. The ability to read and recognize faces will be
compromised.

Macula
Fovea
Normal Macular Edema
CLINICALLY SIGNIFICANT MACULAR EDEMA
(CSME)
• Thickening of the retina at or within 500 µm of the
center of the macula.
• Hard exudates at or within 500 µm of the center of the
macula, if associated with thickening of the adjacent
retina.
• Area of retinal thickening 1 disc area or larger, within 1
disc diameter of the center of the macula.

ETDRS
INTERNATIONAL CLINICAL DIABETIC MACULAR EDEMA
DISEASE SEVERITY SCALE

Proposed disease severity level Findings observable upon dilated


ophthalmoscopy
DME apparently absent No apparent retinal thickening or hard exudates in
posterior pole

DME apparently present Some apparent retinal thickening or hard exudates in


posterior pole

DME present Mild DME (some retinal thickening or hard exudates in


posterior pole but distant from the center of the
macula)

Moderate DME (retinal thickening or hard


exudates approaching the center of the macula but not
involving the center)

Severe DME (retinal thickening or hard exudates


Proposed International Clinical Diabetic involving the center of the macula)
Retinopathy and Diabetic Macular Edema
Disease Severity Scales
Ophthalmology Volume 110, Number 9, September 2003
Imaging of macular edema with optical
coherence tomography
PREVENTION
90 percent of diabetic eye disease can
be prevented simply by proper regular
examinations, treatment and by
controlling blood sugar.
Primary prevention
Strict glycemic control
Blood pressure control

Secondary prevention
Annual eye exams

Tertiary prevention
Retinal Laser photocoagulation
Vitrectomy
DIABETIC RETINOPATHY TREATMENT

The best measure for prevention of


loss of vision from diabetic
retinopathy is strict glycemic control
LASER PHOTOCOAGULATION
Laser Photocoagulation is recommended for eyes with:
• Clinical significant macular edema CSME
• High risk Proliferative diabetic retinopathy
DIABETIC RETINOPATHY TREATMENT
ONCE DR THREATENS VISION TREATMENTS CAN INCLUDE:

Laser therapy to seal leaking blood vessels


(focal laser)

Laser therapy to reduce retinal oxygen


demand (scatter laser)

Surgical removal of blood from the eye


(vitrectomy)
DIABETIC RETINOPATHY TREATMENT
NEWER DEVELOPMENTS:

The use of anti-vascular endothelial growth


factor antibodies has been shown to be
useful in the treatment of DR

Anti-VEGF antibody treatment appears to


be useful for both macular edema and
proliferative retinopathy

Studies to determine the exact role of anti-


VEGF treatment in relation to laser
treatment in specific situations are
underway.
CONCLUSIONS

Diabetic Retinopathy is
preventable through strict
glycemic control and annual
dilated eye exams by an
ophthalmologist.
The Guerrilla Eye Service of the UPMC Eye Center is dedicated
to eliminating barriers to eye care for patients in the Western
Pennsylvania area.
Self-driving cars
Question
How would you define a self-driving car?
Definition: What is an autonomous car?
● Autonomous Car: A driverless vehicle capable of fulfilling the main
transportation capabilities of a traditional car.
Classifications of Autonomy according to the NHTSA.
● Level 0: The driver completely controls the vehicle at all times.
● Level 1: Individual vehicle controls are automated, such as electronic stability
control or automatic braking.
● Level 2: At least two controls can be automated in unison, such as adaptive
cruise control in combination with lane keeping.
● Level 3: The driver can fully cede control of all safety-critical functions in certain
conditions and the car provides a "sufficiently comfortable transition time" for the
driver to do so.
● Level 4: The vehicle performs all safety-critical functions for the entire trip, with
the driver not expected to control the vehicle at any time.
Purpose
What kinds of things does a self-driving car need to be able to do?
Purpose
● navigate to a given destination based on passenger-provided instructions

● avoid environmental obstacles

● safely avoid other vehicles

● obey the laws of the road


History
Linrrican Wonder
● Houdina Radio Control, 1925
● Made by Francis P Houdina
● Traveled up Broadway and down Fifth Avenue through the thick of the traffic
jam
Futurama
● sponsored by General Motors at the 1939
World's Fair
● radio-controlled electric cars
○ propelled via electromagnetic fields
RCA Labs
● 1953- RCA Labs built a miniature car guided and controlled by wires
● 1958- Full sized system made
○ developed in collab. with General Motors
Mercedes Benz
● 1980’s- vision-guided Mercedes-Benz robotic van
○ designed by Ernst Dickmanns and his team at the Bundeswehr University
Munich
● achieved a speed of 39 miles per hour (63 km/h) on streets without traffic
History
● Carnegie Mellon’s Navlab and ALV projects in 1984
● Mercedes-Benz and Budeswehr University Munich’s EUREKA Promethius
Project in 1987
● Others:
○ Continental Automotive Systems, IAV, Autoliv Inc., Bosch, Nissan,
Renault, Toyota, Audi, Volvo, Peugeot, AKKA Technologies, Vislab from
University of Parma, Oxford University, Google
■ these companies were more prevalent 2010-2015
DEMO I, II, and III
● US-funded military efforts
● demonstrated the ability of unmanned ground vehicles to navigate miles of
difficult off-road terrain
The Grand Challenges (I, II, and III)
● a fundamental problem in science or engineering, with broad applications,
whose solution would be enabled by the application of high performance
computing resources that could become available in the near future
● Grand Challenges were US policy terms set as goals in the late 1980s for
funding high-performance computing and communications research
DARPA Grand Challenge (2004)
● DARPA (Defense Advanced Research Projects Agency)
● March 13, 2004 in the Mojave Desert
● No cars finished
● Sandstorm from CMU traveled furthest: 11.78 km (7.32 mi)
Grand Challenge II (2005)
● 6:40am on October 8, 2005
Grand Challenge III (2007) aka Urban Challenge
● November 3, 2007 at the site of the now-closed George Air Force Base
● 96 km (60 mi) urban area course, to be completed in less than 6 hours
● obey all traffic regulations while negotiating with other traffic and obstacles
and merging into traffic
Google
Google’s Technology
● $150,000 in equipment including a $70,000 LIDAR system

● The range finder mounted on the top is a Velodyne 64-beam laser. This laser
allows the vehicle to generate a detailed 3D map of its environment.

● The car uses data collected from these mechanisms to drive itself.
Google’s Technology
How it works: Lidar system
● Laser + radar
● The system detects obstacles and tells the car when to avoid them to
navigate safely.
● It uses a 3D point cloud output provide the necessary data for robot software
to determine where potential obstacles exist in the environment and where
the car is is located relative to those obstacles.
How it works: Velodyne
● Company started experimenting with laser distance in 2005 with the DARPA
Grand Challenge
● Since then, they have vastly reduced the size of the sensor and weight while
improving its performance.
● It is a premier lidar system
How does communication among driverless cars
work?
● vehicles and roadside units as the communicating nodes

○ DSRC devices- 5.9 GHz band with bandwith of 75 MHz- range of 1000m
Communication among driverless cars cont.
● Smart intersections

○ intersections with no lights that communicate for autonomous cars

○ 2012- University of Texas in Austin


Google’s Track Record

● As of July 2015, Google’s cars have been involved in 14 “minor accidents”.


○ only one had resulted in minor injuries
● They’ve logged 1.7 million miles, and Google claims not a single collision was
caused by the self-driving mechanisms
Are we going to see Google on the road soon?
Google plans to make these cars available to the public in 2020.
Other Companies involved (since 1987)
Mercedes-Benz Audi

General Motors Volvo

Bosch Peugeot

Nissan Uber

Renault Google

Toyota Tesla
Mercedes Benz
Audi
Tesla’s Current Auto Pilot
Potential advantages
● being able to get things done while in traffic or on the road

● increase road capacity

● fewer traffic collisions. Experts estimate 300,000 lives can be saved per
decade

● higher speed limits

● reduction in traffic police

● removal of limitations on drivers — age and sobriety won’t be an issue


Potential obstacles
● Liability for damage

● Resistance by individuals to forfeit control of their cars

● Software reliability

● Implementation of legal framework and establishment of government


regulations for self-driving cars

● Drivers being inexperienced if situations arose requiring manual driving

● Loss of driving-related jobs

● Loss of privacy
Legislation
In the United States, state vehicle codes generally do not envisage — but do not
necessarily prohibit — highly automated vehicles.
Public Opinion
What do you think?

Would you be comfortable with an autonomous vehicle?


Public Opinion
● of 2,006 surveyed consumers, 49% would be comfortable
○ Accenture, 2011
● of 17,400 owners, 37% would be interested purchasing a self driving
○ 2012, J.D. Power and Associates
○ dropped to 20% if the technology costs $3000 or more
● of 1,000 German drivers, 10% undecided, 44% skeptical, 24% hostile
○ 2012, automotive researcher Puls
Discussion: Liability
● Situation: If a traditional automobile gets hit by a driverless car, who is
responsible?
● Opinion?
● Take a minute talk with the person next to you and decide what you think.
Discussion: Children
● Situation: Driverless cars may one day be able to pick a child up from school
and take him home if the laws permit
● Opinion?
Discussion: Licenses
● If driverless cars are a thing of the future, will driver licenses be a thing of the
past?
● Opinion?
● Take a minute talk with the person next to you and decide what you think.

Discussion: Morals
● If there was a choice to swerve into a schoolbus and potentially kill the
children onboard but save the driver, or divert the car to kill the driver but save
the children, how should the car be programmed?
● A real life application of The Trolley Problem
● Opinions?
● Take a minute talk with the person next to you and decide what you think.

Discussion: Jobs
● Will there still be a demand for auto insurance? What about public
transportation and taxi jobs, just to name a few?
● Opinion?
Predictions: Possible Developments
● By 2016, Mercedes plans to introduce "Autobahn Pilot" aka Highway Pilot, the
system allows a car to automatically pass someone while driving on a
highway.

● By early 2017, the US Department of Transportation hopes to publish a rule


mandating vehicle-to-vehicle (V2V) communication.

● By 2018, Elon Musk expects Tesla Motors to have developed mature serial
production version of fully self-driving cars, where the driver can fall asleep
behind the wheel.
Predictions: Possible Developments
● By 2018, Nissan anticipates to have a feature that can allow the vehicle
maneuver its way on multi-lane highways.

● By 2020, Volvo envisages having cars in which passengers would be immune


from injuries.

● By 2020, GM, Mercedes-Benz, Audi, Nissan, BMW, Renault, Tesla, Google


and Toyota all expect to sell vehicles that can drive themselves at least part of
the time

● By 2020, Google autonomous car project head's goal to have all outstanding
problems with the autonomous car be resolved.
SMART SPEAKER
CONSUMERADOPTION

REPORT
MARCH2019
U.S.

G I V I N G V OICE T O A RE V O L U TI O N
Table of Contents AboutVoicebot AboutVoicify
Introduction // 3 Voicebot produces the leading online publication, Voicify is the market leader in voice experience
newsletter and podcast focused on the voice and AI management software that combines voice
Smart Speaker Ownership // 6 industries. Thousands of entrepreneurs, developers, optimized content management, cross-platform
investors, analysts and other industry leaders look deployment, and voice-specific customer insights.
Smart Speaker Use Cases //15 to Voicebot each week for the latest news, data, The Voicify Voice Experience Platform™ enables
analysis and insights defining the trajectory of the marketers to connect with their customers by
Voice Assistants on Smart Phones // 22
next great computing platform. At Voicebot, we give creating highly engaging and personalized voice
Voice App Discovery // 25 voice to a revolution. experiences that are automatically deployed to
a broad array of voice platforms such a s voice
Consumer Sentiment about Smart Speakers // 29 assistants (Amazon Alexa, Google Assistant and
Microsoft Cortana), chatbots and other services.
Conclusion // 32 Methodology The platform enables non-technical users to
deploy feature-rich voice applications quickly and
Additional Resources // 33 The survey was conducted online during the first
efficiently while offering the flexibility of unlimited
week of January 2019 and was completed by 1,038
customization.
U.S. adults age 18 or older that were representative
of U.S. Census demographic averages. Because we
reached only online adults which represent 89% of [Link]
the population according to Pew Research Center,
some totals are adjusted downward to provide
device and usage numbers relevant to the entire
adult population. Other findings are relative to device
ownership and do not require adjustment.
SMART SPEAKER CONSUMER ADOPTIONREPORT

One in Four U.S. Consumers Have


Access to a Smart Speaker Today
Smart speakers continued to be popular in 2018 keeping up a torrid pace of consumer adoption. In
January 2018, there were 47.3 million U.S. adults with a smart speaker and by the end of the year
that rose to 66.4 million. That means 26.2% of all U.S. adults have access to a smart speaker.

Moving Past Early Adopters


The number of smart speakers per user also rose more than 10% from 1.8 in 2018 to 2.0 in 2019. That
suggests there are about 133 million smart speakers in use in the U.S. today. However, the expansion
in smart speaker ownership has also brought in more casual users. Whereas over 60% of smart
speaker owners in January 2018 identified themselves a s daily users, less than 50% did so a year
later. And, the number of device owners that claim to use their smart speakers never or rarely doubled
to 26%. That seems like a natural evolution of early adopters being more frequent users than the early
majority users coming afterward.

Regardless, when more than one-in-four consumers are using a device and its voice assistant, the
media, brands, game makers, service providers, independent developers, and even governments are
sure to take notice. This recognition is playing out with more voice apps published. The number of
Alexa skills rose by 2.2 times to nearly 60,000 in the U.S. alone in 2018. During the s ame period Google
Actions grew at a slightly faster rate of 2.5 times to over 4,000.

© [Link] - All Rights Reserved 2019 PAGE


179
SMART SPEAKER CONSUMER ADOPTIONREPORT

A Different Smart Speaker Ecosystem, but the Same Leaders Smart Speakersare Solidly in the Early Majority Market
Voicebot reported in the fall of 2018 that Phase 1 of smart speaker adoption One way we can put the current state of smart speaker adoption in perspective is
was over and we were entering Phase 2. The second phase is characterized by to consider a standard technology adoption lifecycle first developed in the 1950’s
the influx of more casual users but also by the introduction of new product form at Iowa State University and popularized in the 1990’s by Geoffrey Moore.
factors and new manufacturers.
The model posits that about 16% of the user population will be “innovators” and
The most significant of these changes h as been the emergence of smart “early adopters” followed by 34% that will be among the “early majority.” With more
displays. When Amazon was the only manufacturers of these voice-first devices than 26% population adoption, smart speakers are securely in the “early majority”
with display screens, adoption was minimal. However, the introduction of Google segment today.
Assistant enabled smart displays has helped drive sales, including Amazon, a s it
An interesting aspect of moving along the adoption curve is that later adopters
brought more attention to the product category.
have different preferences than early adopters. Two areas of difference are
There are also many more manufacturers today than in 2017. Big names in audio typically placing higher value in broader feature sets and integrations with other
such a s Bose, Bang & Olufsen, and Klipsch all entered the smart speaker segment devices. You should expect to see smart speaker makers emphasize features,
in 2018 offering more consumer choice. However, the most significant new smart convenience of access, and third-party integrations more in the coming year.
speaker launch in 2018 was Apple HomePod. That appears to have captured
a significant number of new sales in Q1 and Q2, but seems to have tapered off
in Q3 and Q4. Although Apple was threatening to break up the smart speaker 2019
duopoly, it appears that Amazon and Google enter 2019 nearly a s strong a s they
did in 2018 by maintaining 85% in total installed base market share.

Innovators Early adopters Early majority Late majority Laggards

U.S. Smart SpeakerAdoption Curve


Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

© [Link] - All Rights Reserved 2019 PAGE


180
Smart Speaker
Ownership
SMART SPEAKER CONSUMER ADOPTIONREPORT

U.S. Smart Speaker Owners Rise 40% in 2018


The percent of U.S. adults that own smart speakers rose 40.3% was under two years from the time when there was more than
in 2018 climbing from 47.3 million to 66.4 million during the one manufacturer participating in the category to reach that
year. This increase means more than one-in-four U.S. adults milestone.
now have access to a smart speaker based voice assistant.
CES 2019 reaffirmed that including a smart voice assistant is no
We have moved past the notion that smart speakers may be longer a differentiator for home speaker makers. Like Bluetooth
a novelty a s they are now in such widespread use that one- before it, making speakers smart has become a must-have
in-three smartphone owners have one. It took fewer than four feature. We now have dozens of smart speaker models from
years for smart speakers to achieve 25% adoption from the numerous manufactures, but soon will have hundreds to choose
initial introduction restricted to Amazon Prime members. And, it from even if 85% of users favor just two device makers.

U.S. Adult Smart Speaker Installed Base January 2019

Total US Adult Population 39.8%


253 MILLION One-Year Growth

J a n 2019 / 66.4 MILLION

J a n 2018 / 47.3 MILLION

Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019


© [Link] - All Rights Reserved 2019 PAGE 6
SMART SPEAKER CONSUMER ADOPTIONREPORT

Amazon Echo Leads But Google Home Narrows the Gap

Amazon continued to have the leading installed base of smart speakers in 2018 despite its market
share shrinking from about 72% to 61%. Google was a big mover shifting from 18.4% to nearly 24%,
accounting for precisely half of Amazon’s market share decline. U.S. Smart Speaker Market Share by Brand
January 2018 &2019
The “Other” category was led by Apple and Sonos, and overall the non-Amazon, non-Google device
market share rose by 50% over 2017. More than half of this growth is attributed to Apple HomePod
2019
which had a strong debut in the first half of 2018, but then tapered off in sales a s the year went on.
There were several new smart speakers introduced in 2018 and many focused on sound quality. It
61.1% 23.9% 15.0%
appears consumers are open to adding these higher end smart speakers to their device collection a s
Amazon Google Other
over three-quarters of “Other” category smart speaker owners also report having either an Amazon
Echo or Google Home device.
2018
Sonos went public in 2018 and was clear in its investor documents that voice assistant integration
was critical to the company’s future competitiveness. However, the inability to launch a Google
Assistant enabled speaker may have hurt its appeal with consumers a s the company’s overall smart 71.9% 18.4% 9.7%
Amazon Google Other
speaker market share fell during the year. We can surmise that most of the Sonos fans that wanted
an Alexa-based speaker already bought their device in 2017. As the overall market expanded in 2018,
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
few additional Sonos One devices were purchased and the company’s relative market share fell.
Adding Google Assistant support in 2019 may help reverse this market share slide.

© [Link] - All Rights Reserved 2019


SMART SPEAKER CONSUMER ADOPTIONREPORT

Low Price Speakers Account for 43% of Market

Amazon Echo Dot is the most widely adopted smart speaker by a significant
margin. The sub $50 list price device is frequently available for less than $30 and
U.S. Smart Speaker Market Share by Device - January 2019
refurbished models can be acquired for under $20. This device has proven more
popular than Amazon’s higher priced offerings such a s the Echo, Echo Plus, Echo
31.4% 11.2% 10.0%
Spot, and Echo Show.
Amazon Echo Dot Google Home Other
In the Google portfolio, the Home and Home Mini appear to be equally popular
with 11.2% share each. There are likely to be more Home Minis in use today in
terms of total devices a s this analysis reflects the number of users with access to
a device. If you have one Home and three Minis, you are counted a s one in each
category. And, this may be common a s 87% of Google smart speaker owners
report having both devices. 11.2%
23.2% Home Mini
Apple HomePod and Sonos One lead with smart speaker market share in the Echo orPlus
“Other” category. It appears that smart displays with Google Assistant along with
2.7%
the introduction of Apple HomePod in February 2018 were the key drivers leading Apple
to a 50% growth in this category during the year. Keep in mind that aside from HomePo
HomePod, the “Other” category devices all have Alexa or Google Assistant on d

board, so the dominance of Amazon and Google voice assistants extends beyond 3.5% // EchoSpot
2.2%
1.2% Sonos One
their own products. 3.0% Voicebot
Source: // Amazon
Smart EchoShow
Speaker Consumer Adoption Report Jan 2019 0.2% ////Home
HomeHub
Max

© [Link] - All Rights Reserved 2019


SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Display Ownership Rises Quickly, Amazon Leads


The Amazon Echo Show debuted a s the first smart speaker
U.S. Smart Display Adoption by Smart Speaker Owners
with a display screen now known a s a smart display in J u n e
2017. Later in 2017, Amazon also launched the smaller Echo
January 2019
Spot. With Amazon a s the single smart display manufacturer, 13.2%
only 2.8% of all smart speaker owners had adopted one of the
devices in 2017. September 2018
7.1%
This figure rose rapidly in early 2018 a s Amazon engaged in
aggressive discounting of the devices and then later in the year May 2018
5.9%
after manufacturers starting introducing smart displays driven
by Google Assistant. By year-end 2018, smart displays were
owned by 13.2% of smart speaker owners, a 558% growth rate January 2018
2.8%
in total installed base from about 1.3 million to 8.7 million.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

[Link] Display Market Share 2018


Today, smart displays in the U.S. are either Alexa or Google
Assistant enabled regardless of the manufacturer. Although
33.0% 67.0%
Amazon doesn’t have many third-party smart display OEM
Google Amazon
Assistant Alexa partners, it has maintained 67% market share in the category.
That means Google Assistant enabled devices rose from zero
to one-third market share in less than six months. This may
have risen faster if Google’s smart display, the Home Hub, had
launched earlier in the year. Despite not appearing for sale until
October 2018, Home Hub captured 38.5% of Google Assistant
smart display sales.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019
SMART SPEAKER CONSUMER ADOPTIONREPORT

U.S. Smart Speaker Frequency of Use 2018 New Smart Speaker Owners are Less
63.6%
Likely to be Daily Users
Maybe the biggest change in the composition of smart speaker owners is the
47.4% influx of more casual users of the devices. Nearly 64% of device owners in
January 2018 reported being daily users. In January 2019, that number fell to only
about 47%. Monthly users were fairly similar with the offsetting difference being
the infrequent users which rose from 13% to over 26%.
26.5% 26.1%
23.5%
This seems like a natural progression. Early
12.9% adopters of technology are more likely to
incorporate them quickly into their daily habits than
consumers that tend to adopt later. However, this
2018 2019 2018 2019 2018 2019 will be a metric to monitor going forward. Three
NEVER ORRARELY MONTHLY DAILY
out of four smart speaker owners still report being
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 monthly active users. As long
a s we see that type of consistent usage along
with continued growth, smart speakers will
continue to grow in importance a s a voice
assistant channel for consumer engagement.

© [Link] - All Rights Reserved 2019


SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Speakers Per Household Rise to 2.0


Over 40% of smart speaker owners now have multiple devices. That is up from 34% in 2018. Households
with two, three, and more than five devices all rose in 2019 a s a percentage of all smart speaker owners. This
change suggests that many households are finding utility in smart speakers being nearby.

The data indicate that the industry sold about 48 million smart speakers in the U.S. in 2018 bringing the total
in use to about 133 million up from about 85 million at the end of 2017. Of the 19 million new smart speaker
owners, 31% have purchased multiple devices. That compares to 49% of U.S. adults that have owned smart
speakers for more than a year and have multiple devices.

Smart Speakers Per Household -U.S.


0.7% / 5-10 devices
2.3% / 5-10 devices
1.4% / 10+ devices 3.2% / 4 devices 0.4% / 10+ devices
3.3% / 4 devices

8.0% 14.4%
3 devices
3 devices

65.7% 58.1%
19.3% 1 device 1 device
2 devices

23.2%
2 devices

2018 2019
© [Link] - All Rights Reserved 2019 Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 PAGE 11
SMART SPEAKER CONSUMER ADOPTIONREPORT

Living Room and Bedroom are Most Popular Locations

For the second straight year, the living room was the Where Consumers Have Smart Speakers
most common location for smart speakers. At just
under 45% it was ahead of the bedroom at 37.6%
2.3% // Garage
which had about the same percentage a s last year 37.6% 14.4%
but moved up from third to the second most popular Bedroom Home Office 32.7%
spot. Third place went to the kitchen. At right around
Kitchen
33%, the kitchen seems to have fallen from favor a
bit among smart speaker owners. It’s still popular, 44.4%
but down from 41% in January 2018. Living Room
Most of the other locations were fairly similar to last
year with the exception of the home office which
grew by about one-third. As consumers have been
adding more smart speakers to their collection,
the home office seems to be a common second
location. 2.0%
6.2% // Bathroom 6.5% // DiningRoom Work
Office
Note: Multiple responses accepted, numbers total more than 100%
© [Link] - All Rights Reserved 2019 Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 PAGE 12
SMART SPEAKER CONSUMER ADOPTIONREPORT

Amazon Prime and Gmail Users More Likely to be Smart Speaker Owners
AMAZON PRIME GMAIL USERS
It will surprise few people that Amazon Prime members are Gmail users are also more likely than all users to own a smart
50% more likely to own a smart speaker and more likely to speaker, in this case by about 31%. However, Gmail users are
own an Echo branded device. Amazon Echo smart speakers no more likely than all users to own a Google Home device.
command a 70% market share among Prime members, but In fact, they are almost exactly representative of all users
surprisingly also adopt Google Home products at almost a when it comes to Amazon, Google, and third-party branded
22% rate. Non-Prime members are more likely to adopt third smart speakers. Whereas a Prime membership and Gmail use
party smart speakers made by manufacturers other than suggests a bias toward early technology adoption, only the
Amazon or Google. Prime membership seems to materially influence consumer
choice of smart speakers.

Amazon Prime Member Smart SpeakerMarket Share


U.S. Smart Speaker OwnershipRates
U.S. 2019

All Smart Speaker Owners All


U.S. 26.2%
23.9% Adults
61.1% 15.0%
Google
Amazon Echo Home Other
Gmail
Users 38.6%
Amazon Prime Members
Amazon
70.0% 22.0%
8.0%
Prime 45.3%
Google Members
Amazon Echo Other
Home Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

© [Link] - All Rights Reserved 2019 PAGE 13


Smart Speaker
UseCases
SMART SPEAKER CONSUMER ADOPTIONREPORT

Questions, Music, & Weather Still Reign Supreme

always. When you look at monthly and daily use, it provides a far more accurate
For the second straight year, asking general questions is the top indication of why consumers are using the devices and in many cases why they
use case most commonly tried by smart speaker owners. may be buying a second or third smart speaker for the home. For example, nearly
However, it is not the top use case employed on a monthly or one-in-four smart speaker owners say they set an alarm to play on their smart
daily basis. That distinction goes to listening to streaming music speaker daily. That would suggest a device location for the bedroom may become
services a s it did in 2018. Third place both years was asking increasingly important.
about the weather which is followed by Timers and Alarms in the
fourth and fifth positions. Number six in 2019 was listening to the The biggest variance is smart home control which is ninth in terms of “ever tried”
radio. and fourth for “daily active use.” You must have a smart home device to use
this feature so that automatically eliminates some people from trial. However,
You may have noticed that four of the top five use cases are what controlling lights or thermostats are already daily functions and if you have smart
are considered first-party services. That means they are provided home devices for these features, then switching your habits from smartphone
by the voice assistant natively. Two of the top six use cases app control to voice interaction is a relatively easy change. What smart speaker
involve music which are third-party entertainment services. and voice assistant developers want to see is consumers using these devices
Positions 7-9 all go to the more traditional third-party services, frequently and incorporating them into daily routines. This not only leads to a
many of which were made by independent developers of Alexa higher perception of value by consumers but also leads to stickiness which
skills and Google Actions. So, the order of use frequency at a means the devices are less likely to be removed or swapped out by consumers for
category level are first-party utilities, third-party entertainment, and a competing product.
third-party apps and services.

Frequency Sometimes
© [Link] - All Rights Reserved 2019 More Important Than Trial PAGE 15

You will notice that many analyses of smart speaker use only
focus on what users have tried. This offers a pretty solid
guide to what users value, but not
SMART SPEAKER CONSUMERADOPTIONREPORT

Smart Speaker Use Case Frequency January


2019
Ask a question 84.0% 66.0% 36.9%
Listen to s tr ea m i n g m u s i c service 83.0% 69.9% 38.
Check the weather 80.1% 61.4% 35.6% 2%

Set a n alarm 62.4% 41.8% 23.5%


Set a timer 62.4% 46.7% 22.9%
Listen to radio 54.9% 40.5% 21.2%
Use a favorite Alexa skill / Google Action 48.7% 35.0% 18.3%
Play g a m e or answer trivia 48.0% 29.1% 10.8%
Control smart home devices 45.8% 33.3% 23.5%
Listen to n ew s or sports 43.8% 28.8% 13.4%
Search for product info 41.2% 27.8% 10.8%
Call someone 40.2% 23.5% 11.4%
Find a recipe / cooking instructions 40.2% 26.1% 7.8%
Listen to podcast / other talk formats 39.9% 26.5% 11.1%

Check traffic 36.9% 22.9% 11.8%


A cce s s my calendar 31.7% 21.2% 11.4% EVER TRIED

S e n d a text m e s s a g e 30.4% 18.3% 10.5% MONTHLY

M ak e a purchase 26.1% 15.0% 3.9% DAILY


© [Link] - All Rights Reserved 2019 Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019 PAGE 16
SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Home and Voice Commerce Rise


SMART HOME Monthly Active [Link] Home Users on Smart Speakers
Smart home use cases were a notable mover between 2018-19.
In 2018 it was the fourteenth most common use case and it
rose to ninth in 2019. Over 45% of smart speaker owners have
used them to control smart home devices up from only 38% in
2018. And, one-third of smart speaker owners report now using
voice for smart home device interactions on a monthly basis,
29.9% 33.3%
up from 30% in 2017. At one time, it was assumed that smart
2017 2018
speaker adoption was being driven by smart home aficionados.
It may be that the large audience of smart speaker owners is
now the key catalyst for further smart home adoption.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

VOICE COMMERCE
Voice commerce was a mover for a different reason. It had
Monthly Active U.S. Smart Speaker Voice Commerce Users
the lowest frequency of use cases tracked this year. However,
it also showed relative growth in monthly active users during
2018. Monthly active users rose 10.5% from 13.6% to 15.0%.
This is still a relatively new use case that consumers are
becoming accustomed to, but the growth is indicative of the 13.6% 15.0%
utility of shopping by voice. And, this isn’t just users searching 2017 2018
for products. The responses were specific to making purchases.
When it comes to product search, over 40% of users have
attempted this use case on smart speakers and 28% do so
monthly. These are figures that are increasingly difficult for
consumer brands toignore.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019 PAGE 17
SMART SPEAKER CONSUMER ADOPTIONREPORT

Music, News & Movies Categories Lead in QuestionTopics


Question Category Frequency on Smart Speakers The most tried smart speaker feature
and second highest frequency
Music
monthly and daily use case is asking
54.9%
general questions. This feature
News 35.6%
turns a smart speaker into a voice
Movies 34.6%
interactive search engine. Music
How to instructions 28.4%
related questions are by far the most
History 25.8%
common at about 55% of users. This
Products 24.8% was followed by news and movies
Restaurants 23.2% which were both topics identified
Sports 22.9% by over one-third of smart speaker
Retail storehours 22.6% owners.
Science 21.2%
Math 17.7% The next tier of questions clusters the
Games 16.7% 20-30% range of users ranging from
Health and wellness 16.0%
asking for how-to instructions and
None of theabove product information to retail store
16.0%
Celebrities hours and science. Topics that are far
15.7%
Politics less common include work-related
11.4%
information, finance and investing,
Local Business 10.1%
Travel
and fashion. These were all registered
8.2%
by 5% or fewer smart speaker
Other 6.9%
owners.
Professional / work related topics 5.2%
Finance, banking, orinvesting 4.3%
Fashion 2.9% Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

© [Link] - All Rights Reserved 2019 PAGE 18


SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Home Devices Popular with Smart Speaker Owners

Smart home devices are popular among Smart Home Devices Used by U.S. Smart Speaker Owners
smart speaker owners. More than 55%
of smart speaker owners say they have
at least one smart home device that they
control by voice. Of course, being able to
interact by voice with your smart home
devices doesn’t mean you are going to
use it a s about one-in-five consumers 33.3% 21.2% 14.4% 12.4%
with smart home devices have never tried Smart media controller,
SmartTV SmartLights SmartThermostat
controlling them with their smart speaker. game console or cable box

The most popular smart home devices by a


wide margin are smart TVs with 33.3%. That
was followed by smart lighting at 21.2%, 55.6%
voice interactive game consoles and cable Have smart
boxes at 14.4% and smart thermostats at home devices 10.5% 8.8% 2.9% 2.0%
12.4%. Rounding out the top five were video
doorbells at 10.5%. Video doorbell Smartcameras Smartappliances Other
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019
SMART SPEAKER CONSUMER ADOPTIONREPORT

Customer Service Coming to a Smart Speaker Near You?


A number of companies are starting to think about
how smart speakers can help address common
customer service questions. At a basic level, this could
be an FAQ from a company’s website repurposed U.S. Consumer Interest in Smart
for voice. The goal would be to provide customer Speaker Use for CustomerService
convenience and deflect inbound call center contacts
when practical. This could include switching the
customer to a mobile, online, or call center channel
30.4% 31.4%
when the voice app cannot adequately address the
inquiry.
Unsure Yes

A more advanced implementation might include


account linking and accessing information specific
to a particular user. Finally, you could enable users to
connect by voice to a live agent to address detailed
issues that are best handled through a phone call.

Today, 31.4% of consumers would like to be able to


use their smart speakers to contact customer service
departments. Another 30% are unsure and just under
40% are not interested. However, this means that more
38.2%
than 20 million U.S. adults are interested in doing this
No
today and a comparable amount are open to it. This Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

could be a significant opportunity for customer service


departments to increase customer satisfaction and
help them resolve issues more quickly.

© [Link] - All Rights Reserved 2019


Voice Assistants on Smart Phones
SMART SPEAKER CONSUMER ADOPTIONREPORT

Voice Assistant Use on Smartphones Rising Quickly


Voice assistant use on smartphones rose quickly
during 2018, particularly in the second half of the Have Used a Voice Assistant on a Smartphone
year. While just under 56% of smartphone owners
had used a voice assistant on their device in
January 2018hat figure rose to over 70% one year
later. That means about 31 million U.S adults tried
using a voice assistant on their smart phone for the
first time in 2018. 70.2%
J ANUARY
More than half of the growth was driven by 2019
consumers trying Amazon Alexa on their
smartphone, rising 124% during 2018. Before
January 2018, users could access Alexa through
the search bar in the Amazon app. However, it was
an obscure location for it and not well publicized.
That changed in early 2018 when an voice activation
button was added to the Alexa app. These features
added to the actual Alexa app likely led to the sharp
increase in use.

Bixby trial by consumers rose 57% and Google


56.7%
J ANUARY
Assistant just 16%, granted from a much larger
2018
base. The smallest incremental gain was Apple Siri
which rose only 4% over all of 2018. It appears that
Apple Siri, which h as been around the longest, did
not have a discovery problem and that few users
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
were introduced to it for the first time in 2018.
© [Link] - All Rights Reserved 2019
SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Speakers Increase Voice Assistant Use on Smartphones

Smart speaker owners are 10% more likely to have used a voice assistant on a Voice Assistant Use Frequency on Smartphones by Smart Speaker Ownership
smartphone. They are also more likely to be daily users. Of all voice assistant
users on smartphones 27.6% report being daily users. Among smart speaker
owners that figure rises to 39.8%. 49.3%
RARELY
One-third of smart speaker owners say after purchasing the device they are using 26.8%
voice assistants on their smartphones more frequently. J u s t over 50% say they
are using smartphone-based voice assistants about the same and only 14% say
A 31.5%
they are using them less. There is a growing consensus that smart speakers are
T
displacing time normally spent on smartphones and many people posit that this L
will help reduce screen time. An accelerant for reducing screen time may be using E
voice assistants on smartphones a s well. This reduces the touch, swipe, and look A
S 1
for many use cases.
T
39.8%
M 33.3%
Smart speaker owners are about a s likely a s non-owners to be monthly voice
O
assistant users on smartphones. However, they are twice a s likely to be daily N Smartphone Voice Assistant Smartphone Voice Assistant
T Use Frequency of Non Smart Use Frequency of Smart
users. Data is consistently showing that usage of voice on one platform increases H Speaker Owners Speaker Owners
usage of voice on others platforms. L
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
Y

© [Link] - All Rights Reserved 2019

A 9.1%
T
L
E
A
S
T
Voice App Discovery
SMART SPEAKER CONSUMER ADOPTIONREPORT

Friends are the Most Common Source of Voice App Discovery

For the second year in a row, just under half of smart speaker owners say they How Smart Speaker Owners Discover Voice Apps
don’t actually discover new voice apps. One-in-four device owners rely on friends I don’t
to introduce them to new voice apps followed by just 15% that note social media 49.7%
a s a discovery channel. Friends

A smaller group of users between 11-14% cite Amazon and Google’s primary 26.8%
promotion channels a s key sources of discovery, such a s their in-app and online Social media
stores and weekly emails. Not far behind these channels is advertising which was 15.4%
less visible in previous years, but now is a source of voice app discovery for one- Alexa skill store / Google Assistant discover section
in-ten smart speaker owners. 13.7%
Email newsletter from Alexa or GoogleAssistant
Discovery is the top issue facing third-party voice app publishers today. The voice
assistant user base is growing quickly, but about half of these users are only 11.1%
discovering first-party solutions provided by the voice assistants themselves Ads / commercials
such a s Alexa and Google Assistant. Many third-parties are having more trouble 10.5%
capturing new users. Word-of-mouth appears to be the most effective channel, Newsmedia
but is the hardest to tap into. So, most voice app publishers should focus on a 7.2%
variety of techniques ranging from news media coverage and social media to Other
advertising to drive discovery today.
2.9%
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019 PAGE 25
SMART SPEAKER CONSUMER ADOPTIONREPORT

Number of 3rd Party Voice Apps Tried on Smart Speakers


The same ratio of smart speaker owners report hav- What does this tell us? Well, there remains a dis-
ing used third-party voice apps (i.e. those not made covery problem where about half of smart speaker
by the voice assistant provider) at the end of 2018 owners have never tried a third-party voice app and
a s the previous year. More than half of consumers relegate their usage to first-party voice assistant
that try third-party voice apps have tried at least interactions. However, once smart speaker owners
three and nearly 12% have used more than six. do employ a third-party voice assistant, they are very
likely to become monthly users.
Even more interesting is about 93% of smart speak-
er owners that try third-party voice apps say they This is good news for third-party voice app publish-
become monthly users. And, 55% of these monthly ers that can drive discovery. If they can induce trial,
users are accessing at least two voice apps from they have a good chance of converting the introduc-
third parties and 35.5% are using three or more. tion to regular use.

Percent of Smart Speaker


Users That Try 3rd Party
Voice Apps and Convert to
Consumers That 51.3% 48.7% 92.6% Monthly Users
Have Have
Have Used 3rd Not Used Used
Become
Monthly
Party Voice Apps Users
3.6%
>10 1.3%
6- 10
6.6%
8.3% >10
6-10
31.0% Number of 3rd
1 27.6% 44.7% Party Voice Apps
Number of 3rd Party 3 -5 1 Accessed by
40.5%
Voice Apps Tried on 3 -5 Monthly Users
Smart Speakers 16.7%
2 19.7%
2

Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019


© [Link] - All Rights Reserved 2019
SMART SPEAKER CONSUMER ADOPTIONREPORT

Voice App Reviews Rise in 2018

Only 14% of smart speaker owners have ever left a review of a third-party voice U.S. Smart Speaker Owners That Have Left a Voice App Review
app. This is up from just 11% at the end of 2017 and is a reflection of the fact that
most reviews must be submitted through a visual interface for a user experience Once
designed for no visual interaction. 6.9% More than once
7.2%
Amazon introduced voice ratings in late 2018 which enabled users to offer a star 85.9%
rating for an Alexa skill by voice after using it. However, this was limited to a few Never
skills and not available for skill publishers to set a s a feature on their own. By
contrast, Google Assistant users are more likely to use the voice assistant both on
smartphones and smart speakers. That multimodal use profile might explain why
Google Home owners are about 11% more likely to have left a review than those
with Amazon Echo devices.

With all of that said, 14% seems like a small number of smart speaker owners
leaving reviews until you consider the fact that only 48.7% say they have even
used a third-party voice app. That means 29% of device owners that have tried a
third-party voice app have left a review. This is a promising figure given the friction
involved in actually leaving a review provided voice app publishers can increase
the proportion of smart speaker owners that try third-party apps.
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019
© [Link] - All Rights Reserved 2019
Consumer Sentiment
About Smart Speakers
SMART SPEAKER CONSUMER ADOPTIONREPORT

Consumers Want to Be Understood What Qualities U.S. Users Value in SmartSpeakers


How well it understands
Two-thirds of consumers say how well a voice assistant me when Ispeak 67.0%
understands them is an important quality. That is followed by
sound quality at 54.9%, how much it can do at 50.7%, and the Sound quality 54.9%
speed of response at 45.1%. Every other quality is well behind
these top four characteristics. How much it can do 50.7%

Amazon, Apple, and Google executives have spoken many times How fast itresponds 45.1%
about their focus on adding personality to voice assistants despite
the fact that it is considered important by only 15.4% of smart Its personality 15.4%
speaker owners. That lower rating m a y be influenced by the fact
that personality is offered by all of the leading voice assistant Whether it has my
favorite mediaentertainment 14.4%
providers, but it is clearly not something having an impact today.
Whether the voice assistant is
the same as my mobile device 10.1%
It is also notable that smartphone ownership only influenced
smart speaker selection for about one-in-ten consumers. Apple I am not interested
in a smart speaker 9.8%
and Google would like that linkage to be higher given their
dominance of smartphone-based voice assistants worldwide.
Whether it has goodgames 4.6%
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

© [Link] - All Rights Reserved 2019 PAGE 29


SMART SPEAKER CONSUMER ADOPTIONREPORT

Privacy Concerns Don’t Preclude Smart Speaker Ownership

There is a lot of talk in the media about consumer


privacy fears related to having a smart speaker
U.S. Consumer Perception of
in the household. Those fears don’t seem to be
Smart Speaker PrivacyRisk
undermining adoption, but it is true that two-thirds
of consumers express at least some privacy
concerns and 26% are very concerned.

Interestingly, the privacy concerns for all consumers


11.8%
I don’tknow
and those that do not own smart speakers are 21.5%
nearly identical. For example, only 27.7% of Notconcerned
consumers without smart speakers said they were
very concerned about privacy issues compared to
21.9% of device owners. This means that even some
26.0%
Very
consumers with privacy concerns went ahead and
purchased smart speakers. Those consumers very
concerned 20.8%
Mildy
concerned about privacy risks are about 16% less concerned
likely to own a smart speaker, but they still make up
a sizeable proportion of users. 20.0%
Moderately
Apple has made a big deal about asserting that Siri concerned
and the iOS ecosystem are more protective of user
privacy than its competitors. However, that doesn’t Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

seem to be making much impression on consumers


a s the most privacy concerned consumers are
actually slightly less likely to own a HomePod.

© [Link] - All Rights Reserved 2019 PAGE 30


SMART SPEAKER CONSUMER ADOPTIONREPORT

About 10% of Consumers Without Smart Speakers are Interested


Only about one-third of consumers without smart Non Smart Speaker Owners Opinions AboutDevices
speakers have no interest in the product category. That
is down from 38% at the end of 2017. The number of
No opinion / not interested
non-owners that say they don’t need smart speakers
because of the capabilities of smartphones is up from 34.2%
21.2% to 24.2%. My smartphone has all the functionality I need

Also, the number of consumers without devices


29.6%
that expect to purchase one fell from 11.8% to 9.8%. Notinterested
However, seven out of 10 of those prospective smart 24.2%
speaker owners say they will acquire their first device Concern that the device will record what I'm saying
this year. That would lift total smart speaker ownership
to about one-third of U.S. adults in 2019.
23.0%
They’re tooexpensive
Another interesting data point is that 12.2% of
12.2%
consumers without smart speakers said this year that
I hope to get one this year
price was a deterrent for acquiring a smart speaker
up from 8.8% last year. That is surprising considering 7.1%
that discounted prices for smart speakers were more I hope to get one after this year
aggressive in 2018 than in 2017. With that said, there
2.7%
were many new smart speakers in the premium
Waiting for Samsung GalaxyHome
category that stretched the pricing umbrella upwards
despite more options at lower prices. 1.0%
Source: Voicebot Smart Speaker Consumer Adoption Report Jan 2019

© [Link] - All Rights Reserved 2019 PAGE 31


SMART SPEAKER CONSUMER ADOPTIONREPORT

Smart Speakers Are a Catalyst for Voice Adoption


The trend line for smart speaker adoption remains on an upward A Catalyst for Smart Home Adoption
slope and has quickly surpassed 25% of the U.S. adult population.
More than three out of four consumers that didn’t own a smart It is widely believed that the first wave of smart speaker adop-
speaker but planned to buy one in 2018 did so. With that same tion was driven in part by smart home early adopters. It was one
follow-through rate, more than one-third of U.S. adults will own a of the first robust use cases for the devices. However, it may be
smart speaker at this time next year. That will reflect 33% penetra- that the catalyst behind the next wave of smart home adoption
tion in five years from the first product launch in the category and will be the growing smart speaker user base. If you exclude
just three years since there were two vendors offering solutions. smart TVs and voice-interactive media boxes and focus just on
home automation, there are now more smart speaker house-
Those first two vendors, Amazon and Google, continue to dom- holds. That means smart speaker ownership can be a fertile
inate the market with 85% market share. That is down 5% from source of new customers for smart home automation vendors.
the previous year, but the share is even higher when you consider
households with multiple devices. When it comes to smart speak- The Battle for New Use Cases and Features
er users in the U.S., Amazon and Google have created a duopoly
Smart speaker competition thus far in the U.S. has been char-
that is likely to last for some time.
acterized by price competition and hardware features such a s
A Catalyst for Voice Adoption Across All Surfaces sound quality and display availability. Over the next year, smart
speaker competition of third-party manufacturers will continue
Smart speakers may have already done their job promoting voice to focus on these areas. However, Amazon and Google will seek
assistants a s one-third of device owners report using them on to differentiate themselves more on new features and use cases
smartphones more frequently after purchase. We are seeing voice enabled by their respective voice assistants that increase the
assistants through smart speakers become more deeply embed- perceived daily value of the devices. Voicebot reported in the
ded in household routines while also increasing use on other sur- fall of 2018 that we have entered phase two of voice assistant
faces such a s smartphones and automobile dashboards. Smart adoption which will be characterized by a focus on daily habit
speakers will continue to be an important consumer touchpoint formation and broader usage by the installed base whereas
and likely the catalyst for increased voice assistant use on other phase one was more about device adoption. That trend will
devices beyond communications, navigation, and alerts. continue in 2019.
© [Link] - All Rights Reserved 2019

You might also like