Deep Learning for Computer Vision
Deep Learning for Computer Vision
Machine Learning
Computer
Vision
Machine Learning
Computer Deep
Vision Learning
Machine Learning
Computer Deep
Vision Learning
Machine Learning
Computer Deep
Vision Learning
Machine Learning
Computer Deep
Vision Learning
Machine Learning
Computer Deep
Vision Learning
Neuroscience
Machine Learning
Computer Deep
Vision Learning
Computer
Science
Physics Psychology
Biolog y Slide inspiration: Justin Johnson
Leonardo da Vinci,
16th Century A D
This work is in the public domain
Left to right:
Image is free to use
Image is CC0 1.0 public
domain
Image by NASA is licensed
under CC BY 2.0
Image is CC0 1.0 public
domain
Simple cells:
Response to specific
rotation and orientation
Complex cells:
Response to light
orientation and
Cat image by CNX OpenStax is licensed movement, some
under CC BY 4.0; changes made
translation invariance
1959
Hubel & Wiesel
Response Stimulus
No
response
Slide inspiration: Justin Johnson
(a) O rig inal picture (b) D ifferentiated picture (c) Feature points selected
1959 1963
Hubel & Wiesel Roberts
Lawrence Gilman Roberts, “Machine Perception of Three-Dimensional Solids”, 1963 Slide inspiration: Justin Johnson
This image is CC0 1.0 public domain This image is CC0 1.0 public domain
1959
1963 1970s 1979 1986
Hubel &
Roberts David Marr Gen. Cylinders Canny
Wiesel
AI Winter
Left Image is CC BY 3.0 Middl Image is public Right Image is CC-BY 2.0; changes made Slide inspiration: Justin Johnson
domain
Left Image is CC BY 3.0 Middl Image is public Right Image is CC-BY 2.0; changes made Slide inspiration: Justin Johnson
domain
1959
Hubel & Wiesel
1963
Roberts
1970s
David Marr
1979
Gen. Cylinders
1986
Canny
1997
Norm. Cuts
1999
SIFT
SIFT, D avid
Lowe, 1999
AI Winter
AI Winter
Train
Person
Airplane
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL
AI Winter
AI Winter
1958
Perceptron Slide inspiration: Justin Johnson
0 1 1
1 0 1
1 1 0
x
AI Winter
1958 1969
Perceptron Minsky & Papert Slide inspiration: Justin Johnson
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL
AI Winter
1958 1969 1980
Perceptron Minsky & Papert Neocognitron Slide inspiration: Justin Johnson
Successfully trained
perceptrons with
multiple layers Illustration of Rumelhart et al., 1986 by Lane McIntosh,
copyright CS231n 2017
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL
AI Winter
1958 1969 1980 1985
Perceptron Minsky & Papert Neocognitron Backprop Slide inspiration: Justin Johnson
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL
AI Winter
1958 1969 1980 1985 1998
Perceptron Minsky & Papert Neocognitron Backprop LeNet Slide inspiration: Justin Johnson
AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning
AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet
AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet
AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet
AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet
AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet
AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet
AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet
fc-4096
fc-4096
fc-1000
softmax
taller than
person person
left of
wear on wear
0
Jan-04 Jul-05 Jan-07 Jul-08 Jan-10 Jul-11 Jan-13 Jul-14 Jan-16 Jul-17 Jan-19 Jul-20
Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 51 4-Apr-23
GFLOP per Dollar
CPU GPU (FP32) GPU (Tensor Core)
200 Recent GPUs have
“Tensor Cores”:
Special hardware
150
for deep learning!
100
0
Jan-04Jul-05Jan-07Jul-08 Jan-10Jul-11 Jan-13Jul-14 Jan-16Jul-17Jan-19Jul-20
Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 52 4-Apr-23
AI’s Explosive Growth & Impact
Stable Diffusion
Dall E 2
Generative AI: Text to Image
Source: Dall E 2
Generative AI: Image Completion
Source: Dall E 2
Generative AI: Video Translation
Source: [Link]/CoDeF_Page/
Generative AI: Video Generation
Text-to-Video Generation
● Education
● Software Engineering
● Productivity
● Law
● Finance
● Healthcare Anthropic’s Claude describes its goal to be a
helpful, harmless, and honest assistant
● Art, and more
60
GPT4 writes an iOS app Midjourney’s AI generates images
Exponential Growth of Large Language Models
GPT4 is estimated
to be here!
[Link] Year
Despite the successes, computer
vision still has a long way to g o
Barocas et al, “The Problem With Bias: Allocative Versus Representational Harms in Machine Learning”, SIGCIS 2017
Kate Crawford, “The Trouble with Bias”, NeurIPS 2017 Keynote Source: [Link]
Source: [Link] (2015) [Link]
Example Credit: Timnit Gebru
cat
cat
Linear Classifier
This image by Nikita is
licensed under CC-BY 2.0
cat
cat
Neural Networks
This image by Nikita is
licensed under CC-BY 2.0
Tasks Models
CAT GRASS, CAT, TREE, DOG, DOG, CAT DOG, DOG, CAT
SKY
No spatial extent No objects, just pixels Multiple Object This image is CC0 public domain
Running? Jumping?
Style Transfer
Beyond 2D Recognition: Generative Modeling
DALL-E 2
ImageGeneration using Diffusion Models You will learn and implement a generative
model in Assignment 3 that generates
emojis from text inputs
[Link]
Contrastive pre-training in CLIP. The blue squares are the pairs for which we want to
optimize the similarity. Image derived from [Link]
Zhou et al., 3D Shape Generation and Completion through Point-Voxel Diffusion (2021) Gkioxari et al., “Mesh R-CNN”, ICCV2019
Li et al., BEHAVIOR-1K:A Benchmark for Embodied AI with 1,000 Everyday Activities and Mandlekar and Xu et al., Learning to Generalize Across Long-
Realistic Simulation (2022) Horizon Tasks from Human Demonstrations (2020)