0% found this document useful (0 votes)
14 views86 pages

Deep Learning for Computer Vision

The document outlines a course on Deep Learning in Computer Vision, highlighting its foundations in machine learning and artificial intelligence. It traces the historical development of computer vision from early experiments in the 1950s to modern techniques such as convolutional networks. Key milestones and figures in the field, including Hubel and Wiesel, are discussed to illustrate the evolution of visual recognition technologies.

Uploaded by

xujiangyuwife
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views86 pages

Deep Learning for Computer Vision

The document outlines a course on Deep Learning in Computer Vision, highlighting its foundations in machine learning and artificial intelligence. It traces the historical development of computer vision from early experiments in the 1950s to modern techniques such as convolutional networks. Key milestones and figures in the field, including Hubel and Wiesel, are discussed to illustrate the evolution of visual recognition technologies.

Uploaded by

xujiangyuwife
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Deep Learning in Computer Vision

COMP 4471 & ELEC 4240

Instructor: Qifeng Chen

Credit: CS231N at Stanford University


Artificial Intelligence

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 2 4-Apr-23


Artificial Intelligence

Machine Learning

Computer
Vision

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 3 4-Apr-23


Artificial Intelligence

Machine Learning

Computer Deep
Vision Learning

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 4 4-Apr-23


This class
Artificial Intelligence

Machine Learning

Computer Deep
Vision Learning

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 5 4-Apr-23


This class
Artificial Intelligence

Machine Learning

Computer Deep
Vision Learning

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 6 4-Apr-23


This class
Artificial Intelligence

Machine Learning

Computer Deep
Vision Learning

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 7 4-Apr-23


This class
Artificial Intelligence

Machine Learning

Computer Deep
Vision Learning

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 8 4-Apr-23


Mathematics Artificial Intelligence This class

Neuroscience
Machine Learning

Computer Deep
Vision Learning
Computer
Science

Physics Psychology
Biolog y Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 9 4-Apr-23


Evolution’s Big Bang:
Cambrian Explosion, 530-540million years, B.C.

This image is licensed under CC-BY 2.5

This image is licensed under CC-BY 2.5


This image is licensed under CC-BY 3.0

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 13 4-Apr-23


Camera Obscura
Gemma Frisius, 1545 Encyclopedia, 18th Century

This work is in the public


domain

Leonardo da Vinci,
16th Century A D
This work is in the public domain

This work is in the public


domain

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 11 4-Apr-23


Computer Vision is everywhere!
Left to right:
Image by Roger H Goun is
licensed under CC BY 2.0
Image is CC0 1.0 public domain
Image is CC0 1.0 public domain
Image is CC0 1.0 public domain

Left to right:
Image is free to use
Image is CC0 1.0 public
domain
Image by NASA is licensed
under CC BY 2.0
Image is CC0 1.0 public
domain

Bottom row, left to right


Image is CC0 1.0 public
domain
Image by Derek Keats is
licensed under CC BY 2.0;
changes made
Image is public domain
Image is licensed under CC-BY
2.0; changes made

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 12 4-Apr-23


Where did we come from?

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 13 4-Apr-23


Hubel and Wiesel, 1959
Measure
brain activity

Simple cells:
Response to specific
rotation and orientation

Complex cells:
Response to light
orientation and
Cat image by CNX OpenStax is licensed movement, some
under CC BY 4.0; changes made
translation invariance

1959
Hubel & Wiesel
Response Stimulus
No
response
Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 14 4-Apr-23


Larry Roberts, 1963

(a) O rig inal picture (b) D ifferentiated picture (c) Feature points selected

1959 1963
Hubel & Wiesel Roberts

Lawrence Gilman Roberts, “Machine Perception of Three-Dimensional Solids”, 1963 Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 15 4-Apr-23


1959 1963
Hubel & Wiesel Roberts

[Link] Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 16 4-Apr-23


2 ½-D sketch 3-D model
Input image Edge image

This image is CC0 1.0 public domain This image is CC0 1.0 public domain

Input Primal 2 ½-D 3-D Model


Imag e Sketch Sketch Representation

Zero crossings, Local surface 3-D models


blobs, edges, orientation and hierarchically
Perceived bars, ends, discontinuities organized in
intensities virtual lines, in depth and in terms of surface
groups, curves surface and volumetric
boundaries orientation primitives
1959 1963 1970s
Hubel & Wiesel Roberts David Marr

Stages of Visual Representation, David Marr, 1970s


Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 17 4-Apr-23


Recognition via Parts (1970s)

Generalized Cylinders, Pictorial Structures,


Brooks and Binford, Fischler and Elshlager, 1973
1979
1959 1963 1970s 1979
Hubel & Wiesel Roberts David Marr Gen. Cylinders

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 18 4-Apr-23


Recognition via Edge Detection (1980s)

1959 1963 1970s 1979 1986 John Canny, 1986


Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny David Lowe, 1987

Image is CC0 1.0 public domain Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 19 4-Apr-23


Arriving at an “AI winter”

- Enthusiasm (and funding!) for AI research dwindled


- ”Expert Systems” failed to deliver on their promises
- But subfields of AI continues to grow
- Computer vision, NLP, robotics, compbio, etc.

1959
1963 1970s 1979 1986
Hubel &
Roberts David Marr Gen. Cylinders Canny
Wiesel
AI Winter

Left Image is CC BY 3.0 Middl Image is public Right Image is CC-BY 2.0; changes made Slide inspiration: Justin Johnson
domain

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 20 4-Apr-23


Recognition via Grouping (1990s)

1959 1963 1970s 1979 1986 1997


Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts
Normalized Cuts, Shi and Malik, 1997
AI Winter

Left Image is CC BY 3.0 Middl Image is public Right Image is CC-BY 2.0; changes made Slide inspiration: Justin Johnson
domain

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 21 4-Apr-23


Recognition via Matching (2000s)

Image is public domain


Image is public domain

1959
Hubel & Wiesel
1963
Roberts
1970s
David Marr
1979
Gen. Cylinders
1986
Canny
1997
Norm. Cuts
1999
SIFT
SIFT, D avid
Lowe, 1999
AI Winter

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 22 4-Apr-23


Face Detection

Viola and Jones, 2001

One of the first successful


applications of machine
learning to vision

1959 1963 1970s 1979 1986 1997 1999 2001


Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J

AI Winter

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 23 4-Apr-23


PASCAL Visual
Object Challenge
Image is CC0 1.0 public domain

Train
Person

Airplane

Image is CC0 1.0 public domain

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter

Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 24 4-Apr-23


Perceptron

Frank Rosenblatt, ~1957


1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958
Perceptron Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 25 4-Apr-23


Minsky and Papert, 1969
X Y F(x,y) y
0 0 0

0 1 1

1 0 1

1 1 0
x

Showed that Perceptrons could not learn the XOR


function
Caused a lot of disillusionment in the field
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969
Perceptron Minsky & Papert Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 26 4-Apr-23


Neocognitron: Fukushima, 1980
Computational model the visual system,
directly inspired by Hubel and Wiesel’s
hierarchy of complex and simple cells

Interleaved simple cells (convolution)


and complex cells (pooling)

No practical training algorithm

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969 1980
Perceptron Minsky & Papert Neocognitron Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 27 4-Apr-23


Backprop: Rumelhart, Hinton, and Williams, 1986
Introduced
backpropagation
for computing recognizable
gradients in neural math
networks

Successfully trained
perceptrons with
multiple layers Illustration of Rumelhart et al., 1986 by Lane McIntosh,
copyright CS231n 2017

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969 1980 1985
Perceptron Minsky & Papert Neocognitron Backprop Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 28 4-Apr-23


Convolutional Networks: LeCun et al, 1998

Applied backprop algorithm to a Neocognitron-like architecture


Learned to recognize handwritten digits
Was deployed in a commercial system by NEC, processed handwritten checks
Very similar to our modern convolutional networks!

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969 1980 1985 1998
Perceptron Minsky & Papert Neocognitron Backprop LeNet Slide inspiration: Justin Johnson

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 39 4-Apr-23


2000s: “Deep Learning”
People tried to train neural networks that
were deeper and deeper

Not a mainstream research topic at this time

Hinton and Salakhutdinov, 2006


Bengio et al, 2007

Slide inspiration: Justin Johnson


Lee et al, 2009
Glorot and Bengio, 2010

1959 1963 1970s 1979 1986 1997 1999 2001 2007


Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 30 4-Apr-23


2000s: “Deep Learning”
People tried to train neural networks that
were deeper and deeper

Not a mainstream research topic at this time

No good dataset to work on

Hinton and Salakhutdinov, 2006


Bengio et al, 2007

Slide inspiration: Justin Johnson


Lee et al, 2009
Glorot and Bengio, 2010
1959 1963 1970s 1979 1986 1997 1999 2001 2007
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL

AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 31 4-Apr-23


Output:
The Image Classification Challenge: Scale
T-shirt
1,000 object classes Steel drum
1,431,167 images Drumstick
Mud turtle

Deng et al, 2009


Russakovsky et al. IJC V 2015

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 33 4-Apr-23


AlexNet, 2012

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 44 4-Apr-23


AlexNet: Deep Learning Goes Mainstream

Slide inspiration: Justin Johnson


Krizhevsky, Sutskever, and Hinton, NeurIPS 2012

1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 35 4-Apr-23


AlexNet vs. Neocognitron: 32 years apart

Slide inspiration: Justin Johnson


1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 36 4-Apr-23


2012 to Present: Deep Learning Explosion
CVPR Papers
8000
6000
Subm…
4000 Acce…
2000
0
1985 1990 1995 2000 2005 2010 2015 2020

Slide inspiration: Justin Johnson


Publications at top Computer Vision conference arXiv papers per month (source)
1959 1963 1970s 1979 1986 1997 1999 2001 2004, 2007 2009
Caltech101;
Hubel & Wiesel Roberts David Marr Gen. Cylinders Canny Norm. Cuts SIFT V&J PASCAL ImageNet

AI Winter
1958 1969 1980 1985 1998 2006 2012
Perceptron Minsky & Papert Neocognitron Backprop LeNet Deep Learning AlexNet

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 37 4-Apr-23


2012 to Present: Deep Learning is Everywhere
Year 2010 Year 2012 Year 2014 Year 2015
NEC-UIUC SuperVision GoogLeNet VGG MSRA
Image
Pooling
Convoluti conv-64
on conv-64
Softmax maxpool
Other conv-
Dense descriptor grid: c1o2n8v-
HOG, LBP 128
maxpool
conv-
Coding: local coordinate, c2o5n6v-
256
super-vector maxpool
conv-
c5o1n2v-
512
Pooling, SPM maxpool
conv-
c5o1n2v-
512
Linear SVM maxpool

fc-4096
fc-4096
fc-1000
softmax

[Lin CVPR 2011] [Krizhevsky NIPS 2012]

Figure copyright Alex Krizhevsky, [Szegedy arxiv 2014]


Ilya and Geoffrey Hinton, [Simonyan arxiv 2014] [He ICCV 2015]
Lion image by Swissfrog Figures copyright Alex Krizhevsky, Ilya Sutskever, 2012. Reproduced with permission.
Sutskever, and Geoffrey Hinton, 2012.
is
Reproduced with permission.
licensed under CC BY 3.0

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 38 4-Apr-23


2012 to Present: Deep Learning is Everywhere
Image Classification Image Retrieval

Slide inspiration: Justin Johnson


Figures copyright Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, 2012. Reproduced with permission.

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 39 4-Apr-23


2012 to Present: Deep Learning is Everywhere
Object Detection Image Segmentation

Slide inspiration: Justin Johnson


Ren, He, Girshick, and Sun, 2015 Fabaret et al, 2012

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 40 4-Apr-23


2012 to Present: Deep Learning is Everywhere

Video Classification Activity Recognition

Slide inspiration: Justin Johnson


Simonyan et al, 2014

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 41 4-Apr-23


2012 to Present: Deep Learning is Everywhere
Pose Recognition (Toshev and Szegedy, 2014)

Playing Atari games (Guo et al, 2014)

Slide inspiration: Justin Johnson


Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 42 4-Apr-23
2012 to Present: Deep Learning is Everywhere
Medical Imaging
Whale recognition

Levy et al, 2016 Figure reproduced with permission

Slide inspiration: Justin Johnson


Galaxy Classification

Dieleman et al, 2014


From left to right: public domain by NASA, usage permitted by
ESA/Hubble, public domain by NASA, and public domain. Kaggle Challenge This image by Christin Khan is in the public domain and
originally came from the U.S. NOAA.

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 43 4-Apr-23


2012 to Present: Deep Learning is Everywhere

A white teddy bear A man in a baseball A woman is holding


Image Captioning sitting in the grass uniform throwing a ball a cat in her hand
Vinyals et al, 2015
Karpathy and Fei-Fei,
2015

Slide inspiration: Justin Johnson


All images are CC0 Public domain:
[Link]
[Link]
[Link]
A man riding a wave A cat sitting on a A woman standing on a
[Link]
[Link]
[Link]
on top of a surfboard suitcase on the floor beach holding a surfboard
Captions generated by Justin Johnson using Neuraltalk2

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 44 4-Apr-23


2012 to Present: Deep Learning is Everywhere
Results:
spatial, comparative, asymmetrical, verb,
prepositional

taller than

person person
left of
wear on wear

shirt snow ski


Krishna*, Lu*, Bernstein, Fei-Fei, ECCV 2016

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 45 4-Apr-23


Slide inspiration: Justin Johnson
Original image is CC0 public domain
Starry Night and Tree Roots by Van Gogh are in the public domain
Bokeh image is in the public domain
Mordvinsev et al, 2015
Stylized images copyright Justin Johnson, 2017;
reproduced with permission Gatys et al, 2016
Figures copyright Justin Johnson, 2015. Reproduced with permission. Generated using the Inceptionism approach from a blog post by Google Research.

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 46 4-Apr-23


2012 to Present: Deep Learning is Everywhere

Slide inspiration: Justin Johnson


Karras et al, “Progressive Growing of GANs for Improved Quality, Stability, and Variation”, ICLR 2018

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 47 4-Apr-23


2012 to Present: Deep Learning is Everywhere

Slide inspiration: Justin Johnson


Ramesh et al, “DALL·E: Creating Images from Text”, 2021. [Link]

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 48 4-Apr-23


2012 to Present: Deep Learning is Everywhere

Slide inspiration: Justin Johnson


Ramesh et al, “DALL·E: Creating Images from Text”, 2021. [Link]

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 49 4-Apr-23


Computation
4-Apr-23 Algorithms Data 60
GFLOP per Dollar
CPU GPU (FP32) RTX 3080
50
RTX 3090
40
Deep Learnin g n
30 Explosio
GTX 1080
Ti RTX 2080
20 GeForce Ti

Slide inspiration: Justin Johnson


GeForce GTX 580
8800 GTX (AlexNet
10
)

0
Jan-04 Jul-05 Jan-07 Jul-08 Jan-10 Jul-11 Jan-13 Jul-14 Jan-16 Jul-17 Jan-19 Jul-20
Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 51 4-Apr-23
GFLOP per Dollar
CPU GPU (FP32) GPU (Tensor Core)
200 Recent GPUs have
“Tensor Cores”:
Special hardware
150
for deep learning!
100

Slide inspiration: Justin Johnson


50

0
Jan-04Jul-05Jan-07Jul-08 Jan-10Jul-11 Jan-13Jul-14 Jan-16Jul-17Jan-19Jul-20
Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 52 4-Apr-23
AI’s Explosive Growth & Impact

Number of attendance Startups Developing AI Enterprise Application AI


At AI conferences Systems Revenue
Source: The Gradient Source: Crunchbase, VentureSource, Sand Source: Statista
Hill Econometrics

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 53 4-Apr-23


Generative AI: Text to Image

Stable Diffusion
Dall E 2
Generative AI: Text to Image

Source: Dall E 2
Generative AI: Image Completion

Source: Dall E 2
Generative AI: Video Translation

Source: [Link]/CoDeF_Page/
Generative AI: Video Generation

Text-to-Video Generation

Sora by OpenAI Kling by Kuaishou Veo3 by Google


Generative AI: Video Generation

Open-source Text-to-Video Generation

Hunyuan Video Step-Video-T2V Wan Video


The Age of Generative AI!
Large Language Models (LLMs) and Generative Vision
Models are rapidly evolving how we work and create:

● Education
● Software Engineering
● Productivity
● Law
● Finance
● Healthcare Anthropic’s Claude describes its goal to be a
helpful, harmless, and honest assistant
● Art, and more

60
GPT4 writes an iOS app Midjourney’s AI generates images
Exponential Growth of Large Language Models
GPT4 is estimated
to be here!

[Link] Year
Despite the successes, computer
vision still has a long way to g o

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 62 4-Apr-23


Computer Vision Can Cause Harm
Harmful Stereotypes Affect people’s lives

Barocas et al, “The Problem With Bias: Allocative Versus Representational Harms in Machine Learning”, SIGCIS 2017
Kate Crawford, “The Trouble with Bias”, NeurIPS 2017 Keynote Source: [Link]
Source: [Link] (2015) [Link]
Example Credit: Timnit Gebru

Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 63 4-Apr-23


Computer Vision Can Save Lives

Slide inspiration: Justin Johnson


Fei-Fei Li & Ruohan Gao & Yunzhu Li CS231n: Lecture 1 - 64 4-Apr-23
Deep Learning Basics
• Image Classification: A core task in Computer Vision

cat

This image by Nikita is


licensed under CC-BY 2.0

April 1,2025 Stanford CS231n 10thLecture


7 1-
Deep Learning Basics
• Image Classification: A core task in Computer Vision

cat

Linear Classifier
This image by Nikita is
licensed under CC-BY 2.0

April 1,2025 Stanford CS231n 10thLecture


8 1-
Deep Learning Basics
• Image Classification: A core task in Computer Vision

cat

Regularization & Optimization


This image by Nikita is
licensed under CC-BY 2.0

April 1,2025 Stanford CS231n 10thLecture


9 1-
Deep Learning Basics
• Image Classification: A core task in Computer Vision

cat

Neural Networks
This image by Nikita is
licensed under CC-BY 2.0

April 1,2025 Stanford CS231n 10thLecture


1 1-
Perceiving and Understanding the Visual World

Tasks Models

April 1,2025 Stanford CS231n 10thLecture


1 1-
Tasks Beyond Image Classification
Semantic Object Instance
Classification
Segmentation Detection Segmentation

CAT GRASS, CAT, TREE, DOG, DOG, CAT DOG, DOG, CAT
SKY

No spatial extent No objects, just pixels Multiple Object This image is CC0 public domain

April 1,2025 Stanford CS231n 10thLecture


1 1-
Tasks Beyond Image Classification
Video Multimodal Video Visualization &
Classification Understanding Understanding

Running? Jumping?

April 1,2025 Stanford CS231n 10thLecture


1 1-
Models Beyond Multi-Layer Perceptron

Illustration of LeCun et al. 1998 from CS231n 2017 Lecture 1

Convolutional neural network

April 1,2025 Stanford CS231n 10thLecture


1 1-
Models Beyond Multi-Layer Perceptron

Recurrent neural network Attention mechanism / Transformers


April 1,2025 Stanford CS231n 10thLecture
1 1-
Large Scale Distributed Training
DataParallelism Model Parallelism

Partitioned Data Shared Data


● Train Large Models on big
datasets faster
● Scale beyond single Worker Worker Worker Worker Worker Worker
GPU/machine limitations
● How?
○ Data Parallelism: Copy the model
to all workers, split the data
○ Model Parallelism: Split model
across devices
○ Synchronous vs. Asynchronous
gradient updates
Shared Model Partitioned Model

April 1,2025 Stanford CS231n 10thLecture


1 1-
Beyond 2D Recognition: Self-supervised
Learning

Image Courtesy of Rohit Kundu

April 1,2025 Stanford CS231n 10thLecture


2 1-
Beyond 2D Recognition: Generative Modeling

Style Transfer
Beyond 2D Recognition: Generative Modeling

“Teddy bears working on new


AI research underwater with
1990s technology”

DALL-E 2

This image is public domain


Beyond 2D Recognition: Generative Modeling

ImageGeneration using Diffusion Models You will learn and implement a generative
model in Assignment 3 that generates
emojis from text inputs

Face with a cowboy hat

[Link]

April 1,2025 Stanford CS231n 10thLecture


2 1-
Beyond 2D Recognition: Vision Language Models

Yasunaga, Michihiro, et al. "Retrieval-augmented multimodal


language modeling." arXiv preprint arXiv:2211.12561 (2022).

Contrastive pre-training in CLIP. The blue squares are the pairs for which we want to
optimize the similarity. Image derived from [Link]

April 1,2025 Stanford CS231n 10thLecture


2 1-
Beyond 2D Recognition: 3D Vision

Choy et al., 3D-R2N2: Recurrent Reconstruction Neural Network (2016)

Zhou et al., 3D Shape Generation and Completion through Point-Voxel Diffusion (2021) Gkioxari et al., “Mesh R-CNN”, ICCV2019

April 1,2025 Stanford CS231n 10thLecture


2 1-
Beyond 2D Recognition: Embodied Intelligence

Li et al., BEHAVIOR-1K:A Benchmark for Embodied AI with 1,000 Everyday Activities and Mandlekar and Xu et al., Learning to Generalize Across Long-
Realistic Simulation (2022) Horizon Tasks from Human Demonstrations (2020)

April 1,2025 Stanford CS231n 10thLecture


2 1-
GPU Resource
• Google Colab

• GPU servers at HKUST


• [Link]
• [Link]
• [Link]
• [Link]
Pre-requisite
• Proficiency in Python, some high-level familiarity
with C/C++/Java
– All class assignments will be in Python (and use numpy),
but some of the deep learning libraries we may look at
later in the class are written in C++.
– A Python tutorial available on course website
• Multivariable Calculus, Linear Algebra
Grading Policy
• 3 Problem Sets: 12% x 3 = 36% (8 Late Days)
• Midterm Exam: 35%
• Course Project: 29% (Separate 8 Late Days)
– Project Proposal: 2%
– Milestone: 3%
– Presentation 4%
– Project Report: 20%
• Late policy
– Each student has 8 free late days – up to 4 late days per
assignment
– Each group has the other 8 free late days – up to 4 late days
for project proposal/milestones/project report
– Afterwards, 25% off per day late
Collaboration Policy
• Rule 1: Don’t look at solutions or code that are not your
own; everything you submit should be your own work

• Rule 2: Don’t share your solution code with others; however


discussing ideas or general strategies is fine and encouraged

• Rule 3: Indicate in your submissions anyone you worked with


Turning in something late / incomplete is better than violating
the honor code
Next Time: Image Classification

K-Nearest Neighbor Linear Classifier

You might also like