0% found this document useful (0 votes)
4 views31 pages

Machine Learning Techniques Overview

Uploaded by

henryhe010713
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views31 pages

Machine Learning Techniques Overview

Uploaded by

henryhe010713
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Machine Learning

Introduction

Rajesh Ranganath
Machine Learning

Conditionals

p(y | x)

[Image of code from Atlantic]


Neuroscience
C

[Manning+ 2014]
Causality

[Pearl+]
Machine Translation

English-German translations
src Orlando Bloom and Miranda Kerr still love each other
ref Orlando Bloom und Miranda Kerr lieben sich noch immer
best Orlando Bloom und Miranda Kerr lieben einander noch immer .
base Orlando Bloom und Lucas Miranda lieben einander noch immer .
src ′′ We ′ re pleased the FAA recognizes that an enjoyable passenger experience is not incompatible
with safety and security , ′′ said Roger Dow , CEO of the U.S. Travel Association .
ref “ Wir freuen uns , dass die FAA erkennt , dass ein angenehmes Passagiererlebnis nicht im Wider-
spruch zur Sicherheit steht ” , sagte Roger Dow , CEO der U.S. Travel Association .
best ′′ Wir freuen uns , dass die FAA anerkennt , dass ein angenehmes ist nicht mit Sicherheit und
Sicherheit unvereinbar ist ′′ , sagte Roger Dow , CEO der US - die .
base ′′ Wir freuen uns über die <unk> , dass ein <unk> <unk> mit Sicherheit nicht vereinbar ist mit
Sicherheit und Sicherheit ′′ , sagte Roger Cameron , CEO der US - <unk> .
German-English translations
src In einem Interview sagte Bloom jedoch , dass er und Kerr sich noch immer lieben .
[Luong+ 2015]
ref However , in an interview , Bloom has said that he and Kerr still love each other .
best In an interview , however , Bloom said that he and Kerr still love .
base However , in an interview , Bloom said that he and Tina were still <unk> .
src Wegen der von Berlin und der Europäischen Zentralbank verhängten strengen Sparpolitik in
Verbindung mit der Zwangsjacke , in die die jeweilige nationale Wirtschaft durch das Festhal-
ten an der gemeinsamen Währung genötigt wird , sind viele Menschen der Ansicht , das Projekt
Europa sei zu weit gegangen
ref The austerity imposed by Berlin and the European Central Bank , coupled with the straitjacket
imposed on national economies through adherence to the common currency , has led many people
to think Project Europe has gone too far .
best Because of the strict austerity measures imposed by Berlin and the European Central Bank in
Image Classification

x ImageNet test set, and


2015 classification com
weight layer
resentations also have e
F(x) relu
on other recognition ta
x
weight layer
identity 1st places on: ImageN
COCO detection, and
F(x) + x
relu COCO 2015 competitio
Figure 2. Residual learning: a building block. the residual learning pr
it is applicable in other
are comparably good or better than the constructed solution
[He+ 2015] (or unable to do so in feasible time). 2. Related Work
In this paper, we address the degradation problem by
introducing a deep residual learning framework. In- Residual Representat
stead of hoping each few stacked layers directly fit a [18] is a representation
desired underlying mapping, we explicitly let these lay- with respect to a dictio
ers fit a residual mapping. Formally, denoting the desired formulated as a probab
underlying mapping as H(x), we let the stacked nonlinear of them are powerful s
layers fit another mapping of F(x) := H(x) x. The orig- trieval and classificatio
inal mapping is recast into F(x)+x. We hypothesize that it encoding residual vect
is easier to optimize the residual mapping than to optimize tive than encoding orig
DSTN @ CMU . EDU
al Laboratory DJSCHLEGEL @ LBL . GOV
ratoryAstrophysics PRABHAT @ LBL . GOV

del of op-
h a varia-
xel inten-
able, with
properties
erties are
r distribu-
data sets.
mages. We
y survey,
he current
stial bod-

Figure 1. An image from the Sloan Digital Sky Survey (SDSS,


nerative model 2015) of a galaxy from the constellation Serpens, 100 million
[Regier+ 2015]
light years from Earth, along with several other galaxies and many
odel to be em-
The work we stars from our own galaxy.
Genetics

LWK YRI ACB ASW CDX CHB CHS JPT KH

[Gopalan+ 2016]
Figure S2: Population structure inferred from
What’s this class about?

Broken into four high level themes


■ Supervised Learning

■ Graphical Models and Approximate Inference

■ Causal Inference

■ Reinforcement Learning
Supervised Learning
Take some input x and predict y
Supervised Learning
Take some input x and predict y

umbrella.98 bus.99

umbrella.98
person1.00

person1.00
person1.00
backpack1.00
person1.00 person.99
handbag.96 person.99
person1.00 person1.00 person1.00
person1.00 person1.00
person.95 person.98
person1.00
person1.00 person1.00 person.94 person1.00 person1.00 person.89

person1.00 sheep.99
backpack.99
sheep.99 sheep.86
backpack.93 sheep.82 sheep.96
sheep.96 sheep.93 sheep.91 sheep.95 sheep.96 sheep1.00
sheep1.00
sheep.99
sheep1.00
sheep.99
sheep.96

sheep.99

person.99
bottle.99
dining table.96

bottle.99
bottle.99

person.99person1.00
person1.00
traffic light.96 tv.99

chair.98 chair.99
chair.90
dining table.99 chair.96 wine glass.97
chair.86
bottle.99wine glass.93 chair.99
bowl.85 wine glass1.00

elephant1.00
wine glass.99
wine glass1.00
person1.00 chair.96 chair.99 fork.95

person1.00 traffic light.95 bowl.81


person1.00
traffic light.92 traffic light.84
person1.00 person.85
person.96 truck1.00 person.99
motorcycle1.00 person.96person1.00
person.83 person1.00
motorcycle1.00 person.98
person.99person.91
person.90 person.87 car.99 car.92
person.99
person.92 car.99 car.93
car1.00
motorcycle.95
knife.83

person.96

Figure 2. Mask R-CNN results on the COCO test set. These results are based on ResNet-101 [19], achieving a mask AP of 35.7 and
[He+running
2015] at 5 fps. Masks are shown in color, and bounding box, category, and confidences are also shown.

a seemingly minor change, RoIAlign has a large impact: it 2. Related Work


improves mask accuracy by relative 10% to 50%, showing
bigger gains under stricter localization metrics. Second, we R-CNN: The Region-based CNN (R-CNN) approach [13]
found it essential to decouple mask and class prediction: we to bounding-box object detection is to attend to a manage-
predict a binary mask for each class independently, without able number of candidate object regions [42, 20] and evalu-
ate convolutional networks [25, 24] independently on each
competition among classes, and rely on the network’s RoI
classification branch to predict the category. In contrast, RoI. R-CNN was extended [18, 12] to allow attending to
FCNs usually perform per-pixel multi-class categorization, RoIs on feature maps using RoIPool, leading to fast speed
which couples segmentation and classification, and based and better accuracy. Faster R-CNN [36] advanced this
on our experiments works poorly for instance segmentation. stream by learning the attention mechanism with a Region
Proposal Network (RPN). Faster R-CNN is flexible and ro-
Without bells and whistles, Mask R-CNN surpasses all bust to many follow-up improvements (e.g., [38, 27, 21]),
previous state-of-the-art single-model results on the COCO and is the current leading framework in several benchmarks.
instance segmentation task [28], including the heavily-
engineered entries from the 2016 competition winner. As Instance Segmentation: Driven by the effectiveness of R-
a by-product, our method also excels on the COCO object CNN, many approaches to instance segmentation are based
with Deep Learning

Supervised Learning
urkar * 1 Jeremy Irvin * 1 Kaylie Zhu 1 Brandon Yang 1 Hershel Mehta 1
y Ding 1Take
Aartisome [Link]
Bagul 1input Ball 2predict y
Curtis Langlotz 3
Katie Shpanskaya 3
Matthew P. Lungren 3 Andrew Y. Ng 1

Abstract
algorithm that can detect
chest X-rays at a level ex-
g radiologists. Our algo-
is a 121-layer convolutional
ained on ChestX-ray14, cur-
publicly available chest X-
aining over 100,000 frontal-
es with 14 diseases. Four
mic radiologists annotate a
Input
Chest X-Ray Image
ch we compare the perfor-
Net to that of radiologists.
eXNet exceeds average ra-
CheXNet
121-layer CNN
ance on the F1 metric. We
to detect all 14 diseases in
d achieve state of the art re-
Output
Pneumonia Positive (85%)
seases.

dults are hospitalized with pneu-


,000 die from the disease every
e (CDC, 2017). Chest X-rays
available method for diagnosing
[Rajpurkar+
01), playing a crucial role 2017]
in clin-
Figure 1. CheXNet is a 121-layer convolutional neural net-
work that takes a chest X-ray image as input, and outputs
Supervised Learning
Take some input x and predict y

We will cover deep variants!


Graphical Models and Approximate Inference

Take some input x understand relationships


Graphical Models and Approximate Inference

Take some input x understand relationships


C
3

9
We have illustrated a small subgraph of this large network, score. Our algorithm assigned it to seven communities, w
centered around a specific article. Across the whole network, we classification tags mostly restricted to “Drug: Bio-Affect
can use the posterior bridgeness to filter and find a collection of Body Treating Compositions” and “Surgery.”
articles that have had interdisciplinary impact. In SI Text we show
the top 10 papers in the arXiv network by posterior bridgeness. Comparisons to Ground Truth on Synthetic Networks. We
Graphical Models and Approximate Inference
The top scientific articles in the arXiv network have a wide im-
pact, as they concern data, parameters, or theory applied in
strated that our algorithm can help explore massive rea
networks. As further validation, we performed a ben
various subfields of physics. For example, the top article, “Maps comparison on synthetic networks where the overlappin
of dust infrared emission for use in estimation of reddening and munities are known. We used the “benchmark” tool
cosmic microwave background radiation foregrounds,” (41) con- synthesize networks with the number of nodes ranging fr
structs an accurate full sky map of the dust temperature useful in thousand to one million.
the estimation of cosmic microwave background radiation. This We compared our algorithm to the best existing algorit
Take some input x understand relationships
filtering demonstrates the practical potential for unsupervised detecting overlapping communities (2, 8, 9, 11–13, 17). E
analysis of large networks. The posterior bridgeness score, a gorithm analyzes the (unlabeled) network and returns b

Fig. 2. The discovered community struct


subgraph of the US Patents network (35).
ure shows subgraphs of the top four com
that include citations to “Process for p
porous products” (42). We visualize the
tween the patents and show titles of som
highly cited patents. Each community is
with its dominant classification; nodes are
their bridgeness (39); the local network
ized using the Fruchterman–Reingold a
(46). This is taken from an analysis of the
million node network.

[Gopalan+ 2014] Gopalan and Blei PNAS Early Edition


Causal Inference

Understand the effect of altering x on y


Causal Inference

Understand the effect of altering x on y

person i

B
Reinforcement Learning

Understand how to take a sequence of actions to meet a goal


Reinforcement Learning

Understand how to take a sequence of actions to meet a goal

[Silver+ 2016]
Reinforcement Learning

Understand how to take a sequence of actions to meet a goal

Figure 2: The PySC2 viewer shows a human interpretable view of the game on the left, and coloured
[Vinayals+ ] of the feature layers on the right. For example, terrain height, fog-of-war, creep, camera
versions
location, and player identity, are shown in the top row of feature layers. A video can be found at
[Link]

Thus, the main observations come as sets of feature layers which are rendered at N ⇥ M pixels
(where N and M are configurable, though in our experiments we always used N = M ). Each of
these layers represents something specific in the game, for example: unit type, hit points, owner,
or visibility. Some of these (e.g., hit points, height map) are scalars, while others (e.g., visibility,
Reinforcement Learning

Understand how to take a sequence of actions to meet a goal

We will talk about deep reinforcement learning


■ Learn how different methods work

■ Learn why different methods need to exist

■ How to think about the data generating distribution

■ Class is mathematical and conceptual


Food for thought: How can something learn?
Food for thought: How can something learn?

To learn you need information.


Logistics

Course website: [Link]

Office Hours: Tuesdays, 3:30-4:30pm. Online.

Ed: You will be added

Gradescope: You will be added


Logistics
Deliverables
■ Class will have 5 homeworks (40%)

■ Homeworks are challenging

■ Every class has reading response due at the start of the next
class(10%)

■ You will be asked to scribe one lecture (5%)

■ Participation (5%)

■ Final project (paper and presentation) due at the end of the term
(40 %)

Questions?
Quiz: Homework -1
Quiz
■ Probabilities are between 0 and 1. Consider the Gaussian
distribution with variance 0.01.
1 1
 ‹
2
p(x) = p exp − x
2π0.01 2 ∗ 0.01

p(0) ≈ 3.98. How is this possible?

∂y
■ Let x ∈ Rn . Let y = sin(Ax) ∈ Rm . What is ∂x? What is its
dimension?

■ Define the covariance. Give an example of two random variables


that have zero covariance but are not independent

■ Define conditional independence. Can variables be independent,


but also conditionally dependent?
Introductions
Particular Subjects of Interest

You might also like