0% menganggap dokumen ini bermanfaat (0 suara)
11 tayangan33 halaman

Metode Pembelajaran Bayesian dalam ML

Diunggah oleh

Eko darmawan
Hak Cipta
© All Rights Reserved
Kami menangani hak cipta konten dengan serius. Jika Anda merasa konten ini milik Anda, ajukan klaim di sini.
Format Tersedia
Unduh sebagai PDF, TXT atau baca online di Scribd
0% menganggap dokumen ini bermanfaat (0 suara)
11 tayangan33 halaman

Metode Pembelajaran Bayesian dalam ML

Diunggah oleh

Eko darmawan
Hak Cipta
© All Rights Reserved
Kami menangani hak cipta konten dengan serius. Jika Anda merasa konten ini milik Anda, ajukan klaim di sini.
Format Tersedia
Unduh sebagai PDF, TXT atau baca online di Scribd

Machine Learning

Metode Bayesian
Entin Martiana
Knowledge Engineering Research Group
Soft Computing Laboratory
Department of Information and Computer Engineering
Politeknik Elektronika Negeri Surabaya
Konten
• Mengapa Metode Bayes
• Probabilitas Bersyarat
• Probabilitas Bersyarat dalam Data
• MAP Hyphotesis
Tujuan Instruksi Umum
Mahasiswa mampu menyelesaikan masalah – masalah
menggunakan metode mesin pembelajaran yang tepat
berdasarkan supervised, unsupervised dan
reinforcement learning, baik secara individu maupun
berkelompok/kerjasama tim.
Tujuan Instruksi Khusus
• Memahami metode naive Bayes
• Memahami metode Bayes Gaussian
• Mampu menerapkan Bayesian
Mengapa Metode Bayes
• Metode Find-S tidak dapat digunakan untuk data yang tidak konsisten
dan data yang bias, sehingga untuk bentuk data semacam ini salah
satu metode sederhana yang dapat digunakan adalah metode bayes.
• Metode Bayes ini merupakan metode yang baik di dalam mesin
pembelajaran berdasarkan data training, dengan menggunakan
probabilitas bersyarat sebagai dasarnya.
Probabilitas Bersyarat
S

X P( X  Y )
XY P( X | Y ) 
Y P(Y )

Probabilitas X di dalam Y adalah probabilitas interseksi X dan Y dari


probabilitas Y, atau dengan bahasa lain P(X|Y) adalah prosentase
banyaknya X di dalam Y
Probabilitas Bersyarat Dalam Data
# Cuaca Temperatur Kecepatan Angin Berolah-raga
1 Cerah Normal Pelan Ya
2 Cerah Normal Pelan Ya
3 Hujan Tinggi Pelan Tidak
4 Cerah Normal Kencang Ya
5 Hujan Tinggi Kencang Tidak
6 Cerah Normal Pelan Ya

Banyaknya data berolah-raga=ya adalah 4 dari 6 data maka dituliskan


P(Olahraga=Ya) = 4/6
Banyaknya data cuaca=cerah dan berolah-raga=ya adalah 4 dari 6 data maka dituliskan
P(cuaca=cerah dan Olahraga=Ya) = 4/6
4/6
P(cuaca  cerah | olahraga  ya)  1
4/6
Probabilitas Bersyarat Dalam Data
# Cuaca Temperatur Berolahraga
1 cerah normal ya
2 cerah tinggi ya
3 hujan tinggi tidak
4 cerah tinggi tidak
5 hujan normal tidak
6 cerah normal ya

Banyaknya data berolah-raga=ya adalah 3 dari 6 data maka dituliskan


P(Olahraga=Ya) = 3/6
Banyaknya data cuaca=cerah, temperatur=normal dan berolah-raga=ya adalah 4 dari 6 data
maka dituliskan
P(cuaca=cerah, temperatur=normal, Olahraga=Ya) = 2/6
2/6 2
P(cuaca  cerah, temperatur  normal | olahraga  ya)  
3/ 6 3
Metode Bayes
X1 X2 …. Xn

P(Y | X k )
P( X k | Y ) 
Y  P(Y | X i )
i

Keadaan Posteriror (Probabilitas Xk di dalam Y) dapat dihitung dari


keadaan prior (Probabilitas Y di dalam Xk dibagi dengan jumlah dari
semua probabilitas Y di dalam semua Xi)
MAP Hyphotesis

MAP (Maximum A propri Probability) Hypothesis menyatakan hipotesa


yang diambil berdasarkan nilai probabilitas berdasarkan kondisi prior
yang diketahui.
Contoh MAP Hypotheses
Diketahui hasil survey yang dilakukan sebuah lembaga kesehatan menyatakan bahwa 90%
penduduk di dunia menderita sakit paru-paru. Dari 90% penduduk yang sakit paru-paru ini 60%
adalah perokok, dan dari penduduk yang tidak menderita sakit paru-paru 20% perokok.

Fakta ini bisa didefinisikan dengan: X=sakit paru-paru dan Y=perokok.


Maka : P(X) = 0.9
P(~X) = 0.1
P(Y|X) = 0.6  P(~Y|X) = 0.4
P(Y|~X) = 0.2  P(~Y|~X) = 0.8
Dengan metode bayes dapat dihitung:
P(Y|X).P(X) = (0.6) . (0.9) = 0.54
P(Y|~X) P(~X) = (0.2).(0.1) = 0.02
P({Y}|X) = 0.54/(0.54+0.02) = 0.96
P({Y}|~X) = 0.54/(0.54+0.02) = 0.04
Bila diketahui seseorang merokok, maka dia menderita sakit paru-paru karana P({Y}|X) lebih besar dari
P({Y}|~X). HMAP diartikan mencari probabilitas terbesar dari semua instance pada attribut target atau semua
kemungkinan keputusan.
Bayes Theorem
• Goal: To determine the most probable hypothesis, given the data D
plus any initial knowledge about the prior probabilities of the various
hypotheses in H.
• Prior probability of h, P(h): it reflects any background knowledge we
have about the chance that h is a correct hypothesis (before having
observed the data).
• Prior probability of D, P(D): it reflects the probability that training data
D will be observed given no knowledge about which hypothesis h
holds.
• Conditional Probability of observation D, P(D|h): it denotes the
probability of observing data D given some world in which hypothesis
h holds.
Bayes Theorem
• Posterior probability of h, P(h|D): it represents the probability that
h holds given the observed training data D. It reflects our confidence
that h holds after we have seen the training data D and it is the
quantity that Machine Learning researchers are interested in.
• Bayes Theorem allows us to compute P(h|D):

P(h|D)=P(D|h)P(h)/P(D)
Maximum A Posteriori (MAP) Hypothesis and Maximum
Likelihood
• Goal: To find the most probable hypothesis h from a set of candidate
hypotheses H given the observed data D.
• MAP Hypothesis, hMAP = argmax h ∈ H P(h|D)
= argmax h ∈ H P(D|h)P(h)/P(D)
= argmax h ∈ H P(D|h)P(h)
• If every hypothesis in H is equally probable a priori, we only need to
consider the likelihood of the data D given h, P(D|h). Then, hMAP
becomes the Maximum Likelihood,
hML= argmax h ∈ H P(D|h)
Bayes Optimal Classifier
• One great advantage of Bayesian Decision Theory is that it gives us
a lower bound on the classification error that can be obtained for a
given problem.
• Bayes Optimal Classification: The most probable classification of
a new instance is obtained by combining the predictions of all
hypotheses, weighted by their posterior probabilities:
argmaxvj∈Vhσ hi∈H P(vh|hi)P(hi|D)
where V is the set of all the values a classification can take and vj
is one possible such classification.
• Unfortunately, Bayes Optimal Classifier is usually too costly to apply!
==> Naïve Bayes Classifier
HMAP Dari Data Training
# Cuaca Temperatur Kecepatan Angin Berolah-raga
1 Cerah Normal Pelan Ya
2 Cerah Normal Pelan Ya
3 Hujan Tinggi Pelan Tidak
4 Cerah Normal Kencang Ya
5 Hujan Tinggi Kencang Tidak
6 Cerah Normal Pelan Ya

Asumsi:
Y = berolahraga,
X1 = cuaca,
X2 = temperatur,
X3 = kecepatan angin.
Fakta menunjukkan:
P(Y=ya) = 4/6  P(Y=tidak) = 2/6
HMAP Dari Data Training
# Cuaca Temperatur Kecepatan Angin Berolah-raga
1 Cerah Normal Pelan Ya Apakah bila cuaca cerah
2 Cerah Normal Pelan Ya dan kecepatan angin
3 Hujan Tinggi Pelan Tidak kencang, orang akan
4 Cerah Normal Kencang Ya berolahraga?
5 Hujan Tinggi Kencang Tidak
6 Cerah Normal Pelan Ya

Fakta: P(X1=cerah|Y=ya) = 1, P(X1=cerah|Y=tidak) = 0


P(X3=kencang|Y=ya) = 1/4 , P(X3=kencang|Y=tidak) = 1/2
HMAP dari keadaan ini dapat dihitung dengan:
P( X1=cerah,X3=kencang | Y=ya )
= { P(X1=cerah|Y=ya).P(X3=kencang|Y=ya) } . P(Y=ya)
= { (1) . (1/4) } . (4/6) = 1/6
P( X1=cerah,X3=kencang | Y=tidak )
= { P(X1=cerah|Y=tidak).P(X3=kencang|Y=tidak) } . P(Y=tidak)
= { (0) . (1/2) } . (2/6) = 0

KEPUTUSAN ADALAH BEROLAHRAGA = YA


Naïve Bayes Algorithm
• Naïve Bayes Algorithm (for discrete input attributes) has two phases
– 1. Learning Phase: Given a training set S,
Learning is easy, just create
For each target value of ci (ci  c1 ,  , c L ) probability tables.
ˆ (C  c )  estimate P(C  c ) with examples in S;
P i i

For every attribute value x jk of each attribute X j ( j  1,  , n; k  1,  , N j )


ˆ ( X  x |C  c )  estimate P( X  x |C  c ) with examples in S;
P j jk i j jk i

Output: conditional probability tables; for elements


Xj , N j  L
– 2. Test Phase: Given an unknown instance , X  ( a ,  , a )
1 n
Look up tables to assign the label c* to X’ if

ˆ (a |c* )    P
[P ˆ (a |c* )]P ˆ (a |c)    P
ˆ (c * )  [ P ˆ (a |c)]P
ˆ (c), c  c* , c  c ,  , c
1 n 1 n 1 L

Classification is easy, just multiply probabilities


Tennis Example
• Example: Play Tennis
The learning phase for tennis example

P(Play=Yes) = 9/14

P(Play=No) = 5/14
We have four variables, we calculate for each we calculate
the conditional probability table

Outlook Play=Yes Play=No Temperature Play=Yes Play=No


Sunny 2/9 3/5
Hot 2/9 2/5
Overcast 4/9 0/5
Mild 4/9 2/5
Rain 3/9 2/5
Cool 3/9 1/5
Humidity Play=Yes Play=No Wind Play=Yes Play=No
Strong 3/9 3/5
High 3/9 4/5 Weak 6/9 2/5
Normal 6/9 1/5
The test phase for the tennis example
• Test Phase
– Given a new instance of variable values,
x’=(Outlook=Sunny, Temperature=Cool, Humidity=High, Wind=Strong)
– Given calculated Look up tables
P(Outlook=Sunny|Play=Yes) = 2/9 P(Outlook=Sunny|Play=No) = 3/5
P(Temperature=Cool|Play=Yes) = 3/9 P(Temperature=Cool|Play==No) = 1/5
P(Huminity=High|Play=Yes) = 3/9 P(Huminity=High|Play=No) = 4/5
P(Wind=Strong|Play=Yes) = 3/9 P(Wind=Strong|Play=No) = 3/5
P(Play=Yes) = 9/14 P(Play=No) = 5/14

– Use the MAP rule to calculate Yes or No

P(Yes|x’): [P(Sunny|Yes)P(Cool|Yes)P(High|Yes)P(Strong|Yes)]P(Play=Yes) = 0.0053


P(No|x’): [P(Sunny|No) P(Cool|No)P(High|No)P(Strong|No)]P(Play=No) = 0.0206

Given the fact P(Yes|x’) < P(No|x’), we label x’ to be “No”.


An Example of the Naïve Bayes Classifier
The weather data, with counts and probabilities
outlook temperature humidity windy play
yes no yes no yes no yes no yes no

sunny 2 3 hot 2 2 high 3 4 false 6 2 9 5


overcast 4 0 mild 4 2 normal 6 1 true 3 3
rainy 3 2 cool 3 1
sunny 2/9 3/5 hot 2/9 2/5 high 3/9 4/5 false 6/9 2/5 9/14 5/14
overcast 4/9 0/5 mild 4/9 2/5 normal 6/9 1/5 true 3/9 3/5
rainy 3/9 2/5 cool 3/9 1/5

A new day
outlook temperature humidity windy play
sunny cool high true ?
• Likelihood of yes
2 3 3 3 9
      0.0053
9 9 9 9 14
• Likelihood of no
3 1 4 3 5
      0.0206
5 5 5 5 14
• Therefore, the prediction is No
The Naive Bayes Classifier for Data Sets with
Numerical Attribute Values

• One common practice to handle numerical attribute values is to assume


normal distributions for numerical attributes.
The numeric weather data with summary statistics
outlook temperature humidity windy play
yes no yes no yes no yes no yes no

sunny 2 3 83 85 86 85 false 6 2 9 5
overcast 4 0 70 80 96 90 true 3 3
rainy 3 2 68 65 80 70
64 72 65 95
69 71 70 91
75 80
75 70
72 90
81 75
sunny 2/9 3/5 mean 73 74.6 mean 79.1 86.2 false 6/9 2/5 9/14 5/14
overcast 4/9 0/5 std dev 6.2 7.9 std dev 10.2 9.7 true 3/9 3/5

rainy 3/9 2/5


• Let x1, x2, …, xn be the values of a numerical attribute in the training data
set.

1 n
   xi
n i 1
n
 
1
  xi   2

n  1 i 1
 w   2
1 
f ( w)  e 2

2 
• For examples,
 66  73 2

f temperature  66 | Yes  
1
e 2  6.2 2
 0.0340
2 6.2
2 3 9
• Likelihood of Yes =  0.0340  0.0221    0.000036
9 9 14

3 3 5
• Likelihood of No =  0.0291  0.038    0.000136
5 5 14
Kelemahan Metode Bayes
• Metode Bayes hanya bisa digunakan untuk persoalan klasifikasi
dengan supervised learning.
• Metode Bayes memerlukan pengetahuan awal untuk dapat
mengambil suatu keputusan. Tingkat keberhasilan metode ini sangat
tergantung pada pengetahuan awal yang diberikan.
Beberapa Aplikasi Metode Bayes
• Menentukan diagnosa suatu penyakit berdasarkan data-data gejala (sebagai contoh hipertensi atau
sakit jantung).
• Mengenali buah berdasarkan fitur-fitur buah seperti warna, bentuk, rasa dan lain-lain
• Mengenali warna berdasarkan fitur indeks warna RGB
• Mendeteksi warna kulit (skin detection) berdarkan fitur warna chrominant
• Menentukan keputusan aksi (olahraga, art, psikologi) berdasarkan keadaan.
• Menentukan jenis pakaian yang cocok untuk keadaan-keadaan tertentu (seperti cuaca, musim,
temperatur, acara, waktu, tempat dan lain-lain)
• Menentukan ekspresi (sedih, gembira, dll) dari kalimat yang diucapkan
Latihan Soal
WAKTU PAKET FREKWEKSI PRIORITAS GANGGUAN
PENDEK BESAR SEDANG RENDAH GANGGUAN
PENDEK KECIL RENDAH TINGGI GANGGUAN
PANJANG BESAR SEDANG TINGGI NORMAL
PANJANG KECIL TINGGI RENDAH NORMAL
PENDEK BESAR TINGGI TINGGI GANGGUAN
PANJANG KECIL RENDAH TINGGI GANGGUAN
PANJANG KECIL TINGGI RENDAH GANGGUAN
PANJANG KECIL SEDANG RENDAH NORMAL
PANJANG BESAR TINGGI TINGGI NORMAL
PANJANG KECIL SEDANG RENDAH GANGGUAN
PENDEK BESAR SEDANG TINGGI NORMAL
PANJANG BESAR RENDAH TINGGI NORMAL

1. Implementasikan fase training untuk data training di atas!


2. Berilah fasilitas untuk memasukkan data test!
3. Berapa persen besarnya error yang terjadi jika data tersebut dimasukkan
pada fase training?
Latihan Soal
USIA KELAMIN MEROKOK OLAHRAGA JANTUNG
TUA PRIA TIDAK YA TIDAK
TUA PRIA YA YA TIDAK
MUDA PRIA YA TIDAK TIDAK
TUA PRIA TIDAK TIDAK TIDAK
MUDA WANITA TIDAK TIDAK YA
MUDA PRIA TIDAK YA YA
MUDA PRIA TIDAK YA TIDAK
TUA WANITA TIDAK TIDAK YA
MUDA PRIA YA TIDAK TIDAK
TUA PRIA YA TIDAK TIDAK
MUDA PRIA YA YA YA
TUA PRIA YA TIDAK TIDAK
MUDA PRIA TIDAK TIDAK TIDAK
TUA PRIA TIDAK YA TIDAK
MUDA PRIA YA TIDAK TIDAK

1. Implementasikan fase training untuk data training di atas!


2. Berilah fasilitas untuk memasukkan data test!
3. Berapa persen besarnya error yang terjadi jika data tersebut dimasukkan
pada fase training?
Referensi
• Modul Ajar Machine Learning, Entin Martiana, Ali
Ridho Barakbah, Nur Rosyid Mubtadaí, Politeknik
Elektronika Negeri Surabaya, 2013.
• Machine Learning, Tom Mitchell, McGraw-Hill. 2008.

Anda mungkin juga menyukai