MACHINE LEARNING
BASIC CLUSTERING
Sindhu Wardhana
Teguh Prasetyo
Aris Budi Santoso
Leonard Yulianus
Definisi
• Cluster adalah kumpulan objek datea
• Memiliki kemiripan dengan antara objek dalam satu cluster;
• Memiliki perbedaan dengan objek lain di luar cluster;
• Clustering adalah metode untuk membagi/ mengelompokan data
berdasarkan properti-properti dari data tersebut;
• Clustering termasuk dalam kategori unsupervised machine
learning
• Data tidak memiliki label/ predefined class
UNSUPERVISED MACHINE LEARNING
Jenis Clustering
Partitional Clustering
▪ Perlu menentukan jumlah kluster
▪ Iterasi untuk menempatkan data ke dalam cluster
▪ K-Means
Hierarchical Clustering
Pembentukan cluster dilakukan secara hirarki
▪ Agglomerative Nested (AGNES) : Bottom-up
Menggabungkan dua titik yang memiliki kemiripan ke dalam sebuah cluster
▪ Divisive Analysis (DIANA) : Top-down
Mulai dari sebuah cluser besar kemudian dibagi
Density Based Clustering
▪ Pembentukan cluster dilakukan berdasar kepadatan titik data pada suatu area;
▪ Antara cluster dipisahkan oleh area dengan kepadatan titik data yang rendah;
▪ Density-Based Spatial Clustering of Applications with Noise (DBSCAN)
Clustering Algorithm Comparison
Al-Raba'nah, Yousef and Mohammed Al-Refai. “Data clustering algorithms: A second look.” International Journal of Advance
Research, Ideas and Innovations in Technology 4 (2018): 1081-1083.
Clustering Algorithm Comparison
[Link]
examples/cluster/plot_cluster_com
[Link]#sphx-glr-auto-exampl
es-cluster-plot-cluster-comparison-
py
Keunggulan
Semakin banyak digunakan pada berbagai bidang
• subgroups of breast cancer patients grouped by their gene expression
measurements
• Pengelompkan karakteristik pelanggan berdasarkan riwayat pencarian dan
pembeliannya
• Pengelompokan film berdasarkan rating yang diberikan pemirsa
• Pemodelan topik dari dokumen-dokumen berupa teks
Data tanpa label lebih mudah diperoleh
Tantangan
more subjective than supervised learning
•• No simple goal for analysis
•• The computer have to learn how to do something that we don’t tell it how to do
Have some issues
•• The number of subgroups (clusters)
•• The different results via K-means with different random initialisations
•• How to assess the performance of the unsupervised learning methods?
The learning (or inference) procedure is hard
Ministry of Finance Data Analytics Community
Clustering Evaluation
Mengukur kedekatan dalam cluster, dan pemisahan antar cluster yang berbeda
Cluster yang baik : Compact, Separated
Ukuran evaluasi clustering:
▪ Silhouette Index
▪ Davies Bouldien Index
▪ Dunn Index
Find the maximum point of average silhouette score
Silhouette Score
Approach
Analytical
The silhouette value
Computatio
measures
n of K how similar a
point is to its own cluster
(cohesion) compared to
other clusters
(separation).
Optimal K value is 5
or 6
WSS (Within-cluster Sum of Find diminishing point (elbow) on WSS curve
Squared) Approach
Analytical
Computatio
Sum of squared error of
n of
each K on a same
datapoint
cluster
Optimal K value is between 4
and 6
Study Case (GOJEK Meeting Points)
Project’s Target
Determine meeting point (gates) of any Place of Interest
(POI)
based on historical driver pick up data
Clustering
POI Example: Blok M Square
Using DBSCAN
After some hyperparameter tuning (eps and
min_samples)
The result is 5 clusters for determining gates on Blok M
Square
Looks ok
POI : Ciputra
World
But it is not
okay on
different
POI
Datapoints are spread so model cant determine
cluster
Using KMeans
Ciputra Blok M
World Square
Different POI may have different number of gates, but Kmean has a fix
value of K
Compare K = 5 and K = 6
• K=5 gives a little
bit more silhouette
point
• K=6 makes 1
additional but not
too significant
cluster (red #5)
• So, better to
choose K = 5
Next Step
Automate the process of finding optimal
K-value for analysing different POI
Give name of each gate (cluster
centroid/mean), so it is easier for drivers
and passengers to locate the gate
NLP case
CONTOH PENERAPAN PADA PYTHON
klik di sini
Ministry of Finance Data Analytics Community
Praktikum
• K-means Clustering
• [Link]
4Yr6Vkg#scrollTo=oSehXCF7ST53
• Clustering Uber Data
• [Link]
q7GHW#scrollTo=q9UyWfWF0FCb
Referensi
• [Link]
[Link]
TERIMA KASIH
Ministry of Finance Data Analytics Community