0% found this document useful (0 votes)
6 views25 pages

K-Means Clustering Basics in Python

The document provides an overview of k-means clustering, emphasizing its efficiency over hierarchical clustering for large datasets. It outlines the steps to generate cluster centers and labels, discusses the calculation of distortion, and introduces the elbow method for determining the optimal number of clusters. Additionally, it highlights the limitations of k-means, including the impact of random seed initialization and bias towards equal-sized clusters.

Uploaded by

Bhuvnesh Verma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views25 pages

K-Means Clustering Basics in Python

The document provides an overview of k-means clustering, emphasizing its efficiency over hierarchical clustering for large datasets. It outlines the steps to generate cluster centers and labels, discusses the calculation of distortion, and introduces the elbow method for determining the optimal number of clusters. Additionally, it highlights the limitations of k-means, including the impact of random seed initialization and bias towards equal-sized clusters.

Uploaded by

Bhuvnesh Verma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Basics of k-means

clustering
C L U S T E R A N A LY S I S I N P Y T H O N

Shaumik Daityari
Business Analyst
Why k-means clustering?
A critical drawback of hierarchical clustering: runtime
K means runs significantly faster on large datasets

CLUSTER ANALYSIS IN PYTHON


Step 1: Generate cluster centers
kmeans(obs, k_or_guess, iter, thresh, check_finite)

obs : standardized observations

k_or_guess : number of clusters

iter : number of iterations (default: 20)

thres : threshold (default: 1e-05)

check_finite : whether to check if observations contain only finite numbers (default: True)

Returns two objects: cluster centers, distortion

CLUSTER ANALYSIS IN PYTHON


How is distortion calculated?

CLUSTER ANALYSIS IN PYTHON


Step 2: Generate cluster labels
vq(obs, code_book, check_finite=True)

obs : standardized observations

code_book : cluster centers

check_finite : whether to check if observations contain only finite numbers (default: True)

Returns two objects: a list of cluster labels, a list of distortions

CLUSTER ANALYSIS IN PYTHON


A note on distortions
kmeans returns a single value of distortions

vq returns a list of distortions.

CLUSTER ANALYSIS IN PYTHON


Running k-means
# Import kmeans and vq functions
from [Link] import kmeans, vq

# Generate cluster centers and labels


cluster_centers, _ = kmeans(df[['scaled_x', 'scaled_y']], 3)
df['cluster_labels'], _ = vq(df[['scaled_x', 'scaled_y']], cluster_centers)

# Plot clusters
[Link](x='scaled_x', y='scaled_y', hue='cluster_labels', data=df)
[Link]()

CLUSTER ANALYSIS IN PYTHON


CLUSTER ANALYSIS IN PYTHON
Next up: exercises!
C L U S T E R A N A LY S I S I N P Y T H O N
How many clusters?
C L U S T E R A N A LY S I S I N P Y T H O N

Shaumik Daityari
Business Analyst
How to find the right k?
No absolute method to find right number of
clusters (k) in k-means clustering

Elbow method

CLUSTER ANALYSIS IN PYTHON


Distortions revisited
Distortion: sum of squared distances of
points from cluster centers

Decreases with an increasing number of


clusters

Becomes zero when the number of clusters


equals the number of points

Elbow plot: line plot between cluster


centers and distortion

CLUSTER ANALYSIS IN PYTHON


Elbow method
Elbow plot: plot of the number of clusters and distortion
Elbow plot helps indicate number of clusters present in data

CLUSTER ANALYSIS IN PYTHON


Elbow method in Python
# Declaring variables for use
distortions = []

num_clusters = range(2, 7)

# Populating distortions for various clusters


for i in num_clusters:
centroids, distortion = kmeans(df[['scaled_x', 'scaled_y']], i)
[Link](distortion)

# Plotting elbow plot data


elbow_plot_data = [Link]({'num_clusters': num_clusters,
'distortions': distortions})

[Link](x='num_clusters', y='distortions',
data = elbow_plot_data)
[Link]()

CLUSTER ANALYSIS IN PYTHON


CLUSTER ANALYSIS IN PYTHON
Final thoughts on using the elbow method
Only gives an indication of optimal k (numbers of clusters)
Does not always pinpoint how many k (numbers of clusters)

Other methods: average silhouette and gap statistic

CLUSTER ANALYSIS IN PYTHON


Next up: exercises
C L U S T E R A N A LY S I S I N P Y T H O N
Limitations of k-
means clustering
C L U S T E R A N A LY S I S I N P Y T H O N

Shaumik Daityari
Business Analyst
Limitations of k-means clustering
How to find the right _K_ (number of clusters)?
Impact of seeds

Biased towards equal sized clusters

CLUSTER ANALYSIS IN PYTHON


Impact of seeds
Initialize a random seed Seed: [Link](1000, 2000)

from numpy import random


Cluster sizes: 29, 29, 43, 47, 52

[Link](12)

Seed: [Link](1,2,3)

Cluster sizes: 26, 31, 40, 50, 53

CLUSTER ANALYSIS IN PYTHON


Impact of seeds: plots
Seed: [Link](1000, 2000) Seed: [Link](1,2,3)

CLUSTER ANALYSIS IN PYTHON


Uniform clusters in k means

CLUSTER ANALYSIS IN PYTHON


Uniform clusters in k-means: a comparison
K-means clustering with 3 clusters Hierarchical clustering with 3 clusters

CLUSTER ANALYSIS IN PYTHON


Final thoughts
Each technique has its pros and cons
Consider your data size and patterns before deciding on algorithm

Clustering is exploratory phase of analysis

CLUSTER ANALYSIS IN PYTHON


Next up: exercises
C L U S T E R A N A LY S I S I N P Y T H O N

You might also like