0% found this document useful (0 votes)
24 views25 pages

Strategic Leadership in Cluster Analysis

This document outlines a case study in unsupervised learning using the R programming language. It begins with objectives to complete analysis using unsupervised learning techniques, reinforce concepts, and emphasize creativity. It then discusses downloading breast cancer cell data, preparing it for modeling, and exploratory data analysis. Next steps include hierarchical and k-means clustering, combining clustering with principal component analysis, and comparing results. The document concludes with a review of key unsupervised learning concepts covered.

Uploaded by

Anto
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views25 pages

Strategic Leadership in Cluster Analysis

This document outlines a case study in unsupervised learning using the R programming language. It begins with objectives to complete analysis using unsupervised learning techniques, reinforce concepts, and emphasize creativity. It then discusses downloading breast cancer cell data, preparing it for modeling, and exploratory data analysis. Next steps include hierarchical and k-means clustering, combining clustering with principal component analysis, and comparing results. The document concludes with a review of key unsupervised learning concepts covered.

Uploaded by

Anto
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Introduction to the

case study
UN SUP ERVISED L EARN IN G IN R

Hank Roark
Senior Data Scientist at Boeing
Objectives
Complete analysis using unsupervised learning

Reinforce what you've already learned

Add steps not covered before (e.g. preparing data, selecting good features for supervised
learning)

Emphasize creativity

UNSUPERVISED LEARNING IN R
Example use case
Human breast mass data:
Ten features measured of each cell nuclei

Summary information is provided for each group of cells

Includes diagnosis: benign (not cancerous) and malignant (cancerous)

1Source: K. P. Benne and O. L. Mangasarian: "Robust Linear Programming Discrimination of Two Linearly
Inseparable Sets"

UNSUPERVISED LEARNING IN R
Analysis
Download data and prepare data for modeling

Exploratory data analysis (# observations, # features, etc.)

Perform PCA and interpret results

Complete two types of clustering

Understand and compare the two types

Combine PCA and clustering

UNSUPERVISED LEARNING IN R
Review: PCA in R
[Link] <- prcomp(x = iris[-5],
scale = FALSE,
center = TRUE)
summary([Link])

Importance of components:
PC1 PC2 PC3 PC4
Standard deviation 2.0563 0.49262 0.2797 0.15439
Proportion of Variance 0.9246 0.05307 0.0171 0.00521
Cumulative Proportion 0.9246 0.97769 0.9948 1.00000

UNSUPERVISED LEARNING IN R
Unsupervised learning is open-ended
Steps in this use case are only one example of what can be done

There are other approaches to analyzing this dataset

UNSUPERVISED LEARNING IN R
Let's practice!
UN SUP ERVISED L EARN IN G IN R
PCA review and
next steps
UN SUP ERVISED L EARN IN G IN R

Hank Roark
Senior Data Scientist at Boeing
Review thus far
Downloaded data and prepared it for modeling

Exploratory data analysis

Performed principal component analysis

UNSUPERVISED LEARNING IN R
Next steps
Complete hierarchical clustering

Complete k-means clustering

Combine PCA and clustering

Contrast results of hierarchical clustering with diagnosis

Compare hierarchical and k-means clustering results

PCA as a pre-processing step for clustering

UNSUPERVISED LEARNING IN R
Review: hierarchical clustering in R
# Calculates similarity as Euclidean distance between observations
s <- dist(x)

# Returns hierarchical clustering model


hclust(s)

Call:
hclust(d = s)

Cluster method : complete


Distance : euclidean
Number of objects: 50

UNSUPERVISED LEARNING IN R
Review: k-means in R

# k-means algorithm with 5 centers, run 20 times


kmeans(x, centers = 5, nstart = 20)

One observation per row, one feature per column

k-means has a random component

Run algorithm multiple times to improve odds of the best model

UNSUPERVISED LEARNING IN R
Let's practice!
UN SUP ERVISED L EARN IN G IN R
Wrap-up and review
UN SUP ERVISED L EARN IN G IN R

Hank Roark
Senior Data Scientist at Boeing
Case study wrap-up
Entire data analysis process using unsupervised learning

Creative approach to modeling

Prepared to tackle real world problems

UNSUPERVISED LEARNING IN R
Types of clustering

UNSUPERVISED LEARNING IN R
Dimensionality reduction

UNSUPERVISED LEARNING IN R
Model selection
# Initialize total within sum of squares error: wss
wss <- 0

# Look over 1 to 15 possible clusters


for (i in 1:15) {
# Fit the model: [Link]
[Link] <- kmeans(pokemon, centers = i, nstart = 20, [Link] = 50)
# Save the within cluster sum of squares
wss[i] <- [Link]$[Link]
}

# Produce a scree plot


plot(1:15, wss, type = "b",
xlab = "Number of Clusters",
ylab = "Within groups sum of squares")

UNSUPERVISED LEARNING IN R
Interpreting PCA results

UNSUPERVISED LEARNING IN R
Importance of scaling data

UNSUPERVISED LEARNING IN R
Course review
[Link] <- prcomp(x = iris[-5],
scale = FALSE,
center = TRUE)
summary([Link])

Importance of components:
PC1 PC2 PC3 PC4
Standard deviation 2.0563 0.49262 0.2797 0.15439
Proportion of Variance 0.9246 0.05307 0.0171 0.00521
Cumulative Proportion 0.9246 0.97769 0.9948 1.00000

UNSUPERVISED LEARNING IN R
Dendrogram

UNSUPERVISED LEARNING IN R
Strengths and weaknesses of each algorithm

UNSUPERVISED LEARNING IN R
Course review
# Repeat for components 1 and 3
plot([Link]$x[, c(1, 3)], col = (diagnosis + 1),
xlab = "PC1", ylab = "PC3")

UNSUPERVISED LEARNING IN R
Hone your skills!
UN SUP ERVISED L EARN IN G IN R

You might also like