0% found this document useful (0 votes)
7 views6 pages

K-Means Clustering Analysis for Students

The document outlines the application of the K-Means algorithm with k=3 to cluster students based on Hours Studied, Attendance, and Marks Scored. It details the cluster assignments, centroid calculations, and Euclidean distances for three iterations, ultimately identifying three distinct clusters representing different study patterns. The final cluster assignments show the convergence of the algorithm after the third iteration.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views6 pages

K-Means Clustering Analysis for Students

The document outlines the application of the K-Means algorithm with k=3 to cluster students based on Hours Studied, Attendance, and Marks Scored. It details the cluster assignments, centroid calculations, and Euclidean distances for three iterations, ultimately identifying three distinct clusters representing different study patterns. The final cluster assignments show the convergence of the algorithm after the third iteration.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

apply k means algorithm

To apply the K-Means algorithm (with k = 3) to the given dataset, the students are clustered
based on the features: Hours Studied, Attendance, and Marks Scored. Here are the cluster
assignments for each student:

Cluster Assignments
Student Hours Studied Attendance Marks Scored Cluster

S1 2 60 45 1

S2 3 65 50 1

S3 5 80 70 2

S4 8 85 78 0

S5 10 90 85 0

S6 4 70 55 1

S7 9 95 88 0

S8 6 75 65 2

S9 1 55 20 1

S10 7 80 72 2

This table shows which cluster (0, 1, or 2) each student belongs to, based on similarity in the
given features [1] .

Interpretation
Cluster 0: Contains students who have the highest hours studied, attendance, and marks
scored.
Cluster 1: Groups students with generally lower values in hours studied, attendance, and
marks.
Cluster 2: Contains students with moderate values in all three features.
These clusters help segment the students into similar study/attendance/marks patterns, as
identified by the K-Means algorithm [1] .

display the euclidean distance of each point from
the centroid. and display all the iterations and
centeroid calculation
Here is the step-by-step K-Means clustering process for your dataset with k=3, including
centroids and Euclidean distances for three iterations.

Iteration Details

Iteration 1
Centroids:
Hours_Studied Attendance Marks_Scored

5 80 70

7 80 72

9 95 88

Distances from each centroid (by student):


Student Centroid 1 Centroid 2 Centroid 3 Assigned Centroid

S1 32.16 33.97 55.88 1

S2 25.08 26.93 48.79 1

S3 0.00 2.83 23.77 1

S4 9.90 7.87 14.18 2

S5 18.71 16.67 5.92 3

S6 18.06 19.95 41.70 1

S7 23.77 22.02 0.00 3

S8 7.14 8.66 30.63 1

S9 56.04 58.01 79.30 1

S10 2.83 0.00 22.02 2

Iteration 2
Centroids:
Hours_Studied Attendance Marks_Scored

3.5 67.5 50.83

7.5 82.5 75.00

9.5 92.5 86.50

Distances:
Student Centroid 1 Centroid 2 Centroid 3 Assigned Centroid

S1 9.62 37.90 53.24 1

S2 2.68 30.85 46.16 1

S3 22.93 6.12 21.18 2

S4 32.63 3.94 11.43 2

S5 41.42 12.75 2.96 3

S6 4.88 23.84 39.10 1

S7 46.56 18.10 2.96 3

S8 16.22 12.59 27.94 2

S9 33.36 61.83 76.82 1

S10 24.83 3.94 19.31 2

Iteration 3
Centroids:
Hours_Studied Attendance Marks_Scored

2.5 62.5 42.50

6.5 80.0 71.25

9.5 92.5 86.50

Distances:
Student Centroid 1 Centroid 2 Centroid 3 Assigned Centroid

S1 3.57 33.31 53.24 1

S2 7.92 26.25 46.16 1

S3 32.69 1.95 21.18 2

S4 42.39 8.53 11.43 2

S5 51.17 17.36 2.96 3

S6 14.65 19.24 39.10 1

S7 56.29 22.62 2.96 3

S8 25.98 8.02 27.94 2

S9 23.76 57.29 76.82 1

S10 34.59 0.90 19.31 2


Final Cluster Assignments
Student Cluster

S1 1

S2 1

S3 2

S4 2

S5 3

S6 1

S7 3

S8 2

S9 1

S10 2

Each cell displays the Euclidean distance to the respective centroid for that iteration, and
assignments update according to the nearest centroid each time until convergence.

show the calcuations of centroid in each iteration


Here are the centroid calculations for each K-Means iteration, showing which data points belong
to each cluster and how the new centroid values are computed for every step:

Iteration 1
Centroid 1:
Members: S1, S2, S3, S6, S8, S9
Coordinates: , , , , , [3]

New centroid:
Hours Studied: (2 + 3 + 5 + 4 + 6 + 1) / 6 = 3.5
Attendance: (60 + 65 + 80 + 70 + 75 + 55) / 6 = 67.5
Marks: (45 + 50 + 70 + 55 + 65 + 20) / 6 ≈ 50.83
Centroid 2:
Members: S4, S10
Coordinates: ,
New centroid:
Hours Studied: (8 + 7) / 2 = 7.5
Attendance: (85 + 80) / 2 = 82.5
Marks: (78 + 72) / 2 = 75.0
Centroid 3:
Members: S5, S7
Coordinates: ,
New centroid:
Hours Studied: (10 + 9) / 2 = 9.5
Attendance: (90 + 95) / 2 = 92.5
Marks: (85 + 88) / 2 = 86.5

Iteration 2
Centroid 1:
Members: S1, S2, S6, S9
Coordinates: , , , [3]

New centroid:
Hours Studied: (2 + 3 + 4 + 1) / 4 = 2.5
Attendance: (60 + 65 + 70 + 55) / 4 = 62.5
Marks: (45 + 50 + 55 + 20) / 4 = 42.5
Centroid 2:
Members: S3, S4, S8, S10
Coordinates: , , ,
New centroid:
Hours Studied: (5 + 8 + 6 + 7) / 4 = 6.5
Attendance: (80 + 85 + 75 + 80) / 4 = 80.0
Marks: (70 + 78 + 65 + 72) / 4 = 71.25
Centroid 3:
Members: S5, S7
Coordinates: ,
New centroid: Same as previous: [9.5, 92.5, 86.5]

Iteration 3 (no further change, so K-Means converges)


Centroid 1: [2.5, 62.5, 42.5]
Centroid 2: [6.5, 80.0, 71.25]
Centroid 3: [9.5, 92.5, 86.5]
Assignments of members and values remain the same as Iteration 2, confirming convergence.
These step-by-step calculations show exactly how and why the cluster centroids move after
each K-Means iteration.

1. selected_image_3336899986120497906.jpg
2. selected_image_3336899986120497906.jpg
3. selected_image_3336899986120497906.jpg

You might also like