FREQUENCY
DISTRIBUTION
FREQUENCY DISTRIBUTION OR
HISTOGRAM
It is a graph plotting values
of observations on the
It is useful for assessing
horizontal axis, with a bar
properties of the
showing how many times
distribution of scores
each value occurred in the
data set.
NORMAL
DISTRIBUTION
• It is characterized by the bell-
shaped curve. This shape basically
implies that the majority of
scores lie around the center of
the distribution (so the largest
bars on the histogram are all
around the central value)
• As it gets further away from the
center the bars get smaller,
implying that as scores start to
deviate from the center their
frequency is decreasing
SKEWED DISTRIBUTIONS
• not symmetrical and instead the most frequent scores (the tall bars
on the graph) are clustered at one end of the scale
2 types of skewness:
• Positively skewed – the frequent scores are clustered at the lower
end and the tail points towards the higher or more positive scores
• Negatively skewed – he frequent scores are clustered at the higher
end and the tail points towards the lower or more negative scores
KURTOSIS
• kurtosis refers to the degree to which scores cluster at the ends of
the distribution (known as the tails)
• this tends to express itself in how pointy a distribution is
2 types of kurtosis
• Leptokurtic – a distribution with positive kurtosis has many scores
in the tails (a so-called heavy-tailed distribution) and is pointy
• Platykurtic – a distribution with negative kurtosis is relatively thin in
the tails (has light tails) and tends to be flatter than normal
H O W T O C R E AT E
A DISTRIBUTION
TA B L E ?
ACHIEVEMENT SCORES
96 74 64 50 76
83 80 92 85 91
59 68 76 75 69
64 87 71 81 83
73 67 68 70 75
COMPUTATIONS
1. Range: R = HS-LS +1
2. Interval Size: i = R / 10
3. Relative frequency: rf = f / N x 100
4. Cumulative proportion: cp = cf / N
5. Cumulative Percentage: CP = cp x100
GRAPHS
• Frequency Polygon
x axis: mid-point
y axis: frequency
• Histogram
x axis: class boundary of the upper limit
y axis: frequency
• Ogive
x axis: upper limit
y axis: Cumulative Percentage
UNGROUPED FREQUENCY DISTRIBUTION TABLE
Table 3
Distribution of a Sample of 1CSF Students According to Age
Age Frequency Percentage (%)
15 1 2
16 6 10
17 13 22
18 11 18
19 12 20
20 6 10
21 4 7
22 5 8
23 2 3
Total 60 100
Mean for Grouped Data
f X
x=
n
Where f – frequency of the class
X – Class Mark
n – sample size
13
Mean for Grouped Data
f X 659
x= = = 32.95
n 20
Classes Freq. (f) Class Mark fXm
(Xm)
12 - 22 4 17 68
23 - 33 7 28 196
34 - 44 6 39 234
45 - 55 2 50 100
56 – 66 1 61 61
Total 20 659
14
COEFFICIENT OF VARIATION
It is used when:
a. The means of the distributions being compared are
far apart, or
b. The data are in different units.
It is usually expressed in percent.
CV is also called measure of relative dispersion.
15
COEFFICIENT OF VARIATION
Example: The mean income of a sample of homeowners in Barangay
A is P40,000 and the standard deviation is P4,000. In Barangay B,
the mean income of a sample of homeowners is P12,000 and the
standard deviation is P1,200. Compare the relative dispersion in the
two groups of incomes.
Solution: Barangay A Barangay B
s 4000 1200
CV = (100) = (100) = 10% CV = (100) = 10%
X 40000 12000
Interpretation: The incomes in both barangays are spread out 10%
from their respective means.
16
COEFFICIENT OF VARIATION
Problem: We will now compare the dispersion in the ages of the
homeowners in Barangay A with the dispersion in their incomes. The
mean age of the sample of homeowners is 40 years with a standard
deviation of 10 years. Recall that for their incomes, the mean is
P40,000 and the standard deviation is 4000.
Solution:
Income Age
4000 10
CV = (100) = 10% CV = (100) = 25%
40000 40
Interpretation: There is greater relative dispersion in the ages of
the homeowners in Barangay A than in their incomes.
17
COEFFICIENT OF VARIATION
Example: The table shows a player’s record of assists and points
in the 10 randomly-selected games which he played. Determine
in which area did he perform more consistently.
game 1 2 3 4 5 6 7 8 9 10 x s
A 7 10 9 1 5 3 4 7 9 4 5.9 2.96
P 25 25 30 22 23 22 16 35 20 18 23.6 5.60
18
MEASURES OF SKEWNESS
• Skewness is a measure of symmetry, or more precisely, the lack of symmetry. A distribution, or
data set, is symmetric if it looks the same to the left and right of the center point.
• The histogram is an effective graphical technique for showing both the skewness and kurtosis
of data set
KARL PEARSON’S MEASURE OF
SKEWNESS
• Pearson’s Coefficient of Skewness #1 uses • Pearson’s Coefficient of Skewness #2 uses
the mode. The formula is: the median. The formula is:
KELLY BOWLEY’S MEASURE OF
SKEWNESS
• Bowley skewness is a way to figure out if you have a positively-skewed or negatively skewed
distribution. One of the most popular ways to find skewness is the Pearson Mode Skewness
formula. However, in order to use it you must know the mean, mode (or median) and standard
deviation for your data.
SKEWNESS: INTERPRETATION
In general:
– The direction of skewness is given by the sign.
– The coefficient compares the sample distribution with a normal distribution. The larger the value,
the larger the distribution differs from a normal distribution.
– A value of zero means no skewness at all.
– A large negative value means the distribution is negatively skewed.
– A large positive value means the distribution is positively skewed.
Z-SCORES
COMPUTING FOR Z-SCORES
observation-mean
Z=____________________
sd