Descriptive
Statistics
Descriptive Statistics
• Here we are interested only in summarizing the data in front
of us, without assuming that it represents anything more.
• We will look at both numerical and categorical data..
• The basic overall idea is to turn data into information.
Numerical Data – Single Variable
• Measures of Location
• Measures of central tendency: Mean, Median, Mode
• Measure of noncentral tendency – Quantiles
• Measure of Dispersion
• Range
• Interquartile range
• Varience
• Standard Deviation
• Coefficient of Variation
• Skewness
Measure of Central Tendency?
• The value which provides a summary of the characteristics of
a given set of data.
• Tells us about the shape and nature of the distribution.
• These include the mean, the median, and the mode.
Ungrouped
Data
The Mean
• The score located at the mathematical center of a
distribution.
• The sample mean is the sum of all the observations (∑Xi)
divided by the number of observations (n):
∑ =1
¯ = where ΣXi = X1 + X2 + X3 + X4 + … + Xn
𝑛
𝑋
𝑖
𝑋
𝑖
𝑛
Example. 1, 2, 2, 4, 5, 10. Calculate the
mean. Note: n = 6 (six observations)
∑Xi = 1 + 2+ 2+ 4 + 5 + 10 = 24
¯ = 24 / 6 = 4.0
𝑋
Example. For the data: 1, 1, 1, 1, 51. Calculate the mean. Note: n = 5
(five observations)
∑Xi = 1 + 1+ 1+ 1+ 51 = 55
¯ = 55 / 5 = 11.0
• Here we see that the mean is affected by extreme values.
𝑋
Properties of Arithmetic Mean
• All values are used
• It is unique
• It is easy to calculate
• The sum of the deviations from the mean is zero. (the only
measure of central tendency)
• It is easily affected by extremes, such as very big or small
numbers in the set (non-robust)
The sum of the deviations from the mean is zero:
Values Deviations
3 -2
4 -1
8 3
Mean 5 0
Weighted Mean
• Arithmetic mean computed by considering relative
importance of each items is called weighted arithmetic mean.
• To give due importance to each item under consideration,
assign number called weight to each item in proportion to its
relative importance.
Example
• Find the Weighted Arithmetic Mean of the given example
below:
Grade
Subjects Weight (W)
Obtained (X)
AE 57 1.0 3
AE 53 1.25 3
AE 130 1.75 3
Physics 41 1.25 4
Math 71 1.75 6
Total
The Median
• The middle value (50th percentile) of the ordered data.
• Therefore, the median is the data value such that half of the
observations are larger and half are smaller.
Determining the Median
• When data are normally distributed, the median is the same
score as the mode.
• When data are not normally distributed, follow the following
procedure:
• Arrange the data into an ordered array (ascending or descending
order)
• If n is odd, the median is the middle observation of the ordered
array. If n is even, it is the average of the two middle observations.
Example 0 2 3 5 20 99 100
Note: Data has been ordered from lowest to highest. Since n is
odd (n=7), the median is the (n+1)/2 ordered observation, or the
4th observation.
Answer: The median is 5.
Unlike the mean, the median is not affected by extreme values.
Q: What happens to the median if we change the 100 to 5,000?
Example 10 20 30 40 50 60
Note: Data has been ordered from lowest to highest. Since n is
even (n=6), the median is the (n+1)/2 ordered observation, or
the 3.5th observation, i.e., the average of observation 3 and
observation 4.
Answer: The median is 35.
• The mean and the median are unique for a given set of data.
There will be exactly one mean and one median.
The Median
• The median has 3 interesting characteristics:
• 1. The median is not affected by extreme values, only by the
number of observations.
• 2. Any observation selected at random is just as likely to be
greater than the median or less than the median.
• 3. Summation of the absolute value of the differences about the
median is a minimum:
∑
− = minimum
=0
𝑖
𝑖
𝑋
𝑀
𝑒
𝑑
𝑖
𝑎
𝑛
Descriptive Statistics I 17
𝑛
The Mode
• The mode is the value of the data that occurs with the greatest
frequency.
• Gives use limited information about a distribution.
• Typically useful in describing central tendency when the data
reflect a nominal scale of measurement.
• It does not make sense to take the average in nominal data.
• Gender: 67 males & 50 females
Example:
• 1, 1, 1, 2, 3, 4, 5
Answer. The mode is 1 since it occurs three times. The other
values each appear only once in the data set.
• 5, 5, 5, 6, 8, 10, 10, 10.
Answer. The mode is: 5, 10.
There are two modes. This is a bi-modal dataset.
Properties of the Mode
• It can be computed for nominal, ordinal, interval or ratio level.
• It is very unstable – it can change radically if the method of
rounding the data is changed.
• The mode may not exist.
• Data: 1, 2, 3, 4, 5, 6, 7, 8, 9, 0
• The mode may not be unique.
• Data: 0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7
• Mode = 1, 2, 3, 4, 5, and 6. There are six modes.
Grouped Data
Calculation of Central Tendency for Grouped
Data
Weight of Zea Maize Ear Number of Zea Maize Ear
40 – 59 6
60 – 79 28
80 – 99 35
100 – 119 50
120 – 139 30
140 – 159 10
160 – 179 12
180 – 199 9
Total
Weight of Zea Number of Zea
Maize Ear Maize Ear
40 – 59 6
60 – 79 28
80 – 99 35
100 – 119 50
120 – 139 30
140 – 159 10
160 – 179 12
180 – 199 9
Total
Weight of Zea Number of Zea
Maize Ear Maize Ear
40 – 59 6
60 – 79 28
80 – 99 35
100 – 119 50
120 – 139 30
140 – 159 10
160 – 179 12
180 – 199 9
Total
Shapes
• A third important property of data is its shape.
• Shape can be described by degree of asymmetry
• Mean > median positive or right – skewness
• Mean = median symmetric or zero - skewness
• Mean < median negative or left – skewness
Positively Skewed Distribution
• In a positively skewed distribution, most of the data values
fall to the left of the mean, and the “tail” of the distribution is
to the right.
Negatively Skewed Distribution
• In a negatively skewed distribution, most of the data values
fall to the right of the mean, and the tail of the distribution is
to the left.
Normal or Symmetrical Distribution
• In a symmetrical distribution, the data values are evenly
distributed on both sides of the mean.
Designed with by
[Link]
The free PowerPoint template library