Chapter 2
Chapter 2
DESCRIPTIVE STATISTICS
1
TABLE OF CONTENTS
2
2.1. Summarizing Data for a Categorical Variable
- Table Statistics;
- Chart Statistics.
3
TABLE STATISTICS
Table statistics for one categorical Variable:
Example:
Having the data of 12 customers according to gender and region.
5
TABLE STATISTICS
Table statistics for one categorical Variable:
Table 2: Customers according to Region
7
TABLE STATISTICS
Table statistics for two categorical variables:
Example:
Having the data of 12 customers according to gender and region.
10
CHARTS STATISTICS
Displaying by charts for one categorical Variable:
The Pie Chart
Region
2.5
1.5
0.5
Male Female
The South The Central The North
12
CHARTS STATISTICS
Displaying by charts for one categorical variable:
The Bar Chart
6
0
The South The Central The North
13
CHARTS STATISTICS - SOFTWARE
Table Statistics;
Chart Statistics.
15
TABLE STATISTICS
Table statistics for one quantitative variable:
Discrete variable and Continuous variable
Example: Having the data of 20 Customers according to worker rank
and Income.
No. Worker Rank Income (USD/a week) No. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160
16
TABLE STATISTICS
Table statistics for one quantitative variable:
Worker rank: Discrete variable
Table 5: Customers according to Worker rank
Worker rank Frequency (Person) Percentage Cumulative Percentage
1.00 2 10.0 10.0
2.00 3 15.0 25.0
3.00 2 10.0 35.0
4.00 4 20.0 55.0
5.00 4 20.0 75.0
6.00 3 15.0 90.0
7.00 2 10.0 100.0
Total 20 100.0
17
TABLE STATISTICS
Table statistics for one quantitative variable:
Data Transform: Continuous Variable
Table 6: Customers according to Income
Income (USD/a week) Frequency (Person) Percentage Cumulative Percentage
70.00 1 5.0 5.0
75.00 2 10.0 15.0
85.00 1 5.0 20.0
It is very 90.00 1 5.0 25.0
difficult 110.00
120.00
1
2
5.0
10.0
30.0
40.0
for data 160.00
165.00
1
1
5.0
5.0
45.0
50.0
summary 173.00 1 5.0 55.0
175.00 1 5.0 60.0
180.00 1 5.0 65.0
190.00 1 5.0 70.0
195.00 1 5.0 75.0
200.00 1 5.0 80.0
205.00 2 10.0 90.0
215.00 1 5.0 95.0
220.00 1 5.0 100.0
Total 20 100.0
18
TABLE STATISTICS
Table statistics for one quantitative variable:
Transform Data: Continuous Variable => Discrete variable
Step 1: Approximate Classes - groups (h)
19
TABLE STATISTICS
Table statistics for one quantitative variable:
20
TABLE STATISTICS
Table statistics for one quantitative variable:
Transform Data: Continuous Variable => Discrete variable
Step 1: Approximate Class (h)
22
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Dot Plot chart
23
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Histogram chart
Income
250
200
150
100
50
0
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
25
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Stem-and-Leaf Display Chart
26
CHARTS STATISTICS
Displaying by charts for two quantitative Variable:
Plot -Trend chart
INCOME
250
200
150
100
50
0
0 1 2 3 4 5 6 7 8
27
2.3. Summarising data for two quantitative
variable using tables and graphics
Table Statistics;
Chart Statistics.
28
TABLE STATISTICS
Table statistics for two quantitative variables:
No. Worker Rank Income (USD/a week) No. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160
29
TABLE STATISTICS
Table statistics for two quantitative variables:
Discrete variable and Continuous Variable
Example:
Having the data of 20 Customers according to worker rank and Income
(transform).
Income (USD/A week)
Variables Total
From 70 to 120 From 120 to 170 From 170 to 220
1.00 2 0 0 2
2.00 3 0 0 3
3.00 1 1 0 2
Worker
4.00 0 2 2 4
Rank
5.00 0 0 4 4
6.00 0 1 2 3
7.00 0 0 2 2
30
Total 6 4 10 20
CHARTS STATISTICS
Displaying by charts for two Variables:
Bar Chart
8
Male Female
The South The Central The North
31
CHARTS STATISTICS
Displaying by charts for two Variables:
Bar Chart
3.5
2.5
1.5
0.5
32
2.4. Summarizing data for a quantitative
variable using numerical measures
- Measures of center location (measures of central
tendency)
- Measures of dispersion (measures of variability)
- Quartiles and box plot
33
CENTRAL TENDENCY- VARIABILITY
VARIABILITY
CENTRAL TENDENCY
Mean
Range Mean absolute
deviation
Median
Variance Coefficient of
Mode
Variation
Standard
Deviation
34
MEAN
Mean:
The mean of a data set is the average of all
the data values.
How to use the Mean?
Comparing multiple groups together.
35
MEAN
Example:
Who has higher income: Americans or Vietnamese?
Americans Vietnamese
Income of Income of
Comparing
a person a person
Right/wrong?
36
MEAN
Example:
Who has higher income: Americans or Vietnamese?
American
Vietnamese
s
Mean Mean
Comparing
of income of income
Right/wrong
?
37
MEAN - Population Mean (μ)
Data without frequency:
38
MEAN - Sample Mean
Data without frequency:
39
SAMPLE MEAN
Example:
Having the data of 10 Customers according to worker rank and
Income:
No. 1 2 3 4 5 6 7 8 9 10
Worker Rank 1 5 4 5 3 2 6 1 2 5
Income
70 195 165 205 110 85 205 75 90 190
(USD/a week)
40
SAMPLE MEAN
- The worker rank mean value:
41
MEAN - Population Mean (μ)
Data with frequency:
43
MEAN - Sample Mean (μ)
Example: Having the data of 10 Customers according to worker rank
Worker Rank Frequency (Person)
No. Percentage Cumulative Percentage
(xi) (fi)
1 1.00 (X1) 2 (f1) 10.0 20.0
2 2.00 2 10.0 40.0
3 3.00 1 5.0 50.0
4 4.00 1 5.0 60.0
5 5.00 3 15.0 90.0
6 6.00 (X6) 1 (f6) 5.0 100.0
Total 10 50.0
- Compute the mean of worker rank.
44
MEAN - Sample Mean (μ)
- The mean of worker rank:
45
WEIGHTED MEAN
Example: Having the Data in an account buying 3 stocks:
46
GEOMETRIC MEAN
47
GEOMETRIC MEAN
- Example: Rate of return
48
MODE (Mo)
- One Mode;
- Multimodal.
49
MODE (Mo)
50
MODE (Mo)
51
MEDIAN (Me)
52
MEDIAN (Me)
-Median:
54
MEDIAN (Me)
For an odd number of observations:
Example: For the following data
26 18 27 12 14 27 19 7 observations
55
MEDIAN (Me)
For an odd number of observations:
Example: For the following data
26 18 27 12 14 27 19 7 observations
- Step 1: in ascending order
12 14 18 19 26 27 27
57
MEDIAN (Me)
For an even number of observations:
Example : For the following data
26 18 27 12 14 27 30 19 8 observations
The median is the average of the middle two values.
58
MEDIAN (Me)
For an even number of observations:
- Step 1: Arranged in ascending order
26 18 27 12 14 27 30 19 8 observations
12 14 18 19 26 27 27 30
Step 2: The median is the average of the middle
two values.
𝒏 𝒏 𝟖 𝟖
𝟐 𝟏 𝟐 𝟏 𝟒 𝟓
𝟐
= 𝟐
=
59
MEASURES OF CENTRAL TENDENCY - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…
60
MEASURES OF CENTRAL TENDENCY
• Example:
Having the data of 10 Customers according to worker rank and Income.
Statistics
N
Mean Median Mode
Valid
Worker Rank 10 3.4 3.5 5
Income 10 139 137.5 205
62
CENTRAL TENDENCY- VARIABILITY
VARIABILITY
CENTRAL TENDENCY
Mean
Range Mean absolute
deviation
Median
Variance Coefficient of
Mode
Variation
Standard
Deviation
63
RANGE
R=xmax-xmin
64
RANGE
• Example: Having the data of 10 Customers according to worker rank
and Income.
N0. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190
66
MEAN ABSOLUTE DEVIATION
• Example: Having the data of 10 Customers according to worker rank
and Income.
N0. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190
- Compute the Mean absolute deviation for worker rank and Income.
67
VARIANCE
- Variance: the average of the squared differences
between each data value and the mean.
68
VARIANCE – POPULATION
- Data without frequency:
72
VARIANCE - STANDARD DEVIATION
- Variance and Standard deviation for worker rank:
N0. Worker Rank (xi) (𝒙𝒊 − 𝒙) (𝒙𝒊 − 𝒙)𝟐
1 1 -2.4 5.76
2 5 1.6 2.56
3 4 0.6 0.36
4 5 1.6 2.56
5 3 -0.4 0.16
6 2 -1.4 1.96
7 6 2.6 6.76
8 1 -2.4 5.76
9 2 -1.4 1.96
10 5 1.6 2.56
Mean 3.4 Total 30.4
=>
73
MEASURES OF VARIABILITY - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…
74
MEASURES OF VARIABILITY
• Example: Having the data of 10 Customers according to worker rank and Income.
No. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190
- Compute the Range; Variance and Standard deviation for worker rank and
Income with SPSS.
75
MEASURES OF VARIABILITY - SPSS
Statistics
Worker Rank Income
N Valid 10 10
Std. Deviation (Standard deviation) 1.84 57.87
Variance 3.38 3348.89
Range 5.00 135.00
Minimum 1.00 70.00
Maximum 6.00 205.00
76
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation in Stock market:
• Example: Have the data about the Closing price of 3 stock in 10 days
(1000 VND).
Day 1 2 3 4 5 6 7 8 9 10
A 18.7 18.2 18.3 18.5 18.8 18.5 18.6 19.4 19.3 19.5
Stock B 17.8 18.85 19.2 20.4 20.25 20.05 21.2 21.3 21.45 21.7
C 17.1 18.15 18.4 19.6 20.45 19.55 20.8 22.2 21.45 22.7
77
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation:
78
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation in services quality:
• Example: The result of quality Rank for Support Services
Services
Standard
No. Code Content Variance quality
deviation
Rank
There are numerous
1 SUP1 2.50 1.581 3
language support services
There are numerous
2 SUP2 0.80 0.894 1
academic support services
There are various social
3 SUP3 1.30 1.140 2
support services
79
COEFFICIENT OF VARIATION
80
COEFFICIENT OF VARIATION
- Coefficient of variation for a Population:
82
QUARTILES - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…
83
2.5. MEASURES OF DISTRIBUTION SHAPE
84
SKEWNESS
- Skewness measure of the shape of a data distribution.
- Skewness value:
+ Skewness =0 => Data is Symmetric distribution;
+ Skewness >0 => Data is right skew distribution;
+ Skewness >0 => Data is left skew distribution. 85
KURTOSIS
- Kurtosis measure of the shape of a data distribution.
- Kurtosis value:
+ Kurtosis ≈ 0 => Data is Normal distribution;
+ Kurtosis ≠ 0 => Data is not Normal distribution.
86
SKEWNESS AND KURTOSIS - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…
87
SUMMARY
Example: Having the data of 20 Customers according to worker rank and Income.
N0. Worker Rank Income (USD/a week) N0. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160
88
SUMMARY
Statistics
Worker Rank Income
N Valid 20 20
Mean 4.1000 151.4000
Median 4.0000 169.0000
Mode 4.00a 75.00
Std. Deviation 1.86096 52.57316
Variance 3.463 2763.937
Skewness -.161 -.360
Kurtosis -.954 -1.475
Range 6.00 150.00
Minimum 1.00 70.00
Maximum 7.00 220.00
25 2.2500 95.0000
Percentiles 50 4.0000 169.0000
75 5.7500 198.7500
89
2.6. Measures of association between two quantitative variables
90
COVARIANCE
- The covariance is a measure of the linear association between two
variables.
- The covariance value:
+ Covariance value = 0 => Two variables are not the linear
association:
+ Covariance value > 0 => The linear association between two
variables are Positive:
X (Y) - grow (reduce) => Y(X) - grow (reduce)
+ Covariance value < 0 => The linear association between two
variables are Negative:
X (Y) - reduce (grow) => Y(X) - grow (reduce)
91
COVARIANCE
- The covariance is computed
𝒊 𝒊 𝒊 𝒊
r 𝟐 𝟐
ρ 𝒙
𝟐
𝒚
𝟐
𝒊 𝒊 𝒊 𝒙 𝒊 𝒚
For Samples For Populations
94
COVARIANCE - CORRELATION COEFFICIENT
Example: The following data show the daily increase percent or daily
decrease percent in the DJIA and S&P 500 for a sample of nine days
over a three-month period (The Wall Street Journal, January 15 to
March 10, 2006) (9 days).
DJIA 0.20 0.82 -0.99 0.04 -0.24 1.01 0.30 0.55 -0.25
S&P 500 0.24 0.19 -0.91 0.08 -0.33 0.87 0.36 0.83 -0.16
95
COVARIANCE - CORRELATION COEFFICIENT
Xi Yi
0.2 0.24 0.04 0.11 0.002 0.012 0.004
0.82 0.19 0.66 0.06 0.436 0.004 0.040
-0.99 -0.91 -1.15 -1.04 1.323 1.082 1.196
0.04 0.08 -0.12 -0.05 0.014 0.003 0.006
-0.24 -0.33 -0.4 -0.46 0.160 0.212 0.184
1.01 0.87 0.85 0.74 0.723 0.548 0.629
0.3 0.36 0.14 0.23 0.020 0.053 0.032
0.55 0.83 0.39 0.7 0.152 0.490 0.273
-0.25 -0.16 -0.41 -0.29 0.168 0.084 0.119
0.16 0.13 2.996 2.486 2.483 96
COVARIANCE - CORRELATION COEFFICIENT
- The covariance of sample between DJIA and S&P 500:
𝒊
𝒙𝒚
Sxy=0.31>0 => The linear association between DJIA and S&P are positive
- The correlation coefficient of sample between DJIA and S&P 500:
2.483
r =
2.996∗2.486
=0.91
97
2.7. MEASURES OF ASSOCIATION BETWEEN TWO
QUALITATIVE VARIABLES
- Cramer's V (V).
- Contingency Coefficient (C)
98
CRAMER'S V
- The Cramer's (V) value:
+ Cramer's (V) = 0 => Two qualitative variables are not
correlated.
+ Cramer's (V) near 1 => Two qualitative variables are strongly
correlated.
99
CRAMER'S V
- Cramer's V (V)
( )
;
∗
101
CONTINGENCY COEFFICIENT
Cramer's V (V)
( ) ∗
;
n: Number of observations in the sample
fij: Frequency of the row i and the column j
ni: Number of observations in row i.
nj: Number of observations in row j.
k: Number of rows
m: Number of columns
102
CRAMER'S - CONTINGENCY COEFFICIENT
SPSS: Analyze => Descriptive statistics => Crosstabs =>Statistics
103
THE END
104