0% found this document useful (0 votes)
3 views104 pages

Chapter 2

Chapter 2 covers descriptive statistics, detailing methods for summarizing data for both categorical and quantitative variables using tables and graphics. It includes examples of frequency distributions, measures of central tendency, and variability, along with visual representations like pie charts and histograms. The chapter also discusses the analysis of relationships between two variables and the transformation of data types.

Uploaded by

ngthanhthu2121
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views104 pages

Chapter 2

Chapter 2 covers descriptive statistics, detailing methods for summarizing data for both categorical and quantitative variables using tables and graphics. It includes examples of frequency distributions, measures of central tendency, and variability, along with visual representations like pie charts and histograms. The chapter also discusses the analysis of relationships between two variables and the transformation of data types.

Uploaded by

ngthanhthu2121
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 2

DESCRIPTIVE STATISTICS

1
TABLE OF CONTENTS

2.1. Summarizing data for a categorical variable


2.2. Summarizing data for a quantitative variable using tables and graphics
2.3. Summarizing data for two variables using tables and graphics
2.4. Summarizing data for a quantitative variable using numerical measures
2.5. Measures of distribution shape
2.6. Measures of association between two quantitative variables
2.7. Measures of association between two qualitative variables

2
2.1. Summarizing Data for a Categorical Variable

- Table Statistics;
- Chart Statistics.

3
TABLE STATISTICS
Table statistics for one categorical Variable:
Example:
Having the data of 12 customers according to gender and region.

No. Gender Region No. Gender Region


1 Male The North 7 Female The South
2 Male The Central 8 Male The South
3 Female The South 9 Female The Central
4 Female The Central 10 Male The Central
5 Male The Central 11 Female The North
6 Male The South 12 Male The North
4
TABLE STATISTICS
Table statistics for one categorical Variable:
Table 1: Customers According to Gender

Frequency Percentage Cumulative Percentage


Gender
(Person) (%) (%)
Male 7 58.3 58.3
Female 5 41.7 100.0
Total 12 100.0

5
TABLE STATISTICS
Table statistics for one categorical Variable:
Table 2: Customers according to Region

Frequency Percentage Cumulative Percentage


Region
(Person) (%) (%)
The South 4 33.3 33.3
The Central 5 41.7 75.0
The North 3 25.0 100.0
Total 12 100.0
6
TABLE STATISTICS
Table statistics for one categorical Variable:
SPSS: Analyze => Descriptive statistics => Frequencies

7
TABLE STATISTICS
Table statistics for two categorical variables:
Example:
Having the data of 12 customers according to gender and region.

No. Gender Region No. Gender Region


1 Male The North 7 Female The South
2 Male The Central 8 Male The South
3 Female The South 9 Female The Central
4 Female The Central 10 Male The Central
5 Male The Central 11 Female The North
6 Male The South 12 Male The North
8
TABLE STATISTICS
Table statistics for one categorical variable:
Table 4: Customers according to Region and Gender
Unit: % within Region
Region
Variables The The Total
The South
Central North
Male 50.0 60.0 66.7 58.3
Gender
Female 50.0 40.0 33.3 41.7
Total 100.0 100.0 100.0 100.0
9
TABLE STATISTICS
Table statistics for two categorical Variable:
SPSS: Analyze => Descriptive statistics => Crosstabs

10
CHARTS STATISTICS
Displaying by charts for one categorical Variable:
The Pie Chart
Region

The South The Central The North


11
CHARTS STATISTICS
Displaying by charts for two categorical variables:
The Bar Chart
3.5

2.5

1.5

0.5

Male Female
The South The Central The North

12
CHARTS STATISTICS
Displaying by charts for one categorical variable:
The Bar Chart
6

0
The South The Central The North

13
CHARTS STATISTICS - SOFTWARE

- MSExcel software: Insert


=> Chart;
- MSWord software: Insert
=> Chart.
14
2.2. Summarizing data for a quantitative variable
using tables and graphics

 Table Statistics;
 Chart Statistics.

15
TABLE STATISTICS
Table statistics for one quantitative variable:
Discrete variable and Continuous variable
Example: Having the data of 20 Customers according to worker rank
and Income.

No. Worker Rank Income (USD/a week) No. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160

16
TABLE STATISTICS
Table statistics for one quantitative variable:
Worker rank: Discrete variable
Table 5: Customers according to Worker rank
Worker rank Frequency (Person) Percentage Cumulative Percentage
1.00 2 10.0 10.0
2.00 3 15.0 25.0
3.00 2 10.0 35.0
4.00 4 20.0 55.0
5.00 4 20.0 75.0
6.00 3 15.0 90.0
7.00 2 10.0 100.0
Total 20 100.0
17
TABLE STATISTICS
Table statistics for one quantitative variable:
Data Transform: Continuous Variable
Table 6: Customers according to Income
Income (USD/a week) Frequency (Person) Percentage Cumulative Percentage
70.00 1 5.0 5.0
75.00 2 10.0 15.0
85.00 1 5.0 20.0
It is very 90.00 1 5.0 25.0
difficult 110.00
120.00
1
2
5.0
10.0
30.0
40.0
for data 160.00
165.00
1
1
5.0
5.0
45.0
50.0
summary 173.00 1 5.0 55.0
175.00 1 5.0 60.0
180.00 1 5.0 65.0
190.00 1 5.0 70.0
195.00 1 5.0 75.0
200.00 1 5.0 80.0
205.00 2 10.0 90.0
215.00 1 5.0 95.0
220.00 1 5.0 100.0
Total 20 100.0
18
TABLE STATISTICS
Table statistics for one quantitative variable:
Transform Data: Continuous Variable => Discrete variable
Step 1: Approximate Classes - groups (h)

xmax: Maximums data value; xmin: Minimums data value;


K: Number of Classes - groups.
Note: Nest distance (h) rounded up to the last unit of measurement (3.4=>4
(person))
Step 2: Frequency distribution with new Discrete variable

19
TABLE STATISTICS
Table statistics for one quantitative variable:

Transform Data: Continuous Variable => Discrete variable


Example: Transform the income data of 20 Customers into 3 groups.

No. Income (USD/a week) No. Income (USD/a week)


1 70 11 180
2 195 12 220
3 165 13 215
4 205 14 200
5 110 15 173
6 85 16 120
7 205 17 175
8 75 18 75
9 90 19 120
10 190 20 160

20
TABLE STATISTICS
Table statistics for one quantitative variable:
Transform Data: Continuous Variable => Discrete variable
Step 1: Approximate Class (h)

Step 2: Frequency distribution with new Discrete variable


Table 7: Customers according to Income

No. Income (USD/A week) Frequency (Person) Percentage Cumulative Percentage


1 From 70 to 120 6 30.0 30.0
2 From 120 to 170 4 20.0 50.0
3 From 170 to 220 10 50.0 100.0
Total 20 100.0
Note: Observe Values is above the nest value are placed in the larger nest.
Quan sát Giá trị nằm trên giá trị tổ thì xếp vào trong tổ lớn hơn
21
TABLE STATISTICS
Table statistics for one quantitative variable:
Transform Data: Continuous Variable => Discrete variable
Note on Number of Classes and Class Width:
+ In practice, number of classes and appropriate class width are determined by trial
and error.
+ Approximate number of Classes –groups:
n: Number of observations
Example: Approximate number of classes groups for 20 customers:

22
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Dot Plot chart

The distribution of Customers according to Income

23
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Histogram chart

SPSS: Analyze => Descriptive statistics =>


Frequencies=>Chart=>… 24
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Line chart

Income
250

200

150

100

50

0
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20

25
CHARTS STATISTICS
Displaying by charts for one quantitative Variable:
Stem-and-Leaf Display Chart

26
CHARTS STATISTICS
Displaying by charts for two quantitative Variable:
Plot -Trend chart
INCOME
250

200

150

100

50

0
0 1 2 3 4 5 6 7 8

27
2.3. Summarising data for two quantitative
variable using tables and graphics

 Table Statistics;
 Chart Statistics.

28
TABLE STATISTICS
Table statistics for two quantitative variables:

Table 7: Customers according to Worker rank and Income


(Transform)

No. Worker Rank Income (USD/a week) No. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160

29
TABLE STATISTICS
Table statistics for two quantitative variables:
Discrete variable and Continuous Variable
Example:
Having the data of 20 Customers according to worker rank and Income
(transform).
Income (USD/A week)
Variables Total
From 70 to 120 From 120 to 170 From 170 to 220
1.00 2 0 0 2
2.00 3 0 0 3
3.00 1 1 0 2
Worker
4.00 0 2 2 4
Rank
5.00 0 0 4 4
6.00 0 1 2 3
7.00 0 0 2 2
30
Total 6 4 10 20
CHARTS STATISTICS
Displaying by charts for two Variables:
Bar Chart
8

Male Female
The South The Central The North

31
CHARTS STATISTICS
Displaying by charts for two Variables:
Bar Chart
3.5

2.5

1.5

0.5

The South The Central The North


Male Female

32
2.4. Summarizing data for a quantitative
variable using numerical measures
- Measures of center location (measures of central
tendency)
- Measures of dispersion (measures of variability)
- Quartiles and box plot

33
CENTRAL TENDENCY- VARIABILITY

VARIABILITY
CENTRAL TENDENCY

Mean
Range Mean absolute
deviation
Median
Variance Coefficient of
Mode
Variation

Standard
Deviation
34
MEAN
Mean:
 The mean of a data set is the average of all
the data values.
How to use the Mean?
 Comparing multiple groups together.

35
MEAN
Example:
Who has higher income: Americans or Vietnamese?

Americans Vietnamese

Income of Income of
Comparing
a person a person
Right/wrong?
36
MEAN
Example:
Who has higher income: Americans or Vietnamese?

American
Vietnamese
s
Mean Mean
Comparing
of income of income
Right/wrong
?
37
MEAN - Population Mean (μ)
Data without frequency:

38
MEAN - Sample Mean
Data without frequency:

39
SAMPLE MEAN
Example:
Having the data of 10 Customers according to worker rank and
Income:

No. 1 2 3 4 5 6 7 8 9 10
Worker Rank 1 5 4 5 3 2 6 1 2 5
Income
70 195 165 205 110 85 205 75 90 190
(USD/a week)

Compute the mean value for worker rank and income.

40
SAMPLE MEAN
- The worker rank mean value:

- The income mean value:

41
MEAN - Population Mean (μ)
Data with frequency:

Xi: The value of the observation i.


Fi: The frequency of the observation i.
N: Number of observations in the population.
42
MEAN - Sample Mean (μ)
Data with frequency:

xi: The value of the observation i.


fi: The frequency of the observation i.
n: Number of observations in the sample.

43
MEAN - Sample Mean (μ)
Example: Having the data of 10 Customers according to worker rank
Worker Rank Frequency (Person)
No. Percentage Cumulative Percentage
(xi) (fi)
1 1.00 (X1) 2 (f1) 10.0 20.0
2 2.00 2 10.0 40.0
3 3.00 1 5.0 50.0
4 4.00 1 5.0 60.0
5 5.00 3 15.0 90.0
6 6.00 (X6) 1 (f6) 5.0 100.0
Total 10 50.0
- Compute the mean of worker rank.
44
MEAN - Sample Mean (μ)
- The mean of worker rank:

45
WEIGHTED MEAN
Example: Having the Data in an account buying 3 stocks:

No. Stock symbol Volume (Wi) Price (1000 VNN) (xi)


1 HPG 500 36.5
2 MSB 700 14.5
3 NVL 800 18.2

=> The mean price of stock:

46
GEOMETRIC MEAN

 It is often used in analyzing growth rates in financial data;


 It should be applied anytime you want to determine the mean rate of change
over several successive periods (be it years, quarters, weeks, . . .).

47
GEOMETRIC MEAN
- Example: Rate of return

48
MODE (Mo)

- One Mode;
- Multimodal.

49
MODE (Mo)

- Mode: The mode of a data set is the value with


greatest frequency.
- Mo=XMo  fMo=Max(fi)
- Example : For the following data : 4; 8; 10; 8; 9
10; 8
=> The mode is 8  F(8)=3 => Max(fi)

50
MODE (Mo)

- If the data have more than two modes, the


data are multimodal
- Example: For the following data: 4; 8; 10;
10; 15; 17; 17.
 Having two modes: 10 and 17

51
MEDIAN (Me)

-For an odd number of observations;

-For an even number of observations.

52
MEDIAN (Me)

-Median:

 The median of a data set is the value in the


middle when the data items are arranged in
ascending order.
53
MEDIAN (Me)

For an odd number of observations:


- Step 1: Arranged in ascending order;
- Step 2: The median is the middle value:

54
MEDIAN (Me)
 For an odd number of observations:
Example: For the following data
26 18 27 12 14 27 19 7 observations

- The median is the middle value:

55
MEDIAN (Me)
For an odd number of observations:
Example: For the following data
26 18 27 12 14 27 19 7 observations
- Step 1: in ascending order

12 14 18 19 26 27 27

- Step 2: The median is the middle value:


=
56
MEDIAN (Me)
For an even number of observations:
- Step 1: Arranged in ascending order;
- Step 2: The median is the average of the middle
two values.

57
MEDIAN (Me)
 For an even number of observations:
Example : For the following data

26 18 27 12 14 27 30 19 8 observations
The median is the average of the middle two values.

58
MEDIAN (Me)
For an even number of observations:
- Step 1: Arranged in ascending order
26 18 27 12 14 27 30 19 8 observations

12 14 18 19 26 27 27 30
Step 2: The median is the average of the middle
two values.
𝒏 𝒏 𝟖 𝟖
𝟐 𝟏 𝟐 𝟏 𝟒 𝟓
𝟐
= 𝟐
=
59
MEASURES OF CENTRAL TENDENCY - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…

60
MEASURES OF CENTRAL TENDENCY
• Example:
Having the data of 10 Customers according to worker rank and Income.

No. Worker Rank Income (USD/a week)


1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190

- Compute: Mean; Mode; Median with SPSS.


61
MEASURES OF CENTRAL TENDENCY
• Example: Mean; Mode; Median

Statistics
N
Mean Median Mode
Valid
Worker Rank 10 3.4 3.5 5
Income 10 139 137.5 205

62
CENTRAL TENDENCY- VARIABILITY

VARIABILITY
CENTRAL TENDENCY

Mean
Range Mean absolute
deviation
Median
Variance Coefficient of
Mode
Variation

Standard
Deviation
63
RANGE

- Range: the difference between the largest and


smallest data values.

R=xmax-xmin

64
RANGE
• Example: Having the data of 10 Customers according to worker rank
and Income.
N0. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190

- The range of worker rank: R1=6-1=5

- The range of income: R2=205-70=135 (USD)


65
MEAN ABSOLUTE DEVIATION

- Data without frequency:

- Data with frequency


+ xi: The value of the observation i.
+ fi: The frequency of the observation i.
+ n: Number of observations in the sample.

66
MEAN ABSOLUTE DEVIATION
• Example: Having the data of 10 Customers according to worker rank
and Income.
N0. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190

- Compute the Mean absolute deviation for worker rank and Income.

67
VARIANCE
- Variance: the average of the squared differences
between each data value and the mean.

68
VARIANCE – POPULATION
- Data without frequency:

- Data with frequency:

+ xi: The value of the observation i.


+ μ: The mean value of the population.
+ fi: The frequency of the observation i.
+ N: Number of observations in the Population.
69
VARIANCE – SAMPLE
- Data without frequency:

- Data with frequency:


+ xi: The value of the observation i.
+ : The mean value of the population.
+ fi: The frequency of the observation i.
+ n: Number of observations in the sample.
70
STANDARD DEVIATION
- Standard Deviation for a Population:

- Standard Deviation for a Sample:

+ : The Variance value of the Population.


+ : The Variance value of the Sample.
71
VARIANCE - STANDARD DEVIATION
• Example: Having the data of 10 Customers according to worker rank and Income.

No. Worker Rank Income (USD/a week)


1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190
- Compute the Variance and Standard deviation for worker rank and Income.

72
VARIANCE - STANDARD DEVIATION
- Variance and Standard deviation for worker rank:
N0. Worker Rank (xi) (𝒙𝒊 − 𝒙) (𝒙𝒊 − 𝒙)𝟐
1 1 -2.4 5.76
2 5 1.6 2.56
3 4 0.6 0.36
4 5 1.6 2.56
5 3 -0.4 0.16
6 2 -1.4 1.96
7 6 2.6 6.76
8 1 -2.4 5.76
9 2 -1.4 1.96
10 5 1.6 2.56
Mean 3.4 Total 30.4

Variance: => Standard deviation:

=>
73
MEASURES OF VARIABILITY - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…

74
MEASURES OF VARIABILITY
• Example: Having the data of 10 Customers according to worker rank and Income.
No. Worker Rank Income (USD/a week)
1 1 70
2 5 195
3 4 165
4 5 205
5 3 110
6 2 85
7 6 205
8 1 75
9 2 90
10 5 190

- Compute the Range; Variance and Standard deviation for worker rank and
Income with SPSS.
75
MEASURES OF VARIABILITY - SPSS
Statistics
Worker Rank Income
N Valid 10 10
Std. Deviation (Standard deviation) 1.84 57.87
Variance 3.38 3348.89
Range 5.00 135.00
Minimum 1.00 70.00
Maximum 6.00 205.00

76
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation in Stock market:
• Example: Have the data about the Closing price of 3 stock in 10 days
(1000 VND).

Day 1 2 3 4 5 6 7 8 9 10
A 18.7 18.2 18.3 18.5 18.8 18.5 18.6 19.4 19.3 19.5
Stock B 17.8 18.85 19.2 20.4 20.25 20.05 21.2 21.3 21.45 21.7
C 17.1 18.15 18.4 19.6 20.45 19.55 20.8 22.2 21.45 22.7

- Compute the Risk about price market of 3 stocks by variance and


standard deviation.

77
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation:

The result about risk rank of 3 stocks

No. Stock Variance Standard deviation Risk Rank


1 A 0.215 0.464 3
2 B 1.630 1.277 2
3 C 3.313 1.820 1

78
VARIANCE - STANDARD DEVIATION
- Application for variance and standard deviation in services quality:
• Example: The result of quality Rank for Support Services

Services
Standard
No. Code Content Variance quality
deviation
Rank
There are numerous
1 SUP1 2.50 1.581 3
language support services
There are numerous
2 SUP2 0.80 0.894 1
academic support services
There are various social
3 SUP3 1.30 1.140 2
support services

79
COEFFICIENT OF VARIATION

- The coefficient of variation indicates how large the


standard deviation is in relation to the mean
- The coefficient of variation is used in cases of
differences in quantity or scale (unit).

80
COEFFICIENT OF VARIATION
- Coefficient of variation for a Population:

- Coefficient of variation for a Sample:

+ : The Standard deviation of the Population.


+ μ: The mean of the Population.
+ : The mean of the Sample.
+ s: The Standard deviation of the Sample. 81
QUARTILES
- The Quartiles of a data set are values to divide the data items
arranged in ascending order to equal parts.
+ Step 1: Arranged in ascending order
+ Step 2: Quartiles are specific percentiles
• First Quartile = 25th Percentile;
• Second Quartile = 50th Percentile = Median
• Third Quartile = 75th Percentile

82
QUARTILES - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…

83
2.5. MEASURES OF DISTRIBUTION SHAPE

84
SKEWNESS
- Skewness measure of the shape of a data distribution.

- Skewness value:
+ Skewness =0 => Data is Symmetric distribution;
+ Skewness >0 => Data is right skew distribution;
+ Skewness >0 => Data is left skew distribution. 85
KURTOSIS
- Kurtosis measure of the shape of a data distribution.

- Kurtosis value:
+ Kurtosis ≈ 0 => Data is Normal distribution;
+ Kurtosis ≠ 0 => Data is not Normal distribution.

86
SKEWNESS AND KURTOSIS - SPSS
SPSS: Analyze =>
Descriptive statistics =>
Frequencies=>Statistics=>…

87
SUMMARY
Example: Having the data of 20 Customers according to worker rank and Income.

N0. Worker Rank Income (USD/a week) N0. Worker Rank Income (USD/a week)
1 1 70 11 6 180
2 5 195 12 7 220
3 4 165 13 7 215
4 5 205 14 5 200
5 3 110 15 4 173
6 2 85 16 3 120
7 6 205 17 4 175
8 1 75 18 2 75
9 2 90 19 4 120
10 5 190 20 6 160

Compute: mean; mode; median; Variance; Standard Deviation;


Quartiles; Skewness; Kurtosis for worker rank and Income with SPSS

88
SUMMARY
Statistics
Worker Rank Income
N Valid 20 20
Mean 4.1000 151.4000
Median 4.0000 169.0000
Mode 4.00a 75.00
Std. Deviation 1.86096 52.57316
Variance 3.463 2763.937
Skewness -.161 -.360
Kurtosis -.954 -1.475
Range 6.00 150.00
Minimum 1.00 70.00
Maximum 7.00 220.00
25 2.2500 95.0000
Percentiles 50 4.0000 169.0000
75 5.7500 198.7500
89
2.6. Measures of association between two quantitative variables

- Association between two quantitative variables:


XY ≈ Y  X
+ If X changed => Y changed
+ If Y changed => X changed
- Covariance
- Correlation Coefficient

90
COVARIANCE
- The covariance is a measure of the linear association between two
variables.
- The covariance value:
+ Covariance value = 0 => Two variables are not the linear
association:
+ Covariance value > 0 => The linear association between two
variables are Positive:
X (Y) - grow (reduce) => Y(X) - grow (reduce)
+ Covariance value < 0 => The linear association between two
variables are Negative:
X (Y) - reduce (grow) => Y(X) - grow (reduce)
91
COVARIANCE
- The covariance is computed

+ xi; yi : The value of the observation i of the variable x and y.


+ The mean value of the sample of the variable x and y
+ μx; μy: The mean value of the population the variable x and y.
+ n: Number of observations in the sample.
+ N: Number of observations in the Population
92
CORRELATION COEFFICIENT
- The Correlation Coefficient is a measure of the linear association
between two variables.
- The Correlation Coefficient value is between [-1, 1]
+ Correlation Coefficient = 0 => Two variables are not the linear
relationship:
+ Correlation Coefficient near +1 => The linear relationship
between two variables are strongly Positive:
X (Y) - grow (reduce) => Y(X) - grow (reduce)
+ Correlation Coefficient near -1 => The linear relationship
between two variables are strongly Negative:
X (Y) - reduce (grow) => Y(X) - grow (reduce) 93
CORRELATION COEFFICIENT
- The Correlation Coefficient is computed

𝒊 𝒊 𝒊 𝒊
r 𝟐 𝟐
ρ 𝒙
𝟐
𝒚
𝟐
𝒊 𝒊 𝒊 𝒙 𝒊 𝒚
For Samples For Populations

+ xi; yi : The value of the observation i of the variable x and y.


+ The mean value of the sample of the variable x and y
+ μx; μy: The mean value of the population the variable x and y.

94
COVARIANCE - CORRELATION COEFFICIENT

Example: The following data show the daily increase percent or daily
decrease percent in the DJIA and S&P 500 for a sample of nine days
over a three-month period (The Wall Street Journal, January 15 to
March 10, 2006) (9 days).

DJIA 0.20 0.82 -0.99 0.04 -0.24 1.01 0.30 0.55 -0.25
S&P 500 0.24 0.19 -0.91 0.08 -0.33 0.87 0.36 0.83 -0.16

Compute: Covariance and Correlation coefficient for these data.

95
COVARIANCE - CORRELATION COEFFICIENT
Xi Yi
0.2 0.24 0.04 0.11 0.002 0.012 0.004
0.82 0.19 0.66 0.06 0.436 0.004 0.040
-0.99 -0.91 -1.15 -1.04 1.323 1.082 1.196
0.04 0.08 -0.12 -0.05 0.014 0.003 0.006
-0.24 -0.33 -0.4 -0.46 0.160 0.212 0.184
1.01 0.87 0.85 0.74 0.723 0.548 0.629
0.3 0.36 0.14 0.23 0.020 0.053 0.032
0.55 0.83 0.39 0.7 0.152 0.490 0.273
-0.25 -0.16 -0.41 -0.29 0.168 0.084 0.119
0.16 0.13 2.996 2.486 2.483 96
COVARIANCE - CORRELATION COEFFICIENT
- The covariance of sample between DJIA and S&P 500:
𝒊
𝒙𝒚

Sxy=0.31>0 => The linear association between DJIA and S&P are positive
- The correlation coefficient of sample between DJIA and S&P 500:

2.483
r =
2.996∗2.486
=0.91

r=0.91>0: DJIA and S&P are strong positive linear relationship.

97
2.7. MEASURES OF ASSOCIATION BETWEEN TWO
QUALITATIVE VARIABLES
- Cramer's V (V).
- Contingency Coefficient (C)

98
CRAMER'S V
- The Cramer's (V) value:
+ Cramer's (V) = 0 => Two qualitative variables are not
correlated.
+ Cramer's (V) near 1 => Two qualitative variables are strongly
correlated.

99
CRAMER'S V
- Cramer's V (V)
( )
;

n: Number of observations in the sample


fij: Frequency of the row i and the column j
ni: Number of observations in row i.
nj: Number of observations in row j.
k: Number of rows
m: Number of column
h=max(k; m) 100
CONTINGENCY COEFFICIENT
- The Contingency coefficient - C value:
+ Contingency coefficient - C = 0 => Two qualitative variables
are not correlated.
+ Contingency coefficient - C near 1 => Two qualitative
variables are strongly correlated.

101
CONTINGENCY COEFFICIENT
Cramer's V (V)
( ) ∗
;
n: Number of observations in the sample
fij: Frequency of the row i and the column j
ni: Number of observations in row i.
nj: Number of observations in row j.
k: Number of rows
m: Number of columns

102
CRAMER'S - CONTINGENCY COEFFICIENT
SPSS: Analyze => Descriptive statistics => Crosstabs =>Statistics

103
THE END

104

You might also like