0% found this document useful (0 votes)
21 views9 pages

Correlation Coefficient Values Explained

Uploaded by

rehaan986750
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views9 pages

Correlation Coefficient Values Explained

Uploaded by

rehaan986750
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SE-CO- EM 3 Prof.

Rajee John
Module 5-Stastistical Techniques
Karl Pearson’s coefficient of correlation (r)
A correlation coefficient is a numerical measure of some type of correlation, meaning a statistical
relationship between two variables. OR
The correlation coefficient is a statistical measure of the strength of the relationship between two
variables.
The mathematical expression for correlation, suggested by Karl Pearson is

∑(𝑥𝑖 −𝑥̅ )(𝑦𝑖 −𝑦̅) ∑(𝑥𝑖 −𝑥̅ )2


𝑟= ……………. (1), where 𝑥 = √ &
𝑛 𝑥 𝑦 𝑛

∑(𝑦𝑖 −𝑦̅ )2
𝑦 = √ 𝑛
∑(𝑥𝑖 −𝑥̅ )(𝑦𝑖 −𝑦̅)
Then, 𝑟 = …………(2)
√∑(𝑥𝑖 −𝑥̅ )2 ∑(𝑦𝑖 −𝑦̅ )2

If 𝑥𝑖 − 𝑥̅ = 𝑑𝑥 & (𝑦𝑖 − 𝑦̅) = 𝑑𝑦 , then


∑ 𝑑𝑥 𝑑𝑦
𝑟= ……..(3)
2 2
√∑ 𝑑𝑥 ∑ 𝑑𝑦

∑ 𝑥𝑖 𝑦𝑖 −𝑛𝑥̅ 𝑦̅
𝑟= ……… (4)
2 2 )(∑ 𝑦 2 −𝑛𝑦
√(∑ 𝑥𝑖 −𝑛𝑥̅ 𝑖 ̅ 2)

∑ 𝑥𝑖 𝑦𝑖 − 𝑛𝑥̅ 𝑦̅
𝑟=
𝑛 𝑥 𝑦
Note:
1. If x and y are independent variables, they are not correlated (r=0)
2. Correlation coefficient r lies between -1 and +1
3. r = 1 means a perfect positive correlation and the value r = -1 means a perfect
negative correlation.
Questions

1. Find from the following values of the demand and the corresponding price of a commodity, the
degree of correlation between the demand and price by computing Karl Pearson’s coefficient of
correlation.

Demand in quintals : 65 66 67 67 68 69 70 72
Price in paise/kg. : 67 68 65 68 72 72 69 71

2. Calculate the coefficient of correlation between price and demand

Price 2 3 4 7 6

Demand 10 7 3 1 2

Spearman's Rank correlation coefficient


Case(i) When the ranks are non-repeated
Spearman's Rank correlation coefficient is a technique which can be used to summarize the
strength and direction (negative or positive) of a relationship between two variables. The result will
always be between 1 and minus 1.
It depends upon the ranks of the items and actual values of items are not required. Hence this can
be used even when actual values are not known.

X 30 33 25 10 33 75 40 85 90 95 65 55
Y 68 65 80 85 70 30 55 18 15 10 35 45
6 ∑ 𝑑𝑖2
It is denoted by R and is given by the formula, 𝑅 =1−
𝑛3 −𝑛
eg: Q1) Calculate Spearman’s rank correlation between x and y
x: 36 56 20 42 33 44 50 15 60
y: 50 35 70 58 75 60 45 80 38
Q2) Calculate the rank correlation coefficient from the following data, relating to the ranks of 10
students in English and Mathematics
Student No: 1 2 3 4 5 6 7 8 9 10
Rank in Eng: 1 3 7 5 4 6 2 10 9 8
Rank in Maths: 3 1 4 5 6 9 7 8 10 2
Q3) Find the coefficient of correlation between x and y for the following data
x: 62 64 65 69 70 71 72 74
y: 126 125 139 145 165 152 180 208
Q4) Calculate the correlation coefficient ‘r’

Marks
Eng 56 75 45 71 62 64 58 80 76 61

Mat
66 70 40 60 65 56 59 77 67 63
h

English Maths Rank Rank


d d2
(mark) (mark) (English) (maths)

56 66 9 4 5 25

75 70 3 2 1 1

45 40 10 10 0 0

71 60 4 7 3 9

62 65 6 5 1 1

64 56 5 9 4 16

58 59 8 8 0 0

80 77 1 1 0 0

76 67 2 3 1 1

61 63 7 6 1 1

Where d = difference between ranks


Case(i) When the ranks are repeated
1 1
6 { ∑ 𝑑𝑖2 + (𝑚13 −𝑚1 )+ (𝑚23 −𝑚2 )+⋯}
12 12
When the ranks are repeated, 𝑅 =1− 3
𝑛 −𝑛
Q 1) Obtain the rank correlation coefficient from the following data
X: 10 12 18 18 15 40
Y: 12 18 25 25 50 25
Q2) Find the rank correlation from the following data
X: 32 55 49 60 43 37 43 49 10 20
Y: 40 30 70 20 30 50 72 60 45 25
Regression
Regression can be defined as ‘a method of estimating the value of one variable when that of the
other is known and when the variables are correlated’
There are two lines of regression:
(i) Regression line of y on x, is given by 𝑦 = 𝑎 + 𝑏𝑥 (This is used to find the value of y for a
given value of x), where b is the coefficient of regression of y on x , denoted by 𝑏𝑦𝑥
(which is the slope of the line y on x)
𝜎𝑦
(ii) Using method of least squares, it can be written as 𝑦 − 𝑦̅ = 𝑟 (𝑥 − 𝑥̅ )
𝜎𝑥
OR 𝑦 − 𝑦̅ = 𝑏𝑦𝑥 (𝑥 − 𝑥̅ )

(iii) Regression line of y on x, is given by 𝒚 = 𝒂 + 𝒃𝒙 , can be found by the following

method also.

𝑦 = 𝑎 + 𝑏𝑥 ∴ ∑𝑦 = 𝑛𝑎 + 𝑏 ∑ 𝑥

& ∑ 𝑥𝑦 = 𝑎 ∑ 𝑥 + 𝑏 ∑ 𝑥 2 , from which a and b can be calculated.

(iv) Regression line of x on y, is given by 𝒙 = 𝒂 + 𝒃𝒚 , where b is the coefficient of

regression of x on y, denoted by 𝑏𝑥𝑦

𝜎𝑥
(v) It can be written as 𝑥 − 𝑥̅ = 𝑟 (𝑦 − 𝑦̅)
𝜎𝑦

𝑥 − 𝑥̅ = 𝑏𝑥𝑦 (𝑦 − 𝑦̅)
(vi) Regression line of x on y, is given by 𝑥 = 𝑎 + 𝑏𝑦 , can be found by the following method
also. 𝑥 = 𝑎 + 𝑏𝑦 ∴ ∑ 𝑥 = 𝑛𝑎 + 𝑏 ∑ 𝑦 &
2
∑ 𝑥𝑦 = 𝑎 ∑ 𝑦 + 𝑏 ∑ 𝑦 , from which a and b can be calculated.
To calculate 𝒃𝒚𝒙 and 𝒃𝒙𝒚
∑ 𝑥𝑦 ∑ 𝑥𝑦
(i) 𝑏𝑦𝑥 = ∑ 𝑥2
& 𝑏𝑥𝑦 = ∑ where 𝑥 = 𝑥 − 𝑥̅ and 𝑦 = 𝑦 − 𝑦̅
𝑦2

(ii) By using actual values of the two variates,


∑ 𝑋. ∑ 𝑌 ∑ 𝑋. ∑ 𝑌
∑ 𝑋𝑌 − ∑ 𝑋𝑌−
𝑁 𝑁
𝑏𝑦𝑥 = (∑ 𝑋)
2 & 𝑏𝑥𝑦 = (∑ 𝑌)
2
∑ 𝑋2− ∑ 𝑌2−
𝑁 𝑁

Result:
𝜎 𝜎
𝑏𝑦𝑥 𝑏𝑥𝑦 = 𝑟 𝜎𝑦 𝑟 𝜎𝑥 = 𝑟 2
𝑥 𝑦

Questions
1. Find the equations of lines of regression for the following data

x : 5 6 7 8 9 10 11
y : 11 14 14 15 12 17 16 Also find r
2. Given 6Y=5X+90, 15X=8Y+130,𝜎𝑥2 = 16. Find (i) 𝑥̅ 𝑎𝑛𝑑 𝑦̅ (ii) r and (iii) 𝜎𝑦2

3. Obtain the rank correlation coefficient from the following data


X: 10 12 18 18 15 40
Y: 12 18 25 25 50 25
4. Find the rank correlation from the following data
X: 32 55 49 60 43 37 43 49 10 20
Y: 40 30 70 20 30 50 72 60 45 25

Lines of regression

Regression can be defined as ‘a method of estimating the value of one variable when that of
the other is known and when the variables are correlated’
There are two lines of regression:
i. Regression line of y on x, is given by 𝑦 = 𝑎 + 𝑏𝑥 (This is used to find the
value of y for a given value of x), where b is the coefficient of regression of y on x , denoted by
𝑏𝑦𝑥 (which is the slope of the line y on x)
𝜎𝑦
ii. Using method of least squares, it can be written as 𝑦 − 𝑦̅ = 𝑟 𝜎 (𝑥 − 𝑥̅ )
𝑥
OR 𝑦 − 𝑦̅ = 𝑏𝑦𝑥 (𝑥 − 𝑥̅ )
iii. Regression line of y on x, is given by 𝒚 = 𝒂 + 𝒃𝒙 , can be found by the following method
also.
𝑦 = 𝑎 + 𝑏𝑥 ∴ Normal equations are ∑ 𝑦 = 𝑛𝑎 + 𝑏 ∑ 𝑥

& ∑ 𝑥𝑦 = 𝑎 ∑ 𝑥 + 𝑏 ∑ 𝑥 2 , from which a and b can be calculated.

iv. Regression line of x on y, is given by 𝒙 = 𝒂 + 𝒃𝒚 , where b is the coefficient of regression


of x on y, denoted by 𝑏𝑥𝑦
𝜎
It can be written as 𝑥 − 𝑥̅ = 𝑟 𝜎𝑦 (𝑦 − 𝑦̅)
𝑥
v.

𝑥 − 𝑥̅ = 𝑏𝑥𝑦 (𝑦 − 𝑦̅)

vi. Regression line of x on y, is given by 𝑥 = 𝑎 + 𝑏𝑦 , can be found by the following method also.

𝑥 = 𝑎 + 𝑏𝑦
∴ Normal equations are ∑ 𝑥 = 𝑛𝑎 + 𝑏 ∑ 𝑦 &

∑ 𝑥𝑦 = 𝑎 ∑ 𝑦 + 𝑏 ∑ 𝑦 2 , from which a and b can be calculated.

To calculate 𝒃𝒚𝒙 and 𝒃𝒙𝒚 :


∑ 𝑥𝑦 ∑ 𝑥𝑦
(i) 𝑏𝑦𝑥 = ∑ 𝑥2
& 𝑏𝑥𝑦 = ∑ where 𝑥 = 𝑥 − 𝑥̅ and 𝑦 = 𝑦 − 𝑦̅
𝑦2

(ii) By using actual values of the two variates,


∑ 𝑋. ∑ 𝑌 ∑ 𝑋. ∑ 𝑌
∑ 𝑋𝑌 − ∑ 𝑋𝑌−
𝑁 𝑁
𝑏𝑦𝑥 = (∑ 𝑋)
2 & 𝑏𝑥𝑦 = (∑ 𝑌)
2
∑ 𝑋2 − ∑ 𝑌2−
𝑁 𝑁

Result:
𝜎𝑦 𝜎
𝑏𝑦𝑥 𝑏𝑥𝑦 = 𝑟 𝜎 𝑟 𝜎𝑥 = 𝑟 2
𝑥 𝑦

1. Find the equations of lines of regression for the following data

X: 5 6 7 8 9 10 11
Y: 11 14 14 15 12 17 16 Also find r
Soln:
Let the line of regression of y on x be 𝑦 − 𝑦̅ = 𝑏𝑦𝑥 (𝑥 − 𝑥̅ ) &

x on y be x − 𝑥̅ = 𝑏𝑥𝑦 (𝑦 − 𝑦̅)

By using actual values of the two variates,


∑ 𝑋. ∑ 𝑌 ∑ 𝑋. ∑ 𝑌
∑ 𝑋𝑌 − ∑ 𝑋𝑌−
𝑁 𝑁
𝑏𝑦𝑥 = (∑ 𝑋)
2 & 𝑏𝑥𝑦 = (∑ 𝑌)
2
∑ 𝑋 2− ∑ 𝑌2−
𝑁 𝑁
X 𝑋2 Y 𝑌2 XY

5 25 11 121 55

6 36 14 196 84

7 49 14 196 98

8 64 15 225 120

9 81 12 144 108

10 100 17 289 170

11 121 16 256 176

∑ 𝑋 = 56 ∑ 𝑋 2 = 476 ∑ 𝑌 = 99 ∑ Y 2 = 1427 ∑ 𝑋𝑌 = 811

Substituting and simplifying the equations, we get the answer

Method (ii) (Otherwise called as method of fitting of curve)

Let Regression line of y on x be 𝑦 = 𝑎 + 𝑏𝑥

∑ 𝑦 = 𝑛𝑎 + 𝑏 ∑ 𝑥 & ∑ 𝑥𝑦 = 𝑎 ∑ 𝑥 + 𝑏 ∑ 𝑥 2

& Regression line of x on y be 𝑥 = 𝑎 + 𝑏𝑦

∑ 𝑥 = 𝑛𝑎 + 𝑏 ∑ 𝑦 & ∑ 𝑥𝑦 = 𝑎 ∑ 𝑦 + 𝑏 ∑ 𝑦 2

Q3) If the tangent of the angle made by the line of regression of y on x is 0.6 and σy = 2 𝜎𝑥

Find the correlation coefficient between x and y

Soln: If the equation of the line of regression of y on x is

𝑦 − 𝑦̅ = 𝑏𝑦𝑥 (𝑥 − 𝑥̅ ), then we know that 𝑏𝑦𝑥 is the slope of the line of regression.
𝜎 𝜎
Hence 𝑏𝑦𝑥 = 0.6 But, 𝑏𝑥𝑦 = 𝑟 𝜎𝑥 ∴ 𝑏𝑦𝑥 = 𝑟 𝜎𝑦 Hence r =0.3
𝑦 𝑥

Q4) Find the (i) lines of regression, (ii) Coefficient of correlation for the following data.
Estimate the value of Y when X=50
X: 78 36 98 25 75 82 90 62 65 39

Y: 84 51 91 60 68 62 86 58 53 47
Equation of line of regression of y on x is 𝑦 − 𝑦̅ = 𝑏𝑦𝑥 (𝑥 − ̅̅̅
𝑥))
∑ 𝑥𝑦 ∑ 𝑥𝑦
and x on y is 𝑥 − 𝑥̅ = 𝑏𝑥𝑦 (𝑦 − 𝑦̅) where 𝑏𝑦𝑥 = ∑ 𝑥 2 & 𝑏𝑥𝑦 = ∑ 𝑦2

& 𝑥 = 𝑥𝑖 − 𝑥̅ and 𝑦 = 𝑦𝑖 − 𝑦̅

Fitting of first- and second-degree curves

Consider the equation y= a+bx, of first order.


Then the normal equations are ∑ 𝑦 = 𝑛𝑎 + 𝑏 ∑ 𝑥 & ∑ 𝑥𝑦 = 𝑎 ∑ 𝑥 + 𝑏 ∑ 𝑥 2

Q 1) Fit a straight line to the given data

x 3 6 5 4 4 6 7 5
y 3 2 3 5 3 6 6 4

Q2) Fit second degree polynomial (parabolic curve) to the following data.

x 1 2 3 4 5 6 7 8 9
y 2 6 7 8 10 11 11 10 9

Consider the equation 𝑦 = 𝑎𝑥 2 + 𝑏𝑥 + 𝑐 , of order 2

Then the normal equations are ∑ 𝑦 = 𝑎 ∑ 𝑥 2 + 𝑏 ∑ 𝑥 + 𝑛𝑐


∑ 𝑥𝑦 = 𝑎 ∑ 𝑥 3 + 𝑏 ∑ 𝑥 2 + 𝑐 ∑ 𝑥

∑ 𝑥2𝑦 = 𝑎 ∑ 𝑥4 + 𝑏 ∑ 𝑥3 + 𝑐 ∑ 𝑥2

Common questions

Powered by AI

The method of least squares is fundamental to regression analysis as it minimizes the sum of the squares of the residuals (the differences between observed and predicted values), thus producing the best-fitting line through the data. This method ensures that the total error between the predicted and actual observations is minimized, making it the most reliable technique for estimating the parameters (slope and intercept) of regression lines .

Lines of regression are used to predict the value of a dependent variable based on the value of an independent variable. The line 'y on x' is used to predict 'y' using 'x', and is formulated as y = a + bx, where b is the slope indicating how much y changes for a unit change in x. The line 'x on y' predicts 'x' using 'y', and is similarly formulated as x = a + by. The differences arise in the orientation and interpretation because 'y on x' focuses on changes in 'y' concerning 'x', whereas 'x on y' does the opposite .

Karl Pearson's coefficient of correlation, denoted as 'r', measures the strength and direction of the linear relationship between two variables. The formula is r = ∑(xi−x̄)(yi−ȳ) / sqrt[∑(xi−x̄)² * ∑(yi−ȳ)²], where xi and yi are the data points, and x̄ and ȳ are the means of x and y respectively. The correlation coefficient 'r' ranges between -1 and 1, indicating perfect negative correlation and perfect positive correlation, respectively, while 0 indicates no correlation .

The normal equations for the regression line are derived from minimizing the sum of the squares of the vertical distances between the observed and estimated points. They are: ∑y = na + b∑x for y = a + bx, and ∑xy = a∑x + b∑x². Using these, parameters 'a' and 'b', which are the intercept and regression slope, respectively, can be calculated for a dataset to establish the best fitting line .

The coefficient of regression is calculated using the formulas b_yx = Σxy / Σx² for 'y on x' and b_xy = Σxy / Σy² for 'x on y', where Σxy represents the sum of the products of 'x' and 'y' deviations. This coefficient measures how much change in the dependent variable is expected per unit change in the independent variable, indicating the strength and direction of the relationship between the variables. It is critical for creating predictions and understanding statistical associations .

Spearman's Rank correlation coefficient is useful when the actual values of the items are unknown because it depends on the ranks of the data rather than the raw data values. It is calculated using rank positions, allowing for the summary of the strength and direction of a relationship between two variables even in the absence of absolute values. This flexibility makes it applicable in situations where only ordinal data is available .

Regression analysis is a statistical method for estimating the relationships among variables, specifically how a dependent variable changes when any one of the independent variables is varied. It enables prediction by creating a model, typically a linear equation, that describes the dependent relationship. When variables are correlated, regression analysis allows for forecasting the dependent variable based on known values of the independent variables, thus providing insight into data trends and making future predictions possible .

The tangent of the angle made by the regression line, referred to as the slope or gradient, is directly related to the coefficient of correlation by the ratio σy/σx times the correlation coefficient. It indicates the rate of change of the dependent variable with respect to the independent variable. Understanding this connection helps quantify how much one variable changes in response to another, further linking geometric representation of data to statistical correlation measures .

Rank correlation is particularly useful over Pearson's correlation in situations where data does not meet the assumptions necessary for Pearson's correlation, such as heteroscedasticity, non-linearity, or when dealing with ordinal data. Since rank correlation (e.g., Spearman's) uses ranks rather than raw data, it is less sensitive to outliers and non-normal data distributions .

The coefficient of determination, denoted as R², is the square of Pearson's correlation coefficient, r. It indicates the proportion of the variance in the dependent variable that is predictable from the independent variable. Essentially, if r is the measure of the strength of a linear relationship, R² quantifies how well the line fits the data, providing insights into the goodness-of-fit for the model .

You might also like