0% found this document useful (0 votes)
131 views2 pages

Correlation and Regression Examples

1. This document provides examples of calculating correlation coefficients, regression coefficients, and probable errors from sample data. It also includes examples of interpreting regression equations and determining most likely values from bivariate data using correlation analysis. 2. Examples include calculating correlation coefficients from price and demand data for a commodity, income and expenditure data across years, and correlation between age and playing habits from sample data. 3. Additional examples demonstrate calculating regression equations and determining most likely values from bivariate data when given statistics like the means, variances, and correlation coefficient.

Uploaded by

Dheephiha m.s
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
131 views2 pages

Correlation and Regression Examples

1. This document provides examples of calculating correlation coefficients, regression coefficients, and probable errors from sample data. It also includes examples of interpreting regression equations and determining most likely values from bivariate data using correlation analysis. 2. Examples include calculating correlation coefficients from price and demand data for a commodity, income and expenditure data across years, and correlation between age and playing habits from sample data. 3. Additional examples demonstrate calculating regression equations and determining most likely values from bivariate data when given statistics like the means, variances, and correlation coefficient.

Uploaded by

Dheephiha m.s
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Correlation and Regression

1.
a) Calculate two regression coefficients when r=0.8, σX = 5and σy = 7.
b) Find the probable error when number of items are 5 and correlation is 0.7 for a given distribution
c) A student calculated the regression coefficient of x on y as -1.25 and regression coefficient of y on x
as -2.40. Comment on his calculation.
d) If n=8 and ∑d2 =4 find rs ?
e) IF bxy= 0.8 and byx = 0.6 find ‘r’ ?
f) Calculate two regression coefficients when r=0.8, σX = 5and σy = 7.
2. Calculate karl pearsons coefficient of correlation for the following data regarding price and demand of a
certain commodity.
Price: 21 22 23 24 25 26 27 28 29
Demand: 20 19 19 17 17 16 16 15 14
3. Compute karl pearson’s coefficient of correlation between per capita income and per capita consumer
expenditure from the following data given below:
Year: 1990 91 92 93 94 95 96 97 98 99
Per capita income: 249 251 248 252 258 269 271 272 280 275
Per capita Expenditure: 237 238 236 240 245 255 254 252 258 251
4. From the following table find correlation coefficient between age and playing habits of students:
Age 15 16 17 18 19 20
No. of students 250 200 150 120 100 80
Regular players 200 150 90 48 30 12
5. The corresponding values of 2 series are given below in the table:
ii) X: 42 44 58 55 89 98 66
iii) Y: 56 49 53 58 65 79 58
iv) Find the coefficient of correlation of the above series. How will you test the significance of ‘r’?
6. Calculate the karl pearson’s coefficient fro the following series of marks scored by ten students in a class
test in Mathematics and Statistics.

Marks in Mathematics 45 70 65 30 90 40 50 75 85 60
Marks in statistics 35 90 70 40 95 40 60 80 80 50
7. From the following calculate coefficient of correlation and probable error between advertisement
expenses and sales of a company during the year 2008:
i. Month jan feb mar apr may jun july aug sept oct nov dec
[Link]. 50 60 70 90 120 150 140 160 170 190 200 150
iii. Sales 1200 1500 1600 2000 2200 2500 2400 2600 2800 2900 3100 3900
8. Calculate rank correlation coefficient between marks in statistics and accountancy.
i. Marks in statistics 48 60 72 62 56 40 39 52 30
[Link] in accountancy 62 78 65 70 38 54 60 32 31
9. From the following data calculate coefficient of rank correlation between X and Y.
X: 36 56 20 65 42 33 44 50 15 60
Y: 50 35 70 25 58 75 60 45 80 38
10. From the data given below compute Karl Pearsons coefficient of correlation:
X series Y series
Number of items 15 15
Mean 25 18
Square of deviation from mean 136 138
Summation of product of deviations from mean of x and y -122.
11. Calculate karl pearsons coefficient of correlation between percentage of pass and failures from the
following data. Also obtain probable error:
No. of students: 80 60 90 70 50 40
No. of students passed 48 30 45 56 45 30
12. For a given distribution the value of correlation is 0.64 and its probable error is 0.13274. Find the number
of items in the series
13. From the data given below find:
I. The two regression equations.
II. The most likely marks in statistics when marks in income tax are 50
III. The coefficient of correlation with bxy and byx
Marks in Income tax: 35 37 28 42 48 54 65 40 35 50
Marks in Statistics: 60 55 75 70 80 35 65 35 50 60
14. Obtain the two regression equation from the following:
X series Y series
Mean 36 85
Variance 121 64
Coefficient of correlation between X and Y is 0.66
15. Obtain the two regression equation from the following:
X series Y series
Mean 20 25
Variance 4 9
Coefficient of correlation between X and Y is 0.75
16. Calculate the two regression coefficients when r= 0.9,σx = 10 and σy = 1.5
17. Find ‘r’ when the two regression coefficients are -0.6 and -1.4
18. Find the most probable value of Y if X is 70 and the most probable value of X when Y is 90 given that
X series Y series
Mean 18 100
SD 14 20
Coefficient of correlation between X and Y is 0.8
19. If covariance between X and Y variables is 12.5 and the variance of X and Y are respectively 16.4 and 13.8
find the coefficient of correlation between them
20. The following data related to the age of husbands and wives. Obtain the two regression equations and
determine the most likely age of husband when age of wife is 25 years.

Age of husbands 25 28 30 32 35 36 38 39 42 55
Age of wives 20 26 29 30 25 18 26 35 35 46
21. Given below is the information about advertisement expenditure and sales:

X Y
Advertisement expenses sales
Mean 20 120
SD 5 25
Given r = 0.6 calculate the two regression equation.

Common questions

Powered by AI

In determining the correlation coefficient, use the relation r = Cov(X,Y) / (SD_X * SD_Y). The standard deviations are the square roots of the variances: SD_X = √16.4 ≈ 4.05 and SD_Y = √13.8 ≈ 3.71. Thus, substitute these into the equation: r = 12.5 / (4.05 * 3.71), which calculates to approximately 0.84. This reflects a strong positive association between the variables X and Y .

Calculating both regression coefficients and correlation provides comprehensive insights into the relationship between variables. Regression coefficients quantify the change in one variable relative to another, while correlation assesses strength and direction. Together, they depict not just linkage strength but predictive capability, making them crucial for understanding dependencies and directional inferences. Their joint application offers a nuanced picture that mere reliance on one statistic might obscure, enriching analysis .

To compute Karl Pearson’s coefficient of correlation, use the formula r = Σ (X - X̄)(Y - Ȳ) / √(Σ(X - X̄)² * Σ(Y - Ȳ)²). Given the number of items (15), means (X̄ = 25, Ȳ = 18), square of deviations (136 for X, 138 for Y), and summed deviations (-122), substitute into the formula: r = -122 / √(136*138). This calculation gives r ≈ -0.297, indicating a weak negative correlation between the variables .

To find the number of observations given a correlation coefficient (r = 0.64) and its probable error (PE = 0.13274), use the formula PE = 0.6745 * (1-r^2)/√n, and solve for n. Rearrange the formula as √n = 0.6745 * (1-0.64^2)/0.13274. Calculating this gives √n ≈ 2.125, therefore n ≈ 4.52. Since the sample size must be a whole number, rounding up suggests there are approximately 5 observations .

Given regression equations of husbands' ages (Y) on wives' ages (X), calculate Y using the derived equation of Y = b_xyX + a, where 'a' is the intercept calculated based on mean adjustments. If wife's age is 25 and assuming a regression equation similar to Y = 0.9X + k (where k is derived from means), substitute wife's age into the equation to find the expected husband age. Assuming for instance k ≈ 5, Y = 0.9(25) + 5 results in a Y value representing the husband's likely age, illustrating how line fitting aids in predictive analysis .

In cases where the price of a commodity increases consistently but the demand decreases, this indicates a negative correlation, meaning as one variable increases, the other decreases. Karl Pearson's coefficient is useful for identifying and quantifying this negative correlation. A strong negative correlation close to -1 would mean that price and demand have an inverse relationship, suggesting that as prices rise, the demand falls significantly following the law of demand .

Probable error (PE) provides a measure of the variability of the correlation coefficient and is used to test its significance. It is calculated using the formula PE = 0.6745 * (1-r^2)/√n. Substituting the given values, r = 0.7, and n = 5, the calculation would be PE = 0.6745 * (1-0.7^2)/√5, which simplifies to approximately 0.1826. This helps understand the reliability of the correlation; if the value of correlation r lies below this range, it implies insignificance .

The regression equation Y on X is given by: Y - Ȳ = b_xy(X - X̄), where b_xy = r * (SD_Y/SD_X). Substitute the given values: b_xy = 0.6 * (25/5) = 3. The regression equation thus is Y - 25 = 3(X - 20), resulting in Y = 3X - 35. For the regression equation X on Y, use the reverse ratio where b_yx = r * (SD_X/SD_Y); thus X = (r * (5/25))(Y - Ȳ) + X̄. Here, X = 0.12Y - 1. Correctly forming these equations helps predict Y based on X and vice-versa .

This situation involves examining the validity of the regression coefficients calculated by the student. The product of the regression coefficients (b_xy and b_yx) should equal the square of the correlation coefficient (r^2). Mathematically, b_xy * b_yx = r^2. In this case, the student's calculated values are b_xy = -1.25 and b_yx = -2.40. Multiplying these gives 3.0, which is not equal to any possible value of r^2 since correlation values range from -1 to 1, implying r^2 should be between 0 and 1. Therefore, the student's calculations seem incorrect, suggesting a miscalculation or a misunderstanding of the regression-coefficient relationship .

Ignoring rank in correlation can lead to misleading results, especially in small samples, where outliers can disproportionately impact analysis. Rank correlation methods, like Spearman's, mitigate outlier effects and better capture ordinal relationships. Without ranking positions, true correlations may be skewed or unnoticed, affecting decision-making or model accuracy. This involves recognizing rank-based methods’ robustness in diverse data conditions, promoting reliable outcomes .

You might also like