0% found this document useful (0 votes)
6 views3 pages

Understanding Correlation in Statistics

The document discusses quantitative methods and examples of calculating and interpreting correlation coefficients (r). It provides examples of calculating r for different datasets and examining the impact of outliers. It demonstrates that r only measures linear relationships and may be low despite clear relationships in scatterplots. It also identifies common mistakes in interpreting r such as applying it to categorical variables or attaching units to r.
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views3 pages

Understanding Correlation in Statistics

The document discusses quantitative methods and examples of calculating and interpreting correlation coefficients (r). It provides examples of calculating r for different datasets and examining the impact of outliers. It demonstrates that r only measures linear relationships and may be low despite clear relationships in scatterplots. It also identifies common mistakes in interpreting r such as applying it to categorical variables or attaching units to r.
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Quantitative Methods

Example 1: (a) Find the correlation between these two variables. r = -0.2013
Relationship Between First Exam and Final Exam
180 170 160 150 140 130 120 110 110.00

Final Exam Score

120.00

130.00

140.00 First Exam Score

150.00

160.00

170.00

(b) Relationship between these two variables is weak. Does your calculation of the correlation support this statement? Explain your answer. The interesting part is that the relationship seems to be negative, most likely caused by the extra ordinary student. Lets remove that point and recalculate r. Without the outlier the r value is now 0.3222, which says that the association while very weak it is positive. Should we eliminate that student? How likely is it that a person can get the worst score on the first and then score the highest on the final? What are we interested in displaying the possible outcomes on the majority of the group or all of the group. You can again see that r is not resistant to outliers thus, we should be careful when outliers appear, and decide whether we need to eliminate them to get a better understanding of the entire group and not just the unusual point. First Final 153 144 162 149 127 118 158 153 145 140 145 170 145 175 170 160
175

Final exam Score

165 155 145 135 125 125

135

145 First Exam Score

155

165

(a) Find the correlation between these two variables. r = 0.5194 (b) The relationship between these two variables is stronger than the relationship between the two variables in the previous exercise. How do the values of the correlations that you calculated support this statement? Explain your answer. The variation about a straight line seems to be less and as the second exam score increases there is a larger increase in the average value of the final exam score. Yet one can still see a quite a lot of variation about the average trend line, thus the reason for the correlation value of 0.5194.

200 190

Score on Final

180 170 160 150 140 130 130

140

150

160

170

180

Score for Second Exam

Example 2: Refer to the previous exercise. Add a ninth student whose scores on the second test and final exam would lead you to classify the additional data point as an outlier. Recalculate the correlation with this additional case and summarize the effect it has on the value of the correlation. Second Final 1 158 145 2 162 140 3 144 145 4 162 170 5 136 145 6 158 175 7 175 170 8 153 160 9 200 200 The ninth student gets two perfect scores on the second and the final (assuming that 200 is the most one can score). The corresponding r value is 0.8015. This is quite a large change from r = 0.5194.

200 190

Score on Final

180 170 160 150 140 130 130

140

150

160

170

180

190

200

210

Score for Second Exam

Example 3: Make a scatterplot find the correlation r. Explain why r is close to zero despite a strong relationship between speed and gas used.

Chart Title 25
Fuel Spent (l/100km)

The value of r = -0.1716

20 15 10 5 0 0 50 100 Speed (km/h) 150 200

The value of r only measures how close our data follows a linear relationship, which this situation does not.
Fuel Linear (Fuel)

Notice r = 0 despite the fact that we do not have a straight line. This shows the importance of looking at the scatterplot.

Example 4:What's wrong? Each of the following statements contains a blunder. Explain in each case what is wrong. (a) "There is a high correlation between the gender of American workers and their income." Gender is a categorical variable. The correlation value r is only to be used to indicate linear association between two quantitative variables. (b) "We found a high correlation (r = 1.09) between students' ratings of faculty teaching and ratings made by other faculty members." The r value can only be a number in the interval [-1, 1]. (c) "The correlation between planting rate and yield of corn was found to be r = 0.23 bushel." The statistic r is a unitless number, thus the statement above, was found to be r = 0.23 bushel." attaches a unit to r which is not correct. The correlation r does not depend on the unit of the measurement.

Common questions

Powered by AI

A scatterplot might suggest a strong relationship due to non-linear patterns, even when the correlation coefficient is close to zero, because correlation measures only linear relationships. For example, the data for speed and fuel usage had an r value of -0.1716, indicating no linear correlation, yet a scatterplot likely showed a non-linear trend, demonstrating why visual inspection is crucial .

Correlation is a unitless measure of linear association between two variables. Misunderstanding arises when units are mistakenly attached to r, such as saying r = 0.23 bushel, as this falsely implies that correlation reflects physical quantities. The nature of correlation is purely statistical, indicating the strength and direction of a relationship, not dependent on specific units .

Considering correlation's non-resistance to outliers is critical, as outliers can heavily skew its value and thus distort the interpreted strength and direction of relationships. For example, an outlier shifted r from -0.2013 to 0.3222, changing interpretation radically. Accurate interpretations require understanding that correlation might not reflect most data if few points are anomalous .

The mathematical domain of the correlation coefficient r is between -1 and 1. This range means values outside of it, such as r = 1.09, are incorrect because they imply impossible relationships. A coefficient greater than 1 suggests a misunderstanding or error in analysis, perhaps in data calculation or misinterpretation of results .

It is essential to distinguish between categorical and quantitative variables because correlation, represented by the coefficient r, only applies to associations between two quantitative variables. Using correlation with categorical variables, like gender, is erroneous as it does not fit its definition, which is to measure linear relationships between numerical data .

Recalculating correlation after adding a new data point is crucial because it can significantly change the value of r, indicating altered relationships. For instance, adding a ninth student scoring perfectly increased r from 0.5194 to 0.8015, suggesting a much stronger linear association. Such changes highlight the sensitivity of correlation to individual data points, particularly those considered outliers .

Recalculating correlation after modifying a data set provides better insights when original data includes errors, misrepresentative outliers, or when exploring hypotheses about data structure changes. For example, removing an extraordinary data point shifted the correlation from -0.2013 to a more meaningful 0.3222. Adding a high-scoring outlier elevated correlation to 0.8015, reflecting enhanced linearity. These recalibrations reveal more accurate inter-variable relationships .

To interpret a correlation coefficient, analyze both its magnitude and sign. A coefficient close to 1 or -1 indicates a strong linear relationship, while near zero suggests a weak one. The sign indicates direction; positive means both variables increase together, while negative means one decreases as the other increases. For example, a coefficient of 0.5194 indicates a moderate positive linear relationship .

The presence of an outlier can significantly affect the correlation coefficient, as seen in the case where an extraordinary student's score shifted the correlation from -0.2013 to 0.3222 when removed. This indicates that the initial negative correlation was influenced by the outlier, and without it, the correlation turns positive. This is because correlation is not resistant to outliers. If the analysis aims to understand the general trend for the majority, removing an outlier may provide a more accurate representation .

Using correlation coefficients in decision-making requires caution, especially when outliers are present because they can distort the true relationship. For instance, an outlier affected correlation from -0.2013 to 0.3222, misleading the analysis. If outliers represent errors or rare cases, one might exclude them for clarity. Decision-makers must consider if they wish to display trends for most or allow outliers to skew interpretation .

You might also like