MODULE IV
Correlation and Regression
Detailed Study Notes for Psychology Statistics
Topics Covered in This Module
1. Correlation – Meaning & Significance in Psychological Studies 2. Types of Correlation – Pearson's
Product Moment Method 3. Types of Correlation – Spearman's Rank Correlation Method 4.
Significance of Scatter Plot 5. Regression – Simple Linear Regression 6. Assumptions and
Limitations of Regression Analysis
Source: Howell, D.C. – Statistical Methods for Psychology (Chapters 9 & 10)
Section 1: Correlation – Meaning and Significance
1.1 What is Correlation?
Correlation is a statistical technique that measures the degree and direction of the linear relationship
between two variables. In psychological research, we are frequently interested not merely in whether
groups differ, but in how strongly two continuous variables are related to each other. Correlation
allows us to quantify this relationship with a single index called the correlation coefficient (r).
Importantly, correlation does not imply causation. A significant correlation tells us that two variables
co-vary, but not that one causes the other.
1.2 Significance of Correlation in Psychological Studies
Correlation is central to psychological research for several reasons:
• Prediction: Knowing the relationship between, say, childhood stress and adult anxiety allows
psychologists to predict outcomes in clinical settings.
• Test Validity & Reliability: The validity of a psychological test is assessed by correlating test
scores with an external criterion. Reliability is assessed by correlating scores across time (test-retest)
or items (internal consistency).
• Theory Building: Correlation is the foundation of factor analysis and structural equation modelling
– key tools in psychology's theoretical arsenal.
• Effect Size: r2 (coefficient of determination) tells us what proportion of variability in one variable is
associated with variability in another – a direct measure of effect magnitude.
• Basis for Regression: Understanding correlation is prerequisite to regression analysis, which
allows prediction of one variable from another.
1.3 The Correlation Coefficient – Key Properties
• Ranges from –1.00 to +1.00.
• +1.00: Perfect positive relationship – as X increases, Y increases proportionally.
• –1.00: Perfect negative relationship – as X increases, Y decreases proportionally.
• 0.00: No linear relationship between X and Y.
• The sign denotes direction; the magnitude denotes strength.
• r = .75 and r = –.75 represent equally strong (but opposite-direction) relationships.
• The correlation coefficient is symmetric: r(X, Y) = r(Y, X).
Section 2: Significance of the Scatter Plot
2.1 What is a Scatter Plot?
A scatter plot (also called a scatter diagram) is a two-dimensional graph in which each data point
represents one observation. The predictor variable (X) is plotted on the horizontal axis and the
criterion variable (Y) on the vertical axis. Each point at coordinates (Xi, Yi) represents one
participant's scores on both variables.
2.2 Why Scatter Plots Matter
• Visualise the direction of a relationship – positive, negative, or no trend.
• Visualise the strength – tightly clustered points indicate a strong relationship; widely scattered
points indicate a weak one.
• Detect non-linearity – correlation only measures linear relationships; a scatter plot reveals
whether the relationship is actually curvilinear.
• Identify outliers – single extreme data points can distort a correlation coefficient; the scatter plot
makes these visible.
• Inspect regression line fit – by superimposing the regression line, you can visually assess how
well it fits the data.
• Check homoscedasticity – the scatter of points around the line should be roughly uniform across
all values of X; a fan-shaped scatter signals heteroscedasticity.
2.3 Types of Scatter Plot Patterns
Pattern Description Approximate r
Strong Positive Points cluster tightly along a line rising left to right .80 to +1.00
Weak Positive Points loosely form an upward trend .20 to .50
No Relationship Points scattered randomly with no visible trend ≈ .00
Weak Negative Points loosely form a downward trend –.20 to –.50
Strong Negative Points cluster tightly along a line falling left to right –.80 to –1.00
Curvilinear Points follow a curved path (e.g., U-shape or inverted U) ≈ 0 despite a clear relationship
r may be
Key point: Always examine a scatter plot before computing a correlation. Pearson's r can be
misleading when the relationship is non-linear or when outliers are present.
Section 3: The Covariance – Foundation of Correlation
3.1 Definition
The covariance (covXY or sXY) is a statistic that reflects the degree to which two variables vary
together. It is defined as:
cov_XY = Σ(X – X■)(Y – ■) / (N – 1)
This formula is structurally similar to the variance: replacing all Y values with X gives s2X, and
replacing all X values with Y gives s2Y.
3.2 Interpreting Covariance
• Large positive covariance: High X scores tend to pair with high Y scores (positive relationship).
• Large negative covariance: High X scores tend to pair with low Y scores (negative relationship).
• Near-zero covariance: No systematic linear co-variation between X and Y.
The limitation of covariance as a correlation measure is that its magnitude depends on the scale
(standard deviations) of both variables. A covariance of 1.34 means something very different
depending on whether the SDs are small or large. This is why we divide by the product of the SDs to
get the standardised correlation coefficient.
Section 4: Pearson's Product-Moment Correlation Coefficient (r)
4.1 Formula and Derivation
To remove the scale-dependency of covariance, we divide by the standard deviations of both
variables:
r = cov_XY / (s_X · s_Y)
Equivalently, in raw-score form (computational formula):
r = [ΣXY – (ΣX·ΣY)/N] / √{[ΣX² – (ΣX)²/N][ΣY² – (ΣY)²/N]}
Since the maximum value of covXY is ±sXsY, dividing ensures r is always bounded between –1 and
+1.
4.2 Step-by-Step Calculation
Step 1: Compute means: X■ and ■
Step 2: Compute standard deviations: s_X and s_Y
Step 3: Compute cov_XY = Σ(X – X■)(Y – ■) / (N – 1)
Step 4: Divide: r = cov_XY / (s_X · s_Y)
4.3 The Coefficient of Determination (r2)
Squaring r gives the coefficient of determination, r2, which tells us the proportion of variability in Y
that is associated with (predictable from) variability in X.
r² = SS_regression / SS_total = (SS_Y – SS_residual) / SS_Y
Example: If r = .53 between stress and psychological symptoms, then r2 = .28, meaning approximately
28% of the variability in symptoms is associated with variability in stress. The remaining 72% is
attributable to other factors.
4.4 Significance Test for r
To test H0: ρ = 0 (that the population correlation is zero), we use:
t = r√(N–2) / √(1–r²) on (N–2) degrees of freedom
If |tobt| > tcritical, we reject H0 and conclude a significant linear relationship exists between X and Y.
Example: With r = .532, N = 28: t = .532√26 / √(1 – .532²) = .532×5.099 / √0.717 = 2.713 / 0.847 =
3.20, df = 26. At α = .05 (two-tailed), tcrit = 2.056, so we reject H0.
4.5 Factors That Affect the Correlation Coefficient
• Range restriction: Restricting the range of X or Y usually reduces r. Example: correlating SAT
scores with college GPA only for admitted students under-estimates the true predictive validity.
• Heterogeneous subsamples: Combining data from distinct subgroups (e.g., males + females)
can artificially inflate r due to between-group differences. Always check for this.
• Outliers: A single extreme data point can substantially change r. The scatter plot will reveal
outliers.
• Non-linearity: r measures only linear association. A perfect curvilinear relationship can yield r ≈ 0.
• Reliability of measurement: Unreliable measures attenuate (reduce) r.
4.6 Adjusted Correlation Coefficient (r_adj)
For small samples, the sample r overestimates ρ (the population correlation). The adjusted
correlation provides a less biased estimate:
r_adj = √[1 – (1–r²)(N–1)/(N–2)]
For large N, radj ≈ r. For small N the difference can be substantial.
Section 5: Spearman's Rank Correlation Method (r_s)
5.1 When to Use Spearman's r_s
• Data are measured on an ordinal scale (ranks, ordered categories).
• Data are severely non-normal and you wish to avoid assumptions of normality.
• The relationship may be monotonic but non-linear (consistent direction but not straight-line).
• Judges have ranked objects or stimuli, and you want to assess inter-rater agreement.
5.2 Calculating Spearman's r_s
The simplest and most accurate method is to apply Pearson's formula directly to the ranked data.
Rank each variable separately, then compute r on those ranks exactly as for Pearson's r.
An alternative formula (valid only when there are NO ties) is:
r_s = 1 – [6ΣD²] / [N(N²–1)]
where D = difference between the two ranks for each observation, and N = number of pairs. When
ties are present, always use Pearson's formula on the ranked data instead, as the alternative formula
requires a correction for ties that is cumbersome.
5.3 Handling Tied Ranks
When values are tied, assign each the mean of the ranks they would have occupied. For example,
if two values are tied for ranks 4 and 5, both receive rank 4.5. If three values are tied for ranks 7, 8,
and 9, all three receive rank 8.
5.4 Worked Example
A teacher ranks 6 students on Maths (X) and Science (Y) ability:
Student Maths Rank (X) Science Rank (Y) D = X–Y D²
A 1 2 –1 1
B 2 1 +1 1
C 3 4 –1 1
D 4 3 +1 1
E 5 6 –1 1
F 6 5 +1 1
ΣD² = 6
r_s = 1 – 6(6) / [6(36–1)] = 1 – 36/210 = 1 – 0.171 = +0.829
Interpretation: There is a strong positive correlation (rs = .83) between students' ranks in Maths and
Science.
5.5 Significance Test for r_s
For N ≥ 10, the significance of rs can be approximated using:
t = r_s√(N–2) / √(1–r_s²) on (N–2) degrees of freedom
For small N (<28), use published tables of critical values of rs. Note that because ranks cannot be
normally distributed, there is no exact standard error for rs for small samples – the t approximation is
used cautiously.
5.6 Pearson's r vs Spearman's r_s
Feature Pearson's r Spearman's r_s
Scale of measurement Interval/Ratio Ordinal or above
Distribution assumption Bivariate normality None (distribution-free)
Type of relationship Linear Monotonic (can be non-linear)
Effect of outliers High sensitivity Reduced sensitivity
Ties in data Not applicable Use mean-rank correction
Preferred use Continuous, normally distributed data
Ranked, ordinal, or skewed data
Section 6: Simple Linear Regression
6.1 What is Regression?
Regression is a statistical method used to predict the value of one variable (the criterion or
dependent variable, Y) from another variable (the predictor or independent variable, X). While
correlation asks 'how strongly are X and Y related?', regression asks 'given a value of X, what value
of Y should we predict?'
6.2 The Regression Equation
The equation of the regression line is:
■ = bX + a
Where:
• ■ (Y-hat) = the predicted value of Y for a given X
• b = the slope – the predicted change in Y for each one-unit increase in X
• a = the intercept – the predicted value of Y when X = 0
6.3 The Least Squares Method
The regression line is determined by the method of least squares – we choose the values of a and
b that minimise the sum of squared residuals Σ(Y – ■)2. This gives us the line of best fit. The
resulting formulas are called the normal equations:
b = cov_XY / s²_X
a = ■ – b·X■
Alternative formula for b: b = r · (s_Y / s_X) (slope = correlation × ratio of standard deviations).
This shows that when sY = sX, then b = r.
6.4 Worked Example
Data: Stress (X) predicting psychological symptoms [ln(Symptoms)] (Y). From a sample of N = 107
students:
X■ = 21.29, ■ = 4.483, s_X = 12.49, s_Y = 0.202, cov_XY = 1.336, r =
.529
b = cov_XY / s²_X = 1.336 / 12.49² = 1.336 / 156.0 = 0.0086
a = ■ – b·X■ = 4.483 – (0.0086)(21.29) = 4.483 – 0.183 = 4.300
Final regression equation: ■ = 0.0086X + 4.300
Interpretation: For each additional unit of stress experienced, we predict an increase of 0.0086 in the
log of psychological symptoms.
6.5 Interpreting the Regression Coefficients
The slope (b):
• Represents the rate of change: one additional unit in X predicts b units change in ■.
• A positive slope: Y increases as X increases. A negative slope: Y decreases as X increases.
• Example: If b = 0.9 for salary regressed on years of experience (in thousands), each additional
year predicts $900 more in salary.
The intercept (a):
• The predicted value of Y when X = 0.
• Often mathematically meaningful but practically uninterpretable (e.g., predicted weight when height
= 0 is nonsensical).
• Centering X at its mean (subtracting X■) makes the intercept more interpretable – it becomes the
predicted Y for the average X.
6.6 The Standardised Regression Coefficient (β / Beta)
When both X and Y are standardised (mean = 0, SD = 1), the slope equals r. This standardised
coefficient β (beta) is interpreted as: "a one standard deviation increase in X is associated with β
standard deviations change in ■." This allows meaningful comparison of predictors measured in
different units.
6.7 Accuracy of Prediction – Standard Error of Estimate
The standard error of estimate (sY·X) measures the typical prediction error – how far actual Y
values fall from the regression line:
s_{Y·X} = √[Σ(Y–■)² / (N–2)]
The relationship between sY·X and r:
s_{Y·X} = s_Y · √(1–r²) · √[(N–1)/(N–2)] ≈ s_Y · √(1–r²) for large N
This shows that as r → ±1, sY·X → 0 (perfect prediction). As r → 0, sY·X → sY (using X gives no
advantage over predicting ■).
r 0.00 0.20 0.40 0.50 0.70 0.80 0.90 1.00
s_Y·X 1.00s_Y 0.98s_Y 0.92s_Y 0.87s_Y 0.71s_Y 0.60s_Y 0.44s_Y 0
Notice: even a correlation of r = .50 still leaves an error that is 87% as large as when X is unknown.
Prediction requires very high correlations to be practically useful.
Section 7: Assumptions and Limitations of Regression Analysis
7.1 Core Assumptions of Linear Regression
1. Linearity
The relationship between X and Y must be linear. The regression equation assumes a straight-line
relationship. Always check the scatter plot for evidence of curvature.
2. Homoscedasticity (Homogeneity of Variance in Arrays)
The variance of Y around the regression line must be constant across all values of X. This means
the spread of residuals should be the same whether X is small or large. Violation (heteroscedasticity)
produces inefficient estimates and distorts standard errors.
3. Normality in Arrays (Conditional Normality)
For hypothesis testing (testing b or setting confidence intervals on ■), the Y values at each value of X
must be approximately normally distributed around ■. This is directly analogous to the normality
assumption in the t-test.
4. Independence of Observations
Each data point must be independent. Repeated measurements on the same individual or cluster
sampling violates this assumption.
5. No Perfect Multicollinearity (for multiple regression)
Predictor variables should not be perfectly linearly related to each other. In simple regression, this is
not an issue.
7.2 Bivariate Normal Model (for Correlation)
When the goal is to estimate the population correlation ρ (rho), we assume the (X, Y) pairs come from
a bivariate normal distribution. This means:
• The conditional distribution of Y for any fixed X is normal.
• The conditional distribution of X for any fixed Y is normal.
• The marginal distributions of both X and Y are normal.
For testing H0: ρ = 0, the bivariate normality assumption is required. Without it, the significance test
on r may not be valid.
7.3 Limitations of Regression Analysis
Cannot establish causation
A significant regression equation shows that X predicts Y, but does not mean X causes Y.
Confounding variables (third variables correlated with both X and Y) can produce spurious regression
relationships.
Extrapolation beyond the data range
Predicting Y for values of X outside the range used to build the model is unreliable. The relationship
may be non-linear beyond the observed range. Example: predicting exam performance for a stress
score of 200 when data range from 0–60.
Outliers and influential points
Single observations can exert high influence (leverage) on the regression line. Always check for
outliers in both scatter plots and residual plots.
Restricted range reduces predictive power
If the predictor range is restricted (e.g., only high-achieving students), r and therefore the regression
slope will be attenuated.
Assumes a single best-fitting line
Standard linear regression fits one line to all data. If the true relationship is non-linear, prediction will
be systematically wrong in parts of the X range.
Residuals should be examined
Always plot residuals (Y – ■) against X and against ■. Patterns in residuals indicate violations of
assumptions (non-linearity, heteroscedasticity).
Section 8: Hypothesis Testing in Regression
8.1 Testing the Significance of the Slope (b)
To test H0: β* = 0 (population slope = 0), we use:
t = b / s_b where s_b = s_{Y·X} / (s_X · √(N–1)) on (N–2) df
In simple regression, this test gives the same result as the test on r. If the slope is significantly
non-zero, then X significantly predicts Y.
8.2 Confidence Interval on the Slope
CI(β*) = b ± t_{α/2} · s_b
If this interval does not contain zero, the slope is significant at level α.
8.3 Confidence Limits on an Individual Prediction (■)
When predicting Y for a new individual at X = Xi:
s'_{Y·X} = s_{Y·X} · √[1 + 1/N + (X_i – X■)² / ((N–1)s²_X)]
CI(Y) = ■ ± t_{α/2} · s'_{Y·X}
Notice three things: (1) the interval is wider for predictions further from X■; (2) the interval narrows as
N increases; (3) even at X = X■, there is prediction error. These are prediction intervals for individual
new cases, not confidence intervals on the line.
Section 9: Module 4 – Summary and Quick Reference
9.1 Key Formulas at a Glance
Statistic / Formula Expression
Covariance cov_XY = Σ(X–X■)(Y–■) / (N–1)
Pearson's r r = cov_XY / (s_X · s_Y)
Coefficient of determination r² = SS_regression / SS_total
Test of H■: ρ = 0 t = r√(N–2) / √(1–r²), df = N–2
Regression slope b = cov_XY / s²_X = r(s_Y/s_X)
Regression intercept a = ■ – b·X■
Regression equation ■ = bX + a
Standard error of estimate s_Y·X = √[Σ(Y–■)² / (N–2)]
Spearman's r_s (no ties) r_s = 1 – 6ΣD² / [N(N²–1)]
Test of r_s t = r_s√(N–2) / √(1–r_s²), df = N–2
Confidence interval on r CI(r') = r' ± z_{α/2} / √(N–3), then back-transform
9.2 Key Conceptual Points to Remember
• Correlation ≠ Causation. Always state this when reporting r.
• Always inspect the scatter plot BEFORE computing a correlation.
• r measures only LINEAR association. Spearman's r_s measures MONOTONIC association.
• r² (not r) is the direct measure of 'effect size' – proportion of shared variance.
• A statistically significant r is not necessarily a practically important r (especially with large
N).
• Regression minimises squared residuals (Method of Least Squares).
• The regression line always passes through the point (X■, ■).
• Prediction intervals on individual Y values are WIDER than confidence intervals on the
regression line.
• Outliers can dramatically change both r and the regression line. Always check.
• Range restriction almost always REDUCES r from its true value.
• Spearman's r_s should be calculated using Pearson's formula on ranked data (not the 'D²'
shortcut) when ties are present.
9.3 Decision Guide: Which Correlation Coefficient to Use?
Situation Use This Coefficient
Both variables continuous, approximately normal Pearson's r
One or both variables ordinal / ranked Spearman's r_s
One variable continuous, one dichotomous Point-biserial r_pb (= Pearson r)
Both variables dichotomous Phi coefficient φ (= Pearson r)
Multiple judges ranking multiple objects Kendall's W (Concordance)
Ordinal data, prefer rank-based coefficient Kendall's τ (Tau)
Notes compiled from: Howell, D.C. (2012). Statistical Methods for Psychology (Chapters 9 & 10). | For Psychology
Statistics – Module IV