Stats Module2
Stats Module2
MODULE – II
Inferential Statistics, Sampling and Hypothesis Testing
Detailed Study Notes
Source: Statistical Methods for Psychology, 7th Ed. — David C. Howell
Inferential Statistics Used to draw conclusions (inferences) about a population on the basis of
sample data. Because we rarely have access to an entire population, we must
use samples and statistical logic to make probabilistic statements about the
population.
The Central Role of The critical link in inferential statistics is between distributions and
Distributions probabilities. If we know the distribution of a statistic, we know the probability
that any particular value of that statistic will occur. This is the foundation for all
hypothesis testing.
Bell-Shaped and The curve is perfectly symmetric around its mean. The left and right halves are
Symmetric mirror images. The mean, median, and mode all coincide at the centre of a
normal distribution.
Asymptotic The tails of the distribution extend infinitely in both directions but never actually
touch the X-axis (limits are ±∞).
Defined by µ and σ The exact shape and location of the normal distribution is completely determined
by its mean (µ) and standard deviation (σ). Different values of µ and σ produce
different normal distributions.
Total Area = 1.0 The total area under the normal curve equals 1.0 (representing 100% of all
possible values). Areas under the curve directly correspond to probabilities.
Standard Normal A normal distribution with mean = 0 and standard deviation = 1, designated as
N(0,1). A single table (Appendix z) works for all normal distributions by
transforming raw scores to z-scores.
Interpretation A score of 60 from N(50,10) becomes z = (60-50)/10 = 1.0. This means the
score is exactly 1 SD above the mean. The shape of the distribution is
unchanged — only the numerical scale on the X-axis changes.
Using z-Tables The z-table gives the area (probability) above or below any z-value. To find the
probability that a randomly chosen score exceeds some value X: (1) convert X
to z, (2) look up the area in the tail from the z-table.
Conditions for Use (1) Fixed number of trials N. (2) Each trial has exactly two outcomes
(success/failure). (3) Probability of success (p) is constant across trials. (4)
Trials are independent of each other.
Mean and Variance Mean = Np | Variance = Npq | SD = √(Npq) Example: N=10, p=0.50 → Mean =
5, SD = √2.5 = 1.58
Discrete Distribution Unlike the normal distribution (which is continuous), the binomial is discrete —
outcomes are whole numbers only (3 heads, not 3.5 heads).
Approaches Normal As N increases and when both Np > 5 and Nq > 5, the binomial distribution
approaches the normal distribution. This allows us to use z-tests as
approximations for large-N binomial problems.
Shape When p = q = 0.50 the distribution is symmetric. As p and q diverge from 0.50,
the distribution becomes increasingly skewed. Positive skew when p < 0.50;
negative skew when p > 0.50.
■ Example: If the true probability of a correct response is p = 0.50, and a participant gets 9 correct out of 10
trials, p(X ≥ 9) = 0.011. This is less than 0.05, so we reject the null hypothesis that the person is guessing
randomly.
Sampling Distribution The distribution of values we would obtain for a given statistic if we drew an
(general) infinite number of samples of the same size from the same population and
calculated the statistic for each sample. Sampling distributions are almost
always derived mathematically, not empirically.
Sampling Distribution The specific distribution of all possible sample means (X■) of size n drawn from
of the Mean a population. This distribution tells us what values of X■ to expect by chance,
and how likely different values are.
Standard Error (SE) The standard deviation of a sampling distribution. For the sampling distribution
of the mean: SE = σ/√n. Reflects the typical amount by which a sample mean
deviates from the population mean µ. Larger n → smaller SE → sample means
cluster more tightly around µ.
Why We Need They allow us to evaluate whether an obtained sample statistic is likely or
Sampling Distributions unlikely to occur if the null hypothesis is true. Without sampling distributions,
statistical hypothesis testing would not be possible.
• Given a population with mean µ and variance σ², the sampling distribution of the mean will have:
• (3) Standard deviation (Standard Error) equal to σ/√n (i.e., σ_X■ = σ/√n)
• (4) A shape that approaches the NORMAL DISTRIBUTION as n increases — regardless of the shape of
the parent population.
If Population is Normal If the parent population is itself normal, the sampling distribution of the mean is
perfectly normal for any sample size, including n = 1.
Mean of Sampling The mean of the sampling distribution of X■ always equals the population mean
Distribution µ. Therefore, X■ is an unbiased estimator of µ — on average, sample means hit
the target.
SE Decreases with n As sample size increases, the standard error decreases (SE = σ/√n). Larger
samples produce sample means that cluster more tightly around µ and are
therefore more precise estimates of the population mean.
Student's t Distribution Derived by William Gosset ('Student', 1908). When s is substituted for σ, the test
statistic follows a t distribution with n−1 degrees of freedom. The t distribution is
wider and heavier in the tails than z, reflecting the extra uncertainty from
estimating σ. As n → ∞, t → z.
Sample A subset of the population that is actually measured. Samples are used because
measuring entire populations is usually impractical or impossible.
Parameter A numerical characteristic of a population. Parameters are the 'true' values that
we ultimately want to know about. Denoted by Greek letters: µ (population
mean), σ² (population variance), σ (population SD), ρ (population correlation).
Estimation The goal of inferential statistics is to use statistics to make accurate inferences
about parameters. A good statistic should be an unbiased estimator of its
parameter (e.g., X■ is an unbiased estimator of µ; s² is an unbiased estimator of
σ²).
Correlation ρ (rho) r
Proportion π or P p or p■
Size N n
Simple Random Every member of the population has an equal and independent chance of being
Sampling selected. Achieved by using a random number table or computer random
number generator. The gold standard for producing representative samples.
Example: choosing 50 students from a list of 500 by lottery.
Stratified Random The population is divided into mutually exclusive strata (subgroups) based on a
Sampling relevant characteristic (e.g., age, gender, diagnosis). Random samples are then
drawn independently from each stratum. Ensures that all important subgroups
are represented. Example: sampling 50 men and 50 women separately to study
gender differences.
Cluster Sampling The population is divided into naturally occurring clusters (e.g., schools,
hospitals, neighbourhoods). A random sample of clusters is selected, and all or
a random sample of members within chosen clusters are included. Efficient for
geographically dispersed populations. Example: randomly selecting 5 schools
and testing all students in those schools.
Multi-Stage Sampling A combination of sampling methods applied in stages. For example: first
randomly select states, then randomly select cities within states, then randomly
select households within cities. Common in large national surveys.
Convenience Sampling Participants are selected because they are easily accessible. Example: testing
students in one's own psychology class. Highly practical but potentially very
unrepresentative. The most common method in psychology research, despite its
limitations.
Purposive (Judgement) The researcher deliberately selects participants who they judge to represent the
Sampling population or who have specific characteristics needed for the study. Example:
selecting expert therapists for a study of treatment decision-making.
Snowball Sampling Existing participants recruit further participants from their networks. Useful for
hard-to-reach or stigmatised populations (e.g., people with rare disorders, illicit
drug users). Can produce highly biased samples.
Quota Sampling The researcher specifies quotas for different subgroups and fills them using
convenience sampling within each group. Resembles stratified sampling but
lacks randomness within strata — still non-probabilistic.
Reducing Sampling Sampling error decreases as sample size increases (SE = σ/√n). With n → ∞,
Error the sample mean X■ → population mean µ exactly. This is why larger samples
give more precise estimates.
Sampling Error vs. Bias Sampling error is random — sometimes the sample mean is above µ,
sometimes below — and averages to zero over many samples. Bias is
systematic error that consistently pushes the estimate in one direction; it is
caused by poor sampling design, not by chance.
Step 2: Null Hypothesis State the null hypothesis (H■) — the specific, testable opposite of the research
hypothesis. It always takes a precise form. Example: H■: µ = 50 (children's
mean does not differ from the population norm).
Step 3: Sampling Construct the sampling distribution of the test statistic assuming H■ is true. This
Distribution Under H■ specifies what values of the statistic we would expect if H■ were correct.
Step 4: Collect Data Draw a random sample and calculate the relevant sample statistics.
Step 5: Compare Calculate the probability (p-value) of obtaining a test statistic as extreme as (or
Statistic to Distribution more extreme than) the one obtained, given that H■ is true.
Step 6: Decision Reject H■ if p ≤ α (significance level). Retain H■ if p > α. Note: we never 'prove'
H■ true; we merely fail to reject it.
Alternative Hypothesis The research hypothesis — the statement we hope the data will support. In
(H■ or H■) practice, H■ is often just the negation of H■: H■: µ ≠ 100; H■: µ■ ≠ µ■. A
specific alternative hypothesis is required for power calculations.
Simple vs. Composite A simple hypothesis specifies a single population value (H■: µ = 50). A
composite hypothesis specifies a range (H■: µ > 50). H■ is usually simple
(allowing us to compute the sampling distribution); H■ is usually composite.
Why Two-Tailed Tests (1) Researchers often cannot be certain of the direction in advance. (2)
Are Preferred Unexpected findings in the other direction would be missed. (3) Changing from
one-tailed to two-tailed after seeing the data is scientifically improper — it
inflates the true α level. Most journals expect two-tailed tests.
Critical Values at α = Two-tailed: z = ±1.96 (for z-test); t value depends on df. One-tailed: z = ±1.645
.05 (for z-test); t value depends on df.
■ Important: The choice between one-tailed and two-tailed must be made BEFORE collecting data, based on
theoretical considerations. Switching after seeing results is equivalent to running an uncontrolled Type I error
rate of 7.5%.
Significance Level of We reject H■ if there is less than a 5% probability that we would obtain these
.05 data (or more extreme data) if H■ were true. Equivalently, we would make a
Type I error (falsely rejecting a true H■) only 5% of the time.
Significance Level of More conservative threshold. Reject H■ only if p ≤ .01. Reduces Type I error
.01 probability to 1% but increases Type II error probability β.
Rejection Region The area(s) in the tails of the sampling distribution where test statistic values
lead to rejection of H■. Determined by α and the number of tails.
Critical Value The boundary of the rejection region. For a two-tailed z-test at α = .05: critical
values are ±1.96. For a t-test, critical values depend on df.
p-Value The exact probability of obtaining a test statistic as extreme as the one
observed, assuming H■ is true. We reject H■ when p ≤ α.
Type I Error (α) Rejecting H■ when it is actually true. A 'false positive.' The probability of Type I
error equals α (the significance level). Set by the researcher in advance.
Example: concluding a drug works when it actually does not.
Type II Error (β) Failing to reject H■ when it is actually false. A 'false negative.' Denoted β (beta).
Example: concluding a drug does not work when it actually does. β depends on
α, effect size, sample size, and σ².
Relationship Between Increasing α (making it easier to reject H■) decreases β and increases power —
α and β but at the cost of more Type I errors. Setting α = .05 typically produces β
somewhere between .20 and .80, depending on n and effect size. There is
always a trade-off between α and β.
Power = 1 − β The probability of correctly rejecting a false H■. A high-power test rarely makes
Type II errors. Recommended standard: power ≥ .80 (Cohen, 1988), meaning β
≤ .20.
Practical Importance Most researchers focus on minimising α but neglect β. Because conducting an
experiment that has only a 30% chance of detecting a real effect (power = .30) is
a waste of resources, power analysis should be conducted before data
collection.
■ α and β cannot both be simultaneously minimised unless sample size is increased. This is why power
analysis is conducted before running a study — to ensure adequate sample size for detecting the expected
effect at acceptable error rates.
■ Assumptions
• The sample is a random sample from the population of interest.
• The dependent variable is measured on at least an interval scale.
• The population is normally distributed, OR n is large enough (≥30) for the CLT to ensure a normal
sampling distribution of X■.
Degrees of Freedom df = n − 1
Standard Error of the SE = s/√n — this is the denominator of the t-formula. It estimates how much
Mean sample means vary around µ due to sampling error alone.
■ Worked Example
Nurcombe et al. studied Psychomotor Development Index (PDI) scores in low-birthweight (LBW)
infants. Population norm: µ = 100, σ unknown.
Sample of n = 56 LBW infants: X■ = 104.125, s = 12.584
H■: µ = 100 H■: µ ≠ 100 α = .05 (two-tailed)
t = (104.125 − 100) / (12.584/√56) = 4.125 / 1.682 = 2.45
Critical value: t.025(55 df) ≈ 2.009. Since 2.45 > 2.009, reject H■.
Conclusion: LBW infants scored significantly differently from the normative population mean on the PDI
(t(55) = 2.45, p < .05, two-tailed).
Point Estimate A single value used to estimate a population parameter. Example: X■ = 1.463
as an estimate of µ for the moon illusion study. Provides no information about
precision or uncertainty.
Confidence Interval A range of values calculated from sample data that will, with a specified
probability (confidence level), contain the true population parameter. It captures
both the estimate and its precision.
95% Confidence CI■■ = X■ ± t.025(df) × (s/√n) Interpretation: If this procedure were repeated
Interval many times (with new random samples), 95% of the intervals calculated would
contain the true µ. Example from moon illusion data: CI■■ = 1.219 ≤ µ ≤ 1.707.
99% Confidence CI■■ = X■ ± t.005(df) × (s/√n) Wider interval because we require higher
Interval certainty of capturing µ. Moon illusion example: CI■■ = 1.112 ≤ µ ≤ 1.814.
Link to Hypothesis If the hypothesised µ■ falls outside the 95% CI, then a two-tailed test at α = .05
Testing would reject H■: µ = µ■. The CI inverts the hypothesis test to show all values of
µ■ that would not be rejected. A CI includes far more information than a single
hypothesis test.
■ Wider confidence intervals indicate less precision (more uncertainty). Precision increases with larger n and
smaller s. The width of a CI is: 2 × t_{α/2} × s/√n.
Variance Sum Law When two variables are independent, the variance of their sum OR difference
equals the SUM of their variances: σ²(X■■ − X■■) = σ²■/n■ + σ²■/n■
Sampling Distribution The sampling distribution of (X■■ − X■■) is approximately normal for
Shape reasonable sample sizes (by CLT). Its mean = µ■ − µ■; under H■, mean = 0.
t Formula t = (X■■ − X■■) / √[s²_p(1/n■ + 1/n■)] Under H■: µ■ − µ■ = 0 (this term drops
out of the formula).
Degrees of Freedom df = n■ + n■ − 2
Assumption of Equal Pooling assumes σ²■ = σ²■ (homogeneity of variance). If equal, s²_p is a better
Variances estimate than either sample variance alone. If variances are unequal (use
Levene's test), use Welch's correction which adjusts the df.
■ Assumptions
• Both samples are independently and randomly drawn.
• The dependent variable is measured on at least an interval scale.
• Both populations are normally distributed (or n is large enough).
Why Power Matters Most researchers focus on controlling Type I errors (α) but neglect Type II
errors. An experiment with power = .30 has only a 30% chance of detecting a
real effect — a costly waste of time and resources. Cohen recommended a
minimum power of .80 for psychological research.
Power = 1 − β If β = .20 (acceptable standard), then power = .80. This means that when H■ is
false (a real effect exists), we correctly reject H■ 80% of the time and make a
Type II error only 20% of the time.
2. Effect Size (d) The larger the true difference between µ■ and µ■ (relative to σ), the further
apart the H■ and H■ distributions are, and the less overlap there is → greater
power. Power depends on the ACTUAL magnitude of the effect, which the
researcher cannot directly control.
3. Sample Size (n) Increasing n decreases the standard error (σ/√n), which shrinks both sampling
distributions → less overlap → greater power. This is the most practical way to
increase power. Doubling n does not double power but increasing n from 10 to
100 dramatically increases power.
4. Population Variance A smaller σ² produces a smaller standard error → smaller, less overlapping
(σ²) sampling distributions → greater power. Researchers can sometimes reduce σ²
by using more precise measurement instruments or tightly controlled laboratory
conditions.
5. The Statistical Test Parametric tests (t, F) are generally more powerful than non-parametric
Used equivalents when their assumptions are met. When assumptions are violated,
non-parametric or resampling tests may be more powerful.
Estimating d Three methods: (1) Prior research — use means and SDs from published
studies. (2) Practical significance — decide the minimum difference that would
be scientifically meaningful. (3) Cohen's conventions.
• Note: Use these conventions only as a last resort when no prior knowledge is available. Using them
mindlessly instead of estimating d from the substantive context simply replaces one form of thoughtlessness
with another (Thompson, 2000).
How to Use (1) Estimate d from prior research or practical importance. (2) Calculate δ = d ×
√n. (3) Look up power for the obtained δ and chosen α in a power table.
Solving for Required n To find n needed for a desired power level: (1) Specify desired power (e.g., .80)
and α (e.g., .05). (2) From the power table, find the required δ. (3) Solve: n =
(δ/d)²
δ = d × √n (one-sample t-test)
■ Example: A clinical psychologist expects d = 0.33 (IQ difference of 5 points, σ = 15). For n = 25: δ = 0.33 ×
√25 = 1.65. Power table gives power ≈ .38. For power = .80 with α = .05, required δ = 2.80, so n = (2.80/0.33)² ≈
72.
Standard Normal (z) Normal distribution with µ=0 and σ=1. All normal distributions can be
converted to z using z = (X−µ)/σ.
Sampling Distribution Distribution of all possible values of a statistic computed from an infinite
number of same-sized samples from the same population.
Standard Error (SE) Standard deviation of a sampling distribution. For means: SE = σ/√n.
Central Limit Theorem As n increases, the sampling distribution of X■ approaches normal with
mean µ and variance σ²/n, regardless of the population shape.
Parameter A numerical characteristic of a population (µ, σ², ρ). The 'true' value that
statistics are used to estimate.
Statistic A numerical value calculated from a sample (X■, s², r) used to estimate the
corresponding population parameter.
Sampling Error Natural, chance variation in a statistic from sample to sample. Not a
mistake; decreases as n increases.
Simple Random Sampling Every population member has an equal, independent chance of selection.
Gold standard for producing representative samples.
Null Hypothesis (H■) Hypothesis of no difference/effect. Always stated precisely so the sampling
distribution can be constructed under it.
Alternative Hypothesis (H■) Research hypothesis; the substantive claim the researcher hopes to
support.
Type I Error (α) Rejecting H■ when it is true. Probability = α (significance level). 'False
positive.'
Type II Error (β) Failing to reject H■ when it is false. 'False negative.' Probability = β; related
to power by: Power = 1 − β.
One-Sample t-Test Tests whether a single sample mean equals a specified population mean
µ■ when σ is unknown. t = (X■ − µ■) / (s/√n), df = n−1.
Independent Samples t-Test Tests whether means of two unrelated groups differ. t = (X■■ − X■■) /
√[s²_p(1/n■ + 1/n■)], df = n■+n■−2.
Confidence Interval Range of plausible values for a parameter. CI■■ = X■ ± t.025 × s/√n. '95%
of intervals calculated this way contain µ.'
Level of Significance (α) Pre-set probability threshold for rejecting H■. Typically .05 or .01. Defines
the maximum acceptable Type I error rate.
Effect Size (d) Cohen's d = (µ■ − µ■) / σ. Standardised measure of effect magnitude
independent of sample size. Small=.20; Medium=.50; Large=.80.
One-Tailed Test Rejection region in one tail only; used for directional hypotheses. Must be
decided before data collection.
Two-Tailed Test Rejection region split between both tails; tests H■: µ ≠ µ■. Most commonly
used in psychological research.
Pooled Variance Weighted average of two sample variances: s²_p = [(n■-1)s²■ + (n■-1)s²■]
/ (n■+n■−2). Used when σ²■ = σ²■.