0% found this document useful (0 votes)
2 views15 pages

W7 LectureNotes

These supplementary notes for Week 7 focus on two critical topics in research design: sample size calculation and confounding. Proper sample size ensures statistical validity, while controlling for confounding is essential for internal validity, as both are necessary to avoid misleading results. The document outlines methods for calculating sample size, the implications of Type I and Type II errors, and the importance of addressing confounding variables in research.

Uploaded by

shahana.bge
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views15 pages

W7 LectureNotes

These supplementary notes for Week 7 focus on two critical topics in research design: sample size calculation and confounding. Proper sample size ensures statistical validity, while controlling for confounding is essential for internal validity, as both are necessary to avoid misleading results. The document outlines methods for calculating sample size, the implications of Type I and Type II errors, and the importance of addressing confounding variables in research.

Uploaded by

shahana.bge
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

1 WEEK 7: SUPPLEMENTARY CLASS NOTES

1.1 About these notes


These supplementary notes accompany the Week 7 lecture. They provide additional depth, worked examples,
and conceptual explanations to support your understanding of two of the most practically important topics in
research design: calculating how many participants you need, and identifying when a third variable is
distorting your results.

Sample size calculation and confounding may seem like separate concerns, but they share a common thread.
Both are about validity. Getting the sample size right protects statistical validity. Controlling for confounding
protects internal validity. A study can be well-powered and still produce a misleading result if confounders
are not addressed — and a study can have confounding under control but lack the power to detect a real
effect. You need both.
2 SECTION 1 | Sample Size Calculation

2.1 Why sample size matters


Every study involves a trade-off. Recruiting more participants costs time and money. Recruiting too few
means your study may fail to detect a real effect — or may detect one that does not actually exist. Sample
size calculation is the process of determining the minimum number of participants needed to give your study
a reasonable chance of producing a valid result.

Underpowered studies are among the most common and costly errors in research. A study that is too small
will report a null result not because there is no effect, but because the study lacked the sensitivity to detect it.
This wastes resources, misleads future researchers, and may prevent effective interventions from being
adopted.

The consequences of getting it wrong


Too small a sample: You miss a real effect. The study reports no association when one truly exists. This is a
false negative — a Type II error.

Too large a sample: You detect statistically significant effects that are clinically trivial. Resources are wasted.
Participants are exposed to study procedures unnecessarily.

Sample size calculation sits at the intersection of statistics, ethics, and resource management.

Key principle: An underpowered study is not just a statistical problem — it is an ethical one. Exposing
participants to research that cannot answer its own question is difficult to justify in any ethics application.

2.2 Type I and Type II errors


Before calculating sample size, you need to understand what you are protecting against. There are two types
of errors in hypothesis testing:

H₀ is TRUE (no real effect) H₀ is FALSE (real effect exists)

Reject H₀ Type I Error (α) — False Positive You conclude an Correct decision — True Positive You
effect exists when it does not correctly detect the real effect

Fail to Correct decision — True Negative You correctly Type II Error (β) — False Negative You miss a
reject H₀ conclude no effect real effect

Alpha (α): The probability of a Type I error. Conventionally set at 0.05 — meaning you accept a 5% chance of
a false positive.

Beta (β): The probability of a Type II error. Conventionally set at 0.20 — meaning you accept a 20% chance of
missing a real effect.

Power (1 − β): The probability of detecting a real effect when it exists. With β = 0.20, power = 0.80 or 80%.

|2|
Common misconception: Many students confuse α and β. Alpha controls false positives (you see
something that is not there). Beta controls false negatives (you miss something that is). They are different
errors with different consequences. Lowering alpha tightens the false positive threshold but does not
automatically improve power.

2.3 Statistical power


Power is the probability that your study will detect an effect of a given size if that effect truly exists. A power
of 80% means that if the same study were repeated many times under identical conditions, it would detect
the true effect in 80 out of 100 repetitions — and miss it in 20.

Power is determined by four factors. Changing any one of them changes the required sample size:

Factor Effect on power when increased Effect on required N when increased

Effect size Power increases N decreases

Alpha (α) Power increases N decreases

Power (1−β) N/A — this is the target N increases

Variability (σ) Power decreases N increases

Practical heuristic: The most common power target in health research is 80% (β = 0.20, α = 0.05). Some
funders and journals now expect 90% power (β = 0.10). Check the specific requirements of your study
design and target journal before committing to a power assumption.

The Z-score values you need to know


Sample size formulas use Z-scores to represent the chosen alpha and power levels. Here are the values used
in practice:

Parameter Common value Z-score

Alpha (two-tailed) 0.05 1.96

Alpha (two-tailed) 0.01 2.576

Power 80% (β = 0.20) 0.842

Power 90% (β = 0.10) 1.282

|3|
2.4 Cochran's formula for cross-sectional studies
The most commonly used formula in health research surveys is Cochran's formula. It applies when you are
estimating a proportion in a population and you want to achieve a given level of precision.

Cochran's formula:

n₀ = (Z² × p × q) / d²

Symbol Term Meaning

n₀ Required sample size The number of participants needed

Z Z-score Reflects your chosen alpha level (1.96 for α = 0.05, two-tailed)

p Expected prevalence Estimated proportion with the characteristic of interest

q 1−p Proportion without the characteristic

d Margin of error Acceptable precision — the half-width of the confidence interval

Choosing your parameters


Prevalence (p): Use the best available estimate from published literature, national surveys, or pilot data. If
you genuinely do not know, use p = 0.50. This is the most conservative value — it maximises the required
sample size and therefore gives you the widest possible margin of safety.

Margin of error (d): This is how precise you need to be. A margin of error of 0.05 means your estimate will
be within ±5 percentage points of the true prevalence. Tighter margins require larger samples.

Alpha level (Z): Use 1.96 for the conventional 5% significance level (two-tailed). If your study requires a more
stringent threshold (e.g., α = 0.01), use Z = 2.576.

Practical heuristic: When prevalence is unknown, use p = 0.50. When you have a reasonable estimate from
prior literature, use that value — but check the sensitivity of your sample size to small changes in p. If the
required N changes dramatically between p = 0.20 and p = 0.30, you need to be confident in your estimate.

How prevalence and precision interact


The required sample size is not simply proportional to any single parameter. The interaction between p, q,
and e can be non-intuitive:

Prevalence (p) q = 1−p p×q Margin of error (d) Required n (α=0.05)

0.50 0.50 0.25 0.05 384

0.30 0.70 0.21 0.05 323

0.10 0.90 0.09 0.05 138

0.50 0.50 0.25 0.03 1068

0.50 0.50 0.25 0.10 97

|4|
Notice that halving the margin of error from 0.10 to 0.05 quadruples the required sample size. This is because
d appears squared in the denominator. Precision is expensive.

2.5 Finite population correction


Cochran's formula assumes an infinitely large population. When your study population is small and your
sample is a substantial fraction of it, the formula overestimates the required sample size. The finite population
correction (FPC) adjusts for this.

FPC formula:

n = n₀ / (1 + (n₀ − 1) / N)

N: the total size of the population you are sampling from.

n₀: the sample size calculated from Cochran's formula before correction.

n: the corrected (smaller) sample size.

Common misconception: The FPC is only worth applying when n₀ is more than about 5% of N. If your
study population is 50,000 and your uncorrected sample size is 384, the correction barely changes the
result. Save the FPC for studies with small, bounded populations — a single hospital ward, a specific school
year cohort, or an occupational group.

Worked example: FPC in practice


Suppose you want to estimate the prevalence of burnout among nurses in a hospital with 400 staff. You
calculate n₀ = 196 using p = 0.40, e = 0.07, Z = 1.96.

Applying FPC:

n = 196 / (1 + (196 − 1) / 400) = 196 / (1 + 0.4875) = 196 / 1.4875 ≈ 132

The corrected sample size is 132 — a saving of 64 participants, which is meaningful in a small hospital setting.

2.6 Design effect (DEFF)


Simple random sampling is assumed in Cochran's formula. In practice, most large surveys use cluster
sampling — selecting groups (villages, clinics, schools) rather than individuals. Cluster sampling is logistically
convenient but statistically less efficient, because individuals within the same cluster tend to be more similar
to each other than to the broader population.

The design effect (DEFF) quantifies this loss of statistical efficiency. It is the ratio of the variance under the
complex design to the variance under simple random sampling.

Adjusted sample size = n × DEFF

|5|
DEFF = 1.0: No efficiency loss — your design is equivalent to simple random sampling.

DEFF = 2.0: You need twice as many participants as a simple random sample to achieve the same precision.
This is common in household cluster surveys.

Practical heuristic: If you are using cluster sampling and have no prior estimate of DEFF, a default of 1.5 to
2.0 is commonly used in health surveys. Document your assumption clearly in your methods section. A
sensitivity analysis showing how your required N changes with different DEFF values is good practice.

2.7 Adjusting for non-response


Even a well-designed study will not achieve 100% participation. People decline to take part, cannot be
contacted, or withdraw before completing the study. If non-response is not accounted for, you will end up
with fewer participants than your sample size calculation required — and the study will be underpowered.

Adjusted sample size = n / (1 − non-response rate)

Example: if your required sample size after all adjustments is 200 and you expect 20% non-response, you
need to recruit 200 / (1 − 0.20) = 200 / 0.80 = 250 participants initially.

Common misconception: Non-response adjustment is not optional — it is part of every sample size
calculation. Even studies with strong engagement strategies rarely achieve 100% response rates. Build in a
buffer. A realistic estimate of non-response should be justified by data from similar studies in the same
population.

2.8 Full worked example


Let's work through a complete sample size calculation, step by step.

Research question: What is the prevalence of hypertension among adults attending primary care clinics in
rural Bangladesh?

Step 1: Define your parameters


p = 0.25 (25% prevalence, estimated from a recent national survey)

q = 0.75 (1 − p)

e = 0.05 (acceptable margin of error: ±5 percentage points)

Z = 1.96 (α = 0.05, two-tailed)

|6|
Step 2: Apply Cochran's formula
n₀ = (Z² × p × q) / e²

n₀ = (1.96² × 0.25 × 0.75) / 0.05²

n₀ = (3.8416 × 0.1875) / 0.0025

n₀ = 0.7203 / 0.0025

n₀ = 288.12 → round up to 289

Step 3: Apply design effect


The study uses cluster sampling (clinics as clusters). DEFF is assumed to be 1.5 based on similar surveys.

n (adjusted) = 289 × 1.5 = 433.5 → round up to 434

Step 4: Adjust for non-response


Expected non-response rate: 15% (based on prior clinic-based surveys in the region).

Final N = 434 / (1 − 0.15) = 434 / 0.85 = 510.6 → round up to 511

Key principle: Always round up, never down. Rounding down reduces your effective sample and may leave
the study underpowered. Be consistent: round up at every step or only at the final step — but document
which approach you took.

Step 5: Report your calculation


Sample size calculations should be reported transparently in your methods section. Include every parameter
and justify each one. Reviewers and ethics committees expect to see this. A one-sentence sample size
statement with no justification will raise questions about your study's rigour.

Reporting template:

"Sample size was calculated using the Cochran formula for estimating a proportion. Assuming a hypertension
prevalence of 25% based on [reference], a margin of error of 5 percentage points, and a significance level of
0.05 (Z = 1.96), the minimum required sample size was 289. Adjusting for a design effect of 1.5 due to cluster
sampling and an anticipated non-response rate of 15%, the final target sample size was 511 participants."

|7|
2.9 Precision vs power: a note on different study designs
Cochran's formula estimates the sample size needed for precision — that is, for estimating a proportion
accurately. It is appropriate for cross-sectional prevalence studies.

For analytical studies that test hypotheses — such as case-control studies, cohort studies, or randomised
trials — a different approach is needed. These designs use power-based sample size calculations that
incorporate effect size estimates and power targets, not just precision parameters.

Study design Sample size goal Key parameters

Cross-sectional survey Estimate prevalence with given precision p, e, α

Cohort study Detect difference in outcome rates Incidence rates, RR, α, power

Case-control study Detect difference in exposure frequency Exposure prevalence, OR, α, power

Randomised trial Detect treatment effect of given size Effect size, α, power, allocation ratio

|8|
3 SECTION 2 | Confounding

3.1 What is confounding?


Confounding occurs when a third variable distorts the apparent association between an exposure and an
outcome. The observed association is not a true causal relationship — it is partially or wholly explained by the
confounder.

Confounding is one of the most important sources of bias in observational research. It is not the same as
chance (random error) or information bias (measurement error). It is a systematic distortion of the exposure-
outcome relationship caused by a variable that is related to both.

Classic example: Studies in the mid-20th century found that coffee drinkers had higher rates of lung cancer
than non-coffee drinkers. This looks alarming — until you realise that coffee drinkers at the time were also
much more likely to smoke. Smoking is associated with both coffee drinking (exposure) and lung cancer
(outcome). It is a confounder. Once smoking was controlled for, the coffee-cancer association disappeared.

Key principle: Confounding is a feature of the real world, not just the data. It exists in the population
regardless of how you collect information. You cannot eliminate it after the study is done — you can only
control for it through design or analysis.

3.2 The three criteria for confounding


A variable is a confounder if it meets all three of the following criteria simultaneously:

# Criterion What it means

1 Associated with the exposure The confounder is more (or less) common in exposed versus
unexposed individuals

2 Independently associated with the The confounder predicts the outcome, even in the absence of the
outcome exposure

3 Not on the causal pathway The confounder is not an intermediate step between the exposure
and the outcome

All three criteria must be met. A variable that only satisfies one or two of them is not a confounder — though
it may be something else of interest, such as a mediator or an effect modifier.

Common misconception: Students often classify any third variable associated with the outcome as a
confounder. This is wrong. The variable must also be associated with the exposure AND must not lie on the
causal pathway. A variable that sits between the exposure and the outcome (a mediator) should not be
adjusted for — doing so will block the very pathway you are trying to estimate.

|9|
3.3 Directed acyclic graphs (DAGs)
A directed acyclic graph (DAG) is a visual tool for mapping causal assumptions about how variables relate to
each other. In a DAG, arrows represent hypothesised causal relationships. The direction of the arrow indicates
which variable causes which.

DAGs help you identify confounders, mediators, and colliders before analysis begins. They make your causal
assumptions explicit — which means reviewers can scrutinise them, and you can defend them.

Reading a DAG:

• E → O: The exposure causes the outcome directly.


• C → E and C → O: C is a confounder. It influences both the exposure and the outcome through
separate pathways.
• E → M → O: M is a mediator. It lies on the causal pathway from exposure to outcome. Do not adjust
for it if you want to estimate the total effect.
• E → Col ← O: Col is a collider. It is caused by both the exposure and the outcome. Adjusting for a
collider opens a spurious pathway and introduces bias.

Key principle: Draw a DAG before every analysis. It forces you to be explicit about your assumptions,
guides your choice of confounders to adjust for, and prevents common mistakes such as adjusting for
mediators or colliders.

3.4 Confounding, mediation, and effect modification


These three concepts are frequently confused. They describe fundamentally different phenomena and require
different analytical responses:

Concept Definition Position of third variable What to do

Confounding Third variable distorts the E-O Outside the causal pathway Control for it — in design or
association (a common cause) analysis

Mediation Third variable transmits the On the causal pathway Do not adjust — unless
effect of E on O (between E and O) estimating direct effect only

Effect The E-O association differs Neither on nor off the Stratify and report
modification across levels of the third pathway — it modifies the separately for each stratum
variable effect

Example — all three in one study:

Suppose you are studying the effect of physical activity on cardiovascular mortality:

• BMI: A confounder. People with higher BMI may exercise less AND have higher mortality for reasons
unrelated to exercise. Adjust for it.

| 10 |
• Reduced blood pressure: A mediator. Physical activity lowers cardiovascular mortality partly by
reducing blood pressure. Do not adjust for it if estimating the total effect of exercise.
• Sex: A potential effect modifier. The association between exercise and cardiovascular mortality may be
stronger in men than in women. Stratify and report effects separately.

3.5 Controlling for confounding: design-stage methods


The most powerful approach to confounding is to prevent it from occurring in the first place. Design-stage
methods are applied before data collection.

Randomisation
In a randomised controlled trial, participants are allocated to exposure groups by chance. If the sample is
large enough, randomisation distributes both known and unknown confounders equally across groups. This is
the most powerful method available for controlling confounding.

The critical advantage of randomisation is that it controls for confounders you do not even know about —
variables you have never measured, or variables whose existence you have not yet recognised. No other
method offers this.

Common misconception: Randomisation does not guarantee that the groups will be identical. In small
trials, imbalances in confounders can occur by chance. This is why balance checks and, if necessary,
stratified randomisation are conducted. In large trials, random imbalances tend to cancel out.

Restriction
Restriction limits eligibility to participants who fall within a specific range of the confounder. For example, if
smoking is a confounder in a study of alcohol and liver disease, you might restrict your sample to non-
smokers only.

The advantage is simplicity and elimination of the confounder from the sample entirely. The disadvantage is
reduced generalisability — findings apply only to the restricted group, not the broader population.

Matching
Matching pairs each exposed participant with one or more unexposed participants who share the same value
(or similar values) of the confounder. For example, matching on age means every exposed participant is
paired with an unexposed participant of the same age.

Matching is commonly used in case-control studies. It controls for the matched variables effectively but
introduces a complication: matched variables cannot be examined as exposures in subsequent analysis, and
matched data requires matched analysis methods (e.g., conditional logistic regression).

Practical heuristic: Restriction and matching both reduce sample size — restriction by narrowing eligibility,
matching by the ratio of cases to controls available. In resource-limited settings, consider whether the
confounding problem is severe enough to justify the reduction in sample size.

| 11 |
3.6 Controlling for confounding: analysis-stage methods
When confounders cannot be controlled at the design stage — as in most observational studies — they must
be addressed in analysis. Analysis-stage methods require that the confounder has been measured.

Stratification
Stratification divides the data into separate groups (strata) defined by the confounder, then estimates the
exposure-outcome association within each stratum. If the association is similar across strata, a pooled
(adjusted) estimate can be calculated.

Mantel-Haenszel method: The standard approach for pooling stratum-specific estimates in stratified
analysis. It produces a weighted average of the stratum-specific odds ratios or risk ratios, giving more weight
to strata with more participants.

Stratification is transparent and intuitive. Its limitation is that it becomes impractical when you have many
confounders — you quickly run out of participants in each cell.

Multivariable regression
Regression models allow you to estimate the association between an exposure and an outcome while
simultaneously adjusting for multiple confounders. The model holds confounders constant, isolating the
independent association between exposure and outcome.

Logistic regression: Used when the outcome is binary (e.g., disease yes/no). Produces an adjusted odds
ratio.

Linear regression: Used when the outcome is continuous (e.g., blood pressure in mmHg). Produces an
adjusted coefficient.

Cox regression: Used when the outcome is time-to-event (e.g., time to death). Produces an adjusted hazard
ratio.

Key principle: Including a variable in a regression model does not automatically make it a confounder.
Only include variables that satisfy the three criteria for confounding. Including variables that are not
confounders — especially mediators or colliders — can introduce bias rather than remove it.

| 12 |
3.7 Residual confounding
Residual confounding occurs when a confounder is measured imprecisely or incompletely, so that adjusting
for it in analysis does not fully remove its effect. It can also occur when an important confounder was never
measured at all.

Sources of residual confounding:

• A confounder is measured with error (e.g., self-reported physical activity as a proxy for actual activity).
• A confounder is categorised too broadly (e.g., adjusting for age in 10-year bands instead of
continuously).
• An important confounder was not included in the study (e.g., socioeconomic status not collected).

Residual confounding is one of the key limitations of observational research. It is why observational studies
are generally considered lower in the evidence hierarchy than randomised controlled trials, and why strong
observational associations should be interpreted with caution before being translated into policy.

Common misconception: Adjusting for measured confounders in a regression model does not mean your
study is free of confounding. The adjustment is only as good as the measurement. A poorly measured
confounder that is adjusted for in the model provides only partial protection against confounding bias.

| 13 |
4 KEY CONCEPTS | Week 7 quick reference

Term Definition

Sample size calculation The process of determining the minimum number of participants needed to detect a
true effect or estimate a population value with acceptable precision

Type I error (α) Rejecting a true null hypothesis — concluding an effect exists when it does not (false
positive)

Type II error (β) Failing to reject a false null hypothesis — missing a real effect (false negative)

Statistical power (1−β) The probability of detecting a true effect. Conventionally 80% (β = 0.20). Influenced by
effect size, alpha, variability, and sample size

Cochran's formula n₀ = (Z² × p × q) / e² — used to calculate sample size for estimating a proportion with
a specified margin of error

Prevalence (p) Expected proportion with the characteristic of interest. Use 0.50 when unknown
(maximises required N)

Margin of error (e) The acceptable half-width of the confidence interval. Tighter precision requires a larger
sample. N is proportional to 1/e²

Z-score Reflects the chosen significance level. Z = 1.96 for α = 0.05 (two-tailed); Z = 2.576 for α
= 0.01

Finite population Adjustment applied when the sample is a substantial fraction of a small, bounded
correction (FPC) population: n = n₀ / (1 + (n₀−1)/N)

Design effect (DEFF) Factor by which the required sample size must be multiplied when using cluster
sampling instead of simple random sampling

Non-response Inflating the required sample size to account for expected non-participation: adjusted
adjustment N = n / (1 − non-response rate)

Confounding Distortion of the exposure-outcome association by a third variable that is associated


with the exposure, independently associated with the outcome, and not on the causal
pathway

Confounder criteria Must satisfy all three: (1) associated with exposure, (2) independently associated with
outcome, (3) not on the causal pathway

Positive confounding The confounder inflates the observed association — making an exposure appear more
strongly associated with an outcome than it truly is

Negative confounding The confounder deflates or reverses the observed association — masking a real effect
or creating a spurious inverse association

Mediator A variable on the causal pathway between exposure and outcome. Adjusting for a
mediator blocks the pathway and underestimates the total effect

Effect modifier A variable that changes the magnitude or direction of the exposure-outcome
association across its own levels. Should be reported by stratum, not adjusted away

Collider A variable caused by both the exposure and the outcome. Adjusting for a collider
opens a spurious association between exposure and outcome

DAG Directed acyclic graph — a causal diagram used to map the relationships between
variables and identify confounders, mediators, and colliders

| 14 |
Randomisation The gold standard for controlling confounding. Distributes known and unknown
confounders equally across groups by chance

Restriction Limiting eligibility to a specific category of the confounder. Eliminates confounding by


that variable but reduces generalisability

Matching Pairing exposed and unexposed participants on confounder values. Common in case-
control studies. Requires matched analysis methods

Stratification Analysing the exposure-outcome association separately within categories of the


confounder. Can be pooled using the Mantel-Haenszel method

Mantel-Haenszel A technique for producing a confounder-adjusted pooled estimate from stratified


method analyses, weighting by stratum size

Multivariable Statistical models that simultaneously adjust for multiple confounders by holding them
regression constant while estimating the exposure-outcome association

Residual confounding Remaining confounding bias after adjustment, due to imprecise measurement, broad
categorisation, or unmeasured confounders

| 15 |

You might also like