CASE CONTROL STUDY
DR. MA. CARMEN C. TOLABING
Professor
Department of Epidemiology and Biostatistics
UP Manila
Session Objectives
General Objective:
Apply the principles in the design, conduct, analysis and
interpretation of Case control (CC)studies
Specific Objectives
1. Define case control study design
2. Describe the steps in doing a CC study
3. Compute and interpret the measure of association
4. Identify strengths and weaknesses of CC study
2
A. Definition
Classification
§ analytic E ? D
§ observational
§ Retrospective
§ Longitudinal
Study population (study base)
§ Cases ( + disease, D)
compare proportion exposed (+ factor)
§ Control (- disease,
Objective
§ To show that the odds of exposure is greater in those with
D than in those
Review: cohort
Factor Disease
Present Future
Past
Measure of association = RR (incidence measure)
Nature of association = > 1 risk factor; <1 protective factor
Case control
Factor Disease
Past Present Future
Exposure Disease
?
Case Control
?
Exposure Disease
?
Retro Cohort
?
Exposure Disease
?
Pros Cohort
?
Past Present Future6
Cross-sectional
E prospective cohort D
E retrospective cohort D
E case control D
past present future
7
? Factor Disease
+
-
+
-
-
past Present
Figure 9. Case control study design
Schematic Diagram
Choice of Pop.
STUDY POPULATION
Classification of subjects
CASES (+ disease) CONTROLS (- disease)
Assessment of Factor
W/ FACTOR W/O FACTOR W/ FACTOR W/O FACTOR
Analysis
9
B. Process
1. Definition and selection of cases
1.1 Establish objective criteria
Diagnostic criteria for the disease
Eligibility criteria
- may be problematic if Dx procedure is
expensive n ≠N
10
1.2 Select cases
Sources:
1. hospitals (secondary or case-defined base)
2. population (primary study base)
Advantage: avoids bias that may arise from selective
factors that guide affected individuals to a particular medical
care facility or physician
Disadvantage: cost and logistical considerations
11
Types:
1. prevalent cases
2. incident cases
Advantages:
2.1 uniform diagnosis
2.2 accurate recall of events
2.3 temporal sequence
2.4 non-survivors
Method of selection
1. total enumeration
2. random sampling
12
2. Definition and selection of controls
2.1 Define control group
comparable to the source pop of the cases
matching
• 1-to-1 matching
• category matching
matching variables: characteristics known to influence
the distribution of the disease/exposure
to miminize confounding
13
2.2 Select controls
(get controls from the same source population
as the cases)
Source: hospital
general population
special groups
Method: 1. random sampling
2. paired sampling (matched)
14
Types of control
hospital control
community control
special control
multiple control groups
---When in doubt as to the correct/ appropriate
control
Potential Problem: inconsistent results
Advantage: if consistent results, could
indicate causality
15
Optimal case control ratio = 1:1
If cases, can controls to increase power of
the study to 1:4
16
3. Ascertainment of Exposure
(data collection)
Operational definition of exposure variable
Sources : subjects (family respondent)
medical records
Method of data collection:
same for the 2 groups (cases and controls)
Reference point
> hypothesized time response
> basis on which an individual should be considered
exposed/unexposed
17
4. Analysis of data
Measure of association : ODDS RATIO (OR)
Odds
the probability that the event will occur divided by
the probability that the event will not occur in a
population
odds of an event is the
number of those who experience the event in a pop
number of those who do not experience the event in same pop
*++, *- . /0*12 34
Odds Ratio OR = 5++, *- . /0*12 36
18
*++, *- . /0*12 34
OR= Disease
5++, *- . /0*12 36 D D
Exposure
Odds of exposure (D)
= (a/a+c) ÷ (c/a+c)
= a/c E a b
Odds of exposure (D-)
= (b/b+d) ÷ (d/b+d)
E c d
= b/d
OR = (a/c) ÷ (b/d) = ad/bc Total a+c b+d
OR = ad/bc
OR = 1, no association
OR ≠ 1, association
19
OR Conclusion/interpretation
=1 No association Exposure to the factor is not associated
with the disease; there is no association
between factor and disease
>1 Association present There is an association between exposure
and disease
odds E among D
Exposure > risk
<1 Association present There is an association between exposure
and disease
odds of E among
Exposure > protective
Odds vs. Risk (incidence)
Odds:
odds of an event is the number of those who experience
the event divided by the number of those who do not
Ex-1: If 2 of 10 people have the E
Exposure Odds = 2/8 or 0.25
Exposure Risk = 2/10 or 0.20
Ex-2: If 2 of 10 people have Disease
Disease Odds = 2/8 or 0.25
Disease Risk = 2/10 or 0.20
Odds ≈ Risk under certain conditions
21
OR Conclusion/Interpretation
=1 Exposure to the factor is not associated with the disease; there is
no association between factor and disease
>1 As odds:
Ø The odds for exposure, relative to non-exposure, was OR
Ø The odds for exposure among those with disease (D) is OR times more
than those without the disease (D)
As estimate of risk
The estimated risk of disease among exposed (E) is OR X times more
compared to the unexposed (E)
Ex-1 OR > 1
OR = 40.4 (smoking and lung cancer) Stellman, 2010
Association present
The OR for smoking for lung cancer, relative to
those without lung cancer, was 40.4X
Interpretation:
1. Smoking was associated with lung cancer
2. The odds for smoking among lung cancer group was
40.4X more compared to the odds for smoking among
non-lung cancer group
3. The estimated risk of lung cancer was 40.4 X more
among smokers compared to non-smokers (under
certain assumptions)
Interpretation
OR ≠ 1
>1 <1
The odds for _____(E)
among _________ (D) was
______ (OR) times more
compared to the _____ ( )
OR
The estimated risk of ____ (D) is
OR times more for ___(E) than
those __( ) )
OR Conclusion/interpretation
<1 As odds
Ø The odds for exposure, relative to non-exposure is OR
Ø The odds of exposure among those with disease (D), is
lower compared to those without the disease ( )
As estimate of risk
For an OR <1, get its reciprocal (1/OR) to obtain the measure of
the estimated risk for those unexposed ( )
Ø The estimated risk of disease D for the is 1/OR times more
compared to the E
Ex-2 OR <1
OR = 0.167 (Vaccination and Disease)
Association present
The odds for vaccination for dengue group, relative to non-dengue
group, was 0.167
Interpretation
1. Vaccination was associated with Disease
2. An OR <1 means the D group was less likely to experience the E
> The odds of vaccination (E) in the Dengue group (D) is less than
the odds in the non-Dengue group ( )
Vaccination X Dengue No Dengue
(+) 10 40
(- ) 90 60
OR = 0.167
total 100 100
3. OR not directly interpretable for risk estimation, so get reciprocal
26
Q: How to interpret the OR < 1
= reciprocal of the OR, 1/OR
= reverse the rows
Reciprocal of OR Reverse the rows
= 1/OR OR = ad/bc
= 3600/600 = 600/3600
= 1/0.167 = 6.0 (E- risk) = 0.167
= 6.0 (E- risk) > E protective
27
Interpretation
OR ≠ 1
>1 <1
The odds of _____( E) The odds of _____ (E)
among _____ (D) was ___(OR) times among _____(D) was less than ____ (D-)
more compared to the ____ ( )
The estimated risk of ____ (D) is __(OR) Reciprocal of OR
times more for ___ (E, exposed) The estimated risk of _____ (D) is
than the __ (E ,unexposed) __ (1/OR) times more for
___(E, unexposed) than the __ (E, exposed)
Determine statistical significance of the OR
+ Disease - Analysis of Single Table
+--------+--------+ Odds ratio = 4.00 (2.59 - 6.20)
+| a b Cornfield 95% confidence limits for OR
+--------+--------+ Relative risk = 2.00 (1.64- 2.43)
-| c d Taylor Series 95% confidence limits for RR
+--------+--------+ Ignore relative risk if case control study.
E
x Chi-Squares P-values
p ----------- --------
o Uncorrected : 45.00 0.0000000 <---
s Mantel-Haenszel: 44.90 0.0000000 <---
u Yates corrected: 43.66 0.0000000 <---
r
e F2 More Strata; <Enter> No More Strata; F10 Quit
29
Review (inferential stat concepts)
• Statistical significance (SS)
Estimation approach
- compute for Confidence Interval
P-value approach
- compute for p-value
Interpretation of test values
Confidence interval
• 95% confidence interval for a OR
= the range within which the true OR is likely
to fall with 95% confidence
p value (simplified def)
probability of observing a difference between the
exposed and unexposed groups due to random
error
>2 exposure categories
Cigarette Cases Controls OR
smoking
1. Non-smoker
2. Moderate
3. Heavy
> IDy a reference group )
- Heavy vs. nonsmoker
32
- Moderate vs. nonsmoker
C. Strengths and limitations
Strengths
Quick and inexpensive compared to cohort
Suited to diseases with long latency
Optimal for rare diseases
Can examine multiple etiologic factors for a single
disease
33
Limitations
Inefficient for rare exposures
Cannot generate incidence of disease
Difficult to establish temporal sequence
Prone to bias –over/underestinate of the OR
34
Example-1
Designing case control
Hypothesis
“Living with a sputum positive adult PTB
case for one year is associated with
development of PTB in children 7 years
old and below”
Study population
<7 year old children residing in Province X
Operational definition of study variables
Exposure variable
– living in a household with an adult who was
diagnosed to be sputum positive for PTB
Outcome – PTB diagnosed based on the WHO
guidelines.
1. Define and select a case
Case:
<7-year old child who has PTB based on the
WHO criteria of the disease
Assume: n1 = 100
2. Define and select a control
Control
<7-year old child who does not have PTB based
on the WHO criteria for presence of disease
Assume: n0=100
3. Ascertain exposure
Data to collect
- History of exposure to the factor
(Note: exposure should have occurred at least 12
months before the diagnosis of PTB – reference
point)
Method
- Face-to-face interview
4. Analyze data
4.1 Construct 2x2 table (hypothetical data)
At the beginning of the study
Factor Disease Total
+ -
+
-
Total 100 100 200
n1 n0 (n)
After data collection
Factor Disease Total
+ -
+ 20 3 23
- 80 97 177
Total 100 100 200
n0 (n)
4.2 Compute and interpret
Odds Ratio(OR)
OR =
Conclusion/Interpretation
= 1: no association
> 1: association, risk factor
< 1: association, protective factor
OR Interpretation/Conclusion
=1
>1
Example -2
Exercise and MI
At the beginning of the study:
Exercise (+)MI (-)MI
(+) ? ?
(- ) ? ?
total 100 100
45
Example: Exercise and MI
At the end of the study:
Exercise (+)MI (-)MI
(+) 20 (a) 40 (b)
(- ) 80 (c) 60 (d)
total 100 100
OR =
46
• OR =
• Conclusion =
• Interpretation =
47
OR Interpretation/Conclusion
Disease
Question: D D
What is the
Exposure
estimated risk of
disease among
those who do not E a b
have the E
E c d
> Get the reciprocal
_____ ?
49
SUMMARY
1. Define case control study design
2. Describe the steps in doing a CC study
3. Compute and interpret the measure of
association
4. Identify strengths and weaknesses of CC
study
Study type Variable being Statistical Measure of
measured from measure association
pop
Cohort Outcome of incidence RR
interest
Case control History of odds OR
exposure
THANK YOU