Multivariable methods
Ojuka
Introduction
• In statistical inferences, whether using estimation or hypothesis
testing, there is a possibility that the differences or associations might
be due to other characteristics or variables
• Regression tests the interrelationships among several risk factors or
exposure variables and a single outcome .
• Multivariable methods addresses confounding and effect
modifications.
Confounding is present when the relationship between a risk factor and an
outcome is modified by a third variable called confounding variable.
Effect modification occurs when there is a different relationship between the
exposure or risk factor and the outcome depending on the level of another
characteristics or variable
On both situations, the third variable can exaggerate or mask the
association between risk factor and outcome
Analytical methods for multivariable
methods
• For confounding factor, an estimate is generated for
the association of exposure to outcome adjusting for the
confounder.
• For effect modification, since the relationship depend on
the level, after the estimate is made , there is a
presentation for different levels .
Confounding factor
• Smoking versus cardiovascular disease- suppose we
tested and found the estimate to be RR of 2.6 (1.5-4.1)
95% CI.
• This means those who smoke has 2.6 risk of developing
cardiovascular disease.
• However, suppose smokers do not exercise hence have
high cholesterol as other factors.
• Multivariable methods can be used to address these by
adjusting the magnitude of association for the impact of
the other variable eg association between smoking and
cardiovascular disease adjusted for exercise and
cholesterol
Effect modification
• Suppose we are interested in efficacy of a drug
designed to lower cholesterol; a trial shows it is
effective compared to placebo.
• Suppose the reduction is only present in participants
with specific genetic marker and there is no reduction
without the genetic marker.
• This is effect modification or statistical interaction.
• The strategy here will be to present separate results
according to the third variable.
Multivariable methods use
• They can also be used to consider several risk factors
simultaneously to assess the relative importance of
each regarding a single outcome variable.
• For example, in the Framingham heart study, these
included age, sex, systolic blood pressure, total and
HDL, current smoking status and diabetes status.
Data analysis plan
• It is important to note that the statistical analysis of any study should
begin with complete description of the study data using summary
statistics ,
• Then proceed to generate CI/hypothesis testing
• These are crude analyses as they focussed exclusively on the
associations between one risk factor and one outcome.
• Multivariable methods are used after the study data are described and
after unadjusted analyses are performed.
• In clinical trials setting, the unadjusted analyses are the final ones due
to the randomization component that eliminate confounding factors.
• Regression should not be relied on to correct the problems in a study.
Confounding
• A confounding variable is • A study designed to
one that related to the assess the association
exposure or risk factor of between
interest and the outcome. obesity(BMI>30kg/m2 and
• Identification requires a incidence of
strong working cardiovascular disease
knowledge of the • Participants aged 35-65
substantive area under years, free from
investigation Incident CVD cardiovascular
No cvd and
Total
Obese 46 followed
254 up for
300 10 years.
Not obese 60 640 700
Total 106 894 1000
Example
• Incidence of CVD among obese Age Obese Not Total
46/300=0.153 Obese
• Proportion of none obese with <50 100 500 600
disease is 60/700=0.086 >50 200 200 400
• 10 years incidence of CVD Total 300 700 1000
0.153/0.086=1.78
• Many studies have shown an Age CVD No CVD Total
increase in the risk of CVD with <50 45 555 600
advancing age. >50 61 339 400
• Is age a confounder? Total 106 894 1000
• We dichotomous age >50 (400),
<50 600
•
Example
Age<50 • Obesity Incidence in >50
Incidence No CVD Total =200/400=0.5
CVD • Incidence in <50=100/600=0.167
Obese 10 90 100 • Obesity in >50 is 200/300=0.667
Not 35 465 500 • Obesity in <50 is 200/700=0.286
obese • CVD in <50 who are obese 10/100=0.1
Total 45 555 600 • None obese 35/500=0.070
• RR=0.1/0.07=1.43
Age>50
• CVD in >50 who are obese 36/200=0.180
Incidence No CVD
CVD • CVD in non obese >50 25/200=0.125
Obese 36 164 200 • RR=0.180/0.125=1.44
• These figure suggest strong association
Not 25 175 200
between obesity and incidenct of CVD
obese RR1.78
Total 61 339 400
Hypothesis
• Set hypothesize • Compute the test statistics
H0: Age and obesity are
• Expected frequency=row total
independent
H1: H0 is false x column total/n
Age Obese Not Total
α=0.05 obese
<50 100(180) 500(420) 600
• Select the appropriate test
statistics >50 200(120) 200(280) 400
Χ2=Σ(O-E)2/E Total 300 700 1000
Age CVD No CVD Total
• Set up the decision rule <50 45(63.6) 555(536.4 600
df =(2-1)(2-1)=1 for 5% )
critical value is 3.84, so Reject >50 61(42.4) 339(357.6 400
H0 if Χ2 ≥3.84 )
Total 106 894 100
u
Χ2=Σ(O-E)2/E
=(100-180)2/180+ (500-420)2/420 + (200-
• Compute the test statistics
120)2/120 +(200-280)2/280 2=Σ(O-E)2/E
=35.56+15.24+53.33+22.86=126.99 =(45-63.6)2/63.6+ (555-
• Conclusion 536.4)2/536.4 + (61-42.4)2/42.4
We reject Ho because 126.99>3.84 +(339-357.6)2/357.6
• For age and CVD =5.44+0.64+8.16+0.97=15.21
• Set hypothesize
H0: Age and incidence of CVD are
independent • Conclusion
H1: H0 is false
α=0.05 • We reject HO because
• Select the appropriate test statistics 15.21>3.84
Χ2=Σ(O-E)2/E
• Age is a confounder because
• Set up the decision rule
df =(2-1)(2-1)=1 for 5% critical value is
it is related to both obesity
3.84, so Reject H0 if Χ2 ≥3.84 and CVD
Effect Modification
• A clinical trial is conducted to evaluate efficacy of a
new drug to increase HDL. 100 participants are enrolled,
randomized to receive either new drug or placebo.
• Background characteristics(age, sex, educational level,
income) and clinical characteristics( height, weight,
blood pressure, total and HDL level) are measured on
baseline. They are equal at baseline
• These are taken again after 8 weeks
HDL by treatment Number % of males in each treatment
N Mean SD group
HDL N (no) %
Drug 50 40.16 4.46 Drug 50 10 20
Placebo 50 39.21 3.91 Placebo 50 9 18
Effect Modification
HDL by treated and sex • Mean HDL is 0.95mg/dl
Female higher in new drug
N Mean SD • 2 sample test z=1.13, not
HDL
significant at 0.05 for HDL
40 38.88 3.97
level
41 39.24 4.21
Male
• χ2=0.06 not significant
hence sex is not confounder.
N Mean SD
HDL • HDL is 0.36mg/dl lower in
10 45.25 1.89 females , and on average it is
9 39.06 2.22 6.19 higher in female on drug
, hence in males the drug has
higher effect than males.
The Cochran-Mantel-Haenszel
method
• When data of age was pooled, effect of obesity was magnified
• While in statistical interaction the effect may not be
• Cochran-Mantel-Haenszel method is a technique that generate an
estimate of an association between an exposure and outcome
accounting for confounding
• The method is used with a dichotomous outcome variable and
dichotomous risk factor and essentially computes a weighted
average of a relative risk or odds ratio across stratum defined by
the confounding variable.
• To implement the CMH method, the confounder must be
categorized so that a series of 2x2 tables showing association
between exposure and outcome in each stratum
Example
• RR=p1/p2 OR=[p1/(1- • CMH
p1)]/[p2/1-p20]
Outcom Outcom Total
• RR=[Σai(ci+di)/ni]/ [Σci(ai+bi)/ni]
e e absent • OR=[Σaidi/ni]/ [Σbici/ni]
present • For example of obesity and CVD
Exposed a b a+b
• RR=[Σai(ci+di)/ni]/ [Σci(ai+bi)/ni]
Unexpose c d c+d
d • =[10(35+365)/
600+36(25+175)/400]/
a+c b+d n
[35(10+90)/600+25(36+164)/
400 =1.44
• OR=[Σaidi/ni]/ [Σbici/ni]
• RR=[a/(a+b)]/[c/(c+d)]
• =[10(465)/600+36(175)/400]/
• OR=(a/b)/(c/d) =ad/bc [90(35)/600+164(25)/400]=1.52
CMH
RR OR • Adjusted is not equal to
Crude , unadjusted 1.78 1.93 unadjusted
Age<50 1.43 1.48
Age >50 1.44 1.52
• The adjusted produces
Adjusted for age 1.44 1.52 estimates of RR and OR
that are much closer to
the stratum-specific
estimates
Regression and correlational
analysis
• Regression is a technique to assess the relationship between an
outcome variable and one or more risk factors or confounding
variables.
• Outcome variable is also called dependent while exposure or risk
factors are called independent variables.
• Predictor or explanatory may be misleading
• Correlational analysis is used to quantify the association between two
continuous variables –independent vs dependent variables
• In correlational analysis we estimate a sample corelation coeffeicient
– Pearson Product moment
• It ranges between -1 to +1 and quantifies direction and strength of
thelinear association between the two varaibles
Correlational analysis
• r=cov(x,y)/√sx2Sy2
• Cov(x,y) is the covariance of x and y defined as
• Cov(x,y)=[Σ(x-X)(y-Y)]/n-1
• And sx2Sy2 are sample variance of x and y
• sx2=Σ(x-X)2/n-1 , Sy2 =[Σ(y-Y))2/n-1
• Meaning correlation coeffeiecient could be from 0.4
positive or negative.
• There are test to assess whether they are statistically
significant correlations
Regression
• Using example of BMI( independent) and total cholesterol (dependent), and
we could consider other predictors( independent variables-age ,sex,smoking
• More than one denoted x1,x2…xp
• Outcome( dependent) is y, while predictors( independent) are x
• Where there is a single independent variable, it is simple linear regression
analysis
• It assumes there is a linear association between two variables
• y=b0+b1x, where y is the predicted value of outcome x and b0 is the
estimated y –intercept and b1 is the estimated slope.
• Slope and intercept are estimated from the data , they minimize the
differences between observed and predicted –called residuals.
• The estimates of y-intercepts and slope minimize the sum of the squared
residuals and are called the least squared estimates
Regression
• The y-intercept is the expected value of dependent variable (y) when the
independent variable ( x) is 0.
• The slope is the expected change in the dependent variable ( y) relative to a
one-unit change in the independent variable (x).
• The least squares estimates of the y-intercept and the slope are computed as
follows
• b1=rsy/sx and b0=Y-b1x, where r is the correlation coefficient, x and Y are the
means and Sx and Sy are the standard deviations of the independent variable x
and the dependent variable y
• Regression equation can be used to estimate dependent variable as a function
of the independent variable.
• There are statistical tests that can be performed to assess whether the
estimated regression coefficients ( b0 and b1 ) provide evidence that the
respective coefficient in the population are statistically different from 0
Linear regression
• The assumption is that the distribution of the dependent
variable is normal at each independent value.
• The independent variable can be continuous or
dichotomous ( indicator variable)
• y=39.21+0.95x for BMI and HDL means is x is coded
BMI, y- intercept is the expected value of y(HDL) when x
is 0. In this example x=0 indicates membership in the
placebo group. Thus the y intercept is exactly equal to
mean HDL level in the placebo group.
• The b1, slope is 0.95 – the difference in mean HDL levels
between treatment group and the placebo.
Multiple regression analysis
• Is an extension of simple regression analysis used to assess the
association between 2 or more independent variables and single
continuous dependent variable.
• y=b0+b1X1+b2X2….bpxp
• X1,x2, xp are distinct independent variables, b0 is the expected value
of y when all the independent variables are equal to 0and b1 and bp are
estimated regression coefficient.
• Each regression coefficient represent the expected change in y relative
to one-unit change in respective independent variable holding the
remaining independent variables constant.
• There are tests to assess whether each regression coefficient is
statistically difference from 0
• Multiple regressin can handle confounder
Multiple logistic analysis
• Logistic regression is similar to linear regression
analysis except the outcome is dichotomous
• Simple logistic regression is when it is between one
dichotomous dependent variable and one independent
variable.
• Multiple logistic regression analysis applies when there
is single dichotomous outcome and more than one
independent variables.
• In logistic regression they use log of [p/1-p] the odds