Classification
Classification
Lawrence M. Healy
COT 711
Introduction
and predict outcomes. One type of regression analysis is known as logistic regression. Logistic
regression is appropriate when the predicted outcome is binary (on/off, pass/fail, infected/not
dichotomous dependent data and the assumptions of ordinary sum of squares regression
methods. The independent variables that are used for outcome prediction may be dichotomous,
related studies. It can be used for any application where binary outcomes can be predicted.
Logistic regression is based on the logit transformation of the dependent variable. The logit
regression model can be developed. The outcome probabilities for each dependent variable
value are the basis for the model. The logit transformation is necessary since dichotomous
dependent data violates ordinary least squares assumptions. Another issue with dichotomous
data is that the error terms are not normally distributed, thus ordinary sum of squares regression
Logistic regression is less restrictive than ordinary sum of squares regression. It does not
ordinary sum of squares regression are based on the observed changes in the independent data
itself. Logistic regression is based on the log of the odds of a particular event occurring with a
given set of observations. Logistic regression’s underlying principles are based on probabilities
and the nature of the log curve. The only assumptions of logistic regression are that the resulting
logit transformation is linear, the dependent variable is dichotomous and that the resultant
Logistic Regression: An Overview 3
logarithmic curve doesn’t include outliers. Discriminant analysis and logistic regression will
produce similar results with dichotomous dependent data except discriminant analysis is more
restrictive and complex. Unlike discriminant analysis, logistic regression does not restrict the
nature of the independent variable. In contrast with discriminant analysis, logistic regression
adherence to normality and the equal variance assumptions while logistic regression does not
have this requirement. Press and Wilson (1978) found in their comparison of discriminant
analysis techniques with logistic regression that the predictions formed by these methods were
consistent. They recommend that logistic regression be used whenever possible, especially when
normality assumptions are violated or when there are a large number of qualitative variables.
Logistic regression techniques are appropriate when dummy variables are required during
analysis. Logistic regression is preferred by many researchers in the analytical fields due to its
robust practical nature, intuitive assumptions and its ability to produce a predictive
curves flexibility. Figure 1(Pampel, 2000) illustrates the contrast between predictive values using
the transformation to continuous data versus a plot for the original non-transformed dichotomous
values. Without the logit transformation, dichotomous value prediction is not appropriate for
ordinary multiple regression techniques. The similarities of logistic regression with other
common multiple regression tools provides the researcher with a relatively easy to use, practical
and less restrictive tool to analyze dichotomous data. Table 4 (Babin, [Link], 2006) depicts a
summary of the comparison between multiple regression and logistic regression techniques.
Logistic Regression: An Overview 4
This paper will present an overview of the logistic regression methodology and illustrate its
As with any research the study objectives must first be defined. The researcher should
then establish the best research design to address the objective. After stating all assumptions, the
researcher should estimate the logistic regression model using the logit transformation and assess
overall model fit and predictive accuracy. The results should then be interpreted and validated.
Once the logistic regression model is estimated there are 3 primary considerations for diagnostic
analysis: 1) How well does the model predict outcome? Does a relationship exist between the
independent variables as a group and dependent variable such that the independent variables
within a given level of confidence actually predict the outcome and that outcome is not random
chance? 2) If the model works well, what is the relative predictive strength of each independent
variable? 3) Are the assumptions of the model completely satisfied? The intent of this paper is to
derivations of logistic regression methods are outside the scope of this paper.
The first task in model estimation is to transform the independent variable and determine
the coefficients of the independent variables. The basic logistic regression analysis begins with
estimation. This is done using the odds ratio. The odds ratio for an event is represented as the
⎡ p ⎤ b 0 + b1 X 1+.......bnXn
Oddsi = ⎢ i ⎥ = e (1 )
⎣1 − pi ⎦
where
It represents all event probabilities, relationships and their exponential nature. The odds ratio has
an event outcome occurrence. If the odds ratio is less than one there is a decreased likelihood of
an event occurring and if the odds ratio is greater than one then there will be an increased
likelihood of the event occurring. The odds ratio provides an intuitive foundation for any
sensitivity analysis of interest between the dependent and independent variable. The odds ratio is
based on the probabilities that a specific binary outcome will occur when using particular model
estimation. It is converted to a continuous function through the logit transformation. The new
plot of the transformation of the independent data into probabilities versus the dichotomous
dependent data will be continuous ranging from infinity to negative infinity. The log of the odds
⎡ p ⎤
logiti = ln ⎢ i ⎥ (2 )
⎣1 − pi ⎦
The maximum likelihood estimation (MLE) is now used to estimate the coefficients
( β 0, β1, ....β p , ) from the logit transformation. MLE is similar to the ordinary least squares used in
multiple regression analysis. The likelihood is the probability that the observed values of the
Logistic Regression: An Overview 6
dependent variable will be predicted by the observed independent variable data. The log
likelihood (LL) is the log of that likelihood and is in the range of infinity to negative infinity.
The logistic curve simplifies the coefficient estimation. The maximum likelihood estimate seeks
to maximize the LL value and estimate the coefficient found at that maximum point. It is
determined through an iterative process that is normally handled by computer software such as
SAS or Minitab. One point worth mentioning is that MLE is extremely accurate for large sample
sizes. Since the LL is the log probability that the dependent variables will be predicted by the
observed independent variables, we should seek to maximize that probability. The coefficient
estimate where the log likelihood is maximized will represent the best probability that the
observed dependent variable is predicted by the observed independent variables. At this juncture
SAS or some other statistical package ahs computed the log likelihood and logit transformations
to estimate the coefficients for the initial model. This paper will now address logistic regression
diagnostics to validate the proposed model derived through the logit transformation.
The first concern is the consistency of the model. How well does the model predict the
outcome? The researcher should ascertain the predictive error occurrence and the model's
sensitivity to that prediction error. The predictive accuracy of the model must be determined.
This is accomplished through goodness-of-fit measures such as log-likelihood and the coefficient
of determination (R2). This section will address these two measures and then briefly discuss
Due to the non linear nature of dichotomous dependent data ordinary sum of squares
methods will not be appropriate. The log likelihood method will be used. Manipulation of the
log likelihood value leads to a test statistic known as the -2LL value. The -2LL value has an
approximate chi-square distribution and therefore can be used to evaluate the significance of the
logistic regression model. This is similar to the sum of squares error analysis used in multiple
regression. The -2LL statistic is known as the likelihood ratio (goodness-of-fit). A perfect fit
will yield a -2LL equal to zero (minimum value). As the -2LL value decreases the model is
interpreted as having a better fit and predictive estimation. The log likelihood test is an
alternative to another test statistic that is often used in logistic regression, the Wald test statistic.
The likelihood ratio should now be used to compare a reduced model against the proposed
logistic regression model. The log likelihood test can be used to test the overall model goodness-
The test of the overall model consists of comparing the -2LL statistic of a baseline null
model (no independent variables, just a constant) against the full model with all proposed
independent variables. The null hypothesis implies that the researcher should accept the baseline
model without any beta coefficients (logit (p) = constant). The alternative hypothesis is that the
full model is significant. The baseline model -2LL statistic is compared with the full model (all
independent variables included). The difference between the baseline -2LL value and the full
model -2LL value is known as the model chi-square test statistic. If the (model chi-square test
statistic) is ≤ .05 then we should reject the null hypothesis. If we reject the null hypothesis that
knowing the independent variables makes no statistical difference in the prediction of the
Logistic Regression: An Overview 8
dependent variable, then the overall model with all independent variables at this point is
statistical significant.
The next step is to evaluate the significance of each individual variable. If removing a
variable form the full model yields no statistically significant change in the -2LL test statistic
then the variable is not significant and can be omitted from the full model. This process is
repeated for all independent variables to determine whether any independent variables can be
2
Goodness-of-Fit for the Estimate – R L , Coefficient of Determination
approach. This method fits the estimation in a similar fashion that the coefficient of
−2 LLnull − ( −2 LLmod el )
R2LOGIT = ; 0 >= R2LOGIT =< 1 (3 )
−2 LLnull
It implies the degree of relative negative impact the independent variables has on the model
versus not having any independent variables at all. It reflects how much the badness-of-fit is
reduced as well as a proportionate reduction in the absolute value of log likelihood. The model
fit improves as R2LOGIT increases from 0 to 1, and is a perfect fit at R2LOGIT = 1. After
determining the significance of the overall model as well as the individual independent
Predictive Accuracy
In logistic regression, the predictive accuracy of a model is an issue that has received
considerable discussion in the literature. It appears that many researchers are confident with
The literature indicates that there are varied ways to estimate predictive accuracy as well as
many logistic regression statistic packages that will analyze the classification and predict the
accuracy of the model. Therefore this paper will only briefly cover the basics of classification
and predictive accuracy. This paper will now present one way to determine predictive accuracy
through classification table assignment. In its most basic form, the classification tables can be
computed by observing the predicted values for each set of independent variables and then
comparing those values with the average predictive values for all observations. If the individual
fitted logit predictive value is less than the average of the predicted values, then it is classified in
the 0 group, else it is classified in the 1 group. To highlights errors in the model, the actual Y
values for each data point are compared to the assigned group (i.e. 0 or 1). For instance, if a data
point was assigned to group 1 and the observed Y for that set of independent variables was 0,
this would indicate an error in the model. This comparison is performed for all data point sets
The acceptable accuracy level is a practical consideration and dependent upon the sensitivity of
the data. Table 1 (Ali, 2000) illustrates a simplistic example of a predictive accuracy
computation. As shown in the example, the model correctly predicted rows 1, 2, 3 and 66
Logistic Regression: An Overview 10
because the assigned group was equal to the observed Y. The model incorrectly predicted in
rows 9 and 52. The average predicted accuracy for this example was 0.50.
The next question to address is: what is the relative predictive strength of each
independent variable? Assuming the overall model is significant then the next step is to
significance as related to the dependent variable we can observe the un-standardized regression
coefficients and ascertain the strength of any causal relationships between the dependent and
regression coefficient must be determined for benchmark purposes. Another way to approach
coefficient. This represents the number of standard deviations a dependent variable changes as a
To determine the significance of the independent variables we can use either the Wald
statistic or the likelihood ratio test. The likelihood ratio test requires extensive repetitive
computations but these can be handled by most statistical software packages. The Wald statistic
is a method to test whether the coefficients are significantly different from 0. A more common
process to test the significance of the coefficients is to evaluate the exponentiated logistic
the logarithmic nature of the logistic coefficient, it can be difficult to interpret, but this difficulty
can be overcome through many statistical software packages. There are two approaches to
estimating the logistic coefficient significance. The original logistic coefficient can be used to
Logistic Regression: An Overview 11
interpret changes in the logit function caused by the coefficient analysis or the exponentiated
logistic expression can be used as a means to interpret changes in the odds. The direction of the
exponentiated coefficient is greater than 1 then there is a positive relationship and if the value it
is less than one, then there is a negative relationship. For example, if the original coefficient is
positive, the transformed exponentiated logistic expression will be greater than one, this means
that the odds will increase for any positive change in the independent variable. This results in a
higher probability of occurrence. If the coefficient is zero then the exponentiated logistic
expression will be equal to one and there will be no change in odds for a change in the
independent variable.
The standardized coefficient will assist in the determination the relative strength of each
change in a dependent variable associated with a 1 standard deviation change in the independent
variable. Menard (1995) depicts a step by step process to compute the standardized coefficients
3. Use the predicted value of Y to calculate the predicted value of logit(Y), using the
∧ ∧ ∧
equation: logit( Y ) = ln[ Y /(1 - Y )]
∧
4. s ∧
log it (Y )
: Calculate descriptive statistics for logit ( Y ), including the standard
deviation.
= (b YX )( s x ) /
* 2 2
b YX s log it (Y )
∧ /R
The final value computed in number 6 above can be interpreted as: for an increase of 1
*
standard deviation in independent variable x there will be a b YX
standard deviation increase
*
(+)/decrease (-) in the dependent variable Y. So for example if b YX
= .591 then for every 1
*
Menard (1995) suggests the following relative strengths for b YX
values:
Weak: 0 to .3
Moderate: .3 to .6
Strong, .7 to 1
Menard (1995) states the standardized coefficient will produce a more accurate picture than the
regression coefficients are more reliable for categorical variables. In essence, the standardized
coefficient should be viewed as a relative ranking, not taken as an absolute value. If a variable
* *
has a b YX
= .7 and another has a b YX
=.1 then we know that one variable has a much stronger
significance relative to the other, a smaller range between the two values might not as
Logistic Regression: An Overview 13
conclusive. The standardized coefficient provides a measure of magnitude of the effects of the
independent variables. The proceeding sections have examined the significance, strength and
predictive accuracy of the proposed logistic regression model and its variables. The final
question to consider is: Are the assumptions of the model completely satisfied?
The final logistic regression diagnostics phase is to determine whether the model exhibits
any significant violations of logistic regression assumptions. This includes issues related to
biased coefficients, inefficient estimates and invalid statistical inferences. Specification error
can lead to biased logistic regression coefficients; the model may be utilizing coefficients that are
immaterial variables during the analysis. Immaterial variables can result in an increase in the
model’s standard error of the parameter estimates. Omitting material variables from the model
will lead to false importance and relative strength of the remaining variables in the model.
Another specification error that can lead to an erroneous model is observed when the model’s
logit(Y) function exhibits non-linearity. This occurs when the change of logit(Y) for a one unit
change in the independent variable is not constant; it is strongly associated to the value of the
dependent variable. The logit function requires a linear relationship between the dependent and
independent variables. Co linearity occurs when the independent variables exhibit correlation
amongst each other. This interaction effect is another potential issue in logistic regression
analysis. If the change in the dependent variable associated with a one unit change in the
independent variable is related to the value of another independent variable then this behavior is
non-additive and it violates the assumptions of the model. The regression coefficient estimates
and their standard error values will be affected. While Co linearity may indicate a biased
coefficient, the degree of correlation between the variables will be the primary consideration
Logistic Regression: An Overview 14
with regards to its impact. Numerical problems such as zero cell count for categorical data can
also cause model assumption violations. When no data exists for a cell in a categorical
independent variable there can be issues with the model. This is not a problem for continuous
variables. The logistic regression analysis must yield residuals for the logi transformation that
have an average equal to zero, be randomly distributed around the mean without pattern, and be
normally distributed. This paper will now present an overview of a portion of a case study to
Dengue fever is an infectious disease that researchers wish to ascertain a model that
predicts the risk of being infected using relevant predictors. All SAS computational results for
this portion of the study are shown in Figures 2 and 3. A two-stage stratified random sample of
196 persons was selected from an area known to have a recent epidemic of the fever (Kleinbaum,
1998). 57 people of the sample were determined to already have the disease. The objective of
this regression analysis is to identify risk factors associated with Dengue fever and create a
regression model of risk prediction factors. The statistical software used for the study was SAS.
The study used Dengue fever status (DENGUE) as the dependent variable for the predicted
outcome with values of 1 for yes and 0 for no. Subject ID, AGE, MOSNET and SECTOR were
the independent variables. MOSNET was defined as an indicator of whether or not mosquito
netting was used by the subject (0=yes; 1=no) and SECTOR was the geographic sector the
subject resided in. They divided the overall region into 5 sectors (1-5). SECTOR was treated as
categorical variable and 4 dummy variables were created using sector 5 as the reference group.
where
where
∧ ∧
The study author computed the coefficients using the SAS statistical software ( β 0 through β 6 ).
MOSNET Significance
The study author then estimated the odds of someone contracting Dengue fever when using
∧
OR ( MOSNET =[Link] =0|age ,sec tor ) ≠ β 2
. (9)
∧
OR ( MOSNET =[Link] =0|age,sec tor ) = e0.3335 = 1.396. (10)
Using the computed standard error computed by the SAS statistical software, Kleinbaum (1998)
β
They were then able to compute the 95% confidence interval for e 2
as
Equation 11 results indicated that the upper confidence limit was 16.89 and the lower confidence
where
Wald (chi-square) statistic for testing null hypothesis for β was 0.0688,
2
At this point, the analysis indicated that the odds of contracting the fever were about 1.4 times
higher for someone who does not use mosquito netting. The Wald statistic is not statistically
significant. The wide range for the 95% confidence interval as well as the inclusion of the null
value of 1 within the range indicates they should reject the null hypothesis. It can therefore be
concluded that the subject’s usage of mosquito netting does not present a statistically significant
To further strengthen the study’s conclusion about the statistical insignificance of the
mosquito netting usage by the subjects (MOSNET, β ), the author utilized the log likelihood
2
statistic (-2LL).
H0: β 2
= 0, (14)
Ha: β 2
≠0 (15)
The -2LL for the full model was computed by SAS to be 203.706 and 203.778 without the
MOSNET variable.
The likelihood ratio for model comparison with and without MOSNET was
By using the same significance values as in equations 14 and 15, with one degree of freedom, chi
square indicates that this is not a statistically significant difference. The subject’s usage of
mosquito netting (MOSNET) does not present a statistically significant affect on the probability
AGE Significance
This study also investigated the affect of age as it related to risk of contracting Dengue
fever controlling for MOSNET and SECTOR. Since AGE is continuous they needed to compare
a larger range of age difference between 2 people so they compared 2 people with 5 yr age
difference. One year difference in age isn’t useful, but a 5 year difference will is more
descriptive.
Logistic Regression: An Overview 18
∧
OR ( AGE1− AGE 0=5|MOSNET ,sec tor ) ≠ e5(0.0243) = 1.13. (17)
β
Further computation for the 95% confidence interval for e 1
found
Equation 18 results indicated that the upper confidence limit was 1.03 and the lower confidence
where
Wald (chi-square) statistic for testing null hypothesis for β was 7.1778,
1
Since the confidence interval does not contain the null tested value of 1, a 5 years AGE
difference between 2 people is statistically significant with regards to being infected with the
fever. The odds ratio of 1.13 indicated that the significance was small. The log likelihood
analysis for age coefficient was not presented in the study. The author indicated that the results
of the log likelihood for AGE supported the conclusions about age referenced above. Table 2
reflects the study’s of risk of contracting the fever versus age significance. They found that as
age difference increases, the odds ratio increases, therefore the significance of the AGE variable
The study also considered the interaction of 2 variables, in this case MOSNET and AGE.
where
∧ ∧
The study author computed the coefficients using the SAS statistical software ( β 0 through β 7 ).
The final logit model with the new variable (MSA) considered was found to be
The author then presented the affect of MSA on the odds ratio for the MOSNET (mosquito
The new adjusted odds ratio for MOSNET considering interaction with AGE was
∧
OR ( MOSNET =[Link] =0|age,sec tor ) = exp [ β + β ( AGE ) ] (23)
2 7
∧
OR ( MOSNET =[Link] =0|age,sec tor ) = e [-.8043 + 0.0306(age)] (24)
The value of the odds ratio when considering the MSA interaction variable will be dependent
upon the AGE difference selected as depicted previously during the analysis of the age affect on
fever contraction risk. The author highlights has concluded that AGE is an effect modifier of the
In Table 3, the age modifying effect is illustrated through comparative analysis of the age
and the MSA interaction variable odds ratio. It depicts that the odds (risk) of getting the fever at
age 20 for someone who does not use mosquito netting is .83 times that of someone at age 20
who uses mosquito netting. He postulates that the likelihood of contracting the fever when
mosquito netting is not used (MOSNET=1) increases with age. For example, at age 40 a person
not using mosquito netting is 1.52 times more likely to contract the fever than someone at age 40
that does use mosquito netting. The SAS computation as stated by Kleinbaum (1998) found that
β +β ( AGE )
They computed the 95% confidence interval for e 2 7
and found to be
∧
(L +/- 1.96 S ∧L )
95% confidence interval = e (25)
where
∧ ∧
S ∧ = VAR( L) ,
L
∧ ∧ ∧
L = β + β (age) ,
2 7
∧ ∧ ∧ ∧ ∧ ∧ ∧ ∧ ∧
VAR( L) = VAR( β ) + (age) 2 VAR( β ) + 2(age) COV ( β , β ) ,
2 7 2 7
∧ ∧ ∧ ∧ ∧ ∧ ∧
VAR( β ) = 2.7004, VAR( β ) = 0.001399, COV ( β , β ) = −0.0435 .
2 7 2 7
∧
L = -0.8043 + 0.0306 + 0.0306(40) = 0.4197 (27)
Logistic Regression: An Overview 21
∧ ∧
VAR( L) = 2.7004 + (40)2 (0.001399) + 2(40)(-0.0435) = 1.4588 (28)
β +β ( AGE )
Adjusted odds ratio = e 2 7
(30)
∧
Adjusted odds = exp( L +/- 1.96 S ∧L ) = exp[0.4197 +/- (1.96)(1.2078)] (31)
Equation 31 results indicated that the upper confidence limit was 0.14 and the lower confidence
limit was 16.23. The author concluded that the confidence interval was wide and thus at age 40
the risk estimate for contracting the disease was not reliable. They performed analysis at various
selected age groups and found outcomes consistent with the results mentioned in this example
(age = 40). The data and computational analysis was not fully noted in the author’s paper.
Finally the author presented the Log likelihood ratio testing for the MSA interaction model
The study concludes with a comparison of the interaction model with the non-interaction model.
Equations 32 and 33 were used to determine the -2LL test statistic was examined to determine
whether the difference between the two models was statistically significant.
H0: β 7
= 0, (non- interaction without MSA variable) (36)
The study researcher assumed a chi-square degree of freedom of 1. The difference of 0.711 is not
significant at 95% confidence and therefore they failed to reject the null hypothesis. The non-
Conclusions
This paper has presented an overview of the logistic regression multivariate statistical
analysis methods, its primary guiding principles and a research study that utilizes a few of the
techniques. The Dengue study presented a reasonable overview of logistic regression methods
with a few useful examples. This author concludes that logistic regression is quite robust,
practical and relatively simple to use. The methodology is ideally suited to manufacturing, the
clinical sciences as well as many analytical fields that require rigid pass/fail results. Logistic
regression’s relaxation of the more rigid normality and variance assumptions of OLS were found
Logistic Regression: An Overview 23
to be advantageous. Often, typical real world data does not exhibit a normal distribution. This
leads to complex and time consuming transformations to normalize the data. Overall, logistic
regression is an excellent tool for a researcher who is attempting to utilize independent data
References
Chatterjee, S., Hadi, A. S., & Price, B. (2000). Regression analysis by example (3rd ed.). New
DeMaris, A. (1992). Logit modeling: Practical applications (Vol. 07-086). Newbury Park: Sage
Publications.
Dielman, T. (2001). Applied regression analysis for business and economics (3rd ed.). New York,
Greenhouse, J. B., Bromberg, J. A., & Fromm, D. (1995). An introduction to logistic regression
Hair, J. F., Black, W. C., Babin, B. J., Anderson, R. E., & Tatham, R. L. (2006). Multivariate
Kleinbaum, D. G., Kupper, L. L., Muller, K. E., & Nizam, A. (1998). Applied regression
analysis and other multivariable methods (3rd ed.). New York, NY: Duxbury Press.
Menard, S. (1995). Applied logistic regression analysis (Vol. 07-106). Thousand Oaks, CA: Sage
Publications.
Pampel, F. C. (2000). Logistic regression: A primer (Vol. 07-132). Thousand Oaks, CA: Sage
Publications.
Logistic Regression: An Overview 25
Press, S. J., & Wilson, S. (1978). Choosing between logistic regression and discriminant
Table 1
1 0 0.00 0
2 0 0.48 0
3 0 -0.12 0
…..
9 0 0.52 1
….
52 1 0.48 0
…..
66 1 0.59 1
Logistic Regression: An Overview 27
Table 2
Odds Ratio
Table 3
10 0.61
20 0.83
30 1.12
40 1.52
50 2.07
60 2.81
Logistic Regression: An Overview 29
Table 4
models
Figure Caption(s)
Figure 1. Dichotomous data without logit transformation (top), dichotomous data with logit
transformation (bottom).
Figure 2. Dengue fever study, SAS output analysis response profile without MSA
Figure 3. Dengue fever study, SAS output analysis response profile with MSA
Logistic Regression: An Overview 31
Logistic Regression: An Overview 32
Logistic Regression: An Overview 33