Chapter One
Chapter One
CHAPTER ONE
Regression Analysis with Qualitative Data: Binary (Dummy) variables
When you had learned econometrics I, you came across estimating & interpreting a linear
regression model in which both the dependent and independent variables have had quantitative
meaning. However in empirical studies, qualitative factors also might affect the dependent
variable so that it has to be incorporated into regression models. Moreover, the dependent
variable by itself may also be categorical (qualitative) in nature and in such a case using the
common OLS estimation is inappropriate. Hence, this chapter introduces you how to incorporate,
estimate & interpret qualitative (dummy) explanatory variables in a regression model. In addition
to this, you will also understand the nature, estimation & interpretation of regression models
which involves dummy (qualitative) dependent variables.
Qualitative factors often come in the form of binary information: a person is female or male; a
person does or does not own a personal computer; a firm offers a certain kind of employee
pension plan or it does not; a state administers capital punishment or it does not. In all of these
examples, the relevant information can be captured by defining a binary variable or a zero-one
variable. In econometrics, binary variables are most commonly known as dummy variables. In
other words, they are variables which are most commonly measured in nominal scale of
measurement. Such variables are thus essentially a device to classify data into mutually exclusive
categories such as male or female. Dummy variables can be incorporated in regression models
just as easily as quantitative variables. As a matter of fact, a regression model may contain
regressers’ that are all exclusively dummy, or qualitative, in nature. Such kinds of models are
known as ANOVA models. On the other hand, econometric models which have a continuous
dependent variable & have a combination of both qualitative and quantitative explanatory
variables are known as ANCOVA models.
In defining a dummy variable, we must decide which event is assigned the value one and which
is assigned the value zero. For example, in a study of individual wage determination, we might
define gender to be a binary variable taking on the value one for females and the value zero for
males. The name in this case indicates the event with the value one. The same information is
captured by defining male to be one if the person is male and zero if the person is female.
Dear students, how do we incorporate dummy variables into regression models? Even though it
is simple to include dummy variables in a regression model, it is not a straightforward task.
Thus, we have to follow a number of guidelines so that our work gets in line with statistical
theories and model assumptions. Some of these basic rules and terminologies regarding inclusion
and interpretation of dummy variables in a regression model are given as follows:
Suppose that we are undertaking an empirical study which tries to determine factors affecting
monthly income of employees of Bahir dar textile factory. To this end, we identified educational
level (years of schooling) and gender of employees are possible factors. From these factors,
gender is a dummy variable which has two categories (male & female). For this time being let’s
take male as a base category, and following other rules mentioned before, the regression model
which includes the dummy variable (gender) is specified as follows:
Where, is the monthly income of the employee in Ethiopian birr, EDU is the educational
level of the employee and FEMALE is a gender dummy variable which takes the value 1 if
the employee is female & takes 0 otherwise. Moreover, shows the effect of education while
measures the effect of gender dummy and it is called the differential intercept coefficient.
Further assume that, after collecting appropriate empirical data, we estimated the above model
and the estimated model is given as follow:
Now keeping everything else constant (i.e. making all variables equals to zero), the model left
with only the intercept term. Mathematically:
(̂|
It is interpreted as: other things remain constant, the average monthly income male employees
(the base category) is 4500 ETB.
On the other hand, if we keep only educational level remain constant and let the dummy variable
vary the expected average monthly income is given as
(̂|
(̂|
Now looking at the results on equation 3 and 4, there is only a 1500 ETB difference which is
exactly the same as the estimated coefficient of i.e. 1500. Therefore, it is this value that we
call the differential intercept coefficient in section 1.2 above. It measures the average monthly
income differential between males and female employees of the factory. This differential
intercept coefficient can be interpreted as: other things remain constant there is a 1500 ETB
monthly income difference between male and female employees of the factory. More
specifically, the average monthly income of male employees is 1500 ETB higher than that of the
average monthly income of females, Or the average monthly income of female employees is
1500 ETB lower than that of the average monthly income of male employees. You have already
learned the interpretation of the coefficient of education (3400) and that is why I overlooked it
here. The graphical description might give you a better clarification so that it is presented on the
succeeding page.
Male: 𝑰𝒏𝒄𝒐𝒎𝒆 𝟒𝟓𝟎𝟎
̂𝟎
𝜷
𝟒𝟓𝟎𝟎 ̂𝟎 + 𝜷
𝟏𝟓𝟎𝟎 → 𝜷 ̂𝟐
Education
Dear students, in empirical studies regression models often contain more than one dummy
explanatory variable. In such a case, apart from measuring & interpreting the individual effects
of each dummy variable, it is very important to evaluate their interaction effects. In other words,
it is worthwhile to estimate & interpret the combined effects of two dummy variables on the
outcome (dependent) variable of interest. Imagine that, in the above example, there are foreign
nationals working in Bahir Dar textile factory. Thus, by considering nationality (categorized in to
Ethiopian and non-Ethiopian) as another dummy variable; we can specify our regression models
as follows:
Where, is the monthly income of the employee in Ethiopian birr, EDU is the educational
level of the employee and FEMALE is a gender dummy variable which takes the value 1 if
the employee is female & takes 0 otherwise; ETHIO is a nationality dummy which takes the
value one if the employee is Ethiopian and it assumes 0 otherwise; is the
interaction dummy variable which is obtained by a simple multiplication of the two dummy
variables (gender and nationality). Moreover, shows the effect of education; measures the
effect of gender dummy; measures the effect of nationality dummy and measures the
combined effects of gender and nationality dummy on the monthly income of employees.
Now, suppose that having collected appropriate data we estimated equation 5 and the result is
given as follow:
When we look at the estimated model, the estimate of coefficient of gender dummy (800) doesn’t
give us any information about the effect of nationality on employees income. Similarly, the
estimate of coefficient of nationality dummy (1200) doesn’t also gives us any information about
the effect of gender dummy on employees monthly income. It means their effect on the average
monthly income of employees may not be simply additive but multiplicative. Thus, we need to
insert and estimate an interaction dummy so that we can get results about the combined effects of
the two variables on employees’ monthly income. Hence, it is the estimate of the coefficient of
the interaction dummy variable (950) which provides us the overall effect of the two dummies on
the monthly income of the employees. Regarding the interpretation of the model, similar to the
previous example the estimate of the coefficient of nationality dummy (1200) indicates that:
Keeping other things constant the average monthly income difference between Ethiopian and
non-Ethiopian employees of the factory is 1200 ETB in which non-Ethiopian employees gets on
average 1200 ETB more than Ethiopian employees. Furthermore, the coefficient of the
interaction dummy variable can be interpreted as: other things remain constant the average
monthly income difference between female Ethiopian and male non-Ethiopian employees of the
factory is 950 ETB in which the latter gets the higher amount than the former.
Reading Assignment! 1) Can we use an interaction term between a dummy and a continuous
explanatory variable in a regression model? If so, what’s its implication? More specifically, in
the above example, what does the coefficient of an interaction term that will be formed by a
gender dummy and educational level of employees indicates?
2) How can we incorporate a categorical explanatory variable, which has an ordinal scale of
measurement, in a regression model?
So far we have considered models in which the regress and is quantitative and the repressors are
quantitative or qualitative or both. But in empirical research there are occasions where the
regress and can also be qualitative or dummy. Consider, for example, the decision of a household
to participate in the capital market. The decision to participate is of the yes or no type, yes if the
household decides to participate and no otherwise. Thus, the capital market participation variable
is a dummy variable. Of course, the decision to participate in the capital market depends on
several factors, such as the expected return, education, and other conditions in the capital market.
Can we still use OLS to estimate regression models where the regress and is dummy? Yes,
mechanically, we can do so. But there are several statistical problems that one faces in such
models. In such cases limited dependent variable regression models are more appropriate than
the common OLS technique.
Limited dependent variable regression models are statistical models used to analyze and predict
outcomes that have limited range or discrete values. These models are designed to handle
dependent variables that are not normally distributed or have specific characteristics that make
traditional linear regression models inappropriate. In such models, the dependent variable is
typically binary (e.g., yes/no, success/failure) or has a limited range of values (e.g., count data,
ordinal data). They incorporate different estimation techniques and assumptions compared to
traditional linear regression models. Some of the varieties of these models are given as follows:
1) Binary choice models: These models are used when the dependent variable is binary
(i.e. it only assumes only two values), and the goal is to explain and predict the
probability of an event occurring. Examples of binary choice models include binomial
logistic regression, probit regression, and complementary log-log models.
2) Ordered choice models: These models are used when the dependent variable has an
ordinal scale of measurement. Examples of ordered choice models include ordered
logistic regression (also known as the proportional odds model) and ordered probit
regression.
3) Multinomial choice models: These models are used when the dependent variable has
multiple (more than two) categories, and the goal is to analyze and predict the probability
of each category. Examples of multinomial choice models include multinomial logistic
regression and multinomial probit regression.
At this level we are going to focus only on binary choice models and you will learn about the rest
when you begin your postgraduate studies. These are linear probability model (LPM), binomial
logistic and probit regression models.
The LPM assumes that the probability of the event occurring is a linear combination of the
independent variables, with coefficients representing the marginal effects of the independent
variables on the probability. Suppose that we are examining households’ capital market
participation (i.e. the household has only two options: to participate (success) and not to
participate (failure)). For such kind of studies, LPM can be specified as:
| |
Where:
| Shows the probability of that a household participates in capital market .
Hence, it is estimating equation 7 using ordinary least square estimation that has been known as
the linear probability model. Since OLS is based on a number of strict assumptions, LPM poses
a number of problems that we address later on. Before that it’s better to look further the nature of
the dependent variable in the above model (equation 7).
𝒀𝒊
1 * *** ** ** ******
0
* *** ** **
𝑿𝒊
We observed that one major assumption (normality of the dependent variable) is violated in
linear probability model which further causes other preconditions also to be violated. In general,
LPM is only a theoretical model and it has no any practical application in empirical studies. This
is because of the fact that it has the following major limitations:
Non-normality of the disturbance term: for the LPMs like that of the dependent
variable, the disturbance term takes only two values; that is, they follow the Bernoulli
distribution.
Hetroscedastic variances of the disturbance term: As statistical theory shows, for a
Bernoulli distribution the theoretical mean and variance are, respectively, and
where is the probability of success (i.e., something happening), showing that the
variance is a function of the mean. Hence the error variance is heteroscedastic.
Mathematically non-theoretical prediction: Since | in the linear probability
model measures the conditional probability of the event occurring given , it must
necessarily lie between 0 and 1. Although this is true a priori, there is no guarantee
that ̂ , the estimators of | will necessarily fulfill this restriction, and this is the
real problem with the OLS estimation of the LPM. There are two ways of finding out
whether the estimated ̂ , lie between 0 and 1. One is to estimate the LPM by the usual
OLS method and find out whether the estimated ̂ , lie between 0 and 1. If some are less
than 0 (that is, negative), ̂ , is assumed to be zero for those cases; if they are greater than
1, they are assumed to be 1. The second procedure is to devise an estimating technique
that will guarantee that the estimated conditional probabilities ̂ , will lie between 0 and
1. The logit and probit models discussed later will guarantee that the estimated
probabilities will indeed lie between the logical limits of and .
Questionable Value of : The conventionally computed is of limited value in the
dichotomous response models. To see why, consider the above figure 1.2. Corresponding
to a given , is either . Therefore, all the values will either lie along the axis
or along the line corresponding to 1. Therefore, generally no LPM is expected to fit such
a scatter well, and as a result, the conventionally computed is likely to be much lower
than for such models.
To this end econometrics literatures make | equal with the cumulative distribution
function (CDF) of the random variable . Since there are a number of random variables with
different CDFs, the CDFs commonly chosen to represent | are the CDF of logistic
and the normal distribution, the former giving rise to the logistic (logit) model and the latter to
the probit (or normit) model. The CDF of both the logistic and standard normal distribution
looks like the following graph:
Now taking the ratio of the probability of participating in capital market (equation 8) to not
participating in capital market (equation 9) gives us:
In logistic regression model equation 10 is known as the odds ratio. It measures the odds in favor
of an event occurring. In our previous example if the odds ratio equals to 0.4 it implies that the
odds in favor of participating in capital market is 2 to 3. Now taking natural logarithm on both
sides of equation 10 and introducing the random disturbance term gives us
( )
It presents the natural logarithm of the odds ratio as a function of linear explanatory variables
and parameters. In statistics and econometrics equation 11 is called the logistic regression model
or simply the logit model. It has the following important properties:
As goes from 0 to 1, the logit goes from That is, although the
probabilities lie between 0 and 1, the logits are not so bounded.
The rate of change of probability with respect to involves not only the parameters but
also the level of probability from which the change is measured. In other words the
probabilities are not a linear function of parameters of the model.
If an estimated parameter associated with a regress or is positive, it means that when the
value of the regress or increases, the odds that the regress and equals 1 (meaning some
event of interest happens) increases. On the other hand if it is negative, the odds that the
regress and equals 1 decrease as the value of increases.
Unlike the traditional linear regression analysis, logistic regression model has three different
estimates namely: (1) logit or logs of odds ratio estimates, (2) odds ratio estimates and (3)
marginal effects estimates.
Numerical Example: Suppose you have a research topic which states that “Determinants of
households poverty in Injibara town”. To achieve your research objectives, you collected data
from 181 sample households & then you categorize those households as “poor” & “non-poor”
based on a certain criterion. Consequently, the procedure makes your dependent variable
(poverty) assume only two values. In other words your dependent variable is binary. Then, since
you are studying to identify factors affecting poverty, if a particular household is found to be
poor it is called a “success” and vice versa. Then finally you decided to use binomial logistic
regression model. Based on this hypothetical example the estimations and interpretations of the
model are given as follows:
The logit or logs of odds ratio estimation result of the above numerical example is given on table
1.2 below. The result indicates that among four explanatory variables, only household head’s
educational level and marital status (dummy) are found to be significantly affecting households’
poverty at 1% and 10% level of significance respectively (i.e.
| | . The coefficient of education ) can be
interpreted as: As a household head’s educational level increases by one year, holding other
things constant, the logs of odds ratio decreases by and the result is statistically significant
at 1% level of significance.
Moreover, the coefficient of marital status can also be interpreted as: As a household
head gets married, keeping all other things constant, the logs of odds ratio decreases by 0.88
and the result is statistically significant at 10% level of significance.
Table 1.3: Estimates of odds ratio of the binomial logistic regression model
Logistic regression Number of obs = 181
LR chi2(4) = 120.46
Prob > chi2 = 0.0000
Log likelihood = -65.005921 Pseudo R2 = 0.4809
Analogous to the coefficient estimates, only household head’s educational level and marital
status (dummy) are found to be significant in odds ratio estimation. The coefficient of education
(0.67) can be interpreted as: As a household head’s educational level increases by one year,
holding other things constant, the odds ratio in favor of being poor decreases by and the
result is statistically significant at 1% level of significance. It can also be interpreted as: As a
household head’s educational level increases by one year, holding other things constant, the
odds ratio in favor of being poor decreases by . Similarly the coefficient of marital status
(0.41) is interpreted as: As a household head gets married, keeping all other things constant, the
odds ratio in favor of being poor decreases by 0.41 and the result is statistically significant at
10% level of significance. In other words, it implies that as a household head gets married,
keeping all other things constant, the odds ratio in favor of being poor decreases by 59%.
Where, is the marginal effect of the explanatory variable, the estimated coefficient
associated with that explanatory variable and is the probability of success.
Table 1.4: Estimates of marginal effects of the binomial logistic regression model
Marginal effects after logit
y = Pr(Poverty) (predict)
= .48256169
Table 1.4 presents marginal effects after the binomial logistic regression model. It shows that
household head’s educational level and marital status are significantly affecting probability of
being poor. The marginal effect associated with education is interpreted as: As a
household head’s educational level increases by one year, holding other things constant, the
probability of being poor decreases by and the result is statistically significant at 1% level
of significance. Furthermore, the marginal effect associated with marital status is
interpreted as: As a household head gets married, keeping all other things constant, the
probability of being poor decreases by and the result is statistically significant at 10%
level of significance.
This model is also known as the normit model. The basic principle of the probit model is
equating the probability of success of the Bernoulli variable with that of the CDF of the
standard normal distribution. Consider the following two equations: equation 13 presents the
probability of success while equation 14 shows the CDF of the standard normal distribution.
Probit regressions basic assumption is that these two equations are equal.
In our determinants of poverty example, consider a latent variable which takes the value 1 if a
household is found to be poor and takes 0 otherwise, the binomial probit regression model is
specified as follows:
∫ ( )
Like that of the logistic model, maximum likelihood estimation technique is used to estimate
probit regression model. Since it assumes a standard normal distribution for a linear predictor,
the coefficients of the probit model don’t directly represent the change in the probability of the
outcome. It has two major estimates: namely (1) coefficients and (2) marginal effects. Hence,
interpreting coefficients of the probit model is not a straightforward activity. Rather, in probit
model we always interpret marginal effect estimates.
Table 1.5 below presents estimates of coefficients of the binomial probit regression model for the
poverty problem discussed in logistic model earlier. Like the result of the logistic model, it
indicates that household head’s educational level and marital status (dummy) are significant
while the rest explanatory variables are insignificant. It also suggests that both of these
significant variables decrease likelihood of households’ probability of being poor. It doesn’t tell
us to what extent those variables reduce likelihood of poverty.
Table 1.6 below presents the estimates of marginal effects after the binomial probit model. Like
the logit model it provides us the rate of change of probability of success (i.e. being poor) as a
particular explanatory variable changes by one unit, holding other things constant. Hence, the
interpretation of the marginal effects estimates is exactly the same as that of the logistic model.
As shown in table 1.6 the marginal effect associated with education can be interpreted
as: As a household head’s educational level increases by one year, holding other things constant,
the probability of being poor decreases by and the result is statistically significant at 1%
level of significance. Furthermore, the marginal effect associated with marital status is
interpreted as: As a household head gets married, keeping all other things constant, the
probability of being poor decreases by and the result is statistically significant at 10%
level of significance. Finally, hence, we can understand that there is only a minor difference in
marginal effect estimates of the logistic and the probit model.
Table 1.6: Estimates of marginal effects of the binomial probit regression model
As mentioned earlier, linear probability model has no any practical significance since it has so
many problems. Thus, as long as the dependent variable of the study is qualitative (dichotomous)
logistic and probit regressions are applicable. In most applications these models gives us quite
similar results, the main difference being that the logistic distribution has slightly fatter tails than
the standard normal distribution, which can be seen in figure 1.4 below. That is to say, the
conditional probability approaches zero or one at a slower rate in logit than in probit.
Therefore, there is no compelling reason to choose one over the other. However, many
researchers choose the logistic model because of its comparative mathematical simplicity.
Figure 1.4: Cumulative distribution function of Logistic and standard normal distribution
with the independent variables. This latent variable is then censored at certain points, resulting in
the observed values. The model estimates the parameters of the linear regression equation while
accounting for the censoring mechanism.
It is typically estimated using maximum likelihood estimation (MLE) techniques, which find the
parameter estimates that maximize the likelihood of observing the censored data given the
model. The estimation procedure takes into account the censoring mechanism and provides
estimates of the regression coefficients, along with standard errors and significance levels. It
allows researchers to account for the limitations imposed by the censoring mechanism and obtain
more accurate estimates of the true underlying relationship.
i) Left-censored Tobit model: It is often used when the values of the dependent variable
are only observed when it falls above a certain threshold. The observed values are
censored at the lower bound, and the model estimates the relationship between the
independent variables and the latent variable that lies below the threshold.
ii) Right-censored Tobit model: In this case, the values of the dependent variable are only
observed below a certain threshold. The observed values are censored at the upper bound,
and the model estimates the relationship between the independent variables and the latent
variable that lies above the threshold.
0 Study Period 1
Example: Suppose that you are undertaking a duration analysis, which begins on November
2023 and lasts up to November 2024, regarding recent economics graduates students of INU.
Your major research objective is to identify factors affecting students’ time to get employed in
the labor market. To this end you start following up 150 students up to the end of your study
period. Further assume that, on the course of your study you identified “time to get employed” as
your dependent variable. However, values for your dependent variable might not be observed for
all of your sample students. Specifically, some students might get employed before the beginning
of the study; hence, the time to get employed for these students is unknown. In such a case, the
value of the dependent variable is only observed above a certain threshold, as represented by line
C of the above figure 1.5. It is such kind of situations we call left censoring and so that we
should use left censoring tobit model.
On the other hand, some of the students from the sample might be remain unemployed at the end
of the study period, thus the time to get employed for those students will not be known. In such a
case the value of the dependent variable is only observed below a certain threshold, as
represented by line B in the above figure 1.5. We call such kinds of things right censoring and it
will be estimated by right censored tobit model. Moreover, there may also be situations in which
the dependent variable is censored from above and below, as represented by line A in the above
figure 1.5, and hence it is both right and left censoring. Such kind of data can be estimated by
tobit model considering both right and left censoring.