0% found this document useful (0 votes)
63 views40 pages

Analyzing Dummy Dependent Variables

- The document discusses using dummy variables when the dependent variable is dichotomous, such as in analyses of war/peace, voting patterns, or yes/no legislative votes. - It presents a linear probability model (LPM) approach where the probability (p) of an outcome (e.g. war) is estimated based on independent variables. - As an example, it analyzes the probability of terrorist group emergence based on a country's electoral system, finding that more majoritarian systems are linked to a lower probability of emergence.

Uploaded by

Daniel Cano
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
63 views40 pages

Analyzing Dummy Dependent Variables

- The document discusses using dummy variables when the dependent variable is dichotomous, such as in analyses of war/peace, voting patterns, or yes/no legislative votes. - It presents a linear probability model (LPM) approach where the probability (p) of an outcome (e.g. war) is estimated based on independent variables. - As an example, it analyzes the probability of terrorist group emergence based on a country's electoral system, finding that more majoritarian systems are linked to a lower probability of emergence.

Uploaded by

Daniel Cano
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Dummy Dependent Variables and Additional Topics

Kåre Vernby

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 1 / 40
Lecture Outline

Dummy Dependent Variables


Outliers and Influential Cases
Multicollinearity

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 2 / 40
Dummy Dependent Variables

We have talked about using dummy variables as independent variables


We can also analyze dummy dependent variables
This happens in examples where we wish to analyze...
...data on whether two nations are at war with each other or not
...data on whether individuals participated in politics or not
...data on whether legislators voted yes or no to a legislative bill
...and in many other situations

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 3 / 40
Dummy Dependent Variables

Since the outcome is dichotomous (war/peace, vote/abstain, yes/no)


we are interested in estimating a probability
The probability that two countries go to war with each other
...that an individual votes
...that a legislator votes yes
...given certain values on the independent variables
We can call this probability p̂

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 4 / 40
Dummy Dependent Variables

The simplest way of doing this is the linear probability model (LPM):

p̂ = a + b1 x1 + b2 x2 + ... + bk xk

where k is the number of independent variables


This is simply your usual regression analysis and can be estimated by
the method of least squares
The only difference is that you have dichotomous/dummy dependent
variable
And you interpret estimates of the slopes (b1 , b2 ...bk ) in terms of
predicted probabilities

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 5 / 40
Dummy Dependent Variables

Emergence of terrorist groups as running example:


Less permissive electoral institutions with majoritarian
electoral formulas are considered to have limited capacity to
appease discontented societal groups. Thus, the
cross-country variation in electoral institutions should also
be particularly relevant to the emergence of terrorism in
democracies (Aksoy and Carter 2012).

Dependent variable takes on the value of 1 if terrorist group has


emerged in a country-year and 0 otherwise
Independent variable is District Magnitude (higher values=more
majoritarian electoral system)
We want to know whether the probability, p̂, that a terrorist group
emerges is related to District Magnitude

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 6 / 40
Dummy Dependent Variables

The results to the right can be


interpreted: Table: LPM of Terrorist Group
Emergence 1974-1975
For a one unit increase in
District Magnitude, the (1)
probability that a terrorist group
emerges goes down by 0.15 District Magnitude -0.15***
Or, alternatively: (0.03)
Constant 0.50***
For a one unit increase in (0.08)
District Magnitude, the
probability that a terrorist group Observations 59
emerges goes down by 15 Standard errors in parentheses
percentage points *** p<0.01, ** p<0.05, * p<0.10

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 7 / 40
Dummy Dependent Variables

Is this a strong or weak effect?


The most majoritarian systems Table: LPM of Terrorist Group
Emergence 1974-1975
have 0 on the District
Magnitude index (1)
The most proportional has 5 on
the District Magnitude index District Magnitude -0.15***
(0.03)
When going from the most
Constant 0.50***
majoritarian to the most
(0.08)
proportional electoral system the
probability that a terrorist group
Observations 59
emerges decreases by 0.75
Standard errors in parentheses
(-.15*5)
*** p<0.01, ** p<0.05, * p<0.10
Since probabilities only go from
0 to 1 this must be considered a
strong effect!
Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 8 / 40
Dummy Dependent Variables

Statistical significance testing as


usual Table: LPM of Terrorist Group
Emergence 1974-1975
For instance, the 99%
Confidence Interval around (1)
coefficient/slope:
−0.15 ± 2.58 ∗ 0.03 District Magnitude -0.15***
−0.15 ± 0.08 (0.03)
Constant 0.50***
Since this interval does not
(0.08)
include 0 we reject H0 that
β=0
Observations 59
Standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.10

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 9 / 40
Dummy Dependent Variables

To get a better grasp of our


result, we can graph the the District Magnitude and the Emergence of Terrorist Groups

1
Probability of Terrorist Group Emergence
predicted probability (p̂) against
District Magnitude
I have added a scatterplot

.5
showing the individual
observations

0
From the latter we can see with
our own eyes that...
...the emergence of terrorist −.5

0 1 2 3 4 5
groups becomes increasingly District Magnitude

uncommon as DM increases phat=0.50−0.15*x

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 10 / 40
Dummy Dependent Variables

The figure also reveals some


problems with the analysis District Magnitude and the Emergence of Terrorist Groups

1
Probability of Terrorist Group Emergence
First, we get predicted
probabilities that are negative
Negative probabilities do not

.5
exist!
Out of bounds predictions

0
(probabilities above 1 and below
0) are common in LPM models
−.5
Method of least squares just fits 0 1 2 3 4 5
the straight line that minimizes District Magnitude

phat=0.50−0.15*x
the RSS

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 11 / 40
Dummy Dependent Variables

Second, the linear model does


not fit the data that well District Magnitude and the Emergence of Terrorist Groups

1
Probability of Terrorist Group Emergence
Looking at the scatterplot we
see that...
As we move along the x-axis...

.5
..the emergence of terrorist
groups first drops dramatically

0
...and then becomes virtually
non-existent (with the exception
−.5
of one case) 0 1 2 3 4 5
District Magnitude
A non-linear model of the
phat=0.50−0.15*x
probability of terrorist group
emergence might be more
appropriate
Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 12 / 40
Dummy Dependent Variables

The most common non-linear model used with dummy dependent


variables is called the Logit

exp(a + b1 x1 + b2 x2 + ... + bk xk )
p̂ =
1 + exp(a + b1 x1 + b2 x2 + ... + bk xk )

where k is the number of independent variables


and β1 through βk are logit coefficients
Where you have just one independent variable the Logit model is

exp(a + b1 x1 )
p̂ =
1 + exp(a + b1 x1 )

E.g. relationship between probability that terrorist group emerges and


district magnitude

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 13 / 40
Dummy Dependent Variables

We need to know the a’s and


b’s to calculate the predicted Table: Logit Model of Terrorist Group
probability Emergence 1974-1976
We can use something called (1)
Maximum Likelihood to Logit
estimate he a’s and b’s
The estimates for our example District Magnitude -1.11***
are shown to the right (0.33)
We can interpret the sign of the Constant 0.18
coefficient/slope for District (0.42)
Magnitude:
Observations 59
The higher the district Standard errors in parentheses
magnitude, the lower the *** p<0.01, ** p<0.05, * p<0.10
probability that terrorist groups Put your notes here.
emerge
Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 14 / 40
Dummy Dependent Variables

But to get the predicted


probability (p̂) we have to plug Table: Logit Model of Terrorist Group
this information back into the Emergence 1974-1976
Logit formula (1)
exp(0.18 − 1.11 ∗ x1 ) Logit
p̂ =
1 + exp(0.18 − 1.11 ∗ x1 )
District Magnitude -1.11***
What is the predicted probability (0.33)
that a terrorist group emerges Constant 0.18
when District Magnitude=2? (0.42)

exp(0.18 − 1.11 ∗ 2) Observations 59


p̂ =
1 + exp(0.18 − 1.11 ∗ 2) Standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.10
≈ 0.12
Put your notes here.

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 15 / 40
Dummy Dependent Variables

We can graph p̂ for the entire


range of District Magnitude District Magnitude and the Emergence of Terrorist Groups

Probability of Terrorist Group Emergence


1
Notice how, with the Logit, p̂ is
a non-linear function of District
Magnitude

.5
When moving along the x-axis,
p̂ at first falls rather quickly...
...and then levels out

0
Seems to fit the data slightly
better (more on this later)
−.5

And does not generate 0 1 2 3 4 5


District Magnitude
out-of-bounds predictions!
Linear Probability Model
Logit

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 16 / 40
Dummy Dependent Variables

We can see from the graph that


the effect of a change in x on District Magnitude and the Emergence of Terrorist Groups

Probability of Terrorist Group Emergence


p̂...

1
...depends on where on x you
start out

.5
We therefore calculate
something called first-differences
The change in p̂ from a certain

0
change in x

−.5

0 1 2 3 4 5
District Magnitude

Linear Probability Model


Logit

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 17 / 40
Dummy Dependent Variables

Say we are interested in the effect on the probability of terrorist


emrgence of moving from a district magnitude of 0 to 1
When District Magnitude=0:

exp(0.18 − 1.11 ∗ 0)
p̂ =
1 + exp(0.18 − 1.11 ∗ 0)
≈ 0.54
When District Magnitude=1:

exp(0.18 − 1.11 ∗ 1)
p̂ =
1 + exp(0.18 − 1.11 ∗ 1)

≈ 0.28

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 18 / 40
Dummy Dependent Variables

So, when moving from a value of 0 on the index of District


Magnitude...
...to the value of 1...
The probability that a terrorist group emerges falls by 0.26 (0.54-0.28)
So, what would be the effect of moving from a value of 1 on the
index of District Magnitude...
...to the value of 2?

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 19 / 40
Dummy Dependent Variables

A different and quicker route is to compute of a one unit change in


our x at a certain p̂
You select a p̂ that is somehow theoretically interesting
And then you calculate the following:

p̂(1 − p̂)b

where b is the estimated slope of the x you are interested in


This is sometimes referred to as the marginal effect of x at p
Let’s see how this works...

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 20 / 40
Dummy Dependent Variables

What is the effect of a one unit


increase in District Magnitude Table: Logit Model of Terrorist Group
when p̂ = 0.5? Emergence 1974-1976
That is, for a country that is at (1)
the tipping point between Logit
having and not having a
terrorist group emerge District Magnitude -1.11***
(0.33)
p̂(1−p̂)b = 0.5∗(1−0.5)∗(−1.11) Constant 0.18
(0.42)
≈ −0.28
A one unit increase in district Observations 59
magnitude at p̂ = 0.5 will Standard errors in parentheses
reduce the probability of having *** p<0.01, ** p<0.05, * p<0.10
a terrorist group emerge by 0.28! Put your notes here.

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 21 / 40
Dummy Dependent Variables

The marginal effect of x is


always highest at p̂ = 0.5 Table: Logit Model of Terrorist Group
Emergence 1974-1976
But what about a scenario
where there is a much lower risk Logit
of the emergence of a terrorist
group, say at p̂ = 0.10 District Magnitude
-1.11***
(0.33)
p̂(1−p̂)b = 0.1∗(1−0.1)∗(−1.11) Constant 0.18
(0.42)
≈ −0.10
Observations 59
A one unit increase in district Standard errors in parentheses
magnitude at p̂ = 0.5 will *** p<0.01, ** p<0.05, * p<0.10
reduce probability of having a Put your notes here.
terrorist group emerge by 0.10

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 22 / 40
Dummy Dependent Variables

Statistical significance testing of


coefficients/slopes as usual Table: Logit Model of Terrorist Group
Emergence 1974-1976
For instance, the 99%
Confidence Interval around Logit
coefficient/slope:
−1.11 ± 2.58 ∗ 0.33 District Magnitude -1.11***
−1.11 ± 0.85 (0.33)
Constant 0.18
Since this interval does not
(0.42)
include 0 we reject H0 that
Observations 59
β=0
Standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.10
Put your notes here.

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 23 / 40
Dummy Dependent Variables

First differences and marginal


effects can be calculated using Table: Logit Model of Terrorist Group
the same methods for Emergence 1974-1976
multivariate logits
To the right, I have added a (1)
measure of Ethnic Logit
Fractionalization to our model
Calculate the effect on p̂ (the District Magnitude -1.06***
probability of terrorist (0.35)
emergence) of moving from a Ethnic Fractionalization 2.96
district magnitude of 0 to 1 (2.32)
when Ethnic Fractionalization is Constant -0.50
0.40 (it’s mean value) (0.68)

Observations 56
Standard errors in parentheses
Kåre Vernby (Uppsala universitet) ***
Dummy Dependent Variables and p<0.01, ** p<0.05,
Additional Topics December 17,*2013
p<0.1024 / 40
Dummy Dependent Variables

When calculating marginal


effects at p̂ one just does it the Table: Logit Model of Terrorist Group
same way as in the bivariate Emergence 1974-1976
case
What is the effect of a one unit (1)
increase in District Magnitude Logit
when, for instance, p̂ = 0.25
(the share having a terrorist District Magnitude -1.06***
group in our sample)? (0.35)
Just use the formula: Ethnic Fractionalization 2.96
(2.32)
p̂(1 − p̂)b Constant -0.50
(0.68)

Observations 56
Standard errors in parentheses
Kåre Vernby (Uppsala universitet) ***
Dummy Dependent Variables and p<0.01, ** p<0.05,
Additional Topics December 17,*2013
p<0.1025 / 40
Dummy Dependent Variables

How is one to judge the goodness of fit of regression models with


dummy dependent variables?
One common approach is to look at % correctly classified
One classifies those observations with p̂> 0.5 as 1s
One classifies those observations with p̂> 0.56 0.5 as 0s
Would a country with District Magnitude=0 be classified as a
predicted 1 or 0 for LPM? Logit?
Do this for all observation and compare with actual outcome

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 26 / 40
Dummy Dependent Variables

Table: Classification Table for Linear Probability Model (LPM)

Predicted Outcome
Actual Outcome No Terrorist Terrorist
No Terrorist 44 0
Terrorist 15 0

Table: Classification Table for Logit Model


Predicted Outcome
Actual Outcome No Terrorist Terrorist
No Terrorist 35 9
Terrorist 2 13

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 27 / 40
Dummy Dependent Variables

The proportion correctly classified for the LPM is then:


#CorrectlyClassifed 44 + 0
Correct = = ≈ 0.75
n 59
The percent correctly classified is 75%
For the Logit Model:
#CorrectlyClassifed 35 + 13
Correct = = ≈ 0.81
n 59
The percent correctly classified is 81%

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 28 / 40
Dummy Dependent Variables

Is our success rate good or bad?


Common practice is to compare to success rate had we predicted all
observations as being in our modal category
In our case the mode is 0 so what if we predicted that all observations
be 0 (as having no group emerge)?
We would be right for those 44 that did not have terrorists and wrong
for the rest:
44
Correct = ≈ 0.75
59
The percent correctly classified would be 75% which is same as LPM
but worse than Logit model

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 29 / 40
Outliers and Influential Cases

When using regression analysis we should be wary of ‘unusual’ cases


Most problematic are cases that have:
1. High leverage which mens that they have atypical values on the
independent variable(s)
2. Large residual values and thus are far away from the regression
line
Cases that both have uncommon values on x and large residual values
are called influential cases
They are called so because they tend to have a substantial impact on
our estimated slope b

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 30 / 40
Outliers and Influential Cases

The Relationship Between Unionization and Inequality, 2002


40

United States
35
Gini Index of Inequality

United Kingdom
Spain Greece
Italy

Ireland Canada Australia


30

France Germany Belgium


Switzerland

Austria Norway
25

Finland Sweden

Netherlands Denmark
20

20 40 60 80 100
Union Membership as Share of Labor Force (%)

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 31 / 40
Outliers and Influential Cases

.3
.25
Detecting Influential Cases
Leverage
.2

United States
.15

France

Netherlands
.1
.05

0 .1 .2 .3
Normalized residual squared

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 32 / 40
Outliers and Influential Cases

As is clear, the Netherlands is an influential case


To some extent, this is true of the United States and France as well
What if they are ‘driving’ our finding that increasing unionization
decreases inequality?
Standard approach is to estimate the model by ‘dummying out’ the
influential cases or by dropping them
If our substantive conclusions remain the same we say that our results
are robust to the influential cases

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 33 / 40
Outliers and Influential Cases

Table: Regression Analyses of Inequality With Dummies for Influential Cases


(1) (2) (3) (4) (5)

Union Membership (%) -0.10** -0.12*** -0.12*** -0.13*** -0.13***


(0.04) (0.04) (0.04) (0.04) (0.04)
Netherlands -9.40** -9.67***
(3.34) (3.09)
United States 3.70
(3.16)
France -5.59*
(3.15)
Intercept 33.12*** 34.66*** 34.66*** 35.04*** 35.04***
(1.87) (1.66) (1.66) (1.82) (1.82)
Observations 18 18 17 18 15
R-squared 0.26 0.52 0.45 0.66 0.49
Adj. R-squared 0.22 0.45 0.41 0.56 0.45
Root MSE 3.76 3.14 3.14 2.82 2.82
Standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.10
Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 34 / 40
Outliers and Influential Cases

We have now ‘controlled for’ the influential cases


We tried both ‘dummying out’ and ‘dropping’ influential cases which
are equivalent approaches
Apparently, our result that unionization decreases inequality is robust
to them
Our slope estimate b even increased when controlling for the
influential cases

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 35 / 40
Outliers and Influential Cases

A formal test of influence is DFBETA


What you do is subtract the estimated b including the influential case
from that estimated excluding or dummying out the case and divide it
by the standard error of the first b
The DFBETA-score of Netherlands is:
−0.10 − (−0.12)
DFBETA = = 0.5
0.04
We say that the exclusion of Netherlands moves the slope estimate
0.5 standard errors
A common ‘rule of thumb’ says that we should ‘dummy out’ or drop

observations with DFBETA > 2/ n

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 36 / 40
Multicollinearity

When the independent variables in a multivariate are too strongly


correlated
Multivariate regression may fail to uncover relationships
Even if they are really there
Intuitively it becomes hard to separate the effects of highly correlated
independent variables
Only solution when there is high multicollinearity among independent
variables is to get more data

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 37 / 40
Multicollinearity

To show how this works, I have created a data 2000 observations in


which the estimated regression line is:

yi = 1.52 + 1.14x1 + 1.32x2

However x1 and x2 are highly correlated (Pearson’s r = 0.95)


If we are drawing random samples from these 2000 observations it is
going to be hard to find significant effects
Especially if samples are small

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 38 / 40
Multicollinearity

Table: Illustrating multicollinearity using simulated data


(1) (2)

x1 1.14*** 1.31
(0.28) (0.91)
x2 1.32*** 1.22*
(0.07) (0.70)
Intercept 1.52*** 1.33***
(0.36) (0.45)

Observations 2,000 60
R-squared 0.73 0.81
Adj. R-squared 0.73 0.80
Root MSE 1.95 1.87
Standard errors in parentheses
*** p<0.01, ** p<0.05, * p<0.10

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 39 / 40
Multicollinearity

Remember three things


1. Multicollinearity’s main consequence is to make it harder to find
significant results (we conclude that there is no effect when there in
fact is)
2. It can be solved by dropping one of the offending variables but we
then fail to do justice to our theories and hypothesis
3. So the only real solution is to collect more data

Kåre Vernby (Uppsala universitet) Dummy Dependent Variables and Additional Topics December 17, 2013 40 / 40

You might also like