Econometric Analysis of Microeconomic Data: Analysis of Duration
data
Example: Suppose that we follow 100 individuals freshly laid-o¤ (we want to
measure time until they …nd a new job)
dur. # of ind. Pr(T=t) # ind. Survivor Hazard
leaving (density) still unemp. P r(T > t) Pr(T = t j T > t 1)
1 10 10/100 90 90/100 10/100
2 10 10/100 80 80/100 10/90
3 10 10/100 70 70/100 10/80
...... ... ... ... ... ...
8 10 10/100 20 20/100 10/30
9 10 10/100 10 10/100 10/20
10 10 10/100 - 0/100 10/10
not stationary with T !
You may verify that
if gave you the hazard rate (or survivor), you could have recovered the
density and the CDF .
Reverse is true (obviously, as long as you have the density or the CDF, you
may get the hazard or survivor)
Example: Suppose we come up with a model of Pr(T = t j T > t 1) =
h(t), then we could explain time needed to …nd a job as follows. Suppose
someone needed t + 1 periods to …nd a job (t periods of unemployment):
Pr(T = t) = (1 h(1)) (1 h(2)):::::(1 h(t)) h(t + 1)
1- Continuous time Models
De…ne T as a random variable denoting the time elapsed until an event
occurres, or Waiting Time,
– T is positive valued
– Usually, we think of T as the time observed to leave one state for
another one
Usually, T is characterized by a CDF, denoted F (:), or a density f (:):For
Duration data analysis, we use
– Hazard rate:
t T <t+ tjT >t
h(t) = lim Pr( )
t!0 t note that this a conditional probability!
f (t) log S (t)
= =
S (t) dt
where S (t) = PrfT > tg = 1 F (t) is the survivor
note it is a
This suggests that
Z t
log S (t) = h(u)du ! S (t) = exp( (t))
0
where (t) is the integrated hazard and is de…ned as:
Z t
(t) = h(u)du
0
Obviously, we can choose to de…ne T in discrete time (example above was
in weeks)
In words, h(u) is the hazard function (instantaneous transition probability),
(t) is the integrated hazard.
The Exponential distribution is one of the most famous distribution used
in hazard models;
f (t) = exp( t) for t > 0 and > 0
f (t) exp( t)
h(t) = = =
S (t) exp( t)
h(t)
! = 0 (memoryless property)
t
Censoring (type 1)
Suppose that observations are censored after a duration of Tc: In other words,
individual durations that are superior to Tc; we have partial information.
De…ne a latent variable Ti ; which plays the role of the actual (true) du-
ration.
Ti = minfTi ; tcg
So, de…ne a censoring indicator,
ci = 1 if Ti > tc
ci = 0 if Ti tc
To form the likelihood, two cases must be considered:
1. Ti = Ti tc
Pr(Ti = ti) = f (ti)
2. Ti = tc < Ti
Pr(Ti > tc) = S (tc)
N
Y
L( ) = f (ti)1 ci S (tc)ci
i=1
Proportional Hazards model
The Proportional hazards model assumes:
proportioanluty part, proportional to the lambda naught,
h(t; X; ) =
THE EXPONENTIAL TAKEN BECAUSE IT HAS THE MEMORYLESS PROPORTY,
THE EXP USTHE ONLY ONE WITH CONSTANT HAARD RATE
(X 0 ) 0 (t; )
= exp(X 0 ) 0 (t; ) (in most cases)
where 0(t) is refered to as the baseline hazard
Most famous example: Weibull Model
h(t; X; ) = exp(X 0 ) t 1 for >0
which is monotonically increasing (decreasing) for > (<) 1.
Easy to write the likelihood:
– i = 1 (no censoring)
– i = 0 (right censoring)
– an observation i contribures f (ti j Xi; ) i S (ti j Xi; )(1 i)
– We can write the likelihood as:
N
X
ln L( ) = [ i ln f (ti j Xi; ) + (1 i ) ln S (ti j Xi ; )
i=1
ln f (ti j Xi; ) = ln[exp(X 0 ) t 1 exp( exp(X 0 )t )]
= X 0 + ln( ) + ( 1) ln(t) exp(X 0 )t
ln S (ti j Xi; ) = ln(exp( exp(X 0 )t )]
= exp(X 0 )t
Note that because ln f (:) = ln h(:) + ln S (:); the log likelihood function
can be expressed generally as follows:
–
N
X
ln L( ) = f i ln h(ti j Xi; ) + (ti j Xi; )g
i=1
Mixture Models and Unobserved Heterogeneity
Objective: Illustrate why unobserved heterogeneity (dynamic self-selection)
is problematic. Show that ignoring unobserved heterogeneity creates bias
toward negative duration dependence
Example: Two types of unemployed with both constant hazard:
– 100 individuals with type F (fast) have hazard rate 0.4
– 100 individuals with type S (slow) with hazard rate 0.1
type F type S Aggregate
Duration freq. hazard freq. hazard freq. hazard
1 40 0.4 10 0.1 50 0.25
SO THE COMPOSITION IS MORE AND MOREBIASED TOWARD PEOPLE OF TYPE S,
2 24 0.4 9 0.1 33 0.22
3 14 0.4 8 0.1 22 0.19
Solution: Incorporate unobserved heterogeneity within the PH model: MPH
h(t j x; ; v ) = 0(t) exp(x0 ) v
= exp(X 0 + log v ) 0(t) where v > 0
E (v ) = 1; var(v ) = 2v and v orthogonal to x
WHAT IS V, THEN IT IS THE SAME AS THETA IN DISCRETE PANEL DATA EFFECT.
EXPECT CONSTANT HAZARD RATE, BUT ACTUALLY THE HAZARD RATE
In Biostatistics v is interpreted as "Frailty"
Orthogonality between x and v is not inconsistent with “Dynamic Self-
Selection”:
Conditional on surviving beyond a speci…c t, x and v are no longer orthog-
onal
To see this: Suppose that x increases mortality and v is unobserved frailty
– At period 0 (birth), x and v are uncorrelated
– Take a sub-population of very old individuals (surving beyond t where
t is large), an individual with large x must be an individual with very
low v
– Hence, as we move with t, orthogonality is lost and there is an increas-
ing negative correlation between x and v
Example with the exponential model:
h(t; xi; ; vi) = exp(x0i ):vi
= exp(x0i ): exp("i)
exp( 0 + x01i 1 + "i)
IF YOU FORGET THIS UNOBSERVED TERM, YOU CREATE BIAS IN THE SHAPE OF LZMBDA NAUGHT T
If we observed v; there would be no problem. We would maximize like-
lihood conditional on v (like in the Panel Binary choice model (Random
E¤ect Probit)
Result 1: The mixture hazard function (aggregate hazard function) is the
expected hazard rate at t, where expectation is over the distribution of
unobserved component at time t given the survival to t.
R
f (t j )g ( )d
h(t) = R
S (t j )g ( )d
R
h(t j )S (t j )g ( )d
= R
S (t j )g ( )d
Z
= h(t j ) (v )dv
with
S (t j )g ( )
(v )dv = R
S (t j )g ( )d
Pr(T > t; v )
=
Pr(T > t)
= g (v j T > t)
So, g (v j T > t) is the probability density of unobserved component v
conditional on survival until t.
Result 2: Neglecting heterogeneity results in an estimated hazard falling
faster (or rising more slowly) than the true hazard rate (Bias toward Neg-
ative Duration Dependence)
@E( jT t)
Example with the Weibull and examine what happens to @t to
see what happens to the hazard functions disclosed in the data (ignoring
unobserved heterogeneity)
Z
S (t) = exp( t )g ( )d
Z
@ ln S (t j )
h(t) = g ( )d
@t
= t 1 E( j T t)
It turns out that
@E ( j T t) 1 fE ( 2 j T
= t t) E( j T t)2g
@t
= t 1 V ar ( jT t)
< 0
This is mean what we mean by Dynamic Self-Selection: as we move have
few and fewer survivors, they become less and less representative.
Identi…cation and Estimation with Unobserved Heterogeneity
We have speci…ed hazard models so to include observed heterogeneity (Xi)
as well as unobserved heterogeneity (vi); but we did not say how to proceed
with vi:
Let’s …rst discuss identi…cation:
S (t; x) provided by the data
S (t; x) = Ev (S (t; x; v )
Z
= exp( v 0 (t) (x))g (v )dv
which means that S (t; x) should convey information about 0 (:); (x)
and g (:)
Non-Parametric Identi…cation Theorems are based on following as-
sumptions (mild):
– independence between v and x
– E (v ) < 1
– (x) > 08x
– 0(:) is continuous and positive over 0; 1)
basic principle guiding estimation: be as ‡exible as possible!
There are two fundamental approaches:
– One approach is to choose a continuous distribution for g ( ) so to get
closed-form expressions (in the end, replace v by var(v ))
– Another approach (more modern) is to try to estimate g (v ) almost
non-parametrically (using mixtures)
Finite Mixture Models
Suppose two populations: One drawn from f1(t j 1 (x)) and f2(t j
2 (x))
The two-components …nite mixture is
f1(t j 1(x)) + (1 ) f2(t j 1(x)) with 0 < <1
If is unknown, we would estimate ; 1 and 2 (each sub-population
de…nes a “type”)
Can easily be extended to more than 2 components; say m components
(for instance, because if m is large enough than it may provide a good
approximation to g (v )
The hazard function analyzed in Heckman and Singer (1984) is
M
X
f (ti j xi; ) = f (ti j xi; vj ; ) j (vj )
j=1
Example: Take the Weibull model and put heterogeneity in the intercept
term (exp( type + X 0 )
f (ti j Xi; type(1)) =
exp( 1 + X 0 ) t 1 exp( exp( 1 + X 0 )t ) (type 1)
f (ti j Xi; type(2)) =
exp( 2 + X 0 ) t 1 exp( exp( 2 + X 0 )t ) (type 2)
The contribution to the likelihood for individual i is
ln f (ti j Xi) = ln[f (ti j Xi; type(1)) (type1)+f (ti j Xi; type(2)) (type2)]
where (type2) = 1 (type1)
N
X
ln L( ) = ln f (ti j Xi; )
i=1
exp(~ )
1 g if no intercept in X 0 :
You estimate f ; 1; 2; (type1) = 1+exp(~
1)
M is …xed (estimation is parametric)
If M is estimated (Heckman and Singer, 1984) then approach is semi-
parametric (but di¢ cult)
Strangely enough empirical applications show that M is usually 2 or 3:
Discrete Time Models
There two ways to analyze discrete duration data
Discrete data arise when failure times (continuous) are observed (recoded)
at aggregated time intervals (Grouped duration data)
– Han and Hausman (1990) is a good example: needs to evaluate the
probability (discrete) that a failure is realized within a speci…c interval
Another approach is link duration data with discrete-time binary choices
– This is especially relevant when choices are naturally set in discrete
time
Suppose you observed A possible discrete outcomes (durations): 1; 2; :::A;
and de…ne t1; t2; ::tA accordingly
The hazard expression is
Pr[ta 1 T < ta j T 0
ta 1 j X ] = F ( a + X(t )
a 1)
The likelihood is
h i
N ai 1 0 0
L( ; 1; 2; :: A) = i=1 s=1 (1 F( s + Xi(ts 1 ) ) F ( ai+X(t )
ai 1 )
Example: Cameron and Heckman (1998) model grade transition level
in schooling
– De…ne Ds = 1 if grade s has been completed.
– The authors model the transition from s 1 to s;
Pr(Ds = 1 j Xs = xs; Ds 1 = 1) = Ps 1;s(xs)
exp(xs s + i)
Ps 1;s(xs) =
1 + exp(xs s + i)
which should be viewed as 1 hazard(:), since the transition probability
(probability of graduating from s 1 to s is a continuation probability.
Cameron and Heckman analyze dynamic self-election within the schooling
model