0% found this document useful (0 votes)
11 views20 pages

Understanding Duration Models in Statistics

The document discusses duration models, which analyze the time between events, incorporating concepts such as censoring, cumulative distribution functions, and hazard functions. It outlines tools for modeling duration data, including continuous and discrete time methods, estimation techniques, and the application of these models to real-world scenarios, such as unemployment exit rates. Additionally, it provides examples of using statistical software for analysis and interpreting results based on various demographic factors.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views20 pages

Understanding Duration Models in Statistics

The document discusses duration models, which analyze the time between events, incorporating concepts such as censoring, cumulative distribution functions, and hazard functions. It outlines tools for modeling duration data, including continuous and discrete time methods, estimation techniques, and the application of these models to real-world scenarios, such as unemployment exit rates. Additionally, it provides examples of using statistical software for analysis and interpreting results based on various demographic factors.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

TOPIC 2.

DURATION MODELS (commands start with st)

1. INTRODUCCIÓN
The duration data provides information about the measure of time that elapses between the start and the end of a
event, for example:

- Movement between work states.


- Marital movements.
- Durations of strikes, training programs, patents…
- Duration until an investment occurs, a stock purchase, the return from emigration...
- Mortality (companies, individuals...).
- Medical treatment (diseases, transplants...).

The observation of an individual who exits the studied state at time t will be a realization of the variable T, it is
to say, we observe that T = t. We interpret this observation as the duration of the individual in the studied state.
is equal to t periods.

Duration models have three components:

The weather
b) Observed variable
c) Explanatory variables

1.1. Censorship

Censoring appears in the duration data when at the moment we make the measurement, the event has not yet occurred.
finished or has disappeared from our sight (normally the duration data is not observed completely).
Ignoring censorship has the same consequences in duration models as in regression.

Censorship from the right: We do not know the final duration, we only observe that T > t (the duration of their stay
in the studied state is greater than that observed at the time of leaving the sample.
Censorship from the left: We do not know the initial duration, and therefore we do not observe the individuals from the
start of the event.

1.2. Other concepts

Period: It is the time interval between the beginning of an event and its end, that is, between the start of a state and the passage.
from this to another state.

Transit: When one or more fundamental characteristics of the state are altered, it is said that a transit has occurred.
between states, where the event transitions from the initial state to a new state.

Calendar effect: Duration data observations may have a common starting point or may begin
at different moments in time. The influence of time in general is indeed relevant.

2. TOOLS OF DURATION MODELS


2.1. Cumulative distribution function

If the duration of a period of time is represented by the random variable T. The end of the event, the transition to
Another state is a dynamic phenomenon. As the period progresses, at every moment, there is always a risk.
of a transition occurring. We describe T by its cumulative distribution function which specifies the probability
that the random variable T is less than some value of t.

2.2. Density function


For continuous random variables, the probability density function provides an equivalent view. This is the approach.
no conditioned.

2.3. Survival function

The survival function gives the probability that the duration of the time period is at least det:

2.4. Hazard function (hazard rate)

It is the probability that the period ends, there is a transition, at time T=t+. t, conditioned upon the individual
it has survived until t. Its expression is:

That is to say, it measures the speed at which the time periods are completed after reaching a duration of t, having
Keep in mind that it has lasted until t. It is also known as output rate or failure ratio among other names.

This function best characterizes the described stochastic process:

Yes If h(t)/dt>0, the process is said to have a positive duration dependence.


If h(t)/dt<0 the process exhibits a negative duration dependency.
Time independence occurs when h(t)/dt=0t , which is a characteristic of processes without
memory. For example, this type of process occurs when it follows an exponential distribution.
3. DURATION MODELS: CONTINUOUS AND DISCRETE TIME
The time that elapses between events is measured on a scale of:

Continuous type: The event can happen at any moment in time.


Discrete type: The event occurs only at specific moments in time or the information.
available are not precise enough to consider continuous durations. Example:
unemployment durations in weeks or quarters.

Discrete-time methods have some advantages such as being easier to estimate.


(related to logit models).

4. ESTIMATION METHODS
Continuous type models

4.1. Componentes del modelo MPH(riegos proporcionales mixto)

Lancaster (1979) developed the MPH, which specifies the exit rate or hazard function as the product of three
components:

h(t) = h0(t) (X, )v

h0(t) is the basic risk;


(X, =exp(X' It is the influence of the observed het. (always in an exponential form);
v is the unobserved heterogeneity.

The basic risk h0(t) captures the effect of the time elapsed on the probability of a transit occurring, under the
assuming that the observed and unobserved heterogeneity is constant. If logarithms are applied, the relationship is established between
the exit rate and the observed and unobserved heterogeneity as follows.

Ln(h(t)) = ln(h0(t))+X' +ln(v) = (t)+ X + u

The most commonly used specifications by researchers for basic risk are parametric distributions. They are
subject to criticism because the parameters of duration models are sensitive to different specifications
parameters of the basic risk function.

Density function, survival and risk of parametric distributions:


The function of regressors (X, ) contains information about the observed heterogeneity that influences the probability
of the event occurring. These regressors are time-variant and time-invariant, and they take into account the effect
calendar in each period. They are placed in the usual way but their influence on the exit rate always obeys to
the exponential distribution:

(X(t), =exp( X).

Unobserved heterogeneity captures the effect of unobserved aspects of individuals or random effects.
specific to the unobservable individuals in the data such as skill, attitude, etc. on the probability of transit.
It is controlled by entering:

- Gamma distribution (most common and criticized).


- Non-parametric (discrete) distribution through support points (more convenient).

The non-parametric method involves specifying the distribution of heterogeneity F(v) as a step function with
J homogeneous groups. The corresponding distribution function of heterogeneity will be discrete with a number
finished known support points, whose locations I saw, i=1,...,J are also known and has a mass of
associated probability, pj, j=1,...,J, where pj=1.

The exit rate model analyzed so far is the mixed proportional hazard model.
MPH), see Lancaster (1979), Heckman and Singer (1984a,b). If unobserved heterogeneity is eliminated, we have the
proportional hazard model (proportional hazard, ph), see Cox (1972).

h(t) = 0 (t) (X, )

If we do not take into account the observed and unobserved heterogeneity, the model is non-parametric:

h(t) = 0 (t)

4.2. Likelihood function

If we have a sample of N individuals (i=1,2,…N) who have duration data T (t=0,1,….ti), which start their
duration at the same moment of time t=0, and continues until the moment of time t i, where if the event occurs
Before that moment, the duration will be complete (di1=1), otherwise censored or incomplete (di1=0). Then
the likelihood function will be:
Applying logarithms to linearize it:

The key is to know what the exit rate f(t) is, and what the survival function S(t) is. We multiply them and we would have the
likelihood function.

4.3. The AFT (Accelerated Failure Time) metric

Another way to estimate duration models could be the following:

= +

It is a log-linear regression model of survival time. It also corresponds to the model:

= exp( )∗

Where =ln ( ) follows an exponential distribution, so it is known as an extreme value distribution.


value distribution) with mean 0 and variance equal to 1.
EXAMPLE

The database corresponding to the MCVL (Continuous Sample of Labor Lives) used in the article is utilized:
Arranz, J.M. and García-Serrano (2013), “The effective measure of unemployment benefit duration: data on spells or
individuals", Applied Economics Letters, vol. 20(14), 1328-1332.

We want to analyze the unemployment exit rate to employment of the beneficiaries of the benefits:

Initial state: All unemployed (t=0).

cd "C:\Users\[Link]\Desktop\Econometrics" = I locate the directory

log using duració[Link]= I save the file with the name 'duració[Link]' in that directory

A) I analyze the transition rate from unemployment to employment of benefit recipients.

- Initial state: All unemployed (t=0).


- Estado final: Encuentran un empleo (duración completa), o continúan desempleados (censura).
- The variable to estimate will be discrete: finding a job (1) or not (0).

We organize the data before the estimation. We have 75,067 individuals of which 94.54%
they find a job (censored=1, full duration) and the remaining 5.46% continue to be unemployed at the end
the analysis period (censor=0, censored duration).

We are interested in measuring the time until they find a job (time). The maximum value is 1800 days (60 months)
in our data.

First. I declare my duration data: always before all commands for duration models. Identify
we are working with this type of data, indicating the censorship and the duration:

stset duration, failure(censu)

stset = always put in front of all commands for duration models;


durtocur = variable name that collects the time;
failure(censu) = variable fija que distingue entre duración completa (no hay censura) o incompleta (censur).

model descriptives

sts graph = to make graphs

Time in unemployment until they find a job

streg= for making regressions


Second. Non-parametric analysis of Kaplan Meier. The probability of surviving after 't', or the
probability of failure after time 't'
sts graph, hazard

Hazard = is an accumulation of exit rates

The more values one has, the more opportunities there are (that is, they have higher employment rates). It occurs
a rebound because as the benefit has a maximum of 2 years, individuals intensify their search again.

In the first phase, the benefits are for reasons of efficiency, but in the second for reasons of equity.
the benefit decreases but is maintained.

If time does not influence, the graph would be a straight line. Individuals change their behavior and characteristics.
they are different in the duration models.

sts list

If we represent this list in Excel, the survival curve should come out the same as before.

Calculate a survival table for me like this:


sts list, hazard

How can I know if the exit rate is higher for men or women?

sts graph, by (gender)

the further to the right the survival curve is, the longer it takes to exit (spends more time unemployed).

In the last part, time practically has no influence because it is the end of the unemployment benefit. If we include more
regressors (age, occupation...), will move to the right or left (because we are collecting the
characteristics of the sample.

To know if the two curves are equal, a test must be done to determine if the rates are equal.
survival (because the curves are so close together, they are almost identical):
Survival equality contrast: sts sex test, logrank("I want to know the gender variable if the exit rates
they are identical)

Since p=0 and therefore it is less than 0.05 -> reject H0

If H0h1(t) = h2(t) that is, both output rates of men (1) and women (2) are identical, as I reject it for
I would accept the alternative hypothesis H1h1(t) ≠ h2(t) that is to say they are different.

REGRESSION FUNCTION

witch sex1 nationality group2-group6, distribution(@@@@) nolog nohr

streg = we are indicating that it is a regression in a duration model (the dependent variable 'time' is not necessary
put it now because I have already defined my duration model before and when I put streg Stata recognizes it already

group2-group6 = I set a range that excludes group1

nohr = option for me to calculate the coefficients

distribution(@@@@) = I indicate the type of distribution, for example distribution(exp) to estimate the exponential.

Class 11/11/19

PARAMETRIC MODELS ESTIMATION

If we have to choose which is the best duration model, we need to make a comparison.

Criterios de Arkaike: teniendo en cuenta numero parámetros, observaciones y ajuste modelo calculo un indicador que
it tells us how good that model is compared to another (the one that fits my data best will be the best model).

The nonlinear model is negative, therefore the one with the highest negative value is the one that fits the least.
To determine which one fits best, we make a graph and based on the shape we see in the graph, we will deduce which one.
what type of model it is.

- Gomprtz; monotonous (they grow or decrease)


- Lognormal: not monotonic (they have an inflection point and change)

According to the criteria, Arkaike, Gompertz, and Weibull are the worst models. The model that fits best is the generalized gamma.
because it encompasses the rest of the distributions.

In Stata, six parametric models can be estimated: Exponential, Weibull, Gompertz, log-normal, log-logistic and the
generalized gamma. To estimate any model:

Streg + explanatory variables (sex1 nationality), distribution(xxxx) nolog nohr

When indicating the program provides the estimated coefficients and not probability ratios, do not log (avoid placing the
iterations). It is important first of all to interpret the coefficient 'p' if it is less than one (as in our
In this case, it indicates that the risk function is decreasing over time (as time passes, the probability of exit decreases).

In duration models, as the dependent variable is time, it is a nonlinear model because we are going to estimate a model.
where the likelihood function is nonlinear (it is raised to a series of coefficients), therefore, logarithms need to be applied
to linearize it. That’s why it should be remembered that the dependent variable is logarithmic.

If we do not indicate to the model that there is censorship and that it is therefore incomplete, the model will assume that the duration is
complete.

EXPONENTIAL MODEL

Duration model where time is not considered, it is a constant variable, meaning a straight line (time does not influence)
whether an individual finds a job or not).

witch sex1 nationality group2-group6 , distribution(exp) nolog nohr

We only have explanatory variables. I have to omit the first one in the groups because if I include all of them, I would have one.
collinear distribution.
AFT and PH metrics (see table, not all models use both metrics):

PH: calculate the probability of a person finding a job (exit rate).


AFT: the time that passes until you find a job (how long it takes until you find a job)
employment.

Interpretation: Being a hazard form, I see that males (sex1) have a higher probability of finding employment than
women (sex0). I interpret it this way because I use PH metric. Furthermore, since they are dichotomous variables (0 and 1), we have to
interpret it in such a way that variable 1 = B0 + B1 and variable 0 = B0 (therefore the coefficient I get in the
table is the difference that exists between B0 which is variable 0 that is the woman, and variable B1 which is the man, therefore
the man is X coefficient above or below the woman.

- Gender: men (1) have 0.1914 more probability than women (0) of finding a job.
- Nationality: foreigners (1) have -0.17826 less likelihood of finding employment than nationals (0).
- Everyone is less likely to find employment than group 1 (reference group). The magnitude
negativity is decreasing

Z: normal distribution (when N approaches infinity, we see significance). The contrasts are made with a normal in
duration models (in linear regression, the contrasts are with a Student's t or Snedecor's F).

All parameters are significant.

If we remove nohr and redo the regression, we will get the hazard ratio that should give us the same as the
previous coefficient B1 through:

Hazard Ratio (HR) = Exp(B)–1


When it is greater than 1, it has a positive effect, and when it is less than 1, it has a negative effect.

Cuando son menores que la unidad tienen menos probabilidad (va a la inversa). Es decir en nacionalidad los extranjeros
They have a 0.17 lower probability than nationals of finding employment (I calculate the inverse of 0.8367). group3
for example, we would say that it has a 0.20 less chance of getting a job than group 1 (which is the group of
reference)
state ic

In the function of the number of observations we have (75,067), compare the similarity function (-121465.6) with the constant (-
122245.4). The degrees of freedom are 8. The Akaike criterion is 242947.1.

The more explanatory variables we have, the better the goodness of fit of the model.

WEIBULL MODEL
This model encompasses the exponential model. It is always a monotonic function, either increasing or decreasing. If we estimate this
model, in addition to the explanatory variables we also collect:

stregsex1 nationality group2-group6, distribution(weibull) nolognohr

There is more information: they are the parameters that capture the basic risk, in the exponential model there isn't because it is
constante.
The results do not change (the signs) but the magnitudes do. The signs do not change because they are sensitive models.
the magnitudes change.

If p > 1, the exit rate has a positive duration dependence.


If p<1, the exit rate has negative duration dependence.

In the data, there is a negative duration difference: as time increases, individuals have less.
probability of finding a job. We had to multiply p (0.9190132) by each of the individuals because
it is a multiplicative model.

If we want to estimate the hazard ratio, we remove the nohr option:

The AIC and BIC are two goodness-of-fit criteria that we use to determine if my model is the best fit.
Therefore, we need to calculate this for all models, and once I have my models with the static applied.
I use the AIC or BIC criterion (indifferent) and the one with the lowest value fits better.
Likelihood = función de ajuste del modelo (log likelihood en la fórmula)

display(-2*(-120996.8)+(2*8)) = 242009.6

In the function of the number of obvs that we have (75,067), it compares the similarity function (-121465.6) with the constant (-
122245.4). The degrees of freedom are 8. The Akaike criterion is 242947.1.

WEIBULL MODEL (AFT)

Accelerated Failure Time (AFT) model.


NOTE, IN THE CASE OF AFT: A higher value means that one remains unemployed for a longer time. If it has a negative effect
appears standing still longer.

If we want to convert from a PH to an AFT, we take the parameter -b/p = - (0.1748843) / 0.919

We can calculate the survival function taking into account the characteristics of the sample:

stcurv, hazard saving (streg2, replace)


As time increases, the probability of finding a job decreases (considering the variables
explanatory gender, nationality...)

GAMMA MODEL

The gamma risk function is very flexible. This function includes:

Weibull model when k=1


Exponential model when k=1 and σ=1,
And the log-normal model when κ=0. This model is mainly used to evaluate and help select.
the most suitable parametric model.
Negative coefficient = less time unemployed = higher probability of finding employment.
Positive coefficient = more time unemployed = lower probability of finding a job.

0.473213 = positive = more time of stay for foreigners (1) in the state (and therefore less probability
to find work compared to the nationals 0).

Since there are no proportional exit rate metrics, I cannot calculate the proportionality ratio (not necessary.
put nohr). Just calculate a type of metric that is AFT (time that elapses until it leaves that state), not the PH.

We have estimated a gamma model, and now we perform a test to determine what model it really is, because the
gamma includes the exponential and the Weibull.

When k=1 it is a Weibull model, therefore I perform the contrast test [/kappa] = 1

According to this result, since p < 0.05 is significant, the null hypothesis is rejected (it is the model of
weibull) and the alternative is accepted (it is not a weibull model).

When k=0 it is a lognormal model, therefore I perform the hypothesis test [/kappa] = 0.
According to this result, since p < 0.05 is significant and therefore the null hypothesis is rejected (it is the model of
Weibull) and the alternative is accepted (it is not a Weibull model).

POST-ESTIMATION: PREDICT

Very useful for setting the average.

COX MODEL

In this model, we assume that there is proportional risk in all the explanatory variables for men and women.
they have the same probability of leaving.

stcox sex1 nationality group2-group6, nohr nolog

The null hypothesis is rejected in the case of sex and group 4.


Class 25/11/19

When we are estimating, we are maximizing a maximum likelihood function. In duration models, there are
censored durations, which is why the density function appears alongside the survival function.

Stata applies logarithms to that function to linearize it. We can assume that the density function and function
survival depending on the distribution a parameterization is made or another:

The metric, when I estimate my coefficients and choose one metric over another, I estimate the probability that it occurs.
event (PH) or the time that elapses until the event occurs (AFT). This is how the
coefficients.

The AFT (Accelerated Failure Time) metric

The dependent variable is not the exit rate, but:

lnTi = Xiγ + e

(the logarithm of time = regressors + error)

We interpret the parameters/coefficients as the positive or negative effect on duration.

If we enter the command 'nohr', it calculates the coefficients (I need you to report them to me in the table), and if we enter 'nohr'
time calculates the coefficients in accelerated output model (we can only put time in all models except the
Gompertz.

The likelihood function is the product of the individual density functions (we are maximizing the product
of all of them), taking into account the characteristics of each regressor.

It doesn't matter which model we use, we always have to know the density function and the survival function.
These are substituted into the likelihood function, maximizing it. Therefore, if we estimate an exponential model in the
PH metric, the estimated coefficient is the same in both PH and AFT, but the sign changes.

Example: What is the exit rate of individuals?

Coeficiente = -5.23 0.05 is the output speed in the PH metric

Coefficient = 5.23 (if I put "time" it's AFT) = estimated survival time on average (because we estimate the
The average that must be included in the intervals is exp(5.2389) = 188.47 days

You might also like