In this chapter we discuss a statistical technique that goes by
various names, such as duration analysis (e.g. the length of time
a person is unemployed or the length of an industrial strike),
event history analysis (e.g. a longitudinal record of events in a
person's life, such as marriage).reliability or failure time analysis
(e.g. how long a light bulb lasts before it burns out), transition
analysis (from one qualitative state to another, such as from
marriage to divorce), hazard rate analysis (e.g. the conditional
probability of event occurrence), or survival analysis (e.g. time
until death from breast cancer). For brevity of exposition, we will
christen all these terms by the generic name of survival analysis
(SA).
The primary goals of survival analysis are: (1) to estimate and
interpret survivor or hazard functions (to be discussed shortly)
from survival data and (2) to assess the impact of explanatory
variables on survival time.
The topic of survival analysis is vast and mathematically
complex. In this chapter our objective is to provide an exposure
to this subject and illustrate it. For further study of this subject,
readers are advised to consult the references.
18.1 An
duration
illustrative
example:
modeling
recidivism
To set the stage, we consider a concrete example. This example
relates to a random sample of 1,445 convicts released from
prison between July 1977 and Jun e 1978 and the time (duration)
until they return to prison. The data were obtained
retrospectively by examining records in April 1984. Because of
different starting times, the censoring times vary from 70 to 81
months. The variables used in the analysis are defined as
follows:
The variable of interest in this study is Dural, the maximum time
until a released convict commits a crime and returns to prison.
We want to find out how Durat is related to the regressors. Also
called covariates, listed above , although we may ~o include all
these variables in the analysis because of collinearity among
some variables, See Table 18.1 on the companion website. .
Before we answer this question, it is essential that we know
some of the terminology used in survival analysis.
18.2 Terminology of survival analysis
Event: "An event consists of some qualitative change that
occurs at a specific point in time The change must consist of a
relatively sharp disjunction between what In precedes what
follows. An obvious example is death. Less obvious, but
nonetheless important. events are job changes. Promotions,
layoffs, retirement, convictions and incarcerations, admission
into a nursing home or hospice facilities, and so on.
Duration spell: It is the length of time before an event occurs,
such as the time when an unemployed person is re-employed, or
the length of time after divorce and the length of time between
successive children. or the person gets remarried , or length of
time before a released prisoner is rearrested
Discrete time analysis: Some events occur only at discrete
time. For example. Presidential elections in the USA take place
every four years and the Census of population is conducted
every 10 years. The unemployment rate in the USA IS published
once a month. There are specialized techniques to handle such
discrete events, such as discrete-time event history.
Continuous time analysis: In contrast to discrete time
analysis, continuous time SA analysis treats time as continuous.
This is often done for mathematical and statistical convenience,
for very few events are observed along a time continuum. In
some cases events can be observed in a small window of time,
such as the weekly unemployment benefit claims. The statistical
techniques used to handle continuous time SA are different from
those used to handle discrete time SA. However, there are no
hard and fast rules about which approach may be appropriate in
a given situation.
The cumulative distribution function (CDF) of time:
Suppose a person is hospitalized and let T denote the time
(measured in days or weeks) until he or she is discharged from
the hospital. If we treat T as a continuous variable, the
distribution of the T is given by the CDF.
Where the expression in the numerator of this function is the
conditional probability of leaving the initial state (e.g. hospital
stay) in the (time) interval {to t + h), given survival up to time t.
Equation (18.4) is known as the hazard function. It gives the
instantaneous rate of leaving the initial state per unit of time.
In simple words, the hazard f unction is the ratio of the density
function to the survivor f unction for a random variable. Simply
stated, it gives the probability that
Someone fails at time t, given that they have survived d up to
that point, failure to be understood in the given context.
Incidentally, note that Eq. (18.7) is also known as the hazard rat
e function, and we will use the terms "hazard function" and
"hazard rate function" interchangeably
Equation (18.7) is an important relations hip. because regardless
of the functional form we choose for the hazard function. H (t).
We can derive the CDF, F (t). from it. Now the question is: how
do we choose fi t) and Set) in practice? We will answer this
question in the next section. In the mean time, we need to
consider some special problems associated with SA.
1 Censoring: A frequently encountered problem in SA is that the data are
often censored. Suppose we follow 100 unemployed people at time t and
follow them until time period (t + h). Depending on the value we choose for
h. there is no guarantee that all 100 people will still be unemployed at time
(t + h); some of them will have been reemployed and some dropped out of
the labor force. Therefore, we will have a censored sample.
Our sample may be right-censored because we stop following our
sample of the unemployed at time t + h. Our sample can also be leftcensored. because we do not know how many of the 100 unemployed were
in that status before time t. In estimating the hazard function we have to
take into account this censoring problem. Recall that we encountered a
similar problem when we discussed the censored and truncated sample
regression models.
2 Hazard function with or without covariates (or regressors): In SA
our interest is not only in estimating the hazard function but also in trying to
find out if it depends on some explanatory variables or covariates. The
covariates for our illustrative example are as given in Section 18.1.
But if we introduce covariates, we have to determine if they are time-variant
or time-invariant. Gender and religion are time-invariant regressors, but
education, job experience, and so on, are time-variant. This complicates SA
analysis.
3 Duration dependence: If the hazard function is not constant, there is
duration dependence. If dh(t ) / dt > O. there is positive duration
dependence. In this case the probability of exiting the initial state increases
the longer is a person in the initial state. For example, the longer a person is
unemployed. his or her probability of exiting the unemployment status
increases in the case of positive duration dependence. The opposite is the
case if there is negative dependency: in this case. dh(t )/ dt < O.
4 Unobserved heterogeneity: No matter how many covariates we
consider, there may be intrinsic heterogeneity among individuals and we
may have to account for this. Recall that we had a similar situation in the
panel data regression models where we accounted for unobserved
heterogeneity by including individual-specific (intercept) dummies, as in the
fixed effects models
With these preliminaries, let us show how survival analysis can be
conducted
18.3 Modeling recidivism duration
There are three basic approaches to analyzing survival data: nonparametric,
parametric, and partially parametric, also known as semi-parametric. In the
nonparametric approach we do not make any assumption about the
probability distribution of survival time, whereas in the parametric approach
we assume some probability distribution.
The nonparametric approach is used in the analysis of life tables, which
have been used for over 100 years to describe human mortality experience.
Actuaries and demographers are obviously interested in life tables, but we
will not pursue this topic in this chapter. The parametric approach is largely
use d for continuous time data
There are several parametric models that are used in du rat ion analysis.
Each depends on the assumed probability distribution, such as the
exponential, Weibull, lognormal, and log logistic. Since the (probability)
density function of each of these distributions is known, we can easily derive
the corresponding hazard and survival functions. We now consider some of
these distributions and apply them to our illustrative example. In each of the
distributions discussed below we assume that h, the hazard rate, can be
explained by on e or more covariates.
But before we consider these models, why not use the traditional normal
linear regression model, regressing Durat on the explanatory variables listed
earlier? The reason why the traditional regression methodology ma y not be
applicable in survival analysis is that, "...the distributions for time to event
might be dissimilar from the normal - they are almost certainly nonsymmetric, they might be bimodal , and linear regression is not robust to
these violations'
18.4 Exponential probability distribution
Suppose the hazard rate h (t) is constant and is equal to h. For our example,
this would mean that the probability of recidivism does not depend on the
duration (time) in the initial state. A constant hazard implies the following
CDF and PDF: