0% found this document useful (0 votes)
32 views5 pages

Appendix 2 An Introduction To The Counting Process Approach To Survival Analysis

hos lem and may app2

Uploaded by

Clancy Birrell
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
32 views5 pages

Appendix 2 An Introduction To The Counting Process Approach To Survival Analysis

hos lem and may app2

Uploaded by

Clancy Birrell
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Applied Survival Analysis: Regression Modeling of

Time-to-Event Data, Second Edition


by David W. Hosmer, Stanley Lemeshow and Susanne May
Copyright 2008 John Wiley & Sons, Inc.

APPENDIX 2
An Introduction to the
Counting Process Approach
to Survival Analysis
We refer to the counting process approach to the analysis of survival time
throughout the text. This method has been the source of many new developments
in the field since it was first used by Aalen |(1975) and (1978)|. Two texts document the mathematical details of this powerful method in a thorough manner.
Fleming and Harrington (1991) and Andersen, Borgan, Gill and Keiding (1993).
Fleming and Harrington (1991) focus primarily on the analysis of survival time
while Andersen et al. (1993) consider analysis of survival time as well as other,
more general statistical problems. We encourage readers of this text to see Andersen et al. (1993, Chapter 1) for an excellent overview of the types of statistical
problems that can be formulated as counting processes. Fleming and Harrington
(1991, Chapter 0) and Andersen et al. (1993, Section II. 1) provide nontechnical
mathematical introductions to the approach. This appendix introduces a few of the
key ideas and constructs used in the counting process approach to the analysis of
survival time. For this reason, many of the more technical mathematical assumptions and details will not be discussed.
Suppose we follow a single subject from time of enrollment, t = 0 , in a study
of a particular cancer until the subject dies from this cancer. Furthermore, we assume it is a 5-year study and that this subject is enrolled on the first day of the
study. Thus the maximum length of follow-up for this subject is 60 months. We
denote the survival time random variable as X. A common approach to modeling
the possibility of right censoring is to assume that there is a second random variable, independent of X, which records the time until observation terminates from
anything other than the event of interest, for example, death from another cause or
loss to follow-up for reasons unrelated to any study factor. We denote this random
variable as Z. The actual observed time random variable is T = min(X,Z) and the
available data for a subject consists of T and an indicator variable C whose value is

359

360

APPENDIX 2

1 if T - X and 0 if T = Z. Thus the variable T records follow-up time and C is


the censoring indicator variable.
Three functions of time central to the counting process approach are: the
counting process
N(t)=l{T<t,C

= l),

the at-risk process


Y(t) =

l(T>t),

and the intensity process


X(t)dt =

Y(t)h{t)dt,

where
h(t)dt = Pr(/ <T <t + dt,C= 1 \T >t)
is the hazard function for survival time. The function /() is the indicator function
whose value is 1 if the argument is true and 0 otherwise.
The counting process records, in our example, whether death from cancer
occurs at time t. The function "counts" this by jumping from a value of 0 to a
value of 1. Suppose our hypothetical subject died from cancer after being in the
study 32 months, (T - 32,C = l). The counting process function for this subject
is equal to zero until 32 months. At exactly 32 months, the function jumps to a
value of 1. The function is equal to 1 for the remaining 28 months of follow-up
time. If the subject's follow-up time is right censored, C = 0, then a death is not
counted and the counting process is equal to 0 for all values of t. If our hypothetical subject was removed from the study at 32 months for reasons unrelated to the
cancer, (7" = 32,C = 0 ) , then the counting process for this subject is equal to zero
over the 60 months of follow-up.
The at-risk process indicates whether the subject is still being followed, at risk
for death, at time t. This function jumps from a value of 1 to a value of 0 when
follow-up ends because of death or censoring. For a hypothetical subject with
follow-up time of 32 months, the at-risk process is equal to 1 from the beginning
of follow-up until 32 months. The function jumps/drops to a value of zero just
after 32 months because the subject is no longer at risk for the remaining 28
months of the study.
The intensity process may be viewed as an "expected number of deaths" at
time t. This follows from the fact that the function is of the form " n x p " (i.e., the

COUNTING PROCESS APPROACH TO SURVIVAL ANALYSIS

361

expected number of events in a binomial distribution). The at-risk process corresponds to "n" and the hazard function to "p".
The process of following the hypothetical subject from time zero to time t
may be thought of as an accumulation of many conditional independent steps,
much like the argument used to construct the Kaplan-Meier estimator in Chapter
2. The total expected number of deaths up to time t is obtained from the intensity
process in the same manner as the cumulative hazard is obtained from the hazard
function, namely by integrating the intensity process over time to obtain
A ( 0 = Jo'A()rfn

j(u)h(u)du,

and this function is called the cumulative intensity process.


Thus one may think of the counting process as the total number of observed
events and the cumulative intensity process as the total number of expected events
up to time /. The difference between these two quantities is a residual-like quantity called the counting process martingale,
M{t) =

N(t)-A{t).

This function is the basis for the martingale residuals that play a central role in
model evaluation methods in Chapter 6. Another way to express the relationship
between the counting, intensity, and martingale processes is via a linear-like model
N(t) = A(t) + M(t).
When expressed in this way, we see that the counting process, the observed part of
the model, is the sum of a systematic component, the cumulative intensity process
and a residual, the martingale process. In our hypothetical example, death or censoring can occur only one time and at the actual follow-up time, T, the value of the
martingale is
. , f l - A ( r ) ifC = I
M(T)
=\
) !
K
' [ O - A ( r ) ifC = 0
It is well beyond the scope of this appendix to explain what makes a process a
martingale and what gives M(t) this quality. We refer the interested reader to the
texts cited above for these technical details. It suffices for the purposes of this
appendix and text to think of M(t) as being similar to a residual.

362

APPENDIX 2

Now suppose we have observations of follow-up time and censoring indicator


variable on n subjects in our hypothetical cancer study. We assume that observations of time are independent and identically distributed. We denote the actual
observed times and right censoring indicator variables in the usual way as (ti,ci ),
i = l,2,...,n. In this setting, a basic result from counting process theory is that the
estimator of the cumulative intensity process for the th subject at time t is

ki(t) = Y,(t)H(t),
where

77, nj

is the Nelson-Aalen estimator of the cumulative hazard at / and

is the number at risk at time tj. The estimator of the martingale residual for the
'th subject at his/her follow-up time is
A/(i / ) = c,-(i,)
= c,.-y, (/,)(/,.)

= c,.-(0
because y,(i.) =/(f, >?,)= 1. We denote this martingale residual as M. We
note that, like residuals from most regression models, 1M = 0.
Assume that we have, in addition to follow-up time and censoring indicator
variables, observations on p fixed (not time-varying) covariates. Assume that we
fit a proportional hazards regression model. The estimator of the cumulative intensity process for the 'th subject at time t is

A(i,x) = -y ( (i) i '*ln[s 0 (f)].

COUNTING PROCESS APPROACH TO SURVIVAL ANALYSIS

363

Thus the value of the martingale residual for the /th subject at his/her follow-up
time is
M, =

ci-e**ln[so(tl)].

The estimated martingale residuals are the basis for many of the diagnostic methods for assessing various aspects of the fitted model described in Chapters 5 and 6.
One, if not the, major theoretical benefit derived from formulating a survival
analysis as a counting process is that a number of theorems from martingale theory
may be used to prove many of the distributional results cited in this text. For example, this theory may be used to prove that the maximum partial likelihood estimators of the coefficients in a proportional hazards model are asymptotically normally distributed with a covariance matrix that may be estimated by the observed
information matrix (Chapter 3). A second example involves the proof that the
Kaplan-Meier estimator and functions of it are asymptotically normally distributed (Chapter 2). The list of applications of this theory in survival analysis is quite
long. The central theme in all of the applications involves proving that a particular
scaled and centered estimator, such as the Kaplan-Meier estimator,
VrtTs(r)-S(r)l, is a martingale.
In summary, we feel that it is important for anyone using the regression methods for the analysis of survival time described in this text to have at least a superficial knowledge of the basics of the counting process paradigm.

You might also like