0% found this document useful (0 votes)
7 views24 pages

Observational Cohort Studies Overview

The document discusses observational cohort studies, highlighting their definition, historical context, and methodology. It contrasts observational studies with experimental cohort studies, emphasizing the importance of group comparability and data analysis techniques. Two illustrative examples, the Framingham Heart Study and the Women’s Health Initiative, demonstrate the application of observational cohort studies in understanding health outcomes related to cardiovascular disease and breast cancer risk.

Uploaded by

suyantounri2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views24 pages

Observational Cohort Studies Overview

The document discusses observational cohort studies, highlighting their definition, historical context, and methodology. It contrasts observational studies with experimental cohort studies, emphasizing the importance of group comparability and data analysis techniques. Two illustrative examples, the Framingham Heart Study and the Women’s Health Initiative, demonstrate the application of observational cohort studies in understanding health outcomes related to cardiovascular disease and breast cancer risk.

Uploaded by

suyantounri2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

7: Observational Cohort Studies DRAFT

7
Observational Cohort Studies
7.1 Introduction
7.2 Historical Perspective
7.3 Assembling and Following a Cohort
7.4 Prospective and Retrospective Cohorts
7.5 Comparisons at Baseline
7.6 Data Analysis
• Arithmetic Principles of Comparison
• Rate Difference and Risk Difference
• Rate Ratio and Risk Ratio
• Relationship Between Rate Ratios and Rate Differences
• Multiple Levels of Exposure (Dose-Response)
7.7 Advanced Topics (Optional): Historically Important Insight

7.1 Introduction

The term cohort derives from the Latin word cohors meaning “an enclosure.” This is an
apt derivation because participants in cohort studies are “enclosed” in a group before
being followed and monitored for relevant events: cohort studies are closed population
studies with individual follow-up of study subjects.

Cohort studies come in experimental and observational forms. * The previous chapter
considered experimental cohort studies in which in which the exposure was randomly
assigned to study subjects by the experimental protocol. This chapter considers
observational cohort studies in which study subjects are classified according to inherent
or environmental attributes or exposure. Thus, the primary distinction between
experimental cohort studies and observational cohort studies is whether the study
exposure is under the direct control of the investigator.

Experimental cohort studies and observational cohort studies share common features.
Both use admissibility criteria to recruit study subjects from a source population; both
follow the experiences of individual study subjects in closed population over time to
monitor health outcomes; both compare the incidence of events in exposed and
nonexposed groups to make causal comparisons. In fact, many of the principals of good
experimental design apply to observational studies. For example, group comparability
must be encouraged through study subject recruitment practices, group comparability
must be assessed before follow-up begins, the ascertainment of study outcomes must be
reliable and valid, data may be analyzed according to initial classifications (intention to
*
When an epidemiologist refers to a “cohort study” without specification, they are usually referring to an
observational cohort study.

© B. Gerstman Page 1 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

treat) or “as treated” (“as exposed”) status, and so no. Principals of experimental studies
thus serve as an important point of reference. It may therefore be useful to review the
prior chapter before beginning this one.

Let us start our consideration of observational cohort studies with two Illustrative
Examples.

Illustrative Example 7.1 (Framingham Heart Study). Many ground-breaking


findings about the causes of heart disease and stroke were identified and confirmed by the
landmark Framingham Heart Study. The initial Framingham study cohort, consisting of
5209 cardiovascular disease-free volunteers between the ages of 20 to 70, was recruited
from the moderately-sized town of Framingham, Massachusetts between 1948 and 1950.
Study participants were followed and examined for coronary disease every two years.
Cases of coronary heart disease (angina, myocardial infarction, and sudden death) were
identified and confirmed using then current recommendations of the New York Heart
Association. Table 6.1 displays data for the 40- to 59-year olds in the study during the
first six years of follow-up. Progressive increases in coronary heart disease incidence
with increasing serum cholesterol levels are noted in both men and women.

TABLE 7.1. Six-Year Incidence Proportions of Coronary Heart Disease According to Initial Serum Cholesterol Level
in 40- to 59-Year old Framingham Heart Study Participants
Serum Cholesterol No. of Incident Cases No. of Individuals Incidence Proportion
(mg/100mL) (%)

Men

< 210 16 454 3.52

210–244 29 455 6.37

≥ 245 51 424 12.03

Women

> 210 8 445 1.80

210–244 16 527 3.04

≥ 245 30 689 4.35

Source: Kannel et al. (1961).

As a more recent example, let us consider an observational study from the The
Women’s Health Initiative (WHI) project. Recall that the WHI project included
both experimental and observational studies. In the previous chapter we considered
the estrogen + progestin WHI experiment. Let us now consider an observational

© B. Gerstman Page 2 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

studies on nonsteroidal anti-inflammatory drugs (NSAID) use and breast cancer


risk.

Illustrative Example 7.2 (WHI NSAIDs and Breast Cancer Study). The
observational component of the Women’s Health Initiative was a long-term ethnically
and geographically diverse, multicenter observational study which studied the experience
of 93,676 women between the ages of 50 and 79 recruited from 40 clinical centers
throughout the United States. The recruitment period started in September 1993 and
ended in July 1998. Participants gave informed consent, were screened for eligibility, and
were followed prospectively for up to 15 years (WHI, 1998).

One of the observational analyses from the Women’s Health Initiative project explored
the relation between analgesic (pain medicine) use and breast cancer occurrence (Harris
et al., 2003). Information about use of analgesics was collected from an interview-
administered questionnaire. Participants were asked if they take aspirin, ibuprofen pills or
tablets, other nonsteroidal anti-inflammatory drug (NSAID) pain pills, or acetaminophen
tablets or capsules. For those individuals who reported using an NSAID or
acetaminophen at least two times in each of the two weeks preceding the interview, the
type of compound, strength (in milligrams), and duration of use (number of years) were
recorded. Use of medication was validated by checking pill bottle labels and prescription
records during interviews.

Breast cancer cases were identified through annual follow-up questionnaires and from
other health care contacts. Cases were confirmed by review or pathology reports,
discharge summaries, operative reports, and radiographic and clinical pathology reports
by physicians and coders blinded to exposure status of potential cases. Follow-up time
for each study subject was accrued from enrollment to the date of diagnosis of breast
cancer, death from a competing (non-breast cancer) cause, loss to follow-up or other
types of withdrawal. Table 7.2 demonstrates trends of decreasing rates of breast cancer
by duration of NSIAD combined (aspirin, ibuprofen, prescription NSAIDS), aspirin
alone, and ibuprofen alone, but no conclusive trend for prescription NSAIDS alone, or
acetaminophen users.

Table 7.2. Breast cancer incidence rates by NSAID type (aspirin, ibuprofen, prescription
NSAID) and acetaminophen use.
Rate
Duration of
Breast Person- per Test for
use at
Group n cancer years 100,000 trend P-
baseline
cases at risk valuea
(years)
p-yrs

Referentb < 1 yr 54,102 955 194,884 490.04 N/A

Any NSAID 1–4 yr 9,000 148 32,127 463.79 0.01


≥5 yr 10,162 83 36,576 404.64

© B. Gerstman Page 3 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Rate
Duration of
Breast Person- per Test for
use at
Group n cancer years 100,000 trend P-
baseline
cases at risk valuea
(years)
p-yrs
Aspirin 1–4 yr 5,124 149 18,231 455.27 0.03
≥5 yr 6,759 99 24,398 405.76
Ibuprofen 1–4 yr 3,469 51 12,553 406.26 0.12
≥5 yr 2,976 42 10,653 394.26

Prescription 1–4 yr 1,615 31 5,552 558.31 0.21


NSAIDs ≥5 yr 947 11 3,388 324.71
Acetaminophen 1–4 yr 2,450 44 8,608 511.18 0.71
≥5 yr 4,675 79 16,698 473.11
Source: Harris et al., 2003
a
Test for trend reported by Harris et al. (2003) based on Wald chi-square test using median duration as an exposure
score.
b
The referent category includes women who reported less than one year of NSAID (aspirin, ibuprofen, prescription
NSAIDs) or acetaminophen use.

7.2 Historical Perspective

The idea of comparing the health of people based on their common characteristics goes
back a very long way. For example, Hippocrates (ca. 460 BC – ca. 370 BC) urged us to
consider

the mode in which the inhabitants live, and what are their pursuits, whether they are fond
of drinking and eating in excess, and given to indolence, or are fond of exercise and labor
and not given to excess in eating and drinking.

The Romans were aware that certain ailments relatively common among occupational
groups. The poet and philosopher Lucretius (ca. 99 BCE – ca. 55 BCE), Martial (ca. 40
AD – ca. 103 AD), and Galen (129 – 199) all wrote on the health of miners, commenting
on their tendency toward pallor and respiratory distress. Martial addressed the diseases of
sulfur workers. But it wasn’t until the 18 th century that Bernardino Ramazzini (1633–
1714) wrote extensively about the diseases of worker-groups in his classical De Morbis
Artificum Diatriba, (1713). De Morbis discusses many different physical (chemicals,
dusts, abrasives) and other work-related hazards, including lifestyle factors:

workers in whom certain morbid affections gradually arise from . . .some particular
posture of the limbs or unnatural movements of the body called for while they work.
Such are the workers who all day long stand or sit, stoop or are bent double…
such as cobblers and tailors . . . [who] become bent, humpbacked, and hold their heads
down like people looking for something on the ground; this is the effect of their sedentary
life and the bent posture of the body as they sit and apply themselves all day to their tasks

© B. Gerstman Page 4 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

in the shops where they sew. . . . Since to do their work they are forced to stoop, the
outermost vertebral ligaments are kept pulled apart and contract a callosity, so that it
becomes impossible for them to return to the natural position. . . . These workers, then,
suffer from general ill-health . . . caused by their sedentary life. . . .

The 18th century English surgeon Percival Pott (1713–1788) was first to identify an
environment carcinogen by describing enormously elevated rates of scrotal cancer in
chimney sweeps which he attributed to “the lodgment of soot in the rugae of the scrotum”
(Pott, 1775).

Eighteenth century French physicians like Pierre Charles Alexandre Louis (1787–
1872) and Phillippe Pinel (1745–1826) brought longitudinal observations of cohorts into
their evaluation of clinical cohorts. In one study, Louis observed superior cure rates in
pneumonia patients who experienced delayed bloodletting compared to those who
experienced early treatment. In a similar vein, Pinel followed the clinical history of
patients with mental illnesses over time to demonstrate superior cure rates at institutions
that practiced humane methods of treatment compared to the standard care of the time.

The Victorian physician and statistician William Farr (1807–1883), who is usually
associated with application of open population statistics, applied longitudinal analyses to
clinical cohorts and noted the importance of cohort analysis by writing “[individual
experiences] should be followed from the beginning to the end; every death or recovery
should be recorded” (Hill, 2003; Gerstman, 2003). The importance of not losing sight of
individual experiences “from the beginning to the end” forms the basis of both cohort and
case-control studies.

In the early 20th century, Joseph Goldberger (1847–1929) made cohort comparisons in
establishing pellagra as a nutrition deficiency disease, rather than an infectious disease, as
was the common belief at the time:

At the State hospital for the insane at Jackson, Miss., there have been recorded 98 deaths
from pellagra for the period between October 1, 1909, and July 1, 1913. At this
institution cases of institutional origin have occurred [entirely] in inmates. … No case…
has developed in a nurse or attendant, although since January 1, 1909, there have been
employed a total of 126 who have served for periods of from 1 to 5 years. In considering
the significance of the foregoing observations it is to be recalled that at all of these
institutions the ward personnel, nurses, and attendants spend a considerable proportion of
the 24 hours, on day or night duty, in close association with the inmates; indeed at many
of these institutions, for lack of a separate building or special residence for the nurses,
these live right in the ward with and of necessity under exactly the same conditions as the
inmates. It is striking therefore that although many inmates develop pellagra after varying
periods of institutional residence…yet nurses and attendants living under identical
conditions appear uniformly to be immune. If pellagra be a communicable disease, why
should there be this exemption of the nurses and attendants? (Goldberger et al., 1913, p.
1684)

The British scientist Janet Elizabeth Lane-Claypon (1877–1967) reported the results of
weight gain in infant cohort fed either boiled cows’ milk (n = 204) human breast milk (n
= 300) (Lane-Claypon, 1912). This investigation revealed that breastfed infants gained

© B. Gerstman Page 5 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

more weight than cows’ milk feed babies during the first 208 days of life, but that the
milk feed rate of weight gain caught up thereafter.

In 1913, the German physician Wilhelm Weinberg (1862–1937) published the results of
a large, retrospective cohort study comparing the experience of an “exposed” cohort of
18,212 children whose fathers and mothers had previously died of tuberculosis to that of
an “nonexposed” cohort of 7,574 children of parents who died of causes other than
tuberculosis. This early cohort study found that the nonexposed cohort had lower
mortality rates and higher fertility rates than did the “exposed” cohort (Morabia &
Guthold, 2007).

It was not until 1935, however, that the first recorded use of the term cohort study was
used (Doll 2001) when Wade Hampton Frost referred to rates of tuberculosis in
generational cohorts. Frost’s generational cohort studies are discussed later in this chapter
(§7.7).

By the middle of the 20th century, the epidemiologic shift from predominantly acute to
chronic causes of morbidity and mortality stimulated the need for long-term follow-up
studies. The refinement of methods was initially addressed to the study of cigarette-
related diseases, cancer, and heart disease, but soon expanded into the study of diseases
with intermediate to long induction. One important heart disease study from, this era—
the Framingham Heart Study—has already been introduced as Illustrative Example 7.1.
Another important early cohort study from this era was the British Doctors Study.

The British Doctors study was launched Richard Doll and Bradford Hill in 1951 when
they sent a seven question questionnaires to all the medical doctors in the United
Kingdom (59,600 initial mailings) asking about their smoking history. The cohort, which
has since been followed for more than half a century (Doll et al., 2004), and has been
responsible for identifying or confirming excess mortality in smokers for dozens of
neoplastic, vascular, respiratory disease. It also confirmed a negative association
between smoking with Parkinson’s disease (Doll et al., 1994). We will consider of the
British Doctor’s Study later in this chapter in Illustrative Example 7.3.

7.3 Assembling and Following a Cohort

The source population for a cohort study is the universe that forms the pool of potential
participants in a study cohort. This universe often shares a common characteristic. For
example, a birth cohort consists of individuals born during a particular period. “Baby
boomers,” for instance, are the pool of individuals born in the years following World War
II. Occupational cohorts comprise people who work in a particular industry or
occupation. The Nurses’ Health Study (Belanger et al., 1978), for example, assembled a
cohort of 122,690 nurses. A cohort of uranium miners may be assembled to study the
effects of radiation (Wagoner et al., 1965). A cohort may also be recruited from a
geographic locale. The Framingham Heart Study, for instance, assembled individuals
from the town of Framingham, Massachusetts (Dawber et al., 1963).

© B. Gerstman Page 6 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

The objective of cohort studies is to accrue accurate evidence about whether specific
exposures cause ill- or good-health. The necessity of unbiased comparisons, therefore,
usually takes precedence over issues concerning generalizability; the need to obtain a
random sample is secondary. In those instances when the objective of the cohort study is
to determine the prevalence or incidence of a condition, random sampling may then
become an issue.

It is essential to obtain the cooperation of the source population, the medical


community that serves it, and its civic leaders before embarking on a cohort study. One
of the reasons the landmark Framingham Heart Study was based in this Massachusetts
town, for example, was because of its supportive population and health care system
(Dawber, 1963).

Even with a highly cooperative population, a certain percentage of individuals invited to


participate in the study will fail to respond or refuse to participate. The non-response
rate in the Framingham study and Nurses Health study were 31% and 29%, respectively
(Dawber et al., 1963; Belanger et al., 1973).

The problem of continued participation must be examined in cohort studies. Subjects who
get lost to follow-up, refuse to continue the study once enrolled, or die from a competing
cause are called withdrawals. Here is an example of a cohort study with a very low
(favorable) withdrawal rate.

Illustrative Example 7.3 (Withdrawal rate, British Doctors Study). The British
Physicians study started in 1951 when the British Medical Association forwarded a
questionnaire about smoking habits to their members. A total of 34,440 men replied,
representing a response rate was about 69% of the men who were alive at the time (Doll
& Peto, 1976). A second questionnaire was sent out in late-1957 / early-1958. By that
time, 3122 of the respondents had died, leaving 31,318 still alive. Of those remaining
alive, 30,810 (98%) replied. By the time third questionnaire was sent out in 1966, an
additional 7,301 had died. Of the remaining 27,139 individuals living, 26,163 (96%)
replied to the survey. The fourth questionnaire, sent out in 1972, had a response rate of
98%. The nonresponse rates of 2%, 4%, and 2% were thought to be this low because of
the cooperative nature of the source population of physicians used to recruit the cohort.

In theory, once an individual is enrolled in a cohort, they are “enclosed” in that cohort for
life. In practice, however, withdrawals from the cohort are inevitable. If withdrawal is
independent of the exposure and disease being studied, the observed exposure–disease
relationship will be unbiased. On the other hand, the exposure–disease relationship will
become distorted if nonresponse and withdrawal are associated with both the study
exposure and disease. Hypothetically, just as an example, had the non-responders and
withdrawers in the Framingham study tended to have both high cholesterol levels and
low heart disease rates, then the observed positive relationship in the Framingham study
(Illustrative Example 7.1, upcoming) would have been biased.

© B. Gerstman Page 7 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

When a rolling period of enrollment is used to recruit study subjects, the experience of
each cohort member is backed-up to “time zero” for the purpose of counting person-time.
Person-time is then tallied for each study subject until they either develop the disease
outcome, are no longer at risk of developing the disease, withdraw from the study, or the
study ends. Then, the cases are tallied and person-time is summed to determine the rate of
disease in the group. For example, in Figure 7.1, there are 2 cases in 43 person-years of
observation for a rate of 0.0465 per year or 4.65 per 100 person-years.

Figure 7.1. (a) Follow-up time in cohort. (b) Same data with person-time backed up to
time zero. [[Link]]

Explanatory factors that distinguish groups in observational cohort studies are


traditionally referred to as exposures. Exposure can represent any personal characteristic
or environmental exposure thought to be related to disease occurrence: the exposure is
the independent variable in the study.

The exposed group in the study may be referred to as the index group. The nonexposed
group may be referred to as the referent group. When there are multiple levels of
exposure in a study, the least exposed group serves as the referent group. In Illustrative
Example 7.1 and Table 7.2, for example, the referent are the women who reported less
than one year of NSAID or acetaminophen use.

Study subjects are periodically assessed for the occurrence of relevant health outcomes
during the follow-up period. Each outcome is confirmed using criteria according to the
study’s case definitions. The criteria that constitute case definitions are based on
established clinical, historical, and laboratory measures (Chapter 12).

The length of the follow-up period of the cohort study reflects the expected ranges of
induction periods for the exposure-disease relationship being evaluated. This may be as
little as a few hours for a disease with a short induction period (e.g., food poisoning) or
may be as much as several decades for diseases with long induction periods (e.g.,
cigarette-related cancers).

© B. Gerstman Page 8 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

7.4 Prospective and Retrospective Cohorts

The Framingham Heart Study followed cohort members in real-time to record events as
they occurred. Cohort studies carried out in this manner are prospective. Cohort studies
can also be carried out by assembling individuals using historical records to reconstruct
longitudinal health histories from the past up until the present. Studies of this second type
are retrospective. Finally, cohort studies that combine prospective and retrospective data
are said to be ambidirectional. Figure 7.2 illustrates these temporal relationships.

The study design feature that determines whether a cohort study is prospective,
retrospective, or ambidirectional is the proximity of data collection to the time events
occurred in real time. Prospective cohort studies use data that are concurrent to the time
of data collection. Retrospective cohort studies use historical data. Ambidirectional
studies use both concurrent and historical data.

Figure 7.2. Proximity of data collection to the occurrence of events. [[Link]]

Also note that the proximity of data collection does not determine whether a study is a
cohort study or case-control study. Cohort studies can be prospective, retrospective, or
ambidirectional. Case-control studies can be retrospective or ambidirectional. For
example, an ambidirectional case-control study can accrue cases as they occur in real
time (prospective accrual of incident cases) with retrospective ascertainment of exposure
information. Experimental studies are always prospective because the investigator must
first assign the exposure before observing its effects.

Retrospective data for cohort studies can be obtained from a variety of sources,
including medical records, administrative data sources, vital records, surveillance
systems, employment records, and through interviewing study subjects or their proxies. A
historically important illustrative example of a retrospective cohort study that used
information from employment records and death certificates follows.

Illustrative Example 7.4 (Retrospective cohort study, Dye Workers) . Case and
colleagues (1954) compiled a retrospective cohort of workmen in Great Britain based on
rosters from 21 companies involved in the manufacture of aniline-based dyes. Data
covered employment histories of 4622 men from 1921 through 1952. Within this cohort,
the investigators found 127 death certificates with mentions of bladder tumors. Based on
vital statistics for the country as a whole, only 3 to 5 such occurrences would have been
expected in a group of similar size and age distribution. Thus, the overall risk of dying of
bladder cancer in the cohort was approximately 30 times the expected rate.

The above example illustrates one of the advantages of retrospective data: the
investigator need not wait the many years required for disease to develop following
exposure to determine its harmful effects. Thus, retrospective cohort studies are well-
suited for the investigation of diseases with long induction. In addition, because the study
used existing records, completion of the study was relatively economical.

© B. Gerstman Page 9 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

7.5 Comparisons at Baseline

The most common objective of observational cohort studies is to provide accurate


evidence about the independent effects of various factors on disease occurrence. To
accomplish this goal, like-to-like comparisons are necessary. This dated but still eloquent
passage by George Bernard Shaw (1911, pp. lxiv–lxv) reminds us that comparisons of
health in different populations do not necessarily provide clear evidence of an effect:

Comparisons which are really comparisons between two social classes with
different standards of nutrition and education are palmed off as comparisons
between the results of a certain medical treatment and its neglect. Thus it is easy
to prove that the wearing of tall hats and the carrying of umbrellas enlarges the
chest, prolongs life, and confers comparative immunity from disease; for the
statistics show that the classes which use these articles are bigger, healthier, and
live longer than the class which never dreams of possessing such things. It does
not take much perspicacity to see that what really makes this difference is not the
tall hat and the umbrella, but the wealth and nourishment of which they are
evidence, and that a gold watch or membership of a club in Pall Mall might be
proved in the same way to have the like sovereign virtues. A university degree, a
daily bath, the owning of thirty pairs of trousers, a knowledge of Wagner’s music,
a pew in church, anything, in short, that implies more means and better nurture
than the mass of laborers enjoy, can be statistically palmed off as a magic-spell
conferring all sorts of privileges.

George Bernard Shaw is referring to the concept of confounding—the mixing-up of


effects and false attribution of cause. To avoid confounding, the hypothetically ideal
referent group would consist of the same individuals as the exposed group had they not
been exposed to the risk factor being studies. However, because individuals cannot be
simultaneously exposed and unexposed to risk factors, this is impossible in fact, and is
therefore referred to as a counterfactual ideal. Counterfactuality suggests a useful way
to think about the suitability of a nonexposed referent group and further reinforces the
importance of making like-to-like comparisons.

Illustrative Example 7.5 (Comparisons at baseline, Nurses Health Study) . The


distribution of risk factors at baseline should be consider in all cohort studies. Table 7.3
lists the distribution of selected risk factors for coronary artery disease for the women
participating in the Nurses Health Study of postmenopausal estrogen use and
cardiovascular disease (Stampfer et al., 1991). This table suggests that the groups were
generally comparable, although some evidence of difference in lifestyle-related risk
factors is also evident, e.g., current hormone users were slightly less likely to have
diabetes and a BMI exceeding 29, and were more likely to engage in vigorous physical
activity at least once per week.

© B. Gerstman Page 10 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Table 7.3. Proportion of Women in the Nurses Health Study with coronary risk
factors according to postmenopausal hormone use cohorts. a The total sample size
was 48,470.

Estrogen Use

Current Former None

Parental MI before 60 10.6 10.0 9.3

Hypertension 23.2 25.0 21.8

Diabetes m. 2.7 3.8 3.5

High serum cholesterol 9.9 11.2 7.6

15 – 24 cigarettes / day 11.2 14.7 14.5

BMI ≥ 29 9.8 13.3 15.0

Surgical menopause 50.3 39.3 9.3

Past use of oral 34.0 27.6 23.9


contraceptives

Vigorous physical 48.2 43.1 42.4


activity ≥ 1 time/week

Mean dietary intake 27.6 26.2 26.7


saturated fats
a
Proportions have been adjusted for age.
Source: Stampfer et al., 1991.

When groups demonstrate differences at baseline, statistical adjustment methods may


be used during data analysis to help “control” for group differences that have the
potential to cause confounding. For example, in the Nurse’s Health Study cited in
Illustrative Example 7.5, a statistical modeling technique known as proportional-hazards
regression was used to help address potential confounding due to age, cigarette smoking,
hypertension, diabetes, high serum cholesterol, parental MI history, BMI, past use of oral
contraceptive, and calendar time trends.

To encourage group comparability, investigator may restrict study participants to


individuals with or without certain characteristics during the recruitment phase of a study.
This is achieved by imposing rigorous admissibility criteria when recruiting study
subjects.† For example, a study may restrict itself to individuals from a specific

Feinstein (1989) points out that, although the application of admissibility criteria are often associated with

© B. Gerstman Page 11 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

socioeconomic group in order to avoid the potential for confounding by socioeconomic


status. As another example, a study that omits smokers will prevent confounding due to
smoking. Thus, by creating a relatively homogenous study base, group comparability is
improved and the researcher is better able to tease out the effects of the risk factors being
studied.

Matching may also be used to limit the potential for confounding. Matching may be
accomplished through individual matching or frequency matching. Individual matching
is achieved by matching subjects on relevant cofactors as part of the recruitment process.
In studying the effects of smoking, for instance, we may match each 30-year old smoker
with a similarly-aged non-smoker. Frequency matching, on the other hand, balances the
number of smokers and non-smokers by age. The intent of both types of matching is to
create group comparability so that the matched-on factors can no longer confounds the
results of the study. Beware, however, that matched data often require specific analytic
techniques.

7.6 Data Analysis


Arithmetic Principles of Comparison

The objective of cohort analysis is to determine the extent to which an exposure increases
or decreases the incidence of disease in affected individuals. The effect of the exposure
can be quantified in absolute terms or in relative terms.

Simple data. To address absolute and relative measures of effect, let us consider a simple
example in which the rate of disease in an exposed group is 2 per 100 person-years and
the rate in the nonexposed group is 1 per 100 person-years.

Absolute measure of effect (rate difference RD): We may say that the rate in exposed
group exceeds the rate in the nonexposed group 1 per 100 person-years. This is an
absolute measure of the effect, since it tells us in absolute terms the amount of disease
that can be attributed to the exposure. Note that the absolute measure of effect is derived
by subtraction: (2 per 100 person-years) MINUS (1 per 100 person-years) = 1 per 100
person-years.

Relative measure of effect (relative risk RR). We may say that the exposed group’s
rate is twice the rate of the nonexposed group’s rate. This describes a relative measure of
effect since it tells us the proportional excess but does not inform us how much additional
disease occurs will occur in absolute terms. Note that the relative measure of effect is

derived by division: = 2.

Difference relative to baseline (“relative risk difference” RRD). The relative effect of
an exposure can also be expressed in terms of a difference relative to baseline. For
recruiting subjects for clinical trials, the method is equally important when recruiting subjects for
observation studies.

© B. Gerstman Page 12 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

example, we can say that the rate in the exposed group is 100% greater than the rate in
the nonexposed group. This merely changes the way the relative comparison is expressed
but does not change its meaning. Notice that it would be incorrect to say that the exposed
rate is 200% greater than the nonexposed rate because this would imply that the exposed
group’s rate is three times that of the nonexposed group’s rate, when in fact it is only
twice as large. To derive an expression of the difference relative to baseline, simply

subtract 1 from the ratio of the two rates. For example, MINUS

1 = 2 – 1 = 1 or 100%.

Rate Difference and Risk Differences

Rate Difference. The rate difference is merely the rate in the exposed group minus the
rate in the nonexposed group:

Rate Difference = Rateexposed – Ratenonexposed (7.1)

This statistic reflects the absolute excess associated with exposure to the risk factor.

Illustrative Example 7.6 (Rate Difference, oral contraceptive estrogen dose).


An observational cohort study of venous thromboembolism (pulmonary embolism and
deep venous thrombosis) among users of different oral contraceptive formulations found
that users of formulations containing 50 µg of estrogen had a rate of
= 7.04 per 10,000 person-years; the rate in users of formulations

with less than 50 µg of estrogen in the source population was =


4.17 per 10,000 person-years (Gerstman et al., 1991). Therefore, the Rate Difference =
7.04 per 10,000 person-years – 4.17 per 10,000 person-years = 2.87 per 10,000 person-
years. This represents an expectation of 2.8 additional cases of venous thromboembolism
per 10,000 users per year with 50 µg formulations compared to lower dose formulations.

The precision the Rate Difference estimate can be gauged by calculating its confidence
interval. Let us use the free online application [Link] (Dean et al., 2011) to
calculate this confidence interval. Go to [Link] and select “Person-time →
Compare two rates” in the left panel. Figure 7.3 exhibits the data input screen for this
example. Figure shows part of the output you will see after clicking the “Calculate”
button. Among the results, the 95% confidence interval for the Risk Difference is shown
as 2.868 with a confidence interval of 0.8622, 4.873. The confidence interval for the Rate
Difference should be reported as (0.8 to 4.9) per 10,000 person-years. ‡ The confidence
provides a range of Risk Differences values compatible with the data at 95% confidence. §

As a general rule, epidemiologic results should be reported with only 2 or 3 significant digits to prevent an
appearance of quasi-precision.

§
The confidence interval addresses imprecision in the data, but does not address nonrandom sources of

© B. Gerstman Page 13 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Note that a Rate Difference of 0 indicates no association between the exposure and
disease, a positive Rate Differences indicates a positive association, and a negative Rate
Differences indicates a negative association. Illustrative Example 7.6 (oral contraceptive
estrogen dose) for example, demonstrates a positive association between oral
contraceptive estrogen dose and venous thromboembolism. An example of a negative
association follows.

Illustrative Example 7.7 (Rate Difference, Physical Fitness And Mortality). A


study of physical fitness and mortality found that men who improved their physical
fitness from the unfit- to the fit-level had an age-adjusted mortality rate of 67.7 per
10,000 persons-years (Blair et al., 1995). Men who were unfit at both examinations had
an age-adjusted mortality rate of 122.0 per 10,000 person-years. Therefore, improved
physical fitness was associated with a rate difference of (67.7 per 10,000 persons-years) –
(122.0 per 10,000 person years) = −54.3 per 10,000 person-years. This suggests 54.3
fewer deaths per 10,000 persons-years associated with improved fitness.

Figure 7.3. OpenEpi’s input screen for comparing two incidence rates, with data
for Illustrative Example 7.6. [[Link]]

Figure 7.4. Output screen from OpenEpi’s comparing two rates module showing results
for Illustrative Example 7.6. [[Link]]

Risk Difference. When the measures of occurrence in the study is an incidence


proportion (“risk”), as opposed to an incidence rate, use the Risk Difference to measure
the effect of the exposure in absolute terms:

Risk Difference = Riskexposed – Risknonexposed (7.2)

Illustrative Example 7.8 (Risk Difference, Framingham Heart Study). Table 7.1
lists incidence proportions (risks) from the first 6-years of follow-up in the Framingham
Heart Study. Because there are multiple levels of exposure in this table, the least exposed
level (<210 mg/dl serum cholesterol) serves as the “nonexposed” referent group. The
Risk Difference associated with intermediate serum cholesterol levels (210–244 mg/dl) =
Riskint – Risklow = 6.37% – 3.52% = 2.85%, indicating an excess of 2.85 cases per 100.
The Risk Difference associated with high serum cholesterol (≥ 245 mg/dl) = Risk high –
Risklow = 12.03% – 3.52% = 8.51%, suggesting an excess of 8.5 cases per 100 people
over the follow-up period.

95% confidence interval for the Risk Difference (Intermediate vs. Low). Figure 7.5
demonstrates the use of [Link] to calculate the 95% confidence interval
for the Risk Difference comparing the intermediate cholesterol group of men in The
Framingham Study to the low cholesterol group. In the left panel of [Link],
select Counts → Two by Two Table. Then enter the number of exposed and
nonexposed cases into the onscreen 2-by-2 table, as shown in Figure 7.5. Note that

error.

© B. Gerstman Page 14 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

29 of the 455 of the men with intermediate cholesterol group developed coronary
heart disease compared to 16 of the 454 men in low cholesterol group (Table 7.1).
We were required to calculate the number of noncases in each group: in the
intermediate cholesterol group there were 455 – 29 = 426 noncases; in the low
cholesterol group, there were 454 – 16 = 438 noncases. Clicking the “Calculate”
button derived the estimates shown in Figure 7.6: the Risk Difference of 2.849% is
associated with a 95% confidence interval of 0.0362% to 5.663%.

Figure 7.5. OpenEpi’s input screen for comparing two proportions using the data from
Illustrative Example 7.8. [[Link]]

Figure 7.6. Output from OpenEpi’s comparisons of two incidence proportions,


Illustrative Example 7.8. [[Link]]

Rate Ratio and Risk Ratio

Rate Ratio. Rate ratios are derived by dividing the rate in the exposed group by the rate
in the nonexposed group:

(7.3)

This statistic quantifies the excess risk associated with the exposure in relative terms.
The Rate Ratio is the risk multiplier associated with the exposure. For example, a Rate
Ratio of 2 indicates that the exposure doubles the risk in the nonexposed group; a Rate
Ratio of 0.5 indicates that the exposure cuts the risk in half; and so on.

Illustrative Example 7.9 (Rate Ratio, Oral Contraceptive Estrogen Dose). For
the data presented in Illustrative Example 7.6, the Rate Ratio is

= 1.69. Thus, the rate associated with the higher dose oral

contraceptive was 1.69 times that of the lower dose formulations. In other words, the rate
was 69% higher (in relative terms) with the higher dose formulations.

Figure 7.4 includes 95% confidence limits for this Rate Ratio, demonstrating a
confidence interval of 1.179 to 2.413, i.e., (1.2 to 2.4).**
The Rate Ratio in the above example is greater than 1, indicating a positive association
between the exposure and disease. Here’s an illustration of a negative association.
Illustrative Example 7.10 (Rate Ratio, Physical Fitness and Mortality). Recall
Illustrative Example 7.7. The exposure for this analysis was improved fitness.
The disease outcome was death. Data are age-adjusted mortality rates of 67.7 per

**
Rounded to two significant digits to avoid an appearance of pseudo-precision.

© B. Gerstman Page 15 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

100,000 person years in the exposed group and 122.0 per 100,000 person-years in

the nonexposed group. The Rate Ratio is = 0.55,

indicating a negative association between improved fitness and mortality. Specifically,


mortality is almost cut in half.

The Rate Ratio of 0.55 can be re-expressed as a relative difference by subtracting 1 from
the RR estimate and multiplying by 100%. Thus, the relative change in mortality
associated with improved fitness is (RR – 1) × 100% = (0.55 – 1) × 100% = ‒45%.
Thus, this Rate Ratio of 0.55 represents a 45% reduction in mortality.

Risk Ratio. The ratio of two incidence proportions (average risks) is a risk ratio:

(7.4)

Both the risk ratio and rate ratio are referred to as “relative risks” because they have the
same interpretation as a “risk multiplier.”

Illustrative Example 7.11 (Risk Ratio, Framingham Heart Study). Recall the
Framingham data from earlier Illustrative Examples and Table 7.1. The Risk Ratio
comparing the coronary heart disease occurrence in the men with intermediate levels
cholesterol (210 – 244 mg/dl) to those with low serum cholesterol (< 210 mg/dl) is

= =1.81. Thus, the intermediate cholesterol group had 81%

greater risk than the low cholesterol group.

The Risk Ratio comparing the men in the cohort with high serum cholesterol levels (≥

245 mg/dl) to low serum cholesterol is = =3.42.

95% confidence interval (intermediate vs. low). In Illustrative Example 7.8 we entered
the data for the men with intermediate cholesterol levels and low cholesterol levels into
[Link]’s 2-by-2 table program to calculate a confidence interval for the Risk
Difference. The output is shown in Figure 7.6. This output also includes the 95%
confidence interval for this Risk Ratio, and is shown as (0.9962, 3.283).

Relation Between Rate Ratios and Rate Differences

Rate Ratios and Rate Differences describe different aspects of the exposure‒disease
relationship. Let us examine the relationship between cigarette smoking, lung cancer, and
coronary disease to demonstrate this point. Table 7.3 reports rates for these outcomes
from the first 20 years of the British Physicians cohort study.

© B. Gerstman Page 16 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

TABLE 7.3. Age-adjusted Mortality Rates for Lung Cancer and Ischemic Heart Disease in Smokers and
Nonsmokers

Smokers Nonsmokers Rate Difference Rate Ratio

Disease (per 100,000 p-yrs) (per 100,000 p-yrs) (per 100,000 p-yrs)

Lung cancer 104 10 94 10.40

Coronary disease 565 413 152 1.37


Source: Doll and Peto (1976).

Notice that the Rate Ratio of smoking and lung cancer (10.40) is much larger than the
Rate Ratio for smoking and coronary disease (1.37). Thus, the relationship between
smoking and lung cancer is stronger than the relationship between smoking and coronary
disease. In contrast, the Rate Difference is much larger for smoking and coronary disease
(152 additional cases per 100,000 smokers per year) than for smoking and lung cancer
(94 additional cases per 100,000 smokers per year). The basis of this apparent paradox is
that coronary disease is much more common than lung cancer: even a modest relative
increase in coronary disease risk affects many more people than a large relative increase
in lung cancer rate.

The algebraic relationship between a Rate Difference (RD) and Rate Ratio (RR) is
understood by noting RD = R1 – R0, where R1 represents the rate in the exposed group and
R0 represents the rate in the nonexposed group. This expression can be re-written

, demonstrating that the Rate Difference is equal to

baseline rate R0 times the segment of the RR above (or below) an RR of 1 (i.e., RR‒ 1).
Thus, a small RR on top of a large R0 will produce many more cases that a large RR on
top of a small R0.

Multiple Levels of Exposure (Dose-Response)

For exposures that are measured at multiple levels, the least exposed group serves as the
referent group. Comparisons are then made to the referent group’s rate or risk in each
instance.

Let R0 represent the rate or risk of disease in the referent group and let Rk represent the
rate or risk or disease associated with the kth level of exposure. The rate ratio or risk
ratio associated with exposure level k is:

(7.5)

© B. Gerstman Page 17 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Illustrative Example 7.11 has already demonstrated that the Framingham men
demonstrate = 1.81 (for intermediate vs. low cholesterol) and

= 3.42 (for high vs. low cholesterol).

Test for trend, proportions. OpenEpi’s “Counts → Dose-Response” application


calculates an extended Mantel-Haensel test for trend (Mantel, 1963). Figure 7.7 exhibits
the data entry screen in OpenEpi for testing the dose-response relation. Data from Table
7.1 for the Framingham have been entered into the on-screen table. Figure 7.9 exhibits
the output showing a Mantel-Haenszel chi-square test statistic for trend of 22.91 with 1
degree of freedom, P = 0.0000017; this suggest the trend observed in the data cannot be
easily ascribed to chance.

Figure 7.8. OpenEpi’s Dose-Response Application for count (proportion) data. Data for
Framingham men in Table 7.1 have been entered. [[Link]]

Figure 7.9. Output from OpenEpi’s Dose-Response program. [[Link]]

As similar approach is used when analyzing rates when the exposure groups can be
ordered. Table 7.4 shows the breast cancer rates by duration of NSIAD use (aspirin,
ibuprofen, perscriptions NSAIDS) for the data from Illustrative Example 7.2 (Harris et
al., 2003). Note that the rates and ratio ratios progressively decline with increased
duration of NSAID use.

Table 7.4. Breast cancer rates in NSAID users by duration of use.


Crude Rate Crude
Duration of Use No. of cases Person-years
per 100,000 p-yrs Rate Ratio
< 1 yr 955 194,884 490.04
(referent) = 1.00

1–4 yrs 149 32,127 463.79


= 0.95

≥ 5 yrs 148 36,576 404.64


= 0.83
Source: Data from Harris et al. 2003.

Test for trend, rates. OpenEpi does not include an application for testing trends in rates,
but WinPEPI’s†† “Describe program B. Appraise a sequence of rates” does. Figure 7.9
exhibits the input screen from this application with the data from Table 7.4 entered.
Figure 7.10 demonstrates part of the output, revealing a P-value of 0.031.‡‡

††
For background information of WinPEPI, see the 2011 article by Abramson posted at [Link]-
[Link]/content/8/1/1.

‡‡
P-value for trend reported in Table 7.2 differs slightly because it was based on a different procedure and
duration scores were weighted based on median duration.

© B. Gerstman Page 18 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Figure 7.10. WinPEPI’s “Describe program B. Appraise a sequence of rates” data


entry screen. [[Link]]

Figure 7.11. Output from WinPEPI’s “Describe B. Appraise a sequence of rates”


program. [[Link]]

7.7 Advanced Topic (Optional): Historically Important Insight

Wade Hampton Frost (Figure 7.11), the first professor of epidemiology in the United
States, coined the term cohort study to describe a studies of tuberculosis rates in birth
cohorts born in different periods. By distinguishing the standard open-population rate
analysis to birth cohort rates, Frost was able to sort out the perplexing shift in age-peaks
observed over time.

Figure 7.12. Wade Hampton Frost (1880 – 1938). Courtesy of Historical Collections &
Services, Claude Moore Health Sciences Library, University of Virginia. [[Link]]

Table 7.5 shows data from Frost’s posthumously published study on tuberculosis
mortality between 1880 and 1930 (Frost, 1939). Data are from the state of Massachusetts
Age- and time-trends can be tracked by reading rates across rows and down columns,
respectively. (Ignore the shaded diagonal for now.) Figure 7.12, which plots these rates
for 1880, 1910, and 1930, demonstrates the following trends:

1. Higher rates in earlier calendar years. The rates in 1880 (within each age group)
are higher than the rates in 1910. The rates in 1910 are higher than the rates in
1930.
2. There are childhood peaks at 0 to 4 years of age. Rates drop precipitously after
age 5.

3. The peak during early and mid-adulthood (identified with a ¤ in Figure 7.12) have
shifted over time. In 1880, the adult peak was in the 20- to 29-year-olds. In 1910,
the peak was at about age 40. In 1930, the adult peak had shifted to 50- to 59-
year-olds.

© B. Gerstman Page 19 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

TABLE 7.5. Tuberculosis Mortality Rates per 100,000 by Age, Year, Males,
Massachusetts, 1880 to 1930a

Age 1880 1890 1900 1910 1920 1930

0–4 760 578 309 209 108 41

5–9 43 49 31 21 24 11

10–19 126 115 90 36 49 21

20–29 444 361 288 207 149 71

30–39 378 368 296 253 164 115

40–49 364 336 253 253 175 118

50–59 366 325 267 252 171 127

60–69 475 346 304 246 172 95

70+ 672 396 343 163 127 95


Source: Frost (1939).
a
The experience of the 1880 birth cohort is shaded along the diagonal.

Figure 7.12. Cross-sectional tuberculosis mortality rates for men, calendar years 1880,
1910, and 1930; ¤ indicates peak rate in adults (Frost, 1939).[[Link]]

That tuberculosis rates were dropping over time (point 1) had been known for some time.
The most likely explanation for this phenomenon was decreasing levels of the agent in
the environment.

In contrast, there was no reason to believe that the precipitous drop in mortality after age
5 (point 2) was due to less exposure to the agent. Neither could the precipitous increase in
occurrence after age 10 be explained in terms of exposure levels. Such downward and
upward shifts in mortality by age is due to changes in host resistance or, as Frost put it,
“the balance established between the destructive forces of the invading tubercle bacillus,
and the sum total of host resistance” (1939, p. 92).

The shifting peak in adulthood (point 3) was initially perplexing. However, this was
shown to be an artifact of the use of cross-sectional open population rates when Frost
rearranged the rates to emulate the longitudinal experience of birth cohorts. This is
accomplished by reading data along diagonals of tables. As an example, the mortality of
the 1880 birth cohort is shaded in Table 7.5. When mortality rates in birth cohorts are
compared, the age peak in tuberculosis mortality is consistently in the 20- to 29-year-old
group (Fig. 7.13). Thus, there was no change in the pattern of age susceptibility over
time. This finding had immediate relevance because there was some fear in the 1930s that
postponement of infection to later ages was causing more serious disease to occur, as is

© B. Gerstman Page 20 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

indeed the case with some other infectious diseases such as measles and chickenpox; it
was postulated that early exposure might afford some immunologic benefits. Frost, on the
other hand, believed that contact with the agent was to be avoided at all ages. His birth
cohort analysis supported this view demonstrating the need to avoid contact at all ages
(Comstock, 2001).

Figure 7.13. Tuberculosis mortality for the birth cohorts of 1870, 1880, 1890, 1900, and
1910. (Frost, 1939). [[Link]]

EXERCISES

7.1 In exercise 5.7 we evaluated a cohort study of risk factors for agriculture-related
injuries among African-American and Caucasian farmers and African-American
farm workers (McGwin et al., 2000). A total of 1,246 subjects (685 Caucasian
owners, 321 African-American owners, and 240 African-American workers) were
enrolled between January 1994 and June 1996. Demographic, farming, and
behavioral information was collected at baseline. Subjects were contacted
biannually to monitor the occurrence of an agriculture-related injury (McGwin et
al, 2000). Some of the data from this study are presented in this Table:

Group Agricultural-related Person-years


of Observation
Injuries

Caucasian Farm Owners 67 2047

Af-American Farm Owners 27 821

Af-American Workers 37 359

(A) Is this a prospective cohort study or retrospective cohort study?

(B) Calculate the Rate Ratios of injury in each group using the Caucasian
owners as the referent group.

(C) The objective of a cohort study is to evaluate the independent contribution


of various risk factors. Is race an independent risk factor for injury? Is
being a worker an independent risk factor? Explain.

REFERENCES
Abramson, J. H. (2011). WINPEPI updated: computer programs for epidemiologists, and
their teaching potential. Epidemiologic Perspectives & Innovations, 8(1), 1
[Link]

© B. Gerstman Page 21 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Belanger, C. F., Hennekens, C. H., Rosner, B., & Speizer, F. E. (1978). The nurses'
health study. The American journal of nursing, 78(6), 1039-1040.

Blair, S. N., Kohl, H. W., 3rd, Barlow, C. E., Paffenbarger, R. S., Jr., Gibbons, L. W., &
Macera, C. A. (1995). Changes in physical fitness and all-cause mortality. A
prospective study of healthy and unhealthy men. JAMA, 273, 1093–1098.

Case, R. A. M., Hosker, M. E., McDonald, D. B., & Pearson, J. T. (1954). Tumors of the
urinary bladder in workmen engaged in the manufacture and use of certain dyestuff
intermediates in the British chemical industry. British Journal of Industrial Medicine,
11, 75–104.

Comstock, G. W. (2001). Cohort analysis: W.H. Frost’s contributions to the


epidemiology of tuberculosis and chronic disease. Sozial- und Präventivmedizin
(Social and Preventive Medicine), 46, 7–12.

Dawber, T. R., Kannel, W. B., & Lyell, L. P. (1963). An approach to longitudinal studies
in a community: the Framingham Study. Annal of the New York Academy of
Science, 107, 539-556.

Dean, A. G., Sullivan, K. M., & Soe, M. M. (2011, 2011/23/06). OpenEpi: Open Source
Epidemiologic Statistics for Public Health, Version 2.3.1. . 2011/07/12, from
[Link].

Dinger, J. C., Heinemann, L. A., & Kuhl-Habich, D. (2007). The safety of a


drospirenone-containing oral contraceptive: final results from the European
Active Surveillance Study on oral contraceptives based on 142,475 women-years
of observation. Contraception, 75(5), 344-354.

Doll, R. (2001). Cohort studies: History of the method. I. Prospective cohort studies.
Sozial- und Präventivmedizin (Social and Preventive Medicine), 46, 75–86.

Doll, R., & Peto, R. (1976). Mortality in relation to smoking: 20 years’ observations on
male British doctors. British Medical Journal, 2(6051), 1525–1536.

Doll, R., Peto, R., Wheatley, K., Gray, R., & Sutherland, I. (1994). Mortality in relation
to smoking: 40 years' observations on male British doctors. British Medical Journal,
309(6959), 901-911.

Doll, R., Peto, R., Boreham, J., & Sutherland, I. (2004). Mortality in relation to smoking:
50 years' observations on male British doctors. [Research Support, Non-U.S. Gov't].
BMJ, 328(7455), 1519.

Feinstein, A. R. (1989). Epidemiologic analyses of causation: the unlearned scientific


lessons of randomized trials. J Clin Epidemiol, 42(6), 481-489; discussion 499-
502.

© B. Gerstman Page 22 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Frost, W. H. (1939). The age selection of mortality from tuberculosis in successive


decades. [Reprinted in the American Journal of Epidemiology, 141(1), 4–9, (1995)].
American Journal of Hygiene, 30, 90–96.

Gerstman, B. B., Piper, J. M., Tomita, D. K., Ferguson, W. J., Stadel, B. V., & Lundin, F.
E. (1991). Oral contraceptive estrogen dose and the risk of deep venous
thromboembolic disease. American Journal of Epidemiology, 133(1), 32-37.

Gerstman, B. B. (2003). Comments regarding "On prognosis" by William Farr (1838),


with reconstruction of his longitudinal analysis of smallpox recovery and death
rates. Soz Praventivmed (Social and Preventive Medicine), 48(5), 285-289.

Goldberger, J. (1914). The Etiology of Pellagra: The Significance of Certain


Epidemiological Observations with Respect Thereto. Public Health Rep, 29(26),
1683-1686.

Harris, R. E., Chlebowski, R. T., Jackson, R. D., Frid, D. J., Ascenseo, J. L., Anderson,
G., et al. (2003). Breast cancer and nonsteroidal anti-inflammatory drugs:
prospective results from the Women's Health Initiative. Cancer Research, 63(18),
6096-6101.

Hill, G. B. (2003). Comments on the paper "On prognosis" by William Farr: a forgotten
masterpiece. Soz Praventivmed, 48, 225-226.

Kannel, W. B., Dawber, T. R., Kagan, A., Revotskie, N., & Stokes III, J. (1961). Factors
of risk in the development of coronary heart disease—six-year follow-up experience.
Annals of Internal Medicine, 55, 33–50.

Lane-Claypon, J. E. (1912). Report to the local government board upon the available
data in regard to the value of boiled milk as a food for infants and young animals.
Number 63. London, United Kingdom: His Majesty′s Stationary Office.

Levin, M. L. (1953). The occurrence of lung cancer in man. Acta Unio Int Contra
Cancrum, 9, 531–541.

Mantel, N. (1963). Chi-square tests with one degree of freedom; extensions of the
Mantel-Haenszel procedure. Journal of the American Statistical Association, 58,
690-700.

Miettinen, O. (1976). Estimability and estimation in case-referent studies. American


Journal of Epidemiology, 103(2), 226-235.

Morabia, A., & Guthold, R. (2007). Wilhelm Weinberg's 1913 Large Retrospective
Cohort Study: a rediscovery. American Journal of Epidemiology, 165(7), 727-
733.

© B. Gerstman Page 23 Last printed 7/13/2011 2:46:00 PM


7: Observational Cohort Studies DRAFT

Ramazzini, B. (1713; 2001 reprint). De morbis artificum diatriba [diseases of workers].


American Journal of Public Health, 91(9), 1380-1382.

Rehn, L. (1895). Blasengeschwulste bei Fuchsin-Arbeitern. Arch Klin Chir, 50, 588-600.

Shaw, G. B. (1911). The doctor’s dilemma, with a Preface on doctors. New York:
Brentano’s.

Stampfer, M., Colditz, G., Willett, W., Manson, J., Rosner, B., Speizer, F., et al. (1991).
Postmenopausal estrogen therapy and cardiovascular disease. Ten-year follow-up
from the nurses' health study. New England Journal of Medicine, 325(11), 756-
762.

[WHI, 1998] Design of the Women's Health Initiative clinical trial and observational
study. The Women's Health Initiative Study Group. (1998). Controlled Clinical
Trials, 19(1), 61-109.

Winkelstein, W., Jr. (2004). Vignettes of the history of epidemiology: Three firsts by
Janet Elizabeth Lane-Claypon. American journal of epidemiology, 160(2), 97-
101.

Williams, D. (1994). Epidemiology. Chernobyl, eight years on. [News]. Nature,


371(6498), 556.

© B. Gerstman Page 24 Last printed 7/13/2011 2:46:00 PM

Common questions

Powered by AI

The Framingham Heart Study marked a pivotal advancement in cohort study methodologies, particularly in chronic disease epidemiology. Its long-term follow-up design and comprehensive data collection on risk factors like serum cholesterol contributed significantly to understanding heart disease causes and prevention. The study's findings on the relationships between cholesterol levels and heart disease helped establish cholesterol as a modifiable risk factor, influencing public health guidelines and preventive measures worldwide .

Confidence intervals provide a range of values within which we can be reasonably sure the true effect size lies, reflecting the precision of the estimate. In the Framingham Heart Study, a Risk Difference was calculated with a 95% confidence interval, indicating the reliability of the effect of different serum cholesterol levels on heart disease risk. By showing the range, confidence intervals help assess the statistical significance and potential variability of the results .

Wade Hampton Frost's cohort studies were foundational in illustrating the value of generational cohort analysis for chronic disease epidemiology. By focusing on birth cohorts, Frost helped demonstrate the temporal relationship between exposures and long-lasting health outcomes. His work emphasized the need to consider long-term exposure effects and life-course influences, shaping modern epidemiological practices by highlighting the importance of prospective cohort studies in monitoring chronic diseases over time .

The McGwin et al. study found differences in injury rates among African-American and Caucasian farm workers and owners, suggesting that race and occupation may play roles in agriculture-related injury risks. The analysis would consider the calculated Rate Ratios, using Caucasian owners as a referent group, to determine if these factors independently contributed to risk. This requires examining whether these factors influenced injury rates when controlling for other variables, acknowledging complex interactions in occupational settings .

The British Doctors Study aimed to evaluate the health effects of smoking among UK medical doctors. Key findings included identifying excess mortality in smokers due to diseases such as cancer, vascular, and respiratory diseases, as well as highlighting a potential protective effect against Parkinson's disease. This study provided significant evidence of the health risks associated with smoking over a long follow-up period .

The lack of pellagra cases among nurses and attendants, despite identical living conditions with inmates who developed the disease, suggests that pellagra may not be a communicable disease. The immunity observed in these staff members could indicate that other factors, such as diet or a genetic predisposition, rather than direct transmission, play a key role in the development of pellagra .

Wilhelm Weinberg's study demonstrated that children of parents who had died of tuberculosis had higher mortality and lower fertility rates compared to those whose parents died of other causes. This early cohort study highlighted the significant health impacts on children linked to parental tuberculosis, suggesting genetic or environmental influences associated with the disease .

Cooperation from the source population and local entities is essential to ensure participant recruitment and retention in cohort studies. The Framingham Heart Study benefited from the supportive local population and healthcare system, which facilitated participant follow-up and data collection. Such cooperation helps mitigate drop-out rates, maintains data quality, and ensures the study's long-term success, leading to robust and reliable findings .

Rate Ratio is a measure used to quantify the relative risk associated with an exposure by comparing the incidence rate of an outcome in the exposed group to that in the nonexposed group. For example, in a study on oral contraceptive estrogen dose, a Rate Ratio of 1.69 indicated that the higher dose was associated with a 69% higher incidence rate of the outcome compared to the lower dose. This statistic is critical in understanding the strength of association between an exposure and a disease .

The Nurses’ Health Study illustrated the importance of assembling occupational cohorts to study specific health outcomes effectively. By recruiting a large cohort of nurses, the study leveraged the shared occupational background to control for certain variables, enhancing the ability to detect associations between health outcomes and lifestyle factors. This approach has provided valuable insights into female health issues and the influence of factors like diet and contraceptive use over time .

You might also like