Understanding Two Sample Problems in Studies
Understanding Two Sample Problems in Studies
Causal inferences
can be drawn
We will begin by considering examples from each freshmen at these two types of schools form the
cell in the above table. First, we will consider units two populations. Independent random samples
that are subjects (distinct individuals). Notice that I of freshmen are selected from each population.
am deliberately not defining the response or, if appli-
cable, treatments. • Lower left: All freshmen at Sun Prairie high
school are selected for study. The students are
• Upper left: The population is all high school divided into treatment groups by randomiza-
freshmen in Wisconsin. A random sample is se- tion.
lected from this population. Once the sample is
obtained, the students are divided into treatment • Lower right: Freshmen at Sun Prairie high
groups by randomization. school are compared to freshmen at Edgewood
high school.
• Upper right: All high schools in Wisconsin are
classified as public or private. The high school Next, I will consider units that are trials.
23
24 CHAPTER 2. THE TWO SAMPLE PROBLEM
• Upper left: A golfer wants to compare two belief that the Clippers and Lakers share an arena; 29
drivers. The trials, individual shots, are as- if I am incorrect). If the researcher selects more than
signed to driver by randomization; we get ran- two levels and the levels are on an interval or ratio
dom samples by assuming a spinner model for scale, then methods of regression might be used.
each driver. After specifying the study factor(s), all other fac-
tors are collectively referred to as background fac-
• Upper right: A golfer has one driver and he tors. I want to mention two ways to handle back-
wants to compare his ability playing at sea ground factors. (This is not an exhaustive list.)
level versus playing at an altitude of 5,000 feet. First, you can block on a background factor. In the
We get random samples by assuming a spinner basketball example, I could block on the opponent.
model at each site. To keep this simple, I will focus on eight games with
• Lower left: Same as upper left, but we no four opponents, as shown below.
longer assume spinner models.
Opponent Home Away H −A
• Lower right: Same as upper right, but we no Dallas 102 97 5
longer assume spinner models. Denver 116 113 3
Toronto 104 102 2
Before we get into inference procedures, formu- Washington 99 100 −1
las for tests and estimation, I want to introduce some Mean 105.25 103.00 2.25
issues of scientific importance. SD 7.46 6.98 2.50
On each unit (subject or trial) we plan to obtain a
response. Typically, the response exhibits some vari- Looking ahead a bit, if the home and away were inde-
ation as we move from unit to unit. (If there is no pendent random samples, then the following analysis
variation, we will not need to do Statistics.) We in- could be appropriate.
vent the notion of factors as the source of the vari-
twos c1 c2;
ation. For example, in its last five basketball games,
pool. %
the Milwaukee Bucks scored 102, 93, 102, 104 and
91 points. Possible factors include strength of oppo-
TWOSAMPLE T FOR C1 VS C2
nent, location of game, and length of time since the
N MEAN SD
previous game. Note the following. These are nat-
C1 4 105.25 7.46
ural factors for a basketball fan to suggest. If you
C2 4 103.00 6.98
know nothing about basketball, then you will be ill
equipped to speculate on the identity of factors. As
95 PCT CI FOR MU C1 - MU C2:
we will see below, if a scientist is bad at suggesting
( -10.2, 14.7)
factors, then he/she will likely not learn much from
collecting and analyzing data.
TTEST MU C1 = MU C2 (VS NE):
From the list of possible factors, the researcher
T= 0.44 P=0.67 DF= 6
chooses one for special status; it is called the study
factor. (In regression and ANOVA, the researcher
POOLED STDEV = 7.22
may choose to have several study factors.) For ex-
ample, for the Bucks I might choose location of game But it is more appropriate to analyze the differences
as my study factor. Next, I must specify the levels of as a one sample problem, as with the following anal-
the study factor. If there are two levels, then we have ysis.
the “two sample problem” which is the title of this
chapter. As a result, I might choose my levels to be ttest c3 %
“home” and “away.” Note that this is not the only
possible choice. I could use the four time zones to TEST OF MU = 0 VS MU N.E. 0
be the levels, or the 28 arenas (if I am correct in my
2.1. OBSERVATIONAL VERSUS RANDOMIZED STUDIES 25
ble reveals the relationship between gender and out- The background factor (job) is statistically related
come. to the study factor (gender) and response (outcome).
Outcome If the background factor fails to be statistically re-
Gender Released Not released Total p̂ lated to either the study factor or response, then
Female 60 40 100 0.60 Simpson’s Paradox will not occur. (This issue will
Male 40 60 100 0.40 be addressed in a future homework assignment.)
Total 100 100 200 If a background factor is strongly (statistically)
related to the response, then you probably want to
Now suppose that the value of a background fac- block on it. If a background factor is strongly (statis-
tor, job type, is available for each person. One could tically) related to the study factor, then it will be dif-
take the data above and stratify it according to job ficult to separate statistically the effect of the study
type, as I have done below. factor from the effect of the background factor.
There is a sampling issue for observational studies
Job A
that I want to address. Years ago I saw a variation
Outcome
of the following example in a really bad introduc-
Gender Released Not released Total p̂
tory Statistics book. Each person in a population of
Female 56 24 80 0.70
college students can be assigned a value on each of
Male 16 4 20 0.80
two dichotomous variables. The first is GPA: high
Total 72 28 100
(A) or low (Ac ); the second is whether the person
Job B smokes tobacco (B) or not (B c ). We can imagine a
Outcome table of population counts (I will follow the notation
Gender Released Not released Total p̂ in Wardrop, Chapter 8).
Female 4 16 20 0.20 B Bc Total
Male 24 56 80 0.30 A NAB NAB c NA
Total 28 72 100 Ac N Ac B N Ac B c N Ac
Note that in the original table, the female release Total NB NB c N
rate is 0.20 larger than the male release rate, but in There are several ways to view this table. You can
each component table (i.e. for each job) the female view it as a single population with two dichotomous
release rate is 0.10 smaller than the male release rate! variables per person (as I have done above). In this
This consistent (across component tables) reversal of case, inference would focus on estimating probabil-
the direction of the relationship is called Simpson’s ities and conditional probabilities. Secondly, you
Paradox. could view smoking status as the response and GPA
We can gain insight into the “why” behind Simp- as the study factor, with levels high and low. This
son’s Paradox by examining the following two ta- means that we have two distinct populations—high
bles. and low GPA. Inference would focus on the propor-
Job tion of smokers in each GPA population. Thirdly, we
Gender A B Total can reverse the roles of smoking and GPA. This gives
Female 80 20 100 two distinct populations—smokers and nonsmokers.
Male 20 80 100 Inference would focus on the proportion of high GPA
Total 100 100 200 in each smoking group. (The bad book thought that
this last perspective was the only one possible and
Outcome compounded its error by suggesting a causal link—
Job Released Not released Total smoking leads to bad grades! One could just as easily
A 72 28 100 argue that anxiety over low grades leads to smoking
B 28 72 100 or that a background factor, time spent partying, is
Total 100 100 100 such that a large amount of time spent partying is
2.1. OBSERVATIONAL VERSUS RANDOMIZED STUDIES 27
0.55, respectively. But now suppose that we pretend 4. English 104, which is taught in 25 sections of
we have independent random samples from the GPA 20 students each.
populations; what happens? Our estimate of smok-
ing given high GPA is 157/445 = 0.353, which is 5. Humanities 105, which is taught in 50 sections
considerably larger than 0.12, the population propor- of 10 students each.
tion. And our estimate of smoking given low GPA is
The college calculates that there are 2500 students in
343/555 = 0.618, which is considerably larger than
the 91 sections offered, for a mean of 27.5 students
0.28, the population proportion.
per section. A rival college reports that for every stu-
As a result, we conclude that it is improper to pre-
dent, the mean class size is 136. Both computations
tend we have a independent random samples from
are correct. What do you think?
the GPA populations. The reason for the strong bias
shown above is simple. By taking equal sample sizes
from each smoking population, we are grossly over- 2.2 Dichotomous response, indepen-
sampling the smokers in the overall population, and
also in the two GPA populations. dent samples
Suppose, however, that I had selected samples of
Later in this chapter we will consider dependent sam-
size 200 from the smokers and 800 from the non-
ples, which arise from pairing.
smokers. (These sample sizes match the proportions
Data from studies of this section can be presented
of smokers and nonsmokers in the overall popula-
in the following manner.
tion.) I did this and obtained the data below.
Smoker? Variable 2
GPA Yes No Total Variable 1 B Bc Total
High 62 434 496 A a b n1
Low 138 366 504 Ac c d n2
Total 200 800 1000 Total m 1 m2 n
If one divides each of the values in the table by 1000, This table is meant to be very general. It can be used
one obtains a very good estimate of the table of pop- for sampling from one population with two dichoto-
ulation proportions. This is a general rule: if we sam- mous responses per unit (remember the GPA and
ple from subpopulations in proportion to occurrence smoking example earlier). This table can be used for
in the overall population, then it is ok to pretend we independent random samples from two populations.
have a random sample from the overall population. Finally, it can be used with a study with randomiza-
Before returning to the two sample problem, I tion. At some point in the analysis I usually view the
want to digress into a common error on sampling. one population, two responses problem as a problem
The point is that it is very important to be careful on conditional probabilities, I will modify the above
about units. table to the following form which I find easier to un-
Suppose that a small college has a freshman class derstand.
of 500 students. Each student enrolls in five courses, Study Response
as detailed below. factor S F Total
1. Social Studies 101, which is taught in one sec- Level 1 a b n1
tions of 500 students. Level 2 c d n2
Total m 1 m2 n
2. Science 102, which is taught in five sections of
100 students each. In order to analyze such data, statisticians typi-
cally begin by arguing that the marginal totals can
3. Math 103, which is taught in 10 sections of 50 (or should) be viewed as fixed numbers. This can be
students each. a bit of a stretch, so some discussion is merited.
2.2. DICHOTOMOUS RESPONSE, INDEPENDENT SAMPLES 29
For the one population, two responses model, only This is an approximate interval, based on using a nor-
the value n is fixed in advance by the researcher; all mal curve approximation. Minitab will not evaluate
other entries in the table are the observed values of this formula for us.
random variables. The statistician then argues that For hypothesis testing, there are several possible
one should perform analysis after conditioning on approaches. The null hypothesis is p 1 = p2 ; there
the other marginal totals. Here is an abridged ver- are three possible alternatives, obtained by replacing
sion of the argument statisticians give. Suppose that ‘=’ in the null hypothesis by >, <, or 6=. An exact
we have the following marginal totals. P-value can be obtained by using the hypergeomet-
ric distribution. This distribution is not in Minitab.
Variable 2 (See me if you want a macro for Version 9.) Ap-
Variable 1 B Bc Total proximate probabilities can be obtained by a normal
A a b 60 or chi-squared approximation. The normal approxi-
Ac c d 40 mation can be written two ways. First, as z = x/σ,
Total 70 30 100 where
s
Only the total number of units, n = 100, is fixed m1 m2
x = p̂1 − p̂2 and σ = .
by the sampling plan. But what do we learn from n1 n2 (n − 1)
the other totals? Well, we get evidence that B is
much more common than B c , and evidence that A This expression can be rewritten as
is somewhat more common than Ac . But we don’t √
n − 1(ad − bc)
learn anything about a relationship between A and z= √ .
n1 n2 m1 m2
B, which is, after all, the primary purpose of the in-
vestigation. Note that with the above margins, we Some people modify z slightly and use
could have a = 60 which would provide evidence of √
a very strong positive association between A and B; 0 n(ad − bc)
z =√ .
or we could have a = 30 which would provide evi- n 1 n2 m1 m2
dence of a very strong negative association between
Finally, others use
A and B; or we could have a = 42 which would pro-
vide no evidence of an association between A and B. χ2 = (z 0 )2 .
In short, knowledge of the marginal totals does not
provide the researcher with evidence of the strength If z or z 0 is used, P-values are obtained from the stan-
or direction of association between A and B; hence, dard normal curve. Literally, χ2 can be used only for
it probably won’t hurt to condition on the margins. the alternative 6= and the P-value is obtained by using
Plus there is the added bonus that conditioning on the chi-squared curve with one degree of freedom.
the margins makes the math much easier. Minitab presents the analysis only for χ 2 .
In the table below, define p̂1 = a/n1 , q̂1 = b/n1 , Exercise 2 on page 252 of Wardrop presents the
p̂2 = c/n2 , and q̂2 = d/n2 . following data.
set c2
301 315 316 317 321 321 323 TWOSAMPLE T FOR C1 VS C2
3(327) N MEAN STDEV
end C1 10 333.00 8.18
desc c1 c2 C2 10 319.50 7.93
Assuming normal pdfs, in this case, does not solve TWOSAMPLE T FOR C11 VS C12
the problem. With normal pdfs, the sampling distri- N MEAN STDEV
bution of W2 can be approximated by, but does not C11 10 50.00 9.40
equal, a t distribution. C12 20 45.00 9.40
There are different opinions about the degrees of
freedom in the approximating distribution. Minitab 95 PCT CI FOR MU C11 - MU C12:
uses a horrendously messy formula for the degrees ( -2.7, 12.7)
of freedom, but since we don’t need to evaluate it,
the fact that it is horrendous is no problem. (See for- TTEST MU C11 = MU C12 (VS NE):
mulas 16.6 and 16.7 on page 592 of Wardrop if you T= 1.37 P=0.19 DF= 18
want to see it!) The above data will be reanalyzed
under this second situation. twos c11 c12 %
strange!
The same phenomenon occurs without pooling, as C3 N = 4 Median = 10.5
shown below. In this case, the lower bound decreases C4 N = 4 Median = 8.0
from 5.9 to 5.7. Point estimate for
ETA1-ETA2 is 2.5
twos c1 c2
97.0 pct c.i. for
TWOSAMPLE T FOR C1 VS C2 ETA1-ETA2
N MEAN STDEV is (-5.000,11.000)
C1 10 334.0 10.4
C2 10 319.50 7.93 W = 21.5
Test of ETA1 = ETA2
95 PCT CI FOR MU C1 - MU C2: vs. ETA1 n.e. ETA2
( 5.7, 23.3) is significant at 0.3865
The test is significant at 0.3836
TTEST MU C1 = MU C2 (VS NE): (adjusted for ties)
T= 3.51 P=0.0029 DF= 16
The data are combined into one set and sorted, and Point estimate for
ranks are assigned to the overall data, as below. ETA1-ETA2 is 13.0
stest c1 % Now, let’s pair the sorted data, matching the smallest
values in each set, the second smallest values, and
SIGN TEST OF MED = 0 VS N.E. 0 so on. Then, after pairing, subtract the values in the
N BELOW EQUAL ABOVE second sample from the corresponding values in the
C1 25 2 3 20 first sample. The 10 differences are below.
P-VALUE MEDIAN Differences of Sorted Data
0.0001 1.500
0.12 −0.61 0.17 0.61 0.43
wint c1 %
EST 0.41 −0.19 −0.10 0.38 −0.48
N MED CONF CONF INT
C1 25 1.50 95.0 (1.0, 2.0) I constructed two 95% confidence intervals for the
difference of the means. Using the two independent
wtest c1 % samples, pooled estimate of variance (an appropriate
analysis), I obtained [−0.85, 1.00]. Using the dif-
TEST OF MED = 0 VS MED N.E. 0 ferences of the sorted data, I obtained [−0.22, 0.37].
N FOR WILC EST We will see that this latter analysis is incorrect. At
N TEST STAT P-VALUE MED this point, however, the latter analysis looks supe-
C1 25 22 220.5 0.002 1.5
rior: both intervals are correct (they contain 0) and
the second interval is much more precise.
36 CHAPTER 2. THE TWO SAMPLE PROBLEM