0% found this document useful (0 votes)
12 views14 pages

Understanding Two Sample Problems in Studies

The document discusses the two sample problem in observational and randomized studies. It defines the key concepts of study factors, levels of study factors, responses, and background factors. For a two level study factor, we have a "two sample problem". The document provides examples of studies that differ based on whether the units are allocated to groups by randomization or not, and whether the units are selected at random. It discusses approaches for handling background factors like blocking and controlling.

Uploaded by

jamesyu
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views14 pages

Understanding Two Sample Problems in Studies

The document discusses the two sample problem in observational and randomized studies. It defines the key concepts of study factors, levels of study factors, responses, and background factors. For a two level study factor, we have a "two sample problem". The document provides examples of studies that differ based on whether the units are allocated to groups by randomization or not, and whether the units are selected at random. It discusses approaches for handling background factors like blocking and controlling.

Uploaded by

jamesyu
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 2

The Two Sample Problem

2.1 Observational Versus Randomized Studies


The table below is taken from The Statistical Sleuth by Ramsey and Schafer. We will discuss it in class.
Allocation of Units to Groups

Selection of Units By Randomization Not by Randomization


At Random A random sample is Random samples are Inferences to
selected from one selected from existing the populations
population; units distinct populations. can be drawn.
are then randomly
assigned to different
treatment groups

Not at Random A group of study Collections of


units is found available units from
units are then distinct groups are
randomly assigned examined.
to treatment groups.

Causal inferences
can be drawn

We will begin by considering examples from each freshmen at these two types of schools form the
cell in the above table. First, we will consider units two populations. Independent random samples
that are subjects (distinct individuals). Notice that I of freshmen are selected from each population.
am deliberately not defining the response or, if appli-
cable, treatments. • Lower left: All freshmen at Sun Prairie high
school are selected for study. The students are
• Upper left: The population is all high school divided into treatment groups by randomiza-
freshmen in Wisconsin. A random sample is se- tion.
lected from this population. Once the sample is
obtained, the students are divided into treatment • Lower right: Freshmen at Sun Prairie high
groups by randomization. school are compared to freshmen at Edgewood
high school.
• Upper right: All high schools in Wisconsin are
classified as public or private. The high school Next, I will consider units that are trials.

23
24 CHAPTER 2. THE TWO SAMPLE PROBLEM

• Upper left: A golfer wants to compare two belief that the Clippers and Lakers share an arena; 29
drivers. The trials, individual shots, are as- if I am incorrect). If the researcher selects more than
signed to driver by randomization; we get ran- two levels and the levels are on an interval or ratio
dom samples by assuming a spinner model for scale, then methods of regression might be used.
each driver. After specifying the study factor(s), all other fac-
tors are collectively referred to as background fac-
• Upper right: A golfer has one driver and he tors. I want to mention two ways to handle back-
wants to compare his ability playing at sea ground factors. (This is not an exhaustive list.)
level versus playing at an altitude of 5,000 feet. First, you can block on a background factor. In the
We get random samples by assuming a spinner basketball example, I could block on the opponent.
model at each site. To keep this simple, I will focus on eight games with
• Lower left: Same as upper left, but we no four opponents, as shown below.
longer assume spinner models.
Opponent Home Away H −A
• Lower right: Same as upper right, but we no Dallas 102 97 5
longer assume spinner models. Denver 116 113 3
Toronto 104 102 2
Before we get into inference procedures, formu- Washington 99 100 −1
las for tests and estimation, I want to introduce some Mean 105.25 103.00 2.25
issues of scientific importance. SD 7.46 6.98 2.50
On each unit (subject or trial) we plan to obtain a
response. Typically, the response exhibits some vari- Looking ahead a bit, if the home and away were inde-
ation as we move from unit to unit. (If there is no pendent random samples, then the following analysis
variation, we will not need to do Statistics.) We in- could be appropriate.
vent the notion of factors as the source of the vari-
twos c1 c2;
ation. For example, in its last five basketball games,
pool. %
the Milwaukee Bucks scored 102, 93, 102, 104 and
91 points. Possible factors include strength of oppo-
TWOSAMPLE T FOR C1 VS C2
nent, location of game, and length of time since the
N MEAN SD
previous game. Note the following. These are nat-
C1 4 105.25 7.46
ural factors for a basketball fan to suggest. If you
C2 4 103.00 6.98
know nothing about basketball, then you will be ill
equipped to speculate on the identity of factors. As
95 PCT CI FOR MU C1 - MU C2:
we will see below, if a scientist is bad at suggesting
( -10.2, 14.7)
factors, then he/she will likely not learn much from
collecting and analyzing data.
TTEST MU C1 = MU C2 (VS NE):
From the list of possible factors, the researcher
T= 0.44 P=0.67 DF= 6
chooses one for special status; it is called the study
factor. (In regression and ANOVA, the researcher
POOLED STDEV = 7.22
may choose to have several study factors.) For ex-
ample, for the Bucks I might choose location of game But it is more appropriate to analyze the differences
as my study factor. Next, I must specify the levels of as a one sample problem, as with the following anal-
the study factor. If there are two levels, then we have ysis.
the “two sample problem” which is the title of this
chapter. As a result, I might choose my levels to be ttest c3 %
“home” and “away.” Note that this is not the only
possible choice. I could use the four time zones to TEST OF MU = 0 VS MU N.E. 0
be the levels, or the 28 arenas (if I am correct in my
2.1. OBSERVATIONAL VERSUS RANDOMIZED STUDIES 25

N MEAN SD T P VALUE It is my opinion that novice statisticians frequently


C3 4 2.25 2.50 1.80 0.17 are overly optimistic on the value of blocking. Un-
less you are pretty certain that the factor has a big im-
tint c3 % pact on the response, it is usually better not to block.
A second way to deal with a background factor is
N MEAN SD 95 % C.I. by controlling for it. This means that you keep the
C3 4 2.25 2.50 (-1.73,6.23) value of the factor constant throughout the study, or
at least for the data you analyze. In a study of her
Notice that for the differences, the P-value is much cat’s consumption of two flavors of treats, Dawn (a
smaller and the confidence interval is much narrower. former student) controlled the cat’s intake of other
Blocking is not always effective. For example, the food, and tried to control his activity level. In addi-
Bucks’ data above was selected deliberately for illus- tion, she presented either treat at the same time each
tration only. There are 14 teams that the Bucks have day, in an attempt to control for time of day effect as
played home and away thus far. The analysis of all well as the cat’s general level of hunger.
the data is very different than that above. After blocking and controlling, if either or both of
twos c1 c2; these are used, there are still lots of background fac-
pool. % tors. If units are assigned to study factor by random-
ization, there is some reason to believe that the ef-
TWOSAMPLE T FOR C1 VS C2 fects of these background factors will be “balanced”
N MEAN SD between the levels of the study factor (this notion can
C1 14 94.60 10.20 be made more precise). But if units are associated
C2 14 101.71 7.16 with the level of a study factor, it is very possible that
the background factors will severely bias the study.
95 PCT CI FOR MU C1 - MU C2: Some examples follow.
( -13.9, -0.2)
1. Yesterday I heard a talk on the effects of “coach-
ing” for the SAT. The subjects are students. The
TTEST MU C1 = MU C2 (VS NE):
response is change in verbal score from PSAT
T= -2.12 P=0.043 DF= 26
to SAT. The study factor is coaching with levels
“yes” and “no.” What are some possible back-
POOLED STDEV = 8.81
ground factors?
The analysis with blocking is below.
2. The subjects are people. The response is
ttest c3 % whether the person develops a particular disease
of interest. The study factor is smoking, with
TEST OF MU = 0 VS MU N.E. 0 levels “yes” and “no.” What are some possible
background factors?
N MEAN SD T P VALUE
C3 14 -7.07 13.04 -2.03 0.063 3. The subjects are men. The response is whether
the man develops a particular disease of inter-
tint c3 % est. The study factor whether the man has had
a vasectomy, with levels “yes” and “no.” What
N MEAN SD 95 % C.I. are some possible background factors?
C3 14 -7.07 13.04 (-14.60, 0.46)
The following hypothetical example illustrates the
Note from above that the independent samples analy- possible effect of a background factor.
sis compared to the block (paired data) analysis gives A company with 200 employees decides it must
a smaller P-value and a narrower confidence interval. reduce its work force by one-half. The following ta-
26 CHAPTER 2. THE TWO SAMPLE PROBLEM

ble reveals the relationship between gender and out- The background factor (job) is statistically related
come. to the study factor (gender) and response (outcome).
Outcome If the background factor fails to be statistically re-
Gender Released Not released Total p̂ lated to either the study factor or response, then
Female 60 40 100 0.60 Simpson’s Paradox will not occur. (This issue will
Male 40 60 100 0.40 be addressed in a future homework assignment.)
Total 100 100 200 If a background factor is strongly (statistically)
related to the response, then you probably want to
Now suppose that the value of a background fac- block on it. If a background factor is strongly (statis-
tor, job type, is available for each person. One could tically) related to the study factor, then it will be dif-
take the data above and stratify it according to job ficult to separate statistically the effect of the study
type, as I have done below. factor from the effect of the background factor.
There is a sampling issue for observational studies
Job A
that I want to address. Years ago I saw a variation
Outcome
of the following example in a really bad introduc-
Gender Released Not released Total p̂
tory Statistics book. Each person in a population of
Female 56 24 80 0.70
college students can be assigned a value on each of
Male 16 4 20 0.80
two dichotomous variables. The first is GPA: high
Total 72 28 100
(A) or low (Ac ); the second is whether the person
Job B smokes tobacco (B) or not (B c ). We can imagine a
Outcome table of population counts (I will follow the notation
Gender Released Not released Total p̂ in Wardrop, Chapter 8).
Female 4 16 20 0.20 B Bc Total
Male 24 56 80 0.30 A NAB NAB c NA
Total 28 72 100 Ac N Ac B N Ac B c N Ac
Note that in the original table, the female release Total NB NB c N
rate is 0.20 larger than the male release rate, but in There are several ways to view this table. You can
each component table (i.e. for each job) the female view it as a single population with two dichotomous
release rate is 0.10 smaller than the male release rate! variables per person (as I have done above). In this
This consistent (across component tables) reversal of case, inference would focus on estimating probabil-
the direction of the relationship is called Simpson’s ities and conditional probabilities. Secondly, you
Paradox. could view smoking status as the response and GPA
We can gain insight into the “why” behind Simp- as the study factor, with levels high and low. This
son’s Paradox by examining the following two ta- means that we have two distinct populations—high
bles. and low GPA. Inference would focus on the propor-
Job tion of smokers in each GPA population. Thirdly, we
Gender A B Total can reverse the roles of smoking and GPA. This gives
Female 80 20 100 two distinct populations—smokers and nonsmokers.
Male 20 80 100 Inference would focus on the proportion of high GPA
Total 100 100 200 in each smoking group. (The bad book thought that
this last perspective was the only one possible and
Outcome compounded its error by suggesting a causal link—
Job Released Not released Total smoking leads to bad grades! One could just as easily
A 72 28 100 argue that anxiety over low grades leads to smoking
B 28 72 100 or that a background factor, time spent partying, is
Total 100 100 100 such that a large amount of time spent partying is
2.1. OBSERVATIONAL VERSUS RANDOMIZED STUDIES 27

linked to smoking and low grades.) Smoker?


A critical point that is often overlooked is the im- GPA Yes No Total
portance of how a sample is selected. Let us imagine High 60 437 497
three possible sampling schemes. We can take a ran- Low 137 366 503
dom sample from the overall population of college Total 197 803 1000
students; we can take independent random samples
Next, I used these data to estimate the three tables
from the populations of smokers and nonsmokers; or
above; the population proportions and the two tables
we can take independent random samples from the
of conditional probabilities. The results are below.
populations of high and low GPA. Suppose that the
population counts are given by the following table. Smoker?
GPA Yes No Total
Smoker? High 0.060 0.437 0.497
GPA Yes No Total Low 0.137 0.366 0.503
High 600 4400 5000 Total 0.197 0.803 1.000
Low 1400 3600 5000
Total 2000 8000 10000 Smoker?
GPA Yes No Total
The table of population proportions is below.
High 0.121 0.879 1.000
Smoker? Low 0.272 0.728 1.000
GPA Yes No Total
High 0.06 0.44 0.50 Smoker?
Low 0.14 0.36 0.50 GPA Yes No
Total 0.20 0.80 1.00 High 0.305 0.544
Low 0.695 0.456
The table of conditional probabilities of smoking Total 1.000 1.000
status given GPA is below.
By inspection, all estimates are quite close to the
Smoker? population proportions. As the comparisons based
GPA Yes No Total on the last two tables suggest, if we have a random
High 0.12 0.88 1.00 sample from the overall population, it is valid to pre-
Low 0.28 0.72 1.00 tend we have either: (a) independent random sam-
ples from the high and low GPA populations, or (b)
Note that 0.12 and 0.28 are p1 and p2 for the perspec- independent random samples from the smoking and
tive of smoking being the response. nonsmoking populations.
The table of conditional probabilities of GPA Second, suppose that I select independent random
given smoking status is below. samples (with replacement) of size 500 each from the
Smoker? smoking and nonsmoking populations. I did this on
GPA Yes No my computer and obtained the results shown below.
High 0.30 0.55 Smoker?
Low 0.70 0.45 GPA Yes No Total
Total 1.00 1.00 High 157 288 445
Low 343 212 555
Note that 0.30 and 0.55 are p1 and p2 for the perspec-
Total 500 500 1000
tive of GPA being the response.
I will consider three ways to sample. First, sup- The estimates of: high GPA for smokers is
pose we select a random sample (with replacement) 157/500 = 0.314, and high GPA for nonsmokers
of size 1000 from the overall population. I did this is 288/500 = 0.576. These numbers are reason-
on my computer and obtained the data below. ably close to the population proportions, 0.30 and
28 CHAPTER 2. THE TWO SAMPLE PROBLEM

0.55, respectively. But now suppose that we pretend 4. English 104, which is taught in 25 sections of
we have independent random samples from the GPA 20 students each.
populations; what happens? Our estimate of smok-
ing given high GPA is 157/445 = 0.353, which is 5. Humanities 105, which is taught in 50 sections
considerably larger than 0.12, the population propor- of 10 students each.
tion. And our estimate of smoking given low GPA is
The college calculates that there are 2500 students in
343/555 = 0.618, which is considerably larger than
the 91 sections offered, for a mean of 27.5 students
0.28, the population proportion.
per section. A rival college reports that for every stu-
As a result, we conclude that it is improper to pre-
dent, the mean class size is 136. Both computations
tend we have a independent random samples from
are correct. What do you think?
the GPA populations. The reason for the strong bias
shown above is simple. By taking equal sample sizes
from each smoking population, we are grossly over- 2.2 Dichotomous response, indepen-
sampling the smokers in the overall population, and
also in the two GPA populations. dent samples
Suppose, however, that I had selected samples of
Later in this chapter we will consider dependent sam-
size 200 from the smokers and 800 from the non-
ples, which arise from pairing.
smokers. (These sample sizes match the proportions
Data from studies of this section can be presented
of smokers and nonsmokers in the overall popula-
in the following manner.
tion.) I did this and obtained the data below.

Smoker? Variable 2
GPA Yes No Total Variable 1 B Bc Total
High 62 434 496 A a b n1
Low 138 366 504 Ac c d n2
Total 200 800 1000 Total m 1 m2 n

If one divides each of the values in the table by 1000, This table is meant to be very general. It can be used
one obtains a very good estimate of the table of pop- for sampling from one population with two dichoto-
ulation proportions. This is a general rule: if we sam- mous responses per unit (remember the GPA and
ple from subpopulations in proportion to occurrence smoking example earlier). This table can be used for
in the overall population, then it is ok to pretend we independent random samples from two populations.
have a random sample from the overall population. Finally, it can be used with a study with randomiza-
Before returning to the two sample problem, I tion. At some point in the analysis I usually view the
want to digress into a common error on sampling. one population, two responses problem as a problem
The point is that it is very important to be careful on conditional probabilities, I will modify the above
about units. table to the following form which I find easier to un-
Suppose that a small college has a freshman class derstand.
of 500 students. Each student enrolls in five courses, Study Response
as detailed below. factor S F Total
1. Social Studies 101, which is taught in one sec- Level 1 a b n1
tions of 500 students. Level 2 c d n2
Total m 1 m2 n
2. Science 102, which is taught in five sections of
100 students each. In order to analyze such data, statisticians typi-
cally begin by arguing that the marginal totals can
3. Math 103, which is taught in 10 sections of 50 (or should) be viewed as fixed numbers. This can be
students each. a bit of a stretch, so some discussion is merited.
2.2. DICHOTOMOUS RESPONSE, INDEPENDENT SAMPLES 29

For the one population, two responses model, only This is an approximate interval, based on using a nor-
the value n is fixed in advance by the researcher; all mal curve approximation. Minitab will not evaluate
other entries in the table are the observed values of this formula for us.
random variables. The statistician then argues that For hypothesis testing, there are several possible
one should perform analysis after conditioning on approaches. The null hypothesis is p 1 = p2 ; there
the other marginal totals. Here is an abridged ver- are three possible alternatives, obtained by replacing
sion of the argument statisticians give. Suppose that ‘=’ in the null hypothesis by >, <, or 6=. An exact
we have the following marginal totals. P-value can be obtained by using the hypergeomet-
ric distribution. This distribution is not in Minitab.
Variable 2 (See me if you want a macro for Version 9.) Ap-
Variable 1 B Bc Total proximate probabilities can be obtained by a normal
A a b 60 or chi-squared approximation. The normal approxi-
Ac c d 40 mation can be written two ways. First, as z = x/σ,
Total 70 30 100 where
s
Only the total number of units, n = 100, is fixed m1 m2
x = p̂1 − p̂2 and σ = .
by the sampling plan. But what do we learn from n1 n2 (n − 1)
the other totals? Well, we get evidence that B is
much more common than B c , and evidence that A This expression can be rewritten as
is somewhat more common than Ac . But we don’t √
n − 1(ad − bc)
learn anything about a relationship between A and z= √ .
n1 n2 m1 m2
B, which is, after all, the primary purpose of the in-
vestigation. Note that with the above margins, we Some people modify z slightly and use
could have a = 60 which would provide evidence of √
a very strong positive association between A and B; 0 n(ad − bc)
z =√ .
or we could have a = 30 which would provide evi- n 1 n2 m1 m2
dence of a very strong negative association between
Finally, others use
A and B; or we could have a = 42 which would pro-
vide no evidence of an association between A and B. χ2 = (z 0 )2 .
In short, knowledge of the marginal totals does not
provide the researcher with evidence of the strength If z or z 0 is used, P-values are obtained from the stan-
or direction of association between A and B; hence, dard normal curve. Literally, χ2 can be used only for
it probably won’t hurt to condition on the margins. the alternative 6= and the P-value is obtained by using
Plus there is the added bonus that conditioning on the chi-squared curve with one degree of freedom.
the margins makes the math much easier. Minitab presents the analysis only for χ 2 .
In the table below, define p̂1 = a/n1 , q̂1 = b/n1 , Exercise 2 on page 252 of Wardrop presents the
p̂2 = c/n2 , and q̂2 = d/n2 . following data.

Study Response Study Response


factor S F Total factor S F Total
Level 1 a b n1 Level 1 46 66 112
Level 2 c d n2 Level 2 30 99 129
Total m 1 m2 n Total 76 165 241
Below is a Minitab analysis of these data.
The confidence interval for p1 − p2 is
s read c1 c2
p̂1 q̂1 p̂2 q̂2 46 66
p̂1 − p̂2 ± z + . (2.1)
n1 n2 30 99
30 CHAPTER 2. THE TWO SAMPLE PROBLEM

end are denoted


chis c1 c2 %
y1,1 , y1,2 , y1,3 , . . . , y1,n1 ,
Expected counts are printed
and are summarized by their mean ȳ1· and standard
below observed counts
deviation s1 . Similarly, data from the second popu-
lation are denoted
C1 C2 Total
1 46 66 112 y2,1 , y2,2 , y2,3 , . . . , y2,n2 ,
35.3 76.7
and are summarized by their mean ȳ2· and standard
2 30 99 129 deviation s2 . (Most authors suppress the comma in
40.7 88.3 the subscript, and I might forget and do that too on
Total 76 165 241 occasion. I like the commas for later work when we
need to know whether, for example y111 is the 11th
ChiSq = 3.230 + 1.488 + observation from the first sample or the first observa-
2.804 + 1.292 = 8.813 tion from the 11th sample.)
df = 1 Attention focuses on the difference µ 1 −µ2 , which
is estimated by ȳ1· − ȳ2· . For inference, we will need
cdf 8.813; the sampling distribution of this estimate. Some ba-
chis 1. sic results in mathematical statistics indicate that the
8.8130 0.9970 standardized form of the estimate is
subt 0.9970 1 k1
(Ȳ1· − Ȳ2· ) − (µ1 − µ2 )
ANSWER = 0.0030 W = r .
σ12 σ22
n1 + n2
let k2=sqrt(8.813)
The mathematical problem arises in trying to deal
cdf k2 k3 with the unknown σ’s in the denominator.
subt k3 1 k4 The first approach is to assume that they are equal
ANSWER = 0.0015 and to estimate them by sp , where

(n1 − 1)s21 + (n2 − 1)s22


s2p = .
n1 + n 2 − 2
2.3 Numerical response, indepen-
Next, we substitute this estimate into W to yield W 1 ,
dent samples
(Ȳ1· − Ȳ2· ) − (µ1 − µ2 )
W1 = .
I am “jumping over” multicategory responses, or-
p
sp 1/n1 + 1/n2
dered or not, and proceeding to numerical. We will
return to multicategory responses soon. If one assumes that the two populations are normal
pdfs, then probabilities for W1 can be obtained from
The most commonly used procedures compare the
the t distribution with (n1 + n2 − 2) degrees of free-
populations by comparing their means. For refer-
dom.
ence, see 16.1 and 16.2 of Wardrop. In fact, the pre-
The following example is taken from exercise 1 on
sentation below is a compression of the ideas in 16.2.
page 401 of Wardrop.
We assume that we have independent random sam-
ples from two populations. The first population has set c1
(unknown) mean µ1 and standard deviation σ1 . The 321 323 329 330 331 332 337 337
second population has (unknown) mean µ 2 and stan- 343 347
dard deviation σ2 . The data from the first population end
2.3. NUMERICAL RESPONSE, INDEPENDENT SAMPLES 31

set c2
301 315 316 317 321 321 323 TWOSAMPLE T FOR C1 VS C2
3(327) N MEAN STDEV
end C1 10 333.00 8.18
desc c1 c2 C2 10 319.50 7.93

N MEAN MEDIAN STDEV 95 PCT CI FOR MU C1 - MU C2:


C1 10 333.0 331.50 8.18 ( 5.9, 21.1)
C2 10 319.5 321.00 7.93
TTEST MU C1 = MU C2 (VS NE):
twos c1 c2; T= 3.75 P=0.0016 DF= 17
pool.
The only difference in the two analyses is that the
TWOSAMPLE T FOR C1 VS C2 latter has 17 degrees of freedom, while the former
N MEAN STDEV has 18. The values of T are identical because for a
C1 10 333.0 8.18 balanced study W1 = W2 .
C2 10 319.5 7.93 It is instructive to consider some artificial data.
95 PCT CI FOR MU C1 - MU C2: twos c11 c12 %
( 5.9, 21.1)
TWOSAMPLE T FOR C11 VS C12
TTEST MU C1 = MU C2 (VS NE): N MEAN STDEV
T= 3.75 P=0.0015 DF= 18 C11 10 50.00 9.40
C12 10 45.00 9.40
POOLED STDEV = 8.06
95 PCT CI FOR MU C11 - MU C12:
The second approach is to make no assumption
( -3.8, 13.8)
about the two standard deviations; simply estimate
each population standard deviation by its corre-
TTEST MU C11 = MU C12 (VS NE):
sponding sample standard deviation. Making this
T= 1.19 P=0.25 DF= 18
change to W , we get W2 ,
Contrast this output with the following four.
(Ȳ1· − Ȳ2· ) − (µ1 − µ2 )
W2 = q .
s21 /n1 + s22 /n2 twos c11 c12 %

Assuming normal pdfs, in this case, does not solve TWOSAMPLE T FOR C11 VS C12
the problem. With normal pdfs, the sampling distri- N MEAN STDEV
bution of W2 can be approximated by, but does not C11 10 50.00 9.40
equal, a t distribution. C12 20 45.00 9.40
There are different opinions about the degrees of
freedom in the approximating distribution. Minitab 95 PCT CI FOR MU C11 - MU C12:
uses a horrendously messy formula for the degrees ( -2.7, 12.7)
of freedom, but since we don’t need to evaluate it,
the fact that it is horrendous is no problem. (See for- TTEST MU C11 = MU C12 (VS NE):
mulas 16.6 and 16.7 on page 592 of Wardrop if you T= 1.37 P=0.19 DF= 18
want to see it!) The above data will be reanalyzed
under this second situation. twos c11 c12 %

twos c1 c2 % TWOSAMPLE T FOR C11 VS C12


32 CHAPTER 2. THE TWO SAMPLE PROBLEM

N MEAN STDEV if each sample size is 30 (or more), then we know


C11 10 50.00 1.00 that we have at least 29 (or more) d.f. As a result, we
C12 10 45 100 might be willing to use the standard normal curve in-
stead of the t curve.
95 PCT CI FOR MU C11 - MU C12: I want to explore the issue of robustness for the
( -66.55, 77) above procedures. I performed a simulation study
with 1000 runs. For each run I selected independent
TTEST MU C11 = MU C12 (VS NE): random samples with n1 = n2 = 10 from exponen-
T= 0.16 P=0.88 DF= 9 tial (1) pdfs. For each random sample I calculated
two 95% confidence intervals for µ1 − µ2 ; one with
twos c11 c12 % pooling and one without. The results were virtually
identical and very close to what one would expect
TWOSAMPLE T FOR C11 VS C12 for normal pdfs. In particular, when pooling, 42 in-
N MEAN STDEV tervals were incorrect (4.2%) and the mean width of
C11 10 50.00 1.00 the intervals is 1.800. When not pooling, 38 inter-
C12 20 45 100 vals were incorrect (3.8%) and the mean width of the
intervals is 1.838.
95 PCT CI FOR MU C11 - MU C12: I now want to address a strange property of the
( -41.83, 52) above procedures. Recall the earlier data from
page 401 of Wardrop. Let us now suppose that the
TTEST MU C11 = MU C12 (VS NE): largest observation from the first population, 347, is
T= 0.22 P=0.83 DF= 19 replaced by 357. This increases the mean of the first
sample by one and clearly has no effect on the sec-
twos c11 c12 ond sample. Thus, we have evidence that µ 1 is even
larger (compared to the evidence in the original data)
TWOSAMPLE T FOR C11 VS C12 and no evidence about µ2 . Thus, it seems “logical”
N MEAN STDEV that our estimate of µ1 − µ2 should “increase,” and
C11 10 50 100 certainly not decrease. But look at the analysis be-
C12 20 45.00 1.00 low.

95 PCT CI FOR MU C11 - MU C12: twos c1 c2;


( -67, 76.55) pool.

TTEST MU C11 = MU C12 (VS NE): TWOSAMPLE T FOR C1 VS C2


T= 0.16 P=0.88 DF= 9 N MEAN STDEV
C1 10 334.0 10.4
I want to remark on a dumb, but increasingly pop- C2 10 319.50 7.93
ular, approach. The suggestion is to use the t dis-
tribution with r − 1 degrees of freedom, where r is 95 PCT CI FOR MU C1 - MU C2:
the minimum of n1 and n2 . The main virtue of this ( 5.8, 23.2)
method is that we avoid having to calculate the d.f.
with the horrendous formula; but if it is done by com- TTEST MU C1 = MU C2 (VS NE):
puter, what is the problem? T= 3.51 P=0.0025 DF= 18
Finally, if n1 and n2 are both large and you must
analyze the data by hand, you might as well use the POOLED STDEV = 9.25
standard normal curve for reference instead of both-
ering with calculating the degrees of freedom. By Earlier, the lower bound for the confidence interval
the “minimum” approach in the previous paragraph, was 5.9; now it has decreased to 5.8! This is very
2.3. NUMERICAL RESPONSE, INDEPENDENT SAMPLES 33

strange!
The same phenomenon occurs without pooling, as C3 N = 4 Median = 10.5
shown below. In this case, the lower bound decreases C4 N = 4 Median = 8.0
from 5.9 to 5.7. Point estimate for
ETA1-ETA2 is 2.5
twos c1 c2
97.0 pct c.i. for
TWOSAMPLE T FOR C1 VS C2 ETA1-ETA2
N MEAN STDEV is (-5.000,11.000)
C1 10 334.0 10.4
C2 10 319.50 7.93 W = 21.5
Test of ETA1 = ETA2
95 PCT CI FOR MU C1 - MU C2: vs. ETA1 n.e. ETA2
( 5.7, 23.3) is significant at 0.3865
The test is significant at 0.3836
TTEST MU C1 = MU C2 (VS NE): (adjusted for ties)
T= 3.51 P=0.0029 DF= 16

The Mann-Whitney-Wilcoxin procedure is an al- Cannot reject at alpha = 0.05


ternative to the above procedures. It assumes that I also ran this command for the earlier data in c1
the pdfs differ in a shift; see the picture in class. and c2.
Mann-Whitney (Wilcoxin is usually suppressed to
avoid confusion with the one-sample procedure) is mann c1 c2 %
a generalization of the normal case with equal stan-
dard deviations. (Discuss.) Mann-Whitney Confidence
The idea behind Mann-Whitney will be illustrated Interval and Test
with a small set of artificial data.
C1 N = 10 Median = 331.5
Sample 1: 8 9 12 15 C2 N = 10 Median = 321.0
Sample 2: 4 7 9 13

The data are combined into one set and sorted, and Point estimate for
ranks are assigned to the overall data, as below. ETA1-ETA2 is 13.0

Data: 4 7 8 9 95.5 pct c.i. for


Ranks: 1 2 3 4.5 ETA1-ETA2 is (5.00,21.00)
Data: 9 12 13 15
Ranks: 4.5 6 7 8 W = 146.5
Note that tied values are given mean ranks. The test
Test of ETA1 = ETA2 vs.
statistic is the sum of the ranks of the data in the first
ETA1 n.e. ETA2 is significant
sample; for these data it is W = 3 + 4.5 + 6 + 8 =
at 0.0019
21.5
The test is significant at
I put the above data into c3 and c4 and ran the
0.0019 (adjusted for ties)
following Minitab command.
I replaced the largest observation in the first sam-
mann c3 c4 % ple, 347, by 999. The analysis is below.

Mann-Whitney Confidence mann c1 c2 %


Interval and Test
34 CHAPTER 2. THE TWO SAMPLE PROBLEM

Mann-Whitney Confidence is called a cross-over design, which I don’t plan to


Interval and Test cover in these notes.
Below are some examples of matching similar
C1 N = 10 Median = 331.5 units.
C2 N = 10 Median = 321.0
1. Sixty students are available for a comparison
Point estimate for of two teaching materials. Students are paired
ETA1-ETA2 is 13.0 based on some criterion (IQ, background in
area, GPA, etc.). In each of the 30 pairs students
95.5 pct c.i. for are assigned to material by randomization.
ETA1-ETA2 is (5.0,22.0) 2. This example is invalid, as demonstrated later
in these notes, but is popular and is advocated
W = 146.5 in some introductory texts. Two classes have 30
students each. Class 1 will use teaching mate-
Test of ETA1 = ETA2 vs. rial A, and class 2 will use teaching material B.
ETA1 n.e. ETA2 is significant Students are paired across classes (i.e. each stu-
at 0.0019 dent in Class 1 is paired with a student in Class
The test is significant at 2). This is invalid.
0.0019 (adjusted for ties)
Matching similar units is valid only if there is ran-
domization.
2.4 Paired data
2.4.1 Dichotomous response
Paired data arises in two ways:
Read Section 8.5 of Wardrop.
• Subdividing (or reusing) units, or

• Matching similar units. 2.4.2 Numerical response


The standard approach is to calculate differences
Below are some examples of subdividing units.
and then use a one sample procedure. Page 405 of
1. The classic “before and after” studies, in which Wardrop presents data on 25-yard backstroke and
a response is obtained before and after some breaststroke times. Below are the first ten pairs; see
event (diet, exercise, training, etc.). Note that Wardrop for complete listing of data.
these studies are observational; i.e. there is no Pair: 1 2 3 4 5
randomization. Bk. 40.0 39.5 39.5 41.0 39.0
Br. 37.0 37.5 37.5 37.0 38.0
2. We want to compare two brands of tires to see
Diff. 3.0 2.0 2.0 4.0 1.0
how they wear on the front wheels of front-
wheel-drive cars. Each car is given one tire of
Pair: 6 7 8 9 10
each brand for its front. For each car the loca-
Bk. 38.0 38.5 38.5 39.0 39.5
tion (left or right) is assigned at random to the
Br. 38.0 38.5 40.5 39.0 39.0
brand.
Diff. 0.0 0.0 −2.0 0.0 0.5
Regarding the second example, if we have, say, 20 The sorted 25 differences are printed and analyzed
cars for study and we randomize we might end up below.
with, say, Brand A being on 12 left wheels and 8
right. If we decide to force these two numbers to prin c1 %
be identical (10 each in this example), we get what
2.4. PAIRED DATA 35

C1 I now want to suggest and investigate an inappro-


-3.5 -2.0 0.0 0.0 0.0 priate way to analyze data. I will do this via com-
0.5 0.5 1.0 1.0 1.0 puter simulation. Suppose that I have independent
1.5 1.5 1.5 1.5 2.0 random samples of size 10 each from two standard
2.0 2.0 2.0 2.0 2.5
normal pdfs. For example, I generated such data on
2.5 2.5 3.0 4.0 7.0
Minitab and got the results below.
tint c1 %
Sample 1
N MEAN STDEV 95.0% C.I.
C1 25 1.44 1.933 (0.64,2.24) 0.30 −0.22 −0.74 0.05 0.79
−1.70 1.15 0.02 −1.85 −0.02
ttest c1 %
Sample 2
TEST OF MU = 0 VS MU N.E. 0

N MEAN STDEV T P VALUE −0.45 0.40 −1.09 0.41 −0.91


C1 25 1.44 1.933 3.73 0.0011 0.24 −1.97 −0.83 1.63 −0.39

sint c1 % Now, let’s sort each sample.

SIGN CONF INT FOR MED Sample 1, Sorted

N MEDIAN −1.85 −1.70 −0.74 −0.22 −0.02


C1 25 1.500 0.02 0.05 0.30 0.79 1.15

ACHIEVED Sample 2, Sorted


CONF CONF INT POS
0.892 (1.0, 2.0) 9
0.950 (1.0, 2.0) NLI −1.97 −1.09 −0.91 −0.83 −0.45
0.957 (1.0, 2.0) 8 −0.39 0.24 0.40 0.41 1.63

stest c1 % Now, let’s pair the sorted data, matching the smallest
values in each set, the second smallest values, and
SIGN TEST OF MED = 0 VS N.E. 0 so on. Then, after pairing, subtract the values in the
N BELOW EQUAL ABOVE second sample from the corresponding values in the
C1 25 2 3 20 first sample. The 10 differences are below.
P-VALUE MEDIAN Differences of Sorted Data
0.0001 1.500
0.12 −0.61 0.17 0.61 0.43
wint c1 %
EST 0.41 −0.19 −0.10 0.38 −0.48
N MED CONF CONF INT
C1 25 1.50 95.0 (1.0, 2.0) I constructed two 95% confidence intervals for the
difference of the means. Using the two independent
wtest c1 % samples, pooled estimate of variance (an appropriate
analysis), I obtained [−0.85, 1.00]. Using the dif-
TEST OF MED = 0 VS MED N.E. 0 ferences of the sorted data, I obtained [−0.22, 0.37].
N FOR WILC EST We will see that this latter analysis is incorrect. At
N TEST STAT P-VALUE MED this point, however, the latter analysis looks supe-
C1 25 22 220.5 0.002 1.5
rior: both intervals are correct (they contain 0) and
the second interval is much more precise.
36 CHAPTER 2. THE TWO SAMPLE PROBLEM

I repeated the above steps 1,000 times. For each


pair of samples I constructed the 95% confidence in-
terval (pooled) for the difference of the means. Of
these intervals, 945 (94.5%) were correct. This is as
expected. But for each pair of samples I also sorted
the data and formed pairs of the sorted data. Then I
calculated differences. Using the one sample t proce-
dure, 517 (51.7%) of the intervals were correct! This
horrible performance demonstrates that the pairing is
not valid!

You might also like