Two Sample Scale Problem
Example (Quality control)
Suppose two machines for bottling coca cola are designed to fill the cans
with 330 ml of the soft drink. It is expected that the observed data on the
amount of coke in each can from two machines are centered around 330
ml but their variability may not be the same (Figure-1)
Figure 1
Seigel-Tukey Test
The data consist two independent samples: 𝑋1 , 𝑋2 , …,𝑋𝑚 denote the
random sample of size 𝑚 from population 1 and 𝑌1 , 𝑌2 , …,𝑌𝑛 denote the
random sample of size 𝑛 from population 2. It is assumed that two
populations are continuous
Null hypothesis 𝐻0 : 𝜎𝑋 = 𝜎𝑌
The test procedure is
- Arrange the combined data from smallest to largest
- Assign rank 1 to the smallest observation, rank 2 to the largest
observation, rank 3 to the next largest observation, rank 4 to the next
smallest observation, rank 5 to the next smallest observation and so
on (Table 1)
The test statistic is
𝑁
𝑆= 𝑎𝑖 𝑍𝑖
𝑖=1
= the sum of weights assigned to 𝑋 sample in combined sample
Note A small rank sum indicates that 𝑋 has more extreme observations
than 𝑌 and 𝑋 tends to have large variability than 𝑌
Table 1 weights of Seigel-Tukey test
i 1 2 3 4 … 𝑁/2
𝑎𝑖 1 4 5 8 … 𝑁
i (𝑁/2)+1 … 𝑁-1 𝑁
𝑎𝑖 𝑁−1 … 3 2
Case 1: Small sample
The alternative hypotheses and their corresponding 𝑝-values are
summaries below
Alternative Critical region 𝑝-value
𝐻1 :𝜎𝑋 > 𝜎𝑌 𝑆 ≤ 𝑠𝛼 𝑃 𝑆 ≤ 𝑠|𝐻0
𝐻1 :𝜎𝑋 < 𝜎𝑌 𝑆 ≥ 𝑠′𝛼 𝑃 𝑆 ≥ 𝑠|𝐻0
𝐻1 :𝜎𝑋 ≠ 𝜎𝑌 𝑆 ≥ 𝑠𝛼/2 or S≤ 𝑠′𝛼/2 2(smaller of the one-tail 𝑝-
values)
where 𝑠 𝛼and 𝑠′𝛼 are the largest and smallest integers such that 𝑃 𝑆 ≤ 𝑠𝛼 |𝐻0 ≤ 𝛼
and 𝑃 𝑆 ≥ 𝑠′𝛼 |𝐻0 ≤ 𝛼; 𝑃 𝑆 ≥ 𝑠𝛼/2 |𝐻0 + 𝑃 𝑆 ≤ 𝑠′𝛼/2|𝐻0 ≤ 𝛼 and 𝑠 is
the observe value of the test statistics
Case 2: large sample
For large sample the null distribution (under 𝐻0 ) of 𝑆 is normal
Alternative Approximate 𝑝-value Critical region
𝐻1 :𝜎𝑋 > 𝜎𝑌 𝑚(𝑁 + 1)
𝑠 + 0.5 −
Φ 2
𝑆
𝑚𝑛(𝑁 + 1)
12 𝑚(𝑁 + 1) 𝑚𝑛(𝑁 + 1)
≤ − 0.5 − 𝑧 𝛼
2 12
𝐻 1 :𝜎 𝑋 < 𝜎𝑌 𝑚(𝑁 + 1)
𝑠 − 0.5 − 𝑆
1−Φ 2
𝑚(𝑁 + 1)
𝑚𝑛(𝑁 + 1) ≥ + 0.5
2
12
where Φ(.) the cdf of standard normal + 𝑧𝛼
𝑚𝑛(𝑁 + 1)
distribution 12
𝐻 1 :𝜎 𝑋 ≠ 𝜎𝑌 2(smaller of the one-tail 𝑝-values) Both of above with 𝑧 𝛼/2
where 𝑧 𝛼 is upper 𝛼th quantile of standard normal probability distribution
Null distribution of 𝑺
Suppose we have 𝑚 = 4 observations from population 1 (𝑥𝑥𝑥𝑥) and 𝑛 = 2
observations from population 2 (𝑦𝑦)
If we combine two samples, we have total (4+2 = 6) observations. There are 6!
=
2!4!
15 possible arrangements of 𝑥’s and 𝑦’s. The possible weights are {1,2,[Link]}
Rank of 𝑋 𝑆 Rank of 𝑋 𝑆
1,2,3,4 10 1,3,5,6 15
1,2,3,5 11 1,4,5,6 16
1,2,3,6 12 2,3,4,5 14
1,2,4,5 12 2,3,4,6 15
1,2,4,6 13 2,3,5,6 16
1,2,5,6 14 2,4,5,6 17
1,3,4,5 13 3,4,5,6 18
1,3,4,6 14
𝑆 =𝑠 10 11 12 13 14 15 16 17 18
𝑃(𝑆 1 1 2 2 3 2 2 1 1
= 𝑠) 15 15 15 15 15 15 15 15 15
=0.067 =0.067
Theorem If the null hypothesis 𝐻0 holds then
𝑚(𝑁 + 1) 𝑚𝑛(𝑁 + 1)
𝐸 𝑆|𝐻0 = , 𝑉 𝑆|𝐻0 =
2 12
Ansari–Bradley Test
The data consist two independent samples: 𝑋1 , 𝑋2 , …,𝑋𝑚 denote the
random sample of size 𝑚 from population 1 and 𝑌1 , 𝑌2 , …,𝑌𝑛 denote the
random sample of size 𝑛 from population 2. It is assumed that two
populations are continuous and 𝑀𝑋 ,𝑀𝑌 are known
Null hypothesis 𝐻0 : 𝜎𝑋 = 𝜎𝑌
The test procedure is
- Arrange the combined data from smallest to largest
- Assign weights according to the following table (Table 2)
The test statistic is
𝑁
𝐴𝐵 = 𝑎𝑖 𝑍𝑖
𝑖=1
= the sum of weights assigned to 𝑋 sample in combined sample
Note A small rank sum indicates that 𝑋 has more extreme observations
than 𝑌 and 𝑋 tends to have large variability than 𝑌
Table 2
I 1 2 3 … 𝑁 … 𝑁-2 𝑁-1 𝑁 (even)
2
𝑎𝑖 1 2 3 … 𝑁 … 3 2 1
2
I 1 2 3 … … 𝑁-1 𝑁 (odd)
𝑎𝑖 1 2 … 𝑁−1 𝑁+1 𝑁−1 … 2 1
2 2 2
Null distribution of 𝑨𝑩
Suppose we have 𝑚 = 2 observations from population 1 (𝑥𝑥) and 𝑛 = 3 observations
from population 2 (𝑦𝑦𝑦)
If we combine two samples, we have total (2+3 = 5) observations. There are 5!
=
2!3!
10 possible arrangements of 𝑥’s and 𝑦’s. The possible weights are {1,1,2,2,3}
Rank of 𝑋 𝐴𝐵 Rank of 𝑋 𝐴𝐵
{1,2} 3 {2,2} 4
{1,3} 4 {1,2} 3
{1,2} 3 {2,3} 5
{1,1} 2 {1,3} 4
{2,3} 5 {1,2} 3
𝐴𝐵 2 3 4 5
1 4 3 2
𝑃(𝐴𝐵) 10 10 10 10
Linear Rank Statistic
Let 𝑍 = (𝑍1 ,𝑍2 , …,𝑍𝑁 ) where 𝑍𝑖 = 1 if the ith random variable in the combined
ordered sample is an 𝑋 and 𝑍 𝑖 = 0 if it is a 𝑌
A linear rank statistic is a linear combination of indicator variables as
𝑁
𝑇𝑁 = 𝑎𝑖 𝑍𝑖
𝑖=1
where 𝑎𝑖 are weights or scores
Theorem
Under the null hypothesis 𝐻0 :𝐹𝑌 𝑥 = 𝐹𝑋 𝑥 = 𝐹(𝑥) for all 𝑥.
𝑁
𝑎𝑖
𝐸(𝑇𝑁 ) = 𝑚
𝑁
𝑖=1
𝑁 𝑁 2
𝑚𝑛
𝑉(𝑇𝑁 ) = 2
𝑁 𝑎2𝑖 − 𝑎𝑖
𝑁 (𝑁 − 1) 𝑖=1 𝑖=1