Is there any differences in number of -Wald-Wolfowitz run test
Facebook friends of male and female -Median test
internet users? -Control Median test
-Mann-Whitney U test
Figure 2.1 a - b
𝐻0 : Two populations are identical
𝐻1 : Two populations are not identical
If 𝐻0 is rejected then we can conclude that population 1 tends to provide
either smaller or larger values than population 2 (Figure 2.1 a -b)
𝐻0 : 𝐹𝑌 (𝑥) = 𝐹𝑋 (𝑥) for all 𝑥
or
𝐻0 : 𝑀𝑌 = 𝑀𝑋
𝐹𝑌 (𝑥) = 𝐹𝑋 (𝑥 − 𝜃) This means 𝑋 + 𝜃 and 𝑌 have the same distribution
Wald-Wolfwitz Run Test
What is a run?
A run is defined to be a succession of one or more symbols which are
followed or preceded by a different symbol or no symbol at all
𝑌𝑌𝑌𝐶𝐶𝑇𝑇𝑇𝐶
Why run?
when the distributions are equal, the number
of runs will likely be large
The data consist two independent samples: 𝑋1 , 𝑋2 , … , 𝑋𝑚 denote the
random sample of size 𝑚 from population 1 and 𝑌1 , 𝑌2 , … , 𝑌𝑛 denote the
random sample of size 𝑛 from population 2
Null hypothesis 𝐻0 : 𝑀𝑋 = 𝑀𝑌
The test statistic is
𝑅 = total number of runs in the combined ordered arrangement of 𝑚 𝑋
and 𝑛𝑌 samples
Case 1: Small sample
The alternative hypotheses and their corresponding 𝑝-values are
summaries below
Alternative Critical region 𝑝-value
𝐻1 : 𝑀𝑋 > 𝑀𝑌 𝑅 ≥ 𝑟𝛼 𝑃(𝑅 ≥ 𝑟|𝐻0 )
𝐻1 : 𝑀𝑋 < 𝑀𝑌 𝑅 ≤ 𝑟′𝛼 𝑃(𝑅 ≤ 𝑟|𝐻0 )
𝐻1 : 𝑀𝑋 ≠ 𝑀𝑌 𝑅 ≥ 𝑟𝛼/2 or 𝑅 ≤ 𝑟′𝛼/2 2(smaller of the one-tail 𝑝-
values)
where 𝑟𝛼 and 𝑟𝛼′ are the smallest and largest integers such that 𝑃(𝑅 ≥ 𝑟𝛼 |𝐻0 ) ≤ 𝛼 and
𝑃(𝑅 ≤ 𝑟′𝛼 |𝐻0 ) ≤ 𝛼; 𝑃(𝑅 ≥ 𝑟𝛼/2 |𝐻0 ) + 𝑃(𝑅 ≤ 𝑟′𝛼/2 |𝐻0 ) ≤ 𝛼 and 𝑟 is the
observe value of the test statistics
Case 2: large sample
For large sample the null distribution (under 𝐻0 ) of 𝑅 is normal
Alternative Approximate 𝑝-value Critical region
𝐻1 : 𝑀𝑋 > 𝑀𝑌
2𝑛1 𝑛2 2𝑛1 𝑛2 2𝑛1 𝑛2 (2𝑛1 𝑛2 − 𝑛1 − 𝑛2 )
𝑟 − 0.5 − −1 𝑅 ≥ + 1 + 0.5 + 𝑧𝛼 √
(𝑛1 +𝑛2 ) (𝑛1 +𝑛2 ) (𝑛1 +𝑛2 )2(𝑛1 +𝑛2 − 1)
1−Φ
2𝑛1 𝑛2 (2𝑛1 𝑛2 − 𝑛1 − 𝑛2 )
√
(𝑛1 +𝑛2 )2 (𝑛1 +𝑛2 − 1)
( )
where Φ(. ) the cdf of standard normal
distribution
𝐻1 : 𝑀𝑋 < 𝑀𝑌
2𝑛1 𝑛2 2𝑛1 𝑛2 2𝑛1 𝑛2 (2𝑛1 𝑛2 − 𝑛1 − 𝑛2 )
𝑟 + 0.5 − −1 𝑅 ≤ + 1 − 0.5 − 𝑧𝛼 √
(𝑛1 +𝑛2 ) (𝑛1 +𝑛2 ) (𝑛1 +𝑛2 )2(𝑛1 +𝑛2 − 1)
Φ
2𝑛1 𝑛2 (2𝑛1 𝑛2 − 𝑛1 − 𝑛2 )
√
(𝑛1 +𝑛2 )2 (𝑛1 +𝑛2 − 1)
( )
𝐻1 : 𝑀𝑋 ≠ 𝑀𝑌 2(smaller of the one-tail 𝑝-values) Both of above with 𝑧𝛼/2
where 𝑧𝛼 is upper 𝛼th quantile of standard normal probability distribution
Null distribution of 𝑹
Suppose we have 𝑛1 = 3 observations from population 1 (𝑥𝑥𝑥) and 𝑛2 = 3
observations from population 2 (𝑦𝑦𝑦)
6!
If we combine two samples, we have total (3+3 = 6) observations. There are =
3!3!
20 possible arrangements of 𝑥’s and 𝑦’s
Arrangements Total Run
𝑥𝑥𝑥 𝑦𝑦𝑦 2
𝑦𝑦𝑦𝑥𝑥𝑥 2
𝑥𝑦𝑦𝑦𝑥𝑥 3
𝑥𝑥𝑦𝑦𝑦𝑥 3
𝑦𝑥𝑥𝑥𝑦𝑦 3 (alt.)
𝑦𝑦𝑥𝑥𝑥𝑦 3(alt)
𝑥𝑦𝑥𝑥𝑦𝑦 4
𝑥𝑥𝑦𝑥𝑦𝑦 4
𝑥𝑦𝑦𝑥𝑥𝑦 4
𝑥𝑥𝑦𝑦𝑥𝑦 4
𝑦𝑥𝑦𝑦𝑥𝑥 4(alt)
𝑦𝑦𝑥𝑦𝑥𝑥 4
𝑦𝑥𝑥𝑦𝑦𝑥 4
𝑦𝑦𝑥𝑥𝑦𝑥 4
𝑥𝑦𝑥𝑦𝑦𝑥 5
𝑥𝑦𝑦𝑥𝑦𝑥 5
𝑦𝑥𝑦𝑥𝑥𝑦 5(alt)
𝑦𝑥𝑥𝑦𝑥𝑦 5
𝑥𝑦𝑥𝑦𝑥𝑦 6
𝑦𝑥𝑦𝑥𝑦𝑥 6
𝑅 2 3 4 5 6
𝑃(𝑅 = 𝑟) 2 4 8 4 2
20 20 20 20 20
Theorem 1.
The probability distribution of 𝑅 total number of runs of 𝑛 = 𝑛1 +𝑛2 objects
𝑛1 − 1 𝑛2 − 1
2( 𝑟 )( 𝑟 )
−1 −1
2 2 if 𝑟 even
𝑛1 +𝑛2
( )
𝑓𝑅 (𝑟) = { 𝑛1
𝑛1 − 1 𝑛 −1 𝑛 −1 𝑛 −1
{( )( 2 )+( 1 )( 2 )}
(𝑟 − 1)/2 (𝑟 − 3)/2 (𝑟 − 3)/2 (𝑟 − 1)/2
if 𝑟 odd
𝑛 +𝑛2
( 1
𝑛1 )
𝒓 = 𝟐, 𝟑, … , 𝑛1 +𝑛2
For example, 𝑛1 = 3 𝑛2 = 3
2 2
2( )( ) 8
𝑓𝑅 (4) = 1 1 =
6
( ) 20
3
2 2 2 2
( )( ) + ( )( )
𝑓𝑅 (5) = 2 1 1 2 = 4
6
( ) 20
3
Theorem 2 If the null hypothesis 𝐻0 holds then
2𝑛1 𝑛2 2𝑛1 𝑛2 (2𝑛1 𝑛2 − 𝑛1 − 𝑛2 )
𝐸 (𝑅|𝐻0 ) = +1, 𝑉 (𝑅|𝐻0 ) =
(𝑛1 +𝑛2 ) (𝑛1 +𝑛2 )2 (𝑛1 +𝑛2 − 1)
Median Test
The data consist two independent samples. Let 𝑋1 , 𝑋2 , … , 𝑋𝑚 denote the
random sample of size 𝑚 from population 1 and let 𝑌1 , 𝑌2 , … , 𝑌𝑛 denote
the random sample of size 𝑛 from population 2
Null hypothesis 𝐻0 : 𝑀𝑋 = 𝑀𝑌
The test statistic can be written as
𝑈 = # of 𝑋 observations that precede 𝛿 (median of combined samples)
Case 1: Small sample
The alternative hypotheses and their corresponding 𝑝-values are
summaries below
Alternative Critical value 𝑝-value
𝐻1 : 𝑀𝑋 > 𝑀𝑌 𝑈 ≥ 𝑢𝛼 𝑃(𝑈 ≥ 𝑢|𝐻0 )
𝐻1 : 𝑀𝑋 < 𝑀𝑌 𝑈 ≤ 𝑢′𝛼 𝑃(𝑈 ≤ 𝑢|𝐻0 )
𝐻1 : 𝑀𝑋 ≠ 𝑀𝑌 𝑈 ≥ 𝑢𝛼/2 or 𝑈 ≤ 𝑢′𝛼/2 2(smaller of the one-tail 𝑝-
values)
where 𝑢𝛼 and 𝑢𝛼′ are the smallest and largest integers such that 𝑃(𝑈 ≥ 𝑢𝛼 |𝐻0 ) ≤ 𝛼 and
𝑃(𝑈 ≤ 𝑢′𝛼 |𝐻0 ) ≤ 𝛼; 𝑃(𝑈 ≥ 𝑢𝛼/2 |𝐻0 ) + 𝑃(𝑈 ≤ 𝑢′𝛼/2 |𝐻0 ) ≤ 𝛼 and 𝑢 is
the observe value of median test statistics
Case 2: large sample
For large sample the null distribution (under 𝐻0 ) of 𝑈 is normal
Alternative Approximate 𝑝-value Critical region
𝐻1 : 𝑀𝑋 > 𝑀𝑌 𝑢 − 0.5 − 𝑚𝑡/𝑁
1 − Φ( ) 𝑈≥
𝑚𝑡
+ 0.5 + 𝑧𝛼 √𝑚𝑛𝑡(𝑁 − 𝑡)𝑁 3
√𝑚𝑛𝑡(𝑁 − 𝑡)𝑁 3 𝑁
where Φ(. ) the cdf of standard normal
distribution
𝐻1 : 𝑀𝑋 < 𝑀𝑌 𝑢 + 0.5 − 𝑚𝑡/𝑁
Φ( ) 𝑈≤
𝑚𝑡
− 0.5 − 𝑧𝛼 √𝑚𝑛𝑡(𝑁 − 𝑡)𝑁 3
√𝑚𝑛𝑡(𝑁 − 𝑡)𝑁 3 𝑁
𝐻1 : 𝑀𝑋 ≠ 𝑀𝑌 2(smaller of the one-tail 𝑝-values) Both of above with 𝑧𝛼/2
(𝑁−1) 𝑁
where 𝑡 = # of observations less than 𝛿 where 𝑡 equals or according to 𝑁 = (𝑚 + 𝑛) odd
2 2
or even
Null distribution The null distribution of 𝑈 is
𝑚 𝑛
( )( )
(
𝑓 𝑢 =) 𝑢 𝑡 − 𝑢
𝑚+𝑛
( )
𝑡
𝑢 = max(0, 𝑡 − 𝑛) , … , min (𝑚, 𝑡)
Theorem If the null hypothesis 𝐻0 holds then
𝑚𝑡 𝑚𝑛𝑡
𝐸 (𝑈|𝐻0 ) = , 𝑉 (𝑈|𝐻0 ) =
𝑁 𝑁 2 (𝑁 − 1)
Proof
The null distribution
𝑚 𝑛
( )( )
𝑢 𝑡−𝑢
𝑓(𝑢) = 𝑚+𝑛 𝑢 = max(0, 𝑡 − 𝑛) , … , min (𝑚, 𝑡)
( )
𝑡
is a pmf of hypergeometric distribution
𝑚𝑡 𝑚𝑛𝑡
𝐸 (𝑈|𝐻0 ) = , 𝑉 (𝑈|𝐻0 ) =
𝑁 𝑁 2 (𝑁 − 1)
For large N= m+n
𝑚𝑛𝑡
𝑉(𝑈|𝐻0 ) =
𝑁3