1
Introduction to Advanced Non-Parametric Tests
Advanced tests go beyond the Mann-Whitney or Wilcoxon signed-rank. They handle:
• Multiple groups (Kruskal-Wallis, Friedman)
• Categorical dependence (McNemar, Cochran)
• Trend detection (Cox-Stuart)
• Variance differences (Square rank test)
• Distribution percentiles (Quantile test)
These tests are distribution-free (no normality assumption) and robust to outliers, but they still
require independence and, in some cases, identical distribution shapes (except for location shift).
Robustness of a non-parametric test
The robustness of a non-parametric test refers to its ability to remain valid (correct Type I error
rate) and powerful (able to detect true effects) even when the assumptions of a comparable
parametric test are violated — particularly normality, homogeneity of variance, or the presence of
outliers.
General Robustness Properties of Non-Parametric Tests
• No normality assumption – Most non-parametric tests (Mann–Whitney U, Wilcoxon
signed-rank, Kruskal–Wallis, etc.) do not require the data to come from a normal
distribution. This makes them valid for ordinal data, heavily skewed data, or data with
unknown distributions.
• Outlier resistance – Because many non-parametric tests are based on ranks, extreme
outliers have bounded influence. For example, in the Mann–Whitney U test, the largest
value is just rank n, no matter how extreme it is.
• Valid under heteroscedasticity – Some non-parametric tests (e.g., Brunner–Munzel test)
are more robust to unequal variances than parametric tests (like the t-test with equal
variance assumption).
• Valid for ordinal data – Non-parametric tests only require that observations can be
ranked, not that they are measured on an interval scale.
When is a Non-Parametric Test Not Robust?
• Violations of independence → Both parametric and non-parametric tests fail.
• Very different distribution shapes (e.g., one group skewed right, another skewed left)
→ Mann–Whitney U might not test medians properly.
• Discrete data with many ties → Can inflate Type I error if not corrected.
Mst. Tanmin Nahar, Lecturer, KU
2
Asymptotic Relative Efficiency (ARE)
The purpose of asymptotic relative efficiency is to compare two statistical procedures by
comparing the sample sizes, n1 and n2, say, at which those procedures achieve some given measure
of performance; the ratio 𝑛2 /𝑛1 is called the relative efficiency of procedure one with respect to
procedure two. Finite-sample evaluations being difficult or impossible, a sequence of measures of
performances requiring that those sample sizes go to infinity is generally considered. If those
measures of performance are indexed by n, say, so that n1 and n2 take the form n1(n) and n2(n), the
limit limn→∞n2(n)/n1(n), if it exists, is called the asymptotic relative efficiency of procedure one
with respect to procedure two.
ARE measures how well are statistical method performs relative to another as the sample size
becomes large.
The factor c is called the asymptotic efficiency of the test. The asymptotic relative efficiency is
denoted by 𝑒12 ,
𝑐 2
𝑒12 = (𝑐1 ) which is ARE of test 1 with respect to test 2. It is use to compare the efficiency of
2
non-parametric and parametric test.
Example: Suppose, we are comparing two tests for the mean a population. A parametric t-test and
a nonparametric Wilcoxon rank-sum test (Mann-Whitney U test).
ARE of the Wilcoxon rank-sum test:
It is possible to work out the ARE of Wilcoxon rank-sum test, although it is much easier to work
with the algebraically equivalent to Mann-Whitney U test statistic:
1 1
𝑇𝑚𝑤 = [𝑇𝑤 − 𝑛(𝑛 − 1)]
𝑚𝑛 2
Where, 𝑇𝑤 is the Wilcoxon rank-sum test statistic. Here, 𝑚 is the sample of size of 𝑥 and n in the
sample size of 𝑦.
𝑚
Let, 𝑁 = 𝑚 + 𝑛 and 𝜆 = 𝑁 .
𝑁+1
𝜎𝑁 (𝜃0 ) = √
12𝑚𝑛
𝜆(1 − 𝜆)
𝑐1 = √
𝜎𝑓
Mst. Tanmin Nahar, Lecturer, KU
3
Suppose, the true distribution of the data is exponential with mean 𝛽.
2
𝑐𝑤 2 √3𝜆(1 − 𝜆). 𝛽 −1
𝑐𝑤𝑡 =( ) =[ ] =3
𝑐𝑡 √𝜆(1 − 𝜆). 𝛽 −1
Interpretation: When the true distribution of the data is exponential the Wilcoxon rank-sum test
is 3 times more efficient than the t-test.
ARE of the Mann-Whitney test:
The Mann-Whitney U Test (also known as the Wilcoxon Rank-Sum Test) is a non-parametric
alternative to the independent samples t-test. It is used to compare differences between two
independent groups when the dependent variable is either ordinal or continuous, but not normally
distributed.
The Asymptotic Relative Efficiency (ARE) measures the power of the Mann-Whitney test
compared to the t-test as the sample size approaches infinity.
ARE of the Sign test
The ARE of the Sign Test is calculated relative to the t-test as sample sizes approach infinity. It
represents the ratio of the sample sizes required by each test to achieve the same statistical power.
The ARE measures how efficient the Sign test is in comparison to another test — typically the t-
test or the Wilcoxon signed-rank test — when sample size goes to infinity.
The efficiency of the Sign Test is highly sensitive to the "peakedness" (kurtosis) of the underlying
distribution.
When comparing the Sign test to the t-test under a normal distribution, the ARE is
2
ARE(Sign test, t − test) = = 0.637 ≈ 0.637
𝜋
1
This means the Sign test needs roughly 0.637 ≈ 1.57 times as many observations as the t-test to
achieve the same power when data are normal.
When comparing the Sign test to the Wilcoxon signed-rank test, the ARE is about 0.67 (under
normality), so the Sign test is again less efficient.
Uses of ARE:
i. Comparing estimators.
ii. Model selection.
Mst. Tanmin Nahar, Lecturer, KU
4
iii. Test comparisons
iv. Designing Experiments.
v. Evaluating Robustness.
vi. Choosing sampling Plans
vii. Resource Allocation.
viii. Economic Decision Making.
ARE between two nonparametric tests:
Testing the ARE between two non-parametric test is less straightforward to comparing parametric
and non-parametric tests. It is used to compare the efficiency of different statistical tests under
certain conditions and it is often applied when comparing a parametric test to a nonparametric test.
Cramer's Contingency Coefficient
The Cramer’s contingency coefficient also known as Cramer’s V or RXC contingency of table. It
is a measure of association used to assess the strength and significance of the relationship between
two categorical variables in a contingency table. It is an extension of the phi- coefficient (𝜙) for
larger contingency tables.
Assumptions:
• The observations are independent
• It used categorical variable.
• Shouldn't have missing data.
• Most homogeneity of variance for large sample.
𝑠2
The test statistic, 𝑉 = √𝑁[(𝑟−1)(𝑐−1)]
Where, 𝑉 is the Cramer’s contingency coefficient.
𝑠 2 is the chi-square statistic for the contingency table.
𝑁 is the total number of observation
𝑟 is number of rows.
c is the number of columns.
Interpretation: The range of V is 0 to 1.
0 = no association between the categorical variable.
0-0-95 = less association.
Mst. Tanmin Nahar, Lecturer, KU
5
0-25-0.75 = moderately association
0.75 -1 = stronger association.
1 = perfect association.
Example: The contingency table,
OK Not OK
Staff 30 70
Student 50 50
Find the Cramer’s V and interpret it.
Solution: From the given table, we can get,
OK Not OK
Staff 40 60
Student 40 60
Now,
(30 − 40)2 (70 − 60)2 (50 − 40)2 (50 − 60)2
𝑠2 = + + +
40 60 40 60
(−10)2 (10)2 (10)2 (−10)2
= + + +
40 60 40 60
= 8.33
Then.
𝑠2 8.33
𝑉=√ == √ = 0.20
𝑁[(𝑟 − 1)(𝑐 − 1)] 200 × [(2 − 1)(2 − 1)]
So, this is less associated.
CRS (Completely Randomized Design) in non-parametric context:
Kruskal-Wallis = non-parametric ANOVA for CRS.
Mann-Whitney U = pairwise comparison in CRS.
Square rank test for variances = CRS for dispersion effects.
Mst. Tanmin Nahar, Lecturer, KU