MODULE 8: HYPOTHESIS TESTING
MODULE 8: HYPOTHESIS
TESTING 268
LEARNING OUTCOMES
Eplain hypothesis testing and its components, including statistical
8.a.
significance, Type I and Type II errors, and the power of a test
Construct hypothesis tests and determine their statistical
8.b. significance, the associated Type I and Type II errors, and power of
the test given a significance level
Compare and contrast parametric and nonparametric tests, and
8.c.
describe situations where each is the more appropriate type of test
MODULE 8: HYPOTHESIS
TESTING 269
OVERVIEW MODULE 8 + MODULE 9
Test of variance
• Test of single variance
• Test of difference in variances
P-value Test of independence
Another way to test whether H0 is Using contingency table data
rejected
-LOS 8.a -LOS 8.b -LOS 9.b
-LOS 8.a -LOS 9.a
The process of hypothesis testing Test of correlation
6 steps • Parametric test
• Nonparametric tests
(Spearman rank correlation test)
Test of mean
• Test of single mean
• Test of Difference in means vs. Mean of differences
MODULE 8: HYPOTHESIS
TESTING 270
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors, and the power of a
test
Why we need to use Hypothesis testing?
Returns for investment One and Two over 33 years
The plot above represents the returns on two investments over 33 years,
but what can we actually glean from this plot?
• Can we tell if each investment’s returns are different from an average of
5%?
• Can we tell whether the returns are different for Investment One and
Investment Two?
• Can we tell whether the variability is different for the two investments?
Hypothesis testing is used to address these questions
MODULE 8: HYPOTHESIS
TESTING 271
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
The process of hypothesis testing
Hypothesis testing is the process of evaluating the accuracy of a statement
regarding a population parameter given sample information.
Step 1: State the hypotheses
Step 2: Identify the appropriate test statistic
Step 3: Specify the level of significance
Step 4: State the decision rule
Step 5: Collect data and calculate the test statistic
Step 6: Make a decision
MODULE 8: HYPOTHESIS
TESTING 272
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
1. Step 1: State the hypotheses
Null hypothesis (H0) Alternative hypothesis (Ha)
Definition: The null hypothesis is Definition: The alternative
the hypothesis that the researcher hypothesis is the hypothesis
wants to reject. concluded if H0 is rejected, also
called the hoped-for hypothesis.
H0 can be stated as (considering Ha can be stated as (considering
population mean): population mean):
μ = μ0 μ ≠ μ0
μ ≤ μ0 μ > μ0
μ ≥ μ0 μ < μ0
Note: The most common null hypothesis will include the “equal to” sign and
the alternative will include the “not equal to” sign. However, the null
hypothesis sometimes can be a “not equal to” hypothesis, combined with an
“equal to” alternative.
MODULE 8: HYPOTHESIS
TESTING 273
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
1. Step 1: State the hypotheses
Example 1: Calculating Roy’s safety first criterion
State the hypotheses of the test if we want to test whether:
(i) the population mean return is equal to 6%.
(ii) the population mean return is greater than 6%.
(iii) the population mean return is less than 6%.
MODULE 8: HYPOTHESIS
TESTING 274
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
1. Step 1: State the hypotheses
Example 1: Calculating Roy’s safety first criterion
Answer:
Recall the basis statement:
• The null hypothesis (H0) is the hypothesis that we are interested in
rejecting.
• The alternate hypothesis (Ha) is essentially the statement whose validity
we are trying to evaluate.
(i) The population mean return is equal to 6%. We would state the
hypotheses as:
H0: µ = 6
Ha: µ ≠ 6.
(ii) The population mean return is greater than 6%. We would state the
hypotheses as:
H0: µ ≤ 6 → this is the hypothesis we want to reject.
Ha: µ > 6 → this is the “hope-for” hypothesis, which is also the hypothesis
we want to test.
(ii) The population mean return is less than 6%. We would state the
hypotheses as:
H0: µ ≥ 6
Ha: µ < 6.
MODULE 8: HYPOTHESIS
TESTING 275
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
1. Step 1: State the hypotheses
One-tailed test Two-tailed test
Test structure is either: Test structure is:
• Right tail: H0: μ = μ0 versus Ha: μ ≠ μ0
H0: μ ≤ μ0 versus Ha: μ > μ0
• Left tail:
H0: μ ≥ μ0 versus Ha: μ < μ0
Note: The null and alternative hypotheses must be mutually exclusive and
collectively exhaustive; in other words, all possible values are contained in
either the null or the alternative hypothesis.
Illustration 2:
Continue with the previous example 1, we classify test structure as follow:
One-tailed test Two-tailed test
• Right tail: H0: µ = 6 versus Ha: µ ≠ 6
H0: µ ≤ 6 versus Ha: µ > 6
• Left tail:
H0: µ ≥ 6 versus Ha: µ < 6
MODULE 8: HYPOTHESIS
TESTING 276
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
2. Step 2: Identify the appropriate test statistic
Definition: A test statistic is a value calculated using a sample, when used in
conjunction with a decision rule, is the basis for deciding whether to reject
the null hypothesis.
Formula:
Test statistic = Sample statistic − Hypothesized value
Standard error of the sample statistics
Illustration 3:
Consider the sample mean X calculated from a sample of returns drawn
from the population. We have test statistic of a population mean in two
cases:
When the population’s standard When the population’s standard
deviation (σ) is known deviation (σ) is unknown
X − μ0 X − μ0
Test statistic = σ/ n Test statistic = s/ n
where σ is the population standard where s is the sample standard
deviation deviation
MODULE 8: HYPOTHESIS
TESTING 277
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
2. Step 2: Identify the appropriate test statistic
The key to hypothesis testing is identifying the appropriate test statistic for
the hypotheses and the underlying distribution of the population.
Test statistic
Distribution of test statistic
Student F- Chi-square
z-statistic
t-statistic distribution distribution
All of these distributions will be further discussed in this MODULE
MODULE 8: HYPOTHESIS
TESTING 278
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
3. Step 3: Specify the level of significance
Since there is always a possibility that the sample may not be perfectly
representative of the population, resulting in the conclusions drawn from
the test may be wrong, there are two types of errors that can be made
when conducting a hypothesis test:
• Type I error: rejection of the null hypothesis when it is actually true (false
positive).
• Type II error: failure to reject the null hypothesis when it is actually false
(false negative).
→ These errors are mutually exclusive.
True situation
Decision
H0 true H0 false
Correct decision
Incorrect: Type II error
Fail to reject H0 P = confidence level
P(type II error) = 𝛽
=1– α
Incorrect: Type I error Correct decision
Reject H0 P(type I error) = α P = power of the test
(level of significance) =1–𝛽
MODULE 8: HYPOTHESIS
TESTING 279
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.1. Determining the critical value
Concept of decision rule:
• The decision we take is based on comparing the calculated test statistic
with a specified value or values, which we refer to as critical values.
• The critical value or values we choose are based on the level of
significance and the probability distribution associated with the test
statistic.
Note:
• If we find that the calculated value of the test statistic is more extreme
than the critical value or values, then we reject the null hypothesis; we say
the result is statistically significant. Otherwise, we fail to reject the null
hypothesis; there is not sufficient evidence to reject the null hypothesis.
• It is statistically incorrect to say “accept” the null hypothesis.
MODULE 8: HYPOTHESIS
TESTING 280
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.1. Determining the critical value
For one-tailed test One-tailed test using standard
normal distribution with α = 5%
• Right tail:
H0: μ ≤ μ0 versus Ha: μ > μ0 5%
Right one-
Reject H0 if: test statistic > upper critical tailed test 95%
value
+1.645
Do not reject H0 if: test statistic ≤ upper
Reject H0 Do not reject H0 Reject H0
critical value
• Left tail: 5%
Left one-
H0: μ ≥ μ0 versus Ha: μ < μ0 tailed test
95%
Reject H0 if test statistic < lower critical value
Do not reject H0 if: test statistic ≥ lower –1.645
critical value
Two-tailed test using standard
For two-tailed test normal distribution with α = 5%
H0: μ = μ0 versus Ha: μ ≠ μ0
• Reject H0 if:
test statistic > upper critical value or 2.5% 2.5%
test statistic < lower critical value 95%
• Do not reject H0 if:
lower critical value ≤ test statistic ≤ upper –1.96 +1.96
Reject H0 Do not reject H0 Reject H0
critical value
MODULE 8: HYPOTHESIS
TESTING 281
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.1. Determining the critical value
Illustration 4:
Suppose we conduct a two-tailed test at 5% significance level, and the test
statistic follows the z-value distribution.
The critical value is therefore determined on the z-distribution at α/2 = 5%/2
= 2.5% for both left and right tail.
2.5% 2.5%
95%
Lower critical value
We look this up in a cumulative z-table (P(Z) ≤ z) for z ≤ 0 with a P(Z) = 2.5%
MODULE 8: HYPOTHESIS
TESTING 282
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.1. Determining the critical value
Cumulative z-table with P(Z ≤ z) = N(z) for z ≤ 0
P(Z ≤ z)]
−Z
Cdf values for the standard normal distribution: The z-table
z .00 .01 .02 .03 .04 .05 .06 .07 .08 .09
−1.7 .0446 .0436 .0427 .0418 .0409 .0401 .0392 .0384 .0375 .0367
−1.8 .0359 .0351 .0344 .0336 .0329 .0322 .0314 .0307 .0301 .0294
−1.9 .0287 .0281 .0274 .0268 .0262 .0256 .0250 .0244 .0239 .0233
According to the z-table, from P(Z) = 2.5% we have z-critical = −1.96, because
it is standard normal distribution, so we have z-critical for each tail is:
• Upper z-value = +1.96
• Lower z-value = −1.96
MODULE 8: HYPOTHESIS
TESTING 283
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.2. The relation between confidence intervals and hypothesis tests
Illustration 5:
Suppose we conduct a two-tailed test of a population mean at 5%
significance level, and the test statistic follows the z-value distribution.
Hypothesis tests
2.5% 2.5%
95%
−1.96 1.96
Reject H 0 Do not reject H0 Reject H 0
MODULE 8: HYPOTHESIS
TESTING 284
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.2. The relation between confidence intervals and hypothesis tests
Confidence intervals
−1.96s X 1.96s
95%
Conclusion:
We reach a similar result between hypothesis test and confidence intervals:
We can make a decision by either comparing the calculated test statistic
with the critical values or comparing the hypothesized population
parameter (μ = μ0) with the bounds of the confidence interval.
A significance level in a two-sided hypothesis test can be interpreted in the
same way as a (1 − α) confidence interval.
MODULE 8: HYPOTHESIS
TESTING 285
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
4. Step 4: State the decision rule
4.3. Decision rules
Making a decision based on Critical values and Confidence intervals for a
two-tailed hypothesis:
Method 1: Based on critical Method 2: Based on
values confidence intervals
Compare the calculated test Compare the calculated test
Procedure statistic with the critical statistic with the bounds of
values. the confidence interval.
If the calculated test If the hypothesized value of
statistic is less than the the population parameter
lower critical value or under the null is outside the
Decision
greater than the upper corresponding confidence
critical value, reject the null interval, the null hypothesis
hypothesis. is rejected.
lower critical upper critical Confidence interval
value value of population mean
Illustration
reject Not reject reject reject Not reject reject
MODULE 8: HYPOTHESIS
TESTING 286
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
5. Step 5: Collect data and calculate the test statistic
Since we state the appropriate test statistics and their distributions in step
2, in this step, we calculate the formula of test statistics.
Example 6:
Suppose that the basketball player’s average score in a sample of 49 games
is 36 points with a standard deviation of 9 points. Determine the test
statistic to test whether his career scoring average is greater than 30 points.
Answer:
• Recalling the formula of test statistic which is stated in Step 2:
Test statistic = Sample statistic − Hypothesized value
Standard error of the sample statistics
• Sample statistic = Sample average score = 36
• Hypothesized value = 30
• Standard error = s/ n with s (sample standard deviation) = 9 and n = 49
X − μ0 36 − 30
→ Test statistic = s/ n = 9/ 49 = 4.67.
MODULE 8: HYPOTHESIS
TESTING 287
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Example of confidence intervals and two-tailed hypothesis tests
Example 7:
A researcher has gathered data on the daily returns on a portfolio of call
options over a recent 250-day period. The mean daily return has been 0.1%,
and the sample standard deviation of daily portfolio returns is 0.25%. The
researcher believes that the mean daily portfolio return is not equal to zero.
1. Construct a hypothesis test of the researcher’s belief at 5% significance
level.
2. Construct a 95% confidence interval for the population mean daily return
over the 250-day sample period.
Answer:
1. Approach 1: Using test statistic
Step 1: State the hypotheses
H0: μ = 0
Hα: μ ≠ 0
Step 2: Identify the appropriate statistic
We use t-distributed test statistic to test whether the population mean
is different than 0% with degrees of freedom = n – 1 = 250 – 1 = 249.
Step 3: Specify the level of significance
α = 5% (two tail)
MODULE 8: HYPOTHESIS
TESTING 288
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Example of confidence intervals and two-tailed hypothesis tests
Example 7:
1. Approach 1: Using test statistic
Step 4: State the decision rule
With df =249 (large), we approximate the critical value of t-value as z-value.
At the 5% level of significance of a two-tailed test, the critical z-values for
the confidence intervals are z0.025 = 1.96 and −z0.025 = −1.96.
→ Reject H0 if test statistic > 1.96 or test statistic < −1.96.
2.5% 2.5%
95%
−1.96 1.96
Reject H0 Not reject H0 Reject H0
Step 5: Calculate the test statistic
X − μ0 0.001 − 0
Test statistic = s/ n = 0.0025/ 250 = 6.325
Step 6: Make a decision
Test statistic = 6.325 > critical z-values = 1.96
→ Reject H0 → The evidence indicates that mean daily portfolio return is
not equal to zero.
MODULE 8: HYPOTHESIS
TESTING 289
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Example of confidence intervals and two-tailed hypothesis tests
MODULE 8: HYPOTHESIS
TESTING 290
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Example of confidence intervals and two-tailed hypothesis tests
Example 7:
2. Approach 2: Using confidence intervals
2.5% 2.5%
0.069% 0.1% 0.131%
In conclusion, there is 95% probability that population mean daily return is
ranged from 0.069% to 0.131%, which is different from 0.
MODULE 8: HYPOTHESIS
TESTING 291
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Statistical decision Economic decision
If the calculated test statistic is
The economic or investment
more extreme than the critical
decision takes into consideration
value, statistical decision is that
not only the statistical decision but
there is sufficient evidence to reject
also all pertinent economic issues.
the null hypothesis.
MODULE 8: HYPOTHESIS
TESTING 292
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Illustration 8:
An analyst is testing whether there are positive risk-adjusted returns to a
trading strategy. He collects a sample and tests the hypotheses of H0: μ ≤
0% versus Ha: μ > 0%, where µ is the population mean risk-adjusted return.
The mean risk-adjusted return for the sample is 0.7%. The calculated t-
statistic is 2.428, and the critical t-value is 2.345. He estimates that the
transaction costs are 0.3%.
t-statistic = 2.428 > critical t-values Risk-adjusted return exceeds the
= 2.345. transaction cost associated with
→ We rejected a null hypothesis this strategy by 0.4% (= 0.7 − 0.3).
and conclude that there is evidence
indicates that the mean risk-
adjusted return is greater than 0%.
This result is economically
This result is statistically significant
significant
MODULE 8: HYPOTHESIS
TESTING 293
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
6. Step 6: Make a decision
Statistically significant but not Economically significant
Although a strategy provides a statistically significant, the results may not
be economically significant when we account for transaction costs, taxes,
and risk.
Illustration 9:
Continue with previous illustration, but in case that the analysist estimates
that the transaction costs are 0.7%.
t-statistic = 2.428 > critical t-values Risk-adjusted return does not
= 2.345. exceed the transaction cost
→ We rejected a null hypothesis associated with this strategy as
and conclude that there is 0.7 − 0.7 = 0%.
evidence indicates that the mean
risk-adjusted return is greater than
0%.
This result is not economically
This result is statistically significant
significant
MODULE 8: HYPOTHESIS
TESTING 294
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Note: Beside hypothesis test, in this LOS we approach another way to test
whether H0 is rejected.
Basic characteristics of p-value
The p-value is the smallest level of significance for
which the null hypothesis can be rejected, or we can
Definition
say that it is the probability of obtaining a test statistic
that would lead to a rejection of the null hypothesis.
The p-value is the area in the probability distribution
Position
outside the calculated test statistic.
The smaller the p-value, the stronger the evidence
against the null hypothesis and in favor of the
Interpretation
alternative hypothesis.
→ The smaller the chance of making a Type I error.
P-value • Reject H0 when p-value < significance level.
approach • Do not reject H0 when p-value > significance level.
MODULE 8: HYPOTHESIS
TESTING 295
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Illustration 10: The position of p-value in each type of test
One-tailed test (right tail)
P-value represents probability
p that lies above the computed
test statistic.
Not reject Reject
Two-tailed test
P-value represents probability
α/2 α/2 that lies above the positive
value of the computed
p/2 p/2 test statistic plus the
probability that lies
below the negative value of
Reject Not reject Reject the computed test statistic.
MODULE 8: HYPOTHESIS
TESTING 296
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Example 11:
An analyst is testing the hypotheses:
H0: µ = 6
Ha: µ ≠ 6.
Using software, she determines that the p-value for the test statistic is 0.03,
or 3%. Make a decision based on p-value in case of:
(i) Level of significance α = 5%.
(ii) Level of significance α = 1%.
Answer:
We have decision based on p-value is:
Reject H0 when p-value < significance level
and Do not reject H0 when p-value > significance level.
Case 1 : Level of significance α = 5%.
p-value = 3% < significant level = 5% → Reject H0.
Case 2: Level of significance α = 1%.
p-value = 3% > significant level = 1% → Do not reject H0.
→ The null is rejected at the 5% level of significance but not at the 1% level
of significance.
MODULE 8: HYPOTHESIS
TESTING 297
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Interpret the significance of a test in the context of multiple tests
– Additional reading
Illustration 12:
Suppose we test the hypothesis that the population mean is equal to 6% at
5% significance level and repeat the sampling process by drawing 20
samples and calculating 20 test statistics. We get 6 test statistics with
lowest p-values, which are shown in table below.
Xഥ 0.0664 0.0645 0.0642 0.0642 0.0641 0.0637
Calculated
3.1966 2.2463 2.0993 2.0756 2.0723 1.8627
z-statistic
P-value 0.0014 0.0247 0.0358 0.0379 0.0382 0.0625
Reject H0? Yes Yes Yes Yes Yes No
• Using 5% level of significance, if we simply relied on each test and its p-
value, then there are 5 tests in which we would reject the null.
• Of the 20 samples tested, we should expect: 20 x 5% = 1 false positive
(this is known as type I error (level of significance), which is the
probability that reject the null hypothesis when it is actually true).
• Now we will use False discovery approach (using BH criteria) to adjust p-
value to decide the tests in which H0 is true rejected.
MODULE 8: HYPOTHESIS
TESTING 298
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Interpret the significance of a test in the context of multiple tests
– Additional reading
Illustration 12:
The process of False discovery approach is presented in the following steps:
Step 2: Calculate p-value Step 3: Compare
Step 1: Rank the
according to BH criteria: original p-value
p-values from
adjusted p-value with p-value
various test from
= α x Rank of p−value according to BH
lowest to highest
Number of tests criteria.
P-value Rank α x Rank of p−value Is p-value ≤ adjusted
Number of tests p-value?
1
0.0014 1 0.05 x = 0.0025 Yes
20
2
0.0247 2 0.05 x = 0.005 No
20
0.0358 3 0.05 x 3 = 0.0075 No
20
4
0.0379 4 0.05 x = 0.01 No
20
5
0.0382 5 0.05 x = 0.0125 No
20
6
0.0625 6 0.05 x = 0.015 No
20
MODULE 8: HYPOTHESIS
TESTING 299
[LOS 8.a] Explain hypothesis testing and its components, including
statistical significance, Type I and Type II errors,…
Interpret the significance of a test in the context of multiple tests
– Additional reading
Illustration 12:
According to the table, there is only 1 actual rejection based on comparison
of their p-values with their adjusted p-values.
→ The number of significant sample results is the same as would be
expected, given the 5% level of significance.
Conclusion:
The results for the samples with Ranks 2 to 5 are false discoveries, and we
have not uncovered any evidence from our testing that supports rejecting
the null hypothesis, or we can say that, we will not reject the null
hypothesis.
In multiple testing (testing more than one time), we can see that if we simply
relied on each test and its p-value, then there is a risk of a false discovery.
→ This is referred to as the multiple testing problem.
The false discovery approach to testing requires using BH criteria in order to
avoid misleading (data snooping bias) from multiple testing problems.
MODULE 8: HYPOTHESIS
TESTING 300
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors, and power of
the test given a significance level
Hypothesis test concerning the population mean
Purpose: The purpose of this test is to check whether there is a difference
between the population mean and hypothesized value (μ0).
Step 1: • H0: μ = μ0 versus Hα: μ ≠ μ0
State the • H0: μ ≥ μ0 versus Hα: μ < μ0
hypotheses • H0: μ ≤ μ0 versus Hα: μ > μ0
• This test statistic follows either t-distribution or z-
distribution.
• With t-distribution, we use t-statistic.
• With z-distribution, we use z-statistic.
(Condition to use each test is presented in next slide)
T-test Z-test
Step 2:
Identify the
appropriate
test statistic
X is sample mean
μ0 is hypothesized population mean
σ is population standard deviation
s is sample standard deviation
MODULE 8: HYPOTHESIS
TESTING 301
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the population mean
Step 2: Identify the appropriate test statistic (cont)
Test statistic
When sampling from a:
Small sample (n < 30) Large sample (n ≥ 30)
Normal population with Z-statistic Z-statistic
known variance
Normal population with T-statistic T-statistic (*)
unknown variance
Non-normal population Not available Z-statistic
with known variance
Non-normal population Not available T-statistic (*)
with unknown variance
* The z-statistic is theoretically acceptable here, but use of t-statistic is more
conservative.
MODULE 8: HYPOTHESIS
TESTING 302
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the population mean
• α% level of significant states that there is a α%
Step 3: Specify
probability of rejecting a true null hypothesis.
the level of
• Level of significant would be given in each exercises,
significance
but most common are 10%, 5% and 1%.
• We use appropriate table (z-table or t-table) with
Step 4: State consistent significance level to determine critical
the decision value.
rule • Reject H0 if: test statistic > upper critical value
or test statistic < lower critical value
Step 5: Test As we state in Step 2, the test statistic can be
statistic calculated using t-statistic or z-statistic.
Step 6: Make a We compare calculated Test statistic with critical value
decision and draw appropriate conclusions.
MODULE 8: HYPOTHESIS
TESTING 303
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 13: Hypothesis test of single mean
Suppose we want to test whether the daily return in the ACE High Yield
Total Return Index is different from zero. Collecting a sample of 1,304 daily
returns, we find a mean daily return of 0.0157%, with a sample standard
deviation of 0.3157%.
1. Using the t-distributed test statistic to test whether the mean daily return
is different from zero at the 5% level of significance.
2. Using the z-distributed test statistic as an approximation, test whether
the mean daily return is different from zero at the 5% level of significance.
Answer:
According to the question, as the population standard deviation is unknown
and sample size is large (n ≥ 30), we can use either z-statistic or t-statistic.
Solution 1: Using t-statistic
Step 1: State the hypotheses
H0: μ = 0%
Hα: μ ≠ 0%
Step 2: Identify the appropriate statistic
Here, we use t-distributed test statistic to test whether the population
mean is different than 0% with degrees of freedom = n – 1 = 1,304 – 1 =
1,303.
Step 3: Specify the level of significance
α = 5% (two tail)
MODULE 8: HYPOTHESIS
TESTING 304
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 13: Hypothesis test of single mean
Step 4: State the decision rule
• With df = 1,303 we use t-table for two-tailed test with significance level of
5%
The t-table
df 0.1 0.05 0.025 0.01 0.005
120 1.289 1.658 1.980 2.358 2.617
200 1.286 1.653 1.972 2.345 2.601
∞ 1.282 1.645 1.960 2.326 2.576
MODULE 8: HYPOTHESIS
TESTING 305
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 13: Hypothesis test of single mean
Step 4: State the decision rule
• With df = 1,303 we use t-table for two-tailed test with significance level of
5%
The t-table
df 0.1 0.05 0.025 0.01 0.005
120 1.289 1.658 1.980 2.358 2.617
200 1.286 1.653 1.972 2.345 2.601
∞ 1.282 1.645 1.960 2.326 2.576
MODULE 8: HYPOTHESIS
TESTING 306
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the equality of two population means
Purpose: The purpose of this test is to check if there is a difference between
the means of two populations.
Requirement: The samples are required to be independent and be taken
from two populations that are normally distributed.
Assumption: Our focus in discussing the test of the differences of means is
using the assumption that the population variances are equal.
Step 1: State • H0: μ1 − μ2 = 0 versus Hα: μ1 − μ2 ≠ 0
the • H0: μ1 − μ2 ≥ 0 versus Hα: μ1 − μ2 < 0
hypotheses • H0: μ1 − μ2 ≤ 0 versus Hα: μ1 − μ2 > 0
Step 2:
Identify the
appropriate
test statistic
MODULE 8: HYPOTHESIS
TESTING 307
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the equality of two population means
Step 3: • α% level of significant states that there is a α%
Specify the probability of rejecting a true null hypothesis.
level of • Level of significant would be given in each exercises,
significance but most common are 10%, 5% and 1%.
• We use t-table with consistent significance level to
Step 4: State
determine t-critical value.
the decision
• Reject H0 if: t-statistic > upper t-critical
rule
or t-statistic < lower t-critical
Step 5: Test
statistic
Step 6: Make We compare calculated Test statistic with critical value
a decision and draw appropriate conclusions.
MODULE 8: HYPOTHESIS
TESTING 308
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 14: Hypothesis test of the difference in means
Continue with the previous example of the returns in the ACE High Yield
Total Return Index, suppose we want to test whether there is a difference
between the mean daily returns in Period 1 and in Period 2, which is shown
in table below, using a 5% level of significance.
Period 1 Period 2
Mean 0.01775% 0.01134%
Standard deviation 0.31580% 0.38760%
Sample size 445 days 859 days
Answer:
Step 1: State the hypotheses
H0: μ1 − μ2 = 0
Hα: μ1 − μ2 ≠ 0
Step 2: Identify the appropriate statistic
Here, we use t-distributed test statistic to test whether there is a difference
between the mean daily returns in Period 1 and in Period 2 with degrees of
freedom = n1+ n2 − 2 = 445 + 859 – 2 = 1,302.
Step 3: Specify the level of significance
α = 5% (two tail)
MODULE 8: HYPOTHESIS
TESTING 309
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 14: Hypothesis test of the difference in means
Answer:
Step 4: State the decision rule
• With df = 1,302 we use t-table for two-tailed test with significance level of
5%
The t-table
df 0.1 0.05 0.025 0.01 0.005
120 1.289 1.658 1.980 2.358 2.617
200 1.286 1.653 1.972 2.345 2.601
∞ 1.282 1.645 1.960 2.326 2.576
MODULE 8: HYPOTHESIS
TESTING 310
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 14: Hypothesis test of the difference in means
Answer:
Step 6: Make a decision
t-statistic = 0.3009 < t-critical = +1.960
→ Do not reject H0
→ There is not sufficient evidence to indicate that the mean daily returns
are different for the two time periods at 5% level of significance.
MODULE 8: HYPOTHESIS
TESTING 311
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the mean difference of two dependent samples
Purpose:
• This test (also called paired comparison test) is conducted when our
samples may be dependent and the observations in the two samples both
depend on some other factor.
• The purpose of this test is to check whether the means of the differences
between observations for the two samples are different.
Requirement: Populations must be normally distributed.
Step 1: • H0: μd = μdz versus Hα: μd ≠ μdz
State the • H0: μd ≥ μdz versus Hα: μd < μdz
hypotheses • H0: μd ≤ μdz versus Hα: μd > μdz
Step 2:
Identify the
appropriate
test statistic
MODULE 8: HYPOTHESIS
TESTING 312
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Hypothesis test concerning the mean difference of two dependent samples
Step 3: • α% level of significant states that there is a α%
Specify the probability of rejecting a true null hypothesis.
level of • Level of significant would be given in each exercises, but
significance most common are 10%, 5% and 1%.
Step 4: • We use t-table with consistent significance level to
State the determine t-critical value.
decision • Reject H0 if: t-statistic > upper t-critical
rule or t-statistic < lower t-critical
Step 5: Test
statistic
Step 6:
We compare calculated Test statistic with critical value and
Make a
draw appropriate conclusions.
decision
MODULE 8: HYPOTHESIS
TESTING 313
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 15: Hypothesis test of the mean of differences
Suppose we want to compare the returns of the ACE High Yield Index with
those of the ACE BBB Index. We collect data over 1,304 days for both
indexes and calculate the means and standard deviations as shown in table
below. Using a 5% level of significance, determine whether the mean of the
differences is different from zero.
ACE index (%) ACE BBB index (%) Difference (%)
Mean return 0.0157 0.0135 −0.0021
Standard deviation 0.3157 0.3645 0.3622
Answer:
Step 1: State the hypotheses
H0: μd = 0
Hα: μd ≠ 0
Step 2: Identify the appropriate statistic
Here, we use t-distributed test statistic to test whether there is a difference
between two dependent population mean with degrees of freedom = n − 1
= 1,304 – 1 = 1,303.
Step 3: Specify the level of significance
α = 5% (two tail)
MODULE 8: HYPOTHESIS
TESTING 314
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 15: Hypothesis test of the mean of differences
Answer:
Step 4: State the decision rule
• With df = 1,302 we use t-table for two-tailed test with significance level of
5%
The t-table
df 0.1 0.05 0.025 0.01 0.005
120 1.289 1.658 1.980 2.358 2.617
200 1.286 1.653 1.972 2.345 2.601
∞ 1.282 1.645 1.960 2.326 2.576
MODULE 8: HYPOTHESIS
TESTING 315
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
Example 15: Hypothesis test of the mean of differences
Answer:
Step 6: Make a decision
t-statistic = −0.20937 > t-critical = −1.960
→ Do not reject H0
→ There is not sufficient evidence to indicate that the mean of the
differences is different from 0 at 5% level of significance.
MODULE 8: HYPOTHESIS
TESTING 316
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Recall the important features of the chi-square distribution
Three important features of the chi-
square distribution are:
• It is asymmetrical.
• It is bounded by zero. Chi-square
values cannot be negative.
• It approaches the normal
distribution in shape as the degrees
0
of freedom increase.
Chi-square test
Purpose: The purpose is to test the value of a single population variance.
Requirement: Populations must be normally distributed.
Step 1: State
the hypotheses
MODULE 8: HYPOTHESIS
TESTING 317
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
This test statistic follows chi-square distributed, which
is referred to as chi-square test.
(n − 1)s2
Step 2: Identify X2 =
σ2
the 0
appropriate with degrees of freedom = n − 1
test statistic where:
n is sample size
s2 is sample variance
σ20 is hypothesized value for the population variance
• α% level of significant states that there is a α%
Step 3: Specify
probability of rejecting a true null hypothesis.
the level of
• Level of significant would be given in each exercises,
significance
but most common are 10%, 5% and 1%.
MODULE 8: HYPOTHESIS
TESTING 318
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
Rejection regions for chi-square
• We use chi-square distribution for two-tailed test at
α = 5%
table with consistent
significance level to
determine chi-
Step 4: State square critical value.
the decision • Reject H0 if:
rule 2.5% 95% 2.5%
test statistic > upper
critical value 0
or test statistic < lower Lower Upper
critical value critical value critical value
Reject H0 Fail to reject H0 Reject H0
As we state in Step 2, test statistic is calculated as:
Step 5: Test (n − 1)s2
statistic X2 =
σ2
0
Step 6: Make a We compare calculated Test statistic with critical value
decision and draw appropriate conclusions.
MODULE 8: HYPOTHESIS
TESTING 319
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
Example 16: Hypothesis test of single variance
Suppose we are analyzing Sendar Equity Fund, a midcap growth fund that
has been in existence for 24 months. During this period, Sendar Equity
achieved a mean monthly return of 1.50% and a standard deviation of
monthly returns of 3.60%.
Using a 5% level of significance, test whether the standard deviation of
returns is less than 4%.
Answer:
Step 1: State the hypotheses
H0: σ2 ≥ 16 (%2)
Hα: σ2 < 16 (%2)
Step 2: Identify the appropriate statistic
Here, we use chi-square distributed test statistic to test whether the
standard deviation of returns is less than 4% with degrees of freedom = n
− 1 = 24 – 1 = 23.
Step 3: Specify the level of significance
α = 5% (one tail, left side)
MODULE 8: HYPOTHESIS
TESTING 320
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
Example 16: Hypothesis test of single variance
Answer:
Step 4: State the decision rule
• This is one-tailed test and we only focus on the left tail side, so the lower
critical value has a 95% probability to its right, which means 5%
probability to its left.
5% 95%
0
Lower
critical value
Reject H0 Fail to reject H0
MODULE 8: HYPOTHESIS
TESTING 321
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
Example 16: Hypothesis test of single variance
Answer:
Step 4: State the decision rule
• With df = 23 we use chi-square table for one-tailed test with significance
level of 5%.
• The chi-square values in chi-square table correspond to the probabilities
in the right tail of the distribution.
→ The lower critical value is calculated as the critical value that has a 95%
probability to its right (which effectively means 5% to its left).
The chi-square table
Probability in right tail
df
0.99 0.975 0.95 0.9 0.1 0.05 0.025 0.01 0.005
22 9.542 10.982 12.338 14.041 30.813 33.924 36.781 40.289 42.796
23 10.196 11.689 13.091 14.848 32.007 35.172 38.076 41.638 44.181
24 10.856 12.401 13.848 15.659 33.196 36.415 39.364 42.980 45.558
→ We have lower critical value = 13.091.
→ We reject H0 when Test statistic < 13.091.
MODULE 8: HYPOTHESIS
TESTING 322
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
1. Hypothesis test of single variance (Chi-square test)
Chi-square test
Example 16: Hypothesis test of single variance
Answer:
Step 5: Calculate the test statistic
2 2
X2 = (n − 1)s = (24 − 1) x (3.60%) = 18.63
σ2 (4%)2
0
Step 6: Make a decision
Test statistic = 18.63 > lower critical value = 13.091
→ Do not reject H0
→ There is not sufficient evidence to indicate that the standard deviation of
returns is less than 4% at 5% level of significance.
MODULE 8: HYPOTHESIS
TESTING 323
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
Recall the important features of the F-distribution
The important features of the F-
distribution are:
• It is skewed to the right.
• It is bounded by zero on the left.
• It is defined by two separate
degrees of freedom. (df1; df2)
0
F-test
Purpose: The purpose is to test the difference between two
population variances.
Requirement:
• Populations must be normally distributed.
• The samples are independent.
MODULE 8: HYPOTHESIS
TESTING 324
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
F-test
• H0: σ2 = σ2 versus Hα: σ2 ≠ σ2
Step 1: State 1 2 1 2
the • H0: σ2 ≤ σ2 versus Hα: σ2 > σ2
1 2 1 2
hypotheses • H0: σ2 ≥ σ2 versus Hα: σ2 < σ2
1 2 1 2
This test statistic follows F-distributed, which is referred to
as F-test.
s2
F= 1
s2
2
Step 2: dfnumerator = df1 = n1− 1
Identify the dfdenominator = df2 = n2 − 1
appropriate
where: s21, s22 are variance of two samples taken from
test statistic
population 1 and population 2.
Note: For easier to determine critical value, we put the
larger variance in the numerator when calculating the F-
test statistic.
MODULE 8: HYPOTHESIS
TESTING 325
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
F-test
Step 3: • α% level of significant states that there is a α%
Specify the probability of rejecting a true null hypothesis.
level of • Level of significant would be given in each exercises,
significance but most common are 10%, 5% and 1%.
• We use F-table with Rejection regions for F-
distribution for two-tailed test
consistent significance
level to determine at α = 5%
F-critical value.
• Since we put the larger
Step 4: variance in the
State the numerator, the F-
decision statistic is always 2.5% 95% 2.5%
rule greater than 1 so we 0
consider only the upper
Upper critical value
critical value.
→ Reject H0 if: Fail to reject H0 Reject H0
test statistic > upper
critical value
MODULE 8: HYPOTHESIS
TESTING 326
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
F-test
As we state in Step 2, F-statistic is calculated as:
Step 5: Test s2
statistic F = __1
s2
2
Step 6:
We compare calculated Test statistic with critical value
Make a
and draw appropriate conclusions.
decision
MODULE 8: HYPOTHESIS
TESTING 327
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
Example 16: Hypothesis test of the difference in variances
Annie Cower is examining the earnings for two different industries. Cower
suspects that the variance of earnings in the textile industry is different
from the variance of earnings in the paper industry. To confirm this
suspicion, Cower has looked at a sample of 31 textile manufacturers and a
sample of 41 paper companies. She measured the sample standard
deviation of earnings across the textile industry to be $4.30 and that of the
paper industry companies to be $3.80. Using a 5% significance level,
determine if the earnings of the textile industry have a different standard
deviation than those of the paper industry.
Answer:
We assume that the sample of 31 textile manufacturers refers to sample 1
and the sample of 41 paper companies refers to sample 2.
Step 1: State the hypotheses
H0: σ21 = σ22
Hα: σ21 ≠ σ22
MODULE 8: HYPOTHESIS
TESTING 328
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
Example 16: Hypothesis test of the difference in variances
Answer:
Step 2: Identify the appropriate statistic
Here, we use F-distributed test statistic to test whether the variance of
earnings for companies in the textile industry is equal to the variance of
earnings for companies in the paper industry with:
dfnumerator = df1= n1− 1 = 31 − 1 = 30
dfdenominator = df2 = n2 − 1 = 41 − 1 = 40.
Step 3: Specify the level of significance
α = 5% (two tail)
Step 4: State the decision rule
As F-table only shows the critical value for right-hand tail, we use F-table
with significance level of 2.5% to find the upper critical value.
F-distribution table: Critical value for right-hand tail equal to 0.025
df2 df1 24 25 30 40 60
30 2.14 2.12 2.07 2.01 1.94
40 2.01 1.99 1.94 1.88 1.80
→ Upper critical value = 1.94
→ With 2 degrees of freedom at level of significant = 5%, we have F-critical
value equals 1.94 → We reject H0 when F-statistic > 1.94.
MODULE 8: HYPOTHESIS
TESTING 329
[LOS 8.b] Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors,…
2. Hypothesis test concerning the equality of two variances (F-test)
Example 16: Hypothesis test of the difference in variances
Answer:
Step 5: Calculate the test statistic
s2 4.32
F = 12 = = 1.28
s 3.8 2
2
Step 6: Make a decision
F-statistic = 1.2805 < F-critical = 1.94
→ Do not reject H0
→ There is not sufficient evidence to indicate that the earnings variances of
the textile industry are significantly different from paper industry at 5%
level of significance.
MODULE 8: HYPOTHESIS
TESTING 330
[LOS 8.c] Compare and contrast parametric and nonparametric
tests, and describe situations where each is the more appropriate
type of test
1. Compare and contrast parametric and nonparametric tests
Parametric tests Nonparametric tests
A parametric test has at least one A non-parametric test is not
of the following two characteristics: concerned with a parameter, and
• It is concerned with parameters, makes only a minimal set of
or defining features of a assumptions regarding the
distribution. population.
• It makes a definite set of
assumptions.
2. Situations when to use nonparametric tests
A non-parametric test is used when:
• The data do not meet distributional assumptions.
• There are outliers.
• The data are given in ranks or use an ordinal scale.
• The hypotheses do not concern a parameter.
This LOS will be discussed further in next module – Module 9,
Parametric and non-parametric tests of independence
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF
INDEPENDENCE
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 332
LEARNING OUTCOMES
Explain parametric and nonparametric tests of the hypothesis that
9.a. the population correlation coefficient equals zero, and determine
whether the hypothesis is rejected at a given level of significance
9.b. Explain tests of independence based on contingency table data
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 333
[LOS 9.a] Compare and contrast parametric and nonparametric
tests, and describe situations where each is the more appropriate
type of test
(*) Recall from LOS 8.c, Module 8 – Hypothesis testing
1. Compare and contrast parametric and nonparametric tests (*)
Parametric tests Nonparametric tests
A parametric test has at least one of A non-parametric test is not
the following two characteristics: concerned with a parameter, and
• It is concerned with parameters, or makes only a minimal set of
defining features of a distribution. assumptions regarding the
• It makes a definite set of population.
assumptions.
2. Situations when to use nonparametric tests
A non-parametric test is used when:
• The data do not meet distributional assumptions.
• There are outliers.
• The data are given in ranks or use an ordinal scale.
• The hypotheses do not concern a parameter.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 334
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero,
and determine whether the hypothesis is rejected at a given level of
significance
Parametric test of the population correlation coefficient
Purpose:
The purpose of this test is to assess the strength of the linear relationship
(correlation) between two variables.
• H0: ρ = 0 versus Hα: ρ ≠ 0
Step 1: State
• H0: ρ ≤ 0 versus Hα: ρ > 0
the hypotheses
• H0: ρ ≥ 0 versus Hα: ρ < 0
This test statistic follows a t-distributed, which is
referred to as t-test.
Step 2: Identify r n−2
the appropriate t=
1 − r2
test statistic with degrees of freedom = n – 2
r is the sample correlation
n is the number of sample observations
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 335
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Parametric test of the population correlation coefficient
• α% level of significant states that there is a α%
Step 3: Specify
probability of rejecting a true null hypothesis.
the level of
• Level of significant would be given in each exercises,
significance
but most common are 10%, 5% and 1%.
• We use t-table with consistent significance level to
Step 4: State
determine t-critical value.
the decision
• Reject H0 if: t-statistic > upper t-critical
rule
or t-statistic < lower t-critical
As we state in Step 2, t-statistic is calculated as:
Step 5: Test r n−2
statistic t=
1 − r5
Step 6: Make a We compare calculated Test statistic with critical value
decision and draw appropriate conclusions.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 336
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Example: Hypothesis test of a correlation
An analyst is examining the annual returns for Investment One and
Investment Two over 33 years. As the analyst is most interested in
quantifying how the returns of these two series are related, so she
calculates the correlation coefficient, equal to 0.43051.
Is there a significant positive correlation between these two return series if
she uses a 1% level of significance?
Answer:
Step 1: State the hypotheses
H0: ρ ≤ 0
Hα: ρ > 0
Step 2: Identify the appropriate statistic
Here, we use t-distributed test statistic to test whether the correlation is
significantly different than 0 with degrees of freedom = n – 2 = 33 – 2 = 31.
Step 3: Specify the level of significance
α = 1% (one tail, right side)
Step 4: State the decision rule
• With df = 31 we use t-table for one-tailed test with significance level of
1%.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 337
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Example: Hypothesis test of a correlation
The t-table
df 0.1 0.05 0.025 0.01 0.005
31 1.309 1.696 2.040 2.453 2.744
32 1.309 1.694 2.037 2.449 2.738
33 1.308 1.692 2.035 2.445 2.733
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 338
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Nonparametric test of the population correlation coefficient
(Spearman rank correlation test)
Purpose:
• We use nonparametric test of the population correlation coefficient,
which is a test based on the Spearman rank correlation coefficient when
we believe that the population under consideration meaningfully departs
from normality.
• Spearman rank correlation test is calculated on the ranks of the two
variables within their respective samples.
• H0: ρ = 0 versus Hα: ρ ≠ 0
Step 1: State • H0: ρ ≤ 0 versus Hα: ρ > 0
the hypotheses • H0: ρ ≥ 0 versus Hα: ρ < 0
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 339
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Nonparametric test of the population correlation coefficient
(Spearman rank correlation test)
Step 2: Identify
the appropriate
test statistic
• α% level of significant states that there is a α%
Step 3: Specify
probability of rejecting a true null hypothesis.
the level of
• Level of significant would be given in each exercises,
significance
but most common are 10%, 5% and 1%.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 340
[LOS 9.a] Explain parametric and nonparametric tests of the
hypothesis that the population correlation coefficient equals zero...
Nonparametric test of the population correlation coefficient
(Spearman rank correlation test)
• We use t-table with consistent significance level to
Step 4: State
determine t-critical value.
the decision
• Reject H0 if: t-statistic > upper t-critical
rule
or t-statistic < lower t-critical
Step 5: Test
statistic
Step 6: Make a We compare calculated Test statistic with critical
decision value and draw appropriate conclusions.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 341
[LOS 9.b] Explain tests of independence based on
contingency table data
Hypothesis test of independence using contingency table data
Purpose: The purpose of this test is to check whether the two
characteristics in a contingency table are independent of each other when
the data is discrete (using a nonparametric test statistic that is chi-square
distributed).
• H0: two characteristics are independent.
Step 1: State
• Hα: two characteristics are not independent.
the hypotheses
→ This is always an one-tailed test (right side).
Step 2: Identify
the appropriate
test statistic
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 342
[LOS 9.b] Explain tests of independence based on
contingency table data
Hypothesis test of independence using contingency table data
• α% level of significant states that there is a α%
Step 3: Specify
probability of rejecting a true null hypothesis.
the level of
• Level of significant would be given in each exercises,
significance
but most common are 10%, 5% and 1%.
• We use chi-square table for one-tailed test with
Step 4: State
consistent significance level to determine chi-square
the decision
critical value.
rule
• Reject H0 if: test statistic > upper critical value
Step 5: Test
statistic
Step 6: Make a We compare calculated Test statistic with critical
decision value and draw appropriate conclusions.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 343
[LOS 9.b] Explain tests of independence based on
contingency table data
Example: Hypothesis test of independence
Suppose we observe the following frequency table of 1,594 exchange-
traded funds (ETFs) based on two classifications: size (that is, market
capitalization) and investment type (value, growth, or blend).
Using a 5% level of significance, test whether the two classifications in a
contingency table are independent of each other.
Investment Size based on market capitalization (j)
Total
type (i) Small (j = 1) Medium (j = 2) Large (j = 3)
Value (i = 1) 50 110 343 503
Growth (i = 2) 42 122 202 366
Blend (i = 3) 56 149 520 725
Total 148 381 1,065 1,594
Before testing the two classifications, we index our three categories of
investment type with i = 1, 2, or 3, and our three categories of size from
small to large with j = 1, 2, or 3. (as illustrated above)
→ Total of row i = r = 3
Total of column j = c = 3
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 344
[LOS 9.b] Explain tests of independence based on
contingency table data
Example: Hypothesis test of independence
Answer:
Step 1: State the hypotheses
H0: Size and Investment type are independent
Hα: Size and Investment type are not independent
Step 2: Identify the appropriate statistic
Here, we use chi-square distributed test statistic to test whether the two
classifications in a contingency table are independent of each other with
degrees of freedom = (r – 1) × (c – 1) = (3 – 1) x (3 – 1) = 4.
Step 3: Specify the level of significance
α = 5% (one tail, right side)
Step 4: State the decision rule
• With df = 4 we use chi-square table for one-tailed test with significance
level of 5%.
The chi-square table
Probability in right tail
df
0.99 0.975 0.95 0.9 0.1 0.05 0.025 0.01 0.005
3 0.1148 0.2158 0.3518 0.5844 6.251 7.815 9.348 11.345 12.838
4 0.297 0.484 0.711 1.064 7.779 9.488 11.143 13.277 14.860
5 0.554 0.831 1.145 1.610 9.236 11.070 12.832 15.086 16.750
→ We have upper critical value = 9.488.
→ We reject H0 when Test statistic > 9.488.
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 345
[LOS 9.b] Explain tests of independence based on
contingency table data
Example: Hypothesis test of independence
Answer:
Step 5: Calculate the test statistic
Firstly, we calculate the expected number of observations (expected
total for row i x total for column j
frequency) in cell i;j using formula Ei,j = total for all colums and rows
Investment Size based on market capitalization (j)
type (i) Small (j = 1) Medium (j = 2) Large (j = 3)
Value 148x503 381x503 1,065x503
E1,1 = 1,594 = E1,2 = 1,594 E1,3 = 1,594
(i = 1) 46.703 = 120.228 = 336.070
Growth 148x366 381x366 1,065x366
E2,1 = 1,594 = E2,2 = 1,594 E2,3 = 1,594
(i = 2) 33.982 = 87.482 = 244.836
Blend 148x725 381x725 1,065x725
E3,1 = 1,594 = E3,2 = 1,594 E3,3 = 1,594
(i = 3) 67.315 = 173.290 = 484.395
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 346
[LOS 9.b] Explain tests of independence based on
contingency table data
Example: Hypothesis test of independence
Answer:
Step 5: Calculate the test statistic
Secondly, we calculate the scaled squared deviation for each combination
(Oi,j − Ei,j)2
of size and investment type using formula Ei,j
Investment Size based on market capitalization (j)
type (i) Small (j = 1) Medium (j = 2) Large (j = 3)
Value (50 – 46.703)2 (110 – 120.228)2 (343 – 336.070)2
46.703 120.228 336.070
(i = 1)
= 0.233 = 0.870 = 0.143
Growth (42 – 33.982)2 (122 – 87.482)2 (202 – 244.836)2
33.982 87.482 244.836
(i = 2)
= 1.892 = 13.620 = 7.399
Blend (56 – 67.315)2 (149 – 173.290)2 (520 – 484.395)2
67.315 173.290 484.395
(i = 3)
= 1.902 = 3.405 = 2.617
MODULE 9: PARAMETRIC AND
NON-PARAMETRIC TESTS OF INDEPENDENCE 347
[LOS 9.b] Explain tests of independence based on
contingency table data
Example: Hypothesis test of independence
Answer:
Step 6: Make a decision
Test statistic = 32.081 > Upper critical value = 9.488
→ Reject H0
→ There is sufficient evidence to indicate that the Size and Investment type
of ETF are not independent at 5% level of significance.
348
HYPOTHESIS TESTING
Probability
What we want Degrees of
Test statistic distribution of the
to test freedom
statistic
Test of single
mean
Test of the
difference in
means
Test of the
mean of
differences
Test of a single
variance
Test of the
difference in
variances
Test of a
correlation
Test of
independence
(categorical
data)
MODULE 8 + 9:
HYPOTHESIS TESTING 349
KNOWLEDGE CHECKLIST
The process of hypothesis testing Hypothesis test of Difference in
□ Step 1: State the hypotheses means vs. Mean of differences
o Null hypothesis (Pair comparison test)
o Alternative hypothesis □ Purpose
□ Step 2: Identify the appropriate test □ Contrast two types of test
statistic □ Process of testing
□ Step 3: Specify the level of
significance Hypothesis test of single variance
o Type I error □ Purpose
o Type II error □ Process of testing
□ Step 4: State the decision rule
o Critical value (one-tailed vs. Hypothesis test of difference in
two-tailed test) variances
o 2 methods for making decision □ Purpose
□ Step 5: Collect data and calculate the □ Process of testing
test statistic
□ Step 6: Make a decision Hypothesis test of correlation
o Statistical decision vs. Economic □ Purpose
decision □ Process of testing
Hypothesis test of single mean Hypothesis test of independence
□ Purpose (categorical data)
□ Process of testing □ Purpose
□ Process of testing