Stat II Chapter 3
Stat II Chapter 3
HYPOTHESIS TESTING
Basic concepts
Researchers are interested in answering many types of questions. For example, a scientist
might want to know whether the earth is warming up. A physician might want to know
whether a new medication will lower a person’s blood pressure. An educator might wish to
see whether a new teaching technique is better than a traditional one. A retail merchant
might want to know whether the public prefers a certain color in a new line of fashion.
Automobile manufacturers are interested in determining whether seat belts will reduce the
severity of injuries caused by accidents. These types of questions can be addressed
through statistical hypothesis testing. In hypothesis testing, the researcher must define
the population under study, state the particular hypotheses that will be investigated, give
the significance level, select a sample from the population, collect the data, perform the
calculations required for the statistical test, and reach a conclusion.
Hypothesis and hypothesis testing defined
1
Alternative hypothesis (H1) is a statement describes what we will believe if we reject
the null hypothesis. It is designated H1 (H sub – one) it indicates existence of a difference
between a parameter and a specific value, or states that there is a difference between two
parameters. The alternative hypothesis will be accepted if the sample data provide us
with evidence that the null hypothesis is false. It is a statement that will be accepted if our
sample data provide us with ample evidence that the null hypothesis is false.
The following key points summarize the null and alternative hypotheses:
The null hypothesis represents the current belief in a situation.
The alternative hypothesis is the opposite of the null hypothesis and represents a
research claim or specific inference you would like to prove.
If you reject the null hypothesis, you have statistical proof that the alternative
hypothesis is correct.
If you do not reject the null hypothesis, you have failed to prove the alternative
hypothesis. The failure to prove the alternative hypothesis, however, does not mean
that you have proven the null hypothesis.
The null hypothesis, always refers to a specified value of the population parameter
The statement of the null hypothesis always contains an equal sign regarding the
specified value of the population parameter.
The statement of the alternative hypothesis never contains an equal sign regarding
the specified value of the population parameter.
Step II: Determine the level of significance and Find the critical value(s).
After setting up the null hypothesis and alternative hypothesis, the next step is to state
the level of significance. It is the probability of rejecting the null hypothesis when it is
actually true. Level of significance is the risk we assume of rejecting the null hypothesis
when it is actually true. The level of significance is designated by the Greek letter alpha
(), it is also referred to as the level of risk. Traditionally three levels of significance are
known
0.05. Level is selected for consumer research
0.01. for quality assurance
0.10. for political polling
A significance level of say 5 percent means that in the long run, the risk of making the
wrong decision is about 5 percent. In other words, one is likely to be committing an error
by accepting a false hypothesis or in rejecting a true hypothesis in 5 out of 100 occasions.
A significance level of, say, 1 percent implies that there is a risk of being wrong in
accepting or rejecting the hypothesis in 1 out of every 100 occasions. So a null hypothesis
that is rejected in 5% level of significance (LOS) may be accepted in 1% LOS because the
area of acceptance at 99% confidence level is more. Your choice of significance level will
be based on the criticality of the decision.
2
The probability of committing another type of error, Type II error, is designated , beta,
failure to reject Ho when it is actually false. We often refer to those two possible errors as
the alpha error , and the beta error ,
Error – the probability of making a type I error
Error – the probability of making type II error
The following table shows the decision the researcher could make and the possible
consequences.
The critical region (or rejection region) is the set of all values of the test statistic that
cause us to reject the null hypothesis. The rejection region refers to the values of the test
statistic for which we will reject the null hypothesis. A region corresponding to a statistic
which amounts to rejection of the H0 (null hypothesis) is termed as the critical region or
region of rejection. The value which separates the critical region from the acceptance
region is called the critical value (also known as the test statistic).
Step III: select the test statistic
Test statistic – A value, determined from sample information, used to reject or not to reject
the null hypothesis. Z distribution is used as test statistic when the sample size is large, n
30. When a sample is small (less than 30), then t – test are more appropriate. Based on
the sample size and the parameter to be tested the statistician will select the appropriate
test statistic.
Step IV: Determine the decision rule
A decision rule is a statement of the conditions under which the null hypothesis is rejected
and the conditions under which it is not rejected. The region or area of rejection defines
the location of all those values that are so large or so small that the probability of their
occurrence under a true null hypothesis is rather remote.
Summary table for decision rule.
Non-rejection
Region or do not reject H 0 Rejection region
3
Scale of Z 0 1.6 45
0.95 Probability 0.05 Probability
The above chart portrays the rejection region for a test of significance. The level of
significance selected is 0.05.
1. The area where the null hypothesis is not rejected includes the area to the left of 1.645
2. The area of rejection is to the right of 1.645
3. A one – tailed test is being applied
4. The 0.05 level of significant was chosen
5. The sampling distribution is for the test statistic Z , the standard normal deviate.
6. The value 1.645 separates the regions where the null hypothesis is rejected and where
it is not rejected
7. The value 1.645 is called the critical value. It is the corresponding value of the test
statistic for the selected level of significance i.e. Z value at the 0.05 level of
significance is 1.645.
Critical value: The dividing point between the region where the null hypothesis is
rejected and the region where it is not rejected.
Steps V: Take a sample and made a decision
At this step a decision is made to reject or not to reject the null hypothesis.
One – tailed and two – tailed tests
One tailed test
A one-tailed test is either a right tailed test or left-tailed test, depending on the
direction of the inequality of the alternative hypothesis. The region of rejection is only in
one tail of the curve.
Left-tailed test:
test: When the population parameter is believed to be lower than the
assumed one, the hypothesis test carried out is the left-tailed test.
Right-tailed test:
test: When the population parameter is supposed to be greater than the
assumed one, the statistical test conducted is a right-tailed test.
One way to determine the location of the rejection region is to look at the direction in
which the inequality sign in the alternate hypothesis is pointing. A test is one - tailed if H 1
state a direction. Test is one – tailed, if H 1 states >, the it is one tailed (lower/left tailed
test) or < it is one tailed ( upper tailed test).
Two-tailed test
A test is two - tailed if H1 does not state a direction. Consider the following example:
Ho: there is no difference between the mean income of males and the mean income of
females.
H1: there is a difference in the mean income of males and the mean income of females.
If Ho is rejected and H1 accepted the mean income of males could be greater than that of
females or vice versa. To accommodate these two possibilities, the 5 level of significance
representing the area of rejection is divided equally in to two tails of the sampling
distribution. If the level of significant is 0.05 each rejection region will have 0.025
probability.
Note that the total area under the normal curve is one found by 0.95 + 0.025 + 0.025.
4
Key Differences Between One-tailed and Two-tailed Test
One-tailed test, as the name suggest is the statistical hypothesis test, in which the
alternative hypothesis has a single end. On the other hand, two-tailed test implies the
hypothesis test; wherein the alternative hypothesis has dual ends.
In the one-tailed test, the alternative hypothesis is represented directionally.
Conversely, the two-tailed test is a non-directional hypothesis test.
In a one-tailed test, the region of rejection is either on the left or right of the sampling
distribution. On the contrary, the region of rejection is on both the sides of the
sampling distribution.
A one-tailed test is used to ascertain if there is any relationship between variables in a
single direction, i.e. left or right. As against this, the two-tailed test is used to identify
whether or not there is any relationship between variables in either direction.
In a one-tailed test, the test parameter calculated is more or less than the critical
value. Unlike, two-tailed test, the result obtained is within or outside critical value.
When an alternative hypothesis has ‘≠ ‘≠’ sign, then a two-tailed test is performed. In
contrast, when an alternative hypothesis has ‘> or <‘sign, then one-tailed test is
carried out.
Right (upper tail Left (lower tail test) Two tail test
test) = ≠
Is greater than Is less than Is equal to Is not equal to
Is above Is lower than Is the same as Is different from
Is higher than Is below Has changed Has not changed
from from
Is longer than Is shorter than Is the same as Is not the same as
Is bigger than Is smaller than
Is increased Is decreased or reduced
from
5
The mean of the sample was computed to be [Link] the 0.01 level of significance;
test the hypothesis that the mean is still 200.
Solution:
Step 1.1. State the null and alternative hypothesis
The null hypothesis is “The population mean is still 200 “the alternative hypothesis is “The
mean is different from 200” or "The mean is not 200” the two hypotheses are written as:
Ho: =200
H1: 200
This is a two - tailed test because the alternate hypothesis does not state the direction of
the difference. That is, it does not state whether the mean is greater than or less than
200.
Step 2: determine the level of significance
As noted the 0.01 level of significance is to be used. Since this is a two - tailed test, half of
0.01 or 0.005 is in each tail. Each rejection region will have a probability of 0.005. The
area where Ho is not rejected located between the two tails, is therefore, 0.99. 0.5000-
0.005= 0.4950 so 0.4950 is the area between 0 and the critical value. The value nearest
to 0.4950 is 0.495. The value for this probability is 2.58.
Step 3: select the test statistic
The test statistic for this type of problem is Z, because the sample is large (n=100).
X−μ
σ
Therefore, Z= √n
Compute Z
X−μ 203 .5−200 203 .5−200
=
σ 16 1.6
Z= √ n = √ 100 = 2.19
Step 4:
4: The decision rule is rejecting Ho if Z>Za
Non-rejection
Rejection region with Region or do not reject H 0 Rejection region
probability 0.99 Probability with probability 0.01÷2=0.005
0.01÷2=0.005 0.4950=0.5-0.005 0.4950=0.5-0.005
Z
It is not rejected 2.19 2.58
The decision rule is therefore: Reject the null hypothesis and accept the alternate
hypothesis if the computed value of Z does not fall in the region between +2.58 and -2.58.
Otherwise do not reject the null hypothesis.
Step 5: Make a decision
Since 2.19 do not fall in the rejection region, Ho is not rejected. So we conclude that the
difference between 203.5, the sample mean and 200 can be attributed to chance
variation.
Example 2
A researcher wishes to see if the mean number of days that a basic, low-price, small
automobile sits on a dealer’s lot is 29. A sample of 30 automobile dealers has a mean of
30.1 days for basic, low-price, small automobiles. At a = 0.05, test the claim that the
mean time is greater than 29 days. The standard deviation of the population is 3.8 days.
6
Solution
Step 1 State the hypotheses and identify the claim.
H0: =29 and H1: >29 (claim)
Step 2 Find the critical value. Since a =0.05 and the test is a right-tailed test, the critical
value is z=1.65.
Step 3 Compute the test value.
Step 4 Make the decision. Since the test value, 1.59, is less than the critical value, 1.65,
and is not in the critical region, the decision is to not reject the null hypothesis.
Example 3: The management of chain of restaurants claims that the mean waiting time
of customers for service is normally distributed with a mean of 3 minutes and a standard
deviation of one minute. The quality assurance department found a sample of 50
customers at a restaurant and that the mean waiting time was 2.75 minutes. At the 0.05
significance level test the mean waiting time is less than 3 minutes?
Step 1.1. State the null and alternative hypothesis
The null hypothesis is “The population mean is still 3 “the alternative hypothesis is “The
mean is less than 3. The two hypotheses are written as:
Ho: =3
H1: < 3
This is a left - tailed test because the alternate hypothesis states left direction.
Step 2: determine the level of significance
As noted the 0.05 level of significance is to be used. The value for this probability is -1.65.
Step 3: select the test statistic
The test statistic for this type of problem is Z, because the sample is large (n=50).
X−μ
σ
Therefore, Z = √n
Compute Z value
X−μ 2. 75−3 2 .75−3
=
σ 1 1/7 . 071
Z= √ n = √ 50 = -1.77
Step 4:
4: The decision rule is rejecting Ho if –Z <-Za
Step 5: the decision is to reject the null hypothesis.
7
In the preceding problems, we knew population standard deviation, . In most cases,
however, it is unlikely that would be known. Thus it must be estimated using the sample
X−μ
S
standard deviation, S. Then the test statistic Z = √n
Example:
A department store issues it own credit card. The credit manger wants to find out if the
mean monthly unpaid balance is more than Br. 400. The level of significance is set at 0.05.
A random check of 172 unpaid balances revealed the sample mean to be 407 and the
standard deviation of the sample 38. Should the credit manager conclude that the
population mean is greater than 400?
Solution
Ho: =400
Hi: > 400
Because Hl states a direction, a one tailed test is applied. The critical value of Z is 1.65 for
0.05 level
X−μ 407−400
S 38
Z= √ n = √ 172 = 2.42
A value of this large (2.42) will occur less than 5% of the time. So the credit manager
would reject the null hypothesis, Ho. that the mean unpaid balance is greater than 400, in
favor of H1, which states that the mean is greater than 400.
Check Your Progress –3
At the time a server was haired at a restaurant was told by the manager that she can
average more than 20 br a day in tips. Over the first 35 days she was employed at the
restaurant, the mean daily amount of her tips was 24.85 br with a standard deviation of
3.24 br. At the 0.01 significance level, can the manager conclude that she is earning more
than 20 br. per day in tips?
t= S/ √ n
t= √ Pc (1−Pc ) P c (1−P c )
n1
+
n2 = -1.530
8
X = 57
= 67
S = 10
n = 26
Step 4: The critical value
There are n -1 degrees of freedom for the test df (26-1= 25). The critical value for df = 25,
a one tailed test and 0.01 level is 2.485. The decision rule for this one tailed test is reject
Ho if the computed value of t falls in any part of the tails to the left of –2.485 otherwise do
not reject Ho.
Step 5: arrive at a decision: Because -1.530 lies in the region to the right of the critical
value –2.485 Ho is not rejected at the 0.01 level. This indicates that the cost cutting
measures have not reduced the mean cost per claim to less than 60 based on sample
results.
Check Your Progress –4
From past records it is known that the mean life of a battery used in a digital clock is 305
days. The lives of the batteries is normally distributed. The battery was recently modified
to last longer. A sample of 20 modified batteries were tested. It was discovered that the
man life was 311 days and the sample standard deviation was 12 days. At the 0.05 level
of significance, did the modification increases the mean life of the battery?
Testing for Population Proportion
A hypothesis test involving a population proportion can be considered as a binomial
experiment when there are only two outcomes and the probability of a success does not
change from trial to trial. In testing hypothesis for the population proportion the
assumptions of the binomial distribution should be met.
There are a fixed number of trials. (represented by the variable n)
Each trial has 2 outcomes (called "success" and "failure" for convenience)
The trials are all independent of one another
The probability of success is the same for all trials. (this probability is represented
by p, and q represents the probability of failure --the opposite of success)
Example: suppose prior elections in a region indicated that it is necessary for a candidate
for governor to receive at least 80% of the majority vote. The current governor is
interested in assessing his chance of returning to office and plans to have a survey
conducted consisting of 2000 registered voters.
voters The sample survey of 2000 potential
voters revealed that 1550 planned to vote for the present governor. At a=0.05 a= test if the
proportion of voters less than 80%?
Step 1: The null hypothesis Ho is that the population proportions is 0.80
The alternate hypothesis, H1 is that the proportion is less than 0.80.
The incumbent governor is concerned only when the sample proportion is less than 0.8. If
it is equal to or greater than 0.8 he will have no problem; that is the sample data would
indicate he will be probably is reelected.
Ho: P = 0.80
H1: P<0.80
Step 2: The level of significance is 0.05
The critical value is, -1.65 obtained from the Z table
Step 3: Z is the appropriate statistic
P−P
Z=
σp where P – is the population proportion and P is the sample proportion,
p is the standard error of the proportion
σP = √
p(1−p )
n so the formula for Z becomes:
9
p−p
Z=
n =2000
√ P(1−P )
n
1550
P=
2000 = 0.775
p = 0.80, the hypothesized population proportion
P−P 0. 775−0 .80
=
Z= √
√ P(1−P)/n 0 . 8(1−0 .801
2000 = -2.80
Step 4: The decision rule is therefore reject the null hypothesis and accept the alternate
hypothesis if the computed value of Z falls to the left of -1.65 otherwise do not
reject Ho.
Step 5. make a decision with reject to Ho.
The computed value of Z (-2.80) is in the rejection region. So the null hypothesis is
rejected at the 0.05 level of significance. The difference of 2.5 percentage points between
the sample (77.5) and the hypothesized population percentage (80.0) is statistically
significance. It is probably not due to sampling variation. To put it another way the
evidence at this point does not support the claim that the incumbent governor will return
to the office.
Example: The Family and Medical Leave Act provide job protection and unpaid time off
from work for a serious illness or birth of a child. In 2000, 60% of the respondents of a
survey stated that it was very easy to get time off for these circumstances. A researcher
wishes to see if the percentage who said that it was very easy to get time off has changed.
A sample of 100 people who used the leave said that 53% found it easy to use the leave.
At a =0.01, has the percentage changed?
Example: A dietitian claims that 60% of people are trying to avoid trans fats in their diets.
She randomly selected 200 people and found that 128 people stated that they were trying
to avoid trans fats in their diets. At a =0.05, is there enough evidence to reject the
dietitian’s claim?
Check Your Progress –5
This Claim is to be investigated at the 0.02 level “Forty percent of those persons who
retired from an industrial job before the age of 60 would return to work if a suitable job
were available” 74 persons out of the 200 sampled said they would return to work. Can we
conclude that the fraction returning to work is different from 0.40?
Z test for the difference between two means; Independent population
The basic concepts of hypothesis testing were explained with the z, t, tests, a sample
mean, or proportion can be compared to a specific population mean or proportion to
determine whether the null hypothesis should be rejected. There are, however, many
instances when researchers wish to compare two sample means, using experimental and
control groups. For example, the average lifetimes of two different brands of bus tires
might be compared to see whether there is any difference in tread wear. Two different
brands of fertilizer might be tested to see whether one is better than the other for growing
plants. Or two brands of drugs might be tested to see whether one brand is more effective
than the other.
Assumption for two-sample test
1. The population should be normally distributed
2. The population standard deviations for both populations should be known. If they are
not known, then both samples should contain at least 30 observations so that the
10
sample standard deviation can be used to approximate the population standard
deviation
3. The samples should be drawn from independent population.
If we select random samples from two normal populations the distribution of the
differences between the two means is also normal or if a large number of independent
random samples are selected from two populations, the difference between the two
means will be normally distributed. If these differences are divided by the standard error
of the difference, the result is the standard normal distribution.
The formula for the test statistic Z is
x 1 −x 2
√
S S
12 22
+
Therefore, Z = n1 n2
These tests can also be one-tailed or two tailed test
Right tailed Left tailed Two tailed
Ho: 1 = 2 Ho: 1 = 2 Ho: 1 = 2
H1: 1 > 2 H1: 1 < 2 H1: 1 ≠ 2
11
Patient type Sample mean
mean Sample standard Deviation
Sample Size
Senior Citizens 5.5 Minutes 0.40 minuets 50
Other 5.3 Minutes 0.30 minutes 100
Solution:-
The testing procedure is the same as for one sample test except the formula for the test
statistic, Z:
Step 1: Ho: there is no difference in the mean response time between the two groups of
patients. i: e The difference of 0.2 minute, in the arithmetic mean response time is due to
chances.
H1: the mean response time is greater for the senior citizens
Because the quality assurance department is concerned that the response time is greater
for senior citizens, he wants to conduct a one – tailed test. Therefore the null and alternate
hypotheses are stated as follows.
Ho: 1 = 2
H1: 1 > 2
Step 2: The 0.01 significance level is selected.
x 1 −x 2
√
S S
12 22
+
Step 3: the test statistics is Z, the standard normal distribution, Z = n1 n2
Step 4: The decision rule is:
Reject the null hypothesis if the computed value of Z is greater than 2.33.
Step 5: Calculate the test statistic and make a decision.
x 1 −x 2
√
S S
12 22
+
The test statistic is Z = n1 n2
5 .5−5 .3
√
(0 . 40 )2 (0 .30 )2
Z = 50
+
100 = 3.13
The computed value of 3.13 is beyond the critical value of 2.33. Therefore, the null
hypothesis is rejected and the alternate hypothesis is accepted at the 0.01 significant
level. The quality assurance department will report to the administrator that the mean
response time of the nurses and resident physicians is longer for senior citizens than for
other patients.
Exercise: A survey found that the average hotel room rate in New Orleans is $88.42 and
the average room rate in Phoenix is $80.61. Assume that the data were obtained from two
samples of 50 hotels each and that the standard deviations of the populations are $5.62
and $4.83, respectively. At 0.05, can it be concluded that there is a significant difference
in the rates?
t-test for the difference between two means independent samples
The z test was used to test the difference between two means when the population
standard deviations were known and the variables were normally or approximately
normally distributed, or when both sample sizes were greater than or equal to 30. In many
situations, however, these conditions cannot be met that is, the population standard
deviations are not known. In these cases, a t test is used to test the difference between
means when the two samples are independent and when the samples are taken from two
normally or approximately normally distributed populations. Samples are independent
12
samples when they are not related. Also it will be assumed that the variances are not
equal.
P 1−P 2
Z=
√
Pc (1−Pc ) P c (1−P c )
n1
+
n2
N1 is the number of observations in the first sample.
N2 is the number of observations in the second sample.
P1 is the proportion in the first sample possessing the trait.
P2 is the proportion in the second sample possessing the trait.
Pc is the pooled proportion possessing the trait in the combined samples. It is
called the pooled estimate of the population proportion and is computed from the
following formula
x1 + x 2
Pc n1 +n2
=
Example: - a company has developed a new perfume. One of the questions is whether
the perfume is preferred by a larger proportion of younger women or larger proportions of
13
older women. A standard smell test is used. Women selected at random are asked to
sniff several perfumes in succession, including the new. Each woman selects the perfume
she likes best. A total of 100 young women selected at random, and each was given the
standard smell test. Forty of the 100 young women chose the perfume, as they liked best.
200 older women were selected at random and each was given the same standard smell
test of the 200 women 100 preferred the perfume. At a= 0.05 is there is a difference in
the proportion of younger and older women who prefer the perfume?
Step 1
Ho “There is no difference between the proportion of younger women who prefer the
perfume and the proportion of older women who prefer it” If the proportion of younger
women in the population is designated as P 1 and the proportion of older women is P 2 then;
Ho: P1= P2
The alternate hypothesis is that the two proportions are not equal or:
Hi: P1 P2
Step 2: It was decided to use the 0.05 level.
Step 3: The test statistic is Z and the formula is:
P 1−P 2 Where: n1, is the number of young women selected in
the sample n2 is the number of older women selected
Z= √ Pc (1−Pc ) P c (1−P c )
n1
+
n2
P
in the sample, c = is the weighted mean of the two
sample proportion computed by
Total number of successe x1 + x 2
Pc =Total number of samples = n1 +n2
Where x1 is the number of younger women
(sample 1) who prefer the perfume, x 2 is the number
of older women (sample 2) who prefer the perfume.
Pc Is generally referred to as the pooled
estimate of the population proportion or it is a
combined estimate, combined proportion.
X1 =40
n1=100 x 1 40
X2 = 100 P1 = = =0 . 40
n2=200 n1 100
x 100
P2 = 2 = =0. 50
Pc is n2 200
The pooled or weighted proportion
x1 + x 2 40+ 100
Pc = n1 +n2 = 100+200 = 140 / 300 = 0.4667
P 1−P2 0 . 40−0 . 50
= =−1 . 64
Z= √
Pc (1−Pc ) P c (1−P c )
n 1
+
n 2
100 √
0 . 4667(0 . 5333) 0 . 4667 (0 .5333 )
+
200
14
Two – tailed test, Areas of rejection and Non-rejection 0.05 level of significance.
Step 5: The decision
The computed value of Z (-1.64) falls in the non-rejection region. Therefore we concluded
that there is no difference in the proportion of younger and older women who prefer the
perfume.
Example: A researchers found that 12 out of 34 small nursing homes had a resident
vaccination rate of less than 80%, while 17 out of 24 large nursing homes had a
vaccination rate of less than 80%. At a = 0.05, test the claim that there is no difference in
the proportions of the small and large nursing homes.
15