HYPOTHESIS TESTING
Language of Hypothesis Testing
What is a hypothesis test?
A hypothesis test uses a sample of data in an experiment to test a statement made
about the value of a population parameter
A hypothesis test is used when the value of the assumed population parameter is
questioned
The hypothesis test will look at the which outcomes are unlikely to occur if assumed
population parameter is true
The probability found will be compared against a given significance level to determine
whether there is evidence to believe that the assumed population parameter is not true
What are the key terms used in statistical hypothesis testing?
Every hypothesis test must begin with a clear null hypothesis (what we believe to
already be true) and alternative hypothesis (how we believe the data pattern or
probability distribution might have changed)
A hypothesis is an assumption that is made about a particular population parameter
- A population parameter is a numerical characteristic which helps define a population
-One example of a population parameter is the probability, p of an event occurring
-Another example is the mean of a population
-The null hypothesis is denoted H0 and sets out the assumed population parameter
given that no change has happened
-The alternative hypothesis is denoted H1 and sets out how we think the population
parameter could have changed
-When a hypothesis test is carried out, the null hypothesis is assumed to be true and
this assumption will either be accepted or rejected
A hypothesis test could be a one-tailed test or a two-tailed test
The null hypothesis will always be H0 : θ = ...
The alternate hypothesis will depend on if it is a one-tailed or two-tailed test
-A one-tailed test would test to see if the population parameter, θ, has either increased
or decreased
-The alternative hypothesis, H1 will be H1 : θ > ... or H1 : θ < ...
-A two-tailed test would test to see if the population parameter, θ , has changed
-The alternative hypothesis, H1 will be H1 : θ ≠ ...
-It is important to read the wording of the question carefully to decide whether your
hypothesis test should be one-tailed or two-tailed
To carry out a hypothesis test an experiment will be carried out on a sample of data, the
result of this experiment will be the observed value
- A sample of data is a subset of data taken from the population
- The observed value is a numerical value calculated from the of data
A hypothesis test will always be carried out at an appropriate significance level
- The significance level sets the smallest probability that an event could have occurred
by chance.
-Any probability smaller than the significance level would suggest that the event is
unlikely to have happened by chance
- The significance level must be set before the hypothesis test is carried out
- The significance level will usually be 1%, 5% or 10%, however it may vary
Worked example
A hypothesis test is carried out at the 5% level of significance to test if a normal coin is fair or
not.
(i) Describe what the population parameter could be for the hypothesis test.
(ii) State whether the hypothesis test should be a one-tailed test or a two-tailed test,
give a reason for your answer.
(iii) (iii) Clearly defining your population parameter, state suitable null and alternative
hypotheses for the test.
(iv)
Critical Regions
How do we decide whether to reject or accept the null hypothesis?
The null hypothesis would be rejected if the observed value falls within the critical
region
- The critical region is the range of values that the observed value could take which will
lead to the null hypothesis being rejected
The critical value is the boundary of the critical region
- It is the least extreme value that would lead to the rejection of the null hypothesis
- The critical value is determined by the significance level
- In a two-tailed test the significance level is halved and both the upper and the lower
tails are tested
- For discrete distributions the critical value is the first value that falls within the critical
region and so the probability of the observed value falling within the critical region may
be lower than the given significance level
- This probability will be known as the actual significance level
- The actual significance level is the probability of incorrectly rejecting the null
hypothesis
- Finding the critical region will be different for a two-tailed test than it is for a one-
tailed test
For an α% significance level
- In a one-tailed test the critical region will consist of alpha% in the tail that is being
tested for
- In a two-tailed test the critical region will consist of alpha over 2 percent sign in each
tail
Do we always need to find the critical region?
In most cases the best method of conducting a hypothesis test is to find the critical
region
- It allows you to see how far the observed value is from the critical value and make
decisions about whether further testing is necessary
In some cases a hypothesis test can be carried out without finding the critical region
The null hypothesis would be rejected if the probability of a value being at least as
extreme as the observed value, assuming that the null hypothesis is true, is less than the
significance level
- If the test is looking for a decrease then extreme values are smaller than the observed
value, so find the probability of less than or equal to the observed value
- If the test is looking for an increase then extreme values are bigger than the observed
value, so find the probability of greater than or equal to the observed value
- This probability is called the "p-value"
In a two-tailed test it is common to half the significance level and compare this with the
probability found in one of the tails
Worked Example
For the following situations, state at the 1% and 5% significance levels whether the null
hypothesis should be rejected or not.
(i) The critical region is X≤3and the observed value is 4.
(ii) Assuming the null hypothesis is true, the probability of a value being at least as
extreme as the test statistic in a one-tailed hypothesis test is 0.0429.
(iii) Assuming the null hypothesis is true, the probability of a value being at least as
extreme as the test statistic in a two-tailed hypothesis test is 0.00705.
Conclusions of Hypothesis Testing
How is a hypothesis test carried out?
There are a number of ways that a hypothesis test can be carried out for different models,
however the following steps should form the base for your test:
Step 1: Define the test statistic and population parameter
Step 2: Write the null and alternative hypotheses clearly
Step 3: Calculate the critical value(s) or the necessary probability for the test
Step 4: Compare the observed value with the critical value(s) or the probability with the
significance level
Step 5: Decide whether there is enough evidence to reject H0 or whether it has to be
accepted
Step 6: Write a conclusion in context
How should a conclusion be written for a hypothesis test?
Your conclusion must be written in the context of the question
Use the wording in the question to help you write your conclusion
- If rejecting the null hypothesis your conclusion should state that there is sufficient
evidence to suggest the alternative hypothesis is true at this level of significance
- If accepting the null hypothesis your conclusion should state that there is not
enough evidence to suggest the alternative hypothesis is true at this level of
significance
Your conclusion must not be definitive
- There is a chance that the test has led to an incorrect conclusion
- The outcome is dependent on the sample, a different sample might lead to a
different outcome
The conclusion of a two-tailed test can state if there is evidence of a change
- You should not state whether this change is an increase or decrease
Worked Example
A teacher carried out a hypothesis test at the 10% significance level to test if her students
perform better in exams after using a new revision technique. Under the null hypothesis she
calculates the probability that a value will be at least as extreme as the observed value to be
0.09142. Write a conclusion for her hypothesis test.
Type I & Type II Errors
Any hypothesis test will only provide evidence about whether a parameter has changed or not.
A conclusion can not claim with certainty whether to accept or reject the null hypothesis as the
test is based on probability, and therefore errors are possible.
What is a Type I error?
A Type I error occurs when the null hypothesis is rejected incorrectly
- In order for a Type I error to happen, the null hypothesis must have been rejected
If a Type I error has been made, the hypothesis test has provided evidence that there is
a change when in fact there is not a change
- Think about the impact of this in some scenarios
- For example a test saying that a student had cheated in an exam when in fact they
had not
The probability of a Type I error occurring in any hypothesis test is the same as the
probability of rejecting a true null hypothesis
- This is the probability of the observed value being at least as extreme as the critical
value(s)
- It is the same or a little bit less than the significance level
In a true hypothesis test you would not need to calculate the probability of a Type I
error as it would be the same as the actual significance level
What is a Type II error?
A Type II error occurs when the null hypothesis is accepted incorrectly
- In order for a Type II error to happen, the null hypothesis must not have been rejected
If a Type II error has been made, the hypothesis test has provided evidence that there is
no change when in fact there was a change
- Think about the impact of this in some scenarios
- For example a test saying that a car’s brakes have not worn down, when in fact they
have
To find probability of a Type II error occurring in any hypothesis test you would need to
be given the true value of the population parameter being tested
- For example, you would be given the true probability of the event occurring or the true
population mean
- The probability of a Type II error would be the probability of the observed value being
outside of the rejection region, given the true value of the population parameter
REJECT H0 ACCEPT H1
H0 TRUE TYPE I No Error
H0 FALSE No Error TYPE II
Can the probabilities of making the errors be manipulated?
It is possible to reduce the probability of making a Type I error by reducing the
significance level before carrying out the test
- However, this would decrease the size of the rejection region and therefore could
increase the probability of a Type II error
It is possible to reduce the probability of making a Type II error by increasing the
significance level before carrying out the test
- This would increase the size of the rejection region, making it easier to reject the null
hypothesis
- As the probability of rejecting the null hypothesis has increased, this would increase
the probability of making a Type I error
Before setting the significance level a researcher could consider which error they would
want to reduce the likelihood of
- For example, if the test is for a company advertising that their product works 90% of
the time, but customers believe it may be less than this:
- the company would want to reduce the probability of a Type I error (incorrectly
declaring a change)
- the customers would want to reduce the probability of a Type II error (incorrectly
declaring no change)
Here are two tips if you cannot remember which error is which but are asked to
calculate one on the exam:
Look to see if you are given a new population parameter, this will be a Type II error.
Check the number of marks, a Type I error is normally only 1 mark whilst a Type II error
needs to be calculated and so will be more.
Worked Example
In the following scenarios, decide whether a Type I error or Type II error could have occurred
(i) A farmer is testing for a change in crop growth after trying a new fertiliser. The test
concludes that there is no evidence of change at the 5% significance level.
(ii) A dentist’s receptionist believes that the waiting times have been reduced due to a
new scheduling system. They conduct a hypothesis test and will reject the null
hypothesis if no more than two customers wait more than ten minutes. Exactly two
customers have to wait more than ten minutes.
Binomial Distribution
How is a hypothesis test carried out with the binomial distribution?
The population parameter being tested will be the probability, p in a binomial
distribution B(n , p)
A hypothesis test is used when the assumed probability is questioned
The null hypothesis, H0 and alternative hypothesis, H1 will always be given in terms of p
- Make sure you clearly define p before writing the hypotheses
- The null hypothesis will always be H0 : p = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed tes
- A one-tailed test would test to see if the value of p has either increased or decreased
- The alternative hypothesis, H1 will be H1 : p > ... or H1 : p < ...
- A two-tailed test would test to see if the value of p has changed
- The alternative hypothesis, H1 will be H1 : p ≠ ...
To carry out a hypothesis test with the binomial distribution, the random variable used
in the test will be the number of successes in a defined number of trials
When defining the distribution, remember that the value of p is being tested, so this
should be written as p in the original definition, followed by the null hypothesis stating
the assumed value of p
The binomial distribution will be used to calculate the probability of the test statistic
taking the observed value or a more extreme value
The hypothesis test can be carried out by
- either calculating the probability of the test statistic taking the observed or a more
extreme value (p – value) and comparing this with the significance level
- or by finding the critical region and seeing whether the observed value of the test
statistic lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing
How is the critical value found in a hypothesis test with the binomial distribution?
The critical value will be the first value to fall within the critical region
- The binomial distribution is a discrete distribution so the probability of the observed
value being within the critical region, given a true null hypothesis may be less than the
significance level
- This is the actual significance level and is the probability of incorrectly rejecting the null
hypothesis
For a one-tailed test use your calculator to find the first value for which the probability
of that or a more extreme value is less than the given significance level
- Check that the next value would cause this probability to be greater than the
significance level
- For H1 : p < ... if straight P left parenthesis X less or equal than c right parenthesis
less or equal than alpha percent sign and straight P left parenthesis X less or equal than
c plus 1 right parenthesis less or equal than alpha percent sign then c is the critical value
- For H1 : p > ... if straight P left parenthesis X greater or equal than c right
parenthesis less or equal than alpha percent sign and straight P left parenthesis X
greater or equal than c minus 1 right parenthesis greater than alpha percent sign then c
is the critical value
For a two-tailed test you will need to find both critical values, one at each end of the
distribution
- Find the first value for which the probability of that or a more extreme value is less
than half of the given significance level in both the upper and lower tails
- Often one of the critical regions will be much bigger than the other
- If the probability in the null hypothesis is 0.5 the critical regions will have an equal
size
What steps should I follow when carrying out a hypothesis test with the binomial
distribution?
Step 1: Define the probability, p
Step 2: Write the null and alternative hypotheses clearly using the form
H0 : p = ...
H1 : p = ...
Step 3: Define the distribution, usually X tilde B left parenthesis n comma p right parenthesis
where n is a defined number of trials and p is the population parameter to be tested
Step 4: Calculate either the critical value(s) or the necessary probability for the test
Step 5: Compare the observed value of the test statistic with the critical value(s) or the
probability with the significance level
Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted
Step 7: Write a conclusion in context
Worked Example
Jacques, a breadmaker, claims that more than 60% of people that shop in a particular
supermarket buy his brand of bread. Jacques takes a random sample of 12 customers that have
purchased bread and asks them which brand of bread they have purchased. He records that 10
of them had purchased his brand of bread. Test, at the 10% level of significance, whether
Jacques’ claim is justified.
Poisson Distribution
How is a hypothesis test carried out for the mean of a Poisson distribution?
The population parameter being tested will be the mean, λ , in a Poisson distribution
- As it is the population mean, sometimes μ will be used instead
A hypothesis test is used when the mean is questioned
The null hypothesis, H0 and alternative hypothesis, H1 will be given in terms of λ (or μ)
- Make sure you clearly define λ before writing the hypotheses
- The null hypothesis will always be H0 : λ = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed test
- A one-tailed test would test to see if the value of λ has either increased or decreased
- The alternative hypothesis, will be H1 will be H1 : λ > ...or H1 : λ < ...
- A two-tailed test would test to see if the value of λ has changed
- The alternative hypothesis, H1 will be H1 : λ ≠ ...
To carry out a hypothesis test with the Poisson distribution, the random variable will be
the mean number of occurrences of the event within the given time/space interval
- Remember you may need to change the mean to fit the interval of time or space for
your observed value
When defining the distribution, remember that the value of λ is being tested, so this
should be written as λ in the original definition, followed by the null hypothesis stating
the assumed value of λ
The Poisson distribution will be used to calculate the probability of the random variable
taking the observed value or a more extreme value
The hypothesis test can be carried out by
- either calculating the probability of the random variable taking the observed or a more
extreme value and comparing this with the significance level
- or by finding the critical region and seeing whether the observed value of the test
statistic lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing
How is the critical value found in a hypothesis test with the Poisson distribution?
The critical value will be the first value to fall within the critical region
- The Poisson distribution is a discrete distribution so the probability of the observed
value being within the critical region, given a true null hypothesis may be less than the
significance level
- This is the actual significance level and is the probability of incorrectly rejecting the null
hypothesis (a Type I error)
For a one-tailed test use the formula to find the first value for which the probability of
that or a more extreme value is less than the given significance level
- Check that the next value would cause this probability to be greater than the
significance level
- For H1 : λ <… if P(X ≤ c) ≤ α% and P(X ≤ c+1) > α% then c is the critical value
- For H1 : λ >… if P( X ≥ c) ≤ α% and P(X ≥ c-1) > α % then c is the critical value
- Using the formula for this can be time consuming so only use this method if you need
to
- otherwise compare the probability of the random variable being at least as extreme
as the observed value with the significance level
For a two-tailed test you will need to find both critical values, one at each end of the
distribution
Take extra care when finding the critical region in the upper tail, you will have to find
the probabilities for less than and subtract from one
What steps should I follow when carrying out a hypothesis test with the Poisson
distribution?
Step 1: Define the mean, λ
Step 2: Write the null and alternative hypotheses clearly using the form
H0 : λ = ...
H1 : λ = ...
Step 3: Define the distribution, usually begin mathsize 16px style X tilde Po left parenthesis
lambda right parenthesis end style where λ is the mean to be tested
Step 4: Calculate the probability of the random variable being at least as extreme as the
observed value
- Or if told to find the critical region
Step 5: Compare this probability with the significance level
- Or compare the observed value with the critical region
Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted
Step 7: Write a conclusion in context
Worked Example
Mr Viajo believes that his travel blog receives an average of 8 likes per day (24 hour period). He
tries a new advertising campaign and carries out a hypothesis test at the 5% level of significance
to see if there is an increase in the number of likes he gets. Over a 6-hour period chosen at
random Mr Viajo’s travel blog receives 5 likes.
(i) State null and alternative hypotheses for Mr Viajo’s test.
(ii) Find the rejection region for the test.
(iii) Find the probability of a Type I error.
(iv) Carry out the hypothesis test, writing your conclusion clearly.
Normal Hypothesis Testing
What steps should I follow when carrying out a hypothesis test for the mean of a
normal distribution?
Following these steps will help when carrying out a hypothesis test for the mean of a normal
distribution:
Step 1: Define the distribution of the population mean usually X~ N (μ, σ2)
Step 2: Write the null and alternative hypotheses clearly
Step 3: Assuming the null hypothesis to be true, define the statistic
Step 4: Calculate either the critical value(s) or the probability of the observed value for the test
Step 5: Compare the observed value of the test statistic with the critical value(s) or the
probability with the significance level
- Or compare the z-value corresponding to the observed value with the z-value corresponding
to the critical value
Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted
Step 7: Write a conclusion in context
How should I define the distribution of the population mean and the statistic?
The population parameter being tested will be the population mean, μ in a normally
distributed random variable N (μ, σ2)
How should I define the hypotheses?
A hypothesis test is used when the value of the assumed population mean is questioned
The null hypothesis, H0 and alternative hypothesis, H1 will always be given in terms of µ
- Make sure you clearly define µ before writing the hypotheses, if it has not been
defined in the question
- The null hypothesis will always be H0 : µ = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed test
- A one-tailed test would test to see if the value of µ has either increased or decreased
- The alternative hypothesis, H1 will be H1 : µ > ... or H1 : µ < ...
- A two-tailed test would test to see if the value of µ has changed
- The alternative hypothesis, H1 will be H1 : µ ≠ ..
How do I define the statistic?
The population mean is tested by looking at the mean of a sample taken from the
population
- The sample mean is denoted
- For a random variable ~ N (μ, σ2) the distribution of the sample mean would be
~ N (μ, σ2/n)
To carry out a hypothesis test with the normal distribution, the statistic used to carry
out the test will be the sample mean,
- Remember that the variance of the sample mean distribution will be the variance of
the population distribution divided by n
- the mean of the sample mean distribution will be the same as the mean of the
population distribution
How should I carry out the test?
The hypothesis test can be carried out by
- either calculating the probability of a value taking the observed or a more extreme
value and comparing this with the significance level
- The normal distribution will be used to calculate the probability of a value of the
random variable taking the observed value or a more extreme value
- or by finding the critical region and seeing whether the observed value lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing
A third method is to compare the z-values of your observed value with the z-values
at the boundaries of the critical region(s)
̅
- Find the z-value for your sample mean using
√
- This is sometimes known as your test statistic
- Use the table of critical values to find the z-value for the significance level
- If the z-value for your test statistic is further away from 0 than the critical z-value
then reject H0
How is the critical value found in a hypothesis test for the mean of a normal
distribution?
The critical value(s) will be the boundary of the critical region
- The probability of the observed value being within the critical region, given a true null
hypothesis will be the same as the significance level
For an α% significance level
- In a one-tailed test the critical region will consist of α% in the tail that is being tested
- In a two-tailed test the critical region will consist of α/2% in each tail
To find the critical value(s) use the standard normal distribution:
Step 1: Find the distribution of the sample means, assuming H0 is true
̅
Step 2: Use the coding to standardize to z
√
Step 3: Use the table to find the z - value for which the probability of Z being equal to or more
extreme than the value is equal to the significance level
- You can often find this in the table of the critical values
Step 4: Equate this value to your expression found in step 2
Step 5: Solve to find the corresponding value of
If using this method for a two-tailed test be aware of the following:
The symmetry of the normal distribution means that the z - values will have the same absolute
value
You can solve the equation for both the positive and negative z – value to find the two critical
values
Check that the two critical values are the same distance from the mean
Worked Example
The time, minutes, that it takes Amelea to complete a 1000-piece puzzle can be modelled using
X~N(204, 81). Amelea gets prescribed a new pair of glasses and claims that the time it takes her
to complete a 1000-piece puzzle has decreased. Wearing her new glasses, Amelea completes
12 separate 1000-piece puzzles and calculates her mean time on these puzzles to be 201
minutes. Use these 12 puzzles as a sample to test, at the 5% level of significance, whether
there is evidence to support Amelea’s claim. You may assume the variance is unchanged.