0% found this document useful (0 votes)
9 views23 pages

Understanding Hypothesis Testing Basics

Hypothesis testing involves using sample data to evaluate a statement about a population parameter, starting with a null hypothesis (H0) and an alternative hypothesis (H1). The test assesses the likelihood of observed outcomes under the assumption that H0 is true, comparing probabilities against a significance level to determine if H0 should be rejected. Key concepts include critical regions, Type I and Type II errors, and the application of binomial distribution for hypothesis tests involving probabilities.

Uploaded by

7wqf8smmx8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views23 pages

Understanding Hypothesis Testing Basics

Hypothesis testing involves using sample data to evaluate a statement about a population parameter, starting with a null hypothesis (H0) and an alternative hypothesis (H1). The test assesses the likelihood of observed outcomes under the assumption that H0 is true, comparing probabilities against a significance level to determine if H0 should be rejected. Key concepts include critical regions, Type I and Type II errors, and the application of binomial distribution for hypothesis tests involving probabilities.

Uploaded by

7wqf8smmx8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

HYPOTHESIS TESTING

Language of Hypothesis Testing


What is a hypothesis test?

 A hypothesis test uses a sample of data in an experiment to test a statement made


about the value of a population parameter
 A hypothesis test is used when the value of the assumed population parameter is
questioned
 The hypothesis test will look at the which outcomes are unlikely to occur if assumed
population parameter is true
 The probability found will be compared against a given significance level to determine
whether there is evidence to believe that the assumed population parameter is not true

What are the key terms used in statistical hypothesis testing?

 Every hypothesis test must begin with a clear null hypothesis (what we believe to
already be true) and alternative hypothesis (how we believe the data pattern or
probability distribution might have changed)
 A hypothesis is an assumption that is made about a particular population parameter
- A population parameter is a numerical characteristic which helps define a population
-One example of a population parameter is the probability, p of an event occurring
-Another example is the mean of a population
-The null hypothesis is denoted H0 and sets out the assumed population parameter
given that no change has happened
-The alternative hypothesis is denoted H1 and sets out how we think the population
parameter could have changed
-When a hypothesis test is carried out, the null hypothesis is assumed to be true and
this assumption will either be accepted or rejected
 A hypothesis test could be a one-tailed test or a two-tailed test
 The null hypothesis will always be H0 : θ = ...
 The alternate hypothesis will depend on if it is a one-tailed or two-tailed test
-A one-tailed test would test to see if the population parameter, θ, has either increased
or decreased
-The alternative hypothesis, H1 will be H1 : θ > ... or H1 : θ < ...
-A two-tailed test would test to see if the population parameter, θ , has changed
-The alternative hypothesis, H1 will be H1 : θ ≠ ...
-It is important to read the wording of the question carefully to decide whether your
hypothesis test should be one-tailed or two-tailed
 To carry out a hypothesis test an experiment will be carried out on a sample of data, the
result of this experiment will be the observed value
- A sample of data is a subset of data taken from the population
- The observed value is a numerical value calculated from the of data
 A hypothesis test will always be carried out at an appropriate significance level
- The significance level sets the smallest probability that an event could have occurred
by chance.
-Any probability smaller than the significance level would suggest that the event is
unlikely to have happened by chance
- The significance level must be set before the hypothesis test is carried out
- The significance level will usually be 1%, 5% or 10%, however it may vary
Worked example
A hypothesis test is carried out at the 5% level of significance to test if a normal coin is fair or
not.

(i) Describe what the population parameter could be for the hypothesis test.
(ii) State whether the hypothesis test should be a one-tailed test or a two-tailed test,
give a reason for your answer.
(iii) (iii) Clearly defining your population parameter, state suitable null and alternative
hypotheses for the test.

(iv)
Critical Regions
How do we decide whether to reject or accept the null hypothesis?

 The null hypothesis would be rejected if the observed value falls within the critical
region
- The critical region is the range of values that the observed value could take which will
lead to the null hypothesis being rejected
 The critical value is the boundary of the critical region
- It is the least extreme value that would lead to the rejection of the null hypothesis
- The critical value is determined by the significance level
- In a two-tailed test the significance level is halved and both the upper and the lower
tails are tested
- For discrete distributions the critical value is the first value that falls within the critical
region and so the probability of the observed value falling within the critical region may
be lower than the given significance level
- This probability will be known as the actual significance level
- The actual significance level is the probability of incorrectly rejecting the null
hypothesis
- Finding the critical region will be different for a two-tailed test than it is for a one-
tailed test
 For an α% significance level
- In a one-tailed test the critical region will consist of alpha% in the tail that is being
tested for
- In a two-tailed test the critical region will consist of alpha over 2 percent sign in each
tail

Do we always need to find the critical region?

 In most cases the best method of conducting a hypothesis test is to find the critical
region
- It allows you to see how far the observed value is from the critical value and make
decisions about whether further testing is necessary
 In some cases a hypothesis test can be carried out without finding the critical region
 The null hypothesis would be rejected if the probability of a value being at least as
extreme as the observed value, assuming that the null hypothesis is true, is less than the
significance level
- If the test is looking for a decrease then extreme values are smaller than the observed
value, so find the probability of less than or equal to the observed value
- If the test is looking for an increase then extreme values are bigger than the observed
value, so find the probability of greater than or equal to the observed value
- This probability is called the "p-value"
 In a two-tailed test it is common to half the significance level and compare this with the
probability found in one of the tails

Worked Example
For the following situations, state at the 1% and 5% significance levels whether the null
hypothesis should be rejected or not.

(i) The critical region is X≤3and the observed value is 4.


(ii) Assuming the null hypothesis is true, the probability of a value being at least as
extreme as the test statistic in a one-tailed hypothesis test is 0.0429.
(iii) Assuming the null hypothesis is true, the probability of a value being at least as
extreme as the test statistic in a two-tailed hypothesis test is 0.00705.
Conclusions of Hypothesis Testing
How is a hypothesis test carried out?
There are a number of ways that a hypothesis test can be carried out for different models,
however the following steps should form the base for your test:

Step 1: Define the test statistic and population parameter

Step 2: Write the null and alternative hypotheses clearly

Step 3: Calculate the critical value(s) or the necessary probability for the test

Step 4: Compare the observed value with the critical value(s) or the probability with the
significance level

Step 5: Decide whether there is enough evidence to reject H0 or whether it has to be


accepted

Step 6: Write a conclusion in context

How should a conclusion be written for a hypothesis test?

 Your conclusion must be written in the context of the question


 Use the wording in the question to help you write your conclusion
- If rejecting the null hypothesis your conclusion should state that there is sufficient
evidence to suggest the alternative hypothesis is true at this level of significance
- If accepting the null hypothesis your conclusion should state that there is not
enough evidence to suggest the alternative hypothesis is true at this level of
significance
 Your conclusion must not be definitive
- There is a chance that the test has led to an incorrect conclusion
- The outcome is dependent on the sample, a different sample might lead to a
different outcome
 The conclusion of a two-tailed test can state if there is evidence of a change
- You should not state whether this change is an increase or decrease
Worked Example
A teacher carried out a hypothesis test at the 10% significance level to test if her students
perform better in exams after using a new revision technique. Under the null hypothesis she
calculates the probability that a value will be at least as extreme as the observed value to be
0.09142. Write a conclusion for her hypothesis test.
Type I & Type II Errors
Any hypothesis test will only provide evidence about whether a parameter has changed or not.
A conclusion can not claim with certainty whether to accept or reject the null hypothesis as the
test is based on probability, and therefore errors are possible.

What is a Type I error?

 A Type I error occurs when the null hypothesis is rejected incorrectly


- In order for a Type I error to happen, the null hypothesis must have been rejected
 If a Type I error has been made, the hypothesis test has provided evidence that there is
a change when in fact there is not a change
- Think about the impact of this in some scenarios
- For example a test saying that a student had cheated in an exam when in fact they
had not
 The probability of a Type I error occurring in any hypothesis test is the same as the
probability of rejecting a true null hypothesis
- This is the probability of the observed value being at least as extreme as the critical
value(s)
- It is the same or a little bit less than the significance level
 In a true hypothesis test you would not need to calculate the probability of a Type I
error as it would be the same as the actual significance level

What is a Type II error?

 A Type II error occurs when the null hypothesis is accepted incorrectly


- In order for a Type II error to happen, the null hypothesis must not have been rejected
 If a Type II error has been made, the hypothesis test has provided evidence that there is
no change when in fact there was a change
- Think about the impact of this in some scenarios
- For example a test saying that a car’s brakes have not worn down, when in fact they
have
 To find probability of a Type II error occurring in any hypothesis test you would need to
be given the true value of the population parameter being tested
- For example, you would be given the true probability of the event occurring or the true
population mean
- The probability of a Type II error would be the probability of the observed value being
outside of the rejection region, given the true value of the population parameter

REJECT H0 ACCEPT H1
H0 TRUE TYPE I No Error
H0 FALSE No Error TYPE II

Can the probabilities of making the errors be manipulated?

 It is possible to reduce the probability of making a Type I error by reducing the


significance level before carrying out the test
- However, this would decrease the size of the rejection region and therefore could
increase the probability of a Type II error
 It is possible to reduce the probability of making a Type II error by increasing the
significance level before carrying out the test
- This would increase the size of the rejection region, making it easier to reject the null
hypothesis
- As the probability of rejecting the null hypothesis has increased, this would increase
the probability of making a Type I error
 Before setting the significance level a researcher could consider which error they would
want to reduce the likelihood of
- For example, if the test is for a company advertising that their product works 90% of
the time, but customers believe it may be less than this:
- the company would want to reduce the probability of a Type I error (incorrectly
declaring a change)
- the customers would want to reduce the probability of a Type II error (incorrectly
declaring no change)

Here are two tips if you cannot remember which error is which but are asked to
calculate one on the exam:

 Look to see if you are given a new population parameter, this will be a Type II error.
 Check the number of marks, a Type I error is normally only 1 mark whilst a Type II error
needs to be calculated and so will be more.
Worked Example
In the following scenarios, decide whether a Type I error or Type II error could have occurred

(i) A farmer is testing for a change in crop growth after trying a new fertiliser. The test
concludes that there is no evidence of change at the 5% significance level.
(ii) A dentist’s receptionist believes that the waiting times have been reduced due to a
new scheduling system. They conduct a hypothesis test and will reject the null
hypothesis if no more than two customers wait more than ten minutes. Exactly two
customers have to wait more than ten minutes.
Binomial Distribution
How is a hypothesis test carried out with the binomial distribution?

 The population parameter being tested will be the probability, p in a binomial


distribution B(n , p)
 A hypothesis test is used when the assumed probability is questioned
 The null hypothesis, H0 and alternative hypothesis, H1 will always be given in terms of p
- Make sure you clearly define p before writing the hypotheses
- The null hypothesis will always be H0 : p = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed tes
- A one-tailed test would test to see if the value of p has either increased or decreased
- The alternative hypothesis, H1 will be H1 : p > ... or H1 : p < ...
- A two-tailed test would test to see if the value of p has changed
- The alternative hypothesis, H1 will be H1 : p ≠ ...
 To carry out a hypothesis test with the binomial distribution, the random variable used
in the test will be the number of successes in a defined number of trials
 When defining the distribution, remember that the value of p is being tested, so this
should be written as p in the original definition, followed by the null hypothesis stating
the assumed value of p
 The binomial distribution will be used to calculate the probability of the test statistic
taking the observed value or a more extreme value
 The hypothesis test can be carried out by
- either calculating the probability of the test statistic taking the observed or a more
extreme value (p – value) and comparing this with the significance level
- or by finding the critical region and seeing whether the observed value of the test
statistic lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing

How is the critical value found in a hypothesis test with the binomial distribution?

 The critical value will be the first value to fall within the critical region
- The binomial distribution is a discrete distribution so the probability of the observed
value being within the critical region, given a true null hypothesis may be less than the
significance level
- This is the actual significance level and is the probability of incorrectly rejecting the null
hypothesis
 For a one-tailed test use your calculator to find the first value for which the probability
of that or a more extreme value is less than the given significance level
- Check that the next value would cause this probability to be greater than the
significance level
- For H1 : p < ... if straight P left parenthesis X less or equal than c right parenthesis
less or equal than alpha percent sign and straight P left parenthesis X less or equal than
c plus 1 right parenthesis less or equal than alpha percent sign then c is the critical value
- For H1 : p > ... if straight P left parenthesis X greater or equal than c right
parenthesis less or equal than alpha percent sign and straight P left parenthesis X
greater or equal than c minus 1 right parenthesis greater than alpha percent sign then c
is the critical value
 For a two-tailed test you will need to find both critical values, one at each end of the
distribution
- Find the first value for which the probability of that or a more extreme value is less
than half of the given significance level in both the upper and lower tails
- Often one of the critical regions will be much bigger than the other
- If the probability in the null hypothesis is 0.5 the critical regions will have an equal
size

What steps should I follow when carrying out a hypothesis test with the binomial
distribution?
Step 1: Define the probability, p

Step 2: Write the null and alternative hypotheses clearly using the form

H0 : p = ...

H1 : p = ...

Step 3: Define the distribution, usually X tilde B left parenthesis n comma p right parenthesis
where n is a defined number of trials and p is the population parameter to be tested

Step 4: Calculate either the critical value(s) or the necessary probability for the test

Step 5: Compare the observed value of the test statistic with the critical value(s) or the
probability with the significance level

Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted

Step 7: Write a conclusion in context


Worked Example
Jacques, a breadmaker, claims that more than 60% of people that shop in a particular
supermarket buy his brand of bread. Jacques takes a random sample of 12 customers that have
purchased bread and asks them which brand of bread they have purchased. He records that 10
of them had purchased his brand of bread. Test, at the 10% level of significance, whether
Jacques’ claim is justified.
Poisson Distribution
How is a hypothesis test carried out for the mean of a Poisson distribution?

 The population parameter being tested will be the mean, λ , in a Poisson distribution
- As it is the population mean, sometimes μ will be used instead
 A hypothesis test is used when the mean is questioned
 The null hypothesis, H0 and alternative hypothesis, H1 will be given in terms of λ (or μ)
- Make sure you clearly define λ before writing the hypotheses
- The null hypothesis will always be H0 : λ = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed test
- A one-tailed test would test to see if the value of λ has either increased or decreased
- The alternative hypothesis, will be H1 will be H1 : λ > ...or H1 : λ < ...
- A two-tailed test would test to see if the value of λ has changed
- The alternative hypothesis, H1 will be H1 : λ ≠ ...
 To carry out a hypothesis test with the Poisson distribution, the random variable will be
the mean number of occurrences of the event within the given time/space interval
- Remember you may need to change the mean to fit the interval of time or space for
your observed value
 When defining the distribution, remember that the value of λ is being tested, so this
should be written as λ in the original definition, followed by the null hypothesis stating
the assumed value of λ
 The Poisson distribution will be used to calculate the probability of the random variable
taking the observed value or a more extreme value
 The hypothesis test can be carried out by
- either calculating the probability of the random variable taking the observed or a more
extreme value and comparing this with the significance level
- or by finding the critical region and seeing whether the observed value of the test
statistic lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing
How is the critical value found in a hypothesis test with the Poisson distribution?

 The critical value will be the first value to fall within the critical region
- The Poisson distribution is a discrete distribution so the probability of the observed
value being within the critical region, given a true null hypothesis may be less than the
significance level
- This is the actual significance level and is the probability of incorrectly rejecting the null
hypothesis (a Type I error)
 For a one-tailed test use the formula to find the first value for which the probability of
that or a more extreme value is less than the given significance level
- Check that the next value would cause this probability to be greater than the
significance level
- For H1 : λ <… if P(X ≤ c) ≤ α% and P(X ≤ c+1) > α% then c is the critical value
- For H1 : λ >… if P( X ≥ c) ≤ α% and P(X ≥ c-1) > α % then c is the critical value
- Using the formula for this can be time consuming so only use this method if you need
to
- otherwise compare the probability of the random variable being at least as extreme
as the observed value with the significance level
 For a two-tailed test you will need to find both critical values, one at each end of the
distribution
 Take extra care when finding the critical region in the upper tail, you will have to find
the probabilities for less than and subtract from one

What steps should I follow when carrying out a hypothesis test with the Poisson
distribution?
Step 1: Define the mean, λ

Step 2: Write the null and alternative hypotheses clearly using the form

H0 : λ = ...

H1 : λ = ...

Step 3: Define the distribution, usually begin mathsize 16px style X tilde Po left parenthesis
lambda right parenthesis end style where λ is the mean to be tested
Step 4: Calculate the probability of the random variable being at least as extreme as the
observed value
- Or if told to find the critical region

Step 5: Compare this probability with the significance level


- Or compare the observed value with the critical region

Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted

Step 7: Write a conclusion in context

Worked Example
Mr Viajo believes that his travel blog receives an average of 8 likes per day (24 hour period). He
tries a new advertising campaign and carries out a hypothesis test at the 5% level of significance
to see if there is an increase in the number of likes he gets. Over a 6-hour period chosen at
random Mr Viajo’s travel blog receives 5 likes.

(i) State null and alternative hypotheses for Mr Viajo’s test.


(ii) Find the rejection region for the test.
(iii) Find the probability of a Type I error.
(iv) Carry out the hypothesis test, writing your conclusion clearly.
Normal Hypothesis Testing
What steps should I follow when carrying out a hypothesis test for the mean of a
normal distribution?
Following these steps will help when carrying out a hypothesis test for the mean of a normal
distribution:

Step 1: Define the distribution of the population mean usually X~ N (μ, σ2)

Step 2: Write the null and alternative hypotheses clearly

Step 3: Assuming the null hypothesis to be true, define the statistic

Step 4: Calculate either the critical value(s) or the probability of the observed value for the test

Step 5: Compare the observed value of the test statistic with the critical value(s) or the
probability with the significance level
- Or compare the z-value corresponding to the observed value with the z-value corresponding
to the critical value

Step 6: Decide whether there is enough evidence to reject H0 or whether it has to be accepted

Step 7: Write a conclusion in context

How should I define the distribution of the population mean and the statistic?
The population parameter being tested will be the population mean, μ in a normally
distributed random variable N (μ, σ2)

How should I define the hypotheses?

 A hypothesis test is used when the value of the assumed population mean is questioned
 The null hypothesis, H0 and alternative hypothesis, H1 will always be given in terms of µ
- Make sure you clearly define µ before writing the hypotheses, if it has not been
defined in the question
- The null hypothesis will always be H0 : µ = ...
- The alternative hypothesis will depend on if it is a one-tailed or two-tailed test
- A one-tailed test would test to see if the value of µ has either increased or decreased
- The alternative hypothesis, H1 will be H1 : µ > ... or H1 : µ < ...
- A two-tailed test would test to see if the value of µ has changed
- The alternative hypothesis, H1 will be H1 : µ ≠ ..

How do I define the statistic?

 The population mean is tested by looking at the mean of a sample taken from the
population
- The sample mean is denoted
- For a random variable ~ N (μ, σ2) the distribution of the sample mean would be
~ N (μ, σ2/n)
 To carry out a hypothesis test with the normal distribution, the statistic used to carry
out the test will be the sample mean,
- Remember that the variance of the sample mean distribution will be the variance of
the population distribution divided by n
- the mean of the sample mean distribution will be the same as the mean of the
population distribution

How should I carry out the test?

 The hypothesis test can be carried out by


- either calculating the probability of a value taking the observed or a more extreme
value and comparing this with the significance level
- The normal distribution will be used to calculate the probability of a value of the
random variable taking the observed value or a more extreme value
- or by finding the critical region and seeing whether the observed value lies within it
- Finding the critical region can be more useful for considering more than one
observed value or for further testing
 A third method is to compare the z-values of your observed value with the z-values
at the boundaries of the critical region(s)
̅

- Find the z-value for your sample mean using



- This is sometimes known as your test statistic
- Use the table of critical values to find the z-value for the significance level
- If the z-value for your test statistic is further away from 0 than the critical z-value
then reject H0
How is the critical value found in a hypothesis test for the mean of a normal
distribution?

 The critical value(s) will be the boundary of the critical region


- The probability of the observed value being within the critical region, given a true null
hypothesis will be the same as the significance level
 For an α% significance level
- In a one-tailed test the critical region will consist of α% in the tail that is being tested
- In a two-tailed test the critical region will consist of α/2% in each tail

To find the critical value(s) use the standard normal distribution:


Step 1: Find the distribution of the sample means, assuming H0 is true
̅

Step 2: Use the coding to standardize to z


Step 3: Use the table to find the z - value for which the probability of Z being equal to or more
extreme than the value is equal to the significance level
- You can often find this in the table of the critical values

Step 4: Equate this value to your expression found in step 2

Step 5: Solve to find the corresponding value of


If using this method for a two-tailed test be aware of the following:

The symmetry of the normal distribution means that the z - values will have the same absolute
value

You can solve the equation for both the positive and negative z – value to find the two critical
values

Check that the two critical values are the same distance from the mean

Worked Example
The time, minutes, that it takes Amelea to complete a 1000-piece puzzle can be modelled using
X~N(204, 81). Amelea gets prescribed a new pair of glasses and claims that the time it takes her
to complete a 1000-piece puzzle has decreased. Wearing her new glasses, Amelea completes
12 separate 1000-piece puzzles and calculates her mean time on these puzzles to be 201
minutes. Use these 12 puzzles as a sample to test, at the 5% level of significance, whether
there is evidence to support Amelea’s claim. You may assume the variance is unchanged.

You might also like