0% found this document useful (0 votes)
23 views15 pages

Hypothesis Testing in Statistics Guide

This document discusses hypothesis testing in statistics, emphasizing the use of probability to estimate parameters and test hypotheses. It covers key concepts such as null and alternative hypotheses, significance levels, and various statistical tests including chi-squared tests and t-tests. The document also highlights the importance of understanding the limitations and appropriate applications of statistical methods.

Uploaded by

dasica7603
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views15 pages

Hypothesis Testing in Statistics Guide

This document discusses hypothesis testing in statistics, emphasizing the use of probability to estimate parameters and test hypotheses. It covers key concepts such as null and alternative hypotheses, significance levels, and various statistical tests including chi-squared tests and t-tests. The document also highlights the importance of understanding the limitations and appropriate applications of statistical methods.

Uploaded by

dasica7603
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Applications and interpretation:

15 Hypothesis testing
ESSENTIAL UNDERSTANDINGS
n Statistics uses the theory of probability to estimate parameters, discover empirical laws, test
hypotheses and predict the occurrence of events.
n Both statistics and probability provide important representations which enable us to make
predictions, valid comparisons and informed decisions.
n The fields of statistics and probability have power and limitations and should be applied with care and
critically questioned, in detail, to differentiate between the theoretical and the empirical/observed.

In this chapter you will learn…


n about the concept and underlying principles of a hypothesis test
how to conduct a χ test for goodness of fit
2
n
how to conduct a χ test for independence using contingency tables
2
n
n how to use a t-test to compare population means
n about Spearman’s rank correlation coefficient for non-linear correlation
n about the appropriateness and limitations of Pearson’s and Spearman’s rank correlation coefficients.

■■ Figure 15.1 How sure can we be of the assumptions we make about populations based on data samples?

f f

30 30

20 20

10 10

9.1 9.2 9.3 9.4 0


1 2 3 4 5
mass (g) school year
f

300

250

200

150

100

50

0
0 1 2 3 4 5 6 7 8 9 10
number of sixes
Applications and interpretation: Hypothesis testing 371

CONCEPTS
The following key concepts will be addressed in this chapter:
n Organizing, representing, analysing and interpreting data, and utilizing different
statistical tools facilitates prediction and drawing of conclusions.
n Different statistical techniques require justification and the identification of their
limitations and validity.
n Approximation in data can approach the truth but may not always achieve it.
n Correlation and regression are powerful tools for identifying patterns.

PRIOR KNOWLEDGE
Before starting this chapter, you should already be able to complete the following:
1 There are 263 students in a year group. The probability of a student being absent is 0.07.
Find the expected number of absent students.
2 Given that X ∼ B(30, 0.23) find P(12  X  18).
3 Given that X ~ N(25, 4.82) find P(X  27).
4 a Find the Pearson’s product-moment correlation coefficient for this data:
32 51 87 22 45 57 45 73
12 18 15 11 15 17 12 19

b Interpret the value found in part a, given that the critical value of the correlation
coefficient is 0.549.

In statistics, we use sample data to make inferences about a population. This always
involves some degree of uncertainty, for example, the mean of a sample will not be,
in general, exactly the same as the mean of the population. We can use probability to
decide whether a value obtained from a sample is significantly different from what we
expect.
Hypothesis testing is used to determine whether something about a population has
changed, or whether two populations have significantly different characteristics.
Different types of test are appropriate in different situations and it is important to
understand the advantages and limitations of each test. Hypothesis tests are commonly
used in a range of subjects, including Biology, Geography, Psychology and Exercise and
Health Science.

Starter Activity
Each of the histograms in Figure 15.1 represent a sample from a different population. What
can you suggest about the distribution of each population? How confident are you in your
statement?

Now look at this problem:


The lengths of leaves of a certain plant are distributed normally with mean 18 cm and
standard deviation 3 cm. Mo has measured the lengths of a sample of leaves. For each of
these measurements decide whether it is possible, unusual or probably a measurement error.
a 23 cm
b 16 cm
c 30 cm
What if, instead, Mo’s measurements are for a plant that has leaf lengths that are
distributed normally with mean 18 cm and standard deviation 4 cm?
372 15 Applications and interpretation: Hypothesis testing

15A Chi-squared tests


■■ Introduction to hypothesis testing
Consider the following questions:
n Has the mean temperature increased in the past fifty years?
n Does r = 0.631 suggest significant correlation?
n Is a normal distribution a good model for the data?

One common statistical technique for answering such questions is hypothesis testing.
This procedure starts with a default position, called the null hypothesis, and determines
whether sample data provides sufficient evidence against it, and in favour of an
alternative hypothesis.

KEY POINT 15.1


The null hypothesis, denoted by H0, is the default position, assumed to be true unless there is
significant evidence against it.
The alternative hypothesis, denoted by H1, specifies how you think the position may have
changed.

WORKED EXAMPLE 15.1


For each situation, write down the null and the alternative hypotheses for a hypothesis test.
a The mean January temperature in Dubai from 1968 to 1998 was 29 °C. In January 2018 it was 31 °C. Is there
significant evidence that the mean temperature has increased?
b Hans collects data on the average house price and the distance from the nearest train station for eight villages. He
calculates the value of the correlation coefficient to be −0.631. Does this provide significant evidence of negative
correlation between the average house price and the distance from the nearest train station?
c Aranyaa knows that heights of trees in her local park follow a normal distribution with mean 12 m and standard
deviation 2.5 m. She wants to test whether the heights of trees in a nearby forest follow the same distribution. She
measures the heights of 50 randomly selected trees from the forest and obtains the following results:

Height (h m) ℎ  7 7  ℎ  9.5 9.5  ℎ  12 12  ℎ  14.5 14.5 < ℎ  17 ℎ  17


Frequency 3 7 20 15 4 1
Does this provide significant evidence that the heights of the forest trees do not follow the same normal distribution?

The default position is that the a H0 : μ = 29, H1: μ  29


temperature has not increased
where μ is the mean January temperature in Dubai, in °C
In this case it is easiest to write the hypotheses
using equations and inequalities, defining
the meaning of any letters you use
You can also write the hypotheses in b H0 : There is no correlation between the average house price
words. In this case, the default position and the distance from the train station.
is that there is no correlation

H1 : There is a negative correlation between the average
house price and the distance from the train station.
The default position is that the trees in the forest c H
 0 : The heights of the trees in the forest come from the
follow the same distribution as the trees in the park distribution N(12, 2.52)
Remember that N(μ, σ 2) stands for a normal 
H1: The heights of the trees in the forest do not come from
distribution with mean μ and standard deviation σ the distribution N(12, 2.52)
15A Chi-squared tests 373

In examples a and b on the previous page, the alternative hypotheses specified the
anticipated direction of change: whether the temperature has increased, or whether
the correlation is negative. In some situations, you may have a reason to believe that
the mean value, for example, has changed, but you might not have any indication in
which direction it has changed. In this case you need to use ≠, instead of  or , in the
alternative hypothesis.

KEY POINT 15.2


In a one-tailed test, the alternative hypothesis specifies whether a population parameter
has increased or decreased.
In a two-tailed test, the alternative hypothesis just states that the parameter has changed.

WORKED EXAMPLE 15.2


For each situation, write down the null and alternative hypotheses, and state whether the test
is one-tailed or two-tailed.
a The proportion of students at the college taking Maths SL was 72%. Elsa wants to test
whether this proportion has changed.
b Asher has collected some data on the heights of trees and the lengths of their leaves and
wants to test whether there is significant positive correlation.
c Theo wants to find out whether there is any correlation between the number of hours of
homework done by a student and their final examination grade.
a two-tailed test
The direction of change
H0: p = 72, H1: p ≠ 72
is not specified
where p is the percentage of students taking Maths SL
b one-tailed test

H0: There is no correlation between the height of a tree
Asher is only testing
and the length of its leaves
for positive correlation

H1: There is a positive correlation between the height of a
tree and the length of its leaves
c two-tailed test
H0: There is no correlation between amount of homework
Theo has not specified
and grade
the type of correlation

H1: There is a correlation between amount of homework
and grade

In a hypothesis test, you start with a default position (the null hypothesis) and ask
whether there is significant evidence against it. Such evidence is found if, assuming
the null hypothesis is true, it would be very unlikely to observe the sample data you
collected. What ‘very unlikely’ means is a matter of judgement. It may be different in
different contexts and is given by the significance level of the test.
When conducting a hypothesis test, you need to calculate the probability of obtaining
the observed sample value, or a more extreme value, assuming the null hypothesis
is true.
374 15 Applications and interpretation: Hypothesis testing

KEY POINT 15.3


The significance level of a hypothesis test specifies the probability that is sufficiently small
to provide significant evidence against H0.
Assuming H0 is correct, the probability of the observed sample value, or more extreme, is
called the p-value.
If the p-value is smaller than the significance level, there is sufficient evidence to reject H0 in
favour of H1. Otherwise, there is insufficient evidence to reject H0.

When writing the conclusion to a hypothesis test, it is important to state it in context,


making it clear that it is not certain, but referring to the significance level.

WORKED EXAMPLE 15.3


State the conclusion of each hypothesis test.
a Daniel wants to find out whether he spends more time on average playing computer
games than his friends do. His hypotheses are H0: μ = 1.2, H1: μ  1.2, where μ is the
mean time (in hours) per day spent playing computer games. He uses a 5% significance
level for his test. Daniel collects some data and calculates the p-value to be 0.065.
b Alessia is investigating whether the times Daniel spends playing computer games
each day come from a normal distribution with mean 1.2 hours and standard deviation
0.3 hours. She uses a 10% significance level for her test. Alessia collects some data and
calculates the p-value to be 0.07.

Compare the p-value to


the significance level
a 0.065  0.05, so there is insufficient evidence to reject H0.
Write the conclusion There is insufficient evidence, at the 5% significance
in context, making it level, that the mean time Daniel spends playing computer
clear that it is uncertain games is more than 1.2 hours.
Compare the p-value to
the significance level
b 0.07  0.10, so there is sufficient evidence to reject H0.
Write the conclusion There is sufficient evidence, at the 10% significance level,
in context, making it that the times Daniel spends playing computer games do
clear that it is uncertain not come from the distribution N(1.2, 0.32).

It is also possible to use critical values, rather than p -values, in establishing a


conclusion of a hypothesis test. You met this approach in Chapter 6, Section D.

Be the Examiner 15.1


Each of these conclusions has been incorrectly written. Explain why.

A Accept H0. There is insufficient evidence that the mean has changed.
B Do not reject H0. The mean daily temperature is still 23 °C.
C Do not reject H0. There is sufficient evidence that the mean height of 12-year-olds is
135 cm.
D Reject H0. There is no correlation between height and weight of puppies.
E Reject H0. The mean running time has decreased.
15A Chi-squared tests 375

The exact details of calculating the p-value depend on the type of test you are
conducting. You will learn about several different types of test in the following sections.

42 TOOLKIT: Problem Solving



3
A hypothesis test rejects H0 at the 5% significance level. What can you say about
924
π what the conclusion would be at the 10% significance level? What about the 2%
significance level?

CONCEPTS – REPRESENTATION
The sample is just a snapshot of the whole population. If it has been collected in a
sensible way, we hope it will be a representative sample. However, it is rarely of interest
in its own right. Most mathematical statistics is about using information from a sample to
infer properties of the population. Hypothesis testing is one of the most common ways of
doing this.

■ χ 2 test for goodness of fit


The χ test (pronounced chi-squared) is used to test whether sample data comes from
2

a given distribution. The null hypothesis is that the data comes from the specified
distribution. The χ 2 statistic measures how far the observed data values are from what
would be expected if the null hypothesis were true. You can use your GDC to calculate
the χ statistic and the corresponding p-value from observed and expected frequencies.
2

You can then compare the χ statistic to the critical value, or the p-value to the
2

significance level of the test.

KEY POINT 15.4


χ 2 test for goodness of fit:
l The data is recorded in a frequency or grouped frequency table. These are the
observed frequencies.
l Expected frequencies are the frequencies for each group assuming the null
hypothesis is true.
l The χ 2 statistic and the p-value can be obtained from a GDC.
l If the χ 2 statistic is larger than the critical value, or if the p-value is smaller than the
significance level, there is sufficient evidence to reject H0.

The critical value depends on the significance level and the number of groups in the
table.

KEY POINT 15.5


2
The critical value for the χ test depends on the significance level and the number of
degrees of freedom (v).
For a frequency table with n groups, v = n − 1.

 lthough in the SL course you will only encounter situations where the number
A
of degrees of freedom is n − 1, this number can in fact change depending on the
number of parameters that have been estimated from the data. You will learn more
about this if you are studying Mathematics: applications and interpretation HL.
376 15 Applications and interpretation: Hypothesis testing

WORKED EXAMPLE 15.4


A school offers a choice of three different languages. Igor knows that in his year, 1 of the students selected Spanish,
3
1 selected German and 5 selected Russian. In order to investigate whether the choices in the year below follow the same
4 12
distribution, he selected a sample of 60 students. Their choices were:

Language Frequency
Spanish 26
German 16
Russian 18
2
Igor conducts a χ test at the 5% significance level.
a State the null and alternative hypotheses.

b State the number of degrees of freedom.


2
c Fill out a table of observed and expected frequencies and calculate the χ value and the p-value.

d The critical value for this test is 5.991. What should Igor conclude?

The null hypothesis is that the two a 


H0 : The data comes from the same distribution as Igor’s year.
distributions are the same H1 : The data does not come from the same distribution.
Use v = n − 1
b v=3−1=2
The table has three groups
c Language Observed frequency Expected frequency

Spanish 26 1 × 60 = 20
The expected frequencies are calculated 3
assuming H0 is true, so the proportions 1 × 60 = 15
are the same as in Igor’s year German 16
4
5 × 60 = 25
Russian 18
12
Use your GDC to find the critical value. You need From GDC,
to enter the observed and expected χ 2 = 3.8267, p = 0.1476
frequencies (into two different lists), and
the number of degrees of freedom
15A Chi-squared tests 377

2
Compare the χ value to the critical value
2
A large χ value provides evidence d 3.83  5.991, so there is insufficient evidence to reject H0.
against H0
There is insufficient evidence, at the 5% significance level,
Interpret the conclusion in context that the choices in the year below come from a different
distribution.

Sometimes you need to calculate probabilities before you can find expected frequencies.
You do this by using the distribution specified in the null hypothesis.

WORKED EXAMPLE 15.5


The times taken by eight-year-old children to solve a puzzle can be modelled by a normal distribution with mean
12 minutes and standard deviation 2.5 minutes. The times taken to solve the same puzzle by a random sample of
50 ten-year-old children are as follows:

Time (min) t9 9 < t  11 11 < t  13 3 < t  15 t > 15


Frequency 10 11 20 5 4
Test, using a 10% significance level, whether the times of the ten-year-old children come from the same distribution.

H0: The times come from the distribution N(12, 2.52 )


State the hypotheses
H1: The times do not come from the distribution N(12, 2.52 )
Calculate the expected frequencies. You need to find
Observed Expected
the probability for each group, using the distribution
Time frequency Probability frequency
N(12, 2.52), and then multiply the probabilities by 50
9 10 0.1151 5.76
For example, the first probability is
9 − 11 11 0.2295 11.5
P(X  9) where X ∼ N(12, 2.52)
11 − 13 20 0.3108 15.5
13 − 15 5 0.2295 11.5
> 15 4 0.1151 5.76
Use your GDC to calculate the critical value and the
v=5−1=4
p-value. You need the number of degrees of freedom
From GDC:
χ 2 = 8.62, p-value = 0.0713
You are not given the critical value, so compare
0.0713 < 0.1, so there is sufficient evidence to reject H0.
the p-value to the significance level
There is sufficient evidence, at the 10% significance level,
Write the conclusion in context that the times for the ten-year-old children do not come from
the distribution N(12, 2.52).

TOK Links
What type of error rate is acceptable to you in decisions you make? For example, when
choosing which subjects to study, which people to be friends with, or which clothes to
buy, what kind of mistakes do you make and how often do you make them? Are you ever
completely certain that you are making the right decision? What would life be like if you
could never make ‘mistakes’?
378 15 Applications and interpretation: Hypothesis testing

■ χ for independence: contingency tables


2
You met
independent
The χ test can also be used to test whether two variables are independent. The principle
2
events
is the same as for the goodness of fit test: the χ statistic is a measure of difference
2
in probability in
Chapter 7, Section B. between observed and expected frequencies. The null hypothesis is that the two variables
are independent.
To carry out the test, the observed frequencies need to be recorded in a contingency
table. Your GDC can then calculate the expected frequencies, the χ value and the
2

p-value.

42 TOOLKIT: Problem Solving



3
You can use the fact that, under the null hypothesis, the two variables are independent
924
π to work out the expected frequencies:
Consider the contingency table:
A B
X 1 1
Y 2 4

a Show that the probability of A happening is 3. Find the probability of X


happening. 8
b Assuming that A and X occur independently, find the probability of A and X
occurring.
c Given that the total sample size is fixed, find the expected frequency of A ∩ X.
d Hence complete the expected frequency contingency table.

The number of degrees of freedom depends on the size of the contingency table.

KEY POINT 15.6


For a m × n contingency table the number of degrees of freedom is v = (m − 1)(n − 1).

WORKED EXAMPLE 15.6


Julio investigates whether people’s favourite sport depends on their age. He asks 30 adults and 50 children to select their
favourite sport out of football, basketball and baseball. He records the results in the contingency table:
Adults Children
Football 8 23
Basketball 12 11
Baseball 10 16
2
Julio conducts a χ test for independence.
a State the hypotheses for this test.

b Find the number of degrees of freedom.


2
c Find the expected frequencies and the χ value.

d Julio looks up a list of critical values for his test and finds that the appropriate critical value is 9.21. What should Julio
conclude from the test?
15A Chi-squared tests 379

The null hypothesis is that the two a H0: Age and favourite sport are independent.
variables are independent
H1: Age and favourite sport are not independent.
v = (m − 1)(n − 1) b v = (3 − 1)(2 − 1) = 2
Enter the observed frequency into a matrix c Expected frequencies:
The GDC will record the expected Adults Children
frequencies in a different matrix Football 11.6 19.4
Basketball 8.63 14.4
Baseball 9.75 16.3

2
From GDC, χ = 3.93

2
Compare the χ value to the critical value. You
2 d 3.93  9.21, so insufficient evidence to reject H0.
need a large χ value in order to reject H0
There is insufficient evidence that age and favourite
State the conclusion in context
sport are not independent.

Tip
You are the Researcher
If a conclusion
There are some situations when the χ test needs to be adapted. For example, the
2
suggests things are not
independent, then it above procedure is not valid if any of the expected frequencies are smaller than
implies a relationship five, or if the number of degrees of freedom is 1 (so for a 2 × 2 table). One possible
alternative to the χ test that you might like to investigate is Fisher’s exact test.
2
and so they can also be
called dependent.

Exercise 15A
For questions 1 to 3, use the method demonstrated in Worked Example 15.1 and 15.2. For each situation, write down
suitable hypotheses and state whether the test is one-tailed or two-tailed.
1 a The mean temperature in Perth in 1980 was 23.4 °C. Sonya wants to test whether the average temperature in
2018 was higher.
b Paulo’s mean time for a 2 km run was 8.3 minutes. After following a new training regime, he wants to test
whether his mean time has decreased.
2 a A bottle claims to contain 300 ml of water on average. Tamara wants to test whether the mean amount of water
in a bottle is different.
b The mean height of adult residents of Nairobi is 163 cm. Amandla wants to test whether the mean height of
residents of Boston is different.
380 15 Applications and interpretation: Hypothesis testing

3 a Leila is investigating whether there is a negative correlation between maximum daily temperature and the
number of people at a seaside café.
b Bashir is investigating whether there is any correlation between daily rainfall and minimum daily temperature.
For questions 4 to 6, use the method demonstrated in Worked Example 15.3 to state the conclusion of the hypothesis test.
4 a H0: μ = 13.4, H1: μ > 13.4, 5% significance level; p-value = 0.0638
b H0: μ = 26, H1: μ > 26, 5% significance level; p-value = 0.105
5 a H0: μ = 2.6, H1: μ ≠ 2.6, 10% significance level; p-value = 0.0723
b H0: μ = 8.5, H1: μ ≠ 8.5, 10% significance level; p-value = 0.103
6 a H0: The data comes from the distribution N(12, 3.42)
H1: The data does not come from the distribution N(12, 3.42)
5% significance level, p-value = 0.0492
b H0: The data comes from the distribution N(53, 82)
H1: The data does not come from the distribution N(53, 82)
5% significance level, p-value = 0.102
In questions 7 to 9, you are given expected probabilities and observed frequencies. The null hypothesis is that the
observed data comes from the same distribution as the expected values. Use the method from Worked Example 15.4 to
2
i calculate the expected frequencies and the χ value
ii state the number of degrees of freedom
iii find the p-value and state the conclusion of the test at 5% significance level.
7 a
Data value Red Blue Yellow
1 1 1
Expected probability
4 2 4
Observed frequency 34 51 15

b Data value Apple Orange Banana


1 3 1
Expected probability
2 10 5
Observed frequency 45 33 22

8 a
Data value 1 2 3 4 5
Expected probability 0.2 0.2 0.2 0.2 0.2
Observed frequency 12 16 8 10 4

b
Data value 0 2 4 6
Expected probability 0.25 0.25 0.25 0.25
Observed frequency 6 18 13 23

9 a
Data value French Spanish English Mandarin
1 1 1 13
Expected probability
3 4 5 60
Observed frequency 52 51 30 67

b Data value Biology Chemistry Physics Geology


1 1 1 3
Expected probability
3 6 5 10
Observed frequency 52 31 42 35
15A Chi-squared tests 381

In questions 10 to 12, you are given the null hypothesis, observed frequencies and the critical value for the goodness of
fit test. Using the method demonstrated in Worked Example 15.5, conduct the test and state the conclusion.
10 a H0: The data comes from the distribution B(4, 0.5)
Critical value = 9.49
Data value 0 1 2 3 4
Observed frequency 10 65 68 43 14
b H0: The data comes from the distribution B(5, 0.6)
Critical value = 12.8
Data value 0 1 2 3 4 5
Observed frequency 7 28 103 188 132 42
11 a H0: The data comes from the distribution B(10, 0.77)
Critical value = 9.24
Data value 5 6 7 8 9 10
Observed frequency 7 16 28 34 12 3
b H0: The data comes from the distribution B(20, 0.08)
Critical value = 7.78
Data value 0 1 2 3 4
Observed frequency 19 35 22 15 9
12 a H0: The data comes from the distribution N(5.5, 22)
Critical value = 7.78
Data value 2 2 to 4 4 to 6 6 to 8 >8
Observed frequency 7 21 44 17 11
b H0: The data comes from the distribution N(36, 112)
Critical value = 9.24
Data value  20 20–30 30–40 40–50 50–60 > 60
Observed frequency 25 92 140 112 29 2
In questions 13 to 15, you are given a contingency table of observed values of two variables and the critical value
2
for a χ test. Use the method demonstrated in Worked Example 15.6 to test whether the two variables are independent.
2
State the χ value and the conclusion of each test.
13 a b
12 15 21 63 81 35
22 18 12 27 32 62
Critical value = 9.21 Critical value = 7.38

14 a 8 12 16 b 18 13 8
9 15 21 13 16 21
13 14 22 8 18 22
Critical value = 7.78 Critical value = 9.49
382 15 Applications and interpretation: Hypothesis testing

15 a 62 37 81 b 13 11 11
88 53 30 12 8 3
26 15 32 21 16 9
81 73 55 13 13 13
Critical value = 12.6 Critical value = 10.6

16 Ruby wants to test whether students’ food preferences depend on their age. She conducts a survey in the school
canteen, recording which option each student chooses.
Veggie burger Fish fingers Peperoni pizza
Junior school 18 26 51
Senior school 53 38 47
2
Conduct a suitable χ test at the 5% significance level, stating your hypotheses and conclusion clearly.
17 A zoologist investigates whether different types of insect are more common in different locations. She collects a
random sample of insects from a meadow and a forest and counts the number of ants, bees and flies.
Ants Bees Flies
Meadow 26 15 21
Forest 32 6 18
Use a χ2 test with a 10% significance level to test whether the type of insect found depends on the location.
18 A theory predicts that three different types of flower should appear in the ratio 1 : 2 : 3. A sample of 60 flowers
2
contains 14 flowers of type A, 18 flowers of type B and 28 flowers of type C. A χ goodness of fit test is used
to test the theory at the 10% significance level.
a Calculate the expected frequencies and state the number of degrees of freedom.
2
b Find the χ value.
c The critical value is 4.605. State the conclusion of the test.
19 Zhao thinks that students at her large college are equally likely to study any of the four mathematics courses.
She asks a random sample of 80 students and obtains the following results:
Analysis and Analysis and Applications and Applications and
Course
approaches HL approaches SL interpretation HL interpretation SL
Number of students 18 21 8 33
2
Zhao conducts a χ test using a 5% significance level.
a State the null and alternative hypotheses.
b Write down the expected frequencies and the number of degrees of freedom.
c Find the p-value for the test and hence state the conclusion.
20 The table shows information about the mode of transport students use to get to school in four different cities. Use
2
a χ test to find out whether there is evidence, at the 5% significance level, that there is a relationship between the
mode of transport and the city.
Amsterdam Athens Houston Johannesburg
Car 12 25 48 24
Bus 18 33 12 18
Bicycle 46 12 7 53
Walk 38 8 3 21
15A Chi-squared tests 383

21 A six-sided dice is rolled 120 times, giving the following results:

Outcome 1 2 3 4 5 6
Frequency 26 12 16 28 14 24
Is there evidence, at the 2% significance level, that the dice is not fair?
22 Rajesh is practising tennis serves. He takes three serves at a time and records the number of successful serves.
He believes that this number can be modelled by the binomial distribution B(3, 0.7).
Number of successful
0 1 2 3
serves out of three
Frequency 7 28 95 70
2
a State the hypotheses for a χ goodness of fit test.
b Find the expected frequencies and write down the number of degrees of freedom.
2
c Calculate the χ value.
d The critical value for the test is 6.25. State the conclusion of the test.
23 Michelle tosses six coins simultaneously and records the number of tails. She repeats this 600 times. The results
are shown in the table.
Number of tails 0 1 2 3 4 5 6
Frequency 9 62 120 178 152 67 12
a State the distribution of the number of tails for an unbiased coin.
b Hence work out the expected frequencies for Michelle’s experiment.
c Test at the 5% significance level whether there is evidence that the coins are biased.
24 Four friends are guessing answers to maths questions. They think that they each have the probability of 0.5 of
guessing the correct answer to any question, independently of each other.
a Let X be the number of correct answers to a single question. Accepting the assumptions above are correct, state
the distribution of X.
In order to test whether their assumptions are correct, the friends guess answers to 100 questions and record the
number of correct answers to each question.
Number of correct answers 0 1 2 3 4
Frequency 12 13 45 22 8
b State appropriate hypotheses for a χ2 goodness of fit test.
c State the number of degrees of freedom.
d Conduct the test at the 2% significance level and state your conclusion.
25 An athlete believes that her long jump distances follow a normal distribution with mean 5.8 m and standard
deviation 0.8 m.
a Assuming her belief is correct, copy the table below and fill in the missing probabilities:

Distance (m) <5 5 to 6 6 to 7 >7


Probability 0.159
In order to test her belief, she records her distances from a random sample of 100 jumps, obtaining the following
results:
Distance (m) <5 5 to 6 6 to 7 >7
Frequency 17 42 38 3
2
b State the number of degrees of freedom for a χ goodness of fit test.
c State suitable hypotheses.
d Conduct the test at the 10% significance level.
384 15 Applications and interpretation: Hypothesis testing

26 A train company claims that times for a particular journey are distributed normally with mean 23 minutes and
standard deviation 2.6 minutes. Sumaya takes this train to school and wants to test the company’s claim. She
2
decides to conduct a χ test and records the durations of 50 randomly selected journeys:
Time (min) < 21.5 21.5–22.5 22.5–23.5 23.5–24.5 > 24.5
Frequency 3 8 14 17 8
a Find the expected frequencies.
b Write down the number of degrees of freedom.
2
c Calculate the χ value.
d The critical value for Sumaya’s test is 9.49. State the conclusion in context.
27 A teacher suggests that exam grades at their college can be modelled by the following distribution:
g (11 − g )
P(G = g ) = for g = 3, 4, 5, 6, 7
140
A random sample of 40 students had the following grades:
Grade 3 4 5 6 7
Frequency 2 10 9 12 7
Test, using a 10% significance level, whether the teacher’s model is appropriate for these data.
28 Katya wants to find out whether diet choices are dependent on age. She collects data from 200 pupils at her school
and records it in the contingency table:
Vegetarian Vegan Eats meat
11–13 12 32 21
14–15 22 16 30
16–19 26 18 23
2
a Conduct a χ test for independence, using a 1% significance level. State your conclusion in context.
In order to get a better understanding of the data, Katya decides to split the 16–19 age group in two. The new data
table is:
Vegetarian Vegan Eats meat
11–13 12 32 21
14–15 22 16 30
16–17 13 12 10
18–19 13 6 13
b Repeat the test, stating the new conclusion clearly.
29 A teacher asked a group of 80 students for their preference out of three science subjects. The table shows the
fraction of students in each group.
Boys Girls
3 1
Biology
20 5
1 3
Chemistry
10 20
1 3
Physics
4 20

a Test, using a 10% significance level, whether favourite science is independent of gender.
b The teacher decides to check her results by using a larger sample. She asks a group of 240 students and finds
that the fractions in each group are the same. Repeat the test and state the new conclusion.

You might also like