0% found this document useful (0 votes)
12 views44 pages

Statistics

The document contains a series of statistics problems with varying maximum marks, covering topics such as probability, regression analysis, and data interpretation. Each problem requires specific calculations or justifications based on provided data or scenarios. The problems are designed for students to demonstrate their understanding of statistical concepts and methods.

Uploaded by

Yunus Altın
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views44 pages

Statistics

The document contains a series of statistics problems with varying maximum marks, covering topics such as probability, regression analysis, and data interpretation. Each problem requires specific calculations or justifications based on provided data or scenarios. The problems are designed for students to demonstrate their understanding of statistical concepts and methods.

Uploaded by

Yunus Altın
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistics [408 marks]

1. [Maximum mark: 5] [Link].TZ0.2


Let A and B be events such that P (A) ,
= 0.5 P (B) = 0.4 and
P (A ∪ B) = 0.6.

Find P (A |B). [5]

2. [Maximum mark: 15] [Link].TZ0.7


A large company surveyed 160 of its employees to find out how much time they
spend traveling to work on a given day. The results of the survey are shown in
the following cumulative frequency diagram.
(a) Find the median number of minutes spent traveling to work. [2]

(b) Find the number of employees whose travelling time is within


15 minutes of the median. [3]

Only 10% of the employees spent more than k minutes traveling to work.

(c) Find the value of k. [3]

The results of the survey can also be displayed on the following box-and-
whisker diagram.
(d) Write down the value of b. [1]

(e.i) Find the value of a. [2]

([Link]) Hence, find the interquartile range. [2]

(f ) Travelling times of less than p minutes are considered outliers.

Find the value of p. [2]

3. [Maximum mark: 7] [Link].TZ0.4


The following table below shows the marks scored by seven students on two
different mathematics tests.

Let L1 be the regression line of x on y. The equation of the line L1 can be written in
the form x = ay + b.

(a) Find the value of a and the value of b. [2]

Let L2 be the regression line of y on x. The lines L1 and L2 pass through the same
point with coordinates (p , q).

(b) Find the value of p and the value of q. [3]

(c) Jennifer was absent for the first test but scored 29 marks on the
second test. Use an appropriate regression equation to estimate
Jennifer’s mark on the first test. [2]
4. [Maximum mark: 6] [Link].TZ0.6
In a city, the number of passengers, X, who ride in a taxi has the following
probability distribution.

After the opening of a new highway that charges a toll, a taxi company
introduces a charge for passengers who use the highway. The charge is $ 2.40 per
taxi plus $ 1.20 per passenger. Let T represent the amount, in dollars, that is
charged by the taxi company per ride.

(a) Find E(T). [4]

(b) Given that Var(X) = 0.8419, find Var(T). [2]

5. [Maximum mark: 6] [Link].TZ0.3


The following table shows the probability distribution of a discrete
random variable X where x = 1, 2, 3, 4.

Find the value of k, justifying your answer. [6]

6. [Maximum mark: 4] [Link].TZ0.1


A data set consisting of 16 test scores has mean 14. 5 . One test score of
9 requires a second marking and is removed from the data set.

Find the mean of the remaining 15 test scores. [4]

7. [Maximum mark: 8] [Link].TZ0.4


The following table shows the systolic blood pressures, p mmHg, and the ages, t
years, of 6 male patients at a medical clinic.

(a.i) Determine the value of Pearson’s product‐moment correlation


coefficient, r, for these data. [2]

([Link]) Interpret, in context, the value of r found in part (a) (i). [1]

The relationship between t and p can be modelled by the regression line of p on


t with equation p = at + b .

(b) Find the equation of the regression line of p on t. [2]

A 50‐year‐old male patient enters the medical clinic for his appointment.

(c) Use the regression equation from part (b) to predict this
patient’s systolic blood pressure. [2]

(d) A 16‐year‐old male patient enters the medical clinic for his
appointment.

Explain why the regression equation from part (b) should not
be used to predict this patient’s systolic blood pressure. [1]

8. [Maximum mark: 6] [Link].TZ0.9


A biased coin is weighted such that the probability, p, of obtaining a
tail is 0. 6. The coin is tossed repeatedly and independently until a tail
is obtained.

Let E be the event “obtaining the first tail on an even numbered toss”.
Find P(E). [6]

9. [Maximum mark: 9] [Link].TZ0.2


A set of data comprises of five numbers x 1 , x2 , x3 , x4 , x5 which have been
placed in ascending order.

(a) Recalling definitions, such as the Lower Quartile is the th


n+1

piece of data with the data placed in order, find an expression


for the Interquartile Range. [2]

(b) Hence, show that a data set with only 5 numbers in it cannot
have any outliers. [5]

(c) Give an example of a set of data with 7 numbers in it that does


have an outlier, justify this fact by stating the Interquartile
Range. [2]

10. [Maximum mark: 15] [Link].TZ0.1


The principal of a high school is concerned about the effect social media use
might be having on the self-esteem of her students. She decides to survey a
random sample of 9 students to gather some data. She wants the number of
students in each grade in the sample to be, as far as possible, in the same
proportion as the number of students in each grade in the school.

(a) State the name for this type of sampling technique. [1]

The number of students in each grade in the school is shown in table.


(b.i) Show that 3 students will be selected from grade 12. [3]

([Link]) Calculate the number of students in each grade in the sample. [2]

In order to select the 3 students from grade 12, the principal lists their names in
alphabetical order and selects the 28th, 56th and 84th student on the list.

(c) State the name for this type of sampling technique. [1]

Once the principal has obtained the names of the 9 students in the random
sample, she surveys each student to find out how long they used social media
the previous day and measures their self-esteem using the Rosenberg scale. The
Rosenberg scale is a number between 10 and 40, where a high number
represents high self-esteem.

(d.i) Calculate Pearson’s product moment correlation coefficient, r. [2]

([Link]) Interpret the meaning of the value of r in the context of the


principal’s concerns. [1]

([Link]) Explain why the value of r makes it appropriate to find the


equation of a regression line. [1]

(e) Another student at the school, Jasmine, has a self-esteem value


of 29.

By finding the equation of an appropriate regression line,


estimate the time Jasmine spent on social media the previous
day. [4]

11. [Maximum mark: 6] [Link].TZ1.5


Consider events A and B such that P(A′) = P(A ∪ B) =
3

4
and
P(B A) =
2
3
.

This question is duplicated in #15

(a) Find P (A ∩ B). [3]

This question is duplicated in #15

(b) Show that events A and B are independent. [3]

12. [Maximum mark: 8] [Link].TZ1.6


Consider a sequence of ten rectangular picture frames F , F , ... , F , F .1 2 9 10

Picture frame F has width 4 cm and height 5 cm.


1

The width and height of picture frame F , are each increased by 50 % to


n

generate the width and height of the next picture frame F , for n ∈ Z , n+1
+

1 ≤ n ≤ 9.

This question is duplicated in #16

(a.i) Show that the area of picture frame F is 20( n


9
)
n−1
cm
2
. [2]
4

This question is duplicated in #16

([Link]) Hence, find the mean area of the ten picture frames, giving your
a
answer in the form p(( 9

4
) − 1) cm
2
, where p ∈ Q
+
,
a ∈ Z
+
. [3]
This question is duplicated in #16

(b) Find the median area of the ten picture frames, giving your
4
answer in the form q( 9

4
) cm
2
, where q ∈ Q
+
. [3]

13. [Maximum mark: 13] [Link].TZ2.7


A discrete random variable, X, has the following probability distribution, where
a > 0 and k is a constant.

(a) Show that k =


1

3
. [5]

(b) Find P (X < 3a) . [2]

(c) Find P (X ≥ a X < 3a) . [3]

(d) Given that E (X) = 20 , find the value of a. [3]

14. [Maximum mark: 4] [Link].TZ3.1


The scores achieved by 80 golfers in a competition are summarized in the
following box and whisker diagram.
(a) Find the interquartile range. [2]

(b) Find the number of golfers that scored between 70 and 74. [2]

15. [Maximum mark: 6] [Link].TZ1.4


Consider events A and B such that P(A′) = P(A ∪ B) =
3

4
and
P(B A) =
2

3
.

This question is duplicated in #11

(a) Find P (A ∩ B). [3]

This question is duplicated in #11

(b) Show that events A and B are independent. [3]

16. [Maximum mark: 8] [Link].TZ1.5


Consider a sequence of ten rectangular picture frames F , F , ... , F , F .
1 2 9 10

Picture frame F has width 4 cm and height 5 cm.


1

The width and height of picture frame F , are each increased by 50 % to


n

generate the width and height of the next picture frame F , for n ∈ Z ,
n+1
+

1 ≤ n ≤ 9.
This question is duplicated in #12

(a.i) Show that the area of picture frame F is 20( n


9
)
n−1
cm
2
. [2]
4

This question is duplicated in #12

([Link]) Hence, find the mean area of the ten picture frames, giving your
a
answer in the form p(( 9

4
) − 1) cm
2
, where p ∈ Q
+
,
a ∈ Z . [3]
+

This question is duplicated in #12

(b) Find the median area of the ten picture frames, giving your
4
answer in the form q( 9

4
) cm
2
, where q ∈ Q
+
. [3]

17. [Maximum mark: 7] [Link].TZ2.4


Events A and B are such that P(A ∪ B) =
5

8
and P (A ∩ B′) =
24
7
.

(a) Find P (B). [3]

(b) Given that events A and B are independent, find P (A′ B) . [4]

18. [Maximum mark: 8] [Link].TZ1.5


In a study, measurements for arm span, A cm, and foot length, F cm, are taken

from a large group of adults.

For this group, the regression line of F on A is found to be


F = 0. 335A − 32. 6, and the regression line of A on F is found to be

A = 2. 89F + 99. 3. Each regression line passes through the mean point.
This question is duplicated in #23

(a) By using an appropriate regression line, find an estimate of the


arm span for an adult with a foot length of 19. 8 cm. [2]

This question is duplicated in #23

(b) For this group of adults, find the mean arm span and the mean
foot length. [3]

The heights, H cm, of adults in the group can be modelled by a normal


distribution with mean 163 cm and standard deviation σ cm.

It is found that 88% of the group have a height between 153 cm and 173 cm.

This question is duplicated in #23

(c) Find the value of σ. [3]

19. [Maximum mark: 14] [Link].TZ1.7


Lynn is playing a game with two unbiased six-sided dice, each with faces marked
with the integers from 1 to 6.

In each round, she throws both dice once. The outcomes can be displayed in the
following sample space diagram, which has been partially completed:
Lynn scores points according to the following rules.

If the two dice show the same score, she scores 10 points.
If the two dice show scores which have a difference of one, for example the
scores 4 and 5 in any order, she scores 5 points.
Otherwise, she scores 0 points.
(a) Show that the probability that Lynn scores 5 points in one
round is 5
18
. [2]

(b) Find the probability that Lynn scores no points in one round. [2]

The random variable X represents the number of points Lynn scores in one
round.

(c) Find E (X). [4]

(d) Hence, estimate the total number of points that Lynn scores if
she plays 90 rounds. [2]

A prize is awarded to any player who scores more than 40 points in total.

Lynn plays exactly five rounds.


(e) Find the probability that Lynn wins a prize. [4]

20. [Maximum mark: 4] [Link].TZ2.1


The following table shows the number of hours of play time, x, and sleep time, y,
for a group of six children, over the period of one week.

The regression line of y on x for this data can be written in the form y = ax + b.

This question is duplicated in #24

(a) Find the value of a and the value of b. [2]

This question is duplicated in #24

(b) Use the equation of the regression line to estimate the sleep
time of a child whose weekly play time is 20 hours. [2]

21. [Maximum mark: 6] [Link].TZ3.4


A supermarket analyses the shopping habits of its customers.

The number of times, X, each customer visits the supermarket in a week is given
by the following probability distribution.
This question is duplicated in #26

(a.i) Find the value of a. [2]

This question is duplicated in #26

([Link]) Write down the mode of X. [1]

This question is duplicated in #26

(b) Find the mean of X. [2]

The manager wants to know why customers come to their supermarket. They
survey the first 50 customers to arrive at the supermarket on a particular day.

This question is duplicated in #26

(c) Identify which one of the following best describes the


manager’s sampling method.

Circle your answer.

[1]

22. [Maximum mark: 17] [Link].TZ3.8


Consider a discrete random variable X.
(a) State two conditions required for X to be modelled by a
binomial distribution. [2]

A water theme park has two rides: Daifong and Torbellino. Each visitor’s decision to
ride on either Daifong or Torbellino is made independently of any other person.

From previous records, it is expected that 37 % of the visitors on any particular


day will ride Daifong.

On Saturday, 1 900 people will visit the theme park.

(b) Find the number of people that are expected to ride Daifong. [2]

(c) Find the probability that

(c.i) 712 people will ride Daifong; [2]

([Link]) between 684 and 712 people, inclusive, will ride Daifong. [2]

(d) Given that between 684 and 712 people, inclusive, will ride
Daifong, find the probability that at most 692 people will ride
Daifong. [4]

The ride Torbellino is more popular at the theme park. It is expected that 61 % of the
visitors on any particular day will ride Torbellino.

It can be assumed that the probability a person will ride Daifong is independent of
them riding Torbellino.

(e) Find the probability that a person will ride both Daifong and
Torbellino. [2]

Next Tuesday n people will visit the theme park. The probability that at most 500
people will ride Torbellino is approximately 0. 693.

(f ) Find the value of n. [3]


23. [Maximum mark: 8] [Link].TZ1.4
In a study, measurements for arm span, A cm, and foot length, F cm, are taken

from a large group of adults.

For this group, the regression line of F on A is found to be


F = 0. 335A − 32. 6, and the regression line of A on F is found to be

A = 2. 89F + 99. 3. Each regression line passes through the mean point.

This question is duplicated in #18

(a) By using an appropriate regression line, find an estimate of the


arm span for an adult with a foot length of 19. 8 cm. [2]

This question is duplicated in #18

(b) For this group of adults, find the mean arm span and the mean
foot length. [3]

The heights, H cm, of adults in the group can be modelled by a normal


distribution with mean 163 cm and standard deviation σ cm.

It is found that 88% of the group have a height between 153 cm and 173 cm.

This question is duplicated in #18

(c) Find the value of σ. [3]

24. [Maximum mark: 4] [Link].TZ2.1


The following table shows the number of hours of play time, x, and sleep time, y,
for a group of six children, over the period of one week.
The regression line of y on x for this data can be written in the form y = ax + b.

This question is duplicated in #20

(a) Find the value of a and the value of b. [2]

This question is duplicated in #20

(b) Use the equation of the regression line to estimate the sleep
time of a child whose weekly play time is 20 hours. [2]

25. [Maximum mark: 7] [Link].TZ2.8


The marks obtained by students in a class quiz are shown in the
following table where p, q ∈ Z . +

The mean and variance of the marks are 31 and 124 respectively.
Find the value of p and the value of q. [7]

26. [Maximum mark: 6] [Link].TZ3.3


A supermarket analyses the shopping habits of its customers.

The number of times, X, each customer visits the supermarket in a week is given
by the following probability distribution.

This question is duplicated in #21

(a.i) Find the value of a. [2]

This question is duplicated in #21

([Link]) Write down the mode of X. [1]

This question is duplicated in #21

(b) Find the mean of X. [2]

The manager wants to know why customers come to their supermarket. They
survey the first 50 customers to arrive at the supermarket on a particular day.

This question is duplicated in #21


(c) Identify which one of the following best describes the
manager’s sampling method.

Circle your answer.

[1]

27. [Maximum mark: 18] [Link].TZ3.11


Amanda enters data from surveys into a database. It can be assumed that the
accuracy of any survey entered is independent of all other surveys entered.

From previous records, it is known that Amanda enters 8 % of the surveys


inaccurately.

(a) On a particular day Amanda enters data from 50 surveys.

(a.i) Find the probability that Amanda entered at most six surveys
inaccurately. [2]

([Link]) Given that at most six surveys were entered inaccurately, find
the probability that exactly four surveys were entered
inaccurately. [3]

On a different day Amanda enters data from n surveys. On this day, the
probability that at most six surveys were entered inaccurately is approximately
0. 367.

(b) Find the value of n. [3]

Bryce and Carmen also enter data from surveys into the same database. It is
known that surveys entered by Bryce and Carmen are inaccurate 6% and 11% of
the time respectively. It can again be assumed that the accuracy of any survey
entered is independent of all other surveys entered.
From the surveys assigned to the three of them, Amanda enters 55%, Bryce 25%
and Carmen 20%.
(c) Find the probability that a randomly selected survey was

(c.i) entered inaccurately; [3]

([Link]) entered by Amanda, given that the survey was entered


inaccurately. [3]

The following year, the accuracy of Amanda’s and Bryce’s work remained the
same, as did the percentage of surveys entered by each of the three employees.
However, Carmen’s accuracy had improved and the probability that she entered
a survey inaccurately was now x%.

The probability that a randomly selected survey had been entered inaccurately
was now the same as the probability that Carmen made an error when entering a
survey.

(d) Find the value of x. [4]

28. [Maximum mark: 23] [Link].TZ1.1


This question asks you to use polynomial functions to model some situations
in probability.

Two unbiased tetrahedral (four-sided) dice with faces labelled 1, 2, 3 and 4 are
thrown and the scores recorded.

The random variable M denotes the maximum of these two scores.

The probability distribution of M is given in the following table.


(a) Find E (M ). [2]

An alternative way to represent the probability distribution of M is to use a


4

polynomial function, G, where G (t) = Σ P (M = m) t


m
.
m=1

Hence, for the distribution of M , G (t) =


1

16
t +
3

16
t
2
+
5

16
t
3
+
7

16
t
4
.

(b) Find G(1). [1]

(c.i) Find G′ (t). [2]

([Link]) Hence, show that G′(1) = E (M ) . [3]

A bag contains two red balls and three yellow balls.

Two balls are selected at random without replacement from the bag.

The random variable X denotes the total number of red balls selected.

The probability distribution of X can be represented by the polynomial


function, G , where
X

G x (t) = Σ P (X = x)t
x
.
x=0

(d) Show that G X (t) =


3

10
+
3

5
t +
1

10
t
2
, making it clear how the
coefficients of G X (t) have been determined. [5]

An unbiased coin and a biased coin are tossed.

The probability of obtaining a tail on the biased coin is p.

The random variable Y denotes the total number of tails obtained from tossing
both coins.
The probability distribution of Y can be represented by the polynomial
function, G , where
Y

G Y (t) = Σ P (Y = y)t
y
.
y=0

(e) Given that the coefficient of t in G2


Y (t) is 1

3
, find

(e.i) the value of p; [2]

([Link]) an expression for G Y .


(t) [4]

The random variable Z denotes the sum of the total number of red balls
selected, X, and the total number of tails obtained from tossing both coins, Y .

The probability distribution of Z can be represented by the function, G , where Z

G Z (t) = G X (t)G Y (t) .

(f ) For random variable Z , it can be shown that G Z ′(1) .


= E (Z)

Use this result to find E (Z). [4]

29. [Maximum mark: 6] [Link].TZ1.1


Consider the following set of ordered data.

This question is duplicated in #31

(a) Write down

This question is duplicated in #31


(a.i) the mode; [1]

This question is duplicated in #31

([Link]) the range; [1]

This question is duplicated in #31

([Link]) the median. [1]

This question is duplicated in #31

(b) Find the interquartile range. [3]

30. [Maximum mark: 6] [Link].TZ1.3


Two events A and B are such that P(A) = 0. 45 , P(B) = 0. 65 and
P(A ∪ B) = 0. 8.

(a) Find P(A ∩ B). [3]

(b) Find P(A ′ ′


B ) . [3]

31. [Maximum mark: 6] [Link].TZ2.1


Consider the following set of ordered data.

This question is duplicated in #29


(a) Write down

This question is duplicated in #29

(a.i) the mode; [1]

This question is duplicated in #29

([Link]) the range; [1]

This question is duplicated in #29

([Link]) the median. [1]

This question is duplicated in #29

(b) Find the interquartile range. [3]

32. [Maximum mark: 6] [Link].TZ2.3


Two events A and B are such that P (A) = 0. 65 , P (B) = 0. 45 and
P (A ∪ B) = 0. 85.

(a) Find P (A ∩ B). [3]

(b) Find P (A′ B′). [3]

33. [Maximum mark: 20] [Link].TZ0.12


Consider the equation z 4
= 16i , where z .
∈ C
The equation has four roots z , z , z , z , where z 1 2 3 4 i = r (cos θ i + i sin θ i ) ,
r > 0 and 0 ≤ θ < θ < θ < θ < 2π.
1 2 3 4

(a) Find z , z , z and z .


1 2 3 4 [6]

The roots z , z , z and z form a geometric sequence.


1 2 3 4

(b) Find the common ratio of the sequence, expressing your


answer in Cartesian form. [3]

The roots z , z , z and z are represented by the points A, B, C and D


1 2 3 4

respectively on an Argand diagram.

(c) Plot the points A, B, C and D on an Argand diagram. [3]

The equation v 4
= a + bi , where v ∈ C and a, b ∈ R has roots z , z , z1
*
2
*
3
*

and z .
4
*

(d) Determine the value of a and the value of b. [3]

The midpoint of [AB] is A , the midpoint of [BC] is B , the midpoint of [CD] is


′ ′

C and the midpoint of [DA] is D .


′ ′

Consider the equation w p


= 2
q
, where w ∈ C and p, q ∈ Z
+
.

Four of the roots of w p


= 2
q
are represented by the points A , B , C and D . ′ ′ ′ ′

(e) Find the least possible value of p and the corresponding value
of q. [5]

34. [Maximum mark: 15] [Link].TZ1.8


The following table shows the population of Canada t years after the year 2000.
A student uses linear regression to model the population of Canada using these
data.
The student model is p = at + b.

This question is duplicated in #36 and #38

(a.i) Write down the value of a and the value of b. [2]

This question is duplicated in #36 and #38

([Link]) Interpret, in context, the value of a. [1]

The student uses this model to predict the population of Canada in the year
2030, where t = 30, and calculates a population of approximately 41. 4 million

people.

This question is duplicated in #36 and #38

(b) Comment on the reliability of the student’s prediction. [1]

A data scientist, Benoit, uses additional information to develop an exponential


model for Canada’s future population.

In this model, B(t) = 30. 6(1. 007) represents the millions of people in
t

Canada t years after the year 2000, where 25 ≤ t ≤ 100.

This question is duplicated in #36 and #38


(c.i) Use Benoit’s model to predict the population of Canada in the
year 2100. [2]

This question is duplicated in #36 and #38

([Link]) Interpret, in context, the value 1. 007 in Benoit’s model. [1]

Another data scientist, Cecilia, develops a third model for the Canadian
population.

In this model, C(t) =


1+e
61
−0.03t
represents the millions of people in Canada t
years after the year 2000, where 25 ≤ t ≤ 100 .

This question is duplicated in #36 and #38

(d) Use Cecilia’s model to predict the population of Canada in the


year 2100. [1]

This question is duplicated in #36 and #38

(e) Determine the year in which the difference between the


predictions from Benoit’s model and Cecilia’s model is greatest. [3]

This question is duplicated in #36 and #38

(f ) Find the value of

This question is duplicated in #36 and #38

(f.i) B′(40); [1]


This question is duplicated in #36 and #38

([Link]) C′(40). [1]

This question is duplicated in #36 and #38

(g) Compare and interpret, in context, the values of B′(40) and


C′(40). [2]

35. [Maximum mark: 5] [Link].TZ2.4


A discrete random variable, X, has the following probability distribution:

P(X = x) =
kx

20
for x ∈ {3, 5, 8, 11} .

This question is duplicated in #37

(a) Find the value of k. [2]

This question is duplicated in #37

(b) Find E(X). [3]

36. [Maximum mark: 15] [Link].TZ2.8


The following table shows the population of Canada t years after the year 2000.
A student uses linear regression to model the population of Canada using these
data.
The student model is p = at + b.

This question is duplicated in #34 and #38

(a.i) Write down the value of a and the value of b. [2]

This question is duplicated in #34 and #38

([Link]) Interpret, in context, the value of a. [1]

The student uses this model to predict the population of Canada in the year
2030, where t = 30, and calculates a population of approximately 41. 3 million

people.

This question is duplicated in #34 and #38

(b) Comment on the reliability of the student’s prediction. [1]

A data scientist, Benoit, uses additional information to develop an exponential


model for Canada’s future population.

In this model, B(t) = 33. 5(1. 005) represents the millions of people in
t

Canada t years after the year 2000, where 25 ≤ t ≤ 100.

This question is duplicated in #34 and #38

(c.i) Use Benoit’s model to predict the population of Canada in the


year 2100. [2]
This question is duplicated in #34 and #38

([Link]) Interpret, in context, the value 1. 005 in Benoit’s model. [1]

Another data scientist, Cecilia, develops a third model for the Canadian
population.

In this model, C(t) =


1+e
62
−0.02t
represents the millions of people in Canada t
years after the year 2000, where 25 ≤ t ≤ 100 .

This question is duplicated in #34 and #38

(d) Use Cecilia’s model to predict the population of Canada in the


year 2100. [1]

This question is duplicated in #34 and #38

(e) Determine the year in which the difference between the


predictions from Benoit’s model and Cecilia’s model is greatest. [3]

This question is duplicated in #34 and #38

(f ) Find the value of

This question is duplicated in #34 and #38

(f.i) B′(75); [1]

This question is duplicated in #34 and #38

([Link]) C′(75). [1]


This question is duplicated in #34 and #38

(g) Compare and interpret, in context, the values of B′(75) and


C′(75). [2]

37. [Maximum mark: 5] [Link].TZ0.3


A discrete random variable, X, has the following probability distribution:

P(X = x) =
kx

20
for x ∈ {3, 5, 8, 11} .

This question is duplicated in #35

(a) Find the value of k. [2]

This question is duplicated in #35

(b) Find E(X). [3]

38. [Maximum mark: 15] [Link].TZ0.10


The following table shows the population of Canada t years after the year 2000.

A student uses linear regression to model the population of Canada using these
data.
The student model is p = at + b.
This question is duplicated in #34 and #36

(a.i) Write down the value of a and the value of b. [2]

This question is duplicated in #34 and #36

([Link]) Interpret, in context, the value of a. [1]

The student uses this model to predict the population of Canada in the year
2030, where t = 30, and calculates a population of approximately 41. 3 million

people.

This question is duplicated in #34 and #36

(b) Comment on the reliability of the student’s prediction. [1]

A data scientist, Benoit, uses additional information to develop an exponential


model for Canada’s future population.

In this model, B(t) = 33. 5(1. 005) represents the millions of people in
t

Canada t years after the year 2000, where 25 ≤ t ≤ 100.

This question is duplicated in #34 and #36

(c.i) Use Benoit’s model to predict the population of Canada in the


year 2100. [2]

This question is duplicated in #34 and #36

([Link]) Interpret, in context, the value 1. 005 in Benoit’s model. [1]


Another data scientist, Cecilia, develops a third model for the Canadian
population.

In this model, C(t) =


1+e
62
−0.02t
represents the millions of people in Canada t
years after the year 2000, where 25 ≤ t ≤ 100 .

This question is duplicated in #34 and #36

(d) Use Cecilia’s model to predict the population of Canada in the


year 2100. [1]

This question is duplicated in #34 and #36

(e) Determine the year in which the difference between the


predictions from Benoit’s model and Cecilia’s model is greatest. [3]

This question is duplicated in #34 and #36

(f ) Find the value of

This question is duplicated in #34 and #36

(f.i) B′(75); [1]

This question is duplicated in #34 and #36

([Link]) C′(75). [1]

This question is duplicated in #34 and #36


(g) Compare and interpret, in context, the values of B′(75) and
C′(75). [2]

39. [Maximum mark: 6] [Link].TZ1.2


Claire rolls a six-sided die 16 times.

The scores obtained are shown in the following frequency table.

It is given that the mean score is 3.

This question is duplicated in #42

(a) Find the value of p and the value of q. [5]

Each of Claire’s scores is multiplied by 10 in order to determine the final score for
a game she is playing.

This question is duplicated in #42

(b) Write down the mean final score. [1]


40. [Maximum mark: 17] [Link].TZ1.9
A bag contains buttons which are either red or blue.

Initially, the bag contains three red buttons and one blue button.

Francine randomly selects one button from the bag. She then replaces the
button and adds one extra button of the same colour.

For example, if she selects a red button, she then replaces it and adds one extra
red button so that the bag then contains four red buttons and one blue button.

Francine then randomly selects a second button from the bag.

The following tree diagram represents the probabilities of the first two
selections.
(a) Find the value of p and the value of q. [2]

(b) Show that the probability that Francine selects two buttons of
the same colour is 10
7
. [2]

(c) Given that Francine selects two buttons of the same colour, find
the probability that she selects two red buttons. [3]

The random variable X is defined as the number of red buttons selected by


Francine.

The following table shows the probability distribution of X.

(d) Find the value of a and the value of b. [2]

(e) Hence, find the expected number of red buttons selected by


Francine. [2]

Francine restarts the process with three red buttons and one blue button in the
bag. She selects buttons as before, replacing the button and adding one extra
button of the same colour each time. She repeats this until she selects a blue
button.

(f ) Given that the first two buttons she selects are red, write down
the probability that the next button she selects is blue. [1]

The probability that she selects the first blue button after n selections in total is
3

56
.

(g) Find the value of n. [5]


41. [Maximum mark: 6] [Link].TZ2.5
A species of bird can nest in two seasons: Spring and Summer.

The probability of nesting in Spring is k.

The probability of nesting in Summer is k

2
.

This is shown in the following tree diagram.

This question is duplicated in #43

(a) Complete the tree diagram to show the probabilities of not


nesting in each season. Write your answers in terms of k. [2]

It is known that the probability of not nesting in Spring and not nesting in
Summer is 5

9
.

This question is duplicated in #43


(b.i) Show that 9k 2
− 27k + 8 = 0 . [3]

This question is duplicated in #43

([Link]) Both k =
1

3
and k =
8

3
satisfy 9k 2
− 27k + 8 = 0 .

State why k =
1
3
is the only valid solution. [1]

42. [Maximum mark: 6] [Link].TZ1.1


Claire rolls a six-sided die 16 times.

The scores obtained are shown in the following frequency table.

It is given that the mean score is 3.

This question is duplicated in #39

(a) Find the value of p and the value of q. [5]

Each of Claire’s scores is multiplied by 10 in order to determine the final score for
a game she is playing.
This question is duplicated in #39

(b) Write down the mean final score. [1]

43. [Maximum mark: 6] [Link].TZ2.4


A species of bird can nest in two seasons: Spring and Summer.

The probability of nesting in Spring is k.

The probability of nesting in Summer is k

2
.

This is shown in the following tree diagram.

This question is duplicated in #41

(a) Complete the tree diagram to show the probabilities of not


nesting in each season. Write your answers in terms of k. [2]
It is known that the probability of not nesting in Spring and not nesting in
Summer is 5

9
.

This question is duplicated in #41

(b.i) Show that 9k 2


− 27k + 8 = 0 . [3]

This question is duplicated in #41

([Link]) Both k =
1
3
and k =
8
3
satisfy 9k 2
− 27k + 8 = 0 .

State why k =
1

3
is the only valid solution. [1]

44. [Maximum mark: 7] [Link].TZ1.1


Janie claims that rabbits in Australia have longer ears than rabbits in Spain.

To test her claim, a randomly selected sample of rabbits was collected in each
country.

The length of one ear of each rabbit was measured and the value recorded
correct to the nearest millimetre (mm).

In the Australian sample, the median recorded value was 80 mm and the
interquartile range was 11 mm.

The recorded values for the Spanish sample are shown in the following box and
whisker diagram.
(a) Complete the following table for the recorded values of the
lengths of the rabbits’ ears in each sample.

[3]

(b) Justifying your answers, compare the distributions of the


lengths of rabbits’ ears in Australia and Spain using

(b.i) the median; [2]

([Link]) the interquartile range. [2]

45. [Maximum mark: 7] [Link].TZ1.6


A class is given two tests, Test A and Test B. Each test is scored out of a total of 100
marks. The scores of the students are shown in the following table.

Let x be the score on Test A and y be the score on Test B.

The teacher finds that the equation of the regression line of y on x for these
scores is y = 0. 822x + 18. 4.
(a) Find the value of Pearson’s product-moment correlation
coefficient, r. [2]

Giovanni was absent for Test A and Paulo was absent for Test B.

The teacher uses the regression line of y on x to estimate the missing scores.

Paulo scored 10 on Test A.

The teacher estimated his score on Test B to be 27 to the nearest integer using
the following calculation:

y = 0. 822(10) + 18. 4 ≈ 27

(b) Give a reason why this method is not appropriate for Paulo. [1]

Giovanni scored 90 on Test B.

The teacher estimated his score on Test A to be 87 to the nearest integer using
the following calculation:

90 = 0. 822x + 18. 4 , so x =
90−18.4

0.822
≈ 87

(c.i) Give a reason why this method is not appropriate for Giovanni. [1]

([Link]) Use an appropriate method to show that the estimated Test A


score for Giovanni is 86 to the nearest integer. [3]

46. [Maximum mark: 6] [Link].TZ0.4


A six-sided biased die is weighted in such a way that the probability of obtaining
a “six” is 7

10
.

(a) The die is tossed five times. Find the probability of obtaining at
most three “sixes”. [3]
(b) The die is tossed five times. Find the probability of obtaining
the third “six” on the fifth toss. [3]

© International Baccalaureate Organization, 2026

You might also like