Statistics & Probability
Statistics & Probability
Lecture notes by
Ahmed Ibrahim
2025
i
TABLE OF CONTENTS
i
CHAPTER 4 ............................................................................................................................ 38
4. Probability ........................................................................................................................ 38
4.1 Sample Spaces and Events ............................................................................................. 38
4.1.2 Sample space ........................................................................................................... 39
CHAPTER 5 ........................................................................................................................... 67
References ................................................................................................................................ 73
ii
CHAPTER 1
1. REPRESENTATION OF DATA
1
1.3 Types of Data
There are two types of data: qualitative (or categorical) data are described by words
and are non-numerical, such as blood types or colors. Quantitative data take numerical values
and are either discrete or continuous. As a general rule, discrete data are counted and cannot
be made more precise, whereas continuous data are measurements that are given to a chosen
degree of accuracy. Discrete data can take only certain values, as shown in the diagram. The
number of letters in the words of a book is an example of discrete quantitative data. Each word
1
has 1 or 2 or 3 or 4 or… letters. There are no words with 3 or 4.75 letters. Discrete
2
quantitative data can take non-integer values. For example, United States coins have dollar
values of 0.01, 0.05, 0.10, 0.25, 0.50 and 1.00. In Canada, the United Kingdom and other
1 1
countries, shoe sizes such as 62 , 7 and 72 are used. Continuous data can take any value
(possibly within a limited range), as shown in the diagram. The times taken by the athletes to
complete a 100-metre race is an example of continuous quantitative data. We can measure these
to the nearest second, tenth of a second or even more accurately if we have the necessary
equipment. The range of times is limited to positive real numbers. It can be concluded that
Discrete data can take only certain values. Continuous data can take any value, possibly
within a limited range.
2
To present the data in a stem-and-leaf diagram, we first group the scores into suitable
equal-width classes. Class widths of 10 are suitable here, as shown below.
5 8583
6 171492
7 297
8
9 72
Next, we arrange the scores in each row in ascending order from left to right and add a
key to produce the stem-and-leaf diagram shown below.
Key:
5 3588 5 3
6 112479 Represents a score of
7 279 53%
8
9 27
In a back-to-back stem-and-leaf diagram, the leaves to the stem's right ascend left to
right, and the leaves on the left of the stem ascend right to left (as shown in Worked example
1.2).
Example 1.2
The number of days on which rain fell in a certain town in each month of 2016 and
2017 are given.
Year 2016
Jan: 17 Feb: 20 Mar: 13 Apr: 12 May: 10 Jun: 8
Jul: 0 Aug: 1 Sep: 5 Oct: 11 Nov: 16 Dec: 9
Year 2017
Jan: 9 Feb: 13 Mar: 11 Apr: 8 May: 6 Jun: 3
Jul: 1 Aug: 2 Sep: 2 Oct: 4 Nov: 8 Dec: 7
Display the data in a back-to-back stem-and-leaf diagram and briefly compare the
rainfall in 2016 with the rainfall in 2017.
3
Key:
5 0 6
2016 2017
Represents days in a
98510 0 1223467889
month of 2016 and 6
763210 1 13
0 2 days in a month of 2017
It rained on more days in 2016 (122 days) than it did in 2017 (74 days).
Note That: If rows of leaves are particularly long, repeated values may be used in the stem.
However, if there were, say, 30 leaves in one of the rows, we might consider grouping the data
into narrower classes of 0–4, 5–9,10–14,15–19 and 20–24. This would require 0, 0, 1,1 and 2
in the stem.
4
1.5 Representation of continuous data: histograms
Continuous data are given to a certain degree of accuracy, such as 3 significant figures,
2 decimal places, to the nearest 10 and so on. We usually refer to this as rounding. When values
are rounded, gaps appear between classes of values, and this can lead to a misunderstanding of
continuous data because those gaps do not exist. For example, consider heights to the nearest
centimeter, given as 146: 150, 151: 155 and 156: 160. Gaps of 1cm appear between classes
because the values are rounded. Using h for height, the actual classes should be 145.5≤h<150.5,
150.5≤h<155.5 and 155.5≤h < 160.5cm.
The classes are shown in the diagram below, with the lower and upper boundary
values and the class mid-values (also called midpoints) indicated.
A histogram is best suited to illustrating continuous data, but it can also be used to
illustrate discrete data. We might have to group the data ourselves or it may be given to us in a
grouped frequency table. For example, the presented in the tables below show the ages and
the percentage scores of 100 students who took an examination.
A histogram is a chart that uses bars to show the distribution of values for a numeric
variable. Each bar represents a range of values, and its height shows how many data points fall
within that range.
5
The first table shows three classes of continuous data; there are no gaps between the
classes and the classes have equal-width intervals of 2 years. This means that we can represent
the data in a frequency diagram by drawing three equal-width columns with column heights
equal to the class frequencies, as shown below.
The following table shows the areas of the columns and the frequency of each of the
three classes presented in the diagram on the previous page.
From this table, we can see that the ratio of the column areas, 68: 92: 40, is the same as
the ratio of the frequencies, 34: 46: 20. In a histogram, the area of a column represents the
frequency of the corresponding class so that the area must be proportional to the frequency. We
may see this written as ‘area _ frequency’. This also means that in every histogram, just as in
the example above, the ratio of column areas is the same as the ratio of the frequencies, even
if the classes do not have equal widths. Also, there can be no gaps between the columns in a
histogram because one class's upper boundary equals the neighboring class's lower boundary.
A gap can appear only when a class has zero frequency. The axis showing the measurements
is labeled as a continuous number line, and the width of each column is equal to the width of
the class it represents. When we construct a histogram, since the classes may not have equal
widths, the height of each column is no longer determined by the frequency alone but must be
6
calculated so that area _ frequency. The vertical axis of the histogram is labeled frequency
density, which measures frequency per standard interval.
The simplest and most commonly used standard interval is 1 unit of measurement. For
example, a column representing 85 objects with masses from 50 to 60 kg has a frequency
density of 85 objects / (60-50) kg = 8.5 objects per kilogram, and so on.
Note That: For a standard interval of 1 unit of measurement, Frequency density = class
frequency/ class width, which can be re-arranged to give
Class frequency= class width * frequency density
Example 1.3
The masses (m), kg, of 100 children are grouped into two classes, as shown in the table.
7
b. There are children with masses from 45 to 63kg in both classes, so we must split this
interval into two parts: 45–50 and 50–63.
Frequency= width* frequency density
For first part (45–50) = (50-45) * 4= 20 children
For second part (50–63) = (63-50) * 3= 39 children
The estimation number of children = 20 + 39 = 59 children
Example 1.4
Consider the times taken, to the nearest minute, for 36 athletes to complete a race, as
given in the table below.
Time taken (min) 13 14-15 16-18
No. athletes (f) 4 14 18
8
a. the number of athletes who took less than 13.0 minutes
Frequency= width* frequency density = (13-12.5) *4= 2 athletes
b. the number of athletes who took between 14.5 and 17.5 minutes
Frequency= width* frequency density = (15.5-14.5) *7 + (17.5-15.5)6= 19 athletes
c. the time taken to run the race by the slowest of three athletes.
"When we refer to the slowest athlete, we are talking about the one who takes the longest
time."
Frequency= width* frequency density
3 = width * 6 3 = (18.5 - t) * 6
so, the time taken= 18.5-0.5=18min
9
Example 1.5
The following table shows the lengths of 80 leaves from a particular tree, given to the
nearest centimeter.
Lengths (cm) 1-2 3-4 5-7 8-9 10-11
No. leaves (f) 8 20 38 10 4
Draw a cumulative frequency curve and a cumulative frequency polygon. Use each of these
to estimate:
a. the number of leaves that are less than 3.7cm long
b. the lower boundary of the lengths of the longest 22 leaves.
Answer:
Lengths (cm) Addition of frequencies No. leaves
L<0.5 0 0
L<2.5 0+8 8
L<4.5 0+8+20 28
L<7.5 0+8+20+38 66
L<9.5 0+8+20+38+10 76
L<11.5 0+8+20+38+10+4 80
a:
➢ The polygon gives an estimate of 20 leaves.
➢ The curve gives an estimate of 18 leaves.
10
b
➢ The polygon gives an estimate of 6.9cm.
➢ The curve gives an estimate of 6.7cm.
1.7 Comparing different data representations
Pictograms, bar charts and pie charts are useful ways of displaying qualitative data and
ungrouped quantitative data, and people generally find them easy to understand. Nevertheless,
it may be of benefit to group a set of raw data so that we can see how the values are distributed.
Knowing the proportion of small, medium and large values.
Pictograms are types of charts and graphs that use icons and images to represent data.
A bar chart or bar graph is a chart or graph that presents categorical data with
rectangular bars with heights or lengths proportional to the values that they represent.
A pie chart is a type of graph representing data in a circular form, with each slice of
the circle representing a fraction or proportionate part of the whole.
Dot plot, Bar charts, and vertical line:
➢ Each dot represents a specific number of observations.
➢ The dots are stacked in a column over a category. The height of the column
represents the absolute frequency of observation.
➢ Used most often to plot frequency counts within a small number of categories.
Line graph is a type of graph that shows information that is connected in some ways.
11
CONCLUSION
Variable
quantitative Qualitative
(Numeric) (Categorical)
12
SHEET 1
1. Twenty people leaving a cinema are each asked, “How many times have you attended
the cinema in the past year?” Their responses are:
6, 2, 13, 1, 4, 8, 11, 3, 4, 16, 7, 20, 13, 5, 15, 3, 12, 9, 26 and 10.
Construct a stem-and-leaf diagram for this data and include a key.
2. A shopkeeper takes 12 bags of coins to the bank. The bags contain the following
numbers of coins:
150, 163, 158, 165, 172, 152, 160, 170, 156, 162, 159 and 175.
a. Represent this information in a stem-and-leaf diagram.
b. Each bag contains coins of the same value, and the shopkeeper has at least one bag
containing coins with dollar values of 0.10, 0.25, 0.50 and 1.00 only.
What is the greatest possible value of all the coins in the 12 bags?
4. In a particular city there are 51 buildings of historical interest. The following table
presents the ages of these buildings, given to the nearest 50 years.
Age (years) 50-150 200-300 350-450 500-600
No. buildings (f) 15 18 12 6
a. Write down the lower and upper boundary values of the class containing the greatest
number of buildings.
b. State the widths of the four class intervals.
c. Illustrate the data in a histogram.
d. Estimate the number of buildings that are between 250 and 400 years old.
5. The masses, m grams, of 690 medical samples, are given on the following table.
13
Mass (m grams) 4≤ m<12 12≤ m<24 24≤ m<28
No. medical samples (f) 200 400 p
6. A university investigated how much space on its computers’ hard drives is used for
data storage. The results are shown below. It is given that 40 hard drives use less than 20GB
for data storage.
7. The following table shows the widths of the 70 books in one section of a library,
given to the nearest centimeter.
Width (cm) 10-14 15-19 20-29 30-39 40-44
No. books (f) 3 13 25 24 5
a. Given that the upper boundary of the first class is 14.5cm, write down the upper boundary
of the second class.
b. Draw up a cumulative frequency table for the data and construct a cumulative frequency
graph.
c. Use your graph to estimate:
i. the number of books that have widths of less than 27cm
ii the widths of the widest 20 books.
14
CHAPTER 2
15
Score on die 1 2 3 4 5 6
Frequency (f) 5 6 5 3 2 4
In a set of grouped data in which raw values cannot be seen, we can find the modal
class, which is the class with the highest frequency density.
EXAMPLE 2.1
Find the modal class of the 270 pencil lengths, given to the nearest centimeter in the
following table.
Length (cm) 4-7 8-10 11-12
No. pencils (f) 100 90 80
Answer
Length (cm) 3.5-7.5 7.5-10.5 10.5-12.5
No. pencils (f) 100 90 80
Width 4 3 2
Frequency density 25 30 40
The modal class is 11–12cm (or, more accurately, 10.5<x < 12.5cm).
EXAMPLE 2.2
Two classes of data have interval widths in the ratio 3:2. Given that there is no modal
class, and that the frequency of the first class is 48, find the frequency of the second class.
Answer
Note That: No modal class means that the frequency densities of the two classes are equal
16
You will soon be performing calculations involving the mean, so here we introduce
notation that is used in place of the word definition used above. We use the upper-case Greek
letter ‘sigma’, written ∑, to represent ‘sum’ and 𝑥̅ to represent the mean, where x represents
our data values. The notation used for ungrouped and for grouped data are shown on separate
rows in the following table.
` Sum Data Frequency Number of Sum of Mean
values of Data data
data values values values
Ungrouped ∑ x - n ∑x ∑x
𝐱̅ =
n
EXAMPLE 2.3
Five persons, whose mean mass is 70.2kg, wish to go to the top of a building in a lift
with some cement. Find the greatest mass of cement they can take if the lift has a maximum
weight allowance of 500 kg.
Answer
17
EXAMPLE 2.4
Find the mean of the 40 values of x given in the following table.
x 31 32 33 34 35
F 5 7 9 8 11
Answer
x 31 32 33 34 35
F 5 7 9 8 11 ∑f=40
X* F 155 224 297 272 385 ∑xf=1333
∑𝐱𝐟 𝟏𝟑𝟑𝟑
̅=
𝒙 = = 33.325
∑𝐟 𝟒𝟎
EXAMPLE 2.5
A large bag of sweets claims to contain 72 sweets, having a total mass of 852.4g. A
small bag of sweets claims to contain 24 sweets, having a total mass of 282.8g. What is the
mean mass of all the sweets together?
Answer
Total number of sweets=72+24=96.
Total mass of sweets= 852.4 + 282.8 = 1135.2g
Mean mass= 1135.2 / 96 = 11.825g
18
2.3.2 Means from grouped frequency tables
When data are presented in a grouped frequency table or illustrated in a histogram or
cumulative frequency graph, we lose information about the raw values. For this reason, we
cannot determine the mean exactly, but we can calculate an estimate of the mean. We do this
∑xf
̅=
by using mid-values to represent the values in each class. We use the formula: ∑𝐱 to
∑𝐟
calculate an estimate of the mean, where x now represents the class mid-values.
EXAMPLE 2.6
Coconuts are packed into 75 crates, with 40 of a similar size in each crate. 46 crates
contain coconuts with a total mass of 20 up to but not including 25 kg. 22 crates contain
coconuts with a total mass of 25 up to, but not including 40 kg. 7 crates contain coconuts with
a total mass of 40 up to but not including 54 kilograms.
a Calculate an estimate of the mean mass of a crate of coconuts.
b Use your answer to part a to estimate the mean mass of a coconut.
Answer
Mass (kg) 20- 25- 40-54
Frequency (f) 46 22 7 ∑=75
Mid value (x) 22.5 32.5 47
X* F 1035 715 329 ∑=2079
∑xf
̅=
The mean mass of a crate = ∑𝐱 = 2079 / 75 = 27.72 kg
∑𝐟
19
EXAMPLE 2.7
Calculate an estimate of the mean age of a group of 50 students, where there are sixteen
18-year-olds, twenty 19-year-olds and fourteen who are either 20 or 21 years old.
Answer
Age (year) 18- 19- 20-22
Frequency (f) 16 20 14 ∑=50
Mid value (x) 18.5 19.5 21
X* F 296 390 294 ∑=980
We subtracted 100 from each x value, so we simply add 100 to the mean of the coded
values to find the mean of x.
Mean(x)=mean(x–100) +100 = 106 or
∑(𝒙− 𝟏𝟎𝟎)
𝑥̅ = + 100 =106
𝟓
20
It can be concluded that:
∑(𝒙− 𝐛)
For ungrouped data, 𝑥̅ = +b
𝐧
∑(𝒙− 𝐛)𝐟
For grouped data, 𝑥̅ = +b
∑𝐟
̅=mean(x–b) +b.
These formulae can be summarized by writing 𝒙
EXAMPLE 2.8
The exact age of an individual boy is denoted by b, and the exact age of an individual
girl is denoted by g. Exactly 5 years ago, the sum of the ages of 10 boys was 127.0 years, so∑(
(b−5) =127.0. In exactly 5 years’ time, the sum of the ages of 15 girls will be 351.0 years, so
∑(𝑥+ b) ∑(𝑔+ 5)
b. 𝑥̅ = –b= – 5 = 351 / 15 – 5 = 18.4 years
n 15
∑(b − 5) = ∑b − (n ∗ 5)
so ∑b = 127 + 10 ∗ 5 = 177
∑(g + 5) = ∑g + (n ∗ 5)
so ∑g = 351 − 15 ∗ 5 = 276
∑b+ ∑g
c. the mean age of combined =
n
177+ 276
= = 18.2 years
10+15
21
EXAMPLE 2.8
Forty values of x are coded in the following table.
x–3 0– 18– 24–32
Frequency 9 13 18
EXAMPLE 2.9
For the 20 values of x summarized by ∑ (2x−3) = 104, find 𝑥
̅.
𝟏𝟎𝟒
2𝑥̅ = + 3 =8.2
𝟐𝟎
𝑥̅ = 4.1
22
2.4 The median
You will recall that the median splits data into two parts with an equal number of values
in each part: a bottom half and a top half. In a set of n-ordered values, the median is the value
halfway between the 1st and the nth. Consider a DIY store that opens for 12 hours on Monday
and 15 hours on Saturday. The following back-to-back stem-and-leaf diagram shows the
numbers of customers served during each hour on Monday and Saturday last week.
Monday (12) Saturday (15)
863100 2 2346 Key:
43110 3 556899 0 2 2
1 4 01379 Represents customers in a
Monday of 20 and in a
Saturday of 22
To find the median number of customers served on each of these days, we need to find
their positions in the ordered rows of the back-to-back stem-and-leaf diagram.
For Saturday, there are n = 15 values arranged in ascending order from top to bottom
𝑛+1 15+1
and from left to right. The median is at the { } th = { } = 8th value. In the first row, we
2 2
have the 1st to 4th values, and in the second row we have the 5th to 10th values, so the 8th
value is 38. The median number of customers on Saturday was 38. For Monday, there are n =
12 values arranged in ascending order from top to bottom and from right to left. The median is
𝑛+1 12+1
at the { } th = { } = 6.5th value, so we locate the median mid-way between the 6th and
2 2
7th values. In the first row, we have the 1st to 6th values and the 6th is 28. The first value in
the second row is the 7th value, which is 30. The median number of customers on Monday
28+30
was { } = 29. When data appear in an ordered frequency table of individual values, we can
2
use cumulative frequencies to investigate the positions of the values, knowing that the median
𝑛+1
is at the { } value.
2
EXAMPLE 2.10
The following table shows 65 ungrouped readings of x. Cumulative frequencies, and
the positions of the readings are also shown. Find the median value of x.
X F Cf
40 11 11
41 23 34
23
42 19 53
43 8 61
44 4 65
𝐧+𝟏 𝟔𝟓+𝟏
The total frequency is 65, and = = 33, so the median is at the 33rd value. From
𝟐 𝟐
the table, we see that the 12th to 34th values are all equal to 41.
24
Sheet 2
1. Find the mode(s) of the following sets of numbers.
a. 12,15,11, 7, 4,10, 32,14, 6,13,19, 3
b. 19, 21, 23,16, 35, 8, 21,16,13,17,12,19,14, 9
y -4 -3 -2 -1 0
f 27 28 29 27 25
4. Find the modal class for x and for y in the following tables.
x 0- 4- 14- 20
f 5 9 8
y 3- 6 7- 11 12- 20
f 66 80 134
6. The mean of 15, 31, 47, 83, 97, 119 and p2 is 63. Find the possible values of p.
7. The mean of 6, 29, 3, 14, q, (q+ 8), q2 and (10 – q) is 20. Find the possible values of q.
9. An examination was taken by 50 students. The 22 boys scored a mean of 71% and
the girls scored a mean of 76%. Find the mean score of all the students.
25
10. The following table summarizes the number of tomatoes produced by the plants in
the plots on a farm.
No. tomatoes 20-29 30-49 50-79 80-100
No. plots (f) 329 413 704 258
a. Calculate an estimate of the mean number of tomatoes produced by these plots.
̅ = 7.4. Find:
11. 10. For 10 values denoted x, it is given that 𝒙
a. ∑x b. ∑ (x+2) c. ∑ (x−1)
̅.
12. Twenty-five values of z are such that ∑ (z−7) = 275. Find 𝒁
13. Given 𝐪
̅ = 22 and ∑ (q−4) = 3672, find the number of values of q.
14. The lengths of 2500 bolts, x mm, are summarized by ∑ (x−40) = 875. Find the mean
length of the bolts.
26
CHAPTER 3
3. MEASURES OF VARIATION
A measure of central tendency alone does not describe or summarize a set of data fully.
Although it may tell us the location of the more central values or the most common values, it
tells us nothing about how widely spread out the values are. Two sets of data can have the same
mean, median or mode, yet they can be completely different. A better description of a set of
data is given by a measure of central tendency and a measure of variation. Variation is also
known as spread or dispersion. Consider the runs scored by two batters in their past eight
cricket matches, which are given in the following table.
The mean number of runs scored by A and by B is the same; namely, 224÷8=28.
However, the patterns of the number of runs are clearly very different. The numbers for batter
A are quite consistent, whereas the numbers for batter B are quite varied. This consistency (or
lack of it) can be indicated by a measure of variation, which shows how spread out a set of data
values are. Three commonly used measures of variation are the range, interquartile range
and standard deviation.
27
For example, in a test for which the lowest mark is 6 and the highest mark is 19, the
range is 19 – 6 = 13. For grouped data, we can find a minimum and maximum possible range,
using the lower and upper boundary values of the data.
EXAMPLE 3.1
To the nearest centimeter, the tallest and shortest pupils in a class are 169cm and 150cm.
Find the least and greatest possible range of the students’ heights.
Answer
The intervals in which the given heights, h, lie are 168.5≤h < 169.5cm and 149.5≤h
<150.5cm.
Least possible range 168.5 –150.5= 18cm
Greatest possible range= 169.5 –149.5= 20cm
28
Ungrouped data
The positions of the lower and upper quartiles depend on whether there are an odd or
even number of values in the set of data. One method that we can use to find the quartiles is as
follows. For an even number of ordered values: we split the data into a lower half and an upper
half. Then Q1 and Q3 are the medians of the lower half and upper half, respectively. For an odd
number of ordered values: we split the data into a lower half and an upper half at the median,
which we then discard. Again, Q1 and Q3 are the medians of the lower half and upper half,
respectively.
EXAMPLE 3.2
Find the interquartile range of the eight ordered values 2, 5, 9, 13, 29, 33, 49 and 55.
Answer
2 5 9 13 29 33 49 55
Q1 Q2 Q3
5+ 9
Q1 = = 7,
2
13+29
Q2 = = 21,
2
33+ 49
Q3 = = 41
2
33+ 49 5+ 9
IQR= Q3- Q1= - = 41 – 7 = 34
2 2
29
EXAMPLE 3.3
Find the interquartile range of the seven values 69, 17, 43, 6, 73, 77 and 39.
6 17 39 43 69 73 77
Q1 Q2 Q3
Answer
IQR= Q3- Q1 = 73-17= 56
EXAMPLE 3.4
Find the interquartile range of the 13 grouped values shown in the following stem-and-
leaf diagram.
Answer:
➢ We identify the median as 153, which we now discard. This leaves a lower half
(142 to 151) and an upper half (155 to 168), with six values in each.
Q1
14 2 2 4 8 9 Q3
15 1 3 5 6 7 9
16 5 8
144+ 148
Q1 = = 146,
2
Q2 = 153,
157+ 159
Q3 = = 158
2
30
Grouped data
We can use a cumulative frequency graph to estimate values in any position in a set of
data. This includes the lower quartile, the upper quartile and any chosen percentile.
For grouped data with total frequency n = ∑f, the positions of the quartiles are shown
in the following table.
The nth percentile is the value that is n% of the way through a set of data. Q1, Q2 and
Q3 are the 25th, 50th and 75th percentiles, respectively. In an ordered dataset with, say, 320
values, Q1, Q2 and Q3 are at the 80th, 160th and 240th values, and the 90th percentile is at the
(0.90×320) =288th value. The range of the middle 80% of a dataset is the difference between
the 10th and 90th percentiles.
EXAMPLE 3.5
The following graph illustrates the times, in minutes, taken by 500 people to complete
a task. Use the graph to find an estimate of:
a. the greatest possible range b. the interquartile range c. the 95th
percentile.
31
Answer:
a. The greatest possible range is equal to the width of the polygon
30 – 2 = 28min
b. We locate the quartiles, then estimate their values by reading from the graph.
n 500
Lower quartile: = = 125th value
4 4
3n 1500
Upper quartile: = = 375𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4
Q1 = 8.0 min
Q3 = 14.5min
IQR = 14.5 – 8 = 6.5 min
c. The 95th percentile is at the (0.95×500) = 475th value
=24.0min
32
Box-and-whisker diagrams
A box-and-whisker diagram (or box plot) is a graphical representation of data, showing
some of its key features. These features are its smallest and largest values, its lower and upper
quartiles, and its median. If drawn by hand, the diagram is best drawn on graph paper and must
include a scale. It takes the form shown in the following diagram, which shows some features
of a dataset denoted by x.
33
3.3 Variance and standard deviation
If we want a measure of variation around the mean, we need to ensure that each
deviation is positive or zero. We can do this by calculating the mean distance of the data values
̅|
∑|𝑿− 𝑿
from the mean, which we call the ‘mean absolute deviation from the mean’,
𝒏
However, it is hard to calculate this accurately or efficiently for large sets of data and it
is difficult to work with algebraically, so this approach is not used in practice. Alternatively,
̅ )2 for all data values and find their mean. This
we can calculate the squared deviation, (𝑿 − 𝑿
is the ‘mean squared deviation from the mean’, which we call the variance of the data.
̅ )2
∑(𝑿− 𝑿
Var(X) =
𝒏
For measurements and deviations in metres, say, the variance is in m2. So, to get a
measure of variation that is also in metres, we take the square root of the variance, which we
call the standard deviation.
∑(𝑿− 𝑿) ̅ 2
Standard deviation of x = √𝐕𝐚𝐫 = √ 𝒏
The formula for variance can be simplified (see appendix at the end of this chapter) to give:
∑𝑿2
Var(X)= ̅2
-𝑿
𝐧
We can find the variance and standard deviation from n, x and x2, which are the number
of values, their sum and the sum of their squares, respectively. We often use the abbreviation
SD(X) to represent the standard deviation of X.
Standard deviation of x for ungrouped ̅)
∑(𝑿− 𝑿
2
data √𝐕𝐚𝐫 = √ 𝒏
Standard deviation of x for grouped data
̅ )2 ∗𝒇 2
√𝐕𝐚𝐫 = √
∑(𝑿− 𝑿
= ̅2
√∑𝑿 ∗𝒇 − 𝑿
∑𝒇 ∑𝒇
EXAMPLE 3.6
For the set of five numbers 3, 9,15, 24 and 29, find the standard deviation
Answer:
2 2
̅)
∑(𝑿− 𝑿 ∑(𝑿) 2
the standard deviation = √𝐕𝐚𝐫 = √ =√ 𝒏 ̅
− 𝑿
𝒏
34
EXAMPLE 3.7
Find the standard deviation of the values of x given in the following table, correct to 3
significant figures.
x f
12 13
14 28
16 10
Answer
x X2 f X*f X2*f
12 144 13 156 1872
14 196 28 392 5488
16 256 10 160 2560
∑ - 51 708 9920
̅ )2 ∗𝒇 2
SD= √𝐕𝐚𝐫 = √ ∑(𝑿− 𝑿
= ̅ 2 =√9920 − (𝟕𝟎𝟖/𝟓𝟏)2 =1.34
√∑𝑿 ∗𝒇 − 𝑿
∑𝒇 ∑𝒇 𝟓𝟏
EXAMPLE 3.8
Calculate an estimate of the standard deviation of the heights of the 20 children given
in the following table.
Height (metres) No. 1.2– 1.4– 1.5–1.7
children (f) 2 12 6
35
Sheet 3
1. For each ordered set of data, A to D, write down these five values: the smallest value;
the lower quartile; the median; the upper quartile; and the largest value.
2. Find the range and the interquartile range of the following sets of data.
3. Find the range and the interquartile range of the dataset represented in the following
box plot.
36
b Illustrate the data in a box-and-whisker diagram on graph paper and include a
scale.
6. Find the mean and the standard deviation for these sets of numbers.
7. The following table shows the number of pets owned by each of 35 families.
No. pets 0 1 2 3 4 5
No. families (f) 6 12 9 4 3 1
Find the mean and variance of the number of pets.
8. The times spent, in minutes, by 30 girls and by 40 boys on an assignment are detailed
in the following table.
37
CHAPTER 4
4. PROBABILITY
Random Experiment
An experiment <with known outcomes> that can result in different outcomes, even
though it is repeated in the same manner every time, is called a random
experiment.
38
4.1.2 Sample space
To model and analyze a random experiment, we must understand the set of possible
outcomes from the experiment. In this introduction to probability, we use the basic concepts
of sets and operations on sets. It is assumed that the reader is familiar with these topics.
Sample Space
The set of all possible outcomes of a random experiment is called the sample space
of the experiment. The sample space is denoted as S.
sample space, or simply a sample point. If the sample space has a finite number
of elements, we may list the members separated by commas and enclosed in
braces. Thus, the sample space S, of possible outcomes when a coin is flipped,
may be written S = {H, T}, where H and T correspond to heads and tails,
respectively.
Example 4.1:
Consider the experiment of tossing a die. If we are interested in the number that shows
40
Example 4.3:
Suppose that three items are selected at random from a manufacturing process. Each
item is inspected and classified defective, D, or non-defective, N.
To list the elements of the sample space providing the most information, we construct
the tree diagram as been show in the Figure. Now, the various paths along the branches of the
tree give distinct sample points. Starting with the first path, we get the sample point DDD,
indicating the possibility that all three items inspected are defective. As we proceed along the
other paths, we see that the sample space is
S = {DDD, DDN, DND, DNN, NDD, NDN, NND, NNN}.
41
Sample spaces with a large or infinite number of sample points are best described by a
statement or rule method. For example, if the possible outcomes of an experiment are the set
of cities in the world with a population over 1 million, our sample space is written S = {x | x is
a city with a population over 1 million}, which reads “S is the set of all x such that x is a city
with a population over 1 million.” The vertical bar is read “such that.” Similarly, if S is the set
of all points (x, y) on the boundary or the interior of a circle of radius 2 with center at the origin,
we write the rule
S = {(x, y) | x2 + y2 ≤ 4}.
Example 4.4:
Consider an experiment that selects a cell phone camera and records the recycle time
of a flash (the time taken to ready the camera for another flash). The possible values for this
time depend on the resolution of the timer and on the minimum and maximum recycle times.
However, because the time is positive, it is convenient to define the sample space as simply the
positive real line
42
S = R+ = {x | x > 0}
If it is known that all recycle times are between 1.5 and 5 seconds, the sample space
can be S = {x | 1.5 < x < 5}.
It is useful to distinguish between two types of sample spaces.
Discrete and Continuous Sample Spaces
A sample space is discrete if it consists of a finite or countable infinite set of outcomes.
A sample space is continuous if it contains an interval (either finite or infinite) of real
numbers.
4.2 Events
For any given experiment, we may be interested in the occurrence of certain events
rather than in the occurrence of a specific element in the sample space. For instance, we may
be interested in the event A that the outcome when a die is tossed is divisible by 3. This will
occur if the outcome is an element of the subset A = {3, 6} of the sample space S1 in Example
4.1.
As a further illustration, we may be interested in the event B that the number of
defectives is greater than 1 in Example 4.3. This will occur if the outcome is an element of the
subset
B = {DDN, DND, NDD, DDD} of the sample space S.
To each event we assign a collection of sample points, which constitute a subset of the
sample space. That subset represents all of the elements for which the event is true.
Event
An event is a subset of the sample space of a random experiment.
Example 4.5:
A dice is rolled twice. What is the Event that the sum of the facesisgreaterthan7, given
that the first outcome was a 4?
𝑆= {11, 12, 13, 14, 15, 16, 21, 22, 23, 24, 25, 26, 31, 32, 33, 34, 35, 36, 𝟒𝟏, 𝟒𝟐, 𝟒𝟑, 𝟒𝟒, 𝟒𝟓,
𝟒𝟔, 51, 52, 53, 54, 55, 56, 61, 62, 63, 64, 65, 66}
𝐸= {44, 45, 46}
We can also be interested in describing new events from combinations of existing
events. Because events are subsets, we can use basic set operations such as unions,
intersections, and complements to form other events of interest. Some of the basic set
operations are summarized here in terms of events:
43
• The union of two events is the event that consists of all outcomes that are contained in
either of the two events. We denote the union as E1 ∪ E2.
• The intersection of two events is the event that consists of all outcomes that are
contained in both of the two events. We denote the intersection as E1 ∩ E2.
• The complement of an event in a sample space is the set of outcomes in the sample
space that are not in the event. We denote the complement of the event E as E′. The
notation EC is also used in other literature to denote the complement.
44
Example 4.6:
Consider an experiment that selects a cell phone camera and records the recycle time
of a flash (the time taken to ready the camera for another flash). The possible values for this
time depend on the resolution of the timer and on the minimum and maximum recycle times.
However, because the time is positive, it is convenient to define the sample space as simply the
positive real line
S = R+ = {x | x > 0}
Let, E1 = {x | 10 ≤ x < 12} and E2 = {x | 11 < x < 15}
Then, E1 ∪ E2 = {x | 10 ≤ x < 15}
And E1 ∩ E2 = {x | 11 < x < 12}
Also, E′1 = {x | x < 10 or 12 ≤ x}
And E′1 ∩ E2 = {x | 12 ≤ x < 15}
Example 4.7:
In the tossing of a die, we might let 𝐴 be the event that an even number occurs and B
the event that a number greater than 3 shows.
Then the subsets A={2,4,6} and 𝐵={4,5,6}are subsets Of the same sample space
S={1,2,3,4,5,6}.
𝐴∩𝐵= {4,6}
𝐴∪𝐵= {2,4,5,6}
𝐴′= {1,3,5}
𝐵′= {1,2,3}
45
Mutually Exclusive Events, or Disjoint:
Two events, denoted as E1 and E2, such that: E1 ∩ E2 = Ø
are said to be mutually exclusive.
, that is, if A and B have no elements in common.
𝐴 = {2,4,6}, and 𝐵 = {1,3,5}
𝐴∩𝐵={ } = ∅
Additional results involving events are summarized in the following. The definition of
the complement of an event implies that:
(E′)′ = E
The distributive law for set operations implies that
(A ∪ B) ∩ C = (A ∩ C) ∪ (B ∩ C) and (A ∩ B) ∪ C = (A ∪ C) ∩ (B ∪ C)
DeMorgan’s laws imply that
(A ∪ B)′ = A′ ∩ B′ and (A ∩ B)′ = A′ ∪ B′
Also, remember that
A ∩ B = B ∩ A and A ∪ B = B ∪ A
A ∩ φ = φ. A ∪ φ = A.
A ∩ A′= φ. A ∪ A′ = S.
S′ = φ. φ′ = S.
46
Venn Diagrams:
Diagrams are often used to portray relationships between sets, and these diagrams are
also used to describe relationships between events. We can use Venn diagrams to represent a
sample space and events in a sample space.
S E
47
Example 4.8:
𝑆= {1,2,3,4,5,6,7}, A= {1,2,4,7}, B= { 1,2,3,6}, C= { 1,3,4,5}
A ∪ C = regions 1, 2, 3, 4, 5, and 7,
B′ ∩ A = regions 4 and 7,
A ∩ B ∩ C = region 1,
(A ∪ B) ∩ C′ = regions 2, 6, and 7.
48
4.3 Counting Techniques
In many of the examples in this chapter, it is easy to determine the number of outcomes
in each event. In more complicated examples, determining the outcomes in the sample space
(or an event) becomes more difficult. In these cases, counts of the numbers of outcomes in the
sample space and various events are used to analyze the random experiments. These methods
are referred to as counting techniques. Some simple rules can be used to simplify the
calculations.
Multiplication Rule (for counting techniques)
Assume an operation can be described as a sequence of k steps, and the number of ways to
complete step 1 is n1, and the number of ways to complete step 2 is n2 for each way to
complete step 1, and• the number of ways to complete step 3 is n3 for each way to
complete step 2, and so forth.
The total number of ways to complete the operation is
n1 × n2 × · · · × nk
Example 4.9:
How many sample points are there in the sample space when a pair of dice is thrown once?
Solution:
The first die can land face-up in any one of n1 = 6 ways. For each of these 6 ways, the second
die can also land face-up in n2 = 6 ways.
Therefore, the pair of dice can land in n1*n2 = (6) *(6) = 36 possible ways.
Example 4.10:
How many sample points are there in the design for a website is to consist of four colors,
three fonts, and three positions for an image?
Solution:
From the multiplication rule, n1 = 4, n2 = 3, n3 = 3
Therefore, different designs possible are
n1*n2* n3 = 4 × 3 × 3 = 36.
Example 4.11:
If a 22-member club needs to elect a chair and a treasurer, how many different ways can these
two to be elected?
Solution:
For the chair position, there are 22 total possibilities. For each of those 22 possibilities, there
are 21 possibilities to elect the treasurer.
Using the multiplication rule, we obtain n1 × n2 = 22 × 21 = 462 different ways.
Example 4.12:
Sam is going to assemble a computer by himself. He has the choice of chips from two
brands, a hard drive from four, memory from three, and an accessory bundle from five local
stores. How many different ways can Sam order the parts?
Solution:
49
Since n1 = 2, n2 = 4, n3 = 3, and n4 = 5, there are
nl × n2 × n3 × n4 = 2× 4 × 3 × 5 = 120 different ways to order the parts.
Example 4.13:
How many even four-digit numbers can be formed from the digits 0, 1, 2, 5, 6, and
9 if each digit can be used only once?
Solution:
Since the number must be even, we have only n1 = 3 choices for the unit’s position.
However, for a four-digit number the thousands position cannot be 0. Hence, we consider the
unit’s position in two parts, 0 or not 0.
If the unit’s position is 0 (i.e., n1 = 1), we have n2 = 5 choices for the thousands position,
n3 = 4 for the hundreds position, and n4 = 3 for the tens position. Therefore, in this case we
have a total of n1*n2*n3*n4 = (1)(5)(4)(3) = 60 even four-digit numbers.
On the other hand, if the unit’s position is not 0 (i.e., n1 = 2), we have n2 = 4 choices for
the thousands position, n3 = 4 for the hundreds position, and n4 = 3 for the tens position. In this
situation, there are a total of n1*n2*n3*n4 = (2)(4)(4)(3) = 96 even four-digit numbers.
Since the above two cases are mutually exclusive, the total number of even four-digit
numbers can be calculated as 60 + 96 = 156.
50
This result follows from the multiplication rule. A permutation can be constructed as
follows. Select the element to be placed in the first position of the sequence from the n
elements, then select the element for the second position from the remaining n − 1 elements,
then select the element for the third position from the remaining n − 2 elements, and so forth.
Permutations such as these are sometimes referred to as linear permutations.
The number of permutations of n objects is n!.
The number of permutations of the four letters a, b, c, and d will be 4! = 24. Now
consider the number of permutations that are possible by taking two letters at a time from four.
These would be ab, ac, ad, ba, bc, bd, ca, cb, cd, da, db, and dc. Using Multiplication Rule
again, we have two positions to fill, with n1 = 4 choices for the first and then n2 = 3 choices for
the second, for a total of n1n2 = (4)(3) = 12 permutations. In general, n distinct objects taken r
at a time can be arranged in n(n − 1)(n − 2) ・ ・ ・ (n − r + 1) ways. We represent this product
by the symbol
Permutations of Subsets
The number of permutations of subsets of r elements selected from a set of n different
𝑛!
elements is nPr = Prn = n × (n − 1) × (n − 2) × · · · × (n − r + 1) =
(𝑛 − 𝑟)!
51
Example 4.14:
In one year, three awards (research, teaching, and service) will be given to a class of
25 graduate students in a statistics department. If each student can receive at most one
award, how many possible selections are there?
Solution:
Since the awards are distinguishable, it is a permutation problem. The total number of
sample points is
𝟐𝟓! 𝟐𝟓!
25P3 = = = (25)(24)(23) = 13, 800.
(𝟐𝟓 − 𝟑)! 𝟐𝟐!
Example 4.15:
A president and a treasurer are to be chosen from a student club consisting of 50 people. How
many different choices of officers are possible if
(a) there are no restrictions.
(b) A will serve only if he is president.
(c) B and C will serve together or not at all.
(d) D and E will not serve together?
Solution:
(a) The total number of choices of officers, without any restrictions, is
𝟓𝟎!
50P2 = = (50)(49) = 2450.
𝟒𝟖!
(b) Since A will serve only if he is president, we have two situations here:
(i) A is selected as the president, which yields 49 possible outcomes for the treasurer’s
position, or
(ii) officers are selected from the remaining 49 people without A, which has the number of
choices
49P2 = (49)(48) = 2352.
Therefore, the total number of choices is 49 + 2352 = 2401.
(c) The number of selections when B and C serve together is 2. The number of selections
when both B and C are not chosen is
48P2 = 2256.
Therefore, the total number of choices in this situation is 2 + 2256 = 2258.
(d) The number of selections when D serves as an officer but not E is (2)(48) = 96, where 2 is
the number of positions D can take and 48 is the number of selections of the other officer
from the remaining people in the club except E.
The number of selections when E serves as an officer but not D is also (2)(48) = 96.
The number of selections when both D and E are not chosen is
52
48P2 = 2256.
Therefore, the total number of choices is (2)(96) + 2256 = 2448.
This problem also has another short solution: Since D and E can only serve together in 2
ways, the answer is 2450 − 2 = 2448.
Example 4.16:
A printed circuit board has eight different locations in which a component can be
placed. If four different components are to be placed on the board, how many different designs
are possible?
Solution:
Each design consists of selecting a location from the eight locations for the first
component, a location from the remaining seven for the second component, a location from the
remaining six for the third component, and a location from the remaining five for the fourth
component. Therefore,
𝟖!
8P4 = = 8 × 7 × 6 × 5 = 1680 different designs are possible.
𝟒!
Example 4.17:
In a college football training session, the defensive coordinator needs to have 10 players
standing in a row. Among these 10 players, there are 1 freshman, 2 sophomores, 4 juniors,
and 3 seniors. How many different ways can they be arranged in a row if only their class
level will be distinguished?
Solution:
We find that total number of players n = 10, n1=1, n2= 2, n3= 4, n4= 3 So,
𝟏𝟎!
the total number of arrangements is = = 12, 600.
𝟏! 𝟐! 𝟒! 𝟑!
Example 4.18:
In how many ways can 7 graduate students be assigned to 1 triple and 2 double hotel rooms
during a conference?
Solution:
53
We find that total number of graduate n = 7, n1=3, n2= 2, n3= 2
𝟕!
The total number of possible partitions would be is = = 210.
𝟑! 𝟐! 𝟐!
Example 4.19:
A young boy asks his mother to get 5 Game-BoyTM cartridges from his collection of 10
arcades and 5 sports games. How many ways are there that his mother can get 3 arcade and 2
sports games?
Solution:
The number of ways of selecting 3 cartridges from 10 is
(10
𝟏𝟎!
3
) = 𝟑! (𝟏𝟎 − 𝟑)! = 120.
54
4.4 Probability of an Event
The probability of event A is the sum of the weights of all sample points in A. Therefore,
0 ≤ P(A) ≤ 1, P(φ) = 0, and P(S) = 1.
Furthermore, if A1, A2, A3, . . . is a sequence of mutually exclusive events, then
P(A1 ∪ A2 ∪ A3 ∪ ・・ ・) = P(A1) + P(A2) + P(A3) + ・ ・ ・ .
Example 4.20:
A coin is tossed twice. What is the probability that at least 1 head occurs?
Solution :
The sample space for this experiment is S = {HH, HT, TH, TT}.
If the coin is balanced, each of these outcomes is equally likely to occur. Therefore, we
assign a probability of ω to each sample point. Then 4ω = 1, or ω = 1/4. If A represents
the event of at least 1 head occurring, then
𝟏 𝟏 𝟏 𝟑
A = {HH, HT, TH} and P(A) = 𝟒 + 𝟒 + =
𝟒 𝟒
Example 4.21:
A die is loaded in such a way that an even number is twice as likely to occur as an odd
number. If E is the event that a number less than 4 occurs on a single toss of the die, find P(E),
then let A be the event that an even number turns up and let B be the event that a number
divisible by 3 occurs. Find P (A ∪ B) and P(A ∩ B).
Solution:
The sample space is S = {1, 2, 3, 4, 5, 6}. We assign a probability of w to each odd
number and a probability of 2w to each even number. Since the sum of the probabilities must
be 1, we have 9w = 1 or w = 1/9. Hence, probabilities of 1/9 and 2/9 are assigned to each odd
and even number, respectively. Therefore,
𝟏 𝟐 𝟏 𝟒
E = {1, 2, 3} and P(E) = 𝟗 + 𝟗 + =
𝟗 𝟗
If an experiment can result in any one of N different equally likely outcomes, and if exactly
n of these outcomes corresponds to event A, then the probability of event A is
𝐧
P(A) = 𝐍
55
Example 4.22:
A statistics class for engineers consists of 25 industrial, 10 mechanical, 10 electrical,
and 8 civil engineering students. If a person is randomly selected by the instructor to answer a
question, find the probability that the student chosen is
(a) an industrial engineering major and
(b) a civil engineering or an electrical engineering major.
Solution:
Denote by I, M, E, and C the students majoring in industrial, mechanical, electrical, and
civil engineering, respectively. The total number of students in the class is 53, all of whom are
equally likely to be selected.
(a) Since 25 of the 53 students are majoring in industrial engineering, the probability of event
I, selecting an industrial engineering major at random, is
𝟐𝟓
P(I) =𝟓𝟑
(b) Since 18 of the 53 students are civil or electrical engineering majors, it follows that
𝟏𝟖
P (C ∪ E) =𝟓𝟑
56
P (A ∪ B ∪ C) = P(A) + P(B) + P(C) – P (A ∩ B) − P (A ∩ C) – P (B ∩ C) + P (A ∩ B ∩
C)
Example 4.23:
John is going to graduate from an industrial engineering department in a university by
the end of the semester. After being interviewed at two companies he likes, he assesses that his
probability of getting an offer from company A is 0.8, and his probability of getting an offer
from company B is 0.6. If he believes that the probability that he will get offers from both
companies is 0.5, what is the probability that he will get at least one offer from these two
companies?
Solution:
Using the additive rule, we have
P (A ∪ B) = P(A) + P(B) – P (A ∩ B) = 0.8 + 0.6 − 0.5 = 0.9.
Example 4.24:
What is the probability of getting a total of 7 or 11 when a pair of fair dice is tossed?
Solution:
Let A be the event that 7 occurs and B the event that 11 comes up. Now, a total of 7
occurs for 6 of the 36 sample points, and a total of 11 occurs for only 2 of the sample points.
Since all sample points are equally likely, we have
P(A) = 1/6 and
P(B) = 1/18.
The events A and B are mutually exclusive, since a total of 7 and 11 cannot both occur on the
same toss. Therefore,
1 1 2
P (A ∪ B) = P(A) + P(B) =6 + =9
18
This result could also have been obtained by counting the total number of points
for the event A ∪ B, namely 8, and writing
𝑛 8 2
P (A ∪ B) = 𝑁 = =9
36
Example 4.25:
57
If the probabilities are, respectively, 0.09, 0.15, 0.21, and 0.23 that a person purchasing
a new automobile will choose the color green, white, red, or blue, what is the probability that
a given buyer will purchase a new automobile that comes in one of those colors?
Solution:
Let G, W, R, and B be the events that a buyer selects, respectively, a green, white, red, or blue
automobile. Since these four events are mutually exclusive, the probability is
P (G ∪W ∪ R ∪ B) = P(G) + P(W) + P(R) + P(B)
= 0.09 + 0.15 + 0.21 + 0.23 = 0.68.
.
If A and A′ are complementary events, then
P(A) + P(A)′ = 1
Example 4.26:
If the probabilities that an automobile mechanic will service 3, 4, 5, 6, 7, or 8 or more
cars on any given workday are, respectively, 0.12, 0.19, 0.28, 0.24, 0.10, and 0.07, what is the
probability that he will service at least 5 cars on his next day at work?
Solution:
Let E be the event that at least 5 cars are serviced. Now, P(E) = 1 − P(E′),
where E′ is the event that fewer than 5 cars are serviced. Since
P(E′) = 0.12 + 0.19 = 0.31,
P(E) = 1 − 0.31 = 0.69.
58
1, or w = 1/5. Relative to the space A, we find that B contains the single element 4. Denoting
this event by the symbol B|A, we write B|A = {4}, and hence
𝟐
P(B|A) =𝟓.
This example illustrates that events may have different probabilities when considered
relative to different sample spaces. We can also write
𝟐
𝐏(𝐀 ∩ 𝐁) 𝟐
P(B|A) = 𝟗𝟓 = =𝟓
𝐏(𝐀)
𝟗
where P (A ∩ B) and P(A) are found from the original sample space S. In other words,
a conditional probability relative to a subspace A of S may be calculated directly from the
probabilities assigned to the elements of the original sample space S.
The conditional probability of B, given A, denoted by P(B|A), is defined by
𝐏(𝐀 ∩ 𝐁)
P(B|A) = 𝐏(𝐀) , provided P(A) > 0.
Example 4.27:
Suppose that our sample space S is the population of adults in a small town who have
completed the requirements for a college degree. We shall categorize them according to gender
and employment status. The data are given in the Table
Employed Unemployed Total
Male 460 40 500
Female 140 260 400
Total 600 300 900
One of these individuals is to be selected at random for a tour throughout the country
to publicize the advantages of establishing new industries in the town. We shall be concerned
with the following events:
M: a man is chosen,
E: the one chosen is employed.
Using the reduced sample space E, we find that
460
460
P(M|E) = 900
600 = 600.
900
Example 4.28:
The probability that a regularly scheduled flight departs on time is P(D) = 0.83; the
probability that it arrives on time is P(A) = 0.82; and the probability that it departs and arrives
on time is P (D ∩A) = 0.78. Find the probability that a plane
(a) arrives on time, given that it departed on time, and
59
(b) departed on time, given that it has arrived on time.
(c) arrives on time, given that it did not depart on time.
Solution:
(a) The probability that a plane arrives on time, given that it departed on time, is
P (D ∩ A) 0.78
P(A|D) = = 0.83 = 0.94
P(D)
(b) The probability that a plane departed on time, given that it has arrived on
time, is
P (D ∩ A) 0.78
P(D|A) = = 0.82 = 0.95
P(A)
(c) The probability that a plane arrives on time, given that it did not depart on time., is
P (A ∩ D‵) P (A)−P (A ∩ D) 0.82 − 0.78
P(A|D‵) = = = = 0.24
P(D‵) P(D‵) 1− 0.83
60
4.7 Independence
In some cases, the conditional probability of P (B | A) might equal P(B). In this special
case, knowledge that the outcome of the experiment is in event A does not affect the probability
that the outcome is in event B.
Independence (two events)
Two events are independent if any one of the following equivalent statements is true:
(1) P (A | B) = P(A)
(2) P (B | A) = P(B)
(3) P (A ∩ B) = P(A)P(B)
Example 4.29:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?
Solution
Let L and R denote the events that the left and right devices operate, respectively. There
is a path only if both operate. The probability that the circuit operates is
P (L and R) = P (L ∩ R) = P(L)P(R) = (0.80) *(0.90) = 0.72
Practical Interpretation: Notice that the probability that the circuit operates degrades to
approximately 0.7 when all devices are required to be functional. The probability that each
device is functional needs to be large for a circuit to operate when many devices are connected
in series.
Example 4.30:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?
Solution
Let T and B denote the events that the top and bottom devices operate, respectively.
There is a path if at least one device operates. The probability that the circuit operates is
P (T or B) = P (T U B) = P (T) + P (B) - P (T)* P (B) = 0.95+0.9-0.95*0.9= 0.995
Another solution
61
P (T or B) = 1 – P [(T or B)′] = 1 − P(T′ and B′) = 1- [P(T′)* P(B′)]= 1- (0.05*0.1) =0.995
Example 4.31:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?
Solution:
The solution can be obtained from a partition of the graph into three columns. Let L
denote the event that there is a path of functional devices only through the three units on the
left. From the independence and based on the previous example,
P(L) = 1 − 0.13
Similarly, let M denote the event that there is a path of functional devices only through
the two units in the middle. Then,
P(M) = 1 − 0.052
The probability that there is a path of functional devices only through the one unit on
the right is simply the probability that the device functions, namely, 0.99. Therefore, with the
independence assumption used again, the solution is
(1 − 0.13) (1 − 0.052) (0.99) = 0.987
Example 4.32:
Suppose that we have a fuse box containing 20 fuses, of which 5 are defective. If 2
fuses are selected at random and removed from the box in succession without replacing the
first, what is the probability that both fuses are defective?
Solution:
We shall let A be the event that the first fuse is defective and B the event that the second
fuse is defective; then we interpret A ∩ B as the event that A occurs and then B occurs after A
has occurred.
The probability of first removing a defective fuse is 5/20 = 1/4; then the probability of
removing a second defective fuse from the remaining 4 is 4/19. Hence,
1 4 1
P (A ∩ B) = P (A) * P (B) = 4* 19 = 19
62
Example 4.32:
One bag contains 4 white balls and 3 black balls, and a second bag contains 3 white balls and
5 black balls. One ball is drawn from the first bag and placed unseen in the second bag. What
is the probability that a ball now drawn from the second bag is black?
Solution:
Let B1, B2, and W1 represent, respectively, the drawing of a black ball from bag 1, a
black ball from bag 2, and a white ball from bag 1. We are interested in the union of the mutually
exclusive events B1 ∩ B2 and W1 ∩ B2. The various possibilities and their probabilities are
illustrated in Figure. Now
P [(B1 ∩ B2) or (W1 ∩ B2)] = P (B1 ∩ B2) + P (W1 ∩ B2)
3 6 4 5 38
= P(B1) P(B2|B1) + P(W1) P(B2|W1) = (7 ∗ 9) + (7 ∗ 9) = 63
++
63
Sheet 4
1. List the elements of each of the following sample spaces:
A. the set of integers between 1 and 50 divisible by 8
B. the set S = {x | x2 + 4x − 5 = 0}
C. the set of outcomes when a coin is tossed until a tail, or three heads appear
D. the set S = {x | 2x − 4 ≥ 0 and x < 1}
E. the set S = {x | 2x − 4 ≥ 0 or x < 1}
2. Which of the following events are equal?
64
6 If an experiment consists of throwing a die and then drawing a letter at random from
the English alphabet, how many points are there in the sample space?
7 Students at a private liberal arts college are classified as being freshmen, sophomores,
juniors, or seniors, and also according to whether they are male or female. Find the total
number of possible classifications for the students of that college.
8 A certain brand of shoes comes in 5 different styles, with each style available in 4
distinct colors. If the store wishes to display pairs of these shoes showing all of its
various styles and colors, how many different pairs will the store have on display?
9 A California study concluded that following 7 simple health rules can extend a man’s
life by 11 years on the average and a woman’s life by 7 years. These 7 rules are as
follows: no smoking, get regular exercise, use alcohol only in moderation, get 7 to 8
hours of sleep, maintain proper weight, eat breakfast, and do not eat between meals. In
how many ways can a person adopt 5 of these rules to follow
65
17 In how many ways can 3 oaks, 4 pines, and 2 maples be arranged along a property line
if one does not distinguish among trees of the same kind?
18 How many ways are there to select 3 candidates from 8 equally qualified recent
graduates for openings in an accounting firm?
19 The probability that an American industry will locate in Shanghai, China, is 0.7, the
probability that it will locate in Beijing, China, is 0.4, and the probability that it will
locate in either Shanghai or Beijing or both is 0.8. What is the probability that the
industry will locate
++
66
CHAPTER 5
We shall use a capital letter, say X, to denote a random variable and its corresponding
small letter, x in this case, for one of its values. In the electronic component testing illustration
above, we notice that the random variable X assumes the value 2 for all elements in the subset
E = {DDN, DND, NDD}
of the sample space S. That is, each possible value of X represents an event that is a
subset of the sample space for the given experiment.
67
Example 5.1:
Two balls are drawn in succession without replacement from an urn containing 4 red
balls and 3 black balls. The possible outcomes and the values y of the random variable Y, where
Y is the number of red balls, are
Sample Space y
RR 2
RB 1
BR 1
BB 0
Example 5.2:
Let X be the random variable defined by the waiting time, in hours, between successive
speeders spotted by a radar unit. The random variable X takes on all values x for which x ≥ 0.
The outcomes of some statistical experiments may be neither finite nor countable. Such
is the case, for example, when one investigates measuring the distances that a certain make of
automobile will travel over a prescribed test course on 5 liters of gasoline. Assuming distance
to be a variable measured to any degree of accuracy, then clearly, we have an infinite number
of possible distances in the sample space that cannot be equated to the number of whole
numbers. Or, if one were to record the length of time for a chemical reaction to take place, once
again the possible time intervals making up our sample space would be infinite in number and
uncountable. We see now that all sample spaces need not be discrete.
If a sample space contains a finite number of possibilities or an unending sequence with as
many elements as there are whole numbers, it is called a discrete sample space.
A random variable is called a discrete random variable if its set of possible outcomes
is countable. The random variables in Example 5.1 are discrete random variables. But a random
variable whose set of possible values is an entire interval of numbers is not discrete. When a
random variable can take on values on a continuous scale, it is called a continuous random
variable. Often the possible values of a continuous random variable are precisely the same
values that are contained in the continuous sample space. Obviously, the random variables
described in Example 5.2 are continuous random variables.
In most practical problems, continuous random variables represent measured data, such
as all possible heights, weights, temperatures, distance, or life periods, whereas discrete random
variables represent count data, such as the number of defectives in a sample of k items or the
number of highway fatalities per year in a given state.
68
If a sample space contains an infinite number of possibilities equal to the number of points
on a line segment, it is called a continuous sample space.
Note that the values of x exhaust all possible cases, and hence the probabilities add to
1. Frequently, it is convenient to represent all the probabilities of a random variable X by a
formula. Such a formula would necessarily be a function of the numerical values x that we shall
denote by f(x), g(x), r(x), and so forth. Therefore, we write f(x) = P(X = x); that is, f(3) = P(X
= 3). The set of ordered pairs (x, f(x)) is called the probability function, probability mass
function, or probability distribution of the discrete random variable X.
The set of ordered pairs (x, f(x)) is a probability function, probability mass function, or
probability distribution of the discrete random variable X if, for each possible outcome x,
1. f(x) ≥ 0,
2. ∑f(x) = 1,
3. P (X = x) = f(x).
Example 5.3:
The shipment of 20 similar laptop computers to a retail outlet contains 3 that are
defective. If a school makes a random purchase of 2 of these computers, find the probability
distribution for the number of defectives.
Solution:
Let X be a random variable whose values x are the possible numbers of defective
computers purchased by the school. Then x can only take the numbers 0, 1, and 2. Now
(3 17
0)( 2 ) 136
f(0) = P(X = 0) = 20 =
(2) 190
69
(3)(17) 51
1 1
f(1) = P(X = 1) = 20 =
(2) 190
(3 17
2)( 0 ) 3
f(2) = P(X = 2) = 20 =
(2) 190
There are many problems where we may wish to compute the probability that the
observed value of a random variable X will be less than or equal to some real number x. Writing
F(x) = P (X ≤ x) for every real number x, we define F(x) to be the cumulative distribution
function of the random variable X.
The cumulative distribution function F(x) of a discrete random variable X with
probability distribution f(x) is
F(x) = P(X ≤ x) = ∑𝐱≤∞ f(x), for −∞ < x < ∞.
70
Although the probability distribution of a continuous random variable cannot be
presented in tabular form, it can be stated as a formula. Such a formula would necessarily be a
function of the numerical values of the continuous random variable X and as such will be
represented by the functional notation f(x). In dealing with continuous variables, f(x) is usually
called the probability density function, or simply the density function, of X. Since X is
defined over a continuous sample space, it is possible for f(x) to have a finite number of
discontinuities. However, most density functions that have practical applications in the analysis
of statistical data are continuous and their graphs may take any of several forms, some of which
are shown in Figure 5-1. Because areas will be used to represent probabilities and probabilities
are positive numerical values, the density function must lie entirely above the x axis.
71
Figure 0-2: P (a < X < b)
Example 5.4:
Suppose that the error in the reaction temperature, in ◦C, for a controlled laboratory
experiment is a continuous random variable X having the probability density function
72
(b) P(0 < X ≤ 1)
Example 5.5:
For the density function of Example 5.4, find F(x), and use it to evaluate P (0 < X ≤ 1).
Solution:
For −1 < x < 2,
Therefore,
REFERENCES
73