0% found this document useful (0 votes)
7 views76 pages

Statistics & Probability

Explanation Statistics & Probability

Uploaded by

snsn2010snsnsnsn
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views76 pages

Statistics & Probability

Explanation Statistics & Probability

Uploaded by

snsn2010snsnsnsn
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

STATISTICAL and PROBABILITY

Lecture notes by

Ahmed Ibrahim
2025

i
TABLE OF CONTENTS

TABLE of Contents .................................................................................................................... i


CHAPTER 1 .............................................................................................................................. 1
1. Representation of data ........................................................................................................ 1
1.1 Introduction ................................................................................................................ 1
1.2 Descriptive and Inferential Statistics ......................................................................... 1
1.3 Types of Data ............................................................................................................. 2
1.4 Representation of discrete data: stem-and-leaf diagrams .......................................... 2
1.5 Representation of continuous data: histograms ......................................................... 5
1.6 Representation of continuous data: cumulative frequency graphs ............................ 9
1.7 Comparing different data representations ................................................................ 11
Conclusion ............................................................................................................................... 12
Type of variable flow chart .................................................................................................. 12
Sheet 1 ...................................................................................................................................... 13
CHAPTER 2 ............................................................................................................................ 15
.2 Measures of central tendency ........................................................................................... 15
2.1 Three types of average ............................................................................................. 15
2.2 The mode and the modal class ................................................................................. 15
2.3 The mean .................................................................................................................. 16
2.3.1 Combined sets of data .............................................................................................. 18
2.3.2 Means from grouped frequency tables ..................................................................... 19
2.3.3 Coded data ............................................................................................................... 20
2.4 The median............................................................................................................... 23
Sheet 2 .................................................................................................................................. 25
CHAPTER 3 ............................................................................................................................ 27
3. Measures of variation ....................................................................................................... 27
3.1 The range ................................................................................................................. 27
3.2 The interquartile range and percentiles .................................................................... 28
Ungrouped data .................................................................................................................... 29
Grouped data ........................................................................................................................ 31
Box-and-whisker diagrams .................................................................................................. 33
3.3 Variance and standard deviation .............................................................................. 34
Sheet 3 .................................................................................................................................. 36

i
CHAPTER 4 ............................................................................................................................ 38
4. Probability ........................................................................................................................ 38
4.1 Sample Spaces and Events ............................................................................................. 38
4.1.2 Sample space ........................................................................................................... 39
CHAPTER 5 ........................................................................................................................... 67
References ................................................................................................................................ 73

ii
CHAPTER 1

1. REPRESENTATION OF DATA

In this chapter you will learn how to:


• display numerical data in stem-and-leaf diagrams, histograms and cumulative
frequency graphs
• interpret statistical data presented in various forms
• select an appropriate method for displaying data.
1.1 Introduction
You are probably asking yourself the question, "When and where will I use statistics?"
If you read any newspaper, watch television, or use the Internet, you will see statistical
information. There are statistics about crime, sports, education, politics, and real estate.
Typically, when you read a newspaper article or watch a television news program, you are
given sample information. With this information, you may decide about the correctness of a
statement, claim, or "fact." Statistical methods can help you make the "best educated guess." It
has become accepted in today’s world that in order to learn about something, you must first
plan and collect data. Statistics is the art of learning from data. It is concerned with the
collection of data, its subsequent description, and its analysis, which often leads to the
drawing of conclusions.

1.2 Descriptive and Inferential Statistics


Descriptive and inferential statistics are two fields of statistics. Descriptive statistics are
used to describe data, and inferential statistics is used to make predictions. Descriptive and
inferential statistics have different tools that can be used to conclude the data. In descriptive
and inferential statistics, the former uses tools such as central tendency, and dispersion, while
the latter uses hypothesis testing, regression analysis, and confidence intervals. The purpose
of descriptive and inferential statistics is to analyze different types of data using different tools.
Descriptive statistics help to describe and organize known data using charts, bar graphs, etc.,
while inferential statistics aim at making inferences and generalizations about the population
data. Both descriptive and inferential statistics are equally important to analyze data.
Descriptive statistics are used to order data and describe the sample using the mean, standard
deviation, charts, etc. Inferential statistics use this sample data to predict the trend of the
population data.

1
1.3 Types of Data
There are two types of data: qualitative (or categorical) data are described by words
and are non-numerical, such as blood types or colors. Quantitative data take numerical values
and are either discrete or continuous. As a general rule, discrete data are counted and cannot
be made more precise, whereas continuous data are measurements that are given to a chosen
degree of accuracy. Discrete data can take only certain values, as shown in the diagram. The
number of letters in the words of a book is an example of discrete quantitative data. Each word
1
has 1 or 2 or 3 or 4 or… letters. There are no words with 3 or 4.75 letters. Discrete
2

quantitative data can take non-integer values. For example, United States coins have dollar
values of 0.01, 0.05, 0.10, 0.25, 0.50 and 1.00. In Canada, the United Kingdom and other
1 1
countries, shoe sizes such as 62 , 7 and 72 are used. Continuous data can take any value

(possibly within a limited range), as shown in the diagram. The times taken by the athletes to
complete a 100-metre race is an example of continuous quantitative data. We can measure these
to the nearest second, tenth of a second or even more accurately if we have the necessary
equipment. The range of times is limited to positive real numbers. It can be concluded that
Discrete data can take only certain values. Continuous data can take any value, possibly
within a limited range.

Discrete data Continuous data


1.4 Representation of discrete data: stem-and-leaf diagrams
A stem-and-leaf diagram is a type of table best suited to representing small amounts
of discrete data. The last digit of each data value appears as a leaf attached to all the other
digits, which appear in a stem. The digits in the stem are ordered vertically, and the digits on
the leaves are ordered horizontally, with the smallest digit placed nearest to the stem. Each
row in the table forms a class of values. The rows should have intervals of equal width to allow
for easy visual comparison of sets of data. A key with the appropriate unit must be included to
explain what the values in the diagram represent. Stem-and-leaf diagrams are particularly
useful because raw data can still be seen, and two sets of related data can be shown back-to-
back for the purpose of making comparisons.
Example 1.1
Consider the raw percentage scores of 15 students in a Physics exam, given in the
following list: 58, 55, 58, 61, 72, 79, 97, 67, 61, 77, 92, 64, 69, 62 and 53.

2
To present the data in a stem-and-leaf diagram, we first group the scores into suitable
equal-width classes. Class widths of 10 are suitable here, as shown below.

5 8583
6 171492
7 297
8
9 72

Next, we arrange the scores in each row in ascending order from left to right and add a
key to produce the stem-and-leaf diagram shown below.
Key:
5 3588 5 3
6 112479 Represents a score of
7 279 53%
8
9 27

In a back-to-back stem-and-leaf diagram, the leaves to the stem's right ascend left to
right, and the leaves on the left of the stem ascend right to left (as shown in Worked example
1.2).

Example 1.2
The number of days on which rain fell in a certain town in each month of 2016 and
2017 are given.

Year 2016
Jan: 17 Feb: 20 Mar: 13 Apr: 12 May: 10 Jun: 8
Jul: 0 Aug: 1 Sep: 5 Oct: 11 Nov: 16 Dec: 9
Year 2017
Jan: 9 Feb: 13 Mar: 11 Apr: 8 May: 6 Jun: 3
Jul: 1 Aug: 2 Sep: 2 Oct: 4 Nov: 8 Dec: 7

Display the data in a back-to-back stem-and-leaf diagram and briefly compare the
rainfall in 2016 with the rainfall in 2017.

3
Key:
5 0 6
2016 2017
Represents days in a
98510 0 1223467889
month of 2016 and 6
763210 1 13
0 2 days in a month of 2017

It rained on more days in 2016 (122 days) than it did in 2017 (74 days).

Note That: If rows of leaves are particularly long, repeated values may be used in the stem.
However, if there were, say, 30 leaves in one of the rows, we might consider grouping the data
into narrower classes of 0–4, 5–9,10–14,15–19 and 20–24. This would require 0, 0, 1,1 and 2
in the stem.

4
1.5 Representation of continuous data: histograms
Continuous data are given to a certain degree of accuracy, such as 3 significant figures,
2 decimal places, to the nearest 10 and so on. We usually refer to this as rounding. When values
are rounded, gaps appear between classes of values, and this can lead to a misunderstanding of
continuous data because those gaps do not exist. For example, consider heights to the nearest
centimeter, given as 146: 150, 151: 155 and 156: 160. Gaps of 1cm appear between classes
because the values are rounded. Using h for height, the actual classes should be 145.5≤h<150.5,
150.5≤h<155.5 and 155.5≤h < 160.5cm.
The classes are shown in the diagram below, with the lower and upper boundary
values and the class mid-values (also called midpoints) indicated.

Lower class boundaries are 145.5, 150.5 and 155.5cm.


Upper class boundaries are 150.5, 155.5 and 160.5cm.
Class widths are 150.5 –145.5=5, 155.5 –150.5=5 and 160.5 –155.5=5.
145.5+150.5 150.5+155.5 155.5+160.5
Class mid-values are = 148, = 153, = 158
2 2 2

A histogram is best suited to illustrating continuous data, but it can also be used to
illustrate discrete data. We might have to group the data ourselves or it may be given to us in a
grouped frequency table. For example, the presented in the tables below show the ages and
the percentage scores of 100 students who took an examination.

Age (A years) 16 ≤A < 18 18 ≤A < 20 20 ≤A < 22


No. students (f) 34 46 20

Score (%) 10–29 30–59 60–79 80–99


No. students (f) 6 21 60 13

A histogram is a chart that uses bars to show the distribution of values for a numeric
variable. Each bar represents a range of values, and its height shows how many data points fall
within that range.

5
The first table shows three classes of continuous data; there are no gaps between the
classes and the classes have equal-width intervals of 2 years. This means that we can represent
the data in a frequency diagram by drawing three equal-width columns with column heights
equal to the class frequencies, as shown below.

The following table shows the areas of the columns and the frequency of each of the
three classes presented in the diagram on the previous page.

First Second Third


Area 2×34=68 2×46=92 2×20=40
Frequency 34 46 20

From this table, we can see that the ratio of the column areas, 68: 92: 40, is the same as
the ratio of the frequencies, 34: 46: 20. In a histogram, the area of a column represents the
frequency of the corresponding class so that the area must be proportional to the frequency. We
may see this written as ‘area _ frequency’. This also means that in every histogram, just as in
the example above, the ratio of column areas is the same as the ratio of the frequencies, even
if the classes do not have equal widths. Also, there can be no gaps between the columns in a
histogram because one class's upper boundary equals the neighboring class's lower boundary.
A gap can appear only when a class has zero frequency. The axis showing the measurements
is labeled as a continuous number line, and the width of each column is equal to the width of
the class it represents. When we construct a histogram, since the classes may not have equal
widths, the height of each column is no longer determined by the frequency alone but must be

6
calculated so that area _ frequency. The vertical axis of the histogram is labeled frequency
density, which measures frequency per standard interval.
The simplest and most commonly used standard interval is 1 unit of measurement. For
example, a column representing 85 objects with masses from 50 to 60 kg has a frequency
density of 85 objects / (60-50) kg = 8.5 objects per kilogram, and so on.

Note That: For a standard interval of 1 unit of measurement, Frequency density = class
frequency/ class width, which can be re-arranged to give
Class frequency= class width * frequency density

Example 1.3
The masses (m), kg, of 100 children are grouped into two classes, as shown in the table.

Mass (m) 40 ≤m < 50 50 ≤m < 70


No. children (f) 40 60

a. Illustrate the data in a histogram.


b. Estimate the number of children with masses between 45 and 63kg.
Answer:
Frequency density is calculated for the unequal-width intervals in the table. The masses
are represented in the histogram, where frequency density measures the number of children per
1kg or simply children per kg.

Mass (m) 40 ≤m < 50 50 ≤m < 70


No. children (f) 40 60
Class width (kg) 50-40 =10 70-50=20
Frequency density 40/10 =4 60/20 =3

7
b. There are children with masses from 45 to 63kg in both classes, so we must split this
interval into two parts: 45–50 and 50–63.
Frequency= width* frequency density
For first part (45–50) = (50-45) * 4= 20 children
For second part (50–63) = (63-50) * 3= 39 children
The estimation number of children = 20 + 39 = 59 children
Example 1.4
Consider the times taken, to the nearest minute, for 36 athletes to complete a race, as
given in the table below.
Time taken (min) 13 14-15 16-18
No. athletes (f) 4 14 18

Use the histogram of race times to estimate:


a. the number of athletes who took less than 13.0 minutes
b. the number of athletes who took between 14.5 and 17.5 minutes
c. the time taken to run the race by the slowest of three athletes.
Answer:
Gaps of 1 minute appear between classes because the times are rounded. Frequency
densities are calculated in the following table.
Time taken (min) 12.5 ≤t < 13.5 13.5 ≤t < 15.5 15.5≤t< 18.5
No. athletes (f) 4 14 18
Class width (min) 1 2 3
Frequency density 4 7 6

8
a. the number of athletes who took less than 13.0 minutes
Frequency= width* frequency density = (13-12.5) *4= 2 athletes
b. the number of athletes who took between 14.5 and 17.5 minutes
Frequency= width* frequency density = (15.5-14.5) *7 + (17.5-15.5)6= 19 athletes
c. the time taken to run the race by the slowest of three athletes.
"When we refer to the slowest athlete, we are talking about the one who takes the longest
time."
Frequency= width* frequency density
3 = width * 6 3 = (18.5 - t) * 6
so, the time taken= 18.5-0.5=18min

1.6 Representation of continuous data: cumulative frequency graphs


A cumulative frequency graph can be used to represent continuous data. Cumulative
frequency is the total frequency of all values less than a given value. If we are given grouped
data, we can construct the cumulative frequency diagram by plotting cumulative frequencies
against upper class boundaries for all intervals. We can join the points consecutively with
straight-line segments to give a cumulative frequency polygon or with a smooth curve to give
a cumulative frequency curve. For example, a set of data that includes 100 values below 7.5
and 200 values below 9.5 will have two of its points plotted at (7.5,100) and at (9.5, 200). We
plot points at upper boundaries because we know the total frequencies up to these points are
precise. From a cumulative frequency graph, we can estimate the number or proportion of
values that lie above or below a given value, or between two values.

9
Example 1.5
The following table shows the lengths of 80 leaves from a particular tree, given to the
nearest centimeter.
Lengths (cm) 1-2 3-4 5-7 8-9 10-11
No. leaves (f) 8 20 38 10 4

Draw a cumulative frequency curve and a cumulative frequency polygon. Use each of these
to estimate:
a. the number of leaves that are less than 3.7cm long
b. the lower boundary of the lengths of the longest 22 leaves.
Answer:
Lengths (cm) Addition of frequencies No. leaves
L<0.5 0 0
L<2.5 0+8 8
L<4.5 0+8+20 28
L<7.5 0+8+20+38 66
L<9.5 0+8+20+38+10 76
L<11.5 0+8+20+38+10+4 80

a:
➢ The polygon gives an estimate of 20 leaves.
➢ The curve gives an estimate of 18 leaves.

10
b
➢ The polygon gives an estimate of 6.9cm.
➢ The curve gives an estimate of 6.7cm.
1.7 Comparing different data representations
Pictograms, bar charts and pie charts are useful ways of displaying qualitative data and
ungrouped quantitative data, and people generally find them easy to understand. Nevertheless,
it may be of benefit to group a set of raw data so that we can see how the values are distributed.
Knowing the proportion of small, medium and large values.
Pictograms are types of charts and graphs that use icons and images to represent data.
A bar chart or bar graph is a chart or graph that presents categorical data with
rectangular bars with heights or lengths proportional to the values that they represent.
A pie chart is a type of graph representing data in a circular form, with each slice of
the circle representing a fraction or proportionate part of the whole.
Dot plot, Bar charts, and vertical line:
➢ Each dot represents a specific number of observations.
➢ The dots are stacked in a column over a category. The height of the column
represents the absolute frequency of observation.
➢ Used most often to plot frequency counts within a small number of categories.

Line graph is a type of graph that shows information that is connected in some ways.

11
CONCLUSION

Type of variable flow chart


The following chart is a guide to some of the most commonly used methods of data
representation.

Variable

quantitative Qualitative
(Numeric) (Categorical)

Continous Discrete Ordinal Nominal

➢ Non-numerical data are called qualitative or categorical data.


➢ Numerical data are called quantitative data and are either discrete or continuous.
➢ Discrete data can take only certain values.
➢ Continuous data can take any value, possibly within a limited range.
➢ Data in a stem-and-leaf diagram are ordered in rows with intervals of equal width.
➢ In a histogram, column area _ frequency, and the vertical axis is labelled frequency
density.
➢ Frequency density = class frequency / class width
➢ Class frequency = class width × frequency density.
➢ In a cumulative frequency graph, points are plotted at class upper boundaries.

12
SHEET 1

1. Twenty people leaving a cinema are each asked, “How many times have you attended
the cinema in the past year?” Their responses are:
6, 2, 13, 1, 4, 8, 11, 3, 4, 16, 7, 20, 13, 5, 15, 3, 12, 9, 26 and 10.
Construct a stem-and-leaf diagram for this data and include a key.

2. A shopkeeper takes 12 bags of coins to the bank. The bags contain the following
numbers of coins:
150, 163, 158, 165, 172, 152, 160, 170, 156, 162, 159 and 175.
a. Represent this information in a stem-and-leaf diagram.
b. Each bag contains coins of the same value, and the shopkeeper has at least one bag
containing coins with dollar values of 0.10, 0.25, 0.50 and 1.00 only.
What is the greatest possible value of all the coins in the 12 bags?

3. This stem-and-leaf diagram shows the number of employees at 20 companies.


a. What is the most common number of employees?
1 0888899
b. How many of the companies have fewer than 25 employees?
2 05667789
c. What percentage of companies have more than 30 employees? 3 01129
d. Determine which of the three rows in the stem-and-leaf diagram contains the smallest
number of: i. Companies ii. Employees.

4. In a particular city there are 51 buildings of historical interest. The following table
presents the ages of these buildings, given to the nearest 50 years.
Age (years) 50-150 200-300 350-450 500-600
No. buildings (f) 15 18 12 6

a. Write down the lower and upper boundary values of the class containing the greatest
number of buildings.
b. State the widths of the four class intervals.
c. Illustrate the data in a histogram.
d. Estimate the number of buildings that are between 250 and 400 years old.

5. The masses, m grams, of 690 medical samples, are given on the following table.

13
Mass (m grams) 4≤ m<12 12≤ m<24 24≤ m<28
No. medical samples (f) 200 400 p

a. Find the value of p that appears in the table.


b. Draw a histogram to represent the data.
c. Calculate an estimate of the number of samples with masses between 8 and 18 grams.

6. A university investigated how much space on its computers’ hard drives is used for
data storage. The results are shown below. It is given that 40 hard drives use less than 20GB
for data storage.

a. Find the total number of hard drives represented.


b Calculate an estimate of the number of hard drives that use less than 50GB.
c Estimate the value of k, if 25% of the hard drives use k GB or more.

7. The following table shows the widths of the 70 books in one section of a library,
given to the nearest centimeter.
Width (cm) 10-14 15-19 20-29 30-39 40-44
No. books (f) 3 13 25 24 5
a. Given that the upper boundary of the first class is 14.5cm, write down the upper boundary
of the second class.
b. Draw up a cumulative frequency table for the data and construct a cumulative frequency
graph.
c. Use your graph to estimate:
i. the number of books that have widths of less than 27cm
ii the widths of the widest 20 books.

14
CHAPTER 2

2. MEASURES OF CENTRAL TENDENCY

In this chapter you will learn how to:


• Find and use different measures of central tendency
• calculate and use the mean of a set of data (including grouped data) either from the
data itself or from a given total ∑x or a coded total ∑(x−b) and use such totals in
solving problems that may involve up to two datasets.

2.1 Three types of average


There are three measures of central tendency that are commonly used to describe the
average value of a set of data. These are the mode, the mean and the median.
➢ The mode is the most commonly occurring value.
➢ The mean is calculated by dividing the sum of the values by the number of values.
➢ The median is the value in the middle of an ordered set of data.

We use an average to summarize the values in a set of data. As a representative value,


it should be fairly central to, and typical of, the values it represents. If we investigate the annual
incomes of all the people in a region, then a single value, i.e., an average income, would be a
convenient number to represent our findings. However, choosing which average to use needs
to be considered, as one measure may be more appropriate to use than the others. Deciding
which measure to use depends on many factors. Although the mean is the most familiar
average, a shoemaker would prefer to know which shoe size is the most popular, i.e., the mode.
A farmer may find the median number of eggs laid by their chickens to be the most useful
because they could use it to identify which chickens are profitable and which are not. As for
the average income in our chosen region, we must also consider calculating an average for the
workers and managers together or separately; and, if separately, then we need to decide who
fits into which category.

2.2 The mode and the modal class


As you will recall, a set of data may have more than one mode or no mode at all. The
following table shows the scores on 25 rolls of a die, where 2 is the mode because it has the
highest frequency.

15
Score on die 1 2 3 4 5 6
Frequency (f) 5 6 5 3 2 4

In a set of grouped data in which raw values cannot be seen, we can find the modal
class, which is the class with the highest frequency density.
EXAMPLE 2.1
Find the modal class of the 270 pencil lengths, given to the nearest centimeter in the
following table.
Length (cm) 4-7 8-10 11-12
No. pencils (f) 100 90 80
Answer
Length (cm) 3.5-7.5 7.5-10.5 10.5-12.5
No. pencils (f) 100 90 80
Width 4 3 2
Frequency density 25 30 40

The modal class is 11–12cm (or, more accurately, 10.5<x < 12.5cm).
EXAMPLE 2.2
Two classes of data have interval widths in the ratio 3:2. Given that there is no modal
class, and that the frequency of the first class is 48, find the frequency of the second class.
Answer

Note That: No modal class means that the frequency densities of the two classes are equal

Let the frequency of the second class be x.


48/x=3/2 X = 48*2/3 =32 The second class has a frequency of 32.
A small company sells glass, which it cut to size to fit into window frames. How could the
company benefit from knowing the modal size of glass its customers purchase?
2.3 The mean
The mean is referred to more precisely as the arithmetic mean and it is the most
commonly known average. The sum of a set of data values can be found from the mean.
Suppose, for example, that 12 values have a mean of 7.5:
𝐬𝐮𝐦 𝐨𝐟 𝐯𝐚𝐥𝐮𝐞𝐬
Mean = , so
𝐧𝐮𝐦𝐛𝐞𝐫 𝐨𝐟 𝐯𝐚𝐥𝐮𝐞𝐬
𝐬𝐮𝐦 𝐨𝐟 𝐯𝐚𝐥𝐮𝐞𝐬
7.5 = , and
𝟏𝟐

sum of values=7.5 × 12 = 90.

16
You will soon be performing calculations involving the mean, so here we introduce
notation that is used in place of the word definition used above. We use the upper-case Greek

letter ‘sigma’, written ∑, to represent ‘sum’ and 𝑥̅ to represent the mean, where x represents
our data values. The notation used for ungrouped and for grouped data are shown on separate
rows in the following table.
` Sum Data Frequency Number of Sum of Mean
values of Data data
data values values values
Ungrouped ∑ x - n ∑x ∑x
𝐱̅ =
n

Grouped ∑ x f ∑f ∑xf ∑xf


̅=
∑𝐱
∑𝐟

EXAMPLE 2.3
Five persons, whose mean mass is 70.2kg, wish to go to the top of a building in a lift
with some cement. Find the greatest mass of cement they can take if the lift has a maximum
weight allowance of 500 kg.
Answer

𝑥̅ = 70.2 kg, n = 5, ∑x = 𝑥̅ * n = 70.2 * 5 = 351 kg


Weight of cement + weight of five persons = 500 kg
Weight of cement = 500 – 351 = 149 kg

17
EXAMPLE 2.4
Find the mean of the 40 values of x given in the following table.
x 31 32 33 34 35
F 5 7 9 8 11

Answer
x 31 32 33 34 35
F 5 7 9 8 11 ∑f=40
X* F 155 224 297 272 385 ∑xf=1333

∑𝐱𝐟 𝟏𝟑𝟑𝟑
̅=
𝒙 = = 33.325
∑𝐟 𝟒𝟎

2.3.1 Combined sets of data


There are many different ways to combine sets of data. However, here we do this by
simply considering all of their values together. To find the mean of two combined sets, we
divide the sum of all their values by the total number of values in the two sets. For example,
by combining the dataset 1, 2, 3, 4 with the dataset 4, 5, 6, we obtain a new set of data that has
seven values in it: 1, 2, 3, 4, 4, 5, 6. Note that the value 4 appears twice. Individually, the sets
have means of (1+2+3+4)/4 = 2.5, and (4+5+6)/3 = 5 where the combined sets have a mean of
(1 +2+ 3+ 4+ 4+ 5+ 6) / 7 = 3.57

EXAMPLE 2.5
A large bag of sweets claims to contain 72 sweets, having a total mass of 852.4g. A
small bag of sweets claims to contain 24 sweets, having a total mass of 282.8g. What is the
mean mass of all the sweets together?
Answer
Total number of sweets=72+24=96.
Total mass of sweets= 852.4 + 282.8 = 1135.2g
Mean mass= 1135.2 / 96 = 11.825g

18
2.3.2 Means from grouped frequency tables
When data are presented in a grouped frequency table or illustrated in a histogram or
cumulative frequency graph, we lose information about the raw values. For this reason, we
cannot determine the mean exactly, but we can calculate an estimate of the mean. We do this
∑xf
̅=
by using mid-values to represent the values in each class. We use the formula: ∑𝐱 to
∑𝐟

calculate an estimate of the mean, where x now represents the class mid-values.

EXAMPLE 2.6
Coconuts are packed into 75 crates, with 40 of a similar size in each crate. 46 crates
contain coconuts with a total mass of 20 up to but not including 25 kg. 22 crates contain
coconuts with a total mass of 25 up to, but not including 40 kg. 7 crates contain coconuts with
a total mass of 40 up to but not including 54 kilograms.
a Calculate an estimate of the mean mass of a crate of coconuts.
b Use your answer to part a to estimate the mean mass of a coconut.

Answer
Mass (kg) 20- 25- 40-54
Frequency (f) 46 22 7 ∑=75
Mid value (x) 22.5 32.5 47
X* F 1035 715 329 ∑=2079

∑xf
̅=
The mean mass of a crate = ∑𝐱 = 2079 / 75 = 27.72 kg
∑𝐟

the mean mass of a coconut= 27.72 / 40 = 0.693 kg = 693 gm

19
EXAMPLE 2.7
Calculate an estimate of the mean age of a group of 50 students, where there are sixteen
18-year-olds, twenty 19-year-olds and fourteen who are either 20 or 21 years old.
Answer
Age (year) 18- 19- 20-22
Frequency (f) 16 20 14 ∑=50
Mid value (x) 18.5 19.5 21
X* F 296 390 294 ∑=980

the mean age =∑xf / ∑f = 980/ 50 = 19.6 years

2.3.3 Coded data


To code a set of data, we can transform all of its values by addition of a positive or
negative constant. The result of doing this produces a set of coded data. One reason for coding
is to make the numbers easier to handle when performing manual calculations. Also, it is
sometimes easier to work with coded data than with the original data (by arranging the mean
to be a convenient number, such as zero, for example). To find the mean of 101, 103, 104, 109
and 113, for example, we can use the values 1, 3, 4, 9 and 13.
Our x values are 101, 103, 104, 109 and 113, so 1, 3, 4, 9 and 13 are corresponding
values of (x–100).
[∑ (𝑥− 100)] (1+3+4+9+13)
Mean of the coded values is = =6
𝟓 𝟓

We subtracted 100 from each x value, so we simply add 100 to the mean of the coded
values to find the mean of x.
Mean(x)=mean(x–100) +100 = 106 or
∑(𝒙− 𝟏𝟎𝟎)
𝑥̅ = + 100 =106
𝟓

20
It can be concluded that:
∑(𝒙− 𝐛)
For ungrouped data, 𝑥̅ = +b
𝐧
∑(𝒙− 𝐛)𝐟
For grouped data, 𝑥̅ = +b
∑𝐟

̅=mean(x–b) +b.
These formulae can be summarized by writing 𝒙

EXAMPLE 2.8
The exact age of an individual boy is denoted by b, and the exact age of an individual
girl is denoted by g. Exactly 5 years ago, the sum of the ages of 10 boys was 127.0 years, so∑(
(b−5) =127.0. In exactly 5 years’ time, the sum of the ages of 15 girls will be 351.0 years, so

∑( (g+5) =351.0. Find the mean age today of


a. the 10 boys
b. the 15 girls
c. the 10 boys and 15 girls combined.
Answer:
∑(𝑥− b) ∑(𝑏− 5)
a. 𝑥̅ = +b= +5= 127/10 + 5 =17.7 years
n 10

∑(𝑥+ b) ∑(𝑔+ 5)
b. 𝑥̅ = –b= – 5 = 351 / 15 – 5 = 18.4 years
n 15

∑(b − 5) = ∑b − (n ∗ 5)
so ∑b = 127 + 10 ∗ 5 = 177
∑(g + 5) = ∑g + (n ∗ 5)
so ∑g = 351 − 15 ∗ 5 = 276

∑b+ ∑g
c. the mean age of combined =
n
177+ 276
= = 18.2 years
10+15

21
EXAMPLE 2.8
Forty values of x are coded in the following table.
x–3 0– 18– 24–32
Frequency 9 13 18

Calculate an estimate of the mean value of x.


Answer

x–3 0– 18– 24–32


Frequency 9 13 18 =40
Mid value 9 21 28
9*9 13*21 18*28 = 858

∑(𝒙− 𝐛)∗𝐟 𝟖𝟓𝟖


̅=
𝒙 +b= + 3= 24.45
∑𝐟 𝟒𝟎

EXAMPLE 2.9
For the 20 values of x summarized by ∑ (2x−3) = 104, find 𝑥
̅.
𝟏𝟎𝟒
2𝑥̅ = + 3 =8.2
𝟐𝟎

𝑥̅ = 4.1

22
2.4 The median
You will recall that the median splits data into two parts with an equal number of values
in each part: a bottom half and a top half. In a set of n-ordered values, the median is the value
halfway between the 1st and the nth. Consider a DIY store that opens for 12 hours on Monday
and 15 hours on Saturday. The following back-to-back stem-and-leaf diagram shows the
numbers of customers served during each hour on Monday and Saturday last week.
Monday (12) Saturday (15)
863100 2 2346 Key:
43110 3 556899 0 2 2
1 4 01379 Represents customers in a
Monday of 20 and in a
Saturday of 22

To find the median number of customers served on each of these days, we need to find
their positions in the ordered rows of the back-to-back stem-and-leaf diagram.
For Saturday, there are n = 15 values arranged in ascending order from top to bottom
𝑛+1 15+1
and from left to right. The median is at the { } th = { } = 8th value. In the first row, we
2 2

have the 1st to 4th values, and in the second row we have the 5th to 10th values, so the 8th
value is 38. The median number of customers on Saturday was 38. For Monday, there are n =
12 values arranged in ascending order from top to bottom and from right to left. The median is
𝑛+1 12+1
at the { } th = { } = 6.5th value, so we locate the median mid-way between the 6th and
2 2

7th values. In the first row, we have the 1st to 6th values and the 6th is 28. The first value in
the second row is the 7th value, which is 30. The median number of customers on Monday
28+30
was { } = 29. When data appear in an ordered frequency table of individual values, we can
2

use cumulative frequencies to investigate the positions of the values, knowing that the median
𝑛+1
is at the { } value.
2

EXAMPLE 2.10
The following table shows 65 ungrouped readings of x. Cumulative frequencies, and
the positions of the readings are also shown. Find the median value of x.
X F Cf
40 11 11
41 23 34

23
42 19 53
43 8 61
44 4 65

𝐧+𝟏 𝟔𝟓+𝟏
The total frequency is 65, and = = 33, so the median is at the 33rd value. From
𝟐 𝟐

the table, we see that the 12th to 34th values are all equal to 41.

24
Sheet 2
1. Find the mode(s) of the following sets of numbers.
a. 12,15,11, 7, 4,10, 32,14, 6,13,19, 3
b. 19, 21, 23,16, 35, 8, 21,16,13,17,12,19,14, 9

2. Which of the eleven words in this sentence is the mode?

3. Identify the mode of x and of y in the following tables.


x 4 5 6 7 8
f 1 5 5 6 4

y -4 -3 -2 -1 0
f 27 28 29 27 25

4. Find the modal class for x and for y in the following tables.
x 0- 4- 14- 20
f 5 9 8

y 3- 6 7- 11 12- 20
f 66 80 134

5. Calculate the mean of the following sets of numbers:


a. 28,16, 83, 72,105, 55, 6 and 35
b. 7.3, 8.6, 11.7, 9.1, 1.7 and 4.2

6. The mean of 15, 31, 47, 83, 97, 119 and p2 is 63. Find the possible values of p.

7. The mean of 6, 29, 3, 14, q, (q+ 8), q2 and (10 – q) is 20. Find the possible values of q.

8. Find the mean of x of given in the following table.


X 18 18.5 19 19.5 20
f 8 10 17 24 1

9. An examination was taken by 50 students. The 22 boys scored a mean of 71% and
the girls scored a mean of 76%. Find the mean score of all the students.

25
10. The following table summarizes the number of tomatoes produced by the plants in
the plots on a farm.
No. tomatoes 20-29 30-49 50-79 80-100
No. plots (f) 329 413 704 258
a. Calculate an estimate of the mean number of tomatoes produced by these plots.

̅ = 7.4. Find:
11. 10. For 10 values denoted x, it is given that 𝒙

a. ∑x b. ∑ (x+2) c. ∑ (x−1)

̅.
12. Twenty-five values of z are such that ∑ (z−7) = 275. Find 𝒁

13. Given 𝐪
̅ = 22 and ∑ (q−4) = 3672, find the number of values of q.

14. The lengths of 2500 bolts, x mm, are summarized by ∑ (x−40) = 875. Find the mean
length of the bolts.

26
CHAPTER 3

3. MEASURES OF VARIATION

In this chapter you will learn how to:


• find and use different measures of variation
• use a cumulative frequency graph to estimate medians, quartiles and percentiles
• calculate and use the standard deviation of a set of data (including grouped data) either
from the data itself or from given totals ∑x and ∑x2, or coded totals ∑(x−b) and ∑(x−b)2
and use such totals in solving problems that may involve up to two datasets.

A measure of central tendency alone does not describe or summarize a set of data fully.
Although it may tell us the location of the more central values or the most common values, it
tells us nothing about how widely spread out the values are. Two sets of data can have the same
mean, median or mode, yet they can be completely different. A better description of a set of
data is given by a measure of central tendency and a measure of variation. Variation is also
known as spread or dispersion. Consider the runs scored by two batters in their past eight
cricket matches, which are given in the following table.

Batter A 25 30 31 26 31 28 29 24 Total: 224


Batter B 2 70 1 0 43 1 104 3 Total: 224

The mean number of runs scored by A and by B is the same; namely, 224÷8=28.
However, the patterns of the number of runs are clearly very different. The numbers for batter
A are quite consistent, whereas the numbers for batter B are quite varied. This consistency (or
lack of it) can be indicated by a measure of variation, which shows how spread out a set of data
values are. Three commonly used measures of variation are the range, interquartile range
and standard deviation.

3.1 The range


As you will recall, the range is the numerical difference between the largest and smallest
values in a set of data. One advantage of using the range is that it is easy to calculate. However,
it does not take the more central values into account but uses only the most extreme values. It
is often more informative to state the minimum and maximum values rather than the difference
between them.

27
For example, in a test for which the lowest mark is 6 and the highest mark is 19, the
range is 19 – 6 = 13. For grouped data, we can find a minimum and maximum possible range,
using the lower and upper boundary values of the data.
EXAMPLE 3.1
To the nearest centimeter, the tallest and shortest pupils in a class are 169cm and 150cm.
Find the least and greatest possible range of the students’ heights.
Answer
The intervals in which the given heights, h, lie are 168.5≤h < 169.5cm and 149.5≤h
<150.5cm.
Least possible range 168.5 –150.5= 18cm
Greatest possible range= 169.5 –149.5= 20cm

3.2 The interquartile range and percentiles


The lower quartile, median and upper quartile, as you will recall, divide the values in
a dataset into four parts, with an equal number of values in each part. These three measures are
commonly abbreviated by:
● Q1 for the lower quartile

● Q2 for the median (or middle quartile)

● Q3 for the upper quartile.


The interquartile range is the numerical difference between the upper quartile and the
lower quartile and gives the range of the middle half (50%) of the values, as shown in the
following diagram.

Interquartile range = upper quartile – lower quartile or IQR=Q3–Q1.


The interquartile range is often preferred to the range because it gives a measure of how
varied the more central values are. It is relatively unaffected by extreme values, also called
outliers, and can be found even when the exact values of these are not known.

28
Ungrouped data
The positions of the lower and upper quartiles depend on whether there are an odd or
even number of values in the set of data. One method that we can use to find the quartiles is as
follows. For an even number of ordered values: we split the data into a lower half and an upper
half. Then Q1 and Q3 are the medians of the lower half and upper half, respectively. For an odd
number of ordered values: we split the data into a lower half and an upper half at the median,
which we then discard. Again, Q1 and Q3 are the medians of the lower half and upper half,
respectively.

EXAMPLE 3.2
Find the interquartile range of the eight ordered values 2, 5, 9, 13, 29, 33, 49 and 55.

Answer
2 5 9 13 29 33 49 55

Q1 Q2 Q3
5+ 9
Q1 = = 7,
2
13+29
Q2 = = 21,
2
33+ 49
Q3 = = 41
2
33+ 49 5+ 9
IQR= Q3- Q1= - = 41 – 7 = 34
2 2

29
EXAMPLE 3.3
Find the interquartile range of the seven values 69, 17, 43, 6, 73, 77 and 39.
6 17 39 43 69 73 77

Q1 Q2 Q3

Answer
IQR= Q3- Q1 = 73-17= 56

EXAMPLE 3.4
Find the interquartile range of the 13 grouped values shown in the following stem-and-
leaf diagram.
Answer:
➢ We identify the median as 153, which we now discard. This leaves a lower half
(142 to 151) and an upper half (155 to 168), with six values in each.
Q1

14 2 2 4 8 9 Q3
15 1 3 5 6 7 9
16 5 8

144+ 148
Q1 = = 146,
2

Q2 = 153,
157+ 159
Q3 = = 158
2

IQR= Q3- Q1 = 158 – 146 = 12

30
Grouped data
We can use a cumulative frequency graph to estimate values in any position in a set of
data. This includes the lower quartile, the upper quartile and any chosen percentile.
For grouped data with total frequency n = ∑f, the positions of the quartiles are shown
in the following table.

Quartile lower (Q1) median (Q2) upper (Q3)


Position 𝐧 𝟏
or ∑ F
𝐧 𝟏
or ∑ F
𝟑𝐧 𝟑
𝐨𝐫 ∑ F
𝟒 𝟒 𝟐 𝟐 𝟒 𝟒

The nth percentile is the value that is n% of the way through a set of data. Q1, Q2 and
Q3 are the 25th, 50th and 75th percentiles, respectively. In an ordered dataset with, say, 320
values, Q1, Q2 and Q3 are at the 80th, 160th and 240th values, and the 90th percentile is at the
(0.90×320) =288th value. The range of the middle 80% of a dataset is the difference between
the 10th and 90th percentiles.

EXAMPLE 3.5
The following graph illustrates the times, in minutes, taken by 500 people to complete
a task. Use the graph to find an estimate of:
a. the greatest possible range b. the interquartile range c. the 95th
percentile.

31
Answer:
a. The greatest possible range is equal to the width of the polygon
30 – 2 = 28min
b. We locate the quartiles, then estimate their values by reading from the graph.
n 500
Lower quartile: = = 125th value
4 4
3n 1500
Upper quartile: = = 375𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
4 4

Q1 = 8.0 min
Q3 = 14.5min
IQR = 14.5 – 8 = 6.5 min
c. The 95th percentile is at the (0.95×500) = 475th value
=24.0min

32
Box-and-whisker diagrams
A box-and-whisker diagram (or box plot) is a graphical representation of data, showing
some of its key features. These features are its smallest and largest values, its lower and upper
quartiles, and its median. If drawn by hand, the diagram is best drawn on graph paper and must
include a scale. It takes the form shown in the following diagram, which shows some features
of a dataset denoted by x.

33
3.3 Variance and standard deviation
If we want a measure of variation around the mean, we need to ensure that each
deviation is positive or zero. We can do this by calculating the mean distance of the data values
̅|
∑|𝑿− 𝑿
from the mean, which we call the ‘mean absolute deviation from the mean’,
𝒏
However, it is hard to calculate this accurately or efficiently for large sets of data and it
is difficult to work with algebraically, so this approach is not used in practice. Alternatively,
̅ )2 for all data values and find their mean. This
we can calculate the squared deviation, (𝑿 − 𝑿
is the ‘mean squared deviation from the mean’, which we call the variance of the data.
̅ )2
∑(𝑿− 𝑿
Var(X) =
𝒏

For measurements and deviations in metres, say, the variance is in m2. So, to get a
measure of variation that is also in metres, we take the square root of the variance, which we
call the standard deviation.

∑(𝑿− 𝑿) ̅ 2
Standard deviation of x = √𝐕𝐚𝐫 = √ 𝒏

The formula for variance can be simplified (see appendix at the end of this chapter) to give:
∑𝑿2
Var(X)= ̅2
-𝑿
𝐧

We can find the variance and standard deviation from n, x and x2, which are the number
of values, their sum and the sum of their squares, respectively. We often use the abbreviation
SD(X) to represent the standard deviation of X.
Standard deviation of x for ungrouped ̅)
∑(𝑿− 𝑿
2
data √𝐕𝐚𝐫 = √ 𝒏
Standard deviation of x for grouped data
̅ )2 ∗𝒇 2
√𝐕𝐚𝐫 = √
∑(𝑿− 𝑿
= ̅2
√∑𝑿 ∗𝒇 − 𝑿
∑𝒇 ∑𝒇

EXAMPLE 3.6
For the set of five numbers 3, 9,15, 24 and 29, find the standard deviation
Answer:
2 2
̅)
∑(𝑿− 𝑿 ∑(𝑿) 2
the standard deviation = √𝐕𝐚𝐫 = √ =√ 𝒏 ̅
− 𝑿
𝒏

∑𝟑𝟐 + 𝟗𝟐 +𝟏𝟓𝟐 +𝟐𝟒𝟐 +𝟐𝟗𝟐 (𝟑+𝟗+𝟏𝟓+𝟐𝟒+𝟐𝟗) 𝟐 𝟏𝟕𝟑𝟐 𝟖𝟎 𝟐


=√ − ( ) =√ − ( 𝟓 ) = √𝟗𝟎. 𝟒 = 9.5
𝟓 𝟓 𝟓

34
EXAMPLE 3.7
Find the standard deviation of the values of x given in the following table, correct to 3
significant figures.
x f
12 13
14 28
16 10
Answer
x X2 f X*f X2*f
12 144 13 156 1872
14 196 28 392 5488
16 256 10 160 2560
∑ - 51 708 9920

̅ )2 ∗𝒇 2
SD= √𝐕𝐚𝐫 = √ ∑(𝑿− 𝑿
= ̅ 2 =√9920 − (𝟕𝟎𝟖/𝟓𝟏)2 =1.34
√∑𝑿 ∗𝒇 − 𝑿
∑𝒇 ∑𝒇 𝟓𝟏

EXAMPLE 3.8
Calculate an estimate of the standard deviation of the heights of the 20 children given
in the following table.
Height (metres) No. 1.2– 1.4– 1.5–1.7
children (f) 2 12 6

Height (metres) 1.2– 1.4– 1.5–1.7 ∑


No.
children (f) 2 12 6 20
Mid of class (X) 1.3 1.45 1.6
X2 1.69 2.1 2.56
X*f 2.6 17.4 9.6 29.6
X2*f 3.38 25.23 15.36 43.97
̅ )2 ∗𝒇 2
SD= √𝐕𝐚𝐫 = √ ∑(𝑿− 𝑿
= √
∑𝑿 ∗𝒇 ̅ 2 =√43.97 − (𝟐𝟗. 𝟔/𝟐𝟎)2 = 0.09m
− 𝑿
∑𝒇 ∑𝒇 𝟐𝟎

Advantages of standard deviation Disadvantages of standard deviation


Much simpler to calculate than the IQR. Far more affected by extreme values than
the IQR.
Data values do not have to be ordered.
Gives greater emphasis to large deviations
Easier to work with algebraically when than to small deviations.
doing more advanced work.
Takes account of all data values.

35
Sheet 3
1. For each ordered set of data, A to D, write down these five values: the smallest value;
the lower quartile; the median; the upper quartile; and the largest value.

Set A: 2, 2, 3, 11, 11, 21, 22.


Set B: 6, 6, 6, 11, 13, 17, 19, 20.
Set C: 9, 15, 28, 32, 35, 49.
Set D: 5, 7, 9, 10, 11, 12, 12, 16, 17.

2. Find the range and the interquartile range of the following sets of data.

a. 5, 8, 13, 17, 22, 25, 30


b. 7, 13, 21, 2, 37, 28, 17, 11, 2
c. 42, 47, 39, 51, 73, 18, 83, 29, 41, 64
d. 113, 97, 36, 81, 49, 41, 20, 66, 28, 32, 17, 107
e. 4.6, 0, –2.6, 0.8, –1.9, –3.3, 5.2, –3.2

3. Find the range and the interquartile range of the dataset represented in the following
box plot.

4. The following stem-and-leaf diagram shows the marks out of 50 obtained by 15


students in a science test.
0 9
1 4 Key:
2 5 8
3 0 1 1 7 9 2 5
4 3 5 6 8 represents a mark of
5 0 0 25
out of 50
a Find the range and interquartile range of the marks.

36
b Illustrate the data in a box-and-whisker diagram on graph paper and include a
scale.

5. The following table shows the cumulative frequencies for values of x.

x <0 < 10 < 15 < 25 < 30 < 40


cf 0 12 30 90 102 120
Without drawing a cumulative frequency graph, find:
a. the interquartile range
b. the 85th percentile.

6. Find the mean and the standard deviation for these sets of numbers.

a. 27, 43, 29, 34, 53, 37,19 and 58.


b. 6.2,−8.5, 7.7,−4.3, 13.5 and −11.9.

7. The following table shows the number of pets owned by each of 35 families.
No. pets 0 1 2 3 4 5
No. families (f) 6 12 9 4 3 1
Find the mean and variance of the number of pets.

8. The times spent, in minutes, by 30 girls and by 40 boys on an assignment are detailed
in the following table.

Time spent (min) 20– 30– 40– 60–80


No. girls (f) 6 14 7 3
No. boys (f) 15 11 7 7
For the boys and for the girls, calculate estimates of the mean and standard
deviation.
9. Given that n = 25, ∑x=275 and Var(x)=7, find ∑x2.

37
CHAPTER 4

4. PROBABILITY

In this chapter you will learn how to:


• evaluate probabilities by means of enumeration of equiprobable (i.e. equally likely)
elementary events
• use addition and multiplication of probabilities appropriately
• use the terms mutually exclusive and independent events
• determine whether two events are independent
• calculate and use conditional probabilities.
4.1 Sample Spaces and Events
Statisticians use the word experiment to describe any process that generates a set of
data. A simple example of a statistical experiment is the tossing of a coin. In this experiment,
there are only two possible outcomes, heads or tails. Another experiment might be the
launching of a missile and observing of its velocity at specified times. The opinions of voters
concerning a new sales tax can also be considered as observations of an experiment. We are
particularly interested in the observations obtained by repeating the experiment several times.
In most cases, the outcomes will depend on chance and, therefore, cannot be predicted with
certainty. If a chemist runs an analysis several times under the same conditions, he or she will
obtain different measurements, indicating an element of chance in the experimental procedure.
Even when a coin is tossed repeatedly, we cannot be certain that a given toss will result in a
head. However, we know the entire set of possibilities for each toss.

Random Experiment
An experiment <with known outcomes> that can result in different outcomes, even
though it is repeated in the same manner every time, is called a random
experiment.

38
4.1.2 Sample space

To model and analyze a random experiment, we must understand the set of possible
outcomes from the experiment. In this introduction to probability, we use the basic concepts
of sets and operations on sets. It is assumed that the reader is familiar with these topics.

Sample Space

The set of all possible outcomes of a random experiment is called the sample space
of the experiment. The sample space is denoted as S.

Each outcome in a sample space is called an element or a member of the

sample space, or simply a sample point. If the sample space has a finite number
of elements, we may list the members separated by commas and enclosed in
braces. Thus, the sample space S, of possible outcomes when a coin is flipped,
may be written S = {H, T}, where H and T correspond to heads and tails,
respectively.
Example 4.1:
Consider the experiment of tossing a die. If we are interested in the number that shows

on the top face, the sample space is


S1 = {1, 2, 3, 4, 5, 6}.
If we are interested only in whether the number is even or odd, the sample
space is simply
39
S2 = {even, odd}.
Example 4.1 illustrates the fact that more than one sample space can be
used to describe the outcomes of an experiment. In this case, S1 provides more
information than S2. If we know which element in S1 occurs, we can tell which
outcome in S2 occurs; however, a knowledge of what happens in S2 is of little help
in determining which element in S1 occurs. In general, it is desirable to use the
sample space that gives the most information concerning the outcomes of the
experiment. In some experiments, it is helpful to list the elements of the sample
space systematically by means of a tree diagram.
Example 4.2:
An experiment consists of flipping a coin and then flipping it a second time if a head
occurs. If a tail occurs on the first flip, then a die is tossed once.
To list the elements of the sample space providing the most information, we construct
the tree diagram as shown in the Figure. The various paths along the branches of the tree give
distinct sample points. Starting with the top left branch and moving to the right along the first
path, we get the sample point HH, indicating the possibility that heads occur on two successive
flips of the coin. Likewise, the sample point T3 indicates the possibility that the coin will show
a tail followed by a 3 on the toss of the die. By proceeding along all paths, we see that the
sample space is
S = {HH, HT, T1, T2, T3, T4, T5, T6}.

40
Example 4.3:
Suppose that three items are selected at random from a manufacturing process. Each
item is inspected and classified defective, D, or non-defective, N.
To list the elements of the sample space providing the most information, we construct
the tree diagram as been show in the Figure. Now, the various paths along the branches of the
tree give distinct sample points. Starting with the first path, we get the sample point DDD,
indicating the possibility that all three items inspected are defective. As we proceed along the
other paths, we see that the sample space is
S = {DDD, DDN, DND, DNN, NDD, NDN, NND, NNN}.

41
Sample spaces with a large or infinite number of sample points are best described by a
statement or rule method. For example, if the possible outcomes of an experiment are the set
of cities in the world with a population over 1 million, our sample space is written S = {x | x is
a city with a population over 1 million}, which reads “S is the set of all x such that x is a city
with a population over 1 million.” The vertical bar is read “such that.” Similarly, if S is the set
of all points (x, y) on the boundary or the interior of a circle of radius 2 with center at the origin,
we write the rule
S = {(x, y) | x2 + y2 ≤ 4}.

Example 4.4:
Consider an experiment that selects a cell phone camera and records the recycle time
of a flash (the time taken to ready the camera for another flash). The possible values for this
time depend on the resolution of the timer and on the minimum and maximum recycle times.
However, because the time is positive, it is convenient to define the sample space as simply the
positive real line

42
S = R+ = {x | x > 0}
If it is known that all recycle times are between 1.5 and 5 seconds, the sample space
can be S = {x | 1.5 < x < 5}.
It is useful to distinguish between two types of sample spaces.
Discrete and Continuous Sample Spaces
A sample space is discrete if it consists of a finite or countable infinite set of outcomes.
A sample space is continuous if it contains an interval (either finite or infinite) of real
numbers.

4.2 Events
For any given experiment, we may be interested in the occurrence of certain events
rather than in the occurrence of a specific element in the sample space. For instance, we may
be interested in the event A that the outcome when a die is tossed is divisible by 3. This will
occur if the outcome is an element of the subset A = {3, 6} of the sample space S1 in Example
4.1.
As a further illustration, we may be interested in the event B that the number of
defectives is greater than 1 in Example 4.3. This will occur if the outcome is an element of the
subset
B = {DDN, DND, NDD, DDD} of the sample space S.
To each event we assign a collection of sample points, which constitute a subset of the
sample space. That subset represents all of the elements for which the event is true.

Event
An event is a subset of the sample space of a random experiment.
Example 4.5:
A dice is rolled twice. What is the Event that the sum of the facesisgreaterthan7, given
that the first outcome was a 4?
𝑆= {11, 12, 13, 14, 15, 16, 21, 22, 23, 24, 25, 26, 31, 32, 33, 34, 35, 36, 𝟒𝟏, 𝟒𝟐, 𝟒𝟑, 𝟒𝟒, 𝟒𝟓,
𝟒𝟔, 51, 52, 53, 54, 55, 56, 61, 62, 63, 64, 65, 66}
𝐸= {44, 45, 46}
We can also be interested in describing new events from combinations of existing
events. Because events are subsets, we can use basic set operations such as unions,

intersections, and complements to form other events of interest. Some of the basic set
operations are summarized here in terms of events:

43
• The union of two events is the event that consists of all outcomes that are contained in
either of the two events. We denote the union as E1 ∪ E2.
• The intersection of two events is the event that consists of all outcomes that are
contained in both of the two events. We denote the intersection as E1 ∩ E2.
• The complement of an event in a sample space is the set of outcomes in the sample
space that are not in the event. We denote the complement of the event E as E′. The
notation EC is also used in other literature to denote the complement.

44
Example 4.6:
Consider an experiment that selects a cell phone camera and records the recycle time
of a flash (the time taken to ready the camera for another flash). The possible values for this
time depend on the resolution of the timer and on the minimum and maximum recycle times.
However, because the time is positive, it is convenient to define the sample space as simply the
positive real line
S = R+ = {x | x > 0}
Let, E1 = {x | 10 ≤ x < 12} and E2 = {x | 11 < x < 15}
Then, E1 ∪ E2 = {x | 10 ≤ x < 15}
And E1 ∩ E2 = {x | 11 < x < 12}
Also, E′1 = {x | x < 10 or 12 ≤ x}
And E′1 ∩ E2 = {x | 12 ≤ x < 15}
Example 4.7:
In the tossing of a die, we might let 𝐴 be the event that an even number occurs and B
the event that a number greater than 3 shows.
Then the subsets A={2,4,6} and 𝐵={4,5,6}are subsets Of the same sample space
S={1,2,3,4,5,6}.
𝐴∩𝐵= {4,6}
𝐴∪𝐵= {2,4,5,6}
𝐴′= {1,3,5}
𝐵′= {1,2,3}

45
Mutually Exclusive Events, or Disjoint:
Two events, denoted as E1 and E2, such that: E1 ∩ E2 = Ø
are said to be mutually exclusive.
, that is, if A and B have no elements in common.
𝐴 = {2,4,6}, and 𝐵 = {1,3,5}
𝐴∩𝐵={ } = ∅
Additional results involving events are summarized in the following. The definition of
the complement of an event implies that:
(E′)′ = E
The distributive law for set operations implies that
(A ∪ B) ∩ C = (A ∩ C) ∪ (B ∩ C) and (A ∩ B) ∪ C = (A ∪ C) ∩ (B ∪ C)
DeMorgan’s laws imply that
(A ∪ B)′ = A′ ∩ B′ and (A ∩ B)′ = A′ ∪ B′
Also, remember that
A ∩ B = B ∩ A and A ∪ B = B ∪ A
A ∩ φ = φ. A ∪ φ = A.
A ∩ A′= φ. A ∪ A′ = S.
S′ = φ. φ′ = S.

46
Venn Diagrams:
Diagrams are often used to portray relationships between sets, and these diagrams are
also used to describe relationships between events. We can use Venn diagrams to represent a
sample space and events in a sample space.

S E

47
Example 4.8:
𝑆= {1,2,3,4,5,6,7}, A= {1,2,4,7}, B= { 1,2,3,6}, C= { 1,3,4,5}

A ∪ C = regions 1, 2, 3, 4, 5, and 7,
B′ ∩ A = regions 4 and 7,
A ∩ B ∩ C = region 1,
(A ∪ B) ∩ C′ = regions 2, 6, and 7.

48
4.3 Counting Techniques
In many of the examples in this chapter, it is easy to determine the number of outcomes
in each event. In more complicated examples, determining the outcomes in the sample space
(or an event) becomes more difficult. In these cases, counts of the numbers of outcomes in the
sample space and various events are used to analyze the random experiments. These methods
are referred to as counting techniques. Some simple rules can be used to simplify the
calculations.
Multiplication Rule (for counting techniques)
Assume an operation can be described as a sequence of k steps, and the number of ways to
complete step 1 is n1, and the number of ways to complete step 2 is n2 for each way to
complete step 1, and• the number of ways to complete step 3 is n3 for each way to
complete step 2, and so forth.
The total number of ways to complete the operation is
n1 × n2 × · · · × nk
Example 4.9:
How many sample points are there in the sample space when a pair of dice is thrown once?
Solution:
The first die can land face-up in any one of n1 = 6 ways. For each of these 6 ways, the second
die can also land face-up in n2 = 6 ways.
Therefore, the pair of dice can land in n1*n2 = (6) *(6) = 36 possible ways.
Example 4.10:
How many sample points are there in the design for a website is to consist of four colors,
three fonts, and three positions for an image?
Solution:
From the multiplication rule, n1 = 4, n2 = 3, n3 = 3
Therefore, different designs possible are
n1*n2* n3 = 4 × 3 × 3 = 36.
Example 4.11:
If a 22-member club needs to elect a chair and a treasurer, how many different ways can these
two to be elected?
Solution:
For the chair position, there are 22 total possibilities. For each of those 22 possibilities, there
are 21 possibilities to elect the treasurer.
Using the multiplication rule, we obtain n1 × n2 = 22 × 21 = 462 different ways.
Example 4.12:
Sam is going to assemble a computer by himself. He has the choice of chips from two
brands, a hard drive from four, memory from three, and an accessory bundle from five local
stores. How many different ways can Sam order the parts?
Solution:

49
Since n1 = 2, n2 = 4, n3 = 3, and n4 = 5, there are
nl × n2 × n3 × n4 = 2× 4 × 3 × 5 = 120 different ways to order the parts.
Example 4.13:
How many even four-digit numbers can be formed from the digits 0, 1, 2, 5, 6, and
9 if each digit can be used only once?
Solution:
Since the number must be even, we have only n1 = 3 choices for the unit’s position.
However, for a four-digit number the thousands position cannot be 0. Hence, we consider the
unit’s position in two parts, 0 or not 0.
If the unit’s position is 0 (i.e., n1 = 1), we have n2 = 5 choices for the thousands position,
n3 = 4 for the hundreds position, and n4 = 3 for the tens position. Therefore, in this case we
have a total of n1*n2*n3*n4 = (1)(5)(4)(3) = 60 even four-digit numbers.
On the other hand, if the unit’s position is not 0 (i.e., n1 = 2), we have n2 = 4 choices for
the thousands position, n3 = 4 for the hundreds position, and n4 = 3 for the tens position. In this
situation, there are a total of n1*n2*n3*n4 = (2)(4)(4)(3) = 96 even four-digit numbers.
Since the above two cases are mutually exclusive, the total number of even four-digit
numbers can be calculated as 60 + 96 = 156.

Permutations Another useful calculation finds the number of ordered sequences of


the elements of a set. Consider a set of elements, such as S = {a, b, c}. A permutation of the
elements is an ordered sequence of the elements. For example, abc, acb, bac, bca, cab, and cba
are all of the permutations of the elements of S. Thus, we see that there are 6 distinct
arrangements. Using Multiplication Rule, we could arrive at answer 6 without actually listing
the different orders by the following arguments: There are n1 = 3 choices for the first position.
No matter which letter is chosen, there are always n2 = 2 choices for the second position. No
matter which two letters are chosen for the first two positions, there is only n3 = 1 choice for
the last position, giving a total of
n1*n2*n3 = (3)(2)(1) = 6
permutations In general, n distinct objects can be arranged in n (n − 1)(n − 2) ・ ・ ・
(3)(2)(1) ways.

Permutations Rule (for counting techniques)


The number of permutations of n different elements is n! where
n! = n × (n − 1) × (n − 2) × · · · × 2 × 1

50
This result follows from the multiplication rule. A permutation can be constructed as
follows. Select the element to be placed in the first position of the sequence from the n
elements, then select the element for the second position from the remaining n − 1 elements,
then select the element for the third position from the remaining n − 2 elements, and so forth.
Permutations such as these are sometimes referred to as linear permutations.
The number of permutations of n objects is n!.
The number of permutations of the four letters a, b, c, and d will be 4! = 24. Now
consider the number of permutations that are possible by taking two letters at a time from four.
These would be ab, ac, ad, ba, bc, bd, ca, cb, cd, da, db, and dc. Using Multiplication Rule
again, we have two positions to fill, with n1 = 4 choices for the first and then n2 = 3 choices for
the second, for a total of n1n2 = (4)(3) = 12 permutations. In general, n distinct objects taken r
at a time can be arranged in n(n − 1)(n − 2) ・ ・ ・ (n − r + 1) ways. We represent this product
by the symbol
Permutations of Subsets
The number of permutations of subsets of r elements selected from a set of n different
𝑛!
elements is nPr = Prn = n × (n − 1) × (n − 2) × · · · × (n − r + 1) =
(𝑛 − 𝑟)!

51
Example 4.14:
In one year, three awards (research, teaching, and service) will be given to a class of
25 graduate students in a statistics department. If each student can receive at most one
award, how many possible selections are there?
Solution:
Since the awards are distinguishable, it is a permutation problem. The total number of
sample points is
𝟐𝟓! 𝟐𝟓!
25P3 = = = (25)(24)(23) = 13, 800.
(𝟐𝟓 − 𝟑)! 𝟐𝟐!

Example 4.15:
A president and a treasurer are to be chosen from a student club consisting of 50 people. How
many different choices of officers are possible if
(a) there are no restrictions.
(b) A will serve only if he is president.
(c) B and C will serve together or not at all.
(d) D and E will not serve together?
Solution:
(a) The total number of choices of officers, without any restrictions, is
𝟓𝟎!
50P2 = = (50)(49) = 2450.
𝟒𝟖!

(b) Since A will serve only if he is president, we have two situations here:
(i) A is selected as the president, which yields 49 possible outcomes for the treasurer’s
position, or
(ii) officers are selected from the remaining 49 people without A, which has the number of
choices
49P2 = (49)(48) = 2352.
Therefore, the total number of choices is 49 + 2352 = 2401.
(c) The number of selections when B and C serve together is 2. The number of selections
when both B and C are not chosen is
48P2 = 2256.
Therefore, the total number of choices in this situation is 2 + 2256 = 2258.
(d) The number of selections when D serves as an officer but not E is (2)(48) = 96, where 2 is
the number of positions D can take and 48 is the number of selections of the other officer
from the remaining people in the club except E.
The number of selections when E serves as an officer but not D is also (2)(48) = 96.
The number of selections when both D and E are not chosen is

52
48P2 = 2256.
Therefore, the total number of choices is (2)(96) + 2256 = 2448.
This problem also has another short solution: Since D and E can only serve together in 2
ways, the answer is 2450 − 2 = 2448.

Example 4.16:
A printed circuit board has eight different locations in which a component can be
placed. If four different components are to be placed on the board, how many different designs
are possible?
Solution:
Each design consists of selecting a location from the eight locations for the first
component, a location from the remaining seven for the second component, a location from the
remaining six for the third component, and a location from the remaining five for the fourth
component. Therefore,
𝟖!
8P4 = = 8 × 7 × 6 × 5 = 1680 different designs are possible.
𝟒!

Permutations of Similar Objects


The number of permutations of n = n1 + n2 + · · · + nr objects of which n1 are
of one type, n2 are of a second type, …, and nr are of an rth type is
𝒏!
𝒏𝟏! 𝒏𝟐! 𝒏𝟑! … 𝒏𝒓!

Example 4.17:
In a college football training session, the defensive coordinator needs to have 10 players
standing in a row. Among these 10 players, there are 1 freshman, 2 sophomores, 4 juniors,
and 3 seniors. How many different ways can they be arranged in a row if only their class
level will be distinguished?
Solution:
We find that total number of players n = 10, n1=1, n2= 2, n3= 4, n4= 3 So,
𝟏𝟎!
the total number of arrangements is = = 12, 600.
𝟏! 𝟐! 𝟒! 𝟑!

Example 4.18:
In how many ways can 7 graduate students be assigned to 1 triple and 2 double hotel rooms
during a conference?
Solution:
53
We find that total number of graduate n = 7, n1=3, n2= 2, n3= 2
𝟕!
The total number of possible partitions would be is = = 210.
𝟑! 𝟐! 𝟐!

Combinations Another counting problem of interest is the number of subsets of r


elements that can be selected from a set of n elements. Here, order is not important. These are
called combinations.

Example 4.19:
A young boy asks his mother to get 5 Game-BoyTM cartridges from his collection of 10
arcades and 5 sports games. How many ways are there that his mother can get 3 arcade and 2
sports games?
Solution:
The number of ways of selecting 3 cartridges from 10 is

(10
𝟏𝟎!
3
) = 𝟑! (𝟏𝟎 − 𝟑)! = 120.

The number of ways of selecting 2 cartridges from 5 is

(𝟓𝟐) = 𝟐! (𝟓 − 𝟐)! = 10.


𝟓!

54
4.4 Probability of an Event
The probability of event A is the sum of the weights of all sample points in A. Therefore,
0 ≤ P(A) ≤ 1, P(φ) = 0, and P(S) = 1.
Furthermore, if A1, A2, A3, . . . is a sequence of mutually exclusive events, then
P(A1 ∪ A2 ∪ A3 ∪ ・・ ・) = P(A1) + P(A2) + P(A3) + ・ ・ ・ .
Example 4.20:
A coin is tossed twice. What is the probability that at least 1 head occurs?
Solution :
The sample space for this experiment is S = {HH, HT, TH, TT}.
If the coin is balanced, each of these outcomes is equally likely to occur. Therefore, we
assign a probability of ω to each sample point. Then 4ω = 1, or ω = 1/4. If A represents
the event of at least 1 head occurring, then
𝟏 𝟏 𝟏 𝟑
A = {HH, HT, TH} and P(A) = 𝟒 + 𝟒 + =
𝟒 𝟒

Example 4.21:
A die is loaded in such a way that an even number is twice as likely to occur as an odd
number. If E is the event that a number less than 4 occurs on a single toss of the die, find P(E),
then let A be the event that an even number turns up and let B be the event that a number
divisible by 3 occurs. Find P (A ∪ B) and P(A ∩ B).
Solution:
The sample space is S = {1, 2, 3, 4, 5, 6}. We assign a probability of w to each odd
number and a probability of 2w to each even number. Since the sum of the probabilities must
be 1, we have 9w = 1 or w = 1/9. Hence, probabilities of 1/9 and 2/9 are assigned to each odd
and even number, respectively. Therefore,
𝟏 𝟐 𝟏 𝟒
E = {1, 2, 3} and P(E) = 𝟗 + 𝟗 + =
𝟗 𝟗

For the events A = {2, 4, 6} and B = {3, 6}, we have


A ∪ B = {2, 3, 4, 6} and A ∩ B = {6}.
By assigning a probability of 1/9 to each odd number and 2/9 to each even number, we have
𝟐 𝟏 𝟐 𝟐 𝟕
P (A ∪ B) = = 𝟗 + 𝟗 + 𝟗 + =
𝟗 𝟗
𝟐
and P (A ∩ B) = 𝟗

If an experiment can result in any one of N different equally likely outcomes, and if exactly
n of these outcomes corresponds to event A, then the probability of event A is
𝐧
P(A) = 𝐍

55
Example 4.22:
A statistics class for engineers consists of 25 industrial, 10 mechanical, 10 electrical,
and 8 civil engineering students. If a person is randomly selected by the instructor to answer a
question, find the probability that the student chosen is
(a) an industrial engineering major and
(b) a civil engineering or an electrical engineering major.
Solution:
Denote by I, M, E, and C the students majoring in industrial, mechanical, electrical, and
civil engineering, respectively. The total number of students in the class is 53, all of whom are
equally likely to be selected.
(a) Since 25 of the 53 students are majoring in industrial engineering, the probability of event
I, selecting an industrial engineering major at random, is
𝟐𝟓
P(I) =𝟓𝟑

(b) Since 18 of the 53 students are civil or electrical engineering majors, it follows that
𝟏𝟖
P (C ∪ E) =𝟓𝟑

4.5 Additive Rules


Often it is easiest to calculate the probability of some event from known probabilities
of other events. This may well be true if the event in question can be represented as the union
of two other events or as the complement of some event. Several important laws that frequently
simplify the computation of probabilities follow. The first, called the additive rule, applies to
unions of events.

If A and B are two events, then


P(A ∪ B) = P(A) + P(B) − P(A ∩ B).

Three or More Events More complicated probabilities, such as P (A ∪ B ∪ C), can be


determined by repeated use of the previous Equation and by using some basic set operations.
For example,

56
P (A ∪ B ∪ C) = P(A) + P(B) + P(C) – P (A ∩ B) − P (A ∩ C) – P (B ∩ C) + P (A ∩ B ∩
C)

If A and B are mutually exclusive, then


P (A ∪ B) = P(A) + P(B).
If A1, A2, . . ., An are mutually exclusive, then
P (A1 ∪ A2 ∪ · · · ∪ An) = P(A1) + P(A2) + · · · + P(An).

Example 4.23:
John is going to graduate from an industrial engineering department in a university by
the end of the semester. After being interviewed at two companies he likes, he assesses that his
probability of getting an offer from company A is 0.8, and his probability of getting an offer
from company B is 0.6. If he believes that the probability that he will get offers from both
companies is 0.5, what is the probability that he will get at least one offer from these two
companies?
Solution:
Using the additive rule, we have
P (A ∪ B) = P(A) + P(B) – P (A ∩ B) = 0.8 + 0.6 − 0.5 = 0.9.

Example 4.24:
What is the probability of getting a total of 7 or 11 when a pair of fair dice is tossed?
Solution:
Let A be the event that 7 occurs and B the event that 11 comes up. Now, a total of 7
occurs for 6 of the 36 sample points, and a total of 11 occurs for only 2 of the sample points.
Since all sample points are equally likely, we have
P(A) = 1/6 and
P(B) = 1/18.
The events A and B are mutually exclusive, since a total of 7 and 11 cannot both occur on the
same toss. Therefore,
1 1 2
P (A ∪ B) = P(A) + P(B) =6 + =9
18

This result could also have been obtained by counting the total number of points
for the event A ∪ B, namely 8, and writing
𝑛 8 2
P (A ∪ B) = 𝑁 = =9
36

Example 4.25:

57
If the probabilities are, respectively, 0.09, 0.15, 0.21, and 0.23 that a person purchasing
a new automobile will choose the color green, white, red, or blue, what is the probability that
a given buyer will purchase a new automobile that comes in one of those colors?
Solution:
Let G, W, R, and B be the events that a buyer selects, respectively, a green, white, red, or blue
automobile. Since these four events are mutually exclusive, the probability is
P (G ∪W ∪ R ∪ B) = P(G) + P(W) + P(R) + P(B)
= 0.09 + 0.15 + 0.21 + 0.23 = 0.68.
.
If A and A′ are complementary events, then
P(A) + P(A)′ = 1

Example 4.26:
If the probabilities that an automobile mechanic will service 3, 4, 5, 6, 7, or 8 or more
cars on any given workday are, respectively, 0.12, 0.19, 0.28, 0.24, 0.10, and 0.07, what is the
probability that he will service at least 5 cars on his next day at work?
Solution:
Let E be the event that at least 5 cars are serviced. Now, P(E) = 1 − P(E′),
where E′ is the event that fewer than 5 cars are serviced. Since
P(E′) = 0.12 + 0.19 = 0.31,
P(E) = 1 − 0.31 = 0.69.

4.6 Conditional Probability


Conditional Probability The probability of an event B occurring when it is known that
some event A has occurred is called a conditional probability and is denoted by P(B|A). The
symbol P(B|A) is usually read “the probability that B occurs given that A occurs” or simply
“the probability of B, given A.” Consider the event B of getting a perfect square when a die is
tossed. The die is constructed so that the even numbers are twice as likely to occur as the odd
numbers. Based on the sample space S = {1, 2, 3, 4, 5, 6}, with probabilities of 1/9 and 2/9
assigned, respectively, to the odd and even numbers, the probability of B occurring is 3/9. Now
suppose that it is known that the toss of the die resulted in a number greater than 3. We are now
dealing with a reduced sample space A = {4, 5, 6}, which is a subset of S. To find the probability
that B occurs, relative to the space A, we must first assign new probabilities to the elements of
A proportional to their original probabilities such that their sum is 1. Assigning a probability
of w to the odd number in A and a probability of 2w to the two even numbers, we have 5w =

58
1, or w = 1/5. Relative to the space A, we find that B contains the single element 4. Denoting
this event by the symbol B|A, we write B|A = {4}, and hence
𝟐
P(B|A) =𝟓.

This example illustrates that events may have different probabilities when considered
relative to different sample spaces. We can also write
𝟐
𝐏(𝐀 ∩ 𝐁) 𝟐
P(B|A) = 𝟗𝟓 = =𝟓
𝐏(𝐀)
𝟗

where P (A ∩ B) and P(A) are found from the original sample space S. In other words,
a conditional probability relative to a subspace A of S may be calculated directly from the
probabilities assigned to the elements of the original sample space S.
The conditional probability of B, given A, denoted by P(B|A), is defined by
𝐏(𝐀 ∩ 𝐁)
P(B|A) = 𝐏(𝐀) , provided P(A) > 0.
Example 4.27:
Suppose that our sample space S is the population of adults in a small town who have
completed the requirements for a college degree. We shall categorize them according to gender
and employment status. The data are given in the Table
Employed Unemployed Total
Male 460 40 500
Female 140 260 400
Total 600 300 900
One of these individuals is to be selected at random for a tour throughout the country
to publicize the advantages of establishing new industries in the town. We shall be concerned
with the following events:
M: a man is chosen,
E: the one chosen is employed.
Using the reduced sample space E, we find that
460
460
P(M|E) = 900
600 = 600.
900

Example 4.28:
The probability that a regularly scheduled flight departs on time is P(D) = 0.83; the
probability that it arrives on time is P(A) = 0.82; and the probability that it departs and arrives
on time is P (D ∩A) = 0.78. Find the probability that a plane
(a) arrives on time, given that it departed on time, and

59
(b) departed on time, given that it has arrived on time.
(c) arrives on time, given that it did not depart on time.
Solution:
(a) The probability that a plane arrives on time, given that it departed on time, is

P (D ∩ A) 0.78
P(A|D) = = 0.83 = 0.94
P(D)

(b) The probability that a plane departed on time, given that it has arrived on
time, is
P (D ∩ A) 0.78
P(D|A) = = 0.82 = 0.95
P(A)

(c) The probability that a plane arrives on time, given that it did not depart on time., is
P (A ∩ D‵) P (A)−P (A ∩ D) 0.82 − 0.78
P(A|D‵) = = = = 0.24
P(D‵) P(D‵) 1− 0.83

60
4.7 Independence
In some cases, the conditional probability of P (B | A) might equal P(B). In this special
case, knowledge that the outcome of the experiment is in event A does not affect the probability
that the outcome is in event B.
Independence (two events)
Two events are independent if any one of the following equivalent statements is true:
(1) P (A | B) = P(A)
(2) P (B | A) = P(B)
(3) P (A ∩ B) = P(A)P(B)

Example 4.29:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?

Solution
Let L and R denote the events that the left and right devices operate, respectively. There
is a path only if both operate. The probability that the circuit operates is
P (L and R) = P (L ∩ R) = P(L)P(R) = (0.80) *(0.90) = 0.72
Practical Interpretation: Notice that the probability that the circuit operates degrades to
approximately 0.7 when all devices are required to be functional. The probability that each
device is functional needs to be large for a circuit to operate when many devices are connected
in series.
Example 4.30:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?

Solution
Let T and B denote the events that the top and bottom devices operate, respectively.
There is a path if at least one device operates. The probability that the circuit operates is
P (T or B) = P (T U B) = P (T) + P (B) - P (T)* P (B) = 0.95+0.9-0.95*0.9= 0.995
Another solution

61
P (T or B) = 1 – P [(T or B)′] = 1 − P(T′ and B′) = 1- [P(T′)* P(B′)]= 1- (0.05*0.1) =0.995
Example 4.31:
The following circuit operates only if there is a path of functional devices from left to
right. The probability that each device functions is shown on the graph. Assume that devices
fail independently. What is the probability that the circuit will operate?

Solution:
The solution can be obtained from a partition of the graph into three columns. Let L
denote the event that there is a path of functional devices only through the three units on the
left. From the independence and based on the previous example,
P(L) = 1 − 0.13
Similarly, let M denote the event that there is a path of functional devices only through
the two units in the middle. Then,
P(M) = 1 − 0.052
The probability that there is a path of functional devices only through the one unit on
the right is simply the probability that the device functions, namely, 0.99. Therefore, with the
independence assumption used again, the solution is
(1 − 0.13) (1 − 0.052) (0.99) = 0.987
Example 4.32:
Suppose that we have a fuse box containing 20 fuses, of which 5 are defective. If 2
fuses are selected at random and removed from the box in succession without replacing the
first, what is the probability that both fuses are defective?
Solution:
We shall let A be the event that the first fuse is defective and B the event that the second
fuse is defective; then we interpret A ∩ B as the event that A occurs and then B occurs after A
has occurred.
The probability of first removing a defective fuse is 5/20 = 1/4; then the probability of
removing a second defective fuse from the remaining 4 is 4/19. Hence,
1 4 1
P (A ∩ B) = P (A) * P (B) = 4* 19 = 19

62
Example 4.32:
One bag contains 4 white balls and 3 black balls, and a second bag contains 3 white balls and
5 black balls. One ball is drawn from the first bag and placed unseen in the second bag. What
is the probability that a ball now drawn from the second bag is black?
Solution:
Let B1, B2, and W1 represent, respectively, the drawing of a black ball from bag 1, a
black ball from bag 2, and a white ball from bag 1. We are interested in the union of the mutually
exclusive events B1 ∩ B2 and W1 ∩ B2. The various possibilities and their probabilities are
illustrated in Figure. Now
P [(B1 ∩ B2) or (W1 ∩ B2)] = P (B1 ∩ B2) + P (W1 ∩ B2)
3 6 4 5 38
= P(B1) P(B2|B1) + P(W1) P(B2|W1) = (7 ∗ 9) + (7 ∗ 9) = 63

++

63
Sheet 4
1. List the elements of each of the following sample spaces:
A. the set of integers between 1 and 50 divisible by 8
B. the set S = {x | x2 + 4x − 5 = 0}
C. the set of outcomes when a coin is tossed until a tail, or three heads appear
D. the set S = {x | 2x − 4 ≥ 0 and x < 1}
E. the set S = {x | 2x − 4 ≥ 0 or x < 1}
2. Which of the following events are equal?

(a) A = {1, 3}.


(b) B = {x | x is a number on a die}.
(c) C = {x | x2 − 4x+3 = 0}.
(d) D = {x | x is the number of heads when six coins are tossed}.
3. An experiment consists of tossing a die and then flipping a coin once, if the number on
the die is even. If the number on the die is odd, the coin is flipped twice, construct a
tree diagram to show the 18 elements of the sample space S, and the find
(a) list the elements corresponding to the event A that a number less than 3 occurs on
the die.
(b) list the elements corresponding to the event B that two tails occur.
(c) list the elements corresponding to the event A′;
(d) list the elements corresponding to the event A′∩B;
(e) list the elements corresponding to the event A∪B.
4. If S = {0, 1, 2, 3, 4, 5, 6, 7, 8, 9} and A = {0, 2, 4, 6, 8}, B = {1, 3, 5, 7, 9}, C = {2, 3,
4, 5}, and D = {1, 6, 7}, list the elements of the sets corresponding to the following
events:
(a) A ∪ C;
(b) A ∩ B;
(c) C′;
(d) (C′ ∩ D) ∪ B;
(e) (S ∩ C)′;
(f) A ∩ C ∩ D′.
5 If S = {x | 0 < x < 12}, M = {x | 1 < x < 9}, and N = {x | 0 < x < 5}, find
(a) M ∪ N;
(b) M ∩ N;
(c) M′ ∩ N′.

64
6 If an experiment consists of throwing a die and then drawing a letter at random from
the English alphabet, how many points are there in the sample space?
7 Students at a private liberal arts college are classified as being freshmen, sophomores,
juniors, or seniors, and also according to whether they are male or female. Find the total
number of possible classifications for the students of that college.
8 A certain brand of shoes comes in 5 different styles, with each style available in 4
distinct colors. If the store wishes to display pairs of these shoes showing all of its
various styles and colors, how many different pairs will the store have on display?
9 A California study concluded that following 7 simple health rules can extend a man’s
life by 11 years on the average and a woman’s life by 7 years. These 7 rules are as
follows: no smoking, get regular exercise, use alcohol only in moderation, get 7 to 8
hours of sleep, maintain proper weight, eat breakfast, and do not eat between meals. In
how many ways can a person adopt 5 of these rules to follow

(a) if the person presently violates all 7 rules?


(b) if the person never drinks and always eats breakfast?
10 In a fuel economy study, each of 3 race cars is tested using 5 different brands of gasoline
at 7 test sites located in different regions of the country. If 2 drivers are used in the
study, and test runs are made once under each distinct set of conditions, how many test
runs are needed?
11 In how many different ways can a true-false test consisting of 9 questions be answered?
12 A witness to a hit-and-run accident told the police that the license number contained the
letters RLH followed by 3 digits, the first of which was a 5. If the witness cannot recall
the last 2 digits, but is certain that all 3 digits are different, find the maximum number
of automobile registrations that the police may have to check.
13 In how many ways can 5 starting positions on a basketball team be filled with 8 men
who can play any of the positions?
14 Find the number of ways that 6 teachers can be assigned to 4 sections of an introductory
psychology course if no teacher is assigned to more than one section.
15 Three lottery tickets for first, second, and third prizes are drawn from a group of 40
tickets. Find the number of sample points in S for awarding the 3 prizes if each
contestant holds only 1 ticket.
16 How many distinct permutations can be made from the letters of the word INFINITY ?

65
17 In how many ways can 3 oaks, 4 pines, and 2 maples be arranged along a property line
if one does not distinguish among trees of the same kind?
18 How many ways are there to select 3 candidates from 8 equally qualified recent
graduates for openings in an accounting firm?
19 The probability that an American industry will locate in Shanghai, China, is 0.7, the
probability that it will locate in Beijing, China, is 0.4, and the probability that it will
locate in either Shanghai or Beijing or both is 0.8. What is the probability that the
industry will locate

(a) in both cities?


(b) in neither city?
20 An automobile manufacturer is concerned about a possible recall of its best-selling four-
door sedan. If there were a recall, there is a probability of 0.25 of a defect in the brake
system, 0.18 of a defect in the transmission, 0.17 of a defect in the fuel system, and 0.40
of a defect in some other area.
(a) What is the probability that the defect is the brakes or the fueling system if the
probability of defects in both systems simultaneously is 0.15?
(b) What is the probability that there are no defects in either the brakes or the fueling
system?

++

66
CHAPTER 5

RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS

5.1 Concept of a Random Variable


Statistics is concerned with making inferences about populations and population
characteristics. Experiments are conducted with results that are subject to chance. The testing
of a number of electronic components is an example of a statistical experiment, a term that is
used to describe any process by which several chance observations are generated. It is often
important to allocate a numerical description to the outcome. For example, the sample space
giving a detailed description of each possible outcome when three electronic components are
tested may be written S = {NNN,NND,NDN,DNN,NDD,DND,DDN,DDD}, where N denotes
non defective and D denotes defective. One is naturally concerned with the number of
defectives that occur. Thus, each point in the sample space will be assigned a numerical value
of 0, 1, 2, or 3. These values are, of course, random quantities determined by the outcome of
the experiment. They may be viewed as values assumed by the random variable X, the number
of defective items when three electronic components are tested.
A random variable is a function that associates a real number with each element in the
sample space.

We shall use a capital letter, say X, to denote a random variable and its corresponding
small letter, x in this case, for one of its values. In the electronic component testing illustration
above, we notice that the random variable X assumes the value 2 for all elements in the subset
E = {DDN, DND, NDD}
of the sample space S. That is, each possible value of X represents an event that is a
subset of the sample space for the given experiment.

67
Example 5.1:
Two balls are drawn in succession without replacement from an urn containing 4 red
balls and 3 black balls. The possible outcomes and the values y of the random variable Y, where
Y is the number of red balls, are
Sample Space y
RR 2
RB 1
BR 1
BB 0

Example 5.2:
Let X be the random variable defined by the waiting time, in hours, between successive
speeders spotted by a radar unit. The random variable X takes on all values x for which x ≥ 0.
The outcomes of some statistical experiments may be neither finite nor countable. Such
is the case, for example, when one investigates measuring the distances that a certain make of
automobile will travel over a prescribed test course on 5 liters of gasoline. Assuming distance
to be a variable measured to any degree of accuracy, then clearly, we have an infinite number
of possible distances in the sample space that cannot be equated to the number of whole
numbers. Or, if one were to record the length of time for a chemical reaction to take place, once
again the possible time intervals making up our sample space would be infinite in number and
uncountable. We see now that all sample spaces need not be discrete.
If a sample space contains a finite number of possibilities or an unending sequence with as
many elements as there are whole numbers, it is called a discrete sample space.

A random variable is called a discrete random variable if its set of possible outcomes
is countable. The random variables in Example 5.1 are discrete random variables. But a random
variable whose set of possible values is an entire interval of numbers is not discrete. When a
random variable can take on values on a continuous scale, it is called a continuous random
variable. Often the possible values of a continuous random variable are precisely the same
values that are contained in the continuous sample space. Obviously, the random variables
described in Example 5.2 are continuous random variables.
In most practical problems, continuous random variables represent measured data, such
as all possible heights, weights, temperatures, distance, or life periods, whereas discrete random
variables represent count data, such as the number of defectives in a sample of k items or the
number of highway fatalities per year in a given state.

68
If a sample space contains an infinite number of possibilities equal to the number of points
on a line segment, it is called a continuous sample space.

5.2 Discrete Probability Distributions


A discrete random variable assumes each of its values with a certain probability. In the
case of tossing a coin three times, the variable X, representing the number of heads, assumes
the value 2 with probability 3/8, since 3 of the 8 equally likely sample points result in two heads
and one tail. The probability distribution of a random variable X is a description of the
probabilities associated with the possible values of X. For a discrete random variable, the
distribution is often specified by just a list of the possible values along with the probability of
each. For other cases, probabilities are expressed in terms of a formula.
x 0 1 2 3
P(X=x) 1/9 3/9 3/9 1/9

Note that the values of x exhaust all possible cases, and hence the probabilities add to
1. Frequently, it is convenient to represent all the probabilities of a random variable X by a
formula. Such a formula would necessarily be a function of the numerical values x that we shall
denote by f(x), g(x), r(x), and so forth. Therefore, we write f(x) = P(X = x); that is, f(3) = P(X
= 3). The set of ordered pairs (x, f(x)) is called the probability function, probability mass
function, or probability distribution of the discrete random variable X.
The set of ordered pairs (x, f(x)) is a probability function, probability mass function, or
probability distribution of the discrete random variable X if, for each possible outcome x,
1. f(x) ≥ 0,
2. ∑f(x) = 1,
3. P (X = x) = f(x).

Example 5.3:
The shipment of 20 similar laptop computers to a retail outlet contains 3 that are
defective. If a school makes a random purchase of 2 of these computers, find the probability
distribution for the number of defectives.
Solution:
Let X be a random variable whose values x are the possible numbers of defective
computers purchased by the school. Then x can only take the numbers 0, 1, and 2. Now

(3 17
0)( 2 ) 136
f(0) = P(X = 0) = 20 =
(2) 190

69
(3)(17) 51
1 1
f(1) = P(X = 1) = 20 =
(2) 190

(3 17
2)( 0 ) 3
f(2) = P(X = 2) = 20 =
(2) 190

Thus, the probability distribution of X is


X 0 1 2
P(X=x), f(x) 136 51 3
190 190 190

There are many problems where we may wish to compute the probability that the
observed value of a random variable X will be less than or equal to some real number x. Writing
F(x) = P (X ≤ x) for every real number x, we define F(x) to be the cumulative distribution
function of the random variable X.
The cumulative distribution function F(x) of a discrete random variable X with
probability distribution f(x) is
F(x) = P(X ≤ x) = ∑𝐱≤∞ f(x), for −∞ < x < ∞.

5.3 Continuous Probability Distributions


A continuous random variable has a probability of 0 of assuming exactly any of its
values. Consequently, its probability distribution cannot be given in tabular form. At first this
may seem startling, but it becomes more plausible when we consider a particular example. Let
us discuss a random variable whose values are the heights of all people over 21 years of age.
Between any two values, say 163.5 and 164.5 centimeters, or even 163.99 and 164.01
centimeters, there are an infinite number of heights, one of which is 164 centimeters. The
probability of selecting a person at random who is exactly 164 centimeters tall and not one of
the infinitely large set of heights so close to 164 centimeters that you cannot humanly measure
the difference is remote, and thus we assign a probability of 0 to the event. This is not the case,
however, if we talk about the probability of selecting a person who is at least 163 centimeters
but not more than 165 centimeters tall. Now we are dealing with an interval rather than a point
value of our random variable.
Note that when X is continuous,
P (a < X ≤ b) = P (a < X < b) + P (X = b) = P (a < X < b).
That is, it does not matter whether we include an endpoint of the interval or not.

70
Although the probability distribution of a continuous random variable cannot be
presented in tabular form, it can be stated as a formula. Such a formula would necessarily be a
function of the numerical values of the continuous random variable X and as such will be
represented by the functional notation f(x). In dealing with continuous variables, f(x) is usually
called the probability density function, or simply the density function, of X. Since X is
defined over a continuous sample space, it is possible for f(x) to have a finite number of
discontinuities. However, most density functions that have practical applications in the analysis
of statistical data are continuous and their graphs may take any of several forms, some of which
are shown in Figure 5-1. Because areas will be used to represent probabilities and probabilities
are positive numerical values, the density function must lie entirely above the x axis.

Figure 5-0-1: Typical density functions


A probability density function is constructed so that the area under its curve bounded
by the x axis is equal to 1 when computed over the range of X for which f(x) is defined. Should
this range of X be a finite interval, it is always possible to extend the interval to include the
entire set of real numbers by defining f(x) to be zero at all points in the extended portions of
the interval. In Figure 5-2, the probability that X assumes a value between a and b is equal to
the shaded area under the density function between the ordinates at x = a and x = b, and from
integral calculus is given by

71
Figure 0-2: P (a < X < b)

Example 5.4:
Suppose that the error in the reaction temperature, in ◦C, for a controlled laboratory
experiment is a continuous random variable X having the probability density function

(a) Verify that f(x) is a density function.


(b) Find P(0 < X ≤ 1).
Solution:
(a) Obviously, f(x) ≥ 0. To verify condition 2, we have

72
(b) P(0 < X ≤ 1)

Example 5.5:
For the density function of Example 5.4, find F(x), and use it to evaluate P (0 < X ≤ 1).
Solution:
For −1 < x < 2,

Therefore,

P (0 < X ≤ 1) = F(1) − F(0) =2/9-1/9=1/9

REFERENCES

73

You might also like