Statistics
Statistics
Learning Outcomes
Introduction
Statistics is a branch of Applied Mathematics which involves collecting, organizing, analyzing and
interpreting data to make suitable and informed decisions. Statistics is applicable to real life in a wide
range of situations; some of which are given below:
1) Medical Field: Statistics help monitor and predict disease spread. In the COVID-19
pandemic, statistical techniques were used to measure the spread of the virus and analyze
infection rates, mortality rates and vaccine efficacy.
2) Weather Forecasting: Meteorologists use statistics to predict weather patterns and the
trajectory of hurricanes.
3) Sports: In sports such as cricket and football, player performance is measured using statistics
such as batting averages and the number of goals scored per season.
4) Business: Statistics is used to identify market trends, analyze sales data and determine
satisfaction levels of customers.
5) Elections: Statistical techniques can be used to predict election outcomes through opinion
polls, voter turnout rates and voting patterns.
Data
In statistics, the term ‘data’ refers to a collection of observations, measurements or facts gathered
through research. Data can be numerical: such as measurements or counts, or categorical such as labels
or classifications.
Discrete data is data which can be counted. Discrete data values are distinct values. Examples of
discrete data include the number of students in a class (e.g. 30), goals scored in a football match (e.g. 5),
number of books in your bag (e.g. 15)
Population: For a particular study, the population is the entire set of items or individuals being studied.
Sample: For a particular study, a sample is selected from the population. That is, a sample is a subset of
the population.
Sample statistics are numerical values calculated from a sample. A population parameter is a value that
describes a characteristic of the entire population. For example, if we wish to study the prevalence of
diabetes in adults in Barbados, it is impractical to measure the blood sugar level of all adults in Barbados
(this may be impossible for a large population). To address this, we usually take a sample of adults from
the population and study the sample. The average blood sugar level of adults in the sample is referred to
as a sample statistic. The average blood level of the population is estimated by mathematical techniques,
using the sample statistic, and this average value is regarded as a population parameter.
• Sample statistic: The average cholesterol level in 300 adults over age 65 in Kingston
Population parameter: The average cholesterol level of all people over age 65 in Kingston
• Sample statistic: The percentage of 250 trial patients recovering fully from cataract surgery in
POS
Population parameter: The percentage of all patients recovering fully from cataract surgery in
POS
The following are examples of a population and a sample: (NOTE: A sample is always selected from the
population):
Population Sample
1) All students in Couva East Secondary 30 students selected randomly from Couva
School East Secondary School to go on a field trip.
2) All Corolla cars manufactured by 50 Corolla cars selected from Trinidad for
Toyota in 2023 brake testing in 2023
3) All trees in Queens Park Savannah 10 trees tagged in Queen Park Savannah to
examine air pollution effects
4) All Samsung smartphones sold 10,000 Samsung smartphones selected
worldwide in 2020 randomly to investigate software update issues.
Frequency
Frequency refers to the number of times a specific value or group of values appears within a given
dataset.
For example: In the dataset:
2, 3, 2, 3, 2, 2, 2, 5, 6, 5, 3, 2, 6, 8, 10, 3
Frequency of 2 = 6 (2 occurs 6 times)
Frequency of 3 = 4 (3 occurs 4 times)
Frequency of 5 = 2 (5 occurs 2 times)
Frequency of 6 = 2 (6 occurs 2 times)
Frequency of 8 = 1 (8 occurs 1 time)
Frequency of 10 =1 (10 occurs 1 time)
We note the sum of the frequencies is equal to the total number of data points. That is, 6 + 4 + 2
+ 2 + 1 + 1 = number of data points = 16.
To determine the frequency of a particular value or characteristic, we can use tally marks to make
counting easier. In using tally marks, 4 strokes are drawn vertically, and a fifth stroke is drawn
diagonally to make a group of five.
Frequency Tables
A frequency table is a table used to organize data to show how often a value appears in a given data set.
We will examine and construct frequency tables for “raw” (ungrouped) data
To ensure that no data value in the given data set is omitted, we find the sum of the frequencies to
ensure: sum of frequencies = total number of data points
The following example illustrates constructing a frequency table for “raw” data.
Example 8.1
Construct a frequency table for the number of sixes AR hit in 20 games where the number of sixes were:
9, 2, 3, 5, 10, 3, 3, 3, 2, 10, 7, 9, 10, 5, 3,8, 2, 5, 3, 10
Solution:
Exercise 8.1:
Cumulative Frequency
The cumulative frequency is the running total of frequencies in a frequency table. It shows the number of
points that fall below or within a particular class interval.
Statistical Diagrams
A statistical diagram is a diagram used to represent data, making it easier to identify trends, patterns,
compare different data sets, or make predictions. In our theory, we will examine, analyze and construct
the following statistical diagrams:
i) Pie Charts
ii) Bar Charts
i) Pie Charts
A Pie Chart is sometimes referred to as a circle graph or circle chart as it involves dividing a circle into
various sectors to represent data proportions. Pie charts show how data, divided into categories
contributes to the overall total, or whole. Each sector angle or area is directly proportional to the
magnitude of the information it is representing.
1) Collect and arrange data: Ensure that the data is placed into the desired categories; then find
the total value of each category.
2) Calculate angles: Divide each category total by the overall total and multiply this result by 360°
to find the number of degrees of the circle to assign to a particular category. We note the sum of
the number of degrees assigned to all categories must be equal to 360°.
3) Use a protractor to mark the angles for each sector and draw the circle to complete the
diagram.
4) Label each sector according to the category it represents.
The following examples illustrate the representation of statistical data as pie charts:
Example 8.2
In a school, 100 students were surveyed to determine the type of fruit they prefer. The responses are
given below:
Apples = 40 students
Oranges = 30 students
Bananas = 20 students
Grapes = 10 Students
Draw and label a pie chart to represent this information.
Solution:
Using a compass, we draw a circle and using a protractor, we measure and mark the angles: 144°, 108°,
72°, 36°. Finally, we label the sectors according to the fruit each sector represents to get:
Grapes
10%
Apples
Bananas
40%
20%
Oranges
30%
Example 8.3
The type of transportation used by students in a Secondary School is given in the table below:
Solution:
First, we get the overall (total) number of students: Total = 54 + 29 + 18 + 19 = 120 students
Calculating angles:
54 360° 54 100
Car: 120 × = 162° (120 × = 45%)
1 1
29 360° 29 100
Bus: 120 × = 87° (120 × = 24%)
1 1
18 360° 18 100
Maxi Taxi: 120 × = 54° (120 × = 15%)
1 1 Checking calculation for accuracy:
19 360° 19 100 162° + 87° + 54° + 57° = 360°
Taxi: 120 × = 57° (120 × = 16%)
1 1
Using a compass, we draw a circle and using a protractor, we measure and mark the angles 162°, 87°,
54° and 57°. Finally, we label the sectors corresponding to each angle to get:
Transportation Type Survey
Taxi, 16%
Car, 45%
Maxi Taxi, 15%
Bus, 24%
Example 8.4
Given the pie chart below represents the main ingredients used to make a “butter cake,” determine for
2454g of cake, how many grams (g) of each ingredient; butter, egg, sugar and flour are needed.
Solution:
From the given pie chart, the fraction of each ingredient (of the total; 360°) is given by:
36 1
Butter: 360 = 10
72 1
Sugar: 360 = 5
108 3
Flour: 360 = 10
144 2
Egg: =
360 5
Therefore, to find the amount of each ingredient needed to make 2454g of cake, we multiply each
fraction by 2454g.
This gives:
1 2454
Butter: 10 × 𝑔 = 245.4 𝑔
1
1 2454
Sugar: 5 × 𝑔 = 490.8 𝑔
1
3 2454
Flour: 10 × 𝑔 = 736.2 𝑔
1
2 2454
Egg: 5 × 𝑔 = 981.6 𝑔
1
Example 8.5
The following pie chart represents how money was spent at a school’s funfair.
Given that the total money spent at the funfair was $150,000, determine how much money was spent on:
a) Rides
b) Ice cream
c) Popcorn
d) Food
e) Face painting
Solution:
From the given pie chart, the fraction of each category out of the whole/ total (360°) is given by:
180 1
Rides: 360 = 2
55 11
Ice cream: 360 = 72
35 7
Popcorn: 360 = 72
70 7
Food: =
360 36
20 1
Face painting: 360 = 18
Therefore, to determine the amount of money spent on each feature offered at the funfair we multiply
each of the fractions above by the total among of money spent ($150,000)
1
Rides: × $150,000 = $75,000
2
11
Ice cream: 72 × $150,000 = $22,916.67
7
Popcorn: 72 × $150,000 = $14,583.33
7
Food: × $150,000 = $29,166.67
36
1
Face painting: 18 × $150,000 = $8,333.33
Exercise 8.2:
1) Construct a pie chart to represent the favorite flavor ice-cream of students in a class based on the
table below:
Ice Cream Flavor Number of Students
Chocolate 45
Cookies n Cream 30
Vanilla 15
Strawberry 30
Coconut 30
2) In a form 4 Mathematics class 20 students like football, 12 students like cricket, 10 students like
swimming, 8 students like volleyball, 5 students like badminton and 5 students like tennis.
Draw and label a pie chart to represent the information given.
a) The pie chart below shows the percentages of types of exercise preferred by 500 fitness
enthusiasts. Use the pie chart to find: How many people prefer weightlifting?
b) How many people prefer cycling?
c) How many people DO NOT prefer weightlifting?
d) What percentage of people DO NOT prefer jogging?
3) The following pie chart shows the various activities done by Arjun in a day (24 hours).
4) A pie chart is divided into 3 parts with the angles measuring as x, 4x and 5x respectively. Find the
value of x in degrees.
Bar Charts
A Bar Chart (or bar graph) is a chart that is used to represent data which is discrete (or categorical)
using rectangular bars (all of the same width) where the length or height of each rectangular bar is
proportional to the value it represents. The bars in a bar chart can be drawn horizontally (in this case the
chart is called a horizontal bar chart) or vertically (in this case the chart is called a vertical bar chart).
Apart from being used to represent data which is discrete it is useful for representing data that does not
need to be in any specific order when being presented.
The bars in a bar chart provide an easy way of comparing quantities in different categories. Bar graphs
consist of the x and y axes, title, scale, labels and horizontal or vertical bars representing the given data.
1) Organize the data into a table (this is not a necessary step but proves to be quite useful)
2) a) Draw and label the x-axis with categories and the y-axis with the scale (or frequencies), for a
vertical bar chart.
OR
b) Draw and label the x-axis with the scale (or frequencies) and the y-axis with the categories
for a horizontal bar chart.
3) Draw bars for each category where the height/ length of each bar is proportional to the value it
represents. Ensure bars are equal width and are evenly spaced in the graph.
4) Label the bar chart.
The following examples illustrate constructing bar charts from given data and analyzing given bar
charts.
Example 8.6
In a class of 40 students, 20 like cricket, 10 like football, 2 like swimming, 3 like badminton and the
remaining students like basketball. Draw:
Solution:
Example 8.7
The following bar chart shows the most bought fruits at Mike’s fruit stall on a particular day:
y
12
10
8
Number of Fruits
0
Mangoes Apples Cherries Plums Bananas Oranges Watermelons x
Types Of Fruits
Exercise 8.3:
1) The table below shows the number of coconut trees planted by a coconut farmer in Guyana over
a 6-year period.
Years Number of Coconut Trees
2006 150
2007 220
2008 350
2009 150
2010 300
2011 400
a) Represent the information on the table as a:
i) Vertical bar chart
ii) Horizontal bar chart
b) What percentage of coconut trees were planted before 2009?
c) What percentage of coconut trees were planted in 2010 – 2011?
2) The following table provides the number of sponge cakes a bakery produces in a particular day:
Days Numbers of Cakes
Monday 65
Tuesday 25
Wednesday 40
Thursday 55
Friday 100
Saturday 120
4) The number of people visiting a mall during lunchtime in Trinidad (on a particular day) is given
below:
Mall Number of People
C3 Centre 1050
SouthPark 450
West Mall 325
Grand Bazaar 500
Trincity 400
Gulf City 625
Laura gave her class of 30 students a mathematics test and Sandra also gave her class of 30 students a
mathematics test. During conversation, they attempted to find out which class performed better at the
test. It would be time consuming and somewhat difficult to compare 30 ‘raw’ scores between the two
classes to determine the better performing class. It would be significantly easier to compare one mark. A
measure of central tendency is a single value that allows for such comparisons to be easily made.
A measure of central tendency is a value which provides a “summary measure” of an entire data set.
It is a value that is representative of the entire given set of data, allowing for easy interpretation and
comparison of different sets of data.
There are three main measures of central tendency: namely the:
• Mean
• Median
• Mode
We will examine each of these for “raw” ungrouped data sets and for ungrouped and grouped data
frequency distributions.
Example 8.8
Find the mean score of 9 students representing Barbados in a mathematics competition given the following
scores:
100 81 64 55 88 93 98 95 90
Solution:
∑𝑥 100+81+64+55+88+93+98+95+90
Mean = =
𝑛 9
764
Mean = = 84.89 (to 2 d.p)
9
Example 8.9
25 30 36 45 51 63 75 100
Solution:
∑𝑥 25+30+36+45+51+63+75+100
Mean = =
𝑛 8
425
Mean = = 53.13 (to 2 d.p)
9
b) Median (Middle Value)
The median is the middle value when the given data set is arranged either in ascending order
(smallest value to largest value) or descending order (largest value to smallest value). We
note, if the data set consists of an odd number of values, the median is the middle value. If the data
set consists of an even number of values the median is the average of the two middle values.
Example 8.10
a) 9, 4, 7, 15, 12
b) 6, 2, 8, 4, 10, 12
Solution:
c) Mode
The mode is the value that occurs most often in a given data set.
Example 8.11
Solution:
1) The following are marks David scored in exams (out of 10) in a month.
4 5 5 5 4 3 2 1 4 5
a) State the mode of his marks (the modal mark)
b) Find the mean mark.
2) Peter rolled a 6-sided dice ten times. The following are his scores:
3 2 4 6 3 3 4 2 5 4
a) Find the mean score.
b) State the modal score.
c) Find the median score.
3) The weights of 8 people (in kg) are given below:
72 63 97 65 90 65 86 70
a) Find the mean weight.
b) State the modal weight.
c) Find the median weight.
4) The ages of 11 students going on a field trip are given below:
17 16 14 17 13 15 14 16 17 14 17
a) Find the mean age.
b) State the modal age.
c) Find the median.
5) Mr. Smith kept a record of the number of times each student is late for school in term 1. The
following are his results.
0 0 0 8 4 5 3 2 1 5
a) Find the median.
b) Find the mean.
c) State the mode.
6) The following are marks scored by 6 girls and 4 boys in a class:
Girls- 5 3 10 2 7 3
Boys- 2 5 9 3
a) Find the mode of the 10 marks.
b) Find the median mark of the boys.
c) Find the mean mark of the girls.
d) Find the mean mark of the 10 students.
a) Mean
The mean of an ungrouped frequency distribution is found by using the formula below:
∑𝒇×𝒙 • Σ means “the sum of”
Mean = • 𝑥 represents the given (individual) data values.
𝜮𝒇
• f is the frequency (given)
To find the mean of a given frequency distribution (table), it is helpful to extend the table by adding a
column for “f × x.” This allows for a more straightforward process of finding the mean. The following
example illustrates the method:
Example 8.12
Find the mean age of 15 students in a class given the following distribution of ages:
**∑(𝒇 × 𝒙) = 𝟑𝟔 + 𝟓𝟐 + 𝟗𝟖 + 𝟏𝟓 = 𝟐𝟎𝟏
Substituting “Σ𝑓 ” and “∑(𝑓 × 𝑥 )” into the formula for mean gives:
∑(𝑓×𝑥) 201
Mean = = = 13.4 to 1 d.p
𝛴𝑓 15
b) Median
The median or middle value can be found from an ungrouped frequency distribution. To achieve
this, we first ensure the data values are in order (increasing or decreasing).
We note the following formula:
𝒏+𝟏
Position of the median = where n is equal to the total frequency. To find the median we
𝟐
follow the steps below:
1) Find the cumulative frequency of the given ungrouped frequency distribution.
2) Find the position of the median using the formula.
3) Use the cumulative frequency together with the result from step 2 to find the median value.
Example 8.13
The frequency distribution below shows the number of cars owned by 13 families:
c) Mode
The mode of an ungrouped frequency distribution is simply the data value that has the highest
frequency.
Example 8.15
Solution:
∑𝒇×𝒙
Mean = 𝜮𝒇 gives:
398
Mean = = 23.41 (to 2 decimal places)
17
Therefore, mean temperature = 23.41℃
b) We draw a table as follows to assist with the calculation or we can add a column for cumulative
frequency to our table from part (a):
Temperature Frequency Cumulative Frequency
21 1 1
22 2 3
23 5 8
24 7 15
25 2 17
We note the given data values are in ascending order.
𝑛+1
Position of the median = value where n = 17
2
17+1
Position of the median = = 9th data value
2
By examining the cumulative frequency, we can deduce that the value is in the fourth row. The row
contains the 9th, 10th, 11th, 12th, 13th, 14th and 15th data values; Therefore, the median is 24℃
c) The modal temperature is 24℃, since the value 24℃ corresponds to the highest frequency (7).
Exercise 8.5:
Mean
In simple terms, when a data set contains a value or values which are significantly different from all other
values in the data set; such different values are referred to as outliers. For example, if in a class of 30
students, 28 students scored marks between 70 and 80, and one student scored 98 and one scored 04, the
marks 04 and 98 are referred to as outliers as they are significantly different from the general set of data
values (70 – 80). We use the mean as a measure of central tendency when the data set has minimal or
no outliers, that is the data is symmetrically distributed.
Example: for values such as: 85, 90, 95, 92, 88 the mean is a good measure of central tendency. Mean =
85+90+92+95+88 450
= = 90
5 5
Median
We can see from the above that the median is the better measure of central tendency for this case
(where an outlier is present) as it more accurately represents the age category of gym users at 6am
(majority of persons using the gym are in their 20’s). The mean is not a good measure of central
tendency when outlier(s) are present in a data set.
Therefore, we use the median as a measure of central tendency when any given data set contains outliers
(the median is less affected by outliers than the mean).
Mode
The mode is the preferred measure of central tendency for qualitative data sets. It is most useful
when analyzing data categorically in nature such as favorite colors, model of car, or cricket team.
Exercise 8.6:
Determine the most appropriate measure of central tendency (mean, median or mode) to use for each of
the following:
1) A company wants to determine the most frequently sold product in their inventory.
2) A principal wants to find a representative age for all students in form 5 given that there are no
outliers.
3) The number of pets owned by 7 students are as follows: 0, 1, 1, 2, 2, 8, 10. What measure best
represents the data?
4) A clothing store wishes to identify the most common shirt size purchased (S, M, L, XS, XL) in
their last sale event.
5) A scientist wishes to analyze the temperature each day at 12:00pm for a month.
6) A store tracks daily revenue, including one day of exceptionally high sales due to a national
holiday. Which measure should the manager use to report usual daily sales?
7) A car dealership wishes to determine the most popular car model sold in 2023.
8) A teacher wants to report the grade for an exam in which one person got zero for being absent.
Scales of Measurement
a) Nominal Scale of Measurement
A nominal scale of measurement is used for qualitative data. If values are used to describe the
data where no numerical meaning is assigned to the values. This means in using a nominal scale
of measurement, the data can be assigned values, but these cannot be added, subtracted,
multiplied, or divided.
Examples of nominal scales are given below:
1) What is your gender? 2) What is your hair color?
• 1 – Male • 1 – Black
• 2 – Female • 2 – Brown
• 3 – Other • 3 – Blonde
• 4 – Gray
• 5 – Other
Here the numbers serve as “tags” or “labels” only.
b) Ordinal Scale of Measurement
The ordinal scale of measurement groups the data into order (or ranks the data values). The
ordinal scale contains the property of the nominal scale as well (where data is classified and
“tagged” or “labelled”)
Examples of ordinal scales are given below:
Example
The study of customer satisfaction with a company’s product or service with a scale of 1 – 5 where:
#1 is very happy
#2 is satisfactory
#3 is neutral
#4 is unhappy
#5 is extremely dissatisfied
Example: Example:
Movie ratings where: Place in class (for student performance in exams):
1 star: Poor 1st place
2 stars: Fair 2nd place
3 stars: Good 3rd place, and so on.
4 stars: Very Good
5 stars: Excellent
Here we see that ordinal measurement scales display the order or rating of the variables. It however,
does not give any numerical value to the data; therefore, it is used for qualitative data in a similar way to
the nominal measurement scale. (we cannot add, subtract, multiply or divide numbers)
Exercise 8.7:
Classify each as nominal, interval ordinal or ratio:
1) The preferred mode of transport of your class
2) The education level of everyone in your school
3) The current temperature in degrees Celsius
4) The number of hours of sleep per night you get on average.
5) Your weight in kilograms.
6) The distance from your home to school in kilometers
7) On a scale of 1 – 10, how much do you enjoy algebra?
8) How often do you exercise (rarely, sometimes, often, everyday)
9) Your top three favorite colors
10) How satisfied are you with your internet service provider (very dissatisfied to very satisfied)
Example 8.17
Find the range of the following data sets:
a) 50, 60, 40, 55, 45, 75
b) 14, 7, 2, 9, 6, 25
Solution:
a) Range = highest value – lowest value
Range = 75 – 40 = 35
b) Range = highest value – lowest value
Range = 25 – 2 = 23
For an ungrouped frequency distribution, the frequency is not needed to find the range as the range is
calculated simply by identifying the maximum (largest) and minimum (smallest) values in the data
set and using the same formula as “raw” data. That is:
Example 8.18
Find the range of each of the following distribution of scores:
a) Score Frequency (f) b) Score Frequency (f)
15 1 4 1
20 3 5 3
25 8 7 8
30 20 8 9
Solution: 10 6
a) Range = maximum score – minimum score
Range = 30 – 15 = 15
b) Range = maximum score - minimum score
Range = 10 – 4 = 6
Exercise 8.8:
2)
a) Score Frequency (f) b) Score Frequency (f)
9 4 20 5
19 3 40 4
60 3
29 5
80 7
39 11
The range as a measure of dispersion has many limitations. As it is calculated from only two values (the
highest and lowest), if any of these are outliers, the range will be unreliable as a measure of dispersion.
Additionally, the distribution of other data points in the data set are not considered, hence a complete
‘picture’ of the spread is not provided by the range.
Quartiles
Quartiles are three values that divide a given data set into four equal parts with each part representing a
quarter of the data. Quartiles give a better understanding of the dispersion (spread) and central tendency
of data. The three quartiles are given below:
The Interquartile Range (IQR) is a value which gives a measure of how spread out the middle (50%)
of a given data set is. It gives the range of the middle (50%) of the data, identifying the spread and
variability of the data in the central portion, without being affected by outliers.
The Semi-Interquartile Range (SIQR) is a value which also gives a measure of the spread of data in
the central portion of a data set. However, the SIQR serves as a more focused measure of spread around
the median (𝑄2 ) than the Interquartile Range (IQR). (The SIQR is more compact than the IQR providing
a clearer understanding of dispersion around the middle value 𝑄2 ) The IQR and SIQR are found from the
formulas below:
𝟏
Semi-Quartile Range = 𝟐 (Third quartile – First quartile)
𝟏
SIQR = 𝟐 (𝑸𝟑 – 𝑸𝟏 )
Before finding 𝑄1 , 𝑄2 , 𝑄3 , IQR or SIQR, the data set must be ordered (in ascending order or
descending order).
Example 8.19
a) 𝑄1 , 𝑄2 , 𝑄3
b) IQR and SIQR
1)
6 47 49 15 42 41 7 39 43 40 36
2)
3 12 8 5 9 16 14 20 17 22
Solution:
1) The given data set is not in any order. In ascending order, the set is:
6 7 15 36 39 40 41 42 43 47 49
Finding the median (𝑄2 ) first:
6 7 15 36 39 40 41 42 43 47 49
Median (𝑄2 ) = 40
We note the median separates the entire data set into halves.
6 7 15 36 39
𝑄1 = 15
The upper half is: 41, 42, 43, 47, 49
Finding the middle value of the upper half (𝑄3 ):
41 42 43 47 49
𝑄3 = 43
Therefore, 𝑄1 = 15, 𝑄2 = 40, 𝑄3 = 43
b) IQR = 𝑄3 – 𝑄2 = 43 – 15 = 28
1 1
SIQR = 2 (𝑄3 – 𝑄1 ) = 2 (28) = 14
b) IQR = 𝑄3 – 𝑄1 = 17 – 8 = 9
1 1
SIQR = 2 (𝑄3 – 𝑄1 ) = 2 (9) = 4.5
• A smaller standard deviation means that data points are closer to the mean.
• A larger standard deviation means that data points are spread out over a wide range.
The following examples illustrate how standard deviation is used to compare two sets of data:
Example 8.20
Given data set A: 10, 15, 20, 25, 30 has a standard deviation of 7.07 and data set B: 5, 20, 35, 50, 65 has
a standard deviation of 23.36. determine which data set is spread closer about the mean.
Solution:
A smaller standard deviation means that data points are closer to the mean. Since the standard deviation
of data set A is smaller than the standard deviation of data set B, the data points in set A are spread
closer to the mean. We say that the values in data set A are more consistent than the data values in set B.
Exercise 8.9:
3) The two data sets below provide the ages of employees in two small companies.
Company 1: 22, 24, 26, 28, 30 (standard deviation = 2.83)
Company 2: 20, 30, 40, 50, 60 (standard deviation = 15.81)
Which company has a more diverse age range of employees working?
Probability
Probability is a topic within Mathematics that deals with predicting how likely events are to happen. It
helps us understand and make decisions in situations involving uncertain outcomes/results, enabling us
to assess any risks involved and make any necessary predictions. Probability plays a significant role in
various real-life situations and decision-making processes such as in:
1) Weather Forecasting:
Meteorologists use probability models to predict the likelihood of rain, sunshine, snow and to
track the path of storms and hurricanes.
2) Medical Fields:
Doctors use probabilities to determine the effectiveness of treatments and medications.
3) Gaming/ Gambling:
Casinos and lottery systems use probability to design games of chance in such a way to ensure
that the “house always wins”.
4) Sports:
Probability models help managers/ coaches make informed decisions regarding player selections.
Also, the outcomes of matches can in some cases, be predicted by probability structures.
5) Traffic / Navigation
Apps like Google Maps use probabilities to estimate traffic and suggest fastest routes.
1) Experiment: this is a process with uncertain outcomes/results (e.g. rolling a die, flipping a coin).
The process of doing an experiment is referred to as a Trial.
2) Outcome: this is a possible result from an experiment (e.g. getting a Head (H) when flipping a
coin)
3) Event: this is a collection of one or more outcomes (e.g. rolling an even number with a dice/die)
4) Sample Space: this is the set of all possible outcomes.
5) Impossible Event: an event is impossible if it will never occur. (e.g. scoring a negative number
of runs in a cricket match).
The probability of an event which is impossible is 0.
6) Certain Event: an event that will definitely occur (or happen) (e.g. getting a number between 1
and 6 inclusive, is certain to occur if a die is rolled).
The probability of an event which is certain to occur is 1.
7) Notation: If E is an event, then P(E) represents the probability of the event E occuring.
We note for any event E: 𝟎 ≤ 𝑷(𝑬) ≤ 𝟏 (probability is always a value between 0 and 1)
1) Coin: A coin has two sides, which we refer to as heads (H) and tails (T). The sample space S is
given by: S = {Heads, Tails}
2) Die: A standard six-sided die (sometimes referred to as ‘a dice’ has six possible outcomes. The
sample space S is given by: S = {1,2,3,4,5,6}
3) Standard Deck of Playing Cards: A standard deck of playing cards has 52 cards with four suits:
a) Hearts b) Diamonds c) Clubs d) Spades
Each suit has 13 cards: 2, 3, 4, 5, 6, 7, 8, 9, 10 and Ace, Jack, Queen and King. It is important to
note that Jack, Queen and King are referred to as ‘face’ cards.