0% found this document useful (0 votes)
6 views46 pages

Data Management in Statistics Module

Module 3 focuses on the use of mathematics as a tool for data management, covering concepts from descriptive and inferential statistics, including measures of central tendency, variation, and linear regression. The module aims to enhance students' ability to process and manage numerical data effectively, particularly in the context of the ongoing pandemic. Students will engage in activities to apply statistical methods to real-world data, fostering skills necessary for informed decision-making.

Uploaded by

kwonnarraaa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views46 pages

Data Management in Statistics Module

Module 3 focuses on the use of mathematics as a tool for data management, covering concepts from descriptive and inferential statistics, including measures of central tendency, variation, and linear regression. The module aims to enhance students' ability to process and manage numerical data effectively, particularly in the context of the ongoing pandemic. Students will engage in activities to apply statistical methods to real-world data, fostering skills necessary for informed decision-making.

Uploaded by

kwonnarraaa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 3

Mathematics as a Tool

Data Management

Mathematics in the Modern World


Mathematics as a Tool


Overview

Hello students! How are you keeping up with our online classes? I hope that you are all doing fine. If not,
please feel free to ask for our help  In this module, we are going to review some concepts that you have
learned from the Statistics courses you took in your basic education. Specifically, this will cover the
following topics: Descriptive Statistics, Inferential Statistics and Planning or Conducting an Experiment
or Study. We will briefly discuss the measures of central tendency, measures of variation, normal
distributions, and linear regression and correlation.

This module is designed for you to finish in three weeks. Time management and self-motivation is the
key to accomplishing all designated tasks/activities of this module, to achieve every objectives and
outcomes of this course.

I hope you will be able to appreciate the importance of managing data as it is relevant to the pandemic
that we are struggling with right now. It can help us fight through this battle and all other disasters which
could strike us in the future. Stay safe everyone!

Learning Outcomes

After completing the study of this module, you should be able to:

Use variety and appropriate statistical tools to process and manage numerical data

Apply the methods of linear regression and correlation to predict the value of a variable given
certain conditions

Make use of statistical data in drawing important decisions

Initial Activity (Accessing Prior Knowledge)

I want you to access the DOH COVID-19 Tracker through the link
[Link]/covid19tracker and browse to the webpage. If you can’t access this page,
refer to a screenshot below of a portion the webpage. Of all the data being shown in this
platform, can you identify which data show average values? Can you interpret these? Do the graphs
make sense to you? Do they display a trend which could help us predict its behavior for in the next 7-14
days or more?

1
Mathematics as a Tool


Review of Descriptive Statistics

Data management is an administrative process that includes acquiring, validating, storing,


protecting, and processing required data to ensure the accessibility, reliability, and timeliness of the data
for its users. Data are individual pieces of factual information recorded and used for the purpose of
analysis. It is the raw information from which statistics are created. Statistics are the results of data
analysis - its interpretation and presentation. Statistics has two branches. The one that involves
collection, organization, summarization and presentation is called Descriptive Statistics. The branch of
statistics which involves interpretation and drawing conclusions is called Inferential Statistics.

Typically, there are two general types of statistic that are used to describe data: measures of
central tendency and measures of variance. A measure of central tendency, which is sometimes
referred to as “averages”, describes a set of data by identifying its central position. In this module, will
consider three types of averages: mean, median and mode. In the first section of discussion, we will take
a look at these measures specifically on how to calculate them.

2
Mathematics as a Tool


Mean

The mean, also called the arithmetic mean, is the most frequently used measure of central tendency.

The mean is obtained by getting the sum of all values and divide it
Did you know?
by the number of values in the data set. So if we have 𝑛 values in
the data set and they have values 𝑥1 , 𝑥2 , 𝑥3 , … , 𝑥𝑛 , the mean usually
denoted by 𝑥̅ (read as x bar) is: Seaweed can grow up to 12
The summation notation “Σ” inches per day averagely.
𝑥1 , 𝑥2 , 𝑥3 , … , 𝑥𝑛 ∑ 𝑥
𝑥̅ = = denotes the sum of all
𝑛 𝑛 values in a given data set.
An average person eats almost
1500 pounds of food a year.
To be precise, 𝑥̅ is known as sample mean. In some situations, data
are collected from small portions of a large group in order to
The human eye blinks an
represent the determined information about the group. The whole average of 4,200,000 times a
group under consideration is called population, while any subset of year.
the population is called a sample.
The most children born to one
woman as 69, she was a
So if we want to compute for the value of the population mean, peasant who lived a 40 year
denoted by µ, we use the formula: life, in which she had 16 twins, 7
∑𝑥 triplets, and 4 quadruplets.
𝜇=
𝑁
*Source: [Link]
where N is the size of the population.

Sometimes a data set may contain a few very small or a few very large values; such values are called
outliers or extreme values. A major shortcoming of the mean as a measure of central tendency is that it
is very sensitive to outliers.

Example

1.) In a grocery store, price (₱) of bath soap products of different brands are the following:
35 46 40 31 39
Find the mean of the prices of these five brands of bath soap.

Solution

The five brands are all of the bath soap brands in the grocery store (population size N = 5).
We use μ to represent the mean.
∑ 𝑥 35 + 46 + 40 + 31 + 39
𝜇= =
𝑁 5
Thus, the mean of the prices of these five brands of bath soap is ₱38.20.

3
Mathematics as a Tool


Example (cont.)

2.) The following are the ages (in years) of eight the 67 employees of a small company:
25 32 61 42 39 48 56 29
Find the mean age of the employees.

Solution
Because the given data set only includes eight of the 67 employees of the company, it
represents the population. Hence, N = 8. The population mean is
∑𝑥 25 + 32 + 61 + 42 + 39 + 48 + 56 + 29 = 332 332
𝑥̅ = = = = 41.5 ≈ 42
𝑛 8 8
Thus, the mean age of all eight employees of this company is 42 years.

3.) The following are the savings (₱) of five siblings of a certain family (arranged from the
youngest to the eldest sibling):

2500 3000 2800 2900 35900


Notice that the eldest sibling’s savings (35900) is very large compared to the others. Hence,
this is an outlier. Show how this affects the value of the mean.

Solution
If we do not include the data of the eldest sibling, then the mean is
2500 + 3000 + 2800 + 2900
𝑀𝑒𝑎𝑛 = = ₱2800
4
Now, if we include it to our data, then the mean is
2500 + 3000 + 2800 + 2900 + 35900
𝑀𝑒𝑎𝑛 = = ₱9420
5
Thus, including the savings of the eldest sibling causes more than three times increase in the
value of the mean, which changes from ₱2800 to ₱9420.

Learning check

Activity 1
Compute the mean of the following set of data values
a.) Anika went to the supermarket to buy some packed potatoes, with her uncertainty of
estimating the difference between the sizes of the packed potatoes, she looks at the
price tags and finds the following prices (₱)
125 142 132.5 201.25 160 172.75
b.) Archie collects stamps and glues them to his notebook. The following shows the
tally of his monthly collection from June to November
16 25 18 30 29 23

c.) The following shows the grades of a student in Physics, Math, English, Biology,
Chemistry and Filipino, respectively. 4
86 89 91 93 88 90
Mathematics as a Tool


Median
Another important measure of central tendency is the median. It is defined as the value of the middle
term in a data set that has been ranked in increasing order.

As is obvious from the definition of the median, it divides a ranked data set into two equal parts. The
calculation of the median consists of the following two steps:
1. Rank the data set in increasing order.
2. Find the middle term. The value of this term is the median.

Note that if the number of observations in a data set is odd, then the median is given by the value of the
middle term in the ranked data. However, if the number of observations is even, then the median is
given by the average of the values of the two middle terms.

Example

1.) The following data give the age of students in a ballet class. Find the median age.
12 15 17 11 13 16 19
Solution

First, we rank the given data in an increasing order. Since there are seven values in this data
set as it was arranged accordingly, the fourth term is the middle term. Hence,
11 12 13 15 16 17 19
Median
The median age of the students in a ballet class is 15.

2.) The following are the number of smart phone users in 10 households in a certain
barangay:
7 9 4 10 5 6 5 8 12 5

Solution

Since there are ten values in this data set as it was arranged accordingly, the fifth term and
the sixth term are the middle term:
4 5 5 5 6 7 8 9 10 12
Median
To get the median, we get the average of the two middle terms
6+7
𝑀𝑒𝑑𝑖𝑎𝑛 = = 6.5
2
The median number of smart phone users among the 10 households in the barangay is 6.5.

5
Mathematics as a Tool


Learning check

Activity 2
Compute the median of the following set of data
a.) 3, 4, 7, 11, 12, 12, 15, 16
b.) -8, -5, -12, -1, 4, 7, 11
c.) 6, 4, 8.5, 9, 11, 8.25, 6.5, 8.75
d.) The following are the ages of the 5 siblings of the Reyes family. Find the age of the
middle child.
25 21 23 18 27
e.) In a supermarket, the following table shows the number of tissue roll of different
brands sold in a certain week (Brand 1 to Brand 8, respectively)
25 12 5 17 20 13 9 21

Mode

Another measure of central tendency is concerned with the count of each data value which we can refer
to as the mode. Moreover, it is defined as the value that occurs with the highest frequency in a data set.

A data set may have none or may have more than one mode, whereas it will have only one mean
and only one median. For instance, a data set with each value occurring only once has no mode.

Example

1.) The following gives the general weighted average of the top 10 students with the highest
grades in a certain class
95 93.5 91 89 92 94 91 93 90 91.5
Solve for the mode.

Solution
In this data set, all values appeared only once except for 91 which appeared twice. Because
91 has the highest frequency, hence
𝑴𝒐𝒅𝒆 = 91

2.) The following are the number of ball pens that each of the eight students owned in a
certain group of friends
1 3 2 0 2 1 3 5
Solve for the mode.

Solution
There are three data values with the highest frequency in this data set which are 1, 2, and 3.
Therefore,
𝑴𝒐𝒅𝒆 = 1, 2, 3

6
Mathematics as a Tool


Example

3.) A statistician conducted a survey in a certain barangay to obtain the profile of households
wherein the statistician gathered information about the appliances owned by each household.
The table below shows the summary of the number of appliances of by the households
Name of Appliance No. of Items
Television 35
Refrigerator 26
Microwave oven 10
Kitchen Stove 36
Washing Machine 15
Clothing Iron 30
Solve for the mode.

Solution
In this data set, all values have the same frequency. Hence, there is no mode.

Learning check

Activity 3
Solve for the mode of the given data set in the following items

a.) 13, 15, 16, 12, 11, 17, 13, 15, 10, 14
b.) 7.5, 8.5, 6.5, 3.5, 5, 5.5, 5, 9.5, 10.5
c.) -11, -4, -5, |-1|, 0, 1, 5, 4, 11
d.) The following shows the number of hops in a skipping rope that 10 students can
complete in one minute
60 35 55 78 56 55 69 48 58 57
e.) The following table shows an estimate of the number of strawberries that a farmer
was able to harvest per day in a certain week
Day No. of Strawberries Harvested
Sunday 560
Monday 980
Tuesday 760
Wednesday 670
Thursday 950
Friday 980
Saturday 1000
7
Mathematics as a Tool


Mean for Grouped Data

If the data are given in the form of a frequency table, we no longer know the values of individual
observations. In such cases, we cannot obtain the sum of individual values. We find an approximation
for the sum of these values using the formulas shown below

∑ 𝑚𝑓
𝜇=
𝑛

∑ 𝑚𝑓
𝑥̅ =
𝑁

where m is the midpoint and f is the frequency of a class

Example

The table below gives the frequency distribution of the daily commuting times (in minutes)
from home to work for all 40 employees of a company.
𝑫𝒂𝒊𝒍𝒚 𝑪𝒐𝒎𝒎𝒖𝒕𝒊𝒏𝒈 𝑻𝒊𝒎𝒆 𝑵𝒖𝒎𝒃𝒆𝒓 𝒐𝒇 𝑬𝒎𝒑𝒍𝒐𝒚𝒆𝒆𝒔
1-9 4
10-18 12
19-27 9
28-36 7
37-45 5
46-54 3

Let x denote the daily commuting times (in minutes) from home to work of the 40 employees
of a company, and f denote the frequency. The values of m and mf are calculated in the table
below
𝑥 𝑓 𝑚 𝑚𝑓
1-9 4 5 20
10-18 12 14 168
19-27 9 23 207
28-36 7 32 224
37-45 5 41 205
46-54 3 50 150
N = 40 ∑ 𝒎𝒇 = 𝟗𝟕𝟒

974
𝜇== 𝟐𝟒. 𝟑𝟓
40
Hence, using the mean daily commuting times of the 40 employees of the company is
24.35 minutes.

8
Mathematics as a Tool


Learning check

Activity 4

The table below shows the frequency distribution of the number of customers received in a
cafe each day during the past 24 days.
Number of Orders Number of Days
10-12 12
13-15 20
16-18 14
19-21 25
22-24 15

Median of Grouped Data

To compute for the median of a grouped set of data, the median class shall be identified by locating the
𝑛
𝑡ℎ data at the >cf column. Then, we will use the formula below:
2

𝑛
− 𝑐𝑓𝐵
𝑀𝑒𝑑𝑖𝑎𝑛 = 𝐿 + 2 ×𝑤
𝑓𝑚

where L is the lower class boundary of the median class, 𝑐𝑓𝐵 is the cumulative frequency of the class
before the median class, 𝑓𝑚 is the frequency of the median class, and w is the class width.

9
Mathematics as a Tool


Example

The table below gives the frequency distribution of the number of hours spent in studying by
50 students before a quarter exam.

Hours Number of Students


0-3 9
4-7 13
8-11 10
12-15 11
16-19 7

Solution

Let x denote the number of hours spent in studying by 50 students before a quarter exam,
and f denote the frequency. The values of 𝑚 − 𝑥̅ and (𝑚 − 𝑥̅ )2 , and 𝑓(𝑚 − 𝑥̅ )2 are
calculated in the table below.
x f >cf CB
0-3 9 9 -0.5-3.5
4-7 13 22 3.5-7.5 Median
8-11 10 32 7.5-11.5 class
12-15 11 43 11.5-15.5
16-19 7 50 15.5-19.5

𝑛 50
First, we have to identify the median class. In this example, = = 25. Hence, the 25th data
2 2
can be found at the third class. (Since the second class contains only 22 data, and the third
class already contains 32 data which includes the 25𝑡ℎ one)

With the third class as the median class, = 7.5 , 𝑐𝑓𝐵 = 22, 𝑓𝑚 = 10, and 𝑤 = 4

25 − 22
𝑀𝑒𝑑𝑖𝑎𝑛 = 7.5 + × 4 = 𝟖. 𝟕
10

Therefore, the median of the number of hours spent in studying by 50 students before
a quarter exam is 8.7 hours.

10
Mathematics as a Tool


Learning check

Activity 5
Calculate the median score in a Math quiz as shown in the table below:

Score Number of Students


12-15 4
16-19 6
20-23 4
24-27 5
28-31 7
32-35 4

Mode of Grouped Data

To compute for the mode of a grouped set of data, the modal class shall be identified by locating the
class with the highest frequency at the f column. If in case there is more than 1 class with the highest
frequency, then we shall compute for more than 1 mode. Then, use the formula below:

𝑓𝑚𝑜 − 𝑓𝑚𝑜−1
𝑀𝑜𝑑𝑒 = 𝐿 + ×𝑤
(𝑓𝑚𝑜 − 𝑓𝑚𝑜−1 ) + (𝑓𝑚𝑜 − 𝑓𝑚𝑜+1 )

Where L is the lower boundary of the modal class, 𝑓𝑚𝑜 is the frequency of the modal class, 𝑓𝑚𝑜−1 is the
frequency of the class before the modal class, and 𝑓𝑚𝑜+1 is the frequency of the class afer the modal
class and w is the class width.

11
Mathematics as a Tool


Example
Calculate the mode of the number of hours spent in studying by 50 students before a quarter
exam.
Hours Number of Students
0-3 9
4-7 13
8-11 10
12-15 11
16-19 7

Solution
Let x denote the number of hours spent in studying by 50 students before a quarter exam,
and f denote the frequency. The values of f and 𝑚𝑓 are calculated in the table below.
x f >cf CB
0-3 9 9 -0.5-3.5
4-7 13 22 3.5-7.5 Modal
8-11 10 32 7.5-11.5 class
12-15 11 43 11.5-15.5
16-19 7 50 15.5-19.5

In this frequency distribution table, the modal class is the second class because it has the
highest frequency which is 13. So, L = 3.5, 𝑓𝑚𝑜 = 13, 𝑓𝑚𝑜−1 = 9, 𝑓𝑚𝑜+1 = 10, 𝑤 = 4

13 − 9
𝑀𝑜𝑑𝑒 = 3.5 + × 4 = 𝟓. 𝟕𝟗
(13 − 9) + (13 − 10)

Hence, the mode of the number of hours spent in studying by 50 students before a
quarter exam is 𝟓. 𝟕𝟗.

Learning check
Activity 6
Calculate the modal score in a Math quiz as shown in the table below:

Score Number of Students


12-15 4
16-19 6
20-23 4
24-27 5
28-31 7
32-35 4

12
Mathematics as a Tool


Relationships among the Mean, Median, and Mode

A histogram or a frequency distribution curve can assume shapes which are symmetric and skewed. The
shape of a frequency distribution curve can be identified or described using the knowledge of the values
of the mean, median, and mode of a certain set of data.

1. If The values of the mean, median, and mode are identical, and they lie at the center of the
distribution, then, a symmetric histogram and frequency distribution curve has one peak.

2. If the value of the mean is the largest, that of the mode is the smallest, and the value of the median
lies between these two, then a histogram and a frequency distribution curve is skewed to the right
(Notice that the mode always occurs at the peak point.)

The value of the mean is the largest in this case because it is sensitive to outliers that occur in the right
tail. These outliers pull the mean to the right.

3.) If the value of the mean is the smallest and that of the mode is the largest, with the value of the
median lying between these two, then, histogram and a frequency distribution curve are skewed to the
left.

In this case, the outliers in the left tail pull the mean to the left.

13
Mathematics as a Tool


Measures of Variation

In examining averages, some characteristics of a set of data may not be evident. For instance, the given
example from the discussion of mean about the savings of siblings in a family, the given shows that the
mean of the savings are equal despite of having an inconsistent value which is quite far from the others.
This example shows that the average do not reflect the dispersion, spread or scatter of data.

In this section, we will be introducing the four measures of variability, namely, thee range, average
deviation, variance and standard deviation, to tell us how much the data tend to disperse, spread or
scatter.

Range

The range of a set of data values is obtained by getting the difference between the largest data value
and the lowest data value.

𝑅𝑎𝑛𝑔𝑒 = 𝐿𝑎𝑟𝑔𝑒𝑠𝑡 𝐷𝑎𝑡𝑎 − 𝑆𝑚𝑎𝑙𝑙𝑒𝑠𝑡 𝐷𝑎𝑡𝑎

Example

1.) In a cafeteria, there are three vendo machines which are set to dispense 10 ounces of
coffee to a cup. For the first five trials, the given table below shows the volume (in oz.)
of coffee dispensed by the machines to each cup. Find the range for each machine.
Machine 1: Range = 10.07 – 5.85 = 4.22 oz.
Machine 1 Machine 2 Machine 3 Machine 2: Range = 8.03 – 7.95 = 0.08 oz.
11.52 10.01 8.02 Machine 3: Range = 10.00 – 6.00 = 4.00 oz.
8.41 9.99 12.03
This indicates that Machine 1 varies the most
12.07 9.95 7.94
7.85 10.03 11.99
in volume of dispensing soft drinks, which
10.15 10.02 10 shows its inconsistency unlike for Machine 2,
the volume dispersed is quite consistent.
2.) Calculate the range of the following set of data values:
a.) 10, 2, 5, 6, 7, 3, 4 Range= 10 – 2 = 8
b.) 99, 45, 23, 67, 45, 91, 82, 78, 62, 51 Range = 99 – 23 =76
c.) 15, 14, 17, 13, 11, 19, 25, 22, 21 Range = 25 –13 = 12

Learning check
Activity 7
Compute the range of the following data values
a.) 1, 2, 5, 7, 8, 19, 22
b.) 3, 4, 7, 11, 12, 12, 15, 16
c.) -8, -5, -12, -1, 4, 7, 11
d.) 2, 4, 6, 7, 9, -5, -7
e.) 12, 16, 12, 16, 14, 12, 11 14
Mathematics as a Tool


Average Deviation for Ungrouped Data

A measure of variation that takes into consideration of the deviations of the individual data scores from
an average is generally considered more accurate and reliable than those determined by only the single
value like the range. One such measure is the average deviation.

In calculating the average deviation of a data set of values, one must obtain the mean of the data set
first. The average deviation can then be obtained by using the equation:

∑ |𝑥 − 𝜇|
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 =
𝑁

Average Deviation for Grouped Data

In calculating the average deviation for a set of grouped data, we have to use the formula:

∑ 𝑓|𝑚 − 𝜇|
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 =
𝑁
where m is the midpoint of a class.

Example

1.) Solve the average deviation of the scores: 17, 15, 18, 19, 15, 13, 14, 16

Score |𝒙 − 𝝁|
17 1.13
15 0.88
18 2.13
19 3.13
15 0.88
13 2.88
14 1.88
16 0.13
Mean = 15.88 Sum = 13.04

13.04
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = = 1.63
8

Hence, the average deviation of the scores is 1.63.

15
Mathematics as a Tool


Example
2.) Solve the average deviation of the scores of students in a Math Quarter exam with the
given data from the frequency distribution below
Score Number of Students
50-59 5
60-69 3
70-79 5
80-89 10
90-99 12

Solution
Let x denote the scores of the scores of students in a Math Quarter exam. The values of 𝑚 and −𝜇
, |𝑚 − 𝜇| and 𝑓|𝑚 − 𝜇| are calculated in the table below.

x f 𝒎 𝒎𝒇 |𝒎 − 𝝁| 𝑓|𝑚 − 𝜇|
51-59 5 55 275 23.4 117
60-68 3 64 192 14.4 43.2
69-77 5 73 365 5.4 27
78-86 10 82 820 3.6 36
87-95 12 91 1092 12.6 151.2
N = 35 𝜇 = 78.4 374.4
374.4
𝐴𝑣𝑒𝑟𝑎𝑔𝑒 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 = = 10.70
35
Hence, the average deviation of the scores of students in a Math Quarter exam is 10.70.

Learning check
Activity 8
Compute the average deviation of the following data values
a.) 1, 2, 5, 7, 8, 19, 22
b.) 3, 4, 7, 11, 12, 12, 15, 16
c.) -8, -5, -12, -1, 4, 7, 11
d.)
x f
16-18 3
19-21 2
22-24
Variance and Standard Deviation 5
25-27 8
28-30 7 16
e.)
Mathematics as a Tool


Variance and Standard Deviation for Ungrouped Data

The standard deviation is the most-used measure of variability. The value of the standard deviation
indicates how close the values of a data set are clustered around the mean. In general, a lower value of
the standard deviation for a data set indicates that the values of that data set are spread over a relatively
smaller range around the mean. Otherwise, it indicates that data set are spread over a relatively larger
range around the mean.

The standard deviation is obtained using the following are the basic formulas which are used to calculate
the variance:

∑(𝑥 − 𝜇)2 ∑(𝑥 − 𝑥̅ )2


𝜎2 = 𝑠2 =
𝑁 𝑛−1

where 𝜎 2 is the population variance and 𝒔𝟐 is the sample variance.

Consequently, to obtain the formula for calculating the standard deviation is:

𝜎 = √𝜎 2 𝑠 = √𝑠 2
where 𝜎 is the population standard deviation and 𝒔 is the sample standard deviation.

Example

The following table gives the estimation of 2008 market values of five international companies
(rounded to trillion of pesos). Compute for the population variance and standard deviation.

Company Market Value


PepsiCo 3.75
Google 5.35
PetroChina 13.55
Johnson and Johnson 6.90
Intel 3.55

Solution
Let x denote the 2008 market value (in billions of dollars) of a company. The values of x, 𝑥 − 𝜇,
and (𝑥 − 𝜇)2 are calculated in the table below
𝑥 𝑥−𝜇 (𝑥 − 𝜇)2
3.75 -2.87 8.24 67.37
5.35 -1.27 1.61 𝜎2 = = 𝟏𝟑. 𝟒𝟕
5
13.55 6.93 48.02 𝜎 = √13.47 = 𝟑. 𝟔𝟕
6.90 0.28 0.08
3.55 -3.07 9.42
Hence, the population variance of the
𝜇 = 6.62 67.37 2008 market value of the international
companies is 13.47, while the standard
deviation is 3.67.

17
Mathematics as a Tool


Learning check
Activity 9
The following table shows the grades of a Grade 7 student in the previous quarter. Calculate
for its population variance and standard deviation.

Subject Grade
Math 91
Filipino 94
English 89
Araling Panlipunan 90
PE 95

Variance and Standard Deviation of Grouped Data

Following are the basic formulas used to calculate the population and sample variances for a set of
grouped data respectively:

∑ 𝑓 (𝑚 − 𝜇 )2 ̅ )2
∑ 𝑓(𝑥 − 𝑥
𝜎2 = 𝑠2 =
𝑁 𝑛−1

where m and 𝑥̅ are the midpoint of a class. Also, the standard deviation is obtained by taking the
positive square root of the variance.
𝜎 2 = √𝜎 𝑠 = √𝑠

18
Mathematics as a Tool


Example
The table below gives the frequency distribution of the daily commuting times (in minutes)
from home to work for all 25 employees of a company. Compute for the sample variance
and sample standard deviation.

Daily Commuting Time of Number of Employees


Employees (minutes)
0-10 4
11-20 9
21-30 6
31-40 4
41-50 2

Solution
(1) Let x denote the market daily commuting time of employees of a company. The values of
and , and are calculated in the table below.
x f m mf
1-10 4 5.5 22 -16.4 268.96 1075.84
11-20 9 15.5 139.5 -6.4 40.96 368.64
21-30 6 25.5 153 3.6 12.96 77.76
31-40 4 35.5 142 13.6 184.96 739.84
41-50 2 45.5 91 23.6 556.96 1113.92
21.9 ∑ = 3376

Hence, the sample variance of the daily commuting times of the employees of a
company is 140.67, and the sample standard deviation is 11.86.

Learning check

Activity 10
Calculate the sample variance and sample standard deviation of the scores of students in a
Chemistry quiz as shown in the table below:

Score Number of Students


15-18 9
19-21 6
22-24 7
25-27 5
28-30 3

19
Mathematics as a Tool


Review of Inferential Statistics


 Descriptive statistics are applied to populations, and the properties of populations, like the mean or
standard deviation, are called parameters as they represent the whole population.
 Properties of samples, such as the mean or standard deviation, are not called parameters, but
statistics.

Inferential statistics are techniques that allow us to use samples to make generalizations about the
populations from which the samples were drawn. The process of accurately achieving this is called
sampling. Inferential statistics arise out of the fact that sampling naturally incurs sampling error and thus
a sample is not expected to perfectly represent the population.
The methods of inferential statistics are (1) the estimation of parameter(s) and (2) testing of statistical
hypotheses.

In this module we are only going to focus on testing of statistical hypothesis.

Testing of Statistical Hypothesis

The main purpose of statistics is to test a hypothesis. For example, you might run an experiment and find
that a certain drug is effective at treating mild symptoms of COVID-19. But if you can’t repeat that
experiment, no one will take your results seriously.

A hypothesis is an educated guess about something in the world around you. It should be testable,
either by experiment or observation. For example:

 A new medicine you think might work.


 A new mode of learning you think might be better.
 A possible location of new species.
 A fairer way to administer standardized tests.

It can really be anything at all as long as you can put it to the test.

Statistical Hypothesis

It is an assertion or conjecture concerning one or more populations.

Two types:

1. Null Hypothesis (H0) – a hypothesis of no difference, ordinarily formulated with the purpose of being
accepted or rejected.
2. Alternative Hypothesis (H1) – it is accepted if the null hypothesis is rejected. It is also classified as
directional and nondirectional hypotheses.

a.) Nondirectional – asserts that there is a significant difference between two measures (uses =
or ≠ )
b.) Directional – asserts that one measure is greater (or less) than another measure of similar
nature (uses > or <)

20
Mathematics as a Tool


Example
The following are examples of statistical hypothesis:

1. H0: The average monthly salary of nurses in a government hospital is ₱12,000


(μ = ₱12,000).
H1: The average monthly salary of nurses in a government hospital is ₱12,000
(μ ≠ ₱12,000).

2. H0: There is no significant difference between the average battery life of smartphone
brand X and that of smartphone brand Y (μx = μy).
H1: There is a significant difference between the average battery life of smartphone
brand X and that of smartphone brand Y (μx ≠ μy).

3. H0: The percentage of Puerto Princesa City college students who prefer modular
learning is 60% (p = 30%).
H1: The percentage of Puerto Princesa City college students who prefer modular
learning is greater than 60% (p > 30%).

4. H0: The percentage of twitter uses posting tweets from 6:00 to 7:00 in the evening is
the same on Fridays and Saturdays (p1 = p2).
H1: The percentage of twitter uses posting tweets from 6:00 to 7:00 in the evening is
less on Fridays than Saturdays (p1 = p2).

*We can observe that the first two examples are directional while the last two examples are
nondirectional.

Learning check

Activity 11
Create two statistical hypothesis (can be about any topic).

Two types of errors

There are four possible consequences of decisions in testing hypotheses and it is summarized in the
table below:

DECISION H0 IS TRUE H0 IS FALSE


Accept H0 Correct decision Type II error
Reject H0 Type I error Correct decision

If we accept a true null hypothesis or reject a false null hypothesis, we arrive at a perfect decision. When we
reject H0 when, in fact, it is true is a type I error. Meanwhile, if we accept H0 when, in fact, it is false it is
a type II error.

21
Mathematics as a Tool


The probability of committing a type I error is denoted by α and for committing type II error is denoted by
β, which commonly has a value of .01 and .05. The lower the value of α is, the lesser the probability of
rejecting a true null hypothesis and higher probability of accepting a false null hypothesis.

Level of Significance

In specifying the probability of committing a type I error α, which is commonly known as the level of
significance, we assume some degree of confidence with our decision.

The use of level of significance can determine a critical value which defines a rejection region (a.k.a.
critical region) and acceptance region, which serves as basis for either accepting or rejecting a
hypothesis.

The level of significance gives the area of the rejection region. (e.g. when α= 0.05, the area of the
rejection region is 0.05 and the area of the acceptance region is 0.95)

One-tailed and Two-tailed Tests

One-tailed test is a test of any statistical hypothesis where H1 is directional. Two-tailed test is where H1
is nondirectional. In the first test, the rejection region lies entirely in one side of the distribution, while in
the second test, the rejection region is split into two equal areas placed in the tails of the distribution.

In determining the critical value, the type of test and level of significance is considered in connection
with the rejection or acceptance of H0.

22
Mathematics as a Tool


Level of Significance One-sided Two-sided


α = .05 z > 1.645 z > 1.96 or z < -1.96
(or z < -1.645)
α = .01 z > 2.33 z > 2.575 or z < -2.575
(or z < -2.33)

For a two-tailed test the critical value is z = ±1.645, respectively, when α= .05 and z = ±2.33, respectively,
when α= .01.

Meanwhile, the critical value for a one-tailed test which requires the rejection region to lie on the right/left
tail of the distribution is z = ±1.96 when α= .05 and z = ±2.575, respectively, when α= .01.

Steps in Hypothesis Testing

The following are the basic steps in testing a hypothesis:

1. Formulate the null hypothesis (H0)


2. Specify the level of significance α.
3. Choose the appropriate test statistic.
4. Establish the critical region.
5. Compute for the value of the statistical test.
6. Make a decision and draw a conclusion, if possible.

Testing a Hypothesized Value of the Mean

We let 𝜇0 be the hypothesized mean of a normal population with a variance of 𝜎 2 . We take a random
sample size n from this population and obtain a sample mean 𝑥̅ . To determine whether the computed
mean 𝑥̅ is significantly different with the hypothesized mean 𝜇0 , we formulate the following hypothesis:

H0 : There is no significant difference between the computed mean and the hypothesized mean
(𝜇 = 𝜇0 ).
H1 : There is a significant difference between the computed mean and the hypothesized mean
(𝜇 ≠ 𝜇0 ).

23
Mathematics as a Tool


We will use z statistic as the test statistic since the parameter 𝜎, and the z-score corresponding to 𝑥̅ is:
𝑥̅ − 𝜇0
𝑧=
𝜎𝑥̅
where 𝑥̅ is the standard deviation of the sample mean which is computed using the equation:
𝜎
𝜎𝑥̅ =
√𝑛

Assuming the significance level is at α= .05, the critical values are ±1.96. Then we apply the following
decision rules:

1. We reject H0 and accept H1 the computed value of z is greater than 1.96 or less than -1.96.
2. We accept H0 if the computed value of z is within the interval between -1.96 and 1.96.

Example

A doctor at a certain hospital claims that the average blood glucose levels for obese patients
have a mean of 100 with a standard deviation of 15. A researcher thinks that a diet high in raw
cornstarch will have a positive or negative effect on blood glucose levels. A sample of 30
patients who have tried the raw cornstarch diet have a mean glucose level of 140. Test the
hypothesis that the raw cornstarch had an effect. (Use α= .05)

Solution

H0 : The average blood glucose level of obese patients is 100. (μ = 100)

H1 : The average blood glucose level of obese patients is not 100. (μ ≠100)

Significance level: α= .05


Test statistic: z statistic
Critical region: t > 2.998 (use t-table in the appendix of the module)
Computations:

n = 30
𝑥̅ = 140
𝜇0 = 100
𝜎 = 15
𝜎 15
𝜎𝑥̅ = =( ) = 2.7386
√𝑛 √30

𝑥̅ − 𝜇0 (140 – 100)
𝑧= = = 14.60
𝜎𝑥̅ 2.7386
The computed z value is greater than 1.96 so we reject H0 and accept H1. Therefore we can
say that the average blood glucose level of obese patients is not 100, and the raw
cornstarch diet had an effect.
.

24
Mathematics as a Tool


In some situations, σ is not given. So to estimate 𝜎𝑥̅ we use the sample standard deviation s, so the
estimated standard deviation is given by

𝑠
𝑠𝑥̅ =
√𝑛
Then to compute the z statistic we use (given that n ≥ 30)
𝑥̅ − 𝜇0
𝑠
√𝑛

Example

An instructor gives his class a midterm exam, in which from all the midterm exams he’s given
to his previous students, his students gets an average score of 81. His current class of 40
students gets a mean of 88 and a standard deviation mean of 8.5. Can he claim that his current
class is superior to his previous classes? (Use α= .01)

Solution

H0 : μ = 81

H1 : μ > 81

Significance level: α= .01, one-tailed test


Test statistic: z statistic
Critical region: z > 2.33
Computations:

n = 40
𝑥̅ = 81
𝜇0 = 88
𝜎 = 8.5
𝑠 8.5
𝑠𝑥̅ = =( ) = 1.3439
√𝑛 √40

𝑥̅ − 𝜇0 (88 – 81)
𝑧= = = 5.21
𝑠𝑥̅ 1.3439
The computed z value is greater than 1.96 so we reject H0 and accept H1. Therefore we can
say that the instructor’s claim is justified that his current class is superior.
.

In the previous examples, we can observe that z statistic is applicable as the test statistic in testing
the hypothesis of a mean such that:
1. The parameter σ is given.
2. The parameter σ is not given, but the sample size n ≥ 30

25
Mathematics as a Tool


If the parameter σ is unknown and the sample size n is less than 30, t statistic is applicable, where
the t value is obtained by
𝑥̅ − 𝜇0
𝑡=
𝑠𝑥̅
with df = n – 1, where 𝑠𝑥̅ is the standard deviation of the sample mean.

Example

Maharlika Corporation replaced their old machines to new ones (latest model) which produces
parts of computer monitors. The following data are the numbers of computer monitors
produced at the company for a sample of 10 days.

28 31 40 23 28 27 23 35 33 29

If the average number of computer monitors produced per day using the old machines is 30,
is the management justified in stating that the number of monitors produced per day can be
increase with the new machines? (Use α= .01)

Solution

H0 : The new machines does not change the average number of monitors produced per day.
(μ = 30)

H1 :The new machines increases the average number of monitors produced per day.
(μ > 30)

Significance level: α= .01, one-tailed test


Test statistic: t statistic with df = 10 – 1 = 9
Critical region: t > 2.998
Computations:

H0 and accept H1. Therefore we can say that the average blood glucose level of obese
patients is not 100, and the raw cornstarch diet had an effect.
.
𝒙 ̅)𝟐
(𝒙 − 𝒙
28 2.89
31 1.69
40 106.09
23 44.89
28 2.89
27 7.29
23 44.89
35 28.09
33 10.89
29 0.49
Σ = 297, n = 10 Σ = 250.1

∑ 𝑥 297
̅=
𝒙 = = 29.7
𝑛 10

26
Mathematics as a Tool


Solution (continutation)

̅)𝟐
∑(𝒙−𝒙 250.1
Computations: 𝑠2 = = = 27.78
𝑛−1 9
𝑠 = √27.78 = 5.27
𝑠 5.27
𝑠𝑥̅ = =( ) = 1.3439
√𝑛 √10

𝑥̅ − 𝜇0 (88 – 81)
𝑧= = = 1.66
𝑠𝑥̅ 1.3439
Since the computed t-value does not lie in the critical region, we accept the H0 and reject
the H1.
.

Learning check

Activity 12
1.) A pharmaceutical firm claims that the average time for a drug to take effect is 15 minutes
with a standard deviation of 1.5 minutes. In a sample of 40 trials, the average time was 23
minutes. Test thee claim that the alternative that the average time is not equal to 15
minutes, using a .01level of significance.

2.) Past experience has shown that the average length of time for students to enroll in a
university is 6.4 hours. A new online enrollment procedure is being tested. If a random
sample of 7 students spent 3.5, 5.3, 4.8, 6, 5, 4.7, and 2 hours for the new online enrollment,
can it be concluded that average number of hours is less than the old procedure?

Testing the Difference Between Means

Sometimes, we need to determine whether or not the mean of one population (μ1) is equal to the
mean of another population (μ2). We can formulate statistical hypothesis in a two-sample case in
three ways:

1. Ho : μ1 = μ2
H1 : μ1 ≠ μ2 (two-tailed)

2. Ho : μ1 = μ2
H1 : μ1 < μ2 (one-tailed)

3. Ho : μ1 = μ2
H1 : μ1 > μ2 (one-tailed)

Where μ1 = mean of population 1 and μ2 = mean of population 2

27
Mathematics as a Tool


Just like in the previous section, we use z-statistic when

Standard Error of the difference between means z-statistic


1.) σ1 and σ2 are 𝑥̅1 − 𝑥̅2 𝑥̅1 − 𝑥̅2
known 𝜎12 𝜎22 𝑧= =
𝜎𝑥̅ 1 −𝑥̅ 2 = √ + 𝜎𝑥̅ 1 −𝑥̅ 2 𝜎 𝜎
𝑛1 𝑛2 √𝑛1 + 𝑛2
1 2
2.) σ1 and σ2 are 𝑥̅1 − 𝑥̅2 𝑥̅1 − 𝑥̅2
unknown 𝑠12 𝑠22 𝑧= =
𝑠𝑥̅ 1−𝑥̅ 2 = √ + 𝑠𝑥̅ 1 −𝑥̅2 𝑠 𝑠
n ≥ 30 𝑛1 𝑛2 √𝑛1 + 𝑛2
1 2
3.) σ1 and σ2 are 𝑥̅1 − 𝑥̅2 𝑥̅1 − 𝑥̅2
unknown
(𝑛1 − 1)𝑠12 + (𝑛1 − 1)𝑠22 1 1 𝑡= =
𝑠𝑥̅ 1 −𝑥̅ 2 = √ ( + ) 𝑜𝑟 𝑠𝑥̅ 1−𝑥̅ 2 𝑠 𝑠
n < 30 𝑛1 + 𝑛2 𝑛1 𝑛2 √𝑛1 + 𝑛2
1 2

∑(𝑥1 − 𝑥̅1 )2 + ∑(𝑥2 − 𝑥̅2 )2 1 1 𝑑𝑓 = 𝑛1 + 𝑛2 − 2


𝑠𝑥̅ 1−𝑥̅ 2 = √ ( + )
𝑛1 + 𝑛2 𝑛1 𝑛2

Example

An instructor wishes to determine which of two methods of teaching A or B is more effective


in teaching a certain concept in Biology. In a class of 40 students, he used method A and in
another class where there are 37 students, he used method B. He gave the two classes the
same examination and obtained the following results:

Method A Method B
x1 = 87 x2 = 85
s1 = 6.1 s2 = 5.2

Is he correct in assuming that method A is more effective than method B? Use α=.05

Solution

H0 : There is no significant difference between methods A and B. (μ1 = μ2)

H1 : Method A is more effective than method B. (μ1 = μ2)

Significance level: α= .05, one-tailed test


Test statistic: z statistic
Critical region: z > 1.645
Computations:

𝑠12 𝑠22 6.12 5.22


𝑠𝑥̅ 1 −𝑥̅ 2 = √ + =√ + = √0.93025 + 0.730811 = 𝟏. 𝟐𝟖𝟖𝟖𝟐
𝑛1 𝑛2 40 37
𝑥̅1 − 𝑥̅2 87 − 85
𝑧= = = 𝟏. 𝟓𝟓𝟏𝟖
𝑠𝑥̅ 1−𝑥̅ 2 1.5518

We may reject H0 and accept H1 since the computed value of z is not in the critical region.

28
Mathematics as a Tool


Example

For a sample of 16 PinoyPhone smartphones, the mean average battery life is 42 hours with
a standard deviation of 5.6 hours. For a sample of 12 Doon smartphones, the mean average
battery life is 45 hours with a standard deviation of 6.7 hours. (Use α= .01)

Solution

H0 : There is a significant difference between the average battery lives of two smartphone
brands. (μ1 = μ2)

H1 : There is a significant difference between the average battery lives of two smartphone
brands. (μ1 ≠ μ2)

Significance level: α= .01, two-tailed test


Test statistic: t statistic with df = 16+12 – 2 = 26
Critical region: t > 2.998 or t < -2.998
Computations:
(𝑛1 − 1)𝑠12 + (𝑛1 − 1)𝑠22 1 1
𝑠𝑥̅ 1−𝑥̅ 2 = √ ( + )
𝑛1 + 𝑛2 𝑛1 𝑛2

(16 − 1)5.6 + (12 − 1)6.7 1 1


=√ ( + ) = 0.906286
16 + 12 16 12
42 − 45
𝑡= = −𝟑. 𝟑𝟏
0.906286

We accept the null hypothesis, since the computed value of t is in the region of acceptance.

Learning check

Activity 13
A History instructor wants to test if online learning mode is more effective than modular
learning mode. In the final exam, the average score of the 15 students (normal population) in
online mode is 87 with a standard deviation σ 1=9.2. The average score of the 18 students
(normal population) in modular mode is 75 with a standard deviation σ 2=7.6. Can he claim that
her students perform online learning mode do a better performance compared to those who
are on modular learning mode?

29
Mathematics as a Tool


Correlation

 Two variables are correlated if they share a statistical dependence / relationship. For example,
measurements of temperature at noon and 1pm every day are correlated, because they both lie
consistently above the mean daily temperature

 Correlations between variables are important because they indicate some underlying physical
relationship between those variables

Types of Correlation
The scatter plot explains the correlation between the two attributes or variables. It represents how
closely the two variables are connected. There can be three such situations to see the relation between
the two variables –

 Positive Correlation – when the value of one variable increases with respect to another.
 Negative Correlation – when the value of one variable decreases with respect to another.
 No Correlation – when there is no linear dependence or no relation between the two variables.

Correlation coefficient (r) Formula:

𝒏(∑ 𝒙𝒚) − (∑ 𝒙)(∑ 𝒚)


𝒓=
√[𝒏 ∑ 𝒙𝟐 − (∑ 𝒙)𝟐 ][𝒏 ∑ 𝒚𝟐 − (𝒚)𝟐 ]

Where:

n = number of pairs of scores


∑𝒙𝒚 = sum of the products of paired scores
∑𝒙 = sum of x scores
∑𝒚 = sum of y scores
∑𝒙𝟐 = sum of squared x scores
∑𝒚𝟐 = sum of squared y scores

30
Mathematics as a Tool


Here’s the correlation coefficient interpretation:

To interpret its value, see which of the following values your correlation r is closest to:

R-value Interpretation
(±) Exactly 1 A perfect uphill/downhill linear relationship
(±)0.70 A strong uphill/downhill linear relationship
(±)0.50 A moderate uphill/downhill relationship
(±)0.30 A weak uphill/downhill linear relationship
0 No linear relationship

Correlation Example
The table below gives the number of years of formal education (X) and the age of entry into the labour
force (Y), for 12 males from the Regina Labour Force Survey. Both variables are measured in years, a
ratio level of measurement and the highest level of measurement. All of the males are aged 30 or over,
so that most of these males are likely to have completed their formal education.

Respondent
X Y
Number
1 10 16
2 12 17
3 15 18
4 8 15
5 20 18
6 17 22
7 12 19
8 15 22
9 12 18
10 10 15
11 8 18
12 10 16

Table 1. Years of Education and Age of Entry into Labour Force for 12 Regina Males

Since most males enter the labour force soon after they leave formal schooling, a close relationship
between these two variables is expected. By looking through the table, it can be seen that those
respondents who obtained more years of schooling generally entered the labour force at an older age.
The mean years of schooling is 𝑋̅ = 12.4 years and the mean age of entry into the labour force is 𝑌̅ =
17.8, a difference of 5.4 years.

31
Mathematics as a Tool


This difference roughly reflects the age of entry into formal schooling, that is, age five or six. It
can be seen through that the relationship between years of schooling and age of entry into the labour
force is not perfect. Respondent 11, for example, has only 8 years of schooling but did not enter the
labour force until age 18. In contrast, respondent 5 has 20 years of schooling but entered the labour force
at age 18. The scatter diagram provides a quick way of examining the relationship between X and Y.

32
Mathematics as a Tool


Example
Based on the given data from Table 1, determine the correlation between the years of
education and age of entry into labour force for 12 Regina males.

Solution

We have to compute for ∑𝒙𝒚,∑𝒙,∑𝒚,∑𝒙𝟐 , and ∑𝒚𝟐

Respondent
x y 𝒙𝒚 𝒙𝟐 𝒚𝟐
Number
1 10 16 160 100 256
2 12 17 204 144 289
3 15 18 270 225 324
4 8 15 120 64 225
5 20 18 360 400 324
6 17 22 374 289 484
7 12 19 228 144 361
8 15 22 330 225 484
9 12 18 216 144 324
10 10 15 150 100 225
11 8 18 144 64 324
12 10 16 160 100 256
∑ ∑𝒙= 149 ∑𝒚=214 ∑𝒙𝒚=2716 ∑𝒙𝟐 =1999 𝟐
∑𝒚 =3876

Substituting this value to the r formula:

𝟏𝟐(𝟐𝟕𝟏𝟔) − (𝟏𝟒𝟗)(𝟐𝟏𝟒)
𝒓=
√[𝟏𝟐(𝟏𝟗𝟗𝟗) − (𝟏𝟒𝟗)𝟐 ][𝟏𝟐(𝟑𝟖𝟕𝟔) − (𝟐𝟏𝟒)𝟐 ]

𝟕𝟎𝟔 𝟕𝟎𝟔
𝒓= = ≈ 𝟎. 𝟔𝟐𝟒
√𝟏𝟕𝟖𝟖(𝟕𝟏𝟔) √𝟏𝟏𝟏𝟐𝟕𝟗𝟒𝟗𝟐

With the r-value of 0.624, which is closest to +0.70, there is a strong uphill linear
relationship between the years of education and age of entry into labour force for 12
Regina males.

33
Mathematics as a Tool


Learning check

Activity 13
In a Chemistry quiz, the table below shows the scores of 10 male and female students.

Male (X) Female (Y)


10 12
15 14
8 10
13 11
7 8
9 9
11 15
12 10
10 9
5 11

Compute the correlation coefficient and make an interpretation.

34
Mathematics as a Tool


Linear Regression

 Technique used for the modeling and analysis of


numerical data

 Exploits the relationship between two or more variables


so that we can gain information about one of them
through knowing values of the other

 Regression can be used for prediction, estimation,


hypothesis testing, and modeling causal relationships

For instance, a student wants to know whether there is a


relationship between the number of hours of sleep and his/her
grade in Math. The table below gives the data.

Number of hrs of sleep (x) 8.5 6.8 7.2 8.6 5.9 6.1 5.2 6.3 6.5 8
Grade in Math (y) 2.0 2.25 1.75 2.50 2.25 1.75 3.0 2.75 2.0 1.0

Using these data, a scatter plot can be drawn, as shown below.

(5.2, 3.0)

Figure 1 Figure 2

 One way to create a model of relationship between the two variables is to find a line that approximates
the data points plotted in the scattered plot.
 Out of all these possible lines, on that is closest to the behavior of the data is called the best-fit line or
least-squares regression line. This line is best fits the data better than any other line that can be
drawn.
 It can be defined as a set of data that minimizes the sum of squares of the vertical deviations from
each data point to the line. From this definition we get that the linear equation minimizes the sum
𝑑12 + 𝑑22 + 𝑑32 + 𝑑42 + 𝑑52 + 𝑑62 + 𝑑72 + 𝑑82 + 𝑑92 + 𝑑10
2

is the equation of the best-fit line, where dn represents the distance from the data point n to the line.

35
Mathematics as a Tool


3.5

3
d6
d9 d3

Grade in Math
2.5
d7 d4
d2
2 d8
d10
1.5 d1

0.5

0
0 5 10
Number of Hours of Sleep

The Formula for the Least-Squares Line

The equation of the least-squares line for the n ordered pairs (x1, y1), (x2, y2), (x3, y3), … , (xn, yn) is
𝑦 = 𝑎𝑥 ̂
+ 𝑏, 𝑤ℎ𝑒𝑟𝑒

𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦)


𝑎= 2 𝑎𝑛𝑑 𝑏 = 𝑦̅
𝑛 ∑ 𝑥 2 − (∑ 𝑥)

36
Mathematics as a Tool


Example
Apply the Least-Squares Line formula to the given data of the students’ number of hours of
sleep and grade in Math. Predict the student’s grade in Math if he/she sleeps for 10 hours.

Solution

First we find the value of each summation.

∑ 𝑥 = 69.1 ∑ 𝑦 = 21.25 ∑ 𝑥 2 = 489.29 ∑ 𝑥𝑦 = 144.275

We use the values to find the value of a.

𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦) 10(144.275) − (69.1)(21.25)


𝑎= 2 = = −0.216995512
𝑛 ∑ 𝑥 2 − (∑ 𝑥) 10(489.29) − (69.1)2

When then find the values of 𝑥̅ and 𝑦̅,

∑𝑥 69.1 ∑𝑦 21.25
𝑥̅ = = = 6.91 and 𝑦̅ = = = 2.125
𝑛 10 𝑛 10

We use them to find the x-intercept b

𝑏 = 𝑦̅ − 𝑎𝑥̅
𝑏 = 2.125 − (−0.216995516)(6.91) = 3.624439016

The regression equation is

𝒚 = −𝟎. 𝟐𝟏𝟔𝟗𝟗𝟓𝟓𝟏𝟐𝒙 + 𝟑. 𝟔𝟐𝟒𝟒𝟑𝟗𝟎𝟏𝟔

Using this equation, to predict the student’s grade we substitute 10 hrs to x


𝑦 = −0.216995512𝑥 + 3.624439016 = −0.216995512(10) + 3.624439016 = 𝟏. 𝟒𝟓
We can say that a student who sleeps for 10 hours is likely to get a grade of 1.45 in Math.

37
Mathematics as a Tool


Learning check

Activity 14

A researcher wants to find out the relationship between the speed of an adults man and a
cow’s pace (stride/walk) and his length of pace. The following gives the data of the two
variables for eight men.

a. Adult man

Pace length (meters) x 2.6 2.9 3.2 3.6 3.7 3.9 4.3 4.4
Speed (meters/seconds) y 3.3 4.8 5.6 6.7 7.1 7.6 8.2 8.6

b. Cow

Pace length (meters) x 2.7 3.2 3.4 3.6 3.7 4.0 4.2 4.4
Speed (meters/seconds) y 2.5 3.9 4.6 5.2 5.7 6.4 7.3 7.8

Predict the speed of the adult man and the cow if their pace lengths are 3.0 m and 5.0 m.

38
Mathematics as a Tool


Reflection
Express briefly your learning, points of clarification, insights, and feelings towards the given learning
material.

Rubrics %
Substance
40%
(depth and validity of the content)
Relevance
30%
(connection to the topic)
Comprehensiveness
20%
(extensiveness of the content)
Clarity
10%
(organization of though)
Total 100%

Evaluation

I. Solve the puzzle below by encircling the words identified by the descriptions in the following
items.

X A B N O I T A L E R R O C
S Q E X U U N I M O D A L R
Z V O Y A H I B D D V N E E
H J K U W F B M R E L N H G
Y D E W T F H C D B A R N R
P Y J G D L E Z E I D F D E
O M E N E R I X D T O H Q S
T O M R A H Y E D O M L E S
H A K A I Q M W R E I R T I
E B U V E C X Z P O T N I O
S I R O C E S Y R R L A E N
I R M O D E U B N A U E R C
S N I N A N Y O E O M M A I
L U T O P R S E N N A E M RS

39
Mathematics as a Tool


1.) This measure of central tendency is defined as the data with the highest frequency
2.) Statistical dependence / relationship between two variables.
3.) This measure of central tendency can be identified by dividing the sum of the data
values by the total frequency of the sample or population.
4.) This data is comparatively large or small to the other values among the data set.
5.) An educated guess.

II. Find the mean, median, mode, and average deviation of the given set of data.

1.) -5, 44, 66, 44, 33, 55, 77


2.) 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5
3.) -2, -11, -31, -5, -4, -18, -17
4.) 1, 4, 7, 11, 1, -5, -8, -9, -11
5.) 12.7, 13.7, 11.7, 14.7, 13.7

III. Compute for the mean, median, mode, and variance of the given set of sample data. Round off to
the nearest tenth.

1.) The following data shows the number of players there are in each university basketball
teams in a certain state.
18 12 16 13 14 12 13 10
11 15 17 8 9 10 21 10

2.) A survey of 10 households noted the average number of visitors they receive in a week.
The results are given below.

12 5 8 7 13 20 18 3 5 1

3.) The box below shows a random sample of the number of candies 7-year old children
eat in a single day.
6 10 7 6 5 12 9 8

40
Mathematics as a Tool


IV. Solve the mean, median, mode, variance and standard deviation of the given data set in the
following problems. Write all necessary solutions and final answer(s).

1. The following is a distribution for the number of runners in a 21 kilometer category of marathon,
with the respective distance which they have ran in the marathon.
Distance Ran Number of Runners

7–9 8

10 – 12 6

13 – 15 5

16 – 18 10

19 – 21 6

2. The following is a distribution for the number of quizzes given by a Statistics teacher in the entire
school year
Number of Weeks Number of Companies

16 – 20 5

21 – 25 4

26 – 30 6

31 – 35 7

36 – 40 3

41
Mathematics as a Tool


IV. Compute for the correlation coefficient of the given data and make an interpretation.
The following table represents the survey results from the 7 online stores.

Online Monthly E-commerce Sales (x) Online Advertising Sales (y)


Store (in 1000 s) (1000 s)
1 9200 42.5
2 8500 37.5
3 13625 70
4 13850 125
5 8275 32.5
6 11900 55
7 9400 32.5

V. Using the same data in the previous part, find the equation least-squares line. Predict the online
advertising sales of an online store if its monthly-ecommerce sales is 10,500 and 7,250.

VI. Solve the following word problems.

1.) The average amount of cellular data usage of teenagers daily are normally distributed with a mean
of 625 mb. What can you say about this claim if a random sample of 1,254 16-year olds spends
743 mb with a standard deviation of 60 mb? Use a .05 level of significance.

2.) A certain brand of canned sardines is advertised as having a net weight of 15 oz. If the net weights
of a random sample of 10 cans are 13.9, 15.2, 14.5, 14.7, 14.3, 15.5, 14.9, 16, 15.2, and 13.5 oz,
can it be concluded that the average weight of cans is less than the advertised amount? Use a .01
level of significance.

3.) In a certain university in Iloilo, a study was conducted to determine whether the IQ scores of
students who came from provincial high schools differ significantly from those students who came
from city high schools. An IQ test was given to 300 (100 from each group) college freshmen and
the results are as follows:

Students from provincial high schools: 𝑥̅1 = 99, 𝑠12 = 5


Students from city high schools: 𝑥̅2 = 102, 𝑠22 = 8

4.) In order to determine the effect of music on the speed of students in answering exams, an
instructor administered a comprehensive test while listening to music to his students and recorded
the time that they completed and submitted the test. He divided the class of 32 students into 2
groups. 17 were in the exam room with music playing and 15 were in the exam room without music.
The 17 students who listened to music had an average score of 59 with a standard deviation of 10.
The other had an average of 65 with a standard deviation of [Link] the instructor conclude that
music distracts students in taking exams?

42
Mathematics as a Tool


Design

In a group of 5 members, conduct a study and create a min-research paper wherein


you have to collect data from at least 30 respondents. Apply either Pearson correlation
or Linear regression for data analysis. The format of the paper will be provided to you
by your teacher.

References

1. Bluman, A. G. (2003). Elementary Statistics: A Step by Step Approach. 5 th Ed. McGraw Hill, Inc.
2. CENGAGE (2018). Mathematics in the Modern World.
3. Guillermo, R.M. (2018). Mathematics in the Modern World. Quezon City: Nieme Publishing House
Co. Ltd.
4. Lactuan, I. R. et. al. (2018). Instructional Material in Mathematics in the Modern World. Puerto
Princesa City: Palawan State University.
5. Walpole, M. and M. (2002). Probability and Statistics for Engineers and Scientists. 7 th Ed. Prentice
Hall Int’l. Inc.
6. [Link] [Link]
7. [Link]
8. [Link]

43
Mathematics as a Tool


Appendices

[Link]
Appendix 1. Critical values of the t-distribution

44
Mathematics as a Tool


[Link]
Appendix 2. Critical values of the z-distribution

45

Common questions

Powered by AI

A data set might have no mode if no number occurs more than once, meaning that every value has the same frequency .

The steps involve (1) formulating the null hypothesis (H0), (2) specifying the level of significance α, (3) choosing the appropriate test statistic, (4) establishing the critical region, (5) computing the value of the statistical test, and (6) making a decision and drawing a conclusion .

The z-statistic is used when the parameter σ is known and the sample size n is large, typically n ≥ 30. The t-statistic is used when the population variance is unknown and n < 30, where the t value is obtained considering the degrees of freedom .

The mode is the value that occurs with the highest frequency in a data set. Unlike the mean and median, a data set can have more than one mode or no mode at all, depending on the frequency of the values .

In the formula for the median of grouped data, class width (w) is a multiplier used after determining the position of the median in the table of cumulative frequencies, affecting the interval width over which the median lies .

The critical region is the set of all values for the test statistic that leads to the rejection of the null hypothesis. It is determined by the level of significance, which decides how much evidence against the null hypothesis is needed to reject it .

In a one-tailed test, the rejection region is entirely on one side of the distribution, used when H1 is directional. In contrast, a two-tailed test has the rejection region split into two equal areas at both tails, used when H1 is nondirectional .

For a data set with an odd number of observations, the median is given by the value of the middle term in the ranked data .

A two-tailed hypothesis test is more appropriate when the alternative hypothesis (H1) is nondirectional, implying that the effect could be in either direction, not specifically greater or less .

First, calculate the standard error: σx̅ = (8.5 / √40) = 1.3439. Then, z = (88 - 81) / 1.3439 = 5.21 .

You might also like