Understanding Statistics in Data Management
Understanding Statistics in Data Management
Module 8
Author: Engr. Ida E. Esquierdo Section 2: The Mathematics as a Tool
Data Management
Overview
Statistical tools derived from mathematics are useful in processing and
managing numerical data in order to describe a phenomenon and predict values. If
you invest in financial markets, you may want to predict the price of a stock in six
months from now on the basis of company performance measures and other
economic factors. As a college student, you may be interested in knowing the
dependence of the mean starting salary of a college graduate, based on your GPA.
These are just some examples that highlight how statistics are used in our modern
society.
The purpose of this module is to introduce you to the subject of statistics as
a science of data. There is data abound in this information age; how to extract
useful knowledge and gain a sound understanding in complex data sets has been
more of a challenge.
The lessons of this module will focus on the overview of descriptive and
inferential statistics. These include the basic concept of statistics, measures of
central tendency, measures of dispersion, and measures of relative position.
Learning Outcomes
After working on this module, you will be able to:
1. define statistics;
2. differentiate descriptive and inferential statistics;
3. evaluate the different measures of the average;
4. explain the appropriate measure of the average to use in a given data set;
5. use a variety of statistical tools to process and manage numerical data, and
6. appreciate the use of statistical data in making important decisions.
Activities To Do
Watch the video on this link: [Link]
Question To Ponder
What is the importance of data in understanding and dealing with various aspects of the
past and our present-day living?
Statistics defined in its plural sense is a set of numerical data, while in its singular sense
refers to the scientific discipline consisting of theory and methods in processing numerical
information that one can use when making decisions in the face of uncertainty. Others define
statistics as the art and science of collecting, presenting, analyzing, and interpreting data.
Statistics aids in decision making, summarizes or describes data, helps forecast or predict
future outcomes, aids in making inferences, and helps in comparisons or establishing
relationships. In education, for instance, statistical techniques and methods are used to get
information on enrolment, finance, physical facilities, dropout rate, proficiency level, and many
Department of Mathematics, College of Science, University of Eastern Philippines 143
GE 1 – Mathematics in the Modern World
Module 8
Author: Engr. Ida E. Esquierdo Section 2: The Mathematics as a Tool
others. In researches, statistical tools are used to test differences, effectiveness, impact,
relationships or interdependence of variables. In management, statistics is used in decision
making such as labour relations, human resource allocation, performance assessment, and many
others. In economics, it determines trends, helps financial analysts make investment decisions.
These are among the many uses of statistics.
Areas of Statistics
Classification of Variables
Levels of Measurement
Nominal, ordinal, interval, and ratio data are the different levels of measurement. They
only differ in the property of numbers (identity, order, additivity) that they possess.
Levels of Measurement
4 The following are examples of nominal, ordinal, interval, and ratio level of measurement
Nominal Ordinal Interval Ratio
Example
gender, eye color, level of educational Celsius scale weight, height, time
smoking status, attainment, military measurement of spent in watching
nationality ranks, academic temperature and television, number of
ranks, class standing intelligence score students
Measures of center, or, more colloquially, averages, is the most well-known measures of
numeric data. They are single values, intended as representatives, which can neatly characterize
the whole group. In other words, a measure of the center is any single value that is used to identify
the center of the given set of data. It is often referred to as the average. Thus, trade unions, when
negotiating pay increases, often use an ‘average wage’ as a standard on which to base those
increases. There are three most commonly used measure of the center, namely, mean, median, and
mode.
The Mean
5
1. The numbers of employees at 5 different gift shops are 4, 8, 10, 12, and 6. Find the
Example
1. x
x 4 8 10 12 6 40 8
n 5 5
2. x
x 84 75 90 98 88 79 95 86 93 89 877 87.7
n 10 10
6
The following data relates to the number of successful sales made by the salesmen
Example
Solution:
Since the given problem is presented in terms of number of salesmen for each range of
number of sales, we can use the formula for the mean of the grouped data. That is,
x
fx , where: f is the number of salesmen and x is the midpoint between each
n
number of sales and n is the total number of salesmen.
2 1 7 14 12 23 17 21 22 15 27 6
x
80
2 98 276 357 330 162
x
80
1225
x
80
x 15.31
A weighted mean is a kind of average. Instead of each data point contributing equally to the
final mean, some data points contribute more “weight” than others. If all the weights are equal,
then the weighted mean equals the arithmetic mean. The formula is given by:
x
wx
wT
The following example will show how the weighted mean is used in computing the grade
point average (GPA).
7
A university uses 4 – point grading system, A = 4, B = 3, C = 2, D = 1 and F = 0. Dillon’s 2nd
Example
x
wx 3 4 4 3 1 3 2 4 35
wT 14 14
x 2.5
The mean may not be an actual value of observation in the data set because it could be
the median or the mode that would give the actual value of the average depending on the
extreme values of the data set.
It can be applied in at least an interval level of measurement.
It is easy to compute.
Every observation contributes to the value of the mean.
Subgroup mean can be combined to come up with a group mean.
The mean is easily affected by extreme values.
Median may also be computed using the formula for the grouped data.
a. 4, 8, 1, 14, 9, 21, 12
b. 46, 23, 92, 89, 77, 108
2. The median of the ranked list 3, 4, 7, 11, 17, 29, 37 is 11. If the maximum value 37 is
increased to 55, what effect will this have on the median?
Solution:
1. The data given must be arranged either increasing or decreasing,
then get the middle value for the value of the median.
a. The number of observation is 7, therefore the middle value is 9
when arranged in increasing order. Therefore, the median is 9.
b. Since the number of observation is 6, there are two middle
values, 77 and 89. The arithmetic mean: = 83. Therefore,
the median is 83.
2. The median will remain the same because 11 will still be the middle
number in the rank list.
3. To compute the median of the given set of data, we can use the formula for the grouped
data given by:
n
F No. of Sales
Md L 2 i No. of Salesmen
f
0–4 1 1
Table 1 shows the no. of salesmen and the 5–9 14 15
10 – 14 23 38
no. of sales. Therefore, from the formula, 15 – 19 21 59
n 20 – 24 15 74
40 , F is the cumulative frequency
2 25 – 29 6 80
preceding the median class whose value is
38, f is the frequency within the median
Table 1
class equals 21, L is the lower limit within
the median class (the true lower limit is
L 0.5 14 0.5 13.5 ).
n
F
Md L 2 i 13.5 40 38 5
f 21
M d 13.5 0.48
M d 13.98
Therefore, the median of the number of sales is 13.98
The mode is the value that appears the most number of times or that value with the
greatest frequency. The mode may not exist, and even if it does exist it may not be unique. A
distribution having only one mode is called unimodal.
9
Find the mode for the data in the following lists.
Example
Data that has two modes is called bimodal. On the other hand, data that has three
modes is called trimodal.
The mean, the median, and the mode are all averages; however, they are generally not
equal. The mean of a set of data is the most sensitive of the averages. A change in any of the
numbers changes the mean, and the mean can be changed drastically by changing an extreme
value.
In contrast, the median and the mode of a set of data are usually not changed by changing
an extreme value. When a data set has one or more extreme values that are very different from
the majority of data values, the mean will not necessarily be a good indicator of an average value.
In the following example, we compare the mean, the median, and the mode for the salaries
of 5 employees of a small company.
,
= 101,200
The median is the middle number, ₱36,000. Because the ₱20,000 salary occurs the most,
the mode is ₱20,000. The data contain one extreme value that is much larger than the other
values. This extreme value makes the mean considerably larger than the median. Most of the
employees of this company would probably agree that the median of ₱36,000 better represents
the average of the salaries than does either the mean or the mode.
Department of Mathematics, College of Science, University of Eastern Philippines 151
GE 1 – Mathematics in the Modern World
Module 8
Author: Engr. Ida E. Esquierdo Section 2: The Mathematics as a Tool
Self-Assessment Activity 1
I. Identify whether the following statements is a descriptive or inferential statistics.
_________________ 1. Last semester, the ages of students at a certain college ranged from 16 to
25 years old.
_________________ 2. Based on the survey conducted by the National Statistics Office, it is
estimated that 24% of unemployed people are women.
_________________ 3. A survey says that 1 out of 10 Filipinos is a member of a fitness center.
_________________ 4. A recent study showed that eating garlic can lower blood pressure.
_________________ 5. After studying the effects of the gradual increase in dosage of a certain
drug on cancer patients, a scientist concludes that the drug can arrest
cancer cell growth at increased quantities in all cancer patients.
II. Classify the following variables as qualitative or quantitative. If the variable is quantitative,
identify if it is discrete or continuous. Then, identify further whether each variable is
nominal, ordinal, interval, or ratio.
__________________________________________________ 1. Room temperature
__________________________________________________ 2. Color of cellular phone casing
__________________________________________________ 3. Number of dining customers
__________________________________________________ 4. Serial number of car motors
__________________________________________________ 5. Flavor of ice cream
__________________________________________________ 6. Height of mercury level in a barometer
__________________________________________________ 7. Educational attainment of teachers
Activity
Compare the dispensed of soda of both machine in
Table 2. Explain the variation of data. What effect would
this make to any of the machines?
A measure of dispersion is the statistical name for the spread or variability of data. This
measure determines whether the set of observations tend to be quite similar (homogeneous) or
whether they vary considerably (heterogeneous). To measure the spread or dispersion of data,
we must introduce statistical values known as the range and the standard deviation.
Range
The range of a set of data values is the difference between the greatest data
value and the least data value.
10
Find the range of the numbers of ounces dispensed by Machine 1 in Table 2
Example
Solution:
The greatest number of ounces dispensed is 10.07 and the least is 5.85. The range of the
numbers of ounces dispensed is 10.07 5.85 = 4.22 .
Standard Deviation
The range of a set of data is easy to compute, but it can be deceiving. The range is a
measure that depends only on the two most extreme values, and as such it is very sensitive. A
measure of dispersion that is less sensitive to extreme values is the standard deviation. The
standard deviation of a set of numerical data makes use of the individual amount that each data
value deviates from the mean.
The standard deviation indicates how closely the values of a given data are set clustered
around the mean. A lower value of the standard deviation means that the values of that given data
set are spread over a smaller range around the mean. On the other hand, a larger value of the
standard deviation means that the values of the data set are spread over a larger range around the
mean.
Standard Deviation
The following examples illustrate how the standard deviation should be computed
following the procedure.
2, 4, 7, 12, 15.
Find the standard deviation of the sample.
Solution:
2 4 7 12 15
Step 1: Calculate the mean: x 8
5
Step 2: Calculate the deviation between the number and the mean.
̅
2 2 8= 6
4 4 8= 4
7 7 8= 1
12 12 8 = 4
15 15 8 = 7
Step 3: Calculate the square of each deviation in Step 2, and find the sum of these squared
deviation
̅ ( ̅)
2 2 8= 6 ( 6) = 36
4 4 8= 4 ( 4) = 16
7 7 8= 1 ( 1) = 1
12 12 8 = 4 (4) = 16
15 15 8 = 7 (7) = 49
Sum of the squared deviations
118
results of the tests are shown in the following table. According to these tests, which
company produces batteries for which the values representing hours of constant use have
the smallest standard deviation?
EverSoBright : ̅ = 7 , = 1.328
From the result, the batteries from Dependable have the smallest standard deviation.
According to these results, the Dependable company produces the most consistent
batteries with regard to life expectancy under constant use.
The Variance
Variance
A statistic known as the variance is also used as a measure of dispersion. The
variance for a given set of data is the square of the standard deviation of the
data.
13
The following numbers were obtained by sampling a population.
Example
2, 4, 7, 12, 15.
Find the variance of the sample.
Solution:
Since the standard deviation has been computed in Example 2 whose value is = 5.43, the
variance is = (5.43) = 29.48
Self-Assessment Activity 2
1. Find the range, the standard deviation, and the variance for the given samples. Round non-
integer results to the nearest tenth.
a. 1, 2, 5, 7, 8, 19, 22
b. 3, 4, 7, 11, 12, 12, 15, 16
c. 2.1, 3.0, 1.9, 1.5, 4.8
d. 5.2, 11.7, 19.1, 3.7, 8.2, 16.3
2. A mountain climber plans to buy some rope to use as a lifeline. Which of the following would
be the better choice? Explain why you think your choice is the better choice.
Rope A: Mean breaking strength: 500 lb; standard deviation of 100 lb
Rope B: Mean breaking strength: 500 lb; standard deviation of 10 lb
3. Evaluate the accuracy of the following statement: When the mean of a data set is large, the
standard deviation will be large. Explain.
There are times when we want to know the position of a value relative to the other
observations in a data set. For instance, you took a 100 – item test. You might want to know how
your score of 88 compares to the scores of the other.
Z - Score
Consider an Internet site that offers movie downloads. Based on data kept by the site, an
estimate of the mean time to download a certain movie is 12 min with a standard deviation of 4
min. When you download this movie, the download takes 20 min, and you think that is an
unusually long time for the download. On the other hand, when your friend downloads the movie,
the download takes only 6 min, and your friend is pleasantly surprised at how quickly she
receives the movie. The point here is that, in each case, a data value far from the mean is
unexpected.
Measuring the distance of a data value from the mean in standard deviation units instead
in the units of the data (minutes in this example) is quite useful. The number of standard
deviations a data value is from the mean is known as its -score or standard score.
𝒁-Score
The 𝒛-score for a given data value 𝑥 is the number of standard deviations that
𝑥 is above or below the mean of the data. The following formulas show how to
calculate the 𝑧-score for a data value 𝑥 in a population and in a sample.
mean of all scores was 65 and the standard deviation was 8. He received a 60 on a second
test, for which the mean of all scores was 45 and the standard deviation was 12. In
comparison to the other students, did Raul do better on the first test or the second test?
Solution:
Find the z – score for each test:
72 65 60 45
First test: z 0.875 Second test: z 1.25
8 12
Raul scored 0.875 standard deviation above the mean on the first test and 1.25 standard
deviations above the mean on the second test. These - scores indicate that, in comparison
to his classmates, Raul scored better on the second test than he did on the first test.
15
Which of the following exam grades has better relative position?
Example
43 40 75 72
Algebra: z 1 Geometry: z 0.6
3 5
Since, the score z - score for Algebra test is larger, the position in the Algebra test is higher
that the position in the Geometry test.
of the bulbs was 842 h, with a standard deviation of 90. One particular light bulb from the
DuraBright Company had a - score of 1.2. What was the life span of this light bulb?
Solution:
Substitute the given values into the z - score equation. Here, z 1.2 , x 842, and s 90 .
x 842
1.2
90
x 842 108 Solve for x .
x 108 842
x 950
Percentile
Most standardized examinations provide scores in terms of percentiles, which are defined
as follows:
Percentile
A percentile, p, is a measure used to indicate the value below which a given
percentage of observation fall.
For example, the 90th percentile indicates the value below which 90% of observation may
be found. And it is also the value above which 10% of the observations may be found. The
following example illustrates how this will be done.
17
The median annual salary for a physical therapist is ₱ 3,724,000. If the 90th percentile for
Example
the annual salary of a physical therapist was ₱ 5,295,000, find the percent of physical
therapists whose annual salary is
a. more than ₱ 3,724,000.
b. less than ₱ 5,295,000.
c. between ₱ 3,724,000 and ₱ 5,295,000.
Solution:
a. By definition, the median is the 50th percentile. Therefore, 50% of the physical
therapists earned more than ₱ 3,724,000 per year.
b. Because ₱ 5,295,000 is the 90th percentile, 90% of all physical therapists made less
than ₱ 5,295,000.
c. From parts a and b, 90% 50% = 40% of the physical therapists earned between
₱ 3,724,000 and ₱ 5,295,000.
18
Find the percentile rank of a test score of 49 in the data set: 12, 28, 35, 42, 47, 49, 50.
Example
Solution:
Arrange the data in order from lowest to highest. Notice that there are 5 values of
scores that is less than 49. Then substitute in the formula.
( 49)
= 100
5
Percentile 100
7
Percentile 71.4th
71.4 or approximately 71th, this means that 71% of the scores in the distribution
are less than 49 or 29% of the scores in the distribution are greater than 49.
Quartiles
2, 5, 5, 8, 11, 12, 19, 22, 23, 29, 31, 45, 83, 91, 104, 159, 181, 312, 354
Self-Assessment Activity 3
1. A data set has a mean of ̅ = 75 and a standard deviation of 11.5. Find the for each
of the following.
a. = 85
b. = 95
c. = 50
d. = 75
2. A blood pressure test was given to 450 women ages 20 to 36. It showed that their mean
systolic blood pressure was 119.4 mm Hg, with a standard deviation of 13.2 mm Hg.
a. Determine the z-score, to the nearest hundredth, for a woman who had a systolic blood
pressure reading of 110.5 mm Hg.
b. The - score for one woman was 2.15. What was her systolic blood pressure reading?
3. The following data give the scores of 18 students in a Statistics class:
84 92 87 77 98 58 63 97 82
94 84 90 68 64 92 78 75 84
Summary
This module served as a review of SHS Math 2 (Statistics and Probability). We have
discussed the differences between descriptive and inferential statistics. We have learned to
classified variables as qualitative and quantitative (discrete and continuous); levels of
measurement; measures of central tendency (mean, median, and mode); measures of
dispersion (range; standard deviation, and variance); and measures of relative position (z-
score, percentile, and quartile).
Population Mean:
Sample Mean:
Standard Deviation
Responses To Consider
After working with this module, were you able to grasp the lessons? If you are having
problems with the computation, you are encouraged to have a consultation with your
professor or instructor.
References
Aufmann, R., Lockwood, J., [Link], Mathematics in the Modern World, Rex Bookstore, Inc., 2018.
Lerner, K.L., Lerner, B.W., Real-life Math, Vol. 2, Thomson Gale, 2006.
Nocon, R., Nocon, E., Essential Mathematics for the Modern World, C & E Publishing, Inc. 2018.
Post, T.R., The Role of Manipulative Materials in the Learning Mathematical Concepts.
Retrieved from: [Link]
Note To Students
Deadline of submission of Worksheet and Reflection Paper to the Municipal Link:
MARCH 05, 2021 (FRIDAY)
Student’s Information:
Student Number: Last Name, First Name M.I.: Course – Year: