Data Management in Statistics Module
Data Management in Statistics Module
Statistics is a branch of applied mathematics that deals with gathering, organizing, presenting, analyzing, and interpreting the collected data. There are two major fields of
applied statistics – descriptive and inferential statistics. Descriptive statistics involve the collecting, organizing, describing, summarizing and presenting of gathered data in a
meaningful and informative way while inferential statistics refers to the process of drawing conclusion and making decision on the population based on evidence obtained from a
sample. Inferential statistics include estimation and hypothesis testing.
In performing all these processes involved, the application of statistical tools and techniques is necessary. Statistical tools derived from mathematics are useful on processing
and managing numerical data in order to describe a phenomenon and predict values.
The essential processes arrange the data to b analyzed and interpreted. These refer to gathering and organizing data that can be done using the frequency distribution or
grouped data and series of values in the case of few data or values. The use of the measures of central tendency is very much important to help us determine central value which can
be used to describe the general or overall performance of a certain group of values like the mean, the median, and the mode. On the other hand, the measures of dispersion can also
be utilized in order to know how close or far the data or values from each other like the range, the standard deviation, and the variance. There are also helpful in describing whether
the groups being studied or the data gathered are heterogenous or homogeneous or they are dispersed, scattered, varied, distant or spread, or they are just clustered or close to each
other. These measures include also the measures of relative position which include the z-scores, percentiles, quartiles, deciles, and box-and-whiskers plots.
The probability and the normal distributions are also discussed in this module. This is to equip the students with further knowledge and skills on how to obtain the value of
probability and problems about normal distributions.
The topics on linear regression and correlation like least-squares line and linear correlation shall be covered in the module so to help students interpret data and prove
assumptions based on the problem given.
Lesson 1: Gathering and Organizing Data, Representing Data using Graphs and Charts and Interpreting Organized Data
In today’s world we process enormous amount of data almost everyday. In schools, laboratories, and companies, volumes of data are processed. Data management plays a
very important role in processing this data. To help analyze a certain phenomenon, we need to manage data with the help of statistics. The use of statistics in predicting outcomes
and possibly explain what is happening is very evident. When data are managed efficiently, it results to understanding the nature of such phenomenon. This will further improve the
lives in the modern world.
The bar graph below depicts the confirmed COVID-19 cases, deaths, and recoveries in Southeast Asian countries.
1. 2.
When conducting a statistical study, the researcher must gather data for the particular variable under study. For example, if a researcher wishes to study the number of people
who were bitten by poisonous snakes in a specific geographic area over the past several years, he or she has to gather the data from various doctors, hospitals, or health departments.
To describe situations, draw conclusions, or make inferences about events, the researcher must organize the data in some meaningful way. The data gathered shall be presented,
analyzed and interpreted that can be easily understood by the reader. Data may be presented in textual, tabular, graphical or a combination of these.
Textual presentation uses statements with numerals in order to describe the data for the concrete information and in expository form. It is to discuss the data and the information
and interpretation it carries. For example, the math test scores of 15 students out of 50 items are 47, 48, 49, 42, 36, 38, 40, 35, 50, 26, 25, 31, 34, 19, 41.
Tabular presentation uses statistical table to directly display the quantities or values collected as data. It is a systematic arrangement of information into columns and rows.
Examples of tabular presentation are simple frequency distribution (or it can be just called a frequency distribution), cumulative frequency distribution, grouped frequency
distribution, and cumulative grouped frequency distribution
Graphical presentation illustrates data in a form of graphs aiding readers to understand the text easily. It is the most attractive, effective and convincing way in describing the
data. There are various types of graphs we can prepare like bar graph, circle graph (pie chart), line graph, pictograph, histogram, frequency polygon, and a scatter diagram.
Example 1. Twenty-five army inductees were given a blood test to determine their blood type. The data set is
c) When the range of the data is large, the data must be grouped into classes that are more than one unit in width, in what is called a grouped or quantitative frequency
distribution. It is a frequency distribution table where the data are grouped according to some numerical or quantitative characteristics. For example, a distribution of the
blood glucose levels in milligrams per deciliter (mg/dL) for 50 randomly selected college students is shown.
Step 2. Determine the number of classes (𝑘): (The rough method is between 5-20 or using 𝑘 = √𝑁 where 𝑁 is the total number of observations in the data set.
Note: Sometimes the number of classes (𝑘) is not followed. An extra class will be added to accommodate the highest observed value in the data set and a class will be deleted if it
turns out to be empty.
𝑅
Step 3. Determine the class size (𝑐) by dividing the range (𝑅) by the number of classes 𝑘 or 𝑐 = 𝑘 where 𝑐 is preferably but not absolutely necessary be an odd number and should
have the same number of decimal places in the raw data; i.e. if the observations in the data set are all whole numbers, then your c should also be a whole number.
Step 4. Enumerate the classes or categories. The classes must be mutually exclusive. Mutually exclusive classes have nonoverlapping class limits so that data cannot be placed into
two classes. The classes must be continuous. Even if there are no values in a class, the class must be included in the frequency distribution. There should be no gaps in a frequency
distribution. The only exception occurs when the class with a zero frequency is the first or last class. A class with a zero frequency at either end can be omitted without affecting
the distribution. The classes must be exhaustive. There should be enough classes to accommodate all the data.
Step 5. Tally the observations and determine the frequency of each class interval.
Step 6. Compute for the values in the other columns of the FDT as deemed necessary.
Example 2. Suppose a researcher wished to do a study on the ages of the 40 patients confined at a certain hospital. The researcher first would have to get the data on the ages of the
participants. When the data are in original form, they are called raw data and are listed next. Construct the FDT of the given data set.
The only allowable calculation on nominal data is to count the frequency of each value of the variable. We can display the counts in four ways: pie charts, bar charts, scatter
plot, time series graph, and pictograph.
1. Pie Chart (Circle graph). Pie chart is a circular graph that is useful in showing how a total quantity is distributed among a group of categories. The “pieces of the pie”
represent the proportion of the total that fall into each category. It is useful for data sorted into categories for a specific period. Its emphasis is to show the components parts
with respect to the total in terms of the percentage distribution. It uses the pie chart if there are less than 8 categories in the data set. The purpose of the pie graph is to show
the relationship of the parts to the whole by visually comparing the sizes of the sections. Percentages or proportions can be used. The variable is nominal or categorical. It is
a circle that is divided into sections or wedges according to the percentage of frequencies in each category of the distribution.
This frequency distribution shows the number of pounds of each snack food eaten during the Super Bowl. Construct a pie graph for the data.
Solution
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
12
Module 4: Data Management
Step 1. Since there are 360° in a circle, the frequency for each class must be converted to a proportional part of the circle. This conversion is done by using the formula
𝑓
Degrees = × 360°
𝑛
where 𝑓 = frequency for each class and 𝑛 = sum of the frequencies. Hence, the following conversions are obtained. The degrees
should sum to 360°.
𝑓
Step 2. Each frequency must also be converted to a percentage. Using the formula % = 𝑛 × 100. Hence, the following
percentages are obtained. The percentages should sum to 100%.
Step 3. Next, using a protractor and a compass, draw the graph, using the appropriate degree measures found in Step 1, and
label each section with the name and percentages, as shown in the figure below.
2. Bar chart (Column graph). Like pie charts, column graphs or bar charts are applicable only to grouped data. They should be used for discrete, grouped data of ordinal and
ordinal scale. Column chart is appropriate for comparing the magnitudes of variable in the x-axis for the different categories of variable in the y-axis. For time series data,
its emphasis is on the magnitude and not the movement or trend. The usual space between bars is around one-fourth of the width of the column. When the data are
qualitative or categorical, bar graphs can be used to represent the data. A bar chart can be drawn using either horizontal or vertical bars.
Horizontal bar chart is used for qualitative types of data given a specific time. It is to compare the magnitudes of the different categories of a qualitative variable. It places
the categories of the qualitative variable on the y-axis and the amount or number is on the horizontal axis. The spaces in between the bars may be one-fifth to one-half the
width of the bar.
The graphs show that first-year college students spend the most on electronic equipment.
Bar charts can also be used to compare data for two or more groups. These types of bar graphs are called compound bar graphs. Consider the following data for the number (in
millions) of never married adults in the United States.
3. Scatter plot (Scatter graph) is a graph used to represent the measurements or values that are thought to be related. It is used to examine possible relationships between two
numerical variables. The two variables are plot in 𝑥-aixs and 𝑦-axis.
4. Time series graph represents data that occur over specific period of time under observation. It shows trends, patterns, forecasts and applicable for one or more time series
data for comparison purposes.
5. Pictograph (Pictogram) immediately suggests the nature of the data being shown. It gives an approximation only of the actual figures and compares the different categories.
The symbols selected should be self-explanatory and easy to understand. Each symbol represents a number.
Example 5. Construct a histogram to represent the data shown for the record high temperatures (in ℉) for each of the
50 provinces in Luzon.
Solution
Step 1. Draw and label the x and y axes. The x axis is always the horizontal axis, and the y axis is always the vertical axis.
Step 2. Represent the frequency on the y axis and the class boundaries on the x axis.
Step 3. Using the frequencies as the heights, draw vertical bars for each class.
Step 2. Draw the x and y axes. Label the x axis with the midpoint of each class, and then use a suitable scale on the y axis for the frequencies.
Step 3. Using the midpoints for the x values and the frequencies as the y values, plot the points.
3. Cumulative frequency polygon (Ogive) is a graph that displays the cumulative frequencies for the classes in a frequency distribution. The vertical axis represents the
cumulative frequency for the classes in a frequency distribution. The vertical axis represents the cumulative frequency of the distribution while the horizontal axis
represents the upper-class boundaries of the frequency distribution.
The less than cumulative frequency polygon (less than ogive) is plotted against upper-class boundaries while the greater than cumulative frequency polygon (greater than
ogive) is plotted against lower class boundaries.
Step 2. Draw the x and y axes. Label the x axis with the class boundaries. Use an appropriate scale for the y axis to represent the cumulative frequencies. (Depending on the
numbers in the cumulative frequency columns, scales such as 0, 1, 2, 3, ..., or 5, 10, 15, 20, ..., or 1000, 2000, 3000, ... can be used. Do not label the y axis with the numbers in the
cumulative frequency column.) In this example, a scale of 0, 5, 10, 15, . . . will be used.
Step 3. Plot the cumulative frequency at each upper-class boundary, as shown in the figure below. Upper boundaries are used since the cumulative frequencies represent the
number of data values accumulated up to the upper boundary of each class.
Step 4. Starting with the first upper class boundary, 104.5, connect adjacent points with line segments, as shown in the figure below. Then extend the graph to the first lower class
boundary, 99.5, on the x axis.
21 26 25 19 29 35 34 26 25 23
27 29 24 29 22 24 28 20 20 27
35 38 25 31 19 25 27 28 22 33
34 25 32 26 26 24 23 28 26 30
23 25 22 25 29 34 34 30 17 25
Solve the following problems and show complete solution (15 points each)
1. Cereal Calories. The number of calories per serving for selected ready-to-eat cereals is listed here. Construct a frequency distribution, using 7 classes. Draw a histogram, a
frequency polygon, and an ogive for the data, using relative frequencies. Describe the shape of the histogram.
130 190 140 80 100 120 220 220 110 100
210 130 100 90 210 120 200 120 180 120
190 210 120 200 130 180 260 270 100 160
190 240 80 120 90 190 200 210 190 180
115 210 110 225 190 130
2. The table below presents the COVID-19 death cases in the Philippines by region as of September 2021. Sketch the bar chart and pie chart of the given data and interpret the
data.
Region / Location Deaths Region / Location Deaths
Metro Manila 9,142 CAR 805
Central Luzon 4,445 Caraga 692
Calabarzon 4,073 Northern Mindanao 610
Central Visayas 3,388 Zamboanga Peninsula 580
Western Visayas 1,971 SOCCSKSARGEN 544
Cagayan Valley 1,426 Bicol 506
Davao Region 1,259 Eastern Visayas 474
Ilocos Region 986 MIMAROPA 430
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
22
Module 4: Data Management
Lesson 2: Measures of Central tendency, Dispersion, and Relative Position
Any given data in statistics are useless if we don’t interpret them. The most appropriate measures found to be useful in describing a distribution of observations are the
measures of central tendency, measures of dispersion, measures of relative position, 𝑧-scores, box and whisker plot, probability and normal curve, linear regression and correlation.
For each item, answer the following questions. After that, create a small group of four members and share your answer or thoughts about the 34 35 40 40 48 21 9
questions. Then synthesizing the answers of your group, choose one representative to present the answer to the class.
1. Why a certain size of a pair shoes or a brand of shirt is made more available than the other sizes? 21 20 19 34 45 21 20
2. Why a certain basketball player gets more playing time than the rest of his teammates?
19 17 18 15 16 20 28
3. Have you ever experienced to compute the average of your grades for your wanted to compare it with your other classmates’ grades?
How did you compute it? 21 20 18 17 10 45 48
4. The set of data shows a score of 35 students in their periodical test.
19 17 29 45 50 48 25
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
23
Module 4: Data Management
a) What score is typical to the group of students? Why?
b) What score frequently appears?
c) What score appears to be in the middle? How many students fall below this score?
A measure of central tendency is any single value that is used to identify the “center” of the data or the typical value. It is called measure of central tendency because when
the data points are arranged according to magnitude, it tends to lie centrally within the set. It is the representative value of the data set. It is the value around which most of the data
points are found.
Mean
The mean represents the center of the data. It is the most important measure if the distribution is symmetric and the most stable measure of location. It is used when the data
is at least interval. When n is small, the mean is very sensitive to extreme values.
It is computed by summing all the observations in the sample and dividing the sum by the number of observations.
Properties of Mean
a) A set of data has only one mean.
b) Mean can be applied for interval and ratio data.
c) All values in the data set are included in computing the mean.
d) The mean is very useful in comparing two or more data sets.
e) Mean is affected by the extreme small or large values on a data set.
f) Mean is most appropriate in symmetrical data.
For the ungrouped data, the following are the formulas of the mean.
∑𝑥
Population Mean (𝜇): 𝜇 = 𝑁 𝑖 , where 𝑥𝑖 is the i𝑡ℎ score or observation, and 𝑁 is the number of observations in the population.
∑𝑥
Sample Mean: (𝑋̅): 𝑋̅ = 𝑛 𝑖 , where 𝑥𝑖 is the i𝑡ℎ score or observation, and 𝑛 is the number of observations in the sample.
Example 2. Determine mean age (in years) of a sample group of children whose ages are 9, 11, 7, 10, 9, 8, 8, 7, 12, 7 and 13.
∑𝑥 9 + 11 + 7 + 10 + 9 + 8 + 8 + 7 + 12 + 7 + 13 101
Solution: 𝑋̅ = 𝑖 = = = 9.18 years
𝑛 11 11
∑ 𝑓𝑀 ∑ 𝑓𝑀
For the grouped data, we have 𝜇 = or 𝑋̅ = , where 𝑓 is the frequency of the class interval and 𝑀 is the midpoint of the class interval.
𝑁 𝑛
Example 3. Calculate the mean grade of 50 students in statistics below and give its description or interpretation.
Solution.
First, determine the midpoint (𝑀) of each interval and the total frequency (∑ 𝑓) or 𝑛.
Grade 𝑓 𝑀
90 – 94 7 92
85 – 89 13 87
80 – 84 16 82
75 – 79 8 88
70 – 74 6 72
𝑛 = 50
Grade 𝑓 𝑀 𝑓𝑀
90 – 94 7 92 644
85 – 89 13 87 1131
80 – 84 16 82 1312
75 – 79 8 88 616
70 – 74 6 72 432
𝑛 = 50 ∑ 𝑓𝑀 = 4135
∑ 𝑓𝑀
Using the formula 𝑋̅ = , solve for the mean.
𝑛
∑ 𝑓𝑀 4135
𝑋̅ = 𝑛 = 50 = 82.7 or 83. Hence, the mean grade of 50 students in statistics is Satisfactory.
Weighted mean (𝑿 ̅ 𝐰 or 𝝁𝒘 ) is the sum of the mean of each group multiplied by its respective weight divided by the sum of the weights. (For mean alone, the weight values
in each distribution are equal). Example of weighted mean is solving the weighted average of a student in a semester to determine whether he or she belongs to the dean’s list. Each
of his or her grade has a corresponding number of units (Example, GECMAT is 3 units, major subject is 4 or 5 units, and so on.)
The formula of the weighted mean is
𝑋1 (𝑤1 ) + 𝑋2 (𝑤2 ) + … + 𝑋𝑛 (𝑤𝑛 )
𝑋̅w =
𝑤1 + 𝑤2 + ⋯ + 𝑤𝑛
1
Example 4. Francis answered 20 calculus problems. He spent 1 hours for the first 6 problems; 45 minutes for the next 3; and 3 hours for the last 11 problems. What was the average
2
time (in minutes) he spent for the 20 problems?
Solution
This problem requires the weighted average time because each set of problems has a weight (which is time).
𝑋1 (𝑤1 ) + 𝑋2 (𝑤2 ) + … + 𝑋𝑛 (𝑤𝑛 ) 6(90) + 3(45) + 11(180) 540 + 135 + 1980 2655
𝑋̅w = = = = ≈ 8.42 minutes
𝑤1 + 𝑤2 + ⋯ + 𝑤𝑛 90 + 45 + 180 315 315
Median is the positional middle of the data array. In the data array, one-half of the values precede the median and one-half follow it. When the data set is ordered, whether ascending
or descending, it is called a data array. Median is an appropriate measure of central tendency for data that are ordinal or above, but is more valuable in an ordinal type of data.
Properties of Median
a) The median is unique, there is only one median for a set of data.
b) The median is found by arranging the set of data from lowest or highest (or highest to lowest) and getting the value of the middle observation.
c) Median is not affected by the extreme small or large values.
d) Median can be applied for ordinal, interval and ratio data.
e) Median is most appropriate in a skewed data.
For ungrouped data, the first step in calculating the median, denoted by (𝑋̃), is to arrange the data in an array. Let X(𝑖) the 𝑖 𝑡ℎ observation in the array, 𝑖 = 1, 2, … 𝑁.
𝑁+1 𝑁+1 𝑡ℎ
If 𝑁 is odd, the median position equals ( ), and the value of the ( ) observation in the array is taken as the median, i.e. 𝑋̃ = X(𝑁+1) .
2 2 2
If 𝑁 is even, the mean of the two middle values in the array is the median, i.e.
X(𝑁) + X(𝑁+1)
2 2
𝑋̃ =
2
Example 5. Find the median of the given data set: 75, 67, 71, 75, and 72
Solution
First, arrange the data set in ascending order: 67, 71, 72, 75, 75
̃ ̃
Since 𝑁 = 5, we will use 𝑋 = X(𝑁+1) , hence, 𝑋 = X (𝑁+1)
2 2
= X5+1
2
= X3
= 72
67, 71, 72, 75, 75.
Therefore, 𝑋̃ = 72.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
27
Module 4: Data Management
Example 6. The reaction times for a random sample of 9 subjects to a stimulant were recorded as 2.5, 3.6, 3.1, 4.3, 2.9, 2.3, 2.6, 4.1, and 3.4 seconds. Calculate the median.
Solution
Array: 2.3, 2.5, 2.6, 2.9, 3.1, 3.4, 3.6, 4.1, 4.3
𝑛
−<𝑐𝑓
For grouped data, the formula for the median is 𝑋̃ = 𝑋𝐿𝐵 + ( 2 𝑓 )𝑖
𝑚
where
𝑋𝐿𝐵 = lower boundary of class containing the median
𝑛 = sample size
< 𝑐𝑓 = cumulative frequency of classes preceding class containing the median
𝑓𝑚 = number of observations in class containing the median
𝑖 = width of the interval containing the median
Example 7. Calculate the median grade of 50 students in statistics in Example 3 and give its description or interpretation.
Solution
First, we add two columns for class boundaries and less than cumulative frequency (< 𝑐𝑓).
Grade 𝑓 Class boundaries < 𝑐𝑓
90 – 94 7 89.5 – 94.5 50
85 – 89 13 84.5 – 89.5 43
80 – 84 16 79.5 – 84.5 30
75 – 79 8 74.5 – 79.5 14
70 – 74 6 69.5 – 74.5 6
𝑋𝐿𝐵 = 79.5
𝑛 = 50
< 𝑐𝑓 = 14
𝑓𝑚 = 16
𝑖=5
By substitution,
𝑛 50
−<𝑐𝑓 −14 55
𝑋̃ = 𝑋𝐿𝐵 + ( 2 𝑓 ) 𝑖 = 79.5 + ( 2
16
) 5 = 79.5 + 16 = 82.94 or 83.
𝑚
Mode is the observed value the occurs most frequently. It locates the point where the observation values occur with the greatest density. It does not always exist, and if it
does, it may not be unique. A data set is said to be unimodal if there is only one mode, bimodal if there are two modes, multimodal if there three or more. There are some cases when
a data set values have the same number frequency. When this occurs, the data set is said to be no mode.
Properties of Mode
a) The mode is found by locating the most frequently occurring value.
b) The mode is the easiest average to compute.
c) There can be more than one mode or even no mode in any given data set.
d) Mode is not affected by the extreme small or large values.
e) Mode can be applied for nominal, ordinal, interval, and ratio data.
Example 9. The reaction times for a random sample of 9 objects described in Example 6 were recorded as 2.5, 3.6, 3.1, 4.3, 2.9, 2.3, 2.6, 4.1, and 3.4 seconds. Calculate the mode.
Solution
𝜇̂ or 𝑋̂ does not exist since all values have the same frequency.
𝑑1
For grouped data, the formula of the mode is 𝑋̂ = 𝑋𝑀𝑜 + (𝑑 )𝑖
1 +𝑑2
where
𝑋𝑀𝑜 = lower class boundary of the modal class
𝑑1 = difference between the frequency of the modal class and that of the immediately preceding lower class
𝑑2 = difference between the frequency of the modal class and that of the immediately following the higher class
𝑖 = class width or size
Example 10. Calculate the modal grade of 50 students in statistics in Example 3 and give its description or interpretation.
Solution
First, determine the modal class of the distribution. The modal class of the distribution has the highest frequency. Hence, 80-84 is the modal class.
𝑋𝑀𝑜 = 79.5
𝑑1 = 8
𝑑2 = 13
𝑖=5
Solution
a) 3, 4, 5, 5, 6, 7, 9, 10, 14
∑𝑥 3+4+5+5+6+7+9+10+14
Mean: 𝑋̅ = 𝑖 = 𝑛 9
63
= 9
= 7 years
Median: Since 𝑁 is 9 (which is odd), use the formula 𝑋̃ = 𝑋(𝑁+1) .
2
𝑋̃ = 𝑋(𝑁+1) = 𝑋(9+1)
2 2
= 𝑋10
2
= 𝑋5
Hence, 𝑋5 = 6.
Mode: The mode is 5 since it has the highest frequency (it appears twice in the distribution)
∑ 𝑥𝑖 7+8+9+9+10+10+11+12
Mean: 𝑋̅ = =
𝑛 8
76
= 8
= 9.5 years
X 𝑁 +X 𝑁
( ) ( +1)
Median: Since 𝑁 is 8 (which is even), use the formula 𝑋̃ = 2
2
2
Measures of Dispersion
Computing a measure of variability is important because without it a measure of central tendency provides an incomplete description of a distribution. The mean, for example,
only indicates the central score and where the most frequent scores are. Thus, to completely describe a set of data, we need to know not only the central tendency but also how much
the individual scores differ from each other and from the center. We obtain this information by calculating statistics called measures of variability.
Measures of variability/dispersion indicate the extent to which individual items in a series are scattered about an average. It is used to determine the extent of the scatter so
that steps may be taken to control the existing variation. It is also used as a measure of reliability of the average value.
Measures of variability describe the extent to which scores in a distribution differ from each other. With many, large differences among the scores, our statistic will be a
larger number, and we say the data are more variable or show greater variability. Measures of variability communicate three related aspects of the data. First, the opposite of variability
is consistency. Small variability indicates few and/or small differences among the scores, so the scores must be consistently close to each other (and reflect that similar behaviors are
occurring). Conversely, larger variability indicates that scores (and behaviors) were inconsistent. Second, recall that a score indicates a location on a variable and that the difference
between two scores is the distance that separates them. From this perspective, by measuring differences, measures of variability indicate how spread out the scores and the distribution
are. Third, a measure of variability tells us how accurately the measure of central tendency describes the distribution. Our focus will be on the mean, so the greater the variability,
the more the scores are spread out, and the less accurately they are summarized by the one, mean score. Conversely, the smaller the variability, the closer the scores are to each other
and to the mean.
One way to describe variability is to determine how far the lowest score is from the highest score. The descriptive statistic that indicates the distance between the two most
extreme scores in a distribution is called the range.
Range
Probably the simplest and easiest way to determine measure of dispersion is the range. The range of a set of measurements is the difference between the largest value and the
smallest value. Range (𝑅) = Maximum value − Minimum value
Example 12. The IQ scores of 5 members of CHMSC Basketball men varsity are 108, 112, 127, 116, and 113. Find the range.
Solution: R = 127– 108 = 19
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
32
Module 4: Data Management
Variance and Standard Deviation
The variance and standard deviation are two measures of variability that indicate how much the scores are spread out around the mean.
Mathematically, the distance between a score and the mean is the difference between them. Recall that this difference is symbolized by 𝑋 − 𝑋̅, which is the amount that a score
deviates from the mean. Thus, a score’s deviation indicates how far it is spread out from the mean. Of course, some scores will deviate by more than others, so it makes sense to
compute something like the average amount the scores deviate from the mean. Let’s call this the “average of the deviations.” The larger the average of the deviations, the greater the
variability.
The sample variance is the average of the squared deviations of scores around the sample mean.
The symbol for the sample variance is 𝑠 2 . Always include the squared sign because it is part of the symbol.
The formula for the variance is similar to the previous formula for the average deviation except that we add the squared sign. The definitional formula for the sample variance is
̅ )2
∑(X−X
𝑠2 = 𝑛−1
where 𝑠 2 is the sample variance, 𝑠 is the sample standard deviation, 𝑋 is the value of any particular observation or measurement, 𝑋̅ is the sample mean, and 𝑛 is the sample size.
The measure of variability that more directly communicates the “average of the deviations” is the standard deviation. The symbol for the sample standard deviation is 𝑠 (which
is the square root of the symbol for the sample variance: √𝑠 2 = 𝑠).
To create the definitional formula here, we simply add the square root sign to the previous defining formula for variance. The definitional formula for the sample standard
̅ )2
∑(X−X
deviation is 𝑠 = √ .
𝑛−1
Example 13: A sample of 5 households showed the following number of household members: 3, 8, 5, 4, and 4. Find the variance and standard deviation.
Solution
First, solve for the sample mean (𝑋̅) and add the columns for (𝑋 − 𝑋̅ ) and (𝑋 − 𝑋̅)2 .
3+8+5+4+4
𝑋̅ =
5
24
= 5
= 4.8
̅)2
∑(Xi − X
𝑠2 =
𝑛−1
14.8
= 5−1
= 3.7
̅ )2
∑(Xi − X
𝑠= √
𝑛−1
= √3.7
= 1.92
Example 14: Find the measures of variability for the grades in Mathematics of the two sample groups of students.
Male: 100, 65, 75, 85, 95 Female: 84, 86, 85, 82, 83
Solution
For Range R,
Male Group, Range R = 100 − 65 = 35 Female Group, Range R = 86 − 82 = 4
For Sample Variance and Sample Standard Deviation
65+75+85+95+100
̅=
The mean for male group is X = 84.
5
82+83+84+85+86
The mean for female group is ̅
X= = 84.
5
For female group, the sample variance and sample standard deviation is
∑(Xi − ̅
X)2
s2 =
n−1
10
= 5−1
= 2.5
Findings. Both groups have the same mean but differ on all measures of variability. The male group is more variable than female group.
Conclusion. The grades of the female group are less variable than that of the male because it has smaller standard deviation. The female group has a more uniform set of grades in
Statistics than the male group.
1. Percentiles
Percentiles are values that divide a set of observations in an array into 100 equal parts. Thus, P 1, read as first percentile, is the value below which 1% of the values fall P 2,
read as second percentile, is the value below which 2% of the values fall,…, P 99, read as ninety – ninth percentile, is the value below which 99% of the fall.
Example. The 80th percentile of a distribution is a value such that at least 80 percent of the ordered observations are less than its value and at least 20 percent of the ordered
observations are larger than its value. If 𝑃80 = 75: At least 80% of the ordered observations are less than 75 or at least 20% of the ordered observations are larger than 75. So any
observation that is smaller than 𝑃80 value belongs in the lower 80% of the distribution while any observation greater than 𝑃80 value belongs in the upper 20% of the distribution.
i(n+1) th
Pi = the value of the [ ] observation in the array
100
Note:
➢ If Pi is a whole number, the ith percentile is the average of the Pi observation and the P(i+1) observation.
➢ If Pi has a fractional value, the ith percentile is the P(i+1) observation, or, round up the value of Pi to the next integer.
Example 15. The following were the scores of 10 students in a short quiz. Find the 64th percentile.
2 8 6 9 7 5 8 10 10 1
Solution: First arrange the data from lowest to highest.
1 2 5 6 7 8 8 9 10 10
i(n+1) th
Then, using Pi = [ ] observation. We have
100
Example 16. Find the 35th percentile of the given frequency distribution of 110 scores in achievement test below.
Score Frequency
50 – 54 10
55 – 59 3
60 – 64 8
65 – 69 13
70 – 74 17
75 – 79 19
80 – 84 22
85 – 89 13
90 – 94 4
95 – 99 1
TOTAL 110
By substitution, we have
𝑖𝑛
−< 𝑐𝑓𝑃𝑖
Pi = 𝑋𝑃𝑖 + 𝑐( 100 )
𝑓𝑃𝑖
38.5 − 34
P35 = 69.5 + 5 ( ) = 70.82
17
Hence, thirty-five percent of the scores in the achievement test are below 70.82.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
38
Module 4: Data Management
2. Deciles
Deciles are values that divide the array into 10 equal parts. Thus, D1, read as first decile, is the value below which is 10% of the values fall, D 2, read as second decile, is the
value below which 20% of the values fall,…, D9, read as ninth decile, is the value below which 90% of the values fall.
To compute for the ith decile, we have
i(n+1) th
Di = the value of the [ ] observation in the array
10
Example 17. From the given set scores in a quiz find the 4thdecile or D4.
3 8 9 11 12 18 19
Solution
Since the data is already arranged from lowest to highest then we may proceed in finding the 4thdecile.
3 8 9 11 12 18 19
i(n+1) th
Using Di = [ ] , we have
10
(
4 7+1)
D4 = [ 10 ]thobservation = 3.02th or 4th observation (always round up to the nearest whole number)
Since, the 4th observation in an ordered array of the given distribution is 11, therefore, the 4th decile of the distribution is 11, which is interpreted as 40% of the scores
are below 11.
Approximating the ith Decile from a Frequency distribution
To solve for the decile in grouped data, we have
𝑖𝑛
−< 𝑐𝑓𝐷𝑖
Di = 𝑋𝐷𝑖 + 𝑐( 10 )
𝑓𝐷𝑖
where
𝑖𝑛
The Dith class is the class where the 10 falls.
𝑋𝐷𝑖 = the lower-class boundary of the Dith class
𝑐 = class size of the Dith class
< 𝑐𝑓𝐷𝑖 = less than cumulative frequency of the class preceding the Dith class
𝑓𝐷𝑖 = frequency of the Dith class
Score Frequency
50 – 54 10
55 – 59 3
60 – 64 8
65 – 69 13
70 – 74 17
75 – 79 19
80 – 84 22
85 – 89 13
90 – 94 4
95 – 99 1
TOTAL 110
𝑖𝑛 𝑖𝑛 6(110)
Using 10, we have 10 = = 66. Since 66 falls on the class interval 75 – 79, hence, the D6th class is 75 – 79. Therefore, we have
10
𝑖𝑛
= 66
10
𝑋𝐷6 = 74.5
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
40
Module 4: Data Management
𝑐=5
< 𝑐𝑓𝐷6 = 51
𝑓𝐷6 = 19
By substitution, we have
𝑖𝑛
−< 𝑐𝑓𝐷6
D6 = 𝑋𝐷6 + 𝑐( 10 )
𝑓𝐷6
66 − 51
D6 = 74.5 + 5 ( ) = 78.45
19
Hence, sixty percent of the scores in the achievement test are below 78.45.
3. Quartiles
Quartiles are values that divide the array into 4 equal parts. Thus, Q1, read as first quartile, is the value below which 25% of the values fall Q2, read as second quartile, is the
value below which 50% of the values fall Q3, read as third quartile, is the value below which 75% of the values fall.
Example 19. From the given set scores in a quiz find the 3rd quartile or Q3
3 8 9 11 12 18 19
Solution
Since the data is already arranged from lowest to highest then we may proceed in finding the 3 rd quartile.
3 8 9 11 12 18 19
i(n+1)
Using Q i = [ 4 ]th, we have
3(7+1) th
Q3 = [ ] observation = 6th observation.
4
Since, the 6th observation in an ordered array of the given distribution is 18, therefore, the 3rd quartile of the distribution is 18, which is interpreted as 75% of the scores
are below 18.
Approximating the ith Quartile from a Frequency distribution
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
41
Module 4: Data Management
To solve for the quartile in grouped data, we have
𝑖𝑛
−< 𝑐𝑓𝑄𝑖
Q i = 𝑋𝑄𝑖 + 𝑐( 4 )
𝑓𝑄𝑖
where
𝑖𝑛
The Qith class is the class where the falls.
4
𝑋𝑄𝑖 = the lower-class boundary of the Qith class
𝑐 = class size of the Qith class
< 𝑐𝑓𝑄𝑖 = less than cumulative frequency of the class preceding the Qith class
𝑓𝑄𝑖 = frequency of the Qith class
Example 20. Find the 1st quartile of the given frequency distribution of 110 scores in achievement test below.
Score Frequency
50 – 54 10
55 – 59 3
60 – 64 8
65 – 69 13
70 – 74 17
75 – 79 19
80 – 84 22
85 – 89 13
90 – 94 4
95 – 99 1
TOTAL 110
𝑖𝑛 𝑖𝑛 1(110)
Using 4 , we have = = 27.5th. Since 27.5 falls on the class interval 65 – 69, hence, the Q1st class is 65 – 69. Therefore, we have
4 4
𝑖𝑛
= 27.5
4
𝑋𝑄1 = 64.5
𝑐=5
< 𝑐𝑓𝑄1 = 21
𝑓𝑄1 = 13
By substitution, we have
𝑖𝑛
−< 𝑐𝑓𝑄1
Q1 = 𝑋𝑄1 + 𝑐( 4 )
𝑓𝑄1
27.5 − 21
Q1 = 64.5 + 5 ( ) = 67
13
Hence, 25% of the scores in the achievement test are below 67.
Example 21: The monthly expenditures of a large group of households has a mean of ₱48,700 and a standard deviation of ₱10,400. What is the 𝑧 −value of monthly expenditures
of ₱59,400 and ₱38,300?
Solution
Let 𝜇 = ₱48,700 and 𝜎 = ₱10,400
Using the formula of 𝑧 to determine 𝑧 −values for the two 𝑥 values (₱59,400 and ₱38,300) are computed as follows:
𝑋−𝜇 ₱59,400−₱48,700
For ₱59,400: 𝑧= = = 1.00
𝜎 ₱10,400
𝑋−𝜇 ₱38,300−₱48,700
For ₱38,300: 𝑧= = = −1.00
𝜎 ₱10,400
The 𝑧 of 1.00 indicates that a monthly expenditure of ₱59,400 for households is one standard deviation above the mean, and a 𝑧 of −1.00 shows that a ₱38,300 monthly expenditure
is one standard deviation below the mean. Note that both household monthly expenditures (₱59,400 and ₱38,300) are the same distance (₱48,700) from the mean.
Example 22: Raul has taken two tests in his mathematics class. He scored 72 on the first test, for which the mean of all scores was 65 and the standard deviation was 8. He received
a 60 on a second test, for which the mean of all scores was 45 and the standard deviation was 12. In comparison to the other students, did Raul do better on the first test or the second
test?
Solution: Find the 𝑧 −score for each test.
72−65 60−45
𝑧72 = = 0.875 𝑧60 = = 1.25
8 12
Raul scored 0.875 standard deviation above the mean on the first test and 1.25 standard deviations above the mean on the second test. These 𝑧 −scores indicate that, in comparison
to his classmates, Raul scored better on the second test than he did on the first test.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
44
Module 4: Data Management
Example 23: A consumer group tested a sample of 100 light bulbs. It found that the mean life expectancy of the bulbs was 842 h, with a standard deviation of 90. One particular
bulb from the DuraBright Company had a 𝑧 −score of 1.2. What was the life span of this light bulb?
Solution: Substitute the given values into the 𝑧 −score equation and solve for 𝑥.
𝑋 − 𝑋̅
𝑧=
𝑠
𝑋 − 842
1.2 =
90
108 = 𝑥 − 842
950 = 𝑥
The light bulb had a life span of 950 h.
5. Box-and-Whisker Plot
A box-and-whisker plot (sometimes called a boxplot) is often used to provide a visual summary of a set of data. It is a graph of a data set obtained by drawing a horizontal line from
the minimum data value to first quartile (𝑄1), drawing a horizontal line to third quartile (𝑄3 ) to the maximum data value, and drawing a box whose vertical line passes through 𝑄1
and 𝑄3 with a vertical line inside the box passing through the median or second quartile (𝑄2 ).
Using the same scale, draw a box-and-whisker plot for each of the two data sets, placing the second plot below the first. Write a valid conclusion based on the data.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
46
Module 4: Data Management
Evaluate: Gauge Your Learning!
1. A survey of 16 energy drinks noted the caffeine concentration of each drink in milligrams per ounce. The results are given below. Find the mean, median, mode, range, variance,
and standard deviation of these data. Concentration of caffeine (mg/oz): 9.1, 7.8, 7.5, 8.9, 9.0, 8.2, 9.1, 8.7, 9.0, 7.7, 8.8, 8.9, 9.0, 9.1, 8.2, 8.9, 7.0
2. Given the data set below, find the mean, median and mode.
Frequency Distribution of Grades in College Algebra
Grade Number of Students
90 – 100 9
80 - 89 30
70 – 79 35
60 – 69 8
50 – 59 9
40 – 49 2
30 – 39 3
20 – 29 1
10 – 19 2
0–9 1
Total 100
3. A professor grades his students on 4 tests, a term paper, and a final examination. Each test counts as 15% of the course grade. The term paper counts as 20% of the course grade.
The final examination counts as 20% of the course grade. Alan has test scores of 80, 78, 92, and 84. Alan received an 84 on his term paper. His final examination score was 88.
Use the weighted mean (average) formula to find Alan’s average for the course. (Hint: The sum of all weights is 100% or 1)
Neighborhood 1 3.97, 3.91, 3.98, 3.70, 4.13, 3.97, 4.01, 3.88, 4.11, 3.70,
3.96, 3.77, 4.30, 4.08, 4.12, 4.93, 3.93, 3.94, 3.85, 3.83
Neighborhood 2 4.31, 4.22, 3.78, 4.10, 4.34, 4.20, 4.35, 4.20, 4.01, 4.04,
4.28, 4.12, 4.59, 4.12, 4.01, 3.85, 3.96, 4.28, 4.39, 4.13
Using the same scale, draw a box-and-whisker plot for each of the two data sets, placing the second plot below the first. Considering that high blood lead concentrations are
harmful to humans, in which of the two neighborhoods would you prefer to live?
7. Find the P20, D4, D6 and Q3 of the following distribution of the ages of the members of a labor union. Interpret the values.
AGE (Years) FREQUENCY
15 – 19 18
20 – 24 42
25 – 29 78
30 – 34 115
35 – 39 178
40 – 44 107
45 – 49 88
50 – 54 52
55 – 59 30
60 – 64 11
TOTAL 719
Work with a partner. Conduct two experiments. Make a frequency table and a histogram for each experiment. Compare and contrast the results of the two experiments and answer
the questions that follow.
First experiment: Toss one number cube 36 times. Record the numbers.
Second experiment: Toss two number cubes 36 times. Record the sums of the two numbers.
1. In your own words, how do histograms show the differences in distributions of data?
2. Describe an experiment that you can conduct to collect data. Predict the type of data distribution the results will create.
The accompanying table shows the scores of 50 students in Mathematics examination. Using the data, construct a histogram and answer the
questions that follow.
1. What is the mean, median and the mode of the given data? What can you say about the values of three central tendencies?
2. What is the maximum data value as shown on the histogram? (What is the largest value on the data axis?)
3. What is the minimum data value as shown on the histogram? (What is the smallest value on the data axis?)
4. Is the histogram symmetric, skewed to the left, skewed to the right, bell-shaped, uniform or does it have no special shape?
5. How many peaks does the histogram have, and where are they located? (Peaks are bars with shorter bars on each side. First bars that are taller than second
bars or last bars that are taller than the preceding are also called peaks. Two or more adjacent bars of the same height with neighboring shorter bars - a plateau - would be considered one peak.)
Example 1: The area under a bell curve gives the probability that a randomly caught fish’s weight is somewhere in a given interval. Thus, there is a big chance of catching a fish
whose weight is somewhere between 450 grams to 550 grams. Moreover, there is a small chance of catching a fish that weighs more than 550 grams. The question is how does one
get the area under that curve? The z-Table gives the answer to this question.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
51
Module 4: Data Management
To do that, standardize the normal curve. Then refer to the z-Table to obtain the value. There is
need to standardize a normal variable. Without this process, finding the area under a particular curve
is close to impossible. Setting up table of values for a normal variable just like the z-table is very
difficult; added to this burden is the number of countless possibilities for the mean and standard
deviation of a normal variable. To simplify the task of getting area, refer to the z-table (area under the
normal curve) above. The z-table (area under the normal curve) has the following properties.
1) The total area under the normal curve is 1 or 100%.
2) Since the normal curve is symmetrical about the mean, then half the normal curve has an area
of 0.5.
3) The table on the next page gives only the area to the right of the mean.
4) The given area in the table is the area from 𝑧 = 0 to ±𝑧.
5) Area is always + but 𝑧 can either be positive or negative.
6) Always draw the curve and shade the given region.
7) Simple arithmetic, addition and subtraction are the only operations needed to get the correct
area.
The normal random variable of a standard normal distribution is called a standard score or a 𝑧-score.
Every normal random variable 𝑋 can be transformed into a 𝑧-score using the following equation
𝑋−𝜇
𝑧=
𝜎
where 𝑋 is a normal random variable; 𝜇 is the mean of 𝑋; 𝜎 is the standard deviation of 𝑋
Example 3: Determine the area under the normal curve between 𝑧 = 0 and 𝑧 = −1.15.
Solution: Draw the figure and represent the area.
The area between 𝑧 = 0 and 𝑧 = −1.15 or
𝑃(−1.15 < 𝑧 < 0) is 0.3749. Therefore, the area is
0.3749 or 37.49%.
The required area is the right tail of the normal curve. Since z-table gives the area between 𝑧 = 0 and 𝑧 = 1.15, first find the area.
𝑃 (0 < 𝑧 < 1.15) = 0.3749
Then subtract 𝑃 (0 < 𝑧 < 1.15) = 0.374 from 0.5000, since half of the area under the normal is to right of 𝑧 = 0.
𝑃 (𝑧 > 1.15) = 0.5000 − 𝑃(0 < 𝑧 < 1.15)
= 0.5000 − 0.3749
= 0.1251
Therefore, the area to the right of 𝑧 = 1.15 is 0.1251 or 12.51%.
Solution: 𝑃(0 < 𝑧 < 1.85) = 0.4678 and 𝑃 (0 < 𝑧 < 0.75) = 0.2734
Hence, 𝑃 (0.75 < 𝑧 < 1.85) = 𝑃(0 < 𝑧 < 1.85) − 𝑃(0 < 𝑧 < 0.75) = 0.4678 − 0.2734 = 0.1944
Therefore, the area is 0.1944 or 19.44%.
Example 6: Find the area under the normal curve between 𝑧 = 1.15 and 𝑧 = −1.85.
Solution: 𝑃(−1.85 < 𝑧 < 0) = 0.4678 and 𝑃 (0 < 𝑧 < 1.15) = 0.3749
Since two areas are on the opposite sides of 𝑧 = 0, we must find both areas and add them.
𝑃(−1.85 < 𝑧 < 1.15) = 𝑃(−1.85 < 𝑧 < 0) + 𝑃(0 < 𝑧 < 1.15) = 0.4678 + 0.3749 = 0.8427
Hence, the total area is 0.8427 or 84.27%.
Example 9: The average Pag-ibig salary loan for RFS Pharmacy Inc. employees is ₱23,000. If the debt is
normally distributed with a standard deviation of ₱2,500, find the probability that the employee owes less than
₱18,500.
Solution:
First, draw a figure and represent the area.
Example 10: The average age of bank managers if 40 years. Assume the variable is normally distributed. If the standard deviation is 5 years, find the probability that the age of a
randomly selected bank manager will be in the range between 35 and 46 years old.
Solution:
Assume that ages of bank managers are normally distributed; then cut off points are as shown in the figure below.
First, draw a figure and represents the area.
Second, find the two 𝑧 −values
𝑋−𝜇 35−40 −5 𝑋−𝜇 46−40 6
𝑧= 𝜎
= 5
= 5
= −1.00 𝑧= 𝜎
= 5
= 5 = 1.20
Third, find the appropriate area for 𝑧 = −1.00 and 𝑧 = 1.20, in 𝑧 −table.
𝑃(−1.00 < 𝑧 < 0) = 0.3414 𝑃 (0 < 𝑧 < 1.20) = 0.3849
2. In a math competition, past records show that the average score of competitors is 78 with a standard deviation of 8.
a) What is the chance that a competitor’s score is less than 70?
b) Suppose an outstanding merit award is offered for any competitor who scored more than 90. What is the chance of winning the award?
c) What’s the chance of scoring greater than 75but less than 85?
3. A company produces different types of energy drinks. The filling machines are adjusted to pour 500 milliliters (ml) of energy drinks into each plastic bottle. Nonetheless, the
actual amount of energy drink poured into each bottle is not exactly 500 ml., it varies from bottle to bottle. It has been observed that the amount of energy drink in a bottle is
normally distributed with a mean of 500 ml and a standard deviation of 4.75 ml. What percentage of the energy drink bottles contains 505 to 513 milliliters?
4. The average daily jail population in the New Bilibid Prison in Muntinlupa City is 36,290. If the distribution is normal and the standard deviation is 3,750, find the probability
that on a randomly selected day, the jail population is greater than 40,145.
1) Write down the number of the diagram that is most likely to represent
a) the heights of a group of boys on the x axis and the sizes of shirts worn by
them on the y axis,
b) the mean temperature during the day on the x axis and the amount of gas
used to heat a house on that day on the y axis,
c) the shoe sizes of a group of adults on the x axis and their ages on the y axis.
3) Classify the variables in each item whether categorical, ordinal, interval or ratio level (each item contains two variables).
a)
b)
c)
1. Construct a scatter plot for the data shown for car rental companies in the Philippines for a recent year.
2. Construct a scatter plot for the data obtained in a study on the number of absences and the final
examination results of seven randomly selected students from a statistics class. The data are shown here.
3. Construct a scatter plot for the data obtained in a study on the number of hours that nine people exercise each week and the amount of
milk (in ounces) each person consumes per week. The data are shown.
4. Analyze the three scatter plots (#1, #2 and #3) and determine which type of relationship (positive, negative, no), if any, exists. Compare
the three scatter plots.
5. Give one real-life example of positive relationship, negative relationship, and no relationship between two variables.
Linear Regression
When performing research studies, scientists often wish to know whether two variables are related. If the variables are determined to be related. A scientist may then wish to
find an equation that can be used to model the relationship. For instance, the zoology professor R. McNeill Alexander wanted to determine whether the stride length of a dinosaur,
as shown by its fossilized footprints, could be used to estimate the speed of the dinosaur. Stride length for an animal is defined as the distance x from a particular point on a footprint
to that same point on the next footprint of the same foot. (See the figure below.) Because no dinosaurs were available, Alexander and fellow scientist A. S. Jayes carried out
experiments with many types of animals, including adult men, dogs, camels, ostriches, and elephants. The results of these experiments tended to support the idea that the speed y of
an animal is related to the animal’s stride length x. To better understand this relationship, examine the data in the table below, which are similar to, but less extensive than, the data
collected by Alexander and Jayes.
Table 1.a
Adult men
Table 1.b
Dogs
Table 1.c
Camels
In the formula for the least-squares regression line, 𝑥 represents the sum of all the 𝑥 values, 𝑦 represents the sum of all the 𝑦 values, and 𝑥𝑦 represents the sum of the 𝑛 products
𝑥1 𝑦1 , 𝑥2 𝑦2 , … , 𝑥𝑛 𝑦𝑛 . The notation 𝑥̅ represents the mean of the 𝑥 values, and 𝑦̅ represents the mean of the y values. The following example illustrates a procedure that can be used
to calculate efficiently the sums needed to find the equation of the least-squares line for a given set of data.
Example 1: Find the equation of the least-squares line for the ordered pairs in Table 1.a on page.
Solution
The ordered pairs are (2.5, 3.4), (3.0, 4.9), (3.3, 5.5), (3.5, 6.6), (3.8, 7.0), (4.0, 7.7), (4.2, 8.3), (4.5, 8.7)
The number of ordered pairs is 𝑛 = 8. Organize the data in four columns, as shown in Table 2. Then find the
sum of each column.
Find the slope 𝑎.
𝑛 ∑ 𝑥𝑦 − (∑ 𝑥)(∑ 𝑦) (8)(195.86) − (28.8)(52.1)
𝑎= = ≈ 2.7303
𝑛 ∑ 𝑥 2 − (∑ 𝑥)2 (8)(106.72) − (28.8)2
Example 2: Use the equation of the least-squares line from Example 1 to predict the average speed of
an adult man for each of the following stride lengths. Round your results to the nearest tenth of a meter
per second.
a) 2.8 m
b) 4.8 m
Solution
a) In Example 1, we found the equation of the least-squares line to be 𝑦̂ = 2.7𝑥 − 3.3. Substituting 2.8 for 𝑥 gives
𝑦̂ = 2.7𝑥 − 3.3 = 2.7(2.8) − 3.3 = 4.26
Rounding 4.26 to the nearest tenth produces 4.3. Thus 4.3 m/s is the predicted average speed for an adult man with a stride length of 2.8 m.
The procedure in Example 2a made use of an equation to determine a point between given
data points. This procedure is referred to as interpolation. In Example 2b, an equation was used
to determine a point to the right of the given data points. The process of using an equation to
determine a point to the right or left of given data points is referred to as extrapolation. See
Figure 4.
Example 3: Find the linear correlation coefficient for stride length versus speed of an adult man. Use the data in Table 1.a. Round your result to the nearest hundredth.
Solution: The ordered pairs are (2.5, 3.4), (3.0, 4.9), (3.3, 5.5), (3.5, 6.6), (3.8, 7.0), (4.0, 7.7), (4.2, 8.3), (4.5, 8.7). The number of ordered pairs is 𝑛 = 8. In Table 2, we found:
∑ 𝑥 = 28.8, ∑ 𝑦 = 52.1, ∑ 𝑥 2 = 106.72, ∑ 𝑥𝑦 = 195.86
The only additional value that is needed is
Substituting the above values into the equation for the linear correlation coefficient gives us
𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦) 8(195.86) − (28.8)(52.1)
𝑟= = ≈ 0.993715
√[𝑛(∑ 𝑥 2 ) − (∑ 𝑥)2 ][𝑛(∑ 𝑦 2 ) − (∑ 𝑦)2 ] √[8(106.72) − (28.8)2 ][8(362.25) − (52.1)2
To the nearest hundredth, the linear correlation coefficient is 0.99.
Carlos Hilado Memorial State College Module GECMAT
College of Arts and Sciences, Mathematics Department Revision 02
68
Module 4: Data Management
What is the significance of the fact that the linear correlation coefficient is positive in Example 3? (Answer: It indicates a positive correlation between a man’s stride length
and his speed. That is, as a man’s stride length increases, his speed also increases.)
The linear correlation coefficient indicates the strength of a linear relationship between two variables; however, it does not indicate the presence of a cause-and-effect relationship.
For instance, the data in Table 3 show the hours per week that a student spent playing pool and the student’s weekly algebra test scores for those same weeks.
Table 3. Algebra Test Scores vs. Hours Spent Playing Pool
The linear correlation coefficient for the ordered pairs in the table is 𝑟 ≈ 0.98. Thus there is a strong positive linear relationship between the student’s algebra test scores
and the time the student spent playing pool. This does not mean that the higher algebra test scores were caused by the increased time spent playing pool. The fact that the student’s
test scores increased with the increase in the time spent playing pool could be due to many other factors or it could just be a coincidence. In your work with applications that
involve the linear correlation coefficient 𝑟, it is important to remember the following properties of 𝑟.
In Exercises 1 and 2, find the equation of the least-squares line and the linear correlation coefficient for the given data. Round the constants, 𝑎, 𝑏, and 𝑟, and to the nearest hundredth.
In Exercises 1 and 2, find the equation of the least-squares line and the linear correlation coefficient for the given data. Round the constants, 𝑎, 𝑏, and 𝑟,
and to the nearest hundredth.
1. {(−7, −11.7), (−5, −9.8), (−3, −8.1), (1, −5.9), (2, −5.7)}
2. {(1, 4.1), (2, 6.0), (4, 8.2), (6, 11.5), (8, 16.2)}
3. The following table shows the percent of water and the number of calories in various canned soups to which 100 grams of water are
added.
a) Find the equation of the least-squares line for the data. Round constants to the nearest hundredth.
b) Use the equation in part 𝑎 to find the expected number of calories in a soup that is 89% water. Round to the nearest whole number.
c) Determine the linear correlation coefficient between the percent of water and the number of calories in various canned soups.
Finding the linear correlation coefficient (r) involves assessing the degree to which two variables share a linear relationship. The process includes calculating sums of products of deviations from the mean ([nΣxy - ΣxΣy]) and standardizing against the product of standard deviations ([√n(Σx² - (Σx)²)][√n(Σy² - (Σy)²]). For example, using stride length and speed, a calculated r-value close to 1 at 0.99 demonstrates a strong positive correlation, confirming that larger strides typically correlate with faster speeds . This relationship has implications in performance enhancement where stride adjustments may improve speed.
Constructing histograms and other graphical forms like frequency polygons and cumulative frequency polygons from frequency distributions allows for a visual interpretation of data. These visuals elucidate patterns, distributions, central tendencies, and variation directly from data, often revealing insights not immediately obvious from raw numbers alone . For example, a histogram offers a view of data symmetry or skewness, while cumulative frequency plots visualize percentiles and thresholds for decision-making. Overall, such graphical tools enhance comprehension and decision-making based on quantitative characteristics .
A positive linear correlation coefficient indicates a strong direct relationship between stride length and speed, where an increase in one variable corresponds to an increase in the other . For instance, in the dataset evaluating different stride lengths against speed, a positive coefficient of approximately 0.99 suggests that as an adult man's stride length increases, his speed increases proportionately . Such insights are valuable in athletic training and rehabilitation settings, where optimizing stride length could directly enhance performance and efficiency.
Cumulative frequency polygons for 'less than' are plotted using upper-class boundaries and indicate how many observations fall below each boundary . In contrast, 'greater than' polygons use lower class boundaries and show how many observations exceed each boundary . These differences affect data interpretation significantly; 'less than' plots help predict the distribution of values below a point (e.g., the median or a specific percentile), while 'greater than' plots are useful for analyzing outliers and extreme values. Understanding these nuances allows for more precise data analysis and prediction .
Constructing a frequency distribution table involves several steps. Step 1: Determine the range by subtracting the lowest value from the highest. This identifies the spread of data . Step 2: Decide the number of classes, usually between 5-20, adjusting for the dataset's size . Step 3: Calculate the class size by dividing the range by the number of classes, ensuring it matches the data's decimal places . Step 4: Enumerate the classes ensuring they're mutually exclusive and exhaustive, avoiding overlaps and ensuring all data points are included . Step 5: Tally observations in each class and compute frequency . This step ensures accuracy in capturing data distribution. These structured steps create a comprehensive view of data distribution, aiding analytical insights.
The z-score measures an individual's data point's deviation from the group mean, expressed in standard deviation units. A high z-score indicates the value is significantly above the mean, suggesting it might be an outlier or unusually high within the dataset. Conversely, a low z-score represents a value substantially below the mean, potentially indicating an unusually low or negative deviation . Calculating a z-score involves subtracting the mean from the individual's value and dividing by the standard deviation, offering insights into individual standing, normality, and potential outlier identification.
Ensuring class intervals are mutually exclusive prevents data from being counted in more than one category, which would lead to inaccurate frequency counts and subsequent distortions in statistical analyses . Similarly, exhaustive intervals cover the entire range of data, ensuring no data point is left unclassified, preserving the dataset's integrity and allowing comprehensive statistical evaluation. These principles maintain data accuracy and allow for reliable insights into trends and distribution characteristics .
To determine the jth quartile in a grouped data frequency distribution, use the formula Qi = XQi + c((in/4) - <cfQi/fQi), where Qi is the jth quartile, XQi is the lower-class boundary of the jth quartile class, c is the class size, <cfQi is the less than cumulative frequency of the class before the jth class, and fQi is the frequency of the jth quartile class. For example, to find Q3 using a dataset of scores, identify the 3n/4th position, locate the corresponding class, substitute the values in the formula and calculate .
True class boundaries ensure that each data point fits neatly into its respective class without overlapping into neighboring classes, which maintains accuracy in data classification and analysis. They are computed by adjusting the lower and upper class limits: true lower-class boundary equals the lower boundary minus 0.5 unit of measure, and the true upper-class boundary equals the upper limit plus 0.5 unit of measure . This methodology prevents data errors and ensures that class intervals are mutually exclusive and exhaustive, critical for precise statistical analysis and interpretation.
In least-squares regression, the slope and y-intercept define the line's equation, reflecting the relationship between variables. The slope (a) indicates the change in the dependent variable for a one-unit increase in the independent variable . It's calculated as a = (nΣxy - ΣxΣy) / (nΣx² - (Σx)²), showing correlation strength and direction . The y-intercept (b) is the value of the dependent variable when the independent variable equals zero, calculated as b = ȳ - ax̄, anchoring the line on the y-axis . This method facilitates predictive analysis by integrating these components in understanding variable dynamics.