Chapter 4
Data
Management
Organization of Data
z conducting a statistical research, investigation or study, the research must
When
gather data for the particular variable under investigation. To describe situations, make
conclusions, and draw inferences about events, the researcher must organize the data
gathered in some meaningful way.
The data gathered shall be presented, analyzed and interpreted that can be easily
understood by the reader. Data may be presented in textual, tabular, graphical or a
combination of these presentations.
Textual Presentation uses statements with numerals in order to describe a data for
the concrete information and in expository form. It is to discuss the data and the
information and interpretation it carries.
Tabular Presentation uses statistical table to directly display the quantities or values
collected as data.
Graphical Presentation illustrates data in a form of graphs aiding readers to
understand the text easily.
These Presentations are formed using a frequency distribution by tabulating the
grouping ofzdata into categories showing the number of observations in each of the non-
overlapping classes.
Pie Chart Bar Graph
6
5
4
3
2
1
0
Category 1 Category 2 Category 3 Category 4
Series 1 Series 2 Series 3 Linear (Series 1)
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr
Line Graph
6
5
4
3
2
1
0
Category 1 Category 2 Category 3 Category 4
Series 1 Series 2 Series 3
z
Terms and Definitions - FREQUENCY DISTRIBUTION
Raw data is the data collected in original form.
Range is the difference of the highest value and the lowest value in a distribution. In
Formula: range = HV - LV
Frequency distribution is the organization of data in a tabular form, using mutually
exclusive classes showing the number of observations in each.
Class limits (or apparent limits) is the highest and lowest values describing a class.
Class boundaries (or real limits) is the upper and lower values of a class for group
frequency distribution whose values has additional decimal place more than the class
limits and end with the digit 5.
z
Terms and Definitions
Interval (or width, Class Interval, and Class Size) is the distance between the class
lower boundary and the class upper boundary and it is denoted by the symbol “ i “ or
sometimes as “ c “.
Frequency or values (denotes as “ f ” or “ n ”) is the number of values in a specific
class of a frequency distribution.
Percentage is obtained by multiplying the relative frequency by 100%.
Cumulative frequency (cf) Is the sum of the frequencies accumulated up to the upper
boundary of a class in a frequency distribution.
Midpoint is the point halfway between the class limits of each class and is
representative of the data within that class.
Steps to Construct a Frequency
z Distribution
1. Determine as to estimate number of classes “ k “,
where: k = 1 + 3.322 log n
and “ n “ is the total number of frequency values.
2. Determine the range “ r “, r = Highest Value – Lowest Value.
𝒓
3. Obtain the Class Size “ c “, c =
𝒌
NOTE: Round the value of the interval or class size up to the nearest whole number if there is a
remainder.
4. Set the lowest value as the first lower limit and get the upper limit which is equal to first lower limit
+ class size – 1.
5. Do the same process again until you reach the last class limit that includes the highest value from
the data.
z
Example 1
Twenty applicants were given a performance evaluation
appraisal. The data set is
High High High Low Average
Average Low Average Average Average
Low Average Average High High
Low Low Average High High
Construct a frequency distribution for the data.
z
Example 2
Construct a frequency distribution for the following data.
11 19 11 15 16 10
16 16 15 17 10 27
21 11 13 21 10 16
11 19 24 12 22 13
19 13 18 20 21 11
19 15 11 25 29 23
16 23 10 17 11 27
16 24 12 21 13 12
26 15 11 14 10 12
11 15 18 12 20 13
Graphing Statistical Data
z
When the data set contains large number of values, making conclusions from
an ordered array or stem-and-leaf plot is often difficult. We will need graphs or
charts in such situations. There are a number of graphs or charts to visually
show numerical data. These include:
1. Histogram,
2. Frequency Polygon, and
3. Cumulative Frequency (Ogive).
In this section, we will discuss several graphical methods that are used for
interval data. The most important of these graphical methods is the histogram.
Histogram is a powerful graphical technique used to summarize interval data,
but it also helps explain an important aspect of probability.
z
Histogram
A histogram is a graph in which the classes are marked on the
horizontal axis (x- axis) and the class frequencies on the vertical axis (y-
axis). The height of the bars represents the class frequencies, and the
bars are drawn adjacent to each other. Nevertheless, the histogram
focuses on the frequency of each class and sacrifices whatever
information is contained in the actual observation.
z
Frequency Polygon
A frequency polygon is a graph that displays the data using points
which are connected by lines. The frequencies are represented by the
heights of the points at the midpoints of the classes. The vertical axis
represents the frequency of the distribution while the horizontal axis
represents the midpoints of the frequency distribution.
Steps to Construct a Frequency Polygon
z
1. Prepare a Frequency Table: Organize the data into classes and calculate their
frequencies.
2. Find Class Midpoints: Calculate the midpoint of each class by averaging its upper
and lower boundaries.
3. Plot Points: Plot points on a graph, where the x-axis represents the class midpoints
and the y-axis represents the frequencies.
4. Connect the Points: Use straight lines to connect the plotted points.
5. Add Zero Frequencies at Ends: Optionally, extend the polygon to the x-axis by
adding points at the midpoints before the first class and after the last class with a
frequency of zero.
Cumulative Frequency Polygon (Ogive)
z
A cumulative frequency polygon or ogive (read as Oh'-jive) is a graph
that displays the cumulative frequencies for the classes in a frequency
distribution. The vertical axis represents the cumulative frequency of the
distribution while the horizontal axis represents the upper class
boundaries (real upper limits) of the frequency distribution.
z Interpretation of Data
The most appropriate measures found to be useful in describing a distribution of
observations are:
1. Measures of Central Tendency
2. Measures of Variation
3. Measures of Relative Position
4. Z-scores
5. Box and Whisker Plot
6. Probability and Normal Curve Distribution
7. Linear Regression and Correlation
z
Measures of Central Tendency
Central Tendency determines a numerical value in the central region of a distribution
of scores.
It refers to the center of a distribution of observations or data.
Three Measures of Central Tendency:
1. Mean
2. Median
3. Mode
z MEAN
The mean, Mn is also called the arithmetic mean in statistics or average in common
name.
The mean, Mn of n numbers is the sum of the numbers divided by n.
𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑣𝑎𝑙𝑢𝑒𝑠 𝑥
𝑴𝒏 = =
𝑡ℎ𝑒 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑛
Example 1:
Jeffrey has been working on programming and updating a Web site for his company for the
past 24 months. The following numbers represent the number of hours Jeffrey has worked
on his Web site for each of the past 7 months: 24, 25, 33, 50, 53, 66, 78. What is the
mean(average) number of hours that Jeffrey worked on his Web site each month?
z
The Weighted Mean
A value called the weighted mean is often used when some data values are more important
than others.
𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑝𝑟𝑜𝑑𝑢𝑐𝑡 𝑜𝑓 𝑡ℎ𝑒 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑎𝑛𝑑 𝑠𝑐𝑜𝑟𝑒 𝑓𝑋
𝑾𝑴𝒏 = =
𝑡ℎ𝑒 𝑡𝑜𝑡𝑎𝑙 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑁
Where: 𝑾𝑴𝒏 = weighted mean
f = frequency
X = Score or Class
N = Total Frequency
• Example 2:
There are 1,000 notebooks sold at Php10 each; 500 notebooks at
Php20 each; 500 notebooks at Php25 each, and 100 notebooks at
30Php each. Compute the weighted mean.
There are two ways on how to
z
solve for the value of mean given Mean of a Group Data
the grouped data or frequency
distribution.
𝑓𝑋𝑚 𝑓𝑋𝑐 𝑖
1. 𝑴𝒏 = = 𝑀𝑒𝑎𝑛 2. 𝑴𝒏 = 𝑋𝑜 +
𝑁
= Mean
𝑁
Where: Where:
f = frequency f = frequency
𝑋𝑚 = class mark 𝑋𝑜 = assumed mean to be pick any
from the 𝑋𝑚 values.
𝑓𝑋𝑚 = sum of the product of
frequencies and class marks 𝑋𝑐 = deviation from the assumed mean
N = total frequency 𝑖 = size of class interval
N = total frequency
z
MEDIAN
The median, Md, is the value in the distribution that divides and
arranged (ascending / descending) set into two equal parts.
It is the midpoint or the middlemost of the a distribution of scores.
50% of scores falls above it and 50% falls below it.
This is used when the distribution of scores is skewed.
The median separates the distribution into two equal parts.
z Median of a Single Data
The MEDIAN on a single data is obtained by inspecting the middlemost value of the
arranged distribution either in ascending or descending order.
It can also be solved using the formula:
𝐍+𝟏
Md =
𝟐
Where: It will be the Mdth Position after being arranged.
Example 1:
Find the median of the following prices:
Php 50, Php 55, Php 60, Php 65, Php 12, Php 35, Php 48.
z
Solution: (Arrange the following sets by ascending values)
Php 12, Php 35, Php 48, Php 50, Php 55, Php 60, Php 65
By this given we get, N = 7
𝐍+𝟏 𝟕+𝟏 𝟖
Md = = = = 4th Score
𝟐 𝟐 𝟐
Hence, the median is the 4th score from the arranged sets of given values.
Md = Php 50
z
MODE
The mode is the value with the largest frequency.
This used when the quickest estimate of typical performance is wanted.
A distribution can be unimodal with one mode value, bimodal with two mode values
and trimodal with three mode values. In other words it can have more than one mode
values.
Example:
1. Find the mode of the following discounts.
4%, 7%, 7%, 7%, 8%, 9%, 10%, 11%, 11%, 13%
Solution:
By inspection, the mode is 7% since it has the largest frequency.
Assignment
z
1. The daily salaries of a sample of eight employees at NEMSU are ₱ 550, ₱ 420, ₱
560, ₱ 500, ₱ 700, ₱ 860, ₱ 480. Find the mean daily rate of employees.
2. At the mathematics department of NEMSU Lianga Campus there are 18 instructors, 12
assistant professors, 7 associate professors and 3 professors. Their monthly salaries are
₱ 30,500, ₱ 33,700, ₱ 38,600, and ₱ 45,000. What is the weighted mean salary?
3. Find the median of the ages of 9 middle-management employees of a certain company.
The ages are 53, 45, 59, 48, 54, 46, 51, 58, and 55.
4. Find the mode of the ages from the previous question (number 3 problem).
5. An operations manager in charge of a company’s manufacturing keeps track of the
number of manufactured LED television in a day. For the past three weeks, the operation
manager gathers data that represent the number of LED television being manufactured:
20, 18, 19, 25, 20, 21, 20, 25, 30, 29, 28, 29, 25, 25, 27, 26, 22 and 20. Find the mode of
the given data set.