0% found this document useful (0 votes)
7 views25 pages

Understanding Data Management in Statistics

Module 8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views25 pages

Understanding Data Management in Statistics

Module 8
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

GE 4- Mathematics in the Modern World La Carlota City College

MODULE 8
DATA MANAGEMENT
STATISTICS is a scientific body of knowledge that deals with the collection, organization or presentation,
analysis, and interpretation of data.
• Collection refers to the gathering of information or data.
• Organization or presentation involves summarizing data or information in textual, graphical, or tabular
forms.
• Analysis involves describing the data by using statistical methods and procedures.
• Interpretation refers to the process of making conclusions based on the analyzed data.

BRANCHES OF STATISTICS
▪ Descriptive Statistics – is a statistical procedure concerned with describing the characteristics and
properties of a group of persons, places or things.

For example, we may describe a collection of persons by stating how many are poor and how
many are rich, how many are literate and how many are illiterate, how many fall into various categories
of age, height, civil status, IQ, and many more. We may also describe a particular barangay in terms of
the number of families it has, the number of grade-schoolers, the number of professionals, the number
of households with certain kinds of appliances, the number of siblings in each household, or the rate of
unemployment.
Generally, descriptive statistics involve gathering, organizing, presenting and describing data.
▪ Inferential Statistics – is a statistical procedure that is used to draw inferences or information about the
properties or characteristics by a large group of people, places, or things or the basis of the information
obtained from a small portion of a large group.

Suppose we want to know the most favorite brand of toothpaste of a certain barangay and we do
not have enough time and money to interview all the residents of that barangay; we may just ask selected
residents. With the data obtained from the interviews, we shall draw or make conclusions as to
barangay’s favorite brand of toothpaste. This example involves the use of inferential statistics.

TERMINOLOGIES IN STATISTICS
Some important terms are commonly used in the study of Statistics. These terms should be understood fully in
order to facilitate the study of statistics.
1. Population refers to a large collection of objects, places, or things. To illustrate this, suppose a
researcher wants to determine the average income of the residents of a certain barangay, and there are
1500 residents in the barangay. Then all of these residents comprise the population. A population is
usually denoted or represented by N. Hence, this case, N = 1500.
GE 4- Mathematics in the Modern World La Carlota City College

2. Sample is a small portion or part of a population. It could also be defined as a sub-group, subset, or
representative of a population. For instance, suppose the above-mentioned researcher does not have
enough time and money to conduct the study using the whole population, and he wants to use only 200
residents. These 200 residents comprise the sample. A sample is usually denoted by n, thus n = 200.
3. Parameter is any numerical or nominal characteristic of a population. It is a value or measurement
obtained from a population. It is usually referred to as the actual value. If, in the preceding illustration,
the researcher uses the whole population (N=1500), then the average income obtained is called a
parameter.
4. Statistic is an estimate of a parameter. It is a value or measurement obtained from the sample. If the
researcher in the preceding illustration makes use of the sample (n=200), then the average income
obtained is called a statistic.
5. Data (singular form is datum) are facts, or a set of information or observation under study. More
specifically, data are gathered by the researcher from a population or from a sample. Data may be
classified into two categories: qualitative or quantitative
a. Qualitative data are data that can assume values that manifest the concepts of attributes. These are
sometimes called categorical data. Data falling in this category cannot be subjected to meaningful
arithmetic. They cannot be added, subtracted, or divided. Gender and nationality are qualitative
data.
Gender is a qualitative dichotomous variable since an individual may take one of the two
values “male or female”. In an opinion poll, the response of an individual towards an issue, whether
to “go” for it, “against” it, or “undecided,” is an example of a qualitative trichotomous variable.
Smoking habits of an individual in different situations may be classified as “Always/Very Often”,
“Often”, “Seldom”, “Very Seldom”, or “Never”. This set of qualitative values is called a
multinomous variable.
b. Quantitative Data are data that are numerical in nature. These are data obtained from counting or
measuring. In addition, meaningful arithmetic operations can be done with this type of data. Test
scores and height are quantitative data.
6. A variable is a characteristic or property of a population or sample that makes the members different
from each other. If a class consists of boys and girls, then gender is a variable in this class. Height is
also a variable because different people have different heights. Variables may be classified based on
whether they are discrete or continuous and whether they are dependent or independent.
a. Discrete Variable
A discrete variable can assume a finite number of values. In other words, it can assume specific
values only. The values of a discrete variable are obtained through the process of counting. The number
of students in a class is a discrete variable. If there are 40 students in a class, it cannot be reported that
there are 40.2 students or 40.5 students, because a fractional part of a student can't be in the class.
GE 4- Mathematics in the Modern World La Carlota City College

b. Continuous Variable
A continuous variable can assume infinite values within a specified interval. The values of a
continuous variable are obtained through measurement. Height is a continuous variable. If one reports
that the height of a building is 15 m, it is also possible that another person reports that the height of the
same building is 15.1m or 15.12m, depending on the precision of the measuring device used. In other
words, the height of the building can assume several values.
c. Dependent Variable
A dependent variable is a variable that is affected or influenced by another variable.
d. Independent Variable
An independent variable affects or influences the dependent variable. To illustrate Independent
and dependent variables, consider the problem entitled, The Effect of Computer-Assisted Instruction
on the Students’ Achievement in Mathematics. Here, the independent variable is the computer-
assisted instruction, while the dependent variable is the achievement of students in mathematics.
7. Constant refers to the fundamental quantities that do not change in value; fixed costs and acceleration
due to gravity are examples of such.

SCALES OF MEASUREMENT
1. Nominal Scale – This is the most primitive level of measurement. The nominal level of
measurement is used when we want to distinguish one object from another for identification
purposes. In this level, we can only say that one object is different from another, but the amount of
difference between them cannot be determined. We cannot tell that one is better or worse than the
other. Gender, nationality, and civil status are of a nominal scale.
2. Ordinal scale – In the ordinal level of measurement, data are arranged in some specified order or
rank. When objects are measured at this level, we can say that one is better or greater than the other.
But we cannot tell how much more or how much less of the characteristic one object has than the
other. The ranking of contestants in a beauty contest, or siblings in the family, or of honor students
in the class is of an ordinal scale.
3. Interval Scale – If data are measured at the interval level, we can say not only that one object is
greater or less than another, but we can also specify the amount of difference. The scores in an
examination are on an interval scale of measurement. To illustrate, suppose Kensly Kyle got 50 in
a Math examination while Kwenn Anne got 40. We can say Kensly Kyle got a higher score than
Kwenn Ann by 10 points.
4. Ratio Scale – The ratio level of measurement is like the interval level. The only difference is that
the ratio level always starts from an absolute or true zero point. In addition, in the ratio level, there
is always the presence of units of measure. If data are measured in this level, we can say that one
object is so many times as large or as small as the other. For example, suppose Mrs. Reyes weighs
GE 4- Mathematics in the Modern World La Carlota City College

50 kg, while her daughter weighs 25 kg. We can say that Mrs. Reyes is twice as heavy as her
daughter. Thus, weight is an example of data measured in a ratio.

ORGANIZATION & PRESENTATION OF DATA

▪ Ungrouped data are data that are either not organized, or if arranged, could only be from highest to
lowest or lowest to highest.

▪ Grouped data - are data that are organized and arranged into different classes or categories.

Forms of Presentation of Data


A. Textual- this form of presentation combines text and numerical facts in a statistical report.

Example1. Below are the test scores of 50 students in Calculus:


25 30 18 17 50 12 43 35 40 9
33 37 41 21 20 31 35 46 10 36
28 19 18 13 28 16 42 27 28 31
40 48 40 39 32 32 26 13 3 50
26 15 14 10 38 35 34 29 30 20

Arranging the scores from the lowest to highest will facilitate the enumeration of important characteristics of
the data. The test scores of the 50 students in Calculus arranged from lowest to highest are shown below:

3 13 17 20 27 30 32 35 40 43
9 13 18 21 28 30 33 36 40 46
10 14 18 25 28 31 34 37 40 48
10 15 19 26 28 31 35 38 41 50
12 16 20 26 29 32 35 39 42 50

The highest scores obtained is 50 and the lowest is 3. Ten students got a score of 40 and above, while only 4
got ten and below. Generally, the students performed well in the test with 33 students or 66% getting a score of
25 and above.

B. Stem – and – leaf plot which sorts data according to a certain pattern. It involves separating a number
into two parts. In a two-digit number, the stem consists of the first digit, and the leaf consists of the
second digit. While in the three digit number, the stem consists of the first two digits, and the leaf
consists of the last digit. In a one-digit number, the stem is zero.

Table 1.1
Stem-and-leaf Plot of an arranged Test Scores in Calculus of 50 Students

Stem Leaves
0 3,9
1 0,0,2,3,3,4,5,6,7,8,8,9
2 0,0,1,5,6,6,7,8,8,8,9
GE 4- Mathematics in the Modern World La Carlota City College

3 0,0,1,1,2,2,3,4,5,5,5,6,7,8,9
4 0,0,0,1,2,3,6,8
5 0,0
By looking at the stem-and –leaf plot, we can easily rank the data or put them in order. Thus, the ten
lowest scores are 3,9,10,10,12,13,13,14,15 and 16 while the ten highest scores are
40,40,40,41,42,43,46,48,50 and 50.

C. Tabular- this form of presentation is better than textual form because it provides numerical facts in a
more concise and systematic manner. Statistical tables are constructed to facilitate the
analysis of relationships. Each class/subclass is assigned to a particular row or column and
figures for various classifications are noted in appropriate cells.
Advantages of Tabular Presentation
1. It is brief, it reduces the matter to the minimum.
2. It provides the reader a good grasp of the meaning of the quantitative relationship indicated in the
report.
3. It tells the whole story without the necessity of mixing textual matter with figures.
4. The systematic arrangement of columns and rows makes them easily read and readily understood.
5. The column and rows make comparison easier.

The table has the following parts:


1. Table number: This is for easy reference to the table.
2. Table Title: It briefly explains the content of the table.
3. Column header: It describes the data in each column.
4. Row classifier: It shows the classes and categories.
5. Body: This is the main part of the table.
6. Source note: This is placed below the table when the data written are not original.

FREQUENCY DISTRIBUTION

A frequency distribution table is a table that shows the data arranged into different classes and the number
of cases that fall into each class.

Parts of the Frequency Table


1. CLASS LIMITS/CLASS INTERVAL – grouping or categories defined by lower and upper limits
Example: 16-20
21-25
26-30
2. CLASS SIZE – width of each class interval
Lower limit upper limit
L.L. U.L.
16 - 20
} class size = 5
21 - 25
GE 4- Mathematics in the Modern World La Carlota City College

3. CLASS BOUNDARIES or REAL or EXACT CLASS LIMITS are the numbers used to separate
class but without gaps created by class limits. The number to be added or subtracted is half the difference
between the upper limit of one class and the lower limit of the preceding class.

Example
Class interval Class boundaries
L.L – U.L L.C.B – U.C.B
16 – 20 15.5 - 20.5
21 – 25 20.5 - 25.5
26 - 30 25.5 - 30.5

4. CLASS MARKS are the midpoints of the classes. They can be formed by adding the lower and upper
limits and then dividing by 2.
Example:
Class interval class mark/midpoint (X)
16-20 18
21-25 23
26-30 28

STEPS IN CONSTRUCTING A FREQUENCY DISTRIBUTION TABLE


1. Compute the range. The range R is defined as the difference between the highest score and the lowest
score.

2. Decide on the number of class intervals. There should not be too many to avoid many empty classes,
and there should not be too few to avoid long details. Use the formula suggested by Sturge.
k = 1 + 3.3 log N

3. Divide the range R by the number of class intervals (k) to obtain the size of the class interval:
i= R/ k or c = R / k

4. Starting from the larger integer less than or equal to the minimum score, construct class intervals of
size i until the maximum score is reached.
5. Set up the class boundaries.
6. Tally the scores in appropriate classes and then add tallies for each class to obtain the frequency.
7. Solve the class mark or midpoint of each class. This is obtained by adding the lowest class limit and
the upper class limit, then dividing by 2.

Example 1. The following are the entrance examination scores of 60 students.


19 31 36 26 34 32
44 33 37 39 45 21
24 38 40 42 39 32
43 18 24 32 49 33
33 33 40 24 46 22
29 33 37 30 43 43
26 39 57 30 40 33
GE 4- Mathematics in the Modern World La Carlota City College

25 33 48 39 34 29
29 37 39 35 41 29
23 32 48 28 45 19

Range = 57 – 18 = 39

K = 1 + 3.3 log N c = Range / k


= 1 + 3.3 log 60 = 39 / 7
= 1 + 3.3 (1.77815) = 5.57
= 6.867 ≈6
≈7
Table 1.2
Grouped Frequency Distribution for the Entrance Examination Scores of 60 Students

Class Limits Class Boundaries Tally Frequency Class mark/midpoint


18-23 17.5 - 23.5 1111-1 6 20.5
24-29 23.5 - 29.5 1111-1111-1 11 26.5
30-35 29.5 - 35.5 1111-1111-1111-11 17 32.5
36-41 35.5 - 41.5 1111-1111-1111 14 38.5
42-47 41.5 - 47.5 1111-111 8 44.5
48-53 47.5 - 53.5 111 3 50.5
54-59 53.5 - 59.5 1 1 56.5
N = 60
Example 2. The following are the entrance examination scores of 60 students
19 31 36 26 34 32
44 33 37 39 45 21
24 38 40 42 39 32
43 18 24 32 49 33
33 33 40 24 46 22
29 33 37 30 43 43
26 39 57 30 40 33
25 33 48 39 34 29
29 37 39 35 41 29
23 32 48 28 45 19
Construct a frequency distribution table of 13 classes. In this problem, c = 39/13, c= 3

Table 1.3
Grouped Frequency Distribution for the Entrance Examination Scores of 60 Students
Class Limits Class Boundaries Tally Frequency Classmark
18-20 17.5 - 20.5 111 3 19
21-23 20.5 - 23.5 111 3 22
24-26 23.5 - 26.5 1111-1 6 25
27-29 26.5 - 29.5 1111 5 28
30-32 29.5 - 32.5 1111-11 7 31
33-35 32.5 - 35.5 1111-1111 10 34
GE 4- Mathematics in the Modern World La Carlota City College

36-38 35.5 - 38.5 1111 5 37


39-41 38.5 - 41.5 1111-1111 9 40
42-44 41.5 - 44.5 1111 5 43
45-47 44.5 - 47.5 111 3 46
48-50 47.5 - 50.5 111 3 49
51-53 50.5 - 53.5 0 52
54-56 53.5 - 56.5 0 55
57-59 56.5 - 59.5 1 1 58
N = 60
Notice that the class interval 54-56 is already the 13th class interval, but the highest value which is 57
has not been reached. Yet addition one more class is allowed in cases like this to accommodate all the
values in the set.

Cumulative Frequency Distribution


The “less than” cumulative frequency distribution (<cf) is obtained by adding successively from the
lowest to the highest interval, while the “greater than” cumulative frequency distribution (>cf) is obtained by
adding frequencies from the highest class interval to the lowest class interval.
Class limits frequency <cf >cf
18-23 6 6 60
24-29 11 17 54
30-35 17 34 43
36-41 14 48 26
42-47 8 56 12
48-53 3 59 4
54-59 1 60 1
N=60
Example 2.
Class limits freq <cf >cf
2-4 3 3 60
5-7 9 12 57
8-10 14 26 48
11-13 18 44 34
14-16 10 54 16
17-19 6 60 6
N=60

Relative Frequency Distribution


The relative frequency distribution of a class is the frequency divided by the total frequency of all
classes and is generally expressed as a percentage.

Relative Frequency
Relative Frequency= 𝑇𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛

Class limits frequency Relative Frequency rf (%)


18-23 6 0.10 10
GE 4- Mathematics in the Modern World La Carlota City College

24-29 11 0.183 18.3


30-35 17 0.283 28.3
36-41 14 0.233 23.3
42-47 8 0.133 13.3
48-53 3 0.05 5
54-59 1 0.017 1.7
N=60
Example 2.
Class limits frequency Relative Freq rf(%)
2-4 3 0.05 5
5-7 9 0.15 15
8-10 14 0.233 23.3
11-13 18 0.30 30
14-16 10 0.167 16.7
17-19 6 0.10 10
N=60

D. Graphical Presentation – this form is the most effective means of organizing and presenting statistical
data because the important relationships are brought out more clearly and creatively in virtually solid
and colorful figures.

GRAPHICAL REPRESENTATION OF THE FREQUENCY DISTRIBUTION


The following graphs can be constructed to represent a frequency distribution:
1. Histogram is a graph represented by vertical or horizontal rectangles whose bases are the class
marks and whose heights are the frequencies.
2. Frequency polygon is a line graph whose bases are the class marks and whose heights are the
frequencies.
3. Ogives is obtained by plotting the cumulative frequency by connecting points of intersection
between the class boundaries versus cumulative frequencies “less than” or “greater than”
-it is the graph of a cumulative frequency distribution.
4. A pie chart is a circle graph showing the proportion of each class through either the relative or
percentage frequency.
5. A bar chart is a graph represented by either vertical or horizontal rectangles whose bases represent
the class intervals and whose heights represent the frequencies.
GE 4- Mathematics in the Modern World La Carlota City College

MEASURES OF CENTRAL TENDENCY

Measures of Central Tendency are numerical descriptive measures which indicate or locate the center of the
distribution or data set.

In layman’s term, a measure of central tendency is the average.

MEAN
The mean of the set of values or measurements is the sum of all the measurements divided by the
number of measurements in the set.

MEAN FOR UNGROUPED DATA


Σx
̅=
𝒙 ; where, 𝒙
̅ = mean, Σx= sum of the measurements or values, n = number of measurements
𝑛

Example 1. Below are the travel times in minutes spent by Kenneth in going to school last week.
Day Time Spent in Travelling
Monday 60 min.
Tuesday 45 min.
Wednesday 50 min.
Thursday 53 min.
Friday 47 min.

Σx 60 + 45 + 50 + 53 + 47
̅=
𝒙 = ̅=
𝒙 = 51 minutes
𝑛 5

Example 2. Find the mean of the following scores 85, 79, 90, 92, 80, 78, 75,93

Mean = 85+ 79+90+ 92+ 80+ 78+75+93 / 8 = 672/8=84

WEIGHTED AVERAGE MEAN


ΣxW
̅ = -----
𝒙
ΣW

Example 1. Below are Dona’s subjects and the corresponding number of units and grades she got for the first
grading period. Compute her grade point average. (GPA)
Subject Unit Grade
Math 1 80
English 1 82
Filipino 1 83
Science 2 81
Social Studies 1 80
GE 4- Mathematics in the Modern World La Carlota City College

PEHM 1.5 85
Technology & 2 82
HE

ΣxW 80(1) + 82(1) + 83(1) + 81(2) + 80(1) + 85 (1.5) + 82 (2) 778.5


̅ = ----- = ------------------------------------------------------------------------- = ----------- = 81.95
𝒙
ΣW 1+1+1+2+1+1.5+2 9.5

Therefore, Dona’s has the GPA of 81.95 for the first grading period.

Example 2.

Subject Unit Grade


GE 1 3 2.5 7.5
GE 2 3 1.5 4.5
GE3 3 2.0 6
GE4 3 1.75 5.25
GE 5 3 1.25 3.75
Pan 12 3 1.5 4.5
GE 9 3 2.5 7.5
Pan 11 3 3.0 9
24 48

GWA = 48/24=2.0

MEAN FOR GROUPED DATA


To compute the mean for grouped data, we can use two formulas, namely:
1. The classmark formula and
2. The Coded formula

CLASSMARK FORMULA
ΣfX
̅ =
𝒙 where, x =measurement or score, f=frequency, X=class mark, n= total frequency
𝑛

example: Below is the frequency distribution of the scores of 40 students in Mathematics.


Class F X fX
Interval (classmark)
16-23 1 19.5 19.5
24-31 3 27.5 82.5
32-39 6 35.5 213
40-47 12 43.5 522
48-55 10 51.5 515
56-63 8 59.5 476
N=40 ΣfX= 1828
GE 4- Mathematics in the Modern World La Carlota City College

ΣfX 1828
X= = = 45.7
𝑛 40

CODED FORMULA:
Σfd
̅ = Xam +( ) i where, Xam = assumed mean, f=frequency, d=coded deviation, N=total frequency, i=class
𝒙 𝑛
size

Class F X d d fd fd
Interval
16-23 1 19.5 -3 0 -3 0
24-31 3 27.5 -2 1 -6 3
32-39 6 35.5 -1 2 -6 12
40-47 12 43.5 0 3 0 36
48-55 10 51.5 1 4 10 40
56-63 8 59.5 2 5 16 40
N=40 Σfd= 11 Σfd= 131
If we assign 0 in the class limits with the highest frequency
Σfd 11
̅ = Xam +(
𝒙 )i , = 43.5+ (40) 8 = 43.5 + 2.2 = 45.7
𝑛

If we assign 0 in the lowest class limits.


Σfd 131
̅ = Xam +(
𝒙 )i , = 19.5+ ( 40 ) 8 = 19.5 + 26.2 = 45.7
𝑛

MEDIAN
Median is the middle value of a given set of measurements, provided that the values or measurements
are arranged in an array. An array is an arrangement of values in increasing or decreasing order.

MEDIAN FOR UNGROUPED DATA


Example 1. The following are the ages of the Mathematics teachers in Pontevedra North Elementary School:
21, 23, 32, 28, 25, 50, 48. Compute the median.
Arrange the data in an array. 21, 23, 25, 28, 32, 48, 50. The median is 28.

1. In an English test, eight students obtained the following scores: 10, 15, 12, 18, 16, 20, 12, 14. Find
the median.

Arrange the scores in an array, that is,


10, 12, 12, 14, 15, 16, 18, 20

The median is (14+15)/2= 14.5


Mode =12

MEDIAN FOR GROUPED DATA


For grouped data, we have the following formula in finding the median:
GE 4- Mathematics in the Modern World La Carlota City College

𝑁
− <𝑐𝑓𝑏
Median = lb + ( 2
)i
𝑓
where lb = lower class boundary of the median class
n = total frequency
<cfb = less than cumulative frequency before the median class
i = size of the class interval
f = frequency of the median class
Let us illustrate how to compute the median of grouped data using the distribution of the test scores of
40 students in Mathematics given in the example of the preceding section.

Example 1:

Class Interval f <cf


16-23 1 1
24-31 3 4
32-39 6 10
40-47 median class 12 22
48-55 10 32
56-63 8 40
n=40
N/2=40/2=20
𝑁
− <𝑐𝑓𝑏
Median = lb + ( 2
)i
𝑓

= 39.5 + (20 – 10) 8


12
80
=39.5 + 12
= 39.5 + 6.67
Median = 46.17 or 46

MODE
Mode is the value that occurs most frequently in a set of measurements or values.

MODE FOR UNGROUPED DATA


The mode for ungrouped data is fairly easy to find. It is just the value or measurement which occurs the
most number of times. In other words, it is the most popular value.

A distribution may have only one mode. In this case, the distribution is said to be unimodal. Data that
have two values for the mode are said to be bimodal. It is also possible that the set of data is multimodal if
there are more than two values for the mode. If all the scores in a set of data occur only once, then the set of
data has no mode.

Example 1: The data on the number of times 10 mothers go to market every week are shown below.

Mother A B C D E F G H I J
GE 4- Mathematics in the Modern World La Carlota City College

No. of times mother goes to 2 1 3 3 1 3 2 3 3 2


market

Find the mode.

Solution: The mode is 3. This means that the majority of the mothers go to market three times a week.

MODE FOR GROUPED DATA


For grouped data, we have the formula to find the mode:

fm− fb
Mode = lbmo + ( 2𝑓𝑚−𝑓𝑎−𝑓𝑏 ) i

Where lbmo = lower class boundary of the modal class


fm = frequency of the modal class
fa = frequency after the modal class
fb = frequency before the modal class
i = size of the class interval
Below are the steps in finding the mode of grouped data:
1. Find the modal class. This is the class interval with the highest frequency.
2. Use the formula to find the mode.

It is important to note that the formula for the mode given above holds only for unimodal distribution. For
multimodal distribution, the rough mode is given by the formula
Mode = 3(Median) – 2(Mean)

Let us use the distribution of scores of 40 students in Mathematics to illustrate how to compute the mode for
grouped data.

Example: Find the mode of the data whose frequency distribution is given below.

Class Interval f
16-23 1
24-31 3
32-39 6
40-47modal class 12
48-55 10
56-63 8
n=40

Notice that the class intervals are arranged from lowest to highest group. The modal class is the class
interval 40-47. The lower class boundary of the modal class is 39.5, the frequency of the modal class is 12, the
frequency below the modal class is 6, the frequency above the modal class is 10, and the size of the class
interval is 8. Substituting these values in the formula, we have

fm− fb
Mode = lbmo + (2𝑓𝑚−𝑓𝑎−𝑓𝑏 ) i
GE 4- Mathematics in the Modern World La Carlota City College

= 39.5 + ( 12- 6 )8
2(12)-10-6
6
= 39.5 +( 8)8
= 39.5 + 6
= 45.5 or 46

In case a distribution has at least 2 modes, a rough mode can be computed as follows:
Moderough = 3 Mdn – 2Mn
where Mdn = median
Mn = mean

Another example for Ungrouped Data

Scores of 8 students in Algebra: 12, 8, 5,10,5,7,13,4


Mean = 12+8+5+10+5+7+13+4= 64/8=8
Median = 4,5,5,7,8,10,12,13 =( 7+8 )/ 2 = 7.5
Mode =5

OTHER MEASURES OF LOCATION


The measure of central tendency is a measure of location. It indicates the center of a given data. Other
descriptive measures that are used to locate the position of values or scores in the distribution are quartiles,
deciles, and percentiles. In this book, we shall not discuss these measures comprehensively, but instead we
shall just define them.

Recall that the median is the value where 50% of the distribution fails or lies above it while 50% of the
distribution lies below it. The midpoint of the line segment is the median.
In other words, the median is the value that divides the distribution into two equal parts. We define the
quartiles, deciles, and percentiles similarly. These descriptive measures – quartiles, deciles, and percentiles
– are called fractiles.

▪ Quartiles are values that divide the distribution into four equal parts.
▪ Deciles are values that divide the distribution into ten equal parts.
▪ Percentiles are values that divide the distribution into 100 equal parts.

The first quartile (Q1) is the value where 25% of the distribution lies below it, while 75% of the
distribution lies above it. The third quartile (Q3) is the value where 75% of the distribution lies below it
while 25% of the distribution lies above it. Observe that the median is equal to the second quartile (Q2).

Computation of the Quantiles for Ungrouped Data


To determine any quantile, change it first to a percentile and follow the steps below.
Step 1. Arrange first the scores according to magnitude or size.
Step 2. Find the position of the given percentile/decile/quartile in the distribution using the formula P(n
+ 1)/100
GE 4- Mathematics in the Modern World La Carlota City College

Step 3. Locate the score corresponding to the obtained position in the distribution starting from the
lowest score.
Step 4. Interpolate to get the score if the obtained position from step 2 is not exact.
𝑛+1 𝑛+1 𝑛+1
Formula: Px = x( 100 ) Dx = x( 10 ) Qx = x( )
4

Example 1. Find the 20th percentile or P20 of the following scores.


25
22
20
16
17
12
8
6
5
𝑛+1
Steps: 1. Locate the position of the score corresponding to the 20th percentile using Px = x( 100 )
9+1 10 200
Solution: P20 = 20( 100 )= 20(100)= 100 = 2
2. Locate the 2nd score from the lowest. The answer is 6.
3. Hence, the 20th percentile or P20 is 6.

This means that 20% of the cases scored below 6.

Example 2. Find the 60th percentile or P60 of the following scores.


99
95
80 5
X =79 4.8
75 4
70 3
60 2
40 1
𝑛+1
Steps: 1. Compute for the position of the 60th percentile using Px = x ( 100 )
7+1 8 480
Solution: P60 = 60 ( ) = 60 ( ) = = 4.8
100 100 100
2. Since 4.8 is between the 4th and the 5th score from the bottom, we have to interpolate to find the
answer.
Take the 4th score from the bottom, which is 75 and the 5th score, which is 80.
3. Solve the difference between these two scores, 80-75, which is 5.
4. Multiply the difference 5 by the decimal part obtained in step 1, which is 0.8. The product is 4.
5. Finally, add this product to the lower score, 75. Thus, P60 = 75 + 4 = 79.

This means that 60 % of the cases fall 79 and below.


GE 4- Mathematics in the Modern World La Carlota City College

Computations of the Quantiles for Grouped Data


The computation of any quantile for grouped data is similar to that of the median. The formula is
𝑥𝑛
−<𝑐𝑓𝑏
Px = lb + (100 𝑓 ) 𝑖
Where:
Px = the desired percentile
lb = lower boundary of the class interval containing Px
n = number of cases
cf = cumulative frequency immediately below the class interval containing Px
f = frequency of the class interval containing Px
i = class size
𝑥𝑛
−<𝑐𝑓𝑏
Dx = lb + ( 10 𝑓
)𝑖

𝑥𝑛
−<𝑐𝑓𝑏
Qx = lb + ( 4 𝑓
)𝑖

Find the Q1, P36, & D8 of the data whose frequency distribution is given below.

Class Interval f <cf


16-23 1 1
24-31 3 4
32-39 Q1 6 10
40-47 12 22
48-55 10 32
56-63 8 40
n=40

1. Q1 Solution:
𝑥𝑛 1(40) 40
= = = 10
4 4 4
<cf = 4
f =6
i =8
lb = 31.5

Substitute the values in the formula:


𝑥𝑛
−<𝑐𝑓
Q1 = lb + ( 4 𝑓
)𝑖 = 31.5 + (10−4
6
) 8 = 31.5 + 8 = 39.5

Class Interval f <cf


16-23 1 1
24-31 3 4
32-39 6 10
40-47 P36 12 22
48-55 10 32
56-63 8 40
GE 4- Mathematics in the Modern World La Carlota City College

n=40

2. P36 Solution:
𝑥𝑛 36(40) 1440
= 100 = 100 =14.4
100
<cf = 10
f = 12
i =8
lb = 39.5

Substitute the values in the formula:


𝑥𝑛
−<𝑐𝑓
P36 = lb + (100 𝑓 )𝑖
= 39.5 + (14.4−10
12
)8
= 39.5 + ( 12 ) 8
4.4

=39.5 + (35.2
12
)
= 39.5 + 2.9
= 42.4
Class Interval f <cf
16-23 1 1
24-31 3 4
32-39 6 10
40-47 12 22
48-55 D8 10 32
56-63 8 40
n=40

3. D8 Solution:
𝑥𝑛
10
= 8(40)
10
= 320
10
= 32
<cf = 22
f = 10
i =8
lb = 47.5

Substitute the values in the formula:


𝑥𝑛
−<𝑐𝑓
D8 = lb + ( 10 𝑓 ) 𝑖
= 47.5 + (32−22
10
)8
=47.5+ (10)8
10

=47.5 + 8
= 55.5

CHARACTERISTICS AND USES OF THE QUARTILE


There are three quartiles, actually: the first, the second, and the third quartiles. The second quartile is
the median. The quartiles are computed and used when a group is to be divided into four equal
subgroups according to some traits or characteristics, such as ability. The quartiles are also used in the
GE 4- Mathematics in the Modern World La Carlota City College

computation of the quartile deviation, a measure of variability. The quartiles are also used to determine
who of a group belong to the lower quartile, middle 50%, or upper quartile.

USES OF PERCENTILE
The percentile is computed and used when:
1. A scaled group of scores are to be divided into 100 equal parts of subgroups.
2. percentile bands are needed, that is, when a group of scores is to be divided into a number of subgroups
with equal or unequal number of scores in the subgroups. This is used especially in the transmutation of
raw scores into school marks or grades.
3. percentile ranks of scores are desired. When there is a large number of scores, percentile ranks are
more useful than ordinal ranks.

MEASURES OF VARIABILITY OR DISPERSION

Measures of Variability or Dispersion are measures of the average distance of each observation from the
center of the distribution. They measure the homogeneity or heterogeneity of a particular group.

A small measure of variability would indicate that the data are


1. Clustered closely around the mean;
2. More homogeneous;
3. Less variable;
4. More consistent and;
5. More uniformly distributed.
Consider the following sets of grades in Mathematics of two groups of 5 students each:
Male Group Female Group
Vincent Troy:70 Marianne Joy: 82
Kensly Kyle:95 Kwenn Anne: 80
Patrick:60 Patricia: 83
Gabriel: 80 Olivia: 81
Jerod: 100 Lurice: 79
Mean: 81 Mean: 81
The mean grade of both groups is 81. By just looking at their mean grade, we can only conclude that both
groups performed equally well in the said test, but this does not explain how far apart the grades are from one
another. Let us picture the position of each grade in a number line.

MALES

! ! ! ! ! !
60 70 80 90 95 100
FEMALES

! ! ! ! !
60 70 80 90 100
GE 4- Mathematics in the Modern World La Carlota City College

Notice that the grades of the males are far apart from each other, while the grades of the female are more
compressed or clustered together. Thus the measure of the center of the distribution is of little help in describing
and comparing these two sets of data. By getting the average distance of each item from the center of the
distribution, the group can be described more completely and, likewise, similarities and differences can be easily
identified.

DIFFERENT MEASURES OF VARIABILITY

RANGE
The range refers to the difference between the highest and the lowest score. If the highest score is 25
and the lowest score in a distribution is 10, then the range is equal to 25-10 = 15. This range is specifically
called as the exclusive range. If the difference between the exact lower limit of the lowest score and the exact
upper limit of the highest score is solved, the result is called the inclusive range. Using the same example, we
get 16 from 25.5-9.5. The inclusive range is simply determined by adding 1 to the exclusive range. Generally,
the exclusive range is used for ungrouped data while the inclusive range is used for grouped data.
The range is the easiest and simplest to determine among the measures of variability because it depends
only on the pair of extreme values. However, it is also the most unstable because its value easily fluctuates with
the change in either of the highest or lowest scores. It is also considered the most unreliable because it does not
give the dispersion or spread of the scores in between the two extreme values. A more reliable measure should
involve all the values in a distribution to provide us an adequate spread of all the scores from the average.

Example 1. Find the range of the grades in Math of the two groups of students in the preceding example.
Male: 100-60 = 40
Female: 83-79=4
The range of grades of the male group is 40 while that of female group is 4. This shows that the grades of the
males are scattered while the grades of the female group are close to each other. It shows further that females
are more homogeneous than the males in their math ability.

AVERAGE OR MEAN ABSOLUTE DEVIATION (MAD)

Mean Absolute Deviation is the average of the summation of the absolute deviation of each observation
from the mean.
x is a value or score from the raw data
Σ/x−𝑥̅ /
MAD= 𝑁 , where 𝑥̅ is the mean
N is the total number of scores

Example 1. A. Find the mean absolute deviation of the male group in example 1.
Solution: The mean of the male group is 81.
Score Mean /x-𝑥̅ /
x 𝑥̅
GE 4- Mathematics in the Modern World La Carlota City College

70 81 11
95 81 14
60 81 21
80 81 1
100 81 19
Σx=405 Σ/x-𝑥̅ / =66
Σ/x−𝑥̅ / 66
MAD= = = 13.2
𝑁 5

B. Find the mean absolute deviation of the female group.

Solution: The mean of the female group is 81.


Score Mean /x-𝑥̅ /
x 𝑥̅
82 81 1
80 81 1
83 81 2
81 81 0
79 81 2
Σx=405 Σ/x-𝑥̅ /
=6

Σ/x−𝑥̅ / 6
MAD= 𝑁 = 5 = 1.2
The male group has a MAD of 13.2 while the female group has 1.2. These results confirm our earlier findings
using the range that the female group is more homogeneous than the male group.

VARIANCE

Variance is the average of the squared deviation from the mean. Formulas for finding the variance for
ungrouped data are shown below:

Σ/x−𝑥̅ /2 Σ/x−𝑥̅ /2
Population Variance : Vp= Sample variance Vs=
𝑁 𝑛−1

STANDARD DEVIATION

Standard Deviation is the square root of the average deviation from the mean, or simply the square root
of the variance.
Σ/x−𝑥̅ /2 Σ/x−𝑥̅ /2
Population Standard Deviation: SDp= √ Sample Standard Deviation: SDs = √
N n−1
GE 4- Mathematics in the Modern World La Carlota City College

Example 4. Find the variance and the standard deviation of male groups in Example 1.

Male group:
x 𝑥̅ x-𝑥̅ (x-𝑥̅ )2
70 81 -11 121
95 81 14 196
60 81 -21 441
80 81 -1 1
100 81 19 361
Σx=405 Σ (x-𝑥̅ )2= 1120

a. Treating the data as population, the variance and the standard deviation is
1120 1120
Vp= 5 = 224 square units Sample variance Vs= 4 = 280 square units

SDp= √224 = 14.97


SDs= √280 = 16.73

• Find the variance and the standard deviation of the female group.
An equivalent formula gives another way of computing the standard deviation. It does away with
computing for the mean and the deviation of the scores. We can use the equivalent formula below.

2 (Σx)2
Σx −
SDs = √ n − 1n

Male Group:
x X2
70 4900
95 9025
60 3600
80 6400
100 10,000
Σx=405 Σx2= 33925

2 (405)2
2 − (Σx) 164025

SDs = √Σx n
= √33925− 5
= √
33925−
5
= √
33925 − 32805
= √
1120
= √280 =
n−1 5−1 4 4 4

16.73 units
variance = 280, standard deviation = 16.73
GE 4- Mathematics in the Modern World La Carlota City College

STANDARD DEVIATION AND VARIANCE OF GROUPED DATA

Formula:
Σf(x−𝑥̅ ) 2
Vp = Σf (x-𝑥̅ )2 / N and SDp =√ N

̅̅̅ 2
Σf(x−𝑥)
Vs= Σf (x-x)2 / n-1 and SDs = √ n−1

Example: The grouped data which gives the wages per day of the laborers in a certain construction.

Wages No. of Mean


Laborers X fX 𝑥̅ x-𝑥̅ (x-𝑥̅ )2 f(x-𝑥̅ )2
(f)
80-84 8 82 656 71.74 10.26 105.2676 842.1408
75-79 12 77 924 71.74 5.26 27.6676 332.0112
70-74 16 72 1152 71.74 0.26 0.0676 1.0816
65-69 13 67 871 71.74 -4.74 22.4676 292.0788
60-64 9 62 558 71.74 -9.74 94.8676 853.8084
Σf= fx=4161 2321.1208
58
Mean = Σfx/ N =4161/58=71.74

Population Variance Population Standard Deviation


Σf(x−𝑥̅ ) 2
Vp = Σf (x-𝑥̅ )2 / N and SDp = √ = √40.02 = 𝟔. 𝟑𝟐𝟔
N
= 2321.1208/ 58 = 40.02

Sample Variance Sample Standard Deviation


Σfx−𝑥̅ 2
Vs= Σf (x-𝑥̅ )2 / n-1 and SDs = √ = √40.72 = 𝟔. 𝟑𝟖
n−1
= 2321.1208 / 57 = 40.72
GE 4- Mathematics in the Modern World La Carlota City College

Activity 3:
1. A group of second-year Business Administration students took the qualifying examination for admission to the
course Bachelor of Science in Accounting. The results are as follows:
64 50 59 73 88 75 78 52 75 53
54 95 68 58 71 57 76 37 49 90
50 49 90 87 42 84 68 56 70 60
44 61 71 66 74 71 39 34 68 65
60 75 74 85 65 46 64 77 77 78
a. Prepare a frequency distribution using 7 classes starting with 34.
b. Include the columns of % relative frequency, less than cumulative frequency, and greater than cumulative
frequency.

2. The data below is the frequency distribution of the Intelligence Quotient (IQ) of 200 students taking up
Behavioral Psychology and Business Statistics at a certain university.
Classes frequency (f)
131-135 7
126-130 10
121-125 14
116-120 35
111-115 42
106-110 21
101-105 18
96-100 29
91-95 10
86-90 14
N=200
Find the following:
1. Class interval size
2. Lower class boundary of the highest class interval.
3. Number of students whose IQ is greater than 115.5
4. Number of students whose IQ is less than 96
5. Lowest upper class limit
6. Upper class boundary of the 3rd class interval
7. Upper class limit of the 5th class interval
8. Number of students whose IQ is between 126-130
9. Highest upper class limit
10. Class mark of the 4th class of the highest frequency

3. The following are the test scores obtained by III-1 students in Statistics. Compute the mean using the:
a. Classmark formula
b. Coded formula
c. What is the average score obtained by the students?
Class Interval f x fx <cf
20-24 4
25-29 6
30-34 7
35-39 10
40-44 5
45-49 8
N=
GE 4- Mathematics in the Modern World La Carlota City College

4. The following data give the time (in minutes) taken to commute from home to school for 20 students of LCCC.
10 50 64 33 48 5 11 23 37 26
26 32 17 7 13 19 29 43 21 22

Find the mean, median, mode, Q1, D5, P70, SDp, Vp, SK, KU

5. Below is the frequency distribution showing the result of an IQ test of a group of students in a certain college.
Classes frequency(f)
80-85 2
86-91 8
92-97 19
98-103 21
104-109 25
110-115 52
116-121 12
122-127 11

1. Determine the following using the frequency distribution above,


a. size of the class interval f. modal class
b. median class g. lower limit of the modal class
c. lower boundary of the median class h. number of students with an IQ greater than 115.5
d. midpoint of the median class i. midpoint of the modal class
e. upper limit of the median class

2. Compute the value of the following:


a. Mean b. Mode c. Q2 d. P75 e. D3

6. The NSAT scores of 12 students in a certain college were taken and are shown below.
86, 95, 84, 87, 91, 90, 99, 84, 83, 88, 96, 92
Determine:
a. Mean b. Median c. Q3 d. P40
7. The IQ’s of 5 members of the family are 108, 112, 127, 118 and 113.
Find the
a. Range b. Variance c. Standard Deviation
8. Below is a frequency distribution showing the result of an IQ test of a group of students in a certain college.
Classes f
80-85 2
86-91 8
92-97 19
98-103 21
104-109 25
110-115 52
116-121 12
122-127 11
Find:
a. Range b. Mean Absolute Deviation
c. Population Variance d. Population Standard Deviation

You might also like