Republic of the Philippines
Province of Cotabato
Municipality of Makilala
MAKILALA INSTITUTE OF SCIENCE AND TECHNOLOGY
Makilala, Cotabato
Public Administration Department
ELEMENTARY STATISTICS
(Course)
Course Code : Stat 111 Instructor : Bernadeth L. Bañ
ados Course Title : Elementary Statistics Duration : 1 week
Credit Unit :3 Class Code : sx6r7g7
I. LEARNING OUTCOMES: At the end of this module, you are expected to:
construct and interpret frequency tables, bar graphs, line graph, pie chart,
frequency histogram, frequency polygon, cumulative frequency ogive, relative
frequency and stem-and-leaf plots;
interpret tabular and graphical data;
solve mean, median and mode for grouped and ungrouped; and
describe variables using appropriate measures of central tendency.
II. TOPIC/SUBJECT MATTER
a. Organization and Presentation of Data
b. Measures of Central Tendency
III. REFERENCES:
Parreño, E and Jimenez, R. (2014) Basic Statistics (2 nd edition).C&E Publishing,
Inc.
IV. COURSE CONTENTs
Organization and Presentation of Data
Collating the Data – after collecting data, they must be recorded in a format that is
useful for your purpose.
Organization of Data
1. Alphabetical arrangement is a kind of arrangement makes any item easy to locate
among the many collected items.
2. Ranked form is arranging of the data in a numerical order either from highest to
lowest or vice versa.
3. Chronological arrangement is the arrangement of data according to the order of their
occurrence, especially when data were gathered at different times or the researcher
wants to see the trend of the progress.
Presentation of data
An effective presentation of data is necessary in any investigation. Data which have
been collected and organized well but not presented clearly would be of little use. In this
premise, there is really a need to effectively present the data so that the readers will
have a clear picture of various relationships.
Three Different Ways of Presenting Data
a. Textual presentation uses both text and figures to convey statistical information. This
is commonly found in statistical reports in magazines and newspapers. The readers are
directed to pay particular attention on specific data such as comparisons, contrasts,
syntheses, generalizations or findings. When this method employed alone elicits
boredom. It is a weak means of presenting the quantitative comparisons or relations
among quantitative or numerical data attractively and interestingly.
Example:
The Faculty Strength of DOSTCOST
The college has 49 full time faculty members, 4 are Baccalaureate degree holders, 1 has
over 40 masters' degree units, 7 have earned 30 to 39 masters' degree units, 17 have earned 20
to 29 masters' degree units, 2 have 10 to 19 masters' degree units, 6 have earned less than 10
masters' degree holders, with 1 having 54 Ph.D. units and 2 with 9 Ed. D. units each and one with
Ph. D. Degree.
b. Tabular presentation makes use of tables in which classifies data are arranged
systematically and orderly by rows and columns with headings and descriptions. Tables
are brief, easy to comprehend, and more convenient to readers than textual
presentation.
Example:
Table 1. Cases of Dengue Grouped According to Sex and Religion in Region D, 2025
Sex Religion A Religion B Total
Male 308 132 440
Female 361 240 601
Total 669 372 1041
When to Use Tables
You want a simple way of presenting data
Only some data are significant
Actual data are important than trends
There is a need to emphasize details
Your reader might want to interpret data for himself; should be understood
without reading the text.
Length of Tables
Avoid tables with only one column of data-combine similar or related data in one
table
Break long tables into shorter tables or related data
Look for other related data to combine with data from very short tables – 2
columns and 2 rows
Guidelines for Making Table Title
The Table Title should tell what the table is all about
Make the table title as short as possible but be specific
Stem and Leaf Plots
It is a special table where each data value is split into a “stem” (the first digit or digits)
and a “leaf” usually (the last digit). Like in these examples:
Example 1. Sam got his friends to do a long jump and got these results:
2.3, 2.5, 2.5, 2.7, 2.8 3.2, 3.6, 3.6, 4.5, 5.0
And here is the stem-and-leaf plot: Stem “2” Leaf “3” means 2.3
Stem Leaf Note:
Say what the stem and leaf mean {Stem
“2” Leaf “3” means 2.3)
2 35578
3 266
4 5
5 0
Example 2: For instance, suppose you have the following list of values: 12, 13, 21, 27,
33, 34, 35, 37, 40, 40, 41.
Stem Leaf The "stem" is the left-hand column which contains the tens digits.
The "leaves" are the lists in the right-hand column, showing all the
1 2 3 ones digits for each of the tens, twenties, thirties, and forties. As
you can see, the original values can still be determined; you can
2 1 7
tell, from that bottom leaf, that the three values in the forties
3 3 4 5 7 were 40, 40, and 41.
4 0 0 1
Frequency Distribution Table
A frequency distribution table is a chart that summarizes values and
their frequency. It is a useful way to organize data if you have a list of numbers that
represent the frequency of a certain outcome in a sample.
*Steps on how to construct a frequency distribution table*
1. Estimate the number of class intervals (k)
k = 1 + 3 log n, where n is the number of population
2. Determine the range, r = highest value – lowest value.
3. Obtain the class size, c = range / k.
4. Set the lowest value as the first lower limit and get the upper limit which is equal to first lower
limit +
class size – 1.
5. Do the same process again until you reach the last limit that includes the highest value from
the
data.
Example 1. Construct a frequency distribution for the following data :
11 19 11 15 16 10 16 16 15 17 10 27
21 11 13 21 10 16 11 19 24 12 22 13
19 13 18 20 21 11 19 15 11 25 29 23
16 23 10 17 11 27 16 24 12 21 13 12
26 15 11 14 10 12 11 15 18 12 20 13
Solution:
1. Determine the value of k = 1 + 3 log n where n = 60. Log 60 = 1.77815125, k = 1 + 3
(1.7781512)
k = 1 + 5.3344536
k=6 Therefore, 6 is the estimated number of classes in these data.
2. r = Highest – lowest =
= 29-10
= 19
3. Class size = 19/6 = 3.16 0r 3
Class limits Frequency
28-30 1
25-27 4
22-24 5
19-21 10
16-18 10
13-15 11
10-12 19
Total, n 60
4. Set the lowest value as the first lower limit. The first upper limit will computed as
Lower limit + class size – 1 = 10 + 3 – 1 = 12.
5. Do the same process again until you reach the last limit that includes the highest value from
the
data.
*For the Frequency, count the number of values from the given data that are included
for each class interval.
c . Graphical presentation provides the reader with a picture of the significant
relationship of the facts presented. The graph can summarize and show visually what
would have been expressed in so many words in a clear and appealing way. The most
common graphs are bar graph, pie graph, and pictogram. Limitations of this method are
as follows:
* graphs are not precise as the tables;
* graphs need more skills and time to prepare; and
* graphs can only be made after data have been shown in tabular from.
Several Methods of Graphing
a. Bar Graph
Bar graphs are usually applied to compare data and to determine which class or interval
is more common or frequently appears in the text. Rectangular figures or bars are used
to show variations in frequencies of observations.
How to Make Bar Graph in Excel
1st : Present the data first in tabular form.
Table 2 Test Scores of Students in Physics Exam
A B C D E F
1 12 7 8 15 8 10
2 56 78 65 45 77 60
3 15 23 44 25 44 65
4 14 15 10 9 5 10
5 34 45 28 33 30 41
2nd: Highlight the data in the table then click “Insert”, “Scatter” . The Switch row is for
your desired variable in x and y axis.
This must be your output.
b. Line Graph
To show trends and increase in sales, improvement
of scores, rise or fall of temperature of patients,
enrolment of students in certain courses, and
comparison of population per year, the line graph is
more appropriate to be used than the bar graph.
c. Pie Chart
Pie chart is helpful in comparing different parts of a
whole at a glance. A circular chart is divided into
sectors and the size of each sector is proportional to
the quantity it represents. Consider the following
example of the monthly expenses of a Filipino family.
d. Frequency Histogram
For grouped data, the frequeny histogram is one of the graph styles which can be
applied. The frequency is represented by points in the vertical axis and the class
intervals in the horizontal axis. The ordered pair of points in the vertical and horizontal
axes is plotted by placing bars in the graph area. Each bar represents the class interval
with its corresponding frequency.
e. Frequency Polygon
Unlike the frequency histogram where bars drawn side by side are used, points
connected by line segments are utilized in the frequency polygon. It looks like the
ordinary line graph except for the labels in the horizontal axis whixh are class intervals.
f. Cumulative Frequency Ogive
In a frequency distribution table, add another column which will be headed as <CF (less
than cumulative frequency). Each entry now is obtained by accumulating the frequencies
starting from the frequency of the interval containing the lowest score up to the interval
containing the highest score.
Table 3. Age of patients of Hospital X, June 2019
Age in Years f Class Boundary <CF
60-68 1 59.5 – 68.5 35+1=3
6
51-59 2 50.5 – 59.5 33+2=3
5
42-50 3 41.5 – 50.5 30+3=3
3
33-41 2 32.5 – 41.5 28+2=3
0
24-32 20 23.5 – 32.5 8+20=2
Start at the bottom frequency
8 found in the 2nd column.
15-23 5 14.5 – 23.5 3+5=8
6-14 3 5.5 – 14.5 3
The graph of the frequency distribution versus the <CF is called the < cumulative
frequency ogive.
6-14 15-23 24-32 33-41 42-50 51-59
g. Relative Frequency
This is also known as the percentage frequency. Using Table 3, add another column for
relative frequency.
Table 4. Frequency Distribution of Patients According to Age
Age in f Class <CF Relative Frequency
Years Boundary
60-68 1 59.5 – 68.5 36 (1/ 36)x 100 = 2.78
51-59 2 50.5 – 59.5 35 (2/ 36)x 100 = 5.56
42-50 3 41.5 – 50.5 33 (3/ 36)x 100 = 8.33
33-41 2 32.5 – 41.5 30 (2/ 36)x 100 = 5.56
24-32 20 23.5 – 32.5 28 (20/ 36)x 100 = 55.56
15-23 5 14.5 – 23.5 8 (5/ 36)x 100 = 13.89
6-14 3 5.5 – 14.5 3 (3/ 36)x 100 = 8.33
36
6-14 15-23 24-32 33-41 42-50 51-59
Measure of Central Tendency
Finding the average grade of studentss, average daily sales of a department store, or
average salary per month of employees in a company etc., are common procedures
done by each one of us. Without knowing it, we have been doing statistical tasks quite
often. Average is another term for central tendency or central location.
Three Types of Measures of Central Tendency
1. Mean to find the mean or arithmetic mean, add all the items or observations then
divide the sum by the total number of observations. In symbol,
µ=
∑ x the symbol “µ” is read as mu. This is use to represent population
n
mean.
x=
∑x the symbol “ x ” is read as x bar. This is use to represent sample mean.
n
Characteristics of mean:
most reliable and accurate measure
sum of the differences of the scores from the mean is 0.
greatly affected by extreme values
Example: Scores of the students: 45, 59, 58, 62,49
µ=
∑ x = 273 = 54.6
n 5
Weighted mean is also a mean where some values contribute more than the others. In
symbol, µ =
∑ fx (f = frequency, x = raw data ).
n
Example: There are 1000 notebooks sold at P10.00 each, 500 notebooks at P20.00 each
and 500 at P25.00 each and 100 notebooks at P30.00 each.
Price f fx
µ=
10 ∑ fx = 1000
35,500
=
10,000
16.90 average price
n 2,100
20 500 10,000
25 500 12,500
30 100 3,000
Total 2,100 35,500
2. Median is a the midpoint of an array of numbers or observations. The symbol Md
n+1
represents the median. In symbol, Md = ( ¿.
2
Characteristics of median:
arranged from lowest to highest
not affected by extreme values
known as the 50th percentile which means 50% falls above it and 50% falls below it
separates the distribution into two equal parts
Case 1. 3, 6, 7, 9 (There are four values given. The number of data is even.)
n+1 4+ 1
Md = ( ¿= = 2.5 , average of the 2nd and 3rd data
2 2
Note: Data must be arranged from highest to lowest. 6 is the second data and 7 is
the third data.
6+7 13
2
= 2 = 6.5 is the median
Case 2. 3, 6, 7, 9,12
n+1 5+1
Md = ( ¿= = 3 , the median is the third value which is 7
2 2
3. Mode is the value with the highest frequency.
Characteristics of mode:
used when the quickest and easiest value is wanted
the only value that can be calculated or determined for nominal data
Example 1: A store owner wants to know from his employee what size of ladies’ shoes
is mostly bought by the women for the whole month of June. Using the sales
report for June, the employee told the owner the most common size is 6.5
since this size appears the most number of times in the list. What the
employee was telling the owner is the MODE of the shoe sizes which is 6.5.
Example 2: What is the mode of the scores of the students in Biostatistics test? The
scores are as follows: 12, 13, 12, 11, 10, 20, 24, 25, 10, 22, 20, 13,16, 18,
20, 20, 20, 20.
The mode is 20 since it is the score which appeared the most
number of times.
Example 3: What is the mode of the following heights of freshman college students?
152 cm, 166 cm, 176 cm, 150 cm, 168 cm, 155 cm, and 149 cm.
In this case, there is no mode.
Example 4: What is the mode of the following heights of freshman college students?
152 cm, 166 cm, 166 cm, 176 cm, 176 cm, 150 cm, 168 cm, 155 cm, 166 cm,
176 cm and 149 cm.
There are two modes, 166 cm and 176 cm.
Computation of Mean for Grouped Data
n
The mean for grouped data is computed by the formula µ =
∑ xi f i , where xi stands for
i=1
n
the class mark and fi is the corresponding frequency.
Example:
Table 4.1. Age of Patients of Hospital X, June 2004
Age in Years fi (frequency) Xi (class mark) fx
60-68 1 64 64
51-59 2 55 110
42-50 3 46 138
33-41 2 37 74
24-32 20 28 560
15-23 5 19 95
6-14 3 10 30
∑ f i= 36 ∑ xi f i =
1,071
Steps:
1. Find the class mark by adding upper and lower class intervals then divide the sum by 2.
Example: class mark 64 on the first row ( 68 + 60 )/ 2= 128 / =64
2. Multiply the class mark by its corresponding frequency to obtain x ifi.
3. Get the sum of xifi.
4. Substitute the values of ∑ x i f i and n in the formula.
n
µ=
∑ xi f i
i=1
n
Computation of Median for Grouped Data
The formula used in getting the median for grouped data in the form of frequency
distribution is
N
−cf
Md = L + ( 2 )W
f med
where
L = exact lower limit of the class interval containing the median
N/2 = one half of the total number of cases
cf = cumulative total of the frequencies up to, but not including the frequency of the
median
fmed = frequency of the class interval containing the median
W= width of the median class interval
N= total number of frequencies
Example:
Age Frequency f Cumulative
Frequency
460-469 16 176
450-459 22 160
440-449 19 138
430-439 14 119
420-429 15 105
410-419 20 90
400-409 13 70
390-399 18 57
380-389 14 39
370-379 15 25
360-369 10 10
N or f = 176
N=176, for you to identify the median class interval use this formula N/2.
So, 176/2 = 88. Referring to the cumulative frequency column, look for the interval containing
the frequency 88. In this example the 6 th class interval which is 410 – 419 has a cumulative
frequency of 90, below it is 70 meaning the 6 th class interval contains a frequency starting
from 71-90.
Steps:
1. Find the value of N/2, which will identify the location of the median term (median class
interval).
2. Find the value of L.
3. Make a cumulative frequency(cf) by adding successively, starting from the bottom, the
individual frequencies.
4. Determine the cumulative frequency (cf) by adding successively, starting from the
bottom, the individual frequencies.
5. Solve W by subtracting the lower limit of the median class interval from its upper class
interval. W = 419 - 410 = 10
N
−cf
6. Substitute the values in the formula. Md = L + ( 2 )W
f med
176
−70
= 409.5 + ( 2 ) 10
20
88−70
= 409.5 + ( ) 10
20
18
Computation of Mode = 409.5 + ( ) 10
for Grouped20 Data
The mode in grouped data is the class mark or midpoint of the class interval with the highest
frequency. This class interval is known as modal class. The mode obtained in this manner is
called a crude mode because it is just a rough approximation of the actual mode. To compute the
formula below is used.
d1 = 22 – 19 , d2 = 22 – 16
d1
Mo = L + ( ) W where di
d 1+ d 2 Mo = L + ( )W
d 1+ d 2
L = exact lower limit of the class interval containing the modal class 22−19
= 449.5 + ( )
d1 = difference between the frequency of the modal class ( 22−19 ) +(22−16)
and the next class lower in value 10
d2 = difference between the frequency of the modal class
3
and the next class higher in value = 449.5 + ( ) 10
W= width of the modal class interval 3+6
(upper limit – lower limit of the modal class) 3
= 449.5 + ( ) 10
9