Statistics Lecture
Statistics Lecture
Statistics is a branch of Mathematics that deals with the scientific collection, organization, presentation,
analysis, and interpretation of numerical data in order to obtain useful and meaningful information.
Example: The NSO conducts surveys to determine the average age, income, and other
characteristics of the Filipino people.
Examples: (a) It only took 2 years for Facebook to reach a market audience of 50 million.
Therefore, Facebook is very influential; (b) A teacher determines if the mean score of the
students in her class in a Mathematics test is significantly related to their scores in a Science
test.
POPULATION VS SAMPLE
In both qualitative and quantitative research, studies are conducted on a selected population. In
many cases only a portion of the total population is selected to study. The selected portion is
called a sample.
POPULATION - the totality of all the elements or persons for which one has interest at a particular
time; a group that includes all the cases (individuals, objects, or groups) in which the researcher is
interested.
Examples: the faculty members of AECMIS, the athletes of UST, Facebook users worldwide
SAMPLE - a subset of the population chosen to represent the population in a statistical analysis;
a small part or portion of a population that represents their characteristics or traits.
Examples: selected students from grade 7, selected employees in a company, monthly budget of
10 families
SAMPLING - refers to the method or process of selecting the members of sample from a population
Interview – direct method of gathering data because this requires a face-to-face inquiry with the
respondent. The interview method is done when a person solicits information from another person. The
person gathering the data is called the interviewer while the person supplying the data is the
interviewee.
Questionnaire – indirect method of gathering data because this makes use of written questions to be
answered by the respondent.
Observation – makes use of the different human senses in gathering information. The data are
gathered either individually or collectively. The person who gathers the data is called investigator while
the person being observed is called the subject.
Registration or Census – it requires the enactment of law to take effect because it needs the
participation of a large, if not the entire population.
Experimentation – usually conducted in laboratories where specimens are subjected to some aspects
of control to find out cause and effect relationships.
ORGANIZATION OF DATA
Frequency Distribution Table for Ungrouped Data
Table is a systematic arrangement of data usually arranged in rows or columns for ready reference of
information. It has a title and appropriate headings for the columns and rows.
Frequency Distribution is a table that lists each data point and its frequency. Data is often
described as ungrouped or grouped.
Ungrouped data is data given as individual data points. It is also called raw data. Examples 1 and 2 are
ungrouped data.
Grouped data is data given in intervals and has been organized into groups (into frequency distribution).
Examples 3 and 4 are grouped data (see page 4)
Example 1: Construct a frequency distribution table for the test scores of 20 students in Statistics.
6 9 3 8 9 6 7 9 5 7 8
5 10 4 5 4 8 10 9 8
Frequency Distribution Table
Scores Tally Frequency
3 I 1
4 II 2
5 III 3
6 II 2
7 II 2
8 IIII 4
9 IIII 4
10 II 2
N=20
Before we go in too deep with grouped frequency distribution, let us first define some terms. Don’t
worry! This is as easy as the frequency distribution for ungrouped data.
𝑹𝒂𝒏𝒈𝒆 = 𝑯𝑺
− 𝑳𝑺
Range is the difference between the highest score and the lowest score.
This term is important to take note because we always began our table by computing the range.
When the range of the data is large, the data must be grouped into classes that are more than one unit in
𝑹𝒂𝒏𝒈𝒆
interval. We get the interval by dividing the range by the desired class size.
𝑰𝒏𝒕𝒆𝒓𝒗𝒂𝒍 (𝒊) =
𝑪𝒍𝒂𝒔𝒔
𝒔𝒊𝒛𝒆
Since we will not make a grouped frequency distribution from scratch, we may just subtract two
consecutive lower class limits. The process will be shown later in the examples.
In the class interval, we have two values in there – the lower class limit and the upper class limit.
The lower class limit is the smallest data value that can be included in the class.
The upper class limit is the largest data value that can be included in the class.
We also have to find the class mark which is the midpoint of the class. It can be found by getting
𝒍𝒊𝒎𝒊𝒕
𝑪𝒍𝒂𝒔𝒔 𝒎𝒂𝒓𝒌 =
𝟐
We have already encountered tally and frequency on the previous part, but we will still be
encountering them again.
Example 3: Solve for the interval, and class mark using the table below.
Daily allowance of Blue 2
Class
Class
Interval Tally Frequency
Mark
(Scores )
281 - 309 | 1
252 – 280 || 2
223 – 251 || 2
194 – 222 || 2
165 – 193 | 1
136 – 164 |||| 4
107 – 135 |||| 4
78 – 106 |||| - |||| - | 11
49 – 77 |||| - |||| - || 12
20 – 48 |||| - |||| - | 11
Total 50
To solve for the interval, we will be subtracting two consecutive lower class limits or upper class limits.
We may use either of the two because we will still get the same answer.
𝒄𝒍𝒂𝒔𝒔 𝒍𝒊𝒎𝒊𝒕
𝟐
You may try to solve for the 4th to 10th class interval’s class mark to check if the answers given are
correct.
Example 4: Solve for the interval, and class mark using the table below.
The Time (in minutes) Green Block Spend in Studying and Accomplishing Daily
Class
Class
Interval Tally Frequency
Mark
(Scores)
|||| - |||| - |||| - |||| - |||| - |||| - |||| - |||| -
374 – 459 53
|||| - |||| - |||
288 – 373 |||| - |||| - |||| - |||| - |||| - | 26
|||| - |||| - |||| - |||| - |||| - |||| - |||| - |||| -
202 – 287 |||| - |||| - |||| - |||| - |||| - |||| - |||| - |||| - 99
|||| - |||| - |||| - ||||
116 – 201 |||| - |||| - |||| 15
30 – 115 |||| - || 7
Total 200
Solution:
Interval:
Class Boundaries or true limits of a given score is the score plus or minus one-half or
0.5 of the unit of measure or the place value of the given score.
It has also two values; the Lower Boundary and Upper Boundary.
Lower boundary can be obtained by subtracting 0.5 from the lower limit of the given interval.
Let us have the table of the daily allowance of Blue 2 in example 3. Identify the boundaries in each class
interval.
To solve for the boundaries, we subtract 0.5 from the lower limit and add 0.5 to the upper limit for each
class interval for lower boundary and upper boundary respectively.
Example 4. Prepare a grouped frequency distribution on the scores of 40 students in a Math quiz. Shown below
are their scores. Use 10 intervals.
86 83 81 81 86 91 79 82 81 87 87 83 82
72 73 78 87 71 94 91 90 82 85 88 71 99
76 96 80 89 98 89 82 80 75 90 72 83 74
85
Solution:
Procedure:
PRESENTATION OF DATA
[Link]
One way we can present the data is by using a bar graph. A bar graph uses the height or length of
the bar to represent how often a particular category was observed.
To draw a bar graph, plot the frequency against the categories as shown below.
[Link]
We can also present the data using a pie chart. A pie chart is a circular graph that shows how the
categories are distributed.
To draw a pie chart, assign one sector of the circle to each category. The angle of each sector should
be proportional to the relative frequency in that category. Since one full circle has 360∘, we can find
the angle for each category by multiplying the relative frequency by 360∘.
[Link]
You may color your pie chart to distinguish the categories from each other.
[Link]
There are data that are best presented by charts aside from bar graphs and pie charts. These data
are quantitative in nature; thus, a simple bar graph or pie chart would not suffice. Look closely at the
data below.
[Link]
The data set is a time series since the variable is recorded over time. Time series is best presented
on a line graph.
A line graph uses dots and lines to discern a pattern or trend that could continue into the future. This
pattern could then be used to predict future events.
To draw a line chart, plot the time (horizontal) against the observed phenomena (vertical) then
connect them using lines as shown below.
[Link]
. A histogram resembles a bar graph and shows how often measurements fall in a particular class or
subinterval.
Let us consider the following data gathered by Rocky about the age of his neighbors:
[Link]
1. Choose the number of classes, usually between 5 and 20. The more data you have the more
classes you should use.
2. Calculate the class width by dividing the range (largest value – smallest value) by the number
range highest value−lowest value
of classes. This will form the class interval. w= =
desired no . of classes desired no . of classes
We round up the class width to a whole number.
3. Record the number of scores (frequency) that fall under each class interval.
4. Construct a statistical table containing the classes, class interval, frequencies, and relative
frequencies.
5. Construct the histogram like a bar graph. Class intervals are on the horizontal and frequencies
are on the vertical. Note that there should be no space in between the bars unless there is no
value that falls under a certain interval.
Thus, if we start from 1, the first class interval must be until 10. The second interval must be 11 - 20,
and so on, until we create 8 intervals. Note that the smallest value must be included in the first class
and the highest value must be included in the last class.
Next, we construct a statistical table and record the frequency of data that fall under the classes.
[Link]
[Link]
An ogive uses cumulative frequencies instead of frequency. Cumulative frequency is used to
determine the number of observations that lie above (or below) a particular value in a data set. Ogive
allows us to quickly estimate the number of observations that are less than or equal to a particular
value. Ogives are similar to histogram in nature; the difference is instead of bar graphs we use points
and lines instead. To get the cumulative frequency of a class, we must add its frequency and all the
preceding frequencies. The first frequency is also the first cumulative frequency.
[Link]
Note that the cumulative frequency of the last class is the total number of subjects. In the given
example, the last cumulative frequency is 20 and the total number of subjects is also 20.
Now we can construct an ogive by plotting the points and connecting them with lines using the
cumulative frequencies on the vertical axis and the class intervals on the horizontal axis. Each point
must be plotted at the upper limit of the class boundary.
[Link]
MEASURES OF CENTRAL TENDENCY
A measure of central tendency is a numerical descriptive measure which locates the center of the
distribution. It is a single, central value that summarizes a set of numerical data.
Three of the most commonly used measures of central tendency are the mean, the median, and the
mode.
Mean, Median, and Mode of Ungrouped Data
1. Mean – is the score obtained if all the scores are “evened out”. It is affected by extreme values. It is the
sum of the data divided by the number of data.
2. Median – is the middle value of the set of data, provided that the data are arranged in an array. An array is
an arrangement of values in decreasing or increasing order. The median is not affected by extreme values
because its position in an ordered list stays the same.
3. Mode – is the number which occurred most frequently in the data. It is useful if the interest is to know the
most common value.
Example 1: The scores of 10 students in Mr. Garcia’s class are 64, 59, 63, 60, 65, 66, 66, 66, 61, 70. Find the
following: a. mean b. median c. mode
Solution:
a. For ungrouped data, the mean has the following formula.
❑
∑❑x
= ❑
n
Where: = mean
❑
∑
❑
❑ x = sum of the measurements or values
n = total number of measurements or values
= 64 or Mean=64
b. To find the median of ungrouped data, we first arrange the values or measurements in an array (either
increasing or decreasing order), and then get the middle value.
Array: 59, 60, 61, 63, 64, 65, 66, 66, 66, 70
We can get the middle value after arranging the data in an array. Since the number of scores is even
(10), there are two middlemost scores; namely 64 and 65. To get the median, we get the mean of the two
middlemost scores. Thus, the median is
64 +65
Md=
2
Md=64.5
c. The mode is the number which occurred most frequently. Since three students got a score of 66, then
66 is the modal score.
Mo=66
Example 2: The picture card numbers of different countries which Joshua has collected are 97, 98, 98, 104, 37,
86, 95, 93, and 105. Find the mean, median, and the mode.
Solution:
97+98+ 98+104+ 37+86+ 95+93+105
a. =
9
= 90.33 or Mean=90.33
b. Arrange the data in array: 36, 86, 93,95, 97, 98, 98, 104, 105
Since the number of scores is odd (9), there is only one middlemost score. Hence the median is 97.
Md=97
c. The number which occurred most frequently is 98. Then, the mode is 98. Hence, the given set of data
is unimodal.
Mo=98
Example 3: The following are the scores of students in the English test: 40, 40, 44, 45, 45, 47, 47. Find the
mean, median, and the mode.
Solution:
40+ 40+44 +45+ 45+ 47+ 47
a. =
7
= 44 or Mean=44
b. Arrange the data in array: 40, 40, 44, 45, 45, 47, 47
Md=45
c. There are three numbers which occurred most frequently, 40, 45 and 47. Then, the modes are 40, 45
and 47. Hence, the given set of data is trimodal.
Mo=40 , 45 ,∧47
The grouped data are data which have been arranged in a frequency distribution table. To compute for the
measures of central tendency of grouped data, we use the following formulas:
a. MEAN:
¿
∑ f xm
n
Where:
X = measurement or score
f = frequency
x m = class mark
n = total frequency
b. MEDIAN:
( )
n
−¿ cf
2
Md=x LB + i
fm
Where:
x LB = lower class boundary of the median class
n = total frequency
¿ cf = less than cumulative frequency above the median class
i = size of the class interval
f m = frequency of the median class
c. MODE:
Mo=x LB +
( d1
)
d 1 +d 2
i
Where:
x LB = lower class boundary of the modal class
d 1 = difference between the frequency in the modal class and the frequency in
the preceding class interval
d 2 = difference between the frequency in the modal class and the frequency in
the succeeding class interval
i = size of the class interval
Example 1: Below is a frequency distribution of scores of 30 students in Mathematics. Compute for the mean,
median, and the mode.
1030
¿
30
¿ 34.33
b. MEDIAN:
( )
n
−¿ cf
2
Md=x LB + i
fm
n
Since =15, the class interval 31 – 35 is the median class because it contains one-half of the total
2
frequency in the ¿ cf column.
Class Interval Frequency (f) Class Mark ( x m ¿ fx m Cumulative frequency
(¿ cf ¿
6 – 10 1 8 8 1
11 – 15 1 13 13 2
16 – 20 2 18 36 4
21 – 25 2 23 46 6
26 – 30 3 28 84 9
31 – 35 6 33 198 15 Middle class
36 – 40 7 38 266 22
41 – 45 4 43 172 26
46 – 50 2 48 96 28
51 – 55 1 53 53 29
56 – 60 1 58 58 30
Notice that the lower class boundary of the median class is 30.5, the frequency of the median
class is 6 and the less than cumulative frequency (¿ cf ¿ above the median class is 9. Substituting the
values in the formula, we have:
Md=30.5+ ( 15−9
6 )
5
Md=30.5+ ( 66 )5
Md=30.5+5
Md=35.5
c. MODE:
Mo=x LB +
( d1
)
d 1 +d 2
i
Notice that the class intervals are arranged from lowest to highest. The modal class is the class
interval 36 – 40 because it has the highest frequency.
Class Interval Frequency (f) Class Mark ( x m ¿ fx m Cumulative frequency
(¿ cf ¿
6 – 10 1 8 8 1
11 – 15 1 13 13 2
16 – 20 2 18 36 4
21 – 25 2 23 46 6
26 – 30 3 28 84 9
31 – 35 6 33 198 15
Modal class
36 – 40 7 38 266 22
41 – 45 4 43 172 26
46 – 50 2 48 96 28
51 – 55 1 53 53 29
56 – 60 1 58 58 30
The lower class boundary of the modal class is 35.5, the frequency above (preceding) the modal
class is 6, and the frequency below (succeeding) the modal class is 4. Substituting these values in the
formula, we have:
Mo=35.5+
( (7−6)+(7−4
7−6
))
5
Mo=35.5+ ( 1+31 )5
Mo=35.5+ ( 14 )5
Mo=35.5+1.25
Mo=36.75
Example 2: The table below shows the age distribution of the contestants in a raffle draw sponsored by a
popular noontime game show.
Class Interval Frequency (f) Class Mark ( x m ¿ fx m Cumulative frequency
(¿ cf ¿
7–9 1 8 8 1
10 – 12 1 11 11 2
13 – 15 4 14 56 6
16 – 18 6 17 102 12
19 – 21 14 20 280 26
22 – 24 10 23 230 36
25 – 27 3 26 78 39
28 – 30 1 29 79 40
❑
i=3 n=40 ∑
❑
❑ fx m =794
a. MEAN:
❑
∑
❑
❑ fx m
¿
n
794
¿
40
¿ 19.85
b. MEDIAN:
( )
n
−¿ cf
2
Md=x LB + i
fm
n
Since =20, the class interval 19 – 21 is the median class because it contains one-half of the total
2
frequency in the ¿ cf column (26 contains 20).
Notice that the lower class boundary of the median class is 18.5, the frequency of the median class is 14
and the less than cumulative frequency (¿ cf ¿ above the median class is 12. Substituting the values in the
formula, we have:
Md=18.5+ ( 20−12
14
3 )
Md=18.5+ ( 148 )3
Md=18.5+ 1.71
Md=20.21
d. MODE:
Mo=x LB +
( d1
)
d 1 +d 2
i
Notice that the class intervals are arranged from lowest to highest. The modal class is the class
interval 36 – 40 because it has the highest frequency.
Class Interval Frequency (f) Class Mark ( x m ¿ fx m Cumulative frequency
(¿ cf ¿
7–9 1 8 8 1
10 – 12 1 11 11 2
13 – 15 4 14 56 6
16 – 18 6 17 102 12
19 – 21 14 20 280 26 Modal class
22 – 24 10 23 230 36
25 – 27 3 26 78 39
28 – 30 1 29 79 40
The lower class boundary of the modal class is 18.5, the frequency above (preceding) the modal class
is 6, and the frequency below (succeeding) the modal class is 10. Substituting these values in the formula,
we have:
Mo=18.5+
( 14−6
(14−6)+(14−10)
3
)
Mo=18.5+ ( 8+48 ) 3
Mo=18.5+ ( 128 )3
Mo=18.5+2
Mo=20.5
MEASURES OF VARIABILITY
Measures of dispersion or variability refer to the spread of the values about the mean. These are
important quantities used by statisticians in evaluation. Smaller dispersion of scores arising from the comparison
often indicates more consistency and more reliability.
The most commonly used measures of dispersion are the range, the average deviation, the standard
deviation, and the variance.
R=H–L
Examples:
Comparing the two wages, you will note that wages of workers of factory B have a higher range than
wages of workers of factory A. These ranges tell us that the wages of workers of factory B are more scattered than
the wages of workers of factory A.
4. the range of the set of scores is 29 and the lowest score is 18, what is the highest score?
Given: R= 29, L= 18 , H= ?
R=H–L 29 = H – 18
29 + 18 = H H=
47 The highest score is 47.
The dispersion of a set of data about the average of these data is the average deviation or mean
deviation. To compute the average deviation of an ungrouped data, we use the formula:
[Link] the absolute difference between each score and the mean.
To find the range, variance and standard deviation of grouped data, take note of the following:
Range = Upper Class Boundary of the Highest Interval – Lower Class Boundary of the Lowest
Interval
Illustrative Example:
Scores Frequency
46 – 50 1
41 – 45 10
36 – 40 10
31 – 35 16
26 – 30 9
21 – 25 4
Solutions:
Range = Upper Class Boundary of the Highest Interval – Lower Class Boundary of the Lowest
Interval
Variance is the average of the square deviation from the mean. For large quantities, the variance is
computed using frequency distribution with columns for the midpoint value, the product of the frequency and
midpoint value for each interval; the deviation and its square; and the product of the frequency and the squared
deviation.
Illustrative Example:
Scores Frequency
46 – 50 1
41 – 45 10
36 – 40 10
31 – 35 16
26 – 30 9
21 – 25 4
Nature of Sensitivity to
Usability Nature of Data
Computation Other Data
The Easily affected Most widely used average and Measure for interval
computational or by extreme subject for further mathematical scales such as
calculated values or outliers computation scores, grades,
Mean
average temperature, and
The mean is the best measure for population
symmetrical data distributions.
Rank or positional May or may not Less widely used than the mean Measure for ordinal
average be affected by and but can be subjected to a scales such as test
extreme values few mathematical computations scores, salary
Median
The median is helpful when
describing asymmetrical/skewed
data.
Inspectional or May or may not Rarely used and cannot be Measure for
commercial be affected by mathematically manipulated nominal scales such
Mode average an introduction as number of a
of other data certain brand of
commodities