Sampling Techniques in Research Methods
Sampling Techniques in Research Methods
SAMPLING TECHNIQUES
Introduction
Inferential statistics basically concerned with making conclusions and predictions about the
population based on the examined samples.
A good survey research paper relies on the precision of the methods and procedures of
conducting the study. This includes reliability of the selected subjects or respondents of the study, the
validity of the information gathered out of the distributed questionnaires, and the accuracy of the
measurements used in answering the research questions and other observations. A study, which conducted
in the entire population assures us of 100% reliability since the responses are obtained from all members
of the population. This means that the data was collected by a complete enumeration method or the so-
called census taking. However, it is impossible for many types of research to conduct a survey to all
members of the population especially if the population size is infinite or finite but very large. To
minimize the time and cost involved in conducting the survey to a large population, it has been accepted
that the information about the population will be based only from a small portion of the population called
[Link] the other hand, considering only the responses of a small portion of the population may result
into some possible biases due to improper selection of samples and errors due to the manner of measuring
the desired observations since the selected sample may not have equally represented the characteristics of
the entire population.
Hence, it is very important to consider the method used in selecting the sample and the statistics
involved in the sampling distribution such as the mean, standard deviation, proportion, standard error of
mean, and standard error of the [Link] module deals with the different sampling technique . It
covers the different types of random and non-random sampling.
Learning Outcomes
2. Explain the importance and uses of the different types of random and non-random techniques
sampling.
LEARNING CONTENTS
For practical reasons such as to economize on time, money and effort, it is not necessary
for the researcher to examine every member of the population to get data or information about
the population. Cost and time constraints will prohibit one from understanding a study of the
entire population. At any rate, all that he needs to do is to draw sample units systematically or at
random. If sampling is done in this way, we can validly infer conclusions about the entire
population from our sample.
Often, when we talk about picking things at random, we mean picking things without bias
or any predetermined choice. In a TV Program, for example, participants may be asked to pick a
prize at random. Listing all the possible prizes and assigning a number to each prize can do this.
The numbers are then written on pieces of paper and placed in a box or container where they are
shaken thoroughly. When the participants draw a number from the box, he would have drawn a
number at random. The practice of awarding prizes through the “raffle” system follows the
principle of random sampling.
Random Sampling is the method of selecting a sample size (n) from a universe (N) such
that each member of the population has an equal chance of being included in the sample and all
possible combinations of size (n) have an equal chance of being selected as the sample.
A prerequisite for the randomness of the selection is a complete listing of the population.
Thus, prior to the actual picking of sample units, the complete listing of the enumeration of the
population has to be undertaken. This phase provides the researcher the list where he would
randomly pick his sample units.
2. Jimmy
3. Jeany
3. Jojie
4. Jenilyn
5. Jerome
6. Jessica
Mr. Gonzales drops the folded pieces of paper into a container and picks from the
container the lucky names who are to receive the tickets. The first four names to be picked are
the lucky ones. If the first name to be drawn is Jerome, he gets the first ticket. If the second name
to be drawn is Jojie, she gets the second ticket. If the third name to be drawn is Jerry, he gets the
third ticket. And if the fourth name he picks is Jessica, then she gets the fourth ticket. This
simple illustration is one way of using the lottery method. Drawing prizes through the raffle
system follows the principle of random sampling.
2. Table of Random Numbers- The use of the Table of Random Numbers is another technique
of random sampling wherein the selection of each member of the population is left adequately to
chance. Every member of the population has an equal chance of being chosen.
To illustrate this technique, refer to Table 2 The use of the Table of Random Numbers
can be illustrated as follows.
a) Direct Selection Method. This method is used when there are only few samples units
to be selected. From the table, the sample units can be selected right away.
EXAMPLE 1.9
Taking the case of Mr. Gonzales, let us find out how the direct selection method differs
from the lottery method. As mentioned before, Mr. Gonzales has four complimentary tickets to a
movie and he wants to distribute them to his seven children without being accused of favoritism.
For fairness’ sake, he can use the Table of Random Numbers.
First, we have to enumerate the children and assign a number to each one of them. The
number to each one of them. The numbers serves as codes, such corresponding to one name.
These are the numbers which will be the basis for the objective choice of winners.
Numbers Name
1 Jerry
2 Jimmy
3 Jeany
4 Jojie
5 Jenilyn
6 Jerome
7 Jessica
Techno Box
Aside from the Table of Random Numbers, random numbers between 0 and 1 can also be
generated using a scientific calculator or computer.
2. In a Microsoft Excel Worksheet, simply enter the function = RND ( ) on any blank cell.
N =123
Table 2
Table of Random Numbers
Referring to the Random Table, any of the four columns may be used. Since Mr.
Gonzales has seven children, we consider 7 as the total population (N) which is only a one-digit
number. So, we can only take the first digit of any of the columns from the table. Let us take the
numbers under the first column one at a time. What do we get?
613238 – The first digit is 6, so, Jerome gets the first ticket.
990056- Not included, since N=7
068217- Also not included.
655118- Skip. Jerome already has this number.
253755- The first digit is 2, so, Jimmy gets the second ticket.
715080- The first digit is 7, so, Jessica gets the third ticket.
48228 - The first digit is 4, so, Jojie gets the fourth and last ticket.
Thus, using the Table of Random Numbers, the lucky four children who can get the
tickets to the movie are Jerome, Jimmy, Jessica and Jojie. As illustrated, numbers are used
through the direct selection method. Similar patterns can be utilized.
b) The Remainder Method. Often, however, we cannot rely only on the direct selection
method in picking the sample units. This is due to the fact that when the number picked from the
Random Table is greater than N (for example, N=120 while the number picked from the Table is
greater than 120) the number that will have to be skipped. If we rely on the direct selection
method, there will be many number skipped. If this happen, we may have gone through the entire
problem; the remainder method is used whenever the direct selection method cannot be
applied. Whenever the direct method can be used, however, we should proceed to use it. There
are two ways of conducting the remainder method.
(1) When the number taken from the table of Random numbers is subtracted from the upper limit
within this number falls, the remainder is the sample unit.
(2) When the upper limit of the set is subtracted from the number taken from the Random table
and yields a number equal or less than N, the remainder is the sample unit.
Example 1
Using the last column of the Table of Random Numbers, pick 10 sample units from a
population of 123. Thus, N= 123 and n=10
The first step is to determine how many sets of 123 items we can construct. Since our N
is a three-digit number, we can only make use of numbers not exceeding three digits, that is, 1-
999. The following will show us how to proceed with this method,
SETS
First set (123x1 = 123); 1 – 123
Second Set (123x2 = 246); 124-246
Third Set (123x3 = 369); 247-369
Fourth Set (123x4 = 492); 370-492
Fifth Set (123x5 = 615); 493-615
Sixth Set (123x6 = 738); 616-738
Seventh Set (123x7 = 861); 739-861
Eight Set (123x8 = 984); 862-984
Ninth Set (123x9 = 1,170); 985-999- rejected, since this set is not complete, i.e., does
not contain 123 items.
We can use the first eight sets because each number in these sets in given eight equal
chances of being selected. We do not use the ninth set because it does not include the total 123
items to make a complete set.
The Table contains numbers, each with six digits. But our N (which is 123) is only a
three-digit number. The numbers we can, therefore, use must no exceed three digits, hence from
1 to 999. We have just distributed these numbers into seven sets of 123. The ninth set was
rejected for not containing the entire 123 elements.
The first number in our Table (using the last column) is 475358. We only need to use the
first three digits of this number. Thus, we pick 475 as our number.
The second step is to look for the set, which contains the number 475.
The third step is to subtract this number from the upper limit, which contains this
number.
When we subtract this number, i.e. 492, the remainder is the number of our first sample
unit.
The 10 sample units we need are, therefore, picked in the following manner.
615
-543
72 – our fourth sample unit
The number we get from the Table is 039218. Since this case the direct selection method
can be applied, we use it to pick number 039 as our fifth sample unit.
The number registered is the table is 008326. Since the direct selection method can also
be applied in this case, we use it to pick the number 8 as our sixth sample unit.
066624 – Thus we pick number 66 as our seventh sample unit, again through the direct
selection method.
246
-132
114 – our eighth sample unit
861
-857
4 – our ninth sample unit
738
-656
82 – our tenth sample unit
Example 2
Using the last column of the Table of Random Numbers, pick 10 sample units from a
population of 123. Thus, N = 123 and n = 10.
Since our population is a three-digit number, we can only make use of the numbers from
1 to 99 because they do not exceed three digits. We are able to construct eight sets with 123
elements each:
The first step is to look for the first number (using the last column) from the Table of
Random Numbers, the number is 475358, but we pick out 475 because we only need the first
three digits.
The second step is to look for the upper limit of the set which when subtracted from the
number obtained from the Table, i.e., 475, would give us a number equal or less than our
population N.
The third step to subtract this upper limit from the number obtained from the Table.
039218 – Number 39 is our fifth sample unit through the direct selection method.
008326 - number 8 is our sixth sample unit through the direct selection method.
Using the third column of the Table of Random Numbers, pick 100 sample units from a
population of 1,150. Thus, N=1,150 and n=10.
Since N= 1,150, a four-digit number, the numbers that can be picked from the Table of
Random Numbers will be 1-9,999.
As we have stated, a fundamental rule in random sampling is to give all numbers equal
chances of being chosen. How can we make use of numbers 1-9,999 so that each number has an
equal chance of being selected? We have to construct several sets of not more than four digits
(since our population is a four-digit number) with each set containing 1,150 elements (N= 1,150.
The first stew sample units under the third column of the Table of Random Numbers are
the following:
796267
492824
329285 The first four digit numbers are the numbers that can be used.
240271
434045
827812
The following procedures will illustrate how we can use these numbers in picking our
sample units.
FIRST METHOD
When the number taken from the Random Table is subtracted from the upper limit within
which this number falls, the remainder is the sample unit.
3450
-3292
158 - our third sample unit
3450
-2402
1048 - our fourth sample unit
The process goes on until we pick 100 sample units.
SECOND METHOD
When the number taken from the Random Table is subtracted from the upper limit within
this number falls, the remainder is the sample unit.
When the upper limit of the set is subtracted from the number taken from the Roman
Table and yields a number equal or less than N, the remainder is the sample unit.
The process goes until the desired sample units are picked out.
This method involves selecting every nth element of a series representing the population.
A complete listing is required in this method. Under this system, the sample units may be picked
in the following manner, where, for example,
N=100
n=10
The value of n may be obtained by dividing the total number of elements in the
population by the desired sample size. Thus,
N=100
n 10= 10th
The sample units would, therefore, be the persons holding the following numbers; 10, 20,
30, 40, 50, 60, 70, 80, 90 and 100 at 10 sample units. Variation may be added by choosing a
random start. Let us take 10 pieces of paper and number them 1-10. We put these pieces of paper
in a box or container and shake them thoroughly. If the number 7 is picked as the random start,
the 10 sample units should be: 7, 17, 27, 37, 47, 57, 67, 77, 87, 97, = 10 sample units.
4. Stratified Sampling - This is a random sampling technique in which the population is divided
into non-overlapping subpopulations called [Link] this method, the population is first divided
into group-based on homogeneity -in order to avoid to possibility of drawing samples whose
members come only from one stratum. In stratified sampling , the distribution of sampling units
is proportionate to the total number of units of each stratum. The bigger the population , the more
sample units are drawn, the less population , the less sample units. That is why this method is
often called stratified proportional sampling. In contrast , simple random takes the proportion to
chance
The formula below is used to determine the sample sizes for proportional allocation.
ni ⟦ NN ⟧n for i= 1, 2, 3….
n = is the total size of the stratified random sample
N = total population
N1 = number of 1st stratum elements
N2 = number of 2nd stratum elements
N3 = number of 3rd stratum elements
Example 1
At a BSHRM in ISU Cauayan , the students may be classified according to the following
scheme:
Table 1.3
Proportional Sample Size Allocation
3rd 210 60
2nd 325 93
1st 346 99
Total 1,000 286
1st year
r = ns
N
r = 286(346)
1000
r= 98.956
r= 99
r = ns
N
r = 286(325)
1000
r= 92.95
r= 93
r = ns
N
r = 286(210)
1000
r= 60.06
r= 60
Note: To determine the appropriate sample size without resorting to your subjective decision,
you may use the Slovin’s formula.
n= N
1+ N e²
Where:
n= sample size
N= population size
e= 0.05 (the sampling error)
n= N
1+ N e²
n= 1000
1 + 1,000(0.05)2
n = 1,000
1 + 1,000(0.0025)
n = 1,000
1 + 2.5
n = 1,000
3.5
n = 285.71 0r 286
r = ns
N
4th Year
r = ns
N
r = 286(119)
1000
r= 34.034 or 34
Using the proportional allocation to select an appropriate sample of size n=286, how
large a sample must be taken for each stratum.
Solutions: n= 286, N1= 119, N2=210, N3= 325, N4=346 and n=100
Computations:
Example 2
A certain corporation has always been plagued with the problem of fast worker turnover.
Every year, a sizeable number of its workers leave the corporation to look for work elsewhere.
Since it gives adequate salaries to its employees compare to the wages of other corporations of
the same class and size, executives of the corporation were perplexed by the situation and wanted
to find out the reason for the high rate of worker turnover by doing a survey.
The corporation is divided into different departments : Administrative, manufacturing,
finance, warehousing, and research and development. The total number of workers is 490. How
will we pick the sample units if the corporation wants 20% of the employees to be surveyed?
STEP 2. Determine the percent share of each stratum with respect to the total population, as
shown in table 3.
STEP 3. Management wants 20% of the population, which is 98% of the sample units, to be
surveyed. The third step will be to multiply the percent share of each department in elation to
the total population by 98, which is the number of sample units to be surveyed. To get the actual
number of sample units for each department . This is shown in Table 4
Another example, in studying the investment habits of working parents in a given region,
it is much cheaper to interview and collect data from individuals living close together in
randomly selected clusters for provinces or cities than to select a sample random sample for the
entire region. Considering geographic areas as clusters, this kind of sampling is also called Area
Sampling.
Example 1.
K=N = 50
N = 12
K= 4.667 =5
1 2 3 4 5 6 7 8 9 10
11 12 13 14 15 16 17 18 19 20
21 22 23 24 25 26 27 28 29 30
31 32 33 34 35 36 37 38 39 40
41 42 43 44 45 46 47 48 49 50
Choose every 5th unit. Thus, if the random start r= 9th unit, then the sample comprises of
students numbers.
Stage 1. Enumerate all 16 regions of the Philippines including their respective municipalities.
Stage 2. From the 16 regions, select three at random. This can be done through Table of Randon
Numbers or by lottery.
Stage 3. We have now three regions out of 16 we had before. From these three regions, we select
two provinces from each region. The process of selection should also be done at random.
Stage 4. With two provinces from the three regions, we have our list right now six provinces. We
enumerate all the municipalities and cities in all of these provinces. The process of selection
should be done at random.
Thus, we have six provinces from where we can pick our sample. From each of these
provinces we select three municipalities or cities. In the final analysis, we will only survey 18
municipalities/cities ( 3 x 6 =18) in our study.
Non- Random Sampling does not involve random selection of sample elements. Some
elements of the population do not have a chance to be included in the sample.
Try this
4. In a town of 25,000 people, it is but practical to choose a representative group for certain
research project. How do you get the random sample?
6. An educator wants to find out information about the families represented in a school system
and decides to pick a random sample of children from those registered in the schools for survey.
He will ask such questions as: How large is your family? How much education do your parents
have? And so on. Is there anything wrong with this sampling plan.
After the data have been presented in tabular or graphical form, the researcher must be able to
describe them in terms of a single number. This number which gives a summary or the
characteristics of a given set of data is called a measure of central tendency or measure of central
location.
You often read the word “average” in books newspapers, magazines, and journals. You hear this
word on radio and television. For example, the average salary of employees of the different
establishments of Isabela is P15,000. The batting average of a baseball player is 375. The
average IQ of students in a particular school is [Link] this module, you will learn that the term
average could imply the mean, median, or the mode, which are referred to as the measures of
central tendency.
The most commonly used measures of central tendency are the mean or arithmetic average, the
median and the mode. Such measures of central tendency can be computed in two data which are
not yet organized or arranged in frequency distribution otherwise such is called group data.
Objectives
Ungrouped data or Raw Data. This refer to data which are not yet organized or arranged into
frequency distribution.
[Link]
Characteristics of Mean
1. An interval statistics
2. A calculated average
3. Value is determined every case in the distribution
4. Affected by extreme values
5. Can be subjected to numerous mathematical computations.
6. Most widely use
7. Represents average quantity
A. Simple Arithmetic [Link] arithmetic mean or simply mean ( popularly called the
average) is the sum of the separate scores or measures divided by the number of the
scores. T
Arithmetic [Link] arithmetic mean or simply mean ( popularly called the average) is the sum
of the separate scores or measures divided by the number of the scores.
Where, ∑ (the uppercase Greek letter sigma), X refers to summation, refers to the individual
value and n is the number of observations in the sample (sample size)
Example 1.
The ages in years of 10 sales clerk of a mall are:32, 41, 28, 54, 35, 26, 23, 33, 38, [Link] is the
mean age of these employees?
Solution :
= 350 / 10
= 35 years
Example 2
Solution :
= (2 + 4 + 6 + 8 + 10 + 12 + 14 + 16) / 8
= 72 / 8
=9
Example 3.
John worked in a food chain as part time for 4 hours, 5 hours and 3 hours respectively on three
consecutive days. How many hours does he work daily on an average?
Solution :
= (4 + 5 + 3) / 3
= 12 / 3
= 4 hours
Thus, we can say that John work for 4 hours daily on an average.
[Link] Mean
Weighted mean is calculated when certain values in a data set are more important than
the others. A weight wi is attached to each of the values xi to reflect this importance.
Consider the proper weights assigned to the observed values according to their relative
importance.
In formula,
n
Xw =
∑ WiXi
i=1
N
Where
W¡ = weight of each item
X¡ = value of each item
X = mean
∑ = means of the sum of
Example 1.
Suppose that a marketing firm conducts a survey of 1,000 households to determine the average
number of TVs each household owns. The data show a large number of households with two or
three TVs and a smaller number with one or four. Every household in the sample has at least one
TV and no household has more than four.
Solution:
Here’s the sample data for the survey:
1 73
2 378
3 459
4 90
As many of the values in this data set are repeated multiple times, you can easily compute the
sample mean as a weighted mean. Follow these steps to calculate the weighted arithmetic mean:
Step 1: Assign a weight to each value in the dataset:
x1=1,w1=73
x2=2,w2=378
x3=3,w3=459
x4=4,w4=90
Step 2: Compute the numerator of the weighted mean formula.
Multiply each sample by its weight and then add the products together:
∑4i=1wixi=w1x1+w2x2+w3x3+w4x4
= (1)(73)+(2)(378)+(3)(459)+(4)(90)
=2566
Step 3: Now, compute the denominator of the weighted mean formula by adding the weights
together.
∑4i=1wi=w1+w2+w3+w4
= 73 + 378 + 459 + 90
=1000
Step 4: Divide the numerator by the denominator
∑4i=1wixi∑4i=1wi
=25661000
=2.566
The mean number of TVs per household in this sample is 2.566.
Example 2
A man bought 10 liters of premium gasoline at 11.50 per liter, 12 at P12.01 per liter and at
P11.78 per liter from three different gasoline stations. Find the average price per liter.
Solution
WiXi+W 2 X 2+W 3 X 3
Xw=
W 1+ W 2+ W 3
To compute for the arithmetic mean of grouped data, we need to determine the midpoint
of each class interval. The mean assumed that the class mark of each is the average value of all
items falling in that class.
1. Long Method:
∑ fiXi
X= i=1
n
Where;
X = sample mean
X¡= the class midpoint or class mark
f¡= the corresponding frequencies
n = total number of items.
X=
( )
∑ fidi
i=1
n
i
Where:
X = sample mean
XA= assumed mean (usually the value of the highest midpoint)
Di = deviation of the values from the assumed mean,
Di = ( Xi−i Xa)
f¡ = the corresponding frequency
I = class interval or class size
N= number of items
Example 1
The following are the distribution of length of service in years of 50 employees of
United laboratories Inc.
1-5 5
6-10 7
11-15 12
16-20 13
21-25 6
26-30 4
31-35 3
SOLUTION:
1-5 5 3 15 -3 -15
6-10 7 8 56 -2 -14
11-15 12 13 156 -1 -12
16-20 13 18 234 0 0
21-25 6 23 138 1 6
26-30 4 28 112 2 8
31-35 3 33 99 3 9
Long Method:
n
x =∑
f ¡ x¡
i=1
n
810
x=
50
= 16.2
Short Method
x = XA + ( ∑
f ¡d ¡
i=1 )i
n
−18
= 18+ ( )5
50
= 18- 1.8
X= 16.2
where x is the midpoint of the interval, f is the frequency for the interval, fx is the product of the
midpoint times the frequency, and n is the number of values.
For example, if 8 is the midpoint of a class interval and there are ten measurements in the
interval, fx = 10(8) = 80, the sum of the ten measurements in the interval.
Σ fx denotes the sum of all the products in all class intervals. Dividing that sum by the number of
measurements yields the sample mean for grouped data.
Therefore, the average price of items sold was about $15.19. The value may not be the exact
mean for the data, because the actual values are not always known for grouped data.
MEDIAN
The median is the value of the middle item after arranging the data in an ascending or
descending order
Characteristics of Median
1. An ordinal statistic
2. A rank or position average
3. Value is determined by scores near the middle of the distribution.
4. Not affected by extreme values
5. Can be subjected to only a few mathematical computations.
6. Less widely used than mean
7. Represents typical score.
Examples
1. Compute for the median from the following set of scores; 6, 4, 5, 3 and 2.
Solution. Arrange the set of items and compute for the median.
2, 3, 4, 5, 6,
The median is 4, which is the middle item.
5, 6, 8, 12, 13, 15
8+12 20
Mdn = = = 10
2 2
n
−cfB
Mdn = Lme + ( 2 )i
Fme
Where;
Mdn = median
Lme = lower class boundary of median class
N = total number of observations
fMe = frequency of the median class
cfB = cumulative frequency preceding the median class
I = class size or size of the class interval
nth
value
2
Example
Solution:
n 50
First solve for = = 25th
2 2
Then, locate where the 25th item is equal or nearest but not greater than the value in the
less than cumulative frequency (<cf) distribution.
n
−cfB
Me = Lme + ( 2 )i
Fme
50
−24
= 15.5 + ( 2 )5
13
25−24
= 15.5 + 5
13
= 15.5 + 1.667
Me = 17.167 / 17.2
MODE
The mode is the simplest measure of central tendency. It is easily identified merely
looking at an ungrouped set of scores or data and locating the score or data which occurs most
frequently.
Characteristics of Mode
1. A nominal statistics
2. An inspection average
3. The most frequent occurring value
4. Usually occurs near the center of distribution.
5. Cannot be manipulated mathematically.
6. Some distribution have more than one mode.
7. Rarely used
8. Most “ popular” score
When all values appears with the same frequency, the mode does not exist. However, for
some sets of data there may be several values appearing with the greatest frequency in which
case we have more than one mode. When two modes are present in a distribution, such is called a
bimodal, trimodal for three modes and so forth.
Example 1
Find the mode of the following set of items.
4, 7, 11, 6, 4, 8, 3, 5, 2, 9
Example 2
Determine the mode of the following distribution
12, 15, 21, 9, 6, 15, 11, 8, 9, 5
Example 3
Find the mode of the following sets of data.
12,14,17,18,19,20,26, 26,19,20, 29
Solution
The modes are 19, 20, 26, trimodal.
Example 4
Find the mode of the following data.
6, 5, 11, 21, 8, 16, 19, 20, 7, 9
Solution: The mode does not exist, because all the time are nor repeated.
A distribution with only one mode is said to be UNIMODAL. Distribution with two
modes is BIMODAL, with three modes is TRIMODAL while a distribution with two or more
modes is described as MULTIMODAL.
The mode of a grouped data is defined as the midpoint of the class interval with the
highest frequency (modal class). The mode obtained in this manner is called a crude mode,
because it is just a rough approximation of the actual mode. So, to determine the true mode, we
use the formula
d1
Mo = Lmo + ( )i
d 1+ d 2
Where:
Mo = mode
Lmo = lower class boundary of modal class
D1 = difference between the frequency of the modal class and the frequency
of the class next lower in value.
D2 = difference between frequency of the modal class
i = class interval or class with
The modal class is the class interval with the highest frequency.
Example
Find the mode of the frequency distribution of length of service in years 50 employees of United
Laboratories Inc. (from example 1)
Solution
Length of Service Number of Employees CB
CL f
1-5 5 0.5 – 5.5
6-10 7 5.5 – 10.5
11-15 12 10.5 – 15.5
Mo Class» 16-20 13 15.5 – 20.5
21-25 6 20.5 – 25
26- 30 4 25.5 – 30.5
31 – 35 3 30.5 – 35.5
n=50
The modal class is the 16-20, since it is the class interval with highest frequency.
d1
Mo = Lmo + ( )i
d 1+ d 2
1
= 15.5 + ( )5
1+ 7
5
= 15.5 +
8
Mo = 16.125
Example
Solution
CL f x fx
21-30 8 25.5 204
31-40 11 35.5 390.5
41-50 15 45.5 682.5
51-60 18 55.5 999
61-70 20 65.5 1310
71-80 12 75.5 906
81-90 9 85.5 769.5
91-100 7 95.5 668.5
N=100 ∑fx=5,930
Mean x
Long Method
n
x =∑
fixi
i=1
n
5930
=
100
x=59.3
Short Method
CL f X CB <cf D fd
21-30 8 25.5 20.5-30.5 8 -4 -32
31-40 11 35.5 30.5-40.5 19 -3 -33
41-50 15 45.5 40.5-50.5 34 -2 -30
51-60 18 55.5 50.5-60.5 52 -1 -18
61-70 20 65.5 60.5-70.5 72 0 0
71-80 12 75.5 70.5-80.5 84 1 12
81-90 9 85.5 80.5-90.5 93 2 18
91-100 7 95.5 90.5-100.5 100 3
N=100 ∑fd=-62
x = XA + ( ∑
f ¡d ¡
i=1 )i
n
= 65.5 + (-6.2)
= 65.5 – 6.2
= 59.3
MEDIAN (Mdn)
n
−cfB
Mdn= Lme+( 2 )i
Fme
n 100
Median class = = = 50th
2 2
50−34
= 50.5 + ( )10
18
= 50.5 +8.89
Mdn = 59.39 59.4
MODE (Mo)
ADVANTAGES
The mean uses every value in the data and hence is a good representative of the data. The irony
in this is that most of the times this value never appears in the raw data.
Repeated samples drawn from the same population tend to have similar means. The mean is
therefore the measure of central tendency that best resists the fluctuation between different
samples.[6]
It is closely related to standard deviation, the most common measure of dispersion.
Go to:
DISADVANTAGES
The important disadvantage of mean is that it is sensitive to extreme values/outliers, especially
when the sample size is small.[7] Therefore, it is not an appropriate measure of central tendency
for skewed distribution.[8]
Mean cannot be calculated for nominal or nonnominal ordinal data. Even though mean can be
calculated for numerical ordinal data, many times it does not give a meaningful value, e.g. stage
of cancer.
References
Petrie A, Sabin C. Medical statistics at a glance. 3rd ed. Oxford: Wiley-Blackwell;
2009. [Google Scholar]. Retrived from
[Link]
[Link]
tendency#:~:text=Measures%20of%20central%20tendency%20are,values%20shown%20in
%20Table%201.
+
Characteristics of Mean
8. An interval statistics
9. A calculated average
10. Value is determined every case in the distribution
11. Affected by extreme values
12. Can be subjected to numerous mathematical computations.
13. Most widely use
14. Represents average quantity
15. The mode is a nominal statistic which means that is used for nominal data. Its
computation does not depend on the values of the variable nor on their order, but merely
on their frequency of occurrence. It is rarely used with interval , ratio, and ordinal
variables, where means and medians can be calculated.
16. It is usually employed as a simple, inspectional measure which indicates roughly the
center of concentration of a distribution. As such, there is no need to calculate it as
exactly as the median or the mean.
Consideration in the Use of a Measure of Central Tendency
Try this
1. Find the mean , median and mode of the following sets of numbers.
a. 4, 6, 7, 8,8,9,3, 2 , 4, 5,9
b. 89, 52, 47, 30, 37, 79, 89, 90 , 83, 56, 62, 73,
c. 24, 29, 40, 28, 32, 33, 22, 27, 32,33
2. A company pays its 45 employees an average a mean monthly salary of P 12,500; a
second company pays it 52 employees of P 13, 750 ; and the third company pay its 100
employees an average monthly salary of P 11, 950. What is the average salary per
employee in the tree companies.
3. The salaries of the different managers of food chain in Isabela are shown below:
Monthly Salary Number of Manager
50,000- 59,999 3
40,000-49,999 15
30,000-39,999 27
20,000-29,999 12
10,000-19,999 11
Assessment Task
1
2 45 12 44 22 45
3 43 13 76 23 55
4 36 14 47 24 32
5 38 15 64 25 30
6 42 16 62 26 28
7 44 17 35 27 75
8 56 18 29 28 74
9 55 19 36 29 32
10 29 20 40 30 30
4. When do you use the mean, median and mode . Explain each by using an example.
ASSESSMENT TASK
1. A sample of chicken meat from 7 supermarkets produced the following data on their
process.
P 170.00 P153.00 P 140.00 P180.00 P185.00 P164.00
P185.00
Calculate the range, variance, standard deviation and coefficient variation for these data.
2. The following table lists the prices of 5 different brand of oven toaster. Calculate the
range, variance, standard deviation and coefficient variation .
Brand Price
Sharp P1,000.00
National P1,150.00
Samsung P920.00
Philips P 850.00
G.E. P1,250.00
3. The following table gives the 2018 gross sales ( rounded to million of pesos for a sample
of food companies. Find the range, variance, standard deviation,and coefficient variation.
ASSESSMENT TASK
Measures of Variability
1. Compute the range, mean absolute deviation, variance and standard deviation of the daily
wages of the 9 employees : P 450, P350, P 425, P360,P370, P580, P320, P 410, and
P650.
2. The following are wages per day of workers in a certain [Link] the range,
mean absolute deviation, variance and standard deviation.
Wages Number
P410-414 22
415-419 19
420-424 10
425-429 15
430-434 20
435-439 11
440-444 14
445-449 12
450-454 15
Key Terms
Hypothesis
Null Hypothesis
Alternative Hypothesis
Directional Alternative Hypothesis
Non-directional Alternative Hypothesis
Learning Objective
In the chapter, the concept of hypothesis testing shall be discussed and the methods used
in testing hypothesis about the population mean when the population standard deviation is given
and other conditions about hypothesis testing will be developed. In this test, the critical value
method is testing hypothesis will be discussed. Other topics related to hypothesis testing shall
also be discussed.
DEFINITION OF HYPOTHESES
1. The proportion of the consumers who purchased Brand X of female facial wash
exceeds 60%.
2. The main daily allowance of high school students in rural areas is at most ₱150.00.
3. The average lifetime of light bulb manufactured by RL Company will last at most
5,500 hours.
4. There is no significant difference between the monthly expenditures of families in the
rural and urban areas.
5. The average of nicotine content of Y cigarette does not exceed 3.25 milligrams.
The statements above are subject to statistical testing in order to determine the
truthfulness of them. When it failed to reject the hypothesis, then the hypothesis should be
accepted.
1. Null Hypothesis
2. Alternative Hypothesis
When dealing with hypothesis tests, there are four possible outcomes; the two outcomes
lead to incorrect decision and the other two lead to correct decision. The outcomes are described
in the given table.
Table 6.1
Possible Outcomes for a Hypothesis Test
Based from the given table, a researcher commits an error if a true H0. When a researcher
rejected a true H0, he commits a Type I error or alpha error (ɑ). When a researcher accepted a
false H0, he commits a Type II error or beta error (β).
1. Type I error or alpha error (ɑ). A Type I error is committed when the researcher rejects a
null hypothesis when in fact it is true.
2. Type II error or beta error (β). A Type II error is committed when the researcher accepts a
null hypothesis when in fact it is false.
Level of Significance
When a researcher tests the hypotheses, he is not certain that the decisions 100% correct.
However, he is confident at a certain level that the decision is correct, say 99% of the decision he
made is a correct one. The confidence level is 99% or the level of significance is 1%. When the
confidence level is 95%, the level of significance is 5%. On the other hand, when the confidence
level is 90%, the level of significance is 10%. In this case, the higher the confidence level, the
more certain that the decision of rejecting the null hypothesis is correct.
Level if significance is the probability of committing a Type I error or alpha (oc) or the
probability of rejecting the correct null hypothesis.
Power of A Test
Power of a test is the probability of not committing a Type II error or beta (β).
Tests Statistic
The tests statistic is used as a basis for deciding whether to reject or accept the null
hypothesis. The rejection region lies at either left or right tail of the normal curve if one-tailed
test is being used. On the other hand, the rejection lies at both end tails of the normal curve if
two-tailed test will be utilized.
ẋ
Figure 6.1
Rejection Region
When the test statistic lies on the rejection region, then the null hypothesis will be
rejected.
Non-Rejection Region
The non-rejection region is the probability of making a Type I error equals to the level of
significance. Non-rejection region is also known as the acceptance region. When the tests
statistic lies within the non-rejection region, the null hypothesis will be accepted or the critical
value is greater than the computed value of the test statistic. The null hypothesis will be rejected
otherwise it will be accepted.
Critical value
The critical value is a value that separates the non-rejection region and the rejection
region.
Rejection region
Non-rejection
region
ẋ Z= 1.645
Critical Value
Figure 6.2
The use of the one-tailed test or two-tailed test will depend on how the alternative is
formulated. If the alternative hypothesis is expressed in non-directional, utilize the two-tailed
test. However, use the one-tailed test if the alternative hypothesis is directional. In two-tailed
test, the two rejection regions lie at both end tails of the normal curve each part will be half of
the alpha value. If ɑ= 0.05, the area in each end tail is ɑ= 0.025. In one-tailed test, the rejection
region lies either left end tail of the normal curve or right end tail of the normal curve.
ẋ
Figure 6.3
One-Tailed Test
Rejection region
Figure 6.4
Two-Tailed Test
Z = -1.645 ẋ
Figure 6.5
ẋ Z = -1.645
Figure 6.6
Non-rejection
region
Rejection region
Figureẋ 6.7
Z = 2.33
Non-rejection
region
Figure 6.8
Z = -1.96 Z = -1.96
oc = 0.025 oc = 0.025
Figure 6.9
Table 6.2
Critical Value of Z-Test
Type of Test
One-Tailed Test Two-Tailed Test
Level of
Left-Tailed Right-Tailed
Significance
Reject Ho if z ≥ 1.645 Reject Ho if z ≥ 1.96 or
ɑ = 0.05 Reject Ho if z ≤ -1.645
or reject Ho if z ≤ -1.96
Reject Ho if z ≥ 2.575 or
ɑ = 0.01 Reject Ho if z ≤ -2.33 Reject Ho if z ≥ 2.33
reject Ho if z ≤ -2.575
Reject Ho if z ≥ 1.645 or
ɑ = 0.10 Reject Ho if z ≤ -1.28 Reject Ho if z ≥ 1.28
reject Ho if z ≤ -1.645
NOTE: The level of significance usually determines by the statistician or the researcher.
To determine whether to accept or reject the null hypothesis based from a sample data, a
statistician usually follows a certain process. This process is known as hypothesis testing.
Hypothesis testing is type of statistical inference, which examines the claim about a
population based from the information obtained in the random sample.
Sampling distribution is approximately normal when any of the following conditions are
applied.
In this case, z-test will be used when the population deviation is known.
Whereas, utilize the t-test when the population standard deviation is not known in the given
distribution or problem and the number of cases is less than 30.
Using z-test, consider the following assumptions: the distribution is normal; n > 30;
known σ (z-value is the distance from the mean in relation to the standard deviation).
In this section, some of the different cases in testing hypothesis shall be discussed. The
first case is hypothesis testing about means (comparing population mean and sample mean when
n ≥ 30 or n< 30). The second case is testing difference between the last case is hypothesis testing
about two proportions.
CASE I. Hypothesis Testing About Means (Comparing Population and Sample Means)
1. z-test
ẋ−μ
z=
σ
√n
Where:
z is the z-test value
μ is the value of the population mean
ẋ is the sample mean
σ is the population standard deviation
n is the number of case, n ≥ 30
2. T-test
ẋ−μ
t=
s
√n
t is the z-test value
μ is the value of the population mean
ẋ is the sample mean
s is the population standard deviation
n is the number of case, n < 30
Example:
1. PJL Corporation is a company that produces RGC brand of laundry soap that uses a
machine to package 425 grams per pack. Assume that the net weight is normally
distributed with a population standard deviation of 8.5 grams. A researcher randomly
selected 32 packs of RGC brand of laundry soap with net weight of 430 grams. Can he
conclude that the packaging machine function properly? Test the significance at 0.05
level.
Given:
μ = 425 grams
ẋ = 430 grams
σ =8.5 grams
ɑ = 0.05
Step 1. State the null and alternative hypothesis.
Ho : µ = 425 grams or the mean weight is 425 grams.
H1 : µ > 425 grams or the mean weight is greater than 425 grams.
Step 2: Decide the level of significance (ɑ).
ɑ = 0.05
Step 3: Select and compute the appropriate test statistic when it is not stated in the
problem.
Use the z-test because the population standard deviation is given and n =
32. Consider one-tailed test because the H1 is directional alternative hypothesis.
Solution:
ẋ−μ
z=
σ
√n
430−425
z=
8.5
√32
5
z=
1.5026
z = 3.33
Step 4: Compare the value of the test statistic and the critical value obtained from ɑ.
The critical value of z = 1.645 at ɑ = 0.05 and the computed value of z = 3.33.
ẋ z = 2.33 z = 3.33
oc = .05
Figure 6.11
ẋ−μ
t=
s
√n
Given:
ẋ = 430
μ = 445
s = 25.5
n = 15
Step 1. State the null and alternative hypotheses.
Ho : µ = 445 grams
H1 : µ< 445 grams
Step 3. Select and compute the appropriate test statistic when it is not stated in the
problem.
Use the test t-test because the sample standard deviation is given and n =
15. Consider one-tailed test because the H1 is directional alternative hypothesis.
ẋ−μ
t=
s
√n
430−445
t=
25.5
√ 15
−15
t=
6.5841
t = -2.278 or -2.28, get the absolute value
t = 2.28
Step 4. Compare the value of the test statistic and the critical value obtained from ɑ.
The critical value of toc = 0.01 = 2.624, df = 14, and the computed value of t =
2.28.
Step 5. Make decision.
The computed value of t = 2.28, which is less than the critical value of toc =
0.01 = 2.624. Hence, based from the given information it failed to reject the null
hypothesis. Therefore, accept the null hypothesis.
Step 6. Interpret the result.
There is no significant difference between the population mean and the
sample mean.
Where:
t is the t-value
ẋ 1 is the mean of the first sample
ẋ 2 is the mean of the second sample
s12 is the population variance of the first sample
s22 is the population variance of the second sample
n1 is the number of cases of the first sample
n2 is the number of cases of the second sample
LINEAR CORRELATION
Correlation analysis is concerned with the relationship in the changes and movements of two
variables. The relationship has a computed value and may be visually illustrated through the
scatter diagram.
The measure of the degree of relationship or association between two variables may be further
classified as linear and non-linear. When height and weight are plotted in a graph called scatter
diagram, the relationship is a simple or linear correlation. If on the other hand, the graph of the
points is a curve the relationship is said to be non-linear.
Temperature
Volume is constant
Pressure
Height
Weight
Accidents
Road width
Die A no correlation
Die B
Summarizing
-1 ←PerfecNegativeCorrelation
} Some Negative Correlation
0 ---No Correlation
} some positive correlation +1 ---PerfectPositiveCorrelation
7.1 Correlation Methods
There are several ways of measuring the relationship between two variables depending upon the
nature of data. The most frequently encountered measure is the Pearson-Moment Correlation
Coefficient (r).
A. Computation of the Person r from the Deviation from the Means. This is computed using
the formula.
r = ∑xy / √(∑x^2)(∑y^2)
where:
Given:
x y x y x^2 y^2 xy
13 19 -2 -5 4 25 10
14 73 -1 -1 1 1 1
15 25 0 1 0 1 0
16 26 1 2 1 4 2
17 27 2 3 4 9 9
A correlation coefficient of 0.95 is very close to the perfect correlation +1. Hence it is indicate of high positive linear correlation between the two variables.
x y x^2 y^2 xy
13 19 169 361 247
Rank the first set of values giving the highest score a rank of 1.
1.
Do the same with the second set of values.
2.
Get the difference between these two ranks for each individual and record in the D column.
3.
Square each of these difference and get their sum.
4.
Substitute the value in the form.
5.
R = 1 - 6∑D^2 / n(n-1)
Example 3. For the data below, compute the Spearman Rank – Order Correlation.
2 18 27 2 2 0 0
3 15 29 3 1 2 4
4 12 25 4 3.5 .5 .25
5 11 21 5 5 0 0
6 9 19 6 6 0 0
7 7 14 7 7 0 0
8 5 13 9 8 1 1
9 6 12 8 9 -1 1
∑D^2=12.5
Substituting:
p = 1 - 6∑D^2 / n(n-1)
= 89.58 = .90
The value of the Pearson Product –Moment Correlation Coefficient can be interpreted as follow: (according to Garrett).
From the above interpretation an r of .85 is regarded as a high correlation coefficient and r of .48 is regarded as moderate while r of .25 denotes as low correlation
coefficient.
We often think of correlation and causation as going together. This is reasonable because when one thing causes another, the two tend to be associated and therefor
correlated.
However, there can be correlation without causation. Think of this way : The correlation is just a number that reveals whether large values of one variable tend to with
large (or with small) values of the other. The correlation cannot explain why the two are associated. Indeed, the correlation provides no sense of whether the investment
is producing the return or vice versa. The correlation just indicates that the numbers seem to go together in some way.
One possible basis for correlation without causation is that there is some hidden , un observed, third factor that makes one of the variables seem to cause the other
when, iin fact, each is being – caused by the missing variable.
The term spurious correlation refers to a high correlation that is actually due to some third factor. For example, you might find high correlation between hiring new
managers and building new facilities. Are the newly hired managers causing new plant investment? or does the act constructing new building causes new managers to
hired? Probably there is a third factor, namely, high long-term demand for the firms products, that is causing both.
EXERCISE 7.1
The following table show the final grades of ten students in algebra and statistics
1.
Algebra(x) 77 84 68 98 71 87 65 93 80 75
Statistics(y) 74 89 72 95 80 91 72 86 78 82
Solve the Pearson Product-Moment correlation, (r) for the data interpret the result.
4. A group of 5 students took test before and after training and obtained the following score.
Before 19 18 20 22 25
(x)
After(y) 20 22 25 30 28