0% found this document useful (0 votes)
14 views69 pages

Sampling Techniques in Research Methods

This module discusses sampling techniques in inferential statistics, emphasizing the importance of reliable sample selection for accurate population conclusions. It outlines various random sampling methods, such as lottery sampling and the use of random number tables, to ensure unbiased representation of the population. The learning outcomes include solving probability sampling problems and understanding the significance of different sampling techniques.

Uploaded by

santuamerlyann
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views69 pages

Sampling Techniques in Research Methods

This module discusses sampling techniques in inferential statistics, emphasizing the importance of reliable sample selection for accurate population conclusions. It outlines various random sampling methods, such as lottery sampling and the use of random number tables, to ensure unbiased representation of the population. The learning outcomes include solving probability sampling problems and understanding the significance of different sampling techniques.

Uploaded by

santuamerlyann
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

MODULE 7

SAMPLING TECHNIQUES
Introduction

Inferential statistics basically concerned with making conclusions and predictions about the
population based on the examined samples.

A good survey research paper relies on the precision of the methods and procedures of
conducting the study. This includes reliability of the selected subjects or respondents of the study, the
validity of the information gathered out of the distributed questionnaires, and the accuracy of the
measurements used in answering the research questions and other observations. A study, which conducted
in the entire population assures us of 100% reliability since the responses are obtained from all members
of the population. This means that the data was collected by a complete enumeration method or the so-
called census taking. However, it is impossible for many types of research to conduct a survey to all
members of the population especially if the population size is infinite or finite but very large. To
minimize the time and cost involved in conducting the survey to a large population, it has been accepted
that the information about the population will be based only from a small portion of the population called
[Link] the other hand, considering only the responses of a small portion of the population may result
into some possible biases due to improper selection of samples and errors due to the manner of measuring
the desired observations since the selected sample may not have equally represented the characteristics of
the entire population.

Hence, it is very important to consider the method used in selecting the sample and the statistics
involved in the sampling distribution such as the mean, standard deviation, proportion, standard error of
mean, and standard error of the [Link] module deals with the different sampling technique . It
covers the different types of random and non-random sampling.

Learning Outcomes

At the end of this section, expect that you will be able to


1. Solve problems involving probability sampling.

2. Explain the importance and uses of the different types of random and non-random techniques
sampling.
LEARNING CONTENTS

For practical reasons such as to economize on time, money and effort, it is not necessary
for the researcher to examine every member of the population to get data or information about
the population. Cost and time constraints will prohibit one from understanding a study of the
entire population. At any rate, all that he needs to do is to draw sample units systematically or at
random. If sampling is done in this way, we can validly infer conclusions about the entire
population from our sample.

Often, when we talk about picking things at random, we mean picking things without bias
or any predetermined choice. In a TV Program, for example, participants may be asked to pick a
prize at random. Listing all the possible prizes and assigning a number to each prize can do this.
The numbers are then written on pieces of paper and placed in a box or container where they are
shaken thoroughly. When the participants draw a number from the box, he would have drawn a
number at random. The practice of awarding prizes through the “raffle” system follows the
principle of random sampling.

Random Sampling is the method of selecting a sample size (n) from a universe (N) such
that each member of the population has an equal chance of being included in the sample and all
possible combinations of size (n) have an equal chance of being selected as the sample.

A prerequisite for the randomness of the selection is a complete listing of the population.
Thus, prior to the actual picking of sample units, the complete listing of the enumeration of the
population has to be undertaken. This phase provides the researcher the list where he would
randomly pick his sample units.

A. Random Sampling Techniques

There are several ways of drawing sample units at random.

1. Lottery Sampling- Assigning number to each member of the population usually


carries out the lottery sampling method. For example, writing down the name of each member of
the population on pieces of paper. These papers are then placed in a box or a container drum. The
box or lottery drum, the required number of sample units is picked. This method can be
illustrated though the following example: Let us say Mr. Gonzales has four complimentary
tickets to a movie. However, he has seven children and he wants to distribute the three tickets
without being accused of favoritism. So, what he does is to write the names of his children in
their presence, of course, to avoid blame later on) in pieces of paper of the same size. He then
folds the pieces of paper uniformly and places them in a container. Chronologically, the names
of his seven children are as follows:
1. Jerry

2. Jimmy

3. Jeany

3. Jojie

4. Jenilyn

5. Jerome

6. Jessica

Mr. Gonzales drops the folded pieces of paper into a container and picks from the
container the lucky names who are to receive the tickets. The first four names to be picked are
the lucky ones. If the first name to be drawn is Jerome, he gets the first ticket. If the second name
to be drawn is Jojie, she gets the second ticket. If the third name to be drawn is Jerry, he gets the
third ticket. And if the fourth name he picks is Jessica, then she gets the fourth ticket. This
simple illustration is one way of using the lottery method. Drawing prizes through the raffle
system follows the principle of random sampling.

2. Table of Random Numbers- The use of the Table of Random Numbers is another technique
of random sampling wherein the selection of each member of the population is left adequately to
chance. Every member of the population has an equal chance of being chosen.

To illustrate this technique, refer to Table 2 The use of the Table of Random Numbers
can be illustrated as follows.

a) Direct Selection Method. This method is used when there are only few samples units
to be selected. From the table, the sample units can be selected right away.

EXAMPLE 1.9

Taking the case of Mr. Gonzales, let us find out how the direct selection method differs
from the lottery method. As mentioned before, Mr. Gonzales has four complimentary tickets to a
movie and he wants to distribute them to his seven children without being accused of favoritism.
For fairness’ sake, he can use the Table of Random Numbers.

First, we have to enumerate the children and assign a number to each one of them. The
number to each one of them. The numbers serves as codes, such corresponding to one name.
These are the numbers which will be the basis for the objective choice of winners.
Numbers Name

1 Jerry

2 Jimmy

3 Jeany

4 Jojie

5 Jenilyn

6 Jerome

7 Jessica

Techno Box

Aside from the Table of Random Numbers, random numbers between 0 and 1 can also be
generated using a scientific calculator or computer.

1. On a scientific calculator, just press the following keys:

Shift . or 2nd RAN#

2. In a Microsoft Excel Worksheet, simply enter the function = RND ( ) on any blank cell.
N =123

Table 2
Table of Random Numbers

613238 946267 983341 475358


990065 028290 796267 759112
067217 252131 492824 556579
655118 613844 329285 543481
253755 019182 240271 039218
715080 381570 434045 008326
482228 615413 827812 066624
011007 596369 691115 132954
15444 26165 381865 857213
52784 86456 377015 656580
91696 37285 846438 990130
77541 69780 235305 457608
54691 72785 288861 696203
93488 49661 575566 734237
47513 49719 127639 656391
41435 11717 501117 956078
35253 54569 738965 991743
84319 51574 604060 246201
11194 61668 686436 155950
40294 83236 693078 480345

Referring to the Random Table, any of the four columns may be used. Since Mr.
Gonzales has seven children, we consider 7 as the total population (N) which is only a one-digit
number. So, we can only take the first digit of any of the columns from the table. Let us take the
numbers under the first column one at a time. What do we get?

613238 – The first digit is 6, so, Jerome gets the first ticket.
990056- Not included, since N=7
068217- Also not included.
655118- Skip. Jerome already has this number.
253755- The first digit is 2, so, Jimmy gets the second ticket.

715080- The first digit is 7, so, Jessica gets the third ticket.
48228 - The first digit is 4, so, Jojie gets the fourth and last ticket.
Thus, using the Table of Random Numbers, the lucky four children who can get the
tickets to the movie are Jerome, Jimmy, Jessica and Jojie. As illustrated, numbers are used
through the direct selection method. Similar patterns can be utilized.

b) The Remainder Method. Often, however, we cannot rely only on the direct selection
method in picking the sample units. This is due to the fact that when the number picked from the
Random Table is greater than N (for example, N=120 while the number picked from the Table is
greater than 120) the number that will have to be skipped. If we rely on the direct selection
method, there will be many number skipped. If this happen, we may have gone through the entire
problem; the remainder method is used whenever the direct selection method cannot be
applied. Whenever the direct method can be used, however, we should proceed to use it. There
are two ways of conducting the remainder method.

(1) When the number taken from the table of Random numbers is subtracted from the upper limit
within this number falls, the remainder is the sample unit.

(2) When the upper limit of the set is subtracted from the number taken from the Random table
and yields a number equal or less than N, the remainder is the sample unit.

Example 1

Using the last column of the Table of Random Numbers, pick 10 sample units from a
population of 123. Thus, N= 123 and n=10

The first step is to determine how many sets of 123 items we can construct. Since our N
is a three-digit number, we can only make use of numbers not exceeding three digits, that is, 1-
999. The following will show us how to proceed with this method,

SETS
First set (123x1 = 123); 1 – 123
Second Set (123x2 = 246); 124-246
Third Set (123x3 = 369); 247-369
Fourth Set (123x4 = 492); 370-492
Fifth Set (123x5 = 615); 493-615
Sixth Set (123x6 = 738); 616-738
Seventh Set (123x7 = 861); 739-861
Eight Set (123x8 = 984); 862-984
Ninth Set (123x9 = 1,170); 985-999- rejected, since this set is not complete, i.e., does
not contain 123 items.

We can use the first eight sets because each number in these sets in given eight equal
chances of being selected. We do not use the ninth set because it does not include the total 123
items to make a complete set.

We do it in the following way:

The Table contains numbers, each with six digits. But our N (which is 123) is only a
three-digit number. The numbers we can, therefore, use must no exceed three digits, hence from
1 to 999. We have just distributed these numbers into seven sets of 123. The ninth set was
rejected for not containing the entire 123 elements.

The first number in our Table (using the last column) is 475358. We only need to use the
first three digits of this number. Thus, we pick 475 as our number.

The second step is to look for the set, which contains the number 475.

The third step is to subtract this number from the upper limit, which contains this
number.

When we subtract this number, i.e. 492, the remainder is the number of our first sample
unit.

The 10 sample units we need are, therefore, picked in the following manner.

First sample unit

492 – upper limit of set 370-492


-475 – the number obtained from the Table (475358)
17 – the remainder which is our first sample unit

Second sample unit

861 – upper limit of set 739-861


-759 – the number obtained from the Table (759112)
102 – the remainder which is our second sample unit
Third sample unit

615 – upper limit


-556 – the number from the Table
59 – our third sample unit

Fourth sample unit

615
-543
72 – our fourth sample unit

Fifth sample unit

The number we get from the Table is 039218. Since this case the direct selection method
can be applied, we use it to pick number 039 as our fifth sample unit.

Sixth sample unit

The number registered is the table is 008326. Since the direct selection method can also
be applied in this case, we use it to pick the number 8 as our sixth sample unit.

Seventh sample unit

066624 – Thus we pick number 66 as our seventh sample unit, again through the direct
selection method.

Eighth sample unit

246
-132
114 – our eighth sample unit

Ninth sample unit

861
-857
4 – our ninth sample unit

Tenth sample unit

738
-656
82 – our tenth sample unit

Example 2

Using the last column of the Table of Random Numbers, pick 10 sample units from a
population of 123. Thus, N = 123 and n = 10.

Since our population is a three-digit number, we can only make use of the numbers from
1 to 99 because they do not exceed three digits. We are able to construct eight sets with 123
elements each:

1 – 123 first set


124 – 246 second set
247 – 369 third set
370 – 492 fourth set
493 – 615 fifth set
616 – 738 sixth set
739 – 861 seventh set
862 – 984 eighth set
985 – 999 we reject this set since it is not complete, i.e., it does not contain
123 items.

How do we then pick the 10 sample units?

The first step is to look for the first number (using the last column) from the Table of
Random Numbers, the number is 475358, but we pick out 475 because we only need the first
three digits.

The second step is to look for the upper limit of the set which when subtracted from the
number obtained from the Table, i.e., 475, would give us a number equal or less than our
population N.

The third step to subtract this upper limit from the number obtained from the Table.

We thus pick the 10 sample units we need in the following way.


First sample unit

475 – The number obtained from the Table


-369 – Upper limit of the set which when subtracted from the number obtained
from
The Table will yield a number equal or less than N.
106 – The remainder, which is our first sample

Second sample Unit


759- The number obtained from the Table
-738- Upper limit of the set which when subtracted from the Table will yield the
number equal or less than N
21 – the remainder, which is our second sample unit.

Third sample unit

556 – The number


-492 – The upper limit
64 - Our third sample unit

Fourth sample Unit


543
-492
64- Our fourth sample unit

Fifth sample unit

039218 – Number 39 is our fifth sample unit through the direct selection method.

Sixth sample unit

008326 - number 8 is our sixth sample unit through the direct selection method.

Seventh sample unit

0066624- we pick number 66 as our seventh sample unit

Eight sample unit

132 – The number from the table


-123 – The upper limit
119 - our ninth sample unit

Tenth sample unit


656 – The number from the Table
-615 – The upper limit
41 - Our tenth and last sample unit

Four- digit numbers

Using the third column of the Table of Random Numbers, pick 100 sample units from a
population of 1,150. Thus, N=1,150 and n=10.

Since N= 1,150, a four-digit number, the numbers that can be picked from the Table of
Random Numbers will be 1-9,999.

As we have stated, a fundamental rule in random sampling is to give all numbers equal
chances of being chosen. How can we make use of numbers 1-9,999 so that each number has an
equal chance of being selected? We have to construct several sets of not more than four digits
(since our population is a four-digit number) with each set containing 1,150 elements (N= 1,150.

Firs set (1,150 x 1 = 1,150 = 1- 1,150


Second Set (1,150 x 2 = 2,300 = 1,151-2,300
Third Set (1,150 x 3 = 3,450 = 2,301- 3,450
Fourth Set (1,150 x 4 = 4,600 = 3,451 – 4,600
Fifth Set (1,150 x 5 = 5,750 = 4,601 – 5,750
Sixth Set (1,150 x 6 = 6,900 = 5,571 – 6,900
Seventh Set (1,150 x 7 = 8,050 = 6,901 – 8,050
Eight Set (1,150 x 8 = 9,200 = 8,051 – 9,200
Ninth Set (1,150 x 9 = 10,350= 9,201- 9,999
- We reject this set since it does not contain 1,150 elements.

The first stew sample units under the third column of the Table of Random Numbers are
the following:

983341 - falls in the rejection region

796267

492824

329285 The first four digit numbers are the numbers that can be used.

240271

434045

827812
The following procedures will illustrate how we can use these numbers in picking our
sample units.

FIRST METHOD

When the number taken from the Random Table is subtracted from the upper limit within
which this number falls, the remainder is the sample unit.

First sample Unit

8050 - upper limit within which 7962 falls


-7692 - number taken from Random Table
88 - our first sample unit

Second sample unit

5750 - upper limit


-4928 - number taken from random table
822 - our second sample unit

Third sample Unit

3450
-3292
158 - our third sample unit

Fourth sample unit

3450
-2402
1048 - our fourth sample unit
The process goes on until we pick 100 sample units.

SECOND METHOD

When the number taken from the Random Table is subtracted from the upper limit within
this number falls, the remainder is the sample unit.

When the upper limit of the set is subtracted from the number taken from the Roman
Table and yields a number equal or less than N, the remainder is the sample unit.

First sample unit

7962 - number taken from Random Table


-6900 -Upper limit which subtracted from the number taken from Random table
less than N
1062 - Our first sample unit.

Second sample unit

4928 - number taken from Random table


-4600- upper limit
328- our second sample unit

Third sample unit

3292- number taken from Random table


-2300- upper limit
992 - our third sample unit

The process goes until the desired sample units are picked out.

[Link] [Link] addition to the method of random sampling, a number of


methods have been developed which may be called systematic methods. These methods use prior
knowledge of the individuals comprising a universe with the end view to increasing precision
and representation of samples. When sample units are obtained by drawing every, say, 4th or 7th
or 10th item on a list, the process of selecting the samples is called systematic sampling.

This method involves selecting every nth element of a series representing the population.
A complete listing is required in this method. Under this system, the sample units may be picked
in the following manner, where, for example,

N=100
n=10

The value of n may be obtained by dividing the total number of elements in the
population by the desired sample size. Thus,

N=100
n 10= 10th

every nth element

The sample units would, therefore, be the persons holding the following numbers; 10, 20,
30, 40, 50, 60, 70, 80, 90 and 100 at 10 sample units. Variation may be added by choosing a
random start. Let us take 10 pieces of paper and number them 1-10. We put these pieces of paper
in a box or container and shake them thoroughly. If the number 7 is picked as the random start,
the 10 sample units should be: 7, 17, 27, 37, 47, 57, 67, 77, 87, 97, = 10 sample units.

4. Stratified Sampling - This is a random sampling technique in which the population is divided
into non-overlapping subpopulations called [Link] this method, the population is first divided
into group-based on homogeneity -in order to avoid to possibility of drawing samples whose
members come only from one stratum. In stratified sampling , the distribution of sampling units
is proportionate to the total number of units of each stratum. The bigger the population , the more
sample units are drawn, the less population , the less sample units. That is why this method is
often called stratified proportional sampling. In contrast , simple random takes the proportion to
chance

The formula below is used to determine the sample sizes for proportional allocation.

ni ⟦ NN ⟧n for i= 1, 2, 3….
n = is the total size of the stratified random sample
N = total population
N1 = number of 1st stratum elements
N2 = number of 2nd stratum elements
N3 = number of 3rd stratum elements

Example 1

At a BSHRM in ISU Cauayan , the students may be classified according to the following
scheme:

Table 1.3
Proportional Sample Size Allocation

Classifications Population Sample


N n
4th 119 34

3rd 210 60
2nd 325 93

1st 346 99
Total 1,000 286

1st year

r = ns
N

r = 286(346)
1000

r= 98.956
r= 99

2nd year (325)

r = ns
N

r = 286(325)
1000

r= 92.95
r= 93

3rd year (210)

r = ns
N

r = 286(210)
1000

r= 60.06
r= 60

Note: To determine the appropriate sample size without resorting to your subjective decision,
you may use the Slovin’s formula.

n= N
1+ N e²
Where:
n= sample size
N= population size
e= 0.05 (the sampling error)

n= N
1+ N e²

n= 1000
1 + 1,000(0.05)2

n = 1,000
1 + 1,000(0.0025)

n = 1,000
1 + 2.5

n = 1,000
3.5

n = 285.71 0r 286

r = ns
N

r=respondents in each strata/year level


n=sample
s=total population in each strata/year level
N=population

4th Year

r = ns
N

r = 286(119)
1000

r= 34.034 or 34

Using the proportional allocation to select an appropriate sample of size n=286, how
large a sample must be taken for each stratum.
Solutions: n= 286, N1= 119, N2=210, N3= 325, N4=346 and n=100
Computations:

N1= 119 x 286 = 34


1000
N2= 210 x 286 = 60
1000
N3= 325 x 286 = 93
1000
N4= 346 x 286 = 99
1000

Example 2
A certain corporation has always been plagued with the problem of fast worker turnover.
Every year, a sizeable number of its workers leave the corporation to look for work elsewhere.
Since it gives adequate salaries to its employees compare to the wages of other corporations of
the same class and size, executives of the corporation were perplexed by the situation and wanted
to find out the reason for the high rate of worker turnover by doing a survey.
The corporation is divided into different departments : Administrative, manufacturing,
finance, warehousing, and research and development. The total number of workers is 490. How
will we pick the sample units if the corporation wants 20% of the employees to be surveyed?

STEP 1. Identify the population and its different strata.

Table 3. Distribution of Population


Departments Employees
Administrative 250
Manufacturing 120
Finance 85
Warehousing 20
Research and Development 15
Total 490

STEP 2. Determine the percent share of each stratum with respect to the total population, as
shown in table 3.

Table 3. Distribution of Population


Departments % Share
Administrative 250/490 = 51.02
Manufacturing 120/490= 24.49
Finance 85/490 = 17.35
Warehousing 20/490 = 4.08
Research and Development 15/490= 3.06
Total 100.00

STEP 3. Management wants 20% of the population, which is 98% of the sample units, to be
surveyed. The third step will be to multiply the percent share of each department in elation to
the total population by 98, which is the number of sample units to be surveyed. To get the actual
number of sample units for each department . This is shown in Table 4

Table 4. Distribution of the Sample Units


Departments Number of Sample Units
Administrative 51.02 % of 98 = 50.00
Manufacturing 24.49 % of 98 = 24.00
Finance 17.35 % of 98 = 17.00
Warehousing 4.08 % of 98 = 4.00
Research and Development 3.06% of 98 = 3.00
Total 100.00

[Link] SAMPLING. This sample techniques is sometimes referred to as an area sampling


because it is frequently applied on a geographical . Selects a sample containing either all, or a
random selection, of the element for clusters that have been selected randomly from the
[Link] sampling for the advantage of being more cost efficient when the population
is wildly [Link] is useful in selecting the sample when blocks in a community or city are
occupied by heterogeneous groups. Example , if a community in Metro Manila has lower,
middle, and upper income residents living side by side, we may use this community as source of
a sample to study the different socio-economic group s in Manila. By concentrating on this
particular area, we can save more time, effort, and money that if we covered different
communities throughout Manila.

Another example, in studying the investment habits of working parents in a given region,
it is much cheaper to interview and collect data from individuals living close together in
randomly selected clusters for provinces or cities than to select a sample random sample for the
entire region. Considering geographic areas as clusters, this kind of sampling is also called Area
Sampling.

6. SYSTEMATIC SAMPLING. In this technique, members of a sample are chosen at regular


intervals of a population. It requires selection of a starting point for the sample and sample size
that can be repeated at regular intervals. This type of sampling method has a predefined interval
and hence this sampling technique is the least time-consuming.

Example 1.

A researcher intends to collect a systematic sample of 500 people in a population of 5000.


Each element of the population will be numbered from 1-5000 and every 10th individual will be
chosen to be a part of the sample (Total population/ Sample Size = 5000/500 = 10).
Example 2

Suppose we select n= 12 students from a population of size N=50.

To employ systematic sampling, divide N by n to get k, that is

K=N = 50
N = 12
K= 4.667 =5

From an ordered population

1 2 3 4 5 6 7 8 9 10
11 12 13 14 15 16 17 18 19 20
21 22 23 24 25 26 27 28 29 30
31 32 33 34 35 36 37 38 39 40
41 42 43 44 45 46 47 48 49 50

Choose every 5th unit. Thus, if the random start r= 9th unit, then the sample comprises of
students numbers.

9, 14, 19, 24, 29, 34, 39, 44, 49 and 54

[Link]-STAGE SAMPLING. A sampling procedure wherein the population is divided into


a sequence of sampling units corresponding to the different sampling stages. However, selection
of the sample is still done at random. This is applicable and useful in conducting nationwide
surveys or any-survey involving a large universe.
As an example , let us see how this technique can be used by picking a sample from the
regional division of the Philippines.

Stage 1. Enumerate all 16 regions of the Philippines including their respective municipalities.

Stage 2. From the 16 regions, select three at random. This can be done through Table of Randon
Numbers or by lottery.

Stage 3. We have now three regions out of 16 we had before. From these three regions, we select
two provinces from each region. The process of selection should also be done at random.

Stage 4. With two provinces from the three regions, we have our list right now six provinces. We
enumerate all the municipalities and cities in all of these provinces. The process of selection
should be done at random.
Thus, we have six provinces from where we can pick our sample. From each of these
provinces we select three municipalities or cities. In the final analysis, we will only survey 18
municipalities/cities ( 3 x 6 =18) in our study.
Non- Random Sampling does not involve random selection of sample elements. Some
elements of the population do not have a chance to be included in the sample.

1. Voluntary Sampling. It is usually done by television or radio programs asking


people to participate in an on-line poll or bay answering a question regarding a
national issue through [Link] is mad up of people who self-select into the survey.
Often , these folks have a strong interest in the main topic of the survey. Therefore,
the sample is chosen by the viewers , not by the survey administrator.
2. Convenience Sampling. It is done by considering these people around him as the
qualified samples. It is made up of people who are easy o reach. Consider the
following example. A pollster interviews shoppers at a local mall. If the mall was
chosen because it was a convenient sire from which to solicit survey participants
and /or because it was close to the pollster’s home business, this would be
convenience sampling.
3. Quota Sampling. This is a relatively quick and inexpensive method to operate. Each
interviewer is given definite instructions about the section of the public he is to
question, but the final choice of the actual persons is left to his own convenience or
preference, and is not predetermined by some carefully operated randomizing plan.
Each interviewer then proceeds to fill the prescribed quota. Suppose there is a survey
to estimate what perecent of the population of Cauayan City consider basketball as
one of their favorite sports. One of the Interviewer might report that 100% of his
quota of 70 people are basketball fans. However, it may later be found out that this
interviewer reached his quota by going to the Araneta Coliseum to enjoy watching his
favorite team and at the same time to interview some thrilled viewers during the
game.
4. Purposive Sampling. This is based on a certain criteria laid down by the researcher.
People who satisfy the criteria are interviewed. A researcher might want to find out ,
for example, the reaction of the bankers association regarding a particular Central
Bank circular. Instead of interviewing the executives of all banks, he purposely can
chose to interview the key executives of only the five biggest banks in the country if
he believes that it is the reaction of these big ones that counts anyway. Of course, the
answers obtained though this procedure are not representative of the entire banking
system.

Try this

1. Why are sampling techniques necessary?

2. When do you use each sampling technique?

[Link] of the following involves a population? A sample?

a) A new feed for piglets that yields a better market


b) The number of people contaminated with SARS

c) The most productive venture in business.

4. In a town of 25,000 people, it is but practical to choose a representative group for certain
research project. How do you get the random sample?

6. An educator wants to find out information about the families represented in a school system
and decides to pick a random sample of children from those registered in the schools for survey.
He will ask such questions as: How large is your family? How much education do your parents
have? And so on. Is there anything wrong with this sampling plan.
After the data have been presented in tabular or graphical form, the researcher must be able to
describe them in terms of a single number. This number which gives a summary or the
characteristics of a given set of data is called a measure of central tendency or measure of central
location.

You often read the word “average” in books newspapers, magazines, and journals. You hear this
word on radio and television. For example, the average salary of employees of the different
establishments of Isabela is P15,000. The batting average of a baseball player is 375. The
average IQ of students in a particular school is [Link] this module, you will learn that the term
average could imply the mean, median, or the mode, which are referred to as the measures of
central tendency.

The most commonly used measures of central tendency are the mean or arithmetic average, the
median and the mode. Such measures of central tendency can be computed in two data which are
not yet organized or arranged in frequency distribution otherwise such is called group data.

Objectives

1. Differentiate one measure of central tendency from the others;


2. Compute for the measure of central tendency for ungrouped and grouped data; and
3. Describe data using measures of central tendency

Measures of Central Tendency of Ungrouped Data

Ungrouped data or Raw Data. This refer to data which are not yet organized or arranged into
frequency distribution.

[Link]

Characteristics of Mean

1. An interval statistics
2. A calculated average
3. Value is determined every case in the distribution
4. Affected by extreme values
5. Can be subjected to numerous mathematical computations.
6. Most widely use
7. Represents average quantity

A. Simple Arithmetic [Link] arithmetic mean or simply mean ( popularly called the
average) is the sum of the separate scores or measures divided by the number of the
scores. T
Arithmetic [Link] arithmetic mean or simply mean ( popularly called the average) is the sum
of the separate scores or measures divided by the number of the scores.
Where, ∑ (the uppercase Greek letter sigma), X refers to summation, refers to the individual
value and n is the number of observations in the sample (sample size)

Example 1.

The ages in years of 10 sales clerk of a mall are:32, 41, 28, 54, 35, 26, 23, 33, 38, [Link] is the
mean age of these employees?

Solution :

Mean age of the sales clerk

= Sum of age of sales clerk / Number of sales clerk

= (23 + 26 + 28 + 32 + 33 + 35 + 38 + 40 + 41 + 54) /10

= 350 / 10

= 35 years

Example 2

Find the mean of 2, 4, 6, 8, 10 , 12, 14, 16.

Solution :

Mean = Sum of given numbers / 8

= (2 + 4 + 6 + 8 + 10 + 12 + 14 + 16) / 8

= 72 / 8
=9

Example 3.

John worked in a food chain as part time for 4 hours, 5 hours and 3 hours respectively on three
consecutive days. How many hours does he work daily on an average?

Solution :

The average work time of John

= Total number of work hours / Number of days for which he work

= (4 + 5 + 3) / 3

= 12 / 3

= 4 hours

Thus, we can say that John work for 4 hours daily on an average.

[Link] Mean

Weighted mean is calculated when certain values in a data set are more important than
the others. A weight wi is attached to each of the values xi to reflect this importance.
Consider the proper weights assigned to the observed values according to their relative
importance.

In formula,
n

Xw =
∑ WiXi
i=1
N
Where
W¡ = weight of each item
X¡ = value of each item
X = mean
∑ = means of the sum of
Example 1.
Suppose that a marketing firm conducts a survey of 1,000 households to determine the average
number of TVs each household owns. The data show a large number of households with two or
three TVs and a smaller number with one or four. Every household in the sample has at least one
TV and no household has more than four.
Solution:
Here’s the sample data for the survey:

Number of TVs per Household Number of Households

1 73

2 378

3 459

4 90
As many of the values in this data set are repeated multiple times, you can easily compute the
sample mean as a weighted mean. Follow these steps to calculate the weighted arithmetic mean:
Step 1: Assign a weight to each value in the dataset:
x1=1,w1=73
x2=2,w2=378
x3=3,w3=459
x4=4,w4=90
Step 2: Compute the numerator of the weighted mean formula.
Multiply each sample by its weight and then add the products together:
∑4i=1wixi=w1x1+w2x2+w3x3+w4x4
= (1)(73)+(2)(378)+(3)(459)+(4)(90)
=2566
Step 3: Now, compute the denominator of the weighted mean formula by adding the weights
together.
∑4i=1wi=w1+w2+w3+w4
= 73 + 378 + 459 + 90
=1000
Step 4: Divide the numerator by the denominator
∑4i=1wixi∑4i=1wi
=25661000
=2.566
The mean number of TVs per household in this sample is 2.566.

Example 2

A man bought 10 liters of premium gasoline at 11.50 per liter, 12 at P12.01 per liter and at
P11.78 per liter from three different gasoline stations. Find the average price per liter.

Solution
WiXi+W 2 X 2+W 3 X 3
Xw=
W 1+ W 2+ W 3

10 (11.50 ) +12 ( 12.01 )=18(11.78)


=
40

Mean for grouped data

To compute for the arithmetic mean of grouped data, we need to determine the midpoint
of each class interval. The mean assumed that the class mark of each is the average value of all
items falling in that class.

These are two methods of computing for the mean:

1. Long Method:

∑ fiXi
X= i=1
n
Where;

X = sample mean
X¡= the class midpoint or class mark
f¡= the corresponding frequencies
n = total number of items.

2. Short Method: (the coded formula)

X=
( )
∑ fidi
i=1
n
i
Where:

X = sample mean
XA= assumed mean (usually the value of the highest midpoint)
Di = deviation of the values from the assumed mean,

Di = ( Xi−i Xa)
f¡ = the corresponding frequency
I = class interval or class size
N= number of items

Example 1
The following are the distribution of length of service in years of 50 employees of
United laboratories Inc.

Length of Service Number of Employees


(Class Limits);CL (Frequency);f

1-5 5
6-10 7
11-15 12
16-20 13
21-25 6
26-30 4
31-35 3

Determine the mean using long method and short method.

SOLUTION:

Length Of Number of Class fXi d fd


Service Employees xi midpoint
CL F

1-5 5 3 15 -3 -15
6-10 7 8 56 -2 -14
11-15 12 13 156 -1 -12
16-20 13 18 234 0 0
21-25 6 23 138 1 6
26-30 4 28 112 2 8
31-35 3 33 99 3 9

n= 50 ∑fx=810 ∑fd= -18

Long Method:
n

x =∑
f ¡ x¡
i=1
n
810
x=
50
= 16.2

Short Method

x = XA + ( ∑
f ¡d ¡
i=1 )i
n

−18
= 18+ ( )5
50

= 18- 1.8

X= 16.2

The formula for a sample mean for grouped data is

where x is the midpoint of the interval, f is the frequency for the interval, fx is the product of the
midpoint times the frequency, and n is the number of values.

For example, if 8 is the midpoint of a class interval and there are ten measurements in the
interval, fx = 10(8) = 80, the sum of the ten measurements in the interval.
Σ fx denotes the sum of all the products in all class intervals. Dividing that sum by the number of
measurements yields the sample mean for grouped data.

For example, consider the information shown in Table 3.

Substituting into the formula:

Therefore, the average price of items sold was about $15.19. The value may not be the exact
mean for the data, because the actual values are not always known for grouped data.

MEDIAN
The median is the value of the middle item after arranging the data in an ascending or
descending order

Characteristics of Median
1. An ordinal statistic
2. A rank or position average
3. Value is determined by scores near the middle of the distribution.
4. Not affected by extreme values
5. Can be subjected to only a few mathematical computations.
6. Less widely used than mean
7. Represents typical score.

Median for Ungrouped Data


The median of ungrouped data arranged in array (increasing or decreasing order of
magnitude) is the middle value when the number of items is odd or the arithmetic average of the
two middle values when the number of items in the distribution is even. The median is usually
denoted by Mdn.

Examples

1. Compute for the median from the following set of scores; 6, 4, 5, 3 and 2.

Solution. Arrange the set of items and compute for the median.
2, 3, 4, 5, 6,
The median is 4, which is the middle item.

2. Find the median of the following set of item;


8, 12, 5, 6, 13 and 15
Solution- Arrange the data and find the median

5, 6, 8, 12, 13, 15

8+12 20
Mdn = = = 10
2 2

Median for Grouped Data

The median of a grouped data could be determined by the following formula:

n
−cfB
Mdn = Lme + ( 2 )i
Fme

Where;
Mdn = median
Lme = lower class boundary of median class
N = total number of observations
fMe = frequency of the median class
cfB = cumulative frequency preceding the median class
I = class size or size of the class interval

The median class is the limit which the

nth
value
2

Example

Find the median of the frequency distribution of length of service in years of 50


employees of United Laboratories Inc. (from example1.)

Solution:

Length of Service Number of Employees CB <cf


CL F
1-5 5 0.5 – 5.5 5
6-10 7 5.5 – 10.5 12
11-15 12 10.5 – 15.5 24
16-20 13 15.5 – 20.5 37
21 – 25 6 20.5 – 25.5 43
26-30 4 25.5 – 30.5 47
31-35 3 30.5 – 35.5 50
N=50

To obtain the median class:

n 50
First solve for = = 25th
2 2

Then, locate where the 25th item is equal or nearest but not greater than the value in the
less than cumulative frequency (<cf) distribution.

The median class is the 16-20 class intervals.

n
−cfB
Me = Lme + ( 2 )i
Fme
50
−24
= 15.5 + ( 2 )5
13
25−24
= 15.5 + 5
13
= 15.5 + 1.667
Me = 17.167 / 17.2

MODE
The mode is the simplest measure of central tendency. It is easily identified merely
looking at an ungrouped set of scores or data and locating the score or data which occurs most
frequently.

Characteristics of Mode
1. A nominal statistics
2. An inspection average
3. The most frequent occurring value
4. Usually occurs near the center of distribution.
5. Cannot be manipulated mathematically.
6. Some distribution have more than one mode.
7. Rarely used
8. Most “ popular” score

Mode for Ungrouped Data


The mode for ungrouped data is defined as the value that appears with the highest
frequency. That is, the item that appears most often, usually denoted by Mo.

When all values appears with the same frequency, the mode does not exist. However, for
some sets of data there may be several values appearing with the greatest frequency in which
case we have more than one mode. When two modes are present in a distribution, such is called a
bimodal, trimodal for three modes and so forth.

Example 1
Find the mode of the following set of items.
4, 7, 11, 6, 4, 8, 3, 5, 2, 9

Solution: The mode is 4. It is unimodal.

Example 2
Determine the mode of the following distribution
12, 15, 21, 9, 6, 15, 11, 8, 9, 5

Solution:The modes are 15 and 9, bimodal .

Example 3
Find the mode of the following sets of data.
12,14,17,18,19,20,26, 26,19,20, 29

Solution
The modes are 19, 20, 26, trimodal.

Example 4
Find the mode of the following data.
6, 5, 11, 21, 8, 16, 19, 20, 7, 9

Solution: The mode does not exist, because all the time are nor repeated.

A distribution with only one mode is said to be UNIMODAL. Distribution with two
modes is BIMODAL, with three modes is TRIMODAL while a distribution with two or more
modes is described as MULTIMODAL.

Mode for Grouped Data

The mode of a grouped data is defined as the midpoint of the class interval with the
highest frequency (modal class). The mode obtained in this manner is called a crude mode,
because it is just a rough approximation of the actual mode. So, to determine the true mode, we
use the formula
d1
Mo = Lmo + ( )i
d 1+ d 2

Where:
Mo = mode
Lmo = lower class boundary of modal class
D1 = difference between the frequency of the modal class and the frequency
of the class next lower in value.
D2 = difference between frequency of the modal class
i = class interval or class with

The modal class is the class interval with the highest frequency.

Example

Find the mode of the frequency distribution of length of service in years 50 employees of United
Laboratories Inc. (from example 1)

Solution
Length of Service Number of Employees CB
CL f
1-5 5 0.5 – 5.5
6-10 7 5.5 – 10.5
11-15 12 10.5 – 15.5
Mo Class» 16-20 13 15.5 – 20.5
21-25 6 20.5 – 25
26- 30 4 25.5 – 30.5
31 – 35 3 30.5 – 35.5
n=50

The modal class is the 16-20, since it is the class interval with highest frequency.

d1
Mo = Lmo + ( )i
d 1+ d 2

1
= 15.5 + ( )5
1+ 7
5
= 15.5 +
8
Mo = 16.125

Example

The following is a frequency distribution of an entrance examination. Compute the mean,


median and mode.

Solution

CL f x fx
21-30 8 25.5 204
31-40 11 35.5 390.5
41-50 15 45.5 682.5
51-60 18 55.5 999
61-70 20 65.5 1310
71-80 12 75.5 906
81-90 9 85.5 769.5
91-100 7 95.5 668.5
N=100 ∑fx=5,930

Mean x

Long Method
n

x =∑
fixi
i=1
n
5930
=
100
x=59.3

Short Method

CL f X CB <cf D fd
21-30 8 25.5 20.5-30.5 8 -4 -32
31-40 11 35.5 30.5-40.5 19 -3 -33
41-50 15 45.5 40.5-50.5 34 -2 -30
51-60 18 55.5 50.5-60.5 52 -1 -18
61-70 20 65.5 60.5-70.5 72 0 0
71-80 12 75.5 70.5-80.5 84 1 12
81-90 9 85.5 80.5-90.5 93 2 18
91-100 7 95.5 90.5-100.5 100 3
N=100 ∑fd=-62

x = XA + ( ∑
f ¡d ¡
i=1 )i
n
= 65.5 + (-6.2)
= 65.5 – 6.2
= 59.3

MEDIAN (Mdn)

n
−cfB
Mdn= Lme+( 2 )i
Fme

n 100
Median class = = = 50th
2 2

50−34
= 50.5 + ( )10
18
= 50.5 +8.89
Mdn = 59.39 59.4

MODE (Mo)

Modal Class is the 61-70 ( with highest frequency)


d1
Mo= LMo + ( )i
d 1+ d 2
2
= 60.5 + ( ) 10
2+ 8
= 60.5 + 2
1
= 60.5 + ( ) 10
5
Mo = 62.5

ADVANTAGES
The mean uses every value in the data and hence is a good representative of the data. The irony
in this is that most of the times this value never appears in the raw data.
Repeated samples drawn from the same population tend to have similar means. The mean is
therefore the measure of central tendency that best resists the fluctuation between different
samples.[6]
It is closely related to standard deviation, the most common measure of dispersion.
Go to:

DISADVANTAGES
The important disadvantage of mean is that it is sensitive to extreme values/outliers, especially
when the sample size is small.[7] Therefore, it is not an appropriate measure of central tendency
for skewed distribution.[8]
Mean cannot be calculated for nominal or nonnominal ordinal data. Even though mean can be
calculated for numerical ordinal data, many times it does not give a meaningful value, e.g. stage
of cancer.
References
Petrie A, Sabin C. Medical statistics at a glance. 3rd ed. Oxford: Wiley-Blackwell;
2009. [Google Scholar]. Retrived from
[Link]

[Link]
tendency#:~:text=Measures%20of%20central%20tendency%20are,values%20shown%20in
%20Table%201.

+
Characteristics of Mean

8. An interval statistics
9. A calculated average
10. Value is determined every case in the distribution
11. Affected by extreme values
12. Can be subjected to numerous mathematical computations.
13. Most widely use
14. Represents average quantity

15. The mode is a nominal statistic which means that is used for nominal data. Its
computation does not depend on the values of the variable nor on their order, but merely
on their frequency of occurrence. It is rarely used with interval , ratio, and ordinal
variables, where means and medians can be calculated.
16. It is usually employed as a simple, inspectional measure which indicates roughly the
center of concentration of a distribution. As such, there is no need to calculate it as
exactly as the median or the mean.
Consideration in the Use of a Measure of Central Tendency

1. We use mean when


a. The greatest reliability is desired; that is, of a measure of central location that varies
less sample to sample drawn from the same population is desired.
b. There is a need for further statistical computations, as finding the measure of standard
deviations, correlation ( which will be discussed later); etc.
c. Data are ratio or interval measurement.

2. We use the median when


a. The distribution is badly skewed. This includes the case when there are extremely
high or extremely low score in a given distribution.
b. An open -end distribution is given.
c. We are interested in determining whether cases fall within the upper or lower laves of
the distribution and not on how far the case are from the central point or midpoint of
the distribution.
d. Data are ordinal or ranked.
3. We use mode when
a. The quickest estimate of central value is wanted.
b. A rough estimate of the central value is sufficient.
c. We wish to know the most typical case.
d. The data are in nominal scale or categorical in nature.

Try this

1. Find the mean , median and mode of the following sets of numbers.
a. 4, 6, 7, 8,8,9,3, 2 , 4, 5,9
b. 89, 52, 47, 30, 37, 79, 89, 90 , 83, 56, 62, 73,
c. 24, 29, 40, 28, 32, 33, 22, 27, 32,33
2. A company pays its 45 employees an average a mean monthly salary of P 12,500; a
second company pays it 52 employees of P 13, 750 ; and the third company pay its 100
employees an average monthly salary of P 11, 950. What is the average salary per
employee in the tree companies.
3. The salaries of the different managers of food chain in Isabela are shown below:
Monthly Salary Number of Manager
50,000- 59,999 3
40,000-49,999 15
30,000-39,999 27
20,000-29,999 12
10,000-19,999 11

Compute the mean, median and mode of their monthly salary

[Link] following is the distribution of length of service in years of 50 employees of United


Laboratories Inc.. Compute the mean using the long and short method, the median and the
mode.

Length of Service Number of Employees


1-5 5
6-10 15
11-15 17
16-20 12
21-25 11
26-30 9
31-35 10

Assessment Task

1. An owner of a certain company obtained 5 kinds of loan from 5 commercial banks at


different rate, P 1,000,000 at 12% , 2,500,000 at 11.5%, 3,700,00 at 11%, 4,200,000 at
10% and 5,000,000 at 9% per annum. What is the mean interest rate?
2. The number of customers per day visiting in a restaurant was recorded in the security’s
logbook over a month period

Day No. of Day No. of Day [Link]


Customers Customers Customers
65 11 33 21 42

1
2 45 12 44 22 45
3 43 13 76 23 55
4 36 14 47 24 32
5 38 15 64 25 30
6 42 16 62 26 28
7 44 17 35 27 75
8 56 18 29 28 74
9 55 19 36 29 32
10 29 20 40 30 30

2.1. Determine the mean, median, and mode.


2.2. What is the least number of customers visiting in a month period?
2.3. What is the greatest number of customers visiting in a month period?
2.4. Based on the given data, what conclusion can you draw about the number
of customers visiting in a day?
2.5. Consider that there is only an average of 8 business operating hours for the
restaurant, how many customers do you expect to see in your random visit?
2.6. Assume that the seating capacity of the restaurant is 30 persons, Do you
still have a recommend for additional chairs and tables? Why?
3. Consider the data set { 8,5,1,10,3,6,4,2,6,5}
3.1. Find the mean, media and the mode of the data set.
3.2. If each data point is increased by 3, what will be the mean, median and the
mode?
3.3. If each data point is multiplied by 3, what will be the mean, median and
the mode?
3.4. If the data 10 is removed from the data set, what will be the mean, median
and mode?

4. When do you use the mean, median and mode . Explain each by using an example.
ASSESSMENT TASK
1. A sample of chicken meat from 7 supermarkets produced the following data on their
process.
P 170.00 P153.00 P 140.00 P180.00 P185.00 P164.00
P185.00
Calculate the range, variance, standard deviation and coefficient variation for these data.

2. The following table lists the prices of 5 different brand of oven toaster. Calculate the
range, variance, standard deviation and coefficient variation .

Brand Price
Sharp P1,000.00
National P1,150.00
Samsung P920.00
Philips P 850.00
G.E. P1,250.00

3. The following table gives the 2018 gross sales ( rounded to million of pesos for a sample
of food companies. Find the range, variance, standard deviation,and coefficient variation.

Company 2018 Gross Sales (Millions of Pesos)


DONGKYS 70
AQUINOS Chiken Inasal 65
ISIDRO’s Milktea 15
ENTENG Kabisote’s Bakery 15
HAROLDs Burger 20
ALFREDs Food chain 32

ASSESSMENT TASK

Measures of Variability

1. Compute the range, mean absolute deviation, variance and standard deviation of the daily
wages of the 9 employees : P 450, P350, P 425, P360,P370, P580, P320, P 410, and
P650.
2. The following are wages per day of workers in a certain [Link] the range,
mean absolute deviation, variance and standard deviation.
Wages Number
P410-414 22
415-419 19
420-424 10
425-429 15
430-434 20
435-439 11
440-444 14
445-449 12
450-454 15

Introduction to Hypothesis Testing

Key Terms

Hypothesis
Null Hypothesis
Alternative Hypothesis
Directional Alternative Hypothesis
Non-directional Alternative Hypothesis
Learning Objective

At the end of the chapter, the student are expected to:


1. Define the different concepts related to hypothesis testing correctly
2. Explain the difference between the Type I error and Type II error
3. Differentiate directional and non-directional alternative hypothesis
4. Determine the critical value of test statistics at a given alpha level
5. Apply the different steps in testing hypothesis
6. Compute the z-test and t-test values correctly
INTRODUCTION

In the chapter, the concept of hypothesis testing shall be discussed and the methods used
in testing hypothesis about the population mean when the population standard deviation is given
and other conditions about hypothesis testing will be developed. In this test, the critical value
method is testing hypothesis will be discussed. Other topics related to hypothesis testing shall
also be discussed.
DEFINITION OF HYPOTHESES

Hypothesis is an assertion or conjecture about a population parameter or parameters. It


also called statistical hypothesis. An example of parameter is the population mean or population
standard deviation. Examples of hypothesis are;

1. The proportion of the consumers who purchased Brand X of female facial wash
exceeds 60%.
2. The main daily allowance of high school students in rural areas is at most ₱150.00.
3. The average lifetime of light bulb manufactured by RL Company will last at most
5,500 hours.
4. There is no significant difference between the monthly expenditures of families in the
rural and urban areas.
5. The average of nicotine content of Y cigarette does not exceed 3.25 milligrams.

The statements above are subject to statistical testing in order to determine the
truthfulness of them. When it failed to reject the hypothesis, then the hypothesis should be
accepted.

Types of Statistical Hypothesis

1. Null Hypothesis

A null hypothesis is a hypothesis of no difference and it is denoted by Ho. it is usually


designated by a ‘not’ or ‘no’ term in the null hypothesis, meaning that the variables
compared are the same or have no difference. The null hypothesis can be expressed in
statement from stating no significant difference between the variables compared or using
mathematical statements such as: Ho : µ = µo’ Ho : µ1 = µ2’ Ho : ẋ1 = : ẋ2. Consider this statement,
the average monthly salary of college professors in public and private schools is the same.
The Ho is, there is no significant difference between the average monthly salary of college
professors in the public and private schools. (Ho : µ1 = µ2).

2. Alternative Hypothesis

Alternative hypothesis is a hypothesis to be considered as an alternate to the null


hypothesis. It is usually denoted by H1 and read as “H sub 1.” The alternative hypothesis
is that there is a significant difference between the average monthly salary of the college
professors in public and private schools. (H1 : µ1 ≠ µ2 when H1 is non-directional; H1 : µ1
> µ2; H1 : µ1 < µ2 when H1 is directional).
Accept the alternative hypothesis if the sample data provides enough evidence
that the null hypothesis it is implied that the alternative hypothesis will be accepted.

Types of Decision Errors

When dealing with hypothesis tests, there are four possible outcomes; the two outcomes
lead to incorrect decision and the other two lead to correct decision. The outcomes are described
in the given table.

Table 6.1
Possible Outcomes for a Hypothesis Test

Fact Ho is true Ho is false


Decision

Failed to reject Ho Correct decision Type II error


Reject Ho Type I error Correct decision

Based from the given table, a researcher commits an error if a true H0. When a researcher
rejected a true H0, he commits a Type I error or alpha error (ɑ). When a researcher accepted a
false H0, he commits a Type II error or beta error (β).

1. Type I error or alpha error (ɑ). A Type I error is committed when the researcher rejects a
null hypothesis when in fact it is true.
2. Type II error or beta error (β). A Type II error is committed when the researcher accepts a
null hypothesis when in fact it is false.

Level of Significance

When a researcher tests the hypotheses, he is not certain that the decisions 100% correct.
However, he is confident at a certain level that the decision is correct, say 99% of the decision he
made is a correct one. The confidence level is 99% or the level of significance is 1%. When the
confidence level is 95%, the level of significance is 5%. On the other hand, when the confidence
level is 90%, the level of significance is 10%. In this case, the higher the confidence level, the
more certain that the decision of rejecting the null hypothesis is correct.

Level if significance is the probability of committing a Type I error or alpha (oc) or the
probability of rejecting the correct null hypothesis.
Power of A Test
Power of a test is the probability of not committing a Type II error or beta (β).

Tests Statistic
The tests statistic is used as a basis for deciding whether to reject or accept the null
hypothesis. The rejection region lies at either left or right tail of the normal curve if one-tailed
test is being used. On the other hand, the rejection lies at both end tails of the normal curve if
two-tailed test will be utilized.

Non-rejection Rejection region


region


Figure 6.1

Graphical Representation of Rejection and Non-Rejection region

Rejection Region

When the test statistic lies on the rejection region, then the null hypothesis will be
rejected.

Non-Rejection Region
The non-rejection region is the probability of making a Type I error equals to the level of
significance. Non-rejection region is also known as the acceptance region. When the tests
statistic lies within the non-rejection region, the null hypothesis will be accepted or the critical
value is greater than the computed value of the test statistic. The null hypothesis will be rejected
otherwise it will be accepted.

Critical value

The critical value is a value that separates the non-rejection region and the rejection
region.

Rejection region
Non-rejection
region

ẋ Z= 1.645
Critical Value
Figure 6.2

Graphical Representation of the Critical Value

One-Tailed Test and Two-Tailed Test

The use of the one-tailed test or two-tailed test will depend on how the alternative is
formulated. If the alternative hypothesis is expressed in non-directional, utilize the two-tailed
test. However, use the one-tailed test if the alternative hypothesis is directional. In two-tailed
test, the two rejection regions lie at both end tails of the normal curve each part will be half of
the alpha value. If ɑ= 0.05, the area in each end tail is ɑ= 0.025. In one-tailed test, the rejection
region lies either left end tail of the normal curve or right end tail of the normal curve.

Non-rejection Rejection region


region


Figure 6.3

One-Tailed Test

Rejection region
Figure 6.4

Two-Tailed Test

Rejection region Non-rejection


region

Z = -1.645 ẋ
Figure 6.5

Graphical Representation of Left-Tailed Test at ɑ = 0.05

Non-rejection Rejection region


region

ẋ Z = -1.645
Figure 6.6

Graphical Representation of Right-Tailed Test at ɑ = 0.05

Non-rejection
region

Rejection region

Figureẋ 6.7
Z = 2.33

Graphical Representation of Left-Tailed Test at ɑ = 0.01

Non-rejection
region
Figure 6.8

Graphical Representation of Left-Tailed Test at ɑ = 0.01

Z = -1.96 Z = -1.96

oc = 0.025 oc = 0.025
Figure 6.9

Graphical Representation of Two-Tailed Test at ɑ = 0.01


Figure 6.10

Graphical Representation of Left-Tailed Test at ɑ = 0.01

Table 6.2
Critical Value of Z-Test
Type of Test
One-Tailed Test Two-Tailed Test
Level of
Left-Tailed Right-Tailed
Significance
Reject Ho if z ≥ 1.645 Reject Ho if z ≥ 1.96 or
ɑ = 0.05 Reject Ho if z ≤ -1.645
or reject Ho if z ≤ -1.96

Reject Ho if z ≥ 2.575 or
ɑ = 0.01 Reject Ho if z ≤ -2.33 Reject Ho if z ≥ 2.33
reject Ho if z ≤ -2.575

Reject Ho if z ≥ 1.645 or
ɑ = 0.10 Reject Ho if z ≤ -1.28 Reject Ho if z ≥ 1.28
reject Ho if z ≤ -1.645
NOTE: The level of significance usually determines by the statistician or the researcher.

DEFINITION OF HYPOTHESIS TESTING

To determine whether to accept or reject the null hypothesis based from a sample data, a
statistician usually follows a certain process. This process is known as hypothesis testing.

Hypothesis testing is type of statistical inference, which examines the claim about a
population based from the information obtained in the random sample.

Steps in Hypothesis Testing


1. State the null and alternative hypotheses.
2. Decide the level of significance (ɑ) when it is not given in the problem.
Usually 1%, 5%, or 10% or any value between 0 and 1.
3. Select and compute the appropriate test statistic when it is not stated in the problem.
4. Compare the value if the test statistic and the critical value obtained from ɑ.
5. Make a decision.
6. Interpret the result.

Hypothesis Test About Means


In testing hypothesis about means, the following conditions must be met: first, the
sampling procedure used is simple random sampling; and second, the sample is drawn from
normal or approximately normal population.

Sampling distribution is approximately normal when any of the following conditions are
applied.

1. The population is normally distributed.


2. The sample size is at most 15 and the sampling distribution is symmetric, unimodal,
without outliers.
3. The sample size between 16 and 40 and the sampling distribution is moderately skewed,
unimodal, without outliers.

In this case, z-test will be used when the population deviation is known.

Whereas, utilize the t-test when the population standard deviation is not known in the given
distribution or problem and the number of cases is less than 30.

Using z-test, consider the following assumptions: the distribution is normal; n > 30;
known σ (z-value is the distance from the mean in relation to the standard deviation).

In this section, some of the different cases in testing hypothesis shall be discussed. The
first case is hypothesis testing about means (comparing population mean and sample mean when
n ≥ 30 or n< 30). The second case is testing difference between the last case is hypothesis testing
about two proportions.

CASE I. Hypothesis Testing About Means (Comparing Population and Sample Means)

1. z-test
ẋ−μ
z=
σ
√n

Where:
z is the z-test value
μ is the value of the population mean
ẋ is the sample mean
σ is the population standard deviation
n is the number of case, n ≥ 30

2. T-test
ẋ−μ
t=
s
√n
t is the z-test value
μ is the value of the population mean
ẋ is the sample mean
s is the population standard deviation
n is the number of case, n < 30

Example:

1. PJL Corporation is a company that produces RGC brand of laundry soap that uses a
machine to package 425 grams per pack. Assume that the net weight is normally
distributed with a population standard deviation of 8.5 grams. A researcher randomly
selected 32 packs of RGC brand of laundry soap with net weight of 430 grams. Can he
conclude that the packaging machine function properly? Test the significance at 0.05
level.
Given:
μ = 425 grams
ẋ = 430 grams
σ =8.5 grams
ɑ = 0.05
Step 1. State the null and alternative hypothesis.
Ho : µ = 425 grams or the mean weight is 425 grams.
H1 : µ > 425 grams or the mean weight is greater than 425 grams.
Step 2: Decide the level of significance (ɑ).
ɑ = 0.05
Step 3: Select and compute the appropriate test statistic when it is not stated in the
problem.
Use the z-test because the population standard deviation is given and n =
32. Consider one-tailed test because the H1 is directional alternative hypothesis.
Solution:
ẋ−μ
z=
σ
√n
430−425
z=
8.5
√32
5
z=
1.5026

z = 3.33
Step 4: Compare the value of the test statistic and the critical value obtained from ɑ.

The critical value of z = 1.645 at ɑ = 0.05 and the computed value of z = 3.33.

ẋ z = 2.33 z = 3.33

oc = .05
Figure 6.11

Step 5. Make a decision.


The computed value of z = 3.33 is greater than the critical value = 1.645,
hence the null hypothesis is rejected. In other words, the computed value of z =
3.33 lies within the rejection region.
Step 6. Interpret the result.
Based from the given information, the packaging machine functions
properly.
Example:
2. According to the supervisor if a certain factory, the mean weight of its product is 445
grams. A group of marketing researchers randomly selected 15 packs of product and
found out that the mean weight is 430 grams with standard deviation of 25.5 grams. Test
this claim at ɑ = 0.01 level of significance.

ẋ−μ
t=
s
√n

Given:

ẋ = 430
μ = 445
s = 25.5
n = 15
Step 1. State the null and alternative hypotheses.
Ho : µ = 445 grams
H1 : µ< 445 grams

Step 2. Decide the level of significance (ɑ).


ɑ = 0.01

Step 3. Select and compute the appropriate test statistic when it is not stated in the
problem.
Use the test t-test because the sample standard deviation is given and n =
15. Consider one-tailed test because the H1 is directional alternative hypothesis.

ẋ−μ
t=
s
√n
430−445
t=
25.5
√ 15
−15
t=
6.5841
t = -2.278 or -2.28, get the absolute value
t = 2.28
Step 4. Compare the value of the test statistic and the critical value obtained from ɑ.
The critical value of toc = 0.01 = 2.624, df = 14, and the computed value of t =
2.28.
Step 5. Make decision.
The computed value of t = 2.28, which is less than the critical value of toc =
0.01 = 2.624. Hence, based from the given information it failed to reject the null
hypothesis. Therefore, accept the null hypothesis.
Step 6. Interpret the result.
There is no significant difference between the population mean and the
sample mean.

CASE II. Test Between Two Sample Means


In testing hypothesis when two means are given, use either z-test or t-test. Use z-test
when the population standard deviation is known and n ≥ 30. Use t-test when the population
standard deviation is unknown and < 30. In this text, there are two cases presented.
1. z-test for independent samples
Where:
z is the z-value
ẋ 1 is the mean of the first sample
ẋ 2 is the mean of the second sample
σ12 is the population variance of the first sample
σ22 is the population variance of the second sample
n1 is the number of cases of the first sample
n2 is the number of cases of the second sample
2. t-test independent samples

Where:
t is the t-value
ẋ 1 is the mean of the first sample
ẋ 2 is the mean of the second sample
s12 is the population variance of the first sample
s22 is the population variance of the second sample
n1 is the number of cases of the first sample
n2 is the number of cases of the second sample
LINEAR CORRELATION

Correlation analysis is concerned with the relationship in the changes and movements of two
variables. The relationship has a computed value and may be visually illustrated through the
scatter diagram.

There are degree of relationship or correction between two variables.


1. Perfect Correlation (positive and negative)
2. Some degree of correlation (positive and negative)
3. No correlation

The measure of the degree of relationship or association between two variables may be further
classified as linear and non-linear. When height and weight are plotted in a graph called scatter
diagram, the relationship is a simple or linear correlation. If on the other hand, the graph of the
points is a curve the relationship is said to be non-linear.

Temperature

Volume is constant
Pressure

Height

Weight
Accidents

Some degree of negative correlation

Road width

Die A no correlation
Die B

Perfect negative correlation

Summarizing

-1 ←PerfecNegativeCorrelation
} Some Negative Correlation
0 ---No Correlation
} some positive correlation +1 ---PerfectPositiveCorrelation
7.1 Correlation Methods

There are several ways of measuring the relationship between two variables depending upon the
nature of data. The most frequently encountered measure is the Pearson-Moment Correlation
Coefficient (r).

A. Computation of the Person r from the Deviation from the Means. This is computed using
the formula.

r = ∑xy / √(∑x^2)(∑y^2)

where:

X= deviation of x – value from its mean

Y = deviation of y – value from its mean

Given:

x y x y x^2 y^2 xy

13 19 -2 -5 4 25 10

14 73 -1 -1 1 1 1

15 25 0 1 0 1 0

16 26 1 2 1 4 2

17 27 2 3 4 9 9

∑x=75 ∑y= 120 ∑x^2 = 10 ∑y^2 = 40 ∑xy = 19

After substituting the obtained values, we have:

r = 19/√(10)(40) = 19/√400 = 19/20 = 0.5

A correlation coefficient of 0.95 is very close to the perfect correlation +1. Hence it is indicate of high positive linear correlation between the two variables.

Computation of peason r from Raw Scores


B.
Another most frequently encountered method in the computation of the pearson r is the method based on the raw scores. The formula is:

r = n∑xy – (∑x)(∑y) / √((n∑x^2) – (∑x)^2 – (∑y^2)

x y x^2 y^2 xy
13 19 169 361 247

14 23 196 529 322

15 25 225 625 375

16 26 226 676 416

17 27 289 729 459

∑x = 75 ∑y = 120 ∑x^2 = 1135 ∑y^2 = 2920 ∑xy = 1819

Substituting the values , we have

r = 5(1819) – 75(120) / √ (5675-5625) (14600-1440)


= 95 / √1000
= 0.95
A correlation coefficient of 0.95 indicates degree or strong correlation between the values presented.

Correlation from Ranks


C.
A simpler method obtaining the correlation coefficient is the Spearman Rank Difference Method (RHO). It indicates the relation between two
variables each of which is arranged in rank order. It is used when the number of cases is 25 to 30 less.

Rank the first set of values giving the highest score a rank of 1.
1.
Do the same with the second set of values.
2.
Get the difference between these two ranks for each individual and record in the D column.
3.
Square each of these difference and get their sum.
4.
Substitute the value in the form.
5.

R = 1 - 6∑D^2 / n(n-1)

Example 3. For the data below, compute the Spearman Rank – Order Correlation.

STUDENT TEST X TEST Y Rx Ry D D^2

1 19 25 1 3.5 -2.5 6.25

2 18 27 2 2 0 0

3 15 29 3 1 2 4

4 12 25 4 3.5 .5 .25

5 11 21 5 5 0 0

6 9 19 6 6 0 0

7 7 14 7 7 0 0

8 5 13 9 8 1 1

9 6 12 8 9 -1 1

∑D^2=12.5
Substituting:

p = 1 - 6∑D^2 / n(n-1)

= 6(12.5) / 9(81-1) = - 75/720

= 89.58 = .90

Interpretation of the Pearson Product-Moment Correlation

The value of the Pearson Product –Moment Correlation Coefficient can be interpreted as follow: (according to Garrett).

r from .00 to ± 20 denotes indifferent, or negligible relationship.

r from ± .21 to ± .40 denotes low correlation or slight relationship.

r from ± .41 to ± .70 denotes substantial or marked relationship.

r from ± .71 to ± 99 denotes high to very high relationship.

R from ± 1 denotes perfect relationship.

From the above interpretation an r of .85 is regarded as a high correlation coefficient and r of .48 is regarded as moderate while r of .25 denotes as low correlation
coefficient.

We often think of correlation and causation as going together. This is reasonable because when one thing causes another, the two tend to be associated and therefor
correlated.

However, there can be correlation without causation. Think of this way : The correlation is just a number that reveals whether large values of one variable tend to with
large (or with small) values of the other. The correlation cannot explain why the two are associated. Indeed, the correlation provides no sense of whether the investment
is producing the return or vice versa. The correlation just indicates that the numbers seem to go together in some way.

One possible basis for correlation without causation is that there is some hidden , un observed, third factor that makes one of the variables seem to cause the other
when, iin fact, each is being – caused by the missing variable.

The term spurious correlation refers to a high correlation that is actually due to some third factor. For example, you might find high correlation between hiring new
managers and building new facilities. Are the newly hired managers causing new plant investment? or does the act constructing new building causes new managers to
hired? Probably there is a third factor, namely, high long-term demand for the firms products, that is causing both.

EXERCISE 7.1

The following table show the final grades of ten students in algebra and statistics
1.
Algebra(x) 77 84 68 98 71 87 65 93 80 75

Statistics(y) 74 89 72 95 80 91 72 86 78 82

Draw the scatter diagram.


a.
Find the correlation coefficient of Algebra and Statistics and interpret your result
b.
Ten employees in one company have the following characteristics of
2.
number of years of experience and yearly salary (given in thousands of pesos):

No. of years of:


Expeience 7 11 3 24 5 18 35 19 10
3
Year Salary 18 16 2 22 19 23 24 21 26
5

Solve the Pearson Product-Moment correlation, (r) for the data interpret the result.

3. For the data below, compute the Spearson Rank-Order Correlation.

Student Subject Subject B


A
1 29 35
2 28 37
3 25 39
4 22 35
5 21 31
6 19 29
7 17 24
8 15 23
9 16 22

4. A group of 5 students took test before and after training and obtained the following score.

Before 19 18 20 22 25
(x)
After(y) 20 22 25 30 28

Calculating the Pearson r by using deviation from the means.

You might also like