0 ratings 0% found this document useful (0 votes) 3 views 16 pages Sampling Methods
Chapter 2 discusses the importance of sampling methods in statistical data analysis, emphasizing the need for reliable data collection to inform decisions. It distinguishes between probability and non-probability sampling methods, highlighting that probability sampling allows for valid inferences about a population, while non-probability sampling does not. The chapter also outlines various sampling techniques, including simple random sampling and stratified random sampling, and their respective advantages and disadvantages.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content,
claim it here .
Available Formats
Download as PDF or read online on Scribd
Go to previous items Go to next items
Save Sampling Methods For Later
CHAPTER 2
SAMPLING METHODS
2.4 Introduction
Data collection is an integral part of statistical data analysis. To derive trustworthy conclusions from
information (data), we need to ensure that suitable and correct data collection methods are applied.
Governments, industry and society need reliable information to make better decisions and to establish an
informed society. Therefore individuals and organizations collect data because information is needed for
many purposes. For example, records for administrative purposes are kept to make decisions about
important issues, or to formulate policy, or to adapt to new situations. Whatever the specific reason, data
have to be collected to provide the needed information.
The nature of the information (data) which is collected is determined by the particular problem being
analysed and the factors associated with the study. It is usually impossible to obtain complete information
of all the possible items of interest within a particular study, usually due to lack of all or some of the following
factors: time, money, energy, equipment, labour (such as manpower), access to work places or lack of
access to the complete sampling frame. The sampling frame is defined as follows:
De ee ean ue an Cm dd
ee Ra
Some examples of sampling frames include lists of all eligible voters held by the Independent Electoral
Committee, the complete list of all matric learners writing the 2014 matric final exam across South Africa,
a list of registered participants at a conference, and the South African Revenue Service's list of all tax
payers.
In spite of the restrictions listed above, the sample must also comply with the necessary requirement that
it represents the underlying population as closely and correctly as possible to ensure that the end results
of the study are as reliable as possible, since the results relate to the unknown properties of the population
(which is what we are ultimately interested in). Before continuing our discussion of sampling and sampling
methods, we will first formally define the concepts of populations, samples, and a census:
13A population is the complete group of elements from which one would like to gain information
ee tse
Teena
The following notation for the sample size and population size will be used throughout the text:
‘The number of elements in a sample (ie.. the sample size) is denoted by n.
‘The number of elements in the population (i... the population size) is denoted by NV.
Mf a census is conducted then no statistical inference (See Chapter 1 for the definition of inference) is
Fequired since a census will reveal all the information inherent in the population. Descriptive methods will
Still be required to represent and order the data.
In general it is reasonable to expect that reliable conclusions conceming a population can only be made
from samples that are representative of the population for the variables being studied. Therefore, samples
cannot be simply chosen in any arbitrary fashion. The composition and nature of the population, as well
as some other properties of the population, will influence the choice of the sampling procedure. For
example, one should know who or what exactly is included in your population of interest and to whom or
which group you want to generalize your results to. If a sample of doctors is drawn from all doctors working
in Potchefstroom, do we wish to generalise our results obtained from this sample to all doctors in
Potchefstroom, or to all doctors in North West? For a sample of 50 women in the age group 30-40, can
the results extracted from this sample be generalised to all women, or only to all women aged 30-407
This chapter provides an overview of some of the more popular sampling procedures which are used in
practice. Advantages and disadvantages of each procedure are also briefly discussed.
4MM SELF-EVALUATION EXERCISES
‘Study the two concepts
1. Population
i. Sample
With regard to the relation between the two concepts given above, which one of the following statements
is correct?
i. tis part of H.
usually has more items than |.
{is studied more intensely than Il.
‘The knowledge of Il is used to gain knowledge of I.
J and Il have no connection.
(“sz 5F
2.2 Sampling methods
From the previous section it should be clear that sampling plays an important role in statistical analysis
and its implementation should be carefully considered and planned. A sample survey costs less than a
census and results are obtained far more quickly for a sample survey than for a census because fewer
units are consulted and less data need to be processed.
‘A sample should be representative of the population no matter what the circumstances regarding the
Population may be or which sampling method is applied. In order to make a sample representative, close
attention should be paid to the procedures involved for sampling and analyses of the data, since insufficient
consideration regarding this aspect of statistics will lead to untrustworthy results and often disastrous
Consequences, such as unscientific, false conclusions from wrong results. Therefore, selecting the best
sampling method is @ primary step in statistical analyses, as well as using the correct sample size.
Insufficient attention to sampling and sampling methods can also mean that conclusions drawn from these
samples are scientifically questionable and so the best method for data collection must be selected for
each case. Keep in mind that cost and data quality will be directly impacted by the method you choose.
Since every survey will differ from almost every other survey, there are no strict rules for determining the
size of the sample required. The factors that will influence size of the survey operations are time, cost,
operational constraints and the desired precision of the results. Itis important to evaluate and assess each
of these issues in order to determine suitable sample sizes. Determining the sample size will be discussed
in Chapter 12.
‘Sampling methods can be divided into two groups, namely probability procedures and non-probability
procedures.
Probability sampling involves drawing a sample from a population based on the principle of randomization
‘or chance. Probability sampling is more complex, more time-consuming and usually more costly than non-
probability sampling. However, because units from the population are randomly selected and each unit's
15probability of inclusion can be calculated, reliable estimates for population parameters can be produced,
and inferences can be made about the population (The terms “parameter” and “estimates” will be
discussed in later chapters.)
In non-probability sampling, since elements are chosen using subjective methods, there is no way to
estimate the probability of any one element being included in the sample. This inability to calculate these
probabilities prevents valid inferences from the sample to the population being considered. Statisticians
are reluctant to use these methods, but in some simple situations these non-probability methods can
Useful, quick, inexpensive and convenient,
The difference between probability and non-probabilty sampling has to do with a basic assumption about
the nature of the population under study. In probability sampling, every item has a chance of being
selected. In non-probability sampling, there is an assumption that there is an even distribution of
characteristics within the population, The researcher thus believes that any sample would be
Tepresentative of the population with respect to the variable being considered. This means that, if the
assumption is true, his non-probability sample will be trustworthy. For probability sampling, randomization
is a feature of the selection process, rather than an assumption made conceming the structure of the
population.
Deon eRe oe cy
pee Rt ear eect an es
A non-probability sample is taken by making use of more subjecth
The sampling procedures which will be discussed can be classified as follows:
© Simple random sampling * Convenience sampling
© Stratified random sampling Judgement sampling
* Clustered sampling * Quota sampling
Several other probability sampling methods exist, such as, for example systematic sampling and multi-
phase sampling, but these will not be discussed in this text.
162.2.1. Simple random sampling
Simple random sampling is the most general probability procedure since the principle used here is a’~
found, and used, in stratified sampling and clustered sampling. In order to randomly select an elem» ——
from a population, one must assume that each element in the population has the same chance of bel
chosen, Also, each combination of members of the population has an equal chance of composing the
sample, No element should be favoured above another in the selection process,
ee a eee ee a me
Re eu ee ene cum nen Sac
‘Any sampling method exhibiting the following properties, can be classified as a simple random sample
+ The population should consist of NV objects,
+ the sample should consist of n objects and
* all possible samples of n objects should be equally likely to occur.
Example 2.1: The national lottery draw, where a sample of 6 numbers is randomly generated from a
population of 49, is a good example of simple random sampling. Each number has an equal chance of
being selected and each combination of 6 numbers has the same chance of being the winning
‘combination. Even though people tend to avoid combinations such as 1-2-3-4-5-6, it has the same chance
of being the winning set of numbers as the combination of 8-15-21-28-32-40 for example.
on
Whenever personal preference comes into play (both consciously and subconsciously) when drawing
samples it will, almost without exception, lead to a non-random sample and so some mechanical system
for selecting these samples is thus preferable. In practice, if you wanted to select a simple random sample
using one of these mechanical systems, you would need to first construct a list all of the units in the
population,
One example of producing a simple random sample is by applying the so-called lottery method (see
‘example 2.1) for choosing n objects from a group of size NV. Here, each of the Nv population members is
assigned a unique number. The numbers are placed in a container whereafter it is shuffled thoroughly.
Then, n numbers are blindly selected randomly and independently, identifying the numbers of the
population members selected to be included in the sample. If it is agreed that when a number is chosen,
it cannot be repeated, the method is referred to as sampling without replacement. If it is agreed that a
number may be chosen more than once, the number must be available in the container after being drawn,
each time. This process is known as sampling with replacement.
‘Another such mechanical method is to make use of generated random numbers. Random numbers are
simply a random ordering of the numbers 0, 1, 2...9. Table A1 in Appendix A is an example of a table of
such numbers.
7Table 2.1; Extract from Table A1
96599 17254 79613
37448 66591 15245
90991 48809
59414 79005
78962
Table 2.1 represents an extract from Table A1. The grouping of the numbers is only used to aid readability
and for no other reason, The numbers are used in the following manner;
Suppose a sample of 10 elements must be obtained from a finite population of 800 elements, and that
each element in the population has a unique number assigned to it. Suppose further that the numbers
assigned to the elements in the population are the numbers beginning at 001 and ending at 800. The first
step is then to arbitrarily choose a starting point in the random number table. Starting at this point, read
the numbers in groups of three, The first 10 numbers found in this way which are both smaller than 800
and non-repeating (.¢., there are no two numbers which are the same in the group of 10 numbers),
represent those numbered values in the population which will be included in the sample.
Note: in the above example the numbers are read off in groups of three because the population size (NV)
was a three digit number. If N’ was, say, a five digit number, then one would make use of groupings of five
random numbers and so on.
Example 2.2(a): A random sample of size five is to be taken from a population of 250 people (N = 250).
The following two rows of random numbers are to be used:
127356127345955561020301771
345156314751245625890242345
By beginning in the first row on the left hand side, the following three-digit number groups are obtained:
127, 356, 127, 345, 955, 561, 020, 301, 771, 345, 156, etc. (Remember that NV has 3 digits). Some of the
numbers obtained are greater than 250, making them useless for this population. f some of the numbers
repeat themselves, and we keep the original and repeated numbers, this sampling is equivalent to
sampling with replacement. Usually the repeating numbers are removed, implying sampling without
replacement. The remaining numbers then represent the numbers of the elements in the population which
should be chosen from. The numbers of the five people to be included in the sample are: 127, 20, 156,
245 and 242
ann
18One can also easily generate random numbers by making use of the “random” function on a pocket
calculator or by using various computer packages.
The random numbers in Table 2.1 can be read-off in a number of different ways. The next two examples
illustrate two other possible ways of doing this.
Example 2.2(b): Suppose a random sample of size four is to be taken from the same population of 250
people (N=250). As before, the following two rows of random numbers are to be used:
127356127345955561020301771
345156314751245625890242345
We now start off in the first row on the right hand side, reading right to left, taking one number at a time
until 3 numbers are identified. Then the following three-digit number groups are obtained: 177, 103, 020,
165, 559, 543, 721, 653, 721, etc. Remember that N has 3 digits. Again, numbers greater than 250 and
numbers that repeat will be discarded (i.¢., we are sampling without replacement), but the first occurrence
of a repeating number is retained, The remaining numbers represent the numbers of the elements in the
population which should be chosen from. The numbers of the four people to be included in the sample
are then: 177, 103, 20, 165.
an
Example 2.2(c): If the numbers were chosen from the right, in groups of three at a time, the following
three-digit number groups are obtained: 771, 301, 020, 561, 955, 345, 127, 356, 127, 345, 242, 890, 625,
245, 751, 314, 156, 345. To select the four desired numbers, we had to continue to the second row, from
the right side, three numbers at a time, The four numbers in case of without replacement sampling, are
20, 127, 242 and 245. With replacement sampling will produce the numbers 20, 127, 127 and 242.
These exampies illustrate that an agreement is needed on how to use the generated random numbers
before sampling is attempted.
Example 2.3: To draw a simple random sample from a home owners directory, each entry would need to
be numbered sequentially. If there were 10 000 entries in the directory and if the required sample size was
2.000, then 2.000 numbers between 1 and 10 000 would have to be randomly generated by a computer
or from Table A1. Each number should have the same chance of being generated by the computer (in
‘order to fulfil the simple random sampling requirement of an equal chance for every unit). The 2.000 home
‘owners corresponding to the 2 000 computer-generated random numbers would make up the sample.
19‘Simple random sampling can therefore be done with or without replacement. A sample with replacement
‘means that there is a possibility that the sampled elements may be selected twice or more. Usually, the
simple random sampling approach is conducted without replacement because it is more convenient and
gives more precise results. For the purposes of this text, we will only make use of sampling without
replacement
‘Simple random sampling is the easiest method of sampling and itis the most commonly used. Advantages
of this technique are that it does not require any additional information other than the complete list of
members of the survey population. Disadvantages include that the method can lead to unsatisfactory
representation of the population or area if large areas are not reached by the random numbers generated,
for whatever reason.
2.2.2 Stratified random sampling
Stratified random sampling is used when a population exhibits a large amount of heterogeneity (i.e., when
the elements in the population exhibit /arge differences with respect to a particular property being studied).
To create this type of sample one must begin by first dividing the population into a number of “natural” and
non-overlapping subgroups or strata which are more or less homogeneous with respect to the property
being studied. Usually the sub-sets have known size. From within each of these strata a few elements
are randomly selected. The number of elements chosen from each group may be different. If the number
of elements chosen from each stratum for the sample is proportional to the number of elements with each
stratum of the population then the sample is known as a stratified random sample.
Example 2.4: If a large area of forest is the study site, different types (sub-sets) of trees exist within the
study area. Random sampling may altogether miss or partially miss one or more of these groups. Stratified
‘sampling would take into account the proportion of the total area occupied by each type of tree within the
total study area of plantations; each tree type could then be sampled proportionally to ensure each type
taken up into the sample. For example, aerial photography shows that 60% of the study area is covered
by eucalyptus trees, 25% by pine trees and 15% by shrubby natural bushes. If a sample of 1000 trees
must be selected from the area for some research reason (e.g., investigating their vulnerability or
resistance to certain diseases), 60% of the 1 000, |.e. 600 plants, must be eucalyptus trees, 25% of the
1 000, i.e. 250, must be pine trees and 15% of the 1 000, i.e. 150, must be shrubby natural bushes (the
formulas used to determine the correct number of items to be selected from each stratum are provided
immediately after this example).
This process may not be easy. One way to select 600 eucalyptus trees in a large area can be achieved
approximately correctly by using some graphical scheme, by dividing the eucalyptus area from aerial
photos into small sections which include only one tree (the average ground area covered by a full-grown
eucalyptus tree can be determined before the time). By numbering these small sections, 600 numbers can
now be selected from the estimated total of eucalyptus trees which may be a large number. This will of
20course not be exact since some trees are small and some large and wide. But at least we will have
identified 600 random points in the eucalyptus plantations. Once the researcher is on the ground at the
particular spot, the nearest tree to the selected point can be included in the study, The same method can
bbe repeated for the other groups, to produce a final approximately representative sample.
oun
In the previous example the number of trees drawn from each stratum (tree type) was determined
proportional to the sizes of each stratum. The following general formula can be used to determine these
sample sizes.
‘Suppose a population of size N can be divided into L non-overlapping strata. Denote the total number
‘of population elements in each of these L. strata by Ny, Nz, ...N,-
‘Suppose now that we want to draw a proportional stratified sample of size n. Let the sample sizes
drawn from each of the stratum be denoted by m,,mz,...,m%, then
ama, [AQ ark
Note that the sum of the sample sizes in each stratum is equal to the total sample size that was
requested, that is, ny +n +--+, =n.
Example 2.5: Suppose the inhabitants of a small town (with population size N = 5 000) can be divided
into four natural subgroups or strata, The number of people in each stratum is as follows:
‘School Learners 1422
University Students 521
Middle-aged People 2730
Pensioners 327
‘The favourite recreational place of the people within each of these four strata is of interest. It was decided
that this phenomenon would be studied by creating a stratified random sample of size n = 800. Each of
the four groups above represents a stratum.
In order to determine the stratum sample size (n,), i.¢., the number of elements to draw from each stratum,
the formula described above is used,
21Ny
menxst, 121,234,
N
where:
ny = sample size drawn from stratum /,
N, = Population stratum size of stratum J.
n= Total sample size
N= Population size
For School learners:
‘The population stratum size is N; = 1422, and so the sample size drawn from this stratum is:
N
ny =n x 2 a0 x 142 2228,
N 5000
Now a simple random sample of size 228 is obtained from the 1 422 school learners.
Note that, since sample size cannot be a fraction, we need to round off to the nearest whole number when
reporting the sample sizes.
For University Students:
‘The population stratum size is N, = 521, and so the sample size is
521
5000
m= nx = 800% = 83,
The following table summarizes the sample sizes for these and the remaining strata:
Table 2.2: Sample sizes of the four strata
Stratum Stratum size | Sample size
1. School Learners 1422(N,) | 228 (m1)
2. University Students 521 (N2) | 83 (m2)
3. Middle-aged People | 2730(Ns) | 437 (ns)
Pensioners 327 (Ns) 52 (ms)
Example 2.6: A sports analyst is conducting an opinion survey, sampling from a list of 10000 recent
bicycle buyers. The list includes 4 types of bike buyers: 2500 Typet buyers, 2500 Type2 buyers, 2500
Type buyers, and 2500 Type4 buyers. The analyst selects a sample of 400 bike buyers, by randomly
sampling 100 buyers of each brand. Is this an example of a simple random sample? Choose the correct
alternative below.A , because each buyer in the sample was randomly sampled.
(8) Yes, because each buyer in the sample had an equal chance of being sampled.
(C) Yes, because bike buyers of every brand were equally represented in the sample.
(0) No, because every possible 400-buyer sample did not have an equal chance of being chosen.
(E) No, because the population consisted of purchasers of four different brands of bicycle.
The correct answer is (D). A simple random sample requires that every possible sample of size n (in this
problem, n is equal to 400) have an equal chance of being selected. In this problem, there was a 100
Percent chance that the sample would include 100 purchasers of each brand of bike. There was zero
Percent chance that the sample would include, for example, 99 Typet buyers, 101 Type2 buyers, 100
Type3 buyers, and 100 Typed buyers. Thus, all possible samples of size 400 did not have an equal chance
of being selected; so this cannot be a simple random sample.
ona
‘Advantages of stratified sampling include the fact that it can be used not only with random sampling, but
also with sampling methods such as the systematic sampling method, which will be discussed in later
courses. Furthermore, if the proportions of the sub-sets are known, the results can be more representative
of the whole population. The method is very flexible and applicable to many areas and fields.
2.2.3. Clustered sampling
Clustered sampling is applied when the population elements are naturally grouped together to form so-
called homogeneous groups referred to as clusters, whose internal distributions are as heterogeneous as
the population itself. This implies that all clusters are representative of the population (therefore they are
homogeneous among each other) but they all include all type of elements in similar proportions as the
population (internally heterogeneous). Each cluster is a smaller version of the population. For example,
sometimes it is too expensive to draw a sample from the population as a whole. Travel costs can become
expensive if interviewers have to survey people from one end of the country to the other. To reduce costs,
statisticians may choose a cluster sampling technique. Cluster sampling begins by dividing the population
into groups or clusters, then a number of clusters are selected randomly to represent the total population,
and finally then all elements within the selected clusters are included in the sample. Elements that belong
to clusters that were not selected are thus not included in the sample — it is the hope that these unselected
elements will be fairly represented by those units from the selected clusters, This differs from stratified
sampling, where some units are selected from each group. This method is therefore largely aimed at
reducing the costs associated with sampling. The first stage of the sampling method consists of randomly
selecting a number of clusters. As we mentioned above, all the elements within the chosen clusters are
used in the sample (in the case of one-stage clustered sampling), or one can go further and randomly
select elements from within each cluster to form the sample. The latter total process is known as two-
stage clustered sampling. It is usually better to survey a large number of small clusters instead of a small
number of large clusters.
Examples of clusters include factories, schools and geographic areas such as electoral subdivisions, The
selected clusters are used to represent the population.
23Example 2.7; The Mathematics grades of grade 12 students in rural and urban schools are being studied.
It is advised that two-stage clustered sampling be used since the rural and urban schools’ students are
expected to be reasonably homogeneous with respect to the variable being studied. Schools are thus
presented as clusters. The sample elements are obtained by first choosing a number of schools randomly,
and then by choosing a number of grade 12 students from within each of these randomly chosen schools.
In this way even a large sample can be chosen with relatively low costs because it will only be necessary
to visit the chosen schools and not every schoo! in the study, thus reducing travelling and lodging costs of
the fieldworker.
an
Example 2.8: Suppose you are a representative from the Sports Federation wishing to find out which
sports 16-year-old school leamers are participating in across the country. It would be too costly and lengthy
to survey every 16-year-old leamer, or even a couple of students from every school's 16-year-old learners.
Instead, 100 schools are randomly selected from all over the country. These schools represent the
clusters. Every 16-year-old leamer in all 100 clusters is then surveyed. In effect, the students in these
clusters represent all 16-year-old leamers in the country.
Example 2.9: Imagine that the municipal council of a small city wants to investigate the use of health care
services by residents. First, the council obtains electoral subdivision maps that identify and label each city
block. From these maps, the council creates a list of all city blocks, Every household in the city belongs to
a city block, and each city block therefore represents a cluster of households. The council randomly picks
a number of city blocks. Using the simple random sample approach, the council then creates a list of all
households in the selected city blocks; these households make up the sample. For two-stage sampling, a
sample can be taken randomly (or in another responsible way) from within the selected clusters.
Strata and clusters are both non-overlapping sub-sets of the population but they differ in several ways:
sub-samples of all strata are represented in the final sample but only a subset of clusters are part of the
final sample. With stratified sampling, the strata are intemally homogeneous and with cluster sampling,
the clusters are internally heterogeneous.
24MM SELF-EVALUATION EXERCISES
Question 1: A statistician is interested in the average height of grade 12 learners in South Africa. He
drew a sample by randomly selecting three South African schools, and his sample consisted of all of the
grade 12 leamers in each of these three schools. What type of sampling did he use?
i, Simple random sampling.
ii, Stratified sampling,
ili, Quota sampling.
iv. Cluster sampling.
‘Question 2: Which of the following sampling methods is/are examples of random sampling?
|. An auditor chooses the ten biggest transactions everyday to check the business's books.
ii. By using numbered balls which are drawn blindly, a sample of 200 statistic students are drawn from
all statistics students on the PUK to determine statistics students’ attitude towards the subject.
ii, Questionnaires are handed out to all Law (LLB) students. The aim of the questionnaires is to
determine the knowledge of all PUK-students about The South African Collaboration of Human
Rights.
iv. To determine the effect of a certain substance which is meant to decrease blood pressure, the 20
best athletes were administered the drug and then studied.
Question 3: There are 730 people working at a specific company, of which 482 are male. A proportional
stratified random sample of size 80 is to be drawn from this set of employees. The gender of the
employees is used as strata.
Calculate the number of women that should be included in the sample.
Question 4: Match column A to column B:
A 8
Is obtained if each element in the population
Sample has @ known equal chance of being taken
Up into the sample.
Probability sample ‘A subset of the population.
1s obiained if each element in the
Population, not already in the sample, has
‘Simple random sample
ml aan equal chance of being taken up into the
‘sample on the next draw,
252.3 Non-probability sampling methods
The three non-probability sampling methods will now be discussed, namely:
* Convenience sampling
‘Judgement sampling
* Quota sampling
Convenience sampling is a method that basically involves selecting a sample which is most convenient
for the sample taker. It is not normally representative of the target population because sample units are
only selected if they can be accessed easily and conveniently. The average person will often find
themselves making use of convenience sampling. A food critic, for example, may try several appetizers or
entrees to judge the quality and variety of a menu. Television reporters often seek so-called ‘people-on-
the-street interviews’ to find out how people view an issue. In both these examples, the sample is chosen
arbitrarily, without use of a specific survey method. The obvious advantage is that the method is easy to
use, but that advantage is greatly offset by the presence of bias. Although useful applications of the
technique are limited, it can deliver accurate results when the population is homogeneous. For example,
a scientist could use this method to determine whether a lake is polluted. Assuming that the lake water is
well-mixed, any sample would yield the same information. A scientist could safely draw water anywhere
on the lake without fretting about whether or not the sample is representative, Also, in Example 2.7 the
fieldworker could just visit the schools in Pretoria because those schools might be closer to his own office.
Judgement sampling is used by researchers to include only the “best” sample elements, These “best”
elements are determined by the researcher's own subjective judgement. This approach is used when a
sample is taken based on certain judgments about the overall population. The undertying assumption is
that the investigator will select units that are representative (or characteristic) of the population, The critical
issue here is objectivity: how much can judgment be relied upon to arrive at a typical sample? Judgment
sampling is subject to the researcher's biases. Since any preconceptions the researcher may have are
reflected in the sample, large biases can be introduced if these preconceptions are inaccurate. Statisticians
often use this method in exploratory studies like pre-testing of questionnaires. They also prefer to use this
method in laboratory settings where the choice of experimental subjects (.e., animal, human, vegetable)
reflects the investigator's pre-existing beliefs about the population. One advantage of judgment sampling
is the reduced cost and time involved in acquiring the sample.
Quota sampling is also an artificial method and requires that the population be divided into segments. A
quota system is then implemented to ensure that a number of elements from each segment are included
in the sample. There are a number of ways in determining the quota, but prescribed methods do exist. The
quotas may be based on population proportions. For example, if there are 100 men and 100 women in a
population and a sample of 20 are to be drawn to participate in a wine taste challenge, you may want to
divide the sample evenly between the sexes, namely 10 men and 10 women. Quota sampling can be
considered preferable to other forms of non-probabilty sampling (e.g., judgment sampling) because it
forces the inclusion of members of different sub-populations.
262.4, Notes on the responsible application of Statistics in practice
+ Nole that the non-probability methods are more subjective in nature and thus caution should be
exercised when reporting results generated by these studies.
+ Sampling from populations of people is a possible source of ethical problems. Since the elements
involved in these situations are real people the data acquired is considered sensitive. It is sometimes
necessary, in the interest of reducing bias, to limit the amount of information a respondent receives.
For example, if a candidate for a political party is interested in taking an opinion poll, itis not advisable
to tell the respondents who paid for the poll since this may influence the answers made by the
respondent. Nevertheless, the respondent should be given information, such as the address and
telephone number, of the organisation conducting the study since this gives him open channels for
feedback and complaints. Courtesy also demands that an indication of the allotted time for the study,
the type of information required and the confidentiality of the information be provided.
* tis important that one does not promise anonymity when one only wants to promise confidentiality.
There is a difference. Anonymity is only possible when a response is given without any form of
identification. Therefore, anonymity implies that there is absolutely no chance of a follow-up study to
try and gain responses from people who did not provide answers the first time round. Thus itis usually
more sensible to only promise confidentiality, i.e., no one individual's response will be made public,
but it will be used along with many others in some form of calculation. It is unethical to promise
anonymity when only confidentiality is intended. Methods (usually used by market researchers) that
involve hidden codes in the questionnaire used to identify individuals are naturally also unethical.
Honesty and courtesy in the sampling process are necessities.
* When handling data and related aspects honesty and objectivity are very important. There are many
ways of representing and interpreting data in order to suit one’s own selfish needs. A great
responsibility rests on the shoulders of the statistician to distinguish between that which is wrong and
that which is right when handling data. Also one must respect the confidentiality of the data. Many
such ethical problems exist in the realm of statistical research and, where applicable, these issues will
be addressed in the remainder of the study material.
2.5 Errors and bias in sampling
In the case where different samples of the same size are taken from a single population and the sample
mean is calculated for each one of these samples it will be found that this sample mean will differ from
sample to sample (but usually not by much). It is necessary to define a measure that describes the
difference between the sample means and that of the true population mean. This measure is known simply
as the sample error.
Seay
Ce et
27The sample error is usually expressed in terms of the standard error of estimation, which will be discussed
later in this text book. In most cases, the degree of sample error depends on the sample size, n. From the
previous description it is clear that the sample error cannot be controlled because it depends on other,
uncontrollable factors. However, it can be made somewhat smaller by choosing to draw larger samples
from the population.
A second type of error which is of some importance is the sampling observation error. This error has to do
with faulty measurements, unreliable questionnaires, and unclear responses. Sampling observation error
is primarily attributed to “human error” and is unrelated to the sampling method employed.
Bae eee eee eS
Ua RCA Rue kel eau Cay
Hee ais ST
‘Sample bias is another factor which should always be considered. There are a multitude of situations,
which will be discussed throughout the course, which can cause this sort of error. in short it will be defined
as follows:
SE eee Ree nec eee eee) One
Cec Ee
MM SELF-EVALUATION EXERCISES
Consider the 3 errors that can occur during sampling in column A as well as the possible solutions for
these errors in column B. Match column A to column 8:
A 8
of
ps proper planning of sampling
‘Sample bias choose a large sample size
‘Sampling observation error well-planned questionnaires.
28