0% found this document useful (0 votes)
4 views42 pages

Chapter 10sampling

Chapter 10 focuses on sampling methods in research, discussing the differences between probability and non-probability sampling, and their implications for generalization and bias. It outlines various sampling techniques, including simple random, stratified, and purposive sampling, and highlights the importance of selecting an appropriate sample to ensure accurate representation of the population. The chapter also addresses potential sources of error and bias, emphasizing the need for careful planning in the sampling process.

Uploaded by

jonkc
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views42 pages

Chapter 10sampling

Chapter 10 focuses on sampling methods in research, discussing the differences between probability and non-probability sampling, and their implications for generalization and bias. It outlines various sampling techniques, including simple random, stratified, and purposive sampling, and highlights the importance of selecting an appropriate sample to ensure accurate representation of the population. The chapter also addresses potential sources of error and bias, emphasizing the need for careful planning in the sampling process.

Uploaded by

jonkc
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 10Sampling

10.1 CHAPTER GUIDE

This chapter deals with sampling issues: how to select a sample – the segment of the
population that you select for your research – and the issues that need to be considered
when assessing what can be inferred from different kinds of sample. This chapter also
illustrates the variety of sampling methods that researchers may decide to employ
according to their research aims and questions, whether they follow quantitative,
qualitative or mixed methods approaches.

Since we are dealing with general principles of sampling, the chapter begins with a brief
comparison of the principles and practices associated with probability sampling and
non-probability sampling and an overview of other sampling basics.

We then explore the various methods used in probability sampling, which enables the
researcher to generalize their findings to a well-defined population from which they drew
the sample. This is associated primarily with quantitative research methods like surveys
but can also be used in qualitative research.

The main sampling methods used by quantitative researchers can be used when
qualitative researchers sample people, documents or organizations, but qualitative
researchers tend to use non-probability sampling. In general, qualitative researchers
tend to emphasize purposive sampling, which focuses attention on answering their
research questions, not generalizing findings to a population. We therefore also
consider the various non-probability sampling methods.

Consequently, when dealing with sampling methods in later chapters, you will need to
refer back to this chapter to plan your sampling. The chapter explores the following
specific sampling issues:

probability sampling – that is, sampling that applies a random selection process

the main types of probability sampling – simple random sampling, systematic sampling,
stratified random sampling and multi-stage cluster sampling, which are typically used in
quantitative data analysis and structured observation

the related ideas of generalization (also known as external validity) and of a


representative sample, which allows the researcher to
generalize findings from a probability sample to the population from which it was drawn

potential sources of error, which is important in survey research

different types of non-probability sampling, including strategic purposive sampling as the


master concept around which different sampling approaches mainly used in qualitative
research can be distinguished, such as theoretical, quota and convenience sampling

more details relating to theoretical sampling, which is a key ingredient of the grounded
theory approach and used in ethnography, and the nature of theoretical saturation,
which is one of the main elements of this sampling strategy

the main issues involved in deciding on sample size

sampling in a variety of specific research methods like online surveys, ethnography and
participant observation, and content analysis.

Key concept 10.1

Icon representing a key concept, showing a cloud with a brain inside.

Basic terms and concepts in sampling

Universe – an undefined group of units, for example, entrepreneurs in Johannesburg.

Population – a defined set of units in the universe from which the sample will be
selected – for example, registered small firms in Gqeberha.

Sampling frame – a list of all elements in the population from which the sample will be
selected

Sample – the segment or subset of the population that is to be selected for


investigation, using probability or a non-probability approach (see lower down in this
list).

Representative sample – an accurate reflection of the demographics of the population.

Sampling bias – distortion that occurs if some members of the population (that is, in the
sampling frame) are likely to be excluded from the sample.

Probability sampling – a selection method so that every unit in the population has a
known chance of being selected. A researcher is more likely to get a representative
sample using this method of selection. The aim of probability sampling is to
keep sampling error (see lower down in this list) to a minimum.

Non-probability sampling – a selection method where units have not been selected at
random, implying that some units in the population are more likely to be selected than
others.

Sampling error – the difference between the profiles of those in the sample and the
population from which it is selected.

Non-response – a problem that occurs whenever some members of the sample refuse
to cooperate, fail to answer, cannot be contacted or cannot supply the required data.

Non-sampling error – differences between the population and the sample because of (a)
deficiencies in the sampling approach, an inadequate sampling frame or non-response
or (b) problems such as poor question wording, poor interviewing or flawed data
processing.

Random selection – a method to select participants at random from a sampling frame.

Census – the collection of data from all units in an entire population. As a result, the
researcher collects census data from the entire population, not a sample.

Universe

Population

Sample

Figure 10.1 Sampling terms

10.2 INTRODUCTION TO SAMPLING

Imagine a scenario where you are interested in finding out about the behaviour,
attitudes and backgrounds of fellow students on any subject.

You might consider sending out questionnaires or conducting structured interviews. You
will have to consider how best to design your interviews or questionnaires and how to
administer them. You will also need to decide how to select the students who will be
part of your study. If your campus is quite large and has around 15 000 students or
more, the only practical way to conduct your study will be by selecting a sample from
the total student population.

If you want to be able to generalize your findings to the whole population, the sample
needs to represent the diversity of the students on your campus. Alternately, you might
decide to go to the cafeteria and approach students who seem to be friendly. This is an
example of a convenience sample (see section 10.4.2), which is unlikely to represent
the population as you will exclude students who do not visit the cafeteria as well as
those who you might regard as unapproachable, perhaps based on their race or gender.
As a result, the sample you collect will be biased.

In this chapter, we describe probability and non-probability sampling separately as both


apply to quantitative and qualitative methods. In mixed methods approaches,
researchers should select an appropriate combination of sampling methods.

10.2.1 Probability and non-probability sampling

Sampling is a critical stage of research as you have to make fundamental choices. You
could approach this stage as having a menu of choices. Your choices are between two
basic approaches, as indicated in Table 10.1 where you also find the main types of
sampling associated with each approach.

10.2.2 Sampling process

Where does sampling fit into the research process that we have already covered? The
researcher identifies a field of interest, scans the literature identifies a specific research
topic, formulate research questions and reviews the relevant literature. Then the
planning of the research

Table 10.1 Overview of key sampling concepts

Probability sampling

Non-probability sampling

Definition

A sample is selected using random selection, so that each unit in the population has a
known chance of being selected.

A representative sample is more likely the result when this method of selection is used.

The aim of probability sampling is to keep sampling error (see above) to a minimum.

A sample that has not been selected using a random selection method. This implies that
some units in the population are more likely to be selected than others.

Sample types

• Simple random sampling

• Systematic sampling
• Stratified random sampling

• Multi-stage cluster sampling

• Quota sampling

• Purposive sampling

• Convenience sampling

• Theoretical sampling

• Opportunistic sampling

• Snowball sampling

design can begin. For example, if a social survey is an appropriate research design, the
next stage of planning of fieldwork, including sampling, can begin.

Figure 10.2 outlines the main steps involved in doing survey research from identifying
the general research issues – through the steps described in Chapter 7 – up to
formulating the findings. The different steps of a survey, other than those to do with
sampling, which will be covered in this chapter, will be addressed in Chapter 11.

Although we present the decisions related to sampling and the research instrument
sequentially in Figure 10.2, in practice these decisions will overlap. The survey
researcher needs to (a) decide what population is suited to the investigation of the topic,
(b) formulate a research instrument – for example, a structured interview schedule or a
self-completion questionnaire and (c) decide how to administer the research instrument.
Figure 10.3 outlines the main ways that surveys can be administered.

10.2.3 Sampling bias

A biased sample is one that does not represent the population from which the sample
was selected. If your decisions about whom to sample are influenced primarily by the
availability of possible respondents, by your underlying criteria for inclusion or by
personal judgements, then your sample is likely

to be biased. In practice, it is incredibly difficult to remove bias altogether and to get a


truly representative sample. You need to ensure that you keep bias to an absolute
minimum.

Three key sources of bias are the following (see Key concept 10.1 for an explanation of
the terms):
Use of a non-probability or non-random sampling method. When the method is not
random, human judgement will affect the selection process, making some members of
the population more likely to be selected than others. This source of bias can be
eliminated through the use of probability or random sampling, the procedure described
in section 10.3.

An inadequate sampling frame. If the sampling frame or list is deficient, for example, if it
is not comprehensive or accurate, the sample cannot represent the population, even if a
random or probability sampling method is used correctly.

Non-response bias occurs when some sample members refuse to participate or cannot
be contacted. Those who agree to participate may differ in various, possibly significant,
ways from those who do not participate. It may be possible to compare the
demographics of those who responded and the sample as a whole. However, usually it
is impossible to determine whether or not there are differences in terms of 'deeper'
factors, such as attitudes or patterns of behaviour.

graph TD 1[1. Review literature/theories relating to topic.] --> 2[2. Formulate research
questions.] 2 --> 3[3. Decide on an appropriate survey method (face-to-face, Internet,
etc. – see Figure 10.3).] 3 --> 4[4. Decide on an appropriate population.] 4 --> 5[5.
Decide on the sample design (probability or non-probability).] 5 --> 6[6. Establish if there
is an accessible sampling frame (or list).] 6 --> 7[7. Evaluate survey types and select
one (see Figure 10.3).] 7 --> 8[8. Develop survey questions (and answers for closed-
ended questions).] 8 --> 9[9. Pilot-test and revise the questionnaire if necessary.] 9 -->
10[10. Select the sampling method.] 10 --> 11[11. Administer questionnaire to sample.]
11 --> 12[12. Capture answers for data analysis and interpretation.]

Flowchart illustrating the 12 steps in conducting a social survey. The steps are: 1.
Review literature/theories relating to topic; 2. Formulate research questions; 3. Decide
on an appropriate survey method (face-to-face, Internet, etc. – see Figure 10.3); 4.
Decide on an appropriate population; 5. Decide on the sample design (probability or
non-probability); 6. Establish if there is an accessible sampling frame (or list); 7.
Evaluate survey types and select one (see Figure 10.3); 8. Develop survey questions
(and answers for closed-ended questions); 9. Pilot-test and revise the questionnaire if
necessary; 10. Select the sampling method; 11. Administer questionnaire to sample; 12.
Capture answers for data analysis and interpretation. A feedback arrow connects step 6
to step 7.

Figure 10.2 Steps in conducting a social survey

10.2.4 Main reasons for using a sample rather than a census

For a small-scale research project, it might seem feasible to sample the whole
population, for example, all workers in a small factory. However, in a larger research
project, for example, of workers in all factories in a town, it might not be feasible to
involve every worker. Cooper and Schindler (2011) identify the following reasons for
using a sample rather than a census:

The cost will be lower.

The quality and accuracy of the results will be improved because of better, more in-
depth interviewing and better supervision of the process.

The time between identifying the need for the information and the collection of the data
will be reduced.

Access to the entire population via a census is only feasible when the population is
small and the elements are clearly distinct from one another.

graph TD Survey --> Structured interview Survey --> Self-completion questionnaire


Structured interview --> Face-to-face Structured interview --> Telephone Face-to-face --
> P1["Paper + pencil 1"] Face-to-face --> C2["CAPI 2"] Telephone --> P3["Paper +
pencil 3"] Telephone --> C4["CATI 4"] Self-completion questionnaire --> S5["Supervised
5"] Self-completion questionnaire --> P6["Postal 6"] Self-completion questionnaire -->
Online Online --> Email Online --> W9["Web 9"] Email --> C7["Web-embedded CAWI 7"]
Email --> A8["Attached form (rarely used) 8"]

A hierarchical diagram showing the main modes of administering a survey. The root is
'Survey', which branches into 'Structured interview' and 'Self-completion questionnaire'.
'Structured interview' branches into 'Face-to-face' and 'Telephone'. 'Face-to-face'
branches into 'Paper + pencil 1' and 'CAPI 2'. 'Telephone' branches into 'Paper + pencil
3' and 'CATI 4'. 'Self-completion questionnaire' branches into 'Supervised 5', 'Postal 6',
and 'Online'. 'Online' branches into 'Email' and 'Web 9'. 'Email' branches into 'Web-
embedded CAWI 7' and 'Attached form (rarely used) 8'.

CAPI: Computer-Assisted Personal Interviews – Google Meet/Zoom, etc.

CATI: Computer-Assisted Telephone Interviews

CAWI: Computer-Assisted Web Interviews

Figure 10.3 Main modes of administering a survey

10.3 PROBABILITY SAMPLING

10.3.1 The qualities of a probability sample


Probability sampling is important in most types of survey research because it enables
the researcher to generalize findings derived from a sample to the population from
which it was selected. This does not imply that we should treat population data and
sample data as the same. If the level of skills development in a sample of 450 out of 9
000 employees was measured as the number of training days completed in the previous
12 months, we can calculate the average (or mean) number of training days undertaken
by the sample to estimate the mean of the population of 9 000, with a known margin of
error using some basic statistical measures. These are presented in Tips and skills
‘Generalizing from a random sample to the population’.

If probability sampling principles were used to select a sample, we can be 95% certain
that the population mean will lie between the sample

mean + or – 1.96 multiplied by the Standard Error (SE). This is known as the confidence
interval.

If the mean (average) number of training days in our sample of 450 employees is 6.7
and the SE is 1.3, we can be 95% certain that the population mean will lie between
$4.152 (= 6.7 - (1.96 \times 1.3))$ and $9.248 (= 6.7 + (1.96 \times 1.3))$ days.

If the SE was smaller, the range of possible values of the population mean would be
narrower; if the SE was larger, the range of possible values of the population mean
would be wider.

If a proportionate stratified random sample (see below) is selected, the SE will be


smaller because the variation between strata is essentially eliminated, and the sample
will represent the population more accurately. This demonstrates how stratification can
improve precision since a possible source of sampling error is eliminated.

By contrast, a cluster sample (see section 10.3.5) that is not stratified will exhibit a
larger SE of the mean than a comparable simple random sample. This occurs because
a possible source of variability between employees (i.e. membership of different
departments, which may affect levels of training undertaken) is disregarded.

Tips and skills

Generalizing from a random sample to the population

Using our illustrative example of training and skills development of 9 000 employees, let
us say that the sample mean is 6.7 days of training per employee in the previous 12
months. How confident can we be that the mean of 6.7 training days measured among
450 employees is likely to be found in the population of 9 000 when we have only
sampled 5% of the population?

If we take an infinite number of samples from the population, the sample estimates of
the mean for the variable 'Training days' will vary in relation to the population mean.
This variation will take the form of a bell-shaped curve known as a normal distribution
(see Figure 10.4). The shape of the distribution implies that there is a clustering of
sample means at or around the population mean. As we move to the left or the right of
the population mean the curve tails off, implying fewer and fewer samples generate
means that are considerably different from the actual mean of the population. The
variation of sample means around the population mean is the sampling error and is
measured using a statistic known as the standard error of the mean (SE). This is an
estimate of the amount that a sample mean is likely to differ from the population mean.

Figure 10.4: The distribution of sample means. A normal distribution curve is shown.
The x-axis is labeled 'Value of the mean' and has three points marked: -1.96 SE,
Population mean, and +1.96 SE. The area under the curve between -1.96 SE and +1.96
SE is shaded, representing 95% of the sample means.

Note: 95% of sample means will lie within the shaded area. SE = standard error of the
mean.

Figure 10.4 The distribution of sample means

10.3.2 Simple random sampling

Simple random sampling is the most basic form of probability sample. Each unit of the
population has an equal probability of inclusion.

Let us return to our previous example. Imagine that we are interested in (a) levels of
training, skills development and learning among those employees in a large local
company and (b) variables that relate to the levels of training they have undertaken. Our
population will all be employees in that company, so we will only be able to generalize
our findings to that company. Assume that our research will only relate to the company's
9 000 full-time employees.

Imagine that we have enough money to interview 450 employees at the company. This
means that the probability of inclusion in the sample is:

$$\frac{450}{9,000}$$

i.e. 1 in 20 or 5%

This is known as the sampling fraction and is expressed as:

$$\frac{n}{N},$$

where $n$ is the sample size and $N$ is the population size.
The key steps in devising our simple random sample can be represented as follows:

Define the population. This will be all full-time employees at the company. This is our
$N$ and in this case it is 9 000.

Select or devise a sampling frame. From the company's employee records kept by the
human resources department, we will exclude the part-time employees and sub-
contractors who do not meet our criteria for inclusion.

Decide your sample size ( $n$ ). This will be 450.

List all the employees in the population and give them consecutive numbers from 1 to
$N$ . In our case, this will be 1 to 9 000.

Using a table of random numbers, or a computer program to generate random numbers,


select $n$ (450) different random numbers that lie from 1 to $N$ (9 000).

The employees with the $n$ (450) random numbers form the sample.

Note two points about this process. Firstly, it is mechanical and there is almost no space
for human bias as a result of subjective criteria. Secondly, the process is not dependent
on the employees' availability. The selection process is done without their knowledge.

Step 5 mentions the possible use of a table of random numbers. Such tables can be
found in many books on statistics. The tables are made up of columns of five-digit
numbers. Since these are five-digit numbers and the maximum number that we can
sample from is 9 000, which is a four-digit number, we now take just the last four digits
in each number.

An alternative way to generate random numbers is to use one of the many random
number generators that can be found online or a program like Microsoft Excel, SPSS, or
other statistical packages. Select $n$ random numbers (in our case 450) that lie
between 1 and $N$ (in our case 9 000). Ignore any random number that occurs more
than once.

10.3.3 Systematic sampling

Systematic sampling is a variation on simple random sampling where you start with a
random number and then select every $n$ th record thereafter directly from the
sampling frame. This approach makes it unnecessary to assign numbers to employees'
names and then to look up their names using random numbers. It is important to ensure
that there is no inherent ordering of the sampling frame, since this may bias the
resulting sample.

As we want to select 1 employee in 20, we would make a random start between 1 and
20 inclusive, possibly by using the first two or last two digits in a table of random
numbers. If the 16th employee on our sampling frame is the first in our sample, we then
take every 20th employee on the list. So, we would select the following sequence from
the list: 16, 36, 56, 76, 96, 116, etc.

10.3.4 Stratified random sampling

In our study of employees, we may want our sample to contain a proportional


representation of the different departments of the company. We may be interested in
the department in which an employee works because this relates to a wide range of
attitudinal features that are relevant to our study of skills development and training.
Generating a simple random sample or a systematic sample might yield such a
representation, but some departments may be under- or over-represented. As a result,
there may be too few respondents from some departments to draw conclusions at the
departmental level.

The sampling frame of company records will show the department in which employees
work, so it will be possible to stratify the population by this criterion and select either a
simple random sample or a systematic sample from each of the resulting strata
(departments). If there are five departments, we would have five strata, with the
numbers in each stratum being one-twentieth of the total for each department, as in
Table 10.2. This table shows a hypothetical outcome of using a simple random sample,
which results in a distribution of employees across departments that does not mirror the
population all that well.

A proportionate stratified sample ensures that the resulting sample will be distributed in
the same way as the population in terms of the stratifying criterion. Note that you can
only conduct stratified sampling when it is relatively easy to identify and allocate units to
strata, based on available, relevant information (for example, Departments in Table
10.2).

In some cases, the sizes of the strata differ significantly. It might be desirable to have a
minimum number of respondents in each stratum, so that you can draw conclusions
about each department. As a result, a larger proportion may be selected from the
smaller strata (for example, human resource management), resulting in a
disproportionate stratified sample as shown in the right-hand column of Table 10.2.

You can use two or more criteria to stratify the sample. You may want to stratify by
department,

Table 10.2 Applying stratified sampling

Department

Population

Possible simple random or systematic sample


Proportionate stratified sample (1 in 20 = 5%)

Disproportionate stratified sample

Sales and marketing

800

55

40

80

Finance and accounts

700

30

35

70

Human resource management

500

18

25

50

Technical, research, and new-product development

800

42

40

80

Production
6 200

305

310

170

TOTAL

9 000

450

450

450

gender and salary levels. Note that stratified sampling may be time-consuming if there
is no available listing in terms of strata, as a great deal of work may be required to
identify members of the population for stratification purposes.

10.3.5 Multi-stage cluster sampling

If a researcher wants to conduct a national sample of employees dispersed throughout


the country, interviewers would then have to travel the length and breadth of the
country. This would be extremely costly and time-consuming. Cluster sampling is one
way of dealing with this potential problem. The primary sampling unit is not the units of
the population to be sampled but groups or clusters of those units.

Imagine that Ayanda needs a nationally representative sample of 5 000 employees who
are working for the 100 largest companies listed on the Johannesburg Stock Exchange
(JSE). One solution might be to select a sample of companies and employees from
each of the selected companies. Using probability sampling at each stage, Ayanda
might randomly sample ten of the population of 100 largest companies, thus yielding ten
clusters, and would then interview or survey 500 randomly selected employees at each
of the ten companies.

This cluster sampling method does not guarantee that these ten companies reflect the
diverse range of

business activities that the population as a whole are engaged in. One solution to
Ayanda's problem would be to group the 100 largest companies by Standard Industrial
Classification (SIC) codes and then randomly sample companies from each of the major
SIC groups (see Tips and skills 'A guide to industry classification systems').
Tips and skills

Icon of a gear with a lightbulb inside, representing a tip or skill.

A guide to industry classification systems

Researchers use industry classification systems to divide firms and other organizations
into groups based on the type of business they are in or the kind of products they make.
The Standard Industrial Classification of All Economic Activities (SIC) consists of a
classification structure of economic activities. SIC edition 7, used in South Africa since
2012, is derived from an international Standard Industrial Classification. It provides a
standard, comprehensive framework within which economic data can be collected and
reported for purposes such as economic analyses, research and decision-making. More
information is available from the national statistics website: [Link]
page\_id=377

Ayanda might decide to exclude the less common SIC groups, and might then sample
one company from each of the twelve remaining major SIC code

categories. As a result, approximately 400 employees from each of the twelve


companies would be interviewed. Thus, there are three separate stages:

Group the 100 largest companies listed on the JSE in SIC categories.

Sample one company from each of the major SIC categories.

Sample 400 employees from each of the companies in the sample.

In a sense, cluster sampling is always a multi-stage approach, because a researcher


always samples clusters first and then samples something else within each cluster –
either further clusters or population units. Research in focus 10.1 provides an example
of a multi-stage cluster sample. It involved three stages: (1) the sampling of provinces
within South Africa, (2) the sampling of major workplaces within provinces and (3) the
sampling of individuals within firms. The advantage of multi-stage cluster sampling is
that it allows interviewers to be far more concentrated than would be the case if a
simple random or stratified sample was selected. Researchers can capitalize upon the
advantages of stratification because they can then stratify the clusters in terms of strata.

10.3.6 Sampling for structured observation

Structured observation (see Chapter 13) necessitates decisions about sampling. It is


more usual not only to sample a larger number of people, but also to incorporate
several other sampling issues.
When people are being sampled for a structured observation study, random sampling is
the ideal. Structured observation can be based on a stratified sampling method, such as
in Jenkins et al's (1975) study of job characteristics (see Research in focus 10.2), or
using non-probability sampling (see Research in focus 10.7).

It is often necessary to ensure that, if certain individuals are sampled on more than one
occasion, they are not always observed at the same time of the day. This means that, if
particular individuals are selected randomly for observation on several different
occasions for short periods, it is desirable for the observation periods also to be
selected randomly. For example, it would not be desirable for a certain manager
working in his office always to be observed at the end of the day. He or she might be
tired, and this would give a false impression of that manager's behaviour.

Research in focus 10.1

An example of a multi-stage cluster sample

Hirschsohn (2011) was interested in changes in union democracy and how it functions
at the workplace. The study was based on information from nationwide surveys of
members of trade unions affiliated to South Africa's pre-eminent labour federation, the
Congress of South African Trade Unions (COSATU).

He conducted surveys of approximately 630 participants every five years before the
country's national elections held in 1994, 1999, 2004 and 2009. The study compared
the attitudes of ordinary trade union members with the attitudes of their shop-stewards,
who are their elected representatives.

The study adopted a multi-stage cluster sampling strategy, which focused on major
urban centres in five provinces. In the first iteration of the survey in 1994, Hirschsohn
identified and then randomly selected large,

unionized workplaces in the selected provinces to ensure that the study covered the
major COSATU unions operating in each province.

Every five years, research teams returned to the same sites to conduct surveys using
face-to-face interviews. In a few cases, the workplaces had to be replaced by
businesses in the same sector to ensure consistent coverage with previous surveys.
After employers granted access, a random sample of ten union members and two shop-
stewards were interviewed face-to-face at work during working hours.

To conduct part of the 2008/9 survey, Hirschsohn travelled to workplaces around the
country during the university vacation period. In each location, he met students from the
University of the Western Cape who resided in the area and they jointly conducted the
face-to-face interviews with the trade union members.
Magnifying glass icon

Research in focus 10.2

Observing jobs

Jenkins et al. (1975) conducted an exploratory study to measure the nature of a wide
variety of job types in three different organizations. An observation schedule was
devised to assess the nature of twenty dimensions of these jobs. Most of the
dimensions were measured through more than one indicator, each of which took the
form of a question that observers had to answer on a six- or seven-point scale. These
were then aggregated for each dimension. Many of the twenty dimensions relate to
issues that have been raised in the sociology of work by labour process theorists and
others (for example, Braverman 1974).

One dimension relates to 'Worker pace control' and comprises three observational
indicators, such as 'How much control does the employee have in setting the pace of his
or her work?' Another dimension was 'Autonomy', which comprised four indicators, such
as whether or not the job allows the individual to make a lot of decisions on his or her
own.

The procedure followed by the observers, most of whom were university students, was
as follows: 'Each respondent was observed twice for an hour. The observations were
scheduled so that the two different observations were separated by at least 2 days,
were usually made at different times of the day, and were always made by two different
observers' (Jenkins et al. 1975).

Various reasons may limit the application of probability sampling in observation studies.
For example, we cannot construct a sampling frame of people walking along a street or
of interactions and other types of meetings between a manager and subordinates.

Concerns about probability sampling largely relate to the external validity of findings.
However, if a researcher conducts a structured observation study over a relatively short
time-span, issues of representivity are also likely to arise. The researcher also needs to
consider what time of

the day, week, month or year the observations are made. Consequently, the researcher
has to consider the timing of observation. Furthermore, how should the researcher
select observation sites, so that they are representative? It is likely that, when access is
secured, the organizations that are studied may not be representative of the relevant
population.

10.3.7 Sampling for online surveys


Online surveys have become increasingly popular for business researchers in recent
decades, as the majority of employees in many organizations work online and a large
proportion of the population are now familiar with using email and the Internet. In
addition, suitable sampling frames of email addresses may be available or can relatively
easily be compiled. (Chapter 11 discusses them further.) In such circumstances,
researchers can conduct surveys using essentially the same probability sampling
procedures as those outlined above. However, email-based surveys of organizational
members may have similar sampling problems to offline surveys, including the
possibility of high levels of non-response.

Surveys of members of commercially relevant online groups, such as those using social
networking technology for business purposes, can also be conducted using probability
sampling principles. For example, C. B. Smith (1997) conducted a survey of people or
organizations involved in creating and maintaining Internet content, using a directory of
providers as her sampling frame.

Challenges with online surveys

How does the above material on sampling apply to online surveys? Certain features of
online communications can make sampling in online surveys problematic:

Not everyone residing in an area or a country is online.

Of those who are online, some may not have the technical ability to handle online
questionnaires. There is clear evidence that there are differences

in Internet access based on income levels, socio-economic status, personal


characteristics and attitudes.

Many people have more than one email address, which may change more often than
physical or postal addresses.

Many people use more than one Internet service provider (ISP).

Several users in the same household may share one computer.

Internet users are a biased sample of the population, in that they tend to be better
educated, wealthier, younger and more often male than female (Blasius & Brandt 2010).
This is a particular problem in South Africa where Internet access is still influenced by
the digital divide between rural and urban areas as well as socio-economic status.

Few sampling frames exist of the general online population, and most of these are likely
to be expensive to acquire, since they are controlled by ISPs or may be confidential.

Such issues seem to limit the use of online surveys using probability sampling
principles. Despite this, business researchers may have more online opportunities than
researchers in other areas. For certain kinds of business research, such as
investigations involving surveys of organizational members, email-based surveys may
present sampling problems that are similar to offline modes of administration, other than
the possibility of higher levels of non-response.

Similarly, business researchers can conduct surveys of members of commercially


relevant online groups using probability sampling principles. If you are interested in
collecting data from members of online communities, such as those using social
networking technology for business purposes, then they would need to be contacted
online to generate a sample.

As Couper (2000: 485) noted about surveys of populations using probability sampling
procedures:

Intra-organizational surveys and those directed at users of the Internet were among the
first to adopt this new survey technology. These restricted populations typically have no
coverage problems ... or very high rates of coverage. Student surveys are a particular
example of this approach that are growing in popularity.

Hewson and Laurent (2008) suggest that when there is no sampling frame, which is
typically the case with samples to be drawn from the general population, the main
approach taken to generating an appropriate sample is to post an invitation:

to answer a questionnaire on relevant newsgroup message boards

to suitable mailing lists, or

on web pages and social media.

The result will be a sample that does not represent a defined population. It is also
impossible to know what the response rate to the questionnaire is, since the size of the
population is also unknown.

However, given that we have limited knowledge and understanding of online behaviour
and attitudes relating to some online issues, we can argue that some information about
these areas is a lot better than none at all, provided that the researcher understands the
limitations of generalizing their findings.

More recently, researchers increasingly use online mechanisms to recruit participants.


These include systems for recruiting paid participants, for example on Google AdWords,
as well as for seeking volunteers, for example, on Facebook. It is important to note that
such mechanisms generate convenience samples, rather than random samples.

There is little research on their effectiveness in generating representative samples, but


one recent study reports that volunteer-seeking methods appear to deliver more
heterogeneous, and thus perhaps more representative, samples than systems that
recruit paid participants (Antoun et al. 2016)

A further issue in relation to online sampling and sampling-related error is the matter of
non-response (see Key concept 10.4). There is growing international evidence that
online surveys typically generate lower response rates than postal questionnaire
surveys (Tse 1998; Sheehan 2001; Pedersen & Nielsen 2016). This has limited
relevance in a developing country like South Africa where there is minimal postal
delivery in many high-density urban areas.

However, as previously noted, with many online surveys, it is impossible to calculate a


response rate. When participants are recruited by invitations and postings on discussion
boards, for example, it is almost impossible to determine the size of the population from
which the sample is drawn.

Response rates can be boosted by the following strategies:

Basic 'netiquette' implies that the researcher contacts prospective respondents before
they are sent a questionnaire. This could enhance the response rates to a web-based
panel survey. Researchers have found pre-notifications sent by text (SMS) message to
be more effective than when sent by email but that a combination of both can be
effective than text messages alone (Bosnjak et al. 2008).

As with postal questionnaire surveys, follow up non-respondents at least once.

There is evidence that incentives can increase response rates in online surveys
(Pedersen & Nielsen 2016).

10.3.8 Probability sampling in qualitative research

Probability sampling may be used in qualitative research, although it is more likely to


occur in interview-based rather than in ethnographic qualitative studies (see Chapter
13). Instead, ethnographers should seek to ensure that they gain access to as wide a
range of individuals, different perspectives and ranges of activity relevant to the
research question as possible.

There is no obvious guide to help qualitative researchers decide when it might be


appropriate to employ probability sampling, but two criteria may be applied.

Firstly, if it is highly significant or important for the qualitative researcher to be able to


generalize to a wider population, probability sampling is likely to be a more compelling
sampling approach. This might occur when the audience for the research views
research that is generalizable as more valuable.
Secondly, if the research questions do not suggest that particular categories of people
(or whatever the unit of analysis is) should be sampled, there may be a case for
sampling randomly.

In many cases, probability sampling is not feasible, because of the constraints of


ongoing fieldwork and also because it can be difficult and often impossible to create a
sampling frame to map

the population from which a random sample might be taken. One way to create a
sampling frame is to define the population narrowly. For example, a researcher may
confine their study to operators of Airbnb accommodation in Khayelitsha, Cape Town.

10.3.9 Limits to generalization

Even when a sample has been selected using probability sampling, the researcher can
only generalize the findings to the population from which that sample was taken. In our
imaginary study above of employee training and skills development (see section
10.3.1), the findings could only be generalized to that company. We should also be
cautious of overgeneralizing in terms of locality and time. A frequent criticism in of
research on employee motivation relates to the extent to which it can be assumed to be
generalizable beyond the confines of the national culture on which the study is based.

For example, a famous book The motivation to work by Herzberg, Mausner and
Snyderman (1959) is based on semi-structured interviews with 203 accountants and
engineers in the Pittsburgh area in the USA. Although respondents were chosen
randomly according to certain criteria for stratification, most were employed in heavy
industry. The authors acknowledge that this 'will inevitably raise questions about the
degree to which the findings are applicable in other areas of the country' (1959: 31).

We can reasonably assume that there is likely to be a male bias to the study, given that
this was a study of accountants and engineers in the late 1950s. The findings may also
reflect the values of high individualism, self-interest and high masculinity that were
identified as characteristic of American culture (Hofstede 1984). This is part of the
reason there have been so many attempts to replicate the study on other occupational
groups and in other localities, including different cultures and countries.

Another issue is whether or not there is a time limit on the findings that were generated
more than 60 years ago.

10.3.10 Sampling errors

To appreciate the significance of sampling error for achieving a representative sample,


look at Figures 10.5–10.8. Imagine there is a population of 200 employees and we want
a sample of 50. Imagine also that one of the variables we are interested in is whether or
not employees receive regular performance appraisals from their immediate supervisor
and whether the population is divided equally between those who do and those who do
not.

This split is represented by the vertical line that divides the population into two halves. If
the sample is precisely representative of the population, the sample of 50 would be
equally split in terms of this variable (see Figure 10.5). If there is a small sampling error,
it will look like Figure 10.6.

A grid of 200 dots arranged in 10 rows and 20 columns. A vertical line down the middle
divides the dots into two equal groups of 100. A rectangular box highlights a sample of
50 dots, consisting of 25 dots from the left group and 25 dots from the right group.

Figure 10.5: A sample with no sampling error. A grid of 200 dots representing a
population, split vertically into two halves. A 10x5 rectangular box highlights a sample of
50 dots, with exactly 25 dots in each half of the population.

Figure 10.5 A sample with no sampling error

A grid of 200 dots, identical to Figure 10.5. The rectangular sample box is shifted one
dot to the left, now containing 26 dots from the left half and 24 dots from the right half.

Figure 10.6: A sample with very little sampling error. Similar to Figure 10.5, but the 50-
dot sample box is slightly shifted, resulting in 26 dots from the left half and 24 dots from
the right half.

Figure 10.6 A sample with very little sampling error

In Figure 10.7, there is a more serious degree of over-representation of employees who


do not receive appraisals, while Figure 10.8 represents a sample where there is a very
serious over-representation of employees who do not receive performance appraisals.

While probability sampling cannot eliminate sampling error, it stands a better chance
than non-probability sampling of keeping sampling error in check. Moreover, probability
sampling allows the researcher to use tests of statistical significance to make inferences
about the population from which the sample was selected. These will be addressed in
Chapter 16.

A grid of 200 dots. The rectangular sample box is shifted further left, now containing 22
dots from the left half and 28 dots from the right half.
Figure 10.7: A sample with some sampling error. A grid of 200 dots with a 50-dot
sample box that is shifted further to the left, containing 22 dots from the left half and 28
dots from the right half.

Figure 10.7 A sample with some sampling error

A grid of 200 dots. The rectangular sample box is shifted even further left, now
containing 18 dots from the left half and 32 dots from the right half.

Figure 10.8: A sample with a lot of sampling error. A grid of 200 dots with a 50-dot
sample box shifted even further left, containing only 18 dots from the left half and 32
dots from the right half.

Figure 10.8 A sample with a lot of sampling error

Errors in survey research

'Errors' in survey research are classified into four categories, two of which are
associated with sampling:

Sampling error (see Key concept 10.1)

Sampling-related errors fall under the category non-sampling error and include an
inaccurate sampling frame and non-response. These arise from activities or events that
are related to the sampling process and are connected with the issue of generalizability
or external validity of findings.

Data collection errors are connected with the implementation of the research process,
and include poor question wording, poor interviewing techniques and flaws in the
administration of research instruments.

Data processing errors arise from data management and errors in coding answers.

We will address the kinds of steps that need to be taken to keep these sources of error
to a minimum in the context of social survey research in Chapters 11 and 12.

10.4 NON-PROBABILITY SAMPLING

Non-probability sampling covers many types of sampling strategy that are not
conducted using the norms of probability sampling, even though some practitioners
regard quota sampling – one type of non-probability sampling – as almost as good as a
probability sample. The practice of surveying one individual per organization, often a
human resources or senior manager, in order to find out about the organization is also
discussed below (see Tips and skills 'Using a single respondent to represent an
organization').

10.4.1 Quota sampling

Quota sampling is used intensively in commercial research, such as market research


and political opinion polling. The aim of quota sampling is

to produce a sample that reflects a population in terms of the relative proportions of


people both in different categories such as gender, ethnicity, age groups, socio-
economic groups and region of residence, and in combinations of these categories.

Unlike a stratified random sample, the final selection of individuals is not at random, but
is left up to the interviewer. Information about the stratification of the South African
population or about certain regions can be obtained from sources like the national
census.

Once the researcher decides on the categories and the number of people to be
interviewed within each category (known as quotas), then the interviewers will typically
be interrelated. In a manner similar to stratified random sampling, the population may be
divided into strata in terms of criteria like gender, social class, age and ethnicity. The
interviewers may use census data to identify the number of people who should be in
each subgroup. The numbers to be interviewed in each subgroup will reflect the
population.

Each interviewer will probably seek out individuals who fit several subgroup quotas. An
interviewer may have been asked to find and interview five Indian, 25 to 34-year-old,
middle-class females among the various subgroups of people, which she has been
assigned. Once a subgroup quota (or a combination of subgroup quotas) has been
achieved, the interviewer will no longer be concerned to locate individuals for that
subgroup.

If you have been approached on the street by a person toting a clipboard and interview
schedule and have then been asked about your age, occupation, and so on, before
being asked a series of questions about a product, you have probably encountered an
interviewer with a quota sample to fill. Sometimes, he or she will decide not to interview
you because you do not meet the criteria required to fill a quota or sub-quota.

A number of criticisms are frequently levelled at quota samples:

Because the choice of respondent is left to the interviewer, the proponents of probability
sampling argue that a quota sample cannot

be representative. When choosing people to approach, interviewers may be influenced


by their perceptions of how friendly people are or by whether they make eye contact
with the interviewer.
People who are available at the time when the interviews are conducted and are in an
interviewer's vicinity may not be typical, for example, because they are not at work
during the day.

The interviewer is also likely to make incorrect judgements about certain characteristics
in deciding whether or not to approach a person; for example, that the person is
younger than he or she looks. This may introduce some bias into the sample.

When using a quota sample, the researcher cannot calculate various population values,
such as a standard error of the mean, because the method of selection is non-random.

Despite these challenges, quota sampling has some arguments in its favour:

It is useful when the researcher has the demographic profile of the population but does
not have access to a database from which to draw a representative sample.

It is cheaper and quicker than using a comparable probability sample for interview
surveys. For example, interviewers do not have to spend a lot of time travelling between
interviews.

It is easier to manage as it is not necessary (and indeed it is not always possible) to


keep track of people who need to be re-contacted or to keep track of refusals.

Interviewers do not have to keep calling people back if they were not available when
they were first approached.

When speed is important, a quota sample is invaluable when compared to the more
cumbersome probability sample.

It is useful for conducting development work on new measures or on research


instruments. It can also be used for exploratory work to generate new theoretical ideas.

There is some evidence that suggests that quota samples often result in biases, when
compared

to random samples (Yang & Banamah 2014). They under-represent people in lower
social strata, people who work in the private sector and manufacturing and people at the
extremes of income, and they over-represent women in households with children and
people from larger households (Marsh & Scarbrough 1990; Butcher 1994).

10.4.2 Convenience sampling

A convenience sample is one that is available to the researcher because of


accessibility. A researcher who teaches at a university business school and is interested
in how managers deal with ethical issues when making business decisions might
administer a questionnaire to several classes of part-time MBA students, who are all
managers. The chances are that there will be a good response rate.

While the findings may prove interesting, it is impossible to generalize the findings,
because we do not know what population this sample represents – the fact that they are
taking the MBA degree programme distinguishes them from managers in general.
Research in focus 10.3 provides an example of the use of such a convenience sample.

This is not to suggest that convenience samples should never be used. A kind of
context in which it may be at least fairly acceptable to use a convenience sample is
when an opportunity too good to miss presents itself to gather data from a convenience
sample. The data will not allow definitive findings to be generated, because of the
problem of generalization, but they could provide a springboard for further research or
allow links to be forged with existing findings in an area.

Convenience sampling plays a relatively prominent role in the field of business and
management, and has become the norm in fields such as consumer behaviour.
Nonetheless, there is evidence that we should be cautious in generalizing from
convenience samples, particularly when they are samples of undergraduate students
(Peterson & Merunka 2014).

Research in focus 10.3

Convenience sampling in a study of discrimination in hiring Deroos et al. (2017) studied


how ethnic cues influenced the outcomes of curriculum vitae (CV) screening during
recruitment processes. While most previous studies of CV screening used samples of
university students (2017: 862), they focused on a sample of human resource (HR)
professionals. The aim was to increase the likelihood that the conclusions drawn were
representative of processes and outcomes involving respondents who undertake
recruitment in their daily working lives.

The authors needed a large sample for their study in Belgium, but there was no
sampling frame available from which to draw a random sample of respondents, so the
sample was drawn from membership lists of professional HR associations, business
publications and the researchers' own networks.

From a sample of 1 463 respondents, 424 agreed to participate. Participants were all
Caucasian. Clearly this sampling strategy did not generate a random sample and the
sample cannot be regarded as broadly representative of the profession statistically.
However, it was a practical way to draw a large sample of actual HR professionals that
could not otherwise have been generated.

The researchers provided each respondent with job advertisements for two different
jobs in two different kinds of businesses. They each also provided with four fictional
CVs, including a photo of a fictional applicant. The CVs showed a variety of
combinations of skin tone (light versus dark) and name (Flemish versus Arab/Maghreb).
Respondents were asked to rate each applicant using a three-item measure with Likert
scale responses. For example:

'Given all the information you read about this applicant, how likely is it that you would
invite this applicant for a job interview?' (1 = not likely at all; 7 = very likely).

The findings showed the equally qualified candidates with a dark skin tone were
systematically rated less suitable for jobs than those with light skin tone, with the effect
varying depending on the nature of the job. The name of the candidate did not appear to
matter. The findings suggested that there was systematic ethnically based
discrimination in CV screening and that there were subtle effects that arose from the
kinds of jobs involved.

10.4.3 Purposive sampling

In much the same way that the discussion of sampling in quantitative research revolves
around probability sampling, discussions of sampling in qualitative research tend to
revolve around the notion of purposive sampling (see Key concept 10.2). This type of
sampling deals essentially with the strategic selection of units where the research will
be conducted (which may be people, organizations, documents, departments, and so
on) with direct reference to the research questions being asked. The research questions
should give an indication of what categories of people (or another unit of analysis) are
the focus of attention and therefore need to be sampled.

As purposive sampling is a non-probability form of sampling, the researcher cannot


generalize the findings

Key concept 10.2

What is purposive sampling?

Purposive sampling is a non-probability form of sampling so the researcher cannot


generalize the results to a population. The researcher does not seek to sample research
participants on a random basis.

The goal of purposive sampling is to sample cases/participants in a strategic way, so


that those sampled are relevant to the research questions. Usually, the researcher
wants to ensure that there is a good deal of variety, so that members of the sample
differ from each other in terms of key characteristics.

to a population and does not aim to sample participants at random. Usually, the
researcher wants to ensure that there is sufficient variety, so that sample members
differ from each other in terms of key characteristics.

The researcher needs to be clear about which criteria will be relevant to including or
excluding cases. Examples of purposive sampling in qualitative research are theoretical
sampling and snowball sampling (see Research in focus 10.4 and 10.5). In quantitative
research, quota sampling is regarded as a form of purposive sampling.

Teddlie and Yu (2007) draw a useful distinction between:

sequential sampling*, which implies an evolving process in that the researcher usually
begins with a sample and step-by-step adds units to the sample as it suits the research
questions, and

non-sequential sampling* approaches or 'fixed sampling strategies' where the sample is

determined when the research starts, guided by the research questions, and

is more or less fixed early on in the research process.

Types of purposive sampling

Purposive sampling is the master concept around which different qualitative sampling
approaches can be distinguished. The following types of purposive sampling have been
identified by researchers such as Patton (1990) and Palys (2008). Three types are
particularly useful to select cases or contexts:

Typical case sampling* – sampling of one or more cases that exemplify a concept or
relationship of interest. For example, a researcher may plan to study how Takealot, the
South African ecommerce business, re-organised its supply chain in response to the
rapid increase in demand for groceries during the Covid-19 crisis.

Extreme or deviant case sampling* – sampling cases that are unusual or that are
unusually at the far end(s) of a particular dimension of interest.

Critical case sampling* – sampling a crucial case that permits a logical inference about
the matter of interest, such as choosing a case so that you can test a theory.

The following types of sampling can be used to sample individuals, cases or contexts:

Theoretical sampling* – sampling by which the researcher 'collects, codes, and


analyses his data and decides what data to collect next and where to find them, in order
to develop his theory as it emerges' (Glaser & Strauss 1967: 45)

Snowball sampling* – initial sampling of a small group of relevant people, who


recommend or refer the researcher to other participants or to connections who have had
the experience or characteristics relevant to the research. These participants will then
suggest others and so on.

Maximum variation sampling* – sampling to ensure the widest possible variation in


terms of the dimension of interest
Criterion sampling* – only sampling units (cases or individuals) that meet particular
criteria

Stratified purposive sampling* – sampling typical cases or individuals within subgroups


of interest

Opportunistic sampling* – capitalizing on opportunities to collect data from certain


individuals, contact with whom is largely unforeseen but who may provide data relevant
to the research question.

Using more than one purposive sampling approach

Purposive sampling often involves combining approaches outlined above. For example,
it is quite common for snowball sampling to be preceded by another form of purposive
sampling. This process can entail sampling initial participants without using a snowball
approach and then using these initial contacts to broaden out through a snowballing
method. More than one sampling approach may also be employed when researchers try
to introduce a clear purposive intent into a snowball sample.

As an example of using two of the purposive approaches mentioned above, Marzano


and Scott (2009) studied the role of power in the branding of Australia's Gold Coast as a
tourist destination. They began by purposively sampling key individuals who worked in
advertising agencies that were responsible for, and interested in, the branding of the
destination. Using the snowballing process, they were able to identify and conduct semi-
structured interviews with senior managers in hotels and theme parks.

10.4.4 Theoretical sampling and saturation

Theoretical sampling is an integral part of the grounded theory approach to qualitative


data analysis. In grounded theory, the researcher continues to collect data (by
observing, interviewing, collecting documents, and so on) until they achieve theoretical
saturation (see Key concept 10.3). Once researchers believe that theoretical saturation
has been achieved, they should decide to (a) stop collecting new data on that particular
theoretical idea and (b) move on to investigating some ramifications of the emerging
theory.

Theoretical sampling is a type of purposive sampling that is used to develop a theory


based on collecting and analyzing data (Glaser & Strauss 1967). The researcher's
approach to data gathering is 'driven by concepts derived from the evolving theory and
based on the concept of "making comparisons", whose purpose is to go to places,
people, or events that will maximize opportunities to discover variations among
concepts and to densify categories in terms of their properties and dimensions.'
(Strauss & Corbin 1998: 201).

Key concept 10.3


Icon for Key concept 10.3, showing a brain and a gear.

What is theoretical saturation?

The key idea is that the researcher carries on sampling theoretically until a category has
been saturated with data, that is, until:

no new or relevant data seem to be emerging regarding a category

the properties and dimensions of the category are well enough developed to
demonstrate variation

'the relationships among categories are well established and validated' (Strauss &
Corbin 1998: 212).

In grounded theory, a category is more abstract than a concept, as it may group


together several concepts that have common defining features. Saturation means that
the researcher recognizes what people say in interviews but that the new data no longer
suggests new insights into an emergent theory or new dimensions of theoretical
categories.

This emphasizes that theoretical sampling is an iterative, on-going process that involves
several stages, rather than a distinct and single stage of the research process.

Theoretical sampling also differs from purposive sampling outlined earlier, because the
emphasis is on generating theory and developing theoretical categories that emerge
while analysing the data that was collected, rather than the exception being just on
boosting sample size (Corbin 2000: 519). Consequently, successive interviews and
observations form the basis for creating a category and confirming its importance.

Once established, there is no need to collect more data regarding that category or
cluster of categories. The researcher should then shift focus to generating propositions
or hypotheses out of the categories that they are developing and then continue on to
collecting data in relation to these hypotheses.

Figure 10.9 outlines the main steps in theoretical sampling. Remember that, in
ethnographic research,

graph TD A[General research question] --> B[Sample based on theory] B --> C[Data
collection] C --> D[Data analysis (concepts, categories)] D --> E[Theoretical saturation]
E --> F[Hypotheses or propositions] F --> B
Flowchart illustrating the iterative process of theoretical sampling. The steps are:
General research question, Sample based on theory, Data collection, Data analysis
(concepts, categories), Theoretical saturation, and Hypotheses or propositions. An
arrow loops back from Hypotheses or propositions to Sample based on theory.

Figure 10.9 The iterative process of theoretical sampling

(see Chapter 12), it is not just people who are being sampled but events and contexts
as well, as explained below (see section 10.6.2). As Glaser and Strauss put it,
'Theoretical sampling is done in order to discover categories and their properties and to
suggest the interrelationships into a theory' (1967: 62).

Research in focus 10.4

An example of theoretical sampling: Transformation and black empowerment in South


African National Parks

Maguranyanga (2009) explored the institutional and political transformation of South


African National Parks (SANParks) from 1991 to 2008 through a multi-disciplinary lens.
Based on a single longitudinal case study, he examined SANParks' transformation by
analyzing transformation strategies and initiatives related to deracialization, black
empowerment, social justice and people-oriented conservation. He used multiple
methods to collect the data – archival documents, interviews with key-informants,
observation and SANParks' official organizational climate survey data set.

Maguranyanga used purposive, theoretical sampling of key informants, designed to


confirm the emerging theoretical framework. He sampled interviewees until the defined
categories achieved theoretical saturation, and chose further interviewees based on his
emerging theoretical focus.

During the course of the interviews, he also used snowball sampling by following
additional leads to whom initial contacts had referred him. After theoretical saturation,
no additional interviews were necessary.

10.4.5 Snowball sampling

Snowball sampling is a technique where the researcher initially makes contact with a
small group of people who are relevant to the research topic, and then requests that
they help to establish contact with other potential participants in the study. In certain

respects, snowball sampling is a form of convenience sampling (see section 10.4.2


'Convenience sampling'). An example of snowball sampling is given in the study by
Venter, Boschoff and Maas (2005; see Research in focus 10.5), where researchers
used the technique to identify owner-managers and successors of small and medium-
sized family businesses in South Africa. Another South African example came from
Grobler and De Villiers (2017) who explored how the information needs of women may
be identified more effectively by focusing on domestic workers in Johannesburg and
Pretoria.

Research in focus 10.5

A snowball sample

Venter et al. (2005) were interested in factors that influence the succession process in
small and medium-sized family businesses. Their initial intention was to obtain access
to a mailing list of small and medium-sized family businesses in South Africa from banks
and other large organizations that had family businesses as clients. This would then
have formed the basis of their sample.

However, these large organizations declined to share their client information with the
research team, which instead used snowball sampling. This involved research
associates, who were employed in different regions of the country, contacting small and
medium-sized businesses with the aim of identifying those that were family businesses.
Potential respondents were then asked to refer the researchers on to other family
businesses that they knew about.

As the researchers explain, 'following up on referrals proved to be the most effective


approach and eventually yielded the majority of the potential respondents listed on the
sampling frame' (Venter et al. 2005: 291).

A questionnaire survey was then mailed to 2 458 respondents, comprising current


owner-managers, potential successors, successors and retiring owner-managers in 1
038 family businesses. A total of 332 usable questionnaires were returned.

Snowball sampling is not an example of random sampling, because the extent of the
population from which the sample would be drawn cannot be determined, as there is no
accessible sampling frame or list. The difficulty of accessing or creating such a sampling
frame means that the snowball sampling approach is often the only feasible one.
Consequently, it is unlikely that the sample will be representative of the population.

Snowball sampling tends to be used in qualitative research, rather than quantitative


studies. Concerns about external validity and the ability to generalize are not as critical
with qualitative research as in quantitative research (see Chapter 3), as qualitative
sampling is more likely to be guided by a preference for theoretical rather than statistical
sampling.

This does not imply that snowball sampling is entirely irrelevant to quantitative research.
When a researcher needs to focus upon or reflect relationships among people, tracing
connections through snowball
sampling may be a better approach than conventional probability sampling (Coleman
1958).

The sampling of informants in ethnographic research (see also section 10.6.2


'Ethnography and participant observation') sometimes combines opportunistic sampling
and snowball sampling. Often, ethnographers are forced to gather information from
whatever sources are available to them. Very often they face opposition or at least
indifference to their research and are relieved to glean information or views from
whoever is prepared to divulge such details.

Jackall (1988) provides an example of opportunistic sampling. See Research in focus


10.6.

Tips and skills

Should a single respondent represent an organization?

It is fairly common practice in business survey research for one respondent, often a
senior manager, to be asked to complete a questionnaire or to be interviewed about
issues that are related to their organization or workplace. This enables a larger number
of organizations to be surveyed with less investment of time and resources than if
multiple respondents were surveyed within each organization.

It may be unwise to rely on a single respondent to know everything about the


organization. If the respondent is a senior manager, he or she may represent
organizational practices in a way that portrays his or her own role and responsibilities
more favourably than other respondents in the organization would.

It is therefore important to acknowledge the potential limitations associated with such a


sampling strategy.

Research in focus 10.6

Research using opportunistic sampling

Jackall (1988) went into several large organizations to study how bureaucracy shapes
moral consciousness. Analysis of the occupational ethics of corporate managers was
based on core data of 143 intensive, semi-structured interviews with managers at every
level of the organization. This formed the basis for selecting a smaller stratified group of
12 managers, who were re-interviewed several times and asked to interpret materials
that Jackall was collecting.

As the study progressed, Jackall realized that any investigation of organizational


morality should also explore managerial dissenters, or 'whistleblowers'. Between 1982
and 1988, he conducted case studies by interviewing 18 'whistleblowers' and reviewing
large amounts of documentary evidence. To explore managerial morality further, he
then presented these cases to the stratified group of 12 managers and asked them 'to
assess the dissenters' actions and motives by their own standards' (1988: 206).

10.4.6 Levels of sampling

Different sampling approaches may be found in qualitative research (see 'Types of


purposive sampling' on page 225). While these are useful, researchers may intermingle
two different levels of sampling, an issue that is particularly relevant to the consideration
of sampling in qualitative research based on a single case study or a multiple case
study design.

With such research designs, the researcher must first select the case or cases as
primary sampling units; subsequently, the researcher must sample secondary units
within the case and, possibly in a third phase, final sampling units. When sampling
contexts or cases, qualitative researchers have a number of principles of purposive
sampling on which to draw. To a significant extent, the ideas and principles behind
these were introduced in Chapter 5 (see section 5.5 'Case study design') in connection
with the different types of case.

10.5 SAMPLE SIZE

The decision about sample size depends on a number of considerations, including time
and cost – there is no one definitive answer and invariably a compromise will have to be
reached between the constraints of time and cost, the need for precision and other
considerations.

10.5.1 Sample size and probability sampling

In sampling, size really does matter. The bigger the sample, the more representative it
is likely to be. However, students do their research with limited resources, so you should
find out if your department has guidelines about minimum sample sizes. If there are no
such guidelines, you will need to conduct your survey so as to maximize the number of
interviews you can manage or the number of postal or electronic questionnaires you can
send out. In most cases, it will not be feasible for you to create a truly random sample.

The crucial point is to be clear about and to justify what you have done. Explain the
difficulties that you would have encountered in generating a random sample. Explain
why you really could not enlarge your sample. Do not make claims about your sample
that are not sustainable. There should still be lots of good features about your sample:
the range of people included, the response rate and the level of cooperation you
received from participants.

Absolute and relative sample size

Contrary to what you might expect, the absolute size of a sample is important, not its
relative size:
a national probability sample of 1 000 individuals in South Africa has as much validity as
a national probability sample of 1 000 individuals in the USA, even though the latter has
a much larger population. Increasing the size of a sample cannot guarantee increased
precision, but it does increase the likely precision of the sample and reduces sampling
error.

Therefore, an important component of any decision about sample size should be how
much sampling error you are prepared to tolerate. The less sampling error you are
prepared to tolerate, the larger the sample will need to be.

Heterogeneity of the population

Another consideration is the homogeneity and heterogeneity of the population from


which the sample is to be taken. When a sample is very diverse, such as a sample of a
whole country or city, the population is likely to be highly varied. When it is relatively
homogeneous, such as shareholders of a company or members of an occupation, the
amount of variation is less.

Therefore, the greater the heterogeneity of a population, the larger a sample will need to
be.

Time and cost

By and large, up to a sample size of around 1 000, the gains in precision are noticeable
as the sample size climbs from low figures of 50, 100, 150, and so on, upwards. After a
certain point, often in the region of 1 000, there is a slowing-down in the extent to which
precision increases (and hence the extent to which the sample error of the mean
declines). Striving for smaller and smaller increments of precision then becomes a
question of time and cost.

Non-response

Most sample surveys attract a certain amount of non-response, which could cause
sampling error. Thus, if we aim to ensure that we survey 450 people and expect a non-
response rate of 20%, it may be advisable to sample 550 to 560 individuals, on the
grounds that approximately 90 will be non-respondents (see Chapter 11 for a discussion
of acceptable response rates and strategies to improve survey response

rates). Response rates (see Key concept 10.4) for mail surveys typically range between
15% and 30%, so you may need to send out three to six times as many questionnaires
as you require.

Icon of a person thinking, representing a key concept.


Key concept 10.4

What is a response rate?

When conducting social survey research, some people in the sample refuse to
participate. The response rate is the percentage of a sample that does agree to
participate. However, not everyone who replies will be included: if a large number of
questions are not answered by a respondent, or if there are clear indications that the
respondent did not take the interview or questionnaire seriously, it is better to exclude
the respondent and only count the interviews or questionnaires that can be used as the
numerator. Other people that were originally in the sample may not be suitable, are
inappropriate to the research purpose or cannot be contacted. Thus the response rate is
calculated as follows:

$$\frac{\text{number of usable questionnaires}}{\text{total sample} -


\text{unsuitable/uncontrollable members of the sample}} \times 100$$

10.5.2 Sample size in qualitative research

If theoretical considerations guide the selection of participants in your research, then it


can be difficult to establish at the outset of your research how many people should be
interviewed. This was confirmed by Guest et al. (2006), who conducted a literature
review of guidelines for qualitative research. In their paper titled 'How many interviews
are enough?', they found:

no description of how saturation might be determined and no practical guidelines for


estimating sample sizes for purposively sampled interviews. This dearth led us to carry
out another search through the social and behavioural science literature ... our
suspicions were confirmed; very little headway has been made in this regard (2006: 60).

This implies that it is impossible to know how many people or groups should be
interviewed before you have achieved theoretical saturation (see Key concept 10.3).
Unlike the use of statistical tests in probability sampling, researchers rarely specify their
criteria for saturation in detail.

As a rule of thumb, the broader the scope of the study and the more comparisons
required between the groups in the sample, the more interviews will need to be carried
out (Warren 2002; Morse 2004). The convincing conclusions is likely to vary from
situation to situation in purposive sampling terms. It is also likely that the orientation of
the researchers and the purposes of their research will be significant.

Warren (2002: 99) also noted that for a study of qualitative interviews to be published,
journals expect the minimum to be between 20 and 30 interviews. Although there is an
emphasis on the importance of sampling purposively in qualitative research, this
suggests that minimum levels of acceptability operate. The exceptions to Warren's rule
would include intensive interviews such as life story interviews, where there may be less
than a handful of interviewees.

Qualitative researchers have to recognize that they are engaged in a delicate balancing
act (Onwuegbuzie & Collins 2007: 289):

In general, sample sizes in qualitative research should not be so small as to make it


difficult to achieve data saturation, theoretical saturation, or informational redundancy.
At the same time, the sample should not be so large that it is difficult to undertake a
deep, case-oriented analysis.

Consequently, it is important to fully justify the sample size and be clear about the
sampling method you used, why you used it, and why the sample size you achieved is
appropriate. If you use saturation as the criterion for sample size, there is no need to
specify minimum or maximum sample sizes because the sample size depends on when
saturation was achieved and where no new theoretical insights are gained.

You also need to be sure that you do not generalize from your data inappropriately. If
you conduct a study on SMEs in a particular area and business sector, you need to be
wary of generalizing your findings to other regions or sectors. Essentially,

into groups in terms of stratifying criteria and then selected randomly or through a
snowball sampling method.

Most of the issues raised in connection with sampling in ethnographic research (below)
apply more or less equally to sampling in qualitative interviewing.

10.6.2 Ethnography and participant observation

The sampling of informants in ethnographic research is often a combination of


convenience sampling and snowball sampling. Typically, ethnographers have little
option but to gather information from whatever sources are accessible. They may also
rely on a form of snowball sampling, by asking for names of others who might be
relevant and who could be contacted.

Very often ethnographers are relieved to glean information or views from whoever is
prepared to divulge such details. For example, Dalton (1959) refers to the importance of
'conversational interviewing' as the basis for his data collection strategy. These are not
interviews in the usual sense, but a series of broken and incomplete conversations that,
when written up, may be 'tied together as one statement' (1959: 280).

In other instances, a stratified sampling approach may be used to emphasize how


representative interviewees are of the overall population. During her research on the
nature of work, Casey (1995) interviewed 60 people at the Hephaestus Corporation to
gain a wide sample from various occupational and demographic groupings. Some
individuals were also chosen on the basis of their strategic importance within the team
or division.

Ethnographers also need to consider time and context as units in the context of
sampling (Hammersley & Atkinson 1995). They may need to ensure that they observe
people or events at different times of the day, and at different days of the week. By
observing in different locations, the researcher can also identify how people's behaviour
is influenced by contextual factors.

In his study of masculinity and workplace culture in a British factory, Collinson (1992)
emphasized how shopfloor workers resisted managerial control by spending time
chatting and joking. Collinson

spent time with workers during lunch and unofficial breaks, in the toilet, the canteen, the
car park, on the bus, in the pub and occasionally in people's homes. As a result, he was
able to explore people's cultural practices in far more detail than if he had confined his
study to observing practices within formal workplace settings. An outstanding South
African example is Phakati's (2002) study of a gold mine where he worked underground
and lived with mineworkers to complete his Master's thesis.

10.6.3 Content analysis

Content analysis is a data analysis method (see Chapters 14 and 15) that can be
applied to many kinds of documents. There are several phases in the sample selection
process. The method of sampling content in the mass media is explored briefly here as
the principles can be applied to a broad range of applications.

Many studies of the mass media specify the research problem as 'the representation of
X [any relevant theme] in the mass media'. The X may be trade unions, affirmative
action, or women and leadership.

The researcher needs to decide which mass media source or sources to focus upon –
newspapers, television, radio, magazines, news websites, blogs or a number of
sources? If newspapers, will the focus be on all newspapers of a type or a sample, the
business press, printed newspapers or online media, or both? Will all news items be
analyzed, including feature articles and letters to the editor? And will newspapers from
more than one province be included?

Typically, researchers will opt for one or possibly two of the mass media and may
sample within that type or types (see Research in focus 10.7).

Research in focus 10.7

Selecting advertisements in South African media for content analysis


Holtzhausen (2010) conducted a descriptive, cross-sectional study of the visual images
of women in magazine advertisements and television commercials to determine the
roles in which they are portrayed in marketing communication. She used quantitative
content analysis to focus on how frequently adverts portrayed women in particular roles.

Purposive, non-probability sampling of monthly and weekly magazines included all


general interest, male and female South African magazines with audited readership
figures of over 500 000. Specialist publications were excluded, as their target audiences
are too specialized.

The sample included all full-page and double-page advertisements that featured at least
one woman in the selected magazines over a two-month period. For the weekly
magazines, he only selected the first weekly issue of the month.

Television advertisements featuring women were only selected from free-to-air


channels, to which the majority of South Africans have access. The research only
included commercials during prime time (from 18:00 to 22:00) on Mondays,
Wednesdays and Fridays, in line with previous research practices. Fridays were
included as they are part of the weekend and may feature different commercials than
weekdays. She selected all the television commercials that featured female models in
the time frame.

10.6.4 Sampling dates

Sometimes, deciding when a sample will be conducted depends on the occurrence of a


phenomenon. With a research question that entails an ongoing general phenomenon,
such as the representation of women in advertisements, the matter of dates is more
open. The principles of probability sampling can readily be adapted for sampling dates –
for example, generating a systematic sample of dates by randomly selecting one day of
the week and then selecting every $n$ th day thereafter. Alternatively, Monday
newspapers could provide the first set, followed by Tuesday the following week,
Wednesday the week after, and so on.

KEY POINTS

Icon of a key inside a circle, representing key points.

Probability sampling is a mechanism for reducing bias in the selection of samples.

Ensure that you become familiar with key technical terms in the literature on sampling
such as: representative sample, random sample, non-response, population, sampling
error.
Randomly selected samples from a comprehensive database permit generalizations to
the population because they have certain known qualities.

Sampling error decreases as sample size increases.

Quota samples can provide reasonable alternatives to random samples, but they suffer
from some deficiencies.

Convenience samples may provide interesting data, but it is crucial to be aware of their
limitations in terms of generalizability.

Sampling and sampling-related error are just two sources of error in social survey
research.

Online surveys pose specific sampling challenges.

It is crucial to be clear about your research questions in order to be certain about your
units of analysis and what exactly is to be analyzed.

Sampling considerations for qualitative research usually differ from those for
quantitative research in that issues of representativeness are less emphasized.

Purposive sampling is the fundamental principle in qualitative research.

Purposive sampling places research questions at the forefront of sampling


considerations.

It is important to remember that purposive sampling entails considering the levels at


which sampling needs to take place.

Theoretical saturation is a useful principle for decisions about sample size, but the
evidence is that it is often claimed by researchers but not clearly demonstrated.

Both ethnographic research and investigations using qualitative interviews seldom


employ random sampling to select participants.

QUESTIONS FOR REVIEW

Icon of a question mark inside a circle, representing questions for review.

What does each of the following terms mean: population, probability sampling, non-
probability sampling, sampling frame, representative sample, sampling and non-
sampling error, purposive sampling, theoretical sampling, snowball sampling?
What are the goals of sampling?

What are the main areas of bias in sampling?

Types of probability sample

What is probability sampling and why is it important?

What are the main types of probability sample?

To what extent does a stratified random sample offer greater precision than a simple
random or systematic sample?

If you were planning to conduct an interview survey of around 500 people in Gqeberha
(previously Port Elizabeth), what type of probability sample would you choose and why?

A researcher positions herself on a street corner and asks one passerby in five to be
interviewed; she continues doing this until she has a sample of 250. Is this a
representative sample? Explain.

The qualities of a probability sample

A researcher is interested in levels of job satisfaction among manual workers in a firm


that is undergoing change. The firm has 1 200 manual workers. The researcher selects
a simple random sample of 10% of the population. He measures job satisfaction on a
Likert scale comprising ten items. A high level of satisfaction is scored 5 and a low level
is scored 1. The mean job satisfaction score is 34.3. The standard error of the mean is
8.57. What is the 95% confidence interval?

Sampling for structured observation

Name some sampling strategies in structured observation.

Sampling for online surveys

What are the main challenges in selecting a sample for an online survey?

Limits to generalization

'The problem of generalization to a population is not just to do with the matter of getting
a representative sample.' Discuss.

Sampling error

What is the significance of sampling error for achieving a representative sample?


Error in survey research

'Non-sampling error, as its name implies, is concerned with sources of error that are not
part of the sampling process.' Discuss.

Types of non-probability sampling

Are non-probability samples useless? Explain.

'Quota samples are not true random samples, but in terms of generating a
representative sample there is little difference, which accounts for their use in market
research and opinion polling.' Discuss.

Purposive sampling

How does purposive sampling differ from probability sampling, and why do many
qualitative researchers prefer to use the former?

Why is theoretical sampling such an important facet of grounded theory?

Why is theoretical saturation such an important ingredient of theoretical sampling?

What are the main reasons for considering the use of snowball sampling?

Compare theoretical and snowball sampling.

Levels of sampling

Why might it be significant to distinguish between the different levels at which sampling
can take place in a qualitative research project?

Sample size

What factors determine how large a probability sample should be?

What is non-response and why is it important to the question of whether or not you will
end up with a representative sample?

Not just people

Why might it be important to remember in purposive sampling that it is not just people
who are candidates for consideration in sampling issues?

Sample size
Why do writers seem to disagree so much on what is a minimum acceptable sample
size in qualitative research?

Does theoretical sampling assist qualitative researchers? Explain.

Using more than one sampling approach

How might it be useful to select people purposively following a survey?

Content analysis: Selecting a sample

What special sampling issues does content analysis pose

You might also like