Sampling
By [Link] ullah khan
Definition
Sampling is a process of selecting observations (a sample)
to provide adequate description and interference of the
populations.
PROBABILITY SAMPLING & Non PROBABILITY SAMPLING
Probability sampling refers to the selection of a sample
from a population, when this selection is based on the
principle of randomization, that is, random selection or
chance. Probability sampling is more complex, more
time-consuming and usually more costly than non-
probability sampling.
Probability sampling is the most common form of sampling
for public opinion studies, election polling, and other
studies in which results will be applied to a wider
population.
Non-probability sampling is when a sample is created
through a non-random process. This could include a
researcher sending a survey link to their friends or
Stopping people on the street .
PROBABILITY SAMPLING & Non PROBABILITY
SAMPLING
Non-probability samples are often used during the
exploratory stage of a research project, and in
qualitative research, which is more subjective than
quantitative research, but are also used for research
with specific target populations in mind, such as farmers
that grow maize.
Non-probability sampling involves non-random selection
based on convenience or other criteria, allowing you to
easily collect data.
Types of sampling
PROBABILITY SAMPLING
Four types of probability sampling are presented
1. Simple Random Sampling
2. Systematic Sampling
3. Stratified Sampling
4. Cluster Sampling
NONPROBABILITY SAMPLING
5. Samples of Convenience or convenient sampling
6. Snowball Sampling
7. Purposive Sampling
Simple Random Sampling
Simple random sampling is a procedure in which each
member of the population has an equal chance of being
selected for the sample, and selection of each subject is
independent of selection of other participants.
Assume that we wish to draw a random sample of 300
participants
from the accessible population of 3000 patients
Simple Random Sampling
To literally “draw” the sample, we would
write each patient’s name on a slip of paper, put the
3000 slips of paper in a rotating cage, mix the slips
thoroughly,
and draw out 300 of the slips. This is an example
of sampling without replacement because each slip of
paper is not replaced in the cage after it is drawn.
It is also possible to sample with replacement, in which case
the selected unit is placed back in the population so
that it may be drawn again. In clinical research, it is not
feasible to use the same person more than once for a
sample, so sampling without replacement is the norm.
The preferred method for generating a simple random
sample is to use random numbers that are provided
in a table or generated by a computer.
Before consulting the table, the
researcher numbers the units in the sampling frame. In
our TKA study, the patients would be numbered from
0001 to 3000. Starting in a random place on the table,
and moving in either a horizontal or vertical direction.
we would include in our sample any four-digit numbers
from 0001 to 3000 that we encounter. Any four-digit
numbers greater than 3000 are ignored, as are duplicate
numbers.
The process is continued until the required
number of units is selected. From within the boldface
portion of Table 9-2 in column 7, rows 76 through 80,
the following numbers, which correspond to individual
participants, would be selected for our TKA sample,
1945, 2757, and 2305. Alternatively, if we used an
Internet random number generator, we would enter
the upper and lower limits of the accessible population,
3000 and 1, respectively, and then select the option that
did not give duplicate selections.
Upon then activating the generator, all 3000 numbers
would be listed in random
order. To obtain our 300-number sample, we would
simply select the first 300 numbers from the list of random
ordered numbers.
For example, on the referenced
date that the random number generator was accessed,
a listing of random numbers with a lower and upper
limit of 1 and 3000 was selected. The first 10 numbers
produced on that list were as follows: 250, 1890, 403,
1360, 2541, 1921, 1458, 45, 170, and 1024.
The remaining 290 numbers for the sample would follow
from that
list. On a second effort of generating a random listing
of numbers with the same parameters, the following
first 10 numbers were produced: 1036, 848, 1734, 1898,
2533, 1708, 254, 1631, 2221, and 1281. Thus, a second
effort of obtaining a random order of 3000 numbers
produced an entirely different order than from the first
effort.
Simple random sampling is easy to comprehend, but
it is sometimes difficult to implement. If the population
is large, the process of assigning a number to each
population
unit becomes extremely time consuming when
using a random numbers table.
Alternatively, employing
a random numbers generator on the Internet makes
the simple random sampling a more reasonable process.
The other probability sampling techniques are easier
to implement than simple random sampling and may
control sampling error as well as simple random sampling.
Simple Random Sampling
Simple Random Sampling
If the random starting number for a systematic sample of
our TKA population
is 1786, and the sampling interval is 10, then
the first four participants selected would be the 1786th,
1796th, 1806th, and 1816th individuals on the list.
Systematic sampling is an efficient alternative to simple
random sampling, and it often generates samples
that are as representative of their populations as simple
random sampling.
Systematic Sampling
Systematic sampling is a process by which the
researcher
selects every nth person on a list. To generate a
systematic
sample of 300 participants from the TKA
population
of 3000, we would select every 10th person. The
list
of 3000 patients might be ordered by patient
number,
Social Security number, date of surgery, or birth
date.
Systematic Sampling
To begin the systematic sampling procedure, a random
start within the list of 3000 patients is necessary. To get
a random start, we can, for example, point to a number
on a random numbers table, observe four digits of the
license plate number on a car in the parking lot, reverse
the last four digits of the accession number of a library
book, or ask four different people to select numbers
between 0 and 9 and combine them to form a starting
Number.
Systematic Sampling
There are endless ways to select the random starting
number for a systematic sample.
If the random starting number for a systematic sample of
our TKA population
is 1786, and the sampling interval is 10, then
the first four participants selected would be the 1786th,
1796th, 1806th, and 1816th individuals on the list.
Systematic Sampling
Systematic sampling is an efficient alternative to simple
random sampling, and it often generates samples
that are as representative of their populations as simple
random sampling.
The exception to this is if the ordering
system used somehow introduces a systematic error
into the sample. Assume that we use dates of surgery to
order our TKA sample and that for most weeks during
the 5-year period 10 surgeries were performed. Because
the sampling interval is 10, and there were usually 10
surgeries performed per week, systematic sampling
would tend to over represent patients who had surgery
on a certain day of the week.
Systematic Sampling
Systematic Sampling
This above table shows an example
of how patients with surgery on Monday might be
overrepresented in the systematic sample; the boldface
entries indicate the units chosen for the sample. If certain
surgeons usually perform their TKAs on Tuesday,
their patients would be underrepresented in the sample.
If patients who are scheduled for surgery on Monday
typically have fewer medical complications than those
scheduled for surgery later in the week, this will also
bias the sample.
Stratified Sampling
Stratified sampling is used when certain subgroups must
be represented in adequate numbers within the sample
or when it is important to preserve the proportions of
subgroups in the population within the sample.
In our TKA study, if we hope to make generalizations across
the eight hospitals within the study, we need to be sure
there are enough patients from each hospital in the
sample to provide a reasonable basis for making statements
about the outcomes of TKA at each hospital. On
the other hand, if we want to generalize results to the
“average” patient undergoing a TKA, then we need to
have proportional representation of participants from
the eight hospitals.
Stratified Sampling
Stratified Sampling
Stratified sampling from the accessible population
is implemented in several steps. First, all units
in the accessible population are identified according
to the stratification . Second, the appropriate
number of participants is selected from each stratum.
Participants may be selected from each stratum through
simple random sampling or systematic sampling.
Stratified Sampling
More than one stratum may be identified. For instance, we
might want to ensure that each of the eight hospitals
and each of the 5 years of the study period are equally
represented in the sample. In this case, we first stratify
the accessible population into eight groups by hospital,
then stratify each hospital into five groups by year,
and finally draw a random sample from each of the 40
Hospital × Year subgroups.
Stratified Sampling
Stratified sampling is easy to accomplish if the stratifying
characteristic is known for each sampling unit.
In our TKA study, both the hospital and year of surgery
are known for all elements in the sampling frame. In
fact, those characteristics were required for placement
of participants into the accessible population.
A much different situation exists, however, if we decide that
it
is important to ensure that certain knee replacement
models are represented in the sample in adequate numbers.
Stratifying according to this characteristic would require that
someone read all 3000 medical charts to
determine which knee model was used for each subject.
Stratified Sampling
Because of the inordinate amount of time it would take
to determine the knee model for each potential subject,
we should consider whether simple random or systematic
sampling would likely result in a good representation
of each knee model
Stratified Sampling
In summary, stratified sampling is useful when a
researcher believes it is imperative to ensure that certain
characteristics are represented in a sample in specified
numbers. Stratifying on some variables will prove to be
too costly and must therefore be left to chance. In many
cases, simple random or systematic sampling will result
in an adequate distribution of the variable in question.
Cluster Sampling
Cluster sampling is the use of naturally occurring
groups as the sampling units. It is used when an
appropriate
sampling frame does not exist or when logistical
constraints limit the researcher’s ability to travel widely.
There are often several stages to a cluster sampling
procedure.
For example, if we wanted to conduct a nationwide
study on outcomes after TKA, we could not use
simple random sampling because the entire population
of patients with TKA is not enumerated—that is, a
nationwide sampling frame does not exist.
Cluster Sampling
In addition we do not have the funds to travel to all the
states and
cities that would be represented if a nationwide random
sample were selected. To generate a nationwide cluster
sample of patients who have undergone a TKA, therefore,
we could first sample states, then cities within
each selected state, then hospitals within each selected
city, and then patients within each selected hospital.
Cluster Sampling
Sampling frames for all of these clusters exist: The
50 states are known, various references list cities and
populations
within each state,and other references list hospitals by city and
size.
Each step of the cluster sampling procedure can be
implemented through simple random, systematic, or
stratified sampling. Assume that we have the money
and time to study patients in six states.
To select six states, we might stratify according to region and
then randomly select one state from each region. From
each of the six states selected, we might develop a list
of all cities with populations greater than 50,000 and
randomly select two cities from this list.
Cluster Sampling
The selection could be random or could be stratified
according to city
size so that one larger and one smaller city within each
state are selected. From each city, we might select two
hospitals for study. Within each hospital, patients who
underwent TKA in the appropriate time frame would be
selected randomly, systematically, or according to specified
strata.
The following pic shows the cluster sampling procedure
with all steps illustrated for one state, city, and
hospital. The same process would occur in the other
selected states, cities, and hospitals
Cluster Sampling
Cluster Sampling
Cluster sampling can save time and money compared
with simple random sampling because participants
are clustered in locations. A simple random sample of
patients after TKA would likely take us into 50 states and
hundreds of cities. The cluster sampling procedure just
described would limit the study to 12 cities in six states.
Cluster Sampling
Cluster sampling can occur on a smaller scale as well.
Assume that some researchers wish to study the
effectiveness
of a speech therapy approach for children who
stutter, and the accessible population consists of children
who seek treatment for stuttering in a single city.
This example is well suited to cluster sampling because
a sampling frame of this population does not exist.
In addition, it would be difficult to train all speech-
language
pathologists within the city to use the new
approach with their clients.
Cluster Sampling
Cluster sampling for a few speech therapy departments or
practices, and then a
therapists within each department or practice, would be
an efficient use of the researcher’s time for both
identifying
appropriate children and training therapists with
the new modality.
NONPROBABILITY SAMPLING
Non probability sampling is widely used in rehabilitation
research and is distinguished from probability sampling
by the absence of randomization. One reason for
the predominance of non probability sampling in
rehabilitation
research is limited funding.
Samples of Convenience
Samples of convenience involve the use of readily available
participants. Rehabilitation researchers commonly
use samples of convenience of patients in certain diagnostic
categories at a single clinic. If we conducted our study of
patients after TKA by
using all patients who underwent the surgery from 2010 to
2014 at a given
hospital, this would represent a sample of convenience.
If patients who undergo TKA at this hospital are different
in some way from the overall population of patients
who have this surgery, our study would have little
generalizability
beyond that facility.
The term “sample of convenience” seems to give the
negative implication that the researcher has not worked
hard enough at the task of sampling.
In addition we already know that probability samples tend to
have less
sampling error than no probability samples.
An accessible population that consists of “patients post-TKA
who are
60 years old or older and had the surgery at one of eight
hospitals in Indianapolis from 2010 to 2014” is technically
a large sample of convenience from the population
of all the individuals in the world who have undergone
TKA.
The white dots represent elements within the
population. The large gray ellipse represents the target population. The black ellipse
represents the accessible population. The small gray ellipse represents a sample of
convenience. The cross-hatched dots represent a random sample from the accessible
population.
The following fig shows the distinctions among a target
population, an accessible population, a random sample,
and a sample of convenience.
Consecutive sampling is a form of convenience sampling.
Consecutive samples are used in a prospective
study in which the population does not exist at the
beginning of the study; in other words, a sampling
frame does not exist. If researchers plan a 2-year
prospective
study of the outcomes of TKA at a particular
hospital beginning January 1, 2010, the population of
interest does not begin to exist until the first patient has
surgery in 2010. In a consecutive sample, all patients
who meet the criteria are placed into the study as they
are identified. This continues until a specified number
of patients are collected, a specified time frame has
passed, or certain statistical outcomes are seen.
Snowball Sampling
A snowball sample may be used when the potential
members of the sample are difficult to identify. In a
snowball sample, researchers identify a few participants
who are then asked to identify other potential members
of the sample. If a team of researchers wishes to study
patients who return to sports activities earlier than
recommended after ligament reconstruction surgery,
snowball sampling is one way to generate a sufficient
sample.
Snowball Sampling
The investigators will not be able to purchase
a mailing list of such patients, nor will they be able to
determine return-to-sport dates reliably from medical
records because many patients will not disclose their
early return to the health care providers who advised
against it. However, if the researchers can use personal
contacts to identify a few participants who returned
early, it is likely that those participants will be able to
identify other potential participants from among their
teammates or workout partners.
Purposive Sampling
Purposive sampling is a specialized form of non probability
sampling that is typically used for qualitative
research. Purposive sampling is used when a researcher
has a specific
reason for selecting particular participants
for study. Whereas convenience sampling uses
whatever units are readily available, purposive sampling
uses handpicked units that meet the researcher’s
needs.
Purposive Sampling
Random, convenience, and purposive samples
can be distinguished if we return to the hypothetical
study of different educational modes for teaching children
about physical disabilities.
If there are 40 elementary
schools in a district and the researchers randomly
select two of them for study, this is clearly a random
sample of schools from the accessible population of a
single school district. For a sample of convenience, the
researchers might select two schools in close proximity.
Purposive Sampling
For a purposive sample, the researchers might pick one
school because it is large and students are from families
with high median incomes and pick a second school
because it is small and draws students from families
with modest median incomes. Rather than selecting for
a representative group of participants, the researchers
deliberately pick participants who illustrate different levels of
variables they believe may be important to the question at hand.
Patton succinctly contrasts random sampling with purposive
sampling
The logic and power of random sampling derive from
statistical probability theory. A random and statistically
representative sample permits confident generalization
from a sample to a larger population.