0% found this document useful (0 votes)
12 views36 pages

Characteristics of Effective Sampling

The document provides an overview of sampling methods, including definitions, advantages, and types of sampling techniques such as nonprobability and probability sampling. It discusses errors in sampling, including sampling error and non-sampling error, and emphasizes the importance of determining an appropriate sample size for studies. Additionally, it outlines the calculation methods for estimating sample sizes based on population characteristics and desired confidence levels.

Uploaded by

eldana.endale77
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views36 pages

Characteristics of Effective Sampling

The document provides an overview of sampling methods, including definitions, advantages, and types of sampling techniques such as nonprobability and probability sampling. It discusses errors in sampling, including sampling error and non-sampling error, and emphasizes the importance of determining an appropriate sample size for studies. Additionally, it outlines the calculation methods for estimating sample sizes based on population characteristics and desired confidence levels.

Uploaded by

eldana.endale77
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Sampling methods

1
Sampling methods

Outlines

 Introduction
 Sampling methods
 Errors in sampling
 Sample size

2
I) Introduction

♣ Sampling involves the selection of a number of study units


from a defined population.
♣ The population is too large for us to consider collecting
information from all its members.
♣ If the whole population is taken there is no need of
statistical inference.
♣ Usually, a representative subgroup of the population
(sample) is included in the investigation.
♣ A representative sample has all the important
characteristics of the population from which it is drawn.

3
Introduction
Advantages of samples
♣ Cost - sampling saves time, labour and money

♣ Quality of data - more time and effort can be


spent on getting reliable data on each individual
sampled.
♣ Due to the use of better trained personnel, more
careful supervision and processing a sample can
actually produce precise results.

4
Introduction

If we have to draw a sample, we will be confronted with


the following questions:

a) What is the group of people ( population) from which we


want to draw a sample?

b) How many people do we need in our sample?

c) How will these people be selected?

 Apart from persons, a population may consist of


mosquitoes, villages, institutions, etc.

5
Introduction

Definitions:
• Reference population (also called source
population or target population) :

the population of interest, to which the investigators would


like to generalize the results of the study, and from which a
representative sample is to be drawn.

• Study or sample population - the population included in


the sample

6
Introduction

♣ Sampling unit - the unit of selection in the sampling process


♣ Study unit - the unit on which information is collected.

- the sampling unit is not necessarily the same as the study


unit.
- if the objective is to determine the availability of latrine,
then the study unit would be the household; if the objective
is to determine the prevalence of trachoma, then the study
unit would be the individual.
♣ Sampling frame - the list of all the units in the reference
population, from which a sample is to be picked.
♣ Sampling fraction (Sampling interval) - the ratio of the number
of units in the sample to the number of units in the reference population
(n/N).

7
II) Sampling methods (Two broad divisions)

A) Nonprobability Sampling Methods


♣ used when a sampling frame does not exist
♣ no random selection(unrepresentative of the given
population)

♣ inappropriate if the aim is to measure variables


and generalize findings obtained from a sample to
the population.

8
Nonprobability…
Two such nonprobability sampling methods are:

1)Convenience sampling: is a method in which for


convenience sake the study units that happen to be available
at the time of data collection are selected.

2) Quota sampling: is a method that ensures that a certain


number of sample units from different categories with
specific characteristics are represented. In this method the
investigator interviews as many people in each category of
study unit as he can find until he has filled his quota.
• Both the above methods do not claim to be representative
of the entire population.

9
Nonprobability…
3)Judgmental sampling or Purposive sampling
• The researcher chooses the sample based on who they think would be
appropriate for the study.
• This is used primarily when there is a limited number of people that
have expertise in the area being researched
4) Snowball sampling (friend of friend….etc.)

10
B) Probability Sampling methods

- a sampling frame exists or can be compiled.

- involve random selection procedures. All units of the population


should have an equal or at least a known chance of being included in
the sample.

- generalization is possible (from sample to population)

11
1. Simple random sampling (SRS)
- this is the most basic scheme of random sampling.
- each unit in the sampling frame has an equal chance of being selected
- representativeness of the sample is ensured.
However, it is costly to conduct SRS. Moreover, minority subgroups of interest
in the population my not be present in the sample in sufficient numbers for
study.

To select a simple random sample you need to:


 make a numbered list of all the units in the population from which you want to
draw a sample.
each unit on the list should be numbered in sequence from 1 to N (where N is
the size of the population)
decide on the size of the sample
select the required number of study units, using e.g. a “lottery” method

12
2. Systematic Sampling
 Individuals are chosen at regular intervals ( for example, every
kth) from the sampling frame.

 The first unit to be selected is taken at random from among the


first k units.

For example, a systematic sample is to be selected from 1200


students of a school. The sample size is decided to be 100. The
sampling fraction is: 100 /1200 = 1/12.

13
Systematic Sampling

• The number of the first student to be included in the sample is chosen


randomly, for example by blindly picking one out of twelve pieces of
paper, numbered 1 to 12. If number 6 is picked, every twelfth student
will be included in the sample, starting with student number 6, until 100
students are selected. The numbers selected would be 6,18,30,42,etc.

Merits
 Systematic sampling is usually less time consuming and easier to
perform than simple random sampling. It provides a good approximation
to SRS.
 Unlike SRS, systematic sampling can be conducted without a sampling
frame (useful in some situations where a sampling frame is not readily
available).
Eg., In patients attending a health center, where it is not possible to
predict in advance who will be attending.

14
Systematic Sampling

Demerits
 If there is any sort of cyclic pattern in the ordering of the
subjects which coincides with the sampling interval, the
sample will not be representative of the population.

Example:
- list of married couples arranged with men's names
alternatively with the women's names will result in a sample
of all men or women.

15
3. Stratified Sampling:
appropriate when the distribution of the characteristic to be
studied is strongly affected by certain variable (heterogeneous
population).
the population is first divided into groups (strata) according to
a characteristic of interest (eg., sex, geographic area,
prevalence of disease, etc.)
A separate sample is taken independently from each stratum,
by simple random or systematic sampling.
 Proportional allocation - if the same sampling fraction is used
for each stratum.
Non- proportional allocation - if a different sampling fraction
is used for each stratum or if the strata are
unequal in size and a fixed number of units is selected from
each stratum.

16
Stratified Sampling:
Merit
- The representativeness of the sample is improved.
That is, adequate representation of minority
subgroups of interest can be ensured by
stratification and by varying the sampling fraction
between strata as required.

Demerit
- sampling frame for the entire population has to
be prepared separately for each stratum.

17
4. Cluster sampling
 the selection of groups of study units (clusters) instead of the selection of
study units individually

 the sampling unit is a cluster, and the sampling frame is a list of these
clusters.

 procedure - the reference population (homogeneous) is divided into


clusters. These clusters are often geographic units (eg
districts, villages, etc.).
- a sample of such clusters is selected
- all the units in the selected clusters are studied.

 it is preferable to select a large number of small clusters rather than a small


number of large clusters.

18
Cluster sampling

Merit - A list of all the individual study units in the


reference population is not required. It is sufficient
to have a list of clusters.

Demerit - It is based on the assumption that the


characteristic to be studied is uniformly distributed
throughout the reference population, which may
not always be the case. Hence, sampling error is
usually higher than for a simple random sample of
the same size.

19
5. Multi-stage sampling
 This method is appropriate when the reference population is large and widely
scattered
 selection is done in stages until the final sampling unit (eg., households or persons)
are arrived at.
 The primary sampling unit (PSU) is the sampling unit (usually large size) in the first
sampling stage.
 The secondary sampling unit (SSU) is the sampling unit in the second sampling stage.
 etc.

Example - The PSUs could be kebeles and the SSUs could be households.

Merit - Cuts the cost of preparing sampling frame


Demerit - Sampling error is increased compared with a simple random sample.
• Multistage sampling gives less precise estimates than simple random sampling for
the same sample size, but the reduction in cost usually far outweighs this, and
allows for a larger sample size.

20
III) Errors in sampling

• When we take a sample, our results will not exactly equal the correct
results for the whole population. That is, our results will be subject to
errors.
A) Sampling error (random error)

 A sample is a subset of a population. Because of this property of


samples, results obtained from them cannot reflect the full range of
variation found in the larger group (population).

 This type of error, arising from the sampling process itself, is called
sampling error, which is a form of random error.

 Sampling error can be minimized by increasing the size of the sample.

21
Errors in sampling

B) Non-sampling error (bias)


 Systematic error in the design or conduct of a
sampling procedure which results in distortion of
the sample , so that it is no longer representative of
the reference population.

 We can eliminate or reduce the non-sampling


error (bias) by careful design of the sampling
procedure and not by increasing the sample size.

22
Errors in sampling

Example : If you take male students only from a student dormitory in Ethiopia in
order to determine the proportion of smokers, you would result in an
overestimate, since females are less likely to smoke. Increasing the number of
male students would not remove the bias.

 There are several possible sources of bias in sampling ( eg., accessibility bias,
volunteer bias, etc.)

 The best known source of bias is nonresponse. It is the failure to obtain


information on some of the subjects included in the sample to be studied.

 Noneresponse results in significant bias when the following two conditions are
both fulfilled
- When non-respondents constitute a significant proportion of the
sample (about 15% or more)
- When non-respondents differ significantly from respondents.

23
Errors in sampling

 There are several ways to deal with this problem and reduce the
possibility of bias:

a) Data collection tools (questionnaire) have to be pre-tested.

b) If nonresponse is due to absence of the subjects, repeated attempts


should be considered to contact study subjects who were absent at
the time of the initial visit.

c) To include additional people in the sample, so that non-respondents


who were absent during data collection can be replaced (make sure
that their absence is not related to the topic being studied).

N.B. : The number of nonresponses should be documented according


to type, so as to facilitate an assessment of the extent of bias
introduced by nonresponse.

24
IV) Sample size determination
♣ In planning any investigation we must decide how many people need to be
studied in order to answer the study objectives. If the study is too small we
may fail to detect important effects, or may estimate effects too imprecisely.
If the study is too large then we will waste resources.

♣ In general, it is much better to increase the accuracy of data collection (by


improving the training of data collectors and data collection tools) than to
increase the sample size after
a certain point.

♣ The eventual sample size is usually a compromise between what is desirable


and what is feasible.

♣ The feasible sample size is determined by the availability of resources. It is


also important to remember that resources are not only needed to collect the
information, but also to analyse it.

25
Sample size determination
In order to calculate the required sample size, you need to know the following
facts:

The reasonable estimate of the key proportion to be studied. If you cannot guess the
proportion, take it as 50%.
The degree of accuracy required. That is, the allowed deviation from the true
proportion in the population as a whole. It can be within 1% or 5%, etc.

The confidence level required, usually specified as 95%.


The size of the population that the sample is to represent. If it is more than 10,000 the
precise magnitude is not likely to be very important; but if the population is less than
10,000 then a smaller sample size may be required

The difference between the two sub-groups and the value of the likelihood or the
power that helps in finding a statistically significant difference.

Note that number 5 is required when there are two population groups and the
interest is to compare between two means or proportions.

26
Estimating a proportion
estimate how big the proportion might be (P)
 choose the margin of error you will allow in the estimate of the proportion
(say  w)
choose the level of confidence that the proportion in the whole population
is indeed between (p-w) and (p+w). We can never be 100% sure. Do you
want to be 95% sure?
the minimum sample size required, for a very large population (N>10,000)
is:
n = Z2 p(1-p) / w2
Show how the above formula is obtained.
A 95% C.I. for P = p  1.96 se , if we want our confidence interval to have a
maximum width of  w,
1.96 se = w
1.96 p(1-p)/n = w
(1.96)2 p(1-p)/n = w2 , Hence, n = (1.96)2 p(1-p)/w2

27
Example 1
a) p = 0.26 , w = 0.03 , Z = 1.96 ( i.e., for a 95% C.I.)
n = (1.96)2 (.26  .74) / (.03)2 = 821.25  822
Thus , the study should include at least 822 subjects.

b) If the above sample is to be taken from a relatively


small population (say, N = 3000) , the required minimum
sample will be obtained from the above estimate by
making some adjustment .

821.25 / (1+ (821.25/3000)) = 644.7  645 subjects

28
Example 2

♣ A hospital administrator wishes to know what proportion of


discharged patients are unhappy with the care received
during hospitalization . If 95% Confidence interval is desired
to estimate the proportion within 5%, how large a sample
should be drawn ?

♣ n = Z2 p(1-p)/w2 =(1.96)2(.5.5)/(.05)2 =384.2  385 patients

♣ NB If you don’t have any information about P, take it as


50% and get the maximum value of PQ which is 1/4 (25%).

29
Estimating a mean

♣ The same approach is used but with SE =  / n


The required (minimum) sample size for a very large
population is given by :
n = Z2 2 / w2

Eg. A health officer wishes to estimate the mean serum


cholesterol in a population of men. From previous similar
studies a standard deviation of 40 mg/100ml was reported.
If he is willing to tolerate a marginal error of up to 5
mg/100ml in his estimate, how many subjects should be
included in his study ? ( =5%, two sided)

30
Comparison of two proportions
n (in each region) = (p1q1 + p2q2) (f(,)) / ((p1 - p2)²

 = type I error (level of significance)


 = type II error ( 1- = power of the study)
power = the probability of getting a significant result
f (,) =10.5, when the power = 80% and the level of significance = 5%

Eg. The proportion of nurses leaving the health service is compared


between two regions. In one region 30% of nurses is estimated to
leave the service within 3 years of graduation. In other region it is
probably 15%.

31
Solution

♣ The required sample to show, with a 90% likelihood


(power), that the percentage of nurses is different in these
two regions would be: (assume a confidence level of 95%)

n = (1.28+1.96)2 ((.3.7) +(.15 .85)) / (.30 - .15)2 = 158


158 nurses are required in each region

Comparison of two means (sample size in each group)


n = (s12 + s22) f(,) / (m1 - m2)2
m1 and s12 are mean and variance of group 1 respectively.
m2 and s22 are mean and variance of group 2 respectively.

32
Demonstration of sample estimations
Use the following software's to calculate sample size
 Epi info
Open EPI

33
Exercises

1) A nutritionist wants to determine the prevalence of malnutrition


among under 5 children in Amhara region. If a sample of 3000
children is required, what is the sampling technique he should use to
select the required subjects. Write a short note on the procedures
(steps) he should follow in selecting these subjects.

34
Exercises

2) In a school there are about 1800 students and the investigator wants
to determine the prevalence of a certain character (eg., KAP on
HIV/AIDS) by taking 450 students. Details are given below:

Grade Number of students Number of sections


9 600 8
10 500 7
11 400 6
12 300 5
Total 1800 26

How do you select the subjects who will be included in your sample?

35
Exercises
3) A multi-national clinical trial is proposed to investigate the value of a gradually
increasing dose schedule of a beta blocker in the treatment of severe heart failure.

The trial will be randomised, double-blind and placebo controlled. Each patient is to be
followed for 2 years, and the main treatment comparison is for all cause mortality.

Previous experience suggests a 2 year mortality rate of around 30%. The investigators
propose that a one-third reduction in mortality due to beta-blockade would be important
to detect. They suggest that type I and type II errors be set at 0.05 and .1, respectively.

Calculate the required number of patients to be recruited.

Suppose one anticipates that 10% of patients randomised to

beta-blockade will fail to comply with the intended treatment

policy. What change in required sample size would you suggest?

36

You might also like