0% found this document useful (0 votes)
10 views9 pages

Chapter One

Chapter One discusses sampling and sampling distribution, defining key concepts such as population, sample, and sampling error. It outlines the reasons for sampling, including cost, timeliness, and accessibility, and describes various sampling methods, including probability and non-probability sampling techniques. The chapter also explains the importance of sample size and introduces the concept of sampling distribution, emphasizing its properties and significance in statistical analysis.

Uploaded by

kemalmehamed84
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views9 pages

Chapter One

Chapter One discusses sampling and sampling distribution, defining key concepts such as population, sample, and sampling error. It outlines the reasons for sampling, including cost, timeliness, and accessibility, and describes various sampling methods, including probability and non-probability sampling techniques. The chapter also explains the importance of sample size and introduces the concept of sampling distribution, emphasizing its properties and significance in statistical analysis.

Uploaded by

kemalmehamed84
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

CHAPTER ONE

Sampling and Sampling distribution


Sampling theory
Sampling is simply the process of learning about the population on the basis of the sample drawn from
it. Thus in sampling technique instead of every unit of the population only part of the population is
studied and the conclusions are drawn on that basis for the entire population. The process of sampling
involves three elements: selecting the sample, collecting the information and making an inference about
the population.

Basic concepts of sampling theory


Population: In statistics term population is used to mean the totality of cases (items) under
consideration in a given investigation or research. In other word, the largest collection of observation on
a variable constitutes the population
Census: The process of gathering data from every element in the population
Sample: Is part of population of interest. Any non-empty subset of a population is called sample. There
are different possible samples that can be selected from a single population. Nevertheless, the one that
best reflects or represents the behavior of the population is considered to be the most appropriate one.
Sampling: The method of selecting a sample from a population.
Statistics: It is a measurable characteristic of the sample. In short it is a sample result.
Parameter: It is a measurable characteristic of the population or it is a numerical result obtained as
measuring the population.
Sampling frame: The list of all possible units in the reference population
Sample size: The number of elements observations in a specific sample
Sampling error: The difference between sample statistic and population parameters
Sampling unit: Elements of the population to be sampled or the unit of selection in the sampling process
Sample design: Is the set of procedures for selecting the sample elements from the population

Reason for sampling


The following are the major reasons for sampling technique:

Cost/Economy
Unit cost of collecting data in the case of census is significantly less than in the case of sampling.
However, due to the larger number of items in the population, the total cost involves in the case of
census is significantly higher than in the case of sampling. Suppose it takes birr 200 per unit to make a
census of 10,000,000 individuals but the unit cost of sampling 5,000 individuals is birr 1000. Thus, the
total cost is: 10,000,000 *200 = 2,000,000,000 but that of sample is 5,000*1000 = 50,000,000

Timeliness
Due to the large size of population total time involves in the case of census is significantly higher than
that of sampling (i.e., the sample may provide us with necessary information quickly)
Large population size
Sometimes, many populations about which inferences must be made are quite large implying that it is
impossible to cover all the items in the population. Thus, the solution is to take sample from such a
population

Inaccessibility of the entire population


In some cases the entire population may not be accessible due to disease, death, conflict, mental
abnormality, prisoners, etc. in the case sampling is necessary

Destructive nature of many tests


Due to destructive nature of many tests, the resources are completed to collect information only from
part of the population. For example: blood test for a patient, life hours of a tube light, strength of wires,
etc.

Accuracy
Non-sampling error in the case of census is higher than the non-sampling error committed in the case of
a sample survey (as less qualified investigator are involve in the case of census and the supervision,
monitoring and quality control mechanism in the case of census may be poor). The higher the degree of
non-sampling error, the less reliable your result may be.

Sampling methods
There are two principal methods of drawing a sample from a population: probability sampling and non-
probability.

1. Probability sampling

In the case of probability sampling each observation in the population has an equal chance of being
selected to become part of the population. There is no human judgment in the case of probability
sampling

There are four basic types of probability sampling techniques

i. Simple random sampling


ii. Stratified sampling
iii. Systematic sampling
iv. Cluster sampling

i. Simple random sample


Simple random sampling is a method of probability sampling in which every unit in the population has
an equal non-zero chance of being selected (or part of the sample). In other words, each element of the
population has an equal and independent chance of being included into the sample. The probability is
given by n/N
There are two methods to select a sample of random sample
Lottery method- In this method, each population item is numbered 1 to N on slips of identical cards
(size, shape and color). Then place numbered cards in a bowl, mix them thoroughly, and select as many
cards as needed in a blind fold selection. The subjects whose numbers are selected constitute the
sample. Since it is difficult to mix the cards thoroughly, there is a chance of obtaining a biased sample.
Thus we need other method of selecting sample elements.

Random number method- due to the problem of lottery method, statisticians uses another method
known as the random number method where numbers are generated using computers.

How to use random number table method


A) Assign a unique number to each population elements in the sampling frame. Start with serial
number 1, or 01, or 001, etc. depending on the number of digits required
B) Choose a random starting position by closing your eyes (blind fold selection) and placing your
finger on a number in the table.
C) Select serial numbers across rows or down columns or diagonally from the starting point
D) Discard numbers that are not assigned to any population elements and ignore numbers that
have already been selected
E) Repeat the selection process until the required number of sample elements is selected.

Advantage of simple random sampling


 It ensures that the sample is unbiased

Disadvantages of simple random sampling


 It requires a sampling frame, and this is sometimes impossible (the case of fish population)
 If the population is very large, it is tedious and time consuming to number and select the
sample
 Minority subgroups of the population may not be represented in the sample.

ii. Stratified sampling


In stratified sampling, a population is first divided into subgroups, called strata (singular stratum). And a
sample is selected from each stratum based on simple random or systematic sampling method. The
strata are made according to various homogeneous characteristics such as sex, race, region or
institutional affiliation such as faculty. Stratified sampling is applied if the population is heterogeneous

Stratified random sampling method is a three-step process:


Step 1- divide the population into homogeneous, mutually exclusive and collectively exhaustive groups
or strata using some stratification variable (e.g. income level, sex, education level, etc.);
Step 2- select an independent simple random sample from each stratum (using simple random sample)
Step 3- from the final sample by consolidating all sample elements chosen in step 2
Stratified sample can be
Proportionate: Involving the selection of sample elements from each stratum, such that the ratio of
sample elements and total number of population elements (n/N) is constant/equal for all strata

Disproportionate: the sample is disproport0ionate when the above mentioned ratio is unequal

Example: to select a proportionate stratified sample of 20 households from Addis Ababa that belongs to
three income groups: low (50), middle (30) and high (20) (N = 50+30+20=100).

 Sub-divide the club members into three homogeneous sub-groups or strata by the income
groups: low, middle and high.
 Calculate the overall sampling fraction, f, in the following manner: f=n/N = 20/100 = 0.2 where
n= sample size and N = population size: n1 = 0.2*50 = 10, n2 = 0.2* 30 = 6 and n3 = 0.2*20 = 4.
Thus, n = n1+n2+n3 = 10 + 6 + 4 = 20

Advantages of stratified sampling


 The representation of the sample is improved

Disadvantages of stratified sampling


 If there are many variables of interest, dividing a large population in to representative
subgroups requires a great deal of effort,
 If variables are somewhat complex or ambiguous (such as beliefs, attitudes, etc), it is difficult to
separate individuals in to the sub-groups according to these variables.

iii. Systematic sampling


In systematic sampling only one random number is needed throughout the entire sampling process.
Elements of the population will be arranged in some order and the elements to be included in the
sample will be selected at a constant interval.

To use systematic sampling, a researcher needs:

a. A sampling frame of the population


b. A skip interval (K) calculated as follows

Skip interval (K) = population list/sample = N/n = K

The first element (number), which is between 1 and K, determined using simple random sampling and
then the next items are selected using the skip interval. For instance, the j th unit is selected at first and
then (j+k)th, (j+2k)th …. etc until the required sample size is obtained.

Example: suppose there are 2000 subjects in the population and a sample size is 50 subjects are needed.
The sampling interval (k) is 40 (2000/50). Select the starting point, say “x”, from 1 through 40 using
simple random sampling, and then include every 40th element starting from “x”.
Advantages of systematic sampling
 Less time consuming and easier to perform than SRS
 It is more convenient to use as compared to SRS, and
 It provides a good approximation to SRS

Disadvantage of systematic sampling


If there is any sort of cyclic ordering of the subjects, the sample will not be representative of the
population. Example: if subjects in the population are arranged in a manner such as:

 Defective item
 Non-defective item etc,

Iv. Cluster sampling


Cluster sampling can be used if the population is homogeneous and very large in size. It is a type of
sampling in which the population is divided into non-overlapping heterogeneous groups called clusters
or groups and clusters/groups of elements are sampled as the sampling units using simple random
sampling technique in the first phase (if it is the two-phase cluster sampling). In other words, cluster
sampling is a type of sampling which involves dividing the population into groups (clusters). Then, one or
more clusters are chosen at random and individual with in the chosen cluster is sampled.

A two-step process
Step 1- define population is divided into number of mutually exclusive and collectively exhaustive
heterogeneous groups/clusters:

Step 2- select an independent simple random sample of clusters using simple random sampling

Advantages of cluster sampling


 A list of all individual study units in the reference population is not required
 Reduces cost, and
 Simplifies field work and it is convenient

Disadvantages of cluster sampling


 The members of the clusters are often more homogeneous than the members of the whole
population and therefore, it may not be representative.
 The elements in a cluster may not have the same variation in characteristics as elements
selected individually from the population

2) Non-probability sampling
In the case of non-probability sampling, not every unit in the population has a chance of being included
in the sample. It involves at least some degree of personal subjectivity instead of following
predetermined, probabilistic rules for selection.

There are three basic types of non-probability techniques

i. Convenience sampling
ii. Judgmental sampling
iii. Quota sampling

i. Convenience sampling
Convenience sampling implies sample drawn at the convenience of the research. It is common in
exploratory research. Does not lead to any conclusion

ii. Judgmental sampling


Sampling based on some judgment, gut-feelings or experience of the researcher. It is common in
commercial marketing research projects. If inference drawing is not necessary, these samples are quite
useful.

iii. Quota sampling


In this method, the decision maker requires the sample to contain a certain number of items with a
given characteristics. It is something like judgmental sampling

Example: suppose we know that 54% of the adults in a community are females, and the study requires
100 respondents as a sample. In quota sampling, we might interview the first 54 females and the first 46
males.

Determinant factors of the sample size


Size of sample means the number of sampling units selected from the population for investigation. If the
size of sample is small it may not represent the population and the inference drawn about the
population may be misleading. On the other hand, if the size of sample is very large, it may be too
burdensome financially, requires a lot of time and may have a serious problem of managing it. Hence
the sample size should be neither too small nor too large. Rather it should be optimal

The following factors should be considered while deciding the sample size

i. The size of the population: the larger the size of the population, the bigger should be the
sample size
ii. The resource available: if the resources available are vast a large sample size could be
taken. However, in most cases resources constitute a big constraint on sample size.
iii. The degree of accuracy or precision desired: the greater the degree of accuracy desired, the
larger should be the sample size. However, it does not necessarily mean that bigger samples
always ensure greater accuracy.
iv. Homogeneous or heterogeneous of the population: if the population consists of
homogeneous units a small sample may serve the purpose, but if the population consists of
heterogeneous units a large sample may be inevitable.
v. Nature of study: for an intensive and continuous study a small sample may be suitable. But
for studies which are not likely to be repeated and are quite extensive in nature, it may be
necessary to take a large sample size.
vi. Method of sampling adopted: the size of sample is also influenced by the type of sampling
plan adopted. For example if the sample is a simple random sample it may necessitate a
bigger sample size. However, in a properly drawn stratified sampling plan, even a small
sample may give better results.
vii. Nature of respondent: where it is expected a large number of respondents will not co-
operate and send back the questionnaire, a large sample should be selected.

Sampling distribution
Note: The normal probability distribution is used to determine probabilities for the normally
distributed individual measurements, given the mean and the standard deviation. Symbolically, the
variable is the measurement X, with the population mean µ and population standard deviation δ. In
contrast to such distributions of individual measurements, a sampling distribution is a probability
distribution for the possible values of a sample statistics.

Population distribution: Is the distribution of measured values of its members and have mean by µ
and variance δ2 and standard deviation describes the variation among values of members of the
population: whereas the standard deviation of sampling distribution measures the variability among
values of the statistics (sample) such as mean values, proportion values due to sampling errors.

Sample distribution: is the distribution of measured values of sample in random samples drawn
from a given population. Each sample mean would vary from sample to sample. This variability
serves as the basis for random sampling distribution. A sampling distribution is a probability
distribution for the possible values of a sample statistics, such as a sample mean.

Sampling distribution of the mean


Sampling distribution of the mean: is the probability distribution of all possible values of a given
statistics (sample) from all district possible sample of equal size drawn from a population or a
process. The sampling distribution of the mean values has its own arithmetic mean denoted by µx
(read as mu sub Ⱦ) and standard deviation δx (sigma sub Ⱦ). The sampling distribution of the mean is
the probability distribution of the means, Ⱦ of all simple random samples of a given sample size n
that can be drawn from the population.
NB: The sampling distribution of the mean is not the sample distribution, which is the distribution of
the measured values of X in one random sample. Rather the sampling distribution of the mean is the
probability distribution for Ⱦ, the sample mean.

For any given sample size n taken from a population with mean µ and standard deviation δ, the
value of the sample mean Ⱦ would vary from sample to sample if several random samples were
obtained from the population. This variability serves as the basis for sampling distribution.

The sampling distribution of the mean is described by two parameters: the expected value (Ⱦ) = Ⱦ
bar, or mean of the sampling distribution of the mean, and the standard deviation of the mean δ Ⱦ,
the standard error of the mean.

Properties of the sampling distribution of means


1. The arithmetic mean µx of sampling distribution of mean values is equal to the population mean
µ regardless of the form of population distribution, i.e. µx = µ
2. The sampling distribution has a standard deviation (also called standard error) equal to the
population standard deviation divided by the square root of the sample size i.e. δx = δ/(N) 1/2.
This hold true if and only if n < 0.05N and N is very large. If N is finite and

n>0.05N , δȾ = S/(n)1/2 or , δȾ = S/(n)1/2 *(N-n/N-1)1/2

3. A sample size n≥30 is generally said to be considered to be a large sample for statistical analysis
where as a sample of size n <30 is considered to be a small sample. The sampling distribution of
means is approximately normal for sufficiently large sample sizes (n≥30).
4. When standard deviation of population δ is not known, the standard deviation of the sample s
which closely approximates δ value is used to compute standard error, i.e. δx = S/(n) 1/2

Example 1. A population consists of the following ages: 10, 20, 30, 40, and 50. A random sample of three
is to be selected from this population and mean computed. Develop the sampling distribution of the
mean.

Solution: The number of sample random samples of size n that can be drawn without replacement from
the population of size N is NCn = N!/ n!(N-n)!, with N=5 and n=3,5C3 = 10 samples can be drawn from
the population as:

Sample items Sample means (Ⱦ)


10, 20, 30 20.00
10, 20, 40 23.33
10, 20, 50 26.67
10, 30, 40 26.67
10, 30, 50 30.00
10, 40, 50 33.33
20, 30, 40 30.00
20, 30, 50 33.33
20, 40, 50 36.67
30, 40, 50 40.00
300.00
A systematic organization of the above frequency gives the following
Sample mean (Ⱦ) frequency Prob. (relative freq.) of Ⱦ
20.00 1 0.1
23.33 1 0.1
26.67 2 0.2
30.00 2 0.2
33.33 2 0.2
36.67 1 0.1
40.00 1 0.1
TOTAL 10 1.00
 Columns 1 and 2 show frequency distribution of sample means.
 Columns 1 and 3 show sampling distribution of the mean
µ = ⅀ X /N = ⅀ Ⱦ / n = 30 regardless of the sample size µ =Ⱦ

X (observation) x-µ (x - µ)2


10 -20 400
20 -10 100
30 0 0
40 10 100
50 20 400
⅀ (x - µ)2 1,000

You might also like