Sampling is the process of selecting a part of a population to make inferences about the whole, often used in research for efficiency and accuracy. Key concepts include population parameters, sampling errors, and various sampling distributions such as the normal distribution and Student's t-distribution. Understanding these concepts is essential for conducting valid and reliable statistical analyses.
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0 ratings0% found this document useful (0 votes)
4 views32 pages
9.sampling & Statistical Inference
Sampling is the process of selecting a part of a population to make inferences about the whole, often used in research for efficiency and accuracy. Key concepts include population parameters, sampling errors, and various sampling distributions such as the normal distribution and Student's t-distribution. Understanding these concepts is essential for conducting valid and reliable statistical analyses.
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Sampling and Statistical
Inference Chapter
EE
Sampling is defined as the selection of some part of an aggregate or totality on the basis of which a
judgement or inference about the aggregate or totality is made. In other words, it is the process of
obtaining information about an entire population by examining only a part of it. In most of the research
work and surveys, the usual approach happens to be to make generalisations or to draw inferences
based on samples about the parameters of population from which the samples are taken. The
researcher quite often selects only a few items from the universe for his study purposes. All this is
done on the assumption that the sample data will enable him to estimate the population parameters.
The items so selected constitute what is technically called a sample, their selection process or technique
is called sample design and the survey conducted on the basis of sample is described as sample
survey: Sample should be truly representative of population characteristics without any bias so that it
may result in valid and reliable conclusions.
Sampling is used in practice for a variety of reasons such as:
1. Sampling can save time and money. A sample study is usually less expensive than a census
study and produces results at a relatively faster speed.
>. Sampling may enable more accurate measurements fora sample study is generally conducted
by trained and experienced investigators. :
3, Sampling remains the only way when population contains infinitely many members.
4. Sampling remains the only choice when a test involves the destruction of the item under
study.
Sampling usually
information concerning some
[A detailed discussion on the need for sampling and various sampling techniques is presented in
Chapter 4.
enables to estimate the sampling errors and, thus, assists in obtaining
characteristic of the population.
(@ scanned with OKEN ScannerResearch Methodology
9.1" PARAMETERIAND STATISTIC
‘The collection of all the items about which the information (on one or more chai
cristics) is desire
is called as population. For example, if we want to measure the average cellphone bill of the people
ina particular ci
¥ in a particular month, our population would consist of the specified month’s cellphone
bills of all the people in that city. The population or universe can be finite or infinite. The population
is said to be finite if it consists of a fixed number of elements so that it is possible to enumerate tin
its totality. For
tance, the population of a city, the number of workers in a factory are examples of
finite populations
rhe symbol ‘A” is generally used to indicate how many elements (or items) are
there in ease of a finite population. An infinite population is that population in which itis theoretically
impossible to observe all the elements. Thus, in an infinite population the number of items is infinite
i-e., we cannot have any idea about the total number of items. The number of stars in a sky, possible
rolls of a pair of dice are examples of infinite population, One should remember that no truly infinite
population of physical objects does actually exist i
n spite of the fact that many such populations
appear to be very
very large. Froma practical consideration, we then use the term infinite population
for a population that cannot be enumerated in a reasonable period of time. This way we use the
‘heoretical concept of infinite population as an approximation of very large finite population
Any characteristic or measure of population units is known as a parameter. Population mean,
population standard deviation and population proportion are commonly studied parameters. These
Parameters are denoted by U, o and 1 respectively. Parameters are usually unknown as we do not
always study the whole population. Unknown parameter are estimated by studying a subpart of the
whole population which is known as a sample.
(@ scanned with OKEN ScannerSampling and Statistical inference gq
Population
o
oe
Sampling wi
frame a
¢
=
G
Response
error
Fig. 9.4 ‘
‘Sampling error = Frame error + Chance error + Response error '
(If'we add measurement error or the non-sampling error to sampl
8 error, we get total error). '
Sampling errors occur randomly and are equally likely to be in either direction. The magnitude of :
the sampling error depends upon the nature of the universe; the more homogeneous the universe, '
smaller the sampling error. Sampling error is inversely related to the size of the sample i.e., sampling i
error decreases as the sample size increases and vice-versa. A measure of the random sampling {
error can be calculated for a given sample design and size and this measure is often called the
precision of the sampling plan, Sampling error is usually worked out as the product of the critical
Value at a certain level of significance and the standard error.
s opposed to sampling errors, we may have non-sampling errors which may creep in during the
process of collecting actual information and such errors occur in all surveys whether census or
sample. We have no way to measure non-sampling errors.
3 SAMPLING DISTRIBUTION
Distribution of. statistic may not be the same as the distribution of population. We are often concerned |
With sampling distribution in sampling analysis. If we take certain number of samples and for each
sample compute various statistical measures such as mean, standard deviation, ete., then we can find
that each sample may give its own value for the statistic under consideration. All such values of a
Particular statistic, say mean, together with their relative frequencies will constitute the sampling
distribution of the particular statistic, say mean, Accordingly, we can have sampling distribution of
Mean, or the sampling distribution of standard deviation or the sampling distribution of any other
Statistical measure.
(@ scanned with OKEN Scanner7
Research Methodology |
C150)
Here, it is important to discuss normal distribution briefly.
the bell shape of the histogram of given data-set as below :
ution is represented hy
Normal dis
Fig, 9.2
In many statistical analyses, it is important that the given data has normal distribution, Normal
distribution is denoted by N(11,0) , where 11 is mean and ¢ is standard deviation of the distribution
Normal distribution can be used to model any variable taking values ~2°to%, provided the valueson
that variable have a bell shaped histogram. Reader may refer to some standard text on statistics to
know more about normal distribution.
Some commonly used sampling dist
uutions are given below:
9.3.1 Sampling Distribution of Mean
Sampling distribution of mean refers tothe probability distribution ofall the possible means of random
samples ofa given size that we take from a population. If samples ae taken from a normal population
1N (1,9), the sampling distribution of mean would also be normal with mean wu, = w and stand!
deviation =o,/n , where 1 isthe mean ofthe population, is the standard deviation of the poplai®
and means the number of items in a sample. But when sam;
normal (may be positively or negatively skewed), even the
sampling distribution of mean tends quite closer to the no
sample items is large i.c., more than 30. In case we want
pling is from a population which is"
N, as per the central limit theorem, the
mal distribution, provided the number
to reduce the sampling distribution of m=#"
x- hing
for the samp!
olin 7
: « ceveltl
is very useful in seve™
to unit normal distribution i.e., (0, 1), we can write the normal vari
variate z:
distribution of mean. This characteristic of the sampling distri;
decision situations. ipling distribution of mean
9.3.2 Sampling Distribution of Proportion
Proportion is a measure of attribute. Let us consi
: ; - msider that bo yh iy
exclusive and collectively exhaustive classes one Ga aad is divided ino twos
‘sessing a particular attribute “
4
(@ scanned with OKEN Scanner‘sampling and Statistical Inference @®
sing that attribute. For example, people in city could be divided into “Smokers”
Let \ = population size, X= number of people out of V possessing a particular
attribute, «= X/N = actual proportion of the people possessing the specified attribute. Let a sample
is selected from this population with m = sample size,.x = number of people in the sample possessing
the specified particular attribute, p =.x/n = sample proportion. Note that, X and P are population
parameters, while x and p are sample statistics. Also, p provides an estimate of m. It can be showed
that the distribution of x is Binomial(n, m). Using the property of binomial distribution, for sufficiently
large sample size n, we have
and “Non-smokers’
mm. _Pt__ No,
Inm(\—n) fx(l—m)/n
Practically, this result is true for n230 or, when nn25 as well as n(I—m)25
CHAPTER 9
9.3.3 Student's t-Distribution
When population standard deviation (0) is not known and the sample is of a small size (i.e., $30),
we use distribution for the sampling distribution of mean and workout /-variable as:
J
Ech where sf = > is sample variance. Student’s f-distribution (or
ution) is also symmetrical and is very close to the distribution of standard normal variate, 2,
except for small values of 7. The variable 1 differs from = in the sense that we use sample standard
deviation (s,) in the calculation of ¢, whereas we use standard deviation of population (6) in the
calculation of z, There is a different /-distribution for every possible sample size i.e., for different
degrees of freedom, The degrees of freedom for a sample of size » isn ~1. As the sample size gets
larger, the shape of the /-distribution becomes apporximately equal to the normal distribution. In fact
for sample sizes of more than 30, the /-distribution is so close to the normal distribution that we can
use the normal to approximate the -distribution. Like normal distribution , the range of -distribution
is also (0,0) . The applications of ¢-distribution are discussed in further chapters.
4
9.3.4 Chi-square (x?) Distribution
Chi-square distribution is encountered when we deal with collections of values that involve adding up
squares. Variances of samples require us to add a collection of squared quantities and thus have
distributions that are related to chi-square distribution. Suppose we have a random sample x,, x,,
"
x, from N(j1,0) population, then the distribution of the statistic) ( =*) is chi-square distribution
a
with n degree of freedom. Alternatively, the distribution of the statistic >» (
ial
is chi-square distribution with (n— /) degree of freedom. Clearly, the range of chi-square distribution
(@ scanned with OKEN Scanner