0% found this document useful (0 votes)
4 views32 pages

9.sampling & Statistical Inference

Sampling is the process of selecting a part of a population to make inferences about the whole, often used in research for efficiency and accuracy. Key concepts include population parameters, sampling errors, and various sampling distributions such as the normal distribution and Student's t-distribution. Understanding these concepts is essential for conducting valid and reliable statistical analyses.

Uploaded by

Chandu Pothina
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
4 views32 pages

9.sampling & Statistical Inference

Sampling is the process of selecting a part of a population to make inferences about the whole, often used in research for efficiency and accuracy. Key concepts include population parameters, sampling errors, and various sampling distributions such as the normal distribution and Student's t-distribution. Understanding these concepts is essential for conducting valid and reliable statistical analyses.

Uploaded by

Chandu Pothina
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
Sampling and Statistical Inference Chapter EE Sampling is defined as the selection of some part of an aggregate or totality on the basis of which a judgement or inference about the aggregate or totality is made. In other words, it is the process of obtaining information about an entire population by examining only a part of it. In most of the research work and surveys, the usual approach happens to be to make generalisations or to draw inferences based on samples about the parameters of population from which the samples are taken. The researcher quite often selects only a few items from the universe for his study purposes. All this is done on the assumption that the sample data will enable him to estimate the population parameters. The items so selected constitute what is technically called a sample, their selection process or technique is called sample design and the survey conducted on the basis of sample is described as sample survey: Sample should be truly representative of population characteristics without any bias so that it may result in valid and reliable conclusions. Sampling is used in practice for a variety of reasons such as: 1. Sampling can save time and money. A sample study is usually less expensive than a census study and produces results at a relatively faster speed. >. Sampling may enable more accurate measurements fora sample study is generally conducted by trained and experienced investigators. : 3, Sampling remains the only way when population contains infinitely many members. 4. Sampling remains the only choice when a test involves the destruction of the item under study. Sampling usually information concerning some [A detailed discussion on the need for sampling and various sampling techniques is presented in Chapter 4. enables to estimate the sampling errors and, thus, assists in obtaining characteristic of the population. (@ scanned with OKEN Scanner Research Methodology 9.1" PARAMETERIAND STATISTIC ‘The collection of all the items about which the information (on one or more chai cristics) is desire is called as population. For example, if we want to measure the average cellphone bill of the people ina particular ci ¥ in a particular month, our population would consist of the specified month’s cellphone bills of all the people in that city. The population or universe can be finite or infinite. The population is said to be finite if it consists of a fixed number of elements so that it is possible to enumerate tin its totality. For tance, the population of a city, the number of workers in a factory are examples of finite populations rhe symbol ‘A” is generally used to indicate how many elements (or items) are there in ease of a finite population. An infinite population is that population in which itis theoretically impossible to observe all the elements. Thus, in an infinite population the number of items is infinite i-e., we cannot have any idea about the total number of items. The number of stars in a sky, possible rolls of a pair of dice are examples of infinite population, One should remember that no truly infinite population of physical objects does actually exist i n spite of the fact that many such populations appear to be very very large. Froma practical consideration, we then use the term infinite population for a population that cannot be enumerated in a reasonable period of time. This way we use the ‘heoretical concept of infinite population as an approximation of very large finite population Any characteristic or measure of population units is known as a parameter. Population mean, population standard deviation and population proportion are commonly studied parameters. These Parameters are denoted by U, o and 1 respectively. Parameters are usually unknown as we do not always study the whole population. Unknown parameter are estimated by studying a subpart of the whole population which is known as a sample. (@ scanned with OKEN Scanner Sampling and Statistical inference gq Population o oe Sampling wi frame a ¢ = G Response error Fig. 9.4 ‘ ‘Sampling error = Frame error + Chance error + Response error ' (If'we add measurement error or the non-sampling error to sampl 8 error, we get total error). ' Sampling errors occur randomly and are equally likely to be in either direction. The magnitude of : the sampling error depends upon the nature of the universe; the more homogeneous the universe, ' smaller the sampling error. Sampling error is inversely related to the size of the sample i.e., sampling i error decreases as the sample size increases and vice-versa. A measure of the random sampling { error can be calculated for a given sample design and size and this measure is often called the precision of the sampling plan, Sampling error is usually worked out as the product of the critical Value at a certain level of significance and the standard error. s opposed to sampling errors, we may have non-sampling errors which may creep in during the process of collecting actual information and such errors occur in all surveys whether census or sample. We have no way to measure non-sampling errors. 3 SAMPLING DISTRIBUTION Distribution of. statistic may not be the same as the distribution of population. We are often concerned | With sampling distribution in sampling analysis. If we take certain number of samples and for each sample compute various statistical measures such as mean, standard deviation, ete., then we can find that each sample may give its own value for the statistic under consideration. All such values of a Particular statistic, say mean, together with their relative frequencies will constitute the sampling distribution of the particular statistic, say mean, Accordingly, we can have sampling distribution of Mean, or the sampling distribution of standard deviation or the sampling distribution of any other Statistical measure. (@ scanned with OKEN Scanner 7 Research Methodology | C150) Here, it is important to discuss normal distribution briefly. the bell shape of the histogram of given data-set as below : ution is represented hy Normal dis Fig, 9.2 In many statistical analyses, it is important that the given data has normal distribution, Normal distribution is denoted by N(11,0) , where 11 is mean and ¢ is standard deviation of the distribution Normal distribution can be used to model any variable taking values ~2°to%, provided the valueson that variable have a bell shaped histogram. Reader may refer to some standard text on statistics to know more about normal distribution. Some commonly used sampling dist uutions are given below: 9.3.1 Sampling Distribution of Mean Sampling distribution of mean refers tothe probability distribution ofall the possible means of random samples ofa given size that we take from a population. If samples ae taken from a normal population 1N (1,9), the sampling distribution of mean would also be normal with mean wu, = w and stand! deviation =o,/n , where 1 isthe mean ofthe population, is the standard deviation of the poplai® and means the number of items in a sample. But when sam; normal (may be positively or negatively skewed), even the sampling distribution of mean tends quite closer to the no sample items is large i.c., more than 30. In case we want pling is from a population which is" N, as per the central limit theorem, the mal distribution, provided the number to reduce the sampling distribution of m=#" x- hing for the samp! olin 7 : « ceveltl is very useful in seve™ to unit normal distribution i.e., (0, 1), we can write the normal vari variate z: distribution of mean. This characteristic of the sampling distri; decision situations. ipling distribution of mean 9.3.2 Sampling Distribution of Proportion Proportion is a measure of attribute. Let us consi : ; - msider that bo yh iy exclusive and collectively exhaustive classes one Ga aad is divided ino twos ‘sessing a particular attribute “ 4 (@ scanned with OKEN Scanner ‘sampling and Statistical Inference @® sing that attribute. For example, people in city could be divided into “Smokers” Let \ = population size, X= number of people out of V possessing a particular attribute, «= X/N = actual proportion of the people possessing the specified attribute. Let a sample is selected from this population with m = sample size,.x = number of people in the sample possessing the specified particular attribute, p =.x/n = sample proportion. Note that, X and P are population parameters, while x and p are sample statistics. Also, p provides an estimate of m. It can be showed that the distribution of x is Binomial(n, m). Using the property of binomial distribution, for sufficiently large sample size n, we have and “Non-smokers’ mm. _Pt__ No, Inm(\—n) fx(l—m)/n Practically, this result is true for n230 or, when nn25 as well as n(I—m)25 CHAPTER 9 9.3.3 Student's t-Distribution When population standard deviation (0) is not known and the sample is of a small size (i.e., $30), we use distribution for the sampling distribution of mean and workout /-variable as: J Ech where sf = > is sample variance. Student’s f-distribution (or ution) is also symmetrical and is very close to the distribution of standard normal variate, 2, except for small values of 7. The variable 1 differs from = in the sense that we use sample standard deviation (s,) in the calculation of ¢, whereas we use standard deviation of population (6) in the calculation of z, There is a different /-distribution for every possible sample size i.e., for different degrees of freedom, The degrees of freedom for a sample of size » isn ~1. As the sample size gets larger, the shape of the /-distribution becomes apporximately equal to the normal distribution. In fact for sample sizes of more than 30, the /-distribution is so close to the normal distribution that we can use the normal to approximate the -distribution. Like normal distribution , the range of -distribution is also (0,0) . The applications of ¢-distribution are discussed in further chapters. 4 9.3.4 Chi-square (x?) Distribution Chi-square distribution is encountered when we deal with collections of values that involve adding up squares. Variances of samples require us to add a collection of squared quantities and thus have distributions that are related to chi-square distribution. Suppose we have a random sample x,, x,, " x, from N(j1,0) population, then the distribution of the statistic) ( =*) is chi-square distribution a with n degree of freedom. Alternatively, the distribution of the statistic >» ( ial is chi-square distribution with (n— /) degree of freedom. Clearly, the range of chi-square distribution (@ scanned with OKEN Scanner

You might also like