CHAPER TWO
STATISTICAL INFERENCE MANUAL DRAFT
COMPILED BY DR MASOUD SALEH (unedited)
SCHOOL OF BUSSINESS (SUZA)
INTRODUCTION TO STATISTICAL INFERENCE
1. Variables
A characteristic that varies from one person or thing to another
a. Qualitative variable is a non-numerically valued variable while quantitative variable is a
numerically valued variable.
b. Discrete variable is a quantitative variable whose possible values can be listed while
continuous variable is a quantitative variable whose possible values form some interval
of numbers.
2. TYPES OF DATA
Data are values of a variable
i. Qualitative data are values of a qualitative variable while quantitative data are
values of quantitative variable
ii. Discrete data are values of discrete variables/quantitative data that can be listed
(take the integers number only) while continuous data are values of continuous
variable/quantitative data whose possible values form some interval of numbers
(not restricted to integers, can take a decimal points)
iii. Primary and Secondary data
Primary data is the data that are used for the specific purpose for which they were
collected main sources is senses or samples while secondary data are data that are
being used for some purpose other than that for which they were originally
collected. Or
Primary data is new data collected by an organization/person its/himself for a
specific purpose while Secondary data is existing data collected by other/another
organizations/people or for other purpose
iv. Data can be classified on how easy it can be measured.
a. Nominal data is the kind of data that we really cannot quantify with any
meaningful units. For example a person is an economist, or a country operates a
market economy, or a cake has a bream in it a car is black, here we measure by
number of observations, also it can be called categorical or descriptive data
b. Ordinal data is more step more quantitative than nominal, here we can rank the
categories of observations into some meaningful orders, and for example we
describe sweaters as large, medium or small.
c. Cardinal data has some attribute tha can be directly measured, for example,
weight, time, temperature this is relevant to quantitative data.
v.
3. Sampling
There are two ways /method of acquiring information (data); Census and Sampling
In census, data are collected for each and every unit belonging to the population while in
sampling method a few units from the whole population are selected and the results
obtained generalized to the whole population.
I. Population/Universe is the total numbers of items in a specific field of enquiry, for
example, total number of lecturers at SUZA , numbers of SME in Zanzibar, number
of first years students who study Quantitative Method (QM) at the school of
Business, SUZA. While a sample is the part/portion of the population for example,
fraction of lecturers from SUZA, some of SME in Zanzibar, some of first year students
who study QM.
Advantage of sampling methods
- It is cheaper to collect data
- Saves a lot of time
- A good quality of labour with better supervision can be provided
- We can have more detailed information during investigation
ii. Sampling design
Is a definite statistical plan concerned with all principal steps taken in the selection
of a sample and the estimation procedure.
Sampling frame is a complete list of all items of a population for example the names
of all lecturers at SUZA
iii. Simple random sampling. A sampling procedures for which each possible sample of a
given size is equally likely to be the one obtained while a simple random sample is a
sample obtained by simple random sampling
a. Simple random sampling with replacement whereby a number of population can be
selected more than once while simple random sampling without replacement a
member of the population can be selected at most once, unless otherwise explained
we assume the simple random sampling is done without replacement.
b.
How
By picking slips of paper out of a box
Using random number tables
Using random number generator by statistical software packages or graphing calculators
c. Simple random sampling is the most natural and easily understood method of
probability sampling, however it has drawback;
It may fail to provide sufficient coverage when information about subpopulations is
required and may be impractical when members of the population are widely scattered
geographically
d. Systematic Random Number
Easier to execute than the simple random sampling and usually provides comparable results
Steps
i. Divide the population size by the sample size and round the result down to the
nearest whole umber m
ii. Use a random number table (or similar device) to obtain a number, k, between
1 and m
iii. Select for the sample those members of the population that are numbered k, k
+ m, k + 2m….
Example: Dr Issa wanted a sample of 20 of the 300 students who studied QM at
SOB.
We divide the population size by the sample size and round the answer down
to the nearest whole number. 300/20 = 15. Then we select a number at
random between 1 and 15, by using say, a table of random numbers. Suppose
that we do so, and obtain the number 7. Then, we list every 15th number, start
at 7, until we have 20 numbers.
Exrc: A production line makes 5,000 units day. How can the quality control
Department take a systematic sample of 2% of these?
Stratified Sampling
- Divide the population into subpopulations (strata)
- From each stratum, obtain a simple random sample of size proportional to the size
of the stratum; that is, the sample size for a stratum equals the total sample size
times the stratum size divided by the population size
- Use all the members obtained in step 2 as the sample.
*stratification may be age, occupation, income group etc.
Example1
A bakery make three different types of loaf. Assume that the bakery’s output is
50% large loaves, 40% small loaf and 10% cottage loaves.( The different loaves
divide the population into three strata) If a sample of 50 loaves is required it
should contain:
0.5*50 = 25 large loaves,
0.4*50 = 20 small loaves and
0.1*50 = 5 cottage loaves
*within this constraints, selection should be done in a random basis.
Excr
The strata consisting of three income groups, upper, middle and low, comprise
respectively 10%, 70% and 10% of the town. Then how would you select a
stratified sample of 60?
Multistage Sampling
Where a population is spread over a relatively wide geographical area, random
sampling will almost certainly necessitate traveling to all part of the area and thus
could be excessively expensive. Therefore multistage sampling will be needed;
which required the following steps
- Splitting the area up into a number of regions;
- Randomly selecting a small number of the regions;
- Restricting sub-samples to these regions alone, with the size of each – subsample
proportional to the size of the area
- The above procedure can repeated for sub- regions within regions…. and so on
- Once the final regions (or sub-regions etc.) have been selected, the final sampling
technique could be (simple or stratified) random or systematic, depending on the
existence or otherwise of the sample frame.
Quota sampling
When investigators are told to interview all the people they meet up to a certain
quota. Most often used in market research where the data is collected by
enumerators armed with questionnaires. Such a quota is nearly always divided up
into different types of people with sub quotas for each type.
For example, out of a quota of 400, the enumerator may be told to interview 250
working wives, 100 non-working wives and 50 unmarried women,.
*is a type of judgement sampling.
e. Cluster sampling
i. Divide the population into groups (clusters)
ii. Obtain a simple random sample of the clusters
iii. Use all the members of the clusters obtained in step 2 as the sample.
Can save time and money however if members of clusters are more homogenous
than the members of the entire population, can cause problems (especially for the
small clusters)
Examples constructing a swimming pool
4. Sampling Size
When you adopt a sample techniques it is very important to decide on the size of your
sample. There are some suggestions on how to find your sample size, however the most
important to keep in your mind is that
- The size of the sample should increase as the variations in the individuals outcome
increases
- The greater the degree of accuracy desired, the larger should be the sample size.
5. Sampling and non-Sampling error
i. Sampling error is the difference between the values of sample statistic and the true
value of the corresponding population parameters.
For example if 𝑋̅ is the mean obtained from a sample of size n and µ is the
corresponding population parameters, then
𝑠𝑎𝑚𝑝𝑙𝑖𝑛𝑔 𝑒𝑟𝑟𝑜𝑟 = 𝑋̅ - µ
As the sample size increases, the sampling error is reduced, and in a complete enumeration
(census) there is no sampling error since ̅̅̅
𝑋 = µ. Sampling error is zero.
ii. Non sampling occurs during the process f gathering data regardless of whether a
sample or a complete census is taken.
6. Parameters and Statistics
i. Any statistical measure based on all units in population is called Parameters. E.g.
population mean, population S.D., proportion defective in the whole lot, etc.
ii. Statistical measures computed from sample observations alone are termed as
statistic, e.g. sample mean, sample S.D., the proportion of defectives observed in the
sample.
iii.
Population parameter Unbiased estimator
Population mean µ Sample mean , 𝑋̅ =∑x/n
2
Population Variance, 𝜎 Sample variance. S2= ∑(x-𝑥̅ )2 /n-1
Population standard deviation, 𝜎 Sample standard deviation, S=√ ∑(x-𝑥̅ )2 /n-1
Population proportion, P Sample proportion, ṗ = x/n
X is the number of observation under the
category of interest
Total value in a population of size N, N µ N 𝑋̅
Total number of observations falling Nṗ
under the category of interest in a
population, NP
7. Sampling Distribution Is the probability distribution of the values of a statistic such as a
mean, a standard deviation, a proportion, etc. computed from all possible samples of the
same size, which might be selected with or without replacement from a population. While
sample distribution is the distribution of individual values of a single sample
The most frequently used in statistical inference are the binomial, the normal, the t-
distribution, the chi- square distribution and the F distribution.
8. Statistical Inference
The process of drawing inferences about a population on the basis of information contained
in a sample taken from the population. This can be divided into estimation of parameters
and testing of hypothesis.
i. Statistical Estimation is the procedure of using statistic to estimate a population
parameter.
Statistical estimation is divided into two main categories; Point estimation and
Interval estimation.
A statistic used to estimate a parameter is called an estimator and the value taken
by the estimator is called an estimate.
ii. Criteria for good estimator
- Efficiency: is the one which has a relatively smaller variance or standard deviation
- Consistency: is one whose standard deviation decreases with the increase of the
size of the sample
Or
It is an estimator that yields values more closely approaching the population
parameters as the sample size increases.
- Sufficiency: is the one which utilizes all the information in the sample to arrive to
the estimate.
- Unbiasedness: is the one which its expected value (mean) is exactly equal to the
parameter being estimated.
Once we established that a certain estimator is unbiased, we conclude that it is
also efficient, consistent, sufficient, and then a good estimator.
II. Point Estimation
Is an estimate of a population parameter given by a single number. The criteria we
used to base our argument is that of unbiasedness
Example
A marketing research analyst collects data for a random sample of 100 customers of
the 500 who purchased a particular item under promotion. The 100 people in the
sample spent an average (mean) of TZS 27,000 in the store with a std deviation of
TZS 8000, and 70% of the costumers in the sample made at least one other purchase
in addition to the term under promotion. Based on this results, estimate the value of
the following parameters.
a. Mean purchase amount by all 500 customers who purchased the item that is
under promotion
b. The std deviation of the distribution of purchase amount by the 500 customers
c. The total amount of purchases made by the 500 customers
d. The numbers of customers out of 500 who made at least one other purchase in
addition to the product under promotion
Central Limit Theorem
If we select a large number of simple random samples, say from any population
and determine the mean of each sample, the distribution of these sample
means will trend to be described by the normal probability distribution with a
mean µ and variance 𝜎2 /n or
The sampling distribution of sample means approaches to a normal distribution,
irrespective of the distribution of population from where sample is taken, and
approximation to normal distribution becomes increasingly close with increase
in sample size.
iii. Interval estimation
An estimation of population parameter given by two numbers between which the
parameter may be considered to lie. This indicate the precision or accuracy of an
estimate, and are, therefore, preferable to point estimates.
The interval estimate or a “confidence interval” consists of an upper confidence limit
and lower confidence limit and we assign a probability that this interval contain the
true population value.
Procedures for interval estimation
- The particular statistic say the mean of the sample or std deviation of the sample is
determined
- The confidence value is decided, i.e. 95%, 99%, etc
- The standard error of the particular statistic is calculated
- Finally, we state with a known degree of confidence that the parameters lies in this
interval.
a. Confidence Limits and Interval
Confidence limits are the outer limits to confidence interval. This is a range of
values within which we may be confident that the population mean (or
parameters being considered) does lie.
The interval between confidence limits is called the confidence interval.
A normal distribution has the following characteristics:
I. Sample mean ± 1.96 𝜎 includes 95% of the population
II. Sample mean ± 2.58 𝜎 includes 99% of the population
These characteristics can be used to calculate confidence limits for the
population mean when we have established the sample mean and the
standard error.
b. Estimation of Population Mean
To estimate a population mean, the following procedure is followed:
i. Take a random sample of n items. Where n represents sample size and
it is not less than 30
ii. Compute the sample mean (𝑋̅ ) and std deviation ( 𝜎x )
iii. Compute the std error of the mean by using the following formula
σx
σ𝑥̅ = 𝑛
√
Where,
σ𝑥̅ = 𝑠𝑡𝑑 𝑒𝑟𝑟𝑜𝑟 𝑜𝑓𝑡ℎ𝑒 𝑚𝑒𝑎𝑛
𝜎x = std deviation of the sample
n = sample size
iv. Choose a confidence level e.g 90% or 99%
v. Estimate the population mean as under
Population mean ( µ ) =𝑋̅ ± 𝑎𝑝𝑝𝑟𝑜𝑝𝑟𝑖𝑎𝑡𝑒 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 σ𝑥̅
Appropriate number means confidence level.
At 95% confidence level, it is 1.96 and for 99% level confidence this
number is 2.58.
Example 1.
The quality department of a wire manufacturing company periodically
selects a sample of wire specimens in order to test for breaking
strength. Past experience has shown that the breaking strengths of a
certain type of wire are normally distributed with std deviation of
200kg. a random sample of 64 specimens gave a mean of 6,200kg. Find
out the population mean at 95% level of confidence.
Solution.
σx
Population mean ( µ ) =𝑋̅ ± 𝑛 where
√
𝑋̅ = 6200
σx 200
σ𝑥̅ = = =25
√𝑛 √64
Hence, population mean = 6200 ± 1.96 (25)
= 6200 ± 49
= 6151 to 6249
Therefore at 95% level of confidence, population mean will be in
between 6151 and 6249
vi. Finite Population Factor
If large sample is drawn from infinitely large populations then std error
of mean reduces and it increases reliability of the sample mean as an
estimate of population mean. But if a given population is relatively of
small size and sample size is more than 5% of the population then the
std error should be adjusted by multiplying it by finite population
correction factor.
𝑁−𝑛
Finite population correction factor = √ 𝑁−1
Where N and n are population and sample size, respectively.
Example2
A manager wants an estimate of sales of salesmen in his company. A
random sample 100 out of 500 salesmen is selected and average sales
are found to be TZS 75,000. If a sample std deviation is TZS 15,000
then find out the population mean at 99% level of confidence.
Solution
We have N = 500, n = 100, 𝑋̅ = 𝑇𝑍𝑆75,000, 𝜎x = TZS 15,000
σx 𝑁−𝑛
Then, std error of mean = √
√𝑛 𝑁−1
15000 500−100
= √
√100 500−1
= 1500(0.895)
= 1342.50
At 99% level of confidence, population mean = 𝑋̅ ± 2.58σ𝑥̅
= 75000 ± 2.58(1342.50)
=71536 to 78464
c. Estimation of Population Proportion
Sometimes, the statistical information is given in the form of proportions
instead of actual measures. A proportion represents an attribute of a
population rather than the value of a variable. The procedure for
estimating population proportions from sample proportions is similar to
that for estimating a population mean. However, the relevant formula for
the std error of a sample proportion is as under:
𝑝𝑞
Std error of sample proportion (𝜎p) = √ 𝑛
Where
P = population proportion
Q = 1-p
n = sample size
Population proportion is estimated as under:
Population proportion = sample proportion ± appropriate number of
𝜎p
Example 3
In a sample of 800 candidates, 560 were male. Estimate the
population proportion at 95% confidence level.
Solution
Sample proportion p = 560/800 = 0.70, n = 800
q = 1 – p = 1- 0.70 = 0.30
Population proportion = p ± 1.96𝜎p
0.70(0.30)
= 0.70 ± 1.96√ 800
= 0.70 ± 0.03
= 0.67 to 0.73
= between 67% to 73%
Example 4.
A sample of 600 accounts was taken to test the accuracy of posting
and balancing of accounts wherein 45 mistakes were found. Find out
the population proportion. Use 99% level of confidence.
Answer, between 4.7% to 10%, hint n = 45/600 = 0.075
d. Determination of proper sample size
In most of the practical situation the sample size is not known. Instead, one
may prefer to specify the width of the interval and use this information to
solve for n.
i. Sample size for estimating a population mean
σ
The formula for confidence interval =𝑋̅ ± z 𝑛 or 𝑋̅ ± E
√
σ
Where E = z
√𝑛
Which is the maximum allowable error, i.e difference between the
population and the sample mean. Or
σ2
n = z2𝐸2
Where Z = Appropriate number i.e. 1.96 for 95% and 2.58 for 99% level
of confidence.
The value of of Z and E must be given and the value of the population
may be actual or estimated
Example 5
A cigarette manufacturer wishes to use a random sample to estimate
the average nicotine content. The sampling error should not be more
than one milligrams above or below the true mean, with a 99 percent
confident efficient. The population std deviation is 4 milligrams, what
sample size should the company used in order to satisfy the
requirement?
Solution
E = 1, Z = 2.58, 𝜎 = 4, n =?
Answ 106.50 or 107
ii. Sample Size for Estimating a population proportion
The confidence interval formula for proportion is given by
𝑝𝑞
P ± Z√ 𝑛
Using E to represent the maximum allowable sampling error, we may
write the above equation as
P ± E
Where E is the difference between the sample proportion and the
population proportion.
𝑝𝑞
Hence, E = Z√
𝑛
Solving for n, we get
pq
n = z2𝐸2
Where the value of Z and E are predetermined. The value of
population proportion p may be actual or estimated from the past
experience.
Example 6
A firm wishes to estimate with a maximum allowable error of 0.05 and
a 98% level of confidence, the proportion of consumers who prefer its
product. How large a sample will be required in order to make such
an estimate if the preliminary sales report indicate that 25 percent of
all consumers prefer the firm’s product?
Solution
E = 0.05, Z = 2.33, p = 0.25, then q = 1 – 0.25 = 0.75 n = ?
The required sample size n = 407