0% found this document useful (0 votes)
2 views57 pages

Sampling Distribution

Uploaded by

xamel48611
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views57 pages

Sampling Distribution

Uploaded by

xamel48611
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Sampling Distribution

• It is a distribution of all possible values of a statistic for a given size sample


selected from a population.
• Since a statistic is a random variable that depends only on the observed sample, it
must have a probability distribution.

Definition

• The sampling distribution of a statistic depends on the distribution of the


population, the size of the samples, and the method of choosing the samples.
• The probability distribution of is called the sampling distribution of the
mean.
Types of Sampling Distribution
Developing a Sampling Distribution
Developing a Sampling Distribution
• Suppose if you consider all possible sample of size n, size n here means we are
going to select 2 people with the replacement there is a possibility first
observation may be 18 20 22 24 second observation may be 18 20 22 24 so
possibility is 18 18, 18 20, 18 22, 18 24, 20 18, 20 20, 20 22 and so on.
• So, there are 16 possible samples here we are doing sampling with replacement
that is why it is coming 20 20, 22 22, 24 24
• Now summary measure of this sampling distribution is:
• Different sample of the sample size from the same population yield
different sample means.
• A measure of variability in the mean from the sample to sample is given by
standard error of the mean.
• So standard error is σ/√n,
• The standard error of the mean decreases when the sample sizes increases.
• if the population is normal with the mean u and standard deviation σ the
sampling distribution of X bar is also normally distributed.
• The normal approximation for will generally be good if n ≥ 30, provided the
population distribution is not terribly skewed.

• If n < 30, the approximation is good only if the population is not too different
from a normal distribution and, as stated above, if the population is known to be
normal, the sampling distribution of will follow a normal distribution
exactly, no matter how small the size of the samples.
• The sample size n = 30 is a guideline to use for the Central Limit Theorem.
• However, as the statement of the theorem implies, the presumption of normality on
the distribution of becomes more accurate as n grows larger.
• In fact, Figure 8.1 illustrates how the theorem works. It shows how the distribution
of becomes closer to normal as n grows larger, beginning with the clearly
nonsymmetric distribution of an individual observation (n = 1). It also illustrates that
the mean of remains μ for any sample size and the variance of gets
smaller as n increases.
• Example :

An electrical firm manufactures light bulbs that have a length of life that is
approximately normally distributed, with mean equal to 800 hours and a standard
deviation of 40 hours. Find the probability that a random sample of 16 bulbs will have
an average life of less than 775 hours.
Solution :
• The sampling distribution of will be approximately normal, with and The desired probability is
given by the area of the shaded region in Figure 8.2.
Sample Variance
• Let X1,X2,........Xn be the random sample from a population. The
sample variance is

• The sample variance is different for different random samples from


the same population.
Sampling Distribution of Sample Variances
• The sampling distribution of s2 has mean σ2.

E(s2)= σ2
If the population distribution is normal then
(n-1) s2/ σ2
Has a chi-square distribution with n-1 degrees of freedom.
Degrees of freedom(df)
• Idea: Number of observations that are free to vary after sample mean
has been calculated.
• Example suppose the mean of 3 numbers is 8.0
Let X1=7 If the mean of these values is 8.0, then
• X2=8 x3 must be 9
What is x3=? (i.e. X3 is not free to vary

Here n=3 , so degrees of freedom is n-1=2


(2 values can be any numbers, but third is not free to vary for a given mean.
Chi-square distribution Example-
• A commercial freezer must hold a selected temperature with little
variation. Specification calls for a standard deviation of no more
than 4 degrees(a variance of 16 degrees2)
• A sample of 14 freezers is to be tested.
• What is the upper limit (K) for the sample variance such that the
probability of exceeding this limit, given that the population standard
deviation is 4, is less than 0.05?
Finding chi-square value
• 213 =22.36 ( =0.05 and 14-1 = 13 d.f.)
• So, P(s2>K)= P[ (n-1) s2 / σ2 > 213 ]
• Or (n-1)K/16=22.36
• So K=(22.36)16/13
• K=27.52
• i.e. The result is give the sample variance from the sample size of 14
is greater than 27.52 there is a strong evidence to suggest that the
population variance exceeds 16
Confidence Interval Estimation
• Confidence Interval:
• Confidence interval for the Population Mean, µ
• 1. when population variance σ2 is known
• 2. when population variance σ2 is unknown
Estimator:
• Definition:
• A estimator of population parameter is
• - a random variable that depends on sample information
• -whose value provides an approximation to this unknown parameter
• A specific value of that random variable is called an estimate.
• E.g. Xbar µ
• s2 σ 2
Point and Interval Estimates
• A point estimate is a single number,
• A confidence interval provides additional information about
variability

Upper confidence
Lower confidence Point Estimate limit
limit

Width of confidence interval


Unbiasedness
• A point estimator θ hat is said to be unbiased estimator of the
parameter θ if the expected value, or mean of the sampling
distribution of θ hat is θ.
• E(θhat)=θ
• Example:
• The sample mean xbar is unbiased estimator of µ
• The sample variance s2 is an unbiased estimator of σ2
• The sample proportion p hat is an unbiased estimator of P
Unbiasedness:
Bias:
• Let θ hat be an estimator of θ
• The bias in θ hat is defined as the difference between its mean and θ
• So,
• Bias(θ hat)= E(θ hat)- θ
• The bias of unbiased estimator is 0.
Most Efficient Estimator:
• Suppose there are several unbiased estimators of θ
• The most efficient estimator of θ is the unbiased estimator with the
smallest variance.
• Let θ1 hat and θ2hat be the two unbiased estimators of θ based on
the same number of sample observations. Then,
• Θ1 hat is said to be most efficient than θ2 hat if
• Var(θ1 hat) < var(θ2 hat)
• The relative efficiency of θ1 hat with respect to θ2 hat is the ratio of
their variances:
• Relative efficiency= Var(θ2 hat)/var(θ1 hat)
Confidence Interval:
• How much uncertainty is associated with the point estimate of the
population parameter?
• E.g. The temperature is 35 degree how much uncertainty is
associated with that point estimate.
• That uncertainty is expressed with the help of confidence interval.
• The interval estimate provides more information about a population
characteristic than does a point estimate.
• Such interval estimates are called confidence interval.
Confidence Interval Estimate
• An interval gives a range values:
• - Takes into consideration variation in sample statistics from sample
to sample
• -Based on observation from 1 sample
• -Gives information about closeness to unknown population
parameters
• -stated in terms of level of confidence
• Can never be 100% confident
Confidence Interval and Confidence Level
• If P(a< θ <b) = 1- then the interval from a to b is called a 100(1- )%
confidence interval of θ.
• The quantity (1- ) is called the confidence level of the interval (
between 0 and 1).
• In repeated samples of population, the true value of the parameter θ
would be contained in 100(1- )% of intervals calculated this way.
• The confidence interval calculated in this manner is written as a< θ <b
with 100(1- )% confidence.
Estimation Process:
Confidence Level:
• Suppose confidence level is 95%
• Also written as (1- ) =0.95
• A relative frequency interpretation:
• - From repeated samples, 95% of all the confidence intervals that can
be constructed will contain the unknown true parameter
• A specific interval either will contain or will not contain the true
parameter.
• General formula:
• Point Estimate+-(Reliability Factor)(Standard error)
• xbar +- Z σ / sqrt(n)
Confidence Interval
2
Confidence interval of mu - σ is unknown
2
Confidence interval of mu - σ is unknown
• If the population standard deviation σ is unknown, we can substitute
the sample standard deviation s
• This introduces extra uncertainty, since s is variable from sample to
sample
• So we use the t-distribution instead of normal distribution.
Student’s t distribution
• The t is a family distribution
• The t value depends on degrees of freedom (d.f.)
• Number of observations that are free to vary after sample mean has
been calculated
• D.f.=n-1
Example
Example 2
• Suppose you do a study of acupuncture to determine how effective it
is in relieving pain. You measure sensory rates for 15 subjects with
the results given. Use the sample data to construct a 95% confidence
interval for the mean sensory rate for the population (assumed
normal) from which you took the data.
• 8.6, 9.4, 7.9, 6.8, 8.3, 7.3, 9.2, 9.6, 8.7, 11.4, 10.3, 5.4, 8.1, 5.5, 6.9
• Solution:
• Xbar=8.2267
• s = 1.6722
• n = 15
• df = 15 – 1 = 14
• Confidence Level so α = 1 – CL = 1 – 0.95 = 0.05
• The area to the right of to.o25 is 0.025, and the area to the left of to.o25 is 1 –
0.025 = 0.975
• tα/2=to.o25=2.14
• Now the confidence interval is
• Xbar-t(n-1), α/2 s/sqrt(n)< µ <Xbar+t(n-1), α/2 s/sqrt(n)
• 7.3028< µ <9.1506
Example3
• You do a study of hypnotherapy to determine how effective it is in
increasing the number of hourse of sleep subjects get each night. You
measure hours of sleep for 12 subjects with the following results.
Construct a 95% confidence interval for the mean number of hours
slept for the population (assumed normal) from which you took the
data.
• 8.2; 9.1; 7.7; 8.6; 6.9; 11.2; 10.1; 9.9; 8.9; 9.2; 7.5; 10.5
Example 4
• The Human Toxome Project (HTP) is working to understand the scope of industrial
pollution in the human body. Industrial chemicals may enter the body through pollution
or as ingredients in consumer products. In October 2008, the scientists at HTP tested
cord blood samples for 20 newborn infants in the United States. The cord blood of the
“In utero/newborn” group was tested for 430 industrial compounds, pollutants, and
other chemicals, including chemicals linked to brain and nervous system toxicity,
immune system toxicity, and reproductive toxicity, and fertility problems. There are
health concerns about the effects of some chemicals on the brain and nervous system.
This table shows how many of the targeted chemicals were found in each infant’s cord
blood.
• 79 145 147 160 116 100 159 151 156 126 137 83 156 94 121 144 123
114 139 99
• Use this sample data to construct a 90% confidence interval for the mean number of
targeted industrial chemicals to be found in an in infant’s blood.
• Solution A:
• From the sample, you can calculate xbar=127.45
• and s = 25.965. There are 20 infants in the sample, so n = 20, and df =
20 – 1 = 19.
• You are asked to calculate a 90% confidence interval: CL = 0.90,
• so α= 1-CL = 1-0.90 = 0.10
• α/2= 0.05
• The 90% confidence interval is (117.41, 137.49).
Confidence Interval- Population Proportion
• An interval estimate for population proportion (P) can be calculated
by adding an allowance for uncertainty to the sample proportion (p
hat).
• The distribution of sample proportion is approximately normal if the
sample size is large, with standard deviation.
• σp

We will estimate this with sample data:


Confidence Interval Endpoints
• Upper and lower confidence limits for the population proportion are
calculated with the formula

Where ,
Z α/2 –is the standard normal value for the level of confidence desired

P hat- is the sample proportion


n is the sample size
But there is condition
nP(1-P)> 5 , then only it can be approximated to normal distribution.
Example:
• A random sample of 100 people shows that 25 are left handed.
• Form a 95% confidence interval for the true population of
left-handers.
• Solution-
• n=100
• P-hat= 25/100
• Z value for 95% is= 1.96
• Interpretation:
• We are 95% confident that the true percentage of left-handers in the
population is between
• 16.51% and 33.49%
• Although the interval from 0.1651 and 0.3349 may or may not
contain the true proportion, 95% of intervals formed from samples of
size 100 in this manner will contain the true proportion.
Confidence Interval for Population Variance
• The confidence interval is based on sample variance s2
• Assumed: The population is normally distributed
• The random variable
• 2n-1 = (n-1) s2/ σ2
• Follows a chi-square distribution with (n-1) degrees of freedom
• The (1- )% confidence interval for the population variance is:
• (n-1) s2/ 2n-1, /2< σ2< (n-1) s2/ 2n-1,1- /2
Example-
• You are testing the speed of a batch of computer processors. You
collect the following data(in Mhz)
• Sample size- 17
• Sample mean-3004
• Sample std dev- 74
• Assume the population is normal. Determine the 95% confidence
interval for σ2x
Solution-
• n=17 so the chi-square distribution has (n-1) degrees of freedom
• =0.05, so use the chi-square values with the area 0.025 in each tail
• 2n-1, /2 = 216,0.025 = 28.85
• 2n-1,1- /2 = 216,0.975 = 6.91
• Hence, the confidence interval is,
• (n-1) s2/ 2n-1, /2< σ2< (n-1) s2/ 2n-1,1- /2
• (17-1)(74)2/28.85 < σ2< (17-1) (74)2/ 6.91
• 3037< σ2<12683
• Converting to standard deviation, we are 95% confident that the population
standard deviation of CPU speed is between 55.1 and 112.6 Mhz.

You might also like