Chapter 4
Chapter 4
Chapter 4: ESTIMATION
1 ESTIMATION (STATISTICAL METHODOLOGY)
In chapter 3, we defined a random variable X and had a new interpretation of
data ----- the actual values of the random variable. Recall that if the distribution
of X is known, then we can answer any probability problem of X.
If the distribution is unknown, then we need to use STATISTICS to make an
inference about the underlying distribution or the parameter of our interest.
~1~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
~2~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝑿 ̅
𝑿
To get a better understanding of their relationship, we need to learn the following
ninjutsu!
~3~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Suppose the original rv X can use the multiple shadow clone and create n
copies of itself. Note that the copies look like the original identically
(Correspondingly, we would say that the copied rvs have the same
distribution as the original rv X) and they are independent (just as all
shadow clones can act independently!). Thus, the reasonable notations of
those n copies of X are 𝑿𝟏 , 𝑿𝟐 , … , 𝑿𝒏 . ALL OF THEM ARE RV.
Then, the n data of X now can be interpreted as the data of these n copies.
In other words, 𝒙𝒊 can be interpreted as the single actual value of 𝑿𝒊 ,
where 𝑖 = 1,2, … , 𝑛.
̅:
According to the n copies of X, we have the following definition of 𝑿
𝒏
𝟏 𝟏
̅
𝑿 = ∑ 𝑿𝒊 = (𝑿𝟏 + 𝑿𝟐 + ⋯ + 𝑿𝒏 )
𝒏 𝒏
𝒊=𝟏
~4~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝑥1 𝑥3 𝜇𝑋 𝑥2 𝑥5 𝑥4
Therefore, when we ask for the goodness of the sample mean or want to see if
̅
the sample mean is close to 𝜇𝑋 , we should consider its random counterpart --- 𝑿
(rv, methodology).
ACCURACY (準確度) --- HOW CLOSE THE ESTIMATOR IS TO THE UNKNOWN PARAMETER.
If the estimator is more centered on the parameter to be estimated, then we would say
that it is more accurate. Recall that the population mean is one measure of the central
tendency of the rv. Thus, we can say that the estimator has high accuracy if its
population mean is equal to the unknown parameter being estimated. In statistics, such
an estimator is said to be UNBIASED (無偏的).
~5~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Thus, we now can use this criterion ---- unbiasedness (無偏性) --- to find a good
estimator (good estimation method) of the unknown parameter being estimated,
or to determine whether or not a given statistic(s) is a good estimator of the
parameter being estimated.
EXAMPLE
Consider a discrete random variable X with pmf given by
The distribution
of X is
UNKNOWN.
It is easy to check that E(𝑋) = 10 − 40𝑝 and Var(𝑋) = 200𝑝 – 1600𝑝2 , and
both of them are unknown. So, in this case, are the common estimators,
̅ , 𝑺𝟐𝒏 , 𝑺𝟐𝒏−𝟏 , unbiased for 𝜇𝑋 and
𝑿 𝜎𝑋2 , respectively?
Note that all of these three estimators in this example are discrete.
i) If we can find their respective sampling distributions (which must be based on
the unknown parameter p in this case) --- the distribution of estimator, then we
can get the population means of these common estimators and see if they are
~6~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
unbiased. Alternatively, if we know more properties of rv, then we can also check
the unbiasedness without finding the sampling distribution.
ii) If we do not know how to find the sampling distribution or we only know less
properties of rv to check the unbiasedness, then we can do a simulation to guess
the answer with a user-specified value of the unknown parameter being
estimated by using the average of B sample means 𝒙 ̅𝟏 , 𝒙
̅𝟐 , … , 𝒙
̅𝑩 to be an
approximation of 𝐸(𝑋̅) (Similar manner for 𝐸(𝑆𝑛2 ) and 𝐸(𝑆𝑛−1 2
)).
Before we check the unbiasedness in case (i), we first discuss the simulation with
R for case (ii). To simplify our explanation, we only consider the sample mean and
sample variance of two copies, 𝑋1 and 𝑋2 , of X. Thus, we have
𝑋1 + 𝑋2
𝑋̅ =
2
𝑛
1 (𝑋1 − 𝑋2 )2
2
𝑆𝑛−1 = ∑(𝑋𝑖 − 𝑋̅) =
2
𝑛−1 2
𝑖=1
𝑛
( )2
1 𝑋1 − 𝑋2
𝑆𝑛2 = ∑(𝑋𝑖 − 𝑋̅)2 =
𝑛 4
𝑖=1
Note that we can use a random experiment of selecting a ball from a box
containing #2, #4, #6, #8 and #10 balls with their respective probabilities and a
user-specified p (pick a real number between 0 and 0.1 exclusively) TWO times
with replacement to generate the actual values 𝒙𝟏 and 𝒙𝟐 of X.
~7~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
R codes:
~8~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
p 𝑥̿ (𝜇𝑋 ) ̅̅̅̅̅̅
2
𝑠𝑛−1 (𝜎𝑋2 ) ̅̅̅
𝑠𝑛2 (𝜎𝑋2 )
Note that roughly speaking, if B is large enough, then 𝑥̿ would be very close to
𝐸(𝑋̅), and similarly, ̅̅̅̅̅̅
2
𝑠𝑛−1 and ̅̅̅ 2
𝑠𝑛2 would be very close to 𝐸(𝑆𝑛−1 ) and 𝐸(𝑆𝑛2 ),
respectively.
respectively, but 𝑺𝟐 2
𝒏 is not unbiased for 𝜎𝑋 .
In the following, I will only show you how to determine the sampling distribution
̅ and check its unbiasedness for this example.
of 𝑿
~9~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Our next step is to use the sampling distribution this discrete rv 𝑋̅ to find its
population mean. Thus, we have 𝜒 = {2, 3, 4, … , 10} and
~ 10 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
EXAMPLE
̅ without
In this example, I would show you how to check the unbiasedness of 𝑿
finding its sampling distribution.
At this time we can use a random experiment of flipping a coin with a user-
specified p (pick a real number between 0 and 1 exclusively) n=10 times to
generate 10 actual values of X and get an actual value of sample mean 𝑿̅.
Similar to the previous example, I used R to generate B = 5000 sample means with
p = 0.72 by
~ 11 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
respectively, but 𝑺𝟐 2
𝒏 is not unbiased for 𝜎𝑋 . To get the correct answer, we prove
̅.
it in the following without finding the sampling distribution of 𝑿
Recall that we have 10 independent shadow clones 𝑋1 , 𝑋2 , … , 𝑋10 of X. Thus, their
common distribution is also Binominal (1, p). Theoretically, we can show that their
sum (also a rv) would have Binominal (10, p). In other words,
(Theorem)
Consider a random sample (independent and identically distributed, in short iid)
of X with sample size n, say 𝑋1 , 𝑋2 , … , 𝑋𝑛 . Then,
𝑛
1
̅
𝐸 (𝑋) = 𝐸 ( ∑ 𝑋𝑖 ) = 𝜇𝑋 .
𝑛
𝑖=1
𝟐
Fifth Question: Is 𝑺𝒏−𝟏 always an unbiased estimator for 𝜎𝑋2 ?
(Theorem)
Consider a random sample (independent and identically distributed, in short iid)
of X with sample size n, say 𝑋1 , 𝑋2 , … , 𝑋𝑛 . Then,
𝑛
1
2 )
𝐸 (𝑆𝑛−1 = 𝐸( ∑(𝑋𝑖 − 𝑋̅)2 ) = 𝜎𝑋2 .
𝑛−1
𝑖=1
~ 13 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Proof (optional)
X X X
X
X X X X
X
2
𝜎
In the proof, we use the result that 𝑉𝑎𝑟(𝑋̅) = 𝑋 (which will be proved later).
𝑛
2
Note that if 𝑆𝑛−1 is unbiased for 𝜎𝑋2 , then we can ensure that 𝑆𝑛2 is NOT unbiased
for 𝜎𝑋2 . (WHY?)
Here I would like to emphasize that in practice we would have more than one
2
estimator for one unknown parameter, like 𝑆𝑛−1 and 𝑆𝑛2 for 𝜎𝑋2 . Unbiasedness can
2
be used to find a good estimator, like 𝑆𝑛−1 for 𝜎𝑋2 . However, it is still possible for
us to have more than one unbiased estimator, like 𝑋1 , 𝑋5 , 𝑋̅ , etc, for 𝜇𝑋 .
Thus, we would have the following natural question:
~ 14 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
PRECISION (精確度) --- HOW CLOSE THE ESTIMATES ARE TO EACH OTHER
If the estimator is more concentrated on a particular value, then we would say that it
has a higher precision on this value. Recall that the population variance is one measure
of the spread of the rv. Note that more spread means lower precision, and less spread
means higher precision. Thus, we can say that the estimator has higher precision if its
population variance is smaller. In statistics, we would say that it is MORE EFFICIENT.
Thus, variance can be used for the comparison of different UNBIASED estimators.
Intuitively, we would expect that 𝑋̅ has a smaller variance than 𝑋1 and 𝑋5 . Then,
how do we verify our intuition? Indeed, we have the following theorem:
(Theorem)
Consider a random sample (independent and identically distributed, in short iid)
of X with sample size n, say 𝑋1 , 𝑋2 , … , 𝑋𝑛 . Then,
𝜎𝑋2
𝑉𝑎𝑟(𝑋̅) = .
𝑛
Note that the variance of 𝑋̅ is proportional to 1/n. This result thus suggests that
we should use data as many as possible so that we can get a more precise sample
mean!!
~ 15 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Suppose there are four different archers, each with varying degree of ability. The bull's-eye in
the target represents the true value of the unknown parameter being estimated.
~ 16 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Proof (optional):
In the following, let’s simply use 𝜇 and 𝜎 2 for 𝜇𝑋 and 𝜎𝑋2 , respectively.
~ 17 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Again, the sample means 𝑥 (number) are NOT always the same when different
sets of sample are used. So, how do they vary? The sampling distribution of 𝑋 ̅ can
give us an information how the numerical guesses vary. Interestingly, including
the variability with a particular guess 𝑥 can give us another popular estimation
method for the unknown population mean 𝜇𝑋 ----- an interval-valued guess or an
interval-valued estimate.
We would first study the methodology to get the so-called a RANDOM interval,
and then discuss its actual “value” which is often called CONFIDENCE interval.
Point-valued Interval-valued
estimation estimation
Methodology, RV Estimator Random Interval
Numerical Estimate Confidence interval
~ 18 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝑋̅−𝜇𝑋
Or ∼ 𝑁(0, 1) .
𝜎2
√ 𝑋
𝑛
∑ 𝑋𝑖 = 𝑋1 + 𝑋2 + ⋯ + 𝑋𝑛 ∼ 𝐵𝑖𝑛𝑜𝑚𝑖𝑎𝑙 (𝑛𝑚, 𝑝) .
𝑖=1
∑ 𝑋𝑖 = 𝑋1 + 𝑋2 + ⋯ + 𝑋𝑛 ∼ 𝑃𝑜𝑖𝑠𝑠𝑜𝑛(𝑛𝜆) .
𝑖=1
~ 19 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
In general, for any other distribution of X, it is very difficult to determine the exact
sampling distribution of 𝑋̅. However, luckily, if n is sufficiently large, then we can
get a good approximation of the exact distribution by the following famous result
----- Central Limit Theorem.
(Theorem)
If 𝑋 follows any distribution with finite 𝝁𝑿 and positive finite 𝝈𝟐𝑿 , then
we have
𝑋̅ − 𝜇𝑋
→ 𝑁(0, 1) ,
2
𝜎
√ 𝑋
𝑛
if n is sufficiently large.
Throughout our course, we ONLY consider the RANDOM INTERVAL in the case
that X follows a Normal distribution, i.e. 𝑋 ∼ 𝑁(𝜇𝑋 , 𝜎𝑋2 ).
𝜎𝑋2
𝑋̅ ∼ 𝑁 (𝜇𝑋 , ) .
𝑛
~ 20 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝜇𝑋 − 𝑘1 𝜇𝑋 𝜇𝑋 + 𝑘2
According to this curve, which interval should we use such that 𝜇𝑋 can be
included? Obviously, any interval in form of [ 𝝁𝑿 − 𝒌𝟏 , 𝝁𝑿 + 𝒌𝟐 ],
where 0 < 𝑘1 ≤ +∞ 𝑎𝑛𝑑 0 < 𝑘2 ≤ +∞ , i.e. including (− ∞, +∞), can include
the parameter 𝝁𝑿 .
Moreover, we also observe that unlike the point estimate or estimator, we can
talk more about the interval with PROBABILITY.
Note that the area of the shaded region in the picture above is
𝑃(𝜇𝑋 − 𝑘1 ≤ 𝑋̅ ≤ 𝜇𝑋 + 𝑘2 )
say 𝐶. Thus, we have
𝑃(𝑋̅ − 𝑘2 ≤ 𝜇𝑋 ≤ 𝑋̅ + 𝑘1 ) = 𝐶 .
The end points of the interval [ 𝑋̅ − 𝑘2 , 𝑋̅ + 𝑘1 ] include 𝑋̅. That is, the interval is
RANDOM. Such an interval in practice is often called a RANDOM INTERVAL.
~ 21 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
In addition, we also notice that the above result actually tells us that
The largest value of 𝐶 is 1. Then, that’s easy! We just find the interval with
probability 1. That is, the interval always contains the unknown parameter!!
Therefore, we can only consider a large value (but not 1) of C for the random
interval. The commonly used values of C for random intervals are 0.9, 0.95 and
0.99.
~ 22 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
In the following, we would study how to get the random interval exactly with
C = 0.95. First, we don’t want the interval to bias towards any either side, thus we
would consider the random interval with 𝑘1 = 𝑘2 .
Recall that
Therefore,
𝜎𝑋 𝜎𝑋
𝑃 (𝜇𝑋 − 2 ≤ 𝑋̅ ≤ 𝜇𝑋 + 2 ) = 0.9544997 .
√𝑛 √𝑛
~ 23 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝜎𝑋 𝜎𝑋
𝑃 (𝜇𝑋 − 1.96 ≤ 𝑋̅ ≤ 𝜇𝑋 + 1.96 ) = 0.9500042 .
√𝑛 √𝑛
Let’s make the value 0.9500042 correct to 4 d.p. for our following steps. Thus, we
have
𝜎𝑋 𝜎𝑋
𝑃 (𝜇𝑋 − 1.96 ≤ 𝑋̅ ≤ 𝜇𝑋 + 1.96 ) = 0.9500 (𝑜𝑟 0.95)
√𝑛 √𝑛
and
𝜎𝑋 𝜎𝑋
𝑃 (𝑋̅ − 1.96 ≤ 𝜇𝑋 ≤ 𝑋̅ + 1.96 ) = 0.95 .
√𝑛 √𝑛
In other words,
𝜎 𝜎
[ 𝑋̅ − 1.96 𝑋 , 𝑋̅ + 1.96 𝑋 ] is the random interval
√𝑛 √𝑛
(estimation method) with probability 0.95 including
the unknown parameter 𝝁𝑿 being estimated.
This estimation method induced by the estimator 𝑋̅ and its variability can provide
us another way to guess the true value of the unknown parameter 𝜇𝑋 .
~ 24 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Recall that confidence interval is an actual “value” of the random interval. Thus,
according to a particular set of collected data, we have the following general
formal for the so-called 95% confidence interval of 𝝁𝑿 :
𝜎𝑋 𝜎𝑋
[ 𝑥 − 1.96 , 𝑥 + 1.96 ],
√𝑛 √𝑛
when 𝝈𝑿 is known.
Indeed, if we have 𝐵 different sets of data, then we would have 𝐵 sample means
𝜎 𝜎
and according to the estimation method [ 𝑋̅ − 1.96 𝑋 , 𝑋̅ + 1.96 𝑋 ] we can get
√𝑛 √𝑛
𝐵 confidence intervals. Some of these confidence intervals would contain 𝜇𝑋 , and
some would not.
𝜇𝑋
~ 26 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
If B is large enough, we would see that the relative frequency --- proportion ---
that the confidence interval contains 𝜇𝑋 would tend to 0.95 in some probabilistic
sense.
This is also consistent with our material on page 8 of Chapter 3.
When 𝝈𝑿 is known, the 100(1 − 𝛼)% confidence interval (hereafter, C.I.) for
the unknown parameter 𝜇𝑋 is given by
𝜎𝑋 𝜎𝑋
[ 𝑥 − 𝑧𝛼 , 𝑥 + 𝑧𝛼 ],
2 √𝑛 2 √𝑛
where 𝑧𝛼 is defined as
2
𝛼
𝑃 (𝑍 ≤ 𝑧𝛼 ) = 1 − .
2 2
~ 27 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝑠𝑛−1 𝑠𝑛−1
𝑃 (𝑋̅ − 1.96 ≤ 𝜇𝑋 ≤ 𝑋̅ + 1.96 )
√𝑛 √𝑛
2
𝑠𝑛−1
𝑋̅ ∼ 𝑁 (𝜇𝑋 , ).
𝑛
̅ − 𝝁𝑿
𝑿
∼ 𝒕𝒏−𝟏
𝟐
𝑺
√ 𝒏−𝟏
𝒏
which means that the random variable on the left follows
a 𝒕 distribution with 𝒏 − 𝟏 degrees of freedom.
Remark that
Unlike the normal distribution having two terms 𝑎 and 𝑏 (referring to the
very beginning form of the pdf of normal distribution), a 𝑡 distribution only
has one term, often called a degree of freedom.
~ 28 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
Note that the sample variance in the denominator above is a rv, not a
number.
~ 29 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
̅ − 𝝁𝑿
𝑿
𝑷 −𝒕𝒏−𝟏,𝜶 ≤ ≤ 𝒕𝒏−𝟏,𝜶 =𝟏−𝜶,
𝟐
𝟐 𝑺
√ 𝒏−𝟏
𝟐
( 𝒏 )
𝛼
𝑃 (𝑇𝑛−1 > 𝑡𝑛−1,𝛼 ) =
2 2
and can be determined by the t table posted on the course webpage. I would
show you more details how to find the value of 𝑡𝑛−1,𝛼 in lecture.
2
𝑺𝟐𝒏−𝟏 𝑺𝟐𝒏−𝟏
̅−𝒕
𝑷 (𝑿 𝜶√ ̅+𝒕
≤ 𝝁𝑿 ≤ 𝑿 𝜶√ )=𝟏−𝜶.
𝒏−𝟏,
𝟐 𝒏 𝒏−𝟏,
𝟐 𝒏
~ 30 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
𝑆𝑛−1 𝑆𝑛−1
[ 𝑋̅ − 𝒕𝒏−𝟏,𝜶 , 𝑋̅ + 𝒕𝒏−𝟏,𝜶 ],
𝟐 √𝑛 𝟐 √𝑛
𝑠𝑛−1 𝑠𝑛−1
[ 𝑥 − 𝒕𝒏−𝟏,𝜶 , 𝑥 + 𝒕𝒏−𝟏,𝜶 ].
𝟐 √𝑛 𝟐 √𝑛
~ 31 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
EXAMPLE
Frequencies, in hertz (Hz), of 12 elephant calls:
14, 16, 17, 17, 24, 20, 32, 18, 29, 31, 15, 35
Assume that the population of possible elephant call frequencies (𝑋) is a normal
distribution, Now a scientist is interested in the expected frequency 𝜇𝑋 of 𝑋 and
want to find a 95% confidence interval for 𝝁𝑿 .
QUESTION
A paint manufacturer wants to determine the average drying time of a new
interior wall paint. Assume that the drying time follows a normal distribution. If
for 12 test areas of equal size he obtained a mean drying time of 66.3 minutes
and a standard deviation of 8.4 minutes, then what is the 95% confidence interval
for the population mean?
~ 32 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
to construct a random interval for 𝜎𝑋2 and then get a confidence interval.
We use the following theoretical result:
(𝒏 − 𝟏)𝑺𝟐𝒏−𝟏 𝟐
∼ 𝝌𝒏−𝟏
𝝈𝟐𝑿
which means that the random variable on the left follows
a 𝝌𝟐 distribution with 𝒏 − 𝟏 degrees of freedom.
Remark that
Like a 𝑡 distribution having one term (degree of freedom), a 𝜒 2 distribution
also has one term and this term is also often called a degree of freedom.
Note that the sample variance in the numerator above is a rv, not a number.
~ 33 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
(𝒏 − 𝟏)𝑺𝟐𝒏−𝟏
𝑷 (𝝌𝟐𝒏−𝟏,𝟏−𝜶 ≤ 𝟐 ≤ 𝝌𝟐
𝜶) = 𝟏 − 𝜶 ,
𝟐 𝝈𝑿 𝒏−𝟏,
𝟐
2 2
where 𝜒𝑛−1,𝛼 (similar for 𝜒
𝑛−1,1−
𝛼 ) is defined as
2 2
2 𝛼
𝑃 (𝑈𝑛−1 > 𝜒𝑛−1, 𝛼) =
2 2
and can be determined by the 𝝌𝟐 table posted on the course webpage. I would
2 2
show you more details how to find the values of 𝜒𝑛−1,𝛼 and 𝜒
𝑛−1,1−
𝛼 in lecture.
2 2
~ 34 ~
MATH2411: Applied Statistics | Dr. YU, Chi Wai
(𝒏 − 𝟏)𝑺𝟐𝒏−𝟏 (𝒏 − 𝟏)𝑺𝟐𝒏−𝟏
𝑷( ≤ 𝝈𝟐𝑿 ≤ )=𝟏−𝜶.
𝝌𝟐 𝜶 𝝌𝟐 𝜶
𝒏−𝟏, 𝒏−𝟏,𝟏−
𝟐 𝟐
(𝒏 − 𝟏)𝑺𝟐𝒏−𝟏 (𝒏 − 𝟏)𝑺𝟐𝒏−𝟏
[ , ],
𝝌𝟐 𝜶 𝝌𝟐 𝜶
𝒏−𝟏, 𝒏−𝟏,𝟏−
𝟐 𝟐
(𝒏 − 𝟏)𝒔𝟐𝒏−𝟏 (𝒏 − 𝟏)𝒔𝟐𝒏−𝟏
[ , ].
𝝌𝟐 𝜶 𝝌𝟐 𝜶
𝒏−𝟏, 𝒏−𝟏,𝟏−
𝟐 𝟐
~ 35 ~