Module 6
(Statistical Properties of MLEs and Fisher
Information)
Invariance Property of MLEs
Theorem:
If 𝜃 is the MLE of 𝜃, then for any function 𝜏(𝜃), the MLE of 𝜏(𝜃) is
𝜏(𝜃).
For this proof we need to know about the concept of induced likelihood
function 𝐿∗ , which is defined as
𝐿∗ 𝜂; 𝒙 = sup 𝐿(𝜃; 𝒙)
{𝜃:𝜏 𝜃 =𝜂}
The value 𝜂Ƹ that maximizes 𝐿∗ (𝜂; 𝒙) will be called the MLE of 𝜂 = 𝜏(𝜃),
and it is clear that the maxima of 𝐿∗ and 𝐿 coincide.
Invariance Property of MLEs
Proof:
Let 𝜂Ƹ denote the value that maximizes 𝐿∗ (𝜂; 𝒙). We must show that
𝐿∗ 𝜂;Ƹ 𝒙 = 𝐿∗ [𝜏 𝜃 ; 𝒙]. Now, as stated before, the maxima of 𝐿 and 𝐿∗
coincide, so we have
𝐿∗ 𝜂;Ƹ 𝒙 = sup sup 𝐿 𝜃; 𝒙
𝜂 𝜃:𝜏 𝜃 =𝜂
= sup 𝐿(𝜃; 𝒙)
𝜃
𝒙)
= 𝐿(𝜃;
where the second equality follows because the iterated maximization is
equal to the unconditional maximization over 𝜃, which is attained at 𝜃.
𝒙 =
Furthermore, 𝐿 𝜃; sup 𝒙)].
𝐿 𝜃; 𝒙 = 𝐿∗ [𝜏(𝜃;
𝜃;𝜏 𝜃 =𝜏 𝜃
𝒙)] and that
Hence, the string of inequalities show that,𝐿∗ 𝜂;Ƹ 𝒙 = 𝐿∗ [𝜏(𝜃;
is the MLE of 𝜏(𝜃).
𝜏(𝜃)
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝑁(𝜃, 1) distribution.
Then find the MLE for 𝜏 = 𝑒 𝜃 .
Solution :
The probability density function for 𝑋𝑖′ 𝑠 is given by
1 −1 𝑥𝑖 −𝜃 2
𝑓 𝑥𝑖 ; 𝜃 = 𝑒 2 , 𝑥𝑖 ∈ ℝ
2𝜋
First, we will find the MLE for 𝜃.
The Likelihood function can be given as
𝐿 𝜃; 𝒙 = ς𝑛𝑖=1 𝑓(𝑥𝑖 ; 𝜃)
1 1
−2 σ𝑛 𝑥𝑖 −𝜃 2
𝐿 𝜃; 𝒙 = 𝑛𝑒
𝑖
2𝜋
Examples
The Log -Likelihood function can be given as
𝑛
𝑛 1 2
𝑙 𝜃 = − log 2𝜋 − 𝑥𝑖 − 𝜃
2 2
𝑖
Now, we have to maximize 𝑙 𝜃 with respect to 𝜃.
For this, we differentiate 𝑙(𝜃) with respect to 𝜃 and equate it to zero.
Hence, we get
𝑛
𝑑𝑙 𝜃
= (𝑥𝑖 − 𝜃) = 0
𝑑𝜃
𝑖
𝑛
⇒ 𝑥𝑖 − 𝑛𝜃 = 0
𝑖
𝑛
1
⇒ 𝜃 = 𝑥𝑖
𝑛
𝑖
Examples
𝑑2𝑙 𝜃
Now, = −𝑛 < 0,
𝑑𝜃 2
Hence, 𝜃 = 𝑋,
ത maximizes 𝑙 𝜃 .
Therefore, MLE for 𝜃 is 𝜃 = 𝑋.
ത
By Invariance property of MLEs,
MLE of 𝜏 = 𝑒 𝜃 is 𝜏Ƹ = 𝑒 𝜃 = 𝑒 𝑋 .
Examples
• Using the theorem, we now see that the MLE of 𝜃 2 , the square of a
normal mean is 𝑋ത 2 .
• We can also apply this theorem to more complicated functions to see
that, for example, the MLE of 𝑝(1 − 𝑝), where 𝑝 is a binomial
probability, is given by 𝑝(1
Ƹ − 𝑝).
Ƹ
Invariance Property of MLEs
• The invariance property of MLEs also holds in the multivariate case.
• If the MLE of (𝜃1 ,…, 𝜃𝑘 ) is 𝜃መ1 , … , 𝜃መ𝑘 , and if 𝜏(𝜃1 ,…, 𝜃𝑘 ) is any function
of the parameters, the MLE of 𝜏(𝜃1 ,…, 𝜃𝑘 ) is 𝜏(𝜃መ1 , … , 𝜃መ𝑘 ).
• If 𝜽 = (𝜃1 ,…, 𝜃𝑘 ) is multidimensional, then the problem of finding an
MLE is that of maximizing a function of several variables.
Module 6
(Consistency Property of MLEs)
Regularity Conditions
Let 𝑋1 , 𝑋2 , … be a sequence of i.i.d. random variables from a population
having PMF/PDF 𝑓(𝑥; 𝜃), where 𝜃 ∈ Θ ⊆ ℛ. Let the true value of 𝜃 is
𝜃0 . Consider the following assumptions :
1. The parameters are identifiable i.e. if 𝜃 ≠ 𝜃′, then 𝑓 𝑥; 𝜃 ≠ 𝑓(𝑥; 𝜃′).
2. The densities 𝑓(𝑥; 𝜃) have common support, and 𝑓(𝑥; 𝜃) is
differentiable in 𝜃.
3. The parameter space Θ contains an open set of which the true
parameter value 𝜃0 is an interior point.
Regularity Conditions
4. The density function 𝑓(𝑥; 𝜃) is three times differentiable with
respect to 𝜃 for all 𝑥 ∈ ℛ and for all 𝜃 ∈ Θ , the third derivative is
continuous in 𝜃 , and ∫ 𝑓 𝑥; 𝜃 𝑑𝑥 can be differentiated three
times under the integral sign.
5. For any 𝜃𝑜 ∈ Θ, there exists a positive number 𝑐 and a function
𝑀(𝑥)(both of which may depend on 𝜃0 ) such that,
𝜕3
ln 𝑓 𝑥; 𝜃 < 𝑀 𝑥 , ∀𝑥 ∈ ℛ, 𝜃0 − 𝑐 < 𝜃 < 𝜃0 + 𝑐,
𝜕𝜃 3
with 𝐸𝜃0 𝑀 𝑋1 < ∞.
Consistency Property of MLEs
Theorem:
Under the regularity conditions stated before, the likelihood
equation has a solution denoted by 𝜃𝑛 (x), such that 𝜃
𝑛 (𝑿) is
consistent estimator for 𝜃.
In other words,
The Maximum Likelihood Estimator (MLE) for 𝜃 is also
consistent for 𝜃.
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝑈 0, 𝜃
distribution then suggest a consistent estimator for 𝜃.
Solution:
We know that MLEs are consistent estimators, so we find MLE
for 𝜃.
Now, the probability distribution function of 𝑋𝑖′ 𝑠 is given as
1
𝑓𝑋𝑖 (xi , 𝜃) = ቐ𝜃 if 0 ≤ xi ≤ 𝜃
0 otherwise
Examples
Therefore, the likelihood function is
𝐿𝑋1 𝑋2 ⋯𝑋𝑛 𝜃; 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 = 𝑓𝑋1 𝜃; 𝑥1 𝑓𝑋2 𝜃; 𝑥2 ⋯ 𝑓𝑋𝑛 𝜃; 𝑥𝑛
Now,
1
if 0 ≤ 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ 𝜃
𝑓𝑋1 𝜃; 𝑥1 𝑓𝑋2 𝜃; 𝑥2 ⋯ 𝑓𝑋𝑛 𝜃; 𝑥𝑛 = ൝𝜃 𝑛
0 otherwise
1
if 0 ≤ min 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ max(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ) ≤ 𝜃
= ൝𝜃𝑛
0 otherwise
Examples
Thus, the likelihood function attains maximum at 𝜃 when
0 ≤ min 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ max(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ) ≤ 𝜃
When 𝜃 > 𝑋(𝑛) the likelihood function 𝐿𝑋1 𝑋2 ⋯𝑋𝑛 𝜃; 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛
decreases as 𝜃 increases and when 𝜃 < 𝑋(𝑛) , the likelihood function is
zero.
Hence, the maximum is attained precisely at 𝜃 = 𝑋(𝑛) .
Therefore, we have 𝜃 = 𝑋(𝑛) as MLE and consequently a consistent
estimator for 𝜃.
Module 6
(Fisher Information)
Fisher Information
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random variables with common pdf (pmf)
𝑓𝜃 (𝑥). Then,
2 𝑛 2
𝜕 log 𝑓𝜃 𝑿 𝜕 log 𝑓𝜃 𝑋1
𝐼 𝜃 = 𝐸𝜃 = 𝐸𝜃
𝜕𝜃 𝜕𝜃
𝑖=1
2
𝜕 log 𝑓𝜃 𝑋1
= 𝑛𝐸𝜃 = 𝑛𝐼1 (𝜃)
𝜕𝜃
𝜕 log 𝑓𝜃 𝑋1 2
where 𝐼1 𝜃 = 𝐸𝜃 .
𝜕𝜃
Fisher Information
Definition:
The quantity
2
𝜕 log 𝑓𝜃 𝑋1
𝐼1 𝜃 = 𝐸𝜃
𝜕𝜃
is called Fisher Information in 𝑋1 and
2
𝜕 log 𝑓𝜃 𝑿
𝐼𝑛 𝜃 = 𝐸𝜃 = 𝑛𝐼1 (𝜃)
𝜕𝜃
is called Fisher Information in random sample 𝑋1 , 𝑋2 , … , 𝑋𝑛 .
𝜕 log 𝑓𝜃 𝑋 2 𝜕2 log 𝑓𝜃 𝑋
Remark: I 𝜃 = 𝐸𝜃 = −𝐸𝜃
𝜕𝜃 𝜕𝜃 2
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information Matrix
Definition:
If 𝑋 is a random variable with probability density function 𝑓 𝑥; 𝜃 ,
where 𝜃 = (𝜃1 , 𝜃2 , … , 𝜃𝑛 ) is an unknown parameter vector then
the Fisher information, 𝐼(𝜃), I a 𝑛 × 𝑛 matrix given by
𝜕 2 log 𝑓𝜃 𝑋
𝐼 𝜃 = (𝐼𝑖𝑗 (𝜃)) = −𝐸𝜃
𝜕𝜃𝑖 𝜕𝜃𝑗
Examples
Question:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be a random sample from a normal population
with mean 𝜇 and variance 𝜎 2 . What is the Fisher Information
Matrix 𝐼𝑛 (𝜇, 𝜎 2 ), of the sample size 𝑛 about the parameters 𝜇 and
𝜎2?
Solution:
Let us take 𝜃1 = 𝜇 and 𝜃2 = 𝜎 2 . Then the Fisher Information,
𝐼𝑛 (𝜃), in sample of size 𝑛 about the parameter (𝜃1 , 𝜃2 ) is equal to
𝑛 times the Fisher Information in the population about 𝜃1 , 𝜃2 , i.e.
𝐼𝑛 𝜃1 , 𝜃2 = 𝑛𝐼(𝜃1 , 𝜃2 )
Examples
Since there are two parameters 𝜃1 = 𝜇 and 𝜃2 = 𝜎 2 ,the Fisher
Information, 𝐼(𝜃1 , 𝜃2 ) is a 2 × 2 matrix given by
𝐼11 (𝜃1 , 𝜃2 ) 𝐼12 (𝜃1 , 𝜃2 )
𝐼 𝜃1 , 𝜃2 =
𝐼21 (𝜃1 , 𝜃2 ) 𝐼22 (𝜃1 , 𝜃2 )
where
𝜕 2 log 𝑓(𝑋; 𝜃1 , 𝜃2 )
𝐼𝑖𝑗 𝜃1 , 𝜃2 = −𝐸𝜃
𝜕𝜃𝑖 𝜕𝜃𝑗
For 𝑖 = 1,2 and 𝑗 = 1,2. Now we want to compute 𝐼𝑖𝑗 .
𝑥−𝜃 2
1 − 2𝜃 1
Since, 𝑓 𝑥; 𝜃1 , 𝜃2 = 𝑒 2
2𝜋𝜃2
1 𝑥−𝜃1 2
we have log 𝑓 𝑥; 𝜃1 , 𝜃2 = − log 2𝜋𝜃2 − .
2 2𝜃2
Examples
1 𝑥−𝜃1 2
We have log 𝑓 𝑥; 𝜃1 , 𝜃2 = − log 2𝜋𝜃2 − .
2 2𝜃2
Calculating the partials of log 𝑓 𝑥; 𝜃1 , 𝜃2 , we have,
𝜕 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 𝑥−𝜃1
= ,
𝜕𝜃1 𝜃2
𝜕 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1 𝑥−𝜃1 2
= − + ,
𝜕𝜃2 2𝜃2 2𝜃22
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1
=− ,
𝜕𝜃12 𝜃2
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1 𝑥−𝜃1 2
= − 2 − ,
𝜕𝜃22 2𝜃2 𝜃23
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 𝑥−𝜃1
=− .
𝜕𝜃1 𝜕𝜃2 𝜃22
Examples
Hence,
1 1 1
𝐼11 𝜃1 , 𝜃2 = −𝐸 − =− = .
𝜃2 𝜃2 𝜎2
Similarly, we have
𝑋−𝜃1
𝐼21 𝜃1 , 𝜃2 = 𝐼12 𝜃1 , 𝜃2 = −𝐸 − =0
𝜃22
and
1 𝑋−𝜃1 2 1 1
𝐼22 𝜃1 , 𝜃2 = −𝐸 − 2 − = = .
2𝜃2 𝜃23 2𝜃22 2𝜎4
Hence, the Fisher Information matrix is given by
1 𝑛
2
0 2
0
𝐼𝑛 𝜇, 𝜎 2 = 𝑛 𝜎 = 𝜎 𝑛 .
1
0 0
2𝜎 4 2𝜎 4
Module 6
(Asymptotic Normality of MLEs)
Asymptotic Normality of MLEs
Theorem:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random variables with common pdf (pmf)
𝑓𝜃 (𝑥). Let 𝜃 be MLE of 𝜃, and let 𝜏(𝜃) be a continuous function of 𝜃.
Then, under some regularity conditions,
1
𝑛 𝜏 𝜃 − 𝜏 𝜃 → 𝑁 0, ,
𝐼1 (𝜃)
where 𝐼1 (𝜃) is the Fisher Information in 𝑋1 .
Asymptotic Normality of MLEs
Remark:
• 𝜃 is asymptotically unbiased i.e. 𝐸 𝜃 → 𝜃 as 𝑛 → ∞ .
1
• The variance of 𝜃 is approximately .
𝑛𝐼1 (𝜃)
• If the true parameter is 𝜃0 ,then the sampling distribution of 𝜃 is
1
approximately 𝑁 𝜃0 , .
𝑛𝐼1 (𝜃0 )
• 𝜃 is asymptotically efficient.
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝐵𝑒𝑟(𝑝) population.
Then find the Fisher Information 𝐼1 𝑝 and show that, when true value
of parameter is 𝑝0 , 𝑛 𝑋ത − 𝑝0 → 𝑁(0, 𝑝0 (1 − 𝑝0 )). Also, what will be
the approximate sampling distribution of 𝑋. ത
Solution:
The probability mass function of 𝑋𝑖′ 𝑠 can be given as
𝑓𝑝 𝑥 = 𝑝 𝑥 1 − 𝑝 1−𝑥
Examples
Now,
𝑓𝑝 𝑋 = 𝑝 𝑥 1 − 𝑝 1−𝑥
Taking log, we get
log 𝑓𝑝 𝑥 = 𝑥 log 𝑝 + 1 − 𝑥 log(1 − 𝑝)
Now, differentiating with respect to 𝑝, we have,
𝜕 log 𝑓𝑝 (𝑥) 𝑥 1−𝑥
= −
𝜕𝑝 𝑝 1−𝑝
𝜕2 log 𝑓𝑝 (𝑥) 𝑥 1−𝑥
Hence, =− −
𝜕𝑝2 𝑝2 1−𝑝 2
Examples
𝜕2 log 𝑓𝑝 𝑋 𝑋 1−𝑋
Since, 𝐼1 𝑝 = −𝐸 = −𝐸 − 2 −
𝜕𝑝2 𝑝 1−𝑝 2
𝑋 1−𝑋
=𝐸 +𝐸
𝑝2 1−𝑝 2
1 1
= 𝐸 𝑋 + 𝐸 1−𝑋
𝑝2 1−𝑝 2
𝑝 1−𝑝
= +
𝑝2 1−𝑝 2
1 1 1
Hence, 𝐼1 𝑝 = + = .
𝑝 1−𝑝 𝑝(1−𝑝)
Examples
We know that 𝑋ത is MLE for 𝑝.
Therefore, by asymptotic normality of MLEs,
1
ത
𝑛 𝑋 − 𝑝 → 𝑁 0,
𝐼1 𝑝
⇒ 𝑛 𝑋ത − 𝑝 → 𝑁(0, 𝑝(1 − 𝑝)).
Now, asymptotic sampling distribution of 𝑋ത means sampling distribution
of 𝑋ത as 𝑛 → ∞ ,
When true value of parameter is 𝑝0 , the sampling distribution of 𝑋ത is
1 𝑝(1−𝑝0 )
𝑁 𝑝0 , i.e. 𝑁 𝑝0 , .
𝑛𝐼1 (𝑝0 ) 𝑛