0% found this document useful (0 votes)
8 views41 pages

Module 6

The document discusses the invariance property of Maximum Likelihood Estimators (MLEs) and the consistency property of MLEs, detailing theorems and proofs related to these concepts. It provides examples of finding MLEs for specific distributions and explains the regularity conditions necessary for MLE consistency. Additionally, it introduces Fisher Information and its matrix representation for random variables with unknown parameter vectors.

Uploaded by

Ananya Kukreja
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views41 pages

Module 6

The document discusses the invariance property of Maximum Likelihood Estimators (MLEs) and the consistency property of MLEs, detailing theorems and proofs related to these concepts. It provides examples of finding MLEs for specific distributions and explains the regularity conditions necessary for MLE consistency. Additionally, it introduces Fisher Information and its matrix representation for random variables with unknown parameter vectors.

Uploaded by

Ananya Kukreja
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 6

(Statistical Properties of MLEs and Fisher


Information)
Invariance Property of MLEs

Theorem:
If 𝜃෠ is the MLE of 𝜃, then for any function 𝜏(𝜃), the MLE of 𝜏(𝜃) is

𝜏(𝜃).

For this proof we need to know about the concept of induced likelihood
function 𝐿∗ , which is defined as
𝐿∗ 𝜂; 𝒙 = sup 𝐿(𝜃; 𝒙)
{𝜃:𝜏 𝜃 =𝜂}

The value 𝜂Ƹ that maximizes 𝐿∗ (𝜂; 𝒙) will be called the MLE of 𝜂 = 𝜏(𝜃),
and it is clear that the maxima of 𝐿∗ and 𝐿 coincide.
Invariance Property of MLEs
Proof:
Let 𝜂Ƹ denote the value that maximizes 𝐿∗ (𝜂; 𝒙). We must show that
𝐿∗ 𝜂;Ƹ 𝒙 = 𝐿∗ [𝜏 𝜃෠ ; 𝒙]. Now, as stated before, the maxima of 𝐿 and 𝐿∗
coincide, so we have
𝐿∗ 𝜂;Ƹ 𝒙 = sup sup 𝐿 𝜃; 𝒙
𝜂 𝜃:𝜏 𝜃 =𝜂
= sup 𝐿(𝜃; 𝒙)
𝜃
෠ 𝒙)
= 𝐿(𝜃;
where the second equality follows because the iterated maximization is

equal to the unconditional maximization over 𝜃, which is attained at 𝜃.
෠ 𝒙 =
Furthermore, 𝐿 𝜃; sup ෠ 𝒙)].
𝐿 𝜃; 𝒙 = 𝐿∗ [𝜏(𝜃;

𝜃;𝜏 𝜃 =𝜏 𝜃

෠ 𝒙)] and that


Hence, the string of inequalities show that,𝐿∗ 𝜂;Ƹ 𝒙 = 𝐿∗ [𝜏(𝜃;
෠ is the MLE of 𝜏(𝜃).
𝜏(𝜃)
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝑁(𝜃, 1) distribution.
Then find the MLE for 𝜏 = 𝑒 𝜃 .
Solution :
The probability density function for 𝑋𝑖′ 𝑠 is given by
1 −1 𝑥𝑖 −𝜃 2
𝑓 𝑥𝑖 ; 𝜃 = 𝑒 2 , 𝑥𝑖 ∈ ℝ
2𝜋
First, we will find the MLE for 𝜃.
The Likelihood function can be given as
𝐿 𝜃; 𝒙 = ς𝑛𝑖=1 𝑓(𝑥𝑖 ; 𝜃)
1 1
−2 σ𝑛 𝑥𝑖 −𝜃 2
𝐿 𝜃; 𝒙 = 𝑛𝑒
𝑖

2𝜋
Examples
The Log -Likelihood function can be given as
𝑛
𝑛 1 2
𝑙 𝜃 = − log 2𝜋 − ෍ 𝑥𝑖 − 𝜃
2 2
𝑖
Now, we have to maximize 𝑙 𝜃 with respect to 𝜃.
For this, we differentiate 𝑙(𝜃) with respect to 𝜃 and equate it to zero.
Hence, we get
𝑛
𝑑𝑙 𝜃
= ෍(𝑥𝑖 − 𝜃) = 0
𝑑𝜃
𝑖
𝑛

⇒ ෍ 𝑥𝑖 − 𝑛𝜃 = 0
𝑖
𝑛
1
⇒ 𝜃෠ = ෍ 𝑥𝑖
𝑛
𝑖
Examples

𝑑2𝑙 𝜃
Now, = −𝑛 < 0,
𝑑𝜃 2

Hence, 𝜃෠ = 𝑋,
ത maximizes 𝑙 𝜃 .

Therefore, MLE for 𝜃 is 𝜃෠ = 𝑋.


By Invariance property of MLEs,



MLE of 𝜏 = 𝑒 𝜃 is 𝜏Ƹ = 𝑒 𝜃 = 𝑒 𝑋෠ .
Examples
• Using the theorem, we now see that the MLE of 𝜃 2 , the square of a
normal mean is 𝑋ത 2 .

• We can also apply this theorem to more complicated functions to see


that, for example, the MLE of 𝑝(1 − 𝑝), where 𝑝 is a binomial
probability, is given by 𝑝(1
Ƹ − 𝑝).
Ƹ
Invariance Property of MLEs

• The invariance property of MLEs also holds in the multivariate case.

• If the MLE of (𝜃1 ,…, 𝜃𝑘 ) is 𝜃መ1 , … , 𝜃መ𝑘 , and if 𝜏(𝜃1 ,…, 𝜃𝑘 ) is any function
of the parameters, the MLE of 𝜏(𝜃1 ,…, 𝜃𝑘 ) is 𝜏(𝜃መ1 , … , 𝜃መ𝑘 ).

• If 𝜽 = (𝜃1 ,…, 𝜃𝑘 ) is multidimensional, then the problem of finding an


MLE is that of maximizing a function of several variables.
Module 6
(Consistency Property of MLEs)
Regularity Conditions

Let 𝑋1 , 𝑋2 , … be a sequence of i.i.d. random variables from a population


having PMF/PDF 𝑓(𝑥; 𝜃), where 𝜃 ∈ Θ ⊆ ℛ. Let the true value of 𝜃 is
𝜃0 . Consider the following assumptions :

1. The parameters are identifiable i.e. if 𝜃 ≠ 𝜃′, then 𝑓 𝑥; 𝜃 ≠ 𝑓(𝑥; 𝜃′).


2. The densities 𝑓(𝑥; 𝜃) have common support, and 𝑓(𝑥; 𝜃) is
differentiable in 𝜃.
3. The parameter space Θ contains an open set of which the true
parameter value 𝜃0 is an interior point.
Regularity Conditions

4. The density function 𝑓(𝑥; 𝜃) is three times differentiable with


respect to 𝜃 for all 𝑥 ∈ ℛ and for all 𝜃 ∈ Θ , the third derivative is
continuous in 𝜃 , and ∫ 𝑓 𝑥; 𝜃 𝑑𝑥 can be differentiated three
times under the integral sign.

5. For any 𝜃𝑜 ∈ Θ, there exists a positive number 𝑐 and a function


𝑀(𝑥)(both of which may depend on 𝜃0 ) such that,
𝜕3
ln 𝑓 𝑥; 𝜃 < 𝑀 𝑥 , ∀𝑥 ∈ ℛ, 𝜃0 − 𝑐 < 𝜃 < 𝜃0 + 𝑐,
𝜕𝜃 3

with 𝐸𝜃0 𝑀 𝑋1 < ∞.


Consistency Property of MLEs

Theorem:
Under the regularity conditions stated before, the likelihood
equation has a solution denoted by 𝜃෢𝑛 (x), such that 𝜃
෢𝑛 (𝑿) is
consistent estimator for 𝜃.
In other words,
The Maximum Likelihood Estimator (MLE) for 𝜃 is also
consistent for 𝜃.
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝑈 0, 𝜃
distribution then suggest a consistent estimator for 𝜃.

Solution:
We know that MLEs are consistent estimators, so we find MLE
for 𝜃.
Now, the probability distribution function of 𝑋𝑖′ 𝑠 is given as
1
𝑓𝑋𝑖 (xi , 𝜃) = ቐ𝜃 if 0 ≤ xi ≤ 𝜃
0 otherwise
Examples
Therefore, the likelihood function is
𝐿𝑋1 𝑋2 ⋯𝑋𝑛 𝜃; 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 = 𝑓𝑋1 𝜃; 𝑥1 𝑓𝑋2 𝜃; 𝑥2 ⋯ 𝑓𝑋𝑛 𝜃; 𝑥𝑛
Now,
1
if 0 ≤ 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ 𝜃
𝑓𝑋1 𝜃; 𝑥1 𝑓𝑋2 𝜃; 𝑥2 ⋯ 𝑓𝑋𝑛 𝜃; 𝑥𝑛 = ൝𝜃 𝑛
0 otherwise

1
if 0 ≤ min 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ max(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ) ≤ 𝜃
= ൝𝜃𝑛
0 otherwise
Examples

Thus, the likelihood function attains maximum at 𝜃 when


0 ≤ min 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ≤ max(𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛 ) ≤ 𝜃
When 𝜃 > 𝑋(𝑛) the likelihood function 𝐿𝑋1 𝑋2 ⋯𝑋𝑛 𝜃; 𝑥1 , 𝑥2 , ⋯ , 𝑥𝑛
decreases as 𝜃 increases and when 𝜃 < 𝑋(𝑛) , the likelihood function is
zero.
Hence, the maximum is attained precisely at 𝜃෠ = 𝑋(𝑛) .
Therefore, we have 𝜃෠ = 𝑋(𝑛) as MLE and consequently a consistent
estimator for 𝜃.
Module 6
(Fisher Information)
Fisher Information

Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random variables with common pdf (pmf)


𝑓𝜃 (𝑥). Then,
2 𝑛 2
𝜕 log 𝑓𝜃 𝑿 𝜕 log 𝑓𝜃 𝑋1
𝐼 𝜃 = 𝐸𝜃 = ෍ 𝐸𝜃
𝜕𝜃 𝜕𝜃
𝑖=1
2
𝜕 log 𝑓𝜃 𝑋1
= 𝑛𝐸𝜃 = 𝑛𝐼1 (𝜃)
𝜕𝜃
𝜕 log 𝑓𝜃 𝑋1 2
where 𝐼1 𝜃 = 𝐸𝜃 .
𝜕𝜃
Fisher Information
Definition:
The quantity
2
𝜕 log 𝑓𝜃 𝑋1
𝐼1 𝜃 = 𝐸𝜃
𝜕𝜃
is called Fisher Information in 𝑋1 and
2
𝜕 log 𝑓𝜃 𝑿
𝐼𝑛 𝜃 = 𝐸𝜃 = 𝑛𝐼1 (𝜃)
𝜕𝜃
is called Fisher Information in random sample 𝑋1 , 𝑋2 , … , 𝑋𝑛 .

𝜕 log 𝑓𝜃 𝑋 2 𝜕2 log 𝑓𝜃 𝑋
Remark: I 𝜃 = 𝐸𝜃 = −𝐸𝜃
𝜕𝜃 𝜕𝜃 2
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information
Fisher Information Matrix

Definition:
If 𝑋 is a random variable with probability density function 𝑓 𝑥; 𝜃 ,
where 𝜃 = (𝜃1 , 𝜃2 , … , 𝜃𝑛 ) is an unknown parameter vector then
the Fisher information, 𝐼(𝜃), I a 𝑛 × 𝑛 matrix given by

𝜕 2 log 𝑓𝜃 𝑋
𝐼 𝜃 = (𝐼𝑖𝑗 (𝜃)) = −𝐸𝜃
𝜕𝜃𝑖 𝜕𝜃𝑗
Examples
Question:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be a random sample from a normal population
with mean 𝜇 and variance 𝜎 2 . What is the Fisher Information
Matrix 𝐼𝑛 (𝜇, 𝜎 2 ), of the sample size 𝑛 about the parameters 𝜇 and
𝜎2?

Solution:
Let us take 𝜃1 = 𝜇 and 𝜃2 = 𝜎 2 . Then the Fisher Information,
𝐼𝑛 (𝜃), in sample of size 𝑛 about the parameter (𝜃1 , 𝜃2 ) is equal to
𝑛 times the Fisher Information in the population about 𝜃1 , 𝜃2 , i.e.
𝐼𝑛 𝜃1 , 𝜃2 = 𝑛𝐼(𝜃1 , 𝜃2 )
Examples
Since there are two parameters 𝜃1 = 𝜇 and 𝜃2 = 𝜎 2 ,the Fisher
Information, 𝐼(𝜃1 , 𝜃2 ) is a 2 × 2 matrix given by
𝐼11 (𝜃1 , 𝜃2 ) 𝐼12 (𝜃1 , 𝜃2 )
𝐼 𝜃1 , 𝜃2 =
𝐼21 (𝜃1 , 𝜃2 ) 𝐼22 (𝜃1 , 𝜃2 )
where
𝜕 2 log 𝑓(𝑋; 𝜃1 , 𝜃2 )
𝐼𝑖𝑗 𝜃1 , 𝜃2 = −𝐸𝜃
𝜕𝜃𝑖 𝜕𝜃𝑗
For 𝑖 = 1,2 and 𝑗 = 1,2. Now we want to compute 𝐼𝑖𝑗 .
𝑥−𝜃 2
1 − 2𝜃 1
Since, 𝑓 𝑥; 𝜃1 , 𝜃2 = 𝑒 2
2𝜋𝜃2
1 𝑥−𝜃1 2
we have log 𝑓 𝑥; 𝜃1 , 𝜃2 = − log 2𝜋𝜃2 − .
2 2𝜃2
Examples
1 𝑥−𝜃1 2
We have log 𝑓 𝑥; 𝜃1 , 𝜃2 = − log 2𝜋𝜃2 − .
2 2𝜃2

Calculating the partials of log 𝑓 𝑥; 𝜃1 , 𝜃2 , we have,


𝜕 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 𝑥−𝜃1
= ,
𝜕𝜃1 𝜃2
𝜕 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1 𝑥−𝜃1 2
= − + ,
𝜕𝜃2 2𝜃2 2𝜃22
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1
=− ,
𝜕𝜃12 𝜃2
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 1 𝑥−𝜃1 2
= − 2 − ,
𝜕𝜃22 2𝜃2 𝜃23
𝜕2 log 𝑓(𝑥;𝜃1 ,𝜃2 ) 𝑥−𝜃1
=− .
𝜕𝜃1 𝜕𝜃2 𝜃22
Examples
Hence,
1 1 1
𝐼11 𝜃1 , 𝜃2 = −𝐸 − =− = .
𝜃2 𝜃2 𝜎2

Similarly, we have
𝑋−𝜃1
𝐼21 𝜃1 , 𝜃2 = 𝐼12 𝜃1 , 𝜃2 = −𝐸 − =0
𝜃22

and
1 𝑋−𝜃1 2 1 1
𝐼22 𝜃1 , 𝜃2 = −𝐸 − 2 − = = .
2𝜃2 𝜃23 2𝜃22 2𝜎4

Hence, the Fisher Information matrix is given by


1 𝑛
2
0 2
0
𝐼𝑛 𝜇, 𝜎 2 = 𝑛 𝜎 = 𝜎 𝑛 .
1
0 0
2𝜎 4 2𝜎 4
Module 6
(Asymptotic Normality of MLEs)
Asymptotic Normality of MLEs

Theorem:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random variables with common pdf (pmf)
𝑓𝜃 (𝑥). Let 𝜃෠ be MLE of 𝜃, and let 𝜏(𝜃) be a continuous function of 𝜃.
Then, under some regularity conditions,
1
𝑛 𝜏 𝜃෠ − 𝜏 𝜃 → 𝑁 0, ,
𝐼1 (𝜃)

where 𝐼1 (𝜃) is the Fisher Information in 𝑋1 .


Asymptotic Normality of MLEs

Remark:
• 𝜃෠ is asymptotically unbiased i.e. 𝐸 𝜃෠ → 𝜃 as 𝑛 → ∞ .
1
• The variance of 𝜃෠ is approximately .
𝑛𝐼1 (𝜃)

• If the true parameter is 𝜃0 ,then the sampling distribution of 𝜃෠ is


1
approximately 𝑁 𝜃0 , .
𝑛𝐼1 (𝜃0 )

• 𝜃෠ is asymptotically efficient.
Examples
Example:
Let 𝑋1 , 𝑋2 , … , 𝑋𝑛 be i.i.d. random sample from 𝐵𝑒𝑟(𝑝) population.
Then find the Fisher Information 𝐼1 𝑝 and show that, when true value
of parameter is 𝑝0 , 𝑛 𝑋ത − 𝑝0 → 𝑁(0, 𝑝0 (1 − 𝑝0 )). Also, what will be
the approximate sampling distribution of 𝑋. ത

Solution:
The probability mass function of 𝑋𝑖′ 𝑠 can be given as
𝑓𝑝 𝑥 = 𝑝 𝑥 1 − 𝑝 1−𝑥
Examples
Now,
𝑓𝑝 𝑋 = 𝑝 𝑥 1 − 𝑝 1−𝑥

Taking log, we get


log 𝑓𝑝 𝑥 = 𝑥 log 𝑝 + 1 − 𝑥 log(1 − 𝑝)
Now, differentiating with respect to 𝑝, we have,
𝜕 log 𝑓𝑝 (𝑥) 𝑥 1−𝑥
= −
𝜕𝑝 𝑝 1−𝑝
𝜕2 log 𝑓𝑝 (𝑥) 𝑥 1−𝑥
Hence, =− −
𝜕𝑝2 𝑝2 1−𝑝 2
Examples

𝜕2 log 𝑓𝑝 𝑋 𝑋 1−𝑋
Since, 𝐼1 𝑝 = −𝐸 = −𝐸 − 2 −
𝜕𝑝2 𝑝 1−𝑝 2
𝑋 1−𝑋
=𝐸 +𝐸
𝑝2 1−𝑝 2
1 1
= 𝐸 𝑋 + 𝐸 1−𝑋
𝑝2 1−𝑝 2
𝑝 1−𝑝
= +
𝑝2 1−𝑝 2
1 1 1
Hence, 𝐼1 𝑝 = + = .
𝑝 1−𝑝 𝑝(1−𝑝)
Examples
We know that 𝑋ത is MLE for 𝑝.
Therefore, by asymptotic normality of MLEs,
1

𝑛 𝑋 − 𝑝 → 𝑁 0,
𝐼1 𝑝
⇒ 𝑛 𝑋ത − 𝑝 → 𝑁(0, 𝑝(1 − 𝑝)).
Now, asymptotic sampling distribution of 𝑋ത means sampling distribution
of 𝑋ത as 𝑛 → ∞ ,
When true value of parameter is 𝑝0 , the sampling distribution of 𝑋ത is
1 𝑝(1−𝑝0 )
𝑁 𝑝0 , i.e. 𝑁 𝑝0 , .
𝑛𝐼1 (𝑝0 ) 𝑛

You might also like