0% found this document useful (0 votes)
17 views126 pages

Bayesian Theory and Decision Rules

The document discusses Bayesian statistics, focusing on Bayes theorem, prior distributions, and decision rules. It outlines the framework of Bayesian theory, which integrates classical estimation, hypothesis testing, and interval estimation. The text also addresses the assessment of decision rules through risk functions and the concept of minimax decision rules.

Uploaded by

erobamanuela
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views126 pages

Bayesian Theory and Decision Rules

The document discusses Bayesian statistics, focusing on Bayes theorem, prior distributions, and decision rules. It outlines the framework of Bayesian theory, which integrates classical estimation, hypothesis testing, and interval estimation. The text also addresses the assessment of decision rules through risk functions and the concept of minimax decision rules.

Uploaded by

erobamanuela
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Introduction

Bayes theorem
Prior distribution

Bayesian Statistics.

Ivivi Joseph Mwaniki


Department of Mathematics, University of Nairobi
Email: jimwaniki@[Link]

November 16, 2023

1/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction

Outline Bayes theorem


Prior distribution

1 Introduction
Decision Rules
Admissibility
2 Bayes theorem
Bayesian Inference
Loss functions
3 Prior distribution
Informative Priors
Conjugate Prior
Jeffery’s Prior
2/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

Bayesian theory places classical estimation, hypothesis testing and


interval estimation, within the same wide framework

3/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

Bayesian theory places classical estimation, hypothesis testing and


interval estimation, within the same wide framework
In classical approach, we observe data generated from a certain
population and on the basis of this data a decision has to be
made(estimation,rejection or not reject) about an unknown
parameter, say θ.

3/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

Bayesian theory places classical estimation, hypothesis testing and


interval estimation, within the same wide framework
In classical approach, we observe data generated from a certain
population and on the basis of this data a decision has to be
made(estimation,rejection or not reject) about an unknown
parameter, say θ.
Let X be the space spanned by the random sample and Θ be the
space spanned by the unknown parameter.

3/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

Bayesian theory places classical estimation, hypothesis testing and


interval estimation, within the same wide framework
In classical approach, we observe data generated from a certain
population and on the basis of this data a decision has to be
made(estimation,rejection or not reject) about an unknown
parameter, say θ.
Let X be the space spanned by the random sample and Θ be the
space spanned by the unknown parameter.
In addition to these two spaces we know we have space actions
A = {a1 , a2 , ..., ak } on the basis of the observed data
x = (x1 , x2 , .., xn )

3/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

The decision maker chooses an action ai . Thus once the data


is observed, we are able to take some action such as to
estimate; reject or not reject hypothesis.

4/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

The decision maker chooses an action ai . Thus once the data


is observed, we are able to take some action such as to
estimate; reject or not reject hypothesis.
All such actions have some cost (loss) associated with them.
This implies that by taking a particular action ′ a′ based on
some observed data x the decision maker incurs some loss
that depends on Θ and a hereafter denoted by L(θ, a)

4/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction Bayes theorem
Prior distribution
Admissibility

The decision maker chooses an action ai . Thus once the data


is observed, we are able to take some action such as to
estimate; reject or not reject hypothesis.
All such actions have some cost (loss) associated with them.
This implies that by taking a particular action ′ a′ based on
some observed data x the decision maker incurs some loss
that depends on Θ and a hereafter denoted by L(θ, a)
See the schematic diagram showing linkages of the spaces
Sample Parameter
spaceX space Θ

Action
space A
4/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction.. Bayes theorem
Prior distribution
Admissibility

Let a = d(x) where d is some decision rule, eg, sample mean as an


estimator.

5/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction.. Bayes theorem
Prior distribution
Admissibility

Let a = d(x) where d is some decision rule, eg, sample mean as an


estimator.
L(θ, a) = L(θ, d(x)) which represents the cost incurred by taking
such action ”a” on the basis of the parameter θ.

5/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction.. Bayes theorem
Prior distribution
Admissibility

Let a = d(x) where d is some decision rule, eg, sample mean as an


estimator.
L(θ, a) = L(θ, d(x)) which represents the cost incurred by taking
such action ”a” on the basis of the parameter θ.
In estimation, d(x) is an estimate of θ and thus L(θ, a) is the loss
occurred when θ is the true value.

5/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction.. Bayes theorem
Prior distribution
Admissibility

Let a = d(x) where d is some decision rule, eg, sample mean as an


estimator.
L(θ, a) = L(θ, d(x)) which represents the cost incurred by taking
such action ”a” on the basis of the parameter θ.
In estimation, d(x) is an estimate of θ and thus L(θ, a) is the loss
occurred when θ is the true value.
In hypothesis testing, a test statistic is an example of a decision rule
and the action taken on the basis of the observed value is to either
reject or not rejet

5/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Introduction.. Bayes theorem
Prior distribution
Admissibility

Let a = d(x) where d is some decision rule, eg, sample mean as an


estimator.
L(θ, a) = L(θ, d(x)) which represents the cost incurred by taking
such action ”a” on the basis of the parameter θ.
In estimation, d(x) is an estimate of θ and thus L(θ, a) is the loss
occurred when θ is the true value.
In hypothesis testing, a test statistic is an example of a decision rule
and the action taken on the basis of the observed value is to either
reject or not rejet
Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.

5/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules Bayes theorem
Prior distribution
Admissibility

Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.

6/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules Bayes theorem
Prior distribution
Admissibility

Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.


Basic requirement of the loss functions are
1 L(θ, d(x)) ≥ 0 loss functions are non-negative.
2 L(θ, d(x)) = 0 if θ = a ie no loss
Several criteria are used to assess decision rules

6/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules Bayes theorem
Prior distribution
Admissibility

Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.


Basic requirement of the loss functions are
1 L(θ, d(x)) ≥ 0 loss functions are non-negative.
2 L(θ, d(x)) = 0 if θ = a ie no loss
Several criteria are used to assess decision rules
Equivalence: Two decision rules d1 and d2 are said to be
equivalent if their risk function are equal i.e.
R(θ, d1 (x)) = R(θ, d2 (x)), ∀θ ∈ Θ

6/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules Bayes theorem
Prior distribution
Admissibility

Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.


Basic requirement of the loss functions are
1 L(θ, d(x)) ≥ 0 loss functions are non-negative.
2 L(θ, d(x)) = 0 if θ = a ie no loss
Several criteria are used to assess decision rules
Equivalence: Two decision rules d1 and d2 are said to be
equivalent if their risk function are equal i.e.
R(θ, d1 (x)) = R(θ, d2 (x)), ∀θ ∈ Θ
As good as A decision rule d1 (x) is said to be as good as d2 (x) iff
R(θ, d1 (x)) ≤ R(θ, d2 (x)), for some θ ∈ Θ

6/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules Bayes theorem
Prior distribution
Admissibility

Decision rules are assessed via the risk function namely

R(θ, d(x)) = Ex [L(θ, d(x))]

which is the average loss function.


Basic requirement of the loss functions are
1 L(θ, d(x)) ≥ 0 loss functions are non-negative.
2 L(θ, d(x)) = 0 if θ = a ie no loss
Several criteria are used to assess decision rules
Equivalence: Two decision rules d1 and d2 are said to be
equivalent if their risk function are equal i.e.
R(θ, d1 (x)) = R(θ, d2 (x)), ∀θ ∈ Θ
As good as A decision rule d1 (x) is said to be as good as d2 (x) iff
R(θ, d1 (x)) ≤ R(θ, d2 (x)), for some θ ∈ Θ
Better: A decision rule d1 (x) is said to be better than d2 (x) iff
R(θ, d1 (x)) < R(θ, d2 (x)), ∀θ ∈ Θ
6/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Minimax decision rule Consider the class D of all decision rules of


interest. Then a decision rule d ∗ (x) is said to be minimax iff

max R(θ, d ∗ (x)) = min max R(θ, d(x))


θ d∈V θ

7/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Minimax decision rule Consider the class D of all decision rules of


interest. Then a decision rule d ∗ (x) is said to be minimax iff

max R(θ, d ∗ (x)) = min max R(θ, d(x))


θ d∈V θ

We pick the maximum of all decision rules. Among these maxima


the minimum among them is the minimax decision rule.

7/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Minimax decision rule Consider the class D of all decision rules of


interest. Then a decision rule d ∗ (x) is said to be minimax iff

max R(θ, d ∗ (x)) = min max R(θ, d(x))


θ d∈V θ

We pick the maximum of all decision rules. Among these maxima


the minimum among them is the minimax decision rule.
Example: Suppose x1 , x2 , ..., xn is a random sample from
f (x|θ) = θx (1 − θ)1−x , x = 0, 1 Consider the class D of unbiased
estimators for θ defined as
n
X n
X
d(x) = ci xi , such that ci = 1
j=1 i=1

7/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Minimax decision rule Consider the class D of all decision rules of


interest. Then a decision rule d ∗ (x) is said to be minimax iff

max R(θ, d ∗ (x)) = min max R(θ, d(x))


θ d∈V θ

We pick the maximum of all decision rules. Among these maxima


the minimum among them is the minimax decision rule.
Example: Suppose x1 , x2 , ..., xn is a random sample from
f (x|θ) = θx (1 − θ)1−x , x = 0, 1 Consider the class D of unbiased
estimators for θ defined as
n
X n
X
d(x) = ci xi , such that ci = 1
j=1 i=1

Pn
E(d(x)) = θ i=1 ci = 1. We show that within this class x̄ is the
minimax. One can get the Risk function, minimize and get the
optimal solution.
7/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Solution: Taking examples of unbiased estimators within this class

d1 (x) = x̄, d2 (x) = x1 , d3 (x) = x1 + x4 − x6

8/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Solution: Taking examples of unbiased estimators within this class

d1 (x) = x̄, d2 (x) = x1 , d3 (x) = x1 + x4 − x6

We determine their risk functions respectively


θ(1 − θ)
R(θ, d1 (x)) = var (x̄) =
n
R(θ, d2 (x)) = var (x1 ) = θ(1 − θ)
R(θ, d3 (x)) = Ex [θ − d3 (x)]2 = Ex [d3 (x) − θ]2 = var (d3 (x))
= var (x1 + x4 − x6 ) = var (x1 ) + var (x4 ) + var (x6 )
= 3θ(1 − θ)

8/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules. Bayes theorem
Prior distribution
Admissibility

Solution: Taking examples of unbiased estimators within this class

d1 (x) = x̄, d2 (x) = x1 , d3 (x) = x1 + x4 − x6

We determine their risk functions respectively


θ(1 − θ)
R(θ, d1 (x)) = var (x̄) =
n
R(θ, d2 (x)) = var (x1 ) = θ(1 − θ)
R(θ, d3 (x)) = Ex [θ − d3 (x)]2 = Ex [d3 (x) − θ]2 = var (d3 (x))
= var (x1 + x4 − x6 ) = var (x1 ) + var (x4 ) + var (x6 )
= 3θ(1 − θ)

d( x) = x̄ is the minimax decision rule since

θ(1 − θ)
R(θ, d1 (x)) = var (x̄) =
n
is the minimum of all the other decision rules
8/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules.. Bayes theorem
Prior distribution
Admissibility

Admissibility:
Definition
Consider two decision rules d1 (x) and d2 (x) for θ. The decision rule
d1 (x) is said to dominate d2 (x) if

R(θ, d1 (x)) ≤ R(θ, d2 (x)), ∀ θ and R(θ, d1 (x)) < R(θ, d2 (x)), for some θ

9/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Decision rules.. Bayes theorem
Prior distribution
Admissibility

Admissibility:
Definition
Consider two decision rules d1 (x) and d2 (x) for θ. The decision rule
d1 (x) is said to dominate d2 (x) if

R(θ, d1 (x)) ≤ R(θ, d2 (x)), ∀ θ and R(θ, d1 (x)) < R(θ, d2 (x)), for some θ

A decision rule d ∗ (x) within the class of decision rules D is said to


be admissibleif there does not exist any other decision rule in this
class that dominates it, i.e. d ∗ (x) is admissible if ∄ d(x) ∈ D such
that

R(θ, d(x) ≤ R(θ, d ∗ (x)), ∀ θ and R(θ, d(x) < R(θ, d ∗ (x)), for some θ

9/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples Bayes theorem
Prior distribution
Admissibility

Suppose that f (x|θ) ∝ θx (1 − θ)1−x , x = 0, 1 and a random sample


of size n is drawn from this population. Let x1 , x2 , ..., xn be the
observed sequence of 0′ s and 1′ s.

10/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples Bayes theorem
Prior distribution
Admissibility

Suppose that f (x|θ) ∝ θx (1 − θ)1−x , x = 0, 1 and a random sample


of size n is drawn from this population. Let x1 , x2 , ..., xn be the
observed sequence of 0′ s and 1′ s.
Consider a class of unbiased estimators for θ and d1 (x) = X̄ and
d2 (x) = x1 . Show that d1 (x) is better than d2 (x) in estimating θ
under the loss function L(θ, a) = (θ − a)2 .

10/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples Bayes theorem
Prior distribution
Admissibility

Suppose that f (x|θ) ∝ θx (1 − θ)1−x , x = 0, 1 and a random sample


of size n is drawn from this population. Let x1 , x2 , ..., xn be the
observed sequence of 0′ s and 1′ s.
Consider a class of unbiased estimators for θ and d1 (x) = X̄ and
d2 (x) = x1 . Show that d1 (x) is better than d2 (x) in estimating θ
under the loss function L(θ, a) = (θ − a)2 .
Solution: We need to show that R(θ, d1 (x)) < R(θ, d2 (x)), ∀θ ∈ Θ

R(θ, d1 (x)) = Ex (θ − x̄)2 = var (x̄)


θ(1 − θ)
=
n
R(θ, d2 (x)) = Ex (θ − d2 (x))2 = var (x1 )
= θ(1 − θ)
∴ R(θ, d1 (x)) < R(θ, d2 (x)), ∀θ ∈ Θ

10/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples. Bayes theorem
Prior distribution
Admissibility

Let x1 , x2 , ..., xn be a random sample from a N(µ, σ 2 ). consider the


class of estimators defined as
n
k X
d(k) = kS 2 = (xi − x̄)2
n−1
j=1

11/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples. Bayes theorem
Prior distribution
Admissibility

Let x1 , x2 , ..., xn be a random sample from a N(µ, σ 2 ). consider the


class of estimators defined as
n
k X
d(k) = kS 2 = (xi − x̄)2
n−1
j=1

which decision rule minimizes the risk function in estimating θ = σ 2


under the squared loss function L(θ, a) = (θ − a)2 (which estimator
has the least risk function?)

11/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples. Bayes theorem
Prior distribution
Admissibility

Let x1 , x2 , ..., xn be a random sample from a N(µ, σ 2 ). consider the


class of estimators defined as
n
k X
d(k) = kS 2 = (xi − x̄)2
n−1
j=1

which decision rule minimizes the risk function in estimating θ = σ 2


under the squared loss function L(θ, a) = (θ − a)2 (which estimator
has the least risk function?)
Solution: L(θ, a) = [θ − d(k)]2
R(θ, d(k)) = Ex (θ − d(k))2 = E(θ2 − 2θd(k) + d(k)2 )
= E(θ2 − 2θkS 2 + (kS 2 )2 ) = θ2 − 2θkE(S 2 ) + k 2 E(S 2 )2

11/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples. Bayes theorem
Prior distribution
Admissibility

Let x1 , x2 , ..., xn be a random sample from a N(µ, σ 2 ). consider the


class of estimators defined as
n
k X
d(k) = kS 2 = (xi − x̄)2
n−1
j=1

which decision rule minimizes the risk function in estimating θ = σ 2


under the squared loss function L(θ, a) = (θ − a)2 (which estimator
has the least risk function?)
Solution: L(θ, a) = [θ − d(k)]2
R(θ, d(k)) = Ex (θ − d(k))2 = E(θ2 − 2θd(k) + d(k)2 )
= E(θ2 − 2θkS 2 + (kS 2 )2 ) = θ2 − 2θkE(S 2 ) + k 2 E(S 2 )2
Since x1 , x2 , ..., xn is a random variable from N(µ, σ 2 ) it follows that
n  2
X xi − x̄
Q = ∼ χ2 (n − 1)
σ
i=1
σ2 Q
E(Q) = n − 1, var (Q) = 2(n − 1), S 2 = 11/37
n−1
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.. Bayes theorem
Prior distribution
Admissibility

It can be shown quite easily that


σ2
E(S 2 ) = E(Q) = σ 2
n−1
 2 2
σ σ4 2σ 4
var (S 2 ) = var (Q) = 2
var (Q) =
n−1 (n − 1) n−1
4
 
2σ n+1
E[(S 2 )2 ] = var (S 2 ) + [E(S 2 )]2 = + σ4 = θ2
n−1 n−1

12/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.. Bayes theorem
Prior distribution
Admissibility

It can be shown quite easily that


σ2
E(S 2 ) = E(Q) = σ 2
n−1
 2 2
σ σ4 2σ 4
var (S 2 ) = var (Q) = 2
var (Q) =
n−1 (n − 1) n−1
4
 
2σ n+1
E[(S 2 )2 ] = var (S 2 ) + [E(S 2 )]2 = + σ4 = θ2
n−1 n−1
Therefore
2θ2
 
2 2 2
R(θ, d(k)) = θ − 2kθ + k + θ2
n−1

12/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.. Bayes theorem
Prior distribution
Admissibility

It can be shown quite easily that


σ2
E(S 2 ) = E(Q) = σ 2
n−1
 2 2
σ σ4 2σ 4
var (S 2 ) = var (Q) = 2
var (Q) =
n−1 (n − 1) n−1
4
 
2σ n+1
E[(S 2 )2 ] = var (S 2 ) + [E(S 2 )]2 = + σ4 = θ2
n−1 n−1
Therefore
2θ2
 
2 2 2
R(θ, d(k)) = θ − 2kθ + k + θ2
n−1
Minimizing R(θ, d(k)) with respect to k, we get
2θ2
 
′ 2 2 n−1
R (θ, d(k)) = −2θ + 2k + θ ,⇛ k =
n−1 n+1
Pn 2
(x
j=1 j − x̄)
∴ d(k) =
n+1 12/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples... Bayes theorem
Prior distribution
Admissibility

Consider the loss function


Θ, D a1 a2
L(θ, a) = θ1 0.0 1.0
θ2 1.0 0.0

Let x be a binary random variable taking values 0 or 1 with the


following probabilities

13/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples... Bayes theorem
Prior distribution
Admissibility

Consider the loss function


Θ, D a1 a2
L(θ, a) = θ1 0.0 1.0
θ2 1.0 0.0

Let x be a binary random variable taking values 0 or 1 with the


following probabilities
The decision rules are defined as
 
a1 , if x = 0; a2 , if x = 0;
d1 (x) = d2 (x) =
a2 , if x = 1. a1 , if x = 1.

13/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples... Bayes theorem
Prior distribution
Admissibility

Consider the loss function


Θ, D a1 a2
L(θ, a) = θ1 0.0 1.0
θ2 1.0 0.0

Let x be a binary random variable taking values 0 or 1 with the


following probabilities
The decision rules are defined as
 
a1 , if x = 0; a2 , if x = 0;
d1 (x) = d2 (x) =
a2 , if x = 1. a1 , if x = 1.

Determine the risk function for these decision rules. Which decision
rule is minimax?. Which decision rule is admissible?

13/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples... Bayes theorem
Prior distribution
Admissibility

Now R(θ, di (x)) = E[L(θ, di (x))], for i = 1, 2.

14/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples... Bayes theorem
Prior distribution
Admissibility

Now R(θ, di (x)) = E[L(θ, di (x))], for i = 1, 2.


We determine the risk function for each θ
1
X
R(θ1 , d1 (x)) = Ex (L(θ1 , d1 (x))) = L(θ1 , d1 (x))pr (X = x|θ1 )
x=0
= L(θ1 , a1 )pr (X = 0|θ1 ) + L(θ1 , a2 )pr (x = 1|θ1 )
= 0∗1+1∗0=0
1
X
R(θ2 , d1 (x)) = Ex (L(θ2 , d1 (x))) = L(θ2 , d1 (x))pr (X = x|θ2 )
x=0
= L(θ2 , a1 )pr (X = 0|θ2 ) + L(θ2 , a2 )pr (X = 1|θ2 )
= 1 ∗ 1/2 + 0 ∗ 1/2 = 1/2
1
X
R(θ2 , d2 (x)) = Ex (L(θ2 , d2 (x))) = L(θ2 , d2 (x))pr (X = x|θ2 )
x=0
= L(θ2 , a2 )pr (X = 0|θ2 ) + L(θ2 , a1 )pr (X = 1|θ2 )
= 0 ∗ 1/2 + 1 ∗ 1/2 = 1/2
14/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.... Bayes theorem
Prior distribution
Admissibility

Minimax decision

max R(θ, di (x)) = min max R(θ, d(x))


θ d∈D θ
1
max R(θ, d1 (x)) = max R(θ, d2 (x)) = 1
θ 2 θ
1
min max R(θ, d(x)) = max R(θ, d1 (x)) =
d∈D θ θ 2

therefore d1 (x) is the minimax decision.

15/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.... Bayes theorem
Prior distribution
Admissibility

Minimax decision

max R(θ, di (x)) = min max R(θ, d(x))


θ d∈D θ
1
max R(θ, d1 (x)) = max R(θ, d2 (x)) = 1
θ 2 θ
1
min max R(θ, d(x)) = max R(θ, d1 (x)) =
d∈D θ θ 2

therefore d1 (x) is the minimax decision.


Admissibility. For θ = θ1 , R(θ1 , d1 (x)) < R(θ1 , d2 (x))

15/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Decision Rules
Examples.... Bayes theorem
Prior distribution
Admissibility

Minimax decision

max R(θ, di (x)) = min max R(θ, d(x))


θ d∈D θ
1
max R(θ, d1 (x)) = max R(θ, d2 (x)) = 1
θ 2 θ
1
min max R(θ, d(x)) = max R(θ, d1 (x)) =
d∈D θ θ 2

therefore d1 (x) is the minimax decision.


Admissibility. For θ = θ1 , R(θ1 , d1 (x)) < R(θ1 , d2 (x))
For θ = θ2 , R(θ2 , d1 (x)) = R(θ2 , d2 (x)) Hence d1 (x) dominates
d2 (x). Hence d1 (x) is admissible.

15/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian theory Bayes theorem
Prior distribution
Loss functions

The fundamental difference between Bayesian and classical method


is that the parameter θ is considered to be a random variable
Bayesian statistics.

16/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian theory Bayes theorem
Prior distribution
Loss functions

The fundamental difference between Bayesian and classical method


is that the parameter θ is considered to be a random variable
Bayesian statistics.
In classical statistics θ is fixed but unknown quantity

16/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian theory Bayes theorem
Prior distribution
Loss functions

The fundamental difference between Bayesian and classical method


is that the parameter θ is considered to be a random variable
Bayesian statistics.
In classical statistics θ is fixed but unknown quantity
An advantage of bayesian statistics is that the it enables the
researcher to make use of any information that we already have
about the situation under investigation.

16/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian theory Bayes theorem
Prior distribution
Loss functions

The fundamental difference between Bayesian and classical method


is that the parameter θ is considered to be a random variable
Bayesian statistics.
In classical statistics θ is fixed but unknown quantity
An advantage of bayesian statistics is that the it enables the
researcher to make use of any information that we already have
about the situation under investigation.
Often researchers investigating an unknown population parameter
have information available from other sources in advance of the
study that provides a strong indication of what parameter is likely to
take.

16/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian theory Bayes theorem
Prior distribution
Loss functions

The fundamental difference between Bayesian and classical method


is that the parameter θ is considered to be a random variable
Bayesian statistics.
In classical statistics θ is fixed but unknown quantity
An advantage of bayesian statistics is that the it enables the
researcher to make use of any information that we already have
about the situation under investigation.
Often researchers investigating an unknown population parameter
have information available from other sources in advance of the
study that provides a strong indication of what parameter is likely to
take.
Consider a case where an insurance company is reviewing its
premium rates for a particular type of a policy and has access to
results from other insurer, as well as from his policy holders, this
may consider use of Bayesian statistics.
16/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayes theorem. Bayes theorem
Prior distribution
Loss functions

Let A and B be events. If D1 , D2 , ..., Dk consists of a partition of a


sample space say S and Pr {Di } =
̸ 0, ∀ i Then for any set A in S
such that Pr {A} =
̸ 0
Pr (A|Dr )Pr (Dr )
Pr {Dr |A} =
Pr (A)
Pr (A|Dr )Pr (Dr )
= Pk , r = 1, 2, ..., k
i=1 Pr (A|Dr )Pr (Dr ))

17/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayes theorem. Bayes theorem
Prior distribution
Loss functions

Let A and B be events. If D1 , D2 , ..., Dk consists of a partition of a


sample space say S and Pr {Di } =
̸ 0, ∀ i Then for any set A in S
such that Pr {A} =
̸ 0
Pr (A|Dr )Pr (Dr )
Pr {Dr |A} =
Pr (A)
Pr (A|Dr )Pr (Dr )
= Pk , r = 1, 2, ..., k
i=1 Pr (A|Dr )Pr (Dr ))

Note that for any two sets A and B


Pr (A ∩ B) Pr (A ∩ B)
Pr (A|B) = ; Pr (B|A) =
Pr (B) Pr (A)
Pr (A ∩ B) = Pr (A|B)Pr (B); Pr (A ∩ B) = Pr (B|A)Pr (A)
1
∴ Pr (A|B) = Pr (B|A)Pr (A)
Pr (B)
Pr (A|B) ∝ Pr (B|A)Pr (A)

17/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference Bayes theorem
Prior distribution
Loss functions

Consider a general problem where we have data X and require


inference about parameter θ.

18/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference Bayes theorem
Prior distribution
Loss functions

Consider a general problem where we have data X and require


inference about parameter θ.
In a Bayesian analysis θ is unknown and given a density function
f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

18/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference Bayes theorem
Prior distribution
Loss functions

Consider a general problem where we have data X and require


inference about parameter θ.
In a Bayesian analysis θ is unknown and given a density function
f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

Typically θ and x are continuous [e,g x|θ ∼ N(θ, σ 2 )] σ 2 known; θ


continuous or θ continuous and x discrete. eg [x|θ ∼ bin(n, θ).] In
exceptional cases θ could be discrete.

18/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference Bayes theorem
Prior distribution
Loss functions

Consider a general problem where we have data X and require


inference about parameter θ.
In a Bayesian analysis θ is unknown and given a density function
f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

Typically θ and x are continuous [e,g x|θ ∼ N(θ, σ 2 )] σ 2 known; θ


continuous or θ continuous and x discrete. eg [x|θ ∼ bin(n, θ).] In
exceptional cases θ could be discrete.
The Bayesian method comprises of the following principle steps,
that is prior density, likelihood function, posterior distribution and
statistical inference.
18/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference steps
Bayes theorem
Prior distribution
Loss functions

Note that In a Bayesian analysis θ is unknown and given a density


function f (θ), = π(θ) we have

π(θ|x) ∝ f (x|θ)π(θ)

Prior: Obtain prior density say π(θ). Express knowledge about θ


prior to observing data

19/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference steps
Bayes theorem
Prior distribution
Loss functions

Note that In a Bayesian analysis θ is unknown and given a density


function f (θ), = π(θ) we have

π(θ|x) ∝ f (x|θ)π(θ)

Prior: Obtain prior density say π(θ). Express knowledge about θ


prior to observing data
Likelihood: Obtain likelihood function f (x|θ). This step simply
describes process of giving rise to data x in terms of θ

19/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference steps
Bayes theorem
Prior distribution
Loss functions

Note that In a Bayesian analysis θ is unknown and given a density


function f (θ), = π(θ) we have

π(θ|x) ∝ f (x|θ)π(θ)

Prior: Obtain prior density say π(θ). Express knowledge about θ


prior to observing data
Likelihood: Obtain likelihood function f (x|θ). This step simply
describes process of giving rise to data x in terms of θ
Posterior: Apply Bayes theorem to describe posterior
f (θ|x) = π(θ|x). Express knowledge about θ after observing the
data.

19/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Bayesian Inference steps
Bayes theorem
Prior distribution
Loss functions

Note that In a Bayesian analysis θ is unknown and given a density


function f (θ), = π(θ) we have

π(θ|x) ∝ f (x|θ)π(θ)

Prior: Obtain prior density say π(θ). Express knowledge about θ


prior to observing data
Likelihood: Obtain likelihood function f (x|θ). This step simply
describes process of giving rise to data x in terms of θ
Posterior: Apply Bayes theorem to describe posterior
f (θ|x) = π(θ|x). Express knowledge about θ after observing the
data.
Inference: Derive appropriate statements from the posterior
distribution. eg. Point estimates, probability of an hypothesis.

19/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)

20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)
Now f (x, y |θ) = f (y |x, θ)f (x|θ)

20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)
Now f (x, y |θ) = f (y |x, θ)f (x|θ)
So that
f (θ|x, y ) ∝ f (y |x, θ)f (x|θ)f (θ)
∝ f (y |x, θ)f (θ|x)

20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)
Now f (x, y |θ) = f (y |x, θ)f (x|θ)
So that
f (θ|x, y ) ∝ f (y |x, θ)f (x|θ)f (θ)
∝ f (y |x, θ)f (θ|x)
First update by x and then by y .

20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)
Now f (x, y |θ) = f (y |x, θ)f (x|θ)
So that
f (θ|x, y ) ∝ f (y |x, θ)f (x|θ)f (θ)
∝ f (y |x, θ)f (θ|x)
First update by x and then by y .
Note that if x and y are conditionally independent given θ then
f (x, y |θ) = f (x|θ)f (x|θ), i.e. f (y |x, θ) = f (y |θ)

20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Sequential data updates
Bayes theorem
Prior distribution
Loss functions

Suppose we have two sources of data x and y . We can add data


sequentially. f (θ|x, y ) ∝ f (x, y |θ)f (θ)
Now f (x, y |θ) = f (y |x, θ)f (x|θ)
So that
f (θ|x, y ) ∝ f (y |x, θ)f (x|θ)f (θ)
∝ f (y |x, θ)f (θ|x)
First update by x and then by y .
Note that if x and y are conditionally independent given θ then
f (x, y |θ) = f (x|θ)f (x|θ), i.e. f (y |x, θ) = f (y |θ)
For several sources of data x, y , z, t we update it sequentially
f (θ|x) ∝ f (x|θ)π(θ)
f (θ|x, y ) ∝ f (y |x, θ)π(θ|x)
f (θ|x, y , z) ∝ f (z|y , x, θ)π(θ|x, y )
f (θ|x, y , z, t) ∝ f (t|z, y , x, θ)π(θ|x, y , z)
20/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are

21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are


Squared loss function where L(θ, a) = (θ − a)2 . Its commonly
used in estimation of a location parameters

21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are


Squared loss function where L(θ, a) = (θ − a)2 . Its commonly
used in estimation of a location parameters
Absolute error loss function |L(θ, a)| = |θ − a|. can be used in
similar situations like squared loss function.

21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are


Squared loss function where L(θ, a) = (θ − a)2 . Its commonly
used in estimation of a location parameters
Absolute error loss function |L(θ, a)| = |θ − a|. can be used in
similar situations like squared loss function.
Loss functions used in estimating scale parameters such as
variance include
θ−a 2
 
L(θ, a) =
θ

21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are


Squared loss function where L(θ, a) = (θ − a)2 . Its commonly
used in estimation of a location parameters
Absolute error loss function |L(θ, a)| = |θ − a|. can be used in
similar situations like squared loss function.
Loss functions used in estimating scale parameters such as
variance include
θ−a 2
 
L(θ, a) =
θ
Another type of loss function for scale parameters
a a
L(θ, a) = − loge −1
θ θ

21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Types of loss functionsBayes theorem
Prior distribution
Loss functions

Typical loss function used in Bayesian theory are


Squared loss function where L(θ, a) = (θ − a)2 . Its commonly
used in estimation of a location parameters
Absolute error loss function |L(θ, a)| = |θ − a|. can be used in
similar situations like squared loss function.
Loss functions used in estimating scale parameters such as
variance include
θ−a 2
 
L(θ, a) =
θ
Another type of loss function for scale parameters
a a
L(θ, a) = − loge −1
θ θ
0 − 1 type of loss function commonly used for hypothesis
testing
21/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples
Bayes theorem
Prior distribution
Loss functions

For example 0 − 1 loss function type is used in hypothesis testing.


H0 : θ = θ0 vs H1 : θ = θa
a1 = reject a = do not reject
2
0, if i = 2;
L(θ0 , ai ) =
1, if i = 1.

22/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples
Bayes theorem
Prior distribution
Loss functions

For example 0 − 1 loss function type is used in hypothesis testing.


H0 : θ = θ0 vs H1 : θ = θa
a1 = reject a = do not reject
2
0, if i = 2;
L(θ0 , ai ) =
1, if i = 1.
Example: Let x1 , x2 , ..., xn be a random sample from a normal
distribution N(µ, σ 2 ). Let d(x) = x̄. Find the risk function for this
decision rule under the squared loss function. Let θ ̸= µ

22/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples
Bayes theorem
Prior distribution
Loss functions

For example 0 − 1 loss function type is used in hypothesis testing.


H0 : θ = θ0 vs H1 : θ = θa
a1 = reject a = do not reject
2
0, if i = 2;
L(θ0 , ai ) =
1, if i = 1.
Example: Let x1 , x2 , ..., xn be a random sample from a normal
distribution N(µ, σ 2 ). Let d(x) = x̄. Find the risk function for this
decision rule under the squared loss function. Let θ ̸= µ
σ2
Solution: Obviously E(x̄) = µ and var (x̄) = n

L(θ, a) = (θ − d(x))2 = (θ − x̄)2


R(θ, d(x)) = Ex (L(θ, a)) = Ex (θ − x̄)2
= E[x̄ − µ]2 + 2(µ − θ)E(x̄ − µ) + (µ − θ)2
σ2
= var (x̄) + (µ − θ)2 = + (µ − θ)2
n

22/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples
Bayes theorem
Prior distribution
Loss functions

For example 0 − 1 loss function type is used in hypothesis testing.


H0 : θ = θ0 vs H1 : θ = θa
a1 = reject a = do not reject
2
0, if i = 2;
L(θ0 , ai ) =
1, if i = 1.
Example: Let x1 , x2 , ..., xn be a random sample from a normal
distribution N(µ, σ 2 ). Let d(x) = x̄. Find the risk function for this
decision rule under the squared loss function. Let θ ̸= µ
σ2
Solution: Obviously E(x̄) = µ and var (x̄) = n

L(θ, a) = (θ − d(x))2 = (θ − x̄)2


R(θ, d(x)) = Ex (L(θ, a)) = Ex (θ − x̄)2
= E[x̄ − µ]2 + 2(µ − θ)E(x̄ − µ) + (µ − θ)2
σ2
= var (x̄) + (µ − θ)2 = + (µ − θ)2
n
σ2
If θ = µ then R(θ, d(x)) = n 22/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples.
Bayes theorem
Prior distribution
Loss functions

Example: Let x1 , x2 , ..., xn be a random sample from a normal


distribution N(µ, σ 2 ). let
a a 
L(θ, a) = − loge −1
θ θ

23/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples.
Bayes theorem
Prior distribution
Loss functions

Example: Let x1 , x2 , ..., xn be a random sample from a normal


distribution N(µ, σ 2 ). let
a a 
L(θ, a) = − loge −1
θ θ
Find the decision rule that minimizes the risk function within the
class d(k) = kS 2 for estimating σ 2 = θ

23/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples.
Bayes theorem
Prior distribution
Loss functions

Example: Let x1 , x2 , ..., xn be a random sample from a normal


distribution N(µ, σ 2 ). let
a a 
L(θ, a) = − loge −1
θ θ
Find the decision rule that minimizes the risk function within the
class d(k) = kS 2 for estimating σ 2 = θ
Solution:Let Qσ 2 = (xi − x̄)2 , E(Q) = n − 1, var (Q) = 2(n − 1)
P
a a 
L(θ, a) = − loge −1
θ θ h a a i
R(θ, d(x)) = Ex (L(θ, a)) = Ex − loge −1
θ θ
kS 2

k 2
= ES − E loge −1
θ θ
S2
 
∴ R(θ, d(k)) = k − loge k − E loge −1
θ

23/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples.
Bayes theorem
Prior distribution
Loss functions

Example: Let x1 , x2 , ..., xn be a random sample from a normal


distribution N(µ, σ 2 ). let
a a 
L(θ, a) = − loge −1
θ θ
Find the decision rule that minimizes the risk function within the
class d(k) = kS 2 for estimating σ 2 = θ
Solution:Let Qσ 2 = (xi − x̄)2 , E(Q) = n − 1, var (Q) = 2(n − 1)
P
a a 
L(θ, a) = − loge −1
θ θ h a a i
R(θ, d(x)) = Ex (L(θ, a)) = Ex − loge −1
θ θ
kS 2

k 2
= ES − E loge −1
θ θ
S2
 
∴ R(θ, d(k)) = k − loge k − E loge −1
θ
1
Minimizing with respect to k we get R(θ, d(x)) = 1 − k
23/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples..
Bayes theorem
Prior distribution
Loss functions

Equating to zero and solving for k we get k = 1

24/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction
Bayesian Inference
Loss functions,Examples..
Bayes theorem
Prior distribution
Loss functions

Equating to zero and solving for k we get k = 1


under the loss function
a a 
L(θ, a) = − loge −1
θ θ
the decision rule that minimises the class d(k) = kS 2 is
n
1 X
d(1) = S 2 = (xj − x̄)2
n−1
j=1

24/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

In a Bayesian analysis θ is unknown and given a density function


f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

25/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

In a Bayesian analysis θ is unknown and given a density function


f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

π(θ) is known as prior distribution.

25/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

In a Bayesian analysis θ is unknown and given a density function


f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

π(θ) is known as prior distribution.


A prior distribution of a parameter is the probability distribution
that presents ones uncertainty about a parameter before the current
data are examined.

25/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

In a Bayesian analysis θ is unknown and given a density function


f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

π(θ) is known as prior distribution.


A prior distribution of a parameter is the probability distribution
that presents ones uncertainty about a parameter before the current
data are examined.
Multiplying the prior distribution and the likelihood function
together leads to posterior distribution of the parameter

25/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

In a Bayesian analysis θ is unknown and given a density function


f (θ), we have

f (x|θ)f (θ)
f (θ|x) =
f (x)
∝ f (x|θ)f (θ)
π(θ|x) ∝ f (x|θ)π(θ)

π(θ) is known as prior distribution.


A prior distribution of a parameter is the probability distribution
that presents ones uncertainty about a parameter before the current
data are examined.
Multiplying the prior distribution and the likelihood function
together leads to posterior distribution of the parameter
We use the posterior distribution to carry out all inference. One
cannot carry out modeling without using prior distribution.
25/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior
Non-informative priors(vague,diffuse,flat prior). A prior distribution
is non-informative if the prior is ”flat” relative to the likelihood
function. It is non-informative if it has minimal impact on the
posterior distribution of θ.

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior
Non-informative priors(vague,diffuse,flat prior). A prior distribution
is non-informative if the prior is ”flat” relative to the likelihood
function. It is non-informative if it has minimal impact on the
posterior distribution of θ.
In some cases non-informative prior can lead to an improper
posterior (non-integrable posterior density). One cannot make
inference with improper posterior distributions.

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior
Non-informative priors(vague,diffuse,flat prior). A prior distribution
is non-informative if the prior is ”flat” relative to the likelihood
function. It is non-informative if it has minimal impact on the
posterior distribution of θ.
In some cases non-informative prior can lead to an improper
posterior (non-integrable posterior density). One cannot make
inference with improper posterior distributions.
A common choice for non-informative prior is the flat prior that
assigns equal likelihood on all possible values of the parameter.

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior
Non-informative priors(vague,diffuse,flat prior). A prior distribution
is non-informative if the prior is ”flat” relative to the likelihood
function. It is non-informative if it has minimal impact on the
posterior distribution of θ.
In some cases non-informative prior can lead to an improper
posterior (non-integrable posterior density). One cannot make
inference with improper posterior distributions.
A common choice for non-informative prior is the flat prior that
assigns equal likelihood on all possible values of the parameter.
R
A prior π(θ) is said to be improper if Ω π(θ)dθ = ∞

26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Prior distribution Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

There are several types of prior in general.


They include but not limited to true priors,non-informative priors,
improper priors,informative priors conjugate priors and Jeffrey’s prior
Non-informative priors(vague,diffuse,flat prior). A prior distribution
is non-informative if the prior is ”flat” relative to the likelihood
function. It is non-informative if it has minimal impact on the
posterior distribution of θ.
In some cases non-informative prior can lead to an improper
posterior (non-integrable posterior density). One cannot make
inference with improper posterior distributions.
A common choice for non-informative prior is the flat prior that
assigns equal likelihood on all possible values of the parameter.
R
A prior π(θ) is said to be improper if Ω π(θ)dθ = ∞
Uniform prior distribution on the real line π(θ) ∝ 1, − ∞ < θ < ∞
is an improper prior
26/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Informative Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Is a prior that is not dominated by the likelihood and thus has an


impact on the posterior distribution.

27/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Informative Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Is a prior that is not dominated by the likelihood and thus has an


impact on the posterior distribution.
These types of prior distribution must be specified with care.

27/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Informative Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Is a prior that is not dominated by the likelihood and thus has an


impact on the posterior distribution.
These types of prior distribution must be specified with care.
Proper use of prior distribution illustrates the power of Bayesian
method.

27/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Informative Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Is a prior that is not dominated by the likelihood and thus has an


impact on the posterior distribution.
These types of prior distribution must be specified with care.
Proper use of prior distribution illustrates the power of Bayesian
method.
Information gathered from previous studies, past experience or
expert opinion can be combined with current information in a
natural way.

27/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

A prior is said to be a conjugate prior for a family of distributions if


the posterior distributions are from the same family. This means
that the form of the posterior has the same distributional form as
the prior

28/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

A prior is said to be a conjugate prior for a family of distributions if


the posterior distributions are from the same family. This means
that the form of the posterior has the same distributional form as
the prior
Commonly used conjugate prior/likelihood include
normal/normal;gamma/poisson;gamma/gamma;gamma/beta

28/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

A prior is said to be a conjugate prior for a family of distributions if


the posterior distributions are from the same family. This means
that the form of the posterior has the same distributional form as
the prior
Commonly used conjugate prior/likelihood include
normal/normal;gamma/poisson;gamma/gamma;gamma/beta
The development of conjugate prior was partially driven by a desire
for computational convenience.

28/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

A prior is said to be a conjugate prior for a family of distributions if


the posterior distributions are from the same family. This means
that the form of the posterior has the same distributional form as
the prior
Commonly used conjugate prior/likelihood include
normal/normal;gamma/poisson;gamma/gamma;gamma/beta
The development of conjugate prior was partially driven by a desire
for computational convenience.
Example: Let x1 , x2 , ..., xn be a random variable from an exponential
distribution f (x|θ) = θe −θx , x > 0, θ > 0 show that the gamma
distribution with parameters α and β is a conjugate prior for this
density.

28/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Solution: f (x|θ) = θe −θx therefore


n
Y Pn
L(x|θ) = θe −θxi = θn e −θ j=1 xi
= θn e −nx̄
j=1
α
β
π(θ) = θα−1 e −βθ
Γ(α)
f (θ|x) ∝ θn e −θx̄n θα−1 e θβ
∝ θn+α−1 e −θ(β+nx̄) ∼ Gamma(n + α, β + nx̄)

thus prior and posterior are from the same parent distribution hence
conjugate prior.

29/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Solution: f (x|θ) = θe −θx therefore


n
Y Pn
L(x|θ) = θe −θxi = θn e −θ j=1 xi
= θn e −nx̄
j=1
α
β
π(θ) = θα−1 e −βθ
Γ(α)
f (θ|x) ∝ θn e −θx̄n θα−1 e θβ
∝ θn+α−1 e −θ(β+nx̄) ∼ Gamma(n + α, β + nx̄)

thus prior and posterior are from the same parent distribution hence
conjugate prior.
If
1
θα−1 e −( β )
θ
π(θ) =
Γαβ α
The corresponding posterior distribution has the same gamma
distribution but with different parameters Gamma(n + α, nx̄ + β1 )
29/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors.. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Example:Let x1 , x2 , ..., xk be a random sample from a binomial


distribution with parameters n and θi.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x

30/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors.. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Example:Let x1 , x2 , ..., xk be a random sample from a binomial


distribution with parameters n and θi.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior

30/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors.. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Example:Let x1 , x2 , ..., xk be a random sample from a binomial


distribution with parameters n and θi.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior
Solution:
Qk
j=1 f (xj |θ)π(θ)
fx (θ|x) =
f (x)
k
Y
∝ f (xj |θ)π(θ) = θk x̄ (1 − θ)nk−k x̄ θα−1 (1 − θ)β−1
j=1

∝ θk x̄+α−1 (1 − θ)nk−k x̄+β−1

30/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Conjugate Priors.. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Example:Let x1 , x2 , ..., xk be a random sample from a binomial


distribution with parameters n and θi.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior
Solution:
Qk
j=1 f (xj |θ)π(θ)
fx (θ|x) =
f (x)
k
Y
∝ f (xj |θ)π(θ) = θk x̄ (1 − θ)nk−k x̄ θα−1 (1 − θ)β−1
j=1

∝ θk x̄+α−1 (1 − θ)nk−k x̄+β−1

The posterior is an incomplete Beta distribution with parameters


α′ = α + k x̄ and β ′ = nk + β − k x̄
30/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Consider the following distribution f (x|θ) obtain the posterior


distribution of θ given

 X, Θ θ1 θ2
0.8, if θ = θ1 ; x1 0.6 0.1
π(θ) = f (x|θ) =
0.2, if θ = θ2 . x2 0.3 0.2
x3 0.1 0.7

31/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Consider the following distribution f (x|θ) obtain the posterior


distribution of θ given

 X, Θ θ1 θ2
0.8, if θ = θ1 ; x1 0.6 0.1
π(θ) = f (x|θ) =
0.2, if θ = θ2 . x2 0.3 0.2
x3 0.1 0.7
solution:
f (x1 ) = f (x1 |θ1 )π(θ1 ) + f (x1 |θ2 )π(θ2 ) = 0.6(.8) + .2(.1) = 0.5
f (x2 ) = f (x2 |θ1 )π(θ1 ) + f (x2 |θ2 )π(θ2 ) = 0.3(.8) + .2(.2) = 0.28
= f (x3 |θ1 )π(θ1 ) + f (x3 |θ2 )π(θ2 ) = 0.1(.8) + .2(.7) = 0.22
f (x3 )
π(θ|x) = π(θ = θ1 |x = xi ); π(θ = θ2 |x = xi ), i = 1, 2, 3.
f (x|θ)π(θ) .6(.8)
π(θ|x) = ; π(θ = θ1 |x = x1 ) = = .96
f (x) .5
f (x1 |θ2 )π(θ2 ) .1(.2)
π(θ = θ2 |x = x1 ) = = = .04
f (x1 ) .5
31/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Similarly
f (x1 |θ2 )π(θ2 ) .1(.2)
π(θ = θ2 |x = x1 ) = = = .04
f (x1 ) .5
f (x2 |θ1 )π(θ1 ) .3(.8)
π(θ = θ1 |x = x2 ) = = = .86
f (x2 ) .28
f (x2 |θ2 )π(θ2 ) .2(.2)
π(θ = θ2 |x = x2 ) = = = .14
f (x2 ) .28
f (x3 |θ1 )π(θ1 ) .1(.8)
π(θ = θ1 |x = x3 ) = = = .36
f (x3 ) .22
f (x3 |θ2 )π(θ2 ) .7(.2)
π(θ = θ2 |x = x3 ) = = = .64
f (x3 ) .22

32/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Similarly
f (x1 |θ2 )π(θ2 ) .1(.2)
π(θ = θ2 |x = x1 ) = = = .04
f (x1 ) .5
f (x2 |θ1 )π(θ1 ) .3(.8)
π(θ = θ1 |x = x2 ) = = = .86
f (x2 ) .28
f (x2 |θ2 )π(θ2 ) .2(.2)
π(θ = θ2 |x = x2 ) = = = .14
f (x2 ) .28
f (x3 |θ1 )π(θ1 ) .1(.8)
π(θ = θ1 |x = x3 ) = = = .36
f (x3 ) .22
f (x3 |θ2 )π(θ2 ) .7(.2)
π(θ = θ2 |x = x3 ) = = = .64
f (x3 ) .22
The resulting posterior distribution is summarized in table form
X, Θ θ1 θ2
x1 0.96 0.04
π(θ|x) =
x2 0.86 0.14
x3 0.36 0.64 32/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example... Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Let x ∼ N(θ, σ 2 ), σ 2 is assumed to be known, θ is unknown given


prior π(θ) ∼ N(θ0 , σ02 ). Show that f (θ|x) is conjugate prior.

33/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example... Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Let x ∼ N(θ, σ 2 ), σ 2 is assumed to be known, θ is unknown given


prior π(θ) ∼ N(θ0 , σ02 ). Show that f (θ|x) is conjugate prior.
Solution: We show that the resulting posterior and prior distribution
are from normal distribution f (θ|x) ∝ f (x|θ)π(θ)
n
"  2 # "  2
Y 1 1 xj − θ 1 1 θ − θ0
f (θ|x) ∝ √ exp − p exp −
2πσ 2 2 σ 2πσ02 2 σ0
j=1
 
n  2  2
1 X xj − θ 1 θ − θ0 
∝ exp − −
2 σ 2 σ0
j=1

1 nθ2 − 2θx̄n θ2 − 2θθ0


  
∝ exp − +
2 σ2 σ02
1 nσ02 + σ 2 nx̄σ02 + θ0 σ 2
    
2
∝ exp − θ − 2 θ
2 σ02 σ 2 nσ02 + σ 2
nx̄σ02 + θ0 σ 2 σ02 σ 2
∼ N(γ, β 2 ); where γ = 2 ; β2 =
nσ0 + σ 2 nσ02 + σ 2
33/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Jeffrey’s Prior Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Jeffrey’s prior is defined in terms of fischer information I (θ), which


tells us how well we can measure a parameter, given a certain
amount data.

34/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Jeffrey’s Prior Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Jeffrey’s prior is defined in terms of fischer information I (θ), which


tells us how well we can measure a parameter, given a certain
amount data.
p
Let πJ (θ) denote Jeffrey’s prior thus πJ (θ) ∝ detI (θ) where fisher
information is given by
 2 
d
I (θ) = −Eθ log e f (x|θ)
dθ2

34/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Jeffrey’s Prior Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Jeffrey’s prior is defined in terms of fischer information I (θ), which


tells us how well we can measure a parameter, given a certain
amount data.
p
Let πJ (θ) denote Jeffrey’s prior thus πJ (θ) ∝ detI (θ) where fisher
information is given by
 2 
d
I (θ) = −Eθ log e f (x|θ)
dθ2
Example: Suppose x is binomially distributed
x ∼ Bin(n, θ), 0 ≤ θ ≤ 1 determine its Jeffrey’s prior
 
n x
f (x|θ) = θ (1 − θ)n−x ⇒ loge f (x|θ) ∝ x ln θ + (n − x) ln(1 − θ)
x
d x n − x d2 x n−x
ln f (x|θ) = − ; 2
ln f (x|θ) = − 2 −
dθ θ 1 − θ dθ θ (1 − θ)2
nθ n − nθ n
I (θ) = + = ; πJ (θ) ∝ θ−1/2 (1 − θ)−1/2 ; ∴
θ2 (1 − θ)2 θ(1 − θ)
πJ (θ) = Beta(3/2, 3/2)
34/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

The resulting posterior f (θ|x) ∝ f (x|θ)πJ (θ)


 
3 3
f (θ|x) ∝ θx (1 − θ)n−x θ−1/2 (1 − θ)−1/2 ∼ Beta x − , n − x −
2 2

35/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

The resulting posterior f (θ|x) ∝ f (x|θ)πJ (θ)


 
3 3
f (θ|x) ∝ θx (1 − θ)n−x θ−1/2 (1 − θ)−1/2 ∼ Beta x − , n − x −
2 2
Let y1 , y2 , ..., yn be from N(θ, σ 2 ) suppose that θ, σ 2 are unknown.
Derive Jeffrey’s prior for θ and σ 2 and determine the corresponding
posterior distribution.

35/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

The resulting posterior f (θ|x) ∝ f (x|θ)πJ (θ)


 
3 3
f (θ|x) ∝ θx (1 − θ)n−x θ−1/2 (1 − θ)−1/2 ∼ Beta x − , n − x −
2 2
Let y1 , y2 , ..., yn be from N(θ, σ 2 ) suppose that θ, σ 2 are unknown.
Derive Jeffrey’s prior for θ and σ 2 and determine the corresponding
posterior distribution.
Solution:We get the log likelihood estimate first then find the
respective derivatives, let σ 2 = α
n  2 !
2
Y 1 1 yj − θ
f (y |θ, σ ) = √ exp −
2πσ 2 2 σ
j=1
 
  n2 n  2
1 1 X y j − θ
= exp − 
2πσ 2 2 σ
j=1
n  2
2 n 1 X yj − θ
ln f (y |θ, σ ) = − ln(2πσ 2 ) − =L
2 2 σ
j=1 35/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Differentiating with respect to respective variables we have I (θ)


∂2L ∂2L
 
∂θ 2 ∂θ∂α
I (θ) = −E  
∂2L ∂2L
∂α∂θ ∂α2
Pn
j=1 (yj −θ)
 
− αn − α2
= −E 
 
Pn Pn 
2
j=1 (yj −θ) n j=1 (yj −θ)
− α2 2α2 − α3
n
 
α 0 1 1
det(I (θ)) = n = , ∴ πJ (θ) ∝
0 2α2 2α3 σ3

36/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

Example. Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

Differentiating with respect to respective variables we have I (θ)


∂2L ∂2L
 
∂θ 2 ∂θ∂α
I (θ) = −E  
∂2L ∂2L
∂α∂θ ∂α2
Pn
j=1 (yj −θ)
 
− αn − α2
= −E 
 
Pn Pn 
2
j=1 (yj −θ) n j=1 (yj −θ)
− α2 2α2 − α3
n
 
α 0 1 1
det(I (θ)) = n = , ∴ πJ (θ) ∝
0 2α2 2α3 σ3

the resulting posterior is f (θ|y ) ∝ f (y |θ)πJ (θ)


" n
!#
1 1 X
2
f (θ|y ) ∝ exp − 2 n(θ − ȳ ) + (yi − ȳ )2
(σ 2 )(n+3)/2 2σ
i=1

36/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

General exercise Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

If a random variable x has the gamma distribution


x ∼ Gamma(α, β) then the random variable Z = 1/X has the
inverted gamma distribution i.e. Z ∼ IG (α, β) with the density
( β 
α 1 1+β − α
e z , α > 0, β > 0;
f (z|α, β) = β z
0, elsewhere.

37/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

General exercise Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

If a random variable x has the gamma distribution


x ∼ Gamma(α, β) then the random variable Z = 1/X has the
inverted gamma distribution i.e. Z ∼ IG (α, β) with the density
( β 
α 1 1+β − α
e z , α > 0, β > 0;
f (z|α, β) = β z
0, elsewhere.

It can be shown that


α2
 
α
E(z) = , β > 1, var (z) = ; β>2
β (β − 1)2 (β − 2)

37/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023
Introduction Informative Priors

General exercise Bayes theorem


Prior distribution
Conjugate Prior
Jeffery’s Prior

If a random variable x has the gamma distribution


x ∼ Gamma(α, β) then the random variable Z = 1/X has the
inverted gamma distribution i.e. Z ∼ IG (α, β) with the density
( β 
α 1 1+β − α
e z , α > 0, β > 0;
f (z|α, β) = β z
0, elsewhere.

It can be shown that


α2
 
α
E(z) = , β > 1, var (z) = ; β>2
β (β − 1)2 (β − 2)

Let the random variable Y and Z with Y ∼ G (b, α) and


Z ∼ G (b, β) be independent, then the random variable
X = Y /(Y + Z ) has beta distribution Beta(α, β), thus
(
Γ(β+β) α−1
Γ(α)Γ(β) x (1 − x)β−1 , 0 < x < 1, α, β > 0.;
fx (x|α, β) =
0, elsewhere.
37/37
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian theory November 16, 2023

You might also like