0% found this document useful (0 votes)
14 views114 pages

Bayesian Inference and Decision Rules

The document discusses Bayes decision rule, Bayes estimator, minimax procedures, and predictions within the context of Bayesian inference. It outlines the concepts of Bayes risk, posterior risk function, and how the Bayes estimator minimizes expected loss. The content is presented by Ivivi Joseph Mwaniki from the University of Nairobi and includes mathematical formulations and definitions related to decision theory.

Uploaded by

erobamanuela
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views114 pages

Bayesian Inference and Decision Rules

The document discusses Bayes decision rule, Bayes estimator, minimax procedures, and predictions within the context of Bayesian inference. It outlines the concepts of Bayes risk, posterior risk function, and how the Bayes estimator minimizes expected loss. The content is presented by Ivivi Joseph Mwaniki from the University of Nairobi and includes mathematical formulations and definitions related to decision theory.

Uploaded by

erobamanuela
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Bayes decision rule

Bayes Estimator
Minimax procedures
Predictions

Bayesian Inference

Ivivi Joseph Mwaniki


Department of Mathematics, University of Nairobi
Email: jimwaniki@[Link]

SOMAS ENROLMENT KEY hcyup7

December 2, 2023

1/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator

Outline Minimax procedures


Predictions

1 Bayes decision rule


Bayes risk
2 Bayes Estimator
Squared loss functions
Absolute loss function
Weighted squared loss function
3 Minimax procedures
minimax Bayes
Admissibility
Credibility intervals
4 Predictions
Posterior Predictive distribution
2/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Bayes risk
Bayes decision rule Minimax procedures
Predictions

Bayes risk: Let θ be a parameter of interest with the prior


distribution π(θ) and let d(x) be a decision rule then the Bayes risk
for the decision rule w.r.t. π(θ) denoted by R(π, d(x))

3/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Bayes risk
Bayes decision rule Minimax procedures
Predictions

Bayes risk: Let θ be a parameter of interest with the prior


distribution π(θ) and let d(x) be a decision rule then the Bayes risk
for the decision rule w.r.t. π(θ) denoted by R(π, d(x))
It is defined as R(π, d(x)) = Eθ R(θ, d(x))

3/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Bayes risk
Bayes decision rule Minimax procedures
Predictions

Bayes risk: Let θ be a parameter of interest with the prior


distribution π(θ) and let d(x) be a decision rule then the Bayes risk
for the decision rule w.r.t. π(θ) denoted by R(π, d(x))
It is defined as R(π, d(x)) = Eθ R(θ, d(x))
Recall that

R(π, d(x)) = Eθ R(θ, d(x))


Z
= R(θ, d(x))π(θ)dθ; or
θ
X
= R(θ, d(x))π(θ)dθ
θ

3/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Bayes risk
Bayes decision rule Minimax procedures
Predictions

Bayes risk: Let θ be a parameter of interest with the prior


distribution π(θ) and let d(x) be a decision rule then the Bayes risk
for the decision rule w.r.t. π(θ) denoted by R(π, d(x))
It is defined as R(π, d(x)) = Eθ R(θ, d(x))
Recall that

R(π, d(x)) = Eθ R(θ, d(x))


Z
= R(θ, d(x))π(θ)dθ; or
θ
X
= R(θ, d(x))π(θ)dθ
θ

Remarks: R(π, d(x)) is the average risk and a single value unlike
R(θ, d(x)) and

3/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Bayes risk
Bayes decision rule Minimax procedures
Predictions

Bayes risk: Let θ be a parameter of interest with the prior


distribution π(θ) and let d(x) be a decision rule then the Bayes risk
for the decision rule w.r.t. π(θ) denoted by R(π, d(x))
It is defined as R(π, d(x)) = Eθ R(θ, d(x))
Recall that

R(π, d(x)) = Eθ R(θ, d(x))


Z
= R(θ, d(x))π(θ)dθ; or
θ
X
= R(θ, d(x))π(θ)dθ
θ

Remarks: R(π, d(x)) is the average risk and a single value unlike
R(θ, d(x)) and
Bayes decision rule is the decision with minimum Bayes risk
3/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator Minimax procedures
Predictions
Weighted squared loss function

Let d ∗ (x) be the decision rule that minimizes the Bayes risk within
class D of decision rules.

4/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator Minimax procedures
Predictions
Weighted squared loss function

Let d ∗ (x) be the decision rule that minimizes the Bayes risk within
class D of decision rules.
Then d ∗ (x) is the Bayes estimator or the Bayes decision rule

R(π, d ∗ (x)) = min R(π, d(x))


d∈D

4/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator Minimax procedures
Predictions
Weighted squared loss function

Let d ∗ (x) be the decision rule that minimizes the Bayes risk within
class D of decision rules.
Then d ∗ (x) is the Bayes estimator or the Bayes decision rule

R(π, d ∗ (x)) = min R(π, d(x))


d∈D

Posterior risk function: Let π(θ) be the prior of θ and π(θ|x) be its
posterior distribution.

4/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator Minimax procedures
Predictions
Weighted squared loss function

Let d ∗ (x) be the decision rule that minimizes the Bayes risk within
class D of decision rules.
Then d ∗ (x) is the Bayes estimator or the Bayes decision rule

R(π, d ∗ (x)) = min R(π, d(x))


d∈D

Posterior risk function: Let π(θ) be the prior of θ and π(θ|x) be its
posterior distribution.
Then the posterior risk for the decision rule denoted by r (π, d(x)) is
defined as
r (π, d(x)) = Eθ|x [L(θ, d(x))]

4/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator Minimax procedures
Predictions
Weighted squared loss function

Let d ∗ (x) be the decision rule that minimizes the Bayes risk within
class D of decision rules.
Then d ∗ (x) is the Bayes estimator or the Bayes decision rule

R(π, d ∗ (x)) = min R(π, d(x))


d∈D

Posterior risk function: Let π(θ) be the prior of θ and π(θ|x) be its
posterior distribution.
Then the posterior risk for the decision rule denoted by r (π, d(x)) is
defined as
r (π, d(x)) = Eθ|x [L(θ, d(x))]

In estimation theory and decision theory, a Bayes estimator or Bayes


action is an estimator that minimizes the posterior expected value
of a loss function.
4/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator. Minimax procedures
Predictions
Weighted squared loss function

Proof: Let L(θ, a) be a loss function and d(x) be a decision rule.


Then the Bayes decision rule depends on L(θ, a)

5/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator. Minimax procedures
Predictions
Weighted squared loss function

Proof: Let L(θ, a) be a loss function and d(x) be a decision rule.


Then the Bayes decision rule depends on L(θ, a)
Bayes estimator minimizes the Bayes risk R(π, d(x))
Z
R(π, d(x)) = R(θ, d(x))π(θ)dθ
θ
Z Z
= L(θ, d(x))f (x|θ)π(θ)dθdx
θ x
Z Z  
f (x|θ)π(θ)
= L(θ, d(x)) f (x)dθdx
θ x f (x)
Z Z 
= L(θ, d(x)) [f (θ|x)] dθ f (x)dx
Zx θ
∴ R(π, d(x)) = r (π, d(x))f (x)dx
x

5/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Bayes Estimator. Minimax procedures
Predictions
Weighted squared loss function

Proof: Let L(θ, a) be a loss function and d(x) be a decision rule.


Then the Bayes decision rule depends on L(θ, a)
Bayes estimator minimizes the Bayes risk R(π, d(x))
Z
R(π, d(x)) = R(θ, d(x))π(θ)dθ
θ
Z Z
= L(θ, d(x))f (x|θ)π(θ)dθdx
θ x
Z Z  
f (x|θ)π(θ)
= L(θ, d(x)) f (x)dθdx
θ x f (x)
Z Z 
= L(θ, d(x)) [f (θ|x)] dθ f (x)dx
Zx θ
∴ R(π, d(x)) = r (π, d(x))f (x)dx
x

The Bayes estimator d (x) which minimizes R(π, d(x)) must also
minimize the posterior risk function r (π, d(x))
5/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Squared Loss function Minimax procedures
Predictions
Weighted squared loss function

Suppose f (θ|x) is the posterior distribution of θ. Then the Bayes


estimator (decision rule) under the squared loss function is the
mean of the posterior distribution of θ

6/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Squared Loss function Minimax procedures
Predictions
Weighted squared loss function

Suppose f (θ|x) is the posterior distribution of θ. Then the Bayes


estimator (decision rule) under the squared loss function is the
mean of the posterior distribution of θ
Now L(θ, a) = (θ − a)2 noting that the Bayes estimator minimizes
the the posterior risk function. Let d(x) be a decision rule.
Z Z
r (π, d(x)) = L(θ, d(x)) [f (θ|x)] dθ = (θ − d(x))2 f (θ|x)dθ
Zθ Z
= θ dθ − 2d(x) θf (θ|x)dθ + (d(x))2
2
θ θ

6/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Squared Loss function Minimax procedures
Predictions
Weighted squared loss function

Suppose f (θ|x) is the posterior distribution of θ. Then the Bayes


estimator (decision rule) under the squared loss function is the
mean of the posterior distribution of θ
Now L(θ, a) = (θ − a)2 noting that the Bayes estimator minimizes
the the posterior risk function. Let d(x) be a decision rule.
Z Z
r (π, d(x)) = L(θ, d(x)) [f (θ|x)] dθ = (θ − d(x))2 f (θ|x)dθ
Zθ Z
= θ dθ − 2d(x) θf (θ|x)dθ + (d(x))2
2
θ θ

6/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Squared Loss function Minimax procedures
Predictions
Weighted squared loss function

Suppose f (θ|x) is the posterior distribution of θ. Then the Bayes


estimator (decision rule) under the squared loss function is the
mean of the posterior distribution of θ
Now L(θ, a) = (θ − a)2 noting that the Bayes estimator minimizes
the the posterior risk function. Let d(x) be a decision rule.
Z Z
r (π, d(x)) = L(θ, d(x)) [f (θ|x)] dθ = (θ − d(x))2 f (θ|x)dθ
Zθ Z
= θ dθ − 2d(x) θf (θ|x)dθ + (d(x))2
2
θ θ
R
Minimizing w.r.t. d(x), − 2 θf (θ|x)dθ + 2d(x) = 0 thus
Z
d(x) = θf (θ|x)dθ = Eθ|x (θ)

6/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Squared Loss function Minimax procedures
Predictions
Weighted squared loss function

Suppose f (θ|x) is the posterior distribution of θ. Then the Bayes


estimator (decision rule) under the squared loss function is the
mean of the posterior distribution of θ
Now L(θ, a) = (θ − a)2 noting that the Bayes estimator minimizes
the the posterior risk function. Let d(x) be a decision rule.
Z Z
r (π, d(x)) = L(θ, d(x)) [f (θ|x)] dθ = (θ − d(x))2 f (θ|x)dθ
Zθ Z
= θ dθ − 2d(x) θf (θ|x)dθ + (d(x))2
2
θ θ
R
Minimizing w.r.t. d(x), − 2 θf (θ|x)dθ + 2d(x) = 0 thus
Z
d(x) = θf (θ|x)dθ = Eθ|x (θ)

The Bayes estimator d(x) is the mean of θ under posterior


distribution.
6/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x

7/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior hence find the Bayes estimate.

7/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior hence find the Bayes estimate.
Solution:
k
Y
fx (θ|x) ∝ f (xj |θ)π(θ) = θk x̄ (1 − θ)nk−k x̄ θα−1 (1 − θ)β−1
j=1

∝ θk x̄+α−1 (1 − θ)nk−k x̄+β−1

7/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Let π(θ) be a Beta distribution with parameters α and β. Show that
π(θ) is a conjugate prior hence find the Bayes estimate.
Solution:
k
Y
fx (θ|x) ∝ f (xj |θ)π(θ) = θk x̄ (1 − θ)nk−k x̄ θα−1 (1 − θ)β−1
j=1

∝ θk x̄+α−1 (1 − θ)nk−k x̄+β−1


The posterior is an incomplete Beta distribution with parameters
α′ = α + k x̄ and β ′ = nk + β − k x̄
α′ α + nx̄
Eθ|x (θ) = ′ ′
=
α +β β + nk + α
7/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Absolute loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a| The Bayes estimator of θ given the absolute


loss function is the median of the posterior distribution.

8/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Absolute loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a| The Bayes estimator of θ given the absolute


loss function is the median of the posterior distribution.
Proof: The Bayes estimator minimizes the posterior risk r (θ, d(x))
Let ′ a′ be a value such that θ − d(x) > 0 ∀ θ > a
Z Z ∞
r (π, d(x)) = L(θ, d(x))f (θ|x)dθ = |θ − d(x)|f (θ|x)dθ
θ −∞
Z a Z ∞
= (d(x) − θ)f (θ|x)dθ − (θ − d(x))f (θ|x)dθ
−∞ a

8/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Absolute loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a| The Bayes estimator of θ given the absolute


loss function is the median of the posterior distribution.
Proof: The Bayes estimator minimizes the posterior risk r (θ, d(x))
Let ′ a′ be a value such that θ − d(x) > 0 ∀ θ > a
Z Z ∞
r (π, d(x)) = L(θ, d(x))f (θ|x)dθ = |θ − d(x)|f (θ|x)dθ
θ −∞
Z a Z ∞
= (d(x) − θ)f (θ|x)dθ − (θ − d(x))f (θ|x)dθ
−∞ a

Minimizing w.r.t. d(x) and equating to zero we get


Z a Z ∞
f (θ|x)dθ = f (θ|x)dθ
−∞ a

8/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Absolute loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a| The Bayes estimator of θ given the absolute


loss function is the median of the posterior distribution.
Proof: The Bayes estimator minimizes the posterior risk r (θ, d(x))
Let ′ a′ be a value such that θ − d(x) > 0 ∀ θ > a
Z Z ∞
r (π, d(x)) = L(θ, d(x))f (θ|x)dθ = |θ − d(x)|f (θ|x)dθ
θ −∞
Z a Z ∞
= (d(x) − θ)f (θ|x)dθ − (θ − d(x))f (θ|x)dθ
−∞ a

Minimizing w.r.t. d(x) and equating to zero we get


Z a Z ∞
f (θ|x)dθ = f (θ|x)dθ
−∞ a

Hence the only value of θ that would satisfy this inequality is the
median of the posterior distribution. This implies that ”a” is the
median.
8/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a|. Obtain the Bayes estimator for θ under loss
function given that
 −3
2θ , θ > 1;
f (θ|x) =
0, elsewhere.

9/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a|. Obtain the Bayes estimator for θ under loss
function given that
 −3
2θ , θ > 1;
f (θ|x) =
0, elsewhere.

Solution: The Bayes estimator under this loss function is the media
of the posterior distribution. Let Md be the median
Z Md
pr [θ < Md ] = 0.5 ⇒ 2θ−3 dθ = 0.5
1
Md 1 √
θ−2 1
= 0.5; ⇒1− 2 = 0.5; ∴ Md = 2
Md

9/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = |θ − a|. Obtain the Bayes estimator for θ under loss
function given that
 −3
2θ , θ > 1;
f (θ|x) =
0, elsewhere.

Solution: The Bayes estimator under this loss function is the media
of the posterior distribution. Let Md be the median
Z Md
pr [θ < Md ] = 0.5 ⇒ 2θ−3 dθ = 0.5
1
Md 1 √
θ−2 1
= 0.5; ⇒1− 2 = 0.5; ∴ Md = 2
Md

Suppose f (θ|x) = N(5, 1). Obtain the Bayes estimator for θ under
the absolute loss function. Solution: Md = 5 This is because under
the normal pdf, the mean partitions it in to two equal sections
9/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Weighted squared loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = w (θ)(θ − a)2

10/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Weighted squared loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = w (θ)(θ − a)2


The the Bayes estimator for θ under this loss function given the
prior π(θ) is
Eθ|x [θ.w (θ)]
dπ (x) =
Eθ|x [w (θ)]

10/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Weighted squared loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = w (θ)(θ − a)2


The the Bayes estimator for θ under this loss function given the
prior π(θ) is
Eθ|x [θ.w (θ)]
dπ (x) =
Eθ|x [w (θ)]
Proof: We know that
Z Z
r (π, d(x)) = L(θ, d(x))f (θ|x)dθ = w (θ)(θ − d(x))2 f (θ|x)dθ
θ θ
Z
w (θ) θ − 2θd(x) + [d(x)]2 f (θ|x)dθ
 2 
=
θ

10/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Weighted squared loss function Minimax procedures
Predictions
Weighted squared loss function

Let L(θ, a) = w (θ)(θ − a)2


The the Bayes estimator for θ under this loss function given the
prior π(θ) is
Eθ|x [θ.w (θ)]
dπ (x) =
Eθ|x [w (θ)]
Proof: We know that
Z Z
r (π, d(x)) = L(θ, d(x))f (θ|x)dθ = w (θ)(θ − d(x))2 f (θ|x)dθ
θ θ
Z
w (θ) θ − 2θd(x) + [d(x)]2 f (θ|x)dθ
 2 
=
θ

Minimizing w.r.t. d(x) and equating to zero while solving for d(x)
we get R
θ.w (θ)f (θ|x)dθ Eθ|x θw (θ)
d(x) = Rθ =
θ
w (θ)f (θ|x)dθ Eθ|x w (θ)
10/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x

11/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Find the Bayes estimator for θ under the loss function
(θ − a)2
L(θ, a) = , given that π(θ) = Beta(α, β)
θ(1 − θ)

11/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Find the Bayes estimator for θ under the loss function
(θ − a)2
L(θ, a) = , given that π(θ) = Beta(α, β)
θ(1 − θ)
Solution: Let w (θ) = 1/(θ(1 − θ)) and f (θ|x) = Beta(α′ , β ′ ) where
k
X k
X
α′ = α + xj , β ′ = β + nk − xj
j=1 j=1

11/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Squared loss functions
Bayes Estimator
Absolute loss function
Example Minimax procedures
Predictions
Weighted squared loss function

Let x1 , x2 , ..., xk be a random sample from a binomial distribution


with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Find the Bayes estimator for θ under the loss function
(θ − a)2
L(θ, a) = , given that π(θ) = Beta(α, β)
θ(1 − θ)
Solution: Let w (θ) = 1/(θ(1 − θ)) and f (θ|x) = Beta(α′ , β ′ ) where
k
X k
X
α′ = α + xj , β ′ = β + nk − xj
j=1 j=1

It can be shown that the Bayes estimator


Eθ|x θw (θ) α′ − 1 α + k x̄ − 1
dπ (x) = = ′ =
Eθ|x w (θ) α + β′ − 2 α+β−2
11/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax procedures. Minimax procedures
Predictions
Credibility intervals

Let dπ (x) be the Bayes estimator for θ with respect to the prior
π(θ) and R(π, dπ (x)) be the Bayes risk for dπ (x).i.e.

R(π, dπ (x)) = Eθ Ex [L(θ, dπ (x))]

12/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax procedures. Minimax procedures
Predictions
Credibility intervals

Let dπ (x) be the Bayes estimator for θ with respect to the prior
π(θ) and R(π, dπ (x)) be the Bayes risk for dπ (x).i.e.

R(π, dπ (x)) = Eθ Ex [L(θ, dπ (x))]

The prior π(θ) is said to be least favourable if

R(π, dπ (x)) ≥ R(π, dπ′ (x))


R(π, dπ′ (x)) = sup R(π, dπ (x))

for all other prior distribution π ′ (θ) for θ

12/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax procedures. Minimax procedures
Predictions
Credibility intervals

Let dπ (x) be the Bayes estimator for θ with respect to the prior
π(θ) and R(π, dπ (x)) be the Bayes risk for dπ (x).i.e.

R(π, dπ (x)) = Eθ Ex [L(θ, dπ (x))]

The prior π(θ) is said to be least favourable if

R(π, dπ (x)) ≥ R(π, dπ′ (x))


R(π, dπ′ (x)) = sup R(π, dπ (x))

for all other prior distribution π ′ (θ) for θ


Note that
Z Z
R(π, dπ′ (x)) = L(θ, dπ (x))f (x|θ)π ′ (θ)dθdx
θ x

12/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes. Minimax procedures
Predictions
Credibility intervals

Theorem: Suppose π(θ) is the prior for θ and dπ (x) is the Bayes
estimator with respect to π(θ) such that the following holds
Z
R(π, dπ′ (x)) = R(θ, dπ (x))π(θ)dθ = sup R(θ, dπ (x))
θ θ

13/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes. Minimax procedures
Predictions
Credibility intervals

Theorem: Suppose π(θ) is the prior for θ and dπ (x) is the Bayes
estimator with respect to π(θ) such that the following holds
Z
R(π, dπ′ (x)) = R(θ, dπ (x))π(θ)dθ = sup R(θ, dπ (x))
θ θ

Then (i) dπ (x) is minimax Bayes; (ii) If dπ (x) is unique, then its the
unique minimax Bayes (iii) π(θ) is least favourable

13/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes. Minimax procedures
Predictions
Credibility intervals

Theorem: Suppose π(θ) is the prior for θ and dπ (x) is the Bayes
estimator with respect to π(θ) such that the following holds
Z
R(π, dπ′ (x)) = R(θ, dπ (x))π(θ)dθ = sup R(θ, dπ (x))
θ θ

Then (i) dπ (x) is minimax Bayes; (ii) If dπ (x) is unique, then its the
unique minimax Bayes (iii) π(θ) is least favourable
Proof: Let d(x) be another decision rule.
Z
sup R(θ, d(x)) ≥ R(θ, d(x))π(θ)dθ
θ

≥ R(θ, dπ (x))π(θ)dθ
θ
sup R(θ, dπ (x)) ≤ sup R(θ, d(x)), ∴ dπ (x) is minimax
θ θ

13/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes. Minimax procedures
Predictions
Credibility intervals

Theorem: Suppose π(θ) is the prior for θ and dπ (x) is the Bayes
estimator with respect to π(θ) such that the following holds
Z
R(π, dπ′ (x)) = R(θ, dπ (x))π(θ)dθ = sup R(θ, dπ (x))
θ θ

Then (i) dπ (x) is minimax Bayes; (ii) If dπ (x) is unique, then its the
unique minimax Bayes (iii) π(θ) is least favourable
Proof: Let d(x) be another decision rule.
Z
sup R(θ, d(x)) ≥ R(θ, d(x))π(θ)dθ
θ

≥ R(θ, dπ (x))π(θ)dθ
θ
sup R(θ, dπ (x)) ≤ sup R(θ, d(x)), ∴ dπ (x) is minimax
θ θ

Unique minimax we replace ≥ by > and follow the same procedure.


13/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes.. Minimax procedures
Predictions
Credibility intervals

Least favourable: Let π ′ (θ) be another prior, then


Z
R(π, dπ (x)) =
′ R(θ, dπ (x))π ′ (θ)dθ
θ
Z
≤ sup R(θ, dπ′ (x)) ≤ R(θ, dπ (x))π ′ (θ)dθ
θ θ
Z
= R(θ, dπ (x))π(θ)dθ
θ
R(θ, dπ′ (x)) ≤ R(θ, dπ (x)). ∴ Π(θ) is least favourable.

14/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes.. Minimax procedures
Predictions
Credibility intervals

Least favourable: Let π ′ (θ) be another prior, then


Z
R(π, dπ (x)) =
′ R(θ, dπ (x))π ′ (θ)dθ
θ
Z
≤ sup R(θ, dπ′ (x)) ≤ R(θ, dπ (x))π ′ (θ)dθ
θ θ
Z
= R(θ, dπ (x))π(θ)dθ
θ
R(θ, dπ′ (x)) ≤ R(θ, dπ (x)). ∴ Π(θ) is least favourable.
Corollary: If a Bayes estimator has a constant risk function, then it
is minimax.

14/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Minimax Bayes.. Minimax procedures
Predictions
Credibility intervals

Least favourable: Let π ′ (θ) be another prior, then


Z
R(π, dπ (x)) =
′ R(θ, dπ (x))π ′ (θ)dθ
θ
Z
≤ sup R(θ, dπ′ (x)) ≤ R(θ, dπ (x))π ′ (θ)dθ
θ θ
Z
= R(θ, dπ (x))π(θ)dθ
θ
R(θ, dπ′ (x)) ≤ R(θ, dπ (x)). ∴ Π(θ) is least favourable.
Corollary: If a Bayes estimator has a constant risk function, then it
is minimax.
Example. Let x1 , x2 , ..., xk be a random sample from a binomial
distribution with parameters n and θ i.e.
 
n x
fX (x|θ) = θ (1 − θ)n−x , x = 0, 1, 2, ..., n; 0 ≤ θ ≤ 1
x
Suppose π(θ) = Beta(α, β) and L(θ, a) = (θ − a)2 . Get the
14/33
minimax Bayes estimator and least favourable prior.
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example. Minimax procedures
Predictions
Credibility intervals

Solution:Under the squared loss function, the Bayes estimator


Pk
α + j=1 xj α + k x̄
dπ (x) = ≡
α + β + nk α + β + nk

15/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example. Minimax procedures
Predictions
Credibility intervals

Solution:Under the squared loss function, the Bayes estimator


Pk
α + j=1 xj α + k x̄
dπ (x) = ≡
α + β + nk α + β + nk
We determine the risk function for dπ (x)
R(θ, dπ (x)) = Ex [L(θ, dπ (x)) = Ex [θ − dπ (x)]2
= var (dπ (x)) + [biasθ dπ (x)]2 ,
α + knθ βθ − α(1 − θ)
biasθ dπ (x) = [θ − E[dπ (x)]] = θ − ≡
α + β + nk α + β + nk
2 2
k k nθ(1 − θ)
var (dπ (x)) = var (x̄) ≡
(α + β + nk)2 k(α + β + nk)2

15/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example. Minimax procedures
Predictions
Credibility intervals

Solution:Under the squared loss function, the Bayes estimator


Pk
α + j=1 xj α + k x̄
dπ (x) = ≡
α + β + nk α + β + nk
We determine the risk function for dπ (x)
R(θ, dπ (x)) = Ex [L(θ, dπ (x)) = Ex [θ − dπ (x)]2
= var (dπ (x)) + [biasθ dπ (x)]2 ,
α + knθ βθ − α(1 − θ)
biasθ dπ (x) = [θ − E[dπ (x)]] = θ − ≡
α + β + nk α + β + nk
2 2
k k nθ(1 − θ)
var (dπ (x)) = var (x̄) ≡
(α + β + nk)2 k(α + β + nk)2
This implies that the risk function R(θ, dπ (x)) is given by
 2
knθ(1 − θ) βθ − α(1 − θ)
R(θ, dπ (x)) = +
(α + β + nk)2 α + β + nk
15/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

From the risk function, since θ exists, then the Bayes estimator is
not minimax. We then determine values of α and β that makes
constant risk function.

16/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

From the risk function, since θ exists, then the Bayes estimator is
not minimax. We then determine values of α and β that makes
constant risk function.
Thereafter we can find minimax Bayes and least favourable prior.
 2
knθ(1 − θ) βθ − α(1 − θ)
R(θ, dπ (x)) = +
(α + β + nk)2 α + β + nk
[α + θ(kn − 2βα − 2α2 )] + θ2 [β 2 − kn + 2αβ + α2 ]
2
=
(α + β + nk)2

16/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

From the risk function, since θ exists, then the Bayes estimator is
not minimax. We then determine values of α and β that makes
constant risk function.
Thereafter we can find minimax Bayes and least favourable prior.
 2
knθ(1 − θ) βθ − α(1 − θ)
R(θ, dπ (x)) = +
(α + β + nk)2 α + β + nk
[α + θ(kn − 2βα − 2α2 )] + θ2 [β 2 − kn + 2αβ + α2 ]
2
=
(α + β + nk)2
Equating then coefficients of θ and θ2 to zero, we solve for
√ α and β.
kn − 2α(β + α) = 0 and (α + β)2 − kn = 0, ∴ α = β = kn/2

16/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

From the risk function, since θ exists, then the Bayes estimator is
not minimax. We then determine values of α and β that makes
constant risk function.
Thereafter we can find minimax Bayes and least favourable prior.
 2
knθ(1 − θ) βθ − α(1 − θ)
R(θ, dπ (x)) = +
(α + β + nk)2 α + β + nk
[α + θ(kn − 2βα − 2α2 )] + θ2 [β 2 − kn + 2αβ + α2 ]
2
=
(α + β + nk)2
Equating then coefficients of θ and θ2 to zero, we solve for
√ α and β.
kn − 2α(β + α) = 0 and (α + β)2 − kn = 0, ∴ α = β = kn/2
The corresponding Bayes estimator has a constant risk function is
our minimax Bayes estimator and the resulting prior is the least
favourable i.e.

kn
√ √ !
+ k x̄ kn kn
dπ (x) = √2 , π(θ) = Beta ,
kn + nk 2 2
16/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility Minimax procedures
Predictions
Credibility intervals

Definition: A decision rule d(x) is said to be admissible if there does


not exist another decision rule d ∗ (x) that dominates it. i.e. ∄ a
d ∗ (x) such that R(θ, d ∗ (x)) ≤ R(θ, d(x)), ∀ θ

17/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility Minimax procedures
Predictions
Credibility intervals

Definition: A decision rule d(x) is said to be admissible if there does


not exist another decision rule d ∗ (x) that dominates it. i.e. ∄ a
d ∗ (x) such that R(θ, d ∗ (x)) ≤ R(θ, d(x)), ∀ θ
Theorem: Absolute Bayes decision rule. Any unique decision rule is
admissible

17/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility Minimax procedures
Predictions
Credibility intervals

Definition: A decision rule d(x) is said to be admissible if there does


not exist another decision rule d ∗ (x) that dominates it. i.e. ∄ a
d ∗ (x) such that R(θ, d ∗ (x)) ≤ R(θ, d(x)), ∀ θ
Theorem: Absolute Bayes decision rule. Any unique decision rule is
admissible
Proof: Let dπ (x) be the Bayes decision rule w.r.t. the prior π(θ)
and assume that dπ (x) is inadmissible.
∄, a d ∗ (x) s.t. R(θ, d ∗ (x)) ≤ R(θ, d(x))∀θ

17/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility Minimax procedures
Predictions
Credibility intervals

Definition: A decision rule d(x) is said to be admissible if there does


not exist another decision rule d ∗ (x) that dominates it. i.e. ∄ a
d ∗ (x) such that R(θ, d ∗ (x)) ≤ R(θ, d(x)), ∀ θ
Theorem: Absolute Bayes decision rule. Any unique decision rule is
admissible
Proof: Let dπ (x) be the Bayes decision rule w.r.t. the prior π(θ)
and assume that dπ (x) is inadmissible.
∄, a d ∗ (x) s.t. R(θ, d ∗ (x)) ≤ R(θ, d(x))∀θ
We take the average of the risk functions to get Bayes risks of two
decision rules
Z Z

R(θ, d (x))π(θ)dθ ≤ R(θ, dπ (x))π(θ)dθ
θ θ
R(π, d ∗ (x)) ≤ R(π, dπ (x))

17/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility Minimax procedures
Predictions
Credibility intervals

Definition: A decision rule d(x) is said to be admissible if there does


not exist another decision rule d ∗ (x) that dominates it. i.e. ∄ a
d ∗ (x) such that R(θ, d ∗ (x)) ≤ R(θ, d(x)), ∀ θ
Theorem: Absolute Bayes decision rule. Any unique decision rule is
admissible
Proof: Let dπ (x) be the Bayes decision rule w.r.t. the prior π(θ)
and assume that dπ (x) is inadmissible.
∄, a d ∗ (x) s.t. R(θ, d ∗ (x)) ≤ R(θ, d(x))∀θ
We take the average of the risk functions to get Bayes risks of two
decision rules
Z Z

R(θ, d (x))π(θ)dθ ≤ R(θ, dπ (x))π(θ)dθ
θ θ
R(π, d ∗ (x)) ≤ R(π, dπ (x))
But dπ (x) is the Bayes decision rule, then it has minimum Bayes
risk. Thus R(π, d ∗ (x)) ≤ R(θ, dπ (x)) contradicts this fact. Thus ∄
a d ∗ (x) that dominates dπ (x) thus dπ (x) is admissible.
17/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility. Minimax procedures
Predictions
Credibility intervals

Theorem: (Assessing admissibility of decision rules) Suppose x is a


random variable with mean θ and variance σ 2 > 0 then
d(x) = ax + b where a and b are constants is inadmissible under the
squared loss function if (i) a > 1, (ii) a = 1, b ̸= 0 (iii) a < 0

18/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility. Minimax procedures
Predictions
Credibility intervals

Theorem: (Assessing admissibility of decision rules) Suppose x is a


random variable with mean θ and variance σ 2 > 0 then
d(x) = ax + b where a and b are constants is inadmissible under the
squared loss function if (i) a > 1, (ii) a = 1, b ̸= 0 (iii) a < 0
Proof: Now d(x) = ax + b

R(θ, d(x)) = Ex [L(θ, d(x))] = Ex [(θ − (ax + b))2 ]


= θ2 − 2θaE(x) − 2θb + a2 E(x 2 ) + 2abE(x) − 2θb
= θ2 (1 + a2 − 2a) − 2bθ(1 − a) + a2 σ 2 + b 2
= θ2 + a2 σ 2 + a2 θ2 + b 2 + 2abθ − 2aθ2 − 2bθ
= ((a − 1)θ + b)2 + σ 2 a2 = f (a, b)

18/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility. Minimax procedures
Predictions
Credibility intervals

Theorem: (Assessing admissibility of decision rules) Suppose x is a


random variable with mean θ and variance σ 2 > 0 then
d(x) = ax + b where a and b are constants is inadmissible under the
squared loss function if (i) a > 1, (ii) a = 1, b ̸= 0 (iii) a < 0
Proof: Now d(x) = ax + b

R(θ, d(x)) = Ex [L(θ, d(x))] = Ex [(θ − (ax + b))2 ]


= θ2 − 2θaE(x) − 2θb + a2 E(x 2 ) + 2abE(x) − 2θb
= θ2 (1 + a2 − 2a) − 2bθ(1 − a) + a2 σ 2 + b 2
= θ2 + a2 σ 2 + a2 θ2 + b 2 + 2abθ − 2aθ2 − 2bθ
= ((a − 1)θ + b)2 + σ 2 a2 = f (a, b)

(i) a > 1; then f (a, b) > σ 2 a2 = f (a, 0) ⇒ d ′ (x) = ax dominates


d(x) = ax + b hence d(x) is inadmissible

18/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility.. Minimax procedures
Predictions
Credibility intervals

(ii) a < 0 ⇒ (a − 1)2 > 1 therefore


 
2 2 b
f (a, b) ≥ ((a − 1)θ + b) = (a − 1) θ +
(a − 1)
 2  
b b
> θ+ =f 0, −
(a − 1) a−1
this implies that d ′ (x) = −b/(a − 1) dominated d(x) = ax + b,
thus d(x) is inadmissible.

19/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility.. Minimax procedures
Predictions
Credibility intervals

(ii) a < 0 ⇒ (a − 1)2 > 1 therefore


 
2 2 b
f (a, b) ≥ ((a − 1)θ + b) = (a − 1) θ +
(a − 1)
 2  
b b
> θ+ =f 0, −
(a − 1) a−1
this implies that d ′ (x) = −b/(a − 1) dominated d(x) = ax + b,
thus d(x) is inadmissible.
Let a = 1, b ̸= 0
f (1, b) = σ 2 + b 2 > σ 2 = f (1, 0)
f (1, 0) is the risk function for d ′ (x) = x which dominates
d(x) = ax + b which is inadmissible

19/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility.. Minimax procedures
Predictions
Credibility intervals

(ii) a < 0 ⇒ (a − 1)2 > 1 therefore


 
2 2 b
f (a, b) ≥ ((a − 1)θ + b) = (a − 1) θ +
(a − 1)
 2  
b b
> θ+ =f 0, −
(a − 1) a−1
this implies that d ′ (x) = −b/(a − 1) dominated d(x) = ax + b,
thus d(x) is inadmissible.
Let a = 1, b ̸= 0
f (1, b) = σ 2 + b 2 > σ 2 = f (1, 0)
f (1, 0) is the risk function for d ′ (x) = x which dominates
d(x) = ax + b which is inadmissible
Example: Suppose x ∼ Bernoulli(θ). Is d(x) = (x/2) + 3 an
admissible decision rule for θ under the squared loss function?

19/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Admissibility.. Minimax procedures
Predictions
Credibility intervals

(ii) a < 0 ⇒ (a − 1)2 > 1 therefore


 
2 2 b
f (a, b) ≥ ((a − 1)θ + b) = (a − 1) θ +
(a − 1)
 2  
b b
> θ+ =f 0, −
(a − 1) a−1
this implies that d ′ (x) = −b/(a − 1) dominated d(x) = ax + b,
thus d(x) is inadmissible.
Let a = 1, b ̸= 0
f (1, b) = σ 2 + b 2 > σ 2 = f (1, 0)
f (1, 0) is the risk function for d ′ (x) = x which dominates
d(x) = ax + b which is inadmissible
Example: Suppose x ∼ Bernoulli(θ). Is d(x) = (x/2) + 3 an
admissible decision rule for θ under the squared loss function?
Solution: We determine E(x) = θ, now a = 1/2, b = 3, this
implies that 0 < a < 1 hence d(x) is admissible 19/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals Minimax procedures
Predictions
Credibility intervals

The idea behind credibility intervals is to give an analogue to


confidence interval estimation in classical theory.

20/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals Minimax procedures
Predictions
Credibility intervals

The idea behind credibility intervals is to give an analogue to


confidence interval estimation in classical theory.
From a frequentist point of view, the parameter or the unknown
quantity of interest is non-random leading problems of
interpretation of confidence interval.

20/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals Minimax procedures
Predictions
Credibility intervals

The idea behind credibility intervals is to give an analogue to


confidence interval estimation in classical theory.
From a frequentist point of view, the parameter or the unknown
quantity of interest is non-random leading problems of
interpretation of confidence interval.
In Bayesian theory that parameter is random hence no such
problems arise.

20/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals Minimax procedures
Predictions
Credibility intervals

The idea behind credibility intervals is to give an analogue to


confidence interval estimation in classical theory.
From a frequentist point of view, the parameter or the unknown
quantity of interest is non-random leading problems of
interpretation of confidence interval.
In Bayesian theory that parameter is random hence no such
problems arise.
Definition: A region Cα (x) = C is a (1 − α)100% credibility interval
or region for θ if
Z
f (θ|x)dθ = 1 − θ; i.e. Pr [θ ∈ Cα (x)|x] = 1 − α
C

20/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals. Minimax procedures
Predictions
Credibility intervals

One difficulty with such credibility intervals is that they are not
uniquely defined and any region that satisfies the integral above will
do.

21/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals. Minimax procedures
Predictions
Credibility intervals

One difficulty with such credibility intervals is that they are not
uniquely defined and any region that satisfies the integral above will
do.
To overcome this problem an additional constraint is imposed. This
constraint requires that the interval be as small as possible.

21/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals. Minimax procedures
Predictions
Credibility intervals

One difficulty with such credibility intervals is that they are not
uniquely defined and any region that satisfies the integral above will
do.
To overcome this problem an additional constraint is imposed. This
constraint requires that the interval be as small as possible.
Such a region is known as the highest posterior density (HPD)
region.

21/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Credibility intervals. Minimax procedures
Predictions
Credibility intervals

One difficulty with such credibility intervals is that they are not
uniquely defined and any region that satisfies the integral above will
do.
To overcome this problem an additional constraint is imposed. This
constraint requires that the interval be as small as possible.
Such a region is known as the highest posterior density (HPD)
region.
In general such intervals have to be found numerically. However for
most univariate posterior distribution the values are tabulated for a
range of α values.

21/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ

22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ
Solution: f (θ|x) is skewed to the right;α = 0.05

22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ
Solution: f (θ|x) is skewed to the right;α = 0.05
The lower limit for the HPD is [Link] u be the upper limit.
Z u
2θ−3 dθ = 1 − α = 0.95; ∴ u 2 = 20
1

22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ
Solution: f (θ|x) is skewed to the right;α = 0.05
The lower limit for the HPD is [Link] u be the upper limit.
Z u
2θ−3 dθ = 1 − α = 0.95; ∴ u 2 = 20
1

Suppose f (θ|x) = N(12, 0.1) Obtain the 95% HPD for θ.

22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ
Solution: f (θ|x) is skewed to the right;α = 0.05
The lower limit for the HPD is [Link] u be the upper limit.
Z u
2θ−3 dθ = 1 − α = 0.95; ∴ u 2 = 20
1

Suppose f (θ|x) = N(12, 0.1) Obtain the 95% HPD for θ.


We need two values a and b such that
pr [a ≤ θ ≤ b] = 0.95

22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
minimax Bayes
Bayes Estimator
Admissibility
Example.. Minimax procedures
Predictions
Credibility intervals

Suppose that θ has a posterior distribution given as


 2
f (θ|x) = θ 3 , θ > 1;
0, otherwise.
Determine the 95% HPD for θ
Solution: f (θ|x) is skewed to the right;α = 0.05
The lower limit for the HPD is [Link] u be the upper limit.
Z u
2θ−3 dθ = 1 − α = 0.95; ∴ u 2 = 20
1

Suppose f (θ|x) = N(12, 0.1) Obtain the 95% HPD for θ.


We need two values a and b such that
pr [a ≤ θ ≤ b] = 0.95
The HPD will be symmetric
 
a − 12 b − 12
Pr √ ≤z ≤ √ = 0.95; (a, b) = (11.38, 12.62)
0.1 0.1 22/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

So far we have focused on parameter estimation where we have


specified a probability distribution of the data generating
mechanism.

23/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

So far we have focused on parameter estimation where we have


specified a probability distribution of the data generating
mechanism.
We have shown how sample information is combined with prior
information to form posterior distribution for analysis

23/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

So far we have focused on parameter estimation where we have


specified a probability distribution of the data generating
mechanism.
We have shown how sample information is combined with prior
information to form posterior distribution for analysis
Most often the purpose of formulating a statistical model is to make
predictions about future processes. This is approached as follows in
the Bayesian framework.

23/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

So far we have focused on parameter estimation where we have


specified a probability distribution of the data generating
mechanism.
We have shown how sample information is combined with prior
information to form posterior distribution for analysis
Most often the purpose of formulating a statistical model is to make
predictions about future processes. This is approached as follows in
the Bayesian framework.
Suppose x1 , x2 , ..., xn is from f (x|θ) and that θ has a prior
distribution π(θ). Then we have seen that the posterior distribution
for θ is given as
f (x|θ)π(θ)
f (θ|x) =
f (x)

23/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

So far we have focused on parameter estimation where we have


specified a probability distribution of the data generating
mechanism.
We have shown how sample information is combined with prior
information to form posterior distribution for analysis
Most often the purpose of formulating a statistical model is to make
predictions about future processes. This is approached as follows in
the Bayesian framework.
Suppose x1 , x2 , ..., xn is from f (x|θ) and that θ has a prior
distribution π(θ). Then we have seen that the posterior distribution
for θ is given as
f (x|θ)π(θ)
f (θ|x) =
f (x)
Suppose we wish to make inferences about future values y from the
data generating mechanism. That is we wish to predict y given
current x.
23/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

Such a prediction will be based on the predictive distribution for y


given x denoted by f (y |x)

24/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

Such a prediction will be based on the predictive distribution for y


given x denoted by f (y |x)
It can be shown that the posterior predictive distribution is given by
Z
f (y |x) = f (y |θ)f (θ|x)dθ

= Eθ|x [f (y |θ)]

24/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

Such a prediction will be based on the predictive distribution for y


given x denoted by f (y |x)
It can be shown that the posterior predictive distribution is given by
Z
f (y |x) = f (y |θ)f (θ|x)dθ

= Eθ|x [f (y |θ)]
Proof: By definition we have
f (y , x) = f (y |x)f (x)
f (y , x)
f (y |x) =
f (x)
⇒ f (x, y , θ) = f (x, y |θ)π(θ)

24/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Predictions Minimax procedures
Predictions

Such a prediction will be based on the predictive distribution for y


given x denoted by f (y |x)
It can be shown that the posterior predictive distribution is given by
Z
f (y |x) = f (y |θ)f (θ|x)dθ

= Eθ|x [f (y |θ)]
Proof: By definition we have
f (y , x) = f (y |x)f (x)
f (y , x)
f (y |x) =
f (x)
⇒ f (x, y , θ) = f (x, y |θ)π(θ)
Assumption: Given that θ, y and x are independent.
f (x, y |θ) = f (x|θ)f (θ|y )
∴ f (x, y , θ) = f (x|θ)f (y |θ)π(θ)
24/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Posterior Predictive distributions Minimax procedures
Predictions

Now,
f (x, y , θ) f (x|θ)f (y |θ)π(θ)
f (y , θ|x) = ≡
f (x) f (x)
f (x|θ)π(θ)
= f (y |θ)
f (x)
= f (θ|x)f (y |θ)
Z
∴ f (y |x) = f (y , θ|x)dθ
ZΩ
= f (θ|x)f (y |θ)dθ

= Eθ|x [f (y |θ)]

25/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Posterior Predictive distributions Minimax procedures
Predictions

Now,
f (x, y , θ) f (x|θ)f (y |θ)π(θ)
f (y , θ|x) = ≡
f (x) f (x)
f (x|θ)π(θ)
= f (y |θ)
f (x)
= f (θ|x)f (y |θ)
Z
∴ f (y |x) = f (y , θ|x)dθ
ZΩ
= f (θ|x)f (y |θ)dθ

= Eθ|x [f (y |θ)]

Example: Suppose x1 , x2 , ..., xk are from x ∼ Bin(n, θ) and


π(θ) = Beta(α, β). Let y1 , y2 , ..., ym be future observations from the
same data generating mechanism. Obtain the predictive distribution
of y given x
25/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Solution: We wish to determine f (y |x) = Eθ|x [f (y |θ)]

26/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Solution: We wish to determine f (y |x) = Eθ|x [f (y |θ)]


now, the assumed data generating distribution is from binomial
mass function
f (y |θ) = Bin(n, θ)
 
n y
= θ (1 − θ)n−y , y = 0, 1, 2, ..., n;
y

26/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Solution: We wish to determine f (y |x) = Eθ|x [f (y |θ)]


now, the assumed data generating distribution is from binomial
mass function
f (y |θ) = Bin(n, θ)
 
n y
= θ (1 − θ)n−y , y = 0, 1, 2, ..., n;
y
hence the corresponding likelihood function is given by
m
Y
f (y |θ) = f (yj |θ)
j=1
m 

n yj
Y
= θ (1 − θ)n−yj
yj
j=1
m  
Y n mȳ
= θ (1 − θ)mn−mȳ
yj
j=1
26/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
solution.. Minimax procedures
Predictions

similarly, the density of the prior distribution is beta


f (θ|x) = Beta(α′ , β ′ ); α′ = α + k x̄, β ′ = β + nk − k x̄
Γ(α′ + β ′ ) α′ −1 ′
= θ (1 − θ)β −1
Γ(α′ )Γ(β ′ )

27/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
solution.. Minimax procedures
Predictions

similarly, the density of the prior distribution is beta


f (θ|x) = Beta(α′ , β ′ ); α′ = α + k x̄, β ′ = β + nk − k x̄
Γ(α′ + β ′ ) α′ −1 ′
= θ (1 − θ)β −1
Γ(α′ )Γ(β ′ )

the resulting posterior distribution, can be derived as follows


Z
f (y |x) = f (y |θ)f (θ|x)dθ
θ
m  
Γ(α′ + β ′ ) α′ −1
Z Y
n ′
= θmȳ (1 − θ)mn−mȳ
′ )Γ(β ′ )
θ (1 − θ)β −1 dθ
θ j=1 y j Γ(α
m
Y n  Z
Γ(α + β + nk) ′′ ′′
= θα −1 (1 − θ)β −1 dθ
yj Γ(α + k x̄)Γ(β + nk − k x̄) θ
j=1

α = α′ + mȳ ; β ′′ = β ′ + mn − mȳ
′′

27/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example.. Minimax procedures
Predictions

similarly, it follows that


m   Z 1
Y n Γ(α + β + nk) ′′
−1 ′′
−1
f (y |x) = θα (1 − θ)β dθ
yj Γ(α + k x̄)Γ(β + nk − k x̄) 0
j=1

α = α′ + mȳ ; β ′′ = β ′ + mn − mȳ
′′

m  
Y n Γ(α + β + nk) Γ(α′′ )Γ(β ′′ )
=
yj Γ(α + k x̄)Γ(β + nk − k x̄) Γ(α′′ + β ′′ )
j=1
m 
Γ(α′′ )Γ(β ′′ )

Y n Γ(α + β + nk)
=
yj Γ(α + k x̄)Γ(β + nk − k x̄) Γ(α + β + nk − mn)
j=1
m
n Γ(α′ + β ′ ) Γ(α′ + mȳ )Γ(β ′ + mn − mȳ )
Y  
=
yj Γ(α′ )Γ(β ′ ) Γ(α′ + β ′ − mn)
j=1

28/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example.. Minimax procedures
Predictions

similarly, it follows that


m   Z 1
Y n Γ(α + β + nk) ′′
−1 ′′
−1
f (y |x) = θα (1 − θ)β dθ
yj Γ(α + k x̄)Γ(β + nk − k x̄) 0
j=1

α = α′ + mȳ ; β ′′ = β ′ + mn − mȳ
′′

m  
Y n Γ(α + β + nk) Γ(α′′ )Γ(β ′′ )
=
yj Γ(α + k x̄)Γ(β + nk − k x̄) Γ(α′′ + β ′′ )
j=1
m 
Γ(α′′ )Γ(β ′′ )

Y n Γ(α + β + nk)
=
yj Γ(α + k x̄)Γ(β + nk − k x̄) Γ(α + β + nk − mn)
j=1
m
n Γ(α′ + β ′ ) Γ(α′ + mȳ )Γ(β ′ + mn − mȳ )
Y  
=
yj Γ(α′ )Γ(β ′ ) Γ(α′ + β ′ − mn)
j=1

This implies that f (y |x) is the predictive distribution of y given x.


28/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 2 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. Po(λ) where λ ∼ Gamma(α, β) find


posterior predictive distributions for y .

29/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 2 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. Po(λ) where λ ∼ Gamma(α, β) find


posterior predictive distributions for y .
Solution: It can be shown quite easily that theP corresponding
posterior distribution f (λ|x) is also Gamma( xi + α, n + β)
 
Xn
f (λ|x) = Gamma  xi + α, n + β 
j=1

29/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 2 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. Po(λ) where λ ∼ Gamma(α, β) find


posterior predictive distributions for y .
Solution: It can be shown quite easily that theP corresponding
posterior distribution f (λ|x) is also Gamma( xi + α, n + β)
 
Xn
f (λ|x) = Gamma  xi + α, n + β 
j=1

Posterior predictive distribution is thus determined as follows


λy e −λ
f (y |x) = Eλ|x [f (y |λ)], f (y |λ) =
y!
Z ∞
(β + n)nx̄+α nx̄+α−1
= f (y |λ) λ exp(−[n + β]λ)dλ
0 Γα + nx̄
Z ∞ y −λ
λ e (β + n)nx̄+α nx̄+α−1
= λ exp(−[n + β]λ)dλ
0 y! Γ(α + nx̄)
29/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Posterior predictive distribution is


Z ∞ y −λ
λ e (β + n)nx̄+α nx̄+α−1
f (y |x) = λ exp(−[n + β]λ)dλ
0 y! Γ(α + nx̄)
nx̄+α Z ∞
(β + n)
= λnx̄+y +α−1 e −[n+β+1] dλ
y !Γ(α + nx̄) 0
nx̄+α
(β + n) Γ (nx̄ + y + α)
=
y !Γ(α + nx̄) (n + β + 1)[nx̄+y +α]
 nx̄+α  y
Γ (nx̄ + y + α) n+β 1
=
Γ (nx̄ + α) Γ(y + 1) n + β + 1 n+β+1

30/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Posterior predictive distribution is


Z ∞ y −λ
λ e (β + n)nx̄+α nx̄+α−1
f (y |x) = λ exp(−[n + β]λ)dλ
0 y! Γ(α + nx̄)
nx̄+α Z ∞
(β + n)
= λnx̄+y +α−1 e −[n+β+1] dλ
y !Γ(α + nx̄) 0
nx̄+α
(β + n) Γ (nx̄ + y + α)
=
y !Γ(α + nx̄) (n + β + 1)[nx̄+y +α]
 nx̄+α  y
Γ (nx̄ + y + α) n+β 1
=
Γ (nx̄ + α) Γ(y + 1) n + β + 1 n+β+1
Which is a negative binomial with mean and variance given as
nx̄ + α nx̄ + α
E[y |x] = var (y |x) = (n + β + 1)
n+β (α + β)2

30/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Solution. Minimax procedures
Predictions

Posterior predictive distribution is


Z ∞ y −λ
λ e (β + n)nx̄+α nx̄+α−1
f (y |x) = λ exp(−[n + β]λ)dλ
0 y! Γ(α + nx̄)
nx̄+α Z ∞
(β + n)
= λnx̄+y +α−1 e −[n+β+1] dλ
y !Γ(α + nx̄) 0
nx̄+α
(β + n) Γ (nx̄ + y + α)
=
y !Γ(α + nx̄) (n + β + 1)[nx̄+y +α]
 nx̄+α  y
Γ (nx̄ + y + α) n+β 1
=
Γ (nx̄ + α) Γ(y + 1) n + β + 1 n+β+1
Which is a negative binomial with mean and variance given as
nx̄ + α nx̄ + α
E[y |x] = var (y |x) = (n + β + 1)
n+β (α + β)2
Note that the posterior predictive distribution has the same mean as
posterior distribution but greater variance since we are drawing a
new data value. 30/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 3 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. N(µ, σ 2 ) where σ 2 is known and


µ ∼ N(δ, τ 2 ). Determine posterior predictive distributions for y |x
(show this !!)

31/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 3 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. N(µ, σ 2 ) where σ 2 is known and


µ ∼ N(δ, τ 2 ). Determine posterior predictive distributions for y |x
(show this !!)
The posterior distribution for µ|x is normally distributed with mean
and variance
δ/τ 2 + nx̄/σ 2 2 τ 2 σ2
µpost = , σ post =
1/τ 2 + n/σ 2 σ 2 + nτ 2

31/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 3 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. N(µ, σ 2 ) where σ 2 is known and


µ ∼ N(δ, τ 2 ). Determine posterior predictive distributions for y |x
(show this !!)
The posterior distribution for µ|x is normally distributed with mean
and variance
δ/τ 2 + nx̄/σ 2 2 τ 2 σ2
µpost = , σ post =
1/τ 2 + n/σ 2 σ 2 + nτ 2

Where y |µ ∼ N(µ, σ 2 ), the posterior predictive distribution is


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

31/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Example 3 Minimax procedures
Predictions

Let x1 , x2 , ..., xn ∼ i.i.d. N(µ, σ 2 ) where σ 2 is known and


µ ∼ N(δ, τ 2 ). Determine posterior predictive distributions for y |x
(show this !!)
The posterior distribution for µ|x is normally distributed with mean
and variance
δ/τ 2 + nx̄/σ 2 2 τ 2 σ2
µpost = , σ post =
1/τ 2 + n/σ 2 σ 2 + nτ 2

Where y |µ ∼ N(µ, σ 2 ), the posterior predictive distribution is


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

Sometimes, the form of f (y |x) can be derived directly but it is often


easier to sample from f (y |x) using Monte-Carlo methods
31/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Sampling Minimax procedures
Predictions

For example one can sample form the expression


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

using the following two or three steps

32/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Sampling Minimax procedures
Predictions

For example one can sample form the expression


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

using the following two or three steps


For j = 1, 2, ..., J sample
step1: µ[j] from f (µ|x) i.e. from the posterior distribution.
step2: X ∗[j] from f (y |µ[j] ) assumed data generating mechanism

32/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Sampling Minimax procedures
Predictions

For example one can sample form the expression


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

using the following two or three steps


For j = 1, 2, ..., J sample
step1: µ[j] from f (µ|x) i.e. from the posterior distribution.
step2: X ∗[j] from f (y |µ[j] ) assumed data generating mechanism
Then X ∗[1] , X ∗[2] , ..., X ∗[n] are i.i.d sample for f (y |x)

32/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023
Bayes decision rule
Bayes Estimator
Posterior Predictive distribution
Sampling Minimax procedures
Predictions

For example one can sample form the expression


Z ∞
f (y |x) = f (y |µ)f (µ|x)dµ
−∞

using the following two or three steps


For j = 1, 2, ..., J sample
step1: µ[j] from f (µ|x) i.e. from the posterior distribution.
step2: X ∗[j] from f (y |µ[j] ) assumed data generating mechanism
Then X ∗[1] , X ∗[2] , ..., X ∗[n] are i.i.d sample for f (y |x)
For more information about sampling predictive posterior
distribution, read more about MARKOV CHAIN MONTE-CALRO
(MCMC) sampling techniques.

32/33
Ivivi Joseph Mwaniki The University of Nairobi DOM STA 402 Bayesian inference December 2, 2023

You might also like