0% found this document useful (0 votes)
11 views91 pages

Lecturer Notes

The document contains lecture notes for a course on Loss Distribution and Credibility Theory aimed at students in actuarial science. It covers various statistical distributions, extreme value theory, credibility theory, ruin theory, generalized linear models, and run-off triangles, providing foundational knowledge for data analysis in insurance. Prerequisites for the course include BAS260 and BAS310, and it emphasizes the application of risk theories to environmental problem-solving.

Uploaded by

Richard シ
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views91 pages

Lecturer Notes

The document contains lecture notes for a course on Loss Distribution and Credibility Theory aimed at students in actuarial science. It covers various statistical distributions, extreme value theory, credibility theory, ruin theory, generalized linear models, and run-off triangles, providing foundational knowledge for data analysis in insurance. Prerequisites for the course include BAS260 and BAS310, and it emphasizes the application of risk theories to environmental problem-solving.

Uploaded by

Richard シ
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SCHOOL OF BUSINESS, ECONOMICS AND

MANAGEMENT
LOSS DISTRIBUTION AND CREDIBILITY THEORY
BAS380

LECTURE NOTES

MWALE RICHARD

January 27, 2025


Abstract
This course is designed for students from various backgrounds, specifically in actuarial science,
who need to understand data and aspire to develop new methods for data analysis, especially with
applications to solving environmental problems. It introduces the risk theories used in short-term in-
surance and will also equip students with the ability to price short-term products and perform simple
short-term policy valuations.

A certain amount of mathematical maturity is necessary to study and enjoy this course. This
course has BAS260 and BAS310 as prerequisites.

1
Contents
1 Loss Distribution 4
1.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.2 Poisson Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.3 Gamma Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.4 Exponential Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
1.5 Pareto Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
1.6 Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.7 Lognormal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
1.8 Weibull Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
1.9 Burr Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
1.10 Mixture Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17
1.11 Tutorial sheet 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20

2 Extreme Value Theory 22


2.1 Extreme Value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
2.2 Extreme Value Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23
2.2.1 Choosing Between the Gumbel and Frechet distribution . . . . . . . . . . . . . 24
2.3 Peak Over Threshold Approach (Generalized Pareto Distribution) . . . . . . . . . . . . 25

3 Credibility Theory 28
3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.2 Credibility Premium Formula . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28
3.3 Credibility Factor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
3.4 Bayesian Credibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30
3.4.1 Poisson- Gamma Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31
3.4.2 Normal-Normal Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38
3.5 Empirical Bayes Credibility Theory: Model 1 . . . . . . . . . . . . . . . . . . . . . . . 42
3.6 Empirical Bayes Credibility Theory: Model 2 . . . . . . . . . . . . . . . . . . . . . . . 48
3.7 Tutorial sheet 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53

4 Ruin Theory 55
4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
4.2 Surplus Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56
4.3 Probability of Ruin in Continuous time . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
4.4 Probability of Ruin in Discrete time . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57
4.5 Poisson and Compound Poisson Process . . . . . . . . . . . . . . . . . . . . . . . . . . 58
4.5.1 Poisson Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
4.5.2 Compound Poisson Process . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
4.5.3 Probability of Ruin in the short term . . . . . . . . . . . . . . . . . . . . . . . . 60
4.5.4 Premium Security Loading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
4.6 Adjustment Coefficient and Lundberg’s Inequality . . . . . . . . . . . . . . . . . . . . . 62
4.6.1 Lundberg’s Inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62
4.6.2 Adjustment coefficient- Compound process . . . . . . . . . . . . . . . . . . . . 63
4.7 Tutorial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64

5 Generalized Linear Models 66


5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
5.2 Exponential families Distibution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66
5.2.1 Normal Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67

2
5.2.2 Mean and Variance of Exponential families distribution . . . . . . . . . . . . . . 68
5.2.3 Poison Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
5.2.4 Binomial Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69
5.2.5 Gamma Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71
5.3 Components of GLM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73
5.4 Tutorial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76

6 Run-Off Triangles 77
6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
6.2 Types of Reserves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77
6.3 Presentation of claims data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
6.3.1 Components of the table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
6.4 Estimating future claims . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 78
6.5 Basic chain ladder method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79
6.6 Inflation adjusted chain ladder method . . . . . . . . . . . . . . . . . . . . . . . . . . . 82
6.6.1 Dealing with past inflation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82
6.6.2 Dealing with future inflation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85
6.7 Tutorial . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87

3
1 Loss Distribution
Objectives:

• Describe the properties of the statistical distributions which are suitable for modelling individual
and aggregate losses.

• Derive moments and moment generating functions (where defined) of loss distributions including
the gamma, exponential, Pareto, generalised Pareto, normal, lognormal, Weibull, Burr distributions
and mixture distribution

• Apply the principles of statistical inference to select suitable loss distributions for sets of claims.

• Estimate the parameters of a failure time or loss distribution when the data is complete, or when it
is incomplete, using maximum likelihood and the method of moments.

1.1 Introduction
General insurance companies need to investigate claims experience and apply mathematical techniques
for many purposes. These include;
Premium rating (i.e. deciding what premium rates to charge policyholders)
Reserving (i.e. assessing how much money should be set aside to cover the cost of claims)
Reviewing reinsurance arrangements
Testing for solvency (i.e. assessing the company’s financial position)
The total amount of claims in a particular period is a quantity of fundamental importance to the proper
management of an insurance company. The key assumption in all the models studied here is that the
occurrence of a claim and the amount of a claim can be studied separately. Thus, a claim occurs accord-
ing to some simple model for events occurring in time, then the amount of the claim is chosen from a
distribution describing the claim amount.
In this chapter we will look at loss distributions, which are a mathematical method of modelling indi-
vidual claims. We will introduce some new statistical distributions and we will see how these can be
“fitted” to observed claims data. We can then test for goodness of fit, and use the fitted loss distributions
to estimate probabilities. These distribution helps the insurance companies model either the number of
claims or the amount of individual claims (price associated with the claim). Basically every insurance
company is very much concerned with the total number of claim it receives in a given period (say 1 year).
Knowing the approximate number of claims and the associated amounts (aggregated amounts) in a given
period helps the insurance company plan well and ensure smooth running of the company. There are two
major concerns when we looks at the claims

• How frequent the claims occur

• the claim amount (i.e the amount associated with the claims)

However, in this course we are more concerned with the severity of the claim that is the claim amount,
and the model or distribution that can be used to determine the claim amounts. There are various statis-
tical models that can be used to understand the distribution on these losses. The most important thing is
the need to find model or distribution to best fit the data. Basically we assume that the claims follows a
particular known distribution. The major challenge in all this process is knowing exactly which distribu-
tion the claim comes from.
Below are the commonly known probability distributions and their properties that can be used to deter-
mine the number of claims and the associated amount to each individual claim.

4
1.2 Poisson Distribution
When a random variable X has a Poisson distribution with parameter λ > 0, its probability function is
given by
e−λ λ x
P(X = x) = x = 0, 1, 2, . . . (1)
x!
The moment generating function is
t
MX (t) = eλ (e −1)

The moment of X can be found from the moment generating function. For example,

Mx′ (t) = λ et MX (t)

and
MX′′ (t)λ et MX (t) + (λ et )2 Mx (t)
From which it follows that E[X] = λ and E[X] = λ 2 + λ so that Var[X] = λ . We use the notation Poi(λ )
to denote a Poisson distribution with parameter λ .

Task 1.2.1 Use the MGF and obtain expression for the mean and variance of the Poisson distribution.

Example 1.2.1 An insurance company uses an Poisson distribution to model the cost of repairing insured
vehicles that are involved in accidents. Find the maximum likelihood estimate of the mean cost, given
that the total cost of repairing a sample of 2500 vehicles was $35, 000.

Solution 1.2.1 Given that


n
∑ xi = $35, 000 and n = 2500
i=1
Then using the maximum likelihood estimator method
n
L(λ |x) = ∏ p(λ |x)
i=1

e−λ λ x
n
L(λ |x) = ∏
i=1 x!

e−nλ λ ∑ x
L(λ |x) = n
∏i=1 xi !
Taking the natural log on both side we have
" #
e−nλ λ ∑ x
ln[L(λ |x)] = ln
∏ni=1 xi !
!
n n
ln[L(λ |x)] = −nλ + ∑ x ln(λ ) − ln ∏ xi!
i=1 i=1

Differentiating w.r.t λ
" !#
n n
∂ ∂
ln[L(λ |x)] = −nλ + ∑ x ln(λ ) − ln ∑ x!
∂λ ∂λ i=1 i=1

5
∂ ∑n xi
= −n + i=1
∂λ λ

Letting ∂λ =0
∑ni=1 xi
=⇒ −n + =0
λ
Solving for λ we get
1 n
λ̂ = ∑ xi = X
n i=1
Thus
1
λ̂ = (35000) = 14
2500
Therefore maximum likelihood estimate of the mean cost λ̂ = 14.

1.3 Gamma Distribution


Suppose a random variable X ∼ Gamma(α, λ ) then its probability density function is given by is given
by
λ α α−1 −λ x
f (x) = x e , x > 0, α, λ > 0 (2)
Γ(α)
where Γ(α) is the gamma function is defined as
Z ∞
Γ(α) = xα−1 e−x dx (3)
0

Properties of Gamma Function

(a) Γ(α) = (α − 1)Γ(α − 1) for α > 1.

(b) Γ(α) = (α − 1)! if α is a positive integer



(c) Γ( 21 ) = π

In the special case when α is an integer the distribution is known as an Erlang distribution, and repeated
integration by parts gives the distribution function as
α−1 j
−λ x (β x)
F(x) = 1 − ∑e , x ≥ 0.
j=0 j!

The moments and moment generating function of the gamma distribution can be found by noting that
Z ∞
f (x)dx = 1 yields
0
Z ∞
Γ(α)
xα−1 e−λ x = (4)
0 λα
The nth moment is
λ α α−1 −λ x
Z ∞
n n
E[X ] = x x e dx
0 Γ(α)
λ α Z ∞
xn+α−1 e−λ x dx
Γ(α) 0

6
and from equation (4) it follows that
Γ(α + n)
E[X n ] =
λ n Γ(α)
For n = 1

Γ(α + 1) αΓ(α)
E[X 1 ] = E[X] = =
λ 1 Γ(α) λ Γ(α)
α
E[X] = as the mean
λ
and for n = 2
Γ(α + 2) α(α + 1)Γ(α)
E[X 2 ] = =
λ 2 Γ(α) λ 2 Γ(α)
α(α + 1)
E[X 2 ] =
λ2
So that
Var[X] = E[X 2 ] − (E[X])2
α(α + 1) α
Var[X] = −
λ2 λ
α
Var[X] = 2 as the variance
λ
We can find the moment generating function in a similar fashion. As

λ α α−1 −λ x
Z ∞
Mx (t) etx x e dx
0 Γ(α)

λα
Z ∞
MX (t) = xα−1 e−(λ −t)x dx
Γ(α) 0
By application of equation (4) we get
 
λα Γ(α)
MX (t) =
Γ(α) (λ − t)α

λα
MX (t) =
(λ − t)α
 α
λ
MX (t) = (5)
λ −t
NB: Equation (4), λ > 0. Hence in order to apply equation (4) to equation (5) we require that λ −
t > 0, so that the moment generating function exists when λ > t and does not exist when t ≥ λ . The
gamma distribution is one of the most important distributions for modelling because it has very tractable
mathematical properties. As we will see, it is also very useful in creating other distributions, but by itself
is rarely a reasonable model for insurance claim sizes.

Example 1.3.1 Based on an analysis of past claims, an insurance company believes that individual
claims in a particular category for the coming year will follows a Gamma distribution with parameters
α = 0.44 and λ = 0.0008. Estimate the probability of claims that will exceed K25, 000,

7
 
α α
Solution 1.3.1 Since X ∼ Gamma(α, λ ) then we can have that X ≈ N λ , λ2 , given that α = 0.44 and
λ = 0.0008 so that
α 0.44
µ= = = 550
λ 0.0008
α 0.44
σ2 = 2 = = 687500
λ (0.0008)2
P(X > 25000)
 
x−µ 25000 − 550
P > √
σ 687500
P(Z > 29.49) = 1 − P(Z < 29.49)
1 − Φ(29.49)
Since Φ(29.49) = 1
1−1 = 0
Therefore the estimated probability of the claims that will exceed K25, 00 is 0. Hence we conclude that it
is impossible to have claims that will exceed K25, 00.

Task 1.3.1 Use the χ 2 -distribution and workout example (1.3.1)

1.4 Exponential Distribution


The exponential distribution is a special case of the gamma distribution. It is just a gamma distribution
with α = 1. Hence , the exponential distribution with parameter λ > 0 is specified as follows from the
gamma distribution.
Since
λ α α−1 −β x
f (x) = x e , x > 0, α, λ > 0
Γ(α)
Let α = 1, then
λ 1 1−1 −λ x
f (x) = x e , x>0
Γ(1)
Since Γ(1) = 0! = 1, then
f (x) = λ e−λ x , for x>0 (6)
as the probability density function and has the cumulative density function

F(x) = 1 − e−λ x for x>0

n!
E[X n ] = (7)
λn
which give the nth moment of the distribution. In a similar way from equation (5), with α = 1

λ
MX (t) = t <λ
λ −t
thus we get the moment generating function. From E[X] = αλ and Var[X] = α
λ2
with α = 1 we get the
mean and the variance of the exponential distribution given by
1
E[X] =
λ

8
and
1
Var[X] =
λ2
respectively.
The exponential distribution is often used in developing models of insurance risks. This usefulness stems
in a large part from its many and varied tractable mathematical properties. However, a disadvantage
of the exponential distribution is that its density is monotone decreasing, a situation which may not be
appropriate in some practical situations.
1
Example 1.4.1 The random variable X has an exponential distribution with mean λ. It is found that
MX (−λ 2 ) = 0.2. Find λ .

Solution 1.4.1 Since X is having an exponential distribution with mean λ1 , we have that
Z ∞
tx
MX (t) = E[e ] = etx f (x)dx
0
Z ∞
MX (t) = etx λ e−λ x dx
0
Z ∞
MX (t) = λ e−(β −t)x dx
0
" #∞
e−(λ −t)x
MX (t) = λ −
λ −t
0
 
1 λ
MX (t) = λ 0 + =
λ −t λ −t
Then
λ 1
Mx (−λ 2 ) = 2
= = 0.2
λ −λ 1+λ
=⇒ λ = 4
Therefore λ = 4

Example 1.4.2 A ZISC general insurance uses an exponential distribution to model the cost of repairing
insured house that are damaged by accidents. Find the maximum likelihood estimate of the mean cost,
given that the average cost of repairing a sample of 50 houses was $2, 200.

Solution 1.4.2 Since f (x) = λ eλ x , then


n
L(x|λ ) = ∏ f (x|λ )
i=1

n
L(x|λ ) = ∏ λ eλ xi
i=1

L(x|λ ) = λ eλ x1 · λ eλ x2 · · · · β eλ xn
L(x|λ ) = λ n eλ ∑ x1
Taking the natural log on both sides

ln[L(x|λ )] = ln[λ n eλ ∑ x1 ]

9
ln[L(x|λ )] = n ln(λ ) − λ ∑ x1
Differentiating the above equation w.r.t λ and equating it to zero.
∂ ln[L(x|λ )] n
= − ∑ x1 = −0
∂λ λ
1
λ= 1
n ∑ xi
1
λ̂ =

since we have n = 50 and x̄ = 2, 200 then we have
1
λ̂ =
2200
Therefore the maximum likelihood estimate of the mean cost, given that the average cost of repairing a
1
sample of 50 houses was $2, 200 is λ̂ = 2200

1.5 Pareto Distribution


Suppose a random variable X has a Pareto distribution with α > 0 and λ > 0, then its density function is
given by
αλ α
f (x) = , x>0 (8)
(λ + x)α+1
By integrating this density function we find that the cumulative density function is
 α
λ
F(x) = 1 − , x≥0
λ +x
Clearly, the shape parameter α and the scale parameter λ are both positive.
NB; For the general Pareto distribution the density function is given by

Γ(α + k)λ α xk−1


fk (x) =
Γ(α)Γ(k)(λ + x)α+k
when k = 1, we get the Pareto distribution
αλ α
f (x) =
(λ + x)α+1
Whenever moments of the distribution exist they can be found from the expression below
Z ∞
n
E[X ] = xn f (x)dx
0

by using integration by parts Thus, we can use the following approach to determine them individually.
Since the integral of the density function over (0, ∞) equals 1 we have
dx 1
Z ∞
=
0 (λ + x)α+1 αλ α
as an identity which holds provided that α > 0.
The nth -moment is given by
Γ(α − n)Γ(n + 1) n
E[X n ] = λ , (9)
Γ(α)

10
So that from (9) we get mean E[X] and variance Var[X] of the Pareto distribution are given by

λ
E[X] =
α −1
And
2λ 2 αλ 2
E[X 2 ] = , α > 2 =⇒ Var[X] =
(α − 1)(α − 2) (α − 1)2 (α − 2)
respectively
The Pareto distribution is often appropriate model for claim size distribution particularly where excep-
tionally large claim may occur, this is due to its extremely thick tail

Example 1.5.1 A random sample of claims with n = 20 from a distribution believed to be Pareto with
parameters α and λ gives values such that

∑ x = 1508 ∑ x2 = 257212
use the methods of moments to estimate the parameters

Solution 1.5.1 Since we know that


λ 1 1508
E[X] = = Σx1 =
α −1 n 20
λ 1508
=⇒ =
α −1 20
λ = 75.4(α − 1)
λ 2 = (75.4)2 (α − 1)2 (10)
Also
2 2λ 2 1 257212
E[X ] = = Σxi2 =
(α − 1)(α − 2) n 20
2λ 2
=⇒ = 12860.6 (11)
(α − 1)(α − 2)
By substituting equation (10) in equation (11) we get

2(75.4)2 (α − 1)
= 12860.6
(α − 2)

By solving for α we get


α ≈ 9.630
and using equation (10) we get
λ ≈ 650.67
Therefore α = 9.630 ad λ = 650.67.
alternatively you can also use the
αλ 2
Var[X] =
(α − 1)2 (α − 2)
instead of E[X 2 ]

11
1.6 Normal Distribution
When a random variable X has a normal distribution with parameters µ and σ 2 , its density function is
given by
1 (x−µ)2

f (x) = √ e 2σ 2 − ∞ < x < +∞
σ 2π
We use the notation N(µ, σ 2 ) to denote a normal distribution with parameters µ and σ 2 The standard
normal distribution has parameters 0 and 1 and its distribution function is denoted Φ where
Z z
1 z2
Φ(z) = √ e− 2 dz
0 2π

A key relationship is that if X ∼ N(µ, σ 2 ) and if z = x−µ


σ , then Z ∼ N(0, 1).
The moment generating function of the normal distribution is
1 2t 2
Mx (t) = eµt+ 2 σ (12)

From which it can be shown that E[X] = µ and Var[X] = σ 2

Example 1.6.1 Derive the formula for the MGF of the standard normal distribution

Solution 1.6.1 The PDF of N(0, 1) is


1 x2
f (x) = √ e− 2

So th MGF is Z +∞
1 x2
MX (t) = E[e ] = tx
etx √ e− 2 dx
−∞ 2π
Z +∞
1 1 2
Mx (t) = √ e− 2 (x −2tx) dx
−∞ 2π
By using completing the square we have
Z +∞
1 1 2 2 2
Mx (t) = √ e− 2 (x −2tx−t +t ) dx
−∞ 2π
Z +∞
1 2 1 1 2 2
Mx (t) = e 2 t √ e− 2 (x −2tx+t ) dx
−∞ 2π
Z +∞
1 2 1 1 2
Mx (t) = e 2 t √ e− 2 (x−t) dx
−∞ 2π

Since the integral is just the area under the graph of the PDF of a N(t, 1) distribution which equals 1.
Hence the MGF of the standard normal distribution is
1 2 1 2
Mx (t) = e 2 t (1) = e 2 t

1.7 Lognormal Distribution


Suppose a random variable X has a lognormal distribution with parameters µ and σ , where −∞ < µ < +∞
and σ > 0, then its density function is given by

1 (ln x−µ)2

f (x) = √ e 2σ 2 , x > 0.
xσ 2π

12
The distribution function or CDF is obtained by integrating the density function as follows
(ln y−µ)2
Z x
1 −
F(x) = √ e 2σ 2 dy
0 yσ 2π
By substituting z = ln y it gives
(z−µ)2
Z ln x
1 −
F(x) = √ e 2σ 2 dz
−∞ yσ 2π
As the integrand is N(µ, σ 2 ) density function
 
ln x − µ
F(x) = φ
σ
However, probabilities under the lognormal distribution can be calculated from the standard normal dis-
tribution function. We use the notation LN(µ, σ ) to denote a lognormal distribution with parameters µ
and σ .
From the proceeding argument it follows that if X ∼ LN(µ, σ 2 ) then ln x ∼ N(µ, σ 2 ). This relationship
between the normal and lognormal distribution is extremely usefully, particularly in deriving moments.
If X ∼ LN(µ, σ 2 ) and Y = ln X, then
1 2 n2
E[X n ] = E[enY ] = MY (n) = eµn+ 2 σ
where the final equality follows by equation (12).
The log-normal distribution is very useful in modelling of claim sizes. For large σ its tail is (semi-)heavy
– heavier than the exponential. For small σ the log-normal resembles a normal distribution, although this
is not always desirable.
Example 1.7.1 If the the individual claim amount X follows a lognormal distribution with PDF
"  #
1 ln x − µ 2

1
f (x) = √ exp − , x>0
xσ 2π 2 σ
Derive a formula for average claim amount.
Solution 1.7.1
"  #
1 ln x − µ 2

1
Z ∞ Z ∞
E[X] = x f (x)dx = x √ exp − dx
0 0 xσ 2π 2 σ
ln x−µ dx
Letting z = σ =⇒ dz = xσ and x = eµ+σ z gives
Z +∞
1 1 2
E[X] = eµ+σ z √ e− 2 z dz
−∞ σ 2π
Z +∞
1 1 2
E[X] = eµ √ e− 2 (z −2σ z) dz
−∞ σ 2π
Using completing the square to evaluate this integral we have that
Z +∞
1 1 2 2 2
E[X] = e µ
√ e− 2 (z −2σ z+σ −σ ) dz
−∞ σ 2π
Z +∞
µ+ 21 σ 2 1 1 2
E[X] = e √ e− 2 (z−σ ) dz
−∞ σ 2π
The function in the last integral is the PDF of a random variable that has a N(σ , 1) distribution which is
equal to 1 then we have that
1 2 1 2
E[X] = eµ+ 2 σ (1) = eµ+ 2 σ
Therefore the formula for average claim amount is given by
1 2
E[X] = eµ+ 2 σ

13
1.8 Weibull Distribution
1
If V is an exponential variate, then the distribution of X = V γ , γ > 0 is called the Weibull distribution.
Its density function is is given by
γ
f (x) = cγxγ−1 e−cx x > 0, (13)

where c = λ −γ and the CDF or distribution function is given by


γ
F(x) = 1 − e−cx x>0

For γ = 2 it is known as the Rayleigh distribution. The Weibull distribution is roughly symmetrical for
the shape parameter γ ≈ 3.6.
When γ is smaller the distribution is right-skewed and when γ is large it is left-skewed. If γ > 1 the
distribution tail is lighter than the exponential distribution and if γ < 1 the tail becomes heavier than the
exponential distribution but lighter when compared to the Pareto distribution like in the case of a Pareto
distribution the moment generating function is infinite. The kth raw moment can be shown to be
 
− γt t
Mx (t) = c Γ 1 +
γ

So that nth moment is given by  


n − nγ n
E[X ] = c Γ 1+ (14)
γ
For the Weibull distribution, the maximum likelihood and method of moment estimators can be only be
evaluated numerically.
Task 1.8.1 Obtain expression for the mean and variance of the Weibull distribution.

Example 1.8.1 A random sample of 100 claim amounts x1 , x2 , . . . , x100 observed from a Weibull distribu-
tion with parameter γ = 2 where c is unknown. For these data:
n n
∑ xi = 487, 926 ∑ xi2 = 976, 444, 000 and median = 4, 500
i=1 i=1

(a) Determine the maximum likelihood estimator for c.

(b) Estimate the value of c using the method of moments.

(c) Calculate the method of percentiles estimate of c.

Solution 1.8.1 (a) Since


γ
f (x) = γcxγ−1 e−cx x > 0,
But γ = 2 then we have
2
f (x) = 2ce−cx x > 0,
Likelihood function is
n
L(x|c) = ∏ f (xi |c)
i=1
n
2
L(x|c) = ∏ 2ce−cxi
i=1
n 2
L(x|c) = 2n cn e−c ∑i=1 xi

14
Taking the natural log we have
n
ln L(x|c) = n ln 2 + n ln c − c ∑ xi2
i=1

Differentiating w.r.t c
n
∂ ln L(x|c) n
= − ∑ xi2
∂c c i=1
Equating to 0 and solving for c
n
ĉ =
∑ni=1 xi2
Since n = 100 and ∑ni=1 xi2 = 976444000 then

100
ĉ = = 1.024 × 10−7
976444000
Thus ĉ = 1.024 × 10−7

(b) Since  
− 1γ 1
E[X] = c Γ 1+
γ
But γ = 2 we get

1 n
 
− 12 1 487926
E[X] = c Γ 1+ = ∑ xi = = 4879.26
2 n i=1 100

1√
Using Γ(1.5) = 12 Γ 1

2 = 2 π, then

1 1√
c− 2 π = 4879.26
2
Solving for c we get
!−2
4879.26
ĉ = 1√
= 3.299 × 10−8
2 π

Therefore ĉ = 3.299 × 10−8

(c) For the median M, which satisfies F(M) = 0.5 using CDF we have
2
F(M) = 1 − e−cM = 0.5
2
F(4500) = 1 − e−c(4500) = 0.5
2
1 − ec(4500) = 0.5
2
e−c(4500) = 0.5
Taking the natural log and solving for c we get
ln 0.5
ĉ = = 3.423 × 10−8
(4500)2

Therefore ĉ = 3.423 × 10−8

15
1.9 Burr Distribution
1
Suppose Y is a Pareto random variable, then the distribution of X = Y γ is known as the Burr distribution
and its density and distribution functions are given by

xγ−1
f (x) = γαλ α x>0 (15)
(λ + xγ )α+1
And  α
λ
F(x) = 1 − x>0
λ + xγ
The t th raw moment
   
1 t t t
Mx (t) = λ Γ 1+
γ Γ α− exist only for t < γα
Γ(α) γ γ

So that nth moment is given by


n    
n λγ n n
E[X ] = Γ 1+ Γ α− (16)
Γ(α) γ γ
Just like the weibull distribution that the moment generating function do not exist. The maximum likeli-
hood and method of moments estimators for the Burr distribution can only be evaluated numerically.
Task 1.9.1 Obtain expression for the mean and variance of the Burr distribution.

Example 1.9.1 A random variable X has a Burr distribution with parameters γ = 2 and λ = 500 . Prove
that the maximum likelihood estimate of the parameter α, based on a random sample x1 , x2 , . . . , xn is
given by
n
α̂ = n
∑i=1 ln (500 + xi2 ) − n ln 500
Hence evaluate this estimator based on a sample consisting of six values 53, , 54, 109, 114, 163 and 181.

Solution 1.9.1 The likelihood function is


n
L(α, λ , γ) = ∏ f (x|α, λ , γ)
i=1

n γ−1
xi
L(α, λ , γ) = ∏ γαλ α γ
i=1 (λ + xi )α+1
γ−1
n n nα ∑ni=1 xi
L(α, λ , γ) = γ α λ γ
∏ni=1 (λ + xi )α+1
Taking the natural log we have
" #
n n
γ
ln L(α, λ , γ) = n ln γ + n ln α + nα ln λ + (γ − 1) ln ∑ xi − (α + 1) ln ∏(λ + xi )
i=1 i=1

n n
γ
ln L(α, λ , γ) = n ln γ + n ln α + nα ln λ + (γ − 1) ln ∑ xi − (α + 1) ∑ ln(λ + xi )
i=1 i=1
Differentiating w.r.t λ
n
∂ ln L(α, λ , γ) n γ
= + n ln λ − ∑ ln(λ + xi )
∂λ α i=1

16
Equating to 0 and solving for α we get
n
α= γ
∑ni=1 ln (λ + xi ) − n ln λ

But since γ = 2 and λ = 500 we have that


n
α̂ = hence proven
∑ni=1 ln (500 + xi2 ) − n ln 500

Given 53, , 54, 109, 114, 163 and 181. then we have that
n
∑ ln (500 + xi2) = 45.58683462392
i=1

so that
6
α̂ = ≈ 0.7230
45.58683462392 − 6 ln 500
Therefore α̂ = 0.7230

1.10 Mixture Distribution


The exponential distribution is one of the simplest models for insurance losses. Suppose that each indi-
vidual in a large insurance portfolio incurs losses according to an exponential distribution.
Practical knowledge of almost any insurance portfolio reveals that the means of these various distribu-
tions will differ among the policyholders. Thus the description of the losses in the portfolio is that each
loss follows its own exponential distribution, i.e. the exponential distributions have means that differ
from individual to individual.
A description of the variation among the individual means must now be found. One way to do this is
to assume that the exponential means themselves follow a distribution. In the exponential case, it is
convenient to make the following assumption. Let
1
λi =
θi

be the reciprocal of the mean loss for the ith policyholder. Suppose we assume that the variation among the
λi can be described by a known gamma distribution Gamma(α, β ). i.e. assume that λ ∼ Gamma(α, β )
where
β α α−1 −β λ
f (x) = λ e , λ >0
Γ(α)
Thus note that the PDF is in terms of λ with known values of α and β This formulation has much in
common with that used in Bayesian estimation. The fundamental idea in Bayesian estimation is that the
parameter of interest (here, λ ) can be treated as a random variable with a known distribution. However,
that the purpose here is not to estimate the individual λi , but to describe the aggregate losses over the
whole portfolio.
Estimation of the individual λi can be treated by Bayesian estimation, when the Gamma(α, β ) distribution
would be referred to as a prior distribution. In this problem of describing the losses over the whole port-
folio, the Gamma(α, β ) distribution is used to average the exponential distributions; the Gamma(α, β )
distribution is referred to as the mixing distribution and the resulting loss distribution as a mixture distri-
bution.
The random variable X represents the amount of a single randomly selected claim and E[X] represents

17
the mean claim amount for all risks in the portfolio. X.
To find the overall distribution of claim amounts, we need to work out the marginal distribution of X .
This is obtained by integrating the joint density function fX,λ , over all possible values of λ
The PDF of the mixture (or marginal) distribution of X is
Z ∞
fX (x) = fX,λ (x, λ )dλ
0
Z ∞
fλ (λ ) fX,λ (x|λ )dλ
0
Z ∞ α
β
λ α−1 e−β λ · λ e−λ x dλ
0 Γ(α)
βα
Z ∞
λ α e−(x+β )λ dλ
Γ(α) 0

We now make the integrand look like the PDF of gamma(α + 1, x + β ) distribution:

βα Γ(α + 1) (x + β )α+1 α −(x+β )λ


Z ∞
fX (x) = · λ e dλ
Γ(α) 0 (x + β )α+1 Γ(α + 1)

βα Γ(α + 1) (x + β )α+1 α −(x+β )λ


Z ∞
fX (x) = · λ e dλ
Γ(α) (x + β )α+1 0 Γ(α + 1)
The integration the PDF over all possible values of λ will yield 1, so that

β α Γ(α + 1)
fX (x) =
Γ(α) (x + β )α+1

αβ α
fX (x) = , x>0
(x + β )α+1
which can be recognised as the PDF of the Pareto distribution, Pa(α, β ). This result gives a very nice in-
terpretation of the Pareto distribution: Pa(α, β ) arises when exponentially distributed losses are averaged
using a Gamma(α, β ) mixing distribution.

Example 1.10.1 The annual number of claims an individual policy in a portfolio has a Poisson(µ) dis-
tribution. The variability in µ among policies is modelled by assuming that over the portfolio, individual
values of µ have a Ga(α, β ) distribution. Derive mixture distribution for the annual number of claims
from each policy in the portfolio.

Solution 1.10.1 Let X be the total number of claims, we have


Z ∞
P(X = x) = Px|µ (x) f µ (µ)dµ
0
Z ∞ −µ x α
e µ β
= µ α−1 e−β µ dµ
0 x! Γ(α)
βα
Z ∞
= µ x+α−1 e−β µ−µ dµ
0 x!Γ(α)
βα
Z ∞
= µ (x+α)−1 e−(β +1)µ dµ
0 x!Γ(α)

18
We make this integral look like that of a gamma(x + α, β + 1) distribution.

Γ(x + α) (β + 1)x+α β α
Z ∞
P(X = x) = · µ (x+α)−1 e−(β +1)µ dµ
0 (β + 1)x+α Γ(x + α) x!Γ(α)

Γ(x + α) β α (β + 1)x+α
Z ∞
P(X = x) = µ (x+α)−1 e−(β +1)µ dµ
(β + 1)x+α x!Γ(α) 0 Γ(x + α)
The integral yields 1 so that we have

Γ(x + α) β α
P(X = x) =
(β + 1)x+α x!Γ(α)

Γ(x + α) βα
P(X = x) = x = 0, 1, 2, . . .
x!Γ(α) (β + 1)x+α
By rearranging the above expression we get
  α  x
x+α −1 β 1
P(X = x) =
x β +1 β +1
β
This is a negative binomial distribution with parameter P = β +1 and k = α

Summary
• Individual claim amounts can be modelled using a loss distribution. Loss distributions are often
positively skewed and long-tailed.

• Distributions such as the exponential, normal, lognormal, gamma, Pareto, generalised Pareto, Burr
or Weibull distribution are commonly used to model individual claim amounts.

• Once the form of the loss distribution has been decided upon, the values of the parameters must be
estimated. This may be done using the method of maximum likelihood, the method of moments or
the method of percentiles. Goodness of fit can then be checked using a chi square test.

19
1.11 Tutorial sheet 1
1. Let X be a continuous random variable with density function
(
f (x) = θ e−θ x for x > 0
0, elsewhere

If the median of this distribution is 26, determine θ .

2. The claim size of an insurance portfolio follows the Pareto distribution with mean and variance of
20 and 900, respectively. Find

(a) the shape α and scale parameters λ .


(b) 90th percentile of this distribution.

3. Show that the mean and variance of the weibull distribution is given by
 !
1 2
     
− 1γ 1 − 2γ 2
E[X] = λ Γ 1+ and Var[X] = λ Γ 1+ − Γ 1+
γ γ γ

4. Suppose that the probability distribution of the lifetime of cancer patients (in months) from the
time of diagnosis is described by the Weibull distribution with shape parameter γ = 1.2 and scale
parameter λ = 33.33.

(a) Find the probability that a randomly selected person from this population survives at least 12
months.
(b) A random sample of 10 patients will be selected from this population. What is the probability
that at most two will die within one year of diagnosis with a success chance of 0.254
(c) Find the 99th percentile of the distribution of lifetimes.

5. Given the single-parameter Pareto with distribution function:

500 α
 
F(x) = 1 − x > 500
x

has the following five observations: 521, 658, 702, 819, 1217. Determine the maximum likelihood
estimator of α

6. Losses have a Pareto distribution with parameters α and θ . The 10th percentile is θ − k. The 90th
percentile is 5θ − 3k. Determine the value of θ .

7. The cdf of a random variable is


F(x) = 1 − x−2 , x≥1
Determine the mean, median, and mode of this random variable.

8. Losses follows a single-parameter Pareto distribution with density function given by

f (x) = λ x−(λ +1), x>1, 0<λ <∞


.

A random sample of size five produced three losses with values 2, 4 and 9, the other two losses
exceeds 20. Determine the maximum likelihood estimate of λ .

20
9. The claim amount follows a shifted exponential distribution with probability density function given
by
x−µ
f (x) = θ −1 e− θ , for µ < x < ∞.
A random sample of claims amount X1 , X2 , . . . , X10 which can be summarized as
10 10
∑ Xi = 100 and ∑ Xi2 = 1306
i=1 i=1

Estimate the parameters θ and µ using the method of moment.

10. The annual number of claim an individual policy in a portfolio has a Poisson (θ ) distribution. The
variability in θ among policies is modelled by assuming that over the portfolio, individual values
of θ have a gamma(α, λ ) distribution.

(i) Derive the mixture distribution for the annual number of claim from each policy in the port-
folio.
(ii) Given that α = 4 and λ = 9. Find the probability that a policy holder chosen at random will
experience 5 claims over the next year.

11. Claim amounts Xi from a portfolio of insurance policies follow a gamma distribution with param-
eters k and λi . Each λi also follows a gamma distribution with parameters α and β

(a) Show that the mixture distribution of losses is a generalised Pareto, with parameters α, β and
k.
(b) Claim amounts are now assumed to be exponentially distributed with parameter λi . Show,
using your answer to part (a), that the mixture distribution of losses is now a standard Pareto
distribution with parameters α, β .

21
2 Extreme Value Theory
Objectives:

• Define extreme events and state/explain the concept of extreme events.

• Use Generalized extreme value distribution (GEV) to model extreme event.

• Use Generalized Pareto Distribution (GPD) to model extreme events.

• Measures of tail weight.

2.1 Extreme Value


There are many application in the world finance and risk management where we come across extreme
value or extreme events. are event that are unlikely to occur but if they occur they are severe and very
expensive to handle. These events create a big financial problem (damage) in other words, these event
can be called low frequency high severity kind of events.
Extreme value events can be divided into two sections theses are;

(a) Disaster caused by nature.

(b) Disaster caused by man.

Examples of disaster caused by nature are hurricanes, earthquakes, storm, tornadoes.


Examples of disasters caused by man are; pollution, stock market, war accidents.
The objective is to understand the behaviour of these events, because their occurrences cause a very big
damage, Thus the estimation of these events is very important. Now the question is that;

(i) How do we estimate the frequencies of the claims(how frequent these event occur).

(ii) How do we estimate the claim amount or severity associated to the damages caused by these ex-
treme events.

In general it is very difficult to make these estimations because very few data points are available and
hence the estimation will be very much uncertain. In this way, the estimation is purely based on the
assumptions that are being made by the analyst and statisticians.
The lack of enough data is really being covered for by the various assumptions made. These assumption
are always subjected to questions. Generally what happens in these kind of processes is that an analyst
or a statistician will chose some distribution e.g. Lognormal, Gamma or Pareto distribution,. Once the
distribution is chosen then data is fitted to that particular loss distribution.
In most case, the central values will get fitted quite appropriately, but the extreme values will not get
fitted properly. This implies that the tail kind of mechanism does not work, which means that whenever
we are dealing with extreme values, we need to use a completely different approach. For this reason we
can not combine this values(extreme values) with the regular one because the regular values can fit to any
distribution.
This is where the concept of extreme value theory comes from. The extreme value theory targets on the
extreme values or observations. In extreme value theory we discuss what kind of distribution is fitted to
data and how can we estimate the parameter from the data.

22
2.2 Extreme Value Theory
Suppose we have random variables x1 , x2 , . . . , xn which are independent and identically distributed (iid)
but we have no knowledge about the distribution of the xi ’s. Now we need to determine the particular
distribution where theses random variables (xi ’s) are coming from. It is therefore difficult to use the
traditional way of finding specific distribution where these random variables are coming from.
Suppose we have some data for some distribution , now we need to look at some extreme bar for this
particular data. For a given sample size of n, we need to find the maximum observation. In the very
general way we try to understand the behaviour of extreme values as the sample size n approaches infinity
(∞). Below is the theorem called Fisher-Tippet theorem which helps us to understand the behaviour of
the extreme values.

Theorem 2.2.1 Fisher Tippet Theorem states that the distribution of the maximum values Mn = max(x1 , x2 , . . . , xn )
approaches the generalized extreme values distribution as n approaches ∞.

The generalized extreme value distribution (GEV) cdf is given by;


 h i− 1
 − 1+ζ (x−µ)
 ζ

F(x) = e σ
for ζ ≠ 0
(x−µ)
e−e σ

for ζ = 0

Thus

 −[1+ζ y]− ζ1
F(x) = e y ]for ζ ̸= 0 (17)
e−e for ζ = 0

(x−µ)
Where y = σ and 1 + ζ (x−µ)
σ > 0. The GEV has three parameters

(a) Zeta (ζ ): is the tail index and it deals with the shape of the tail. Thus it is called the shape
parameter. The shape parameter(ζ ) ranges from −∞ to +∞.

(b) Mu (µ): is the location parameter and it deals with the central tendency of extreme values. The
location parameter (µ) ranges from −∞ to +∞.

(c) Sigma (σ ): is the scale parameter and it deals with the dispersion of the extreme value and it ranges
from 0 to +∞.

The notation used to denote the generalized extreme value (GEV) is X ∼ GEV (µ, σ , ζ ).
The generalized Extreme value distribution is categorized into three groups of distribution which are;

(a) Frechet Distribution

(b) Gumbel Distribution

(c) Weibull Distribution

By analyzing the shape parameter ζ we have the following cases;


If ζ > 0, then the tail is very heavy and the GEV distribution is called Frechet Distribution and density
function is given by
i 1
(x−µ) − ζ
h
− 1+ζ σ
F(x) = e ,x > 0 (18)

23
For µ = 0 and σ = 1 then we have
−1
F(x) = e−(1+ζ x)
ζ
,x > 0 (19)

Which is called the standardized Frechet distribution.


In this case, the power function increases faster that the tail is very heavy and generally the underlying
data comes from a t-distribution or a Pareto distribution. Thus the Frechet distribution can be comfortably
used in finance and insurance because of it heavy tail.
Task: Research on the graph of a Pareto distribution and state why it has a heavy tail.
If ζ = 0 then the tail is moderate and the GEV distribution is now called the Gumbel distribution and its
density function id given by

−y (x − µ)
F(x) = e−e , where y = , x∈R (20)
σ
And for µ = 0 and σ = 1 then
−x
F(x) = e−e (21)

which called the standard Gumbel distribution. Under the Gumbel distribution, the tail follows an expo-
nential pattern which implies that that the tail is relatively lighter. The Gumbel distribution has a very few
application in finance and insurance because of its relatively light tail. In this distribution, the underlying
data follows a normal or a lognormal distribution.
If ζ < 0 then the tail is very lighter and the GEV distribution becomes the Weibull distribution with the
distribution function given by
(x−µ)
F(x) = e− σ

The Weibull distribution has few application in finance because of its lighter tail.
Generally focus on Frechet and Gumbel distribution due to their heavy tails. Both Frechet and Gumbel
distribution are positively skewed but the Frechet is more positively skewed. The tail of the Frechet
distribution is much longer.
Suppose we are looking at a very extreme value of X and we need a higher probability of generating that
extreme value, then we can probably consider the Frechet distribution.

2.2.1 Choosing Between the Gumbel and Frechet distribution


The following is the procedure used when we want to decide which distribution to use;

(a) Knowing the Parent distribution helps to decide which distribution to use.

(i) If the parent distribution is a t-distribution or a Pareto distribution then we have to choose the
Frechet distribution to model the extreme value.
(ii) If the parent distribution is the normal or lognormal distribution then we choose the Gumbel
distribution.

(b) Tail index of ζ . We carry out the test of significant, if ζ is insignificant the we choose Gumbel
otherwise we go with the Frechet distribution.

(c) We choose the Frechet distribution because it is more conservative. Thus it is more positively
skewed and has a heavy tail.

24
Once we know what type of distribution we are dealing with, then we can estimate the parameter by using
the following method

(i) Maximum likelihood Estimate (µ, σ , ζ ).

(ii) Methods of Moments

(iii) Regression (OLS)

2.3 Peak Over Threshold Approach (Generalized Pareto Distribution)


When we are interested in looking at the distribution of modelling excess loss value over threshold. From
the previous section we saw that the Generalized Extreme Value (GEV) distribution is more concerned
with modelling the maximum or minimum values of the data set (Extreme values). Basically we were
concerned with determining a distribution that can be used to model these extreme values. In a similar
way, in this section we are interested in modelling maximum value which are great than a specific thresh-
old say u.
In order to achieve this we use a method called the Peak Over Threshold approach (POT).
The Peak Over Threshold approach method is one way to model extreme values. The main concept of
this method is that it uses a threshold to seclude values considered extreme to the rest of the data and then
create a model for the extreme value by modelling the tail of all the value exceeding the defined threshold.
This approach we then lead us to a specific distribution called the Generalized Pareto distribution (GPD).
In simpler terms, if we are interested in modelling the maximum or minimum values of a particular dis-
tribution, we are going to employ the GEV distribution to do that. But if our objective is to model the
values that are greater than a particular set threshold, thus we use the Pot approach where we use the
Generalized Pareto distribution to model such values.
Suppose X is iid with a distribution function of f (x) and a threshold u. The choices of the value u is
completely arbitrary, for whatever value of u we choose, the distribution of excess is given by

F(x + u) − F(u)
Fu (x) = P(X − u ≤ x|X > u) = (22)
1 − F(u)
Where X is a random variable and u is a given threshold with x = y − u being the excesses. This gives the
distribution of excess losses about a particular threshold u.
Suppose threshold (u) is getting large and larger then the distribution of Fu (x) approaches the Generalized
Pareto distribution. That is; Fu (x) ≈ F(x, ζ , σ ) u −→ ∞

Definition 2.3.1 A random variable X has a generalized Pareto distribution GPD(ζ , σ ), then its cumu-
lative distribution function is given by
  − 1
1 − 1 + ζσx
ζ
, ζ ̸= 0,

F(x) = x
(23)
1 − e− σ , ζ = 0, σ > 0

Where ζ is the shape parameter and its ranges for −∞ to ∞ and σ is the scale parameter. Here if ζ > 0
then it gives a heavy tail(longer tail). If ζ = 0 then it gives a moderates tail. and ζ < 0 then it gives a
lighter tail.
The advantage about this distribution is that x can come from any distribution. This is an important
advantage which implies that if we are choosing a reasonable threshold u, then the excess always follow
a GPD. Basically irrespective of whatever distribution the underlying data follows, the excess value
(losses) always follow a GPD.

25
This calls for careful choice of the values of the threshold(u), that will help us determine the values above
the threshold u. If u is very high then we will have few values above the threshold, in this case the
estimation of the parameters (ζ , σ ) will be very [Link] on the choice of u, the number of values
(observations) exceeding u will be known. The idea is to have the threshold higher enough for the excess
value to follow the GPD and lower enough to estimate the values of the parameters.
This just tells us that the choice of u is very useful in this entire process. If we manage to get the data for
excess loss, we can use any of the methods discussed to estimate the parameter ζ and σ .

Task 2.3.1 List and explain the methods used to determine the threshold in the POT approach.

Suppose we want to determine the value at risk (VaR) we use the formula
 −ζ
σ n
VaR = u + (1 − α) −1
ζ Nu
Where;
Nu = Number of excess loss
n=Sample size
u= Threshold
α = Significance level.

Example 2.3.1 A sports scientist is interested in analysing the probability that the javelin world record
may be broken next year and is intending to use EVT to do this. The sports scientist has obtained data
for the distances of all javelin throws from all javelin competitions in 2019. The total number of throws
recorded was 3, 000. The sports scientist has carried out an EVT analysis using the Generalised Pareto
Distribution by selecting only those throws that exceeded 50 metres. This resulted in the longest 150
throws being selected for the analysis. The following parameters were obtained from the EVT analysis:

σ = 15,

ζ =3
Determine the percentage of javelin throws that would be expected to exceed 70 metres next year.

Solution 2.3.1 Let u be the threshold so that u = 50 metre


Then
P(X > 70/X > 50) = 1 − P(X − u ≤ x/X > u)
= 1 − P(X − 50 ≤ x/X > 50)
= 1 − P(70 − 50 ≤ 20/X > 50)
=⇒ GPD(20) = P(70 − 50 ≤ 20/X > 50)
Using
 1
ζx −ζ

F(x) = 1 − 1 +
σ
Thus
 1
(3)(20) − 3

GDP(20) = F(20) = 1 − 1 + = 0.4152
15
Thus
P(X > 70/X > 50) = 1 − GDP(20) = 1 − 0.4152 = 0.5848

26
Now we are told that the sports scientist has selected 150 throws for the analyst out of 3000 so that
150
= 0.05
3000
of the data was selected. Hence we only have 0.05 = 5% of the data exceeding the threshold which
implies that
P(X > 50) = 0.05
So that
P(X > 70) = P(X > 70/X > 50) × P(X > 50)
P(X > 70) = (0.5848)(0.005) = 0.02924
Therefore the percentage of javelin throws that would be expected to exceed 70 metres next year is 2.924%

27
3 Credibility Theory
Objectives:

• Explain what is meant by the credibility premium formula and describe the role played by the
credibility factor.

• Explain the Bayesian approach to credibility theory and use it to derive credibility premiums in
simple cases.

• Explain the fundamental concepts of Bayesian statistics and use these concepts to calculate Bayesian
estimators.

– Explain the empirical Bayes approach to credibility theory, in particular its similarities with
and its differences from the Bayesian approach.
– State the assumptions and calculate credibility premiums for underlying the two models

3.1 Introduction
In this chapter we will look at the credibility theory which is a technique that can be used to determine
premiums or claim frequencies in general insurance.

Definition 3.1.1 Credibility is defined as a set of quantitative tool that allows an insurer to perform
prospective experience rating that is adjusting future premiums based on past experience on a risk or
group of risk.

In other words, credibility theory refers to the policies, tools or procedures used when analysing data in
order to estimate the risk and thus adjust the future premiums.
Here we are going to introduce the Bayesian approach to credibility and we will proceed to study credi-
bility theory from a practical stand point.

Task 3.1.1 Review the Bayesian estimation.

3.2 Credibility Premium Formula


The basic idea behind the credibility premium formula is very simple and very appealing. Our main goal
is to estimate the expected aggregate claim or possibly just the expected number of [Link] order to
achieve this, the following information is available:
X̄ is an estimate of expected aggregate claim or number of claim for the coming year based solely on the
data from the risk itself.
µ is an estimate of expected aggregate claim or number of claim for the coming year based on collateral
data i.e. data for risk similar to, but not necessarily identical to the particular risk under consideration.
The credibility premium formula or credibility estimate of the aggregate claims for this risk is

Z X̄ + (1 − Z)µ (24)

Where Z is called the credibility factor and it ranges between 0 and 1.

28
Example 3.2.1 A specialist insurer that provides insurance against breakdown of photocopying equip-
ment calculates its premiums using a credibility formula. Based on the company’s recent experience
of all models of flyer printer, the premium for this year should be K2000 per machine. The company’s
experience for a new model of flyer printer, which is considered to be more reliable, indicates that the
premium should be K600 per machine. If the credibility factor is 0.85, Determine premium should the
insurer charge for insuring the new model.

Solution 3.2.1 The premium based on the collateral data (including all machines) is µ = K2, 000
The premium based on the direct data from the risk itself the new model) is µ = K600
Given that the credibility factor isZ = 0.85, then we have
Premium Future = ZX̄ + (1-Z)µ
Premium Future = 0.85(600)+(1-0.85)(2000)
Premium Future = 810
Therefore the insurer charge K810 for insuring the new model.

3.3 Credibility Factor


The credibility factor Z is just a weighted factor. its values represents how much trust is placed on the
data from the risk itself (X̄) compared with the data from the large group (µ) as an estimate of next years
expected aggregate claim or number of claims.
The higher the value of Z, the more trust is placed on X̄ as compared with µ and vice-versa. Basically the
more data are from the risk itself, the higher the value of the credibility factor. Again the more relevant
the collateral data, the lower the value of the credibility factor. Another vital point about the credibility
factor is that its value should represent the amount of data available from the risk itself, it is independent
of the actual data from the risk itself (i.e. the value of X̄).
Now suppose Z was allowed to be dependent on X̄ then any estimate of the aggregate number of claims
say φ , taking a value between X̄ and µ could be written in the form of (24), by choosing Z to be equal to
φ −µ
Z=
X̄ − µ

Example 3.3.1 Verify if the expression of the credibility factor Z given above is correct.

Solution 3.3.1 To verify if the credibility factor is correct, we need to show that
Z X̄ + (1 − Z)µ = φ
where
φ −µ
Z= ,
X̄ − µ
Thus        
φ −µ φ −µ φ −µ X̄ − µ
Z X̄ + (1 − Z)µ = Z X̄ + 1 − µ= X̄ + X̄
X̄ − µ X̄ − µ X̄ − µ X̄ − µ
(φ − µ)X̄ + (X̄ − φ )µ φ X̄ − φ µ φ (X̄ − µ)
= = = =φ
X̄ − µ X̄ − µ X̄ − µ
Therefore Z X̄ + (1 − Z)µ = φ

verified.

29
3.4 Bayesian Credibility
In this section, we study credibility theory using the Bayesian approach. The Bayesian approach to
credibility constitute the following steps:

Step1: Prior Distribution


A prior distribution is adopted to describe the possible values of the unknown parameter under consider-
ation e.g the claim frequency or claim amount. The form of this prior distribution is obtained or should
be derived from the information provided by the collateral data.

Step 2: Likelihood Function


For any given value of the parameter there exist a certain probability of incurring the particular pattern of
claim observed in the direct data. This helps in determining the likelihood of a given pattern of claim as
a function of the unknown parameter.

Step 3: Posterior Distribution


The prior distribution can be combined with the likelihood function by simple multiplication to form a
posterior distribution for the parameter. Here the idea of the underlying proportional method applies in
the case where the unknown parameter is assumed to have a continuous distribution. Thus, we use the
proportional method formula given by

Posterior distribution ∝ Prior Distribution × Likelihood

to determine the posterior parameter.

Step 4: Loss Function


Once the posterior distribution is determined, then we use an appropriate loss function to minimize the
error in estimating the parameter. Ideally the loss function is employed to quantify how serious mis-
judging the parameter would be. The loss function should be based on the commercial consideration of
the financial effect of mis-estimating the parameter and hence the premium rates would have on the in-
surer’s business. We can therefore use any of the three loss function which are; Quadratic loss function,
Absolute loss function and All or nothing loss function.

Quadratic loss function


The Bayesian estimate for µ under the quadratic loss function is just the mean of the posterior distribution.

All -or-Nothing loss function


The Bayesian estimate under the all-or-nothing loss function is given by the mode of the posterior dis-
tribution, we need to different the PDF or equivalently differentiate the log of the PDF and equate it to
zero.

Absolution error loss function


The Bayesian estimate under absolute loss is given by the median M of the posterior distribution.

30
Step 5: Parameter Estimation
The loss function is then applied to the posterior distribution to determine the Bayesian estimate of the
unknown parameter. This Bayesian estimate is the parameter that minimizes the expected loss based on
the posterior distribution.
Now, the study of Bayesian credibility theory is demonstrated by considering two models these are;

(a) Poisson- Gamma Model

(b) Normal-Normal Model

3.4.1 Poisson- Gamma Model


This model is used to model the claim frequencies or the expected number of claim. Suppose the claim
frequency(expected number of claims) for a risk in the coming year needs to be estimated. This problem
can be divided into two part that is;

(i) The number of claims each year is assumed to follow a Poisson distribution with a parameter λ .
The parameter λ is unknown (X ∼ Poi(λ )). Thus the distribution function is given by

e−λ λ x
p(x) = x = 0, 1, 2, . . .
x!

(ii) The parameter λ which is unknown follows a gamma distribution with parameter α and β (λ ∼
Gamma(α, β )). So ideally we see that the gamma distribution is a prior distribution for the param-
eter λ . Hence the density function of λ is given by

β α α−1 −β λ
f (x) = λ e λ >0
Γ(α)

Now if the data from this risk is available showing the number of claim arising in each of the past n years.
This data will help us determine the likelihood function. Suppose we collected claims x1 , x2 , . . . , xn and
the mean of the actual distribution looking at the historical data is given by

1 n
∑ xi = X
n i=1

But if we are following the assumption that the parameter λ follows a gamma distribution then its mean
is given by
α

β
Since the random variable X represents the number of claims in the coming year from the risk, and the
distribution of X is a Poisson(λ ) distribution. Given that the prior distribution of the parameter λ is the
Gamma(α, β ) then the posterior distribution is obtained as follows;
Since
β α α−1 −β λ
Posterior distribution = λ e
Γ(α)
ne−λ λix
Likelihood = ∏
i=1 xi !
So that using
Posterior distribution ∝ Prior Distribution × Likelihood

31
We have that !
n e−λ λix
 
β α α−1 −β λ
Posterior distribution ∝ λ e ∏ xi!
Γ(α) i=1
  −nλ ∑ xi !
β α α−1 −β λ e λ
∝ λ e n
Γ(α) ∏i=1 xi !
 α  
β 1
∝ λ α−1 e−β λ e−nλ λ ∑ xi
Γ(α) ∏ni=1 xi !
| {z }
constant can be ignored

Posterior distribution = λ α−1+∑ xi e−λ (β +n)


Thus the posterior distribution is a gamma distribution with parameter α ∗ = (α + ∑ xi ) and β ∗ = (β + n).
Since the posterior distribution, then its mean can be easily obtained and is given by

α ∗ α + ∑ xi
=
β∗ β +n

This is called the expected value of λ given X = (x1 , x2 , . . . , xn ) that is

α + ∑ xi
E[λ |X] =
β +n
Now the objective is to break down this mean into a credibility theory mechanism. That is we want to
give some Z a weighted to X̄ and some (1 − Z) to µ such as

ZX + (1 − Z)µ

Here, X is the mean we get when using the data from the risk itself (x1 , x2 , . . . , x3 ) so that its mean is
1 α
n ∑ xi and µ is based on the prior distribution and its mean is β . So we break the mean of the posterior
distribution into the mean obtained from the data of the risk itself and the mean of the collateral data thus;
α + ∑ xi α ∑ xi
E[λ |X] = = +
β +n β +n β +n
α ∑ xi
β n
= β +n
+ β +n
β n

µ X
= β +n
+ β +n
β n

β n
= µ+ X
β +n β +n
By comparison with (24), we have that Z is given by

n β
Z= and 1−Z =
β +n β +n

where n is the number of years of historical data we are collecting, and β is an estimate from the prior
distribution.

32
 
α
We have seen that the estimate is a weighted average of the prior mean β and the maximum likelihood
 
estimate ∑nxi . The weighting factor Z applied to the MLE is

n
Z= , (25)
β +n
and is called the credibility factor.
Now suppose no data is available from the risk itself this implies n = 0, so that Z = 0. The we are
completely relying on the collateral data only, that is the only information available to help estimate the
parameter λ would be its prior
 distribution. In this case, the best estimate of λ would be the mean of the
prior distribution which is αβ .
On the other hand if only the data from the risk were available to estimate λ , that is we have n and
β = 0 this yields Z = 1, which means that the full weight we are giving in terms of estimating the claim
frequencyfor the coming year is purely based on the historical data and obviously the estimate of λ
∑ xi
would be n .
The value of Z depends on the amount of data available for the risk, n and the collateral information
through β . Recall that n is the number of year of data available whilst β is the variable of the prior
distribution for λ . Increasing n means the value of Z is higher which implies that we put much of our
trust on the data from the risk itself. On the other hand increasing β will make the variance decrease
(since the variance is βα2 , which depends on β ), thus giving the collateral data more confidence which
makes the credibility factor lower.

Example 3.4.1 Given that the prior distribution for λ is taken to be Gamma(100, 1). The actual number
of claims arising each year from the risk follows a Poisson distribution is given below;

Years(n) 1 2 3 4 5 6 7 8 9 10
Number of claims (xn ) 144 144 174 148 151 156 168 147 140 161

(a) Graph the credibility factor in successive years, and explain what is happening to the graph.

(b) Show that the estimated number of claims for Year 9 is lower than the estimated number of claims
for Year 8. State the reason why we have such a case.

(c) Graph the actual number of claim and the estimated number of claims on the same plot.

(d) Given that he prior distribution of λ is Gamma(500, 5). Show the graph of the credibility factor
and the estimated number of claim in successive years.

(e) Graph the credibility factor of the under the prior distribution Gamma(100, 1) and that of Gamma(500, 5)
on the same plot and state your observations.

Solution 3.4.1 (a) Since it is a Poisson-Gamma model then we recall that λ ∼ Gamma(100, 1). Now
determine the credibility factor for each successive year using the formula
n
Z=
β +n

here β = 1 and α = 100 so that


n
Zn =
1+n

33
Years(n) 1 2 3 4 5 6 7 8 9 10
Number of claims (xn ) 144 144 174 148 151 156 168 147 140 16
1 2 3 4 5 6 7 8 9
Credibility factor (Zn ) 0 2 3 4 5 6 7 8 9 10
Mean (X n ) 144 144 154 152.5 152.2 152.8 155 154 152.4 153
Estimate(λ ∼ Gamma(100, 1)) 100 122 129.3 140.5 142 143.5 145.3 148.1 148 147

By graphing the credibility factor in successive years we have;

From the graph we see that as we increase the value of n our credibility factor, this mean sense
because the credibility formula depends on which is the number of years of the historical data (n).

(b) Estimate for Year 8: Z8 = 78 ,µ = α


β = 100
1 = 100 and X̄7 = 155
 
o 7 7
N Estimated claims = Z8 X̄7 + (1 − X̄7 )µ = (155) + 1 − (100) = 148.125
8 8
Estimate for Year 9: Z9 = 89 ,µ = 100 and X̄8 = 154
 
o 8 8
N Estimated claims = Z8 X̄7 + (1 − X̄7 )µ = (154) + 1 − (100) = 148
9 9
Thus the estimate for year 9 is lower than that of year 8, this is because the actual value for year
8 is relatively small and hence it lowers the estimate for year 9.

(c) The graph of the actual and the estimated number of claims.

34
(d) Since λ ∼ Gamma(500, 5) so that
n
Zn =
5+n

Years(n) 1 2 3 4 5 6 7 8 9 10
No claims (xn ) 144 144 174 148 151 156 168 147 140 161
1 2 3 4 5 6 7 8 9
Credibility factor (Zn ) 0 6 7 8 9 10 11 12 13 14
Estimate(λ ∼ Ga(500, 5)) 100 107.3 112.6 120.3 123.3 126.1 128.8 132.1 133.2 133.7

The graphs of credibility factor in successive years of the parameter λ where λ ∼ Gamma(500, 5)
and λ ∼ Gamma(500, 5) are the prior distributions.

From the figure above we see that the credibility factor increase as we increase the number of years
of history data n, but the credibility factor associated with the prior distribution λ ∼ Gamma(100, 1)
is higher than that of the prior distribution λ ∼ Gamma(500, 5). This due to the fact that the pa-
rameter β was increased to 5 which therefore lowers the credibility factor.

(e) We we plot the graphs for the estimates for λ ∼ Gamma(100, 1), λ ∼ Gamma(500, 5) and the
actual number of claims.

35
We can observe that the estimates for the number of claim under the prior distribution λ ∼
Gamma(100, 1) are much closer to the actual number of claims as compared to estimates with
the prior distribution λ ∼ Gamma(500, 5).

Example 3.4.2 An Actuarial scientist wishes to find a Bayesian estimate of the mean of an exponential
distribution with a density function
1
f (x) = e−x/µ .
µ
He is proposing to use a prior distribution of the form

θ α e−θ /µ
Prior(µ) = µ >0
µ α+1 Γ(α)

Given that the mean of this distribution is


θ
.
α −1
(a) Determine the likelihood function for µ, based on a random sample values x1 , x2 , . . . , xn from an
exponential distribution.

(b) Find the form of the posterior distribution for µ and hence show that an expression for the Bayesian
estimate for µ under the square error loss function is
θ + ∑ xi
µ̂ = .
n+α −1

(c) Show that the Bayesian estimate for µ can be written in the form of the credibility estimate and
write down a formula for the credibility factor.

(d) The Actuarial scientist now decide to use a prior of this form with parameter θ = 40 and α = 1.5.
The Actuary sample data have statistics n = 100, ∑ xi = 9826 and ∑ xi2 = 1, 200, 00. Find the
posterior estimate for µ and the value of the credibility factor in this case.

Solution 3.4.2 (a)


n
L(µ) = ∏ f (xi )
i=1

36
where
1 −x/µ
f (x) = e
µ
So that
n
1
L(µ) = ∏ e−xi /µ
i=1 µ
1 −x1 /µ 1 −x2 /µ 1
L(µ) = e × e · · · × e−xn /µ
µ µ µ
 n − µ1 ∑ xi
1 1
−∑ xi
e
L(µ) = eµ =
µ µn
Therefore the likelihood function is given by
1
e− µ ∑ xi
L(µ) =
µn

(b) Recall that


Posterior distribution ∝ Prior distribution × Likelihood function
1
θ e−θ /µ e− µ ∑ xi
Posterior distribution ∝ α+1 ×
µ Γ(α) µn
1 !
θα e−θ /µ e− µ ∑ xi
Posterior distribution ∝ ×
Γ(α) µ α+1 µn
| {z }
constant can be ignored
1
e− µ (θ +∑ xi )
Posterior distribution =
µ α+1+n
We see that this has the same form as the prior distribution, but with different parameters. Thus the
posterior has he same distribution as the prior but with parameter.

θ ∗ = θ + ∑ x1 and α∗ = n + α

that is
θ∗
e− µ
Posterior distribution = α ∗ +1
µ
Hence the mean is given by
θ∗ θ + ∑ Xi
E[µ] = ∗
=
α −1 n+α −1
Thus
θ + ∑ Xi
µ̂ =
n+α −1
(c) From
θ + ∑ Xi θ ∑ Xi
µ̂ = = +
n+α −1 n+α −1 n+α −1
θ ∑ xi
α−1 n
= n+α−1
+ n+α−1
α−1 n
     
α −1 θ n ∑ xi
= +
n+α −1 α −1 n+α −1 n

37
   
α −1 n
= µ+ X
n+α −1 n+α −1
So that our credibility factor
n α −1
Z= and 1−Z =
n+α −1 n+α −1
θ

where µ is a weighted average of the prior distribution mean α−1 and the maximum likelihood
 
∑ xi
estimate of µ or X is n .

(d) Posterior estimate for µ


θ + ∑ Xi 40 + 9826
µ̂ = = = 98.1692.
n + α − 1 100 + 1.5 − 1
The value of the credibility factor is
n 100
Z= = = 0.9950
n + α − 1 100 + 1.5 − 1
Comment: By observation, the value of Z is close to 1. So our credibility estimate comes out to
be very close to the sample mean (98.26) and takes little account of the prior mean (80). This is
because n is much bigger than α.

3.4.2 Normal-Normal Model


This model is basically used to model the severity of the claim or the claim amounts. Suppose X is a
random variable representing the aggregate claim amounts in the coming year for a particular risk that
follows a normal distribution with parameters θ and σ12 i.e.[X ∼ N(θ , σ12 )], whereas θ ∼ N(µ, σ22 ) which
is the prior distribution. The values of µ, σ12 and σ22 are known while θ is unknown.
Again we need to determine the posterior distribution so that
(θ −µ)2
1 −
2σ22
Prior distribution = √ e
σ2 2π
And
n (xi −θ )2
−∑

1 2σ12
Likelihood Function = √ e
σ1 2π
Thus
Posterior distribution ∝ Prior distribution × Likelihood function
(θ −µ)2  n ∑(xi −θ )2
1 − 1 −
2σ22 2
Posterior distribution ∝ √ e × √ e 2σ1
σ2 2π σ1 2π
(θ −µ) 2  n ∑(xi −θ )2
1 − 1 −
2σ22 2
Posterior distribution ∝ √ e × √ e 2σ1
σ2 2π σ1 2π
n (θ −µ)2 (x −θ )2
−∑ i 2
 
1 1 − 2
Posterior distribution ∝ √ √ e 2σ2 × e 2σ1
σ2 2π σ1 2π
| {z }
constant can be ignored
 
(θ −µ)2 ∑(xi −θ )2
− 21 2 + 2
Posterior distribution = e σ2 σ1

38
Now since the exponents in this expression is a quadratic function of θ , and it is proportional to
 
(θ −µ ∗ )2
− 12
σ∗2
e

which is a normal distribution with parameters µ∗ and σ∗2 i.e. [N(µ∗ , σ∗2 )]. Now we determine the
parameters µ∗ and σ∗2
So from  
(θ −µ)2 ∑(xi −θ )2
− 21 +
σ22 σ12
Posterior distribution = e
2 2
 
(θ 2 −2θ µ+µ 2 ) ∑(xi −2xi θ +θ )
− 12 +
σ22 σ12
=e
2
 
2
θ 2 2θ µ µ ) ∑ xi 2θ x 2
− 12 2 − 2 + 2 + 2 − ∑2 i + nθ2
=e σ2 σ2 σ2 σ1 σ1 σ1

2 ∑ x2
     
x
− 12 θ 2 1
+ n −2θ µ
+ ∑ 2i + µ 2 + 2i
σ22 σ12 σ22
=e σ1 σ2 σ1

By comparison with    
(θ −µ∗ )2 2
θ 2 2θ µ∗ µ∗
− 21 − 21 − 2 + 2
σ∗2 σ∗2
e =e σ∗ σ∗

We have that
θ2
 
2 1 n 1 1 n
2
=θ 2
+ 2 =⇒ 2 = 2 + 2 (26)
σ∗ σ2 σ1 σ∗ σ2 σ1
 
2θ µ∗ µ ∑ xi µ∗ µ ∑ xi
− 2 = −2θ 2
+ 2 =⇒ 2 = 2 + 2 (27)
σ∗ σ2 σ1 σ∗ σ2 σ1
Solving equation (26) and (27) for the values of µ∗ and σ∗2
From (26) we have;
1 σ12 + nσ22
=
σ∗2 σ22 σ12
σ22 σ12
σ∗2 = 2
σ1 + nσ22
Using (27) we have
µ∗ µσ12 + σ22 ∑ xi
=
σ∗2 σ22 σ12
µσ12 + σ22 ∑ xi 2
=⇒ µ∗ = σ∗
σ22 σ12
µσ12 + σ22 ∑ xi σ22 σ12 µσ 2 + σ22 ∑ xi
µ∗ = × 2 =
σ22 σ12 σ1 + nσ22 σ12 + nσ22
µσ12 + σ22 ∑ xi
µ∗ =
σ12 + nσ22
Since n1 ∑ xi = X =⇒ ∑ xi = nX then
µσ12 + nσ22 X
µ∗ =
σ12 + nσ22

39
Thus the posterior distribution of θ given X = (x1 , x2 , . . . xn ) is a normal distribution with parameter

µσ12 + nσ22 X σ22 σ12


and
σ12 + nσ22 σ12 + nσ22

Now the mean of the posterior distribution is given by

µσ12 + nσ22 X
E[θ |X] =
σ12 + nσ22

Thus
σ12 nσ22
E[θ |X] = 2 µ+ 2 X
σ1 + nσ22 σ1 + nσ22
By comparison with (24) we have that

nσ22 σ12
Z= and 1−Z =
σ12 + nσ22 σ12 + nσ22

In this model the credibility factor is given by

nσ22 n
Z= = (28)
σ12 + nσ22 σ12
+n
σ22

Here n is the number of years of historical data available and σ22 is the variability in the collateral data. If
the variability in the collateral data σ22 is less then the Z value will be lower giving more confidence to the
collateral data. On the other hand increasing n, increases the value of Z therefore giving more confidence
to the data coming from the risk itself.

Example 3.4.3 A certain portfolio has is claim amount following a N(θ , 4) distribution. Assuming that
the distribution of θ is N(1, 9). Derive the Bayesian estimate of θ . Hence show that the credibility factor
Z is given by
9n
Z=
4 + 9n
Solution 3.4.3 The prior distribution is given by;
(θ −µ)2
1 −
2σ22
f (θ ) = √ e ∼ N(1, 9)
σ2 2π

Since µ = 1 and σ22 = 9 then


1 (θ −1)2
f (θ ) = √ e− 18
3 2π
Then the likelihood function is;

n n (xi θ )2
1 −
2σ12
L(θ ) = ∏ f (xi ) = ∏ √ e ∼ N(θ , 4)
i=1 i=1 σ1 2π

Thus  n
1 ∑(xi −θ )2
L(θ ) = √ e− 8
2 2π
Posterior distribution ∝ Prior(θ ) · L(θ )

40
   n  
1 (θ −1)2 1 (x −θ )2
− 18 − ∑ i8
∝ √ e · √ e
3 2π 2 2π
  n
1 1 (θ −1)2 (x −θ )2
− 18 − ∑ i8
∝ √ √ e ·e
3 2π 2 2π
| {z }
constant can be ignored
 
(θ −1)2 ∑(xi −θ )2
− 18 + 8
=e  
4(θ −1)2 +9 ∑(xi −θ )2
− 72
= e
2

(4+9n)θ 2 2(4+9 ∑ xi )θ 4+9 ∑ xi
− 12 36 − 36 + 36
=e
Since the posterior distribution is a normal distribution with parameters µ∗ and σ∗2 and is proportional to
 
2
(θ −µ∗ )2 θ 2 2µ∗ θ µ∗
− 12 − 12 2 − 2 + 2
e σ∗2 =e σ∗ σ∗ σ∗

By comparison we have;
θ2 (4 + 9n)θ 2
=
σ∗2 36
1 (4 + 9n)
=⇒ 2 =
σ∗ 36
36
σ∗2 =
4 + 9n
2µ∗ θ 2(4 + 9 ∑ xi )θ
− 2 =−
σ∗ 36
µ∗ 4 + 9 ∑ xi
=⇒ 2 =
σ∗ 36
4 + 9 ∑ xi 2
µ∗ = σ∗
36
4 + 9 ∑ xi 36 4 + 9 ∑ xi
µ∗ = · =
36 4 + 9n 4 + 9n
Thus the posterior distribution is normal with parameters
36 4 + 9 ∑ xi 36
µ∗ = = and σ∗2 =
4 + 9n 4 + 9n 4 + 9n
Hence the estimate for θ is
36 4 + 9 ∑ xi
E[θ |X] = =
4 + 9n 4 + 9n
Thus
4 9 ∑ xi
4 + 9 ∑ xi 4 9 ∑ xi 4 n
E[θ |X] = = + = 4+9n + 4+9n
4 + 9n 4 + 9n 4 + 9n 4 n
4 n
·1+ ·X
4 + 9n 4 + 9n
Since
ZX + (1 − Z)µ
here µ = 1 thus
ZX + (1 − Z) · 1
Therefore
9
Z= hence shown
4 + 9n

41
3.5 Empirical Bayes Credibility Theory: Model 1
This section describes two models of empirical Bayesian reliability theory (EBCT). Similar to the Poisson-
Gamma and Normal-Normal models described earlier. This section uses these models to estimate the
"true" claim frequency or risk premium. It is based on total claim accounts in successive periods. Model
1 gives equal weights to all risks each year. Model 2 is more demanding and takes volume into account
of the business written under each risk of each year.
The key similarities and differences between Bayesian and EBCT models are described below.

Risk parameter
Both methods make use of an auxiliary risk parameter θ . The quantity we want to estimate with the
Bayesian approach is θ , whereas with the EBCT models, is m(θ ), which is a function of θ . Unlike the
Bayesian approach, the EBCT models do not make any assumptions, a particularly statistical distribution
for θ .

Conditional claim distribution


Both approaches are based on the assumption that the conditional variables X j |θ ’s are independent and
identically distributed. Unlike the pure Bayesian approach, the EBCT models make no assumptions
about the statistical distribution of X j |θ . Formulae are instead developed solely on the assumption that
the mean m(θ ) and variance s2 (θ ) of X j |θ can be expressed as functions of θ . The models then estimate
the values of m(θ ) and s2 (θ ).

Credibility formula
The resulting formula for estimating the claim frequency or risk premium for a given risk can be expressed
using a credibility formula in either approach. The credibility formula is a linear combination of past
claim averages and an overall average.

Model 1: Specification
The assumptions for Model 1 of EBCT will be laid out in this section. The EBCT Model 1 is a generali-
sation of the normal/normal model. This point will be discussed in greater depth later in this section. The
problem at hand is estimating the pure premium or, possibly, the claim frequency for a risk. Let X1 , X2 . . .
represent the total number of claims for this risk in successive periods. A more specific statement of the
problem is that after observing the values of X1 , X2 . . . Xn , the expected value of Xn+1 must be estimated.
X will now be used to represent X1 , X2 . . . , Xn . The following assumptions are made:

(i) The distribution of each X j depends on a parameter, denoted θ , whose value is fixed (and the same
for all the X j ’s) but is unknown.

(ii) Given θ , the X j ’s are independent and identically distributed.

Following that, some notation is introduced. m(θ ) and s2 (θ ) are defined as follows:

m(θ ) = E[X j |θ ] and s2 (θ ) = Var[X j |θ ]

There are two things to note about m(θ ) and s2 (θ ). The first is that, because the X j ’s are identically
distributed given θ , neither m(θ ) nor s2 (θ ) depend on j, as their notation implies. The second point is
that, because θ is a random variable, then both m(θ ) and s2 (θ ) are random variables.
The following are the similarities between EBCT Model 1 and the normal/normal model:

42
(i) For both models, the role of θ is the same: it characterises the underlying distributions of the
processes being modelled, such as the aggregate claim distribution for each year of business.

(ii) The assumptions regarding the unconditional distribution of X j ’s are the same: they are identically
distributed in each case.

(iii) The assumptions regarding the conditional distribution of X j ’s given θ are the same: they are
conditionally independent in each case.
The EBCT model 1 can be considered as a generalisation of the normal/normal model. Specific points
where it differs, i.e. general, normal/normal model is:
(i) E[Xi |θ ] is some function of θ , m(θ ), for EBCT Model 1 but is simply θ for the normal/normal
model. Hence:

Var[m(θ )] {EBCT } corresponds to σ22 (= Var(θ )){normal/normal}.

(ii) Var[X j |θ ] is a function of θ , s2 (θ ), for EBCT Model 1 but is constant, σ12 , for the normal/normal
model. Hence:
E[s(θ )] corresponds to σ21 (= Var|([X j |θ ])
.

(iii) The normal/normal model makes very precise distributional assumptions about both X j given θ ,
which is N(θ , σ22 ), and θ , which is N(µ, σ22 ). EBCT Model 1 makes no such distributional as-
sumptions.

(iv) The risk parameter, θ , is a real number for the normal/normal model but could be a more general
quantity for EBCT Model 1.

Model 1: Credibility Premium


The derivation of the credibility premium under this model is beyond the scope of this course syllabus
and is not covered here. Thus the result of EBCT model are given below;
The estimate of m(θ ) given X given by EBCT Model 1 is:

ZX + (1 − Z)E[m(θ )]

where
n
1
X= ∑ Xj and
n j=1
n
Z= E[s (θ )]2 (29)
n + Var[m(θ )]
The solution’s first and most important feature is that it takes the form of a credibility estimate. In other
words, it is in the form of (24) with
n
1
E[m(θ )] playing a role of µ and ∑ X j playing a role of X
n j=1

The second feature to notice is the similarity between the solution above and the solution in the nor-
mal/normal model, specifically the formulae for the credibility factor. Formula (29) is a straight general-
isation of formula (28), since E[s2 (θ )] and Var[m(θ )] are considered to be the generalisation of the σ12
and σ22 respectively.

43
Model 1: Parameter Estimation
The solution to the problem of estimating m(θ ) given X will be completed in this section by showing
how to estimate E[m(θ )], Var[m(θ )], and E[s2 (θ )].
To do so, another assumption must be made: that data from other risks similar, but not identical, to the
original risk is available. This necessitates resolving the problem more broadly, making some additional
assumptions, and changing the notation slightly. A key distinction between a pure Bayes approach to
credibility and EBCT is that the former requires no data to estimate parameters, whereas the latter requires
data to estimate parameters.
Consider the table below

Let Xi j denote the aggregate claims, or number of claims, for Risk Number i, i = 1, 2, . . . , N, in year j,
j = 1, 2, 3 . . . , n.
Each row of Table 1 represents observations for a different risk; the first row, relating to Risk Number 1,
is the set of observed values denoted by X11 , X12 . . . , X1n = X
Each row in Table 1 represents a fixed value of θ . Keeping this in mind, as well as the definitions of
m(θi ) and s2 (θi ), the obvious estimators for m(θ ) and s2 (θi ) are:
n n
1 1
Xi =
n ∑ Xi j and n−1 ∑ (Xi j − X i)2 respectively
j=1 j=1

E[m(θ )] is now the "average" (across the θ distribution) of the values of m(θ ) for different values of
θ . The obvious estimator for E[m(θ )] is the average of the m(θ ) estimates for i = 1, 2, . . . , N. In other
words, X is the estimator for E[m(θ )]. i.e.

1 N
X= ∑ Xi (30)
N i=1

Similarly, E[s2 (θ ) is the "average" value of s2 (θ ), and thus an obvious estimator is the average of the
estimates of s2 (θ ), which is ( )
1 N 1 n 2
∑ n − 1 ∑ (Xi j − Xi)
N i=1
(31)
j=1

For each row of Table 1, X i is an estimate of m(θi ), i = 1, 2, . . . , N. As a result, it could be argued that the
observed variance of these values, i.e:

1 N
∑ (Xi − X)2 (32)
N − 1 i=1

would be an obvious estimator for Var[m(θ )]. Unfortunately, this can be shown to be a biased estimate of
Var[m(θ )]. So, it can be shown that an unbiased estimate of Var[m(θ )] can be formulated by subtracting

44
a correction term from the formula (32). Thus an unbiased estimator for Var[m(θ )] is:
( )
1 N 1 N
1 n
∑ (Xi − X)2 − Nn ∑ n − 1 ∑ (Xi j − Xi)2
N − 1 i=1
(33)
i=1 j=1

1
NB: The correction term is n times the estimator for E[s2 (θ ).
An important point to note is that although Var[m(θ )] is a non-negative parameter since it is a variance,
the estimator given by (33) could be negative. Formula (33) is the difference between two terms. Each of
these terms is non-negative but their difference need not [Link] practice, if (33) gives a negative value, the
accepted procedure is to estimate Var[m(θ )] as 0.
The parameter E[s2 (θ ) must also be non-negative, but its estimator, given by formula (31), will always
be non-negative and so no adjustment to (31) is required.
It can be shown that the estimators for E[m(θ )], E[s2 (θ ), and Var[m(θ )] are unbiased. The proof is
beyond the scope of the course syllabus.
Now considering equation (29) for the credibility factor for the EBCT 1 model. The various "compo-
nents" of this formula have the following definitions or interpretations:
n is the number years of data values in respect to the risk.
E[m(θ )] is the average variability of data values from year to year for a single risk, i.e the average vari-
ability within the rows of Table 1
var[m(θ )] is the variability of the average data values for different risks, i.e the variability of the row
means in Table 1.
Looking at formula (29) for Z , the following observations can be made:

• Z is always between zero and one (0 ≤ Z ≤ 1).

• Z is an increasing function of n . This is to be expected – the more data there are from the risk
itself, the more it will be relied on when the credibility estimate of the pure premium or number of
claims is calculated.

• Z is a decreasing function of E[s2 (θ )]. This is to be expected – the higher the value of E[s2 (θ )]
relative to var[m(θ )], the more variable, and hence less reliable, are the data from the risk itself
relative to the data from the other risks in the collective.

• Z is an increasing function of var[m(θ )]. This is to be expected – the higher the value of var[m(θ )],
relative to E[s2 (θ )], the more variability there is between the different risks in the collective and
hence the less likely it is that the other risks in the collective will resemble the risk that is of interest,
and the less reliance should be placed on the data from these other risks.

The last point seem to contradicts with point observed in the previous section of this chapter. Previously,it
was stated that the credibility factor should not depend on the data from the risk being rated. However,
here these data have been used to estimate var[m(θ )] and E[s2 (θ )], whose values are then used to cal-
culate Z. The explanation is that in principle the credibility factor, Z , as given by formula (24) does
not depend on the actual data from the risk being rated. Unfortunately, the formula for Z involves two
parameters, var[m(θ )] and E[s2 (θ )], whose values are unknown but which can, in practice, be estimated
from data from the risk itself and from the other risks in the collective.

Example 3.5.1 The table below shows the aggregate claim amounts (in ZMK million) for an interna-
tional insurer’s fire portfolio for a 5-year period, together with some summary statistics.

45
(i) Complete the table the and calculate E[m(θ )], E[s2 (θ )] and var[m(θ )]
(ii) Find the credibility factor and hence calculate the EBCT premium for each of the countries for the
coming year
Solution 3.5.1 (i) In this example, there are n = 5 years and N = 4 risks (= countries). Xi is the
average claims for risk i for the 5-year period. So the missing entry is:
1 n n
 
1
X4 = ∑ Xi j = 44 + 52 + 69 + 55 + 71 = ∑ = 58.2
n j=1 5 j=1

X4 = 58.2
With the calculation above, we can therefore find the other missing entry:
1 n
 
2 1 2 2 2 2 2
∑ ((Xi j −Xi) = 4 (44−58.2) +(52−58.2) +(69−58.2) +(55−58.2) +(71−58.2) = 132.7
n − 1 j=1

To estimate E[m(θ )], E[s2 (θ )] and var[m(θ )] we need to first find X that is
1 N
 
1
X = ∑ Xi = 50.4 + 68.4 + 74 + 58.2 = 62.75
N i=1 4
Now we can estimate the parameter thus;
E[m(θ )] ≈ X = 62.75
( )
N n
1 1
E[s2 (θ )] ≈ ∑ ∑ (Xi j − Xi)2
N i=1 n − 1 j=1
1
= (39.317.3215.5 + 132.7) = 101.2
4 ( )
1 N 1 N 1 n
var[m(θ )] ≈ ∑ (Xi − X)2 − ∑ 2
∑ (Xi j − Xi)
N − 1 i=1 Nn i=1 n−1 j=1
Thus
1 N
 
2 1 2 2 2 2
∑ (Xi −X) = 3 (50.4−62.75) +(68.4−62.75) +(74−62.75) +(58.2−62.75) = 110.57
N − 1 i=1
and, ( )
1 N n  
1 2 1 2 1

Nn i=1 n−1 ∑ (Xi j − Xi) = E[s (θ )] =
n 5
101.2 = 20.24
j=1
Therefore
Var[m(θ )] ≈ 110.57 − 20.24 = 90.33

46
(ii)
n 5
Z= = ≈ 0.8169
E[s2 (θ )]
n + Var[m(θ )] 5 + 101.2
90.33

Now since we have the same number of time periods for each country, the credibility factor is the
same for each. Hence use the basic credibility formula

P = ZXi + (1 − Z)E[m(θ )]

Country 1: P = (08169)50.4 + (1 − 08169)62.75 = 52.66


Country 2: P = (08169)68.4 + (1 − 08169)62.75 = 67.37
Country 3: P = (08169)74.0 + (1 − 08169)62.75 = 71.94
Country 4: P = (08169)58.2 + (1 − 08169)62.75 = 59.03

Example 3.5.2 The following data represents the total claim amounts per year, Xi j , over a six year period
(6) n = for five fleets of buses (5) N =, A to E:

(i) Determine the credibility factor.

(ii) Calculate the EBCT premium for each of the fleet of buses for the coming year.

Solution 3.5.2 (i) We first need to determine the following functions from the data

Xi = 1n ∑6j=1 Xi j ∑6j=1 (Xi j − Xi )2 (Xi − X)2


A 1,375 973,150 5,569,600
B 3,090 11,218,200 416,025
C 2,810 4,562,000 855,625
D 5,325 18,039,550 2,528,100
E 6,075 19,130,550 5,475,600
Totals 18,675 53,923,450 14,844,950

From the table above, we can determine the estimates of the following parameters E[m(θ )], E[s2 (θ )]
and Var[m(θ )] thus;

1 N
 
1
E[m(θ )] ≈ X = ∑ Xi = 18675 = 3735
N i=1 5
This is the overall mean, which we obtain by finding the average of the means for each fleet m(θi ).
( )
N n   
2 1 1 2 1 1
E[s (θ )] ≈ ∑ ∑ (Xi j − Xi) = 5 5 53923450 = 2156938
N i=1 n − 1 j=1

47
This is the average of the fleet variances, obtain determining the average variance of each fleet
s2 (θi ) ( )
1 N 1 N
1 n
Var[m(θ )] ≈ ∑ (Xi − X)2 − ∑ ∑ (Xi j − Xi )2
N − 1 i=1 Nn i=1 n − 1 j=1
  
1 1 1
= (14844950) − 53923450
4 (5)(6) 5
Var[m(θ )] ≈ 3351748
This gives the variance between the fleet mean claims
Therefore, we determine the credibility Z
n 6
Z= E[s2 (θ )]
= 2156938
≈ 0.90313
n + Var[m(θ 6+ 3351748
)]

(ii) Now for each fleets (or risks) in the table above, we can determine the credibility premium using
the formula Credibility premium = ZXi + (1 − Z)E[m(θ )]
Fleet A: Credibility premium = 0.90313(1, 375) + (1 − 0.90313)3735 ≈ 1605
Fleet B: Credibility premium = 0.90313(3, 090) + (1 − 0.90313)3735 = 3152
Fleet c: Credibility premium = 0.90313(2, 810) + (1 − 0.90313)3735 = 2900
Fleet D: Credibility premium = 0.90313(5, 325) + (1 − 0.90313)3735 = 5171
Fleet E: Credibility premium = 0.90313(6, 075) + (1 − 0.90313)3735 = 5848

3.6 Empirical Bayes Credibility Theory: Model 2


Model 2 is a generalisation of model 1, it is used to estimate the pure premium, or the expected number
of claims, in the coming year for a risk.

Model 2: Specification
The problem is estimating the total expected claims or the expected number of claims over the next year
for a given risk. Let Y1 ,Y2 , . . . ,Yn random variables represent the total claims or the number of claims in
successive years for this risk. It is assumed that the values of Y1 ,Y2 , . . . ,Yn have already been observed
and the expected value of Yn+1 needs to be estimated. The difference between EBCT model 1 and EBCT
model 2 is that model 2 involves an extra parameter known as the risk volume, Pj =. Ideally the value of
Pj measures the “amount of business” in year j.
Next a new sequence of random variables, X1 , X2 , . . . is defined as follows:
Yj
Xj = for j = 1, 2, 3, . . .
Pj

The random variable X j represents the aggregate claims, or the number of claims, in year j standardised
to remove the effect of different levels of business in different years. The assumptions that specify EBCT
Model 2 are as follows;

(i) The distribution of each X j depends on the value of a parameter, θ , whose value is the same for
each j but unknown.

48
(ii) Given θ , the X j ’s are independent (but not necessarily identically distributed).

(iii) E[X j |θ ] does not depend on j

(iv) PjVar[X j |θ ] does not depend on j


From assumptions (iii) and (iv),m(θ ) and s2 (θ ) can be defined as follows:

m(θ ) = E[X j |θ ] (34)

s2 (θ ) = PjVar[X j |θ ] (35)

Model 2: Credibility Premium


As for Model 2 the derivation of the credibility premium under this model is beyond the scope of the
course syllabus and will not be covered here. Thus the results of EBCT model 2 are as follows;

ZX + (1 − Z)E[m(θ )]
( )−1 ( )−1
n n n n
X= ∑ Pj X j ∑ Pj = ∑ Y j ∑ Pj
j=1 j=1 j=1 j=1

∑nj=1 Pj
Z= E[s2 (θ )]
(36)
∑nj=1 Pj + Var[m(θ )]
Here are some additional points of note about this solution:
(i) If all the Pi ’s were equal to 1, the solution given by (36) is exactly the same as the solution given
by (29). This is as it should be since if all the Pi ’s are equal to 1 then EBCT Model 2 is exactly the
same as EBCT.

(ii) As for EBCT Model 1, the solution given by (36) involves three parameters, E[m(θ )], Var[m(θ )]
and E[s2 (θ )] Model 1. The estimate for these parameter are given in the next sub-section.

Model 2: Parameter estimation


Let Yi j be a random variable denoting the aggregate claims, or the number of claims, for risk number i in
year j , j = 1, 2, . . . , n, i = 1, 2, . . . , N, and let Pi j be the corresponding risk volume. Then for each i and j
we define
Yi j
Xi j =
Pi j
Consider the data are summarised in the following table

49
For simplicity it is assumed, that the risk that is of particular interest is Risk Number 1 in this collective.
This means that previously what were denoted Y j , Pj and X j are now denoted by Y1 j , P1 j and X1 j respec-
tively in this section. The problem is to estimate the expected value of X1,1+n and the solution to this
problem has already been given by (36).
The purpose of the data from the other risks in the collective is purely to help to estimate the parameters
E[m(θ )], Var[m(θ )] and E[s2 (θ )] that appear in (36). From the given data int the table we can obtain the
following values
n n N  
∗ 1 Pi
Pi = ∑ Pi j , P = ∑ Pi , P = ∑ Pi 1 − P
Nn − 1 i=1
j=1 i=1
n Pi j Xi j N n P X
ij ij
Xi = ∑ and X = ∑ ∑
j=1 Pi i=1 j=1 P

With this new notation the credibility estimate of the pure premium, or number of claims, per unit of risk
volume for the coming year for Risk Number i in the collective originally given by formula (36) can be
reformulated as
ZXI + (1 − Z)E[m(θ )]
N n
Pi j Xi j
X=∑ ∑
i=1 j=1 P

∑nj=1 Pi j
Zi = E[s2 (θ )]
(37)
∑nj=1 Pi j + Var[m(θ )]
where the quantity E[m(θ )], E[s2 (θ )] and Var[m(θ )] are estimated by the following unbiased estimators
( )
1 n 1 n 2
X, ∑ n − 1 ∑ Pi j (Xi j − Xi)
N i=1 j=1

and ( )!
N n
1 1 2 1 N 1 n

P∗
P
∑ ∑ ij ij
Nn − 1 i=1
(X − X) − ∑
N i=1 n−1 ∑ Pi j (Xi j − Xi)2
j=1 j=1

respectively

Example 3.6.1 The table below shows the same claim amounts Yi j and corresponding number of buses,
Pi j over a six-year period ( n = 6) for the five fleets of buses ( N = 5), A to E:

(i) Determine the credibility factor for each risk (fleet of buses).

(ii) Calculate the EBCT premium for each of the fleet of buses for the coming year.

50
Solution 3.6.1 (i) Recall that
Yi j
Xi j =
Pi j
Thus we have the following table

2005 2006 2007 2008 2009 2010


A 250 196 450 340 200 236
B 154.55 236.92 170 235 384 248.57
C 683.33 890 700 533.33 1400 1325
D 521.11 485.56 600 1133.75 418.89 525
E 1021.43 497.14 626.25 601.25 971.11 726

In order to determine the credibility factors we need to computer the following quantities;

P X
∑6j=1 Pi j Xi j Pi = ∑6j=1 Pi j Xi = ∑nj=1 i Pj i j
i
∑6j=1 Pi j (Xi j − Xi )2 ∑6j=1 Pi j (Xi j − X)2
A 8,250 30 275 217,910 1,680,436
B 18,540 75 247.2 437,922 5,072,925
C 16,860 19 887.4 1,812,796 4,726,031
D 31,950 53 602.8 2,804,027 3,411,220
E 36,450 49 743.9 1,706,742 4,722,410
Totals 112,050 226 2756.3 6,979,397 19,613,022

And
 
Pi
Pi 1 −
∑5i=1 Pi
A 26.018
B 50.111
C 17.403
D 40.571
E 38.376
Totals 172.478

From the above calculation we have


N n
Pi j Xi j 112, 050
E[m(θ )] ≈ X = ∑ ∑ = = 495.796
i=1 j=1 P 226
!
N  
∗ 1 Pi 1
P = ∑ Pi 1 − ∑5 P = 29 172.478 = 5.9475
Nn − 1 i=1 i=1 i
( )
n n
1 1 6, 979, 397
E[s2 (θ )] ≈ ∑ ∑ Pi j (Xi j − Xi )2 = = 279,175.88
N i=1 n − 1 j=1 25
and
( )!
N n
1 1 2 1 N 1 n
Var[m(θ )] ≈ ∗
P ∑∑
Nn − 1 i=1
Pi j (Xi j − X) − ∑
N i=1 n−1 ∑ Pi j (Xi j − Xi)2
j=1 j=1
   
1 1
= 19, 613, 022 − 279, 175.88 = 66,774
5.9475 29

51
Therefore using (37) to determine the credibility factor for each risk thus;
30
Z1 = 279175.88
≈ 0.87768
30 + 66774

75
Z2 = 279175.88
≈ 0.94720
75 + 66774
19
Z3 = 279175.88
≈ 0.81964
19 + 66774
53
Z4 = 279175.88
≈ 0.92688
53 + 66774
49
Z5 = ≈ 0.92138
49 + 279175.88
66774
Thus

Zi
A 0.87768
B 0.94720
C 0.81964
D 0.92688
E 0.92138

(ii) Now for each fleets (or risks) in the table above, we can determine the credibility premium per unit
of risk volume using the formula

Credibility premium = Zi Xi + (1 − Z)E[m(θ )]

Fleet A: Credibility premium per unit risk volume

[Link] = 0.87768(275) + (1 − 0.87768)495.795 ≈ 302.0


Fleet B: Credibility premium per unit risk volume

[Link] = 0.94720(247.2) + (1 − 0.94720)495.795 ≈ 260.3


Fleet C: Credibility premium per unit risk volume

[Link] = 0.81964(887.4) + (1 − 0.81964)495.795 ≈ 816.7


Fleet D: Credibility premium per unit risk volume

[Link] = 0.92688(602.8) + (1 − 0.92688)495.795 ≈ 595.0


Fleet E: Credibility premium per unit risk volume

[Link] = 0.92138(743.9) + (1 − 0.92138)495.795 ≈ 724.4


Note: By multiplying these figures by the risk volumes for 2011 would then give the credibility
premiums for 2011.

52
3.7 Tutorial sheet 2
1. The number of claims in one year has a Binomial distribution with n = 3 and θ unknown. Given
that the prior distribution of θ is a Beta distribution with pdf

π(θ ) = 280θ 3 (1 − θ )4 , 0<θ <1

Two claims were observed. Determine the

(i) posterior distribution for θ .


(ii) expected value of θ ) from the posterior distribution.
(iii) use the expected value of the parameter θ and obtain the credibility factor z. Hence comment
on your answer.

2. The number of claims is distributed according to a Gamma-Poisson mixed distribution. The prior
distribution is assumed to be a Gamma distribution with parameter α = 4 and β = 2. Over a
three-year period, 5 claims were observed. Determine the parameter α ∗ and β ∗ of the posterior
distribution.

3. Suppose the likelihood of s claim is given by s Poisson distribution with parameter θ . The prior
density function of θ is given by
f (θ ) = 32θ 2 e−4θ .
If 3 claims are observed in 2 years, what is the posterior density function of θ ?

4. Claims X each year from a portfolio of insurance policies are normally distributed with mean θ
and variance τ. Prior information is that θ is normally distributed with known mean µ and known
variance σ 2
Aggregate claims over the last n years have been xi for i = 1 to n, and you should assume that these
are independent.

(i) Derive the posterior distribution of θ .


(ii) Write down the Bayesian estimate of θ under quadratic loss.
(iii) Show that the estimate in your answer to part (ii) can be expressed in the form of a credibility
estimate, including statement of the credibility factor Z

5. For three years an insurance company has insured buildings in three different towns against the risk
of fire damage. Aggregate claims in the jth year from the ith town are denoted by Xi j for i = 1, 2, 3
and j = 1, 2, 3. The data is given in the table below.

Calculate the expected claims from each town for the next year using the assumptions of Empirical
Bayes Credibility Theory model 1.

53
6. The table below shows annual aggregate claim statistics for 3 risks over 4 years. Annual aggregate
claims for risk i , in year j , are denoted by Xi j .

Determine the value of the credibility factor for Empirical Bayes Model 1. Hence Using the num-
bers calculated obtained, describe the way in which the data affect the value of the credibility
factor.

7. A shipping insurance company has insured ships for six years, and classifies the ships it insures
into three types.
Let:
Pi j be the number of ships insured in the jth year from type i,
Yi j be the corresponding number of claims.
The six years of data are summarised as follows:

Given that there are 120 ships of Type 3 to be insured in year seven. Estimate the number of claims
from Type 3 ships in year seven using empirical Bayes credibility theory (EBCT) Model 2

8. The total claim amount per year on a particular insurance policy follows a normal distribution with
unknown mean θ and variance 2 . Prior beliefs about θ are described by a normal distribution with
mean 30 and variance 12. Claim amounts x1 , x2 , . . . , xn are observed over n years.

(i) Derive the posterior distribution of θ and clearly state its parameters .
(ii) Show that the mean of the posterior distribution of θ can be written in the form of a credibility
estimate and write down the credibility factor expression.
(iii) Now suppose that total claims over the five years were 170. Calculate the posterior probability
that θ is greater than 30.

54
4 Ruin Theory
Objectives:
• Explain the concept of ruin for a risk model
• Calculate the adjustment coefficient and state Lundberg’s inequality
• Describe the effect on the probability of ruin of changing parameter values and of simple reinsur-
ance arrangements

4.1 Introduction
Ruin theory is practically motivated by the issue of solvency. Although solvency is a complex topic but
in simpler terms an insurance company is said to be solvent if it has sufficient asset to meet its liabilities.
Thus ruin theory is very much concerned with the level of an insurer’s surplus for a portfolio of insurance
policies.
Suppose we consider the evolution of an insurance fund over time, taking account of the times at which
the claim occur, as well as their [Link] make this mathematically track-able, we simplify a rel life
insurance operation by assuming that the insurer’s starts with a non-negative amount of money, collects
premiums and pays claims as they occur. The model of an insurance surplus process is thus deemed to
have three components these are;
(i) Initial surplus (surplus at time zero(t=0))
(ii) Premiums received
(iii) Claims paid
For this model, if the insurer’s surplus falls to zero or below then we say that ruin has occurred.
Recall from risk theory, we used the collective risk model to study the aggregate claim S arising during a
fixed period of time. We saw that S was given by the equation
S = X1 + X2 + X3 + · · · + XN (38)
where N is the number of claims arising in a specified period of time. In the section, we will extend this
model by treating S(t) as a function of time. This gives us the equation
S(t) = X1 + X2 + X3 + · · · + XN(t) (39)
where N(t) represent the number of claims occurring before time t. N(t) is called a Poisson process
and S(t) is called a compound Poisson process. S(t) is used to model claims received by an insurance
company and hence consider the probability that this insurance company is ruined.
Here, we will consider the claim generated by a portfolio over successive time period. Some notation
needed are ;
N(t) : the number of claims generated by the portfolio in the interval [0,t] ∀t ≥ 0
Xi : the amount of the ith claim. i = 1, 2, 3, . . .
S(t) : the aggregate claim in the interval [0,t] ∀t ≥ 0
{Xi }∞i=1 : is a sequence of random variables
{N(t)}t≥0 and {S(t)}t≥0 are both stochastic processes.
Now it can be seen that;
N(t)
S(t) = ∑ Xi (40)
i=1

55
with the understanding that S(t) is zero if N(t) is zero.
The stochastic process defined in (40) is known as the aggregate claim process for the risk. For instance
the random variable N(1) and S(1) represents the number of claims and the aggregate claim respectively
from a portfolio at time t = 1. The insurer of this portfolio will receive premium from the policyholders.
At this stage it is convenient to assume that the premium income is received continuously at a constant
rate c( rate of premium income per unit time). So that the total premium income received in the time
interval [0,t] is ct, where c is assumed to be strictly positive.

4.2 Surplus Process


The surplus process is mainly concerned with the amount premium collected and the claim payment that
is
Surplus = Premium collected (income) − Claim payment(expenses)
Suppose that at time time t = 0, the insurer has an amount of money set aside for this portfolio. This
amount of money is called the initial surplus and is denoted by U, and also it is assumed that U ≥ 0. The
initial surplus is needed by the insurer because the future premium income on its own may not be enough
to cover the future claims (here the expense are ignored). The insurer’s surplus at any future time (t ≥ 0)
is a random variable since its value depends on the claims experience up to time t.
Consider the graph below which shows the surplus process

Figure 1
The graph above shows one possible outcome of the surplus process. Claims occur at times T1 , T2 , T3 , T4
and T5 and at these times the surplus immediately falls by the amount of the claim. Between the claims,
the surplus increases at a constant rate c per unit time. The model for insurer’s surplus incorporates
simplification,like a model of a complex real-life operation. Here, it is assumed that claims are settled
as soon as they are received and also no interest is earned on the insurer’s surplus. Then the insurer’s
surplus at time t is denoted by U(t) and is given by

U(t) = U + ct − S(t), t ≥0 (41)

In words the formula means that the insurer’s surplus (U(t)) at time t is the initial surplus(U) plus the
premium income up to time t (ct) minus the aggregate claim up to time t (S(t)). Note that the premium
income are not not random variable since they are determined before the risk process starts. The formula
(41) given is valid for t ≥ 0 with the understanding that U(0) = U. At any time t, U(t) is a random

56
variable because S(t) is a random variable. Thus {U(t)}t≥0 is a stochastic process which is known as the
surplus process or cash flow process

4.3 Probability of Ruin in Continuous time


From the graph above we saw that the insurer’s surplus fall below zero as a result of the claim at time T3 .
Basically, when the surplus fall below zero the insurer has run out of money and it is said that ruin has
occurred.
Ideally every insurance company’s main objective is to keep the probability of this event (i.e. the probabil-
ity of ruin) as small as possible, or atleast below a predetermined bound. A way to look at the probability
of ruin is to think of it as the probability that, at some future time, the insurance company will need to
provide more capital to finance this particular portfolio.
With this said, we now define the two probability of ruin as follows;

Ψ(U) = P[U(t) < 0], for some o < t < ∞ (42)

Ψ(U) is the probability of ultimate ruin given initial surplus U. This is sometimes called the probability
of ruin in infinite time.
Ψ(U,t) = P[U(t) < 0], for some o < τ < t (43)
Ψ(U,t) is the probability of ruin within time t given initial surplus U, and sometimes referred to as the
probability of ruin in finite time.
Here are some important logical relationship between these probabilities for 0 < t1 ≤ t2 < ∞ and for
0 ≤ U1 ≤ U2 then
Ψ(U2 ,t) ≤ Ψ(U1 ,t) (44)
Ψ(U2 ) ≤ Ψ(U1 ) (45)
The bigger the initial surplus, the less likely it is that ruin will occur either in a finite time period (44) or
infinite time period (45).
Ψ(U1 ,t1 ) ≤ Ψ(U1 ,t2 ) (46)
For a given initial surplus U, the longer the period considered when checking for ruin, the more likely it
is that ruin will occur.
lim Ψ(U,t) = Ψ(U) (47)
t→∞
The limit of the probability of ruin in finite time t can be approximately the probability of ultimate ruin
provided t is sufficiently large.

4.4 Probability of Ruin in Discrete time


So far the two probability of ruin considered have been in continuous time probabilities of ruin. There are
so-called because they check for ruin in continuous time. Practically it may be possible or even desirable
to check for ruin only at discrete time interval. For a given interval of time, denoted by h, the following
two discrete time probabilities of ruin are defined as;

Ψh (U) = P[U(t) < 0, for some t,t = h, 2h, 3h . . . ] (48)

Ψh (U,t) = P[U(τ) < 0, for some τ, τ = h, 2h, 3h . . . ,t − h,t] (49)


For convenience seek, the definition of Ψh (U,t) t is assumed to be an integer multiple of h. The figure
below shows the same realisation of the surplus process as given in figure 1, but the process is checked
at discrete time intervals.

57
Figure 2

The black dots shows the values of the surplus process at integer time interval (h = 1), the black and
white dots shows the value of the surplus process at time interval of length 21 . From the above graph in
discrete time with h = 1, ruin does not occur for this realisation of the surplus process before times but
ruin does occur (at time 2.5) in discrete time with h = 21
Below are five relationships between different discrete time probabilities of ruin for 0 ≤ U1 ≤ U2 and for
0 ≤ t1 ≤ t2 < ∞ That is
Ψh (U2 ,t) ≤ Ψh (U1 ,t) (50)
Ψh (U2 ) ≤ Ψh (U1 ) (51)
Ψh (U,t1 ) ≤ Ψh (U,t2 ) ≤ Ψh (U) (52)
lim Ψh (U,t) ≤ Ψh (U) (53)
t→∞
Ψh (U,t) ≤ Ψ(U,t) (54)
Intuitively, it is expected that the following two relationship are true since the probability of ruin in
continuous the could be approximated by the probability of ruin in discrete time with the same initial
surplus U, and time horizon t, provided ruin is checked for sufficiently often i.e. provided h is sufficiently
small
lim Ψh (U,t) ≤ Ψh (U,t) (55)
t→0+
lim Ψh (U) ≤ Ψh (U) (56)
t→0+
Formula (55) and (56) are true but their proofs is beyond the scope of the course.

4.5 Poisson and Compound Poisson Process


Here some assumptions will be made about the claim process, {N(t)}t≥0 , and the claim amount, {Xi }∞
i=1 .
The claim number process will be assumed to be a Poisson process, which will then result to a compound
Poisson process {S(t)}t≥0 for aggregate claims.

58
4.5.1 Poisson Process
Poisson process is just an example of a counting process. In this case the number of claims arising from
a risk is of interest. Since the number of claims is being counted over a period of time then the claim
number process {N(t)}t≥0 must satisfy the condition below

(i) N(0) = 0, there are zero claims at time t = 0.

(ii) For any t > 0,N(t) must be integer valued.

(iii) When s < t, N(s) ≤ N(t), that is the number of claims over time is non-decreasing.

(iv) When s < t, N(t) − N(s) represent the number of claims occurring in the time interval (s,t]

The claim number process {N(t)}t≥0 is defined to be a Poisson process with parameterλ if it satisfies the
following conditions

(i) N(0) = 0 and N(s) ≤ N(t) with s < t.

(ii)
P[N(t + h) = r|N(t) = r] = 1 − λ h + o(h),
P[N(t + h) = r + 1|N(t) = r] = λ h + o(h),
P[N(t + h) > r + 1|N(t) = r] = o(h) (57)

(iii) when s < t, the number of claims in the time interval (s,t] is independent of the number of claims
up to time s.

Condition (ii) implies that there can be a maximum of one claim in a very short time interval h. It also
implies that the number of claims in a time interval of length h does not depend on when that time interval
starts.

4.5.2 Compound Poisson Process


Having looked at the Poisson process, here, the Poisson process for the number of claims will be com-
bined with a claim amount distribution to yield a compound Poisson process for the aggregate claims.
The following three important assumptions are made;

(i) the random variables {Xi }∞


i=1 are independent and identically distributed (iid).

(ii) the random variables {Xi }∞


i=1 are independent of N(t) for all t ≥ 0.

(iii) The stochastic process {N(t)}t≥0 is a Poisson process whose parameter is denoted by λ . This
implies that for t ≥ 0, the random variable N(t) has a Poisson distribution with parameter λt such
that
e−λt (λt)k
P(N(t) = k) = , for k = 0, 1, 2, 3, . . . (58)
k
With these assumption the aggregate claim process, {S(t)}t≥0 is called a compound [Poisson process
with a Poisson parameter λ . Thus by comparing these assumption with those made under the Poisson
process, it can be observed that there is a connection between the two,i.e. {S(t)}t≥0 is a compound
Poisson process with a Poisson parameter λ .
Note that there is a slight change in terminology here, ’Poisson parameter λ ’ because "Poisson parameter
λt’ when a change is made from the process to the distribution. The cumulative distribution function
of the xi ’s will be denoted by F(x) and it is assumed that F(0) = 0, so that all claims are for positive

59
amounts.
The probability density function of the Xi ’s, if it exists will be denoted by f (x) and the kth moment about
zero of the Xi , if it exist will be denoted by mk such that

mk = E[Xik ] k = 1, 2, 3 . . .

Whenever the common generating function of the Xi exists, then its value at the point r will be denoted
by mX (r). Since for a fixed value of t, s(t) has a compound Poisson distribution. Recall from risk models
that the process {S(t)}t≥0 has mean
E[S(t)] = λtm1 ,
variance
Var[S(t)] = λtm2
and moment generating function ms (r) which is given by
 
λt MX (r)−1
MS (r) = e (59)

Here, it will be assumed that c > λ m1 , where c is the rate of premium income. This assumption is made
so that the insurer’s premium income (per unit time) is greater then the expected claim outgo (per unit
time).

4.5.3 Probability of Ruin in the short term


Ideally, if we know the distribution of the aggregate claim S(t), we can determine the probability of
ruin for the discrete model over a finite time interval directly( without referring to the models), just by
observing the cash flows involved.

Example 4.5.1 The aggregate claims arising during each year from a particular type of annual insur-
ance policy are assumed to follow a normal distribution with mean 1.35P and standard deviation 1.5P
, where P is the annual premium. The claims are assumed to arise independently. Insurers assess their
solvency position at the end of each year. A small insurer with an initial surplus of K100, 000 expects to
sell 50 policies at the beginning of the coming year in respect of identical risks for an annual premium of
K2, 500. The insurer incurs expenses of 0.1P at the time of writing each policy. Calculate the probability
that the insurer will prove to be insolvent at the end of the coming year. Ignore interest.

Solution 4.5.1 From the given data, the insurer’s surplus at the end of the coming year is given by

Surplus at th end of year one = initial surplus + premium − expenses − claims

U(1) = U(0) + kP + 0.1kP + −S(1)


where k is the number of policies sold thus

U(1) = 100, 000 + (50)(2500) − (0.1)(50)(2500) − S(1)

U(1) = 100, 000 + (50)(2500) − (0.1)(50)(2500) − S(1)


U(1) = 212, 500 − S(1)
We know that the distribution of S(1) is

S(1) ∼ N(1.35P, 1.5P)

60
Since 50 policies were sold then
 
2
S(1) ∼ N (1.35(2500)(50), (1.5(2500)) (50)

S(1) ∼ N(168750, 703125000)


So from
U(1) = 212, 500 − S(1)
we have that
U(1) = 212, 500 − S(1) < 0
=⇒ S(1) > 212, 500
Thus  
P S(1) > 212, 500

Using the z-score we have  


S(1) − µ 212, 500 − 168, 750
P > √
σ 703, 125, 000
P(Z > 1.65) = 1 − P(Z < 1.65) = 1 − φ (1.65) = 1 − 0.95053 = 0.04947
There the probability that the insurer will prove to be insolvent at the end of the coming year is 0.04947
or 4.947%

Task 4.5.1 Suppose that the insurer’s expects to sell 200 polices during the second year for the same
premium and expect to incur expenses at the same rate as in example (4.5.1). Determine the probability
that the insurer will prove to be insolvent at the end of the second year.

Example 4.5.2 The number of claims from a portfolio of policies has a Poisson distribution with param-
eter 60 per year. The individual claim amount distribution is lognormal with parameters µ = 2.5 and
σ 2 = 1.25. The rate of premium income from the portfolio is 1400 per year. If the insurer has an initial
surplus of 1000, estimate the probability that the insurer’s surplus at time 2 will be negative, by assuming
that the aggregate claims distribution is approximately normal.

Solution 4.5.2 In order to solve this, firstly we need to determine the mean and variance of the aggregate
claims in a two-year period.
Since we are looking the insurer’s surplus for a two-year period, thus

Expected value(λ ) = 2(60) = 120

Using this , we then determine the mean and variance using the formula for the first two moments of a
lognormal distribution.
Since 1 2 2
E[X r ] = erµ+ 2 r σ

with µ = 6 and σ 2 = 2.2,


For r = 1 1 2 (1.25)
E[X 1 ] = e(1)(2.5)+ 2 (1) = e3.125
For r = 2
1 2 (1.25)
E[X 2 ] = e(2)(2.5)+ 2 (2) = e7.5

61
Now we determine the mean and variance of S(2), that is

Mean: E[S(2)] = λ E[X] = 120e3.125

Variance: Var[S(2)] = λ E[X 2 ] = 120e7.5


Recall that ruin will occur if S(2) is greater than the initial surplus plus premium received such that

U(t) = U(0) + ct − S(2) < 0

=⇒ S(2) > U(0) + ct


Such that
P(S(2) > U(0) + ct)
P(S(2) > 1000 + 2(1400) = 3800)
=⇒ P(S(2) > 3800
Since S(2) ∼ N(E[s(2)],Var[s(2)]) , so that;
!
S(2) − E[s(2)] 3800 − 120e3.125
P(S(2) > 3800) = P p > √
Var[s(2)] 120e7.5

= P(Z > 2.29) = 1 − P(Z > 2.29) = 1 − φ (2.29) = 1 − 0.98899 = 0.01101


Therefore the probability that the insurer’s surplus at time 2 will be negative is approximately 0.01101
or 1.101%.

4.5.4 Premium Security Loading


So far, we have being denoting the rate of premium income per unit time as c, independent of the claim
outgo. In some circumstance it is more useful to think of the rate of premium income as being related to
the rate of claim outgo.
In order for an insurer to survive, the rate at which premium income comes in needs to be greater than
the rate at which claim are paid out. Thus if this does not hold then the insurer is certain to be ruined at
some point. Sometimes c will be written as

c = (1 + θ )λ m1 (60)

Where θ (> 0) is the premium loading factor, this represent the percentage by which the rate of premium
income exceeds the rate of claims outgo.

4.6 Adjustment Coefficient and Lundberg’s Inequality


In the section, introduce the adjustment coefficient and its association with the probability of ruin. Adjust-
ment coefficient is a parameter associated with risk, and the letters R and r will be used interchangeably
to denote the adjustment coefficient.

4.6.1 Lundberg’s Inequality


Lundberrg’s inequality states that
Ψ(U) ≤ e−RU , (61)
where U is the insurer’s initial surplus and Ψ(U) is the probability of ultimate ruin, R is a parameter
associated with the surplus process called the adjustment coefficient The value of R depends on the
distribution of aggregate claims and the rate of premium income loading.

62
4.6.2 Adjustment coefficient- Compound process
The adjustment coefficient is defined as a parameter associated with a surplus process which takes into
account two factor aggregate claims and premium.
This coefficient gives a measure of risk for a surplus process, When aggregate claims are a compound
Poisson process, the adjustment coefficient is defined in terms of the Poisson parameter, the mgf of the
individual claim amount and the premium income per unit time.
The adjustment coefficient R is therefore defined as the positive root of

λ MX (r) − λ − cr = 0 (62)

so, R is given by
λ MX (R) = λ + cR (63)
Both equation (62) and (63) implies that the value of the adjustment coefficient depends on the Poisson
parameter, the individual claim amount distribution and the rate of premium income. Thus using equation
(61), the adjustment coefficient becomes

MX (R) = 1 + (1 + θ )m1 R (64)

such that R is not dependent of the Poisson parameter but simply depends on the loading factor θ and the
individual claim amount distribution. In actuarial, e−RU is mostly used as an approximation to Ψ(U). R
can be interpreted as measuring risk. That is R is viewed as an inverse measure of risk.
The larger the value of R, the smaller the upper bound for Ψ(U) will be. Thus Ψ(U) would be expected
to decrease as R increases. R is a function of the parameter that affect the probability of ruin.
Example 4.6.1 An insurer knows from past experience that the number of claims received per month has
a Poisson distribution with mean 15, and that claim amounts have an exponential distribution with mean
500. The insurer uses a security loading of 30%. Calculate the insurer’s adjustment coefficient and give
an upper bound for the insurer’s probability of ruin, if the insurer sets aside an initial surplus of $1, 000.

Solution 4.6.1 Given that the equation for the adjustment coefficient is

MX (R) = 1 + (1 + θ )m1 R (65)


1

Now since X ∼ exp 500 such that mgf is
  −1
1
MX (R) = 1 − R = (1 − 500R)−1
λ
Also we have that θ = 30% and that
m1 = E[X] = 500,
then by substituting in equation (65) thus;

(1 − 500R)−1 = 1 + (1 + 0.30)500R

=⇒ 1 = (1 + 650R)(1 − 500R)
325000R2 − 150R = 0
3
R = 0( ignore) R =
6500
The upper bound is given by

Ψ(U) ≤ e−RU using Lundberg’s inequality

63
3
Ψ(U) ≤ e− 6500 (1000) = 0.63031386
Ψ(U) ≤ 0.63031
Reason why we ignored R = 0: Firstly R is defined to be the unique positive root, so R cannot be zero
Secondly, is that R = 0 should always be a solution to the equation, but since

MX (0) = e[e0 ] = 1 and 1 + (1 + θ )m1 × 0 = 1,

so R = 0 is always a solution , but it si ignore since from the Lundberg’s inequality Ψ(U) ≤ e−RU for
R = 0, then the upper bound is 1 for the probability of ruin which is obvious, since the upper bound for
any probability is 1.

4.7 Tutorial
1. Given that 0 < t1 < t2 < ∞ and 0 < U1 < U2 , below are some logical relation of the probabilities of
ruin. Explain meaning of each relationship given below.

(i) Ψ(U2 ,t) ≤ Ψ(U1 ,t) or Ψ(U2 ) ≤ Ψ(U1 )


(ii) Ψ(U,t1 ) ≤ Ψ(U,t2 ) ≤ Ψ(U)
(iii) limt→∞ Ψ(U,t) ≤ Ψ(U)
(iv) Ψh (U,t) ≤ Ψ(U,t)

2. A fourth year Actuarial Science Student, believe that claims on a portfolio of insurance policies
arrive as a Poisson process with parameter 100. Individual claim amounts follow a normal distri-
bution with mean 30 and variance 16 . The insurer calculates premiums using a premium loading
of 20% and has initial surplus of £100.

(i) Explain carefully the meaning of the ruin probabilities Ψ(100), Ψ(100, 1) and Ψ1 (100, 1)
(ii) Define the adjustment coefficient R.
(iii) Show that for this portfolio the equation to determine the adjusted coefficient R is given by

30R + 12.5R2 − ln(1 + 37.5R) = 0

(iv) Suppose we solve the equation in part a(ii) using Newton-Raphson method and R ≈ 0.0115,
compute an upper bound for Ψ(100) and an estimate of Ψ(100, 1).
(iv) Comment on the results in part a(iv)

3. Claims on portfolio of insurance policies arise as a Poisson process with parameter λ = 125. Indi-
vidual claim amounts, Xi follow a gamma distribution with mean 4 and variance 8. The insurance
company calculates premiums using a premium loading factor of 25% and has an initial surplus of
£50.

(i) Determine the adjustment coefficient value(s) for this portfolio


(ii) Use smaller value of R to calculate an upper bound for Ψ(50).
(iii) Calculate an estimate of Ψ(50, 1), using a Normal approximation.
(iv) Suppose the the second parameter of the Gamma distribution now reduces to 0.4, explain
what would happen to the estimate of Ψ(50, 1), without carrying out any further calculations.
(v) Suggest two ways in which the insurance company could reduce Ψ(50, 1).

64
4. Claim events on a portfolio of insurance policies follow a Poisson process with parameter λ . In-
dividual claim amounts, X, follow a Normal distribution with parameters µ = 500 and σ 2 = 200.
Assume the insurance company calculates premiums using a premium loading factor of 20%

(i) Show that the equation for obtaining the adjustment coefficient is given by

100r2 + 500r − ln(1 + 600r) = 0

(ii) Show that r = 0.000708 to three significant figures and obtain an upper bound on the proba-
bility of ruin, using Lundberg’s inequality given that the insurance company’s initial surplus
is $5000

65
5 Generalized Linear Models
Objectives:

• Introduction to GLMs

• Exponential families of Distributions

• Components of GLMs

5.1 Introduction
The generalized linear model (GLM) can be considered an extension of the linear model discussed in the
Statistical Inference course. In particular, it will be seen that the regression model is just a simple GLM,
and many of the ideas from regression modeling will be used in the discussion of GLMs.
The main difference between a linear model and a GLM is that a GLM allows the distribution of the data
to be non-normal. This is important in the actuarial field because, very often, data do not follow a normal
distribution. For instance, in modeling the force of mortality, µx , and the binomial distribution for the
initial rate of mortality, qx , In addition, in general insurance, the Poisson distribution is often used for
modeling the claim frequency (i.e., number of claims received) and the Gamma, Lognormal, or Pareto
distribution for claim sererity (i.e., size of the claim or claim amounts). Basically, GLMs are widely used
both in general and in life insurance. They are used to:

(i) Determine rating factors to use in the GLMs(i.e rating factors are measurable or categorical factors)
that are used as proxies for risk in setting premium.

(ii) Estimate an appropriate premium to charge for a particular policy given the level of risk present.

The main objective of the data analysis exercise is usually to decide which variables or factors are im-
portant predictors for the risk being considered and then quantify the relationship between the predictors
and the risk in order to set appropriate premium levels. The aim is to cover the basic theory of GLMs
that is necessary for application, such as those highlighted above. The GLM relates a variable called
the response variable or dependent variable which we want to predict to variables or factors called
predictor, covariates, or independent variables.
In order to do this, it is required that we first define the distribution of the response variable. Then the
covariates can be related to the response, allowing for the random variation of the data. Thus, the first step
is to consider the general form of distribution known as exponential families, which are used in GLMs.

5.2 Exponential families Distibution


A distribution for a random variable Y belongs to an exponential family if its density has the following
form: h i
(yθ −b(θ ))
+c(y,ϕ)
f (y; θ , ϕ) = e a(ϕ)
(66)
where a, b and c are functions.
It should be noted that this is not a unique form, and it is possible to see an exponential family defined in
slightly different ways elsewhere.
There are two parameters in the above density function. θ , which is called the natural parameter is the
one that is relevant to the model for relating the response (Y) to the covariates, and ϕ is known as the
scale parameter or dispersion parameter.
When attempting to demonstrate that a distribution belongs to the exponential family, keep in mind that θ

66
is merely a function of µ = E[Y ]. We will see later in the chapter how it is used to relate the answer to the
variables. When a distribution has two parameters, such as the N(µ, σ 2 ), one approach to determining
the scale parameter ϕ is to take ϕ to be the "other" parameter in the distribution, i.e the parameter other
than the mean. For example, in the case of the normal distribution, we take ϕ = σ 2 . In the case where a
distribution has one parameter, such as the Poi(λ ) , we take ϕ = 1.
In order to motivate these definitions and the subsequent developments, we first consider the normal
distribution.

5.2.1 Normal Distribution


Given that
1 (y−µ)2

f (y) = √ e σ2
2πσ 2
Then !−1  
√ y2 −2yµ+µ 2
− 12
σ2
f (y) = 2πσ 2 e

!− 1  
2 y2 −2yµ+µ 2
− 12
σ2
f (y) = 2πσ 2 e

Writing every term in term of an exponent we have


   
y2 −2yµ+µ 2 y2 −2yµ+µ 2
1 2 − 12 − 12 log 2πσ 2 − 12
f (y) = e− 2 log 2πσ e σ2
=e σ2

µ2  
yµ+ 2 y2
− 21 +log 2πσ 2
σ2 σ2
f (y) = e
By comparison with equation (66) we have that θ = µ and ϕ = σ 2 so that

µ2 θ 2
b(θ ) = =
2 2
a(ϕ) = σ 2 = ϕ
and
y2 1 y2
   
1 2
c(y, ϕ) = − + log 2πσ = − + log 2πϕ
2 σ2 2 ϕ
Therefore
µ2  
yµ+ 2 y2
− 12 +log 2πσ 2
σ2 σ2
f (y; θ , ϕ) = e (67)
which is in the form of an exponential families distribution with

θ =µ
ϕ = σ2
b(θ ) = ϕ
and
y2
 
1
c(y, ϕ) = − + log 2πϕ
2 ϕ
Thus, the natural parameter for the normal distribution is µ and the scale parameter is σ 2 .

67
5.2.2 Mean and Variance of Exponential families distribution
!
Consider the log -likelihood function, l(y; θ , ϕ) = log f (y; θ , ϕ) . Then using the results.

"  #
∂ 2l ∂l 2
   
∂l
E =0 E + E =0
∂θ ∂θ2 ∂θ

We can show that the mean and variance of an exponential family distribution are

E[Y ] = b′ (θ ) and Var[Y ] = a(ϕ)b”(θ )

Derivation of the mean and Variance


From equation (66), we apply the log operator we have
yθ − b(θ )
l= + c(y, ϕ) (68)
a(ϕ)
On differentiating l with respect to θ
∂l y − b′ (θ )
=
∂θ a(ϕ)
For the mean,we take the expectation of the above equation we get
y − b′ (θ )
   
∂l
E =E =0
∂θ a(ϕ)
E[y] − b′ (θ )
=⇒ =0
a(ϕ)
E[Y ] = b′ (θ )
Differentiating w.r.t θ we get
∂ 2l −b′′ (θ )
=
∂θ2 a(ϕ)
For variance, using "  #
∂ 2l ∂l 2
 
E + E =0
∂θ2 ∂θ
"  #
y − b′ (θ ) 2
 
b”(θ )
E − +E =0
a(ϕ) a(ϕ)
b′′ (θ ) E[(y − b′ (θ ))2 ]
=⇒ − + =0
a(ϕ) a(ϕ)2
E[(y − b′ (θ ))2 ] = a(ϕ)b′′ (θ )
Now recall that Var[Y ] = E[(y − b′ (θ ))2 ] then

Var[Y ] = a(ϕ)b′′ (θ )

Task 5.2.1 Using the expression of mean and variance under the exponential family distribution, verify
that the mean and variance of the normal distribution random variable are E[X] = µ and Var[X] = σ 2

68
5.2.3 Poison Distribution
Suppose the
λ y e−λ
f (y) = , y = 0, 1, 2, . . .
y!
f (y) = λ y e−λ (y!)−1
By writing every term as a power of an exponent we have

f (y) = ey log λ e−λ e− log y! = ey log λ −λ −log y!


y log λ −λ
−log y!
=⇒ f (y) = e 1

Comparing with
yθ −b(θ )
+c(y,ϕ)
f (y) = e a(ϕ)

We have;
θ log λ =⇒ λ = eθ
ϕ =1 a(ϕ) = 1
b(θ ) = λ =⇒ b(θ ) = eθ
and
C(y, ϕ) = − log y!
Therefore
y log λ −λ
−log y!
f (y) = e 1

which is in the form of an exponential families distribution with

θ log λ

ϕ =1 a(ϕ) = 1
b(θ ) = eθ
and
C(y, ϕ) = − log y!

Task 5.2.2 Using the expression of mean and variance under the exponential family distribution, verify
that the mean and variance of the Poisson distribution random variable are E[X] = Var[X] = λ

5.2.4 Binomial Distribution


Suppose that z ∼ Bin(n, µ) then
z
y= =⇒ z = ny (69)
n
Recall that  
n
f (z) = µ z (1 − µ)n−z
z
Using equation (69) we have that
 
n
f (y) = µ ny (1 − µ)n−ny
ny

69
Writing every term as a power of an exponent we get
!
n
log
f (y) = e ny eny log µ e(n−ny) log(1−µ)
!
n
log +ny log µ+(n−ny) log(1−µ)
f (y) = e ny

!
n
log +ny log µ+(n−ny) log(1−µ)
f (y) = e ny

  !
n
n y log µ+(1−y) log(1−µ) +log
f (y) = e ny
  !
n
n y log µ−y log(1−µ)+log(1−µ) +log
f (y) = e ny
    !
µ n
n y log 1−µ +log(1−µ) +log
f (y) = e ny
   
µ
y log 1−µ +log(1−µ) !
n
1 +log
f (y) = e n ny

We Compare with
yθ −b(θ )
+c(y,ϕ)
f (y) = e a(ϕ)

We get; 
µ
θ = log , ϕ =n
1−µ
1
a(ϕ) =
n
1
=⇒ a(ϕ) =
ϕ
b(θ ) = − log(1 − µ)
Since b(θ ) is a function of θ , to achieve that we need to manipulate the expression of θ that is;
Letting
µ µ
eθ = =⇒ 1 + eθ = 1 +
1−µ 1−µ
1
1 + eθ = = (1 − µ)−1
1−µ
log(1 + eθ ) = − log(1 − µ)
=⇒ b(θ ) = log(1 + eθ )
 
n
c(y, ϕ) = log
ny
 
ϕ
=⇒ c(y, ϕ) = log
ϕy

70
Therefore    
µ
y log 1−µ +log(1−µ) !
n
1 +log
f (y) = e n ny

which is in the form of an exponential families distribution with


 
µ
θ = log
1−µ

1
ϕ =n a(ϕ) =
ϕ
b(θ ) = log(1 + eθ )
and  
ϕ
c(y, ϕ) = log
ϕy

Task 5.2.3 Using the expression of mean and variance under the exponential family distribution, verify
that the mean and variance of the Binomial distribution random variable are E[X] = np Var[X] =
np(1 − p)

5.2.5 Gamma Distribution


Let Y ∼ Gamma(α, µ) where
α α
µ= =⇒ λ =
λ µ
Recall that the Gamma distribution probability distributional function is

λ α α−1 −λ y
f (y) = y e , α > 0, λ > 0, y > 0
Γ(α)

By substituting λ with αµ , we get


 α
α
µ α
f (y) = yα−1 e− µ y
Γ(α)
αα α
α−1 − µ y
f (y) = y e
µ α Γ(α)
α
f (y) = α α µ −α Γ(α)−1 yα−1 e− µ y
Writing every term on the right hand side as a power of the exponent we get;
α
f (y) = eα log α e−α log µ e− log Γ(α) e(α−1)y e− µ y
α
f (y) = eα log α−α log µ−log Γ(α)+(α−1) log y− µ y
α
f (y) = e− µ y−log µ+α log α−log Γ(α)+(α−1) log y
 
α − µy −log µ +α log α−log Γ(α)+(α−1) log y
f (y) = e

71
y
− µ −log µ
1 +α log α−log Γ(α)+(α−1) log y
f (y) = e α

Compare with
yθ −b(θ )
+c(y,ϕ)
f (y) = e a(ϕ)

We get;
1
θ =− , ϕ =α
µ
1
a(ϕ) =
α
1
=⇒ a(ϕ) =
ϕ
b(θ ) = log µ
Since b(θ ) is a function of θ , to achieve that we need to manipulate the expression of θ that is;
From
1
θ =−
µ
1
−θ = = µ −1 =⇒ −θ = µ −1
µ

log(−θ ) = − log µ
=⇒ log µ = − log(−θ )
=⇒ b(θ ) = − log(−θ )
c(y, ϕ) = α log α − log Γ(α) + (α − 1) log y
=⇒ c(y, ϕ) = ϕ log ϕ − log Γ(ϕ) + (ϕ − 1) log y
Therefore y
− µ −log µ
1 +α log α−log Γ(α)+(α−1) log y
f (y) = e α

which is in the form of an exponential families distribution with


1
θ =−
µ
1
ϕ =α a(ϕ) =
ϕ
b(θ ) = − log(−θ )
and
c(y, ϕ) = ϕ log ϕ − log Γ(ϕ) + (ϕ − 1) log y

Task 5.2.4 Using the expression of mean and variance under the exponential family distribution, verify
that the mean and variance of the Binomial distribution random variable are E[X] = αλ Var[X] = λα2

Task 5.2.5 Show that the exponential can be written in the form of a member of the exponential family.
Hence use the expression of mean and variance under the exponential family distribution, verify that the
mean and variance of the Exponential distribution random variable are E[X] = λ1 Var[X] = λ12

72
5.3 Components of GLM
The Generalized linear models is consisted of three components namely; distribution of the data, linear
predictors and link functions.

Distribution of the data


In this situation, the data can follow any of the distributions that can be expressed as exponential families.
For example, we may use a gamma distribution to model the sizes of motor insurance claims( claim
amounts), a Poisson distribution to model the number of claims, or a binomial distribution to model the
likelihood of developing a certain disease.

Linear predictors
The predictor linear η is a function of the covariates, that is, it is a function that associates the covariates
or predictors. For example,
η = β0 + β1 x
where "η" is a Greek letter.
As an example, if the response variable is weight, this linear predictor would be adequate for a model
in which height, x, was assumed to be the only co-variate. It is important to note that a linear predictor
is linear in the parameters β0 and β1 . It is not required to be linear in the covariates or the independent
variables; for example,
η = β0 + β1 x2
is also a linear predictor. The other main type of covariate is a factor, which takes a categorical value.
For example, the gender of the policyholder is either male or female, which constitutes a factor with 2
categories (or levels). This type of covariate can be parameterised so that the linear predictor has a term
α1 for a male, and a term α2 for a female.
Thus, a model which includes an age effect and an effect for the gender of the policyholder could have a
linear predictor
αi + β x
where i = 1 for male and i = 2 for female.
The next step would be to use MLE on a set of past data to come up with estimates for each of the
parameters.
Below are some examples of linear predictors

Model Linear predictor


age β0 + β1 x
age+age 2 β0 + β1 x + β2 x2
age+duration β0 + β1 x + β2 x1
log(age) β0 + β1 log x
gender αi
vehicle rating group βj
gender+ vehicle rating group αi + β j
gender×vehicle rating group αi + β j + γi j or αi j

Link functions
A link function links the mean of the response variable to the linear predictor. It is necessary to connect
the mean response to the linear predictor. E[Y ] defines the relationship between the response and the

73
covariates. Consider first a straight line regression model for normally distributed data, which can be
represented as:
Y ∼ N(µ, σ )
with µ = β0 + β1 x. Both µ and E[Y ] will be used interchangeably. In this situation, the link is straight-
forward:
E[Y ] = Linear predictors
As a result,
µi = ηi = β0 + β1 x
In general we take some function of the mean response and this function is called the link function.
Equating the link function to the linear predictors we have in general the following relationship

g(µ) = η

where g is the link function and η is the linear predictors.


There are a number of functions that are appropriate for the distributions above. For each distribution,
the natural, or canonical, link function is defined by. Examples of common canonical link functions
g(µ) = θ (µ)

Distribution Type Link function


Normal Identity g(µ) = µ
Poisson Log g(µ) = log
 µ 
µ
Binomial Logit g(µ) = log 1−µ
1
Gamma Inverse(reciprocal) g(µ) = µ

Under the Gamma distribution showed that


1
θ =−
µ
The minus sign is dropped in the canonical link function. This doesn’t affect anything since constants
will be absorbed into the parameters in the linear predict. These link functions work well for each of the
above distributions, but it is not obligatory that they be used in each case. For example, you could use
the identity link function in conjunction with the Poisson distribution, you could use the log link function
for data which had a gamma distribution, and so on. However, you need to consider the implications of
the choice of the link function on the possible values for µ . For example, if the data have a Poisson
distribution then µ must be positive.
The link function, as the name implies, is the missing link. Remember that the goal of a GLM is to find a
relationship between the mean of the response variable and the covariates. We can make the mean µ the
subject of the formula by setting the link function

g(µ) = η

by assuming that the link function is invertible:

µ = g−1 (η)

To further appreciate the concept consider the following example.

74
Example 5.3.1 Suppose that we are trying to model the number of claims on car insurance policies. The
response variable, Yi , is the number of claims on Policy i. We decide that a Poisson distribution would
be appropriate:
Yi ∼ Poi
Consider a model where we believe that the only covariate is the age, xi , of the policyholder. The linear
predictor is
ηi = α + β x1
Show that the relationship between the mean of the response variable and the covariate is given by

µi = eα+β xi

Solution 5.3.1 Recall that the link function under the poisson distribution is given by

g(µ) = log µ

Now we set the link function equal to the linear predictor we have

g(µi ) = ηi

=⇒ log µi = α + β xi
We invert the formula so that µ is the subject of the formula that is

µi = eα+β xi

We now have a relationship between the mean of the response variable and the covariate.

Example 5.3.2 Suppose that we are setting up a model to predict the pass rate for a particular student
in BAS380 course exam. We might expect there to be many factors that affect whether a student is likely
to pass or not. We might decide to set up a three-factor model, so that the probability of passing is a
function of:

(i) the number of test x1 written by the student (a value from 0 to 4)

(ii) the student’s mark on the mock exam x2 (on a scale from 0 to 100)

(iii) whether the student had attended tutorials or not (Yes/No).

We therefore decide to use the linear predictor:

η = αi + β1 x1 + β2 x2

where x1 is the number of test written, x2 is the mark on the mock exam, and αi takes one value for those
attending tutorials and a different value for those who do not.
Suppose that the data follow a binomial distribution and the MLE of the parameters are

αYes = −1.501 αNo = −3.196 β1 = 0.5459 β2 = 0.0251

Given that a student attended tutorials, sat for three test and scores 85% on the mid-semester examina-
tion. Using the model, determine the student probability of passing BAS380 course in the upcoming end
of semester final examinations.

Solution 5.3.2

75
5.4 Tutorial
1. A discrete probability distribution is defined by
 
n 1 2
f (y, µ) = µ ny (1 − µ)n−ny y = 0, , , . . . , 1
ny n n

where µ is a parameter that lies between 0 and 1.

(i) Show that this distribution belongs to an exponential family.


(ii) State the three main components that need to be taken into account when constructing a
generalised linear model.
(iii) Suggest a natural choice of link function if the response variable followed the distribution
defined above.

2. The probability density function of a gamma distribution is given in the following parameterised
form:
αα α
f (y) = α yα−1 e− µ x x>0
µ Γ(α)
where µ = αλ .

(i) Express this density in the form of a member of the exponential family, clearly specifying all
the parameters, and hence state its canonical link function.
(ii) Hence show that the mean and variance of the distribution are given by
α α
and 2 respectively.
λ λ
3. Show that if a random variable has the exponential distribution it belongs to an exponential family.

76
6 Run-Off Triangles
Objectives:
• Define a development factor and show how a set of assumed development factors can be used to
project the future development of a delay triangles.
• Describe and apply the basic chain ladder method for completing the delay triangle.
• Show how the basic chain ladder method can be adjusted to make explicit allowance for inflation.

6.1 Introduction
Run-off triangles (or delay triangles) are an essential topic in the practical work of general insurance
actuaries who predict future claim numbers and amounts using spreadsheets and other computer software.
In this section, we will look at only two standard ways for projecting run-off triangles: the basic chain
ladder approach and the inflation-adjusted chain ladder method.
Run-off triangles or delay triangles usually arise in types of insurance, particularly non-life insurance,
where it may take some time after a loss until the full extent of the claims that have to be paid are known.
It is important that the claims are attributed to the year in which the policy was written. The insurance
firm must know how much it is obligated to pay in claims in order to determine its surplus. However, it
may be many years before the exact claim totals are known. There are several reasons for the delays in
finalising claim totals. The delay may occur prior to claim notification and/or between claim notification
and the final settlement.
It is obvious that, while the insurance company does not know the actual number of total claims each
year, it must attempt to estimate that figure with as much confidence and accuracy as possible.

6.2 Types of Reserves


General insurers must be able to estimate the final cost of claims for a variety of reasons. For instance, in
order to compute future premium rates, they must know the total cost of paying claims. They must also
establish reserves in their accounts to ensure that they have enough assets to cover all of their obligations.
The figure shows the normal phases required in settling a general insurance claim:
Claim Event Occurs → Claim Reported→ Claim Payments Made→ Claim File closed
After the claim event occurs (for example, a policyholder being involved in a car accident or being bur-
gled), the policyholder will notify the insurer. In due course, the insurer will make the payments that are
required (for example, paying for car repairs, reimbursement to an injured person, or the cost of replac-
ing stolen things). A single claim may result in many payments. When the insurer determines that no
additional payments are required for this claim, the claim file will be closed.
A general insurer will need to set up reserves to cover its liabilities for future payments in respect of acci-
dents that have already occurred. These reserves will relate to claims at different stages in the settlement
process. Basically reserves will be required for outstanding reported claims and IBNR claims.

Definition 6.2.1 An IBNR (pronounced "I.B.N.R.") claims reserve is necessary for claims that have been
incurred but not reported, which means that the claim event has occurred but has not yet been reported
to the insurer.
Definition 6.2.2 An outstanding reported claims reserve is required in respect of claims that have been
reported but have not yet been closed.

77
6.3 Presentation of claims data
There are various ways of presenting claim data, each emphasising a distinct component of the data. They
will be shown as a triangle here, as is the most common format.

6.3.1 Components of the table


The following are terminologies used in the table presenting the claims data; accident year and delay,
or development years

Definition 6.3.1 The accident year is the year in which an incident occurred and the insurer was put at
risk.

Definition 6.3.2 The development years sometimes called delay years, is the number of years until a
payment is made.

The claims data are organized by accident year and development year. The table below shows an example
of claim data organized by accident year and development year. The development of claims by month or
quarter may be relevant in some types of insurance, but the principles remain the same.

Example 6.3.1 The figures in the table below are in units of ’000’, but for convenience we will not write
them out in full in our calculations

Cumulative claim payments

Accident Year Development Year


0 1 2 3 4
2019 786 1410 2216 2440 2519
2020 904 1575 2515 2796
2021 995 1814 2880
2022 1220 2142
2023 1182

Table 1: Runoff Triangles

The data provided are cumulative and indicate total payments made by the end of each development year.
They were compiled following the end of the 2023 accident year. Only payments with a delay of 0 have
been reported for the accident year 2023. Payments with delays of 0 and 1 have been reported for the
2022 accident year, and so on.

6.4 Estimating future claims


We are now looking at some methods for estimating the missing figures in the table. In this section we
will discuss only two methods namely basic chain ladder method and inflation adjusted chain ladder
method.
The fundamental assumption in assessing outstanding claims is the runoff pattern. The most basic as-
sumption is that payments will be distributed similarly in each accident year. The proportionate increases
in known cumulative payments from one development year to the next can then be used to compute pre-
dicted cumulative payments for subsequent development years. However, as the example below shows,
there are other options for determining which ratio should be used to forecast future claims. The ratios
used to forecast future claims are known as development factors, or link ratios.

78
A statistical model for run-off triangles is basically written as

Ii j = r j · s j · xi+ j + ei j (70)

Where
Ii j represents the incremental claims
r j is the development factor for year j, representing the proportion of claim payments in Development
Year j. Each r j is independent of the Origin Year i.
s j is a parameter varying by Origin Year, i, representing the exposure, for instance the number of claims
(or claim amount) incurred in the Origin Year i.
xi+ j is a parameter varying by calendar year, for example representing inflation, and
ei j is an error term.

6.5 Basic chain ladder method


This section illustrate how to use the basic chain ladder approach to complete the runoff triangle calcula-
tions. Consider the following example.
Example 6.5.1 Consider to example (6.3.1) and use basic chain ladder method to determine reserve that
needs to be held at the end of 2023.

Solution 6.5.1 We start by calculating the development factor or link ratios.


1410 + 1575 + 1814 + 2142
r1 = ≈ 1.777
786 + 904 + 995 + 1220

2216 + 2525 + 2880


r2 = ≈ 1.586
1410 + 1575 + 1814 + 2142
2440 + 2796
r3 = ≈ 1.107
2216 + 2525
2519
r4 = ≈ 1.032
2440
Using the development factors, we can compute the cumulative payment missing in the table. Let Ci j be
the cumulative claim amount of ith accident year under jth development year.
2020
C2020,4 = (C2020,3 )(r4 ) = (2796)(1.032) = 2885
2021
C2021,3 = (C2021,2 )(r3 ) = (2880)(1.107) = 3188
C2021,4 = (C2021,2 )(r3 )(r4 ) = (2880)(1.107)(1.032) = 3290
2022
C2022,2 = (C2022,1 )(r2 ) = (2142)(1.586) = 3397
C2022,3 = (C2022,1 )(r2 )(r3 ) = (2142)(1.586)(1.107) = 3761
C2022,4 = (C2022,1 )(r2 )(r3 )(r4 ) = (2142)(1.586)(1.107)(1.032 = 3881
2023
C2023,1 = (C2022,0 )(r1 ) = (1182)(1.777) = 2100

79
C2023,2 = (C2023,0 )(r1 )(r2 ) = (1182)(1.777)(1.586) = 3331
C2023,3 = (C2023,0 )(r1 )(r2 )(r3 ) = (1182)(1.777)(1.586)(1.107) = 3688
C2023,4 = (C2023,0 )(r1 )(r2 )(r3 )(r4 ) = (2142)(1.777)(1.586)(1.107)(1.032) = 3806
Thus the complete table will be

Cumulative claim payments

Accident Year Development Year


0 1 2 3 4
2019 786 1410 2216 2440 2519
2020 904 1575 2515 2796 2885
2021 995 1814 2880 3188 3290
2022 1220 2142 3397 3761 3881
2023 1182 2100 3331 3688 3806

Table 2: Complete table

With the information in the above about we can now determine the insurer’s total reserves Therefore

Accident Year Reserves


2019 -
2020 89
2021 410
2022 1739
2023 2624
Total 4862

Table 3: Reserves

the insurer’s need to reserve a sum of K4, 862, 000 in order to settle all the outstanding claims.

Example 6.5.2 The table below shows the number of home insurance claims reported in each develop-
ment year from 2021 to 2024 for accident years. Estimate the total final number of claims resulting from
accidents that occurred between January 1, 2021 and December 31, 2024 using the basic chain ladder
method.

Number of claims reported

Accident Year Development Year


0 1 2 3
2021 17500 5000 2250 750
2022 21000 6200 2750
2023 18800 5500
2024 21300

Solution 6.5.2 We first cumulate the figure in the table given, we get

80
Number of claims reported

Accident Year Development Year


0 1 2 3
2021 17500 22500 24750 25500
2022 21000 27200 29950
2023 18800 24300
2024 21300

We calculate the development factors or link ratios.


22500 + 27200 + 24300
r1 = ≈ 1.2914
17500 + 21000 + 18800
24750 + 29950
r2 = ≈ 1.1006
22500 + 27200
25500
r3 = ≈ 1.0303
24750
We project future claim
2022
C2022,3 = (C2022,2 )(r3 ) = (29950)(1.0303) = 30857
2023
C2023,2 = (C2023,1 )(r2 ) = (24300)(1.1006) = 26745
C2023,3 = (C2023,1 )(r2 )(r3 ) = (24300)(1.0303)(1.1006) = 27555
2024
C2024,1 = (C2024,0 )(r1 ) = (21300)(1.2914) = 27507
C2024,2 = (C2024,0 )(r1 )(r2 ) = (21300)(1.2914)(1.1006) = 30274
C2024,3 = (C2024,0 )(r1 )(r2 )(r3 ) = (21300)(1.2914)(1.1006)(1.0303 = 31191
Thus the complete table is

Number of claims reported

Accident Year Development Year


0 1 2 3
2021 17500 22500 24750 25500
2022 21000 27200 29950 30857
2023 18800 24300 26745 27555
2024 21300 27507 30274 31191

81
Therefore the estimated total final number of claim is just the summation of column under development
year 3 that is;
Estimated total final number of claims = 25500 + 30857 + 27555 + 31191 = 115103 claims.

6.6 Inflation adjusted chain ladder method


6.6.1 Dealing with past inflation
Payments in the run-off triangle will be affected by claims inflation by the calendar year of payment.
In our statistical model, the claims inflation were represented by the x’s in equation (70). The inflation-
adjusted chain ladder approach works by adjusting the values in the triangle to account for inflation
effects.
In the model addressed here, claims inflation is assumed to be the same yearly rate for all claims within
a given calendar year of payment. Each payment year corresponds to a diagonal in the triangle.
When adjusting for inflation, the payments in each calendar year, rather than the cumulative totals, must
be considered. The first step is to determine incremental payments from the cumulative totals by dif-
ferencing along each row. Consider the following example for an illustration of the inflation adjusted
chain.

Example 6.6.1 Consider example (6.3.1) and suppose that the annual claim payments inflation rates
over the 12 months up to the middle of the given year are as follows

2020 5.1%

2021 6.4%
2022 7.3%
2023 5.4%
Use the inflation adjusted chain ladder method to compute the insurer’s claim reserves.

Solution 6.6.1 The first step under this method is to compute the incremental payments from the cumu-
lative total, by diferencing along each row.

Incremental claim claim payments

Accident Year Development Year


0 1 2 3 4
2019 786 624 806 224 79
2020 904 671 940 281
2021 995 819 1066
2022 1220 922
2023 1182

82
Since we have the inflation rates, we can then determine the index for each year so that we can adjust
the increments to the mid of 2023 prices.
Acident Year Index
2019 (1 + 5.1%)(1 + 6.4%)(1 + 7.3%)(1 + 5.4%) = 1.265
2020 (1 + 6.4%)(1 + 7.3%)(1 + 5.4%) = 1.203
2021 (1 + 7.3%)(1 + 5.4%) = 1.131
2022 (1 + 5.4%) = 1.054
Now we adjust each increment by multiplying with the corresponding year index i.e

I2019,0 = 1.265(786) = 994


I2019,1 = 1.203(624) = 781
I2019,2 = 1.131(806) = 912
I2019,3 = 1.054(224) = 236

I2019,4 = 79 remains the same because we do not have the index for that year.

I2020,0 = 1.203(904) = 1088


I2020,1 = 1.131(671) = 759
I2020,2 = 1.054(904) = 991

I2020,3 = 281 remains the same.

I2021,0 = 1.131(995) = 1125


I2021,1 = 1.054(819) = 863

I2021,2 = 1066 remains the same.

I2022,0 = 1.054(1220) = 1286

I2022,1 = 922 remains the same.


I2023,0 = 1182 remains the same.
Thus the adjusted increment will be

Adjusted Incremental claim payments

Accident Year Development Year


0 1 2 3 4
2019 994 751 912 236 79
2020 1088 759 991 281
2021 1125 863 1066
2022 1286 922
2023 1182

83
We get the adjusted cumulative claim payment.

Adjusted Cumulative claim payments

Accident Year Development Year


0 1 2 3 4
2019 994 1745 2657 2893 2972
2020 1088 1847 2838 3119
2021 1125 1988 3054
2022 1286 2208
2023 1182

We start by calculating the development factor or link ratios.


1745 + 1847 + 1988 + 21208
r1 = ≈ 1.7334
994 + 1088 + 1125 + 1286

2657 + 2838 + 3054


r2 = ≈ 1.5321
1745 + 1847 + 11988
2893 + 3119
r3 = ≈ 1.0941
2657 + 2838
2972
r4 = ≈ 1.0273
2893
Using the development factors, we determine the expected cumulative claim payment.
C2020,4 = (C2020,3 )(r4 ) = (3119)(1.0273) = 3204
2021
C2021,3 = (C2021,2 )(r3 ) = (3054)(1.0941) = 3341
C2021,4 = (C2021,2 )(r3 )(r4 ) = (3054)(1.0941)(1.0273) = 3432
2022
C2022,2 = (C2022,1 )(r2 ) = (2208)(1.5321) = 3383
C2022,3 = (C2022,1 )(r2 )(r3 ) = (2208)(1.5321)(1.0941) = 3701
C2022,4 = (C2022,1 )(r2 )(r3 )(r4 ) = (2208)(1.5321)(1.0941)(1.0273) = 3802
2023
C2023,1 = (C2022,0 )(r1 ) = (1182)(1.7334) = 2049
C2023,2 = (C2023,0 )(r1 )(r2 ) = (1182)(1.7334)(1.5321) = 3139
C2023,3 = (C2023,0 )(r1 )(r2 )(r3 ) = (1182)(1.7334)(1.5321)(1.0941) = 3434
C2023,4 = (C2023,0 )(r1 )(r2 )(r3 )(r4 ) = (1182)(1.7334)(1.5321)(1.0941)(1.0273) = 3528

Thus the complete table will be

84
Cumulative claim payments

Accident Year Development Year


0 1 2 3 4
2019 994 1745 2657 2893 2972
2020 1088 1847 2838 3119 3204
2021 1125 1988 3054 3341 3432
2022 1286 2208 3383 3701 3802
2023 1182 2049 3139 3434 3528

Table 4: Complete table

6.6.2 Dealing with future inflation


The forecasts of cumulative payments in the above example do not account for future inflation. An
assumed rate of future inflation will be required to forecast the actual payments. Again, non-cumulative
data must be converted rather than cumulative totals before being adjusted for future inflation in the same
manner as when dealing with past inflation.

Example 6.6.2 Consider solution (6.6.1) and apply an annual inflation rate of 10% (at 30 June) to the
data in table 4.

Solution 6.6.2 Applying an annual future inflation of 10%, we get the index for the years 2024, 2025,
2026 and 2027.
Acident Year Index
2024 (1 + 10%) = 1.1
2025 (1 + 10%)2 = (1.1)2
2026 (1 + 10%)3 = (1.1)3
2027 (1 + 10%)4 = (1.1)4
We get the adjusted increment by multiplying with the corresponding year index.

Incremental claim payments

Accident Year Development Year


0 1 2 3 4
2019
2020 85
2021 289 91
2022 1175 318 101
2023 867 1090 295 94

85
Getting the adjusted increment using the future inflation rates

I2020,4 = (85)(1.1) = 94
I2021,3 = (289)(1.1) = 318
I2021,4 = (91)(1.1)2 = 110
I2022,2 = (1175)(1.1) = 1293
I2022,3 = (318)(1.1)2 = 385
I2022,4 = (101)(1.1)3 = 134
I2023,1 = (867)(1.1) = 954
I2023,2 = (1090)(1.1)2 = 1319
I2023,3 = (295)(1.1)3 = 393
I2023,4 = (94)(1.1)4 = 138

Thus the new future adjusted increments.

Adjusted Incremental claim payments

Accident Year Development Year


0 1 2 3 4
2019
2020 94
2021 318 110
2022 1293 385 134
2023 954 1319 393 138

Now we accumulate the claim payment using the original figures

Cumulative claim payments

Accident Year Development Year


0 1 2 3 4
2019 786 1410 2216 2440 2519
2020 904 1575 2515 2796 2890
2021 995 1814 2880 3196 3306
2022 1220 2142 3435 3820 3954
2023 1182 2136 3455 3848 3986

Table 5: Complete table

86
We can now determine the insurer’s total reserves
Accident Year Reserves
2019 -
2020 94
2021 426
2022 1812
2023 2804
Total 5136

Table 6: Reserves

Therefore the insurer’s need to reserve a sum of K5, 136, 000 in order to settle all the outstanding claims.

6.7 Tutorial
1. The table below shows claims paid on a portfolio of general insurance policies. Claims from this
portfolio are fully run off after 3 years

Underwriting Year Development Year


0 1 2 3
2020 85 42 30 7
2021 103 65 25
2022 93 47
2023 111

Estimate the outstanding claims using the basic chain ladder approach.
2. Apply the chain ladder method to the triangle of loss ratios shown below. Estimate the ultimate
loss ratio in each year, and the amount that needs to be set aside to meet future claims. The
amounts of premium received in each of the years 2017 to 2020 were K2.84m, K3.28m, K3.46m
and K3.64m.
Loss ratio

Accident Year Development Year


0 1 2 3
2017 47% 63% 70% 74%
2018 48% 62% 71%
2019 49% 60%
2020 50%

3. The table below shows the incremental claims paid on a portfolio of insurance policies together
with an extract from an index of prices at a certain insurance company in Lusaka. Claims are fully
paid by the end of development year 3.

Accident Year Development Year


Year Price Index(mid year)
0 1 2 3
2017 100
2017 103 32 29 13
2018 104
2018 88 21 16
2019 109
2019 110 35
2020 111
2020 132

87
Calculate the reserve for unpaid claims using the inflation-adjusted chain ladder approach, assum-
ing that future claims inflation will be 3% per annum.

4. The following table shows incremental claims data from a portfolio of insurance policies for the
accident years 2010, 2011 and 2012. Claims from this type of policy are fully run off after the end
of development year two

Incremental Claims

Accident Year Development Year


0 1 2
2010 2328 1484 384
2011 1749 1188
2012 2117

Estimate the total claims outstanding using the basic chain ladder technique.

5. The table below shows claims paid on a portfolio of general insurance policies. You may assume
that claims are fully run off after three years.

Underwriting Year Development Year


0 1 2 3
2020 450 312 117 41
2021 503 389 162
2022 611 438
2023 555

Use the inflation adjusted chain ladder method to calculate the outstanding claims on the portfolio.
Given that the past claims inflation has been 5% p.a. However, it is expected that future claims
inflation will be 10% p.a.

6. The table below shows cumulative claims (not adjusted for inflation) from a portfolio of insurance
policies for 4 accident years.

Underwriting Year Development Year


0 1 2 3
2021 2047 3141 3209 3310
2022 2471 3712 3810
2023 2388 3750
2024 2580

It may be assumed that payments are made in the middle of a calendar year. It is estimated that
the inflation rate applicable to these data has been 5% per annum over the relevant period. Use the
inflation adjusted chain ladder method to estimate the total outstanding payments, up to the end of
Development Year 3, for Accident Year 2024 in mid-2024 prices.

7. An actuary using this model has estimated the parameters for a run-off triangle as follows:

s2015 = K39.0m, r0 = 0.6, x2015 = 1.00

s2016 = K45.5m, r1 = 0.3, x2016 = 1.10


s2017 = K41.6m, r2 = 0.1, x2017 = 1.20

88
x2018 = 1.25
x2019 = 1.30
Use these estimates (ignoring error terms) to construct the complete table of incremental claim
amounts for Accident Years 2015-2017, and hence estimate the amount of outstanding claims at
the end of 2017.

89
References
1. ActEd, (2019). CS2 Combined Pack, London: The Actuarial Education Company.

2. An introduction to statistical modelling. - Dobson, Annette J. - Chapman & Hall, 1983. - Viii, 125
pages. - ISBN: 0 412 24860 3.

3. Introductory statistics with applications in general insurance. - Hossack, Ian B; Pollard, John H;
Zehnwirth, Benjamin. - 2nd ed. - Cambridge University Press, 1999.- xi, 282 pages. - ISBN: 0 521
65534 X.

4. Loss models: from data to decisions. - Klugman, Stuart A; Panjer, Harry H; Willmot, Gordon E;
Venter, Gary G. - John Wiley & Sons, 1998. - Xiii, 644 pages. - ISBN: 0471 23884 8.

5. Practical risk theory for actuaries. - Daykin, Chris D; Pentikainen, Teivo; Pesonen, Martti. - Chap-
man & Hall, 1994. - 545 pages. - ISBN: 0 412 42850 4.

Recommended Texts
ActEd, (2019). CS2 Combined Pack, London: The Actuarial Education Company.

90

You might also like