Quantitative Methods Notes
Quantitative Methods Notes
Offered by AnalystPrep
1
©2024 AnalystPrep “This document is protected by International copyright laws. Reproduction and/or distribution of this document is
2
© 2014-2024 AnalystPrep.
Learning Module 1: Rate and Return
The time value of money is a concept that states that cash received today is more valuable than
cash received in the future. If a person agrees to receive payment in the future, he foregoes the
An interest rate or yield, usually denoted by r, is a rate of return that reflects the connection
Assume you currently possess $100. Next, consider depositing this money into a savings account,
expecting it to grow to $110 after one year. Intuitively, the compensation required for deferring
the consumption of $100 now in favor of receiving $110 in one year is $10 (equal to 110 minus
100). This compensation is equivalent to a 10% rate of return (calculated as 10 divided by 100).
1. Required rate of return: The minimum return an investor expects to earn to accept an
investment.
2. Discount rate: The rate used to discount future cash flows to allow for the time value of
money (to determine the present value equivalent of some money to be received
sometime in the future). Discount rates and interest rates are used almost
interchangeably.
3. Opportunity cost: The value of the best-forgone alternative; the most valuable
Economics postulates that the forces of supply and demand determine interest rates. In this
case, the investors (lenders) supply the money, and the borrowers demand money for their
3
© 2014-2024 AnalystPrep.
consumption.
As such, interest is a borrower's reward for using an asset, usually capital, belonging to a lender.
It is compensation for the loss or value depreciation occasioned by the use of the asset.
Therefore, an interest rate is composed of a real risk-free interest rate plus a set of four
The real risk-free interest rate is the single-period interest rate for a completely risk-free security
if no inflation is expected. According to economic theory, the real risk-free rate reflects people’s
Inflation risk is the loss of money's purchasing power caused by increased consumer goods
prices.
The inflation premium compensates investors for expected inflation. It represents the average
inflation rate expected over the debt's maturity. The risk of a decrease in purchasing power
Liquidity refers to the ease with which an investment can be converted into cash without
4
© 2014-2024 AnalystPrep.
The liquidity premium compensates investors for the risk of loss relative to an investment’s fair
Default risk describes a situation in which a borrower may fail to repay borrowed funds due to
The default risk premium compensates investors for the possibility that the borrower will fail
to make a promised payment at the contracted time and in the contracted amount.
The maturity risk premium is the additional return an investor requires for assuming interest
rate and reinvestment risk resulting from a longer investment maturity timeline. Maturity risk
premium increases with an increase in the maturity timeline. In other words, the longer the
The nominal risk-free interest rate is defined as the sum of the real risk-free interest rate and the
inflation premium. In other words, the nominal risk-free interest rate can be seen as the
combination of the real risk-free rate plus an inflation premium, as shown by the following
equation:
Most rates quoted on short-term government debts can be taken as nominal risk-free interest
5
© 2014-2024 AnalystPrep.
Question
investing?
A. Discount rate.
B. Opportunity cost.
Solution
Opportunity cost is a key factor in interpreting interest rates. It refers to the interest
foregone when investors opt for an alternate option, such as spending on current
A is incorrect. The discount rate is the interest rate used to discount future cash
C is incorrect. The required rate of return is the minimum rate of return an investor
6
© 2014-2024 AnalystPrep.
LOS 1b: calculate and interpret different approaches to return
measurement over time and describe their appropriate uses
Financial assets are primarily defined based on their return-risk characteristics. This definition
approach helps when building a portfolio from all the assets available. It's noteworthy that there
Financial market assets generate two types of returns: Income from cash dividends or interest
payments and capital gains or losses from changes in the prices of financial assets.
Some financial assets give only one stream of return. For instance, headline stock market indices
typically only report on price appreciation. They do not include the dividend income unless the
A holding period return is earned from holding an asset for a specified period, such as a day,
(P 1 − P0 ) + I1
R=
P0
An investor purchased 100 shares of a stock at $50 per share and held the investment for one
year. During that period, the stock paid dividends of $2 per share. At the end of the year, the
7
© 2014-2024 AnalystPrep.
The holding period return is closest to:
Solution
Therefore,
Holding period returns can also be calculated for periods longer than a year. For instance, if we
need to calculate the holding period return for a five-year period, we should compound the five
(P5 − P 0 ) + I(1−5)
R=
P0
Arithmetic Return
When we have assets with multiple holding periods, we must aggregate the returns into one
overall return.
Denoted by R̄ i arithmetic mean for an asset i is a simple process of finding the average holding
Where:
R it = Return of asset i in period t.
8
© 2014-2024 AnalystPrep.
T = Total number of periods.
1 T 1
R̄ i = ∑ R it = (15% + 10% + 12% + 3%) = 10%
T t=1 4
Geometric Return
Computing a geometric mean follows a principle similar to the one used to compute compound
interest. It involves compounding returns from the previous year to the initial investment's value
at the start of the new period, allowing you to earn returns on your returns.
A geometric return provides a more accurate representation of the portfolio value growth than
an arithmetic return.
T
= ⎷T ∏ (1 + R t) − 1
t=1
Using the same annual returns of 15%, 10%, 12%, and 3% as shown above, we compute the
1
Geometric mean = [(1 + 15%) × (1 + 10%) × (1 + 12%) × (1 + 3%)] 4 − 1
= 9.9%
Note that the geometric return is slightly less than the arithmetic return. Arithmetic returns tend
to be biased upwards unless the holding period returns are all equal.
Harmonic Mean
9
© 2014-2024 AnalystPrep.
The harmonic mean measures central tendency. It's especially useful for rates or ratios such as
P/E ratios. The harmonic mean's formula is derived from the harmonic series, a specific
mathematical sequence.
n
X̄ H = , Xi > 0 for all i = 1, 2, … , n
∑ni=1 X1
i
The harmonic mean is handy for averaging ratios when those ratios are consistently applied to a
fixed quantity, resulting in varying unit numbers. For instance, it's applied in cost-averaging
An investor practices cost averaging by investing in a particular stock over a three-month period.
The investor decides to allocate different amounts of money each month. In the first month, the
investor invests $2,000; in the second month, $3,000; and in the third month, $4,000. The share
prices of the stock for these three months are $10, $12, and $15, respectively.
Calculate the average price paid per share for the three-month period.
Solution
n 3
X̄ H = = = 12
∑ ni=1 X1 1 1 1
+ +
i 10 12 15
Trimmed and Winsorized means seek to lower the effect of outliers in a data set.
Trimmed Mean
10
© 2014-2024 AnalystPrep.
The trimmed mean is a measure of central tendency. We calculate it after excluding a small
For example, a data set consists of 10 observations: 12, 15, 18, 20, 22, 25, 27, 30, 35, and 40. We
can calculate the trimmed mean after removing the highest and lowest values.
After removing these values, the remaining data set is 15, 18, 20, 22, 25, 27, 30, and 35.
Now, let’s calculate the trimmed mean by taking the average of these remaining values:
15 + 18 + 20 + 22 + 25 + 27 + 30 + 35192
= = 24
8 8
Winsorized Mean
The Winsorized mean is a central tendency measure. It works by replacing extreme values at
both ends of the data with the values of their closest observations. This process is similar to the
11
© 2014-2024 AnalystPrep.
Question 2
What are the arithmetic mean and geometric mean, respectively, of an investment
that returns 8%, -2%, and 6% each year for three years?
Solution
8% + (−2%) + 6%
Arithmetic mean = = 4%
3
12
© 2014-2024 AnalystPrep.
LOS 1c: Compare the money-weighted and time-weighted rates of return
and evaluate the performance of portfolios based on these measures
The money-weighted return considers the money invested and gives the investor information on
The money-weighted rate of return (MWRR) is like the portfolio's internal rate of return (IRR).
It's the rate at which the present value of cash flows equals zero. It's a way to measure how well
a portfolio performs.
T
CFt
∑ =0
t=0 (1 + I RR)t
Where:
T = Number of periods.
The money-weighted rate of return (MWRR) looks at a fund’s starting and ending values and all
the cash flows in between. In an investment portfolio, cash inflows are a part of it. These inflows
could be from deposits or investments made during a certain period. The MWRR considers these
inflows and calculates the overall rate of return for the portfolio:
Dividends/interest reinvested.
Contributions made.
13
© 2014-2024 AnalystPrep.
Withdrawals made.
At the end of the first year, after the portfolio's value increases to $12,000, the investor
At the end of the second year, the portfolio value further increases to $25,000.
The money-weighted rate of return for the investor’s portfolio is closest to:
Solution
We need to calculate the internal rate of return (IRR) considering the following cash flows:
CF2 = +$25, 000 (Final portfolio value at the end of year two).
To find the money-weighted rate of return, solve the equation for IRR:
14
© 2014-2024 AnalystPrep.
Calvin Hair purchased a share of Superior Car Rental Company for $85 at the beginning of the
first year. He bought an additional unit for $87 at the end of the first year. At the end of the
second year, he sold both shares at $90. During both years, Hair received a dividend of $4 per
Solution
To calculate the money-weighted return in this example, we need to consider the timing and
Number of shares sold × Selling price = 2 shares × $90 = $180 + 8( Dividend received for
both shares)
= $188
As such, we have:
CF0 = −85 .
CF1 = −83 .
CF2 = 188.
Using the BA II Plus calculator, you will get IRR = 7.71% , equivalent to the money-weighted rate
of return.
15
© 2014-2024 AnalystPrep.
Shortcomings of the Money-weighted Rate of Return
The money-weighted rate of return (MWRR) considers all cash flows, such as withdrawals or
contributions. If an investment spans multiple periods, MWRR gives more importance to the
fund's performance when the account is at its largest. This can be a problem for fund managers
because it might make their performance seem worse due to factors they can't control.
The time-weighted rate of return (TWRR) calculates an investment's compound growth. Unlike
the money-weighted rate, it doesn't care about withdrawals or contributions. TWRR is like
finding the average return of different time periods within your investment.
Step 1: Value the portfolio immediately before any significant cash inflow or outflow of funds.
Divide the evaluation period into subperiods based on dates of significant additions or
withdrawals of funds.
Step 2: Compute the holding period return on the portfolio for each period.
Step 3: Compound or link the holding period returns to the annual rate of return, which is the
TW RR = (1 + HP R 1 × (1 + H P R2 ) × (1 + HP R3 ) … × (1 + HP Rn −1 ) × (1 + H P Rn ) – 1
If the evaluation period is more than one year, compute the geometric mean of the annual
16
© 2014-2024 AnalystPrep.
An investor purchases a share of stock at t = 0 for $200. At the end of the year (at t = 1), the
investor purchases an additional share of the same stock, this time for $220. She then sells both
shares at the end of the second year for $230 each. She also received annual dividends of $3 per
share at the end of each year. Calculate the annual time-weighted rate of return on her
investment.
Solution
First, we break down the two years into two one-year periods.
Holding period 1:
Dividends paid = 3.
Holding period 2:
(220 − 200 + 3)
H P R1 = = 11.5%
200
(460 − 440 + 6)
H P R2 = = 5.9%
440
Lastly, we need to find the geometric mean of the HPRs since we are dealing with a period of
1
T WRR = [(1 + HPR 1 ) × (1 + HPR 1 ) … × (1 + HPRn )] n − 1
0.5
= (1.115 × 1.059) − 1 = 8.7%
17
© 2014-2024 AnalystPrep.
than One Year)
The beginning value of a portfolio as of January 1, 2020, was $1,000,000. On February 10, the
portfolio’s value was $1,100,000, including an additional contribution of the $50,000 injected
into the portfolio on this date. The portfolio’s ending value at the beginning of April was
$1,350,000.
Solution
18
© 2014-2024 AnalystPrep.
Question
an extra share of the same stock for $53. The share gives a dividend of $0.50 per
share for the first year and $0.60 per share for the second year. He sells the shares at
the end of the second year for $55 per share. Calculate the annual time-weighted rate
of return.
A. 5.90%.
B.12.24%.
C. 7.00%.
HP 1 HP 2
P 0 = 50 P 0 = 106
D = 0.5 D = 1.2
P 1 = 53 P 1 = 110
(53 − 50 + 0.5)
HP R 1 = = 7%
50
(110 − 106 + 1.2)
HP R 2 = = 4.9%
106
⇒ T WRR = 1.07 × 1.049 − 1 = 12.24%
Therefore,
19
© 2014-2024 AnalystPrep.
LOS 1d: Calculate and interpret annualized return measures and
continuously compounded returns and describe their appropriate uses
To compare returns over different timeframes, we need to annualize them. This means
Non-Annual Compounding
Interest may be paid semiannually, quarterly, monthly, or even daily – interest payments can be
made more than once a year. Consequently, the present value formula can be expressed as
R S −m N
P V = FVN (1 + )
m
Where:
N = Number of years.
Example: Calculating the Present Value of a Lump Sum (More than One
Compounding Period)
Jane Doe wants to invest money today and have it become $500,000 in five years. The annual
interest rate is 8%, and it's compounded quarterly. How much should Jane invest right now?
F V N = $500, 000 .
R S = 8% .
m = 4.
20
© 2014-2024 AnalystPrep.
8%
R s /m = = 2% = 0.02 .
4
N = 5.
mN = 4 × 5 = 20 .
Therefore,
RS −m N −20
P V = F V N (1 + ) = $500, 000 × (1.02) = $336, 485.67
m
Press the [2nd] button, then the [FV] button to clear the financial registers. The display
Enter the future value (FV). This is the amount Jane wants to have in five years, which
Enter the interest rate (I/Y). This is the annual interest rate, which is 8%. However,
since interest is compounded quarterly, we need to divide this by 4. To do this, type “8”,
press the [÷] button, type “4”, then press the [ENTER] button, and finally press the
[I/Y] button.
Enter the number of periods (N). This is the number of quarters in five years, which is
5*4 = 20. To do this, type “20” and press the [N] button.
Compute the present value (PV). To do this, press the [CPT] and then the [PV] buttons.
The display should show the amount Jane needs to invest today, approximately
$336,485.49.
Annualized Returns
To annualize a return for a period shorter than a year, you need to account for how many times
that period fits into a year. For example, if you have a weekly return, you would compound it 52
21
© 2014-2024 AnalystPrep.
Generally, we can annualize the returns using the following formula:
Returnannual = (1 + Returnperiod )c − 1
Where:
If the monthly return is 0.7%, then the compound annual return is:
12
Returnannual = (1 + Returnmonthly) − 1
12
= (1.007) − 1 = 0.0873 = 8.73%
For a period of more than one year, for example, a 15-month return of 16% can be annualized as:
12
Returnannual = (1 + Return 15 month ) 15 − 1
4
= (1.16)5 − 1 = 12.61%
We may apply the same procedure to convert weekly returns to annual returns for comparison
52
Returnannual = (1 + Returnweekly ) −1
For comparison with weekly returns, we can convert annual returns to weekly returns by making
An investor is evaluating the returns of two recently formed bonds. Selected return information
22
© 2014-2024 AnalystPrep.
Bond Time Since Issuance Return Since Issuance (%)
A 120 days 2.50
B 8 months 6.00
To compare the annualized rate of return for both bonds, you can use the formula for annualizing
365
Return Since Issuance
Annualized Return = (1 + ) Time Since Issuance
−1
100
For Bond A:
365
2.50 120
Annualized Return for Bond A = (1 + ) − 1.
100
For Bond B:
365
6.00 240
Annualized Return for Bond B = (1 + ) − 1.
100
23
© 2014-2024 AnalystPrep.
Bond B has an annualized return of approximately 9.28%.
The continuously compounded return is calculated by taking the natural logarithm of one plus
the holding period return. For example, if the monthly return is 1.2%, you'd calculate it as
P t+1
rt,t+1 = ln ( ) = ln (1 + Rt,t+1 )
Pt
Assume now that the investment horizon is from time t = 0 to time t = T then the continuously
PT
r0 ,T = ln ( )
P0
If we apply the exponential function on both sides of the equation, we have the following:
PT = P 0 er0,T
PT
Note that can be written as:
P0
PT PT P T−1 P1
=( )( ) …( )
P0 P T −1 P T−2 P0
PT PT P T −1 P
ln ( ) = ln ( ) + ln ( ) + … + ln ( 1 )
P0 P T−1 P T −2 P0
⇒ r0 ,T = rT−1, T + rT −2,T −1 + … + r0, 1
24
© 2014-2024 AnalystPrep.
Therefore, the continuously compounded return to time T is equivalent to the sum of one-period
25
© 2014-2024 AnalystPrep.
Question
A. 0.40%.
B. 0.92%.
C. 0.41%.
Recall that:
52
Returnannual = (1 + Returnweekly ) −1
1
Returnweekly = (1 + Return annual) 52 − 1
1
= (1 + 0.23)52 − 1
≈ 0.40%
26
© 2014-2024 AnalystPrep.
LOS 1e: calculate and interpret major return measures and describe
their appropriate uses
The gross return is what an asset manager earns before subtracting various costs such as
management fees, custody fees, taxes, and other administrative expenses. However, it does
Gross return does not consider management or administrative costs. For this reason, it is a
suitable metric for assessing and comparing the investment expertise of asset managers.
Net return is a metric for how much an investment has earned for the investor. It considers all
Unless otherwise stated, all returns are nominal pre-tax returns in general. Depending on the
jurisdiction, different rates apply to capital gains and income. Long-term and short-term taxes
The after-tax nominal return is determined by subtracting any tax deductions applied to
Real Returns
Returns are typically presented in nominal terms, which consist of three components: the real
risk-free return as compensation for postponing consumption, inflation as compensation for the
loss of purchasing power, and a risk premium. Real returns are useful in comparing returns over
Recall the relationship between the nominal rate and the real rate:
27
© 2014-2024 AnalystPrep.
(1+ Nominal Risk-free rate)=(1+Real risk free rate)(1+Inflation premium)
We can find the connection between nominal and real returns by considering the real risk-free
rate of return and the inflation premium. This relationship can be expressed as:
Real returns become particularly useful when you want to compare returns across various time
periods and different countries. This is especially important when returns are shown in local
currencies and when inflation rates vary from one country to another.
After-tax real return is the amount the investor receives as payment for delaying consumption
Leveraged Returns
If an investor uses derivative instruments within a portfolio or borrows money to invest, then
leverage is introduced into the portfolio. The leverage amplifies the returns on the investor's
The leveraged return considers the actual return on the investment and the cost of the borrowed
money. The cost of borrowing and financing fees are subtracted from the overall return produced
Using the borrowed capital (debt) increases the size of the leveraged position by the additional
borrowed capital.
Portfolio return
RL =
Portfolio equity
[R P × (VE + VB ) − (V B × rD )]
=
VE
VB
= RP + (R P − rD )
VE
28
© 2014-2024 AnalystPrep.
Where:
For a $250,000 equity portfolio with an annual 9% total investment return, 40% financed by debt
VB $100,000
RL = R P + (R P − rD ) = 9% + (9% − 6%) = 11%
VE $150,000
29
© 2014-2024 AnalystPrep.
Question
the equity portfolio generates a 9% annual total investment return, the leverage
A. 11.15%.
B. 14.00%.
C. 8. 25%.
VB
R L = RP + (R P − rD )
VE
$2, 625, 000
= 9% + (9% − 5%) = 11.15%
$4, 875, 000
30
© 2014-2024 AnalystPrep.
Learning Module 2: The Time Value of Money in Finance
LOS 2a: calculate and interpret the present value (PV) of fixed-income
and equity instruments based on expected future cash flows
The time value of money (TVM) is a fundamental financial concept. It emphasizes that a sum of
money is worth more in the present than in the future. There are three key reasons supporting
this principle:
The concept of opportunity cost suggests that money available today can be invested
and generate interest, increasing its value over time. By delaying the use of money, one
Inflation poses a threat to the purchasing power of money in the future. Due to
inflation, the same amount of money may buy fewer goods or services in the future
compared to its present value. Consequently, having money now is advantageous since
reliable. Until the money is obtained, there is a level of uncertainty attached to its
Time value of money calculations allow us to establish a given amount's future value.
Discount rate or interest rate: The rate of discounting or compounding that you
Time periods: The whole number of time periods over which a sum's present or future
weekly, etc.
Present value (PV): The amount of money you have today (or at time T = 0) is
31
© 2014-2024 AnalystPrep.
referred to as the present value.
Future value (FV): The accumulated amount of money you get after investing the
original sum at a specific interest rate and for a given time period, say, two years.
Let,
F V = Future value.
P V = Present value.
F V =P V (1 + r)N
F V = P V erN
To find the present value of the investment, we rewrite the above formula so that:
PV = F V (1 + r)−N
P V = F Vte−r N
A fund continuously accumulates to $4,000 over ten years at a 10% annual interest rate.
32
© 2014-2024 AnalystPrep.
Solution
So,
Frequency of Compounding
When the frequency of compounding is more than once per year (quarterly, monthly, etc.), the
rs mN
F V N = P V (1 + )
m
Where:
N = Number of years.
r s −m N
P V = F V (1 + )
m
In the following discussion, we shall let t = mN denote the number of compounding periods and
rs
m
= r denote the stated discount rate per period.
For calculating F V and PV using the BA II PlusTM Financial Calculator, use the following keys:
33
© 2014-2024 AnalystPrep.
P V = Present value.
F V = Future value.
P MT = Payment.
CP T = Compute.
It is important to note that the sign of P V and F V will be opposite. For example, if PV is
negative, then F V will be positive. Generally, an inflow is entered with a positive sign, while an
Fixed-income instruments are debt securities where an issuer borrows money from an investor
(lender) in exchange for a promised future payment. Examples of fixed-income instruments are
The market discount rate for fixed-income instruments, also known as yield-to-maturity (YTM), is
Cash flows in fixed-income instruments follow three general patterns: discount, periodic interest,
For discount cashflow patterns, an investor pays an initial discounted price (P V ) for the
instrument (such as a bond or a loan) and gets one payment (F V ) at the end maturity. The
investor's return is the interest earned, that is, the difference between the initial price and
principle (F V − P V ).
Discount bonds are also called zero-coupon bonds, and they do not have periodic interest
payments.
34
© 2014-2024 AnalystPrep.
The price of a discount bond can be calculated using the formula for the present value (P V ) of a
P V = F Vt(1 + r)−t
Where:
F V = Future value.
P V = Present value.
Assume Chad invests $8,000 in a zero-coupon bond that yields 8% annually and matures in four
35
© 2014-2024 AnalystPrep.
Solution
Recall that:
F V = PV (1 + r)t
Note that zero-coupon bonds can be issued at negative interest rates. In this case, the price (PV)
Example: Calculating the Price of a Discount Bond Issued at Negative Interest Rates
In January 2018, the Swiss government issued 15-year sovereign bonds at a negative yield of
-0.08%. The present value (PV) of the bond per CHF100 of principal (FV) at the time of issuance
is closet to:
Solution
PV = F V t(1 + r)−t
= 100(1 − 0.0008)−15 = 101.21
36
© 2014-2024 AnalystPrep.
A coupon instrument is a fixed-income investment. It includes periodic cash flows called coupons
and repays the principal at maturity. People often use these in coupon bond investments. These
The pricing of a coupon bond involves calculating its present value (PV) based on the market
discount rate. The general formula for calculating the bond's price is derived from the
discounted cash flow model. It considers the coupon payments (PMTs) and the final principal
payment (FV) at maturity. The bond's price is determined by discounting each cash flow using
The formula used to calculate the present value (PV) of a coupon bond is as follows:
P MT PMT (P MT N + F VN )
PV(Coupon Bond) = + +. . . +
(1 + r) 1 (1 + r) 2 (1 + r)N
Where:
P MT = Coupon payment.
37
© 2014-2024 AnalystPrep.
F V = Future value.
N = Number of periods.
Suppose we have a 5-year bond with a face value of $1,000 and an annual coupon rate of 5%.
The market discount rate is 6%. The bond’s price is closest to:
Solution
P MT P MT (PMT N + F VN )
PV(Coupon Bond) = + +⋯+
(1 + r)1 (1 + r) 2 (1 + r)N
Example 2: Pricing a Coupon Bond With a Single Cash Flow on a semi-annual Basis
Assume an investor has a 2-year bond with a face value of $1000 and an annual coupon rate of
38
© 2014-2024 AnalystPrep.
6%, paid semi-annually. The market discount rate is 5%. The price of the bond is closest to:
Solution
Recall that:
P MT PMT (P MT N + F VN )
PV(Coupon Bond) = + +. . . +
(1 + r)1 (1 + r)2 (1 + r)N
Where:
6%
P MT = Coupon payments ($1, 000 × ) = $30 in this case.
2
(5%)
r = Market discount rate (YTM), ( 2
= 2.5%) per period in this case.
You can easily use the BA II Plus calculator (or any other allowed financial calculator) to solve
Perpetual Bonds
39
© 2014-2024 AnalystPrep.
Perpetual bonds are rare types of coupon bonds that do not have a stated date of maturity. They
are generally issued by firms seeking equity-like financing and usually include redemption
provisions.
The formula present value of perpetual bonds is obtained as follows: As N → ∞, the formula for
PV (perpetual bond)
lim PMT P MT (P MT N + F VN )
= [ + +⋯ + ]
(N → ∞) (1 + r) 1 (1 + r) 2 (1 + r)N
P MT
=
r
P MT
PV =
r
In 2021, XYZ Financial (the holding company for XYZ Bank) issued $500 million in perpetual
bonds with a 4.00 percent semi-annual coupon. Calculate the bond's yield to maturity (YTM) if
Solution
Recall,
P MT
PV =
r
Hence,
P MT
r=
PV
To solve this problem, we first need to calculate the semi-annual coupon payment, which is,
$100 × 4%
40
© 2014-2024 AnalystPrep.
$100 × 4%
PMT(semi-annual coupon payment) = = $2 , PV = $98.50
2
Therefore,
$2
r= = 0.0203 = 2.03%
$98.50
r = 0.0203 × 2 ≈ 4.06%
An annuity is a finite series of cash flows, all with the same value. A fixed-income instrument
with annuity payments provides a stream of periodic equal cash inflows over a finite period.
Level payments consist of interest and principal. Fixed income instruments with level payments
There are two types of annuities: ordinary annuities and annuities due. Annuity due is a type of
annuity where payments start immediately at the beginning of time, at time t = 0. In other
On the other hand, an ordinary annuity is an annuity where the cashflows occur at the end of
each period. Such payments are said to be made in arrears (beginning at time t = 1). We shall
Ordinary Annuity
Remember that in an ordinary annuity, the series of payments does not begin immediately.
Instead, payments are made at the end of each period. It is further worth noting that the present
value of an annuity is equal to the sum of the current value of each annuity payment:
41
© 2014-2024 AnalystPrep.
Where:
r(P V )
A=
(1 − (1 + r)−N
Consider a fully amortizing mortgage loan. In this case, the borrower receives the mortgage loan
now and promises to make periodic payments equal to the sum of interest and principal
payments.
42
© 2014-2024 AnalystPrep.
Note that the periodic mortgage payment is constant, but the proportion of the interest payment
The cash flow pattern of a fully amortizing mortgage follows the pattern of an ordinary annuity
with a series of equal cash flows. As such, the periodic annuity (periodic payment) of a fully
r(P V )
A=
1 − (1 + r)−t
Where:
43
© 2014-2024 AnalystPrep.
A = Periodic cash flow.
Jake is looking to secure a fixed-rate 25-year mortgage to finance 75% of the value of an
$800,000 residential property. If the annual interest rate on the mortgage is 4.5%, Jake's monthly
Solution
Remember,
r(P V )
A=
1 − (1 + r)−t
Where:
4.5%
r = 0.375% (= )
12
the issuing company. This gives investors the right to receive a share of the company's available
In the context of equity instruments, the time value of money (TVM) is used to discount expected
future cash flows to determine their present value. This allows investors to value the company
shares.
The present value of expected future cash flows is calculated using a discount rate, r, which
Valuing equity investments depends on dividend cashflows, which can take one of three forms:
constant dividends, constant dividend growth rate, or changing dividend growth rate.
1. Valuing Equity Instruments based on Constant Dividend: The Constant Dividends model
values stocks based on the assumption that dividends will remain constant over time. The
preferred or common share dividend cash flows are in the form of an infinite series that is valued
like perpetuity. The formula for the constant dividends model is as follows:
45
© 2014-2024 AnalystPrep.
∞
Dt Dt
P Vt = ∑ =
i=1 (1 + r) i r
r = Discount rate.
Assuming we have a preferred stock with a dividend payment of $5 per year. The discount rate is
Solution
Recall,
Dt
P Vt =
r
So,
5
PV = = $62.5
0.08
2. Valuing Equity Instruments Based on Constant Dividend Growth Rate: The constant
dividend growth model is a method used to estimate the value of a stock based on its future
dividends. This model assumes that dividends will grow at a constant rate (g) forever. To derive
the formula for this model, we start by considering that the present value of a stock is equal
to the sum of its future dividends, discounted by the required rate of return r . If dividends
are assumed to grow at a constant rate, then each future dividend can be calculated by
46
© 2014-2024 AnalystPrep.
Let Dt represent the expected dividend in the next period. The present value of the stock can
(1+r )
This is an infinite geometric series with a common ratio of . Using the formula for the sum of
(1+g)
Dt(1 + g) Dt+1
P Vt = =
r −g r−g
Where:
r−g>0
Therefore, this is the formula for calculating the present value of a stock using the constant
dividend growth rate. This model can help estimate the value of a stock when its future dividends
Suppose a stock currently pays an annual dividend of $2.00 per share. The required rate of
return for this stock is 10%, and the dividends are expected to grow at a constant rate of 5% per
year indefinitely. Using the constant dividend growth model, the present value of this stock is
closest to:
Solution
47
© 2014-2024 AnalystPrep.
Recall that,
Dt(1 + g) Dt+1
P Vt = =
r −g r−g
So,
2 × 1.05 2.10
PV = = = $42
r −g 0.10 − 0.05
process. It begins with the investor buying a stock at an initial price and getting an initial
dividend. The unique aspect is that the dividend is expected to grow at a rate that evolves as the
company matures and shifts from high growth to slower growth. This valuation doesn't have a
single formula because it relies on assumptions about future dividend growth. However, a
common method is to use a multi-stage dividend discount model. This model assumes that
dividends will grow at different rates during various stages of the company's growth. To find the
stock's present value, you sum up the present values of dividends at each stage.
The Multi-Stage Dividend Discount Model builds on the Constant Dividend Growth Model. It
accommodates a company's transition from high initial growth to lower, more stable growth.
Let's say a company has a high short-term growth rate gs followed by a perpetual lower growth
rate gl . To find the present value (PV) of the stock at time t using this model, we compute it in
two stages:
I. First Part: The first part calculates the present value of dividends during the initial n
periods of higher growth (g s ). This is done by discounting the dividends for each period
n Dt (1 + gs )i
P Vt = ∑
i =1 (1 + r)i
48
© 2014-2024 AnalystPrep.
Where: PV = Present value. n = Number of [Link] = Dividend at time (t). gs = Initial
II. Second Part: The second part calculates the present value of dividends after the initial n
periods, assuming constant growth at a lower long-term rate (gl ). This can be simplified
using the geometric series simplification, where E(St + n) represents the terminal value
E(St + n)
P Vt =
(1 + r)n
D t+ n+1
Where: E(St + n) = and g l is the lower, more stable dividend growth rate.
( r−g l)
Assuming we have a stock with an expected dividend payment of $2 in one period, and the
discount rate is 10%. The stock is expected to have a high dividend growth rate of 20% for the
first three years, followed by a slower growth rate of 5% thereafter. Calculate the present value
of the stock.
Solution
First, we calculate the present value of the dividends during the high growth period:
Recall that,
n Dt (1 + gs )i
P Vt = ∑
i =1 (1 + r)i
So,
2 2 × (1 + 0.20)1 2 × (1 + 0.20)2
P V1 = + +
(1 + 0.10)1 1 + 0.10)2 1 + 0.10)3
P V1 = 1.818 + 1.983 + 2.163 = 5.965
P V1 = $5.97
49
© 2014-2024 AnalystPrep.
Next, we calculate the present value of the dividends during the slower growth period, assuming
Recall that,
Dt+n+1
E(St + n) =
(r − gl )
So,
$72.578
P V2 = = $54.527
(1 + 0.10)3
P Vtotal = P V1 + P V2
= $5.965 + 54.527
= $60.493 ≈ $60.49
50
© 2014-2024 AnalystPrep.
Question
Five years ago, Milton Inc. issued corporate bonds with a 15-year maturity. The bonds
have a semi-annual coupon rate of 7.8% per annum, and the current yield to maturity
is 8.5% per annum. The current price of Milton Inc's bonds (per CAD100 of par value)
is closest to:
A. CAD91.23.
B. CAD95.35.
C. CAD96.15.
Solution
P MT PMT (P MT N + F VN )
PV(Coupon Bond) = + +. . . +
(1 + r)1 (1 + r)2 (1 + r)N
7.8%
The semi-annual coupon rate is 2 = 3.9%.
8.5%
The semi-annual yield to maturity is 2 = 4.25%.
Next, we find the number of periods remaining until the bond matures:
Since the bonds were issued 5 years ago and have a 15-year maturity, 10(= 15 − 5)
10 × 2 = 20 periods.
You can plug the above values into the general formula, consuming valuable time.
51
© 2014-2024 AnalystPrep.
Steps Explanation Display
[2nd][QUIT] Return to standard calc Mode 0
[2nd][CLR TVM] Clears TVM Worksheet 0
20[N ] Years/periods N = 20
4.25[1/Y ] Set interest rate I /Y = 4.25
3.9[P MT ] Set the periodic P MT = 3.90
coupon payment
100[P V ] Set the face F V = 100.00
value of the bond
[CP T ][P V ] Compute the present value PV = −95.35
52
© 2014-2024 AnalystPrep.
LOS 2b: Calculate and interpret the implied return of fixed-income
instruments and the required return and implied growth of equity
instruments given the present value (PV) and cash flows
The growth rate is the rate at which the market expects an asset to grow. On the other hand,
implied return reflects a return based on the current price and future security cash flows.
Consider a fixed-income instrument. If we have its present value and assume all future cash
flows happen as expected, the discount rate, rrr, or yield-to-maturity, YTM, shows the implied
Now, take an equity investment. If we have the present value, future value, and discount rate, we
can find the implied growth rate that aligns with these values.
The implied return or growth rate provides a view of the market expectations incorporated into
an asset's market price. Understanding these expectations is critical for investors when making
investment decisions.
Discount Bond
In the case of a discount bond or instrument, recall that an investor receives a single principal
To solve for the implied return earned over the life of an instrument (N periods), we can
Recall that the single cash flow present value formula is:
P V = F Vt(1 + r)−t
53
© 2014-2024 AnalystPrep.
Where:
F V = Future value.
P V = Present value.
1
F Vt F Vt t
r=√ t −1 = ( ) −1
PV PV
We use this formula to calculate the periodic return earned during the instrument's life (t
periods) based on its present value (or price) and future value.
Consider a zero-coupon bond with price of $900, a future value of $1,000, and a maturity of 5
Solution
Recall that,
1
F Vt t
r=( ) −1
PV
So,
1
1000 5
r=( ) − 1 = 2.13%
900
This means that an investor who purchases this zero-coupon bond at a price of $900 and holds it
54
© 2014-2024 AnalystPrep.
Coupon Bonds
Recall that fixed-income instruments that pay periodic interest have cash flows throughout their
life until maturity. The yield-to-maturity (YTM) is a single implied market discount rate for all
cash flows, regardless of timing. It assumes an investor expects to receive all promised cash
flows through maturity and reinvest any cash received at the same YTM.
The present value of a fixed-income instrument with periodic interest can be calculated using the
following formula:
P MT 1 P MT2 (P MTN + FV N )
PV = + +⋯+
(1 + r) 1 (1 + r)2 (1 + r)N
Where:
P MT = Periodic payment.
F V = Bond's principal.
Consider a five-year corporate bond issued in 2023 with a 4.00 percent annual coupon and a
price of USD110.00 per USD100 principal three years later. If Milka can reinvest periodic
interest at the original YTM of 4.00 percent, the implied three-year return is closest to:
Solution
We can calculate the future value (FV) after three years, including the future price of 110.00 and
55
© 2014-2024 AnalystPrep.
We can then solve for Milka's annualized return, r, using the formula for implied return since we
F Vt 122.49
r = √t − 1 = √3 − 1 = 3.65%
PV 110
This means that Milka, who purchased the corporate bond at a price of 100 and held it for three
CityGroup Corp. issued a corporate bond 7 years ago with a face value of $1,000 and a 20-year
maturity. The bond pays annual interest at a coupon rate of 6%. Currently, the bond is trading at
$1,120. The yield to maturity (YTM) of CityGroup Corp.'s bond is closest to:
Solution
We have
The value of a stock is determined by both the expected return and the growth of its cash flows.
56
© 2014-2024 AnalystPrep.
By assuming a constant growth rate for dividends, we can use the formula for the present value
Implied Return
Recall that the present value of a stock for constant growth of dividends is given by:
Dt(1 + g) Dt+1
P Vt = =
r −g r−g
Where:
r−g>0
Therefore, we can calculate the implied return on a stock given its expected dividend yield and
Dt (1 + g) Dt+1
r= +g = +g
P Vt PV t
In simple terms, if we assume a stock's dividends will grow at a steady rate forever, the implied
return is the combination of its expected dividend yield and the constant growth rate.
Assume Apple Inc. stock is trading at a share price of USD150.00, and its annualized expected
Moh, an analyst, projects that Apple's dividend per share will increase at a constant rate of 5%
per year indefinitely. The required return expected by investors on the stock is closest to:
57
© 2014-2024 AnalystPrep.
Solution
Dt (1 + g) Dt+1
r= +g = +g
P Vt PV t
Therefore,
2.00(1.05)
r= + 0.05 = 6.4%
150
Implied Growth
We can also solve for a stock's implied growth rate, which is given by the following formula:
r × P Vt − Dt r − Dt+1
g= + Dt =
P Vt P Vt
Consider the previous example. Suppose Moh believes that Apple stock investors should expect a
return of 8%, calculate the implied dividend growth rate for Apple Inc.
Solution
r × P Vt − Dt r − Dt+1
g= + Dt =
P Vt P Vt
So,
2.00 × 1.05
g = 0.08 − = 0.066 = 6.60%
150
Price-to-Earnings Ratio
58
© 2014-2024 AnalystPrep.
In equity instruments, it is common practice to compare the price-to-earnings ratio.
The price-to-earnings (P/E) ratio is a valuation metric that compares the current share price
of a stock to its earnings per share. Investors and analysts use it to determine the relative value
A stock with a higher price-to-earnings ratio is more expensive than a lower one, as investors are
willing to pay more for each unit of earnings. This ratio is also a valuation metric for stock
The P/E ratio can relate to our earlier discussion on a stock's price (PV) to the expected future
D t × (1 + g)
P Vt =
r−g
By dividing both sides of the equation by E t, which represents earnings per share for period t , we
Dt
× (1 − g)
PV t Et
=
Et r−g
Where:
P Vt
= Price-to-earnings (P/E) ratio.
Et
Dt
= Dividend payout ratio.
Et
g = Growth rate.
The dividend payout ratio represents the percentage of a company's earnings paid out to
59
© 2014-2024 AnalystPrep.
Typically, the forward P/E ratio, which is based on a projection of a company's earnings per
share for the next period (t + 1), is used. This ratio is positively correlated with higher expected
dividend payouts and growth rates but negatively correlated with the required return.
Dt
× (1 − g)
PV t Et
=
Et r−g
D t+1
P Vt E t+1 Dt+1 1
= = ×
Et+1 r−g Et+1 r −g
Suppose a company has a forward P/E ratio of 15, a dividend payout ratio of 40%, and a required
return of 10%. The implied dividend growth rate for this company is closest to:
Solution:
First, we can use the formula for the forward P/E ratio to solve for the implied dividend growth
rate:
P Vt Dt+1 1
= ×
Et+1 Et+1 r −g
Where:
r = Required return.
60
© 2014-2024 AnalystPrep.
Substituting the given values into the formula, we get:
0.4
15 =
0.1 − g
0.4
g = 0.1 − = 0.0733
15
Therefore, the implied dividend growth rate for this company is 7.33%.
Let's assume you are not given the required rate of return in the question above so that the
company has a forward P/E ratio of 15, a dividend payout ratio of 40%, and an implied dividend
growth rate of 7.33%. What is the required return for this company?
Solution
Recall the formula for the forward P/E ratio to solve for the required return:
P Vt Dt+1 1
= ×
Et+1 Et+1 r −g
0.4
15 =
r − 0.0733
0.4
r= + 0.0733 = 0.1000
15
61
© 2014-2024 AnalystPrep.
Question
Edmund company's stock trades at USD50.00. The company pays an annual dividend
to its shareholders, and its most recent payment of USD 2.00 occurred yesterday.
Analysts following the company expect its dividend to grow at a constant rate of 4
A. 8.16%.
B. 8.48%.
C. 9.16%.
Solution
Recall that:
Dt+1 × (1 + g)
PV =
r−g
Where:
r = Required return.
2 × (1 + 0.04)
50 =
r − 0.04
2 × (1 + 0.04)
62
© 2014-2024 AnalystPrep.
2 × (1 + 0.04)
r= + 0.04 = 0.0816
50
r = 0.0816
63
© 2014-2024 AnalystPrep.
LOS 2c: explain the cash flow additivity principle, its importance for the
no-arbitrage condition, and its use in calculating implied forward
interest rates, forward exchange rates, and option values
A timeline is a physical illustration of the amounts and timing of cashflows associated with an
investment project. For cashflows that are regular and of equal amounts, the standard annuity
formula or the financial calculator can be used. However, a timeline is preferred for irregular,
Remember that the general formula that relates the present value and the future value of an
F VN = P V (1 + r)N
Where:
64
© 2014-2024 AnalystPrep.
In a particular timeline, a time index, t, represents a particular point in time, a specified number
of periods from today. Therefore, the present value is the investment amount today (t = 0), and
by using this amount, we can calculate the future value (t = N). Alternatively, we can use the
The above argument can be written in terms of the present value. That is:
P V = F VN (1 + r)−N
received in perpetuity. Payments are to be made at the end of each year, starting at the end of
year 4. If the discount rate is 9%, then what is the present value of the perpetuity at t = 0?
Solution
65
© 2014-2024 AnalystPrep.
Here, we can see that the investor is receiving $6,500 in perpetuity. Recall that the PV of
C
PV of a perpetuity =
r
$6, 500
PV 3 = = $72, 222
9%
This is the value of the perpetuity at t = 3, so we need to discount it for three more periods to
P V0 = F V N (1 + r)−N
$72, 222
The PV at time zero is = $55 , 769
(1+0.09)3
There are many instances in real life when cashflows are uneven. A good example is a pension
66
© 2014-2024 AnalystPrep.
contribution that varies with age. Applying one of the basic time value formulae is impossible in
such cases. You are advised to draw a timeline even if the question appears relatively
straightforward. It will help you understand the question structure better. A timeline also helps
candidates add cashflows indexed to the same period and apply the value additivity principle.
According to the cashflow additivity principle, the present value of any stream of cashflows
indexed at the same point equals the sum of the present values of the cashflows. This principle
has different applications in time value of money problems. Besides, this principle can be applied
The principle of cash flow additivity can be applied to scenarios involving different currencies by
converting all cash flows to a common currency using the appropriate exchange rates. Doing so
allows us to compare and combine cash flows from different currencies and make investment
For example, suppose we have two investment opportunities, one in US dollars and one in Euros.
We can convert the expected cash flows from the Euros investment into US dollars using the
appropriate exchange rates. Then, we can compare the combined value of the two investments
Dealing with different currencies assumes continuous compounding. Recall that the present
PV = F VN e−N r s
Consider an investor with USD 2,000 who wants to invest it for three months. The investor can
67
© 2014-2024 AnalystPrep.
choose between two options: investing in the US government debt or German government debt.
The investor can invest his USD 2,000 in a three-month US Treasury bill. This means that he
lends the government USD 2,000, and it promises to pay him back with interest in three months.
Recall that,
3
F V = 2 , 000 × e0.03 × 12 = USD 2 , 015
The investor chooses to invest in German government debt. To do this, the investor must convert
his USD 2,000 into Euros at the current exchange of EUR/USD = 0.92 (1 USD =0.92EUR). This
means that the investor will receive (2, 000 × 0.92) = EUR 1840 . He can then lend this money to
the German government by investing in a three-month German Treasury bill. Assuming the
interest rate is 0.06 percent, after three months, the investor will receive:
0.06 × 3
F V = EUR 1840 × e 12
= EUR 1, 867.81
Assuming the investor wants his money in US dollars, we need to convert the EUR1867.81 back
into USD at the forward exchange rate of USD/EUR = 1.0788. This means that the investor will
receive:
Both options give you the same amount of money after three months: USD 2,015. The difference
is that one option involves investing in US dollars, and the other involves converting your money
The forward exchange rate of 1.0788 USD/EUR is important because it determines how much
money you, the investor, will receive when converting your Euros back into US dollars. If this
rate differs from 1.0788, there would be an arbitrage opportunity in converting Euros to dollars.
68
© 2014-2024 AnalystPrep.
More on foreign exchange rates will be discussed later in the curriculum.
Consider two zero-coupon bonds. Bond A has a maturity of two years and a yield of 2% per
annum, while bond B has a maturity of four years and a yield of 3%. An investor, who doesn't
seek to take advantage of price differences and is risk-neutral, has $1,000 to invest. The investor
Option 1: The investor can put their money into bond B now, which has an annual yield of 3%,
and will pay out at the end of the four years. The Future Value (FV) of this investment in four
F V4 = P V0 (1 + r4 )4 = 1000(1.03)4 = 1, 125.51
Option 2: Alternatively, the investor can initially invest in bond A and, after two years, reinvest
the proceeds at a forward rate F2 ,2 which represents a two-year forward rate starting in year
two.
By the principle of cash flow additivity, a risk-neutral investor will not prefer one option over the
other - they are indifferent between Options 1 and 2. This is because the Future Values of both
F V4 = P V0 (1 + r4 )4 = PV 0 (1 + r2 )(1 + F2 ,2 )
1, 125.51
⇒ F2,2 = − 1 = 8.18%
1 , 000(1.02)2
Therefore, to prevent arbitrage opportunities, the forward rate F2 ,2 should be set to 8.18%. This
69
© 2014-2024 AnalystPrep.
ensures that there is no potential for risk-free profits, maintaining market efficiency.
Cash flow additivity can be used to determine the fair price of an option contract. An option
contract gives the buyer the right, but not the obligation, to buy (call) or sell (put) an underlying
Cash flow additivity allows investors to compare different strategies and determine a no-
Consider a stock that costs $100 now. Its price might increase by 30% to $130 or decrease by
Let's say an investor wants to sell a call on the stock that gives the buyer the right, but not the
obligation, to purchase the asset for $120. The principle of cash flow additivity can be used to
70
© 2014-2024 AnalystPrep.
determine the contract's no-arbitrage price.
If the stock price goes up, the contract is worth cu1 = 10. That's because the buyer can use the
contract to buy the asset for $120 and then sell it for $130, making a profit of $10. But if the
price goes down, the contract is worth nothing. The buyer wouldn't want to use the contract to
buy the asset for $120 when they could just buy it for $85 without the contract.
The underlying argument here is that the value of the option is each movement of the stock
option may be used to construct a risk-free portfolio (the value of the portfolio is the same in
both scenarios).
Denote the initial value of the call option by c0 , which we wish to determine using cash flow
additivity and no-arbitrage pricing. Also, denote the value of the portfolio at t = 0 by V0 , when
the stock price increases by V1u and when the stock price decreases by V1d
Assume that at t = 0 creates a risk-free portfolio by selling a call option at c0 and buying 0.22
units of underlying assets. Then, the value of the portfolio at inception is:
V0 = 0.22 × 100 − c 0
Intuitively, the value replicating portfolio equals 18.89, whether the stock prices rise or decline.
As such, the replicating portfolio is risk-free and can be discounted as a risk-free asset. Assuming
At this point, we can calculate the value of c 0 rearranging the initial portfolio value equation:
71
© 2014-2024 AnalystPrep.
V0 = 0.22 × 100 − c0
⇒ 18.43 = 0.22 × 100 − c0
∴ c 0 = 3.57
As such, the fair price of the call option is $3.57, which the seller expects to receive from the
buyer.
72
© 2014-2024 AnalystPrep.
Question
The current USD/CHF exchange rate is 0.9. The risk-free interest rates for one year
are 2% for the US dollar and 1% for the Swiss franc. Which of the following one-year
A. USD/CHF 0.909.
B. USD/CHF 0.099.
C. USD/CHF 1.122.
Solution
Dealing with different currencies assumes continuous compounding. Recall that the
F V = P VN eNr s
So,
In one year, a single unit of Swiss Franc converted to US dollars and then invested
risk-free is worth;
1
1 CHF = ≈ 1.1111 USD
0.9
Therefore, to convert USD 1.1334 into CHF 1.0101 requires a forward exchange rate
of:
1.1334
73
© 2014-2024 AnalystPrep.
1.1334
= USD/CHF 1.1221
1.0101
74
© 2014-2024 AnalystPrep.
Learning Module 3: Statistical Measures of Asset Returns
The center of any data is identified via a measure of central tendency. A measure of central
tendency for a series of returns reveals the center of the empirical distribution of returns. They
Measures of location help us understand where data points tend to cluster. These measures
include central tendency measures such as mean, median, and mode. There are also other
measures that provide different insights into how the data is spread out or located within a
distribution.
Arithmetic Mean
The arithmetic mean is the sum of the values of the observations in a dataset divided by the
75
© 2014-2024 AnalystPrep.
number of observations.
Recall the formula: denoted by R̄i arithmetic mean for an asset i is a simple process of finding the
Where:
1 T 1
R̄ i = ∑ R it = (15% + 10% + 12% + 3%) = 10%
T t=1 4
The population mean is the summation of all the observed values in the population, ∑ Xi divided
by the total number of observations, N . The population mean differs from the sample mean,
which is based on a few observed values chosen from the population. Thus:
∑ Xi
Population mean =
N
∑ Xi
Sample mean =
n
Analysts use the sample mean to estimate the actual population mean.
The population mean and the sample mean are both arithmetic means. The arithmetic mean for
any data set is unique and is computed using all the data values. Among all the measures of
central tendency, it is the only measure for which the sum of the deviations from the mean is
76
© 2014-2024 AnalystPrep.
zero.
Median
The median is the statistical value located at the center of a data set organized in ascending or
descending order.
Unlike the arithmetic mean, the median resists the effects of extreme observations. However, it
only gives the relative position of the ranked observations without considering all observations
The following are the annual returns on a given asset realized between 2005 and 2015.
{12% 13% 11.5% 14% 9.8% 17% 16.1% 13% 11% 14%}
77
© 2014-2024 AnalystPrep.
The median is closest to:
Solution
{9.8% 11% 11.5% 12% 13% 13% 14% 14% 16.1% 17%}
Since the number of observations is even, the median return will be the middle point of the two
n ( n +2)
middle values in the positions and .
2 2
n 10 ( n+2) 12
The value occupying = = 5th position is 13, and the value located in = = 6th
2 5 2 2
13% + 13%
= 13%
2
Mode
The mode is the value that appears most often in a dataset. Sometimes, a dataset has a mode;
sometimes, it doesn't. If all the observations in a dataset are different and no value repeats more
A dataset with one mode is called unimodal. When there are two modes, it's called bimodal. If the
An interval with the highest frequency is called the modal interval (or intervals) in a frequency
distribution. For instance, in the frequency distribution below, the modal interval is -1.0 to 0.0
78
© 2014-2024 AnalystPrep.
Return Bin Absolute Relative Cumulative Cumulative]
(%) Frequency Frequency(%) Absolute Relative
Frequency Frequency (%)
-6.0 to -5.0 2 0.16 2 0.16
-5.0 to -4.0 8 0.64 10 0.80
-4.0 to -3.0 27 2.16 37 2.96
-3.0 to -2.0 80 6.40 117 9.36
-2.0 to -1.0 485 38.80 602 48.16
-1.0 to 0.0 520 41.60 1 , 122 89.76
0.0 to 1.0 100 8.00 1 , 222 97.76
1.0 to 2.0 24 1.92 1 , 246 99.68
2.0 to 3.0 3 0.24 1 , 249 99.92
3.0 to 4.0 1 0.08 1 , 250 100.00
The mode is the only measure of central tendency that can be used with nominal data. Nominal
data refers to a type of data that is categorized into distinct categories or groups without any
inherent order or numerical value. Examples of nominal data include gender (male, female), eye
color (blue, brown, green), and marital status (single, married, divorced).
79
© 2014-2024 AnalystPrep.
Determine the mode from the following data set:
{20% 23% 20% 16% 21% 20% 16% 23% 25% 27% 20%}
Solution
The mode is 20%. It occurs four times, a frequency higher than any other value in the data set.
An outlier may represent a distinct value in a population. In addition, it may show that there was
When working with a sample with outliers, we can potentially transform the variable or choose
another variable that achieves the same objective. If these observations prove to be impossible,
Option 1: Take no action and use the data as it is. If these observations are accurate, then this is
appropriate.
Option 2: Remove outliers using the trimmed mean. For example, when calculating the central
tendency, a 4 percent trimmed mean excludes the lowest 2% and highest 2% of values.
Option 3: Substitute a different value for the outliers. The winsorized mean is an illustration of a
central tendency that does this. For instance, when computing a 96% winsorized mean, the value
at or above the lowest and highest 2% is assigned the lowest and highest 2% values.
Measures of Location
Quartiles, quintiles, deciles, and percentiles are values or cut points that partition a finite
number of observations into nearly equal-sized subsets. The number of partitions depends on the
80
© 2014-2024 AnalystPrep.
Quartiles
They divide data into four parts. The first quartile, Q1, is referred to as the lower quartile, and
the last quartile, Q4, is the upper quartile. Q1 splits the data into the lower 25% and upper 75%
values. Similarly, the upper quartile subdivides the data into the lower 75% of the values and the
upper 25%. The difference between the upper and lower quartiles is known as the interquartile
range, which indicates the spread of the middle 50% of the data.
Quintiles
Though rarely used in practice, quintiles split a set of data into five equal parts, i.e., fifths.
Therefore, the second quintile splits data into the lower 40% of the values and the upper 60%.
Deciles
The deciles subdivide data into ten equal parts. There are 10 deciles in any data set. For
example, the fourth decile splits data into the lower 40% of the values and the upper 60%.
Percentiles
Percentiles split data into 100 equal parts, i.e., hundredths. So, for instance, the 77th percentile
splits the data into the lower 77% of the values and the upper 23%.
Financial analysts commonly use the four types of subdivisions to rank investment performance.
You should note that quartiles, quintiles, and deciles can all be expressed as percentiles. For
instance, the first quartile is just the 25th percentile. Similarly, the fourth decile is simply the
(n + 1) y
Position of percentile =
100
81
© 2014-2024 AnalystPrep.
{10% 23% 12% 21% 14% 17% 16% 11% 15% 19%}
Solution
{10% 11% 12% 14% 15% 16% 17% 19% 21% 23%}
Next, we establish the position of the first quartile. This is simply the 25th percentile. Therefore:
(10 + 1) 25
P 25 = = 2.75 th value
100
Since the value is not straightforward, we have to extrapolate between the 2nd and the 3rd data
points. The 25th percentile is three-fourths (0.75) of the way from the 2nd data point (11%) to the
Box and whisker plot is used to display the dispersion of data across quartiles. A box and whisker
plot consists of a "box" with "whiskers" connected to the box. It shows the following five-number
82
© 2014-2024 AnalystPrep.
Example: Box and Whisker Plot
83
© 2014-2024 AnalystPrep.
A. 10.
B. 10.5.
C. 11.
Solution
(10 + 11)
Median or Quartile 2 (Q2) = = 10.5
2
A. 4.5.
B. 6.5.
C. 11.
Solution
Interquartile range is Q3 − Q1 = 15 − 4 = 11
A. 4.
B. 15.
C. 18.
Solution
3 × (n + 1) 3 × (16 + 1)
Quartile 3 (Q3) = = = 12.75 th term, which is 14.75
4 4
84
© 2014-2024 AnalystPrep.
Quantiles in Investment Practice
Quantiles have two main purposes in investment practice. First, quantiles are used to rank
performance. Secondly, quantiles can be used in investment research for comparison purposes.
For example, companies can be clustered into deciles to compare the performance of small
companies with the large ones. In this case, the first decile will contain the portfolio of
companies with the smallest market values, while the tenth decile will contain the companies
85
© 2014-2024 AnalystPrep.
Question
A mutual fund achieved the following rates of growth over an 11-month period:
A. 2%.
B. 3%.
C. 4%.
Solution
Secondly, you should establish the 5th decile. This is simply the 50th percentile and is
(1 + 11) 50
P 50 =
100
= 12 × 0.5
= 6 , i.e., the 6th data point
86
© 2014-2024 AnalystPrep.
LOS 3b: calculate, interpret, and evaluate measures of dispersion to
address an investment problem
Measures of dispersion are used to describe the variability or spread in a sample or population.
They are usually used in conjunction with measures of central tendency, such as the mean and
the median. Specifically, measures of dispersion are the range, variance, absolute deviation, and
standard deviation.
Measures of dispersion are essential because they give us an idea of how well the measures of
central tendency represent the data. For example, if the standard deviation is large, then there
are large differences between individual data points. Consequently, the mean may not be
Range
The range is the difference between the highest and the lowest values in a dataset, i.e.,
{78 56 67 51 43 89 57 67 78 50}
Range = 89 − 43 = 46
The range is not a reliable dispersion measure. It provides limited information about
87
© 2014-2024 AnalystPrep.
The range is sensitive to outliers.
MAD is a measure of dispersion representing the average of the absolute values of the
∑ ∣∣Xi − X̄∣∣
MAD =
n
Remember that the sum of deviations from the arithmetic means is always zero, which is why we
Six financial analysts have reported the following returns on six different large-cap stocks over
2021:
Solution
{|6%– 6.83%| + |7%– 6.83%| + |12%– 6.83%| + |2%– 6.83%| + |3%– 6.83%| + |11%– 6.83
MAD =
6
0.83 + 0.17 + 5.17 + 4.83 + 3.83 + 4.17
=
6
= 3.17%
88
© 2014-2024 AnalystPrep.
Interpretation: On average, an individual return deviates by 3.17% from the mean return of
6.83%.
The sample variance,s2 , is the measure of dispersion that applies when working with a sample
instead of a population.
2
2
∑ (Xi − X̄)
s =
n −1
Where:
X̄ = Sample mean.
n = Number of observations.
The sample standard deviation, s, is simply the square root of the sample variance.
2
(Xi − X̄ )
s = √s = ⎷
2
n− 1
Assume that the returns realized in the previous example were sampled from a population
comprising 100 returns. The sample mean and the corresponding sample variance are closest to:
Solution
Hence,
89
© 2014-2024 AnalystPrep.
{(6% − 6.83%)2 + (7% − 6.83%)2 + (12% − 6.83%)2 + (2% − 6.83%)2 + (3% − 6.83%)2 + (11%
2
s =
5
= 0.001656
Therefore,
1
s = 0.001656 2
= 0.0407
When trying to estimate downside risk (i.e., returns below the mean), we can use the
following measures:
semi-variance.
Target semi-variance: The sum of the squared deviations from a specific target
return.
n (Xi − B)2
sTarget = ∑
⎷for all X n−1
i ≤B
90
© 2014-2024 AnalystPrep.
Month Return %
2010 36%
2011 29%
2012 10%
2013 52%
2014 41%
2015 16%
2016 10%
2017 23%
2018 −10%
2019 −19%
2020 2%
Solution
Squared
Deviation Deviation
Return deviations
Month from the 20% below the
% below the
target target
target
2010 36.00 16.00 − −
2011 29.00 9.00 − −
2012 10.00 (10.00) (10.00) 100
2013 52.00 32.00 −
2014 41.00 21.00 −
2015 16.00 (4.00) (4.00) 16
2016 10.00 (10.00) (10.00) 100
2017 23.00 3.00 −
2018 (10.00) (30.00) (30.00) 900
2019 (19.00) (39.00) (39.00) 1, 521
2020 2.00 (18.00) (18.00) 324
Sum 2,961
0.5
2961
Target semi-deviation = ( ) = 17.21%
10
Coefficient of Variation
91
© 2014-2024 AnalystPrep.
The coefficient of variation, CV , is a measure of spread that describes the amount of variability
of data relative to its mean. It has no units, so we can use it as an alternative to the standard
deviation to compare the variability of data sets that have different means. The coefficient of
s
CV =
X̄
Where:
?
Note: The formula can be replaced with ?
when dealing with a population.
What is the relative variability for the samples 40, 46, 34, 35, and 45 of a population?
Solution
Step 2: Calculate the sample standard deviation. (Start with the variance, s2.)
Note: Since it is the sample standard deviation (not the population standard deviation), we use
n − 1 as the denominator.
Therefore,
s = √30.5 = 5.52268
Mean 5.52268
= = 0.13806 or 13.81%
s 40
In finance, the coefficient of variation is used to measure the risk per unit of return. For
example, imagine that the mean monthly return on a T-Bill is 0.5% with a standard deviation of
0.58%. Suppose we have another investment, say, Y, with a 1.5% mean monthly return and
0.58
CVT−Bill = = 1.16
0.5
6
CVY = =4
1.5
Interpretation: The dispersion per unit monthly return of T-Bills is less than that of Y.
93
© 2014-2024 AnalystPrep.
Question 1
If a security has a mean expected return of 10% and a standard deviation of 5%, its
A. 0.005.
B. 0.500.
C. 2.000.
Solution
S 0.05
CV = = = 0.5
x? 0.10
Where:
0.05
CV = = 0.005
10
10
CV = =2
5
Question 2
94
© 2014-2024 AnalystPrep.
{12 13 54 56 25}
Assuming that this is a sample from a certain population, the sample standard
A. 21.62.
B. 374.00.
C. 1,870.00.
Hence,
{(12 − 32)2 + (13 − 32)2 + (54 − 32)2 + (56 − 32)2 + (25 − 32)2 }
s2 =
4
1870
= = 468
4
Therefore,
s = √468 = 21.62
95
© 2014-2024 AnalystPrep.
LOS 3c: interpret and evaluate measures of skewness and kurtosis to
address an investment problem
Since the deviations from the mean are squared when calculating variance, we cannot determine
whether significant deviations are more likely to be positive or negative. In order to identify
other crucial distributional traits, we must look beyond measures of central tendency, location,
and dispersion.
Skewness
Skewness refers to the degree of deviation from a symmetrical distribution, such as the normal
distribution. A symmetrical distribution has identical shapes on either side of the mean.
Distributions that are nonsymmetrical have unequal shapes on either side of the mean, leading to
skewness. This is because nonsymmetrical distributions depart from the usual bell shape of the
96
© 2014-2024 AnalystPrep.
normal distribution.
Skewness can be positive, negative, or, in some cases, undefined. The shape of a skewed
distribution depends on outliers, which are extremely negative and positive observations.
Positive Skewness
A positively skewed distribution has a long right tail because of many outliers or extreme
values on the right side. Perhaps the best way to remember its shape is to consider its points in a
positive direction. Most data points are concentrated on the left side.
specific country.
Negative Skewness
A negatively skewed distribution has a long left tail resulting from many outliers on the left side
of the distribution. Therefore, we could say that it points in the negative direction. This is
97
© 2014-2024 AnalystPrep.
because the right side harbors most of the data points.
Application of Skewness
Skewness matters in finance. Market data often show positive or negative skewness, like stock
prices or mortgage costs. Investors can predict if future prices will be above or below the mean
The approximate sample skewness when sample is large (n ≥ 100) is given by:
3
1 ∑n (Xi − X̄)
Skewness = ( ) i=1
n s3
Where:
X̄ = Sample mean.
98
© 2014-2024 AnalystPrep.
s = Sample standard deviation.
n = Number of observations.
A positive value indicates positive skewness. A 'zero' value indicates that the data is not skewed.
{12 13 54 56 25}
Solution
First, we must determine the sample mean and the sample standard deviation:
Therefore,
s = √467.5 = 21.62
3
1 ∑ni 1 (X i − X̄)
Skewness = ( ) =
n s3
3 3 3
1 (−20) + (−19) + 22 3 + 243 + (−7)
Skewness = ( )
5 21.62 3
Skewness = 0.1835
99
© 2014-2024 AnalystPrep.
Kurtosis
Kurtosis refers to the measurement of the degree to which a given distribution is more or less
'peaked' relative to the normal distribution. The concept of kurtosis is instrumental in decision-
Leptokurtic.
Mesokurtic.
Platykurtic.
Leptokurtic
A leptokurtic distribution is more peaked than the normal distribution. The higher peak results
from the clustering of data points along the x-axis. The tails are also fatter than those of a normal
The term "lepto" means thin or skinny. When analyzing historical returns, a leptokurtic
distribution means that small changes are less frequent since historical values are clustered
100
© 2014-2024 AnalystPrep.
around the mean. However, there are also large fluctuations represented by the fat tails.
Platykurtic
A platykurtic distribution has extremely dispersed points along the x-axis, resulting in a lower
peak when compared to a normal distribution. "Platy" means broad. Hence, the prefix fits the
distribution's shape, which is wide and flat. The points are less clustered around the mean
Returns that follow this type of distribution have fewer major fluctuations compared to
leptokurtic returns. However, you should note that fluctuations represent the riskiness of an
asset. More fluctuations represent more risk and vice versa. Therefore, platykurtic returns are
Mesokurtic
Lastly, mesokurtic distributions have a curve that is similar to that of a normal distribution. In
The majority of equity return series are found to have fat tails. Suppose a return distribution has
fat tails, and we apply statistical models that do not consider distribution. In that case, we will
Investors often study a stock's daily trading volume distribution to assess its trading liquidity. It
helps them see if the market can handle a large trade in that stock. This is useful for investors
who want to make big investments or exit their positions in a particular stock.
Sample kurtosis is always measured relative to the kurtosis of a normal distribution, which is 3.
Where:
4
¯
101
© 2014-2024 AnalystPrep.
n 4
1 ∑ (Xi − X̄)
Sample Excess Kurtosis = ( ) i =1
n s4
Positive excess kurtosis indicates a leptokurtic distribution. A zero value indicates a mesokurtic
Using the data from the example above (12, 13, 54, 56, and 25), determine the type of kurtosis
present.
Therefore,
n 4
1 ∑i 1 (Xi − X̄ )
Excess Kurtosis = ( ) = −3
n s4
4 4 4 4 4
1 (−20) + (−19) + 22 + 24 + (−7)
Excess Kurtosis = ( ) −3
5 21.624
Excess Kurtosis = 2.2139
102
© 2014-2024 AnalystPrep.
Question 1
A. Zero.
B. Positive.
C. Negative.
Solution
Since the normal curve is symmetric about its mean, its skewness is zero.
Question 2
A frequency distribution in which there are too few scores at the extremes of the
A. Platykurtic.
B. Leptokurtic.
C. Mesokurtic.
Solution
distribution. It implies that there are fewer scores at the extremes of the distribution,
103
© 2014-2024 AnalystPrep.
Question 3
When most of the data are concentrated on the left of the distribution, it is most likely
called:
A. Symmetric distribution.
Solution
A distribution is said to be skewed to the right or positively skewed when most of the
skewed to the left or negatively skewed if most of the data are concentrated on the
right of the distribution. The left tail clearly extends farther from the distribution's
104
© 2014-2024 AnalystPrep.
A is incorrect. A symmetric distribution is one in which the left and right sides
most of the data are concentrated on the right of the distribution. The left tail extends
105
© 2014-2024 AnalystPrep.
LOS 3d: Interpret the correlation between two variables to address an
investment problem
Covariance
∑N ¯
i=1 (Xi − X̄) (Yi − Y )
sX Y =
n −1
The formula above implies that the sample covariance is the mean of the product of the
deviations in the two random variables and from their sample means.
If the covariance between two random variables is positive, it means they move in the same
direction. When one is below its mean, the other is also below its mean, and vice versa.
A major drawback of covariance is that it is difficult to interpret since its value can vary from
Correlation
Correlation is a standardized measure of the linear relationship between two variables. It takes
the covariance and divides it by the product of the standard deviations of both variables. As a
Sxy
r xy =
Sx × Sy
Where:
106
© 2014-2024 AnalystPrep.
sX = Standard deviation of variable X.
Properties of Correlation
107
© 2014-2024 AnalystPrep.
A positive correlation close to +1 indicates a strong positive linear relationship.
Two variables can have a very low correlation despite having a strong nonlinear
relationship.
Correlation can be an unreliable measure when outliers are present in the data.
Correlation does not imply causation. This implies that the correlation may be spurious.
dataset.
variable.
variable.
108
© 2014-2024 AnalystPrep.
Question
The correlation coefficient between X and Y is 0.7, and the covariance is 29. If the
A. 8.29.
B. 29.
C. 68.65.
Solution
sX Y
rX Y =
sX × SY
29
⇒ 0.7 =
X ×5
∴ X = 8.2857
109
© 2014-2024 AnalystPrep.
Learning Module 4: Probability Trees and Conditional Expectations
Expected Value
Expected value is an essential quantitative concept investors use to estimate investment returns
and analyze any factor that may impact their financial position.
Mathematically, the expected value is the probability-weighted average of the possible outcomes
of the random variable. For a random variable X, the expected value of X is denoted E(X). More
specifically,
Where,
Note that the expected can be a forecast (looking into the future) or the true value of the
population mean.
The sample mean differs from the expected value. The sample mean is a central value for a
Return Probability
5% 65%
7% 25%
8% 10%
110
© 2014-2024 AnalystPrep.
The expected value of the investment is closest to:
Solution
Recall that,
n
E(X) = ∑ P(Xi )Xi
i=1
Consider expected value as a forecast of the outcome of an investment. Then, variance and
standard deviation measure the risk of an investment. That is the dispersion of outcomes around
the mean.
The variance of a random variable is the expected value (the probability-weighted average) of
squared deviations from the random variable's expected value. Denoted by σ 2 (X) or V ar(X), its
Since variance is in squared terms, it can take any number greater than or equal to
0(V ar(X) ≥ 0) . Intuitively, if V ar(X) = 0, there is no risk (dispersion). On the other hand, if
Moreover, V ar(X) is a quantity given in square units of X. That is, if the X is given in percentage,
111
© 2014-2024 AnalystPrep.
σ(X) = √σ 2(X) = √V ar(X)
The standard deviation is given in the same units as the random variance; hence it is easy to
interpret.
112
© 2014-2024 AnalystPrep.
Question
Return Probability
5% 65%
7% 25%
8% 10%
The variance and standard deviation of the investment are closest to:
Solution
We know that,
n
V ar(X) = ∑ P(Xi )[Xi − E(X)]2
i=1
113
© 2014-2024 AnalystPrep.
LOS 4b: Formulate an investment problem as a probability tree and
explain the use of conditional expectations in investment application
A tree diagram is a visual representation of all possible future outcomes and the associated
probabilities of a random variable. Tree diagrams are handy when we have several possible
outcomes.
They facilitate the recording of all the possibilities in a clear, uncomplicated manner. Each
Let's consider a scenario where we toss a fair coin twice. The outcomes of these tosses are
independent, meaning the first toss doesn't influence the second one.
For the first toss, we have two possibilities: It can result in either a head or a tail. Similarly, we
still have the same two possibilities for the second toss: head or tail. Importantly, the outcome of
the second toss is not affected by what happened in the first toss because coins don't have
memory.
114
© 2014-2024 AnalystPrep.
Please, note the following:
To calculate probabilities, we follow the tree branches from left to right and multiply any
probabilities we encounter.
So, to find the probability of getting two heads (HH), we multiply the probabilities along the
path.
1 1 1
115
© 2014-2024 AnalystPrep.
1 1 1
P(HH ) = × =
2 2 4
Conditional Expectations
based on specific real-world events. Analysts consider the probability and impact of future events
Competitors, governments, and other financial institutions keep releasing new information. Such
pieces of information may have a positive or a negative impact on investment. This means that a
In statistics, the conditional expected value is the expected value of a random variable X given an
Now, assume that X can take on any of n different outcomes X1 , X2 , … , Xn, which are outcomes
To state the unconditional expected values in terms of conditional expected value, we use the
total probability rule for expected value, which is built from the following formula:
116
© 2014-2024 AnalystPrep.
Now assume that S can take on any of S1 , S2 , … , Sn mutually exclusive and exhaustive scenarios
or events, then.
The probability of relaxed trade restrictions in a given country is 40%. Therefore, shareholders of
XYZ Company Limited expect a 5% share return if trade restrictions are maintained and a loss of
Solution
We must take every possibility into account. We have a 40% chance of relaxed trade restrictions
in this case. Intuitively, this means there is a 60% chance that the current restrictions will be
maintained. Therefore:
n
E(X) = ∑ P (Si) ⋅ E(X ∣ Si )
i =1
= 0.6(0.05) × 0.4(−0.08)
= −0.002
BlueChip Inc.'s profits are sensitive to economic growth, benefitting significantly during periods
of high economic growth. Suppose there is a 0.70 probability that BlueChip Inc. will operate in a
high-growth economic environment in the next fiscal year and a 0.30 probability that it will
If a high-growth economic environment occurs, the probability that EPS will be USD 3.00 is
estimated at 0.40, and the probability that EPS will be USD 2.80 is estimated at 0.60.
On the other hand, if the company operates in a moderate-growth environment, the probabilities
that the EPS will be USD 2.50 and USD 2.25 are 20% and 80%, respectively.
117
© 2014-2024 AnalystPrep.
Calculate the expected value of EPS for BlueChip Inc. in the next fiscal year.
Solution
We first need to calculate the conditional expectations of EPS for each scenario: High-growth
118
© 2014-2024 AnalystPrep.
E(EP S) = P (High-growth environment)
⋅ E(EP S ∣ High-growth environment)
+ P (High-growth environment)
⋅ E(EP S?Moderate-growth environment)
= 0.70 × 2.88 + 0.30 × 2.30
= USD 2.71
119
© 2014-2024 AnalystPrep.
Question 1
replacement. He then picks another one. Draw a probability tree and use it to
25
A. 102
.
13
B. 51
.
26
C. 51
.
Solution
120
© 2014-2024 AnalystPrep.
Question 2
There is a 20% chance that the government will impose a tariff on imported cars. A
company that assembles cars locally expects returns of 14% if the tariff is imposed
and returns of 11% otherwise. The (unconditional) expected return is closest to:
A. 11.6%.
B. 12.8%.
C. 12.5%.
Solution
1. The expected return given no tariff times the probability that a tariff
2. The expected return given tariff times the probability that the tariff
n
E(X) = ∑ P (Si) + E(X ∣ Si )
i =1
= 0.11(0.8) + 0.14(0.2)
= 0.116 = 11.6%
121
© 2014-2024 AnalystPrep.
LOS 4c: calculate and interpret an updated probability in an investment
setting using Bayes’ formula
Investors make investment decisions based on their experience and expertise. Their decisions
Bayes' formula allows us to update our decisions as we receive new information. In other words,
Bayes' formula is used to calculate an updated or posterior probability given a set of prior
Given a set of prior probabilities for an event, if we receive new information, the updated
probability is as follows:
P(Information ∣ Event)
P (Event ∣ Information) = ⋅ P (Event)
P (Information)
P (Bi ∩ A)
P (Bi ∣ A) = … … (1)
P (A)
122
© 2014-2024 AnalystPrep.
n n
P (A) = ∑ P (A ∩ B i) = ∑ P (B i) ⋅ P (A ∣ Bi ) … … (3)
i=1 i=1
(P (A ∣ B i)
P (Bi ∣ A) = ⋅ P (Bi )
∑ni=1 P (Bi ) ⋅ P(A ∣ Bi )
This is the Bayes' formula, and it allows us to ‘turnaround’ conditional probabilities, i.e., we can
Note that:
of stocks listed on different exchanges. In the sample, 50% of stocks were listed on the New York
Stock Exchange (NYSE), 30% on the London Stock Exchange (LSE), and 20% on the Tokyo Stock
Exchange (TSE).
The probability of a stock posting a negative return on the NYSE, LSE, and TSE is 40%, 35%, and
25%, respectively.
If the Analyst picks a stock at random from this group, what is the probability that it has a
Solution
123
© 2014-2024 AnalystPrep.
LSE is the event “A stock chosen at random is listed on the LSE.”
Finally, let NR be the event “A randomly chosen stock posts a negative return.”
Therefore,
124
© 2014-2024 AnalystPrep.
Question
You have developed a set of criteria for assessing potential investments in growth-
stage companies. Companies not meeting these criteria are predicted to be insolvent
within 24 months. You gathered the following information when validating your
criteria:
Fifty percent of the companies that have been assessed will become
The probability that a company will meet the criteria given that it remains
The probability that a company will remain solvent, given that it meets the criteria,
A. 20%.
B. 50%.
C. 62%.
Solution
Clearly, we need to calculate the P (meet criteria ∣ insolvency). Using the total
probability:
125
© 2014-2024 AnalystPrep.
P (meet criteria) = P (meet criteria ∣ solvency)P (solvency)
+ P (meet criteria ∣ insolvency)P (insolvency)
⇒ 0.65 = 0.80 × 0.50 + P(meet criteria ∣ insolvency) × 0.50
As such,
0.80 × 0.50
P (solvency ∣ meet criteria) = = 0.6153 ≈ 62%
0.80 × 0.50 + 0.50 × 0.50
126
© 2014-2024 AnalystPrep.
Learning Module 5: Portfolio Mathematics
LOS 5a: calculate and interpret the expected value, variance, standard
deviation, covariances, and correlations of portfolio returns
portfolio consists of assets such as stocks, bonds, or cash equivalents. Financial professionals
To calculate the portfolio's expected return, you take the expected returns of each security in the
portfolio. Then, you multiply each security's expected return by its proportion in the portfolio and
add them up. The formula below helps you find the portfolio's expected return:
Where:
1, 2 ,… , n.
Assume we have a simple portfolio of two mutual funds, one invested in bonds and the other
invested in stocks. Let us further assume that we expect a stock return of 8% and a bond return
of 6%, and our allocation is equal in both funds. The expected return would be calculated as
follows:
Portfolio Variance
The variance of a portfolio's return is a function of the individual asset covariances as well as the
127
© 2014-2024 AnalystPrep.
covariance between each of them.
Consider a portfolio with three assets: A, B, and C. The portfolio variance is given by:
Portfolio Variance
= WA2 σ 2(R A) + WB2 σ 2(R B) + WC2 σ 2(R C ) + 2(WA )(W B)Cov(RA , RB )
+ 2(WA )(WC )Cov(RA , R C ) + 2(WB )(WC )Cov(RB , RC )
Where:
Portfolio variance is a measure of risk. The higher the variance, the higher the risk. Investors
usually reduce the portfolio variance by choosing assets with low or negative covariance, e.g.,
Portfolio standard deviation is simply the square root of the portfolio variance. It is a measure of
Considering a portfolio with two assets, A and B, the portfolio standard deviation is given by:
Covariance
128
© 2014-2024 AnalystPrep.
Covariance is a measure of the degree of co-movement between two random variables. The
general formula used to calculate the covariance between two random variables, X and Y is:
Where:
This formula calculates the population covariance. It does this by taking the probability-weighted
average of the cross-products of the random variables' deviations from their expected values for
Sample Covariance
The sample covariance between two variables, X and Y , based on a sample data of size n is:
n (Xi − X̄ )(Y i − Y¯ )
Cov(X, Y ) = ∑
i=1 n− 1
Where:
X̄ = Sample mean of X .
Y¯ = Sample mean of Y .
The covariance between two random variables can be positive, negative, or zero.
A positive number indicates co-movement. The variables tend to move in the same
direction.
129
© 2014-2024 AnalystPrep.
A negative value shows that the variables move in opposite directions.
Covariance Matrix
A covariance matrix displays a complete list of covariances between assets needed to calculate
the portfolio variance. Consider a portfolio with three assets A, B, and C. The covariance matrix
is as follows:
Asset A B C
A Cov(RA , RA ) Cov(RA , RB ) Cov(RA , R C )
B Cov(RB , RA ) Cov(RB , RB ) Cov(RB , RC )
C Cov(RC , R A) Cov(RC , R B) Cov(R C , RC )
Asset A B C
A σA2 Cov(R A, R B) Cov(RA ,R C )
B Cov(R B, R A) σB2 Cov(RB , RC )
C C ov(R C , RA ) Co v(R C , RB ) σC2
not count the off-diagonal terms since they contain the individual variances of the assets. As
Note that:
C ov(R B, R A) = Cov(RA ,R B )
Cov(RA ,R C ) = Cov(RA ,R C )
Cov(R C , RB ) = Cov(RB , RC )
6
Therefore, there are = 3 distinct covariance terms in the above covariance matrix.
2
n( n −1)
In general, if we have n securities in a portfolio, there are distinct covariances and n
2
130
© 2014-2024 AnalystPrep.
variances to estimate.
Correlation
Correlation is the covariance ratio between two random variables and the product of their two
standard deviations. The correlation formula for random variables X and Y is:
Correlation measures the strength of the linear relationship between two variables. While the
covariance can take on any value between negative infinity and positive infinity, the correlation is
+1 indicates a perfect linear relationship (i.e., the two variables move in the same
-1 indicates a perfect inverse relationship, i.e., a unit change in one means that the
Harrison is a portfolio manager who oversees three assets: A, B, and C. The covariance matrix of
Asset A B C
A 0.04 0.02 0.01
B 0.02 0.05 0.015
C 0.01 0.015 0.09
Using this information, what is the correlation coefficient between assets B and C ?
Solution
131
© 2014-2024 AnalystPrep.
Note:
Cov(B , C)
Correlation (B, C) =
σBσC
0.015
= = 0.224
√0.05 × 0.09
We expect a 15% chance that ABC Corp's stock returns for the next year will be 6%. There's a
60% probability that they will be 8% and a 25% probability of a 10% return. The expected return
We also anticipate that the same probabilities and states are associated with a 4%, 5%, and 5.5%
return for XYZ Corp. The expected value of returns is then 4.975%, and the standard deviation is
0.46%.
To calculate the covariance and the correlation between ABC and XYZ returns, then:
Correlation(Ri, Rj)
Covariance(R ABC , RX YZ )
=
Standard deviation(RABC) × Standard deviation(RXYZ)
Therefore:
0.0000561
Correlation = = 0.976
(0.01249 × 0.0046)
The correlation between the returns of the two companies is very strong (almost +1), and the
An analyst studied five years of historical data to examine how changes in Central Bank interest
132
© 2014-2024 AnalystPrep.
rates affect the country's inflation rate. The covariance between the interest rate and inflation
rate is -0.00075. The standard deviation of the interest rate is 5.5%, and the inflation rate is
12%. Now, let's calculate and interpret the correlation between these two variables.
Solution
A correlation of -0.11364 indicates a negative correlation between the interest rate and the
inflation rate.
Cov(A, B)
Corr(A, B) = ρ(A, B) =
σA σB
⇒ Cov(A, B) = σA σBρ(A, B)
Consequently, in the formula for calculating portfolio variance, consisting of two assets, A and B,
133
© 2014-2024 AnalystPrep.
Question
Assume that we have investments in two companies, ABC and XYZ. For ABC, there's a
10% return. The expected return for ABC is 8.2%, and the standard deviation is
1.249%. For XYZ, there are similar probabilities of 4%, 5%, and 5.5% returns. The
expected return for XYZ is 4.975%, and the standard deviation is 0.46%.
A. 0.0000561.
B. 0.00007234.
C. 0.00851.
Since we already have the weight and the standard deviation of each asset, we can
√0.00007234 = 0.00851
134
© 2014-2024 AnalystPrep.
LOS 5b: Calculate and interpret the covariance and correlation of
portfolio returns using a joint probability function for returns
Historical covariance or other techniques, such as market model regression with historical
return data, can help us forecast return covariance and correlation. We use the joint probability
The probability that values of the two random variables X and Y will occur simultaneously is
given by the joint probability function of X and Y , denoted as P (X, Y ). For instance,
This formula calculates the covariance between random variables X and Y, such as portfolio
returns.
To find it, we take the sum of the products of the deviations of X and Y from their expected
Two random variables, X and Y , are independent if P (X, Y ) = P (X) ⋅ P (Y ) . That is, X and Y are
The independence property is stronger than correlation because the correlation coefficient
If random variables X and Y are uncorrelated (also holds for independent random variables),
135
© 2014-2024 AnalystPrep.
then:
Assume we wish to find the variance of each asset and the covariance between the returns of
ABC and XY Z, given that the amount invested in each company is $1,000.
Solution
For us to find the covariance, we must calculate the expected return of each asset as well as
1000
W ABC = = 0.5
2000
1000
WXY Z = = 0.5
2000
Finally, we can compute the covariance between the returns of the two assets:
136
© 2014-2024 AnalystPrep.
A portfolio manager is considering the following two possible economic growth of a country and
Solution
Note: For the rest of the calculation, your curriculum sometimes ditches the percentage signs so
Since covariance is negative, the two returns show some co-movement in opposite signs.
137
© 2014-2024 AnalystPrep.
Question
The following table represents the estimated returns for two motor vehicle
Given the above joint probability function, the covariance between T Y and Ford
A. 0.054.
B. 0.1542.
C. 0.1442.
Solution
First, we must start by calculating the expected return for each brand:
138
© 2014-2024 AnalystPrep.
Covariance = 0.5(6% − 3.7%)(10% − 5.4%)
+ 0.3(3% − 3.7%)(4%– 5.4%)
+ 0.2(−1% − 3.7%)(−4%– 5.4%)
= 5.29% + 0.294% + 8.836%
= 0.1442
The covariance is positive. This means that the returns for the two brands show some
In real life, this scenario is highly likely because the companies belong to the same
139
© 2014-2024 AnalystPrep.
LOS 5c: define shortfall risk, calculate the safety-first ratio, and identify
an optimal portfolio using Roy’s safety-first criterion
Modern Portfolio Theory (MPT) evaluates investment options based on mean return and return
variance. This approach is applicable when investors are risk-averse, meaning they seek to
ii. Investors have quadratic utility functions, a mathematical model representing the
The mean-variance analysis can be reasonably accurate even if the two assumptions aren't
entirely met. Professionals prefer using observable data, such as returns. The assumption that
returns roughly follow a normal distribution has played a crucial role in applying MPT.
Mean-variance analysis only considers risk symmetrically. This implies that standard deviation
reflects variability above and below the mean. An alternative strategy is focusing on downside
Shortfall Risk
Shortfall risk refers to the probability that a portfolio will not exceed the minimum (benchmark)
return an investor sets. In other words, it is the risk that a portfolio will fall short of the level of
return considered acceptable by an investor. As such, shortfall risks are downside risks. While a
shortfall risk focuses on the downside economic risk, the standard deviation measures the overall
140
© 2014-2024 AnalystPrep.
Safety-First Ratio
Roy's safety-first criterion states that the optimal portfolio is the one that minimizes the
probability that a portfolio return, denoted by RP , may fall below the threshold level of return, RL
As such, if returns are distributed normally, the optimal portfolio is the one with the highest
E(R P ) − R L
SFRatio =
σP
141
© 2014-2024 AnalystPrep.
The numerator, E(RP − RL ), represents the distance from the mean return to the threshold level,
i.e., it measures the excess return over and above the threshold level of return per unit risk.
Intuitively, if the returns are normally distributed, the safety-first optimal portfolio maximizes the
SFRatio.
Given a portfolio SFRatio, the probability that its return will be less than RL is:
An investor sets a minimum threshold of 3%. There are three portfolios from which he is to
choose one. The expected return and the standard deviation for each portfolio are given below:
Solution
Compute the safety-first ratio for each of the three portfolios and then compare them.
For portfolio A:
5− 3
SFRatio A = = 0.1333
15
10 − 3
SFRatio B = = 0.35
20
Lastly:
20 − 3
142
© 2014-2024 AnalystPrep.
20 − 3
SFRatioC = = 0.68
25
The optimal portfolio should maximize the safety-first ratio. Comparing the three ratios, it is easy
to notice that the safety-first ratio for portfolio C is the highest. Therefore, the investor should
choose portfolio C.
143
© 2014-2024 AnalystPrep.
Question
The returns on a fund are distributed normally. At the end of year t, the fund has a
value of $100,000. At the end of year t + 1, the fund manager wishes to withdraw
$10,000 for further funding but is reluctant to tap into the $100,000. There are two
investment options:
Portfolio A Portfolio B
Expected return 14% 13%
Standard deviation 17% 20%
A. Portfolio A.
B. Portfolio B.
Solution
First, you should calculate the threshold return from the information given. Since
there should be no tapping into the fund, the threshold return is:
10 , 000
= 10% or 0.1
100, 000
You should then calculate the safety-first ratio for each portfolio:
14 − 10
SFRatioA = = 0.24
17
13 − 10
SFRatioB = = 0.15
20
Portfolio A has the highest safety-first ratio. This is the reason it is the most desirable.
You can also go a step further and calculate P (R P < R L ). To do this, you would have to
144
© 2014-2024 AnalystPrep.
negate each safety-first ratio and then find the CDF of the standard normal
P (R P < R L ) = N (−SFRatio)
N (−0.24) = 1 − N (0.24)
= 1 − 0.5948 = 0.4052
N (−0.15) = 1 − N (0.15)
= 1 − 0.5596 = 0.4404
the threshold return. For portfolio B , this probability rises to 44%. Therefore, we
choose the option for which the chance of not exceeding the benchmark return is
lowest – portfolio A.
145
© 2014-2024 AnalystPrep.
Learning Module 6: Simulation Methods
distributed.
The lognormal distribution is positively skewed, meaning it's skewed to the right and has a long
right tail. In this distribution, values are bounded by 0. Typically, the mean is greater than the
mode.
Consider the following graph of two probability density functions (pdfs) of two lognormal
distributions.
146
© 2014-2024 AnalystPrep.
Like the normal distribution, two parameters – the mean and variance of the associated normal
Assume that X is normally distributed with the mean μ and variance σ 2. Also, define the variable
Y = eX .
Then ln Y = ln (eX ) = X is lognormally distributed with the following mean and variance
expressions:
147
© 2014-2024 AnalystPrep.
1
(μ+ σ 2 )
Mean = μL = e 2
2 2
Variance = σL2 = e2u +σ (eσ − 1)
The lognormal distribution works well for modeling asset prices that cannot be negative because
When the continuously compounded returns on a stock follow a normal distribution, the stock
prices follow a lognormal distribution. Note that even if returns do not follow a normal
distribution, the lognormal distribution is still the most appropriate for stock prices.
Remember that given the investment horizon from time t = 0 to time t = T , the continuously
PT
r0 ,T = ln ( )
P0
If we apply the exponential function on both sides of the equation, we have the following:
PT = P 0 er0,T
PT
Note that can be written as:
P0
PT PT P T−1 P
=( )( ) …( 1 )
P0 P T −1 P T−2 P0
PT PT P T−1 P
ln ( ) = ln (( )( ) … ( 1 ))
P0 PT −1 P T−2 P0
⇒ r0,T = rT −1,T + rT−2 ,T−1 + … + r0,1
148
© 2014-2024 AnalystPrep.
Therefore, the continuously compounded return to time T equals the sum of one-period
Remember that a linear combination of normal random variables is also normal. Therefore, if the
rT−1 ,T , rT −2,T −1 , … , r0,1 are independently and identically distributed (i.i.d) random variables with
The expected value of the continuously compounded return over a holding period of T
The variance of the continuously compounded return over a holding period is given by:
σ 2 (r0 ,T ) = σ 2T
The standard deviation of the continuously compounded returns, also known as volatility, is given
by:
σ(r0, T ) = σ√ T
In other words, if rT−1 ,T , rT −2, T−1 , … , r0, 1 are normally distributed with the mean of μ and variance
PT = P 0 er0,T
If X is normally distributed with the mean μ and variance σ2 and that Y = eX then,
formula, it would be easy to see that we can model P T as a lognormally distributed random
149
© 2014-2024 AnalystPrep.
variable since r0 ,T is approximately normally distributed.
Volatility measures the standard deviation of the continuously compounded returns on the
We calculate volatility using the historical series of continuously compounded returns. Another
method is converting daily holding returns into continuously compounded daily returns and then
We base annualizing volatility on 250 trading days in a year, which is an estimate of the business
days the financial markets operate. The formula we use for annualizing volatility is:
σ(r0, T ) = σ√T
For example, if the daily volatility is 0.05, then the annual volatility is:
Jess Kasuku is analyzing the stock of ABC Company, which is listed on the London Stock
Exchange under the ABC ticker symbol. Kasuku wants to understand how the stock's price
changed during a particular week when significant developments in the global economy
impacted the UK stock market. To do this, she calculates the stock's volatility for that week using
Using the information in Table 1, calculate the annualized volatility of ABC Company's stock for
150
© 2014-2024 AnalystPrep.
that week, assuming 250 trading days in a year.
Solution
Step 1: Calculate the continuously compounded daily returns for each day using the formula
Ending Price
ln( Beginning Price
):
78
r1 = ln ( ) = 0.03922
75
72
r2 = ln ( ) = −0.08004
78
70
r3 = ln ( ) = −0.02817
72
68
r4 = ln ( ) = −0.02899
70
r1 + r2 + r3 + r4
μ=
4
0.03922 + (−0.08004) + (−0.02817) + (−0.02899)
=
4
= −0.024495
Step 4: Calculate the standard deviation of the continuously compounded daily returns:
σ = √σ 2
= √ 0.001795 = 0.042363
Step 5: Annualize the volatility by multiplying the daily volatility by the square root of the
151
© 2014-2024 AnalystPrep.
number of trading days in a year.
We know that:
σ(r0 ,T ) = σ√T
∴ σannualized = σdaily × √250
= 0.042363 × √ 250
= 0.6698 ≈ 67%
So, the annualized volatility of ABC Company's stock for that week was 67.23 percent.
152
© 2014-2024 AnalystPrep.
Question
distributions?
C. They are less suitable for describing asset prices than asset returns.
Solution
are less suitable for describing asset prices than asset returns.
153
© 2014-2024 AnalystPrep.
LOS 6b: describe Monte Carlo simulation and explain how it can be used
in investment applications
Monte Carlo simulations are about producing many random variables based on specific
Imagine an investor who wants to predict the results of a 70% stock and 30% bond portfolio over
The quantity of interest here could be the final portfolio value after 20 years, denoted as ViT . In
this case, this is the final portfolio value at time T resulting from ith simulation trial.
The underlying variable is the return on the portfolio. The starting portfolio value is $100,000,
Assume we're interested in yearly returns, so the time horizon is 20 years. Divide the calendar
time into sub-periods. In this case, we will assume yearly returns so that the number of sub-
periods is K = 20, and the time increment Δ t is, therefore, one year.
Step 3: Specify the method for generating the data used in the simulation
Here, we need to make distributional assumptions. We might assume that the annual portfolio
return follows a normal distribution. Let's say we estimate an average return μ of 7% for stocks,
3% for bonds, a standard deviation σ of 15% for stocks, and 5% for bonds. We can model changes
154
© 2014-2024 AnalystPrep.
ΔPortfolio value = 0.7 ∗ (μ stock × Prior portfolio value × Δt
+ σstock × Prior portfolio value × Zk )
+ 0.3 ∗ (μ bond × Prior portfolio value × Δt
+ σbond × Prior portfolio value × Zk )
Here, Zk is a standard normal random variable representing the uncertainty in the portfolio
return (risk factor). We can use a computer program to draw 20 random values of Zk .
This step involves converting the standard normal random numbers (Zk) generated in step 3 into
yearly changes in portfolio value (ΔPortfolio value) using our model from step 3. This gives us 20
observations of possible changes in portfolio value over the 20-year period. From these
observations, we create a sequence of 20 portfolio values, starting with the initial value of
$100,000.
The average portfolio value at the end of 20 years (V iT ) is calculated by summing up the portfolio
values at the end of each year and dividing by 20. We then calculate the present value (Vi0 ) of
this average value by discounting it to the present using an appropriate interest rate. The
subscript i in V iT and Vi0 indicates that these values are from the ith simulation trial. This
Finally, we repeat steps 4 and 5 multiple times, say, 1,000 times. We then calculate summary
statistics, such as the mean, median, and percentiles of the distribution of Vi 0 values. These
summary statistics provide a range of potential outcomes for the portfolio value after 20 years,
helping the investor understand the risks and rewards of the investment strategy.
155
© 2014-2024 AnalystPrep.
It can also be used to value complex securities such as American or European options.
It is fairly complex and can only be carried out using specially designed software that
may be expensive.
The complexity of the process may cause errors, leading to wrong results that can be
potentially misleading.
156
© 2014-2024 AnalystPrep.
Question
Which of the following is a correct statement about the use of Monte Carlo
Solution
Monte Carlo simulations can assess how changes in assumptions, such as interest
rates or market volatility, affect a financial model. This allows analysts to understand
call options. Instead, they can estimate the value of these options by simulating their
potential outcomes.
potential returns, they do not simply simulate its performance. Instead, they use
157
© 2014-2024 AnalystPrep.
LOS 6c: describe the use of bootstrap resampling in conducting a
simulation based on observed data in investment applications
Resampling
Resampling means repeatedly drawing samples from the original observed sample to make
statistical inferences about population parameters. There are two common methods: Bootstrap
Bootstrap Resampling
Bootstrap resampling relies on computer simulations for statistical inferences, bypassing the
need for conventional analytical formulas like z-statistics. The bootstrap technique is
underpinned by a strategy that mirrors the random sampling process from a population to create
a sampling distribution.
158
© 2014-2024 AnalystPrep.
Note that in bootstrapping, we do not have information about the population. Our only insight
The core concept is that a random sample can effectively stand in for the entire population. So,
we can mimic drawing samples from the population by repeatedly resampling from the initial
sample. Essentially, the bootstrap method treats the initially obtained sample as a stand-in for
Both bootstrap and Monte Carlo simulation techniques heavily rely on repetitive sampling.
Bootstrap considers the resampled dataset a proxy for the true population and infers population
parameters such as mean, variance, skewness, and kurtosis from the statistical distribution of
these samples.
159
© 2014-2024 AnalystPrep.
Conversely, Monte Carlo simulation is centered on the generation of random data with a pre-
Simulation using bootstrapping is similar to Monte Carlo Simulation except for the source of
random variables. In bootstrapping, the random variables are taken from a bootstrap sample
Let's say an investor wants to understand the potential outcomes of investing in a portfolio with
a 70-30 split between stocks and bonds over a 20-year period. Here's how a Monte Carlo
The simulation steps using the bootstrap sampling distribution are as follows:
The quantity of interest here could be the final portfolio value after 20 years, denoted as ViT . The
underlying variable is the return on the portfolio. The starting portfolio value is $100,000, with
Assume we're interested in yearly returns, so the time horizon is 20 years. Divide the calendar
time into sub-periods. In this case, we will assume yearly returns so that the number of
Step 3: Generate bootstrap samples from the empirical distribution of portfolio returns
Here, we use the historical return data as our empirical distribution. Instead of assuming that
the annual portfolio return follows a specific theoretical distribution, we will use the bootstrap
procedure to draw the K = 20 yearly returns from the observed empirical distribution.
160
© 2014-2024 AnalystPrep.
Running the Simulation Over a Given Number of Trials
Step 4: Use the bootstrap samples to produce portfolio values used to value the
contingent claim
This step involves using the bootstrap samples drawn in Step 3 to compute the yearly changes in
portfolio value. From there, we create a sequence of 20 portfolio values, starting with the initial
value of $100,000.
The average portfolio value at the end of 20 years (V iT ) is calculated by summing up the portfolio
values at the end of each year and dividing by 20. We then calculate the present value (Vi0 ) of
this average value by discounting it to the present using an appropriate interest rate. The
subscript i in V iT and Vi0 indicates that these values are from the ith bootstrap sample. This
Finally, we repeat steps 4 and 5 multiple times, say, 1,000 times. We then calculate summary
statistics, such as the mean, median, and percentiles of the distribution of Vi 0 values. These
summary statistics provide a range of potential outcomes for the portfolio value after 20 years,
helping the investor understand the risks and rewards of the investment strategy based on the
161
© 2014-2024 AnalystPrep.
Question
analysis?
probability distributions for primary risk factors that govern the underlying
random variables.
Solution
from a set of unknown population parameters. Although the actual distribution of the
B is incorrect. In bootstrap analysis, the analyst repeatedly samples from the initial
sample, not the entire population. Each resample has the same size as the original
sample, and for each new draw, selected items go back into the sample.
establish probability distributions for the key risk factors that govern the underlying
162
© 2014-2024 AnalystPrep.
Learning Module 7: Estimation and Inference
LOS 7a: compare and contrast simple random, stratified random, cluster,
convenience, and judgmental sampling and their implications for
sampling error in an investment problem
Sampling is the systematic process of selecting a subset or sample from a larger population.
Sampling is essential because it is costly and time-consuming to analyze the whole population.
Sampling methods can be broadly categorized into probability sampling and non-probability
sampling.
In probability sampling, every population member has an equal chance of being chosen for the
factors such as the sampler's judgment or data accessibility, increasing the risk of an
unrepresentative sample.
Simple random sampling means selecting a sample from a population where each element has an
equal chance of being chosen. This method aims to create an unbiased sample that accurately
Imagine we wish to come up with a sample of 50 CFA level I candidates out of 100,000.
One approach may involve numbering each of the 100,000 candidates, placing them in a basket,
and shaking the basket to jumble up the numbers. Next, we would randomly draw 50 numbers
163
© 2014-2024 AnalystPrep.
A more scientific approach may also involve the use of random numbers. All the 100,000
candidates are numbered in a sequence (from 1 to 100,000). We may then use a computer to
randomly generate 50 numbers between 1 and 100,000, where a given number represents a
The underlying feature of random sampling is that all elements in the population must have
In stratified random sampling, analysts subdivide the population into separate groups known as
strata (singular stratum). Each stratum comprises elements with a common characteristic
(attribute) that distinguishes them from all the others. The method is most appropriate for large
heterogeneous populations.
A simple random sample is then drawn from within each stratum and combined to form the
overall, final sample that takes heterogeneity into account. The number of members chosen from
any one stratum depends on its size relative to the population as a whole.
An advertising firm wants to determine the extent to which it needs to invigorate television
advertisements in a district. The company decides to conduct a survey to estimate the mean
number of hours households spend watching TV per week. The district has three distinct towns –
A, B, which are urbanized, and C , located in a rural area. Town A is adjacent to a major factory
where most residents work, with most having kids of school-going age. Town B mainly harbors
There are 160 households in town A, 60 in town B, and 80 in C . Given the differences in the
composition of each region, the firm decides to draw a sample of 50 households, considering the
What is the number of homes that have been sampled in each town?
Solution
164
© 2014-2024 AnalystPrep.
We have three strata: towns A, B, and C. We use the following formula to determine the number
160
= × 50 = 27 (approximately)
300
60
= × 50 = 10
300
80
Finally, the firm would need ( 300 × 50) = 13 households in town C.
Stratification enables analysts to estimate the population parameter, say, the mean for
Cluster Sampling
Cluster sampling involves categorizing all population elements into distinct and all-encompassing
groups called clusters. Then, you can either choose a random sample of entire clusters or select
a random subset from each cluster. So, there are two cluster sampling approaches:
One-stage (or single-stage) cluster sampling: All the members in each sampled cluster
are sampled.
165
© 2014-2024 AnalystPrep.
each cluster.
In cluster sampling, a cluster serves as a single sampling unit, and only specific
In stratified sampling, you select members from within each stratum and then draw a
Non-probability samples are selected based on judgment or the convenience of accessing data.
As such, non-probability sampling depends on the researchers' sample selection skills. There are
access. This method may not provide a fully representative sample, limiting sampling
accuracy.
ii. Judgmental sampling: Researchers select elements subjectively, often based on their
own knowledge and expertise. However, this approach can introduce bias and result in a
166
© 2014-2024 AnalystPrep.
non-representative sample.
Sampling error refers to the difference between the observed value (results obtained from
analyzing a sample of investment data) and the true values that would have been obtained from
For instance, when we take a sample to estimate a population's mean, there's typically a
difference between the sample mean and the true population mean. This difference, known as
sampling error, emerges due to natural variation in sampling and because we work with data
Therefore, any conclusions or predictions drawn based on the sample data may deviate from the
167
© 2014-2024 AnalystPrep.
Question 1
according to the annual family income: Less than $30,000, $31,000 – $40,000,
$41,000 to $50,000, and $51,000 to $60,000. He then selects a sample from each
distinct group to form a whole sample. The sampling method used by the analyst is
most likely:
A. Cluster sampling.
B. Stratified sampling.
Solution
Dividing the population into different strata/groups and selecting a sample from each
entire population such that each member or element of the population has an equal
Question 2
A Ph.D. student is conducting research related to her thesis, and for this purpose, she
uses some students from her university to constitute a sample. The sampling method
168
© 2014-2024 AnalystPrep.
A. Simple random sampling.
B. Convenience sampling.
C. Judgmental sampling.
Solution
The researcher has selected the students from her university because she can
entire population such that each member or element of the population has an equal
Question 3
An analyst wants to estimate the downtime of ABC Bank's ATMs in a city for the last
six months. For this purpose, he selects 20 locations or areas within the city and then
selects 50% of the ATMs in each area. The sampling method used by the analyst is
most likely:
A. Cluster sampling.
Solution
169
© 2014-2024 AnalystPrep.
In cluster sampling, all population elements are categorized into mutually exclusive
and exhaustive groups called clusters. A simple random sample of the cluster is
selected, and then the elements in each of these clusters are sampled.
elements that have a common characteristic (attribute) that distinguishes them from
entire population such that each member or element of the population has an equal
170
© 2014-2024 AnalystPrep.
LOS 7b: explain the central limit theorem and its importance for the
distribution and standard error of the sample mean
The central limit theorem asserts that “given a population described by any probability
distribution having mean μ and finite variance σ 2, the sampling distribution of the sample mean
X̄ computed from random samples of size n from this population will be approximately normal
2
σ
with mean μ (the population mean) and variance (the population variance divided by n ) when
n
The answer to this question might not be straightforward. Nevertheless, the widely accepted
value is n ≥ 30. The truth is that the value of n depends on the shape of the population involved,
In a non-normal but fairly symmetric distribution, n = 10 can be considered large enough. With a
Remember that from the central limit theorem, the variance of the sample mean distribution is
given by:
σ2
σX̄2 =
n
The standard error is the standard deviation of the statistic (sample mean).
σ
σX̄ =
√n
Formally defined, for a sample mean X̄ computed from a sample generated by a population with
171
© 2014-2024 AnalystPrep.
standard deviation σ, the standard error of the sample mean is given by:
σ
σX̄ =
√n
When the population standard deviation, σ, is unknown, the following formula is used to estimate
s
sX̄ =
√n
The formula above is applicable where we do not know the population standard deviation. Note
that the sample standard deviation is the square root of the sample variance, s2, given by:
2
∑ ni 1 (Xi − X̄ )
2 =
s =
n−1
2
∑ni=1 (Xi − X̄)
⇒s=⎷
n− 1
The standard error of the sample mean estimates the variation that would occur if you took
multiple samples from the same population. While the standard deviation measures variation
within one sample, the standard error estimates variation across many samples. So, standard
The standard error of the sample mean gives analysts an idea of how precisely the sample mean
estimates the population mean. A lower standard error value indicates a more precise estimation
of the population mean. On the other hand, a larger standard error value indicates a less precise
It is also important to note that the standard error becomes smaller as the sample size increases.
This can be seen from its formula. This happens because increasing the sample size ultimately
brings the sample mean closer to the true value of the population mean.
172
© 2014-2024 AnalystPrep.
Example 1
In a certain property investment company with an international presence, workers have a mean
hourly wage of $12 with a population standard deviation of $3. Given a sample size of 30, the
σ
σX̄ =
√n
3
= = $0.55
√30
If we were to draw several samples of size 30 from the employee population and construct a
sampling distribution of the sample means, we would end up with a mean of $12 and a standard
error of $0.55.
Example 2
A sample of 30 latest returns on XYZ stock reveals a mean return of $4 with a sample standard
deviation of $0.13. The standard error of the sample mean is closest to:
s
sX̄ =
√n
0.13
= $0.02
√30
If we were to draw more samples from the population of yearly returns on XYZ stock and
construct a sample mean distribution, we would end up with a mean of $4 and a standard error
of $0.02.
173
© 2014-2024 AnalystPrep.
Question
Emma Johnson wants to know how finance analysts performed last year. Johnson
returns is 8 percent and that the returns are independent across analysts.
The random sample size that Johnson needs if she wants the standard deviation of the
A. 4.
B. 16.
C. 72.
Solution
Remember that,
σ
σX̄ =
√n
0.08
⇒ 0.02 =
√n
∴ n = 16
174
© 2014-2024 AnalystPrep.
LOS 7c: describe the use of resampling (bootstrap, jackknife) to estimate
the sampling distribution of a statistic
Resampling refers to the act of repeatedly drawing samples from the original observed data
sample for the statistical inference of population parameters. The two commonly used methods
Bootstrap
Using a computer, the bootstrap resampling method simulates drawing multiple random samples
from the original sample. Each resample is the same size as the original sample. These
175
© 2014-2024 AnalystPrep.
In the bootstrap method, the number of repeated samples drawn is at the researcher's
Furthermore, we can calculate the standard error of the sample mean. This is done by
resampling and calculating the mean of each sample. The following formula is used to estimate
B 2
1
sX̄ = (^ − )
⎷ B − 1 b=1 θ b θ̄
∑
Where:
The bootstrap resampling method can also be used to estimate the confidence intervals for
because it doesn't rely on an analytical formula for estimating distributions. This makes
it versatile for complex estimators and especially useful when analytical formulas are
unavailable.
that effectively handles complicated estimators. It can handle a wide range of statistical
models, making it suitable for various applications in finance where complex estimations
are common.
datasets and estimating population parameters on each. This helps understand estimator
176
© 2014-2024 AnalystPrep.
variability and robustness, ultimately improving result accuracy.
Jackknife
Jackknife is a resampling method in which samples are drawn by omitting one observation at a
time from the original data sample. This process involves drawing samples without replacement.
For a sample size of n, we need n repeated samples. This method can be used to reduce the bias
of an estimator or to estimate the standard error and the confidence interval of an estimator.
177
© 2014-2024 AnalystPrep.
Question
Assume that you are studying the median height of 100 students in a university. You
draw a sample of 1000 students and obtain 1000 median heights. The mean across all
resample means is 5.8. The sum of squares of the differences between each sample
2
mean, and the mean across all resample means ∑B
b=1
(θ^b − θ̄) is 2.3.
The Estimate of the standard error of the sample mean is closest to:
A. 0.05.
B. 0.08.
C. 0.10.
Solution
B 2
1
sX̄ = ⎷ ∑ (θ^b − θ̄)
B − 1 b=1
1
=√ × 2.3 = 0.04798 ≈ 0.05
1000 − 1
178
© 2014-2024 AnalystPrep.
Learning Module 8: Hypothesis Testing
opinion or claim about an issue. To determine if a hypothesis is accurate, statistical tests are
used. Hypothesis testing uses sample data to evaluate if a sample statistic reflects a population
“The mean return of small-cap stock is higher than that of large-cap stock.”
Hypothesis testing involves collecting and examining a representative sample to verify the
accuracy of a hypothesis. Hypothesis tests help analysts to answer questions such as:
Whenever a statistical test is being performed, the following procedure is generally considered
ideal:
2. Selection of the appropriate test statistic, i.e., what's being tested, e.g., the population
4. A clear statement of the decision rule to guide the choice of whether to reject or approve
179
© 2014-2024 AnalystPrep.
6. Arrival at a decision based on the sample results.
The null hypothesis, denoted as H0 , signifies the existing knowledge regarding the population
parameter under examination, essentially representing the "status quo." For example, when the
U.S. Food and Drug Administration inspects a cooking oil manufacturing plant to confirm that the
cholesterol content in 1 kg oil packages doesn't exceed 0.15%, they might create a hypothesis
like:
A test would then be carried out to confirm or reject the null hypothesis.
H 0 : μ = μ0
H 0 : μ ≤ μ0
H 0 : μ ≥ μ0
Where:
The alternative hypothesis, denoted as H a , contradicts the null hypothesis. Therefore, rejecting
the H0 makes H a valid. We accept the alternative hypothesis when the “status quo” is discredited
Using our FDA example above, the alternative hypothesis would be:
180
© 2014-2024 AnalystPrep.
One-tailed Test
A one-tailed test (one-sided test) is a statistical test that considers a change in only one direction.
In such a test, the alternative hypothesis either has a < (less than sign) or > (greater than sign),
A one-tailed test directs all the significance levels (α) to test statistical significance in one
direction. In other words, we aim to test the possibility of a change in one direction and
If we have a 5% significance level, we shall allot 0.05 of the total area in one tail of the
Let us assume we are using the standardized normal distribution to test the hypothesis that the
population mean equals a given value X . Further, let us assume we are using data from a sample
drawn from the population of interest. Our null hypothesis can be expressed as:
H0 : μ = X
If our test is one-tailed, the alternative hypothesis will test if the mean is either significantly
Ha : μ < X
The mean is significantly less than X if the test statistic is in the bottom 5% of the probability
distribution. This bottom area is known as the critical region (rejection region). We will reject the
181
© 2014-2024 AnalystPrep.
Case 2: Still at the 95% Confidence Level
Ha 1 : μ > X
We would reject the null hypothesis only if the test statistic is greater than the upper 5% point of
the distribution. In other words, we would reject H 0 if the test statistic is greater than 1.645.
182
© 2014-2024 AnalystPrep.
A Two-tailed Test
A two-tailed test considers the possibility of a change in either direction. It looks for a statistical
relationship in both a distribution's positive and negative directions. Therefore, it allows half the
value of α to test statistical significance in one direction and the other half to test the same in the
opposite direction. A two-tailed test may have the following set of hypotheses:
H0 : μ = X
H1 : μ ≠ X
Refer to our earlier example. If we were to carry out a two-tailed test, we would reject H0 if the
test statistic turned out to be less than the lower 2.5% point or greater than the upper 2.5%
183
© 2014-2024 AnalystPrep.
Step 2: Identify the Appropriate Test Statistic and Distribution
Test Statistic
A test statistic is a standardized value computed from sample information when testing
hypotheses. It compares the given data with what an analyst would expect under a null
hypothesis. As such, the null hypothesis is a major determinant of the decision to accept or reject
We use test statistics to gauge the degree of agreement between sample data and the null
hypothesis. Analysts use the following formula when calculating the test statistic for most tests:
The test statistic is a random variable that varies with each sample. The table below provides an
overview of commonly used test statistics, depending on the presumed data distribution:
184
© 2014-2024 AnalystPrep.
Hypothesis Test Test Statistic
Z-test Z- statistic (Normal distribution)
Chi-Square Test Chi-square statistic
t-test t-statistics
ANOVA F-statistic
We can subdivide the set of values that the test statistic can take into two regions: The non-
rejection region, which is consistent with the H 0, and the rejection region (critical region), which
is inconsistent with the H 0. If the test statistic has a value found within the critical region, we
reject the H 0.
As is the case with any other statistic, the distribution of the test statistic must be completely
Means
d̄−μ d0
Mean of t= t − distribution n −1
s d̄
Differences
s2 ( n −1)
Single χ2 = Chi-square n −1
σ2
0
Variance Distribution
S12
Difference in F= F-distribution n1 − 1,n 2 − 1
S22
variances
r√n −2
Correlation t= t-distribution n −2
√1−r2
2
( O ij−Eij)
Independence χ2 = ∑ m
i=1
Chi-square (r − 1) (c − 1)
Eij
Where:
μ0 , μd 0 , and σ02 denote hypothesized values of the mean, mean difference, and variance in that
185
© 2014-2024 AnalystPrep.
order.
X̄ , b̄ s2 , s and r denote the sample mean of the differences, sample variance, sample standard
Oij and E ij are observed and expected frequencies, respectively, with r indicating the number of
The significance level represents the amount of sample proof needed to reject the null
When using sample statistics to draw conclusions about an entire population, the sample might
not accurately represent the population. This can result in statistical tests giving incorrect
results, leading to either erroneous rejection or acceptance of the null hypothesis. This
Type I Error
Type I error occurs when we reject a true null hypothesis. For example, a type I error would
Type II Error
Type II error occurs when we fail to reject a false null hypothesis. In such a scenario, the
evidence the test provides is insufficient and, as such, cannot justify the rejection of the null
186
© 2014-2024 AnalystPrep.
Decision True Null Hypothesis False Null Hypothesis
(H0 ) (H 0 )
Fail to reject the Correct decision Type II error
null hypothesis
Reject null Type I error Correct decision
hypothesis
The level of significance, denoted by α, represents the probability of making a type I error, i.e.,
rejecting the null hypothesis when it is true. The confidence level complements the significance
level, (1 − α).
We use α to determine critical values that subdivide a distribution into the rejection and the non-
rejection regions. The figure below gives an example of the critical regions under a two-tailed
Consequently, β, the direct opposite of α, is the probability of making a type II error within the
bounds of statistical testing. The ideal but practically impossible statistical test would be one
187
© 2014-2024 AnalystPrep.
The Power of a Test
The power of a test is the direct opposite of the significance level. The level of significance gives
us the probability of rejecting the null hypothesis when it is, in fact, true. On the other hand, the
power of a test gives us the probability of correctly discrediting and rejecting the null hypothesis
when it is false. In other words, it gives the likelihood of rejecting H 0 when, indeed, it is false.
Expressed mathematically,
In a scenario with multiple test results for the same purpose, the test with the highest power is
The decision rule is the procedure that analysts and researchers follow when deciding whether
to reject or not reject a null hypothesis. We use the phrase "not to reject" because it's statistically
incorrect to "accept" a null hypothesis. Instead, we can only gather enough evidence to support
it.
The decision to reject or not reject a null hypothesis relies on the distribution of the test statistic.
The decision rule compares the calculated test statistic to the critical value.
If we reject the null hypothesis, the test is considered statistically significant. If not, we fail to
If a variable follows a normal distribution, we use the test's significance level to find critical
values corresponding to specific points on the standard normal distribution. These critical values
guide the decision-making process for rejecting or not rejecting a null hypothesis.
Before deciding whether to reject or not reject a null hypothesis, it's crucial to determine
188
© 2014-2024 AnalystPrep.
whether the test should be one-tailed or two-tailed. This choice depends on the nature of the
research question and the direction of the expected effect. Notably, the number of tails
determines the value of α (significance level). The following is a summary of the decision rules
H a : Parameter < X
Decision rule: Reject H 0 if the test statistic is less than the critical value. Otherwise, do not
reject H 0 .
H a : Parameter > X
Decision rule: Reject H 0 if the test statistic exceeds the critical value. Otherwise, do not reject
H 0.
189
© 2014-2024 AnalystPrep.
Two-tailed Test
Decision rule: Reject H0 if the test statistic is greater than the upper critical value or less than
190
© 2014-2024 AnalystPrep.
The p-Value in Hypothesis Testing
The p-value is the lowest level of significance at which we can reject a null hypothesis. The
probability of generating a test statistic would justify our rejection of a null hypothesis, assuming
When carrying out a statistical test with a fixed significance level (?) value, we merely compare
the observed test statistic with some critical value. For example, we might “reject an H 0 using a
5% test” or “reject an H 0 at a 1% significance level.” The problem with this ‘classical’ approach
is that it does not give us details about the strength of the evidence against the null
hypothesis.
testing. The p-value is the lowest level at which we can reject an H 0. This means that the
In one-tailed tests, the p-value is the probability below the calculated test statistic for left-tailed
tests or above the test statistic for right-tailed tests. For two-tailed tests, we find the probability
below the negative test statistic and add it to the probability above the positive test statistic. This
Example: p-value
θ represents the probability of obtaining a head when a coin is tossed. Assume we tossed a coin
200 times, and the head came up in 85 out of the 200 trials. Test the following hypothesis at a
5% level of significance.
H 0 : θ = 0.5
H 1 : θ < 0.5
Solution
191
© 2014-2024 AnalystPrep.
Our p-value will be given by P (X < 85), where X follows a binomial (200,0.5), assuming the H0 is
true.
(85– 100)
= P ⌈Z < ⌉
√50
= P (Z < −2.12) = 1 − 0.9834 = 0.01660
(We have applied the Central Limit Theorem by taking the binomial distribution as approximately
normal.)
Since the probability is less than 0.05, the H0 is extremely unlikely, and we have strong evidence
against an H 0 that favors H1 . Therefore, clearly expressing this result, we could say:
“There is very strong evidence against the hypothesis that the coin is fair. We, therefore,
Remember, failure to reject a H0 does not mean it is true. It means there is insufficient evidence
192
© 2014-2024 AnalystPrep.
Question
A CFA candidate conducts a statistical test about the mean value of a random
variable X .
H 0 : μ = μ0 vs H1 : μ ≠ μ 0
She obtains a test statistic of 2.2. Given a 5% significance level, determine the p-
value.
A. 1.39%.
B. 2.78.
C. 2.78%.
Solution
193
© 2014-2024 AnalystPrep.
Interpretation: The p-value (2.78%) is less than the significance level (5%). Therefore,
we have sufficient evidence to reject the H0 . In fact, the evidence is so strong that we
would also reject the H 0 at significance levels of 4% and 3%. However, at significance
levels of 2% or 1%, we would not reject the H0 since the p-value surpasses these
values.
194
© 2014-2024 AnalystPrep.
LOS 8b: Construct hypothesis tests and determine their statistical
significance, the associated Type I and Type II errors, and the power of
the test given a significance level
The z-test is the ideal hypothesis test when the sample's sampling distribution is normally
Given a random sample of size n from a normally distributed population with mean μ, variance
(X̄ − μ0 )
z − statistic =
σ
( )
√n
Where:
Once computed, the z-statistic is compared to the critical value corresponding to the test's
significance level. For example, if the significance level is 5%, the z-statistic is screened against
the upper or lower 95% point of the normal distribution (±1.96). The decision rule is to reject the
Example: z-test
195
© 2014-2024 AnalystPrep.
Academics carried out a study on 50 former United States presidents and found an average IQ of
135. You are required to carry out a 5% statistical test to determine whether the average IQ of
presidents is greater than 130. (IQs are distributed normally, and previous studies indicate that
σ = 25.)
Solution
H 0 : μ ≤ 130
H 1 : μ > 130
(X̄ − μ0 )
z − statistic =
σ
( )
√n
( X̄−130)
Assuming the H 0 is true, ∼ N (0, 1)
( σ )
√n
This is a right-tailed test. Therefore, we compare our test statistic to the upper 95% point of the
(135– 130)
The z-statistic is = 1.414
( 25 )
√50
Since 1.414 is less than 1.645, we do not have sufficient evidence to reject the H 0 . As such, it
would be reasonable to conclude that the average IQ of U.S. presidents is not more than 130.
196
© 2014-2024 AnalystPrep.
The t-test
The t-test is based on the t-distribution. The test is appropriate for testing the value of a
σ is unknown.
The sample size is large (n ? 30), and if n < 30, the distribution must be normal or
approximately normal.
(X̄ − μ 0 )
tn−1 =
s
( )
√n
Where:
Example: t-Test
Financial analysts in a certain equatorial country are interested in evaluating the potential
impact of rainfall on agricultural investments. They have gathered data on the annual rate of
Previously, the recorded average rainfall was 23 cm. The analysts want to find out if there's been
an increase in the average rainfall rate, which could impact agricultural investment. Conduct a
197
© 2014-2024 AnalystPrep.
statistical test at a 5% significance level to investigate this.
Solution
H 0 : μ ≤ 23
H 1 : μ > 23
If we assume that the annual rainfall quantities are distributed normally and recorded
independently, then:
(X̄ − μ0 )
∼ tn−1
s
( )
√n
α = 5% (right − tailed)
Reject the null hypothesis if the t-statistic is greater than t0.05 ,9 = 1.833
(26.6 − 23)
Therefore, our t-statistic = = 7.96
( 1.43)
√10
Our test statistic (7.96) is greater than the upper 95% point of the t0.05 ,9 distribution (1.833).
198
© 2014-2024 AnalystPrep.
Therefore, we have sufficient evidence to reject the H0. As such, it is reasonable to conclude
that the average annual rainfall has increased from its former long-term average of 23.
199
© 2014-2024 AnalystPrep.
Question
What is the value of t in the example above if the significance level is reduced from
Solution
A quick glance at the t0.005,9 distribution when α = 0.5% gives a value of 3.25.
However, the evidence against the H 0 is overwhelming since our test statistic (7.96)
is still greater than 3.25. As such, the conclusion would remain unchanged.
Analysts are often interested in establishing whether there exists a significant difference
between the means of two different populations. For instance, they might want to know whether
200
© 2014-2024 AnalystPrep.
the average returns for two subsidiaries of a given company exhibit a significant variance. Such
a test may then be used to make decisions regarding resource allocation or the reward of the
directors. Before embarking on such an exercise, ensuring that the samples taken are
independent and sourced from normally distributed populations is paramount. It can either be
assumed that the population variances are equal or unequal. In this reading, we will assume that
Assume that μ 1 is the mean of the first population while μ2 is the mean of the second population.
In testing the equality of two population means, we wish to determine whether they are equal. As
I. Two-sided:
H0 : μ 1 − μ 2 = 0 vs. H a ?μ 1 − μ 2 ≠ 0,
H 0 : μ1 = μ 2 vs. H a?μ1 ≠ μ 2
H0 : μ 1 − μ 2 ≤ 0 vs. H a ?μ 1 − μ 2 > 0,
H0 : μ 1 − μ 2 ≥ 0 vs. H a ?μ 1 − μ 2 < 0,
201
© 2014-2024 AnalystPrep.
H 0 : μ1 ≥ μ 2 vs. H a?μ1 < μ 2
However, note that tests such as H 0 : μ1 − μ 2 = 3 vs. H a?μ1 − μ 2 ≠ 3 are valid. The
When testing for the difference between two population means, we assume that the two
populations are distributed normally. Further, we assume that they have equal and unknown
variances. We always make use of the student's t-distribution where the test statistic is given by:
Where s 2p is the pooled estimator of the common variance and is given by:
202
© 2014-2024 AnalystPrep.
Nutritionists want to establish whether obese patients on a new special diet have a lower weight
than the control group. After six weeks, the average weight of 10 patients (group A) on the
special diet is 75kg, while that of 10 more patients of the control group (B) is 72kg. Carry out a
5% test to determine if the patients on the special diet have a lower weight.
Solution
H 0 : μ1 − μ 2 = 0 Vs Ha ?μ 1 − μ2 ≠ 0,
We assume that the two samples have equal variances, are independent, and are normally
B̄ − Ā
∼ tm+n −2
S√ m1 + 1
n
2
∑X 2 − n X̄
2
s =
n− 1
So,
{59520 − (10 ∗ 75 2 )}
SA2 = = 363.33
9
{56430 − (10 ∗ 72 2 )}
SB2 = = 510
9
Therefore,
(9 × 363.33 + 9 × 510)
S2p = = 436.665
(10 + 10 − 2)
And
(75 − 72)
203
© 2014-2024 AnalystPrep.
(75 − 72)
Test statistic = = 0.3210
1 1
{√439.665 × √ ( 10 + )}
10
Our test statistic (0.3210) is less than the upper 5% point (1.734) of the t-distribution with 18
degrees of freedom.
reasonable to conclude that the special diet has the same effect on body weight as the placebo.
Note to candidates: You could choose to work with the p-value and determine P (t18 > 0.937)
and then establish whether this probability is less than 0.05. Working out the problem this way
204
© 2014-2024 AnalystPrep.
Suppose we replace ‘>’ with ‘≠’ in H1 in the example above, would the decision rule change?
Replacing ‘>’ with ‘≠’ in H1 would change the test from a one-tailed one to a two-tailed test. We
would compute the test statistic just as demonstrated above. However, we would have to divide
the significance level by two and compare the test statistic to the lower and upper 2.5% points of
Since our test statistic lies within these limits (non-rejection region), the decision rule would
remain unchanged.
There are some challenges when testing the difference between two population means using
205
© 2014-2024 AnalystPrep.
independent samples. Variability within each sample, caused by factors unrelated to the
research, can obscure the real difference of interest. Random variation within a sample might be
so substantial that it obscures the actual difference caused by the specific phenomenon the
analyst is studying.
When we want to test the differences between means with dependent samples, we use the paired
Assume that we have observations for the random variables XA and XB and that the samples are
dependent.
Organize the observations in pairs and denote the differences between the two paired
1, 2 ,… , n.
Also, let μ d be the population mean difference and μd0 be the hypothesized value for the
Practically, μd0 = 0.
d̄ − μ d0
t=
sd̄
Where:
206
© 2014-2024 AnalystPrep.
1 n
d̄ = ∑ d
n i=1 i
sd
s d̄ = sd̄ = = standard error of the mean differences
√n
sd = standard deviation of the differences
Note that the degree of freedom is n − 1 where n is the number of the paired observations.
An analyst aims to compare the performance of the BCD High Growth Index and the BCD
Investment Grade Index. They collect data for both indexes over 2,050 days and calculate the
Using a 5% significance level, determine whether the mean of the differences is different from
zero.
Solution
H0 : μ d0 = 0 vs Ha : μd0 ≠ 0
d̄ − μ d0
t=
sd̄
α = 5% (two-tailed test)
207
© 2014-2024 AnalystPrep.
Step 4: State the decision rule:
The degrees of freedom amount to n − 1 = 2 , 050 − 1 = 2 , 049; thus, the critical values are
±1.960 . Therefore, we will reject the null hypothesis if the calculated t-statistic is less than -1.96
Note that from the table d̄ = −0.0022 and sd = 0.3321 so that the t-statistic is given:
d̄ − μd0 −0.0022 − 0
t= = = −0.30
s d̄ 0.3321
√2050
-0.30 falls within the bounds of the critical values of ±1.960 . As such, there is insufficient
evidence to show that the mean of the differences in returns differs from zero.
208
© 2014-2024 AnalystPrep.
Testing of a Single Variance
A chi-square test helps determine if a hypothesized variance value matches the true population
variance. Unlike other distributions in the CFA® Program, the chi-square distribution is
As a natural consequence, the chi-square distribution has no negative values and is bounded by
(n − 1) S 2
χ2n −1 =
σ02
Where:
n = Sample size.
S 2 = Sample variance.
209
© 2014-2024 AnalystPrep.
Example: Chi-square Test
For the 15-year period between 1995 and 2010, ABC's monthly return had a standard deviation
of 5%. John Matthew, CFA, wishes to establish whether the standard deviation witnessed during
that period still adequately describes the long-term standard deviation of the company's return.
To achieve this end, he collects data on the monthly returns recorded between January 1, 2015,
and December 31, 2016, and computes a monthly standard deviation of 4%.
Carry out a 5% test to determine if the standard deviation computed in the latter period differs
Solution
H 0 : σ02 = 0.0025
H 1 : σ2 ≠ 0.0025
Since the latter period has 24 months, n = 24 , the test statistic is:
(24 − 1) 0.0016
χ 224−1 = = 14.72
0.0025
This is a two-tailed test. As such, we have to divide the significance level by two and screen our
test statistic against the lower and upper 2.5% points of χ223 .
Consulting the chi-square table, the test statistic (14.72) lies between the lower (11.689) and the
210
© 2014-2024 AnalystPrep.
Note that you will be given a simplified critical value table in the exam situation.
211
© 2014-2024 AnalystPrep.
Evidently, we have insufficient evidence to reject the H 0. It is, therefore, reasonable to conclude
that the latter standard deviation value is close enough to the 15-year value.
To test the equality concerning the variances, we use the F-test. Assume that we have 2
independent random samples of sizes n1 and n2 from N(μ 1 , σ12 ) and N(μ 2 , σ22 ).
Also, let us consider a scenario where we have the sample variances as S12 and S22. The basic
H 0 : σ12 = σ22
H a : σ12 ≠ σ22
S 12
The test statistic is ∼ Fn 1–1,n 2–1 under H 0.
2
S2
212
© 2014-2024 AnalystPrep.
The decision rule is to reject the null hypothesis if the test statistic falls within the critical region
of the F-distribution.
Example: F-test
An analyst is studying whether the population variance of returns on a commodity index changed
after the introduction of new trading guidelines. The first 320 weeks elapsed before the
guidelines were introduced, and the second 320 weeks came after the introduction. The analyst
gathers the data in the table below for 320 weeks of returns before and after the change in
guidelines.
Do the variances of returns differ before and after the guideline change? Employ a 5 percent
significance level.
213
© 2014-2024 AnalystPrep.
Solution
2
H 0 : σBefore 2
= σAfter
2 vs Ha : σBefore ≠ σAfter
2
s2Before
t=
s2After
α = 5% (two-tailed)
Reject the null if the calculated t-statistic is less than 0.803 and reject the null if the calculated t-
s2Before 3.520
t= = = 1.1864
s2After 2.967
Fail to reject the null hypothesis because 1.1864 falls within the bounds of the critical values of
[0.80,1.246]. There is insufficient evidence to indicate that the weekly variances of returns are
214
© 2014-2024 AnalystPrep.
LOS 8c: Compare and contrast parametric and nonparametric tests, and
describe situations where each is the more appropriate type of test
Parametric Tests
Parametric tests are statistical tests in which we make assumptions regarding population
distribution. Such tests involve the estimation of the key parameters of a distribution. For
When conducting statistical tests, the choice of distribution directly influences how the test
statistic is calculated. The tests we've discussed are considered parametric tests. For example,
assuming a parameter follows a normal distribution leads to the computation of the z-statistic.
During parametric testing, approximating normal distributions for non-normal data may be
required. This approximation is valid due to the central limit theorem, which states that as
Nonparametric Tests
Nonparametric tests, sometimes known as distribution-free tests, don't assume anything about
the parameter's distribution being studied. Researchers turn to nonparametric testing when they
The following table gives the alternative nonparametric tests for the parametric tests.
215
© 2014-2024 AnalystPrep.
Situations Where Nonparametric Tests are Appropriate
This happens when the distributional assumptions of the parametric tests are not met. For
instance, we may find parametric tests such as t-test are inappropriate because the sample size
is small and may be drawn from non-normally distributed. As such, a nonparametric test is
appropriate.
Outliers can affect the parametric statistics. On the other hand, outliers do not affect parametric
tests.
Consider a situation where we want to establish the center of a rather skewed distribution, such
as that of the income of the residents of a given city. While the majority of the residents could be
categorized as the middle class, the presence of just a few billionaires in a sample can greatly
increase the mean income. Such a mean, therefore, may not provide a very reliable or realistic
measure of income.
Instead, it may be more appropriate to use the median. Compared to the mean, the median can
better represent the center of the income distribution. This is due to the fact that 50% of the
residents will be above the median and the remaining 50% below it.
In summary, “outliers” affect the mean when dealing with skewed data. The median, on the other
Although nonparametric tests are usually easier to conduct than parametric ones, they do not
have as much statistical power. Nonetheless, they provide an efficient tool for analyzing ordinal,
216
© 2014-2024 AnalystPrep.
We often use nonparametric tests, such as the runs test, when our goal is to determine if a
sample from a population isn't random. Since randomness isn't a parameter, nonparametric tests
Nonparametric inference:
Nonparametric methods make our statistical analysis broader. They work with limited
assumptions and can be used for ordered data. Plus, they handle questions that aren't tied to
specific parameters.
Nonparametric tests are commonly used alongside parametric tests. They help analysts
understand how sensitive the statistical results are to the assumptions of parametric tests. But
when the conditions for a parametric test are met, we usually choose it over nonparametric tests.
We prefer parametric tests because they often have more statistical power, which means they are
217
© 2014-2024 AnalystPrep.
Learning Module 9: Parametric and Non Parametric Tests of
Independence
A parametric test is a hypothesis test concerning a population parameter used when the data has
specific distribution assumptions. If these assumptions are not met, non-parametric tests
are used.
We frequently compare the population correlation coefficient to zero when testing for
correlation. This helps us determine whether there's a relationship between the variables. The
population correlation coefficient, represented by ρ, is used to test the relationship. There are
Two-sided; H0 : ρ = 0 versus H a : ρ ≠ 0.
Let's assume that we have variables X and Y. The sample correlation, rXY , tests the above
218
© 2014-2024 AnalystPrep.
hypotheses.
used to test the correlation in a parametric test. The formula for the sample correlation involves
the sample covariance between the X and Y variables and their respective standard deviations,
SXY
r=
SX SY
Where:
A t-test can determine if the null hypothesis should be rejected using the sample correlation, r if
the two variables are normally distributed. The formula for the t-test is:
r√n − 2
t=
√(1 − r2 )
Where:
r= Sample correlation.
n= Sample size.
The test statistic follows a t-distribution with n − 2 degrees of freedom. From the equation above,
it is easy to see that the sample size, n, increases, and the degrees of freedom increase. In other
words, as the sample size n increases, the power of the test increases. This implies that a false
219
© 2014-2024 AnalystPrep.
null hypothesis will likely be rejected as the sample size increases.
The table below shows the sample correlations between the monthly returns of five different
sector-specific exchange-traded funds (ETFs) and the overall market index (Market 1). There are
48 monthly observations, and the following ETFs are included in the analysis:
the t-statistic for the correlation between ETF 2 and ETF 4. Based on the calculated t-statistic,
draw a conclusion about the significance of the correlation using the following sample t-table:
Solution
220
© 2014-2024 AnalystPrep.
To test the significance of the correlation between ETF 2 and ETF 4, we will use the t-test
formula:
r√n − 2
t=
√(1 − r2 )
Where:
r√n − 2 0.5789√48 − 2
t= = = 4.815
√(1 − r2 ) √1 − 0.57892
The calculated t-statistic for the correlation between ETF2 and ETF4 is 4.815.
Conclusion: We reject the null hypothesis since our calculated t-statistic (4.815) is greater than
the critical value (+2.687). This indicates sufficient evidence to suggest that the correlation
The Spearman rank correlation coefficient, rS , is a non-parametric test used to examine the
relationship between two data sets when the population deviates from normality.
The Spearman rank correlation coefficient is like the Pearson correlation coefficient. The
difference is that the Spearman coefficient is calculated based on the ranks of variables in the
samples.
221
© 2014-2024 AnalystPrep.
Consider two variables, X and Y . We need to calculate Spearman's Rank Correlation rS.
Rank the observations of each variable X and Y in descending order. Note that when
there are tied values in the data, their ranks are calculated by taking the average of the
ranks that would have been assigned to those values if they were not tied.
Find the difference between the ranks for each pair of observations.
Square the difference and calculate the sum of the difference, that is, ∑ di .
6 ∑ni=1 d i2
rs = 1 −
n (n 2 − 1)
Where; di =The difference between the ranks for each pair of observationsn= Sample
size.
An analyst is studying the relationship between returns for two sectors, steel and cement, over
the past 5 years using Spearman's rank correlation coefficient. The hypotheses are H 0 : rS = 0
Solution
222
© 2014-2024 AnalystPrep.
Year Steel Cement Rank Rank D d2
sector sector order order
returns returns for X for Y
(X) (Y)
1 10% 8% 2 2 0 0
2 6% 7% 5 3 2 4
3 9% 5% 3 5 −2 4
4 12% 6% 1 4 −3 9
5 8% 9% 4 1 3 9
Sum = 26
6 ∑ ni=1 d i2 (6 × 26)
rs = 1 − = 1− [ ] = 1 − 1.3
n (n 2 − 1) 5 × (5 2 − 1)
rs = −0.3
This indicates a very weak negative correlation between the returns of the steel and cement
sectors.
The hypothesis test on the Spearman Rank depends on the sample size. If the sample size is
small (n ≤ 30), we would need a specialized table of critical value. On the other hand, if the
sample size is large (n > 30), we can perform a t-test using the test statistic similar to that of
Pearson correlation:
rs √n − 2
t=
√(1 − r2s )
Consider the above example. Assume we want to conduct a hypothesis test at a 5% significance
223
© 2014-2024 AnalystPrep.
Question
Assume an investment analyst, John Smith, is studying the relationship between two
stocks, X and Y . Based on 100 observations, he has found that SXY = 10, SX = 2, and
SY = 8. Smith needs to find the sample correlation rXY and use it to perform a t-test
Y . The critical value for the test statistic at the 0.05 level of significance is
and Y is:
A. Significant because the test statistic falls outside the range of the critical
values.
B. Significant, because the absolute value of the test statistic is less than the
critical value.
C. Insignificant because the test statistic falls outside the range of the critical
values.
Solution
Note that the sample correlation coefficient, rXY is calculated using the following
formula:
SXY
rXY =
SX SY
10
rXY = = 0.625
2× 8
To test the significance of the sample correlation, we can use a t-test with the
224
© 2014-2024 AnalystPrep.
The test statistic for this test is calculated using the following formula:
r√ n − 2
t=
√1 − r2
Where:
The critical value for the test statistic at the 0.05 level of significance is
approximately 1.96.
Since our calculated test statistic (7.9262) is greater than the upper bound of the
critical values for the test statistic (1.96), we reject the null hypothesis. This indicates
Therefore, John Smith should conclude that the statistical relationship between X and
Y is significant because the test statistic falls outside the range of the critical values
(Option A).
225
© 2014-2024 AnalystPrep.
LOS 9b: Explain tests of independence based on contingency table data
With categorical or discrete data, correlation is not suitable for assessing relationships between
variables. Instead, we use a non-parametric test called the chi-square test of independence,
We employ a contingency table to structure the data when examining the connection between
distribution to assess whether a noteworthy relationship exists between these variables. The test
2
m (Oij − Eij )
χ=∑
i=1 (E ij )
Where:
m= Number of cells in the table, the Number of groups in the first class, multiplied by the
Oij = Number of observations in each cell of row i and column j (i.e., observed frequency).
Ei j= Expected number of observations in each cell of row i and column j, assuming independence
Where:
r= Number of rows.
c= Number of columns.
226
© 2014-2024 AnalystPrep.
The following contingency table shows the responses of two categories of investors (employed
vs. retired) with regard to their primary investment objectives (growth, income, or both). The
Use a 95% significance level to test whether there is any significant difference between
Solution
H 0: There is no significant difference between employed and retired investors with regard to
H α: There is a significant difference between employed and retired investors with regard to
Step 1: We calculate the expected frequency of investors by their category (employed vs.
Step 2: We calculate the scaled squared deviation for each combination of investor category and
227
© 2014-2024 AnalystPrep.
Growth Income Both
(52−42) 2 (25−36) 2 (10−9)2
Employed 42
= 2.254 36
= 0.469 9
= 0.246
(32−42) 2 (47×36) 2 (7−8) 2
Retired 42
= 2.280 36
= 3.510 8
= 0.349
Total 4.534 6.979 0.495
228
© 2014-2024 AnalystPrep.
Decision rule: The calculated value of χ2 = 12.008 is greater than the critical value of 5.99. As
such, sufficient evidence supports the conclusion that retired and employed investors have
229
© 2014-2024 AnalystPrep.
Question
Regarding the chi-square test of independence, which statement is accurate? The chi-
B. Used to test whether two categorical variables are related to each other.
C. Used to test whether two continuous variables are related to each other.
Solution
hypothesis test that can be used to test whether two categorical variables are related.
parametric.
230
© 2014-2024 AnalystPrep.
Learning Module 10: Simple Linear Regression
LOS 10a: describe a simple linear regression model, how the least
squares criterion is used to estimate regression coefficients, and the
interpretation of these coefficients
Linear regression is a mathematical method used for analyzing how the variation in one variable
Let Y be the variable we wish to explain. As such, the observation of this variable is Yi , and Y¯ is
n
2
Variation of Y = ∑ (Y i − Y¯ )
i=1
Our main objective is to explain what causes this variation, usually called the sum of squares
total (SST).
By definition of the regression, we need to explain the variation of Y with another variable. Let X
be the explanatory variable. As such, the observations of X will be denoted by Xi and X̄ sample
n
2
Variation of X = ∑ (Xi − X̄ )
i=1
To visualize the relationship between variables X and Y, you can use a scatter plot, also known as
a scattergram. In this type of plot, the variable you want to explain (Y) is usually plotted on the
vertical axis. In contrast, the explanatory variable (X) is placed on the horizontal axis to show the
For example, consider the following table. We wish to use linear regression analysis to forecast
231
© 2014-2024 AnalystPrep.
Year Unemployment Rate Inflation Rate
2011 6.1% 1.7%
2012 7.4% 1.2%
2013 6.2% 1.3%
2014 6.2% 1.3%
2015 5.7% 1.4%
2016 5.0% 1.8%
2017 4.2% 3.3%
2018 4.2% 3.1%
2019 4.0% 4.7%
2020 3.9% 3.6%
In this scenario, the Y variable is the inflation rate, and the X axis is the unemployment rate. A
scatter plot of the inflation rates against unemployment rates from 2011 to 2020 is shown in the
following figure.
A dependent variable, often denoted as YYY, is the variable we want to explain. In contrast, an
independent variable, typically denoted as XXX, explains variations in the dependent variable.
232
© 2014-2024 AnalystPrep.
The independent variable is also referred to as the exogenous, explanatory, or predicting
variable.
In our example above, the inflation rate is the dependent variable, and the unemployment rate is
linear relationship, usually a straight line. When there's one independent variable, we use simple
linear regression. If there are multiple independent variables, we use multiple regression.
In simple linear regression, we assume linear relationships exist between the dependent and
independent variables. The aim is to fit a line to the observations of X (Xi s) and Y (Y is) to
minimize the squared deviations from the line. To accomplish this, we use the least squares
criterion.
Y = b 0 + b 1 X1 + εi, i = 1, 2, … , n
Where:
Y = Dependent variable.
b0 = Intercept.
b1 = Slope coefficient.
X = Independent variable.
b0 and b 1 are known as regression coefficients. The equation above implies that the dependent
233
© 2014-2024 AnalystPrep.
is equivalent to the intercept (b0 ) plus the product of the slope coefficient (b 1 ) and the
The error term is equal to the difference between the observed value of Y and the one expected
As stated earlier, linear regression calculates a line that best fits the observations. In the
following image, the line that best fits the regression is clearly the blue one:
234
© 2014-2024 AnalystPrep.
Note that we cannot directly observe the population parameters b0 and b 1. As such, we observe
their estimates, ^
b 0 and ^
b 1 . They are the estimated parameters of the population using a sample.
minimized.
Specifically, we concentrate on the sum of the squared differences between observations Yi and
the respective estimated value Y^i on the regression line, also called the sum of squares error
(SSE).
Note that,
Y^i = ^
b 0 +^
b 1 Xi + e2i
As such,
235
© 2014-2024 AnalystPrep.
n 2 2 n n
SSE = ∑ (Y i − ^
b 0 −^
b 1 Xi) = ∑ (Yi − Y^i ) = ∑ e2i
i =1 i =1 i=1
Note that the residual for the ith observation (ei = Y i − Y^i ) is different from the error term (εi ).
The error term is based on the underlying population, while the residual term results from
Conventionally, the sum of the residuals is zero. As such, the aim is to fit the regression line in a
simple linear regression that minimizes the sum of squared residual terms.
For a simple linear regression, the slope coefficient is estimated as the ratio of the Cov(X, Y ) and
V ar(X):
The slope coefficient is defined as the change in the dependent variable caused by a one-unit
^
b 0 = Y¯ − ^
b 1 X̄
Where:
Y^ = Mean of Y .
X^ = Mean of X.
236
© 2014-2024 AnalystPrep.
The intercept is the estimated value of the dependent variable when the independent variable is
zero. The fitted regression line passes through the point equivalent to the means of the
Let us consider the following table. We wish to estimate a regression line to forecast inflation,
2 2
Year Unemployment Inflation (Y i − Y¯ ) (Xi − X̄) (Yi − Y¯)
Rate% (Xi s) Rate% (Y i s) (Xi − X̄ )
2011 6.1 1.7 0.410 0.656 −0.518
2012 7.4 1.2 1.300 4.452 −2.405
2013 6.2 1.3 1.082 0.828 −0.946
2014 6.2 1.3 1.082 0.828 −0.946
2015 5.7 1.4 0.884 0.168 −0.385
2016 5.0 1.8 0.292 0.084 0.157
2017 4.2 3.3 0.922 1.188 −1.046
2018 4.2 3.1 0.578 1.188 −0.828
2019 4.0 4.7 5.570 1.664 −3.044
2020 3.9 3.6 1.588 1.932 −1.751
Sum 52.90 23.4 13.704 12.989 −11.716
Arithmetic 5.29 2.34
Mean
Cov (X Y ) ∑n (Y − Y¯ ) (X − X̄ )
237
© 2014-2024 AnalystPrep.
Cov (X , Y ) ∑ni=1 (Y i − Y¯ ) (Xi − X̄ ) −11.716
^
b1 = = = = −0.9020
V ar (X) ∑ni=1 (Xi − X̄ )
2 12.989
^
b 0 = Y¯ − ^
b 1X̄ = 2.34 − (−0.9020) × 5.29 = 7.112
Y^ = 7.112 − 0.9020Xi + εi
If the unemployment rate increases (decreases) by one unit, say, from 2% to 3%–the
In general,
Furthermore, with the estimated regression model, we can predict the values of the dependent
variable based on the value of the independent variable. For instance, if the unemployment rate
In practice, analysts perform regression analysis using statistical functions in software like
Regression analysis is commonly used with cross-sectional and time series data. In cross-
238
© 2014-2024 AnalystPrep.
sectional analysis, you compare X and Y observations from different entities, like various
companies in the same time period. For instance, you might analyze the link between a
company's R&D spending and stock returns across multiple firms in a year.
Time-series regression analysis involves using data from various time periods for the same entity,
like a company or an asset class. For instance, an analyst might examine how a company's
quarterly dividend payouts relate to its stock price over multiple years.
239
© 2014-2024 AnalystPrep.
Question
A. Predicted variable.
B. Predicting variable.
C. Endogenous variable.
Solution
or endogenous variable.
240
© 2014-2024 AnalystPrep.
LOS 10b: explain the assumptions underlying the simple linear
regression model, and describe how residuals and residual plots indicate
if these assumptions may have been violated
Assume that we have samples of size n for dependent variable Y and independent variable X . We
wish to estimate the simple regression of Y and X. The classic normal linear regression model
Linearity: A linear relationship implies that the change in Y due to a one-unit change in X is
constant, regardless of the value X takes. If the relationship between the two is not linear, the
regression model will not accurately capture the trend, resulting in inaccurate predictions. The
model will be biased and underestimate or overestimate Y at various points. For example, the
model Y = b 0 + b1 eb1x is nonlinear in b1 . For this reason, we should not attempt to fit a linear
model between X and Y . It also follows that the independent variable, X , must be non-stochastic
(must not be random). A random independent variable rules out a linear relationship between
the dependent and independent [Link] addition, linearity means the residuals should not
exhibit an observable pattern when plotted against the independent variable. Instead, they
should be completely random. In the example below, we're looking at a scenario where the
residuals appear to show a pattern when plotted against the independent variable, X . This
241
© 2014-2024 AnalystPrep.
Normality Assumption: This assumption implies that the error terms (residuals) must follow
a normal distribution. It's important to note that this doesn't mean the dependent and
independent variables must be normally distributed. However, it's crucial to check the
distribution of the dependent and independent variables to identify any outliers. A histogram of
the residuals can be used to detect if the error term is normally distributed. A symmetric bell-
Homoskedasticity: Homoskedasticity implies that the variance of the error terms is constant
E (ϵ 2i ) = σϵ2, i = 1 , 2 , … , n
If the variance of residuals varies across observations, then we refer to this as heteroskedasticity
242
© 2014-2024 AnalystPrep.
(not homoscedasticity). We plot the least square residuals against the independent variable to
test for heteroscedasticity. If there is an evident pattern in the plot, that is a manifestation of
heteroskedasticity.
In case residuals and the predicted values increase simultaneously, then such a situation is
243
© 2014-2024 AnalystPrep.
Independence Assumption: The independence assumption implies that the observations Xi
and Y i are independent of each other. Failure to satisfy this assumption implies the variables are
not independent, and thus, residuals will be correlated. To ascertain this assumption, we visually
and statistically analyze the residuals to check whether residual shows exhibit a pattern.
244
© 2014-2024 AnalystPrep.
Question
A regression model with one independent variable requires several assumptions for
valid conclusions. Which of the following statements most likely violates those
assumptions?
C. There exists a linear relationship between the dependent variable and the
independent variable.
Solution
Linear regression assumes that the independent variable, X, is NOT random. This
ensures that the model produces the correct estimates of the regression coefficients.
B is incorrect. The assumption that the error term is distributed normally allows us
variables have a linear relationship is the key to a valid linear regression. If the
parameters of the dependent and independent variables are not linear, then the
245
© 2014-2024 AnalystPrep.
LOS 10c: calculate and interpret measures of fit and formulate and
evaluate tests of fit and of regression coefficients in a simple linear
regression
The sum of Squares Total (total variation) is a measure of the total variation of the dependent
variable. It is the sum of the squared differences of the actual y-value and mean of y-
observations.
n
2
S ST = ∑ (Y i − Y¯ )
i=1
i. The sum of Squares Regression (SSR): The sum of squares regression measures the
explained variation in the dependent variable. It is given by the sum of the squared
n 2
S SR = ∑ (Y^i − Y¯ )
i=1
ii. The Sum of Squared Errors (SSE): The sum of squared errors is also called the
unexplained by the independent variable. SSE is given by the sum of the squared
n 2
SSE = ∑ (Y i − Y^i )
i=1
246
© 2014-2024 AnalystPrep.
Sum of Squares Total = Explained Variation + Unexplained Variation
= SSR + SSE
The components of the total variation are shown in the following figure.
For example, consider the following table. We wish to use linear regression analysis to forecast
247
© 2014-2024 AnalystPrep.
Year Unemployment Rate (%) Inflation Rate (%)
2011 6.1 1.7
2012 7.4 1.2
2013 6.2 1.3
2014 6.2 1.3
2015 5.7 1.4
2016 5.0 1.8
2017 4.2 3.3
2018 4.2 3.1
2019 4.0 4.7
2020 3.9 3.6
Remember that we had estimated the regression line to be Y^ = 7.112 − 0.9020Xi + εi . As such,
n
2
SST = ∑ (Y i − Y¯) = 13.704
i=1
n 2
SSR = ∑ (Y^ i − Y¯ ) = 10.568
i=1
n 2
SSE = ∑ (Y i − Y^i) = 3.136
i=1
248
© 2014-2024 AnalystPrep.
Measures of Goodness of Fit
We use the following measures to analyze the goodness of fit of simple linear regression:
I. Coefficient of determination.
Coefficient of Determination
The coefficient of determination (R 2 ) measures the proportion of the total variability of the
dependent variable explained by the independent variable. It is calculated using the formula
below:
Explained Variation
249
© 2014-2024 AnalystPrep.
Explained Variation
R2 =
Total Variation
Sum of Squares Regression (SSR)
=
Sum of Squares Total (SST)
10.568
= = 76.61%
13.794
R 2 lies between 0% and 100%. A high R 2 explains variability better than a low R2 . If R 2=1%, only
1% of the total variability can be explained. On the other hand, if R 2 =90%, over 90% of the total
variability can be explained. In a nutshell, the higher the R 2 , the higher the model's explanatory
power.
For simple linear regression (R 2 ) is calculated by squaring the correlation coefficient between
2
C ov (X, Y )
2 ∑ni=1 (Y^i − Y¯ )
r2 = R 2 = ( ) =
σX σY ∑ni=1 (Y i − Y¯ )
2
Where:
2 2
An analyst determines that (∑6i=1 (Yi − Y¯) = 13.704) and (∑6i=1 (Yi − Y^i ) = 3.136) from the
Solution
= 0.7712 = 77.12%
Note that the coefficient of determination discussed above is just a descriptive value. To check
the statistical significance of a regression model, we use the F-test, which requires us to
In simple linear regression, the F-test confirms whether the slope (denoted by (b 1 )) in a
regression model is equal to zero. In a typical simple linear regression hypothesis, the null
hypothesis is formulated as: (H 0 : b1 = 0) against the alternative hypothesis (H 1 : b1 ≠ 0). The null
hypothesis is rejected if the confidence interval at the desired significance level excludes zero.
The Sum of Squares Regression (SSR) and Sum of Squares Error (SSE) are employed to calculate
the F-statistic. In the calculation, the Sum of Squares Regression (SSR) and Sum of Squares
The Sum of Squares Regression(SSR) is divided by the number of independent variables (k) to
2
∑ni=1 ( Ŷi − Y¯ )
SSR
MSR = =
k k
Since we only have (k = 1), in a simple linear regression model, the above formula changes to:
2
∑ni=1 ( Ŷi − Y¯) n 2
SSR
MSR = = = ∑ (Y^i − Y¯)
1 1 i =1
Also, the Sum of Squares Error (SSE) is divided by degrees of freedom given by (n − k − 1) (this
251
© 2014-2024 AnalystPrep.
translates to (n − 2) for simple linear regression) to arrive at Mean Square Error (MSE). That is,
2
Sum of Squares Error (SSE)∑i=1 (Yi − Y^)
n
MS E = =
n −k− 1 n− k− 1
2
Sum of Squares Error(SSE)∑i=1 (Y i − Y^)
n
MSE = =
n −2 n −2
Finally, to calculate the F-statistic for the linear regression, we find the ratio of MSR to MSE.
That is,
2
∑n ¯
i=1 ( Ŷi − Y )
SS R
MSR k k
F − statistic = = =
MSE SS E
∑n=1 (Y i−Y^)
2
n−k−1 i
n−k−1
2
SS R ∑ni=1 ( Ŷi − Y¯ )
MSR k
F − statistic = = =
MSE SS E
∑n ^
2
i=1 (Y i− Y )
n−k−1
n −2
The F-statistic in simple linear regression is F-distributed with (1) and (n − 2) degrees of
MS R
∼ F1, n−2
MS E
Note that the F-test regression analysis is a one-side test, with the rejection region on the right
side. This is because the objective is to test whether the variation in Y explained (the numerator)
252
© 2014-2024 AnalystPrep.
A large F-statistic value proves that the regression model effectively explains the variation in the
dependent variable and vice versa. On the contrary, an F-statistic of 0 indicates that the
independent variable does not explain the variation in the dependent variable.
We reject the null hypothesis if the calculated value of the F-statistic is greater than the critical
F-value.
It is worth mentioning that F-statistics are not commonly used in regressions with one
independent variable. This is because the F-statistic is equal to the square of the t-statistic for
the slope coefficient, which implies the same thing as the t-test.
Standard Error of Estimate, Se or SEE, is alternatively referred to as the root mean square error
or standard error of the regression. It measures the distance between the observed dependent
variables and the dependent variables the regression model predicts. It is calculated as follows:
2
∑ ni=1 (Y i − Y^i )
Standard Error of Estimate (Se ) = √MSE = ⎷
n −2
The standard error of estimate, coefficient of determination, and F-statistic are the measures
that can be used to gauge the goodness of fit of a regression model. In other words, these
measures tell the extent to which a regression model syncs with data.
The smaller the Standard Error of Estimate is, the better the fit of the regression line. However,
the Standard Error of Estimate does not tell us how well the independent variable explains the
Note that the F-statistic discussed above is used to test whether the slope coefficient is
253
© 2014-2024 AnalystPrep.
significantly different from 0. However, we may also wish to test whether the population slope
differs from a specific value or is positive. To accomplish this, we use the t-distributed test.
H 0 : b1 = 0 versus Ha : b1 ≠ 0
H 0 : b1 ≤ 0 versus Ha : b1 > 0
2. Identify the appropriate test statistic: The test statistic for the t-distributed test on
^
b 1 − B1
t=
s^b1
freedom. Since we are dealing with simple linear regression, we will deal with n − 2
ratio of the standard error of estimate (se) and the square root of the variation of the
independent variable:
se
s^
b1 =
2
√∑ni (Xi − X̄ )
=1
Where:
se = √MSE
3. Specify the level of significance: Note the level of significance level, usually denoted
4. State the decision rule: Using the significance level, find the critical values. You can
programming languages such as Python. In an exam situation, such critical values will be
254
© 2014-2024 AnalystPrep.
provided. Compare the t-statistic value to the critical t-value (tc) . Reject the null
hypothesis if the absolute t-statistic value is greater than the upper critical t-value or less
than the lower critical value, i.e., t > +tcritical or t < −tcritical
5. Calculate the test statistic: Using the formula above, calculate the test statistic.
Intuitively, you might need to calculate the standard error of the slope coefficient (s^
b1
)
first.
6. Make a decision: Make a decision whether to reject or fail to reject the null hypothesis.
Recall the example where we regressed inflation rates against unemployment rates from 2011 to
2020.
Y^ = 7.112 − 0.9020Xi + εi
Assume that we need to test whether the slope coefficient of the unemployment rates is positive
at a 5% significance level.
H 0 : b1 < 0 versus Ha : b1 ≥ 0
255
© 2014-2024 AnalystPrep.
Next, we need to calculate the test statistic given by:
^
b 1 −B1
t= s^
b1
Where:
se
s^
b1 =
2
√∑ni (Xi − X̄ )
=1
Recall that,
2
∑ni=1 (Y i − Y^)
SSE 3.136
se = √ MSE = √ =⎷ =√ = 0.6261
n −k− 1 n−2 8
So that,
se 0.6261
s^
b1 = = = 0.1737
√∑ni 1 (Xi − X̄ )
2 √12.989
=
Therefore,
^
b 1 − B1 −0.9020 − 0
t= = = −5.193
s^b1
0.1737
Next, we need to find critical t-values. Note that this is a one-sided test. As such, we need to find
256
© 2014-2024 AnalystPrep.
From the table, t8 , 0.05 = 1.860. We fail to reject the null hypothesis since the calculated test
statistic is less than the critical t-value (?5.193 < 1.860). There is sufficient evidence to indicate
In simple linear regression, a distinct characteristic exists: the t-test statistic checks if the slope
coefficient equals zero. This t-test statistic is the same as the test-statistic used to determine if
257
© 2014-2024 AnalystPrep.
H 0 : b1 ≤ 0 versus Ha : ρ > 0 or H 0 : ρ > 0 versus H a : ρ ≤ 0 and H 0 : b 1 > 0 versus H a : ρ ≤ 0).
Note that the test -statistic to test whether the correlation is equal to zero is given by:
r√ n − 2
t=
√1 − r2
Consider our previous example, where we regressed inflation rates against unemployment rates
from 2011 to 2020. Assume we want to test whether the pairwise correlation between the
In the example, the correlation between unemployment and inflation rates is -0.8782. As such,
−0.8782√10 − 2
t= ≈ −5.19
2
√1 − (−0.8782)
Note this is equal to the test statistic t-test statistic used to perform the hypothesis test whether
^
b 1 − B1 −0.9020 − 0
t= = = −5.193
s^b1
0.1737
Similar to the slope coefficient, we may also want to test whether the population intercept equals
a certain value. The process is similar to that of the slope coefficient. However, the test statistic
^
b 0 − B0
t=
s^b0
Where:
258
© 2014-2024 AnalystPrep.
B1 = Hypothesized intercept coefficient.
s^
b0
= Standard error of the intercept.
2
1 X̄
s^
b0 =
+
⎷n ∑ ni=1 (Xi − X̄ )
2
Recall the example where inflation rates were regressed against unemployment rates from 2011
to 2020.
Y^ = 7.112 − 0.9020Xi + εi
Assume that we need to test whether the intercept is greater than 1 at a 5% significance level.
H 0 : b0 ≤ 1 versus Ha : b0 > 1
^
b −B
259
© 2014-2024 AnalystPrep.
^
b 0 − B0
t=
s^
b0
Where:
2
1 X̄ 1 5.292
s^
b0 =
+ = √ + = 1.501
⎷n ∑ni=1 (Xi − X̄ )
2 10 12.989
Therefore,
7.112 − 1
t= = 4.0719
1.501
Note that this is a one-sided test. From the table, t8 , 0.05 = 1.860. Since the calculated test
statistic is less than the critical t-value (4.0179 > 1.860), we reject the null hypothesis. There is
Dummy variables, also known as indicator variables or binary variables, are used in regression
analysis to represent categorical data with two or more categories. They are particularly useful
for including qualitative information in a model that requires numerical input variables.
(ESG) focused fund affects its monthly stock returns. In this case, we'll analyze the monthly
We can use a simple linear regression model to explore this. In the model, we regress monthly
returns, denoted as R, on an indicator variable, ESG. This indicator takes the value of 0 if the
R = b 0 + b 1 ESG + εi
260
© 2014-2024 AnalystPrep.
Note that we estimate the simple linear regression in a way similar to if the independent variable
was continuous.
The intercept β0 is the predicted value when the indicator variable is 0. On the other hand, the
slope when the indicator variable is 1 is the difference in the means if we grouped the
Assume that the following table is the results of the above regression analysis:
Additionally, we have the following information regarding the means and variances of the
variables.
The intercept (0.5468) equals the mean of the returns for the non-ESG stocks.
The slope coefficient (1.1052) is the difference in means of returns between ESG-
Now, assume we want to test whether the slope coefficient equals 0 at a 5% significance level.
48 − 2 = 46. As such, the critical t-values (usually given in the table above) is t46, 0.025 = ±2.013.
From the first table above, the calculated test statistic for the slope is greater than the critical t-
value (9.9532 > 2.013). As a result, we reject the null hypothesis that the slope coefficient is
equal to zero.
261
© 2014-2024 AnalystPrep.
p-Values and Level of Significance
The p-value is the smallest level of significance level at which the null hypothesis is rejected.
Therefore, the smaller the p-value, the smaller the probability of rejecting the true null
hypothesis (type I error) and, hence, the greater the validity of the regression model.
Software packages commonly offer p-values for regression coefficients. These p-values help test
a null hypothesis that the true parameter equals 0 versus the alternative that it's not equal to
zero.
We reject the null hypothesis if the p-value corresponding to the calculated test statistic is less
An analyst generates the following output from the regression analysis of inflation on
unemployment:
Regression Statistics
R Square 0.7684
Standard Error 0.0063
Observations 10
At the 5% significant level, test the null hypothesis that the slope coefficient is significantly
H0 : b 1 = 1 vs. H a : b1 ≠ 1
Solution
^
b 1 −b1
The calculated t-statistic, t = is equal to:
^
S b1
−0.9041 − 1
t= = −10.85
0.1755
262
© 2014-2024 AnalystPrep.
The critical two-tail t-values from the table with n − 2 = 8 degrees of freedom are:
tc = ±2.306
Therefore, we reject the null hypothesis and conclude that the estimated slope coefficient is
Note that we used the confidence interval approach and arrived at the same conclusion.
263
© 2014-2024 AnalystPrep.
Question 1
Samantha Lee, an investment analyst, is studying monthly stock returns. She focuses
explains how stock returns vary concerning the indicator variable RENEW. RENEW
equals 1 when there's a positive policy change towards renewable energy during that
month, and 0 if not. The total variation in the dependent variable amounted to
220.34. Of this, 94.75 is the part explained by the model. Samantha's dataset includes
36 monthly observations.
A. R 2=43.00%;F=26.07;Standard deviation=2.51.
B. R 2=53.00%;F=26.41;Standard deviation=2.55.
C. R 2=33.00%;F=36.07;Standard deviation=3.55.
Solution
Coefficient of determination:
F-statistic:
Standard deviation:
Note that,
264
© 2014-2024 AnalystPrep.
n
2
Total Variation = ∑ (Yi − Y¯) = 220.34
i=1
2
∑ni=1 (Y i − Y¯ )
Standard deviation = ⎷
n−1
As such,
Question 2
Neeth Shinu, CFA, is forecasting the price elasticity of supply for a specific product.
Shinu uses the quantity of the product supplied for the past 5months as the
dependent variable and the price per unit of the product as the independent variable.
Regression Statistics
R Square 0.9941
Standard Error 3.6515
Observations 5
Coefficients Standard Error t Stat P-value
Intercept −159 10.520 (15.114) 0.001
Slope 0.26 0.012 22.517 0.000
Which of the following most likely reports the correct value of the t-statistic for the
slope and most accurately evaluates its statistical significance with 95% confidence?
Solution
265
© 2014-2024 AnalystPrep.
The correct answer is A.
^
b 1 − b1
t=
S^b 1
Where:
^
b 1 = Point estimator for B 1 .
0.26 − 0
t= = 21.67
0.012
The critical two-tail t-values from the t-table with n − 2 = 3 degrees of freedom are:
tc = ±3.18
266
© 2014-2024 AnalystPrep.
Notice that | t| > tc (i.e., 21.67 > 3.18).
Therefore, the null hypothesis can be rejected. Further, we can conclude that the
267
© 2014-2024 AnalystPrep.
LOS 10d: describe the use of analysis of variance (ANOVA) in regression
analysis, interpret ANOVA results, and calculate and interpret the
standard error of estimate in a simple linear regression
The sum of squares of a regression model is usually represented in the Analysis of Variance
(ANOVA) table. The ANOVA table contains the sum of squares (SST, SSE, and SSR), the degrees
mean square error or standard error of the regression. It measures the distance between the
observed and dependent variables predicted by the regression model. The Standard Error of
Estimate is easily calculated from the ANOVA table using the following formula:
2
∑ ni=1 (Y i − Y^)
Standard Error of Estimate (Se ) = √MSE = ⎷
n− 2
The standard error of estimate, coefficient of determination, and F-statistic are the measures
that can be used to gauge the goodness of fit of a regression model. In other words, these
measures are used to tell the extent to which a regression model syncs with data.
The smaller the Standard Error of Estimate is, the better the fit of the regression line. However,
the Standard Error of Estimate does not tell us how well the independent variable explains the
268
© 2014-2024 AnalystPrep.
Example: Calculating and Interpreting F-Statistic
The completed ANOVA table for the regression model of the inflation rate against the
b. Test the hypothesis that the slope coefficient equals a 5% significance level.
Solution
significance level is roughly 5.32. Note that this is a one-tail test, so we use the 5% F-
table.
269
© 2014-2024 AnalystPrep.
Remember that the null hypothesis is rejected if the calculated value of the F-statistic is
greater than the critical value of F. Since 26.960 > 5.32, we reject the null hypothesis
and conclude that the slope coefficient is significantly different from zero. Notice that we
also rejected the null hypothesis in the previous examples. We did so because the 95%
confidence interval did not include [Link] F-test duplicates the t-test in regard to the
slope coefficient significance for a linear regression model with one independent
variable. In this case, t2 = 2.3062 ≈ 5.32. Since the F-statistic is the square of the t-
statistic for the slope coefficient, its inferences are the same as the t-test. However, this
270
© 2014-2024 AnalystPrep.
Question
The value of R 2 and the F-statistic for the test of fit of the regression model are
closest to:
A. 6% and 16.
Solution
271
© 2014-2024 AnalystPrep.
LOS 10e: calculate and interpret the predicted value for the dependent
variable, and a prediction interval for it, given an estimated linear
regression model and a value for the independent variable
We calculate the predicted value of the dependent variable, Y , by inserting the estimated value
of the independent variable, X , into the regression equation. The predicted value of the
Y^ = ^
b0 +^
b 1X
Where:
Refer to the example of regressed inflation rates against unemployment rates from 2011 to 2020.
Y^ = 7.112 − 0.9020Xi + εi
Calculate the predicted inflation rate value if the forecasted value of the unemployment rate is
272
© 2014-2024 AnalystPrep.
4.5%.
Solution
The confidence interval calculation for the predicted value of a dependent variable is the same as
that of the confidence interval for regression coefficients. The confidence interval for a predicted
Where:
2 2
⎡ 1 (Xf − X̄) ⎤ ⎡ 1 (Xf − X̄) ⎤
s2f = s2e 1+ + = s2e 1 + +
⎣ n (n − 1) sx ⎦
2 ⎣ n 2
∑ni = 1 (X i − X̄) ⎦
Where:
n = Number of observations.
273
© 2014-2024 AnalystPrep.
We can, therefore, calculate the standard error of forecast as shown below:
2
1 (Xf − X̄ )
sf = se 1 + +
⎷ n 2
∑ni=1 (Xi − X̄ )
A better fit of the regression analysis leads to a smaller standard error of the estimate
When the sample size (n) in the regression calculation increases, it directly
variable (X̄ ) utilized in the regression analysis, it decreases the standard error of the
forecast.
Refer to the example of regressed inflation rates against unemployment rates from 2011 to 2020.
Consider the results of the regression analysis of inflation rates on unemployment rates:
274
© 2014-2024 AnalystPrep.
Regression Statistics
R Square 0.7711
Standard Error 0.6261
Observations 10
ANOVA
df Sum of Mean F
Squares Square
Regression 1 10.568 10.568 26.9565
Residual 8 3.136 0.392
Total 9 13.704
Given that the forecasted unemployment rate is 4.5%, calculate the 95% confidence interval for
Solution
2
⎡ 1 (X f − X̄ ) ⎤
s 2f = s 2e 1 + +
⎣ n (n − 1) s2X ⎦
2
⎡ 1 (X f − X̄ ) ⎤
= s 2e 1+ +
⎣ n 2
∑ni=1 (Xi − X̄ ) ⎦
2
1 (4.5 − 5.29)
= 0.62612 [1 + + ] = 0.450
10 12.989
sf = √0.450 = 0.6708
The predicted value of the inflation rate given an unemployment rate of 4.5% is:
275
© 2014-2024 AnalystPrep.
Y^ = 7.112 − 0.9020 × 4.5 = 3.05%
The two-tailed critical t-value with 8 (n − 2) degrees of freedom at the 5% significance level is
2.306.
276
© 2014-2024 AnalystPrep.
PI = 3.05 ± 2.306 × 0.6708 = 1.50% to 4.60%
Interpretation
Given an unemployment rate of 4.5%, we are 95% confident that the inflation rate will lie
277
© 2014-2024 AnalystPrep.
Question 1
The regression equation of the quantity of goods against the price is given by:
Y = −159 + 0.26X
Where:
Y = Quantity supplied.
The predicted value of the quantity supplied when the price equals 1,200 is closest
to:
A. 153.
B. 155.
C. 471.
278
© 2014-2024 AnalystPrep.
LOS 10f: Describe different functional forms of simple linear regressions
the data for linear regression. Here are three commonly used log transformation functional
forms:
1. Log-lin model: In this log transformation, the dependent variable is logarithmic, while
lnY = b0 + b1 Xi .
The slope coefficient in the log-lin model is the relative change in the dependent variable
When utilizing a log-lin model, caution must be exercised when making forecasts. For
lnY = −3 , then,
Y = e−3 = 0.0498
Moreover, the lin-lin model cannot be compared with the log-lin model without the
2. Lin-log model: In this case, the dependent variable is linear, while the independent
Y i = b 0 + b 1 lnXi .
The slope coefficient in the lin-log model is responsible for the absolute change in the
3. Log-log model: In this log transformation, both the dependent and independent
the log-log model is the relative change in the dependent variable for a relative change in
279
© 2014-2024 AnalystPrep.
Selecting the Correct Functional Form
To settle on the correct functional form, consider the following goodness of fit measures:
In addition to the factors cited above, the patterns in residuals can also be analyzed when
280
© 2014-2024 AnalystPrep.
Question 1
Which of the following statements about the log-lin model is most likely correct:
logarithmic.
linear.
lnY = b 0 + b 1 Xi
A is incorrect. It describes the lin-log model, where the dependent variable is linear
B is incorrect. It describes the log-log model, where both the dependent and
281
© 2014-2024 AnalystPrep.
Learning Module 11: Introduction to Big Data Techniques
LOS 11b: describe Big Data, artificial intelligence, and machine learning
Big data is a term that describes large, complex datasets. These datasets are analyzed with
computers to uncover patterns and trends, particularly those related to human behavior. Big
data includes traditional sources like company reports and government data and non-
traditional sources like social media, sensors, electronic devices, and data generated as a
Volume: The amount of data collected in various forms, including files, records, tables, etc.
Velocity: The speed of data processing can be extremely high. In most cases, we deal with real-
time data.
Variety: The number of types/formats of data. The data could be structured (e.g., SQL tables or
Veracity: This is the trustworthiness and reliability of data sources. Veracity is crucial when
using big data for making predictions or drawing conclusions. Big data makes it challenging to
Structured data refers to information with a high degree of organization. Items can be
organized in tables and stored in a database where each field represents the same type of
information.
282
© 2014-2024 AnalystPrep.
Unstructured data refers to information with a low degree of organization. Items such as text
messages, tweets, emails, voice recordings, pictures, blogs, scanners, and sensors are
Semi-structured data may have the qualities of both structured and unstructured data.
transactions.
Individuals: Product reviews, credit card purchases, social media posts, etc.
The Internet of Things: data generated by ‘smart’ buildings through fittings such as
Professional investors, particularly quantitative ones, use alternative data sources in their
financial analysis and decision-making. These sources significantly influence how they conduct
their processes. They use alternative data to support data-driven investment models and
decisions.
Commercial operations data: This includes data on credit cards and corporate
283
© 2014-2024 AnalystPrep.
Data produced by sensors: This data is typically unstructured and is gathered
Investment professionals must consider legal and ethical aspects when they use non-public
information. Web data scraping can gather personal data that might be legally protected or
Quality: Important questions include, but are not limited to, "Does the dataset contain
Experts have created artificial intelligence (AI) and machine learning methods to handle large
and intricate alternative datasets. These technologies help in understanding and evaluating this
Artificial Intelligence
In broad terms, artificial intelligence refers to machines that can perform tasks in “intelligent”
ways. It has much to do with developing computer systems that exhibit cognitive and decision-
being able to carry out tasks in a way that we would consider “smart.”
Early AI took the shape of expert systems, using "if-then" computer programming to mimic
human knowledge and analysis. Neural networks, another early form, mimicked human brain
284
© 2014-2024 AnalystPrep.
Machine Learning
Machine learning is a current application of AI that revolves around the idea that we should
really just give machines access to data and let them learn by themselves without making any
The idea is that when exposed to more data, machines can make changes independently and
come up with solutions to problems without reliance on human expertise – find and apply the
pattern.
In the context of investment, machine learning requires big data for training. The growth of big
In machine learning (ML), a computer algorithm receives inputs, which can be datasets or
variables, as well as outputs, representing the target data. The algorithm then learns how to
effectively model inputs into outputs or describe a data structure. It learns by identifying data
The ML divides the dataset into three unique types: a training dataset, a validation dataset,
and a test dataset. A training dataset allows the algorithm to identify the link between inputs
and outputs based on the historical pattern in the data. These relationships are then validated,
As the name suggests, the test dataset tests the model's strength in predicting well on the new
data. Note that machine learning still needs human intervention to understand the underlying
data and choose suitable techniques for data analysis. In other words, before data is utilized, it
The model overfits the data when it discovers “false” associations or “unsubstantiated” patterns
that cause prediction errors and wrong forecasts. In other words, overfitting happens when the
ML model is overtrained on the data and considers the noise in the data as true parameters.
285
© 2014-2024 AnalystPrep.
Underfitting the Data
Underfitting of data occurs when the model considers the true parameters as noise and is unable
to identify the relationship within the training data. In other words, the model is too simple to
Machine learning models don't use explicit rules like traditional software. They learn from lots of
data during training. This makes ML models, such as black boxes, sometimes give results that
Supervised Learning
Under supervised learning, computers learn to model data based on labeled training data
containing inputs and the desired outputs. After “learning” how best to model the relationships
for the labeled data, the algorithms are employed to predict the results for the new datasets.
Unsupervised Learning
In unsupervised learning, computers get input data without labels and have to describe it, often
by grouping data points. They learn from unlabeled data and react based on commonalities. For
Deep Learning
Deep learning involves computers using neural networks to process data in multiple stages,
identifying complex patterns. It employs both supervised and unsupervised machine learning
methods.
286
© 2014-2024 AnalystPrep.
Question
Solution
programs, enabling them to solve problems without human input. It's about
287
© 2014-2024 AnalystPrep.
LOS 11c: Describe applications of Big Data and Data Science to
investment management
Data science is an interdisciplinary field that uses developments in computer science, statistics,
and other fields to extract information from Big Data or data in general.
Data analysts and scientists in big data analysis use different data management approaches.
Capture: Describes the method by which data is gathered and put into a form that the
Curation: Data curation ensures the quality and accuracy of the data by undertaking a
data cleaning activity. This procedure finds data inaccuracies, and any missing data is
compensated for.
Search: Involves querying data to locate specific information. With big data,
content.
Transfer: Describes the process of transferring data from the underlying data source
Data Visualization
representations. Tables, charts, and trends are commonly used for traditional structured data,
while non-traditional unstructured data demands innovative techniques like interactive three-
288
© 2014-2024 AnalystPrep.
dimensional (3D) graphics, tag clouds, and mind maps.
Text analytics employs computer programs to analyze and extract insights, primarily from
unstructured text- or voice-based datasets like company filings, written reports, quarterly
earnings calls, and social media content. Text analytics can be utilized in predictive analysis to
Natural language processing (NLP) is an area of study that involves creating computer programs
to decipher and analyze human language. Essentially, NLP combines computer science, AI, and
linguistics.
Translation, speech recognition, text mining, sentiment analysis, and topic analysis are examples
of automated tasks that use NLP. Annual reports, call transcripts, news articles, social media
posts, and other text- and audio-based data may all be analyzed using natural language
processing (NLP), allowing NLP to discover trends more quickly and accurately than is humanly
possible.
Using natural language processing data, earnings projections for a company's near-term
prospects can be created. X (formerly Twitter) sentiments have also been used to gauge an initial
Python, R, and Excel VBA are frequently used programming languages, whereas SQL, SQLite,
289
© 2014-2024 AnalystPrep.
Question
Which of the five data processing methods refers to the process of ensuring data
A. Data search.
B. Data storage.
C. Data curation.
Data curation refers to the process of ensuring data quality and accuracy through a
data cleaning exercise. It involves uncovering data errors and adjusting for missing
data.
A is incorrect. Data search refers to how to query data. Big data requires advanced
B is incorrect. Data storage refers to how the data will be recorded, archived, and
290
© 2014-2024 AnalystPrep.
LOS 11a: describe aspects of “fintech” that are directly relevant for the
gathering and analyzing of financial data
Fintech refers to technological innovation in designing and delivering financial services and
products. At its core, fintech has helped companies, business owners, and investment managers
Note that the term fintech is commonly used to refer to companies that develop new
technologies and their applications and also the business sector that encompasses such
companies.
Initially, financial innovation was limited to simple tasks such as data processing and automation
of routine tasks. Today, fintech encompasses more advanced systems that can analyze
information and make decisions based on machine-learning logic. Machines have been developed
to “learn” how to perform tasks over time. Using such systems has brought about high levels of
efficiency that surpass human capabilities. Fintech covers a broader range of services and
applications. As such, services and applications of fintech relevant to the investment industry
include:
1. Analysis of Large Datasets: Apart from traditional data such as corporate financial
analysis techniques. Diverse approaches to data analysis are now possible because of
through the vast volumes of data from corporate filings and annual reports to produce
insights.
291
© 2014-2024 AnalystPrep.
Question
A. at its most advanced state, using systems that follow specified rules and
instructions.
B. limited to simple tasks such as automating routine processes and data processing.
advancement.
The availability of vast amounts of data and technological advancements have been
the primary drivers of fintech's expansion. The rapid growth in data, including
diverse types, large quantities, and improved quality, has provided valuable insights
have made it possible to analyze and interpret this data effectively, creating
A is incorrect. While this may be true for some aspects of fintech, it doesn't directly
address the two most important reasons behind its growth - the rapid growth in data
financial solutions.
various complex financial activities. While it does automate routine processes and
data processing, it does so through advanced technologies that enable the handling of
292
© 2014-2024 AnalystPrep.