QuintEdge Notes
QuintEdge Notes
Edition
2025
QUANTITATIVE METHODS
[ Pre requisite + Core Curriculum]
Quantitative Methods
Study Notes
ii
Preface
Quintedge's series of CFA Exam Notes is crafted with the vision of guiding aspiring financial analysts through the rigorous and
intellectually demanding Chartered Financial Analyst (CFA) Program. Understanding the depth and breadth of knowledge
required to excel in this program, our company, operated by CFA Qualified Professionals, aims to demystify the complex world
of finance and investment.
The CFA credential symbolizes a commitment to excellence, ethical standards, and a deep understanding of financial analysis
and portfolio management. It's a path that demands not just theoretical knowledge but the ability to apply these concepts in
real-world scenarios. Every Study Note in our series is dedicated to one of the ten subjects within the CFA curriculum, ensuring a
comprehensive and focused study experience.
Our approach sets us apart. Quintedge is committed to real-life learning, integrating practical examples and contemporary case
studies to illustrate theoretical concepts. We believe that true understanding comes from seeing these theories in action, which
is why each note is filled with real-world scenarios and examples from the current financial markets. This method ensures that
learners are not just memorizing information but are truly comprehending and able to apply it.
Our team, comprising CFA Charterholders and industry experts, brings a wealth of knowledge and personal experience to the
table. They've navigated the challenges of the CFA Program themselves and understand what it takes to succeed. Their insights
have shaped these Study Notes into more than just summaries.
Quintedge is dedicated to your success. We understand the commitment you're making and the challenges ahead. As such, our
series is designed to be a companion in your journey, transforming the daunting task of preparing for the CFA exams into an
achievable goal. We stand with you, ready to turn hard work into achievement and ambition into success.
Disclaimer: Quintedge’s CFA Exam Notes should be used in conjunction with the original readings as set forth by CFA Institute in
their 2024 Level I CFA Study Guide. The information contained in these Notes covers topics contained in the readings referenced
by CFA Institute and is believed to be accurate. However, their accuracy cannot be guaranteed nor is any warranty conveyed as
to your ultimate exam success.
iii
Table of Contents
Note: The prerequisites for “Quantitative Methods” are not presented separately. Instead, they are combined with the core
curriculum and are jointly presented in the following pages.
iv
6.8. Guide to Selecting among Visualization Types............................................................................................................ 18
7. Measures of Central Tendency ............................................................................................................................................ 19
7.1. The Arithmetic Mean .................................................................................................................................................. 19
7.2. The Median ................................................................................................................................................................ 19
7.3. The Mode .................................................................................................................................................................... 20
7.4. Other Concepts of Mean ............................................................................................................................................. 20
8. Quantiles .............................................................................................................................................................................. 21
8.1. Quartiles, Quintiles, Deciles, and Percentiles ............................................................................................................. 21
8.2. Quantiles in Investment Practice ................................................................................................................................ 22
9. Measures of Dispersion ....................................................................................................................................................... 22
9.1. The Range.................................................................................................................................................................... 22
9.2. The Mean Absolute Deviation ..................................................................................................................................... 22
9.3. Sample Variance and Sample Standard Deviation ...................................................................................................... 23
10. Downside Deviation and Coefficient of Variation ........................................................................................................... 23
10.1. Coefficient of Variation ............................................................................................................................................... 23
11. The Shape of the Distributions ........................................................................................................................................ 24
11.1. The Shape of the Distributions: Kurtosis ..................................................................................................................... 24
12. Correlation between Two Variables ................................................................................................................................ 25
12.1. Properties of Correlation ............................................................................................................................................ 26
12.2. Limitations of Correlation Analysis ............................................................................................................................. 26
v
5. Applications of the Normal Distribution .............................................................................................................................. 43
6. Lognormal Distribution and Continuous Compounding ...................................................................................................... 44
6.1. The Lognormal Distribution ........................................................................................................................................ 44
6.2. Continuously Compounded Rates of Return ............................................................................................................... 44
7. Student’s t-, Chi-Square, and F-Distributions ...................................................................................................................... 45
7.1. Student’s t-Distribution............................................................................................................................................... 45
7.2. Chi-Square and F-Distribution ..................................................................................................................................... 45
8. Monte Carlo Simulation ....................................................................................................................................................... 46
vi
5.2. Decision Rules and Confidence Intervals .................................................................................................................... 63
5.3. Collect the Data and Calculate the Test Statistic ........................................................................................................ 63
6. Make a Decision ................................................................................................................................................................... 64
6.1. Make a Statistical Decision.......................................................................................................................................... 64
6.2. Make an Economic Decision ....................................................................................................................................... 64
6.3. Statistically Significant but Not Economically Significant? .......................................................................................... 64
7. The Role of p-Values ............................................................................................................................................................ 64
8. Multiple Tests and Significance Interpretation .................................................................................................................... 65
9. Tests Concerning a Single Mean .......................................................................................................................................... 66
10. Test Concerning Differences between Means with Independent Samples..................................................................... 67
11. Test Concerning Differences between Means with Dependent Samples ....................................................................... 67
12. Testing Concerning Tests of Variances ............................................................................................................................ 68
12.1. Tests of a Single Variance ............................................................................................................................................ 68
12.2. Test Concerning the Equality of Two Variances (F-Test) ............................................................................................. 69
13. Parametric vs. Nonparametric Tests ............................................................................................................................... 70
13.1. Uses of Nonparametric Tests ...................................................................................................................................... 70
14. Tests Concerning Correlation .......................................................................................................................................... 70
14.1. Parametric Test of a Correlation ................................................................................................................................. 71
14.2. Tests Concerning Correlation: The Spearman Rank Correlation Coefficient .............................................................. 71
15. Test of Independence Using Contingency Table Data ..................................................................................................... 71
Appendices ........................................................................................................................................... 88
1. Appendix A - CUMULATIVE Z-TABLE .................................................................................................................................... 88
1.1. Standard Normal Distribution – For positive Z value .................................................................................................. 88
1.2. Standard Normal Distribution – For negative Z value ................................................................................................. 89
2. Appendix B: STUDENT’S T-DISTRIBUTION ............................................................................................................................ 90
3. Appendix C: F-TABLE AT 5% (UPPER TAIL) ........................................................................................................................... 91
4. Appendix D: F-TABLE AT 2.5% (UPPER TAIL) ........................................................................................................................ 92
5. Appendix E: CHI-SQUARED TABLE........................................................................................................................................ 93
viii
Quantitative Methods - Learning Module 1
for the CFA exam
The Time Value of Money
Learning Outcomes:
a) Interpret interest rates as required rates of return, discount rates, or opportunity costs
b) Explain an interest rate as the sum of a real risk-free rate and premiums that compensate investors for bearing distinct
types of risk
c) Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
d) Demonstrate the use of a time line in modelling and solving time value of money problems
e) Calculate the solution for time value of money problems with different frequencies of compounding
f) Calculate and interpret the effective annual rate, given the stated annual interest rate and the frequency of compounding
1. Introduction
The time value of money is a crucial concept for both individuals and investment analysts. It revolves around the idea that
money's worth changes over time, with people valuing money received sooner more than money received later. This principle is
fundamental when making decisions about saving or borrowing money. For investment analysts, understanding the time value
of money is vital when evaluating financial transactions with present and future cash flows. It helps in determining the value of
future cash flows and establishing equivalence relationships between cash flows with different dates. In summary, grasping the
mathematics of time value of money is essential for accurate decision-making and financial analysis
2. Interest Rates
LOS (a) : Interpret interest rates as required rates of return, discount rates, or opportunity costs
LOS (b) : Explain an interest rate as the sum of a real risk-free rate and premiums that compensate investors for bearing distinct
types of risk
The time value of money deals with the concept of equivalence between cash flows occurring at different times. It's a
straightforward idea: if you paid $10,000 today and received only $9,500 today in return, you wouldn't agree to the deal.
However, if you received the $9,500 today and paid the $10,000 a year from now, these amounts could be considered
equivalent. This is because $10,000 in the future is worth less than $10,000 today. To account for this difference, an interest rate
(denoted as "r") is used to express the relationship between cash flows at different dates. If $9,500 today and $10,000 in a year
are considered equivalent, the interest rate required to make them so is 5.26 percent, meaning you'd need a 5.26 percent
return to compensate for receiving $10,000 in one year rather than today.
o they represent the minimum return an investor requires to accept an investment, which is the required rate of return.
o they act as discount rates, helping determine the present value of future amounts, making "interest rate" and "discount
rate" almost interchangeable.
o interest rates can be seen as opportunity costs, signifying the value an investor gives up by choosing a specific course of
action. For instance, in the example given, the 5.26 percent interest rate represents the opportunity cost of not investing
the $9,500 today and missing out on that return.
Economics teaches us that interest rates in the marketplace are determined by the interplay of supply and demand, with
investors supplying funds and borrowers demanding funds. From the investor's perspective, an interest rate (r) is comprised of
several components, including:
1. The real risk-free interest rate: This is the interest rate for a risk-free investment if inflation were not a factor. It
reflects people's preferences for current versus future consumption.
2. The inflation premium: This compensates investors for expected inflation over the maturity of the investment. It
accounts for the decrease in the purchasing power of money due to inflation.
1
3. The default risk premium: It compensates investors for the risk that the borrower might fail to make payments as
promised.
4. The liquidity premium: This compensates investors for the risk of loss if an investment needs to be quickly converted to
cash. Investments like US T-bills have low liquidity premiums, while less liquid bonds have higher ones.
5. The maturity premium: It compensates investors for the increased sensitivity of debt's market value to changes in
market interest rates as maturity lengthens. Longer-term debt often comes with a positive maturity premium.
Overall, these components together make up the interest rate, reflecting the return required by investors in exchange for their
funds and the risks associated with the investment i.e.:
➔ r = Real risk-free interest rate + Inflation premium + Default risk premium + Liquidity premium + Maturity premium
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS (d) : Demonstrate the use of a time line in modelling and solving time value of money problems
Suppose you invest $100 (initial investment or Present Value = $100) in an interest-bearing bank account paying 5 percent
annually. At the end of the first year, you will have the $100 plus the interest earned, 0.05 × $100 = $5, for a total of $105. If we
invest the same amount for two years, and interest is credited into the account annually (i.e annual compounding), at the end of
the first year (the beginning of the second year), your account will have $105, which you will leave in the bank for another year.
Thus, with a beginning amount of $105 (PV = $105), the amount at the end of the second year will be $105(1.05) = $110.25.
Note that the $5.25 interest earned during the second year is 5 percent of the amount invested at the beginning of Year 2.
The $5 interest earned on the initial $100 investment is simple interest, calculated as the interest rate multiplied by the
principal amount. Over two years, you accumulate $10 in simple interest. The extra $0.25 at the end of Year 2 represents
interest earned on the Year 1 interest of $5 that was reinvested. This concept illustrates the phenomenon of compounding.
While the interest earned on the initial investment is important, it remains constant in size over periods. In contrast,
compounded interest on reinvested interest is much more powerful because it grows in size with each period and its
significance is greater with higher interest rates and longer investment periods.
The general formula relating the present value of an initial investment to its future (final) value after N periods is :
Where, r is the interest rate per period and N is the number of compounding periods. For our example, this becomes:
Note : When using the future value equation, it's crucial to ensure that the interest rate (r) and the number of compounding
periods (N) are in the same time units. If N is expressed in months, then r should be the monthly interest rate, not the
annualized rate. Compatibility in time units is essential for accurate calculations.
Example
The Future Value of a Lump Sum
A pension fund manager estimates that his corporate sponsor will make a $10 million contribution five years from now. The rate of return on
plan assets has been estimated at 9 percent per year. The pension fund manager wants to calculate the future value of this contribution 15
years from now, which is the date at which the funds will be distributed to retirees. What is that future value?
2
By positioning the initial investment, PV, at t = 5, we can calculate the future value of the contribution using the following data :
From the standpoint of today (t = 0), the future amount of $23,673,636.75 is 15 years into the future. Although the future value is 10 years
from its present value, the present value of $10 million will not be received for another five years.
For this illustration, we have followed the convention of indexing today as t = 0 and indexing subsequent times by adding 1 for each period.
The additional contribution of $10 million is to be received in five years, so it is indexed as t = 5 and appears as such in the figure. The future
value of the investment in 10 years is then indexed at t = 15; that is, 10 years following the receipt of the $10 million contribution at t = 5. Time
lines like this one can be extremely useful when dealing with more-complicated problems, especially those involving more than one cash flow.
LOS (e) : Calculate the solution for time value of money problems with different frequencies of compounding
This section discusses investments that pay interest more frequently than once a year, such as monthly compounding offered by
many banks. Instead of specifying the monthly interest rate, financial institutions often provide an annual interest rate, known
as the stated annual interest rate or quoted interest rate, denoted as rs. For example, if a bank advertises an 8 percent annual
interest rate compounded monthly, the monthly interest rate is calculated as 0.08/12 = 0.0067 or 0.67 percent. However, it's
important to note that this rate is a quoting convention, as (1 + 0.0067) 12 equals 1.083, not 1.08. The term (1 + rs) is not
intended to be a future value factor when compounding occurs more frequently than annually.
With more than one compounding period per year, the future value formula becomes :
Where, rs is the stated annual interest rate, m is the number of compounding periods and N is the number of years. The periodic
rate, rs/m, and the number of compounding periods, mN, must be compatible
5. Continuous Compounding
LOS (e) : Calculate the solution for time value of money problems with different frequencies of compounding
LOS (f) : Calculate and interpret the effective annual rate, given the stated annual interest rate and the frequency of
compounding
The previous explanation focused on discrete compounding, where interest is credited after a specific period. However, if the
number of compounding periods per year approaches infinity, it leads to continuous compounding. To apply the future value
formula in this scenario, we calculate the limiting value of the future value factor as the number of compounding periods per
year goes towards infinity. This results in the following expression :
➔ FVN = PV ersN
The term ersN is the transcendental number e ≈ 2.7182 raised to the power r sN.
The effects of the number of compounding periods for the same interest rate is summarized below :
3
The above implies that an 8% rate for semi-annual compounding is equivalent to 8.16% rate for annual compounding. This
concept is referred to as the effective annual rate (EAR) i.e. for annual 8% rate with semi-annual compounding, the EAR is
8.16%.
The effective annual rate concept extends to all compounding periods (even continuous) and is calculated as follows:
Where, periodic interest rate is the stated annual interest rate divided by m, where m is the number of compounding periods in
one year. We can also reverse the formulas for EAR with discrete and continuous compounding to find a periodic rate that
corresponds to a particular effective annual rate.
LOS () : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS () : Demonstrate the use of a time line in modeling and solving time value of money problems
This section discusses series of cash flows, both regular and irregular. The key terms used in valuing cash flows distributed over
multiple time periods are
In this scenario, there are five separate deposits of $1,000 each, occurring at yearly intervals, with the first payment at t = 1. The
objective is to determine the future value of this ordinary annuity after the last deposit at t = 5, with an annual interest rate of 5
percent. Using FVN = PV(1 + r)N), the future value of each $1,000 deposit is calculated as of t = 5. For instance, the first $1,000
deposit made at t = 1 will have grown to $1,215.51 by t = 5. The future values of all deposits are calculated similarly. Since we're
finding the future values at t = 5, the last payment does not earn any interest. When all these future values at t = 5 are added
together, the future value of the annuity is $5,525.63.
The general annuity formula for annuity amount A, N number of periods and interest rate per period as r is :
(1+𝑟)𝑁 – 1 )
➔ FVN = A [ ]. Where the bracketed term is the future value annuity factor.
𝑟
4
In many cases, cash flow streams are unequal, precluding the simple use of the future value annuity factor. In such cases, we
always find the future value of a series of unequal cash flows by compounding each individual cash flow separately, one at a
time and summing them up to find the total future value.
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS (d) : Demonstrate the use of a time line in modelling and solving time value of money problems
The present value factor, like the future value factor, helps connect present value to future value. It enables us to discount a
future value back to its present value. For instance, if an investment at 5 percent interest grows to $105 in one year, what
amount invested now at the same interest rate will result in $105 in one year? To calculate present value, given a future cash
flow to be received in N periods and an interest rate per period of r, we can use the formula for future value in reverse :
➔ PV = FVN (1 + r)-N
(1 + r)-N is the present value factor. Notice the negative sign in this case.
Present value problems involve calculating the present value factor, which is represented as (1 + r) -N. The relationship between
present values, the discount rate (r), and the number of periods (N) is as follows:
1. When the discount rate is fixed, the further in the future an amount is to be received, the smaller its present value
becomes.
2. Keeping the time period constant, if the discount rate is higher, the present value of a future amount will be smaller. In
other words, higher discount rates result in lower present values for future cash flows
LOS () : Calculate the solution for time value of money problems with different frequencies of compounding
In general, with more than one compounding period in a year, we can express the formula for present value as
➔ PV = FVN ( 1 + rs/m) - mN
LOS () : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS () : Demonstrate the use of a time line in modeling and solving time value of money problems
Investment management frequently deals with assets that provide a sequence of cash flows over time. These cash flows can
vary in terms of frequency, consistency, and duration, ranging from irregular to regular and short-term to long-term, and even
continuing indefinitely. This section explores the methods for calculating the present value of such a series of cash flows.
An ordinary annuity consists of equal payments, and the first payment occurs one period into the future. The annuity involves N
payments in total, with the first payment at t = 1 and the last payment at t = N. The present value of an ordinary annuity can be
found by summing up the present values of each individual annuity payment:
5
➔ PV = A/(1+r) + A/(1+r)2 + A/(1+r)3 +... + A/(1+r)N−1 + A/(1+r)N . This can be re-written as :
When we have unequal cash flows, we must first find the present value of each individual cash flow and then sum the respective
present values
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
Consider the case of an ordinary annuity that extends indefinitely. Such an ordinary annuity is called a perpetuity (a perpetual
annuity). The formula for the present value of a perpetuity is a modified version of the earlier equation to account for the
infinite cash flow series :
As long as interest rates are positive, the sum of present value factors converges to :
Example
The Present Value of a Perpetuity
The British government once issued a type of security called a consol bond, which promised to pay a level cash flow indefinitely. If a consol
bond paid £100 per year in perpetuity, what would it be worth today if the required rate of return were 5 percent?
Solution: To answer this question, we can use Equation 13 with the following data:
A = £100
r = 5% = 0.05
PV = A / r = £100 / 0.05 = £2, 000. The bond would be worth £2,000.
Consider a level perpetuity of £100 per year with its first payment beginning at t = 5. What is its present value today (at t = 0), given a 5
percent discount rate?
Solution: First, we find the present value of the perpetuity at t = 4 and then discount that amount back to t = 0. (Recall that a perpetuity or an
ordinary annuity has its first payment one period away, explaining the t = 4 index for our present value calculation.)
A = £100
r = 5% = 0.05
PV = A/r= £100/0.05 = £2, 000
2. Find the present value of the future amount at t = 4. From the perspective of t = 0, the present value of £2,000 can be considered a future
value. Now we need to find the present value of a lump sum:
6
Today’s present value of the perpetuity is £1,645.40.
11. Solving for Interest Rates, Growth Rates, and Number of Periods
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
In previous examples, specific details like the interest rate (r), number of time periods (N), annuity amount (A), and either the
present value (PV) or future value (FV) were provided. In practical situations, you might need to determine the interest rate, the
number of periods, or the annuity amount when you have the present or future values
Say a €100 bank deposit is known to result in a €111 payoff in one year. Using FV N = PV(1 + r)N with N = 1, we can calculate the
interest rate (r) that equates the present value of €100 with the future value of €111:
So, the interest rate that links €100 at t = 0 to €111 at t = 1 is 11 percent. This means that €100 grows to €111 with a growth rate
of 11 percent. This example illustrates that an interest rate can also be viewed as a growth rate. The choice of whether to refer
to it as an "interest rate" or a "growth rate" depends on the specific context. Solving for r in the present value equation and
replacing it with the growth rate (g) results in the formula for determining growth rates:
➔ g = (FVN/PV)(1/N) - 1
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS (d) : Demonstrate the use of a time line in modelling and solving time value of money problems
Example
The Projected Annuity Amount Needed to Fund a Future-Annuity Inflow
Jill Grant is 22 years old (at t = 0) and is planning for her retirement at age 63 (at t = 41). She plans to save $2,000 per year for the next 15 years
(t = 1 to t = 15). She wants to have retirement income of $100,000 per year for 20 years, with the first retirement payment starting at t = 41.
How much must Grant save each year from t = 16 to t = 40 in order to achieve her retirement goal? Assume she plans to invest in a diversified
stock-and-bond mutual fund that will earn 8 percent per year on average.
Solution: To help solve this problem, we set up the information on a time line which shows, Grant will save $2,000 (an outflow) each year for
Years 1 to 15. Starting in Year 41, Grant will start to draw retirement income of $100,000 per year for 20 years. In the time line, the annual
savings is recorded in parentheses ($2) to show that it is an outflow. The problem is to find the savings, recorded as X, from Year 16 to Year 40.
Solving this problem involves satisfying the following relationship: the present value of savings (outflows) equals the present value of
retirement income (inflows). We could bring all the dollar amounts to t = 40 or to t = 15 and solve for X.
Let us evaluate all dollar amounts at t = 15. As of t = 15, the first payment of X will be one period away (at t = 16). Thus we can value the
stream of X’s using the formula for the present value of an ordinary annuity. This problem involves three series of level cash flows. The basic
idea is that the present value of the retirement income must equal the present value of Grant’s savings. Our strategy requires the following
steps:
7
1. Find the future value of the savings of $2,000 per year and index it at t = 15. This value tells us how much Grant will have saved.
2. Find the present value of the retirement income at t = 15. This value tells us how much Grant needs to meet her retirement goals (as of t
= 15). Two substeps are necessary. First, calculate the present value of the annuity of $100,000 per year at t = 40. Use the formula for the
present value of an annuity. (Note that the present value is indexed at t = 40 because the first payment is at t = 41.) Next, discount the
present value back to t = 15 (a total of 25 periods).
3. Now compute the difference between the amount Grant has saved (Step 1) and the amount she needs to meet her retirement goals (Step
2). Her savings from t = 16 to t = 40 must have a present value equal to the difference between the future value of her savings and the
present value of her retirement income.
Our goal is to determine the amount Grant should save in each of the 25 years from t = 16 to t = 40. We start by bringing the $2,000 savings to
t = 15, as follows:
A = $2,000
r = 8 %= 0.08
N = 15
FV=A[(1+r)N−1] / r = $2,000 [(1.08)15 – 1]/ 0.08 = $2, 000 (27.152114) = $54, 304.23
At t = 15, Grant’s initial savings will have grown to $54,304.23. Now we need to know the value of Grant’s retirement income at t = 15. As
stated earlier, computing the retirement present value requires two sub- steps. First, find the present value at t = 40 with the formula in
Equation 11; second, discount this present value back to t = 15. Now we can find the retirement income present value at t = 40:
A = $100,000
r = 8% = 0.08
N=20
PV = A [ 1 − 1/(1+r)N]/ r = $100,000[1 – 1/(1.08)20]/ 0.08 = $100, 000 (9.818147) = $981, 814.74
The present value amount is as of t = 40, so we must now discount it back as a lump sum to t = 15:
Now recall that Grant will have saved $54,304.23 by t = 15. Therefore, in present value terms, the annuity from t = 16 to t = 40 must equal the
difference between the amount already saved ($54,304.23) and the amount required for retirement ($143,362.53). This amount is equal to
$143,362.53 − $54,304.23 = $89,058.30. Therefore, we must now find the annuity payment, A, from t = 16 to t = 40 that has a present value of
$89,058.30. We find the annuity payment as follows:
PV = $89,058.30
r = 8% = 0.08
N = 25
Present value annuity factor = [1 – 1/ (1 + r) N ] / r = [1 – 1/(1.08)25] = 10.674776
A = PV / Present value annuity factor = $89, 058.30 / 10.674776 = $8, 342.87
Grant will need to increase her savings to $8,342.87 per year from t = 16 to t = 40 to meet her retirement goal of having a fund equal to
$981,814.74 after making her last payment at t = 40.
13. Present and Future Value Equivalence and the Additivity Principle
LOS (c) : Calculate and interpret the future value (FV) and present value (PV) of a single sum of money, an ordinary annuity, an
annuity due, a perpetuity (PV only), and a series of unequal cash flows
LOS (d) : Demonstrate the use of a time line in modelling and solving time value of money problems
The cash flow additivity principle is the idea that amounts of money indexed at the same point in time are additive—is one of
the most important concepts in time value of money mathematics. An illustration is given below :
8
Two series of cash flows labeled A and B are presented. Assuming an annual interest rate of 2 percent, we can find the future
value of each series as follows:
By the method used so far, the future value of (A + B) is found by adding these values: $202 + $404 = $606.
Alternatively, the future value can be determined by adding the cash flows of each series (A and B) to create a combined series,
labeled A + B. The future value of this combined cash flow, as shown in the third timeline, results in the same value: $300(1.02) +
$300 = $606.
9
Quantitative Methods - Learning Module 2
for the CFA exam
Organizing, Visualizing and Describing Data
Learning Outcomes :
1. Introduction
The availability of large volumes of diverse data, along with advances in machine learning, is reshaping the investment industry.
While this presents significant opportunities, effectively using this data requires careful organization, cleaning, and analysis. In
fact, a majority of an analyst's time is often spent on these tasks. Once data is properly prepared, it can reveal important
relationships and insights, laying the foundation for effective investment strategies. This reading introduces fundamental
concepts essential for investment professionals, setting the stage for more advanced tools and techniques in the CFA
curriculum.
2. Data Types
Properly categorizing data types is crucial for selecting appropriate statistical methods and visualizations. This involves
considering distinctions such as numerical versus categorical data, cross-sectional versus time-series versus panel data, and
structured versus unstructured data.
From a statistical perspective, data can be classified into two basic groups: numerical data and categorical data:
Another data classification standard is based on how data are collected, and it categorizes data into three types: cross-sectional,
time series, and panel.
In the context of data types, it's essential to understand two key terms: "variable" and "observation." A variable is a
characteristic or quantity that can be measured, counted, or categorized, like stock price or dividend yield. An observation, on
the other hand, is a specific value of a variable collected at a particular point in time or over a defined period, such as DEF, Inc.'s
EPS of $7.50 last year, which showed a 15% annual increase.
Cross-sectional data involve observations of a specific variable from various entities at a single point in time. These entities can
be individuals, companies, regions, etc. For instance, the inflation rates of different European countries in January.
Time-series data, on the other hand, consist of a sequence of observations for one entity over time at regular intervals, like
daily, weekly, or annually. For example, the daily closing prices of a stock in a particular month form time-series data.
Panel data combine elements of both time-series and cross-sectional data. They include observations over time for one or more
variables across multiple entities and are often organized in a data table. Panel data, as shown below, are commonly used in
financial analysis and modelling.
Structured data are well-organized and typically follow established patterns. They often take the form of one-dimensional
arrays (like time series) or two-dimensional data tables with rows and columns. Structured data are easy to input, store, query,
and analyze without much manual processing. Examples include market data from stock exchanges, fundamental financial data
from statements, and analytical data generated from forecasts and projections.
Unstructured data lack conventional organization and can include text (like news articles and social media posts) as well as
audio and video content (such as earnings calls). This category has emerged with the growth of alternative data, collected from
non-traditional sources like electronic devices and social media. Unstructured data can be classified by source: individual-
generated (e.g., social media posts), business-generated (e.g., credit card transactions), and sensor-generated (e.g., satellite
imagery). While unstructured data can provide unique market insights, incorporating them into financial analysis is challenging.
Financial models typically work with structured data, so unstructured data needs to be transformed for use in these models.
LOS (b) : Describe how data are organized for quantitative analysis
Quantitative analysis and modelling require data to be clean and organized. Raw data are often unsuitable for direct use by
analysts. Depending on the number of variables, data can be organized into two common formats for quantitative analysis: one-
dimensional arrays and two-dimensional rectangular arrays.
A one-dimensional array is used for representing a single variable, often in a time-ordered sequence, making it easy to update
with new data. This format preserves valuable information beyond basic statistics, helping to identify trends and patterns over
time.
A two-dimensional rectangular array, similar to an Excel spreadsheet, is a popular way to organize data for computer
processing or human consumption. It uses columns to represent variables and rows to hold observations. When used for a single
entity (like a company), each column represents a variable, and each row holds observations for those variables over successive
time periods. It must be ensured that observations for different variables are sorted and aligned on the same time scale.
A frequency distribution is a tabular presentation of data, categorized either by counting observations for distinct values or
grouping numerical values into ordered bins. It's a valuable tool for summarizing data initially, making it easier to interpret.
Creating a frequency distribution for a categorical variable involves two basic steps:
1. Counting the number of observations for each unique value of the variable.
2. Constructing a table listing these unique values and their respective counts, often sorted by count for easier display.
For example, the table below shows a frequency distribution of a portfolio's stock holdings by sectors, counted and summarized
in absolute and relative frequencies. Absolute frequency represents the actual number of observations for each sector, while
relative frequency is calculated as a percentage of each unique value compared to the total number of observations. This table
provides a snapshot of the data, aiding in identifying patterns and comparisons between datasets.
Frequency distributions are valuable for analyzing large sets of numerical data. Constructing a frequency distribution for
numerical data involves several steps:
When rounding the bin width in Step 4, it's advisable to round up to ensure the final bin includes the maximum data value. In
practice, refinements like starting bins at whole numbers and using the nearest whole number below the minimum value can
improve interpretation. A frequency distribution groups data into distinct bins, with each observation assigned to one bin. The
frequency distribution is essentially a list of these bins and their associated measures of frequency. For example, suppose we
have data for an equity index EAA, daily returns of the EAA Equity Index spans a five-year period and consists of 1,258
observations with a minimum value of −4.1% and a maximum value of 5.0% ( range = 5- (-4.1) = 9.1%). To determine an
appropriate value for k (the number of bins), it's important to consider the usefulness of the resulting bin width. If there are a
large number of empty bins, it may indicate that we're trying to provide too much detail. Starting with a small bin width, we can
assess if most bins are empty and if k is too large. For instance, if we prefer whole number bin widths for ease of interpretation,
in the case of daily EAA Equity Index returns, a 1% bin width would be associated with approximately 10 bins, covering a range
of 10%. For k = 10, the frequency distribution for the daily returns of the EAA Equity Index are as follows :
Return Bin (%) Absolute Frequency Relative Frequency (%) Cumulative Absolute Frequency Cumulative Relative
Frequency (%)
−5.0 to −4.0 1 0.08 1 0.08
−4.0 to −3.0 7 0.56 8 0.64
−3.0 to −2.0 23 1.83 31 2.46
−2.0 to −1.0 77 6.12 108 8.59
−1.0 to 0.0 470 37.36 578 45.95
0.0 to 1.0 555 44.12 1,133 90.06
1.0 to 2.0 110 8.74 1,243 98.81
2.0 to 3.0 13 1.03 1,256 99.84
3.0 to 4.0 1 0.08 1,257 99.92
4.0 to 5.0 1 0.08 1,258 100.00
12
Note : Readers should mark this table as its data will be used in later parts of this module.
As the above table shows, two additional useful ways to present data are cumulative absolute frequency and cumulative relative
frequency. Cumulative absolute frequency adds up the absolute frequencies from the first bin to the last bin, while cumulative
relative frequency is a sequence of partial sums of the relative frequencies. In the last bin, the cumulative absolute frequency
will match the total number of observations, and the cumulative relative frequency will reach 100%.
We have shown that the frequency distribution table is a powerful tool to summarize data for one variable. Using a contingency
table, we can summarize data for two variables simultaneously.
A contingency table is a tabular format that simultaneously displays the frequency distributions of two or more categorical
variables, aiming to identify patterns between them. When used for two categorical variables, it's called a two-way table.
Contingency tables list one variable's levels as rows and the other variable's levels as columns. An R × C table has R levels for one
variable and C levels for the other. Each variable in a contingency table must have a finite number of levels, whether ordered
(ordinal data) or unordered (nominal data). The data in the table cells can represent either frequencies (counts) or relative
frequencies (percentages), based on overall totals, row totals, or column totals. A sample of this type of table is shown below :
In a contingency table, the cells display the number of stocks in each sector with a specific market capitalization level. For
instance, there are 275 small-cap health care stocks, which is the most frequent subgroup in the portfolio. These values are
referred to as joint frequencies because they combine one variable from the rows (sector) and another variable from the
columns (market cap) to count observations. The joint frequencies are summed across rows and columns, and these totals are
called marginal frequencies. For example, the marginal frequency of health care stocks is 435, which is the sum of the joint
frequencies across all market cap levels. Similarly, the marginal frequency of small-cap stocks is 575, calculated by adding the
joint frequencies across all sectors.
Also, similar to the one-way frequency distribution table, we can express frequency in percentage terms as relative frequency
depending on the values involved. For example, Small-cap healthcare form 27.5% of the overall data, healthcare forms 43.5% of
the overall data while small-cap forms 57.5% of the overall data. Inferences can be made depending on the values/totals
involved.
Contingency tables have various applications, including assessing the performance of a classification model, known as a
confusion matrix. For instance, if we have a model for classifying companies into two groups (defaulting on bond payments or
not), the confusion matrix would be a 2 × 2 table displaying the actual defaults versus the model's predicted defaults, as shown
below :
Another application of contingency tables is to investigate potential association between two categorical variables. One way to
test for a potential association between categorical variables is to perform a chi-square test of independence. This involves
using the marginal frequencies in the contingency table to create a table with expected values for observations. The actual and
expected values are then used to calculate the chi-square test statistic, which is compared to a value from the chi-square
13
distribution at a specific significance level. If the test statistic is greater than the chi-square distribution value, it suggests a
significant association between the categorical variables, and the claim of independence can be rejected.
Example
Contingency Tables and Association between Two Categorical Variables
Suppose we randomly pick 315 investment funds and classify them two ways: by fund style, either a growth fund or a value fund; and by risk
level, either low risk or high risk. Growth funds primarily invest in stocks whose earnings are expected to grow at a faster rate than earnings
for the broad stock market. Value funds primarily invest in stocks that appear to be undervalued relative to their fundamental values. Risk
here refers to volatility in the return of a given investment fund, so low (high) volatility implies low (high) risk. The data are summarized in a 2
× 2 contingency table shown
1. Calculate the number of growth funds and number of value funds out of the total funds.
Solution : The task is to calculate the marginal frequencies by fund style, which is done by adding joint frequencies across the rows. Therefore,
the marginal frequency for growth is 73 + 26 = 99, and the marginal frequency for value is 183 + 33 = 216.
2. Calculate the number of low-risk and high-risk funds out of the total funds.
Solution : The task is to calculate the marginal frequencies by fund risk, which is done by adding joint frequencies down the columns.
Therefore, the marginal frequency for low risk is 73 + 183 = 256, and the marginal frequency for high risk is 26 + 33 = 59.
3. Describe how the contingency table is used to set up a test for independence between fund style and risk level.
Solution : Based on the procedure mentioned for conducting a chi-square test of independence, we would perform the following three steps.
Step 1: Add the marginal frequencies and overall total to the contingency table. We have also included the relative frequency table for
observed values.
Step 2: Use the marginal frequencies in the contingency table to construct a table with expected values of the observations. To determine
expected values for each cell, multiply the respective row total by the respective column total, then divide by the overall total. So, for cell i,j
(in ith row and jth column): Expected Value i,j = (Total Row i × Total Column j)/Overall Total
For example, Expected value for Growth/Low Risk is: (99 × 256)/ 315 = 80.46; and Expected value for Value/High Risk is: (216 × 59) / 315 =
40.46. The table of expected values (and accompanying relative frequency table) are:
Step 3: Use the actual values and the expected values of observation counts to derive the chi-square test statistic, which is then compared to a
value from the chi-square distribution for a given level of significance. If the test statistic is greater than the chi-square distribution value, then
there is evidence of a significant association between the categorical variables.
6. Data Visualization
LOS (e) : Describe ways that data may be visualized and evaluate uses of specific visualizations
14
LOS (f) : Describe how to select among visualization types
Visualization is the presentation of data in a pictorial or graphical format for the purpose of increasing understanding and for
gaining insights into the data.
A histogram is a chart that displays the distribution of numerical data. It represents each data interval with a bar or column,
where the height of the bar corresponds to the absolute or relative frequency of that interval. For continuous variables, data is
first divided into intervals or bins and summarized in a frequency distribution table as we did for EAA index earlier. The y-axis
typically represents absolute or relative frequency (in percentage terms), and the x-axis shows the bins of the variable. The bars
are of equal width, often with no gaps between them, although small gaps can enhance readability. The height of each bar can
indicate the absolute frequency or relative frequency.
Histograms are useful for visualizing a large amount of numerical data that has been grouped into a frequency distribution. They
allow for a quick assessment of the distribution's shape, centre, and spread. Whether to use absolute or relative frequencies in a
histogram depends on the specific question being addressed. An absolute frequency histogram reveals the number of items in
each bin, while a relative frequency histogram shows the proportion or percentage of total observations in each bin.
If we join the mid-point of all the histogram bars, the resulting shape is called a frequency polygon. The following illustration
shows a histogram for the EAA index data used earlier, with a frequency polygon laid over it :
Histograms efficiently display the frequency distribution of numerical data, while categorical data's frequency distribution can
be visualized using a bar chart. In a bar chart, each bar represents a category, with the height of the bar corresponding to the
category's frequency. When categories in a bar chart are ordered by frequency in descending order and include a line showing
cumulative relative frequency, it's known as a Pareto Chart. This chart is often used to emphasize dominant categories or
essential groups.
15
When dealing with two categorical variables and wanting to display their joint frequencies, an enhanced version called a
grouped bar chart (or clustered bar chart) is used. An alternative way to display the joint frequency distribution of two
categorical variables is through a stacked bar chart. In this chart, bars representing sub-groups are stacked on top of each other
to create a single bar, offering a visual representation of the relationship between the variables. Both these are illustrated below
uses data from the contingency table discussed above : Left - Grouped bar chart, Right - Stacked bar chart
Bar charts are a clear and efficient way to present the frequency distribution of categorical data. They can also be extended for
cases where categorical data are related to numerical data. For instance, a vertical bar chart can depict a company's quarterly
profits over the past year, with each bar representing a quarter and its height indicating the profit value for that period.
6.3. Tree-Map
16
6.5. Line Chart
A scatter plot is a graph used to visualize the joint variation in two numerical
variables, making it a valuable tool for exploring potential relationships
between these variables. It's constructed with one variable on the x-axis and
the other on the y-axis, with dots indicating the values of both variables for
specific data points. They offer important insights, helping to identify potential
associations between variables. The pattern in the scatter plot can suggest no
clear relationship, a linear relationship, or a non-linear relationship. Randomly
distributed data points indicate no apparent association, while data points
forming a straight line imply a significant relationship. The slope of the line
indicates whether the association is positive or negative, and the closeness of
data points to the line reveals the strength of the relationship.
A sample plot is
shown alongside to
compare the relative
performance of IT
sector with the
broader market and to assess the strength of the relationship between
their performances.
Scatter plots are valuable for identifying patterns and extreme values
between two variables. In scenarios where we need to assess pairwise
associations among many variables, like feature selection for
predictive modelling, a scatter plot matrix is a useful tool. It organizes
scatter plots for pairs of variables, allowing inspection of all pairwise
relationships in one visual display. However, it's essential to remember
that scatter plots and scatter plot matrixes should complement, not
replace, robust statistical tests for the best results. A small section of a
scatter plot matrix is shown on the right, where along with IT and S&P
500, utilities and financials were included for further analysis between
the sectoral relationships.
17
6.7. Heat Map
A flow chart is given below to help guide an analyst select the right visualization technique :
Data visualization is a powerful tool for gaining insights from data, but it can be misleading if not used carefully. There are four
common pitfalls to avoid:
1. Choosing an inappropriate chart type can hinder accurate data interpretation. For example, using a scatter plot is more
effective for investigating the correlation between two variables compared to plotting them separately in a line chart.
2. Selectively presenting data for a short time period can create the illusion of trends that are actually just noise. It's important
to consider longer timeframes for a more accurate assessment.
3. Using a truncated graph with a y-axis that doesn't start at zero can exaggerate differences and lead to incorrect conclusions.
This is particularly relevant when comparing close values.
4. Improperly scaling axes, such as setting an unnecessarily high maximum on the y-axis, can distort the appearance of the
data making it look less/more steep or less/more volatile than it actually might be. It's crucial to use appropriate scaling to
maintain accuracy.
Analysts should be mindful of these pitfalls to ensure accurate and ethical use of data visualization.
18
7. Measures of Central Tendency
In this section, we explore quantitative measures used to analyze data, focusing on central tendency and other location-related
measures. Central tendency measures indicate where data are centred and are widely used for their simplicity and practicality.
We discuss various central tendency measures like the arithmetic mean, median, mode, weighted mean, geometric mean, and
harmonic mean. Additionally, we cover other measures of location, including quartiles, quintiles, deciles, and percentiles. These
measures provide a deeper understanding of data beyond summarization techniques like frequency distributions, histograms,
and contingency tables.
Statistics are summary measures of a set of observations, and descriptive statistics help summarize central tendency and the
spread of data distribution. When a statistic summarizes all possible observations of a population, it's called a parameter. If it
summarizes a subset of the population, it's a sample statistic. We typically focus on sample statistics because investment
managers often work with samples rather than entire populations of data.
The arithmetic mean is the sum of the values of the observations divided by the number of observations.
∑𝑛
𝑖 𝑋𝑖
➔ 𝑋 = where n is the number of observations in the sample
𝑛
7.1.3. Outliers
When dealing with extreme values in a financial dataset, you have three options:
➔ Do nothing and use the data without any adjustment, suitable when the values are legitimate and important to reflect
the entire sample distribution.
➔ Delete all the outliers, and consider using a trimmed mean by excluding a certain percentage of the lowest and highest
values to compute the mean of the remaining values.
➔ Replace the outliers with other values, and consider using a winsorized mean, which involves assigning a percentage of
the lowest values to a specified low value and a percentage of the highest values to a specified high value before
computing the mean from the restated data.
Choice among these options depends on the nature of the data and the specific requirements of your analysis.
The median is the middle value in a sorted set of data. In an odd-numbered sample, it's the value at the (n + 1)/2 position. In an
even-numbered sample, it's the average of the values at the n/2 and (n + 2)/2 positions.
The median is less affected by extreme values compared to the mean, making it useful for data with outliers. However, it
doesn't consider the size of the observations , only the relative position of the observation counts, and median can be more
complex to calculate than the mean. This is because it requires sorting the data and different calculations depending on the
sample size.
19
7.3. The Mode
The mode is the most frequently occurring value in a distribution. A distribution can have one mode (unimodal), two modes
(bimodal), or three modes (trimodal), or it can have no mode if all values are different. Continuous data may not have a clear
mode, but when grouped into intervals, the interval with the highest frequency is considered the modal interval.
where the sum of the weights equals 1. If each weight becomes 1/n, then this becomes the formula for the arithmetic mean.
➔ 𝑋G = 𝑛√𝑋1, 𝑋2, . . . , 𝑋𝑛
with Xi ≥ 0 for I = 1 , 2,..., n. The geometric mean has a solution if all observations (Xi) are greater than or equal to zero, and it
exists when the product under the square root sign is non-negative.
When calculating the geometric mean for risky assets with potentially negative returns, we transform the returns to make them
positive by adding 1.0 to each return expressed as a decimal (1 + R t). This transformation ensures that the observations will
never be negative because the most negative return is -100%. The geometric mean of 1 + Rt is computed, and then we subtract
1.0 from this result to obtain the geometric mean of the individual returns :
𝑛
➔ 1 + RG = √(1 + 𝑅1 )(1 + 𝑅2 ) … . (1 + 𝑅𝑇 )
The geometric mean is always less than or equal to the arithmetic mean. They are equal only when there is no variability in the
observations, meaning that all the observations in the series are the same. Typically, the difference between the arithmetic and
geometric means increases with the variability within the sample. In other words, the more dispersed or variable the
observations, the greater the difference between the arithmetic and geometric means.
The harmonic mean is obtained by summing the reciprocals of the observations—terms of the form 1/Xi—then averaging
using the number of observations n, and, finally, taking the reciprocal of the average. It can be viewed as a special type of
weighted mean in which an observation’s weight is inversely proportional to its magnitude. A well-known application
arises in the investment strategy known as cost averaging, which involves the periodic investment of a fixed amount of
money.
Example
Cost Averaging and the Harmonic Mean
Suppose an investor purchases €1,000 of a security each month for n = 2 months. The share prices are €10 and €15 at the two purchase dates.
What is the average price paid for the security?
The average price paid is in fact the harmonic mean of the asset’s prices at the purchase dates. Using the harmonic mean equation, the price is
2/[(1/10) + (1/15)] = €12. The value €12 is less than the arithmetic mean purchase price (€10 + €15)/2 = €12.5.
However, we could find the correct value of €12 using the weighted mean formula, where the weights on the purchase prices equal the shares
purchased at a given price as a proportion of the total shares purchased. In our example, the calculation would be (100/166.67)€10.00 +
(66.67/166.67)€15.00 = €12.
Note : If we had invested varying amounts of money at each date, we could not use the harmonic mean formula. We could, however, still use
the weighted mean formula.
The arithmetic, geometric, and harmonic means are mathematically related to one another in the following manner :
The choice of central measure to use depends on factors like : Are there outliers that we want to include? Is the distribution
symmetric? Is there compounding? Are there extreme outliers? For an easy way to decide which measure to use, we can use
the adjoining illustration as a guide :
8. Quantiles
We now discuss a method for describing the location of data by identifying values at or below which specified proportions of the
data are found. We introduce the concept of quantiles (or fractiles) as a general term for these values and mentions commonly
used quantiles such as quartiles, quintiles, deciles, and percentiles and their relevance in investment analysis
Percentiles are a way to divide a data distribution into smaller segments. The median splits the data in half, and other
percentiles like quartiles, quintiles, and deciles divide it into quarters, fifths, or tenths. Each percentile represents the value at
or below which a certain percentage of observations lie. For instance, the first quartile is also the 25th percentile, the second
quartile is the 50th percentile, and the third quartile is the 75th percentile. The interquartile range (IQR) is the difference
between the third and first quartiles, which measures the spread of the middle 50% of the data.
21
If Py is the value below which y% of the distribution lies, then the location of that value L y, in an array of n entries sorted in
ascending order is given by :
➔ Ly = (n+1) y/100 where y is the percentage point at which we are dividing the distribution.
A box and whisker plot shown along, is a useful way to visualise the percentile relate data
Quantiles are useful in investments for portfolio performance evaluation, investment strategy development, and research.
Investment analysts often use quantiles to rank performance, such as the percentile or quartile in which portfolios or
investment managers fall compared to their peers. Quantiles are also employed in investment research to evaluate the impact
of characteristics like market value on investment outcomes. For example, companies are ranked and sorted into deciles based
on their market value to compare the performance of small and large companies in empirical finance studies.
9. Measures of Dispersion
In investments, understanding both the expected return and how returns are dispersed around that mean is crucial. Expected
return indicates where returns are centered, while dispersion measures the variability around that center, addressing the
concept of risk. This section discusses common measures of dispersion, including range, mean absolute deviation, variance, and
standard deviation. These measures assess absolute dispersion, which quantifies variability without reference to any specific
point or benchmark.
The range is the difference between the maximum and minimum values in a dataset:
Although range is easy to calculate but it does have some limitations like It only considers the maximum and minimum values in
the data, providing no information about the distribution's shape. This simplicity means it can be influenced by extreme outliers,
potentially giving an inaccurate representation of the overall data.
Our previous discussion on the arithmetic mean emphasized the importance of deviations from the mean (Xi − 𝑋) in statistics.
Calculating dispersion as the average of these deviations presents a problem because they always sum to 0. To overcome this
issue, we use absolute deviations around the mean, as seen in the mean absolute deviation, which is also referred to as the
average absolute deviation. For a data sample, it is given by :
Where 𝑋 is the sample mean, n is the number of observations in the sample, and “| |” represents the absolute value i.e. even if
negative inside the bars, the value is taken as positive for computation ( for eg: | -1 | = 1 )
22
The mean absolute deviation uses all of the observations in the sample and is thus superior to the range as a measure of
dispersion. One technical drawback of MAD is that it is difficult to manipulate mathematically compared with sample variance
discussed next.
As discussed in previous section, simply averaging the deviation will give zero value. To solve it we can either take the mean-
absolute value as before or we can square the individual deviations and then take their average as a mathematical measure of
dispersion. Variance is defined as the average of the squared deviations around the mean. Standard deviation is the positive
square root of the variance.
Notice the n-1, we do this because we are using a sample statistic as a measure for an entire population. Using n-1 instead
of n improves the estimate for the population.
9.3.3. Dispersion and the Relationship between the Arithmetic and the Geometric Means
The gap between the arithmetic mean and geometric mean is related through the sample variance as :
➔ 𝑋𝐺 ≈ 𝑋 - ( s2/ 2 )
Larger the variance of the sample, the wider the difference between the geometric mean and the arithmetic mean.
Asset risk is often measured using the variance or standard deviation of returns, which consider both upside and downside risks.
However, investors are typically more concerned with downside risk, which focuses on returns below a specified target or
minimum acceptable return. Analysts have developed measures of downside risk, such as the target downside deviation (or
target semi deviation), which quantifies the dispersion of observations (e.g., returns) below the target. This measure is
calculated by identifying observations below the target, summing the squared negative deviations from the target, dividing by
the sample size minus one, and taking the square root, as shown below :
The standard deviation is easier to interpret than variance because it uses the same units as the data. However, it can be
challenging to compare the relative variability of datasets with different means or units. To address this, the coefficient of
variation is introduced, which measures relative dispersion, allowing for meaningful comparisons, especially in situations
involving datasets with different means or units of measurement. Co-efficient of variation (CV) formula is given as :
23
The coefficient of variation (CV) measures the risk (standard deviation) per unit of reward (mean return), particularly when
dealing with returns. However, it can be meaningless if the mean return is negative. The CV can be expressed as a multiple or a
percentage, allowing for direct comparisons of variability across different datasets. As a scale-free measure, the CV is unit-
agnostic, making it suitable for comparing datasets with different units of measurement.
Skewness refers to the lack of symmetry in a distribution. A positively skewed return distribution means it has more small losses
and a few large gains, while a negatively skewed distribution has more small gains and a few large losses. The mode, median,
and mean have specific relationships in these skewed distributions. Positive skew has the mean > median > mode, making it
more attractive to investors with limited but frequent downside returns compared to less frequent but potentially unlimited
upside returns. A negative skew has mean < median < mode. Skewed distributions are illustrated below :
The approximation for computing sample skewness when n is large (100 or more) is:
where s is the sample standard deviation. A positive values indicates positive skewness, and likewise for negative. A zero value
indicates a symmetric distribution. Principle behind this calculation is that cubing the deviations from the mean preserves their
sign. In a positively skewed distribution, where the mean is greater than the median, more than half of the deviations are
negative, but the sum of cubed deviations is positive when losses are small and likely, and gains are less likely but more extreme.
This indicates that in positively skewed distributions, the average magnitude of positive deviations is greater than that of
negative deviations.
24
called platykurtic, suggesting a lower likelihood of extreme deviations. A distribution similar to the normal distribution in terms
of tail weight is called mesokurtic. Fat-tailed distributions generate more frequent extremely large deviations, while thin-tailed
ones generate less frequent large deviations compared to the normal distribution.
A normal distribution has kurtosis value of 3. Excess kurtosis is the kurtosis relative to the normal distribution. For a large
sample size (n = 100 or more), sample excess kurtosis (K ) is approximately as follows:
Most equity return series have been found to be fat-tailed. If a return distribution is fat-tailed and we use statistical models that
do not account for the distribution, then we will underestimate the likelihood of very bad or very good outcomes.
Correlation is a measure of the linear relationship between two random variables. Understanding this metric begins with how
two variables vary with respect to each other i.e. their covariance. The sample covariance (sXY) is a measure of how two
variables in a sample move together and it is given by:
The sample covariance is the average product of the deviations of observations of two random variables (Xi and Yi) from their
sample means. It measures the joint variability of the variables, with a positive covariance indicating they vary in the same
direction and a negative covariance indicating opposite variation relative to their means. Using n-1 in the denominator makes it
an unbiased estimate of population covariance.
In itself, the covariance measure is difficult to interpret as it is not normalized and so depends on the magnitude of the
variables. The normalized version of covariance is the correlation coefficient which expresses the strength of the linear
relationship between the two random variables given by :
Scatter plots discussed earlier can be used to depict and make sense of correlation between two variables. Some illustrations
are shown below :
25
Important to note that Panel D shows a scatter plot of two variables that have a non-linear relationship. Because the correlation
coefficient is a measure of the linear association between two variables, it would not be appropriate to use the correlation
coefficient in this case.
o Correlation, denoted as rXY, ranges between -1 and +1 for two random variables, X and Y: -1 ≤ rXY ≤ +1.
o A correlation of 0 means there is no linear relationship between the variables (uncorrelated).
o A positive correlation close to +1 signifies a strong positive linear relationship, and a correlation of 1 means a perfect linear
relationship.
o A negative correlation close to -1 indicates a strong negative linear relationship, and a correlation of -1 represents a perfect
inverse linear relationship.
o Two variables can have a strong nonlinear relation and still have a very low correlation. Panel D in the illustration above
presents such a case.
o Correlation may be quite sensitive to outliers which is why it might be an unreliable measure when they are present in one
or both of the variables. It's essential to consider whether it makes sense to exclude outliers and determine if they contain
meaningful information about the relationship between the variables. To decide whether to include or exclude outliers, one
should assess if removing them drastically changes the computed correlation. Techniques like trimming or winsorizing the
dataset can be used to handle outlier observations when excluding them is deemed appropriate.
o Correlation does not imply causation, meaning a strong correlation between two variables doesn't indicate that one causes
the other. Visualizations like scatter plots can also lead to incorrect assumptions about causal relationships, so it's crucial to
avoid drawing unwarranted conclusions based solely on the data.
Investment professionals should be wary of relying on high correlations, as they can lead to strategies that seem profitable but
may not be in practice.
26
Quantitative Methods - Learning Module 3
for the CFA exam
Probability Concepts
Learning Outcomes :
Investment decisions are often made in a risky environment, and probability concepts are essential tools for making consistent
and logical decisions in this context. This reading introduces key probability tools that are crucial for addressing real-world
problems related to risk, such as predicting investment manager performance, forecasting financial variables, and pricing bonds
to fairly compensate bondholders for default risk. The focus is on practical application, emphasizing concepts like independence,
expectation, and variability, which are vital for investment research and practice.
Investors primarily focus on returns when making investment decisions. Returns on risky assets are considered random
variables, which are quantities with uncertain future outcomes. These outcomes are the possible values that a random variable
can take.
Our interest in investments often extends beyond single outcomes, and this is where the concept of an "event" comes into play.
An event represents a specific set of outcomes. For instance, an event could be as specific as the portfolio earning a return of
exactly 10%. We can also define another event that represents returns below 10%, encompassing all possible returns greater
than or equal to -100% (the worst possible return) but less than 10%. After these comes Probability which helps answer the
question : how likely is it that the portfolio will earn a return below 10%? Or the likelihood of the event occurring ? It is always a
number between 0 to 1.
27
In investments, we frequently estimate the probability of an event by examining its historical occurrence as a relative frequency.
This approach results in empirical probability. For instance, if 51 out of 60 stocks in a large-cap equity index pay dividends, the
empirical probability of a stock in the index paying a dividend is calculated as 51/60, which is equal to 0.85. In the realm of
investments, empirical probabilities rely on stable historical relationships to be accurate, and they can't be calculated for events
absent from historical records or very rare events. In some cases, we may adjust empirical probabilities to accommodate
changing perceptions of relationships, or we may need to rely on personal judgments to assess probabilities when no empirical
data is available. This leads to subjective probabilities, which play a significant role in investment decisions. Alternatively, for
well-defined problems, we can deduce probabilities through logical analysis, resulting in a priori probabilities. A priori and
empirical probabilities, as they tend to be consistent across individuals, are often grouped as objective probabilities.
Probabilities encountered in investments can also be expressed in terms of odds, like "the odds for E" or "the odds against E."
The probability of an event is the expected frequency of its occurrence, and the odds for an event are calculated as the
probability of the event happening divided by the probability of it not happening. For example, if a football team has a 0.25
probability of winning the World Cup and a 0.75 probability of losing, the odds for winning are 0.25/0.75, which equals 0.33. If
the probability of an event is P(E), the odds relations are :
➔ Odds for E = P(E)/[1 − P(E)]. Given odds for E of “a to b,” the implied probability of E is a/(a + b).
➔ Odds against E = [1 − P(E)]/P(E). Given odds against E of “a to b,” the implied probability of E is b/(a + b).
In the context of probability, there are unconditional and conditional probabilities. Unconditional probability, denoted as P(A),
answers the question, "What is the probability of event A?" It's a straightforward probability calculation where the numerator is
the sum of probabilities of A occurring, and the denominator is 1, representing all possible outcomes.
Conditional probability, denoted as P(A|B), addresses the question, "What is the probability of A, given that B has occurred?"
This involves restricting the context to event B. The conditional probability is calculated as the ratio of the sum of probabilities of
A happening (numerator) to the sum of probabilities of all outcomes related to B (denominator). It is used when one event
depends on another and can differ from the unconditional probability, which is the standalone probability without any
restrictions. To calculate it, we introduce the concept of joint probability, denoted as P(AB), which represents the probability of
both A and B happening. The conditional probability of A given B is then calculated as :
For example, suppose B happens half the time, P(B) = 0.50, and A and B both happen 10% of the time, P(AB) = 0.10. What is the
probability that A happens, given that B happens? That is P(A | B) = P(AB)/P(B) = 0.10 / 0.50 = 0.20.
If we re-arrange the above equation, we get Multiplication Rule for Probability which we can write as :
If B happens 50% of the time, and the probability of A given that B happens is 20%, the joint probability of A and B happening is
P(AB) = P(A | B)P(B) = 0.20 × 0.50 = 0.10.
Example
Conditional Probabilities and Predictability of Mutual Fund Performance
An analyst conducts a study of the returns of 200 mutual funds over a two-year period. For each year, the total returns for the funds were
ranked, and the top 50% of funds were labelled winners; the bottom 50% were labelled losers. Exhibit 3 shows the percentage of those funds
that were winners in two consecutive years, winners in one year and then losers in the next year, losers then winners, and finally losers in both
years. The winner–winner entry, for example, shows that 66% of the first-year winner funds were also winners in the second year. The four
entries in the table can be viewed as conditional probabilities.
28
Year 1 Winner 66% 34%
Year 1 Loser 34% 66%
1. State the four events needed to define the four conditional probabilities.
Solution to 1: The four events needed to define the conditional probabilities are as follows:
2. State the four entries of the table as conditional probabilities using the form P(this event | that event) = number.
Solution to 2: From Row 1: P(fund is a Year 2 winner | fund is a Year 1 winner) = 0.66 P(fund is a Year 2 loser | fund is a Year 1 winner) = 0.34
From Row 2: P(fund is a Year 2 winner | fund is a Year 1 loser) = 0.34 P(fund is a Year 2 loser | fund is a Year 1 loser) = 0.66
Solution to 3: These probabilities are calculated from data, so they are empirical probabilities.
4. Using information in the table, calculate the probability of the event a fund is a loser in both Year 1 and Year 2. (Note that because 50% of
funds are categorized as losers in each year, the unconditional probability that a fund is labelled a loser in either year is 0.5.)
Solution to 4: The estimated probability is 0.33. Let A represent the event that a fund is a Year 2 loser, and let B represent the event that the
fund is a Year 1 loser. Therefore, the event AB is the event that a fund is a loser in both Year 1 and Year 2. From Exhibit 3, P(A | B) = 0.66 and
P(B) = 0.50.
P(AB) = P(A | B)P(B) = 0.66(0.50) = 0.33 or a probability of 0.33. Note that this equation states that the joint probability of A and B equals the
probability of A given B times the probability of B. Because P(AB) = P(BA), the expression P(AB) = P(BA) = P(B | A)P(A) is equivalent to the one
discussed before.
Independence and dependence are crucial concepts for investment analysts, influencing decisions related to variables for
investment analysis, predictability of asset returns, and selecting superior investment managers based on their records.
The definition of independent events is that two events, A and B, are independent if and only if P(A | B) = P(A) or P(B | A) = P(B).
In simpler terms, independence implies that knowing the occurrence of one event provides no information about the likelihood
of the other event, due to which the conditional probability is the same as the unconditional probability. In this case, the
multiplication rule for independent events simplifies to : P(AB) = P(A)P(B).
When two events are not independent, they are dependent, meaning the occurrence of one is related to the occurrence of the
other. Information about a dependent event can be useful for forecasting, unlike an independent event which won't provide any
useful insights. For example, if a biotech company is announced to be acquired at an attractive price, and the stock prices of
pharmaceutical companies rise as a result, then we can say the stock prices are not independent of the takeover announcement
event (i.e. the occurrence of event A – biotech takeover, does provide some information for event B : stock price movements in
other same sector firms). On the other hand, if two events are mutually exclusive, the occurrence of one event means the other
mutually exclusive event cannot occur.
29
In practical problem-solving, we often formulate scenarios to assess the likelihood of an event. When these scenarios are
mutually exclusive and exhaustive, covering all possible outcomes, we can use the total probability rule. This rule calculates the
unconditional probability of the event in terms of probabilities conditional on the scenarios. If S represents an event, then its
complement event, i.e. event containing all other outcomes not in S, is denoted by S C and we get P(S) + P(SC) = 1 since either ‘S’
or ‘not-S’ must occur. The total probability rule is expressed as:
Where S1, S2, ..., Sn are mutually exclusive and exhaustive scenarios or events. This would mean that the probability of any event
would be weighted average of probabilities of that particular event given certain scenarios. The weights here refer to the
scenario probabilities PS1, PS2 etc. A visualisation is shown along :
LOS (h) : Calculate and interpret the expected value, variance, and standard deviation of random variables
LOS (i) : Explain the use of conditional expectation in investment applications
LOS (j) : Interpret a probability tree and demonstrate its application to investment problems
The expected value of a random variable is the probability-weighted average of the possible outcomes of the random variable.
For a random variable X, the expected value of X is denoted as E(X).
where Xi is one of n possible outcomes of the random variable X. A simple way to look at expected value is from a mean
perspective, what is the average outcome of a particular event. However, since it is a forecasted value, a risk measure is also
important. Variance and standard deviations can be used to measure the dispersions :
➔ Variance of a random variable - σ2(X) = E { [ X – E (X)] 2} and standard deviations would be its square root. Not going into the
detailed mathematics here, this variance equation can be re-written as :
➔ σ2(X) = P(X1) [X1 − E(X)]2 + P(X2) [X2 − E(X)]2 + ... + P(Xn) [ Xn − E(X)]2 = ∑𝑛𝑖=1 𝑃(Xi) [Xi − E(X)] 2
In investments, we utilize available information to make forecasts, and when we adjust our expectations based on new
information or events, we work with conditional expected values. The conditional expected value of a random variable X given
an event or scenario S is denoted as E(X | S). It's calculated as :
For expected values, we can also apply the total probability rule, known as the total probability rule for expected value. Given
as
Where, S and SC are complements to each other and S1, S2, ..., Sn are mutually exclusive and exhaustive scenarios or events.
30
Example
BankCorp’s Earnings per Share
As part of work as a banking industry analyst, we build models for forecasting earnings per share of the banks you cover. Today we are
studying BankCorp EPS data. We have recorded a probability distribution for BankCorp’s EPS for the current fiscal year as :
What are the variance and standard deviation of BankCorp’s EPS for the current fiscal year?
Solution : The order of calculation is always expected value, then variance, then standard deviation. Expected value has already been
calculated. Following the definition of variance above, calculate the deviation of each outcome from the mean or expected value, square each
deviation, weight (multiply) each squared deviation by its probability of occurrence, and then sum these terms.
σ2(EPS) = P($2.60)[$2.60 − E(EPS)]2 + P($2.45)[$2.45 − E(EPS)]2 + P ($2.20) [$2.20 − E (EPS) ]2 + P ($2.00) [$2.00 − E (EPS) ]2
= 0.15 (2.60 − 2.34)2 + 0.45 (2.45 − 2.34)2 + 0.24 (2.20 − 2.34)2 + 0.16 (2.00 − 2.34)2
= 0.01014 + 0.005445 + 0.004704 + 0.018496 = 0.038785
Standard deviation is the positive square root of 0.038785: σ(EPS) = 0.0387851/2 = 0.196939, or approximately 0.20.
LOS (k) : Calculate and interpret the expected value, variance, standard deviation, covariances, and correlations of portfolio
returns
Given a portfolio with n securities, the expected return on the portfolio (E(Rp)) is given by :
Where Ri is the expected return and wi is the weight of the ith security in the portfolio. For risk related measures, we need to
consider covariances between the portfolio assets. Covariance as a measure, describes how two variables move relative to each
other and was discussed in a previous module. Portfolio variance, in general can be given as :
The general expression for variance of a portfolio of size n (which can be obtained by using the above equations and applying
substitutions) is given by :
Note : An important property of covariances which comes in handy when solving problems is that Cov(R2,R1) = Cov(R1,R2)
A related concept is correlation between two variables which was discussed in an earlier module. As a recap, correlation
between two random variables is defined as :
31
As was discussed, like covariance, the correlation coefficient is also measure of linear association. However, the division in the
definition makes correlation a pure number (without a unit of measurement) and places bounds on its largest and smallest
possible values, which are +1 and –1, respectively.
LOS (l) : Calculate and interpret the covariances of portfolio returns using the joint probability function
Covariance between two random variables X and Y can be calculated using their joint probability function, if it can be
estimated. This function, denoted as P(X, Y), provides the probability of both X and Y taking specific values simultaneously, such
as P(X=3, Y=2), which represents the probability of X being 3 and Y being 2. If we know the joint probabilities, then the
covariance random variables RA and RB can be calculated as :
Example
Two banks : BankCorp stock returns (RA) and NewBank stock returns (RB)
The expected return on BankCorp stock is 0.20(25%) + 0.50(12%) + 0.30(10%) = 14%. The expected return on NewBank stock is 0.20(20%) +
0.50(16%) + 0.30(10%) = 15%. These values can help us in inferring the overall sectoral conditions with good conditions being when both the
banks returns are higher than expected values, average when the returns are similar and poor when returns are lower than expected. The
covariance calculation is as follows :
Banking Industry Deviations Deviations Product of Probability of Probability-Weighted
Condition BankCorp NewBank Deviations Condition Product
Good 25−14 20−15 55 0.20 11
Average 12−14 16−15 −2 0.50 −1
Poor 10−14 10−15 20 0.30 6
Cov(RA,RB) = 16
If the random variables are independent, then the joint probability function simplifies to : P(X,Y) = P(X)P(Y).
The multiplication rule for expected value of the product of independent random variables is given by :
Since, independence is a stronger property than uncorrelatedness because correlation only deals with linear relationships, the
above relation also holds for uncorrelated random variables.
6. Bayes' Formula
LOS (m) : Calculate and interpret an updated probability using Bayes’ formula
In investment decisions, our initial viewpoints are shaped by experience and knowledge. These viewpoints can be adjusted or
confirmed by new information. Bayes' formula is a rational method for updating our perspectives when we encounter new data.
Bayes' formula utilizes the total probability rule discussed before, to calculate the probability of an event based on a set of
scenarios. It essentially reverses the "given that" information, using the occurrence of the event to infer the probability of the
scenario that caused it. Bayes' formula is often referred to as inverse probability and is commonly used when updating beliefs
about potential causes that might explain new observations. Theoretically, its formula is given as :
32
Updated probability of event Probability of the new information given event
➔ = X Prior probability of event
given new information Unconditional probability of the new information
P(Information | Event)
➔ P ( Event | Information) = x P (Event)
P (Information)
Example
Bayesian analysis
Q. What is the probability a firm is a tech firm given that it has a return
of > 10% or P(tech | R > 10%)?
If we employ the bayes rule, we first need to find the probability that a firm has a return of > 10% and then the probability that a firm with a
return of > 10% is a tech firm, P(tech | R > 10%). We do these steps as follows :
P(R > 10%) = P(R > 10% | tech) × P(tech) + P(R > 10% | non-tech) × P(non-tech) = 0.60×0.20 + 0.25×0.80 = 0.32.
From the above example, we can generalise the Bayesian formula for two events A and B as :
𝑃(𝐵 | 𝐴)
➔ P(A | B) = x P(A)
𝑃(𝐵)
Which helps us calculate the probability of event A occurring, given event B has occurred. This updated probability is called
posterior probability because it reflects or comes after the new information.
7. Principles of Counting
LOS (n) : Identify the most appropriate method to solve a particular counting problem and analyze counting problems using
factorial, combination, and permutation concepts
33
In counting, the basic method is to enumerate or count outcomes one by one. However, this section focuses on shortcuts and
principles that make the counting process easier and less error-prone. These techniques are handy for complex counting
problems.
The first and basic principle of counting is the multiplication rule which states that if there are k tasks, each with a specific
number of ways it can be done (n1, n2, n3, ..., nk), then the total number of ways to accomplish all k tasks is found by multiplying
the individual ways for each task together, resulting in (n 1)(n2)(n3)...(nk).
Another counting principle relates to the labelling problems where we want to give each object in a group a label, to place it in a
category. The number of ways that n objects can be labelled with k different labels, with n1 of the first type ,n2 of the second
type and so on with n1 +n2 +...+ nk =n is given by :
For example, lets assume that we want to label 18 mutual funds into different categories, 4 into high risk category, 4 above-
average risk category , 3 average risk category, 4 below-average risk, and 3 low risk designations. For this we have approximately
13 billion different ways. This is calculated by considering the total number of possible sequences, which is 18!, ( For first
position in the first category, we have 18 choices, for the next we have 17 and so on ….) but accounting for the fact that the
order of assignment within each category doesn't matter. For example if mutual fund A, B, C, D are in high risk category, then
whether A was assigned first or D was assigned first won’t make any difference. But the simple 18! counts these as different. To
eliminate redundancies, we divide 18! by (4!)(4!)(3!)(4!)(3!).
If there are only 2 labels, then the above formula simplifies into something known as combination formula ( also known an
Binomial Formula). The number of ways that we can choose r objects from a total of n objects, when the order in which the r
objects are listed does not matter, is :
If we label the r objects as belongs to the group and the remaining objects as does not belong to the group, whatever the group
of interest, the combination formula tells us how many ways we can select a group of size r. Similar to the mutual fund example
above.
So far in our discussion, we have neglected the order in which the entries on objects are arranged within each group. If the order
matters, then we use the permutation formula ( An ordered listing is known as a permutation ) which states that the number
of ways that we can choose r objects from a total of n objects, when the order in which the r objects are listed does matter, is :
Example
Permutations and Combinations for Two Out of Four Outcomes
There are four balls numbered 1, 2, 3, and 4 in a basket. You are running a contest in which two of the four balls are selected at random from
the basket. To win, a player must have correctly chosen the numbers of the two randomly selected balls. Suppose the winning numbers are
numbers 1 and 3. If the player must choose the balls in the same order in which they are drawn, she wins if she chose 1 first and 3 second. On
the other hand, if order is not important, the player wins if the balls drawn are 1 and then 3 or if the balls drawn are 3 and then 1. The number
of possible outcomes for permutations and combinations of choosing 2 out of 4 items is illustrated below. If order is not important, for
choosing 2 out of 4 items, the winner wins twice as often.
A flow chart shown below can be useful while applying the counting methods presented so far:
34
35
Quantitative Methods - Learning Module 4
for the CFA exam
Common Probability Distributions
Learning Outcomes :
a) Define a probability distribution and compare and contrast discrete and continuous random variables and their probability
functions
b) Calculate and interpret probabilities for a random variable given its cumulative distribution function
c) Describe the properties of a discrete uniform random variable, and calculate and interpret probabilities given the discrete
uniform distribution function
d) Describe the properties of the continuous uniform distribution, and calculate and interpret probabilities given a continuous
uniform distribution
e) Describe the properties of a Bernoulli random variable and a binomial random variable, and calculate and interpret
probabilities given the binomial distribution function
f) Explain the key properties of the normal distribution
g) Contrast a multivariate distribution and a univariate distribution, and explain the role of correlation in the multivariate
normal distribution
h) Calculate the probability that a normally distributed random variable lies inside a given interval
i) Explain how to standardize a random variable
j) Calculate and interpret probabilities using the standard normal distribution
k) Define shortfall risk, calculate the safety-first ratio, and identify an optimal portfolio using Roy’s safety-first criterion
l) Explain the relationship between normal and lognormal distributions and why the lognormal distribution is used to model
asset prices
m) Calculate and interpret a continuously compounded rate of return, given a specific holding period return
n) Describe the properties of the Student’s t-distribution, and calculate and interpret its degrees of freedom
o) Describe the properties of the chi-square distribution and the F-distribution, and calculate and interpret their degrees of
freedom
p) Describe Monte Carlo simulation
LOS (a) : Define a probability distribution and compare and contrast discrete and continuous random variables and their
probability functions
LOS (b) : Calculate and interpret probabilities for a random variable given its cumulative distribution function .
This reading discusses the importance of understanding probability distributions in investment decisions. It introduces seven key
probability distributions: uniform, binomial, normal, lognormal, Student's t-, chi-square, and F-distributions, and explains their
relevance in investment analysis. Normal and binomial distributions are used in valuation models, while Student's t-, chi-square,
and F-distributions are used in statistical significance and hypothesis testing. A grasp of these distributions helps in quantitative
methods like regression analysis and time-series analysis, and the reading concludes by introducing Monte Carlo simulation as a
tool for tackling complex investment problems.
A random variable is a quantity whose future outcome is uncertain, and there are two types: discrete and continuous. Discrete
random variables have a countable number of possible values, like the number of "yes" votes at a board meeting. They can be
finite or infinite. For instance, the number of trades at a stock exchange is countable but infinite. Discrete random variables are
represented as X, and specific outcomes as x1, x2, etc. Discrete random variables can be finite or infinite, and they encompass
situations where we can count the possible outcomes.
Continuous random variables cannot be counted or listed because they have an infinite number of possible values. For
example, the volume of water in a glass is continuous because it can't be listed on a discrete scale, only measured. Investment
returns are an example of continuous random variables. In finance, continuous distributions are often used unless a variable
displays truly discrete behavior.
36
Every random variable is characterized by a probability distribution. This distribution can be understood in two ways: the
probability function, which tells us the likelihood of the variable taking a specific value, and is represented as P(X = x); for
discrete variables, it is denoted as p(x), and for continuous variables, it is represented as f(x) and known as the probability
density function (pdf). A probability function has two key properties :
The cumulative distribution function (cdf) is used when we want to find the probability of a range of outcomes instead of a
specific one. It tells us the probability that a random variable X is less than or equal to a particular value x, denoted as P(X ≤ x) or
F(x) depending on discrete or continuous variable. The cdf is derived by accumulating values of the probability function for all
outcomes less than or equal to x. It is similar to cumulative relative frequency.
LOS (c) : Describe the properties of a discrete uniform random variable, and calculate and interpret probabilities given the
discrete uniform distribution function
LOS (d) : Describe the properties of the continuous uniform distribution, and calculate and interpret probabilities given a
continuous uniform distribution
The discrete uniform distribution is the simplest probability distribution. It occurs when possible outcomes are finite and
discrete and the probability of each outcome is the same (uniform). Suppose outcomes are integers from 1-8. With eight
outcomes, the probability p(x) is 1/8 (0.125) for all values of X (1 to 8). Thus this distribution is defined by a finite number of
equally likely specified outcomes. The following table and illustration help infer the pdf and cdf for this distribution.
Note : In these kinds of questions, do keep an eye on the equality sign. There is a difference between P(4 ≤ X ≤ 6) and P(4 < X ≤
6)
From the above table we can say that the cdf has two other characteristic properties:
The continuous uniform distribution is the simplest continuous probability distribution. It serves two main purposes: it is used in
generating random numbers for Monte Carlo simulation and represents equally likely outcomes, making it a suitable model for
uncertainty when all outcomes are perceived as equally likely. The pdf for a uniform random variable is :
37
➔
For example if a =0 and b = 8, then f(x) = 1/8 and the pdf comes out as a straight line as shown :
For a discrete uniform random variable with finite outcomes, we sum individual probabilities to find the cumulative probability.
In contrast, for a continuous uniform random variable (or any continuous variable), the probability of assuming a specific fixed
value is 0. Instead, we calculate cumulative probabilities by finding the area under the probability density function (pdf) curve,
𝑏
which involves integrating the pdf { P(a ≤ X ≤ b) = ∫𝑎 𝑓(𝑥)𝑑𝑥 }. For example, to find F(3), we calculate the area under the curve
between 0 and 3 on the x-axis, which is 3/8 or 0.375. (note : integration is nothing but the area under the function, here the
function is a straight line so calculations are easier)
In the case of continuous random variables, the probability of the variable being equal to any specific point is 0. This has a
significant implication for the cumulative distribution function (cdf) of continuous random variables. For any continuous random
variable X, the probabilities of ranges like P(a ≤ X ≤ b), P(a < X ≤ b), P(a ≤ X < b), and P(a < X < b) are all equal, as the probabilities
at the endpoints (a and b) are all 0. However, as discussed before, this equality does not hold for discrete random variables
because they accumulate probability at specific points.
Example
Probability That a Lending Facility Covenant Is Breached
We are evaluating the bonds of a below-investment-grade borrower at a low point in its business cycle. We have many factors to consider,
including the terms of the company’s bank lending facilities. The contract creating a bank lending facility such as an unsecured line of credit
typically has clauses known as covenants. These covenants place restrictions on what the borrower can do. The company will be in breach of a
covenant in the lending facility if the interest coverage ratio, EBITDA/interest, calculated on EBITDA over the four trailing quarters, falls below
2.0. EBITDA is earnings before interest, taxes, depreciation, and amortization. Compliance with the covenants will be checked at the end of the
current quarter. If the covenant is breached, the bank can demand immediate repayment of all borrowings on the facility. That action would
probably trigger a liquidity crisis for the company. With a high degree of confidence, we forecast interest charges of $25 million. Our estimate
of EBITDA runs from $40 million on the low end to $60 million on the high end. Address two questions (treating projected interest charges as a
constant):
1. If the outcomes for EBITDA are equally likely, what is the probability that EBITDA/interest will fall below 2.0, breaching the covenant?
Solution to 1: EBITDA/interest is a continuous uniform random variable because all outcomes are equally likely. The ratio can take on values
between 1.6 = ($40 million)/($25 million) on the low end and 2.4 = ($60 million/$25 million) on the high end. The range of possible values is
2.4 − 1.6 = 0.8. The fraction of possible values falling below 2.0, the level that triggers default, is the distance between 2.0 and 1.6, or 0.40; the
value 0.40 is one-half the total length of 0.8, or 0.4/0.8 = 0.50. So, the probability that the covenant will be breached is 50%.
38
Estimate the mean and standard deviation of EBITDA/interest. For a continuous uniform random variable, the mean is given by μ = (a + b)/2
and the variance is given by σ2 = (b − a)2/12.
Solution to 2: In Solution 1, we found that the lower limit of EBITDA/interest is 1.6. This lower limit is a. We found that the upper limit is 2.4.
This upper limit is b. Using the formula given previously, μ = (a + b)/2 = (1.6 + 2.4)/2 = 2.0.
The variance of the interest coverage ratio is σ2 = (b − a)2/12.= (2.4 − 1.6)2/12 = 0.053333.
The standard deviation is the positive square root of the variance, 0.230940 = (0.053333)1/2. However, the standard deviation is not
particularly useful as a risk measure for a uniform distribution. The probability that lies within various standard deviation bands around the
mean is sensitive to different specifications of the upper and lower. Here, a one standard deviation interval around the mean of 2.0 runs from
1.769 to 2.231 and captures 0.462/0.80 = 0.5775, or 57.8%, of the probability. A two standard deviation interval runs from 1.538 to 2.462,
which extends past both the lower and upper limits of the random variable.
3. Binomial Distribution
LOS (e) : Describe the properties of a Bernoulli random variable and a binomial random variable, and calculate and interpret
probabilities given the binomial distribution function
The binomial distribution is used in investment contexts to model outcomes as either successes or failures. It is particularly
useful when dealing with binary outcomes. The binomial distribution is built on the Bernoulli random variable, which represents
a trial with two outcomes. Such a trial is called a Bernoulli trial. If we have a Bernoulli random variable Y which equates to 1
when the outcome is success and to 0 when the outcome is failure, then the probability function of Y is :
where p is the probability that the trial is a success. In a series of n Bernoulli trials, the number of successes can range from 0 to
n. Since the outcome of each trial is random, the total number of successes in these n trials is also random. This random
variable, denoted as X, represents the count of successes in n Bernoulli trials and is known as a binomial random variable. The
probability distribution of this random variable X, called the binomial distribution, makes these assumptions:
The two parameters p, n completely describe a binomial distribution if the above two assumptions hold. If n=1, then binomial
becomes a Bernoulli random variable.
Now, the general expression for the probability that a binomial random variable shows x successes in n consecutive trials (also
known as the probability mass function) is given as :
What this essentially says is that out of n trials, x trials will be a success with a probability of p, and n-x trials are failures with
probability 1-p. Since the order/sequence in which results of the trials are studied is not relevant, we have used the combination
formula ( n! / (n-x)! x! ) to account for redundancies. (This concept was discussed in the previous module).
Example
A Trading Desk Evaluates Block Brokers
Block brokers connect with your trading desk to offer large blocks of stocks for sale. These transactions carry inherent risks, including potential
losses due to unfavourable information or incomplete disclosure by the seller. Your firm assesses broker performance by calculating post-
trade, market-risk-adjusted returns and categorizes trades as profitable or unprofitable. An excerpt from a spreadsheet summarizing broker
performance for November is provided, with brokers labelled BB001 and BB002.
You now want to evaluate the performance of the block brokers, and you begin with two questions:
39
1. If you are paying a fair price on average in your trades with a broker, what should be the probability of a profitable trade?
Solution to 1: If the price you trade at is fair, then 50% of the trades you do with a broker should be profitable.
Solution to 2: Your firm has logged 3 + 9 = 12 trades with block broker BB001. Since 3 of the 12 trades were profitable, the portion of
profitable trades was 3/12, or 25%. With broker BB002, the portion of profitable trades was 5/8, or 62.5%. The rate of profitable trades with
broker BB001 of 25% clearly missed your performance expectation of 50%. Broker BB002, at 62.5% profitable trades, exceeded your
expectation.
3. You also realize that the brokers’ performance has to be evaluated in light of the sample sizes, and for that you need to use the binomial
probability function. Under the assumption that the prices of trades were fair, A. calculate the probability of three or fewer profitable trades
with broker BB001. B. calculate the probability of five or more profitable trades with broker BB002.
Solution to 3A : For broker BB001, n = 12, and 3 trades were profitable. Probability of three or fewer profitable trades, implies F(3) = p(3) +
p(2) + p(1) + p(0). Assume underlying probability of a profitable trade with BB001 is p = 0.50. With n = 12 and p = 0.50, according to binomial
equation, the probability of three profitable trades is:
P(3) = { 12! / (12-3)! 3! } x px(1 – p) n−x where p = 0.5 which gives us 0.053711. Thus, the probability of exactly 3 profitable trades out of 12 is
5.4% if broker BB001 were giving you fair prices. Now you need to calculate the other probabilities: p(2), p(1) and so on.
Adding all the probabilities, F(3) = 0.053711 + 0.016113 + 0.00293 + 0.000244 = 0.072998, or 7.3%. The probability of making 3 or fewer
profitable trades out of 12 would be 7.3% if your trading desk were getting fair prices from broker BB001.
Solution to 3B. : For broker BB002, you are assessing the probability that the underlying probability of a profitable trade with this broker was
50%, despite the good results. The question was framed as the probability of making five or more profitable trades if the underlying probability
is 50%: 1 − F(4) = p(5) + p(6) + p(7) + p(8). You could calculate F(4) and subtract it from 1, but you can also calculate p(5) + p(6) + p(7) + p(8)
directly.
You begin by calculating the probability that exactly five out of eight trades would be profitable if BB002 were giving you fair prices:
p (5) = (85) (0.505) (0.503) = 56 (0.003906) = 0.21875. The probability is about 21.9%. The other probabilities can be found similarly.
So, p(5) + p(6) + p(7) + p(8) = 0.21875 + 0.109375 + 0.03125 + 0.003906 = 0.363281, or 36.3%. A 36.3% probability is substantial; the underlying
probability of executing a fair trade with BB002 might well have been 0.50 despite your success with BB002 in November of last year. If one of
the trades with BB002 had been reclassified from profitable to unprofitable, exactly half the trades would have been profitable. In summary,
your trading desk is getting at least fair prices from BB002; you will probably want to accumulate additional evidence before concluding that
you are trading at better-than-fair prices. The magnitude of the profits and losses in these trades is another important consideration. If all
profitable trades had small profits but all unprofitable trades had large losses, for example, you might lose
Two descriptors of a distribution that are often used in investments are the mean and the variance. For the Bernoulli and
binomial case, these are presented in the table below :
Mean Variance
Bernoulli, B(1, p) P p(1 − p)
Binomial, B(n, p) Np np(1 − p)
An important observation to note is that some distributions are always symmetric, such as the normal, and others are always
asymmetric or skewed, such as the lognormal. The binomial distribution is symmetric when the probability of success on a trial
is 0.50, but it is asymmetric or skewed otherwise.
Example
The Expected Number of Defaults in a Bond Portfolio
Suppose as a bond analyst you are asked to estimate the number of bond issues expected to default over the next year in an unmanaged high-
yield bond portfolio with 25 US issues from distinct issuers. The credit ratings of the bonds in the portfolio are tightly clustered around
Moody’s B2/Standard & Poor’s B, meaning that the bonds are speculative with respect to the capacity to pay interest and repay principal. The
40
estimated annual default rate for B2/B rated bonds is 10.7%.
1. Over the next year, what is the expected number of defaults in the portfolio, assuming a binomial model for defaults?
Solution to 1:For each bond, we can define a Bernoulli random variable equal to 1 if the bond defaults during the year and zero otherwise.
With 25 bonds, the expected number of defaults over the year is np = 25(0.107) = 2.675, or approximately 3.
2. Estimate the standard deviation of the number of defaults over the coming year.
Solution to 2: The variance is np(1 − p) = 25(0.107)(0.893) = 2.388775. The standard deviation is (2.388775)1/2 = 1.55. Thus, a two standard
deviation confidence interval (±3.10) about the expected number of defaults (≈ 3), for example, would run from approximately 0 to
approximately 6.
Solution to 3: An assumption of the binomial model is that the trials are independent. In this context, a trial relates to whether an individual
bond issue will default over the next year. Because the issuing companies probably share exposure to common economic factors, the trials
may not be independent. Nevertheless, for a quick estimate of the expected number of defaults, the binomial model may be adequate.
4. Normal Distribution
The descriptions provided so far are related to a single variable or univariate distribution, which characterizes the distribution
of one random variable. However, in investment contexts, you'll also come across multivariate distributions that describe the
probabilities for groups of related random variables. When dealing with a group of assets, we can model the distribution of
returns for each asset individually or for the entire group, accounting for their statistical interrelationships. The multivariate
normal distribution is a model often used for security returns, and it requires three lists of parameters:
The pdf for the curves shown in the normal distribution illustration above is given as :
When μ = 0 and σ = 1, it is called the standard normal distribution (or unit normal distribution). For the sake of efficiency, we
refer all probability statements to a single normal distribution. The standard normal distribution fits that role perfectly. Usually
we associate a normal distribution with the “bell curve,” which, in fact, is the probability density function of the normal
distribution, depicted in Panel A below. The cumulative distribution function, depicted in Panel B, in fact plots the size of the
shaded areas of the pdfs:
o Normal distribution is usually considered an approximate model for asset returns although it is not 100% accurate. This
becomes most evident when we approximate equity return distributions with the normal distribution, where normal
distribution tends to underestimate the probability of extreme returns.
o since option returns are skewed, we should be cautious in using the symmetrical normal distribution to model the returns
on portfolios containing significant positions in options.
o The normal distribution is more appropriate for modelling returns than asset prices. Asset prices can't go below zero,
making the normal distribution less suitable for them. Practitioners often use the lognormal distribution to model asset
prices instead.
The following probability statements, related to normal distribution are used frequently in practise :
42
Many a times, the random variable we observe has a non-zero mean and std. deviation different from 1. Standardizing a
normal variable helps simplify calculations and comparisons as it results in a new random variable which then has a mean = 0
and std. deviation of 1. Standardizing a normal random variable X involves two steps: first, subtract the mean of X from X, and
then divide the result by the standard deviation of X, which is known as computing the Z-score. When working with a list of
observations on a normal random variable X, you subtract the mean from each observation to obtain deviations from the mean
and then divide each deviation by the standard deviation. This results in the standard normal random variable, Z, where Z
represents a standard normal distribution.
The standardization formula is Z = (X - μ) / σ, where X ~ N(μ, σ2), with μ being the mean and σ being the standard deviation. For
instance, if X has μ = 5 and σ = 1.5, standardizing X with Z = (X - 5) / 1.5 yields Z values. For example, if X = 9.5, Z is calculated as
(9.5 - 5) / 1.5 = 3. The probability of observing a value as small as or smaller than 9.5 for X ~ N(5, 1.5) is the same as the
probability of observing a value as small as or smaller than 3 for Z ~ N(0, 1). Z-tables are available to find these probability
values.
To answer probability questions about a random variable X, we typically use standardized values discussed before. Important to
note that we often substitute the sample mean (X̄) for the population mean (μ) and the sample standard deviation (s) for the
population standard deviation (σ) since these population parameters are usually unknown.
LOS (k) : Define shortfall risk, calculate the safety-first ratio, and identify an optimal portfolio using Roy’s safety-first criterion
Modern portfolio theory (MPT) relies on measuring investment opportunities based on mean return and variance of return. In
economic theory, mean-variance analysis is applicable when investors are risk-averse, maximize expected utility, and when
returns are normally distributed or when investors have quadratic utility functions. Even if these assumptions are not perfectly
met, mean-variance analysis can still be useful. Practitioners often work with observable data like returns, and the assumption
of returns being approximately normally distributed is crucial in MPT.
Mean-variance analysis considers risk symmetrically, capturing both upside and downside variability with standard deviation.
However, an alternative approach, such as safety-first rules, focuses only on downside risk. Safety-first rules assess the risk of a
portfolio falling below a minimum acceptable level over a specified time frame, making them suitable for evaluating practical
investment problems. An example is the risk that assets in a defined benefit plan may fall below plan liabilities, illustrating the
concept of shortfall risk.
Roy's safety-first criterion is an approach where an investor considers any return below a specific level (R L) as unacceptable. The
optimal portfolio, according to this criterion, minimizes the probability of the portfolio's return (R P) falling below the threshold
level RL. In cases where portfolio returns are normally distributed, this probability (P(R P < RL)) can be calculated using the
number of standard deviations by which RL is below the expected portfolio return (E(RP)). The portfolio for which (E(RP) - RL) is
43
the largest relative to the standard deviation (σP) minimizes P(RP < RL). In this context, the safety-first optimal portfolio
maximizes the safety-first ratio (SFRatio), defined as:
The numerator represents the distance from the mean return to the shortfall level, which is then divided by σ P to measure it in
standard deviation units. When using Roy's criterion to choose among portfolios, assuming normality, we follow two steps:
For a portfolio with a specific SFRatio, the probability of its return falling below RL is given by the normal distribution N(–
SFRatio) { which can be solved using the property of N(-Z) = 1 – N(Z) }, and the safety-first optimal portfolio minimizes this
probability, aiming for the lowest shortfall risk.
LOS (l) : Explain the relationship between normal and lognormal distributions and why the lognormal distribution is used to
model asset prices
LOS (m) : Calculate and interpret a continuously compounded rate of return, given a specific holding period return
The lognormal distribution, similar to the normal distribution, is characterized by two parameters. However, it's unique in that
it's defined in terms of the parameters of a different distribution - its associated normal distribution. The two parameters for a
lognormal distribution are the mean and standard deviation (or variance) of the associated normal distribution, which is the
mean and variance of ln(Y), given that Y is lognormal. This means we need to track two sets of means and standard deviations
(or variances) - one for the associated normal distribution (the parameters) and one for the lognormal variable itself.
Without going into mathematical complexities, we can simply note the following observations :
If a stock's continuously compounded return is normally distributed, it implies that the future stock price will be lognormally
distributed. Even when the continuously compounded returns are not normally distributed, stock prices can still be effectively
described by the lognormal distribution. This provides the theoretical basis for using the lognormal distribution to model asset
prices.
The relationship is established by expressing the future stock price (S T) as : the current stock price (S0) multiplied by e raised to
the power of the continuously compounded return (r 0,T) from time 0 to T, represented as ST = S0 * exp(r0,T). When shorter-term
continuously compounded returns are assumed to be normally distributed, it follows that r 0,T is also normally distributed (under
44
specific assumptions) or approximately so (when not making those assumptions). Since S T is proportional to the log of a normal
random variable, it is lognormally distributed.
o In investment applications, a critical assumption is that returns are independently and identically distributed (i.i.d.).
Independence means past returns don't predict future returns, while identical distribution assumes stationarity, where the
mean and variance of returns remain constant from period to period.
o Volatility measures the standard deviation of the continuously compounded returns on the underlying asset. It is typically
an annualised measure. In some cases, daily volatility is mentioned. To annualise it, we need to multiply it by the square
root of the number of days in the target period. For example, annualised volatility = daily volatility X (250) 1/2 assuming
there are 250 days where the market is in a year.
LOS (n) : Describe the properties of the Student’s t-distribution, and calculate and interpret its degrees of freedom
LOS (o) :Describe the properties of the chi-square distribution and the F-distribution, and calculate and interpret their degrees
of freedom
The standard t-distribution is a probability distribution characterized by a single parameter called degrees of freedom (df),
representing the number of independent variables used in defining sample statistics. It can be used to model asset returns,
similar to the normal distribution, but with "longer tails," making it more conservative for estimating downside risk. Each specific
value for degrees of freedom defines a unique distribution within this family of distributions. To understand the concept of
degrees of freedom we examine the calculation of the sample variance, which is given as :
In estimating the population variance and using the t-distribution for reliability factors, the denominator "n - 1" (sample size
minus 1) represents the degrees of freedom. This term accounts for the fact that, when calculating the sample variance, the
sample mean is used, which limits the number of observations that can be independently chosen. For instance, with a sample
size of 10 and a mean of 10%, only 9 observations can be freely selected, as the 10th observation can be determined to maintain
the mean at 10%. Therefore, the degrees of freedom for the sample variance formula are "n - 1."
Probability calculations for the t-distribution are easily done with spreadsheets, statistical software, and programming
languages, like the R programming language.
45
The chi-square distribution is asymmetrical and a family of distributions. It represents the sum of squares of k independent
standard normally distributed variables and doesn't have negative values. There's a different chi-square distribution for each
possible value of degrees of freedom, which is n - 1, where n is the sample size.
Similar to the chi-square distribution, the F-distribution is also a family of asymmetrical distributions. It's bounded from below
by 0 and defined by two values of degrees of freedom, known as the numerator and denominator degrees of freedom.
The relationship between the chi-square and F-distributions is such that if you have one chi-square random variable χ12 with m
degrees of freedom and another chi-square random variable χ22 with n degrees of freedom, then F = (χ 12/m) / (χ22/n) follows an
F-distribution with m numerator and n denominator degrees of freedom.
Both chi-square and F-distributions are asymmetric, and as their degrees of freedom increase, their probability density functions
become more bell curve-like, resembling the normal distribution. For the F-distribution, as both the numerator (df1) and the
denominator (df2) degrees of freedom increase, the density function will also become more bell curve–like
Monte Carlo simulation is a technique used in finance that employs computer software to model complex financial systems. It
involves generating a large number of random samples from specified probability distributions to represent risk within the
system. This simulation is commonly used in investment applications to estimate risk and return, simulating a portfolio's profit
and loss performance over a defined time period. It is also employed to value complex securities without available pricing
formulas, such as mortgage-backed securities with embedded options, allowing for sensitivity analysis by controlling
assumptions within the simulation.
To better understand this procedure, we consider a contingent claim with a value equal to the difference between the
underlying stock's price at maturity and the average stock price during the claim's life or $0, whichever is greater. The process
involves simulating stock prices monthly over 12 months in 1,000 scenarios to evaluate the claim. The payoff diagram at
maturity (Panel A) shows that if the final stock price is less than or equal to the average, the payoff is zero; otherwise, it's the
difference between the final and average prices. Histograms of simulated final and average stock prices are displayed (Panel B),
with the claim's value depending on this difference. Panel C shows the histogram of contingent claim payoffs, where 65.4% of
trials resulted in zero pay-outs, and the maximum was $11.
Now, once we have the results of all the possible trials, we can calculate an expected value from all the outcomes, and then
apply present value technique to determine a value for this contingent claim.
46
Example
Valuing a Lookback Contingent Claim Using Monte Carlo Simulation
A standard lookback contingent claim on stock has a value at maturity equal to (Value of the stock at maturity − Minimum value of stock
during the life of the claim prior to maturity) or $0, whichever is greater. If the minimum value reached prior to maturity was $20.11 and the
value of the stock at maturity is $23, for example, the contingent claim is worth $23 − $20.11 = $2.89. Briefly discuss how you might use Monte
Carlo simulation in valuing a look- back contingent claim.
Solution: We previously described how to use Monte Carlo simulation to value a certain type of contingent claim. Just as we can calculate the
average value of the stock over a simulation trial to value that claim, for a lookback contingent claim, we can also calculate the minimum value
of the stock over a simulation trial. Then, for a given simulation trial, we can calculate the terminal value of the claim, given the minimum
value of the stock for the simulation trial. We can then discount this terminal value back to the present to get the value of the claim today (t =
0). The average of these t = 0 values over all simulation trials is the Monte Carlo simulated value of the lookback contingent claim.
47
Quantitative Methods - Learning Module 5
for the CFA exam
Sampling and Estimation
Learning Outcomes :
a) Compare and contrast probability samples with non-probability samples and discuss applications of each to an investment
problem
b) Explain sampling error
c) Compare and contrast simple random, stratified random, cluster, convenience, and judgmental sampling
d) Explain the central limit theorem and its importance
e) Calculate and interpret the standard error of the sample mean
f) Identify and describe desirable properties of an estimator
g) Contrast a point estimate and a confidence interval estimate of a population parameter
h) Calculate and interpret a confidence interval for a population mean, given a normal distribution with 1) a known population
variance, 2) an unknown population variance, or 3) an unknown population variance and a large sample size
i) Describe the use of resampling (bootstrap, jackknife) to estimate the sampling distribution of a statistic
j) Describe the issues regarding selection of the appropriate sample size, data snooping bias, sample selection bias,
survivorship bias, look-ahead bias, and time-period bias
1. Introduction
In investment analysis, we regularly use stock market indexes like the S&P 500 and Nikkei 225 as samples to understand how
various markets are performing. These indexes are considered valid indicators of the overall market behavior. However, any
statistics computed using sample data are just estimates of the underlying population parameters, as a sample is a subset of the
entire population.
This module covers sampling, where we obtain a sample, and how we use the mean as a measure of central tendency for
variables in investment analysis. The central limit theorem lets us make probability statements about the population mean even
when we don't know the variable's probability distribution. The module also discusses statistical estimation, which seeks precise
answers about parameter values. These concepts are crucial for investment analysis, helping us assess investment strategies. It
also addresses interpreting statistical results in financial data and potential pitfalls in the process.
2. Sampling Methods
LOS (a) : Compare and contrast probability samples with non-probability samples and discuss applications of each to an
investment problem
LOS (b) : Explain sampling error
LOS (c) : Compare and contrast simple random, stratified random, cluster, convenience, and judgmental sampling
This section discusses methods for gathering information about a population through samples, which are smaller subsets of the
population. It aims to estimate population parameters using sample statistics. Sampling is used either because it's impractical to
study every member of the population or for efficiency in time and cost.
Sampling methods fall into two categories: probability and non-probability. Probability sampling ensures an equal chance for
every population member to be chosen, resulting in a representative sample. Non-probability sampling relies on factors other
than chance, like the sampler's judgment or data accessibility, risking non-representative samples. Generally, probability
sampling tends to offer more accuracy and reliability. The section explores simple and stratified random sampling within
probability sampling before delving into non-probability sampling.
Suppose an analyst wants to estimate the average spending of major telecommunications equipment customers for the coming
year. He can either survey the entire population or take a representative sample of companies. Surveying all companies is costly
and time-consuming, while sampling offers a quicker and more cost-effective approach. However, sampling introduces some
error, as not all companies are surveyed, so the analyst is essentially trading off time and money for this sampling error.
48
To use sampling in data analysis, an analyst must create a sampling plan, which outlines the rules for selecting a sample. The
most reliable way to draw valid statistical conclusions about a population is to use a simple random sample (or random sample),
which means selecting a subset of the population in a manner that ensures every element in the population has an equal chance
of being included in the sample.
Simple random sampling is the process of selecting a sample which conforms to the above definition. For finite populations, this
is typically done using random numbers. The population members are numbered sequentially, and random numbers are
generated using a computer or table. These random numbers are then matched with population member codes to create a
sample of the desired size. Simple random sampling works well when the population is homogeneous, meaning the data
characteristics are broadly similar. If the data is not homogeneous, other sampling methods may be more appropriate.
Systematic sampling is employed when it's challenging to code or identify all population members. In this method, every kth
member is selected until the desired sample size is achieved. While this approach creates an approximately random sample,
practical sampling situations sometimes necessitate such an approximation.
Continuing with our earlier example, when an analyst surveys a random sample of telecom equipment customers to estimate
the population's average expenditure, the sample mean provides an estimate, but it won't match the true population mean. This
discrepancy is known as sampling error, which arises due to sampling variation. In essence, it's the difference between the
observed statistic and what it's meant to estimate, caused by using only a subset of the population for data collection.
A random sample provides an unbiased representation of the population, and sample statistics like the sample mean are valid
estimates of population parameters. Sample statistics are random variables, meaning they have their own distribution, known
as the sampling distribution. The sampling distribution of a statistic is the distribution of all the possible values that the statistic
can take when computed from samples of the same size randomly drawn from the same population. For example, the sampling
distribution of the sample mean represents the distribution of sample means when repeatedly taken from the same population.
This concept will be discussed in a lot more detail later in the reading.
Stratified random sampling is an alternative to simple random sampling, where the population is divided into subpopulations
(strata) based on specific criteria. Simple random samples are then taken from each stratum in proportion to their size in the
population, combining to form a stratified random sample.
Compared to simple random sampling, stratified random sampling ensures that important population subdivisions are
represented in the sample. It also offers more precise parameter estimates with smaller variance. This method is particularly
valuable in areas like bond indexing, where investors aim to replicate a specified bond index but the sheer number of
constituent securities makes it impossible to fully replicate. In this context, stratified sampling is used to group bonds based on
factors like duration, cash flow distribution, sector, credit quality, and call exposure. Then the investor chooses a sample from
each stratum proportional to the relative market weighting of the stratum in the index to be replicated.
Example
Bond Indexes and Stratified Sampling
Suppose that in selecting among the securities that qualify for selection within each strata, you apply a criterion concerning the liquidity of the
security’s market. Is the sample obtained random? Explain your answer.
Solution: Applying any additional criteria to the selection of securities for the strate, not every security that might be included has an equal
probability of being selected. As a result, the sampling is not random. In practice, indexing using stratified sampling usually does not strictly
involve random sampling because the selection of bond issues within cells is subject to various additional criteria. Because the purpose of
sampling in this application is not to make an inference about a population parameter but rather to index a portfolio, lack of randomness is not
in itself a problem in this application of stratified sampling.
Cluster sampling involves dividing the population into subpopulation groups called clusters. Some clusters are selected using
simple random sampling. If every member within the chosen clusters is sampled, it's one-stage cluster sampling; if a random
subsample is taken from each selected cluster, it's two-stage cluster sampling. Cluster sampling differs from stratified sampling
49
in that it treats entire clusters as sampling units, whereas stratified sampling includes all strata and selects specific elements
within each stratum as sampling units.
Cluster sampling is often used in market surveys, offering a time-efficient and cost-effective approach for large populations.
However, it tends to have lower accuracy compared to other probability sampling methods because a cluster sample may be
less representative of the entire population.
Non-probability sampling methods are not based on a fixed selection process but rather on the researcher's discretion. There
are two major types:
o Convenience Sampling: In this method, elements are chosen from the population based on accessibility or ease of access
for the researcher. These samples may not be representative of the entire population, leading to limited sampling accuracy.
However, it is advantageous for quickly and cost-effectively collecting data, especially in preliminary research stages or
when cost constraints exist.
o Judgmental Sampling: This method involves selectively picking elements based on the researcher's knowledge and
judgment. It can be influenced by the researcher's bias and may not provide a representative sample. However, in time-
constrained situations or when the researcher's expertise is crucial, judgmental sampling allows for targeted selection. For
instance, experienced auditors can use their judgment to choose accounts or transactions for financial statement audits,
ensuring adequate coverage.
50
3. The Central Limit Theorem and Distribution of the Sample Mean
LOS (d) : Explain the central limit theorem and its importance
LOS (e) : Calculate and interpret the standard error of the sample mean
To assess the sampling error in estimating the population mean, an analyst needs to understand the sampling distribution of the
sample mean, which is a probability distribution of random outcomes. This distribution is called the statistic's sampling
distribution. The central limit theorem is a helpful result that aids in understanding the sampling distribution of the mean for
various estimation problems, allowing the analyst to gauge how closely the sample mean is likely to match the underlying
population mean.
The Central Limit Theorem states that when you take random samples of a large enough size (n) from a population with any
probability distribution having a mean (μ) and finite variance (σ²), the sampling distribution of the sample mean (𝑋) will be
approximately normal. This normal distribution will have a mean equal to the population mean (μ) and a variance equal to the
population variance divided by the sample size (σ²/n).
The expression σ²/n represents the variance of the sample mean. As n, the sample size, increases, this fraction's value
decreases. This indicates that with a larger sample size, it becomes less likely to obtain a sample mean that is significantly
different from the population mean, making the estimates more reliable.
The central limit theorem permits us to make accurate probability statements about the population mean using the sample
mean, regardless of the population's distribution (as long as it has finite variance). Generally, when the sample size is 30 or
more, we can assume that the sample mean follows an approximate normal distribution. However, in cases of highly non-
normal populations, a larger sample size may be necessary for this assumption to hold true.
The central limit theorem states that the variance of the sample mean is σ²/n, where σ is the standard deviation of the
population. The standard deviation of a sample statistic is referred to as the standard error of the statistic. In the case of the
sample mean, it is a crucial quantity when applying the central limit theorem. The standard error of the sample mean is
calculated using one of two expressions:
o If we know the population standard deviation (σ), the standard error of the sample mean is σ/√n.
o If we don't know the population standard deviation and have to estimate it using the sample standard deviation (s), the
standard error of the sample mean is s/√n where s is estimated using by the square root of the sample variance given by :
Standard deviation and standard error are not interchangeable terms. While standard deviation measures the spread of data
from the mean, standard error quantifies the inaccuracy of a population parameter estimate resulting from sampling. These
concepts reflect the difference between describing data and making inferences about a population. Standard deviation is used
when assessing data spread, whereas standard error is employed when evaluating the precision of a population parameter
estimate from sampled data relative to its true value. Confidence intervals, which use the sample mean and its standard error,
will be discussed as a technique to make probability statements about the population mean.
Statistical inference involves two main branches: hypothesis testing and estimation. Hypothesis testing assesses whether a
parameter (e.g., population mean) is equal to a specific value, while estimation focuses on determining the actual value of the
parameter. In estimation, we use sample information to calculate point estimates (single values) or confidence intervals (ranges
of values with a specified level of probability) for the population parameter. This section primarily covers point estimates of
parameters and introduces the formulation of confidence intervals for the population mean. Hypothesis testing will be discussed
separately.
51
4.1. Point Estimators
Sample statistics, such as the sample mean, are considered random variables because they depend on random outcomes. The
formulas used to compute these sample statistics are referred to as estimation formulas or estimators. When we calculate a
specific value from sample data using an estimator, that value is called an estimate. An estimator has a sampling distribution,
while an estimate is a fixed number related to a particular sample and does not have a sampling distribution. For example, the
calculated value of the sample mean in a specific sample, which serves as an estimate of the population mean, is known as a
point estimate of the population mean. Different samples drawn from the population will yield varying results when using the
sample mean formula, demonstrating the variability of point estimates.
In choosing an estimator for a parameter, we often consider desirable statistical properties. Three important properties are
unbiasedness, efficiency, and consistency.
o Unbiasedness: An unbiased estimator is one whose expected value (the mean of its sampling distribution) equals the
parameter it aims to estimate. For example, the sample mean (𝑋) is an unbiased estimator of the population mean (μ)
because its expected value equals μ. Similarly, the sample variance (s²) with a divisor of (n - 1) is an unbiased estimator of
the population variance (σ²). If we had used a divisor of (n) for the sample variance, it would be biased because its expected
value would be smaller than the population variance.
o Efficiency is a criterion used to select the best estimator among alternatives when multiple unbiased estimators for a
parameter are available. An unbiased estimator is considered efficient if no other unbiased estimator for the same
parameter has a sampling distribution with smaller variance.
A statistic's sampling distribution is defined for a specific sample size, and different sample sizes result in different sampling
distributions. Unbiasedness and efficiency are properties of an estimator's sampling distribution that remain consistent
regardless of the sample size. For example, an unbiased estimator remains unbiased whether in a sample of size 100 or 1,000.
However, in some cases, finding estimators with desirable properties like unbiasedness in small samples can be challenging. In
such situations, statisticians may choose estimators based on their asymptotic properties in extremely large samples, with
consistency being a crucial one.
52
grows, the estimator tends to produce more accurate estimates of the population parameter. In contrast, an
inconsistent estimator does not become more accurate with increasing sample size.
Note : In a Big Data context, consistency becomes more important than efficiency because the accuracy of population parameter
estimates can be significantly improved with larger sample data. Even if an estimator is biased, it can provide reduced error with
a bigger dataset. For instance, the estimator s2/n for variance is biased, but as the sample size (n) approaches infinity, the
difference between s²/n and the unbiased estimator s²/(n - 1) becomes negligible. In large datasets, consistency is a key factor in
achieving accurate parameter estimates.
Example
The image below plots several sampling distributions of an estimator for the population mean, and the vertical dash line
Solution: C is correct. The chart shows three sampling distributions of the estimator at different sample sizes (n = 50, 200, and
1,000). We can observe that the means of each sampling distribution—that is, the expected value of the estimator—deviates
from the population mean, so the estimator is biased. As the sample size increases, however, the mean of the sampling
distribution draws closer to the population mean with smaller variance. So, it is a consistent estimator.
5. Confidence Intervals for the Population Mean and Sample Size Selection
LOS (g) : Contrast a point estimate and a confidence interval estimate of a population parameter
LOS (h) : Calculate and interpret a confidence interval for a population mean, given a normal distribution with 1) a known
population variance, 2) an unknown population variance, or 3) an unknown population variance and a large sample size
When estimating a population parameter, we often use a point estimate, which is a single number. However, due to sampling
error, the point estimate is unlikely to match the population parameter exactly in any given sample. An alternative and more
practical approach is to calculate a range of values, called a confidence interval, which we expect to contain the parameter with
a specified level of probability.
Confidence Interval is defined as a range where one can assert, with a given degree of confidence (1 - α), that it will include the
parameter being estimated. This interval is commonly referred to as the 100(1 - α)% confidence interval for the parameter.
Confidence intervals can be interpreted both probabilistically and practically. In the probabilistic interpretation, a 95%
confidence interval for the population mean means that in repeated sampling, 95% of such intervals will contain the population
mean. In the practical interpretation, we assert that we are 95% confident that a single 95% confidence interval contains the
53
population mean. This assertion is justified because we know that 95% of all possible confidence intervals constructed in the
same way will include the population mean.
The confidence intervals that we discuss have structures similar to the following basic structure:
➔ A 100(1 − α)% confidence interval for a parameter is given as : Point estimate ± Reliability factor × Standard error
Where, Reliability factor = a number based on the assumed distribution of the point estimate and the degree of confidence (1 –
α) for the confidence interval
Standard error = the standard error of the sample statistic providing the point estimate
The quantity “Reliability factor × Standard error” is sometimes called the precision of the estimator; larger values of the product
imply lower precision in estimating the population parameter.
The most basic confidence interval for the population mean is when sampling from a normal distribution with a known variance.
It relies on the standard normal distribution, with a mean of 0 and a variance of 1, denoted as Z. The notation z α represents the
point in the standard normal distribution such that α of the probability is in the right tail. For instance, z 0.05 equals 1.65, meaning
that 5% of possible values are greater, and z0.025 equals 1.96, indicating 2.5% of values are larger.
From this above discussion we can say that a 100(1 − α)% confidence interval for population mean μ when we are sampling from
a normal distribution with known variance σ2 is given by :
The reliability factors for the most frequently used confidence intervals are as follows :
The above information implies that as we increase the degree of confidence, the confidence interval becomes wider and gives
us less precise information about the quantity we want to estimate.
In practical scenarios, assuming that the sampling distribution of the sample mean is at least approximately normal is often
reasonable due to the underlying distribution being roughly normal or having a large sample size where the central limit
theorem applies. However, knowing the population variance is rare in practice. When the population variance is unknown but
the sample mean is approximately normally distributed, two acceptable methods can be used to calculate the confidence
interval for the population mean.
o The first approach, relies on Student's t-distribution. When constructing confidence intervals for the population mean with
an unknown population variance, we often opt for this conservative approach using the t-distribution. This is especially
crucial for small sample sizes. Even when dealing with large samples, it's appropriate to use the t-distribution when the
population variance is unknown, as it results in more conservative (wider) confidence intervals.
In the case of sampling from a normal distribution, the ratio z = (𝑋 − μ) / (σ/√n) follows a standard normal distribution. On
the other hand, the ratio t = (𝑋 − μ) / (s/√n) follows a t-distribution with a mean of 0 and n - 1 degrees of freedom. The key
difference is that t is not normally distributed because it's the ratio of two random variables: the sample mean and the
sample standard deviation, whereas the standard normal distribution involves only the sample mean.
The form of a 100(1 − α)% confidence interval for the population mean μ using the t-distribution is given as :
where the number of degrees of freedom for tα/2 is n − 1 and n is the sample size.
o The second approach, known as the z-alternative, is based on the standard normal distribution and is applicable only when
the sample size is large, typically considered as 30 or more. This approach employs the sample standard deviation (s) to
calculate the standard error of the sample mean, and the formula for a 100(1 - α)% confidence interval for the population
mean (μ) is:
54
➔
The flowchart illustrated below helps identify the scenarios where the t-distribution and z-alternative should be used :
o Choice of Statistic (t or z): The choice of whether to use the t-distribution or the standard normal distribution (z) impacts the
width of the confidence interval.
o Degree of Confidence: The chosen level of confidence (1 - α) determines the specific value of the reliability factor (t or z)
used in the interval calculation.
o Sample Size: The sample size has a significant impact on the width of the confidence interval. Larger sample sizes result in
narrower intervals.
As we have seen, the standard error of the sample mean is inversely proportional to the square root of the sample size .
Therefore, as the sample size increases, the standard error decreases, leading to a narrower confidence interval. Greater sample
size provides more precise estimates of the population parameter.
To determine the required sample size for a desired confidence interval width (E) at a given level of confidence (1 - α), you can
use the formula:
➔ n = [(t × s) / E]².
This calculation helps ensure that the confidence interval has the desired width. Appropriate sample size is also crucial for
conducting power analysis and determining the minimum detectable effect in hypothesis testing, which will be covered in later
discussions.
6. Resampling
LOS (i) : Describe the use of resampling (bootstrap, jackknife) to estimate the sampling distribution of a statistic
Resampling methods, such as Bootstrap, provide a way to perform statistical inference and estimate population parameters
without relying on analytical formulas like z-statistics or t-statistics. Bootstrap, in particular, is a widely-used resampling
technique that uses computer simulation to generate sampling distributions.
The key idea behind Bootstrap is to simulate the process of random sampling from a population, similar to traditional sampling
methods. However, in Bootstrap, we don't have information about the actual population; instead, we use the observed data
55
sample, which is assumed to represent the population. This means that Bootstrap treats the available sample as if it were the
entire population, allowing us to perform resampling and statistical inference. The following illustration provides the difference.
Bootstrap allows us to perform statistical analysis without relying on specific assumptions about the population's distribution,
making it a powerful tool for various applications in statistics and data analysis.
Example
Bootstrap Resampling Illustration
The following table displays a set of 12 monthly returns of a rarely traded stock, shown in Column A. Our aim is to calculate the standard error
of the sample mean. Using the bootstrap resampling method, a series of bootstrap samples, labelled as “resamples” (with replacement) are
drawn from the sample of 12 returns. Notice how some of the returns from data sample in Column A feature more than once in some of the
resamples (for example, 0.055 features twice in Resample 1).
56
0.296 −0.191 −0.132 0.255 ..... −0.132
0.055 −0.132 −0.132 0.296 ..... 0.055
−0.072 −0.096 0.055 −0.096 ..... −0.096
0.255 0.055 −0.072 0.055 ..... −0.191
−0.157 −0.157 −0.053 −0.157 ..... 0.055
Sample mean −0.019 −0.060 0.001 ..... 0.040
Drawing 1,000 such samples, we obtain 1,000 sample means. The mean across all resample means is −0.01367. Using the above equation, the
standard error of the sample mean can now be calculated.
Jackknife is a resampling technique used in statistical inference. It differs from bootstrap, which involves repeatedly drawing
samples with replacement. In the jackknife method, samples are created by removing one observation at a time from the
original data without replacement. Jackknife is commonly employed to reduce estimator bias and can also be used to determine
the standard error and confidence interval of an estimator. Jackknife consistently produces similar results with each run, while
bootstrap often yields different results due to the randomness of its resampling. Jackknife typically requires n repetitions for a
sample of size n, whereas bootstrap's appropriate number of repetitions must be determined.
LOS (j) : Describe the issues regarding selection of the appropriate sample size, data snooping bias, sample selection bias,
survivorship bias, look-ahead bias, and time-period bias
There are various challenges to ensure valid sampling, including data snooping bias, sample selection biases (such as
survivorship bias), look-ahead bias, and time-period bias. These issues are critical when conducting point and interval estimation
and hypothesis testing. Biased samples can lead to errors in point and interval estimates and any conclusions drawn from the
sample data.
Data snooping involves overusing the same or related data to discover statistically significant patterns, which can lead to data
snooping bias. Investment strategies influenced by data snooping biases often fail in the future. Analysts and researchers must
be cautious of this issue. Data snooping typically involves extensive testing with a small significance level (e.g., 5%), leading to
the risk of finding false patterns due to chance. Using separate datasets for training, validation, and out-of-sample testing can
help identify data snooping bias by evaluating whether a variable or strategy remains significant beyond the data it was derived
from. Caution is necessary as strategies may not work in the future if widely known. Analysts should be sceptical of profitable
strategies rooted in data snooping and their future applicability.
Researchers studying questions relevant to analysts or portfolio managers may omit specific assets or time periods from their
analysis due to data availability issues. This exclusion of assets or data is referred to as sample selection bias.
Survivorship bias, a type of sample selection bias, occurs when databases only contain data from currently active companies or
funds, omitting those that are no longer operational. This bias affects studies using such databases as they overlook the
performance of failed entities. A sample can also be biased because of the removal (or delisting) of a company’s stock from an
exchange.
Sample selection bias is a concern even in markets with high-quality data. Hedge funds, a diverse group of unregulated
investment vehicles have discretion over disclosing performance data, leading to self-selection bias in hedge fund databases, as
those with poor track records may avoid public disclosure. Moreover, hedge fund databases can suffer from survivorship bias
when they exclude funds that cease to operate.
Implicit selection bias can arise due to thresholds that allow self-selection, affecting the quality of the data used in analysis. For
instance, choosing stocks listed on the NYSE, which has higher listing requirements compared to smaller exchanges, can
introduce an implicit quality bias in the analysis. Although this bias may be less apparent, it's essential to consider when making
generalizations. Backfill bias is a variation of selection bias where the past performance of a new hedge fund is added to an
index's database when it wasn't included in the previous year. This can inflate the index's performance, especially when a new
fund begins contributing data after a period of strong performance.
57
7.3. Look-Ahead Bias
Look-ahead bias in test design occurs when the test uses information that wasn't available on the test date, potentially leading
to inaccurate results. To address this bias, one approach is to use point-in-time (PIT) data, which is recorded or released with a
specific date, rather than using data from a different time period. It's important to avoid implicitly introducing look-ahead bias
when normalizing data, such as using the standard deviation of training data as the reference for validation and test data, rather
than the standard deviation of those specific datasets.
Time-period bias in a test design arises when the chosen period leads to results that are specific to that time frame. A short time
series is likely to provide results that are specific to that short period and may not be representative of a longer time frame. In
contrast, a longer time series may offer a more accurate reflection of investment performance but could be affected by
structural changes over time. These changes could result in two different return distributions: one before the change and
another after the change, such as shifts in volatility or interest rates. Such regime changes can significantly impact asset classes,
and any inferences drawn from data influenced by one regime should consider how the regime might bias those inferences.
58
Quantitative Methods - Learning Module 6
for the CFA exam
Hypothesis Testing
Learning Outcomes:
a) Define a hypothesis, describe the steps of hypothesis testing, and describe and interpret the choice of the null and
alternative hypotheses
b) Compare and contrast one-tailed and two-tailed tests of hypotheses
c) Explain a test statistic, Type I and Type II errors, a significance level, how significance levels are used in hypothesis testing,
and the power of a test
d) Explain a decision rule and the relation between confidence intervals and hypothesis tests, and determine whether a
statistically significant result is also economically meaningful
e) Explain and interpret the p-value as it relates to hypothesis testing
f) Describe how to interpret the significance of a test in the context of multiple tests
g) Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the population mean of both
large and small samples when the population is normally or approximately normally distributed and the variance is (1)
known or (2) unknown
h) Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the equality of the
population means of two at least approximately normally distributed populations based on independent random samples
with equal assumed variances
i) Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the mean difference of two
normally distributed populations
j) Identify the appropriate test statistic and interpret the results for a hypothesis test concerning (1) the variance of a
normally distributed population and (2) the equality of the variances of two normally distributed populations based on two
independent random samples
k) Compare and contrast parametric and nonparametric tests, and describe situations where each is the more appropriate
type of test
l) Explain parametric and nonparametric tests of the hypothesis that the population correlation coefficient equals zero, and
determine whether the hypothesis is rejected at a given level of significance
m) Explain tests of independence based on contingency table data
1. Introduction
LOS (a) : Define a hypothesis, describe the steps of hypothesis testing, and describe and interpret the choice of the null and
alternative hypotheses
In the face of overwhelming data, analysts need to organize and analyze it to gain a clearer understanding. While visualizations
are helpful, they may not provide precise answers to questions about data for example if investment returns differ from an
average of x% or from each other, or if their standard deviations are different from y%, or if their variability varies. Hypothesis
testing is essential in these cases. It is a statistical inference technique that allows us to make decisions about a larger
population based on observations from a smaller sample.
As an example, consider a set of 1,000 asset returns with a mean of 6% and a standard deviation of 2%, the chance of a sample
mean being exactly 6% depends on factors like sample size and sampling methodology. Drawing samples of say 30 observations
may yield sample means slightly different from 6%, and this variation is influenced by the population's variability. Hypothesis
testing is a technique used to determine if a sample statistic is likely to come from a population with a specified parameter value.
It provides an objective way to assess whether available evidence supports a hypothesis. The focus of this reading is on the
framework of hypothesis testing and its application to mean, variance, and correlation tests, which are common in investments,
allowing for a clearer understanding of the likelihood that a hypothesis is true. However, it is important to keep in mind that
these conclusions are not absolute certainties.
59
2. The Process of Hypothesis Testing
LOS (b) : Compare and contrast one-tailed and two-tailed tests of hypotheses
Hypothesis testing is a part of statistical inference, which involves estimation and hypothesis testing. Estimation includes point
and interval estimates, while hypothesis testing focuses on how sample statistics inform us about population parameters.
Hypothesis testing starts with formulating a theory, testing it using observations and predictions, and assessing correctness. The
following image illustrates the standard approach to hypothesis testing :
In hypothesis testing, two hypotheses are stated: the null hypothesis (H0) and the alternative hypothesis (Ha). The null
hypothesis assumes an initial population parameter value until the sample data convincingly indicates otherwise. The null
hypothesis is what we aim to reject. If evidence shows the null hypothesis is (likely) untrue, we reject it in favor of the
alternative hypothesis. Both hypotheses refer to population parameters, and sample statistics are used to test them.
In hypothesis testing, we create two types of tests: two-sided and one-sided. In a two-sided test, we assess if the population
mean is different from a specific value (e.g., μ = 6). The hypothesis are stated as :
If the sample mean significantly deviates from this value, we reject the null hypothesis.
In a one-sided test, we check if the mean is greater (or less) than a specified value (e.g., μ > 6). The hypothesis are stated as :
➔ H0 : μ ≤ 6 Ha : μ > 6.
If the sample mean is significantly higher, we reject the null hypothesis, which is stated with "≤" to ensure mutual exclusivity and
exhaustiveness of the null and alternative hypotheses.
Despite the different ways to formulate hypotheses, we always conduct a test of the null hypothesis at the point of equality; for
example, μ = μ0. Whether the null is H0: μ = μ0, H0: μ ≤ μ0, or H0: μ ≥ μ0, we actually test μ = μ0. Taking the population mean
example, for hypothesis testing we have three possible formulations:
Choosing null and alternative hypotheses involves deciding what you want to test. The null is what you aim to reject. The
common alternative is the "not equal to" hypothesis, but economic or financial theory may suggest a one-sided alternative. The
sign in the alternative reflects the researcher's belief, and a two-sided alternative may be chosen to maintain neutrality.
Typically, it's easiest to specify the alternative hypothesis first and then the null.
LOS (c) : Explain a test statistic, Type I and Type II errors, a significance level, how significance levels are used in hypothesis
testing, and the power of a test
A test statistic is a value calculated on the basis of a sample that, when used in conjunction with a decision rule, is the basis for
deciding whether to reject the null hypothesis.
The focal point of our statistical decision is the value of the test statistic, which depends on what we're testing. For example, In
the case of testing a population mean risk premium, the test statistic depends on the population standard deviation (σ) and the
sample size (n). When σ is known, the standard error of the sample means (σ X) is calculated as σ/√n. The test statistic for this
case is a z-distributed test statistic, expressed as z = (Xrp - μ0) / (σ/√n), where Xrp is the sample mean and μ0 is the hypothesized
mean risk premium. If the hypothesized mean risk premium is 6 (μ 0 = 6), the test statistic is calculated as z = (Xrp - 60) / (σ/√n). If
the hypothesized mean risk premium is zero (μ0 = 0), the test statistic simplifies to z = Xrp / (σ/√n).
Important to remember that identifying the appropriate test statistic for the hypotheses and the underlying distribution of the
population is the key to hypothesis testing.
Following the identification of the appropriate test statistic, we must be concerned with the distribution of the test statistic.
Some examples of the test statistics and their corresponding distributions are shown below :
61
Where, μ0 , μd0 , and σ02 denote hypothesized values of the mean, mean difference, and variance, respectively. The 𝑥, 𝑑, s2, s,
and r denote for a sample the mean, mean of the differences, variance, standard deviation, and correlation, respectively, with
subscripts indicating the sample, if appropriate. The sample size is indicated as n, and the subscript indicates the sample, if
appropriate. Oij and Eij are observed and expected frequencies, respectively, with r indicating the number of rows and c
indicating the number of columns in the contingency table..
The level of significance determines how much evidence from a sample is needed to reject the null hypothesis. This threshold
can vary based on the hypotheses and the potential consequences of an error. When testing a null hypothesis, there are four
possible outcomes which are shown below.
True Situation
Decision H0 True H0 False
Type II error: Fail to reject a false null hypothesis.
Fail to reject H0 Correct decision: Do not reject a true null hypothesis.
False negative
Type I error: Reject a true null hypothesis.
Reject H0 Correct decision: Reject a false null hypothesis.
False positive
A Type I error is a false positive, rejecting null when it is true and a Type II error is a false negative, failing to reject null when it is
false. These errors are mutually exclusive
In testing a hypothesis about a population mean return (for e.g. our earlier case of µ = 6%), we determine how different a
sample mean can be from the hypothesized value before we consider that it is significantly different. We do this by setting a risk
tolerance for Type I errors and identifying critical values that indicate significant deviation from the population mean. These
critical values depend on the chosen alternative hypothesis (one-sided or two-sided), the probability distribution of the test
statistic, sample size, and the level of risk tolerance for Type I errors.
The probability of making a Type I error is denoted as α, representing the level of significance, and its complement (1 - α) is the
confidence level. For instance, a 5% significance level corresponds to a 95% confidence level, where there's a 5% chance of
rejecting a true null hypothesis.
Balancing Type I and Type II error probabilities involves a trade-off because reducing Type I error risk (e.g., using a 1%
significance level instead of 5%) increases the risk of Type II errors because we because we will reject the null less frequently,
including when it is false. Deciding which error to accept more depends on the consequences and costs of each error. Increasing
the sample size is the way to reduce both types of errors simultaneously. Quantifying this trade-off is challenging because the
probability of a Type II error is difficult to estimate due to various potential false hypotheses. Therefore, we mainly specify α, the
Type I error probability, when conducting hypothesis tests.
The significance level of a test is the probability of mistakenly rejecting a true null hypothesis (Type I error), while the power of a
test is the probability of correctly rejecting the null when it is false (complement of Type II error). The probability of a Type II
error is typically denoted as beta (β). These probabilities can be classified as shown :
In hypothesis testing, it's better to choose a test statistic with maximum power and then set a significance level beforehand. This
prevents our judgment from being influenced by the test statistic's result if we calculate it first.
LOS (d) : Explain a decision rule and the relation between confidence intervals and hypothesis tests, and determine whether a
statistically significant result is also economically meaningful
62
In hypothesis testing, the fourth step is to establish a decision rule that defines when to reject the null hypothesis and when not
to. This rule is based on comparing the calculated test statistic to a critical value, the selection of which depends on the level of
significance and the probability distribution of the test statistic. If the test statistic is more extreme than the critical value, the
null hypothesis is rejected, indicating statistical significance. If not, the null hypothesis is not rejected, indicating insufficient
evidence to do so.
If we refer to the case A in the above image, we can say that a 95% confidence interval for the population mean, μ, based on
sample mean, 𝑋 is given by 𝑋 ± 1.96 σ / √n . Now if z = ( 𝑋 – µ0 ) / (σ / √n), which becomes our test statistic, then the
conditions for rejecting the null hypothesis is z < -1.96 or z > 1.96 which implies that the sample value is far enough from the
hypothesized value for us to reject it with a 95% confidence.
The fifth step in hypothesis testing involves collecting data and calculating the test statistic. The quality of our conclusions relies
on both the suitability of the statistical model and the data's quality. This requires ensuring that the sampling procedure is
unbiased and free from errors like sample selection or time bias. Additionally, data must be cleansed to identify and correct
inaccuracies and measurement errors. Once we have unbiased and accurate data, we use it to calculate the relevant test
statistic.
Example
Using a Confidence Interval in Hypothesis Testing
Consider the hypotheses H0: μ = 3 versus Ha: μ ≠ 3. If the confidence interval based on sample information has a lower bound of 2.75 and an
upper bound of 4.25, the most appropriate decision is:
Solution C is correct. Since the hypothesized population mean (μ = 3) is within the bounds of the confidence interval (2.75, 4.25), the correct
63
decision is to fail to reject the null hypothesis. It is only when the hypothesized value is outside these bounds that the null hypothesis is
rejected. Note that the null hypothesis is never accepted; either the null is rejected on the basis of the evidence or there is a failure to reject
the null hypothesis.
6. Make a Decision
LOS (d) : Explain a decision rule and the relation between confidence intervals and hypothesis tests, and determine whether a
statistically significant result is also economically meaningful
The sixth step in hypothesis testing is making the decision. For instance, in a test comparing the population mean risk premium
with zero, if the calculated z-statistic is 2.5 and there's a two-sided alternative hypothesis with a 5% significance level, we reject
the null hypothesis because 2.5 falls outside the bounds of ±1.96. This is a statistical decision, indicating that the evidence
suggests the mean risk premium is not equal to zero.
In addition to the statistical decision, making an economic or investment decision is a crucial part of hypothesis testing. This
decision incorporates both statistical results and relevant economic factors. For instance, if we reject the null hypothesis of a
zero risk premium in favor of the alternative that it's greater than zero, we find evidence that the US risk premium is not zero.
The key question is whether this difference is economically significant. Based on these factors, an investor may choose to invest
in US equities. Various non-statistical factors, like the investor's risk tolerance and financial situation, also play a role in the
decision-making process.
Sometimes, small differences between a variable and its hypothesized value are statistically significant but lack economic
significance. For example, we might reject a null hypothesis of a zero mean return for an investment strategy based on a large
sample. In such cases, the smaller the standard error of the mean and the larger the test statistic, the more likely we are to
reject the null. This standard error decreases with larger sample sizes, making it possible to reject the null even for small
deviations. However, a statistically significant positive mean return may not be economically significant when considering
transaction costs, taxes, and risk. It's essential to assess the logic of why a strategy might work in the future before
implementation, but these considerations go beyond what can be addressed in a hypothesis test.
LOS (e) : Explain and interpret the p-value as it relates to hypothesis testing
The p-value, often reported in hypothesis tests, represents the area in the probability distribution outside the calculated test
statistic. For a two-sided test, it's the area outside ± the test statistic, while for a one-sided test, it's the area outside the test
statistic on the relevant side of the distribution. The p-value is the smallest level of significance at which the null hypothesis can
be rejected. For instance, with a calculated z-statistic of 2.33
in a two-sided test, the p-value is the area outside ±2.33,
typically calculated by software. If the p-value is smaller than
the chosen significance level, we reject the null hypothesis,
indicating stronger evidence for the alternative hypothesis.
64
8. Multiple Tests and Significance Interpretation
LOS (f) : Describe how to interpret the significance of a test in the context of multiple tests
A Type I error, also known as a false positive, occurs when we mistakenly reject a true null hypothesis. The false discovery rate
(FDR) represents the expected proportion of false positives. In the context of multiple hypothesis testing, like when drawing
1,000 samples of 50 observations each, some samples lead to the rejection of a true null hypothesis. With a significance level of
0.05, you would reject the null hypothesis about 5% of the time, even if it is true. This situation is called the multiple testing
problem.
The false discovery approach to testing involves adjusting the p-values when conducting a series of tests. Benjamini and
Hochberg introduced this approach in 1995. The researcher ranks the p-values from lowest to highest and makes the following
comparison starting with the lowest p-value p(1):
The highest ranked p-value that still satisfies this comparison is declared significant, and all the tests are considered significant.
Suppose a scenario where we repeatedly test the hypothesis that the population mean is 6% by drawing 20 samples and
calculating 20 test statistics, the (BH) criteria can help control the false discovery rate. If we used a 5% significance level and
relied solely on each test's p-value, we might end up rejecting the null hypothesis in five tests. However, when applying the BH
criteria, only one test is considered significant (shown in the table below). This approach helps mitigate the risk of making
excessive false positive (Type I) errors when conducting multiple hypothesis tests.
The conclusions regarding p-values and the multiple testing problem are as follows:
o A non-significant result from a single test is not necessarily incorrect, and the null hypothesis may indeed be true.
o Caution is needed when the test's statistical power is low or the sample size is small, as there is a higher risk of
encountering false positives (Type I errors).
o Critical values used in hypothesis tests are based on the assumption of a single test. Repeatedly testing the same data
increases the likelihood of spurious results due to data snooping. It's advisable to perform the test once and avoid repeated
testing in search of statistically significant results to prevent chance findings (false positives).
o In very large samples, it's common to obtain statistically significant results for many tests. To address this, one can draw
different samples and assess the robustness of the results. Consistency across different samples suggests more reliable
findings.
Example
False Discovery and Multiple Tests
A researcher is examining the mean return on assets of publicly traded companies that constitute an index of 2,000 mid-cap stocks and is
testing hypotheses concerning whether the mean is equal to 15%: H0: μ ROA = 15 versus Ha: μROA ≠ 15. She uses a 10% level of significance and
collects a sample of 50 firms. She wants to examine the robustness of her analysis, so she repeats the collection and test of the return on
assets 30 times. The results for the samples with the five lowest p-values are given below :
65
1. Of the 30 samples tested, how many should the researcher expect, on average, to have p-values less than the level of significance?
Solution to 1 : Of the 30 samples tested, she should expect 30 × 0.10 = 3 to have significant results just by chance. Consider why she ended up
with more than three. Three is based on large sample sizes and large numbers of samples. Using a limited sample size (i.e., 50) and number of
samples (i.e., 30), there is a risk of a false discovery with repeated samples and tests.
2. What are the corrected p-values based on her selected level of significance, and what is the effect on her decision?
Solution to 2 : Applying the BH criteria, the researcher determines the adjusted p-values shown below :
Calculated z- p-Value Rank of p-Value (lowest to ( α x Rank of i ) / ( Number of Is value in (3) less than or equal to
Statistic highest) tests) value in (5)?
3.203 0.00136 1 0.00333 Yes
3.115 0.00184 2 0.00667 Yes
2.987 0.00282 3 0.01000 Yes
2.143 0.03211 4 0.01333 No
1.903 0.05704 5 0.01667 No
On the basis of the above results, there are three samples with p-values less than their adjusted p-values. So, the number of significant sample
results is the same as would be expected from chance, given the 10% level of significance. The researcher concludes that the results for the
samples with Ranks 4 and 5 are false discoveries, and she has not uncovered any evidence from her testing that supports rejecting the null
hypothesis.
LOS (g) : Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the population mean of
both large and small samples when the population is normally or approximately normally distributed and the variance is (1)
known or (2) unknown
Hypothesis tests involving the mean are common in practice. When the population standard deviation is unknown, the sampling
distribution of the mean follows a t-distribution, and when it's known, it follows a z-distribution. Since the population standard
deviation is typically unknown, we often use a t-distributed test statistic. The t-distribution has a parameter called degrees of
freedom (df) and is similar to the standard normal distribution but has a greater standard deviation ( > 1) and fatter tails. As the
degrees of freedom increase with the sample size, the t-distribution approaches the standard normal distribution.
The t-statistic is the theoretically correct test statistic for hypothesis tests concerning the population mean of a normally
distributed population with an unknown variance. The t-statistic is also robust for the population’s moderate deviations from
normality, except for outliers and strong skewness. For larger sample sizes, concerns about these deviations from normality are
less significant. While a common rule of thumb suggests using the normal distribution for sample sizes over 30, it's better to rely
on the t-distribution and statistical software for precise testing.
If the population sampled has unknown variance, then the test statistic for hypothesis tests concerning a single population
mean, μ, is :
This test statistic is t-distributed with n − 1 degrees of freedom, which we can write as tn-1. If this is outside the bounds of the
critical values based on the level of significance, we will reject the null hypothesis in favor of the alternative.
When dealing with a single population mean and an unknown population variance, we typically use a t-test, especially for small
samples with at least approximate normality. For large samples, the central limit theorem ensures that the sample mean is
approximately normally distributed, making the t-statistic still suitable. However, in large samples, some practitioners opt for a
z-test due to the following reasons:
66
o The sample mean tends to follow a normal distribution in large samples, fulfilling the normality assumption of the z-test.
o The difference between the rejection points for t-tests and z-tests becomes negligible in large samples.
Even though the z-test is theoretically correct when we know the population variance, the t-test is often presented as the
preferred choice because it's readily available in statistical software and applicable when the population variance is unknown,
which is the case in most scenarios. Nevertheless, whenever we use the z test, the calculation for z statistic is similar to that
discussed for t, just in the denominator, the sample standard deviation (s) is replaced with the population standard deviation (σ)
for the z-case. Rest all remains the same.
LOS (h) : Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the equality of the
population means of two at least approximately normally distributed populations based on independent random samples with
equal assumed variances
In hypothesis testing, we often want to determine if there's a significant difference in the mean values of two groups, such as
mean returns. To test this, we draw samples from each group, assuming that the samples are roughly normally distributed and
independent. When conducting tests of differences in means, we may assume that the population variances are either equal or
unequal. This discussion primarily focuses on the case where population variances are assumed to be equal. In the test statistic
calculation, we combine observations from both samples to estimate the common population variance.
Typically, we aim to test whether the population means are equal or if one is larger than the other. Thus, we formulate the
following sets of hypotheses:
Two sided: H0:μ1 −μ2 =0 versus Ha:μ1 −μ2 ≠ 0, or, equivalently, H :μ1 =μ2 versus Ha : μ1 ≠ μ2
One sided (right side): H0:μ1 −μ2 ≤0 versus Ha : μ1 −μ2 >0, or, equivalently, H0 : μ1 ≤ μ2 versus Ha:μ1 > μ2
One sided (left side): H0:μ1 −μ2 ≥ 0 versus Ha: μ1 −μ2 < 0, or, equivalently, H0:μ1 ≥μ2 versus Ha:μ1 <μ2
We can, however, formulate other hypotheses, where the difference is something other than zero, such as H0:μ1 −μ2 =2 versus
Ha:μ1 −μ2 ≠ 2. The procedure is the same.
Assuming that the two populations are normally distributed and that the unknown population variances are equal, we use a t-
distributed test statistic based on independent random samples:
LOS (i) : Identify the appropriate test statistic and interpret the results for a hypothesis test concerning the mean difference of
two normally distributed populations
When comparing two independent samples, we use a t-distributed test statistic that considers the difference in means and a
pooled variance. However, if we're working with samples we believe are dependent, we use the test of the mean of the
differences.
This section focuses on the t-test for paired observations, which is sometimes called the paired comparisons test. Paired
observations are dependent because they share something in common. For example, we might examine dividend policy changes
in companies before and after a tax law change, resulting in pairs of observations for the same companies. In such cases, the
test is used to assess the mean of the observed differences.
67
The key difference between the test of paired differences and the test of differences in means for independent samples is that
the paired comparisons test is more powerful. It leverages the common element (e.g., same periods or companies) to eliminate
variations between samples caused by factors other than what is being tested. This increases its statistical power.
In cases where we have dependent samples, we pair observations from random variables, say XA and XB, denoting the difference
between paired observations as di = xAi − xBi, where xAi and xBi represent the ith pair of observations. Here, n is the number of
pairs.
We aim to test hypotheses regarding the population mean difference, denoted as μ d. These hypotheses can be formulated in
three ways:
o Two-sided: Null Hypothesis (H0): μd = μd0 and Alternative Hypothesis (Ha): μd ≠ μd0
o One-sided (right side): Null Hypothesis (H0): μd ≤ μd0 and Alternative Hypothesis (Ha): μd > μd0
o One-sided (left side): Null Hypothesis (H0): μd ≥ μd0 and Alternative Hypothesis (Ha): μd < μd0
These hypotheses are used to test whether there is a significant difference in the population mean difference, μ d, considering a
hypothesized value μd0 which is usually 0.
We use a t-distributed test statistic because we are concerned with the case of normally distributed populations with unknown
population variances. We start with the mean of differences :
The sample standard deviation, sd, is the standard deviation of the differences, so the standard error of the mean differences =
sd /√n
When we have data consisting of paired observations from samples generated by normally distributed populations with
unknown variances, the t-distributed test statistic, with n-1 degrees of freedom (n being the number of observations) is :
LOS (j) : Identify the appropriate test statistic and interpret the results for a hypothesis test concerning (1) the variance of a
normally distributed population and (2) the equality of the variances of two normally distributed populations based on two
independent random samples
Often, we are interested in the volatility of returns or prices, and one approach to examining volatility is to evaluate variances.
We examine two types of tests involving variance: tests concerning the value of a single population variance and tests
concerning the difference between two population variances.
When the goal is to assess whether the variance of a fund's returns matches a specified target, we conduct a test of a population
variance. This test necessitates specifying a hypothesized value for the variance, denoted as σ0². We can formulate hypotheses
regarding whether the variance is equal to a specific value or if it is greater or less than the hypothesized value. There are three
possible formulations for these hypotheses:
o Two-sided alternative: Null Hypothesis (H0): σ² = σ0² and Alternative Hypothesis (Ha): σ² ≠ σ0²
o One-sided alternative (right tail): Null Hypothesis (H0): σ² ≤ σ0² and Alternative Hypothesis (Ha): σ² > σ0²
o One-sided alternative (left tail): Null Hypothesis (H0): σ² ≥ σ0² and Alternative Hypothesis (Ha): σ² < σ0²
For tests concerning the variance of a single normally distributed population, we employ a chi-square test statistic, represented
as χ². The chi-square distribution is not symmetrical, unlike the normal and t-distributions. It consists of a family of distributions,
68
each associated with a specific degree of freedom (n - 1, where n is the sample size). Importantly, the chi-square distribution is
bounded from zero and does not produce negative values.
For n independent observations from a normally distributed population, the appropriate test statistic is :
Where s2 is the sample variance and σ0² is the hypothesized variance. The chi-square test is more sensitive to violations of its
assumptions compared to tests like the t-test. If the sample is not drawn randomly or does not originate from a normally
distributed population, using a chi-square test may result in unreliable inferences. Additionally, because the chi-square
distribution is both asymmetric and bounded from zero, we cannot employ the straightforward ± notation for critical values as
seen in the z- and t-distributions. Instead, we must rely on chi-square tables or statistical software to calculate and obtain the
critical values for chi-square tests. A sample is shown for 25 observations (df = 24) and at 5% significance level. The shaded
region represents the null hypothesis rejection regions for two-sided tests :
In cases where we want to compare the volatility of two samples, we can conduct tests to determine the equality of their
variances. For instance, this might involve comparing the volatility of different baskets of securities to indexes or benchmarks or
comparing volatility in various time periods. When comparing variances of two normally distributed populations with variances
σ₁² and σ₂² (distinguishing populations 1 and 2), we can formulate hypotheses with two-sided or one-sided alternatives:
o Two-sided alternative: Null Hypothesis (H0): σ₁² = σ₂² and Alternative Hypothesis (Ha): σ₁² ≠ σ₂²
o One-sided alternative (right side): Null Hypothesis (H0): σ₁² ≤ σ₂² and Alternative Hypothesis (Ha): σ₁² > σ₂²
o One-sided alternative (left side): Null Hypothesis (H0): σ₁² ≥ σ₂² and Alternative Hypothesis (Ha): σ₁² < σ₂²
These tests, based on independent random samples from the populations, use an F-test, which calculates the ratio of sample
variances. F-tests for differences in population variances rely on the F-distribution which like the chi-square distribution, is a
family of asymmetric distributions bounded from zero and is characterized by numerator and denominator degrees of freedom.
However, it's important to note that the F-test is not robust to violations of its assumptions.
When we have two samples with the first sample containing n₁ observations and a sample variance s₁², and the second sample
containing n₂ observations with a sample variance s₂², and these samples are random and independent, generated by normally
distributed populations, we can conduct a test to compare the differences between the variances of these two populations. This
test is based on the ratio of sample variances :
➔ F = s12/s22
with df1 = (n1 − 1) numerator degrees of freedom and df2 = (n2 − 1) denominator degrees of freedom. Once we have the F-stat,
we can compare it with the critical values and make inferences.
69
13. Parametric vs. Nonparametric Tests
LOS (k) : Compare and contrast parametric and nonparametric tests, and describe situations where each is the more appropriate
type of test
The hypothesis-testing procedures discussed so far have two common characteristics: they deal with parameters, such as mean
and variance, and rely on specific assumptions about the population distribution. These are known as parametric tests.
However, there are cases where we are interested in quantities other than population parameters or when the assumptions of
parametric tests are not met. In such situations, nonparametric tests, which are not concerned with parameters and make
minimal population distribution assumptions, can be useful. The following table gives examples of nonparametric alternatives to
the parametric, t-distributed tests concerning means:
Parametric Nonparametric
Tests concerning a single mean t-distributed test z-distributed test Wilcoxon signed-rank test
Tests concerning differences between means t-distributed test Mann–Whitney U test (Wilcoxon rank
sum test)
Tests concerning mean differences (paired comparisons t-distributed test Wilcoxon signed-rank test Sign test
tests)
o When the data do not meet distributional assumptions: Nonparametric tests are employed when the data suggest that the
distributional assumptions of parametric tests are not met. This situation is common when dealing with small samples or
data from non-normally distributed populations. Nonparametric tests often involve converting observations into ranks and
working with "greater than" or "less than" relationships.
o When there are outliers: Even when the underlying population distribution is normal, the presence of extreme values or
outliers can affect parametric statistics. In such cases, nonparametric tests, particularly those involving the median, are
preferred.
o When the data are ranked or use an ordinal scale: Nonparametric tests are suitable when the data consist of rankings or
ordinal scales, as parametric tests typically require a stronger measurement scale. For example, when comparing
investment manager rankings, nonparametric procedures are preferred.
o When the hypotheses do not concern a parameter: Nonparametric procedures are used when the research question does
not involve a parameter, such as testing whether a sample is random (using a runs test) or assessing whether a sample
follows a specific probability distribution.
Nonparametric statistical procedures are valuable because they make minimal assumptions, accommodate ranked data, and can
address questions that do not involve parameters. They are often reported alongside parametric tests to evaluate the sensitivity
of statistical conclusions to the underlying assumptions of the parametric tests. However, when the assumptions of parametric
tests are met, parametric tests are generally preferred due to their higher statistical power, allowing for better detection of false
null hypotheses.
LOS (l) : Explain parametric and nonparametric tests of the hypothesis that the population correlation coefficient equals zero,
and determine whether the hypothesis is rejected at a given level of significance
A significance test of a correlation coefficient helps assess the strength of the linear relationship between two variables and
whether it is due to chance. If the correlation is zero, there is no linear relationship between the variables. Significance tests
evaluate if the estimated correlation is significantly different from zero, indicating a meaningful relationship. The correlation
coefficient ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 representing no linear
relationship.
Common hypotheses regarding correlation compare the population correlation coefficient to zero. Hypotheses can be two-sided
(no relationship) or one-sided (positive or negative relationship). The sample correlation is used to test these hypotheses on the
population correlation. Let ρ represent the population correlation coefficient. The possible hypotheses are as follows:
70
o Two sided: H0: ρ = 0 versus Ha: ρ ≠ 0
The parametric pairwise correlation coefficient, often called the Pearson correlation, is used for testing correlation. It measures
the relationship between two variables, X and Y, and is calculated as r = s XY / (sX * sY), where sXY is the sample covariance, sX is
the standard deviation of X, and s Y is the standard deviation of Y. If the two variables are normally distributed, we can test to
determine whether the null hypothesis (H0: ρ = 0) should be rejected using the sample correlation, r. The formula for the t-test,
with n-2 degrees of freedom is:
The Spearman rank correlation coefficient, denoted as rS, is used when the population under consideration significantly departs
from normality. It is similar to the usual correlation coefficient but is calculated based on the ranks of two variables, X and Y,
within their respective samples. The calculation of rS involves the following steps:
o Rank the observations for X and Y separately, from largest to smallest. Assign ranks such that the largest value receives a
rank of 1, the second-largest a rank of 2, and so on. In the case of ties (equal values), assign an average rank to the tied
observations. For instance, if the third and fourth largest values are tied, assign them both a rank of 3.5 (the average of 3
and 4).
o Calculate the difference, di, between the ranks of each pair of observations on X and Y. Then, compute d i2, which represents
the squared difference in ranks.
o With n as the sample size, the Spearman rank correlation is given by
This process is used to calculate the Spearman rank correlation coefficient, which assesses the relationship between two
variables based on their ranked data.
The test of hypothesis for the Spearman rank correlation depends on whether the sample is small or large (n > 30). For small
samples, the researcher requires a specialized table of critical values, but for large samples, we can conduct a t-test using the
test statistic , which is t-distributed with n − 2 degrees of freedom.
When dealing with categorical or discrete data, the methods discussed so far for hypothesis testing may not be applicable
to test the independence of different classifications within the data. For instance, consider a frequency table that
categorizes 1,594 exchange-traded funds (ETFs) based on two classifications: size (market capitalization) and investment
type (value, growth, or blend) :
Since the classification of investment type in this example is discrete, it is not suitable to employ correlation to evaluate
the relationship between size and investment type. Such tables are often referred to as contingency tables or two-way
71
tables because they involve two classifications or classes (size and investment type). To test for the independence of these
classifications, different statistical methods are employed.
To test whether there is a relationship between the size and investment type, we can perform a test of independence
using a nonparametric test statistic that is chi-square distributed :
Where, m = the number of cells in the table, which is the number of groups in the first class multiplied by the number of groups
in the second class
Oij = number of observations in each cell of row i and column j (i.e., observed frequency)
Eij = expected number of observations in each cell of row i and column j, assuming independence ([Link]
frequency)
This test statistic has (r − 1)(c − 1) degrees of freedom, where r is the number of rows and c is the number of columns.
For our case, m = 3x3 = 9. we need to estimate Eij, the expected frequency, which is the number of ETFs we would expect to be
in each cell if size and investment type are completely independent. The expected number of ETFs (E ij) is calculated using :
For example, for small cap value cell, Eij = 503 x 148 / 1594 = 46.703. We do this for the rest of the entries. Next we calculate (O ij
– Eij)2 / Eij for each cell and sum them up to get the chi-squared statistic.
In the ETF example, we conduct a hypothesis test to assess whether the two classifications (size and investment type) are
independent or dependent. The null hypothesis states that there is no relationship between these two classes, while the
alternative hypothesis suggests there is a relationship. The test statistic is calculated based on the differences between observed
and expected values. If the observed and expected values are identical, the test statistic is zero. In cases where differences exist,
these differences are squared, resulting in a positive chi-square statistic. Therefore, in the test of independence using a
contingency table, there is only one rejection region on the right side.
72
Quantitative Methods - Learning Module 7
for the CFA exam
Introduction to Linear Regression
Learning Outcomes :
a) Describe a simple linear regression model and the roles of the dependent and independent variables in the model
b) Describe the least squares criterion, how it is used to estimate regression coefficients, and their interpretation
c) Explain the assumptions underlying the simple linear regression model, and describe how residuals and residual plots
indicate if these assumptions may have been violated
d) Calculate and interpret the coefficient of determination and the F-statistic in a simple linear regression
e) Describe the use of analysis of variance (ANOVA) in regression analysis, interpret ANOVA results, and calculate and interpret
the standard error of estimate in a simple linear regression
f) Formulate a null and an alternative hypothesis about a population value of a regression coefficient, and determine whether
the null hypothesis is rejected at a given level of significance
g) Calculate and interpret the predicted value for the dependent variable, and a prediction interval for it, given an estimated
linear regression model and a value for the independent variable
h) Describe different functional forms of simple linear regressions
LOS (a) : Describe a simple linear regression model and the roles of the dependent and independent variables in the model
Financial analysts use regression analysis which is a statistical tool that helps assess the relationship and predictive power of
one variable in relation to another.
Let Y represent the variable for a company that we would like to explain, like return on assets, profitability etc. Let Yi represent
an observation/value of that variable and let 𝑌 represent the mean for the sample of size n. Now, we can either study the
variable values individually, in isolation, to make inferences or we can try to understand why the value differs from the mean
value. The variation in Y in our case can be written as :
2
➔ Variation of Y = ∑𝑛𝑖=1(𝑌𝑖 − 𝑌)
The variation of Y is often referred to as the sum of squares total (SST), or the total sum of squares.
Now, let’s suppose that our experience tells us that another variable X is a driving force behind the variable Y (for example
capital expenditures helps explain the return on assets values). We can write the variation in X in a similar manner :
2
➔ Variation of X = ∑𝑛𝑖=1(𝑋𝑖 − 𝑋)
Let’s take X to be capital expenditure levels and Y to be ROA (return on assets). The tabular data is presented below and to
understand this relation better, we have used a scatter plot.
73
In this example we are trying to use capital expenditures (CAPEX) to explain ROA. The variable being explained, which is ROA in
this case, is called the dependent variable (Y), while the variable used to explain it, which is CAPEX, is the independent variable
(X). To relate dependent and independent variables, a common approach is to establish a linear relationship, which represents
the connection between the two variables as a straight line. When there's just one independent variable, it's called simple linear
regression (SLR), but if there are multiple independent variables, it's termed multiple regression. SLR involves a single
independent variable, while multiple regression involves several.
LOS (b) : Describe the least squares criterion, how it is used to estimate regression coefficients, and their interpretation
Linear regression assumes a linear relationship between dependent and independent variables, aiming to fit a line that
minimizes squared deviations from the data points. This approach is known as least squares regression or ordinary least squares
(OLS) regression due to its widespread use. The linear relation between the dependent and independent variables is described
as :
Fitting the line requires minimizing the sum of the squared residuals, the sum of squares error (SSE), also known as the residual
sum of squares :
̂𝑖 )2 = ∑𝑛𝑖=1(𝑌𝑖 − (̂
➔ Sum of squares error = ∑𝑛𝑖=1(𝑌𝑖 − 𝑌 𝑏0 + 𝑏̂1 X𝑖 ))2 = ∑𝑛𝑖=1 𝑒𝑖 2
The slope of this regression line can be calculated using the following equation :
74
Which can be simplified to :
Since the regression line passes through the mean of the observations, we can use the slope value to find the intercept value as :
➔ ̂
𝑏0 = 𝑌 - 𝑏̂1 𝑋𝑖
We may find the slope formula similar to correlation formula discussed in earlier modules which was :
The difference is the denominator which for slope in the variance of the dependent variable but is the product of standard
deviations for the correlation calculations.
Regression coefficients convey the meaning of the intercept and slope in a regression model. The intercept represents the value
of the dependent variable when the independent variable is zero, which may be irrelevant or impractical in certain contexts. For
instance, in a model explaining GDP growth with money supply, the intercept lacks meaning because zero money supply isn't
feasible. The slope indicates the change in the dependent variable for a one-unit change in the independent variable. A positive
slope means both variables change in the same direction, while a negative slope implies they change in opposite directions.
With the estimated regression coefficients, we can predict the dependent variable if it follows the average relationship between
the dependent and the independent variable. Mathematics behind least squares regression ensures that residual term's
expected value is zero: E(ε) = 0
Regression analysis involves two main types of data: cross-sectional and time series. Cross-sectional regression uses multiple
observations of X and Y for the same time period, which can come from various sources like companies, asset classes, or
countries. Time series data, on the other hand, involves multiple observations from different time periods for the same entity,
such as a company or country. While time series data is often denoted as t = 1, 2, ..., T, this text primarily uses the notation i = 1,
2, ..., n for both cross-sectional and time series data in the following sections.
LOS (c) : Explain the assumptions underlying the simple linear regression model, and describe how residuals and residual plots
indicate if these assumptions may have been violated
In a simple linear regression model, there are four key assumptions that need to be met for valid conclusions:
o Linearity: The relationship between the dependent variable (Y) and the independent variable (X) is linear.
o Homoskedasticity: The variance of the regression residuals is consistent for all observations.
o Independence: Observations (pairs of Ys and Xs) are independent, meaning the regression residuals are uncorrelated across
observations.
o Normality: The regression residuals follow a normal distribution.
In linear regression, it's crucial to assume a linear relationship between dependent and independent variables. If the relationship
is nonlinear in the model parameters, using a simple linear regression will yield biased results and can lead to underestimation
and/or overestimation of the dependent variable at different points. Additionally, the independent variable, X, must not be
random, as randomness would break the linear relationship assumption.
75
When examining model residuals, they should ideally be random with no discernible pattern concerning the independent
variable.
Assumption 2 in linear regression, known as homoskedasticity, requires that the variance of residuals is consistent across all
observations. If the residuals have varying variances among observations, it's termed heteroskedasticity. For instance, in a time
series of short-term interest rates (Y) and inflation rates (X) spanning 16 years, if central bank actions lead to artificially low
interest rates for the last eight years, the residuals may appear to come from two different models: Regime 1 (normal rates) and
Regime 2 (low rates). This results in different model fits and varying variances in the residuals. This heteroskedasticity is evident
in the spread of residuals for the two regimes ( as shown in the following image), indicating a violation of the homoskedasticity
assumption. When we estimate separate regression lines for each regime, we can observe significant differences in the model
parameters, with distinct slopes for each regime, highlighting the presence of distinct regimes in the relationship between short-
term interest rates and inflation rates.
In linear regression, we assume that the observations (Y and X pairs) are independent and uncorrelated with each other. If there
is correlation between observations, known as autocorrelation, it means they are not independent, and this can lead to
correlated residuals. The independence of residuals across observations is crucial for accurate estimation of the variances of the
estimated parameters (b0 and b1 or ̂ 𝑏0 and 𝑏̂1 ) used in hypothesis tests for the intercept and slope. To ensure this assumption
holds, it's essential to visually and statistically assess the residuals of a regression model for any patterns that might indicate a
violation of this assumption.
The assumption of normality in linear regression pertains to the normal distribution of residuals, not the variables themselves.
It's good practice to check the distribution of both dependent and independent variables for outliers as outliers can significantly
impact the model's fit. Normally distributed residuals allow hypothesis testing for a linear regression model. In large samples,
the assumption of normality can sometimes be relaxed, relying on the central limit theorem and asymptotic theory, which
suggest that test statistics remain valid even when residuals are not normally distributed.
4. Analysis of Variance
LOS (d) : Calculate and interpret the coefficient of determination and the F-statistic in a simple linear regression
LOS (e) : Describe the use of analysis of variance (ANOVA) in regression analysis, interpret ANOVA results, and calculate and
interpret the standard error of estimate in a simple linear regression
Our primary objective is to understand the variation in the dependent variable based on our chosen independent variable. For
this, we need to assess how well we have achieved this goal with our selected independent variable.
76
4.1. Breaking down the Sum of Squares Total into Its Components
The sum of squares total (SST) measure discussed before can be broken into two components :
o Sum of squares regression (SSR) which represents the sum of the squared differences between the predicted value of the
dependent variable and the mean of the dependent variable. This can be interpreted as the explained variation in Y.
o Sum of squares error (SSE) which represents the sum of the squared differences between the actual value and the
predicted value. This is the unexplained variation in Y.
Using the above we can write Total variation in Y = Explained variation + Unexplained Variation which can also be written as :
To evaluate how well the regression fits our data we can use measures like coefficient of determination, the F-statistic for the
test of fit, and the standard error of the regression.
The coefficient of determination, also referred to as the R-squared or R2 , is the percentage of the variation of the dependent
variable that is explained by the independent variable. It ranges from 0-100% and is given by :
➔ = R2
The coefficient of determination describes the relationship between dependent and independent variables but is not a statistical
test. To determine the statistical significance of a regression model, we use an F-distributed test statistic. This statistic helps us
compare variances and assess whether the slopes in a regression (bi) are equal to zero or not.
The null hypothesis is H0: b1 = b2 = b3 = . . . = bk= 0. and the alternate hypothesis is Ha: At least one bk is not equal to zero. For
our simple linear regression case, these simplify to H0: b1 = 0. Ha : b1 0.
The F-statistic is a ratio of two variances : sum of squares regression and the sum of squares error, each adjusted for degrees of
freedom (which is equal to the number of independent variables and is represented by k). The formula for the F-stat is given as :
For simple linear regression, k = 1. The MSR and MSE and short for mean square regression and mean square error respectively
and for k=1 case are given as :
➔
77
➔
The F-statistic in regression analysis is one-sided, with the rejection region on the right side, as it helps determine if the
explained variation in Y is greater than the unexplained variation, which is of interest.
Analysis Of Variance (ANOVA) table is basically a summary of all the linear regression parameters, calculations and is usually
produced as an output of linear regression in many software packages. A sample table is given below :
The standard error of the estimate (se) in regression analysis, also known as the standard error of the regression or root mean
square error, measures the discrepancy between observed and predicted values of the dependent variable. A smaller s e
indicates a better model fit. It, along with the coefficient of determination and the F-statistic, assesses the goodness of fit of the
regression model. Unlike the other two, se is an absolute measure of the difference between observed and predicted values and
is vital for model evaluation, prediction intervals, and coefficient tests. Calculating s e is simple once you have the ANOVA table
as it's the square root of the Mean Square Error (MSE) i.e.
Example
Using ANOVA Table Results to Evaluate a Simple Linear Regression
Suppose you run a cross-sectional regression for 100 companies, where the dependent variable is the annual return on stock and the
independent variable is the lagged percentage of institutional ownership (INST). The results of this simple linear regression estimation are
shown
Solution to 1 : The coefficient of determination is sum of squares regression/sum of squares total: 576.148 ÷ 2,449.71 = 0.2352, or 23.52%.
2. What is the standard error of the estimate for this regression model?
Solution to 2 The standard error of the estimate is the square root of the mean square error: (19.1180)1/2 = 4.3724.
3. At a 5% level of significance, do we reject the null hypothesis of the slope coefficient equal to zero if the critical F-value is 3.938?
Solution to 3 : Using a six-step process for testing hypotheses, we get the following:
78
Step 3 Specify the level of significance. α = 5% (one tail, right side).
Step 4 State the decision rule. Critical F-value = 3.938. Reject the null hypothesis if the calculated F-statistic is greater
than 3.938.
Step 5 Calculate the test statistic. F = 576.1485/ 19.1180 = 30.1364
Step 6 Make a decision. Reject the null hypothesis because the calculated F-statistic is greater than the critical F-
value. There is sufficient evidence to indicate that the slope coefficient is different from
0.0.
4. Based on your answers to the preceding questions, evaluate this simple linear regression model.
Solution to 4 The coefficient of determination indicates that variation in the independent variable explains 23.52% of the variation in the
dependent variable. Also, the F-statistic test confirms that the model’s slope coefficient is different from 0 at the 5% level of significance. In
sum, the model seems to fit the data reasonably well.
LOS (f) : Formulate a null and an alternative hypothesis about a population value of a regression coefficient, and determine
whether the null hypothesis is rejected at a given level of significance
The F-statistic is used to test the significance of the slope coefficient in regression analysis. However, there are cases where we
want to perform other hypothesis tests for the slope coefficient, such as determining if it differs from a specific value or if it is
positive. This allows us to test hypotheses about regression slopes, like assessing the average systematic risk of a stock or testing
whether economists' inflation rate forecasts are unbiased. To do this, we employ a t-distributed test statistic, subtracting the
hypothesized population slope (B1) from the estimated slope coefficient (𝑏̂ 1) and dividing this difference by the standard error
of the slope coefficient 𝑠𝑏̂1 :
This test statistic is t-distributed with n − k − 1 or n − 2 degrees of freedom because two parameters (an intercept and a slope)
were estimated in the regression. The standard error of the slope coefficient 𝑠𝑏̂1 for a simple linear regression is the ratio of the
model’s standard error of the estimate (se) to the square root of the variation of the independent variable:
In hypothesis testing for regression, we use the t-statistic to compare with critical values. A lower standard error of the slope
(related to the variability of the independent variable) results in a higher t-statistic. If the calculated t-statistic falls outside the
critical t-values, we reject the null hypothesis. If it falls within the critical values, we fail to reject the null hypothesis. These tests
can have a two-sided or one-sided alternative hypothesis.
A feature of simple linear regression is that the t-statistic used to test if the slope coefficient is zero is the same as the t-statistic
used to test if the pairwise correlation is zero (H 0: ρ = 0 vs. Ha: ρ ≠ 0). We can use both two-sided and one-sided alternative
hypotheses for testing correlation, similar to how you can for testing a slope (e.g., H 0: ρ ≤ 0 vs. Ha: ρ > 0). The test-statistic to test
whether the correlation is equal to zero is :
Another feature of simple linear regression is that there's an interesting relationship between the test-statistic for assessing
model fit (F-statistic) and the t-statistic used to test if the slope coefficient is zero: t2 = F.
In some cases, we need to test if the population intercept has a specific value. For example, in simple linear regression, if we
have a company's revenue growth rate (Y) as the dependent variable and the GDP growth rate of its home country (X) as the
independent variable, the intercept represents the company's revenue growth rate when the GDP growth rate is 0%. The
equation for the standard error of the intercept, 𝑠𝑏̂0 : , is:
79
➔
Suppose we want to study the influence of a company's quarterly earnings announcements on its monthly stock returns. To do
this we can use a simple linear regression model with an indicator variable (dummy variable) as the independent variable. This
indicator variable, say EARN in our case, takes the values 0 or 1, representing the absence or presence of an earnings
announcement in a given month. The regression equation is as follows :
This allows us to test whether there are distinct returns for months with earnings announcements (EARN=1) compared to those
without earnings announcements (EARN=0). The data and the regression results are illustrated below :
In the context of analyzing monthly returns with an indicator variable for earnings announcements, there are noticeable
differences in returns between announcement and non-announcement months. We conduct hypothesis testing in a similar way
as with continuous independent variables. In this regression, the intercept represents the predicted value of the dependent
variable when the indicator variable is zero. Additionally, when the indicator variable is one, the slope signifies the difference in
means if observations were grouped by the indicator variable.
The choice of significance level in hypothesis testing is subjective. A common choice is 0.05, signifying a 5% chance of rejecting a
null hypothesis when it is actually true, thus making a Type I error (false positive). Lowering the significance level to 0.01 reduces
the Type I error but increases the Type II error (false negative) which is failing to reject the null hypothesis when, in fact, it is
false. The p-value represents the smallest significance level at which the null hypothesis can be rejected. A smaller p-value
indicates a lower chance of making a Type I error and suggests stronger support for the validity of the regression model. For
instance, a p-value of 0.005 corresponds to a 0.5% significance level (99.5% confidence). If for example, we calculate a t-statistic
of 4.00131 which leads to a p-value of 0.008. If we are testing it at a 5% significance level, since this is less than that 5%, it easily
allows us to reject the null hypothesis.
How do we determine the p-values? Since this is the area in the distribution outside the calculated test statistic, we need to
resort to software tools
80
6. Prediction Using Simple Linear Regression and Prediction Intervals
LOS (g): Calculate and interpret the predicted value for the dependent variable, and a prediction interval for it, given an
estimated linear regression model and a value for the independent variable
As analysts, whenever we are predicting the dependent variable value using regression, we need to consider that the estimated
regression line does not describe the relation between the dependent and independent variables perfectly; it is an average of
the relation between the two variables. This is evident because the residuals are not all zero. Thus, an interval estimate of the
forecast is needed to reflect this uncertainty.
o A better-fitting regression model results in a smaller se and, consequently, a smaller standard error of the forecast.
o A larger sample size (n) in the regression estimation leads to a smaller standard error of the forecast.
o The standard error of the forecast is smaller when the forecasted independent variable (X f) is closer to the mean of
the independent variable (𝑋) used in the regression estimation.
Using this standard error of forecast value, the interval around the predicted value of the dependent variable is given as:
For example, if predicted value of Y is 12.375, and the standard error of the forecast is 3.736912. Assuming a 5%
significance level (α), two sided, with n − 2 degrees of freedom (Say n = 6 in our case so, df = 4), the critical values for the
prediction interval are ±2.776. Then the 95% prediction interval then becomes 12.375 ± 2.776 (3.736912) leading us to
the following inference : {2.0013 < 𝑌̂𝑓 < 22.7487}
In regression analysis, not all sets of independent and dependent variables exhibit a linear relationship, especially in economic
and financial data. However, you can still apply the simple linear regression model by modifying either the dependent or
independent variables. Common transformations include using natural logarithms (log), reciprocals, squares, or differences of
variables. Three frequently used transformations involve log-based forms: the log-lin model (dependent variable is logarithmic,
independent is linear), the lin-log model (dependent variable is linear, independent is logarithmic), and the log-log model (both
variables are in logarithmic form). These transformations help make non-linear relationships suitable for linear regression
analysis.
In the log-lin model, only the dependent variable is in logarithmic form as follows:
➔ lnYi = b0+b1Xi
81
The slope coefficient in this model is the relative change in the dependent variable for an absolute change in the independent
variable. In this model, care must be taken when calculating the dependent variables value. The equation gives lnY to us. We
need to find Y using exponent of that value. Also, comparing a log-lin model with a lin-lin model (where both variables are in
their original form) is not straightforward. To make a comparison, we’d need to transform R-squared and the F-statistic.
However, examining the residuals can provide valuable insights in such cases.
The lin-log model is similar to the log-lin model, but only the independent variable is in logarithmic form:
➔ Yi = b0 + b1 lnXi
The slope coefficient in this regression model provides the absolute change in the dependent variable for a relative change in
the independent variable.
The log-log model, in which both the dependent variable and the independent vari- able are linear in their logarithmic forms, is
also referred to as the double-log model.
➔ lnYi = b0 + b1 lnXi.
This is useful in calculating elasticities as the slope coefficient is the relative change in the dependent variable for a relative
change in the independent variable.
To determine the right functional form for a simple linear regression, it's crucial to assess goodness of fit measures like R-
squared, the F-statistic, and the standard error of the estimate (s e), and check for patterns in the residuals. Most statistical
software also offers residual plots, allowing visual inspection. It's important that the plots show random residuals
82
Quantitative Methods - Learning Module 8
for the CFA exam
Introduction to Big Data Techniques
Learning Outcomes :
a) Describe aspects of “fintech” that are directly relevant for the gathering and analyzing of financial data.
b) Describe Big Data, artificial intelligence, and machine learning
c) Describe applications of Big Data and Data Science to investment management
1. Introduction
The convergence of finance and technology, or fintech, is transforming investment management by incorporating Big Data,
artificial intelligence, and machine learning. This shift is influencing both quantitative and fundamental asset managers, enabling
them to utilize advanced tools for evaluating investments, optimizing portfolios, and managing risks.
LOS (a) : Describe aspects of “fintech” that are directly relevant for the gathering and analyzing of financial data.
LOS (b) : Describe Big Data, artificial intelligence, and machine learning
Fintech broadly refers to technology-driven innovation in the financial services industry, specifically in the design and delivery of
financial products and services. It encompasses the development of new technologies and applications by companies in the
sector, challenging traditional business models. Early fintech involved data processing and task automation, evolving into
sophisticated decision-making applications based on machine learning. This technological evolution has significantly impacted
various aspects of the financial services industry, introducing new systems for investment advice, financial planning, business
lending, and payments.
Fintech, while encompassing various services, particularly impacts quantitative analysis in investment with a focus on:
o Analysis of Large Datasets: Traditional and alternative data, including social media and sensor networks, are integrated
into portfolio managers' decision-making processes. This enables the utilization of diverse data sources to enhance
investment decisions and manage risks.
o Analytical Tools and Artificial Intelligence (AI): AI, capable of tasks previously requiring human intelligence, is employed for
extremely large datasets. This facilitates the identification of complex, non-linear relationships more effectively than
traditional quantitative methods. AI-based techniques assist in sorting through vast amounts of data from various sources,
such as company filings and reports, to determine key information, uncover trends, and generate insights related to human
sentiment and behavior.
The term Big Data refers to the vast and diverse information generated by various sources since the late 1990s. This includes
traditional data from stock exchanges, companies, and governments, as well as alternative data from electronic devices, social
media, sensor networks, and business operations. Traditional data sources encompass corporate data like annual reports and
financial market data, while the increasingly connected world provides data from devices such as smartphones, cameras, and
satellites. The use of non-traditional or alternative data sources, like social media, emails, web traffic, and online news, has
grown with the expansion of the internet and networked devices.
o Volume : Involves a vast amount of data, often in the range of millions or billions of data points, stored in files, records, and
tables.
o Velocity : Data recording and transmission occur at an accelerated speed. Real-time or near-real-time data collection has
become common in many areas.
o Variety : Data is collected from diverse sources and comes in various formats, including structured data (e.g., SQL tables),
semi structured data (e.g., HTML code), and unstructured data (e.g., video messages). Structured data, organized in tables,
83
is typically stored in databases with consistent field types. Unstructured data, like that from social media, emails, and
images, lacks organization and requires specialized applications or custom programs for usability. Analyzing data from
sources like emails or texts may necessitate specially developed code. Semi structured data exhibits characteristics of both
structured and unstructured data.
In the context of using Big Data for inference or prediction, a crucial aspect is veracity, representing the credibility and reliability
of different data sources. Assessing the trustworthiness of data sources is essential in empirical investigations, with veracity
becoming particularly critical in Big Data due to the diverse and numerous sources involved. Big Data accentuates the
longstanding challenge of distinguishing quality from quantity in data analysis.
o Financial Markets: Data from equity, fixed income, futures, options, and other derivatives.
o Businesses: Information such as corporate financials, commercial transactions, and credit card purchases.
o Governments: Data covering trade, economic indicators, employment, and payroll.
o Individuals: Personal data like credit card purchases, product reviews, internet search logs, and social media posts.
o Sensors: Information from satellite imagery, shipping cargo details, and traffic patterns.
o Internet of Things (IoT): Data generated by "smart" buildings, providing information on climate control, energy
consumption, security, and operational details.
Traditionally, business intelligence relied on statistical methods and traditional data sources. However, the analysis of Big Data
now involves leveraging alternative data sources, such as retail sales, social media sentiment, and satellite imagery. These
alternative datasets offer additional insights into consumer behavior, firm performance, and trends, transforming the approach
of professional investors, especially quantitative investors, in financial analysis and decision-making processes.
The three primary sources of alternative data are data generated by individuals, business processes, and sensors:
o Individual Data: Produced in text, video, photo, and audio formats, often unstructured, and includes online activities like
social media interactions and e-commerce transactions.
o Business Process Data: Structured data from corporations, including direct sales information like credit card data and
corporate exhaust such as supply chain information and retail scanner data. Business process data can serve as leading or
real-time indicators of business performance.
o Sensor Data: Collected from devices like smartphones, cameras, RFID chips, and satellites, typically unstructured. The
volume of sensor data is significantly larger than individual or business process data, fueled by the widespread integration
of microprocessors and networking technology, especially in the Internet of Things (IoT) network arrangement.
84
Alternative data is increasingly employed to identify new factors impacting security prices, enhance asset selection, improve
trade execution, and uncover trends in data-driven investment models. The growing interest in alternative data has led to the
rise of specialized firms collecting and selling such datasets. However, as the market expands, investment professionals must be
aware of potential legal and ethical issues, especially concerning information not in the public domain. This includes
considerations related to privacy regulations, as the scraping of web data may capture personal information without explicit
consent. Best practices are still evolving, and regulatory approaches vary, posing potential conflicts in guidance across
jurisdictions.
LOS (b) : Describe Big Data, artificial intelligence, and machine learning
Artificial intelligence (AI) involves computer systems performing tasks traditionally requiring human intelligence, exhibiting
cognitive and decision-making abilities comparable to or superior to humans. Early AI, such as expert systems, simulated human
expertise using "if-then" rules. By the late 1990s, AI expanded into logistics, data mining, finance, and medical diagnosis. In
finance, AI, especially neural networks inspired by the human brain, has been employed since the 1980s, such as in credit card
fraud detection systems.
Machine learning (ML) involves computer-based techniques that extract knowledge from data without assumptions about its
underlying distribution. ML aims to automate decision-making by learning from known examples, emphasizing the algorithm's
ability to generate patterns or predictions without human assistance. ML historically faced limitations due to insufficient data for
training, but the growth of Big Data has improved modelling accuracy. In ML, algorithms learn from inputs and, if available,
outputs, refining their learning process by identifying relationships in the data.
ML involves dividing the dataset into training, validation, and test subsets. The training dataset helps the algorithm identify
relationships in historical data, validated with the validation set, and tested for predictive accuracy with the test set. Clean and
unbiased data are crucial for ML, and the model's performance may suffer without sufficient training data. Analysts must be
cautious about overfitting, where the model learns noise as true parameters, and underfitting, where true parameters are
treated as noise. Overfitted models may be too complex, while underfitted ones may fail to recognize data patterns. ML
techniques, being not explicitly programmed, may be perceived as opaque or black box approaches with outcomes that are not
entirely understood or explainable.
ML encompasses techniques for identifying relationships, detecting patterns, and creating structure from data. It includes
supervised learning (using labelled data to predict outcomes), unsupervised learning (describing data and structure without
labels), and deep learning (utilizing neural networks for multistage, non-linear processing to identify patterns). Neural networks,
existing since 1958, have seen advances in algorithms, leading to more accurate models for activities like image and speech
recognition. These improvements require less computing power, allowing analysts to uncover insights and relationships that
were previously challenging or time-consuming to discover.
The integration of machine learning (ML) techniques with traditional statistical methods for analyzing Big Data is a significant
development in investment research. This is facilitated by increased data availability, advanced algorithms, improved computing
power, faster software processing, and reduced storage costs. ML is employed in predicting trends or market events, such as
successful mergers or election outcomes. Image recognition algorithms, analyzing data from satellite imaging, offer insights into
various aspects like retail store parking lot activity, shipping, manufacturing, and agricultural crop yields. This information can
inform valuation or economic models at individual, national, or global levels.
LOS (c) : Describe applications of Big Data and Data Science to investment management
85
Data science is an interdisciplinary field that utilizes advances in computer science, including machine learning, and statistics to
extract information from Big Data. Companies depend on data scientists/analysts to derive insights for various business and
investment purposes. The structure of the data is a crucial consideration for data scientists, especially when dealing with the
unstructured nature of alternative data, which often requires specialized treatment before analysis.
Data scientists employ various data processing methods for Big Data analysis, including capture, curation, storage, search, and
transfer:
o Capture: Involves collecting and transforming data into a usable format. Low-latency systems are crucial for real-time
applications like automated trading, while high-latency systems don't require real-time data.
o Curation: Focuses on ensuring data quality and accuracy through a cleaning process, detecting errors, and adjusting for
missing data.
o Storage: Involves recording, archiving, and accessing data, considering whether it's structured or unstructured and whether
low-latency solutions are needed.
o Search: Addresses how to query data, necessitating advanced applications for examining large data quantities to find
requested content.
o Transfer: Deals with moving data from the source or storage to analytical tools, which can be done through direct data
feeds, such as a stock exchange's price feed.
Data visualization is crucial for understanding Big Data, involving the formatting, display, and summarization of data in graphical
form. Traditional structured data can be visualized using tables, charts, and trends, while non-traditional unstructured data
require innovative visualization techniques. Examples include interactive 3D graphics for exploring specified data ranges and
multidimensional data analysis using color, shapes, and sizes. Various solutions like heat maps, tree diagrams, and network
graphs represent data structures through geometry in interactive graphics. For textual data, techniques like "tag clouds" size
words based on frequency, and "mind maps" depict relationships between different concepts. These visualisations were
discussed in a previous module.
Fintech is being used in numerous areas of investment management. Applications for investment management include text
analytics and natural language processing, risk analysis, and algorithmic trading.
Text analytics entails utilizing computer programs to analyze and derive meaning from large, unstructured text- or voice-based
datasets, such as company filings, reports, social media, and surveys. It involves automated information retrieval from diverse
sources to aid decision-making, encompassing lexical analysis and pattern recognition based on keywords and phrases.
Predictive analysis can leverage text analytics to identify indicators of future performance, such as consumer sentiment.
Natural Language Processing (NLP) is a computer science, AI, and linguistics field focused on developing programs to analyze
and interpret human language. In text analytics, NLP performs tasks like translation, speech recognition, text mining, sentiment
and topic analysis. It is used in compliance functions to review communications for policy adherence, fraud detection, and
privacy protection. NLP, especially with ML, can analyze vast amounts of textual and audio data, providing insights and trends
more efficiently than human assessment.
For investment decision-making, NLP can monitor analyst commentary, assign sentiment ratings, and detect shifts ahead of
recommendations. It can analyze communications from policymakers, extracting insights from subtle messages. NLP models
may use non-traditional information, such as social media sentiment, to identify trends and short-term indicators, impacting
investment performance. Past research has explored Twitter sentiment's predictive power on IPO performance and the
influence of news sentiment on stock returns.
86
• Excel VBA: Bridges programming and manual data processing in Excel, automating tasks like updating data tables and running queries.
• SQL: Query language for structured data stored in tables, accessed via server-based SQL queries.
• SQLite: Embedded database for structured data, commonly used in mobile apps.
• NoSQL: Database for unstructured data, not organized in traditional tables with rows and columns.
87
Appendices
for the CFA exam
88
1.2. Standard Normal Distribution – For negative Z value
89
2. Appendix B: STUDENT’S T-DISTRIBUTION
90
3. Appendix C: F-TABLE AT 5% (UPPER TAIL)
91
4. Appendix D: F-TABLE AT 2.5% (UPPER TAIL)
92
5. Appendix E: CHI-SQUARED TABLE
93