Understanding Probability Concepts
Understanding Probability Concepts
Subjective
Addition rule
Solving for I/Y, CAGR, PMT, N
Bayes' formula
Structured (numbers) / Unstructured (text, audio, video, images)
Organizing data
Combination (tổ hợp)
Principles of counting
Frequency distributions
Permutation (chỉnh hợp)
Contingency table
Discrete => Probability mass function: p(x) = P(X = x)
Visualization
Random variables Continuous => Probability density function (pdf): f(x) = P(X = x)
Median
Discrete and continuous uniform
Mode
Bernoulli trial/random variable
Binomial
Outliers
Binomial random variable (x): number of successes in Bernoulli trial Probability of 'x' successes in 'n' trials
Organizing, Visualizing, and Describing Data
Winsorized mean
Univariate/Multivariate distribution
Other means
Geometric mean
'n' Means
Harmonic mean
Common Probability Distributions Multivariate normal distribution for 'n' variables 'n' Variances
Trimmed mean
Suitable for modelling quarterly/yearly returns; NOT suitable for modelling asset prices
Winsorized mean
Standardizing a random variable
Sample variance and standard deviation Measures of Central Tendency & Dispersion
Modern portfolio theory (MPT): the
value of investment opportunities can
Range be meaningfully measured in terms of
mean return and variance of return.
Application of normal distribution
Mean absolute deviation (MAD)
Shortfall risk: The risk that portfolio value
or portfolio return will fall below some
Sample variance and standard deviation minimum acceptable level over some time
horizon.
L = (n+1)y/100 Location (L) of the 'y'th percentile with 'n' data entries Quartiles / Quintiles / Deciles / Percentiles
Stress testing: A specific type of scenario
analysis that estimates losses in rare and
Box and whisker plot
extremely unfavorable combinations of
events or scenarios.
Target downside deviation (target semi- Mean (μL) of a lognormal random variable = e^(μ + 0.5σ^2)
deviation) and coefficient of variation Lognormal distribution
Normal
Kurtosis Volatility measures the standard deviation of the continuously compounded returns on the
underlying asset (by convention, it is stated as an annualized measure typically done on the
basis of 250 days in a year - the approximate number of days markets are open for trading).
Correlation
Quantitative
methods t-distribution has fatter tails than normal distribution => more reliable and
conservative downside risk estimate
t-test, Chi square, F-test
Weaknesses:
Provides only statistical estimates, not exact results.
Analytic methods, where available, provide more insights into cause-effect relationship.
Convenience
Non-probability Small scale pilot studies
Judgmental
Auditing
Simple random
Stratified random: population is divided into strata based on classification criteria; simple
Appropriate test statistics random samples are then drawn from each stratum proportionally to the relative size of
Sampling methods each stratum in the population to form a large sample.
Level of significance
Cluster: divides a population into clusters representative of the population and then
randomly draws certain clusters to form a sample.
1-tailed test
Relatively less accurate but more time-efficient and cost-efficient
Decision rule
2-tailed test
Sampling error: difference between the observed value of a statistic and the quantity it is
Sampling and Estimation intended to estimate as a result of sampling.
Sampling Distribution of a Statistic: the distribution of all the distinct possible values that
the statistic can assume when computed from samples of the same size randomly drawn
from the same population.
Hypothesis Testing
CLT: Given a population described by any probability distribution having mean μ and finite
variance σ^2, the sampling distribution of the sample mean computed from random
samples from this population will be approximately normal with mean μ (the population
Making a decision Central limit theorem (CLT) & Distribution of the sample mean
mean) and variance σ^2/n (the population variance divided by n) when the sample size n is
Statistical significance ≠ economic significance
large (n ≥ 30), regardless of the population's distribution.
False discovery rate (FDR): The rate of Type I errors in testing a null hypothesis multiple Consistent: one for which the probability of estimates getting close to the value of the
times for a given level of significance. population parameter increases as sample size increases.
Tests of variances Jackknife: repeatedly draws samples by taking the original observed data sample and
leaving out one observation at a time (without replacement) from the set.
2 variances
Sample selection bias: Bias introduced by systematically excluding some members of the
When there are outliers. Usage Parametric vs Non-parametric tests Sampling biases
population according to a particular attribute—for example, the bias introduced when data
availability leads to certain observations being excluded from the analysis.
When the data are given in ranks or use an ordinal scale. implicit selection bias: selection bias introduced through the presence of a threshold
that filters out some unqualified members.
Time-period bias: statistical conclusion may be sensitive to the starting and ending dates of
the sample.
Homoskedasticity
Introduction to Linear Regression
Assumptions
Independence
Normality
Analysis of variance
Sum of squared regressions (SSR): (ANOVA)
measures variation in observed values
attributable to the relationship between
the dependent and independent variables
Slope coefficient
Hypothesis testing of linear regression coefficients
Intercept
Log-lin
Lin-log
Functional forms of SLR
Log-log: useful in calculating elasticities because the slope coefficient is the relative change
in the dependent variable for a relative change in the independent variable.
Selecting functional forms: examining the goodness of fit measures (R-squared, F-statistic,
and the standard error of the estimate), and whether there are patterns in the residuals.