0% found this document useful (0 votes)
1 views10 pages

02 QuantMethods Expanded (1)

The document provides expanded notes on CFA Level I Quantitative Methods, focusing on key concepts such as interest rates, holding period returns, and time value of money. It explains various financial metrics, including money-weighted vs. time-weighted returns, and details statistical measures of asset returns, portfolio mathematics, and probability trees. Additionally, it covers essential formulas and exam tips to aid in understanding and applying these concepts effectively.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
1 views10 pages

02 QuantMethods Expanded (1)

The document provides expanded notes on CFA Level I Quantitative Methods, focusing on key concepts such as interest rates, holding period returns, and time value of money. It explains various financial metrics, including money-weighted vs. time-weighted returns, and details statistical measures of asset returns, portfolio mathematics, and probability trees. Additionally, it covers essential formulas and exam tips to aid in understanding and applying these concepts effectively.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

CFA Level I — Quantitative

Methods
Expanded Notes — explanations, worked examples, and formula walkthroughs added to the original condensed
notes.

1. Rates and Returns


What is "the" interest rate, really?
Your notes list about six names for the same underlying number: required rate of return, yield, market discount
rate, opportunity cost, time value of money, YTM, I/Y. These are all the SAME rate viewed from different
angles — the rate that connects a price today to cash flows in the future. The exam likes to test whether you
recognize these are interchangeable labels depending on context (bond pricing calls it YTM, capital budgeting
calls it required return, your calculator calls it I/Y).

Building up an interest rate from its parts


Think of an interest rate as a stack of compensations, each one paying you for a specific risk you're taking on:

Nominal risk-free rate = Real risk-free rate + Inflation premium


Required rate of return = Nominal risk-free rate + Default premium +
Liquidity premium + Maturity premium
● Real risk-free rate: compensation for just giving up money today, assuming zero inflation and zero
risk — the pure "time value."
● Inflation premium: compensation because the money you get back later will buy less.
● Default premium: compensation for the chance the borrower doesn't pay you back.
● Liquidity premium: compensation for holding something that's hard to sell quickly without a price
discount.
● Maturity premium: compensation for the extra price risk that comes with longer time to maturity
(long bonds swing more when rates change).
Exact vs. additive version: The precise (non-approximated) relationship between nominal and real rates
is: (1 + R_nominal) = (1 + R_real)(1 + inflation premium). The simple "add them up" version most people
use is just an approximation that works when rates are small.
● Sunk cost: money already spent, which cannot be recovered by any future decision. A sunk cost
should NEVER factor into a forward-looking decision — a classic behavioral trap the exam tests (e.g.,
"I already spent $10,000 on this project, so I have to keep funding it" is flawed reasoning).

Holding Period Return (HPR)


HPR answers: "what did I actually earn over the period I held this, including any income I received?"

1-year HPR = (P1 − P0 + Income) / P0


Multi-year HPR = (1+r1)(1+r2)(1+r3)... − 1 [do NOT take a root
here]
Exam tip: The most common mistake: taking the geometric root of a multi-year HPR because it "feels
like" an average return question. HPR over multiple years is a straight compounded total — the root only
comes in when you're finding an ANNUALIZED average (geometric mean), which is a different question.

Money-weighted return (IRR) vs. Time-weighted return


We worked through a full numeric example of this earlier — here's the formal recap:
● Money-weighted return (IRR): the discount rate that sets the present value of all cash flows (in and
out) equal to zero. Because it's built directly from the actual dollar cash flows, both the SIZE and the
TIMING of your contributions/withdrawals affect this number.
● Time-weighted return: (1) calculate the HPR for each sub-period between cash flows, including any
dividends even if not reinvested, then (2) chain-link (multiply) those HPRs together, or use the
geometric mean if you want an annualized figure. This measure deliberately ignores the size and timing
of cash flows — it only reflects the sequence of period-by-period returns.
Exam tip: Use money-weighted whenever the question emphasizes an investor's personal dollar
experience or a manager who controls cash flow timing (private equity). Use time-weighted whenever the
question is judging a manager's pure stock-picking skill, independent of when clients added/withdrew
money (mutual funds).

Continuously compounded return


From time 0 to t: R = ln(1 + r)
From time t to t+1: R(t,t+1) = ln(Price_t+1 / Price_t)
Think of continuous compounding as compounding "infinitely often" — every instant, not just
annually/quarterly. Taking the natural log of a price relative converts a simple holding return into its
continuously compounded equivalent.

Annualizing and comparing rates


R_annual = (1 + R_period)^c − 1, where c = number of periods per
year
Example: If a fund has earned a 6% return since inception 15 months ago, c = 12/15 = 0.8. Annualized
return = (1.06)^0.8 − 1 ≈ 4.79%. Because c is less than 1, you're compressing a 15-month return down
into a 12-month equivalent.
● Nominal rate (stated annual rate): quoted per annum, but stated alongside its compounding
frequency (e.g., "6% compounded quarterly").
● Effective annual yield (EAY): the TRUE annual rate once you account for compounding — this is
what you actually earn, and is always ≥ the nominal rate whenever compounding happens more than
once a year.
● C/Y: compounding periods per year — how many times interest is credited annually (2 = semiannual,
4 = quarterly, 12 = monthly).

2. Time Value of Money (TVM)


Annuities: ordinary vs. due
● Ordinary annuity: cash flows occur at the END of each period (t = 1, 2, 3...). This is the
default/standard setting on your calculator ("END" mode) and the default assumption unless a question
says otherwise.
● Annuity due: cash flows occur at the START of each period (t = 0, 1, 2...) — think rent payments,
which are usually paid in advance. On a BA II Plus, switch to "BGN" mode (2ND, then BGN/END,
then 2ND, SET) before computing.
Exam tip: Forgetting to switch your calculator back to END mode after an annuity-due question is one of
the single most common careless errors on quant/fixed income problems. Always double check the mode
indicator before moving to the next question.

Continuous compounding (future value)


FV = PV × e^(r × N)
Here N is typically expressed in years (or fraction of a year). This is the FV counterpart to the continuously
compounded RETURN formula above — same concept (infinite compounding frequency), applied to growing a
dollar amount instead of measuring a return.

Perpetuities and constant-growth values


Perpetual bond: PV = PMT / (rs / m)
Constant (non-growing) dividend: PV_t = D_t / r
A perpetuity is just an annuity that never ends — the level payment (PMT or dividend D) divided by the
appropriate periodic discount rate gives you its value today.

Forward rates from two risk-free spot rates


F(1,1) = (1+r2)^2 / (1+r1) − 1
This comes directly from the no-arbitrage idea that investing for 2 years directly should equal investing for 1
year and then reinvesting at the 1-year-forward rate:

FV2 = PV0 (1+r2)^2 = PV0 (1+r1)(1+F1,1)


Reading the notation: F(1,1) means: the 1-year rate, 1 year from now (i.e., the forward rate for the second
year). This is a core building block for bond pricing and the term structure of interest rates — you'll see
this idea again in Fixed Income.

3. Statistical Measures of Asset Returns


A parameter describes a POPULATION characteristic (the true, usually unknowable value); a statistic describes
a SAMPLE characteristic (what you actually calculate from data you have).

Measures of central tendency


● Arithmetic mean: simple average — best for a single-period expected value, but overstates the TRUE
compound growth rate over multiple periods.
● Geometric mean: the Nth root of the product of (1+return) terms, minus 1. This is the number that

Geometric mean return = [(1+r1)(1+r2)...(1+rn)]^(1/n) − 1


correctly reflects actual compounded multi-period growth.

● Harmonic mean: used for averaging ratios (like cost-per-share when investing a fixed dollar amount
each period, e.g., dollar-cost averaging). It's less sensitive to outliers than the other means.
Ranking rule to memorize: Harmonic mean ≤ Geometric mean ≤ Arithmetic mean, whenever there is any
variability in the data. If there's NO variability (every return is identical), all three are equal. Bonus
identity: Arithmetic mean × Harmonic mean = (Geometric mean)².
● Trimmed mean: remove a set percentage of extreme values from BOTH ends, then average what's
left.
● Winsorized mean: instead of removing extreme values, REPLACE them with the nearest remaining
(non-extreme) value, then average.
● Median: the true middle value. Odd n: the (n+1)/2th observation. Even n: average of the n/2th and (n/2
+ 1)th observations.
● Mode: the most frequently occurring value.

Measures of dispersion (how spread out the data is)


● Absolute dispersion: range, variance, standard deviation — all measured in the SAME units as the

Mean Absolute Deviation (MAD) = Σ|Xi − X̄| / n


original data.
Target semideviation (target B, only values ≤ B) = same formula as
SD but using (Xi − B), with denominator n−1
● Relative dispersion — Coefficient of Variation (CV): standard deviation per unit of mean return —
i.e., risk per unit of reward. Useful for comparing riskiness across investments with very different

CV = s / X̄
average return levels.

Measures of shape (skewness and kurtosis)


● Positive (right) skew: mode < median < mean. There's a long right tail — a few big gains pull the
mean up. Generally preferred by investors, since it means more upside surprises than downside ones.
● Negative (left) skew: frequent small gains punctuated by a few extreme losses — the tail risk investors
fear most.
● Kurtosis: measures how "fat" the tails of a distribution are relative to a normal distribution — i.e., how
likely extreme outcomes are.
○ Leptokurtic (fat-tailed): kurtosis > 3 (excess kurtosis > 0). Higher peak, fatter tails, MORE
extreme outcomes than normal. Most real-world equity return data is leptokurtic — a key
reason risk models that assume normality underestimate tail risk.
○ Mesokurtic: kurtosis = 3 (excess kurtosis = 0) — this is the normal distribution itself.
○ Platykurtic (thin-tailed): kurtosis < 3 (excess kurtosis < 0). Lower peak, thinner tails,
FEWER extreme outcomes (mnemonic: "platy's no fatty").

Measures of location (quantiles)


● Percentile: e.g., the 10th percentile is the value below which 10% of observations fall.
● Interquartile range (IQR): = Q3 − Q1, the spread of the "middle 50%" of the data — a robust

Position of the yth percentile: Ly = (n+1) × (y/100)


measure of dispersion that isn't distorted by extreme outliers.

Value at that position (linear interpolation): Py = X(Ly's whole number


position) + (Ly's decimal part) × [X(next position) − X(that position)]
Worked example: Data: 12, 15, 18, 22, 30 (n=5). Find the 50th percentile (median, y=50). Ly = (5+1)
(50/100) = 3.0 exactly → the 3rd observation → 18. If instead Ly had come out to, say, 3.4, you would
interpolate 40% of the way between the 3rd and 4th observations.
● Box and whisker plot: the line inside the box marks the median (not the mean).

4. Probability Trees and Conditional Expectations


Variance of a random variable
σ²(X) = E[X − E(X)]² = Σ P(Xi) × (Xi − E(X))²
In words: for each possible outcome, find how far it is from the expected value, square that distance (so negative
and positive deviations don't cancel out), weight by its probability, and sum.

Conditional expectation and the total probability rule


E(X|S) = expected value of X, given that scenario S has occurred
Total Probability Rule for Expected Value: E(X) = E(X|S)×P(S) + E(X|
S^c)×P(S^c)
This lets you build up an unconditional (overall) expected value from a set of scenario-conditional expected
values, weighted by how likely each scenario is. Probability trees work backward: you estimate the probability
of reaching each end node, then average the outcomes weighted by those probabilities to get today's expected
value.
Key nuance: The branches of a probability tree are mutually exclusive and exhaustive (every path is
covered, no overlap), but the branches are NOT independent — the probabilities at a second-stage node
are CONDITIONAL on which first-stage branch you took.

Bayes' formula
P(A|B) = [P(B|A) × P(A)] / P(B)
Bayes' formula updates your probability estimate for event A once you learn new information B. Read it as: "my
updated belief about A, given I just observed B, equals how likely B would be if A were true, times how likely
A was to begin with, scaled by how likely B is overall."
Exam tip: On the exam, identify which piece is the 'new information' (B) and which is the 'event you care
about' (A) before plugging into the formula — mixing these up is the most common error.

5. Portfolio Mathematics
Portfolio variance (two assets)
σ²(Rp) = w1²σ²(R1) + w2²σ²(R2) + 2w1w2 Cov(R1,R2)
For three assets, you add a third variance term and two more covariance cross-terms (covering every unique
pair):

σ²(Rp) = w1²σ²(R1) + w2²σ²(R2) + w3²σ²(R3) + 2w1w2 Cov(R1,R2) +


2w1w3 Cov(R1,R3) + 2w2w3 Cov(R2,R3)
Counting rule: With n assets, there are n variance terms and n(n−1)/2 unique covariance terms. This
combinatorial count is a common standalone exam question — e.g., '10 stocks in a portfolio: how many
covariance terms?' → 10×9/2 = 45.
● Large-portfolio intuition: as the number of assets (N) grows very large, the AVERAGE
COVARIANCE between holdings dominates portfolio variance, while the influence of any single
asset's own variance shrinks toward zero. This is the mathematical basis for why diversification
reduces risk but can never eliminate it — you can't diversify away the average covariance.

Covariance and correlation


Cov(Ri,Rj) = σXY = E[(Ri − E(Ri))(Rj − E(Rj))] = ρ(Ri,Rj) × σ(Ri) ×
σ(Rj)

Correlation: ρ = σXY / (σX × σY)


Correlation is just covariance RESCALED to always fall between −1 and +1, which makes it far easier to
interpret than covariance's unbounded raw units.
● ρ = +1: no diversification benefit at all — combining the assets doesn't reduce risk, it can even increase
volatility relative to holding the lower-risk asset alone.
● ρ = 0: assets are unrelated — this gives strong diversification benefit.
● ρ = −1: perfect negative correlation — theoretically, you could combine the two assets in the right
weights to build a fully risk-free (zero variance) portfolio.
Correlation only captures the LINEAR relationship between two variables — two variables can have a strong
nonlinear relationship while showing near-zero correlation.
Independence vs. uncorrelated — a critical distinction
Independence: P(X,Y) = P(X) × P(Y) [the joint probability equals the
product of individual probabilities]
Uncorrelated: E(XY) = E(X) × E(Y)
Exam tip: Independence is a STRONGER condition than being uncorrelated. Two variables can be
uncorrelated (no LINEAR relationship) while still being dependent on each other through a nonlinear
relationship. Independent variables are always uncorrelated, but uncorrelated variables are not necessarily
independent.

6. Simulation Methods
Lognormal distribution
● Continuously compounded return = ln(Ending value / Beginning value).

Monte Carlo simulation


Uses the inverse transformation method to turn randomly generated, uniformly distributed numbers into
simulated values drawn from whatever probability distribution you actually want to model. It's extremely useful
for risk quantification and predictive modeling under uncertainty, but it's fundamentally a statistical estimation
tool — it cannot tell you WHY something happens (no cause-and-effect insight), only what range of outcomes
is plausible given your assumed inputs.

Resampling methods
● Bootstrap method: repeatedly draws samples from your existing data WITH replacement (the same
observation can be picked more than once).
● Jackknife method: repeatedly leaves out observations WITHOUT replacement — sample size and
repetition count both equal n (the original sample size), systematically leaving out one observation at a
time.

7. Estimation and Inference (Sampling)


Probability (random) sampling methods
● Simple random sampling: every member of the population has an equal chance of selection.
● Systematic sampling: select every kth member from a list.
● Stratified random sampling: divide the population into strata/subgroups (e.g., by sector or market
cap), THEN simple-random-sample from within each stratum. This ensures every subgroup is
proportionally represented.
● Cluster sampling: divide the population into clusters, then randomly select entire CLUSTERS (not
individual members within them) to include.
● Non-probability sampling: convenience sampling (whatever data is easiest to access) and judgmental
sampling (based on the researcher's own judgment about what's representative) — both introduce
potential selection bias.

The Central Limit Theorem (CLT)


As sample size grows (a commonly cited rule of thumb threshold is around n ≥ 30), the distribution of the
SAMPLE MEAN approaches a normal distribution — regardless of the shape of the underlying population
distribution. This is what allows us to use normal-distribution-based confidence intervals and hypothesis tests
even when we don't know (or the data isn't) normally distributed.

Sample variance of the mean = σ² / n


Standard error of the sample mean = σ / sqrt(n)

8–9. Hypothesis Testing


The core logic
Hypothesis testing exists to answer one question: is a pattern we observe in a SAMPLE real (reflects the true
population), or could it just be random noise? The whole framework is built around trying to REJECT a
skeptical default position.
● Null hypothesis (H0): the skeptical default — "there is NO relationship / no difference / no effect."
This is the hypothesis you're trying to find evidence AGAINST.
● Alternative hypothesis (Ha): what you'll conclude if you successfully reject H0 — "there IS a
relationship/difference."
● Critical value: the threshold — if your test statistic is more extreme than this, you reject H0.
● P-value: the smallest significance level at which you could still reject H0. Smaller p-value = stronger
evidence against H0. Easiest rule: if p-value < your chosen significance level (α), reject H0.

Test statistics — quick reference


● Z-test (single mean): Z = (X̄ − μ) / (σ/√n). Use when the population variance IS known, with a large or
normally distributed sample. Key z-values: 90% CI → z=1.65; 95% CI → z=1.96; 99% CI → z=2.58
(these map to the empirical rule: ~68% within 1σ, ~95% within 2σ, ~99% within 3σ).
● T-test (single mean): t = (X̄ − μ0) / (s/√n), with n−1 degrees of freedom. Use when population
variance is UNKNOWN (this is the far more common real-world case) — becomes closer to the Z-
distribution as sample size grows.
● Difference between means (2 independent samples): degrees of freedom = n1+n2−2. Uses a
POOLED variance estimate, sp², IF you can assume the two populations share equal variance. If

sp² = [(n1−1)s1² + (n2−1)s2²] / (n1+n2−2)


variances are unequal/unknown, s1 and s2 are used separately instead of a pooled estimate.

● Mean of differences (2 DEPENDENT/paired samples): the paired comparison t-test, df = n−1. Use
this when the two samples aren't independent — e.g., testing the SAME stocks' returns before vs. after

t = (d̄ − μd0) / (sd / sqrt(n)), where d̄ = mean of the paired


an event.

differences, sd = std error of those differences = s_d / sqrt(n)


● Single variance (chi-square, 𝛘²): df = n−1, one-sided/asymmetric test. Tests whether a sample's
variance differs from a hypothesized value σ0². As n grows, the chi-square distribution's shape
becomes more bell-like.
● Difference in variances (F-test): F = s²(larger)/s²(smaller). Degrees of freedom = (n1−1, n2−1) for the
two samples being compared.
● Correlation significance (t-test for Pearson ρ): df = n−2. Tests whether a sample correlation r is
statistically different from zero.
● Test of independence (chi-square contingency table): used with categorical/discrete data in a two-
way (contingency) table, df = (r−1)(c−1) where r=rows, c=columns. H0: the two categorical variables
are independent (not related).

Type I and Type II errors


This 2×2 table is one of the highest-yield memorization items in the whole quant chapter:
● Fail to reject H0, and H0 is actually TRUE: correct decision. Probability = 1 − α (the confidence
level).
● Reject H0, but H0 is actually TRUE: Type I error ("false positive"). Probability = α, the significance
level you chose.
● Fail to reject H0, but H0 is actually FALSE: Type II error ("false negative"). Probability = β.
● Reject H0, and H0 is actually FALSE: correct decision — this is called the POWER of the test.
Probability = 1 − β.
Exam tip: Lowering your significance level (α) — e.g., testing at 1% instead of 5% — makes it HARDER
to reject H0, which reduces Type I errors but increases Type II errors. There's always a trade-off; you
cannot minimize both error types simultaneously by adjusting α alone.

Non-parametric tests
Used when: (1) the data doesn't meet the distributional assumptions a parametric test requires, (2) there are
significant outliers, (3) data is ranked/ordinal rather than continuous, or (4) the hypothesis isn't really about a
specific population parameter at all.
● Single mean (nonparametric alternative to Z/t-test): Wilcoxon signed-rank test.
● Difference between means (nonparametric alternative to t-test): Mann-Whitney U test.
● Mean differences, paired data: Wilcoxon signed-rank test or the (simpler) sign test.
● Spearman rank correlation (rs): rank each variable's observations from largest (1) to smallest (n);
tied values get the average of their joint ranks (e.g., a 3-way tie for 3rd/4th/5th place all get rank 4).

rs = 1 − [6 × Σdi²] / [n(n² − 1)]


Compute the difference di between each pair's ranks, square it, then:

10. Simple Linear Regression (SLR)


The model
Yi = b0 + b1·Xi + εi
● Y: the dependent/explained variable — the thing you're trying to predict. (On a financial calculator's
regression function, Y is entered SECOND.)
● X: the independent/explanatory variable — what you're using to predict Y.
● b0 (intercept): the estimated value of Y when X = 0. Formula: b̂ 0 = Ȳ − b̂ 1X̄ .
● b1 (slope): how much Y changes for a one-unit change in X. Formula: b̂ 1 = Cov(Y,X) / Var(X).
● εi (residual/error): the gap between the actual observed Yi and the model's predicted Ŷi. By
construction (least squares), the residuals sum to zero.
Indicator (dummy) variables as X: When the independent variable X is a 0/1 indicator (e.g., 1 = after a
policy change, 0 = before): b0 becomes the average value of Y in the '0' group, and b1 becomes the
DIFFERENCE in average Y between the '1' group and the '0' group.

Breaking down variation: SST, SSR, SSE


SST (Total) = Σ(Yi − Ȳ)² = SSR + SSE
SSR (Regression/explained) = Σ(Ŷi − Ȳ)²
SSE (Error/unexplained/residual) = Σ(Yi − Ŷi)² = Σei²
Think of SST as the total variation in Y that exists in your data. Your regression model explains SOME of it
(SSR) and leaves the rest unexplained (SSE, the residual noise). A smaller SSE means tighter, more accurate
predictions — which translates to a narrower prediction interval and a better-fitting model.
Regression assumptions (memorize all four)
● Linearity: the true relationship between X and Y is linear.
● Homoskedasticity: the variance of the residuals is CONSTANT across all levels of X (if it isn't —
heteroskedasticity — your standard errors become unreliable).
● Independence: observations (X,Y pairs) are independent of one another, and residuals are uncorrelated
across observations (no autocorrelation).
● Normality: the RESIDUALS (not the raw X or Y variables themselves) are normally distributed.

Goodness of fit
R² (coefficient of determination) = SSR / SST
R² tells you the percentage of the variation in Y that's explained by X. Useful shortcut for SLR specifically: if
the slope is positive, the correlation ρ(X,Y) = √R².

F-statistic = MSR / MSE


● MSR (Mean Square Regression) = SSR / k, where k = number of independent variables. In SLR,
k=1, so MSR simply equals SSR.
● MSE (Mean Square Error) = SSE / (n−k−1). In SLR this becomes SSE/(n−2).
The F-test's null hypothesis is that ALL slope coefficients equal zero simultaneously (for SLR with only one
slope, this collapses to the same test as testing whether b1=0).

T-test for a regression coefficient


t = (b̂1 − B1) / s(b̂1)
Where b̂ 1 is your estimated slope, B1 is the hypothesized population slope (usually 0, testing "is there really a
relationship at all"), and s(b̂ 1) is the standard error of the slope estimate:

s(b̂1) = se / sqrt[Σ(Xi − X̄)²]


se (Standard Error of the Estimate) = sqrt(MSE)
se measures the typical distance between actual observed Y values and the values your regression line predicts
— essentially the "average size of your mistakes." A smaller se means a tighter fit.

Prediction interval: Ŷf ± (t-critical for α/2) × Sf, where Sf = standard


error of the forecast

Non-linear regression forms


● Log-lin model: ln(Yi) = b0 + b1Xi — used when Y grows at a roughly constant PERCENTAGE rate
as X increases (common for modeling variables that grow exponentially, like GDP or stock prices over
long periods).
● Lin-log model: Yi = b0 + b1·ln(Xi) — used when Y's response to X flattens out at higher levels of X
(diminishing effect).
● Log-log model: ln(Yi) = b0 + b1·ln(Xi) — used when you want to interpret b1 as an elasticity (the %
change in Y for a 1% change in X).

11. Introduction to Big Data Techniques


● Fintech: technology-driven innovation reshaping the financial services industry.
● Big data: large volumes of financial data collected from many different sources, in multiple formats
(structured and unstructured).
A useful mental hierarchy of increasing sophistication:
● Expert systems: simple computer if-then rule sets — no real "learning" involved.
● Neural networks: loosely modeled on how the human brain processes information, using layers of
interconnected nodes.
● Machine learning: systems that train, validate, and test on data to improve performance. Can behave
like a "black box" — the outcome/logic isn't always fully interpretable even to its creators. Two failure
modes to know: overfitting (the model is TOO closely fit to the training data's noise, and performs
poorly on new data) and underfitting (the model is too simple to capture the real pattern at all).
○ Supervised learning: uses labeled data (inputs paired with known correct outputs).
○ Unsupervised learning: uses only unlabeled data and looks for hidden structure on its own.
○ Deep learning: uses many-layered neural networks to detect complex patterns.

Data science processing steps


● 1. Capture: collecting and formatting data for analysis — needs low-latency systems if real-time
calculation is required.
● 2. Curation: cleaning the data to ensure quality (removing errors, duplicates, inconsistencies).
● 3. Storage: recording, archiving, and enabling access to the data.
● 4. Search: querying the specific data you need.
● 5. Transfer: moving data from storage to the analytical tool that will actually process it.

Data visualization tools


● Tag cloud: visual where text size reflects the importance/frequency of each term.
● Text analytics: summarizing/extracting insight from unstructured text, often used for short-term trend
detection.
● Natural Language Processing (NLP): techniques for interpreting and translating human language
computationally.

You might also like