Chapter 1:
Exercise 1:
Answer:
(a) Continuous and Discrete Data
Continuous data can take on any value and is not confined to specific numbers, meaning
its values are limited only by the precision of the measurement (e.g., a rental yield of
6.238%). Discrete data, however, can only take on certain specific values, which are
usually integers or "count numbers," such as the number of shares traded in a day where a
fractional value would not make sense.
(b) Ordinal and Nominal Data
Ordinal data provides a position or a sequence where the order matters, but the distance
between the values is not necessarily equal; for example, being in second place is better
than fourth, but it is not "twice as good". Nominal data has no natural ordering at all, and
numerical values are assigned to categories arbitrarily, such as using "1" for the NYSE
and "2" for the NASDAQ.
(c) Time Series and Panel Data
Time series data consists of observations on one or more variables collected over a period
of time at a particular frequency, such as daily stock prices. Panel data possesses both
time series and cross-sectional dimensions, measuring the same set of entities (e.g.,
several different firms) over a period of time.
(d) Noisy and Clean Data
Noisy data is characterized by random and uninteresting features that make it difficult for
a researcher to isolate underlying trends or patterns. Clean data refers to the ideal, high-
quality "ingredients" required by perfect econometric theory; however, in the real world,
the "statistical cook" often lacks these and must substitute them with "messy" materials
that are available.
(e) Simple and Continuously Compounded Returns
Simple returns measure the percentage change in an asset's price from the previous period
and have the advantage of being additive across a portfolio of assets. Continuously
compounded returns (or log returns) are calculated as the natural logarithm of the price
ratio; while they are not additive across a portfolio, they are time-additive, meaning a
weekly return can be found by simply summing the daily log returns.
(f) Nominal and Real Series
Nominal series are those expressed in terms of the actual prices that existed at the time
the data was recorded. Real series are "inflation-adjusted" or at "constant prices,"
obtained by dividing a nominal series by a price deflator (like the CPI) to allow for like-
for-like comparisons across different time periods.
(g) Bayesian and Classical Statistics
Classical statistics follows a philosophy where a researcher postulates a theory and then
estimates a model to test whether that theory is upheld or refuted by the data. Bayesian
statistics starts with an assessment of existing beliefs (priors) formulated as probabilities,
which are then updated as more data becomes available to form "posterior probabilities".
--------------------------------------------------------------------------------
Analogy for Nominal vs. Real Series: Comparing nominal prices over time without
adjusting for inflation is like comparing the height of two children using different rulers;
one ruler might be marked in inches while the other is in centimeters. Converting the data
into a "real" series provides a standardized ruler, allowing for a fair, like-for-like
comparison by removing the "stretch" caused by the rising general level of prices over
time
Question 2: Answer
Present and explain a problem that can be approached using a time series
regression, another one using cross-sectional regression, and another using
panel data
Based on the provided sources, here are three problems that can be approached using
time series, cross-sectional, and panel data regressions, respectively:
1. Time Series Regression
Problem: Analyzing how the value of a country's stock index varies with that country's
macroeconomic fundamentals.
• Explanation: Time series data involves observations on variables collected over a
period of time at a specific frequency (e.g., daily, monthly, or quarterly). In this problem,
a researcher would collect data for a single country (the entity) over many time periods.
For instance, one might regress the monthly returns of the FTSE 100 index against
monthly data on inflation, interest rates, and GDP growth. The analysis focuses on the
time dimension to understand dynamic relationships or to forecast future values.
2. Cross-Sectional Regression
Problem: Investigating the relationship between company size and the return to investing
in its shares.
• Explanation: Cross-sectional data are collected at a single point in time across multiple
entities. In this scenario, a researcher would gather data on the size (e.g., market
capitalization) and the stock returns for a large number of different companies (e.g., all
firms on the NYSE) for just one specific year or day. The goal is to compare differences
between these entities at that fixed moment, rather than tracking them over time. Unlike
time series, there is no natural ordering of the observations in the sample.
3. Panel Data Regression
Problem: Modelling the investment behavior of companies based on their profits over a
decade.
• Explanation: Panel data (or longitudinal data) combines both time series and cross-
sectional dimensions, measuring the same set of entities over a period of time. In this
problem, the researcher would collect investment and profit data for a set of companies
(e.g., Company 1 to Company N) for each year from year 1 to year T. This approach
allows the researcher to control for unobserved heterogeneity (factors specific to each
company that do not change over time, like the sector they operate in) which would be
missed by a simple pooled regression,. It provides a richer dataset that can address more
complex issues than either time series or cross-sectional data alone.
Question 3: What are the key features of asset return time
series?
Answer:
The key features of asset return time series can be categorized into their statistical
properties, volatility characteristics, and time-series behavior.
1. Stationarity vs. Non-Stationarity
A fundamental distinction in financial econometrics is between asset prices and asset
returns.
• Prices are Non-Stationary: Asset prices (or the logarithms of asset prices) typically
follow a random walk or a random walk with drift. This means they are non-stationary
(containing a unit root), and shocks to the system persist infinitely.
• Returns are Stationary: Asset returns, calculated as the first difference of the log
prices, are generally stationary. This means they have a constant mean and variance over
time and tend to revert to their mean, making them suitable for statistical modeling.
2. Non-Normality and Leptokurtosis
While many standard econometric techniques assume data follows a normal (Gaussian)
distribution, financial asset returns almost always violate this assumption.
• Leptokurtosis (Fat Tails): Financial returns are typically leptokurtic. This means the
distribution has "fatter tails" and is more peaked at the mean than a normal distribution.
• Implication: Extreme events (very large positive or negative returns) occur more
frequently in real financial markets than a standard normal distribution would predict.
Consequently, assuming normality can lead to systematic underestimation of risk.
3. Volatility Clustering (Pooling) (Sự tích tụ biến động)
Volatility in financial markets is not constant (homoscedastic) but rather time-varying
and clustered.
• The Phenomenon: Large returns (of either sign) tend to follow large returns, and small
returns tend to follow small returns. Volatility appears in bunches (từng đoạn, từng cụm)
or bursts (các đợt bùng phát) rather than being evenly spaced (Phân bổ đồng đều) over
time.
• Autocorrelation in Volatility ( tính tự tương quan của biến động): While raw returns
may show little correlation, the volatility of returns is strongly autocorrelated. A period of
high volatility is likely to be followed by another period of high volatility. This motivates
the use of ARCH (Autoregressive Conditionally Heteroscedastic) and GARCH models.
4. Leverage Effects – Tác động đòn bẩy (Asymmetry – Tính bất đối xứng)
Volatility often responds asymmetrically to price shocks, a phenomenon known as the
leverage effect.
• The Effect: A drop in asset prices (bad news) tends to cause volatility to rise more than
an equivalent increase in asset prices (good news).
• Theoretical Explanation: As equity prices fall, a firm's debt-to-equity ratio increases,
making the firm's cash flows appear riskier to equity holders, thereby increasing future
volatility. Standard GARCH models enforce symmetry, so asymmetric models like GJR
or EGARCH are often required to capture this feature.
5. Noise and Microstructure
Financial data, particularly at high frequencies, is considered very "noisy."
• Signal-to-Noise Ratio: It is often difficult to distinguish underlying trends or patterns
from random, uninteresting features in the data.
• Market Microstructure: High-frequency data (e.g., daily or intra-day) often contain
specific patterns resulting from the mechanics of trading, such as bid-ask spreads and the
way prices are recorded.
6. Weak Autocorrelation in Raw Returns (Tỷ suất sinh lời gốc)
Unlike volatility, the raw returns themselves often exhibit little to no autocorrelation.
• Efficiency: If markets are informationally efficient, price changes should be largely
unpredictable, meaning the current return is not significantly related to previous returns.
• Random Walk: This is consistent with the efficient markets hypothesis, which suggests
asset prices follow a random walk, making returns unpredictable
Master's Tip: Trong bài thi Econometrics, nếu bạn chạy kiểm định Ljung-Box test trên
chuỗi Raw Returns, kết quả thường sẽ là "Fail to reject $H_0$" (nghĩa là không có tự
tương quan). Nhưng nếu chạy trên chuỗi Squared Returns ($r_t^2$), bạn sẽ thấy tự
tương quan rất mạnh. Đây chính là bằng chứng thép cho thấy lợi nhuận thì ngẫu nhiên
nhưng rủi ro thì có tính hệ thống. (Gemini explain)
--------------------------------------------------------------------------------
Analogy for Volatility Clustering: Think of asset return volatility like traffic jams. Traffic
doesn't usually slow down randomly for one minute and then speed up immediately. If
you hit a patch of heavy traffic (high volatility), it is highly probable that the next mile of
driving will also be in heavy traffic. Conversely, if the road is clear (low volatility), it is
likely to remain clear for a while. Just as traffic clusters together, so does financial
turbulence.
Chapter 3:
Question 1:
Answer:
OLS: minimize total sum of the squared errors
a. OLS minimizes vertical distances because of the fundamental assumption
regarding the independent variable (x). The regression model assumes
that the x variable is non-stochastic, meaning it has fixed values in
repeated samples. Consequently, the statistical problem is defined as
determining the appropriate model for the dependent variable (y) given (or
conditional upon) the observed values of x.
b. Why are the vertical distances squared before being added together?
The distances are squared to prevent positive and negative deviations
from cancelling one another out. If the method simply minimized the sum
of the residuals (∑u^t), points lying above the line (positive errors) would
cancel out points lying below the line (negative errors). This would allow
virtually any line passing through the mean of the observations to satisfy
the condition of summing to zero, resulting in no unique solution for the
estimated coefficients. Squaring ensures that all deviations contribute a
positive quantity to the total sum
c. Why are the squares of the vertical distances taken rather than the absolute
values? While the text acknowledges that minimizing absolute values is a possible
alternative (and is sometimes preferred when outliers are present), OLS uses
squares for the following reasons:
Quadratic Loss Function (Hàm tổn thất bậc 2): Squaring the distances
penalizes large errors disproportionately more than small errors. This
implies that having large outliers is considered much more serious than
having small deviations.
Derivation of Estimators: The OLS parameter estimates are derived by
differentiating the Residual Sum of Squares function (L) with respect to the
coefficients (α and β) and setting the first derivatives to zero. This
mathematical derivation relies on the properties of the squared function
Các tham số ước lượng của OLS được tìm ra bằng cách đạo hàm hàm Tổng
bình phương sai số (L) theo các hệ số (α và β) và cho các đạo hàm bậc nhất
bằng không. Phép toán này dựa trên các đặc tính của hàm bình phương.
Question 2: Explain, with the use of equations, the difference between the sample
regression function and the population regression function
The equation for the PRF is: yt=α+βxt+ut
Where:
• α and β are the true population parameters (coefficients) which are generally unknown
fixed constants.
• ut is the disturbance (error) term. It represents the unobservable random factors that affect y
but are not included in the model
The equation for the fitted SRF is: y^t=α^+β^xt
Alternatively, when expressing the actual observed value of y using the SRF, the equation is: yt
=α^+β^xt+u^t
Where:
• α^ and β^ are the estimates of the true parameters (α and β). These are numerical values
calculated from the sample data.
• y^t is the fitted value (the value predicted by the model).
• u^t is the residual. This is the observable counterpart to the disturbance term (ut). It is the
difference between the actual value (yt) and the fitted value (y^t)
Summary of Differences
• Parameters: The PRF uses true parameters (α,β), while the SRF uses estimates (α^,β^)
derived from sample data.
• Errors: The PRF contains the unobservable disturbance term (ut), whereas the SRF contains
the calculated residual (u^t),.
• Purpose: The SRF is used to infer the nature of the PRF. While the PRF represents the true
DGP, the SRF is the best "guess" available based on the sample
Question 3: What is an estimator? Is the OLS estimator superior to all other
estimators? Why or why not?
What is an estimator?
An estimator (Biến ước lượng) is the formula or rule used to calculate the coefficients (hệ số)
of a model. It is distinct from an "estimate (Giá trị ước lượng)" which is the actual numerical
value for the coefficients obtained from a specific sample of data. For example, the mathematical
expression β^=∑(xt−xˉ)2∑(xt−xˉ)(yt−yˉ) is the OLS estimator for the slope, whereas the resulting
number (e.g., 0.5091) is the estimate.
Is the OLS estimator superior to all other estimators?
No, the Ordinary Least Squares (OLS) estimator is not superior to all other estimators under all
circumstances. It is considered the optimal estimator only within a specific class of estimators
and under a specific set of assumptions. OLS is the Best Linear Unbiased Estimator (BLUE)
(Ước lượng Tuyến tính Không chênh lệch Tốt nhất), meaning it is superior only to other
linear and unbiased estimators, provided the assumptions of the Classical Linear Regression
Model (CLRM) hold.
Why or why not?
The superiority of OLS depends entirely on whether the underlying assumptions of the model are
met and what class of estimators is being considered:
1. The Gauss-Markov Theorem (When OLS is "Best") OLS is considered "Best" (efficient)
because, according to the Gauss-Markov theorem, it has the minimum variance among all linear
unbiased estimators, provided that the CLRM assumptions hold. These assumptions include:
• The errors have a mean of zero.
• The variance of the errors is constant (homoscedasticity).
• The errors are statistically independent of one another (no autocorrelation).
• The regressors are not correlated with the error term.
2. Violations of Efficiency (Inefficiency): If the assumptions of homoscedasticity or no
autocorrelation are violated, OLS is no longer the superior estimator:
• Heteroscedasticity and Autocorrelation: If these issues are present, OLS estimates remain
unbiased, but they are no longer BLUE (they become inefficient),. In these cases, estimators like
Generalised Least Squares (GLS) or the use of robust standard errors (like Newey-West) may be
superior because OLS standard errors will be incorrect, leading to misleading inferences,.
3. Violations of Consistency and Unbiasedness (Bias): In certain situations, OLS produces
biased and inconsistent estimates, making other estimators superior:
• Endogeneity/Simultaneity: If the explanatory variables are correlated (tương quan) with the
error term (sai số) (e.g., due to reverse causality (quan hệ nhân quả ngược) or omitted (bỏ sót)
variables), OLS is biased and inconsistent,. In such simultaneous equation systems, techniques
like Two-Stage Least Squares (2SLS) or Instrumental Variables (IV) are required and are
superior to OLS,.
• Measurement Error: If the independent variables contain measurement error, OLS estimates are
biased and inconsistent (attenuation bias), making IV estimators potentially preferable.
4. Non-Linearity: OLS minimises the residual sum of squares (Tổng bình phương phần dư
(RSS)), which is appropriate for linear models. However, for non-linear models (such as
ARCH/GARCH models for volatility or Logit/Probit models for binary choices), OLS is not
appropriate,. In these cases, Maximum Likelihood (ML) estimation is the superior (and often
necessary) technique because it finds parameter values that maximise the probability of
observing the sample data
Question 4: What five assumptions are usually made about the unobservable error terms
in the classical linear regression model (CLRM)? Briefly explain the meaning
of each. Why are these assumptions made?
Classical linear regression model (CLRM): yt=α+βxt+ut.
1. E(ut)=0
- Meaning: The errors have a zero mean. On average, the disturbance term is zero, meaning the
regression line passes through the mean of the data points.
2. var(ut)=σ2 < ∞
◦ Meaning: The variance of the errors is constant and finite (hữu hạn) over all values of xt.
This property is known as homoscedasticity (phương sai không đổi). It implies that the spread
of the data points around the regression line remains the same regardless of the value of the
independent variable
Variance: phương sai
3. cov(ui,uj)=0 for i=j
◦ Meaning: The errors are linearly independent of one another. This means there is no
pattern in the errors over time (no autocorrelation); an error in one period is not correlated
with the error in a previous period.
4. cov(ut,xt)=0
◦ Meaning: There is no relationship between the error term and the corresponding x
variable (independent variable). This expresses that the independent variables are exogenous
5. ut∼N(0,σ2)
or non-stochastic.
◦ Meaning: The disturbances are normally distributed.
Why Are These Assumptions Made?
These assumptions are critical for establishing the properties of the Ordinary Least Squares
(OLS) estimators and for conducting statistical inference:
• To ensure OLS is BLUE: If assumptions 1 through 4 hold, the OLS estimators (denoted
α^ and β^) are determined to be BLUE (Best Linear Unbiased Estimators).
◦ Unbiasedness: Assumptions 1 and 4 are sufficient to prove that the OLS estimates are
consistent and unbiased (E(β^)=β),.
◦ Efficiency ("Best"): Assumptions 2 (homoscedasticity) and 3 (no autocorrelation) are
required to show that OLS estimators have the minimum variance among all linear
unbiased estimators (the Gauss-Markov theorem). If these are violated, OLS estimates are
still unbiased but are no longer efficient, meaning standard errors could be wrong,.
• To enable Hypothesis Testing: Assumption 5 (Normality) is required to make valid
inferences about the population parameters from the sample parameters. Without the
assumption that errors are normally distributed, we cannot validly use t-ratios or F-tests for
hypothesis testing, particularly in small samples
Question 5:
Answer:
3.39. Can it be estimated by OLS? Yes. • Reason: This is the standard simple linear
regression model. It is linear in the parameters α and β and linear in the variables. No
rearrangement is necessary
3.40 • Can it be estimated by OLS? Yes (after rearrangement).
• Reason: This is an exponential regression model. As written, it is non-linear in parameters.
However, by taking the natural logarithm (ln) of both sides, it becomes linear:
This "double log" form is linear in the parameters α and β and can be estimated using OLS
3.41
Can it be estimated by OLS? No.
• Reason: This model is not linear in the parameters because two parameters, β and γ, are
multiplied together (βγ). OLS can estimate the combined slope (let's call it δ=βγ), but it
cannot separate β from γ to provide unique estimates for each
3.42
Can it be estimated by OLS? Yes.
• Reason: This model is already linear in the parameters α and β. The fact that the variables y
and x are in logarithmic form does not violate the OLS assumption of linearity in parameters
3.43
• Can it be estimated by OLS? Yes.
• Reason: This model is linear in the parameters α and β. The term xtzt is simply the product
of two variables. You can create a new variable (e.g., wt=xt×zt) and regress yt on wt. OLS
handles transformed or interaction variables easily as long as the parameter β itself is linea
Question 6: