0% found this document useful (0 votes)
19 views6 pages

Understanding Simple Linear Regression

Chapter Two discusses the concepts of covariance, correlation, and regression, focusing on simple linear regression (SLR) as a method to model relationships between variables. It explains the significance of the disturbance term in regression models and outlines key assumptions of the classical linear regression model. The chapter emphasizes the importance of understanding the relationships between dependent and independent variables for accurate predictions and analyses.

Uploaded by

yonasbekele05
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views6 pages

Understanding Simple Linear Regression

Chapter Two discusses the concepts of covariance, correlation, and regression, focusing on simple linear regression (SLR) as a method to model relationships between variables. It explains the significance of the disturbance term in regression models and outlines key assumptions of the classical linear regression model. The chapter emphasizes the importance of understanding the relationships between dependent and independent variables for accurate predictions and analyses.

Uploaded by

yonasbekele05
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter Two

Simple linear regression function


# What is SLR?

 Covariance
is a statistical measure that indicates the extent to which two random variables change
together. Specifically, it measures whether increases in one variable correspond with
increases or decreases in another variable. Here‟s a breakdown of the concept:

Positive Covariance: If the covariance between two variables is positive, it means that as
one variable increases, the other variable tends to increase as well. Conversely, when one
variable decreases, the other tends to decrease.

Negative Covariance: If the covariance is negative, it indicates that as one variable


increases, the other tends to decrease, and vice versa.

 Correlation
is a statistical measure that describes the strength and direction of a linear relationship
between two random variables. Unlike covariance, correlation is standardized, meaning its
value ranges from -1 to 1, making it easier to interpret. Here are the key points about
correlation:

Types of Correlation

1. Positive Correlation: A positive correlation (close to +1) indicates that as one variable
increases, the other variable also tends to increase. For example, height and weight might
have a positive correlation in a population.

2. Negative Correlation: A negative correlation (close to -1) indicates that as one variable
increases, the other tends to decrease. For example, the amount of time spent watching TV
and academic performance might have a negative correlation in students.

3. No Correlation: a correlation close to 0 indicates any linear relationship between the


variables. For example, the number of hours spent studying and the number of movies
watched in a month might show no correlation.

Interpretation of the Correlation Coefficient

 +1: Perfect positive linear relationship.


 0: No linear relationship.
 -1: Perfect negative linear relationship.
 Regression
is a statistical technique used to model and analyze the relationships between a dependent
variable and one or more independent variables. The primary goal of regression analysis is
to understand the nature of the relationship between variables, make predictions, and infer
causal relationships. Here are the key concepts and types of regression:

Key Concepts
1. Dependent Variable: The variable you are trying to predict or explain, often denoted as Y.
2. Independent Variable(s): The variable(s) used to make predictions or explain the dependent
variable, often denoted as X (or X1,X2,…,Xn in the case of multiple predictors).
3. Regression Coefficient: Represents the change in the dependent variable for a one-unit change in an
independent variable, holding other variables constant.
 Simple linear regression function
Simple linear regression function is the simplest form of a regression analysis having a
single explanatory variable related in linear form.
# Simple: SLR consists only two variables (one dependent and one independent variable)

- If the number of independent or explanatory variables is greater than one, we don‟t


use the term „simple‟. Instead we use the term „multiple‟.

- Example: …………………………….Simple

- ……..multiple
# Linear
- linear” regression will always mean a regression that is linear in the parameters;
the β’s (that is, the parameters are raised to the first power only).
………linear

…….Non linear

Example: - stock price


Financial theory indicates there is a relationship between stock price and dividend per
[Link] dependent variable is stock price and the independent variable is dividend per
share. So

Where:
- : stock price - : Dividend per share
# General form of SLR

𝒀 = 𝜷 𝜷 + 𝒖
Variation in Y = Systematic variation + Random variation

Variation in Y = Deterministic variation + Non stochastic variation

Variation in Y = Explained variation + Unexplained variation

Where,
Y is dependent variable 𝒖 is error (disturbance) term
X is independent variable is number of cases or observations

𝜷 &𝜷 regression coefficients


# THE SIGNIFICANCE OF THE STOCHASTIC DISTURBANCE TERM

- The disturbance term ui is a surrogate for all those variables that are omitted from the model but
that collectively affect Y.
- Error term is a term added to a regression model to capture all the variation in Y that can‟t be
explained by Xs.

 It is a proxy for all variables that are not included in the regression model, but may
collectively affect Y.

 Why we include 𝒖 In the model?


1. It captures the effect of omitted variables.

a model is a simplification of reality. It is not always possible to include all relevant variables in a
functional form. For instance, we may construct a model relating ROE and Capital structure. But
ROE is influenced not only by Capital structure: operating leverage, size of assets and several
other variables also influence it. The omission of these variables from the model introduces an
error. In addition we may omit variable due to the following factors.

 Lack of data and limited knowledge: we may not have information about variables

 Difficulty in measuring some factors

 Poor proxy variables


2. Vagueness of theory: The theory, if any, determining the behavior of Y may be, and often is,
incomplete. We might be ignorant or unsure about the other variables affecting Y.

3. Principle of parsimony: we would like to keep our regression model as simple as possible. If we
can explain the behavior of Y “substantially” with two or three explanatory variables and if our
theory is not strong enough to suggest what other variables might be included, why introduce
more variables? Let ui represent all other variables.

4. Errors of Measurement: errors of measurements of variables which are inevitable due to the
method of collecting and processing statistical information.
# Assumptions of Classical Linear Regression Model

The classicals‟ made important assumption in their analysis of regression .The most imporntant of
these assumptions are discussed below.

 Assumption 1: The model is linear in parameters.

The classicals assumed that the model should be linear in the parameters regardless of whether the
explanatory and the dependent variables are linear or not. This is because if the parameters are non-
linear it is difficult to estimate them since their value is not known.

 Assumption 2: The mean value of the random variable(U) in any particular period is
zero

= =
This means that for each value of x, the random variable(u) may assume various values, some
greater than zero and some smaller than zero, but if we considered all the possible and negative
values of u, for any given value of X, they would have on average value equal to zero. In other
words the positive and negative values of u cancel each other.

 Assumption 3: The variance of the random variable(U) is constant in each


period (The assumption of homoscedasticity)

For all values of X, the u‟s will show the same dispersion around their mean or error terms should be
homoscedastic. Homoscedasticity describes a situation in which the error term (that is, the “noise”
or random disturbance in the relationship between the independent variables and the dependent
variable) is the same across all values of the independent variables.

 [ ] …………

 Put simply, the variation around the regression line (which is the line of average relationship
between Y and X) is the same across the X values; it neither increases or decreases as X varies.
 If this condition is not fulfilled or if the variance of the error terms varies as sample size
changes or as the value of explanatory variables changes, then this leads to
Heteroscedasticity problem.

 Assumption 4:The random variable (U) has a normal distribution

This means the values of u (for each x) have a bell shaped symmetrical distribution about their zero
mean and constant variance , i.e.

 Assumption 5: The random terms of different observations ( )are


independent.(The assumption of no autocorrelation)

This means the value which the random term assumed in one period does not depend on the
value which it assumed in any other period.

 ( )

 Error terms of different observations are independent.

 The error terms across observations are NOT correlated with each other

 The error term in one time period never affects the error term in the next.

- Covariance and correlation measure the relationship and the dependency between two
variables. Covariance indicates the direction of the linear relationship between variables
while correlation measures both the strength and direction of the linear relationship between
two variables.
 Assumption 6: The explanatory variable Xi is fixed in repeated samples.
Each value of Xi does not vary for instance owing to change in sample size. This means the
explanatory variables are non- random.
 Assumption 7: zero covariance between 𝒖 (No autocorrelation)

 𝒖
 No correlation between regressors and error terms.

 𝒖 Assumed to have separate and additive effect on Y.

 Assumption 8: X-values in a given sample must not be the same (within a sample)

 Var (X) must be a finite positive number.

 If ̅ , it is impossible to estimate the parameters.

 Assumption9: Randomness of 𝒖 :

The error terms „ ‟ are randomly distributed. is a random real variable. The value which
may assume in any period depends on chance: some may be positive, some may be negative or
some may be zero.
 Assumption 10: Explanatory variables should not be perfectly, linearly and/or
highly correlated.

Using explanatory variables which are highly or perfectly correlated in a regression function causes
a biased function or model. It also results in multicollinearity problem.

 Assumption 11:The variables are measured without error (the data are error free).

Since wrong data leads to wrong conclusion, it is important to make sure that our data is free from
any type of error.

 Assumption 12:No model specification error: the model is correctly specified


 Assumption 13:The number of observations must be greater than the number of
explanatory variables

Common questions

Powered by AI

Incorrectly specifying a regression model can have severe consequences, including biased and inconsistent estimates, misleading inference, and poor predictive performance. This might occur due to omitting relevant variables, including irrelevant ones, or using incorrect functional forms. To avoid such errors, modelers should ensure a thorough understanding of the theoretical framework and the system being studied. They should conduct specification tests, consider appropriate transformation of variables, and validate model assumptions thoroughly. Engaging in exploratory data analysis can also help detect specification errors .

Including multiple highly correlated explanatory variables in a regression model leads to multicollinearity, which is problematic because it undermines the ability to determine the independent effect of each variable on the dependent variable. Multicollinearity inflates the standard errors of the coefficients, which reduces the precision of the estimated parameters and makes them very sensitive to changes in the model. This can lead to unreliable estimates and make it difficult to ascertain the true relationship between variables, hindering the interpretability and predictive power of the regression model .

Autocorrelation, which occurs when the residuals (error terms) are correlated across observations, can significantly affect the validity of a regression model's estimates. It violates the assumption that error terms are independent, leading to potentially biased standard error estimates. Consequently, hypothesis tests about the relationship between variables might become invalid, as the presence of autocorrelation inflates the t-statistics, giving a false impression of significance. It can also lead to inefficient estimates and suboptimal predictions of the dependent variable, undermining the model's reliability .

The assumption of homoscedasticity is fundamental in regression analysis because it stipulates that the variance of the error terms should be constant across all levels of the independent variable(s). If homoscedasticity holds, the regression model delivers reliable and consistent estimates. When this assumption is violated, leading to heteroscedasticity, the estimated standard errors might be biased, resulting in unreliable statistical tests and confidence intervals. This makes it difficult to assess the significance of predictor variables, potentially leading to incorrect conclusions about relationships between variables .

The assumption that a regression model is linear in its parameters affects the estimation process by ensuring that the model parameters can be estimated using linear algebra techniques. If the parameters are non-linear, the estimation becomes complex and often requires iterative, non-linear optimization techniques which could be computationally intensive and less robust. A linear parameter model allows for straightforward estimation, hypothesis testing, and interpretation, making the regression analysis more practical and reliable in capturing linear relationships and providing valid inferential statistics .

Covariance and correlation both measure the linear relationship between two variables, but they differ significantly in their properties. Covariance indicates the direction of the linear relationship—positive when both variables tend to move in the same direction, and negative when they move inversely. However, it does not provide the strength of this relationship and is not standardized, making its interpretation dependent on the units of measurement. In contrast, correlation not only shows the direction but also the strength of the relationship, and it is standardized to a range from -1 to 1, making it more interpretable regardless of units. Correlation tells us how strongly the variables are related linearly .

Measurement errors in regression variables can severely affect the validity of the analysis by introducing bias and inconsistency in the estimated coefficients. If the errors are in the explanatory variables, it leads to attenuation bias, where the estimated coefficients are systematically underestimated. Measurement errors in the dependent variable, although less problematic, can still inflate standard errors, reducing the precision of estimates. These errors make it difficult to correctly interpret the relationships being modeled, potentially leading to false conclusions about causality and effect size .

Explanatory variables must not perfectly correlate in regression analysis because perfect correlation makes it impossible to disentangle the separate effects of these variables on the dependent variable. This scenario is known as perfect multicollinearity and leads to computational issues, where the matrix needed to estimate the regression coefficients does not invert, resulting in undefined estimates. It prevents assessing each variable's impact and interpreting the regression coefficients, essentially invalidating the regression results .

The stochastic disturbance term, denoted as 'u', plays a critical role in simple linear regression models by capturing all the variation in the dependent variable Y that cannot be explained by the independent variables X. It acts as a proxy for all the omitted variables, measurement errors, vagueness in theory, and unforeseen influences that affect Y. Its inclusion is essential as it accounts for the reality that all relevant variables cannot always be included due to missing data, the principle of parsimony, and other factors. Failure to include it would mean ignoring the random or unexplained variability inherent in empirical data, which could lead to biased estimates .

The principle of parsimony guides the selection of variables in a regression model by advocating for simplicity without sacrificing explanatory power. It suggests including only those variables that provide substantial explanatory insight into the dependent variable while omitting redundant or irrelevant variables. This not only simplifies the model, making it more understandable and interpretable, but also enhances its generalizability to other datasets. Parsimony helps in avoiding overfitting and unnecessary complexity, thus leading to more robust and reliable models .

You might also like