0% found this document useful (0 votes)
7 views17 pages

Ecotrix Group Project

The document outlines a group project focused on analyzing the relationship between GDP and its influencing factors such as exports, imports, and per capita income using regression analysis. It discusses the methodologies of linear and multiple regression models, their assumptions, and includes a literature review of past studies on the topic. Additionally, it presents data from 1960 to 2022 related to these economic indicators in India.

Uploaded by

munjalgunjan5
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views17 pages

Ecotrix Group Project

The document outlines a group project focused on analyzing the relationship between GDP and its influencing factors such as exports, imports, and per capita income using regression analysis. It discusses the methodologies of linear and multiple regression models, their assumptions, and includes a literature review of past studies on the topic. Additionally, it presents data from 1960 to 2022 related to these economic indicators in India.

Uploaded by

munjalgunjan5
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Econometrics

GROUP PROJECT

Mentored by : amit sir

Team members :
1)
INTRODUCTION
To analyse whether there is a relation between the GDP and the factors
affecting GDP. There were several factors affecting GDP some of them are
exports, per capita income and import.
1. Exports: The value of goods and services exported from the country or
region.
2. Imports: The value of goods and services imported into the country or
region.
3. GDP (in ruppee): The Gross Domestic Product of the country or region,
measured in the local currency (ruppee).
4. Per Capita Income: The average income per person in the country or
region.
The data spans multiple years, with each row representing the values of these
economic indicators for a specific year. This type of cross sectional data is
commonly used in regression analysis to investigate the relationships between
different economic variables and their trends over time.
Regression analysis is a statistical technique used to model the relationship
between a dependent variable (GDP) and one or more independent variables
(e.g., Per Capita , Exports, Imports). It can help identify the influence of each
independent variable on the dependent variable, predict future values of the
dependent variable based on the independent variables, and understand the
overall trends and patterns in the data.

Linear Regression MODEL


linear regression is a statistical method that models the relationship between a
dependent variable (the variable to be predicted or explained) and one or more
independent variables (predictor variables) using a linear equation of the form:
Y = β0 + β1X1 + ε
Where:
- Y is the dependent variable
- X1 is a independent variable
- β0 is the intercept (the value of Y when all independent variables are zero)
- β1 is a coefficients representing the change in Y for a one-unit change in the
corresponding independent variable
- ε is the random error term
The goal of linear regression is to estimate the values of the coefficients (β0, β1)
that best fit the observed data, minimizing the sum of squared differences
between the predicted and observed values of the dependent variable.
Linear regression provides a straightforward way to quantify the relationship
between variables and make predictions based on the estimated linear model.

Multiple Regression Model


A multiple regression model is a statistical technique used to analyze the
relationship between a dependent variable and two or more independent
variables. It extends the simple linear regression model, which examines the
relationship between a dependent variable and a single independent variable, to
accommodate multiple predictors simultaneously.
Assumptions of the multiple regression model include linearity (the relationship
between the dependent and independent variables is linear), independence of
errors, homoscedasticity (constant variance of errors), and normality of errors.
Violations of these assumptions may affect the validity of the model and its
predictions.
Y = β0 + β1X1 + β2X2 + ... + βnXn + ε
Where:
- Y is the dependent variable
- X1, X2, ..., Xn are the independent variables
- β0 is the intercept (the value of Y when all independent variables are zero)
- β1, β2, ..., βn are the coefficients representing the change in Y for a one-unit
change in the corresponding independent variable
- ε is the random error term

Variables
GDP (Y) 🡪 Gross Domestic Product (GDP) is a key indicator used to
measure the economic performance of a country or region. It represents the total
value of all goods and services produced within the borders of a country within
a specific time period, typically annually or quarterly. GDP is often used as a
broad measure of a country's economic health and is widely tracked by
policymakers, economists, investors, and the general public
Per capita income (X1i) 🡪 Gross Domestic Product (GDP) is a key
indicator used to measure the economic performance of a country or region. It
represents the total value of all goods and services produced within the borders
of a country within a specific time period, typically annually or quarterly. GDP
is often used as a broad measure of a country's economic health and is widely
tracked by policymakers, economists, investors, and the general public
Exports(X2i) 🡪 Exports refer to goods and services produced domestically
within a country and sold to buyers in other countries. They are an essential
component of international trade and play a crucial role in a country's economy.
Exports generate revenue for domestic producers, contribute to economic
growth, create jobs, and can improve a country's balance of trade.
IMPORTS(X3i) 🡪 Imports refer to goods and services purchased from
foreign countries and brought into a country's domestic market. They are a
crucial component of international trade and play a significant role in a
country's economy by providing access to goods and services that may not be
domestically produced or available at competitive prices.

METHODOLOGY
OLS models assume that the analysis is fitting a model of a relationship
between one or more explanatory variables and a continuous or at least interval
outcome that minimizes that sum of square errors, where an error is that
difference between the actual and the predicted value of the outcome variable.
That most common analytical method that utilizes OLS models in linear
regression (with a single or multiple predictor variables).
The importance of OLS assumption cannot be overemphasized. Following are
the assumption of the OLS method:
1. The linear regression model is “linear in parameter.” And correctly specified.
2. The values of Explanatory variables are stochastic.
3. Given Values of Explanatory variables mean of the error term is 0. 4.
Existence oh homoscedasticity.
5. There is no multi co linearity (or perfect co linearity).
6. Number of observations should be more than number of explanatory
variables
7. Error terms should be normally distributed
LITERATURE REVIEW
Sachin N. Mehta, (2015). This study looks at the relationship between Gross
Domestic Product (GDP), Exports and imports in India using the time series
data of 1976 to 2014 using ADF unit root tests, Johansen cointegration and
vector error correction techniques. The results of the above study can be
interpreted as follows: The ADF unit root test shows that the series of GDP,
Exports and imports becomes stationary when the first difference is taken into
account. The empirical result shows that there is a long-term co-integration
relationship between Gross Domestic Products (GDP), exports and imports in
India. We found a unidirectional causal relationship between GDP and
imports, which means that in long term, GDP leads to export but not to GDP.
We also found a causal relationship between imports and exports, which
means that imports lead to GDP but not to exports.
Pradhan (2010) from the Indian perspective described that there is long run
stability between the exports, financial development and the economic growth
and the future expectations are also Correlated to the situation. Akram, khan,
Atif & shafique (2011) shared the same idea in the perspective of Canadian
economy stating the positive association between export and economic growth.
Below is the table, Which enlightens the relationship between export and
economic growth. Most of the researchers conclude a Positive relation, while
some have the negative linkage between both of these.
DATA
GDP(in Per Capita Export Imports
Year ruppee) in cr. Income in cr. s in cr. in cr.
1960 307348 0.000689 6.42E+02 961
1961 325629.2 0.000714 6.60E+02 1122
1962 349940.3 0.000749 6.85E+02 1090
1963 401902 0.000841 7.93E+02 1131
1964 468786.4 0.000959 8.16E+02 1223
1965 494305.3 0.000988 8.10E+02 1349
1966 380683.3 0.000745 1.16E+03 1409
1967 416120 0.000797 1.20E+03 2078
1968 440609.3 0.000826 1.36E+03 2008
1969 485118.4 0.00089 1.41E+03 1909
1970 518106.6 0.000929 1.54E+03 1582
1971 559016.7 0.000981 1.61E+03 1634
1972 593144.5 0.001018 1.97E+03 1825
1973 709776.7 0.001191 2.52E+03 1867
1974 826065 0.001355 3.33E+03 2955
1975 817324.2 0.001311 4.04E+03 4519
1976 852552.5 0.001337 5.14E+03 5265
1977 1008339 0.001547 5.41E+03 5074
1978 1139592 0.00171 5.73E+03 6020
1979 1269831 0.001864 6.42E+03 6811
1980 1546500 0.002219 6.71E+03 9143
1981 1605978 2.25E-03 7.81E+03 12549
1982 1665940 2.28E-03 8.80E+03 13608
1983 1811576 2.43E-03 9.77E+03 14293
1984 1.76E+06 2.31E-03 1.17E+04 15831
1985 1.93E+06 2.47E-03 1.09E+04 17134
1986 2.07E+06 2.59E-03 1.25E+04 19658
1987 2.32E+06 2.84E-03 1.57E+04 20096
1988 2.46E+06 2.95E-03 2.02E+04 22244
1989 2.46E+06 2.88E-03 2.77E+04 28235
1990 2.66E+06 3.06E-03 3.26E+04 35328
1991 2.24E+06 2.52E-03 4.40E+04 43198
1992 2.39E+06 2.64E-03 5.37E+04 47851
1993 2.32E+06 2.50E-03 6.98E+04 63375
1994 2.72E+06 2.87E-03 8.27E+04 73101
1995 2.99E+06 3.10E-03 1.06E+05 89971
1996 3.26E+06 3.32E-03 1.19E+05 122678
1997 3.45E+06 3.44E-03 1.30E+05 138920
1998 3.50E+06 3.42E-03 1.40E+05 154176
1999 3.81E+06 3.66E-03 1.59E+05 178332
2000 3.89E+06 3.67E-03 2.01E+05 215529
2001 4.03E+06 3.73E-03 2.09E+05 228307
2002 4.27E+06 3.89E-03 2.55E+05 245200
2003 5.04E+06 4.51E-03 2.93E+05 297206
2004 5.89E+06 5.18E-03 3.75E+05 359108
2005 6.81E+06 5.90E-03 4.56E+05 501065
2006 7.80E+06 6.66E-03 5.72E+05 660409
2007 1.01E+07 8.49E-03 6.56E+05 840506
2008 9.95E+06 8.25E-03 8.41E+05 1012312
2009 1.11E+07 9.10E-03 8.46E+05 1374436
2010 1.39E+07 1.12E-02 1.14E+06 1363736
2011 1.51E+07 1.20E-02 1.47E+06 1683467
2012 1.52E+07 1.19E-02 1.63E+06 2345463
2013 1.54E+07 1.19E-02 1.91E+06 2669162
2014 1.69E+07 1.29E-02 1.90E+06 2715434
2015 1.75E+07 1.32E-02 1.72E+06 2737087
2016 1.90E+07 1.42E-02 1.85E+06 2490306
2017 2.20E+07 1.63E-02 1.96E+06 2577675
2018 2.24E+07 1.64E-02 2.31E+06 3001033
2019 2.35E+07 1.70E-02 2.22E+06 3594675
2020 2.22E+07 1.59E-02 2.16E+06 3360954
2021 2.61E+07 1.86E-02 3.15E+06 2915958
2022 2.84E+07 2.00E-02 2.27E+06 4572775
RESULTS
> data=[Link]("[Link]")
> head(data)
Exports imports [Link]. [Link]
1 642 961 307348.0 0.000689191
2 660 1122 325629.2 0.000713549
3 685 1090 349940.3 0.000749298
4 793 1131 401902.0 0.000840916
5 816 1223 468786.4 0.000958547
6 810 1349 494305.3 0.000988385
> # data of India's GDP, Per Capita income, Import, Export from 1960-2022
> # all of these values are in crore
> plot(data$[Link],data$[Link].)

#the above is the scatter plot of the GDP(on the y-axis) and per capita
income(on the x-axis) of India from 1960-2022

MODEL 1 ( 2 variable model)


> model1=lm(formula([Link].~[Link]),data=data)
> model1

Call:
lm(formula = formula([Link]. ~ [Link]), data = data)

Coefficients:
(Intercept) [Link]
-1281513 1428199076

#we have run the codes for the basic two variable regression. Where GDP is our
dependent variable and per capita income as independent variable
here our β 1=−1281513 ( the intercept parameter )
β 2=1428199076(the slope parameter )

> abline(model1,col="blue")

# we are plotting our regressed model of the scatter plot to check the fit of the
model. For almost each point our regression line is overlapping
Which tells us that our model is a good fit
> summary(model1)

Call:
lm(formula = formula([Link]. ~ [Link]), data = data)

Residuals:
Min 1Q Median 3Q Max
-814316 -321047 -77548 372756 1117532

Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) -1281513 75953 -16.87 <2e-16 ***
[Link] 1428199076 10199780 140.02 <2e-16 ***
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Residual standard error: 430900 on 61 degrees of freedom


Multiple R-squared: 0.9969, Adjusted R-squared: 0.9968
F-statistic: 1.961e+04 on 1 and 61 DF, p-value: < 2.2e-16

> anova(model1)
Analysis of Variance Table

Response: [Link].
Df Sum Sq Mean Sq F value Pr(>F)
[Link] 1 3.6410e+15 3.6410e+15 19606 < 2.2e-16 ***
Residuals 61 1.1328e+13 1.8571e+11
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

#The Pr(>|t|) tells us about the pvalue of the respective parameter


# for testing whether our beta2 is significant
# H0 🡪 b2= 0; H1 🡪 b2 not= 0
#Here our beta parameter is having pvalue less than 5% which means it
is a significant parameter

# we are having a very high R^2 which is almost near to 1 which means
that the model is a good fit to the data
# when we run the anova we got to know about the F-statistics and when
observe its pvalue which is resulting less than 5% which determines that
our model is containing significant parameters

MODEL 2 (multi variable model)

>
model2=lm(data$[Link].~data$[Link]+data$imports+data$
Exports)
> summary(model2)

Call:
lm(formula = data$[Link]. ~ data$[Link] + data$imports +
data$Exports)

Residuals:
Min 1Q Median 3Q Max
-1109587 -165142 44623 165230 826416

Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) -7.439e+05 9.322e+04 -7.980 5.91e-11 ***
data$[Link] 1.145e+09 4.051e+07 28.276 < 2e-16 ***
data$imports 7.878e-01 1.575e-01 5.003 5.39e-06 ***
data$Exports 7.960e-01 2.587e-01 3.077 0.00316 **
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Residual standard error: 316300 on 59 degrees of freedom


Multiple R-squared: 0.9984, Adjusted R-squared: 0.9983
F-statistic: 1.215e+04 on 3 and 59 DF, p-value: < 2.2e-16

> anova(model2)
Analysis of Variance Table

Response: data$[Link].
Df Sum Sq Mean Sq F value Pr(>F)
data$[Link] 1 3.6410e+15 3.6410e+15 36391.5452 < 2.2e-16 ***
data$imports 1 4.4775e+12 4.4775e+12 44.7523 9.047e-09 ***
data$Exports 1 9.4754e+11 9.4754e+11 9.4706 0.003165 **
Residuals 59 5.9030e+12 1.0005e+11
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

# Model 2 by taking exports and imports as an additional parameter to


our previous model
# Our parameter values are coming out as
¿ β 1=−7.439e+05 ; β 2=1.145e+09 ; β 3=7.878e-01;

β 4 =7.960e-01

# for parameter testing we need to know about there pvalue and here all the
variables is having a pvalue less than 5%
# all of the parameters are having a pvalue <5% which means all of them are
good fit for the model.
# adjusted R^2 have increased slightly which makes this the best model in the
above 3 models.
# F-statistics’s pvalue is coming out as 2.2e-16 which is less than 5% which
makes our parameter significant.
Using the best model to do further analysis i.e. Model 2

MULTICOLLINEARITY
Multicollinearity in regression analysis occurs when two or more explanatory
variables are highly correlated to each other, such that they do not provide
unique or independent information in the regression model. If the degree of
correlation is high enough between variables, it can cause problems when
fitting and interpreting the regression model. Fortunately, it’s possible to
detect multicollinearity using a metric known as the variance inflation factor
(VIF), which measures the correlation and strength of correlation between the
explanatory variables in a regression model.

> vif(model3)
data$[Link] data$imports data$Exports
29.27602 20.62129 27.22797

If we are having VIF>5 which means that high multicollinearity is there


So, all the independent variable is having a VIF >5 which implies that there is
multicollinearity present in our model

Heteroscedasticity
Heteroscedasticity refers to the situation where the variability of the errors
(residuals) in a regression model is not constant across all levels of the
independent variables. In other words, the spread of the residuals changes as
the value of the independent variables change. Heteroscedasticity violates one
of the assumptions of linear regression, which assumes that the variance of the
errors is constant (homoscedasticity).
> bptest(model3)

> bptest(model3,studentize = F)

Breusch-Pagan test

data: model3
BP = 80.106, df = 3, p-value < 2.2e-16

We have done white’s test for whether there is heteroscedasticity present or


not
H0 🡪 Heteroscedasticity is not existing in our model
H1 🡪 Heteroscedasticity is existing in our model
So as our pvalue is less than 5% which means that we have sufficient evidence
to reject H0
So there is heteroscedasticity present in our model

AUTOCORRELATION
Autocorrelation, also known as serial correlation, is a statistical phenomenon
where the values of a variable in a time series or a regression model are
correlated with their past or future values. In simpler terms, autocorrelation
occurs when there is a pattern of correlation between successive observations
or residuals within a dataset.
H0 : no positive autocorrelation
H1 : positive autocorrelation
> durbinWatsonTest(model3)
lag Autocorrelation D-W Statistic p-value
1 0.811138 0.2505243 0
Alternative hypothesis: rho != 0
We are having pvalue as 0 which means that is less than 5% which tells
us that we have sufficient evidence to reject H0
CONCLUSION
In our project we have shown whether there is a significant role of per Capital
Income, exports and imports in GDP or not. We have taken the data of more
than past years.
We have made 2 models under our data. 1 st one is a 2 variable linear
regression model
And the 2nd one is a multivariable regression model. Under both the models all
the β ' s are playing a significant role. Which means that all these factors are
significantly affecting our GDP.
We are having the adjusted R2 under the 1st model as 0.9968 which is a very
high value we are getting a almost perfect correlation between the 2 variables
and
under the 2nd model our adjusted R2 is resulting as 0.9983 which is even more
closer to a perfect correlation than model 1
and as we know the higher the adjusted R^2 the better the model is. So, here
model 2 is a better model than model 1
With the help of anova table we can conclude that
Under 1st test our pvalue is less than 2e-16 which is less than 5% which tells us
that overall our model is significant
Under 2nd test too our pvalue is less than 2e-16 which is less than 5% which
shows that our model is overall significant
We can see the presence of multicollinearity. As our VIF is having a very high
value
There is also heteroscedasticity in our model. We are saying this because we
have conducted white’s test to check whether there is heteroscedasticity
present in our model or not
There is positive autocorrelation also in the model
REFERENCES
[Link]
[Link]

You might also like