19CSE305
Machine Learning
Assignment :
GST Prediction for a Financial
year
In partial fulfilment for the award of the
degree of
BACHELOR OF TECHNOLOGY
IN
“COMPUTER SCIENCE AND ENGINEERING”
AMRITA SCHOOL OF ENGINEERING,
BANGALORE
AMRITA VISHWA VIDYAPEETHAM
1
BENGALURU-560035
October -2024
Submitted by:
Registration ID Name
[Link].U4CSE22104 Akarsh Surya Venkat
[Link].U4CSE22114 Tejesh Chowdary
[Link].U4CSE22124 Navya Reddy
[Link].U4CSE22134 Namratha Akshaya
[Link].U4CSE22144 Nithya Shree
[Link].U4CSE22154 Soumish Ghosh
[Link].U4CSE22164 Nishanth
[Link].U4CSE22174 Keshav Padmakumar
[Link].U4CSE22184 Krishna Bhagat
2
Contents
Problem Statement............................................................................4
Introduction and Motivation.............................................................4
Literature Survey................................................................................6
Solution................................................................................................8
High-Level Plan................................................................................8
Detailed Plan....................................................................................8
Phase 1: Setup and State-level Data Collection...............................8
Phase 2: Data Processing and Model Development on a State Level
.........................................................................................................9
Phase 3: National Aggregation and Model Training........................11
Phase 4: Dashboard and Reporting................................................12
Phase 5: Continuous Monitoring and Refinement...........................13
Requirements....................................................................................14
Client-Side Requirements:.........................................................14
Server-Side Requirements:.......................................................15
Software Requirements:............................................................16
Cost Estimation.................................................................................17
Conclusion..........................................................................................18
References.........................................................................................19
3
Problem Statement
Develop an appropriate and robust Goods and Services Tax
revenue forecasting model for the coming financial year to
aid and improve fiscal planning, economic policy
formulation, and resource allocation
Introduction and Motivation
As one of the major sources of revenue to the government
the role that GST must play cannot be over emphasized in
the financial health of India. It is also regarded one of the
most crucial factors that might affect a state’s economic
status. But it is always a task to forecast it with high
accuracy because of the prevailing economic factors,
changes in policies and the marketplace factors. The current
methods sometimes prove ineffective leading to a variance
of the projected revenue collection. Such trends, in turn,
result in budgetary shortages, ineffective resource
distribution, and the presentation of inefficient policies in the
field of finance.
Today’s advancement in analytics, machine learning, and
artificial intelligence offer a new approach to predicting GST
revenue. The combined application of these tools with
overall economic pointers and past records can create a
sound and swift model of prediction.
4
The image below shows the trend of GST collection in India
over recent months:
Figure 1: Trends in GST Collection (Rs. in Crore) for FY 2022-23 and 2023-24
This chart shows the monthly GST collections for two
consecutive financial years, highlighting the variability and
growth trends that must be accounted for by any prediction
model.
These challenges explain the need for better forecasting
techniques. Here are several key factors that illustrate the
importance of enhancing GST prediction models:
Better Budget Planning: Accurate estimation of the GST
revenues assist the government in planning for the right
amounts and the right things in the best manner and using
the public funds in the most proper manner.
Improved Policy Making: Reliable forecasts of future GST
trends facilitate well-informed economic policy decisions. For
instance, expected growth in GST collections might give
policymakers confidence to increase investment in public
services.
Better Business Planning: Accurate GST predictions help
the government stay ahead of potential financial issues by
giving them a clearer picture of future revenue, which in turn
supports overall economic stability.
5
Public Transparency: The general public’s trust in the
government can also be achieved through accurate revenue
predictions as people will be assured of how their monies are
utilized.
Thus, the enhancement of GST revenue forecasting model
that is capable of expressing precise and effective results is
important for effective fiscal management, economic
stability and public confidence. The government should be
able to create the best possible budget and make the best
decisions about the policies to be implemented, as well as be
able to distribute the resources more effectively through the
recent advancement in machine learning and data analysis.
Forcing of actual GST revenues accurately would at last
deliver good governance, boosting the level of trust of the
business community as well as the public.
Literature Survey
The paper by Purnama et al. [1] where the research
conducts a comparison of the different Performance of
Support Vector Machine (SVM) and Linear Regression (LR)
algorithm for stock price prediction of PT. Vale Indonesia.
Their work shows that LR performs better in this case for
stock price prediction. Alex et al. [2] study focuses on the
quarterly data effect on the stock prices of ten Indian IT
midcap firms. This study reveals the predictive factors which
include P/E ratio, Face Value, Net Profit, and EPS to have a
positive influence on stock prices, however, Sales, Expenses,
and Operating Profit have a no impact on stock prices by
using multiple linear regression and random forest models. A
major implication of the study is that there is a need to
refine the model and the importance of features in the
accuracy of the prediction.
6
Abdi et al. [3] present a basic feature dataset for stock
market prediction, with focus and difference between time-
dependent and time-independent one. Comparing the
effectiveness of the four methods used in the study, random
forest, SVM, logistic regression and gradient boosting, the
authors note that logistic regression and gradient boosting
highlighted the highest accuracy when forecasting return
and risk, respectively. According to Vyas et al. [4], sales
forecasting is significant in boosting profit and operations
planning in businesses such as Walmart with the Business to
Consumer (B2C) model. The paper compares various
forecasting models and specifically compares how they fit
sales seasonality that is typical in the holiday period and
indicates that depending on their use case, businesses must
choose between high accuracy and model run time.
The startup capital requirement and profitability: prediction
for e-Commerce ventures using machine learning, linear
regression model by Adebiyi et al. [5] The research
evaluates past data to explain patterns associated with
these financial measures and advocates for the use of big
data to improve the resilience of e-commerce start-ups. The
results point to the need for model enhancement as well as
explore other approaches to enhance predictive accuracy
that may prove beneficial for financial decision-making for
entrepreneurs and investors. The paper of Das et al. [6]
deals with a marketing strategy based on the computation of
potential customers by applying the logistic regression on
transaction history and other demographic variables.
However, the results have highlighted the importance of the
type of model to be used based on the characteristics of the
data set used, and pointed to further research as to how one
can improve the overall accuracy of the model same as the
positive samples, but they did not consider the negatives.
To predict India’s GST collection, which is the research
question of Thayyib et al. [7] linear models like the TBATS
7
and more complex forms are hybrid ones, where the linear
and non-linear components are combined into one for more
precise results. Based on the findings of the study, the
Hybrid Theta-TBATS model claimed better specification when
compared with other neural network-based models such as
ANN and NNAR on GST revenue prediction. In their paper
Nayyar et al. [8] explain the adoption of GST in India that is
attempting to eliminate VAT, Service Tax, and Central Excise
into one easily understandable taxation method. It clears
how GST is likely to enhance efficiencies and effectiveness of
tax administration, enhance economic growth rate by
between 1-2%, cut on corruption and innovation. Advanced
data analytical work is recommended to overcome these
issues and better implement the GST system to reap the
benefits in multifarious sectors in India.
In their paper, Antad et al. [9] has identified that Linear
Regression is one of the powerful machine learning
algorithms that can be used for stock price prediction
through data analysis. This is evident where Linear
Regression models a slightly better prediction model than
Deep Learning and Neural networks even when tested
during a volatile stock market. The paper also concludes by
stating that the choice of data set is crucial to the success of
stock market forecasts. Margaret, et al. [10] have published
a paper in which the authors have sought to use the
Evaluated Linear Regression-based Machine Learning (ELR-
ML) technique in predicting stock price movement within the
Standard and Poor’s 500 (S&P 500) index. The conclusion
strengthens this argument by insisting on the efficiency of
correct stock market predictions, especially with ELR-ML as a
way of informing trading decisions that could be profitability
yielding for investors.
8
Solution
High-Level Plan
The solution focuses on the prediction of the amount of GST
at the state level, and then sums up the predicted amounts
to come up with the national prediction. The plan includes
establishing data collection systems for gathering state
level data, data preprocessing, normalization of the data,
algorithm development - in this case, individual machine
learning models for each state, integrating the state-level
forecasts to come up with a national estimate, and finally
visualizing the results through a dashboard. The approach
improves accuracy of the predictions because it takes into
account regional differences in the economic environments
and tax regimes.
Detailed Plan
Phase 1: Setup and State-level Data Collection
Objective: Establish data conduits and framework to gather
and process state level data.
Strategy:
1. Data Collection Pipeline Setup
i) Identify key data sources for each state, including
state tax departments, local government databases,
and publicly available economic indicators.
ii) Establish APIs or data feeds to pull historical GST
data and economic indicators such as state GDP,
industrial growth, inflation rates, and sector-specific
economic activity.
iii) Define data update frequency for each source to
ensure timely and consistent data collection.
9
iv) Perform data ingestion with either using Python or
any other data ingestion tools such as Apache NiFi or
Talend.
2. Infrastructure Setup
i) Decide if you want to go cloud or physical leveraging
(AWS, Azure or Google Cloud) or physical hardware
(Physical Server, Dell PowerEdge R740) depending on
the volume of data and the scalability and costs
implications.
ii) Determine arrangements for storage of data as it is
received using distributed storage techniques such as;
cloud storages such as Amazon S3 or Hadoop
Distributed File System (HDFS for physical servers).
iii) Established arrangements for collecting and
processing data (such as a data warehouse Apache
Hive in ’big data’, or the standard RDBMS such as
PostgreSQL).
3. Data Security and Compliance
i) Develop data encryption on data which are static and
dynamic.
ii) Implement the user authentication together with
various options for role-based access control.
iii) To guarantee treatment of the law on data
protection (for example; general data protection
regulation, California consumer privacy act, or the
Indian information Technology Act).
10
Phase 2: Data Processing and Model Development on a State
Level
Objective: Clean and analyze the state level data and
develop state wise machine learning model for estimating
GST collections.
Strategy:
1. Data Cleaning and Normalization
i) Handle missing data using methods like mean
imputation or forward filling.
ii) Scale the data to a standard range in order to avoid
the possibility of variation of different states across the
analysis methods which include; min max scaling and z
score scaling.
iii) Perform feature engineering to get numerous
inspired and meaningful features (e.g., sectoral
contribution, monthly/seasonal patterns,
macroeconomic indicators).
2. Exploratory Data Analysis (EDA)
i) Study the historical trend analysis of the GST amount
collected up to the current fiscal year at the state level
to draw the pattern.
ii) Analyzing correlation between different features can
also be done by plotting heatmaps, time series charts, etc.
3. Model Development for Each State
i) Select the right kind of machine learning models
according to the type of data present in our dataset (for
11
instance, tabular data will require XGBoost, while time-
series data call for LSTM).
ii) Divide data into training data set, the validation data
set and the test data set.
iii) Trim models on each state data set, using
hyperparameters to get the best predictability.
iv) Calculate accuracy parameters like RMSE –Root
mean square error, MAE –Mean Absolute Error and R
Squared to evaluate model performance.
v) Cross validation should be used to make sure the
model of selection is very stable and ideal for predicting
state-level GST collections.
4. Model Deployment for State Predictions
i) Deploy and operate the models through model
deployment frameworks (e.g., Flask, FastAPI, or
MLflow etc).
ii) Packaging the models as Docker containers requires
portability, versatility and easier scaling.
iii) Setting up automated retraining schedules to train
models periodically in the new data that may come
up from time to time.
Phase 3: National Aggregation and Model Training
Objective: Summarize the state-level predictions to arrive
at a national GST collection forecast.
Strategy:
1. Develop Aggregation Pipeline
12
i) Develop scripts or workflows to obtain predictive
information from all state models where this is applicable.
ii) Sum the predictions to develop a preliminary
estimate of national GST collection.
iii) Introduce more calculations into the system to make
state predictions more weighted by specific factors
such as the state’s GDP addition or previous GST
proportion.
2. Secondary Aggregation Model
i) Train regression models and produce better state and
national-level aggregates utilizing the state predictions
and the other measures of the national level aggregate
of the economy (i.e., overall rate of economic growth,
overall rate of inflation, etc.).
ii) Analyze the efficiency of the models used to predict
future occurrences and fine tune the parameters for
improved accuracy.
3. Aggregation Model Deployment
i) Many states require county-level information, and the
aggregation model will include these data injected into
the state prediction pipeline so that there is an efficient flow
from state predictions to national forecasts.
ii) Set up a timetable for regular updates of the
aggregation model to ensure it includes the latest data.
Phase 4: Dashboard and Reporting
Objective: To visualize GST predictions at both the state
and national levels to support better decision-making.
Strategy:
13
1. Dashboard Development
i) Select a visualization tool/ Software like Power BI,
Tableau or Web Applications using libraries like [Link],
Plotly etc
ii) Design the model in a way that the predictions for
each state should be shown in addition to the
aggregated national forecast.
iii) Integrate additional functionalities of data drilling to
highlight state level specifics, historical data and data
according to different sectors.
2. Reporting Features
i) Automate the preparation of GST report for monthly,
quarterly and yearly forecast.
ii) Provide user defined warning to inform the interested
party of changes that are of special importance with
reference to estimated values.
iii) Incorporate real-time data updates to reflect the
latest state and nationwide GST estimations.
Phase 5: Continuous Monitoring and Refinement
Objective: To Continuously enhance prediction accuracy
by regularly updating models and monitoring their
performance.
Strategy:
1. Model Monitoring
i) Implementing monitoring systems for model
performance metrics such as accuracy of predictions and
data shift.
ii) Utilize tools like Prometheus or Grafana for real-time
monitoring and alerting.
14
2. Model Refinement
i) Continuously update the state models by training with
new data, and testing state models against updated
datasets.
ii) Perform hyperparameter tuning from time to time so
that it may be ready to capture changes in the patterns
against updated datasets.
3. Feedback Loops
i) Integrate user feedback (e.g., from state tax
departments) systems to modify models or enhance the
data quality.
ii) Use an error logging file and develop a method for
checking the magnitude of frequency of errors made by
the model such that if the number of errors surpass a
certain limit, then the model re-training should commence.
15
Requirements
Client-Side Requirements:
1. User Devices:
a. Laptops/Desktops/any similar gadgets:
i. Processor: Intel i7/i9 or AMD equivalent.
ii. RAM: 16-24 GB.
iii. Storage: SSD with at least 1TB.
iv. GPU (Optional): For reducing time while
representing data in graph and other visual
formats a powerful GPU would cut down time
and be helpful
b. Smaller portables and handheld devices
(smartphones etc): For viewing the dashboard
or reports via web browsers.
i. Devices should support modern browsers
(Chrome, Firefox, Safari, edge, brave etc).
2. Internet Connection:
a. A stable internet connection is required to make
and access requests from the server
Server-Side Requirements:
1. Cloud or Physical Server:
a. when using cloud, services like AWS EC2,
Google Cloud, or Azure can handle scalable data
processing and machine learning model training:
i. Type: General purpose (e.g., AWS or Azure).
ii. CPUs: 16-32 vCPUs.
iii. RAM: 32-64 GB.
iv. Storage: 1TB SSD for fast read/write when
compared to HDD.
b. If using physical servers, specifications can
include:
16
i. Processor: Dual Intel Xeon Silver/Gold series
with 16-24 cores.
ii. RAM: 128-256 GB DDR4.
iii. Storage: At least 2 TB SSD for high-speed
data handling.
iv. GPU: NVIDIA Tesla T4 (if using deep learning
models like LSTMs).
2. Big Data Framework Hardware:
a. If using frameworks like Hadoop for parallel state-
wise data processing, additional servers with
similar configurations may be needed to ensure
performance for handling large datasets.
Software Requirements:
1. Data Processing & Machine Learning:
a. Python with libraries like:
i. pandas (for data manipulation),
ii. numpy (for numerical computations),
iii. scikit-learn (for basic machine learning
models),
iv. tensorflow/keras (for LSTM or deep learning
models, if needed),
v. statsmodels (for time series analysis).
b. Hadoop for processing large datasets and
handling state-wise data in parallel.
c. Jupyter Notebooks or PyCharm for
development and testing of machine learning
models.
2. Cloud/Infrastructure:
a. AWS or Google Cloud Storage for storing state-
level data and model artifacts.
17
b. AWS, Google Cloud Compute Engine, or Azure
Virtual Machines for model training and
deployment.
c. Docker for containerizing the model training
pipeline and ensuring consistent deployments
across environments.
3. Model Deployment and Management:
a. Kubernetes for scaling machine learning services
across multiple states.
b. DVC (Data Version Control) for managing machine
learning experiments and model versions.
4. Database and APIs:
a. MySQL for managing state-specific GST collection
data and results.
b. RESTful APIs (using Flask) for collecting data
from state tax databases and integrating state-
wise models.
5. Visualization and Reporting:
a. Power BI for dashboard development, providing
state-wise and national-level GST predictions.
These specified requirements can be taken as a minimum
threshold for collection, training the model and displaying it
from a national or state point of view, for faster results we
can make use of more expensive cloud services but keeping
in mind the minimum cost we have come up with the above
requirements.
Cost Estimation
18
CATEGORY OBJECTS MINIMUM MAXIMUM
Client-Side End-User Devices
Hardware
Laptops/Desktops of $720 $1650
intel i7/i9
Smaller Portables $300 $700
like smartphones
Total Cost of $1020 $2350
Client-Side
Hardware
Internet Connection $50 $150
(per month per
location)
Total client- $1070 $2500
Side Cost
Sever-Side Cloud Server
Hardware
General Purpose $200 $450
Server
Total Cloud $200 $450
Server-Side
Cost
Physical Server $6700 $11400
Big-Data Framework $4200 $7900
Hardware
Total Physical $10900 $19300
Server-Side
Cost
Software
Requirements
Python libraries free free
Hadoop setput $1000 $5000
Jupyter Notebooks free free
Cloud Storage(1TB) $240 $360
19
Kubernetes $1000 $1500
Data Version Control free free
MySQL free free
Flask API $20 $200
Power BI $20 $4995
Total Software $2280 $12055
Requirements
Cost
Conclusion
In conclusion, this GST prediction model helps in the tax
revenues for our booming country because of the lack of
accurate prediction models and by utilizing machine learning
and data analytics. Thus, by building state specific models
and then pooling them together to produce a national
forecast it improves fiscal strategy, resource management,
and policy making. The exact identification of data
processing, models to be used, and model improvement
steps make the system more responsive to fluctuating
economic situations. Besides, dashboard and reporting
improve GST real-time visualization and facilitates effective
decision making due to timely information. The cost of
developing this system covers all expenses, including
overhead costs while developing the structure guarantees
extensibility, offering the system suitability for large
numbers of data across most states. Finally, this model can
act as a useful tool for increasing the level of openness and
rationality in the economic governance of India.
References
1. I. P. C. Purnama, N. L. W. S. R. Ginantra, I. W. A. S. Darma
and I. P. A. E. D. Udayana, "Comparison of Support Vector
20
Machine (SVM) and Linear Regression (LR) for Stock Price
Prediction," 2023 Eighth International Conference on
Informatics and Computing (ICIC), Manado, Indonesia,
2023, pp. 1-6, doi: 10.1109/ICIC60109.2023.10381982.
2. S. Alex, S. Purakayastha, S. Chattaraj and A. Jadhav,
"Stock Price Prediction of IT Midcaps Companies Using ML
Models," 2024 Second International Conference on
Emerging Trends in Information Technology and
Engineering (ICETITE), Vellore, India, 2024, pp. 1-6, doi:
10.1109/ic-ETITE58242.2024.10493586.
3. K. Abdi, H. Rezaei and M. Hooshmand, "Machine Learning-
Based Fundamental Stock Prediction Using Companies'
Financial Reports," 2024 32nd International Conference on
Electrical Engineering (ICEE), Tehran, Iran, Islamic
Republic of, 2024, pp. 1-5, doi:
10.1109/ICEE63041.2024.10668367.
4. R. Vyas and R. As, "Seasonal Sales Prediction and
Visualization for Walmart Retail Chain Using Time Series
and Regression Analysis: A Comparative Study," 2022
International Conference on Smart Technologies and
Systems for Next Generation Computing (ICSTSN),
Villupuram, India, 2022, pp. 1-6, doi:
10.1109/ICSTSN53084.2022.9761294.
5. M. O. Adebiyi, S. A. Ajayi, D. Olaniyan, J. Olaniyan and D.
Kikelomo, "Start-Up Capital Estimation and Profitability
Prediction for E-Commerce Start-Ups Using Machine
Learning," 2024 International Conference on Science,
Engineering and Business for Driving Sustainable
Development Goals (SEB4SDG), Omu-Aran, Nigeria, 2024,
pp. 1-7, doi: 10.1109/SEB4SDG60871.2024.10630338.
6. T. Das, B. K. Agarwal and V. M. Gayathri, "Maximizing
Customer Base By Forecasting The Most Profitable
Customers Using Logistic Regression," 2023 International
Conference on Data Science and Network Security
(ICDSNS), Tiptur, India, 2023, pp. 1-6, doi:
10.1109/ICDSNS58469.2023.10245842.
21
7. Thayyib, P. V. et al. (2023) ‘Forecasting Indian Goods and
Services Tax revenue using TBATS, ETS, Neural Networks,
and hybrid time series models’, Cogent Economics &
Finance, 11(2). doi: 10.1080/23322039.2023.2285649.
8. Nayyar, Anand & Singh, Inderpal. (2018). A
Comprehensive Analysis of Goods and Services Tax (GST)
in India. Indian Journal of Finance. 12. 57.
10.17010/ijf/2018/v12i2/121377.
9. Antad, Sonali & Khandelwal, Saloni & Khandelwal, Anushka
& Khandare, Rohan & Khandave, Prathamesh & Khangar,
Dhawal & Khanke, Raj. (2023). Stock Price Prediction
Website Using Linear Regression - A Machine Learning
Algorithm. ITM Web of Conferences. 56.
10.1051/itmconf/20235605016.
10. J. Margaret Sangeetha, K. Joy Alfia, Financial stock
market forecast using evaluated linear regression-based
machine learning technique, Measurement: Sensors,
Volume 31, 2024, 100950, ISSN 2665-9174,
doi:10.1016/[Link].2023.100950.
22