International Journal of Pure and Applied Mathematics
Volume 119 No. 10 2108, 81-86
ISSN: 1311-8080 (printed version); ISSN: 1314-3395 (on-line version)
url: [Link]
Special Issue
[Link]
Analysis of stock market by using Big Data Processing Environment
M.D. Jaweed* and J. Jebathangam**
*Computer Science student, **Assistance Prof, Department of IT
VISTAS, Chennai - 600 117, India
E-mail: [Link]@[Link]
ABSTRACT The remainder of this paper is organized as
The purpose of this study is to apply Hadoop Big Data to follows.
financial analysis and to identify top companies Section 2 summaries the NYSE data set.
whose volume are traded highest in past years. For this Section 3 describes the data analysis view. In
research, historical data of NSE. analyzing the data in Section 4, we describe the components of our data
QlikView. The result contains top 11 companies by its analysis systems.
highest volume traded for each industry It is shown that Section 5 discusses the results of our experimental
financial data analysis can be done efficiently and evaluation. We conclude the research in Section 6.
easily using big data technologies like Hadoop and its
II. QLIKVIEW [3, 4]
ecosystem QLIKVIEW by using cloud services like
Microsoft Azure. QlikView is a self-service BI for all users in
organizations. With QlikView you can analyze data
Keywords : Hadoop, big data, QlikView, and use your data discoveries to support decision
Data analysis, NSE, financial data. making. QlikView lets you ask and answer your own
questions and follow your own paths to insight.
I. INTRODUCTION QlikView enables you and your colleagues to
This data is generated in social networking sites via reach decisions collaboratively.
posts from many users, sensor data: to get climate 2.2 NYSE Historical Data [1, 2]
information, purchase transaction records in large
Daily stock data of each company is available live
industry and many more. With the help of normal
on NSE for each stock exchange. We have
legacy systems, it becomes very difficult and
taken the NSE stock exchange data for this
expensive to store and analyze large scale data for
study. The data set is composed of: company
data analyst. It is also time consuming process. This
symbol, date, open of the day, high of the day, low
kind of large scale data with structured and
of the day, close of the day and volume.
unstructured format is called Big Data.
This paper is to present main 2 objectives: (1) Top
However, Hadoop framework is growing now a
11 companies which have been traded highest by its
days to store and analyze data and it is convenient
volume by each Industry. (2) Top 5 highest volume
for its functions.
traded of a specific company by its date.
Many companies have been using big data framework
to analyze the data and find some patterns and There are csv (comma separated values)
relationship among the data to target customer and market files containing following fields:
competition. In this study we have used • NSE: Company Symbol, Date, Open of the
NYSE historical data of 10,000 companies from Day, High of the Day, Low of the Day, Close
January 2007 to December 2018 (10 years). of the Day, Volume.
We collect the data to keep in HDFS III. IMPLICATIONS OF DATA ANALYSIS
(Hadoop Distributed File Systems) and to
analyze it using QlikView on Azure blob In this paper, we make the following contributions:
storage to find top 11 companies by its highest • We download and combine all 2,480 csv
traded volume for each industry. files in a single text file.
365
81
International Journal of Pure and Applied Mathematics Special Issue
INTERNATIONAL CONFERENCE ON
INNOVATIONS AND RECENT TRENDS IN COMPUTING TECHNIQUES AND APPLICATIONS
• We create one storage account and one
• It can convert the data into graphical
Hadoop cluster having 4 nodes on Microsoft Azure.
analytics.
• Simple to use and you don’t need any
We use ‘import’ all csv files into online
additional training to work on it.
server/ Hard drive through QlikView.
• Easy implementation, flexibility and high
scalability.
• It is one stop shop for dashboards,
graphical analytics and easy reporting
4.2 Experimental Results
We have taken the list of companies by sector wise from
the NSE We include this list in our query to find the
Figure 1: companies whose volume is traded highest in that
particular industry. There are total 11 sectors in NSE
stock exchange: Basic, Consumer Durables,
Technology, Energy, Transportation, Public Utilities,
Consumer Nondurables, Consumer Services, Capital
Goods, Finance, Healthcare, and Miscellaneous.
Some example are presented as follows for industries:
Eleven Sectors
SECTOR SYMBOL SUM OF VOLUME (LACS)
Figure 2: AUTO 11677.55
Hence we have one file .Qvw file format FMCG 27517.1
containing data for all companies we create a single BANK 26909.5
PVT BANK 15049.9
table using QlikView.
IT 12633.85
• This research finds the companies whose
volume of stock traded maximum from all the FIN SERVICE 11073.25
PHARMA 9637.8
industries of NYSE stock exchange.
METAL 4033.4
• Additionally we can have highest volumes
MEDIA 3540.55
of a single company till December 2018..
REALTY 347.3
4. METHODOLOGY Table 1 and Figure 2 illustrate the 10 sectors in
• This section describes the QlikView to create all industry who trade maximum volumes in
the table and store the data and to visualize the 2007 and 2018.
result. company till December 2018. Sector Symbol Max Volume (lacs)
4.1 Features of QlikView
• QlikView uses in-memory data model.
• It is capable of manipulating huge data
sets instantly with accuracy.
• Hardware cost for QlikView is not too high.
• It is a business discovery platform with
fast and powerful visualization capabilities.
• Automated data integration is possible
with QlikView.
Figure 3: Automobile Symbol Volume (lacs)
366
82
International Journal of Pure and Applied Mathematics Special Issue
INTERNATIONAL CONFERENCE ON
INNOVATIONS AND RECENT TRENDS IN COMPUTING TECHNIQUES AND APPLICATIONS
Second objective is to find 10 highest
volume of a company. As a result, we get the 10
highest volumes companies.
Third objective is to find highest volume by last
10 years.
Fourth objective is to find highest volume by
month for the particular year.
PHARMACY Symbol Volume (lacs) in Ascending
Figure 4: Order
Finance Symbol Volume (lacs) in Ascending Order
IT Symbol Volume (lacs)
VOLUME
SYMBOL SECTOR
(LACS)
UNITECH 488.9 Reality
SBIN 465.5 auto
ASHOKLEY 380.49 FMGC
ITC 350.08 PVT Bank
SOUTHBANK 264.46 bank
BANKBARODA 244.2 Reality
HDIL 242.1 bank
PHARMACY by Volume (lacs) HINDALCO 217.37 fin
PNB 176.22 metal
10 highest volume from the 10 sectors
367
83
International Journal of Pure and Applied Mathematics Special Issue
INTERNATIONAL CONFERENCE ON
INNOVATIONS AND RECENT TRENDS IN COMPUTING TECHNIQUES AND APPLICATIONS
Our second result for WFC shows that company
Month Volume (lacs) moved fast in 2008 and 2009 among 2000 through
Jan 441.15 2014. Its highest volumes are traded in these 2
Feb 835.6 years.
Mar 541.4 In this paper, we showed the possibility that
Apr 280.7 Big Data Hadoop and QlikView can be adopted for
May 299.35 financial industry. This approach should be very
Jun 528.1 useful in the field of Business Intelligence where
Jul 442.5 the company act upon the past performance and data. As
Aug 365.8 a future work, we will find more financial data set to
Sep 965.8 find out more useful relationship or pattern trends.
Oct 522
REFERENCES
Nov 422
Dec 321 [1] Jay Mehta, 2 Jongwook Woo 1 Graduate
Student,
[2] Prof, Department of Computer Information
Systems California State University Los
Angeles, USA
[3] David J. Hand, “Data, Not Dogma: Big Data,
Open Data,and the Opportunities Ahead”,
LNCS 8207, pp. 1-12,2013.
[4] Emerson, J.W., Kane, M. J.: Don’t drown in
[Link] Vol. 9(4), pp. 38–39, 2012.
[5] FergusonGT. Strategy in the digital age:
VOLUME (LACS) VOLUME YEARLY (LACS) Role ofinformation technology in corporate
2009 6003.6 strategic [Link] of Business
2010 7203 Strategy, Vol. 17(6), pp. 28–31, 1996.
2011 3368.4 [6] Gartner (2012), [Link]
2012 3592.2 DisplayDocument? id=2057415&ref =client
2013 5293.8 FriendlyUrl.
2014 6337.2 [7] Hand, D.J., Blunt, G., Kelly, M.G., Adams,
2015 6496.8 N.M.: Datamining for fun and profit.
2016 4389.6 Statistical Science 15, pp. 111–131, 2000.
2017 5310 [8] Hand, D.J.: Mining the past to determine the
2018 4023.6 future:problems and possibilities. International
Journal ofForecasting 25, pp. 441–451, 2008.
5. CONCLUSION
[9] Hoplin HP. Re-engineering information
From the above analysis we find the technology: Anenabler for the new business
companies who have made profits from each strategy. IndustrialManagement & Data
industries. Most of the Investors prefer to invest in System, Vol. 95(2), pp. 24–27, 1995.
such company, which is performing well in the [10] Kamran Rezaie, Vahid Majazi Dalfard,
equity market. The data should be useful for Loghman HatamiShirkouhi,Salman Nazari-
financial analyst and users, who works in the Shirkouhi, “Efficiency appraisal and ranking
stock market and analyze all the past records of a of decision-making units using data envelopment
company to advice their clients for investments. analysis in fuzzy environment: a case study of
Tehran stock exchange”, Neural Comput &
368
84
International Journal of Pure and Applied Mathematics Special Issue
INTERNATIONAL CONFERENCE ON
INNOVATIONS AND RECENT TRENDS IN COMPUTING TECHNIQUES AND APPLICATIONS
Applic DOI 10.1007/s00521-012-1209-6.
[11] Krishna Kumar Singh, Dr. Priti Dimri and
Madhu Rawat, “Fractal Market Hypothesis in
Indian Stock Market”,IJARCSSE, Vol.
3(11), pp. 739-743, Nov 2013.
[12] Manyika, J., Chui, M., Brwon, B., Bughin, J.,
Dobbs, R.,Roxburgh, C.,Byers, R.H.: Big data:
the next frontier for innovation, competition, and
productivity (2011), [Link]
com/insights/business technology/big data
the next frontier for innovation
[13] Chien-Jen Huang, Peng-Wen Chen, Wen-Tsao
Pan, “Usingmulti-stage data mining technique
to build forecast model for Taiwan stocks”,
Springer-Verlag London Limited 2011.
369
85
86