0% found this document useful (0 votes)
2 views20 pages

Abstract Data mining

The report outlines the progress in the field of computer science, particularly focusing on data mining and its potential to create employment opportunities for students. It highlights the importance of Business Intelligence and Analytics in enhancing organizational performance and discusses the growing demand for IT professionals due to a significant gap in the workforce. Additionally, the report emphasizes the role of data mining in e-commerce and its impact on decision-making and business intelligence.

Uploaded by

muirujohn14
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views20 pages

Abstract Data mining

The report outlines the progress in the field of computer science, particularly focusing on data mining and its potential to create employment opportunities for students. It highlights the importance of Business Intelligence and Analytics in enhancing organizational performance and discusses the growing demand for IT professionals due to a significant gap in the workforce. Additionally, the report emphasizes the role of data mining in e-commerce and its impact on decision-making and business intelligence.

Uploaded by

muirujohn14
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Computer Science Progress: Data Mining 1

Computer Science Progress: Data Mining


Name:
Institution:
Computer Science Progress: Data Mining 2

Abstract
As a computer engineer at the Dynamite Hardware and Software services and the purpose of the report
is to show and provide more information about our progress so far. The field of computer science has
become one of the most vibrant enterprises across the globe. A wide range of college institutions,
government organs as well as private institutions have embraced computer science in one way or the other.
The purpose of this progress report is to provide insights into computer science, especially data mining and
how it can provide innovative approaches tailored to create more employment opportunities for students
enrolled in the field.
Computer Science Progress: Data Mining 3

Executive Summary
Business Intelligence and Analytics (BI&A) is now recognized and understood as a very significant
factor that can enhance organizational performance through intelligently taken decisions and meticulous
ideas and perspectives. Software solutions are instrumental in utilizing business applications. Risk
management and enterprise decision-making cannot be separated from mining tools. Business Intelligence
(BI) is acquired by using mining. Use of data warehousing and Information Systems (IS) have made it
possible for enterprise datasets to proliferate. Credit card companies generally log millions of transactions
in any given year (Bhagyashree, & Borkar, 2012). Mobile operators and telecommunications companies
usually generate the largest datasets. User accounts exceed 100 million which make billions of data per
year. Against such numbers, OLAP or other analytical processing and manual operation do not have any
chance to stand up. BI, on the other hand, makes the task possible.

With the aim of getting to the core of the subject, the period covered in this research was from October
24, 2014, through November 15, 2015. At the moment, this study has already covered 45% of the final
project as scheduled. More specifically, in the Computer Science field, advanced research should be carried
out to find means of improving the process of data mining for higher education institutions, as a strategy for
the creation of employment opportunities which will encourage more students to enrol in Computer Science
courses under the branch of software development (Bhagyashree, & Borkar, 2012). After surveying the
current state of the literature in the history of computing, this paper discusses some of the benefits of
Computer Science, primarily under the field of Data Mining and its application. It suggests aspects of the
development of computing which are pertinent to those benefits and hence for which that recent work could
provide models of historical analysis. As a new scientific or statistical technology with unique features,
computing, in turn, can offer new perspectives on the history of technology (Bhagyashree, & Borkar, 2012).

Introduction and Background Information


The term “data mining" has been used in a variety of contexts in data analysis in recent years. It is
likely that you already think of data mining from the computer science viewpoint, namely, as a broad set of
techniques and algorithms for extracting useful patterns and models from extensive data sets (Bhagyashree,
& Borkar, 2012). Since the early 1990s, there has been an immense surge of both research and commercial
activity in this area, driven mainly by a computationally-efficient massive search for patterns in data (such
as association rules) and by business applications such as analysis of extensive transactional data archives.

As evidenced by well-known business-oriented data mining texts (e.g., [BL00]), as well as different
published proceedings of data mining research conferences, much current work in data mining is focused on
Computer Science Progress: Data Mining 4

algorithmic issues such as computational efficiency and data engineering issues such as data representation
(Kusiak, et al., 2016). While these are relevant topics in their own right, statistical considerations and
analyses are often conspicuous by their absence in data mining texts and papers. The is a natural
consequence of the fact that research in data mining is primarily practiced by computer scientists who
naturally focus on algorithmic and computational issues rather than on statistical matters. In this case, I will
explore the interplay of computer science and statistics and how this has impacted and will impact, the
development of data mining.

Everyone knows that ‘‘computing has changed the world,’’ but, strangely enough, our existing
historiography of computing faces numerous difficulties in addressing this question directly. Examples and
models for a historical understanding of this key question are surprisingly scarce (Kusiak, et al., 2016). I
believe this is because historians’ disciplinary preferences for subject specificity and archival virtuosity
have encouraged us to do detailed studies of individual machines, programs, and companies, and
occasionally to examine the ‘‘social construction’’ of specific computing technologies (Rao, T.K.R.K.,
Khan, Begun, and Divakar, 2013). However, our primary focus on specifics has made it difficult to
conceive and conduct the wide-ranging and long-duration studies that can show the longer-term
consequences of technical changes for society, culture, economics, and politics.

The Case for Computer Science Education


Every academic subject, from Latin to art and from math to the humanities, has dedicated advocates
(including the teachers who teach the subjects) who believe that the U.S. education system can be improved
by committing more time to its study and conversely that the system would be gravely weakened by
reducing exposure to the subject. Just look at the recent controversy President Obama ignited when he had
the nerve to (rightly) say that it was more important to teach students advanced manufacturing skills than
art history; art historians came out in force to express their righteous indignation (Rao, T.K.R.K., Khan,
Begun, and Divakar, 2013). Given that there are a limited number of hours in the academic year, not every
subject matter advocate can be right. Choices have to be made. Not making decisions and continuing with
status quo is in itself a choice. However, there is a strong argument to be prepared for putting relatively
more focus on computer science (CS) and featuring it in every high school in the country.

Computer science challenges students and teaches them to approach problems in new and rigorous
ways. If appropriately trained, computer science courses instill creativity, critical thinking skills, and logical
reasoning. Its core concepts are broadly transferable, giving students the ability to apply skills to myriad
problems, enabling them to pursue cross-disciplinary pursuits, and allowing them to learn about the world
Computer Science Progress: Data Mining 5

they live (Rao, T.K.R.K., Khan, Begun, and Divakar, 2013). And, perhaps most importantly, computer
science provides computational literacy and problem-solving skills that are desperately needed by the
workforce. CS ensures that students are competitive and adaptable in the labour market, not just for jobs in
computer science, but for many occupations that increasingly require “double-deep” skills.

Demand for Computer Science in IT Profession


As technology plays a more significant role in our world, growth in IT jobs has outstripped overall job
growth. The widening and deepening of demand for computer scientists have led to above-average wages
and faster wage growth in this field relative to the others. In the last decade, IT occupations have grown by
36 percent (Wu, Zhang, and Li, 2013). Demand for these jobs has increased even faster, but there are not
enough IT professionals to meet the rapidly expanding market (Wu, Zhang, and Li, 2013). This gap
between the supply and demand for IT workers is a significant component of the national STEM shortage.
There are over 545,000 unfilled jobs requiring technology skills; while these jobs demand a diverse set of
STEM skills, many are jobs requiring the ability to solve problems with computers (Wu, Zhang, and Li,
2013). This line chart shows the 10-year projected employment growth (from 2014 to 2024) for Computer,
engineering, & science occupations. This profession is expected to grow faster than 7.4%, the average rate
of national job growth (Wu, Zhang, and Li, 2013).

Figure 1: Job Growth Projections in Computer, Engineering and Science Professional


Computer Science Progress: Data Mining 6

Source: DataUSA. (2018). Computer, Engineering and Science Occupations. From:


[Link]

Employment in the Field of Computer Science


Information on the businesses and industries that employ Computer and Information Sciences and
Support Services graduates and on wages and locations for those in the field (DataUSA., 2018). Computer
Systems Design is the industry that uses the most Computer and Information Sciences and Support
Services majors both by share and by number, though the highest paying industry for Computer and
Information Sciences and Support Services majors, by average wage, is Internet publishing, broadcasting &
web search portals (DataUSA., 2018).

Wages and Salaries


The closest comparable data for the 6 Digit Course Computer Science is from the 2 Digit Course
Computer and Information Sciences and Support Services. According to DataUSA (2018), the average
wage in the workforce is $92,592 (+/- $948), which reflects an annual growth of 4.33% (+/- 1.66%)
(DataUSA., 2018). The average salary for those who major in Computer and Information Sciences and
Support Services is $92,592, and the most common occupations are Software developers, applications &
systems software; Computer programmers; and Other Computer Occupations (DataUSA., 2018). The
profession that pays the highest of the five common jobs filled by Computer and Information Sciences and
Support Services majors is Physicians & surgeons at $196,298 per year (DataUSA., 2018).

Figure 2: Annual Salaries of the most common occupations for Computer and Information Sciences
Computer Science Progress: Data Mining 7

Source: DataUSA. (2018). Computer Science. From: [Link]

Database Management
Another use of computers, having a significant effect on the field of public health, is in storing and
manipulating large databases. Secondary storage costs have decreased from approximately $500 per
megabyte in the mid-1970s to $30 per megabyte today, while electromagnetic disks are increasing in
capacity and access speed (Srinniva, Srinivas, and Harsh, 2013). All predictions indicate that these
developments will continue, accompanied by additional technologic improvements in disks. Other forms of
storage are becoming available and affordable. In particular, optical discs, which are slower than
conventional devices but can hold massive amounts of information and can be readily duplicated, are
providing an excellent alternative to tape for archival storage and information sharing (Srinniva, Srinivas,
and Harsh, 2013).

Coupled with the economies of scale afforded by the large-scale integration, the use of computers for
diverse data storage requirements is now viable. Software developers are following the lead of the hardware
manufacturers and developing database management systems to handle massive amounts of information
(Srinniva, Srinivas, and Harsh, 2013). Relational model-based systems, for example, provide a mechanism
for representing complex data with a conceptually simple structure. Advances are also occurring in an
environment where many researchers analyze the standard bank of data. Sophisticated query mechanisms
are also being designed, as are improved methods of security, setting the groundwork for future needs
(Srinniva, Srinivas, and Harsh, 2013).
Computer Science Progress: Data Mining 8

Concepts in database management hardly fall in the category of come-and-go, as the cost of shifting
between technical approaches overwhelms producers, managers, and designers. However, there are several
trends in database management, and knowing how to take advantage of them will benefit your organization.
Following are some of the current trends:

1. Databases that bridge SQL/NoSQL


The latest trends in database products are those that don’t merely embrace a single database structure.
Instead, the databases span SQL and NoSQL, giving users the best capabilities offered by both. For
example, it includes products that allow users to access a NoSQL database in the same way as a relational
database.

2. Databases in the cloud/Platform as a Service


As developers continue pushing their enterprises to the cloud, organizations are carefully weighing the
trade-offs associated with public versus private. Developers are also determining how to combine cloud
services with existing applications and infrastructure. Providers of cloud service offer many options to
database administrators. Making the move towards the cloud doesn’t mean changing organizational
priorities, but finding products and services that help your group meet its goals.

3. Automated management
Automating database management is another emerging trend. The set of such techniques and tools
intended to simplify maintenance, patching, provisioning, updates and upgrades — even project workflow.
However, the pattern may have limited usefulness since database management frequently needs human
intervention.

4. An increased focus on security


While not exactly a trend given the constant emphasis on data security, recent ongoing retail database
breaches among US-based organizations show with sufficient clarity the importance for database
administrators to work hand-in-hand with their IT security colleagues to ensure all enterprise data remains
safe. Any organization that stores data is vulnerable. Database administrators must also work with the
security team to eliminate potential internal weaknesses that could make data vulnerable. These could
include issues related to network privileges, even hardware or software misconfigurations that could be
misused, resulting in data leaks.
Computer Science Progress: Data Mining 9

5. In-memory databases
Within the data warehousing community there are similar questions about columnar versus row-based
relational tables; the rise of in-memory databases, the use of flash or solid-state disks (which also applies
within transaction processing), clustered versus no-clustered solutions and so on.

6. Big Data
In other words, to be clear, big data does not necessarily mean lots of data. What it refers to is the
ability to process any data: what is typically referred to as semi-structured and unstructured data as well as
structured data. Current thinking is that these will live alongside conventional solutions as separate
technologies, at least in large organizations, but this will not always be the case.

Current Progress in the field of Computer Science


Based on the current findings, computer science is considered to be one of the areas that are growing
very fast all over the world. In this report, much has already been covered. For example, comprehensive
insights on ways through which computer science has helped the global economies to fight unemployment
among the youth are provided. While reflecting on the study conducted by Lidström, Holm, and Lundström
(2014), it was revealed that the CGPS technology could be useful in helping students to maximize the
chances of getting enrolled in their courses of choice. In most cases, this can be done by using the computer
system, which employs the technique of CGPS to analyze the cumulative grades for the number of years a
student takes to complete a given course.

The statistical data mining process for obtaining the CGPA information of every student in a
geographical location requires the use of sophisticated analytical data mining procedures and algorithms.
By finding the best suitable means of data mining for students’ grades, there will be increased efficiency in
the process of student selection for advanced certificate courses such as Ph.D. depending on the standard
cutoff grades for specific higher learning institutions (Lidström, Holm, and Lundström, 2014).

The graphic below is a bar chart representation of enrolment growth by year. The research has proven
that computer science directly empowers innovative youths in societies, paves the way for a more equitable
world, accelerates health progress (Kusiak, et al., 2016). Similarly, computer science expands
communications which lead to better delivery of goods and services, it has improved individuals' creativity
in different fields and led to inventions. This has resulted in a fall in the number of unemployed youths
(Flavin, 2015). However, it should be recalled that the whole of the study is soon to be completed so as to
compile it and get its full picture and recommendations.
Computer Science Progress: Data Mining 10

Figure 1 Bar Chart for Growth of Graduate Enrolment From 2011 to 2017

Source: Flavin, B. (2015). Figure 4.15: How unemployment rates are affected by problem-solving
proficiency and lack of computer experience. Doi: 10.1787/888933231788

Data Mining in E-commerce


Data mining in e-commerce is a vital way of repositioning the e-commerce company for supporting the
enterprise with the required information concerning the business. Recently, most companies adopt e-
business and being in possession of big data in their data repositories. The only way to get the most out of
this data is to mine it to increase decision making or to enable business intelligence (Guo, Wang, Sun, Li,
Xin, and Zhang, 2012). In e-commerce data mining there are three critical processes that data must pass
before turning into knowledge or application. Figure 1 shows the steps for data mining in e-commerce.

Figure 1. Data mining process in e-commerce


Computer Science Progress: Data Mining 11

Source: Ismail, M., Ibrahim, M. M., Sanusi, Z. M., & Nat, Muesser. (2015). Data Mining in Electronic
Commerce Benefits and Challenges. Management Information Systems Department, Cyprus International
University, Turkey.

Benefits of Data Mining in E-commerce


Application of data mining in e-commerce refers to possible areas in the field of e-commerce where
data mining can be utilized for enhancements in business. As we all know while visiting an online store for
Computer Science Progress: Data Mining 12

shopping, users usually leave behind specific facts that companies can store in their database (Zhao, Dong,
Zhao, and Wong, 2008). These facts represent unstructured or structured data that can be mined to provide a
competitive advantage to the company. The following areas are where data mining can be applied in the
field of e-commerce for the benefits of companies:

1) Customer Profiling
It is also known as the customer-oriented strategy in e-commerce. This allows companies to use
business intelligence through the mining of customer’s data to plan their business activities and operations
as well as develop new research on products or services for prosperous e-commerce. Classifying the
customers of great purchasing potentially from the visiting data can help companies to lessen the sales cost
(Zhao, Dong, and Li, 2007). Companies can use users’ browsing data to identify whether they are
purposefully shopping or just browsing or buying something they are familiar with or something new. It
helps companies to plan and improve their infrastructure.

2) Personalization of Service
Personalization is the act to provide contents and services geared to individuals by information of their
needs and behaviour. Data mining research related to customization has focused mostly on recommender
systems and related subjects such as collaborative filtering (Lu, Dong, and Li, 2015). Recommender
systems have been explored intensively in the data mining community. This system can be divided into
three groups: Content-based, social data mining and collaborative filtering. These systems are cultured and
learned from explicit or implicit feedback of users and are usually represented as the user profile. Social
data mining, based on the source of data that are created by a group of individuals as part of their daily
activities, can be an indispensable source of valuable information for companies (Lu, Dong, and Li, 2015).
Contrarily, personalization can be achieved by the aid of collaborative filtering, where users are matched
with particular interest and in the same vein the preferences of these users to make recommendations.

3) Basket Analysis
Every shoppers’ basket has a story to tell, and market basket analysis (MBA) is common retail,
analytics and business intelligence tool that helps retailers to know their customers better. There are
different ways to get the best out of market basket analysis and these include:
• Identification of product affinities; tracking not so apparent product affinities and leveraging on them
is the real challenge in retail (Lu, Dong, and Li, 2015). Walmart customers purchasing Barbie dolls shows
an affinity towards one of three candy bars, an obscure connection such as this can be discovered with an
advanced market basket analytics for planning more effective marketing efforts (Kusiak, et al., 2016).
Computer Science Progress: Data Mining 13

• Cross-sell and up-sell campaigns; these shows the products purchased together, so customers who buy
the printer can be persuaded to pick up high-quality paper or premium cartridges (Lu, Dong, and Li, 2015).
• Planograms and product combos are used for better inventory control based on product affinities,
developing combo offers and design effective user-friendly planograms in focusing on products that sell
together.
• Shoppers profile; in analyzing market basket with the aid of data mining over time to get a glimpse of
who your shoppers are, gaining insight to their ages, income range, buying habits, likes and dislikes,
purchase preferences, levering this and giving the customer experience.

4) Sales Forecasting
Sales forecasting involves the aspect of the time an individual customer spend to buy an item and in
this process trying to predict if the customer will buy again. This type of analysis can be used to determine a
strategy of planned obsolescence or figure out complimentary products to sell. In sales forecasting, cash
flow can be projected into three which include the pessimistic, optimistic and the realistic (Fugon, Juban,
and Kariniotakis, 2008). This helps to have a plan on the adequate amount of capital available to endure the
worst possible scenario that is if sales do not go as planned.

4) Merchandise Planning
Merchandise planning is useful for both online and offline retail companies. In the case of online
business, merchandise planning will help to determine stocking options and the inventory warehousing,
while in the case of offline companies, business that are looking to boost by adding stores can assess the
required amount of merchandise they will be adequately needing by having a foresight at the exact layout of
the current store (Fugon, Juban, and Kariniotakis, 2008). Using the right approach to merchandise planning
will undoubtedly lead to answers on what to do with:
• Pricing: the aspect of database mining will help to determine the suited best price of products or
services in the processes of revealing customer sensitivity (Fugon, Juban, and Kariniotakis, 2008).
• Deciding on products; data mining provides e-commerce businesses with the aspect of which products
customers desire, which includes the element of intelligence on competitor’s merchandise (Fugon, Juban,
and Kariniotakis, 2008).
• Balancing of stocks; in mining the local database, it helps determine the right and specific amount of
stocks needed, i.e. not too much and not too less, throughout the business year and also during the buying
seasons (Fugon, Juban, and Kariniotakis, 2008).

6) Market Segmentation
Computer Science Progress: Data Mining 14

Customer segmentation is one of the best uses of data mining. From the lots of data gotten, it can be
broken down into different and meaningful segments like income, age, gender, the occupation of customers,
and this can be used when either the companies are running email marketing campaigns or SEO strategies
(Devi, and Manonmani, 2012). The aspect of market segmentation can also help a company identify its
competitors. This provided information alone can help the retail company determine that the periodic
respondents are usually not the only ones pointing the same customer money as the present company is
(Devi, and Manonmani, 2012). Segmenting the database of a retail company will improve the conversion
rates as the company can focus their promotion on a close-fitted and highly wanted market (Devi, and
Manonmani, 2012). It also helps the retail company to understand the competitors that are involved in every
segment in the process permitting the customization of products that will generically satisfy the target
audience.

7) Fight against Terrorism


After the 9-11 attacks in the United States, many countries imposed new laws against fighting
terrorism. These laws allow intelligence agencies to fight against terrorist organizations effectively.
The USA launched the Total Information Awareness program with the goal of creating a massive
database of that consolidate all the information on population. Similar projects were also started in
European countries and the rest of the world (Devi, and Manonmani, 2012).

Data Mining and its Progress in the field of Health Care


Healthcare-based business intelligence systems are complicated to build, maintain and face the
knowledge engineering and computing difficulties. Intelligent data mining techniques provide effective
computational methods and robust environment for business intelligence in the healthcare decision making
systems (Devi, and Manonmani, 2012). Intelligent data mining (IDM) approach aims to extract useful
knowledge and discover some hidden patterns from a huge amount of databases, which statistical methods
cannot discover.

IDM and knowledge discovery (KD) is not a coherent field; it is a dwells upon already well-established
technologies including data cleaning, data pre-processing, machine learning, pattern recognition, statistics,
neural networks, fuzzy sets, rough sets, clustering, etc. Healthcare organization uses the business
intelligence (BI) solution not only for analysis but also to change business processes and drive toward the
value-driven healthcare vision (Edwards, Wilkinson, Canny, Pearce, & Coates, 2014). BI provides an
integrated view of data that can be used to monitor key performance indicators, identify hidden patterns in
diagnosis, illuminate anomalies in processes, and identify variations in cost factors, all of which facilitate
Computer Science Progress: Data Mining 15

accountability and visibility and can drive an organization towards efficiency (Edwards, Wilkinson, Canny,
Pearce, & Coates, 2014).

Data Mining in Higher Education


Nowadays, higher learning institutions encounter many problems, which keep them away from
achieving their quality objectives — most of these problems caused by a knowledge gap. The knowledge
gap is the lack of significant knowledge at the main educational processes such as advising, planning,
registration, evaluation and marketing. For example, many learning institutions do not have access to the
necessary information to advise students. Therefore, they are not able to give a suitable recommendation for
them.

Data mining is a powerful tool for academic intervention. Through data mining, a university could, for
example, predict with 85% accuracy which students will or will not graduate. The university could use this
information to concentrate academic assistance on those students most at risk. Therefore, to understand how
and why data mining works, it’s important to understand a few fundamental concepts (Zhao, Dong, Li, and
Wong, 2007). First, data mining relies on four essential methods: Classification, categorization, estimation,
and visualization. Classification identifies associations and clusters and separates subjects under study.
Categorization uses rule induction algorithms to handle categorical outcomes, such as “persist” or
“dropout,” and “transfer” or “stay” (Zhao, Dong, Li, and Wong, 2007).

Estimation includes predictive functions or likelihood and deals with continuous outcome variables,
such as GPA and salary level. Visualization uses interactive graphs to demonstrate mathematically induced
rules and scores and is far more sophisticated than pie or bar charts. Visualization is mainly used to depict
three-dimensional geographic locations of mathematical coordinates. Higher education institutions can use
classification, for example, for a comprehensive analysis of student characteristics, or use estimation to
predict the likelihood of a variety of outcomes, such as transferability, persistence, retention, and course
success.

Educational Data Mining (EDM) is an emerging discipline, concerned with developing methods for
exploring the unique types of data that come from educational settings and using those methods to
understand better the students, and the environments which they learn in. A critical area of EDM is a
mining student’s performance. Another key area is mining enrollment data. Vital uses of EDM include
predicting student performance and studying learning to recommend improvements to current educational
practice. EDM can be considered one of the learning sciences, as well as an area of data mining. The main
Computer Science Progress: Data Mining 16

applications of EDM are listed as follows:

1. Analysis and Visualization of Data


It is used to highlight useful information and support decision making. In the educational environment,
for example, it can help educators and course administrators to analyze the students’ course activities and
usage information to get a general view of a student’s learning. Statistics and visualization information are
the two main techniques that have been most widely used for this task. Statistics is a mathematical science
concerning the collection, analysis, interpretation or explanation, and presentation of data (Wu, Zhou, Mo,
and Zhu, 2006). It is relatively easy to get basic descriptive statistics from statistical software, such as
SPSS. Statistical analysis of educational data (logs files/databases) can tell us things such as where students
enter and exit, the most popular pages students browse, number of downloads of e-learning resources,
number of different web pages browsed and total time for browsing different pages.

It also provides knowledge about usage summaries and reports on weekly and monthly user trends,
amount of material students might go through and the order in which students study topics, patterns of
investigating activity, timing and sequencing of events, and the content analysis of students notes and
summaries (Wu, Zhou, Mo, and Zhu, 2006). Statistical analysis is also instrumental in obtaining reports
assessing how many minutes student worked, some problems he resolved and his exact percentage along
with our prediction about his score and performance level.

2. Predicting Student Performance


In this case, we estimate the unknown value of a variable that describes the student. In education, the
values usually predicted are the student’s performance, their knowledge, score, or marks (Wu, Zhou, Mo,
and Zhu, 2006). This value can be a numerical/continuous (regression task) or categorical/discrete
(classification task). Regression analysis is used to find a relation between a dependent variable and one or
more independent variables (Carbone, 2015). Classification is used to group individual items based on
quantitative characteristics inherent in the elements or on a training set of previously labelled items.

Prediction of a student’s performance is the most popular applications of DM in education. Different


techniques and models are applied like neural networks, Bayesian networks, rule-based systems, regression,
and correlation analysis to analyze educational data (Carbone, 2015). This analysis helps us to predict
student’s performance, i.e. to predict his success in a course and to predict his final grade based on features
extracted from logged data.
Computer Science Progress: Data Mining 17

3. Grouping Students
In this case groups of students are created according to their customized features, personal
characteristics, etc. These clusters/groups of students can be used by the instructor/developer to build a
personalized learning system which can promote active group learning (Carbone, 2015). The DM
techniques used in this task are classification and clustering. Different clustering algorithms that are used to
group students are hierarchical agglomerative clustering, K-means and model-based clustering.

Conclusion and Recommendations


Finally, this progress report provides excellent insights into the current study of computer science,
especially data mining and how it can provide innovative approaches tailored to create more employment
opportunities for students enrolled in the field. More so, significant progress on the key findings has been
highlighted as well as a few drawbacks (Cao, Li, & Yu, 2011). Nonetheless, it was indicated that the study
is still on the right track as far as the schedule is concerned. Business intelligence provides an integrated
view of data that can be used to monitor key performance indicators, identify hidden patterns in diagnosis,
illuminate anomalies in processes, and identify variations in cost factors, all of which facilitate
accountability and visibility and can drive an organization towards efficiency.

The efficiency of such systems is determined by the effectiveness of intelligent techniques and
methodologies. Intelligent data mining techniques provide effective computational methods and robust
environment for business intelligence in the healthcare decision-making domain. Data mining techniques
can be seen as an enabler for storing, managing, representing, analyzing, visualizing, and giving efficient
access to a massive amount of data (Cao, Li, & Yu, 2011).

The data collected about customers and their transactions, which are the greatest assets of e-commerce
companies, needs to be used consciously for the benefits of the companies. For such companies, data
mining plays an essential role in providing customer-oriented services to increase customer satisfaction
(Cao, Li, & Yu, 2011). It has become apparent that utilizing data mining tools is a necessity for e-commerce
companies in this globally competitive environment. Although the complexity and granularity of the
mentioned challenges differ, e-commerce companies can overcome these problems by using and applying
the right techniques. For example, developing an e-commerce website in a way that search engines can read
and access the latest version of the site, help companies to overcome the Search Engine Spider
Identification Problem.

References
Computer Science Progress: Data Mining 18

Bellazzi, R., & Zupan, B. (2008). Predictive data mining in clinical medicine: Current issues and guidelines.
International Journal of Medical Informatics 77 , 81–97.

Bhagyashree, A. and Borkar, V. (2012). Data Mining in Cloud Computing. Multi-Conference (MPGINMC-
2012). [Link]

Cao, L., Li, Y. & Yu, H. (2011). Research of Data Mining in Electronic Commerce. IEEE Computer
Society, Hebei.

Carbone, P.L. (2015) Expanding the Meaning and Application of Data Mining. International Conference on
Systems, Man and Cybernetics, 3, 1872-1873.
[Link]

DataUSA. (2018). Computer, Engineering and Science Occupations. From:


[Link]

DataUSA. (2018). Computer Science. From: [Link]

Devi, M. R. and Manonmani, R. (2012). Forecasting Using Data mining techniques in Tamil Nadu and
Other Countries - A Survey. International Journal of Emerging Trends in Engineering and
Development, 6(2), 295–302.

Edwards, D., Wilkinson, D., Canny, B. J., Pearce, J., & Coates, H. (2014). Developing outcomes
assessments for collaborative, cross-institutional benchmarking: Progress of the Australian
Medical Assessment Collaboration. Medical Teacher, 36(2), 139-147.

Flavin, B. (2015). Figure 4.15: How unemployment rates are affected by problem-solving proficiency and
lack of computer experience. Doi: 10.1787/888933231788

Fugon, L., Juban, J. and Kariniotakis, G. (2008). Data Mining for Wind Power Forecasting. In: European
Wind Energy Conference & Exhibition EWEC 2008 Brussels,Belgium, 1–6.

Guo, Q., Wang, Y., Sun, H., Li, Z., Xin, S. and Zhang, B. (2012). Factor Analy-sis of the Aggregated
Electric Vehicle Load Based on Data Mining. Energies, 5(6),2053–2070.
Computer Science Progress: Data Mining 19

Ismail, M., Ibrahim, M. M., Sanusi, Z. M., & Nat, Muesser. (2015). Data Mining in Electronic Commerce
Benefits and Challenges. Management Information Systems Department, Cyprus
International University, Turkey.

Kusiak, A., Kernstine, K., Kern, J., McLaughlin, K., & Tseng, T. (2016). Data Mining: Medical and
Engineering Case Studies. Industrial Engineering Research 2016Conference, (pp. 1-7).
Cleveland, Ohio.

Lidström, L., Holm, A. S., & Lundström, U. (2014). Maximizing opportunity and minimizing risk? Young
people’s upper secondary school choices in Swedish quasi-markets. Young,
22(1), 1-20.

Lu, X., Dong, Y. Z. and Li, X. (2015). Electricity Market Price Spike Forecast with Data Mining
Techniques. Electric Power Systems Research,73(1), 19–29.

Rao, T.K.R.K., Khan, S.A., Begun, Z. and Divakar, Ch. (2013). Mining the E-Commerce Cloud: A Survey
on Emerg-ing Relationship between Web Mining, E-Commerce and Cloud Computing.
IEEE International Conference on Computational Intelligence and Computing Research,
Enathi, 26-28 December 2013, 1-4. [Link]

Srinniva, A., Srinivas, M.K. and Harsh, A.V.R.K. (2013). A Study on Cloud Computing Data Mining.
International Journal of Innovative Research in Computer and Communication
Engineering, 1, 1232-1237.

Salmasi, F., Sattari, M. T. and Pal, M. (2012). Application of Data Mining on Evaluation of Energy
Dissipation Over Low Gabion-Stepped Weir. Turkish Journal of Agriculture and
Forestry,36(1), 95–106.

Weiss, G. M., & Davison, B. D. (2010). Data Mining. Handbook of Technology Management, H. Bidgoli
(Ed.), John Wiley and Sons.

Wu, W., Zhou, J., Mo, L. and Zhu, C. (2006). Forecasting electricity market price spikes based on Bayesian
expert with support vector machine. Advanced Data Mining and Applications Lecture
Computer Science Progress: Data Mining 20

Notes in Computer Science, 4093, 205–212.

Wu, M., Zhang, H. and Li, Y. (2013). Data Mining Pattern Valuation in Apparel Industry E-Commerce
Cloud. IEEE 4th International Conference on Software Engineering and Service Science
(ICSESS), 689-690.

Zhao, J. H., Dong, Z. Y., Li, X. and Wong, K. P. (2007). A General Method for Electricity Market Price
Spike Analysis. IEEE Transactions on Power Systems, 22(1),376–385.

Zhao, J. H., Dong, Z. Y. and Li, X. (2007). Electricity Market Price Spike Forecasting and Decision
Making. IET Generation, Transmission & Distribution, 1(4), 647–654.

Zhao, J. H., Dong, Z. Y., Zhao, X. and Wong, K. P. (2008). A Statistical Approach forInterval Forecasting
of the Electricity Price. IEEE Transactions on Power Systems, 23(2), 267–276.

You might also like