Understanding Data for Business Analytics
Understanding Data for Business Analytics
MODULE-2
Data can be small or very large, structured (organized for computers) or unstructured (like
text or videos), and it can arrive in small streams or in large batches. This variety and
complexity are what we often call Big Data. While it makes data harder to process, it also
increases its value by allowing us to gain deeper insights. Modern technologies, such as the
Internet and sensors, have replaced traditional manual data collection, making data gathering
faster, larger in volume, and more accurate.
However, not all data is useful. To be valuable, data must be “analytics-ready.” This means it
should be relevant to the problem, meet quality and quantity requirements, and follow a
proper structure. Organizations also need agreed definitions for common terms (like how to
define a customer). Some analytics methods require specific formats—for example,
predictive models often need flat files with a target variable, while neural networks require all
variables to be numeric.
According to Delen (2015), several dimensions define whether data is suitable for analytics:
Data Reliability: Data should come from trustworthy sources to avoid errors or
misrepresentation.
Data Accuracy: The data should correctly represent what it is supposed to, such as
recording customer information exactly as provided.
Data Accessibility: Data should be easy to obtain when needed, even if stored across
different systems.
Data Security and Privacy: Only authorized people should access the data,
especially sensitive information like health records.
Data Richness: The dataset should include all necessary variables to fully describe
the subject and build reliable models.
Data Consistency: Data collected from different sources should match and merge
correctly to avoid errors, like mixing records of different people.
Data Timeliness (Currency): Data must be up-to-date and captured close to the time
of the event for accuracy.
Data Granularity: Data should be detailed enough for the analysis, such as recording
lab test results with the required decimal precision.
Data Validity: Values must fall within defined acceptable ranges (e.g., gender
recorded as male, female, or unknown).
Data Relevancy: Only data relevant to the problem should be included, since
irrelevant data can mislead the analysis.
Structured Data
Structured data is most often used in data mining and analytics. It is divided into categorical
and numeric types.
1. Nominal data are labels or names without order, such as marital status (single,
married, divorced) or eye color (brown, blue, green). They can be binary (yes/no) or
multinomial (more than two categories).
2. Ordinal data have an order or ranking, such as credit score (low, medium, high) or
education level (school, college, graduate). The ranking adds extra meaning beyond
simple labels.
1. Interval data are measured on scales with equal intervals but no true zero point. For
example, Celsius temperature has meaningful differences but does not have an
absolute zero.
2. Ratio data include measurements with a true zero, such as mass, length, time, or
Kelvin temperature. These values can be compared using ratios (e.g., twice as heavy,
half as long).
Data can also come as text, images, videos, or sounds. These need to be converted into
categorical or numeric forms for analytics. Data may also be static (fixed) or dynamic/time-
series (changing over time).
Different analytics and machine learning methods work with different data types. Some
require all variables to be numeric, while others work better with categorical values. In such
cases, data scientists convert categorical values into numeric codes (like one-hot encoding).
On the other hand, decision tree algorithms may prefer categorical inputs, so numeric data
may need to be grouped into ranges.
Application Case 2.1 Medical Device Company Ensures Product Quality While Saving
Money
Few technologies are advancing faster than those in the medical field—so having the right
advanced analytics software can be a game changer. Instrumentation Laboratory is a leader in
the development, manufacturing, and distribution of medical devices and related
technologies, including technology that is revolutionizing whole blood and hemostasis
testing. To help ensure its continued growth and success, the company relies on data analytics
and Dell Statistica.
Problem
As a market leader in diagnostic instruments for critical care and hemostasis, Instrumentation
Laboratory must take advantage of rapidly evolving technologies while maintaining both
quality and efficiency in its product development, manufacturing, and distribution processes.
In particular, the company needed to enable its research and development (R&D) scientists
and engineers to easily access and analyze the wealth of test data it collects, as well as
efficiently monitor its manufacturing processes and supply chains. “Like many companies,
we were data-rich but analysis-poor,” explains John Young, Business Analyst for
Instrumentation Laboratory. “It’s no longer viable to have R&D analysts running off to IT
every time they need access to test data and then doing one-off analyses in Minitab. They
need to be able to access data quickly and perform complex analyses consistently and
accurately.” Implementing sophisticated analytics was especially critical for Instrumentation
Laboratory because of the volume and complexity of its products. For example, every year,
the company manufactures hundreds of thousands of cartridges containing a card with a
variety of sensors that measure the electrical signals of the blood being tested. “Those sensors
are affected by a wide range of factors, from environmental changes like heat and humidity to
inconsistencies in materials from suppliers, so we’re constantly monitoring their
performance,” says Young. “We collect millions of records of data, most of which is stored in
SQL Server databases. We needed an analytics platform that would enable our R&D teams to
quickly access that data and troubleshoot any problems. Plus, because there are so many
factors in play, we also needed a platform that could intelligently monitor the test data and
alert us to emerging issues automatically.”
Solution
Instrumentation Laboratory began looking for an analytics solution to meet its needs. The
company quickly eliminated most tools on the market because they failed to deliver the
statistical functionality and level of trust required for the healthcare environment. That left
only two contenders: another analytics solution and Dell Statistica. For Instrumentation
Laboratory, the clear winner was Statistica. “Choosing Statistica was an easy decision,”
recalls Young. “With Statistica, I was able to quickly build a wide range of analysis
configurations on top of our data for use by analysts enterprise-wide. Now, when they want
to understand specific things, they can simply run a canned analysis from that central store
instead of having to ask IT for access to the data or remember how to do a particular test.”
Moreover, Statistica was far easier to deploy and use than legacy analytics solutions. “To
implement and maintain other analytics solutions, you need to know analytics solutions
programming,” Young notes. “But with Statistica, I can connect to our data, create an
analysis and publish it within an hour—even though I’m not a great programmer.” Finally, in
addition to its advanced functionality and ease of use, Statistica delivered world-class support
and an attractive price point. “The people who helped us implement Statistica were simply
awesome,” says Young. “And the price was far below what another analytics solution was
quoting.”
Results
With Statistica in place, analysts across the enterprise now have easy access to both the data
and the analyses they need to continue the twin traditions of innovation and quality at
Instrumentation Laboratory. In fact, Statistica’s quick, effective analysis and automated
alerting is saving the company hundreds of thousands of dollars. “During cartridge
manufacturing, we occasionally experience problems, such as an inaccuracy in a chemical
formulation that goes on one of the sensors,” Young notes. “Scrapping a single batch of cards
would cost us hundreds of thousands of dollars. Statistica helps us quickly figure out what
went wrong and fix it so we can avoid those costs. For example, we can marry the test data
with electronic device history record data from our SAP environment and perform all sorts of
correlations to determine which particular changes—such as changes in temperature and
humidity—might be driving a particular issue.” Manual quality checks are, of course,
valuable, but Statistica runs a variety of analyses automatically for the company as well,
helping to ensure that nothing is missed and issues are identified quickly. “Many analysis
configurations are scheduled to run periodically to check different things,” Young says.“If
there is an issue, the system automatically emails the appropriate people or logs the violations
to a database.” Some of the major benefits of advanced data analytics with Dell Statistica
included the following:
• Supply chain monitoring. Instrumentation Laboratory manufactures not just the card with
the sensors but the whole medical instrument, and therefore it relies on suppliers to provide
parts. To further ensure quality, the company is planning to extend its use of Statistica to
supply chain monitoring.
• Saving time. In addition to saving money and improving regulatory compliance for
Instrumentation Laboratory, Statistica is also saving the company’s engineers and scientists
valuable time, enabling them to focus more on innovation and less on routine
matters.“Statistica’s proactive alerting saves engineers a lot of time because they don’t have
to remember to check various factors, such as glucose slope, all the time. Just that one test
would take half a day,” notes Young. “With Statistica monitoring our test data, our engineers
can focus on other matters, knowing they will get an email if and when a factor like glucose
slope becomes an issue.”
Future Possibilities
Instrumentation Laboratory is excited about the opportunities made possible by the visibility
Statistica advanced analytics software has provided into its data stores. “Using Statistica, you
can discover all sorts of insights about your data that you might not otherwise be able to
find,” says Young. “There might be hidden pockets of money out there that you’re just not
seeing because you’re not analyzing your data to the extent you could. Using the tool, we’ve
discovered some interesting things in our data that have saved us a tremendous amount of
money, and we look forward to finding even more.”
1. What were the main challenges for the medical device company? Were they market or
technology driven? Explain.
3. What were the results? What do you think was the real return on investment (ROI)?
The first step is handling inconsistencies or errors in the data. For example, unusual values
or mistakes must be corrected using expert knowledge or domain understanding. Once
cleaned, the data often goes through transformation. Common transformations include
normalization (scaling values to a fixed range so that large numbers like income do not
overshadow smaller ones like years of service), discretization (converting numbers into
categories like low, medium, high), or aggregation (grouping values into broader categories,
such as combining individual states into regions). Sometimes, new variables are created to
simplify the dataset, such as combining donor and recipient blood types into a single
match/no-match variable.
Next comes data reduction, which deals with very large datasets. Data is usually represented
in rows (records) and columns (variables). When the number of variables is too high,
dimensionality reduction techniques like Principal Component Analysis (PCA) help
simplify them while keeping important information. When the dataset has too many records
(millions or billions), sampling methods are used to select a smaller but representative subset.
To ensure fairness, random or stratified sampling is preferred over simply picking data from
the top or bottom, which can introduce bias.
Another challenge is imbalanced data, where some categories have far more examples than
others. For better model performance, analysts often balance the data by oversampling the
smaller class or undersampling the larger one. Research shows that balanced data usually
leads to better predictions.
Data preprocessing is one of the most important and time-consuming steps in analytics.
The more effort invested in cleaning, transforming, and reducing the data, the more accurate
and meaningful the results will be. Far from being a waste of time, preprocessing sets the
foundation for successful analytics and ensures trustworthy insights.
Student attrition has become one of the most challenging problems for decision makers in
academic institutions. Despite all the programs and services that are put in place to help retain
students, according to the U.S. Department of Education, Center for Educational Statistics
([Link]), only about half of those who enter higher education actually earn a bachelor’s
degree. Enrollment management and the retention of students has become a top priority for
administrators of colleges and universities in the United States and other countries around the
world. High dropout of students usually results in overall financial loss, lower graduation
rates, and inferior school reputation in the eyes of all stakeholders. The legislators and policy
makers who oversee higher education and allocate funds, the parents who pay for their
children’s education to prepare them for a better future, and the students who make college
choices look for evidence of institutional quality and reputation to guide their decision-
making processes.
Proposed Solution
To improve student retention, one should try to understand the nontrivial reasons behind the
attrition. To be successful, one should also be able to accurately identify those students that
are at risk of dropping out. So far, the vast majority of student attrition research has been
devoted to understanding this complex, yet crucial, social phenomenon. Even though these
qualitative, behavioral, and survey-based studies revealed invaluable insight by developing
and testing a wide range of theories, they do not provide the much-needed instruments to
accurately predict (and potentially improve) student attrition. The project summarized in this
case study proposed a quantitative research approach where the historical institutional data
from student databases could be used to develop models that are capable of predicting as well
as explaining the institution-specific nature of the attrition problem. The proposed analytics
approach is shown in Figure.
Although the concept is relatively new to higher education, for more than a decade now,
similar problems in the field of marketing management have been studied using predictive
data analytics techniques under the name of “churn analysis,” where the purpose has been to
identify among the current customers to answer the question, “Who among our current
customers are more likely to stop buying our products or services?” so that some kind of
The data for this research project came from a single institution (a comprehensive public
university located in the Midwest region of the United States) with an average enrollment of
23,000 students, of which roughly 80% are the residents of the same state and roughly 19%
of the students are listed under some minority classification. There is no significant difference
between the two genders in the enrollment numbers. The average freshman student retention
rate for the institution was about 80%, and the average 6-year graduation rate was about 60%.
The study used 5 years of institutional data, which entailed to 16,000+ students enrolled as
freshmen, consolidated from various and diverse university student databases. The data
contained variables related to students’ academic, financial, and demographic characteristics.
After merging and converting the multidimensional student data into a single flat file (a file
with columns representing the variables and rows representing the student records), the
resultant file was assessed and preprocessed to identify and remedy anomalies and unusable
values. As an example, the study removed all international student records from the data set
because they did not contain information about some of the most reputed predictors (e.g.,
high school GPA, SAT scores). In the data transformation phase, some of the variables were
aggregated (e.g., “Major” and “Concentration” variables aggregated to binary variables
MajorDeclared and ConcentrationSpecified) for better interpretation for the predictive
modeling. In addition, some of the variables were used to derive new variables (e.g.,
Earned/Registered ratio and YearsAfterHighSchool).
The Earned/Registered ratio was created to have a better representation of the students’
resiliency and determination in their first semester of the freshman year. Intuitively, one
would expect greater values for this variable to have a positive impact on retention/
persistence. The YearsAfterHighSchool was created to measure the impact of the time taken
between high school graduation and initial college enrollment. Intuitively, one would expect
this variable to be a contributor to the prediction of attrition. These aggregations and derived
variables are determined based on a number of experiments conducted for a number of
logical hypotheses. The ones that made more common sense and the ones that led to better
prediction accuracy were kept in the final variable set. Reflecting the true nature of the
subpopulation (i.e., the freshmen students), the dependent variable (i.e., “Second Fall
Registered”) contained many more yes records (~80%) than no records (~20%). Research
shows that having such an imbalanced data has a negative impact on model performance.
Therefore, the study experimented with the options of using and comparing the results of the
same type of models built with the original imbalanced data (biased for the yes records) and
the wellbalanced data.
The study employed four popular classification methods (i.e., artificial neural networks,
decision trees, support vector machines, and logistic regression) along with three model
ensemble techniques (i.e., bagging, busting, and information fusion). The results obtained
from all model types were then compared to each other using regular classification model
assessment methods (e.g., overall predictive accuracy, sensitivity, specificity) on the holdout
samples.
Results
In the first set of experiments, the study used the original imbalanced data set. Based on the
10-fold cross-validation assessment results, the support vector machines produced the best
accuracy with an overall prediction rate of 87.23%, the decision tree came out as the runner-
up with an overall prediction rate of 87.16%, followed by artificial neural networks and
logistic regression with overall prediction rates of 86.45% and 86.12%, respectively (see
Table 2.1). A careful examination of these results reveals that the predictions accuracy for the
“Yes” class is significantly higher than the prediction accuracy of the “No” class. In fact, all
four model types predicted the students who are likely to return for the second year with
better than 90% accuracy, but they did poorly on predicting the students who are likely to
drop out after the freshman year with less than 50% accuracy. Because the prediction of the
“No” class is the main purpose of this study, less than 50% accuracy for this class was
deemed not acceptable.
Such a difference in prediction accuracy of the two classes can (and should) be attributed to
the imbalanced nature of the training data set (i.e., ~80% “Yes” and ~20% “No” samples).
The next round of experiments used a wellbalanced data set where the two classes are
represented nearly equally in counts. In realizing this approach, the study took all the samples
from the minority class (i.e., the “No” class herein) and randomly selected an equal number
of samples from the majority class (i.e., the “Yes” class herein) and repeated this process for
10 times to reduce potential bias of random sampling. Each of these sampling processes
resulted in a data set of 7,000+ records, of which both class labels (“Yes” and “No”) were
equally represented. Again, using a 10-fold crossvalidation methodology, the study
developed and tested prediction models for all four model types. The results of these
experiments are shown in Table 2.2. Based on the holdout sample results, support vector
machines once again generated the best overall prediction accuracy with 81.18%, followed by
decision trees, artificial neural networks, and logistic regression with an overall prediction
accuracy of 80.65%, 79.85%, and 74.26%. As can be seen in the per-class accuracy figures,
the prediction models did significantly better on predicting the “No” class with the well-
balanced data than they did with the unbalanced data. Overall, the three machine-learning
techniques performed significantly better than their statistical counterpart, logistic regression.
Next, another set of experiments were conducted to assess the predictive ability of the three
ensemble models. Based on the 10-fold crossvalidation methodology, the information fusion–
type ensemble model produced the best results with an overall prediction rate of 82.10%,
followed by the bagging-type ensembles and boosting-type ensembles with overall prediction
rates of 81.80% and 80.21%, respectively (see Table 2.3). Even though the prediction results
are slightly better than the individual models, ensembles are known to produce more robust
prediction systems compared to a single-best prediction model.
In addition to assessing the prediction accuracy for each model type, a sensitivity analysis
was also conducted using the developed prediction models to identify the relative importance
of the independent variables (i.e., the predictors). In realizing the overall sensitivity analysis
results, each of the four individual model types generated its own sensitivity measures
ranking all the independent variables in a prioritized list. As expected, each model type
generated slightly different sensitivity rankings of the independent variables. After collecting
all four sets of sensitivity numbers, the sensitivity numbers are normalized and aggregated
and plotted in a horizontal bar chart.
Conclusions
The study showed that, given sufficient data with the proper variables, data mining methods
are capable of predicting freshmen student attrition with approximately 80% accuracy.
Results also showed that, regardless of the prediction model employed, the balanced data set
(compared to unbalanced/ original data set) produced better prediction models for identifying
the students who are likely to drop out of the college prior to their sophomore year. Among
the four individual prediction models used in this study, support vector machines performed
the best, followed by decision trees, neural networks, and logistic regression. From the
usability standpoint, despite the fact that support vector machines showed better prediction
results, one might choose to use decision trees because compared to support vector machines
and neural networks, they portray a more transparent model structure. Decision trees
explicitly show the reasoning process of different predictions, providing a justification for a
specific outcome, whereas support vector machines and artificial neural networks are
mathematical models that do not provide such a transparent view of “how they do what they
do.”
2. What were the traditional methods to deal with the attrition problem?
3. List and discuss the data-related challenges within context of this case study.
4. What was the proposed solution? And, what were the results?
Big Data refers to the large and complex sets of data that are difficult to process using
traditional tools and methods. Businesses today generate and collect massive amounts of
information from different sources, making it challenging to manage, store, and analyze.
While the term is often overused as a buzzword, at its core, Big Data is about uncovering new
value and insights from both conventional and unconventional data sources.
Data comes from almost everywhere: social media, web logs, GPS, RFID tags, sensors,
medical records, scientific experiments, videos, photos, and even online shopping activities.
Previously, these sources were ignored because of technical limitations, but modern
technologies have turned them into valuable assets for organizations.
Big Data is not entirely new. Companies have been working with large datasets since the
1990s when data warehouses were introduced. What has changed is the size and structure of
data. Earlier, terabytes of data seemed massive, but now organizations deal with zettabytes.
The growth of the Internet of Things (IoT) and sensors means that data will continue to grow
at a rapid pace.
Big Data is often explained using several key characteristics, known as the “V”s:
The promise of Big Data is its ability to provide deeper insights and better decision-making.
By analyzing massive and diverse data sets, organizations can identify patterns, trends, and
opportunities that would not be visible in smaller datasets. This is why businesses and
researchers are investing heavily in Big Data technologies and analytics.
Getting a good forecast and understanding of the situation is crucial for any scenario, but it is
especially important to players in the investment industry. Being able to get an early
indication of how a particular retailer’s sales are doing can give an investor a leg up on
whether to buy or sell that retailer’s stock even before the earnings reports are released. The
problem of forecasting economic activity or microclimates based on a variety of data beyond
the usual retail data is a very recent phenomenon and has led to another buzzword—
“alternative data.” A major mix in this alternative data category is satellite imagery, but it
also includes other data such as social media, government filings, job postings, traffic
patterns, changes in parking lots or open spaces detected by satellite images, mobile phone
usage patterns in any given location at any given time, search patterns on search engines, and
so on. Facebook and other companies have invested in satellites to try to image the whole
globe every day so that daily changes can be tracked at any location and the information can
be used for forecasting. In the last 6 to 12 months, many interesting examples of more
reliable and advanced forecasts have been reported. Indeed, this activity is being led by start-
up companies. Here are some of the examples:
• Facebook used its image recognition engine to analyze over 14.6 billion images of every
corner of the world to identify areas of low connectivity.
• RS Metrics monitored parking lots across the United States for various hedge funds. In
2015, based on an analysis of the parking lots, RS Metrics predicted a strong second quarter
in 2015 for JC Penney. Its clients (mostly hedge funds) profited from this advanced insight. A
similar story has been reported for Wal-Mart using car counts in its parking lots to forecast
sales.
• Orbital Insights uses satellite imagery data to provide macroeconomic indicators for various
industry sectors. For example, by analyzing shadows of the oil storage tanks around the
world, it claims to have produced a better daily estimate of worldwide oil storage than is
available from the International Energy Agency (IEA).
• Spaceknow keeps track of changes in factory surroundings for over 6,000 Chinese factory
sites. Using this data, the company has been able to provide a better idea of China’s industrial
economic activity than what the Chinese government has been reporting.
• Descartes Labs uses satellite data to predict U.S. corn harvests with more accuracy than the
U.S. Department of Agriculture does. Better forecasts can have huge financial impacts on
futures trading. An older example of this was a company called Lanworth that also predicted
corn crop estimates. Lanworth was acquired by Thomson Reuters and is integrated in their
Eikon service.
• DigitalGlobe is able to analyze the size of a forest with more accuracy because its software
can count every single tree in a forest. This results in a more accurate estimate because there
is no need to use a representative sample.
• Kensho, a company backed by Goldman Sachs, is reportedly analyzing data from multiple
sources (mentioned earlier) to build a trading engine. These examples illustrate just a sample
of ways data can be combined to generate new insights. Of course, there are privacy concerns
in some cases. For example, a story in the Wall Street Journal in 2015 reported that Yodlee, a
company that provides personal finance tools to many large banks and thus has access to
millions of customers’ credit card transactions, sells such data to other analytics firms that
can use the information to develop early predictions of how sales are trending for a particular
retailer.
Such information is highly sought by stock market traders. This story led to an uproar about
the customer information being used in ways not authorized. There is also a concern in some
circles about the legality of developing such advanced predictions about a particular
commodity or company. Although such concerns will eventually be resolved by policy
makers, what is clear is that new and interesting ways of combining satellite data and many
other data sources are spawning a new crop of analytics companies. All of these
organizations are working with data that meets the three V’s—variety, volume, and velocity
characterizations. Some of these companies also work with another category of data—
sensors. We will discuss those in the next chapter when we review emerging trends in
analytics. But this group of companies certainly also falls under a group of innovative and
emerging applications.
2. Can you think of other data streams that might help give an early indication of sales at a
retailer?
3. Can you think of other applications along the lines presented in this application case?
However, managing Big Data is not easy. Traditional systems for storing and analyzing data
are not capable of handling the volume, variety, and velocity of data we see today.
Organizations must adopt new technologies and strategies to take on this challenge. For
example, if a company cannot process all the data it collects, or wants to include new data
sources like social media or sensor data, then it must consider Big Data solutions.
To succeed with Big Data analytics, organizations need certain key factors:
1. Clear business need – Big Data projects should be driven by business goals, not just
technology.
4. Fact-based culture – Decisions should be based on data and evidence, not gut
feeling.
High-Performance Computing
Because Big Data requires heavy computation, new techniques have been developed:
In-database analytics: Performs analysis inside databases to save time and avoid
moving data.
Despite the challenges, Big Data analytics provides huge value. It helps improve process
efficiency, cost reduction, and customer experience across industries. For example,
manufacturing and healthcare use Big Data for efficiency, while retail and insurance focus on
customer experience. Banks and schools often use it for risk management. Other benefits
include: brand management, revenue growth, churn reduction, better customer service, new
market opportunities, regulatory compliance, and stronger security.
Application Case 2.4 Top Five Investment Bank Achieves Single Source of the Truth
The bank’s highly respected derivatives team is responsible for over one-third of the world’s
total derivatives trades. Their derivatives practice has a global footprint with teams that
support credit, interest rates, and equity derivatives in every region of the world. The bank
has earned numerous industry awards and is recognized for its product innovations.
Challenge
With its significant derivatives exposure, the bank’s management recognized the importance
of having a real-time global view of its positions. The existing system, based on a relational
database, was comprised of multiple installations around the world. Due to the gradual
expansions to accommodate the increasing data volume varieties, the legacy system was not
fast enough to respond to growing business needs and requirements. It was unable to deliver
real-time alerts to manage market and counterparty credit positions in the desired time frame.
Solution
The bank built a derivatives trade store based on the MarkLogic (a Big Data analytics
solution provider) Server, replacing the incumbent technologies. Replacing the 20 disparate
batch-processing servers with a single operational trade store enabled the bank to know its
market and credit counterparty positions in real time, providing the ability to act quickly to
mitigate risk. The accuracy and completeness of the data allowed the bank and its regulators
to confidently rely on the metrics and stress test results it reports. The selection process
included upgrading existing Oracle and Sybase technology. Meeting all the new regulatory
requirements was also a major factor in the decision as the bank looked to maximize its
investment. After the bank’s careful investigation, the choice was clear—only MarkLogic
could meet both needs plus provide better performance, scalability, faster development for
future requirements and implementation, and a much lower total cost of ownership. Below
Figure illustrates the transformation from the old fragmented systems to the new unified
system.
FIGURE: Moving from Many Old Systems to a Unified New System. (Source: MarkLogic.)
Results
MarkLogic was selected because existing systems would not provide the subsecond updating
and analysis response times needed to effectively manage a derivatives trade book that
represents nearly one-third of the global market. Trade data is now aggregated accurately
across the bank’s entire derivatives portfolio, allowing risk management stakeholders to
know the true enterprise risk profile, to conduct predictive analyses using accurate data, and
to adopt a forwardlooking approach. Not only are hundreds of thousands of dollars of
technology costs saved each year, but the bank does not need to add resources to meet
regulators’ escalating demands for more transparency and stress-testing frequency. Here are
the highlights:
• An alerting feature keeps users appraised of upto-the-minute market and counterparty credit
changes so they can take appropriate actions.
• Derivatives are stored and traded in a single MarkLogic system requiring no downtime for
maintenance, a significant competitive advantage.
• Complex changes can be made in hours versus days, weeks, and even months needed by
competitors.
• Replacing Oracle and Sybase significantly reduced operations costs: one system versus 20,
one database administrator instead of up to 10, and lower costs per trade.
Next Steps
The successful implementation and performance of the new system resulted in the bank’s
examination of other areas where it could extract more value from its Big Data—structured,
unstructured, and/or polystructured. Two applications are under active discussion. Its equity
research business sees an opportunity to significantly boost revenue with a platform that
provides real-time research, repurposing, and content delivery. The bank also sees the power
of centralizing customer data to improve onboarding, increase cross-selling opportunities, and
support know-your-customer requirements.
2. How did the MarkLogic infrastructure help ease the leveraging of Big Data?
3. What were the challenges, the proposed solution, and the obtained results?
MapReduce
MapReduce, developed by Google, is a programming model that splits large data into smaller
chunks and processes them in parallel across many machines. It works in two stages: the
Map step organizes the data (e.g., grouping colored squares), and the Reduce step
summarizes results (e.g., counting each color). Programmers only write the map and reduce
functions; the system handles parallelization. It is widely used for tasks like text analysis,
machine learning, and indexing. Libraries like Apache Mahout provide ready-to-use
MapReduce-based algorithms.
Hadoop
Hadoop Ecosystem
The biggest advantage of Hadoop is its ability to process massive volumes of unstructured
and semi-structured data at low cost. It scales to petabytes or exabytes and lets data scientists
analyze all available data, not just samples. However, Hadoop is still evolving and requires
skilled experts to set up and manage. It is mainly batch-oriented, so it is not ideal for real-
time processing.
NoSQL
NoSQL databases emerged to handle Big Data in real-time applications. Unlike relational
databases, NoSQL systems manage large volumes of varied data while delivering fast
performance. Examples include MongoDB, Cassandra, HBase, CouchDB, and
DynamoDB. Many NoSQL databases trade strict ACID compliance for speed and scalability.
Some, like HBase, can work with Hadoop to provide low-latency lookups on massive
datasets.
eBay is the world’s largest online marketplace, enabling the buying and selling of practically
anything. Founded in 1995, eBay connects a diverse and passionate community of individual
buyers and sellers, as well as small businesses. eBay’s collective impact on e-commerce is
staggering: In 2012, the total value of goods sold on eBay was $75.4 billion. eBay currently
serves over 112 million active users and 400+ million items for sale.
One of the keys to eBay’s extraordinary success is its ability to turn the enormous volumes of
data it generates into useful insights that its customers can glean directly from the pages they
frequent. To accommodate eBay’s explosive data growth—its data centers perform billions of
reads and writes each day—and due to the increasing demand to process data at blistering
speeds, eBay needed a solution that did not have the typical bottlenecks, scalability issues,
and transactional constraints associated with common relational database approaches. The
company also needed to perform rapid analysis on a broad assortment of the structured and
unstructured data it captured.
Its Big Data requirements brought eBay to NoSQL technologies, specifically Apache
Cassandra and DataStax Enterprise. Along with Cassandra and its high-velocity data
capabilities, eBay was also drawn to the integrated Apache Hadoop analytics that come with
DataStax Enterprise. The solution incorporates a scale-out architecture that enables eBay to
deploy multiple DataStax Enterprise clusters across several different data centers using
commodity hardware. The end result is that eBay is now able to more cost effectively process
massive amounts of data at very high speeds, at very high velocities, and achieve far more
than they were able to with the higher cost proprietary system they had been using. Currently,
eBay is managing a sizable portion of its data center needs—250TBs+ of storage—in Apache
Cassandra and DataStax Enterprise clusters. Additional technical factors that played a role in
eBay’s decision to deploy DataStax Enterprise so widely include the solution’s linear
scalability, high availability with no single point of failure, and outstanding write
performance.
eBay employs DataStax Enterprise for many different use cases. The following examples
illustrate some of the ways the company is able to meet its Big Data needs with the extremely
fast data handling and analytics capabilities the solution provides. Naturally, eBay
experiences huge amounts of write traffic, which the Cassandra implementation in DataStax
Enterprise handles more efficiently than any other RDBMS or NoSQL solution. eBay
currently sees 6 billion+ writes per day across multiple Cassandra clusters and 5 billion+
reads (mostly offline) per day as well. One use case supported by DataStax Enterprise
involves quantifying the social data eBay displays on its product pages. The Cassandra
distribution in DataStax Enterprise stores all the information needed to provide counts for
“like,” “own,” and “want” data on eBay product pages. It also provides the same data for the
eBay “Your Favorites” page that contains all the items a user likes, owns, or wants, with
Cassandra serving up the entire “Your Favorites” page. eBay provides this data through
Cassandra’s scalable counters feature. Load balancing and application availability are
important aspects to this particular use case. The DataStax Enterprise solution gave eBay
architects the flexibility they needed to design a system that enables any user request to go to
any data center, with each data center having a single DataStax Enterprise cluster spanning
those centers. This design feature helps balance the incoming user load and eliminates any
possible threat to application downtime. In addition to the line of business data powering the
Web pages its customers visit, eBay is also able to perform high-speed analysis with the
ability to maintain a separate data center running Hadoop nodes of the same DataStax
Enterprise ring.
Another use case involves the Hunch (an eBay sister company) “taste graph” for eBay users
and items, which provides customer recommendations based on user interests. eBay’s Web
site is essentially a graph between all users and the items for sale. All events (bid, buy, sell,
and list) are captured by eBay’s systems and stored as a graph in Cassandra. The application
sees more than 200 million writes daily and holds more than 40 billion pieces of data. eBay
also uses DataStax Enterprise for many time-series use cases in which processing
highvolume, real-time data is a foremost priority. These include mobile notification logging
and tracking (every time eBay sends a notification to a mobile phone or device it is logged in
Cassandra), fraud detection, SOA request/response payload logging, and RedLaser (another
eBay sister company) server logs and analytics. Across all of these use cases is the common
requirement of uptime. eBay is acutely aware of the need to keep their business up and open
for business, and DataStax Enterprise plays a key part in that through its support of high
availability clusters. “We have to be ready for disaster recovery all the time. It’s really great
that Cassandra allows for active-active multiple data centers where we can read and write
data anywhere, anytime,” says eBay architect Jay Patel.
2. What were the challenges, the proposed solution, and the obtained results?
On the Internet today all users have the power to contribute as well as consume information.
This power is used in many ways. On social network platforms such as Twitter, users are able
to post information about their health condition as well as receive help on how best to
manage those health conditions. Many users have wondered about the quality of information
disseminated on social network platforms. Whereas the ability to author and disseminate
health information on Twitter seems valuable to many users who use it to seek support for
their disease, the authenticity of such information, especially when it originates from lay
individuals, has been in doubt. Many users have asked, “How do I verify and trust
information from nonexperts about how to manage a vital issue like my health condition?”
What types of users share and discuss what type of information? Do users with a large
following discuss and share the same type of information as users with a smaller following?
The number of followers of a user relate to the influence of a user. Characteristics of the
information are measured in terms of quality and objectivity of the Tweet posted. A team of
data scientists set out to explore the relationship between the number of followers a user had
and the characteristics of information the user disseminated (Asamoah & Sharda, 2015).
Solution
Data was extracted from the Twitter platform using Twitter’s API. The data scientists adapted
the knowledge-discovery and data management model to manage and analyze this large set of
data. The model was optimized for managing and analyzing Big Data derived from a social
network platform and included phases for gaining domain knowledge, developing an
appropriate Big Data platform, data acquisition and storage, data cleaning, data validation,
data analysis, and results and deployment.
Technology Used
The tweets were extracted, managed, and analyzed using Cloudera’s distribution of the
Apache Hadoop. The Apache Hadoop framework has several subprojects that support
different kinds of data management activities. For instance, the Apache Hive subproject
supported the reading, writing, and managing of the large tweet data. Data analytics tools
such as Gephi were used for social network analysis and R for predictive modeling. They
conducted two parallel analyses; social network analysis to understand the influence network
on the platform and text mining to understand the content of tweets posted by users. What
Was Found? As noted earlier, tweets from both influential and noninfluential users were
collected and analyzed. The results showed that the quality and objectivity of information
disseminated by influential users was higher than that disseminated by noninfluential users.
They also found that influential users controlled the flow of information in a network and that
other users were more likely to follow their opinion on a subject. There was a clear difference
between the type of information support provided by influential users versus the others.
Influential users discussed more objective information regarding the disease management—
things such as diagnoses, medications, formal therapies. Noninfluential users provided more
information about emotional support and alternative ways of coping with such diseases. Thus
a clear difference between influential users and the others was evident. From the nonexperts’
perspective, the data scientists portray how healthcare provision can be augmented by helping
patients identify and use valuable resources on the Web for managing their disease condition.
This work also helps identify how nonexperts can locate and filter healthcare information that
may not necessarily be beneficial to the management of their health condition.
1. What was the data scientists’ main concern regarding health information that is
disseminated on the Twitter platform?
2. How did the data scientists ensure that nonexpert information disseminated on social media
could indeed contain valuable health information?
3. Does it make sense that influential users would share more objective information whereas
less influential users could focus more on subjective information? Why?
Stream analytics, also called real-time analytics or data-in-motion analytics, is the process of
analyzing data while it is being created and flowing into the system. Unlike traditional
methods where data is stored first and analyzed later, stream analytics extracts useful insights
immediately. A stream is a continuous flow of data elements (often called tuples), which can
be compared to rows in a database. Sometimes, to get meaningful insights, we use a
window—a small set of recent data—so that patterns and correlations can be detected
quickly.
Although the terms sound similar, stream and perpetual analytics are slightly different.
Streaming analytics looks at data within a short, moving window (e.g., the last 5 seconds or
last 10,000 records). Perpetual analytics, however, compares each new observation with all
past data, without limiting to a window. Streaming is faster and better for high-volume data,
while perpetual analytics is more powerful for critical tasks where every detail matters.
One of the key uses of stream analytics is critical event processing. This involves detecting
important or unusual events in real time, such as fraud, system failures, or cyberattacks. By
analyzing multiple data sources together, organizations can predict or respond to events
instantly. For example, an e-commerce site can offer personalized promotions while you
browse, or a bank can block a suspicious transaction as it happens.
Data stream mining focuses on finding patterns and making predictions from fast-flowing
data. Unlike traditional data mining, which works on stored datasets, stream mining must
handle continuous, high-speed data that can only be processed once or a few times. Examples
include analyzing sensor data, financial transactions, or web clicks. Machine learning models,
such as decision trees, are adapted for streaming environments to make real-time predictions.
Energy Industry: Smart grids use streaming data from meters and sensors to balance
supply and demand in real time.
Financial Services: Banks and trading firms analyze streaming market data to make
split-second buy/sell decisions.
Healthcare: Medical devices generate continuous data (like ECG or blood pressure
readings) that can be analyzed instantly to save lives.
Application Case 2.7 Salesforce Is Using Streaming Data to Enhance Customer Value
Salesforce has expanded their Marketing Cloud services to include Predictive Scores and
Predictive Audience features called the Marketing Cloud Predictive Journey. This addition
uses real-time streaming data to enhance the customer engagement online. First, the
customers are given a Predictive Score unique to them. This score is calculated from several
different factors, including how long their browsing history is, if they clicked an e-mail link,
if they made a purchase, how much they spent, how long ago did they make a purchase, or if
they have ever responded to an e-mail or ad campaign. Once customers have a score, they are
then segmented into different groups. These groups are given different marketing objectives
and plans based on the predictive behaviors assigned to them. The scores and segments are
updated and changed daily and give companies a better road map to target and achieve a
desired response. These marketing solutions are more accurate and create more personalized
ways companies can accommodate their customer retention methods.
2. Besides customer retention, what are other benefits of using predictive analytics?
Through the analysis of data acquired in the here and now, companies are able to make
predictions and decisions about their consumers more rapidly. This ensures that businesses
target, attract, and retain the right customers and maximize their value. Data acquired last
week is not as beneficial as the data companies have today. Using relevant data makes our
predictive analysis more accurate and efficient.
Mean (average): The sum of all values divided by the number of values. Commonly
used but sensitive to outliers.
Median: The middle value when data is sorted. Best when data has outliers or is
skewed.
Mode: The most frequent value. Useful for categorical or nominal data.
Each measure has its use, but it’s often best to look at all three together for a clearer picture.
Measures of Dispersion
Standard deviation: Square root of variance, shows how far data points are from the
mean.
Mean Absolute Deviation (MAD): Average absolute difference from the mean.
Quartiles & Interquartile Range (IQR): Spread of the middle 50% of the data,
useful for skewed data.
Dispersion is important because it shows whether the mean or other central measures truly
represent the data.
A box plot graphically shows both central tendency (median, sometimes mean) and
dispersion (IQR, range, outliers). It’s easy to interpret and now widely used in business
analytics.
Shape of a Distribution
Skewness: Shows if data leans left or right. Positive skew = longer right tail, negative
skew = longer left tail.
Kurtosis: Shows if the distribution is more peaked (tall and skinny) or flat compared
to a normal distribution.
Descriptive and inferential statistics can be calculated using software like SAS, SPSS,
Minitab, R, Excel, etc. Excel is the easiest tool for most business users to compute these
statistics.
Application Case 2.8 Town of Cary Uses Analytics to Analyze Data from Sensors,
Assess Demand, and Detect Problems
A leaky faucet. A malfunctioning dishwasher. A cracked sprinkler head. These are more than
just a headache for a home owner or business to fix. They can be costly, unpredictable, and,
unfortunately, hard to pinpoint. Through a combination of wireless water meters and a data-
analytics-driven, customeraccessible portal, the Town of Cary, North Carolina, is making it
much easier to find and fix water loss issues. In the process, the town has gained a bigpicture
view of water usage critical to planning future water plant expansions and promoting targeted
conservation efforts. When the Town of Cary installed the wireless meters for 60,000
customers in 2010, it knew the new technology wouldn’t just save money by eliminating
manual monthly readings; the town also realized it would get more accurate and timely
information about water consumption. The Aquastar wireless system reads meters once an
hour—that’s 8,760 data points per customer each year instead of 12 monthly readings. The
data had tremendous potential, if it could be easily consumed. “Monthly readings are like
having a gallon of water’s worth of data. Hourly meter readings are more like an Olympic-
size pool of data,” says Karen Mills, Finance Director for the Town of Cary. “SAS helps us
manage the volume of that data nicely.” In fact, the solution enables the town to analyze a
halfbillion data points on water usage and make them available, and easily consumable, to all
customers. The ability to visually look at data by household or commercial customer, by the
hour, has led to some very practical applications:
• Customers can set alerts that notify them within hours if there is a spike in water usage.
• Customers can track their water usage online, helping them to be more proactive in
conserving water.
Through the online portal, one business in the Town of Cary saw a spike in water
consumption on weekends, when employees are away. This seemed odd, and the unusual
reading helped the company learn that a commercial dishwasher was malfunctioning, running
continuously over weekends. Without the wireless water-meter data and the customer-
accessible portal, this problem could have gone unnoticed, continuing to waste water and
money. The town has a much more accurate picture of daily water usage per person, critical
for planning future water plant expansions. Perhaps the most interesting perk is that the town
was able to verify a hunch that has far-reaching cost ramifications: Cary residents are very
economical in their use of water. “We calculate that with modern high-efficiency appliances,
indoor water use could be as low as 35 gallons per person per day. Cary residents average
45 gallons, which is still phenomenally low,” explains town Water Resource Manager Leila
Goodwin. Why is this important? The town was spending money to encourage water
efficiency—rebates on low-flow toilets or discounts on rain barrels. Now it can take a more
targeted approach, helping specific consumers understand and manage both their indoor and
outdoor water use. SAS was critical not just for enabling residents to understand their water
use, but also in working behind the scenes to link two disparate databases. “We have a billing
database and the meter-reading database. We needed to bring that together and make it
presentable,” Mills says. The town estimates that by just removing the need for manual
readings, the Aquastar system will save more than $10 million above the cost of the project.
But the analytics component could provide even bigger savings. Already, both the town and
individual citizens have saved money by catching water leaks early. As the Town of Cary
continues to plan its future infrastructure needs, having accurate information on water usage
will help it invest in the right amount of infrastructure at the right time. In addition,
understanding water usage will help the town if it experiences something detrimental like a
drought. “We went through a drought in 2007,” says Goodwin. “If we go through another, we
have a plan in place to use Aquastar data to see exactly how much water we are using on a
day-by-day basis and communicate with customers. We can show ‘here’s what’s happening,
and here is how much you can use because our supply is low.’ Hopefully, we’ll never have to
use it, but we’re prepared.”
4. What other problems and data analytics solutions do you foresee for towns like Cary?
Although both correlation and regression deal with the relationship between variables, they
are different. Correlation only measures the strength and direction of association between two
variables, without assuming one causes the other. Regression, on the other hand, assumes that
one or more variables influence (cause changes in) the output variable and models that
relationship with an equation.
Simple regression: The relationship between one input (e.g., height) and one output
(e.g., weight).
Multiple regression: Involves two or more inputs (e.g., height, gender, BMI) to
predict the output.
Both aim to fit a straight-line equation to the data using a method called ordinary least
squares (OLS), which minimizes the difference between actual and predicted values.
Not every regression model gives good results. To check its quality, we use measures like:
R² (R-squared) – shows how much of the variability in the output is explained by the
model (ranges from 0 to 1).
RMSE (Root Mean Square Error) – measures the average error in predictions.
4. Constant variance (homoscedasticity) – errors should have the same spread across
all input values.
Violating these assumptions can weaken the model, but there are techniques to detect and
correct issues.
Logistic Regression
Unlike linear regression, logistic regression is used when the output variable is categorical
(e.g., yes/no, pass/fail). It predicts the probability of an outcome rather than a continuous
value. The logistic function transforms input values into probabilities between 0 and 1.
Logistic regression is especially popular in medicine, social sciences, and classification
problems like fraud detection or student pass/fail prediction.
Sometimes, the only useful input is time. A time series is a sequence of values measured
over time (e.g., daily stock prices, monthly sales). Time series forecasting uses past patterns
(trends, seasonality, randomness) to predict future values.
Exponential smoothing.
The accuracy of forecasts is usually measured with error metrics such as MAE, MSE, or
MAPE.
Predicting the outcome of a college football game (or any sports game, for that matter) is an
interesting and challenging problem. Therefore, challengeseeking researchers from both
academics and industry have spent a great deal of effort on forecasting the outcome of
sporting events. Large quantities of historic data exist in different media outlets (often
publicly available) regarding the structure and outcomes of sporting events in the form of a
variety of numerically or symbolically represented factors that are assumed to contribute to
those outcomes. The end-of-season bowl games are very important to colleges both
financially (bringing in millions of dollars of additional revenue) as well as reputational—for
recruiting quality students and highly regarded high school athletes for their athletic programs
(Freeman & Brewer, 2016). Teams that are selected to compete in a given bowl game split a
purse, the size of which depends on the specific bowl (some bowls are more prestigious and
have higher payouts for the two teams), and therefore securing an invitation to a bowl game
is the main goal of any division I-A college football program. The decision makers of the
bowl games are given the authority to select and invite bowl-eligible (a team that has six wins
against its Division I-A opponents in that season) successful teams (as per the ratings and
rankings) that will play in an exciting and competitive game, attract fans of both schools, and
keep the remaining fans tuned in via a variety of media outlets for advertising. In a recent
data mining study, Delen, Cogdell, and Kasap (2012) used 8 years of bowl game data along
with three popular data mining techniques (decision trees, neural networks, and support
vector machines) to predict both the classification-type outcome of a game (win versus loss)
as well as the regression-type outcome (projected point difference between the scores of the
two opponents). What follows is a shorthand description of their study.
Methodology
In this research, Delen and his colleagues followed a popular data mining methodology called
CRISP-DM (Cross-Industry Standard Process for Data Mining), which is a six-step process.
This popular methodology, which is covered in detail in Chapter 4, provided them with a
systematic and structured way to conduct the underlying data mining study and hence
improved the likelihood of obtaining accurate and reliable results. To objectively assess the
prediction power of the different model types, they used a cross-validation methodology,
called k-fold crossvalidation. Below Figure graphically illustrates the methodology employed
by the researchers.
The sample data for this study is collected from a variety of sports databases available on the
Web, including [Link], [Link], [Link], ncaa. org, and [Link]. The
data set included 244 bowl games, representing a complete set of eight seasons of college
football bowl games played between 2002 and 2009. We also included an outof-sample data
set (2010–2011 bowl games) for additional validation purposes. Exercising one of the
popular data mining rules-of-thumb, they included as much relevant information into the
model as possible. Therefore, after an in-depth variable identification and collection process,
they ended up with a data set that included 36 variables, of which the first 6 were the
identifying variables (i.e., name and the year of the bowl game, home and away team names
and their athletic conferences—see variables 1–6 in Table), followed by 28 input variables
(which included variables delineating a team’s seasonal statistics on offense and defense,
game outcomes, team composition characteristics, athletic conference characteristics, and
how they fared against the odds— see variables 7–34 in Table 2.4), and finally the last two
were the output variables (i.e., ScoreDiff—the score difference between the home team and
the away team represented with an integer number, and WinLoss—whether the home team
won or lost the bowl game represented with a nominal label). In the formulation of the data
set, each row (a.k.a. tuple, case, sample, example, etc.) represented a bowl game, and each
column stood for a variable (i.e., identifier/input or output type). To represent the game-
related comparative characteristics of the two opponent teams, in the input variables, we
calculated and used the differences between the measures of the home and away teams. All
these variable values are calculated from the home team’s perspective. For instance, the
variable PPG (average number of points a team scored per game) represents the difference
between the home team’s PPG and away team’s PPG. The output variables represent whether
the home team wins or loses the bowl game. That is, if the ScoreDiff variable takes a positive
integer number, then the home team is expected to win the game by that margin, otherwise (if
the ScoreDiff variable takes a negative integer number) then the home team is expected to
lose the game by that margin. In the case of WinLoss, the value of the output variable is a
binary label, “Win” or “Loss” indicating the outcome of the game for the home team.
In this study, three popular prediction techniques are used to build models (and to compare
them to each other): artificial neural networks, decision trees, and support vector machines.
These prediction techniques are selected based on their capability of modeling both
classification as well as regression-type prediction problems and their popularity in recently
published data mining literature. To compare predictive accuracy of all models to one
another, the researchers used a stratified k-fold cross-validation methodology. In a stratified
version of k-fold cross-validation, the folds are created in a way that they contain
approximately the same proportion of predictor labels (i.e., classes) as the original data set. In
this study, the value of k is set to 10 (i.e., the complete set of 244 samples are split into 10
subsets, each having about 25 samples), which is a common practice in predictive data
mining applications. A graphical depiction of the 10-fold cross-validations was shown earlier
in this chapter. To compare the prediction models that were developed using the
aforementioned three data mining techniques, the researchers chose to use three common
performance criteria: accuracy, sensitivity, and specificity. The simple formulas for these
metrics were also explained earlier in this chapter. The prediction results of the three
modeling techniques are presented in Table 2.5 and Table 2.6. Table 2.6 presents the 10-fold
cross-validation results of the classification methodology where the three data mining
techniques are formulated to have a binary-nominal output variable (i.e., WinLoss). Table 2.6
presents the 10-fold cross-validation results of the regression-based classification
methodology, where the three data mining techniques are formulated to have a numerical
output variable (i.e., ScoreDiff). In the regression-based classification prediction, the
numerical output of the models is converted to a classification type by labeling the positive
WinLoss numbers with a “Win” and negative WinLoss numbers with a “Loss,” and then
tabulating them in the confusion matrixes. Using the confusion matrices, the overall
prediction accuracy, sensitivity, and specificity of each model type are calculated and
presented in these two tables. As the results indicate, the classification-type prediction
methods performed better than regression-based classification-type prediction methodology.
Among the three data mining technologies, classification and regression trees produced better
prediction accuracy in both prediction methodologies. Overall, classification and regression
tree classification models produced a 10-fold crossvalidation accuracy of 86.48%, followed
by support vector machines (with a 10-fold cross-validation accuracy of 79.51%) and neural
networks (with a 10-fold cross-validation accuracy of 75.00%). Using a t-test, researchers
found that these accuracy values were significantly different at 0.05 alpha level, that is, the
decision tree is a significantly better predictor of this domain than the neural network and
support vector machine, and the support vector machine is a significantly better predictor
than neural networks. The results of the study showed that the classification-type models
predict the game outcomes better than regression-based classification models. Even though
these results are specific to the application domain and the data used in this study, and
therefore should not be generalized beyond the scope of the study, they are exciting because
decision trees are not only the best predictors but also the best in understanding and
deployment, compared to the other two machine-learning techniques employed in this study.
More details about this study can be found in Delen et al. (2012).
1. What are the foreseeable challenges in predicting sporting event outcomes (e.g., college
bowl games)?
2. How did the researchers formulate/design the prediction problem (i.e., what were the
inputs and output, and what was the representation of a single sample—row of data)?
3. How successful were the prediction results? What else can they do to improve the
accuracy?
Big Data complexity challenges traditional systems due to its massive volume, diverse variety, rapid velocity, and need for high veracity. Traditional systems struggle with processing speeds and real-time analysis. Necessary strategies include adopting modern technologies like in-memory processing, robust data infrastructures, and data governance frameworks, along with strategies for integrating high-performance computing and Big Data tools to manage complexity effectively .
The study targeted the class imbalance because it significantly affected prediction accuracy, particularly for the 'No' class—students likely to drop out. By balancing the dataset, both classes had equal representation, leading to improved prediction performance of the 'No' class. This balance addressed biased results from the original imbalanced dataset, where models had high accuracy for predicting 'Yes' but performed poorly on 'No' predictions .
Researchers compared model efficacy using a balanced data set to predict freshman student attrition. They employed a 10-fold cross-validation methodology across various models. Support vector machines yielded the best overall accuracy followed by decision trees, neural networks, and logistic regression. Decision trees, albeit slightly less accurate, offer more transparency, making them preferable for their ease of interpretation .
The sensitivity of predictor variables was assessed by conducting leave-one-out assessments to measure the sensitivity of models to specific variables. The ratio of the model's error without the variable to the error with the variable indicates sensitivity. This analysis is crucial as it identifies the relative importance of variables and helps in refining models by understanding which inputs most significantly affect outcomes .
Ensemble models, such as information fusion, bagging, and boosting, offered improved predictive accuracy by combining the strengths of different individual predictions, leading to more robust systems. Ensembles outperformed individual models by providing diverse perspectives on the data. Pros include enhanced accuracy and robustness, while cons may involve increased computational costs and complexity in model interpretation .
Transparency in predictive models, such as decision trees, facilitates understanding of the logical reasoning behind predictions, which drives their adoption in practical scenarios. Unlike opaque models like SVMs and neural networks, transparent models provide clear justifications for outcomes, making stakeholders more comfortable using them in decision-making processes .
Organizations invest in Big Data technologies to harness opportunities for deeper insights and improved decision-making. Big Data analytics enables the identification of patterns, trends, and new opportunities that smaller datasets cannot reveal. Despite the challenges of data integration, processing speed, and governance, the enhanced efficiency, cost savings, and competitive advantages make the investment appealing .
Statistica is poised to play a crucial role in supply chain monitoring for Instrumentation Laboratory by ensuring the quality of supply chain operations through advanced analytics. It also aids in time management by proactively alerting engineers to potential issues, allowing them to focus on innovation rather than routine checks, thus significantly saving valuable time .
Organizations should focus on several critical success factors: aligning Big Data projects with clear business needs and objectives, ensuring strong executive sponsorship, aligning analytics with IT strategy, fostering a culture committed to fact-based decision-making, and maintaining a robust data infrastructure that integrates traditional and modern Big Data tools .
Instrumentation Laboratory ensures consistency and quality in its data analysis processes by using Statistica across the enterprise. This approach minimizes variability by standardizing how analyses are performed, despite different scientists having potentially different methods for data trimming or analysis. As a result, consistent outcomes are achieved with all scientists performing analyses in a harmonized way .