Chapter 8
Quality and risks
Assurance and control in BO
INTRODUCTION
This chapter discusses the quality and privacy of data in optimized digital
business processes. The characteristics of Big Data discussed in Chapter 2
+ 1 Vs) present unique challenges in terms of quality.
(3
Big Data characteristics include lack of format, schemaless-ness, high
volume, and high velocity – each of these characteristics leads to a chal-
lenge in handling the quality. The user communities on social media pro-
duce alternative data. These are the customers solving their own problems
and helping each other via communities. This data is not owned by the
organization analyzing it. There is no opportunity to verify the authenticity
of such alternative data. The question of privacy is continually challenging
the use of such data. Additionally, artificial intelligence (AI) solutions are
exposed to data without initially establishing the context. Context around
a data point is crucial for quality, as it provides the basis for analytics and
subsequent testing. Therefore, simply testing an AI solution thoroughly
using a range of data is not enough. Think data (Chapter 2) and all four
of its aspects come into play in ensuring holistic quality. Governance, risk,
and compliance (GRC) provide necessary controls to enhance the quality of
decisions and reduce the risks associated with data usage.
Quality initiatives establish trust and reliability of A I-based solutions in
the minds of the users. Robust quality approaches to AI go beyond the
technological quality of the solution and provide business benefits. Coupled
with robust governance frameworks, visibility of quality and services helps
in establishing business value.
Quality impacts business decisions. Relying on data analytics in
decision-making depends on trust. Insights from analytics are of value only
when the perception of quality and service assurance are met. The size,
speed, and variety of data and the corresponding analytics become irrel-
evant if the users start doubting the analytics outputs. Quality is an overall
function to establish that trust by preventing and detecting errors well in
advance and resolving the problems that occur in usage.
189
190 Artificial Intelligence for Business Optimization
Quality initiatives include validating and verifying the quality of data,
analytics (intelligence), model, processes, and usability. These quality ini-
tiatives are well supported by testing the source of data, applying feed-
back and results iteratively, and providing high visibility of the changes.
Managing the intricacy associated with data acquisition, storage, cleans-
ing, usage, and retirement is included in this all-encompassing aspect of
quality.
Database systems, information systems, and knowledge-based systems
could function independently before BO. The systems, hidden behind the
firewalls and in silos, could not be easily hacked. Thus, by their very
nature, pre-AI application in BO, enterprise systems were relatively
secure. BO requires systems to work together. Customer call centers,
enterprise web pages, online shopping stores, banking, and e-commerce
cannot work in silos. Analytics and services are based on an intercon-
nected world. This interconnectedness leads to quality and security
issues. These issues are of overall quality environment and cannot be
solved through testing alone.
The quality function includes quality of data, analytics, models, algo-
rithms, code, processes, and people. Quality is also a recursive func-
tion as it is also responsible for itself. The quality management function
is itself subject to the quality criteria of the quality environment. Tools
and techniques related to quality and testing support the quality process.
Transitions, training, project selections, prioritizations, and documentation
are quality issues associated with BO. The quality environment is an itera-
tive and incrementally operating environment that contributes directly to
the value-generating effort of BO.
Two practical aspects of quality are testing and assurance. These quality
aspects rest on detection and prevention. Detection occurs through testing
and prevention is through assurance. Quality control deals with the detec-
tion of errors. This is also called testing. Testing focuses on identifying
errors as compared with a reference point or an ideal situation. Testing
requires a valid data, and a model against which the new artifact being
tested is judged for its quality.
Quality assurance primarily deals with the prevention of errors so as
to provide an excellent user experience. This assurance is the discipline
of quality processes and models to ensure the prevention of semantic and
aesthetic issues. The aim of assurance is an end product that is free from
defects.
Quality assurance in BO includes activities such as setting data filters at
the source of the data, proactively managing the meta-data in addition to
the transactional data, and use of agile iterations and increments in devel-
oping AI solutions. Quality control or testing examples include the data and
verification of algorithms, testing the nonfunctional or performance aspect
of the solution, and fixing the “bugs” and retesting.
Quality and risks 191
Direct and indirect impact of bad quality
There are numerous direct and indirect impacts of bad data quality on busi-
ness decision-making. These impacts of quality specific to business optimi-
zation are as follows:
• Direct, short-term tactical operational decisions are affected by poor
quality of data. Defective products get installed and used by customer.
Other examples include accounting (payment and invoicing) errors.
These errors are rectified by undertaking a data cleansing and testing
exercise.
• Direct, long-term strategic decisions are impacted by poor data and
analytics: for example, loss of partnership or acquiring wrong part-
ners. These decisions cannot be easily corrected by simply correcting
the data; all underlying business processes need to be revisited, and so
also the algorithms analyzing the data – in order to improve the qual-
ity of these decisions. The more strategic is the decision, the greater
the need for agility.
• Indirect, short-term tactical operational decisions have poor quality:
for example, poor inventory management, slow production, or poor
customer service that are hidden in the processes. Handling these
indirect poor decisions requires improvement in data and analytical
quality, coupled with coaching and training for the staff and custom-
ers using the insights to make those decisions.
• Indirect, long-term strategic decisions are impacted due to poor
historical data and no control over alternative data: for example,
the development of a completely wrong product or inefficient and
undirected marketing, compliance issues, and loss of goodwill
of the organization in the market. The quality situations can be
improved by a holistic, strategic approach to quality assurance
that includes improving data, information, algorithm, knowledge,
and decision – coupled with enhancing people skills, attitudes, and
influences.
Risks and governance policies
Risk management closely accompanies quality initiatives. Governance,
risk, and compliance (GRC) provide the framework to control risks and
ensure compliance. Quality is further augmented by GRC, which enables
the creation of an overall quality environment that is strongly focused on
prevention rather than detection of errors. A GRC initiative is thus also use-
ful in creating services quality.
Good governance is a balancing act. While maintaining required busi-
ness outcomes, it also needs to ensure that the business is not so overloaded
192 Artificial Intelligence for Business Optimization
that it ceases to exist. Pragmatic governance balances risks with opportu-
nities, cost, and the need to deliver services and operations. An important
responsibility of risk management1 is to keep business viable. Governance
should not negatively impact quality, cost, and the operational efficiency of
the organization. 2
GRC is a part of the quality environment. The GRC effectiveness and effi-
ciency are enhanced by understanding the business environment, type and
size of business, existing capabilities and limitations of the business, and the
available tools and technologies for the governance function. Furthermore,
due to the prominence of compliance acts, such as the Sarbanes–Oxley Act
(SOX) 3 and Health Insurance Portability and Accountability Act (HIPAA),4
AI applications are required to provide explanations for the decisions
arrived at. Regulations also make it mandatory to provide auditability and
accountability of AI-based solutions.
General data protection regulation (GDPR)
General Data Protection Regulation (GDPR)5 dictates rules for the collec-
tion and processing of personal data. The regulation also provides the user
with the right to port their data from one organization to another – such as
a patient moving from one hospital to another.
While GDPR jurisdiction is the EU (European Union), it has a much
wider applicability in the digital world that transcends geographical bound-
aries. GDPR enforces organizations to justify their collection of data and
explain the reasons and methods of its processing. Electronic data can-
not be collected without unambiguous and explicit consent from the user.
Furthermore, the business must provide the users with a copy of their own
records, correct the details therein when asked by the user, and erase the
data if the user wishes so.
Quality and ethics
The ethical, social, and legal consequences of decisions impact the qual-
ity of BO. These issues cannot be handled by AI and its automated ML
systems on their own. AI systems do not understand the context in which
decisions are made. Therefore, it is crucial for an enterprise to validate the
quality of predictions and decisions based on AI. Chapter 10 describes
a model of combining human natural intelligence (NI), experience, and
intuition in every stage of the automated ML decision-making pipeline.
Human subjectivity which is an indispensable parameter for producing
and validating quality decisions and customer value plays an important
complementary role in supporting and enhancing BO with AI solutions.
Quality and risks 193
Big Data-specific challenges to quality and testing
The following are Big Data–specific challenges to quality and testing:
• Analytics use a wide variety of structured and unstructured data that
require different testing techniques to ensure their quality. Structured,
transactional data can be subjected to traditional testing techniques
based on equivalence partitioning and boundary values. Unstructured,
schemaless data cannot be tested with the same techniques. The auto-
mated data preprocessing and the dynamic online data monitoring (AI
tools and frameworks presented in Chapters 4 and 5, respectively) are
ideal for quality assurance and control of this kind of unstructured data.
• The data quality practices include data profiling that is based on the
potential users. Data profiling maps the data to the desired outcomes.
Therefore, in a way, the data profile starts to provide the context of
data usage.
• Changes to the context in which data is used are based on the changes
in outcomes desired. This context change results in different percep-
tions of the same dataset and, therefore, requires a different quality
strategy (including testing) for the same dataset. For example, a data-
set on interest rates has to be tested for algorithms predicting changes
to the interest rate versus predicting risks to the loan amount.
• Changing levels of granularity. Analytics will have to be tested against a
range of granularity, and that can be challenging. Fine granular analytics
needs to be tested with high-velocity data on a very narrow requirement.
• Externally created and externally sourced data cannot be easily fil-
tered for quality. Furthermore, should a filter spot an error, there are
very limited means available to correct such data. And on correction,
that data may end up with a different format to the original data –
creating challenges of inconsistent data formats. Externally, data that
is not relevant to the business context or one whose source is dubious
must be kept outside the firewall.
• The need to continuously and rapidly align the new, incoming Big
Data with existing transactional data. New data potentially contain
anomalies and security risks that are identified only when an attempt
is made to align it to the existing data.
• The need to test the concurrency of data processing that requires the
balancing of loads to ensure operational performance.
• Operational parameters can wildly vary. They may not always pro-
vide the opportunity for a satisfactory user experience (in terms of
time and visuals).
• Cybersecurity of data that is spread within and outside the organiza-
tion is a major challenge. Strategize for use of security analytics tools.
194 Artificial Intelligence for Business Optimization
QUALITY OF “ DATA TO DECISIONS”
The pyramid shown in Figure 8.1 summarizes the various items, processes,
and their relationships in the business world. Data management, data qual-
ity, data consistency, data access, data continuity, data completeness, and
data permissions form the basis for data quality. Each layer in the evolution
of data to decisions requires attention to quality.
The pyramid layers based on earlier discussion in Chapter 2 are stacked
from bottom to top with data, information, services, knowledge, and
decisions. Traditionally, business organization is set up h ierarchically –
with teams engaged with the duties in the lower rungs of the pyramid,
while the managers occupied the higher rungs, chiefly as knowledge
heads involved in decision-making. Data contains raw figures obtained
from sensor measurements. Information is the interpretation and mean-
ing imposed on the data by humans. Services are orchestration of vari-
ous information systems at a higher level. Knowledge is the association
of various types of information and services making action possible.
Decisions deal with the selection of a particular course of action, based
on knowledge.
AI requires businesses to change from the traditional pyramid hierarchy.
AI and ML have been extensively applied to business, and closer inspection
reveals that AI and ML algorithms limited to the bottom and top layers of
the pyramid is a lost opportunity. Large quantities of data (bottom layer)
are collected, preprocessed, and directly fed into ML algorithms, the result
of which are predictions and/or decisions (top layer).
action
ml Decisions meta ml
ES Knowledge Ml for es
services
Analytics & Services Ml for services
apps Information Ml for apps
Data Ml for DB
database
Reality
Observation
Figure 8.1 AI impacts and is impacted by the evolution of data from observations to
decisions.
Quality and risks 195
Quality of data
Data is the first layer of quality that ascertains its veracity. Quality data is
essential for quality AI and ML. Techniques such as equivalence partition-
ing and boundary values can test the veracity of data. These techniques
make use of sampling, checking, and correcting the data based on the
parameters provided by the business. The data quality initiatives surround-
ing Big Data need to handle factors in addition to testing just the “data.”
For example, large volumes of unstructured data simply cannot be tested
on their own. An understanding of the context enhances the value of data
analytics/semantics before the data can be tested.
The peculiar nuances of Big Data bring additional challenges to quality.
For example, consider how data is viewed as an “aggregate” in the unstruc-
tured data. This viewing of an aggregate implies a move away from the
structured rows and columns of a relational database. Testing and ensuring
the quality of such unstructured data cannot be undertaken by the tradi-
tional sampling from a set of data and testing it. Sampling of data based
on equivalence partitioning and boundary values presumes a semblance of
structure within the underlying data. Such structures are often not available
in Big Data. Hence, testing of Big Data sets may have to occur over an entire
data set rather than a sample.
Another important consideration in testing and quality assurance of Big
Data is that the data on its own may have very limited parameters that
can be tested. For example, basic filtering of input data can ensure that
numbers and texts are in their respective fields within a form on a Web
page. Beyond that, basic filtering may not be able to ascertain and test the
semantics behind the number or text. Additionally, with Big Data, the data
values may not have a format, and therefore even the format-level filtering
may not apply.
Merged data used for the calculation of derived values includes account
management, permissions management, database indexing, and log file
management. Testing this data requires an integrated test strategy with the
required functional testing, user acceptance testing, penetration testing,
accessibility testing, and load testing.
In addition to the sheer size of data, there are other factors impacting the
quality of data. These factors include infiltration of missing, misplaced, and
distorted datapoints. ML tools can also be used to address the data quality
problem along the following lines:
• Missing data is a common problem in many datasets. A couple of attri-
butes in a couple of records or some entire records could be missing in
a dataset for several reasons, including human error. ML algorithms
trained on relevant datasets can easily detect the missing datapoints
and fill in appropriate values obtained through the learning process.
196 Artificial Intelligence for Business Optimization
• Data augmentation techniques include ML to solve the problem
of insufficient data in numerical, text, and image data formats
(Chapter 4). In a similar way, ML algorithms look for distorted or
inappropriate datapoints.
• Numerical data in which values are misplaced or distorted or out of
range are very difficult to detect when dealing with large datasets. For
example, the data entry of someone’s year of birth as 1794 (instead of
1974) goes unnoticed when data scientists must check tens of thou-
sands of data records. ML algorithms trained on the validity of datas-
ets can pinpoint incorrect values of data attributes and, in some cases,
even suggest solutions.
• Text data product reviews by customers, customer sentiment analysis,
blogs, and so on are the modern sources of data in the text format.
Misspellings, wrong usage of words, and grammatical and syntactical
errors can distort the meaning of sentences or make them meaningless.
Conveying correct syntax and semantic meaning is important because
text-based information is the source of knowledge. For example, in
the course of typing the contents of this book, several typographical
errors like machine leaning (instead of learning) go unnoticed. Such
subtle errors in text data can be traced and fixed by ML algorithms
trained on large language corpora.
• Image data. Human eyes soon get exhausted looking at the details in
colored images. It is humanly impossible to check millions of images
that are routinely crunched by image processing programs in com-
puter vision-related tasks. In addition, some of the latest security
attacks come from Generative Adversarial Networks (GAN),6 which
fool even the state-of-the-art image recognition algorithms. For exam-
ple, introducing an infinitesimal perturbation at the right intersec-
tion of pixels in the image of a panda can mislead even the highest
performing algorithm to recognize the panda image as nematode or
gibbon with exceedingly great confidence, even though nothing has
changed on the surface when viewed by human eyes.7 ML algorithms
trained with GANs are capable of detecting such errors.
Quality of information
Level 2 in Figure 8.1 represents the processing of data to create informa-
tion. Information is a systematic identification of patterns and trends within
those data. Data, on its own, is not meaningful, whereas information based
on the data provides meaning or semantics. Ensuring the authenticity of
this meaning is the responsibility of quality of information. The e-world
comprises billions of pages that are increasing exponentially. There is no
centralized agency monitoring the quality of information available on the
internet. Individual information systems on the web are constantly trying
Quality and risks 197
to adjust to the flow of information on the web so as not to be outmoded.
It is difficult to test the quality of the information fed into these systems
from the web. ML algorithms placed at the front-end of web-based infor-
mation can learn to test and validate the quality of the incoming informa-
tion. News, podcasts, articles, blogs, auctions, advertisements, shopping
sites, and mobile apps are subject to testing via ML algorithms.
A caveat in the BO world is that anomaly in information is not always a
lack of quality. Anomalies can signal security breaches which require detec-
tion and filtering. Anomalies can also indicate emergence of a new idea.
For example, consider how spam filters in late 2019 filtered out mails based
on certain keywords like “virus, infection, epidemic, disease, and so on.”
Filters embedded with ML functions would have possibly noticed the emer-
gence of a new entity (like COVID-19) by analyzing the anomalies in the
email information.
Quality of analytics and services (collaborations)
Analytics quality is the quality of its algorithms. Analytics are complex,
and their algorithms need to be verified from both statistical techniques
and a programming viewpoint, including exceptions and error handling.
The quality of analytics includes verification and validation of their syntax,
logic, standards (e.g., naming of attributes and operations), and the purpose
they serve (semantics). Analytics establish correlations between data and
information. The more the data, the better is the output. Also, data analyt-
ics is not limited to analytics on a singular type of data. Data is sourced with
the help of services from multiple and widely varied databases (typically on
the Cloud) and a relationship established between them in order to perform
data analytics. Analytics itself is offered as a Service8 on the Cloud.
Since the algorithmic code deals with and manipulates the data, the qual-
ity of that code also influences the quality of the data. A white box method
for quality of analytics approach to quality of analytics consists in meticu-
lously checking the code for syntax and semantic errors. White box meth-
ods can be tiresome and not foolproof. ML offers a black box approach to
quality testing of analytics and corresponding services. Continuous testing
based on Agile iterations enable black box testing. An ML algorithm pro-
vided with known output for a given analytics and services can learn (in a
supervised way) to detect errors and faults in similar analytics and services
on the web.
These tests (quality control) are also meant to detect performance and
reliability issues. Quality assurance of the algorithms happens through
models and architectures and following a development process (i.e., Agile
in the solution space).
198 Artificial Intelligence for Business Optimization
depend on a number of factors, such as the ease of configurability of the
analytics, the expertise and experience of the user, and the urgency and
criticality of the analytics. Most of these factors are outside the control
of the organization, and testing them requires considerable assumptions.
These assumptions include the reliability of the source of analytics, their
compliance with the legal requirements of their jurisdiction, and the use of
SSA by other business partners.
Data usability is the ease of use of data within analytics. The accessibil-
ity, profiling, cleansing, and staging of data play a part in data usability
and, in turn, its quality. These factors affect the quality of the solution
being developed.
Quality of knowledge and insights
Knowledge is rationalization and correlation of information through reflec-
tion, learning, and logical reasoning. Analytics form the basis of such cor-
relation. The body of explicit knowledge available in an enterprise along
with the implicit knowledge in the form of intuition and experience of busi-
ness experts is optimized in BO. End-users and domain experts seek this
knowledge.
Use of knowledge and insights is based on analytics from expert systems.
These expert systems in the past were insular, with the body of knowledge
occasionally updated to keep abreast with the latest domain knowledge.
ML-based systems analyze and filter online information and engage in
knowledge discovery. Newly discovered knowledge is automatically added
to the knowledge base. Increasing the reliability requires testing the discov-
ery capabilities through test data and on a real-time basis. Tested results
are inputted via feedback loop back into the system and further tested to
enhance its capabilities.
Quality of decisions
While data, information, analytics, and (to a large extent) knowledge are
considered objective, decisions are a human subjective trait. Quality of
decision-making depends on tacit human factors such as personal expe-
rience, value system, time and location of decision-making, sociocultural
environment, and ability to make estimates and take risks.
Quality of decision is a result of the quality of the previous layers shown
in Figure 8.1 and combined with NI. AI-based systems can suffer data
bias, algorithm bias, and decision bias. Large amounts of skewed input
data introduce data bias. Algorithms can be designed by developers with
preconceived notions that can be biased. Eventually, decisions based on
data and algorithm biases can themselves be biased. The feedback of conse-
quences from these decisions creates further biases. Therefore, the quality
Quality and risks 199
of decisions can be improved by iterative feedback, by evaluation of con-
sequences by multiple people and by keeping mind the context, time and
place of decisions.
Finally, intelligence is actionable knowledge based on insightful use of
Big Data. The decision-maker’s ability to distinguish decisions based on
their important, relevance, context, and organizational principles is a cru-
cial quality factor for intelligent decision-making.
Testing is performed at best with software tools. The ML technol-
ogy discussed here is applied to assure quality of the middle layers of the
data-decision pyramid. ML can be used as a tool to test and refine the qual-
ity. ML algorithms need the application of quality techniques to enhance
their quality.
QUALITY ENVIRONMENT IN AI AND ML
Assuring ML quality
Analytics can be used to ascertain the quality of data. This data and the cor-
responding analytical algorithms are subjected to quality assurance activi-
ties using the capabilities of AI. For example, an analytical algorithm can
be applied to the incoming weather data to identify potentially unrealistic
spikes (e.g., a temperature variation of 500°F at the same location within
a few minutes). Such identification of spikes can lead to an investigation of
the algorithm itself to ensure that it has passed the testing and is secured.
Quality assurance and testing of ML algorithms/systems are carried on
at three levels:
1. Testing performance and logic of ML algorithms
2. Testing for bias in the ML predictions
3. Reducing the inexplicability of ML algorithms
The first is discussed here, and the second and third are discussed in
Chapter 10.
Testing of ML algorithms starts by separating the data into testing and
training. Testing the algorithms also requires a strategic approach includ-
ing black- and white-box testing. Unstructured data, in particular, is chal-
lenging to test. Creating sample data and testing the execution of logic are
recommended. When the questions themselves are not known and the cor-
relations are produced by the machines, quality assurance becomes a risk
management exercise.
The opportunity for feedback loop in testing Big Data and its analytics
is much less. This is so because the topography of themes provided by the
Big Data is not known at the onset. Testing requires validation of data and
algorithms. Testing of BD systems requires a certain amount of guesswork.
200 Artificial Intelligence for Business Optimization
The quality of data and the validation of algorithms are never entirely com-
plete. Quality assurance and testing of ML systems require validating the
test data, validating the training data, and then validating the performance
of the algorithms against open data.
GRC positively influences the transformation and delivery of BO.
Leadership defines, leads, and manages BO and its overall quality. Leaders
not only set the objective and strategy, but they also set parameters for qual-
ity. Leaders balance the cost of making bad quality decisions and also not
making decisions. BO comes with a risk and the risk is higher when quality
of data and algorithm is not verified and people are involved.
The quality and value from Big Data are also based on the perception of
the users of the analytics and business processes. Therefore, ensuring the
quality of Big Data goes beyond technologies and also includes sociology,
presentations, and user experience. These are some important aspects of
quality that are specific to Big Data and that need to be handled on a con-
tinuous basis for the analytics and business processes to provide the neces-
sary confidence and value in business decision-making.
Assuring quality of business processes
The quality activities here include modeling of business processes and,
thereafter, the verification and validation of those business processes using
techniques similar to those used in the quality assurance of models and
architectures. Process quality depends on the way in which the processes
are executed by the users and their end goals.
The quality of business processes is verified and validated by creating visual
models and then applying the quality techniques of walk-throughs, inspec-
tions, and audits. Business processes use applications and analytics to help
the end users achieve their goals. Therefore, the quality of business processes
depends on the way the users perceive their achievements. Understanding,
documenting, and presenting the visual models of the business processes to
the end users and incorporating their feedback in an iterative manner are
Agile ways to enhance the quality of business processes. Process modeling
standards (such as Unified Modeling Language [UML]9,10,11 and Business
Process Model and Notation [BPMN]12) and corresponding modeling tools
further help in improving the quality of business processes. In addition to
the processes that form part of the business, there are the processes that deal
with producing the solutions. Project management, business analysis, and
solutions development life cycles are examples of these processes.
The quality of these adoption and solutions development processes is also
important and needs to be subject to the same quality techniques as those
used for quality in business processes. A set of well-thought-out activi-
ties and tasks combined with the Agile techniques produce accurate and
higher-quality analytics which, in turn, enhances the quality of the business
process.
Quality and risks 201
Developing the quality environment
The following are the strategic considerations in developing a quality envi-
ronment in a Big Data initiative:
• Identifying key business outcomes: Defined in the corporate strategy
in order to create a common understanding of the Big Data adoption
exercise and the role of quality in helping achieve those outcomes.
These business outcomes have a bearing on each of the data quality
characteristics. The more detailed the outcomes and the higher the
risks associated with those outcomes, the greater is the demand on
data quality.
• Modeling the range or coverage of Big Data: Its sources, types, stor-
age mechanisms, and costs. Clarity in understanding the range of data
enables the formulation of quality strategies based on the depth and
breadth of the incoming data, associated risks, and costs associated
with analyzing that data.
• Extent of tool usage: Most Big Data quality initiatives need tools
and technologies that complement the technologies of Big Data. For
example, quality assurance and control activities on a NoSQL data-
base will need tools that can verify the extraction of data, match
the extracted data against reference data, and provide feedback to
the quality personnel. This verification exercise can be challenging
because of the unstructured nature of the data; therefore, tools that
can handle the testing of unstructured data and its performance
are required. Tools are also a must because of the high velocity of
incoming data and the varying levels of granularity in analyzing that
data. Quality tools need to be able to operate within the Big Data
environment.
• Ensuring a balance between the rigors of quality and correspond-
ing value: The business decision-makers need to collaborate with the
quality personnel to ensure that quality efforts are balanced with the
business outcomes. Standardization of data and its cleansing in order
to enable processing can either go overboard or be carried out over
data that may not be used at all. Therefore, it is important to keep the
ultimate usage of the data and the desired business outcomes in mind
before undertaking quality activities on the data
Assurance activities
Quality activities are carried out over the key phases that are transitioned
by data entry, storage, cleansing, and retiring. In each of these phases, the
volume, velocity, and variety of Big Data (including myriad data sources)
add to the challenge of data quality. These challenges include complica-
tions of data governance and risks associated with the management of
202 Artificial Intelligence for Business Optimization
data. Big Data in particular needs continuous filtering and standardization.
Following are the data assurance activities for Big Data:
• The entry point for Big Data has “presumptions” in their use. These
presumptions are needed to create and apply filters to that data. These
presumptions can be based on prefabricated (i.e., halfway completed)
analytics. The fuzzy and uncertain nature of the use of the data pres-
ents input filtering challenges.
• Sources of data (social media, crowd, other systems) each requiring
a specific filtering before entry, and each of these data sources is not
always under the organization’s control.
• Lack of context requires presumptions about the use of data before
filtering.
• Velocity of data is a big challenge requiring the use of tools for filtering.
• Complexity of analytics applied to the data presents quality challenges.
• Data is identified, secure, and stored – presumably on the Cloud
requiring data assurance effort to shift to the Cloud.
• Need to maintain data entities as separately and identifiable as pos-
sible due to the 3V of data but the challenge of data quality is further
exacerbated when data is mixed types. A NoSQL database requires
dynamic modeling and design before data can be stored in it.
• Strategies for backing up and mirroring of databases should be in
place. This allows for efficient restoration that is important for data
assurance.
• Data is continuously checked for redundancies and abnormalities.
Tests are used to remove spikes and prepare the data for staging area
where analytics can be performed.
• Ongoing monitoring of data as new data is integrated/ interfaced with
existing data for analytics.
• Data retired after use (and when it has lost its currency and relevance)
has to be formally archived with the help of tools even in a Hadoop
environment.
Developing the testing environment
Testing in the Big Data space requires due consideration to the testing
of data, scripts, algorithms, and tools. Following is a list of these testing
considerations:
• Testing of algorithms – Development of test harnesses to test algo-
rithms that cannot be executed and tested on their own.
• Test scripts – Writing of test scripts based on use cases in order to test
the data algorithms and repeat those tests on a continuous basis in an
agile environment.
Quality and risks 203
• Repository of test cases – Which can be used, reused, and augmented
as the continuous testing progresses.
• Testing tools – Adopt a standard test tool for test plans, test cases,
and results tracking. These tools (e.g. Silk, Chapter 9) also provide
security analytics.
• Test planning – A standard process for the creation of test plans and
test cases, as well as an outline of the test schedule and resource needs.
These plans are based on industry experience, test frameworks, and
CAMS agile.
• Test result tracking – Tools and processes for tracking test results and
tracking back to requirements and release versions. Analytics on test
results indicate areas for rework and regression testing.
• Test approvals – Processes and tools for tracking test approvals and
tracking them back to releases and authorizations for releases. These
approvals can happen on a Kanban board.
• Testing processes – The required processes, knowledge base, training,
and support associated with testing Big Data Testing Types. This is a
visible and highly interactive agile process.
Ongoing monitoring of data that is being used for analysis. The quality of
data here depends on factors such as changes to the existing data, addition
of new data (typical of high-velocity Big Data), and loss of currency of data
while it was being used for analytics.
New data is integrated or interfaced with existing data for analytics –
and this integration needs to be modeled, tested, and then executed. The
integration of Big Data occurs between various data sets (structured and
unstructured) owned by the business, data sets provided by external entities
(third party), and those being made available through open data initiatives.
ADDITIONAL QUALITY CONSIDERATIONS
All quality efforts are directed at improving the eventual quality of business
decisions. Therefore, the basic data quality eventually impacts the business
processes. There are a number of quality techniques that are applied at the
data and process levels.
• Cleansing and standardizing data – This is a technique to identify and
remove the spikes and troughs within a given set of data. In the Big
Data domain, this technique becomes important from the point of
view of standardizing the data in preparation for its use in analytics.
Data editing tools are used in this exercise of cleansing data.
• Applying syntax, semantics, and aesthetic checks to data, algorithms,
and code quality. While these techniques are immensely helpful in
204 Artificial Intelligence for Business Optimization
ensuring the quality of models and processes, they can also help in
reviewing and improving the code and data quality.
• Identifying the source of data and tracing data to that source in order
to enable filtering and cleaning – as much as possible. Identification of
the source of data may not always be possible beforehand (especially if
those sources are identified automatically through a Web service), and
in some cases, precise identification of the source may not be permit-
ted (as in the case of identifying the crowdsource).
• Using standard architectural reference models and data patterns in
order to provide the basis for the matching of data and thereby spot-
ting mismatches.
• Controlling the business process quality through timely checks and
balances at specific activities and steps within the business processes
using tools and standards.
• Continuous testing effort as Big Data streaming results in changes to
the incoming data points and their context.
• Using Agile techniques, such as showcasing and daily stand-ups, to
enable a high level of visibility and feedback in developing high-quality
analytics. Some analytics can also be used in identifying errors within
a database.
• Using processes and tools for implementing governance policies for
data. These can be the automated implementation of electronic poli-
cies through algorithms that are specifically created and executed
for that purpose. The quality of data stored within the organization
boundary can be subjected to greater controls than those acquired
from outside. In both cases, though, good data may not always result
in good decisions, and vice versa. All the checks and balances can-
not guarantee total accuracy. Besides, Big Data input is from sources
other than humans. Therefore, cross-checking and filtering out human
input is not a guarantee for erroneous data entry (e.g., wrong data cre-
ated). The greater the number of analyses and manipulation of data,
the greater are the chances of loss of quality. This loss of quality of
data is primarily felt in the quality of decision-making resulting from
that data.
Nonfunctional testing
Slow, unoptimized, and bug-ridden applications suffer a lack of accuracy
and performance. Databases, information systems, analytics and services,
knowledge-based systems, and decision-aiding systems are composed of AI
analytics. These applications need intense quality control of their function-
ality as well as performance.
Such testing includes performance, volume, and scalability testing. This
type of testing is called nonfunctional. The interest here is in the run time
Quality and risks 205
performance of the solution (as against its step-by-step function accuracy).
This type of testing requires an executable system with fully loaded opera-
tional data (or its equivalent synthetic data). An Agile approach to develop-
ing solutions helps here because in Agile, testing is a continuous process.
Quality of metadata
Ensuring specific focus on metadata: This focus is on the quality of the
context surrounding a data point. Each data point has many contextual
parameters that provide additional information about that data point. This
additional information, also called metadata, provides a filter for capturing
and sharing data elements. For example, metadata around a temperature
data point can be the location from where that temperature is being cap-
tured. A dramatic change to the next weather data point, in the next minute
from the same location, is indicative of bad data quality. The metadata
around the collection of data points produces a common reference model
that helps filter incoming data. In addition to the quality of the incom-
ing data, there is also a need to ensure the quality of and test the refer-
ence model itself. This quality assurance and testing of the model requires
cross-functional collaboration, iterative development of prototypes, and
incorporation of the feedback back into the metamodel.
In addition to the quality of the data and processes, the quality of
metadata is another important element that impacts the quality of busi-
ness decision-making. Metadata starts to provide an initial context to the
incoming data. For example, a tag is basic metadata that is assigned to
incoming unstructured data. This tagging provides an identity to the data.
This identity and the parameters (metadata) surrounding the data provide
a hook for testing – as they improve the chances of filtering out bad data.
For example, consider a weather data point showing 800°F. The metadata
around this data point provides the location (latitude and longitude) and
the time (say, 1:00 p.m.) at which this temperature is recorded. If the next
data point shows, say, 350°F, then the tools filtering this incoming data
can cross-check against the location (latitude and longitude) and the time.
And if the location is the same and the time is similar (say, 3:00 p.m.), then
the incoming data is wrong. Big Data quality at the metadata level implies
improved consistency across the data suite. Reference to metadata provides
a common basis for data validation, standardization, enhancement, and
transformations.
Quality of alternative data
Alternative data holds the promise of niche analytics but, at the same time,
presents the biggest challenge in terms of quality. Alternative data is neither
owned nor controlled by the users of that data. For example, discussion
206 Artificial Intelligence for Business Optimization
blogs, tweets, and likes on a social media page are all contributors to the
Alternative data. These data can very easily comprise fake data, news,13
and unsubstantiated information on social media platforms. Verifying their
quality using traditional tools and techniques is an almost impossible task.
Suggestions on handling the quality of alternative data include critically
reading the material, checking the sources, and comparing with others.
Automation and optimization of processes can include ML algorithms to
take over some of the basic activities of verifying alternative data.
Sifting value from noise in Big Data
The available data can come in with a lot of noise that is irrelevant to
the outcomes and does not gel with other data points within a data set.
Analytics on this data will be embedded within business processes. Thus,
the quality effort is in identifying the relevant data, providing it with
some structure to make it analyzable, and then decision-making (explicit).
Supporting data quality is the quality of the enterprise architecture, ana-
lytical algorithms, and visualizations. Quality initiative is an effort to sift
value from the chatter and noise of data and make it available to business,
verifying and validating its results through a business process. Techniques
such as data sourcing and profiling, data standardization, matching and
cleansing (scrubbing), and data enrichment (plugging the gaps and correct-
ing the errors) are applied here to make the data analyzable.
Retiring the data safely and securely after use – data retirement after
use (and when it is no longer current and relevant) needs to be undertaken
carefully to ensure that the retired data cannot be abused by unauthorized
parties. Besides, the data that is retired for one business can still have some
potent value in it for another business – such as being able to track the his-
tory of decisions made by a business.
Quality in retiring data
A formal archival process of retired data is another important ingredient
of quality. Spent or unusable data has to retire in a safe and secure man-
ner. Audit trail needs to be preserved in order to provide the proof of data
destruction in a controlled manner. Big retirement of Big Data is controlled
by governance and compliance requirements. Governance and policies
around the management of risks help control the retirement and detection
of data sets. Attention to the external sources of data and systems is also
required. The policies and processes provide checks and balances in data
handling – including its cleansing, storage, usage, and eventually retire-
ment. Unique data types such as meta- and alternative data may not belong
to the organization. Hence, these data types may not retire. This archival
process is undertaken with the help of data manipulation tools.
Quality and risks 207
Velocity testing
Velocity of data at the entry is an important quality challenge requiring the
use of tools for filtering. This is particularly so when the data is being gener-
ated through the Internet of Things (IoT) and machine sensors. The effect
of bad data on quality is exacerbated if that data is generated by sensors
and, as a result, is inundating the entire analytical systems.
Testing the velocity of data is important because of the impact of velocity
on performance. This is the performance of the storage systems, as well as
that of the analytics to keep up the processing with the velocity.
The following are the characteristics of velocity testing:
• Velocity testing includes testing the speed with which data is being
produced and received (e.g., data generated by IoT devices).
• The rate of change of data (including speed of transmission, storage,
and retrieval) and its impact on the analytics is also tested here.
• Also included is testing the speed of analytics – that is, the rate of pro-
cessing of data and the creation of information or knowledge.
• Velocity testing will require the creation of a test environment that
mirrors the production environment, as this is a part of operational
testing of the performance of the system.
Visualizations are presentations on various user devices. These are the
graphic user interfaces presenting the dashboard of analytics. The quality
of visualizations plays an important role in providing a satisfactory user
experience.
G OVERNANCE–RISK–COMPLIANCE
AND DATA QUALITY
Effective governance is based on a body that comprises both business and
technical decision-makers of the organization. Underneath this group of
decision-makers is the business capability competency group that helps
align the capabilities of the organization to the business outcomes.14 These
groups synergize operational, strategy, and IT professionals to ensure that
the relevant IT systems, services, and platforms support the organization’s
outcomes.15
The GRC returns a significant value when it is carefully mapped to busi-
ness capabilities. This is so because apart from ensuring compliance, GRC
is also geared to ensure an ROI for the business. GRC ensures that BO
adoption is of value to the organization.
GRC is helpful in maintaining compliance with both external and inter-
nal legal, audit, and accounting requirements. GRC coupled with business
capabilities is vital to pave the path for AI technology investments.
208 Artificial Intelligence for Business Optimization
Business compliance and quality
Business compliance is the need for the business to develop capabilities
to meet regulatory compliances. These compliances enhance quality and
reduce risks. The external demands for government and regulatory require-
ments also need the business to reorganize itself internally. An Agile inter-
nal business structure is able to respond better to ever-changing legislation.
Consider, for example, the Sarbanes–Oxley (SOX) legislation. This legisla-
tion provides protection from fraudulent practices to shareholders and the
general public and, at the same time, also pins the responsibility for internal
controls and financial reporting on the chief executive officer (CEO) and
the chief financial officer (CFO) of the company.
AI-enabled Agile business carries out this accountability and responsibil-
ity through changes in the internal processes, updating of ICT-based sys-
tems to enable accurate collection and timely reporting of business data,
and changes in the attitude and practices of senior management. Another
example of the need for the business to comply is the rapid implementation
of regulations related to carbon emissions. This legislation requires busi-
nesses to update and implement their carbon collection procedures, analy-
sis, control, audit, and internal and external reporting.
GRC complements the BO initiatives in an organization. GRC provides a
consolidated and comprehensive approach to controlling an organization’s
business. GRC helps control existing enterprise data and functionality as
much as it helps in handling the new Big Data. With a formal GRC in place,
an organization can monitor its activities, provide necessary controls around
the activities, conduct audits, and prepare reports. As a result of GRC, an
organization improves its ability to prevent fraud by providing transparency
and enabling executive-level control of data and business processes. This
makes it imperative to discuss GRC in terms of the quality and value.
GRC, Business, and Big Data Governance within a business imply the
following:
Governance is the overall management approach to control and direct
the activities of an organization. This direction in turn requires an
understanding of the desired business outcomes, and the capabilities16
of the advent of Big Data require even more governance than before
because of the uncertainty of data sources and the collaborations
required among business partners.
Risk management supports governance through which management
identifies, analyzes, and where necessary, responds appropriately to
risks. Risks in the Big Data age are extremely dynamic because of the
dynamicity of the underlying data and the depth of analytics. While
Big Data analytics can help identify risks, there are also risks associ-
ated with the analytics themselves. The need to test the algorithm
Quality and risks 209
for syntax, semantics, and aesthetics, as well as for performance and
other nonfunctional parameters, is acute in the Big Data world.
Compliance means conforming to stated requirements, standards, and
regulations both external and internal to the organization. Due to the
complexity of Big Data, compliance assumes greater importance than
before. This is because compliance requirements (especially exter-
nal and legal) can be potentially broken at any of the layers of the
organization at which Big Data analytics are aiding decision-making.
Analytics enable decision-making at the lowest rung of the organi-
zation, but it is the senior-most directors of the company that are
responsible for the ultimate outcome.
Quality of service
Analytics are “served” through various services that are predefined. Quality
of these services is assured by GRC. For example, the ITIL standards ensure
areas of service management. Big Data services require the management of
requests. The request management processes are modeled, reviewed, and
tested to support a new service, process, and operations. The ability of ser-
vice in the organization considers the following:
• Accounts and permissions – Request for new accounts, closed
accounts, and permission changes.
• Projects – New development to be managed as a project. Development
that requires complex management, multiple stakeholder engagement,
and taking typically more than five business days of work. May have
own release cycle or be released as part of other releases.
• Enhancements – Additions, extensions, and enhancements that
can be done in typically less than five business days. This work is
clearly defined, easily accomplished, and simple, testing overhead.
Enhancements are mainly released as part of a planned cycle but may
be released out of cycle.
• Defects – May be remediated as part of incident management or
within problem management. Defects may take more than five days to
remediate and may require project management. Defects are mainly
released as part of a planned cycle but may be released out of cycle.
Defects are mainly managed as incidents. For a request to be pro-
cessed, the service will need actionable (all required information) and
authorized (from the correct party) requests, especially when working
with vendors and outsourced operations.
• Metrics and measurements associated with Big Data can provide
performance information, risks, and opportunities for operational
optimization. These metrics can be applied to measure analytics and
management.