BDA Module11 Notes
BDA Module11 Notes
MODULE 11
Data Analysis and Modelling
Sections: 11.1 Process, Benefits & Types | 11.2 Data Mining | 11.3 Analytics & Model
Building (Descriptive, Diagnostic, Predictive, Prescriptive) | 11.4 XML & XBRL | 11.5
Cloud, BI, AI, RPA, ML | 11.6 Model vs Data-driven Decision Making
Benefit Description
(i) Improves Companies can use information from data analytics to base their
Decision Making decisions, resulting in enhanced outcomes. Reduces guesswork
involved in preparing marketing plans and deciding what materials to
produce. Continuously collect and analyse new data to gain a deeper
understanding of changing circumstances.
(ii) Increase in Data analytics assists firms in streamlining their processes,
Operational conserving resources, and increasing profitability. When firms have a
Efficiency better understanding of their audience's demands, they spend less
time creating advertising that does not fulfil those needs.
(iii) Improved Data analytics gives organisations a more in-depth understanding of
Service to their customers, employees and other stakeholders. This enables the
Stakeholders company to tailor stakeholders' experiences to their needs, provide
more personalisation and build stronger relationships.
# Step Description
(i) Setting the The most difficult step. Data scientists and business stakeholders
Business Objective together identify the business challenge, which informs the data
queries and parameters for a specific project. Analysts may also
need to conduct further study to understand the company
environment.
(ii) Preparation of Once the scale of the problem is established, data scientists
Data determine which data collection will assist the company. The data
is then cleansed by eliminating any noise, such as repetitions,
missing numbers and outliers. Dimensionality reduction may also
be performed to maintain only the most essential predictors.
(iii) Model Building & Data scientists study intriguing relationships such as frequent
Pattern Mining patterns, clustering algorithms or correlations. Deep learning
algorithms may be used for classification (supervised learning
with labelled input data) or clustering (unsupervised learning with
unlabelled data).
(iv) Result Evaluation After aggregating data, findings must be analysed. Results must
& Implementation be valid, original, practical and comprehensible. When this
criterion is satisfied, companies can execute new strategies based
on this understanding to attain their intended goals.
Technique Description
(i) Association Rule-based technique for discovering associations between variables
Rules inside a given dataset. Commonly employed for market basket analysis
— understanding linkages between various items. Helps create
effective cross-selling tactics and recommendation engines.
(ii) Neural Primarily used for deep learning algorithms. Replicates the
Networks interconnection of the human brain through layers of nodes to process
training data. Every node has inputs, weights, a bias (or threshold) and
an output. If output value exceeds a predetermined threshold, the node
'fires' and passes data to the subsequent layer. Learns using supervised
learning and gradient descent.
(iii) Decision Tree Uses classification or regression algorithms to classify or predict likely
outcomes based on a collection of decisions. Employs a tree-like
representation to depict the potential results of actions.
(iv) K-Nearest Classifies data points depending on their closeness and correlation with
Neighbour (KNN) other accessible data. Assumes similar data points exist in close
proximity to one another. Measures distance between data points (often
by Euclidean distance), then assigns the most common category or
average.
Application Description
(i) Detecting Money Money laundering is the illegal conversion of black money to
Laundering & Financial white money. Data mining techniques have advanced to detect
Crimes money laundering. The methodology provides a mechanism for
bank customers to detect or verify the detection of the anti-
money laundering impact.
(ii) Loan Repayment & Loan Distribution is the core business function of every bank.
Credit Policy Analysis The loan prediction system automatically computes the size of
the characteristics it employs and examines data pertaining to its
size. Data mining aids in management of all critical data and
massive databases by utilising its models.
(iii) Target Marketing Data mining and marketing work together to target a certain
market and assist in determining market decisions. With data
mining, it is possible to keep track of earnings, margins etc. and
determine which product is optimal for various types of
customers.
(iv) Design & The business is able to retrieve or move data into several huge
Construction of Data data warehouses, allowing a vast volume of data to be correctly
Warehouses and reliably evaluated with the aid of various data mining
methodologies and techniques.
Businesses utilise analytics to study and evaluate their data and translate their discoveries into insights
that aid executives, managers and operational personnel in making more educated and prudent business
choices.
Data tagging and reporting standards ensure that financial and business data is communicated in a
standardised, machine-readable format. The two most important standards are XML and XBRL.
Definition Definition
A file format and markup language for A data description language that facilitates
storing, transferring and recreating arbitrary the interchange of standard, comprehensible
data. Specifies standards for encoding texts corporate data. It is based on XML and
in a format understandable by both humans enables automated interchange and
and machines. Defined by the 1998 XML trustworthy extraction of financial data
1.0 Specification of the World Wide Web across all software types and advanced
Consortium. technology, including the Internet.
storing, sending and rebuilding arbitrary on the precise data content of financial
data. documents.
● Labels, categorises and arranges ● XBRL tags may designate data as
formats: RSS, Atom, Office Open XML, from a single source of information —
OpenDocument, SVG, XHTML. reduces erroneous data entry and
Communication protocols: SOAP, XMPP. increases data reliability.
11.5 Cloud Computing, Business Intelligence, AI, RPA and Machine Learning
Type Description
(i) Private Cloud Offers a cloud environment exclusive to a single corporate
organisation. Physical components housed on-premises or in a vendor's
datacenter. High level of control since the private cloud is available to
just one enterprise.
Benefits: Customizable architecture, enhanced security procedures,
capacity to expand computer resources as needed.
Deployment: Business maintains private cloud infrastructure on-
premises (via intranet), or engages a third-party cloud service provider
to host and operate servers off-site.
Type Description
(ii) Public Cloud Stores and manages access to data and applications through the
internet. Fully virtualized — shared resources may be utilised as
necessary.
Benefits: Enables enterprises to grow with more ease; option to pay for
cloud services on an as-needed basis (significant benefit over local
servers). Rigorous security measures prevent unauthorised access to
user data by other tenants.
(iii) Hybrid Cloud Blends private and public cloud models — enterprises exploit benefits
of shared resources while leveraging their existing IT infrastructure for
mission-critical security needs.
Benefits: Store sensitive data on-premises and access it through apps
hosted in the public cloud. Example: Keep sensitive user data in a
private cloud (for privacy compliance) and execute resource-intensive
computations in a public cloud.
9 BI METHODS
BI Method Description
(i) Data Mining Large datasets may be mined for patterns using databases,
analytics and machine learning (ML).
(ii) Reporting The dissemination of data analysis to stakeholders in order for
them to form conclusions and make decisions.
BI Method Description
(iii) Performance Comparing current performance data to previous performance
Metrics & data in order to measure performance versus objectives,
Benchmarking generally utilising customised dashboards.
(iv) Descriptive Utilizing basic data analysis to determine what transpired.
Analytics
(v) Querying BI extracts responses from data sets in response to data-specific
queries.
(vi) Statistical Analysis Taking the results of descriptive analytics and using statistics to
further explore the data, such as how and why this pattern
occurred.
(vii) Data Visualisation Data consumption is facilitated by transforming data analysis
into visual representations such as charts, graphs and histograms.
(viii) Visual Analysis Exploring data using visual storytelling to share findings in real-
time and maintain the flow of analysis.
(ix) Data Preparation Multiple data source compilation, dimension and measurement
identification, and data analysis preparation.
DEFINITION OF AI
John McCarthy of Stanford University defined Artificial Intelligence as: 'It is the science and
engineering of making intelligent machines, especially intelligent computer programs. It is
related to the similar task of using computers to understand human intelligence, but AI does
not have to confine itself to methods that are biologically observable.' Alan Turing's
landmark paper 'Computing Machinery and Intelligence' marked the genesis of AI discourse.
Turing, the 'father of computer science,' posed the question 'Can machines think?' and
proposed the famous Turing Test — where a human interrogator attempts to differentiate
between a machine and a human written answer. In its simplest form, AI combines
computer science and substantial datasets to allow problem-solving. It includes the subfields
of machine learning and deep learning.
ARTIFICIAL INTELLIGENCE IN
FINANCE
Both deep learning and machine learning are subfields of AI. However, deep learning is a subfield of
machine learning. The hierarchy is: AI → Machine Learning → Deep Learning (innermost).
Often requires more structured data to learn. Automates feature extraction — reduces the
need for manual human involvement.
Can ingest unstructured data in raw form
(text, images) without human interaction.
Automatically establishes the hierarchy of
characteristics. Also known as 'scalable
machine learning' (Lex Fridman).
WHAT IS RPA?
With RPA, software users develop software robots or 'bots' that are capable of learning,
simulating and executing rules-based business processes. By studying human digital
behaviours, RPA automation enables users to construct bots. Robotic Process Automation
software bots can communicate with any application or system in the same manner that
humans can — but operate around the clock without fatigue, errors or breaks.
RPA bots are simple to configure, utilise and distribute — as easy as pressing record, play
and stop buttons and using drag-and-drop.
RPA bots may be scheduled, copied, altered and shared to conduct enterprise-wide business
operations.
RPA bots function continuously, around-the-clock, with 100 percent accuracy and
dependability.
7 BENEFITS OF RPA
Benefit Description
(i) Higher RPA bots operate continuously around-the-clock, 24/7, without
Productivity breaks or fatigue — dramatically increasing overall output.
(ii) Higher Accuracy Bots function with 100% accuracy and dependability, eliminating
human errors in repetitive tasks.
(iii) Saving of Cost Automating repetitive, labour-intensive tasks reduces operational
costs and frees up human resources for higher-value work.
(iv) Integration Bots can communicate with any application or system — no need to
Across Platforms modify existing corporate systems, apps or processes in order to
automate.
(v) Better Customer Faster processing and fewer errors lead to improved service delivery
Experience and a better experience for customers.
(vi) Harnessing AI RPA can be combined with AI capabilities (such as NLP and ML) to
handle more complex, cognitive tasks beyond simple rule-based
automation.
(vii) Scalability RPA bots can be scheduled, copied, altered and shared to conduct
enterprise-wide business operations — easily scaled up or down as
business needs change.
Approach Description
(i) Supervised Algorithms construct a mathematical model of a dataset that includes
Learning both inputs and expected outcomes (training data). Each training
example has one or more inputs and the expected output (supervisory
signal). By optimising an objective function iteratively, the algorithm
discovers a function to predict outputs for new inputs.
Key concepts: Feature vector, training data matrix, objective function
optimisation.
Sub-types: Classification (outputs limited to a certain set of values,
e.g., email spam filter), Regression (outputs may take any value
within a range), Similarity learning (quantifies how similar/related
two items are — used in ranking, recommendation systems, face
verification).
(ii) Unsupervised Uses a dataset comprising just inputs (unlabeled, unclassified and
Learning uncategorised test data) to identify data structure such as grouping
and clustering. Algorithms identify similarities in the data and
respond based on the presence or absence of such similarities in each
new dataset.
Applications: Density estimation (calculating probability density
function), data feature summary and explanation.
Cluster analysis: Assigning a set of data to subsets (clusters) so that
observations within the same cluster are similar, while observations
from other clusters are different.
(iii) Semi- Intermediate between unsupervised learning (without labelled
Supervised training data) and supervised learning (with completely labelled
Learning training data). When unlabeled data is combined with a small
quantity of labelled data, there is a significant gain in learning
accuracy.
In poorly supervised learning, training labels are noisy, restricted or
inaccurate; yet, these labels are frequently less expensive to acquire,
resulting in larger effective training sets.
Approach Description
(iv) Reinforcement The learning system (agent) interacts with an environment, observing
Learning the results of its actions and learning from rewards or penalties. The
goal is to maximise cumulative reward over time.
Key feature: The agent learns from trial and error — it is not told
which actions to take but must discover which actions yield the most
reward.
Applications: Game playing (AlphaGo), robotics, algorithmic
trading, autonomous vehicles.
(v) Dimensionality Dimensionality reduction is the process of acquiring a set of major
reduction variables in order to reduce the number of random variables under
consideration. In other words, it is the process of lowering the size of
the feature set, which is also referred to as the “number of features.”
NOTE
This section covers: Cloud Computing, Business Intelligence (BI), Artificial Intelligence
(AI), Robotic Process Automation (RPA) and Machine Learning — and their applications in
business and finance. These topics are covered in the complete ICAI textbook (pages not
included in the current upload). Please refer to your ICAI study material for the full content
of Section 11.5.
Key Terms to Know:
◆ Cloud Computing: Delivery of computing services (servers, storage, databases,
networking, software) over the internet ('the cloud') to offer faster innovation, flexible
resources and economies of scale.
◆ Business Intelligence (BI): Technologies, applications and practices for the collection,
integration, analysis and presentation of business information to support better business
decision making.
◆ Artificial Intelligence (AI): Simulation of human intelligence processes by computer
systems, including learning, reasoning and self-correction.
◆ Robotic Process Automation (RPA): Use of software robots or 'bots' to automate highly
repetitive, routine tasks normally performed by knowledge workers.
◆ Machine Learning (ML): A subset of AI — a computer system's ability to learn from
and adapt to data without being explicitly programmed.