Business Analytics: Data Types & Importance
Business Analytics: Data Types & Importance
Module 2
• Data is the main ingredient for Business Intelligence (BI), Data Science, and Business
Analytics.
o Information
o Insight
o Knowledge
• Structure:
• Flow:
• Automated collection:
[AUTHOR NAME] 2
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
[AUTHOR NAME] 3
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
• Requires merging/transformation.
• With new tech (data lakes, Hadoop), accessibility is more critical.
➢ Data Security & Privacy
• Only authorized people should access data.
• Prevents unauthorized access & misuse.
• Very critical in sensitive areas (e.g., healthcare).
• Example: HIPAA law ensures patient health records are secure and accessible
only to authorized users.
➢ Data Richness (Comprehensiveness)
• Data should contain all required elements for meaningful analysis.
• Provides dimensionality for deeper insights.
• Should be complete enough to support predictive/prescriptive models.
➢ Data Consistency
• Data should be accurately collected & merged.
• Problem: wrong merging can mix up records (e.g., merging two patient
records).
• Consistency ensures data from multiple sources aligns correctly.
➢ Data Currency / Timeliness
• Data should be up-to-date and recent.
• Recorded close to the event → avoids memory/recall errors.
• Accurate & timely data = reliable analytics.
➢ Data Granularity
• Data should have the required level of detail.
• Example:
o Lab test results must have correct decimal precision.
o Demographic data should be detailed enough to differentiate
subpopulations.
• Rule: Aggregated data cannot be disaggregated, but granular data can be
aggregated.
➢ Data Validity
• Actual data values must match expected values/ranges.
• Example: gender variable → valid entries = male, female, other.
[AUTHOR NAME] 4
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
➢ Data Relevancy
• Only include variables relevant to the study.
• Relevancy = spectrum (least relevant → most relevant).
• Avoid irrelevant data → can mislead algorithms and reduce accuracy.
• Data (datum in singular form) refers to a collection of facts usually obtained as the
result of experiments, observations, transactions, or experiences.
• Data may consist of numbers, letters, words, images, voice recordings, and so on,
as measurements of a set of variables (characteristics of the subject or event that we
are interested in studying).
• Data are often viewed as the lowest level of abstraction from which information and
then knowledge is derived. At the highest level of abstraction, one can classify data as
structured and unstructured (or semi structured).
• Unstructured data/semi structured data is composed of any combi nation of textual,
imagery, voice, and Web content..
• Structured data is what data mining algorithms use and can be classified as categorical
or numeric.
The categorical data can be subdivided into nominal or ordinal data, whereas numeric data can be
subdivided into intervals or ratios.
➢ Categorical Data
[AUTHOR NAME] 5
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
• Labels used to divide variables into groups (e.g., race, age group, education).
• Age and education can be numeric but are often better expressed as categories (e.g.,
“teen,” “adult”).
• Numbers used as codes are only symbols, not for calculations like fractions.
• Types:
➢ Ordinal Data
• Examples:
• Use in analytics: Algorithms like ordinal multiple logistic regression use rank-order
info to improve classification.
• Examples:
o Age
[AUTHOR NAME] 6
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
o Number of children
o Temperature (°F)
• Can be:
➢ Interval Data
• Key points:
o Only differences between values are meaningful; ratios are not meaningful
➢ Ratio Data
• Examples:
• Key points:
[AUTHOR NAME] 7
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
Algorithm Compatibility:
Examples of Requirements:
Data Transformation:
[AUTHOR NAME] 8
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
o Analytics tools help convert and represent data so algorithms can process it
properly.
• Data in its original form (real-world data) is usually not ready for analytics tasks.
• Common issues:
o Dirty
o Misaligned
o Overly complex
o Inaccurate
• Converts raw real-world data into a well-refined form suitable for analytics
algorithms.
• Time factor:
o Often takes longer than the rest of the analytics tasks (model building and
assessment).
[AUTHOR NAME] 9
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
1. Data Collection
Data Blending
• Importance:
[AUTHOR NAME] 10
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
o Data blending is a critical part of data science, the most popular job of the 21st
century.
▪ Database blending
▪ Time blending
▪ Tool blending
o Discretization / Aggregation:
[AUTHOR NAME] 11
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
The final phase of data preprocessing is data reduction, which addresses two main dimensions:
o Dimensional reduction (variable selection) helps retain only the most relevant
features.
o Sampling is used, but the sample must fairly represent the whole dataset.
[AUTHOR NAME] 12
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
o Often expanded to Big Data Analytics, emphasizing its use for generating
insights.
[AUTHOR NAME] 13
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
o Everywhere: web logs, sensors, GPS, RFID, social media, search indexes, call
records, scientific research, medical records, e-commerce, multimedia
archives, etc.
• Evolution:
o The scale has grown from terabytes → exabytes, driven by demand for deeper
insights.
o Variability (inconsistencies/changes)
Big Data is typically described by three core “V”s—Volume, Variety, Velocity—but other
solution providers have added Veracity, Variability, and Value Proposition.
1. Volume
o Refers to the massive size of data, growing from petabytes (PB) to zettabytes
(ZB).
2. Variety
[AUTHOR NAME] 14
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
3. Velocity
o Refers to both the speed of data generation and the speed of data processing
required.
o Real-time data streams (RFID, sensors, GPS, smart meters) demand fast
analytics.
4. Veracity (IBM)
5. Variability (SAS)
o Refers to inconsistent data flows with sudden peaks (e.g., trending social
media events, seasonal surges).
6. Value Proposition
o Large, rich datasets provide more patterns, anomalies, and insights than small
data.
[AUTHOR NAME] 15
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
• Big Data analytics enables deeper, on-demand exploration with new data sources and
technologies.
2. You want to integrate new data types (social media, RFID, sensors, web, GPS, text,
etc.) that don’t fit traditional schemas.
As is the case with any other large IT investment, the success in Big Data analytics depends on a
number of factors. Figure 2.4 shows a graphical depiction of the most critical success factors.
1. Clear Business Need – Big Data investments must align with business vision and
strategy (strategic, tactical, or operational).
3. Business–IT Alignment – Analytics must support business strategy, not dictate it.
5. Strong Data Infrastructure – Combine traditional data warehouses with new Big
Data technologies.
• Grid computing – Shared pool of IT resources for efficiency and lower costs.
[AUTHOR NAME] 17
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
• Solution Cost – High experimentation costs; need cost-effective solutions for ROI.
• Industry-Specific Priorities:
• Brand management
• Enhanced security
[AUTHOR NAME] 18
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
Hadoop
• Key Features:
1. Data Input: Structured, semi-structured, and unstructured data (e.g., logs, social
media, internal stores).
2. Storage (HDFS):
o Name Node: Manages metadata (which nodes hold what data, failures, etc.).
[AUTHOR NAME] 19
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
3. Processing (MapReduce):
o Map Phase: Client submits a “Map” job (query, often Java). Job Tracker
assigns tasks to nodes holding relevant data. Nodes process their share in
parallel.
o Reduce Phase: Results from map tasks are aggregated across nodes to
produce a final answer.
4. Output: Results are returned to the client and can be further analyzed in analytic
tools.
Once the MapReduce phase is complete, the processed data is ready for advanced analysis by data
scientists and analytics professionals. They can:
• Transfer and model the data in relational databases, warehouses, or IT systems for
further analysis or transactional support.
MapReduce
[AUTHOR NAME] 20
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
4. Shuffle/Sort: System merges map outputs and prepares them for reduction.
5. Reduce Phase: Reduce program aggregates counts (e.g., total number of squares by
colour).
Programmers can optimize performance by providing custom shuffle/sort logic or adding combiners
to minimize remote file transfers.
• Application areas: search, graph analysis, text analysis, ML, data transformation.
• Advantages:
Beyond the core of HDFS and MapReduce, Hadoop includes a rich set of ecosystem tools that
extend its functionality for data warehousing, querying, integration, monitoring, and
machine learning.
1. Hive
2. Pig
• Suitable for complex, long data pipelines, where SQL may fall short.
3. HBase
[AUTHOR NAME] 22
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
4. Flume
• Uses distributed agents (e.g., on web servers, apps, mobile devices) to collect and
push data into HDFS.
5. Oozie
• Allows defining and chaining jobs written in MapReduce, Hive, Pig, etc.
• Supports dependencies (e.g., a job starts only after required previous jobs complete).
6. Ambari
• Developed by Hortonworks.
7. Avro
• Useful for parsing, data exchange, and remote procedure calls (RPCs).
8. Mahout
9. Sqoop
• A connectivity tool for transferring data between Hadoop and traditional data stores.
[AUTHOR NAME] 23
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
• Moves data from relational databases (Oracle, Teradata, MySQL, etc.) into Hadoop
(and vice versa).
10. HCatalog
• Allows tools like Pig and Hive to access data without knowing its physical storage
location.
Descriptive Statistics
Descriptive statistics describe the basic features of data. It helps summarize and present data in a
simple, understandable way using numbers, tables, or graphs. It doesn’t make conclusions
about a larger population—only about the sample data available.
These show the center or average value in data. Common measures include Mean, Median, and
Mode.
Mean (Average)
Mean = (Sum of all values) ÷ (Number of values)
It is most commonly used, but it can be affected by extreme values (outliers).
Median
The middle value when data is arranged in order. It is not affected by outliers.
Mode
The value that appears most often in a data set. It is useful for categorical or nominal data.
Measures of Dispersion
These show how spread out or varied the data is. Common measures include Range, Variance,
Standard Deviation, Mean Absolute Deviation, and Quartiles.
[AUTHOR NAME] 24
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
Range
Range = Maximum value – Minimum value. It shows total spread of data.
Variance
Variance measures how far each data point is from the mean. Higher variance means more spread-out
data.
Standard Deviation
It is the square root of variance. It tells how much data values deviate from the mean.
Box-and-Whiskers Plot
This graphical tool shows median, quartiles, minimum, maximum, and outliers. It helps visualize the
spread and skewness of data.
Shape of Distribution
The shape of data distribution shows how data values are spread. The most common shape is the
normal distribution, which is symmetric around the mean.
[AUTHOR NAME] 25
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
Logistic Regression
Used when the output is categorical (like Yes/No, Pass/Fail). It predicts the probability of a
class using a sigmoid (S-shaped) curve. It is widely used in business, health, and social
sciences.
[AUTHOR NAME] 26
RLJIT Dept Of CSE (Data science)
BUSINESS ANALYTICS
[AUTHOR NAME] 27
RLJIT Dept Of CSE (Data science)