BCSE402L
Big Data Analytics
Module -1
Overview of Big data
analytics
Introduction to Big Data
• Big Data refers to vast, complex datasets - structured and
unstructured that traditional tools (Example: Excel) cannot
handle.
Why it matters:
• Extract actionable insights to drive innovation.
• 2025 Stat: Global data volume is expected to hit 182 zettabytes
this year
The Need for Big Data
• Explosion of Data Sources:
Over 75 billion IoT devices generate continuous data
streams
• Real-Time Insight Demands:
Organizations must act immediately on fresh data - from
preventing fraud to optimizing traffic
• Business Intelligence:
3 in 5 companies use analytics to spark innovation;
90% report ROI on their investments
"90% report ROI on their
investments"
90% of companies (or respondents) said they gained measurable
financial or strategic benefit (Return on Investment – ROI) from what
they invested in.
• Companies investment in Big Data technologies, tools, training, or
infrastructure resulted in:
• Improved efficiency
• Cost savings
• Better customer insights
• Increased sales or productivity
• Essentially, 9 out of 10 companies saw real, valuable outcomes from
adopting Big Data strategies.
Core Characteristics – The Big
Data 5 Vs
1. Volume: Massive scale (zettabytes)
2. Velocity: Real-time streaming analysis
3. Variety: Mix of textual, sensor, video, IoT data
4. Veracity: Ensuring trust, accuracy and quality via
governance
5. Value: Unlocking benefits like personalization, cost
reduction, and predictive insights.
Industry Applications & Trends
(2025)
• Fast Food & Retail: AI‑driven supply chain optimization
reduces waste (McDonald’s, Starbucks)
• Energy & Utilities: AI for predictive grid maintenance; GenAI
assistants aiding field technicians
• Disaster Management: Satellite and AI models forecasting
bushfires and storms globally
• Consulting & Finance: KPMG invests $100 M in AI analytics;
* Many more
Emerging Trends & Future
Outlook
• Edge Computing & Hybrid Pipelines: Close-to-source
processing mixed with batch systems.
• Data-as-a-Service (DaaS): Access scalable, on-demand
datasets.
• Quantum & GenAI: Advancing compute power and AI
generative capabilities.
• Ethics & Governance: Reporting, data privacy, and
responsible AI are becoming essential.
Types of Big Data
1. Structured data:
This kind of data is the simplest to organize and search. It can include
things like financial data, machine logs and demographic details. An
Excel spreadsheet with its layout of pre-defined columns and rows is a
good way to structure data.
2. Unstructured data:
This category of data can include things like social media posts, audio
files, images and open-ended customer comments. This kind of data
cannot be easily captured in standard row-column relational databases.
3. Semi-structured data:
As it sounds, semi-structured data is a hybrid of structured and
unstructured data. E-mails are a good example as they include
unstructured data in the body of the message as well as more
organizational properties such as sender, recipient, subject and date.
Multi-structured data
• Multi-structured data in big data analytics refers to the combination
of structured, semi-structured, and unstructured data types within a
single environment.
• This data comes from diverse sources and applications, posing unique
challenges for analysis and integration.
• Effectively managing and analyzing multi-structured data is crucial for
gaining comprehensive insights and driving informed business
decisions.
Example
Structured component of data:
Each chess moves is recorded in a table in each match that players refer in
future. The records consist of serial numbers (row numbers, which mean
move numbers) in the first column and the moves of White and Black in
two subsequent vertical columns. Volume of data, i.e. data used for
analyzing erroneous or best moves in the matches, keeps growing with
more and more tables, and may eventually become 'voluminous data'.
Unstructured component of data
Social media generates data after each international match. The media
publishes the analysis of classical matches played between Grand
Masters. The data for analyzing chess moves of these matches are thus
in a variety of formats.
Multi-structured data
The voluminous data of these matches can be in a structured format
(i.e. tables) as well as in unstructured formats (i.e. text documents,
news columns, blogs, Facebook etc.). Tools of multi-structured data
analytics assist the players in designing better strategies for winning
chess championships.
Yet other Classification - Big data
types
1. Social networks and web data
2. Transactions data and Business Processes (BPs) data
3. Customer master data
4. Machine-generated data
5. Human-generated data
Sources of Big Data
• Social Media: Billions of tweets, posts, videos and photos are
uploaded every day.
• IoT Devices: Sensors in vehicles, factories, wearable devices and
smart appliances generate continuous streams of data.
• Scientific Instruments: Particle accelerators, space telescopes and
genome sequencers produce huge datasets for research.
• Business Transactions: Every time we shop online or swipe a credit
card, we produce transaction data.
• Public Sector Records: Health records, weather data, public safety
databases add to the data landscape.
Big data handling techniques
• Huge data volumes storage, data distribution, high-speed networks
and High-Performance Computing
• Applications scheduling using open source, reliable, scalable,
distributed file system, distributed database, Parallel and Distributed
computing systems such as Hadoop and Spark.
• Open source tools which are scalable, elastic and provide virtualized
environment, clusters of data nodes, task and thread management.
• Data management using NoSQL, Document database or in-memory
data management using column or Parquet format.
• Data mining and analytics, data retrieval, data reporting , data
visualization and ML Big data tools.
• Columnar storage is ideal for aggregation queries, such as calculating
sums, averages, or counts. These types of operations often focus on
specific columns.
• With Parquet, only the relevant columns need to be read into
memory, which not only improves query speed but also reduces the
overall resource usage.
• Apache Parquet is an efficient columnar storage format.
• Compared to saving this dataset in csvs using parquet: Greatly
reduces the necessary disk space. Loads the data into Pandas with
memory efficient datatypes.
Reading and Writing the Apache
Parquet Format
• import [Link]
• Refer:
[Link]
Big Data – Life cycle
• Business Case/Problem Definition
• Data Identification
• Data Acquisition and filtration
• Data Extraction
• Data Munging(Validation and Cleaning)
• Data Aggregation & Representation(Storage)
• Exploratory Data Analysis
• Data Visualization(Preparation for Modeling and Assessment)
• Utilization of analysis results
Classification of Big Data
Analytics
• Descriptive Analytics: This type helps us understand past events. In
social media, it shows performance metrics, like the number of likes
on a post.
• Diagnostic Analytics: In Diagnostic analytics delves deeper to uncover
the reasons behind past events. In healthcare, it identifies the causes
of high patient re-admissions.
• Predictive Analytics: Predictive analytics forecasts future events
based on past data. Weather forecasting, for example, predicts
tomorrow's weather by analyzing historical patterns.
• Prescriptive Analytics: However, this category not only predicts
results but also offers recommendations for action to achieve the best
results. In e-commerce, it may suggest the best price for a product to
achieve the highest possible profit.
• Real-time Analytics: The key function of real-time analytics is data
processing in real time. It swiftly allows traders to make decisions
based on real-time market events.
• Spatial Analytics: Spatial analytics is about the location data. In urban
management, it optimizes traffic flow from the data under the sensors
and cameras to minimize the traffic jam.
• Text Analytics: Text analytics delves into the unstructured data of
text. In the hotel business, it can use the guest reviews to enhance
services and guest satisfaction.
Evolution of Big Data and its
characteristics
• EDP, IS, and MIS all relate to the use of technology in
business, but with slightly different focuses.
• EDP (Electronic Data Processing) refers to the basic use of
computers to process data, essentially the automation of
data handling tasks.
• IS (Information Systems) is a broader term encompassing the
systems that collect, process, store, and distribute
information, including hardware, software, and people.
• MIS (Management Information Systems) specifically focuses
on using information systems to support management
decision-making, providing relevant data for planning,
control, and operational activities.
• ERP (Enterprise Resource Planning) is a type of software that
integrates all aspects of a business, including finance, HR,
and supply chain management, into a unified system. In
essence, EDP is a foundational concept, while MIS and ERP
are more specific applications or systems that build upon it.
Scalability and Parallel Processing
• Co-exist with traditional data store (RDBMS tables / data warehouse)
• Big data requires Scale up and Scale out capabilities.
• i.e. Vertical and Horizontal Computing resources
• Therefore, Computing and storage when run in parallel, enables scale
out.
• Scalability – Increase/decrease in the capacity of data storage,
processing and analytics.
• With increased workload – system capability has to be auto adjusted.
1. Analytics Scalability to Big data
• Vertical Scalability
- Scaling up the given system’s resources and increasing the system’s
analytics, reporting and visualization capabilities.
- Designing the algorithm according to the architecture that uses
resources efficiently.
Real-World Analogy:
Think of vertical scalability like upgrading from a bicycle to a
motorbike. You are still using one vehicle, but it's now more powerful
and faster.
Use case
Imagine you have a retail analytics system running on a single server.
This system:
• Collects sales data from stores
• Generates reports
• Visualizes inventory trends
Initially, the server has:
• 4 CPU cores
• 8 GB RAM
As your business grows, the system struggles to process real-time
reports.
To Do (Vertical Scaling):
Upgrade the server to:
• 16 CPU cores
• 64 GB RAM
Now the system can handle more data, generate faster reports, and
show real-time dashboards with better visualization.
Designing Algorithms Efficiently:
While scaling up, the algorithm must be designed to make use of these
new resources.
For example:
• Use multi-threading to utilize all CPU cores.
• Optimize memory usage to process large datasets in RAM.
• Use in-memory processing tools like Apache Spark for faster analytics.
• Horizontal Scalability
- Increasing the number of systems working in coherence and scaling
out the workload.
- Using more resources and distributing the processing and storage
tasks in parallel.
Use case
Imagine the same retail analytics system. It’s now being used by
hundreds of stores across the country.
A single server, no matter how powerful, can’t handle the increasing
load.
To Do (Horizontal Scaling):
Add multiple servers and connect them in a distributed system:
• Each server handles data from a group of stores.
• A load balancer distributes incoming requests across servers.
• Data is stored and processed in distributed databases (e.g., MongoDB,
Cassandra, Hadoop HDFS).
Now, the system:
• Supports more users
• Processes more transactions in parallel
• Generates reports faster by splitting the workload
Designing Algorithms Efficiently:
Algorithms must be designed to:
• Distribute tasks among multiple servers
• Handle synchronization and data consistency
• Avoid bottlenecks (e.g., using MapReduce for batch processing)
• For example, Apache Hadoop or Apache Spark splits a data analytics job
into chunks and processes them across a cluster of machines.
Real-World Analogy:
Think of horizontal scalability like adding more delivery trucks instead of
just upgrading one. You deliver more packages faster by dividing the work.
Key Differences
Vertical Scaling Horizontal Scaling
Add more power to one machine Add more machines
Easier to implement More complex (needs
coordination)
Limited by hardware capacity Virtually unlimited scalability
Cons of using more CPUs
• Implement using bigger machines (More CPU cores)
• High config CPUs, RAM modules, Hard disks, Faster and Bigger
Motherboards would be expensive as well as the software that is used
for design of algorithms does not exploit the advantage of them
• So no performance would be increased due to additional CPUs.
• Alternative - MPPs
2. Massively Parallel Processing
Platforms (MPP)
• Scaling - Parallel Processing Systems
• Computational problem is broken into discrete pieces of sub-tasks to
get processed simultaneously.
• Parallelization of tasks can be done at:
1. Distributing separate tasks onto separate threads on the same CPU
2. Distributing separate tasks onto separate CPUs on the same
computer
3. Distributing separate tasks onto separate computers
Distributed computing paradigms
• Uses grid, cloud , cluster computing model
3. Cloud computing
• Shared processing resources and data to the computers and other
devices on demand.
• To perform parallel and distributed computing in a cloud-computing
environment.
• Cloud resources - AWS, EC2, Azure or Apache CloudStack, Amazon
Simple Storage Service (S3)
Cloud services
• IaaS
• PaaS
• SaaS
Key differences
IaaS (Infrastructure PaaS (Platform as a SaaS (Software as
Aspect as a Service) Service) a Service)
Virtualized Ready-to-use
What it provides hardware (servers, Development software
storage, networks) platform and tools applications
OS, apps, Just use the app
User controls middleware, Apps and data (no control over
runtime backend)
For whom? IT Admins / DevOps Developers End Users
AWS EC2, Microsoft Google App Engine,
Example services Azure VM, Google AWS Elastic Gmail, Microsoft
Compute Engine Beanstalk 365, Salesforce
Questions for you
Imagine that You are part of an IT consulting team. A client approaches you
with the following requirements.
Identify whether they should use IaaS, PaaS, or SaaS.
1) Startup launching a new mobile app
2) Large enterprise wants to migrate its internal IT infrastructure to reduce
hardware costs
3) School wants to adopt an online platform for email, file sharing, and
collaboration for students and teachers
4) Gaming company wants a backend infrastructure that can auto-scale and
allows them to control game logic and security
5) Small business wants to use customer relationship management (CRM)
tools like Salesforce without technical headaches
1) Startup launching a new mobile app - PAAS
2) Large enterprise wants to migrate its internal IT
infrastructure to reduce hardware costs - IAAS
3) School wants to adopt an online platform for email, file
sharing, and collaboration for students and teachers - SAAS
4) Gaming company wants a backend infrastructure that can
auto-scale and allows them to control game logic and
security - IAAS
5) Small business wants to use customer relationship
management (CRM) tools like Salesforce without technical
headaches - SAAS
4. Grid and Cluster Computing
• Grid computing
Group of computers from several locations are connected with each
other to achieve a common task.
The computer resources are heterogeneously and geographically
disperse.
A group of computers that might spread over remotely comprise a
grid.
A grid is used for a variety of purposes.
Features of Grid computing
• Scalable
• Sharing of resources to attain coordination and coherence among
resources similar to grid computing
• Grid is a form of distributed network for resource integration.
Drawbacks of Grid computing
• Single point, which leads to failure in case of underperformance or
failure of any of the participating nodes.
• A system’s storage capacity varies with the number of users, instances
and the amount of data transferred at a given time.
• Sharing resources among a large number of user helps in reducing
infrastructure costs and raising load capabilities.
Skeleton of Grid computing
Cluster computing
• A cluster is a group of computers connected by a network.
• Works together to accomplish the same task.
• Load balancing
• They shift processes between nodes to keep an even load on the
group of connected computers.
Drawbacks of cluster computing
• Setting up and managing a cluster involves complex configurations,
network management, and synchronization between nodes are
complex
• High Initial Cost
• Scalability Limitations
• Fault Tolerance Challenges
• Not all software applications are optimized for cluster environments;
rewriting or adapting code for parallel execution may be required.
Grid computing and related
paradigms
5. Volunteer computing
• Provide computing resources to projects of importance that use
resources to do distributed computing and/or storage.
• A distributed computing paradigm which uses computing resources of
the volunteers.
• Volunteers – Organizations / members who own personal computers.
• Distributed computing encompasses the general concept of breaking
down a task to be processed across multiple computers
• Grid and cluster computing represent specific architectures for
achieving this.
Volunteer computing
• BOINC – an open source
software for Volunteer and grid
computing
• Berkeley Open Infrastructure for
Network Computing
• https://
[Link]/BOINC/boinc/wiki/V
olunteerComputing
Data Storage and analysis
• Data store with Structured and Semi-structured data
• SQL
• Large data storage using RDBMS
• DDBMS
• In-Memory Column Formats Data (OLAP)
• In-Memory Row Formats Data (OLTP)
• Enterprise Data-Store Server and Data Warehouse (Slide 57)
import pandas as pd
# Read only specific columns from a Parquet file
df = pd.read_parquet('[Link]', columns=['column_name'])
print(df)
import [Link] as pq
# Open the Parquet file
table = pq.read_table('[Link]', columns=['column_name'])
# Convert to pandas DataFrame if needed
df = table.to_pandas()
print(df)
Big Data Storage
• Amazon S3 – uninterrupted key value
import boto3
s3 = [Link]('s3')
# Upload an object with an uninterrupted key
s3.put_object(Bucket='my-bucket', Key='userlog20250719', Body='Sample data')
# Retrieve it
response = s3.get_object(Bucket='my-bucket', Key='userlog20250719')
print(response['Body'].read().decode())
• Unordered Keys - MongoDB (JSON)
• Ordered keys – Cassandra (JSON)
• No fixed table schema
• No Joins
• Fault tolerant
• CAP theorem (atleast 2 out of 3)