0% found this document useful (0 votes)
6 views10 pages

Module 5

Uploaded by

hamdaayoob51
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views10 pages

Module 5

Uploaded by

hamdaayoob51
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Nasscom Digital Edge 101

Module - 5
Module 5: Big Data and Data Science

Topic 1: Evolution of Big Data Analytics

Big Data: A Short Overview

Big Data refers to an enormous amount of data sets that are too large and complex to be
stored, altered, or investigated using classic data processing tools. Today, data is
generated at unprecedented speeds from numerous sources, particularly social media
platforms. For instance, Facebook alone generates more than 500 terabytes of data
daily, encompassing photos, text, voice messages, and videos.

Big data is used extensively in predictive modeling, machine learning, and advanced
analytics to solve complex business problems.

Evolution of Big Data

The evolution of Big Data can be mapped across three distinct phases:

• Big Data Phase 1.0 (1970–2000): Rooted in traditional database management. It


relied heavily on Relational Database Management Systems (RDBMS) and data
warehousing. It focused on structured content, Extract-Transform-Load (ETL)
processes, Online Analytical Processing (OLAP), and standard statistical data
mining.

• Big Data Phase 2.0 (2000–2010): Sparked by the expansion of the internet and
web traffic. Companies like Amazon and Yahoo began analyzing unstructured
and semi-structured data like search logs, IP locations, and click-rates. This
phase popularized web analytics, opinion mining, and social network analysis.

• Big Data Phase 3.0 (2010–Present): Defined by the rise of mobile and sensor-
based devices (the Internet of Things or IoT). It focuses on location-aware
analysis (GPS data), person-centered behavioral tracking, and spatial-temporal
analysis, generating zettabytes of data daily.

Analytics5 Industries Being Reshaped by Big Data

1. Healthcare: Digitizing medical systems to record health metrics and provide


precision medicine.

1|Page
2. Retail: Gaining deeper insights into consumer behavior and preferences to
enhance target marketing.

3. Manufacturing: Becoming heavily data-driven via the Industrial Internet of


Things (IIoT).

4. Finance: Transforming the financial landscape through rapid transaction


analysis and risk modeling.

5. Energy: Optimizing rapidly growing utilities and renewable energy deployment.

9 Industries that Benefit the Most from Big Data Science

1. Retail: Predicting customer desires and personalizing shopping experiences.

2. Medicine: Using wearable trackers to monitor patient vitals and medication


adherence.

3. Banking and Finance: Employing Natural Language Processing (NLP) and


predictive analytics for virtual assistants (e.g., Bank of America's Erica).

4. Construction: Tracking project timelines and material-based costs for optimized


building.

5. Transportation: Mapping passenger journeys and handling unexpected


scenarios (e.g., Transport for London).

6. Media & Entertainment: Understanding real-time consumption patterns to


recommend targeted content (e.g., Spotify).

7. Education: Measuring student engagement, teacher performance, and


demographic trends.

8. Natural Resources/Manufacturing: Conducting geographic, seismic, and


reservoir analyses.

9. Government: Identifying fraud (e.g., Social Security claims) and analyzing


patterns in food-related disorders.

Top 10 Best Big Data Companies

1. AWS (Amazon Web Services): Offers massive cloud data lakes, business
intelligence, and machine learning scaling.

2. Google Cloud: Provides modular cloud services, prominently featuring BigQuery


for serverless, planet-scale ML models.

2|Page
3. Microsoft (Azure): Offers over 200 cloud products for scalable computing,
storage, and analytics.

4. IBM: Provides a highly secure hybrid cloud platform with advanced AI


capabilities.

5. Hewlett Packard Enterprise (HPE): Supplies edge-to-cloud advanced analytics


solutions.

6. Teradata: Delivers analytics at scale, aiding in asset optimization and product


innovation.

7. Cloudera: A hybrid cloud data company offering practitioner-focused analytic


services built by industry veterans.

8. Alteryx: Provides a no-code analytics platform functioning directly within


popular databases.

9. Snowflake: A cloud-native platform offering a data warehouse-as-a-service and


secure Data Exchange.

10. Informatica: Focuses on intelligent, AI-powered data management and


transformation platforms.

10 Latest Trends in BDA (Big Data Analytics)

1. Real-Time Data and Insights: Moving beyond historical data to execute instant
decision-making in finance and media.

2. Automated Decision Making: AI systems diagnosing medical conditions or


rerouting manufacturing lines automatically.

3. Heightened Veracity: Utilizing data observability platforms to flag anomalies


and ensure data integrity.

4. Data Governance: Enforcing strict compliance (GDPR, CCPA) through data


catalogs and certification programs.

5. Cloud Storage and Analytics Platforms: Utilizing infinite cloud storage


(Snowflake, Redshift) to prevent performance bottlenecks.

6. Processing Data Variety: Integrating complex unstructured data via connectors


like Fivetran.

7. Democratization of Data: Empowering non-technical stakeholders to perform


self-service visual analytics.

3|Page
8. No-Code Solutions: Removing coding barriers so business teams can directly
engage with data.

9. Microservices and Data Marketplaces: Breaking monolithic apps into


independent services and augmenting internal data via external marketplaces.

10. Data Mesh: Decentralizing monolithic data lakes into domain-specific "data
products" owned by cross-functional teams.

11. Leveraging GenAI and RAG: Using Generative AI to create synthetic data and
Retrieval-Augmented Generation to ensure contextual accuracy.

Topic 2: An Overview of Big Data Analytics

What Is Big Data?

Big data describes enormous and diverse datasets characterized by high volume,
velocity, and variety. Data formats fall into three main categories: structured
(databases), unstructured (video, text), and semi-structured (JSON, XML). It relies on
data integration, cloud management, and continuous analysis to derive value.

What Is Big Data Analytics? Types, Tools, and Applications

Big data analytics is the process of examining massive volumes of disparate data to
uncover hidden patterns, market trends, and correlations.

Types of Analytics:

• Descriptive Analytics: Summarizes past data into readable formats, such as


revenue reports or social media metrics tables.

• Diagnostic Analytics: Analyzes data to comprehend the root source of an issue


via data mining and drill-down techniques.

• Predictive Analytics: Relies on AI and historical data to forecast future trends,


such as anticipating customer market demands.

• Prescriptive Analytics: Recommends specific solutions to problems using


machine learning, such as an airline algorithm automatically adjusting ticket
prices based on weather and demand.

Tools & Applications: Top tools include Hadoop (storage/analysis), MongoDB (handling
frequently changing data), Talend (integration), Cassandra (handling mass data), and

4|Page
Spark (examining large sets). BDA fuels customer acquisition, targeted ads, price
optimization, and risk management.

Big Data Analytics for Manufacturing

Manufacturers deploy analytics alongside IoT and robotics. They utilize it to improve
product yield, drastically reduce time to market, enhance overall quality, and build
digital prototypes before a physical launch. Furthermore, it enables predictive
maintenance—analyzing machine performance to detect equipment failures before
they cause operational downtime.

What Is Data Visualization?

Data visualization involves representing numerical and abstract data through charts,
graphs, and interactive dashboards. In the context of big data, human brains cannot
easily parse millions of rows in a spreadsheet. Visualization makes complex data
universally understandable across business departments, increasing creativity and
allowing stakeholders to spot outliers and trends rapidly.

Topic 3: Applications of Big Data Analytics

Decoding the 4 Vs of Big Data (With Examples)

Big data is fundamentally characterized by the 4 Vs:

• Volume: The sheer scale of data. Example: Financial institutions managing


colossal transaction records and market data daily.

• Velocity: The speed at which data is generated and processed. Example: Stock
market data updating in milliseconds, requiring instant algorithmic responses.

• Variety: The heterogeneity of data formats. Example: Finance tracking structured


SQL databases alongside unstructured client emails and regulatory filings.

• Veracity: The accuracy, trustworthiness, and quality of the data. Messy or noisy
data can lead to dangerous business decisions.

Using AI and Data for Predictive Planning and Supply Chain

Organizations leverage predictive analytical models to manage complex B2B supplier


networks. Supply chain analytics allow for preemptive inventory replenishment,

5|Page
intelligent route optimizations for shipping, and automated notifications regarding
potential delivery delays.

In-demand Big Data Skills

• Data Analysis: The ability to examine raw datasets to extract actionable insights
and refine business strategies.

• Programming Skills: Proficiency in programming languages like Python, R, and


Java, as well as a strong grasp of data structures and algorithms.

• Big Data Tools Mastery: Familiarity with distributed systems and frameworks
like Hadoop, Spark, and Hive.

• Data Visualization: The ability to translate findings into communicative


dashboards.

Big Data Analytics Tools and Their Key Features

• Apache Hadoop: A highly configurable open-source framework utilizing HDFS


for storage, MapReduce for processing, and YARN for resource management.

• Apache Spark: Offers in-memory data processing up to 100 times faster than
disk-based MapReduce, supporting Java, Scala, and Python APIs.

• Apache Storm: A real-time computation system processing millions of


messages per second with built-in fault tolerance.

• Talend: Automates big data integration via graphical wizards generating native
code for ETL/ELT pipelines.

• Apache CouchDB: A document-oriented NoSQL database storing data in JSON


format, written in Erlang for distributed scaling.

• R & Python: R is tailored for complex statistical computing and matrix


calculations , while Python is highly versatile for ML applications.

• Tableau & Plotly: Leading tools for turning data into eye-catching visualizations
and cloud-hosted graphics.

• Splunk: The premier tool for parsing and analyzing machine-generated log data.

Big Data Analytics in Healthcare: Possibilities and Challenges

Possibilities: Big data revolutionizes healthcare by enabling personalized medicine


based on historical patient conditions. It reduces costs by predicting hospital

6|Page
admissions to optimize resource allocation and shortens hospital stays. It also aids in
advanced R&D for rapid drug discovery and detects fraudulent billing claims.

Challenges: The industry faces severe hurdles, including strict data privacy and
security mandates. Inconsistent data formats lead to integration and interoperability
issues between hospitals. Furthermore, implementing these tools requires massive
infrastructure costs and highly skilled personnel, which are currently in short supply.

Big Data Analytics in Aviation Industry

Airlines utilize big data to maximize profitability and efficiency:

• Centralized Customer Views: Integrating siloed data from online ticket


bookings, search apps, and past travel histories into a single profile.

• Route Optimization: Analyzing travel behaviors to allocate more aircraft to


highly profitable routes, preventing revenue loss from unsold seats.

• Demand Forecasting: Utilizing predictive analytics to estimate future ticket


demand, optimizing fleet and crew allocation.

• Differential Pricing Strategy: Segmenting time-sensitive versus price-sensitive


customers and dynamically pricing tickets to generate maximum revenue per
demographic.

Topic 4: Database Management for Data Science

Database Management Systems (DBMS) Explained

A DBMS is a software application designed to create, optimize, store, and manage


databases, acting as an interface between the data and the end-user. It manipulates file
structures and handles all data requests.

• Core Components: Data repository, database access/query language (like SQL),


management resources (run-time managers), and a query processing engine.

• Key Benefits: Reduces data redundancy, enforces strong data security via
access controls, eliminates data inconsistency, enables secure multi-user
sharing, and provides automated recovery.

• Types of DBMS: Hierarchical (tree-like, one-to-many) , Network (interconnected,


many-to-many) , Relational (tables) , and Object-Oriented (stores data as
objects/classes).

7|Page
Relational Database Management Systems (RDBMS) Explained

An RDBMS stores related data points in a highly structured format utilizing tables
containing columns (attributes) and rows (records). It relies heavily on SQL for data
access and manipulation.

• Core Functions: Facilitates CRUD operations (Create, Read, Update, Delete)


and adheres to ACID properties (Atomicity, Consistency, Isolation, Durability) to
guarantee transaction validity.

• Table Constraints: Relies on constraints to maintain data integrity. A Primary


Key uniquely identifies a row. A Foreign Key links tables together. Constraints
like Not Null and Check ensure valid data entry.

• Differences from Standard DBMS: A standard DBMS generally supports single


users and hierarchical structures, while an RDBMS supports massive datasets,
distributed databases, and concurrent multi-user access.

• Examples: Oracle Database, MySQL, Microsoft Azure SQL, and IBM Db2.

What Is MySQL and How Does It Work?

MySQL is a highly popular, open-source RDBMS written in C and C++. It operates on a


Client-Server Model: clients (devices) connect to a central server via a network and
send requests using SQL statements via a GUI.

• Note on SQL vs. MySQL: SQL is the language used to communicate, whereas
MySQL is the branded software executing the commands.

• Why it is popular: It is highly flexible (open-source GPL license), delivers high


performance for massive eCommerce data clusters, possesses industry-
standard support, and features robust host-based security encryption.

MongoDB: Introduction, Features, and Advantages

MongoDB is a leading NoSQL, open-source database that utilizes a flexible, non-


relational document-oriented data model. Instead of rows and columns, it stores data
in BSON (Binary JSON) format.

• Architecture: The core units are Databases (physical containers), Collections


(equivalent to RDBMS tables, but schema-less), and Documents (sets of
dynamic key-value pairs).

• Key Features: It supports dynamic ad-hoc queries, automated sharding for load
balancing, MapReduce, and GridFS for dividing massive files.

8|Page
• Advantages: Offers immense scalability without downtime , flexible dynamic
schemas for rapid iterative development , and reduced Total Cost of Ownership
(TCO) via its Atlas cloud service running on commodity hardware.

• Drawbacks: It consumes very high memory, limits individual document sizes,


and notably lacks transaction support.

• Use Cases: Ideal for building single-view customer repositories, handling high-
speed IoT data ingestion, real-time analytics, backend gaming architecture, and
high-volume payment processing.

9|Page

You might also like