0% found this document useful (0 votes)
3 views51 pages

01 Introduction

The document provides an introduction to big data analytics, covering key concepts such as data, data analytics, data mining, and the significance of big data in today's digital age. It discusses various sources of data, the importance of analyzing data to extract useful knowledge, and the challenges posed by the vast amounts of data generated. Additionally, it highlights advanced techniques used in big data analysis to uncover hidden patterns and make informed decisions.

Uploaded by

relentlessboy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views51 pages

01 Introduction

The document provides an introduction to big data analytics, covering key concepts such as data, data analytics, data mining, and the significance of big data in today's digital age. It discusses various sources of data, the importance of analyzing data to extract useful knowledge, and the challenges posed by the vast amounts of data generated. Additionally, it highlights advanced techniques used in big data analysis to uncover hidden patterns and make informed decisions.

Uploaded by

relentlessboy
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

VNUHCM - University of Science

Faculty of Information Technology

CS435 – Lecture 1
Introduction

Le Thi Nhan
ltnhan@[Link]

Designed by SlidesCarnival
Content
» Data
» Data analytics
» Data mining
» Big data
» Big data analysis

» Why do we study these?

Intro2 Big Data Analytics - FIT - HCMUS 2


1.
Data
We are actually living in the data age
Where data come from
» Services of Internet
» Internet of things
» Mobile devices
» Explosion of biological data
» Internal data of enterprises
» Machine logging data
» …

Intro2 Big Data Analytics - FIT - HCMUS 4


Where data come from (cont.)
» Services of Internet
⋄ Personal user information
⋄ User-generated content
⋄ Social media data
− Tweets of microblog Twitter
− Comments from Facebook, LinkedIn
− Social graphs, social networks
− Digital pictures
− Clips, videos
− Audio

Intro2 Big Data Analytics - FIT - HCMUS 5


Where data come from (cont.)
» Bio-medical data
⋄ Omics data
− Genomics data
− Proteomics data
− Metabolomics data
⋄ Clinical data

» Internet of things
⋄ Data about environment, location, movement,
temperature, weather… from sensors/devices

Intro2 Big Data Analytics - FIT - HCMUS 6


Where data come from (cont.)
» Internal data of enterprises
⋄ Historically static data
− Online trading data
− Online analysis data
− …
⋄ Production/inventory/sales/financial data
⋄ Customer feedback
⋄ Emails

Intro2 Big Data Analytics - FIT - HCMUS 7


What do we do with data?
» We know
⋄ Analyzing data is an important need

» We want to
⋄ Exploit the data
⋄ Extract new and useful knowledge from the
data

» In short
⋄ Understand and use the data

Intro2 Big Data Analytics - FIT - HCMUS 8


Jiawei Han, 2012

Intro2 Big Data Analytics - FIT - HCMUS 9


2.
Data analytics
The science of analyzing data
Definition
» A process of
⋄ Inspecting, cleansing, transforming and
modeling data

» With the goal of


⋄ Discovering useful information,
⋄ Informing conclusion, and
⋄ Supporting decision making

Intro2 Big Data Analytics - FIT - HCMUS 11


Category

From Gartner 2017

Intro2 Big Data Analytics - FIT - HCMUS 12


Example of Descriptive
» Dr. John Snow and the Broad Street Pump

» Problem
⋄ How cholera is contracted and spread

» Analysis methods
⋄ Observation, visualization, comparison

Intro2 Big Data Analytics - FIT - HCMUS 13


Example (cont.)
» London in the 1850’s
» Disease was rife in the poorer parts of the city,
and cholera was among the most feared
⋄ People died within a day or two of contracting it
⋄ Hundreds could die in a week
⋄ The total number of deaths in a single wave could
reach tens of thousands
» It was not yet known that germs cause disease
» People thought that “miasmas” were the main
culprit
⋄ Bad smells arising out of decaying matter
Intro2 Big Data Analytics - FIT - HCMUS 14
Dr. John Snow’s observation
» While entire households were wiped out by cholera,
neighbors sometimes remained completely
unaffected
» As they were breathing the same air as their
neighbors
⋄ There was no association between bad smells and
cholera
» The disease always involved vomiting and diarrhea
⋄ The infection was carried by something people ate or
drank, not by the air that they breathed
» A prime suspect was water contaminated by
sewage

Intro2 Big Data Analytics - FIT - HCMUS 15


Dr. John Snow’s visualization

Intro2 Big Data Analytics - FIT - HCMUS From wiki 16


Dr. John Snow’s comparison
» Only contaminated water was causing the
spread of cholera???

» Gathering data on cholera deaths in an area of


London that was served by two water companies
⋄ The Lambeth water company drew its water
upriver from where sewage was discharged into
the River Thames à clean water
⋄ The Southwark and Vauxhall (S&V) company
drew its water below the sewage discharge à
contaminated water

Intro2 Big Data Analytics - FIT - HCMUS 17


John Snow’s comparison (cont.)

Supply Area Number of Cholera Deaths per


houses deaths 10,000 houses
S&V 40,046 1,263 315
Lambeth 26,107 98 37
Rest of London 256,423 1,422 59

Intro2 Big Data Analytics - FIT - HCMUS 19


John Snow’s comparison (cont.)
» There is no difference whatever in
⋄ The houses or the people receiving the supply of
the two water companies
⋄ Any of the physical conditions with which they are
surrounded
» One group being supplied with water containing
the sewage of London, and amongst it, whatever
might have come from the cholera patients, the
other group having water quite free from
impurity…

» The only difference was in the water supply

Intro2 Big Data Analytics - FIT - HCMUS 20


In brief
» Dr. John Snow established a method for the field of
epidemiology
⋄ Whether the treatment has an effect on the outcome?
» Randomized controlled trial (RCT)
⋄ Step 1 - Choose a group of volunteers (study objects)
⋄ Step 2 - Apply statistical methods to divide objects into
2 groups randomly
− Treatment group receives the treatment
− Control group does not receive the the treatment
⋄ Step 3 - Keep track objects for a while to collect data
(outcome)
⋄ Step 4 – Compare outcomes between the 2 groups to
know whether the treatment is effective or not

Intro2 Big Data Analytics - FIT - HCMUS 21


3.
Data mining
Knowledge Discovery from Data (KDD)
What is data mining?
» Given lots of data
» Discover patterns and models
⋄ Valid
− Hold on new data with some certainty
⋄ Useful
− Should be possible to act on the item
⋄ Unexpected
− Non-obvious to the system
⋄ Understandable
− Humans should be able to interpret the pattern

Intro2 Big Data Analytics - FIT - HCMUS 23


Pattern and Model
» Pattern
⋄ A low level summary of a relationship which
holds only for a few records/variables in data

» Model
⋄ A mathematical representation to recognize
patterns

Intro2 Big Data Analytics - FIT - HCMUS 24


Pattern and Model (cont.)

[Link]/2014/pattern-recognition-systems-convey-learning-1205

Intro2 Big Data Analytics - FIT - HCMUS 25


Pattern and Model (cont.)

[Link]/reactive-lda-library/
Intro2 Big Data Analytics - FIT - HCMUS 26
Data mining & Machine learning
DM ML
» To find new and » To build computer
useful knowledge systems that learn
from large datasets as well as human
does

» Knowledge » Effort of building


discovery from artificial
databases intelligence

Intro2 Big Data Analytics - FIT - HCMUS 27


Data mining
» Database
⋄ Large-scale data, simple queries
⋄ Data mining is a form of analytic processing
− Result is the query answer

» Machine learning
⋄ Small data, complex models
⋄ Data mining is inference of models
− Result is the parameters of the model

» Computer science theory


⋄ Algorithms

Intro2 Big Data Analytics - FIT - HCMUS 28


Disciplines

Ar
Pattern

tifi
c
Recognition

ial
s
tic

Int
tis

ellig
Sta

en
ce
DATA Machine
MINING Learning

Mathematical
Modeling Databases

Management Science &


Information Systems
Sharrda, Business Intelligence, 3rd

Intro2 Big Data Analytics - FIT - HCMUS 29


Data mining process

Jiawei Han, 2012

Intro2 Big Data Analytics - FIT - HCMUS 30


Data mining tasks
» Descriptive methods
⋄ Find human-interpretable patterns that describe
the data
⋄ E.g. clustering, association analysis,
summarization…

» Predictive methods
⋄ Use some variables to predict unknown or future
values of other variables
⋄ E.g. decision tree, neural network, support vector
machine, hidden Markov model…

Intro2 Big Data Analytics - FIT - HCMUS 31


Data mining tasks (cont.)
Data Mining Learning Method Popular Algorithms

Classification and Regression Trees,


Prediction Supervised
ANN, SVM, Genetic Algorithms

Decision trees, ANN/MLP, SVM, Rough


Classification Supervised
sets, Genetic Algorithms

Linear/Nonlinear Regression, Regression


Regression Supervised
trees, ANN/MLP, SVM

Association Unsupervised Apriory, OneR, ZeroR, Eclat

Link analysis Unsupervised Expectation Maximization, Apriory


Algorithm, Graph-based Matching

Sequence analysis Unsupervised Apriory Algorithm, FP-Growth technique

Clustering Unsupervised K-means, ANN/SOM

Outlier analysis Unsupervised K-means, Expectation Maximization (EM)

Intro2 Big Data Analytics - FIT - HCMUS Sharrda, Business Intelligence, 3rd 32
Meaningfulness of Answers
» Risk : you will “discover” patterns that are
meaningless

» Statisticians call it Bonferroni’s principle


⋄ If you look in more places for interesting patterns
than your amount of data will support, you are
bound to find crap

» When looking for a property, make sure that the


property does not allow so many possibilities that
random data will surely produce facts “of
interest.”

Intro2 Big Data Analytics - FIT - HCMUS 33


4.
Big data
It's not big, it's just bigger
IDC’s report
(International Data Corporation)

» In 2011, the overall created and copied data


volume in the world is 1.8 Zettabytes (1021B)

» Increased by nearly 9 times within 5 years


» Will double at least every other 2 years in the
future

» Expect 175 Zettabytes of data worldwide by


2025
Intro2 Big Data Analytics - FIT - HCMUS 35
IBM’s report

» 2.5 Exabytes (1018B) of data was generated


every day in 2012

» About 75% of data is unstructured, coming


from sources such as text, voice and video

Intro2 Big Data Analytics - FIT - HCMUS 36


[Min Chen, 2014]
» Services of Internet
⋄ Google processes data of hundreds of
Petabyte (1015B)
⋄ Facebook generates log data of over 10PB per
month
⋄ 72 hours of video are uploaded to YouTube in
every minute
⋄ Baidu processes data of tens of PB
⋄ Taobao, Alibaba generates data of tens of
Terabyte (1012B) for online trading per day

Intro2 Big Data Analytics - FIT - HCMUS 37


What is big data?
» 2001, Doug Laney, Gartner’s analyst
» The 3V model
⋄ Volume
− The quantity of data
− The size of the data
⋄ Velocity
− The speed of generation of data
− How fast the data is generated and processed to meet
the demands
⋄ Variety
− The category which data belongs to

Intro2 Big Data Analytics - FIT - HCMUS 38


What is big data? (cont.)
» 2010, Apache Hadoop
⋄ Datasets which could not be captured, managed, and
processed by general computers within an acceptable
scope
» 2011, McKinsey & Company
⋄ Datasets which could not be acquired, stored and
managed by classic database software

» Criterions for big data


⋄ Volume of a dataset
⋄ Growing data scale
⋄ Big data management
− Not be handled by traditional database technologies

Intro2 Big Data Analytics - FIT - HCMUS 39


What is big data? (cont.)
» 2011, IDC
⋄ Big data technologies describe a new generation
of technologies and architectures, designed to
economically extract value from very large
volumes of a wide variety of data, by enabling the
high-velocity capture, discovery, and/or analysis

» Definition has 4Vs

» “You could only own a bunch of data other than


big data if you do not utilize the collected data” -
Jay Parikh, Facebook

Intro2 Big Data Analytics - FIT - HCMUS 40


What is big data? (cont.)
» Currently, IBM, Gartner, Cisco…

Data in many
Data at Scale Forms Data Uncertainty
Data in Motion
Terabytes to Structured, Precise data &
Streaming data
Petabytes unstructured, text, correct analyses
multimedia

Intro2 Big Data Analytics - FIT - HCMUS 41


In brief

» Datasets that are typically beyond the ability of a single machine


or a modern data management system to store and analyze

» Brings new opportunities for discovering new values


» Helps us to gain an in-depth understanding of the hidden values
» Incurs new challenges
⋄ How to effectively organize and manage

» In future
⋄ Data will exceed the strength of IT architecture and infrastructure
⋄ Requirements of real-time data processing
⋄ Putting pressure on the existing systems

Intro2 Big Data Analytics - FIT - HCMUS 42


In brief (cont.)
» New technologies, methods, skills
⋄ Data acquisition
− Oracle NoSQL database
− Amazon Simple Storage Service (S3)
− Google BigTable
⋄ Data organization
− Hadoop/MapReduce
− Spark
⋄ Data analysis
⋄ Data decision making
Intro2 Big Data Analytics - FIT - HCMUS 43
5.
Big data analysis
What is big data analysis?
» The use of advanced analytic techniques
against very large & diverse datasets to
uncover
⋄ Hidden patterns, unknown correlations,
⋄ Trends, preferences,
⋄ Other useful information

» Purpose
⋄ To make better and faster decisions

Intro2 Big Data Analytics - FIT - HCMUS 45


Advanced techniques
» Text analytics
» Machine learning
» Predictive analytics
» Data mining
» Statistics
» Natural language processing

Intro2 Big Data Analytics - FIT - HCMUS 46


Solution
» Push computation to the data, instead of
pushing data to a computing node
⋄ The MapReduce programming paradigm

» Some examples
⋄ Apache, HortonWorks, Cloudera, and Teradata

Intro2 Big Data Analytics - FIT - HCMUS 47


6.
Why study this course
The fact
» We are living in the information era
» Information/data is very important

» We should know
⋄ How to own the data
⋄ How to analyze data

Intro2 Big Data Analytics - FIT - HCMUS 49


In my opinion
» For those who will work
⋄ I am not sure all companies in VN where have
applied data mining
− But a few do
⋄ You will not use data mining right after leaving
school
− But maybe in the future you will

» For those who continue studying


⋄ Have strong background in CS field
⋄ Create something new that was used by people

Intro2 Big Data Analytics - FIT - HCMUS 50


Some companies
» trustingsocial ([Link])
» East Agile ([Link])
» Sentifi ([Link])
» Inspectrorio ([Link])
» VNG ([Link])
» Zalora ([Link])
» Momo ([Link])
» …

Intro2 Big Data Analytics - FIT - HCMUS 51


THANKS!
Any questions?

52

You might also like