VNUHCM - University of Science
Faculty of Information Technology
CS435 – Lecture 1
Introduction
Le Thi Nhan
ltnhan@[Link]
Designed by SlidesCarnival
Content
» Data
» Data analytics
» Data mining
» Big data
» Big data analysis
» Why do we study these?
Intro2 Big Data Analytics - FIT - HCMUS 2
1.
Data
We are actually living in the data age
Where data come from
» Services of Internet
» Internet of things
» Mobile devices
» Explosion of biological data
» Internal data of enterprises
» Machine logging data
» …
Intro2 Big Data Analytics - FIT - HCMUS 4
Where data come from (cont.)
» Services of Internet
⋄ Personal user information
⋄ User-generated content
⋄ Social media data
− Tweets of microblog Twitter
− Comments from Facebook, LinkedIn
− Social graphs, social networks
− Digital pictures
− Clips, videos
− Audio
Intro2 Big Data Analytics - FIT - HCMUS 5
Where data come from (cont.)
» Bio-medical data
⋄ Omics data
− Genomics data
− Proteomics data
− Metabolomics data
⋄ Clinical data
» Internet of things
⋄ Data about environment, location, movement,
temperature, weather… from sensors/devices
Intro2 Big Data Analytics - FIT - HCMUS 6
Where data come from (cont.)
» Internal data of enterprises
⋄ Historically static data
− Online trading data
− Online analysis data
− …
⋄ Production/inventory/sales/financial data
⋄ Customer feedback
⋄ Emails
Intro2 Big Data Analytics - FIT - HCMUS 7
What do we do with data?
» We know
⋄ Analyzing data is an important need
» We want to
⋄ Exploit the data
⋄ Extract new and useful knowledge from the
data
» In short
⋄ Understand and use the data
Intro2 Big Data Analytics - FIT - HCMUS 8
Jiawei Han, 2012
Intro2 Big Data Analytics - FIT - HCMUS 9
2.
Data analytics
The science of analyzing data
Definition
» A process of
⋄ Inspecting, cleansing, transforming and
modeling data
» With the goal of
⋄ Discovering useful information,
⋄ Informing conclusion, and
⋄ Supporting decision making
Intro2 Big Data Analytics - FIT - HCMUS 11
Category
From Gartner 2017
Intro2 Big Data Analytics - FIT - HCMUS 12
Example of Descriptive
» Dr. John Snow and the Broad Street Pump
» Problem
⋄ How cholera is contracted and spread
» Analysis methods
⋄ Observation, visualization, comparison
Intro2 Big Data Analytics - FIT - HCMUS 13
Example (cont.)
» London in the 1850’s
» Disease was rife in the poorer parts of the city,
and cholera was among the most feared
⋄ People died within a day or two of contracting it
⋄ Hundreds could die in a week
⋄ The total number of deaths in a single wave could
reach tens of thousands
» It was not yet known that germs cause disease
» People thought that “miasmas” were the main
culprit
⋄ Bad smells arising out of decaying matter
Intro2 Big Data Analytics - FIT - HCMUS 14
Dr. John Snow’s observation
» While entire households were wiped out by cholera,
neighbors sometimes remained completely
unaffected
» As they were breathing the same air as their
neighbors
⋄ There was no association between bad smells and
cholera
» The disease always involved vomiting and diarrhea
⋄ The infection was carried by something people ate or
drank, not by the air that they breathed
» A prime suspect was water contaminated by
sewage
Intro2 Big Data Analytics - FIT - HCMUS 15
Dr. John Snow’s visualization
Intro2 Big Data Analytics - FIT - HCMUS From wiki 16
Dr. John Snow’s comparison
» Only contaminated water was causing the
spread of cholera???
» Gathering data on cholera deaths in an area of
London that was served by two water companies
⋄ The Lambeth water company drew its water
upriver from where sewage was discharged into
the River Thames à clean water
⋄ The Southwark and Vauxhall (S&V) company
drew its water below the sewage discharge à
contaminated water
Intro2 Big Data Analytics - FIT - HCMUS 17
John Snow’s comparison (cont.)
Supply Area Number of Cholera Deaths per
houses deaths 10,000 houses
S&V 40,046 1,263 315
Lambeth 26,107 98 37
Rest of London 256,423 1,422 59
Intro2 Big Data Analytics - FIT - HCMUS 19
John Snow’s comparison (cont.)
» There is no difference whatever in
⋄ The houses or the people receiving the supply of
the two water companies
⋄ Any of the physical conditions with which they are
surrounded
» One group being supplied with water containing
the sewage of London, and amongst it, whatever
might have come from the cholera patients, the
other group having water quite free from
impurity…
» The only difference was in the water supply
Intro2 Big Data Analytics - FIT - HCMUS 20
In brief
» Dr. John Snow established a method for the field of
epidemiology
⋄ Whether the treatment has an effect on the outcome?
» Randomized controlled trial (RCT)
⋄ Step 1 - Choose a group of volunteers (study objects)
⋄ Step 2 - Apply statistical methods to divide objects into
2 groups randomly
− Treatment group receives the treatment
− Control group does not receive the the treatment
⋄ Step 3 - Keep track objects for a while to collect data
(outcome)
⋄ Step 4 – Compare outcomes between the 2 groups to
know whether the treatment is effective or not
Intro2 Big Data Analytics - FIT - HCMUS 21
3.
Data mining
Knowledge Discovery from Data (KDD)
What is data mining?
» Given lots of data
» Discover patterns and models
⋄ Valid
− Hold on new data with some certainty
⋄ Useful
− Should be possible to act on the item
⋄ Unexpected
− Non-obvious to the system
⋄ Understandable
− Humans should be able to interpret the pattern
Intro2 Big Data Analytics - FIT - HCMUS 23
Pattern and Model
» Pattern
⋄ A low level summary of a relationship which
holds only for a few records/variables in data
» Model
⋄ A mathematical representation to recognize
patterns
Intro2 Big Data Analytics - FIT - HCMUS 24
Pattern and Model (cont.)
[Link]/2014/pattern-recognition-systems-convey-learning-1205
Intro2 Big Data Analytics - FIT - HCMUS 25
Pattern and Model (cont.)
[Link]/reactive-lda-library/
Intro2 Big Data Analytics - FIT - HCMUS 26
Data mining & Machine learning
DM ML
» To find new and » To build computer
useful knowledge systems that learn
from large datasets as well as human
does
» Knowledge » Effort of building
discovery from artificial
databases intelligence
Intro2 Big Data Analytics - FIT - HCMUS 27
Data mining
» Database
⋄ Large-scale data, simple queries
⋄ Data mining is a form of analytic processing
− Result is the query answer
» Machine learning
⋄ Small data, complex models
⋄ Data mining is inference of models
− Result is the parameters of the model
» Computer science theory
⋄ Algorithms
Intro2 Big Data Analytics - FIT - HCMUS 28
Disciplines
Ar
Pattern
tifi
c
Recognition
ial
s
tic
Int
tis
ellig
Sta
en
ce
DATA Machine
MINING Learning
Mathematical
Modeling Databases
Management Science &
Information Systems
Sharrda, Business Intelligence, 3rd
Intro2 Big Data Analytics - FIT - HCMUS 29
Data mining process
Jiawei Han, 2012
Intro2 Big Data Analytics - FIT - HCMUS 30
Data mining tasks
» Descriptive methods
⋄ Find human-interpretable patterns that describe
the data
⋄ E.g. clustering, association analysis,
summarization…
» Predictive methods
⋄ Use some variables to predict unknown or future
values of other variables
⋄ E.g. decision tree, neural network, support vector
machine, hidden Markov model…
Intro2 Big Data Analytics - FIT - HCMUS 31
Data mining tasks (cont.)
Data Mining Learning Method Popular Algorithms
Classification and Regression Trees,
Prediction Supervised
ANN, SVM, Genetic Algorithms
Decision trees, ANN/MLP, SVM, Rough
Classification Supervised
sets, Genetic Algorithms
Linear/Nonlinear Regression, Regression
Regression Supervised
trees, ANN/MLP, SVM
Association Unsupervised Apriory, OneR, ZeroR, Eclat
Link analysis Unsupervised Expectation Maximization, Apriory
Algorithm, Graph-based Matching
Sequence analysis Unsupervised Apriory Algorithm, FP-Growth technique
Clustering Unsupervised K-means, ANN/SOM
Outlier analysis Unsupervised K-means, Expectation Maximization (EM)
Intro2 Big Data Analytics - FIT - HCMUS Sharrda, Business Intelligence, 3rd 32
Meaningfulness of Answers
» Risk : you will “discover” patterns that are
meaningless
» Statisticians call it Bonferroni’s principle
⋄ If you look in more places for interesting patterns
than your amount of data will support, you are
bound to find crap
» When looking for a property, make sure that the
property does not allow so many possibilities that
random data will surely produce facts “of
interest.”
Intro2 Big Data Analytics - FIT - HCMUS 33
4.
Big data
It's not big, it's just bigger
IDC’s report
(International Data Corporation)
» In 2011, the overall created and copied data
volume in the world is 1.8 Zettabytes (1021B)
» Increased by nearly 9 times within 5 years
» Will double at least every other 2 years in the
future
» Expect 175 Zettabytes of data worldwide by
2025
Intro2 Big Data Analytics - FIT - HCMUS 35
IBM’s report
» 2.5 Exabytes (1018B) of data was generated
every day in 2012
» About 75% of data is unstructured, coming
from sources such as text, voice and video
Intro2 Big Data Analytics - FIT - HCMUS 36
[Min Chen, 2014]
» Services of Internet
⋄ Google processes data of hundreds of
Petabyte (1015B)
⋄ Facebook generates log data of over 10PB per
month
⋄ 72 hours of video are uploaded to YouTube in
every minute
⋄ Baidu processes data of tens of PB
⋄ Taobao, Alibaba generates data of tens of
Terabyte (1012B) for online trading per day
Intro2 Big Data Analytics - FIT - HCMUS 37
What is big data?
» 2001, Doug Laney, Gartner’s analyst
» The 3V model
⋄ Volume
− The quantity of data
− The size of the data
⋄ Velocity
− The speed of generation of data
− How fast the data is generated and processed to meet
the demands
⋄ Variety
− The category which data belongs to
Intro2 Big Data Analytics - FIT - HCMUS 38
What is big data? (cont.)
» 2010, Apache Hadoop
⋄ Datasets which could not be captured, managed, and
processed by general computers within an acceptable
scope
» 2011, McKinsey & Company
⋄ Datasets which could not be acquired, stored and
managed by classic database software
» Criterions for big data
⋄ Volume of a dataset
⋄ Growing data scale
⋄ Big data management
− Not be handled by traditional database technologies
Intro2 Big Data Analytics - FIT - HCMUS 39
What is big data? (cont.)
» 2011, IDC
⋄ Big data technologies describe a new generation
of technologies and architectures, designed to
economically extract value from very large
volumes of a wide variety of data, by enabling the
high-velocity capture, discovery, and/or analysis
» Definition has 4Vs
» “You could only own a bunch of data other than
big data if you do not utilize the collected data” -
Jay Parikh, Facebook
Intro2 Big Data Analytics - FIT - HCMUS 40
What is big data? (cont.)
» Currently, IBM, Gartner, Cisco…
Data in many
Data at Scale Forms Data Uncertainty
Data in Motion
Terabytes to Structured, Precise data &
Streaming data
Petabytes unstructured, text, correct analyses
multimedia
Intro2 Big Data Analytics - FIT - HCMUS 41
In brief
» Datasets that are typically beyond the ability of a single machine
or a modern data management system to store and analyze
» Brings new opportunities for discovering new values
» Helps us to gain an in-depth understanding of the hidden values
» Incurs new challenges
⋄ How to effectively organize and manage
» In future
⋄ Data will exceed the strength of IT architecture and infrastructure
⋄ Requirements of real-time data processing
⋄ Putting pressure on the existing systems
Intro2 Big Data Analytics - FIT - HCMUS 42
In brief (cont.)
» New technologies, methods, skills
⋄ Data acquisition
− Oracle NoSQL database
− Amazon Simple Storage Service (S3)
− Google BigTable
⋄ Data organization
− Hadoop/MapReduce
− Spark
⋄ Data analysis
⋄ Data decision making
Intro2 Big Data Analytics - FIT - HCMUS 43
5.
Big data analysis
What is big data analysis?
» The use of advanced analytic techniques
against very large & diverse datasets to
uncover
⋄ Hidden patterns, unknown correlations,
⋄ Trends, preferences,
⋄ Other useful information
» Purpose
⋄ To make better and faster decisions
Intro2 Big Data Analytics - FIT - HCMUS 45
Advanced techniques
» Text analytics
» Machine learning
» Predictive analytics
» Data mining
» Statistics
» Natural language processing
Intro2 Big Data Analytics - FIT - HCMUS 46
Solution
» Push computation to the data, instead of
pushing data to a computing node
⋄ The MapReduce programming paradigm
» Some examples
⋄ Apache, HortonWorks, Cloudera, and Teradata
Intro2 Big Data Analytics - FIT - HCMUS 47
6.
Why study this course
The fact
» We are living in the information era
» Information/data is very important
» We should know
⋄ How to own the data
⋄ How to analyze data
Intro2 Big Data Analytics - FIT - HCMUS 49
In my opinion
» For those who will work
⋄ I am not sure all companies in VN where have
applied data mining
− But a few do
⋄ You will not use data mining right after leaving
school
− But maybe in the future you will
» For those who continue studying
⋄ Have strong background in CS field
⋄ Create something new that was used by people
Intro2 Big Data Analytics - FIT - HCMUS 50
Some companies
» trustingsocial ([Link])
» East Agile ([Link])
» Sentifi ([Link])
» Inspectrorio ([Link])
» VNG ([Link])
» Zalora ([Link])
» Momo ([Link])
» …
Intro2 Big Data Analytics - FIT - HCMUS 51
THANKS!
Any questions?
52