0% menganggap dokumen ini bermanfaat (0 suara)
16 tayangan117 halaman

Data Mining: Konsep dan Penerapan

Hak Cipta
© All Rights Reserved
Kami menangani hak cipta konten dengan serius. Jika Anda merasa konten ini milik Anda, ajukan klaim di sini.
Format Tersedia
Unduh sebagai PDF, TXT atau baca online di Scribd
0% menganggap dokumen ini bermanfaat (0 suara)
16 tayangan117 halaman

Data Mining: Konsep dan Penerapan

Hak Cipta
© All Rights Reserved
Kami menangani hak cipta konten dengan serius. Jika Anda merasa konten ini milik Anda, ajukan klaim di sini.
Format Tersedia
Unduh sebagai PDF, TXT atau baca online di Scribd

[Link]

be/h82NuHDNhKI
Romi Satria Wahono
romi@[Link]
[Link]
08118228331

1
Romi Satria Wahono
• SMA Taruna Nusantara Magelang (1993)
• [Link], [Link] and Ph.D in Software Engineering
Saitama University Japan (1994-2004)
Universiti Teknikal Malaysia Melaka (2014)
• Research Interests in Software Engineering and
Machine Learning
• LIPI Researcher (2004-2007)
• Founder and CEO:
• PT Brainmatics Cipta Informatika (2005)
• PT IlmuKomputerCom Braindevs Sistema (2014)
• Professional Member of IEEE, ACM and PMI
• IT and Research Award Winners from WSIS (United Nations), Kemdikbud,
Ristekdikti, LIPI, etc
• SCOPUS/ISI Indexed Journal Reviewer: Information and Software Technology,
Journal of Systems and Software, Software: Practice and Experience, etc
• Industrial IT Certifications: TOGAF, ITIL, CCNA, etc
• Enterprise Architecture Consultant: KPK, RistekDikti, INSW, LIPI, Kemenkeu
(Itjend, DJBC, DJPK), Kemsos, Telkom, PLN PJB, Pertamina EP, FIF, etc
2
Learning Design

Criterion Referenced
Educational Objectives Minimalism
Instruction
(Benjamin Bloom) (John Carroll)
(Robert Mager)

Start Immediately
Cognitive Competencies

Minimize the Reading


Affective Performance
Error Recognition

Psychomotor Evaluation
Self-Contained

3
Textbooks

4
Pre-Test
1. Jelaskan perbedaan antara data, informasi dan pengetahuan!
2. Jelaskan apa yang anda ketahui tentang data mining!
3. Sebutkan peran utama data mining!
4. Sebutkan pemanfaatan dari data mining di berbagai bidang!
5. Pengetahuan atau pola apa yang bisa kita dapatkan dari data
di bawah?
NIM Gender Nilai Asal IPS1 IPS2 IPS3 IPS 4 ... Lulus Tepat
UN Sekolah Waktu
10001 L 28 SMAN 2 3.3 3.6 2.89 2.9 Ya
10002 P 27 SMAN 7 4.0 3.2 3.8 3.7 Tidak
10003 P 24 SMAN 1 2.7 3.4 4.0 3.5 Tidak
10004 L 26.4 SMAN 3 3.2 2.7 3.6 3.4 Ya
...
11000 L 23.4 SMAN 5 3.3 2.8 3.1 3.2 Ya

5
1.1. Pengantar • 1.1 Apa dan Mengapa Data Mining?
• 1.2 Peran Utama dan Metode Data Mining
Data Mining • 1.3 Sejarah dan Penerapan Data Mining
Course Outline
• 2.1 Proses dan Tools Data Mining
2. Proses • 2.2 Penerapan Proses Data Mining
Data Mining •

2.3 Evaluasi Model Data Mining
2.4 Proses Data Mining berbasis CRISP-DM

• 3.1 Data Cleaning


• 3.2 Data Reduction
3. Persiapan Data • 3.3 Data Transformation and Data Discretization
• 3.4 Data Integration

• 4.1 Algoritma Klasifikasi


4. Algoritma • 4.2 Algoritma Klastering
Data Mining •

4.3 Algoritma Asosiasi
4.4 Algoritma Estimasi dan Forecasting

• 5.1 Text Mining Concepts


5. Text Mining • 5.2 Text Clustering
• 5.3 Text Classification

6
1. Pengantar Data Mining
1.1 Apa dan Mengapa Data Mining?
1.2 Peran Utama dan Metode Data Mining
1.3 Sejarah dan Penerapan Data Mining

7
1.1 Apa dan Mengapa Data Mining?

8
Manusia Memproduksi Data
Manusia memproduksi beragam
data yang jumlah dan ukurannya
sangat besar
• Astronomi
• Bisnis
• Kedokteran
• Ekonomi
• Olahraga
• Cuaca
• Financial
• …

9
10
Pertumbuhan Data kilobyte (kB) 103
megabyte (MB) 106
Astronomi gigabyte (GB) 109
• Sloan Digital Sky Survey terabyte (TB) 1012
• New Mexico, 2000 petabyte (PB) 1015
• 140TB over 10 years exabyte (EB) 1018
zettabyte (ZB) 1021
• Large Synoptic Survey Telescope yottabyte (YB) 1024
• Chile, 2016
• Will acquire 140TB every five days

Biologi dan Kedokteran


• European Bioinformatics Institute (EBI)
• 20PB of data (genomic data doubles in size each year)
• A single sequenced human genome can be around 140GB in size

11
Perubahan Kultur dan Perilaku

12
Perubahan Kultur dan Perilaku

13
Datangnya Tsunami Data kilobyte (kB)
megabyte (MB)
103
106
gigabyte (GB) 109

• Mobile Electronics market terabyte (TB)


petabyte (PB)
1012
1015
• 4.43B mobile phone users in 2015 exabyte (EB) 1018
• 7B mobile phone subscriptions in 2015 zettabyte (ZB) 1021
yottabyte (YB) 1024

• Web and Social Networks generates


amount of data
• Google processes 100 PB per day, 3 million servers
• Facebook has 300 PB of user data per day
• Youtube has 1000PB video storage
• 235 TBs data collected by the US Library of Congress
• 15 out of 17 sectors in the US have more data stored
per company than the US Library of Congress
14
Kebanjiran Data tapi Miskin Pengetahuan

We are drowning in data, but


starving for knowledge!
(John Naisbitt, Megatrends, 1988)

15
Mengubah Data Menjadi Pengetahuan
• Data harus kita olah menjadi pengetahuan
supaya bisa bermanfaat bagi manusia

• Dengan pengetahuan
tersebut, manusia dapat:
• Melakukan estimasi dan prediksi
apa yang terjadi di depan
• Melakukan analisis tentang
asosiasi, korelasi dan
pengelompokan antar data dan atribut
• Membantu pengambilan keputusan dan
pembuatan kebijakan
16
Apa itu Data Mining?

17
Apa itu Data Mining?
• Disiplin ilmu yang mempelajari metode untuk
mengekstrak pengetahuan atau menemukan pola dari
suatu data yang besar

• Ekstraksi dari data ke pengetahuan:


1. Data: fakta yang terekam dan tidak membawa arti
2. Pengetahuan: pola, rumus, aturan atau model yang muncul
dari data

• Nama lain data mining:


• Knowledge Discovery in Database (KDD)
• Knowledge extraction
• Pattern analysis
• Information harvesting
• Business intelligence
• Big data
18
Apa itu Data Mining?

Himpunan Metode Data


Pengetahuan
Data Mining

19
Contoh Data di Kampus
• Puluhan ribu data mahasiswa di kampus yang
diambil dari sistem informasi akademik
• Apakah pernah kita ubah menjadi pengetahuan
yang lebih bermanfaat? TIDAK!
• Seperti apa pengetahuan itu? Rumus, Pola, Aturan

20
Prediksi Kelulusan Mahasiswa

21
Contoh Data di Komisi Pemilihan Umum
• Puluhan ribu data calon anggota legislatif di KPU
• Apakah pernah kita ubah menjadi pengetahuan
yang lebih bermanfaat? TIDAK!

22
Prediksi Calon Legislatif DKI Jakarta

23
From Stupid Apps to Smart Apps

Stupid Smart
Applications Applications
• Sistem Informasi • Sistem Prediksi
Akademik Kelulusan Mahasiswa
• Sistem Pencatatan • Sistem Prediksi Hasil
Pemilu Pemilu
• Sistem Laporan • Sistem Prediksi
Kekayaan Pejabat Koruptor
• Sistem Pencatatan • Sistem Penentu
Kredit Kelayakan Kredit

24
Revolusi Industri 4.0

25
Perusahaan Pengolah Pengetahuan
• Uber - the world’s largest taxi company,
owns no vehicles
• Google - world’s largest
media/advertising company, creates no
content
• Alibaba - the most valuable retailer, has
no inventory
• Airbnb - the world’s largest
accommodation provider, owns no real
estate
• Gojek - perusahaan angkutan umum,
tanpa memiliki kendaraan
26
Definisi Data Mining
• Melakukan ekstraksi untuk mendapatkan informasi
penting yang sifatnya implisit dan sebelumnya tidak
diketahui, dari suatu data (Witten et al., 2011)

• Kegiatan yang meliputi pengumpulan, pemakaian


data historis untuk menemukan keteraturan, pola
dan hubungan dalam set data berukuran besar
(Santosa, 2007)

• Extraction of interesting (non-trivial, implicit,


previously unknown and potentially useful)
patterns or knowledge from huge amount of data
(Han et al., 2011)

27
Data - Informasi – Pengetahuan

NIP TGL DATANG PULANG


1103 02/12/2004 07:20 15:40
1142 02/12/2004 07:45 15:33
1156 02/12/2004 07:51 16:00
1173 02/12/2004 08:00 15:15
1180 02/12/2004 07:01 16:31
1183 02/12/2004 07:49 17:00

Data Kehadiran Pegawai


28
Data - Informasi – Pengetahuan

NIP Masuk Alpa Cuti Sakit Telat

1103 22

1142 18 2 2

1156 10 1 11

1173 12 5 5

1180 10 12

Informasi Akumulasi Bulanan Kehadiran Pegawai


29
Data - Informasi – Pengetahuan

Senin Selasa Rabu Kamis Jumat

Terlambat 7 0 1 0 5

Pulang 0 1 1 1 8
Cepat
Izin 3 0 0 1 4

Alpa 1 0 2 0 2

Pola Kebiasaan Kehadiran Mingguan Pegawai


30
Data - Informasi – Pengetahuan - Kebijakan

• Kebijakan penataan jam kerja karyawan khusus


untuk hari senin dan jumat

• Peraturan jam kerja:


• Hari Senin dimulai jam 10:00
• Hari Jumat diakhiri jam 14:00
• Sisa jam kerja dikompensasi ke hari lain

31
Data Mining Tasks and Roles

Increasing potential
to support business
End User
decisions Decision
Making

Data Presentation Business Analyst


Visualization Techniques
Data Mining Data Analyst
Information Discovery

Data Exploration
Statistical Summary, Querying, and Reporting

Data Preprocessing/Integration, Data Warehouses


DBA
Data Sources
Paper, Files, Web documents, Scientific experiments, Database Systems

32
Hubungan Data Mining dan Bidang Lain

Computing
Statistics
Algorithms

Machine Database
Learning Technology

Pattern Data High


Performance
Recognition
Mining Computing

33
Data Mining vs Text Mining
1. Text Mining:
• Mengolah data tidak terstruktur dalam bentuk text,
web, social media, dsb
• Menggunakan metode text processing untuk
mengkonversi data tidak terstruktur menjadi terstruktur
• Kemudian diolah dengan data mining

2. Data Mining:
• Mengolah data terstruktur dalam bentuk tabel yang
memiliki atribut dan kelas
• Menggunakan metode data mining, yang terbagi
menjadi metode estimasi, forecasting, klasifikasi,
klastering atau asosiasi
• Yang dasar berpikirnya menggunakan konsep statistika atau
heuristik ala machine learning

34
Text Mining

Text Processing

35
3
6

Text Mining
Jejak Pornografi di
Indonesia
Text Mining: AHY-AHOK-ANIES

37
Masalah-Masalah di Data Mining
• Tremendous amount of data
• Algorithms must be highly scalable to handle such as tera-
bytes of data
• High-dimensionality of data
• Micro-array may have tens of thousands of dimensions
• High complexity of data
• Data streams and sensor data
• Time-series data, temporal data, sequence data
• Structure data, graphs, social networks and multi-linked data
• Heterogeneous databases and legacy databases
• Spatial, spatiotemporal, multimedia, text and Web data
• Software programs, scientific simulations
• New and sophisticated applications
38
Latihan
1. Jelaskan dengan kalimat sendiri apa
yang dimaksud dengan data mining?

2. Sebutkan alur proses data mining!

39
1.2 Peran Utama dan Metode Data
Mining

40
Peran Utama Data Mining

1. Estimasi

5. Asosiasi 2. Forecasting

Data Mining Roles


(Larose, 2005)

4. Klastering 3. Klasifikasi

41
Dataset (Himpunan Data)
Attribute/Feature/Dimension
Class/Label/Target

Record/
Object/
Sample/
Tuple/
Data

Nominal
Numerik
42
Tipe Data

(Kontinyu)

(Diskrit)

43
Tipe Data Deskripsi Contoh Operasi
Tipe Data
Ratio • Data yang diperoleh dengan cara • Umur geometric mean,
(Mutlak) pengukuran, dimana jarak dua titik • Berat badan harmonic mean,
pada skala sudah diketahui • Tinggi badan percent variation
• Mempunyai titik nol yang absolut • Jumlah uang
(*, /)
Interval • Data yang diperoleh dengan cara • Suhu 0°c-100°c, mean, standard
(Jarak) pengukuran, dimana jarak dua titik • Umur 20-30 tahun deviation,
pada skala sudah diketahui Pearson's
• Tidak mempunyai titik nol yang correlation, t and
absolut F tests
(+, - )
Ordinal • Data yang diperoleh dengan cara • Tingkat kepuasan median,
(Peringkat) kategorisasi atau klasifikasi pelanggan (puas, percentiles, rank
• Tetapi diantara data tersebut sedang, tidak puas) correlation, run
terdapat hubungan atau berurutan tests, sign tests
(<, >)
Nominal • Data yang diperoleh dengan cara • Kode pos mode, entropy,
(Label) kategorisasi atau klasifikasi • Jenis kelamin contingency
• Menunjukkan beberapa object • Nomer id karyawan correlation, 2
yang berbeda • Nama kota test
(=, ) 44
Peran Utama Data Mining

1. Estimasi

5. Asosiasi 2. Forecasting

Data Mining Roles


(Larose, 2005)

4. Klastering 3. Klasifikasi

45
1. Estimasi Waktu Pengiriman Pizza
Label
Customer Jumlah Pesanan (P) Jumlah Traffic Light (TL) Jarak (J) Waktu Tempuh (T)

1 3 3 3 16
2 1 7 4 20
3 2 4 6 18
4 4 6 8 36
...
1000 2 4 2 12

Pembelajaran dengan
Metode Estimasi (Regresi Linier)

Waktu Tempuh (T) = 0.48P + 0.23TL + 0.5J


Pengetahuan
46
Contoh: Estimasi Performansi CPU
• Example: 209 different computer configurations

Cycle time Main memory Cache Channels Performance


(ns) (Kb) (Kb)
MYCT MMIN MMAX CACH CHMIN CHMAX PRP
1 125 256 6000 256 16 128 198
2 29 8000 32000 32 8 32 269

208 480 512 8000 32 0 0 67
209 480 1000 4000 0 0 0 45

• Linear regression function


PRP = -55.9 + 0.0489 MYCT + 0.0153 MMIN + 0.0056
MMAX
+ 0.6410 CACH - 0.2700 CHMIN + 1.480 CHMAX

47
Output/Pola/Model/Knowledge
1. Formula/Function (Rumus atau Fungsi Regresi)
• WAKTU TEMPUH = 0.48 + 0.6 JARAK + 0.34 LAMPU + 0.2 PESANAN

2. Decision Tree (Pohon Keputusan)

3. Korelasi dan Asosiasi

4. Rule (Aturan)
• IF ips3=2.8 THEN lulustepatwaktu

5. Cluster (Klaster)

48
2. Forecasting Harga Saham
Label Time Series

Dataset harga saham


dalam bentuk time
series (rentet waktu)

Pembelajaran dengan
Metode Forecasting (Neural Network)

49
Pengetahuan berupa
Rumus Neural Network

Prediction Plot

50
Forecasting Cuaca

51
Exchange Rate Forecasting

52
Inflation Rate Forecasting

53
3. Klasifikasi Kelulusan Mahasiswa
Label

NIM Gender Nilai Asal IPS1 IPS2 IPS3 IPS 4 ... Lulus Tepat
UN Sekolah Waktu
10001 L 28 SMAN 2 3.3 3.6 2.89 2.9 Ya
10002 P 27 SMA DK 4.0 3.2 3.8 3.7 Tidak
10003 P 24 SMAN 1 2.7 3.4 4.0 3.5 Tidak
10004 L 26.4 SMAN 3 3.2 2.7 3.6 3.4 Ya
...
...
11000 L 23.4 SMAN 5 3.3 2.8 3.1 3.2 Ya

Pembelajaran dengan
Metode Klasifikasi (C4.5)

54
Pengetahuan Berupa Pohon Keputusan

55
Contoh: Rekomendasi Main Golf
• Input:

• Output (Rules):
If outlook = sunny and humidity = high then play = no
If outlook = rainy and windy = true then play = no
If outlook = overcast then play = yes
If humidity = normal then play = yes
If none of the above then play = yes
56
Contoh: Rekomendasi Main Golf
• Output (Tree):

57
Contoh: Rekomendasi Contact Lens
• Input:

58
Contoh: Rekomendasi Contact Lens
• Output/Model (Tree):

59
Klasifikasi Sentimen Analisis

60
Bankruptcy Prediction

61
4. Klastering Bunga Iris
Dataset Tanpa Label

Pembelajaran dengan
Metode Klastering (K-Means)

62
Pengetahuan (Model) Berupa Klaster

63
Klastering Jenis Pelanggan

64
Klastering Sentimen Warga

65
Poverty Rate Clustering

66
5. Aturan Asosiasi Pembelian Barang

Pembelajaran dengan
Metode Asosiasi (FP-Growth)

67
Pengetahuan Berupa Aturan Asosiasi

68
Contoh Aturan Asosiasi
• Algoritma association rule (aturan asosiasi) adalah
algoritma yang menemukan atribut yang “muncul
bersamaan”
• Contoh, pada hari kamis malam, 1000 pelanggan
telah melakukan belanja di supermaket ABC, dimana:
• 200 orang membeli Sabun Mandi
• dari 200 orang yang membeli sabun mandi, 50 orangnya
membeli Fanta
• Jadi, association rule menjadi, “Jika membeli sabun
mandi, maka membeli Fanta”, dengan nilai support =
200/1000 = 20% dan nilai confidence = 50/200 = 25%
• Algoritma association rule diantaranya adalah: A
priori algorithm, FP-Growth algorithm, GRI algorithm
69
Aturan Asosiasi di [Link]

70
Heating Oil Consumption
Korelasi antara jumlah konsumsi minyak
pemanas dengan faktor-faktor di bawah:

1. Insulation: Ketebalan insulasi rumah


2. Temperatur: Suhu udara sekitar rumah
3. Heating Oil: Jumlah konsumsi minyak
pertahun perrumah
4. Number of Occupant: Jumlah penghuni rumah
5. Average Age: Rata-rata umur penghuni rumah
6. Home Size: Ukuran rumah

71
72
Tingkat Korelasi Faktor-Faktor terhadap
Konsumsi Minyak
Jumlah
Penghuni
Rumah
Rata-Rata
Umur
0.381
0.848
Konsumsi
Ketebalan Minyak
Insulasi 0.736
Rumah

-0.774
Temperatur

73
Insight Law (Data Mining Law 6)
Data mining amplifies perception in the
business domain
• How does data mining produce insight? This law
approaches the heart of data mining – why it must be a
business process and not a technical one
• Business problems are solved by people, not by algorithms
• The data miner and the business expert “see” the
solution to a problem, that is the patterns in the domain
that allow the business objective to be achieved
• Thus data mining is, or assists as part of, a perceptual process
• Data mining algorithms reveal patterns that are not normally
visible to human perception
• Within the data mining process, the human problem
solver interprets the results of data mining algorithms
and integrates them into their business understanding
74
Metode Learning Algoritma Data Mining

Supervised Semi- Unsupervised


Supervised
Learning Learning Learning

Association based Learning

75
1. Supervised Learning

• Pembelajaran dengan guru, data set memiliki


target/label/class
• Sebagian besar algoritma data mining
(estimation, prediction/forecasting,
classification) adalah supervised learning
• Algoritma melakukan proses belajar
berdasarkan nilai dari variabel target yang
terasosiasi dengan nilai dari variable prediktor

76
Dataset dengan Class
Attribute/Feature/Dimension Class/Label/Target

Nominal

Numerik
77
2. Unsupervised Learning

• Algoritma data mining mencari pola dari


semua variable (atribut)
• Variable (atribut) yang menjadi
target/label/class tidak ditentukan (tidak ada)
• Algoritma clustering adalah algoritma
unsupervised learning

78
Dataset tanpa Class
Attribute/Feature/Dimension

79
3. Semi-Supervised Learning
• Semi-supervised learning
adalah metode data mining
yang menggunakan data
dengan label dan tidak
berlabel sekaligus dalam
proses pembelajarannya

• Data yang memiliki kelas


digunakan untuk membentuk
model (pengetahuan), data
tanpa label digunakan untuk
membuat batasan antara
kelas

80
Algoritma Data Mining

1. Estimation (Estimasi):
• Linear Regression, Neural Network, Support Vector Machine, etc
2. Prediction/Forecasting (Prediksi/Peramalan):
• Linear Regression, Neural Network, Support Vector Machine, etc
3. Classification (Klasifikasi):
• Naive Bayes, K-Nearest Neighbor, C4.5, ID3, CART, Linear Discriminant
Analysis, Logistic Regression, etc
4. Clustering (Klastering):
• K-Means, K-Medoids, Self-Organizing Map (SOM), Fuzzy C-Means, etc
5. Association (Asosiasi):
• FP-Growth, A Priori, Coefficient of Correlation, Chi Square, etc

81
Output/Pola/Model/Knowledge

1. Formula/Function (Rumus atau Fungsi Regresi)


• WAKTU TEMPUH = 0.48 + 0.6 JARAK + 0.34 LAMPU + 0.2 PESANAN

2. Decision Tree (Pohon Keputusan)

3. Tingkat Korelasi

4. Rule (Aturan)
• IF ips3=2.8 THEN lulustepatwaktu

5. Cluster (Klaster)

82
Latihan
1. Sebutkan 5 peran utama data mining!
2. Jelaskan perbedaan estimasi dan forecasting!
3. Jelaskan perbedaan forecasting dan klasifikasi!
4. Jelaskan perbedaan klasifikasi dan klastering!
5. Jelaskan perbedaan klastering dan association!
6. Jelaskan perbedaan estimasi dan klasifikasi!
7. Jelaskan perbedaan estimasi dan klastering!
8. Jelaskan perbedaan supervised dan unsupervised
learning!
9. Sebutkan tahapan utama proses data mining!

83
1.3 Sejarah dan Penerapan Data
Mining

84
Evolution of Sciences
• Sebelum 1600: Empirical science
• Disebut sains kalau bentuknya kasat mata

• 1600-1950: Theoretical science


• Disebut sains kalau bisa dibuktikan secara matematis atau eksperimen

• 1950s-1990: Computational science


• Seluruh disiplin ilmu bergerak ke komputasi
• Lahirnya banyak model komputasi

• 1990-sekarang: Data science


• Kultur manusia menghasilkan data besar
• Kemampuan komputer untuk mengolah data besar
• Datangnya data mining sebagai arus utama sains

Jim Gray and Alex Szalay, The World Wide Telescope:


An Archetype for Online Science, Comm. ACM, 45(11): 50-54, Nov. 2002
85
86
Revolusi Industri 4.0

87
1. Big Data

Digital
4. Enterprise 2. Internet of
Architecture Transformation Things
Trends

3. Business
Process
Automation
88
89
Data Mining Use Cases
across Industries

90
Business

Knowledge

Methods

Technology
91
Business Goals Law (Data Mining Law 1)
Business objectives are the origin of every data
mining solution

• This defines the field of data mining: data mining is


concerned with solving business problems and
achieving business goals
• Data mining is not primarily a technology; it is a
process, which has one or more business objectives
at its heart
• Without a business objective, there is no data
mining
• The maxim: “Data Mining is a Business Process”

92
Business Knowledge Law (Data Mining Law 2)
Business knowledge is central to every step of the
data mining process

• A naive reading of CRISP-DM would see business


knowledge used at the start of the process in
defining goals, and at the end of the process in
guiding deployment of results
• This would be to miss a key property of the data
mining process, that business knowledge has a
central role in every step

93
Private and Commercial Sector
• Marketing: product recommendation, market basket
analysis, product targeting, customer retention
• Finance: investment support, portfolio management,
price forecasting
• Banking and Insurance: credit and policy approval,
money laundry detection
• Security: fraud detection, access control, intrusion
detection, virus detection
• Manufacturing: process modeling, quality control,
resource allocation
• Web and Internet: smart search engines, web
marketing
• Software Engineering: effort estimation, fault
prediction
• Telecommunication: network monitoring, customer
churn prediction, user behavior analysis
94
Use Case: Product Recommendation
4.000.000
[Link]
3.500.000
[Link]
3.000.000 [Link]
2.500.000

2.000.000

1.500.000

1.000.000

500.000

0
0 5 10 15 20 25 30 35

Cluster - 2 Cluster - 3 Cluster - 1

95
Use Case: Penentuan Kelayakan Kredit
20

15

10 Jumlah kredit
macet

0
2003 2004

96
Use Case: Software Fault Prediction

• The cost of capturing and


correcting defects is expensive
• $14,102 per defect in post-release
phase (Boehm & Basili 2008)
• $60 billion per year (NIST 2002)
• Industrial methods of manual
software reviews activities can
find only 60% of defects
(Shull et al. 2002)
• The probability of detection of
software fault prediction
models is higher (71%) than
software reviews (60%)

97
Public and Government Sector
• Finance: exchange rate forecasting, sentiment analysis
• Taxation: adaptive monitoring, fraud detection
• Medicine and Healt Care: hypothesis discovery, disease
prediction and classification, medical diagnosis
• Education: student allocation, resource forecasting
• Insurance: worker’s compensation analysis
• Security: bomb, iceberg detection
• Transportation: simulation and analysis, load estimation
• Law: legal patent analysis, law and rule analysis
• Politic: election prediction

98
Use Case: Deteksi Pencucian Uang

99
Use Case: Prediksi Kebakaran Hutan
FFMC DMC DC ISI temp RH wind rain ln(area+1)
93.5 139.4 594.2 20.3 17.6 52 5.8 0 0
92.4 124.1 680.7 8.5 17.2 58 1.3 0 0
90.9 126.5 686.5 7 15.6 66 3.1 0 0
85.8 48.3 313.4 3.9 18 42 2.7 0 0.307485
91 129.5 692.6 7 21.7 38 2.2 0 0.357674
90.9 126.5 686.5 7 21.9 39 1.8 0 0.385262
95.5 99.9 513.3 13.2 23.3 31 4.5 0 0.438255
12
9,648
10

8
5,9 5,615
SVM SVM+GA 6
C 4.3 1,840 4,3

4
Gamma (𝛾) 5.9 9,648 3,9

Epsilon (𝜀) 3.9 5,615 1,840 1,391

2
RMSE 1.391 1.379
0 1,379
C Gamma Epsilon RMSE
SVM SVM+GA
100
Use Case: Prediksi Koruptor

Aktivitas Penindakan Prediksi dan klastering


calon tersangka koruptor
Aktivitas Intelijen

Asosiasi atribut
DATA tersangka koruptor
DATA DATA Pengetahuan

DATA
Prediksi pencucian uang

Aktivitas Pendukung
Estimasi jenis dan
Aktivitas Pencegahan jumlah tahun hukuman

101
Contoh Penerapan Data Mining
• Penentuan kelayakan kredit pemilihan rumah di bank
• Penentuan pasokan listrik PLN untuk wilayah Jakarta
• Prediksi profile tersangka koruptor dari data pengadilan
• Perkiraan harga saham dan tingkat inflasi
• Analisis pola belanja pelanggan
• Memisahkan minyak mentah dan gas alam
• Penentuan pola pelanggan yang loyal pada perusahaan
operator telepon
• Deteksi pencucian uang dari transaksi perbankan
• Deteksi serangan (intrusion) pada suatu jaringan

102
Data Mining Society
• 1989 IJCAI Workshop on Knowledge Discovery in Databases
• Knowledge Discovery in Databases (G. Piatetsky-Shapiro and W. Frawley, 1991)

• 1991-1994 Workshops on Knowledge Discovery in Databases


• Advances in Knowledge Discovery and Data Mining (U. Fayyad, G.
Piatetsky-Shapiro, P. Smyth, and R. Uthurusamy, 1996)

• 1995-1998 International Conferences on Knowledge Discovery


in Databases and Data Mining (KDD’95-98)
• Journal of Data Mining and Knowledge Discovery (1997)

• ACM SIGKDD conferences since 1998 and SIGKDD Explorations


• More conferences on data mining
• PAKDD (1997), PKDD (1997), SIAM-Data Mining (2001), (IEEE) ICDM
(2001), WSDM (2008), etc.

• ACM Transactions on KDD (2007)


103
Conferences dan Journals Data Mining

Conferences Journals
• ACM SIGKDD Int. Conf. on • ACM Transactions on Knowledge
Knowledge Discovery in Databases
and Data Mining (KDD) Discovery from Data (TKDD)
• SIAM Data Mining Conf. (SDM) • ACM Transactions on
• (IEEE) Int. Conf. on Data Mining Information Systems (TOIS)
(ICDM)
• IEEE Transactions on Knowledge
• European Conf. on Machine and Data Engineering
Learning and Principles and
practices of Knowledge Discovery • Springer Data Mining and
and Data Mining (ECML-PKDD) Knowledge Discovery
• Pacific-Asia Conf. on Knowledge
Discovery and Data Mining • International Journal of Business
(PAKDD) Intelligence and Data Mining
• Int. Conf. on Web Search and Data (IJBIDM)
Mining (WSDM)
104
Data Mining Law

Tom Khabaza, Nine Laws of Data Mining, 2010


([Link]

105
Data Mining Law
1. Business objectives are the origin of every data mining
solution
2. Business knowledge is central to every step of the data
mining process
3. Data preparation is more than half of every data mining
process
4. There is no free lunch for the data miner
5. There are always patterns
6. Data mining amplifies perception in the business domain
7. Prediction increases information locally by generalisation
8. The value of data mining results is not determined by the
accuracy or stability of predictive models
9. All patterns are subject to change
Tom Khabaza, Nine Laws of Data Mining, 2010
([Link]
106
1 Business Goals Law
Business objectives are the origin of every data
mining solution

• This defines the field of data mining: data mining is


concerned with solving business problems and
achieving business goals
• Data mining is not primarily a technology; it is a
process, which has one or more business objectives
at its heart
• Without a business objective, there is no data
mining
• The maxim: “Data Mining is a Business Process”

107
2 Business Knowledge Law
Business knowledge is central to every step of the
data mining process

• A naive reading of CRISP-DM would see business


knowledge used at the start of the process in
defining goals, and at the end of the process in
guiding deployment of results
• This would be to miss a key property of the data
mining process, that business knowledge has a
central role in every step

108
2 Business Knowledge Law
1. Business understanding must be based on business
knowledge, and so must the mapping of business
objectives to data mining goals
2. Data understanding uses business knowledge to
understand which data is related to the business problem,
and how it is related
3. Data preparation means using business knowledge to
shape the data so that the required business questions
can be asked and answered
4. Modelling means using data mining algorithms to create
predictive models and interpreting both the models and
their behaviour in business terms – that is, understanding
their business relevance
5. Evaluation means understanding the business impact of
using the models
6. Deployment means putting the data mining results to
work in a business process

109
3 Data Preparation Law
Data preparation is more than half of every data
mining process

• Maxim of data mining: most of the effort in a data


mining project is spent in data acquisition and
preparation, and informal estimates vary from 50 to
80 percent
• The purpose of data preparation is:
1. To put the data into a form in which the data mining
question can be asked
2. To make it easier for the analytical techniques (such as
data mining algorithms) to answer it

110
4 No Free Lunch Theory
There is No Free Lunch for the Data Miner (NFL-DM)
The right model for a given application can only be discovered by
experiment

• Axiom of machine learning: if we knew enough about a problem


space, we could choose or design an algorithm to find optimal
solutions in that problem space with maximal efficiency
• Arguments for the superiority of one algorithm over others in data
mining rest on the idea that data mining problem spaces have one
particular set of properties, or that these properties can be
discovered by analysis and built into the algorithm
• However, these views arise from the erroneous idea that, in data
mining, the data miner formulates the problem and the algorithm
finds the solution
• In fact, the data miner both formulates the problem and finds the
solution – the algorithm is merely a tool which the data miner uses
to assist with certain steps in this process

111
4 No Free Lunch Theory
• If the problem space were well-understood, the data mining
process would not be needed
• Data mining is the process of searching for as yet unknown
connections
• For a given application, there is not only one problem space
• Different models may be used to solve different parts of the
problem
• The way in which the problem is decomposed is itself often the
result of data mining and not known before the process begins
• The data miner manipulates, or “shapes”, the problem space
by data preparation, so that the grounds for evaluating a
model are constantly shifting
• There is no technical measure of value for a predictive
model
• The business objective itself undergoes revision and
development during the data mining process
• so that the appropriate data mining goals may change completely

112
5 Watkins’ Law
There are always patterns

• This law was first stated by David Watkins


• There is always something interesting to be found in
a business-relevant dataset, so that even if the
expected patterns were not found, something else
useful would be found
• A data mining project would not be undertaken
unless business experts expected that patterns
would be present, and it should not be surprising
that the experts are usually right

113
6 Insight Law
Data mining amplifies perception in the
business domain
• How does data mining produce insight? This law approaches
the heart of data mining – why it must be a business process
and not a technical one
• Business problems are solved by people, not by algorithms
• The data miner and the business expert “see” the solution to a
problem, that is the patterns in the domain that allow the
business objective to be achieved
• Thus data mining is, or assists as part of, a perceptual process
• Data mining algorithms reveal patterns that are not normally visible to
human perception
• The data mining process integrates these algorithms with the
normal human perceptual process, which is active in nature
• Within the data mining process, the human problem solver
interprets the results of data mining algorithms and integrates
them into their business understanding
114
7 Prediction Law
Prediction increases information locally by
generalisation
• “Predictive models” and “predictive analytics” means “predict the
most likely outcome”
• Other kinds of data mining models, such as clustering and
association, are also characterised as “predictive”; this is a much
looser sense of the term:
• A clustering model might be described as “predicting” the group into
which an individual falls
• An association model might be described as “predicting” one or more
attributes on the basis of those that are known
• What is “prediction” in this sense? What do classification,
regression, clustering and association algorithms and their resultant
models have in common?
• The answer lies in “scoring”, that is the application of a predictive model to
a new example
• The available information about the example in question has been
increased, locally, on the basis of the patterns found by the algorithm and
embodied in the model, that is on the basis of generalisation or induction
115
8 Value Law
The value of data mining results is not determined by
the accuracy or stability of predictive models

• Accuracy and stability are useful measures of how


well a predictive model makes its predictions
• Accuracy means how often the predictions are correct
• Stability means how much the predictions would change
if the data used to create the model were a different
sample from the same population
• The value of a predictive model arises in two ways:
• The model’s predictions drive improved (more effective)
action
• The model delivers insight (new knowledge) which leads
to improved strategy

116
9 Law of Change
All patterns are subject to change

• The patterns discovered by data mining do not last


forever
• In marketing and CRM applications of data mining, it is
well-understood that patterns of customer behaviour
are subject to change over time
• Fashions change, markets and competition change, and the
economy changes as a whole; for all these reasons, predictive
models become out-of-date and should be refreshed
regularly or when they cease to predict accurately
• The same is true in risk and fraud-related applications of data
mining. Patterns of fraud change with a changing
environment and because criminals change their behaviour in
order to stay ahead of crime prevention efforts

117

Anda mungkin juga menyukai