4 - Chapter 4. Machine Learning Basics
4 - Chapter 4. Machine Learning Basics
Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 289
Chapter Outline
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 290
What is Machine Learning?
“The field of study that gives computers the ability to
learn without being explicitly programmed“
Arthur Samuel
(1901 – 1990)
Founder of ML at
IBM in 1959
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 291
Machine Learning: What & Why?
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 292
Machine Learning: An Overview
Machine Learning
Supervised Unsupervised
Learning Learning
Classification Clustering
Regression Dimensionality
Reduction
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 293
Overview of Supervised Learning
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 294
Overview of Unsupervised Learning
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 295
Supervised vs Unsupervised Learning
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 296
Introduction to scikit-learn
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 297
Supervised Learning Algorithms
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 298
Unsupervised Learning Algorithms
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 299
Reinforcement Learning
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 300
Reinforcement Learning: Key Concepts
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 301
Reinforcement Learning
Principle: Learning through Trial and Error
• Agent: The robot
• Action: Move
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 302
Reinforcement Learning
Agent
Reward
Action
State
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 303
Reinforcement Learning: Flappy Bird Game
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 304
Regression Algorithms
Introduction
Linear Regression
Polynomial Regression
Logistic Regression
Hands-On Practices
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 305
Linear Regression - Example
Linear regression example:
import numpy as np
from sklearn.linear_model import LinearRegression
from [Link] import mean_squared_error
import [Link] as plt
# Sample data
X = [Link]([[1], [2], [3], [4], [5]]) # Input features
y = [Link]([2, 4, 5, 4, 5]) #Target variable
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 306
Linear Regression - Example
Linear regression example (cont’d):
# Calculate the mean squared error
mse = mean_squared_error(y, y_pred)
print("Mean Squared Error:", mse)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 307
Linear Regression - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 308
Linear Regression - Quiz
Quiz: Regression regression
Input:
X is an array of 100 random values between 0 and 1.5.
y is generated by the formula 𝑦 = 2 + 1.5X + noise, where
the noise is added using [Link](100,1)
Task Requirements:
a. Train a Linear Regression model using the provided
dataset.
b. Predict the values of y for X = 0 and X = 2 using the
trained model.
c. Plot the following: the original dataset, the predicted
values from the model
d. Output the model's intercept and coefficient.
e. Calculate the MSE between the true target variable y and
the predicted
11/14/2025 Dr. Maivalues
Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 309
Linear Regression - Solution
Linear regression quiz (solution)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 310
Linear Regression - Solution
Linear regression quiz (solution) – cont’d
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 311
Linear Regression - Solution
Linear regression quiz (solution) – cont’d
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 312
Linear Regression - Solution
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 313
Linear Regression - Solution
Linear regression quiz (solution) – cont’d
# Print the model's intercept and coefficient
print("Intercept:", lin_reg.intercept_)
print("Coefficient:", lin_reg.coef_)
>>
Intercept: [2.51794653]Coefficient: [[1.49144611]]Mean Squared
Error (MSE) for the model: 0.08105176427547395
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 314
Multiple Linear Regression
Example:
Dataset: A synthetic dataset will be generated to predict
car prices based on two features:
Engine size (in liters)
Horsepower
⇒ The target variable will be the car's Price.
The price of a car is calculated using the following
formula:
Price = (Engine Size×5000) + (Horsepower×200) + ϵ
Where, ϵ represents random noise, modeled as a normal
distribution with a mean of 0 and standard deviation of
1000
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 315
Multiple Linear Regression
Instruction:
1. Dataset Creation: A synthetic dataset with the specified
features will be generated.
2. Data Preprocessing: The dataset will be split into training and
testing sets.
3. Model Building: A Multiple Linear Regression model will be
trained using the training data.
4. Model Evaluation: The model's performance will be assessed
using R-squared and Mean Squared Error (MSE) on the test
data.
5. Prediction: The trained model will be used to predict prices
for new data.
6. Results visualization: A plot comparing predicted prices to
actual prices will be created.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 316
Multiple Linear Regression
Multiple Linear Regression example – Task 1: Create the Dataset
import pandas as pd
import numpy as np
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 317
Multiple Linear Regression
Multiple Linear Regression example – Task 1: Create the Dataset
# Create a DataFrame
df = [Link]({
'Engine Size': engine_size,
'Horsepower': horsepower,
'Price': price
})
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 318
Multiple Linear Regression
Multiple Linear Regression example – Task 2: Preprocess the
Data
from sklearn.model_selection import train_test_split
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 319
Multiple Linear Regression
Multiple Linear Regression example – Task 3: Build a Multiple
Linear Regression Model
from sklearn.linear_model import LinearRegression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 320
Multiple Linear Regression
Multiple Linear Regression example – Task 4: Evaluate the Model
from [Link] import mean_squared_error, r2_score
>>
R-squared: 0.97
Mean Squared Error: 992085.17
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 321
Multiple Linear Regression
Multiple Linear Regression example – Task 5. Predict
# Make a prediction
prediction = [Link](new_data)
>>
Predicted price: $57,807.80
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 322
Multiple Linear Regression
Multiple Linear Regression example – Task 6: Plot the Results
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 323
Multiple Linear Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 324
Polynomial Regression
Where:
𝑐 , 𝑐 , …, 𝑐 are the coefficients of the polynomial,
𝑥 is the value of the independent variable (input
feature).
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 325
Polynomial Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 326
Polynomial Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 327
Polynomial Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 328
Quiz x y
-5 -228.75
-4.5 -179.81
Given the following dataset, determine the -4 -141
coefficients of the cubic polynomial: -3.5 -108.31
y = c0 + c1*x + c2*x^2 + c3*x^3 that best fits -3 -81
-2.5 -58.06
the data using Excel. -2 -38
-1.5 -20.75
-1 -6
-0.5 5.81
0 9
0.5 9.81
1 8
1.5 4.75
2 3
2.5 4.06
3 7
3.5 11.31
4 16
4.5 20.81
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 329
Quiz
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 330
Polynomial Regression workflow with
Scikit-learn
Step 1: Import required libraries
Step 2: Generate the dataset
Step 3: Split data into training and testing sets
Step 4: Polynomial feature transformation
Step 5: Train the linear regression model
Step 6: Make predictions and evaluate the model
Step 7: Visualize the results
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 331
Polynomial Regression workflow with
Scikit-learn
Step 1: Import required libraries
Step 2: Generate the dataset
Step 3: Split data into training and testing sets
Step 4: Polynomial feature transformation
Step 5: Train the linear regression model
Step 6: Make predictions and evaluate the model
Step 7: Visualize the results
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 332
Polynomial Regression – Quiz 1
Create a dataset with 100 data points for x ranging from -3
to 3.
Add random noise to the y-values, where the relationship
between x and y is governed by the equation: y = 2*x^3 -
3*x^2 + 5*x -1 + noise
Use the seed value 42 for reproducibility.
Split the dataset into training and testing sets with an 80-
20% split ratio, ensuring the results are reproducible by
using a fixed random state.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 333
Polynomial Regression – Quiz 1
(solution)
Polynomial regression
import numpy as np quiz 1:
import [Link] as plt
from [Link] import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from [Link] import mean_squared_error, r2_score
from sklearn.model_selection import train_test_split
# Step 7: Visualization
[Link](figsize=(10, 6))
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 336
Polynomial Regression – Quiz 1
(solution)
Polynomial regression quiz 1:
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 337
Polynomial Regression – Quiz 1
(solution)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 338
Polynomial Regression
Polynomial regression example:
from sklearn.linear_model import LinearRegression # liner
regression model
from [Link] import PolynomialFeatures #
polynommial features(extended features)
import numpy as np
import [Link] as plt
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 340
Polynomial Regression
Example: Polynomial regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 341
Polynomial Regression
Polynomial regression example (cont’d):
poly_features = PolynomialFeatures(degree=2) # decide the
maximal degree of the polynomial feature
X_poly = poly_features.fit_transform(X) # convert the original
feature to polynomial feature
# check the extened polynomial features of the first data point
print('original feature:', X[0])
print('polynomial features',X_poly[0])
lin_reg = LinearRegression()
lin_reg.fit(X_poly,y)
lin_reg.intercept_, lin_reg.coef_ # check the bais term and
feature weights of the trained model
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 342
Polynomial Regression
Polynomial regression example (cont’d):
X_new = [Link](X,axis = 0) # in order to plot the line of the
model, we need to sort the the value of x-axis
X_new_poly = poly_features.fit_transform(X_new) # compute the
polynomial features
y_predict = lin_reg.predict(X_new_poly) # make predictions
using trained Linear Regression model
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 343
Polynomial Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 344
Polynomial Regression
Quiz: Polynomial regression
Input:
A dataset consisting of n = 100 data points:
X: A 1D array of random values between -3 and 3.
y: A target variable, generated using the formula: y = X^3
+ 2*X^2 + 3*X + noise, where noise is a Gaussian noise
added to the data points.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 345
Polynomial Regression
Quiz: Polynomial regression
Task Requirements:
a. Polynomial Regression (Degree 2 and 3):
Implement linear regression for the dataset.
Implement polynomial regression with a degree of 2
and 3.
Fit a linear model and two polynomial models (degree
2 and degree 3) to the dataset.
b. Predictionpredict the target:
For each model (linear, quadratic, cubic), variable y for a new
set of X values (X_new) in the range from -3 to 3.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 346
Polynomial Regression
Quiz: Polynomial regression
Task Requirements:
c. Visualization:
Plot the original dataset
Plot the predicted values for the three model
d. Calculate the MSE for each model (linear regression,
quadratic regression, cubic regression) using the true target
variable y and the predicted values from each model.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 347
Polynomial Regression
Polynomial regression quiz (solution)
import numpy as np
import [Link] as plt
from sklearn.linear_model import LinearRegression
from [Link] import PolynomialFeatures
from [Link] import mean_squared_error
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 348
Polynomial Regression
Polynomial regression quiz (solution) – cont’d
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 349
Polynomial Regression
Polynomial regression quiz (solution) – cont’d
# Step 3: Create and fit the models
lin_reg = LinearRegression() # Linear Regression model
poly_reg_2 = LinearRegression() # Polynomial Regression of
degree 2
poly_reg_3 = LinearRegression() # Polynomial Regression of
degree 3
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 350
Polynomial Regression
Polynomial regression quiz (solution) – cont’d
# Step 4: Make predictions
X_new = [Link](-3, 3, 100).reshape(100, 1) # New X values
for prediction
X_new_poly_2 = poly_features_2.transform(X_new) # Polynomial
features of degree 2
X_new_poly_3 = poly_features_3.transform(X_new) # Polynomial
features of degree 3
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 351
Polynomial Regression
Polynomial regression quiz (solution) – cont’d
# Step 5: Plot the results
[Link](figsize=(10, 6))
[Link](X, y, color='blue', label='Original Dataset') #
Plot original dataset points
[Link](X_new, y_pred_lin, color='red', label='Linear
regression (Degree 1)')
[Link](X_new, y_pred_poly_2, color='green', label='Polynomial
regression (Degree 2)')
[Link](X_new, y_pred_poly_3, color='orange',
label='Polynomial regression (Degree 3)')
[Link]('x')
[Link]('y')
[Link]('Polynomial Regression Comparisons')
[Link]()
[Link](True)
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 352
Polynomial Regression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 353
Polynomial Regression
Polynomial regression quiz (solution) – cont’d
# Step 6: Calculate and print Mean Squared Error for each model
mse_lin = mean_squared_error(y, lin_reg.predict(X))
mse_poly_2 = mean_squared_error(y,
poly_reg_2.predict(X_poly_2))
mse_poly_3 = mean_squared_error(y,
poly_reg_3.predict(X_poly_3))
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 355
Logistic Regression - Example
Objective of the prediction:
The goal is to predict whether a student will pass a course using
logistic regression, focusing on the following:
Training the model: Fit the logistic regression model to the
data (study and sleep hours) to learn the pass/fail
relationship.
Evaluating the model: Assess the model's accuracy in
predicting pass/fail on test data.
Making predictions: Use the trained model to predict the
probability of passing for new students based on their study
and sleep hours.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 356
Logistic Regression - Example
Logistic regression
import pandas as pd example: Task 1. Create the Dataset
import numpy as np
# Create a DataFrame
df = [Link]({
'study_hours': study_hours,
'sleep_hours': sleep_hours,
'passed': [Link](int)
})
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 358
Logistic Regression - Example
Logistic regression example: Task 3. Build a Logistic
Regression Model
from sklearn.linear_model import LogisticRegression
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 359
Logistic Regression - Example
Logistic regression example: Task 4. Evaluate the Model
from [Link] import accuracy_score
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 360
Logistic Regression - Example
Logistic regression example: Task 5. Predict
# Make a prediction
prediction = [Link](new_data)
>>
Predicted class: 1
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 361
Logistic Regression - Example
Logistic regression example: Task 6. Plot the Results
import [Link] as plt
from [Link] import ListedColormap
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 362
Logistic Regression - Example
Logistic regression example: Task 6. Plot the Results – cont’d
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 363
Logistic Regression - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 364
Navie Bayes - Example
Dataset Description: A synthetic dataset is created to predict a
patient's medical condition based on two features:
Age (in years)
Blood Pressure (in mmHg)
⇒ The target variable, Condition: Binary (1 for Positive
condition, 0 for Negative condition).
Condition: 'Positive' if (Age > 50) and (Blood Pressure > 120);
else 'Negative'.
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 365
Navie Bayes - Example
Instructions::
1. Step 1: Generate synthetic data for age, blood pressure, and
medical condition.
2. Step 2: Separate features (X) and target (y), then split into
training and testing sets..
3. Step 3: Train a Naive Bayes model using the training data..
4. Step 4: Predict on test data and calculate accuracy..
5. Step 5: Plot results to visualize predictions versus actual
data.
6. Step 6: Predict the condition of a new patient based on age
and blood pressure..
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 366
Navie Bayes - Example
Navie Bayes example: Step 1. Data Generation
import pandas as pd
import numpy as np
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 367
Navie Bayes - Example
Navie Bayes example: Step 1. Data Generation
# Create a DataFrame
df = [Link]({
'Age': age,
'Blood Pressure': blood_pressure,
'Condition': condition
})
>>
Age Blood Pressure Condition
0 58 157 1
1 71 166 1
2 48 141 0
3 34 119 0
4 62 164 1
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 368
Navie Bayes - Example
Navie Bayes example: Step 3. Data Splitting
# Split the data into training and testing sets (80% training,
20% testing)
X_train, X_test, y_train, y_test = train_test_split(X, y,
test_size=0.2, random_state=42)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 369
Navie Bayes - Example
Navie Bayes example: Step 4. Model Training
from sklearn.naive_bayes import GaussianNB
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 370
Navie Bayes - Example
Navie Bayes example: Step 5. Model Evaluation
from [Link] import accuracy_score
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 371
Navie Bayes - Example
Navie Bayes example: Step 6. Visualization
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 372
Navie Bayes - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 373
Navie Bayes - Example
Navie Bayes example: Step 7. New Prediction
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 374
Decision tree - Example
Dataset Description: Iris Dataset:
Features (Input variables):
Sepal length
Sepal width
Petal length
Petal width
⇒ Target variable (Output variable): Species (setosa, versicolor,
virginica)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 375
Decision tree - Example
Decision tree - Example
import [Link] as plt
from [Link] import plot_tree
from sklearn.model_selection import train_test_split
from [Link] import DecisionTreeClassifier
from sklearn import metrics
from [Link] import load_iris
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 376
Decision tree - Example
Decision tree - Example
# Train the model
clf = [Link](X_train,y_train)
# Model Accuracy
print("Accuracy:",metrics.accuracy_score(y_test, y_pred))
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 377
Decision tree - Example
Decision tree - Example
# Plot the confusion matrix
from [Link] import confusion_matrix
import seaborn as sns
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 378
Decision tree - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 379
Decision tree - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 380
K-nearest neighbors (KNN)- Example
K-nearest neighbors (KNN)- Example
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from [Link] import KNeighborsClassifier
from [Link] import accuracy_score
import [Link] as plt
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 381
K-nearest neighbors (KNN)- Example
K-nearest neighbors (KNN)- Example
# Create a DataFrame
df = [Link]({
'Hours of Study': hours_of_study,
'Attendance Rate': attendance_rate,
'Final Grade': final_grade
})
# Split the data into training and testing sets (80% training, 20%
testing)
X_train, X_test, y_train, y_test = train_test_split(X, y,
test_size=0.2,
11/14/2025 random_state=42)
Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 382
K-nearest neighbors (KNN)- Example
K-nearest neighbors (KNN)- Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 384
K-nearest neighbors (KNN)- Example
K-nearest neighbors (KNN)- Example
# Plot test data predictions (Pass: Blue, Fail: Yellow)
test_pass = [Link](X_test['Hours of Study'][y_pred==1],
X_test['Attendance Rate'][y_pred==1], color='blue',
edgecolors='k', marker='x', s=80, label='Test Pass')
test_fail = [Link](X_test['Hours of Study'][y_pred==0],
X_test['Attendance Rate'][y_pred==0], color='yellow',
edgecolors='k', marker='x', s=80, label='Test Fail')
# Create legend
[Link](loc='best')
[Link]('Hours of Study')
[Link]('Attendance Rate')
[Link]('KNN for Student Final Grade Prediction')
[Link]()
>>
Model Accuracy on Test Data: 0.97
Predicted Final Grade: Pass
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 385
K-nearest neighbors (KNN)- Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 386
Support Vector Machine (SVM) - Example
Support Vector Machine (SVM) - Example
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from [Link] import SVC
from [Link] import accuracy_score
import [Link] as plt
# Split the data into training and testing sets (80% training, 20%
testing)
X_train, X_test, y_train, y_test = train_test_split(X, y,
test_size=0.2, random_state=42)
# Example new data for prediction (patient with age 50 and BMI 28)
new_data = [Link]({
'Age': [50],
'BMI': [28]
})
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 389
Support Vector Machine (SVM) - Example
Support Vector Machine (SVM) - Example
# Optional: Visualizing the SVM decision boundary
[Link](figsize=(8, 6))
[Link](X_train['Age'], X_train['BMI'], c=y_train,
cmap='coolwarm', label='Training Data', edgecolors='k', s=50)
[Link](X_test['Age'], X_test['BMI'], c=y_pred, cmap='winter',
marker='x', label='Test Predictions', s=80)
# Create grid for decision boundary visualization
x_min, x_max = X['Age'].min() - 1, X['Age'].max() + 1
y_min, y_max = X['BMI'].min() - 1, X['BMI'].max() + 1
xx, yy = [Link]([Link](x_min, x_max, 0.1), [Link](y_min,
y_max, 0.1))
# Plot decision boundary
Z = [Link](np.c_[[Link](), [Link]()])
Z = [Link]([Link])
[Link](xx, yy, Z, alpha=0.3, cmap='coolwarm')
[Link]('Age')
[Link]('BMI')
[Link]('SVM for Disease Prediction')
[Link]()
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 390
Support Vector Machine (SVM) - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 391
Neural Network (NN) Model- Example
Neural
import Network (NN)
pandas as pd Model - Example
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.neural_network import MLPClassifier
from [Link] import accuracy_score
import [Link] as plt
from [Link] import StandardScaler
# Preprocess the data: Split the data into training and testing sets
(80% training, 20% testing)
X_train, X_test, y_train, y_test = train_test_split(X, y,
test_size=0.2, random_state=42)
# Example new data for prediction (customer with age 35 and income
80K)
new_data = [Link]([[35, 80]])
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 395
Neural Network (NN) Model- Example
Neural Network (NN) Model - Example
# Create grid for decision boundary visualization
x_min, x_max = X_train_scaled[:, 0].min() - 1, X_train_scaled[:,
0].max() + 1
y_min, y_max = X_train_scaled[:, 1].min() - 1, X_train_scaled[:,
1].max() + 1
xx, yy = [Link]([Link](x_min, x_max, 0.1), [Link](y_min,
y_max, 0.1))
[Link]('Age')
[Link]('Income')
[Link]('Neural Network for Customer Purchase Prediction')
[Link]()
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 396
Neural Network (NN) Model- Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 397
K-Means Clustering - Example
K-Means Clustering - Example
import pandas as pd
import numpy as np
from [Link] import StandardScaler
from [Link] import KMeans
import [Link] as plt
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 398
K-Means Clustering - Example
K-Means Clustering - Example
# Create a DataFrame
df = [Link]({
'Annual Income': annual_income,
'Spending Score': spending_score
})
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 399
K-Means Clustering - Example
K-Means Clustering - Example
# Determine the optimal number of clusters using the Elbow Method
inertia = []
K = range(1, 11)
for k in K:
kmeans = KMeans(n_clusters=k, random_state=42)
[Link](X)
[Link](kmeans.inertia_)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 400
K-Means Clustering - Example
K-Means Clustering - Example
# Based on the Elbow Method, choose the optimal number of clusters
(e.g., k=4)
optimal_k = 4
kmeans = KMeans(n_clusters=optimal_k, random_state=42)
[Link](X)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 401
K-Means Clustering - Example
K-Means Clustering - Example
# Example new data for clustering (customers with new income and
spending scores)
new_data = [Link]([[50, 75], [85, 20]])
new_data_scaled = [Link](new_data)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 402
K-Means Clustering - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 403
K-Means Clustering - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 404
Hierarchical Clustering - Example
Hierarchical Clustering - Example
import pandas as pd
import numpy as np
from [Link] import StandardScaler
from [Link] import dendrogram, linkage, fcluster
import [Link] as plt
# Create a DataFrame
df = [Link]({
'Study Hours': study_hours,
'Exam Scores': exam_scores
})
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 405
Hierarchical Clustering - Example
Hierarchical Clustering - Example
# Print the first few rows of the dataset
print([Link]())
# Create a dendrogram
[Link](figsize=(10, 7))
dendrogram(Z, truncate_mode='level', p=5)
[Link]('Dendrogram for Hierarchical Clustering')
[Link]('Data Points')
[Link]('Euclidean Distance')
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 406
Hierarchical Clustering - Example
Hierarchical Clustering - Example
# Based on the dendrogram, choose the optimal number of clusters
(e.g., 3)
optimal_clusters = 3
clusters = fcluster(Z, optimal_clusters, criterion='maxclust')
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 407
Hierarchical Clustering - Example
Hierarchical Clustering - Example
# Example new data for clustering (students with new study hours and
exam scores)
new_data = [Link]([[5, 80], [9, 95]])
new_data_scaled = [Link](new_data)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 409
Hierarchical Clustering - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 410
Gaussian Mixture Model- Example
Gaussian Mixture Model - Example
import pandas as pd
import numpy as np
from [Link] import StandardScaler
from [Link] import GaussianMixture
import [Link] as plt
# Create a DataFrame
df = [Link]({
'Annual Spending': annual_spending,
'Frequency of Visits': frequency_of_visits
}) 11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 411
Gaussian Mixture Model- Example
Gaussian Mixture Model - Example
# Based on the BIC, choose the optimal number of clusters (e.g., n=3)
optimal_n = 3
gmm = GaussianMixture(n_components=optimal_n, random_state=42)
[Link](X)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 413
Gaussian Mixture Model- Example
Gaussian Mixture Model - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 414
Gaussian Mixture Model- Example
Gaussian Mixture Model - Example
# Example new data for clustering (customers with new annual spending
and visit frequencies)
new_data = [Link]([[30, 10], [70, 40]])
new_data_scaled = [Link](new_data)
>>
Predicted Clusters for New Data: [2 0]
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 415
Gaussian Mixture Model- Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 416
Gaussian Mixture Model- Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 417
DBSCAN Clustering - Example
DBSCAN Clustering- Example
import numpy as np
from [Link] import DBSCAN
import [Link] as plt
# Compute DBSCAN
db = DBSCAN(eps=0.3, min_samples=10).fit(X)
core_samples_mask = np.zeros_like(db.labels_, dtype=bool)
core_samples_mask[db.core_sample_indices_] = True
labels = db.labels_
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 418
DBSCAN Clustering - Example
DBSCAN Clustering - Example
# Plot result
unique_labels = set(labels)
colors = [[Link](each)
for each in [Link](0, 1, len(unique_labels))]
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 419
DBSCAN Clustering - Example
DBSCAN Clustering - Example
for k, col in zip(unique_labels, colors):
if k == -1:
# Black used for noise.
col = [0, 0, 0, 1]
class_member_mask = (labels == k)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 420
DBSCAN Clustering - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 421
DBSCAN Clustering - Example
DBSCAN Clustering - Example
for k, col in zip(unique_labels, colors):
if k == -1:
# Black used for noise.
col = [0, 0, 0, 1]
class_member_mask = (labels == k)
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 422
Principal components analysis - Example
Principal components analysis - Example
import pandas as pd
import numpy as np
from [Link] import load_iris
from [Link] import StandardScaler
from [Link] import PCA
import [Link] as plt
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 424
Principal components analysis - Example
Principal components analysis - Example
# Visualize the results
[Link](figsize=(10,8))
[Link](x=finalDf['principal component 1'],
y=finalDf['principal component 2'],
c=finalDf['target'],
cmap='viridis')
[Link]('Principal Component 1')
[Link]('Principal Component 2')
[Link]('2D PCA of Iris Dataset')
[Link]()
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 426
Factor analysis - Example
Factor analysis - Example
import pandas as pd
from [Link] import StandardScaler
from [Link] import FactorAnalysis
import [Link] as plt
import numpy as np
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 427
Factor analysis - Example
Factor analysis - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 428
Factor analysis - Example
Factor analysis - Example
# Visualize the factor loadings
[Link](figsize=(10, 8))
loadings = fa.components_.T
[Link](loadings, cmap='viridis', aspect='auto')
[Link]()
[Link](range(n_components), [f'Factor {i+1}' for i in
range(n_components)])
[Link](range([Link][1]), [Link])
[Link]('Factor Loadings')
[Link]('Factors')
[Link]('Features')
[Link]()
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 429
Factor analysis - Example
Factor analysis - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 430
Factor analysis - Example
11/14/2025 Dr. Mai Cao Lan - Faculty of Geology & Petroleum Engineering, HCMUT 431