0% found this document useful (0 votes)
4 views10 pages

Exercise Advanced Classification

The document outlines various machine learning exercises involving Naive Bayes, k-NN, and Logistic Regression algorithms for predicting events such as flash floods, cyberbullying, fake news, and classifications of environmental and health metrics. Each exercise provides historical data and requires calculations of probabilities or distances to classify new instances. The tasks involve applying statistical methods to assess outcomes based on given features.

Uploaded by

OUI OUI
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views10 pages

Exercise Advanced Classification

The document outlines various machine learning exercises involving Naive Bayes, k-NN, and Logistic Regression algorithms for predicting events such as flash floods, cyberbullying, fake news, and classifications of environmental and health metrics. Each exercise provides historical data and requires calculations of probabilities or distances to classify new instances. The tasks involve applying statistical methods to assess outcomes based on given features.

Uploaded by

OUI OUI
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Exercise Advance Classification

1. The Malaysian Meteorological Department is developing a prediction model for flash floods
in Kuala Lumpur. You are provided with a historical summary of 10 days of weather data and
the occurrence of floods. A new day is recorded with Rainfall Intensity: High and Monsoon
Season: No. Using the Naive Bayes algorithm, calculate the posterior probabilities and
determine whether a Flash Flood is predicted to occur.

Day Rainfall Intensity Monsoon Season Flash Flood (Target)


1 High Yes Yes
2 High Yes Yes
3 High No Yes
4 Low Yes Yes
5 High Yes Yes
6 High No Yes
7 High Yes No
8 Low No No
9 Low Yes No
10 Low No No

2. A Malaysian social media platform is developing an automated filter to detect Cyberbullying


in comments. The system analyzes comments based on two features: Profanity Content and
Targeted Mentions. A new comment is flagged with Profanity Content: High and Targeted
Mentions: No. Using the Naive Bayes algorithm, calculate the posterior probabilities and
determine whether the comment should be classified as Bullying: Yes or Bullying: No.

Comment ID Profanity Content Targeted Mentions Bullying? (Target)


1 High Yes Yes
2 High Yes Yes
3 High Yes Yes
4 High Yes Yes
5 High No Yes
6 High No Yes
7 High No Yes
8 Low Yes Yes
9 Low Yes Yes
10 High No No
11 Low Yes No
12 Low Yes No
13 Low No No
14 Low No No
15 Low No No
3. A global news agency is building a machine learning tool to flag potential Fake News articles.
The model evaluates headlines based on two features: Capitalization Level (e.g.,
"SENSATIONAL TITLES") and Source Reliability (based on historical verification). A new article
headline is detected with Capitalization Level: High and Source Reliability: High. Using the
Naive Bayes algorithm, calculate the posterior probabilities and determine whether the
article should be classified as Fake News: Yes, or Fake News: No.

Article ID Capitalization Level Source Reliability Fake News? (Target)


1 High Low Yes
2 High Low Yes
3 High Low Yes
4 High Low Yes
5 High Low Yes
6 High Low Yes
7 High Low Yes
8 High Low Yes
9 High Low Yes
10 High High Yes
11 Normal High Yes
12 Normal High Yes
13 High Low No
14 High Low No
15 Normal Low No
16 Normal Low No
17 Normal Low No
18 Normal High No
19 Normal High No
20 Normal High No
4. A global environmental agency uses a k-NN model to classify city zones into "Healthy" or
"Unhealthy" based on two sensor readings: Particulate Matter (PM2.5) levels and Nitrogen
Dioxide ( NO2 ) concentrations. A new industrial zone in a developing city reports a reading of
PM2.5 = 35 and NO2 = 40 .

Zone ID PM2.5 Level NO2 Level Class (Target)


A 10 15 Healthy
B 20 25 Healthy
C 50 60 Unhealthy
D 45 55 Unhealthy
New Zone 35 40 ?

(i) Calculate the Euclidean distance between the new zone and all historical zones.
(ii) Using k = 3 , determine the classification of the new industrial zone.

5. The Ministry of Natural Resources and Environmental Sustainability is classifying city blocks
in the Klang Valley into "Heat Vulnerable" or "Climate Resilient" based on two environmental
metrics: Vegetation Density (NDVI score) and and Impervious Surface Percentage
(Pavement/Concrete cover). A new residential development in Cheras reports a Vegetation
Density of 40 and an Impervious Surface Percentage of 60.

Block ID Vegetation Density ( X 1 ) Impervious Surface ( X 2 ) Class (Target)


A 80 20 Climate Resilient
B 70 30 Climate Resilient
C 20 85 Heat Vulnerable
D 30 75 Heat Vulnerable
New
40 60 ?
(Cheras)

(i) Calculate the Euclidean distance between the new development and all historical
blocks.
(ii) Using k = 3 , determine the classification of this new development.
6. A conservation group in Sabah is using k-NN to classify forest patches as "Suitable" or
"Unsuitable" for Orangutan translocation based on two metrics: Canopy Height (m) and
Distance to Human Settlement (km). A new forest patch is surveyed with a Canopy Height of
20m and a Distance to Human Settlement of 5km.

Site ID Canopy Height ( X 1 ) Distance to Settlement ( X 2 ) Class (Target)


A 25 10 Suitable
B 30 8 Suitable
C 10 2 Unsuitable
D 15 3 Unsuitable
New
20 5 ?
Site

(i) Calculate the Euclidean distance between the new patch and all historical sites.
(ii) Using k = 3 , determine if the new patch is classified as Suitable or Unsuitable.

7. A retail chain in Malaysia, Pasaraya Jaya, is using a Logistic Regression model to predict the
probability that a customer will sign up for a "Premium Membership" ( y = 1 ) based on their
Average Monthly Spend (in hundreds of RM) and Years as a Customer. A customer has an
Average Monthly Spend of RM 500 (Value = 5) and has been a customer for 4 years.

Feature Input Value ( x ) Coefficient ( w )


Intercept 1 -4.5
Monthly Spend (Units of 100) 5 0.6
Loyalty Years 4 0.3

(i) Calculate the linear combination ( z ) for this customer.


(ii) Calculate the probability ( P ) of this customer signing up using the Sigmoid function.
(iii) Based on a decision threshold of 0.6 (60%), should the marketing team send a
premium invitation to this customer?
8. A global health organization is using Logistic Regression to predict the probability of a patient
developing Type 2 Diabetes ( y = 1 ) based on their Body Mass Index (BMI) and Daily Sugar
Intake (in grams). A patient has a BMI of 30 and a Daily Sugar Intake of 80 grams.

Feature Patient Value ( x ) Coefficient ( w )


Intercept 1 -8.0
Body Mass Index (BMI) 30 0.15
Sugar Intake (Grams) 80 0.05

(i) Calculate the linear combination ( z ) for this patient.


(ii) Calculate the probability ( P ) of this patient being classified as "High Risk" for
Diabetes using the Sigmoid function.
(iii) Based on a clinical decision threshold of 0.5 (50%), classify the patient as "High
Risk" ( y = 1 ) or "Low Risk" ( y = 0 )

9. Sustainable Energy Development Authority (SEDA) Malaysia is using a Logistic Regression


model to predict the probability that a household will install Rooftop Solar Panels ( y = 1 )
based on their Average Monthly Electricity Bill (in hundreds of RM) and the House Age (in
years). A homeowner in Subang Jaya has an Average Monthly Electricity Bill of RM 600 (Value
= 6) and a House Age of 8 years.

Feature Household Value ( x ) Coefficient ( w )


Intercept 1 -6.2
Monthly Bill (Units of 100) 6 0.8
House Age (Years) 8 -0.1

(i) Calculate the linear combination ( z ) for this household.


(ii) Calculate the probability ( P ) of this household adopting solar energy using the
Sigmoid function.
(iii) Based on a decision threshold of 0.5 (50%), classify whether this household is a
"Likely Adopter" ( y = 1 ) or "Unlikely Adopter" ( y = 0 ).

You might also like