SUBJECT NAME – Elements of AI - ML
LAB ACTIVITY – 5
Implementation of Different Techniques for
Handling Imbalanced Data
SUBMITTED BY:
NAME – SHIVANG RATURI
SAP ID – 590022331
BATCH – 46
SUBMITTED TO:
DR. SANDEEP CHAND KUMAIN
1. Aim
To implement and compare different techniques for
handling imbalanced datasets using machine learning.
2. Theory
In machine learning, an imbalanced dataset occurs
when one class has significantly more samples than the
other.
Example:
Class 0 → 90%
Class 1 → 10%
This causes models to become biased toward the
majority class.
Techniques to Handle Imbalance
1. Random Under Sampling
2. Random Over Sampling
3. SMOTE (Synthetic Data Generation)
4. Class Weighting
3. Requirements
pip install numpy pandas matplotlib seaborn scikit-learn
imbalanced-learn
4. Implementation
Step 1: Import Libraries
Step 2: Create Imbalanced Dataset
Step 3: Visualize Data
Step 4: Train-Test Split
[Link] Training
6. Techniques Implementation
A. Random Under Sampling
B. Random Over Sampling
C. SMOTE
D. Class Weighting
[Link] of Results
8. Observations
Original model shows high accuracy but poor
minority prediction
Under Sampling loses data
Over Sampling may cause overfitting
SMOTE improves generalization
Class Weighting balances learning effectively
9. Conclusion
Handling imbalanced data is crucial in ML
SMOTE and Class Weighting give better
performance
F1-score is a better metric than accuracy for
imbalanced data
OUTPUT:
LEARNING OUTCOMES
Understand imbalanced datasets and their impact on machine
learning model performance.
Use appropriate evaluation metrics like precision, recall, F1-
score, and ROC-AUC instead of accuracy.
Apply techniques to handle imbalance, such as oversampling,
undersampling, and class weighting.
Compare and interpret model performance to select the most
effective method for a given problem.