Retail Demand
Forecasting using Machine
Learning
Predicting SKU-Level Sales Across Stores
Your Name, Date, Organization/Institution
Problem Statement
Our goal is to forecast SKU-level units sold using historical retail data.
This will optimize inventory, reduce stockouts, and improve
promotional effectiveness, directly impacting business growth.
Dataset Overview
Data Size Key Features
150,150 rows covering 2011–2013. store_id, sku_id, week, total_price, base_price,
is_featured_sku, is_display_sku, units_sold.
The target variable for our prediction is units_sold.
EDA: Trend Over Time
Analysis of Units Sold & Total Price over Time reveals significant
seasonal demand spikes.
Recurring peaks suggest a clear cyclical behavior in sales patterns,
crucial for accurate forecasting.
EDA: Top Performers
Top 10 Stores Top SKUs
Ranked by total units sold, Identified based on cumulative
exceeding 300K+ units. revenue.
Insight: A small number of stores and SKUs disproportionately drive
overall demand.
EDA: Price & Promotion
Impact
Histograms of base_price and total_price show their distribution.
Only ~10% of SKUs are featured, yet they exhibit 34–36% higher
correlation with units_sold, highlighting promotion effectiveness.
EDA: Correlation Heatmap
Strong correlation between total_price and base_price (~0.96).
Moderate negative correlation between base_price and units_sold (-24%).
Price and visibility are key sales drivers; is_featured_sku and
is_display_sku show mild positive correlation with demand.
Data Cleaning & Feature Engineering
Temporal Features
Encoding Categoricals
Extracted week and year to
Outlier Removal
One-hot encoded store_id and capture time-based patterns.
Removed top 1% outliers in sku_id; target encoding used for
units_sold to improve model 10K+ SKUs.
accuracy.
Model Building
We trained two types of models:
Linear Regression (Base & Optimized)
Random Forest Regressor (Base & Optimized)
Data split 80/20 for training/testing. Evaluation metrics: R², MSE, RMSE.
Model Performance
Linear Regression (Base) 0.26 32.00
Random Forest (Base) 0.73 31.99
Linear Regression (Tuned) 0.57 32.00
Random Forest (Tuned) 0.78 32.00
Best Model: Random Forest with optimized features and categorical handling.