Business Analytics
Q.1
Introduction
In today‟s retail world, data is basically the heartbeat of the business. Every small decision, like when to
launch a discount, how much to charge for delivery or which store is underperforming, depends on
accurate numbers showing up on the sales dashboards. If the numbers are wrong, decisions can quickly
go off track.
In this case, the retail chain is ready to launch a brand new analytics dashboard to track sales. But there‟s
a problem, in the dataset, several entries in the delivery amount column are missing. These gaps exist
because of system glitches or simple human typing errors. If the analyst ignores them, the dashboard
will give a distorted picture of performance. If they are filled carelessly, management could be misled.
This is where imputation comes in, the process of filling in missing data with estimated values so the
dataset remains consistent and useful. The tricky part is figuring out how to fill those blanks in a way
that actually reflects business reality. The most common ways are mean, median and mode imputation.
Each sounds simple, but in practice, the choice can completely change the story the dashboard tells.
Why filling the gaps matters?
Missing values are not just empty boxes on a spreadsheet. They can skew the entire picture. Think of it
this way, if 20% of your delivery amounts are missing, your average revenue per order could look far
lower than reality. If the missing ones are mostly from high value orders, it may appear that delivery
charges are always cheap. On the other hand, filling those blanks with unrealistically high numbers
might trick managers into thinking delivery is a huge profit center.
For retail, this is dangerous. Delivery charges impact:
Pricing strategy- If delivery seems too expensive, management might cut charges unnecessarily.
Customer satisfaction- A wrong picture could make it look like customers are overpaying.
Operations planning- Misjudging delivery revenue could lead to wrong budgeting for logistics.
So, imputation isn‟t just about math. It‟s about protecting decision making, ensuring managers act on
truth, not illusions.
Mean Imputation
The most straightforward fix is to replace missing delivery amounts with the average of all available
ones. Example: If recorded deliveries are Rs. 50, Rs. 60, Rs. 70 and one is missing, we take the average-
Rs. 60 and use it for the blank. Why it works? Simple, quick and keeps the dataset intact. The overall
mean doesn‟t change. Why it‟s risky? Delivery amounts in retail often have outliers. Imagine one rare
case where a delivery was Rs. 5,000 for a bulk order. The average will shoot up and imputing missing
values with that inflated number will create false impressions. So, mean works best if delivery charges
are usually clustered together without extremes. But retail rarely works that neatly.
Median Imputation
Here, the missing value is filled with the „middle‟ number in the sorted list. Example: Say the delivery
amountd are Rs. 40, Rs. 50, Rs. 55, Rs. 60 and Rs. 500. The median is Rs. 55. If one value is missing,
we use Rs. 55. Why it wirks? Resistant to outliers. That one Rs. 500 doesn‟t drag the result into absurb
territory. Why it‟s risky? By replacing blanks with the same middle value again and again, we may lose
the natural spread of data. But in retail, this method is usually safer. Delivery charges often depend on
distance or order value, meaning skewed data is normal. Median gives a realistic typical charge.
Mode Imputation
Here, missing values are replaced with the most common delivery charge. Example: If most customers
pay Rs. 50 and some data is missing, we simply fill in Rs. 50. Why it works? Perfect when delivery
pricing follows slabs like Rs. 50 for local, Rs. 100 for regional. Why it‟s risky? If delivery amounts vary
widely, repeatedly imputing the same number flattens diversity and hides patterns. This works well only
if the company already has fixed delivery has fixed delivery slabs. Otherwise, it oversimplifies.
Worked Example
Let‟s test this with a small dataset:
[50, 60, 55, 500, 45, NaN, 50, NaN, 55, 60]
Mean Imputation: Average- Rs. 159.5. The missing values become 159.5. But almost all real
charges are around Rs. 50- Rs. 60, so this massively distorts reality.
Mediain imputation: Middle value- Rs. 55. The missing values become 55. This better reflects
reality because most values cluster around that range.
Mode imutation: Mode- Rs. 50. Missing values become 50. If the retailer mostly charges Rs.
50, this is valid.
Looking at this, median or mode clearly tells the truer story than mean.
Beyond Simple Imputation
Retail analysts today also use smarter methods:
1. Regression imputation- Predict delivery amount based on order value, distance or weight. For
example, larger orders might automatically carry higher delivery charges.
2. K- Nearest Neighbors- Find orders similar to the missing one and average their delivery
charges.
3. Multiple imputation- Fill in several possible values and combine them reducing bias.
4. Business rule logic- Sometimes, the best method is not statistical but operational. For this retail
chain, sucj rules might capture reality better than pure math.
How to decide: Key considerations
Before filling gaps, the analyst should pause and ask, What‟s the data distribution? If normal, mean
works fine. If skewed, median is safer. Are there outliers? If yes, avoid mean. What‟s the dashboard‟s
purpose? If management wants to study averages, don‟t distort them. If the goal is customer behavior,
reflect the most common patterns. The method must fit both the data and the business context.
Conclusion
Imputation is the bridge that keeps the dataset whole, but the choice of method matters. Mean is quick
but risky if there are outliers. Median is stable, especially when data is skewed. Mode works best when
delivery charges follow fixes slabs. Beyond these, advanced methods and business logic can improve
accuracy even further. At the end of the day, data analysis is not just about crunching numbers. The data
analysis is about telling the truth of the business.
Q. 2 (A)
Introduction
In e- commerce, customer satisfaction is the real currency. Happy customers come back, spend more and
spread the word. After running statistical analysis, Mehta E- Commerce figured out that three thingd
matter most: Product quality, delivery speed and customer support. Now comes the harder part,
deciding how to spend limited money and resources on these areas. Data might point to what‟s
significant, but management has to weight that against business priorities, budget limits and even the
risks of making the wrong call.
Statistic vs. Reality
Here‟s where many companies trip. Just because something is statistically significant doesn‟t mean it‟s
the smartest thing to act on right away. Delivery speed may rank highly, but if most deliveries are
already within 2 days, cutting it to one day could double logistics costs while adding only a tiny bump in
satisfaction. Product quality might take less investment but has a ripple effect, fewer returns, fewer
complaints and more trust. That‟s big for long term loyalty. Customer support won‟t fix a bad product or
late delivery, but it can save a relationship. A single positive call can turn an angry customer into a
repeat buyer. So the company must decide not just what matters statistically, but what matters most for
survival and growth.
Strategic Choices and Trade offs
With a limited budget, Mehta E- commerce has to play a balancing game.
1. Double down on product quality
Every rupee spent here prevents complaints and refund later. It‟s the long term play, a reputation
for “things just work” keeps customers coming back without extra marketing spend.
2. Boost customer support
It‟s cheaper to improve than logistics. Training staff, adding self-help tools or AI chatbots can
create quick, visible improvements. Great for short term wins which customer notices
immediately.
3. Re- evaluate delivery speed
Re- evaluating delivery speed does matter. But trying to match Amazon level speed could drain
profits. Improve consistency instead of chasing extremes. Promise 2 day delivery and actually
deliver 2 day delivery every time. Reliability can sometimes beat raw speed.
Unintended Consequences
Every shiny investment has a shadow side. Over focus on delivery could burn money and pressure
logistics staff into mistakes. Overspending on quality could raise prices too much, scaring off cost
conscious buyers. Expanding customer support without fixing root problems could just mean more
people apologizing for the same recurring issues.
Conclusion
For Mehta E- commerce, the statistics are like a compass pointing north- quality, speed and support
matter. But strategy decides the journey. Invest in product quality to build unshakable trust. Strengthen
customer support for quick visible wins. Tweak delivery speed smartly instead of blowing the budget
chasing impossible targets. Numbers highlight the factors, but leadership decides the priorities. When
Mehta E- commerce balances data with strategy, customers stay happy and the business sails forward-
strong, steady and future ready.
Q. 2 (B)
Introduction
A retail company built a simple linear regression model to forecast monthly sales using advertising
spend as the key driver. The analytics team celebrated a high R squared value, claiming the model was
highly reliable. Management quickly took that as proof that the forecast was trustworthy. But a few team
members questioned whether R squared alone could capture the fully reality of how sales actually work.
After all, business performance depends on far more than ad spend- factors like competition, pricing,
consumer trends and seasonality also shape results. The debate raises an important question: how
effective R squared as the main tool to judge a model‟s success?
The illusion of Accuracy
R squared measures how much of the variation in sales is explained by advertising spend. A value of
0.85 means the model explains 85% of the changes in sales. Impressive on paper, but in business,
correlation doesn‟t equal causation. Just because sales rise when ad spend increases doesn‟t prove ads
caused those sales. Maybe there was a festive season or a viral product launch that drove both at once.
That‟s where the illusion begins. A high R squared only shows how well the model fits past data, not
whether it will predict the future accurately. It‟s like saying someone drives perfectly on a simulator and
assuming they‟ll do just as well on a busy road. The number looks great, but the context matters more.
The Over fitting Trap
Another hidden danger is over fitting- when the model becomes too perfect for its own good. It
memorizes every detail of historical data instead of learning general patterns. So while R squared soars,
real world accuracy collapses. Like, if sales spiked one month due to a celebrity endorsement, the model
might assume that spike will happen every year, leading to unrealistic forecasts. R squared can also rise
when irrelevant variables are added, giving a false sense of improvement. It doesn‟t test whether those
variables genuinely matter or if they just clutter the model. In short, R squared rewards complexity, not
necessarily correctness.
A smarter way to evaluate Models
To truly assess how well a regression model performs, R- squared must be balanced with other
diagnostic tools. Adjusted R squared keeps the score honest by penalizing unnecessary predictors.
Residual analysis checks if errors are randomly scattered or reveal hidden biases. Error metrics such as
Mean Absolute Error show how close predictions are to reality. Cross validation tests how the model
performs on unseen data, ensuring it can adapt to new situations. Together these give a more rounded
view of performance, combining statistical strength with real world reliability.
Conclusion
R squared is a helpful starting point, but it‟s not the final verdict. It tells you how well your model fits
history, not how wisely it forecasts the future. For the retail company, trusting R squared alone could
lead to costly missteps. The real goal is to blend data accuracy with business insight- understanding
what the numbers mean in the marketplace. When R squared is balanced with other diagnostic checks
and strategic judgment, the model becomes more than a mathematical fit; it becomes a dependable
decision making tool.