0% found this document useful (0 votes)
12 views9 pages

Business Analytics

The document discusses the importance of data imputation in retail analytics, emphasizing that filling missing values accurately is crucial for informed decision-making. It outlines various imputation methods such as mean, median, and mode, highlighting their advantages and risks, and suggests advanced techniques for better accuracy. Additionally, it addresses the significance of balancing statistical insights with strategic business decisions in e-commerce and the limitations of relying solely on R squared for evaluating regression models.

Uploaded by

dharavparmar22
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views9 pages

Business Analytics

The document discusses the importance of data imputation in retail analytics, emphasizing that filling missing values accurately is crucial for informed decision-making. It outlines various imputation methods such as mean, median, and mode, highlighting their advantages and risks, and suggests advanced techniques for better accuracy. Additionally, it addresses the significance of balancing statistical insights with strategic business decisions in e-commerce and the limitations of relying solely on R squared for evaluating regression models.

Uploaded by

dharavparmar22
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Business Analytics

Q.1

Introduction

In today‟s retail world, data is basically the heartbeat of the business. Every small decision, like when to

launch a discount, how much to charge for delivery or which store is underperforming, depends on

accurate numbers showing up on the sales dashboards. If the numbers are wrong, decisions can quickly

go off track.

In this case, the retail chain is ready to launch a brand new analytics dashboard to track sales. But there‟s

a problem, in the dataset, several entries in the delivery amount column are missing. These gaps exist

because of system glitches or simple human typing errors. If the analyst ignores them, the dashboard

will give a distorted picture of performance. If they are filled carelessly, management could be misled.

This is where imputation comes in, the process of filling in missing data with estimated values so the

dataset remains consistent and useful. The tricky part is figuring out how to fill those blanks in a way

that actually reflects business reality. The most common ways are mean, median and mode imputation.

Each sounds simple, but in practice, the choice can completely change the story the dashboard tells.

Why filling the gaps matters?

Missing values are not just empty boxes on a spreadsheet. They can skew the entire picture. Think of it

this way, if 20% of your delivery amounts are missing, your average revenue per order could look far

lower than reality. If the missing ones are mostly from high value orders, it may appear that delivery
charges are always cheap. On the other hand, filling those blanks with unrealistically high numbers

might trick managers into thinking delivery is a huge profit center.

For retail, this is dangerous. Delivery charges impact:

 Pricing strategy- If delivery seems too expensive, management might cut charges unnecessarily.

 Customer satisfaction- A wrong picture could make it look like customers are overpaying.

 Operations planning- Misjudging delivery revenue could lead to wrong budgeting for logistics.

So, imputation isn‟t just about math. It‟s about protecting decision making, ensuring managers act on

truth, not illusions.

Mean Imputation

The most straightforward fix is to replace missing delivery amounts with the average of all available

ones. Example: If recorded deliveries are Rs. 50, Rs. 60, Rs. 70 and one is missing, we take the average-

Rs. 60 and use it for the blank. Why it works? Simple, quick and keeps the dataset intact. The overall

mean doesn‟t change. Why it‟s risky? Delivery amounts in retail often have outliers. Imagine one rare

case where a delivery was Rs. 5,000 for a bulk order. The average will shoot up and imputing missing

values with that inflated number will create false impressions. So, mean works best if delivery charges

are usually clustered together without extremes. But retail rarely works that neatly.
Median Imputation

Here, the missing value is filled with the „middle‟ number in the sorted list. Example: Say the delivery

amountd are Rs. 40, Rs. 50, Rs. 55, Rs. 60 and Rs. 500. The median is Rs. 55. If one value is missing,

we use Rs. 55. Why it wirks? Resistant to outliers. That one Rs. 500 doesn‟t drag the result into absurb

territory. Why it‟s risky? By replacing blanks with the same middle value again and again, we may lose

the natural spread of data. But in retail, this method is usually safer. Delivery charges often depend on

distance or order value, meaning skewed data is normal. Median gives a realistic typical charge.

Mode Imputation

Here, missing values are replaced with the most common delivery charge. Example: If most customers

pay Rs. 50 and some data is missing, we simply fill in Rs. 50. Why it works? Perfect when delivery

pricing follows slabs like Rs. 50 for local, Rs. 100 for regional. Why it‟s risky? If delivery amounts vary

widely, repeatedly imputing the same number flattens diversity and hides patterns. This works well only

if the company already has fixed delivery has fixed delivery slabs. Otherwise, it oversimplifies.

Worked Example

Let‟s test this with a small dataset:

[50, 60, 55, 500, 45, NaN, 50, NaN, 55, 60]

 Mean Imputation: Average- Rs. 159.5. The missing values become 159.5. But almost all real

charges are around Rs. 50- Rs. 60, so this massively distorts reality.
 Mediain imputation: Middle value- Rs. 55. The missing values become 55. This better reflects

reality because most values cluster around that range.

 Mode imutation: Mode- Rs. 50. Missing values become 50. If the retailer mostly charges Rs.

50, this is valid.

Looking at this, median or mode clearly tells the truer story than mean.

Beyond Simple Imputation

Retail analysts today also use smarter methods:

1. Regression imputation- Predict delivery amount based on order value, distance or weight. For

example, larger orders might automatically carry higher delivery charges.

2. K- Nearest Neighbors- Find orders similar to the missing one and average their delivery

charges.

3. Multiple imputation- Fill in several possible values and combine them reducing bias.

4. Business rule logic- Sometimes, the best method is not statistical but operational. For this retail

chain, sucj rules might capture reality better than pure math.

How to decide: Key considerations

Before filling gaps, the analyst should pause and ask, What‟s the data distribution? If normal, mean

works fine. If skewed, median is safer. Are there outliers? If yes, avoid mean. What‟s the dashboard‟s

purpose? If management wants to study averages, don‟t distort them. If the goal is customer behavior,

reflect the most common patterns. The method must fit both the data and the business context.
Conclusion

Imputation is the bridge that keeps the dataset whole, but the choice of method matters. Mean is quick

but risky if there are outliers. Median is stable, especially when data is skewed. Mode works best when

delivery charges follow fixes slabs. Beyond these, advanced methods and business logic can improve

accuracy even further. At the end of the day, data analysis is not just about crunching numbers. The data

analysis is about telling the truth of the business.

Q. 2 (A)

Introduction

In e- commerce, customer satisfaction is the real currency. Happy customers come back, spend more and

spread the word. After running statistical analysis, Mehta E- Commerce figured out that three thingd

matter most: Product quality, delivery speed and customer support. Now comes the harder part,

deciding how to spend limited money and resources on these areas. Data might point to what‟s

significant, but management has to weight that against business priorities, budget limits and even the

risks of making the wrong call.

Statistic vs. Reality

Here‟s where many companies trip. Just because something is statistically significant doesn‟t mean it‟s

the smartest thing to act on right away. Delivery speed may rank highly, but if most deliveries are

already within 2 days, cutting it to one day could double logistics costs while adding only a tiny bump in
satisfaction. Product quality might take less investment but has a ripple effect, fewer returns, fewer

complaints and more trust. That‟s big for long term loyalty. Customer support won‟t fix a bad product or

late delivery, but it can save a relationship. A single positive call can turn an angry customer into a

repeat buyer. So the company must decide not just what matters statistically, but what matters most for

survival and growth.

Strategic Choices and Trade offs

With a limited budget, Mehta E- commerce has to play a balancing game.

1. Double down on product quality

Every rupee spent here prevents complaints and refund later. It‟s the long term play, a reputation

for “things just work” keeps customers coming back without extra marketing spend.

2. Boost customer support

It‟s cheaper to improve than logistics. Training staff, adding self-help tools or AI chatbots can

create quick, visible improvements. Great for short term wins which customer notices

immediately.

3. Re- evaluate delivery speed

Re- evaluating delivery speed does matter. But trying to match Amazon level speed could drain

profits. Improve consistency instead of chasing extremes. Promise 2 day delivery and actually

deliver 2 day delivery every time. Reliability can sometimes beat raw speed.
Unintended Consequences

Every shiny investment has a shadow side. Over focus on delivery could burn money and pressure

logistics staff into mistakes. Overspending on quality could raise prices too much, scaring off cost

conscious buyers. Expanding customer support without fixing root problems could just mean more

people apologizing for the same recurring issues.

Conclusion

For Mehta E- commerce, the statistics are like a compass pointing north- quality, speed and support

matter. But strategy decides the journey. Invest in product quality to build unshakable trust. Strengthen

customer support for quick visible wins. Tweak delivery speed smartly instead of blowing the budget

chasing impossible targets. Numbers highlight the factors, but leadership decides the priorities. When

Mehta E- commerce balances data with strategy, customers stay happy and the business sails forward-

strong, steady and future ready.

Q. 2 (B)

Introduction

A retail company built a simple linear regression model to forecast monthly sales using advertising

spend as the key driver. The analytics team celebrated a high R squared value, claiming the model was

highly reliable. Management quickly took that as proof that the forecast was trustworthy. But a few team

members questioned whether R squared alone could capture the fully reality of how sales actually work.
After all, business performance depends on far more than ad spend- factors like competition, pricing,

consumer trends and seasonality also shape results. The debate raises an important question: how

effective R squared as the main tool to judge a model‟s success?

The illusion of Accuracy

R squared measures how much of the variation in sales is explained by advertising spend. A value of

0.85 means the model explains 85% of the changes in sales. Impressive on paper, but in business,

correlation doesn‟t equal causation. Just because sales rise when ad spend increases doesn‟t prove ads

caused those sales. Maybe there was a festive season or a viral product launch that drove both at once.

That‟s where the illusion begins. A high R squared only shows how well the model fits past data, not

whether it will predict the future accurately. It‟s like saying someone drives perfectly on a simulator and

assuming they‟ll do just as well on a busy road. The number looks great, but the context matters more.

The Over fitting Trap

Another hidden danger is over fitting- when the model becomes too perfect for its own good. It

memorizes every detail of historical data instead of learning general patterns. So while R squared soars,

real world accuracy collapses. Like, if sales spiked one month due to a celebrity endorsement, the model

might assume that spike will happen every year, leading to unrealistic forecasts. R squared can also rise

when irrelevant variables are added, giving a false sense of improvement. It doesn‟t test whether those
variables genuinely matter or if they just clutter the model. In short, R squared rewards complexity, not

necessarily correctness.

A smarter way to evaluate Models

To truly assess how well a regression model performs, R- squared must be balanced with other

diagnostic tools. Adjusted R squared keeps the score honest by penalizing unnecessary predictors.

Residual analysis checks if errors are randomly scattered or reveal hidden biases. Error metrics such as

Mean Absolute Error show how close predictions are to reality. Cross validation tests how the model

performs on unseen data, ensuring it can adapt to new situations. Together these give a more rounded

view of performance, combining statistical strength with real world reliability.

Conclusion

R squared is a helpful starting point, but it‟s not the final verdict. It tells you how well your model fits

history, not how wisely it forecasts the future. For the retail company, trusting R squared alone could

lead to costly missteps. The real goal is to blend data accuracy with business insight- understanding

what the numbers mean in the marketplace. When R squared is balanced with other diagnostic checks

and strategic judgment, the model becomes more than a mathematical fit; it becomes a dependable

decision making tool.

You might also like