NMIMS Centre for Distance and Online Education (NCDOE)
Course: Quantitative Methods - I
Internal Assignment - June 2026 Examination
Q1. Z-Score and Standard Error Analysis – Chinese Food Preference
Assessment Parameter Weightage
Understanding and usage of the formula 20%
Procedure / Steps 60%
Correct Answer & Interpretation 20%
Introduction
Statistical hypothesis testing is a fundamental tool that helps institutions make objective decisions
based on data rather than intuition. In this case, the university has collected survey data from 100
new students and found that 40 of them prefer Chinese food. The historical preference rate stood at
30%. To determine whether this observed increase is statistically meaningful or simply due to chance
variation in sampling, the university should apply a one-sample Z-test for proportions. This test relies
on two key calculations: the standard error of the proportion and the Z-score, which together tell us
how far the sample result is from what was historically expected.
Concepts and Application
Step 1: Define Hypotheses
Before doing any calculation, the university must clearly state what it is testing. The null hypothesis
(H0) assumes no real change — that the true proportion of students preferring Chinese food is still
30%, same as in prior years. The alternative hypothesis (Ha) states that the preference has shifted —
the proportion is now different from 30%. This makes it a two-tailed test, since we are looking for
any significant difference, not just an increase.
H0: p = 0.30
Ha: p ≠ 0.30
Step 2: Identify the Sample Statistics
From the survey data, the following values are known:
• Sample size (n) = 100 students
• Number preferring Chinese food = 40
• Sample proportion (p-hat) = 40/100 = 0.40
• Historical/population proportion (p0) = 0.30
• Significance level: alpha = 0.05 (standard for academic research)
Step 3: Calculate the Standard Error (SE)
The standard error tells us how much the sample proportion is expected to vary just due to random
sampling. It is calculated using the assumed population proportion under H0.
Formula: SE = sqrt [ p0 × (1 - p0) / n ]
SE = sqrt [ 0.30 × 0.70 / 100 ]
SE = sqrt [ 0.21 / 100 ]
SE = sqrt [ 0.0021 ]
SE = 0.0458 (approximately)
Step 4: Calculate the Z-Score
The Z-score measures how many standard errors the observed sample proportion is away from the
hypothesized population proportion. A higher Z-score means the result is more unusual under H0.
Formula: Z = (p-hat - p0) / SE
Z = (0.40 - 0.30) / 0.0458
Z = 0.10 / 0.0458
Z = 2.183 (approximately)
Step 5: Compare with Critical Value
For a two-tailed test at a 5% significance level (alpha = 0.05), the critical Z-value from the standard
normal table is ±1.96. Since our calculated Z-score of 2.183 is greater than 1.96, it falls in the
rejection region.
|Z calculated| = 2.183 > 1.96 (critical value)
This means we reject the null hypothesis. The observed proportion of 40% is statistically
significantly different from the historical 30% preference rate.
Why Z-Test is the Right Choice
The Z-test for proportions is appropriate here because: the sample size is sufficiently large (n = 100),
both np0 = 30 and n(1-p0) = 70 are greater than 5, which satisfies the conditions for using the normal
approximation. Also, the data collected is binary in nature — each student either prefers Chinese
food or does not. These conditions make the Z-test a statistically sound and defensible choice for this
analysis.
Conclusion
Based on the Z-test analysis, the university can conclude with 95% confidence that the proportion of
new students preferring Chinese food (40%) is statistically different from the historic rate of 30%.
With a Z-score of 2.183, which exceeds the critical value of 1.96, the null hypothesis is rejected. This
is not just a random fluctuation in the data — it reflects a meaningful shift in student food
preferences. The student affairs office should take this finding seriously when making decisions
about future dining options and menu planning. Increasing the availability of Chinese food options in
the cafeteria would be a data-supported response to this change in student preference.
Q2(A). Marginal, Joint, and Conditional Probabilities for Customer
Segmentation
Assessment Parameter Weightage
Introduction 20%
Concepts and Application related to the question 60%
Conclusion 20%
Introduction
In today's data-driven retail environment, understanding customer behavior is not just helpful — it is
necessary for effective campaign targeting. A large retail chain using a contingency table to study
shopping habits based on gender and purchase frequency already has a solid foundation. However,
the real value of this data comes from correctly interpreting the three types of probabilities it offers:
marginal, joint, and conditional. Each type tells a different story about customer behavior, and
choosing the right one to guide marketing decisions can significantly impact both cost efficiency and
campaign accuracy. The internal debate around whether gender and purchasing frequency are
independent or related variables is also critical to resolve before designing cross-selling strategies.
Concepts and Application
Understanding the Three Probability Types
Marginal probability refers to the overall likelihood of a single event — for example, the probability
that a randomly selected customer is female, or the probability that a customer makes more than
three purchases per week, regardless of any other factor. In retail, marginal probabilities are useful
for understanding the composition of the customer base at a broad level. However, they are limited
because they do not capture any relationship between two variables. Relying solely on marginal
probabilities would mean treating all female customers the same, or all high-frequency shoppers the
same, without considering how these characteristics interact.
Joint probability captures the likelihood that two events occur together. For example: what
proportion of all customers are female AND make three or more purchases per week? This is more
useful for identifying specific customer profiles. A retail chain can use joint probability to estimate
the size of a target segment — say, male customers who purchase infrequently — and allocate
budget accordingly. However, joint probability alone does not tell you anything about causation or
the relationship between the variables.
Conditional probability is arguably the most powerful tool for marketing segmentation. It answers
questions like: given that a customer is female, what is the probability that she shops more than three
times a week? This directly enables targeted strategies. If the conditional probability of frequent
purchasing differs significantly between genders, that is a strong signal to design gender-specific
campaigns. Conditional probabilities transform raw data into actionable insights — which is exactly
what the retail chain needs for cross-selling.
Independence vs. Dependence Between Gender and Purchasing Behavior
The internal disagreement about whether gender and purchase frequency are independent is a crucial
one. Two events are independent if knowing one tells you nothing about the other — that is, if
P(Frequent Buyer | Female) = P(Frequent Buyer). If this equality holds in the data, then gender-
based segmentation adds no value, and the company's marketing budget would be better spent on
other segmentation criteria like age, location, or product category.
In real-life retail, complete independence between gender and purchase frequency is quite rare.
Research consistently shows that shopping patterns vary across gender lines — for instance,
frequency of grocery purchases, clothing, or household items often shows clear gender-related
patterns. If the contingency table data reveals that conditional probabilities differ from marginal
probabilities, that is statistical evidence of dependence, and it justifies segmenting campaigns by
gender.
The retail chain can test this formally using the chi-square test of independence on the contingency
table. If the test result shows a significant association, the company should build gender-dependent
models for its marketing campaigns rather than assuming independence.
Recommended Approach for the Retail Chain
The retail chain should adopt conditional probabilities as the primary tool for customer segmentation.
Marginal probabilities should only be used to understand the overall customer mix, while joint
probabilities help in estimating segment sizes and planning resource allocation. Conditional
probabilities, combined with an empirical test for independence, will provide the most accurate and
defensible foundation for cross-selling strategies. Assuming independence without testing it would
risk misallocating the marketing budget and designing campaigns that do not resonate with actual
customer behavior.
Conclusion
The choice between marginal, joint, and conditional probabilities is not just a technical question — it
has real financial and strategic consequences for the retail chain. Marginal probabilities provide a
broad picture but lack the nuance needed for targeted campaigns. Joint probabilities help identify
specific customer groups. Conditional probabilities, however, are the most valuable for campaign
segmentation because they reveal how behavior changes based on a characteristic like gender. The
chain should empirically test whether gender and purchasing frequency are independent before
finalizing its strategy. If dependence is found — which is likely in real retail data — conditional
probability-based segmentation will lead to higher campaign accuracy, better budget utilization, and
improved cross-selling outcomes.
Q2(B). Normal Distribution – Probability of Monthly Sales Exceeding $180,000
Assessment Parameter Weightage
Understanding and usage of the formula 20%
Procedure / Steps 60%
Correct Answer & Interpretation 20%
Introduction
The normal distribution is one of the most widely used probability distributions in business analytics,
particularly for modeling continuous variables like sales, revenue, or delivery times. In this case, a
national retail chain wants to assess the likelihood that a randomly selected store achieves monthly
sales exceeding $180,000, given that the average sales across 250 stores follow a normal distribution
with a mean of $150,000 and a standard deviation of $20,000. This analysis will help management
evaluate whether their premium store classification threshold of $180,000 is a realistic and fair target.
Concepts and Application
Step 1: Define the Distribution Parameters
• Mean (mu) = $150,000
• Standard Deviation (sigma) = $20,000
• Threshold value (X) = $180,000
• Number of stores = 250
Step 2: Calculate the Z-Score
To find the probability that a store earns more than $180,000, we first convert the threshold value
into a standardized Z-score. The Z-score tells us how many standard deviations above (or below) the
mean the value of $180,000 lies.
Formula: Z = (X - mu) / sigma
Z = (180,000 - 150,000) / 20,000
Z = 30,000 / 20,000
Z = 1.50
Step 3: Find the Probability from Z-Table
From the standard normal distribution table, the cumulative probability for Z = 1.50 is:
P(Z <= 1.50) = 0.9332
This gives us the probability that a store earns less than or equal to $180,000. To find the probability
of earning more than $180,000, we subtract from 1:
P(X > 180,000) = 1 - P(Z <= 1.50)
P(X > 180,000) = 1 - 0.9332
P(X > 180,000) = 0.0668 or approximately 6.68%
Step 4: Estimate the Number of Stores Likely to Qualify
Out of 250 stores, the expected number achieving premium classification:
Expected stores = 250 × 0.0668 = 16.7 stores (approximately 17
stores)
2. Managerial Interpretation
The probability that any randomly selected store earns more than $180,000 in a month is
approximately 6.68%. Translated to the chain's 250 stores, this means only around 17 stores would
realistically meet the premium classification threshold in any given month. From a managerial
standpoint, this is an important finding. The company's premium classification is not based on
average or even above-average performance — it targets the top 6-7% of stores. This means the
threshold is set at a level that most stores would genuinely struggle to reach, particularly those in
lower-traffic regions, smaller cities, or stores that have been established for a shorter period.
3. Comment on Whether the Threshold is Too Strict or Reasonable
The premium threshold of $180,000 sits exactly 1.5 standard deviations above the mean. This places
it well into the upper tail of the distribution, which means it is intentionally designed to recognize
only the top-performing stores. Whether this is 'too strict' or 'reasonable' depends on the purpose of
the classification.
If the premium label is meant to reward a small elite group of high performers for benchmarking
purposes or special incentives, then 6.68% is a defensible cutoff. However, if the company expects a
larger portion of its stores to qualify for premium status and uses it as a general performance
standard, then the threshold is likely too strict. In practice, a threshold closer to one standard
deviation above the mean — around $170,000 — would capture approximately 15.87% of stores,
making the target more achievable while still being selective. The management may want to revisit
this threshold in light of regional sales variability, store size differences, and market conditions, to
ensure the premium classification remains a motivating — rather than discouraging — performance
benchmark.
Conclusion
Using the normal distribution framework, the probability that a randomly selected store earns more
than $180,000 per month is approximately 6.68%, meaning only about 17 out of 250 stores would
qualify as premium stores in a typical month. The Z-score of 1.50 confirms that the $180,000
threshold is placed well above the average, making it an aspirational rather than a broadly achievable
target. While this may be appropriate for recognizing exceptional performance, management should
consider whether such a strict threshold aligns with the broader business goals of the premium
classification program. A data-driven reconsideration of this cutoff — perhaps adjusted for regional
and store-level differences — would make the program fairer and more strategically effective across
the chain.
End of Assignment