W12 Assignment - Data Collection & Sampling Methods
Student Name -
Student Number -
Introduction
The below is our approach to the questions asked using a hypothetical
scenario: Regional coffee shop chain analyzing customer preferences for a new line of
seasonal beverages.
1. Data Collection
Our objective of collecting this customer preference data is to identify which
new seasonal beverage flavors are most appealing to the target market. This means
that given the launch, we should expect maximum acclaim for the beverage since our
launch would be informed by an understanding of the price point customers are
willing to pay as well as determining the preferred dietary options (e.g. plant-based,
etc) (Backstrom et al., 2025).
The following variables will help us meet our objective by measuring them: 1)
Favorite Flavor Profile which will be a categorical variable representing four
possible flavours namely sweet, spicy, fruity and bitter. 2) Preferred Milk Type which
will also be categorical and represents the nature of the milk namely whole milk,
skim milk, oat milk and almond milk. Last but importantly, we shall track the
Willingness to Pay as a quantitative/ continuous variable to represent the maximum
amount a customer will pay for a specialty drink (Alford & Teater, 2025)..
To collect the data across our variables, we employ the following steps:
1. Define the Goal: You must clearly state what the analysis should achieve for
the business goal. In this example, we could say the goal is to answer the
question, "Which fall beverage should we launch?".
2. Determine the Method: Data is acquired in many ways and not each is
optimal for every exercise. We must select the best way to gather the data. For
customer preferences, an online survey would work if the shop would have an
email list or a loyalty app.
3. Design the Instrument: We then draft as clear and as unbiased questions as
possible including, say, in our example, multiple-choice questions for flavors
and a scale for willingness to pay.
4. Collect the Data: We then distribute the instrument of collection over a set
period (e.g., two weeks). It is advisable to offer a small incentive such as a
discount code where applicable in order to encourage participation.
5. Clean the Data: Finally, we review the collected responses to remove
incomplete surveys or duplicate entries before analysis begins.
2. Population and Sample
A population represents the entire group that one is analysing for their set
objectives. For our earlier example, the coffee shop wants to draw conclusions about
the logic of launching a new product. In this scenario, the population is all current
and potential customers of the coffee shop.
On the other hand, a sample is a portion of the set population selected through
established techniques called sampling methods. This is the focus of our analysis
because it is often impossible and too expensive to survey every customer in the
greater population. In our earlier scenario, the sample is the subset of customers who
receive and complete the survey.
Through the correct application of chosen established sampling techniques, we
should ideally end up with a sample that accurately reflects the diversity of the entire
population in order for it to be useful. This is called a representative sample. Having a
representative sample allows an analyst to offer the business generalizable analyses
using which the business confidently makes decisions that apply to the broader
customer base on a preference level. On the other hand, if unrepresentative, it means
that the sample is biased and only bases its conclusions on a smaller cluster within the
population hence can lead to poor business decisions. A simple example can be
conducting the survey too early in the morning which may cause the sample to only
represent early-morning risers and commuters and thus misses the preferences of
afternoon students or evening study groups (Creswell & Inoue, 2025).
3. Sampling Methods
As discussed earlier, different statistical methods can be used to select the
customers who will make up the sample. The following are sampling methods and
how they apply to the customer preference data:
Simple Random Sampling is a technique used which ensures that every
customer in the population has an equal chance of being selected. For example, if a
company has 10,000 email subscribers, the analyst would assign a number to all the
email subscribers and use a random number generator to select, say, a sample of 500
people to receive the survey (Celestin et al., 2025).
Stratified Sampling is a method used to divide the population into distinct
subgroups called strata based on shared characteristics. We then perform random
sampling on each strata proportionally for analysis. For example, a coffee shop would
divide its customers by location such as urban versus suburban stores. If 60% of total
sales come from urban stores, they ensure 60% of the sample consists of urban
customers to capture regional preference differences (Cohen, 2025).
Cluster Sampling is an approach whereby the population is divided into pre-
existing, geographically or naturally occurring groups called clusters. For example,
region A, B, C, etc. A few clusters are chosen at random and every individual in those
chosen clusters is surveyed. For example, the company can pick 3 out of 50 store
locations randomly and for a set period, every customer who enters/ purchases from
those 3 specific stores is asked to complete a preference card (Chand, 2025).
Convenience Sampling also known as Non-Probability sampling is a method
whereby data is collected from individuals who are simply the easiest to reach by the
analyst/ business. For example, a barista can ask the first x number of people who
walk into the beverage store on a specific day of the week what their favorite drink is.
Notably, convenience sampling is easy and cheap but highly susceptible to biases
from the collector, and other sources (Conway, 2025).
4. Graphical Representation
Chosen Method: Bar Graph
A bar graph is the most appropriate visual tool for representing customer
preference data. In this scenario, the primary data points are categorical eg favoritte
flavor and preferred milk type. This would enable us to graph in the bar chart the
distinct categories e.g. “Oat Milk”, “Almond Milk” and “Whole Milk” on one axis
and the frequency which represents the percentage/ number of customers who prefer
them each on the other axis.
As such, we would yield a clear visual comparison in our groups that would
give a quick insight into the most and least popular options. A pie chart can become
easily cluttered and hard to read where there are more than 3 or 4 categories. A bar
graph can cleanly handle multiple customer preference options while maintaining ease
of readability for intended/ targeted business stakeholders (Creswell & Inoue, 2025).
References
Alford, S., & Teater, B. (2025). Quantitative research. In Handbook of research
methods in social work (pp. 156-171). Edward Elgar Publishing.
Backstrom, L. J., Callaghan, C. T., Worthington, H., Fuller, R. A., & Johnston, A.
(2025). Estimating sampling biases in citizen science datasets. Ibis, 167(1), 73-87.
Celestin, M., Gidisu, J. A., Boakye, M. M., & Kumar, M. S. (2025). Statistical
Sampling Methods as Tools to Enhance the Accuracy and Reliability of Audits in
Government Budgets and Expenditure. Kings and Queens Journal of Mathematical
Modeling and Applied Analytics, 1(1), 47-57.
Chand, S. P. (2025). Methods of data collection in qualitative research: Interviews,
focus groups, observations, and document analysis. Advances in Educational
Research and Evaluation, 6(1), 303-317.
Cohen, M. P. (2025). Stratified sampling. In International encyclopedia of statistical
science (pp. 2704-2708). Berlin, Heidelberg: Springer Berlin Heidelberg.
Conway, F. (2025). Sampling: An introduction for social scientists. Routledge.
Creswell, J. W., & Inoue, M. (2025). A process for conducting mixed methods data
analysis. Journal of General and Family Medicine, 26(1), 4-11.