Case Study: Sampling Strategy Dilemma at YUVA Connect
YUVA Connect, an initiative focused on “Linking Young Minds to Digital Futures,” has been
actively working to enhance digital literacy among youth across rural and semi-urban regions
in India. After successfully implementing pilot programs in states like Telangana and
Karnataka, the organization is now preparing to scale its operations nationally. To support
funding proposals and demonstrate measurable impact, the leadership team has decided to
undertake a comprehensive research study to evaluate how far their training programs have
improved employability, digital adoption, and income levels among participants aged 18–30.
However, as the research team begins designing the study, a critical challenge emerges—
selecting an appropriate sampling strategy. The total population consists of nearly 1.2 lakh
trained individuals spread across six states, making a complete census impractical due to time
and budget constraints. While the research head advocates for probability sampling to ensure
representativeness and generalizability of results, field coordinators highlight practical
difficulties such as incomplete databases, especially in rural areas where records are
fragmented or outdated. This creates a dilemma between choosing a method that is
statistically rigorous versus one that is operationally feasible.
Further complicating the decision, multiple sampling techniques are proposed within the
team. Some members suggest simple random sampling, considering it unbiased, but it
requires a complete sampling frame, which is not available. Others recommend stratified
sampling by categorizing respondents based on geography—rural, semi-urban, and urban—to
ensure better representation. A few argue for cluster sampling, selecting specific districts or
villages to reduce cost and logistical complexity. In areas where data is scarce, snowball
sampling is also discussed, relying on participant referrals to reach more respondents. Each
approach presents its own trade-offs between accuracy, cost, and ease of execution.
Another concern arises regarding the appropriate sample size. The statistician proposes a
sample of around 1,200 respondents to achieve reliable and statistically significant results.
However, operational managers argue that given the limited budget and tight three-month
timeline, a smaller sample of about 600 respondents may be more realistic. The team must
therefore balance the need for precision with practical limitations.
The issue of constructing a reliable sampling frame further intensifies the challenge. While
urban training centers maintain digital records, rural areas lack structured databases,
increasing the risk of excluding certain groups. This raises concerns about sampling bias,
such as over-representation of easily accessible respondents and under-representation of
marginalized youth. The team also anticipates potential interviewer bias during field surveys.
As the deadline approaches, Dr. Suresh Kamarapu must finalize a sampling design that
ensures adequate representation across regions, minimizes bias, and delivers valid and
reliable findings within the given constraints. The decision he makes will not only influence
the credibility of the research but also impact the organization’s ability to secure funding and
expand its mission.
Discussion Questions
1. In the given context, how should the research team choose between probability and
non-probability sampling methods? Justify your answer.
2. Which sampling technique or combination of techniques would be most suitable for
this study, and why?
3. How should the team determine an optimal sample size considering both statistical
reliability and practical constraints?
4. What steps can be taken to minimize sampling bias and improve the
representativeness of the study?