0% found this document useful (0 votes)
4 views9 pages

Sample Design Unit 2

The document discusses sampling design, contrasting census inquiries with sample surveys, and outlines the steps and considerations for creating a sample design. It details types of sampling methods, including probability and non-probability sampling, and highlights the importance of selecting a representative sample to minimize bias and errors. Additionally, it covers various complex sampling techniques such as stratified, cluster, and multi-stage sampling, emphasizing their advantages and disadvantages.

Uploaded by

mohit1022005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views9 pages

Sample Design Unit 2

The document discusses sampling design, contrasting census inquiries with sample surveys, and outlines the steps and considerations for creating a sample design. It details types of sampling methods, including probability and non-probability sampling, and highlights the importance of selecting a representative sample to minimize bias and errors. Additionally, it covers various complex sampling techniques such as stratified, cluster, and multi-stage sampling, emphasizing their advantages and disadvantages.

Uploaded by

mohit1022005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

SAMPLE DESIGN (Chapter 4 — Sampling

Design)
Census vs. Sample Survey
Universe / Population: All items in any field of inquiry put together.

Census Inquiry: Complete enumeration of every single item in the population.

• Theoretically gives highest accuracy (no chance element).


• But in practice: even a slight bias gets larger and larger as observations increase, and
there's no way to check the extent of bias without a re-survey.
• Extremely time-consuming, expensive, and energy-intensive.
• Practically impossible for large populations.
• Only governments can afford it — and even they do it rarely (e.g., Population Census
once every decade).
• When the universe is small, census is fine. But for large populations, sampling is the
practical answer.

Sample Survey: Studying a representative part (sample) of the population instead of the
whole.

• Population size = N. Sample size = n (where n < N).


• The selected n units = the sample.
• The process of selecting them = sampling technique.
• The survey conducted using this sample = sample survey.

Key condition: The selected respondents/units must be as representative of the total


population as possible — a miniature cross-section of the whole.

What is a Sample Design?


A sample design is a definite plan for obtaining a sample from a given population.

It refers to:

• The technique or procedure the researcher uses to select items.


• The number of items to include (sample size).
• It is determined before data collection begins.

Steps in Sample Design (7 key considerations)


1. Type of Universe

• Clearly define the universe/population to be studied.


• Finite universe: Number of items is certain/countable (e.g., workers in a factory,
population of a city).
• Infinite universe: Number of items is infinite/cannot be known (e.g., stars in the sky,
all possible throws of a die, listeners of a radio programme).

2. Sampling Unit

• Decide what the "unit" of selection will be.


• Can be: geographical (state, district, village), construction (house, flat), social (family,
club, school), or individual.
• Researcher decides one or more units to select.

3. Source List (Sampling Frame)

• The list from which the sample is drawn.


• Contains names of ALL items in the universe (for finite universes).
• Must be: comprehensive, correct, reliable, appropriate, and as representative of the
population as possible.
• If no source list is available, the researcher must prepare one.

4. Size of Sample

• How many items to select — a major decision.


• Should be optimum — not too large, not too small.
• An optimum sample fulfills: efficiency, representativeness, reliability, and
flexibility.
• Factors affecting size: desired precision, confidence level, population variance (larger
variance = larger sample), size of population, parameters of interest, and cost/budget.

5. Parameters of Interest

• What specific population characteristics do you want to estimate?


• E.g., proportion of people with a characteristic, averages, sub-group details.
• This directly influences the sample design you choose.

6. Budgetary Constraint

• Cost considerations have a major impact on both the size and type of sample.
• This practical factor can even force the researcher toward non-probability sampling.

7. Sampling Procedure

• The actual technique used to select items — this IS the sample design.
• Choose the design that, for a given sample size and cost, has the smallest sampling
error.
Criteria for Selecting a Sampling Procedure
Two types of costs in sampling:

1. Cost of collecting the data.


2. Cost of an incorrect inference from the data.

Two types of errors in sampling:

1. Systematic Bias — errors from the sampling procedure itself; CANNOT be reduced
by increasing sample size. Can only be corrected by fixing the cause.
2. Sampling Error — random variations of sample estimates around true population
parameters. These are compensatory (cancel out), expected value = zero. Decreases as
sample size increases.

Causes of Systematic Bias:

1. Inappropriate sampling frame — if the frame is a biased representation of the


universe.
2. Defective measuring device — a biased questionnaire or defective physical
instrument introduces constant error.
3. Non-respondents — if those you couldn't contact/get responses from are
systematically different from those you did reach, bias enters.
4. Indeterminancy principle — people behave differently when they know they're
being observed. E.g., factory workers slow down when they know a work study is
being done (because quotas will be set based on it).
5. Natural bias in reporting — people naturally underreport income when asked for tax
purposes (downward bias) but overstate it to show social status (upward bias). In
psychological surveys, people give what they think is the "correct" answer.

Sampling error is measurable; its measurement = precision of the sampling plan.


Increasing sample size improves precision, BUT also increases cost and possibly systematic
bias. The most effective way to improve precision = choose a better sampling design that
has smaller sampling error at a given cost.

Characteristics of a Good Sample Design


(a) Must result in a truly representative sample. (b) Must result in a small sampling error.
(c) Must be viable within available funds. (d) Must help control systematic bias in a better
way. (e) Results should be generalizable to the universe with a reasonable level of
confidence.

Types of Sample Designs


Two main bases for classification:

• Representation basis: Probability or Non-probability sampling.


• Element selection basis: Unrestricted or Restricted sampling.

Probability Sampling Non-Probability Sampling


Haphazard/Convenience
Unrestricted Simple Random Sampling
Sampling
Complex Random Sampling (cluster, Purposive Sampling (quota,
Restricted
systematic, stratified, etc.) judgement)

Non-Probability Sampling
Also called: deliberate sampling, purposive sampling, judgement sampling.

• Researcher deliberately selects items — their choice is supreme.


• The sample is selected based on the belief that the selected units are representative.
• No basis for calculating the probability of any item being included.

Advantage: Saves time and money; suitable for small inquiries.

Disadvantage:

• Personal element has a great chance of entering; researcher may (consciously or


unconsciously) select a biased sample.
• Sampling error CANNOT be estimated.
• Bias — great or small — is always present.
• Not suitable for large, important inquiries.

Quota sampling is a type of non-probability sampling. Interviewers are given quotas to fill
from different strata, but the actual selection of items is left to the interviewer's discretion.
Convenient and cheap, but not random samples; inferences are not amenable to formal
statistical treatment.

Probability Sampling (Random Sampling)


Also called: random sampling or chance sampling.

• Every item in the universe has an equal chance of being included in the sample.
• Selection is by mechanical/blind chance, not by deliberate choice.
• Based on the Law of Statistical Regularity: a randomly chosen sample will, on
average, have the same composition and characteristics as the universe.
• Results can be measured in terms of probability — errors of estimation can be
calculated.
• This is why probability sampling is considered the best technique for selecting a
representative sample.

Simple random sample from a finite population: Each of the NCn possible samples of size
n has the same probability (1/NCn) of being selected.

Example: Population N=6 (elements a,b,c,d,e,f), sample size n=3. Total possible samples =
6C3 = 20. Each has probability 1/20 of being selected.

How to select a random sample:

1. Lottery method: Write each element's name/number on a slip, mix thoroughly, draw
required number blindly.
2. Random number tables: Tables prepared by statisticians (Tippett, Yates, Fisher).
Tippett gave 10,400 four-figure numbers (41,600 digits from census reports combined
into fours).
o Example: From population of 5000 units numbered 3001 to 8000, selecting 10
units. Read Tippett's table, pick numbers between 3001 and 8000: 6641, 3992,
7979, 5911, 3170, 5624, 4167, 7203, 5356, 7483.

Sampling without replacement (usual): Once an item is selected, it's not put back. This is
the standard approach. Sampling with replacement (less common): Selected item is
returned before the next draw — same item could appear twice.

Random sample from infinite population: Probability of getting a particular result is the
same at each draw, and successive draws are independent.

Complex Random Sampling Designs


(i) Systematic Sampling

• Select every i-th item on a list.


• Introduce randomness by randomly selecting the first item, then selecting every nth
item thereafter.
• Example: For a 4% sample → randomly pick first item from first 25 → then every
25th item automatically.
• Only the FIRST unit is randomly selected; the rest are at fixed intervals.

Advantages:

• Evenly spread over the entire population.


• Easier and less costly.
• Works well for large populations.
• Considered equivalent to random sampling IF the population list is in random order.

Disadvantages:
• Hidden periodicity problem: If population has a hidden periodic pattern coinciding
with sampling interval, results are badly biased. Example: If every 25th item in a
production process is defective, a 4% systematic sample will give either ALL
defectives or NO defectives depending on the starting point.
• Not truly random in strict sense.
• Used when lists of population are available and lengthy.

(ii) Stratified Sampling

Used when the population is not homogeneous — divides it into sub-groups first.

Process:

1. Divide population into strata (sub-populations) — each stratum is internally


homogeneous; strata are heterogeneous from each other.
2. Select a sample from each stratum (usually by simple random sampling).
3. Combine to get the overall sample.

Result: More precise estimates for each stratum → better estimate for the whole.

Three key questions in stratified sampling:

Q1: How to form strata?

• Based on common characteristics.


• Elements WITHIN each stratum → most homogeneous.
• Elements BETWEEN strata → most heterogeneous.
• Based on past experience and researcher's judgement.
• Pilot study may help determine efficient stratification.

Q2: How to select items from each stratum?

• Usually: simple random sampling within each stratum.


• Systematic sampling may be used if more appropriate.

Q3: How many items from each stratum? (Allocation)

Proportional Allocation: Sample size from each stratum is proportional to that stratum's size
in the population.

• Formula: n_i = n × (N_i / N)


• Example: N = 8000, divided into N1=4000, N2=2400, N3=1600. Total sample n=30.
o n1 = 30 × (4000/8000) = 15
o n2 = 30 × (2400/8000) = 9
o n3 = 30 × (1600/8000) = 6
• Best when: cost per item is equal across strata, within-stratum variance is equal,
purpose is to estimate the total population.
Disproportionate (Optimum) Allocation: Used when strata differ in size AND variability.
Take larger samples from more variable strata.

• Formula: n_i = (n × N_i × σ_i) / (N_1σ_1 + N_2σ_2 + ... + N_kσ_k)

Example of Optimum Allocation: N1=5000 (σ1=15), N2=2000 (σ2=18), N3=3000 (σ3=5).


Total sample n=84.

• n1 = 84×5000×15 / (5000×15 + 2000×18 + 3000×5) = 6,300,000 / 126,000 = 50


• n2 = 84×2000×18 / 126,000 = 3,024,000 / 126,000 = 24
• n3 = 84×3000×5 / 126,000 = 1,260,000 / 126,000 = 10

When strata ALSO differ in sampling cost: Use cost-optimal formula: n_i = (n ×
N_i×σ_i/√C_i) / Σ(N_j×σ_j/√C_j)

Cross-stratification: Stratifying on MORE than one characteristic simultaneously (e.g.,


class, sex, and college for a student survey). Increases reliability; widely used in opinion
surveys.

Note: Stratified sampling = purposive (stratification) + random (within-strata selection) =


mixed sampling / stratified random sampling.

(iii) Cluster Sampling

Used when the total area of interest is very large — divide into clusters instead of listing
individuals.

Process:

1. Divide total population into non-overlapping clusters (smaller sub-divisions).


2. Randomly select a number of clusters.
3. Include ALL units in the selected clusters in the sample (or take a sample within each
cluster).

Example: 20,000 machine parts in 400 cases of 50 each. Select clusters (cases) randomly →
examine all machine parts in selected cases.

Advantage: Reduces cost by concentrating surveys geographically/logistically.

Disadvantage: Less precise than simple random sampling. 'n' observations within a cluster
contain less information than 'n' randomly drawn observations (because items in a cluster
may be similar to each other). BUT more reliable per unit cost.

(iv) Area Sampling


• A special case of cluster sampling where the clusters are geographic subdivisions
(areas).
• The primary sampling unit is a geographic cluster.
• Same advantages and disadvantages as cluster sampling.

(v) Multi-Stage Sampling

A further development of cluster sampling — sampling done in multiple stages.

Example: Studying working efficiency of nationalised banks in India:

• Stage 1: Randomly select states.


• Stage 2: From selected states, randomly select districts.
• Stage 3: From selected districts, randomly select towns.
• Stage 4: From selected towns, randomly sample banks.

If random at all stages → multi-stage random sampling design.

Advantages:

1. Easier to administer — sampling frame developed in partial units.


2. Can sample a large number of units for given cost due to sequential clustering.

Used for: Big inquiries extending over a large geographical area (e.g., entire country).

(vi) Sampling with Probability Proportional to Size (PPS)

Used when cluster units do NOT have the same number of elements.

• Probability of selecting each cluster is proportional to its size (larger clusters =


higher chance of selection).
• List number of elements per cluster, calculate cumulative totals.
• Sampling interval = Total elements / Sample size.
• Select systematically from cumulative totals.

Illustration: 15 cities with departmental stores: 35, 17, 10, 32, 70, 28, 26, 19, 26, 66, 37, 44,
33, 29, 28. Total = 500 stores. Sample size = 10. Sampling interval = 500/10 = 50. Start = 10.
Selected numbers: 10, 60, 110, 160, 210, 260, 310, 360, 410, 460. Matching to cumulative
totals → select: 2 stores from city 5, 1 each from cities 1, 3, 7, 9, 10, 11, 12, 14. Results
equivalent to simple random sample; less cumbersome and less expensive.

(vii) Sequential Sampling


• The ultimate sample size is NOT fixed in advance.
• Determined by mathematical decision rules as the survey progresses.
• Used mainly in statistical quality control (acceptance sampling).

Types (based on number of samples before decision):

• Single sampling: Decision based on ONE sample.


• Double sampling: Decision based on TWO samples.
• Multiple sampling: More than two samples, but number is fixed in advance.
• Sequential sampling: More than two samples, number is NEITHER certain NOR
decided in advance. Can keep sampling as long as needed.

Conclusion on Sampling
• Normally prefer simple random sampling — eliminates bias, allows estimation of
sampling error.
• Use purposive sampling when the universe is small and a known characteristic is to
be studied intensively.
• Use other designs when random sampling is not possible or when other designs are
easier, cheaper, or more informative.
• Multiple sampling methods can be combined in a single study (mixed sampling).

You might also like