0% found this document useful (0 votes)
2 views12 pages

Sampling Theory Notes Rewritten

The document provides an overview of the fundamentals of sampling theory, including key concepts such as population, sample, sampling errors, and survey design principles. It discusses the importance of sampling in data collection, the differences between census and sample surveys, and the application of the Central Limit Theorem in statistical analysis. Additionally, it covers standard error and its implications for estimating population parameters.

Uploaded by

sableom10
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views12 pages

Sampling Theory Notes Rewritten

The document provides an overview of the fundamentals of sampling theory, including key concepts such as population, sample, sampling errors, and survey design principles. It discusses the importance of sampling in data collection, the differences between census and sample surveys, and the application of the Central Limit Theorem in statistical analysis. Additionally, it covers standard error and its implications for estimating population parameters.

Uploaded by

sableom10
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Sampling Theory — Unit 1 Notes

UNIT 1
Fundamentals of Sampling Theory
Personal Study Notes — [Link]. Actuarial Science, Semester III
Self-prepared revision notes

Page 1
Sampling Theory — Unit 1 Notes

Table of Contents
TOC \h \o "1-2"

Page 2
Sampling Theory — Unit 1 Notes

1.1 Sampling Concepts — Population and Sample


Before getting into the actual techniques of data analysis, it is necessary to be clear on
two basic ideas that every statistical exercise relies on: the Population and the Sample.
Practically all statistical work starts by deciding exactly who or what is being studied (the
population), and then how many of them can realistically be observed (the sample).

What is a Population?
Population — the complete set of individuals, items, or occurrences that we wish to
study and about which conclusions are to be drawn. It is also referred to as the
Universe.
Example: Every learner enrolled in actuarial science programmes across Maharashtra
in the academic year 2025-26.
Example: All motor insurance claims registered with an insurer over the course of a
calendar year.

What is a Sample?
Sample — a smaller part drawn out of the population for the purpose of actual
observation. Conclusions drawn from the sample are then extended, with appropriate
caution, to the population as a whole.
Example: A randomly chosen group of 300 actuarial students drawn from colleges
across Maharashtra.
Example: 750 motor insurance claim files picked out for a quality audit.
Key Terminology at a Glance:

Term Meaning
A numerical fact about the population (e.g., population mean
Parameter
µ)
A numerical fact computed from the sample (e.g., sample
Statistic
mean x̄ )
Any single element eligible to be picked into the sample
Sampling Unit
(e.g., one claim file)
The full listing of every unit belonging to the population (e.g.,
Sampling Frame
a claims register)
Observing every single member of the population, with no
Census
sampling involved

Page 3
Sampling Theory — Unit 1 Notes

Why Sample Rather Than Study the Whole Population?


Examining an entire population is, in most real situations, either impossible or simply
impractical because of the money, time, or physical limitations involved. Drawing a
sample lets us reach dependable conclusions while using a fraction of the resources,
without giving up scientific rigour.
• Cost: Covering every unit is costly; a sample cuts down expenses sharply.
• Time: A full population study can take years to complete, whereas a sample
produces results much sooner.
• Destructive Testing: Some forms of testing destroy the item tested (e.g.,
checking the bursting strength of every pipe in a batch).
• Infinite Populations: Certain populations are conceptually unlimited in size (e.g.,
every possible outcome of repeatedly tossing a coin).
• Accuracy: A carefully designed sample, measured with care, can sometimes
outperform a carelessly conducted census in terms of accuracy.

1.2 Census vs Sample Survey


Two broad routes exist for gathering data: a Census (where every unit is enumerated)
and a Sample Survey (where only part of the population is enumerated). Each has a
role to play, and the choice between them depends on the resources at hand, the
accuracy that is needed, and the nature of the population under study.

Census Sample Survey


Covers every single unit in the population Covers only a chosen portion of the population
Costlier and slower to carry out More economical and quicker
Works well for populations that are small Suited to populations that are large or
and finite unbounded
Free of sampling error since nothing is left Carries sampling error, but often delivers
out timelier results
Non-sampling errors tend to be high given Quality control is comparatively easier to
the scale maintain
Example: NFHS (National Family Health
Example: The decennial Census of India
Survey) sample rounds

Types of Sample Surveys


Sample surveys can be grouped according to the method used to pick the sample:

Type Description
Probability Sampling Each unit has a known and non-zero chance of being

Page 4
Sampling Theory — Unit 1 Notes

chosen, so the results carry statistical validity. Examples:


Simple Random, Stratified, Cluster, Systematic.
Units are picked based on convenience or judgement rather
than chance, so representativeness is not guaranteed.
Non-Probability Sampling
Examples: Quota sampling, Judgement sampling, Snowball
sampling.

Steps in Conducting a Sample Survey


1. State the objective — what exactly are we trying to find out?
2. Define the population — what is the full universe under study?
3. Build the sampling frame — assemble a complete listing of population units.
4. Pick the sampling technique — simple random, stratified, cluster, and so on.
5. Fix the sample size — guided by the precision and confidence level desired.
6. Carry out data collection — through questionnaires, interviews, or direct
observation.
7. Run the analysis — work out estimates and carry out hypothesis tests.
8. Arrive at conclusions — infer about the population, stating the precision attached.

1.3 Sampling and Non-Sampling Errors


No survey, however carefully conducted, is free of error. The errors that creep into data
collection are broadly split into two categories: Sampling Errors and Non-Sampling
Errors. Knowing the difference helps in designing better studies and in reading results
with the right amount of caution.

Sampling Errors
Sampling Error — the error that shows up purely because only a part of the population
— not the whole of it — was examined. The value of the sample statistic (say, the
sample mean) will generally not match the true population parameter exactly; this gap is
termed the sampling error.
Sampling Error = Statistic − Parameter = x̄ − µ
Key property: sampling error shrinks as the sample size grows larger.

Type Explanation
A systematic, one-directional deviation caused by a flawed
Biased Error sampling design. It does NOT disappear merely by enlarging
the sample.
A random fluctuation around the true value, which averages
Unbiased Error
out over repeated samples and shrinks as n grows.

Page 5
Sampling Theory — Unit 1 Notes

Numerical Illustration — Sampling Error


A logistics firm employs 1,200 delivery riders. Their actual average daily earning
(population mean µ) works out to ₹620.
A sample of 40 riders is drawn, and their average daily earning (sample mean x̄ ) comes
to ₹655.
Sampling Error = x̄ − µ = 655 − 620 = ₹35
∴ Sampling Error = ₹35 (the sample has overstated the true average by this margin,
purely because only 40 out of 1,200 riders were surveyed).

Non-Sampling Errors
Non-Sampling Error — an error that can creep into either a census or a sample
survey, arising not from the act of sampling but from other sources — mistakes during
data collection, faulty measurement, or bias on the part of respondents.
Important: unlike sampling error, this type of error does not necessarily fall as the
sample grows; in fact, it can become worse as the scale of the exercise increases.

Types of Non-Sampling Errors


Type Description & Example
Happens when some population units never make it into the
Coverage Error sampling frame. Example: A purely SMS-based survey leaves out
people who do not own a mobile phone.
The gap between the true value and what gets recorded. Example:
Measurement Error A respondent states monthly rent as ₹8,000 instead of the actual
₹8,750.
Respondents supply wrong answers because of confusion, a wish
Response Error
to appear in a good light, or outright dishonesty.
Certain chosen units fail to respond at all; if these differ
Non-Response Error
systematically from those who do respond, the results get skewed.
Processing Error Errors introduced while entering, coding, or tabulating the data.

Key Distinction to Remember: A census does away with sampling error, yet it remains
exposed to non-sampling errors — and often more so, given the sheer scale at which it
is conducted. A well-planned sample survey can, in practice, turn out more accurate
overall than a poorly executed census.

Page 6
Sampling Theory — Unit 1 Notes

1.4 Sampling Frames & Principles of Survey Design


Sampling Frame
Sampling Frame — a complete and current list (or map) of every unit belonging to the
target population, from which the actual sample is drawn.
Example: For a study covering registered taxpayers in a city, the frame could be the
municipal property-tax database.
Example: For a study of college library usage, the frame is the student enrolment
register.

Problems with Sampling Frames


Problem Meaning
Under-coverage Certain population units are absent from the frame.
The frame contains units that do not actually belong to the
Over-coverage
target population.
Duplication Some units are listed more than once in the frame.
Units appear bundled in groups rather than as separate
Clustering
individual entries.
The frame fails to reflect recent entries or exits from the
Outdated Frame
population.

Principles of Survey Design


A sound survey rests on a handful of core principles that guarantee the results are valid,
dependable, and genuinely useful. These form the backbone of any scientifically
conducted data-collection exercise.

Principle Explanation
The survey must actually capture what it sets out to
Principle of Validity
measure, with questions that are clear and relevant.
Data gathered should mirror the true values; every possible
Principle of Accuracy source of error — sampling and non-sampling — should be
kept to a minimum.
The sample must be large enough to deliver the level of
Principle of Sufficiency
precision and confidence that is required.
The required precision should be achieved at the lowest
Principle of Economy
possible cost and in the shortest possible time.
The survey plan must be workable given the resources and
Principle of Practicability
technology actually available.

Page 7
Sampling Theory — Unit 1 Notes

A sufficiently large random sample tends to reflect the


Principle of Statistical
characteristics of the population — this is what underpins
Regularity
probability sampling.
Principle of Inertia of Large Bigger samples tend to give steadier, more consistent
Numbers results, with less variability in the estimates.

1.5 Central Limit Theorem and its Applications


Central Limit Theorem — Statement — if a sufficiently large random sample (a rough
rule of thumb being n ≥ 30) is drawn from any population, irrespective of the shape of
that population's distribution, the distribution of the sample mean (x̄ ) will be
approximately Normal.
This stands among the most far-reaching results in statistics, forming the theoretical
bedrock for confidence intervals, hypothesis testing, and most of inferential statistics
generally.

Formal Statement of CLT


Let X1, X2, ..., Xn be independent and identically distributed random variables drawn
from a population with mean µ and finite variance σ². Then, as n becomes large:
x̄ ~ N(µ, σ²/n) — the sample mean follows a Normal distribution with mean µ and
variance σ²/n
Standardised form: Z = (x̄ − µ) / (σ / √n), which follows the standard Normal distribution
N(0,1) for large n.

What Does CLT Tell Us?


• Even when the underlying population is skewed, uniform, or otherwise irregular,
the sampling distribution of x̄ still tends towards Normal for large n.
• The mean of that sampling distribution equals the population mean µ — in other
words, sample means are unbiased estimators.
• The spread of the sampling distribution, given by the Standard Error σ/√n,
narrows as n increases — bigger samples yield more precise estimates.
• CLT makes it possible to rely on Normal distribution (Z) tables for inference even
when the shape of the underlying population is not known.

Numerical Illustration — Applying CLT


Monthly mobile-data bills for customers of a telecom operator have a mean µ = ₹560
and a standard deviation σ = ₹90.
A random sample of n = 25 customers is drawn.
By CLT, the sampling distribution of x̄ ~ N(560, 90²/25) = N(560, 324).
Standard Error (SE) = σ/√n = 90/√25 = 90/5 = ₹18.

Page 8
Sampling Theory — Unit 1 Notes

Question: What is P(x̄ > 587)?


Z = (587 − 560) / 18 = 27/18 = 1.5
P(Z > 1.5) = 1 − Φ(1.5) = 1 − 0.9332 = 0.0668
∴ P(x̄ > 587) ≈ 6.68% — there is only around a 6.68% chance that the sample mean bill
exceeds ₹587.

Applications of CLT in Actuarial Work


• Confidence Intervals: Building intervals of the form x̄ ± Z · σ/√n to estimate an
unknown population parameter.
• Hypothesis Testing: Checking claims made about a population mean through Z-
tests.
• Risk Assessment: Modelling the aggregate of many individual claims, which
tends to approximate a Normal distribution as the number of claims grows.
• Portfolio Theory: Returns from a large, well-diversified portfolio tend to be
approximately Normal.
• Insurance Operations: The expected claim frequency across a large book of
policies tends to follow the patterns predicted by CLT.

Numerical Illustration — Confidence Interval Using CLT


A random sample of 64 policyholders shows an average claim amount of x̄ = ₹12,400.
The population standard deviation is known to be σ = ₹2,400.
We wish to construct a 95% Confidence Interval for the true mean claim µ.
SE = σ/√n = 2400/√64 = 2400/8 = ₹300.
Z value at 95% confidence = 1.96
CI = x̄ ± Z · SE = 12,400 ± 1.96 × 300 = 12,400 ± 588
CI = (₹11,812 , ₹12,988)
∴ We can state, with 95% confidence, that the true mean claim amount lies between
₹11,812 and ₹12,988.

1.6 Standard Error and its Interpretation


Standard Error — Definition — the standard deviation of the sampling distribution of a
statistic — most often the sample mean. It tells us how much that statistic is expected to
vary from one sample to the next.
Do not confuse the Standard Error (SE) with the Standard Deviation (σ): SD describes
the spread within the raw data, whereas SE describes the spread of the sample statistic
itself.

Page 9
Sampling Theory — Unit 1 Notes

Formulas for Standard Error


Quantity Formula
SE of Mean SE(x̄ ) = σ / √n, where σ = population SD and n = sample size
SE(p̂ ) = √[P(1−P)/n], where P = population proportion and n =
SE of Proportion
sample size
SE(x̄ ) = s / √n, where s = sample standard deviation, used as
SE when σ is unknown
a substitute when σ is unknown

Interpreting Standard Error


• Smaller SE means a more precise estimate: the sample mean is likely to sit close
to the true population mean.
• Larger SE means a less precise estimate: there is more uncertainty about where
the true mean actually lies.
• SE depends jointly on σ and n: greater underlying variability pushes SE up, while
a larger sample size pulls it down.
• SE = 0 would imply the sample statistic equals the population parameter exactly
— possible only when the entire population has been surveyed.

Numerical Illustration — Effect of Sample Size on SE


Population: heights of saplings in a nursery, with σ = 6 cm.
Sample n = 9: SE = 6/√9 = 6/3 = 2.0 cm
Sample n = 36: SE = 6/√36 = 6/6 = 1.0 cm
Sample n = 144: SE = 6/√144 = 6/12 = 0.5 cm
Observation: doubling the sample size brings SE down by a factor of √2 ≈ 1.41. To cut
SE in half, the sample size must be quadrupled.
∴ SE falls at the rate of 1/√n — increasing the sample size brings diminishing returns.

Numerical Illustration — SE of a Proportion


In a survey of 250 commuters, 90 said they preferred travelling by metro over the bus.
Sample proportion: p̂ = 90/250 = 0.36 (36%)
Estimated SE(p̂ ) = √[p̂ (1−p̂ )/n] = √[0.36 × 0.64 / 250] = √[0.2304/250] = √0.0009216 =
0.0304 (≈ 3.04%)
95% CI for P = p̂ ± 1.96 × SE = 0.36 ± 1.96 × 0.0304 = 0.36 ± 0.0596 = (0.3004,
0.4196), i.e., (30.0%, 42.0%)
∴ We are 95% confident that the true proportion of commuters preferring the metro lies
between 30.0% and 42.0%.

Page 10
Sampling Theory — Unit 1 Notes

Relationship Between SE, CLT and Sampling


The Standard Error is what links the Central Limit Theorem to practical inference. CLT
establishes the shape of the sampling distribution (Normal), while the Standard Error
establishes how spread out that distribution is. Used together, they let us make exact
probability statements about how close a sample estimate is likely to be to the true
population value.

Concept Role in Inference


Population (µ, σ) The true, unknown quantity we are trying to estimate
Sample (x̄ , s) Our observed data, used to estimate the population parameters
Guarantees the sampling distribution of x̄ is approximately Normal
CLT
for large n
Quantifies how precise our estimate is — i.e., how spread out the
Standard Error
sample means are
Combines CLT and SE to provide a range likely to contain the true
Confidence Interval
parameter
Combines CLT and SE to judge whether observed data is
Hypothesis Test
consistent with a stated claim about µ

Unit 1 — Summary at a Glance


Topic Key Takeaway
Population vs Population is the entire group; Sample is the subset actually studied.
Sample Sample statistics are used to estimate population parameters.
A Census covers everyone (no sampling error, but resource-heavy); a
Census vs Survey Sample Survey covers a subset (cheaper and faster, but carries
sampling error).
Arise because the whole population is not studied. They shrink as the
Sampling Errors
sample size grows, and may be biased or unbiased.
Stem from data collection, measurement, or processing issues. These
Non-Sampling Errors
occur in censuses too, and can worsen as scale increases.
The full list of population units from which the sample is drawn; it
Sampling Frame
should be current, complete, and free of duplication.
Survey Design Validity, Accuracy, Sufficiency, Economy, Practicability, Statistical
Principles Regularity, and Inertia of Large Numbers.
For large n, the sampling distribution of x̄ is approximately N(µ, σ²/n),
Central Limit
regardless of the shape of the parent population — the foundation of
Theorem
statistical inference.
Standard Error SE = σ/√n, a measure of how precise the sample estimate is. A

Page 11
Sampling Theory — Unit 1 Notes

smaller SE is better; halving the SE requires quadrupling the sample


size.

End of Unit 1 — Sampling Theory Notes | [Link]. Actuarial Science, Semester III | 2025–26

Page 12

You might also like