0% found this document useful (0 votes)
6 views12 pages

Understanding Sampling Distributions

1) The document discusses sampling distributions and how sample statistics like the sample mean and proportion vary across different samples from the same population. 2) It explains that as sample size increases, the sampling distribution of the sample mean and proportion approaches a normal distribution, even if the population distribution is not normal, according to the central limit theorem. 3) Examples are provided to illustrate how to calculate the sampling distribution of the sample mean when taking multiple samples from a population with a given mean and standard deviation.

Uploaded by

Ravinder Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views12 pages

Understanding Sampling Distributions

1) The document discusses sampling distributions and how sample statistics like the sample mean and proportion vary across different samples from the same population. 2) It explains that as sample size increases, the sampling distribution of the sample mean and proportion approaches a normal distribution, even if the population distribution is not normal, according to the central limit theorem. 3) Examples are provided to illustrate how to calculate the sampling distribution of the sample mean when taking multiple samples from a population with a given mean and standard deviation.

Uploaded by

Ravinder Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 3 : Sampling Distributions

1 Introduction
Statistical infcloncc is a,ll about using a samplc to prcdict charactcristics of a population. The sanrplc chal'a,ctct'istics a,r'c summarizcd using thc sample statistics
whilc thc cha,r'a,ctcristics of

eters.

a,

population arc sumnrurizad by thc population param-

Thus, statistzcal, 'in,fererux; ln'ils down to estimati,ng a populo,tion parameter

us'ing analogous sample stati,st'ic.

ES 1. The traditionally low pclccntagc of fcmalcs in thc IIM studcnt body has
been a topic of hcatcd dcbatc for quitc somc timc. You can cstimatc thc pcrccntagc of
fcmalcs in ther incciming batch of PGP studcnts of all thc IIMs combincd
using the concsporrding pcrcclta,gc irr

a,

(parameter)

ra,ndom samplc of (incoming PGP) studcnts

you sclcct from thc conrbincd studcrrt rcgistcr of all thc IIMs (statistic).

Two most comntonly rmcd sta,tistir:s arc thc sample mean

(r-=)

and sample pro-

portion (p). Thc corrcsponding paramctcrs arc thc population mean (p) and
population proportion (p).
Eg

2.

For thc a,bovc crxn,mplc, thc statistic is thc samplc proportion of fcmalc

PGP-I studcnts in your samplcr whikr thc paramctcr is thc corrcsponding proportion

in thc popula,tion of all PGP-I studcnts across all thc IIMs.


Clcarly, in ordcr for thc infclcntia,l proccdurc to bc "goc1d", thc samplc statistic
should be prettv close to thc unknown population paramctcr. In ordcr to cnsurc that,
we have to study thc

sampling distribution of samplc statistics.

1.1 IIow Does it Work ?


stltistics zuc bzrsctl orr str,rnplcs, tltcy will havc diffblcnt valucs lbr dilfcrcnt samplcs. T'hus, a, sta,tistic is a,lso a, random variablc and will havc a probability
distribution that will assign probabilitics to thc possiblc valucs it can takc fbr thc difSincc sarnplc

fercnt samples. The probability distribution of a sample statisti,c zs called a samplzng


di,stri,buti,on.

Eg 3. Supposc thc tnrc proportion of fcmalcs in thc PGP 13-15 batch across all
the IIMs is p = 6.13. Supposc yorr sckr<;t all possiblc random sa,mplcs of 30 studcnts -

will yicld a valuc of thc samplc proportion (of fcmalcs in that


sample). If you construct a histogra,rn of thosc valucs, what you will gct is prcciscly
cach of thosc santplcs

the sampling distribution of thc: sa,rnplc proportion (of fcmalcs).

Eg 4. IAS Officers With nrga,rcl to thc rcccnt incidcnt about thc suspcnsion of
a female IAS officer (fbr her a,ctions a,ga,inst the sand mafias in UP), it is of interest
to estimate the avera,ge tenure of an IAS officer. Suppose you select ra,ndom sa,mples
of 20 IAS officcrs. For cach samplc, yorr calculatc thc avcragc tcmrrc of an officcr. If
you continue doing this for

a,

largc nurnbcr of samplcs and draw thc histogram (of thc

avcragc tcnurc va,lucs of ca,ch sa,rnpkr), what you

will havc is prcciscly thc sampling

distribution of the sa,rnple rnean tenule of IAS officersl.


Howeuer,
statr,sttc

it

for

is

irnposs'ibl,e

each,.

to col,kr:t

In reality,

eur:ry possible sample and calculate the sample

u)e carl only collect one sample. Howeuer, th,e sarnplzng

distributi,on wi,ll tell us how much, a stati,st'ic would uary from sample to sample and
will help us to predict h,ow close a stat'istic falls to the param,eterit estimates.

I.2

Recapitulation

Let us briefly recapitula,te

a, c;oupk:

of important concepts which a,re essentia,l lor

a,

propcr undcrstauding of sa,mpling distribution.

1. Empirical rule

thc mcan and standard dcviation of a samplc


it can bc a^ssumcd that thc samplc comcs from

Suppos<r

and s lcspcctivcrly2.

If

is
a

distribution that is approxirna,tcly bcll-shapcd and symmctric, thcn

68Ta of

thc obscrvatiorrs would fall within

i.e within

95To of

(" -

",

standard dcviation of thc mcan

r; + s).

thc obscrvations would fall within 2 standard dcviation of thc mcan

i.e within

(7 2s,i

+ 2s).

rHere the popula,tion is the set of ALL IAS officers currently on duty; the parameter is the true

of an IAS offic:er whic;L is unknown (since we cannot survey each and every IAS
officer). In lact, a, rcc:cnt Ilarvtl,rcl stucly has c'orrcluded that the average tenure of a,n IAS officer is
avera,ge tenure

only 16 months
2s is known a,s thc samplc stancla,rrl ckrviation and is givcn
-

bcing thc samplc units and n tht: sanrplc

siztr

,r, ,lf-frt.-i)2,
V n-I

tr;,,i -- L,..',n

of tho obscrva,tions would fall within 3 standard dcviation of thc


mca,n i.c within (n - 3s, e-: + 3s).
99.7%

6el'

1>7,

79 V

$-'3s [Link]
Fts i+zs
i.t3 s
This is impofta,nt to know trcca,usc a lot of variablcs obscrvcd in rcality havc
approximatcly bcll shapcd clistributions.
Central Limit Theorem (CLT) : This is a vcry wcll known rcsult in Statistics
which basically says tha,t a,s thc samplc sizc incrcascs (i.c as wc takc largcr
and largcr satnplcs), thc sa,nrpling distribution of thc samplc sta,tistic (mcan or
proportion) tcnds to a, trortna,l distribution. In fact this holds truc cvcn if wc
sample from a population whoscr distribution is modcratcly skcwcd or discrctc.
The only rcstriction is tha,t thc mcan and standard dcviation of thc population

distribution should cxist.

Note

z ) 30 is good <rnough to ensurc thc normality of thc sa,mplc mcan


(X) if thc populatiol distribution is not too skcwcd. (For ertrem,ely skewed

(a) Usually

dzstrr,brt"ti,ons, we nlalJ nrr:d a l,arger ,sam,pl,e s'ize

to ensure normali,ty of the

sample mean).
(b) Specifically, if the underlying population distribution is approxirnately nor-

mal ([Link]. bcll-shapcd a,nd symmctric), thcn


regard,l,c:ss

(c)

ol the

so,rrt,pl,r: s'i,ze

is also appr<-rximatcly normal

i.c thc CLT is not rcquircd in such a situation.

Thc CLT only assurcs normality of thc samplc mcan providcd thc samplc
sizc is largc;

it

docs not

tcll us anything about thc distribution of thc

individual sanrplc rurits (X,, i = 7, ...,n).


(d)

If thc undcrlying popula,tion distribution is binary with succcss probability, say p, thc sa,mplc mca,n (in this ca;sc the samplc proportion p) will b<r
approxima,tcly Norrntrl providod

np2l0

and nq

10.

Two rcasons why thc CLT is so uscful arc

Whcn thc sampling distribution of a samplc statistic (mcan or proportion)

is a,pploximatclv norrna,l. wc can usc thc trmpirical rulc to prcdict how


closc salnplc sta,tistics will bcr to thc truc population pa,ramctcr.
Sincc thc CLT holds for

zr,

divcrsc rangc of population distlibutions,

it

hclps

us to nrakc inf'clcnccs a,bout thc populaticln paramctcrs lcgardlcss of thc


shapc of thc population distribution. This is oftcn hclpful in practicc sincc
we usua,lly do not knciw thc truc shapc of thc population distribution (and

oftcn

it

is skcwcd).

Now, wc will looh irrto thc sa,rrrpling distributions of samplc nrcan and proportions in
dctail.

Sampling Distribution of the Sample Mean

Sincc mcans or avcragcrs a,rc so ubiquitous in Statistics,

it

is uscful to lcarn about

thcir

sampling distribution i.c how salnplc lncans vary from samplc to samplc and how closc

thcy will bc to thc popula,tion ntcln irt rcpcatcd sampling from thc population. Lct
us consider thc following cxamplc

Eg 5. Thcrc arc 3 catcrics in a, Iocality, say A, B and C. Thc chancc that a


student gcres to tltcm arc 60%, 20% and 20To rcspcctivcly. Assumc that hc chooscs
an ca,tcry cvcly day indcpcndcntly of his prcvious dccisions. It is known that thc
studcnt's favouritc dish in A, B and C costs Rs 100, Rs 140 and Rs 150 rcspcctivcly.

(i) Obtain thc distributiorr cil thc studcnt's daily cxpcnditulc on lunch.

f E"laHarn'

(ii) Obtain thc

ove

sa,mpling distribution of thc studcnt's avcragc cxpcnditurc on lunch

days.

9or-t& $?e '_


x

lcru

"l'

I arc

.tF

l aA_

dF

l,l

t1o

ls)

,fF

"q'

dh

.36 4 ('ry)

'al + ('-!-1')
a1 ) ( l.,? E-'t)
r .o! -+ ( u+-51').
'

'o1
'oI

+ (rg-lJ")
)'

/L-l5)?l!') )\

la,

r, c]

Lt-,Vortsu1

(iii) Obtain thc sampling distribution of thc studcnt's avcragc cxpcnditurc

ovcr + Sar^^f U g,y

oavs.

In order to atrswer (iii), we have to deal with a huge nurnber of clifferent

com-

binations of meals, each havirrg tr clifi'erent probability. Clearly doing this manually
is virtually impossiblc. Howcvcr, thc rcsults of sampling distribution (givcn bclow)
providcs us with :ur autornatcd way of a,chieving thc samc.

you draut sam,ples of

tt from a population wi,th, mearl p anfl stand,arj,


deuiatton o and calcttlate th,e sa,rtt,pl,e m,ean (a) for each. Then th,e m,earls wtl,l, h,aue
Suppose

.s'izr:

a sampling dzstri'bution uh,ich, util,l

bc: cr:'ntered

about the true population mean

Ji . In f act, i,J' n 2 30 , then thi,s


Na"r-ol w,ith mean pt anrl stanrlard, error of Ji.

wi,ll haue a standard deu'r,ation ( or stand,ard error)

distnbution wi,ll be approrim,atel'g


For Eg

5.

oI

thc population distribution of thc studcnt's daily cxpcnditurc


rtca,tl (p,) and sta,ndald dcviation (o) can casily bc shown to bc

a'bovc,

is givcn in (i) whos()

= lcn x'&o r t{[Link] + ls-a,r ,2o

c=lv/

For (iii),

n: 30

(H<-r.

' trg

vo,nic"n

Hcncc thcr rcqnircd sampling

crxpcnditurc ovcr 30 davs)

Note : Sincc

p and

will bcr Novy-a

t (tf

"lJi

to ,r^ru le99 tha,n thc original obscrvat


your samplc sizc (rz), thc sa,nrplc rrrcans will vary
closcr to thc truc popula,tion nrca,n (yr).
mcans tcnd

Eg 6. Rambhai's income Thc

salcs of food and

drink in Rambhai's stall vary

from day to day. The daily sa,les [Link] fluctua,te with mea,n p, : Rs 900 and standard
dcvia,tion o' - Rs 300. R,ambhari u'arts to calculato thc mca,n dailv salcs

fol thc wcck

to chcck how ht, is doirrg.

Ql. What woulcl

t;he rnean dzrily

s21,ls

= 1o

figures for the wcek center around

o ( ?"P .,luFio.." '^'*)

Q2. How much vtr,r'iability woulcl yon expect in the mean daily sales figures fbr the

(i( e

h5l",hli"

dr;$ ir.
+nt

h.o+

ck-o

wt d)

wcek ? Interprct.

Q t+,

l{<r-. ,n= | . He.r",, +t", d*r'"rL,ttln (s{o,.r^J ah e erAAv)


y\^rr\yl ao"; t^j SoJA/t 5l {Ov l, , r k^ [Link]' U b4
.

hbo,
If wc wcrc to obscrvc thc mca,n clzrily salcs for scvcral wccks,
with a standard rkrvia,tion l, tA .31

th

v,rn= =7"V
--tt3.31
round

t too

{Oo

to look at thc monthly salcs. What will bc thc


sampling distribution ? Will his mcan daily salcs for thc month vary morc or lcss
Q3.

Supposcr Rambhai now wa,nts

than thc mcan daily salcs for thc wcck /

1OO *ith rturrdard


rs, thc mcan daily salcs for thc month will tcnd to
vary
salcs fbr thc wcck and thus bc closcr to thc truc
mcan salcs of Rs 900. In fact,, s'i,ncc n : 30, CLT h,olds and th,e mean dai,ly sales Jor
the month r,s approrimately norm,allE di,stri,buted w'ith mean j o o
arr,d, stand,ard,
e,ror 51'T+ fe N(9oor 51'??).
His mean daily salcrs for thc rnonth will bc ccntcrcd around

Q4. What is thc probability ther,t


800 and 1000 Rupcrcs ?

th<,r

rncan daily salcs of thc month

will bc bctwccn

F(eoo 4 i=orrooo)

=F(sff3> ! affi;
(Z l t.s3)
= .q f [Link])- F(Z-(-[Link])
= ? (-t.s3

_ r.g3

The followirrg figurc portrays a possible population distribution for daily

t'gJ

sales

and also thc sarnpling distribution of thc mcan daily salcs for thc wcck and month
rcspcctively.

-1661 -

,o339
N

='1979

L-s
-/

loo

gwlut

d,-j
g^lM

Clcrarly thcrcr

is lAg S.

va,ria,bilit,y

in thc mczrn daily

sa,lcs fi'om nronth-to-nronth

tha'n from [Link] tha,n tl*rr'<r is firlrn day-to-day in thc daily

sa,lcs.

Eg 7. Suppostr in a, kittcry ga,ul(), you bct Rs 1 cln a numbur bctwccrr 0 and g that
you pick at ra'ndom. If you alc <:otrcct, you win Rs 5, othcrwisc you win nothing.
Let E denote yottr ea,rning (NOT plofit) in a pa,rticula,r game. Then, the probabilitv
distribution of E will bc

E=

, I'

'rl' %o

o
and hencc thc nxrar

(trr,)

will

Yr"

bc SX ho t o Tto =' S'


^

tho population standa,rd dcviation is o = 1.50. If you plav oncc a wcck for
thc ncxt ycar ([Link] 52 timcs), what will bc thc sampling distribution of your mcan
Suppose

it

? Can

bc a,pproxirna,tcrl by

a,

Normal distribution ? Explain.

e^hv\jh&9.

Lc.t X L {+1r
Ot.^r'tg

Thc^

X a, Br'v.o trr.,i
'lt"t
a\

=h
$ ow,r

6l -|r..,,,,^.,t c .1o,
(t^,f 4 s-z ltvv..t:r

vtvrarrL<n

+u

3aw,pL

-x

fa,l

v1

^)

^lr

sol

eA)^ h.o t

l-lt-,,

-ra

E=-sh

y*

Fa^qed)

h-Fo. Li" 6\

4-dtr,-t-,

yor, 6rnn

h = .52-

Lc\9\ra,r'vrg

Thu^ vy\t o\Alt e-qK hl't


No V\),

( S Z, o .

Aorr.t- l'\ft't4

@fr[ Sx

t'/n

:S =
S zx y'ro = S'2 1lo

= sl

-^

= E.J lca,.r

b.

"Fl'Y12xr'"-oLJ tJ ^ Uo*h^a, l d.;lh


t,.,ri tJ t^,o+ 1". N o w ,,u, o,)
et'+ tr,UL .

Noti:(i)Tg !A-- [oo, s*1 , {J^r ^h = lo, aa=1o


Itvr, i (1v-u si=E) r^n'r,( [Link] t" a I f vz Nav M4l,
(td )
[o,.r f, t.,r h.,- f- t hn' t t b c ht^l ea,nvrty.{ =sx

-sz

+v

Id ust

cl

Suppose you plav

a) What is thc pr
lcairt Rs 1

P(

b) What is thc pro

P(,
= ? (7

s 3'1o) - P ('z 1D
.1117

thc intcrval which

sampling Distribution of the Sample proportion

As with any proberbilitv distributiorr, wc can dcscribc thc sampling distribution of a


samplc statistic by its mean a,nd standard deviation.

If p is thc

sa'mpltr pt'opot'tion

fbr a la,ndom sa,mplc of

with proportion p, thun p hns

siza

n drawn fi'om

a,

population

trxra,rr

wl|- . r':fcr)r,r.
vG>= 1F(*r)
;
=
+'F
o wF,-l''
') )q = f xi = v'tvwtrt* 5l Svcq-tt> 6* 6l n h,.[Link]. Xat R,'r,o* f (q, p)
Note : (i) If n is suf{icic'tlv l:rr.i1. such t}rat botb np and n(I - p) ara at lcarst 10, E (r)= i4
F
X;

't

thcn this sa,mpling distributiort is zr,pproximatcly normal duc to CLT


(ii) Thc standarcl clcvia,tion of a, sa,rnpling distribution is known a,s thc standard
error. So, thc standa,rd crror of thc sa,mpling distribution of 1.r is
(iii) Sincc p -. Xlrt (X : llurrlrcl of succcss out r-rf zr, trials), wo can restatc thc

, lu-a

V(rC

la

p(r-

f)

Eg 8. Supposc

7B7o

of thc irtt:titning PGP studcnts (a,cross all IIMs) a1c fcmalcs

and you randomly sclcct 100 strrdcnts.

(i) What is thc sampling distribution of thc proportion of fcmalcs in ygu1 samplc

HCe,r rrr=ltFot

F='rB =) ^?-- tB>o j rr(r-p), g2>to

Lle-.-. CLr afftr'<,a q

.o38)

F-n (.,r,

(ii) what is thc sampli'g [Link]' of thc number of fcmalcs in your saqrDlql-:-

Ero.f

d..st:\..tEi.(roo).r8)
vrTr

A11'*. sc^w.f*-j d.;th


So,

lh

.-)

\.--l

if you select different rartdorrt

d"P:?
J

: x^, N (,i, @F;Lsx'rrlll'Tc )


sarnples

of

100 students each, then

the

typical

dcviation of thc sa,rnplc proportions (of cach samplc) from thc population proportion
(p = [Link]) wrII D('

/.
V

[Link]
ldo

__

Note :
(i) Clcarly, thc standard crror will

.O3,g

you incnro$Gyour samplc sizc.

For

c of 200 studcnts, thc standard crror of p

will

dccrcasc

a"s

standard crror, closcr will bc thc samplc pro-

(ii) If the sampling distribution is a,pproximatcly normal, wc can usc thc Empirical rule. For cxa,rnplc, in thc a,bovrr cxamplc, ncarly all thc samplc [Link]
I lic bctwccn

valious probabilitics rcga,r'ding p.

Eg

9. Internship :

Sttpposer

_\

rtttt of all first year strrdents enrollecl in the top

busincss schools across India,, a,borrt 55% wcnt abroad

for summcr intcrnship last


ycar. Supposc you randomly scrlcct a busincss school and it turns out to bc XLRI
which has about 350 students errrolled in the first year.
a) What is thc probability that at hast 50% of thc 350 XLRI studcnts will so abroad

=N

(l gr B,el)

r(z l

[Link]

)'r

c) What is thc probability th:rt btl,wcrn 50% and 70To of thc 350 XLRI studcnts will
go a,broad for intcrnship this ycar '/

('s) i

-. ?")

t- '03 2L
No

r E:

a.\. 5l

ltae

se f r'o Ll' "w s

Lour. A ovt 'lt /

B t',,no ,rz.,r'a,)

. o\^.t exac,l
d,;lxjL,,1.^.; .

Lef {= # dLRl s+-d-!,+B


,\,, Bi." ( 3ga, .-r-f
)
Thc

vt

a>

[Link]

b. Solvcd qs
10

B-t ++-.; ie r*J

l^o,

9o

r'rzr

-JJ

? (x

AL

^rso

= L5 tsi, (s 5?z 'ts )

tedr.o ,,.,o

iJt*,^,,h.

> rls )

3rz

So 1,. h- o

tn

l.
=

5l

3sl

l?s.

Population, Data and Sampling Distributions


r Population distribution

is thc probabilitv distribution fiom which wc draw

thc samplcr. It's pa,ralncrtc' va,lucs (like p and p,) arc gcncrally unknown.

r Data distribution

is thc distribution of thc samplc data -

it is dcscribcd

bv

statistic likc sa,mplc proportions and mcans. With random sampling, largcr
thc samplc siza n, closcr is thc data distribution to thc population distribution.

Sampling distribution is tlrc probability distribution of thc samplc statistics


likc samplc urcralt aud prolrortiols. It tclls us how closc thc samplc sta,tistics arc
cxpcctcd to fall from thc population paramctcr in rcpca,tcd random sampling
from ther population.

Eg 10. Thc 2006 Scnatorial clcction in New York pittcd thc Dcmocratic candidatc
Hillarv Clinton aga,inst thc Rcpublir:an candidatc John Spcnccr. Thc CNN cxit poll,
bascd on 1336 votcrs rcportcd

tbal

of thc sa,mplc has votcd for Clinton whilc


33% votcd for Spcuccr. Whcn 4.1 rnillion votcs wc'rc tallicd, it was found that 68%
67%o

votcd for Clinton and 32To for Sprrrrcur.

What is population distribution, da,ta distribution and sampling distribution of


thc samplc proportion ?

X) will bc thc votc outcomc such that r = 1


for Clinton tr,nd r - 0 for Spcrrccrr. Thus, X ha^s mcan 0'68 . which is th<r

1. Herc, thc randon

va,ria,bl<r (sav,

para,mctcr. Thc popul:rticin clistribution will bc a graph with two stcms with
mean 0.68.

'se

I
'3u

o .6t I

x+

2. Thc data distribution will bc thc graph of thc


i.c 0 with proportion 0.33 arr<l 1 with 0.67.

f
I
11

'6+

.n+

1336 voting outcomc

in thc samplc

Harc, np =

l,336x'69 ) le

a,n<l

n(I - p) = 1336

^'3 2)t0 Su,

thc sampling

distrbution of p (samplc propoltion of votcrs voting for Clinton) will bc


with mcatt b . ' 69
11d standard crror' / '1g,rr'= Z
'ot

ft-

a,

N (.es,.ot3

Thc population atrd data distribtrtiorr

a,r'c

No"*,o
r.c

1336

discrctc, cr-rnccntratcd at 0 and 1. Howcvc1,

thc sampling distribution is bcrll sha,pcd - it dcscribcs thc sprcad of all possiblc samplc
proportion valucs collesponding t<l all possiblc random samplcs of votcrs rcprcscnting
thc voting pcrccntagc of Clinton.

'6s

t2

Common questions

Powered by AI

The standard error is a crucial measure in the context of the sampling distribution of a sample statistic, as it quantifies the variability of the statistic from sample to sample. Essentially, it provides an estimate of how much the sample statistic is expected to fluctuate around the population parameter. A smaller standard error indicates that sample statistics are likely to be closer to the population parameter, enhancing the reliability of statistical inferences made with the sample. Thus, standard error is a pivotal component in confidence intervals and hypothesis testing .

The Central Limit Theorem allows for the use of normal distribution approximations by stating that the sampling distribution of the sample mean (and proportion) will become approximately normal with sufficiently large sample sizes (typically n ≥ 30), regardless of the population's original distribution. This is particularly useful in practical scenarios where the exact distribution of the population is unknown or non-normal. The normality assumption enables the application of various statistical tools and techniques, such as z-tests and t-tests, making inferential processes much more straightforward and widely applicable .

The sampling distribution of the sample mean is the probability distribution of all possible sample means from all possible samples of a specific size from a population. For example, if we draw samples of size 20 from a population of IAS officers with an unknown true average tenure, calculate the mean tenure for each sample, and plot these means, we obtain the sampling distribution of the sample mean. This distribution will be centered around the true population mean and will have a standard deviation known as the standard error. This distribution can be used to make inferences about the population mean .

The sample size directly affects the standard error of the sampling distribution, which is the standard deviation of the sample statistic's distribution. As the sample size increases, the standard error decreases because the sample mean becomes more concentrated around the true population mean. Mathematically, the standard error of the sample mean is calculated as the population standard deviation divided by the square root of the sample size, indicating the inverse relationship between sample size and standard error .

It is impossible to collect every possible sample due to the sheer volume and cost of data collection, as well as time constraints. Additionally, populations can be infinitely large. To overcome this limitation, we rely on sampling distributions, such as those described by the Central Limit Theorem, which predict how sample statistics are expected to behave across different samples. By understanding the probability distribution of these sample statistics, we can make inferences about population parameters using only one or a few samples, rather than all possible samples .

The Central Limit Theorem (CLT) provides a crucial foundation for making statistical inferences about population parameters. It states that as the sample size increases, the sampling distribution of the sample mean (or proportion) approaches a normal distribution, regardless of the original population's distribution, provided the population mean and standard deviation exist. This allows statisticians to apply the empirical rule to predict how close sample statistics are to true population parameters, even when the population distribution is unknown. CLT is particularly valuable because it applies to a wide range of population distributions, facilitating the inference process .

Empirical rules, based on the properties of the normal distribution, help in making predictions about sample statistics by providing a framework to estimate the likelihood that a statistic will fall within specific ranges. For example, empirical rules state that approximately 68% of observations lie within one standard deviation from the mean, 95% within two, and 99.7% within three, in a normal distribution. This allows statisticians to gauge the probable accuracy and precision of sample estimates by determining the range within which most of the sample statistics should fall. Using this, one can evaluate the closeness of sample statistics to the population parameter .

Statistical inference uses sample statistics to estimate population parameters. A sample statistic, such as sample mean or sample proportion, is computed from a randomly selected sample, and these statistics are used to make inferences about the corresponding population parameters, like population mean or proportion. This process involves examining the sampling distribution of the sample statistic to predict how close the statistic is likely to be to the unknown population parameter. For example, to estimate the percentage of females in the incoming PGP batch across all IIMs, one would use the sample proportion of females in a random sample of PGP students .

The empirical rule allows statisticians to predict the percentage of observations that fall within certain ranges in a bell-shaped, symmetric distribution. Specifically, about 68% of observations fall within one standard deviation of the mean, 95% within two, and 99.7% within three. This helps in estimating how much a sample statistic, such as the sample mean, varies from the true population parameter. For instance, if the sampling distribution of the sample mean is approximately normal, the empirical rule can be used to approximate the likelihood that a sample mean will be within certain deviations from the population mean .

A large sample size is critical when the population distribution is not normal because it helps to ensure that the sampling distribution of the sample mean becomes approximately normal, as guaranteed by the Central Limit Theorem (CLT). For moderately skewed or discrete distributions, a larger sample size (typically n ≥ 30) is often sufficient to achieve this approximation. However, for extremely skewed distributions, an even larger sample size may be necessary to reach a normal approximation, which is essential for making reliable inferences about population parameters using the normal distribution framework .

You might also like