0% found this document useful (0 votes)
2 views5 pages

Sampling

The document discusses various sampling methods and concepts in statistics, including definitions of population and census, types of sampling (random, convenience), and the importance of representative samples. It provides exercises to illustrate these concepts, such as identifying biased samples and calculating statistics from given data. Additionally, it covers the implications of sample size and variability in statistical analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views5 pages

Sampling

The document discusses various sampling methods and concepts in statistics, including definitions of population and census, types of sampling (random, convenience), and the importance of representative samples. It provides exercises to illustrate these concepts, such as identifying biased samples and calculating statistics from given data. Additionally, it covers the implications of sample size and variability in statistical analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

6 Sampling

Skills check Exercise 6.2


1. 1,1; 1,2; 1,3; 1,4; 2,1; 2,2; 2,3; 2,4; 1. a) Unlikely to get a representative sample with
3,1; 3,2; 3,3; 3,4; 4,1; 4,2; 4,3; 4,4 this method – any reason which identifies a
group likely to be over, or under-represented in
Exercise 6.1 the sample is a good answer, e.g. people who
1. a) i) A population is the complete group of have more time are more likely to fill out the
people (or observations) of interest in a questionnaires
statistical investigation. b) Unless the route is fully booked the offer costs
ii) A census is when information is gathered the airline almost nothing, and getting people to
about all members of the population. fill out the questionnaire engages their attention
b) No single answer here – your example should to consider the possibility of flying on that
show clearly why it is biased – for example, route at some point (and paying for a ticket!).
asking about health issues outside a fitness 2. Anything where the sample is just done by who
centre. happens to be in a particular place at a particular
time is the easiest way to identify a convenience
2. a) i) A list in which all members of the
sample, e.g. outside a supermarket / train station /
population appear.
school.
ii)  A sampling unit can be an individual, or a
household, or the result of an experiment 3. a) If they were disposable batteries, then a census
(e.g. how many heads occur in 3 tosses of a would not be possible as there would be no
coin) batteries to sell. For rechargeable batteries it
b) i) All adults in the UK, could use the latest would be possible, but very expensive to test
electoral register, or census files if it is close them all.
to the time at which the census was last b) Testing all 50 in one box means that the box
done (only done every 10 years) can be removed. Taking only one or two from a
ii)  The number of adults in a country changes box until you have 50 batteries means a number
every day – people move in and out; of boxes are affected and have to be refilled
some die; others become adults (on their until they have 50 batteries again. That problem
18th birthday in the UK); and any such could be avoided by taking a random sample
register will never have everyone on it that of 50 at a point in the process before they get
should be there, so even on the day the packed into the boxes.
information was collected it would not be 4. a) A census would be very time consuming and
100% accurate. expensive.
3. Each member of the population has an equal b) They would be more concerned with keeping
chance of being selected, and all possible their important customers happy so they may
combinations are equally likely. wish to deliberately set up a sampling method
4. a) All the customers of that bank designed to focus on particular types of
b) Use the account records held by the bank account.
c) An individual customer may hold more than
one account, possibly even in more than one Exercise 6.3
name (a woman might have accounts in her 1. Methods a and e will give a random sample –
maiden name and married name). More in e it does not matter that the sampling frame
than one person can have the same name so has been constructed in a non-random manner
removing duplicate names may actually remove because the method of choosing the members of
customers. the sample is random.

© Oxford University Press 2018: this may be reproduced for class use solely for the purchaser’s institute Sampling 1
2. The following samples for a, b are based on using Exercise 6.4
the two-digit numbers in the table, ignoring 00 and
91–99 and any repeats until 10 different numbers 1. A statistic is just a function of (some or all of)
are obtained: the set of observations collected, so a to d are all
a) 52, 44, 17, 71, 20, 63, 47, 88, 22, 02 statistics but e is not as it involves parameters.
b) 52, 54, 07, 08, 43, 49, 24, 73, 67, 64 20
⎛ 1⎞
c) For a small sample you could just ignore 00
2. a) ∑
i =1
X i ~ B ⎜ 20, ⎟
⎝ 3⎠
and 43–99 but it is very wasteful of the random 2 18
⎛ 20 ⎞ ⎛ 20 ⎞ ⎛ 1 ⎞ ⎛ 2 ⎞
numbers – so you could take 01–42 and also b) P ⎜ ∑ X i = 2 ⎟ = ⎜ ⎟⎜ ⎟ ⎜ ⎟ = 0.0143
⎝ i =1 ⎠ ⎝ 2 ⎠⎝ 3 ⎠ ⎝ 3 ⎠
51–92 (after subtracting 50 from anything in
1 ⎛ 20 ⎞ 20
this range) c) n = 20, p =
⇒E P ⎜ ∑ X i = 2 ⎟ = np = ;
Just using 01–42 gives: 17, 20, 22, 02, 18 3 ⎝ i =1 ⎠ 3
using the subtraction method gives: 02, 17, 21, ⎛ 20
⎞ 40
P ⎜ ∑ X i = 2 ⎟ = npq =
Var
20, 13 – note that rather than taking the next ⎝ i =1 ⎠ 9
block of 42 numbers immediately after the first, 10
⎛ 3⎞
the arithmetic is much simpler to subtract 50
3. a) ∑
i =1
X i ~ B ⎜ 10, ⎟
⎝ 4⎠
– it means that 00, 43–50 and 93–99 are not ⎛ 10 ⎞ ⎛ 10 ⎞
used which looks a little strange, but it is much b) P ⎜ ∑ X i ≤ 7 ⎟ =
1 − P ⎜ ∑ Xi > 7 ⎟
⎝ i =1 ⎠ ⎝ i =1 ⎠
easier.
⎧⎪⎛ 10 ⎞ ⎛ 3 ⎞8 ⎛ 1 ⎞2 ⎛ 10 ⎞ ⎛ 3 ⎞ ⎛ 1 ⎞ ⎛ 3 ⎞ ⎫⎪
9 10
d) As long as you specify what you intend to do = 1 − ⎨⎜ ⎟ ⎜ ⎟ ⎜ ⎟ + ⎜ ⎟ ⎜ ⎟ ⎜ ⎟ + ⎜ ⎟ ⎬

before looking at the table of numbers there is ⎪⎩⎝ 8 ⎠ ⎝ 4 ⎠ ⎝ 4 ⎠ ⎝ 9 ⎠ ⎝ 4 ⎠ ⎝ 4 ⎠ ⎝ 4 ⎠ ⎪⎭
no reason why it has to start at the top left. = 0.474
Just using 01–38 gives: 07, 20, 22, 16, 11, 08 3
c) n = 10, p = ⇒ E(X ) = np = 7.5;
using 01–38 and 41–78 (after subtracting 40 in 4
this range) gives 07, 27, 36, 20, 25, 22. Var(X ) = npq = 1.875
3. The actual samples depend on the numbers
4. a) � = 0.2 × 1 + 0.8 × 3 = 2.6
generated by your calculator. But it should give
b) (1, 1, 1) (1, 1, 3) (1, 3, 1) (1, 3, 3) (3, 1, 1)
you an insight into the amount of variability in
(3, 1, 3) (3, 3, 1) (3, 3, 3)
the answers taking small samples – everyone 5 5
owned a mobile phone so there is no variability These 8 possible samples have means: 1, , ,
3 3
there and the favourite type of film does have, 7 5 7 7
, , , and 3 and the modes are 1, 1, 1, 3, 1,
but it is not numerical so there is no easy way 3 3 3 3
to summarise the variability. The times are all 3, 3, 3 and the medians are 1, 1, 1, 3, 1, 3, 3, 3
relatively small values with a number of repeats
so the averages of your sample times are likely Exercise 6.5
to be closer together than the average number of 4
1. a) E( X 10 ) = 5; Var( X10=
= 0.4
)
DVDs owned. 10
4. and 5. Again, the actual samples depend on the 7.8
b) E( X 10 ) = 26.3; Var(X=10 ) = 0.78
numbers generated by your calculator. These 10
practical activities are designed both to give you a 4 6 6 4
c) E(Z ) = + + + = 2;
little bit of experience in the process of generating 10 10 10 10
samples (pretty tedious to do lots of it …) and 4 12 18 16
E(Z 2) = + + + =5
the opportunity to get some understanding of just 10 10 10 10
how much variability results from taking random ⇒ Var(Z ) = 5 − 22 = 1
samples. 1
⇒ E ( Z 10 ) = 2; Var ( Z 10 ) = = 0.1
10

© Oxford University Press 2018: this may be reproduced for class use solely for the purchaser’s institute Sampling 2
3
1 ⎛ σ2 ⎞
d) E(X ) = (9x – x 3)dx 3. a) X n ~ N ⎜ μ , ⎟
−3 36 ⎝ n ⎠
3
⎡ 1 ⎛9 1 ⎞⎤ P(| X n − μ | > 0.25σ ) < 0.1

= ⎢ ⎜ x 2 − x 4 ⎟⎥
⎣ 36 ⎝ 2 4 ⎠ ⎦ −3 ⎛ ⎞
1 ⎛ 81 81 ⎞ 1 ⎛ 81 81 ⎞ ⎜ Xn − μ ⎟
= ⎜ − ⎟− ⎜ − ⎟=0 ⇒ P ⎜ Z=
⇒ > 0.25 n ⎟ < 0.1
36 ⎝ 2 4 ⎠ 36 ⎝ 2 4 ⎠ ⎜ σ ⎟
⎜ n ⎟
(this could have been observed by symmetry) ⎝ ⎠
3
1
E(X 2) = (9x 2 – x 4)dx Φ(α) = 0.05 ⇒ α = 1.645 ⇒ 0.25 n > 1.645
−3 36
⎡ 1 ⎛ 3 1 5 ⎞⎤
3 ⇒ n > 1.645 × 4 ⇒ n > 43.3

= ⎢ ⎜ 3x − x ⎟ ⎥

⎣ 36 ⎝ 5 ⎠ ⎦ −3 so the smallest sample size is 44.
1 ⎛ 243 ⎞ 1 ⎛ 243 ⎞ b) P(| X n − μ | < 0.1σ )
= ⎜ 81 − ⎟ − ⎜ −81 + ⎟ = 1.8
36 ⎝ 5 ⎠ 36 ⎝ 5 ⎠
⎛ ⎞
⇒ Var(X ) = 1.8 – 0 = 1.8 ⎜ X 44 − μ ⎟

=P⎜ Z = < 0.1 44 =0.663 ⎟
⎜ σ ⎟
1.8 ⎜ ⎟
⇒ E( X10 ) = 0; =
Var( X10 ) = 0.18 ⎝ 44 ⎠
10
23.1 = 2 × (Φ(0.663) – 0.5) = 0.493
2. a) Var (X=
15 ) = 1.54
15
4. a) S 25 ∼ N  85, 9.2 
2
23.1
b) Var =
Xn <1  25 
n
⇒ n > 23.1, so the minimum sample size is 24 ⎛ ⎞
⎜ 83 − 85 ⎟
P(S 25 < 83) =
 P⎜ Z < =
− 1.087 ⎟
1 4 6 9 ⎜ 9.2 ⎟
3. a) E(X ) =
+ + + = 5; ⎜ ⎟
4 4 4 4 ⎝ 25 ⎠
1 16 36 81
E(X 2) = + + + = 33.5 = Φ(–1.087) = 0.139
4 4 4 4
⇒ Var(X ) = 33.5 – 52 = 8.5 ⎛ 9.22 2.12 ⎞
b) D =
S 25 − P 5 ~ N ⎜ 2, + ⎟ = N(2, 4.2676)
8 .5 ⎝ 25 5 ⎠
⇒ 5; Var (X=
b) E( X 12 ) = 12 ) = 0.708
12 0−2
P(D < 0 ) = P  Z < = −0.9681
 4.2676 
Exercise 6.6
= 1 – Φ(0.9681) = 0.166
⎛ 7.222 ⎞ ⎛ 44.7 2 ⎞
1. a) X 66 ∼
~ N ⎜ 600, ⎟ 5. C 8 ~ N ⎜ 1983,
⎝ 6 ⎠ ⎟
⎝ 8 ⎠
⎛ ⎞
⎜ 597 − 600 ⎟  P  Z > 2000.0625 − 1983 = 1.080
P(C 8 > 2000) =
b) P( X 6 < 597) = P⎜ Z < =
−1.021⎟  44.7 
⎜ 7.2 ⎟  
⎜ ⎟ 8
⎝ 6 ⎠
= 1 – Φ(1.080) = 0.140
= Φ(–1.021) = 0.154 The question referred to in the margin note
⎛ 4.52 ⎞ (Summary exercise of Chapter 4, Q7) is the same
2. a) X10 ~ N ⎜ 352, ⎟
⎝ 10 ⎠ question looking at the total number of characters
⎛ ⎞ rather than the average number per page.
⎜ 350 − 352 ⎟
P( X10 > 350) = P⎜ Z > =
− 1.405 ⎟
⎜ 4.5 ⎟
⎜ ⎟
⎝ 10 ⎠
= Φ(−1.405) = 0.920
b) Φ(α) = 0.01 ⇒ α = −2.326
350 − 352
⇒ ⇒ < −2.326 ⇒ n > 27.38
4.5
n
so the smallest sample size is 28.

© Oxford University Press 2018: this may be reproduced for class use solely for the purchaser’s institute Sampling 3
Exercise 6.7 Summary exercise 6
1. While there are three outcomes (home win,

approx
4.32 ⎞
1. C 50 ~ N ⎜ 43.2, ⎟ draw and away win) these are not equi-probable
⎝ 50 ⎠
outcomes (there tend to be more home wins in
  most league formats).
 P Z > 44 − 43.2 = −1.3155
P(C 50 > 44) ≈  4.3 
  2. a) A random sample is one in which each member
50
of the population has an equal chance of being
= 1 – Φ(1.3155) = 0.0942 selected, and all possible combinations are
equally likely.
approx ⎛ 422 ⎞
2. S 45 ~ N ⎜ 912, ⎟ b) Any reason which identifies a group likely to
⎝ 45 ⎠
be over, or under-represented in the sample is a
⎛ ⎞ good answer.
⎜ 900 − 912 ⎟
P(S 45 > 900) =
P⎜ Z > =
− 1.917 ⎟ 3. It would be very expensive to test every battery for
⎜ 42 ⎟
⎜ ⎟ its usable life before it was sold – and they would
⎝ 45 ⎠
need to be charged again.
= Φ(1.917) = 0.972
4. Your answer will depend on the particular random
approx ⎛ 8.9 2 ⎞ sample you have taken, but many sample means
3. a) T 60 ~ N ⎜ 23.3, ⎟
⎝ 60 ⎠ will lie in, or close to, the interval 55–60.
⎛ ⎞ 5. a) AB AC AD BA BC BD CA CB CD DA
⎜ 25 − 23.3 ⎟
P(T 60 < 25) =
P⎜ Z < =
1.480 ⎟ DB DC
⎜ 8.9 ⎟ b) If A, B are 50 cent coins, C is the 20 cent coin
⎜ ⎟
⎝ 60 ⎠ and D is the 10 cent coin then the mean values
= Φ(1.480) = 0.931 of these samples (in the same order) are 50 35
30 50 35 30 35 35 15 30 30 15.
b) The sampling distribution for the mean of a
Sample mean 15 30 35 50
large sample uses the Central Limit Theorem
to allow calculation of probabilities using the 2 4 4 2
Probability
normal distribution as an approximation, but 12 12 12 12
for small samples the distribution is not known
and probabilities cannot be calculated. 6. E(X ) = 22.4 Var (X ) = 7.9
7 .9
a) Var ( X 20 ) = = 0.395
4. X ~ Po(7) ⇒ E(X ) = Var(X ) = 7 20
7 .9

approx
7 ⎞ b) Var ( X=
n) < 1 ⇒ n > 7.9
i) CLT ⇒ X 72 ~ N ⎜ 7, ⎟ n
⎝ 72 ⎠ so smallest sample size is 8.
 2

 6.506 94 − 7  7. a) F 6 ∼ N  505, 6.202 
ii) P( X 72 > 6.5) ≈
 P Z > = −1.5813  6 
 7 
77 ⎛ ⎞
⎜ 500 − 505 ⎟
= Φ(–1.5813) = 0.943 b) P( F 6 < 500) =
P⎜ Z < =
− 2.008 ⎟
⎜ 6.1 ⎟
⎜ ⎟
5. X ~ B(10, 0.3) ⇒ E(X ) = 3, Var(X ) = 2.1 ⎝ 6 ⎠
approx
⎛ 2.1 ⎞ = 1 − Φ(2.008) = 0.0223
i) CLT ⇒ X 80 ~ N ⎜ 3, ⎟
⎝ 80 ⎠ ⎛ 3.52 ⎞
8. a) M 10 ~ N ⎜ 252, ⎟
⎝ 10 ⎠
 3.25 − 3 
ii) P( X 80 > 3.25) ≈
 PZ > = −1.5816  ⎛ ⎞
 2.1 
  ⎜ 250 − 252 ⎟
80 P( M 10 > 250) =
P⎜ Z > =
− 1.807 ⎟
⎜ 3.5 ⎟
= 1 − Φ(1.5816) = 0.0569 ⎜ ⎟
⎝ 10 ⎠
= Φ(1.807) = 0.965

© Oxford University Press 2018: this may be reproduced for class use solely for the purchaser’s institute Sampling 4
250 − 252 b) Again, it is not a random sample, and
b) P( M n < 250) < 0.005 ⇒ < − 2.576
3 .5 attendance on a Monday morning may not be
n typical of the full week.
2.576 × 3.5
⇒ n > =
4.508 c) This is a systematic sampling – and on
2
production lines often one or more processes
⇒ n > 20.3 so the smallest sample size is 21. use a rotation of instruments; if one of
σ2 several instruments is faulty then defective
9. a) Mean = μ and variance = . batteries will occur at regular intervals in
n
b) If the underlying population is normal then the addition to any random faults.
distribution of sample means will be exactly 11. a) It would be far too time-consuming to
normal, and if n is large (> 30) it will be take a census of this information (even
approximately normal even if the underlying if it were possible to be sure of measuring
distribution is not normal. everyone).
10. Note that these answers are illustrative only – b) Any batteries that are tested are no longer
other reasons are possible. able to be used (testing to destruction).
a) It is not a random sample – if a train was
delayed all these people could be affected, and
if the train has not yet arrived, the response is
not the time they will wait for the train.

© Oxford University Press 2018: this may be reproduced for class use solely for the purchaser’s institute Sampling 5

You might also like