0% found this document useful (0 votes)
9 views7 pages

ST205 Summer 2018 Exam: Surveys & Experiments

Uploaded by

Kashif Ahmed
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views7 pages

ST205 Summer 2018 Exam: Surveys & Experiments

Uploaded by

Kashif Ahmed
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Summer 2018 examination

ST205
Sample Surveys and Experiments

Suitable for all candidates

Instructions to candidates

This paper contains four questions. Answer THREE of the four questions. Question 1 is com-
pulsory and two of the other three questions need to be answered. Question 1 carries 40
marks and the other questions carry 30 marks each, so the maximum achievable score is 100.

Time allowed - Reading Time: None


Writing Time: 2 hours
You are supplied with: Graph Paper
Formula Sheet (at the end of the examination paper)

You may also use: No additional materials


Calculators: Calculators are allowed in this examination

c LSE ST 2018/ST205 Page 1 of 7


1. [This question is compulsory.]
Answer the following questions:

(a) Explain briefly why randomisation is used in the design of experiments and how you
could check statistically to see whether randomisation has been implemented prop-
erly after an experiment has been conducted. [6 marks]

(b) Explain briefly by using two examples how measurement error may lead to bias in
surveys. [5 marks]

(c) Explain briefly what is meant by each of the following terms and why they might be
used in an experiment: double-blinding and blocking. [5 marks]

(d) Assume that N universities operate in a certain country. Explain briefly what steps
are needed to select a systematic sample of 12 universities. Also mention one circum-
stance where such a systematic sample may lead to an imprecise estimation. [6 marks]

(e) Discuss briefly the advantages and disadvantages of using (i) the internet and (ii) the
mail to collect data in surveys. [6 marks]

(f) A survey was conducted to estimate t, the total annual expense on campus for all
students across all universities. Assume that this amount is determined for a sample
of universities. Explain how ratio estimation can be used to estimate t if the variable
number of students in the university is known for all universities. Explain which
conditions would be desirable for the precision of the ratio estimator to be superior
to the precision of the unbiased estimator. [6 marks]
(g) An internet provider company wants to estimate the proportion P of its customers
who would be willing to pay more for expanded services. Suppose that a simple ran-
dom sample of n = 250 customers is selected from the population of N customers of
the company. Calculate the largest possible standard error for the estimated propor-
tion of the customers who would be willing to pay more for expanded services. (Note:
It is not necessary to know N .) [6 marks]

c LSE ST 2018/ST205 Page 2 of 7


2. A survey is conducted to investigate whether the curatorial decisions of art museums
in a certain country focus on enriching its cultural heritage through the acquisition of
contemporary art works. The aim is to estimate quantities such as the total number of
contemporary art works acquired during the last ten years by the art museums in this
country. A sampling frame of 207 museums is available and it is assumed to have an
excellent coverage of the population. Also, a stratified sample of 33 museums was selected
by using the size of the museum as the stratifying variable with 3 strata defined as: large
(h = 1), medium (h = 2) and small (h = 3) museums. The numbers of museums in these
strata, both at the population level, Nh , and the sample level, nh , are provided in the table
below, together with the sample mean, ȳh and the standard deviation, sh , of the number
of contemporary art works, y, bought by a museum during the last decade:

h Nh nh ȳh sh
1 60 10 37 8
2 56 8 22 12
3 91 15 13 7

(a) Assume that t is the total number of contemporary art works which were acquired
during the last decade by the museums in this country. Estimate t, the associated
standard error, and obtain a 95% confidence interval for t. [7 marks]

(b) Using the values reported in the table above determine the values of n1 , n2 and n3
for both
(i) proportional allocation, and [3 marks]
(ii) Neyman allocation. [4 marks]
(c) Assume that ȳU is the average number of contemporary art works which were ac-
quired during the last decade by the museums in this country. Estimate the gain
in precision measured by the reduction in variance of the estimator ȳU for Neyman
allocation compared to proportional allocation. [8 marks]

(d) Apart from the size of the museum the other variable which is available in the sampling
frame is the size of the budget available to the museum in the last ten years. Discuss
how suitable each variable is for stratification. Also, explain what other factors should
be considered for selecting the stratification scheme. [8 marks]

c LSE ST 2018/ST205 Page 3 of 7


3. A veterinary scientist wants to examine if obese dairy cows in a county are more prone
to ketosis than normal weight ones. To do this they survey a simple random sample of
10 farms in the county. On each of these farms the number of obese cows is counted and
their weight recorded. The data are given in the table below:

Farm, i Number of obese cows, Mi Total weight (kg), ti


1 13 10, 673
2 16 13, 132
3 9 7, 391
4 15 12, 323
5 18 14, 781
6 12 9, 854
7 10 8, 212
8 14 11, 493
9 20 16, 411
10 19 15, 590
Total 146 119, 860

(a) Assume that the number of obese cows in the county is 7, 300. Obtain the ratio
estimator of t, the total weight of obese cows in the county. Also, explain why the
ratio estimator may be biased. [6 marks]

(b) Suppose that there are 500 farms in the county. Obtain a standard error for your
estimate of t in (a) and an associated 95% confidence interval. [7 marks]

(c) Suppose that the veterinary scientist decided to modify this cluster sampling scheme
by subsampling 1 in 3 cows within sampled farms which are found to contain over
17 obese cows and selecting all obese cows in the remaining sampled farms. Com-
ment on how the Horvitz-Thompson estimation could be used to estimate t. [6 marks]

(d) Suppose that it was possible to obtain a simple random sample of the same number
of obese cows and to compute the associated unbiased estimator of t. Comment on
how you would decide between this estimator and the estimator of t in (a). [6 marks]

(e) The veterinary scientist adopts the unbiased estimator but their new objective is to
estimate t by designing a similar follow up survey so that the standard error is reduced
by one half. How many farms will be needed to achieve this objective? (Note: For
full marks take account of the finite population correction.) [5 marks]

c LSE ST 2018/ST205 Page 4 of 7


4. A higher education institution is interested in evaluating the scientific productivity of its
newly appointed academic staff with no more than four counted years of appointment. To
do this they select a simple random sample of 10 newly appointed members of the academic
staff and record both the number of years which they have been appointed and number
of articles which they have submitted for publication since their first day of appointment.
The results are as follows:

Number of Number of articles


Staff Member, i , xi , yi
years of employment submitted for publication
1 3 6
2 1 4
3 1 7
4 1 6
5 2 7
6 4 8
7 2 6
8 1 8
9 3 8
10 1 5
Total 19 65

The sample standard deviations are sx = 1.1 and sy = 1.35 and the sample correlation
coefficient is rxy = 0.49. It is also known that the institution has 80 members on its
academic staff with no more than four years of employment and that the population total
for x is tx = 152.

(a) Calculate a ratio estimate of the population mean, ȳU , together with an estimated
standard error. [7 marks]

(b) Calculate an unbiased estimate of the population mena, ȳU , together with an esti-
mated standard error. [7 marks]

(c) Based on both the observed pattern of (yi , xi ) values and the relative precision, dis-
cuss whether the ratio estimator or the unbiased estimator is to be preferred as an
estimator of ȳU . [6 marks]

(d) Discuss briefly the potential for bias in the ratio estimator and provide a justification
for your answer. [4 marks]

(e) Suppose that the previously defined sample of 10 observations was obtained by strat-
ified sampling with the strata defined by the four values of x. Estimate ȳU assuming
that you know that 50% of the 80 newly appointed members of the academic staff of
the institution are in their first year of appointment (x = 1), 20% are in their second
year (x = 2), 20% are in their third year (x = 3) and 10% are in their fourth year
(x = 4). [6 marks]

c LSE ST 2018/ST205 Page 5 of 7


Formulae Sheet for Reference

I. Simple Random Sampling (SRS)

Each subset of n out of N units has same chance of selection.


(1−f ) 2
1. E (ȳs ) = ȳU , V ar(ȳs ) = n S , f = n/N.

PN
2. t = i=1 yi , t̂ = N ȳs .

Pn
3. E s2 = S 2 , where s2 = 1
− ȳs )2 .

n−1 i=1 (yi

p(1−p)(1−f )
4. estimated variance of proportion p is n−1 .

II. Sampling with Unequal Probabilities

1 Pn yi
1. with replacement: t̂pps = n i=1 pi , where unit i is selected with probability pi at
each of n draws.

Pn yi
2. without replacement: t̂ppswor = i=1 πi , where πi = probability that population unit
i is selected into sample.

III. Stratified Random Sampling

Select SRS of nh out of Nh from stratum h, h = 1, ..., H. Let Wh = Nh /N, fh = nh /Nh .

2 (1−fh ) 2
PH PH
E s2h = Sh2 .

1. ȳstr = h=1 Wh ȳh , V ar (ȳstr ) = h=1 Wh nh Sh ,

2 (1−fh ) 2
PH
V ar t̂str = H
 P
2. t̂str = h=1 Nh ȳh , h=1 Nh nh Sh .

3. Proportional allocation: nh = nNh /N,


P 
H
Neyman allocation: nh = nNh Sh / h=1 h h ,
N S
General optimal
P allocation: choose nh to minimise √ V ar(ȳstr
P), subject to cost con-
√ 
straint c0 + H c n
h=1 h h = C, requires n h = n N S
h h / ch / H
h=1 N S
h h / ch .

IV. Ratio and Regression Estimation


Pn Pn PN
1. t̂yr = B̂tx , where B̂ = i=1 yi / i=1 xi , tx = i=1 xi .

2. b̄
y r = t̂yr /N.

c LSE ST 2018/ST205 Page 6 of 7


y r = (1−f ) 1 PN 2 (1−f )
Sy2 − 2BRSx Sy + B 2 Sx2 under SRS,
 
3. V ar b̄ n N −1 i=1 (yi − Bxi ) = n
where B = ty /tx , R = corr(y, x).
 2
y r = (1−f ) 1 Pn

V̂ b̄ n n−1 i=1 y i − B̂x i .

4. t̂yr more precise than t̂ = N ȳs if R > CV (x) / {2CV (y)} , where CV (x) = Sx /x̄U ,
CV (y) = Sy /ȳU .
 
5. t̂y,reg = N ȳs + B̂1 (x̄U − x̄) , where B̂1 = (yi − ȳ) (xi − x̄) / (xi − x̄)2 .
P P

V. Cluster Sampling

SRS of n out of N clusters with sizes Mi and totals ti , i = 1, ..., n.

N Pn s2
V̂ t̂unb = N 2 (1 − f ) nt ,

1. t̂unb = n i=1 ti ,
 2
1 Pn t̂unb
where s2t = n−1 i=1 ti − N .
Pn Pn
Pni=1 ti Pni=1 ti ,
PN
2. t̂rat = i=1 Mi , y rat =

i=1 Mi i=1 Mi
Pn 2
(t − y rat Mi )
Pn
1
where M̄ = n−1
i
 i=1

V̂ b̄
y rat = (1 − f ) nM̄ 2 n−1 , i=1 Mi ,
  PN 2
V̂ t̂rat = V̂ y rat
b̄ i=1 Mi .

3. two-stage sampling: use t̂unb or t̂rat as before, but with ti replaced by unbiased
estimator t̂i based upon sampled second-stage units in cluster i.

c LSE ST 2018/ST205 Page 7 of 7

Common questions

Powered by AI

Proportional allocation assigns sample sizes to strata in proportion to their overall population sizes. Using the museum example, nh would be selected based on nh = nNh/N. However, Neyman allocation considers both the size and variability (measured as standard deviation) of the strata, allocating more samples to strata with higher variability for efficiency. Therefore, Neyman allocation will result in a different nh distribution that minimizes overall variance, potentially differing substantially from proportional allocation especially in strata with varied variability .

Randomization is important in experiment design as it helps to eliminate bias by ensuring that each participant or experimental unit has an equal chance of receiving any treatment. This process increases the internal validity of the experiment by balancing out known and unknown confounders across treatment groups. To verify its implementation, one can conduct statistical tests such as comparing baseline characteristics across treatment groups. If the randomization is proper, these characteristics should not differ significantly .

To select a systematic sample of 12 universities, one needs to list all N universities in a random order. Then, calculate a sampling interval k as N/12, and choose a random start from the first k universities. Thereafter, select every k-th university from the start. This method may lead to imprecise estimates if there is a periodic pattern in the list that coincides with the sampling interval, causing certain types of universities to be over or underrepresented .

Internet surveys are cost-effective, quicker, and allow for easier data management and analysis. However, they may exclude respondents without stable internet access, leading to sampling bias. Conversely, mail surveys can reach a broader demographic base, including those who prefer traditional methods, but they are more expensive, slower, and have lower response rates .

For a sample size of 250, the largest possible standard error of the estimated proportion occurs when the proportion is 0.5. This can be calculated using the formula SE = sqrt(p(1-p)/n), where n is the sample size and p is the proportion. Hence, SE = sqrt(0.5*0.5/250) = 0.0316 .

Measurement error can introduce bias in surveys by distorting the true value of the variable of interest, leading to incorrect conclusions. For instance, if respondents inaccurately report their income either by underreporting or overreporting, the survey could result in biased estimates of average income. Similarly, if health survey respondents misreport their medical history due to forgetfulness, it could lead to incorrect associations between health conditions and lifestyle factors .

The ratio estimator is preferred when there is a strong linear relationship between the number of years employed and the number of articles submitted, as it exploits this correlation to increase precision. If the coefficient of variation of employment years is less relative to that of publications, it could yield more precise results than an unbiased estimator. However, if the correlation is weak or non-linear, the unbiased estimator might be more appropriate due to potential biases in the ratio estimator .

Ratio estimation can be used to estimate total expenses by using the known number of students as a correlated auxiliary variable. By calculating the ratio of expenses to student numbers in the sample and applying it to the total number of students, a more precise estimate of total expenses is obtained when the variables are correlated. The method's precision is enhanced if the auxiliary variable is highly correlated with the estimator and the coefficient of variation of the auxiliary variable is less than twice that of the main variable .

Double-blinding is a technique where neither the participants nor the experimenters know who is receiving a particular treatment. This reduces bias from both the subjects and investigators, as expectations about the treatment outcomes are minimized. Blocking involves grouping experimental units with similar characteristics to reduce variation, thus increasing the accuracy and efficiency of comparisons between treatments .

The Horvitz-Thompson estimator can effectively handle unequal probability sampling schemes, such as the modified cluster sampling where only some cows are subsampled. It adjusts for varying inclusion probabilities by weighting the sampled units accordingly, providing an unbiased estimate of the total weight. Its main advantage lies in its ability to mitigate biases introduced by uneven subsampling, thus offering more reliable estimates in complex designs .

You might also like