SELF-LEARNING MODULE
STATISTICS AND PROBABILITY
Quarter 3 | Module 6 | S.Y. 2020 – 2021
TEACHER: MARK JASER R. AQUINO, CoE
I. OBJECTIVE
Identify sampling distribution of the mean.
M11/12SP-IIId-4
Find the mean and variance of a sampling distribution of the mean.
M11/12SP-IIId-5
II. SUBJECT MATTER
Sampling Distribution of the Mean
The Mean and Variance of a Sampling Distribution of the Mean
III. LEARNING RESOURCES
Textbook: Seeing the World Through Statistics and Probability by Enriqueta
D. Reston, et. al., Soline Publishing Company Inc.
IV. LESSON DISCUSSION
The Sampling Distribution of a Statistic
In previous examples, we repeatedly selected random samples of 10 observations on the ages
of 100 Filipino youth delegates to the 2016 WYD as our population of interest. When we sample
repeatedly and measure a variable of interest, there are actually three distinct distributions
involved, namely:
1. The population distribution
2. The distribution of sample data
3. The sampling distribution of the statistics
In that example, our variable of interest is the age of Filipino youth delegates to the 2016 WYD.
The data in Table 1 constitute the population distribution of ages (in years).
NVCFI – Statistics and Probability | Module 6
1
EXAMPLE 1
Suppose we define a smaller population of Filipino delegates to the 2016 WYD with ages
between 15 to 19 years old from the data in Table 1. We find that there are only N = 30
delegates who belong to this newly defined population and a random sample of 4 delegates will
be selected who will be lucky to have audience with the Pope. There are many possible
samples of size 4 that can be selected from this population of 30 delegates. The data in Table 1
(Module 5) shows the distribution of ages for 20 random samples of 4 youth delegates selected.
The data for each sample (represented by a row of the table) illustrates the second type of
distribution mentioned earlier, that is, the distribution of sample data which allows the values of
the variable (ages) for all the individuals in the sample.
The last column of Table 2 shows the mean age for each sample recorded in one decimal place
accuracy. The values of the sample means depend on the sample selected and they vary from
sample to sample. If the samples are randomly selected, we see that nothing influence the
variation of the sample means except chance alone. The possible values of the sample means
and their frequency of occurrence is summarized in Table 2. These results show that the mean
age of the delegate in the samples vary from 16.0 years to 18.5 years. The frequency column
tells us the number of samples with the corresponding value of the sample mean and the
relative frequency column corresponds to the proportion of samples (relative to the total number
of samples) having the specified value of the sample mean.
Table 2
Distribution of Ages (years) and Sample Means of 20 Possible Random Samples of Size 4
selected from 30 World Youth Day Filipino Delegates, Aged 15 – 19
Ages (years) of WYD Filipino Youth Delegates selected Sample
Sample No.
Y1 Y2 Y3 Y4 Mean
1 15 18 15 16 16.0
2 16 15 18 17 16.5
3 15 16 19 18 17.0
4 19 15 17 17 17.0
5 15 17 16 18 16.5
6 16 18 19 17 17.5
7 15 17 16 18 16.5
8 19 17 18 16 17.5
9 15 17 18 16 16.5
10 17 18 18 19 16.5
11 17 18 16 19 17.5
12 15 18 19 16 17.0
13 15 16 18 19 17.0
14 17 16 18 17 17.0
15 17 18 19 18 18.0
16 15 17 19 17 17.0
17 18 16 17 19 17.5
18 18 16 15 17 16.5
19 18 16 19 19 18.0
20 19 18 19 18 18.5
NVCFI – Statistics and Probability | Module 6
2
Table 2
Relative Frequency Distribution of Sample Means for the Ages of the Youth Delegates Selected
in Samples of Size n = 4
Mean Age Frequency
Relative Frequency
(years) (No. of Samples)
16.0 1 0.05
16.5 5 0.25
17.0 6 0.30
17.5 4 0.20
18.0 3 0.15
18.5 1 0.05
Total 20 1.00
Note that we have 20 random samples of size n = 4. The data in Table 1 may be also
represented graphically in a histogram as shown in Figure 1 below. Recall that a histogram is a
special type of bar graph for continuous data where there are no gaps or spaces between the
bars.
Figure 1
Distribution of Sample Means on the Ages of Youth Delegates (years) for 20 Random Samples
of Size 4.
A sampling distribution of the mean X is the probability distribution of the sample means
obtained through a large number of samples drawn from a specific population. It may be
represented in a table, graph or a formula which shows the distribution of a range of different
values of X and their associated probabilities of occurrence.
Based on the foregoing example, a sampling distribution of the mean X can be thought of as a
relative frequency distribution of the X sample mean with a very large number of samples. The
more samples we take, the closer the relative frequency distribution approaches the sampling
distribution as we increase the number of samples infinitely.
While it is easier to visualize a distribution of data (whether from a population or sample)
corresponding to ages, heights, weights, or test scores using tables and graphs, visualizing a
sampling distribution of the mean is more difficult since it entails drawing all possible samples of
a given size, say n, from a larger or infinite population. In reality, we do not necessarily draw all
random samples of a population and study the distribution of the sample means. In practice,
such as in sample surveys and opinion polls, or in conducting statistical experiments, the
researcher or investigator takes only one or a few samples and use the sample information as
NVCFI – Statistics and Probability | Module 6
3
basis for generalizing results. Had the process of selecting respondents for a sample survey
been repeated, a new set of respondents would yield a new set of sample statistics. Thus, these
sampling distributions are hypothetical and abstract. Theoretically, they exist if the means (or
any sample distributions are hypothetical and abstract. Theoretically, they exist if the means (or
any sample statistics) of all possible samples were obtained; empirically, they may be produced
through simulations. Think of the sampling distribution of the mean X as a distribution of
estimates of the population mean if we were to repeat the sampling process over and over
again.
The Mean and Variance of a Sampling Distribution of the Mean
The sampling distribution of the mean X has its own mean and variance. Since the sample
mean X is a random variable then the mean of the sampling distribution denoted by μ x is the
expected value of the random variable X i . From Module 1, we learned that the mean of a
random variable X is the expected value of X which is the sum of the products of the possible
values of Xi and their associated probabilities of occurrence P (Xi); that is,
n
μ= E ( X )=∑ X i ∙ P ( X i )= X 1 ∙ P ( X 1 ) + X 2 ∙ P ( X 2 ) +…+ X n ∙ P( X n )
i=i
By applying the formula where the sample mean X i is a random variable, the mean of the
sampling distribution of the mean X is given by the following formula.
Formula 1
Mean of the sampling distribution of the mean, μ x:
n
μ x =E ( X )=∑ X i ∙ P ( X i ) + X 2 ∙ P ( X 2 ) +…+ X n ∙ P(X n)
i=1
On the other hand, the variability of a sampling distribution of mean is measured by its variance
or its standard deviation. The value of the variance depends on three factors, namely:
The size of the population
The size of the sample
The method of selecting the samples from the population
By using the formula for the variance of random variable, we obtain the variance of the sampling
2
distribution of mean distribution of mean X , denoted by σ x , as follows:
Formula 2
Variance and standard deviation of the sampling distribution of the mean X
n
Variance: σ x =Var ( X ) =∑ [ X i ∙ P( X i) ]−μ x
2 2 2
i=1
Standard Deviation: σ x =√ Var ( X )
NVCFI – Statistics and Probability | Module 6
4
The standard deviation of a sampling distribution of the mean X is also called the standard
error of the mean and is often referred to as S.E.M. It measures the amount of chance
variability in the estimates of the mean that could be generated over all possible samples. We
want estimates with a small standard error (usually less than 1) since the chances are good that
the estimate will be close the true value of the population parameter.
EXAMPLE 2
Finding the mean and variance of a sampling distribution
Random samples of size 25 were drawn from a population of 100 teen-agers who were asked
on the estimated number of videos they download from YouTube on a given day. The table
below shows the sampling distribution of the mean X and their associated probabilities of
occurrence (in relative frequencies) for a large number of random samples of size 25.
x (mean number of video
0 1 2 3 4 5 6
downloads
0.4
P¿ 0.05 0.10 0.30 0.10 0.03 0.02
0
Find the mean and variance of the sampling distribution of X . Also, find its standard error of the
mean (SEM).
Solution
We can organize our solution better in a table form for the values needed in the formula for
finding the mean and variance, as follows:
Xi P( X i) X i ∙ P ( X i ) X 2i ∙ P ( X i )
0 0.05 0.00 0.00
1 0.10 0.10 0.10
2 0.30 0.60 1.20
3 0.40 1.20 3.60
4 0.10 0.40 1.60
5 0.03 0.15 0.75
6 0.02 0.12 0.72
Sum 1.00 2.57 7.97
For the mean of the sampling distribution of X , we use the formula:
n
μ x =E ( X )=∑ X i∗P ( X i) , where the summation value refers to the sum of 2.57 obtained in the
i=1
n
X i∗P ( X i )−column of thetable . Thus, μ x =E ( X )=∑ X i∗P ( X i) =2.57 ≈ 3 video downloads
i=1
Since we are counting number of video downloads which is a discrete quantity, this means that,
on the average, there are around 3 video downloads for the distribution sample means for a
large number of random samples of 25 teen-agers selected from the population.
For the variance of the sampling distribution of X , we use the formula:
n
σ x =Var ( X ) =∑ [ X i ∙ P( X i) ]−μ x
2 2 2
i=1
NVCFI – Statistics and Probability | Module 6
5
Where the summation value refers to the sum of X obtained by taking the product of the
squares of the means and their associated probabilities. This is the sum in the fourth column of
the table, which is 7.97. Thus, we compute the variance as shown below.
n
σ x =Var ( X ) =∑ [ X i ∙ P( X i) ]−μ x = (7.97 )−( 2.57 ) =1.365≈ 1.3
2 2 2 2
i=1
Since the standard deviation is the square root of the variance, then we obtain the standard
error of the mean, σ x , as follows:
σ x =√ Var ( X )= √1.365=1.168 ≈1
This means that on the average, the mean number of video downloads vary from sample to
sample by around 1.168 or roughly 1 video about the expected value or mean of 3 video
downloads.
Table 3 provides a summary of the notations used for the three different types of distributions
discussed in this chapter.
Table 3
Three Types of Distribution and Corresponding Notation
Sampling
Population
Sample Distribution Distribution of the
Distribution
Mean
The distribution of the The distribution of The distribution of
data corresponding data for a variable of values of the mean
Definition to values of a interest in sample obtained from
variable of interest in selected from a samples selected
the target population population from a population
Mean μ x μx
Variance 2
σ s
2
σ 2x
Standard Deviation σ s σx
NVCFI – Statistics and Probability | Module 6
6
STATISTICS & PROBABILITY | MODULE 6
TEACHER: MARK JASER R. AQUINO, CoE
NAME: ______________________________ GRADE & STRAND: ____________
CONCEPTUAL UNDERSTANDING
For each of the following situations, identify the most appropriate method of selecting a random
sample, and describe how you would go about the selection process. Write your answer on the
spaces given.
1. There are 1300 sixth-grade students in Kidapawan National School. Mrs. Palmares, an
English Teacher Coordinator is given the task of estimating their arithmetic achievement
by selecting a sample of 200 students who will take an arithmetic achievement test.
a. Define precisely the target population.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
b. If you were to advise Mrs. Palmares how to select the sample of 200 pupils, what
would be the most appropriate sampling procedure for her to use? Describe briefly
how would you sample from this population based on your suggested sampling
procedure.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
2. When the Department of Health (DOH) announced that they will be giving free COVID-
19 vaccinations to elementary pupils at North Valley College, 400 pupils enlisted with
parent’s consents on a specified day. But due to limited stocks of the vaccines, DOH
announced that they will administer the vaccine by batches and they will administer the
vaccines to the first batch of 80 randomly selected students from the list. The DOH
wanted to select the random sample based on the list using a regular interval. What
would be the most appropriate sampling procedure? Describe in further details how the
selection of these 90 students will be done.
______________________________________________________________________
______________________________________________________________________
______________________________________________________________________
______________________________________________________________________
______________________________________________________________________
______________________________________________________________________
______________________________________________________________________
NVCFI – Statistics and Probability | Module 6
7
Read in advance about The Central Limit Theorem.
NVCFI – Statistics and Probability | Module 6
8