Probability & Statistics for CS Students
Probability & Statistics for CS Students
for
Computer Science Students
Demeke Lakew Workie (Associate Professor in Statistics)
BDU, College of Science, Statistics Program
Email: wadela1606@[Link]
September, 2025
BDU-Ethiopia
Collection,
by
General Procedures
✓ performing estimations & hypothesis
tests about a population, based on
Population Sample Data
information obtained from samples
✓ determining relationships among
variables &
Parameter Infer Statistic
✓ making predictions
the first step is to collect a set of related observations (data) from which
statistical conclusions may be drawn.
Data
1 a set of related information (facts)/ a real value of the variable
These are:
▪Classifications by Sources
▪Classification by the Role of Time
▪Classification by Scale of Measurement
Differences between
measurements, true zero &
ratio exists
e.g.: height, weight, time,
Ratio Data
salary, age
Quantitative Data
Differences between
measurements but no true zero
& ratio Interval Data
e.g.: IQ, T0C, BD
Qualitative Data
Categories (no ordering or
Nominal Data
direction)
e.g.: Gender, Eye color, nationality
Types of
1 Quantitative: It can be measured in the usual sense
Variables
2 Qualitative: Many characteristics are not capable of being measured.
Some of them can be ordered or ranked.
Types of
1 Discrete: is characterized by gaps or interruptions in the values that it
can assume.
quan Var
2 Continuous: can assume any value within a specified relevant interval
of values assumed by the variable.
Observation,
Questionnaire (mailed, telephone, face-to-face, self admin)
Marital status
3 Graphical
Married 2 Diagrammatic E.g.
6.8
19.5 Divorced
E.g. •Histogram
73.8
•Bar chart •Line graph
Not Married
• Pie chart •Box plot
Marital status
284
300
200
75
100 26
0
Married Divorced Not Married
1 Introduction
Tables: is an orderly and systematic presentation of numerical data in rows and columns.
Table 1: distribution of marital status for sample population
Income Level
Low Medium High Total
Single 50 100 50 200
Married 100 250 100 450
Marital Status
Divorced 30 40 30 100
Widowed 20 20 10 50
Total 200 410 190 800
1 Introduction
Frequency distributions
Frequency: is the number of observations belonging to a given value or a group.
Frequency distribution: is a table which contains the values and the corresponding frequencies.
There are three basic types of frequency distributions
frequency distribution in which the data is only nominal or ordinal &
Categorical Number of times that an event occur, is of our interest
a frequency distribution of numerical data in which the values are not grouped
Ungrouped table of all the potential raw score values along with the number of times each actually
occurred.
Example 3: consider the following data & construct grouped FD (Age of 80 adult male)
24 18 14 20 24 24 26 23 21 16
15 19 20 22 14 13 20 19 27 29
22 38 28 34 44 23 19 21 31 16 Since the number of observations are 80,
28 19 18 12 27 15 21 25 16 30 then :
k = 1+3.322(log 80 ) = 7.32 7,
17 22 29 29 18 25 20 16 11 17
R = 44– 10 = 34 &
12 15 24 25 21 22 17 18 15 21
w = 34 / 7 = 4.857 5
20 23 18 17 15 16 26 23 22 11
Then determine the class limits
16 18 20 23 19 17 15 20 10 23
1 Introduction
then the intervals will be in the form:
Now, determine the class boundaries The class mark is also calculated as
❖ UCB1 = UCL1 + ½*(d=1) =14 +1/2 = 14.5 ❖ m1 = ½*(UCL1 +LCL1) = ½*(UCB1 + LCB1) = 12.
❖ LCB1 = LCL1 - ½*(d=1) =10 - 1/2 = 9.5 etc. ❖ Count the frequency in each CL/CB
Then, the complete frequency distribution table with cumulative frequencies is:
CL CB Mi f <CF >CF
10 - 14 9.5 - 14.5 12 5 5 80
15 - 19 14.5 - 19.5 17 10 15 75
20 - 24 19.5 - 24.5 22 20 35 65
25 - 29 24.5 - 29.5 27 22 57 45
30 - 34 29.5 - 34.5 32 10 67 23
35 - 39 34.5 - 39.5 37 1 68 13
40 - 44 39.5 - 44.5 42 12 80 12
1 Introduction
2 Diagrammatic presentation:
These are techniques for presenting data in visual displays using diagrams.
Importance:
The most commonly used diagrammatic presentation for discrete as well as qualitative data are
Pie charts and Bar charts
Marital Status in %
Pie chart is a diagrammatic depiction 1.25
56.25
Each bar represent one category and its high is the frequency.
12.5 3.75
20 10
12.5 12.5 5 1.25
2.5
6.25 3.75 2.5
10 0
1.25 Single Married Divorced Widowed
0
Low Medium High
Single Married Divorced Widowed
It is building, joining the midpoint & frequency at the top of each bar of histogram by
line.
80
70
60
50
40
30
20
10
0
34.5 44.5 54.5 64.5 74.5 84.5
1 Introduction
Graphical Presentation:
the “less than” cfs are plotted against UCBs of their respective classes and they are joined by
lines adjacently.
The “more than” cfs are plotted against LCBs of their respective classes and they are joined
by lines adjacently.
Example: A chemist measures the efficiency (%) of a polymerization reaction for various vessel
temperature and pressures over time as efficiency in % is: 74 81 85 76 85 88 76 82 91.
Then, present the data by a line graph.
1 Introduction
Graphical Presentation:
Box-plot: is a visual description of the distribution based on the five summaries.
The five summaries are:
The pulse rates of 12 individuals arranged in
➢ Minimum increasing order are:
➢ Q1
62, 64, 68, 70, 70, 74, 74, 76, 76, 78, 78, 80
➢ Median
Q1=(68+70)2 = 69,
➢ Q3
M= (74+74)/2=74,
➢ Maximum Q3=(76+78)2 = 77
✓ Useful for comparing large sets of data
2 Measures of Central Tendency & Variation
2.1 Measure of Central Tendency
A data set contain many observations however, we are not always interested in each of the
measured values but rather in a summary which interprets the data.
Statistical functions fulfill the purpose of summarizing the data in a meaningful yet concise
way.
The most important statistical concepts to summarize interval and/or ratio data are measures
of central tendency and measures of variability.
The three most commonly used measures of central tendency are:
the most common measures of All MCT give us an
Mean idea about the
central tendency obtained by sum
of all X & divide by n. location where most
of the data is
MCT Median the value which divides the concentrated.
observations into two equal parts
Find arithmetic mean for the sample birth weights is computed as:
1 1 63338
𝑋ത = 20
σ 𝑋𝑖 = 20
(3265 + 3260 + ….+ 2834) = 20
= 3166.9 g.
2.1 Measure of Central Tendency
If X is a variable having values X1, X2,…, Xm occurring with frequencies f1, f2,…, fm
✓ k is number of classes,
✓ mi is the midpoint of the ith class &
✓ fi is the ith class frequency.
Example: the mean time spent by students for leisure activities is calculated as:
CL CB mi f
10 - 14 9.5 - 14.5 12 8 CL f
15 - 19 14.5 - 19.5 17 28 155 - 160 2
20 - 24 19.5 - 24.5 22 27
160 - 165 6
25 - 29 24.5 - 29.5 27 12
165 - 170 18
30 - 34 29.5 - 34.5 32 3
170 - 175 25
35 - 39 34.5 - 39.5 37 1
175 - 180 9
40 - 44 39.5 - 44.5 42 1
180 - 185 4
Affected by extreme values. Since all values enter into the computation.
𝐧
𝐗 𝟏 . 𝐗 𝟐 … . 𝐗 𝐧 , 𝐟𝐨𝐫 𝐮𝐧𝐠𝐫𝐨𝐮𝐩𝐞𝐝 𝐝𝐚𝐭𝐚 𝐬𝐞𝐭𝐬
GM =൞ 𝐦 𝐟 𝐟 𝐟
.
𝐗 𝟏𝟏 . 𝐗 𝟐𝟐 … . 𝐗 𝐦𝐦 , 𝐟𝐨𝐫 𝐝𝐚𝐭𝐚 𝐬𝐞𝐭𝐬 𝐨𝐟 𝐗𝐢 𝐡𝐚𝐯𝐢𝐧𝐠 𝐟𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐢𝐞𝐬 𝐟𝐢
Example: calculate the geometric mean for the following births per 1000 individuals
dataset. 7, 8, 3,14, 2, 1, 440, 15, 52, 6, 2, 1, 25, 12, 6, 9, 2, 1, 6, 7, 3, 4, 70, 20,
200, 2, 50, 21,15, 10, 120, 8, 4, 70, 3, 1,103, 20, 90, 1, 237
𝐧 42
GM= 𝐗𝟏. 𝐗𝟐 … . 𝐗𝐧, = 7 ∗ 8 ∗ 3 ∗ ⋯ 1 ∗ 237 = 0.999 = 10
2.1 Measure of Central Tendency
Harmonic Mean (HM)
is a suitable MCT when the data pertains to speed, rates & time.
n n
The HM is calculated as: HM = = or
1 + 𝟏 + …+ 𝟏 σ
𝟏
X𝟏 X𝟐 Xn Xi
n
HM = , where X1, X2, …Xk have the corresponding frequencies f1, f2,…fk
fi
σ
Xi
While if the data is grouped, mi‘s are considered as Xi.
Example: Milk is sold at a place with the rates of 1.8, 2, 2.25 & 2.5 birr per liter in four
different months. Then, find the average price paid per liter.
(1) Since n=20 is even, the average of 10th & 11th observation is (3245 + 3248)/2 = 3246.5 g
median
(2) Since n is odd, the 5th = ((9+1)/2)th observation, which is equal to 8 is median.
2.1 Measure of Central Tendency
Median:
𝐧
− 𝐂𝐅
෩=𝐋+
For a grouped dataset, median is calculated as: 𝐗 𝟐
∗ 𝐰, where
𝐟
L = LCB of the interval containing the median class,
w = class width,
n = total frequency of the sample,
CF = Cumulative frequency of all interval below L,
f = Frequency of the interval containing the median.
Note: the median class is first class whose cumulative frequency is at least n/2.
Example: find the median for the following grouped data.
CL CB mi f
10 - 14 9.5 - 14.5 12 8 Less than cf=8, 36, 63, 75, 78, 79,80
15 - 19 14.5 - 19.5 17 28 n/2 = 80/2 =40, then the median class is
20 - 24 19.5 - 24.5 22 27 3rd Therefore, L= 19.5, CF = 36, f=27 &
25 - 29 24.5 - 29.5 27 12 w = 5.
30 - 34 29.5 - 34.5 32 3 40 − 36
35 - 39 34.5 - 39.5 37 1 19 5 + 5 = 20.241
27
40 - 44 39.5 - 44.5 42 1
2.1 Measure of Central Tendency
3 The Mode: is the value of the observation that occurs with the greatest frequency.
Example: Find the mode for the following data:
(a) 22, 66, 69, 70, 73
(b) 1.8, 3.0, 3.3, 2.8, 2.9, 3.6, 3.0, 1.9, 3.2, 3.5
(c) 10, 10, 9, 9, 8, 12, 15, 5 .
Solution Thus, distributions with one mode are called unimodal,
(a) No mode those with two modes are called bimodal, & those with
(b) 3.0 more than two modes are called multimodal.
(c) 9 and 10
∆𝟏
The mode for grouped data is calculated as: =L+
𝐗 * w, where
∆𝟐 + ∆𝟏
L = Lower class boundary of the modal class; where modal class is a class that has max f.
w = the class width;
∆1 = 𝑓 − f1 & ∆2 = 𝑓 − f3 ;
f =frequency of the modal class, & f1 = frequency of the class immediately preceding the
modal class, & f3= frequency of the class immediately succeeding the modal class.
2.1 Measure of Central Tendency
The Mode:
Example: Find the mode for the following data:
Is distribution
Yes Use Median
skewed?
No
Use Mean
2.1 Measure of Central Tendency
4 The Quantiles:
Quantiles are dividing the distribution of ordered values/data into equal-sized parts.
✓ Quartiles: 4 equal parts
✓ Deciles: 10 equal parts
✓ Percentiles: 100 equal parts
Quartiles: Split Ordered Data into 4 Quarters
The first and third quartiles (denoted Q1 and Q3) are defined as follows:
❖ 25% of the data lie below Q1 (and 75% is above Q1),
❖ 25% of the data lie above Q3 (and 75% is below Q3).
( Q1 ) ( Q2 ) ( Q3 )
first quartile 2nd quartile third quartile
Median
2.1 Measure of Central Tendency
The Quantiles:
k(n + 1)𝑡ℎ
formula to get for ungrouped data the Kth quartile is: Qk = , where k=1,2,3.
4
c ( kn − C𝐹)
For a grouped data the kth quartiles can be done: Qk = L + 4
, k =1, 2, 3, where
f
CF = the less than cumulative frequency corresponding to the class immediately preceding
the kth quartile class
N.B.: Deciles and Percentiles can be computed in the same fashion as quartiles by k=1,2,3, …9
and k=1, 2, 3, …99, respectively.
2.2 Measure of Variation/Dispersion
The location measures may not be adequate enough to describe the distribution of the
data.
The Variation/dispersion of observations around any particular value is another property
which characterizes the data and its distribution.
7 7
For instance,
7 8 3 2
Consider the 7 77
following data
7 77 7 8 13
8 7
6 9
Mean = 7 Mean = 7
Mean = 7
Thus, MV/MD gives information on the spread or variability of the data values.
2.2 Measure of Variation/Dispersion
◼ There are different types of measures of variability.
◼ 2. Sample Variance
◼ The sample variance, s2, is the arithmetic mean of the squared deviations from
n
the sample mean: (
ix − x )2
s 2 = i =1
n −1
>
◼ 3. Sample Standard Deviation
◼ The sample standard deviation, s, is the square-root of the variance
n
(xi − x )
2
◼s has the advantage of being in the
i =1
s=
n −1 same units as the original variable x
2.2 Measure of Variation/Dispersion
◼ Variance and Standard Deviation for Grouped Data
The calculation is the same to the formula of data given in frequency distribution except that Xi is
substitute by mi. that is: n
f i (mi − x )
2
s= i =1
n −1
4 Coefficient of Variation (CV):
The coefficient of variation (CV) or relative standard deviation (RSD) is the sample standard
deviation expressed as a percentage of the mean, i.e.
s
CV = 100%
x
The CV is not affected by multiplicative changes in scale
Some times it is also important to compare the variation of the datasets if they have different
mean in magnitude
2.2 Measure of Variation/Dispersion
4 Skewness & Kurtosis
▪ Skewness is a measure of the asymmetry of
the probability distribution.
▪ Roughly speaking, a distribution has positive
skew (right-skewed), Normal, & negative
skew (left-skewed).
▪ Skewness is computed as (mean–mode)/SD
= -
2.2 Measure of Variation/Dispersion
Example: consider the following datasets and find
(a) The sample variance & Standard deviation
(b) the coefficient of variation
(c) the Skewness
(d) Compare the variability of the datasets
<C CL f
CL CB Mi f >CF
F
155 -160 2
10 - 14 9.5 - 14.5 12 5 5 80
15 - 19 14.5 - 19.5 17 10 15 75 160 - 165 6
Randomness & uncertainty exist in our daily lives as well as in every discipline, & hence
Probability is a mathematical framework that allows us to describe and analyse random
phenomena.
Random phenomena we mean?
events or experiments whose outcomes we can not predict with certainty.
Thus, Probability uses the language of sets &
a set is a collection of things (elements) which denoted by a capital letters like A, B, C….
For example, to define a set A that consists of the two elements and ♢ is, A = { ,♢}.
Solution:
Sample space
(a) E1 = {(6,1), (5,2), (4,3), (3,4), (2,5), (1,6)}
(b) E2 = {(6,5), (5,6)}
(c) E3 = {(1,1), (2,1), (1,2)}
(d) E4 = {(6,6)}
Intersection: the intersection of two sets A and B, denoted by A∩B, consists of all
elements in both A and B.
Example: If the universal set is given by S = {1, 2, 3, 4, 5, 6}, and A = {1, 2}, B = {2, 4, 5},C =
{1, 5, 6} are three sets, find the following sets:
(a) A ∪ B (b) A ∩ B (c) Ᾱ (d) A ∪ B ∪ C (e) A ∩ B ∩ C (f) B-A
2. A construction tower crane can operate to height H of 400 ft, a range (radius) R of 60 ft, and an
angle of ± 90o.
What is the sample space of operation of the crane?
Sketch the sample space and also the following event A: 0<H<80 and 0 30
Example: A contractor operates three concrete pumps. A pump is either operational (O) or not
operational (N). Then list all possible sample points states of the pump occurred.
Solution: Because each pump can have two states (either O or N), the sample points of all the
possible states of the pump is listed use tree diagram.
𝒏!
Moreover, we can rewrite nPr in terms of factorials as nPr =
𝒏–𝒓 !
Example: We assume that the bridge is supported by 9 cables, & the failure of 3 cables results in the failure of the
bridge, then what is the number of permutations of 3 out of 9 that can result in bridge failure?
9! 9! 9𝑥8𝑥7𝑥6!
Solution: n= 9 & r=3, then the number of permutations is 9P3 = = 6! = = 9x8x7 = 729 .
9 –3 ! 6!
Remark: If a set consists of n objects of which n1 are of one type, n2 are of a second type, . . . , nk are of a kth type.
𝒏!
Then the number of different permutations of the objects is given by: .
𝒏𝟏!𝒙𝒏𝟐!𝒙𝒏𝟑!𝒙…𝒙𝒏𝒌!
Example: if you consider the bridge example above, how may combinations that the bridge failed in the result of
the 3 cables out of 9?
Exercise:
1. How many horizontal flags can be formed using 3 colors out of 5 when
a) Repetition is allowed? b) Repetition is not allowed?
2. A college team plays 10 football games during a season. In how many ways can it end the season with five wins,
four losses, and one tie?
Frequentist Definition: the probability of an event is simply the “long-run proportion” of times
that the event occurs under many repetitions of the experiment.
Example: when Mendel conducted his famous hybridization experiment with peas, one such
experiment resulted in offspring consisting of 428 peas with green pods and 152 peas with yellow
pods.
Therefore, the probability of yellow pods is 152/580 = 0.262
Example: I roll a fair die. Let A be the event that the outcome is an odd number &
let B be the event that the outcome is less than or equal to 3, then
(a) What is the probability of A?
(b) What is the probability of A given B?
Example: Let I pick a random number from {1, 2, 3, ⋯ , 10}, and call it N. Suppose that all
outcomes are equally likely. Let A be the event that N is less than 7, and let B be the
event that N is an even number, then are A and B independent?
Similarly,
✓ If B1, B2, B3,… form a partition of the sample space S, and A is any event with p(A) > 0,
we have
Example: consider the marble example above, and suppose we observe that the chosen
marble is red, then what is the probability that Bag 1 was chosen?
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
Chapter 4: Random Variables & Probability Distribution
Expected Value, Variance & SD: Let X be a discrete random variable with PMF of p(x)
then,
✓ the expected value of X, denoted by E(X) = µx, is defined as:
𝑥𝑝 𝑋 = 𝑥 = 𝑥𝑝 𝑥
𝑥 𝑥
𝑥 − µx 2𝑝 𝑋 = 𝑥 = 𝑥 − µx 2𝑝 𝑥 = 𝑬 𝑿𝟐 − µ2x
𝑥 𝑥
SD(X) = σ2 = 𝑉𝑎𝑟 𝑋
Example 3: consider the experiment of tossing a fair coin three times, and let X be
defined as the number of heads I observed, then construct its pmf & CDF; check the
properties of pmf & CDF; and find the expected value, variance and SD of X.
Expected Value, Variance & SD: Let X be a continuous r.v. with pdf of f(x) then,
✓ the expected value of X, denoted by E(X) = µx, is defined as:
∞
න 𝑥𝑓 𝑥 𝑑𝑥
−∞
SD(X) = σ2 = 𝑉𝑎𝑟 𝑋
Tossing a coin 20 times to see how many tails occur. probability distribution (b) What is the likelihood that 5
Asking 200 people whether they watch ETV news. will be broken? (c) What is the likelihood that they will
Example: Suppose it is known that in a certain population 10 percent of the population is color blind. If a random
sample of 25 people is drawn from this population, then construct the probability distribution and find
a) The mean and SD.
b) The probability that two or fewer will be color blind.
c) The probability that between six and eight inclusive will be color blind.
Here x is the number of times an event occurs in an interval independently and the average rate at which events
occur is constant called 𝜆 .
If X is a Poisson distribution, then
✓ The expected value (mean) of X is 𝜆
✓ The variance of X is 𝜆
Note that
✓ 𝜆 is the average number of occurrences of the random event
✓ An interesting feature of Poisson distribution is the fact that the mean = variance and it is the only
distribution.
✓ The occurrence of events are independent.
✓ Theoretically, an infinite number of occurrences of the event must be possible in the interval.
2. Suppose you own a coffee shop, and based on your historical data, you know that the average
number of customers arriving per hour is 15. then
a) Construct the pmf of X
b) Find the probability that 10 customers arrive in an hour
c) Find the probability that 10 to 12 customers arrive in an hour.
This bell-shaped curve provides an adequate model for the relative frequency distributions of data
collected from many different scientific areas.
2
1 x−
A continuous r.v. X is said to have a normal distribution, if its pdf is given by: f ( x) = 1 e − 2
,− x
2
If X is a normal distribution, then
To compute the probabilities associated with the normal distribution, we use the standard normal table.
If X is a normal random variable with the mean μ and variance σ then the variable Z = (X - µ)/σ is the standardized
normal random variable. In particular, if μ = 0 and σ = 1, then the density function is called the standardized normal
density .
The equation for the standard normal distribution is written as:
z2
1 −
f ( x) = e 2
,− z
2
Characteristics of the Standard Normal Distribution
➢The highest point occurs at μ=0.
➢It is a bell-shaped curve that is symmetric about the mean, μ=0.
➢The total area under the curve equals one.
➢Empirical Rule:
✓ Approximately 68% of the area under the curve is between -1 and +1.
0
✓ Approximately 95% of the area under the curve is between -2 and +2.
✓ Approximately 99.7% of the area under the curve is between -3 and +3.
Answers
p(Z <2) =0.5 + P( 0<Z < 2)= 0.5 + 0.4772 = 0.9772
-1.74 1.53
P(-1.74 < Z < 1.53) = 0.4591 + 0.4370 = 0.8961.
P(Z > 1.71) =0.5 – 0.4564 = 0.0436.
Example1. Suppose that X N (165, 9), where X = the breaking strength of cotton
fabric. A sample is defective if X<162. Find the probability that a randomly
chosen fabric will be defective.
2. Let the height of BDU student is normally distributed with mean of 170 cm and
the standard deviation is 5 cm, then find the probability that a randomly
1.71
chosen student will between 165 and 160 cm.
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
4.3 Probability Distributions
6. Exponential Distribution: The exponential distribution is one of the widely used continuous distributions.
The exponential distribution is one of the widely used continuous distributions and it is often used to model the
time elapsed between events.
A continuous random variable X is said to have an exponential distribution with parameter λ > 0, if its pdf is given by:
f ( x ) = e − x , x 0 E( X ) =
1
,V ( X ) =
1
2
If X is a exponential distribution, then
Example 1: Let X ∼ Exponential(2), then construct the pdf of X, find the mean and variance of X.
Why sampling?
Get information about large populations
Less costs
Less field time
More accuracy i.e. Can Do A Better Job of Data Collection
When it is impossible to study the whole population (In some cases, it might not be possible to check
100%, like blood test, tasting food)
Select the required number of study units, using a “lottery” method or a “table of random numbers”.
"Lottery” method
it may be possible to use the “lottery” method for a small population.
Each unit in the population is represented by a slip of paper, these are put in a box and mixed, and a
sample of the required size is drawn from the box.
Example: Suppose that you have a population of 100 patients. From this, you want to draw a
sample of 5 patients.
0 8 4 2 5 7 9 5 4 1 2 5 6 3 2 1 4 0
5 8 2 0 3 2 0 5 4 7 8 5 9 6 2 0 2 4
3 6 2 3 3 3 2 5 4 7 8 9 1 2 0 3 2 5
9 8 5 2 6 3 0 1 7 4 2 4 5 0 3 6 8 6
2 9 0 1 1 2 3 4 5 6 8 7 9 4 3 2 3 4
B. Systematic Sampling
A list of N elements in the population is compiled, ordered according to
a specified variable,
• A sampling size n is chosen,
• A systematic step of k=N/n is set,
• A random number i between 1 and k is taken randomly and
represents the first element to be included,
• Then the other elements selected are i+k, i+2k, i+3k…till n (sample
size). –Less representative (biased) if the
–Cheaper and easier than SRS
order is cyclical
–More representative if order is related
to the interest variable (monotone)
–Sampling frame not always necessary
Systematic sampling
D. Cluster sampling
Section 3
Section 5
Section 4
Probability & Statistics for CS by
Demeke L.
(wadela1606@[Link])
5.2 Types of sampling
E. Convenience sampling
• is used in exploratory research where the researcher is
interested in getting an inexpensive approximation.
G. Quota sampling
First identify the stratums (such as sex, age…) and their
proportions as they are represented in the population.
H. Snowball Sampling
A first small sample is selected randomly
• Respondents are asked to identify others who belong to
the population of interests
• The referrals will have demographic and psychographic
characteristics similar to the referrers
–Lower costs –Inference is not possible
–Low variability
–Useful for “rare” populations
Dealing with such situations is the subject of the field of statistical inference.
Thus, Statistical inference is a collection of methods that deal with drawing conclusions from data that are prone
to random variation.
Random Sampling is a key technique for obtaining a representative sample and making valid inferences about the
population.
The main idea behind random sampling is to ensure that every member of the population has an equal probability
of being included in the sample.
This helps to minimize bias and increase the likelihood that the sample is a good representation of the population.
Sampling distribution provides information about the distribution of sample statistic (e.g. sample Mean, sample
Variance, sample proportion), and
It plays a crucial role in hypothesis testing, confidence intervals, and other inferential statistical techniques.
They provide a framework for making statistical inferences about population parameters based on sample data.
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
5.3.1 Introduction to Mean of the Sample Mean
Statistic: is a function of ‘observable’ random variables, which does not contain any unknown
parameters.
1
Example: If X1,…,Xn is a random sample (provided X1, …, Xn are observable), then 𝑋ത𝑛 = σ 𝑋𝑖 is
𝑛
statistic.
Thus, the sampling distribution of the statistic is the tool that tells us how close is the statistic to the
parameter.
As we begin to use sample data to draw conclusions about a wider population, we must be clear
about whether a number describes a sample or a population.
Histogram of 𝑋ത
mean).
This is because:
=𝜇
Notice that 𝜎𝑋ത is smaller than σx.
The larger the sample size the smaller 𝜎𝑋ത . Therefore, 𝑋ത tends to fall closer to μ, as the sample size
increases.
Thus, if a population is normal with mean μ and standard deviation σ,
the sampling distribution of 𝑋ത is also normally distributed with 𝜇𝑋ത = 𝜇 &
𝜎
standard deviation (standard error of the mean) 𝜎𝑋ത = , which measures how the sample statistic
𝑛
𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐−𝐸(𝑋)
That is, the random variable 𝑍 = converges in distribution to the standard normal.
𝑣(𝑋)
𝑋ത −𝜇
Examples: Let's assume that X 's are distributed with mean 0.5 and variance 1/12, then 𝑍 = 𝜎 gets closer to the
ൗ 𝑛
An interesting thing about the CLT is that it does not matter what the distribution of the Xi 's is, e.i., Xi 's can be
Probability & Statistics for CS by Demeke
discrete, continuous. L. (wadela1606@[Link])
5.4 Central limit theorem
Law of large numbers (LLN)
The law of large numbers (LLN) basically states that the average of a large number of i.i.d. random
variables converges to the expected value (which did not needed the normality assumption).
The weak law of large numbers (WLLN) is the main versions of the law of large numbers.
Let X1, X2 , ... , Xn be i.i.d. random variables with a finite expected value E(Xi) = μ < ∞.
𝑉 𝑋ത
That is, lim 𝑃 |𝑋ത − 𝜇 ≥∈ ≤ , by Chebyshev's Inequality
𝑛→∞ ∈2
𝜎2
= = 0, which goes to zero as 𝑛 → ∞
𝑛∈2
Estimation
LCL UCL
PE - (RF)*(SE) PE + (RF)*(SE)
PE
Recall that the general formula for all confidence intervals is: PE± RF *SE
The value of the reliability factor depends on the desired level of confidence
The confidence interval for µ is constructed based on σ is known and unknown, while for π
assume n is large.
Intervals Estimator
Population Proportion
Population Mean (assume n is large & estimate π if it
is unknown)
σ2 Known σ2 Unknown
(n is large/small) (We use Z-distn),if n is large
(We use Z-distn) other wise t-distn)
The width of the confidence interval estimate is a function of the confidence level, the
population standard deviation, and the sample size.
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
1 − = .95
α α
= .025 = .025
2 2
Confidence
Confidence
Coefficient, Z/2 value
Level
1−
80% .80 1.28
90% .90 1.645
95% .95 1.96
98% .98 2.33
99% .99 2.58
99.8% .998 3.08
99.9% .999 3.27
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
Example: The population consists of survival times of cancer patients who have
been treated with a new drug has SD of 43.3 months.
If a random sample of 100 drug-treated patients has a mean survival time of 46.9
months, then
a) What is the point estimate of the population mean?
b) Find a 95% confidence interval for the population mean.
Solution:
a) The point estimates of the population mean is 46.9 months.
b) Sigma is known and n is large, the 95% CI for µ is:
𝜎
𝑋ത ± 𝑍𝛼Τ 2
= 46.9 ± 1.96*43.3/10 = 46.9 ±8.5
𝑛
Thus, the 95% CI for µ is 38.4 < µ < 55.4
Therefore, with 95% confidence the mean survival times of patients in the
population is between 38.4 Probability
months& and 55.4
Statistics months.
for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
where tn-1,α/2 is the critical value of the t distribution with n-1 df and an
area of α/2 in each tail are given
Probability below.
& Statistics for CS by Demeke
L. (wadela1606@[Link])
6.3 Confidence Intervals for µ
Confidence t t t Z
Level (10 d.f.) (20 d.f.) (30 d.f.) ____
Note: t Z as n increases
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
Example: A medical researcher takes the blood pressure from a random sample of 25 (50-
year-old) women and the mean blood pressure of these sampled women is 140 mm Hg
with a standard deviation of 10 mmHg. Then
a) What is the point estimate of the mean blood pressure of all 50-year-old women?
b) Construct the 95% CI for the mean blood pressure of all 50-year-old women.
Solution:
a) The point estimate of the mean blood pressure of all 50-year-old women is 140 mmHg.
b) d.f. = n – 1 = 24, so t n-1, α/2 = t 24, 0.025 = 2.0639
Therefore, with 95% confidence the mean blood pressure of all 50-year-old women is between
Probability & Statistics for CS by Demeke
135.87mmHg and 144.13mmHg. L. (wadela1606@[Link])