0% found this document useful (0 votes)
9 views123 pages

Probability & Statistics for CS Students

The document outlines a course on Probability and Statistics tailored for Computer Science students, detailing its contents, objectives, and methods of data collection and presentation. It covers fundamental concepts such as measures of central tendency, probability theory, sampling techniques, and statistical inference. The course aims to equip students with the skills to organize, summarize, and analyze data effectively.

Uploaded by

andebetmelkam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views123 pages

Probability & Statistics for CS Students

The document outlines a course on Probability and Statistics tailored for Computer Science students, detailing its contents, objectives, and methods of data collection and presentation. It covers fundamental concepts such as measures of central tendency, probability theory, sampling techniques, and statistical inference. The course aims to equip students with the skills to organize, summarize, and analyze data effectively.

Uploaded by

andebetmelkam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Probability & Statistics for Engineers

for
Computer Science Students
Demeke Lakew Workie (Associate Professor in Statistics)
BDU, College of Science, Statistics Program

Email: wadela1606@[Link]

September, 2025
BDU-Ethiopia

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
Course Contents
Chapter 1. Introduction
– Definition & classifications
– Method of data collection and organization
– Method of data representation
Chapter 2. Measures of Central Tendency & Variation
– Measures of Central Tendency
– Measures of Variation
Chapter 3. Introduction to Elementary Probability Theory
– Revision on set theory
– Counting Rules and probability
– Random variables and probability distributions
Chapter 4. Common Probability Distributions and their properties
– Common Discrete Probability Distributions
– Common Continuous probability distribution
Chapter 5. Sampling Theory and Statistical Inference
– Sampling techniques
– Probability and non-probability sampling techniques
– Introduction to sampling distributions

Chapter 6. Statistical Inference about a single population mean


– Estimation of the population mean (μ)
– Hypothesis testing about a single population mean(μ)

Probability & Statistics for Data Science by


Demeke L. (wadela1606@[Link])
Objectives
At the end of the course, students would understand:
 to know definitions and concepts of statistics
 to organize and summarize data,
 to evaluate random variables & some common probability distributions,
 To understand concept sampling techniques
 To understand concept of statistical inferences
After this course you will be able to:
 Familiar with concepts of statistics,

 Manage, present & summarize your data,

 Familiar with RVs & some common probability distributions,

 Familiar with statistical inference, &

 Explore sampling methods

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
1 Introduction
1 Introduction
Statistics is the science of learning from data.
Commonly, the word statistics defined as numerical data/plural senses & a field
of study/singular senses.

aggregates of numerically expressed figures collected


Plural
in a systematic manner for a pre-determined purpose.

Statistics the science of collecting, organizing, presenting,


Singular analyzing & interpreting numerical data for making
wise decisions.

Hence, statistics is a procedural process performing the five major


activities on numerical data called Stages in statistical Investigation.

Collection,

What are Organization,


the 5 major Data Presentation,
activities?
Analysis
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link]) Interpretation
1 Introduction
Thus, statistics is divided into two main areas, depending on how data are used.

consists of the collection, organization,


Classification Descriptive summarization, and presentation of data.
of statistics
consists of generalizing from samples to
Inferential population

by
General Procedures
✓ performing estimations & hypothesis
tests about a population, based on
Population Sample Data
information obtained from samples
✓ determining relationships among
variables &
Parameter Infer Statistic
✓ making predictions

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
1 Introduction
In any statistical investigation,

the first step is to collect a set of related observations (data) from which
statistical conclusions may be drawn.

Data
1 a set of related information (facts)/ a real value of the variable

2 We have different types of data under different basis of classification.

These are:
▪Classifications by Sources
▪Classification by the Role of Time
▪Classification by Scale of Measurement

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
1 Introduction
Measurement Levels

Differences between
measurements, true zero &
ratio exists
e.g.: height, weight, time,
Ratio Data
salary, age
Quantitative Data
Differences between
measurements but no true zero
& ratio Interval Data
e.g.: IQ, T0C, BD

Ordered Categories (rankings, order,


or scaling) Ordinal Data
e.g.: Grade, Juging, rating scale,

Qualitative Data
Categories (no ordering or

Nominal Data
direction)
e.g.: Gender, Eye color, nationality

Statistics & Probability by Demeke L.


(wadela1606@[Link])
1 Introduction
In our discussion of statistics we have frequently used terminologies. The ff are some:
A collection, or set, of individuals or objects or events
Population whose properties are to be analyzed.

A subset of the population.


Sample

A number that describes a population characteristic


Basic Terms Parameter

A number that describes a sample characteristic


Statistic

Is a characteristic that takes on different values in


Variable
different persons, places, or things

Experiment A planned activity whose results yield a set of data

Data a set of related information (facts)/ a real value of the


variable
1 Introduction
There are quantitative & qualitative types of variables.

Types of
1 Quantitative: It can be measured in the usual sense

Variables
2 Qualitative: Many characteristics are not capable of being measured.
Some of them can be ordered or ranked.

Types of
1 Discrete: is characterized by gaps or interruptions in the values that it
can assume.

quan Var
2 Continuous: can assume any value within a specified relevant interval
of values assumed by the variable.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
1 Introduction
Uses, Applications & Limitations of statistics.
✓ It represents the facts in a definite form.
✓ It facilitates comparison of data.
Uses ✓ It helps in formulating and testing hypothesis and develops new theories.
✓ It gives guidance in the formulation of suitable polices.

✓ In Sociology: to examine patterns of social inequality & social networks,


mobility, & family dynamics
✓ In governmental activities: estimating and predicting of air pollution
Application
✓ In agriculture: experiments of crop yields, types of fertilizers & soils
✓ In production: to produce a chemical based on the demand
✓ In pharmaceutical areas: to test the effectiveness of a new drug

✓ It deals on aggregates of facts


✓ Statistical data are only approximately and not mathematically correct.
Limitations
✓ Statistical interpretations require a high degree of skill & subject matter
Probability & Statistics for CS by Demeke L.
knowledge (wadela1606@[Link])
1 Introduction
Method of Data Collection
Before any statistical work can be done data must be collected.
Depending on the objective of the study different data collection methods can be employed.
Various data collection techniques that can be used are:
 Experimental,

 Observation,
 Questionnaire (mailed, telephone, face-to-face, self admin)

 Focus group discussions (FGD), key informed individual (KII)


 Case studies, and money more.
1 Introduction
Method of Data Collection & Presentation
Data in raw form are not easy to use for decision making & should presented by app methods
Therefore, presentation of data is broadly classified in to three.
1 Tabular presentation

Nominal and/or ordinal Data Interval and/or ratio Data

Marital status
3 Graphical
Married 2 Diagrammatic E.g.
6.8
19.5 Divorced
E.g. •Histogram
73.8
•Bar chart •Line graph
Not Married
• Pie chart •Box plot

Marital status
284
300
200
75
100 26
0
Married Divorced Not Married
1 Introduction
Tables: is an orderly and systematic presentation of numerical data in rows and columns.
Table 1: distribution of marital status for sample population

Marital Status n Percent (%) Valid percent


Single 200 25 25
Married 450 56.25 56.25
Divorced 100 12.5 12.5
Widowed 50 1.25 1.25

Table 2: cross-tabulation of income levels by marital status for a sample population

Income Level
Low Medium High Total
Single 50 100 50 200
Married 100 250 100 450
Marital Status
Divorced 30 40 30 100
Widowed 20 20 10 50
Total 200 410 190 800
1 Introduction
Frequency distributions
Frequency: is the number of observations belonging to a given value or a group.
Frequency distribution: is a table which contains the values and the corresponding frequencies.
There are three basic types of frequency distributions
frequency distribution in which the data is only nominal or ordinal &
Categorical Number of times that an event occur, is of our interest

a frequency distribution of numerical data in which the values are not grouped
Ungrouped table of all the potential raw score values along with the number of times each actually
occurred.

We need to divide the data into groups or intervals or classes.


So, we need to determine:
✓ Class size: The number of intervals ( 5 ≤ k ≤ 20; k = 1 + 3.322 (log n),
✓ The range (R=Xmin - Xmax),
✓ The Width of the interval (w=R/k),
Grouped ✓ Class limit: lowest & highest values that can be included in a class,
✓ Class boundaries: UCB = UCL + ½*d & LCB = LCL – ½*d; d be the gap
between two successive classes,
✓ Class mark /class midpoint (mi): is the value located half way between
the lower and upper class limits of that class.
✓ <CF: the cumulation from the lowest size of the variable to the highest
size &
✓ >CF: the cumulation is from the highest to the lowest value.
1 Introduction
Examples of frequency distribution
Frequency: is the number of observations belonging to a given value or a group.
Frequency distribution: is a table which contains the values and the corresponding frequencies.
➢ Categorical Frequency Distribution

Example 1: distribution of marital status for sample population

Marital Status n Percent (%) Valid percent


Single 200 25 25
Married 450 56.25 56.25
Divorced 100 12.5 12.5
Widowed 50 1.25 1.25
1 Introduction
Examples of frequency distribution
➢ Ungrouped frequency distribution

Example 2: Distribution of frequencies in students of a certain school that have smoked at


least once. n=534 It is useful calculate cumulative frequency.
1 Introduction
➢ Grouped frequency distribution
For large samples, we can’t use the simple frequency table to represent the data.
We need to divide the data into groups or intervals or classes.

Example 3: consider the following data & construct grouped FD (Age of 80 adult male)
24 18 14 20 24 24 26 23 21 16
15 19 20 22 14 13 20 19 27 29
22 38 28 34 44 23 19 21 31 16 Since the number of observations are 80,
28 19 18 12 27 15 21 25 16 30 then :
k = 1+3.322(log 80 ) = 7.32  7,
17 22 29 29 18 25 20 16 11 17
R = 44– 10 = 34 &
12 15 24 25 21 22 17 18 15 21
w = 34 / 7 = 4.857 5
20 23 18 17 15 16 26 23 22 11
Then determine the class limits
16 18 20 23 19 17 15 20 10 23
1 Introduction
then the intervals will be in the form:
Now, determine the class boundaries The class mark is also calculated as

❖ UCB1 = UCL1 + ½*(d=1) =14 +1/2 = 14.5 ❖ m1 = ½*(UCL1 +LCL1) = ½*(UCB1 + LCB1) = 12.

❖ LCB1 = LCL1 - ½*(d=1) =10 - 1/2 = 9.5 etc. ❖ Count the frequency in each CL/CB

❖ Construct the < & > CFs

Then, the complete frequency distribution table with cumulative frequencies is:

CL CB Mi f <CF >CF
10 - 14 9.5 - 14.5 12 5 5 80
15 - 19 14.5 - 19.5 17 10 15 75
20 - 24 19.5 - 24.5 22 20 35 65
25 - 29 24.5 - 29.5 27 22 57 45
30 - 34 29.5 - 34.5 32 10 67 23
35 - 39 34.5 - 39.5 37 1 68 13
40 - 44 39.5 - 44.5 42 12 80 12
1 Introduction
2 Diagrammatic presentation:
These are techniques for presenting data in visual displays using diagrams.

Importance:

 They have greater attraction.

 They facilitate comparison.

 They are easily understandable.

 They have greater memorizing value than mere figures.

The most commonly used diagrammatic presentation for discrete as well as qualitative data are
Pie charts and Bar charts
Marital Status in %
Pie chart is a diagrammatic depiction 1.25

of data as slices of a pie. 12.5 25

56.25

Single Married Divorced Widowed


1 Introduction
Diagrammatic presentation: Bar chart is used with categorical or numerical discrete data.

Each bar represent one category and its high is the frequency.

Bar may Simple, Component or Multiple


Marital Status by Income level in %
Marital Status in % 60
60 56.25
50 12.5
50
40
40 30 31.25
30 25 20 6.25

12.5 3.75
20 10
12.5 12.5 5 1.25
2.5
6.25 3.75 2.5
10 0
1.25 Single Married Divorced Widowed
0
Low Medium High
Single Married Divorced Widowed

Marital Status by Income level in %


35 31.25
30
25
20
15 12.5 12.5 12.5
10 6.25 6.25
3.75 5 3.75 2.5 2.5 1.25
5
0
Single Married Divorced Widowed
Low Medium High
1 Introduction
3 Graphical Presentation:
A graphic presentations are the most commonly used devices for presenting a
continuous data. Of which
Histogram is useful to continuous data.
➢ Histogram Properties
➢ Polygon ➢ There are not spaces between bars,
➢ Ogive ➢ Width of the bar represent by w (from LCB to UCB),
➢ Box plot ➢ Height represent for frequencies.
➢ The center of the bar is represent by the midpoint
➢ Line graph

Example: Consider the ff grouped frequency data & draw histogram


80
70
60
50
40
30
20
10
0
34.5 44.5 54.5 64.5 74.5 84.5
1 Introduction
Graphical Presentation:

Frequencies polygon: is another form to show the frequency distribution of a numerical


data.

 It is building, joining the midpoint & frequency at the top of each bar of histogram by

line.
80
70
60
50
40
30
20
10
0
34.5 44.5 54.5 64.5 74.5 84.5
1 Introduction
Graphical Presentation:

Ogive: the cumulative frequencies graph is called Ogive Curves.

 Ogive can be “Less than” Ogive and “more than” Ogive.

 the “less than” cfs are plotted against UCBs of their respective classes and they are joined by
lines adjacently.

 The “more than” cfs are plotted against LCBs of their respective classes and they are joined
by lines adjacently.

Used to facilitate to read Q, D &


P
The intercession point of n/2 &
horizontal axis is M
1 Introduction
Graphical Presentation:
Line graph: is suitable for depicting a consecutive trend of a series over a long period.
✓ useful in showing variations in continuously changing variables.

Example: A chemist measures the efficiency (%) of a polymerization reaction for various vessel
temperature and pressures over time as efficiency in % is: 74 81 85 76 85 88 76 82 91.
Then, present the data by a line graph.
1 Introduction
Graphical Presentation:
Box-plot: is a visual description of the distribution based on the five summaries.
The five summaries are:
The pulse rates of 12 individuals arranged in
➢ Minimum increasing order are:
➢ Q1
62, 64, 68, 70, 70, 74, 74, 76, 76, 78, 78, 80
➢ Median
Q1=(68+70)2 = 69,
➢ Q3
M= (74+74)/2=74,
➢ Maximum Q3=(76+78)2 = 77
✓ Useful for comparing large sets of data
2 Measures of Central Tendency & Variation
2.1 Measure of Central Tendency
A data set contain many observations however, we are not always interested in each of the
measured values but rather in a summary which interprets the data.
Statistical functions fulfill the purpose of summarizing the data in a meaningful yet concise
way.
The most important statistical concepts to summarize interval and/or ratio data are measures
of central tendency and measures of variability.
The three most commonly used measures of central tendency are:
the most common measures of All MCT give us an
Mean idea about the
central tendency obtained by sum
of all X & divide by n. location where most
of the data is
MCT Median the value which divides the concentrated.
observations into two equal parts

Mode the value which occurs the most


compared with all other values
2.1 Measure of Central Tendency
1 The Mean: These are three types of Means which are suitable for a particular type of data
called Arithmetic Mean (Simply Mean), Geometric Mean and Harmonic Mean.
Arithmetic mean for ungrouped data of the sample, denoted by 𝑋ത is given as:
σ𝒏
𝒊=𝟏 𝑿𝒊
ഥ = 𝑿𝟏 + 𝑿𝟐+ …+𝑿𝒏 =
𝑿 .
𝒏 𝒏
Example: Suppose the sample consists of birth weights (in grams) of all live born infants at a
private hospital in a city, during a 1-week period. These samples are:
3265 3323 2581 2759 3260 3649 2841 3248 3245 3200 3609 3314
3484 3031 2838 3101 4146 2069 3541 2834.

Find arithmetic mean for the sample birth weights is computed as:

1 1 63338
𝑋ത = 20
σ 𝑋𝑖 = 20
(3265 + 3260 + ….+ 2834) = 20
= 3166.9 g.
2.1 Measure of Central Tendency
If X is a variable having values X1, X2,…, Xm occurring with frequencies f1, f2,…, fm

𝑿𝟏𝒇𝟏 + 𝑿𝟐𝒇𝟐 + …+𝑿𝒎 𝒇𝒎 σ𝒎


𝒊=𝟏 𝑿𝒊𝒇
ഥ=
respectively, then its arithmetic mean is given by: 𝑿 = 𝒊
.
𝒇𝟏 +𝒇𝟐 +⋯+𝒇𝒎 σ𝒎
𝒊=𝟏 𝒇𝒊

Example: Suppose the X values are 3, 5, 4, 2, 7 and 6 with corresponding frequencies


of 2, 1, 3, 2, 1 and 1 respectively.

Then the mean for the frequent data is computed as:


𝒎
ഥ = σ𝒊=𝟏
𝑿
𝑿𝒊
=
𝟑∗𝟐+𝟓∗𝟏+ …+𝟕∗𝟐 +𝟔∗𝟏 𝟒𝟎
= 𝟏𝟎 = 4.
σ𝒎 𝒇
𝒊=𝟏 𝒊 𝟐+⋯+𝟏
2.1 Measure of Central Tendency
σ𝐤𝐢=𝟏 𝐦𝐢𝐟
ഥ=
The mean for grouped data of the sample data is computed as: 𝐗 𝐢
, where
σ𝐤𝐢=𝟏 𝐟𝐢

✓ k is number of classes,
✓ mi is the midpoint of the ith class &
✓ fi is the ith class frequency.

Example: the mean time spent by students for leisure activities is calculated as:

CL CB mi f
10 - 14 9.5 - 14.5 12 8 CL f
15 - 19 14.5 - 19.5 17 28 155 - 160 2
20 - 24 19.5 - 24.5 22 27
160 - 165 6
25 - 29 24.5 - 29.5 27 12
165 - 170 18
30 - 34 29.5 - 34.5 32 3
170 - 175 25
35 - 39 34.5 - 39.5 37 1
175 - 180 9
40 - 44 39.5 - 44.5 42 1
180 - 185 4

σ𝟕𝐢=𝟏 𝐦𝐢 𝐟𝐢 𝟏𝟐∗𝟖+𝟏𝟕∗𝟐𝟖+⋯+𝟒𝟐∗𝟏 𝟏𝟔𝟓𝟓 185 - 190 1


ഥ=
𝐗 = = 𝟖𝟎 = 20.7 hours.
σ𝟕𝐢=𝟏 𝐟𝐢 𝟖+𝟐𝟖+⋯+𝟏 Totals
2.1 Measure of Central Tendency
Properties of the Mean
 Uniqueness. For a given set of data there is one and only one mean.

 Simplicity. It is easy to understand and to compute.

 Affected by extreme values. Since all values enter into the computation.

Example: computed the mean for the following datasets.

Dataset1: 115, 110, 119, 117, 121, 126

Dataset2: 75, 75, 80, 80, 280.


The mean for both dataset is 118, but a mean value for the second data is not
representative of the set of data as a whole.

 If every value of the variable is multiplies or divided by some constant, then


the new mean can be obtained by multiplying or dividing the old mean by that
constant.
2.1 Measure of Central Tendency
Geometric Mean
If the observed values are measured as ratios, proportions or percentages &
the series of observations contains one or more unusually large values
then geometric mean gives a better measure of MCT than other means.
It is obtained by

𝐧
𝐗 𝟏 . 𝐗 𝟐 … . 𝐗 𝐧 , 𝐟𝐨𝐫 𝐮𝐧𝐠𝐫𝐨𝐮𝐩𝐞𝐝 𝐝𝐚𝐭𝐚 𝐬𝐞𝐭𝐬
GM =൞ 𝐦 𝐟 𝐟 𝐟
.
𝐗 𝟏𝟏 . 𝐗 𝟐𝟐 … . 𝐗 𝐦𝐦 , 𝐟𝐨𝐫 𝐝𝐚𝐭𝐚 𝐬𝐞𝐭𝐬 𝐨𝐟 𝐗𝐢 𝐡𝐚𝐯𝐢𝐧𝐠 𝐟𝐫𝐞𝐪𝐮𝐞𝐧𝐜𝐢𝐞𝐬 𝐟𝐢

Note: if the dataset is grouped, mi of the class intervals are considered as Xi

Example: calculate the geometric mean for the following births per 1000 individuals
dataset. 7, 8, 3,14, 2, 1, 440, 15, 52, 6, 2, 1, 25, 12, 6, 9, 2, 1, 6, 7, 3, 4, 70, 20,
200, 2, 50, 21,15, 10, 120, 8, 4, 70, 3, 1,103, 20, 90, 1, 237
𝐧 42
GM= 𝐗𝟏. 𝐗𝟐 … . 𝐗𝐧, = 7 ∗ 8 ∗ 3 ∗ ⋯ 1 ∗ 237 = 0.999 = 10
2.1 Measure of Central Tendency
Harmonic Mean (HM)
is a suitable MCT when the data pertains to speed, rates & time.

n n
The HM is calculated as: HM = = or
1 + 𝟏 + …+ 𝟏 σ
𝟏
X𝟏 X𝟐 Xn Xi

n
HM = , where X1, X2, …Xk have the corresponding frequencies f1, f2,…fk
fi
σ
Xi
While if the data is grouped, mi‘s are considered as Xi.

Example: Milk is sold at a place with the rates of 1.8, 2, 2.25 & 2.5 birr per liter in four
different months. Then, find the average price paid per liter.

HM = 4/(1/1.8 + ½ + 1/2.25 + 1/2.5 ) = 4/1.9 = 2.11 birr.


2.1 Measure of Central Tendency
2 Median: is an alternative MCT.
Suppose there are n observations in a sample and they are ordered from smallest to
largest, then the median is defined as:
𝑛 + 1 th
✓ The observations if n is odd
2
𝑛 th 𝑛 th
✓ The average of the and 2 + 1 observations if n is even.
2
Example: (1) Compute the sample median for the ordered birth weight data.
2069, 2581, 2759, 2834, 2838, 2841, 3031, 3101, 3200, 3245, 3248,
3260, 3265, 3314, 3323, 3484, 3541, 3609, 3649, 4146.
(2) Number of births (x103) taken in a certain country: 3,5,7,8,8,9,10,12,35.

(1) Since n=20 is even, the average of 10th & 11th observation is (3245 + 3248)/2 = 3246.5 g

median
(2) Since n is odd, the 5th = ((9+1)/2)th observation, which is equal to 8 is median.
2.1 Measure of Central Tendency
Median:
𝐧
− 𝐂𝐅
෩=𝐋+
For a grouped dataset, median is calculated as: 𝐗 𝟐
∗ 𝐰, where
𝐟
L = LCB of the interval containing the median class,
w = class width,
n = total frequency of the sample,
CF = Cumulative frequency of all interval below L,
f = Frequency of the interval containing the median.
Note: the median class is first class whose cumulative frequency is at least n/2.
Example: find the median for the following grouped data.
CL CB mi f
10 - 14 9.5 - 14.5 12 8 Less than cf=8, 36, 63, 75, 78, 79,80
15 - 19 14.5 - 19.5 17 28 n/2 = 80/2 =40, then the median class is
20 - 24 19.5 - 24.5 22 27 3rd Therefore, L= 19.5, CF = 36, f=27 &
25 - 29 24.5 - 29.5 27 12 w = 5.
30 - 34 29.5 - 34.5 32 3 40 − 36
35 - 39 34.5 - 39.5 37 1 19  5 +  5 = 20.241
27
40 - 44 39.5 - 44.5 42 1
2.1 Measure of Central Tendency
3 The Mode: is the value of the observation that occurs with the greatest frequency.
Example: Find the mode for the following data:
(a) 22, 66, 69, 70, 73

(b) 1.8, 3.0, 3.3, 2.8, 2.9, 3.6, 3.0, 1.9, 3.2, 3.5
(c) 10, 10, 9, 9, 8, 12, 15, 5 .
Solution Thus, distributions with one mode are called unimodal,
(a) No mode those with two modes are called bimodal, & those with
(b) 3.0 more than two modes are called multimodal.
(c) 9 and 10

∆𝟏
The mode for grouped data is calculated as: ෡=L+
𝐗 * w, where
∆𝟐 + ∆𝟏

L = Lower class boundary of the modal class; where modal class is a class that has max f.
w = the class width;
∆1 = 𝑓 − f1 & ∆2 = 𝑓 − f3 ;
f =frequency of the modal class, & f1 = frequency of the class immediately preceding the
modal class, & f3= frequency of the class immediately succeeding the modal class.
2.1 Measure of Central Tendency
The Mode:
Example: Find the mode for the following data:

The modal class is 2rd Therefore,


L= 14.5, f1 = 28, f2=8, f3=27 & w = 5.
28 − 8
14  5 +  5 = 19.26
( 28 − 8) + ( 28 − 27)

Empirical Relationship between Mean, Median and mode


If the distribution is symmetrical, then mean=mode=median
While if the distribution is asymmetrical, then mean & mode usually lie on the 2 ends & hence:
(Mean - Mode) = 3(Mean - Median)
2.1 Measure of Central Tendency
Which MCT is Best? Why?

Is the data Use mode


Yes
categorical?
No

Is total of Use Mean


Yes
interest?
No

Is distribution
Yes Use Median
skewed?
No

Use Mean
2.1 Measure of Central Tendency
4 The Quantiles:
Quantiles are dividing the distribution of ordered values/data into equal-sized parts.
✓ Quartiles: 4 equal parts
✓ Deciles: 10 equal parts
✓ Percentiles: 100 equal parts
Quartiles: Split Ordered Data into 4 Quarters
The first and third quartiles (denoted Q1 and Q3) are defined as follows:
❖ 25% of the data lie below Q1 (and 75% is above Q1),
❖ 25% of the data lie above Q3 (and 75% is below Q3).

25% 25% 25% 25%

( Q1 ) ( Q2 ) ( Q3 )
first quartile 2nd quartile third quartile
Median
2.1 Measure of Central Tendency
The Quantiles:

k(n + 1)𝑡ℎ
formula to get for ungrouped data the Kth quartile is: Qk = , where k=1,2,3.
4

c ( kn − C𝐹)
For a grouped data the kth quartiles can be done: Qk = L + 4
, k =1, 2, 3, where
f

L = lower class boundary of the kth quartile class

n= the total number of observations

CF = the less than cumulative frequency corresponding to the class immediately preceding
the kth quartile class

C= the class width of the quartile class

F = frequency of the kth quartile class

N.B.: Deciles and Percentiles can be computed in the same fashion as quartiles by k=1,2,3, …9
and k=1, 2, 3, …99, respectively.
2.2 Measure of Variation/Dispersion
The location measures may not be adequate enough to describe the distribution of the
data.
The Variation/dispersion of observations around any particular value is another property
which characterizes the data and its distribution.
7 7
For instance,
7 8 3 2
Consider the 7 77
following data
7 77 7 8 13
8 7
6 9

Mean = 7 Mean = 7
Mean = 7

Var=0 Var=0.63 Var=4.04

Thus, MV/MD gives information on the spread or variability of the data values.
2.2 Measure of Variation/Dispersion
◼ There are different types of measures of variability.

a measure of dispersion defined as the difference


1 Range between the maximum and minimum value

average of squared deviations of each values from the


mean & it measures the average square distance of
2 variance each values from the mean.
MV/MD
✓ the square root of the variance & it measures the
distance of each values from the mean.
3 SD ✓ It is most commonly used measure of variation .

compare the variability of two or more datasets that


4 CV the mean is differ in magnitude and/or having
different units of measurement.
2.2 Measure of Variation/Dispersion
◼ There are different types of measures of variability.

◼ 2. Sample Variance
◼ The sample variance, s2, is the arithmetic mean of the squared deviations from
n
the sample mean: (
 ix − x )2

s 2 = i =1
n −1
>
◼ 3. Sample Standard Deviation
◼ The sample standard deviation, s, is the square-root of the variance
n
 (xi − x )
2
◼s has the advantage of being in the
i =1
s=
n −1 same units as the original variable x
2.2 Measure of Variation/Dispersion
◼ Variance and Standard Deviation for Grouped Data
The calculation is the same to the formula of data given in frequency distribution except that Xi is
substitute by mi. that is: n

 f i (mi − x )
2

s= i =1
n −1
4 Coefficient of Variation (CV):
The coefficient of variation (CV) or relative standard deviation (RSD) is the sample standard
deviation expressed as a percentage of the mean, i.e.
s
CV =   100%
x
 The CV is not affected by multiplicative changes in scale

 Consequently, a useful way of comparing the dispersion of variables measured on different


scales

 Some times it is also important to compare the variation of the datasets if they have different
mean in magnitude
2.2 Measure of Variation/Dispersion
4 Skewness & Kurtosis
▪ Skewness is a measure of the asymmetry of
the probability distribution.
▪ Roughly speaking, a distribution has positive
skew (right-skewed), Normal, & negative
skew (left-skewed).
▪ Skewness is computed as (mean–mode)/SD

Kurtosis is the degree of measure of peaked-


ness or flatness of a distribution compared
with the normal distribution.

Kurtosis of a sample data set is calculated


directly from the data by the formula:

= -
2.2 Measure of Variation/Dispersion
Example: consider the following datasets and find
(a) The sample variance & Standard deviation
(b) the coefficient of variation
(c) the Skewness
(d) Compare the variability of the datasets

<C CL f
CL CB Mi f >CF
F
155 -160 2
10 - 14 9.5 - 14.5 12 5 5 80
15 - 19 14.5 - 19.5 17 10 15 75 160 - 165 6

20 - 24 19.5 - 24.5 22 20 35 65 165 - 170 18


25 - 29 24.5 - 29.5 27 22 57 45 170 - 175 25
30 - 34 29.5 - 34.5 32 10 67 23 175 - 180 9
35 - 39 34.5 - 39.5 37 1 68 13 180 - 185 4
40 - 44 39.5 - 44.5 42 2 70 12 185 - 190 1
3 Elementary Probability Theory
3.1 Review of Set Theory

Randomness & uncertainty exist in our daily lives as well as in every discipline, & hence
Probability is a mathematical framework that allows us to describe and analyse random
phenomena.
Random phenomena we mean?
events or experiments whose outcomes we can not predict with certainty.
Thus, Probability uses the language of sets &
a set is a collection of things (elements) which denoted by a capital letters like A, B, C….
For example, to define a set A that consists of the two elements and ♢ is, A = { ,♢}.

Examples: the set of: ▪ Roll a fair die: S={1,2,3,4,5,6}


▪ natural numbers, ℕ ={1, 2, 3, ….} ▪ Toss a fiar coin: S={H,T}
▪ integers, ℤ = {…, -3, -2, -1, 0, 1, 2, 3,…}
▪ closed intervals on the real line, [2, 3]
▪ open intervals on the real line, (-1, 4)

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Countable Sets:
✓ It is a finite set, |A| < ∞ ; or
✓ If A is finite Or the elements of A can be enumerated/listed.
Example:
▪Roll a fair die: S={1,2,3,4,5,6}
▪Toss a fiar coin: S={H,T}
Countable infinite:
A set is countable infinite if it is in one-to-one correspondence with ℕ = {1, 2, 3, ….}.
Example:
▪ natural numbers, ℕ ={1, 2, 3, ….}
▪ integers, ℤ = {…, -3, -2, -1, 0, 1, 2, 3,…}
Uncountable Sets:
a set A is called uncountable if it not countable
Example:
▪closed intervals on the real line, [2, 3]
▪open intervals on the real line, (-1, 4)
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
3.1 Review of Set Theory
Random experiment: A phenomenon whose outcome cannot be predicted with
certainty, such as:
Sample Space: is the set of all possible outcomes.
✓ Roll a die
Examples:
✓ Flip a coin
✓ Roll a die → S= {1, 2, 3, 4, 5, 6}
✓ Flip a coin two times
✓ Flip a coin → S= {H, T}
✓ Select a family having 2 C.
✓ Roll a die three times →S = {HH, HT, TH, TT}
✓ Select a family having 2 C. → S={bb, bg, gb, gg}

Outcome: is the result of a random experiment.


Examples Events: is collection of possible outcomes.
✓ Roll a die → 3 Examples:
✓ Flip a coin → (H) Flip a coin & the # of head is 1→E={HT, TH}
✓ Flip a coin two times → (H, H) Roll a die & consider odd #→ E = {1, 3, 5}
✓ Select a family having 2 C. → (gb) A family having at least 1 boy → E= {bb, bg, gb}
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
3.1 Review of Set Theory
Exercise: Consider an experiment of rolling two dice, then what is the event that
(a) A sum of 7 turns up
(b) A sum of 11 turns up
(c) A sum less than 4 turns up
(d) A sum of 12 turns up

Solution:

Sample space
(a) E1 = {(6,1), (5,2), (4,3), (3,4), (2,5), (1,6)}
(b) E2 = {(6,5), (5,6)}
(c) E3 = {(1,1), (2,1), (1,2)}
(d) E4 = {(6,6)}

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Universal set (S= Ω): the set of all things that we could possibly consider in a given context.
Examples of universal sets
✓ S={1,2,3,4,5,6}
✓ S={H,T}

Null set/ empty set (∅={}): the set with no elements


For any set A, ∅ ⊂ A.
Set A is a subset of set B if every element of A is also an element of B.
We write A⊂B, where "⊂" indicates "subset."
Examples: A= {1,4} and B={1,4,6}, then A⊂B.
If two sets are equal, i.e., A= B if and only if A⊂B & B⊂A.
Example: A={1,2,3} & B={3,2,1} is A=B Sample Space S

Venn Diagrams: Venn diagrams are very useful in


B
visualizing relation between sets. C
A
Belonging, non-belonging, and overlaps between events
and sets can be presented by these diagrams
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
3.1 Review of Set Theory
Union: the union of two sets is a set containing all elements that are in A or B (possibly
both).
The union of sets A and B is defined as:
Or
A= {1,2} & B= {2,3}, then A ∪ B ={1,2}∪{2,3}={1,2,3}
Similarly we can define the union of three or more sets

Intersection: the intersection of two sets A and B, denoted by A∩B, consists of all
elements in both A and B.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Complement: The complement of a set A, denoted by Ac or Ᾱ, is the set of all elements in
the S =Ω that not in A.

Example: If the universal set is given by S = {1, 2, 3, 4, 5, 6}, and A = {1, 2}, B = {2, 4, 5},C =
{1, 5, 6} are three sets, find the following sets:
(a) A ∪ B (b) A ∩ B (c) Ᾱ (d) A ∪ B ∪ C (e) A ∩ B ∩ C (f) B-A

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Mutually exclusive set (disjoint): Two sets A and B are mutually
exclusive (or disjoint) if A∩B=∅.
S
B
A

In general, A1, A2, A3, …, An are ME if Ai ∩ Aj = ∅, i ≠ j.

Partition: A collection of sets A1, A2, A3, …, An is a Partition of S if:


a) They are disjoint .
b) A1 ∪ A2 ∪ A3⋯ ∪ An = S

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Exercise: 1. A city has two TV channels, the NILESAT and the EUTELSAT. The following information
was obtained from a survey of 100 residents of the city. 35 people used NILESAT, 60 used the
EUTELSAT, 20 used both channels. Then use Venn Diagram to show the distribution.
Solution:
N E
15 20 40
25

2. A construction tower crane can operate to height H of 400 ft, a range (radius) R of 60 ft, and an
angle  of ± 90o.
What is the sample space of operation of the crane?
Sketch the sample space and also the following event A: 0<H<80 and 0    30

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.1 Review of Set Theory
Here are some rules that are often useful when working with sets.
De Morgan's law: for a set A1, A2, A3, …, An , we have
The complement of the union is the intersection of the complements, i.e.,
✓ (A1 ∪ A2 ∪ A3⋯ ∪ An)c = Ᾱ1 ∩ Ᾱ2 ∩ Ᾱ3⋯ ∩ Ᾱn
The complement of the intersection is the union of the complements, i.e.,
✓ (A1 ∩ A2 ∩ A3⋯ ∩ An )c = Ᾱ1 ∪ Ᾱ2 ∪ Ᾱ3⋯ ∪ Ᾱn

Distributive law: for any set A, B, and C we have


✓ A ∩ (B∪ C) = (A ∩ B) ∪ (A ∩ C)
✓ A ∪ (B ∩ C) = (A ∪ B) ∩ (A ∪ C)
Examples 1: S= {1, 2, 3, 4, 5, 6}, and A = {1, 2}, B = {2, 4, 6}, C= {1, 5, 6}, then find
✓ A ∪ B, A ∩ B, Ᾱ
✓ Check De Morgan’s law by finding (A ∪ B)c and Ac ∩ Bc.
✓ Check the distributive law by finding A ∩ (B ∪ C) & (A ∩ B) ∪ (A ∩ C)
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
3.2 Counting Techniques
In many cases the number of sample points in a sample space is not very large, &
direct counting of sample points used to obtain probabilities.
However, practically direct counting of sample points become impossibility.
To avoid such difficulties we apply the fundamental principles of counting (counting techniques) like:
✓ Tree Diagram
✓ Multiplication Techniques
✓ Permutation Techniques
✓ Combination Techniques
Tree Diagram: is used to describing the sample points and listing them in a systematic way.

Example: A contractor operates three concrete pumps. A pump is either operational (O) or not
operational (N). Then list all possible sample points states of the pump occurred.

Solution: Because each pump can have two states (either O or N), the sample points of all the
possible states of the pump is listed use tree diagram.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.2 Counting Techniques
Multiplication Techniques:
The tree diagram technique is not working with a large number of trials, where each trial has several
possible outcomes.
For example, if an experiment has n trials and the ith trial has mi possible outcomes (i = 1, 2, 3, ...,n),
then there will be:
▪ m1 branches at the starting point,
▪ m2 branches at the end of each of the m1 branches,
▪ m3 branches at the end of each of m1 × m2 branches, and so on.
Therefore, the total number of branches at the end would be m1 × m2 × m3 ×···× mn, which
represents all the sample points in the sample space S of the experiment.
This rule of describing the total number of sample points is known as the Multiplication Technique.
Example: consider the pump example and determine the possible sample sates of the pump
occurred.
Solution: the sample points of all the possible states of the pump is 2*2*2 = 8 events.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.2 Counting Techniques
Permutation Techniques:
Suppose that we have n distinct objects & wish to arrange r of these objects in a line.
Since there are
• n ways of choosing the 1st object,
• n - 1 ways of choosing the 2nd object, . . . , &
• n - r + 1 ways of choosing the rth object
it follows that the number of different arrangements is given by n(n - 1)(n - 2) . . . (n - r + 1) = nPr
We call nPr the number of permutations of n objects taken r at a time.
𝒏!
In the particular case where r = n, the above equation becomes nPn = = n(n - 1)(n - 2) . . . 1 = n!
𝒏−𝒏 !

𝒏!
Moreover, we can rewrite nPr in terms of factorials as nPr =
𝒏–𝒓 !

Example: We assume that the bridge is supported by 9 cables, & the failure of 3 cables results in the failure of the
bridge, then what is the number of permutations of 3 out of 9 that can result in bridge failure?
9! 9! 9𝑥8𝑥7𝑥6!
Solution: n= 9 & r=3, then the number of permutations is 9P3 = = 6! = = 9x8x7 = 729 .
9 –3 ! 6!

Remark: If a set consists of n objects of which n1 are of one type, n2 are of a second type, . . . , nk are of a kth type.
𝒏!
Then the number of different permutations of the objects is given by: .
𝒏𝟏!𝒙𝒏𝟐!𝒙𝒏𝟑!𝒙…𝒙𝒏𝒌!

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.2 Counting Techniques
Combination Techniques:
In a permutation we are interested in the order of arrangement of the objects.
For example, abc is different permutation from bca.
In many problems, however, we are interested only in selecting or choosing objects without regard to order.
For example, abc & bca are the same and Such selections are called combinations.
Thus the total number of combinations of r objects selected from n (also called the combinations of n things taken
𝒏! P
r at a time) is denoted by nCr = =n r.
𝒓! 𝒏−𝒓 ! 𝒓!

Example: if you consider the bridge example above, how may combinations that the bridge failed in the result of
the 3 cables out of 9?
Exercise:
1. How many horizontal flags can be formed using 3 colors out of 5 when
a) Repetition is allowed? b) Repetition is not allowed?
2. A college team plays 10 football games during a season. In how many ways can it end the season with five wins,
four losses, and one tie?

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.3 Approaches of Defining Probability
Recall that Probability is a way of expressing knowledge or belief that an event will occur/has
occurred.
But it has often been noted that the word probability is used in three different senses in ordinary
language that is:
✓ Frequentist Definition
✓ Classical Definition
✓ Subjective Definition

Frequentist Definition: the probability of an event is simply the “long-run proportion” of times
that the event occurs under many repetitions of the experiment.

Example: when Mendel conducted his famous hybridization experiment with peas, one such
experiment resulted in offspring consisting of 428 peas with green pods and 152 peas with yellow
pods.
Therefore, the probability of yellow pods is 152/580 = 0.262

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.3 Approaches of Defining Probability
Classical Definition: based on the assumption of underlying equally likely events.
Example: the probability that a randomly selected family give birth either a boy or girl.
Subjective Definition: based on a personal degree of belief and experience
Example: the probability that the rain will be raining today (we consider things like whether there are
clouds in the sky & the humidity)
What ever it, all definitions lead to the same set of axioms.
Axiom 1: P(A) ≥ 0 for every event A (probability of an event is non-negative)
Axiom 2: For the sure or certain event S, P(S) = 1
Axiom 3: if A1, A2, A3 …, are disjoint events, then P(A1∪ A2∪ A3 …)=p(A1)+p(A2) + p(A3) +…
Note the following are not axioms, but it can be proved
1: P(∅) =0
2: P(Ac) = 1 – p(A)
3: P(A) ≤ 1
4: Let A, B, C be events (may not ME), then P(A ∪ B) = p(A) + p(B) – p(A ∩ B)
5: P(A ∪ B ∪ C) = p(A)+p(B)+p(C ) – p(A∩ B) – p(A∩C) – p(B∩C) +p(A∩B∩C)

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.4 Conditional probability, Independent & Dependent Events
Conditional Probability: Sometimes the chance of a particular event happens depends
on the outcome of some other event.
Thus, if A and B are two events in a sample space, S, then the conditional probability of A
given B is defined as: p(A/B) = p(A ∩ B)/ p(B), p(B) > 0.
Note that conditional probability itself is a probability measure, so it satisfies probability
axioms.

Example: I roll a fair die. Let A be the event that the outcome is an odd number &
let B be the event that the outcome is less than or equal to 3, then
(a) What is the probability of A?
(b) What is the probability of A given B?

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.4 Conditional probability, Independent & Dependent Events
Independent Events: Two events are independent if the occurrence of one does not
affect the probability of the other occurring.
If A has no effect on B, we said that A,B are independent events. Then,
✓ P(A ∩ B)= P(B)P(A)
✓ P(A/B)=P(A) or P(B/A)=P(B)
Warning! Disjoint (mutually exclusive) ≠ Independent
Because, p( A ∩ B) = ∅ ≠ p(A) p(B)
p(A/B) = 0 ≠ p(A)

Example: Let I pick a random number from {1, 2, 3, ⋯ , 10}, and call it N. Suppose that all
outcomes are equally likely. Let A be the event that N is less than 7, and let B be the
event that N is an even number, then are A and B independent?

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.5 Law of Total Probability Rules and the Baye’s Theorem
Law of Total Probability: If B1, B2, . . . , Bn is a partition of the sample space S, then for any
event A we have

Example: I have three bags that each contain 100 marbles:


Bag 1 has 75 red and 25 blue marbles;
Bag 2 has 60 red and 40 blue marbles;
Bag 3 has 45 red and 55 blue marbles.
I choose one of the bags at random and then pick a marble from the chosen bag, also at random.
What is the probability that the chosen marble is red?

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
3.5 Law of Total probability rules and the Baye’s Theorem
Bayes’ Rule: suppose that we know p(A/B), but we are interested in the probability
p(B/A).
✓ For any two event A and B, where p(A) > 0, we have

Similarly,

✓ If B1, B2, B3,… form a partition of the sample space S, and A is any event with p(A) > 0,
we have

Example: consider the marble example above, and suppose we observe that the chosen
marble is red, then what is the probability that Bag 1 was chosen?
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
Chapter 4: Random Variables & Probability Distribution

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
Random Variable: a random variable is a real-valued variable that gets its value from a
random experiment
OR
a random variable X is a function from the sample space to the real numbers.
Examples:
1. Let I toss a coin three times and suppose we are interested in the number of heads,
then
S={HHH, HHT, HTH, THH, TTH, THT, HTT, TTT}.
X= 0, 1, 2, 3.
2. Recording the lifetime of an electronic device
S = (T: t ∈ 0, ∞), then T : t> 0
Thus, a random variable is a real-valued function that assigns a numerical value to each
possible outcome of the random experiment.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
Discrete Random Variable: a random variable X is called discrete, if X takes on a finite or
countable infinite number of values.
Example: Number of heads in three tosses of a coin
Probability Mass Function (pmf): All the information about a discrete r.v. can be summed
up in a function called the probability function/probability mass function (pmf).
Thus, if X is a discrete r.v. then its pmf is defined as p(X) = p(X = x).
Properties of PMF
✓ 0 ≤ p(x) ≤ 1 for all x
✓ σ𝑥 𝑝 𝑥 = 1 𝑓𝑜𝑟 𝑎𝑙𝑙 𝑥

Cumulative Distribution Function (CDF): the cumulative distribution function (CDF) of a


random variable is another method to describe the distribution of random variables.
Thus, the CDF of random variable X is defined as F(x) = P(X ≤ x), for all x.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
Properties of distribution functions, F(x)
The function F(x) is a CDF if and only if the following conditions hold.
✓ 0 ≤ F(x) ≤ 1 for all x
✓ F(x) is non-decreasing [i.e., F(a) ≤ F(b) if a ≤ b].

✓ F(x) is continuous to the right (i.e.,lim 𝐹 𝑥 + ℎ = 𝐹 𝑥 𝑓𝑜𝑟 𝑎𝑙𝑙 𝑥)


ℎ→0

✓ lim 𝐹 𝑥 = 0 and lim 𝐹 𝑥 = 1


𝑥→−∞ 𝑥→∞

Expected Value, Variance & SD: Let X be a discrete random variable with PMF of p(x)
then,
✓ the expected value of X, denoted by E(X) = µx, is defined as:

෍ 𝑥𝑝 𝑋 = 𝑥 = ෍ 𝑥𝑝 𝑥
𝑥 𝑥

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
✓ the variance of X , Var(X) = σ2x , with µx is defined as:

෍ 𝑥 − µx 2𝑝 𝑋 = 𝑥 = ෍ 𝑥 − µx 2𝑝 𝑥 = 𝑬 𝑿𝟐 − µ2x
𝑥 𝑥

✓ the standard deviation of X is defined as:

SD(X) = σ2 = 𝑉𝑎𝑟 𝑋

Example 3: consider the experiment of tossing a fair coin three times, and let X be
defined as the number of heads I observed, then construct its pmf & CDF; check the
properties of pmf & CDF; and find the expected value, variance and SD of X.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
Continuous Random Variable: A random variable X is continuous r.v. if X takes all values
in an interval on the real line.

The theory of continuous random variables is completely analogous to the theory of


discrete random variables. i.e.,
take any formula about discrete random variables, & then replace
✓ sums with integrals, &
✓ replace PMFs with probability density functions (PDFs).
Probability density function (pdf): a function with values f(x), defined over the set of all
real numbers, is pdf of the continuous random variable X.
Properties of pdf, f(x)
✓ f(x) ≥ 0, −∞ < x < ∞

✓ f(x) = ‫׬‬−∞ 𝑓 𝑥 𝑑𝑥 = 1
𝑏
✓ P(a ≤ x ≤ b) = ‫ 𝑥𝑑 𝑥 𝑓 𝑎׬‬, −∞ < a < x < b < ∞
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
4.2 Introduction to random Variables
Cumulative distribution function (CDF): If X is a continuous random variable, then the
cumulative distribution of X is given by:
𝑥
F(x) = p(X≤ x)= ‫׬‬−∞ 𝑓 𝑡 𝑑𝑡.
Note that: If f (x) and F(x) are the values of the probability density and the distribution
function of X at x, then
𝑏
✓ P (a ≤ x ≤ b) = F(b) - F(a) = ‫𝑥𝑑 𝑥 𝑓 𝑎׬‬, 𝑎 ≤ 𝑏 &
d(F(x)
✓ f x = dx
if F(x) is differentiable at x.

Expected Value, Variance & SD: Let X be a continuous r.v. with pdf of f(x) then,
✓ the expected value of X, denoted by E(X) = µx, is defined as:

න 𝑥𝑓 𝑥 𝑑𝑥
−∞

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.2 Introduction to random Variables
✓ the variance of X, Var(X) = σ2x , with µx is defined as:
∞ ∞
න 𝑥 − µx 2𝑓 𝑥 𝑑𝑥 = න 𝑥2𝑓 𝑥 𝑑𝑥 − µx2 = 𝑬 𝑿𝟐 − µ2x
−∞ −∞

✓ the standard deviation of X is defined as:

SD(X) = σ2 = 𝑉𝑎𝑟 𝑋

Example 5: Let X is a r.v. with f(x) = 1, 0 < x <1, then


a) Check properties of pdf
b) Construct the CDF of X & check its properties
c) Find E(X), Var(X) and SD(X)

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
There are a number of specific distributions that are used over and over in practice that possess the properties of a
probability :
✓ mass function for discrete random variables
✓ density function for continuous random variables.
There is a random experiment behind each of these distributions.
Since these random experiments model a lot of real life phenomenon, these ‘special’ distributions are used
frequently in different applications.
Recall that we have two types of distribution
Discrete Distribution Continuous Distribution
✓ Uniform Distribution
✓ Discrete Uniform distribution ✓ Normal (Gaussian) Distribution
✓ Bernoulli Distribution ✓ Exponential Distribution
✓ Gamma Distribution
✓ Binomial distribution ✓ Beta Distribution
✓ Hypergeometric distribution ✓ Chi-squared Distribution
✓ Cauchy Distribution
✓ Poisson distribution ✓ Weibull Distribution
✓ Negative Binomial (Pascal) distribution ✓ The t-Distribution
✓ F-distibution
✓ Geometric distribution
But, in this section we will discuss the nature and applications of Bernoulli, Binomial, and Normal distribution.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
1. Bernoulli Distribution: a Bernoulli random variable is a random variable that can only take two possible values,
usually 0= failure & 1 = success.
A random variable X is said to be a Bernoulli random variable with the probability of success p is defined as:
𝑝 𝑋 = 𝑝 𝑥 (1 − 𝑝)1−𝑥 , 𝑓𝑜𝑟 𝑥 = 0 , 1
Properties of Bernoulli Random Variable
➢The trail results in one of S or F outcomes,
➢No other outcomes are possible,
➢The probability of S and F is denoted by p & q =1 – p.

If X is a Bernoulli distribution, then


✓ The expected value (mean) of X is P
✓ The variance of X is p(1-p) = pq
Example:
➢ Tossing a coin once and considering heads as success and tails as failure.
➢ Checking items from a production line: success = not defective, failure = defective.
➢ Phoning a call center: success = operator free; failure = no operator free.
➢ Student passes exam: success=passing the exam; failure= fail the exam

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
2. Binomial Distribution: A binomial experiment consists of n identical and independent sequencies of Bernoulli
trials and each trials can result in one of the two possible outcomes (success or failure).
The probability of success, p is constant from trial to trial.
Therefore, the probability of a r.v. X having x success out of the n trials is given by:
𝑛 𝑥
𝑝 𝑋 = 𝑝 (1 − 𝑝)𝑛−𝑥 , 𝑓𝑜𝑟 𝑥 = 0, 1, 2, 3, … where n > 0 and 0 ≤ p ≤ 1
𝑥
If X is a binomial distribution, then
A company owns 400 laptops. Each laptop has an 8%
✓ The expected value (mean) of X is np
✓ The variance of X is np(1-p) = npq
probability of not working. You randomly select 20

Example: laptops for your salespeople. (a) construct the

Tossing a coin 20 times to see how many tails occur. probability distribution (b) What is the likelihood that 5

Asking 200 people whether they watch ETV news. will be broken? (c) What is the likelihood that they will

Rolling a die 10 times to see if a 5 appears. all work?

Example: Suppose it is known that in a certain population 10 percent of the population is color blind. If a random
sample of 25 people is drawn from this population, then construct the probability distribution and find
a) The mean and SD.
b) The probability that two or fewer will be color blind.
c) The probability that between six and eight inclusive will be color blind.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
3. Poisson Distribution: A random variable X is defined to have a Poisson distribution if the pmf of X is given by:
𝒆−𝝀 𝝀𝒙
𝒑 𝒙 = , for x =0, 1, 2, …
𝒙!

Here x is the number of times an event occurs in an interval independently and the average rate at which events
occur is constant called 𝜆 .
If X is a Poisson distribution, then
✓ The expected value (mean) of X is 𝜆
✓ The variance of X is 𝜆
Note that
✓ 𝜆 is the average number of occurrences of the random event
✓ An interesting feature of Poisson distribution is the fact that the mean = variance and it is the only
distribution.
✓ The occurrence of events are independent.
✓ Theoretically, an infinite number of occurrences of the event must be possible in the interval.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
Example: Suppose that customers enter a waiting line at random at a rate of 4 per minute. Assuming
that the number entering the line during a given time interval has a Poisson distribution, find the
probability that:
(a) one customer enters during a given one-minute interval of time;
(b) at least one customer enters during a given half-minute time interval.

2. Suppose you own a coffee shop, and based on your historical data, you know that the average
number of customers arriving per hour is 15. then
a) Construct the pmf of X
b) Find the probability that 10 customers arrive in an hour
c) Find the probability that 10 to 12 customers arrive in an hour.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
4. Normal Distribution: The normal (Gaussian) density function was proposed by C.F. Gauss (1777-1855)

as a model for the relative frequency distribution of measurement.

This bell-shaped curve provides an adequate model for the relative frequency distributions of data
collected from many different scientific areas.

2
1  x− 
A continuous r.v. X is said to have a normal distribution, if its pdf is given by: f ( x) = 1 e − 2   

,− x
2 
If X is a normal distribution, then

✓ the expected value (mean) is µ


✓ the variance is σ2.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
4. Normal Distribution:
Characteristics of the Normal Distribution
➢The curve is bell-shaped.
➢The mean, median and mode are equal and located at the center of the distribution.
➢It is uni-modal
➢The curve is symmetrical about the mean.
➢The curve is continuous, i.e., for each X, there is a corresponding Y value.
➢It never touches the X axis.
➢The total area under the curve is 1 and half of it is 0.5000
➢ The areas under the curve that lie within 1, 2 and 3 standard deviations of the mean are approximately 0.68 (68%),
0.95 (95%) and 0.997 (99.7%) respectively. Graphically, it can be shown as:

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
4.3 Probability Distributions
5. Standard Normal Distribution

To compute the probabilities associated with the normal distribution, we use the standard normal table.

If X is a normal random variable with the mean μ and variance σ then the variable Z = (X - µ)/σ is the standardized
normal random variable. In particular, if μ = 0 and σ = 1, then the density function is called the standardized normal
density .
The equation for the standard normal distribution is written as:
z2
1 −
f ( x) = e 2
,− z
2
Characteristics of the Standard Normal Distribution
➢The highest point occurs at μ=0.
➢It is a bell-shaped curve that is symmetric about the mean, μ=0.
➢The total area under the curve equals one.
➢Empirical Rule:
✓ Approximately 68% of the area under the curve is between -1 and +1.
0
✓ Approximately 95% of the area under the curve is between -2 and +2.
✓ Approximately 99.7% of the area under the curve is between -3 and +3.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
z 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09
0.0 0.0000 0.0040 0.0080 0.0120 0.0160 0.0199 0.0239 0.0279 0.0319 0.0359
Standard Normal 0.1 0.0398 0.0438 0.0478 0.0517 0.0557 0.0596 0.0636 0.0675 0.0714 0.0753
Table (Z-Table) 0.2 0.0793 0.0832 0.0871 0.0910 0.0948 0.0987 0.1026 0.1064 0.1103 0.1141
0.3 0.1179 0.1217 0.1255 0.1293 0.1331 0.1368 0.1406 0.1443 0.1480 0.1517
0.4 0.1554 0.1591 0.1628 0.1664 0.1700 0.1736 0.1772 0.1808 0.1844 0.1879
0.5 0.1915 0.1950 0.1985 0.2019 0.2054 0.2088 0.2123 0.2157 0.2190 0.2224
0.6 0.2257 0.2291 0.2324 0.2357 0.2389 0.2422 0.2454 0.2486 0.2517 0.2549
0.7 0.2580 0.2611 0.2642 0.2673 0.2704 0.2734 0.2764 0.2794 0.2823 0.2852
0.8 0.2881 0.2910 0.2939 0.2967 0.2995 0.3023 0.3051 0.3078 0.3106 0.3133
0.9 0.3159 0.3186 0.3212 0.3238 0.3264 0.3289 0.3315 0.3304 0.3365 0.3389
1.0 0.3413 0.3438 0.3461 0.3485 0.3508 0.3531 0.3554 0.3577 0.3599 0.3621
1.1 0.3643 0.3665 0.3686 0.3708 0.3729 0.3749 0.3770 0.3790 0.3810 0.3830
1.2 0.3849 0.3869 0.3888 0.3907 0.3925 0.3944 0.3962 0.3980 0.3997 0.4015
1.3 0.4032 0.4049 0.4066 0.4082 0.4099 0.4115 0.4131 0.4147 0.4162 0.4177
1.4 0.4192 0.4207 0.4222 0.4236 0.4251 0.4265 0.4279 0.4292 0.4306 0.4319
1.5 0.4332 0.4345 0.4357 0.4370 0.4382 0.4394 0.4406 0.4418 0.4429 0.4441
1.6 0.4452 0.4463 0.4474 0.4484 0.4495 0.4505 0.4515 0.4525 0.4535 0.4545
1.7 0.4554 0.4564 0.4573 0.4582 0.4591 0.4599 0.4608 0.4616 0.4625 0.4633
1.8 0.4641 0.4649 0.4656 0.4664 0.4671 0.4678 0.4686 0.4693 0.4699 0.4706
1.9 0.4713 0.4719 0.4726 0.4732 0.4738 0.4744 0.4750 0.4756 0.4761 0.4767
2.0 0.4772 0.4778 0.4783 0.4788 0.4793 0.4798 0.4803 0.4808 0.4812 0.4817
2.1 0.4821 0.4826 0.4830 0.4834 0.4838 0.4842 0.4846 0.4850 0.4854 0.4857
2.2 0.4861 0.4864 0.4868 0.4871 0.4875 0.4878 0.4881 0.4884 0.4887 0.4890
2.3 0.4893Probability
0.4896& Statistics
0.4898 for
0.4901 0.4904 L. 0.4906
CS by Demeke 0.4909 0.4911 0.4913 0.4916
(wadela1606@[Link])
4.3 Probability Distributions
To finding Area under the Standard Normal Curve
I. Draw the picture,
II. Shade the desired area /region,
III. Use standard table to obtain it
Example: find the area under the curve:
✓ P( Z < 2) 0 2
✓ P(-1.74 < Z < 1.53)
✓ P(Z > 1.71)

Answers
p(Z <2) =0.5 + P( 0<Z < 2)= 0.5 + 0.4772 = 0.9772
-1.74 1.53
P(-1.74 < Z < 1.53) = 0.4591 + 0.4370 = 0.8961.
P(Z > 1.71) =0.5 – 0.4564 = 0.0436.

Example1. Suppose that X N (165, 9), where X = the breaking strength of cotton
fabric. A sample is defective if X<162. Find the probability that a randomly
chosen fabric will be defective.
2. Let the height of BDU student is normally distributed with mean of 170 cm and
the standard deviation is 5 cm, then find the probability that a randomly
1.71
chosen student will between 165 and 160 cm.
Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
4.3 Probability Distributions
6. Exponential Distribution: The exponential distribution is one of the widely used continuous distributions.
The exponential distribution is one of the widely used continuous distributions and it is often used to model the
time elapsed between events.

A continuous random variable X is said to have an exponential distribution with parameter λ > 0, if its pdf is given by:

f ( x ) =  e − x , x  0 E( X ) =
1
,V ( X ) =
1
 2
If X is a exponential distribution, then

Example 1: Let X ∼ Exponential(2), then construct the pdf of X, find the mean and variance of X.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
Chapter 5. Sampling Theory
5.1 Introduction Sampling Theory
The population is too large for us to consider collecting information from all its members.
Usually, a representative subgroup of the population (sample) is included in the investigation.
A representative sample has all the important characteristics of the population from which it is drawn.
Sampling involves the selection of a number of study units from a defined population.

Why sampling?
Get information about large populations
 Less costs
 Less field time
 More accuracy i.e. Can Do A Better Job of Data Collection
 When it is impossible to study the whole population (In some cases, it might not be possible to check
100%, like blood test, tasting food)

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
5.1 Introduction Sampling Theory

Common terms used in sampling


the population to be studied
Target Population

are non-overlapping collections of elements from the


Sampling Unit population that cover the entire population.

Common Study unit/element an object on which a measurement is taken.


Terms

Sampling frame a list of sampling units.

the ratio of the number of units in the sample to the


Sampling fraction
number of units in the target population (n/N).
5.2 Types of sampling
Probability sampling
all individuals in the population have an equal likelihood of being chosen.
Non-probability sampling
members are selected from the population in some nonrandom manner.
Probability sampling
Simple random sampling
Systematic sampling Non-Probabilisty sampling

Cluster sampling Convenience sampling


Stratified sampling Judgmental sampling
Multistage sampling Quota sampling
Snowball sampling

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
5.2 Types of sampling
A. Simple Random Sampling
To select a simple random sample you need to:
➢Make a numbered list of all the units in the population from which you want to draw a sample.
➢Each unit on the list should be numbered in sequence from 1 to N (where N is the size of the
population).
➢Decide on the size of the sample

Select the required number of study units, using a “lottery” method or a “table of random numbers”.

"Lottery” method
it may be possible to use the “lottery” method for a small population.
Each unit in the population is represented by a slip of paper, these are put in a box and mixed, and a
sample of the required size is drawn from the box.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
5.2 Types of sampling

Simple random sampling


✓ Lottery Method

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.2 Types of sampling
Simple random sampling
✓ Table of Random Numbers

Example: Suppose that you have a population of 100 patients. From this, you want to draw a
sample of 5 patients.

0 8 4 2 5 7 9 5 4 1 2 5 6 3 2 1 4 0
5 8 2 0 3 2 0 5 4 7 8 5 9 6 2 0 2 4
3 6 2 3 3 3 2 5 4 7 8 9 1 2 0 3 2 5
9 8 5 2 6 3 0 1 7 4 2 4 5 0 3 6 8 6
2 9 0 1 1 2 3 4 5 6 8 7 9 4 3 2 3 4

Samples: 84, 32, 54, 24, 17


Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
5.2 Types of sampling

B. Systematic Sampling
A list of N elements in the population is compiled, ordered according to
a specified variable,
• A sampling size n is chosen,
• A systematic step of k=N/n is set,
• A random number i between 1 and k is taken randomly and
represents the first element to be included,
• Then the other elements selected are i+k, i+2k, i+3k…till n (sample
size). –Less representative (biased) if the
–Cheaper and easier than SRS
order is cyclical
–More representative if order is related
to the interest variable (monotone)
–Sampling frame not always necessary

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.2 Types of sampling

Systematic sampling

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
5.2 Types of sampling
C. Stratified sampling
• Population is partitioned in strata through control variables
(stratification variables), closely related with the target variable, so
that there is homogeneity within each stratum and heterogeneity
between strata.
• A simple random sampling method is applied in each strata of the
population.
– Proportionate sampling: size of the sample from each stratum is proportional
to the relative size of the stratum in the total population.
– Disproportionate sampling: size is also proportional to the standard deviation
of the target variable in each stratum.
–Gains in precision –Stratification variables may not be
–Include all relevant sub-population easily identifiable
even if small –Stratification can be expensive
–reduces sampling error
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
5.2 Types of sampling

D. Cluster sampling

The population is partitioned into clusters (Geographically).

Elements within the cluster should be as heterogeneous as possible


with respect to the variable of interests.

A random sample of clusters is extracted through SRS (with


probability proportional to the cluster size)
✓ All the elements of the cluster are selected (one- stage)

✓ A probabilistic sample is extracted from the cluster (two-stage cluster


sampling)
–Reduced costs –Less precision
–Higher feasibility –Inference can be difficult
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
Introduction Sampling Theory

Section 1 Cluster sampling Section 2

Section 3

Section 5

Section 4
Probability & Statistics for CS by
Demeke L.
(wadela1606@[Link])
5.2 Types of sampling

E. Convenience sampling
• is used in exploratory research where the researcher is
interested in getting an inexpensive approximation.

• The sample is selected because they are convenient.

–Cheapest method –Selection bias


–Quickest method –Non representativeness
–Inference is not possible

 Often used during preliminary research efforts to get an estimate


without incurring the cost or time required to select a random sample.

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.2 Types of sampling
F. Judgmental sampling
Selection based on the judgment of the researcher and is a common non-probability
method.

–Low cost –Non representativeness


–Quick –Inference is not possible
–Subjective

• When using this method, the researcher must be confident


that the chosen sample is truly representative of the entire
population.

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.2 Types of sampling

G. Quota sampling
First identify the stratums (such as sex, age…) and their
proportions as they are represented in the population.

– Then convenience or judgment sampling is used to select


the required number of sample from each stratum.

–There is no guarantee that the


sample is representative (relevance
–Cheapest method of control characteristic chosen)
–Quickest method –Many sources of selection bias
–No assessment of sampling error

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.2 Types of sampling

H. Snowball Sampling
A first small sample is selected randomly
• Respondents are asked to identify others who belong to
the population of interests
• The referrals will have demographic and psychographic
characteristics similar to the referrers
–Lower costs –Inference is not possible
–Low variability
–Useful for “rare” populations

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.3.1 Introduction to Mean of the Sample Mean
When conducting a statistical analysis, researchers often collect data from a sample and use it to make inferences
about the population from which the sample was drawn.

Dealing with such situations is the subject of the field of statistical inference.

Thus, Statistical inference is a collection of methods that deal with drawing conclusions from data that are prone
to random variation.

Random Sampling is a key technique for obtaining a representative sample and making valid inferences about the
population.
The main idea behind random sampling is to ensure that every member of the population has an equal probability
of being included in the sample.
This helps to minimize bias and increase the likelihood that the sample is a good representation of the population.

Sampling distribution provides information about the distribution of sample statistic (e.g. sample Mean, sample
Variance, sample proportion), and
It plays a crucial role in hypothesis testing, confidence intervals, and other inferential statistical techniques.
They provide a framework for making statistical inferences about population parameters based on sample data.
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
5.3.1 Introduction to Mean of the Sample Mean
Statistic: is a function of ‘observable’ random variables, which does not contain any unknown
parameters.
1
Example: If X1,…,Xn is a random sample (provided X1, …, Xn are observable), then 𝑋ത𝑛 = σ 𝑋𝑖 is
𝑛

statistic.
Thus, the sampling distribution of the statistic is the tool that tells us how close is the statistic to the
parameter.
As we begin to use sample data to draw conclusions about a wider population, we must be clear
about whether a number describes a sample or a population.

Summary Statistics Statistic Used to Parameter


Mean 𝑋ത estimates µ
Standard Deviation S estimates σ
Proportion 𝑝Ƹ estimates π
Generally ෡
Θ estimates 𝜃

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.3.2 Sampling Distribution of sample Mean
A Sampling distribution is a probability distribution of a statistic that comes from choosing random
samples of a given population.
It depends on multiple factors like the statistic, sample size, sampling process, and the overall
population.

How Does it Work?


✓ Select a random sample of a specific size from a given population.
✓ Calculate a statistic for the sample, such as the mean, standard deviation ….
✓ Develop a frequency distribution of each sample statistic that you calculated.
✓ Plot/tabulate the frequency distribution of each sample statistic that you developed.
Then the resulting graph/table will be the sampling distribution.

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.3.2 Sampling Distribution Sample Mean
Sampling Distribution of Sample Mean
We take many random samples of a given size n from a population with mean μ and standard
deviation σ.
Then, the sampling distribution of the sample mean is the probability distribution of all possible
values of the random variable 𝑋ത computed from a sample of size n.
Example: Suppose we want to determine the average weight of students in a certain primary school.
The weight of student follows a normal distribution with a mean weight of 25 kg and a standard
deviation of 5 kg.
Sampling
distribution of 𝑋ത

Histogram of 𝑋ത

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
5.3.2 Sampling Distribution Sample Mean
For any population with mean μ and standard deviation σ:
The mean, or center of the sampling distribution of 𝑋ത is equal to the population mean (𝜇𝑋ത = 𝜇 ).
𝜎
The standard deviation of the sample mean is 𝜎𝑋ത = (often called standard error of the
𝑛

mean).
This is because:

=𝜇
Notice that 𝜎𝑋ത is smaller than σx.
The larger the sample size the smaller 𝜎𝑋ത . Therefore, 𝑋ത tends to fall closer to μ, as the sample size
increases.
Thus, if a population is normal with mean μ and standard deviation σ,
the sampling distribution of 𝑋ത is also normally distributed with 𝜇𝑋ത = 𝜇 &
𝜎
standard deviation (standard error of the mean) 𝜎𝑋ത = , which measures how the sample statistic
𝑛

varies from sample to sample.


Probability & Statistics for CS by Demeke
This is because, different random samples would produce different values of 𝑋ത called sampling variability
L. (wadela1606@[Link])
5.4 Central limit theorem
The central limit theorem (CLT) - states that, under some conditions, the random variables of the sample statistic
from a population has an approximately normal distribution.

𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐−𝐸(𝑋)
That is, the random variable 𝑍 = converges in distribution to the standard normal.
𝑣(𝑋)

𝑋ത −𝜇
Examples: Let's assume that X 's are distributed with mean 0.5 and variance 1/12, then 𝑍 = 𝜎 gets closer to the
ൗ 𝑛

normal PDF as n increases.

An interesting thing about the CLT is that it does not matter what the distribution of the Xi 's is, e.i., Xi 's can be
Probability & Statistics for CS by Demeke
discrete, continuous. L. (wadela1606@[Link])
5.4 Central limit theorem
Law of large numbers (LLN)
The law of large numbers (LLN) basically states that the average of a large number of i.i.d. random
variables converges to the expected value (which did not needed the normality assumption).

The weak law of large numbers (WLLN) is the main versions of the law of large numbers.

Let X1, X2 , ... , Xn be i.i.d. random variables with a finite expected value E(Xi) = μ < ∞.

Then, for any ∈ > 0, lim 𝑃 |𝑋ത − 𝜇 ≥∈ = 0.


𝑛→∞

𝑉 𝑋ത
That is, lim 𝑃 |𝑋ത − 𝜇 ≥∈ ≤ , by Chebyshev's Inequality
𝑛→∞ ∈2

𝜎2
= = 0, which goes to zero as 𝑛 → ∞
𝑛∈2

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
Chapter 6. Statistical Inference
6.1 Introduction Statistical Inference
Statistical inference deals with the collected data so as to form conclusions about the
population. It can be Estimation or Hypothesis testing

Estimation

Point estimate Interval estimate


e.g. e.g.
❑ sample mean ❑ confidence interval for mean
❑ sample proportion ❑ confidence interval for proportion

• A point estimate is a single value • a confidence interval provides additional


• Point estimate is always within the IE information about variability

LCL UCL
PE - (RF)*(SE) PE + (RF)*(SE)
PE

Width of confidence interval


Probability & Statistics for CS by Demeke L.
(wadela1606@[Link])
6.1 Introduction Statistical Inference

Hypothesis Testing is a tentative assumption regarding the value of population parameter


whose validity is checked by statistical methods.
Every hypothesis test has the following steps.
Steps Activities

is a claim (assumption) about population parameters


Hypothesis & has null (H0) & alternative (H1) hypothesis

specify the level of significance () & the rejection


Significance region
Hypothesis
Testing Statistic compute appropriate statistic from the sample data

The decision rule may perform by:


✓test statistic- reject H0 if |calc value| is > tab value
Decision
✓p-Value- reject H0 if p-value <  , and/or
✓Confidence interval- reject H0 if zero is in the CI
Conclusion
Make conclusion about the population parameter
based on the decision you made.

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
6.1 Introduction Statistical Inference
Among the statistical inferences, the T-TEST procedure is the most popular that used to
compare mean value.

The t-test procedures can:

can be used to compare a sample mean to a given


One-sample T test
value (Assumed mean).

Indp’t-sample T test is used to comparing two independent groups.

Paried-sample T test is testing the significance of difference in means for


paired samples

Probability & Statistics for CS by Demeke L.


(wadela1606@[Link])
6.2 Point & Interval Estimation for Mean

Recall that the general formula for all confidence intervals is: PE± RF *SE
The value of the reliability factor depends on the desired level of confidence
The confidence interval for µ is constructed based on σ is known and unknown, while for π
assume n is large.
Intervals Estimator

Population Proportion
Population Mean (assume n is large & estimate π if it
is unknown)

σ2 Known σ2 Unknown
(n is large/small) (We use Z-distn),if n is large
(We use Z-distn) other wise t-distn)

Probability & Statistics for CS by Demeke


L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean

1. Confidence Interval for μ (σ2 Known)


The (1 -α )X100% CI for µ is given by:

Usually represented with a


lower confidence limit upper confidence limit
“plus/minus” ( ± ) sign (LCL) (UCL)

The width of the confidence interval estimate is a function of the confidence level, the
population standard deviation, and the sample size.
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean

Finding the Reliability Factor, z/2 from the above formula

• Consider a 95% confidence interval:

1 −  = .95

α α
= .025 = .025
2 2

Z units: z = -1.96 0 z = 1.96


X units: LCL Point Estimate UCL

▪ Find z.025 = 1.96 from the standard normal distribution table


Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean

• Commonly used confidence levels are 90%, 95%, and 99%

Confidence
Confidence
Coefficient, Z/2 value
Level
1− 
80% .80 1.28
90% .90 1.645
95% .95 1.96
98% .98 2.33
99% .99 2.58
99.8% .998 3.08
99.9% .999 3.27
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
Example: The population consists of survival times of cancer patients who have
been treated with a new drug has SD of 43.3 months.
If a random sample of 100 drug-treated patients has a mean survival time of 46.9
months, then
a) What is the point estimate of the population mean?
b) Find a 95% confidence interval for the population mean.
Solution:
a) The point estimates of the population mean is 46.9 months.
b) Sigma is known and n is large, the 95% CI for µ is:
𝜎
𝑋ത ± 𝑍𝛼Τ 2
= 46.9 ± 1.96*43.3/10 = 46.9 ±8.5
𝑛
Thus, the 95% CI for µ is 38.4 < µ < 55.4

Therefore, with 95% confidence the mean survival times of patients in the
population is between 38.4 Probability
months& and 55.4
Statistics months.
for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean

2. Confidence Interval for μ(σ2 Unknown)


• If the population standard deviation σ is unknown, we can substitute it by
the sample standard deviation, s
• So we use the t -distribution instead of the normal distribution
• Assumptions
o Population standard deviation is unknown
o Population is normally distributed
o If population is not normal, use large sample
• Thus, the confidence Interval for µ is given by:
S S
x − t n-1,α/2  μ  x + t n-1,α/2
n n

where tn-1,α/2 is the critical value of the t distribution with n-1 df and an
area of α/2 in each tail are given
Probability below.
& Statistics for CS by Demeke
L. (wadela1606@[Link])
6.3 Confidence Intervals for µ

t distribution values with comparison to the Z value

Confidence t t t Z
Level (10 d.f.) (20 d.f.) (30 d.f.) ____

.80 1.372 1.325 1.310 1.282


.90 1.812 1.725 1.697 1.645
.95 2.228 2.086 2.042 1.960
.99 3.169 2.845 2.750 2.576

Note: t Z as n increases
Probability & Statistics for CS by Demeke
L. (wadela1606@[Link])
6.2 Point & Interval Estimation for Mean
Example: A medical researcher takes the blood pressure from a random sample of 25 (50-
year-old) women and the mean blood pressure of these sampled women is 140 mm Hg
with a standard deviation of 10 mmHg. Then
a) What is the point estimate of the mean blood pressure of all 50-year-old women?
b) Construct the 95% CI for the mean blood pressure of all 50-year-old women.
Solution:
a) The point estimate of the mean blood pressure of all 50-year-old women is 140 mmHg.
b) d.f. = n – 1 = 24, so t n-1, α/2 = t 24, 0.025 = 2.0639

Thus, the 95% CI for µ


S S
= xlj − t n−1,α/2 < μ < xlj + t n−1,α/2
n n
10 10
140 − (2.0639) < μ < 140 + (2.0639)
25 25
135.87 < μ < 144.13

Therefore, with 95% confidence the mean blood pressure of all 50-year-old women is between
Probability & Statistics for CS by Demeke
135.87mmHg and 144.13mmHg. L. (wadela1606@[Link])

You might also like