Engineering Statistics
Engineering Statistics
Engineering Statistics
By:
Statistics: - is the area of science that deals with collection, organization, analysis, and interpretation
of data.
A variable: - is a characteristic, or attribute under study that can assume different values like, length,
weight.
Data: - are the values (measurement or observations) that the variables can assume.
Depending on how data are used, the body of knowledge called statistics is sometimes divided into
two main areas:
1-Descriptive Statistics.
2-Inferential Statistics.
Inferential Statistics: deals with making decision, inferences, predictions, and forecasts about
populations based on results obtained from samples. It uses probability (the chance of on event
occurring) to achieve inferences.
Population: consists of all elements (human or otherwise) that are being studied.
1
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Measurement Scales: how variables are categorized, counted, or measured. There are four basic
levels: nominal, ordinal, interval, and ratio.
2- Ordinal level of measurement: classifies data into categories that can be ranked, however, precise
differences between the ranks do not exist. For example:
- Poor, Good, Excellent
- Small, Medium, Large
- 1st, 2nd, 3rd
3- Interval level of measurement: ranks data and precise differences between units of measure do
exist however, there is no meaningful zero. For example:
- Temperature (No zero because temperature exists even at 0° )
- IQ Scores (No zero because it doesn’t measure people without intelligence)
4- Ratio level of measurement: possess all the characteristics of interval measurement, and there
exists a true zero. In addition, true ratio exists when the same variable is measured on two different
members of the population. For example:
Height, weight, Time, Salary, Age
Data Collection: a variety of ways. One of the most common methods is through the use of surveys.
1- Telephone surveys
- less costly.
-people may be more candid in their opinions.
-not all people have a chance of being surveyed.
2
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Methods of Sampling:
1-Random Sampling:
chance method or random numbers. Generating random numbers with a computer or calculator.
2-Systematic Sampling: each element has an equal probability of selection, but combinations of
elements have different probabilities.
Population size N, sample size= n, sampling interval k=N/n.
Randomly select a number j between 1 and k, sample element j and then every kth element thereafter,
j+k, j+2k, etc.
K =5
3- Stratified Sampling: divided population into groups (strata) according to some specified
characteristics such as age, grade level or income. The subsamples are randomly selected from each
strata.
3
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4- Cluster Sampling:
The population is divided into groups or clusters of elements usually geographic or organizational,
some of the groups are randomly chosen This method is useful when it is difficult or costly to develop
a complete list of the population members or when the population elements are widely dispersed
geographically
Statistical Studies:-
1-Observational Study: the researcher merely observes what is happening or what has happened in the
past and tries to draw conclusions based on these observations.
2-Experimental study: the researcher manipulates one of variables and tries to determine how the
manipulation influences other variables.
4
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
2nd LECTURE: ORGANIZING AND GRAPHING DATA
1- Qualitative data
a- Frequency distribution for qualitative data: a frequency distribution for qualitative data lists
all categories and the number of elements that belong to each category.
Example (1): A sample of 30 employees from large companies was selected, and these employees
were asked how stressful their jobs were. The responses of these employees are recorded. Where
very means very stressful, somewhat means somewhat stressful, and none stands for not stressful
at all.
i- Bar Graph: A graph made of bars whose heights represent the frequencies of respective
categories. The graph bar can be also drawn simply for relative frequency and percentage by marking
the relative frequency or percentage instead of the class frequencies on the vertical axis.
ii- Pie Chart: it is more commonly used to display percentage, although it can use to display
frequencies or relative frequencies. The Whole Pie (or circle) represents the total sample (population).
Then we divide the Pie into different portions that represent the different categories.
5
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Example (2): Construct bar graph and pie chart for data in example (1).
- To construct bar graph, we mark the various categories on horizontal axis with same width and
marked the frequencies on the vertical axis and then draw one bar for each category and leave a
small gap between adjacent bar.
Bar Graph
16
14
12
10
Frequancy
8
6
4
2
0
Very Some what None
Stress in job
- To construct pie chart, we multiply 360 by relative frequency of each category to obtain the degree
measure of angle for the corresponding category.
6
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
2- Quantitative data
a- Frequency distribution for quantitative data: a frequency distribution for quantitative data lists
all the classes and the values that belong to each class. Data presented in the form of a frequency
distribution are called grouped data. When we need to construct the frequency distribution, we need
to make the following steps:
2- Number of classes: usually varies between (5 to 20) and we can use this equation:
C = 1 + 3.3 log n
Where:
Note: usually the approximate class width is rounded to a convenient number, which is then used
as the class width.
Example (3): the following data represent the length of (14) concrete beams, measured to the nearest
(1cm). (use 6 classes):
99, 96, 94, 92, 98, 89, 99, 101, 104, 102, 101, 106, 107, 112
Solution:
- Re- arranged data: 89, 92, 94, 96, 98, 99, 99, 101, 101, 102, 104, 106, 107, 112
- The upper limit of first class = 89 +3 = 92. (class width – accuracy =4-1=3)
7
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Cumulative
Cumulative freq. Cumulative Cumulative
freq. (more than Relative Relative Relative
Class Class Freq. Mid. point
(less than) or equal) Frequency freq. freq.
limit boundaries (fi) (xm)
Upper limit Lower limit (less than) (more than
or equal)
After the data have been organized into a frequency distribution, they can be presented in graphic forms. The
purpose of graphs in statistics is to convey the data to the viewer in pictorial form. Statistical graphs can be
used to describe the data set or analyze it.
Histogram:
It is a graph that can be drawn for frequency distribution, relative frequency distribution, or percentage
distribution. To draw a histogram:
1- Mark classes on the horizontal axis and frequency (or relative frequency or percentages) on vertical
axis.
2- Draw a bar for each class so that its height represents the frequency of that class. In histogram the
bars are drawn adjacent to each other with no gap between them.
Example: Construct a frequency histogram, relative frequency histogram, and percentage histogram to
represent data shown in example (3)
8
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
5
4
Frequency (fi)
0
88.5 92.5 96.5 100.5 104.5 108.5 112.5
Length (cm)
Frequency histogram
0.35
0.3
0.25
Relative frequency
0.2
0.15
0.1
0.05
0
88.5 92.5 96.5 100.5 104.5 108.5 112.5
length (cm)
9
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
35
30
25
Percentage
20
15
10
0
88.5 92.5 96.5 100.5 104.5 108.5 112.5
length (cm)
Percentage histogram
4
freq. (fi)
0
82.5 86.5 90.5 94.5 98.5 102.5 106.5 110.5 114.5 118.5
length (cm)
10
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Ogives
It is a curve drawn for the cumulative frequency distribution, one advantage of an ogive is that it can
be used to approximate the cumulative frequency for any interval. Steps to draw an ogive are:
1- Mark variable on horizontal axis and the cumulative frequencies on the vertical axis.
2- Mark dots above the upper boundaries of various classes at height equal to the cumulative
frequency.
3- Joining consecutive points with straight lines. Note that the ogive starts at the lower boundary of
the first class and ends at the upper boundary of the last class.
Example: Draw an ogive for the cumulative frequency distribution (less than) for Exp (3).
14
12
cumulative frequency
10
8
6
4
2
0
84.5 88.5 92.5 96.5 100.5 104.5 108.5 112.5 116.5
Length (cm)
*********************************************************************************
H.W: Draw an ogive for the cumulative frequency distribution (more than or equal) for Example
(3).
11
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3rd LECTURE: NUMERICAL DESCRIPTIVE MEASURES
Solution:
∑𝑥 53+32+61+27+39+44+49+57
𝑥̅ = → = 45.25 𝑦𝑒𝑎𝑟𝑠
𝑛 8
∑ 𝑓. 𝑥𝑚
𝑥̅ =
𝑛
Where 𝑥𝑚 is the midpoint, and f is the frequency of a class.
Exp (2): Calculate the arithmetic mean for the frequency distribution table below:
12
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Solution:
i Class limit Freq. (fi) xm xm.f
1 10-12 4 11 44
2 13-15 12 14 168
3 16-18 20 17 340
4 19-21 14 20 280
n = 50 ∑= 832
∑ 𝑓.𝑥𝑚 832
𝑥̅ = , 𝑥̅ = = 16.64
𝑛 50
1.1.2. Geometric Mean: The geometric mean (G ̅) of a set of n positive values X1, X2,…,Xn is
defined as the positive nth root of their product..
̅ = n√x1 . x2 . x3 … … xn
G. M. = G
When n is large, the computation of the geometric mean becomes difficult as we have to extract the
nth root of the product of all the values. The arithmetic is simplified by the use of logarithms.
∑ log x
̅) =
log(G n
Example (3): Find the geometric mean of numbers: 45, 32, 37, 46, 39, 36, 41, 48, 36.
Solution:
̅ = n√x1 . x2 . x3 … … xn
G. M. = G
̅ = 9√45 × 32 × 37 × 46 × 39 × 36 × 41 × 48 × 36 = 39.68
G
x 45 32 37 46 39 36 41 48 36
Log(x) 1.653 1.505 1.568 1.663 1.591 1.556 1.613 1.681 1.556 ∑=14.3870
13
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
∑ log x
̅) =
log(G n
̅ = 10(1.5986) = 39.68
G
𝑛
̅ = √𝑥 𝑓1 . 𝑥 𝑓2 . 𝑥 𝑓3 ⋯ ⋯ 𝑥𝑚𝑛
G. M. = G
𝑓𝑛
𝑚1 𝑚2 𝑚3
̅) = ∑ 𝑓.log 𝑥𝑚
log(G
n
Exp. (4): Find the geometric mean for the frequency distribution table of Exp. (2).
Sol:
1.1.3. Harmonic mean is defined as the value obtained when the number of values in the data set is
divided by the sum of its reciprocals. Harmonic mean is applied when the set of observations is in the
form of fractions or has extreme values. Also, stability of the data set with outliers is more when
harmonic mean is applied.
̅=
𝐻. 𝑀 = 𝐻 𝑛
1
∑
𝑥
14
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (5): Find the Harmonic mean of data: 2, 3, 5, 7, and 60.
̅=
𝐻 𝑛
1
∑
𝑥
̅= 𝑛 5 5
𝐻 1 = 1 1 1 1 1 = 1.1929
= 4.1916
∑𝑥 + + + +
2 3 5 7 60
Exp. (6): Find the harmonic mean for the frequency distribution table of Exp (2):
1 10-12 4 11 0.3636
2 13-15 12 14 0.8571
3 16-18 20 17 1.1765
4 19-21 14 20 0.7000
n= 50 ∑= 3.0972
𝑛 50
̅=
𝐻. 𝑀 = 𝐻 𝑓 = = 16.143
∑ 3.0972
𝑥𝑚
Exp.(7): Compare between the arithmetic, geometric, and harmonic Mean for the frequency
distribution table of Exp. (2):
15
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (8): Consider the following set of data 5, 8, 12, 15, and 20. Compare between the arithmetic,
geometric, and harmonic mean for the data:
𝑛 5 5
̅=
𝐻 1 = 1 1 1 1 1 = = 9.524
∑ + + + + 0.525
𝑥 5 8 12 15 20
𝐺̅ = 101.0317 = 10.757
∑𝑥 5+8+12+15+20 60
𝑥̅ = = = = 12
𝑛 5 5
H.W: Compare between the arithmetic, geometric, and harmonic Mean for the following data.
*********************************************************************************
1.2. Median: is another important measure of central tendency, it is the value of the middle term in a
data set that has been ranked in increasing
Solution:
1- Rank data in an increasing order: 3 5 8 10 19
n+1 5+1
2- The position of the middle term: ( ) → =3
2 2
Median = 8
Solution:
9 11 14 15 17 23
n+1 6+1
( )= ( ) = 3.5, the position of the median is between third and fourth value.
2 2
14+15
Median = ( ) = 14.5
2
16
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
** For grouped data, the median is obtained by:
2- Find the class that contain the median. Class Median is the first class with the value of cumulative
frequency equal at least n/2.
n
(2 − F)
Median (M) = Lm + ×∆
fm
Lm: is the lower class boundary of the median class.
n: Total number of data.
F: The cumulative frequency of the class before the median class.
fm: The frequency of the median class.
Δ: the class width.
n/2 = 20/2= 10
(10 – 9)
M = 24.5 + × 5 = 25.5
5
i Class Class Freq. Cumulative
limit boundaries (fi) Frequency
1 10-14 9.5-14.5 2 2
2 15-19 14.5- 19.5 3 5
3 20-24 19.5- 24.5 4 9
4 25-29 24.5- 29.5 5 14
5 30-34 29.5- 34.5 3 17
6 35-39 34.5- 39.5 2 19
7 40-44 39.5- 44.5 1 20
17
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
1.3 Mode: is the value that occurs with highest frequency in a data set.
a- 1, 1, 2, 3, 4, 4, 4, 5, 6, 6. Mode =4 (unimodal )
b- 2, 3, 6, 8, 9 No mode.
fm − fm-1
Mode = L + ×Δ
(fm − fm-1) + (fm − fm+1)
*********************************************************************************
1.4. Mid- range: is the mean of the largest and the smallest values in a data set.
18
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4th LECTURE: MEASURES OF DISPERISON, POSITION AND SHAPE
2. Measures of dispersion: In statistics, to describe the data set accurately, statisticians must know
more than the measures of central tendency. Two data sets with the same mean may have completely
different variation or dispersion, so the measures that help us know about the spread of data set are
called the measures of dispersion such as:
2.1. Range.
2.2. Variance and standard deviation.
2.3. Coefficient of variation.
2.1. Range: The range is the simplest of the three measures and is defined now. The range is the
highest value minus the lowest value. The symbol R is used for the range.
b- Extremely large or extremely small data can significantly affect the range.
*********************************************************************************
Exp. (1): Calculate the range for the following data set: 5 -7 2 0 -9 16 10 7
Sol: -9 -7 0 2 5 7 10 16
R= 16 – (-9) = 25
a- Ungrouped data
∑(𝑥 − 𝜇)2
𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝜎2 =
𝑁
∑(𝑥 − 𝑥̅ )2
𝑆𝑎𝑚𝑝𝑙𝑒 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 =
𝑛−1
19
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (2): Find the sample variance, standard deviation and the range, for the amount of European auto
sales for a sample of 6 years shown. The data are in millions of dollars.
Sol:
∑(𝑥−𝑥̅ )2
1 − 𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 = 𝑛−1
𝑠 2 = 1.278
b- Grouped data
∑ 𝑓(𝑥𝑚 −𝜇)2
𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝜎2 = 𝑁
∑ 𝑓(𝑥𝑚 −𝑥̅ )2
𝑆𝑎𝑚𝑝𝑙𝑒 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 = 𝑛−1
Exp (3): Find the variance and the standard deviation for the data in this frequency distribution table.
The data represent the number of miles that 20 runners ran during one week.
Class
5.5–10.5 10.5–15.5 15.5–20.5 20.5–25.5 25.5–30.5 30.5–35.5 35.5–40.5
boundaries
Freq. (fi) 1 2 3 5 4 3 2
20
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Sol:
i Class Freq. (fi) xm [Link] (xm-x-)2 fi(xm-x-)2
boundaries
1 5.5–10.5 1 8 8 272.25 272.25
2 10.5–15.5 2 13 26 132.25 264.5
3 15.5–20.5 3 18 54 42.25 126.75
4 20.5–25.5 5 23 115 2.25 11.25
5 25.5–30.5 4 28 112 12.25 49.00
6 30.5–35.5 3 33 99 72.25 216.75
7 35.5–40.5 2 38 76 182.25 364.5
∑490 ∑1305
∑ 𝑓.𝑥𝑚 490
𝑥̅ = , 𝑥̅ = = 24.5 ∑ 𝑓.𝑥 490
𝑛 20 𝑥− = 𝑛 𝑚 , 𝑥− = = 24.5
2 ∑ 𝑓(𝑥𝑚 −𝑥 − )2 1305 20
𝑠 = = = 68.68
𝑛−1 19
s = 8.28
𝑠
For samples: 𝐶𝑉𝑎𝑟 = × 100
𝑥̅
𝜎
For populations: 𝐶𝑉𝑎𝑟 = × 100
𝜇
Exp. (4): The mean of the number of sales of cars over a 3-month period is 87, and the standard
deviation is 5. The mean of the commissions is 5225 $, and the standard deviation is 773 $. Compare
the variations of the two.
Solution:
𝑠 5
𝐹𝑜𝑟 𝑠𝑎𝑙𝑒𝑠: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 5.75%
87
𝑠 773
𝐹𝑜𝑟 commissions: 𝐶𝑉𝑎𝑟 = × 100 = × 100 = 14.8%
𝑥̅ 5225
Since the coefficient of variation is larger for commissions, the commissions are more
variable than the sales.
Exp. (5): Suppose we have a sample of executive with mean age of 51 and standard division of 11.74
years, suppose also we know their average IQ is 125 with standard division of 20 points. How can be
compare deviations.
Solution:
𝑠 11.74
𝐹𝑜𝑟 𝑎𝑔𝑒: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 23%
51
𝑠 20
𝐹𝑜𝑟 IQ: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 16%
125
21
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
[Link] of Position
In addition to measures of central tendency and measures of variation, there are measures of
position or location. These measures include:
1-standard scores.
2- Quartiles.
3- Percentiles and deciles, and
They are used to locate the relative position of a data value in the data set.
3.1. Standard score (z score): it represents the number of standard deviations that a data value falls
above or below the mean.
𝑥−𝑥̅
𝐹𝑜𝑟 𝑠𝑎𝑚𝑝𝑙𝑒, 𝑧= 𝑆
𝑥−𝜇
𝐹𝑜𝑟 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛, 𝑧= 𝜎
Exp. (6): A student scored 65 on a calculus test that had a mean of 50 and a standard deviation of 10;
she scored 30 on a history test with a mean of 25 and a standard deviation of 5. Compare her relative
positions on the two tests.
Sol:
𝑥−𝑥̅ 65−50
𝑧1 = = = 1.5
𝑆 10
𝑥−𝑥̅ 30−25
𝑧2 = = =1
𝑆 5
Since the z score for calculus is larger, her relative position in the calculus class is
higher than her relative position in the history class.
*** Note that if the z score is positive, the score is above the mean. If the z score is 0, the
score is the same as the mean. And if the z score is negative, the score is below the mean.
3.2. Quartile
As the name implies, quartiles divide the data set into four equal parts. Therefore, the first quartile, Q1,
is the 25th percentile, the second quartile, Q2 is the 50th percentile (or the median), and the third quartile,
Q3, is the 75th percentile. The difference between the third and first quartiles is inter quartile range
(IQR).
IQR = Q3- Q1
Min. Q1 Q2 Q3 Max.
Median
22
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
For ungrouped data, the quartiles (Q1, Q2, and Q3) are calculated by:
Exp. (7): Find Q1, Q2, and Q3 for the data set 15, 13, 6, 5, 12, 50, 22, 18.
Sol:
1- Arrange the data in increasing order: 5, 6, 12, 13, 15, 18, 22, 50
Q2 is the median of all values → Q2 = (13+15)/2= 14
Q1 is the median of values ( 5, 6, 12, 13 ) → Q1 = (6+12)/2= 9.
Q3 is the median of values (15, 18, 22, 50) → Q3 = (18+22)/2= 20.
Exp.(8) : the following are the ages of nine employees of an insurance company
47 28 39 51 33 37 59 24 33
Box plots give a good graphical image of the concentration of the data. They also show how far the
extreme values are from most of the data. A box plot is constructed from five values:
23
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp.(9): Construct a box- plot for the following dataset
1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5
Ans.
Min.= 1 Q1= 2 Median = 7 Q3= 9 Max.= 11.5
Q1 Q2 Q3
Min Max
Exp.(10): Draw a box- plot for the data set {15, 9, 6, 8, 3, 14, 15, 13, 21}.
3.3. Percentiles: divide the data set into 100 equal groups. Each data set has 99 percentiles; data must
be ranked in increasing order to compute percentiles. The kth percentile is denoted by Pk , where k is
an integer range from (1 –99). For example, the 25th percentile which is denoted by P25, is defined to
be that numerical value such that at most 25% of the values are smaller than it and at most 75% are
larger than it in an ordered data set.
The percentile corresponding to a given value (x) is computed by using the formula:
𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑏𝑒𝑙𝑜𝑤 𝑥+0.5∗𝐹
Percentile = ∗ 100
𝑇𝑜𝑡𝑎𝑙 𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠
F: is the frequency of x.
24
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp.(11): A teacher gives a 20-point test to 10 students. Find the percentile rank of score of 12.
Scores: 18, 15, 12, 6, 8, 2, 3, 5, 20, 10.
Sol:
Ordered set: 2, 3, 5, 6, 8, 10, 12, 15, 18, 20.
69 93 70 53 92 75 85 70 68 76 88 70 77 82 85 82 80 100 96 85
Sol:
Ordered set: 53 68 69 70 70 70 75 76 77 80 82 82 85 85 85 88 92 93 96 100
25
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (13): For the following data set: 2, 3, 5, 6, 8, 10, 12, 15, 18, 20.
Find the values of the 25th and 80th percentile.
Sol:
a. n = 10, p = 25
c = (10×25)/100 = 2.5. Hence round up to c = 3.
Thus, the value of the 25th percentile is x = 5.
b. n = 10, p = 80
c = (10× 80)/100 = 8.
Thus, the value of the 80th percentile is
x = (15 + 18)/2 = 16.5.
3.4 Deciles: divide the distribution into 10 groups. They are denoted by D1, D2, etc.
Note that
D1 = P10 D2 = P20 D3 = P30 D4 = P40 ……etc.
Deciles can be found by using the formulas given for percentiles. Taken altogether then, these are the
relationships among percentiles, deciles, and quartiles.
Exp.(14): The following are test scores for a particular math class. Find the sixth deciles
44 56 58 62 64 64 70 72 72 72
74 74 75 78 78 79 80 82 82 84
86 87 88 90 92 95 96 96 98 100
Sol:
D6 = P60
n = 30, p = 60, c = (30×60)/100 = 18
The average of the 18th and 19th items represents the 6th deciles. D6= 82.
26
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Percentiles, deciles, and quartiles for grouped data: in order to find what value corresponds to a
specified i position such as the positions of Percentile, Quartile, or Decile in grouped data, the
following formulas must be used:
𝑖
n∗ ( )− 𝐹
4
𝑄𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖
n∗( )− 𝐹
100
𝑃𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖
n∗( )− 𝐹
10
𝐷𝑖= 𝐿 + ∗ ∆
𝑓𝑖
Where:
Exp. (15): The time taken by 20 workers in a factory to do a particular job were tabled as follow, find
Q2, P70, and D4.
cumulative
i Class boundaries Freq. (fi)
freq.
1 7.5–10.5 2 2
2 10.5–13.5 4 6
3 13.5–16.5 6 12
4 16.5–19.5 4 16
5 19.5–22.5 3 19
6 22.5–25.5 1 20
𝑖
n∗( )− 𝐹
4
𝑄𝑖= 𝐿 + 𝑓𝑖
∗ ∆
𝑖 2
For Q2 → n ∗ (4) = 20 ∗ 4 = 10
27
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
𝑖
n∗ ( )− 𝐹
100
𝑃𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖 70
for P70 → n ∗ (100) = 20 ∗ 100 = 14
14− 12
𝑃70= 16.5 + ∗ 3 = 18
4
𝑖
n∗( )− 𝐹
10
𝐷𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖 4
For D4 → n ∗ (10) = 20 ∗ 10 = 8
8− 6
𝐷4 = 13.5 + ∗ 3 = 14.5
6
H.W: The airborne speeds in miles per hours for 21 planes are shown in the following table. Find
the value that correspond to the 9th, 20th, 45th, and 75th percentiles.
28
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
[Link] of Shape (Skewness and Kurtosis)
The histogram can give you a general idea of the shape, but two numerical measures of shape
give a more precise evaluation: skewness tells you the amount and direction of skew (departure from
horizontal symmetry), and kurtosis tells you how tall and sharp the central peak is, relative to a standard
bell curve.
4.1. Skewness
The coefficient of Skewness (SK) is a measure for the degree of symmetry in the variable distribution.
1
𝑛
∑(𝑥𝑖 −𝑥̅ )3
**** For un grouped data: 𝑆𝐾 = 𝑆3
1
𝑛
∑ 𝑓.(𝑥𝑚 −𝑥̅ )3
**** For grouped data: 𝑆𝐾 = 𝑆3
1- In a normal distribution (symmetrical, SK= 0): the value of mean, median, and mode are identical,
and they lie at the center of distribution (Fig. 1).
2- In a positively skewed distribution (right skewed, SK > 0), the value of mean is largest, mode is
smallest, and the value of median lies between them (Fig. 2).
3- In a negatively skewed distribution (lift skewed, SK< 0), The value of mean is smallest, mode is
largest, and the value of median lies between them (Fig. 3).
Fig. (1): Normal distribution (SK= 0) Fig. (2): Right skewed SK > 0
29
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4.2. Kurtosis
The coefficient of Kurtosis (K) is a measure for the degree of peakedness/flatness in the variable
distribution curve.
1
𝑛
∑ (𝑥𝑖 −𝑥̅ )4
**** For un grouped data 𝐾=
𝑆4
1
𝑛
∑ 𝑓.(𝑥𝑚 −𝑥̅ )4
*** For grouped data 𝐾=
𝑆4
Types of Kurtosis
Exp. (15): Find the skewness and kurtosis coefficients for the following data set.
68 82 63 86 34 96 41 89 29 51 75 77 56 59 42
∑ 𝑥𝑖 948
𝑥̅ = = = 63.2
𝑛 15
∑(𝑥𝑖 −𝑥̅ )2 6150.4
𝑆2 = = = 439.314 , 𝑆 = 20.96
𝑛−1 14
1
𝑛
∑(𝑥𝑖 −𝑥̅ )3 −12291.4
𝑆𝐾 = = = −0.089
𝑆3 15× 9208.18
1
n
∑ (xi − x̅)4 4616940
K= 4
= = 1.594 …𝐾 < 3
S 15 × 193003.5
The data have left skewness and flat distribution
30
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (16): Describe the shape for distribution curve for this frequency distribution table:
Ans.:
1
n
∑ f.(xm −x̅)4 6415.01
K= = = 2.23
S4 50∗57.525
25
20
15
Freq. (f)
10
0
21.5
9.5 12.5 15.5 18.5
class boundaries
31
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
5th LECTURE: PROBABILITY AND COUNTING RULES
Probability: is a tool to guess the chance of an event will be happened. It denoted by (P)
Experiment: is the process by which an observation (or measurement) is obtained.
Outcome: is the result of a single trial of probability experiment.
Sample space: is the set of all possible outcomes of probability experiment, denote by (S).
Event: an event consists of one or more of the outcomes of an experiment. According to the number
of outcomes, an event may be a simple event or a compound event.
A simple event: is an event which consists of only a single outcome (observation) of the sample
space. Usually, simple events are denoted by E.
Venn diagram: is a closed geometric (such as rectangle, square or circle) that depicts all possible
outcomes for an experiment.
Exp. (1): A class contains a group of students (male, female), If two students are selected randomly:
2. What the outcomes of event " at most one male is selected". Draw the Venn diagram for event A.
Sol:
1- In this experiment, there are four outcomes: (MM), (MF), (FM), (FF)
32
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
S
MF FM
MM
FF
Exp. (2): In a group of people, some are in favor of civil engineering and others are against it. Two
persons are selected at random from this group and asked whether they are in favor of or against civil
engineering.
Solution:
F = event " a person is in favor of civil engineering".
A= event " person is against civil engineering"
AF AA
Venn diagram
33
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3) List all the outcomes included in each of following events:
Rule 4 The sum of the probabilities of all the outcomes in the sample space is 1.
1- Classical probability assumes that all outcomes in the sample space are equally likely to occur.
Number of outcomes in 𝐴
For compound event A p(A) = Total number of outcomes in the sample space
1
For simple event A p(E) = Total number of outcomes in the sample space
34
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Compute the probability of obtaining an even number in one roll of a dice.
Sol.: The outcomes of this experiment are: 1, 2, 3, 4, 5, and 6. All these outcomes are equally likely.
A = event “even number"
The outcomes of event A= {2, 4, 6}
Number of outcomes in 𝐴 3
p(A) = = = 0.5
Total number of outcomes in the sample space 6
Exp. (4): When a coin is tossed what is the probability of getting: a) head b) tail.
Sol: The outcomes of this experiment are: head, tail
E1 = event "getting head" ; E2 = event "getting tail"
1 1
p(E1 ) = =
Total number of outcomes in the sample space 2
1 1
p(E2 ) = =
Total number of outcomes in the sample space 2
*** In real life, the events of probability experiments are not always equally likely, so that, it is
needed to create another technique to compute probability in such experiments.
Relative Frequency Probability: Some time the classical probability rule is not suitable to apply to
compute probabilities, this because the various outcomes for the corresponding experiments are not
equally likely.
Exp. (5): Researcher for the American Automobile Association (AAA) asked 50 people who plan to
travel over the Thanksgiving holiday how they will get to their destination. The results can be
categorized in a frequency distribution as shown. Find the probability that a person will travel by
airplane over the Thanksgiving holiday.
Method Frequency
Drive 41
Fly 6
Train or bus 3
50
Solution:
E = event "person will travel by airplane".
𝑇ℎ𝑒 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝐸 6 3
𝑃(𝐸) = = =
Total number of trials 50 25
35
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3. Subjective probability based on an educated guess or estimate, employing opinions exp.: 80%
probability that an earthquake will occur in a certain area.
Two events are mutually exclusive when one event occurs, the other cannot and vice versa.
For any two mutually exclusive A and B, the probability that A or B will occur is:
P (A or B) = P (A) + P (B)
Exp. (6): Consider the following event for one roll of a die.
A= {2, 4, 6} , B = { 1, 3, 5} , C = {1, 2, 3, 4}
1. Are events A and B mutually exclusive?
2. Are events A and C mutually exclusive?
Sol:
S A
1 2
5 6
3 4
B
2. A and C have two common elements (2,4) → Non -mutually exclusive events
S
5 A
1 2
3 4 6
36
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (7): In a sample of 100 people, 42 had type O blood, 44 had type A blood, 10 had type B blood,
and 4 had type AB blood. Set up a frequency distribution and find the following probabilities for a
person to have:
1. Type A or B blood.
2. Neither type A nor Type O blood.
Type Frequency
A 44
B 10
AB 4
O 42
Sum 100
When probability of an event A is given and we have to find the probability of other event B based on
that event A, then the probability obtained is called conditional probability.
It is denoted by: P(B/A).
P(B/A) is read as probability of B given that A already occurred.
Exp. (8): At a large factory, the employees were classified according to their level of education and
whether they attend a sports event at least once a month as shown in the table. If an employee is
selected at random, find the probability that, the employee does not attend sports events given that
the he is a high school graduate.
Sol:
D= event " don’t attend"
H = event " high school graduate ".
12 3
P(D/H)=𝑃 (𝐷/𝐻) = = ≈ 0.43
28 7
37
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
INDEPENDENT AND DEPENDENT EVENTS:
Two events are said to be independent if the occurrence of one does not affect the
probability of the others. In other words, A and B are independent events if:
** NOTE: if one of these two condition is true, then the second is also true
** NOTE: if one of these two condition is not true, then the second is also not true.
Exp. (9): Suppose 100 employees in factory were asked whether they are in favor or against using
anew electric machine, the responses of these 100 employees:
In favor Against
Male 15 45
Female 4 36
Sol:
F = event " female"
A = event "In favor"
Exp. (10): A compressive strength test for 100 concrete cubes were made by using two machines A
and B as shown:
Are events “Fail test " and “Use machine A " independent?
Sol:
F = event " Fail test"
A = event " Use Machine A".
Fail test Success test Total
Machine A 9 51 60
Machine B 6 34 40
Total 15 85 100
38
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
COMPLEMENTARY EVENTS
The complementary of event, A` is the event that includes all the outcomes for an experiment that
are not in A.
S
A`
A P(A) + P(A`) = 1
Exp. (11): If the probability that a person lives in an industrialized country of the world is 1/5, find
the probability that a person does not live in an industrialized country.
Sol:
A= event " person lives in an industrialized country of the world"
A` = event " not living in an industrialized country"
P(A) + P(A`) = 1
P(A`) = 1- (1/5) = 4/5
39
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (12): A group of 2000 student were asked if they are in favor of or against cloning. The
following table gives the responses.
If one person is selected at random, find the probability that this person is:
1- In favor of cloning.
2- Against cloning.
3- In favor of cloning given the person is a female.
4- Male given the person has no opinion.
5- Are the events “Male” and “In favor” mutually exclusive? What about the events "In favor” and "
Against”?
6- Are the events “Female” and “No. Opinion” independent.
Sol:
M: event "Male" , F: event "Female".
I: event "In Favor", A: event "Against", N: event "No. Opinion"
40
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
6 th LECTURE: INTERSECTION & UNION OF EVENTS
INTERSECTION OF EVENTS
The intersection of event A and event B represents the collection of all outcomes that are common to
both A and B and is denoted by (A and B) or (A ∩ B) or (AB)
Exp. (1): The table below gives the classification of all employees of a company by gender and college
degree. If one of these employees is selected at random, what is the probability that this employee is
female and college graduate
Sol:
M: event "male"
F: event "female"
G: event "College Graduate"
N: event "Not College Graduate"
P (F) = 13/40 = 0.325,
P(F/G)= 4/11 = 0.364
P (F) ≠ P(F/G) → F and G are dependent events
P (F and G) = P(F) . P(G/F)
P(G/F) = 4/13 = 0.307.
P (F ∩ G) = (0.325) (0.307) = 0.1
41
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (2): A coin is flipped and a die is rolled. Find the probability of getting a head on the coin and a
4 on the die.
Sol:
A: event "get head"
B: event “get 4 on the die"
A, B are independent events
P(A) = ½= 0.5
P(B) = 1/6 = 0.167
P (A and B) = P(A) . P(B)
P (A and B) = (0.5) (0.167) = 0.0833
Exp. (3): Approximately 40 % of civil engineers using modern structural software for building
analysis. If three engineers were selected at random. Find the probability that all of them will use
modern structural analysis software
Sol:
A: event " the first engineer uses modern structural software"
B: event " the second engineer uses modern structural software"
C: event " the third engineer uses modern structural software"
A, B, C are independent events
UNION OF EVENTS
The union of two events A and B includes all outcomes that are either in A or in B or in both A and
B. It is denoted by (A or B) or (A U B).
Sol:
C: event "civil engineer". , E : event " electronic engineer".
42
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (5): A day of the week is selected at random. Find the probability that it is a weekend
Sol:
F: event " Friday " , S : event " Saturday"
P (F U S) = P(F) + P(S) – P (F and S)
P (F) = 1/ 7 , P( S ) = 1 / 7 , P( F and S) = 0
P (F or S) = 1/7 + 1/7 = 2/7
Exp. (6): In a hospital unit there are 8 nurses and 5 physicians; 7 nurses and 3 physicians are
females. If a staff person is selected, find the probability that the person is nurse or male.
Sol:
N: event "nurse" , M: event " male"
43
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
7 th LECTURE: PROBABILITY DISTRIBUTION OF A DISCRETE RANDAM
VARIABLE
Discrete Random Variables assumes values that can be counted, such as cars, houses, persons, etc.
Discrete Probability Distribution consists of the values a random variable can assume and the
corresponding probabilities of the values. The probabilities are determined theoretically or by
observation.
The probability distribution of discrete variable can be presented in the form of mathematical
formula , table, or graph.
1. The sum of the probabilities of all the events in the sample space must equal 1; that is :
∑ P(X) = 1.
2. The probability of each event in the sample space must be between or equal to 0 and 1.
0 ≤ P(X) ≤ 1.
b)
X 1 2 3 4
P(X) ¼ ¼ ¼ ¼
c)
X 8 9 12
P(X) 2/3 1/6 1/6
d)
X 1 3 5 7 9
P(X) 0.3 0.1 0.2 0.4 - 0.7
Sol:
a. No. It is not a probability distribution since P(X) cannot be negative or greater
than 1.
b. Yes. It is a probability distribution.
c. Yes. It is a probability distribution.
d. No, since P(X) ≠ -0.7.
44
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
The Binomial Probability Distribution
A binomial experiment is a probability experiment that satisfies the following four requirements:
1. There must be a fixed number of trials (n).
2. Each trial can have only two outcomes.
3. The outcomes of each trial must be independent of one another.
4. The probability of a success must remain the same for each trial.
Binomial formula:
𝑃(𝑋) = 𝑛 𝐶𝑥 𝑝 𝑥 𝑞 𝑛−𝑥
Where:
n : total number of trail.
nCx : number of ways to obtain x successes in n trail. nCx = n!/ x!.(n-x)!
P: probability of success.
q : probability of failure , q = 1-P
x : number of successes in n trail.
Sol:
n = 20
x = No. of defective device. P(x) = 5% .
x- = No. of good device. q (x) = 1- 0.05= 0.95
a) x = 5
𝑃(𝑋 = 5 ) = 𝑛 𝐶𝑥 𝑝 𝑥 𝑞 𝑛−𝑥
c) x ≥ 3
45
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
OR:
P (x ≥ 3) = 1- P (x < 3)
P (x ≥ 3) = 1- [(P (x =2) + P (x=1) + P (x=0)]
P (x ≥ 3) = 1- [ 0.189 + 0.377 + 0.358]
P (x ≥ 3) = 1- 0.924
P (x ≥ 3) = 0.076
Exp. (3): A manufacturer of metal pistons finds that on the average, 12% of his pistons are rejected
because they are either oversize or undersize. What is the probability that a batch of 10 pistons will
contain (a) no more than 2 rejects? (b) at least 2 rejects?
Ans.
Let x = number of rejected pistons (In this case, "success" means rejection!)
n = 10, p = 0.12, q = 0.88.
a.
𝑃(𝑥 = 0 ) = 10 𝐶0 0.12 0 0.8810 = 0.2785
𝑃(𝑥 = 1 ) = 10 𝐶1 0.12 1 0.889 = 0.379
𝑃(𝑥 = 2 ) = 10 𝐶2 0.12 2 0.888 = 0.233
P (x ≤ 2) = 0.2785+ 0.379 + 0.233 = 0.891
b. We could work out all the cases for X = 2, 3, 4, ..., 10, but it is much easier to proceed as follows:
P (x ≥ 2) = 1- P (x < 2)
P (x ≥ 2) = 1- (P (x =1) + P (x =0))
P (x ≥ 2) = 1- (0.379 + 0.2785) = 0.343
**** Note: the interval for λ and x must be equal. If they are not, the mean λ must be redefined to
make them equal.
Exp (4): A compressive strength apparatus in construction laboratory breaks down an average of
three times per month. Using the Poisson probability distribution formula, find the probability that
during the next month this apparatus will have:
a) exactly two breakdowns. b) at most one breakdown
Sol:
x: no. of break down during next month
λ: is the mean number of breaks down in month, λ = 3
a) x = 2
𝑒 −3 32 (0.04979 ).(9)
P( x = 2) = = = 0.224
2! 2
46
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
b) x ≤ 1
P( x ≤ 1) = 𝑝(𝑥 = 1) + 𝑝(𝑥 = 0)
𝑒 −3 . 31 𝑒 −3 .30
P( x ≤ 1) = +
1! 0!
(0.0498) 3 (0.0498) 1
P( x ≤ 1) = + = 0.1992
1 1
Exp. (5): If there are 200 typographical errors randomly distributed in a 500-page manuscript, find
the probability that a given page contains exactly 3 errors.
Sol:
x: no. of error in page.
λ: is the mean number of errors in page
λ = 200/500= 0.4
𝑒 −0.4 0.43
P( x = 3) = = 0.0072
3!
Exp (6): Vehicles pass through a junction on a busy road at an average rate of 300 per hour.
(a) Find the probability that none passes in a given minute.
(b) Find the probability that ten vehicles pass in two-minute period.
Solution:
x: number of cars per minute
The average number of cars per minute is: λ = 300/ 60 = 5
𝑒 −5 .50
a. P( x = 0) = = 0.00674
0!
47
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
8 th LECTURE: NORMAL DISTRIBUTION
1 x
2
1
f ( x) e 2
2
48
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
49
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (1): Find the area under the standard normal curve
between z =0 and z = 1.95
Sol:
Area between 0 and 1.95
= P (0 < z < 1.95)
= 0.4744
Sol:
a- P (1.19 < z < 2.12)
From table:
Area between 0 and 1.19 = 0.3830
Area between 0 and 2.12 = 0.4830
P (1.19 < z < 2.12)
= Area between 1.19 and 2.12
= 0.4830 – 0.3830
= 0.1
c- P (z > - 0.75)
The area on either side of the mean (z > 0) = 0.5
From table:
Area between 0 and -0.75 = 0.2734
= P (z > 0) + P (- 0.75 < z < 0)
= 0.5 + 0.2734
= 0.7734
50
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Let x be a continuous random variable that is normally distributed with mean 40 and standard
deviation equal to 5. Find the following probabilities:
1- P (x >55) 2- P (x < 49)
Sol:
1- For x = 55
𝑥−𝜇 55 − 40
𝑧= = =3
𝜎 5
P (x > 55)
= P (z > 3)
= 0.5 – 0.4987
= 0.0013
2- For x = 49
𝑥−𝜇 49 − 40
𝑧= = = 1.8
𝜎 5
P (x < 49)
= P (z < 1.8)
= 0.5 + 0.4641
= 0.9641
Exp. (4): The speed of vehicles passing through construction zone on a highway are normally
distributed with mean 46 miles per hour and standard deviation 4 miles per hour. Find the probability
of following:
a- The speed of vehicles more than 40 miles per hour.
b- The speed of vehicles between 50 and 55 miles per hour
Sol:
x is the speed of vehicle
1- For x = 40
𝑥−𝜇 40 − 46
𝑧= = = −1.5
𝜎 4
P (x > 40)
= p (z > -1.5)
= 0.4332 + 0.5
= 0.9332 = 93.32%
51
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
b- P (50 < x < 55)
For x = 50 z = (x - µ) / σ = (50 – 46) / 4 = 1
For x = 55 z = (x - µ) / σ = (55 – 46) / 4 = 2.25
P (50 < x < 55)
= p (1 < z < 2.25)
= 0.4878 – 0.3413
= 14.65 %
Exp. (5): A factory products concrete tiles with average length of 60 cm and standard deviation of 2.5
mm, find the probability of the rejected tiles if:
a- The accepted limited length is (59.5 – 60.5) cm.
b- The accepted limited length is to not more than 60.4 cm.
Sol:
x = the length of tile
a. For x = 60.5 z = (x - µ) / σ = (60.5 – 60) / 0.25 = +2
For x = 59.5 z = (x - µ) / σ = (59.5 – 60) / 0.25 = -2
P (59.5 < x < 60.5)
= P ( -2 < z < +2) area between (+2 and -2)
From table: the area between z = 0 and z = +2 is 0.4772
P (-2 < x < +2)
= 0.4772 *2 = 95.44% (Probability of accepted tiles)
The probability of the rejected tiles:
= 1- 0.9544
= 0.0456 = 4.56%
b. x = 60.4
𝑥−𝜇 60.4 − 60
𝑧= = = 1.6
𝜎 0.25
P (x ≤ 60.4)
= P (z ≤ 1.6)
= 0.5 + 0.4452
= 0.9452 (Probability of accepted tiles)
Thus, the probability of the rejected tiles:
=1 – 0.9452
= 0.0548
= 5.48%
52
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
9 th LECTURE: THE T- DISTRIBUTION AND HYPOTHESIS TEST
The t distribution is similar to the standard normal distribution in the following ways:
1. It is bell-shaped.
2. It is symmetric about the mean.
3. The mean, median, and mode are equal to 0 and are located at the center of the distribution.
4. The curve never touches the x axis.
The t distribution differs from the standard normal distribution in the following ways.
1. The variance is greater than 1.
2. The t distribution is a family of curves based on the degrees of freedom, which is a
number related to sample size. (df = n-1).
3. As the sample size increases, the t distribution approaches the normal distribution
*********************************************************************************
Exp. (1): Find the value of t (0.025, 14)
Sol:
α = 0.025, df = 14 from table t (0.025, 14) = 2.145.
Exp. (2): Find the value of t for 16 degrees of freedom and 0.05 area in the right and left tail of a
distribution curve.
53
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
54
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3).: For each of the following, find the area in the appropriate tail of the
t distribution.
a. t = 2.467 and df =28 Ans.: 0.01. right tail.
b. t = -2.878 and df = 18. Ans.: 0.005. left tail.
c. t = -2.145 and n= 15. Ans.: 0.025. left tail.
d. t = 2.508 and n= 23. Ans.: 0.01. right tail.
************************************************************************
Exp. (4).: Find the value of t for the t - distribution for each of the following.
The null hypothesis (H0): it is a statement about the population parameter that is assumed to be true
until it is declared false.
The alternative hypothesis (Ha): is a statement about a population parameter that will be true if the
null hypothesis is false.
H0 is true H0 is false
55
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Two-tailed test:
H0: µ = µo
H1: µ# µo
Left-tailed test:
H0: µ = µo
H1: µ ˂ µo
Right-tailed test
H0: µ = µo
H1: µ > µo
**If the population standard deviation (σ) is known, use the normal distribution to perform
hypothesis test:
(𝑥̅ − 𝜇0 )
𝑧 =
𝜎⁄√𝑛
(𝑥̅ − 𝜇0 )
𝑡 =
𝑠 ⁄ √𝑛
56
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Where:
𝑥̅ : the mean of sample.
s: is the sample standard deviation.
*********************************************************************************
Exp. (5): The average resistance force of the specific manufactured material is 100 units and their
standard deviation 15 units. The factory management claimed an improvement in productivity and the
durability of the materials had increased, so a sample of 36 pieces was taken and found that the
resistance force had actually increased to 106 units. Can the factory management claim be accepted
with α = 0.01?
Sol:
µo= 100
σ = 15
n= 36
𝑥̅ = 106
α = 0.01
Decision:
Reject H0 and accept H1
57
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (6) The mean score of statistics on the first statistics test is 65. A statistics lecturer thinks that the
mean score is higher than 65. He samples ten students and obtains the scores 65, 65, 70, 67, 66, 63,
63, 68, 72, 71. Perform a hypothesis test using a 5% level of significance. The data are assumed to be
from a normal distribution.
Sol:
µo= 65
n= 10
α = 0.05
H0: µ = µo
H1: µ > µo
The population standard deviation, (σ) is not known, and n < 30 → use t- distribution:
𝑥̅ = average score on the first statistics test.
65+65+70+⋯⋯+71
𝑥̅ = = 67
10
∑(𝑥𝑖 −𝑥 − )2
𝑠2 = 𝑛−1
→ s = 3.2
𝑥 − − 𝜇0 67− 65
𝑡 = = = 1.978
𝑠⁄√𝑛 3.2⁄√10
*************************************************************************************************************************
Exp. (7) The average daily amount of scrap from a particular manufacturing process is 25 kg with a
standard deviation of 1.6 kg. A modification process is attempted to reduce this amount. During 10-
day trial period, the average scrap was 23.5 Kg. Does the modification process reduce the scrap
amount, perform a hypothesis test using a 1 % level of significance?
Sol:
µo= 25
σ = 1.6
n= 10
𝑥̅ = 23.5
α = 0.01
Ho: µ = µo
H1: µ < µo
z
(σ) is known, use normal- distribution
𝑥̅ − 𝜇0 23.5 − 25
𝑧 = = = −2.96
σ⁄√𝑛 1.6⁄√10
58
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (8) Aconstruction materials factory produces a specific building material that has average expiry
date 12.5 months, in order to verify, a sample of 18 units were taken and found that the mean age at
which these units expire were 12.9 months with a standard deviation of 0.80 month. Using the 1%
significance level, can you conclude that the mean age of this material is different from 12.5 months?
Assume that the date at which all materials expire have an approximately normal distribution.
Sol:
µo= 12.5
n = 18
𝑥̅ =12.9 months
s = 0.8 month
α = 0 .01
H0: µ = µo
H1: µ# µo
The population standard deviation(σ) is not known, the sample size is small (n < 30) use t-
distribution:
𝑥̅ − 𝜇0 12.9− 12.5
𝑡 = = = 2.121
𝑠⁄√𝑛 0.8⁄√18
The value of the test statistic t = 2.121 falls between the two critical points, -2.898 and 2.898.
59
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (9) The average wind speed in a certain city is 8 miles per hour. A sample of 32 days has an
average wind speed of 8.2 miles per hour. The standard deviation of the sample is 0.6 mile per hour.
can you conclude that the average wind speed is different from 8 miles per hour, using the 5%
significance level?
Sol:
µo= 8
n = 32
𝑥̅ = 8.2
s = 0.6
α = 0 .05
H0: µ = µo
H1: µ# µo
Decision:
accept H0 and reject H1
*********************************************************************************
H.W: A random sample of 50 people was chosen from the population of a country. If the arithmetic
mean of the weekly income of the individuals in the sample is 80 $, if you know that the population
standard deviation of the individuals’ income is 15 $. How can we test the null hypothesis that the
average weekly income for the peoples of this country is equal to 75 $, versus the alternative hypothesis
that it is different from 75$? Use the 5% significance level
60
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
10 th LECTURE: CORRELATION AND LINEAR REGRESSION
Correlation: is a statistical technique used to determine the degree to which two quantitative variables
are related and finding the relation between them without being able to understand causal relationships.
Scatter diagram: is a mathematical diagram using Cartesian coordinates to display the relation
between two quantitative variables, one variable is called independent (X) and the second is called
dependent (Y)
The pattern of data is indicative of the type of relationship between your two variables:
1- positive relationship
2- negative relationship
3- no relationship
*********************************************************************************
Exp. (1): The table below contains the weights and Systolic Blood Pressure (SBP) for 10 persons.
Draw the scatter diagram and explain the type of relationship.
Weight.(kg) 67 69 85 83 74 81 97 92 114 85
SBP (mmHg) 120 125 140 160 130 180 150 140 200 130
61
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Sol: The relationship between variables is positive.
220
200
180
SBP (mm Hg)
160
140
120
100
80
60 70 80 90 100 110 120
Weight ( kg)
*********************************************************************************
Correlation Coefficient: Statistic showing the degree of relation between two variables.
Simple Correlation coefficient (r): It is also called Pearson's correlation or product moment correlation
coefficient. It measures the nature and strength between two variables of the quantitative type.
The value of (r) ranges between ( -1) and (+1).
The value of (r) denotes the strength of correlation between the two variables, as follows:
* If the sign is (+ve) this means the relation is positive or direct (an increase in one variable is
associated with an increase in the other variable and a decrease in one variable is associated with a
decrease in the other variable).
* if the sign is (-ve) this means negative or indirect relationship (which means an increase in one
variable is associated with a decrease in the other).
62
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
The correlation coefficient is calculated as:
∑𝑋∑𝑌
∑ 𝑋𝑌 −( )
𝑛
r= (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛
where n = the number of data points.
*********************************************************************************
Example (2): A sample of 6 children was selected, data about their age in years and weight in
kilograms was recorded as shown in the following table. It is required to find the correlation between
age and weight.
Age (year) 7 6 8 5 6 9
Weight (kg) 12 8 12 10 11 13
Sol:
∑𝑋∑𝑌
∑ 𝑋𝑌 −
𝑛
𝑟 = (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛
(41)(66)
461−
6
𝑟= (41)2 (66)2
= 0.759 (strong direct correlation)
√(291− )(742− )
6 6
14
13
12
Weight (kg)
11
10
9
8
7
4 5 6 7 8 9 10
Age (year)
63
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Compute the value of the correlation coefficient for the data obtained in the study of the
number of absences and the final grade of the seven students in the statistics class.
Sol:
Number of absences
No. Final grade (Y) XY X2 Y2
(X)
1 6 82 492 36 6724
2 2 86 172 4 7396
3 15 43 645 225 1849
4 9 74 666 81 5476
5 12 58 696 144 3364
6 5 90 450 25 8100
7 8 78 624 64 6084
Total ∑X= 57 ∑Y= 511 ∑XY= 3745 ∑ X2= 579 ∑ Y2= 38993
∑𝑋∑𝑌
∑ 𝑋𝑌 −
𝑛
𝑟= (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛
(57)(511)
3745−
7
𝑟= (57)2 (511)2
= −0.944 (strong indirect correlation)
√(579− )(38993− )
7 7
100
90
80
70
Final grade
60
50
40
30
20
10
0
0 2 4 6 8 10 12 14 16
Number of absences
64
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Regression: is a mathematical equation that describes the relationship between two or more variables.
Simple regression: includes only two variables: (y) dependent variable and (x) independent variable.
Linear regression: gives a straight – line relationship between two variables. The equation of linear
relationship:
y = a + bx
Where:
(y) dependent variable
(x) independent variable
(a) is y- intercept.
(b) represents the slope of the line
65
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Least squares method (SSE): it means that the sum of the squares of the vertical distances from each
point to the line is at a minimum.
we are able to construct a best fitting straight line to the scatter diagram points and then formulate a
regression equation in the form of:
y^ = a + bx
(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
2
∑𝑥𝑖 − 𝑛
𝑦 ^ = 𝑦̅ + 𝑏 ( 𝑥 − 𝑥̅ )
∑ 𝑦𝑖 ∑ 𝑥𝑖
𝑦̅ = , 𝑥̅ =
𝑛 𝑛
66
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (4): In a certain type of metal test specimen, the normal stress on a specimen is known to be
functionally related to the shear resistance. The following is a set of coded experimental data on
the two variables:
Normal Stress (X) 26.8 25.4 28.9 23.6 27.7 23.9 24.7 28.1 26.9 27.4 22.6 25.6
Shear Resistance
26.5 27.3 24.2 27.1 23.6 25.9 26.3 22.5 21.7 21.4 25.8 24.9
(Y)
Sol:
X Y XY X2
26.8 26.5 710.2 718.24
25.4 27.3 693.42 645.16
28.9 24.2 699.38 835.21
23.6 27.1 639.56 556.96
27.7 23.6 653.72 767.29
23.9 25.9 619.01 571.21
24.7 26.3 649.61 610.09
28.1 22.5 632.25 789.61
26.9 21.7 583.73 723.61
27.4 21.4 586.36 750.76
22.6 25.8 583.08 510.76
25.6 24.9 637.44 655.36
2
∑x= 311.6 ∑y= 297.2 ∑xy = 7687.76 ∑x = 8134.26
(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
∑𝑥𝑖 2 − 𝑛
(311.6)∗(297.2)
(7687.76)−
𝑏= 12
(311.6)2
= − 0.6861
(8134.26)−
12
∑ 𝑥𝑖 311.6
𝑥̅ = = 12
= 25.967
𝑛
∑ 𝑦𝑖 297.2
𝑦̅ = = 12
= 24.767
𝑛
𝑦^ = 𝑦− + 𝑏 ( 𝑥 − 𝑥
̅)
y ^ = 42.58 − 0.686x
67
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Coefficient of Determination (R2):
The coefficient of determination is a descriptive measure of the utility of the regression equation for
making predictions.
∑(𝑦𝑖 ^ − 𝑦 − )2
𝑅2 =
∑(𝑦𝑖 − 𝑦 − )2
Exp. (5).: Table below displays data on age and price for a sample of cars of a particular make
and model where x represents age, in years, and y represents predicted price, in hundreds of
dollars.
a- Determine the regression equation for the data.
b- Draw the scatter diagram and best fit line.
c- find the coefficient of determination
d- predict the price of a 3-year-old
Age (X) 5 4 6 5 5 5 6 6 2 7 7
Price (Y) 85 103 70 82 89 98 66 95 169 70 48
Sol:
(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
∑𝑥𝑖 2 − 𝑛
68
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
(58)∗(975)
(4732) −
𝑏= 11
(58)2
= -20.26
(326)−
11
∑ 𝑥𝑖 58
𝑥̅ = = = 5.273
𝑛 11
∑𝑦
𝑦̅ = 𝑛 𝑖 = 975
11
= 88.64
𝑦 ^ = 𝑦 − + 𝑏 ( 𝑥 − 𝑥̅ )
𝑦 ^ = 88.64 + (−20.26) ( 𝑥 − 5.273 )
y ^ = 195.47 − 20.26 x
Therefore, the regression line goes through the two points (2, 154.95) and (7, 53.65).
180
160
140
120
Price (100$)
100
80
60
40
20
0
0 1 2 3 4 5 6 7 8
Age(year)
69
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
∑(𝑦𝑖 ^ −𝑦 − )2 8284.80
c) 𝑅2 = ∑(𝑦𝑖 −𝑦 − )2
= = 0.853
9708.55
d)
𝑦 ^ = 195.47 − 20.26 𝑥
H.W: Temperatures (in degrees Fahrenheit) and Precipitation (in inches) are as follows:
(a) Estimate the liner regression equation.
(b) Estimate the precipitation for temperature of 70 𝐹 ° .
Temperatures (x) 86 81 83 89 80 74 64
Precipitation (y) 3.4 1.8 3.5 3.6 3.7 1.5 0.2
Ans.:
a) 𝑦 ^ = −8.994 + 0.1448 𝑥
b) 1.1
70