0% found this document useful (0 votes)
5 views71 pages

Engineering Statistics

The document provides an introduction to engineering statistics, covering key concepts such as types of data, measurement scales, data collection methods, and sampling techniques. It distinguishes between descriptive and inferential statistics and explains how to organize and graph data using various methods, including frequency distributions, histograms, and pie charts. Additionally, it outlines observational and experimental studies as statistical research methods.

Uploaded by

preciousenumah16
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views71 pages

Engineering Statistics

The document provides an introduction to engineering statistics, covering key concepts such as types of data, measurement scales, data collection methods, and sampling techniques. It distinguishes between descriptive and inferential statistics and explains how to organize and graph data using various methods, including frequency distributions, histograms, and pie charts. Additionally, it outlines observational and experimental studies as statistical research methods.

Uploaded by

preciousenumah16
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Ministry of Higher Education

& Scientific Research


University of Kerbala
College of Engineering
Civil Department

Engineering Statistics

By:

Dr. Layla A. Mohammed Saleh


Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
1st LEC TURE: INTRODUCTION OF STATISTICA

Statistics: - is the area of science that deals with collection, organization, analysis, and interpretation
of data.

A variable: - is a characteristic, or attribute under study that can assume different values like, length,
weight.

Data: - are the values (measurement or observations) that the variables can assume.

Random variables: - are variables whose values are determined by chance.

- A collection of data values forms a data set.


- Each value in the data set is called a data value or a datum.

Depending on how data are used, the body of knowledge called statistics is sometimes divided into
two main areas:

1-Descriptive Statistics.
2-Inferential Statistics.

Descriptive Statistics: consists of organization-summarization, and-presentation of data by using


tables, graphs and summary measures.

Inferential Statistics: deals with making decision, inferences, predictions, and forecasts about
populations based on results obtained from samples. It uses probability (the chance of on event
occurring) to achieve inferences.

Population: consists of all elements (human or otherwise) that are being studied.

Sample: is a group of elements selected from a population.

Variables: divided into


1. Qualitative: variables that can't be measured numerically but can classified into different categories
according to some characteristics or attribute such as color, gender of person.
2. Quantitative: variables that can be measured numerically. It can be divided into:
a. discrete: assume values that can be counted such as number of houses, cars accidents.
b- continuous: assume all values between any two specific values such as length, time, and weight.
- They are obtained by measuring. Data must be measured and rounded due to the limits of the
measuring device.
-Data of continuous variables are written in boundaries. boundaries are given in one additional
decimal place and always end with the digit 5.

1
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Measurement Scales: how variables are categorized, counted, or measured. There are four basic
levels: nominal, ordinal, interval, and ratio.

1- Nominal level of measurement: classifies data into mutually exclusive (non-overlapping),


exhausting categories in which no order or ranking can be imposed on the data. Exp., (Democratic,
Republican, Independent), Gender (male, female)

2- Ordinal level of measurement: classifies data into categories that can be ranked, however, precise
differences between the ranks do not exist. For example:
- Poor, Good, Excellent
- Small, Medium, Large
- 1st, 2nd, 3rd

3- Interval level of measurement: ranks data and precise differences between units of measure do
exist however, there is no meaningful zero. For example:
- Temperature (No zero because temperature exists even at 0° )
- IQ Scores (No zero because it doesn’t measure people without intelligence)

4- Ratio level of measurement: possess all the characteristics of interval measurement, and there
exists a true zero. In addition, true ratio exists when the same variable is measured on two different
members of the population. For example:
Height, weight, Time, Salary, Age

Data Collection: a variety of ways. One of the most common methods is through the use of surveys.

1- Telephone surveys
- less costly.
-people may be more candid in their opinions.
-not all people have a chance of being surveyed.

2-Mailed questionnaire surveys.


-cover a wider geographic area, less expensive to conduct.
- a low number of responses and inappropriate answers to questions.
-difficulty reading or understanding the questions.

3-Personal interview surveys.


- obtaining in-depth responses.
- need training in asking questions and recording responses.
- more costly.
-interviewer may be biased in selections of respondents.

4-Other ways of data collection


-surveying records, direct observation of situations.

2
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Methods of Sampling:
1-Random Sampling:
chance method or random numbers. Generating random numbers with a computer or calculator.

2-Systematic Sampling: each element has an equal probability of selection, but combinations of
elements have different probabilities.
Population size N, sample size= n, sampling interval k=N/n.
Randomly select a number j between 1 and k, sample element j and then every kth element thereafter,
j+k, j+2k, etc.

K =5

3- Stratified Sampling: divided population into groups (strata) according to some specified
characteristics such as age, grade level or income. The subsamples are randomly selected from each
strata.

3
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4- Cluster Sampling:
The population is divided into groups or clusters of elements usually geographic or organizational,
some of the groups are randomly chosen This method is useful when it is difficult or costly to develop
a complete list of the population members or when the population elements are widely dispersed
geographically

Statistical Studies:-

1-Observational Study: the researcher merely observes what is happening or what has happened in the
past and tries to draw conclusions based on these observations.

2-Experimental study: the researcher manipulates one of variables and tries to determine how the
manipulation influences other variables.

4
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
2nd LECTURE: ORGANIZING AND GRAPHING DATA

1- Qualitative data

a- Frequency distribution for qualitative data: a frequency distribution for qualitative data lists
all categories and the number of elements that belong to each category.

Example (1): A sample of 30 employees from large companies was selected, and these employees
were asked how stressful their jobs were. The responses of these employees are recorded. Where
very means very stressful, somewhat means somewhat stressful, and none stands for not stressful
at all.

Somewhat none Somewhat Very Very None


Very Somewhat Somewhat Very Somewhat Somewhat
Very Somewhat None very None Somewhat
Somewhat very Somewhat Somewhat Very None
Somewhat very Very Somewhat None Somewhat

For these data:


1-Construct a frequency distribution table.
2- Determine the relative frequency and percentage distributions.
Solution:

1- The variable in this example is stress on Job.


2- Classified the variable into three categories: very stressful, somewhat stressful, and not stressful.
3- Relative Frequency of categories= Frequency of category/Sum of all Frequency.
4- The percentage for a category = Relative Frequency of that category *100

Stress on Job Frequency Relative Frequency percentage


Very 10 10/30= 0.333 0.333*100= 33.3
Some what 14 14/30= 0.467 0.467*100= 46.7
None 6 6/30= 0.200 0.2*100= 20
Sum = 30 Sum = 1 Sum = 100

b - Graphical Presentation of Qualitative data:

i- Bar Graph: A graph made of bars whose heights represent the frequencies of respective
categories. The graph bar can be also drawn simply for relative frequency and percentage by marking
the relative frequency or percentage instead of the class frequencies on the vertical axis.

ii- Pie Chart: it is more commonly used to display percentage, although it can use to display
frequencies or relative frequencies. The Whole Pie (or circle) represents the total sample (population).
Then we divide the Pie into different portions that represent the different categories.

5
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Example (2): Construct bar graph and pie chart for data in example (1).

- To construct bar graph, we mark the various categories on horizontal axis with same width and
marked the frequencies on the vertical axis and then draw one bar for each category and leave a
small gap between adjacent bar.

Bar Graph
16
14
12
10
Frequancy

8
6
4
2
0
Very Some what None
Stress in job

- To construct pie chart, we multiply 360 by relative frequency of each category to obtain the degree
measure of angle for the corresponding category.

Stress on Job Relative Frequency Angle Size

Very 0.333 0.333*360 = 119.88

Some what 0.467 0.467*360 = 168.12

None 0.200 0.200*360 = 72.00

Sum = 1 Sum = 360

6
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
2- Quantitative data

a- Frequency distribution for quantitative data: a frequency distribution for quantitative data lists
all the classes and the values that belong to each class. Data presented in the form of a frequency
distribution are called grouped data. When we need to construct the frequency distribution, we need
to make the following steps:

1- Re- arranged data: data must be arranged ascension or descending.

2- Number of classes: usually varies between (5 to 20) and we can use this equation:

C = 1 + 3.3 log n

Where:

C: is the No. of class


n: is sample size

3- Range = largest value – smallest value.

4 - Approximate Class Width = Range/ Number of class.

Note: usually the approximate class width is rounded to a convenient number, which is then used
as the class width.

5- Class midpoint = (Lower limit+ upper limit)/2

Example (3): the following data represent the length of (14) concrete beams, measured to the nearest
(1cm). (use 6 classes):
99, 96, 94, 92, 98, 89, 99, 101, 104, 102, 101, 106, 107, 112

Solution:

- Re- arranged data: 89, 92, 94, 96, 98, 99, 99, 101, 101, 102, 104, 106, 107, 112

- Range = 112 – 89= 23

- Approximate width of class = 23/6 = 3.8 Say 4.

- Select a starting point for the lowest limit =89.

- The upper limit of first class = 89 +3 = 92. (class width – accuracy =4-1=3)

- Lower boundary of first class =89-( accuracy/2) =88.5.

- Upper boundary of first class =92+( accuracy/2) =92.5.

7
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

Cumulative
Cumulative freq. Cumulative Cumulative
freq. (more than Relative Relative Relative
Class Class Freq. Mid. point
(less than) or equal) Frequency freq. freq.
limit boundaries (fi) (xm)
Upper limit Lower limit (less than) (more than
or equal)

89-92 88.5-92.5 2 90.5 2 14 0.143 0.143 1


93-96 92.5-96.5 2 94.5 4 12 0.143 0.286 0.857
97-100 96.5-100.5 3 98.5 7 10 0.214 0.500 0.714
101-104 100.5-104.5 4 102.5 11 7 0.286 0.786 0.500
105-108 104.5-108.5 2 106.5 13 3 0.143 0.929 0.214
109-112 108.5-112.5 1 110.5 14 1 0.071 1 0.071
∑=14 ∑=1

Graphing quantitative data

After the data have been organized into a frequency distribution, they can be presented in graphic forms. The
purpose of graphs in statistics is to convey the data to the viewer in pictorial form. Statistical graphs can be
used to describe the data set or analyze it.

The three most commonly used graphs are:


1. The histogram.
2. The frequency polygon.
3. The cumulative frequency graph, or ogive.

Histogram:
It is a graph that can be drawn for frequency distribution, relative frequency distribution, or percentage
distribution. To draw a histogram:
1- Mark classes on the horizontal axis and frequency (or relative frequency or percentages) on vertical
axis.
2- Draw a bar for each class so that its height represents the frequency of that class. In histogram the
bars are drawn adjacent to each other with no gap between them.

Example: Construct a frequency histogram, relative frequency histogram, and percentage histogram to
represent data shown in example (3)

8
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
5

4
Frequency (fi)

0
88.5 92.5 96.5 100.5 104.5 108.5 112.5

Length (cm)

Frequency histogram

0.35

0.3

0.25
Relative frequency

0.2

0.15

0.1

0.05

0
88.5 92.5 96.5 100.5 104.5 108.5 112.5
length (cm)

Relative frequency histogram

9
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
35

30

25
Percentage

20

15

10

0
88.5 92.5 96.5 100.5 104.5 108.5 112.5
length (cm)

Percentage histogram

The frequency Polygon


It is another graph that can be used to display the quantitative data in graphic form. To draw the
frequency polygon:
1- Mark a dot above the midpoint of each class at a height equal to the frequency of that class
2- Mark two more classes, one at each end, and mark their midpoints. These two classes have zero
frequency.
3- Join the adjacent dots with straight lines. The resulting line graph is called frequency polygon.

Example: Draw the frequency polygon for example (3):

4
freq. (fi)

0
82.5 86.5 90.5 94.5 98.5 102.5 106.5 110.5 114.5 118.5
length (cm)

10
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Ogives

It is a curve drawn for the cumulative frequency distribution, one advantage of an ogive is that it can
be used to approximate the cumulative frequency for any interval. Steps to draw an ogive are:

1- Mark variable on horizontal axis and the cumulative frequencies on the vertical axis.

2- Mark dots above the upper boundaries of various classes at height equal to the cumulative
frequency.

3- Joining consecutive points with straight lines. Note that the ogive starts at the lower boundary of
the first class and ends at the upper boundary of the last class.

Example: Draw an ogive for the cumulative frequency distribution (less than) for Exp (3).

14
12
cumulative frequency

10
8
6
4
2
0
84.5 88.5 92.5 96.5 100.5 104.5 108.5 112.5 116.5

Length (cm)

Ogive for cumulative frequency distribution (less than)

*********************************************************************************

H.W: Draw an ogive for the cumulative frequency distribution (more than or equal) for Example
(3).

11
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3rd LECTURE: NUMERICAL DESCRIPTIVE MEASURES

The numerical descriptive measures are:


1- Measures of central tendency (mean, median, mode, mid-range)
2- Measures of dispersion (range, variance, standard deviation)
3- Measures of position (percentiles, deciles, quartiles).
4- Measures of shape (kurtosis and Skewness).

[Link] of central tendency


1.1 Mean
1.1.1. Arithmetic Mean: is the most frequently used measure of central tendency, it is denoted by (𝑥̅ )
for sample data and (µ) for population.

** For un grouped data, the arithmetic mean is obtained by:


𝑆𝑢𝑚 𝑜𝑓 𝑎𝑙𝑙 𝑣𝑎𝑙𝑢𝑠𝑒 ∑𝑥
Arithmetic 𝑀𝑒𝑎𝑛 𝑓𝑜𝑟 𝑠𝑎𝑚𝑝𝑙𝑒 𝑑𝑎𝑡𝑎 = → 𝑥̅ =
𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑛

Exp. (1): Following are the ages of eight employees of a company: 53 32 61 27 39 44 49 57


Find the arithmetic mean of these employees age.

Solution:
∑𝑥 53+32+61+27+39+44+49+57
𝑥̅ = → = 45.25 𝑦𝑒𝑎𝑟𝑠
𝑛 8

** For grouped data, the mean is obtained by:

∑ 𝑓. 𝑥𝑚
𝑥̅ =
𝑛
Where 𝑥𝑚 is the midpoint, and f is the frequency of a class.

Exp (2): Calculate the arithmetic mean for the frequency distribution table below:

i Class limit Freq. (fi)


1 10-12 4
2 13-15 12
3 16-18 20
4 19-21 14

12
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Solution:
i Class limit Freq. (fi) xm xm.f
1 10-12 4 11 44
2 13-15 12 14 168
3 16-18 20 17 340
4 19-21 14 20 280
n = 50 ∑= 832

∑ 𝑓.𝑥𝑚 832
𝑥̅ = , 𝑥̅ = = 16.64
𝑛 50

1.1.2. Geometric Mean: The geometric mean (G ̅) of a set of n positive values X1, X2,…,Xn is
defined as the positive nth root of their product..

̅ = n√x1 . x2 . x3 … … xn
G. M. = G

When n is large, the computation of the geometric mean becomes difficult as we have to extract the
nth root of the product of all the values. The arithmetic is simplified by the use of logarithms.

̅) = 1 (log 𝑥1 + log 𝑥2 + ⋯ ⋯ ⋯ + log n)


log(G 𝑛

∑ log x
̅) =
log(G n

̅ = anti log {∑ log x}


G n

Example (3): Find the geometric mean of numbers: 45, 32, 37, 46, 39, 36, 41, 48, 36.

Solution:

̅ = n√x1 . x2 . x3 … … xn
G. M. = G

̅ = 9√45 × 32 × 37 × 46 × 39 × 36 × 41 × 48 × 36 = 39.68
G

Also, we can use of logarithms, as shown below:

x 45 32 37 46 39 36 41 48 36
Log(x) 1.653 1.505 1.568 1.663 1.591 1.556 1.613 1.681 1.556 ∑=14.3870

13
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
∑ log x
̅) =
log(G n

̅ = anti log {14.387}


G 9

̅ = 10(1.5986) = 39.68
G

** Geometric mean for grouped data:

𝑛
̅ = √𝑥 𝑓1 . 𝑥 𝑓2 . 𝑥 𝑓3 ⋯ ⋯ 𝑥𝑚𝑛
G. M. = G
𝑓𝑛
𝑚1 𝑚2 𝑚3

Where 𝑥𝑚 is the midpoint and f the frequency of a class.

̅) = 1 (𝑓1 log 𝑥𝑚1 + 𝑓2 log 𝑥𝑚2 + ⋯ ⋯ ⋯ + 𝑓𝑛 log 𝑥𝑚𝑛 )


log(G 𝑛

̅) = ∑ 𝑓.log 𝑥𝑚
log(G
n

̅ = anti log {∑ 𝑓.log 𝑥𝑚}


G
n

Exp. (4): Find the geometric mean for the frequency distribution table of Exp. (2).

Sol:

i Class limit Freq. 𝑥𝑚 log 𝑥𝑚 𝑓. log 𝑥𝑚


(𝑓)
1 10-12 4 11 1.0414 4.1656
2 13-15 12 14 1.1461 13.7535
3 16-18 20 17 1.2304 24.6090
4 19-21 14 20 1.3010 18.2144
n= 50 ∑=60.7425

̅ = anti log {∑ 𝑓.log 𝑥𝑚 }


G n

̅ = anti log {60.7425} = 101.2149 = 16.402


G 50

1.1.3. Harmonic mean is defined as the value obtained when the number of values in the data set is
divided by the sum of its reciprocals. Harmonic mean is applied when the set of observations is in the
form of fractions or has extreme values. Also, stability of the data set with outliers is more when
harmonic mean is applied.

** For un grouped data, the harmonic mean is obtained by:

̅=
𝐻. 𝑀 = 𝐻 𝑛
1

𝑥

14
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (5): Find the Harmonic mean of data: 2, 3, 5, 7, and 60.

̅=
𝐻 𝑛
1

𝑥
̅= 𝑛 5 5
𝐻 1 = 1 1 1 1 1 = 1.1929
= 4.1916
∑𝑥 + + + +
2 3 5 7 60

** In case of grouped data (data grouped into a frequency distribution):


𝑛
̅=
𝐻. 𝑀 = 𝐻 𝑓

𝑥𝑚

(where 𝑥𝑚 represents the midpoints of the various classes).

Exp. (6): Find the harmonic mean for the frequency distribution table of Exp (2):

i Class limit Freq. (fi) xm f/ xm

1 10-12 4 11 0.3636
2 13-15 12 14 0.8571
3 16-18 20 17 1.1765
4 19-21 14 20 0.7000
n= 50 ∑= 3.0972

𝑛 50
̅=
𝐻. 𝑀 = 𝐻 𝑓 = = 16.143
∑ 3.0972
𝑥𝑚

Relation Between Arithmetic, Geometric and Harmonic means:

Arithmetic Mean > Geometric Mean > Harmonic Mean

Exp.(7): Compare between the arithmetic, geometric, and harmonic Mean for the frequency
distribution table of Exp. (2):

Arithmetic Mean = 16.64


Geometric Mean = 16.402
Harmonic Mean = 16.143
16.64 > 16.402 > 16.143

15
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (8): Consider the following set of data 5, 8, 12, 15, and 20. Compare between the arithmetic,
geometric, and harmonic mean for the data:
𝑛 5 5
̅=
𝐻 1 = 1 1 1 1 1 = = 9.524
∑ + + + + 0.525
𝑥 5 8 12 15 20

∑ 𝑙𝑜𝑔 𝑥 𝑙𝑜𝑔5+𝑙𝑜𝑔8+𝑙𝑜𝑔12+𝑙𝑜𝑔15+𝑙𝑜𝑔20 5.1584


𝑙𝑜 𝑔(𝐺̅ ) = = = = 1.0317
𝑛 5 5

𝐺̅ = 101.0317 = 10.757
∑𝑥 5+8+12+15+20 60
𝑥̅ = = = = 12
𝑛 5 5

12 > 10.757 > 9.524

H.W: Compare between the arithmetic, geometric, and harmonic Mean for the following data.

Class limit 10-14 15-19 20-24 25-29 30-34 35-39


Freq. 2 3 4 3 2 1

*********************************************************************************
1.2. Median: is another important measure of central tendency, it is the value of the middle term in a
data set that has been ranked in increasing

** For un grouped data, the median is obtained by:


- Rank the data set in an increasing order
n+1
- Find the position of the middle term which obtained as: ( )
2

Note: If the data set represents a population, replace n with N.

Exp. (9): Find the median for the Following data: 10 5 19 8 3

Solution:
1- Rank data in an increasing order: 3 5 8 10 19
n+1 5+1
2- The position of the middle term: ( ) → =3
2 2

Median = 8

Exp. (10): Find the median for the Following data: 11 15 23 9 14 17

Solution:

9 11 14 15 17 23
n+1 6+1
( )= ( ) = 3.5, the position of the median is between third and fourth value.
2 2

14+15
Median = ( ) = 14.5
2

16
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
** For grouped data, the median is obtained by:

1- Construct the cumulative frequency distribution.

2- Find the class that contain the median. Class Median is the first class with the value of cumulative
frequency equal at least n/2.

3- Find the median by using the following formula:

n
(2 − F)
Median (M) = Lm + ×∆
fm
Lm: is the lower class boundary of the median class.
n: Total number of data.
F: The cumulative frequency of the class before the median class.
fm: The frequency of the median class.
Δ: the class width.

Exp (10): Calculate the median for the frequency


i Class Freq.
distribution of the following data which represented the limit (fi)
1 10-14 2
average rain precipitation depth in 20-gauge stations 2 15-19 3
3 20-24 4
during one year. Data are record to the nearest mm. 4 25-29 5
5 30-34 3
Solution: 6 35-39 2
7 40-44 1
n = ∑fi = 20

n/2 = 20/2= 10

median class (24.5-29.5)


n
( −F)
2
Median (M) = Lm + ×∆
fm

(10 – 9)
M = 24.5 + × 5 = 25.5
5
i Class Class Freq. Cumulative
limit boundaries (fi) Frequency
1 10-14 9.5-14.5 2 2
2 15-19 14.5- 19.5 3 5
3 20-24 19.5- 24.5 4 9
4 25-29 24.5- 29.5 5 14
5 30-34 29.5- 34.5 3 17
6 35-39 34.5- 39.5 2 19
7 40-44 39.5- 44.5 1 20

17
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
1.3 Mode: is the value that occurs with highest frequency in a data set.

Exp (11): Find the mode for these data sets.

a- 1, 1, 2, 3, 4, 4, 4, 5, 6, 6. Mode =4 (unimodal )

b- 2, 3, 6, 8, 9 No mode.

c- 1, 2, 3, 3, 3, 4, 4, 4, 5, 6. Mode = 3,4 (bimodal )

** For un grouped data, the mode is obtained by:

fm − fm-1
Mode = L + ×Δ
(fm − fm-1) + (fm − fm+1)

L: The lower boundary of the modal class


fm-1: The frequency of the class before the modal class
fm: The frequency of the modal class
fm+1 : The frequency of the class after the modal class.
Δ: class width.
*********************************************************************************
i Class Class Freq.
Exp. (12): Find the mode for data in table. limit boundary (fi)
1 10-14 9.5-14.5 2
Solution: The mode class is (24.5-29.5)
2 15-19 14.5-19.5 3
𝑓 −𝑓𝑚−1 3 20-24 19.5-24.5 4
Mode = L + (𝑓 −𝑓 𝑚 )+(𝑓 ×∆
𝑚 𝑚−1 −𝑓
𝑚 𝑚+1 ) 4 25-29 24.5-29.5 5
(5−4)
5 30-34 29.5-34.5 3
Mode = 24.5 + (5−4)+(5−3) × 5 = 26.17 6 35-39 34.5-39.5 2
7 40-44 39.5-44.5 1

*********************************************************************************

1.4. Mid- range: is the mean of the largest and the smallest values in a data set.

Exp. (13): Find the mid-range of data: 10, 4, 8, 6, 2, 12, 20

Sol: 2, 4, 6, 8, 10, 12, 20

Mid- range = (smallest value + largest value)/2 = (2+20) / 2= 11

18
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4th LECTURE: MEASURES OF DISPERISON, POSITION AND SHAPE

2. Measures of dispersion: In statistics, to describe the data set accurately, statisticians must know
more than the measures of central tendency. Two data sets with the same mean may have completely
different variation or dispersion, so the measures that help us know about the spread of data set are
called the measures of dispersion such as:

2.1. Range.
2.2. Variance and standard deviation.
2.3. Coefficient of variation.

2.1. Range: The range is the simplest of the three measures and is defined now. The range is the
highest value minus the lowest value. The symbol R is used for the range.

R = highest value - lowest value


Disadvantage of range:

a- Based on two values only, largest and smallest.

b- Extremely large or extremely small data can significantly affect the range.

*********************************************************************************

Exp. (1): Calculate the range for the following data set: 5 -7 2 0 -9 16 10 7

Sol: -9 -7 0 2 5 7 10 16

R = highest value - lowest value

R= 16 – (-9) = 25

2.2 Variance and Standard Deviation

a- Ungrouped data

∑(𝑥 − 𝜇)2
𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝜎2 =
𝑁
∑(𝑥 − 𝑥̅ )2
𝑆𝑎𝑚𝑝𝑙𝑒 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 =
𝑛−1

𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛: 𝜎 = √ 𝜎2

𝑆𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛: 𝑠 = √ 𝑠2

19
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (2): Find the sample variance, standard deviation and the range, for the amount of European auto

sales for a sample of 6 years shown. The data are in millions of dollars.

11.2, 11.9, 12.0, 12.8, 13.4, 14.3

Sol:
∑(𝑥−𝑥̅ )2
1 − 𝑉𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 = 𝑛−1

x- = ∑x/n = 11.2+ 11.9+ 12.0+ 12.8+ 13.4+ 14.3/6= 75.6/6= 12.6


∑(x−𝑥̅ )2
s2 = n−1

(11.2−12.6)2 +(11.9−12.6)2 +(12−12.6)2 +(12.8−12.6)2 +(13.4−12.6)2 +(14.3−12.6)2


=
5

𝑠 2 = 1.278

2- Standard deviation: s = √1.278= 1.13


3- The range (R) = 14.3 – 11.2 = 3.1

b- Grouped data

∑ 𝑓(𝑥𝑚 −𝜇)2
𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝜎2 = 𝑁

∑ 𝑓(𝑥𝑚 −𝑥̅ )2
𝑆𝑎𝑚𝑝𝑙𝑒 𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒: 𝑠2 = 𝑛−1

𝑃𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛: 𝜎 = √ 𝜎2

𝑆𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛: 𝑠 = √ 𝑠2

Exp (3): Find the variance and the standard deviation for the data in this frequency distribution table.
The data represent the number of miles that 20 runners ran during one week.

Class
5.5–10.5 10.5–15.5 15.5–20.5 20.5–25.5 25.5–30.5 30.5–35.5 35.5–40.5
boundaries

Freq. (fi) 1 2 3 5 4 3 2

20
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Sol:
i Class Freq. (fi) xm [Link] (xm-x-)2 fi(xm-x-)2
boundaries
1 5.5–10.5 1 8 8 272.25 272.25
2 10.5–15.5 2 13 26 132.25 264.5
3 15.5–20.5 3 18 54 42.25 126.75
4 20.5–25.5 5 23 115 2.25 11.25
5 25.5–30.5 4 28 112 12.25 49.00
6 30.5–35.5 3 33 99 72.25 216.75
7 35.5–40.5 2 38 76 182.25 364.5
∑490 ∑1305

∑ 𝑓.𝑥𝑚 490
𝑥̅ = , 𝑥̅ = = 24.5 ∑ 𝑓.𝑥 490
𝑛 20 𝑥− = 𝑛 𝑚 , 𝑥− = = 24.5
2 ∑ 𝑓(𝑥𝑚 −𝑥 − )2 1305 20
𝑠 = = = 68.68
𝑛−1 19
s = 8.28

2.3 Coefficient of Variation (𝑪𝑽𝒂𝒓 )


A statistic that allows you to compare standard deviations when the units are different, it denoted by
(𝐶𝑉𝑎𝑟 ):

𝑠
For samples: 𝐶𝑉𝑎𝑟 = × 100
𝑥̅
𝜎
For populations: 𝐶𝑉𝑎𝑟 = × 100
𝜇

Exp. (4): The mean of the number of sales of cars over a 3-month period is 87, and the standard
deviation is 5. The mean of the commissions is 5225 $, and the standard deviation is 773 $. Compare
the variations of the two.

Solution:
𝑠 5
𝐹𝑜𝑟 𝑠𝑎𝑙𝑒𝑠: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 5.75%
87
𝑠 773
𝐹𝑜𝑟 commissions: 𝐶𝑉𝑎𝑟 = × 100 = × 100 = 14.8%
𝑥̅ 5225

Since the coefficient of variation is larger for commissions, the commissions are more
variable than the sales.

Exp. (5): Suppose we have a sample of executive with mean age of 51 and standard division of 11.74
years, suppose also we know their average IQ is 125 with standard division of 20 points. How can be
compare deviations.

Solution:
𝑠 11.74
𝐹𝑜𝑟 𝑎𝑔𝑒: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 23%
51
𝑠 20
𝐹𝑜𝑟 IQ: 𝐶𝑉𝑎𝑟 = 𝑥̅ × 100 = × 100 = 16%
125

The age is more variable than IQ

21
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
[Link] of Position

In addition to measures of central tendency and measures of variation, there are measures of
position or location. These measures include:
1-standard scores.
2- Quartiles.
3- Percentiles and deciles, and
They are used to locate the relative position of a data value in the data set.

3.1. Standard score (z score): it represents the number of standard deviations that a data value falls
above or below the mean.
𝑥−𝑥̅
𝐹𝑜𝑟 𝑠𝑎𝑚𝑝𝑙𝑒, 𝑧= 𝑆
𝑥−𝜇
𝐹𝑜𝑟 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛, 𝑧= 𝜎
Exp. (6): A student scored 65 on a calculus test that had a mean of 50 and a standard deviation of 10;
she scored 30 on a history test with a mean of 25 and a standard deviation of 5. Compare her relative
positions on the two tests.

Sol:
𝑥−𝑥̅ 65−50
𝑧1 = = = 1.5
𝑆 10
𝑥−𝑥̅ 30−25
𝑧2 = = =1
𝑆 5
Since the z score for calculus is larger, her relative position in the calculus class is
higher than her relative position in the history class.

*** Note that if the z score is positive, the score is above the mean. If the z score is 0, the
score is the same as the mean. And if the z score is negative, the score is below the mean.

3.2. Quartile

As the name implies, quartiles divide the data set into four equal parts. Therefore, the first quartile, Q1,
is the 25th percentile, the second quartile, Q2 is the 50th percentile (or the median), and the third quartile,
Q3, is the 75th percentile. The difference between the third and first quartiles is inter quartile range
(IQR).

IQR = Q3- Q1

25% of data 25% of data 25% of data 25% of data

Min. Q1 Q2 Q3 Max.
Median

22
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
For ungrouped data, the quartiles (Q1, Q2, and Q3) are calculated by:

1- Arrange the data in order from lowest to highest.


2- Find the median of the data values. This is the value for Q2.
3- Find the median of the data values that fall below Q2. This is the value for Q1.
4- Find the median of the data values that fall above Q2. This is the value for Q3.

Exp. (7): Find Q1, Q2, and Q3 for the data set 15, 13, 6, 5, 12, 50, 22, 18.

Sol:
1- Arrange the data in increasing order: 5, 6, 12, 13, 15, 18, 22, 50
Q2 is the median of all values → Q2 = (13+15)/2= 14
Q1 is the median of values ( 5, 6, 12, 13 ) → Q1 = (6+12)/2= 9.
Q3 is the median of values (15, 18, 22, 50) → Q3 = (18+22)/2= 20.

Exp.(8) : the following are the ages of nine employees of an insurance company
47 28 39 51 33 37 59 24 33

a- Find the values of three quartiles


b- When does the age 28 fall in relation to the ages of these employees.
c- Find the inter quartile range (IQR).
Sol:
Arrange the data in increasing order: 24, 28, 33, 33, 37, 39, 47, 51, 59
Q2 is the median of all values → Q2 = 37
Q1 is the median of values ( 24, 28, 33, 33) → Q1 = (28+33)/2= 30.5.
Q3 is the median of values (39, 47, 51, 59) → Q3 = (47+51)/2=49.
b- The age 28 fall in the first 25% of the ages.
c- The inter quartile range(IQR) = Q3- Q1 = 49-30.5= 18.5 years.

Box plots give a good graphical image of the concentration of the data. They also show how far the
extreme values are from most of the data. A box plot is constructed from five values:

1. The minimum value,


2. The first quartile (Q1)
3. The median (Q2)
4. The third quartile (Q3)
5. The maximum value.
We use these values to compare how close other data values are to them. To construct a box plot, use
a horizontal or vertical number line and a rectangular box. The smallest and largest data values label
the endpoints of the axis. The first quartile marks one end of the box and the third quartile marks the
other end of the box. Approximately the middle 50 percent of the data fall inside the box.

23
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp.(9): Construct a box- plot for the following dataset
1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5
Ans.
Min.= 1 Q1= 2 Median = 7 Q3= 9 Max.= 11.5

The following image shows the constructed box plot.

Q1 Q2 Q3
Min Max

Exp.(10): Draw a box- plot for the data set {15, 9, 6, 8, 3, 14, 15, 13, 21}.

Ans: Re arrange the data: 3, 6, 8, 9, 13, 14, 15, 15, 21

Min.= 3 Q1 = 7 Median: 13 Q3 = 15 Max.= 21.

3.3. Percentiles: divide the data set into 100 equal groups. Each data set has 99 percentiles; data must
be ranked in increasing order to compute percentiles. The kth percentile is denoted by Pk , where k is
an integer range from (1 –99). For example, the 25th percentile which is denoted by P25, is defined to
be that numerical value such that at most 25% of the values are smaller than it and at most 75% are
larger than it in an ordered data set.

For ungrouped data,

The percentile corresponding to a given value (x) is computed by using the formula:
𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑏𝑒𝑙𝑜𝑤 𝑥+0.5∗𝐹
Percentile = ∗ 100
𝑇𝑜𝑡𝑎𝑙 𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠

F: is the frequency of x.

24
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

Exp.(11): A teacher gives a 20-point test to 10 students. Find the percentile rank of score of 12.
Scores: 18, 15, 12, 6, 8, 2, 3, 5, 20, 10.

Sol:
Ordered set: 2, 3, 5, 6, 8, 10, 12, 15, 18, 20.

𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑏𝑒𝑙𝑜𝑤 𝑥+0.5∗𝐹


Percentile = ∗ 100
𝑇𝑜𝑡𝑎𝑙 𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠
6+0.5∗1
Percentile = ∗ 100 = 65 percentile
10

Student did better than 65% of the class.

Exp.(12): Consider the following data set:

69 93 70 53 92 75 85 70 68 76 88 70 77 82 85 82 80 100 96 85

a. Find the percentile rank of 82. b. Find the percentile rank of 68

Sol:
Ordered set: 53 68 69 70 70 70 75 76 77 80 82 82 85 85 85 88 92 93 96 100

𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠 𝑏𝑒𝑙𝑜𝑤 𝑥+0.5∗𝐹


Percentile = ∗ 100
𝑇𝑜𝑡𝑎𝑙 𝑁𝑜.𝑜𝑓 𝑣𝑎𝑙𝑢𝑒𝑠
10+0.5∗2
Percentile 𝑟𝑎𝑛𝑘 𝑜𝑓 82 = ∗ 100 = 55
20
1 + 0.5∗1
Percentile 𝑟𝑎𝑛𝑘 𝑜𝑓 68 = ∗ 100 = 7.5= 8
20

To Finding the value corresponding to a given percentile:

1. Let p be the percentile and n the sample size.


2. Arrange the data in order.
𝑛𝑝
3. Compute c: 𝑐=
100
4. If c is not a whole number, round up to the next whole number. If c is a whole number, use
the value halfway between c and c+1.
5. The value of c is the position value of the required percentile.

25
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

Exp. (13): For the following data set: 2, 3, 5, 6, 8, 10, 12, 15, 18, 20.
Find the values of the 25th and 80th percentile.

Sol:
a. n = 10, p = 25
c = (10×25)/100 = 2.5. Hence round up to c = 3.
Thus, the value of the 25th percentile is x = 5.

b. n = 10, p = 80
c = (10× 80)/100 = 8.
Thus, the value of the 80th percentile is
x = (15 + 18)/2 = 16.5.

3.4 Deciles: divide the distribution into 10 groups. They are denoted by D1, D2, etc.

Note that
D1 = P10 D2 = P20 D3 = P30 D4 = P40 ……etc.
Deciles can be found by using the formulas given for percentiles. Taken altogether then, these are the
relationships among percentiles, deciles, and quartiles.

Deciles are denoted by D1, D2, D3, . . . , D9,


and they correspond to P10, P20, P30, . . . , P90.

Quartiles are denoted by Q1, Q2, Q3


and they correspond to P25, P50, P75.

The median = P50 = Q2 =D5.

Exp.(14): The following are test scores for a particular math class. Find the sixth deciles

44 56 58 62 64 64 70 72 72 72
74 74 75 78 78 79 80 82 82 84
86 87 88 90 92 95 96 96 98 100
Sol:

D6 = P60
n = 30, p = 60, c = (30×60)/100 = 18
The average of the 18th and 19th items represents the 6th deciles. D6= 82.

26
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Percentiles, deciles, and quartiles for grouped data: in order to find what value corresponds to a
specified i position such as the positions of Percentile, Quartile, or Decile in grouped data, the
following formulas must be used:

𝑖
n∗ ( )− 𝐹
4
𝑄𝑖= 𝐿 + ∗ ∆
𝑓𝑖

𝑖
n∗( )− 𝐹
100
𝑃𝑖= 𝐿 + ∗ ∆
𝑓𝑖

𝑖
n∗( )− 𝐹
10
𝐷𝑖= 𝐿 + ∗ ∆
𝑓𝑖

Where:

Qi , Pi, Di are the quartile, Percentile and deciles of i position.


L: lower class boundary for the class contain i position.
n: total number of data.
F: The cumulative frequency of the class before the class contain i position
fm: The frequency of the class contains i position.
Δ: the class width.

Exp. (15): The time taken by 20 workers in a factory to do a particular job were tabled as follow, find
Q2, P70, and D4.

cumulative
i Class boundaries Freq. (fi)
freq.
1 7.5–10.5 2 2
2 10.5–13.5 4 6
3 13.5–16.5 6 12
4 16.5–19.5 4 16
5 19.5–22.5 3 19
6 22.5–25.5 1 20

𝑖
n∗( )− 𝐹
4
𝑄𝑖= 𝐿 + 𝑓𝑖
∗ ∆
𝑖 2
For Q2 → n ∗ (4) = 20 ∗ 4 = 10

the class boundary of Q2 is (13.5–16.5)


10 − 6
𝑄2= 13.5 + ∗ 3 = 15.5
6

27
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

𝑖
n∗ ( )− 𝐹
100
𝑃𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖 70
for P70 → n ∗ (100) = 20 ∗ 100 = 14

the class boundary of P70 is (16.5–19.5)

14− 12
𝑃70= 16.5 + ∗ 3 = 18
4

𝑖
n∗( )− 𝐹
10
𝐷𝑖= 𝐿 + ∗ ∆
𝑓𝑖
𝑖 4
For D4 → n ∗ (10) = 20 ∗ 10 = 8

the class boundary of D4 is (13.5–16.5)

8− 6
𝐷4 = 13.5 + ∗ 3 = 14.5
6

H.W: The airborne speeds in miles per hours for 21 planes are shown in the following table. Find
the value that correspond to the 9th, 20th, 45th, and 75th percentiles.

28
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
[Link] of Shape (Skewness and Kurtosis)

The histogram can give you a general idea of the shape, but two numerical measures of shape
give a more precise evaluation: skewness tells you the amount and direction of skew (departure from
horizontal symmetry), and kurtosis tells you how tall and sharp the central peak is, relative to a standard
bell curve.

4.1. Skewness
The coefficient of Skewness (SK) is a measure for the degree of symmetry in the variable distribution.
1
𝑛
∑(𝑥𝑖 −𝑥̅ )3
**** For un grouped data: 𝑆𝐾 = 𝑆3
1
𝑛
∑ 𝑓.(𝑥𝑚 −𝑥̅ )3
**** For grouped data: 𝑆𝐾 = 𝑆3

1- In a normal distribution (symmetrical, SK= 0): the value of mean, median, and mode are identical,
and they lie at the center of distribution (Fig. 1).

2- In a positively skewed distribution (right skewed, SK > 0), the value of mean is largest, mode is
smallest, and the value of median lies between them (Fig. 2).

3- In a negatively skewed distribution (lift skewed, SK< 0), The value of mean is smallest, mode is
largest, and the value of median lies between them (Fig. 3).

Fig. (1): Normal distribution (SK= 0) Fig. (2): Right skewed SK > 0

Fig. (3): Lift skewed (SK< 0)

29
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
4.2. Kurtosis

The coefficient of Kurtosis (K) is a measure for the degree of peakedness/flatness in the variable
distribution curve.
1
𝑛
∑ (𝑥𝑖 −𝑥̅ )4
**** For un grouped data 𝐾=
𝑆4

1
𝑛
∑ 𝑓.(𝑥𝑚 −𝑥̅ )4
*** For grouped data 𝐾=
𝑆4

Types of Kurtosis

1. K = 3 → Mesokurtic or normal distribution curve.


2. K >3 → leptokurtic (thin) distribution curve.
3. K < 3 → platykurtic (flat) distribution curve.

Exp. (15): Find the skewness and kurtosis coefficients for the following data set.

68 82 63 86 34 96 41 89 29 51 75 77 56 59 42
∑ 𝑥𝑖 948
𝑥̅ = = = 63.2
𝑛 15
∑(𝑥𝑖 −𝑥̅ )2 6150.4
𝑆2 = = = 439.314 , 𝑆 = 20.96
𝑛−1 14
1
𝑛
∑(𝑥𝑖 −𝑥̅ )3 −12291.4
𝑆𝐾 = = = −0.089
𝑆3 15× 9208.18

1
n
∑ (xi − x̅)4 4616940
K= 4
= = 1.594 …𝐾 < 3
S 15 × 193003.5
The data have left skewness and flat distribution

30
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (16): Describe the shape for distribution curve for this frequency distribution table:

Class limit 10-12 13-15 16- 18 19 - 21


Freq. 4 12 20 14

Ans.:

Class Class freq. xm [Link] f.(xm-x-)2 f.(xm-x-)3 f.(xm-x-)4


limit boundaries (fi)
10-12 9.5-12.5 4 11 44 127.24 -717.62 4047.40
13-15 12.5-15.5 12 14 168 83.64 -220.80 582.90
16-18 15.5-18.5 20 17 340 2.59 0.93 0.34
19-21 18.5-21.5 14 20 280 158.05 531.06 1784.37
n= 50 ∑= 832 371.52 -406.43 6415.01
∑ 𝑓.𝑥𝑚 832
𝑥̅ = , 𝑥̅ = = 16.64
𝑛 50

∑ 𝑓.(𝑥𝑖 −𝑥̅ )2 371.52


𝑆2 = = = 7.582 , 𝑆 = 2.754
𝑛−1 49
1
𝑛
∑ 𝑓.(𝑥𝑚 −𝑥̅ )3 −406.43
𝑆𝑘 = = = -0.389
𝑆3 50∗ 20.89

1
n
∑ f.(xm −x̅)4 6415.01
K= = = 2.23
S4 50∗57.525

Sk < 0 and K <3

The distribution curve is flat and the left skewed

25

20

15
Freq. (f)

10

0
21.5
9.5 12.5 15.5 18.5

class boundaries

31
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
5th LECTURE: PROBABILITY AND COUNTING RULES

Probability: is a tool to guess the chance of an event will be happened. It denoted by (P)
Experiment: is the process by which an observation (or measurement) is obtained.
Outcome: is the result of a single trial of probability experiment.
Sample space: is the set of all possible outcomes of probability experiment, denote by (S).

Examples of Experiments, Outcomes, and Sample spaces

Experiments Outcomes Sample spaces


Take a test Pass, Fail S= { Pass, Fail}
Select a student Male , Female S= { Male , Female }
Toss one coin Head, Tail S= { Head, Tail }
Toss two coins Head, Tail S= { HH, HT, TH, TT }
Roll one die 1,2,3,4,5,6 S= { 1,2,3,4,5,6}

Event: an event consists of one or more of the outcomes of an experiment. According to the number
of outcomes, an event may be a simple event or a compound event.

A simple event: is an event which consists of only a single outcome (observation) of the sample
space. Usually, simple events are denoted by E.

A compound event: consists of more than one outcome. It denoted by A, B, C…

Venn diagram: is a closed geometric (such as rectangle, square or circle) that depicts all possible
outcomes for an experiment.

Exp. (1): A class contains a group of students (male, female), If two students are selected randomly:

[Link] are the possible outcomes in this experiment?

2. What the outcomes of event " at most one male is selected". Draw the Venn diagram for event A.

Sol:

Let: M = male, F = female

1- In this experiment, there are four outcomes: (MM), (MF), (FM), (FF)

2- A= event "at most one male student is selected"

(event A will occur if either no male or one male is selected)

A = {MF, FM, FF}

Because event A contains more than one outcome, it is compound event.

32
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
S
MF FM
MM
FF

Venn diagram of event A

Exp. (2): In a group of people, some are in favor of civil engineering and others are against it. Two
persons are selected at random from this group and asked whether they are in favor of or against civil
engineering.

1- How many different outcomes are possible (S).


2- Draw a Venn diagram for this experiment.
3- List all the outcomes included in each of following events and mention whether they are simple or
compound events:
a) both persons are in favor of civil engineering
b) At most one person is against civil engineering
c) Exactly one person is in favor of civil engineering.

Solution:
F = event " a person is in favor of civil engineering".
A= event " person is against civil engineering"

1) All possible Outcomes:

S= {FF, FA, AF, AA}


Where:
FF= both persons are in favor of civil engineering
FA= the first person is in favor and the second is against.
AF = the first person is against and the second is in favor.
AA = both persons are against civil engineering.

2) Draw Venn diagram for this experiment


S FA FF

AF AA

Venn diagram
33
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3) List all the outcomes included in each of following events:

No Event Outcomes Type The reason

"Both persons are in favor of civil


1 {FF} Simple Only one outcome
engineering"

"At most one person is against


2 {FF, FA, AF} Compound More than one outcome
civil engineering"

"Exactly one person is in favor of


3 {FA, AF} Compound More than one outcome
civil engineering"

Basic probability rules which are helpful in solving probability problems:

Rule 1 The probability of any event E is a number between 0 and 1: 0 ≤ P(E) ≤1

Rule 2 If an event E cannot occur → P(E) = 0

Rule 3 If an event E is certain → P(E) = 1.

Rule 4 The sum of the probabilities of all the outcomes in the sample space is 1.

There are three basic interpretations of probability:


1. Classical probability
2. Empirical or relative frequency probability
3. Subjective probability

1- Classical probability assumes that all outcomes in the sample space are equally likely to occur.

Number of outcomes in 𝐴
For compound event A p(A) = Total number of outcomes in the sample space

1
For simple event A p(E) = Total number of outcomes in the sample space

34
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Compute the probability of obtaining an even number in one roll of a dice.

Sol.: The outcomes of this experiment are: 1, 2, 3, 4, 5, and 6. All these outcomes are equally likely.
A = event “even number"
The outcomes of event A= {2, 4, 6}

Number of outcomes in 𝐴 3
p(A) = = = 0.5
Total number of outcomes in the sample space 6

Exp. (4): When a coin is tossed what is the probability of getting: a) head b) tail.
Sol: The outcomes of this experiment are: head, tail
E1 = event "getting head" ; E2 = event "getting tail"

1 1
p(E1 ) = =
Total number of outcomes in the sample space 2

1 1
p(E2 ) = =
Total number of outcomes in the sample space 2

*** In real life, the events of probability experiments are not always equally likely, so that, it is
needed to create another technique to compute probability in such experiments.

Relative Frequency Probability: Some time the classical probability rule is not suitable to apply to
compute probabilities, this because the various outcomes for the corresponding experiments are not
equally likely.

The frequency of event f


P(E) = Relative frequency = =
Total number of trials N

Exp. (5): Researcher for the American Automobile Association (AAA) asked 50 people who plan to
travel over the Thanksgiving holiday how they will get to their destination. The results can be
categorized in a frequency distribution as shown. Find the probability that a person will travel by
airplane over the Thanksgiving holiday.

Method Frequency
Drive 41
Fly 6
Train or bus 3
50
Solution:
E = event "person will travel by airplane".

𝑇ℎ𝑒 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 𝑜𝑓 𝐸 6 3
𝑃(𝐸) = = =
Total number of trials 50 25

35
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
3. Subjective probability based on an educated guess or estimate, employing opinions exp.: 80%
probability that an earthquake will occur in a certain area.

MUTUALLY EXCLUSIVE E VENTS

Two events are mutually exclusive when one event occurs, the other cannot and vice versa.

For any two mutually exclusive A and B, the probability that A or B will occur is:

P (A or B) = P (A) + P (B)

Exp. (6): Consider the following event for one roll of a die.

A= {2, 4, 6} , B = { 1, 3, 5} , C = {1, 2, 3, 4}
1. Are events A and B mutually exclusive?
2. Are events A and C mutually exclusive?

Sol:

1. A and B have no common element → Mutually exclusive events.

S A

1 2
5 6
3 4
B

A and B aer mutually exclusive events

2. A and C have two common elements (2,4) → Non -mutually exclusive events

S
5 A

1 2

3 4 6

A and C are Non-mutually exclusive events

36
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (7): In a sample of 100 people, 42 had type O blood, 44 had type A blood, 10 had type B blood,
and 4 had type AB blood. Set up a frequency distribution and find the following probabilities for a
person to have:
1. Type A or B blood.
2. Neither type A nor Type O blood.

Solution: The frequency distribution is illustrated in the following table

Type Frequency
A 44
B 10
AB 4
O 42
Sum 100

P (A or B) = P(A) + P(B) = 44/100 + 10/100 = 27/50


P (Neither A nor O) = P (B or AB) = P(B) + P(AB) = 10/100 + 4/100= 0.14

THE CONDITIONAL PROBABILITY

When probability of an event A is given and we have to find the probability of other event B based on
that event A, then the probability obtained is called conditional probability.
It is denoted by: P(B/A).
P(B/A) is read as probability of B given that A already occurred.

Exp. (8): At a large factory, the employees were classified according to their level of education and
whether they attend a sports event at least once a month as shown in the table. If an employee is
selected at random, find the probability that, the employee does not attend sports events given that
the he is a high school graduate.

Sol:
D= event " don’t attend"
H = event " high school graduate ".
12 3
P(D/H)=𝑃 (𝐷/𝐻) = = ≈ 0.43
28 7

37
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
INDEPENDENT AND DEPENDENT EVENTS:

Two events are said to be independent if the occurrence of one does not affect the
probability of the others. In other words, A and B are independent events if:

P(A/B) = P(A) Independent events


Or P(B/A) = P(B) Independent events

** NOTE: if one of these two condition is true, then the second is also true
** NOTE: if one of these two condition is not true, then the second is also not true.

Exp. (9): Suppose 100 employees in factory were asked whether they are in favor or against using
anew electric machine, the responses of these 100 employees:

In favor Against
Male 15 45
Female 4 36

Are events "female" and "In favor" independent?

Sol:
F = event " female"
A = event "In favor"

In favor Against Total


Male 15 45 60
Female 4 36 40
Total 19 81 100

P(F)= 40/100= 0.4


P(F/A) = 4/19= 0.2105
P(F) ≠ P(F/A) the two events are dependent.

Exp. (10): A compressive strength test for 100 concrete cubes were made by using two machines A
and B as shown:

Fail test Success test


Machine A 9 51
Machine B 6 34

Are events “Fail test " and “Use machine A " independent?

Sol:
F = event " Fail test"
A = event " Use Machine A".
Fail test Success test Total
Machine A 9 51 60
Machine B 6 34 40
Total 15 85 100
38
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

P(F)= 15/100= 0.15


P(F/A) = 9/60 = 0.15
P(F) = P(F/A) the two events are independent.

COMPLEMENTARY EVENTS
The complementary of event, A` is the event that includes all the outcomes for an experiment that
are not in A.

S
A`

A P(A) + P(A`) = 1

Exp. (11): If the probability that a person lives in an industrialized country of the world is 1/5, find
the probability that a person does not live in an industrialized country.
Sol:
A= event " person lives in an industrialized country of the world"
A` = event " not living in an industrialized country"

P(A) + P(A`) = 1
P(A`) = 1- (1/5) = 4/5

39
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (12): A group of 2000 student were asked if they are in favor of or against cloning. The
following table gives the responses.

In Favor Against No. Opinion


Male 395 405 100
Female 300 680 120

If one person is selected at random, find the probability that this person is:
1- In favor of cloning.
2- Against cloning.
3- In favor of cloning given the person is a female.
4- Male given the person has no opinion.
5- Are the events “Male” and “In favor” mutually exclusive? What about the events "In favor” and "
Against”?
6- Are the events “Female” and “No. Opinion” independent.
Sol:
M: event "Male" , F: event "Female".
I: event "In Favor", A: event "Against", N: event "No. Opinion"

In Favor Against No. Opinion Total


Male 395 405 100 900
Female 300 680 120 1100
Total 695 1085 220 2000

1- P (I) = 695/2000 = 0.3475= 34.75 %


2- P (A) = 1085/2000 = 0.5425 = 54.25 %
3- P (I / F) = 300 / 1100 = 0.2727 = 27.27 %
4- P(M/N) = 100/220 = 0.4545 = 45.45%.
5- event “male” and “in favor " are not mutually exclusive.
events “against” and “in favor " are mutually exclusive.
6- P(F) = 1100/2000= 0.55, P(F/N) = 120/220= 0.54
P(F) ≠ P(F/N) → F and N are dependent events .

40
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
6 th LECTURE: INTERSECTION & UNION OF EVENTS

INTERSECTION OF EVENTS
The intersection of event A and event B represents the collection of all outcomes that are common to
both A and B and is denoted by (A and B) or (A ∩ B) or (AB)

P (A and B) = P(A) . P (B/A) ……….. for dependent events


P (A and B) = P(A) . P (B) ……….. for independent events

Exp. (1): The table below gives the classification of all employees of a company by gender and college
degree. If one of these employees is selected at random, what is the probability that this employee is
female and college graduate

College Graduate Not College Graduate Total


Male 7 20 27
Female 4 9 13
Total 11 29 40

Sol:
M: event "male"
F: event "female"
G: event "College Graduate"
N: event "Not College Graduate"
P (F) = 13/40 = 0.325,
P(F/G)= 4/11 = 0.364
P (F) ≠ P(F/G) → F and G are dependent events
P (F and G) = P(F) . P(G/F)
P(G/F) = 4/13 = 0.307.
P (F ∩ G) = (0.325) (0.307) = 0.1

41
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (2): A coin is flipped and a die is rolled. Find the probability of getting a head on the coin and a
4 on the die.
Sol:
A: event "get head"
B: event “get 4 on the die"
A, B are independent events
P(A) = ½= 0.5
P(B) = 1/6 = 0.167
P (A and B) = P(A) . P(B)
P (A and B) = (0.5) (0.167) = 0.0833

Exp. (3): Approximately 40 % of civil engineers using modern structural software for building
analysis. If three engineers were selected at random. Find the probability that all of them will use
modern structural analysis software

Sol:
A: event " the first engineer uses modern structural software"
B: event " the second engineer uses modern structural software"
C: event " the third engineer uses modern structural software"
A, B, C are independent events

P (A and B and C) = P(A). P(B). P(C) = (0.4) (0.4) (0.4) = 0.064

UNION OF EVENTS
The union of two events A and B includes all outcomes that are either in A or in B or in both A and
B. It is denoted by (A or B) or (A U B).

P (A U B) = P(A) + P(B) – P (A and B)

Exp. (4): An engineering company has the following number of employees:


Civil engineer 110, electronic engineer 750, Material engineer 250. If a research employee is selected
at random, find the probability that the employee is civil or electronic.

Sol:
C: event "civil engineer". , E : event " electronic engineer".

P (C U E) = P(C) + P(E) – P (C and E)


P (C) = 110/ 1110 = 0.1
P (E) = 750/ 1110 = 0.67
P (C and E) = 0
P (C or E) = 0.1 + 0.67 = 0.77.

42
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (5): A day of the week is selected at random. Find the probability that it is a weekend

Sol:
F: event " Friday " , S : event " Saturday"
P (F U S) = P(F) + P(S) – P (F and S)
P (F) = 1/ 7 , P( S ) = 1 / 7 , P( F and S) = 0
P (F or S) = 1/7 + 1/7 = 2/7

Exp. (6): In a hospital unit there are 8 nurses and 5 physicians; 7 nurses and 3 physicians are
females. If a staff person is selected, find the probability that the person is nurse or male.

Sol:
N: event "nurse" , M: event " male"

Staff Females Males Total


Nurses 7 1 8
Physicians 3 2 5
Total 10 3 13

P (N or M) = P(N) + P(M) - P(NM)


P(N) = 8/13
P(N/M) = 1/3
P(N) ≠ P(N/M) ) → N and M are dependent events
P(NM) = P(N). P(M/N) = (8/13) (1/8) = 1/13
P (N or M) = P(N) + P(M) - P(NM)
P (N or M) = (8/13) + (3/13) – (1/13) = 10/13

43
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
7 th LECTURE: PROBABILITY DISTRIBUTION OF A DISCRETE RANDAM
VARIABLE

Discrete Random Variables assumes values that can be counted, such as cars, houses, persons, etc.

Discrete Probability Distribution consists of the values a random variable can assume and the
corresponding probabilities of the values. The probabilities are determined theoretically or by
observation.

The probability distribution of discrete variable can be presented in the form of mathematical
formula , table, or graph.

Two Requirements for a Probability Distribution

1. The sum of the probabilities of all the events in the sample space must equal 1; that is :
∑ P(X) = 1.

2. The probability of each event in the sample space must be between or equal to 0 and 1.
0 ≤ P(X) ≤ 1.

Exp.(1) Determine whether each distribution is a probability distribution.


a)
X 4 6 8 10
P(X) - 0.6 0.2 0.7 1.5

b)
X 1 2 3 4
P(X) ¼ ¼ ¼ ¼

c)
X 8 9 12
P(X) 2/3 1/6 1/6

d)
X 1 3 5 7 9
P(X) 0.3 0.1 0.2 0.4 - 0.7

Sol:
a. No. It is not a probability distribution since P(X) cannot be negative or greater
than 1.
b. Yes. It is a probability distribution.
c. Yes. It is a probability distribution.
d. No, since P(X) ≠ -0.7.

44
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
The Binomial Probability Distribution

A binomial experiment is a probability experiment that satisfies the following four requirements:
1. There must be a fixed number of trials (n).
2. Each trial can have only two outcomes.
3. The outcomes of each trial must be independent of one another.
4. The probability of a success must remain the same for each trial.

Binomial formula:

𝑃(𝑋) = 𝑛 𝐶𝑥 𝑝 𝑥 𝑞 𝑛−𝑥
Where:
n : total number of trail.
nCx : number of ways to obtain x successes in n trail. nCx = n!/ x!.(n-x)!
P: probability of success.
q : probability of failure , q = 1-P
x : number of successes in n trail.

Exp.(2): In a large electronics company 5% of manufactured devices are defective If a random


sample of 20 devices is selected, find these probabilities :
a. There are exactly 5 devices are defective.
b. There are at most 3 devices are defective.
c. There are at least 3 devices are defective.

Sol:
n = 20
x = No. of defective device. P(x) = 5% .
x- = No. of good device. q (x) = 1- 0.05= 0.95

a) x = 5
𝑃(𝑋 = 5 ) = 𝑛 𝐶𝑥 𝑝 𝑥 𝑞 𝑛−𝑥

𝑃(5) = 20 𝐶5 0.05 5 0.95 20−5


𝑃(5) = (15504 )( 0.05 5 ) (0.95 15 ) = 0.002
b) x ≤ 3

p(x ≤ 3) = p(x = 0) + p(x = 1) + p(x = 2) + p(x = 3)

𝑃(𝑋 = 0 ) = 20 𝐶0 0.05 0 0.9520 = 0.358


𝑃(𝑋 = 1 ) = 20 𝐶1 0.05 1 0.9519 = 0.377
𝑃(𝑋 = 2 ) = 20 𝐶2 0.05 2 0.9518 = 0.189
𝑃(𝑋 = 3 ) = 20 𝐶3 0.05 3 0.9517 = 0.060

P (x ≤ 3) = 0.358 + 0.377 + 0.189 + 0.06 = 0.984

c) x ≥ 3

P (x ≥ 3) = p(x = 3) + p(x = 4) + p(x = 5) + p(x = 6) + p(x=7)+…………p(x=20)

45
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
OR:
P (x ≥ 3) = 1- P (x < 3)
P (x ≥ 3) = 1- [(P (x =2) + P (x=1) + P (x=0)]
P (x ≥ 3) = 1- [ 0.189 + 0.377 + 0.358]
P (x ≥ 3) = 1- 0.924
P (x ≥ 3) = 0.076

Exp. (3): A manufacturer of metal pistons finds that on the average, 12% of his pistons are rejected
because they are either oversize or undersize. What is the probability that a batch of 10 pistons will
contain (a) no more than 2 rejects? (b) at least 2 rejects?
Ans.
Let x = number of rejected pistons (In this case, "success" means rejection!)
n = 10, p = 0.12, q = 0.88.
a.
𝑃(𝑥 = 0 ) = 10 𝐶0 0.12 0 0.8810 = 0.2785
𝑃(𝑥 = 1 ) = 10 𝐶1 0.12 1 0.889 = 0.379
𝑃(𝑥 = 2 ) = 10 𝐶2 0.12 2 0.888 = 0.233
P (x ≤ 2) = 0.2785+ 0.379 + 0.233 = 0.891

b. We could work out all the cases for X = 2, 3, 4, ..., 10, but it is much easier to proceed as follows:
P (x ≥ 2) = 1- P (x < 2)
P (x ≥ 2) = 1- (P (x =1) + P (x =0))
P (x ≥ 2) = 1- (0.379 + 0.2785) = 0.343

The Poisson Distribution


A discrete probability distribution that is useful when n is large and p is small and when
the independent variables occur over a period of time is called the Poisson distribution.

Formula for Poisson distribution:


Formula for the Poisson Distribution
𝑒 −𝜆 𝜆𝑥
P(x) =
𝑥!

Where: λ is the mean number of occurrence in interval.

**** Note: the interval for λ and x must be equal. If they are not, the mean λ must be redefined to
make them equal.

Exp (4): A compressive strength apparatus in construction laboratory breaks down an average of
three times per month. Using the Poisson probability distribution formula, find the probability that
during the next month this apparatus will have:
a) exactly two breakdowns. b) at most one breakdown

Sol:
x: no. of break down during next month
λ: is the mean number of breaks down in month, λ = 3
a) x = 2
𝑒 −3 32 (0.04979 ).(9)
P( x = 2) = = = 0.224
2! 2

46
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
b) x ≤ 1
P( x ≤ 1) = 𝑝(𝑥 = 1) + 𝑝(𝑥 = 0)
𝑒 −3 . 31 𝑒 −3 .30
P( x ≤ 1) = +
1! 0!
(0.0498) 3 (0.0498) 1
P( x ≤ 1) = + = 0.1992
1 1

Exp. (5): If there are 200 typographical errors randomly distributed in a 500-page manuscript, find
the probability that a given page contains exactly 3 errors.
Sol:
x: no. of error in page.
λ: is the mean number of errors in page
λ = 200/500= 0.4
𝑒 −0.4 0.43
P( x = 3) = = 0.0072
3!

Exp (6): Vehicles pass through a junction on a busy road at an average rate of 300 per hour.
(a) Find the probability that none passes in a given minute.
(b) Find the probability that ten vehicles pass in two-minute period.

Solution:
x: number of cars per minute
The average number of cars per minute is: λ = 300/ 60 = 5
𝑒 −5 .50
a. P( x = 0) = = 0.00674
0!

b. x= number of cars per 2 minute


λ = 5 *2= 10
𝑒 −10 .1010
P( x = 10) = = 0.1251
10!

47
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
8 th LECTURE: NORMAL DISTRIBUTION

A normal distribution is a continuous, symmetric, bell-shaped distribution of a continues variable.

Properties of normal distribution


1. The mean, median, and mode are equal and are located at the center of the distribution.
2. A normal distribution curve is unimodal (i.e., it has only one mode).
3. The curve is symmetric about the mean, which is equivalent to saying that its shape is the same on
both sides of a vertical line passing through the center.
4. The curve is continuous; that is, there are no gaps or holes
5. The curve never touches the x axis.
6. The total area under a normal distribution curve is equal to 1.00, or 100%.
The mean µ and standard deviation σ are the parameters of normal distribution.

1  x 
2

1   
 
f ( x)  e 2
 2

The Standard Normal Distribution

The standard normal distribution is a normal distribution with µ = 0 and σ = 1.


The units marked on the horizontal axis of the standard normal distribution are denoted by z and
called z score.

48
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

49
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (1): Find the area under the standard normal curve
between z =0 and z = 1.95
Sol:
Area between 0 and 1.95
= P (0 < z < 1.95)
= 0.4744

Exp. (2): Find the following probability


a- P (1.19 < z < 2.12)
b- P ( -1.56 < z < 2.31)
c- P (z > - 0.75)

Sol:
a- P (1.19 < z < 2.12)
From table:
Area between 0 and 1.19 = 0.3830
Area between 0 and 2.12 = 0.4830
P (1.19 < z < 2.12)
= Area between 1.19 and 2.12
= 0.4830 – 0.3830
= 0.1

b- P ( -1.56 < z < 2.31)


From table:
Area between 0 and - 1.56 = 0.4406
Area between 0 and 2.31 = 0.4896
P ( -1.56 < z < 2.31)
= Area between (-1.56 and 2.31)
= 0.4406 + 0.4896
= 0.9302

c- P (z > - 0.75)
The area on either side of the mean (z > 0) = 0.5
From table:
Area between 0 and -0.75 = 0.2734
= P (z > 0) + P (- 0.75 < z < 0)
= 0.5 + 0.2734
= 0.7734

50
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Let x be a continuous random variable that is normally distributed with mean 40 and standard
deviation equal to 5. Find the following probabilities:
1- P (x >55) 2- P (x < 49)

Sol:
1- For x = 55
𝑥−𝜇 55 − 40
𝑧= = =3
𝜎 5
P (x > 55)
= P (z > 3)
= 0.5 – 0.4987
= 0.0013

2- For x = 49
𝑥−𝜇 49 − 40
𝑧= = = 1.8
𝜎 5
P (x < 49)
= P (z < 1.8)
= 0.5 + 0.4641
= 0.9641

Exp. (4): The speed of vehicles passing through construction zone on a highway are normally
distributed with mean 46 miles per hour and standard deviation 4 miles per hour. Find the probability
of following:
a- The speed of vehicles more than 40 miles per hour.
b- The speed of vehicles between 50 and 55 miles per hour
Sol:
x is the speed of vehicle
1- For x = 40
𝑥−𝜇 40 − 46
𝑧= = = −1.5
𝜎 4
P (x > 40)
= p (z > -1.5)
= 0.4332 + 0.5
= 0.9332 = 93.32%

51
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
b- P (50 < x < 55)
For x = 50 z = (x - µ) / σ = (50 – 46) / 4 = 1
For x = 55 z = (x - µ) / σ = (55 – 46) / 4 = 2.25
P (50 < x < 55)
= p (1 < z < 2.25)
= 0.4878 – 0.3413
= 14.65 %

Exp. (5): A factory products concrete tiles with average length of 60 cm and standard deviation of 2.5
mm, find the probability of the rejected tiles if:
a- The accepted limited length is (59.5 – 60.5) cm.
b- The accepted limited length is to not more than 60.4 cm.
Sol:
x = the length of tile
a. For x = 60.5 z = (x - µ) / σ = (60.5 – 60) / 0.25 = +2
For x = 59.5 z = (x - µ) / σ = (59.5 – 60) / 0.25 = -2
P (59.5 < x < 60.5)
= P ( -2 < z < +2) area between (+2 and -2)
From table: the area between z = 0 and z = +2 is 0.4772
P (-2 < x < +2)
= 0.4772 *2 = 95.44% (Probability of accepted tiles)
The probability of the rejected tiles:
= 1- 0.9544
= 0.0456 = 4.56%

b. x = 60.4
𝑥−𝜇 60.4 − 60
𝑧= = = 1.6
𝜎 0.25
P (x ≤ 60.4)
= P (z ≤ 1.6)
= 0.5 + 0.4452
= 0.9452 (Probability of accepted tiles)
Thus, the probability of the rejected tiles:
=1 – 0.9452
= 0.0548
= 5.48%
52
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
9 th LECTURE: THE T- DISTRIBUTION AND HYPOTHESIS TEST

The t distribution is similar to the standard normal distribution in the following ways:
1. It is bell-shaped.
2. It is symmetric about the mean.
3. The mean, median, and mode are equal to 0 and are located at the center of the distribution.
4. The curve never touches the x axis.

The t distribution differs from the standard normal distribution in the following ways.
1. The variance is greater than 1.
2. The t distribution is a family of curves based on the degrees of freedom, which is a
number related to sample size. (df = n-1).
3. As the sample size increases, the t distribution approaches the normal distribution

*********************************************************************************
Exp. (1): Find the value of t (0.025, 14)
Sol:
α = 0.025, df = 14 from table t (0.025, 14) = 2.145.

Exp. (2): Find the value of t for 16 degrees of freedom and 0.05 area in the right and left tail of a
distribution curve.

Sol: from t - distribution table

t (0.05, 16) = 1.746 (right tail)

Because of the symmetric shape


of the t - distribution curve:
t (0.05, 16) = -1.746 (left tail)

53
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

54
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3).: For each of the following, find the area in the appropriate tail of the
t distribution.
a. t = 2.467 and df =28 Ans.: 0.01. right tail.
b. t = -2.878 and df = 18. Ans.: 0.005. left tail.
c. t = -2.145 and n= 15. Ans.: 0.025. left tail.
d. t = 2.508 and n= 23. Ans.: 0.01. right tail.

************************************************************************
Exp. (4).: Find the value of t for the t - distribution for each of the following.

a. Area in the right tail 0.05 and df = 12.


b. Area in the left tail 0.025 and n = 26.
c. Area in the left tail 0.001 and df = 19.
d. Area in the right tail 0.005 and n = 24.
Sol:
a. t (0.05,12) = 1.782. right tail. b. t (0.025,25) = - 2.060 left tail
c. t (0.001, 19) = -3.579 left tail. d. t (0.005, 23) = 2.807 right tail.
*********************************************************************************
*********************************************************

THE HYPOTHESIS TEST

The null hypothesis (H0): it is a statement about the population parameter that is assumed to be true
until it is declared false.

The alternative hypothesis (Ha): is a statement about a population parameter that will be true if the
null hypothesis is false.

**Rejection of the null hypothesis when it is true is called a type I error.


α = probability of a type I error.
**No rejection of the null hypothesis when it is false is called a type II error.
β = probability of a type II error

H0 is true H0 is false

Do not reject H0 Correct decision type II error

Reject H0 type I error Correct decision

55
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------

Two-tailed test:
H0: µ = µo
H1: µ# µo

Left-tailed test:
H0: µ = µo
H1: µ ˂ µo

Right-tailed test
H0: µ = µo
H1: µ > µo

**If the population standard deviation (σ) is known, use the normal distribution to perform
hypothesis test:
(𝑥̅ − 𝜇0 )
𝑧 =
𝜎⁄√𝑛

** If the population standard deviation (σ) is not known:


1- The sample size is large ( n ≥ 30) → use normal - distribution.
2-The sample size is small ( n < 30) → use t- distribution:

(𝑥̅ − 𝜇0 )
𝑡 =
𝑠 ⁄ √𝑛

56
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Where:
𝑥̅ : the mean of sample.
s: is the sample standard deviation.

Steps to Perform a Test of Hypothesis:

[Link] the null and alternative hypotheses.


2. Identify the relevant test statistic and its distribution.
3. Compute from the data the value of the test statistic.
4. Construct the rejection region.
5. Compare the value computed in Step 3 to the rejection region constructed in
Step 4 and make a decision. Formulate the decision in the context of the problem.

*********************************************************************************

Exp. (5): The average resistance force of the specific manufactured material is 100 units and their
standard deviation 15 units. The factory management claimed an improvement in productivity and the
durability of the materials had increased, so a sample of 36 pieces was taken and found that the
resistance force had actually increased to 106 units. Can the factory management claim be accepted
with α = 0.01?

Sol:

µo= 100
σ = 15
n= 36
𝑥̅ = 106
α = 0.01

H0: µ = µo (no improvement)


H1: µ > µo (improvement increased)

σ is known use z- distribution.


𝑥̅ − 𝜇0
𝑧 =
σ⁄√𝑛
106− 100
𝑧 = = 2.4
15⁄√36

From standard normal distribution table:


z critical = z (0.49) = 2.33

Decision:
Reject H0 and accept H1

57
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (6) The mean score of statistics on the first statistics test is 65. A statistics lecturer thinks that the
mean score is higher than 65. He samples ten students and obtains the scores 65, 65, 70, 67, 66, 63,
63, 68, 72, 71. Perform a hypothesis test using a 5% level of significance. The data are assumed to be
from a normal distribution.

Sol:
µo= 65
n= 10
α = 0.05
H0: µ = µo
H1: µ > µo

The population standard deviation, (σ) is not known, and n < 30 → use t- distribution:
𝑥̅ = average score on the first statistics test.
65+65+70+⋯⋯+71
𝑥̅ = = 67
10
∑(𝑥𝑖 −𝑥 − )2
𝑠2 = 𝑛−1
→ s = 3.2
𝑥 − − 𝜇0 67− 65
𝑡 = = = 1.978
𝑠⁄√𝑛 3.2⁄√10

From t distribution table:


tc = t (0.05,9)= 1.833 t

Decision: accept H1 and reject H0

*************************************************************************************************************************
Exp. (7) The average daily amount of scrap from a particular manufacturing process is 25 kg with a
standard deviation of 1.6 kg. A modification process is attempted to reduce this amount. During 10-
day trial period, the average scrap was 23.5 Kg. Does the modification process reduce the scrap
amount, perform a hypothesis test using a 1 % level of significance?

Sol:
µo= 25
σ = 1.6
n= 10
𝑥̅ = 23.5
α = 0.01
Ho: µ = µo
H1: µ < µo
z
(σ) is known, use normal- distribution
𝑥̅ − 𝜇0 23.5 − 25
𝑧 = = = −2.96
σ⁄√𝑛 1.6⁄√10

From standard normal distribution table: z critical = z (0.49) = -2.33

Decision: accept H1 and reject H0

58
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (8) Aconstruction materials factory produces a specific building material that has average expiry
date 12.5 months, in order to verify, a sample of 18 units were taken and found that the mean age at
which these units expire were 12.9 months with a standard deviation of 0.80 month. Using the 1%
significance level, can you conclude that the mean age of this material is different from 12.5 months?
Assume that the date at which all materials expire have an approximately normal distribution.

Sol:
µo= 12.5
n = 18
𝑥̅ =12.9 months
s = 0.8 month
α = 0 .01

H0: µ = µo
H1: µ# µo

The population standard deviation(σ) is not known, the sample size is small (n < 30) use t-
distribution:
𝑥̅ − 𝜇0 12.9− 12.5
𝑡 = = = 2.121
𝑠⁄√𝑛 0.8⁄√18

From t- distribution table:


tc = t (0.005, 17) = 2.898 (right tail)
tc = t (0.005, 17) = -2.898 (left tail).

The value of the test statistic t = 2.121 falls between the two critical points, -2.898 and 2.898.

Decision: accept H0 and reject H1

59
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (9) The average wind speed in a certain city is 8 miles per hour. A sample of 32 days has an
average wind speed of 8.2 miles per hour. The standard deviation of the sample is 0.6 mile per hour.
can you conclude that the average wind speed is different from 8 miles per hour, using the 5%
significance level?

Sol:
µo= 8
n = 32
𝑥̅ = 8.2
s = 0.6
α = 0 .05

H0: µ = µo
H1: µ# µo

n > 30 → use z- distribution


𝑥 − − 𝜇0
𝑧 = s⁄√𝑛
8.2 − 8
𝑧 = = 1.89
0.6⁄√32

From standard normal table:


z c = z (0.475) = ±1.96.

Decision:
accept H0 and reject H1

*********************************************************************************
H.W: A random sample of 50 people was chosen from the population of a country. If the arithmetic
mean of the weekly income of the individuals in the sample is 80 $, if you know that the population
standard deviation of the individuals’ income is 15 $. How can we test the null hypothesis that the
average weekly income for the peoples of this country is equal to 75 $, versus the alternative hypothesis
that it is different from 75$? Use the 5% significance level

60
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
10 th LECTURE: CORRELATION AND LINEAR REGRESSION

Correlation: is a statistical technique used to determine the degree to which two quantitative variables
are related and finding the relation between them without being able to understand causal relationships.

Scatter diagram: is a mathematical diagram using Cartesian coordinates to display the relation
between two quantitative variables, one variable is called independent (X) and the second is called
dependent (Y)

The pattern of data is indicative of the type of relationship between your two variables:
1- positive relationship
2- negative relationship
3- no relationship

1- Positive relationship 2- Negative relationship 3- No relationship

*********************************************************************************

Exp. (1): The table below contains the weights and Systolic Blood Pressure (SBP) for 10 persons.
Draw the scatter diagram and explain the type of relationship.

Weight.(kg) 67 69 85 83 74 81 97 92 114 85

SBP (mmHg) 120 125 140 160 130 180 150 140 200 130

61
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Sol: The relationship between variables is positive.

220

200

180
SBP (mm Hg)

160

140

120

100

80
60 70 80 90 100 110 120
Weight ( kg)

*********************************************************************************
Correlation Coefficient: Statistic showing the degree of relation between two variables.

Simple Correlation coefficient (r): It is also called Pearson's correlation or product moment correlation
coefficient. It measures the nature and strength between two variables of the quantitative type.
The value of (r) ranges between ( -1) and (+1).
The value of (r) denotes the strength of correlation between the two variables, as follows:

Value of (r) Type of correlation


r=0 No correlation
0 < r < 0.25 Weak correlation
(0.25 ≤ r < 0.75) Intermediate correlation
(0.75 ≤ r < 1) Strong correlation
r=1 Perfect correlation

The sign of (r) denotes the nature of association.

* If the sign is (+ve) this means the relation is positive or direct (an increase in one variable is
associated with an increase in the other variable and a decrease in one variable is associated with a
decrease in the other variable).
* if the sign is (-ve) this means negative or indirect relationship (which means an increase in one
variable is associated with a decrease in the other).

62
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
The correlation coefficient is calculated as:

∑𝑋∑𝑌
∑ 𝑋𝑌 −( )
𝑛
r= (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛
where n = the number of data points.
*********************************************************************************

Example (2): A sample of 6 children was selected, data about their age in years and weight in
kilograms was recorded as shown in the following table. It is required to find the correlation between
age and weight.

Age (year) 7 6 8 5 6 9
Weight (kg) 12 8 12 10 11 13

Sol:

Age (year) Weight (kg)


No. XY X2 Y2
(X) (Y)
1 7 12 84 49 144
2 6 8 48 36 64
3 8 12 96 64 144
4 5 10 50 25 100
5 6 11 66 36 121
6 9 13 117 81 169
Total ∑X= 41 ∑Y = 66 ∑XY= 461 ∑ X2= 291 ∑ Y2= 742

∑𝑋∑𝑌
∑ 𝑋𝑌 −
𝑛
𝑟 = (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛
(41)(66)
461−
6
𝑟= (41)2 (66)2
= 0.759 (strong direct correlation)
√(291− )(742− )
6 6

14
13
12
Weight (kg)

11
10
9
8
7
4 5 6 7 8 9 10
Age (year)

63
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (3): Compute the value of the correlation coefficient for the data obtained in the study of the
number of absences and the final grade of the seven students in the statistics class.

Number of absences (X) 6 2 15 9 12 5 8


Final grade (Y) 82 86 43 74 58 90 78

Sol:

Number of absences
No. Final grade (Y) XY X2 Y2
(X)
1 6 82 492 36 6724
2 2 86 172 4 7396
3 15 43 645 225 1849
4 9 74 666 81 5476
5 12 58 696 144 3364
6 5 90 450 25 8100
7 8 78 624 64 6084
Total ∑X= 57 ∑Y= 511 ∑XY= 3745 ∑ X2= 579 ∑ Y2= 38993

∑𝑋∑𝑌
∑ 𝑋𝑌 −
𝑛
𝑟= (∑ 𝑥)2 (∑ 𝑌)2
√(∑ 𝑋 2 − )(∑ 𝑌 2 − )
𝑛 𝑛

(57)(511)
3745−
7
𝑟= (57)2 (511)2
= −0.944 (strong indirect correlation)
√(579− )(38993− )
7 7

100
90
80
70
Final grade

60
50
40
30
20
10
0
0 2 4 6 8 10 12 14 16
Number of absences

64
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Regression: is a mathematical equation that describes the relationship between two or more variables.

Simple regression: includes only two variables: (y) dependent variable and (x) independent variable.

Linear regression: gives a straight – line relationship between two variables. The equation of linear
relationship:
y = a + bx
Where:
(y) dependent variable
(x) independent variable
(a) is y- intercept.
(b) represents the slope of the line

= Y-intercept = Y-intercept a = Y-intercept

As shown in the figure below, a number of


straight lines can be drawn through the scatter
diagram each one of these lines will give
different values for a and b of equation.

In regression analysis, we try to find the best fit


line, such line provides the best relation
between dependent and independent variables
using the least squares method

65
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Least squares method (SSE): it means that the sum of the squares of the vertical distances from each
point to the line is at a minimum.

SSE = ∑d2 = ∑(y –y^ )2


y^ = a + bx
y: observed value
y^: predicted value.
d: "error" or residual.

we are able to construct a best fitting straight line to the scatter diagram points and then formulate a
regression equation in the form of:

y^ = a + bx
(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
2
∑𝑥𝑖 − 𝑛

𝑦 ^ = 𝑦̅ + 𝑏 ( 𝑥 − 𝑥̅ )
∑ 𝑦𝑖 ∑ 𝑥𝑖
𝑦̅ = , 𝑥̅ =
𝑛 𝑛

66
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Exp. (4): In a certain type of metal test specimen, the normal stress on a specimen is known to be
functionally related to the shear resistance. The following is a set of coded experimental data on
the two variables:

Normal Stress (X) 26.8 25.4 28.9 23.6 27.7 23.9 24.7 28.1 26.9 27.4 22.6 25.6
Shear Resistance
26.5 27.3 24.2 27.1 23.6 25.9 26.3 22.5 21.7 21.4 25.8 24.9
(Y)

(a) Estimate the liner regression equation.


(b) Estimate the shear resistance for a normal stress of 24.5.

Sol:

X Y XY X2
26.8 26.5 710.2 718.24
25.4 27.3 693.42 645.16
28.9 24.2 699.38 835.21
23.6 27.1 639.56 556.96
27.7 23.6 653.72 767.29
23.9 25.9 619.01 571.21
24.7 26.3 649.61 610.09
28.1 22.5 632.25 789.61
26.9 21.7 583.73 723.61
27.4 21.4 586.36 750.76
22.6 25.8 583.08 510.76
25.6 24.9 637.44 655.36
2
∑x= 311.6 ∑y= 297.2 ∑xy = 7687.76 ∑x = 8134.26

(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
∑𝑥𝑖 2 − 𝑛

(311.6)∗(297.2)
(7687.76)−
𝑏= 12
(311.6)2
= − 0.6861
(8134.26)−
12

∑ 𝑥𝑖 311.6
𝑥̅ = = 12
= 25.967
𝑛
∑ 𝑦𝑖 297.2
𝑦̅ = = 12
= 24.767
𝑛

𝑦^ = 𝑦− + 𝑏 ( 𝑥 − 𝑥
̅)

𝑦 ^ = 24.767 + (−0.686) ( 𝑥 − 25.967 )

y ^ = 42.58 − 0.686x

(b) The shear resistance for a normal stress of 24.5.

𝑦 ^ = 42.58 − 0.686 ∗ (24.5)


𝑦 ^ = 25.773

67
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
Coefficient of Determination (R2):
The coefficient of determination is a descriptive measure of the utility of the regression equation for
making predictions.

∑(𝑦𝑖 ^ − 𝑦 − )2
𝑅2 =
∑(𝑦𝑖 − 𝑦 − )2

**Note: The coefficient of determination, R 2, always lies between (0 and 1).


** R2 near 0 suggests that the regression equation is not very useful for making predictions.
** R2 near 1 suggests that the regression equation is quite useful for making predictions.

Exp. (5).: Table below displays data on age and price for a sample of cars of a particular make
and model where x represents age, in years, and y represents predicted price, in hundreds of
dollars.
a- Determine the regression equation for the data.
b- Draw the scatter diagram and best fit line.
c- find the coefficient of determination
d- predict the price of a 3-year-old

Age (X) 5 4 6 5 5 5 6 6 2 7 7
Price (Y) 85 103 70 82 89 98 66 95 169 70 48

Sol:

X Y XY X2 y^ (y^- y-)2 (y - y-)2


5 85 425 25 94.17 30.57 13.25
4 103 412 16 114.43 665.10 206.21
6 70 420 36 73.90 217.04 347.45
5 82 410 25 94.17 30.57 44.09
5 89 445 25 94.17 30.57 0.13
5 98 490 25 94.17 30.57 87.61
6 66 396 36 73.90 217.04 512.57
6 95 570 36 73.90 217.04 40.45
2 169 338 4 154.95 4397.23 6457.73
7 70 490 49 53.64 1224.54 347.45
7 48 336 49 53.64 1224.54 1651.61
∑X = 58 ∑Y= 975 ∑XY = 4732 ∑x = 326
2
∑ =8284.80 ∑ =9708.55

(∑𝑥𝑖 ∑𝑦𝑖 )
∑𝑥𝑖 𝑦𝑖 −
𝑛
𝑏= (∑ 𝑥𝑖 )2
∑𝑥𝑖 2 − 𝑛

68
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
(58)∗(975)
(4732) −
𝑏= 11
(58)2
= -20.26
(326)−
11

∑ 𝑥𝑖 58
𝑥̅ = = = 5.273
𝑛 11
∑𝑦
𝑦̅ = 𝑛 𝑖 = 975
11
= 88.64

𝑦 ^ = 𝑦 − + 𝑏 ( 𝑥 − 𝑥̅ )
𝑦 ^ = 88.64 + (−20.26) ( 𝑥 − 5.273 )
y ^ = 195.47 − 20.26 x

b) To graph the regression equation, we need to substitute two different x-values


in the regression equation to obtain two distinct points.

Let’s use the x-values 2 and 7.


at x= 2 ˆy = 195.47 – (20.26 * 2) = 154.95
at x= 7 ˆy = 195.47 – (20.26 * 7) = 53.65.

Therefore, the regression line goes through the two points (2, 154.95) and (7, 53.65).

180
160
140
120
Price (100$)

100
80
60
40
20
0
0 1 2 3 4 5 6 7 8
Age(year)

69
Engineering Statistics Dr. Layla A. Mohammed Saleh
--------------------------------------------------------------------------------------------------------
∑(𝑦𝑖 ^ −𝑦 − )2 8284.80
c) 𝑅2 = ∑(𝑦𝑖 −𝑦 − )2
= = 0.853
9708.55

d)
𝑦 ^ = 195.47 − 20.26 𝑥

𝑦 ^ = 195.47 − 20.26 ∗ 3 = 134.69

The price of a 3-year-old is 13469 $.

H.W: Temperatures (in degrees Fahrenheit) and Precipitation (in inches) are as follows:
(a) Estimate the liner regression equation.
(b) Estimate the precipitation for temperature of 70 𝐹 ° .

Temperatures (x) 86 81 83 89 80 74 64
Precipitation (y) 3.4 1.8 3.5 3.6 3.7 1.5 0.2

Ans.:
a) 𝑦 ^ = −8.994 + 0.1448 𝑥
b) 1.1

70

You might also like