Stat
Stat
Objectives:
Having studied this unit, you should be able to
understand statistics and basic terminologies
understand scales of measurement in statistics
understand the basic methods of data collection
Introduction
Most people become familiar with statistics through radio, television, newspapers and magazines. For instance,
one may find the following statements in a newspaper or reports. “The HIV prevalence rate in Ethiopia among
adults 15-49 years is 1.4 in 2005”; “Among older men, the mortality rate for smokers is twice the rate of those
who never smoked”; “The agricultural production increased by 5 percent this year”.
However, statistics is used in almost all fields of human endeavor to make a scientific decisions based on data.
For example, in public health an administrator would be concerned with the number of residents who contract a
new strain of flu virus during a certain year. In pharmacy, it is used to study the efficacy and potency of drugs.
To study plant life, a botanist has to relay on statistics to know the effect of temperature, rainfall and so on. In
general, statistics can be applied in business, social sciences, natural sciences and engineering.
The word statistics has several meanings. In the first place, it is a plural noun which describes a collection of
numerical data such as employment statistics, accident statistics, population statistics, economic statistics, and
agricultural statistics e t c. It is in this sense that the word 'statistics' is usually understood by a layman.
Secondly the word statistics as a singular noun is used to describe a branch of applied mathematics, whose
purpose is to provide methods of dealing with collections of data and extracting information from them in
compact form by tabulating, summarizing and analyzing the numerical data or a set of observations.
Classification of Statistics
Statistics may be divided into two main branches:
(1) Descriptive Statistics (2) Inferential Statistics
Descriptive statistics includes statistical methods involving the collection, presentation, and characterization of
a set of data in order to describe the various features of the data. In general, methods of descriptive statistics
include graphic methods (bar chart, pie chart, e t c) and numeric measures (mean, median, variance e t c).
Descriptive statistics do not, however, allow us to make conclusions beyond the data we have analyzed. They
are simply a way to describe data. Meaningful and pertinent information cannot be realized from raw data
unless summarized by the tools of descriptive statistics. Descriptive statistics, therefore, allow us to present the
data in a more meaningful way which allows interpretation of the data easily.
Inferential statistics includes statistical methods which facilitate estimation the characteristics of a population or
making decisions concerning a population on the basis of sample results. In this regard, methods like estimation
and hypothesis testing are examples of inferential statistics.
For example, a biologist collected blood samples of 10 students from biology department to study blood types.
Accordingly, the following data is obtained:
O A O AB A A O O B A O
Summary measures, for example, the proportion of students with blood type O in the sample is 50% is an
example of descriptive statistics. We can also describe the data using bar or pie charts.
1
However, if he/she wants to get information on the proportion of students with blood type O in the entire class,
he/she may use the sample proportion (50%) as an estimate of the corresponding value of the entire class. This
is an example of inferential statistics.
3
UNIT TWO: METHODS OF DATA COLLECTION AND PRESENTATION
Objectives:
After completing this unit you should be able to
organize data using frequency distribution.
present data using suitable graphs or diagrams.
Introduction
The amount of data collected in real life situations is often too large, thus we need some methods to organize it.
One of such methods is grouping, that is putting data into groups rather than treating each observation
individually. In fact, raw data provide little, if any, information to decision makers. Thus, we need a means of
converting the raw data into useful information. Hence, the purpose of this unit is to introduce tools used for
data presentation.
3. Higher ordered tables: results when we have more than two characteristics of classification. For instance,
we can classify the students who took introduction to statistics in 1998 by age, gender and faculty.
4
2.3 Frequency distributions
In this section, we will concentrate on some of the frequently used method of organizing data. The easiest
method of organizing data is using a frequency distribution, which converts raw data into a meaningful pattern
for statistical analysis.
The main uses of a frequency distribution are
to organize data in a meaningful, intelligible way.
to enable one to determine the nature or shape of the distribution; how the observations cluster around a
central value; and how the values spread around the center of the data.
to facilitate computational procedures for measures of average and spread.
to enable one to draw charts and graphs for the presentation of data.
to enable one to make comparisons between data sets.
Terminologies
Frequency distribution: a grouping of data into categories showing the number of observations in each
mutually exclusive category.
Array: data put in an ascending or descending order of magnitude.
Grouped data: data presented in the form of a frequency distribution.
Frequency: the number of observations corresponding to a fixed value or to a class of values.
Relative frequency: the number obtained when the frequency of a class is divided by total number of
observations.
Components of a frequency distribution
Class limits: the values of a variable which typically serve to identify the classes of a frequency distribution.
They are sometimes referred to as nominal or apparent limits. The smaller and the larger values are known as
the lower and the upper class limits, respectively. They should be selected in such a way that they have the same
number of significant places or units of measurement as the observations to be classified.
Class boundaries: the precise points which separate various classes rather than the values included in any one
of the classes. They are sometimes referred to as exact or true limits. They leave no space for ambiguity and
overlapping. A class boundary is located mid-way between the upper class limit of a class and the lower class
limit of the next higher class. They are carried out to one more decimal place than the class limits.
Class mark: the point which divides the class into two equal parts. This is also known as class mid-point. This
can be determined by dividing the sum of the two limits or the sum of the two boundaries by 2.
Class width: the length of a class.
Example 2.3: The following data are the weights in kg of 40 individuals participated in a diet program for
weight loss:
70 64 99 55 64 89 87 65 62 38 67 70 60 69 78 39 75 56 71 51
99 68 95 86 57 53 47 50 55 81 80 98 51 36 63 66 85 79 83 70
By grouping data into classes we can make the data much easier to read and understand. We group these data by
10s. The smallest weight is 36 kg, thus the 1rst class of weights is 31 kg up to, including, 40 kg.
Table 2.1: Distribution of weights.
Class Class boundary Count (Frequency)
31 – 40 30.5-40.5 3
41 – 50 40.5-50.5 2
51 – 60 50.5-60.5 8
61 – 70 60.5-70.5 12
71 – 80 70.5-80.5 5
81 – 90 80.5-90.5 6
91 – 100 90.5-100.5 4
Total 40
5
For this example, the first class is ‘31-40’. Lower limit of this class = 31; upper limit = 40. The lower class
boundary = 30.5; upper class boundary = 40.5. The width of the class = upper class boundary - lower class
boundary = 40.5-30.5 = 10. The class mark (class mid-point) of this class is (31+40)/2 = 35.5. The values 36,
39, 38 are included in this class. Therefore, the frequency of this class is 3.
Guidelines for constructing a frequency distribution
Find the range of the data
Range =R=Maximum value – Minimum value
Determine the number of classes.
Let K be the number of classes and n be the number of observations to be classified. There are two
alternatives to determine K.
1. Choose K to be between 5-15;
2. Use the following formula known as Sturgess’ formula given by:
K 1 3.322 log(n) and round K to the nearest integer.
Determine the class width , let it be W , by using
W=Range/K
Tally the observations, count and assign frequencies to the classes.
Example 2.4: The following data are on the number of minutes to travel from home to work for a group of
automobile workers: 28 25 48 37 41 19 32 26 16 23 23 29 36 31 26 21 32 25 31 43 35 42 38 33
28. Construct a frequency distribution for this data.
Solution:
Range = 48 – 16 =32
K=1+3.322 =5.64≈6
W=32/6=5.33 rounding up to the nearest integer i.e W=6.
Let the lower limit of the first class be 16 then the frequency distribution is as follows:
Class limit Class boundaries Tally Frequency
16-21 15.5-21.5 \\\ 3
22-27 21.5-27.5 \\\\\ \ 6
28-33 27.5-33.5 \\\\\ \\\ 8
34-39 33.5-39.5 \\\\ 4
40-45 39.5-45.5 \\\ 3
46-51 45.5-51.5 \ 1
Total 25
The final frequency distribution is shown in table below.
Definition 2.1: A relative frequency distribution is a distribution which specifies the frequency of a class
relative to the total frequency.
Example 2.5: Convert the above absolute frequency distribution in example 2.4 to a relative frequency
distribution.
Solution: First we find the relative frequency of each class. The relative frequency of a class is the frequency of
the class divided by the total number of observations. For instance the relative frequency of the first class is
3/25=0.12, the relative frequency of the second class is 6/25=0.24, and so on. Thus, the relative frequency
distribution is shown in the table below.
Definition 2.2: Cumulative frequency refers to the number of observations that are below a specified value or
that are above a specified value.
Note: Class boundaries are mostly used to obtain cumulative frequencies. Based on whether the observations
are bounded from above or from below, we can have a cumulative less than or a cumulative more than
frequency distributions, respectively.
Example 2.6: Convert the absolute frequency distribution in example 2.4 into:
i) a cumulative less than frequency distribution.
ii) a cumulative more than frequency distribution.
Solution:
i) We use the class boundaries to form cumulative frequencies. For instance, there is no observation which
is less than 15.5, 3 observations are less than 21.5, 9 observations are less than 27.5 and so on. Thus,
the following less than cumulative frequency distribution is obtained.
8
Class frequency Relative frequency
Democratic 13 0.325
Republican 18 0.45
Other 9 0.225
Total 40 1
Frequency polygon: is a graphic form of a frequency distribution. It can be constructed by plotting the class
frequencies against class marks and joining them by a set of line segments.
Note: we should add two classes with zero frequencies at the two ends of the frequency distribution to complete
the polygon.
Example 2.10: Construct a frequency polygon for the frequency distribution of the time spent by the
automobile workers that we have seen in example 2.4.
9
2.4.2 Graphs useful for presenting qualitative data
Bar charts are diagrammatic representation of data in which the data are represented by series of vertical or
horizontal bars, the height (or length) of each bar indicating the size of the figure represented.
Example 2.11: Draw a bar chart for the following coffee production data.
Table: Coffee productions from 1990 to 1995.
Production year 1990 1991 1992 1993 1994 1995
Amounts of coffee (in 1000 tons) 50 75 92 64 100 120
120
80
60
40
20
0
1990 1991 1992 1993 1994 1995
Production year
Pie-chart: it is a circle divided by radial lines into sections or sectors so that the area of each sector is
proportional to the size of the figure represented.
Pie-chart construction:
f
Calculate the percentage frequency of each component. It is i *100 .
n
f
Calculate the degree measures of each sector. It is given by i * 3600 .
n
Draw the circle using protractor and compass
Example 2.13: Draw a pie-chart to represent the following data on a certain family expenditure.
Table: Family expenditure.
Item Food Clothing House rent Fuel & light Miscellaneous Total
Expenditure(in birr) 50 30 20 15 35 150
Percentage 33.33 20 13.33 10 23.33
frequencies
Angles of the sector 1200 720 480 360 840 3600
10
Item
Food
Clothing
House rent
Fuel and light
Miscellaneous
Example 2.14: The following data are the blood types of 50 volunteers at a blood plasma donation clinic:
O A O AB A A O O B A O A AB B O O O A B A A O A A B O B A O AB A O O
A B A A A O B O O A O A B O AB A O
a) Organize this data using a categorical frequency distribution
b) Present the data using both a pie and a bar chart.
Solution
a) The classes of the frequency distribution are A, B, O, AB. Count the number of donors for each of the
blood types.
A 19 38.0
B 8 16.0
O 19 38.0
AB 4 8.0
Total 50 100.0
b) Pie chart
Find the percentage of donors for each blood type. In order to find the angles of the sector for each blood
type, multiply the corresponding percentage by 3600 and divide by 100.
11
Blood type
A
B
O
AB
15
10
0
A B O AB
Blood type
12
UNIT THREE: MEASURES OF CENTERAL TENDENCY
Objectives:
Having studied this unit, you should be able to:
understand the role of descriptive statistics in summarization, description and interpretation of data.
use several numerical methods belonging to measures of central tendency to describe the characteristics
of a data set.
Note that the individual values of the distribution must have a tendency to cluster around an average. In view of
this requirement an average is also referred to as a measure of central tendency.
An average (a measure of central tendency) is considered satisfactory if it possesses all or most of the following
properties. An average should be
based on all the observed values.
simple to understand and easy to interpret.
calculable by reasonable ease and rapidity.
easily manipulated algebraically.
little affected by fluctuations of sampling.
should not unduly be influenced by extreme values.
13
it should be defined rigidly which means that it should have a definite value.
Objectives of measuring central tendency:
To get one single value that describes the characteristics of the entire group.
To facilitate comparison between different data sets.
Rules of summation
and
Definition 3.2:
i) Let be the values of the variable X. The simple arithmetic mean denoted by is the
sum of these observations of X divided by the no values.
ii) If the numbers occur with frequencies , respectively. Then mean can be
Note that if the data refers to a population data the mean is denoted by the Greek letter µ (read as mu).
Arithmetic mean for raw data (ungrouped data)
Example 3.1: The following data is the weight (in Kg) of eight youths: 32,37,41,39,36,43,48 and 36. Calculate
the arithmetic mean of their weight.
14
Solution:
Example 3.2: The ages of a random sample of patients in a given hospital in Ethiopia is given below:
Age 10 12 14 16 18 20 22
Number of patients 3 6 10 14 11 5 4
Calculate the average age of these patients.
Solution:
Age (xi) Number of patients (fi)
10 3 30
12 6 72
14 10 140
16 14 224
18 11 198
20 5 100
22 4 88
Total 53 852
Example 3.3: The GPA or CGPA of a student is a good example of a weighted arithmetic mean. Suppose that
Solomon obtained the following grades in the first semester of the freshman program at Jima University in
2000.
Course Credit hour (wi) Grade
Math101 4 A=4
Bio101 3 C=2
Chem101 3 B=3
Phys101 4 B=3
Flen101 3 C=2
15
Find the GPA of Solomon.
Example 3.4: In a vacancy for a position of botanist in an organization, the criteria of selection were work
experience, entrance exam, and, interview result. The relative importance of these criteria was regarded to be
different. The weights of these criteria and the scores obtained by 3 candidates (out of 100 in each criterion) are
given in the following table. In addition, the selection of a candidate is based on average result on these criteria.
Criterion Weight Candidates
Tesfaye Gutema Kedir
Work experience 4 70 89 85
Entrance exam 3 78 83 89
Interview result 2 90 92 90
Who is the appropriate candidate for the position based on the criteria?
Solution: We use the weighted mean since the relative importances of these criteria are different.
Criterion Weight Candidates
Tesfaye Gutema Kedir
xi xiwi xi xiwi xi xiwi
Work experience 4 70 280 89 356 85 340
Entrance exam 3 78 234 83 249 89 267
Interview result 2 90 180 92 184 90 180
Total 9 238 694 264 789 264 787
The weighted mean and the simple arithmetic mean for the applicants are as follows:
Applicant Tesfaye Gutema Kedir
Weighted mean 694/9=77.11 789/9=87.67 787/9=87.44
Simple arithmetic mean 238/3=79.33 264/3=88 264/3=88
If we use the simple arithmetic mean of the scores, both Gutema and Kedir have got equal chances to be
recruited. However, the relative importance of the criteria is different. So we have to use the weighted mean for
discriminating among the candidates. The weighted mean of the scores obtained by Gutema is larger than the
others. So Gutema should be recruited for the job.
Properties of arithmetic mean
i. It can be computed for any set of numerical data, it always exists, and unique.
ii. It depends on all observations.
iii. The sum of deviations of the observations about the mean is zero i.e.
16
3.3.2 Geometric mean
Definition 3.4: The geometric mean of any n positive numbers is the nth root of the products of the numbers.
Symbolically if are given their geometric (G.M) mean is given by
17
3.3.3 The median
Definition 3.5: the median of a set of data is a value which divides the set in such a way that the number of
observations below it is the same as the number of observations above it.
iii. A shop keeper (sales person) recorded the number of video cassette recorders (VCRs) sold per month
over a two year period. Find the median number VCRs sold.
Number of sets sold Frequency ( months) Cumulative frequency
1 3 3
2 8 11
3 5 16
4 4 20
5 2 22
6 1 23
7 1 24
Properties of median
It is an average of position.
It is affected by the number of observations than by extreme values.
The sum of the deviations about the median, signs ignored, is less than the sum of deviations taken from
any other value or specific average.
3.3.4 The mode
Definition 3.6: The mode (modal value) of an observed set of data is the value that occurs the largest number of
18
times.
Solution:
Note: In case of grouped data if any class interval is open, arithmetic mean cannot be calculated.
The median for grouped data can be approximated by the following formula.
19
where Lm= lower class boundary for the median class.
n= total number of observations in the distribution.
cf= less than cumulative frequency for the class preceding the median class.
w= class width for median class.
fm=frequency for median class.
Note that the median class is the class containing the (n/2) th observation.
Example 3.13: Find the median for the following frequency distribution.
Class boundaries Frequency (f) Cumulative frequency
5.5-10.5 1 1
10.5-15.5 2 3
15.5-20.5 3 6
20.5-25.5 5 11
25.5-30.5 4 15
30.5-35.5 3 18
35.5-40.5 2 20
th th
Solution: The class containing the (n/2) observation or the 10 observation is the median class. This class has
class boundaries 20.5 & 25.5(4th class).
The mode for grouped data can be estimated by the following formula.
The modal value is denoted by . For grouped data we can compute the mode as follows:
20
Therefore, the modal time spent is 29.5 minutes.
Note: The mode can be calculated for distributions with open ended classes.
Deciles are nine points which divide an array into 10 parts in such a way that each part contains equal number
of elements. The 1st, 2nd,…, and the 9th points are known as the 1st, 2nd,…, and the 9th deciles and are usually
denoted by D1,D2,…,D9, respectively.
Percentiles are 99 points which divide an array into 100 parts in such a way that each part consists of equal
number of elements. The 1st, 2nd… and the 99th points are known as the 1st, 2nd… and the 99th percentiles and are
usually denoted by P1, P2… P99, respectively.
Note: The array should be in ascending order in order to get the quantiles.
i. Quantile points for raw data
First form an array in an ascending order and then apply the following procedure.
Example 3.15: The following data relate to sizes of shoes sold at a stock during a week. Find the quartiles, the
seventh decile and the 90th percentile.
Size of shoes 5 5.5 6 6.5 7 7.5 8 8.5 9 9.5
Number of pairs 2 5 15 30 60 40 23 11 4 1
Solution: The total number of observations is 191.
21
Note: Relationships between fractile points
Q1=P25
Q2=P50=D5=
Q3=P75
D1=P10; D2=P20 …D9=P90.
22
UNIT FOUR: MEASURES OF VARIATION
Objectives:
Having studied this unit, you should be able to
understand the importance of measuring the variability (dispersion) in a data set.
measure the scatter or dispersion in a data set.
understand ‘moments’ as a convenient and unifying method for summarizing several descriptive
statistical measures.
measure the extent to which the distribution of values in a data set deviate from symmetry.
4.1 Introduction and objectives of measuring variation
We have seen that averages are representatives of a frequency distribution. But they fail to give a complete
picture of the distribution. They do not tell anything about the spread or dispersion of observations within the
distribution. Suppose that we have the distribution of yield (kg per plot) of two rice varieties from 5 plots each.
Variety 1: 45 42 42 41 40
Variety 2: 54 48 42 33 30
The mean yield of both varieties is 42 kg. The mean yield of variety 1 is close to the values in this variety. On
the other hand, the mean yield of variety 2 is not close to the values in variety 2. The mean doesn’t tell us how
the observations are close to each other. This example suggests that a measure of central tendency alone is not
sufficient to describe a frequency distribution. Therefore, we should have a measure of spreads of observations.
There are different measures of dispersion. In this chapter we shall discus the most commonly used measure of
dispersion or variation like Range, Quartile Deviation, Standard Deviation, coefficient of variation. And
measure of shape such as skewness and kurtosis.
Objectives of measuring variation
To describe dispersion (variability) in a data.
To compare the spread in two or more distributions.
To determine the reliability of an average.
Note: The desirable properties of good measures of variation are almost identical with that of a good measure of
central tendency.
4.2 Absolute and relative measures
Measures of variation may be either absolute or relative. Absolute measures of variation are expressed in the
same unit of measurement in which the original data are given. These values may be used to compare the
variation in two distributions provided that the variables are in the same units and of the same average size.
In case the two sets of data are expressed in different units, however, such as quintals of sugar versus tones of
sugarcane or if the average sizes are very different such as manager’s salary versus worker’s salary, the absolute
measures of dispersion are not comparable. In such cases measures of relative dispersion should be used. A
measure of relative dispersion is the ratio of a measure of absolute dispersion to an appropriate measure of
central tendency. It is a unitless measure.
4.3 Types of measures of variation
The range and relative range
Definition 4.1: Range is defined as the difference between the maximum and minimum observations in a set of
data.
Range is the crudest absolute measures of variation. It is widely used in the construction of quality control
charts and description of daily temperature.
23
Definition 4.2: Relative range (RR) is defined as
Definition 4.3: The variance is the average of the squares of the distance each value is from the mean. The
symbol for the population variance is σ2 (σ is the Greek lower case letter sigma). Let x1,x2,…,xN be the
measurements on N population units then, the population variance is given by the formula:
Example 4.1: The height of members of a certain committee was measured in inches and the data is presented
below.
Height(x): 69 66 67 69 64 63 65 68 72
2 -1 0 2 -3 -4 -2 1 5
4 1 0 4 9 16 4 1 25
.
Definition 4.6: The sample standard deviation, denoted by S, is the square root of the sample variance
.
Example 4.2: For a newly created position, a manager interviewed the following numbers of applicants each
day over a five-day period: 16, 19, 15, 15, and 14. Find the variance and standard deviation.
Solution:
24
Note that the procedure for finding the variance and standard deviation for grouped data is similar to that for
finding the mean for grouped data, and it uses the mid-points of each class.
Properties of variance
The unit of measurement of the variance is the square of the unit of measurement of the observed values.
It is one of its limitations.
The variance gives more weight to extreme values as compared to those which are near to mean value,
because the difference is squared in variance.
It is based on all observations in the data set.
Properties of standard deviation
Standard deviation is considered to be the best measure of dispersion and is used widely.
There is, however, one difficulty with it. If the unit of measurement of variables of two series is not the
same, then their variability cannot be compared by comparing the values of standard deviation.
Uses of the variance and standard deviation
The variance and standard deviations can be used to determine the spread of data, consistency of a
variable and the proportion of data values that fall within a specified interval in a distribution.
If the variance or standard deviation is large, the data is more dispersed. This information is useful in
comparing two or more data sets to determine which is more (most) variable.
Finally, the variance and standard deviation are used quite often in inferential statistics.
Coefficient of variation (CV)
The standard deviation is an absolute measure of dispersion. The corresponding relative measure is known as
the coefficient of variation (CV).
Coefficient of variation is used in such problems where we want to compare the variability of two or more
different series. Coefficient of variation is the ratio of the standard deviation to the arithmetic mean, usually
expressed in percent:
S
CV 100%
x
where S is the standard deviation of the observations.
A distribution having less coefficient of variation is said to be less variable or more consistent or more uniform
or more homogeneous.
Example 4.3: Last semester, the students of Biology and Chemistry Departments took Stat 273 course. At the
end of the semester, the following information was recorded.
Department Biology Chemistry
Mean score 79 64
Standard deviation 23 11
Compare the relative dispersions of the two departments’ scores using the appropriate way.
Solution:
Biology Department Chemistry Department
23 11
CV 100 29.11% CV 100 17.19%
79 64
Since the CV of Biology Department students is greater than that of Chemistry Department students, we can say
that there is more dispersion in the distribution of Biology students’ scores compared with that of Chemistry
students.
Example 4.4: The mean weight of 20 children was found to be 30 kg with variance of 16kg2 and their mean
height was 150 cm with variance of 25cm2. Compare the variability of weight and height of these children.
Example 4.5: Two sections were given an exam in a course. The average score was 72 with standard deviation
of 6 for section 1 and 85 with standard deviation of 5 for section 2. Student A from section 1 scored 84 and
student B from section 2 scored 90. Who performed better relative to his/her group?
Solution: Section 1: x = 72, S = 6 and score of student A from Section 1; x A = 84
Section 2: x = 85, S = 5 and score of student B from Section 2; x B = 90
x x1 84 72
Z-score of student A: Z A 2.00
S1 6
x x 2 90 85
Z-score of student B: Z B 1.00
S2 5
From these two standard scores, we can conclude that student A has performed better relative to his/her section
students because his/her score is two standard deviations above the mean score of selection 1 while the score of
student B is only one standard deviation above the mean score of section 2 students.
Example 4.6: A student scored 65 on a calculus test that had a mean of 50 and a standard deviation of 10; she
scored 30 on a history test with a mean of 25 and a standard deviation of 5. Compare her relative positions on
each test.
Solution: First, find the z-scores.
For calculus the z-score is
Since the z-score for calculus is larger, her relative position in the calculus class is higher than her relative
position in the history class.
4.4 Moments, skewness and kurtosis
Moments
Definition 4.7: The average of deviations from an arbitrary origin raised to an integral power of the observations
of a distribution is defined as a moment. Let x1,x2,…,xn be observations, we define the r-th moment about A as:
The most known moments are moments about the mean also known as the central moments and the moments
about zero (also known as moments about the origin.)
The rth moment about the mean, µr, is given by:
26
.
Special ceases: µ0=1, µ1=0, µ2=s2.
The rth moment about the origin, , is given by:
.
Special cases: , ,
Skweness: it refers to lack of symmetry in a distribution.
Note: for a symmetrical and unimodal distribution:
i) Mean =median =mode
ii) The lower and upper quartiles are equidistant from the median, so also are corresponding pairs of deciles
and percentiles.
iii) Sum of positive deviations from the median is equal to the sum of negative deviations (signs ignored).
iv) The two tails of the frequency curve are equal in length from the central value.
If a distribution is not symmetrical we call it skewed distribution.
Measures of skewness
i) Pearsonian coefficient of skewness (Pcsk) defined as:
Interpretation:
Note: in a negatively skewed distribution larger values are more frequent than smaller values. In a positively
skewed distribution smaller values are more frequent than larger values.
Example 4.7: If the mean, mode and s.d of a frequency distribution are 70.2, 73.6, and 6.4, respectively. What
can one state about its skeweness?
.
This figure suggests that there is some negative skeweness.
27
When the values of a distribution are closely bunched around the mode in such a way that the peak of the
distribution becomes relatively high, the distribution is said to be leptokurtic. If it is flat topped we call it
platykurtic. A distribution which is neither highly peaked nor flat topped is known as a meso-kurtic distribution
(normal).
Measures of kurtosis
where
Interpretation:
28
Examples of random experiments includes throwing a fair coin and observing the outcome, throwing a fair die
and observing the number on the top face, taking a student at random from science class and noting the sex of
the student.
All of these examples satisfy the above characteristics of a random experiment.
Definition 5.2:
Sample point (outcome): The individual result of a random experiment.
Sample space: The set containing all possible sample points (out comes) of the random experiment. The
sample space is often called the universe and denoted by S.
Event: The collection of outcomes or simply a subset of the sample space. We denote events with capital
letters, A, B, C, etc.
Example 5.3: Suppose one wants to purchase a certain commodity and that this commodity is on sale in 5
government owned shops, 6 public shops and 10 private shops. How many alternatives are there for the person
to purchase this commodity?
Solution: Total number of ways =5+6+10=21 ways
Example 5.4: If we can go from Addis Ababa to Rome in 2 ways and from Rome to Washington D.C. in 3
ways then the number of ways in which we can go from Addis Ababa to Rome to Washington D.C. is 2x3
ways or 6 ways. We may illustrate the situation by using a tree diagram below:
W
R W
W
A
W
R
W
W
Example 5.5: If a test consists of 10 multiple choice questions, with each permitting 4 possible answers, how
many ways are there in which a student gives his/her answers?
Solution: There are 10 steps required to complete the test.
First step: To give answer to question number one. He/she has 4 alternatives.
Second step: To give answer to question number two, he/she has 4 alternatives……
Last step: To give answer to last question, he/she has 4 alternatives.
Therefore, he/she has 4x4x4x…x4=410 ways or1, 048, 576 ways of completing the exam. Note that there is only
one way in which he /she can give correct answers to all questions and that there are 310 ways in which all the
answers will be incorrect.
Example 5.6: A manufactured item must pass through three control stations. At each station the item is
inspected for a particular characteristic and marked accordingly. At the first station, three ratings are possible
while at the last two stations four ratings are possible. Hence there are 48 ways in which the item may be
marked.
Example 5.7: Suppose that car plate has three letters followed by three digits. How many possible car plates are
there, if each plate begins with a H or an F?
30
2x 26x 26x 10x 10x 10 or 1, 352, 000 different plates.
Definition 5.4: If n is a positive integer, we define n!= n(n-1)(n-2)…1 and call it n-factorial and 0!=1.
Permutations
Suppose that we have n different objects. In how many ways, say nPn, may these objects be arranged
(permuted)? For example, if we have objects a, b and c we can consider the following arrangements: abc, acb,
bac, bca, cab, and cba. Thus the answer is 6. The following theorem gives general result on the number of such
arrangements.
There are many problems in which we are interested in determining the number of ways in which r objects can
be selected from n distinct objects without regard to the order in which they are selected. Such selections are
called combinations or r-sets. It may help to think of combinations as committees. The key here is without
regard for order.
To obtain the general result we recall the formula derived above: the number of ways of choosing r objects out
of n and permuting the chosen r equals n!/(n-r)!. Let C be the number of ways of choosing r out of n,
disregarding order. C is the number required. Note that once the r items have been chosen, there are r! ways of
permuting them. Hence applying the multiplication principle again, together with the above result, we obtain
n!
C.r! = n!/(n-r)!. Therefore, C . This number arises in many contexts in mathematics and hence a
r!(n r )!
special symbol is used for it. We shall write
n n!
n C r .
r r!(n r )!
Example 5.12: How many different committees of 3 can be formed from Hawa, Segenet, Nigisty and Lensa?
Solution: The question can restated in terms of subsets from a set of 4 objects, how many subsets of 3 elements
are there? In terms of combinations the question becomes, what is the number of combinations of 4 distinct
objects taken 3 at a time? The list of committees:{H,S,N}, {H,S,L}, {H,N,L}, {S,N,L}.Therefore, we have 4C3
or 4 possible number of committees.
32
Example 5.13:
(i) A committee of 3 is to be formed from a group of 20 people. How many different committees are possible?
(ii) From a group of 5 men and 7 women, how many different committees consisting of 2 men and 3 women can
be formed?
20 20!
Solution: (i) There are 1140possible committees.
3 3!17!
5 7 5! 7!
(i) 350 possible committees.
2 3 2!3! 3!4!
Remarks:
n n
i)
r n r
It is rather surprising that with only these three axioms, we can construct the "entire" theory of probability! The
next theorems and definitions help in assigning probabilities of events.
Theorem 5.6 :If A is an event in a discrete sample space S, then P(S) equals the sum of the probabilities of
the individual outcomes comprising A.
Theorem 5.7: Suppose that we have a random experiment with sample space S and probability function P
and A and B are events. Then we have the following results:
i) P( ) = 0
ii) P(Ac) = 1 − P(A)
iii) P(B n Ac) = P(B) − P(A n B)
iv) If A subset of B then P(A) ≤ P(B).
P ( B c ) 1 P ( B ) 1 0 .5 0 .5
P ( S c ) 1 P ( S ) 1 1 0 P ( )
Example 5.16: From a group of 5 men and 7 women, it is required to form a committee of 5 persons. If the
selection is made randomly, then
i) what is the probability that 2 men and 3 women will be in the committee?
ii) what is the probability that all members of the committee will be men?
iii) what is the probability that at least three members will be women?
12 12!
Solution: The total number of possible committees is 792 , i.e. the number of possible out comes
5 5!7!
in the sample space is 792.
i) Let A be the event that the committee will consist of two 2 men and 3 women. We need to know the
number of possible outcomes favoring this event. The number of ways we can select 2 men from 5
5 5!
men is 10 and the number of ways of selecting 3 women out of 7 women is
2 2!3!
7 7!
35 . Using the multiplication principle, the number of elements favoring event A is
3 3!4!
10x35 or 350.
Hence, using the classical definition of probability,
5 7
P( A)
2 3 350
0.44
12 792
5
ii) Let B be the event that all members of the committee will be men. Hence
34
5 7
P( A)
5 0 1
12 792
5
iii) Let C be the event that at least three of the committee members will be women.
Basically, three different compositions of committee members can be formed in terms of sex: 3
women and 2 men, 4 women and 1 man, and all are women. Hence the number of possible outcomes
favoring event C using the principle of combination together with the addition principle
5 7 5 7 5 7
is 350 175 21 546 .
2 3 1 4 0 5
5 7 5 7 5 7
Therefore, P(C )
2 3 1 4 0 5 546
0.69
12 792
5
The above definition of probability is based on empirical data accumulated through time or based on
observations made from repeated experiments for a large number of times.
5.5 Some probability rules
Theorem 5.8: If A and B , then P(A u B) = P(A) + P(B) − P(A n B).
35
Solution: Let A represents that the family owns a car and B represents that the family owns a house. Given
information: P(A)=0.6,P(B)=0.3, and P(AnB)=0.2.
a) Required: P(Ac) = ?
P(Ac)=1-P(A) = 1-0.6 = 0.4
b) Required: P(AUB) = ?
P(AUB) = P(A)+P(B)-P(AnB) = 0.6+0.3-0.2 = 0.7
c) Required: P((AnBc)U(AcnB)) = ?
P((AnBc)U(AcnB)) = P(AnBc)+P(AcnB) = [P(A)-P(AnB)]+[P(B)-P(AnB)]
= [0.6-0.2]+[0.3-0.2]=0.5
d) Required: P(AcnB) =?
P(AcnB) = P(B)-P(AnB) = 0.3-0.2 = 0.1
e) Required: P(AcnBc) = ?
P(AcnBc) = P((AUB)c) = 1-P(AUB) = 1-0.7 = 0.3
We can represent various events by an informative diagram called vein diagram. If properly and correctly
drawn, a vein diagram helps to calculate probabilities of events easily. The figure below shows various events
represented by shaded regions. Note that the rectangle in each figure represents the sample space.
In more precise terms, given an experiment, a corresponding sample space, and a probability law, supposes that
we know that the outcome is within some given event B. We wish to quantify the likelihood that the outcome
also belongs to some other given event A. We thus seek to construct a new probability law, which takes into
account this knowledge and which, for any event A, gives us the conditional probability of A given B, denoted
by P(A|B).
Definition 5.8: If P(B) > 0, the conditional probability of A given B, denoted by P(A|B), is
P ( AnB)
P( A / B) .
P( B)
Example 5.19: Suppose cards numbered one through ten are placed in a hat, mixed up, and then one of the
cards is drawn at random. If we are told that the number on the drawn card is at least five, then what is the
conditional probability that it is ten?
Solution: Let A denote the event that the number on the drawn card is ten, and B be the event that it is at least
five. The desired probability is P(A|B).
36
P( AnB) P({10}n{5,6,7,8,9,10}) P({10}) 1 / 10 1
P( A / B)
P( B) P({5,6,7,8,9,10}) P({5,6,7,8,9,10}) 6 / 10 6
Example 5.20: A family has two children. What is the conditional probability that both are boys given that at
least one of them is a boy? Assume that the sample space S is given by S = {(b, b), (b, g), (g, b), (g, g)}, and all
outcomes are equally likely. (b, g) means, for instance, that the older child is a boy and the younger child is a
girl.
Solution: Letting A denote the event that both children are boys, and B the event that at least one of them is a
boy, then the desired probability is given by
P( AnB) 1 / 4 1
P( A / B)
P( B) 3/ 4 3
Law of Multiplication
The defining equation for conditional probability may also be written as:
P(AnB) = P(B) P(A|B)
This formula is useful when the information given to us in a problem is P(B) and P(A|B) and we are asked to
find P(AnB). An example illustrates the use of this formula. Suppose that 5 good fuses and two defective ones
have been mixed up. To find the defective fuses, we test them one-by-one, at random and without replacement.
What is the probability that we are lucky and find both of the defective fuses in the first two tests?
Example 5.21: Suppose an urn contains seven black balls and five white balls. We draw two balls from the urn
without replacement. Assuming that each ball in the urn is equally likely to be drawn, what is the probability
that both drawn balls are black?
Solution: Let A and B denote, respectively, the events that the first and second balls drawn are black. Now,
given that the first ball selected is black, there are six remaining black balls and five white balls, and so P(B|A)
= 6/11. As P(A) is clearly 7/12 , our desired probability is
7 6 7
P( AnB) P( A) P( B / A) .
12 11 22
Independence
We have introduced the conditional probability P(A|B) to capture the partial information that event B provides
about event A. An interesting and important special case arises when the occurrence of B provides no
information and does not alter the probability that A has occurred, i.e., P(A|B) = P(A). When the above equality
holds, we say that A is independent of B. Note that by the definition P(A|B) = P(A ∩ B)/P(B), this is equivalent
to P(A ∩ B) = P(A)P(B).
Example 5.22: A basket contains 2 black and 2 white balls. Consider selecting two balls at random in two
ways:
a) selecting a second ball without replacing first selected ball in the basket
b) selecting a second ball after replacing first selected ball in the basket
Let A and B represents black ball will be selected in the first and second selection, respectively. In which of the
two ways are A and B independent?
Solution:
First way (a):
S = {B1B2, B2B1, B1W1, B1W2 , B2W1, B2W2, W1W2 , W1B1, W2B1 , W1B2, W2B2,W2W1}
A = {B1B2, B2B1, B1W1, B1W2 , B2W1, B2W2 }
B = {B1B2, B2B1, W1B1, W2B1 , W1B2, W2B2 }
A n B = {B1B2, B2B1}
37
6 1 6 1 2 1
P( A) , P( B) and P( AnB)
12 2 12 2 12 6
1 1
P( AnB) P( A) P( B) .
6 4
Thus A and B are not independent.
Second way (b):
Using this option to select a ball does not affect the composition of the basket.
2 1 2 1 2 2 1
P( A) , P( B) and P( AnB) .
Hence 4 2 4 2 4 4 2
38
a certain locality, the number of bacteria in a cubic mm of agar, etc. If random variable assumes any numerical
value in an interval or collection of intervals, then it is called a continuous random variable. Examples include
body weight of new born baby, life time of a human being, height of a person, etc.
The most important way to characterize a random variable is through the probabilities of the values that it can
take. For a discrete random variable X, these are captured by the probability mass function (p.m.f. for short) of
X, denoted PX(x). For a continuous random variable X it is done by the probability density function (p.d.f.),
denoted fX(x).
Example 6.1: Consider an experiment of tossing two fair coins. Letting X denote the number of heads
appearing on the top face, then X is a random variable taking on one of the values 0, 1, 2 . The random variable
X assigns a 0 value for the outcome (T,T), 1 for outcomes (T ,H) and (H, T ), and 2 for the outcome (H,H).
Thus, we can calculate the probability that X can take specific value/s as follows:
P(X = 0) = P({(T , T )}) = ¼
P(X = 1) = P({(T ,H),(H, T )}) = 2/4,
P(X = 2) = P({(H,H)}) = ¼
The table below shows the probability mass function X.
X 0 1 2
PX(x) ¼ 2/4 ¼
We can justify that PX(x) is probability mass function.
PX(x)≥0 for x=0,1,2 and
P(X = 0) + P(X = 1)+ P(X = 2) = ¼ + 2/4 + ¼=1
Suppose we are interested to calculate the probability that X≥1. The values of X which are greater than or equal
to 1 are 1 and 2. Thus, the probability that X is greater than or equal to 1, denoted P(X≥1), is found as P(X≥1)
= P(X = 1) + P(X = 2)=3/4.
We can use the probability density function to calculate probabilities of events expressed in terms of the random
variable X. For instance, if we are interested in the probability that X lies between two points, say a and b, we
can find it using integration of fX(x) on the interval [a,b],i.e.
b
P(a X b) f X ( x)dx
a
39
Figure: P(a≤ X ≤ b) is the shaded region
Remarks:
i) The area bounded under the graph of a probability density function and below by the horizontal axis is 1.
ii) The probability that a continuous random variable X will assume a specific value is zero, i.e.
c
P( X c) f X ( x)dx 0 where c is a constant.
c
iii) The probability that a continuous random variable X will assume a value in a closed intervals is the same
as the probability that it will assume in open interval or half open intervals, i.e. , P(a≤X≤b) = P(a<X<b) =
P(a≤X<b) = P(a<X≤b), P(X≤c) = P(X<c) , P(X≥c) = P(X>c) where a, b, and c are constants.
It is useful to view the mean of X as a “representative” value of X, which lies somewhere in the middle of its
range. We can make this statement more precise, by viewing the mean as the center of gravity of the
distribution.
Variance
Definition 6.4: The variance of a random variable X denoted V(X) or σ2 is defined as V(X)=E[(X- μ)2] =
E(X2) – μ2.
i) if X is discrete, V ( X ) [ x 2 PX ( x)] 2
ii) if X is continuous, V ( X ) [ x 2 f X ( x)dx] 2
The variance provides a measure of dispersion of X around its mean. Another measure of dispersion is the
standard deviation of X, which is defined as the square root of the variance and is denoted by σ.
Example 6.2: Calculate the mean and variance of the random variable X in example 7.1.
1 1 1
E ( X ) xPX ( x) 0 1 2 1
4 2 4
1 1 1
E ( X 2 ) x 2 PX ( x ) 02 12 2 2 1.5
4 2 4
V ( X ) E ( X ) 1.5 1 0.5
2 2 2
40
6.3 Common discrete probability distributions – binomial and Poisson
The Binomial distribution
Many real problems (experiments) have two possible outcomes, for instance, a person may be HIV-Positive or
HIV-Negative, a seed may germinate or not, the sex of a new born bay may be a girl or a boy, etc. Technically,
the two outcomes are called Success and Failure. Experiments or trials whose outcomes can be classified as
either a “success” or as a “failure” are called Bernoulli trails.
Suppose that n independent trials, each of which results in a “success” with probability p and in a “failure” with
probability 1 − p, are to be performed. If X represents the number of successes that occur in the n trials, then X
is said to have binomial distribution with parameters n and p. The probability mass function of a binomial
distribution with parameters n and p is given by
n
PX ( x) p x (1 p ) n x , x 0, 1, 2, ..., n
x
The mean and variance of the binomial distribution are np and np(1-p), respectively. Note that the binomial
distributions are used to model situations where there are just two possible outcomes, success and failure. The
following conditions also have to be satisfied.
i) There must be a fixed number of trials called n
ii) The probability of success (called p) must be the same for each trial.
iii) The trials must be independent
Example 6.3: A fair coin is flipped 4 times. Let X be the number of heads appearing out of the four trials.
Calculate the following probabilities:
i) 2 heads will appear
ii) No head will appear
iii) At least two heads will appear
iv) Less than two heads will appear
v) At most heads 2 will appear
Solution: We can consider that the outcomes of each trial are independent to each other. In addition the
probability that a head will appear in each trial is the same. Thus, X has a binomial distribution with number of
trials 4 and probability of success (the occurrence of head in a trial) is ½. The probability mass function of X is
given by
n n
PX ( x) 0.5 x (1 0.5) n x 0.5 n , x 0, 1, 2, 3,4 , Note that n = 4 and p = 1/2
x x
4
i) P( X 2) 0.5 2 (1 0.5) 4 2 0.3750
2
4
ii) P( X 0) 0.5 0 (1 0.5) 40 0.0625
0
iii) P( X 2) P( X 2) P( X 3) P( X 4) 0.3750 0.2500 0.0625 0.6875
iv) P( X 2) P( X 0) P( X 1) 0.0625 0.2500 0.3125
v) P( X 2) P( X 0) P( X 1) P( X 2) 0.0625 0.2500 0.3750 0.6875
Example 6.4: Suppose that a particular trait of a person (such as eye color or left handedness) is classified on
the basis of one pair of genes and suppose that d represents a dominant gene and r a recessive gene. Thus a
person with dd genes is pure dominance, one with rr is pure recessive, and one with rd is hybrid. The pure
dominance and the hybrid are alike in appearance. Children receive one gene from each parent. If, with respect
to a particular trait, two hybrid parents have a total of four children, what is the probability that exactly three of
the four children have the outward appearance of the dominant gene?
Solution: If we assume that each child is equally likely to inherit either of two genes from each parent, the
probabilities that the child of two hybrid parents will have dd, rr, or rd pairs of genes are, respectively, ¼, ¼,
½. Hence, because an offspring will have the outward appearance of the dominant gene if its gene pair is either
41
dd or rd, it follows that the number of such children ,say X, is binomially distributed with parameters n equals
4
4 and p equals ¾. Thus the desired probability is P( X 3) 0.753 (1 0.75) 43 0.421875.
3
Example 6.5: Suppose it is known that the probability of recovery for a certain disease is 0.4. If random sample
of 10 people who are stricken with the disease are selected, what is the probability that:
(a) exactly 5 of them will recover?
(b) at most 9 of them will recover?
Solution: Let X be the number of persons will recover from the disease. We can assume that the selection
process will not affect the probability of success (0.4) for each trial by assuming a large diseased population
size. Hence, X will have a binomial distribution with number of trials equal to 10 and probability of success
10
equal 0.4. P ( X k ) 0.4 k 0.610 k , k 0,1,2,... 10
k
10
(a) P( X 5) 0.4 5 0.6105 0.200658
5
10
(b) P( X 9) 1 P( X 10) 1 0.410 0.61010 1 0.000105 0.9999
10
The Poisson Random Variable
A random variable X, taking on one of the values 0, 1, 2, . . . , is said to have a Poisson distribution if its
probability mass function is given by
e x
PX ( x) , x 0, 1, 2, 3, ... and 0 .
x!
λ is the parameter of this distribution. The mean and variance of the poisson distribution are equal and their
values are equal to λ. Note that poisson distributions is used to model situations where the random variable X is
the number of occurrences of a particular event over a given period of time (or space). Together with this , the
following conditions must also be fulfilled: events are independent of each other, events occur singly, and
events occur at a constant rate (in other words for a given time interval the mean number of occurrences is
proportional to the length of the interval).
The poisson distribution is used as a distribution of rare events such as telephone calls made to a switch board in
a given minute, number of misprints per page in a book, road accidents on a particular motor way in one day,
etc. The process that give rise to such events are called poisson processes.
Example 6.6: Suppose that the number of typographical errors on a single page of this lecture note has a
Poisson distribution with parameter λ = 1. if we randomly select a page in this lecture note, calculate the
probability that
a) no error will occur.
b) exactly three errors will occur.
c) less than 2 errors will occur.
d) there is at least one error.
Solution: Let X= Number of errors per page
e k
P( X k ) , 1, k 0,1,2,...
k!
e 110 1
a) Required P(X≥1)=? P( X 0) 0.367879
0! e
1 3
e 1
b) P( X 3) 0.061313
3!
c) P( X 2) P( X 0) P( X 1) 0.73576
D) P( X 1) 1 P( X 0) 1 0.367879 0.632121
42
Example 6.7: If the number of accidents occurring on a highway each day is a Poisson random variable with
parameter λ = 3, what is the probability that no accidents will occur on a randomly selected day in the future?
Solution: Let X= number of accidents per day
e 3 3 k
P( X k ) , k 0,1,2,...
k!
e 3 30
Required P(X= 0) = ? P( X 0) e 3 0.05
0!
Note: The Poisson random variable has a wide range of applications in a diverse number of areas. An important
property of the Poisson random variable is that it may be used to approximate a binomial random variable when
the binomial parameter n is large and p is small. The probability that X will be k can be approximated by
e k
substituting λ by np in the poisson distribution, i.e. P( X k ) , np .
k!
6.4 Common continuous probability distributions
Normal distribution
The normal distribution plays an important role in statistical inference because many real-life distributions are
approximately normal; many other distributions can be almost normalized by appropriate data transformations
(e.g., taking the log) and as a sample size increases, the means of samples drawn from a population of any
distribution will approach the normal distribution.
A continuous random variable X is said to follow normal distribution , if and only if , its probability density
1 x 2
1 ( )
function (p.d.f.) is f X ( x) e 2
where x (-∞,∞ ), μ (-∞,∞ ) and σ (0,∞ ). There are
2
infinitely many normal distributions since different values of μ and σ define different normal distributions. For
1
1 2 z2
instance, when μ= 0 and σ =1 , the above density will have the following form f Z ( z ) e . This
2
particular distribution is called the standard normal distribution and sometimes known as Z-distribution.. The
random variable corresponding to this distribution is usually denoted by Z. If X has a normal distribution with
mean μ and variance σ2, we denote it as X ~ N , 2 .
Properties of normal distribution
i) The normal distribution curve is a bell shaped, symmetrical about μ and mesokurtic. The p.d.f. attains its
maximum value at x= μ.
ii) Since for x= μ divides the area under the normal curve into two equal parts, μ is the mean, the median and
the mode of the distribution.
iii) The mean and variance of the normal distribution are μ, and σ2, respectively.
iv) The total area under the curve and bounded from below by the horizontal axis is 1, i.e. f
X ( x)dx 1
43
Since a normal distribution is a continuous probability distribution, the probability that X lies between a and b is
the area bounded under the curve, from left to right by the vertical lines x = a and x = b and below by the
horizontal axis.
Example 6.8: Let Z be the standard normal random variable. Calculate the following probabilities using the
standard normal distribution table: a) P(0<Z<1.2) b) P(0<Z<1.43) c) P(Z≤0) d) P(-1.2<Z<0) e) P(Z≤-1.43)
f) P(-1.43≤Z<1.2) g) P(Z≥1.52) h)P(Z≥-1.52)
Solution:
a) The probability that Z lies between 0 and 1.2 can be directly found from the standard normal table as
follows: look for the value 1.2 from z column ( first column) and then move horizontally until you find
the value of 0.00 in the first row. The point of intersection made by the horizontal and vertical movements
will give the desired area (probability). Hence P(0<Z<1.2)= 0.3849. Refer the table below as a guide to
find this probability.
44
Figure: P(0<Z<1.2) is the shaded area
b) In a similar way P(0<Z<1.43)= 0.4236.
c) We know that the normal distribution is symmetric about its mean. Hence the area to the left of 0 and the
to the right of zero are 0.5 each. Therefore P(Z≤0)=P(Z≥0)=0.5
Figure: The area to the left and the right of 0 for z-distribution
d) P(-1.2<Z<0)=P(0<Z<1.2)= 0.3849 due to symmetry
e) P(Z<-1.43)= 1- P(Z ≥ -1.43) Using the probability of the complement event.
= 1-[P(-1.43<Z<0)+P(Z≥0)] Since a region can be broken down
=1-[P(0<Z<1.43)+P(Z ≥0)] into non overlapping regions.
=1-[0.4236 + 0.5]
=1-0.9236=0.0764
45
Figure: P(-1.43≤Z<1.2) is the shaded region
g) P(Z≥1.52) = 0.5 – P(0≤ Z<1.52)=0.5 – 0.4357=0.0643
46
P(Z>z*) = 0.8554 = P(z*≤ Z <0) + P( Z ≥ 0) = P(0 ≤ Z ≤ - z*) + 0.5
Implies P(0 ≤ Z ≤ - z*) = 0.8554 – 0.5=0.3554. Hence the value –z* form the table satisfying the above
condition is 1.06. Therefore z* = -1.06.
Example 6.9: If the total cholesterol values for a certain target population are approximately normally
distributed with a mean of 200 (mg/100 ml) and a standard deviation of 20 (mg/100 ml), calculate the
probability that a person picked at random from this population will have a cholesterol value
a) greater than 240 (mg/100 ml)
b) between 180 and 220(mg/100 ml)
c) less 200 (mg/100 ml)
Solution: Let X be the cholesterol values in mg/100 ml, then X ~ N 200, 400
X b
P ( X 240) P ( )
a)
240 200
P( Z ) P ( Z 2) 0.5 P (0 Z 2) 0.5 0.4772 0.0228
20
X X b
P (180 X 220) P ( )
180 200 220 200
b) P ( Z ) P (1 Z 1)
20 20
2 P (0 Z 1) 2 0.3413 0.6826
200 200
c) P( X 200) P( Z ) P( Z 0) 0.5
20
Example 6.10: Assume that the test scores for a large class are normally distributed with a mean of 74 and a
standard deviation of 10.
(a) Suppose that you receive a score of 88. What percent of the class received scores higher than yours?
(b) Suppose that the teacher wants to limit the number of A grades in the class to no more than 20%. What
would be the lowest score for an A?
Solution: Let X be the score of a randomly picked student, then X ~ N 74, 100
X 74 88 74
P( X 88) P( ) P( Z 1.4)
a) 10 10
0.5 P(0 Z 1.4) 0.5 0.4192 0.0808
Hence 8.08 percent of the students score more than you did?
b) Let XA be the lowest mark to get letter grade A. We are given that
X 74 x A 74
P ( X x A ) 0 .2 P ( ) P( Z z A )
10 10
x 74
P (0 Z z A ) 0.5 0.2 0.3 z A 0.85 z A 0.85 A
10
Hence, the lowest mark to get letter grade A is 82.5.
47
The chi-square and t distributions
The chi-square and t distributions are important continuous distributions which are useful in statistical
inference. In this section we will see a brief introduction of these distributions. In later chapters, we are going to
see in detail on how to use these distributions in estimation and hypotheses testing.
Chi-square distribution
A random variable X is said to have a chi-square distribution with n degrees of freedom (denoted by n2 ) if its
probability density function is given by
n x
1 1
f X ( x) n x 2 e 2 , x 0.
2 2 ( n )
2
The chi-square distribution has one parameter called the degrees of freedom, n. Depending on the values of n,
we can have many different chi-square distributions. The mean and the variance of chi-square distribution are
n, and 2n, respectively.
Because of its importance, the chi-square distribution is tabulated for various values of the parameter n (refer
table). Thus we may find in the table that value, denoted by 2 ( n) , satisfying p( X 2 (n)) , 0 1.
The example below helps on how to read chi-square distribution values.
Example 6.11: To read the chi-square value with 2 degrees of freedom where the area to the right of this value
is [Link] the degrees of freedom, 2, in the first column (df column) and then move horizontally until you
find the value of α , 0.005 in the first row. The point of intersection made by the horizontal and vertical
movement will give the desired chi-square value, 10.597. This value satisfies the following:
P( X 10.597) 0.005. In a similar way,The chi-square value with 100 degrees of freedom where the area to
the right of this value is 0.975 is 74.222.
The t distribution
The t distribution is an important distribution useful in inference concerning population mean/means. This
distribution has one parameter called the degrees of freedom. Depending on the values of the degrees of
freedom, we may have different t distributions. The degrees of freedom is usually denoted by n. In inference on
the population mean, the degrees of freedom is related to sample size. As the sample size or degrees of
freedom increases, the t distribution approaches the standard normal distribution.
The t- distribution shares some characteristics of the normal distribution and differs from it in others. The t
distribution is similar to the standard normal distribution in the following ways.
i) it is bell-shaped
ii) it is symmetrical about the mean
iii) the mean, median, and mode are equal to 0 and are located at the center of the distribution.
48
iv) The curve never touches the x-axis
The t distribution differs from the standard normal distribution in the following ways.
i) the variance is greater than 1.
ii) The t distribution is actually a family of curves based on the concept of degrees of freedom.
Objectives:
After a successful completion of this unit, students will be able to:
Differentiate the two major sampling techniques: probabilistic and non-probabilistic
Apply simple random sampling technique to select sample
Define sampling distribution of the sample mean
49
7.1 Methods of sampling
Definition of some basic terms
Sampling: is the technique of selecting representative sample from the whole.
Population: is the totality of elements or units under study.
Sample: is the part of the population.
Sampling Frame: A complete list of all the units of the population is called the sampling frame. A unit of
population is a relative term. If all the workers in a factory make a population, then a worker is a unit of the
population. If all the factories in a country are being studied for some purpose, then a factory is a unit of the
population of factories. The frame provides a base for the selection of a sample.
Major reasons to use sampling
1. Saves Time and Cost: As the size of the sample is small as compared to the population, the time and cost
involved on sample study are much less than the complete counts. Hence a sample study requires less time
and cost.
2. To prevent destruction: The destructive nature of some experiments (or inspection) do not allow to
carryout complete enumeration, for instance, to check quality of beers, to study the efficacy of new drugs,
testing the life length of a bulb, e t c.
3. Sample survey provides higher level of accuracy: This accuracy can be achieved through more selective
recruiting of interviewers and supervisors, more extensive training programs, a closer supervision of the
personnel involved and a more efficient monitoring of the field work.
Types of sampling
Generally, two types of sampling methods exist: probability and non-probability sampling.
Probability Sampling
The term probability sampling (or random sampling) is used when the selection of the sample is purely based on
chance. There is no subjective bias in the selection of units. Every unit of the population has a known nonzero
probability to be in the sample. The following are some of the t random sampling methods: Simple random
sampling, Stratified random sampling, Cluster sampling, Systematic random sampling.
For example, you may use the lottery method to draw a random sample by using a set of 'N' tickets, with
numbers ' 1 to N' if there are 'N' units in the population. After shuffling the tickets thoroughly, the sample of a
required size, say n, is selected by picking the required n number of tickets.
The best method of drawing a simple random sample is to use a table of random numbers. After assigning
consecutive numbers to the units of population, the researcher starts at any point on the table of random
numbers and reads the consecutive numbers in any direction horizontally, vertically or diagonally. If the read
out numbers corresponds with the one written on a unit card, then that unit is chosen for the sample.
Suppose that a sample of 6 study centers is to be selected at random from a serially numbered population of 60
study centers. The following table is portion of a random numbers table used to select a sample.
50
Row 1 2 3 4 5 …… N
Column
1 2315 7548 5901 8372 5993 ….. 6744
2 0554 5550 4310 5374 3508 ….. 1343
3 1487 1603 5032 4043 6223 ….. 0834
4 3897 6749 5094 0517 5853 ….. 1695
5 9731 2617 1899 7553 0870 ….. 0510
6 1174 2693 8144 3393 0862 ….. 6850
7 4336 1288 5911 0164 5623 ….. 4036
8 9380 6204 7833 2680 4491 ….. 2571
9 4954 0131 8108 4298 4187 ….. 9527
10 3676 8726 3337 9482 1569 ….. 3880
11 ….. ….. ….. ….. ….. ….. …..
12 ….. ….. ….. ….. ….. ….. …..
13 ….. ….. ….. ….. ….. ….. …..
14 ….. ….. ….. ….. ….. ….. …..
15 ….. ….. ….. ….. ….. ….. …..
N 3914 5218 3587 4855 4888 ….. 8042
If you start in the first row and first column, centers numbered 23, 05, 14,…, will be selected. However, centers
numbered above the population size (60) will not be included in the sample. In addition, if any number is
repeated in the table, it may be substituted by the next number from the same column. Besides, you can start at
any point in the table. If you chose column 4 and row 1, the number to start with is 83. In this way you can
select first 6 numbers from this column starting with 83.
The sample, then, is as follows:
83 75
53 33
40 01
05 26
Hence, the study centers numbered 53, 40, 05, 33, 01 and 26 will be in the sample.
Simple random sampling ensures the best results. However, from a practical point of view, a list of all the units
of a population is not possible to obtain. Even if it is possible, it may involve a very high cost which a
researcher or an organization may not be able to afford. In addition, it may result an unrepresentative sample by
chance.
Stratified sampling
Stratified random sampling takes into account the stratification of the main population into a number of sub-
populations, each of which is homogeneous with respect to one or more characteristic(s). Having ensured this
stratification, it provides for selecting randomly the required number of units from each sub-population. The
selection of a sample from each subpopulation may be done using simple random sampling. It is useful in
providing more accurate results than simple random sampling.
Systematic sampling
In this method, samples are selected at equal intervals from the listings of the elements. This method provides a
sample as good as a simple random sample and is comparatively easier to draw a sample. For instance, to study
the average monthly expenditure of households in a city, you may randomly select every fourth households
from the household listings
Cluster sampling
Cluster sampling is used when sampling frame is difficult to construct or using other sampling techniques
(simple random sampling) is not feasible or costly. For instance, when the geographic distribution of units is
scattered it is difficult to apply simple random sampling. It involves division of the population of elementary
units into groups or clusters that serve as primary sampling units. A selection of the clusters is then made to
form the sample. The precision of estimates made based on samples taken using this method is relatively low.
51
Non-probabilily sampling techniques
In non-probability sampling, the sample is not based on chance. It is rather determined by personal judgment.
This method is cost effective; however, we cannot make objective statistical inferences. Depending on the
technique used, non-probability samples are classified into quota, judgment or purposive and convenience
samples.
2. If sampling is with replacement we will have Nn = 32 = 9 possible samples: (A, A), (A, B), (A, C), (B,
A), (B, B), (B, C), (C, A), (C, B) and (C, C). Hence the probability distribution (sampling distribution)
of the sample mean is:
x 3 4.5 6 7.5 9
P(X = x ) 1/9 2/9 3/9 2/9 1/9
E ( X ) = x P( x ) = 3(1/9) + 4.5(2/9) + 6(3/9) + 7.5(2/9) + 9(1/9) = 6
V ( X ) = ( x 2 P ( x )) x = (1 + 4.5 + 12 + 12.5 + 9) – 36 = 3
2
Note:
The mean of the sampling distribution of the sample mean is the same as the population mean
irrespective of the sampling procedure.
The variance of the sampling distribution of the sample mean is:
2
, if sampling is with replacement
n
2
N n , if sampling is without replacement
n N 1
The problem with using sample mean to make inferences about the population mean is that the sample
mean will probably differ from the population mean. This error is measured by the variance of the
52
sampling distribution of the sample mean and is known as the standard error. The standard error is the
average amount of sampling error found because of taking a sample rather than the whole population.
As sample size increases, the standard error decreases.
7.3 Central Limit Theorem
If X1, X2, …, Xn is a random sample from a population with mean μ and variance σ2, then as n goes to infinity
the distribution of the sample mean, X , approximates normal distribution with mean μ and variance σ2/n. That
X
is, as n gets large, X N (μ, σ2/n) and its standardized form is Z ~ N (0,1).
/ n
Note: The central limit theorem is useful for approximating the distribution of the sample mean based on a large
sample size and when the population distribution is non normal; however, if the population is normal, then the
sampling distribution of the sample mean will be normal regardless of the sample size.
Example 7.2: If the uric acid values in normal adult males are normally distributed with mean 5.7 mgs and
standard deviation of 1mg. Find the probability that
a) a sample of size 4 will yield a mean less than 5
b) a sample of size 9 will yield a mean greater than 6
Solution: Let X be the amount of uric acids in normal adult males with mean 5.7 and variance 1.
a) If a sample of size 4 is taken, then X ~ N (5.7, 0.25) since the population is normally distributed.
5 5.7
P( X 5) P( Z ) P( Z 1.4)
0.5
0.5 P(0 Z 1.4) 0.0808
b) If a sample of size 9 is taken, then X ~ N (5.7, 1/9) since the population is normally distributed.
6 5.7
P( X 6) P( Z ) P( Z 0.9)
1
3
0.5 P(0 Z 0.9) 0.1841
53
Interval estimate: In most practical problems, a point estimate does not provide information about ‘how close
is the estimate’ to the population parameter unless accompanied by a statement of possible sampling errors
involved based on the sampling distribution of the statistic. Hence, an interval estimate of a population
parameter is a confidence interval with a statement of confidence that the interval contains the parameter value.
An interval estimate of the population parameter consists of two bounds within which the parameter will be
contained:
L U
where L is the lower bound and U is the upper bound.
Case 1: When the population is normal.
If the variance 2 is known, the sampling distribution of the sample mean X is normal with mean and
2 2 X
variance . i.e., X ~ N , and Z ~ N(0,1).
n n
n
X
If the variance 2 is unknown, t will have t-distribution with
S
n
n - 1 degrees of freedom. Moreover, as the sample size increases t is approximately the same as standard
normal.
Consider the case 2 is known, we can derive a (1 )100% confidence interval for the population mean .
Let Z be a point on the standard normal curve that cuts an area of to the right. i.e. P( Z Z ) = . By
2 2 2 2
the symmetric property of the normal distribution, P( Z Z ) = (see the diagram below).
2 2
From the standard normal distribution, we know that
P( Z Z Z ) 1
2 2
To obtain the limit of the interval estimate, we use the standardized form of X in the above probability
X
statement. i.e., letting Z
n
P( Z Z Z ) 1 Becomes
2 2
X
P( Z Z ) 1
2 2
n
P( Z X Z ) 1
2 n 2 n
54
P( X Z X Z ) 1
2 n 2 n
P( X Z X Z ) 1
2 n 2 n
We can assert with probability 1 that the interval ( X Z X Z ) contains the population
2 n 2 n
mean we are estimating.
55
S S
X t (n 1) , X t (n 1) .
2 n 2 n
And from the t distribution table, t (n 1) t 0.025 (5) 2.571
2
0.95 0.95
2.28 (2.571) , 2.28 (2.571)
6 6
(2.28-0.997, 2.28+0.997)
(1.28, 3.27)
We are 95% confident that the mean drop in blood pressure lies in between 1.28 mmHg and 3.27 mmHg for the
sampled population.
Example 8.2: Punctuality of patients in keeping appointment is of interest to a research team. In a study of
patients flow through the office of general practitioners, it was found that a sample of 35 patients were 17.2
minutes late for appointments, on the average. Previous research had shown the standard deviation to be about 8
minutes. The population distribution was felt to be not normal. What is the 90 percent confidence interval for
the true mean amount of time late for appointment?
Solution: Given: X 17.2 , 8 , n 35
(1 )100% 90% 1 0.90 0.1 0.05
2
Since the sample size is fairly large (n > 30), and since the population standard deviation is known, according to
the central limit theorem, the sampling distribution of sample mean is approximately normal. Thus, a
confidence interval of the population mean is given by:
X Z , X Z
2 n 2 n
And from the standard normal distribution table, Z Z 0.05 1.65
2
8 8
17.2 (1.65) , 17.2 (1.65)
35 35
(17.2 – 2.2, 17.2 + 2.2)
(15.0, 19.4)
Therefore, the 90% confidence interval for true mean amount of time late for appointment is between 15.0 and
19.4 minutes.
57
Figure: Area of acceptance and rejection of H 0 (Two-tailed test)
Based on the form of the alternative hypothesis and the test statistic we can make the following decisions:
i. For H 1 : 0 (two-tailed test) reject H 0 if Z Z .
2
58
We can summarize the decsion rules as follows:
Alternative hypotheses
Decision H1 : 0 H1 : 0 H1 : 0
Reject H 0 : 0 if Z Z Z Z Z Z
2
Reject H 0 : 0 if t t ( n 1) t t (n 1) t t (n 1)
2
Null Hypothesis ( H 0 )
Decision True False
Reject H 0 Type I error ( )
Correct decision
Accept H 0 Correct decision Type II error ( )
Type I error is committed if we reject the null hypothesis when it is true. The probability of committing a type I
error, denoted by is called the level of significance. The probability level of this error is decided by the
decision-maker before the hypothesis test is performed. Type II error is committed if we do not reject the null
hypothesis when it is false. The probability of committing a type II error is denoted by (Greek letter beta). As
type one error increases type two error will decrease (they are inversely proportional). Hence we cannot reduce
both errors simultaneously. As the sample size increases both errors will decrease.
Example 8.3: The life expectancy of people in the year 1999 in a country is expected to be 50 years. A survey
was conducted in eleven regions of the country and the data obtained, in years, are given below:
Life expectancy (years): 54.2, 50.4, 44.2, 49.7, 55.4, 47.0, 58.2, 56.6, 61.9, 57.5, and 53.4.
Do the data confirm the expected view? (Assuming normal population) Use 5% level of significance.
Solution: Let be the life expectancy of people in the year 1999 in a country.
1. H 0 : 50 (The life expectancy of people in the year 1999 in a country is 50 years)
H1 : 50 (The life expectancy of people in the year 1999 in a country is different from 50 years)
2. Level of significance, α = 0.05.
3. Since is unknown and the population is normal, the t-test statistic is appropriate.
Given: n = 11; 0 50 and we need to compute X and s .
11
x i
54.2 50.4 ..... 57.5 53.4 598.5
X i 1
54.41
n 11 11
11
1
x i
xi 1
2
32799.91
(598.5) 2
S
2 2
n 1 n 10 11
1
(236.07) 23.607
10
59
S 23.607 4.859
Then, the t-test statistic is calculated as:
X 0 54.41 50 4.41
t 3.01
S 4.859 1.465
n 11
4. For α = 0.05 and two-tailed test, the critical (table) value is:
t (n 1) t 0.05 (11 1) t 0.025 (10) 2.228
2 2
Since t 3.01 t (n 1) 2.228 reject the null hypothesis H 0 . That is, the calculated t value lies in
2
the rejection region (the shaded region).
5. Conclusion: The data do not confirm the expected view. That is, the life expectancy is different from 50
years at 5% level of significance.
Example 8.4: Suppose that we want to test the hypothesis with a significance level of .05 that the climate has
changed since industrialization. Suppose that the mean temperature throughout history is 50 degrees. During
the last 40 years, the mean temperature has been 51 degrees and the population standard deviation is 2 degrees.
What can we conclude?
Solution:
Let be the mean temperature.
1. H 0 : 50 (There is no change in temperature since industrialization)
H1 : 50 (There is change in temperature since industrialization)
2. Level of significance, α = 0.05.
3. Since n = 40 is large, the Z-test statistic is appropriate.
Given: n = 40; = 2; X = 51; 0 50
X 0
51 50 1
Z
= 3.16
2 0.316
n 40
4. For α = 0.05 and two-tailed test, the critical (table) value is:
Z Z 0.05 Z 0.025 1.96
2 2
Since Z 3.16 Z Z 0.025 1.96 reject the null hypothesis H 0 . That is, the calculated Z value
2
lies in the rejection region (the shaded region).
5. Conclusion: There has been a change in temperature since industrialization, at 5% level of significance.
Example 8.5: A study was conducted to describe the menopausal status, menopausal symptoms, energy
expenditure and aerobic fitness of healthy midwife women and to determine relationship among these factors.
60
Among the variables measured was maximum oxygen uptake (Vo2max). The mean Vo2max score for a sample of
242 women was 33.3 with a standard deviation of 12.14. On the basis of these data, can we conclude that the
mean score for a population of such women is greater than 30? Use 5% level of significance.
Solution:
Let be the mean Vo2max score for a population of healthy midwife women.
1. H 0 : 30 (The mean score for a population of healthy midwife women is 30)
H1 : 30 (The mean score for a population of healthy midwife women is greater than 30).
2. Level of significance, α = 0.05.
3. Since n = 242 is large, the Z-test statistic is appropriate.
Given: n = 242; S = 12.14; X = 33.3; 0 30
X 0 33.3 30 3.3
Z = 4.23
S 12.14 0.7804
n 242
4. For α = 0.05 and right-tailed test, the critical (table) value is:
Z Z 0.05 1.65
Since Z 4.23 Z 1.65 reject the null hypothesis H 0 . That is, the calculated Z value lies in the
rejection region (the shaded region).
5. Conclusion: The mean Vo2max score for the sampled population of healthy midwife women is greater
than 30 at 5% level of significance.
8.3 Test of Association (Independence)
Usually we encounter with nominal scale data. The 2 test of association is useful for determining whether
there is any relationship or association exists between two nominal variables. For instance, we might be
interested in the relationship between HIV status with sex, lung cancer and smoking habit, political affiliation
and sex, e t c.
When observations are classified according to two variables or attributes and arranged in a table, the display is
called a contingency table as shown below:
The test of association or independence uses the contingency table format. Here the variables A and B have
been classified into mutually exclusive categories. The values Oij in row i and column j of the table shows the
observed frequency falling in each joint category i and j. The row and column totals are the sums of their
corresponding frequencies. The sum of row or column totals will give grand total n, which represents the
61
sample size. The procedures to test the association between two independent variables is summarized as
follows:
Step 1: State the null and alternative hypotheis
H 0 : There is no association or relationship exists between two variables, that is, the two variables are
independent.
H 1 : There is association or relationship between two variables, that is, the two variables are dependent.
Step 2: State the level of significance, .
Step 3: Calculate the expected frequencies, Eij, corresponding to the observed frequency in row i and column j.
The expected frequencies in each cell are calculated as:
Row i total Column j total Ri C j
Eij
Sample size n
Step 4: Compute the value of test-statistic:
r c (O E ) 2
Cal
2 ij ij
i 1 j 1 E ij
where Oij is the observed frequency of row i and coulumn j and Eij is the expected frequency of row i and
coulumn j.
Step 5: Find the critical (table) value of (df ) (from Appendix..). The value of correponds to an area in
2 2
Example 8.6: The following data on the colour of eye and hair for 6800 individuals were obtained from a
source:
Eye colour
Hair colour Fair Brown Black red Total
Blue 1768 808 190 47 2813
Green 946 1387 746 43 3122
Brown 115 444 288 18 865
Total 2829 2639 1224 108 6800
Test the hypothesis that hair colour and eye colour are independently distributed (there is no association
between colour of eye and colour of hair) at the level of = 0.01.
Solution:
1. H 0 : There is no association between hair colour and eye colour.
H 1 : There is association between hair colour and eye colour.
2. = 0.01.
3. Calculate the expected frequencies, Eij
Ri C j
Eij
n
2813 2829 2813 108
E11 1170.29 ……………….. E14 44.68
6800 6800
865 2829 865 108
E31 359.87 ………………….. E34 13.74
6800 6800
62
Therefore, the contingency table for expected frequencies is as follows:
Eye colour
Hair colour Fair Brown Black red Total
Blue 1170.29 1091.69 506.34 44.68 2813
Green 1298.84 1211.61 561.96 49.58 3122
Brown 359.87 335.70 155.70 13.74 865
Total 2829 2639 1224 108 6800
4. Calculate the test statistic:
r c (O E ) 2
Cal 2
ij ij
i 1 j 1 E ij
(1768 1170.29) 2 (47 44.68) 2 (946 1298.84) 2
Cal 2
..... .....
1170.29 44.68 1298.84
(43 49.58) 2 (115 359.87) 2 (18 13.74) 2
.....
49.58 359.87 13.74
Cal 1074.43
2
df = (r – 1) (c – 1) = (3 – 1) (4 – 1) = (2) (3) = 6
2 (df ) 0.012 (6) 16.812
6. Since Cal 1074.43 > (df ) 16.812 Reject H 0 .
2 2
7. Conclusion: There is association between hair colour and eye colour. That is, hair colour and eye colour
are dependent.
Objectives:
Having studied this unit, you should be able to:
formulate a simple linear regression model.
express quantitatively the magnitude and direction of the association between two variables
Introduction
The statistical methods discussed so far are used to analyze the data involving only one variable. Often an
analysis of data concerning two or more variables is needed to look for any statistical relationship or association
between them. Thus, regression and correlation analysis are helpful in ascertaining the probable form of the
relationship between variables and the strength of the relationship.
63
The first step in regression analysis involving two variables is to construct a scatter plot (diagram) of the
observed data. Scatter diagram is a plot of all ordered pairs ( X i , Yi ) on the coordinate plane which is helpful for
determining an apparent relationship between two variables.
The simple linear regression of Y on X can be expressed with respect to the population parameters and as
Y X
where = y-intercept that represents the mean value of the dependent variable Y when the independent
variable X is zero; = slope of the regression line that represents the change in the mean of Y for a unit
change in the value of X ; = error term
The population parameters and can be estimated from sample data using the least square technique. The
estimators of and are usually denoted by a and b, respectively. The resulting regression line is
Y abX
and the equation is known as the fitted regression line. The estimated values of Y are denoted by Y . The
observed values of Y are denoted by y. The difference between the observed and the estimated values, Y - Y , is
known as error or residual, and is denoted by ˆ . The residual can be positive, negative or zero.
A best fitting line is the one for which the sum of squares of the residuals, ˆ 2 has the minimum value. This is
called the method of least squares. According to this method, one would select a and b such that ˆ 2
=
(Y Y ) 2 is minimum. The solution of this minimization problem using partial differentiation is as follows:
X Y
XY n n XY X Y
b = and a Y bX
( X ) 2 n X 2 ( X ) 2
X 2
n
Example 9.1: A researcher wants to find out if there is any relationship between height of the son and his
father. He took random sample of 6 fathers and their sons. The height in inch is given in the table below:
Height of father (X) 63 65 64 65 67 68
Height of the son (Y) 66 68 65 67 69 70
i) Draw the scatter diagram and comment on the type of relationship.
ii) Fit the regression line of Y on X.
iii) Predict the height of the son if his father’s height is 66 inch.
Solution:
i)
From the scatter plot one can see that the points are roughly on straight line.
ii)
n6 X 392 , Y 405 , X 25628, XY 26476, Y 27355
2 2
64
n XY X Y 6(26476) (392)(405) 405 392
b = 0.923 a Y bX 0.923 = 7.2
n X ( X )
2 2
6(25628) (392) 2
6 6
Then the fitted (regression) line of Y on X is given by:
Y a b X = 7.2+0.923X
The slope of the line, i.e. b=0.923, tells us that a unit (one inch) increase in the height of the father
results in 0.923 inch increase in the height of the son.
The y-intercept of the line, i.e. a=7.2, is the value of Y when the value of X is zero(do you think that
the intercept is meaningful?)
iii) Y=7.2+0.923(66) =68.118, thus the height of the son is 68.118 inch.
n 1 n 1
=
( X X )(Y Y )
( X X ) (Y Y )
2 2
65
Example 9.2: In some locations, there is strong association between concentrations of two different pollutants.
An article reports the accompanying data on ozone concentration x (ppm) and secondary carbon concentration y
( g / m 3 ) :
X 0.066 0.088 0.120 0.050 0.162 0.186 0.057 0.100
Y 4.6 11.6 9.5 6.3 13.8 15.4 2.5 11.8
a. Calculate the correlation coefficient and comment on the strength and direction of the relationship
between the two variables.
Solution: The summary quantities are
n 16, xi 1.656, y i 170.6, xi y i 20.0397, xi 0.196912, y i 2253.56
2 2
66
67
68
69