0% found this document useful (0 votes)
2 views57 pages

Introduction Statistics

The document outlines a course on basic statistical concepts and methods with applications in agriculture. It covers topics such as population and sample definitions, types of variables, data collection methods, descriptive statistics, probability, and hypothesis testing. The course aims to equip students with the ability to analyze agricultural data and make statistical inferences.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views57 pages

Introduction Statistics

The document outlines a course on basic statistical concepts and methods with applications in agriculture. It covers topics such as population and sample definitions, types of variables, data collection methods, descriptive statistics, probability, and hypothesis testing. The course aims to equip students with the ability to analyze agricultural data and make statistical inferences.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

INTRODUCTION TO STATISTICS

Lecture 1: General Introduction:


Definitions and notations; Population,
COURSE DESCRIPTION: sample, statistical concepts and their
The course introduces the basic application in agriculture

statistical concepts and methods and Lecture 2: Variable:


their use in agriculture and other Definition of variables; Types of
variables; Quantitative and
fields. It covers descriptive summary Qualitative variables
and graphical display of data;
Lecture 3: Understanding data:
elementary probability; binomial and Method of data collection – Sampling
normal distribution; estimation and survey and experiment, methods of
sampling; types of data
hypothesis testing; simple linear
regression and correlation. The course Lecture 4: Numerical Description of
Data – Measures of Location:
uses examples and applications to Arithmetic mean (average), trimmed
which students can easily relate. mean, median, mode, Quartile,
advantages and disadvantages of each
method
COURSE OBJECTIVES:
Lecture 5: Numerical Description of
The aim of the course is to introduce Data - Measures of
the basic statistical concepts and Dispersion/Variability:
Range, variance, standard deviation,
methods commonly used in coefficient of variation and inter-
agriculture. In particular, the students quartile range

should be able to: Lecture 6: Graphical Description of


1. Define and apply the basic Data:
Qualitative data; frequency table, bar
statistical terms, notations and chart, pie chart, Quantitative data:
concepts stem-and-leaf display, histogram

2. List and appreciate the Lecture 7: Introduction to


importance of statistics in Probability:
Definition of terms (probability
Agricultural studies experiment, outcome, sample space,
3. Present data using appropriate events); Classical probability,
Empirical probability
descriptive summaries, including
tables, diagrams, graphs and Lecture 8: Laws of Probability:
General probability rules, Additive
descriptive statistics. laws of probability, Multiplicative
4. Make statistical inferences based laws of probability of independent
events
on sample data by constructing
confidence intervals for Lecture 9: Conditional, Joint and
Marginal Probability:
population means and
differences, and testing Lecture 10: Random Variables and
Probability Distribution:
hypotheses. Definition of random variables,
5. Differentiate the use of Probability of random variable,
Binomial distribution
correlation and regression
analysis. Lecture 11: Normal Distribution,
Sampling Distribution and
Confidence Interval
COURSE OUTLINE:

1
Probability distribution of a normal 1.2 FIELDS OF STATISTICS
random variable; the standard normal
In our discussion of statistics, we will look at
variable, use of the standard normal
distribution table and the Student’s t- the two major fields of statistics i.e. descriptive
distribution, the central limit
and inferential statistics. In descriptive
theorem; confidence interval
statistics, we will take collected information
Lecture 12: Hypothesis Testing:
(data) and attempt to summarize the
Definition of hypothesis, the null and
alternative hypotheses, Type I and information into a few numbers or graphs. One
Type II errors, Significance level,
example of a descriptive statistics is the mean.
two-tail and one-tail test of
significance Here, you are simplifying the data set by
looking only at what the average value
Lecture 13: Hypothesis Testing
Continued: happens to be. Yet another example of a
One sample z or normal test and t-
descriptive statistics is a histogram or bar
test, Two sample z test and t-test
chart. In that case, you are summarizing all of
Lecture 14: Relationship between
the information in the data set into a simple
two variables:
Scatter plot and its role; Correlation plot (picture). On the other hand, inferential
analysis and its application
statistics will take a data set and attempt to
Lecture 15 Relationship between two make assertions on the population as a whole.
variables:
Rather than saying, "...75% of students who
Regression analysis and its
application attended an interview to joint S. 1 failed
mathematics." you may wish to talk about
LECTURE ONE:
what this indicates about the population (all
GENERAL INTRODUCTION
primary school leavers) as a whole. Does it
seem very likely that over half of the pupils
1.0 INTRODUCTION
who leave primary schools do not have good
In this lecture, you will be introduced to basic
mathematics background based on the results
concepts of statistics. The lecture will explain
of this interview?
the importance of statistics in agriculture;
define the different terms and notations used
frequently in statistics.
1.3 WHY SHOULD YOU STUDY
STATISTICS
1.1 WHAT STATISTICS IS
Agriculture is the backbone of Uganda’s
We all use statistics in our daily lives such as
economy and most people derive their
when discussing football, election results,
livelihood directly through involvement in it.
businesses etc. In statistics, you study how to
In order to solve agricultural problems such as
gather information (data), summarize
low productivity due to pests and diseases, soil
information, analyze the information collected
degradation and other production constraints
and reporting the results.
scientists carry out research. In the research
You can use statistics in different fields (e.g. process, scientists are involved in collecting,
agriculture, business, etc.). In each field,
summarizing and analysing information thus
statistics is modified accordingly but the basic
concepts are the same. they need to apply the knowledge of statistics.
This sometimes involved comparing different
technologies such as methods of pest and
2
disease control, soil erosion control etc all of population (e.g. number of teachers in your
which involved the use of statistics. Apart school) is referred to as the population size
from its use to support research, statistics is and if finite is denoted by N (capital letter).
also used for monitoring agricultural
production of a country. The status of Activity

agriculture production needs to be monitored 1. Give five examples of finite population

from year to year to assess progress and 2. Give five examples of infinite

capacity of the country to feed her population. population

As some body learnt, you can also use


statistics to evaluate published numerical facts. 1.4.2 Sample/Random Sample
You have often heard politicians criticizing the When you are interested in studying a
results of the opinion poll published in population which is very large, e.g. all the
newspapers. schoolchildren in your county, because of
reasons such as cost and time involved, this
1.4 BASIC DEFINITIONS may not be easy. In statistics, it is
As a student studying statistics, you need to recommended that you select and study only a
acquaint yourself with a number of common subset of that population. This subset of the
terms and notations that are frequently used in population you have selected for your study is
statistics. This will make you feel at home what we called a sample. If you choose your
whenever you are reading statistical texts or sample in such a way that every possible
discussing with friends. sample from the population has an equal
chance of being selected, then the sample is
1.4.1 Population referred to as a random sample. The number
Population is a set of all subjects or units or of subjects or units in your sample is referred
elements that you are interested in studying. to as sample size and is denoted by n (small
The subjects or units or elements may be letter).
human, plants, animals, machines, buildings,
organisations, cities, days of the week etc. If Take note
§ A sample may be defined as a subset of the
you want to study the characteristics of
population that has been selected for a given study
students in your school, then all the students in
to obtain information about the population
the school will constitute your population.
§ A sample is said to be a random sample if it is
Other examples of population are all plants on
selected in such a way that every possible sample
a given plot of land, all teachers in your from the population has an equal chance
school, all grains of sand on the riverbank, etc. (probability) of being selected
We know that in your school, it is possible to § Any sample for your study should be selected such
count the number of teachers and in this case, that it is a true representative of the population you

the population of teachers is called a finite want to study. Statistics help us in selecting a
suitable sample.
population. On the other hand, it is not
possible to count all the grains of sand on the 1.4.2 Parameter
riverbank and this kind of population is In your study, you will always be interested in
referred to as infinite. The total number of a particular characteristics of the population
subjects or units that constitute your e.g. performance in mathematics of
3
schoolchildren in all schools in your county, 1.4.4 Estimate
heights of plants on a given plot of land, milk We now know that in most studies, we use a
production of cows in your district, etc. You sample to obtain information about the
can described or summarised those population population. Since it is generally not possible,
characteristics (mathematics performance, or very difficult, to calculate the value of the
plant height, milk production) using numbers. population parameter directly, we calculate a
The numerical description or summary of your sample statistic that corresponds to that
population characteristics is referred to as a parameter and use it as an estimate. A sample
PARAMETER (denoted by Greek letters). statistic is thus referred to as an ESTIMATE of
Examples include; population mean a population parameter. Sample mean for
(denoted µ (mu)), population variance example estimates population mean and thus it
is its estimate.
(denoted σ2 (sigma square)), population
standard deviation (denoted σ (sigma)). Since
1.4.5 Inference
a parameter is a measure from the entire In your study, you will use the information
population, it is a fixed value i.e. there is only obtained from the sample to draw conclusion
one value for each parameter for a given or generalization about the whole population.
population. For example if you realized that the 100
schoolchildren you selected from your county
1.4.3 Statistic performed very well in mathematics you can
As we have seen from above, in most of our therefore conclude that schoolchildren in your
studies we cannot study every member of the county do well in mathematics. It involves
population of interest and thus resort to generalizing from a part (sample) to the whole
studying a sample. Any numerical description (population). This generalisation from a
or summary of a sample is referred to as a sample to the population is what we called
STATISTIC (denoted by Latin letters). statistical inference.
Examples include; sample mean (denoted
2
by x (x bar)), variance ( S (s square)), 1.4.6 Experiment/Trial/Event

standard deviation ( S ). When you are Sometimes when we are carrying out studies,

selecting a sample from your population, there we are involved directly in generating

are many possible samples that can be selected information. For example, you may plant

and each sample could have different values of different varieties of maize on different plots

a given statistic. For example if you want take and later take measurement on plant height and

a sample of size 100 from the 3000 grain yield for comparison. The whole process

schoolchildren in your county, you will realise of generating information in this case from

that they are many ways to pick 100 from the planting up to measuring of plant height in

3000 children and each sample will of course statistics is referred to as an experiment. An

have different sample mean. Statistic thus experiment can be simple as; tossing a coin to

varies from sample to sample compared to the see whether head or tail turn up, rolling a die

population parameter which is fixed. to see the face turning up, closing your eye and
pointing in the direction of students to see

4
whether you pointing at a male or female 2.2 TYPES OF VARIABLES
student. In your studies, you may be interested in many
characteristics of the population (variables).
Take note
For example in studying the farms in your
§ In a simple term you can define an experiment as
district, you may be interested in knowing the
a process by which an observation (or
following:
measurement) is obtained.
§ Whether the farms made profit the previous
§ The experiment, when performed is called a trial
§ An event is an outcome of the experiment year or not
§ Amount of maize harvest in tones the
previous year
§ The number of goats on the farm
LECTURE TWO: VARIABLES § The number of farm workers employed
2.0 INTRODUCTION § Farmers opinion on the amount of rain in the
In lecture 1, you were introduced to general current year
ideas of statistics. You learnt how to define As seen from above, those variables are of
different statistical terms and notations. In this different types. In statistics, these different
lecture, you will learn about the different types types of variables are put into groups based on
of variables used in statistics and other studies. their measurement scales and other
characteristics.
2.1 WHAT IS A VARIABLE? Broadly, variables may be of two types:
We now know that in our studies, we are quantitative or qualitative.
interested in some particular characteristics of
the population and these characteristics vary 2.2.1 Quantitative variables
from one subject or unit to another. For “These are variables or characteristics that
example, mathematics performance varies represent amount or quantity of something and
from one child to another, milk production can be measured over a range of values”.
varies from one cow to another and plant Those are characteristics that you can quantify
height varies from one plant to another. A or measure using some defined scale. For
characteristic or attribute that varies from one example, maize harvested on a given farm can
subject or unit to another in statistics is be quantified and measured using weighing
referred to as variable. In your study, you will scales, the number of goats on a farm can be
have to take measurements or records on the quantified by physical counting. However,
units that are in the selected sample. although both maize harvested and the
Take note numbers of goats on a farm are quantitative
§ The value of a variable that has been recorded
or measured on a particular subject or unit is variables, they differ. When you dealing with
referred to as an observation maize harvested on a farm, it is possible to get
§ A data is a set of observations usually of several
variables taken on many individual subjects or yield of 4.5 tones but you cannot get 4.5 goats.
units. The individual subjects or units may be
human, plants, animals, machines, buildings, This brings us to another division within the
organisations, etc quantitative variables category. Quantitative
variables may be either discrete or continuous.

5
iii. Nominal/Categorical variables are
Take note
§ A continuous quantitative variable is one associated with some quality,
for which all values in some range are characteristic or attribute which the
possible e.g., yield of maize, height, mass,
etc subject posses; for example, eye
§ A discrete quantitative variable is one for
which only certain values are possible. colour: (blue – grey – green - brown),
They are consecutive integers; for example, vehicle type: (bus – truck – car); etc.
number of goats on a farm, size of
household, number of insects in an insect In this case, there are more than two
trap, number student in your class etc.
categories.

2.2.2 Qualitative variable LECTURE THREE:


UNDERSTANDING DATA
With this type of variables, you will simply
describe the characteristics or attributes of an 3.0 INTRODUCTION

individual subject or unit. It includes variables In lecture 2, you learnt about the different
that can be categorised but not quantified. You types of variables. In this lecture, you will
learn methods of data collection, sources of
can for example categories students using and types of data.
name of their villages, whether they are
From what we learnt from lecture 1, we can
present or absent on a given day. Class grading view statistics as consisting two parts; a)
such as good, very good or excellent are also collection of information (data) and b)
processing information (summarizing and
examples of qualitative variables. Just as we analyzing data), in this lecture, we will dwelt
found out in quantitative variables, qualitative on collection of data.

variables can also be broken down into


different classes. Qualitative variables may be 3.1 METHODS OF DATA COLLECTION
In statistics, methods for collecting data can be
ordinal, binary or Nominal/Categorical.
broadly divided into two .i.e. sample survey
and experiment.
i. Ordinal variables deal with relative
differences, e.g., (short – average -
3.1.1 Sample Survey
tall), (bad – good - excellent), etc.
In a survey, you as an investigator solicit view,
This type of variables is often
opinion, or existing information without
represented by `scores’ such as 1 – 2
control over the factors that you are studying.
– 3, etc. and such are erroneously
For example if you are interested in studying
treated statistically as discrete
“why” schoolchildren come to school late on
variables. Ordinal variables have
Mondays, all you need to do is select a number
natural order between the different
of students (a sample) who come late on
categories (one category is “bigger”
Mondays and ask them the reasons why they
than the other but we can not tell by
are late. In this case, you are interested in
how much)
studying the factors responsible for lateness on
ii. Binary variables are qualitative
Mondays. Your study may reveal that some
variables which have only two
students come late because of distance from
possible categories e.g. a student has
school while others come late because they go
either passed or failed, a farm makes
to bed late. As an investigator, you have not
profit or loss, etc.
done anything (not manipulated the factors

6
responsible) to result to lateness of the the weight of goats in a given village, we can
students. consider farmers or households as our clusters.
Once a household is selected for our study, all
You may ask yourself this question “how do I the goats in the household will be weighted,
select a suitable sample?” In statistics, a goats in non-sampled households will not be
sample is selected through a process referred studied.
to sampling.
Stratified Sampling: This is used when the
population is divided into groups or strata in
Sampling Methods
which there is less variation within groups and
Sampling is a process by which a sample is
more variation between groups. For example if
selected from a given population so as to
we look at two classes (P1 and P7) in a
obtained information about the population. The
primary school, within each class height of
main idea behind sampling is to ensure that the
pupils are similar BUT between the two
sample selected is a good representative of the
classes there a big difference. In stratified
target population. For example if you are
sampling, we would do a simple random
interested in finding out the proportion of
sampling from each group to contribute to our
teachers in your school who drink alcohol, ten
study sample.
teachers picked from “Mama Brown’s bar” is
certainly not a good representative and neither
3.1.2 Experiment
is a group of teachers picked from a born-again As a teacher of agriculture in your school, you
church or a mosque. The sample that may want to demonstrate to the neighboring
misrepresents the population is referred to as a community that application of fertilizer
biased sample whereas a sample that is a good improves the yield of maize or that a certain
representative of the population is unbiased. new variety of maize yield better than the local
variety. To prove your point you need to set

Activity up plots of land on which you plant maize and


apply fertilizer. On some plots, you will apply
Give three examples of a biased sample
fertilizer while on others you will not and then
wait to compare yield from the different plots.

Simple random sampling: It is a type of In this study, you are interested in studying the

sampling where each member of the effect of fertilizer or variety (factor of interest).

population has an equal chance of being Since you can determine which plot of land

selected. The sampling can be done using a receive fertilizer and which ones do not

table of random numbers or by what is called receive, you are playing an active role in

lottery method. controlling or manipulating the factors of


interest. This form of study in which the

Cluster sampling: This type of sampling investigator is actively involved in

applies when the population is grouped into controlling/manipulating the factors of interest

clusters and it is the clusters which are is referred to an experiment.

sampled. Once a particular cluster is selected


all the members of the cluster will be studied.
For example if we are interested in measuring
7
data whereas if the data in question were
Take note
collected for some other purpose other than the
In an experiment, an investigator plays an present study (but being used in this study)
active role in a study by controlling the
environment and administering a treatment(s). then this data is referred to as secondary data.
The treatments are imposed by investigator
using standard protocols and may infer that the
response was due to treatment. 3.2.3 Observational verses Intervention
Data
This is based on the process of data collection.

An experiment is based on three basic Observational data are collected simply by

principles of randomization, replication and observing the process under investigation

local control (blocking). The principle of (conducting sample surveys) whereas

randomization deals with the procedure of intervention data come from doing

assigning the treatments (factors of interest) to experiments.

the different subjects or study units.


Replication is the application of each treatment LECTURE FOUR:

or factor to more than one study unit to ensure NUMERICAL DESCRIPTION OF DATA –

the results that will be obtained can be reliable. MEASURES OF LOCATIONS

For example in your fertilizer demonstration,


you need to have two (2) or more plots treated 4.0 INTRODUCTION

with fertilizers and a similar number without In our last lecture, we looked at the different

fertilizers. In local control, we are concern aspect of data. After collecting a “good data”,

with how to control for other factors that we you need to process that information through

are not interested in studying. summary and analysis of data.

3.2 Types of Data No matter what your final objectives are, you

As you read different books, you will realize must first adequately describe your data. This

that different field of studies or authors applies to both sample and population data.

categorized data differently. We will look


briefly at a few categorization of data. Good descriptive statistics enable you to make
sense of the data by reducing a large set of

3.2.1 Qualitative verses Quantitative Data measurements to a few summary measures that

This categorization is mainly used in social provide a good, rough picture of the original

research. Qualitative data is a data in which the measurements. There are two methods for

information is recorded inform of text, description of data: Numerical and Graphical

audiovisual and pictures whereas quantitative methods. In this lecture, you will learn how to

data has information stored inform of numbers. use numbers to summarize or describe your
data.
3.2.2 Primary verses Secondary Data

This classification is based on the purpose for 4.1 MEASURE OF LOCATION

which the data is collected. If the data have


been collected specifically for the current As a teacher, you may want to group your

study, then the data is referred to as primary class based on marks scored in your subject
8
and the best way to do that is to look at how Example 4.1
the marks are distributed. You need to know
the lowest and highest marks scored, range of If in your demonstration of the effect of
marks scored by most students, the marks that fertilizers on maize yield (see section 3.1.2)
divide the students into two groups, the mark you collected maize yield data from 10 plots
below which the worst 25% performers lie, ( x1 , x 2 , x3 , x 4 , x5 , x 6 , x 7 , x8 , x9 , x10 ) on
etc. What we have just described about your
which fertilizer was applied and would want to
class marks and students’ grouping are what
get the mean of this sample. The sample mean
we call measures of location. We are going to
is
look at arithmetic mean, median, mode and
quartiles as examples of measures of location i =10

of data. x = ( x1 + x 2 + x3 + ... + x10 ) / 10 = ∑ xi / 10


. Now let us try to put some imaginary figures
i =1
to your fertilizer demonstration. Suppose at the
4.1.1 Arithmetic mean (Average)
end of your demonstration you got the
This the most common measure of centre of
following results from your 10 plots onto
the population. The population mean denoted
which you applied fertilizer ( x1 =25, x 2 =20,
by the Greek letter µ , is the numerical value

that locates the balance point or centre of the


x3 =20, x 4 =22, x5 =28, x6 =25, x7 =26, x8 =27
population distribution which is estimated by , x9 =24, x10 =23) i.e. you got 25, 20, 20
the sample mean is denoted by x . The kilograms per meter square from plot 1,2,3
population (sample) mean is got by adding all and so on respectively. The mean maize yield
the values in a given population (sample) and is
dividing this sum by the population (sample)
size.
x = (25 + 20 + 20 + ... + 23) / 10 = 240 / 10 = 24
Take note .
§ If the population consist of observations x1 , Therefore, in this case we conclude that on
x 2 , x3 , …., x N , then the population mean average when fertilizer is used maize yield is
is given by
about 24 kilograms per meter square. This
i=N sample mean can then be compared with the
µ = ( x1 + x 2 + ... + x N ) / N = ∑ xi / N
i =1
mean from plots not treated with fertilizers.
§ If a sample consist of observations, x1 , x 2 ,
x3 , …., x n , then the sample mean is given The mean is supposed to represent the
by characteristics of a typical member of the
i=n population. However, the mean as a measure
x = ( x1 + x 2 + ... + x n ) / n = ∑ xi / n of centre of distribution is often affected by
i =1
i=n extreme values. If in your demonstration you
§ The ∑
1=1
means sum observations from
get from plot 1 a yield of 80 kilogram per
subject one (i =1) to n (sample size), it is a meter square instead of 25, will the mean still
summation sign
be a good measure of the centre of the
§ In most cases we are only able calculate the distribution? The value 80 kilogram per meter
sample mean since it may not be possible to
take observations on all members of the square is an extreme value or outlier in this
population.
9
case and it will pull the value of the men to the ascending order appear as 20, 20, 22, 23,
high side. In case you still insist on using the 24, 25, 26, 27, 28, 80.
mean as a measure of the centre of distribution 3. Cut off (remove) appropriate numbers of
then you need to remove those extreme values. observations from both sides of the
The mean which is calculated after extreme arranged data. In our example, we cut off
value(s) have been removed is called trimmed 20 from the lower end and 80 from the
mean. upper end of the arranged data. In that
way we have managed to remove the
Example 4.2 extreme value (80) but we also remove
Suppose the following observations one none extreme value and our new data

( x1 =80, x 2 =20, appear as 20, 22, 23, 24, 25, 26, 27, 28.
4. Calculate the mean of the remaining data
x3 =20, x 4 =22, x5 =28, x6 =25, x7 =26, x8 =27
(after removing the extreme value(s)). In
, x9 =24, x10 =23) our demonstrations example,
were obtained from your fertilizer
demonstration .i.e. from plot 1 we obtained 80
kilogram per square meters instead of 25. As xT = (20 + 22 + 23 + 24 + 25 + 26 + 27 + 28) / 8
we have noted before, the yield of 80 kilogram
= 187/8 = 23.75. This is a better measure of
per meter squares is an extreme value
centre compared to untrimmed mean of 18.5
(abnormal compared to the other nine plots).
Because of this extreme value, the new mean
is now 18.5 kilogram per meter square instead Activity
of the original 24 kilogram per meter square 1. Using the data in Example 3.2, calculate a

(from example 4.1) and this is no longer the 2% trimmed mean


2. State one disadvantage of using trimmed
centre of the distribution (in fact it is out side
mean
the range of the data). Now let us try to trim
off 10% of the sample and calculate 1%
4.1.2 Median
trimmed mean.
This is another measure of centre of
distribution of the population that is not
Steps in calculating the trimmed mean
affected by outliers (extreme observations).
1. Determine the number of observations to
The population median, denoted by Greek
be cut off from both ends of the
letter θ (theta), is the numerical value that
distribution of the data. For this example,
divides the population distribution in half. If a
we need to cut off 1% (p = 1%) of sample
sample consist of observations x1, x2, x3, …,
(1% of 10 = 1) .i.e. we need to cut off one
xn, then the sample median , denoted by M, is
(1) observation from either end of the
the middle observation if n is odd, or the
distribution.
average of the two middle observations if n is
2. Arrange the observations in ascending
even. In either case, the median is located at
order (from the smallest to largest
the position (n + 1)/2 is the ordered data set.
observation) or descending order (from
the largest to the smallest observation). In
our example, the observations arranged in
10
1. Position of the median = (12+1)/2 =
6.5. The median lies between the 6th
Steps in calculation of the median and 7th observations arranged in
1. Calculate the position of the median. The ascending or descending order.
position of the median is given by (n + 2. Rearrange the data in ascending
1)/2 where n is the sample size. order:
2. Arrange the observations in ascending or 4.2, 4.2, 4.3, 4.4, 4.4, 4.5, 4.7, 4.8, 4.8,
descending order. 4.8, 4.9, 4.9
3. Use the position of the median to 3. Identify the observations on the 6th
identify or calculate the median and 7th positions in the rearranged
data
Example 4.3 4. The median is M = (6th Observation +
Odd number of observations 7th Observation)/2 = (4.5 + 4.7)/2 =
4.6
A physical education teacher in your school
recorded the time in seconds taken to complete
a 10 meters race by 13 (sample size) students 4.1.3 Mode
from his class and asked you to get the median The mode is yet another numerical value that
time. The following is time in seconds run by is used to describe the centre of the distribution
the students. of the population that is also not affected by
extreme values. The mode is defined as the

4.4 4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 observation/measurement
4.4 4.2 that occurs most
often (with the highest frequency). You simply
1. Position of the median = (n + 1)/2 =
get the mode by counting how many times
(13 +1)/2 = 7 .i.e. the median lies in
each value occurs and the one with the highest
the 7th position
is the mode.
2. Rearranging the data in ascending
order:
Example 4.5
4.2, 4.2, 4.3, 4.4, 4.4, 4.4, 4.5, 4.7, 4.8,
Let us go back and revisit the data used in
4.8, 4.8, 4.9, 4.9
example 4.3 (time in seconds run by 13
3. The observation in the 7th position
students) and count the number of times each
from either side of arranged data is
value appears.
our median. Our median is 4.5
4.4 4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 4.2

Example 4.4

Even number of observations Data How many observations with this value in
value the data set
If instead of giving you the time run by 13
students in Example 4.3, you are given the 4.2 2
time for 12 students. Find out the median
4.3 1
4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 4.2 4.4 3
4.5 1
Steps 4.7 1

11
integer part” means that if (n+1)/2 has a
4.8 3
0.5 decimal, we just drop it off before
4.9 2
adding 1 and dividing by 2.

In this example, two values (4.4 and 4.8) have 2. Rearrange the observations in ascending

occurred more times than others. Our mode is order

therefore 4.4 and 4.8. Have heard about 3. The first quartile Q1 is found by
bimodal rainfall? We can have a uni-modal (1 counting the observations from the lower
mode), bimodal (2 modes) or multimodal end of the ordered data until we get the
(more than 2 modes) data. observation in the quartile position. The

third quartile, Q3 , is found by counting


Take note
the number of observations from the
§ The mode as a measure of centre is often a
meaningful and useful characteristics of discrete higher end of ordered data until we get
variables the observation in the quartile position.
§ For most continuous variables, the mode is often
not a meaningful and useful measure of the centre.
However, we can define the modal interval as the
histogram class interval with the highest Example 4.6
frequency. The mode is taken as the midpoint of
the modal interval.
Odd number of observations
Find the lower and upper quartiles of the data
used in example 4.3 (time in second run by 13
4.1.4 Quartile
students).
Just as the population median divides the
4.4 4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 4.2
population distribution in half, the population
quartiles divides it into quarters. The first
Steps
quartile, denoted by θ 1 , is the numerical value
Position of quartiles
that divides the lower half of the population
distribution in half. The third quartile, denoted Let Pm be the position of the median ((n+1)/2)

as θ 3 , is the value that divides the upper half of = (13+1)/2 = 7.0


Let PQ be the position of the quartile
the population distribution in half. The first

and third sample quartiles Q1 ,Q3 are similarly


PQ = ( Pm* + 1) / 2 = (7+1)/2 = 4
defined for the samples. The median is the
second quartile Q2 . Rearranging the data: 4.2, 4.2, 4.3, 4.4, 4.4,

Steps in calculating the quartiles 4.5, 4.7, 4.8, 4.8, 4.8, 4.9, 4.9

1. Calculate the position of the quartile. In


The lower quartile is the 4th observation from
order to get the position of the quartile,
the lower end of the arranged data (4.4)
you first need to determine the position
The upper quartile is the 4th observation from
of the median say Pm ((n+1)/2) after
the upper end of the arranged data (4.8)
which you we can take its integer part of
Pm, add 1 and divide by 2 to find the
Example 4.7
location of the quartiles. “Taking the
Even number of observations
12
Find the lower and upper quartiles of the data will learn how to describe our data using the
used in example 4.4 different measures of variability or dispersions.

4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 The
4.2 measures of variability that we are going
to look at in this lecture will include: range,
Pm = (12 +1)/2 = 6.5 variance and standard deviation

Pm* = 6 (Drop the 0.5 from Pm – taking


Illustration Example
integer part)
A teacher of statistics was one day shocked by
PQ = (6 + 1) / 2 = 3.5.
his five-year-old son who demanded to know
why he was shorter than most of his age mates.
Rearranging the data: 4.2, 4.2, 4.3, 4.4, 4.4, The son provided the father with height in
4.5, 4.7, 4.8, 4.8, 4.8, 4.9, 4.9 centimetres of five of his 5-year-old playmates
together with height of six girls in the same
rd
The lower quartile (Q1) lies between the 3 age group and begged his father to use those
and 4th observation from the lower end of the figures to explain this abnormal situation.
arranged data Height of boys: 68, 70, 69, 70, 71, 72
rd th
Q1 = (3 Observation + 4 Observation)/2 Height of girls: 65, 85, 90, 65, 55, 60
(counting starts from the lower end) The father told his son that although a typical
The upper quartile (Q3) lies between the 3rd 5-year-old is about 70cm, there is variation
th
and 4 observation from the upper end of the among different children because of several
arranged data factors. These factors according to the boy’s
Q3 = (3rd Observation + 4th Observation)/2 father include among others environment, food
(counting starts from the upper end) and the genes of the parents. By looking the at
figures presented by the son the father told his
son that the variation among the girls was
LECTURE FIVE: actually more than among the boys. This
NUMERICAL DESCRIPTION OF DATA – excited the boy more and asked his father how
MEASURES OF VARIABILITY he concluded that the variability was higher
among the girls. We will try to answer the
5.0 INTRODUCTION boy’s question by looking at the different ways
In the previous lecture, we learnt how to of quantifying or measuring variability in our
describe our data using measures of locations. data.
Measures of central location is important
because they describe a “typical” observation 5.1. Range
in the population. Not all observations are This is the simplest measure of variability you
typical; in fact, most deviate from the centre in can use. The range is the difference between
some way. The amount of deviation from the highest and the lowest values.
centre (from typical observation) is an Range = Maximum – Minimum.
important consideration when we are The range of height for boys is 72-68 = 4 and
investigating the properties of a data set. This for girls, 90-55 = 35. The boy’s father could
deviation from the centre is what we describe have used the range to conclude that there was
as variability or dispersions. In this lecture, we
13
more variability among the girls than among the distance between each observation and the
the boys. centre.

Advantages of using range as a measure of The population variance, σ


2
is the average
variability squared distance of all measurements from the
• Easy to calculate population mean. It is calculated with the
• Very simple to understand formula
Disadvantages of using a range as a
measure of variability N

• Uses only 2 values regardless of total 2 2 2 ∑ ( xi − µ ) 2


( x − µ ) + ( x − µ ) + .... + ( x − µ )
number of observations σ2 = 1 2 N
= i =1
N N
• Reveals nothing with respect of the way
in which the bulk of the observations As I stated before, each observation’s
are dispersed within the interval contribution to the variance depends on its

• Very sensitive extremes values deviation from the mean .i.e. ( x1 − µ )


(outliers)
determines the contribution of observation x1 .

To get the variance we just need to square


Both boys and girls have the average height of
these deviations, sum them up and divide by
70cm, but there is a difference between the
the population size (N).
heights of the two groups. The difference is
that the observations in the data set for girls
By now, you know that in most cases it is not
are more spread out from the mean more than
possible to calculate the population parameters
those in the data set for boys.
such as population variance so we have to use
2
the sample statistics. The sample variance, s ,

5.2 Variance and Standard Deviation is the average squared distance of the sample

The range is a useful measure of variability for values from the sample mean. It is calculated

a small data set. For a large data set, however, with the formula

a more sensitive measure of variability is


needed. Remember the range just uses only
Sum of
two observations (smallest and largest) so for a
Squares
n
large data set much information will be lost.
∑ (x i − x)2
The measure of variability that is use most
s2 = i =1

commonly is the variance and its square root- n −1 Degrees of


freedom
standard deviation.

5.2.1 Variance The expression in the numerator is referred to

You will realise shortly that in calculating the as sum of squares (SS), which measures the

variance all observations in a data set will total squared deviation of the whole data. The

contribute. The contribution of each expression in the denominator is referred to as

observation will depend on how different it is the degrees of freedom (df)

from the mean (deviation from the mean) .i.e.


14
n
Take Note
In practice Sums of squares is calculated as ∑ (x i − x ) 2 = 10; n - 1 = (6 – 1) = 5 (our
i =1
(∑ xi ) 2
sample size is 6)
SS = ∑ xi2 − i
and it can proved
n
that
10
(∑ xi ) 2 s2 = =2
n
5
∑ (x i − x ) 2 = ∑ xi2 − i

i =1 n We can conclude that the variability in the


height of the boys is 2 cm2. The variance for
the girls’ height is 200 cm2. Thus, this result
confirmed the conclusion drawn from using
Example 5.1
the range as a measure of variability.
Let us bring the data of the height of the 5-year
old boys and girls and see whether the father
of the young boy would have come to the same
Take Note
conclusion that there was more variation
o Since the variance is calculated from squared deviations,
among girls than boys it is always positive .i.e. the variance can never be
Height of boys: 68, 70, 69, 70, 71, 72 negative
o The unit of the variance is the square of the unit of the
Height of girls: 65, 85, 90, 65, 55, 60 variable under consideration. For example for height in
cm the unit of the variance is cm2
Let us start with the boys first. The sample o The variance is also affect by extreme values but not as
mean x for both boys and girls is 70 much as the range.

Observation
x2 ( x1 − x ) ( x1 − x ) 2
(x) By using an alternative formula
2 2
68 68 = (68 - 70) = - (-2) = 4
⎛ (∑ xi ) 2 ⎞
4624 2 ⎜ 2 i
⎟
⎜ ∑ xi − ⎟ for calculating the
70 4900 (70 - 70) = 0
⎜ n ⎟
0 ⎝ ⎠
69 4761 (69 - 70) = - 1
sums of squares, we can avoid the tedious
1
process of going through the construction of
70 4900 (70 - 70) = 0
0 the table used in example 5.1. Let us now try
71 5041 (71 - 70) = 1 to prove that the two formulae for calculating
1 sums of squares give us the same result. From
72 5184 (72 - 70) = 4
our calculation in Example 5.1, we now know
2
Total 420 29410 10 that the sum of squares is 10.

Applying our formula for calculating the

⎛ n ⎞ ( ∑ xi ) 2
⎜ ∑ ( xi − x ) 2 ⎟
(420) 2 176400
variance ( s = ⎜ i =1
2
⎜ n −1
⎟ )
⎟
∑ xi2 − n
i

n
= (29410 −
6
) = (29410 −
6
) = 10
⎜ ⎟ = ∑ (x i − x)2
⎝ ⎠ i =1

So at last, we have proved that the two


formulae bring the same result. I do advise you
the use the second formulae

15
the mean. Coefficient of variation measures
⎛ (∑ xi ) 2 ⎞
⎜ 2 i
⎟ the variability in the values of the data relative
⎜ ∑ xi − ⎟ for calculating the
⎜ n ⎟ to the mean.
⎝ ⎠
s
sums of squares. CV = x100
x

5.1.2 Standard Deviation • CV has no units, but is often


expressed as a percentage.
Standard deviation as a measure of variability
is derived directly from the variance by taking • Relative measure of variation that is
the positive square root of the variance. The independent of the unit of measure.
sample standard deviation is referred to as
standard error. 5.4 Inter-quartile Range ( IQR )

Inter-quartile range is a measure of variability


which is not affected by extreme values. It is
calculated by getting the difference between
the upper and the lower quartiles

IQR = Q3 − Q1
As a measure of variability, inter-quartile
range is useful for comparing variability of
Example 5.2 two or more data sets,
In example 5.1, we got the variance of the
heights of boys to 10cm2 and for girls to be Example 5.3
200cm2. Using the results obtained from example 4.6
Sample standard deviation (standard error) for Data: 4.2, 4.2, 4.3, 4.4, 4.4, 4.5, 4.7, 4.8, 4.8,
boys = 2 = 1.414 cm 4.8, 4.9, 4.9
Sample standard deviation (standard error) for Q1 = 4.4 ; Q3 = 4.8
girls = 20 0 = 14.14 cm IQR = Q3 − Q1 = 4.8 − 4.4 = 0.4

Take note LECTURE SIX:


GRAPHICAL METHODS FOR DATA
§ Unit of the standard deviation (error) is the same as
DESCRIPTION
that of the observation
§ The standard error is often referred to as a measure of
precision. In which case the smaller the value, the 6.0 INTRODUCTION
higher the precision.
In the last two lectures (4 and 5), we learnt
§ When we calculate the mean, it is good to attach a
measure precision to it. For example the mean height how to describe our data using numbers. You
of boys was 70 ± 3.16 cm now know how to describe a data set using
mean, range, variance, standard error and other
descriptive statistics. In this lecture, we are
going to learn how data are organised and
5.3 Coefficient of Variation (CV)
displayed in the form of tables and graphs for
Coefficient of variation as a measure of
illustrating their distribution .i.e. graphical
variability is derived from standard error and
description of data
16
6.1 DISPLAYING CATEGORICAL Response Frequency Relative
VARIABLES (DATA) Frequency

From section 2.5.2 you learnt that categorical Increased 305 305/500=0.61
variables are associated with some quality,
Decreased 25 0.05
characteristic or attribute which the variable
posses; for example, eye color, vehicle type, Remained 150 0.30

village of origin; etc. The following are the same

examples of graphical methods that can be No 20 0.04


used for categorical data: frequency tables, response
bar graphs and pie chart.
Total 500 1.00

6.1.1 Frequency Tables


Therefore, from this table you can now tell that
A table that lists the different categories of 305 farmers (majority of the farmers – 61%)
categorical variable and the corresponding said that their maize yield was increased while
frequencies with which they occur is called 25 of them stated that they have been
frequency table. experiencing yield reduction over the past two
Look at the following question that was years.
presented to 500 farmers in Gulu district about
yield of maize on their farms:
You will realize that sometimes this kind of
How has maize yield on your farm changed surveys can be done in more than one district
over the past 2 years? or location. For example, the same study
a) Increased b) Decreased c) Remained (asking farmers about their maize yield) could
the same have also been done in Lira, Arua, and
Mbarara and in of each those districts; the
number of farmers interviewed could as well
In this example, we can put the farmers into been different. If we want to compare two or
four categories basing on whether the yield on more studies with different sample sizes, then
their farms increased, decreased, remained the you need to calculate relative frequency. The
same or the farmer has given no response. relative frequencies in our frequency table
The result of the survey is presented in a table (table 6.1) were got by dividing the frequency
showing the different categories and how of each category by the sample (total count)
many farmers belong to each group). Table (e.g. 0.61 = 305/500). Relative frequency is

6.1: Frequency table showing farmers’ important because its value is independent of

responses on status of maize yield on their the size of the sample and this is necessary

farm when we compare two (or more) data sets of


different sizes.

17
6.1.2 Bar Graph Note that although the two bar graphs are
similar in most aspects, the vertical axes are
We all know that tables of numbers are
scaled differently. In the first graph, the
sometimes difficult to interpret especially for
vertical axis shows frequencies whereas in the
those with little or no education background or
second graph we have relative frequencies.
even highly educated people. In those cases, a
The relative frequencies is standardized with
picture might better illustrate the distribution
values ranging from 0 to 1 and this makes it a
of the data. Apart from being easy to interpret,
good tool for comparing two or more groups
graphs are also very good at showing trend and
with different sample sizes.
pattern in the data. For example, trend in
coffee production in Uganda over the last 10 Example 6.1
year can be picked easily from a graph than
Assuming the question about the yield of
from table of numbers.
maize on the farmers’ fields was also asked to
Let us try to convert the information on the selected farmers in two other districts
response of farmers to the maize yield question (Mbarara and Arua) and you are requested to
in Table 6.1 to a bar graph. As the name compare the response from the different
suggest, a bar graph consist a number of districts. Table 6.2 shows the result obtained
disjoint bars whose heights are determined by from the different districts.
the frequency or relative frequency of the
category represented by the bar. The bars can
be vertical or horizontal.

Figure 6.1: Bar graph (Frequency)

Figure 6.1: Bar graph (Relative


Frequency)

18
Gulu

Table 6.2: Frequency table showing farmers Non response

responses on status of maize yield on their


farm from three districts (Gulu, Mbarara and
Constant
Arua)

Response Frequency Relative Frequency


Frequency (Mbarara)
(Gulu)
(Gulu)

Increased 305 0.61 250

Decreased 25 0.05 10 Deccreased


Increased
Remained 150 0.30 38
the
constant
Figure 6.3: Pie Chart showing Responses of
No 20 0.04 02 Farmers from Gulu District
response The pie chart just like relative frequency bar
Total 500 1.00 300 graph can also be used to compare groups. Just
remember that the comparison should be based
on relative frequencies instead frequencies.
We can see from the above that relative
frequency makes comparison very easy since
all values like between 0 and 1 i.e. the fact that 6.2 Displaying Quantitative Data
sample sizes used in the three district does not From section 2.5.2, you learnt that quantitative
affect our comparison. Mbarara has a higher variables are characteristics that can be
proportion of farmers who experience yield quantified or measured using some defined
increase compared to Gulu. scale e.g. maize harvested on a given farm.

6.1.3 Pie Chart


6.2.1 Stem and leaf plot
The pie chart is another useful way to exhibit The stem and leaf plot is a quick and easy way
categorical data. A circle or a pie is divided to display the distribution of continuous data.
into pieces corresponding to categories of the It is extremely useful for arranging the
variable so that size (angle) of the slice observations from the smallest to the largest so
proportional to the relative frequency of the that specific locations within the data set can
category. be found.

Example 6.2

An agriculture teacher gave a test to his class


of 19 students and decided to use stem and leaf
plot to summary the result. The following are
marks scored by his student in the test.

84 59 82 78 96 44 76 85 66 77 91 62 54 72 65
84 38 76 70

19
By observation, we see that the marks are two-
digit numbers that range from the 30s to 90s.
Example 6.3:
To construct the display, we divide each
3 8
observation into a stem and leaf. In this
4 4
example the digit in the tens place of the
5 4 9
number become the stem, and the digits in the
2 5 6
units place becomes the leaf of the stem and
7 0 2 4 6 6 7 8
leaf plot. A vertical line is drawn to separate
8 2 4 4 5
the stems from the leaves.
9 1 6
3 3 8
4 4 4
Using the data of marks from example 6.3, we
5 5 4 9
construct a histogram using each stem as a class.
6 2 5 6
From our stem and leaf plot we see that in the
7 7 0 2 4 6 6 7 8
class limit for stem value 3 are 30 and 39, for
8 4 8 2 4 4 5
stem value 4 are 40 and 49, and so on. Once the
9 9 1 6
class limits are determined, that data can be
(a) (b)
formulated in a grouped frequency table
(a) The first observation (84) plotted
(grouped in the sense that values are grouped
(b) complete and ordered stem and leaf plot into varies classes) see table below

The stem and leaf plot leaves the data intact Table 6.3: Grouped frequency table
for future calculations. In our example above
Class Class Frequency Relative
we can see that most students scored 70 and
Limits Boundaries Frequency
above in this particular test (see the shaded
30 - 29.5 – 39.5 1 1/19
part).
39
40 - 39.5 – 49.5 1 1/19
49
6.2.2 Histograms 50 - 49.5 – 59.5 2 2/19
59
The histogram is a graphical mean of
60 - 59.5 – 69.5 3 3/19
illustrating the distribution of continuous 69
quantitative data. The measurements are 70 - 69.5 - 79.5 6 6/19
79
grouped into classes defines as an interval of
80 - 79.5 – 89.5 4 4/19
values. The class limits are the smallest and 89
largest possible values in the interval. Once the 90 - 89.5 – 99.5 2 2/19
class limits are determined, the data can be 99

formulated into grouped frequency table. It


easy to transform our stem and leaf plot into a The class boundaries are obtained by lowering
histogram in which each stem will constitute a the lower class limits by 0.5 and raising the
class upper limits by 0.5. Thus, the class limits 30
and 39 become class boundaries 29.5 and 39.5
respectively. The class boundaries are given so
that the classes are continuous; that is the first
20
class runs from 29.5 to 39.5, and immediately
the second class picks up at 39.5 and runs to
Example 6.4
49.5, and so on.
A group of agriculture students bought 100
You can now construct the histogram using the
day-old chicks from UgaChick and weighed
grouped frequency table. The class boundaries
them. In order to report the quality of chicks
are scaled off on the horizontal axis. Bars are
they bought to their classmates, they decided
constructed over each class boundary so that
to use the histogram for this purpose. Below
the height of each bar is the frequency or
are the weights of the 100 chicks arranged in
relative frequency of the class which is marked
ascending order. They decided to use 14
off on the vertical axis.
classes. Class interval = (4.9 – 3.6)/14 =
0.0929 ~ 0.1

The following are the general steps that you


Table 6.4: Weight gain for chicks (grams)
should follow in constructing a histogram
ordered (see a separate paper)
Step 1 - Array/order the numerical values
form low to high Table 6.5: Grouped frequency table (see a
Step 2 - Divide the overall range (maximum - separate paper)
minimum) of the values in our data
set into a number of classes and
Histogram
count the number of observations that 14
fall into each of the class - class
12
frequency (generally you should have
10
5-20 classes).
Frequency

Step 3 - Construct frequency histogram or 8

relative frequency histogram 6

4
Helpful hints:
2
• After dividing the range by the
0
desired number of classes to obtain
5

5
3.6

3.7

3.8

3.9

4.0

4.1

4.2

4.3

4.4

4.5

4.6

4.7

4.8

4.9

class interval - round off the result to


5-

5-

5-

5-

5-

5-

5-

5-

5-

5-

5-

5-

5-

5-
3.5

3.6

3.7

3.8

3.9

4.0

4.1

4.2

4.3

4.4

4.5

4.6

4.7

4.8

a convenient unit. For example if the Class intervals


maximum observation is 4.9 and
minimum observation is 3.6 and if we LECTURE SEVEN:
decided to have 10 classes then our INTRODUCTION TO PROBABILITY
class interval is (4.9-3.6)/10 = 0.13 ~
0.1 In lectures 3 to 6 you learnt how to describe
• First class should contain the smallest data or information collected using numerical
measurement and the last class should and graphical methods. So now, you are able
contain the largest measurement. to summarize a large data set into a few
• No measurement should fall in numbers (e.g. mean, median, mode, etc) or
between classes graphs (e.g. bar graphs, pie, histogram, etc) but

21
that is not all about statistics. You learnt from
lecture 1 that statistics has two major
7.1.2 Outcome
fields/branches i.e. descriptive and inferential. An outcome is the result of a single trial of a
In inferential statistics, we are concern with probability experiment. In rolling a die once
how to use information from a sample to draw for example, the outcome can be 1, 2, 3, 4, 5 or
generalization about the population. The 6 i.e. you can get one of the numbers 1 to 6
knowledge of probability will help us to showing up (see Figure 7.1).
understand the concepts of statistical inference
that we are going to cover in the next part of
the course. In this lecture, you will be
introduced to the basic concepts of probability Figure 7.1: Faces of a die-Any one of the 6
numbers can show up when you roll a die
including basic definitions.

7.1.3 Sample Space (S)


The sample space denoted by S is a set of all
7.1 Definitions of terms
possible outcomes of a probability experiment.
Intuitively we think of probability as a
Example of sample spaces of different
numerical value that is associated with some
probability experiments
outcomes and indicates how likely it is that the
a) If the experiment is rolling a die, then
outcome will occur. You heard people
the sample space is
swearing that they are 100% sure that
S = {1, 2, 3, 4, 5, 6}
something will happen. The value of
b) If the experiment is tossing a coin,
probability lies between 0 and 1 but can also
the sample space is
be expressed in percentage. The probability of
S = {head, tail}
0 means it is impossible for something to
c) If the experiment is weighing a person,
happen for example the probability that the sun
the sample space (assuming no one weighs
will not rise tomorrow is 0. The probability 1
more than 250kg)
means we are very certain that something will
S = {x| x is a real number between 0
happened for example the probability that you
and 250}
will die one day is 1 since all humans are
d) If the experiment is taking a malaria
mortal.
test, the sample space, S= {positive,
negative} .i.e., there is only two possible
7.1.1 Probability Experiment
outcomes
This is a process which leads to well-defined
results call outcomes. It is a process of making 7.1.4 Event (E)
An event is one or more outcomes of a
observation or taking measurement. Examples
probability experiment that is of interest to the
of probability experiment including; rolling a
researcher. It is any subset of the sample space.
die, tossing a coin, taking a blood test for
An event is said to have occurred if any one of
malaria, measuring the height of students,
the elements is the outcome when an
measuring crop yield from plots treated
experiment is conducted. When rolling a six-
different fertilizers.
sided die for example you may be interested in
the number 6 or an odd or an even number

22
showing up. In the above case your event are a die at the same time, the outcome of the coin
{6}, {1, 3, 5} and {2, 4, 6} for a 6, an odd or toss experiment and the rolling of a die do not
even number respectively. affect each other.

Take Note
7.1.9 Dependent Events
An event may have more than one outcome.
Two events are dependent if the first event
For example when you are interested in an odd
affects the outcome or occurrence of the
number from rolling a die as your event, then
second event in a way that the probability is
your event has more than one outcome .i.e. if
any of the numbers 1, 3 or 5 appears it will still changed.
be an odd number and your interest will be met.

7. 2 SAMPLE SPACES (ILLUSTRATION)

A sample space is the set of all possible


7.1.5 Equally Likely Events outcomes. From the same probability
Any two or more events are said to be equally
experiment, you can define the sample space in
likely if they have the same chances
more than way. However, some sample spaces
(probability) of occurring. For example if you
are better than others.
toss a fair/balance coin, both the tail and head
sides of the coin, have the same chances Consider the experiment of flipping two coins.
(probability) of occurring. It is possible to get 0 heads, 1 head, or 2 heads.
Thus, the sample space could be {0, 1, 2}.
7.1.6 Complement of an Event Another way to look at it is flip {HH, HT, TH,
All the events in the sample space except the TT}. The second way is better because each
given events. In our example of rolling a die, if event is as equally likely to occur as any other.
the event of A = {1, 3, 5} then the complement
of event A denoted by A’ = {2, 4, 6}.
Take Note

When writing the sample space, it is highly desirable


7.1.7 Mutually Exclusive Events
to have events which are equally likely. This makes
Two events are said to be mutually exclusive if the calculation of probability simpler
they cannot occur at the same .i.e. when one
event occurs, the other one cannot occur at the
same time. For example when you roll a die Another example is rolling two dice and

and the number 1 comes up, there is no way summing up the faces that show up on the 2

that on the same die the number 2, 3, 4, 5 or 6 dice. The sums are {2, 3, 4, 5, 6, 7, 8, 9, 10,

will show thus these events mutually 11, 12}. However, each of these are not

exclusive. Mutually exclusive events are also equally likely. The only way to get a sum 2 is

referred to as disjoint events. to roll a 1 on both dice, but you can get a sum
of 4 by rolling a 1-3, 2-2, or 3-1. Figure 7.2

7.1.8 Independent Events and Table 7.1 illustrate a better sample space

Two events are independent if the occurrence for the sums obtain when rolling two dice.

of one does not affect the probability/chance of When two dices are rolled there are 36 unique

the other occurring. If you toss a coin and roll ways in which the dices can land (see Figure
7.2 and Table 7.1)
23
If you roll a die, the sample space has six
outcomes S = {1, 2, 3, 4, 5, 6}. Let B be the
event that an even number show up. In this
case, our event has three outcomes {2, 4, 6}
.i.e. if any of the three numbers 2, 4 or 6 show
up we still got an even number. The
probability that an even will show up when we
roll a die is 3/6 or ½.
Figure 7.2: Sample space of rolling 2 dices
(see the sums in Table 7.1)
Example 7.3
Consider the example of rolling two dice used
Table 7.1: Sums of two dices
to illustrate the concept of sample space in
Second Die section 7.2. The table below show the possible

First 1 2 3 4 5 6 outcomes you can get from adding the number


Die
on two dice rolled together and associated
1 2 3 4 5 6 7 probabilities
2 3 4 5 6 7 8

3 4 5 6 7 8 9 Table 7.2: Probabilities associated with


4 5 6 7 8 9 10 different events from sum of two dices
5 6 7 8 9 10 11 Sum (Die Number of Probability = (No.
1 + Die possible ways of of outcomes in an
6 7 8 9 10 11 12
2) getting an event event)/ (Total No.
(specified sum of events in

7.3 Classical verses Empirical Probability from two dice) sample space)
2 1 1/36
7.3.1 Classical Probability
3 2 2/36
Classical probability uses the sample space to
4 3 3/36
determine the numerical probability that an 5 4 4/36
event will happen. Classical probability is also 6 5 5/36
called theoretical probability. In this case, to 7 6 6/36
get the probability of a given event you need to 8 5 5/36
divide the number of outcomes that are in your 9 4 4/36

event with the total number outcomes in the 10 3 3/36

sample space. 11 2 2/36


12 1 1/36

Example 7.1
If you toss a coin the sample space is S =
{Head, Tail} or {H, T}. Let A be the event that
the head (H) shows up. In this case, our event
has only one outcome .i.e. H and our sample
space has two possible outcomes H and T. The
probability that H will show up is ½.
Example 7.2
24
Take Note
For classical probability, the probability of an
LECTURE EIGHT:
event occurring is the number outcomes in
LAWS OF PROBABILITY
the event (n(E)) divided by the number of
outcome in the sample space (n(S)). INTRODUCTION
P(E) = n(E) / n(S)
However, this is only true when the In lecture seven, we looked at the various

outcomes are equally likely. terms and concepts used in probability. Now
you can talk comfortably about probability
terms without any fear. In this lecture, you
7.3.2 Empirical Probability will learn the different laws/rules of
Unlike classical probability, empirical probability that will help your further
probability is based on observations. Empirical understanding of probability.
probability is the relative frequency of a
frequency distribution based upon 1. All probabilities are between 0 and 1
observations. Do you remember how we inclusive. In mathematical term, this is
calculated the relative frequency in lecture 4? represented as shown below .i.e. the
If for example you roll a die 120 times and you probability that a given event E will
realise that the number 1 occurs 20 times then occur, lies between 0 and 1.
the empirical probability of 1 is 20/120. 0 ≤ P( E ) ≤ 1
2. The sum of all the probabilities in the
P(E) = (Number of times the event occurs
sample space is 1. Let us recall that when
(frequency)) / (Total number of trials) = f/n
you roll a die the sample space S = {1, 2,
3, 4, 5, 6} what this law tells us is that
Example 7.4
when you roll a die you are certain that
In order to determine the probability of having
one of those numbers will show up.
malaria parasites among students in your
Mathematically we can represent the sum
school, you decided to sample 50 students and
of the probabilities of the sample space as
took them for a laboratory test? If it turns out
that 15 students tested positive for the parasite,
what will you conclude about the probability P( S ) = P(1) + P(2) + P(3) + P(4) + P(5) + P(6) = 1
of occurrence of malaria parasites at your
school? Note that we can use the empirical 3. The probability of an event which cannot
method to determine this probability since this occur is 0.
is based on observations. We just need to the 4. The probability of any event which is not
divide the frequency of malaria parasites (15) in the sample space is 0. For example, if
with the total number of students tested (50). you roll a die the probability that a number
P(Malaria parasites) = (Number who tested 7 will show up is 0 since the number 7 is
positive)/(Total number tested) = 15/50 not in our sample space .i.e. a die only has
six sides numbered from 1 to 6.

25
5. The probability of an event which must occur. First ask yourself, are these two events
occur is 1. For example, the probability (A and B) mutually exclusive? Yes.
that the sun will rise tomorrow is 1 since it
Thus
must happen.

P(A) = 1/6, P(B) = 1/6, A and B are disjoint or


6. The probability of an event not occurring
mutually exclusive
is one minus the probability of it
occurring. If you toss a coin, the
P(A or B) = P(A ∪ B) = P( A) + P( B) = 1 / 6 + 1 / 6 = 2 / 6
probability that a head (H) will not show Example 8.2
up is one minus probability of head
In the example of rolling 2 dice in section 7.2
showing up.
we realised that in summing the values
showing up on the 2 dice our sample space S
P( E ' ) = 1 − P( E )
={2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12}, what is the
8.2 "OR" or Unions probability of a sum that exceeds 8. In this
Remember that two events are mutually case our event E = {9, 10, 11, 12} and since
exclusive if they cannot occur at the same time these event are mutually exclusive, the
and that another word for mutually exclusive is probability of event E is given by the sum of
disjoint. the probability of the occurrence of each
element
If two events are disjoint, then the probability
of them both occurring at the same time is 0.
P ( sum > 8) = P (9) + P (10) + P (11) + P (12) = 4 / 36 + 3 / 36 + 2 / 36 + 1 / 36 = 10 / 36
Let A and B be two mutually exclusive events,
then the probability that both A and B will
General Addition Rule (Additive law of
occur at the same time will be mathematically mutually and non mutually exclusive
represented as below events)
Non-Mutually Exclusive Events
P(A and B) = P(A ∩ B) = 0 In events which aren't mutually exclusive,
there is some overlap .i.e. there is a possibility
Specific Addition Rule (Additive law of
that both events can occur at the same time
mutually exclusive events)
(intersection). When P(A) and P(B) are added,
If two events are mutually exclusive, then the the probability of the intersection is added
probability of either occurring is the sum of the twice (see the figure below).
probabilities of each occurring.

P(A or B) = P(A ∪ B) = P( A) + P( B) To compensate for that double addition, the


intersection needs to be subtracted.
Example 8.1
General Addition Rule
Let us recall that when you roll a die, the
This rule is valid for both mutually exclusive
probability of any number showing up is 1/6.
and non-mutually exclusive events
Let A be the event that when you roll a die a 1
will show up and B be event that a 4 will show
up. What is the probability that A or B will P(A or B) = P(A) + P(B) - P(A and B)
26
A B 8.3 "AND" or Intersections

Let us now recall that two events are said to be


independent if the occurrence of one event
does not change the probability of the other
occurring. For example, when you roll a die
and toss a coin at the same time the result of
A∩ B the die rolling does not affect the outcome of
Example 8.4 the coin flipping or tossing. For sure, rolling a
2 on a die does not affect the probability of
Let us use the experiment of rolling the die to
flipping the head on the coin.
illustrate this rule. Let A be the event that an
even number will show up and B be the event If you are dealing with independent events,
that a 2 will show up when a die is rolled. then the probability of them both occurring is
What is the probability that an even number the product of the probabilities of each
(event A) or a 2 (event B) will show up? occurring. This is referred multiplicative law
Event A comprise of the following outcomes of probability as is given below. The
2, 4 and 6 .i.e. A = {2, 4, 6}; event B has only probability that both event A and B will occur
one outcome B = {2}. is also referred to as the probability of A
The probability that event A will occur is got intersection B ( A ∩ B ).
by dividing the number of outcomes in event A
(3) by the total number of outcomes in the
Multiplicative Law for Independent Events
sample space (6).
Let A and B be any two independent events
P(A) = 3/6
then
The probability that event B will occur is got
by dividing the number of outcomes in event B P(A and B) = P( A ∩ B ) = P(A) * P(B)
(1) by the total number of outcomes in the
sample space (6) Where P(A and B) is the probability of both A
and B occurring at the same time
P(B) = 1/6 P(A) is the probability of event A
occurring
We now need to calculate the probability that P(B) is the probability of event B
both A and B will occur together. For both occurring
events A and B to occur, a 2 must show up
(both events have 2 in common). Example 8.5

P(A and B) = P(2) = 1/6 In a probability experiment involving rolling a


die and flipping a coin at the same time we can
Therefore the probability that event A or B
occurs is calculate the probability of a 2 showing up on
a die and a head showing up on the coin. Let A
P(A or B) = P(A) + P(B)-P(A and B)
be the event that a 2 will show up when a die is
= (3/6 + 1/6) – 1/6
rolled and B be the event that a head will show
= 3/6
up when a coin is flipped.

27
Using the classical probability we know that P(A|B) . This is the conditional probability of
the probability of event A occurring is P(A) = event A.
1/6 and probability that event B occurring is
P(B) = ½ P(A and B)
P(A/B) =
P( B)
Thus P(A and B) = (1/6)*(1/2) = 1/12
From the above formula, we can see that
LECTURE NINE:
P(A and B) = P(B) * P(A|B)
CONDITIONAL, JOINT AND
MARGINAL PROBABILITY This looks like the multiplicative law for
independent events. In fact, it is the general
multiplicative law for both independent and
INTRODUCTION
dependent events.
From the last two lectures we have learnt
terms, concepts and rules at are used in
Example 9.1
understanding and calculations of probability
associated with different types of events. In Suppose that at St. Charles Lwanga College,

lecture 8, we studied rules that are used in the 40% of students take Biology, 25% take

calculation of probabilities of mutually Agriculture and 12% take both. If a student is

exclusive and independent events. In this randomly selected from a Biology class, what

lecture, we will learn conditional events and is the probability that he or she is also taking

probability associated with those events. We Agriculture?

will also look at how to calculate the


Before you start calculating the probability, we
probabilities associated events that occur
need to ask ourselves whether this is a
jointly.
conditional probability problem.

“This is a conditional probability problem


9.1 CONDITIONAL PROBABILITY because we are given that the student is
taking Biology and we are asked to find the
Not all events that you will encounter in
probability that he or she is taking
probability experiments are independent .i.e.
Agriculture”.
the occurrence of some event dependents on
other events. For example, the chances of
Let B be the event that the student is taking
getting a toad moving are higher at night
Biology, and let A be the event that the student
compared to daytime. If A is the event that you
is taking Agriculture.
will find a toad moving and event B defines
the time you set off to look for the toad, then
P(B)=0.4 P(A)=0.25 P(B and A)=0.12
Events A and B are dependent. In this
example we see that the probability of finding
a toad moving (event A) will depend on the P(A and B) 0.12
P(A/B)= = = 0.3
time of the day one go out to look of them. P( B) 0.4

The probability of event A occurring given


Thus, if we know that a student is taking
that event B has already occurred is read "the
Biology, there is a 30% chance that he or she
probability of A given B" and is written:
is also taking Agriculture. If we do not know
28
that the student is taking Biology, then the consider (sex, smoking habit) or (smoking
probability that he or she is taking Agriculture habit, drinking habit) as a single outcome of
is 25%. the probability experiment. We might study
the height H and weight W of some students,
giving rise to (h, w) outcome. Finally, we
Since we are given that event B has occurred might observe the total rainfall R and average
(we are already in a biology class), we have a temperature T at a certain locality during a
reduced sample space. Instead of the entire specified month giving rise to outcome (r, t).
sample space S (all the students), we now have Instead of calculating the probability that the
a sample space of B (only students in biology teacher you pick smokes, we might be
class) since we know B has occurred. So the interested in the probability that you will pick
old rule about being the number in the event female teacher who smokes. These are what
divided by the number in the sample space still we refer to joint events.
applies. It is the number in A and B (must be
in B since B has occurred) divided by the Example 9.2
number in B. If you then divided numerator
The question, "Do you smoke?" was asked of
and denominator of the right hand side by the
100 people. Results are shown in the table.
number in the sample space S, then you have
the probability of A and B divided by the . Yes No Total
probability of A.
Male 19 41 60

Female 12 28 40

Total 31 69 100
Take Note
The following four statements are equivalent 1. What is the probability of a randomly
selected individual being a male who
1. A and B are independent events
smokes? This is just a joint probability
2. P(A and B) = P(A) * P(B)
since we are looking at the sex and
3. P(A|B) = P(A) smoking habits if the individuals jointly.
4. P(B|A) = P(B) The number of "Male and Smoke" divided
by the total = 19/100 = 0.19
The last two are because if two events are
independent, the occurrence of one does not
change the probability of the occurrence of the 2. What is the probability of a randomly
other. This means that the probability of B selected individual being a male? This is
occurring, whether A has happened or not, is
simply the probability of B occurring. the total for male divided by the total =
60/100 = 0.60. Since no mention is made
of smoking or not smoking, it includes all
9.2 Joint Probability the cases. In this case we are only
In many situations, you may be interested in interested in the probability of picking a
observing more than one characteristics of the male and this is referred to as marginal
subject or study unit. For example, the probability
smoking and drinking habit of a group of
teachers may be of interest and we would

29
3. What is the probability that a randomly you are going to be introduced to the concept
selected individual smoking? Again, since of probability distribution and we learn about a
no mention is made of gender, this is a binomial distribution.
marginal probability, the total who smoke
10.1 Random Variable
divided by the total = 31/100 = 0.31.

A random variable is a variable whose value is


4. What is the probability of a randomly
determined by chance. A random variable is a
selected male is a smoker? This time,
rule that represents the possible numerical
you're told that you have a male - think of
values associated with the outcome of an
conditional probability. What is the
experiment. A random variable is denoted by
probability that the male smokes? Well,
the capital letter X or Y but its value is denoted
19 males smoke out of 60 males, so 19/60
by small letters x or y.
= 0.31666...

In the example of rolling a pair of dice, let x


represent the sum of the two faces that show
LECTURE TEN: when we roll a pair of dice. The possible
RANDOM VARIABLES AND values of x are
PROBABILITY DISTRIBUTIONS
Sx = {2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12}

INTRODUCTION
A random variable can be discrete or
In previous lecture (lecture 9) you learnt how continuous. The random variable whose values
to deal with dependent and joint events and the consist of numbers such as {0, 1, 2, 3, 4} or at
probabilities associated with those types of most a countable number of values such as
events. From lecture 7, you learnt that
probability experiment results into outcomes. {1, 2, 3, 4, . . .} is referred to as discrete

You know that for example that when you roll random variables. Continuous random

a die then one of the faces will have to show variables are those that can assume all the

up as an outcome or if you spray an insect with values in an interval of the number line.

insecticide, then the insect will either die or


live as the outcome of that action. When a
probability experiment is performed several 10.2 The Probability Distribution Function
times, certain pattern in the outcome will begin of a Random variable
to emerge. For example if you measure the A table or function that lists all possible values
height of many schoolchildren of 7 – 9 years of a discrete (continuous) random variable and
of age, you will probably realize that most of their associated probabilities is called the
them will belong to some range with just a few probability distribution function of the
being either too short or too tall. The pattern random variable. To make inference about the
that will emerge after repetition of probability population based on a sample we need to know
experiment can be described in terms of the probability of observing a particular
probability of occurrence of the different outcome and this can only be got when we
elements in the sample space. In this lecture, know the probability distribution.
30
The followings are example of Bernoulli
Example 10.1
population:
Consider families with three children. Let the
• The toss of a coin results in a
variable w be the number of girls in the family.
Find the probability distribution for w and Bernoulli population because each
toss results in a head (success) or
graph the distribution. Let g represent a girl
tail(failure)
child and b represent a boy child, thus a family
with all the children girls would have (g, g, g). • On a given day of the week Gulu
town either has electricity (success) or

The possibilities for three children .i.e. the not (failure)

sample space is • When treated for Malaria either you


are cured (success) or not (failure)

S = {(g, g, g), (g, g, b), (g, b, g), (b, g, g), (b, b, • A none twin birth results in either a

g), (b, g, b), (g, b, b), (b, b, b)} boy (success) or a girl baby (failure)

Activity
We see that there are eight possibilities, and if o List five other examples of a Bernoulli population

we assume that a boy is just as likely as a girl, and in each case define what you consider to be
success and failure
all of the eight possibilities are equally likely.
Of the eight only one (b, b, b) correspond to
zero girls, in which w = 0 and the probability
When you tossed a coin once, you have
of zero is 1/8. Three of the eight possibilities
performed a Bernoulli trial but when you
correspond to w = 1(one girl), and therefore
repeat the process more than once then it
the probability is 3/8. Continuing in this
ceases to be a Bernoulli. A sequence Bernoulli
fashion we get the following completed
trials is called a binomial experiment
probability distribution of w.

w 0 1 2 3 10.3.2 Binomial Experiment


P(w) 1/8 3/8 3/8 1/8
For a probability experiment to be called a
binomial experiment, it has to satisfy these
four conditions
10.3 THE BINOMIAL DISTRIBUTION

Before we look at a binomial distribution, we • There has to be a fixed number of


need to define a unique type of population trials
called Bernoulli population.
• Each trial is independent of the others

10.3.1 Bernoulli Population and Trial • There are only two outcomes (success
A Bernoulli population is a population in or failures)
which each element is one of the two
possibilities. The two possibilities are usually • The probability of each outcome

designated as success and failure. A Bernoulli remains constant from trial to trial.

trial is observing one element in a Bernoulli The probability of success is p and

population. probability of failure is 1-p in some


books is referred to as q.
31
These can be summarized as follows: An What is the probability of finding exactly two
experiment that consists of n repeated people with malaria parasites when you test
independent Bernoulli trials in which the three people (Okello, Mukasa and Muhwezi)?
probability of success in each trial is p and the
There are five things you need to do to work a
probability of failure 1-p.
binomial problem
The fact that each trial is independent actually
1. Define Success first. Success must be for
means that the probabilities remain constant.
a single trial. Success = "Testing positive
Examples of binomial experiments for malaria parasite on a single test"
• Tossing a coin 20 times to see how many
2. Define the probability of success (p): In
tails occur.
our example we assume p = 1/6
• Asking 200 people if they listen to the
3. Find the probability of failure (q ): In our
local FM radio everyday
example q = 1-p = 5/6
• Rolling a die to see if a 5 appears.
4. Define the number of trials (n): In our

• Spraying 20 insects with insecticide to see example n = 3 (we are testing three

how many will die people)

Examples which aren't binomial 5. Define the number of successes out of


experiments those trials (x): In our example x = 2

• Rolling a die until a 6 appears (not a fixed (two people testing positive)

number of trials)
Anytime a person test positive, it is a success

• Asking 20 people how old they are (not (denoted S) and anytime negative appears, it is

two outcomes) a failure (denoted F). The number of ways you


can get exactly 2 successes in 3 trials (testing)
• Drawing 5 cards from a deck for a poker are given below. In each of the scenario in the
hand (done without replacement, so not Table 10.1, two successes occur .i.e. 2 people
independent) test positive

10.3.3 Binomial Probability Function


When you plant 20 seeds for example, you
may want to know the probability that 10 seeds
will germinate given the germination rate
specified by the seed company. This
probability can be calculated using the
probability function which we are going to
introduce in this section.

Example 10.2

32
Table 10.1: Possible Scenario on testing Further, note that there are three ways this can
three people for Malaria occur. This is the number of ways 2 successes
can be occur in 3 trials without repetition and
Possible test People tested for malaria
order not being important, or a combination of
results
Okello Mukasa Muhwezi 3 things, 2 at a time.

Scenario 1 Negative Positive Positive


(F) (S) (S) Take Note

Scenario 2 Positive Negative Positive o The number of combinations of k objects taken from n
objects is given by
(S) (F) (S)
⎛ n ⎞ n!
Scenario 3 Positive Positive Negative ⎜⎜ ⎟⎟ = Where n! is read “n factorial”
(S) (S) (F)
⎝ k ⎠ k!(n − k )!
and is given by

n!=n(n-1)(n-2)(n-3). . . (2)(1)
In the first scenario, Okello tested negative
o When calculating n!, we will assume that
(failure) for malaria but Mukasa and Muhwezi
a) 1! = 1 and
tested positive (successes) and the probability
associated with this given as b) 0! =1

P(Okello Negative (F))*P(Mukasa Positive ⎛ 3 ⎞ 3!


(S))*P(Muhwezi Positive(S)) =
Thus ⎜⎜ ⎟⎟ = =3
⎝ 2 ⎠ 2!(3 − 2)!
P(F)*P(S)*P(S)
o In general, the probability of getting exactly k success
= 5/6 * 1/6 * 1/6 = (1/6)2 * (5/6) in n trials
Note that because the trials are independent,
⎛ n ⎞ k
the probability of the event (test results for the P(X=k) = ⎜⎜ ⎟⎟ p (1 − p) n −k
three people) is the product of each probability ⎝ k ⎠
of each outcome (each person’s trial result)
⎛ n ⎞
You can use a calculator to solve ⎜
o
⎜ k ⎟⎟ . Please try to
⎝ ⎠
The probabilities associated with the three read the manual of your calculator BUT on most
scenarios in Table 10.1 is presented below calculators use (n Cr k)

1 FSS 5/6 * 1/6 * 1/6 = (1/6)2 * (5/6)


2 SFS 1/6 * 5/6 * 1/6 = (1/6)2 * (5/6)
The probability of getting 2 success out of 3
3 SSF 1/6 * 1/6 * 5/6 = (1/6)2 * (5/6) trials is

Notice that each of the 3 probabilities are P(X=2)=


2 1
exactly the same: (1/6) 2 * (5/6). ⎛ 3 ⎞ 2 3− 2 ⎛ 3 ⎞⎛ 1 ⎞ ⎛ 5 ⎞
⎜⎜ ⎟⎟ p (1 − p) = ⎜⎜ ⎟⎟⎜ ⎟ ⎜ ⎟ = 0.06944
Also, note that the 1/6 is the probability of ⎝ 2 ⎠ ⎝ 2 ⎠⎝ 6 ⎠ ⎝ 6 ⎠
success and you needed 2 successes. The 5/6 is
Example 10.4
the probability of failure, and if 2 of the 3 trials
were success, then 1 of the 3 must be failure. Using information from Example 10.3 above
Note that 2 is the value of x and 1 is the value suppose 6 people were tested what is the
of n-x. probability that exactly 2 people will test
positive?

33
In this example, the number of trials n = 6 (6 P(X< 2) = P(X =0) + P(X=1)
people are being tested), the number successes
x =2 and using the general probability of ⎛ 6 ⎞ 0 ⎛ 6 ⎞ 1
= ⎜⎜ ⎟⎟ * (0.8) * (0.2) 6 + ⎜⎜ ⎟⎟ * (0.8) * (0.2) 5
formula ⎝ 0 ⎠
= 0.0015 ⎝1 ⎠
c) More than 2 will germinate
⎛ n ⎞ x
P(X= x) = ⎜⎜ ⎟⎟ p (1 − p) n − x = In this case, there are four possible values
x
⎝ ⎠ that x can take .i.e. x can be 3, 4, 5 and 6
2 4 2 since
4 each of them is greater than 2.
⎛ 6 ⎞ 2 3− 2 ⎛ 6 ⎞⎛ 1 ⎞ ⎛ 5 ⎞ ⎛ 1 ⎞ ⎛ 5 ⎞
⎜⎜ ⎟⎟ p (1 − p) = ⎜⎜ ⎟⎟⎜ ⎟ ⎜ ⎟ = 15 * ⎜ ⎟ * ⎜ ⎟P(X>2) = P(X=3) + P(X=4) + P(X=5) +
⎝ 2 ⎠ ⎝ 2 ⎠⎝ 6 ⎠ ⎝ 6 ⎠ ⎝ 6 ⎠ ⎝ 6 ⎠
P(X=6)

⎛ 6 ⎞ 6! 6 * 5 * 4 * 3 * 2 *1 ⎛ 6 ⎞ 3 ⎛ 6 ⎞ 4 ⎛ 6 ⎞ 5 ⎛ 6 ⎞ 6
⎜⎜ ⎟⎟ = = = 15 = ⎜⎜ ⎟⎟ * (0.8) * (0.2)3 + ⎜⎜ ⎟⎟ * (0.8) * (0.2) 2 + ⎜⎜ ⎟⎟ * (0.8) * (0.2)1 + ⎜⎜ ⎟⎟ * (0.8) * (0.2)0
⎝ 2 ⎠ 2!(6 − 2)! (2 *1) * (4 * 3 * 2 *1) ⎝ 3 ⎠ ⎝ 4 ⎠ ⎝ 5 ⎠ ⎝ 6 ⎠

= 0.08192 + 0.24576 + 0.39216 +


0.262144 = 0.9810
Example 10.5

According to Victoria seed company the


germination rate of their pumpkin seeds is 10.3.4 The Mean and Standard Deviation of
the Binomial Distribution
80%, if you plant 6 seeds what is the
probability that The number of trials n and probability of
success p are two values that defines a
Binomial distribution.
a) None will germinate

b) Less than 2 will germinate X ~ Bin(n, p) means X has a Binomial


distribution with n trials and probability of
c) More than 2 will germinate success p.

Solution Mean = n*p; Variance = n*p*(1-p) = n*p*q

General information Standard deviation = np(1 − p) = npq


The number of trial n = 6 (6 seeds were
planted)

Probability of success p = 0.8 (from the Example 10.6


information given by the seed company)
What is the mean and variance of the Victoria
a) Probability that none will germinate company seed germination of Example 10.5?

The number success x = 0 Mean germination for every six seeds planted
= n*p = 0.8*6 = 4.8
P(X=0) =
Variance of seed germination = np*(1-p) =
⎛ n ⎞ x ⎛ 6 ⎞
⎜⎜ ⎟⎟ p (1 − p) n − x = ⎜⎜ ⎟⎟ * (0.8)0 * (0.2) 6 = 6.46*0.8*0.2
x10 −5 = 0.96
⎝ x ⎠ ⎝ 0 ⎠
Standard deviation = np(1 − p) = 0.96 =
0.9798
b) Less than 2 will germinate

In this case, there are two possible values


that x can take .i.e. can either be 0 or 1
since both of them are less than 2.

34
LECTURE ELEVEN: 11.2 The Normal Distribution

The normal distribution is important because it


THE NORMAL DISTRIBUTION,
SAMPLING DISTRIBUTION AND provides a model for many real-world
distributions. Most biological and other studies
CONFIDENCE INTERVAL
generate variables that follow a normal
distribution. For example plant heights, crop
11.0. Introduction yields, animal weights, IQs, etc are all found to
In lecture 10, we looked at the definition of follow a normal distribution. Even variables
random variables and discussed the binomial that are not normally distributed can be
distribution as an example of a discrete assumed to follow a normal distribution when
random variables. In this lecture, we will look the sample size in a given study is large .i.e.
at the distribution of continuous random frequency distributions for many random
variable and especially the normal distribution. variables in nature as well as several statistics
Just as a reminder we learnt from lecture 2 can be approximated by a normal probability
that a continuous variable is a type of distribution.
quantitative variables which can assume any
uncountable number of values on a line A normal distribution curve looks like the
interval. In this lecture, we will learn how to Figure 11.2 below
calculate probability of an interval values (see
figures below).

11.1 Probability Distribution Curve


Do you still remember the difference between
histogram and bar graph? Please remember
that a bar graph is used for presenting
categorical variables (also for discrete
variables) and histogram is used for presenting
continuous data. In the construction of the
Figure 11.2: A normal distribution curve
histogram, if we continue to reduce the size of
the class intervals, the lines joining the mid The centre of the normal distribution locates
points of the classes can result into a smooth the population mean ( µ ) and the spread is
curve and this can result into what we refer to determined by the standard deviation (square
as probability curve for a continuous variable root of the variance)( σ ). If the standard
(see figure 11.1a & b). Different types of deviation is increased, the normal curve
variables results into curves of different shapes become more spread out. Figure 11.3 shows
depending on the probability distribution the three normal curves that have the same mean
variable follow. The total area under the but different standard deviation. Curve 1 is
probability curve is equal to one. The general least spread (has smallest standard deviation)
shape of the probability distribution is while 3 has the largest spread (largest standard
important since it will affect the inferences spread).
about the population parameters.

35
Figure 11.4: Normal distribution curve
showing proportions covered by factors of
SD
Figure 11.3: Three normal distribution
curves with the same mean but different SD
To calculate the probability associated with an
interval we need a probability function.
The following are the characteristics of a
The probability function for a normal
normal distribution curve (normal density distribution is
cure)
Described by two parameters µ and σ 1 1 (x − µ)2
• f ( x) = e−
2πσ 2 2 σ2
• The curve is symmetric and centered on
the mean
Where µ is the population mean

• It is bell shaped with its spread determine σ 2 is the population variance


by the standard deviation σ
• The Area under the curve equals 1 or For a continuous random variable, the
100% probability that a particular outcome falls in an
• The normal density curve continues interval from a to b is equal to the area under
infinitely in both directions. the probability distribution curve for that
• For a normal distribution approximately interval.
a) 68.26% of the distribution lies between
µ -σ and µ +σ Mathematically area under the f(x) curve can

b) 95.44% of the distribution between µ- be found by integration:

2 σ and µ + 2σ b b
1 1 (x − µ)2
P(a ≤ x ≤ b) = ∫ f ( x)dx = ∫ e− dx
c) 99.74% of the distribution between µ- a a 2πσ 2 2 σ2
3 σ and µ + 3σ
The actual areas under the curve bounded by Well, you do not need to be scared by the
the various multiples of σ is shown above formula since in most of our
Figure 11.4
applications we do not use it. Statisticians have
come up with a unique table (standard normal
table) which help us to calculate area under a
normal curve.

If X is a random variable which has a normal


distribution with the population mean µ and

36
variance σ2 then mathematically, this can be The z-score gives the number of standard

represented as deviations that a given value is from the mean.


A negative z-score indicates that the value in
X ~ N (µ ,σ 2 )
question is below the mean, and a positive z-
For example, if the height (X) of students in
score indicates that the score is above the
Agriculture class has a normal distribution
mean.
with mean of 172cm and variance of 4cm2 then
mathematically this can be written as
Activity
X ~ N (172,4) Given that X ~ N (25,9)
a) Standardise (get z-scores) the following
11.2.1 Standard normal variable values of x: 25, 29, 16, 56, 34, 12, 17 and
100.
As stated in the previous section, in order to b) What is common with all x values that
calculate the probability or the area under the result into negative z values?
normal density curve we need to use a special
table of numbers. The most natural question
To Work Probability Problems for
you would ask would be “how will that table
Normally Distributed Populations
handle the different variables measured on
different scales?” For example in one study the a) Draw a graph of a normal curve, label the
variable of interest is weight of grasshoppers mean, and shade the desired area.
(milligrams) while in another study one would
b) Find the number of standard deviations the
be interested in the birth weight of elephants
given score (value) is from the mean by
(kilograms) in a zoo. The solution is to
finding the z-score.
standardise those values so that they are on the
same scale. The standard normal distribution c) Find the associated probability for the z-
table uses those standardized values. score in the standard normal probability
table.
Let X be a random variable which has a
normal distribution with the population mean d) Relate the result to the problem at hand.

µ and variance σ2 and let Z be another

variable derived from X using the formula


Example 11.1
below:
The weights of goats on a certain farm has a

( X − µ) normal distribution with µ = 20 kg and σ =2


Z= , the Z is also normally
σ kg. What is the probability of getting a goat on
distributed with a mean of Zero (0) and this farm which weighs less than 23 kg?

variance of one (1) .i.e. Z ~ N (0,1) and is


Solution
known as the standard normal distribution.
Step 1: Draw a normal curve and shade the
For the sample mean x from a sample of size
desired area
(x − µ)
n, Z=
σ n

37
µ=20 x=23 µ=0 Z=1.5

probability of getting a goat which weighs less


Step 2: Find the z-score (we need to than 23 kg.
standardize the value of interest – 23 kg)
P(X<23) = P(Z<1.5) = 0.9332
(x − µ) (23 − 20)
z= = = 1.5
σ 2 Step 4: Relate the result to the problem at
hand
Step 3: Find the associated probability for the
z-score using the standard normal probability We can conclude that there is a 93.32% chance
of getting a goat which weighs 23 or less
table.
kilograms on this particular farm. If there are
Table: 11.1: A portion of a standard normal
for example 1000 goats on this farm, then we
table
can conclude that about 933 (0.9332*1000)
z 0.00 0.01 0.02 0.03 0.04
goats weigh 23 or less kilograms.
0.0 .5000 .5040 .5080 .5120 .5160 ………
Example 11.2
0.1 .5398 .5438 .5478 .5517 .5557 ………
0.2 .5793 .5832 .5871 .5910 .5948 ……… Suppose the weight of cabbages grown in a
. . . . . .
school garden can be assumed to follow a
. . . . . .
. . . . . . normal distribution with a mean of 1.5kg and
1.4 .9192 .9207 .9222 .9236 .9251 ……… standard deviation of 0.05. What is the
1.5 .9332 .9345 .9357 .9370 .9382 ……… probability that a cabbage picked from this
. . . . . field will weigh somewhere between 0.8 and
. . . . .
. . . . . 1.7 kilograms?

Let us start by sketching the normal density


This table (z-table) has two components .i.e. curve of cabbages in the school garden and
the z-scores and probability values. The first shade the area of the curve we are interested in
column and the first row give us the z-values determining.
while the rest of the values inside the table are
the probabilities. To find the probability
corresponding a given z-value, we need to
round off our z value to 2 decimal places. To
locate a particular z value 1.50 for example,
we need to add 1.5 + 0.00 .i.e. where the row
with 1.5 meets the column of 0.00 gives us the

38
x=0.8 µ=1.5 x=1.7 z = -1.4 µ=0 z =0.4

Since our table only gives us the probability or


area of the curve to the left of a give z value, Take Note
o Because of the symmetric nature of the
we need to work the probability (area) in two
normal distribution curve, it can be proved
steps: that
1. P(Z< -z*) = 1- P(Z< z*) e.g. P(Z< -1.4)
We need to first get A1 the area to the left of x = 1- P(Z<1.40)
2. P(Z> z*) = 1 – P(Z<z*) e.g. P(Z>3.56)
=1.7, and then A2, the area to the left of x =0.8.
= 1 –P(Z<3.56)
We then subtract A2 from A1 to find the o The probability calculated can be used to
desired probability/area. approximate the proportion of individuals
with some characteristics. For example you
can determine the proportion of underweight
Area A1
birth in infants in a district if you know the
distribution of birth weight
We need to compute the z value for x=1.7 as
follows:

(x − µ) (1.7 − 1.5) 11.3 “Students” t-distribution


z= = = 0.40
σ 0.5 From (11.2) we have learnt that for

(x − µ)
Using a standard normal table, A1 (P(X<1.7) = x ~ N (µ , σ 2 ) , the ratios z = and
P(Z<0.4)) = 0.6554
σ
(x − µ)
z= have exact Normal distribution
Area A2 σ n
The z value for x=0.8 is calculated as follows: provided the variance σ2 is known.
Otherwise, these ratios are approximately
(x − µ) (0.8 − 1.5)
z= = = −1.4 normally distributed if the numerator is
σ 0.5
normally distributed and the denominator is

Using a standard normal table, A2 (P(X<0.8) = based on a reasonable number of degrees of

P(Z<-1.4)) = 0.0808 freedom, .i.e. if the samples are ‘large’. For


small samples however this approximation
The probability we are interested in is deteriorates until for very small samples the
probabilities given by the ‘z-tables’ will be
P(0.8<X<1.7) = P(Z<0.4) – P(Z<-1.4) =
seriously in error.
0.6554 – 0.0808 = 0.5746

39
In most practical application in which sample 11.5 The Central Limit Theory
means are used to estimate population means,
Let x be the mean of a sample of size n from
the value of σ 2 is not known and it is a population with an unknown distribution.
2 2
necessary to obtain an estimate s for σ When n is relatively large, the sampling
from a sample data that gives us x . “Student’ distribution of x is approximately normally
(.i.e., W. S., Gosset, 1908) showed that the distributed. The approximation becomes better

(x − µ) (x − µ) as the sample size increases. In most cases, a


ratios t= and t = have
sample size of 30 or more is good enough for
s s n
this approximation.
an exact distribution called the t-distribution.
We discuss more about t-distribution in
lectures 11 and 12.
11.6 The Confidence Interval
11.4 Sampling Distributions Sometime the main objective of a study is to
We know that in most cases we can not study estimate certain quantities e.g. yield of maize
the entire population but take a sample to per hectare, species composition in a botanical
represent the population. From a sample we analysis etc. If you want to estimate maize
can generate a number of statistics such yield per hectare for example you will use
sample mean, variance etc. several maize plots treated alike and calculate
the sample mean x as your estimate. The
The sampling distribution of a sample estimate of this kind is called a point estimate
statistic is the probability distribution .i.e. x is a point estimate of the true mean µ
associated with the various values that the
statistic can assume in repeated sampling. We learnt from earlier lectures that as the
sample increases the sample mean x becomes
The sampling distribution of x when closer and closer to the true mean but with a
sampling from a Normally Distributed
small sample size, the estimate is bound to
Population
deviate in most cases considerably from the
Let x be the mean of a sample of size n from true value. The question is now asked, can an
a normally distributed population that has interval be found within which the true mean
mean µ and standard deviation σ. For a will lie with a stated degree of confidence?
sample of size n, the sampling distribution of Using the sample statistic (point estimate), we
x: can find interval referred to as confidence
a) Is exactly normally distributed interval. We can construction a confidence
b) Is centred at µ, the mean of the interval either using a standard normal

population distribution or t-distribution.

c) Has a standard deviation of σ n,


11.6.1 Confidence Interval based a standard
where σ is the standard deviation of the
normal distribution (when σ2 is known)
population
Let x be normally distributed with mean µ
and variance σ2 then using a single

40
observation as an estimate for the population involves making hypothesis about the
mean µ , then we can construct: population, drawing sample from the

a) A 95% Confidence interval as population, testing the hypothesis by


generalising results from the sample.
µ = x ± 1.96σ when the standard

deviation is known
b) A 99% Confidence interval as
12.2 The Hypothesis
µ = x ± 2.576σ A hypothesis is a supposition, `a general rule’,
made as the basis of reasoning without
Furthermore if we have a random sample of
assuming of its truth. It can also be defined as
2
size n from a population N ( µ , σ ), we know
a tentative prediction of the outcome of the
2
that x is also ND ( µ , σ / n ), thus study. The prediction may be based on theory,
given x , an estimate of µ , and σ , the reasoning or observations. For example, we
(standard deviation) is known then: test the hypothesis that `trees from tropical rain
2 forest are taller than trees from temperate
σ
c) µ = x ± 1.96 is the 95%
n forest’ since theoretically it is known that
Confidence Interval. growth is faster in the tropic than in the

σ2 temperate region. This hypothesis can then be


d) µ = x ± 2.576 is the 99%
tested by collecting data on tree heights from
n
confidence interval. the two types of forest. A well stated
hypothesis thus gives direction to how data
11.6.2 Confidence Interval based on the t- will be collected
2
value (t), i.e., σ is unknown
2
When σ 2 is unknown we use s , it is the There are two types of hypothesis.

sample estimate, and in this case tα values will


Null hypothesis
replace zα in the above formulae.
This is the hypothesis that presumes that there
is no difference between the groups or things
LECTURE TWELVE:
being tested (e.g., there are no differences in
HYPOTHESIS TESTING
height between trees from tropical rain forest
and temperate forest). This is the basis of all
12.0 INTRODUCTON
statistical tests – some people refer to as the
When we were looking at the introduction
statistician’s hypothesis. The null hypothesis
statistics (Lecture 1), we learnt that statistics
is symbolized by H0. After performing the
has two major fields (branches) .i.e.
statistical test, the null hypothesis can be
descriptive and inferential statistics. In lectures
rejected or Accepted (statistician prefer to say
4 – 6, we looked at different ways of
failed to be rejected).
describing a data (descriptive statistics). In this
lecture, we will learn elements of inferential
Alternative hypothesis
statistics. In inferential statistics, we are
This is the hypothesis that presumes that there
concern with generalising result obtained from
are differences between the groups or things
a sample to the whole population. This
being tested. This hypothesis is sometimes
41
refers to as researcher’s hypothesis fertilizer was applied does not differ from
(researcher’s claim). The alternative those from the plots without fertilizers (null
hypothesis is denoted by H1 or HA. hypothesis). We first assume that the null
In statistics, both the null and alternative hypothesis is true .i.e. any observed difference
hypotheses are presented in term of population between yields from the two plots is pure due
parameters and not in term of sample statistics. to chance. To determine whether we reject or
Accept (fail to reject) the null hypothesis, we
Example 12.1 calculate the probability of obtaining by
Let us assume that in your school the chance a value of x at least as different from
agriculture teacher set up an experiment to test µ NF as that actually obtained.
whether addition of fertilizer improves the
yield of maize.
We can find the probability by calculating

( x − µ NF )
a) State the null and alternative hypotheses. z= and referring the result to the
σ
tables of standardised normal variable (z)
Of course, in this case, the theory would
predict that there will be differences in yield
For example if the maize yield from plots not
between plots treated and those not treated
treated with fertilizer is normally distributed
with fertilizers and this will be reflected in the
with a mean of 10 tonnes per hectare and the
alternative hypothesis.
standard deviation of 1 and we wish to test the
hypothesis that the yield from the plot treated
The null and alternative hypotheses are stated
with fertilizer is also 10 tonnes (yields from
as follows:
the two different plots are not different).
H0 µ F = µ NF Suppose the maize yield from a plot treated
H1 µ F ≠ µ NF with fertilizer is 12.3 tonnes per hectare, is the
null hypothesis true (is 12.3 tonnes statistically
Where: µF is the mean yield of maize when
different from 10 tonnes per hectare)?
fertilizer is applied

µ NF is the yield when fertilizer is not Ho: The plot treated with fertilizer has same
applied yield as those not treated ( µ F = µ NF )
H1: The plot treated with fertilizers has
b) Consider a situation where we have maize
different yield compared to those not treated
yield (x) from a plot on which fertilizer was
( µF ≠ µ NF )
applied. Suppose we know that the distribution
of maize yield is Normal and that the standard
deviation of this distribution is σ (i.e. is We can test these hypotheses by probability

known). Suppose we also know that the mean method

maize yield from plots onto which no fertilizer x = 12.3 (yield from plot treated with fertilizer)

was applied to be µ NF µ NF = 10 (assumed mean of plots not treated

Now we wish to test the hypothesis that the with fertilizers)

yield of maize from the plot onto which the σ = 1 (standard deviation)

42
Now if the observed deviation is significant
z = (12.3 − 10) = 2.3
1 one is confronted by one of the following
alternatives, viz:

Using the standard normal (Z) table P (z > 2.3)


Either : H0 is true and the observed deviation
= 0.017. This probability can interpreted as
has occurred by pure CHANCE!
the probability of obtaining maize yield of 12.3 Or : H0 is untrue!
tones per hectare by chance from a plot not
treated with fertilizer. In other words, what we But, if H0 is true, then we must ask ourselves,

are saying is that most often (98.3% of the “Is it not a remarkable coincidence that our

times) there must be other explanation for this particular observed deviation is abnormal?”

abnormally high yield other than mere chance. However, it must be remembered that the
researcher is usually not doing the test in

In this case, we would say that the observed vacuum, .i.e., there are usually possible

maize yield of 12.3 tonnes per hectare is reasons why H0 is untrue, and we may thus

(statistically) different from the mean yield of prefer to disbelieve the coincidence and reject

10 tonnes per hectare at 1.07% level. H0. In other words, if there is some unique
special condition which could account for the

The test we have done above is called the z- significant deviation, it is adopted as the

test (normal test) since it uses the probability reason why H0 is untrue. For example in this

from z values. particular example, we have a strong believe


that it is the fertilizer which is responsible for
the yield of 12.3 tonnes per hectare.
12.3 Logic of ‘tests of significance`

In the computation of the probability that we Take note


used for testing the hypothesis in Example
The deviation d indicates how far away the
12.1 we based our work on the deviation observation is from the centre (population mean µ ).
between the observed values (x) and the Remember the centre represent a typical member of
hypothetical value .i.e. d = (x - µ NF ). the population and thus the further away from the
µ NF centre, the less likely that you belong to the
With reference to that example we may ask, population and that is why we would reject the null
can the observed deviation be reasonably hypothesis.
accounted for by chance alone? This question
may then be argued along the following lines 12.4 Rejection (Type I) and Acceptance
(Type II) Errors
Procedure
In testing hypothesis, we based our test on the
Firstly form a test criterion, e.g. H 0:
results obtained from a sample. Because
{d = (x − µ ) = 0}
sample results vary from sample to sample, the
Then evaluate Prob (d) .i.e. the probability evidence provided by a particular sample can
(under H0) of the deviation as great or greater be misleading .i.e. a test of hypothesis could
than d being chance. Our argument then result in an incorrect decision. Two types of
proceeds as follows error are possible when testing hypothesis.

43
12.4.1 Type I errors α = 0.001 or 0.01%
A type I error is rejecting a null hypothesis that
is true .i.e. a false rejection of H0. The Criteria for Rejection of H0 based a given
probability of committing type I error is α value and calculated probability
denoted by Greek letter α and this correspond 1. If the p-value calculated is greater
to (100 α )% significance level. Thus for than α (p-value > α ), then we fail to
example with α =0.05, there is at most 5% reject H0 and declare the results to be
chance of wrongly rejecting H0 when H0 is insignificant.
true. 2. If the p-value is less or equal to α
(p-value ≤ α ), then we reject H0 and
12.4.2 Type II errors declare the results significant
This is the incorrect non-rejection (acceptance)
of H0 when H0 is false. The probability of Example 12.2
committing type II error is denoted by Greek In example 12.1, we calculated the probability
letter β (P (Type II error) = β ). of obtaining the maize yield of 12.3 tonnes per
hectare by chance from a plot not treated with
Take Note fertilizer to be 0.017.
We may want to keep both errors as low as
a) Is this significantly different at
possible but unfortunately, this is not possible
as you try to reduce Type I error, Type II error α =0.05
increases so the best thing to do to fixed Type I
error. b) Is this significantly different at
α =0.01
At α =0.05; p-value < 0.05 so we reject H0
12.4.3 The power of the test
and declare the results significant
The probability of correct rejection of H0 is
At α =0.01; p-value > 0.01 so we fail to reject
given by 1- β and this is referred to as the
H0 and declare the results insignificant
power of the test. Large value of the power of
the test indicates a good test.
12.6 Two-tail verses One-tail Test of
12.5 Levels of Significance
Significance
Before performing hypothesis testing, there is
In stating our hypothesis, sometimes we are
need to set up criterion for rejecting or
interested in particular direction of the
accepting H0 just like teachers set pass mark
outcome of the study. For example if you are
for exams. In hypothesis testing, this criterion
testing the effect of fertilizer application on the
is referred to as significance level α . We have
yield of maize as a scientist, you would expect
already learnt that α is associated the the plots treated with fertilizer to perform
probability of committing type I error. You better than the ones not treated and this will
will often see in research papers and statistical affect how you will state your alternative
books that H0 was rejected at significance level hypothesis.
α =0.05 etc. The most commonly used
significance levels are: 12.6.1 Two-tail test of Significance
α = 0.05 or 5% Let us once again revisit our example of
α = 0.01 or 1% fertilizer application on maize. Let µF be the

44
population mean yield of maize from plots H0 µ F = µ NF
which received fertilizer and let µ NF be the
H1 µ F < µ NF .i.e. the researcher is
population mean yield of maize from plots
interested in the negative deviation only
which did not receive fertilizer. If the
(lower tail of the
researcher is just interested in testing the
distribution)
prediction that there will be differences
The above two cases (a & b) constitute what is
between yield from plots treated with
referred to as one-tail hypothesis test of
fertilizers and those which were not treated,
significance.
our hypotheses will appear as:

H0 µ F = µ NF LECTURE THIRTEEEN:

H1 µ F ≠ µ NF thus in conclusion, there are HYPOTHESIS TESTING CONT’D

two possible alternatives


INTRODUCTION
In the last lecture, you were introduced to the
either µ F > µ NF
idea of hypothesis testing. We learnt the
or definition and types of hypothesis, the
µ F < µ NF meaning of significance level and the one and
two-tail test of significance. In this lecture, we

This alternative ( µ F will continue with looking at hypothesis. We


≠ µ NF ) is referred to as
will look at the different types of statistical
two-tail alternative hypothesis. In this case
tests available and the condition under which
whether µ F > µ NF or µ F < µ NF will
each test can be used.
still satisfies the condition that the two groups
are different. 13. 1 TESTS OF SIGNIFICANCE

12.6.2 One-tail test of significance In lecture 12, we showed how to use the z

If the researcher is only interested in the value and its associated probability to test the

hypothesis that: hypothesis. This is referred to as the z-test (or

a) The plots treated with fertilizer the normal test). The z-test works on

perform better than those not treated assumptions that the population variance is
known or the sample size is large (n>30). If
the sample size is small (n<30) and the
H0 µ F = µ NF
population variance is not known then another
H1 µ F > µ NF .i.e. the researcher is test called the student’s t-test or simply t-test is
interested in positive deviation only used.
(upper tail of the
distribution) 13.1.1 The Normal test (z-test)
b) The plots treated with fertilizers The normal test (z-test) is use when the sample
perform worst than those not treated size(s) is large and the population variances are
estimated by the sample variances or when the
population variance is known.

45
Instead of going through the tedious process of
calculating the z-values and then looking up
the associated probability, we can use the z-
value calculated directly to test our hypothesis.
Depending on the type of hypothesis under test
(one or two tails), the calculated value of z is
then compared as follows:

Figure 13.2: Normal distribution curve


Showing Acceptance and Rejection Regions
For one-tail test of significance (the tail of
interest should be taken into consideration) for 1-tail test ( α = 0.01 )
a. If (−1.645 < z for left tail) or

( z < 1.645 for right tail) the result is


not significant (.i.e., NS), and we
For two-tail test of significance
ACCEPT H0
d. If (−1.96 < z < 1.96) the result is
b. If (−1.645 > z for left tail) or
not significant (.i.e., NS), and we
( z > 1.645 for right tail) the result is
ACCEPT H0
significant, and we REJECT H0 at 5% e. If (1.96 <| z |< 2.576) the result is
level
significant, and we REJECT H0 at 5%
level
f. If (| z |< 2.576) the result is

significant, and we REJECT H0 at 1%


level
|z| means absolute value of z (consider all

z-values calculate as positive)

Figure 13.1: Normal distribution curve


Showing Acceptance and Rejection Regions
for 1-tail test ( α = 0.05 )

c. If (−2.326 > z for left tail) or

( z > 2.326 for right tail) the result is


Figure 13.3: Normal distribution curve
significant, and we REJECT H0 at 1%
Showing Acceptance and Rejection Regions
level
for 2-tail test ( α = 0.05 and α = 0.01 )

46
Case 1:
If x is N ( µ , σ 2 ) .i.e. with known In case 2 the z-value is calculated using this
2
variance σ , we can use the standard normal formula
table to decide whether an observed value of x x − µ0
z=
is significantly different from some
σ/ n
hypothetical value, µ 0 say.
Example 13.1
Procedure:
A group of 60 male students from Makerere
a) State the hypothesis H0 µ = µ0
University have a mean weight of 70.41 kg
H1 µ ≠ µ0 and estimated variance of 6.05 kg2. Records of
b) Calculate the probability that a value similar British students show a mean of
as extreme as x or more could have 70.0kg. Assuming that weight is normally

occurred by chance. Use the formula distributed, are the local students different in
weight?
x − µ0
z= to standardised the value
σ
Solution
of x and then use the z value to get
In this case we are comparing the weight from
the probability required
a sample of 60 Makerere University students
c) Decide whether the probability (or z-
value) calculated is small (big) with some hypothetical value µ 0 =70.0 kg
enough to reject H0. (see section H0 µ = 70.0 verses H1 µ ≠ 70.0 ;
12.5)
x ~ N (70,6.05 / 60 )
Case 2:
Now let us consider a sample of size n from a x − µ0 70.41 − 70
Then z = = = 1.29
population which is normally distributed with σ/ n 6.05 / 60
mean µ and variance of σ2 .i.e. σ 2 is For us to complete the hypothesis testing we
need to calculate the probability associated
known.
with this z-value
Let x be the mean from this sample of size n.
P(Z>1.291) = 1- P(Z<1.29)
From the sampling distribution (section 11.3)
= 1-0.9015 = 0.0985
we learnt that the variance of the sample mean

σ2 σ
( x ) is and its standard deviation is . Conclusion: since p-value > 0.05 (.i.e.
n n 0.0985>0.05), we accept H0 at significance
In this case we can also use standard normal
level of 5% ( α = 0.05 ) .i.e. there is no
tables to decide whether the sample ( x ) is
evidence to suggest that there is difference in
significantly different from some hypothetical
weight between the local and the British
value µ 0 say. students.
The procedure for testing this hypothesis is the
same as the one followed in case 1 but the only BUT we could have also just used the fact that
except is the formula for calculating the z – the z-value (1.29) we have calculated fall in
value.
47
the range (−1.96 < z < 1.96) which also ( x1 − x2 ) − 0
Applying equation, z=
leads to acceptance of H0 σ 12 σ 22
+
n1 n2
CASE 3:
we get:
In this case, we have two populations that we
need to compare. Consider two samples of size
2
n1 and n2 from populations x1 ~ N ( µ1 , σ 1 )
2 174.46 − 162.23
and x 2 ~ N ( µ 2 , σ 2 ) respectively. Let x1 z= = 45.94 * *
47.6522 43.7625
and x 2 be sample estimates of µ1 and µ 2 +
1164 1456
respectively. Conclusion: Since z-value calculated is
greater than 2.326, Reject H0, the men are
Consider H0: µ1 = µ2 alternatively, H0: significantly taller than the women at
α =0.01.
µ1 − µ 2 = 0

Again as for Case 1, we have very large


( x1 − x2 ) − 0 samples and the `approximate test’ is used.
Hence z=
2 2
σ 1 σ 2
+
n1 n2 13.1.2 The t-test
In most of our studies because of limited
resources, we are more likely to have relatively
For an approximate test, as for CASE 2, no
small samples in which case the approximation
assumption of normality is required provided
of the z-tests given above deteriorates until for
n1 and n2 are large. If σ 12 and σ 22 are not
very small samples the probabilities given by
known they may be replaced by their sample
`z-tables’ will be seriously in error. We
2 2
estimates, s and s respectively.
1 2 therefore need a different test for our
hypothesis.
EXAMPLE 13.2:
In an investigation to determine whether men
are taller on average than women the CASE 4: Exact t-test
following results were obtained. A group of In CASE 2 we looked at how to test whether
1164 men has mean height of 174.46 cm with the sample mean x is significantly different
2
variance ( s ) = 47.6522 cm2. A similar group from a given hypothetical value using the z-
test. By then we were able to use z-test
of 1456 women has mean height of 162.23 cm
because the population variance was assumed
2
with variance ( s ) = 43.7625 cm2.
to be known. In CASE 4, we are looking at
the situation where the population variance is

H 0: µ1 = µ 2 H A: µ1 > µ 2 (.i.e. a one-tail unknown and the sample size is small (n<30).

test, since the question asks if men are taller.)


As for the CASE 2, consider a sample size n
2
from a population N ( µ , σ ) , let
48
x be the mean from this sample. Let s 2 be
Take Note
2 o For an EXACT t-TEST, the following assumptions
the estimate of σ with (n-1) degree of
freedom (df). are made:
i. x must be Normally distributed.
ii. Numerator and denominator in the formula use
From previous lecturers, we now know that
for calculating t statistic must independent.
2
s o By |t-calculated| we mean consider only the
a) Variance of x = , and
n magnitude of the t value calculated for example |-

b) Standard error of x =s n 2.67| = |2.67| = 2.67 .i.e. we ignore the sign.

Consider, H0: µ = µ0
We now calculate a statistic t by the formula
CASE 5:
x − µ0 Just like in CASE 3, In this case, we also have
tf =
s two populations that we need to compare.
n
Consider two independent samples of size n1
Where f is the degree of freedom of
2
denominator, in this case (n-1) and n2 from populations x1 ~ N ( µ1 , σ 1 )
2
After calculating the test statistic t f most often and x 2 ~ N ( µ 2 , σ 2 ) respectively. Let x1
referred to as “t-calculated”, we compare this and x 2 be sample estimates of µ1 and µ 2
value with a t-value from a “t-table” which we respectively. However, in this case, the
will refer to as “t-tabulated”.
population variances σ 12 and σ 22 are not

known and the sample sizes n1 and n2 are


Criteria for Rejecting H0 based on a given
small.
significance level
1. If “|t-calculated|” < “t-tabulated”, then
For an exact t-test, it is necessary to assume
we fail to reject (accept) H0 and declare
that the two random samples are from Normal
the results non significant
population with equal variances (even though
2. If “|t-calculated|” > “t-tabulated”, then
the means may differ). On this assumption it is
we reject H0 and declare the results
logical to form `pooled’ estimate of common
significant
variance from the estimates of variance
“t-tabulated” is obtained from “t table” using
appropriate degrees of freedom and obtained from the separate samples, viz s12 and
significance level ( α ). s 22 with f1 and f2 degrees of freedom
respectively
(f1 = (n1-1) and f2 = (n2-1)).

f1 s12 + f 2 s 22
Thus s 2pooled =
f1 + f 2

49
Some important results In both case we calculate the‘t’ statistic using
It can be proved that if the two samples are this formula
independent then:

i. Variance of ( x1 − x2 ) = ( x1 − x2 ) − 0
tf =
SE ( x1 − x2 )
σ 12 σ 22
n1σ 1 + n2σ 22
2

+ =
n1 n2 n1 n2
Where SE( x1 − x2 ) is the square root of the
ii. Variance of ( x1 − x2 ) =
variance calculated from one of the follows
n +n
2 2 2 2
σ ( 1 2) if ( σ 1 =σ =σ 2
) above. The degrees of freedom f = (n1 + n2 -
n1n2
2). x1 and x 2 are sample estimates of µ1 and
2
σ
iii. Variance of ( x1 − x2 ) = 2 if µ 2 respectively
n
( n1 = n2 = n and σ 12 = σ 22 = σ 2
In general, SE( x1 − x2 ) =
)
2 ⎛ S 2 2 ⎞
Since population variance ( σ ) are rarely ⎜ pooled + S pooled ⎟
known it is replaced by its sample estimate ⎜ n1 n2 ⎟
⎝ ⎠
2
S pooled
Example 13.3
If we are comparing two groups usually the Imagine we measure length of shoot of tree
null hypothesis is stated as: species A and B in a given forest. The sample
H 0: statistics are:
Sample A Sample B
µ1 = µ 2 nA = 6 nB = 8
alternatively as; H0: x A = 74.8 xB = 72.99
µ1 − µ 2 = 0 SA, = 1.04 SB = 1.48
2
S A = 1.08 s B2 = 2.20
Where µ1 and µ 2 are the true means of the
Test the hypothesis that plant species A has a
respective populations.
longer shoot compared to species B

Depending on the test under consideration the Solution:


alternative hypothesis H1 can be either;
State the null and alternative hypothesis
H 1:

µ1 ≠ µ 2 (.i.e., a 2-tail test of significance);


H 0: µ A = µB
or H 1: H 1: (.i.e. 1-tail test of significance)
µ A > µB
µ1 > µ 2
or Let us set our significance level α = 0.05
}(.i.e. 1-tail tests of significance)
Calculations of sample statistics: most of it
H 1:
has already been done for you
µ1 < µ 2
fA = (nA -1) fB = (nB -1)

50
Table 13.1 Amount of α in one-tail

Degrees 0.2 0.1 0.05 0.025 0.01


2 f s 2 + f B s B2 (6 − 1) *1.08 + (8 − 1) * 2.2of 20.8
s pooled = A A = = = 1.733
fA + fB (6 − 1) + (8 − 1) freedom12
1 1.376 3.078 6.314 12.706 31.821 ………
( x1 − x2 ) − 0 (74.8 − 72.99) − 0 2 1.061 1.886 2.920 4.303 6.965 ………
tf = =
SE ( x1 − x2 ) SE ( x1 − x2 ) 3 .978 1.638 2.353 3.182 4.541 ………
. . . . . .
. . . . . .
But SE( x1 − x2 ) = . . . . . .
11 .876 .1.363 1.796 2.201 2.718 ………
⎛ S 2 2 ⎞
⎜ pooled + S pooled ⎟ = 1.733 + 1.733 = 0.505412= 0.7110
.873 .1.356 .1.782 2.179 2.681 ………
⎜ n1 ⎟ . . . . .
n2 6 8 . . . . .
⎝ ⎠
. . . . .

Thus tf =
LECTURE FOURTEEN: CORRELATION
( x1 − x2 ) − 0 (74.8 − 72.99) − 0 1.81
= = = 2.546
SE ( x1 − x2 ) SE ( x1 − x2 ) 0.711 14.0 INTRODUCTION
Previously we have studied only a single
variable at a time .i.e. for a give population or
sample we have just been studying one
The degrees of freedom for test f = (nA + nB -
2) = (f1 + f2)= (6 + 8 - 2) = 12 variable at a time. However, in many research
whether in agriculture, biology or social
To get the “t-tabulated” value from the t-table, sciences it is frequently necessary to study
we need to know the df and α
more than one variable measured on an
In this case our df = 12, α = 0.05 individual in order to get the whole picture. In
The t-tabulated can be represented as t 0.05,12 this lecture, we are turning our attention to

showing the significance level and the degrees studying relationship between two variables.

of freedom. For example, is there relationship between the


amount of fertilizer applied and crop yield? If
t 0.05,12 = 1.782 t 0.01,12 = 2.681 (see these
relationship exists between fertilizer and yield,
two values on a portion “t table” Table 13.1)
then what is the nature of that relationship? We
Conclusion: Since t-calculated (tf)> t-
are going to look at one statistical tool that can
tabulated ( t 0.05,12 ), then we reject H0 and help us to study the nature of linear
declare the results significant at α = 0.05 relationship between two variables. In this

.i.e. plant species has significantly longer shoot lecture, we will only look at linear relationship

length compared to plant species B. However, between quantitative variables.

if we change our significance level to


α = 0.01 (become more strict), then we fail 14.1 Bivariate data and Scatter plot

to reject H0.
Let us consider a situation where you sampled
10 students from your class and then you

51
record the height (x) and weight (y) of each completed on the vertical scale (y-scale). Then
student. This sample is called a bivariate each of the pairs (x, y) of observations is
sample and the data generate is a bivariate plotted as a point in the xy-plane.
data. The 10 pairs of simultaneously sampled
Scatter plot of sleep deprivation versus completion
values are designated (x1, y1), (x2, y2), (x3, of task
y3)…(x10, y10).
20

15
When the bivariate data are quantitative, we

Tasks
can use what is called a scatter plot to visual 10

any relationship that might exist between 5

them. A scatter plot is useful for displaying 0


14 15 16 17 18 19 20 21 22 23 24 25
trends in data and revealing any association
Hours
that might exist between the two variables in a
bivariate data. If the two variables are related, It appears from the scatter plot that there is an
we then ask: association between sleep deprivation and the
number of task completed .i.e. the longer one
What is the form of the relationship? How
goes without sleep, the fewer the task one
strong is the relationship? Can we predict the
completes.
value of one variable from the other?

14.2 Co-relationships
Example 14.1
From Example 14.1, we have seen that there
To study relationship between sleep
seems to be an association between sleep
deprivation and ability to complete a simple
deprivation and the number of tasks
task, 12 people were asked to solve a simple
completed. In many agricultural and biological
task after having been without sleep for 15, 18,
investigations, it is important to measure the
21 and 24 hours. The data were recorded as
strength (intensity) of co-relationships between
follows:
two variables x and y say. By intensity of
Subject 1 2 3 4 5 6 7 8 9 10 11 12
correlation is meant the extent to which
Hours 15 15 15 18 18 18 21 21 21 24 24 24
without deviations from the mean in one variable (x -
sleep µx ) tend to be accompanied by proportional
Tasks 13 9 15 8 12 10 5 8 7 3 5 4
completed deviations in the other variable (y - µ y ).
Source: Exploring Statistics 2nd Edition by Perfect correlation occurs when deviations in
Larry J. Kitchen one variable are exactly proportional to
Use the information to construction a scatter deviations in the other variable.
plot

We learnt from lecture 5 that the variability of


Solution a data set can be summarised using a sample
To construct a scatter plot, a rectangular variance. In a bivariate data set, we have two
coordinate system is drawn with the number of variables (x and y) and we can measure
hours without sleep recorded on the horizontal variability in each variable using sample
scale (x-axis) and the number of tasks variance.

52
Also for some data sets, the covariance can get
2 very large and become difficult to interpret.
Variance of x = s 2
=
∑ (x − x) and for y
x
n −1
2 Take Note
= s 2
=
∑ ( y − y)
y o Covariance cannot be used to show non-linear
n −1 relationship

Since the above two sets of variability are


computed independent of each other, they 14.2.2 Correlation Coefficient
cannot measure how these two variables vary Any measure of correlation must be
together. To assess a possible relationship, we comparable from sample to sample;
need a measure of the covariability between irrespective of the units in which the variables
the two variables. This measure of are measured.
covariability is called covariance

To get a more meaningful and scale


14.2.1 Covariance independent measure of a linear association
Let us consider a situation where you sampled between two variables, we can standardize the
10 students from your class and then you covariance by dividing it by the standard
record the height (x) and weight (y) of each deviation of the two samples. The new
student. Given the 10 pairs of observations measurement obtained after standardizing the
(x1, y1), (x2, y2), (x3, y3)…(x10, y10) of the covariance is referred to as the Pearson
variables height (x) and weight (y), the correlation coefficient. Pearson correlation
covariance between x and y is coefficient measures the strength of linear
relationship between the variables x and y and

s xy =
∑ (x i − x )( y i − y ) its values lies between +1 and -1.
n −1
The sample Pearson correlation coefficient
As the name suggests, the covariance is a denoted by r, is given by the formula below
measure of how two variables vary together.
The covariance measures the linear
∑ (x i − x )( yi − y )
dependence between the two variables. If there n −1
rxy =
is no linear dependence (association) between
∑ ( xi − x ) 2 * ∑ ( y i − y ) 2
the two variables, then the covariance will be
n −1 n −1
zero. From the formula we can see clearly that
the covariance dependent on the scale of
The above formula can be simplified by the
measurement for the two variables x and y and
removal (n-1)
thus it may be not be very useful for measuring
association. For example, values of variable x rxy =
∑ (x i − x )( y i − y )

might be in the range of 0.05 to 1.0 while the ∑ (x i − x ) 2 * ∑ ( yi − y ) 2


values for variable y may be in the range of
500 to 2000 and this differences of
measurement scale can distort the covariance.
53
The numerator of the above equation is the
called the sum of the cross products, SP(x, y) x y (x - (y - y) (x - x )(y - y)
is computed simply as, x)
0 5.1 -3 2.243 -6.729
SP(x,y) = 1 4.9 -2 2.043 -4.086

∑x ∑y
i i 2 3.4 -1 0.5 -0.543
∑ (x i − x )( y i − y ) = ∑ xi y i −
n 3 2.4 0 -0.457 0.00
4 2.0 1 -0.857 -0.857
5 1.2 2 -1.657 -3.314
Example 14.2
6 1.0 3 -1.857 -5.571
A student from the faculty of Economics and
Total -21.1
Management, Makerere University wanted to
study the relationship between house rent and
distance from campus. She selected seven
Covariance
locations north of the campus and recorded the
distance from campus and house rent for a s xy =
∑ (x i − x )( y i − y )
=
− 21.1
= −3.5167
two-room house at each location. The follow
n −1 7 −1
distance and house rents are recorded.

We now need to standardize the covariance to


Distance 0.0 1.0 2.0 3.0 4.0 5.0 6.0
get Pearson correlation coefficient
Rent 5.1 4.9 3.4 2.4 2.0 1.2 1.0
(00000)
Plot the data in a scatter plot and calculate the
∑ ( x − x)( y − y)
i i

n −1 s xy − 21.1
correlation between the two variables rxy = = = = −0.978
( x x
∑ i *∑ i
− ) 2
( y − y ) 2
s 2
x * s 2
y
4.667 * 2.772
Solution
The data appear to fall on a straight line, n −1 n −1
suggesting a strong linear relationship.

Scatter plot for distance from Makerere University This is a very high negative correlation
and House rent
coefficient indicating that there is a strong
6
5
linear relationship between the distance and
House rent (00000)

4 house rent which is consistent with the scatter


3
2
plot. The houses nearer the university are more
1 expensive than those far away
0
0 1 2 3 4 5 6 7
Distance (Kilometers)
Interpretation of r

We need to calculate the statistics for each a) The correlation coefficient (r) measures

variable before we can calculate the correlation the degree of linear dependence or

coefficient. association of two variables and the


values will be between -1 and +1.
Distance (x) x =3 s x2 = 4.667 ;
b) Independent variables have zero
2
Rent (y) y = 2.875 s = 2.772 y correlation (.i.e. they are uncorrelated)

54
(see figure 14.1c) BUT, zero correlation confirm the relationship. In this lecture, we are
does not (necessarily) imply going to look at yet another method for
independence, since the variable could be studying the relationship between two
non-linearly related (see figure 14.1d) quantitative variables. Regression analysis
c) The closer r is to either -1 or +1, the aims at formulating a mathematical
stronger the linear relationship between x relationship between two variables. In general,
and y. If all points fall on a straight line, our objective would be to express one variable
then we have a correlation of +1 if the y as a function a second variable x.
line has a positive slope and -1 if the line
has a negative slope (see figure 14.1 a & 15.1 Classification of variables in
b). Regression Analysis
d) In discussing correlation the following Unlike correlation, with regression analysis we
terminology is often used must distinguish between the response
o | r | < 0.3 - Weak correlation (dependent) variable and the predictor
o | r | < 0.5 – Moderate (independent) variable.
o | r | > 0.7 – Strong correlation
Figure 14.1: Correlation coefficient for
different relationships

A response variable also called a dependent


variable is the variable we wish to predict or
LECTURE FIFTEEN:
describe based on the values of another
REGRESSION ANALYSIS
variable. The predictor variable, also called
the independent variable, is the variable that is
15.0 INTRODUCTION
used to predict the response variable. In the
In the last lecture, we looked a measure of
context of regression, the response variable is
intensity of linear relationship between two
labeled y and the predictor variable will be
quantitative variables. With the help the scatter
labeled x.
diagram, we were able to visualize any linear
relationship between two variables before
computing the correlation coefficient to

55
15.2 Regression Equation
Regression analysis just like correlation is used We can therefore use the above regression
to study linear relationship between two equation for predicting the rent at a given
variables x and y. The first step is to draw a distance from campus. Of course, we do not
scatter plot after which you calculate r. expect our regression equation to predict the
Regression analysis is aimed at getting the rent exactly. We expect some error between
equation of the straight line that pass through the predicted value and the real (observed)
the data in the scatter plot. value since our regression line does not pass
The equation of the straight line is given by: through all the points.

y = b0 + b1 x
Let y i be the observed or recorded value of
Where b0 is the y-intercept and b1 is the slope
of the straight line, the y-intercept is the value rent at distance i from campus and let ŷ i be
of y when x = 0, and the slope gives the the predicted value of that observed rent. The
change in y relative to the change in x. predicted value ŷ i is given by:

yˆ i = bo + b1 xi
Example 15.1
Let us revisit example 14.2 (Relationship The error of prediction of rent at the ith is given

between distance from campus and house by: ei = yi − yˆ i


rent). We can fit a straight line through our
scatter plot and then proceed to determine the Example 15.2
regression equation for that line. Using the regression equation obtained in
example 15.1 to predict the rent at locations 0,
Scatter plot for distance from Makerere
Universityand House Rent 1, 2 and 3 kilometers from campus and
compute the error associated with prediction at
6
5
each distance.
y = 5.1179 - 0.7536x
Rent (00000)

4
3 Solution
2
The regression equation is given as:
1
0 y = 5.1179 − 0.7536 x
0 1 2 3 4 5 6 7
For each distance, the prediction is got by
Distance (Kilometers)
substituting the appropriate value of x in the

In this example, our regression equation is: above equation.

y = 5.1179 − 0.7536 x For x = 0;



y0 = 5.1179 − 0.7536 * 0 = 5.1179
Where the y-intercept ( b0 ) is 5.1179 and the
For x = 1;
slope ( b1 ) is -0.7536. In other words, when 
y1 = 5.1179 − 0.7536*1 = 5.1179 − 0.7536 = 4.3643
you stay within campus (distance x = 0) you
pay about 511,790 and for every 1 kilometer For x = 2;
away from campus the rent reduces by 75,360 
y2 = 5.1179 − 0.7536* 2 = 5.1179 − 1.5072 = 3.6107
(-0.7536 x 100,000).

56
For x = 3;
 b1 =
∑ ( x − x )( y − y )
i i
or
y3 = 5.1179 − 0.7536 * 3 = 5.1179 − 2.2608 = 2.8571 ∑ (x − x) i
2

x y
Prediction error
∑x y − ∑ ∑
i i
n
i i

xi yi ŷ i ei = yi − yˆ i
2 ( ∑ x) 2
0 5.1 5.1179 -0.0179 x
∑ i −
n
1 4.9 4.3643 0.5357
2 3.4 3.6107 -0.2107
bo = y − b1 x
3 2.4 2.8571 -0.4571

Example 15.3
15.3 Estimation of Regression Equation
Now that we seen how useful a regression Find the least squares regression line
equation can be for prediction purposes, one yˆ = bo + b1 x for rent data in Example 14.2
would then ask how do we estimate the
parameters of the regression equation? From x y (x - (x - (y - (x -
the scatter plot, you notice that it is possible to x) x )2 y) x )(y -
draw many lines through the data each of y)
which with its own equation. We need to select 0 5.1 -3 9 2.243 -6.729
one of these lines that will give the best
1 4.9 -2 4 2.043 -4.086
prediction .i.e. the one that will give us the
2 3.4 -1 1 0.5 -0.543
smallest prediction error. The method of
3 2.4 0 0 -0.457 0.00
selection of the regression line that minimizes
4 2.0 1 1 -0.857 -0.857
the “prediction error” ( ei = yi − yˆ i ) is 5 1.2 2 4 -1.657 -3.314
referred as the least squares method. The 6 1.0 3 9 -1.857 -5.571
least square method will select the line that Total 28 -21.1
minimizes the sum of the squared error
2
( ∑e i
). In this lecture, we are not going into

the mathematics of how the least square


Solution
method works but we are going to see the final
formulae for estimating the least squares
Distance (x) x =3 s x2 = 4.667 ;
regression line. Rent (y) y = 2.875 s y2 = 2.772

Our least squares regression line is

yˆ = bo + b1 x and below are the formulae for b1 =


∑ ( x − x )( y − y ) = − 21.1 = 0.7536
i i
2
∑ (x − x) i 28
calculating bo and b1
bo = y − b1 x = 2.875 − (−0.7536) * 3 = 2.875 + 2.2608 = 5.1358

Thus our least square regression line


is yˆ = 5.1358 + 0.7536 x . This line quite
similar to the one obtained by computer in
Example 15.1 the difference is due rounding
off figures.
57

You might also like