Introduction Statistics
Introduction Statistics
1
Probability distribution of a normal 1.2 FIELDS OF STATISTICS
random variable; the standard normal
In our discussion of statistics, we will look at
variable, use of the standard normal
distribution table and the Student’s t- the two major fields of statistics i.e. descriptive
distribution, the central limit
and inferential statistics. In descriptive
theorem; confidence interval
statistics, we will take collected information
Lecture 12: Hypothesis Testing:
(data) and attempt to summarize the
Definition of hypothesis, the null and
alternative hypotheses, Type I and information into a few numbers or graphs. One
Type II errors, Significance level,
example of a descriptive statistics is the mean.
two-tail and one-tail test of
significance Here, you are simplifying the data set by
looking only at what the average value
Lecture 13: Hypothesis Testing
Continued: happens to be. Yet another example of a
One sample z or normal test and t-
descriptive statistics is a histogram or bar
test, Two sample z test and t-test
chart. In that case, you are summarizing all of
Lecture 14: Relationship between
the information in the data set into a simple
two variables:
Scatter plot and its role; Correlation plot (picture). On the other hand, inferential
analysis and its application
statistics will take a data set and attempt to
Lecture 15 Relationship between two make assertions on the population as a whole.
variables:
Rather than saying, "...75% of students who
Regression analysis and its
application attended an interview to joint S. 1 failed
mathematics." you may wish to talk about
LECTURE ONE:
what this indicates about the population (all
GENERAL INTRODUCTION
primary school leavers) as a whole. Does it
seem very likely that over half of the pupils
1.0 INTRODUCTION
who leave primary schools do not have good
In this lecture, you will be introduced to basic
mathematics background based on the results
concepts of statistics. The lecture will explain
of this interview?
the importance of statistics in agriculture;
define the different terms and notations used
frequently in statistics.
1.3 WHY SHOULD YOU STUDY
STATISTICS
1.1 WHAT STATISTICS IS
Agriculture is the backbone of Uganda’s
We all use statistics in our daily lives such as
economy and most people derive their
when discussing football, election results,
livelihood directly through involvement in it.
businesses etc. In statistics, you study how to
In order to solve agricultural problems such as
gather information (data), summarize
low productivity due to pests and diseases, soil
information, analyze the information collected
degradation and other production constraints
and reporting the results.
scientists carry out research. In the research
You can use statistics in different fields (e.g. process, scientists are involved in collecting,
agriculture, business, etc.). In each field,
summarizing and analysing information thus
statistics is modified accordingly but the basic
concepts are the same. they need to apply the knowledge of statistics.
This sometimes involved comparing different
technologies such as methods of pest and
2
disease control, soil erosion control etc all of population (e.g. number of teachers in your
which involved the use of statistics. Apart school) is referred to as the population size
from its use to support research, statistics is and if finite is denoted by N (capital letter).
also used for monitoring agricultural
production of a country. The status of Activity
from year to year to assess progress and 2. Give five examples of infinite
the population of teachers is called a finite want to study. Statistics help us in selecting a
suitable sample.
population. On the other hand, it is not
possible to count all the grains of sand on the 1.4.2 Parameter
riverbank and this kind of population is In your study, you will always be interested in
referred to as infinite. The total number of a particular characteristics of the population
subjects or units that constitute your e.g. performance in mathematics of
3
schoolchildren in all schools in your county, 1.4.4 Estimate
heights of plants on a given plot of land, milk We now know that in most studies, we use a
production of cows in your district, etc. You sample to obtain information about the
can described or summarised those population population. Since it is generally not possible,
characteristics (mathematics performance, or very difficult, to calculate the value of the
plant height, milk production) using numbers. population parameter directly, we calculate a
The numerical description or summary of your sample statistic that corresponds to that
population characteristics is referred to as a parameter and use it as an estimate. A sample
PARAMETER (denoted by Greek letters). statistic is thus referred to as an ESTIMATE of
Examples include; population mean a population parameter. Sample mean for
(denoted µ (mu)), population variance example estimates population mean and thus it
is its estimate.
(denoted σ2 (sigma square)), population
standard deviation (denoted σ (sigma)). Since
1.4.5 Inference
a parameter is a measure from the entire In your study, you will use the information
population, it is a fixed value i.e. there is only obtained from the sample to draw conclusion
one value for each parameter for a given or generalization about the whole population.
population. For example if you realized that the 100
schoolchildren you selected from your county
1.4.3 Statistic performed very well in mathematics you can
As we have seen from above, in most of our therefore conclude that schoolchildren in your
studies we cannot study every member of the county do well in mathematics. It involves
population of interest and thus resort to generalizing from a part (sample) to the whole
studying a sample. Any numerical description (population). This generalisation from a
or summary of a sample is referred to as a sample to the population is what we called
STATISTIC (denoted by Latin letters). statistical inference.
Examples include; sample mean (denoted
2
by x (x bar)), variance ( S (s square)), 1.4.6 Experiment/Trial/Event
standard deviation ( S ). When you are Sometimes when we are carrying out studies,
selecting a sample from your population, there we are involved directly in generating
are many possible samples that can be selected information. For example, you may plant
and each sample could have different values of different varieties of maize on different plots
a given statistic. For example if you want take and later take measurement on plant height and
a sample of size 100 from the 3000 grain yield for comparison. The whole process
schoolchildren in your county, you will realise of generating information in this case from
that they are many ways to pick 100 from the planting up to measuring of plant height in
3000 children and each sample will of course statistics is referred to as an experiment. An
have different sample mean. Statistic thus experiment can be simple as; tossing a coin to
varies from sample to sample compared to the see whether head or tail turn up, rolling a die
population parameter which is fixed. to see the face turning up, closing your eye and
pointing in the direction of students to see
4
whether you pointing at a male or female 2.2 TYPES OF VARIABLES
student. In your studies, you may be interested in many
characteristics of the population (variables).
Take note
For example in studying the farms in your
§ In a simple term you can define an experiment as
district, you may be interested in knowing the
a process by which an observation (or
following:
measurement) is obtained.
§ Whether the farms made profit the previous
§ The experiment, when performed is called a trial
§ An event is an outcome of the experiment year or not
§ Amount of maize harvest in tones the
previous year
§ The number of goats on the farm
LECTURE TWO: VARIABLES § The number of farm workers employed
2.0 INTRODUCTION § Farmers opinion on the amount of rain in the
In lecture 1, you were introduced to general current year
ideas of statistics. You learnt how to define As seen from above, those variables are of
different statistical terms and notations. In this different types. In statistics, these different
lecture, you will learn about the different types types of variables are put into groups based on
of variables used in statistics and other studies. their measurement scales and other
characteristics.
2.1 WHAT IS A VARIABLE? Broadly, variables may be of two types:
We now know that in our studies, we are quantitative or qualitative.
interested in some particular characteristics of
the population and these characteristics vary 2.2.1 Quantitative variables
from one subject or unit to another. For “These are variables or characteristics that
example, mathematics performance varies represent amount or quantity of something and
from one child to another, milk production can be measured over a range of values”.
varies from one cow to another and plant Those are characteristics that you can quantify
height varies from one plant to another. A or measure using some defined scale. For
characteristic or attribute that varies from one example, maize harvested on a given farm can
subject or unit to another in statistics is be quantified and measured using weighing
referred to as variable. In your study, you will scales, the number of goats on a farm can be
have to take measurements or records on the quantified by physical counting. However,
units that are in the selected sample. although both maize harvested and the
Take note numbers of goats on a farm are quantitative
§ The value of a variable that has been recorded
or measured on a particular subject or unit is variables, they differ. When you dealing with
referred to as an observation maize harvested on a farm, it is possible to get
§ A data is a set of observations usually of several
variables taken on many individual subjects or yield of 4.5 tones but you cannot get 4.5 goats.
units. The individual subjects or units may be
human, plants, animals, machines, buildings, This brings us to another division within the
organisations, etc quantitative variables category. Quantitative
variables may be either discrete or continuous.
5
iii. Nominal/Categorical variables are
Take note
§ A continuous quantitative variable is one associated with some quality,
for which all values in some range are characteristic or attribute which the
possible e.g., yield of maize, height, mass,
etc subject posses; for example, eye
§ A discrete quantitative variable is one for
which only certain values are possible. colour: (blue – grey – green - brown),
They are consecutive integers; for example, vehicle type: (bus – truck – car); etc.
number of goats on a farm, size of
household, number of insects in an insect In this case, there are more than two
trap, number student in your class etc.
categories.
individual subject or unit. It includes variables In lecture 2, you learnt about the different
that can be categorised but not quantified. You types of variables. In this lecture, you will
learn methods of data collection, sources of
can for example categories students using and types of data.
name of their villages, whether they are
From what we learnt from lecture 1, we can
present or absent on a given day. Class grading view statistics as consisting two parts; a)
such as good, very good or excellent are also collection of information (data) and b)
processing information (summarizing and
examples of qualitative variables. Just as we analyzing data), in this lecture, we will dwelt
found out in quantitative variables, qualitative on collection of data.
6
responsible) to result to lateness of the the weight of goats in a given village, we can
students. consider farmers or households as our clusters.
Once a household is selected for our study, all
You may ask yourself this question “how do I the goats in the household will be weighted,
select a suitable sample?” In statistics, a goats in non-sampled households will not be
sample is selected through a process referred studied.
to sampling.
Stratified Sampling: This is used when the
population is divided into groups or strata in
Sampling Methods
which there is less variation within groups and
Sampling is a process by which a sample is
more variation between groups. For example if
selected from a given population so as to
we look at two classes (P1 and P7) in a
obtained information about the population. The
primary school, within each class height of
main idea behind sampling is to ensure that the
pupils are similar BUT between the two
sample selected is a good representative of the
classes there a big difference. In stratified
target population. For example if you are
sampling, we would do a simple random
interested in finding out the proportion of
sampling from each group to contribute to our
teachers in your school who drink alcohol, ten
study sample.
teachers picked from “Mama Brown’s bar” is
certainly not a good representative and neither
3.1.2 Experiment
is a group of teachers picked from a born-again As a teacher of agriculture in your school, you
church or a mosque. The sample that may want to demonstrate to the neighboring
misrepresents the population is referred to as a community that application of fertilizer
biased sample whereas a sample that is a good improves the yield of maize or that a certain
representative of the population is unbiased. new variety of maize yield better than the local
variety. To prove your point you need to set
Simple random sampling: It is a type of In this study, you are interested in studying the
sampling where each member of the effect of fertilizer or variety (factor of interest).
population has an equal chance of being Since you can determine which plot of land
selected. The sampling can be done using a receive fertilizer and which ones do not
table of random numbers or by what is called receive, you are playing an active role in
applies when the population is grouped into controlling/manipulating the factors of interest
randomization deals with the procedure of intervention data come from doing
or factor to more than one study unit to ensure NUMERICAL DESCRIPTION OF DATA –
with fertilizers and a similar number without In our last lecture, we looked at the different
fertilizers. In local control, we are concern aspect of data. After collecting a “good data”,
with how to control for other factors that we you need to process that information through
3.2 Types of Data No matter what your final objectives are, you
As you read different books, you will realize must first adequately describe your data. This
that different field of studies or authors applies to both sample and population data.
3.2.1 Qualitative verses Quantitative Data measurements to a few summary measures that
This categorization is mainly used in social provide a good, rough picture of the original
research. Qualitative data is a data in which the measurements. There are two methods for
audiovisual and pictures whereas quantitative methods. In this lecture, you will learn how to
data has information stored inform of numbers. use numbers to summarize or describe your
data.
3.2.2 Primary verses Secondary Data
study, then the data is referred to as primary class based on marks scored in your subject
8
and the best way to do that is to look at how Example 4.1
the marks are distributed. You need to know
the lowest and highest marks scored, range of If in your demonstration of the effect of
marks scored by most students, the marks that fertilizers on maize yield (see section 3.1.2)
divide the students into two groups, the mark you collected maize yield data from 10 plots
below which the worst 25% performers lie, ( x1 , x 2 , x3 , x 4 , x5 , x 6 , x 7 , x8 , x9 , x10 ) on
etc. What we have just described about your
which fertilizer was applied and would want to
class marks and students’ grouping are what
get the mean of this sample. The sample mean
we call measures of location. We are going to
is
look at arithmetic mean, median, mode and
quartiles as examples of measures of location i =10
( x1 =80, x 2 =20, appear as 20, 22, 23, 24, 25, 26, 27, 28.
4. Calculate the mean of the remaining data
x3 =20, x 4 =22, x5 =28, x6 =25, x7 =26, x8 =27
(after removing the extreme value(s)). In
, x9 =24, x10 =23) our demonstrations example,
were obtained from your fertilizer
demonstration .i.e. from plot 1 we obtained 80
kilogram per square meters instead of 25. As xT = (20 + 22 + 23 + 24 + 25 + 26 + 27 + 28) / 8
we have noted before, the yield of 80 kilogram
= 187/8 = 23.75. This is a better measure of
per meter squares is an extreme value
centre compared to untrimmed mean of 18.5
(abnormal compared to the other nine plots).
Because of this extreme value, the new mean
is now 18.5 kilogram per meter square instead Activity
of the original 24 kilogram per meter square 1. Using the data in Example 3.2, calculate a
4.4 4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 observation/measurement
4.4 4.2 that occurs most
often (with the highest frequency). You simply
1. Position of the median = (n + 1)/2 =
get the mode by counting how many times
(13 +1)/2 = 7 .i.e. the median lies in
each value occurs and the one with the highest
the 7th position
is the mode.
2. Rearranging the data in ascending
order:
Example 4.5
4.2, 4.2, 4.3, 4.4, 4.4, 4.4, 4.5, 4.7, 4.8,
Let us go back and revisit the data used in
4.8, 4.8, 4.9, 4.9
example 4.3 (time in seconds run by 13
3. The observation in the 7th position
students) and count the number of times each
from either side of arranged data is
value appears.
our median. Our median is 4.5
4.4 4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 4.2
Example 4.4
Even number of observations Data How many observations with this value in
value the data set
If instead of giving you the time run by 13
students in Example 4.3, you are given the 4.2 2
time for 12 students. Find out the median
4.3 1
4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 4.2 4.4 3
4.5 1
Steps 4.7 1
11
integer part” means that if (n+1)/2 has a
4.8 3
0.5 decimal, we just drop it off before
4.9 2
adding 1 and dividing by 2.
In this example, two values (4.4 and 4.8) have 2. Rearrange the observations in ascending
therefore 4.4 and 4.8. Have heard about 3. The first quartile Q1 is found by
bimodal rainfall? We can have a uni-modal (1 counting the observations from the lower
mode), bimodal (2 modes) or multimodal end of the ordered data until we get the
(more than 2 modes) data. observation in the quartile position. The
Steps in calculating the quartiles 4.5, 4.7, 4.8, 4.8, 4.8, 4.9, 4.9
4.9 4.2 4.4 4.8 4.9 4.8 4.5 4.3 4.8 4.7 4.4 The
4.2 measures of variability that we are going
to look at in this lecture will include: range,
Pm = (12 +1)/2 = 6.5 variance and standard deviation
5.2 Variance and Standard Deviation is the average squared distance of the sample
The range is a useful measure of variability for values from the sample mean. It is calculated
a small data set. For a large data set, however, with the formula
You will realise shortly that in calculating the as sum of squares (SS), which measures the
variance all observations in a data set will total squared deviation of the whole data. The
Observation
x2 ( x1 − x ) ( x1 − x ) 2
(x) By using an alternative formula
2 2
68 68 = (68 - 70) = - (-2) = 4
⎛ (∑ xi ) 2 ⎞
4624 2 ⎜ 2 i
⎟
⎜ ∑ xi − ⎟ for calculating the
70 4900 (70 - 70) = 0
⎜ n ⎟
0 ⎝ ⎠
69 4761 (69 - 70) = - 1
sums of squares, we can avoid the tedious
1
process of going through the construction of
70 4900 (70 - 70) = 0
0 the table used in example 5.1. Let us now try
71 5041 (71 - 70) = 1 to prove that the two formulae for calculating
1 sums of squares give us the same result. From
72 5184 (72 - 70) = 4
our calculation in Example 5.1, we now know
2
Total 420 29410 10 that the sum of squares is 10.
⎛ n ⎞ ( ∑ xi ) 2
⎜ ∑ ( xi − x ) 2 ⎟
(420) 2 176400
variance ( s = ⎜ i =1
2
⎜ n −1
⎟ )
⎟
∑ xi2 − n
i
n
= (29410 −
6
) = (29410 −
6
) = 10
⎜ ⎟ = ∑ (x i − x)2
⎝ ⎠ i =1
15
the mean. Coefficient of variation measures
⎛ (∑ xi ) 2 ⎞
⎜ 2 i
⎟ the variability in the values of the data relative
⎜ ∑ xi − ⎟ for calculating the
⎜ n ⎟ to the mean.
⎝ ⎠
s
sums of squares. CV = x100
x
IQR = Q3 − Q1
As a measure of variability, inter-quartile
range is useful for comparing variability of
Example 5.2 two or more data sets,
In example 5.1, we got the variance of the
heights of boys to 10cm2 and for girls to be Example 5.3
200cm2. Using the results obtained from example 4.6
Sample standard deviation (standard error) for Data: 4.2, 4.2, 4.3, 4.4, 4.4, 4.5, 4.7, 4.8, 4.8,
boys = 2 = 1.414 cm 4.8, 4.9, 4.9
Sample standard deviation (standard error) for Q1 = 4.4 ; Q3 = 4.8
girls = 20 0 = 14.14 cm IQR = Q3 − Q1 = 4.8 − 4.4 = 0.4
From section 2.5.2 you learnt that categorical Increased 305 305/500=0.61
variables are associated with some quality,
Decreased 25 0.05
characteristic or attribute which the variable
posses; for example, eye color, vehicle type, Remained 150 0.30
6.1: Frequency table showing farmers’ important because its value is independent of
responses on status of maize yield on their the size of the sample and this is necessary
17
6.1.2 Bar Graph Note that although the two bar graphs are
similar in most aspects, the vertical axes are
We all know that tables of numbers are
scaled differently. In the first graph, the
sometimes difficult to interpret especially for
vertical axis shows frequencies whereas in the
those with little or no education background or
second graph we have relative frequencies.
even highly educated people. In those cases, a
The relative frequencies is standardized with
picture might better illustrate the distribution
values ranging from 0 to 1 and this makes it a
of the data. Apart from being easy to interpret,
good tool for comparing two or more groups
graphs are also very good at showing trend and
with different sample sizes.
pattern in the data. For example, trend in
coffee production in Uganda over the last 10 Example 6.1
year can be picked easily from a graph than
Assuming the question about the yield of
from table of numbers.
maize on the farmers’ fields was also asked to
Let us try to convert the information on the selected farmers in two other districts
response of farmers to the maize yield question (Mbarara and Arua) and you are requested to
in Table 6.1 to a bar graph. As the name compare the response from the different
suggest, a bar graph consist a number of districts. Table 6.2 shows the result obtained
disjoint bars whose heights are determined by from the different districts.
the frequency or relative frequency of the
category represented by the bar. The bars can
be vertical or horizontal.
18
Gulu
Example 6.2
84 59 82 78 96 44 76 85 66 77 91 62 54 72 65
84 38 76 70
19
By observation, we see that the marks are two-
digit numbers that range from the 30s to 90s.
Example 6.3:
To construct the display, we divide each
3 8
observation into a stem and leaf. In this
4 4
example the digit in the tens place of the
5 4 9
number become the stem, and the digits in the
2 5 6
units place becomes the leaf of the stem and
7 0 2 4 6 6 7 8
leaf plot. A vertical line is drawn to separate
8 2 4 4 5
the stems from the leaves.
9 1 6
3 3 8
4 4 4
Using the data of marks from example 6.3, we
5 5 4 9
construct a histogram using each stem as a class.
6 2 5 6
From our stem and leaf plot we see that in the
7 7 0 2 4 6 6 7 8
class limit for stem value 3 are 30 and 39, for
8 4 8 2 4 4 5
stem value 4 are 40 and 49, and so on. Once the
9 9 1 6
class limits are determined, that data can be
(a) (b)
formulated in a grouped frequency table
(a) The first observation (84) plotted
(grouped in the sense that values are grouped
(b) complete and ordered stem and leaf plot into varies classes) see table below
The stem and leaf plot leaves the data intact Table 6.3: Grouped frequency table
for future calculations. In our example above
Class Class Frequency Relative
we can see that most students scored 70 and
Limits Boundaries Frequency
above in this particular test (see the shaded
30 - 29.5 – 39.5 1 1/19
part).
39
40 - 39.5 – 49.5 1 1/19
49
6.2.2 Histograms 50 - 49.5 – 59.5 2 2/19
59
The histogram is a graphical mean of
60 - 59.5 – 69.5 3 3/19
illustrating the distribution of continuous 69
quantitative data. The measurements are 70 - 69.5 - 79.5 6 6/19
79
grouped into classes defines as an interval of
80 - 79.5 – 89.5 4 4/19
values. The class limits are the smallest and 89
largest possible values in the interval. Once the 90 - 89.5 – 99.5 2 2/19
class limits are determined, the data can be 99
4
Helpful hints:
2
• After dividing the range by the
0
desired number of classes to obtain
5
5
3.6
3.7
3.8
3.9
4.0
4.1
4.2
4.3
4.4
4.5
4.6
4.7
4.8
4.9
5-
5-
5-
5-
5-
5-
5-
5-
5-
5-
5-
5-
5-
3.5
3.6
3.7
3.8
3.9
4.0
4.1
4.2
4.3
4.4
4.5
4.6
4.7
4.8
21
that is not all about statistics. You learnt from
lecture 1 that statistics has two major
7.1.2 Outcome
fields/branches i.e. descriptive and inferential. An outcome is the result of a single trial of a
In inferential statistics, we are concern with probability experiment. In rolling a die once
how to use information from a sample to draw for example, the outcome can be 1, 2, 3, 4, 5 or
generalization about the population. The 6 i.e. you can get one of the numbers 1 to 6
knowledge of probability will help us to showing up (see Figure 7.1).
understand the concepts of statistical inference
that we are going to cover in the next part of
the course. In this lecture, you will be
introduced to the basic concepts of probability Figure 7.1: Faces of a die-Any one of the 6
numbers can show up when you roll a die
including basic definitions.
22
showing up. In the above case your event are a die at the same time, the outcome of the coin
{6}, {1, 3, 5} and {2, 4, 6} for a 6, an odd or toss experiment and the rolling of a die do not
even number respectively. affect each other.
Take Note
7.1.9 Dependent Events
An event may have more than one outcome.
Two events are dependent if the first event
For example when you are interested in an odd
affects the outcome or occurrence of the
number from rolling a die as your event, then
second event in a way that the probability is
your event has more than one outcome .i.e. if
any of the numbers 1, 3 or 5 appears it will still changed.
be an odd number and your interest will be met.
and the number 1 comes up, there is no way summing up the faces that show up on the 2
that on the same die the number 2, 3, 4, 5 or 6 dice. The sums are {2, 3, 4, 5, 6, 7, 8, 9, 10,
will show thus these events mutually 11, 12}. However, each of these are not
exclusive. Mutually exclusive events are also equally likely. The only way to get a sum 2 is
referred to as disjoint events. to roll a 1 on both dice, but you can get a sum
of 4 by rolling a 1-3, 2-2, or 3-1. Figure 7.2
7.1.8 Independent Events and Table 7.1 illustrate a better sample space
Two events are independent if the occurrence for the sums obtain when rolling two dice.
of one does not affect the probability/chance of When two dices are rolled there are 36 unique
the other occurring. If you toss a coin and roll ways in which the dices can land (see Figure
7.2 and Table 7.1)
23
If you roll a die, the sample space has six
outcomes S = {1, 2, 3, 4, 5, 6}. Let B be the
event that an even number show up. In this
case, our event has three outcomes {2, 4, 6}
.i.e. if any of the three numbers 2, 4 or 6 show
up we still got an even number. The
probability that an even will show up when we
roll a die is 3/6 or ½.
Figure 7.2: Sample space of rolling 2 dices
(see the sums in Table 7.1)
Example 7.3
Consider the example of rolling two dice used
Table 7.1: Sums of two dices
to illustrate the concept of sample space in
Second Die section 7.2. The table below show the possible
7.3 Classical verses Empirical Probability from two dice) sample space)
2 1 1/36
7.3.1 Classical Probability
3 2 2/36
Classical probability uses the sample space to
4 3 3/36
determine the numerical probability that an 5 4 4/36
event will happen. Classical probability is also 6 5 5/36
called theoretical probability. In this case, to 7 6 6/36
get the probability of a given event you need to 8 5 5/36
divide the number of outcomes that are in your 9 4 4/36
Example 7.1
If you toss a coin the sample space is S =
{Head, Tail} or {H, T}. Let A be the event that
the head (H) shows up. In this case, our event
has only one outcome .i.e. H and our sample
space has two possible outcomes H and T. The
probability that H will show up is ½.
Example 7.2
24
Take Note
For classical probability, the probability of an
LECTURE EIGHT:
event occurring is the number outcomes in
LAWS OF PROBABILITY
the event (n(E)) divided by the number of
outcome in the sample space (n(S)). INTRODUCTION
P(E) = n(E) / n(S)
However, this is only true when the In lecture seven, we looked at the various
outcomes are equally likely. terms and concepts used in probability. Now
you can talk comfortably about probability
terms without any fear. In this lecture, you
7.3.2 Empirical Probability will learn the different laws/rules of
Unlike classical probability, empirical probability that will help your further
probability is based on observations. Empirical understanding of probability.
probability is the relative frequency of a
frequency distribution based upon 1. All probabilities are between 0 and 1
observations. Do you remember how we inclusive. In mathematical term, this is
calculated the relative frequency in lecture 4? represented as shown below .i.e. the
If for example you roll a die 120 times and you probability that a given event E will
realise that the number 1 occurs 20 times then occur, lies between 0 and 1.
the empirical probability of 1 is 20/120. 0 ≤ P( E ) ≤ 1
2. The sum of all the probabilities in the
P(E) = (Number of times the event occurs
sample space is 1. Let us recall that when
(frequency)) / (Total number of trials) = f/n
you roll a die the sample space S = {1, 2,
3, 4, 5, 6} what this law tells us is that
Example 7.4
when you roll a die you are certain that
In order to determine the probability of having
one of those numbers will show up.
malaria parasites among students in your
Mathematically we can represent the sum
school, you decided to sample 50 students and
of the probabilities of the sample space as
took them for a laboratory test? If it turns out
that 15 students tested positive for the parasite,
what will you conclude about the probability P( S ) = P(1) + P(2) + P(3) + P(4) + P(5) + P(6) = 1
of occurrence of malaria parasites at your
school? Note that we can use the empirical 3. The probability of an event which cannot
method to determine this probability since this occur is 0.
is based on observations. We just need to the 4. The probability of any event which is not
divide the frequency of malaria parasites (15) in the sample space is 0. For example, if
with the total number of students tested (50). you roll a die the probability that a number
P(Malaria parasites) = (Number who tested 7 will show up is 0 since the number 7 is
positive)/(Total number tested) = 15/50 not in our sample space .i.e. a die only has
six sides numbered from 1 to 6.
25
5. The probability of an event which must occur. First ask yourself, are these two events
occur is 1. For example, the probability (A and B) mutually exclusive? Yes.
that the sun will rise tomorrow is 1 since it
Thus
must happen.
27
Using the classical probability we know that P(A|B) . This is the conditional probability of
the probability of event A occurring is P(A) = event A.
1/6 and probability that event B occurring is
P(B) = ½ P(A and B)
P(A/B) =
P( B)
Thus P(A and B) = (1/6)*(1/2) = 1/12
From the above formula, we can see that
LECTURE NINE:
P(A and B) = P(B) * P(A|B)
CONDITIONAL, JOINT AND
MARGINAL PROBABILITY This looks like the multiplicative law for
independent events. In fact, it is the general
multiplicative law for both independent and
INTRODUCTION
dependent events.
From the last two lectures we have learnt
terms, concepts and rules at are used in
Example 9.1
understanding and calculations of probability
associated with different types of events. In Suppose that at St. Charles Lwanga College,
lecture 8, we studied rules that are used in the 40% of students take Biology, 25% take
exclusive and independent events. In this randomly selected from a Biology class, what
lecture, we will learn conditional events and is the probability that he or she is also taking
Female 12 28 40
Total 31 69 100
Take Note
The following four statements are equivalent 1. What is the probability of a randomly
selected individual being a male who
1. A and B are independent events
smokes? This is just a joint probability
2. P(A and B) = P(A) * P(B)
since we are looking at the sex and
3. P(A|B) = P(A) smoking habits if the individuals jointly.
4. P(B|A) = P(B) The number of "Male and Smoke" divided
by the total = 19/100 = 0.19
The last two are because if two events are
independent, the occurrence of one does not
change the probability of the occurrence of the 2. What is the probability of a randomly
other. This means that the probability of B selected individual being a male? This is
occurring, whether A has happened or not, is
simply the probability of B occurring. the total for male divided by the total =
60/100 = 0.60. Since no mention is made
of smoking or not smoking, it includes all
9.2 Joint Probability the cases. In this case we are only
In many situations, you may be interested in interested in the probability of picking a
observing more than one characteristics of the male and this is referred to as marginal
subject or study unit. For example, the probability
smoking and drinking habit of a group of
teachers may be of interest and we would
29
3. What is the probability that a randomly you are going to be introduced to the concept
selected individual smoking? Again, since of probability distribution and we learn about a
no mention is made of gender, this is a binomial distribution.
marginal probability, the total who smoke
10.1 Random Variable
divided by the total = 31/100 = 0.31.
INTRODUCTION
A random variable can be discrete or
In previous lecture (lecture 9) you learnt how continuous. The random variable whose values
to deal with dependent and joint events and the consist of numbers such as {0, 1, 2, 3, 4} or at
probabilities associated with those types of most a countable number of values such as
events. From lecture 7, you learnt that
probability experiment results into outcomes. {1, 2, 3, 4, . . .} is referred to as discrete
You know that for example that when you roll random variables. Continuous random
a die then one of the faces will have to show variables are those that can assume all the
up as an outcome or if you spray an insect with values in an interval of the number line.
S = {(g, g, g), (g, g, b), (g, b, g), (b, g, g), (b, b, • A none twin birth results in either a
g), (b, g, b), (g, b, b), (b, b, b)} boy (success) or a girl baby (failure)
Activity
We see that there are eight possibilities, and if o List five other examples of a Bernoulli population
we assume that a boy is just as likely as a girl, and in each case define what you consider to be
success and failure
all of the eight possibilities are equally likely.
Of the eight only one (b, b, b) correspond to
zero girls, in which w = 0 and the probability
When you tossed a coin once, you have
of zero is 1/8. Three of the eight possibilities
performed a Bernoulli trial but when you
correspond to w = 1(one girl), and therefore
repeat the process more than once then it
the probability is 3/8. Continuing in this
ceases to be a Bernoulli. A sequence Bernoulli
fashion we get the following completed
trials is called a binomial experiment
probability distribution of w.
10.3.1 Bernoulli Population and Trial • There are only two outcomes (success
A Bernoulli population is a population in or failures)
which each element is one of the two
possibilities. The two possibilities are usually • The probability of each outcome
designated as success and failure. A Bernoulli remains constant from trial to trial.
• Spraying 20 insects with insecticide to see example n = 3 (we are testing three
• Rolling a die until a 6 appears (not a fixed (two people testing positive)
number of trials)
Anytime a person test positive, it is a success
• Asking 20 people how old they are (not (denoted S) and anytime negative appears, it is
Example 10.2
32
Table 10.1: Possible Scenario on testing Further, note that there are three ways this can
three people for Malaria occur. This is the number of ways 2 successes
can be occur in 3 trials without repetition and
Possible test People tested for malaria
order not being important, or a combination of
results
Okello Mukasa Muhwezi 3 things, 2 at a time.
Scenario 2 Positive Negative Positive o The number of combinations of k objects taken from n
objects is given by
(S) (F) (S)
⎛ n ⎞ n!
Scenario 3 Positive Positive Negative ⎜⎜ ⎟⎟ = Where n! is read “n factorial”
(S) (S) (F)
⎝ k ⎠ k!(n − k )!
and is given by
n!=n(n-1)(n-2)(n-3). . . (2)(1)
In the first scenario, Okello tested negative
o When calculating n!, we will assume that
(failure) for malaria but Mukasa and Muhwezi
a) 1! = 1 and
tested positive (successes) and the probability
associated with this given as b) 0! =1
33
In this example, the number of trials n = 6 (6 P(X< 2) = P(X =0) + P(X=1)
people are being tested), the number successes
x =2 and using the general probability of ⎛ 6 ⎞ 0 ⎛ 6 ⎞ 1
= ⎜⎜ ⎟⎟ * (0.8) * (0.2) 6 + ⎜⎜ ⎟⎟ * (0.8) * (0.2) 5
formula ⎝ 0 ⎠
= 0.0015 ⎝1 ⎠
c) More than 2 will germinate
⎛ n ⎞ x
P(X= x) = ⎜⎜ ⎟⎟ p (1 − p) n − x = In this case, there are four possible values
x
⎝ ⎠ that x can take .i.e. x can be 3, 4, 5 and 6
2 4 2 since
4 each of them is greater than 2.
⎛ 6 ⎞ 2 3− 2 ⎛ 6 ⎞⎛ 1 ⎞ ⎛ 5 ⎞ ⎛ 1 ⎞ ⎛ 5 ⎞
⎜⎜ ⎟⎟ p (1 − p) = ⎜⎜ ⎟⎟⎜ ⎟ ⎜ ⎟ = 15 * ⎜ ⎟ * ⎜ ⎟P(X>2) = P(X=3) + P(X=4) + P(X=5) +
⎝ 2 ⎠ ⎝ 2 ⎠⎝ 6 ⎠ ⎝ 6 ⎠ ⎝ 6 ⎠ ⎝ 6 ⎠
P(X=6)
⎛ 6 ⎞ 6! 6 * 5 * 4 * 3 * 2 *1 ⎛ 6 ⎞ 3 ⎛ 6 ⎞ 4 ⎛ 6 ⎞ 5 ⎛ 6 ⎞ 6
⎜⎜ ⎟⎟ = = = 15 = ⎜⎜ ⎟⎟ * (0.8) * (0.2)3 + ⎜⎜ ⎟⎟ * (0.8) * (0.2) 2 + ⎜⎜ ⎟⎟ * (0.8) * (0.2)1 + ⎜⎜ ⎟⎟ * (0.8) * (0.2)0
⎝ 2 ⎠ 2!(6 − 2)! (2 *1) * (4 * 3 * 2 *1) ⎝ 3 ⎠ ⎝ 4 ⎠ ⎝ 5 ⎠ ⎝ 6 ⎠
The number success x = 0 Mean germination for every six seeds planted
= n*p = 0.8*6 = 4.8
P(X=0) =
Variance of seed germination = np*(1-p) =
⎛ n ⎞ x ⎛ 6 ⎞
⎜⎜ ⎟⎟ p (1 − p) n − x = ⎜⎜ ⎟⎟ * (0.8)0 * (0.2) 6 = 6.46*0.8*0.2
x10 −5 = 0.96
⎝ x ⎠ ⎝ 0 ⎠
Standard deviation = np(1 − p) = 0.96 =
0.9798
b) Less than 2 will germinate
34
LECTURE ELEVEN: 11.2 The Normal Distribution
35
Figure 11.4: Normal distribution curve
showing proportions covered by factors of
SD
Figure 11.3: Three normal distribution
curves with the same mean but different SD
To calculate the probability associated with an
interval we need a probability function.
The following are the characteristics of a
The probability function for a normal
normal distribution curve (normal density distribution is
cure)
Described by two parameters µ and σ 1 1 (x − µ)2
• f ( x) = e−
2πσ 2 2 σ2
• The curve is symmetric and centered on
the mean
Where µ is the population mean
2 σ and µ + 2σ b b
1 1 (x − µ)2
P(a ≤ x ≤ b) = ∫ f ( x)dx = ∫ e− dx
c) 99.74% of the distribution between µ- a a 2πσ 2 2 σ2
3 σ and µ + 3σ
The actual areas under the curve bounded by Well, you do not need to be scared by the
the various multiples of σ is shown above formula since in most of our
Figure 11.4
applications we do not use it. Statisticians have
come up with a unique table (standard normal
table) which help us to calculate area under a
normal curve.
36
variance σ2 then mathematically, this can be The z-score gives the number of standard
37
µ=20 x=23 µ=0 Z=1.5
38
x=0.8 µ=1.5 x=1.7 z = -1.4 µ=0 z =0.4
(x − µ)
Using a standard normal table, A1 (P(X<1.7) = x ~ N (µ , σ 2 ) , the ratios z = and
P(Z<0.4)) = 0.6554
σ
(x − µ)
z= have exact Normal distribution
Area A2 σ n
The z value for x=0.8 is calculated as follows: provided the variance σ2 is known.
Otherwise, these ratios are approximately
(x − µ) (0.8 − 1.5)
z= = = −1.4 normally distributed if the numerator is
σ 0.5
normally distributed and the denominator is
39
In most practical application in which sample 11.5 The Central Limit Theory
means are used to estimate population means,
Let x be the mean of a sample of size n from
the value of σ 2 is not known and it is a population with an unknown distribution.
2 2
necessary to obtain an estimate s for σ When n is relatively large, the sampling
from a sample data that gives us x . “Student’ distribution of x is approximately normally
(.i.e., W. S., Gosset, 1908) showed that the distributed. The approximation becomes better
40
observation as an estimate for the population involves making hypothesis about the
mean µ , then we can construct: population, drawing sample from the
deviation is known
b) A 99% Confidence interval as
12.2 The Hypothesis
µ = x ± 2.576σ A hypothesis is a supposition, `a general rule’,
made as the basis of reasoning without
Furthermore if we have a random sample of
assuming of its truth. It can also be defined as
2
size n from a population N ( µ , σ ), we know
a tentative prediction of the outcome of the
2
that x is also ND ( µ , σ / n ), thus study. The prediction may be based on theory,
given x , an estimate of µ , and σ , the reasoning or observations. For example, we
(standard deviation) is known then: test the hypothesis that `trees from tropical rain
2 forest are taller than trees from temperate
σ
c) µ = x ± 1.96 is the 95%
n forest’ since theoretically it is known that
Confidence Interval. growth is faster in the tropic than in the
( x − µ NF )
a) State the null and alternative hypotheses. z= and referring the result to the
σ
tables of standardised normal variable (z)
Of course, in this case, the theory would
predict that there will be differences in yield
For example if the maize yield from plots not
between plots treated and those not treated
treated with fertilizer is normally distributed
with fertilizers and this will be reflected in the
with a mean of 10 tonnes per hectare and the
alternative hypothesis.
standard deviation of 1 and we wish to test the
hypothesis that the yield from the plot treated
The null and alternative hypotheses are stated
with fertilizer is also 10 tonnes (yields from
as follows:
the two different plots are not different).
H0 µ F = µ NF Suppose the maize yield from a plot treated
H1 µ F ≠ µ NF with fertilizer is 12.3 tonnes per hectare, is the
null hypothesis true (is 12.3 tonnes statistically
Where: µF is the mean yield of maize when
different from 10 tonnes per hectare)?
fertilizer is applied
µ NF is the yield when fertilizer is not Ho: The plot treated with fertilizer has same
applied yield as those not treated ( µ F = µ NF )
H1: The plot treated with fertilizers has
b) Consider a situation where we have maize
different yield compared to those not treated
yield (x) from a plot on which fertilizer was
( µF ≠ µ NF )
applied. Suppose we know that the distribution
of maize yield is Normal and that the standard
deviation of this distribution is σ (i.e. is We can test these hypotheses by probability
maize yield from plots onto which no fertilizer x = 12.3 (yield from plot treated with fertilizer)
yield of maize from the plot onto which the σ = 1 (standard deviation)
42
Now if the observed deviation is significant
z = (12.3 − 10) = 2.3
1 one is confronted by one of the following
alternatives, viz:
are saying is that most often (98.3% of the “Is it not a remarkable coincidence that our
times) there must be other explanation for this particular observed deviation is abnormal?”
abnormally high yield other than mere chance. However, it must be remembered that the
researcher is usually not doing the test in
In this case, we would say that the observed vacuum, .i.e., there are usually possible
maize yield of 12.3 tonnes per hectare is reasons why H0 is untrue, and we may thus
(statistically) different from the mean yield of prefer to disbelieve the coincidence and reject
10 tonnes per hectare at 1.07% level. H0. In other words, if there is some unique
special condition which could account for the
The test we have done above is called the z- significant deviation, it is adopted as the
test (normal test) since it uses the probability reason why H0 is untrue. For example in this
43
12.4.1 Type I errors α = 0.001 or 0.01%
A type I error is rejecting a null hypothesis that
is true .i.e. a false rejection of H0. The Criteria for Rejection of H0 based a given
probability of committing type I error is α value and calculated probability
denoted by Greek letter α and this correspond 1. If the p-value calculated is greater
to (100 α )% significance level. Thus for than α (p-value > α ), then we fail to
example with α =0.05, there is at most 5% reject H0 and declare the results to be
chance of wrongly rejecting H0 when H0 is insignificant.
true. 2. If the p-value is less or equal to α
(p-value ≤ α ), then we reject H0 and
12.4.2 Type II errors declare the results significant
This is the incorrect non-rejection (acceptance)
of H0 when H0 is false. The probability of Example 12.2
committing type II error is denoted by Greek In example 12.1, we calculated the probability
letter β (P (Type II error) = β ). of obtaining the maize yield of 12.3 tonnes per
hectare by chance from a plot not treated with
Take Note fertilizer to be 0.017.
We may want to keep both errors as low as
a) Is this significantly different at
possible but unfortunately, this is not possible
as you try to reduce Type I error, Type II error α =0.05
increases so the best thing to do to fixed Type I
error. b) Is this significantly different at
α =0.01
At α =0.05; p-value < 0.05 so we reject H0
12.4.3 The power of the test
and declare the results significant
The probability of correct rejection of H0 is
At α =0.01; p-value > 0.01 so we fail to reject
given by 1- β and this is referred to as the
H0 and declare the results insignificant
power of the test. Large value of the power of
the test indicates a good test.
12.6 Two-tail verses One-tail Test of
12.5 Levels of Significance
Significance
Before performing hypothesis testing, there is
In stating our hypothesis, sometimes we are
need to set up criterion for rejecting or
interested in particular direction of the
accepting H0 just like teachers set pass mark
outcome of the study. For example if you are
for exams. In hypothesis testing, this criterion
testing the effect of fertilizer application on the
is referred to as significance level α . We have
yield of maize as a scientist, you would expect
already learnt that α is associated the the plots treated with fertilizer to perform
probability of committing type I error. You better than the ones not treated and this will
will often see in research papers and statistical affect how you will state your alternative
books that H0 was rejected at significance level hypothesis.
α =0.05 etc. The most commonly used
significance levels are: 12.6.1 Two-tail test of Significance
α = 0.05 or 5% Let us once again revisit our example of
α = 0.01 or 1% fertilizer application on maize. Let µF be the
44
population mean yield of maize from plots H0 µ F = µ NF
which received fertilizer and let µ NF be the
H1 µ F < µ NF .i.e. the researcher is
population mean yield of maize from plots
interested in the negative deviation only
which did not receive fertilizer. If the
(lower tail of the
researcher is just interested in testing the
distribution)
prediction that there will be differences
The above two cases (a & b) constitute what is
between yield from plots treated with
referred to as one-tail hypothesis test of
fertilizers and those which were not treated,
significance.
our hypotheses will appear as:
H0 µ F = µ NF LECTURE THIRTEEEN:
12.6.2 One-tail test of significance In lecture 12, we showed how to use the z
If the researcher is only interested in the value and its associated probability to test the
a) The plots treated with fertilizer the normal test). The z-test works on
perform better than those not treated assumptions that the population variance is
known or the sample size is large (n>30). If
the sample size is small (n<30) and the
H0 µ F = µ NF
population variance is not known then another
H1 µ F > µ NF .i.e. the researcher is test called the student’s t-test or simply t-test is
interested in positive deviation only used.
(upper tail of the
distribution) 13.1.1 The Normal test (z-test)
b) The plots treated with fertilizers The normal test (z-test) is use when the sample
perform worst than those not treated size(s) is large and the population variances are
estimated by the sample variances or when the
population variance is known.
45
Instead of going through the tedious process of
calculating the z-values and then looking up
the associated probability, we can use the z-
value calculated directly to test our hypothesis.
Depending on the type of hypothesis under test
(one or two tails), the calculated value of z is
then compared as follows:
46
Case 1:
If x is N ( µ , σ 2 ) .i.e. with known In case 2 the z-value is calculated using this
2
variance σ , we can use the standard normal formula
table to decide whether an observed value of x x − µ0
z=
is significantly different from some
σ/ n
hypothetical value, µ 0 say.
Example 13.1
Procedure:
A group of 60 male students from Makerere
a) State the hypothesis H0 µ = µ0
University have a mean weight of 70.41 kg
H1 µ ≠ µ0 and estimated variance of 6.05 kg2. Records of
b) Calculate the probability that a value similar British students show a mean of
as extreme as x or more could have 70.0kg. Assuming that weight is normally
occurred by chance. Use the formula distributed, are the local students different in
weight?
x − µ0
z= to standardised the value
σ
Solution
of x and then use the z value to get
In this case we are comparing the weight from
the probability required
a sample of 60 Makerere University students
c) Decide whether the probability (or z-
value) calculated is small (big) with some hypothetical value µ 0 =70.0 kg
enough to reject H0. (see section H0 µ = 70.0 verses H1 µ ≠ 70.0 ;
12.5)
x ~ N (70,6.05 / 60 )
Case 2:
Now let us consider a sample of size n from a x − µ0 70.41 − 70
Then z = = = 1.29
population which is normally distributed with σ/ n 6.05 / 60
mean µ and variance of σ2 .i.e. σ 2 is For us to complete the hypothesis testing we
need to calculate the probability associated
known.
with this z-value
Let x be the mean from this sample of size n.
P(Z>1.291) = 1- P(Z<1.29)
From the sampling distribution (section 11.3)
= 1-0.9015 = 0.0985
we learnt that the variance of the sample mean
σ2 σ
( x ) is and its standard deviation is . Conclusion: since p-value > 0.05 (.i.e.
n n 0.0985>0.05), we accept H0 at significance
In this case we can also use standard normal
level of 5% ( α = 0.05 ) .i.e. there is no
tables to decide whether the sample ( x ) is
evidence to suggest that there is difference in
significantly different from some hypothetical
weight between the local and the British
value µ 0 say. students.
The procedure for testing this hypothesis is the
same as the one followed in case 1 but the only BUT we could have also just used the fact that
except is the formula for calculating the z – the z-value (1.29) we have calculated fall in
value.
47
the range (−1.96 < z < 1.96) which also ( x1 − x2 ) − 0
Applying equation, z=
leads to acceptance of H0 σ 12 σ 22
+
n1 n2
CASE 3:
we get:
In this case, we have two populations that we
need to compare. Consider two samples of size
2
n1 and n2 from populations x1 ~ N ( µ1 , σ 1 )
2 174.46 − 162.23
and x 2 ~ N ( µ 2 , σ 2 ) respectively. Let x1 z= = 45.94 * *
47.6522 43.7625
and x 2 be sample estimates of µ1 and µ 2 +
1164 1456
respectively. Conclusion: Since z-value calculated is
greater than 2.326, Reject H0, the men are
Consider H0: µ1 = µ2 alternatively, H0: significantly taller than the women at
α =0.01.
µ1 − µ 2 = 0
H 0: µ1 = µ 2 H A: µ1 > µ 2 (.i.e. a one-tail unknown and the sample size is small (n<30).
Consider, H0: µ = µ0
We now calculate a statistic t by the formula
CASE 5:
x − µ0 Just like in CASE 3, In this case, we also have
tf =
s two populations that we need to compare.
n
Consider two independent samples of size n1
Where f is the degree of freedom of
2
denominator, in this case (n-1) and n2 from populations x1 ~ N ( µ1 , σ 1 )
2
After calculating the test statistic t f most often and x 2 ~ N ( µ 2 , σ 2 ) respectively. Let x1
referred to as “t-calculated”, we compare this and x 2 be sample estimates of µ1 and µ 2
value with a t-value from a “t-table” which we respectively. However, in this case, the
will refer to as “t-tabulated”.
population variances σ 12 and σ 22 are not
f1 s12 + f 2 s 22
Thus s 2pooled =
f1 + f 2
49
Some important results In both case we calculate the‘t’ statistic using
It can be proved that if the two samples are this formula
independent then:
i. Variance of ( x1 − x2 ) = ( x1 − x2 ) − 0
tf =
SE ( x1 − x2 )
σ 12 σ 22
n1σ 1 + n2σ 22
2
+ =
n1 n2 n1 n2
Where SE( x1 − x2 ) is the square root of the
ii. Variance of ( x1 − x2 ) =
variance calculated from one of the follows
n +n
2 2 2 2
σ ( 1 2) if ( σ 1 =σ =σ 2
) above. The degrees of freedom f = (n1 + n2 -
n1n2
2). x1 and x 2 are sample estimates of µ1 and
2
σ
iii. Variance of ( x1 − x2 ) = 2 if µ 2 respectively
n
( n1 = n2 = n and σ 12 = σ 22 = σ 2
In general, SE( x1 − x2 ) =
)
2 ⎛ S 2 2 ⎞
Since population variance ( σ ) are rarely ⎜ pooled + S pooled ⎟
known it is replaced by its sample estimate ⎜ n1 n2 ⎟
⎝ ⎠
2
S pooled
Example 13.3
If we are comparing two groups usually the Imagine we measure length of shoot of tree
null hypothesis is stated as: species A and B in a given forest. The sample
H 0: statistics are:
Sample A Sample B
µ1 = µ 2 nA = 6 nB = 8
alternatively as; H0: x A = 74.8 xB = 72.99
µ1 − µ 2 = 0 SA, = 1.04 SB = 1.48
2
S A = 1.08 s B2 = 2.20
Where µ1 and µ 2 are the true means of the
Test the hypothesis that plant species A has a
respective populations.
longer shoot compared to species B
50
Table 13.1 Amount of α in one-tail
Thus tf =
LECTURE FOURTEEN: CORRELATION
( x1 − x2 ) − 0 (74.8 − 72.99) − 0 1.81
= = = 2.546
SE ( x1 − x2 ) SE ( x1 − x2 ) 0.711 14.0 INTRODUCTION
Previously we have studied only a single
variable at a time .i.e. for a give population or
sample we have just been studying one
The degrees of freedom for test f = (nA + nB -
2) = (f1 + f2)= (6 + 8 - 2) = 12 variable at a time. However, in many research
whether in agriculture, biology or social
To get the “t-tabulated” value from the t-table, sciences it is frequently necessary to study
we need to know the df and α
more than one variable measured on an
In this case our df = 12, α = 0.05 individual in order to get the whole picture. In
The t-tabulated can be represented as t 0.05,12 this lecture, we are turning our attention to
showing the significance level and the degrees studying relationship between two variables.
.i.e. plant species has significantly longer shoot lecture, we will only look at linear relationship
to reject H0.
Let us consider a situation where you sampled
10 students from your class and then you
51
record the height (x) and weight (y) of each completed on the vertical scale (y-scale). Then
student. This sample is called a bivariate each of the pairs (x, y) of observations is
sample and the data generate is a bivariate plotted as a point in the xy-plane.
data. The 10 pairs of simultaneously sampled
Scatter plot of sleep deprivation versus completion
values are designated (x1, y1), (x2, y2), (x3, of task
y3)…(x10, y10).
20
15
When the bivariate data are quantitative, we
Tasks
can use what is called a scatter plot to visual 10
14.2 Co-relationships
Example 14.1
From Example 14.1, we have seen that there
To study relationship between sleep
seems to be an association between sleep
deprivation and ability to complete a simple
deprivation and the number of tasks
task, 12 people were asked to solve a simple
completed. In many agricultural and biological
task after having been without sleep for 15, 18,
investigations, it is important to measure the
21 and 24 hours. The data were recorded as
strength (intensity) of co-relationships between
follows:
two variables x and y say. By intensity of
Subject 1 2 3 4 5 6 7 8 9 10 11 12
correlation is meant the extent to which
Hours 15 15 15 18 18 18 21 21 21 24 24 24
without deviations from the mean in one variable (x -
sleep µx ) tend to be accompanied by proportional
Tasks 13 9 15 8 12 10 5 8 7 3 5 4
completed deviations in the other variable (y - µ y ).
Source: Exploring Statistics 2nd Edition by Perfect correlation occurs when deviations in
Larry J. Kitchen one variable are exactly proportional to
Use the information to construction a scatter deviations in the other variable.
plot
52
Also for some data sets, the covariance can get
2 very large and become difficult to interpret.
Variance of x = s 2
=
∑ (x − x) and for y
x
n −1
2 Take Note
= s 2
=
∑ ( y − y)
y o Covariance cannot be used to show non-linear
n −1 relationship
s xy =
∑ (x i − x )( y i − y ) its values lies between +1 and -1.
n −1
The sample Pearson correlation coefficient
As the name suggests, the covariance is a denoted by r, is given by the formula below
measure of how two variables vary together.
The covariance measures the linear
∑ (x i − x )( yi − y )
dependence between the two variables. If there n −1
rxy =
is no linear dependence (association) between
∑ ( xi − x ) 2 * ∑ ( y i − y ) 2
the two variables, then the covariance will be
n −1 n −1
zero. From the formula we can see clearly that
the covariance dependent on the scale of
The above formula can be simplified by the
measurement for the two variables x and y and
removal (n-1)
thus it may be not be very useful for measuring
association. For example, values of variable x rxy =
∑ (x i − x )( y i − y )
∑x ∑y
i i 2 3.4 -1 0.5 -0.543
∑ (x i − x )( y i − y ) = ∑ xi y i −
n 3 2.4 0 -0.457 0.00
4 2.0 1 -0.857 -0.857
5 1.2 2 -1.657 -3.314
Example 14.2
6 1.0 3 -1.857 -5.571
A student from the faculty of Economics and
Total -21.1
Management, Makerere University wanted to
study the relationship between house rent and
distance from campus. She selected seven
Covariance
locations north of the campus and recorded the
distance from campus and house rent for a s xy =
∑ (x i − x )( y i − y )
=
− 21.1
= −3.5167
two-room house at each location. The follow
n −1 7 −1
distance and house rents are recorded.
n −1 s xy − 21.1
correlation between the two variables rxy = = = = −0.978
( x x
∑ i *∑ i
− ) 2
( y − y ) 2
s 2
x * s 2
y
4.667 * 2.772
Solution
The data appear to fall on a straight line, n −1 n −1
suggesting a strong linear relationship.
Scatter plot for distance from Makerere University This is a very high negative correlation
and House rent
coefficient indicating that there is a strong
6
5
linear relationship between the distance and
House rent (00000)
We need to calculate the statistics for each a) The correlation coefficient (r) measures
variable before we can calculate the correlation the degree of linear dependence or
54
(see figure 14.1c) BUT, zero correlation confirm the relationship. In this lecture, we are
does not (necessarily) imply going to look at yet another method for
independence, since the variable could be studying the relationship between two
non-linearly related (see figure 14.1d) quantitative variables. Regression analysis
c) The closer r is to either -1 or +1, the aims at formulating a mathematical
stronger the linear relationship between x relationship between two variables. In general,
and y. If all points fall on a straight line, our objective would be to express one variable
then we have a correlation of +1 if the y as a function a second variable x.
line has a positive slope and -1 if the line
has a negative slope (see figure 14.1 a & 15.1 Classification of variables in
b). Regression Analysis
d) In discussing correlation the following Unlike correlation, with regression analysis we
terminology is often used must distinguish between the response
o | r | < 0.3 - Weak correlation (dependent) variable and the predictor
o | r | < 0.5 – Moderate (independent) variable.
o | r | > 0.7 – Strong correlation
Figure 14.1: Correlation coefficient for
different relationships
55
15.2 Regression Equation
Regression analysis just like correlation is used We can therefore use the above regression
to study linear relationship between two equation for predicting the rent at a given
variables x and y. The first step is to draw a distance from campus. Of course, we do not
scatter plot after which you calculate r. expect our regression equation to predict the
Regression analysis is aimed at getting the rent exactly. We expect some error between
equation of the straight line that pass through the predicted value and the real (observed)
the data in the scatter plot. value since our regression line does not pass
The equation of the straight line is given by: through all the points.
y = b0 + b1 x
Let y i be the observed or recorded value of
Where b0 is the y-intercept and b1 is the slope
of the straight line, the y-intercept is the value rent at distance i from campus and let ŷ i be
of y when x = 0, and the slope gives the the predicted value of that observed rent. The
change in y relative to the change in x. predicted value ŷ i is given by:
yˆ i = bo + b1 xi
Example 15.1
Let us revisit example 14.2 (Relationship The error of prediction of rent at the ith is given
4
3 Solution
2
The regression equation is given as:
1
0 y = 5.1179 − 0.7536 x
0 1 2 3 4 5 6 7
For each distance, the prediction is got by
Distance (Kilometers)
substituting the appropriate value of x in the
56
For x = 3;
b1 =
∑ ( x − x )( y − y )
i i
or
y3 = 5.1179 − 0.7536 * 3 = 5.1179 − 2.2608 = 2.8571 ∑ (x − x) i
2
x y
Prediction error
∑x y − ∑ ∑
i i
n
i i
xi yi ŷ i ei = yi − yˆ i
2 ( ∑ x) 2
0 5.1 5.1179 -0.0179 x
∑ i −
n
1 4.9 4.3643 0.5357
2 3.4 3.6107 -0.2107
bo = y − b1 x
3 2.4 2.8571 -0.4571
Example 15.3
15.3 Estimation of Regression Equation
Now that we seen how useful a regression Find the least squares regression line
equation can be for prediction purposes, one yˆ = bo + b1 x for rent data in Example 14.2
would then ask how do we estimate the
parameters of the regression equation? From x y (x - (x - (y - (x -
the scatter plot, you notice that it is possible to x) x )2 y) x )(y -
draw many lines through the data each of y)
which with its own equation. We need to select 0 5.1 -3 9 2.243 -6.729
one of these lines that will give the best
1 4.9 -2 4 2.043 -4.086
prediction .i.e. the one that will give us the
2 3.4 -1 1 0.5 -0.543
smallest prediction error. The method of
3 2.4 0 0 -0.457 0.00
selection of the regression line that minimizes
4 2.0 1 1 -0.857 -0.857
the “prediction error” ( ei = yi − yˆ i ) is 5 1.2 2 4 -1.657 -3.314
referred as the least squares method. The 6 1.0 3 9 -1.857 -5.571
least square method will select the line that Total 28 -21.1
minimizes the sum of the squared error
2
( ∑e i
). In this lecture, we are not going into