WEEK 1.
1 NOTES
Introduction
Definition of terms
Statistics is the study of how to collect, organize, analyze and interpret numerical
information (data).
The main divisions of statistics are;
Descriptive statistics – refers to the statistical methods of describing and
organizing datag mean, median, mode, range, variance and standard deviation.
Inferential statistics – refers to the methods of using a sample to obtain
information about a population or making conclusions about a population from
the sample. It is therefore, important to remember that the main role of inferential
statistics is to draw conclusions about a population based on information
obtained from a sample.
Population – refers to all measurements of interest. For example the weights of all
pineapples in a field.
Sample – is simply a representative part of the population. For example, 100
weights of pineapples.
Importance of Statistics
Statistical background is important in understanding research reports within your
area of interest. This is due to the fact that:
1. Statistics permits summarization and presentation of large quantities of
information in a manner that facilitates its
2. Statistics enables the researcher to extend research beyond the restricted setting
in which most research is actually
Statistics enables the formulation and testing of hypothesis. Statistical procedures in
research are used in diverse fields like:
Behavioural sciences (education, psychology, and sociology)
Agriculture
Economics
Medicine
Biology
For example, statistics enables the education sector to draw conclusions concerning
the efficiency of various instructional methods; medical scientists to choose the most
effective medicine etc.
Limitations of statistics
Despite the universal approach of statistics, it has its own limitations:
1. Does not study qualitative phenomena which cannot be expressed in figures like
honesty, beauty, color of hair
2. Does not reveal the entire study of the Some problems have other background
factors e.g. religion, which cannot be covered by statistic.
3. Does not study individuals.
4. Liable to misuse – any person can misuse statistics by the type of conclusions
they make g. opinion polls, incidents of crime employment, census.
5. Statistical laws are true only on Conclusions arrived at one not entirely accurate
and the same conclusions cannot be arrived at under similar conditions at all
times.
Data Variables
Qualitative variables – no numerical values can be assigned e.g the colour of hair.
Quantitative variables – have numerical values, e.g. the number of people in a
Sunday service. Quantitative variables can be either discrete or continuous.
A discrete variable – is one which naturally inherently contains gaps between
successive observation e.g. the number of family members in a family (it is either 2,
3, 4 etc but not 3.5 countries, number of chairs, number of animals, etc.
A continuous variable has the property that between any 2 observable value lies
another observable value.
A continuous variable takes values along a continuum, i.e along a whole interval of
values, e.g. length, weight, height, temperature, speed, PH of a solution etc.
An essential attribute of a continuous variable is that unlike a discrete variable, it can
never be measured exactly.
WEEK 1.2 NOTES
COLLECTION OF DATA
It may be very difficult and tedious to collect data as the population may have a
number of characteristics and may be you are interested in only one or two
characteristics e.g. pineapples in a field, required to estimate the mean weight of the
pineapples.
Data may be obtained through the following methods:
a) Survey
The most common method of collection of data is through surveys which have to be
conducted very carefully otherwise, the results would be of no value. Survey consists
of two parts:
i) Planning
You need to consider:
Nature of
Objective and scope of enquiry
Source of information
Type of enquiry to be conducted
Accuracy
1. ii) Executing the plan - involves various steps to put the plans in
Setting up an organization
Selecting and training field staff
Supervision of field work
Analysis of data
Preparation of a report
Surveys can be done by using a variety of methods. The choice of the method
depends on a number of factors.
Factors affecting the choice of method of survey;
Nature, objective and scope of enquiry – method selected should be such that it suits
the type of enquiry that is being conducted.
Availability of finance – when financial resource are scarce leave a side expensive
methods.
Availability of time – some methods involve a long duration of time while others takes
a shorter time. The time allocated to the enquiry affects the selection of the method.
Methods of survey
The three most common methods are:
i) Personal interview survey
Investigation has to collect the information personally from the source concerned. It
has the advantage of obtaining the depth responses to questions. But the interview
must be trained in asking questions and recording responses, which makes it costly.
There is also the possibility of bias by the interviewer in the selection of respondents.
ii) Telephone surveys
They are less costly than personal surveys. Also people may be more – open since
there is nio face to face contact. It has a disadvantage in that. Some people do not
have telephone or networks not good. The office or home when calls are made
hence nt interviewed.
iii) Mailed questionnaire surveys:
They can be used to cover a wider geographical area. Also, the respondents can
remain anonymous if desired. However, it has low response, in appropriate answers
as well as difficulty in reading or understanding the questions by some people.
b) Sampling methods
This is used by researcher to collect data and information about a particular
character from a large population. Using samples saves time and money as well as
it allows the researcher to get more information about a particular subject.
In order to obtain samples that are unbiased, each subject in the population is given
an equal chance of being selected.
The basic methods of sampling are;
1. Random sampling – is a sample chosen from a population in such a way as to
ensure that every individual in the population has equal chance of being included.
[Link] sampling - samples are obtained by numbering each subject of the
population and then selecting every nth However the first subject would be selected
at random.
[Link] sampling – samples are obtained by dividing the population into groups
called strata according to some characteristic important to the researcher
[Link] sampling – we divide the population area into sections (clusters)
randomly select a few of these sections and then choose all the members from
selected sections
WEEK 2 NOTES
ORGANISATION OF DATA
This deals with the basic techniques of organization and representation of data in a
way that is simple to understand and analyze.
a) REPRESENTATION OF DATA
The main purpose of representation of statistical data is to make collected data more
easily understood. Methods commonly used in representation are:
1. Bar graphs
2. Pie charts
3. line graphs
4. Histograms
5. Frequency polygons
6. Stem and leaf diagram
Bar graphs
A bar graph consist of a number of spaced rectangle which generally have major
axes vertical. Bars are of uniform width. The axes must always be labeled and
scales indicated. For example, the John’s performance in five subjects was as
follows:
Maths 80
English 75
Kiswahili 60
Biology 70
Geograph 75
This information can be represented on a bar graph as shown below.
Pie chart
A pie chart is a circle divided into various sectors. Each sector represents a certain
quantity of the item being considered. The size of the sector is proportional to the
quantity it represents.
Consider the total population of animals in a farm given as 1800 out of these 1200
are chicken, 200 cows, 300 goats and 100 ducks. This information can be
represented in a pie chart as follows:
Line graphs
In line graphs, data is represented using lines. Examples of line graphs are a graph
showing temperature of a patient at various times of the day, and a travel graph.
Consider the following table which shows the temperature of a patient at different
times of the day:
Time [Link] 9a.m [Link] 11 a.m 12 noon 1pm 2 pm
Temp (0c) 35.6 36.4 37.0 37.2 36.8 35.9 37.1
The data can be represented in a line graph as below.
WEEK 3 NOTES
Histograms
In this, the frequency in each class is represented by a rectangular bar whose area is
proportional to the frequency.
When, the bars are of the same width the height of the rectangle is proportional to
the frequency. Note that the bars are joined together.
For example, the following data represents masses to the nearest kilogram, of fifth
caught by a fisherman in a day.
Mas (kg) 5-9 10-14 15-19 20-24 25-29 30-34 35-39 40-44
No. of fish 6 20 12 10 5 6 2 1
The above data can be represented in a histogram as shown: Write the classes
interval of true limits or boundaries.
14.5- 19.5- 24.5- 29.5- 34.5- 39.5-
Mas (kg) 4.5-9.5 9.5-14.5
19.5 24.5 29.5 34.5 39.5 44.5
No. of
6 20 12 10 5 6 2 1
fish
Note
The class boundaries mark the boundaries of the rectangular bars in the histogram.
Histograms can also be drawn when the class, intervals are not the same.
For example, the following table shows the marks obtained in a math test.
Marks 10-14 15-24 25 -29 30 -44
No. of students 5 16 4 15
Marks 9.5-14.5 14.5-24.5 24.5-29.5 29.5 – 44.5
No. of students 5 16 4 15
Frequency 16/10 =
5/5 = 1 4/5 = 0.8 15/15 = 1
density 1.6
Frequency
density 10 16 8 10
×10
Area A1 = 5×10 = 50
Area A2 = 10 ×16 = 160
Area A3 = 5 ×8 = 40
Area A4 = 15 ×10 = 150
Note that the area of each bar is proportional to its frequency.
Frequency polygon
When the midpoints of the tops of the bars in a histogram are joined with straight
lines, the resulting graph is called a frequency polygon.
You can obtain a frequency polygon by plotting the frequency against the mid-points
of the classes.
Consider, this data which represents masses, to the nearest kilograms of fish
caught by a fisherman in a day.
Mas (kg) 5-9 10-14 15-19 20-24 25-29 30-34 35-39 40-44
No. of fish 6 20 12 10 5 6 2 1
The frequency polygon for this data is
Note:
More than one frequency polygon can be drawn on the same axes for the purpose of
comparison. For example, in an examination the marks in mathematics and physics
were grouped as given below. We can draw frequency polygon to represent the data
given:
Marks 10 -19 20 -29 30 – 39 40 – 49 50 – 59 70 – 79 70 -79
Frequency Maths 2 6 7 13 6 4 2
Physics 3 9 10 6 9 2 1
Note
From the above frequency polygon, we notice that:
39. There were more students in physics than in mathematics with less than 39.5
40. There were more students in mathematics than in physics with more than 39.5
This implies that more students had relatively higher marks in mathematics than in
physics.
From the two observations, we can conclude that the performance in mathematics
was better than that in physics
Stem and leaf plot
2 7
3 2
4 1 3 3 4 7 7 8
5 0 1 1 2 3 3 344456689
6 888
7 388
8 5
Key: 2 7 = 27
To construct a stem and leaf plot/ display, the first digit of each response/ number is
listed in a column at the left; and the associated second digits are listed to the right
of each first digit.
For example
Integrated circuit response times
4.6 4.0 3.7 4.1 5.6 4.5 6.0 6.0 3.4 3.4 4.6 3.7 4.2 4.6 4.7 4.1 3.7 3.4 3.3
3.7 4.1 4.5
4.6 4.4 4.8 4.3 4.4 5.1 3.9
The stem and leaf display would be
3 744774379
4 60115626711564834
5 61
6 00
key 3/7 = 3.7
In this display/ plot, each row is a stem and the numbers in the column to the left of
the vertical line are called stem labels. Each number in a stem to the right of the
vertical line is called a leaf.