TOPIC ONE: DESCRIPTIVE STATISTICS
1.1 Definition of Statistics
Statistics is a science that deals with the methods of collecting, organizing, presenting, analyzing
and interpretation of numerical data to assist in making more effective decisions.
Types of Statistics
a) Descriptive statistics: It deals with processing data without attempting to draw any
inferences from it. It refers to the presentation of data in the form of tables and graphs and to
the description of some of its features such as averages.
b) Inferential/Inductive statistics: Refers to methods of using a sample to obtain information
about a population i.e., making conclusions about the population based on information from
the sample.
Population, Sample and Variables
Population: is the totality of all the items or individuals whose characteristics we wish to
study. Examples of a population are all the eligible voters in an election.
Sample: is a subset or section of the population that is used to represent the whole population.
Parameter: is any quantitative measure that describes a characteristic of a population e.g.,
2
population mean (µ) or population variance ( ). Population Standard deviation ()
Statistic: is a quantitative measure that describes a characteristic of a sample e.g. sample mean
(x) or sample variance (s2). Sample standard deviation (S)
E.G. The mean height of the people in Kenya is a parameter, whereas the mean height of a
sample of 500 people is a statistic.
Variable: is the characteristic that is being studied. It is represented by symbols X, Y, or
Z. Height of people, grades in a test etc are examples of variables.
There are two kinds of variables:
a) Qualitative variables: Are variables that are non-numeric i.e., attributes e.g., Gender,
Religion, Color, State of birth etc.
b) Quantitative variables: are numeric variables e.g., the height of an individual when
expressed in feet or inches, etc. Quantitative variables are either discrete or continuous.
1
i) Discrete variables: Are variables, which can only assume certain values i.e., whole
numbers. Are always counted. E.G: number of children in a family, the number of
defective bulbs, etc.
ii) Continuous variables: Are variables, which can assume any value within a specific
range. Are always measured e.g., height, temperature, weight, radius etc.
Exercise 1
1. Explain the functions of Statistics
2. Discuss the applications of Statistical Knowledge in Business Management
3. Explain any four limitations of statistics in decision making
4. Misuse of statistics may take several forms. Explain four ways in which statistics may be
misused
1.2 Data Collection, Organization and Presentation
Data refers to any information or facts collected for reference or analysis.
There are two types of data: secondary data and primary data
Secondary Data
Its data that been gathered earlier for some other purpose
E.G: the demographic statistics collected every ten years are the primary data with the
registrar of persons but the same statistics used by anyone else would-be secondary data with
that individual.
Primary Data
Primary data are measurements observed and recorded as part of an original study.
It is data that are collected first hand by someone specifically for the purpose of facilitating
the study are known as primary data.
2
Organization and Presentation of Data
Data collected in an investigation and not organized systematically is called raw data. The
arrangement of this data in ascending or descending order of magnitude is called an ordered
array.
The difference between the largest and the smallest value is called the range.
E.G: The data below records the heights, in inches, of eight students.
66, 68, 72, 65, 66, 73, 68, 69
Arrange the data into an array
65, 66, 66, 68, 68, 69, 72, 73
Frequency Distributions
Ungrouped data
In forming an array, a value is repeated as many times as it appears. The number of times a
value appears in the listing is referred to as its frequency. In giving the frequency of a value,
we answer the question, “How frequently does the value occur in the listing?”
When the data is arranged in tabular form by giving its frequencies, the table is called a
frequency table. The arrangement itself is called a frequency distribution.
Example:
The following data were obtained when a die was tossed 30 times. Construct a frequency table.
1 2 4 2 2 6 3 5 6 3
3 1 3 1 3 4 5 3 5 3
5 1 6 3 1 2 4 2 4 4
Outcome Frequency Cumulative
(X) (f) Frequency (CF)
1 5 5
2 5 10
3 8 18
4 5 23
5 4 27
6 3 30
Total 30
Grouped Data
When dealing with a huge mass of data and when the observed values consist of too many
distinct values, it is preferable to divide the entire range of values and group the data into
classes.
E.G: If we are interested in the distribution of ages of people, we could form the classes
3
0 – 19, 20 – 39, 40 – 59, 60 – 79 and 80 – 99. A class such as 40 – 59 represents all the
people with ages between 40 and 59 years inclusive.
When data are arranged in this way, they are called grouped data. The number of
individuals in a class is called the class frequency.
The following set of steps are suggested to form a frequency distribution from the raw data
i) Range
Scan through the raw data and find the smallest and the largest value. The largest value minus the
smallest value gives the range.
i) Number of classes
Decide on a suitable number of classes. This could be anywhere from six to twenty.
ii) Class size
Divide the range by the number of classes. Round this figure to a convenient value to
obtain the class size and form the classes.
iii) Frequency
Find the number of observations in each class.
Example
The following data gives the amounts (in dollars) spent on groceries by 40 housewives during a
week.
22 12 9 8 33 32 30 33 8 11
21 16 12 15 37 30 16 22 12 24
18 25 37 16 25 28 25 18 9 28
25 28 26 15 12 35 38 16 24 31
Construct a frequency distribution using seven classes.
Class Interval Frequency (f) CF
8 – 12 8 8
13 – 17 7 15
18 – 22 5 20
23 – 27 8 28
28 – 32 7 35
33 – 37 4 39
38 – 42 1 40
Total 40
4
Class Intervals, Class Marks and Class Boundaries
The blocks 10 – 20, 20 – 30, 30 – 40, etc are called class intervals. The lower ends of the
class intervals are called lower limits and their upper ends are called upper limits.
The number of values specified in a given interval is called its length or width or magnitude.
E.G: The class 1 – 3 has values 1, 2, 3 thus its length is 3.
The class 5 – 9 has values 5, 6, 7, 8, 9; the length or magnitude is 5
There are two types of classes
i) Inclusive type: These are of the type 5 – 9, 10 – 14, 15 – 19, … where both the
upper- and lower-class limits are included in a given class.
ii) Exclusive type: These are of the type 5 – 10, 10 – 15, 15 – 20, … where the upper-
class limit of a given class is the lower-class limit of the next class.
The class 5 – 10 has values 5, 6, 7, 8, 9 and the class 10 – 15 has 10, 11, 12, 13, 14.
NB: The conversion of inclusive type of classes to exclusive type is useful in calculating
certain measures such as mode and median.
A point that represents the halfway or dividing point between successful classes is called a
class boundary. If d is the difference between the lower-class limit of a given class and the
upper-class limit of the succeeding class, then
1
Upper Class Boundary (UCB) = Upper Class Limit (UCL) + d
2
1
Lower Class Boundary (LCB) = Lower Class Limit (LCL) - d
2
The class mark is defined as the midpoint of a class interval. It is computed by adding the
lower and upper class limits of a class and then dividing by 2.
Mid point1=UCB LCB
1=
UCB LCB
2
5
Class Limits Class Boundaries Class Mark (x) Frequency (f)
8 – 12 7.5 – 12.5 10 8
13 – 17 12.5 – 17.5 15 7
18 – 22 17.5 – 22.5 20 5
23 – 27 22.5 – 27.5 25 8
28 – 32 27.5 – 32.5 30 7
33 – 37 32.5 – 37.5 35 4
38 – 42 37.5 – 42.5 40 1
6
Graphical Representation of a Frequency Distribution
The following types of graphical representation are usually used for frequency distribution.
a) Histogram: It is a graph in which classes boundaries are marked on the horizontal axis
and the class frequencies on the horizontal axis and the class frequencies on the vertical
axis. The class frequencies are represented by the heights of the bars and the bars are
drawn adjacent to each other.
b) Frequency polygons and Frequency Curve: A frequency polygon is a line graph where
we plot the class marks or midpoints along the horizontal axis and the corresponding
frequencies along the vertical axis. The class midpoints are connected with a line
segment.
If the classes are very many and the class widths are so small that the midpoints are close
together, the polygon can be formed by free hand to give a smooth curve known as a
frequency curve.
c) Cumulative Frequency Curve or the Ogive. An ogive is a line graph obtained by
representing the upper-class boundaries along the horizontal axis and the corresponding
cumulative frequency along the vertical axis.
Exercise 2
A random sample of 50 auto drivers insured with a company and having similar auto
insurance policies was selected. The following data shows monthly auto insurance
premium (in Kshs.000) paid by them.
54 40 45 20 60 30 35 40 55 70 20 15
45 60 45 25 15 30 25 18 35 25 45 56
59 25 27 39 50 56 20 25 30 30 41 25
56 48 45 25 35 60 55 48 38 34 60 60
60 64
i) Group the above data starting with the class 10 -20 exclusive
ii) Represent the data using a Histogram and an Ogive.
7
Class Interval Frequency (f) Cumulative Frequency (CF)
10 – 20 3 3
20 – 30 11 14
30 – 40 10 24
40 – 50 10 34
50 – 60 8 42
60 – 70 7 49
70 – 80 1 50