0% found this document useful (0 votes)
7 views59 pages

1 Quantitative Methods Note

The document provides an overview of quantitative techniques and spatial analysis in geography, focusing on the definition, classification, and application of statistics. It distinguishes between descriptive and inferential statistics, explains key concepts such as population, sample, and measurement scales, and outlines methods for data collection. Additionally, it highlights the objectives of the quantitative revolution in geography, emphasizing the importance of statistical techniques in enhancing the scientific rigor of the discipline.

Uploaded by

masredf81
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views59 pages

1 Quantitative Methods Note

The document provides an overview of quantitative techniques and spatial analysis in geography, focusing on the definition, classification, and application of statistics. It distinguishes between descriptive and inferential statistics, explains key concepts such as population, sample, and measurement scales, and outlines methods for data collection. Additionally, it highlights the objectives of the quantitative revolution in geography, emphasizing the importance of statistical techniques in enhancing the scientific rigor of the discipline.

Uploaded by

masredf81
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

Kotebe University of Education, Department of Geography and

Environmental studies

Quantitative Techniques & Spatial Analysis (GeES 2021)


1. Introduction
1.1 Definition and classification of statistics
Statistics is a science that deals about the methods of collection, organizing the collected
data, presentation, analysis and interpretation of the data. Kerlinger, (1986) defined it as
the theory and method of analysing quantitative data obtained from samples of
observations in order to study and compare sources of variation of phenomena, to help
make decisions to accept or reject hypothesized relations between phenomena, and to aid
in making reliable inferences from empirical observations.
It is a subject that deals with numbers and figures describing certain situations. It
primarily deals with numerical data taken by surveys and summarizes these data in such a
way that this summary gives a good indication about the nature of the data. The subjects
of statistics, is not a new discipline but it is as old as the human society itself. The sphere
of its utility, however, was very much restricted.
The word “statistics” is derived from Latin for “state” indicating the historical
importance of governmental data gathering, which related to demographic information
(military recruitment and tax collecting). Thus, the scope of statistics in the ancient times
was primarily limited to the collection of demographic, property and wealth data of a
country by governments for framing military and fiscal policies. Nowadays, statistics is
used almost in every field of study, such as natural science, social science engineering,
medicine, agriculture, e t c.
Classification: There are different classifications of statistical methods. A statistical
method can be classified as descriptive and inferential.
1. Descriptive Statistics
Descriptive Statistics deals with describing data without attempting to infer anything that
goes beyond the given set of data that consists of collection, organization, summarization
and presentation of data. Descriptive statistics are quantitative summaries of the
measurements made on a set of objects.
2. Inferential Statistics
Inferential Statistics deals with making inferences and/or conclusions about a population
based on data obtained from a sample of observations which consists of performing
Kotebe University of Education, Department of Geography and
Environmental studies
hypothesis testing, determining relationships among variables and making predictions. It
is the process of making an estimate, prediction, or decision about a population based on
a sample.

Parametric statistics are based on assumptions about the distribution of population from
which the sample was taken. Nonparametric statistics are not based on assumptions, that
is, the data can be collected from a sample that does not follow a specific distribution.

1.2 Definition of some basic terms


a) Population: is the totality (collection) of all objects or items under consideration. This
set of elements is called the universe. A set of values associated with these elements is
called a population of values. Example: If you want to study the mean age of Kotebe
University of Education Geography and Environmental studies students, all Kotebe
University of Education Geography and Environmental studies students constitute the
population of your study.
b) Sample: Is a part of a population taken so that some generation about the population
can be made. A sample should be representative of the population. Example: If you want
to study the mean age of Kotebe University of Education, Geography and Environmental
Studies students through measuring the age of some student, the selected ones constitute
your sample.
To understand the differences between a population and a sample, it is necessary to
consider numbers which are used to describe the two (Parameter and Statistics)
c) Parameter: It is a descriptive measure of a population, or summary value calculated
from a population. Examples: Arithmetic mean (symbolised as μ (Greek small letter mu),
Range, variance value of the population.
d) Statistics: It is a descriptive measure of a sample, or summary value calculated from a
sample. Example: Mean, Range, variance value of the sample.
1.3 Stages in Statistical Investigation
Data Collection: This is a stage where we gather information for our purpose.
Data Organization: It is a stage where we edit our data. A large mass of figures that are
collected from surveys frequently need organization. The collected data involve
irrelevant figures, incorrect facts, omission and mistakes.

2
Kotebe University of Education, Department of Geography and
Environmental studies
Data Presentation: The organized data can now be presented in the form of tables,
charts diagrams and graphs. At this stage, large data are presented in a very summarized
and condensed manner.
Data Analysis: This is the stage where we critically study the data. The purpose of data
analysis is to dig out information useful for decision making.
Data Interpretation: This is the stage where draw valid conclusions from the results
obtained through data analysis. If the data that have been analyzed are not properly
interpreted, the whole purpose of the investigation may be defected and misleading
conclusion may be drawn.
1.4 Application of statistics
The science of statistics is very essential for research and decision making processes in
all aspects of human life. The following are some of the areas for which statistical
analysis is required:
 To represent the facts in the form of numerical data
 To summarize a mass of data into a few presentable understandable and precise
figures.
 To Predict or forecast future trend.
 To help select a course of action among a number of alternatives.
 To help in formulating policies.

1.5 Types of variables and measurement scales


A variable is a characteristic of an object that can have different possible values.
There are two types of variables.
a) Quantitative variables: are variables that can be quantified or can have numerical
values. Quantitative data are the result of counting or measuring attributes of a
population. Amount of money, pulse rate, weight, number of people living in your town,
and number of students who take statistics are examples of quantitative data. Quantitative
data may be either discrete or continuous. All data that are the result of counting are
called quantitative discrete data. On the other hand, all data that are the result of
measuring are continuous data. If you and your friends carry backpacks with books in
them to school, the numbers of books in the backpacks are discrete data and the weights
of the backpacks are continuous data.

3
Kotebe University of Education, Department of Geography and
Environmental studies
b) Qualitative variables are the result of categorizing or describing attributes of a
population. Hair colour, blood type, ethnic group are examples of qualitative data.
Qualitative data are generally described by words or letters. For instance, hair colour
might be black, dark brown, light brown, blonde, grey, or red. Blood type might be AB,
O, or B.
Measurement scales
There are four types of measurement scales for variables:
1. Nominal scale: - “Nominal “is a Latin word for “name”. This is a scale for grouping
individuals into different categories. In such measures numbers are assigned just for
recognition. For example when you go out to collect data from males and females you
may decide to assign the number 1 to males and 2 to females. It does not mean that males
score lower or are inferior to females. Such measures are merely to identify, to name and
to differentiate males from females, one player from another. Nominal scale is the
simplest of all measurement scales. In this scale, one is different from the other and +, -,
*, /, impossible, comparison is impossible
2. Ordinal scale: - “ordinal” is a Latin word, meaning “order”. It is a scale for grouping
and ordering of individuals in to different categories. Data consisting of an ordering or
ranking of measurements are said to be on an ordinal scale of measurements. This scale
has magnitude only. Examples: military ranks, ranks in race, ranks of collage academic
staff, e t c. One is different from and grater /better/ less than the other and +, -, *, / are
impossible, comparison is possible.
Ordinal scales data contain and convey more information than the nominal scale data, for
relative magnitudes are known, however, quantitative comparisons are impossible. Let us
consider students’ performance in a test. Student A scored 80%, B had 75%, C 63% and
D 60%. Based on their performances, student A is first, B is second, C is third and D is
fourth. However, the difference between first and second position is 5%, between second
and third position you have 12% and between third and fourth position is 3%. In each
situation there is no equal interval.
3. Interval scale: This is a scale with magnitude and equal interval but no absolute zero.
There is no true zero point (arbitrary zero point not real) and there is no physical

4
Kotebe University of Education, Department of Geography and
Environmental studies
significance to the zero point. The zero value is defined conventionally since it does not
imply a lack of the quantity under study.
There is a constant interval size between any adjacent units on the measurement scale.
Example: 0C, 0F (Measuring units of temperature) which do not mean that there is no
heat.
In this measurement scale One is different, better/greater and by a certain amount of
difference than another (Possible to add and subtract but multiplication and division are
not possible)
37oC – 35oC = 2oC and 45oC – 43 oC= 2oC
40oC = 2(20oC) But this does not imply that an object which is 40 oC is twice as hot as an
object which is 20 oC. Interval scale data convey better information than nominal and
ordinal scale data.
4. Ratio scale: It is a measurement scale in which there is a constant interval size
between any adjacent units on the measurement scale. There exists a zero point on the
measurement scale and that there is a physical significance to this zero point. This means
that a measured value of zero really implies the absence of the phenomenon being
studied. For example, zero distance means no distance travelled and zero precipitation
means that none has been recorded.
Examples: height, weight, volume, etc.
One is different, larger /taller/ better/ less by a certain amount of difference and so much
times than the other and (+, -, *, / are possible on this scale).This measurement scale
provides better information than interval scale of measurement.
1.6 Sources of data and methods of data collection
Facts or figures, which are numerical or otherwise, collected with a definite purpose, are
called data. The required data can be obtained from either a primary source or a
secondary source.
Primary source: Is a source of data that supplies first-hand information for the use of the
immediate purpose. Primary data is the data that is collected for the first time through
personal experiences or evidence, particularly for research.
The term ‘primary data collection’ is generally used to refer to the strategy, in which
researchers collect information for a specific purpose ‘in the field’.

5
Kotebe University of Education, Department of Geography and
Environmental studies
The main advantage of this approach is that the quality of the data is known and better
understood, i.e., the researchers are better able to assess the effects of potential sources of
error in the data because they have been intimately involved in their collection and
recording.
Its main disadvantages are that it is often slow, arduous, labour intensive, and frequently
expensive. It may take months, if not longer, to collect enough data of the type required
for the study.
Secondary data is a second-hand data that is already collected and recorded by some
researchers for their purpose, and not for the current research problem. Secondary data:
are data collected from a secondary source.
Secondary source: are individuals or agencies, which supply data originally collected for
other purposes by them or others. Usually they are published or unpublished materials. It
is accessible in the form of data collected from different sources such as government
publications, censuses, internal records of the organisation, books, journal articles,
websites and reports, etc.
This method of gathering data is affordable, readily available, and saves cost and time.
However, the one disadvantage is that the information assembled is for some other
purpose and may not meet the present research purpose or may not be accurate.
Discrete Vs continuous data
Discrete data (countable) is information that can only take certain values. These values
don’t have to be whole numbers but they are fixed values – such as shoe size, number of
teeth, number of kids, etc. Discrete data includes discrete variables that are
finite, numeric, countable, and non-negative integers (5, 10, 15, and so on).
Continuous data (measurable) is data that can take any value. Height, weight,
temperature and length are all examples of continuous data. Continuous data changes
over time and can have different values at different time intervals like weight of a person.
Methods of data collection
There are three major methods of data collection
I. Observation or measurement

6
Kotebe University of Education, Department of Geography and
Environmental studies
In this method, data can be obtained through direct observation or measurement. It
requires training of persons who measure in order to insure the use of standard procedure.
It provides accurate information but it is expensive and inconvenient
II. Interviews and Questionnaires
Questionnaire: - are written documents which instruct the readers or listeners to answer
the questions written on it. There are three ways of collecting information under this
method.
a) Face to face interviews (Questionnaires in charge of interviewers)
b) Telephone interviews
c) Mailed questionnaires (Self-administered questionnaires returned by mail)
III. The use of documentary sources
It is extracting of information from existing sources.
1.7 Geographic Data: Spatial and Aspatial Data
By ‘data’, geographers usually mean collections of facts and figures. Data are the
observed values of a variable. Geographers usually refer data as ‘aspatial’, because an
explicit spatial or locational reference is not an integral part of the information they
contain. Aspatial data are attribute data which specifies characteristics at that location
(what, how much, and when).
There is a data type which may be exclusively geographical: spatial data. Such data
consist of observations on geographical individuals which may only be interpreted
satisfactorily when their locations (where) have been taken into consideration. This may
involve a consideration of their absolute locations (site characteristics), or their relative
locations as measured with respect to some benchmark such as the National Grid or sea
level. Spatial data are often collected using maps, plans or charts, but increasingly, they
are to be found as the basic data type of computerised information systems, in which
location provides an obvious and generally tractable method of organisation.
By ‘spatial’, geographers usually mean data which are gathered in the form of points or
dots, lines, areas or surfaces.
1.8 Main Objectives of Quantitative Revolution in Geography
Geography for more than two hundred years was confronted with the problems of
generalization and theory-building. After the Second World War, geographers,

7
Kotebe University of Education, Department of Geography and
Environmental studies
especially those of the developed countries, realized the significance of using
mathematical language rather than the language of literature in the study of geography.
Consequently, empirical descriptive geography was discarded and greater stress was laid
on the formulation of abstract models. Mathematical and abstract models need rigorous
thinking and use of sophisticated statistical techniques.
The diffusion of statistical techniques in geography to make the subject and its theories
more precise is known as the ‘quantitative revolution’ in geography. The application of
statistical and mathematical techniques, theorems and proofs in understanding
geographical systems is known as the ‘quantitative revolution’ in geography
The main objectives of the quantitative revolution in geography are
1. To change the descriptive character of the subject (geo + graphy) and to make it a
scientific discipline;
2. To explain and interpret the spatial patterns of geographical phenomena in a rational,
objective and cogent manner;
3. To use mathematical language instead of the language of literature,
4. To make precise statements (generalizations) about locational order
5. To test hypotheses and formulate models, theories and laws for estimations and
predictions
6. To identify the ideal locations for the various economic activities so that the profit may
be maximized by the resource users; and
7. To provide geography a sound philosophical and theoretical base, and to make its
methodology objective and scientific.
In order to achieve these objectives, the preachers of quantitative techniques stressed on
field surveys for the collection of data and empirical observations.
2. Tabular and graphical display of quantitative data
There are two types of statistical presentation of data - graphical and numerical.
A tabular database, as the name implies is a database that is structured in
a tabular form. It arranges data elements in vertical columns and horizontal rows
Advantages of tabular presentation of data:
 Tabulated data can be easily understood and interpreted.
 Tabulation facilitates comparison as data are presented in compact and organized
form.
 It saves space and time.

8
Kotebe University of Education, Department of Geography and
Environmental studies
 Tabulated data can be presented in the form of diagrams and graphs.
Tables
Data can also be presented by means of tables. This is a systematic organization of data
using columns and rows.
Elements of a table
Ideal table should have Number, Title, Column headings and Foot-notes
• Number: Table number for identification in a report
• Title, place: Describe the body of the table, variables
• Time period: (What, how classified, where and when)
• Column Heading: Variable name, No. , Percentages (%), etc.,
• Foot-note(s): to describe some column/row headings, special cells, source, etc.,
The Frequency Distribution Table
A frequency is the number of times a value of the data occurs. A frequency distribution
is the organization of raw data in table form, using classes and frequencies. Frequency
distribution refers to data classified on the basis of some variable that can be measured
such as prices, weight, height, wages etc. The frequency distribution table shows the
number of items falling into each group.
The tables display counts (frequencies) and percentages or proportions (relative
frequencies). The percent columns make comparing the same categories easier.
Displaying percentages along with the numbers is often helpful, but it is particularly
important when comparing sets of data that do not have the same totals. The following
technical terms are important when a continuous frequency distribution is formed:
1. Class frequency - the number of observations corresponding to a particular class.
2. Class Limits - are the lowest and the highest values that can be included in the class.
The lower limit of a class is the value below which there can be no item in the class. The
upper limit of a class is the value above which there can be no item to that class.
Example: 0 – 10, in this class, the lowest value is zero and highest value is 10. The two
boundaries of the class are called upper and lower limits of the class. Class limit is also called
as class boundaries.
3. Class interval - The difference between upper and lower limit of class .
Example: In the class 0 – 10, the class interval is (10 – 0) = 10.

9
Kotebe University of Education, Department of Geography and
Environmental studies
The formula to find class interval is gives on below

i= where L = Largest value

S = Smallest value
R = the no. of classes
Example: If the mark of 60 students in a class varies between 40 and 100 and if we want
to form 6 classes, the class interval would be
I= (L-S) / R = 100-40/6=10 Therefore, class intervals would be 40 – 50, 50 – 60, 60 – 70,
70 – 80, 80 – 90 and 90 – 100.
Types of class intervals: There are three methods of classifying the data according to
class intervals namely
a) Exclusive method: When the class intervals are so fixed that the upper limit of one
class is the lower limit of the next class; it is known as the exclusive method of
classification. This method makes continuity of data. The following data are classified on
this basis.
Household Expenditure in birr Better way of expression No of families
0-5000 0 to less than 5000 60
5000-10000 5000 to less than 10000 95
10000-15000 10000 to less than 15000 122
15000-20000 15000 to less than 20000 83
20000-25000 20000 to less than 25000 40
Total 400
Thus, a household with the expenditure of 4999.5 birr will be included in the 0 – 5000
class.
b) Inclusive method (Non-overlapping): In this method, the overlapping of the class
intervals is avoided. Both the lower and upper limits are included in the class interval.
Class interval Frequencies
5-9 7
10-14 4
15-19 3
20-24 2
25-29 1
Total 17

10
Kotebe University of Education, Department of Geography and
Environmental studies
c) Open-end classes- are formed when a class limit is missing either at the lower end of
the first class interval or at the upper end of the last class interval or both are not
specified.
Salary range No of workers
Below 2000 7
2000-4000 5
4000-6000 4
6000-8000 6
Above 8000 3
Total 25
4. Class mark (CM) - midpoint of a class interval. It is found out by adding the upper
and lower limits of a class and dividing the sum by 2
Categorical or Qualitative Frequency Distributions
A categorical frequency distribution represents data that can be placed in specific
categories, such as gender, blood group, & hair colour, etc
Example: The blood types of 25 blood donors are given below. Summarize the data
using a frequency distribution. AB, B, A, O, B, O, B, O, A, O, B, O, B, B, B, A, O, AB,
AB, O, A, B, AB, O, and A
Blood type Frequency
A 5
B 8
O 8
AB 4
Total 25

Quantitative Frequency Distributions -- Ungrouped


An ungrouped frequency distribution simply lists the data values with the corresponding
frequency counts with which each value occurs.
Example: The at-rest pulse rate for 16 athletes at a meet were 57, 57, 56, 57, 58, 56, 54,
64, 53, 54, 54, 55, 57, 55, 60, and 58. Summarize the information with an ungrouped
frequency distribution.
Pulse rate 53 54 55 56 57 58 60 64 Total
Frequency 1 3 2 2 4 2 1 1 16

Quantitative Frequency Distributions -- Grouped

11
Kotebe University of Education, Department of Geography and
Environmental studies
A grouped frequency distribution is obtained by constructing classes (or intervals) for the
data, and then listing the corresponding number of values (frequency counts) in each
interval.
Creating a Grouped Frequency Distribution
1. Find the largest and smallest values

2. Compute the Range = Maximum - Minimum

3. Select the number of classes desired. This is usually between 5 and 20. However
there is no rigidity about it.
4. Find the class width by dividing the range by the number of classes and rounding
up. There are two things to be careful of here. You must round up, not off.
Normally 3.2 would round to be 3, but in rounding up, it becomes 4. If the range
divided by the number of classes gives an integer value (no remainder), then you
can either add one to the number of classes or add one to the class width.
Sometimes you're locked into a certain number of classes because of the
instructions.
5. Pick a suitable starting point less than or equal to the minimum value. You will be
able to cover: "the class width times the number of classes" values. You need to
cover one more value than the range. Follow this rule and you'll be okay: The
starting point plus the number of classes times the class width must be greater
than the maximum value. Your starting point is the lower limit of the first class.
Continue to add the class width to this lower limit to get the rest of the lower
limits.
6. To find the upper limit of the first class, subtract one from the lower limit of the
second class. Then continue to add the class width to this upper limit to find the
rest of the upper limits.

7. Find the boundaries by subtracting 0.5 units from the lower limits and adding 0.5
units from the upper limits. The boundaries are also half-way between the upper
limit of one class and the lower limit of the next class. Depending on what you're
trying to accomplish, it may not be necessary to find the boundaries.

8. Tally the data.

12
Kotebe University of Education, Department of Geography and
Environmental studies
9. Find the frequencies.

10. Find the cumulative frequencies. Depending on what you're trying to accomplish,
it
11. If necessary, find the relative frequencies and/or relative cumulative frequencies.
For example, Let us consider the weights in kg of 50 college students. 42, 62, 46, 54, 41,
37,54,44,32,45,47,50,58,49,51,42,46,37,42,39,54,39,51,58,47,64,43,48,49,48,49,61,41,40
,58,49,59,57,57,34,56,38,45,52,46,40,63,41,51, and 41
Sturges formula to find number of classes is given below
K = 1 + 3.322 log N.
Where K = No. of class
log N = Logarithm of total no. of observations
K=1 + 3.322 log 50= 1+3.22*1.6990=6.64 (Rounded off to 7)
To arrange the data in grouped table we have determine size of class interval. Thus, size
of class interval (C) =

64-32/1+3.322log50=32/6.64=4.8193 (Rounded off to 5)


Thus the size of each class is 5 and the number of classes is 7
The required frequency distribution is prepared using tally marks as given below:

Variations of the Frequency Distribution


A relative frequency is the ratio (fraction or proportion) of the number of times a value of
the data occurs in the set of all outcomes to the total number of outcomes. To find the
relative frequencies, divide each frequency by the total number of samples. Relative

13
Kotebe University of Education, Department of Geography and
Environmental studies
frequencies can be written as fractions, percents, or decimals. Sum of relative frequencies
should equal 1.0 or = 100% by percentage.
RF = class frequency ÷ no. of observations
Class (Pulse rate) Frequency Relative frequency
53 1 0.0625
54 3 0.1875
55 2 0.125
56 2 0.125
57 4 0.25
58 2 0.125
60 1 0.0625
64 1 0.0625
Total 16 1.00
Cumulative relative frequency is the accumulation of the previous relative frequencies.
To find the cumulative relative frequencies, add all the previous relative frequencies to
the relative frequency for the current row.
Example of a simple frequency distribution-families with number of children- 5, 7, 8, 1,
5, 9, 3, 4, 2, 2, 3, 4, 9, 7, 1, 4, 5, 6, 8, 9, 4, 3, 5, 2, 1
Class (Pulse rate) Frequency Relative frequency Cumulative frequency
9 3 0.12 3
8 2 0.08 5
7 2 0.08 7
6 1 0.04 8
5 4 0.16 12
4 4 0.16 16
3 3 0.12 19
2 3 0.12 22
1 3 0.12 25
Total 25 1.00 25

Tables are a good way of organizing and displaying data. But graphs can be even more
helpful in understanding the data. Graphic presentation represents a highly developed
body of techniques for elucidating, interpreting, and analyzing numerical facts by means
of points, lines, areas, and other geometric forms and symbols. In graphical Presentation
we look for the overall pattern and for striking deviations from that pattern. Over all
pattern usually described by shape, centre, and spread of the data. An individual value
that falls outside the overall pattern is called an outlier.

14
Kotebe University of Education, Department of Geography and
Environmental studies
There are no strict rules concerning which graphs to use. It is a good idea to look at a
variety of graphs to see which is the most helpful in displaying the data. We might make
different choices of what we think is the “best” graph depending on the data and the
context. Our choice also depends on what we are using the data for.
Common Types of Graphs
Histogram
A histogram is a graphic version of a frequency distribution. The graph consists of bars of
equal width drawn adjacent to each other. The horizontal scale represents classes of
quantitative data values and the vertical scale represents frequencies. The heights of the
bars correspond to frequency values. Histograms are typically used for large, continuous,
quantitative data sets. A frequency polygon can also be used when graphing large data
sets with data points that repeat.
A bar chart consists of a series of rectangular bars where the length of the bar represents
the quantity of frequency for each category if the bars are arranged horizontally. If the
bars are arranged vertically, the height of the bar represents the quantity. In a bar graph,
the length of the bar for each category is proportional to the number or percent of
individuals in each category. Bars may be vertical or horizontal.

A pie chart a circular graph that is useful in visually showing how a total quantity is
distributed among a group of categories. The “pieces of pie” represent the proportions of
the total that fall into each category. In a pie chart, categories of data are represented by
wedges in a circle and are proportional in size to the percent of individuals in each
category.

15
Kotebe University of Education, Department of Geography and
Environmental studies
Annual Estimates of U.S. Population 65 Years and Over by Race, 2003

White
Black
American Indian
Asian
Pacific Islander
Two or more

Time Series Graph


Time series graphs are important tools in various applications of statistics. When
recording values of the same variable over an extended period of time. Time series chart:
a graph displaying changes in a variables at different points in time. It shows time
(measured in units such as years or months) on the horizontal axis and the frequencies
(percentages or rates) of another variable on the vertical axis.
Percentage of Total U. S. Population 65 Years and Over, 1900 to 2050
25

20

15

10

0
1900 1920 1940 1960 1980 2000 2020 2040 2060

3. Measures of Central Tendency


What are the measures of central tendency?
A measure of central tendency (also referred to as measures of centre or central location)
is a summary measure that attempts to describe a whole set of data with a single value
that represents the middle or centre of its distribution. It does not provide information
regarding individual data from the dataset, where it gives a summary of the dataset. An

16
Kotebe University of Education, Department of Geography and
Environmental studies
average represents a whole series and as such, its value always lies between the minimum
and maximum values and generally ’it is located in the centre or middle of the
distribution.
There are five main measures of central tendency: Arithmetic Mean, Geometric Mean,
Harmonic Mean, Mode and Median. Each of these measures describes a different
indication of the typical or central value in the distribution. Of the above mentioned five
important averages Arithmetic Average, Median and Mode are the most popular ones.
Arithmetic Mean or Simply Mean:
Mean for sample is denoted by symbol ‘M or x̅ (‘x-bar’)’ and mean for population is
denoted by ‘µ’ (mu). It is one of the most commonly used measures of central tendency
and is often referred to as average. Arithmetic Mean can be calculated as the sum of all
the values in the dataset divided by the number of values. That is, If X is the observation
which takes the value X1, X2, X3, … Xn and n represents the total number of

observation; AM (= ) equals:

Computation of Mean for Ungrouped Data


Ungrouped data: Any data that has not been categorised in any way is termed as an
ungrouped data. The formula for computing mean for ungrouped data is X̅ = ΣX/ N
Where, X̅ = Mean
ΣX= Summation of scores in the distribution and
N = Total number of scores.
Let us now compute mean with the help of an example: The scores obtained by 10
students on psychology test are as follows: 58, 34, 32, 47, 74, 67, 35, 34, 30 and 39
Step 1: In order to obtain mean for the above data we will first add the marks to obtain
ΣX: 58+ 34+ 32+ 47+ 74+ 67+ 35+ 34+ 30+ 39 = 450
Step 2: Now using the formula, we will compute mean X̅ = ΣX/ N, ΣX= 450, N= 10
(Total number of students)

17
Kotebe University of Education, Department of Geography and
Environmental studies
Thus, X̅ = 450/ 10 = 45
Thus, the mean obtained for the above data is 45
Computation of Mean for Grouped Data
Grouped data: A data that is categories or organised is termed as grouped data. Mainly
such data is organised in frequency distribution. There are discrete and continuous
grouped data. Discrete data is a numerical type of data that includes whole, concrete
numbers with specific and fixed data values determined by counting. Discrete
data includes discrete variables that are finite, numeric, countable, and non-negative
integers (5, 10, 15, and so on). On the other hand, continuous data includes complex
numbers and varying data values that are measured over a specific time interval. Height,
weight, temperature and length are all examples of continuous data.
Computation of Mean for Grouped Discreet Data
In a discrete series the values of the variable are multiplied by their respective
frequencies and the products so obtained are totalled. This total is divided by the number
of items, which in a discrete series, is equal to the total of the frequencies. If we have a
discrete frequency distribution with frequencies f1, f2, … fn associated with the values
X1, X2, …, Xn of the variable it can be seen that the sum of all the items equals
f1x1+f2x2+…+fnxn since there are now f1 items with value x1, f2 items with value x2,
and so on. The total of items is f=N. = Summation

Example: Frequency of marks of students


Marks out of 20 (x) Frequency (f) fx
9 1 9
10 2 20
11 3 33
12 6 72
13 10 130
14 11 154
15 7 105
16 3 48
17 2 34
18 1 18
Total 46 623

18
Kotebe University of Education, Department of Geography and
Environmental studies
x̅ = 623/46=13.54
Computation of Mean for Grouped Continuous Data
If, the data is grouped with class interval we do not know the exact values of each item.
In this case what we do is estimation based on some assumption. The assumption is that
all the items within a particular class are concentrated at the mid value of the class and
thus fx corresponds to the f items of a class equals fm where m is the mid point of class
interval. Thus,

The method of finding the mid value in this class is as follows:

Formula:

Example: Frequency of marks of students


Marks Frequencies (f) Midpoint (X) fX
35- 39 5 37 185
30-34 7 32 224
25-29 5 27 135
20-24 6 22 132
15-19 4 17 68
10- 14 3 12 36
N= 30 ΣfX = 780

The steps followed for computation of mean with grouped data are as follows:
Step 1: The data is arranged in a tabular form with marks grouped in categories with class
interval of 5.
Step 2: Once the categories are created, the marks are entered under frequency column
based on which category they fall under.
Step 3: The midpoints of the categories are computed and entered under X.
Step 4: fX is obtained by multiplying the frequencies and midpoints for each category.
Step 5: fX for all the categories are added to obtain fX, in case of our example it is
obtained as 780
Step 6: The formula M = fX/ N is used, N is equal to 30.
x̅ = fX/ N M = 780/ 30= 26

19
Kotebe University of Education, Department of Geography and
Environmental studies
Computation of Mean by Shortcut Method (with Assumed Mean)
The above direct method of the calculation of arithmetic average can be used only when
the items are few and the size of the figures is small. If it is not so, there would be
considerable difficulty in the calculation of the arithmetic average. In such situations, a
short cut method with the help of assumed mean can be computed. The assumed
mean method is a technique used in statistics to calculate the arithmetic mean. So we can
assume any arbitrary mean to find out the deviations of items from this assumed [Link]
is a useful shortcut method to calculate the mean from a set of data. The formula to
calculate mean of groped and ungrouped data through short cut method is as follows:

The short-cut method of calculating arithmetic average should be used in all cases as it
saves time and gives accurate results. The process of calculating the arithmetic average
by the short-cut method from a continuous groped data can be summarized as follows:
(i) Assume as average, the midpoint of a class which is in the middle of the series.
Technically any class can be chosen, but if the class chosen is in the middle of the
distribution there is considerable facility in calculations.
(ii) Calculate the deviations of the items (midpoints in case of continuous series) from the
assumed mean.
(iii) Divide the deviations by a common factor or magnitude of the class interval. These
deviations are known as step deviation or deviation in class-interval limits
(iv) Multiply the deviations with the respective frequencies of the various classes and
total the products, taking into account the algebraic signs (plus or minus).
(v) Divide this total by the total of the frequencies and if step deviations have been taken
multiply the result by the common factor or the magnitude of the class-interval.
(vi) Add this figure to the assumed average and the resulting figure would be the actual
arithmetic average of the series.
Example: Find the mean of this set of values by the assumed mean method. 91, 48, 9, 37,
6, 42, 23, 45, 63, 84, 88, 29, 28, 10, 8 by taking the assumed mean as 40

20
Kotebe University of Education, Department of Geography and
Environmental studies
Example: Let us discuss the steps followed for computation of mean with the help of an
example given below:
Class Intervals Frequencies (f) Midpoint (X) x′ = {(X −AM)/ i} f x′

35- 39 5 37 3 15

30-34 7 32 2 14

25-29 5 27 1 5

20-24 6 22 0 0

15-19 4 17 -1 -4

10-14 3 12 -2 -6

N=30  fx'=24

Step 1: We will assume mean (AM) as 22.


Step 2: Difference is obtained between each of the midpoints and the assumed mean and
then the same is divided by ‘i’ that is the class interval (5 in this case), these are then
entered under column with heading x'= { (X - AM)/ i}. The x' for 22 will be 0.
Step 3: Frequency (f) is then multiplied with x' to obtain f x'.
Step 4: All f x' are added to obtain  fx', in the present example it is 24.
Step 5: The formula for mean is now applied x̅ =22 + (24/30 x 5) = 22 + 4 = 26
And if you refer to the mean obtained by the direct method and mean obtained with the
shortcut method, the mean is the same that is 26.
Merits of arithmetic average
(a) It is simple to calculate,
(b) It does not need arraying of data,
(e) It is easy to understand
Limitations of Mean
1) Since arithmetic average is calculated from all the items of a series sometimes the
abnormal items may considerably affect this average, particularly when the number of
items is not large.

21
Kotebe University of Education, Department of Geography and
Environmental studies
2) When there are open ended classes, such as 10 and above or below 5, mean cannot be
computed. In such cases median and mode can be computed. This is mainly because in
such distributions midpoint cannot be determined to carry out calculations.
3) If a score in the data is missing or lost or not clear, then mean cannot be computed
unless mean is computed for rest of the data by not considering the lost score and
dropping it all together.
4) Arithmetic average is no doubt easy to calculate but in a relative sense its calculation
may be more difficult than that of mode or median as they can be located merely by
inspection.

5) It is not suitable for data that is skewed or is very asymmetrical as then in such
cases mean will not adequately represent the data.

The median is the half


way point in a data set.
It is the point at which
50% of the values in
the
data set have a value
the size of the median
value or smaller and

22
Kotebe University of Education, Department of Geography and
Environmental studies

50% of the values have


a value
the size of the median
value or larger. The
median is a positional
average and unlike the
mean its
computation is not
based on all items.
The median is the half
way point in a data set.
It is the point at which
50% of the values in
the

23
Kotebe University of Education, Department of Geography and
Environmental studies

data set have a value


the size of the median
value or smaller and
50% of the values have
a value
the size of the median
value or larger. The
median is a positional
average and unlike the
mean its
computation is not
based on all items.
Median
when the observation are arranged in ascending or descending order, then a value, that
divides a distribution into equal parts, is called median. The median is the middle point in
the distribution; 50% of the observations fall on each side of the median or the median is
the value of the point, which has half the data smaller than that point and half the data
larger than that point, thus it is a positional average.
Formula to determine mode in ungrouped data (odd and even observations)

24
Kotebe University of Education, Department of Geography and
Environmental studies
To determine the median:
1. Put the data in order from smallest to largest.
2. Determine the number in the exact center.
i. If there are an odd number of data points, the median will be the number in the absolute
middle.
ii. If there is an even number of data points, the median is the mean of the two center data
points, meaning the two center values should be added together and divided by 2.

Example: Consider the data set: 17, 10, 9, 14, 13, 17, 12, 20, and 14
Step 1: Put the data in order from smallest to largest. 9, 10, 12, 13, 14, 14, 17, 17, 20
Step 2: Determine the absolute middle of the data. 9, 10, 12, 13, 14, 14, 17, 17, 20
th
Since the number of data points is odd then median = size of (9+1/2) = 10/2 item = 5th
item in the data set is 14.
Example: Suppose the values of the observation are 1, 2, 3, 5, 8 and 10
In this case the median is the arithmetic mean of the two middle observations, i.e., n/2=3
and (n/2)+1=4. The median in this case is found by averaging the 3 rd and 4th items which
is (3+5)/2=4.
Median for grouped data (discrete variable)
In a discrete series also the items are first arranged according to the ascending or
descending order of magnitude and their respective frequencies are written against them.
After this, the frequencies are cumulated and then the value of the middle item can be
easily located. The following example illustrates the procedure
Example: Find the median size of the shoe from the following data.
Size of shoes Number of pairs (frequency) Cumulative frequency
5 30 30

25
Kotebe University of Education, Department of Geography and
Environmental studies
6 40 70
7 50 120
8 150 270
9 300 570
10 600 1170
11 950 2120
Total 2120

Solution: Median is the value of [2120/2]th = 1060th observation. Since the observations
are arranged in ascending order (size-wise), the size of 1060 th observation is easily
determined by constructing the cumulative frequency. Thus, the median size of shoes
sold is 10.
Example 2: The following data gives the distribution of the height of students
Height (in cm) 160 150 152 161 156 154 155
Number of students 12 8 4 4 3 3 7

Solution:
Arranging the data in ascending order of magnitude, we obtain
Height (in cm) 160 150 152 161 156 154 155
Number of students 12 8 4 4 3 3 7
Cumulative frequency 8 12 15 22 25 37 41

26
Kotebe University of Education, Department of Geography and
Environmental studies
Here, the total number of items is 41 i.e., an odd number. Hence, the median is [41 + 1] /
2th i.e., 21st item. From the cumulative frequency table, we find that median i.e., 21 st item
is 155, (All items from 16th to 22nd are equal, each 155).
Median for grouped data (continuous variable):
When the median of a continuous frequency distribution has to be determined there is one
difficulty. The value of the median lies in a class interval, and to get a definite figure,
interpolation has to be done. Suppose, for example it is found that the value of the median
lies in the 20 to 30 class interval whose frequency is 40. Now to find out the value of the
median we have to take recourse to interpolation and to apply a particular formula.
The formula of interpolation to find out the median is:-

Example: Consider the following data: (find median income)


Distribution of workers by average monthly income
Group Number Monthly earnings Number of workers Cumulative Frequency
1 27.5-32.5 120 120
2 32.5-37.5 152 272
3 37.5-42.5 170 442
4 42.5-47.5 214 656
5 47.5-52.5 410 1066
6 52.5-57.5 429 1495

27
Kotebe University of Education, Department of Geography and
Environmental studies

7 57.5-62.5 568 2063


Total 2063

Solution: The median of 2063 cases is the income of [N/2] th worker, which is [2063/2]th =
1031.5 which is 1032nd worker arranged in ascending order of income. From the
cumulative frequency this worker has his income in the class 47.5-52.5. But it is
impossible to determine his exact income. We, therefore, resort to approximation by
assuming that the 410 workers of this class are distributed uniformly across the interval
47.5-52.5. The median worker is [1032-656 (the CF of the preceding median class)]
=376th of these 410, and hence, the value corresponding to him can be approximated as
And this is 47.5+ (376/410)*(52.5-47.5) =52.1
Merits of median
(i) It can be easily calculated and it is understood without any difficulty.
(ii) It is not affected by the values of the extreme items and as such is sometimes more
representative than arithmetic average.
(iii) Even if the value of the extremes is not known median can be calculated if the
number of items is known.
(iv) It can be located merely by inspection in many cases.
(v) It gives best results in a study of those phenomena which are incapable of direct
quantitative measurement
Drawbacks of median
(i) Median may not be representative of a series in many cases. This is specially so when
there are wide variations between the values of different items
(ii) It is not suitable for further algebraic treatment.
(iii) When median has to be calculated in continuous series it requires interpolation. The
assumption of the interpolation, that all the frequencies of the class-interval are uniformly
spread over their values in the class-interval, may not be actually true. In most cases it
will not be true.
(iv) If big or small items in a series are to receive greater importance median would be an
unsuitable average. Median ignores the values of extreme items.

28
Kotebe University of Education, Department of Geography and
Environmental studies
(v) Median is more likely to be affected by the fluctuations of sampling than the
arithmetic average.
(vi) The arrangement of items in ascending or descending order is sometimes very
tedious
Mode

Mode is denoted by symbol ‘Mo’ is the value that occurs with the greatest frequency in a
given distribution. The mode is the value which occurs or repeats itself the greatest
number of times. It is the easiest score to spot in a [Link] is most useful as a
measure of central tendency when examining categorical data. It is the only way to
express the central tendency of a nominal level variable. For the normal distribution, the
mode is also the same value as the mean and median . There may be one mode; multiple
modes, if more than one number occurs most frequently; or no mode at all, if every number
occurs only once. To determine the mode:

1. Put the data in order from smallest to largest, as you did to find your median.
2. Look for any value that occurs more than once.
3. Determine which of the values from Step 2 occurs most frequently.
The data having one mode is called uni-modal distribution. For example, in the
following list of numbers, 16 is the mode since it appears more times in the set than any
other number: 3, 3, 6, 9, 16, 16, 16, 27, 27, 37, 48
The data having two modes is called bi-modal distribution. Example: 3, 3, 3, 9, 16, 16,
16, 27, 37, 48 In this above example, both the number 3 and the number 16 are modes as
they each occur three times and no other number occurs more often.
The data having more than two modes is called multi-modal distribution. It should also
be noted that a distribution may not have mode if no number in a set of numbers occurs
more than once, that set has no mode.
Example: 3, 6, 9, 16, 27, 37, 48

Mode for grouped data:


Mode in case of Discrete Grouped Data
In a discrete grouped data a value which has the largest frequency in a set of data is
called mode. If two elements have the highest frequency, then the data is considered to
be having two modes and is called bimodal data.

29
Kotebe University of Education, Department of Geography and
Environmental studies
Example
Find the mode of the following data?
38 39 40 41 42 43
Size
Number of items 32 15 24 27 44 38

Solution: From the above table, it is observable that the size 42 is the entry with highest
frequency (44).
Mode in case of Continuous Grouped Data:
In case of continuous grouped data, the mode would lie in the class that carries the
highest frequency. This class is called the modal class. The exact location of mode in
a class-interval is done by interpolation
In a continuous series the determination of mode involves three steps.
First, prepare the frequency distribution table in such a way that its first column consists
of the observations and the second column the respective frequency.
Second by the process of grouping, the class in which there is maximum concentration
has to be located. This class is called the modal class.
Third, calculate mode, using the formula
The interpolation is made by the use of the following formula:

M0= L1+

Where:
L1 is the lower limit of the modal class,
f0 is the frequency of the preceding class (class next below modal class),
f1 is the frequency of the modal class,
f2 is the frequency of the following class (class next above modal class)
i is the width of the modal class
Example: compute the mode of the following table
Wage group Frequency
14-18 6
18-22 18
22-26 19

30
Kotebe University of Education, Department of Geography and
Environmental studies
26-30 12
30-34 5
34-38 4
38-42 3
42-46 2
46-50 1
50-54 0
54-58 1

In the above given data 22-26 is the modal class, since it has the largest frequency. The
lower limit of the modal class is 22, its upper limit is 26, its frequency is 19, the
frequency of the preceding class is 18, and of the following one is 12. The class interval
is 4. Using the method mentioned above we can determine the mode as follows;
Mode=22+ (19-18)/ (2x19-18-12) x4=22.5
Example: A survey on the heights (in cm) of 30 students of the same batch was
conducted at a university. The data so obtained has been organized in the table given
below. Find the mode
Height (in cm) Number of students
120 – 125 3
125 – 130 5
130 – 135 11
135 – 140 6
140 – 145 5
Total 30

The Modal class = 130 – 135 as its frequency is the highest (11)
The Lower limit of the modal class = (L) = 130
Frequency of the modal class = 11
Frequency of the preceding modal class = 5
Frequency of the next modal class = 6
Size of the modal class interval = (h) = 5
Putting the values in the formula,
Mode = 130 + (11-5)/ (2×11-5-6)5= 130 + (6/11) (5) = 130 + 0.54×5 = 132.72

31
Kotebe University of Education, Department of Geography and
Environmental studies
Therefore, mode = 132.72

Exercise: based on the data presented in the table below, compute the mode
Class boundaries Frequency
29.5---39.5 8
39.5---49.5 87
49.5---59.5 190
59.5---69.5 304
69.5---79.5 211
79.5---89.5 85
89.5---99.5 20

 f =905
Merits of Mode
(i) It possesses the merit of simplicity. It can be determined without much mathematical
calculation. In a discrete series mode can be located even by inspection. In this respect,
like median, it has an advantage over arithmetic average
(ii) It is commonly understood. As has been said earlier, mode an average which people
use in their day-to-day expressions.
(iii) Since mode is the most common item of a series it is not an› isolated example like
the median: Unlike arithmetic average it cannot be a value which is not found in the
series.
(iv) Mode is not affected by the values of extreme items provided. They adhere to the
natural law relating to extremes.
v) For the determination of mode it is not "necessary to know the values of all the items
of a series. If the point of norm or maxi› mum concentration is known it is enough.
Drawbacks of mode
Mode is an unsatisfactory average and has many drawbacks. Some of them are as
follows:
(i) Mode is not capable of further mathematical treatment.
(ii) In many cases it may be impossible to set a definite value of .ode. There may be 2, 3
or more modal values

32
Kotebe University of Education, Department of Geography and
Environmental studies
(iii) Mode may be unrepresentative in many cases. If in a series 1000 items 20 have a
particular value and other values have frequencies is than 20, it does not necessarily mean
that the value whose frequency 20 is the typical or average value. In such cases data
should be converted into class intervals of a bigger magnitude.
The relationship between mean, median and mode
In a normal/ symmetrical distribution the mean, median and mode are identical. In actual
practice, however, symmetrical distributions are very rare, and data usually give a
symmetrical curve.

In distributions which moderately differ from symmetrical distribution, there is an


empirical relationship between mean, median and mode. This relationship holds good for
most of the moderately asymmetrical distributions. The distance between the mean and
the median is about one-third the distance between the mean and the mode. It is
expressed as follows:
Mean−Median=1/3(Mean−Median)
3(Mean−Mode) =Mean−Mode
Mode=Mean−3 (Mean−Median)
Mode=3Median−2Mean

Example: the mean and median of a distribution is 44.6 and 44.05 respectively. Find the
mode. We have, median = 44.05 and mean = 44.6. We know,

Mode=3Median−2Mean
=3(44.05)−2(44.6)

33
Kotebe University of Education, Department of Geography and
Environmental studies
=132.15−89.2=42.95

4. Measures of Scale/Variation/Dispersion/Spreadness
Measuring of central tendency is to determine a single figure to represent a whole series.
Two different distributions may have the same mean, mode and median but can be quite
different in their scatter about the mean. So, the dispersion in itself is a very important
property of a distribution and needs to be measured. Average by itself is not a good
indication of quality of the sample. You need to know the variance to make any educated
assessment. These are statistical procedures for describing the nature and extent of
differences among the information in the distribution. The average income in a
community hides the distribution of income and does not show the quality of the
sample/data. The measures of dispersion (not central tendency) bring out this inequality.
A measure of dispersion/variation is defined as a statistics signifying the extent of the
scatteredness of items around a measure of central tendency. Dispersion‟ or `variation‟ in
statistics is the degree of spread of each individual item or value from the central value in
the given distribution. According to Minium, King and Bear (2001), measures of
variability express quantitatively the extent to which the score in a distribution scatter
around or cluster together. The measures of dispersion are also called the average of
second order, because here we consider the arithmetic mean of the deviations from the
mean of the values of the individual items.
Statistical measures of variation are numerical values that indicate the variability inherent
in a set of data measurements. The greater the similarity of the scores to each other lower
would be the measure of variability or dispersion. The less the similarity of the scores are
to each other, higher will be the measure of variability or dispersion. In general, the more
the spread of a distribution, larger will be the measure of dispersion.
In measuring dispersion, it is imperative to know the amount of variation (absolute
measure) and the degree of variation (relative measure). In the former case, we consider
the range, mean deviation, standard deviation etc. In the latter case, we consider the
coefficient of range, the coefficient of mean deviation, the coefficient of variation etc.

34
Kotebe University of Education, Department of Geography and
Environmental studies
Thus, there are two broad classes of the measures of dispersion or variability. They are
absolute measure of dispersion and relative measure of dispersion.

Absolute dispersion usually refers to the standard deviation, a measure of variation from
the mean. The units of standard deviation are the same as for the data. In other words,
absolute measure is expressed in terms of the original units of a distribution. Therefore,
absolute dispersion is not suitable for comparing the variability of two distributions since
the two variables are expressed and measured in two different units. For instance, the
variability in body height (cm) and body weight (kg) cannot be compared because the
absolute measure (standard deviation) is expressed in cm and kg. The absolute measure is
also not appropriate for two sets of scores expressed in the same units with wide
divergence in means (central value). Nevertheless, absolute measures are widely used,
except in the exceptional cases like above. The absolute measures include range, mean
deviation, standard deviation, and variance.
Relative dispersion, sometimes called the coefficient of variation, is the result of dividing
the standard deviation by the mean and it may be presented as a quotient or as a
percentage. Thus, relative measures are computed from the absolute measures of
dispersion and its corresponding central values. A low value of relative dispersion usually
implies that the standard deviation is small in comparison to the magnitude of the mean.
The six most common measures of variation are the range, quartile deviation/semi-inter-
quartile range, absolute mean deviation, variance, standard deviation, and coefficient of
variation.
Range:
Range is the simplest possible measure of dispersion. The range of any distribution is the
difference between the highest and lowest values in the series.
Symbolically R= X max- X min or R = L - S
Where, R = Range
L = X max= maximum value or largest value
S = X min= minimum or smallest value.
Example: The yields (kg per plot) of a cotton variety from five plots are 8, 9, 8, 10 and
11. Find the range

35
Kotebe University of Education, Department of Geography and
Environmental studies
Solution L=11, S = 8. Range = L – S = 11- 8 = 3
In a frequency distribution, range is given by the difference between the lower limit of
the lowest class and the upper limit of the highest class.

Example: Calculate range from the following distribution.


Size 60-63 63-66 66-69 69-72 72-75
Number 5 18 42 27 8
Solution L = Upper boundary of the highest class = 75
S = Lower boundary of the lowest class = 60
Range = L – S = 75 – 60 = 15
The greater the range is the greater the variation of the values in the group.
Range is a crude measure of dispersion. It is a measure of absolute dispersion and as such
cannot be usefully employed for comparing the variability of two distributions expressed
in different units. For example the range of the weights of students cannot be compared
with the range of their height measurements as the range of weights would be in pounds
and that of heights in inches. So the need of measuring relative dispersion arises for the
purpose of comparison. An absolute measure can be converted to relative measure if we
divide it by some other value regarded as standard for the purposes. If range is divided by
the sum of the extreme items, the resulting figure is called "The Ratio of the Range" or
"The Coefficient of the Scatter."
Coefficient of Range= X max- Xmi= L - S
Xmax + Xmin L + S
Example: Calculate range and its coefficient from the following data relating to Salaries
of 400 Lectures.

Salary 1800 2000 2400 2800 3000 3200 3400 4000 4500 5000 5400 7000
(In birr)
No of 20 25 26 27 30 30 20 40 24 24 94 40
Lectures

Solution:
Range =X max- X min

36
Kotebe University of Education, Department of Geography and
Environmental studies
=7000-1800 = 5200 Birr
Coefficient of Range = Xmax -Xmax = 7000 -1800=5200

Coefficient of Range= X max- Xmi= 7000 -1800 =5200


Xmax + Xmin 7000 + 1800 8800
= 0.591
Example: Calculate range and its coefficient from the mark obtained by distance students
in statistics.
Mark 11-20 21-30 31-40 41-50 51-60 61-70
Number of 15 28 37 10 48 63
Students
Solution: maximum mark is 70 and minimum mark is 11.
X max = 70 , X min=11
Range = X max- Xmin =70 -11 = 59
Coefficient of Range= X max- Xmi= 70 -11 =59 = 0.728
Xmax + Xmin 70 + 11 81
Range is very easy to understand and easy to compute but it is very crude measures of
variation.
Limitations of Range:
The only merits possessed by range are, that it can be easily calculated and readily
understood. As against these, there are many drawbacks from which it suffers.
(i) The most important point against range is that it is affected very greatly by
fluctuations of sampling. A single variation in the value of an extreme item affects the
value of the range. The removal or add of either of the two extreme values may change
the range considerably
(ii) Range is not based on all the observations of the series; it is only based on the two
extreme values Range does not take into account the composition of a series or the
distribution of items within the extremes
The range of a symmetrical and an asymmetrical distribution can be identical. Two such
distributions can never have the same dispersion. In this way we find that range is a very
unsatisfactory measure of dispersion and should be used with extreme caution.

37
Kotebe University of Education, Department of Geography and
Environmental studies
(iii) It cannot be computed if the distribution has open-end class

Quartile Deviation/Semi-inter-quartile range (QD):


Another measure of dispersion, much is better than the range, is the semi-inter-quartile
range. It is the midpoint of the inter-quartile-range (Q3-Q1). In other words, it is one half
of the difference between the third quartile and the first quartile. It can be defined as the
average absolute difference between the third (upper) quartile and the first (lower)
quartile of the frequency distribution. It studies the range of spread various items on
either side of the median and it ignores nearly 50% of the items on either the extreme
ends of the distribution. High degree of quartile deviation means low uniformity, and low
degree of variation.
The semi-inter-quartile range or quartile deviation is given by one-half of the difference
between Q3-Q1, i.e., (Q3-Q1)/2 or Inter quartile range/2. Where Q3 and Q1 stand for the
third quartile (upper quartile) and first quartile (lower quartiles) respectively
First let us see how we can calculate the quartiles
Quartiles: when the observation are arranged in increasing order then
the values, that divide the whole data in to four ( 4 ) equal parts , are
called quartiles.
These values are denoted by Q1 (first or lower quartile), Q2 (second quartile or median)
and Q3 (third or upper quartile). It is to be noted that 25% of the data falls below Q1,
50% of the data falls below Q 2and 75% of the data falls below Q3.

38
Kotebe University of Education, Department of Geography and
Environmental studies

Calculation of quartiles for ungrouped data


Calculate the quartiles for the following the marks obtained by 9 students are given
below:
x 45 32 37 46 39 36 41 48 36
Arranged the observation in ascending order
32, 36, 36, 37, 39, 41, 45, 46, 48.
If “n” is odd then we use the below formula:

Qj =Marks obtained by student, generalized formula of quartiles where j=1, 2,

39
Kotebe University of Education, Department of Geography and
Environmental studies

Q1= Marks obtained by (2+1)th students


Q1= Marks obtained by the third students = 36

Exercise: Calculate Q2 and 3 from the above distribution?


Calculation of quartiles for discrete grouped data

No. of assistants Frequency Cumulative frequency (CF)


0 3 3
1 4 7
2 6 13
3 7 20
4 10 30
5 6 36
6 5 41
7 5 46
8 3 49
9 1 50
Total f=50

40
Kotebe University of Education, Department of Geography and
Environmental studies
Calculation of quartiles for continuous grouped data: Find the first and the third
quartiles of the following data.
Distribution of workers by average monthly income
Group no. Monthly earnings No. workers CF
1 27.5-32.5 120 120
2 32.5-37.5 152 272
3 37.5-42.5 170 442
4 42.5-47.5 214 656
5 47.5-52.5 410 1066
6 52.5-57.5 429 1495
7 57.5-62.5 568 2063
Total 2063
Solution: First quartile:
First locate the Q1 by computing nN/4 = 1(2063)/4=515.75
Q1 class is the class that contains the 515.75 th item. This belongs to (from the cumulative
frequency) 42.5-47.5 class. Substituting into the above formula:

=44.22

Interpretation: first quarter (25 percent) of the workers income is less than or equal to
Birr 44.22.
Third quartile:
First locate the Q3 by computing nN/4 = (3*2063)/4=1547.25
Q3 class is the class that contains the 1547.25 th item. This belongs to (from the
cumulative frequency) 57.5-62.5 class. Substituting into the above formula:

=57.96

Interpretation: Third quarter (75 percent) of the workers income is less than or equal to
Birr 57.96.
The semi-inter-quartile range or quartile deviation of the above example could be
computed as QD= (Q3-Q1)/2 or Inter quartile range/2.
QD= (57.96-44.22)/2= 13.74/2=6.87

41
Kotebe University of Education, Department of Geography and
Environmental studies
Example: Find the quartile deviations for the four distributions given in the four
workshops.
First we have to find Q3, Q2, and Q1 for each workshop Data.
Calculation of Quartile Deviation
WS A WS B WS C WS D
Location of Q2 25.3 25.61 25.07 25.25
Location of Q1 23.41 23.07 22.5 22.75
Location of Q3 27.64 28.0 28.17 28.17
Quartile deviation (27-64-23.41)/2=2.12 (28-23.07)/2=2.46 2.83 2.71
For workshop A the QD is Birr 2.1 and median is 25.3. This means that if the distribution
is symmetrical the number of workers, whose wages vary between (25.3-2.1) Birr 23.2
and (25.3+2.1) Birr 27.4, shall be just half of the total number of workers. As this
distribution is not symmetrical the distance between Q1 and the median (Q2) is not the
same as between Q3 and the median, hence the interval defined by median plus and
minus QD will not be exactly the same as given by the value of the two quartiles. Under
such conditions the range between Birr 23.2 and Birr 27.4 will not include precisely 50%
of the workers.
Quartile deviation is an absolute measure of dispersion and cannot be used for comparing
the variability of any two series. If Quartile deviation is to be used for comparing the
variability of any two series it is necessary to convert the absolute measure to a
coefficient of Quartile Deviation (CQD). If Quartile deviation is divided by the average
value of the two quartiles, a relative measure of dispersion is obtained. It is called the
Coefficient of Quartile Deviation

Symbolically,

By applying this the CQD of the four WS are given below

WS A WS B WS C WS D

Characteristics of QD

42
Kotebe University of Education, Department of Geography and
Environmental studies
The size of the quartile deviation gives an indication about the uniformity or otherwise of
the size of the items of a distribution. If the quartile deviation is small it denotes large
uniformity. Thus, a CQD may be used for comparing uniformity or variation in different
distributions.
The quartile deviation possesses the merits of simple calculation and easy
understandability. It is commonly understood and its calculation does not involve any
mathematical intricacies. These are the points in favor of quartile deviation but there are a
large number of points which go against it.
Quartile deviation is neither based on all the observations of the data, nor is it capable of
further algebraic treatment. It is affected to a considerable extent by the fluctuations of
sampling. A change in the value of a single item may in certain cases affect its value
considerably. Thus quartile deviation is not a very good measure of dispersion,
particularly for series in which the variation is considerable. However, for rough studies
quartile deviation may give an approximate idea of the extent of variability in a series.

Absolute mean deviation (AMD)

The mean deviation will answer the defect of QD. Mean deviation is the arithmetic
average of the variations (deviations) of the individual items of the series from a measure
of their central tendency (commonly used are median and mean).
Calculating AMD involves
(i) Calculate the mean ( ) or median (Me)
(ii) Record the deviations = of each of the items, ignoring the sign
(iii) Find the average value of the deviations

For ungrouped data

For grouped data

Where = for grouped discrete data and = for grouped


discrete/continuous data with M as the mid-value of a particular group
Example: Calculate the mean deviation from the following data
Table: Marks obtained by 11 students in a class test

43
Kotebe University of Education, Department of Geography and
Environmental studies
Serial No. Marks
1 10 8
2 12 6
3 14 4
4 15 3
5 16 2
6 18 0
7 19 1
8 20 2
9 23 5
10 25 7
11 30 12
Total 50

Median= size of = (11+1)/2nd item


=mark of the 6th student is the median=18
AMD from the median = 50/11=4.54 marks.
Merits and demerits of absolute mean deviation:
Merits:
(i) It is easy to understand,
(ii) As compared to standard deviation its computation is simple,
(iii) This measure does not square the distance from the mean, so it is less affected by
extreme observations than are the variance and standard deviation, and
(iv) Since it is based on all values in the distribution, it is better than range or quartile
deviation.
Demerits:
(i) It lacks those algebraic properties which would facilitate its computation and establish
its relation to other measures, it ignores the algebraic signs of deviations, and
(ii) It is not suitable for further mathematical processing.
Coefficient of Mean Deviation (CMD): The coefficient or relative dispersion is found
by dividing the mean deviation by that measure of central tendency about which
deviations were recorded. Thus,

When deviations were recorded from the median

44
Kotebe University of Education, Department of Geography and
Environmental studies

When deviations were recorded from the mean

Standard Deviation (Root Mean Square)


Standard deviation is the most universally used and better measure of dispersion.
Standard deviation may be abbreviated as s and is most commonly represented in
mathematical texts and equations by the lower case Greek letter sigma σ, for the
population standard deviation, or SD, for the sample standard deviation. It is defined as
the square root of the mean of the squares of the deviations of individual items from their
arithmetic mean.
The process of computing a standard deviation always involves computing a variance.
Since standard deviation is the square root of the variance, it is always expressed in the
same units as the raw data. The computation of Standard deviation follows the following
respects:
1. Compute the deviation (distance from the mean) for each score. The deviations are
always recorded from the arithmetic mean, because although the sum of deviations is the
minimum from the median, the sum of squares of deviations is minimum when deviations
are measured from the arithmetic mean. Deviation= (x- ).
2. Square each deviation

SS  xi  x 
2

3. Compute the mean of the squared deviations. Deviations are squared so as to get rid of
negative signs.
Standard deviation for the for ungrouped data is

For grouped data (discrete variable)

45
Kotebe University of Education, Department of Geography and
Environmental studies

SD where n is summation of the frequency.


For grouped data (continuous variable):

Why the sum of the squares in the numerator is divided by n-1. This divisor is called the
degrees of freedom and when estimating a population variance or SD, it is necessary to
divide by the degrees of freedom to have an estimate that, on the average, exactly equal
to population variance and population SD.
The standard deviation is a measure of the amount of variation or dispersion of a set of
values. A low standard deviation indicates that the values tend to be close to
the mean (also called the expected value) of the set, while a high standard deviation
indicates that the values are spread out over a wider range. A SD of zero value or
variance of zero value means that there is no variation at all in the data set. In other
words, the observations have the same values.
For a population, this involves summing the squared deviations (sum of squares, SS) and
then dividing by N. The resulting value is called the variance or mean square and
measures the average squared distance from the mean.
Variance
Variance is the expectation of the squared deviation of a random variable from
its population mean or sample mean. Variance is a measure of dispersion, meaning it is a
measure of how far a set of numbers is spread out from their average value. Variance is
the average of the squared deviations of each observation in the set from the arithmetic

46
Kotebe University of Education, Department of Geography and
Environmental studies
mean of all of observations. Or it is the arithmetic mean of the squared deviation of the
observations about their mean. It is often represented by s2 (for sample variance) or σ2 (for
population variance)

Population variance

When you have collected data from every member of the population that you’re
interested in, you can get an exact value for population variance. The population
variance formula looks like this:

Where: σ2 = population variance


 Σ = sum of…
 Χ = each value
 μ = population mean
 Ν = number of values in the population

Sample variance

When you collect data from a sample, the sample variance is used to make estimates
or inferences about the population variance. The sample variance formula looks like
this:

Variance = S2 = , n is the number of observation in the sample and its value

should beat at least 2.

S2 for grouped discrete data = and continuous data = , where m

is mid point and n is summation of the frequency.


The variance is a measure of spread or dispersion among values in a data set. Therefore,
the greater the variance, the lower the quality.

47
Kotebe University of Education, Department of Geography and
Environmental studies
Steps for calculating the variance
Step 1: Find the mean
Step 2: Find each score’s deviation from the mean
Step 3: Square each deviation from the mean
Step 4: Find the sum of squares
Step 5: Divide the sum of squares by n – 1 or N
Example: Compute the SD for the following data
11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21
The first formula is appropriate: The procedure is
(i) First calculate the mean =x/N=176/11=16
(ii) And then calculate the deviations
x x2 x- (x- )2
11 121 -5 25
12 144 -4 16
13 169 -3 9
14 196 -2 4
15 225 -1 1
16 256 0 0
17 289 1 1
18 324 2 4
19 361 3 9
20 400 4 16
21 441 5 25
Total 176 2926 110

Thus, SD= =3.16 NB: SD is measured in the units of the original observation.

Calculate the standard deviations for the following data.


Class Frequency
1-3 1
3-5 9
5-7 25
7-9 35
9-11 17

48
Kotebe University of Education, Department of Geography and
Environmental studies
11-13 10
13-15 3
This is an example of continuous frequency series and formula 3 is appropriate.
Class m f fm m- (m- )2 f(m- )2
1-3 2 1 2 -6 36 36
3-5 4 9 36 -4 16 144
5-7 6 25 150 -2 4 100
7-9 8 35 280 0 0 0
9-11 10 17 170 2 4 68
11-13 12 10 120 4 16 160
13-15 14 3 42 6 36 108
 100 800 616
=800/100=8

SD= =2.48

Coefficient of Variation (CV)


Variance and SD cannot be used in comparing two data sets. Thus a SD of 5 on an exam
where the mean score is 30 has an altogether different meaning than on an exam where
the mean score is 90. Clearly the variability in the second exam is much less. To take care
of this problem, we define and use a coefficient of variation, V.
CV expresses the standard deviation as a percentage of the arithmetic mean.

V= (expressed as percentage).

The lesser the coefficient of variation is the lesser the variability.


CV is independent of the unit of measurement. In estimation of a parameter when CV is
less than say 10%, the estimate is assumed acceptable. The inverse of CV; namely 1/CV
is called the Signal-to-noise Ratio.
Example: The following are the scores of two batsmen A and B in a series of innings.
A 12 115 6 73 7 19 119 36 84 29
B 47 12 76 42 4 51 37 48 13 0
Who is the better run-getter? Who is more consistent?

49
Kotebe University of Education, Department of Geography and
Environmental studies
In order to decide as to which of the two batsmen, A and B, is the better run-getter, we
should find their batting averages. The one whose average is higher will be considered as
a better batsman.
Batsman A’s average is =500/10=50
Batsman B’s average is =330/10=33
A is a better batsman since his average is 50 as compared to 33 of B.
To determine the consistency in batting we should determine the coefficient of variation.
The less this coefficient the more consistent will be the player. That is the less the
coefficient the less the variability.
A B
2
Score (x) x- (x- ) Score (x) x- (x- )2
12 -38 1444 47 14 196
115 65 4255 12 -21 441
6 -44 1936 76 43 1849
73 23 529 42 9 81
7 -43 1849 4 -29 841
19 -31 961 51 18 324
119 69 4761 37 4 16
36 -14 196 48 15 225
84 34 1156 13 -20 400
29 -21 441 0 -33 1089
Total 500 17498 330 5462

SD of Batsman A= =41.83

V=(41.83/50)*100=83.66 percent.

SD of Batsman B= =23.37

V= (23.37/33)*100=70.8 percent.
B is more consistent (or less variable) since the variation in his case is 70.8 per cent as
compared to 83.66 of A.
EG 2: The following data are given about height of boys and girls.
Boys Girls
Number 72 38
Average height 68m. 61m.
Variance of distribution 9 4

50
Kotebe University of Education, Department of Geography and
Environmental studies
Find
(i) in which sex is there greater variability in individual height? Compare the coefficient
of variation
(ii) Common average height in boys and girls. Combined mean
(iii) SD of boys and girls taken together. Combined SD
(iv) Combined variability.
(v) Did the separation make each group more homogenous?
SOLUTION:

(i) CV of boys height= *100=4.41%

CV of girls height= *100=3.28%

Thus there is a greater variability in the height of boys than that of girls,

(ii) Height of boys and girls combined is

=65.58 meter

(iii) Combined SD

= =4.28 meter
(iv) Combined variability (combined coefficient of variation) is

CV12= *100= (4.28/65.58)*100=6.53 per cent

(v) Did the separation make each group more homogenous? Compare the coefficient of
variation of boys, girls and combined coefficient of variation, i.e., coefficient of variation
of total population. The lesser the coefficient the more homogenous the distribution is.
CV of boys height = 4.41%
CV of girls height = 3.28%
CV of boys and girls or total population = 6.53 per cent

Thus, the separation makes each group more homogenous.

51
Kotebe University of Education, Department of Geography and
Environmental studies
Comparisons of Various Measures of Dispersion
Range is easy to calculate and useful statistic to know, but it cannot stand alone as a
measure of spread since it takes into account only two extreme scores and hence it is
extremely sensitive to the size of the sample, and to the sample variability.
Quartile deviation is also easy to calculate but suffer from the disadvantage that they are
not amenable to algebraic treatment.
Mean Deviations: MD is easy to interpret and easier to calculate than SD. But it is not
suitable because we cannot obtain the MD of a combined series from the deviations of
component series.
Standard Deviation: It lends itself to rigorous algebraic treatment and based on all
observations. It is, therefore, quite insensitive to sample size (provided that the sample
size is ‘large enough’) and is least affected by sampling variations. It is used as a standard
scale for measuring deviations from the mean. The only disadvantage of SD is it involves
to much work in its calculation and the large weight it attaches to extreme values because
of the process of squaring involved in its calculation.

The standard deviation is by far the most widely used measure of spread. It takes every
score into account, has extremely useful properties when used with a normal distribution
and is tractable mathematically and, therefore, it appears in many formulas in inferential
statistics. The standard deviation is not a good measure of spread in highly-skewed
distributions and should be supplemented in those cases by the semi-interqurtile range.
Measures of Shape (Skewness)
So far we discussed mean and standard deviation, which describe the distribution. But
there are other parameters which describe how symmetrical the distribution is about the
mean, how peaky is the distribution, or the shape of the distribution.

Skewness: When a distribution is not symmetrical it is said to be asymmetrical or


skewed.
The following table shows the heights of the students of a college:
Class A B C D
f f f f

52
Kotebe University of Education, Department of Geography and
Environmental studies
56.5-58.5 5 3 0 4
58.5-60.5 25 5 4 8
60.5- 15 20 40 20
62.5 10 44 24 24
64.5 15 20 20 40
66.5 25 5 8 4
68.5-70.5 5 3 4 0
N 100 100 100 100
Mean 63.5 63.5 63.5 63.5
Median 63.5 63.5 63 64
Mode bimo 63.5 61.9 65.1

The shape of the curves and histograms which placed equal items at equal distance on
either side of the median clearly shows that distributions A and B are symmetrical. If we
fold this curves, or histograms on the ordinate at the mean, the two halves of the curve or
histograms will coincide. In distribution B all the three measures are identical and in A
which is a bimodal distribution mean and median are equal.

Distribution C and D are asymmetrical. This is evident from the shape of the figures and
from the fact that (i) items at equal distances from the median are not equal in number,
and (ii) the three measures of central tendency are not equal. There is also a difference in
asymmetry distributions. When mean is greater than the median and mode, the curve is
pulled more to the right (skewed to the right) and when mean is less than median and
mode =skewed to the left. Put differently, if extreme variations in a given distribution are
towards higher values of observation, they give the curve a longer tail to the right
(skewed to the right) and this pulls the median and mean in that direction from the mode.
If, however, extreme variations are towards lower values of observations, the longer tail
is to the left and the mean and median are pulled to the left of the mode.
It could also be shown that in a symmetrical distribution, the lower and upper quartiles
are equidistant from the median, so also are corresponding pairs of deciles and
percentiles.

53
Kotebe University of Education, Department of Geography and
Environmental studies
So we can have the following as tests for the presence of skewness.
(i) graph,
(ii) when three measures of central tendency are not equal in their value,
(iii) when the sum of the positive deviations from the median are not equal to the
negative deviations from the median,
(iv) when distances from the median to the quartiles are not equal, and
(v) when corresponding pairs of deciles and percentiles are not equi-distant from the
median.

Skewness is the degree of asymmetry of a distribution. If the distribution has a longer tail
less than the maximum, the function has negative skewness. Otherwise, it has positive
skewness.

Measures of Skewness:
1. Karl Pearson’s measures of skewness (Relationship between the three measures of
central tendency.
We said that =Me=Mo for symmetrical distribution. As the distribution departs from
symmetry these values are pulled apart, the difference between and Mo being the
greatest. Karl Pearson has suggested the use of this difference in measuring skewness.
Absolute skewness = -Mo with units measurement.
And plus or minus signs obtained in this formula indicates the direction of the skewness.
If it is positive the extreme variation in the given distribution are towards higher values
(skewed to the right).

Coefficient of skewness: The difference between mean and mode is an absolute measure
of skewness. An absolute measure cannot be used for making valid comparison between
the skewness in two or more distributions for the following reasons: (i) the same size of
skewness has different significance in distributions with small variation and in
distributions with large variation, in the two series, and (ii) the unit of measurement in the
two series may be different. To make this measure suitable to compare skewness it is
necessary to eliminate from it the disturbing influence of variation and units of

54
Kotebe University of Education, Department of Geography and
Environmental studies
measurement. Such elimination is accomplished by dividing the difference between mean
and mode by the standard deviation. The resultant coefficient is called Pearsonian
coefficient of skewness.

Coefficient of skewness =

Since, as we have already seen, in moderately skewed distributions


Mo = Mean-3(Mean-Median)
By substituting this for mode in the above formula

Coefficient of skewness =

2. Bowley’s Measures of skewness (Quartile measures of skewness)


In the above two methods of measuring skewness, the whole series is taken into
consideration. But skewness can be secured even for a part of the series. The usual device
is to measure the distance between the lower and the upper quartiles. In a symmetrical
distribution, the quartiles would be equidistant from the value of the median, i.e.,
Median-Q1=Q3-Median
Median=(Q3+Q1)/2

In other words the value of the median is the mean of Q1 and Q3.
In a skewed distributions, quartiles would not be equidistant from the median unless the
entire asymmetry is located at the extremes of the series. Bowley has suggested the
following formula for measuring skewness, based on above facts.
Absolute SK=(Q3-Me)-(Me-Q1)
=Q3+Q1-2Me
If the quartiles are equidistant from the median
SK = 0
If Me-Q1>Q3-Me = negative skewness
And Me-Q1<Q3-Me = positive skewness

55
Kotebe University of Education, Department of Geography and
Environmental studies
If the series expressed in different units are to be compared, it is essential to convert the
absolute amount into the relative. Using the interquartile range as denominator, the
formula for the coefficient of skewness is as follows:

Relative SK=

Or =

If in the series the median and lower quartiles coincide, then the SK becomes (+1). If the
median and the upper quartiles coincide, then the SK becomes (-1).
Merits and demerits of this measurement: (i) easily computable, (ii) it has value limits
between (+1) and (-1). The disadvantages of this measure is it does not consider all items
in the distribution. It neglects extreme items.
Eg: Calculate the skewness for the above table.

3. Kelly’s measure of skewness (Percentile measure of skewness)


To remove the defect of Bowley’s measure that it does not take into account all the
values, it can be enlarged by taking two deciles (or percentiles), equidistant from the
median.

SK=P50-

=D5-

Skewness will take on a value of zero when the distribution is a symmetrical curve. A
positive value indicates the observations are clustered more to the left of the mean with
most of the extreme values to the right of the mean. A negative skewness indicates
clustering to the right.

Significant Level

56
Kotebe University of Education, Department of Geography and
Environmental studies
The amount of evidence required to accept that an event is unlikely to have arisen
by chance is known as the significance level or critical p-value. Significance
levels show you how likely a result is due to chance. In the field of Education, the
most common level, used to mean something is good enough to be believed, is .95.
This means that the finding has a 95% chance of being true. However, this value is
also used in a misleading way. No statistical package will show you "95%" or
".95" to indicate this level. Instead it will show you ".05," meaning that the finding
has a five percent (.05) chance of not being true, which is the converse of a 95%
chance of being true. To find the significance level, subtract the number shown
from one. For example, a value of ".01" means that there is a 99% (1-.01=.99)
chance of it being true.
Hypothesis
The hypotheses are often statements about population parameters like expected value and
variance. A hypothesis might also be a statement about the distributional form of a
characteristic of interest. They are simply informed of intelligent guess about the solution
to a problem. To understand more about hypothesis it will be necessary to consider the
different types.

Types of Hypothesis

Hypothesis can be classified as either research hypothesis or statistical hypothesis. When


postulations are about the relationships between two or more variables it is known as
research hypothesis. Such a hypothesis is not directly testable statistically because it is
not expressed in measurable terms. For an example ‗Students poor performance in
science is due to use of inadequate methods of teaching the subject. However, it is
possible to express the relationship between two or more variables in statistical and
measurable terms. This is known as statistical hypothesis. Statistical hypothesis is
formulated in two forms – null and alternative hypothesis. A null hypothesis is a
hypothesis of no difference, no relationship or no effect while the alternative hypothesis
specifies any ot the possible conditions not anticipated in the null Hypothesis (Nworgu,
1991). The null hypothesis is represented as H0 while the alternative is H1 or Ha.
Examples:

57
Kotebe University of Education, Department of Geography and
Environmental studies
H = There is no significant difference in the mean scores in science of students
0

taught using inquiry method and those taught using no inquiry method
H1 or Ha = There is a significant difference in the mean score in science of students taught
using inquiry and non-inquiry methods.
Alternative hypothesis can further be described as Non-directional and Directional
Alternatives. When an alternative hypothesis gives the direction of the difference or
effect it is known as a Directional Alternative, but when direction is not given it is non-
directional.
Example:
Directional Alternative: Boys taught using inquiry method has higher mean score
in science than girls taught using the same method of teaching is Directional
Alternative.
Non-Directional Alternative: The mean score in science for boys taught
Steps in Hypothesis Testing
Hypothesis testing is the use of statistics to determine the probability that a given
hypothesis is true. The usual process of hypothesis testing consists of four steps:
1. Formulate the null hypothesis and the alternative hypothesis.
2. Identify a test statistic that can be used to assess the truth of the null hypothesis.
3. Compute the P-value, which is the probability that a test statistic at least as significant
as the one observed would be obtained assuming that the null hypothesis were true (the
smaller the p-value, the stronger the evidence against the null hypothesis).
4. Compare the -value to an acceptable significance value (sometimes called an alpha
value). If P< α, that the observed effect is statistically significant, the null hypothesis is
ruled out, and the alternative hypothesis is valid.
Two-tailed and one-tailed Tests
One important concept in significance testing is whether you use a one-tailed or two-
tailed test of significance. The answer is that it depends on your hypothesis. When your
hypothesis states the direction of the difference or relationship, you use a one-tailed
probability. For example, a one-tailed test would be used to test these null hypotheses:
Females will not score significantly higher than males on an IQ test. Blue collar workers

58
Kotebe University of Education, Department of Geography and
Environmental studies
are will not buy significantly more product than white collar workers. Superman is not
significantly stronger than the average person. In each case, the null hypothesis
(indirectly) predicts the direction of the difference. A two-tailed test would be used to test
these null hypotheses: There will be no significant difference in IQ scores between males
and females. There will be no significant difference in the amount of product purchased
between blue collar and white collar workers. There is no significant difference in
strength between Superman and the average person. The one-tailed probability is exactly
half the value of the two-tailed probability.
There is a raging controversy (for about the last hundred years) on whether or not it is
ever appropriate to use a one-tailed test. The rationale is that if you already know the
direction of the difference, why bother doing any statistical tests. While it is generally
safest to use two-tailed tests, there are situations where a one-tailed test seems more
appropriate. The bottom line is that it is the choice of the researcher whether to use one-
tailed or two-tailed research questions.

59

You might also like