1 Quantitative Methods Note
1 Quantitative Methods Note
Environmental studies
Parametric statistics are based on assumptions about the distribution of population from
which the sample was taken. Nonparametric statistics are not based on assumptions, that
is, the data can be collected from a sample that does not follow a specific distribution.
2
Kotebe University of Education, Department of Geography and
Environmental studies
Data Presentation: The organized data can now be presented in the form of tables,
charts diagrams and graphs. At this stage, large data are presented in a very summarized
and condensed manner.
Data Analysis: This is the stage where we critically study the data. The purpose of data
analysis is to dig out information useful for decision making.
Data Interpretation: This is the stage where draw valid conclusions from the results
obtained through data analysis. If the data that have been analyzed are not properly
interpreted, the whole purpose of the investigation may be defected and misleading
conclusion may be drawn.
1.4 Application of statistics
The science of statistics is very essential for research and decision making processes in
all aspects of human life. The following are some of the areas for which statistical
analysis is required:
To represent the facts in the form of numerical data
To summarize a mass of data into a few presentable understandable and precise
figures.
To Predict or forecast future trend.
To help select a course of action among a number of alternatives.
To help in formulating policies.
3
Kotebe University of Education, Department of Geography and
Environmental studies
b) Qualitative variables are the result of categorizing or describing attributes of a
population. Hair colour, blood type, ethnic group are examples of qualitative data.
Qualitative data are generally described by words or letters. For instance, hair colour
might be black, dark brown, light brown, blonde, grey, or red. Blood type might be AB,
O, or B.
Measurement scales
There are four types of measurement scales for variables:
1. Nominal scale: - “Nominal “is a Latin word for “name”. This is a scale for grouping
individuals into different categories. In such measures numbers are assigned just for
recognition. For example when you go out to collect data from males and females you
may decide to assign the number 1 to males and 2 to females. It does not mean that males
score lower or are inferior to females. Such measures are merely to identify, to name and
to differentiate males from females, one player from another. Nominal scale is the
simplest of all measurement scales. In this scale, one is different from the other and +, -,
*, /, impossible, comparison is impossible
2. Ordinal scale: - “ordinal” is a Latin word, meaning “order”. It is a scale for grouping
and ordering of individuals in to different categories. Data consisting of an ordering or
ranking of measurements are said to be on an ordinal scale of measurements. This scale
has magnitude only. Examples: military ranks, ranks in race, ranks of collage academic
staff, e t c. One is different from and grater /better/ less than the other and +, -, *, / are
impossible, comparison is possible.
Ordinal scales data contain and convey more information than the nominal scale data, for
relative magnitudes are known, however, quantitative comparisons are impossible. Let us
consider students’ performance in a test. Student A scored 80%, B had 75%, C 63% and
D 60%. Based on their performances, student A is first, B is second, C is third and D is
fourth. However, the difference between first and second position is 5%, between second
and third position you have 12% and between third and fourth position is 3%. In each
situation there is no equal interval.
3. Interval scale: This is a scale with magnitude and equal interval but no absolute zero.
There is no true zero point (arbitrary zero point not real) and there is no physical
4
Kotebe University of Education, Department of Geography and
Environmental studies
significance to the zero point. The zero value is defined conventionally since it does not
imply a lack of the quantity under study.
There is a constant interval size between any adjacent units on the measurement scale.
Example: 0C, 0F (Measuring units of temperature) which do not mean that there is no
heat.
In this measurement scale One is different, better/greater and by a certain amount of
difference than another (Possible to add and subtract but multiplication and division are
not possible)
37oC – 35oC = 2oC and 45oC – 43 oC= 2oC
40oC = 2(20oC) But this does not imply that an object which is 40 oC is twice as hot as an
object which is 20 oC. Interval scale data convey better information than nominal and
ordinal scale data.
4. Ratio scale: It is a measurement scale in which there is a constant interval size
between any adjacent units on the measurement scale. There exists a zero point on the
measurement scale and that there is a physical significance to this zero point. This means
that a measured value of zero really implies the absence of the phenomenon being
studied. For example, zero distance means no distance travelled and zero precipitation
means that none has been recorded.
Examples: height, weight, volume, etc.
One is different, larger /taller/ better/ less by a certain amount of difference and so much
times than the other and (+, -, *, / are possible on this scale).This measurement scale
provides better information than interval scale of measurement.
1.6 Sources of data and methods of data collection
Facts or figures, which are numerical or otherwise, collected with a definite purpose, are
called data. The required data can be obtained from either a primary source or a
secondary source.
Primary source: Is a source of data that supplies first-hand information for the use of the
immediate purpose. Primary data is the data that is collected for the first time through
personal experiences or evidence, particularly for research.
The term ‘primary data collection’ is generally used to refer to the strategy, in which
researchers collect information for a specific purpose ‘in the field’.
5
Kotebe University of Education, Department of Geography and
Environmental studies
The main advantage of this approach is that the quality of the data is known and better
understood, i.e., the researchers are better able to assess the effects of potential sources of
error in the data because they have been intimately involved in their collection and
recording.
Its main disadvantages are that it is often slow, arduous, labour intensive, and frequently
expensive. It may take months, if not longer, to collect enough data of the type required
for the study.
Secondary data is a second-hand data that is already collected and recorded by some
researchers for their purpose, and not for the current research problem. Secondary data:
are data collected from a secondary source.
Secondary source: are individuals or agencies, which supply data originally collected for
other purposes by them or others. Usually they are published or unpublished materials. It
is accessible in the form of data collected from different sources such as government
publications, censuses, internal records of the organisation, books, journal articles,
websites and reports, etc.
This method of gathering data is affordable, readily available, and saves cost and time.
However, the one disadvantage is that the information assembled is for some other
purpose and may not meet the present research purpose or may not be accurate.
Discrete Vs continuous data
Discrete data (countable) is information that can only take certain values. These values
don’t have to be whole numbers but they are fixed values – such as shoe size, number of
teeth, number of kids, etc. Discrete data includes discrete variables that are
finite, numeric, countable, and non-negative integers (5, 10, 15, and so on).
Continuous data (measurable) is data that can take any value. Height, weight,
temperature and length are all examples of continuous data. Continuous data changes
over time and can have different values at different time intervals like weight of a person.
Methods of data collection
There are three major methods of data collection
I. Observation or measurement
6
Kotebe University of Education, Department of Geography and
Environmental studies
In this method, data can be obtained through direct observation or measurement. It
requires training of persons who measure in order to insure the use of standard procedure.
It provides accurate information but it is expensive and inconvenient
II. Interviews and Questionnaires
Questionnaire: - are written documents which instruct the readers or listeners to answer
the questions written on it. There are three ways of collecting information under this
method.
a) Face to face interviews (Questionnaires in charge of interviewers)
b) Telephone interviews
c) Mailed questionnaires (Self-administered questionnaires returned by mail)
III. The use of documentary sources
It is extracting of information from existing sources.
1.7 Geographic Data: Spatial and Aspatial Data
By ‘data’, geographers usually mean collections of facts and figures. Data are the
observed values of a variable. Geographers usually refer data as ‘aspatial’, because an
explicit spatial or locational reference is not an integral part of the information they
contain. Aspatial data are attribute data which specifies characteristics at that location
(what, how much, and when).
There is a data type which may be exclusively geographical: spatial data. Such data
consist of observations on geographical individuals which may only be interpreted
satisfactorily when their locations (where) have been taken into consideration. This may
involve a consideration of their absolute locations (site characteristics), or their relative
locations as measured with respect to some benchmark such as the National Grid or sea
level. Spatial data are often collected using maps, plans or charts, but increasingly, they
are to be found as the basic data type of computerised information systems, in which
location provides an obvious and generally tractable method of organisation.
By ‘spatial’, geographers usually mean data which are gathered in the form of points or
dots, lines, areas or surfaces.
1.8 Main Objectives of Quantitative Revolution in Geography
Geography for more than two hundred years was confronted with the problems of
generalization and theory-building. After the Second World War, geographers,
7
Kotebe University of Education, Department of Geography and
Environmental studies
especially those of the developed countries, realized the significance of using
mathematical language rather than the language of literature in the study of geography.
Consequently, empirical descriptive geography was discarded and greater stress was laid
on the formulation of abstract models. Mathematical and abstract models need rigorous
thinking and use of sophisticated statistical techniques.
The diffusion of statistical techniques in geography to make the subject and its theories
more precise is known as the ‘quantitative revolution’ in geography. The application of
statistical and mathematical techniques, theorems and proofs in understanding
geographical systems is known as the ‘quantitative revolution’ in geography
The main objectives of the quantitative revolution in geography are
1. To change the descriptive character of the subject (geo + graphy) and to make it a
scientific discipline;
2. To explain and interpret the spatial patterns of geographical phenomena in a rational,
objective and cogent manner;
3. To use mathematical language instead of the language of literature,
4. To make precise statements (generalizations) about locational order
5. To test hypotheses and formulate models, theories and laws for estimations and
predictions
6. To identify the ideal locations for the various economic activities so that the profit may
be maximized by the resource users; and
7. To provide geography a sound philosophical and theoretical base, and to make its
methodology objective and scientific.
In order to achieve these objectives, the preachers of quantitative techniques stressed on
field surveys for the collection of data and empirical observations.
2. Tabular and graphical display of quantitative data
There are two types of statistical presentation of data - graphical and numerical.
A tabular database, as the name implies is a database that is structured in
a tabular form. It arranges data elements in vertical columns and horizontal rows
Advantages of tabular presentation of data:
Tabulated data can be easily understood and interpreted.
Tabulation facilitates comparison as data are presented in compact and organized
form.
It saves space and time.
8
Kotebe University of Education, Department of Geography and
Environmental studies
Tabulated data can be presented in the form of diagrams and graphs.
Tables
Data can also be presented by means of tables. This is a systematic organization of data
using columns and rows.
Elements of a table
Ideal table should have Number, Title, Column headings and Foot-notes
• Number: Table number for identification in a report
• Title, place: Describe the body of the table, variables
• Time period: (What, how classified, where and when)
• Column Heading: Variable name, No. , Percentages (%), etc.,
• Foot-note(s): to describe some column/row headings, special cells, source, etc.,
The Frequency Distribution Table
A frequency is the number of times a value of the data occurs. A frequency distribution
is the organization of raw data in table form, using classes and frequencies. Frequency
distribution refers to data classified on the basis of some variable that can be measured
such as prices, weight, height, wages etc. The frequency distribution table shows the
number of items falling into each group.
The tables display counts (frequencies) and percentages or proportions (relative
frequencies). The percent columns make comparing the same categories easier.
Displaying percentages along with the numbers is often helpful, but it is particularly
important when comparing sets of data that do not have the same totals. The following
technical terms are important when a continuous frequency distribution is formed:
1. Class frequency - the number of observations corresponding to a particular class.
2. Class Limits - are the lowest and the highest values that can be included in the class.
The lower limit of a class is the value below which there can be no item in the class. The
upper limit of a class is the value above which there can be no item to that class.
Example: 0 – 10, in this class, the lowest value is zero and highest value is 10. The two
boundaries of the class are called upper and lower limits of the class. Class limit is also called
as class boundaries.
3. Class interval - The difference between upper and lower limit of class .
Example: In the class 0 – 10, the class interval is (10 – 0) = 10.
9
Kotebe University of Education, Department of Geography and
Environmental studies
The formula to find class interval is gives on below
S = Smallest value
R = the no. of classes
Example: If the mark of 60 students in a class varies between 40 and 100 and if we want
to form 6 classes, the class interval would be
I= (L-S) / R = 100-40/6=10 Therefore, class intervals would be 40 – 50, 50 – 60, 60 – 70,
70 – 80, 80 – 90 and 90 – 100.
Types of class intervals: There are three methods of classifying the data according to
class intervals namely
a) Exclusive method: When the class intervals are so fixed that the upper limit of one
class is the lower limit of the next class; it is known as the exclusive method of
classification. This method makes continuity of data. The following data are classified on
this basis.
Household Expenditure in birr Better way of expression No of families
0-5000 0 to less than 5000 60
5000-10000 5000 to less than 10000 95
10000-15000 10000 to less than 15000 122
15000-20000 15000 to less than 20000 83
20000-25000 20000 to less than 25000 40
Total 400
Thus, a household with the expenditure of 4999.5 birr will be included in the 0 – 5000
class.
b) Inclusive method (Non-overlapping): In this method, the overlapping of the class
intervals is avoided. Both the lower and upper limits are included in the class interval.
Class interval Frequencies
5-9 7
10-14 4
15-19 3
20-24 2
25-29 1
Total 17
10
Kotebe University of Education, Department of Geography and
Environmental studies
c) Open-end classes- are formed when a class limit is missing either at the lower end of
the first class interval or at the upper end of the last class interval or both are not
specified.
Salary range No of workers
Below 2000 7
2000-4000 5
4000-6000 4
6000-8000 6
Above 8000 3
Total 25
4. Class mark (CM) - midpoint of a class interval. It is found out by adding the upper
and lower limits of a class and dividing the sum by 2
Categorical or Qualitative Frequency Distributions
A categorical frequency distribution represents data that can be placed in specific
categories, such as gender, blood group, & hair colour, etc
Example: The blood types of 25 blood donors are given below. Summarize the data
using a frequency distribution. AB, B, A, O, B, O, B, O, A, O, B, O, B, B, B, A, O, AB,
AB, O, A, B, AB, O, and A
Blood type Frequency
A 5
B 8
O 8
AB 4
Total 25
11
Kotebe University of Education, Department of Geography and
Environmental studies
A grouped frequency distribution is obtained by constructing classes (or intervals) for the
data, and then listing the corresponding number of values (frequency counts) in each
interval.
Creating a Grouped Frequency Distribution
1. Find the largest and smallest values
3. Select the number of classes desired. This is usually between 5 and 20. However
there is no rigidity about it.
4. Find the class width by dividing the range by the number of classes and rounding
up. There are two things to be careful of here. You must round up, not off.
Normally 3.2 would round to be 3, but in rounding up, it becomes 4. If the range
divided by the number of classes gives an integer value (no remainder), then you
can either add one to the number of classes or add one to the class width.
Sometimes you're locked into a certain number of classes because of the
instructions.
5. Pick a suitable starting point less than or equal to the minimum value. You will be
able to cover: "the class width times the number of classes" values. You need to
cover one more value than the range. Follow this rule and you'll be okay: The
starting point plus the number of classes times the class width must be greater
than the maximum value. Your starting point is the lower limit of the first class.
Continue to add the class width to this lower limit to get the rest of the lower
limits.
6. To find the upper limit of the first class, subtract one from the lower limit of the
second class. Then continue to add the class width to this upper limit to find the
rest of the upper limits.
7. Find the boundaries by subtracting 0.5 units from the lower limits and adding 0.5
units from the upper limits. The boundaries are also half-way between the upper
limit of one class and the lower limit of the next class. Depending on what you're
trying to accomplish, it may not be necessary to find the boundaries.
12
Kotebe University of Education, Department of Geography and
Environmental studies
9. Find the frequencies.
10. Find the cumulative frequencies. Depending on what you're trying to accomplish,
it
11. If necessary, find the relative frequencies and/or relative cumulative frequencies.
For example, Let us consider the weights in kg of 50 college students. 42, 62, 46, 54, 41,
37,54,44,32,45,47,50,58,49,51,42,46,37,42,39,54,39,51,58,47,64,43,48,49,48,49,61,41,40
,58,49,59,57,57,34,56,38,45,52,46,40,63,41,51, and 41
Sturges formula to find number of classes is given below
K = 1 + 3.322 log N.
Where K = No. of class
log N = Logarithm of total no. of observations
K=1 + 3.322 log 50= 1+3.22*1.6990=6.64 (Rounded off to 7)
To arrange the data in grouped table we have determine size of class interval. Thus, size
of class interval (C) =
13
Kotebe University of Education, Department of Geography and
Environmental studies
frequencies can be written as fractions, percents, or decimals. Sum of relative frequencies
should equal 1.0 or = 100% by percentage.
RF = class frequency ÷ no. of observations
Class (Pulse rate) Frequency Relative frequency
53 1 0.0625
54 3 0.1875
55 2 0.125
56 2 0.125
57 4 0.25
58 2 0.125
60 1 0.0625
64 1 0.0625
Total 16 1.00
Cumulative relative frequency is the accumulation of the previous relative frequencies.
To find the cumulative relative frequencies, add all the previous relative frequencies to
the relative frequency for the current row.
Example of a simple frequency distribution-families with number of children- 5, 7, 8, 1,
5, 9, 3, 4, 2, 2, 3, 4, 9, 7, 1, 4, 5, 6, 8, 9, 4, 3, 5, 2, 1
Class (Pulse rate) Frequency Relative frequency Cumulative frequency
9 3 0.12 3
8 2 0.08 5
7 2 0.08 7
6 1 0.04 8
5 4 0.16 12
4 4 0.16 16
3 3 0.12 19
2 3 0.12 22
1 3 0.12 25
Total 25 1.00 25
Tables are a good way of organizing and displaying data. But graphs can be even more
helpful in understanding the data. Graphic presentation represents a highly developed
body of techniques for elucidating, interpreting, and analyzing numerical facts by means
of points, lines, areas, and other geometric forms and symbols. In graphical Presentation
we look for the overall pattern and for striking deviations from that pattern. Over all
pattern usually described by shape, centre, and spread of the data. An individual value
that falls outside the overall pattern is called an outlier.
14
Kotebe University of Education, Department of Geography and
Environmental studies
There are no strict rules concerning which graphs to use. It is a good idea to look at a
variety of graphs to see which is the most helpful in displaying the data. We might make
different choices of what we think is the “best” graph depending on the data and the
context. Our choice also depends on what we are using the data for.
Common Types of Graphs
Histogram
A histogram is a graphic version of a frequency distribution. The graph consists of bars of
equal width drawn adjacent to each other. The horizontal scale represents classes of
quantitative data values and the vertical scale represents frequencies. The heights of the
bars correspond to frequency values. Histograms are typically used for large, continuous,
quantitative data sets. A frequency polygon can also be used when graphing large data
sets with data points that repeat.
A bar chart consists of a series of rectangular bars where the length of the bar represents
the quantity of frequency for each category if the bars are arranged horizontally. If the
bars are arranged vertically, the height of the bar represents the quantity. In a bar graph,
the length of the bar for each category is proportional to the number or percent of
individuals in each category. Bars may be vertical or horizontal.
A pie chart a circular graph that is useful in visually showing how a total quantity is
distributed among a group of categories. The “pieces of pie” represent the proportions of
the total that fall into each category. In a pie chart, categories of data are represented by
wedges in a circle and are proportional in size to the percent of individuals in each
category.
15
Kotebe University of Education, Department of Geography and
Environmental studies
Annual Estimates of U.S. Population 65 Years and Over by Race, 2003
White
Black
American Indian
Asian
Pacific Islander
Two or more
20
15
10
0
1900 1920 1940 1960 1980 2000 2020 2040 2060
16
Kotebe University of Education, Department of Geography and
Environmental studies
average represents a whole series and as such, its value always lies between the minimum
and maximum values and generally ’it is located in the centre or middle of the
distribution.
There are five main measures of central tendency: Arithmetic Mean, Geometric Mean,
Harmonic Mean, Mode and Median. Each of these measures describes a different
indication of the typical or central value in the distribution. Of the above mentioned five
important averages Arithmetic Average, Median and Mode are the most popular ones.
Arithmetic Mean or Simply Mean:
Mean for sample is denoted by symbol ‘M or x̅ (‘x-bar’)’ and mean for population is
denoted by ‘µ’ (mu). It is one of the most commonly used measures of central tendency
and is often referred to as average. Arithmetic Mean can be calculated as the sum of all
the values in the dataset divided by the number of values. That is, If X is the observation
which takes the value X1, X2, X3, … Xn and n represents the total number of
observation; AM (= ) equals:
17
Kotebe University of Education, Department of Geography and
Environmental studies
Thus, X̅ = 450/ 10 = 45
Thus, the mean obtained for the above data is 45
Computation of Mean for Grouped Data
Grouped data: A data that is categories or organised is termed as grouped data. Mainly
such data is organised in frequency distribution. There are discrete and continuous
grouped data. Discrete data is a numerical type of data that includes whole, concrete
numbers with specific and fixed data values determined by counting. Discrete
data includes discrete variables that are finite, numeric, countable, and non-negative
integers (5, 10, 15, and so on). On the other hand, continuous data includes complex
numbers and varying data values that are measured over a specific time interval. Height,
weight, temperature and length are all examples of continuous data.
Computation of Mean for Grouped Discreet Data
In a discrete series the values of the variable are multiplied by their respective
frequencies and the products so obtained are totalled. This total is divided by the number
of items, which in a discrete series, is equal to the total of the frequencies. If we have a
discrete frequency distribution with frequencies f1, f2, … fn associated with the values
X1, X2, …, Xn of the variable it can be seen that the sum of all the items equals
f1x1+f2x2+…+fnxn since there are now f1 items with value x1, f2 items with value x2,
and so on. The total of items is f=N. = Summation
18
Kotebe University of Education, Department of Geography and
Environmental studies
x̅ = 623/46=13.54
Computation of Mean for Grouped Continuous Data
If, the data is grouped with class interval we do not know the exact values of each item.
In this case what we do is estimation based on some assumption. The assumption is that
all the items within a particular class are concentrated at the mid value of the class and
thus fx corresponds to the f items of a class equals fm where m is the mid point of class
interval. Thus,
Formula:
The steps followed for computation of mean with grouped data are as follows:
Step 1: The data is arranged in a tabular form with marks grouped in categories with class
interval of 5.
Step 2: Once the categories are created, the marks are entered under frequency column
based on which category they fall under.
Step 3: The midpoints of the categories are computed and entered under X.
Step 4: fX is obtained by multiplying the frequencies and midpoints for each category.
Step 5: fX for all the categories are added to obtain fX, in case of our example it is
obtained as 780
Step 6: The formula M = fX/ N is used, N is equal to 30.
x̅ = fX/ N M = 780/ 30= 26
19
Kotebe University of Education, Department of Geography and
Environmental studies
Computation of Mean by Shortcut Method (with Assumed Mean)
The above direct method of the calculation of arithmetic average can be used only when
the items are few and the size of the figures is small. If it is not so, there would be
considerable difficulty in the calculation of the arithmetic average. In such situations, a
short cut method with the help of assumed mean can be computed. The assumed
mean method is a technique used in statistics to calculate the arithmetic mean. So we can
assume any arbitrary mean to find out the deviations of items from this assumed [Link]
is a useful shortcut method to calculate the mean from a set of data. The formula to
calculate mean of groped and ungrouped data through short cut method is as follows:
The short-cut method of calculating arithmetic average should be used in all cases as it
saves time and gives accurate results. The process of calculating the arithmetic average
by the short-cut method from a continuous groped data can be summarized as follows:
(i) Assume as average, the midpoint of a class which is in the middle of the series.
Technically any class can be chosen, but if the class chosen is in the middle of the
distribution there is considerable facility in calculations.
(ii) Calculate the deviations of the items (midpoints in case of continuous series) from the
assumed mean.
(iii) Divide the deviations by a common factor or magnitude of the class interval. These
deviations are known as step deviation or deviation in class-interval limits
(iv) Multiply the deviations with the respective frequencies of the various classes and
total the products, taking into account the algebraic signs (plus or minus).
(v) Divide this total by the total of the frequencies and if step deviations have been taken
multiply the result by the common factor or the magnitude of the class-interval.
(vi) Add this figure to the assumed average and the resulting figure would be the actual
arithmetic average of the series.
Example: Find the mean of this set of values by the assumed mean method. 91, 48, 9, 37,
6, 42, 23, 45, 63, 84, 88, 29, 28, 10, 8 by taking the assumed mean as 40
20
Kotebe University of Education, Department of Geography and
Environmental studies
Example: Let us discuss the steps followed for computation of mean with the help of an
example given below:
Class Intervals Frequencies (f) Midpoint (X) x′ = {(X −AM)/ i} f x′
35- 39 5 37 3 15
30-34 7 32 2 14
25-29 5 27 1 5
20-24 6 22 0 0
15-19 4 17 -1 -4
10-14 3 12 -2 -6
N=30 fx'=24
21
Kotebe University of Education, Department of Geography and
Environmental studies
2) When there are open ended classes, such as 10 and above or below 5, mean cannot be
computed. In such cases median and mode can be computed. This is mainly because in
such distributions midpoint cannot be determined to carry out calculations.
3) If a score in the data is missing or lost or not clear, then mean cannot be computed
unless mean is computed for rest of the data by not considering the lost score and
dropping it all together.
4) Arithmetic average is no doubt easy to calculate but in a relative sense its calculation
may be more difficult than that of mode or median as they can be located merely by
inspection.
5) It is not suitable for data that is skewed or is very asymmetrical as then in such
cases mean will not adequately represent the data.
22
Kotebe University of Education, Department of Geography and
Environmental studies
23
Kotebe University of Education, Department of Geography and
Environmental studies
24
Kotebe University of Education, Department of Geography and
Environmental studies
To determine the median:
1. Put the data in order from smallest to largest.
2. Determine the number in the exact center.
i. If there are an odd number of data points, the median will be the number in the absolute
middle.
ii. If there is an even number of data points, the median is the mean of the two center data
points, meaning the two center values should be added together and divided by 2.
Example: Consider the data set: 17, 10, 9, 14, 13, 17, 12, 20, and 14
Step 1: Put the data in order from smallest to largest. 9, 10, 12, 13, 14, 14, 17, 17, 20
Step 2: Determine the absolute middle of the data. 9, 10, 12, 13, 14, 14, 17, 17, 20
th
Since the number of data points is odd then median = size of (9+1/2) = 10/2 item = 5th
item in the data set is 14.
Example: Suppose the values of the observation are 1, 2, 3, 5, 8 and 10
In this case the median is the arithmetic mean of the two middle observations, i.e., n/2=3
and (n/2)+1=4. The median in this case is found by averaging the 3 rd and 4th items which
is (3+5)/2=4.
Median for grouped data (discrete variable)
In a discrete series also the items are first arranged according to the ascending or
descending order of magnitude and their respective frequencies are written against them.
After this, the frequencies are cumulated and then the value of the middle item can be
easily located. The following example illustrates the procedure
Example: Find the median size of the shoe from the following data.
Size of shoes Number of pairs (frequency) Cumulative frequency
5 30 30
25
Kotebe University of Education, Department of Geography and
Environmental studies
6 40 70
7 50 120
8 150 270
9 300 570
10 600 1170
11 950 2120
Total 2120
Solution: Median is the value of [2120/2]th = 1060th observation. Since the observations
are arranged in ascending order (size-wise), the size of 1060 th observation is easily
determined by constructing the cumulative frequency. Thus, the median size of shoes
sold is 10.
Example 2: The following data gives the distribution of the height of students
Height (in cm) 160 150 152 161 156 154 155
Number of students 12 8 4 4 3 3 7
Solution:
Arranging the data in ascending order of magnitude, we obtain
Height (in cm) 160 150 152 161 156 154 155
Number of students 12 8 4 4 3 3 7
Cumulative frequency 8 12 15 22 25 37 41
26
Kotebe University of Education, Department of Geography and
Environmental studies
Here, the total number of items is 41 i.e., an odd number. Hence, the median is [41 + 1] /
2th i.e., 21st item. From the cumulative frequency table, we find that median i.e., 21 st item
is 155, (All items from 16th to 22nd are equal, each 155).
Median for grouped data (continuous variable):
When the median of a continuous frequency distribution has to be determined there is one
difficulty. The value of the median lies in a class interval, and to get a definite figure,
interpolation has to be done. Suppose, for example it is found that the value of the median
lies in the 20 to 30 class interval whose frequency is 40. Now to find out the value of the
median we have to take recourse to interpolation and to apply a particular formula.
The formula of interpolation to find out the median is:-
27
Kotebe University of Education, Department of Geography and
Environmental studies
Solution: The median of 2063 cases is the income of [N/2] th worker, which is [2063/2]th =
1031.5 which is 1032nd worker arranged in ascending order of income. From the
cumulative frequency this worker has his income in the class 47.5-52.5. But it is
impossible to determine his exact income. We, therefore, resort to approximation by
assuming that the 410 workers of this class are distributed uniformly across the interval
47.5-52.5. The median worker is [1032-656 (the CF of the preceding median class)]
=376th of these 410, and hence, the value corresponding to him can be approximated as
And this is 47.5+ (376/410)*(52.5-47.5) =52.1
Merits of median
(i) It can be easily calculated and it is understood without any difficulty.
(ii) It is not affected by the values of the extreme items and as such is sometimes more
representative than arithmetic average.
(iii) Even if the value of the extremes is not known median can be calculated if the
number of items is known.
(iv) It can be located merely by inspection in many cases.
(v) It gives best results in a study of those phenomena which are incapable of direct
quantitative measurement
Drawbacks of median
(i) Median may not be representative of a series in many cases. This is specially so when
there are wide variations between the values of different items
(ii) It is not suitable for further algebraic treatment.
(iii) When median has to be calculated in continuous series it requires interpolation. The
assumption of the interpolation, that all the frequencies of the class-interval are uniformly
spread over their values in the class-interval, may not be actually true. In most cases it
will not be true.
(iv) If big or small items in a series are to receive greater importance median would be an
unsuitable average. Median ignores the values of extreme items.
28
Kotebe University of Education, Department of Geography and
Environmental studies
(v) Median is more likely to be affected by the fluctuations of sampling than the
arithmetic average.
(vi) The arrangement of items in ascending or descending order is sometimes very
tedious
Mode
Mode is denoted by symbol ‘Mo’ is the value that occurs with the greatest frequency in a
given distribution. The mode is the value which occurs or repeats itself the greatest
number of times. It is the easiest score to spot in a [Link] is most useful as a
measure of central tendency when examining categorical data. It is the only way to
express the central tendency of a nominal level variable. For the normal distribution, the
mode is also the same value as the mean and median . There may be one mode; multiple
modes, if more than one number occurs most frequently; or no mode at all, if every number
occurs only once. To determine the mode:
1. Put the data in order from smallest to largest, as you did to find your median.
2. Look for any value that occurs more than once.
3. Determine which of the values from Step 2 occurs most frequently.
The data having one mode is called uni-modal distribution. For example, in the
following list of numbers, 16 is the mode since it appears more times in the set than any
other number: 3, 3, 6, 9, 16, 16, 16, 27, 27, 37, 48
The data having two modes is called bi-modal distribution. Example: 3, 3, 3, 9, 16, 16,
16, 27, 37, 48 In this above example, both the number 3 and the number 16 are modes as
they each occur three times and no other number occurs more often.
The data having more than two modes is called multi-modal distribution. It should also
be noted that a distribution may not have mode if no number in a set of numbers occurs
more than once, that set has no mode.
Example: 3, 6, 9, 16, 27, 37, 48
29
Kotebe University of Education, Department of Geography and
Environmental studies
Example
Find the mode of the following data?
38 39 40 41 42 43
Size
Number of items 32 15 24 27 44 38
Solution: From the above table, it is observable that the size 42 is the entry with highest
frequency (44).
Mode in case of Continuous Grouped Data:
In case of continuous grouped data, the mode would lie in the class that carries the
highest frequency. This class is called the modal class. The exact location of mode in
a class-interval is done by interpolation
In a continuous series the determination of mode involves three steps.
First, prepare the frequency distribution table in such a way that its first column consists
of the observations and the second column the respective frequency.
Second by the process of grouping, the class in which there is maximum concentration
has to be located. This class is called the modal class.
Third, calculate mode, using the formula
The interpolation is made by the use of the following formula:
M0= L1+
Where:
L1 is the lower limit of the modal class,
f0 is the frequency of the preceding class (class next below modal class),
f1 is the frequency of the modal class,
f2 is the frequency of the following class (class next above modal class)
i is the width of the modal class
Example: compute the mode of the following table
Wage group Frequency
14-18 6
18-22 18
22-26 19
30
Kotebe University of Education, Department of Geography and
Environmental studies
26-30 12
30-34 5
34-38 4
38-42 3
42-46 2
46-50 1
50-54 0
54-58 1
In the above given data 22-26 is the modal class, since it has the largest frequency. The
lower limit of the modal class is 22, its upper limit is 26, its frequency is 19, the
frequency of the preceding class is 18, and of the following one is 12. The class interval
is 4. Using the method mentioned above we can determine the mode as follows;
Mode=22+ (19-18)/ (2x19-18-12) x4=22.5
Example: A survey on the heights (in cm) of 30 students of the same batch was
conducted at a university. The data so obtained has been organized in the table given
below. Find the mode
Height (in cm) Number of students
120 – 125 3
125 – 130 5
130 – 135 11
135 – 140 6
140 – 145 5
Total 30
The Modal class = 130 – 135 as its frequency is the highest (11)
The Lower limit of the modal class = (L) = 130
Frequency of the modal class = 11
Frequency of the preceding modal class = 5
Frequency of the next modal class = 6
Size of the modal class interval = (h) = 5
Putting the values in the formula,
Mode = 130 + (11-5)/ (2×11-5-6)5= 130 + (6/11) (5) = 130 + 0.54×5 = 132.72
31
Kotebe University of Education, Department of Geography and
Environmental studies
Therefore, mode = 132.72
Exercise: based on the data presented in the table below, compute the mode
Class boundaries Frequency
29.5---39.5 8
39.5---49.5 87
49.5---59.5 190
59.5---69.5 304
69.5---79.5 211
79.5---89.5 85
89.5---99.5 20
f =905
Merits of Mode
(i) It possesses the merit of simplicity. It can be determined without much mathematical
calculation. In a discrete series mode can be located even by inspection. In this respect,
like median, it has an advantage over arithmetic average
(ii) It is commonly understood. As has been said earlier, mode an average which people
use in their day-to-day expressions.
(iii) Since mode is the most common item of a series it is not an› isolated example like
the median: Unlike arithmetic average it cannot be a value which is not found in the
series.
(iv) Mode is not affected by the values of extreme items provided. They adhere to the
natural law relating to extremes.
v) For the determination of mode it is not "necessary to know the values of all the items
of a series. If the point of norm or maxi› mum concentration is known it is enough.
Drawbacks of mode
Mode is an unsatisfactory average and has many drawbacks. Some of them are as
follows:
(i) Mode is not capable of further mathematical treatment.
(ii) In many cases it may be impossible to set a definite value of .ode. There may be 2, 3
or more modal values
32
Kotebe University of Education, Department of Geography and
Environmental studies
(iii) Mode may be unrepresentative in many cases. If in a series 1000 items 20 have a
particular value and other values have frequencies is than 20, it does not necessarily mean
that the value whose frequency 20 is the typical or average value. In such cases data
should be converted into class intervals of a bigger magnitude.
The relationship between mean, median and mode
In a normal/ symmetrical distribution the mean, median and mode are identical. In actual
practice, however, symmetrical distributions are very rare, and data usually give a
symmetrical curve.
Example: the mean and median of a distribution is 44.6 and 44.05 respectively. Find the
mode. We have, median = 44.05 and mean = 44.6. We know,
Mode=3Median−2Mean
=3(44.05)−2(44.6)
33
Kotebe University of Education, Department of Geography and
Environmental studies
=132.15−89.2=42.95
4. Measures of Scale/Variation/Dispersion/Spreadness
Measuring of central tendency is to determine a single figure to represent a whole series.
Two different distributions may have the same mean, mode and median but can be quite
different in their scatter about the mean. So, the dispersion in itself is a very important
property of a distribution and needs to be measured. Average by itself is not a good
indication of quality of the sample. You need to know the variance to make any educated
assessment. These are statistical procedures for describing the nature and extent of
differences among the information in the distribution. The average income in a
community hides the distribution of income and does not show the quality of the
sample/data. The measures of dispersion (not central tendency) bring out this inequality.
A measure of dispersion/variation is defined as a statistics signifying the extent of the
scatteredness of items around a measure of central tendency. Dispersion‟ or `variation‟ in
statistics is the degree of spread of each individual item or value from the central value in
the given distribution. According to Minium, King and Bear (2001), measures of
variability express quantitatively the extent to which the score in a distribution scatter
around or cluster together. The measures of dispersion are also called the average of
second order, because here we consider the arithmetic mean of the deviations from the
mean of the values of the individual items.
Statistical measures of variation are numerical values that indicate the variability inherent
in a set of data measurements. The greater the similarity of the scores to each other lower
would be the measure of variability or dispersion. The less the similarity of the scores are
to each other, higher will be the measure of variability or dispersion. In general, the more
the spread of a distribution, larger will be the measure of dispersion.
In measuring dispersion, it is imperative to know the amount of variation (absolute
measure) and the degree of variation (relative measure). In the former case, we consider
the range, mean deviation, standard deviation etc. In the latter case, we consider the
coefficient of range, the coefficient of mean deviation, the coefficient of variation etc.
34
Kotebe University of Education, Department of Geography and
Environmental studies
Thus, there are two broad classes of the measures of dispersion or variability. They are
absolute measure of dispersion and relative measure of dispersion.
Absolute dispersion usually refers to the standard deviation, a measure of variation from
the mean. The units of standard deviation are the same as for the data. In other words,
absolute measure is expressed in terms of the original units of a distribution. Therefore,
absolute dispersion is not suitable for comparing the variability of two distributions since
the two variables are expressed and measured in two different units. For instance, the
variability in body height (cm) and body weight (kg) cannot be compared because the
absolute measure (standard deviation) is expressed in cm and kg. The absolute measure is
also not appropriate for two sets of scores expressed in the same units with wide
divergence in means (central value). Nevertheless, absolute measures are widely used,
except in the exceptional cases like above. The absolute measures include range, mean
deviation, standard deviation, and variance.
Relative dispersion, sometimes called the coefficient of variation, is the result of dividing
the standard deviation by the mean and it may be presented as a quotient or as a
percentage. Thus, relative measures are computed from the absolute measures of
dispersion and its corresponding central values. A low value of relative dispersion usually
implies that the standard deviation is small in comparison to the magnitude of the mean.
The six most common measures of variation are the range, quartile deviation/semi-inter-
quartile range, absolute mean deviation, variance, standard deviation, and coefficient of
variation.
Range:
Range is the simplest possible measure of dispersion. The range of any distribution is the
difference between the highest and lowest values in the series.
Symbolically R= X max- X min or R = L - S
Where, R = Range
L = X max= maximum value or largest value
S = X min= minimum or smallest value.
Example: The yields (kg per plot) of a cotton variety from five plots are 8, 9, 8, 10 and
11. Find the range
35
Kotebe University of Education, Department of Geography and
Environmental studies
Solution L=11, S = 8. Range = L – S = 11- 8 = 3
In a frequency distribution, range is given by the difference between the lower limit of
the lowest class and the upper limit of the highest class.
Salary 1800 2000 2400 2800 3000 3200 3400 4000 4500 5000 5400 7000
(In birr)
No of 20 25 26 27 30 30 20 40 24 24 94 40
Lectures
Solution:
Range =X max- X min
36
Kotebe University of Education, Department of Geography and
Environmental studies
=7000-1800 = 5200 Birr
Coefficient of Range = Xmax -Xmax = 7000 -1800=5200
37
Kotebe University of Education, Department of Geography and
Environmental studies
(iii) It cannot be computed if the distribution has open-end class
38
Kotebe University of Education, Department of Geography and
Environmental studies
39
Kotebe University of Education, Department of Geography and
Environmental studies
40
Kotebe University of Education, Department of Geography and
Environmental studies
Calculation of quartiles for continuous grouped data: Find the first and the third
quartiles of the following data.
Distribution of workers by average monthly income
Group no. Monthly earnings No. workers CF
1 27.5-32.5 120 120
2 32.5-37.5 152 272
3 37.5-42.5 170 442
4 42.5-47.5 214 656
5 47.5-52.5 410 1066
6 52.5-57.5 429 1495
7 57.5-62.5 568 2063
Total 2063
Solution: First quartile:
First locate the Q1 by computing nN/4 = 1(2063)/4=515.75
Q1 class is the class that contains the 515.75 th item. This belongs to (from the cumulative
frequency) 42.5-47.5 class. Substituting into the above formula:
=44.22
Interpretation: first quarter (25 percent) of the workers income is less than or equal to
Birr 44.22.
Third quartile:
First locate the Q3 by computing nN/4 = (3*2063)/4=1547.25
Q3 class is the class that contains the 1547.25 th item. This belongs to (from the
cumulative frequency) 57.5-62.5 class. Substituting into the above formula:
=57.96
Interpretation: Third quarter (75 percent) of the workers income is less than or equal to
Birr 57.96.
The semi-inter-quartile range or quartile deviation of the above example could be
computed as QD= (Q3-Q1)/2 or Inter quartile range/2.
QD= (57.96-44.22)/2= 13.74/2=6.87
41
Kotebe University of Education, Department of Geography and
Environmental studies
Example: Find the quartile deviations for the four distributions given in the four
workshops.
First we have to find Q3, Q2, and Q1 for each workshop Data.
Calculation of Quartile Deviation
WS A WS B WS C WS D
Location of Q2 25.3 25.61 25.07 25.25
Location of Q1 23.41 23.07 22.5 22.75
Location of Q3 27.64 28.0 28.17 28.17
Quartile deviation (27-64-23.41)/2=2.12 (28-23.07)/2=2.46 2.83 2.71
For workshop A the QD is Birr 2.1 and median is 25.3. This means that if the distribution
is symmetrical the number of workers, whose wages vary between (25.3-2.1) Birr 23.2
and (25.3+2.1) Birr 27.4, shall be just half of the total number of workers. As this
distribution is not symmetrical the distance between Q1 and the median (Q2) is not the
same as between Q3 and the median, hence the interval defined by median plus and
minus QD will not be exactly the same as given by the value of the two quartiles. Under
such conditions the range between Birr 23.2 and Birr 27.4 will not include precisely 50%
of the workers.
Quartile deviation is an absolute measure of dispersion and cannot be used for comparing
the variability of any two series. If Quartile deviation is to be used for comparing the
variability of any two series it is necessary to convert the absolute measure to a
coefficient of Quartile Deviation (CQD). If Quartile deviation is divided by the average
value of the two quartiles, a relative measure of dispersion is obtained. It is called the
Coefficient of Quartile Deviation
Symbolically,
WS A WS B WS C WS D
Characteristics of QD
42
Kotebe University of Education, Department of Geography and
Environmental studies
The size of the quartile deviation gives an indication about the uniformity or otherwise of
the size of the items of a distribution. If the quartile deviation is small it denotes large
uniformity. Thus, a CQD may be used for comparing uniformity or variation in different
distributions.
The quartile deviation possesses the merits of simple calculation and easy
understandability. It is commonly understood and its calculation does not involve any
mathematical intricacies. These are the points in favor of quartile deviation but there are a
large number of points which go against it.
Quartile deviation is neither based on all the observations of the data, nor is it capable of
further algebraic treatment. It is affected to a considerable extent by the fluctuations of
sampling. A change in the value of a single item may in certain cases affect its value
considerably. Thus quartile deviation is not a very good measure of dispersion,
particularly for series in which the variation is considerable. However, for rough studies
quartile deviation may give an approximate idea of the extent of variability in a series.
The mean deviation will answer the defect of QD. Mean deviation is the arithmetic
average of the variations (deviations) of the individual items of the series from a measure
of their central tendency (commonly used are median and mean).
Calculating AMD involves
(i) Calculate the mean ( ) or median (Me)
(ii) Record the deviations = of each of the items, ignoring the sign
(iii) Find the average value of the deviations
43
Kotebe University of Education, Department of Geography and
Environmental studies
Serial No. Marks
1 10 8
2 12 6
3 14 4
4 15 3
5 16 2
6 18 0
7 19 1
8 20 2
9 23 5
10 25 7
11 30 12
Total 50
44
Kotebe University of Education, Department of Geography and
Environmental studies
SS xi x
2
3. Compute the mean of the squared deviations. Deviations are squared so as to get rid of
negative signs.
Standard deviation for the for ungrouped data is
45
Kotebe University of Education, Department of Geography and
Environmental studies
Why the sum of the squares in the numerator is divided by n-1. This divisor is called the
degrees of freedom and when estimating a population variance or SD, it is necessary to
divide by the degrees of freedom to have an estimate that, on the average, exactly equal
to population variance and population SD.
The standard deviation is a measure of the amount of variation or dispersion of a set of
values. A low standard deviation indicates that the values tend to be close to
the mean (also called the expected value) of the set, while a high standard deviation
indicates that the values are spread out over a wider range. A SD of zero value or
variance of zero value means that there is no variation at all in the data set. In other
words, the observations have the same values.
For a population, this involves summing the squared deviations (sum of squares, SS) and
then dividing by N. The resulting value is called the variance or mean square and
measures the average squared distance from the mean.
Variance
Variance is the expectation of the squared deviation of a random variable from
its population mean or sample mean. Variance is a measure of dispersion, meaning it is a
measure of how far a set of numbers is spread out from their average value. Variance is
the average of the squared deviations of each observation in the set from the arithmetic
46
Kotebe University of Education, Department of Geography and
Environmental studies
mean of all of observations. Or it is the arithmetic mean of the squared deviation of the
observations about their mean. It is often represented by s2 (for sample variance) or σ2 (for
population variance)
Population variance
When you have collected data from every member of the population that you’re
interested in, you can get an exact value for population variance. The population
variance formula looks like this:
Sample variance
When you collect data from a sample, the sample variance is used to make estimates
or inferences about the population variance. The sample variance formula looks like
this:
47
Kotebe University of Education, Department of Geography and
Environmental studies
Steps for calculating the variance
Step 1: Find the mean
Step 2: Find each score’s deviation from the mean
Step 3: Square each deviation from the mean
Step 4: Find the sum of squares
Step 5: Divide the sum of squares by n – 1 or N
Example: Compute the SD for the following data
11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21
The first formula is appropriate: The procedure is
(i) First calculate the mean =x/N=176/11=16
(ii) And then calculate the deviations
x x2 x- (x- )2
11 121 -5 25
12 144 -4 16
13 169 -3 9
14 196 -2 4
15 225 -1 1
16 256 0 0
17 289 1 1
18 324 2 4
19 361 3 9
20 400 4 16
21 441 5 25
Total 176 2926 110
Thus, SD= =3.16 NB: SD is measured in the units of the original observation.
48
Kotebe University of Education, Department of Geography and
Environmental studies
11-13 10
13-15 3
This is an example of continuous frequency series and formula 3 is appropriate.
Class m f fm m- (m- )2 f(m- )2
1-3 2 1 2 -6 36 36
3-5 4 9 36 -4 16 144
5-7 6 25 150 -2 4 100
7-9 8 35 280 0 0 0
9-11 10 17 170 2 4 68
11-13 12 10 120 4 16 160
13-15 14 3 42 6 36 108
100 800 616
=800/100=8
SD= =2.48
V= (expressed as percentage).
49
Kotebe University of Education, Department of Geography and
Environmental studies
In order to decide as to which of the two batsmen, A and B, is the better run-getter, we
should find their batting averages. The one whose average is higher will be considered as
a better batsman.
Batsman A’s average is =500/10=50
Batsman B’s average is =330/10=33
A is a better batsman since his average is 50 as compared to 33 of B.
To determine the consistency in batting we should determine the coefficient of variation.
The less this coefficient the more consistent will be the player. That is the less the
coefficient the less the variability.
A B
2
Score (x) x- (x- ) Score (x) x- (x- )2
12 -38 1444 47 14 196
115 65 4255 12 -21 441
6 -44 1936 76 43 1849
73 23 529 42 9 81
7 -43 1849 4 -29 841
19 -31 961 51 18 324
119 69 4761 37 4 16
36 -14 196 48 15 225
84 34 1156 13 -20 400
29 -21 441 0 -33 1089
Total 500 17498 330 5462
SD of Batsman A= =41.83
V=(41.83/50)*100=83.66 percent.
SD of Batsman B= =23.37
V= (23.37/33)*100=70.8 percent.
B is more consistent (or less variable) since the variation in his case is 70.8 per cent as
compared to 83.66 of A.
EG 2: The following data are given about height of boys and girls.
Boys Girls
Number 72 38
Average height 68m. 61m.
Variance of distribution 9 4
50
Kotebe University of Education, Department of Geography and
Environmental studies
Find
(i) in which sex is there greater variability in individual height? Compare the coefficient
of variation
(ii) Common average height in boys and girls. Combined mean
(iii) SD of boys and girls taken together. Combined SD
(iv) Combined variability.
(v) Did the separation make each group more homogenous?
SOLUTION:
Thus there is a greater variability in the height of boys than that of girls,
=65.58 meter
(iii) Combined SD
= =4.28 meter
(iv) Combined variability (combined coefficient of variation) is
(v) Did the separation make each group more homogenous? Compare the coefficient of
variation of boys, girls and combined coefficient of variation, i.e., coefficient of variation
of total population. The lesser the coefficient the more homogenous the distribution is.
CV of boys height = 4.41%
CV of girls height = 3.28%
CV of boys and girls or total population = 6.53 per cent
51
Kotebe University of Education, Department of Geography and
Environmental studies
Comparisons of Various Measures of Dispersion
Range is easy to calculate and useful statistic to know, but it cannot stand alone as a
measure of spread since it takes into account only two extreme scores and hence it is
extremely sensitive to the size of the sample, and to the sample variability.
Quartile deviation is also easy to calculate but suffer from the disadvantage that they are
not amenable to algebraic treatment.
Mean Deviations: MD is easy to interpret and easier to calculate than SD. But it is not
suitable because we cannot obtain the MD of a combined series from the deviations of
component series.
Standard Deviation: It lends itself to rigorous algebraic treatment and based on all
observations. It is, therefore, quite insensitive to sample size (provided that the sample
size is ‘large enough’) and is least affected by sampling variations. It is used as a standard
scale for measuring deviations from the mean. The only disadvantage of SD is it involves
to much work in its calculation and the large weight it attaches to extreme values because
of the process of squaring involved in its calculation.
The standard deviation is by far the most widely used measure of spread. It takes every
score into account, has extremely useful properties when used with a normal distribution
and is tractable mathematically and, therefore, it appears in many formulas in inferential
statistics. The standard deviation is not a good measure of spread in highly-skewed
distributions and should be supplemented in those cases by the semi-interqurtile range.
Measures of Shape (Skewness)
So far we discussed mean and standard deviation, which describe the distribution. But
there are other parameters which describe how symmetrical the distribution is about the
mean, how peaky is the distribution, or the shape of the distribution.
52
Kotebe University of Education, Department of Geography and
Environmental studies
56.5-58.5 5 3 0 4
58.5-60.5 25 5 4 8
60.5- 15 20 40 20
62.5 10 44 24 24
64.5 15 20 20 40
66.5 25 5 8 4
68.5-70.5 5 3 4 0
N 100 100 100 100
Mean 63.5 63.5 63.5 63.5
Median 63.5 63.5 63 64
Mode bimo 63.5 61.9 65.1
The shape of the curves and histograms which placed equal items at equal distance on
either side of the median clearly shows that distributions A and B are symmetrical. If we
fold this curves, or histograms on the ordinate at the mean, the two halves of the curve or
histograms will coincide. In distribution B all the three measures are identical and in A
which is a bimodal distribution mean and median are equal.
Distribution C and D are asymmetrical. This is evident from the shape of the figures and
from the fact that (i) items at equal distances from the median are not equal in number,
and (ii) the three measures of central tendency are not equal. There is also a difference in
asymmetry distributions. When mean is greater than the median and mode, the curve is
pulled more to the right (skewed to the right) and when mean is less than median and
mode =skewed to the left. Put differently, if extreme variations in a given distribution are
towards higher values of observation, they give the curve a longer tail to the right
(skewed to the right) and this pulls the median and mean in that direction from the mode.
If, however, extreme variations are towards lower values of observations, the longer tail
is to the left and the mean and median are pulled to the left of the mode.
It could also be shown that in a symmetrical distribution, the lower and upper quartiles
are equidistant from the median, so also are corresponding pairs of deciles and
percentiles.
53
Kotebe University of Education, Department of Geography and
Environmental studies
So we can have the following as tests for the presence of skewness.
(i) graph,
(ii) when three measures of central tendency are not equal in their value,
(iii) when the sum of the positive deviations from the median are not equal to the
negative deviations from the median,
(iv) when distances from the median to the quartiles are not equal, and
(v) when corresponding pairs of deciles and percentiles are not equi-distant from the
median.
Skewness is the degree of asymmetry of a distribution. If the distribution has a longer tail
less than the maximum, the function has negative skewness. Otherwise, it has positive
skewness.
Measures of Skewness:
1. Karl Pearson’s measures of skewness (Relationship between the three measures of
central tendency.
We said that =Me=Mo for symmetrical distribution. As the distribution departs from
symmetry these values are pulled apart, the difference between and Mo being the
greatest. Karl Pearson has suggested the use of this difference in measuring skewness.
Absolute skewness = -Mo with units measurement.
And plus or minus signs obtained in this formula indicates the direction of the skewness.
If it is positive the extreme variation in the given distribution are towards higher values
(skewed to the right).
Coefficient of skewness: The difference between mean and mode is an absolute measure
of skewness. An absolute measure cannot be used for making valid comparison between
the skewness in two or more distributions for the following reasons: (i) the same size of
skewness has different significance in distributions with small variation and in
distributions with large variation, in the two series, and (ii) the unit of measurement in the
two series may be different. To make this measure suitable to compare skewness it is
necessary to eliminate from it the disturbing influence of variation and units of
54
Kotebe University of Education, Department of Geography and
Environmental studies
measurement. Such elimination is accomplished by dividing the difference between mean
and mode by the standard deviation. The resultant coefficient is called Pearsonian
coefficient of skewness.
Coefficient of skewness =
Coefficient of skewness =
In other words the value of the median is the mean of Q1 and Q3.
In a skewed distributions, quartiles would not be equidistant from the median unless the
entire asymmetry is located at the extremes of the series. Bowley has suggested the
following formula for measuring skewness, based on above facts.
Absolute SK=(Q3-Me)-(Me-Q1)
=Q3+Q1-2Me
If the quartiles are equidistant from the median
SK = 0
If Me-Q1>Q3-Me = negative skewness
And Me-Q1<Q3-Me = positive skewness
55
Kotebe University of Education, Department of Geography and
Environmental studies
If the series expressed in different units are to be compared, it is essential to convert the
absolute amount into the relative. Using the interquartile range as denominator, the
formula for the coefficient of skewness is as follows:
Relative SK=
Or =
If in the series the median and lower quartiles coincide, then the SK becomes (+1). If the
median and the upper quartiles coincide, then the SK becomes (-1).
Merits and demerits of this measurement: (i) easily computable, (ii) it has value limits
between (+1) and (-1). The disadvantages of this measure is it does not consider all items
in the distribution. It neglects extreme items.
Eg: Calculate the skewness for the above table.
SK=P50-
=D5-
Skewness will take on a value of zero when the distribution is a symmetrical curve. A
positive value indicates the observations are clustered more to the left of the mean with
most of the extreme values to the right of the mean. A negative skewness indicates
clustering to the right.
Significant Level
56
Kotebe University of Education, Department of Geography and
Environmental studies
The amount of evidence required to accept that an event is unlikely to have arisen
by chance is known as the significance level or critical p-value. Significance
levels show you how likely a result is due to chance. In the field of Education, the
most common level, used to mean something is good enough to be believed, is .95.
This means that the finding has a 95% chance of being true. However, this value is
also used in a misleading way. No statistical package will show you "95%" or
".95" to indicate this level. Instead it will show you ".05," meaning that the finding
has a five percent (.05) chance of not being true, which is the converse of a 95%
chance of being true. To find the significance level, subtract the number shown
from one. For example, a value of ".01" means that there is a 99% (1-.01=.99)
chance of it being true.
Hypothesis
The hypotheses are often statements about population parameters like expected value and
variance. A hypothesis might also be a statement about the distributional form of a
characteristic of interest. They are simply informed of intelligent guess about the solution
to a problem. To understand more about hypothesis it will be necessary to consider the
different types.
Types of Hypothesis
57
Kotebe University of Education, Department of Geography and
Environmental studies
H = There is no significant difference in the mean scores in science of students
0
taught using inquiry method and those taught using no inquiry method
H1 or Ha = There is a significant difference in the mean score in science of students taught
using inquiry and non-inquiry methods.
Alternative hypothesis can further be described as Non-directional and Directional
Alternatives. When an alternative hypothesis gives the direction of the difference or
effect it is known as a Directional Alternative, but when direction is not given it is non-
directional.
Example:
Directional Alternative: Boys taught using inquiry method has higher mean score
in science than girls taught using the same method of teaching is Directional
Alternative.
Non-Directional Alternative: The mean score in science for boys taught
Steps in Hypothesis Testing
Hypothesis testing is the use of statistics to determine the probability that a given
hypothesis is true. The usual process of hypothesis testing consists of four steps:
1. Formulate the null hypothesis and the alternative hypothesis.
2. Identify a test statistic that can be used to assess the truth of the null hypothesis.
3. Compute the P-value, which is the probability that a test statistic at least as significant
as the one observed would be obtained assuming that the null hypothesis were true (the
smaller the p-value, the stronger the evidence against the null hypothesis).
4. Compare the -value to an acceptable significance value (sometimes called an alpha
value). If P< α, that the observed effect is statistically significant, the null hypothesis is
ruled out, and the alternative hypothesis is valid.
Two-tailed and one-tailed Tests
One important concept in significance testing is whether you use a one-tailed or two-
tailed test of significance. The answer is that it depends on your hypothesis. When your
hypothesis states the direction of the difference or relationship, you use a one-tailed
probability. For example, a one-tailed test would be used to test these null hypotheses:
Females will not score significantly higher than males on an IQ test. Blue collar workers
58
Kotebe University of Education, Department of Geography and
Environmental studies
are will not buy significantly more product than white collar workers. Superman is not
significantly stronger than the average person. In each case, the null hypothesis
(indirectly) predicts the direction of the difference. A two-tailed test would be used to test
these null hypotheses: There will be no significant difference in IQ scores between males
and females. There will be no significant difference in the amount of product purchased
between blue collar and white collar workers. There is no significant difference in
strength between Superman and the average person. The one-tailed probability is exactly
half the value of the two-tailed probability.
There is a raging controversy (for about the last hundred years) on whether or not it is
ever appropriate to use a one-tailed test. The rationale is that if you already know the
direction of the difference, why bother doing any statistical tests. While it is generally
safest to use two-tailed tests, there are situations where a one-tailed test seems more
appropriate. The bottom line is that it is the choice of the researcher whether to use one-
tailed or two-tailed research questions.
59