Quantitative Techniques in Geography Manual
Quantitative Techniques in Geography Manual
Department : Geography
Saurbh
DEPARTMENT OF GEOGRAPHY
SCHOOL OF ENVIRONMENT AND EARTH SCIENCES
CENTRAL UNIVERSITY OF PUNJAB, BATHINDA
Geo 525 Quantitative Techniques in Geography
Table of Contents
Contents
Exercise 1: Lorenz Curve..................................................................................................................2
Exercise 2: Ginni’s Coefficient...........................................................................................................5
Exercise 3: Location Quotients..........................................................................................................8
Properties ......................................................................................................................................8
Limitations of Location Quotient......................................................................................................9
Exercise 4: Nearest Neighbour Analysis.......................................................................................... 10
Exercise 5: Transportation Network Analysis & Connectivity Metrix .............................................. 13
Exercise 6:- Simple Linear Regression:- .......................................................................................... 23
Exercise 7 Multiply regression:- ...................................................................................................... 26
Exercise 8:- Gravity model............................................................................................................... 29
Exercise 9 Theissen polygon: - ........................................................................................................ 32
Introduction
Quantitative methods are concerned with objective measurements and statistica l, mathematical,
or numerical analysis of data collected through polls, questionnaires, and surveys, as well as the
modification of pre-existing statistical data using computational tools. Quantitative research is
gathering numerical data and interpreting it across groups of people or understanding a
phenomenon. The goal of a quantitative research study is to establish a relationship within a
population between one item [an independent variable] and another [a dependent or outcome
variable]. Quantitative research designs can be descriptive [subjects are usually only measured
once] or experimental [subjects are measured before and after a treatment].
DATE:
percentage distribution are same, it means there has no inequality. On the other hand, if both are
no same, it means the distribution is not equal.
Graphically, cumulative distribution of the population are plotted on the horizontal axis and
cumulative distribution of the income or wealth are plotted on the vertical axis. All these point
will be pointed on the diagonal line of the square. This diagonal line are to present as line of
equal distribution. When we plot all the point according to data, the shape of the curve are
formed which is called ‘Lorenz Curve’.
Some other measures such as Coefficient of variation, Location Quotient, Sopher’s Index etc. of
inequality that are used in place of Lorenz Curve to show the level of inequality in any
distribution.
Method/formula:-
Solution:-
Cumulative of Cumulative of C % of C%
Region income('000) Population income Population population Income
0 0 0 0
A 10 14 10 14 28 5
B 20 12 30 26 52 15
C 40 10 70 36 72 35
D 50 8 120 44 88 60
E 80 6 200 50 100 100
Geo 525 Quantitative Techniques in Geography
Result:-
Lorenz Curve
Lorenz Equality
100
90
80
70
60
C % Income
50
40
30
20
10
0
0 10 20 30 40 50 60 70 80 90 100
C % Population
Interpretation:-
This Lorenz curve represents income distribution.
In this data, two percentage are not same, so we do not get a straight line, in spite of
we get a curve below the line of equal distribution.
This Lorenz Curve given above shows deficient income inequality in the population
of ‘X’ place.
Curve shows the inequality in the distribution of income in relation to the distribution
of population of the place.
We compared the income distribution of ‘x’ place with a straight line which
representing perfect equality.
At the 28th percentile of population, the value of the Lorenz curve is 5. In other
words, this Lorenz curve estimates that 28% of the population takes 5% of the total
income. If ‘X’ place have a perfect equality, 28% population would get 28% income.
DATE:
Derive the Ginni’s Coefficient using the data given in table 3 (use % area
and % population figures).
Data:-
Ta bl e 2, AREA, AND DENSITY OF POPULATION IN DISTRICTS OF KERALA 2011
Population
District % Area (sq km) % Population (2011)
density (2011)
Here X1 and Y2 are the cumulative percentage of population and area. And 100*100 is the
length and breath. In the bracket, the value is to be taken without plus or minus sign.
Solution (Insert Excel table and figure):-
After plotting the Lorenz Curve which showing the inequality in the population distribution
over the area of the district of Kerala, by using the column 4 and column 5.
Geo 525 Quantitative Techniques in Geography
% Cumulati Cumulative
Population % Area
District Populatio ve % of % of A B
density (2011) (sq km)
n (2011) area Population
X Y X1 Y1 Y1*X1k+1 X1*Y1k+1
1 Idukki 255 11.21 3.32 11.21 3.32 55.4108 64.6817
2 Wayanad 383 5.48 2.45 16.69 5.77 135.6527 156.0515
3 Pathanamthitta 451 6.82 3.58 23.51 9.35 327.7175 417.5376
4 Palakkad 626 11.54 8.41 35.05 17.76 713.4192 759.5335
5 Kasaragod 657 5.12 3.91 40.17 21.67 1035.6093 1173.767
6 Kannur 852 7.62 7.55 47.79 29.22 1562.3934 1678.863
7 Kottayam 895 5.68 5.91 53.47 35.13 2152.0638 2377.811
8 Thrissur 1031 7.79 9.34 61.26 44.47 3008.3955 3207.574
9 Kollam 1061 6.39 7.89 67.65 52.36 3954.7508 4207.154
10 Ernakulam 1071 7.88 9.83 75.53 62.19 5266.2492 5626.985
11 Malappuram 1157 9.15 12.31 84.68 74.5 6758.64 7091.103
12 Kozhikode 1316 6.04 9.24 90.72 83.74 7901.7064 8174.779
13 Alappuzha 1503 3.64 6.37 94.36 90.11 9011 9436
14 Thiruvananthapuram 1508 5.64 9.89 100 100 0 0
Total 100 100 41883.009 44371.84
B-A= 2488.831
Ginni's concentration ratio G= 0.248883
80 74.5
Cumul ative % of Population
62.19
60 52.36
44.47
40 35.13
29.22
21.67
17.76
20
10 9.35
5.77
3.32
0
0 20 40 60 80 100
Cumul ative % of a rea
Lorenze Equality
Geo 525 Quantitative Techniques in Geography
We can also measure the inequality through Ginni’s Concentration ratio ‘G’.
Column 5 and 6 of the table 3, give the cumulative percentage of population X1 and Y1.
Column 7, give the values of Y1*X1+1, By multiplying each value of Y with the value of
following Y that is Y1 with X2, Y2 with X3 and so on. In the next column 8, give the values of
X1*Y1+1, in this value we do reserve of it that is multiply X1 with Y2, X2 with Y3 and so on.
Total of column 7 and 8 of the table 3 will give the values of ⅀ (X1Y1+1) and ⅀ (Y1X2+1).
The value of the ‘G’ can be calculated as:-
1 1
G = 100∗100 (44371.84-41883.009) =100∗100 *2488.831 =0.248883Ans.
Result:-
Here the value of “G” is 0.248883 which indicates less amount of inequality in the distribution of
population over area.
Interpretation:-
In this example, Y-axis denotes ‘Cumulative % of population’ and X-axis denotes
‘cumulative % of area.
In this example, straight line shows here line of equality, it means there is zero
inequality. And curve shows here inequality. Gini index which show the value of
inequality between Zero to 1. In this example, Gini Index is close to zero it means less
amount of inequality or greater equality. If everyone had exactly the same amount of
money, Gini Index would register a reading of 0. But low numbers are not always a
perfect indicator of economic health. If the no. can be multiplied by 100 in order to
express it as a percentage.
Date:-
Location quotient is useful in demographic studies because it shows what makes the region’s
demographics unique in comparison to its state and/or the nation.
Properties:
If LQ is less than 1, then it indicates opportunity to develop business/industry in the local
area as there is less competing area
If it is greater than 1, then the area has less opportunity to develop new project as more
industries or other economic sector is concentrated.
If LQ = 1, the state of nation has a share of the total in accordance with its share of
Geo 525 Quantitative Techniques in Geography
Derive the location quotients value by using the data which give on table
4.
Data:-
Table 4, Native speaking population and total population in different regions of the X
Method/Formula:-
𝑆/⅀(𝑆)
Location Quotients =
𝑇/⅀(𝑇)
Solution:-
Table 5, Native speaking population and total population in different regions of the X
Interpretation:
A Location Quotient less than one indicates that the percentage of native speaking in the
sub region is lower than the national average.
From above result:
D sub- region has the LQ value greater than 1 means more % of native language
speaker than entire region.
Location Quotient equal to one indicates that equal % of native language speaker equal to
the % of native language speaker in entire region.
From above result:
Region E have equal % of native language speaker
Date:-
Set of point of an area can be arranged in a number of ways but three basic patterns are
recognized:-(1) Regular or uniform: - That pattern will occur when interval between points is
similar.
Geo 525 Quantitative Techniques in Geography
Method/Formula:-
Rn=2d*√ (n/A)
Solution:-
Select the point or settlements on a map which we want to analyse, for example,
figure X show the distribution of 7 villages located in a part of various district of
Haryana. These point are collected from ‘Google Earth Pro’. Total area of this
polygon is 35 km seq. met.
Geo 525 Quantitative Techniques in Geography
With the help of ‘RULER’ option in Google Earth Pro, connect the all point with
their nearest neighbour village and measure their crow-fly distances to calculate the
average NND for the area under consideration (table-5)
Calculate the mean of the measure distance, by doing- total sum of ground
distance divided by total villages.
Put all the value in NNA formula. Result will come.
Area= 35
d= 1.885714286
n= 7
Rn= 2d*√(n/A)
1.686634132
Result:-
The value of nearest neighbour analysis is 1.686634132.
Interpretation:-
Nearest neighbour analysis value will be between 0-2.15.
The smaller the nearest neighbour analysis value means clustered pattern. Higher the
value means more regular pattern. If value of NNA is ‘0’ it show complete
clustering that means maximum aggregation of all the point at one location.
If value of NNA is 1, that means random distribution. Value 2.5 indicates a regular
pattern.
Geo 525 Quantitative Techniques in Geography
In the above exercise, value of NNA is 1.6866634132, that means a near random
situation or we can say that pattern is uniform.
Date
A B C D E F G
Geo 525 Quantitative Techniques in Geography
A 0 1 0 0 0 1 1
B 1 0 1 1 0 1 0
C 0 1 0 0 1 0 1
D 0 1 0 0 1 1 0
E 0 0 1 1 0 0 0
F 1 1 0 1 0 0 0
G 1 0 1 0 0 0 0
The least accessible node is ‘G’ only two direct connection is there. The most accessible node is
‘B’ where four direct connection is there.
Data:-
Case 1:-
Method/Formula:-
α= actual circuit/ maximum circuits
Or
α= e-ν+1/ 2ν-5
e= edge, v= vertex
Geo 525 Quantitative Techniques in Geography
Solution:-
E= 10
V= 7
10−7+1
= ;
2∗7−5
4
= ;
9
= 0.44 Ans.
Result:- the value of alpha index is 0.44 Ans.
INTERPRETATION:
For Alpha Index:
The higher the alpha index, the more a network is connected. It means network A (0.44) has
high network connectivity as simple networks will have a value of 0.
Case-2
Method/Formula:-
α= actual circuit/ maximum circuits
Or
α= e-ν+1/ 2ν-5
e= edge, v= vertex
Solution:-
E=4
V= 6
4-6+1/2*6-5
= -0.14 ans.
Data:-
Case -1
Method/Formula:-
β = arcs/ nodes
Arcs: - edges
Nodes: - vertex
Solution:-
V=4
E=3
=3/4
=0.75.
Method/Formula:-
β = arcs/ nodes
Solution:-
V= 5
E= 8
β = E/V
=8/5
= 1.6 Ans
Result:- beta value of data is 1.6 ans.
INTERPRETATION:
For Beta Index:
Here, our Beta value of data is 1.6 and hence this Network has complex network system as it
values is more than [Link] means this network system low degree of connectivity.
Data:-
Geo 525 Quantitative Techniques in Geography
Method/Formula:-
ү=e/3(v – 2)
e= edges
v= vertex
Solution:-
E= 3
V= 4
= 3/3(4-2)
=0.50 ans.
Method/Formula:-
ү=e/3(v – 2)
e= edges
v= vertex
Solution:-
e= 4
v= 4
Y= 4/3(4-2)
= .66 ans.
Result:- the value of gamma is .66 ans.
Geo 525 Quantitative Techniques in Geography
Case-3
Method/Formula:-
ү=e/3(v – 2)
e= edges
v= vertex
Solution:-
E= 8
V= 5
8
Ү=3(5−2)
=8/9
=0.88 Ans. Or
If we want to beta value in percent then multiply by 100
0.88*100= 88.8% Ans.
Data;-
Case 1.
Geo 525 Quantitative Techniques in Geography
Method/Formula: - when two or more sub-region, then the formula for Cyclomatic number will
be:-
=a–n+x
a= number of arcs (edges)
n= number of nodes
x= number of subgraphs
Solution:-
a=4
n=6
= 4-6+2
=0 ans.
Result:- Cyclomatic number is 0.
Case-2
Method/Formula:-
When sub-graphs are not in data then this formula will be used for finding the Cyclomatic
Number.
Cyclomatic number = a – (n – 1)
Or
a –n+1
Geo 525 Quantitative Techniques in Geography
Solution:-
a=10
n=7
= 10-7+1
=4 Ans.
Result:- Cyclomatic number is 4.
Detours: - For shortest distance, travellers use the straight route between two
places that type of route also known ‘desire line’. In the other words, the
detour index is the actual distance calculated as a percentage of the desire line
distance. In the detours index, mostly of the time actual distance is always
longer than the desire line distance. The value of detours index can never be
less than 100. If the value of detours index is lower, it means more direct is a
given route.
Data:-
Method/Formula:-
Solution:-
Detour Index:-
Actual route distance=22
Straight line distance=10
22
= ∗ 100
10
=220 Ans.
Interpretation:-
In fact, the actual route distance is almost always longer than the desire line distance, then the
detour index will be greater, in almost all cases, than 100 and in the nature of things can never be
less than 100. It is obvious that lower the detour index, the more direct is a given route. The
detour index is used for assessing the effects which the addition or abstraction of links produce in
a given network.
Derive the simple linear regression value by using the data which give on
table 5.
Table 9, Ads cost and their sales
Solution:-
Summary output
Regression Statistics
Multiple R 0.985401606
R Square 0.971016325
Adjusted R Square 0.967393366
Standard Error 2.180543935
Observations 10
ANOVA
df SS MS F Significance F
Regression 1 1274.362 1274.362 268.0174 1.95241E-07
Residual 8 38.03817 4.754772
Total 9 1312.4
Residual Output:-
45
40 y = 0.5225x - 5.6521
35 R² = 1
Sales (Y)
30 Sales(Y)
25 Predicted Y
20
Linear (Predicted Y)
15
10
40 50 60 70 80 90 100 110
Interpretation:-
Multiple R: - this is the correlation coefficient that tell us how strong the linear
relationship is between two variables. The range of value can be any value between -1
and 1. If value is 1 it means a perfect positive relationship.-1 value show a strong
negative relationship 0 value means no relationship at all. In the above example,
Geo 525 Quantitative Techniques in Geography
For example, we might use multiple regression to see if exam success can be predicted based on
revision time, test anxiety, lecture attendance, and gender. You may also use multiple regression
to see if daily cigarette use can be predicted based on smoking duration, age at which you first
started smoking, smoker type, income, and gender.
We can also use multiple regression to assess the model's overall fit (variance explained) and the
proportional contribution of each predictor to the total variance explained. You could, for
example, wish to know how much variance in exam performance can be explained by revision
Geo 525 Quantitative Techniques in Geography
time, test anxiety, lecture attendance, and gender "as a whole," as well as the "relative
contribution" of each factor.
Data:-
Y Sex Ratio (Female per thousand male) (Y).
X1 Growth rate of population, 2001-11( in Percentage) (X1)
X2 Levels of Literacy ( in percentage) (X2)
X3 Population Density per square kilo meter of area (X3)
X4 Female work participation rate ( in percentage) (X4)
S.N. District Y X1 X2 X3 X4
1 Indore 928 32.88 80.87 841 20.9
2 Jabalpur 929 14.51 81.07 473 25.3
3 Sagar 893 17.63 76.46 232 28.9
4 Bhopal 918 28.62 80.37 855 19.6
5 Rewa 931 19.86 71.62 375 32.9
6 Satna 926 19.19 72.26 297 29.9
7 Dhar 964 25.6 59 268 40.2
8 Chhindwara 964 13.07 71.16 177 36.6
9 Gwalior 864 24.5 76.65 446 14.5
10 Ujjain 955 16.12 72.34 326 33.8
11 Morena 840 23.44 71.03 394 16.8
12 West Nimar 965 22.85 62.7 233 40.9
13 Chhattarpur 883 19.51 63.74 203 32.7
14 Shivpuri 877 22.76 62.55 171 34.5
15 Bhind 837 19.21 75.26 382 8.4
16 Balaghat 1021 13.6 77.09 184 47
17 Betul 971 12.92 68.9 157 42.9
18 Dewas 942 19.53 69.35 223 38.4
19 Rajgarh 956 23.26 61.21 251 41.8
20 Shajapur 938 17.2 69.09 244 39.1
21 Vidisha 896 20.09 70.53 286 21.6
22 Ratlam 971 19.72 66.78 255 39
23 Tikamgarh 901 20.13 61.43 157 36.9
24 Barwani 982 27.57 49.08 242 41.9
25 Seoni 982 18.22 72.12 157 42.4
26 Mandsaur 963 13.24 71.78 199 42.9
27 Raisen 901 18.35 72.98 178 23.4
28 Sehore 918 21.54 70.06 261 28.7
29 East Nimar 943 21.5 66.39 178 38.6
30 Katni 952 21.41 71.98 173 31
Geo 525 Quantitative Techniques in Geography
Solution:-
Interpretation
Geo 525 Quantitative Techniques in Geography
The intercept value and the equation for each variable in this regression will be the same, but the
slope of each independent variable with the dependable variable will change. Now we'll see how
well the four independent factors can predict the dependent variable under consideration. Among
the four independent variables given above, we must choose the best predictor (independent).
Multiple R: - this is the correlation coefficient that tell us how strong the linear
relationship is between two variables. The range of value can be any value
between -1 and 1. If value is 1 it means a perfect positive relationship.-1 value
show a strong negative relationship 0 value means no relationship at all. In the
above example, correlation coefficient value is 0.866906 that is near to 1. It
means there is positive relationship between two variables.
R square: - this value show us how good your model is. The range of this value
is between0 to 1. It can be convert into percentage. Zero value of R square mean a
terrible model and if value of R square is 1 it means a perfect model. In the above
example, the value of R square is 0.751526 that is showing extremely well model.
So we are confident in our prediction. This value has given in percentage.
Standard error: - The standard error is an absolute metric that displays how
much the data points deviate from the regression line on average. And it is the
measure that tell us accurate mean of any given sample. Larger value of standard
error means inaccurate representation and vice-versa. In the above example.
Standard error value is 21.9373 that is smaller number, it means accurate
representation.
Observations: - it show the total number of the observations in our model. In
given example, observation value is 50.
Significance F: - the value of Significance F indicates how reliable our results
are. Our model is acceptable if Significance F is less than 0.05 (5 percent). If it's
larger than 0.05, we have to choose a different independent variable. In the above
example, standard value is 4.43E-13. When we convert this value then the value
of Significance f is 0.0195241, it means this value is less than 0.05 which means
our results is reliable.
P- VALUE: - If the null hypothesis of your statistical test was true, the p-value informs
you how often you'd expect to see a test statistic as severe as or more extreme than the
one generated by your statistical test. In above example, we predicted some value, for
different independent data.
(1) The total population of the places, and (2) the distance between them, or the time or expense
of travelling that distance.
There will be a positive relationship between flow volume and population size, when two place
have large populations, we expect that a large volume of migrants, but if places are separated by
a large distance, we expect a small volume of migrants.
Newton's law is used to estimate and calculate the relationship between objects. The gravity
model uses this same idea to predict the relationship between places. Instead of gravitational
pull, however, we're interested in the degree of interaction between cities, towns, or regions.
Newton's law of gravity predicts that bodies which are larger and closer will exert more force.
Our main variables are size and distance. In the gravity model of human geography, we can use
these same variables. Size is measured in population, and distance can be measured using any
metric. The idea in the gravity model is the same as in Newton's law. The larger and closer two
places are, the more influence they'll have on each other
Data:-
Population
Sl. No. Urban Centres (‘000)
1 Duliajan 17.017
2 Naharkatia 15.052
3 Namrup 19.74
4 Dibrugarh 120.127
Method/formula:-
Tij = Pi*Pj / (dij )2
Pi= Represent the population size of origin place
Solution:-
Geo 525 Quantitative Techniques in Geography
According to rule: - multiply the two place’s population and divided by the square of the distance
of the two place. For example, if we apply the gravity model on the Dibrugarh and Duliajan, then
multiply 120.127*17.017 and divided by the 45*45, answer will be 1.009482.
Gravity Model
Duliajan and Dibrugarh 120.12*17.017/45*45 1.009482
Duliajan and Naharkatia 17.017*15.052/14.8*14.8 1.16937493
Duliajan and Namrup 17.017*19.74/43*43 0.18167419
Namrup and Naharkatia 19.74*15.052/29.6*29.6 0.33912354
Namrup and Dibrugarh 19.74*120.127/81*81 0.36142463
Result: - There are the values of gravity model of different place:- 1.009482, 1.16937493,
0.18167419, 0.33912354, 0.36142463, 0.67386876
Interpretation:-
Geo 525 Quantitative Techniques in Geography
In the above example we can shows that the cities are separated by a broad variety of distances,
and that the spatial interaction between two places is influenced not just by distance but also by
population size.
We predict a significant volume of migrants or commuters when two regions have large
populations, but we expect the impact of distance, as a mediator, when places are separated by a
considerable distance. In the other words, we can see here population size leads to a positive
relationship, distance leads an inverse correlation.
Based on an area containing at least two points, a Thiessen Polygon is a 2-dimensional shape
whose boundaries contain all space which is closer to a point within the area than any other point
without the area. Such polygons are a critical component of such spatial analysis concepts as
nearest neighbor and proximity assessment. Named after American meteorologist Alfred H.
Thiessen, Thiessen polygons are a more specific application of Voronoi diagram to meteorology
and geophysics
Thiessen polygons are generated from a set of sample points such that each polygon defines an
area of influence around its sample point, so that any location inside the polygon is closer to that
point than any of the other sample points.
Data:-
Geo 525 Quantitative Techniques in Geography
Geo 525 Quantitative Techniques in Geography
Method/formula:-
𝑃1𝐴1+𝑃2𝐴2+⋯+𝑃𝑛𝐴𝑛 ∑𝑛
𝑖=1 𝑃𝑖𝐴𝑖 1
P= = ∑𝑛 𝐴𝑖
= ∑𝑁
𝑖=1 𝐴𝑖𝑃𝑖
𝐴1+𝐴2+⋯+𝐴𝑛 𝑖=1 𝐴
Pi= precipitation in the catchment area
Solution:-
𝟖.𝟖∗𝟓𝟕𝟎+𝟕.𝟔∗𝟗𝟐𝟎+𝟏𝟎.𝟖∗𝟕𝟐𝟎+𝟗.𝟐∗𝟔𝟐𝟎+𝟏𝟑.𝟖∗𝟓𝟐𝟎+𝟏𝟎.𝟒∗𝟓𝟓𝟎+𝟖.𝟓∗𝟒𝟎𝟎+𝟏𝟎.𝟓∗𝟔𝟓𝟎+𝟏𝟏.𝟐∗𝟓𝟎𝟎+𝟗.𝟓∗𝟑𝟓𝟎+𝟕.𝟖∗𝟓𝟐𝟎+𝟓.𝟐∗𝟐𝟓𝟎+𝟓.𝟔∗𝟑𝟓𝟎
𝟓𝟕𝟎+𝟗𝟐𝟎+𝟕𝟐𝟎+𝟔𝟐𝟎+𝟓𝟐𝟎+𝟓𝟓𝟎+𝟒𝟎𝟎+𝟔𝟓𝟎+𝟓𝟎𝟎+𝟑𝟓𝟎+𝟓𝟐𝟎+𝟐𝟓𝟎+𝟑𝟓𝟎
64,850
= 6920
=9.37138728cm Ans.
Interpretation:-
As we all know, the Thiessen polygon is a technique for forecasting average rainfall depths in
watersheds using a rain-gauge network that is particularly well adapted to electronic
computation. In the example data table, the measured rainfall at 13 locations is already provided
(table number). Rainfall variability was calculated using the regions, which was then multiplied
by sub-area rainfall and added. The average annual precipitation is 9.37138728cm, according to
the facts and estimates. It is also clear that the annual precipitation at various places with
recorded mean rainfall is significantly less than the annual total mean rainfall.