The World’s 100 Largest Cities:
A case study of inequality
KEYWORDS: David Drew and Dave Steyne
Teaching; Sheffield Hallam University, UK.
Quality of life;
Poverty; Summary
Third world. Data are presented here for the hundred largest cities in
the World. They form part of a case study to teach stu-
dents about exploratory data analysis but are of added
interest in providing a focus on poverty and underde-
velopment in the Third World and the contrast between
this and the wealth of the First World.
◆INTRODUCTION◆ ◆THE DATASET◆
ONE thing we do not expect from our students is an The data set comprises of ten indicators for the 100 larg-
automatic interest in data analysis or statistics as a est cities of the world taken from a study by a Washing-
subject. In England, experience has shown us that ton DC based environmental organisation (see Camp,
only a small minority of students of the social sci- Barberis and Hinds, 1990, and the Guardian, 1990). The
ences come with a developed interest in statistics. list of areas was prepared by a metropolitan area expert
Their interest has to be won. These students are un- from Rand McNally and Company and the data relates
likely to be strong in mathematics and they will prob- to 1989 for all variables.
ably find the more mathematical parts of the subject The statistics were collected for each metropolitan area
difficult. One way out of this problem is to link the by means of a 13-page in-country questionnaire and there
teaching of statistics and data analysis closely to is an extensive list of acknowledgements in the original
sociology itself, to play down the role of mathemat- report to the planners and statisticians who compiled
ics and to get students involved in practical data the information for each city. The aim of the survey was
analysis. This is the approach used in the excellent to collect data ‘concerning living standards relevant
book by Marsh, (1988) and this approach seems to across national boundaries without cultural bias’ (Camp,
work for our students. Barberis and Hinds, 1990). The two urban areas for
The case study here enables students to practise their which data was not obtainable are Yangon, Myanmar
knowledge of Minitab and to produce boxplots, ta- (Rangoon, Burma) and Bucharest (Romania). A total of
bles, descriptive statistics and a correlation matrix. 162 questionnaires were completed and returned, many
At a statistical level, it enables them to gain experi- with multiple respondents, and often representing offi-
ence in discussing data sources, investigating reli- cial sources.
ability and validity, producing tables and graphics, Data was collected on a number of groups of variables.
integrating these into a report and discussing results. These were:
It allows us to show the importance of good, well
constructed tables and the need for a sociological POPULATION
interpretation of the data; exploring data is about To obtain the most reliable current demographic data
discovering patterns but this cannot be done without for use in the creation of later indicators, the question-
a sociological framework. The case study can also naire asked for official population figures and age, sex
provide an introduction to multivariate analysis us- and household breakdowns. It also asked for the latest
ing cluster analysis. It has been used successfully in estimates of total population available, whether official
a course on Research Methods for undergraduate or not.
social science students.
PUBLIC SAFETY
As a measure of the level of personal security and the
degree of violent crime found in the metropolitan our own designation of area into three groups; The First
area, the questionnaire asked for the annual number World, China and the Third World. The inclusion of
of homicides. China as a separate group was simply because we felt
that Chinese cities might show interestingly different
FOOD COSTS characteristics to those of the other two groups. From
As a general indicator of wealth and poverty, the these groups of variables, the variables given in Table 1
questionnaire asked for the percentage of household were selected.
income spent on food.
Table 1. Variables selected for analysis.
LIVING SPACE AND HOUSING STAND- 1. Population (in millions).
2. Public safety: Murders per 100,000 people.
ARDS
3. Food cost: Percentage of income spent on food.
To determine the level of crowding and provide an
4. Living space: Persons per room.
indicator of the degree of substandard housing, re-
5. Housing standards: Percentage of houses with
spondents were asked questions about the number
water/electricity.
of housing units, rooms per housing unit, and, in sepa-
6. Communications: Telephones per 100 people.
rately ranked variables, the existence of electric and
7. Education: Percentage of children in secondary
water connections.
school.
8. Public health: Infant deaths per 1000 live births.
COMMUNICATIONS 9. Peace and quiet: Levels of ambient noise (1 10). -
To provide a measure of city infrastructure, in par- 10. Traffic flow: Miles per hour in rush hour.
ticular modem communications, respondents were 11. Designation: 1 = First World
asked how many working telephones exist. They 2 = China
were also asked to give an estimate of what percent- 3 = Rest of the World
age of calls actually make a successful connection. 12. Metropolitan area name.
EDUCATION ◆THE ANALYSIS◆
To collect internationally comparable data on edu-
cational standards, specifically secondary school Many different analyses are possible with data of this
enrolment, the questionnaire asked what percentage type. We found it useful to present students with prob-
of children aged 14-17 (or comparable age group) lems of both a closed and open-ended kind, that is we
are in school. asked them to consider particular boxplots and tables,
but we also asked them to analyse the data after this in
PUBLIC HEALTH any way they thought appropriate. This enabled them to
To provide a general indicator of the standard of pub- develop their own hypotheses and test them against the
lic health, respondents were asked the infant mortal- data. This met the objective that weaker students could
ity rate. complete the assignment satisfactorily whilst stronger
students could go further and develop their own ideas.
PEACE AND QUIET At a substantive level the data can be related to the re-
Since there is a lack of specific data on environmen- cent history of colonialism and imperialism in the Third
tal noise pollution, the questionnaire asked respond- World. When Britain, France and the other colonial pow-
ents to give a subjective assessment of the level of ers conceded independence to their colonies they left
ambient noise by ranking their own cities on a 10 behind poorly developed infrastructures in all areas,
point scale. The high, middle and low points of the notably health, education and housing and the economic
scale were defined in non quantitative terms. infrastructure was dependent almost wholly on foreign
owned companies (for an analysis of this see Rodney,
TRAFFIC FLOW 1972 and Miles, 1989). Industry continues to be owned,
In order to develop an internationally comparable in the post-colonial period, by multinational companies
measure of urban traffic congestion, the question- and the newly independent countries developed a reli-
naire asked the distance to the nearest airport and ance on their former colonisers and the banks. They
how long it would take to drive by private car from borrowed to develop the infrastructure they desperately
the airport to the central business district during the needed and debt and further poverty was a result of this.
morning rush hour. The ‘quality of life’ indicators in the case study reflect
As well as these characteristics we ourselves added the effect of this set of conditions. By the late 1980’s the
spiral of debt had continued and in many cases the eco-
where the highest percentages of income are spent on
nomic position of these Third World countries had food then cities in India, Latin America, Africa, South
become even worse than before. East Asia and China are well represented (see Table 4).
The data can be used to discuss issues of reliability This table also suggests that the poorest cities are also
and validity. Considerable detail is given in the the ones where the take-up of places in secondary schools
original data source about the limitations of the is low although there is evidently not a close correlation
data and the ways in which approximations had to between this and the percentage of income spent on food.
be made in order for estimates to be obtained (for If we take the percentage of households with water/elec-
further details see Camp, Barberis and Hinds, tricity as a measure of housing quality and the infant
1990). For example to obtain the murder rate for death rate as a measure of health then we might expect
Johannesburg, data had to be combined for the to find an association between housing conditions and
(largely white) municipality and the (largely black) health. This is indeed the case but the association is not
township of Soweto. Whilst in general, statistical a simple one (see Figure 2). All the First World cities
estimates for socio-economic characteristics will are characterised by relatively low infant death rates and
be prone to a large number of sources of error, the relatively good housing conditions. For the Third World
resources that some governments devote to collect- cities the opposite is frequently the case but there are
ing data will be more than for others and this will exceptions. Johannesburg, for example has poor hous-
be reflected in the quality of the data. ing quality but a relatively low infant death rate. Data is
As far as validity is concerned it is not possible to collected separately there for the municipality of Johan-
tell from the original study what was the exact pur- nesburg and the black townships including Soweto and
pose for which the data was collected. The question the infant death rates in the former are lower than in the
of whether or not the chosen indicators adequately latter (Camp, Barberis and Hinds, 1990). The infant death
reflect the quality of life and what this means could rate in Kanpur, India is extremely high, more than 150
be discussed with students. We have preferred to (i.e. more than 15 percent).
tackle the problem in a different way and that is to
suggest that some of these variables can be used as Table 2. Numerical summary of percentage of income
measures of the particular problems that are experi- spent on food
enced in Third World cities. Poverty, for example, is
measured in an indirect way by the percentage of Percentage of income
household income spent on food, the problems of Depth spent on food
education infrastructure by the percentage of chil- 50.5 35 Median
dren in secondary schooling and the problems of so- 25.5 21 47 26 Midspread
cial conflict by the murder rate. 1 9 80 71 Range
We carried out univariate, bivariate and multivariate
analyses. Taking the percentage of income spent on Note:The format of this numerical summary is that
food as an example it is found that this ranges from 9 given in Marsh (1988).
percent in Washington DC to 80 percent in Ho Chi
Minh City, Vietnam with a median of 35 percent (see
Table 2 and Figure 1). Much higher percentages of
income are spent on food in the Third World and in
China (see Table 3). The boxplots illustrate this and
allow us to see that Katowice in Poland is an outlier
for First World countries. If we consider the cities
Table 3. Medians for selected variables by city has more than one dimension. The standard of living
group may be high in some First World cities in terms of
Group % income % of ch’dren Murders housing and education but these same cities may also
spent on in secondary per 100,000 be characterised by high levels of crime or urban
food school people pollution. A cluster analysis enables us to examine
First World 17 90 3.1 this. Whilst we would not necessarily be expecting
China 54 76 2.5 students to carry out such an analysis at this stage
Rest of World 40 61 6.0 such an analysis could easily be used for illustrative
All Cities 35 76 4.1 purposes here. We carried out a cluster analysis using
SPSS-PC+ and Wards’ method on variables 2-10
given in Table 1. For illustrative purposes we present
Table 4. Selected variables for the ten cities with the cluster analysis for just ten of the cities because it
the highest percentage of income spent on food is interesting to see how these cluster together. The
(ranked). dendrogram for this analysis is shown in Figure 3. If
Group % income % children Murders
we had settled for a four cluster solution the
spent on in secondary per 100,000
food schools people dendrogram shows that Dhaka and Karachi group
Lagos, Nigeria 58 31 * together and this seems intuitively reasonable because
Calcutta, India 60 49 1.1 they are part of the same subcontinent, whilst San
Istanbul, Turkey 60 67 3.5 Francisco, USA forms a cluster on its own. Belo
Guangzhou, China 60 56 2.5 Horizonte, Mexico City and Ahmedabad form a
Bangalore, India 62 60 2.8 separate cluster and one could begin to speculate
Dhaka, Bangladesh 63 37 2.4 about what it is that makes these two cities similar.
Kinshasa, Zaire 63 60 *
It is better still to analyse the full set of data. The
Katowice, Poland 67 87 2.1
Lima, Peru 70 55 *
poorest cluster is characterised by the highest percent-
Ho Chi Minh City, 80 52 2.1 ages of income spent on food, overcrowding in
Vietnam housing and infant deaths which are double the overall
Median All Cities 35 76 4.1 mean. This cluster includes Bangkok, Dhaka and most
of the Indian cities. Many of the characteristics are
A discussion of a bivariate relationship of this kind shared by a further rather interesting cluster although
naturally leads into a discussion of the need for a this cluster is less extreme. This cluster is character-
multivariate analysis. The associations are not ised by a particularly high murder rate, one that is 4.5
simple ones and ‘quality of life’ however defined times the average overall. This cluster includes Cape
Town (South Africa), Manila
(Philippines) and Rio de
Janeiro (Brazil). This suggests
that high levels of urban
crime and social conflict are
an important additional
dimension to the poverty and
decay of some Third World
cities. It may be that large
inequalities in the distribution
of wealth within such cities
contribute to and exaggerate
these problems (Cape Town
may well be an example of
this).
As a teaching vehicle this
data allows a number of points
to be made. Boxplots provide a
succinct and very useful data
summary. Tables need to be
clear and carefully drawn up.
Minitab is an easy and power-
ful package to use, sorting and
ranking data for example is particularly easy to do.
Univariate and bivariate analysis may take us a long ◆NOTES◆
way with the data but multivariate analysis can take
us further. The data set and accompanying documentation can
be obtained free of charge by sending an unformatted
◆CONCLUSION◆ 3.5 inch disk to Dr David Drew, School of Computing
and Management Sciences, Sheffield Hallam Univer-
sity 100 Napier Street, Sheffield, 511 8HD. We would
We feel it is important to present students with chal-
like to thank Elizabeth Coates for drawing our attention
lenging but tractable problems. At this level of teach-
to this data set and Tina Beatty and Paresh Patel for
ing there is a tremendous variability in the back-
originally undertaking the cluster analysis.
ground and experience of the students. Many students
lack confidence but suitable case studies should be References
able to give confidence to the weaker students and Camp, S., Barberis, M., and Hinds, I. (1990). Cities:
stretch the able or more accomplished ones. Life in the World’s 100 Largest Metropolitan Areas.
Computing developments have greatly enhanced Population Crisis Committee. Washington DC.
the potential interest of our courses by removing the Suite 550. 1120 19th Street N.W. Washington DC.
drudgery of computation. Students can see that com- 20036-3605.
puting (in this case Minitab) is interesting, easy and Guardian (1990). Manchester 8th in World quality
fun. We need to meet the challenge to develop case league. Guardian Dec 29th, 1990.
studies which capitalise on this. Marsh, C. (1988). Exploring Data. Cambridge: Polity
We feel it is important that statistics and social sci- Press.
ence problems are seen to be inextricably linked. Only Miles, .R. (1989). Racism. London and New York:
in this way will we get to see social science students Routledge.
becoming interested and involved in data analysis. Rodney, W. (1972). How Europe Underdeveloped
The teaching of data analysis in social science is much Africa. London: Bogle-L ‘Ouverture Publications.
more enjoyable now than it was when we started
nearly twenty years ago. Our graduates need to be
given new competencies in order to come to grips
with the subject.