Guide To Data Visualization
Guide To Data Visualization
visualisation
Legal notice:
1. Introduction 5
3. Methodology 21
3.1. Stage 1 - Strategy 21
3.1.1. Research and analysis 21
3.1.2. Objectives 23
3.1.3. Indicators 25
3.2. Stage 2 - Data 27
3.2.1. Obtaining data 27
3.2.2. Formatting and cleaning 27
3.2.3. Processing 28
3.3. Stage 3 - Design 28
3.3.1. Sketch 28
3.3.2. Create prototype 29
3.3.3. Finalise 30
4. Graphics 31
4.1. Comparisons 32
4.1.1. Bar charts 33
4.1.2. Grouped bar charts 37
4.1.3. Stacked bar charts 38
4.1.4. Radar charts 41
4.1.5. Colour intensity charts 43
4.1.6. Bullet charts 45
4.1.7. Small multiples 47
5. Principles of design 80
5.1. Visual thought 80
5.2. Colour 82
5.3. Shape 84
5.4. Interaction 85
5.5. General recommendations 87
5.5.1. Less is more 87
5.5.2. Balance between functionality and aesthetics 87
5.5.3. Form follows need 88
5.5.4. Text 89
5.5.5. Layout 90
6. Bibliography 91
7. Appendices 92
7.1. Glossary 92
7.2. Tools 93
Data are currently generated regarding virtually any human or natural phenomenon. In areas as
diverse as health, sport, politics, education and marketing, to name just a few, data are used so
that we can better understand reality. More and more people therefore need to develop skills to
work with data and understand their meaning and impact. However, a recent study estimates
that only 17% of European citizens are appropriately educated in the use of data1.
This guide is intended to help everyone to work with data, in particular to explore, analyse and
explain them graphically. It uses the data visualisation framework, a discipline that explains how
to treat information visually. Visualisation offers two great advantages: first of all, visual language
is the best way to make data accessible to non-specialist audiences and, secondly, it makes it
easier to highlight facts in a context of information overload and, therefore, helps to send a clear
message to a specific audience.
However, a poorly designed graphic may lead to an incorrect interpretation of the data, thus
leading to a bad decision. Likewise, a correct graphic that answers an irrelevant question will
not have any impact. That is why this guide addresses both design and methodology, in order to
ensure that the right questions are formulated and that relevant indicators are chosen.
Reference is also made to guidelines for the production of graphics, tables and infographics
that must be in line with the design of other Generalitat de Catalunya communication media.
Corporate identity is a basic element of the institution’s public image and confirms that the data
presented are official and accurate.
The guide is divided into four major blocks: What is data visualisation?, Methodology, Graphics
and Principles of design. Although the guide can be read from beginning to end, some readers
will prefer to use it as a reference work at specific moments. For this reason, there are four info-
graphics below that summarise the four main blocks of content. These infographics work like an
interactive index to parts of the content of the guide that can be consulted when required.
1
Qlik. Data Equality Campaign. <[Link] [electronic resource]. [Accessed: November 2017]
NEWS
DEFINING a strategy
PIB
Població
€€€
A x x
B x x
C x x
D x x
BD c
A 1983 1
0 67% 230€
33,57 ff
12º
Comparisons Trends
To compare different To understand changes in
variables or categories quantitative variables
with each other. over time.
A B
Distribution Correlations
To understand the form To understand the
and properties of a variable. relationship between
different variables.
Connections,
relationships and networks
To understand the
relationship between
the elements of a set
of data.
Shape Interaction
What attributes exist to encode How to use
data through shape? interaction to facilitate
exploration and
understanding.
General recommendations
Other considerations
to produce
TÍTOL
good visualisation
Data visualisation is a very recent discipline and is still being consolidated. Colin Ware was one
of the first researchers to offer a definition of data visualisation that is sufficiently exact and, at
the same time, broad enough not to become outdated in view of the constant technological ad-
vances in the processing and representation of data. According to Ware, visualisation is:
Although for Ware, visualisation is the graphical representation of data or concepts, in this guide
the word visualisation is used to refer only to the visualisation of data.
At a formal level, and in a very general sense, two types of visualisation can be distinguished:
static and interactive.
Static visualisation is normally used to communicate data that have been analysed previously.
This type of visualisation helps to explain patterns and general trends.
In the following image you can see a visualisation of changes in the unemployment rate in the
United States in the form of a graph.
2
Ware, Colin. Information Visualization, Third Edition: Perception for Design. Morgan Kaufmann, 2012
Interactive visualisation enables the user to interact with the data. For example, the interactive
display of The New York Times referred to below allows users to indicate their estimates for a
series of economic and social indicators in the United States:
[Link]
The interaction focuses the user’s attention on the data. In this case, we make the user think
about changes in certain indicators under the presidencies of George Bush and Barack Obama.
A static graph could communicate the same information, but it would certainly not achieve the
same level of reader involvement.
The bidirectional nature of digital media makes it easy for the user to interact with the data. Thus,
instead of telling a story with data, we can make the reader the protagonist. Readers can inter-
rogate the data, so that they help them to respond to particular concerns and answer questions
relevant to their own situation. In these cases, visualisation helps them to explore a set of data.
The visualization of data not only helps users to resolve doubts, but invites
questions they had not even imagined. Exploratory visualisation is a good
example of this: it facilitates interaction with the reader, helps to open a debate
on a topic and answer many questions, while raising new ones.
The visualisation of exploratory data usually requires a multidisciplinary development team, capa-
ble of thinking about a story, collecting and analysing the data and designing interactive inter-
faces. Tools like Tableau facilitate the development of interactive data applications, but the web
format (using libraries such as D3) is the most flexible and efficient, and the one that can best be
adapted to different devices.
In the following example, instead of offering average statistics on the performance of basketball
player Stephen Curry, data for all his throws are shown. The result is an application that allows
you to see the changes over time in the points he scored, in which games he scored most, and
the position of these throws on the court. We can filter by different matches, moments in the
game or types of shot. This visualisation is not intended to answer any specific questions; It
simply allows you to explore a dataset.
30
15
Points 841
Accuracy Attempts
0 2PT 58.7% 305
Oct 24, 15 Nov 8, 15 Nov 23, 15 Dec 8, 15 Dec 23, 15 Jan 7, 16 3PT 44.6% 361
Figure 2.2. Visualisation of throws by player Stephen Curry. Image by permission of OneTandem
Source: [Link]
Apart from telling a story or allowing the user to explore a dataset, visualisation is also useful when
it comes to analysing this data. Data visualisation can thus be used first to analyse a set of data and
extract a series of conclusions, and then to tell a story or allow the data to be explored according
to a series of parameters.
Although interactive visualisation has reached the general public thanks to journalism, its main use
is still the treatment and analysis of complex problems, in fields as diverse as the army, medicine,
biology, public policies and social networks. It is especially useful when the complexity of the pro-
blem to be dealt with or the amount of data exceeds the capacities of traditional visualisation and
analysis techniques.
The following example is an application of visual analytics that allows the user to contextualise
the ranking of different universities, analyse what factors contribute to their position, and what
correlations there are between the different indicators. It is not appropriate for the presentation
of a ranking of universities (for this purpose a bar chart would be sufficient). It would, however,
be appropriate for helping a university to understand why it occupies a specific position and
what it can do to improve it.
3
Cluster analysis is a data analysis technique that can automatically segment a set of data into subgroups (called clusters)
that have values similar to their variables.
2.3.1. Infographics
Infographics became popular in magazines and newspapers as a way of communicating com-
plex themes to the general public. With social networks, infographics have increased their audi-
ence exponentially for three reasons:
• the visual component helps them stand out so that they are shared more often and propaga-
ted quickly via networks;
• they are designed to be consumed in a short time and thus conform to the patterns of rapid
consumption of information associated with social networks;
• they do not usually require prior knowledge of the subject covered and are therefore addres-
sed to a wide audience.
A very good example of infographics is Jaime Serra’s work on the whale:
Figure 2.4. “La ballena franca” (The Right Whale). Infographics by permission of Jaime Serra.
Scrollable narrative is a format designed purely for website pages, usually associated with large
digital media. It can be seen as belonging to the tradition of interactive narrative4. Its creation
calls for many resources because it requires input from contributors with different profiles: scrip-
twriters, designers, programmers, photographers, film-makers, etc.
2.3.3. Presentations
Presentations with tools such as PowerPoint and the like have become a common feature of
corporate communication. As they usually present information concisely, data visualisation plays
an important role.
Generally speaking, it is advisable to incorporate just one idea in each slide. This idea can
be transmitted with a phrase, an image, or the visualisation of data, and must accompany the
spoken presentation naturally. The rule of limiting ourselves to one idea per slide invites us to
introduce the visualisation of data in different steps.
Below we examine the case presented by Cole Nussbaumer, a specialist in storytelling with
data, in his book and blog Storytelling with data5 ([Link]
4
Wikipedia, Interactive Storytelling. <[Link] [Accessed: No-
vember 2017]
5
Nussbaumer, Cole. A Data Visualization Guide for Business Professionals. Wiley, 2015
Figure 2.5. Visualisation used to show the number of active users of a mobile application. Image by permission of Cole
Nussbaumer.
To show the data in a presentation using PowerPoint or a similar application, Cole Nussbaumer
recommends breaking the graph down into a series of slides. First of all, we would begin with
this image to show the initial situation:
Figure 2.6. Image recommended to show the initial situation. Image by permission of Cole Nussbaumer.
Figure 2.7. Series of images used to present and explain changes in the number of active users of a mobile application. Images
by permission of Cole Nussbaumer.
Finally, a slide that summarises the whole explanation can be shown. This will be the slide that
can be included in a PowerPoint (or similar) file sent to the audience by e-mail or printed.
This example shows that the challenge of a presentation is twofold: one needs to know the sub-
ject and know how to explain it. Apart from being good designers and speakers, it is therefore
vital to know the audience. For example, if a presentation is prepared for a class, we need to find
out what pupils want to learn and their level of knowledge about the topic to be addressed. Even
if we design a very complete presentation, it will not work if it does not respond to the audience’s
level and objectives. A common mistake is to use terminology that the audience does not know
or visualisations of excessively complex data.
2.3.4. Video
Video is on its way to becoming the internet format par excellence. The visualisation of data,
so much a part of the internet, is related to this phenomenon. For example, motion charts have
become a common mechanism for transmitting data in a lively way. GIFs also play an important
role in fast consumer information networks such as Twitter. Moreover, many YouTube and digital
media channels opt for video formats that incorporate data and infographics in innovative ways.
A medium that supports this format is Vox ([Link] In its YouTube channel (https://
[Link]/user/voxdotcom) you can see a good sample of videos that make use of the
data visualisation, such as “The real reason American health care is so expensive” [Link]
[Link]/watch?v=k1vE_LVBx4s.
This chapter describes the process of developing data visualisation, in three stages:
• Stage 1: strategy
• Stage 2: data
• Stage 3: design
Once you have defined one or more angles to deal with the topic, it is advisable to consult pu-
blications and experts in the field, as well as other data visualisation done previously, to ensure
that the same work is not repeated.
Finally, it is useful to obtain the testimony of those involved. For example, in the case of pollution,
it would be very enriching to incorporate the opinion of people affected by the problem. Data
helps to quantify a problem, but personal testimonials help to bring it closer and thus increase
the likelihood of causing a greater impact on those seeing the visualisation
Audience
By audience, we understand those who will use the data visualisation. Designing good data
visualisation is similar to designing an application and, therefore, knowing the knowledge and
motivations of the user is vital to ensure its success.
To identify the audience, it is useful to answer these questions:
Relationship
• Do we know them personally?
• Do they know us?
• Is the audience homogeneous or heterogeneous?
In short, it is a question of understanding the mental model of users so that you can design pro-
ducts adapted to their objectives, abilities and the context in which they will be used.
Analysis
Another point to consider when defining strategy is the analysis of the data. When a dataset is
being explored and analysed, the idea of visualisation often arises. It can also happen that the
data show that the initial idea is not valid, or that it is irrelevant.
For the purposes of this first analysis of the data, the following preliminary explorations are ad-
visable:
• Distribution of the main indicators, with histograms and box plots: This will help find atypical
values, which often hide interesting stories.
• Trends in the main indicators: if there are large changes at certain times, this will be evident.
• Correlation between different indicators: the mind tends to establish relationships between
phenomena, so stories that explain relationships work very well.
3.1.2. Objectives
Now that the theme and the audience are known, it is time to define the goals of the data visu-
alisation. Generally speaking, it is said that the objectives are defined if the following questions
can be answered clearly:
• How do we want to help the user?
• Why?
• How will we assess this?
To make this easier to understand, let us imagine, for example, that a visualisation of data regar-
ding pollution in Catalonia is proposed.
In the example of the visualisation of pollution in Catalonia, the question “How do you want to
help the user?” could be answered in three very different ways:
• We want to provide users with data for changes in pollution by county, so that they can explore
how it has varied in the area where they live and work.
• We want to provide users with data on pollution, population density, industry and traffic, plus
mechanisms to establish correlations between different indicators, so that they can analyse
which factors contribute most to the presence of pollution in different areas of Catalonia..
• We want to explain to users how much pollution has been reduced in Catalonia in the last ten
years, and what the main reasons have been.
The three responses can lead to good data visualisation, but it will probably not be possible to
use the same data visualisation for all of them. The same subject can thus be approached from
multiple points of view, for each of which there will be different ways of visualising the data.
Why?
This question is intended to make us reflect on the importance and scope of what we are doing
and, therefore, the resources that should be devoted to it. It also makes us start thinking about
format and channels of communication.
Going back to the case of pollution in Catalonia, we could give the following answers to the
question “Why?”:
In the first case, we want to communicate with the population in general, people who do not
have advanced knowledge of data analysis and interpretation. A series of presentations using
data visualisation in GIF format, accompanied by a campaign in social networks, will surely
be more effective than a sophisticated interactive visualisation. However, in the second case,
it would be better to develop an interactive visualisation with multiple indicators and tools for
analysis that facilitate the interpretation of the data in accordance with the skills of a technician.
3.1.3. Indicators
Data visualisation often fails to achieve the objectives set because it uses data that are not
relevant or do not cover all aspects of the problem. That is why it is vital to define appropriate
indicators. Generally speaking, three types of indicator can be defined:
• Volume
• Quality
• Context
For example, the number of traffic accidents (volume indicator) is not a good indicator for com-
parison between provinces, even if it is presented clearly with a graph like the following:
On the other hand, an indicator such as the accident rate per inhabitant (quality indicator) is
more appropriate, since it is independent of the population of the community and can, therefore,
be compared. In the following chart you can see that Ceuta and Melilla top the classification and
that, for example, the Balearic Islands are above Madrid.
Finally, showing the trend in the accident rate in recent years (context indicator) will help the
audience to interpret the indicators. For example, a 0.2% rate would be positive if it was 0.5%
five years ago, but negative if it was 0.1% at that time.
Apart from using the right indicators, we must bear in mind the level of aggregation of the data.
For example, if you want to analyse the annual trend in traffic accidents, you need to decide if you
want to display data by hours, days or months. Each level of aggregation will allow us to reach
different conclusions and will require a different type of graphic.
The strategy stage thus ends when:
• We have established the topic to be discussed and the target audience.
• We have determined how we want to help the user and why.
• Adequate indicators have been defined (volume, quality, context).
These points should always be borne in mind during the subsequent process of designing the
data visualisation, to ensure that the goals initially identified are achieved.
• General data format: There are numerous formats that express data in a structured way.
Spreadsheets, which contain rows and columns, are a way of organising data that can be
expressed in different formats (XLS, CSV, etc.). Another format used quite often is the JSON
format, although this is more advanced. The data format is chosen based on the data available
and according to the options available to import data into the software used to work on it.
• Format of each of the indicators: It is important to check that all the indicators have a
consistent and, preferably, standard format. For example, if there is a column that expresses
dates, it is important for all of them to be in the same format. For example, day / month / year
(15/10/1981). In this case, the most important thing is consistency, so that all the values of
the indicator are expressed in the same way.
• Sketch
• Create prototype
• Finalise
3.3.1. Sketch
This stage consists of producing drafts with a view to discovering ways to represent the data in
accordance with the strategy defined. It is often referred to as “sketching”.
This stage is particularly useful in defining the appearance of the data visualisation, which must
be consistent with the objectives defined in the strategy. When the first drafts are done, we need
to decide the type of graphics that will be used, which texts should accompany them and what
the layout will be like.
Generally, in this stage it is better to work with pen and paper to have the maximum freedom
when thinking about visualising data. It is advisable to work with notebooks of different sizes, to
have a better appreciation of designs for computer screen, tablet and mobile.
Once the product is finished, a quality check also needs to be carried out, i.e.
• Ensuring that the design works correctly on all the devices where it will be necessary to use
it, with different resolutions and screen sizes.
• Checking each function and making sure it does not give any errors (particularly, filters and
other interactive elements).
• Checking that the data are correct and taking special care to ensure that there are no unex-
pected results for certain combinations of filters or interactive elements.
Although the design has been completed, it has not yet been used in a real environment. New
problems are likely to arise calling for further improvements. It is also likely that the initial require-
ments will change over time. It should thus be borne in mind that a design will never be definitive
and that resources must be reserved to adapt it to different circumstances.
To ensure that the design evolves, it is advisable to:
There are endless possibilities for visualising data, be it for exploring, analysing or explaining
them. However, there is a set of standard graphics, such as bar charts or graphs. In this chapter,
these graphics are classified in 7 different categories, depending on what they can achieve:
• comparisons
• trends
• maps
• parts of a whole
• distributions
• correlations
• connections, relationships and networks
Below we consider the different visualisations in each category. However, it must be borne in
mind that there can always be several valid options for representing a set of data. You need to
decide what you want to represent and know what the data are like to choose the best visuali-
sation to use.
Small multiples
To show data
in different categories
that cannot be seen
correctly in a
single graphic.
Figure 4.1. Example of the 15 services with the largest number of budget items during 2015.
This graphic consists of two axes: a quantitative axis, which shows the scale of the values repre-
sented; and a text axis, which indicates the category to which the data belong. On the text axis
there is a series of bars, the length of which indicates the value of each category.
Figure 4.2. Example of a bar chart ordered to facilitate appreciation of the ranking of the categories.
When data correspond to time intervals (hours, days, months, etc.), they should be ordered ac-
cordingly. In this case, vertical bars are recommended.
Figure 4.3. If data correspond to time intervals, vertical bars are preferable.
Figure 4.4. The quantitative axis of the bar graph must always start at 0.
When the names of categories are long, it is advisable to write them horizontally to make them
easier to read.
Figure 4.5. Example of a horizontal bar chart to make the name of the categories easier to read.
You often want to compare one element with the rest. In these cases, it is useful to highlight the
corresponding bar with a different colour, so that it is easier to identify it and compare it with the
other elements.
Figure 4.7. Example of a bar graph highlighting the category that you want to compare with the others.
This graphic is a standard bar chart. In this case, it consists of a series of bars on the same axis.
Recommendations
For this type of graphic, you should follow the same recommendations as those for the stan-
dard bar chart. It is important not to overload it with too many categories, since this can lead to
problems of comprehension. In such cases, it is advisable to convert the graphic into separate
standard bar graphs.
Figure 4.9. Population of American states segmented by age ranges. Source: [Link]
This type of graphic is another extension of the standard bar chart. In this case, each of the bars
is divided according to the segments that compose it.
Recommendations
If you want to compare the composition of the bars with each other, this is not a good option,
since, as can be seen in the following image, it is difficult to compare the number of budget items
allocated to each subsector.
In this case, it would be better to divide the graphic into different bar charts, as can be seen in
the following example.
Figure 4.11. An alternative to the stacked bar chart is to show each subcategory in a different bar chart.
Figure 4.12. Adaptation of a chart that shows the percentage of Facebook posts by the Democratic and Republican parties in
the United States, dealing with different topical subjects. Source: [Link]
Figure 4.13. Example of a radar chart showing the most valued features in a telephone.
Source: [Link]
Radar charts show a set of axes arranged in a circle, each relative to a variable in our set of data.
The value for each variable is located at a point on the corresponding axis, and then a line that
joins these points is drawn. Often, the area of the resulting polygon is coloured.
The following diagram illustrates the structure of a radar chart:
In essence, this chart presents a set of values in an intuitive and easily interpreted way.
Recommendations
It is not a good idea to show too many variables at the same time, as this may generate shapes
that are too angular or strange, so that understanding and comparing different charts becomes
difficult.
Figure 4.15. Example of a radar chart with too many variables generating polygons with too many angles that are not recognis-
able.
This type of chart is especially useful when the order in which the different variables appear has
some meaning. For example, in football, the variables related to a team’s attacking manoeuvres
(number of goals, shots at goal, etc.) would be placed on the axes at the top of the chart, while
those related to defence (number of goals stopped, fouls, blocked passes, etc.) would be at the
bottom. A radar chart that shows an almost perfect circle will thus indicate a team that is very
good in both attack and defence. If, on the other hand, the chart is more complete at the top, it
will indicate that the team is better in attack than in defence.
This type of graphic does not allow us to represent variables that can have negative and positive
values.
Figure 4.16. Colour intensity chart showing the number of people infected by malaria in different countries. Image by permission
of OneTandem. Source: [Link]
Colour intensity charts (heatmaps) are derived from tables. Instead of representing values using
numbers, they are represented by the intensity of the colour of the cell they occupy.
Recommendations
Ordering the rows and columns of a colour intensity chart according to an established criterion
can facilitate the discovery of elements that have similar data. In the example above, countries
would need to be ordered by the number of people infected and the year.
If the colour intensity chart is used for a set of data where the variables have different scales, it is
advisable to standardise them, so that they all have the same range of values. In this article, dif-
ferent indicators are shown for NBA players: [Link]
a-heatmap-a-quick-and-easy-solution/.
Bullet charts are a variation on bar charts, with two elements that help to provide context:
• A vertical mark that indicates a value with which you can compare the variable shown (for
example, if the bar indicates sales, the vertical mark could indicate the sales target).
• A background with areas in different shades, indicating “acceptable”, “good” or “bad” ranges
for the variable in question.
Recommendations
Often the target marked for an indicator is annual (for example, an annual sales target). If the
chart is produced, for example, at mid-year, the target will probably not have been reached. For
such cases there is a version of the chart that shows both the current value (dark blue in the first
example) and the expected value (light blue).
Figure 4.20. Example of the use of small multiples: A chain of office supplies stores uses this technique to compare profits in
each US state for different product categories and their three types of customer.
More than a type of graphic, the concept of small multiples refers to a composition technique.
The same graphic is used repeatedly, to show the behaviour of the same variable for different
categories. The repetition and the fact that the graphics are very close together facilitates com-
parison between them.
Recommendations
It is important for the type of graphic used to be sufficiently clear for its meaning to be unders-
tood despite its small size.
It is also important to be consistent in the use of ranges and colours, so that it is easy to compare
all the graphics with each other.
Trends
an element in a variable, but also to understand how it has changed over time.
Mini-line graph
To contextualise
an indicator
Figure 4.21. Example of a line graph showing changes in the total sales of two products of a company.
The line graph allows us to represent changes in a variable over time. The time-linked or conti-
nuous variable is placed on the x-axis and the graph is constructed by having a series of points
connected by a line determined by the height corresponding to the variable on the y-axis.
A line graph may contain one or more lines, allowing you to compare changes in data from diffe-
rent categories, often shown in different colours.
If there is only one line, the area between it and the x-axis can be coloured, producing an area
graph.
Recommendations
Great care is needed with the range of values on the y-axis, since there is a risk of overempha-
sising changes in the variable over time. This effect can be seen in the following graphs, which
show changes in the numbers of people infected with malaria over time. The graph on the left
takes 0 as the minimum value of the y-axis, while the one on the right begins with 60,000 infec-
ted. As can be seen, in the graph on the right the slopes of the lines are much more pronounced
and give a sensation of large changes.
Generally speaking, it is a good idea to add reference lines in the form of a grid to help users to
estimate the value of the variables at each point.
The line graph also makes it possible to show more than one line at a time, as in the first example.
However, the inclusion of too many lines can considerably reduce the readability of the graph.
Figure 4.23. Example of a slope chart showing the classification of different football teams in the first division in the 2015-16
and 2016-17 seasons.
Recommendations
So that the slope chart is coherent, the maximum and minimum values of each axis must be the
same.
4.2.3. Sparklines
Figure 4.24. Sparklines used in the header of a dashboard to show changes in a company’s main indicators.
A sparkline is a graphic designed to provide context for a figure. They are usually used in das-
hboards and placed next to the value of a key indicator for the system, so they need to be quite
small.
Although sparklines were initially small graphics that helped to understand trends in a particular
metric, the concept has been extended in recent years and allows the use of other types of gra-
phic, such as bar charts.
Recommendations
As a general rule, the scale of the axes of a sparkline is not shown, since the objective is to give
an overview of the behaviour of the value shown.
7
More information on map projections: Wikipedia, Map projection.
<[Link] [Accessed: 21 November 2017]
Figure 4.25. Choropleth map showing unemployment rates in the United States in August 2016.
Source: [Link]
Choropleth maps show the values of a variable on a map by painting each region involved in a
particular colour. Colours can be used to represent a numerical variable or to show that a region
belongs to a particular category. Depending on our purpose, one colour scale or another will be
used.
Recommendations
Generally speaking, a map should be used to present information of a geographical nature. A
choropleth map would therefore not be suitable to show a ranking of different regions.
Figure 4.26. Proportional symbols map showing the number of people registered in the United States census.
Source: [Link]
In the proportional symbols map an icon or symbol, usually a circle that is proportional to the
variable it represents, is placed over the centre of the region to which it refers.
Unlike the choropleth map, the proportional symbols map allows us to show geographical data
without the need to identify political areas. Colour can also be used to add a second variable to
the map, or to emphasise the size of the variable.
Recommendations
As with choropleth maps, a proportional symbols map should be used only when we are identi-
fying geographical patterns and not to rank geographical locations according to a variable.
Remember that, as you can see in the example, in this type of map the symbols used to repre-
sent the values of the variable often overlap. In these cases, they should be made transparent so
that the map is easier to read and the symbols that appear below each other can be identified.
Parts of a whole
or when we want to analyse the answers to a survey. Below, we describe a series of graphics
that are very useful for communicating this type of information.
Treemaps
To show hierarchical
data and compare
one or two variables
among the different
elements.
The pie chart consists of a circle divided into sectors proportional to the values of each category.
Figure 4.28. In this pie chart we can see how the orange category occupies more than half of the total, a figure not reached by
the other sectors taken together.
Figure 4.29. Pie charts with many divisions are not a good representation as they confuse users and do not help them to understand the
difference between values. Source:
[Link]
The use of 3D is especially popular for this type of graphic; however, this technique makes them
even harder to understand because of the perspective.
Figure 4.30. Sectors in a 3D pie chart represented with a bar chart. Adaptation of an image by permission of Ann K. Emery.
Source: [Link]
Figure 4.31. Ordered data represented with a stacked bar chart. Adaptation of the image provided by permission of Ann K. Emery.
Source: [Link]
If you have data that follows a time sequence, it is preferable to use a line graph or bar chart.
Figure 4.32. It is better to show data that follows a time sequence with a line graph. Adaptation of an image by permission of Ann K.
Emery. Source: [Link]
Figure 4.33. It is preferable to use bars to compare values across various categories. Adaptation of an image by permission of Ann K.
Emery. Source: [Link]
An alternative to the pie chart is the doughnut chart. This is constructed exactly like the pie chart,
but with a hole in the middle, where a value can be shown. This option is often used to show the
progress of a value.
This type of graphic is used to express a specific value, usually a percentage. It takes the form
of a square or rectangle. The outer square represents the maximum value, and the number of
squares in another colour represents the current value.
Figure 4.36. Percentage of the population that drink cava in different continents. Visualisation by OneTandem.
Source: [Link]
Pictograms are a good alternative for displaying a single number, especially when it is a main
indicator ranging from 0% to 100%.
Recommendations
They are ideal for representing indicators with a maximum value of 100, for example, percenta-
ges that will never exceed 100%.
If it is important to include the decimal portion of the values, they are not a good alternative be-
cause each coloured square represents one unit.
Figure 4.37. Treemap representing the volume of sales (represented by size) and the profits (represented by colour) of a ficti-
tious company in different American cities grouped by state.
The treemap allows you to see a hierarchical grouping of values. It is constructed by dividing a
rectangle into smaller rectangles, whose size and colour give specific information.
Recommendations
The size of the rectangles means that it not easy to see the differences between very similar
values. This type of graphic is not suitable for identifying the highest values (i.e. to rank them).
The treemap is especially useful in interactive visualisation. The name of each rectangle can then
be shown when the cursor passes over it. This avoids one of the main problems of this type of
graphic: if the square is too small, the name of the category it represents is not visible.
Distribution
or random distribution of values. This will indicate, for example, whether it is appropriate to use
statistical measurements such as the average or the median.
Figure 4.38. Histogram of a normal distribution showing the ratings of various TV series. Visualisation by OneTandem.
Source: [Link]
A histogram is a bar chart showing the distribution of values of a variable. The height of each bar
represents the frequency with which values occur within the range assigned to it.
Recommendations
By convention, histogram bars must not be separated from each other.
The shape of a histogram will depend on the range of values that each bar represents. This ran-
ge is important as it allows us to group the values and will therefore determine the shape of the
histogram. Generally speaking, programs automatically decide the size of this range according
to the bars involved, although it can be edited. Below we can see two distributions of the same
data, using two different ranges to group the values in the data set. The example on the left uses
a range of 0.5, while the one on the right uses 0.25.
There is no ideal formula for calculating the range, so it will be the data analyst’s job to try different
ranges to get an idea of how the data is distributed.
Figure 4.40. Box plot of ratings of chapters in television series (data extracted from [Link]). Each dot is a different
chapter. Graph prepared by OneTandem.
Recommendations
To make good use of this type of graphic you need to be familiar with the fundamental concepts
of descriptive statistics, in particular, the median and the quartiles. It is not, therefore, suitable
for all audiences.
Like bar charts, box plots can be horizontal or vertical. They are usually vertical, but if the names
of the categories are very long, they will need to be horizontal so that the names fit in comforta-
bly.
The horizontal lines at the ends of the box plot are called “whiskers” and help to identify atypical
values. There are different conventions when it comes to positioning these lines:
• They can be placed at 1.5 times the interquartile distance from the first and third quartiles
respectively.
• At percentiles 9 and 91 of the data.
• Or at percentiles 2 and 98 of the data.
Correlations
noted that a correlation between two variables does not necessarily imply that one causes the
other.
Parallel coordinates
To explore sets of
multidimensional data.
Figure 4.42. Scatter plot showing sales and profit per product. Each dot is a product.
This type of graphic allows you to represent the values of two variables on two axes (x and y
coordinates). Each element is represented as a dot based on its values on each of the axes.
• The shape: If the dots are positioned around a straight line, it means that there is a linear
correlation between the variables. If the dots are distributed around a rising line, it means that
they are positively correlated (when the value of x increases, so does that of y, as in the main
example). If they follow a falling line, they are negatively correlated. However, if they are not
distributed along any line, they are not correlated. Apart from linear correlation, expressed as
a straight line, there could also be exponential, logarithmic or polynomial correlation.
Figure 4.44. From left to right, exponential, polynomial and logarithmic correlation.
• Atypical values: are those that fall outside the main shape in a scatter plot and correspond
to exceptions that need to be explained or removed from the graphic to avoid distorting its
shape. This is the case of the blue dots in the following figure.
Figure 4.46. Example of a scatter plot with atypical values shown in blue.
5
Comerç
Sectors amb molta feina i
Sectors amb més feina i
poc risc d'automatització.
més risc d'automatització.
Són els sectors més
Estan en perill 5 milions
atractius en l'era de la
Sanitat de llocs de feina.
robotització. 4
Nombre de treballadors actualment (milions)
3 Ciència i tecnologia
Educació
Administració
Manufactura
Hostaleria
Construcció
2
Transport
Administració pública
Comunicacions
0% 5% 10% 15% 20% 25% 30% 35% 40% 45% 50% 55% 60% 65%
Figure 4.47. Automation and the labour market. Image by permission of OneTandem.
Source: [Link]
Finally, we should bear in mind that adding reference lines to scatter plots in the form of a grid is
very useful because they help users to estimate the values of the variables at each point.
Figure 4.48. Bubble chart show the relationship between profit, sales and discount for a series of products. The size and colour
of the bubble represents the size of the discount. It can be seen that the products that make a loss are not always the ones that
have the largest discount.
The bubble chart is a scatter plot where the area and the colour of the bubbles can represent
additional variables. In this way, up to 5 different variables can be represented at the same time:
Recommendations
Apart from the recommendations for the scatter plot, in the case of the bubble chart it is impor-
tant to decide which variable is used for each of the visual attributes (the position marked by the
x and y axes, the size and the colour).
The main variables in the analysis are located along the x and y axes so that it can be seen more
easily if they are related to each other or if there are elements that are grouped around certain
values.
Size is normally used to represent a variable for which you want to identify the largest element.
Colour is used to distinguish the categories to which each element belongs or to compare large
and small values for another variable (in this case a numerical range of colours should be used).
Lastly, animation is used to show changes in the values of the elements over time, so that users
can see how the bubbles move up or down, to the right or the left, as time passes.
Parallel coordinates allow you to explore the relationships between different variables for a set
of elements. Each variable is represented by a vertical axis, and the elements in the data set are
represented by joining the points that indicate the value of each variable on each axis.
Recommendations
Parallel coordinates are especially useful in interactive presentations. The most commonly used
interaction is filtering, which allows the user to select a range of values on one axis and see how
they behave with regard to the other variables. In the following image, for example, the elements
have been filtered for a given range of the weight variable. Among other things, it can be seen
that these elements tend to have low values for the economy (mpg) variable and there may, the-
refore, be an inverse relationship between these two variables.
Figure 4.51. Network of characters in Les Miserables. Each link indicates that they appear in the same scene.
Source: [Link]
Node-link diagrams use circles, called nodes, to represent elements, and links shown as lines
between nodes, indicating a relationship between them.
Connections can be established based on different factors, such as friendships between the
members of a social network, or characters that appear in the same scenes in a film. The di-
rection of the relationship can be represented with lines that end in the shape of an arrow. For
example, in a network like Facebook relationships between people are bi-directional, whereas in
Twitter, they are not (user A can follow user B, but user B need not follow user A).
The mathematical form created by establishing the connections between nodes using lines is
called a “graph”.
• Make sure that nodes overlap as little as possible and minimise the number of lines that cross.
• Include the entire network on the same screen.
• Facilitate the discovery of groups of similar elements (clusters).
These aims can usually be achieved by applying force-directed algorithms, which, according to
the structure of the network, determine the best position for each node.
Data set variables are also often used to modify the size or colour of the nodes and / or lines.
Figure 4.52. Sankey diagram showing the production and consumption chain of electrical energy.
Source: [Link]
The Sankey diagram shows different categories or states through which a variable changes in
value. It also shows that a category can be the sum of the values of various categories, as shown
in the example.
Recommendations
The most important thing to do when creating a Sankey diagram is to decide the vertical order
of the different categories that are represented in each of the states. The choice of position will
result in different degrees of overlapping between lines, which will make reading the graph ea-
sier or more difficult.
When preparing data visualisation, as well as analysing the data, we must decide on the best
way to represent them. That is why it is important to have a basic knowledge of the principles
of design, which will allow us to understand why we make certain decisions regarding visual
representation so that we can act accordingly. In this chapter we will discuss the main aspects
of design that must be taken into account in the visualisation of data.
It should also be borne in mind that Government of Catalonia data visualisation is prepared in
accordance with corporate identity standards. All departments, affiliated autonomous bodies
and public companies under its authority must apply the following guidelines, as defined in the
Corporate Identity website:
Figure 5.1. The table shows the quarterly sales of a company to small business customers and large companies.
Figure 5.2. The line graph shows the quarterly sales of the company to small business customers and large companies.
Thanks to the use of the line graph, we can see the cyclical trend and an overall increase in sa-
les. Why does a line graph allow this while a table does not? To understand this, one must bear
in mind the three types of memory that act in the brain:
• iconic memory
• long-term memory
• working memory
Iconic memory deals with the information that the brain receives from visual receptors. It is pro-
cessed quickly without our being aware of it: this is known as “pre-attentive processing”. It does
not carry out an exhaustive analysis of everything we see, but it extracts a subset of significant
elements such as colour or shape.
Long-term memory is accessed most slowly. It enables us to store information that we have dealt
with extensively or that we have studied carefully. For example, the password for our e-mail, or
the material we have studied for an exam.
5.2. Colour
Colour is a pre-attentive visual attribute (processed by the iconic memory), and the most exten-
sively used together with shape. To use it to communicate information, you need to distinguish
between quantitative data and categorical data.
If the data are quantitative, we have two options:
• A scale using a single colour with different levels of saturation, as shown in Figure 5.3
• If we have a central value, such as zero, a dichotomous scale can be used, as we can see in
Figure 5.4
Figure 5.3. Single colour scale, which uses saturation to distinguish elements.
8
Miller, George Armitage. The magical number seven, plus or minus two: Some limits on our capacity for processing information.
Psychological Review, 1956
Although the use of colour is a very intuitive method of encoding, it is not a good choice when
you need to know how much bigger one value is than another. For example, we can see that one
colour is darker than another, but it is virtually impossible for us to see that one colour is twice
as dark as another.
If our data are categorical, colour can be used to distinguish between categories. We therefore
need to look for colours that are very different from each other.
Figure 5.5. Scale of categorical colours provided by the D3js library. Source: [Link]
Although a huge number of colours exist, it is difficult to find more than ten that are markedly dif-
ferent from each other. Colour is, therefore, not a good visual attribute if we have a large number
of categories to represent. Also, we should remember that the working memory has difficulty
dealing with more than 5 to 9 different elements (in particular, more than 5 to 9 colours). In such
cases, data need to be grouped differently or filtering mechanisms can be introduced to allow
users to go from the whole to particular cases interactively.
Finally, it should be taken into account that approximately 5% to 10% of the population suffer
from colour-blindness and, therefore, have difficulty distinguishing between red and green. The-
re are colour scales specially designed for these people, which we recommend you to use
Length is the attribute that our brain processes most easily. This is why bar-charts are so popular.
Size is also used quite a lot, for example in bubble charts. Spatial grouping works very well in a
scatter graph where we can detect a group of elements which are clearly differentiated from the
rest. Other attributes are used to a lesser extent, but they can also be useful in certain situations.
Tooltips
Tooltips are small boxes with text and/or visual elements that appear when we select an item in
our visualisation. They are the basic way of providing details about an element in which the user
is interested.
Figure 5.7. When we move the cursor over the “Europe” bar, a panel appears with a bar chart showing details of the countries
that are beneficiaries.
Figure 5.8. Example of linking and brushing used in the visualisation “De què parlen els partits del 21-D” (What the 21-D parties
say) created by OneTandem. Source: [Link]
tatCatalunya21D/heatmap
“Above all else show the data (...). Graphical excellence is that which gives to
the viewer the greatest number of ideas in the shortest time with the least ink
in the smallest space.”
The following animated GIF shows clearly how to apply this principle to the design of graphics:
[Link]
Figure 5.9. Partial capture of the OECD “Better Life Index” visualisation. Source: [Link]
However, it is very important to understand that the primary objective of visualising data in an
analytical environment must be to generate knowledge.
9
Tufte, Eduard R. The visual display of quantitative information. Graphic Press, 1983
Unemployment Galicia
-0.39
Asturias
-0.82
Cantabria
+0.54
País Vasco
+0.84
Navarra
+0.39
Aragón
+2.48
rates by region
(in October)
Percentage change Castilla Cataluña
compared to previous month y León +1.39
Madrid +0.77
+2.33 La Rioja
+1.02
+0.83% a +3.42%
+0.82% (average) Extremadura Castilla-
La Mancha
+0.82% a -4.27% -1.86 +1.78 Baleares
-4.27
Comunidad
Andalucía Valenciana
Canarias -0.30 +2.08
+3.42
Murcia
Ceuta Melilla +1.78
+1.81 +1.93
Figure 5.10. Unemployment rates by autonomous community. Image by permission of Alberto Cairo from his book El Arte Fun‑
cional 10.
The aim of the visualisation is to show which autonomous communities have improved more and
which are having more problems compared with the previous month. However, the approach
chosen makes this task complicated, as only three shades of colour have been used and within
communities appearing in the same colour we have no choice but to inspect the numbers, try to
memorise them and then work out the order of the values.
Figure 5.11. Redesign proposed by Alberto Cairo. Image by permission of the same author from his book El Arte Funcional.
10
Cairo, Alberto. El arte funcional. Infografía y visualización de información. Alamut 2011.
5.5.4. Text
Sometimes, in data visualisation, the text is not given the importance it deserves. For example,
a suitable heading may be the key to attracting the reader’s attention. The texts that accompany
graphs as headings or notes are necessary to ensure that they are correctly understood. The
text may even be the main feature of the visualisation. We can see this clearly with an example.
The following chart shows the number of steps that a person has walked in a day.
A general text is used to describe the content of the graphic. The text is not incorrect, but it does
not help us to understand the meaning of the data. Now see how the same graphic changes, if
we use other texts:
700
10am 5pm
La llevadora em Arriben els familiars.
mana marxar a Em moc per evitar una
600 12am
esmorzar. "Tens habitació massa plena.
Camino per
un dia molt llarg
l'habitació. Em
per endavant". 9pm
sento inútil al
500 Després de quasi
costat del
patiment de la 20 hores despert,
9am m'adormo. Encara
meva dona.
La meva dona no sé que en Teo
4am
Passes
100
0
1 AM 2 AM 3 AM 4 AM 5 AM 6 AM 7 AM 8 AM 9 AM 10 AM 11 AM 12 PM 1 PM 2 PM 3 PM 4 PM 5 PM 6 PM 7 PM 8 PM 9 PM 10 PM
A simple graphic becomes a personal, even moving, story. The data are exactly the same but,
thanks to the text, we place them in context, we understand what is behind them and what they
mean. As we can see, in a visualisation, the text can be even more important than the numerical
data itself.
5.5.5. Layout
“Layout” refers to the distribution of the different elements in the data visualisation: graphics,
texts, images, filters, etc. As a rule, we recommend you reserve space for:
• information about the methodology and the data sources, which will generally be placed at the
bottom of the visualisation
• filters and selectors, which will generally be grouped in a column to the right or the left, or in
a row at the top
• help mechanisms, which will generally be next to each graphic (a general “Help” icon can also
be placed in the upper right corner).
It should also be remembered that people generally read from top to bottom and from left to
right, so the most important information should be positioned in the upper left corner.
Below we list some books that may be useful if you wish to deepen your knowledge of data
visualisation.
Few, Stephen. Now You See It. Analytics Press, 2009.
Meirelles, Isabel. Design for Information. Rockport Publishers, 2013.
Cotgreave, Andy; Shaffer, Jeffrey; Wexler, Steve. The Big Book of Dashboards: Visualizing Your
Data Using Real‑World Business. Wiley, 2017.
Cairo, Alberto. The Truthful Art. [s. n.] 2016.
7.1. Glossary
Cluster analysis: Cluster analysis is a data analysis technique that allows us to segment a set
of data automatically into subgroups (called clusters) with similar values.
Indicator: Also called metrics. An indicator is a numerical data variable that can be measured
and aggregated.
Quartile: The quartiles of a ranked set of data values are the four subsets whose boundaries are
the three quartile points (definition in Wikipedia: [Link]
Variable: All attributes of our data, whether numerical or text. For example, in a spreadsheet, the
variables in our data set are the columns.
You will find more terms related to this subject in Termcat data visualisation terminology: http://
[Link]/termvisdades
Tableau ([Link] Tableau has become the standard tool in the world of
data visualisation, especially as used in the fields of business intelligence or visual analytics.
Tableau makes it possible to change the type of visualisation with a single click, and facilitates
the creation of interactive graphics. The tool also makes it easy to connect to multiple sources of
data, ranging from simple spreadsheets to data bases and massive data platforms.
Tableau can be used in the prototyping and final stages of visualisation, but it can also be a good
tool to visually explore data in the strategy stage.
Although Tableau is easy to use to create simple graphics, the learning curve is quite steep when
you want to prepare advanced visualisation. Finally, it is interesting to note that, although it is a paid-
for tool, there is also a free version (Tableau Public [Link]
Figure 7.5. Screenshot of one of the dashboards provided by Tableau in its Desktop version.
Source: [Link]
RAL_2017.html#8/41.882/1.208
Figure 7.11. Screenshot of a collage created with D3 using visualisation developed with this library.
Source: [Link]