R Network Visualization Workshop
R Network Visualization Workshop
Contents
2 Colors in R plots 6
1
6 Interactive network visualizations 52
6.1 Simple plot animations in R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
6.2 Interactive JS visualization with visNetwork . . . . . . . . . . . . . . . . . . . . . . . 53
6.3 Interactive JS visualization with threejs . . . . . . . . . . . . . . . . . . . . . . . . . 58
6.4 Interactive JS visualization with networkD3 . . . . . . . . . . . . . . . . . . . . . . . 60
2
1 Introduction: network visualization
The main concern in designing a network visualization is the purpose it has to serve. What are the
structural properties that we want to highlight? What are the key concerns we want to address?
T1 T2
A B
Network maps are far from the only visualization available for graphs - other network representation
formats, and even simple charts of key characteristics, may be more appropriate in some cases.
3
In network maps, as in other visualization formats, we have several key elements that control the
outcome. The major ones are color, size, shape, and position.
Color Position
Size Shape
Modern graph layouts are optimized for speed and aesthetics. In particular, they seek to minimize
overlaps and edge crossing, and ensure similar edge length across the graph.
Layout aesthetics
4
Note: You can download all workshop materials here, or visit [Link]/sunbelt2023.
This tutorial uses several key packages that you will need to install in order to follow along. Other
packages will be mentioned along the way, but those are not critical and can be skipped.
The main packages we are going to use are igraph (maintained by Gabor Csardi and Tamas Nepusz),
sna & network (maintained by Carter Butts and the Statnet team), ggraph(maintained by Thomas
Lin Pederson), visNetwork (maintained by Benoit Thieurmel), threejs (maintained by Bryan W.
Lewis), NetworkD3 (maintained by Christopher Gandrud), and ndtv (maintained by Skye Bender-
deMoll).
[Link]("igraph")
[Link]("network")
[Link]("sna")
[Link]("ggraph")
[Link]("visNetwork")
[Link]("threejs")
[Link]("networkD3")
[Link]("ndtv")
5
2 Colors in R plots
Colors are pretty, but more importantly, they help people differentiate between types of objects or
levels of an attribute. In most R functions, you can use named colors, hex, or RGB values.
In the simple base R plot chart below, x and y are the point coordinates, pch is the point symbol
shape, cex is the point size, and col is the color. To see the parameters for plotting in base R,
check out ?par.
5
4
3
2 4 6 8 10
You may notice that RGB here ranges from 0 to 1. While this is the R default, you can also set it
1:10
to the 0-255 range using something like rgb(10, 100, 100, maxColorValue=255).
We can set the opacity/transparency of an element using the parameter alpha (range 0-1):
5
4
3
0 1 2 3 4 5 6
6
par(bg="gray40")
[Link] <- grDevices::adjustcolor("557799", alpha=0.7)
plot(x=1:5, y=rep(5,5), pch=19, cex=12, col=[Link], xlim=c(0,6))
7
6
rep(5, 5)
5
4
3
0 1 2 3 4 5 6
In many cases, we need a number of contrasting colors, or multiple shades of a color. R comes with
some predefined palette function that can generate those for us. For example:
pal1 <- [Link](5, alpha=1) # 5 colors from the heat palette, opaque
pal2 <- rainbow(5, alpha=.5) # 5 colors from the heat palette, transparent
plot(x=1:10, y=1:10, pch=19, cex=5, col=pal1)
10
8
1:10
6
4
2
2 4 6 8 10
1:10
plot(x=1:10, y=1:10, pch=19, cex=5, col=pal2)
7
10
8
1:10
6
4
2
2 4 6 8 10
We can also generate our own gradients using colorRampPalette. Note that colorRampPalette
1:10
returns a function that we can use to generate as many colors from that palette as we need.
6
4
2
2 4 6 8 10
To add transparency to colorRampPalette, you need to use a parameter alpha=TRUE:
10:1
palf <- colorRampPalette(c(rgb(1,1,1, .2),rgb(.8,0,0, .7)), alpha=TRUE)
plot(x=10:1, y=1:10, pch=19, cex=5, col=palf(10))
10
8
1:10
6
4
2
2 4 6 8 10
10:1
8
3 Data format, size, and preparation
In this tutorial, we will work primarily with two small example data sets. Both contain data about
media organizations. One involves a network of hyperlinks and mentions among news sources. The
second is a network of links between media venues and consumers.
While the example data used here is small, many of the ideas behind the visualizations we will
generate apply to medium and large-scale networks. This is also the reason why we will rarely use
certain visual properties such as the shape of the node symbols: those are impossible to distinguish
in larger graph maps. In fact, when drawing very big networks we may even want to hide the
network edges, and focus on identifying and visualizing communities of nodes.
At this point, the size of the networks you can visualize in R is limited mainly by the RAM of your
machine. One thing to emphasize though is that in many cases, visualizing larger networks as giant
hairballs is less helpful than providing charts that show key characteristics of the graph.
The first data set we are going to work with consists of two files, “Dataset1-Media-Example-
[Link]” and “[Link]” (download here).
head(nodes)
head(links)
Next we will convert the raw data to an igraph network object. To do that, we will use the
graph_from_data_frame() function, which takes two data frames: d and vertices.
• d describes the edges of the network. Its first two columns are the IDs of the source and the
target node for each edge. The following columns are edge attributes (weight, type, label, or
anything else).
• vertices starts with a column of node IDs. Any following columns are interpreted as node
attributes.
library('igraph')
net <- graph_from_data_frame(d=links, vertices=nodes, directed=T)
net
9
## IGRAPH c3e1b57 DNW- 17 49 --
## + attr: name (v/c), media (v/c), [Link] (v/n), [Link] (v/c),
## | [Link] (v/n), type (e/c), weight (e/n)
## + edges from c3e1b57 (vertex names):
## [1] s01->s02 s01->s03 s01->s04 s01->s15 s02->s01 s02->s03 s02->s09 s02->s10
## [9] s03->s01 s03->s04 s03->s05 s03->s08 s03->s10 s03->s11 s03->s12 s04->s03
## [17] s04->s06 s04->s11 s04->s12 s04->s17 s05->s01 s05->s02 s05->s09 s05->s15
## [25] s06->s06 s06->s16 s06->s17 s07->s03 s07->s08 s07->s10 s07->s14 s08->s03
## [33] s08->s07 s08->s09 s09->s10 s10->s03 s12->s06 s12->s13 s12->s14 s13->s12
## [41] s13->s17 s14->s11 s14->s13 s15->s01 s15->s04 s15->s06 s16->s06 s16->s17
## [49] s17->s04
The two numbers that follow (17 49) refer to the number of nodes and edges in the graph. The
description also lists node & edge attributes, for example:
We also have easy access to nodes, edges, and their attributes with:
It is also easy to extract an edge list or matrix back from the igraph network:
10
# Get an edge list or a matrix:
as_edgelist(net, names=T)
as_adjacency_matrix(net, attr="weight")
Now that we have our igraph network object, let’s make a first attempt to plot it.
s11
s15 s05
s12
s13 s01
s04
s02
s14
s03
s17 s06
s10
s07 s09
s16 s08
That doesn’t look very good. Let’s start fixing things by removing the loops in the graph.
We could also use simplify to combine multiple edges by summing their weights with a command
like simplify(net, [Link]=list(Weight="sum","ignore")). Note, however, that this
would also combine multiple edge types (in our data: “hyperlinks” and “mentions”).
Let’s and reduce the arrow size and remove the labels (we do that by setting them to NA):
plot(net, [Link]=.4,[Link]=NA)
11
3.3 DATASET 2: matrix
Our second dataset is a network of links between news outlets and consumers. It includes two
files, “[Link]” and “[Link]” (down-
load here).
head(nodes2)
head(links2)
We can see that links2 is an adjacency matrix for a two-mode network. Two-mode or bipartite
graphs have two different types of actors and links that go across, but not within each type. Our
second media example is a network of that kind, examining links between news sources and their
consumers.
12
node attribute called type that is FALSE (or 0) for vertices in one mode and TRUE (or 1) for those
in the other mode.
head(nodes2)
head(links2)
13
4 Plotting networks with igraph
Plotting with igraph: the network plots have a wide set of parameters you can set. Those include
node options (starting with vertex.) and edge options (starting with edge.). A list of selected
options is included below, but you can also check out ?[Link] for more information.
NODES
[Link] Node color
[Link] Node border color
[Link] One of “none”, “circle”, “square”, “csquare”, “rectangle”
“crectangle”, “vrectangle”, “pie”, “raster”, or “sphere”
[Link] Size of the node (default is 15)
vertex.size2 The second size of the node (e.g. for a rectangle)
[Link] Character vector used to label the nodes
[Link] Font family of the label (e.g.”Times”, “Helvetica”)
[Link] Font: 1 plain, 2 bold, 3, italic, 4 bold italic, 5 symbol
[Link] Font size (multiplication factor, device-dependent)
[Link] Distance between the label and the vertex
[Link] The position of the label in relation to the vertex, where
0 is right, “pi” is left, “pi/2” is below, and “-pi/2” is above
EDGES
[Link] Edge color
[Link] Edge width, defaults to 1
[Link] Arrow size, defaults to 1
[Link] Arrow width, defaults to 1
[Link] Line type, could be 0 or “blank”, 1 or “solid”, 2 or “dashed”,
3 or “dotted”, 4 or “dotdash”, 5 or “longdash”, 6 or “twodash”
[Link] Character vector used to label edges
[Link] Font family of the label (e.g.”Times”, “Helvetica”)
[Link] Font: 1 plain, 2 bold, 3, italic, 4 bold italic, 5 symbol
[Link] Font size for edge labels
[Link] Edge curvature, range 0-1 (FALSE sets it to 0, TRUE to 0.5)
[Link] Vector specifying whether edges should have arrows,
possible values: 0 no arrow, 1 back, 2 forward, 3 both
OTHER
margin Empty space margins around the plot, vector with length 4
frame if TRUE, the plot will be framed
main If set, adds a title to the plot
sub If set, adds a subtitle to the plot
asp Numeric, the aspect ratio of a plot (y/x).
palette A color palette to use for vertex color
rescale Whether to rescale coordinates to [-1,1]. Default is TRUE.
14
We can set the node & edge options in two ways - the first one is to specify them in the plot()
function, as we are doing below.
s05
s15 s02
s01 s09
s03 s10
s04 s08
s16s06
s17 s07
s11
s12
s14
s13
# Set edge color to light gray, the node & border color to orange
# Replace the vertex label with the node names stored in "media"
plot(net, [Link]=.2, [Link]="orange",
[Link]="orange", [Link]="#ffffff",
[Link]=V(net)$media, [Link]="black")
[Link] CNN
Google News MSNBC
Yahoo News FOX News
ABC
15
The second way to set attributes is to add them to the igraph object. Let’s say we want to color
our network nodes based on type of media, and size them based on degree centrality (more links
-> larger node) We will also change the width of the edges based on their weight.
# Compute node degrees (#links) and use that to set node size:
deg <- degree(net, mode="all")
V(net)$size <- deg*3
# We could also use the audience size value:
V(net)$size <- V(net)$[Link]*0.6
16
plot(net, [Link]="orange", [Link]="gray50")
plot(net)
legend(x="bottomleft", c("Newspaper","Television", "Online News"), pch=21,
col="#777777", [Link]=colrs, [Link]=2, cex=.8, bty="n", ncol=1)
Newspaper
Television
Online News
Sometimes, especially with semantic networks, we may be interested in plotting only the labels of
the nodes:
17
plot(net, [Link]="none", [Link]=V(net)$media,
[Link]=2, [Link]="gray40",
[Link]=.7, [Link]="gray85")
Google News
Yahoo News
[Link]
[Link]
[Link] BBC
New York Post
USA Today
FOX News
Let’s color the edges of the graph based on their source node color. We can get the starting node
for each edge with the ends() igraph function. It returns the start and end vertex for edges listed
in the es parameter. The names parameter control whether the function returns edge names or
IDs.
18
4.2 Network layouts
Network layouts are simply algorithms that return coordinates for each node in a network.
For the purposes of exploring layouts, we will generate a slightly larger 100-node graph. We use
the sample_pa() function which generates a simple graph starting from one node and adding more
nodes and links based on a preset level of preferential attachment (Barabasi-Albert model).
plot([Link], layout=layout_randomly)
19
Or you can calculate the vertex coordinates in advance:
l <- layout_in_circle([Link])
plot([Link], layout=l)
l is simply a matrix of x, y coordinates (N x 2) for the N nodes in the graph. For 3D layouts, it
has x, y, and z coordinates (N x 3). You can easily generate your own:
This layout is just an example and not very helpful - thankfully igraph has a number of built-in
layouts, including:
20
# Randomly placed vertices
l <- layout_randomly([Link])
plot([Link], layout=l)
# Circle layout
l <- layout_in_circle([Link])
plot([Link], layout=l)
# 3D sphere layout
l <- layout_on_sphere([Link])
plot([Link], layout=l)
21
Fruchterman-Reingold is one of the most used force-directed layout algorithms out there.
Force-directed layouts try to get a nice-looking graph where edges are similar in length and cross
each other as little as possible. They simulate the graph as a physical system. Nodes are electrically
charged particles that repulse each other when they get too close. The edges act as springs that
attract connected nodes closer together. As a result, nodes are evenly distributed through the
chart area, and the layout is intuitive in that nodes which share more connections are closer to each
other. The disadvantage of these algorithms is that they are rather slow and therefore less often
used in graphs larger than ~1000 vertices.
l <- layout_with_fr([Link])
plot([Link], layout=l)
With force-directed layouts, you can use the niter parameter to control the number of iterations
to perform. The default is set at 500 iterations. You can lower that number for large graphs to get
results faster and check if they look reasonable.
22
l <- layout_with_fr([Link], niter=50)
plot([Link], layout=l)
The layout can also interpret edge weights. You can set the “weights” parameter which increases
the attraction forces among nodes connected by heavier edges.
You will also notice that the Fruchterman-Reingold layout is not deterministic - different runs will
result in slightly different configurations. Saving the layout in l allows us to get the exact same
result multiple times, which can be helpful if you want to plot the time evolution of a graph, or
different relationships – and want nodes to stay in the same place in multiple plots.
23
[Link]()
By default, the coordinates of the plots are rescaled to the [-1,1] interval for both x and y. You can
change that with the parameter rescale=FALSE and rescale your plot manually by multiplying the
coordinates by a scalar. You can use norm_coords to normalize the plot with the boundaries you
want. This way you can create more compact or spread out layout versions.
l <- layout_with_fr([Link])
l <- norm_coords(l, ymin=-1, ymax=1, xmin=-1, xmax=1)
par(mfrow=c(2,2), mar=c(0,0,0,0))
plot([Link], rescale=F, layout=l*0.4)
plot([Link], rescale=F, layout=l*0.6)
plot([Link], rescale=F, layout=l*0.8)
plot([Link], rescale=F, layout=l*1.0)
24
[Link]()
Some layouts have 3D versions that you can use with parameter dim=3. As you might expect, a
3D layout returns a matrix with 3 columns containing the X, Y, and Z coordinates of each node.
Another popular force-directed algorithm that produces nice results for connected graphs is Kamada
Kawai. Like Fruchterman Reingold, it attempts to minimize the energy in a spring system.
l <- layout_with_kk([Link])
plot([Link], layout=l)
25
Graphopt is a nice force-directed layout implemented in igraph that uses layering to help with
visualizations of large networks.
l <- layout_with_graphopt([Link])
plot([Link], layout=l)
The available graphopt parameters can be used to change the mass and electric charge of
nodes, as well as the optimal spring length and the spring constant for edges. The parameter
names are charge (defaults to 0.001), mass (defaults to 30), [Link] (defaults to 0), and
[Link] (defaults to 1). Tweaking those can lead to considerably different graph layouts.
par(mfrow=c(1,2), mar=c(1,1,1,1))
plot([Link], layout=l1)
plot([Link], layout=l2)
26
[Link]()
The LGL algorithm is meant for large, connected graphs. Here you can also specify a root: a node
that will be placed in the middle of the layout.
plot([Link], layout=layout_with_lgl)
The MDS (multidimensional scaling) algorithm tries to place nodes based on some measure of
similarity or distance between them. More similar nodes are plotted closer to each other. By
default, the measure used is based on the shortest paths between nodes in the network. We can
change that by using our own distance matrix (however defined) with the parameter dist. MDS
layouts are nice because positions and distances have a clear interpretation. The problem with
them is visual clarity: nodes often overlap, or are placed on top of each other.
plot([Link], layout=layout_with_mds)
27
Let’s take a look at all available layouts in igraph:
par(mfrow=c(3,3), mar=c(1,1,1,1))
for (layout in layouts) {
print(layout)
l <- [Link](layout, list(net))
plot(net, [Link]=0, layout=l, main=layout) }
28
layout_as_star layout_components layout_in_circle
29
4.3 Highlighting aspects of the network
Notice that our network plot is still not too helpful. We can identify the type and size of nodes,
but cannot see much about the structure since the links we’re examining are so dense. One way
to approach this is to see if we can sparsify the network, keeping only the most important ties and
discarding the rest.
hist(links$weight)
mean(links$weight)
sd(links$weight)
There are more sophisticated ways to extract the key edges, but for the purposes of this exercise
we’ll only keep ones that have weight higher than the mean for the network. In igraph, we can
delete edges using delete_edges(net, edges):
30
Another way to think about this is to plot the two tie types (hyperlink & mention) separately. We
will do that in section 5 of this tutorial: Plotting multiplex networks.
We can also try to make the network map more useful by showing the communities within it:
par(mfrow=c(1,2))
# We can also plot the communities without relying on their built-in plot:
V(net)$community <- clp$membership
colrs <- adjustcolor( c("gray50", "tomato", "gold", "yellowgreen"), alpha=.6)
plot(net, [Link]=colrs[V(net)$community])
[Link]()
Sometimes we want to focus the visualization on a particular node or a group of nodes. In our
example media network, we can examine the spread of information from focal actors. For instance,
let’s represent distance from the NYT.
The distances function returns a matrix of shortest paths from nodes listed in the v parameter
to ones included in the to parameter.
31
# Set colors to plot the distances:
oranges <- colorRampPalette(c("dark red", "gold"))
col <- oranges(max([Link])+1)
col <- col[[Link]+1]
2 1 1
0
2 2 1
1
2 1 2
2 2
3
2
3
3
32
We can highlight the edges going into or out of a vertex, for instance the WSJ. For a single node,
use incident(), for multiple nodes use incident_edges()
We can also point to the immediate neighbors of a vertex, say WSJ. The neighbors function
finds all nodes one step out from the focal [Link] find the neighbors for multiple nodes, use
adjacent_vertices() instead of neighbors(). To find node neighborhoods going more than one
step out, use function ego() with parameter order set to the number of steps out to go from the
focal node(s).
33
# Set colors to plot the neighbors:
vcol[[Link]] <- "#ff9d00"
plot(net, [Link]=vcol)
A way to draw attention to a group of nodes (we saw this before with communities) is to “mark”
them:
par(mfrow=c(1,2))
plot(net, [Link]=c(1,4,5,8), [Link]="#C5E5E7", [Link]=NA)
[Link]()
34
4.5 Interactive plotting with tkplot
R and igraph allow for interactive plotting of networks. This might be a useful option for you if
you want to tweak slightly the layout of a small graph. After adjusting the layout manually, you
can get the coordinates of the nodes and use them for other plots.
tkid <- tkplot(net) #tkid is the id of the tkplot that will open
l <- [Link](tkid) # grab the coordinates from tkplot
plot(net, layout=l)
As you might remember, our second media example is a two-mode network examining links between
news sources and their consumers.
head(nodes2)
head(links2)
plot(net2, [Link]=NA)
35
As with one-mode networks, we can modify the network object to include the visual properties that
will be used by default when plotting the network. Notice that this time we will also change the
shape of the nodes - media outlets will be squares, and their users will be circles.
# Media outlets are blue squares, audience nodes are orange circles:
V(net2)$color <- c("steel blue", "orange")[V(net2)$type+1]
V(net2)$shape <- c("square", "circle")[V(net2)$type+1]
# Media outlets will have name labels, audience members will not:
V(net2)$label <- ""
V(net2)$label[V(net2)$type==F] <- nodes2$media[V(net2)$type==F]
V(net2)$[Link]=.6
V(net2)$[Link]=2
NYT
BBC
USAT
WSJ
LATimes
ABC
WaPo CNN
FOX
MSNBC
In igraph, there is also a special layout for bipartite networks (though it doesn’t always work great,
and you might be better off generating your own two-mode layout).
36
Using text as nodes may be helpful at times:
Jill
MSNBCJim
CNN Ronda
Jo Sheila
Paul
Brian LATimes
NYT
John Mary
FOX
Jason BBC
Sandra
Lisa USAT
Nancy
Dan
ABC
Kate
WSJ
Dave
Anna
Ed
WaPo
Ted
Tom
In this example, we will also experiment with the use of images as nodes. In order to do this, you
will need the png package (if missing, install with [Link]('png')
# [Link]('png')
library('png')
37
V(net2)$raster <- list(img.1, img.2)[V(net2)$type+1]
By the way, we can also add any image we want to a plot. For example, many network graphs can
be largely improved by a photo of a puppy in a teacup.
38
# The numbers after the image are its coordinates
# The limits of your plotting area are given in par()$usr
We can also generate and plot bipartite projections for the two-mode network: co-memberships are
easy to calculate by multiplying the network matrix by its transposed matrix, or using igraph’s
[Link]() function.
par(mfrow=c(1,2))
39
WaPo Mary
Paul
John
ABC
FOX Nancy
WSJ Sandra
MSNBC Ronda
Dan
Anna
CNN Ed
USAT Sheila
Kate
LATimes
Brian Jim
BBC Lisa
Dave Jo
Jason Jill
Tom
NYT Ted
[Link]()
In some cases, the networks we want to plot are multigraphs: they can have multiple edges con-
necting the same two nodes. A related concept, multiplex networks, contain multiple types of ties.
For instance, we can represent friendship, romantic, and work relationships between individuals in
a single multiplex network.
In our example network, we also have two tie types: hyperlinks and mentions. One thing we can
do with them is plot each type of tie separately:
40
net.m <- net - E(net)[E(net)$type=="hyperlink"] # another way to delete edges:
net.h <- net - E(net)[E(net)$type=="mention"] # using the minus operator
41
Tie: Hyperlink Tie: Mention
[Link]()
In our example network, it so happens that we do not have node dyads connected by multiple types
of connections. That is to say, we never have both a ‘hyperlink’ and a ‘mention’ tie between the
same two news outlets. However, this could easily happen in a multiplex network.
One challenge in visualizing multigraphs is that multiple edges between the same two nodes may
get plotted on top of each other in a way that makes impossible to see them clearly. For example,
let us generate a very simple multiplex network with two nodes and three ties between them:
42
Because all edges in the graph have the same curvature, they are drawn over each other so that
we only see one of them. What we can do is assign each edge a different curvature. One useful
function in igraph called curve_multiple can help us here. For a graph G, [Link](G)
will generate a curvature for each edge that maximizes visibility.
It is a good practice to detach packages when we stop needing them. Try to remember that
especially with igraph and the statnet family packages, as bad things tend to happen if you have
them loaded together.
detach('package:igraph')
The igraph package is only one of many available network visualization options in R. This section
provides a few quick examples illustrating other available approaches to static network visualization.
Plotting with the network package is very similar to that with igraph - although the notation
is slightly different (a whole new set of parameter names!). This package also uses less default
controls obtained by modifying the network object, and more explicit parameters in the plotting
function.
43
Here is a quick example using the (by now familiar) media network. We will begin by converting
the data into the network format used by the Statnet family of packages (including network, sna,
ergm, stergm, and others).
As in igraph, we can generate a ‘network’ object from an edge list, an adjacency matrix, or an
incidence matrix. You can get the specifics with ?[Link]. Here we will use the
edge list and the node attribute data frames to create the network object. One specific thing to
pay attention to here is the [Link] parameter. It is set to TRUE by default, and that setting
causes the network object to disregard edge weights.
library('network')
Here again we can easily access the edges, vertices, and the network matrix:
net3[,]
net3 %n% "[Link]" <- "Media Network" # network attribute
net3 %v% "media" # Node attribute
net3 %e% "type" # Node attribute
Note that - as in igraph - the plot returns the node position coordinates. You can use them in
other plots using the coord parameter.
44
l <- plot(net3, [Link]=(net3 %v% "[Link]")/7, [Link]="col")
plot(net3, [Link]=(net3 %v% "[Link]")/7, [Link]="col", coord=l)
detach('package:network')
The network package also offers the option to edit a plot interactively, by setting the parameter
interactive=T:
For a full list of parameters that you can use in the network package, check out ?[Link].
The ggplot2 package and its extensions are known for offering the most meaningfully structured
and advanced way to visualize data in R. In ggplot2, you can select from a variety of visual building
blocks and add them to your graphics one by one, a layer at a time.
45
The ggraph package takes this principle and extends it to network data. In this section, we’ll only
cover the basics without providing a detailed overview of the grammar of graphics approach. For a
deeper look, it would be best to get familiar with ggplot2 first, then learn the specifics of ggraph.
One good news is that we can use our igraph objects directly with the ggraph package. The
following code gets the data and adds separate layers for nodes and links.
library(ggraph)
library(igraph)
ggraph(net) +
geom_edge_link() + # add edges to the plot
geom_node_point() # add nodes to the plot
## Warning: Using the ‘size‘ aesthetic in this geom was deprecated in ggplot2 3.4.0.
## i Please use ‘linewidth‘ in the ‘default_aes‘ field and elsewhere instead.
## This warning is displayed once every 8 hours.
## Call ‘lifecycle::last_lifecycle_warnings()‘ to see where this warning was
## generated.
You will also recognize here some network layouts familiar from igraph plotting: ‘star’, ‘circle’,
‘grid’, ‘sphere’, ‘kk’, ‘fr’, ‘mds’, ‘lgl’, etc.
ggraph(net, layout="lgl") +
geom_edge_link() +
ggtitle("Look ma, no nodes!") # add title to the plot
46
Look ma, no nodes!
Here we can use geom_edge_link() for straight edges, geom_edge_arc() for curved ones, and
geom_edge_fan() when we want to make sure any overlapping multiplex edges will be fanned out.
As in other packages, we can set visual properties for the network plot by using key function
parameters. For instance, nodes have color, fill, shape, size, and stroke. Edges have color,
width, and linetype. Here too the alpha parameter controls transparency.
ggraph(net, layout="lgl") +
geom_edge_fan(color="gray50", width=0.8, alpha=0.5) +
geom_node_point(color=V(net)$color, size=8) +
theme_void()
As in ggplot2, we can add different themes to the plot. For a cleaner look, you can use a minimal
or empty theme with theme_minimal() or theme_void().
47
ggraph(net, layout = 'linear') +
geom_edge_arc(color = "orange", width=0.7) +
geom_node_point(size=5, color="gray50") +
theme_void()
The ggraph package also uses the traditional ggplot2 way of mapping aesthetics: that is to say, of
specifying which elements of the data should correspond to different visual properties of the graphic.
This is done using the aes() function that matches visual parameters with attribute names from
the data. In the code below, the edge attribute type and node attribute [Link] are taken
from our data as they are included in the igraph object.
ggraph(net, layout="lgl") +
geom_edge_link(aes(color = type)) + # colors by edge type
geom_node_point(aes(size = [Link])) + # size by audience size
theme_void()
48
[Link]
20
30
40
50
60
type
hyperlink
mention
One great thing about ggplot2 and ggraph you can see above is that they automatically generate
a legend which makes plots easier to interpret.
We can add a layer with node labels using geom_node_text() or geom_node_label() which cor-
respond to similar functions in ggplot2.
49
[Link]
Google News
ABC
NY Times
Washington Post
FOX News
detach("package:ggraph")
While those are not discussed here, note that ggraph offers a number of other interesting ways to
represent networks, including dendrograms, treemaps, hive plots, and circle plots.
At this point it might be useful to provide a quick reminder that there are many ways to represent
a network not limited to a hairball plot.
For example, here is a quick heatmap of the network matrix:
50
[Link]
[Link]
[Link]
[Link]
Google News
Yahoo News
BBC
ABC
FOX News
MSNBC
CNN
New York Post
LA Times
USA Today
Wall Street Journal
Washington Post
NY Times
[Link]
[Link]
[Link]
[Link]
Google News
Yahoo News
BBC
ABC
FOX News
MSNBC
CNN
New York Post
LA Times
USA Today
Wall Street Journal
Washington Post
NY Times
Depending on what properties of the network or its nodes and edges are most important to you,
simple graphs can often be more informative than network maps.
0.8
0.4
0.0
0 2 4 6 8 10 12
Degree
51
6 Interactive network visualizations
If you have already installed “ndtv”, you should also have a package used by it called “animation”.
If not, now is a good time to install it with [Link]('animation'). Note that this
package provides a simple technique to create various (not necessarily network-related) animations
in R. It works by generating multiple plots and combining them in an animated GIF.
The catch here is that in order for this to work, you need not only the R package, but also an
additional software called ImageMagick ([Link] You probably don’t want to
install that during the workshop, but you can try it at home.
The good news is that once you figure this out, you can turn any series of R plots (network or not!)
into an animated GIF.
library('animation')
library('igraph')
We will now generate 4 network plots (the same way we did before), only this time we’ll do it
within the saveGIF command. The animation interval is set with interval, and the [Link]
parameter controls name of the gif.
l <- layout_with_lgl(net)
52
col[setdiff(step.3, step.2)] <- "#FFDD1F"
plot(net, [Link]=col, layout=l) },
interval = .8, [Link]="network_animation.gif" )
detach('package:igraph')
detach('package:animation')
These days it is fairly easy to export R plots to HTML/JavaScript output. There are a number
of packages like rcharts and htmlwidgets that can help you create interactive web charts right
from R. One thing to keep in mind though is that the network visualizations created that way are
most helpful as a starting point for further work. If you know a little bit of javascript, you can use
them as a first step and tweak the results to get closer to what you want.
Here we will take a quick look at visNetwork which generates interactive network visualizations us-
ing the [Link] javascript library. You can install the package with [Link]('visNetwork').
We can visualize our media network right away: visNetwork() will accept our node and link data
frames. As usual, the node data frame needs to have an id column, and the link data needs to
have from and to columns denoting the start and end of each tie.
library('visNetwork')
visNetwork(nodes, links)
53
If we want to set specific height and width for the interactive plot, we can do that with the height
and width parameters. As is often the case in R, the title of the plot is set with the main parameter.
The subtitle and footer can be set with submain and footer respectively.
Like the igraph package, visNetwork allows us to set graphic properties as node or edge attributes.
We can simply add them as columns in our data before we call the visNetwork() function. Check
out the available options with:
?visNodes
?visEdges
In the following code, we are changing some of the visual parameters for nodes. We start with
the node shape (the available options for it include ellipse, circle, database, box, text, image,
circularImage, diamond, dot, star, triangle, triangleDown, square, and icon). We are also
54
going to change the color of several node elements. In this package, background controls the node
color, border changes the frame color; highlight sets the color on mouse click, and hover sets
the color on mouseover.
# We'll start by adding new node and edge attributes to our dataframes.
[Link] <- nodes
[Link] <- links
visNetwork([Link], [Link])
55
visnet <- visNetwork([Link], [Link])
visnet
We can also set the visualization options directly with visNodes() and visEdges().
visNetwork offers a number of other options in the visOptions() function. For instance, we can
highlight all neighbors of the selected node (highlightNearest), or add a drop-down menu to
select subset of nodes (selectedBy). The subsets are based on a column from our data - here we
use the type label.
56
visNetwork can also work with predefined groups of nodes. The visual characteristics for nodes
belonging in each group can be set with visGroups(). We can add an automatically generated
group legend with visLegend().
57
6.3 Interactive JS visualization with threejs
Another good package exporting networks from R to javascript is threejs, which generates interac-
tive network visualizations using the [Link] javascript library and the htmlwidgets R package.
One nice thing about threejs is that it can directly read igraph objects.
You can install the package with [Link]('threejs'). If you get errors or warnings
using this library with the latest version of R, try also installing the development version of the
htmlwidgets package which may have bug fixes that will help:
devtools::install_github('ramnathv/htmlwidgets')
The main network plotting function here,graphjs, will take an igraph object. We could use
our initial net object with a slight modification: we will delete its graph layout and let threejs
generate one on its own. We cheated a bit earlier by assigning a function to the layout attribute
in the igraph object rather than giving it a table of node coordinates. This is fine by igraph, but
threejs will not let us do it.
library(threejs)
library(htmlwidgets)
library(igraph)
Note that RStudio for Windows may not render the threejs graphics properly. We will save the
output in an HTML file and open it in a browser. Some of the parameters that we can add include
main for the plot title; curvature for the edge curvature; bg for background color; showLabels
to set labels to visible (TRUE) or not (FALSE); attraction and repulsion to set how much nodes
attract and repulse each other in the layout; opacity for node transparency (range 0 to 1); stroke
to indicate whether nodes should be framed in a black circle (TRUE) or not (FALSE), etc.
For the full list of parameters, check out ?graphjs.
58
Once we open the resulting visualization in a browser, we can use the mouse scrollwheel to zoom
in and out, the left mouse button to rotate the network, and the right mouse button to pan.
We can also create simple animations with threejs by using lists of layouts, vertex colors, and
edge colors that will switch at each step.
As an additional example, we can take a look at the Les Miserables network included with the
package:
59
data(LeMis)
[Link] <- graphjs(LeMis, main="Les Miserables", showLabels=T)
print([Link])
saveWidget([Link], file="[Link]")
browseURL("[Link]")
We will also take a quick look at networkD3 which - as its name suggests - generates interactive
network visualizations using the D3 javascript library. If you d not have the networkD3 library,
install it with [Link]("networkD3").
The data that this library needs from is is in the standard edge list form, with a few little twists.
In order for things to work, the node IDs have to be numeric, and they also have to start from 0.
An easy was to get there is to transform our character IDs to a factor variable, transform that to
numeric, and make sure it starts from zero by subtracting 1.
library(networkD3)
The nodes need to be in the same order as the “source” column in links:
60
Now we can generate the interactive chart. The Group parameter in it is used to color the nodes.
Nodesize is not (as one might think) the size of the node, but the number of the column in the node
data that should be used for sizing. The charge parameter controls node repulsion (if negative) or
attraction (if positive).
Here we will create D3 visualizations using the ndtv package. You should not need additional
software to produce web animations with ndtv. If you want to save the animations as video files
(see ?saveVideo), you have to install a video converter called FFmpeg ([Link] To find
out how to get the right installation for your OS, check out ?[Link]. To use all available
layouts, you would also need to have Java installed on your machine.
[Link]('ndtv', dependencies=T)
As ndtv is part of the Statnet family, it will accept objects from the network package such as the
one we created earlier (net3).
library('ndtv')
net3
61
Most of the parameters below are self-explanatory at this point (bg is the background color of the
plot). Two new parameters we haven’t used before are [Link] and [Link]. Those
contain the information that we can see when moving the mouse cursor over network elements. Note
that the tooltip parameters accepts html tags – for example we will use the line break tag <br>.
The parameter launchBrowser instructs R to open the resulting visualization file (filename) in
the browser.
If you are going to embed the plot in a markdown document, use [Link]='inline' above.
Animations are a good way to show the evolution of small to medium size networks over time. At
present, ndtv is the best R package for that – especially since it now has D3 capabilities and allows
easy export for the Web.
In order to work with the network animations in ndtv, we need to understand Statnet’s dynamic
network format, implemented in the networkDynamic package. The format can be used to represent
longitudinal structures, both discrete (if you have multiple snapshots of your network at different
time points) and continuous (if you have timestamps indicating when edges and/or nodes appear
and disappear from the network). The examples below will only scratch the surface of temporal
62
networks in Statnet - for a deeper dive, check out Skye Bender-deMoll’s Temporal network tools
tutorial and the networkDynamic package vignette.
Let’s look at one example dataset included in the package, containing simulation data based on a
network of business connections among Renaissance Florentine families:
data([Link])
[Link]
head([Link]([Link]))
What we see here is a temporal edge list. An edge goes from a node with ID in the tail column
to a node with ID in the head column. Edges exist from time point onset to time point terminus.
As you can see in our example, there may be multiple periods (activity spells) where an edge is
present. Each of those periods is recorded on a separate row in the data frame above.
The idea of onset and terminus censoring refers to start and end points enforced by the beginning
and end of network observation rather than by actual tie formation/dissolution.
We can simply plot the network disregarding its time component (combining all nodes and edges
that were ever present):
plot([Link])
63
We can also use [Link]() to get a network that only contains elements active at a given
point, or during a given time interval. For instance, we can plot the network at time 1 (at=1):
Plot nodes and edges that were active for the entire period (rule=all) from time 1 to time 5:
Plot nodes and edges that were active at any point (rule=any) between time 1 and time 10:
64
Let’s make a quick d3 animation from the example network:
render.d3movie([Link],displaylabels=TRUE)
Next, we will create and animate our own dynamic network. Dynamic network objects can be
generated in a number of ways: from a set of networks/matrices representing different time points;
from data frames/matrices with node lists and edge lists indicating when each is active, or when
they switch state. You can check out ?networkDynamic for more information.
We are going to add a time component to our media network example. The code below takes a
0-to-50 time interval and sets the nodes in the network as active throughout (time 0 to 50). The
edges of the network appear one by one, and each one is active from their first activation until time
point 50. We generate this longitudinal network using networkDynamic with our node times as
[Link] and edge times as [Link].
If we try to just plot the networkDynamic network, what we get is a combined network for the
entire time period under observation – or as it happens, our original media example.
65
plot([Link], [Link]=(net3 %v% "[Link]")/7, [Link]="col")
One way to show the network evolution is through static images from different time points. While
we can generate those one by one as we did above, ndtv offers an easier way. The command to do
that is filmstrip(). As in the par() function controlling base R plot parameters, here mfrow sets
the number of rows and columns in the multi-plot grid.
Next, let’s generate a network animation. We can pre-compute the coordinates for it (otherwise
they get calculated when we generate the animation). Here [Link] is the layout algorithm
- one of “kamadakawai”, “MDSJ”, “Graphviz” and “useAttribute” (user-generated coordinates).
In filmstrip() above and in the animation computation below, [Link] is a list of parameters
controlling how the network visualization moves through time. The parameter interval is the
time step between layouts, [Link] is the period shown in each layout, rule is the rule for
displaying elements (e.g. any: active at any point during that period, all: active during the entire
period, etc).
render.d3movie([Link], usearrows = F,
displaylabels = F, label=net3 %v% "media",
bg="#ffffff", [Link]="#333333",
[Link] = degree(net3)/2,
[Link] = [Link] %v% "col",
[Link] = ([Link] %e% "weight")/3,
[Link] = '#55555599',
[Link] = paste("<b>Name:</b>", ([Link] %v% "media") , "<br>",
"<b>Type:</b>", ([Link] %v% "[Link]")),
66
[Link] = paste("<b>Edge type:</b>", ([Link] %e% "type"), "<br>",
"<b>Edge weight:</b>", ([Link] %e% "weight" ) ),
launchBrowser=T, filename="[Link]",
[Link]=list([Link] = 30, [Link] = F),
[Link]=list(mar=c(0,0,0,0)), [Link]='inline' )
render.d3movie([Link], usearrows = F,
displaylabels = F, label=net3 %v% "media",
bg="#000000", [Link]="#dddddd",
[Link] = function(slice){ degree(slice)/2.5 },
[Link] = [Link] %v% "col",
[Link] = ([Link] %e% "weight")/3,
[Link] = '#55555599',
[Link] = paste("<b>Name:</b>", ([Link] %v% "media") , "<br>",
"<b>Type:</b>", ([Link] %v% "[Link]")),
67
[Link] = paste("<b>Edge type:</b>", ([Link] %e% "type"), "<br>",
"<b>Edge weight:</b>", ([Link] %e% "weight" ) ),
launchBrowser=T, filename="[Link]",
[Link]=list([Link] = 15, [Link] = F), [Link]='inline',
[Link]=list(start=0, end=50, interval=4, [Link]=4, rule='any'))
The example presented in this section uses only base R and mapping packages. If you have experi-
ence with ggplot2, that package does provide a more versatile way of approaching this task. The
code using ggplot() would be similar to what you will see below, but you would use ‘borders()’ to
plot the map and ‘geom_path()’ for the edges.
In order to plot on a map, we will need a few more packages. As you will see below, maps will
let us generate a geographic map to use as background, and geosphere will help us generate arcs
representing our network edges. If you do not already have them, install the two packages, then
load them.
68
[Link]('maps')
[Link]('geosphere')
library('maps')
library('geosphere')
Let us plot some example maps with the maps library. The parameters of maps() include col for
the map fill, border for the border color, and bg for the background color.
[Link]()
The data we will use here contains US airports and flights among them. The airport file includes
geographic coordinates - latitude and longitude. If you do not have those in your data, you can the
geocode() function from package ggmap to grab the latitude and longitude for an address.
69
airports <- [Link]("[Link]", header=TRUE)
flights <- [Link]("[Link]", header=TRUE, [Link]=TRUE)
head(flights)
head(airports)
# Select only large airports: ones with more than 10 connections in the data.
tab <- table(flights$Source)
[Link] <- names(tab)[tab>10]
airports <- airports[airports$ID %in% [Link],]
flights <- flights[flights$Source %in% [Link] &
flights$Target %in% [Link], ]
In order to generate our plot, we will first add a map of the United states. Then we will add a
point on the map for each airport:
70
points(x=airports$longitude, y=airports$latitude, pch=19,
cex=airports$Visits/80, col="orange")
Next we will generate a color gradient to use for the edges in the network. Heavier edges will be
lighter in color.
For each flight in our data, we will use gcIntermediate() to generate the coordinates of the
shortest arc that connects its start and end point (think distance on the surface of a sphere). After
that, we will plot each arc over the map using lines().
for(i in 1:nrow(flights)) {
node1 <- airports[airports$ID == flights[i,]$Source,]
node2 <- airports[airports$ID == flights[i,]$Target,]
71
c(node2[1,]$longitude, node2[1,]$latitude),
n=1000, addStartEnd=TRUE )
[Link] <- round(100*flights[i,]$Freq / max(flights$Freq))
Note that if you are plotting the network on a full world map, there might be cases when the
shortest arc goes “behind” the map – e.g. exits it on the left side and enters back on the right (since
the left-most and right-most points on the map are actually next to each other). In order to avoid
that, we can use greatCircle() to generate the full great circle (circle going through those two
points and around the globe, with a center at the center of the earth). Then we can extract from
it the arc connecting our start and end points which does not cross “behind” the map, regardless
of whether it is the shorter or the longer of the two.
This is the end of our tutorial. If you have comments, questions, or want to report typos, please
e-mail netvis@[Link]. Check for updated versions of the tutorial at [Link].
72