Atlas Computer Networking
Atlas Computer Networking
IT UNIVERSITY OF COPENHAGEN
arXiv:2101.00863v1 [[Link]] 4 Jan 2021
T H E AT L A S F O R T H E
ASPIRING NETWORK
SCIENTIST
Copyright © 2021 Michele Coscia
michele coscia is employed by the it university of copenhagen, rued langgaards vej 7, 2300
copenhagen, denmark
[Link]
Licensed under the Apache License, Version 2.0 (the “License”); you may not use this file except in com-
pliance with the License. You may obtain a copy of the License at [Link]
LICENSE-2.0. Unless required by applicable law or agreed to in writing, software distributed under the
License is distributed on an “as is” basis, without warranties or conditions of any kind, either
express or implied. See the License for the specific language governing permissions and limitations under
the License.
1 Introduction 9
I Basics 21
2 Probability Theory 22
3 Basic Graphs 40
4 Extended Graphs 48
5 Matrices 68
II Simple Properties 89
6 Degree 90
9 Density 132
4 the atlas for the aspiring network scientist
17 Epidemics 237
26 Homophily 364
28 Core-Periphery 384
29 Hierarchies 396
IX Communities 416
47 Glossary 652
Bibliography 665
1
Introduction
2
AR Wallace. On the tendency of
Wallace2 , Darwin3 , and Huxley4 . And, again, genes are nothing more varieties to depart indefinitely from the
than interacting proteins. It’s interactions all the way down5 . In other original type. J. Linn. Soc. Lond. Zool., 3:
53–62, 1858
words: networks. You really should pay attention to them. 3
Charles Darwin. On the origin of
So we have reached the part of the introduction of any self- species. 1859
respecting network science book when we need to address the ques- 4
Thomas Henry Huxley. Evidence as to
tion: how did we get here? How did we discover that networks were Man’s Place in Nature. London, Williams
and Norgate, 1863
a thing, how to think about them, how to hone our tools to tame 5
Before you start protesting about
the complexity of reality, the same complexity you hopefully now preposterous examples, I will ask you
appreciate from the impressionistic picture I just painted? This is to be patient: you’ll discover in due
time that the brain, ecosystems, and
the time for the creation myth of network science. Which is problem- biological protein interactions are
atic, because creation myths are always a lie, as they try to identify classical examples of complex systems
studied via network analysis.
a discontinuity point in the continuous process of the expansion of
knowledge by accumulation.
But we need to start from somewhere, so what the hell.
If graphs have a three century long history, why does 99% of network
science happened from 1999 on? The reason is probably because,
up until the revolutionary invention of the computer, we only really
had a general intuition about the pervasiveness of networks, without
anything tangible to act upon.
Modern network science is a gift from sociology. Before sociology,
graphs were seen as exact and deterministic mathematical objects,
worthy of exploration through the manipulation of abstract symbols.
Sociologists saw the value in using these mathematical objects – sym-
bols – to investigate a statistical and stochastic reality. This was the
first – fundamental and necessary – explosion in possibilities for a
true network science. There were two problems, though: the trivial
12 the atlas for the aspiring network scientist
one was that we could only collect and manipulate data manually.
More importantly, we still didn’t have a unified language to repre-
sent all of reality as a symbol. The representations we had before
computers were ad-hoc, made only for a specific problem.
The value of the computer is not that it can perform lots of opera-
tions quickly, although that certainly helps. One can prove theorems
and lemmas without computers. Rather, the revolution of the com-
puter is in its parallel development of the usage of symbols to repre-
sent reality. Computers seem to be able to allow you to manipulate
anything: with spreadsheets you can tame problems in logistics, with
XML you can map semantic concepts, with media players you can 10
It has been legendarily said that VLC,
appreciate art and videos10 . And yet, inside computers you just have one of the most popular multimedia
a mass of zeroes and ones. software, can open anything – even a
can of tuna.
Thus, the power of the computer is its ability of seeing everything
– anything – as a symbol. Once everything is a symbol, you can
analyze it mathematically and understand its relations with other 11
One of the facts that never fails to
symbols11 . In this sense, the true revolution of the computer was not blow my mind is the realization that
pioneered by Babbage and von Neumann12 , with their mechanical a piece of software is, after all, just a
very cleverly composed number. Thus
inventions; but rather by Lovelace13 and Wittgenstein14 , with their
you can sum Adobe Photoshop to
logical inventions concerning the manipulation of symbols and their Google Chrome, although the result
interactions. won’t probably make much sense. That
is also why there exist such a thing
Thus, the second half of the XX century was the moment when we as an “illegal number” ([Link]
started connecting symbols together in the same place: the memory [Link]/wiki/Illegal_number)
ignored29 . A great favorite of mine combines the computer science university press, 2018b
methodology with applications in economics30 . If you need yet an- 27
Albert-László Barabási. Linked: The
other proof of the breadth of approaches in network science, consider new science of networks, 2003
28
Duncan J Watts. Six degrees: The
that another major book on the topic was authored by a chemist31 . science of a connected age. WW Norton &
The natural question now is: if there are already so many network Company, 2004
science books and they are all great, what is the need of this one 29
Filippo Menczer, Santo Fortunato,
and Clayton A Davis. A First Course in
you’re reading? Is it just for updating with the newest developments Network Science. Cambridge University
in the field? Not really. I hope you noticed that, when presenting the Press, 2020
other network science books, I never introduced them as books writ- 30
David Easley and Jon Kleinberg.
Networks, crowds, and markets: Reasoning
ten by a network scientist. This was not by accident. My impression about a highly connected world. Cambridge
is that these books are aimed at introducing people from a variety of University Press, 2010
disciplines into the skills and tools of network science, rather than 31
Ernesto Estrada. The structure of
complex networks: theory and applications.
examining network science from within. Oxford University Press, 2012
There are now PhD programs for network scientists32 , but they 32
[Link]
are only a handful years old, meaning that there are only a few [Link]/phd
14 the atlas for the aspiring network scientist
graduates coming out of them and they do not have yet the time or
the experience to write a network science book. Worse still, I believe
there are even fewer master and bachelor programs in network
science, if any. This means that every book you can find on network
science really is “something adapted to network science”, with that
something being sociology, physics, computer science, archaeology, or
other.
This book has the – probably overambitious – aim of being really
the first network science book. In other words, the difference between
this and the other books is that this book considers “network science”
not as something one attaches to another discipline, but rather it
is a discipline in itself. People can – and should! – be trained from
scratch in it.
I believe my background is as close as it could be to the right mix
that network analysis requires. I am a digital humanist, a field pio- 33
Roberto Busa. Index thomisticus sancti
neered by Busa33 which focuses on the digital processing of content thomae aquinatis operum omnium
produced by humans. This is to say: the mathematical manipula- indices et concordantiae in quibus
verborum omnium et singulorum
tion and analysis of symbols representing different facets of reality
formae et lemmata cum suis frequentiis
– which are not necessarily mathematical – and their connections. If et contextibus variis modis referuntur.
this sounds familiar, it is because this is the exact characterization of 1974
36
[Link]
long and perilous ladder of fields in descending order of purity36 .
But this is not a given: some phenomena might not be reducible to 37
Philip W Anderson. More is different.
the underlying laws37 . Maybe we do need to toss away a good chunk Science, 177(4047):393–396, 1972
of physics, add a bunch of new tools, to understand this new field
called “chemistry”, because the change in scale causes the emergence
of new phenomena.
We already have a theory for this compartmentalization of the
scientific investigation. This is what Hayek called “division of knowl- 38
Friedrich August Hayek. The use
edge”38 , which is a much more powerful concept than Smith’s classi- of knowledge in society. The American
cal division of labor39 . If I specialize as a chemist and hone my skills economic review, 35(4):519–530, 1945
and tools to that specific task, I can be immensely more productive,
39
Adam Smith. The Wealth of Nations.
1776
because I am outsourcing all other knowledge discovery endeavors to
other specialists. This is how societies grow their pool of knowledge
efficiently. However, the result is that, now, no individual can really
fully grasp a well-rounded picture of reality. The collective society
can, but not its individual components. It is all deformed by the lens
of their specialization.
Network science is the field that gives us an understanding of
“emergence”. If we understand emergence we will know how the
different fields – physics, chemistry, biology, ... – relate and transform
into each other, which is necessary to reconstruct a picture of reality.
Connecting those fields means finding the relations between the
symbols they use – again, that same language returns. This under-
standing can be universal: as mentioned before, it is about quality
relationship. To recall the Frinston example from the very beginning
of this introduction: his theory about how brains work is purely
based on axioms about how information is aggregated in each node
given its neighbors. This is independent of what the nodes actually
are, as long as the conditions hold. You can use his network theory of
intelligence to describe not only how individual brains learn, but how
collectives made of brains learn.
It should now be clear why I consider network science important,
and a truly network science book necessary: it is our best shot at
building a collective understanding of all human knowledge, and
such attempt needs to be approached with the proper humility of
those who are not expert in anything else but gluing together the
pieces created by the real experts.
That said, I don’t want to oversell the importance of network sci-
ence. If what I said is really true, it means that we can represent any
– or at least most – aspects of reality as mathematical symbols and
we can manipulate them with the mathematical tools of computer
and network science. Which means that a complete understanding
of them is necessarily out of reach. Not just because, as Poincaré
would put it, “the head of the scientist, which is only a corner of the
16 the atlas for the aspiring network scientist
40
Henri Poincaré. The foundations of
universe, could never contain the universe entire”40 . Rather, because science. 1913
Gödel taught us that there is a strong bound of what is tractable 41
Kurt Gödel. Über formal unentschei-
in a formal system41 . At some point, even when you represent the dbare sätze der principia mathematica
entirety of reality as interconnected mathematical symbols, you will und verwandter systeme i. Monatshefte
für mathematik und physik, 38(1):173–198,
need to jump out and look at the loops from the outside42 . No book
1931
can really give you a scientific road map on how to do so. 42
Douglas R Hofstadter. Gödel, Escher,
Bach. Harvester press Hassocks, Sussex,
1979
1.5 What is in This Book?
This is all fine and dandy but, at the end of the day, what does this
book contain?
At a general level, it contains the widest possible span of all that
is related to network science that I know. It is the result of twelve
years of experience that I poured on the field. Virtually any concept
that I used or that I simply came to know in these twelve years is
represented in at least a sentence in this book.
As you might expect, this is a lot to include and would not fit a
book, not even a 650+ pages like this one. By necessity, many – if
not all – of the topics included in this book are treated relatively
superficially. I would not say that this book would provide you what
you need to know to be a network scientist. But it would point you 43
Connecting the symbols of network
to what you need to know43 . To borrow from Rumsfeld44 : the book science, maybe?
provides little to no known knowns, but it will provide you with all the 44
[Link]
known unknowns in network science – so that your unknown unknowns Transcripts/[Link]?
TranscriptID=2636
are aligned with those of everyone else. After internalizing this book,
you will know what you don’t know; you will be handed all the tools
you need to ask meaningful questions about network science in 2021.
You can go to the other books or to any other article, and find the
answers.
That is why I decided to call this book an “Atlas”. It is the map
you need to set foot among networks and start exploring. An atlas
doesn’t do the exploration for you, but you can’t explore without an
atlas. This is the book I wished I had twelve years ago.
At a more specific level, the book is divided in thirteen parts.
Part I is about setting the stage for network analysis. It starts with
a quick recap of the main concepts we borrow daily in network
analysis from probability theory. Then it teaches you what a graph
is and how many features to the simple mathematical model were
added over the years, to empower our symbols to tame more and
more complex real world entities. Finally, it pivots perspectives to
show an alternative way of manipulating networks, via matrices
and linear algebra.
Part III uses some of the tools presented in the previous part to build
slightly more advanced analyses. Specifically, it focuses on the
question: which nodes are playing which role in the network?
And: can we say that a node is more important than another? If
you want to answer these questions, you need to relate the entire
network structure to a node, i.e. to use fully what Part II trained
you to do.
Part IV teaches you the main approaches for the creation of synthetic
network data. It explores the main reasons why we want to do
it. Sometimes, it is because we need to test an algorithm and we
need a benchmark. Alternatively, we can use these models to
reproduce the properties of real world networks we investigated in
the previous parts, to see whether we understand the mechanisms
that make them emerge.
Part VIII opens the Pandora’s Box of the level of analysis that is
the most interesting and probably the one with which you will
struggle most of the time: the mesoscale. The mesoscale is what
lies between local node properties and global network statistics.
This includes – but is not limited to – questions such as: does my
network have a hierarchical structure? Is there a densely connected
core surrounded by a sparsely connected periphery? Do nodes
consider other nodes’ properties in their decision to connect to
them?
Part X takes a steep turn into the realm of computer science. It deals
with graph mining: a collection of techniques that allow you to
discover patterns in your graph structure, even if you are not sure
about what these patterns might look like or hint at. It is what we
would call “bottom-up” discovery.
Part XII includes a few tips and tricks for an aspect of network sci-
ence that is rarely covered in other books: how to browse/explore
your network data and how to communicate your results. Specifi-
cally, I will show you some best practices in visualizing networks.
I am a visual thinker and, sometimes, patterns and ideas about
those patterns emerge much more clearly when you see them,
introduction 19
1.6 Acknowledgements
Roberta Sinatra, Yong-Yeol Ahn, and Yu-Ru Lin. All these people
donated hours of their time with no real tangible reward, just to
make sure my book graduated from “incomprehensible mess” to
“almost passable and not misleading”. Thank you.
With their work, some reviewers expressed their intent to support 45
[Link]
charitable organizations. Speciphically, they mentioned TechWomen45
– to support the careers of women in STEM fields –, and Evidence 46
[Link]
Action46 – to expand our de-worming efforts and reaping the sur-
prisingly high societal payoff. You should also consider donating to
them.
If there’s any value in this book, it comes from the hard work of
these people. All the mistakes that remain here are exclusively due to
my ineptitude in properly implementing my reviewers’ valuable com-
ments. I expect there must be many of such mistakes, ranging from
trivial typos and clumsily written sentences, to more fundamental
issues of misrepresentation. If you find some, feel free to drop me an
email to mcos@[Link].
If, for some reason, you only have access to a printed version of
this book – or you found the PDF somewhere on the Internet, know 47
[Link]
that there is a companion website47 with data for the exercises, their
solutions, and – hopefully in the future – interactive visualizations.
Part I
Basics
2
Probability Theory
Frequentist Bayesian
Prior expectation
Figure 2.1: Schematics of the
mental processes used by a fre-
Experiment outcome quentist and a Bayesian when
presented with the results of an
Updating priors
experiment.
Future expectation
2.1 Notation
many values and forms, and we don’t know which of them it will be
before actually running the process. As for events, they are the focus
of all questions in probability theory: you can sum up probability
theory as the set of instruments that allow you to ask and answer
questions about events (sets of X) such as: “What is the probability
that X, the outcome of the process, is this and/or this but not that
and/or that?”
Mathematically one writes such a question as P( X ∈ S), where S
is a set of the values that X takes in our question. X is an outcome,
X ∈ S is an event. For instance, if we were asking about the event
“will the die land on an even number?”, S = {2, 4, 6}. So, P( X ∈ S)
asks what’s the probability of the “die lands on an even number”
event – or for X to take either of the 2, 4, 6 values. Note that elements
in S are all possible alternatives: if we write P( X ∈ {2, 4, 6}), we’re
asking about the probability of landing on 2 or 4 or 6. If you want to
have the probability of two events happening simultaneously, you
have to explicitly specify it with set notation: P( X ∈ {{2, 4, 6} ∩ {1}})
asks the probability of landing on an even side and on 1 at the same
time.
We also need to consider special questions. For instance, there
is the case in which no event happens: P( X ∈ ∅) (here ∅ refers to
the empty set, a set containing no elements). The converse is also
important: the probability of any event happening. In the case of
the die, there are a total of six possible outcomes. Notation-wise, we
define the set of all possible outcomes as Ω = {1, 2, 3, 4, 5, 6}. So
this is represented as P( X ∈ {1, 2, 3, 4, 5, 6}), or P( X ∈ Ω). Figure
2.2 shows how the mathematical notation corresponds to our visual
intuition.
or or
or
or or
2.2 Axioms
Ws
Figure 2.3: The baseline prob-
P(H|W) > P(H) ability of H is 0.5. When you
add feet to the coin (W) the
coin is more likely to land
P(H|W) < P(H) on the opposite side. Thus,
P( H |W ) 6= P( H ) and the two
events are not independent –
P(H) = 0.5
P(H|W) = P(H) unless you add feet on both
sides as in the bottom example.
P ( H |W ) P (W )
P (W | H ) = .
P( H )
P( ) P( | )
P( ) P( | ) = P( ) P( | ) P( | ) =
P( )
I already told you that I’m a pretty good coin rigger (P( H |W ) =
0.9). For the sake of the argument, let’s assume I’m a very honest
person: the probability I cheat is fairly low (P(W ) = 0.3).
Now, what’s the probability of landing on heads (P( H ))? P( H ) is
trickier than it appears, because we’re in a world where people might
cheat. Thus we can’t be naive and saying P( H ) = 0.5. P( H ) is 0.5
if rigging coins is impossible. It’s more correct to say P( H | − W ) =
0.5: a non rigged coin (if W didn’t happen, which we refer to as
−W) is fair and lands on heads 50% of the times. The real P( H ) is
P( H | − W ) P(−W ) + P( H |W ) P(W ). In other words: the probability
of the coin landing on heads is the non rigged heads probability if I
didn’t rig it (P( H | − W ) P(−W )) plus the rigged heads probability if I
rigged it (P( H |W ) P(W )).
The probability of not cheating P(−W ) is equal to 1 − P(W ). This
is because cheating and non cheating are mutually exclusive and
either of the two must happen. Thus we have Ω = {W, −W }. Since
P(Ω) = 1 and P(W ) = 0.3, the only way for P(W, −W ) to be equal to
1 is if P(−W ) = 0.7.
This leads us to: P( H ) = P( H | − W ) P(−W ) + P( H |W ) P(W ) =
0.5 × 0.7 + 0.9 × 0.3 = 0.62. Shocking.
The aim of Bayes’ theorem is to update your prior about me
cheating (P(W )) given that, suspiciously, the toss went in my favor
(P(W ) → P(W | H )). Plugging in the numbers in the formula:
0.9 × 0.3
P (W | H ) = = 0.43.
0.62
A couple of interesting things happened here. First, since the event
went in my favor, your prior about me possibly cheating got updated.
Specifically, the event became more likely: from 0.3 to 0.43. Second,
even if my success probability after cheating is very high, it is still
more likely that I didn’t cheat, because your prior about my lack of
integrity was low to begin with.
This second aspect is absolutely crucial and it’s easy to get it
wrong in everyday reasoning. The textbook example is the cancer
diagnosing machine. Let’s say that 0.1% of people develop a cancer,
and we have this fantastic diagnostic machine with an accuracy of
99.9%: the vast majority of people will be diagnosed correctly (posi-
tive result for people with cancer and negative for people without).
You test yourself and the test is positive. What’s your chance of hav-
ing cancer? 99.9% accuracy is pretty damning, but before working on
your last will, you apply Bayes’ Theorem:
0.999 × 0.001
P(C |+) = = 0.5.
0.999 × 0.001 + 0.001 × 0.999
The probability you have cancer is not 99.9%: it’s a coin toss! (Still
probability theory 29
5
Of course, in the real world, if you
bad, but not that bad).5 took the test it means you thought you
The real world is a large and scary environment. Many different might have cancer. Thus you were not
drawn randomly from the population,
things can alter your priors and have different effects on differ- meaning that you have a higher prior
ent events. The way a Bayesian models the world is by means of a that you had cancer. Therefore, the test
Bayesian network: a special type of network connecting events that is more likely right than not. Bayes’
theorem doesn’t endorse carelessness
influence each other. Exploring a Bayesian network allows you to when receiving a bad news from a very
make your inferences by moving from event to event. I talk more accurate medical test.
about Bayesian networks in Section 4.6.
2.5 Stochasticity
8
Specifically, it is a right stochastic
matrix8 .
Figure 2.6 is a stochastic The rows tell you your current matrix: the rows sum to one, although
there’s a bit of rounding going on. In a
state and the columns tell you your next state. If you are in the first
left stochastic matrix, the columns sum
row, you have a 30% probability of remaining in that state (the value to one.
of the cell in the first row and first column is 0.3). You have a 20%
probability of transitioning to state two (first row, second column), 8%
probability of transitioning to state three, and so on.
A bit more formally, let’s assume you indicate your state at time
t with Xt . You want to know the probability of this state to be a
specific one, let’s say x. x could be the id of the node you visit at the
t-th step of your random walk. If your process is a Markov process,
the only thing you need to know is the value of Xt−1 – i.e. the id of
the node you visited at t − 1. In other words, the probability of Xt = x
is P( Xt = x | Xt−1 = xt−1 ). Note how Xt−2 , Xt−3 , ..., X1 aren’t part of
this estimation. You don’t need to know them: all you care about is
X t −1 .
On the other hand, a non-Markov process is a process for which
knowing the current state doesn’t tell you anything about the next
possible transitions. For instance, a coin toss is a non-Markov process.
The fact that you toss the coin and it lands on heads tells you nothing
about the result of the next toss – under the absolute certainty that
the coin is fair. The probability of Xt = x is simply P( Xt = x ): there’s
no information you can gather from your previous state.
Finally, we have higher-order Markov processes. Higher-order
means that the Markov process now has a memory. A Markov pro-
cess of order 2 can remember one step further in the past. This means
that, now, P( Xt = x | Xt−1 = xt−1 , Xt−2 = xt−2 ): to know the probabil-
ity of Xt = x, you need to know the state value of Xt−2 as well as of
Xt−1 . More generally, P( Xt = x | Xt−1 = xt−1 , Xt−2 = xt−2 , ..., Xt−m =
xt−m ), with m ≤ t.
The classical network examples of a higher order Markov pro-
cess is the non-backtracking random walk (Figure 2.8). In a non-
backtracking random walk, once you move from node u to node
v, you are forbidden to move back from v to u. This means that,
once you are in v, you also have to remember that you came from u.
Higher order Markov processes are the bread and butter of higher
order network problems, which is the topic of Chapter 30.
32 the atlas for the aspiring network scientist
No backtracking!
Outcome i
Sample Space
There are two possible cases in your sample space: either it con-
tains discrete finite outcomes, or it contains effectively infinite contin-
uous ones. The first case is, for instance, a coin toss. There are only
two possible outcomes: heads or tails. The second case is when, for
instance, you’re measuring something that can take any real value as
an outcome. In the first discrete case, we call the probability distri-
bution a “probability mass function”. In the second case, we call it a
“probability density function”.
There are a few common probability distributions you should be
familiar with – Figure 2.10 shows some stylized representations of
each of them –:
probability theory 33
a shorter observation interval usually “cut off” the left side of the
distribution. Interestingly, many examples commonly mentioned
for explaining a Poisson distribution (number of admittances in a
hospital in an hour, number of email written in an hour, and so on)
aren’t actually Poisson distributions, because people making those 9
Albert-Laszlo Barabasi. The origin
examples fail to account for the burstiness of human behavior9 . of bursts and heavy tails in human
dynamics. Nature, 435(7039):207, 2005
• Hypergeometric: this is yet another discrete probability function.
It is very similar to a binomial distribution. If the binomial de-
scribed the success odds in an extraction-with-replacement urn
game, the hypergeometric describes the more common case of
extraction-without-replacement: when you extract a ball from the
urn, you don’t put it back. It is mathematically less tractable, but
much more useful. This is used especially for the task of network
backboning (Chapter 24).
• Exponential: the exponential distribution is the continuous version
of the geometric distribution. The geometric distribution tells you
the probability that the first success of an experiment happens
at trial n. Each experiment is independent and the probability
of a success will determine how steep the distribution is. One
cool property of the exponential/geometric distribution is that is
doesn’t “age”: it doesn’t matter how many trials you did so far –
the likelihood of a success doesn’t change. The classical example
of an “ageless” process is atomic decay: the half life of carbon14 is
the same regardless for how long it had been decaying. To go from
2kg to 1kg takes the exact same amount of time as going from 1kg
to 500 grams.
• Power law: a power law can be both a discrete or a continuous
distribution. It describes the relationship between two quantities,
where a relative change in one corresponds to a proportional
relative change in the other (so the second variable changes as a
power of the first). An example of discrete power law is Zipf’s
10
Mark EJ Newman. Power laws,
law10 . We’ll see more than you want to know about power laws
pareto distributions and zipf’s law.
when talking about fitting degree distributions in Section 6.3. Contemporary physics, 46(5):323–351,
2005b
• Lognormal: a lognormal distribution is the distribution of a con-
tinuous random variable whose logarithm follows a normal dis-
tribution – meaning the logarithm of the random variable, not of
the distribution. This is the typical distribution resulting from the
multiplication of two independent random positive variables. If
you throw a dozen 20-sided dice and multiply the values of their
faces up, you’d get a lognormal distribution. It’s very tricky to tell
this distribution apart from a power law, as we’ll see.
Sometimes, rather than looking at the probability mass/density
function, it’s more useful to look at their cumulative versions. In
probability theory 35
X X
Figure 2.12: A simple example
=0 0 to understand information
9 bits entropy. From left to right: the
0
vector x has six elements tak-
6 values
0 ing three different values. We
!
pij
MIxy = ∑ ∑ pij log pi p j
,
j∈y i ∈ x
where pij is the joint probability of i and j. Even if I don’t give you
the full explanation, you can hopefully see what’s going on here. The
meat is comparing the joint probability of i and j happening with
what you would expect if i and j were completely independent. If
they are, then pij = pi p j , which means we take the logarithm of one,
which is zero and everything collapses into zero if that’s always the
case. Any time the happening of i and j is not independent, we add
something to the mutual information. That something is the number
of bits we save.
2.9 Summary
1. Probability theory gives you the tools to make inferences about un-
certain events. We often use a frequentist approach, the idea that
an event’s probability is approximated by the aggregate past tests
of that event. Another important approach is the Bayesian one,
which introduces the concept of priors: additional information that
you should use to adjust your inferences.
2.10 Exercises
1. Suppose you’re tossing two coins at the same time. They’re loaded
in different ways, according to the table below. Calculate the
probability of getting all possible outcomes:
txt. Follow the process for three steps and reconstruct the correct
answer. (Note, this is a Caesar cipher12 with shift 7 applied three
times, because the Caesar cipher is a Markov process).
Outcome p
1 0.1
2 0.15
3 0.2
4 0.21
5 0.17
6 0.09
7 0.06
8 0.02
5. How many bits do we need to independently encode v1 and v2
from [Link]
How much would we save in encoding v1 if we knew v2 ?
3
Basic Graphs
Every story should start from the beginning and, in this case, in 1
John Adrian Bondy, Uppaluri Siva Ra-
the beginning was the graph1 , 2 , 3 , 4 . To explain and decompose the machandra Murty, et al. Graph theory
elements of a graph, I’m going to use the recurrent example of social with applications, volume 290. Citeseer,
1976
networks. The same graph can represent different networks: power 2
Douglas Brent West et al. Introduction
grids, protein interactions, financial transactions. Hopefully, you can to graph theory, volume 2. Prentice hall
effortlessly translate these examples into whatever domain you’re Upper Saddle River, 2001
going to work.
3
Reinhard Diestel. Graph theory.
Springer Publishing Company, Incorpo-
Let’s start by defining the fundamental elements of a social net- rated, 2018
work. In society, the fundamental starting point is you. The person. 4
Jonathan L Gross and Jay Yellen. Graph
Following Euler’s logic that I discussed in the introduction, we want theory and its applications. CRC press,
2005
to strip out the internal structure of the person to get to a node. It’s
like a point in geometry: it’s the fundamental concept, one that you
cannot divide up into any sub-parts. Each person in a social network
is a node – or vertex; in the book I’ll treat these two terms as syn-
onyms. We can also call nodes “actors” because they are the ones
interacting and making events happen – or “entities” because some-
times they are not actors: rather than making things happen, things
happen to them. “Actor” is a more specific term which is not an
exact synonym of “node”, but we’ll see the difference between the 5
The understatement of the century.
two once we complicate our network model just a bit5 , in Section 4.2.
To add some notation, we usually refer to a graph as G. V indi-
cates the set of G’s vertices. Since V is the set of nodes, to refer to the
number of nodes of a graph we use |V | – some books will use n, but
I’ll try to avoid it. Throughout the book, I’ll tend to use u and v to
indicate single nodes.
So far, so good. However, you cannot have a society with only one
individual. You need more than one. And, once you have at least two
people, you need interactions between them. Again, following Euler,
for now we forget about everything that happens in the internal struc-
ture of the communication: we only remember that an interaction is
basic graphs 41
taking place. We will have plenty of time to make this model more
complicated. The most common terms used to talk about interactions
are “edge”, “link”, “connection” or “arc”. While some texts use them
with specific distinctions, for me they are going to be synonyms, and
my preferred term will always be “edge”. I think it’s clearer if you
always are explicit when you refer to special cases: sure, you can
decide that “arc” means “directed edge”, but the explicit formula “di-
rected edge” is always better than remembering an additional term,
because it contains all the information you need. (What the hell are
“directed edges”? Patience, everything will be clear)
Again, notation. E indicates the set of G’s edges and | E| is the
number of edges – some books will use m as a synonym for | E|.
Usually, when talking about a specific edge one will use the notation
(u, v), because edges are pairs of nodes – unless we complicate the
graph model. Now we have a way to refer to the simplest possible
graph model: G = (V, E), with E ⊆ V × V. A graph is a set of nodes
and a set of edges – i.e. node pairs – established among those nodes.
(a) (b)
(c)
We’re going to talk about how to visualize networks much later
in Part XII, but it’s better to introduce some visual elements now,
otherwise how are we supposed to have figures before then? Nodes
are usually represented as dots, or circles – Figure 3.1(a). Edges are
lines connecting the dots – Figure 3.1(b). When all you have is nodes
and edges, then you have a simple graph – Figure 3.1(c). Note that
these visual elements are basic and widely used, but they are by no
means the only way to visualize nodes and edges. In fact, when you
want to convey a message about a network of non-trivial size, they’re
usually not a great idea.
The first famous graph in history is Euler’s Königsberg graph,
which I show in Figure 3.2. In the graph, each node represents a land-
mass and each edge represents a bridge connecting two landmasses.
Since there were multiple bridges connecting the same landmasses,
we have multiple edges between the same two nodes. This seemingly
trivial fact is actually rather interesting.
“Simple graph” means literally simple: nothing more than nodes
and edges – no attributes, no possibility of having multiple connec-
tions between the same two nodes. If you add any special feature,
it’s not a simple graph any more. Under this light, we discover that
Euler’s first graph wasn’t simple after all. It allowed for parallel
42 the atlas for the aspiring network scientist
edges: multiple edges between the same two nodes. Euler’s first
graph was a multigraph. That’s so non-standard that we’re not even
going to talk about it in this chapter: you’ll have to wait for the next
one, specifically for Section 4.2.
In our simple graph we also assume there are no self loops, which
are edges connecting a node with itself. Our assumption is that we
aren’t psychopaths: everybody is friend with themselves, so we don’t
need to keep track of those connections.
(a) (b)
3-5
3 Figure 3.4: (a) A graph. (b) Its
4-5 linegraph version.
2 4 3-4
5 2-4
1-4
1
1-2
(a)
(b)
line graph. We’ll see how you can use line graphs to represent high
order relationships in Chapter 30, to find overlapping communities in
Chapter 34, and to estimate similarities between networks in Chapter
41.
(a)
(b)
In a message passing game, (u, v) – or u → v – means that node
u can pass a message to node v, but v cannot send it back to u. Di-
rected graphs introduce all sorts of intricacies when it comes to
finding paths in the network, a topic we’re going to dissect in Chap-
44 the atlas for the aspiring network scientist
looking for the shortest path (see Chapter 10) in a road network, your
edge weight could mean different things. It could be a distance if it
represents the length of the trait of road: longer traits will take more
time to cross. Or it can be a proximity: it could be the throughput of
the trait of road in number of cars per minute that can pass through
it – or the number of lanes. If the weight is a distance, the shortest
path should avoid high edge weights. If the weight is a proximity, it
should do its best to include them.
To sum up, “proximity” means that a high weight makes the
nodes closer together; e.g. they interact a lot, the edge has a high
capacity. “Distance” means that a high weight makes the nodes
further apart; e.g. it’s harder or costly to make the nodes interact.
Edge weights don’t have to be positive. Nobody says nodes should
be friends! Examples of negative edge weights can be resistances in
electric circuits or genes downregulating other genes. This observa-
tion is the beginning of a slippery slope towards signed networks,
which is a topic for another time (namely, for Section 4.2, if you want
to jump there).
The network in Figure 3.6 has nice integer weights. In this case,
the edge weights are akin to counts. For instance, in a phone call
network, it could be the number of times two people have called each
other. Unfortunately, not all weighted networks look as neat as the
example in Figure 3.6. In fact, most of the weighted networks you
might work with will have continuous edge weights. In that case,
many assumptions you can make for count weights won’t apply – for
instance when filtering connections, as we will see in Chapter 24.
By far, the most common case is the one of correlation networks.
In these networks, the nodes aren’t really interacting directly with
one another. Instead, we are connecting nodes because they are
similar to each other, for some definition of similarity. For instance, 11
Boris C Bernhardt, Zhang Chen,
we could connect brain areas via cortical thickness correlations11 , or Yong He, Alan C Evans, and Neda
currencies according to their exchange rate12 , or correlating the taxa Bernasconi. Graph-theoretical analysis
reveals disrupted small-world organi-
presence in different biological communities13 . zation of cortical thickness correlation
These cases have more or less the same structure. I provide an ex- networks in temporal lobe epilepsy.
Cerebral cortex, 21(9):2147–2157, 2011
ample in Figure 3.7. In this case, nodes are numerical vectors, which 12
Takayuki Mizuno, Hideki Takayasu,
could represent a set of attributes, for instance. We calculate a corre- and Misako Takayasu. Correlation
lation between the vectors, or some sort of attribute similarity – for networks among currencies. Physica A:
Statistical Mechanics and its Applications,
instance mutual information (Section 2.8). We then obtain continuous
364:336–342, 2006
weights, which typically span from −1 to 1. And, since every pair of 13
Jonathan Friedman and Eric J Alm.
nodes have a similarity (because any two vectors can be correlated, Inferring correlation networks from
genomic survey data. PLoS computational
minus extremely rare degenerate cases), every node is connected to
biology, 8(9):e1002687, 2012
every other node. So, when working with similarity networks, you
will have to filter your connections somehow, a process we call “net-
work backboning” which is far less trivial that it might sound. We
46 the atlas for the aspiring network scientist
0.36
B
for correlation networks: (left to
right) from nodes represented
C
C
as some sort of vectors, to a
graph with a similarity measure
as edge weigth.
3.4 Summary
3.5 Exercises
2. Mr. A considers Ms. B a friend, but she doesn’t like him back. She
has a reciprocal friendship with both C and D, but only C con-
siders D a friend. D has also sent friend requests to E, F, G, and
H but, so far, only G replied. G also has a reciprocal relationship
with A. Draw the corresponding directed graph.
The world of simple graphs is... well... simple. The only thing com-
plicating it a bit so far was adding some information on the edges:
whether they are asymmetric – meaning (u, v) 6= (v, u) – and whether
they are strong or weak. Unfortunately, that’s not enough to deal
with everything reality can throw your way. In this chapter, we
present even more graph models, which go beyond the simple addi-
tion of edge information.
(b)
(a)
Many-to-Many
into the different layers to which it belongs. In this case, your identity
includes multiple personas: you are the union of the “Facebook
you”, the “Linkedin you”, the “Twitter you”. Figure 4.4(a) shows
the visual representation of this model: each layer is a slice of the
network. There are two types of edges: the intra-layer connections –
the traditional type: we’re friends on Facebook, Linkedin, Twitter –,
and the inter-layer connections. The inter-layer edges run between
layers, and their function is to establish that the two nodes in the
different layers are really the same node: they are coupled to – or
dependent on – each other.
Aspects
Do you think we can’t make this even more complicated? Think
again. These aren’t called “complex networks” by accident. To fully
generalize multilayer networks, adding the many-to-many interlayer
coupling edges is not enough. To see why that’s the case, consider
the fact that, up to this point, I considered the layers in a multilayer
network as interchangeable. Sure, they represent different relation-
ships – Facebook friendship rather than Twitter following – but they
are fundamentally of the same type. That’s not necessarily the case:
the network can have multiple aspects.
For instance, consider time. We might not be Facebook friends
now, but that might change in the future. So we can have our mul-
tilayer network at time t and at time t + 1. These are two aspects of
the same network. All the layers are present in both aspects and the
edges inside them change. Another classical example is a scientific
community. People at a conference interact in different ways – by
attending each other talks, by chatting, or exchanging business cards
– and can do all of those things at different conferences. The type of
interaction is one aspect of the network, the conference in which it
happens is another.
I can’t hope to give you here an overview of how many new things
this introduces to graph theory. So I’m referring you to a specialized 22
Ginestra Bianconi. Multilayer Networks:
book on the subject22 . Structure and Function. Oxford University
Press, 2018
Signed Networks
Signed networks are a particular case of multilayer networks. Sup-
pose you want to buy a computer, and you go online to read some
reviews. Suppose that you do this often, so you can recognize the
reviewers from past reviews you read from them. This means that
you might realize you do not trust some of them and you trust others.
This information is embedded in the edges of a signed network: there
are positive and negative relationships.
Signed networks are not necessarily restricted to either a single
positive or a single negative relationship – e.g. “I trust this person”
or “I don’t trust this person”. For instance, in an online game, you
can have multiple positive relationships like being friend or trading
together; and multiple reasons to have a negative relationship, like
fighting each other, or putting a bounty on each other heads.
A key concept in signed networks is the one of structural balance.
Since this is mostly related to the link prediction problem, I expand
on this in Section 21.1.
Positive and negative relationships have different dynamics. For
instance, in a seminal study looking at interactions between players
extended graphs 55
23
Michael Szell, Renaud Lambiotte,
in a massively multiplayer online game23 , the authors studied the dif- and Stefan Thurner. Multirelational
ferent degree distributions (Section 6.2) for each type of relationship. organization of large-scale social
networks in an online world. Proceedings
They uncovered that positive relationships have a marked exponen-
of the National Academy of Sciences, 107
tial cutoff, while negative relationships don’t. You’ll become more (31):13636–13641, 2010
accustomed to what a degree distribution is and all the lingo related
to it in Chapter 6. For now, the meaning of what I just said is: there is
a limit to the number of people you can be friends with, but there is
no limit to the number of people that can be mad at you.
4.3 Hypergraphs
(a) (b)
In the classical definition, an edge connects two nodes – the gray
lines in Figure 4.7(a). Your friendship relation involves you and your
friend. If you have a second friend, that is a different relationship.
There are some cases in which connections bind together multiple
people at the same time. For instance, consider team building: when
you do your final project with some of your classmates, the same
relationship connects you with all of them. When we allow the
same edge to connect more than two nodes we call it a hyperedge –
the gray area in Figure 4.7(b). A collection of hyperedges makes a
24
Vitaly Ivanovich Voloshin. Introduction
hypergraph24 , 25 .
to graph and hypergraph theory. Nova
Graphs with simplicial complexes26 , 27 are related to hypergraphs. Science Publishers Hauppauge, 2009
The difference between the two is that simplicial complexes have a 25
Alain Bretto. Hypergraph theory: An
strong emphasis on geometry. Simplicial complex analysis specializes introduction. Mathematical Engineering.
Cham: Springer, 2013
in systems with many-to-many interactions that are embedded in 26
Vsevolod Salnikov, Daniele Cassese,
real physical spaces. For instance, you can use simplicial complexes and Renaud Lambiotte. Simplicial
to study groups of people interacting at a conference, because social complexes and complex systems.
European Journal of Physics, 40(1):014001,
groups will form in the two dimensional floor of the conference 2018
building. 27
Jakob Jonsson. Simplicial complexes of
graphs, volume 3. Springer, 2008
To make them more manageable, we can put constraints to hy-
peredges. We could force them to always contain the same number
of nodes. In a soccer tournament, the hyperedge representing a
team can only have eleven members: not one more nor one less, be-
28
Shenglong Hu and Liqun Qi. Alge-
cause that’s the number of players in the team. In this case, we call
braic connectivity of an even uniform
the resulting structure a “uniform hypergraph”, and have all sorts hypergraph. Journal of Combinatorial
of interesting properties28 . In general, when simply talking about Optimization, 24(4):564–579, 2012
56 the atlas for the aspiring network scientist
is special: its elements are not forced to be tuples any more. They
can be triples, quartuplets, and so on. For instance, (u, v, z) is a legal
element that can be in E, with u, v, z ∈ V.
Time
A
Figure 4.10: An example of dy-
namic edge information. Time
B flows from left to right. Each
A row represents a possible po-
tential edge between nodes A,
C
B, and C. The moments in time
B
in which each edge is active are
C represented by gray bars.
Time Time
A A
B B
A A
C C
B B
C C
A C A C A C A C A C A C
B B B B B B
B B
A A
C C
B B
C C
A C A C A C A C A C A C
B B B B B B
B 3
Figure 4.12: (a) A network with
A 3
qualitative node attributes,
represented by node labels
A 4
A 4 and colors. (b) A network with
quantitative node attributes,
represented by node labels and
A 5 sizes.
B 4
B 3
A 4
B 2
(a) (b)
are the various countries. They connect together if one country ex-
ports goods to another. We can have multiple quantitative attributes
on each country. For instance, it can be its GDP per capita, its pop-
ulation, its total trade volume. On the other hand, we can also put
countries in different categories: in which world region are they lo-
cated? Are they democracies or not? Of which trade agreement are
they part of?
In this case, our graph changes form again: G = (V, E, A). We
can see each v ∈ V not as a simple entity, but as a vector of attribute
values: v = ( a1 , a2 , a3 , ...). In this representation, a1 is the value for v
of the first attribute in A. a1 can be a real, integer, or a category.
Node attributes are important because nodes might have tenden-
cies of connecting – or refusing to connect – to nodes with similar
attribute values. We’ll explore this topic in the forms of “homophily”
in Chapter 26 for qualitative attributes, and “assortativity” in Chapter
27 for quantitative attributes. This is different from bipartite networks
because in bipartite networks edges between nodes with different
attribute values are forbidden, while in these cases edges are simply
correlated with attribute values. Moreover, bipartite networks are only
defined for qualitative attributes, not quantitative.
To wrap up, no one forces you to use a single of these more com-
plex graph models at a time. You can merge them together to fit your
analytical needs. For instance, you can create this monster graph
type: Gn = (V1 , V2 , E, L, W, A): a bipartite graph with V1 and V2
nodes, each with attributes in A, which is weighted (W) multilayer
with | L| layers and – for good measure – is also a hypergraph, allow-
ing edges in E with more than two nodes. And, of course, you can
observe it at multiple time intervals (G1 , G2 , ...). Yikes.
extended graphs 61
Now that you know more about the various features of different net-
work models, we can start looking at different types of networks. I’m
going to use a taxonomy for this section. I find this way of organizing
networks useful to think about the objects I work with.
Simple Networks
The first important distinction between network types is between
simple and complex networks. A simple network is a network we
can fully describe analytically. Its topological features are exact and
trivial. You can have a simple formula that tells you everything you
need to know about it. In complex networks that is not possible, you
can only use formulas to approximate their salient characteristics.
The difference between a simple network and a complex network
is the same between a sphere and a human being. You can fully
describe the shape of a sphere with a few formulas: its surface is
4
4πr2 , its volume is πr3 . If you know r you know everything you
3
need to know about the sphere. Try to fully describe the shape of
a human being, internal organs included, starting from a single
number. Go on, I have time.
(a) (b)
What do simple networks look like? I think the easiest example
conceivable is a square lattice. This is a regular grid, in which each
node is connected to its four nearest neighbors. Such lattice can either
span indefinitely (Figure 4.13(a)), or it can have a boundary (Figure
4.13(b)). Their fundamental properties are more or less the same.
Knowing this connection rule that I just stated allows you to picture
any lattice ever. That is why this is a simple topology.
Regular lattices can come in many different shapes besides square,
for instance triangular (Figure 4.14(a)) or hexagonal (Figure 4.14(b)).
They also don’t necessarily have to be two dimensional as the exam-
ples I made so far: you can have 1D (Figure 4.14(c)) and 3D (Figure
4.14(d)) lattices – the latter might be a bit hard to see, but it is a cube
of with four nodes per side.
62 the atlas for the aspiring network scientist
Complex Networks
If simple networks were the only game in town, this book would not
exist. That is because, as I said, you can easily understand all their
properties from relatively simple math. That is not the case when the
network you’re analyzing is a complex network. Complex networks
model complex systems: systems that cannot be fully understood if
all you have is a perfect description of all their parts. The interactions
between the parts let global properties emerge that are not the simple
extended graphs 63
sum of local properties. Thus, there isn’t a simple wiring rule and,
even knowing all the wiring, some properties can still take you by
surprise.
Personally, I like to divide complex networks into two further
categories: complex network with fundamental metadata and with-
out fundamental metadata. As we saw so far, there are a number of
metadata you can attach to your nodes and network. You can have
quantitative and qualitative node/edge attributes, layers, bipartite
networks, and so on. The difference between the two types is that,
if the metadata are fundamental, they change the way you interpret
some or all the metadata themselves.
For instance, social networks, infrastructure networks, biological
networks, and so on, model different systems and have different
metadata attached to their nodes and edges. It can be age/gender,
activation types, up- and down-regulation. However, at a funda-
mental level, the algorithms and the analyses you perform on them
are the same, regardless of what the networks represent. They have
nodes and edges and you treat them as such. You perform the Euler
operation: you forget about all that is unnecessary so you can apply
standardized analytic steps.
That is emphatically not true for networks with fundamental
metadata. In that case, you need to be aware of what the metadata
represent, because they change the way you perform the analysis and
you interpret the results. A few examples:
weighted networks. For instance, edges with very low weights are
important here, because a strong negative correlation is interesting,
even if its value (−1) is lower than no correlation at all (0).
Rain
Figure 4.16: (a) A Bayesian
T F
network. (b) The conditional
0.2 0.8
probability tables for the node
states. The tables are referring
Raining Sprinkler
to, from top to bottom: Rain,
Rain T F
Sprinkler, Wet.
T 0.01 0.99
Sprinklers F 0.2 0.8
Wet
Wet Rain Sprinkler T F
T T 0.99 0.01
(a) T F 0.98 0.02
F T 0.97 0.03
F F 0.01 0.99
(b)
applying some of the techniques you will learn later on. For instance,
you might discover set of variables that are independent of each
other, even if, at first glance, it might be difficult to tell.
A not so distant relative of Bayesian networks are neural networks,
the bread and butter of machine learning these days. Notwithstand-
ing their amazing – and, sometimes, mysterious – power, neural
networks are actually much more similar to simple networks than to
complex ones. Differently from Bayesian networks, the wiring rules
of neural networks – of which I show some examples in Figure 4.17 –
are usually rather easy to understand.
The way they work is that the weight on each node of the out-
put layer is the answer the model is giving. This weight is directly
dependent on a combination of the weights of the nodes in the last
hidden layer. The contribution of each hidden node is proportional to
the weight of the edge connecting it to the output node. Recursively,
the status of each node in the hidden layer is a combination of all
its incoming connections – combining the edge weight to the node
weight at the origin. The first hidden layer will be directly dependent
on the weights of the nodes in the input layer, which are, in turn,
determined by the data.
What the model does is simply finding the combination of edge
weights causing the output layer’s node weights to maximize the
desired quality function.
4.7 Summary
1. Bipartite networks are networks with two node types. Edges can
only connect two nodes of different types. You can generalize
them to be n-partite, and have n node types.
4.8 Exercises
Graphs, with their fancy nodes and edges, are not the only way to
represent a network. One can do so also by using matrices. In fact,
ask some people and they will tell you that everything is a matrix.
What’s a number if not a zero-dimensional tensor? I mean, come on!
Unfortunately, I am not one of those people, so this chapter will
contain only the bare minimum for you to smile and nod while
talking to them.
The reason of having this chapter is because sometimes operations
are more natural to understand with the graph models, and some-
times they are just matrix operations. Which perspective is more
useful – graph vs matrix – often depends on the perspective used by
the researcher(s) discovering a given property of developing a given
tool. So in the book I’ll often switch back and forth between these
two representations, and this chapter is your map not to get lost once
I start rambling about “positive semi-definite matrices”, whatever the
hell that means.
We start with the simplest object: the adjacency matrix (Section
5.1). In Section 5.2 we will see what kinds of operations you can do
on them, and then in Section 5.3 some special matrix representations
for graphs, namely the stochastic, the incidence, and the graph
Laplacian matrices. Finally, Section 5.4 shows more advanced matrix
operations, specifically how to decompose complex matrices in
smaller, simpler, and more informative objects.
Yes! Yes!
Do you 5
know each
other?
2 Figure 5.1: (a) A vignette of
7
1 2 3 4 5
9 1 how one would construct an ad-
1
2
x
1 5
3
6 jacency matrix. (b) An example
4 8
5x graph. (c) The adjacency matrix
4 3 of (b). Rows and columns are in
(a) (b) (c) the same order as the node ids
(so the first row/column refers
to node 1, the second to node 2,
The adjacency matrix is the basic representation of a graph as etc).
a matrix. Each row/column corresponds to a node. Each cell rep-
resents an edge, set to one if the edge exists, and zero otherwise.
If the graph is undirected, each edge sets two cells to one. If the
edge connects nodes u and v both the Auv and the Avu entries are
equal to one. Figure 5.1(b) shows a graph and Figure 5.1(c) shows
its adjacency matrix – in the graph view I labeled the nodes with the
order as they appear in the adjacency matrix: the first row/column
represents node 1, the second row/column is for node 2, and so on.
In Figures 5.1(b) and 5.1(c) we have the simplest graph possible:
the unweighted undirected graph. In this case, the adjacency matrix
carries a few properties. For instance, the graph has no self-loops
– edges connecting a node to itself. For this reason, the diagonal
of the adjacency matrix contains zeros. We like to keep it that way,
because we’ll use the diagonal for all sorts of interesting stuff in the
future – for instance later on when dealing with the graph Laplacian.
The adjacency matrix is also square, meaning that it has the same
number of rows and columns. Moreover, it is symmetric, meaning
that ∀u, v Auv = Avu . The diagonal divides the matrix into two
identical triangular halves.
You can calculate the complement of any graph by simply calcu-
lating 1 − A, with 1 being a matrix full of ones – although you might
want to fill its diagonal with zeros to avoid self loops, which are
usually ignored in complement graphs.
We can adapt the adjacency matrix to deal with all the compli-
8
Figure 5.2: (a) A non-symmetric
2 6 adjacency matrix. (b) The corre-
sponding directed graph.
7
3
5
4 1
9
(a) (b)
70 the atlas for the aspiring network scientist
6
Figure 5.3: (a) A non-binary
5 5 adjacency matrix. (b) The corre-
4 4
3 2 sponding weighted graph.
2
3
5
9 3 7 5 2
3
5 6 5
2
4 8
2
1
(a) (b)
1 2 3 4 5 6 7 8 9 10
Figure 5.4: (a) A non-square
adjacency matrix. (b) The corre-
sponding bipartite graph.
4 5 3 6 2 1
(a) (b)
What else? We can make the adjacency matrix not square if we
need to represent a bipartite network. The different numbers of rows
and columns allow us to use one dimension to represent the nodes
in V1 and the other to represent the nodes in V2 . Figure 5.4 depicts
an example. The downside is that we lose the power of the diagonal
we had in the adjacency matrix – which doesn’t seem like a big deal
now, because at the moment I’m being all hush hush about what this
power really is.
Of course, it’s possible to have a square adjacency matrix for a
bipartite network if |V1 | = |V2 |. You can also “squarify” a bipartite
adjacency matrix by dividing it in four blocks. The blocks on the
main diagonal contain zeros, while the blocks in the other diagonal
contain the original adjacency matrix. Such a construct is a (|V1 | +
|V2 |) × (|V1 | + |V2 |) matrix, and they can be useful. Figure 5.5 shows
matrices 71
|V2| AT 0
an example.
Finally we can – and do – represent even multilayer networks
with matrices. Or, to be more precise, we use tensors to represent
them. I’m not going deep into technicalities, so I’m going to give
you a superficial view of tensors: just enough to have an intuition.
Technically speaking, a tensor is a generalized vector. A vector can be
seen as a monodimensional array: a list of values. A matrix could be
said to be a two-dimensional array. A tensor is a multidimensional
array: we can have as many dimensions as we want.
(a) (b)
Transpose
Let’s consider a matrix A, whose rows and columns are your nodes.
If we transpose A it means that all Auv becomes Avu and viceversa.
In this book, for convention, A T will be the transpose of A. Figure
5.7 shows the case of a squared non symmetric matrix transpose. In
practice, transposing is like placing a mirror on the diagonal.
(a) A (b) A T
If your matrix is squared and symmetric – an undirected unipar-
tite graph – transposing has no effect: A T = A. For directed graphs,
A T and A will be different, because of the directionality. To be really
blunt, transposing A in a directed graph means to reverse all edge
directions. In a bipartite network with V1 and V2 node types, the
shape of A will change, from being a |V1 | × |V2 | matrix to a |V2 | × |V1 |
matrices 73
matrix.
Matrix Multiplication
Formally, matrix multiplication is an operation that produces a
matrix C from two matrices A and B. You cannot multiply any two
matrices together, though: they have to have one dimension of equal
size. So, if A is an n × m matrix, it can only be multiplied by B if B is
either an m × x or an x × n matrix, with x being anything. Suppose
that B is m × x: the result of A multiplied to B will be a n × x matrix.
The common dimension “disappears”.
Understanding why this is the case is easy once you know what
matrix multiplication actually does. Each cuv entry of C is equal to
the sum of the products of all entries in the uth row of A and the vth
m
column of B. Formally: cuv = ∑ auk bkv .
k =1
p = (4, 5)
Figure 5.9: An example of
Euclidean distance in m = 2
dimensions. Note that we build
((p - q)T(p - q))1/2
the special p − q vector to have,
p2 - q 2 = 4
at its ith entry, the difference
between the ith entries of p and
q.
p1 - q 1 = 3
q = (1, 1)
(x’, y’)
(λx, λy) Figure 5.10: A graphical depic-
tion of an eigenvector.
w=Av=λv
w'=A'v
(x, y)
Stochastic
Adjacency matrices are nice, but I think most of the times you’ll see
them transformed in various ways to squeeze out all the possible
analytic juice. The simplest makeover we can give to the adjacency
matrix is to convert it into a stochastic matrix. This means that we
normalize it, dividing each entry by the sum of its corresponding row
– this means that each of its rows sums to one. If nodes u and v are
connected, and u has 5 connections, the Auv entry will be 1/5 = 0.2.
Figure 5.11 shows an example of this stochastic transformation.
5
2 Figure 5.11: (a) The original
7
9 1 graph. (b) The adjacency matrix
6 of (a). (c) The corresponding
8 stochastic version.
4 3
now write this matrix as A1 , which is the same thing as A. Let’s say
this again: A1 is the probability of all transitions for random walks of
length 1. Could it be, then, that A2 is the probability of all transitions
for random walks of length 2? And that An is the probability of all
transitions for random walks of length n? Yes, they are!
From the matrix multiplication crash course I gave you in Section
5.2 you know why: A2 ’s uv entry is, as the formula I wrote there
shows, the sum of the multiplication of probabilities of all nodes k
that are connected to both u and v, and thus can be used in a path
of length 2. Multiplying Auk to Akv means asking the probability of
going from u to k and from k to v. Summing Auk1 Ak1 v to Auk2 Ak2 v
means asking the probability of passing through k1 or k2 . See Chapter
2 for a refresher on what multiplying and summing probabilities
mean.
Incidence
Laplacian
The stochastic adjacency matrix is nice, but the real superstar when
it comes to matrix representations of networks is the Laplacian. To
know what that is, we need to introduce the concept of Degree ma-
trix D – which is a very simple animal. It is what we call a “diagonal”
matrix. A diagonal matrix is a matrix whose nonzero values are ex-
clusively on the main diagonal. The other off-diagonal entries in the
matrix are equal to zero. In D the diagonal entries are the degrees of
matrices 81
5
2 Figure 5.14: The degree matrix
7
9 1 (a) of the sample graph (b).
6
8
4 3
(a) (b)
mystical properties.
Matrix Decomposition
One of the easiest ways to perform matrix factorization is what we
call “eigendecomposition”. The adjacency matrix A of an undirected
unweighted graph can always be decomposed as A = ΦΛΦ T . Rather
than being the left and right eyes of a really pissed frowny face, Φ is
the matrix we obtain piling all eigenvectors next to each other, and Λ
is a diagonal matrix with the eigenvalues on its main diagonal and
zeros everywhere else:
λ0 ... 0
Λ=
..
0 . 0
0 ... λn .
The eigendecomposition is useful to solve a set of linear difference
equations. We are mostly interested in it as the special case of the
more general Singular Value Decomposition (SVD) – which can be
applied to any matrix, even non-square ones. In SVD, we simply
replace Φ and Λ with generic matrices. In other words, we say that
we can reconstruct A with the following operation: A = Q1 ΣQ2T . Like
Λ, also Σ is a diagonal matrix. The difference is that Σ contains the
singular values of A, rather than its eigenvalues. While there is only
one valid Σ to solve this equation – that is why it is called “singular”
– there could be multiple Q1 and Q2 matrices that you could plug
in, as long as they’re both unitary matrices. A unitary matrix Q is a
matrix whose transpose is also its inverse: QQ−1 = I = QQ T , with
I being the identity matrix. SVD is especially useful for estimating
node distances on networks (Section 40.2).
Along with eigendecomposition, the two most common and useful
matrix decomposition tools are the Principal Component Analysis
matrices 83
Day Temp (°C) Wind (km/h) Sunlight (%) Rain (mm) Snow (mm)
1 27 10 80 2 0
2 26 1.2 95 1 0
3 32 7.6 100 0 0
4 12 2.3 12 20 0
5 14 3.8 8 25 0
6 6 0.2 24 40 1
7 4 0.1 2 8 30
8 2 0.9 4 1 40
9 −1 1.1 4 0 80
Figure 5.16: A table recording
(PCA) and the Non-Negative Matrix Factorization (NMF). in a matrix the characteristics of
To understand PCA, suppose that your matrix is just a set of some days.
observations and variables. Each row of the matrix is an observation
and each column is a variable. PCA, like NMF, is used to summarize
this matrix of data. If two columns/variables are correlated it means
they contain redundant information. Thus, you’re after a way to
describe your data in such a way that each variable has no redundant
information.
Figure 5.16 shows an example: each row is a day and each col-
umn is some measurement taken in that day – the temperature,
wind speed, the millimeters of rain/snow that fell that day, etc. You
might expect that some of these variables might be correlated. For
instance, it is very difficult to have a single millimeter of snow if
the temperature is above a certain value. Rather than describing a
day by all variables, you want to describe it by its similarity with an
“archetypal” day: is this a snow day or a rain day?
components as you have variables, but usually you want much fewer
– for instance two, so you can plot the data. That is because the first
component explains the most variance in the system, the second a bit
less, and so on, until the last few components which are practically
random. Thus you want to stop collecting components after you’ve
taken the first n, setting n to your delight. In Figure 5.17, I collect the
first two – they’re there for illustrative purposes so don’t be shocked
if you realize they’re not really orthogonal.
9
8 Figure 5.18: A scatter plot with
7 the first principal component
2nd Comp
6
2 4 5
3 1
1st Comp
Tensor Decomposition
A∼ ∑ λ k a k ◦ bk ◦ c k .
k
Here, A = a ◦ b ◦ c → Aijk = ai b j ck which means that ◦ represents
the outer product. a and b are vectors of length |V | and c is a vector
of length | L|. Finally, λk is just a scaling factor that tells us how
much to count the kth element of the sum. The convention is to call
λk ak ◦ bk ◦ ck a component, while the vectors are the factors.
c1 c2 ck
a1 a2 ak
|V| ~ + +…+
|L| b1 b2 bk
|V|
If you find difficult to understand what’s going on by just looking Figure 5.19: A schema of tensor
at the formula, take inspiration from Figure 5.19. What this operation rank decomposition.
does is to find the right set of one-dimensional vectors a, b and c such
that, once they are scaled by factors λ, they can best represent the
full tensor A. At that point, you are working in a lower dimensional
space and all the rest of linear algebra starts making sense again.
How many components does this sum have? Or, in other words,
how big should k be to approximate A? That depends on the rank
of the tensor. Unfortunately, calculating the rank of a tensor isn’t as
easy as calculating the rank of a matrix. The rank of a matrix is the
number of columns (or rows) that are lineraly independent from each
other. The definition is the same for a tensor but, in this case, there
12
Joseph B Kruskal. Three-way arrays:
is no straightforward algorithm to determine it12 . What happens is
rank and uniqueness of trilinear
that, to find the rank of a tensor, you would literally apply the rank decompositions, with application to
decomposition with different k values and find the one that works arithmetic complexity and statistics.
Linear algebra and its applications, 18(2):
the best. 95–138, 1977
Tucker decomposition13 takes a different approach. It decomposes 13
Ledyard R Tucker. Some mathematical
our tensor A into a smaller core tensor and a set of matrices. If we notes on three-mode factor analysis.
Psychometrika, 31(3):279–311, 1966
keep our simplified case of a 3D tensor representing an adjacency
matrix, mathematically speaking the Tucker factorization does:
A ∼ A × X × Y × Z.
Here, A is the core tensor, whose dimensions are smaller than A’s.
X, Y, and Z are matrices which have one dimension in common with
A and the other in common with A – so that the matrix multiplica-
tion of them with A reconstructs a tensor with A’s dimensions.
matrices 87
|V| ~
|L|
|V|
Again, for the visual thinkers, Figure 5.20 might come in handy.
In Tucker decomposition you have the freedom to choose the di-
mensions of the core tensor A. Smaller cores tend to be more inter-
pretable, because they defer most of the heavy lifting to X, Y, and Z.
However, they also tend to make the decomposition less precise in
reconstructing A.
5.5 Summary
5.6 Exercises
Simple Properties
6
Degree
a connection.
The degree is a property of a node. Let’s call k v the degree of node
v. We can aggregate the degrees of all nodes in a network to get a
“global” information about its connectivity. The most common way to
do it is by calculating the average degree of a network. This would
be k̄ = ∑ k v /|V |, however it’s much simpler to remember that
v ∈V
k̄ = 2| E|/|V |. The average degree of a network is twice the number
of edges divided by the number of nodes. Why twice? Because each
edge increases by one the degree of the two nodes it connects.
In a social network, this is how many friends people have on
average. What would that number be in your opinion? If we have
a social network including two billion people, what’s the average
degree? It turns our that this number is usually ridiculously lower
than one would expect, because – as we’ll see in Section 9.1 – real 2
Anna D Broido and Aaron Clauset.
networks are sparse2 . Scale-free networks are rare. Nature
We call a node with zero degree, a person without friends, an communications, 10(1):1017, 2019
isolated node, or a singleton. A node with degree one is a “leaf”
node: this term comes from hierarchies, where nodes at the bottom
– the leaves of the tree – can only have one incoming connection
without outgoing ones. The sum of all degrees is 2| E|, which implies
that any graph can only have an even number of nodes with odd 3
Leonhard Euler. Solutio problematis ad
degree3 – otherwise the sum of degrees would be odd and thus it geometriam situs pertinentis. Commen-
cannot be two times something. tarii academiae scientiarum Petropolitanae,
pages 128–140, 1741
2
Figure 6.2: A graph and its
2 3
degree sequence.
1
2
[3, 2, 2, 2, 1]
(a) (b)
4
Michael Molloy and Bruce Reed.
The size of the giant component of a
A degree sequence is the list of degrees of all nodes in the net- random graph with a given degree
work4 , 5 . Typically, we sort the nodes in descending degree, so you sequence. Combinatorics, probability and
computing, 7(3):295–305, 1998
always start with the node with maximum degree and you go down 5
Béla Bollobás, Oliver Riordan, Joel
until you reach the node with the lowest degree. Figure 6.2 shows an Spencer, and Gábor Tusnády. The degree
example. sequence of a scale-free random graph
process. Random Structures & Algorithms,
Note that not all lists of integers are valid degree sequences. Some 18(3):279–290, 2001
lists cannot generate a valid graph. The easiest case to grasp is if
they contain an odd number of odd numbers. As we just saw, the
6
Gerard Sierksma and Han Hoogeveen.
degree sequence must sum to an even number (2| E|), thus a sequence
Seven criteria for integer sequences
summing to an odd number cannot describe a simple undirected being graphic. Journal of Graph theory, 15
graph6 . We call all valid sequences “graphic”. We’ll see that there are (2):223–231, 1991
Of course, the degree definition I just gave only makes sense in the
world of undirected, unweighted, unipartite, monolayer networks.
We had two whole chapters detailing when such a simple model
doesn’t work in complex real scenarios. We need to extend the
definition of degree to take into account all different graph models
we might have to deal with.
Directed
As we saw in Section 3.2, edges can have a direction, meaning that
the edge going from u to v doesn’t necessarily point back from v
to u. Such is life. In directed graphs you can obviously still keep
counting the degree as simply the number of connections of a node,
but there is a more helpful way to think about it. You might want
to distinguish the people who send a lot of connections – but don’t
necessarily see them reciprocated –, and those who are the target of a
lot of friends requests – whether they accept them or not.
1 1
2 0
0 2
(a) (b)
So we split the concept in two parts, helpfully named in-degree 7
Frank Harary, Robert Zane Norman,
and out-degree7 , 8 . As one can expect, the in-degree is the number of and Dorwin Cartwright. Structural
models: An introduction to the theory of
incoming connections. If we represent a directed edge as an arrow, directed graphs. Wiley, 1965
the in-degree is the number of arrow heads attached to your node. 8
Jørgen Bang-Jensen and Gregory Z
See Figure 6.3(a) for a helpful representation. The out-degree is the Gutin. Digraphs: theory, algorithms and
applications. Springer Science & Business
number of outgoing connections, the number of arrow tails attached Media, 2008
to your node. I show the out-degree of the nodes in my example in
Figure 6.3(b).
A directed graph’s degree sequence is now a list of tuples. The
first element of the tuple tells you the indegree, while the second
degree 93
element tells you the outdegree. Or you can have two sequences, but
you need to make sure that the nth positions of the two sequences re-
fer to the same node. If the two sequences are the same, meaning that
every node has the same in- and out-degree, we have a “balanced”
graph.
Weighted
Most of the time, people would not adapt the definition of degree
when dealing with weighted networks. Many network scientists like
how the standard definition works in weighted graphs, and keep it
that way. The degree is simply the number of connections a node has.
Bipartite
The bipartite case doesn’t need too much treatment: the degree is
still the number of connections of a node. It doesn’t matter much
that for V1 nodes it is gained exclusively via connections to V2 nodes
and viceversa. However, there’s a little change when one uses a
matrix representation that it’s worthwhile to point out. Assuming
A as a binary adjacency matrix (not stochastic), in the regular case
the degree is the sum of the rows: the sum of first row tells you the
degree of the first node, and so on.
Multilayer
The multilayer case is possibly the most complex of them all. At first,
it doesn’t look too bad. The degree is still the number of connections
a node has. Then you realize that there are some connections you
shouldn’t count. For instance, no one – that I know of – counts the
interlayer coupling connections as part of the degree. It’s easy to
see why: these are not connections that lead you to a neighbor in a
proper sense. They lead you to... a different version of yourself.
Consider Figure 6.7. Let’s call u the bottom node, the one with two
orange edges, a blue and a purple one. We can see that its degree
is four (four edges), and its neighbor set is of size three: u has three
neighbors, | Nu | = 3.
Now, we can count the size of the neighbor set per layer too, or
Nu,l . In the orange layer u has two neighbors (| Nu,l | = 2), in the blue
and purple one it has only one (| Nu,l | = 1). There is a difference
between the neighbors in the orange layer and the ones in the other
layers. If u wants to communicate with them, it has to use the orange
layer: there is no alternative. On the other hand, if the blue layer
were to disappear, u could still use the purple one, and vice versa.
This observation is at the basis of the definition of the “exclusive
neighbor” set, or N XOR . Given a node u and a layer l, the Nu,l XOR
purple one, because they have to share fairly the remaining 1/3 of u’s
neighbors.
Hyper
As one might expect, allowing edges to connect an arbitrary number
of nodes – rather than just two – does unspeakable things to your
intuition of the degree. We can still keep our usual definition: the
degree in a hypergraph is the number of hyperedges to which a node 18
Paul Erdős and Miklós Simonovits. Su-
belongs – or: the number of its hyper-connections18 , 19 . However, if persaturated graphs and hypergraphs.
Combinatorica, 3(2):181–192, 1983
you take any step further, all hell breaks loose. The number of neigh- 19
Alain Bretto. Hypergraph theory: An
bors has no relationship whatsoever with the number of connections: introduction. Mathematical Engineering.
with a single hyperedge you can connect a node with the entirety of Cham: Springer, 2013
the network. Also the average degree is something tricky to calculate.
Forget about k̄ = 2| E|/|V |: if a single hyperedge can connect the
entire network, then | E| = 1, but k̄ = |V |.
Things are a bit less crazy for uniform hypergraphs – where we
force hyperedges to always have the same number of nodes. Which
might explain why they’re a much more popular thing to study,
rather than arbitrary hypergraphs.
The degree of a node only gives you information about that node.
The average degree of a network gives you information about the
whole structure, but it’s only a single bit of data. There are many
ways for a network to have the same average degree. It turns out that
looking at the whole degree distribution can shed light on surprising
properties of the network itself. Since degree distributions can be so
important, generating and looking at them is a second nature for a
network scientist. As a consequence, there are a lot of standardized
procedures you want to follow, to avoid confusing your reader by
breaking them.
# Nodes
1
2 2
Figure 6.8: The degree scatter
plot (left) of the graph on the
2
Degree right.
98 the atlas for the aspiring network scientist
0.5 1
0.45
0.4 Figure 6.9: The degree distri-
0.35 0.1
0.3
bution of the protein-protein
p(k)
p(k)
0.25
0.2 0.01
interaction network. The distri-
0.15
0.1
butions are the same, but in (a)
0.05 0.001
0
we have a linear scale for the x
0 10 20 30 40 50 60 70 80 90 100 1 10 100
k k
and y axes, which is replaced in
(b) by a log-log scale.
(a) (b)
Figure 6.9(a) shows you the degree distribution of protein-protein
interaction for the Saccharomyces Cerevisiae, the beer bug. An inter-
esting pattern is that there are lots of nodes with few interactions,
and few nodes with many. As a consequence, we end up with all our
datapoints concentrated in the same part of the plot, and it’s difficult
to appreciate both the low- and the high-degree structure. These
degree patterns are more evident and easy to see when represented
on a log-log scale, as Figure 6.9(b) shows, which stretches out the
low-degree area while compressing the high-degree one.
So, we just discovered that this protein-protein interaction network
has something peculiar. The baseline assumption would be that
nodes connect at random. If that were the case, we would expect the
degree to distribute normally, in a nice bell-shape – see Chapter 13.
But Figure 6.9(b) is not what a normal distribution looks like. The
vast majority of nodes have a very low degree, and a few giant hubs
have a degree much larger than average. Is this common?
Yes it is. Most real world network would show such a broad 20
Holger Ebel, Lutz-Ingo Mielsch, and
distribution: email exchanges20 , synapses in the brain21 , internal cell Stefan Bornholdt. Scale-free topology of
interactions22 . Take a look at the degree distribution zoo in Figure e-mail networks. Physical review E, 66(3):
035103, 2002
6.10. To put it simply: in most networks we have many orders of 21
Victor M Eguiluz, Dante R Chialvo,
magnitude between the minimum and the maximum degree (x axis), Guillermo A Cecchi, Marwan Baliki,
and between the most and least popular degree value (y axis). This is and A Vania Apkarian. Scale-free brain
functional networks. Physical review
not what scientists initially expected. And when things are not as we letters, 94(1):018102, 2005
expected, we all get excited and start wonder why. 22
Reka Albert. Scale-free networks in
Before exploring these questions we need to finish our deep dive cell biology. Journal of cell science, 118(21):
4947–4957, 2005
into how to generate and visualize a proper degree distribution. The
degree 99
100 100
p(k)
networks: (a) coauthorship in
10-3 10-3
scientific publication [Leskovec
10-4 10-4 et al., 2007b]; (b) coappearance
1 10 100 1000 100 101 102 103 104
k k
of characters in the same comic
book [Alberich et al., 2002]; (c)
(a) (b)
100 100
interactions of trust between
10-1
PGP users [Boguñá et al., 2004];
10-1
10-2
(d) connections through the
p(k)
p(k)
10-2
10-3
Slashdot platform [Leskovec
10-3 et al., 2009].
10-4
-4
10
1 10 100 1000 100 101 102 103 104
k k
(c) (d)
100
3 3
-1 10 10
10
102 102
p(k)
p(k)
p(k)
-2
10
1 1
10-3 10 10
10-4 10
0
10
0
0 1 2 3 0 1 2 3
1 10 100 1000 10 10 10 10 10 10 10 10
k k k
That’s why you do power binning. You start with a small bin size,
usually equal to 1. Then each bin becomes progressively larger, by a
constant multiplicative factor. At first, the bins are still small. But, as
you progress, the bins start to be large enough to group a significant 25
An example of power binning,
portion of your space25 . A good power bin choice can make the plot starting with size 1 and increas-
clearer, as the one in Figure 6.11(c). In Figure 6.11(c) we saved the ing the bin size by 10% at each
step: [1, 2, 3, 5, 6, 8, 9, 11, 14, ...,
head of the distribution and further reduced the fat tail.
1410, 1552, 1709, 1881, 2070, 2278, ...]
# Nodes p(k>=x)
Figure 6.12: The degree scatter
plot (left) and its corresponding
complement of the cumulative
distribution (CCDF).
Degree x
1 1
0.01 0.01
interaction network. The distri-
butions are the same and are
0.001 0.001
both in log-log scale, but in (a)
1 10 100 1 10 100
k x
we have the degree histogram,
and in (b) we show the CCDF
(a) (b)
version (with the best fit in
We can see the relationship of our protein-protein network more
blue).
clearly in Figure 6.13. It appears that, in log-log space, the relation-
ship between degree and the number of nodes with a given degree
degree 101
1
Figure 6.14: An example of
power law, showing how the
0.1
red line always goes down by
p(k>=x)
1 10 100
x
in and out the picture and the size distributions would be the same.
This is the scale invariance I’m talking about: no matter the zoom, the
picture looks the same – obviously in reality it doesn’t, because the
moon isn’t an infinite plane, and you cannot zoom in infinitely many
times (in fact, whether finite systems can actually generate power 27
Michael PH Stumpf and Mason A
laws is a controversial topic27 ). Porter. Critical truths about power laws.
Science, 335(6069):665–666, 2012
1
Figure 6.16: An example of
power law in a CCDF. The ver-
0.1
tical gray bar shows that the
p(k>=x)
1
α=2 Figure 6.17: The CCDF degree
α=3
distributions of two random
0.1 networks with different α expo-
p(k>=x) nents.
0.01
0.001
1 10 100
x
(a) α = 2
(b) α = 3
Figure 6.18 confirms this: in Figure 6.18(a) you see that, for α = 2,
you have only one obvious hub that is head and shoulders above the
rest, practically connected to the entire network. In Figure 6.18(b),
instead, you still have a clear winner catching your eye (in the top),
but it is much closer to the second best hub.
The average degree is heavily influenced by the outliers with
thousands of connections. For instance, in Figure 6.16 the average
degree is equal to three, meaning that around 70% of nodes are
below average. This is well illustrated by the stadium example: you
have a stadium with 79, 999 individuals sampled at random from
the US population. If you calculate their average net worth you’ll
obtain a value – it’s difficult to be precise, but let’s say it’s around
104 the atlas for the aspiring network scientist
100 10
0
1
-1
10 -1
10
0.1
10-2
p(k>=x)
p(k>=x)
p(k>=x)
-2
-3
10
10 0.01
-4 -3
10 10
0.001
0 1 2 3 4 0 1 2 3 4
10 10 10 10 10 10 10 10 10 10 1 10 100
x x x
0 0 0
10 10 10
-1
10
-1 10
-1
10
-2
10-2
10
p(k>=x)
p(k>=x)
p(k>=x)
10-3 10
-2
10-3 -4
10
-3
-4 -5 10
10 10
-6
10
-5 10 10
-4
0 1 2 3 4 5 0 1 2 3 4 5 6 0 1 2 3 4
10 10 10 10 10 10 10 10 10 10 10 10 10 10 10 10 10 10
x x x
Early works have found power law degree distributions in many Figure 6.19: A showcase of
networks, prompting the belief that scale free networks are ubiqui- broad degree distributions from
tous. In fact, this seems true. Figure 6.19 shows the CCDFs of many the same networks used in
networks: protein interactions, PGP, Slashdot, DBpedia concept the examples in the previous
network, Gowalla, Internet autonomous system routers. section.
But we need to be aware of our tendency of seeing patterns when
they aren’t there – after all, as Feynman says, the easiest person you
can fool is yourself. So in the next section I’ll give you an arsenal to
defend yourself from your own eyes and brain.
0
10 100
PowerLaw
10
-1 10-1 Lognormal
Figure 6.20: (a) An example
10-2
of a CCDF that is most defi-
p(k>=x)
p(k>=x)
10-2 10-3
10-4
nitely NOT a power law, but
10
-3
??? 10-5 that a researcher with a lack of
-4 10-6 proper training might be fooled
10
1 10 100 1000 100 101 102 103 104 105 106
x x
into thinking it is. (b) Fitting a
power law (blue) and a lognor-
(a) (b)
mal (green) on data (red) can
yield extremely similar results.
want to pass as one perpetuating the myth that “everything that
looks like a straight line in a log-log space is a power law”. That is
equally wrong, even if more subtle and harder to catch.
Seeing the plot in Figure 6.20(b), you might be tempted to perform
a linear fit in the log-log space. This more or less looks like fitting the
logged values with a log( p( x )) = α log( x ) + β. Transforming this back
into the real values, the slope α becomes the scaling factor, and β is
the intercept, in other words: log( p( x )) = α log( x ) + β is equivalent to
p( x ) = 10β x α – assuming you logged to the power of ten.
A small aside: if you were to do this on the distributions from Fig-
ure 6.17, you would expect to recover α ∼ 2 and α ∼ 3, because I told
you I generated the degree distributions with those exponents. In-
stead, you will obtain α ∼ 1 and α ∼ 2, respectively. That is because,
in Figure 6.17, I showed you the CCDF of the degree distribution, not
the distribution itself. The CCDF of a power law is also a power law, 32
Heiko Bauke. Parameter estimation for
but with a different exponent32 . If you’re doing the fit on the CCDF, power-law distributions by maximum
you have to remember to add one to your α to recover the actual likelihood methods. The European
Physical Journal B, 58(2):167–173, 2007
exponent of the degree distribution.
Back to parameter estimation. If you perform a simple linear
regression, you’ll get an unbelievably high R2 associated to a super-
significant p value. Well, of course: you’re fitting a straight line over a
straight-ish line. Does that mean you’re looking at a power law? Not
really.
Just because something looks like a straight line in a log-log plot,
it doesn’t mean it’s a power law. You need a proper statistical test to
confirm your hypothesis. The reason is that other data generating
processes, such as the ones behind a lognormal distribution, can
generate plots that are almost indistinguishable from a power law.
33
Aaron Clauset, Cosma Rohilla Shalizi,
Figure 6.20(b) shows an example. You cannot really tell which of the
and Mark EJ Newman. Power-law
two functions fits the data better. distributions in empirical data. SIAM
What you need to do is to fit both functions and then estimate review, 51(4):661–703, 2009
34
Jeff Alstott and Dietmar Plenz Bull-
the likelihood of each model to explain the observed data33 . This
more. powerlaw: a python package for
can be done with, for instance, the powerlaw package34 , 35 – available analysis of heavy-tailed distributions.
for Python. However, be prepared for the fact that having a signifi- PloS one, 9(1), 2014
35
[Link]
cant difference between the power law and the lognormal model is powerlaw
106 the atlas for the aspiring network scientist
extremely hard.
In most practical scenarios, you’ll have to argue that your network
is a power law. How could you do it? Well, in complex networks
power law degree distributions can arise by many processes, but one
in particular has been observed time and time again: cumulative
advantage. Cumulative advantage in networks says that the more
connections a node has, the more likely it is that the new nodes will
connect to it. For instance, if you write a terrific paper which gathers
lots of citations this year, next year it will likely gain more citations 36
Derek de Solla Price. A general theory
than the less successful papers36 . of bibliometric and other cumulative
This is the same mechanism behind – for instance – Pareto distri- advantage processes. Journal of the
American society for Information science,
butions and the 80/20 rule. Pareto says that 80% of the effects are
27(5):292–306, 1976
generated by 20% of the causes37 . For instance, 20% of people control 37
Vilfredo Pareto. Manuale di economia
80% of the wealth. And, given that it takes money to make money, politica con una introduzione alla scienza
sociale, volume 13. Società editrice
they are likely to hold – or even grow – their share, given their abil- libraria, 1919
ity to unlock better opportunities. In fact, the Pareto distribution is
a power law. Similar to this is Zipf’s Law, the observation that the
second most common word in the English language occurs half of the
time as the most common, the third most common a third of the time, 38
Jean-Baptiste Estoup. Gammes
etc38 , 39 , 40 . In practice, the nth word occurs 1/n as frequently as the sténographiques: méthode et exercices
first, or f (n) = n−1 , which is a power law with α = 1. pour l’acquisition de la vitesse. Institut
sténographique, 1916
This is opposed to the data generating process of a lognormal 39
Felix Auerbach. Das gesetz der
distribution. To generate a lognormal distribution you simply have to bevölkerungskonzentration. Petermanns
multiply many random and independent variables, each of which is Geographische Mitteilungen, 59:74–76,
1913
positive. A lognormal distribution arises if you multiply the results 40
GK Zipf. The psycho-biology of
of many ten-dice rolls. You can see that there is no cumulative advan- language. 1935
tage here: scoring a six on one die doesn’t make a six more likely on
any other die – nor influences subsequent rolls.
So, to sum up, to test for a power law you have to do a few things.
First, make sure that your observations cannot be explained with
an exponential. Confusion between a power law and some other
distribution such as an exponential is easy, and so you should start
by assuming that your distribution is not a power law. Second, try to
see if you can statistically prefer a power law model over a lognormal. 41
Whether this holds true also for
In the likely event of you not being able to mathematically do so, networks is the starting point of
a surprisingly hot debate, see for
you should look at your data generating process. If you have the instance [Broido and Clauset, 2019] and
suspicion that it could be due to random fluctuations, then you might [Voitalov et al., 2018].
have a lognormal. Otherwise, if you can make a convincing argument
42
Gudlaugur Jóhannesson, Gunnlaugur
Björnsson, and Einar H Gudmundsson.
of non-random cumulative advantage, go for it. Afterglow light curves and broken
There are a few more technicalities. Pure power laws in nature are power laws: a statistical study. The
Astrophysical Journal Letters, 640(1):L5,
– as I mentioned earlier – rare41 . Your data might be affected by two 2006
impurities. Your power law could be shifted42 , or it could have an 43
Aaron Clauset, Cosma Rohilla Shalizi,
exponential cutoff43 . In a shifted power law, the function holds only and Mark EJ Newman. Power-law
distributions in empirical data. SIAM
on the tail. In an exponential cutoff the power law holds only on the review, 51(4):661–703, 2009
degree 107
head.
Shifted power laws have an initial regime where the power law
doesn’t hold. Formally, the power law function needs a slowly grow-
ing function on top that will be overwhelmed by the power law for
large values of k – as I show in Figure 6.21(a). So we modify our
master equation as: p(k) ∼ f (k)k−α , with f (k) being an arbitrary
but slowly growing. Slowly growing means that, for low values of
k it will overwhelm the k−α term, but for high values of k, the latter
would be almost unaffected. In power law fitting, this means to find
the k min value of k such that, if k < k min we don’t observe a power
law, but for k > k min we do.
0 0
10 10
10
-1
p(k) ~ f(k)k-α -1
Figure 6.21: (a) An example of
10
10
-2 shifted power law. The area in
p(k>=x)
p(k>=x)
10-2
10-3 which the power law doesn’t
10
-4 10-3 p(k) ~ k-αe-λk hold is shaded in blue. (b) An
10-5 example of truncated power
10-4
100 101 102 103 104 105 1 10 100 1000 law: a power law with an ex-
x x
ponential cutoff. The area in
(a) (b)
which the power law doesn’t
Shifted power laws practically mean that “Getting the first k min hold is shaded in green.
connections is easy”. If you go and sign up for Facebook, you gener-
ally already have a few people you know there. Thus we expect to
find fewer nodes with degree 1, 2, or 3 than a pure power law would
predict. The main takeaway is that, in a shifted power law, we find
fewer nodes with low degrees than we expect in a power law.
Truncated power laws are typical of systems that are not big
enough to show a true scale free behavior. There simply aren’t
enough nodes for the hubs to connect to, or there’s a cost to new
connections that gets prohibitive beyond a certain point. This is
practically a power law excluding its tail, that’s why we call them
“truncated”. Mathematically speaking, this is equivalent to having an
exponential cutoff added to our master equation: p(k) ∼ kα e−λk . The
exponential function is dominated by the power law function for low
values of k, but it becomes dominant for high values of k. See Figure
6.21(b) for an example.
Truncated power laws practically mean that “Getting the last
connections is hard”: the biggest superstar on Twitter has a lot of
followers, but relatively speaking they are not that many more as the
second biggest superstar on Twitter. Thus its degree is not as big as
we would expect. The main takeaway is that, in a truncated power
law, the hubs have lower degrees than we expect in a power law.
At the end of the day, it doesn’t matter too much if your network
has an exponential, lognormal or power law degree distribution. On
108 the atlas for the aspiring network scientist
one thing the brotherhood of network scientists can agree: the vast
majority of networks have broad degree distributions, spanning mul-
tiple orders of magnitude. Most nodes have below-average degree
and hubs lie many standard deviations above the average. Even if
they are not power laws at all, that’s still pretty darn interesting.
6.5 Summary
6.6 Exercises
1. Write the in- and out-degree sequence for the graph in Figure
6.3(a). Are there isolated nodes? Why? Why not?
2. Calculate the degree of the nodes for both node types in the
bipartite adjacency matrix from Figure 6.5(a). Find the isolated
node(s).
degree 109
3. Write the degree sequence of the graph in Figure 6.7. First consid-
ering all layers at once, then separately for each layer.
6. Find a way to fit the truncated power law of the network at http:
//[Link]/exercises/6/6/[Link]. Hint: use the
[Link].curve_fit to fit an arbitrary function and use the
functional form I provide in the text.
7
Paths & Walks
both conventions and you simply use the one you find most natural –
as everybody does in network analysis.
same edge as many times as you want. The length of a walk is the
number of times you’re using the edges in your walk. If you use the
same edge n times, this will increase the walk’s length by n. Figure
7.1(a) shows an example of a walk of length 6 in a network.
In a walk the choice of the next edge to explore is yours. You can
have a slightly more constrained definition of a walk, where you
put rules to choose the next edge to traverse. For instance, in the
random walk you impose to make this choice completely at random.
We already saw a way to calculate node exploration probabilities via
a random walk using powers of the adjacency matrix in Section 5.1.
We’ll see that random walks are a phenomenally powerful way to
explore your network’s properties and are at the basis of countless
methods: Chapter 8 will be but a superficial introduction.
When you impose even more constraints on your walks, then
you can generate a path, or “simple path” (Figure 7.1(b)). This is a
walk that does not repeat nodes nor edges. Again, you can put more
qualifiers on your path to make it special. For instance, recalling
the seven bridges problem, an Eulerian path is a path that travels
through all edges of a connected graph – since it is a path, not only it
has to visit each edge, but it also has to do it exactly once. A cousin
of the Eulerian path is the Hamiltonian path, which instead wants to
visit each node – not edge – exactly once. More interestingly, you can
try to find the shortest path between two nodes. That will be the topic
of Chapter 10.
Similarly to a walk, also a path has a length. This is again defined
as the number of edges the path crosses. Since no edge can be used
twice in a path, this is also the number of distinct edges used.
2
Figure 7.2: (a) A graph. (b-c)
3 Different powers of its binary
4
5 adjacency matrix A.
6
7.2 Cycles
You can make a walk and a path in any graph, no matter its topology.
There is a special path that you cannot always do, though. That is the
cycle. Picking up the social network example as before, now you’re
not happy just by reaching somebody with your message. You want
the message you originally sent to come back to you. Also, you don’t
want anybody to hear it twice. If you manage to do so, then you have
found a cycle in your social network.
A cycle is a path that begins and ends with the same node. Note
that I said “path”, so we don’t have any repeated nodes nor edges –
to that directed tree, except for that little pesky edge at the bottom,
going into the opposite direction. To restore sanity to the network
world, we decided to create a final definition for directed graphs:
arborescences. This is French for “tree”, and in fact the two terms
are often used interchangeably. But, technically speaking, an arbores-
cence is a directed tree in which all nodes have in-degree of one,
except the root. In an arborescence, the root is a special node: the
only one with in-degree of zero. An arborescence must have one and
only one root. Figure 7.5(d) fixes Figure 7.5(c) to be compliant to the
definition of arborescence, and it is a work of art. So satisfying.
7.3 Reciprocity
Walks and paths can help you uncover some interesting properties
in your network. Let’s pick up our game of message-passing. In this
scenario, we might end up in a situation where there is no way for
a message to reach some of the people in the social network. The
people you can reach with your message do not know anybody who
can communicate to your intended targets. In this scenario, it is
natural to divide people into groups that can talk to each other. These
are the network’s “components”.
Unreachable
λ1
2 cency matrix of a disconnected
4
6 1 graph looks like two differ-
9 ent adjacency matrices pasted
5 on the diagonal. Thus, they
λ2 8
both have a (different) leading
7
eigenvalue equal to one.
paths & walks 117
a 0
a 0
λ1
Figure 7.10: If your graph has
a 0 two components, the eigen-
a 0
vectors associated with the
a 0
a 0 largest two eigenvalues of the
0 b stochastic adjacency matrix will
0
0
b
b λ2 tell you to which component
the node belongs, by having a
v1 v2
non-zero value.
5 6 8 9
(a) (b)
not part of the same strongly connected component. SCCs are impor-
tant: if you are playing a message-passing game where messages can
only go in one direction, you can always hear back from the players
in the same strongly connected component as you.
Popular algorithms to find strongly connected components in 5
Robert Tarjan. Depth-first search and
a graph are Tarjan’s5 , Nuutila’s6 , and others that exploit parallel linear graph algorithms. SIAM journal on
computation7 . computing, 1(2):146–160, 1972
The definition of SCC leaves the door open for some confusion.
6
Esko Nuutila and Eljas Soisalon-
Soininen. On finding the strongly
Even by visually inspecting a network that appears to be connected connected components in a directed
in a single component, you will find multiple different SCCs – as in graph. Inf. Process. Lett., 49(1):9–14, 1994
Figure 7.11(b). In the figure, there is no path that respects the edge
7
Sungpack Hong, Nicole C Rodia, and
Kunle Olukotun. On fast parallel detec-
directions and leads from node 1 to node 7 and back. The best one tion of strongly connected components
could do is 1 → 2 → 3 → 7 → 8 → 6 → 5 → 4. (scc) in small-world graphs. In Proceed-
ings of the International Conference on
However, it feels like this network should have one component,
High Performance Computing, Networking,
because we can see that there are no cuts, no isolated vertices. If Storage and Analysis, pages 1–11, 2013
we were to ignore edge directions, Figure 7.11(b) would really look
like a connected component in an undirected network. This feel-
ing of uneasiness led network scientists to create the concept of
“weakly connected components” (WCC). WCCs are exactly what I
just wrote: take a directed network, ignore edge directions, and look
for connected components in this undirected version of it. Under this
definition, Figure 7.11(b) has only one weakly connected component.
know you’re never going to see the same document twice. That
would imply that there is a cycle, and thus that you are in a strongly
connected component with someone. However, what you see in the
document can be radically different. The document might arrive
to you with or without the core’s stamp of approval. These two
scenarios are quite different.
If you are in the first scenario, it means your WCC is positioned
“before” the core. Documents pass through it and they are put in the
core. The flow of information originates from you or from some other
member of the weakly connected component, and it is poured into
the core. This is the scenario in Figure 7.12(a): you are one of the four
leftmost nodes. In this paragraph I highlighted the word in because
we decided to call these special WCCs in-components.
7.5 Summary
2. Cycles are paths which start and end in the same node. Acyclic
graphs are graphs without cycles. An undirected acyclic graph is
called a tree – a graph with |V | nodes and |V | − 1 edges. Otherwise,
you can have directed acyclic graphs which are not trees.
7.6 Exercises
Figure 8.2 shows the result. We see now that the columns are
constant vectors. These numbers have a specific meaning. When we
calculate A∞ , what we’re doing is basically asking the probability
of being in a node after a random walk of infinite length. Since
122 the atlas for the aspiring network scientist
the length is infinite, it does not really matter from which node
you originally started. That’s why all rows of A∞ are the same –
remember that the row indicates the starting point while the column
indicates the ending point.
This row vector – you can pick any of them, since they’re all
the same – is so important that we give it a name. We call it the
“stationary distribution” – or π, for short. π tells us that, if you have
a path of infinite length, the probability of ending up on a destination
is only dependent on the destination’s location and not on your point
of origin. In practice, if you apply the transition probability (A) to the
stationary distribution (π), you still obtain the stationary distribution:
πA = π. Having a high value in the stationary distribution for
a node means that you are likely to visit it often with a random
walker – by the way, this is almost exactly what PageRank estimates,
plus/minus some bells and whistles, see Section 11.4.
Note that it is not necessary to calculate A∞ to know the stationary
distribution. At least for undirected networks, π is quite literally the
normalized degree of the nodes: the degree divided by the sum of all
degrees (2| E|).
But... wait! This stationary distribution formula is oddly familiar:
πA = π. Haven’t we seen something similar to it? This kind of
looks like our eigenvector specification (Av = λv, see Section 5.2),
with a few odd parts. First, where’s the eigenvalue? Well, we can
always multiply a vector to 1 and we won’t change anything in the
equation. So: π A = π1. This is cool, because we already know that
1 is the largest eigenvalue (λ1 ) of a stochastic matrix. Second, the
vector π is on the left, not on the right. Putting these things together:
the stationary distribution π is the vector associated with the largest
eigenvalue, if multiplied on the left of A. Therefore: π is the leading
left eigenvector.
If you’re dealing with an undirected graph, there is a relationship
between right and left eigenvectors. If you were to transpose the
stochastic adjacency matrix, that is making it column-normalized
instead of row-normalized, the left and right eigenvectors would
swap. In different words: the left eigenvectors of A are exactly the
same as the right eigenvectors of A T . Thus the vector of constant
and π are the right and left leading eigenvectors of A, and they swap
roles in A T .
What do you do if your graph is not connected? No matter how
many powers of A you take, how infinitely long your walks are,
some destinations are unreachable from some origins. We end up
with two stationary distributions, one for one component, and one
for the other. Figure 8.3 shows an example. These two stationary
distributions are not directly comparable one with the other. They are
random walks 123
8 7
having a row and a column per node, we instead have a node and 5
Ki-ichiro Hashimoto. Zeta functions
column per edge5 . Each cell contains a one if we can use the edge for of finite graphs and representations of
our non-backtracking walk, zero otherwise. Formally: p-adic groups. In Automorphic forms
and geometry of arithmetic varieties, pages
211–280. Elsevier, 1989
1 if u 6= z
NBuv,vz =
0 otherwise.
2
0.5 × 0.5
1 0.5
= 1/2 and − √ = 0. So, for n = 2 the part on
1 − λ3 1 1
the right side of the sum evaluates to 1 and for n = 3 it evaluates to
zero. Thus the total sum is one. Multiplied to 2| E| you obtain four.
This is super intuitive. How long does it take to get from node 1 to
node 2? Well, node 1 has only one connection and it goes to node 2,
so it will always take one step. But to get from node 2 to node 1, you
only have a 50% chance of doing it in one step. The other 50% of the
times the random walker will go to node 3. It will always come back
after another step, and then we’ll have another 50% chance to go to
node 1. You sum that to infinity, and you get an expected hitting time
of three.
λ1 λ2 λ3
1 k1 = 1 H Figure 8.6: The elements
needed to calculate the hit-
ting time of a graph. From left
to right: the graph and its de-
2
gree vector, the eigenvalues and
k2 = 2
eigenvectors of N, the resulting
hitting time matrix H.
3 k3 = 1
w1 w2 w3
|V |
1
∑ πv Hu,v = ∑ 1 − λn
.
v ∈V n =2
Noticing something weird? The right hand side has no trace of u.
This means that the average time to hit something doesn’t depend on
your starting point u, exactly like, in the stationary distribution, the
probability of ending your random walk somewhere didn’t depend
on from where you started.
random walks 127
.37 2 1
+
.37 Figure 8.7: (a) The Laplacian
.28 matrix of a graph. We show the
3 4
.28 2-cut solution for this graph
.12 5 with the red and blue blocks.
-.23 (b) The second smallest eigen-
- -.38
-.38
7
6
8
vector of (a). (c) The graph view
of (a) – we color the nodes ac-
-.43 cording to the value attached to
v2 9 them in the (b) vector.
(a) (b) (c)
One classical problem in graph theory is to find the minimum
(or normalized) cut of a graph: how to divide nodes in two disjoint
groups such that the number of edges running across groups is
minimized. Turns out that the second smallest eigenvector of the 12
Miroslav Fiedler. Laplacian of graphs
Laplacian is a very good approximation to solve this problem12 . and algebraic connectivity. Banach Center
Publications, 25(1):57–70, 1989
How? Consider Figure 8.7. In Figure 8.7(a) I show the Laplacian
matrix of a graph. I arranged the rows and columns of the matrix
so that the 2-cut solution is evident: by dividing the matrix in two
diagonal blocks there is only one edge outside our block structure
that needs to be cut.
Now, why did I label the two blocks as “+” and “-”? The reason
lies in the magical second smallest eigenvector of the Laplacian – also
known as the Fiedler vector –, which is in Figure 8.7(b). We can see
that the top entries are all positive (in red) and the bottom are all
negative (in blue). This is where L shines: by looking at the sign of
the value of a node in its second smallest eigenvector we know in
which group the node has to be to solve the 2-cut problem!
Not only that, but the values in Figure 8.7(b) are clearly in de-
scending order. If we look at the graph itself – in Figure 8.7(c) – and
128 the atlas for the aspiring network scientist
4
6 0.5
16 12
2 15
0.4 11
10
18 0.3 9
3rd Eigenvector
8
5
0.2
1 13
3
7
17 0.1
14 0
7 13 1
9
-0.1 17
1514 23
18 16 45
8 -0.2 6
11
-0.3
-0.4 -0.3 -0.2 -0.1 0 0.1 0.2 0.3 0.4
10
12 2nd Eigenvector
(a) (b)
One thing that might be left in your head after reading the previous
section is: why? Why do the eigenvectors of the Laplacian help with
the mincut problem? What’s the mechanism? To sketch an answer for
this question we need to look at what we call “consensus dynamics”.
This is a subclass of studying the diffusion of something (a disease,
word-of-mouth, etc) on a network – which we’ll see more in depth 13
Michael T Schaub, Jean-Charles
in Part V. This section is sketched from a paper13 that you should Delvenne, Renaud Lambiotte, and
Mauricio Barahona. Structured networks
read to have a more in-depth explanation of the dynamics at hand. and coarse-grained descriptions: A
Consensus dynamics were originally modeled this way by DeGroot14 . dynamical perspective. Advances in
Network Clustering and Blockmodeling,
In this section I’m going to use the stochastic adjacency matrix
pages 333–361, 2019
of the graph, but what I’m saying also holds for the Laplacian. The 14
Morris H DeGroot. Reaching a
difference between the two – as I also mention in Section 30.2 – is consensus. Journal of the American
Statistical Association, 69(345):118–121,
that the stochastic adjacency matrix describes the discrete diffusion
1974
over a network. In other words, you have a clock ticking and nothing
happens between one tick of the clock and the other. The Laplacian,
instead, describes continuous diffusion: time flows without ticks in
a clock, and you can always ask yourself what happens between two
observations. Besides this difference, the two approaches could be
considered equivalent for the level of the explanation in this section.
How does the stochastic adjacency help us in studying consensus
over a network? Let’s suppose that each node starts with an opinion,
which is simply a number between 0 and 1. We can describe the
status of a network with a vector x of |V | entries, each corresponding
to the opinion of each node. One valid operation we could do is
multiplying x with A, the stochastic adjacency matrix, since A is a
|V | × |V | matrix.
What does this operation mean? Mathematically, from Section 5.2,
|V |
the result is a vector x 0 of length |V | defined as xv0 = ∑ xu Auv . In
u =1
practice, the formula tells you that node v is updating its opinion
by averaging the opinion of its neighbors. Non-neighbors do not
contribute anything because Auv = 0, and this is an average because
we know that the rows of the adjacency matrix sum to 1 – thus each
xu is weighted equally and xv0 will still be between 0 and 1.
We can reach a consensus by multiplying x with A an infinite
number of times. If you do that, x will converge to a vector of con-
stant – which is the average of the initial x, in this case 0.5 since
values were extracted uniformly at random between 0 and 1. Figure
130 the atlas for the aspiring network scientist
1
Figure 8.9: The value of x
0.8 (y-axis) for each node in the
0.6 graph from Figure 8.8 over time
v(T)
8.9 shows an example process on the network from Figure 8.8. Each
node starts with a uniform random value and the line tracks this
value over time.
Now the connection with eigenvectors: from Section 8.1 you
remember that the vector of constant is the leading eigenvector of A.
This operation showed you why: if a network has separate connected
components, it cannot reach a unique consensus; every connected
component will reach its own consensus independently because
there’s no exchange of information across components.
The second eigenvector, instead, tells you how quickly the nodes
will reach the consensus. In the figure, the line color tells you the
community of the node. You might notice that nodes bundle up
with their community mates before reaching the network’s final
consensus. This is described by the inverse of the second eigenvalue
of the Laplacian. For that network, it is 1/λ2 ∼ 5.13. This tells you
that the nodes are expected to converge to their community’s opinion
between step 5 and 6, which is the step highlighted in Figure 8.9. You
might notice that one blue and one red node don’t seem to converge
to their community, but that’s because they are nodes 1 and 13 and,
as you can see from Figure 8.8, they are in between communities, i.e.
they are close to the cut.
So the reason why the Fiedler vector allows you to solve the
mincut is because it tells you how much it will take for the node
to converge to the consensus. Nodes farther from the cut will take
longer time. Moreover, you can use the sign to know on which side
of the cut you are, because the nodes will first tend to converge to the
value of their own community, which is the opposite of the value in
the other community.
8.6 Summary
3. The hitting time is the number of expected steps you need to take
in a random walk to reach one node starting from another. It is
related to a special eigenvector decomposition of the adjacency
matrix.
8.7 Exercises
(a) (b)
To quantify the difference between the two cases, network scien-
tists defined the concept of network density. Informally, this is the
probability that a random node pair is connected. Or, the number of
edges in a network over the total possible number of edges that can
exist given the number of nodes. We can estimate this latter quantity
density 133
quite easily – from now on, I’ll just assume the network is undirected,
unipartite and monolayer, for simplicity.
Here’s the problem, rephrased as a question: how many edges
do we need to connect |V | nodes? Well, let’s start by connecting one
node, v, to every other node. We will need |V | − 1 edges – we’re
banning self loops. Now we take a second node, u. We need to
connect it to the other nodes in V minus itself – seriously, no self
loops! – and v, because we already added the (u, v) edge at the
previous step. So we add |V | − 2 edges. If you go on and perform
the sum for all nodes in v, you’ll obtain that the number of possible
edges connecting |V | nodes is |V |(|V | − 1)/2. In other words, you
need |V | − 1 edges to connect a node in V with all the other nodes,
and you divide by two because each edge connects two nodes. In
fact, the number of possible edges in a directed network is simply
|V |(|V | − 1), because you need an edge for each direction, as (u, v) 6=
(v, u).
Now we can tell the difference between the networks in Figures
9.1(a) and 9.1(b). The first one has three nodes. We would expect
3 ∗ 2/2 = 3 edges, and that’s exactly what we have. Its density is
the maximum possible: 100%. The network on the right, instead,
contains 650 nodes. Since the average degree of the network is also
two, we know it also contains 650 edges. This is a far cry from the
650 ∗ 649/2 = 210, 925 we’d require. Its density is just 0.31%, more
than three hundred times lower than the example on the left!
With all this talk about real world networks having low degree, we
should expect them to be quite sparse. They are, in fact, even sparser
than you think. A few examples.
The network connecting the routers forming the backbone of the
Internet? It contains |V | = 192, 244 nodes. So the possible number of
edges is |V |(|V | − 1)/2 = 18, 478, 781, 646. How many does it have,
really? 609, 066, which is just 0.003% of the maximum. How about
the power grid? The classical studied structure has 4, 941 nodes,
which means – potentially – 12, 204, 270 edges. Yet, it only contains
6, 594 of them, or 0.05%. Well, compared to the Internet that’s quite
the density! Final example: scientific paper citations. A dataset from
arXiV contains 449, 673 papers. The theoretical maximum number
of citations is 101, 102, 678, 628. Physicists, however, are quite stingy:
they only made 4, 689, 479 citations, or 0.004% of the theoretical
maximum. (All figures are from classical network structures studied 1
Albert-László Barabási et al. Network
in the literature1 ) science. Cambridge university press, 2016
You might have spotted a pattern there. The density of a network
seems to go down as you increase the number of nodes. While not
an ironclad rule, you might be onto something. The problem is that
the denominator of the density formula grows quadratically. Each v
134 the atlas for the aspiring network scientist
1400
Potential Edges
1200 Figure 9.2: The red line shows
Actual Edges
1000
800
the number of possible edges
|E|
Density doesn’t solve all ambiguities you had in the case of the
average degree. Two networks can have the same density and the
same number of nodes, but end up looking quite different from each
other. That is why the ever industrious network scientists created
yet another measure to distinguish between different cases: the
clustering coefficient.
that the other two nodes will connect to each other to form a triangle.
“Triangle” is the name we give to a specific network pattern: three
nodes that are all connected to each other, and I show one in Figure
9.4(b).
(a) (b)
So, let’s calculate the global clustering coefficient for the graph in
Figure 9.6(a). We know how many triads there are in the graph. How
many triangles are there? Here I made my life easier, because it is
rather trivial to count the number of triangles in a planar graph – a
graph you can draw on a 2D plane without intersecting any edges.
There are eight triangles and 48 triads in the network. Thus the
global clustering coefficient of the network is CC = 3 × 8/48 = 0.5.
Half of the triads in the network close to form a triangle.
The fact that I keep calling this the “global” clustering coefficient,
density 137
should tip you off about the existence of a “local” clustering coef-
ficient. Its definition is rather similar, but it is focused on a single
node. It is the number of triangles to which the node belongs over
the number of triads centered on it. We do not multiply by three the
numerator, because we’re focusing exclusively on the triads of v, that
can then be closed only in one way.
Looking at Figure 9.6(b), let’s try to calculate the local clustering
coefficient of the node highlighted by the arrow. Again, we use
our planar graph to have an easy time counting triangles: there
are five that include node v. We already counted the number of
triples before: it was 15. Thus, the local clustering coefficient of v is
CCv = 5/15 = 0.3̄.
1/3
1/2
1 2/3
coefficient are two different things, and not to report one as the other.
I closed the previous section with a mantra – “real world networks
are sparse” –, so I want to do it again. However, there’s a surprise
here. If average degrees are low and networks are sparse, wouldn’t
you expect real world networks to have a low clustering too? Instead,
the opposite holds: real world networks are clustered. The power
grid example I used before? It has a CC of 0.1032, which is 150 times
higher than you would expect if its edges were distributed randomly.
The scientific paper citations? It has CC = 0.318, more than 200 times
higher than expected.
This means that these systems might have few connections per
node, but these connections tend to be clustered in the same neigh-
borhood. Nodes tend to close triangles. This is especially true for
social systems. In fact, the protein-protein network I used in Chapter
6 has a clustering of 0.0236, which is still higher than expected. But
in this biological case, we only have a factor of 16, a far cry from the
factor of 200 in the social paper citation system.
9.3 Cliques
7
Tommy R Jensen and Bjarne Toft.
problem7 . Solving graph coloring tells you, for instance, how many Graph coloring problems, volume 39. John
colors you need to use for your map such that no two neighboring Wiley & Sons, 2011
countries share the same hue – in this case you represent countries as
nodes and connect two countries if they share a land border.
When it comes to independent sets, you should not confuse the
maximal independent set with the maximum independent set. A maxi-
mal independent set, just like a maximal clique, is an independent set
to which you cannot add any node and still make it an independent
set. The green set in Figure 9.10 is maximal because the two green
nodes are connected to all other nodes in the graph.
On the other hand, the maximum independent set is simply the
largest possible independent set you can have in your network. In
Figure 9.10, the red set is the maximum independent set, or at least
one of the many possible independent sets of size 3. Finding the
largest possible independent set is an interesting problem, because it
tells you something about the connectivity of the graph. It also has
applications in graph mining – see Section 39.4.
9.5 Summary
9.6 Exercises
Centrality
10
Shortest Paths
1 1
2 12 1922 27 2 3 4 5 6
Random edge access is when you read the file containing your
graph one line at a time, and each line contains an edge. We call this
type of graph storing format an “edge list”, because it’s a list of one
edge per line. In this case, you may or may not have sorted the edges
in a particular way, but the baseline assumption is that they’re in a
random order. See Figure 10.2(a) for an example.
Random node access is the same, but the file records, in each line,
the complete list of a node’s neighbors. We call this type of graph
storing format an “adjacency list”, because it’s a list of the adjacency
of one node per line. Also in this case, you may or may not have
sorted the nodes – and their neighbors – in a particular way, but the
baseline assumption is that they’re in a random order. See Figure
10.2(b) for an example.
The problem of these ways to explore the graph is that they are
not optimal, they cannot find the “best” (shortest) way to go from
v to u. Well, they can, but only in very specific cases under very
specific assumptions. For instance, BFS finds shortest paths only for
undirected unweighted graphs. This can be useful, for instance, to
1
Shimon Even. Graph algorithms.
Cambridge University Press, 2011
find the shortest path out of a maze1 , 2 , 3 . But we still need a better, 2
Edward F Moore. The shortest path
more general way. through a maze. In Proc. Int. Symp.
To see why finding “best paths” is important, suppose you have Switching Theory, 1959, pages 285–292,
1959
to deliver a letter to a person, as I show in Figure 10.3. If you know 3
Konrad Zuse. Der Plankalkül. Num-
them, no problem: you just give it to them. What if you don’t? You ber 63. Gesellschaft für Mathematik und
might know one of their friends, and pass through them. Or you Datenverarbeitung, 1972
might know that one of your friends knows one of theirs. But if
none of this is true, you have to know the shape of the entire social
network – the part in gray in the figure – and discover what’s the
least amount of people you have to bother to get your letter to the
recipient. This we call the “shortest path problem” in networks.
(a) (b)
This is for the case of undirected, unweighted networks. If you
have directed networks you obviously have to respect the edge direc-
tions – see Figure 10.5(a). If you have weighted networks, you might
want to minimize the weight (as in Figure 10.5(b)), assuming that the
edge weight represents its traversal cost. If your edge weights repre-
sent proximities rather than costs, for instance they are the capacity
of a trait of road as explained in Section 3.3, you’d do the opposite.
(a) (b)
How do we find the shortest path? Depending on the properties
of the graph (e.g. direct/undirected, weighted/unweighted), there
are different algorithms for finding shortest paths. We also need to
know if we just want to find a path from v to u (single-origin single-
shortest paths 147
sp(u, v, 0) = Auv .
If K > 0 it means that we are adding a node as a possible member
of the shortest path. When we do it, either of two things can happen:
(i) adding the extra node allowed us to find a better (shorter) path, or
(ii) it didn’t. So:
1 3 2
1 8 4
3 1 1 3
1 6
2 1 3
1 1
8 4
1 3 2 2 4
2 2 4 1 1 3 1 2
3 2
2 3 6 4 3 1 2 2 4 1 1 3 1 2 2 4
Just like with the degree, knowing the length distribution of all
shortest paths in the network conveys a lot of information about its
connectivity. A tight distribution with little deviation implies that all
nodes are more or less at the same distance from the average node in
the network. A spread out distribution implies that some nodes are
in a far out periphery and others are deeply embedded in a core.
# Paths
Figure 10.9: The path length
distribution (left) of a graph
(right). Each bar counts the
number of shortest paths of
length one, two and three,
which is the maximum length
in the network.
Path Length
Diameter
The rightmost column of the histogram in Figure 10.9 is important.
It records the number of shortest paths of maximum length. These
are the “longest shortest paths”. Since this is an important concept,
such a long mouthful name won’t do. We’re busy people and we
got places to be. So we use a different name for them or, to be more
precise, to their length. We call it the diameter of the network.
Why do we care about the diameter? Because that’s the worst case
shortest paths 151
• ...
It’s now easy to see that a network with diameter equal to three
is easy to navigate. As the diameter grows, the number of people to
rely on for a full traversal starts becoming unwieldy.
If your network has multiple connected components (Section 7.4),
we have a convention. Nodes in different components are unreach-
able, and thus we say that their shortest path length is infinite. Thus,
a network with more than one connected component has an infinite
diameter. Usually, in these cases, what you want to look at is the
diameter of the giant connected component.
Average
The diameter is the worst case scenario: it finds the two nodes that
are the most far apart in the network. In general, we want to know
the typical case for nodes in the network. What we calculate, then, is
not the longest shortest path, but the typical path length, which is the
average of all shortest path lengths. This is the expected length of the
shortest path between two nodes picked at random in the network.
If Puv is the path to go from u to v and | Puv | is its length, then the
∑ | Puv |
u,v∈V
average path length of the network is APL = . Figure
|V |(|V | − 1)
10.10 shows that, even in a tiny graph, the diameter and the APL
can take different values, with the former being more than twice the
length of the latter.
With APL, we can fix the origin node. For instance, in a social
network, you can calculate your average separation from the world.
152 the atlas for the aspiring network scientist
Diameter = 3
Figure 10.10: The diameter and
the APL in a graph can be quite
different.
APL ~ 1.4
Mean = 3.57
125
Figure 10.11: The path length
100 distribution for Facebook in
FB Users (M)
2012.
75
50
25
2.5 2.7 2.9 3.1 3.3 3.5 3.7 3.9 4.1 4.3 4.5 4.7
Average Degree of Separation
This would be an APLv , the average path length for all paths starting
at v. Then you can generate the distribution of all lengths for all
origins. How does this APLv distribution look like for a real world
network? One of the most famous examples I know comes from 15
Sergey Edunov, Carlos Diuk, Is-
Facebook15 . I show it in Figure 10.11. The remarkable thing is how mail Onur Filiz, Smriti Bhagat, and
ridiculously short the paths are even in such a gigantic network. Moira Burke. Three and a half degrees
of separation. Research at Facebook, 2016
This is in line with classical results of network science, showing
that the diameter and APL typically grow sublinearly in terms of 16
Mark EJ Newman. The structure and
number of nodes in the network16 . In other words, there are dimin- function of complex networks. SIAM
ishing returns to path lengths: each additional person contributes review, 45(2):167–256, 2003b
origin city, given a list of cities and the distances between each pair of
cities. We can represent the problem as a weighted graph, with city
distances as edge weights. The minimum spanning tree doesn’t solve
the problem: it creates a tree, which has no cycles. Thus, to get back
to the origin city, you have to backtrack all the way through the tree –
not ideal.
(a) (b)
biclique. Look at Figure 10.14 and try to draw those graphs in two
dimensions without having any edge crossing another one. You’ll
find out that is not possible.
The second cousin of spanning trees is the triangulated maximally 28
Guido Previde Massara, Tiziana
filtered graph28 . This was originally proposed as a more efficient Di Matteo, and Tomaso Aste. Network
algorithm to extract planar maximally filtered graphs from larger filtering for big data: Triangulated
maximally filtered graph. Journal of
graphs. However, it also allows to specify different topological con-
complex Networks, 5(2):161–178, 2016
straints, which are not necessarily making the graph planar.
10.6 Summary
2. Shortest paths are the paths connecting two arbitrary nodes in the
network using the minimum possible number of edges. In directed
networks you have to respect the edge’s direction, in weighted
networks you have to minimize (or maximize, depending on the
problem definition) the sum of the edge weights.
10.7 Exercises
The most direct way to find the most important nodes in the network
is to look at the degree. The more friends a person has, the more
important she is. This way of measuring importance works well in
many cases, but can miss important information. What if there is a
person with only few friends, but placed in different communities –
just like in Figure 11.1? The removal of such person will create iso-
lated groups, which are now unable to talk to each other. Shouldn’t
this person be considered a key element in the social network, even
with her puny degree?
11.1 Closeness
1
Alex Bavelas. A mathematical model
If we want to know the closeness centrality1 of a node v, first we for group structures. Human organization,
calculate all shortest paths starting from that node to every possible 7(3):16, 1948
0 .8
0.66
0. 5 0.57
11.2 Betweenness
2
Jac M Anthonisse. The rush in a
directed graph. Stichting Mathematisch
Network scientists developed betweenness centrality2 , 3
to fix some of Centrum. Mathematische Besliskunde, (BN
the issues of closeness centrality. Differently from closeness, with be- 9/71), 1971
tweenness we are not counting distances, but paths. We still calculate
3
Linton C Freeman, Douglas Roeder,
and Robert R Mulholland. Centrality in
all shortest paths between all possible pairs of origins and destina- social networks: Ii. experimental results.
tions. Then, if we want to know the betweenness of node v, we count Social networks, 2(2):119–141, 1979
the number of paths passing through v – but of which v is neither an
origin nor a destination. In other words, the number of times v is in
between an origin and a destination. If there is an alternative way of
equal distance to get from the origin to the destination that does not
use v, we discount the contribution of the path passing through v to
v’s betweenness centrality. I provide an example in Figure 11.4.
The total number of paths that can pass through a node – ex-
cluding the ones for which it is the origin or the destination – are
(|V | − 1)(|V | − 2) in a directed network, and (|V | − 1)(|V | − 2)/2 in
an undirected one.
One intuitive way to think about betweenness centrality is asking
yourself: how many paths would become longer if node v would
disappear from the network? How much is the network structure
dependent on v’s presence? Since real world networks have hubs
which are closer to most nodes, the shortest paths will use them
often. As a result, betweenness centrality distributes over many
162 the atlas for the aspiring network scientist
in one clique can reach the nodes in the other one. All those paths
passing through the removed edge are lost forever. The surviving
edge between the two most central ones will have a much reduced
edge betweenness centrality: it cannot be used to move between
cliques any more. This consideration holds true not only for the edge
betweenness, but also for the node betweenness.
A relaxed version of betweenness centrality does not use shortest
paths, but random walks (Chapter 8). This simulates the spreading
of information into the network. The definition is similar: this “flow”
centrality is the number of random walker passing through the node 4
Mark EJ Newman. A measure
during the information spread event4 , 5 . Just like with the regular of betweenness centrality based on
betweenness centrality, also in this case you can take an edge-centric random walks. Social networks, 27(1):
39–54, 2005a
approach, and count the number of random walks going through a 5
Ulrik Brandes and Daniel Fleischer.
specific edge. This has been used, for instance, to solve the problem Centrality measures based on current
of community discovery6 , 7 . flow. In Annual symposium on theoretical
aspects of computer science, pages 533–544.
Springer, 2005
6
Santo Fortunato, Vito Latora, and
11.3 Reach Massimo Marchiori. Method to
find community structures based on
Reach centrality is only defined for directed networks. The local information centrality. Physical review E,
reach centrality of a node v is the fraction of nodes in a network that 70(5):056104, 2004
7
Vito Latora and Massimo Marchiori.
you can reach starting from v8 . From this definition, one can see why Vulnerability and protection of infras-
it doesn’t make much sense for undirected networks. If your network tructure networks. Physical Review E, 71
has a single connected component, then all nodes have the same (1):015103, 2005
8
Enys Mones, Lilla Vicsek, and Tamás
reach centrality, which is equal to one. That is also the case if your Vicsek. Hierarchy measure for complex
directed network has only one strongly connected component. In a networks. PloS one, 7(3):e33799, 2012
strongly connected component there are no “sinks” where paths get
trapped, thus every node can reach any other node.
11.4 Eigenvector
8 7
(a) (b)
calculate a set of PageRanks, one per topic. There are other possible 15
Arda Halu, Raúl J Mondragón, Pietro
ways to define a multilayer PageRank15 . Panzarasa, and Ginestra Bianconi.
Multiplex pagerank. PloS one, 8(10):
e78293, 2013
Katz
16
Leo Katz. A new status index derived
Another popular variant of eigenvector centrality is Katz centrality16 . from sociometric analysis. Psychometrika,
At a philosophical level, the difference between the two is that Katz 18(1):39–43, 1953
says that nodes that are farther away from v should count less when
estimating v’s importance. So it matters whether v is reached at the
first step of the random walk, rather than at the second, or at the
hundredth. For eigenvector centrality when you meet v in a random
walk makes no difference, for Katz it does.
If we were to write the eigenvector centrality not as an eigenvector,
but as a sum, we would end up with something that looks a bit like
this:
∞
ECv = ∑ ∑ ( Ak )uv ,
k =1 u ∈V
∞
KCv = ∑ ∑ αk ( Ak )uv .
k =1 u ∈V
UBIK
A paper of mine presents UBIK, which is the lovechild between Katz
17
Michele Coscia, Giulio Rossetti, Diego
Pennacchioli, Damiano Ceccarelli, and
centrality and the personalized PageRank I presented before17 . The Fosca Giannotti. You know because
weird acronym is short for “you (U) know Because I Know” and we i know: a multidimensional network
approach to human resources problem.
developed it with networks of professional in mind (like Linkedin). In Proceedings of the 2013 IEEE/ACM
Each professional has skills, which allow her to perform her job. International Conference on Advances in
However, sometimes she is required to do something using a skill she Social Networks Analysis and Mining,
pages 434–441. ACM, 2013b
doesn’t have. In many cases, she might be able to perform the task
anyway because she can ask for help in her social network. Think
about any time you asked a friend to fix something in your script, or
scrape some data, or patch a leaking water pipe.
Of course, if the task you need to perform requires only knowl-
edge you have, you can do it quickly. Every level of social interaction
you add will slow you down. If you’re a computer scientist you can
node ranking 167
∞
ECv = ∑ ∑ ( Ak )uv ,
k =1 u ∈V
∞
ECv = (1 − α)ev + ∑ ∑ α( Ak )uv .
k =1 u ∈V
11.5 HITS
22
Jon M Kleinberg, Ravi Kumar, Prab-
HITS22 , 23is an algorithm designed by Jon Kleinberg and collabora- hakar Raghavan, Sridhar Rajagopalan,
and Andrew S Tomkins. The web as
tors to estimate a node’s centrality in a directed network. It is part a graph: measurements, models, and
of the class of eigenvector centrality algorithms from Section 11.4, methods. In International Computing
and Combinatorics Conference, pages 1–17.
but it deserves its own section due to its interesting characteristics. Springer, 1999
Differently from other centrality measures, HITS assigns two values 23
Jon M Kleinberg. Authoritative sources
to each node. In fact, one can say that HITS assigns nodes to one of in a hyperlinked environment. Journal of
the ACM (JACM), 46(5):604–632, 1999
two roles – we will see more node roles in Chapter 12. The two roles
are “hubs” and “authorities”.
of saying that a central node in a large component is more impor- McGraw-Hill Companies, 1976
tant than a central node in a smaller component. Lin’s centrality26
achieves this by multiplying the closeness centrality of a node by the
size of its connected component, which – incidentally – just means to
square the numerator:
11.7 k-Core
3
3
3 3
3
2 2
3
2 2
1 1 2 1 2
1 1 1 1
11.9 Summary
famous examples.
11.10 Exercises
1. Based on the paths you calculated for your answer in the previous
chapter, calculate the closeness centrality of the nodes in Figure
10.12(a).
4. What’s the most central node in the network used for the previous
exercise according to PageRank? How does PageRank compares
with the in-degree? (for instance, you could calculate the Spear-
man and/or Pearson correlation between the two)
6. Based on the paths you calculated for your answer in the previous
chapter, calculate the harmonic centrality of the nodes in Figure
10.12(a).
Not all nodes perform the same role in the network. Sometimes, the
differences between nodes can be estimated quantitatively. A person
is measurably more or less connected in a social network. That is
what we described in the previous chapter: if you can estimate the
importance of the node (number of connections, centrality, etc), you
do so by calculating the corresponding quantitative measure (degree,
betweenness, etc).
Sometimes you cannot put a number to what you’re trying to de-
scribe. What the person is doing in the social network does not have
a quantity, but a quality: she is playing a specific role, which does
not have a countable result. This could be explicitly represented in
your data as a node attribute (Section 4.5): for instance, in a corporate
network, nodes might be explicitly labeled as managers, executive,
technicians, etc.
If you don’t have explicit qualitative data, you might want to put
a label on the nodes based on the structural network data. Rather
than being a characteristic of the person by itself, the node role is
determined by her position in the network.
There are many ways to define node roles, dependent on the
aspect of the network you want to describe. The main split in the
literature is on the type of procedure you’re following: unsupervised
or supervised. In unsupervised role learning, no node in your data
has a role and you’re making up your own definition of roles de-
pending on what’s meaningful to you. This is the classical approach,
which I dissect in Sections 12.1 and 12.2. In supervised node learning,
you already have partially labeled data and you want to figure out
what are the underlying rules determining the node roles. This is the
theme of Section 12.3.
node roles 177
3
Keith Henderson, Brian Gallagher,
Rolx3 is one of the best known computer science approaches for Tina Eliassi-Rad, Hanghang Tong,
the extraction of node roles in complex networks. The way it works Sugato Basu, Leman Akoglu, Danai
Koutra, Christos Faloutsos, and Lei
is by representing nodes as vectors of attributes. Attributes can be,
Li. Rolx: structural role extraction &
for instance, the degree, the local clustering, betweenness centrality, mining in large graphs. In Proceedings
and so on. In practice, you decide which node features are relevant of the 18th ACM SIGKDD international
conference on Knowledge discovery and
to determine the roles your nodes should be playing in the network. data mining, pages 1231–1239. ACM,
This means that, selecting the right set of features, you can recover all 2012
the roles I discussed so far – core, periphery, broker, gatekeeper.
Rolx works in the way you would expect from a standard machine
node roles 179
• Bridges (red): these are the hubs keeping the network together;
• Tightly knit (blue): these are the authors who have a reliable group
of co-authors, and are usually embedded in cliques;
Note that, with Rolx, you can also estimate how much each role
tends to connect with nodes in a similar role. For instance, by their
very nature, bridges tend to connect to nodes with different roles,
while tightly knit nodes band together. This is related to the concepts
of homophily and disassortativity, which we’ll explore in Chapter 26.
Structural Equivalence
When two nodes have the same role in a network they are, in a sense,
similar to each other. Researchers have explored this observation
and derived measures of “node similarity”. We can also call this
180 the atlas for the aspiring network scientist
1 1
Figure 12.5: (a) An example
of two structurally equivalent
nodes (nodes 1 and 2). (b) Here,
3 4 5 6 3 4 5 6 nodes 1 and 2 are not struc-
turally equivalent, because
node 1 has a neighbor that
2 2 node 2 does not.
(a) (b)
This is important to point out in this section, because it helps
understanding the definition of structural equivalence. For two nodes
to be structurally equivalent they have to be connected to the same 4
Robert A Hanneman and Mark Riddle.
neighbors4 . If they do, they are indistinguishable from one another, Introduction to social network methods.
therefore they cannot be any more similar. Consider Figure 12.5(a): 2005
nodes 1 and 2 have the same neighbors and no other additional one.
If I were to flip their IDs, you would not be able to tell. There is
no extra information for you to do so, because all you have is their
neighbor set.
On the other hand, we can tell the difference between nodes 1
and 2 in Figure 12.5(b). That is because we know that node 1 also
connects to node 3, which node 2 does not. So the two nodes are not
structurally equivalent. You can use any vector similarity measure
to estimate structural equivalence. For instance, you can calculate
the Jaccard similarity of the neighbor sets between u and v. In Figure
12.5(b), nodes 1 and 2 have three common neighbors out of four
possible, thus their structural equivalence is 0.75.
Alternatively, one could use cosine similarity, Pearson correlation
coefficients, or inverse Euclidean distance. In all these cases, you
have to transform the neighbor set into a numerical vector. If you sort
the nodes consistently, each node can be represented as a vector of
zeros and ones. Zeros correspond to nodes not connected to u, while
ones are u’s neighbors. These are the rows in the adjacency matrix
corresponding to the nodes, as Figure 12.6 shows. You can input the
node roles 181
Automorphic Equivalence
Automorphic equivalence is a more relaxed version of structural equiv-
alence. To understand it, we need to introduce the concepts of iso-
morphism and automorphism. We call two graphs “isomorphic” if they
have the same topology: the graph in Figures 12.5(a) and 12.5(b)
would be isomorphic if you were to add an edge in Figure 12.5(b)
between nodes 2 and 3. As a mnemonic trick: “iso” = same, and
“morph” = “shape” – two isomorphic graphs have the same shape.
We’ll see how to determine whether two graphs are isomorphic in
Section 39.2.
A graph is “automorphic” if it is isomorphic with itself. This
means that you can shuffle all node IDs of your graph such that you
preserve the neighborhoods of all nodes. If node 1 was connected
only to nodes 2 and 3 in G, you can only swap its ID with another
node that only has two connections and those connections lead to
nodes that have swapped their IDs with nodes 2 and 3. The graph
in Figure 12.7(a) is automorphic because we can swap around labels
respecting this rule – as I do in Figure 12.7(b).
So, for automorphic equivalence, two nodes are equivalent if you
1 2
3 4 5 6 6 4 5 3
Figure 12.7: An example of rela-
beling of the graph to highlight
2 1 the automorphic equivalence
between nodes 1 and 2.
(a) (b)
182 the atlas for the aspiring network scientist
Regular Equivalence
2 3
Figure 12.8: A graph with three
4 5 6
regularly equivalent classes of
nodes.
node roles 183
class one and class three, even if the number of connections to the
members of class three can vary. Class one is the mirror of class three:
it connects to nodes in class two, but to no node in class three.
To sum up the difference between structural, automorphic, and
regular equivalence, consider familial bonds. Two women with the
same husband and the same children are structurally equivalent: they
have the same relationships with the same people (in fact, they would
be the same person – although this is not a necessary requirement
for structural equivalence). Two women can be automorphically
equivalent if they have the same number of husbands and the same
number of children. Finally, to be regularly equivalent to a married
woman with children you have to be a married woman with children,
even if you have a different number of relations – the woman with
fewer husbands will definitely lead a less frustrating life.
There are many other measures of node similarity. Researchers
usually define new ones to better solve a problem called “link predic-
tion”, under the assumption that two similar nodes are more likely
to connect to each other. We will see more node similarity measures
than you want to know in Chapter 20. Node similarity can be used to
estimate network similarity, which is the topic of Chapter 41.
So far we’ve been pretty rigid in the way we wanted to classify nodes
into roles. Either we explicitly defined the roles with strict rules, or
we adopted the similarity approach, finding which node plays a
similar structural role to which other node. This is a sort of “zero-
dimensional” approach, where everything collapses in a single label.
One could use instead node embeddings, which determine the role
of a node with a vector of numbers and then classifies the node with
it. Recently, the most common way to discover such roles has become
the use of graph convolutional techniques. These techniques are not
the only way to go about solving this task, but I’ll focus on them for
the remaining of this chapter.
As an initial extreme simplification, graph convolutional uses
machine learning techniques – and, specifically, neural networks – 5
Jie Zhou, Ganqu Cui, Zhengyan
to learn a function classifying nodes5 , 6 . Graph convolutional is a Zhang, Cheng Yang, Zhiyuan Liu, and
supervised method, meaning that you already have a set of nodes Maosong Sun. Graph neural networks:
A review of methods and applications.
for which the label is known and you’re trying to infer the function arXiv preprint arXiv:1812.08434, 2018
behind this label assignment. In this way, you can assign a role to the 6
Zonghan Wu, Shirui Pan, Fengwen
nodes for which you don’t know the label yet. Chen, Guodong Long, Chengqi Zhang,
and Philip S Yu. A comprehensive
This should not be confused with graph embedding techniques, survey on graph neural networks. arXiv
which perform a similar task, but to learn graph embeddings. The preprint arXiv:1901.00596, 2019
difference between the two is that graph embeddings are then used
184 the atlas for the aspiring network scientist
intersection (node). 17
Mathias Niepert, Mathias Ahmed,
Other possible applications are the generation of a realistic net- and Konstantin Kutzkov. Learning
convolutional neural networks for
work topology (Chapter 15), the prediction of a link (Chapter 20), or graphs. In ICML, pages 2014–2023, 2016
summarizing the graph (Chapter 38). Given that these are not related 18
Yaguang Li, Rose Yu, Cyrus Shahabi,
to node roles, I’ll deal with such applications in the proper chapters. and Yan Liu. Diffusion convolutional
recurrent neural network: Data-driven
traffic forecasting. arXiv preprint
1,2,1,2,1,? 1,2,1,2,1,2 arXiv:1707.01926, 2017b
Figure 12.11: A schema for
spatial-temporal neural net-
works. We have an activation
timeline for each node (here
Learning f
showing only one). The task is
predicting the activation state in
the next timestep.
19
Bing Yu, Haoteng Yin, and Zhanxing
Input Layer Hidden Layers Output Layer Zhu. Spatio-temporal graph convo-
lutional networks: a deep learning
framework for traffic forecasting. In
Moreover, graph neural network techniques are not necessar- IJCAI, pages 3634–3640. AAAI Press,
ily limited to the convolutional approach. For instance, different 2018
20
Sijie Yan, Yuanjun Xiong, and Dahua
approaches have been used also to solve the graph isomorphism Lin. Spatial temporal graph convo-
problem22 , the problem of deciding whether two graphs are the same lutional networks for skeleton-based
action recognition. In Thirty-Second
graph – we’ll see more about this in Chapter 39.
AAAI Conference on Artificial Intelligence,
2018
21
Ashesh Jain, Amir R Zamir, Sil-
12.4 Summary vio Savarese, and Ashutosh Saxena.
Structural-rnn: Deep learning on spatio-
temporal graphs. In Proceedings of the
1. Going beyond node centrality, we can attach to nodes qualitative IEEE Conference on Computer Vision and
roles, rather than quantitative estimations of their importance. Pattern Recognition, pages 5308–5317,
These qualitative roles are dependent on the node’s position in 2016
22
Christopher Morris, Martin Ritzert,
the network topology. Traditionally, this is an “unsupervised” Matthias Fey, William L Hamilton,
learning task, in which you don’t know any node role and you’re Jan Eric Lenssen, Gaurav Rattan, and
Martin Grohe. Weisfeiler and leman
substantially inventing your own definition.
go neural: Higher-order graph neural
networks. In Proceedings of the AAAI
2. There are many ways to define roles, some well known are brokers Conference on Artificial Intelligence,
– in between communities –, gatekeepers – on the border of a volume 33, pages 4602–4609, 2019
12.5 Exercises
Before diving deep into these more complex models, I need to spend
some time with the grandfather of all graph models. It is the family
of network generating processes created by Paul Erdős and Alfréd 2
P Erdős and A Rényi. On random
Rényi in their seminal set of papers2 , 3 , 4 (some credit goes also to graphs. Publicationes Mathematicae
Gilbert5 for a few variants of the model). Debrecen, 6:290–297, 1959
These are simply known colloquially as “Random graphs”. I
3
Paul Erdos and Alfréd Rényi. On the
evolution of random graphs. Publ. Math.
can divide them fundamentally in two categories: Gn,p and Gn,m Inst. Hung. Acad. Sci, 5(1):17–60, 1960
models. The way they work is slightly different, but their results 4
Paul Erdos and Alfred Renyi. On
are mathematically equivalent. The difference is there simply for random matrices. Magyar Tud. Akad. Mat.
Kutató Int. Közl, 8(455-461):1964, 1964
convenience in what you want to fix: Gn,p allows you to define the 5
Edgar N Gilbert. Random graphs. The
probability p that two random nodes will connect, while Gn,m allows Annals of Mathematical Statistics, 30(4):
to fix the number of edges in your final network, m. 1141–1144, 1959
Ok, but... what is a random graph? We’re all familiar to the con-
cept of random number. You toss a die, the result is random. But
what does “random” mean in the context of a graph? What’s ran-
domized here? For this chapter, I will answer these questions assum-
ing uncorrelated random graphs – meaning that you can mentally
replace “random” with “statistical independence”. This is not strictly
speaking necessary: in the same way that you can study the statistics
of correlated coins, you can study correlated random graphs. How-
ever, that would make for a nasty math, and it isn’t super useful for
the aim of this book.
This means that you don’t know anything about the connections that
are already there.
If you’re tired to read a book, this is the perfect occasion for a
physical exercise. Here’s a process you can follow to make your
own random network. First, take a box of buttons and pour it on the
floor. Yup, you heard me: just make a mess. Then take a yarn, cut
some strings, and drop them on the buttons. The buttons are now
your nodes, and the strings the edges. Congratulations! You have a
random network! Time to calculate!
Figure 13.3 depicts the process. Note that throwing coins for each
node pair isn’t exactly the most efficient way to go about generating a
Gn,p – although I invite you to try. A few smart folks determined an 6
Vladimir Batagelj and Ulrik Brandes.
algorithm to generate Gn,p efficiently6 . Efficient generation of large random
Since Gn,m and Gn,p generate graphs with the same properties, it networks. Physical Review E, 71(3):036113,
2005
means that p and m must be related. Since the graphs have n nodes,
we can derive the number of edges (m) from p. p is applied to each
pair of nodes independently. We know how many pairs of nodes the
graph has, which is n(n − 1)/2. Thus we have an easy equation to
derive m from p: the number of edges is the probability of connecting
any random node pair times the number of possible node pairs, or
n ( n − 1)
p = m.
2
This is useful if you use Gn,p but you want to have, more or less,
control on how many edges you’re going to end up with.
By the way this gives you an idea of what’s the typical density
of a random Gn,p graph. The density (see Section 9.1) is the number
of links over the total possible number of links. We just saw that
n ( n − 1)
the number of links in a random graph is p and the total
2
n ( n − 1)
possible number of links is . One divided by the other gives
2
you p. So, if you want to reproduce the sparseness of real world
networks, you can do that at will. Just tune the p parameter to be
exactly the density you want.
In the following sections we explore each property of interest of
random graphs, to see when they model the properties of real world
networks well, and when they don’t. The latter is the starting point
of practically any subsequent graph model developed after Erdős and
Rényi.
random graphs 193
% nodes % nodes
with k = x with k = x
Real World
Figure 13.4: (a) The typical
degree distribution of Gn,p net-
k works. (b) Comparing a Gn,p
p(|V| - 1)
G(n, p)
degree distribution with one
that you would typically get
x x
from a real world network.
(a) (b)
x
Low p p
(a) (b)
13.5 Clustering
p * # possible
edges among
neighbors
k * (k - 1) k * (k - 1)
p*
2 2
Much lower than the
one of real world networks! p Every node has the same
(So it’s also the average cc)
If we say, for simplicity, that k̄(k̄ − 1)/2 = x, then our CCv formula
looks like: CCv = px/x = p. This means that the clustering coeffi-
cient doesn’t depend on any node characteristic, and it’s expected
to be the same – equal to p – for all nodes. Figure 13.7 provides a
graphical version of this derivation.
Compared to real world networks, p is usually a very low value
for the clustering coefficient. In real world networks, it’s more likely
to close a triangle than to establish a link with a node without com-
mon neighbors. Thus the clustering coefficient tends to be higher
than simply the probability of connection. This is a second pain point
of Erdős-Rényi graphs when it comes to explain real world prop-
erties, after the lack of a realistic degree distribution, as we saw in
Section 13.2. Such a low clustering usually implies also the absence of
random graphs 197
13.6 Summary
2. The oldest and most venerable random graph model is the Gn,p
(or Gn,m ) model, where we fix the number of nodes n and the
probability p of connecting a random node pair, and we extract
edges uniformly at random.
5. Random graphs have a short average path length just like real
world networks typically have. However, they have a much lower
clustering coefficient than what you find in the wild.
13.7 Exercises
the largest connected component on the y axis. Can you find the
phase transition?
14.1 Clustering
Before the Stone Age, a caveman society was very simple. You had
tribes living in their own caves. The tribes were very small, they were
families. Everybody knew everyone else in their cave, but between
caves there was almost no communication. Maybe there could have
been one weak link if the two caves were close enough.
This metaphor was the starting point for Watts in developing 1
Duncan J Watts. Networks, dynamics,
his “cavemen” model1 . The cavemen model is part of the family of and the small-world phenomenon.
simple networks (see Section 4.6). It takes two parameters: the cave American Journal of sociology, 105(2):
493–527, 1999
size (Figure 14.1(a)) and the number of caves (Figure 14.1(b)). The
cave size is the number of people living in each cave. A cave is a
clique: as said, everyone in the cave knows every cavemate (Figure
200 the atlas for the aspiring network scientist
nodes.
However, high clustering does not necessarily mean that you are
going to have communities. In fact, a small world model typically
doesn’t have them. The triangles are distributed everywhere uni-
formly in the network. There are no discontinuities in the density,
no differences between denser and sparser areas. This is especially
evident for high k and low p: as I show in Figure 14.5, small world
networks with such parameter combinations just look like odd
snakes without clear groups of densely connected nodes. This is a
precondition to have communities, and so you cannot find them in a
small world model.
5
Jonathan R Cole and Stephen Cole.
tage5 . This is a fundamentally dynamic model. You have an initial Social stratification in science. 1974
condition and then you keep adding one element at a time. Each
element you add does not contribute to the preexisting ones uni-
formly at random, but prefers to contribute to specific older elements,
according to a rule you determine.
The preferential attachment model starts from the assumption
that the rich get richer. For instance, suppose you have one coin and
you invest in the stock market. If you’re lucky, after a while, you
will have another coin. Consider instead somebody who has a lot
of coins. Not only she can match your returns, she can probably do
better, because she can have a diversified portfolio which is resilient
to market shocks and black swans – highly improbable but also
massively impactful events, like the one at the basis of the mortgage 6
Nassim Nicholas Taleb. The black
crisis6 . Moreover, she can probably pay better advisers, and capital – swan: The impact of the highly improbable,
according to Piketty7 – just has better returns at scale. In the time it volume 2. Random house, 2007
takes for you to make a coin, she makes hundreds. Being already rich
7
Thomas Piketty. Capital in the 21st
century. 2014
makes her proportionally richer than you.
1/6
1/2
1/6 Figure 14.8: Adding a new
node with the link selection
1/2 1/2
1/3 model. (a) Pick an existing
1/3
1/6 link at random. (b) Pick one
1/3
of the two nodes connected by
(a) (b) (c) that link. (c) Connect to it. The
probabilities of connecting to
comers have an ever decreasing chance to get the new connections. each node in this process are
This is not the only way to create a cumulative advantage. In fact, the ones floating next to it.
the model has some defects. Preferential attachment requires the
newcoming nodes to have global information about all the existing
nodes’ degree. This might be unrealistic in some cases – you may not
know the number of citations of every paper when you are making
a citation. It is also not a necessary feature to generate a cumulative
advantage. 13
Sergey N Dorogovtsev and Jose FF
An alternative to preferential attachment is link selection13 . In Mendes. Evolution of networks.
link selection, the newcoming node selects a link at random from the Advances in physics, 51(4):1079–1187,
2002
ones that exist in the network (Figure 14.8(a)). Then, it connects with
one of the two nodes connected by that edge – choosing uniformly
at random between the two (Figure 14.8(b)). Cumulative advantage
arises because nodes with more links are more likely to be selected,
thus getting more links on average (Figure 14.8(c)). No matter which
link you select, the central hub is connected to it.
1/4 1 / 12
1/4 1 / 12 Figure 14.9: Adding a new
node with the copying model.
1/4 3/4
(a) Pick an existing node at
1/4 1 / 12
random. (b) Copy one of its
connections. (c) The probabili-
(a) (b) ties of connecting to each node
(c)
in this process are the ones
A third alternative from preferential attachment and link selection floating next to it.
is the copying model. Just like in link selection, the newcoming node
has no information about the network, it just picks something uni-
formly at random. Differently from the link selection model, here it
picks another existing node, rather than a link (Figure 14.9(a)). It then
copies one of its connections (Figure 14.9(b)). You can see again how
it’s more likely to connect to the hub: the hub has more neighbors,
thus it is more likely to select one of its neighbors. Moreover, the
14
Jon M Kleinberg, Ravi Kumar, Prab-
hakar Raghavan, Sridhar Rajagopalan,
neighbor of a hub is likely to be low degree, increasing the chances and Andrew S Tomkins. The web as
of selecting the hub in the copying step (Figure 14.9(c)). The copy- a graph: measurements, models, and
methods. In International Computing
ing model is based on an analogy on how webmasters create new and Combinatorics Conference, pages 1–17.
hyperlinks to pre-existing content on the web14 . Springer, 1999
206 the atlas for the aspiring network scientist
6
Gn,m Figure 14.10: The average short-
5
4 PA
est path length (y axis) for
APL
3
2
increasing number of nodes (x
1 axis) for Gn,m (blue) and prefer-
10 100 1000
ential attachment (red) models,
|V|
with the same average degree.
0
10
Figure 14.11: The age effect in
10-1
the degree distribution of the
-2
10 preferential attachment model.
p(k>=x) “Young”
-3
10 nodes
-4
10
“Old” nodes
10-5
-6
10
0 1 2 3 4 5 6
10 10 10 10 10 10 10
x
network will not contain a single node with degree equal to one.
Thus the clustering in these cumulative advantage models is much
lower than real world networks, and there are no communities –
because everything connects to hubs which make up a single core.
There are some extensions of the model which try to include clus- 18
Petter Holme and Beom Jun Kim.
tering18 . At every step of this model you have a choice. You either Growing scale-free networks with
add a node with its links, or you just add links between existing tunable clustering. Physical review E, 65
(2):026107, 2002
nodes without adding a new one. The probability of taking that step
regulates the clustering coefficient of the network.
However, triangles close randomly, thus we have no communities
just like in the small world model. If we want to look at models
which generate more realistic network data, we have to look at the
ones I discuss in the next chapter.
14.4 Summary
14.5 Exercises
Both the small world and the preferential attachment models are
useful because they give us ideas on how some real world network
properties arise. The small world model tells us that small diameters
happen because a clustered network might have some random short-
cuts. The preferential attachment model tells us that broad degree
distributions arise because of cumulative advantage: having many
links is the best way to attract more links.
Yet, neither of them is able to reproduce all the features of a real
world network. If we want to do so, we have to sacrifice the explana-
tory power of a model. We have to fine tune the model so that we
force it to have the properties we want, regardless of what realistic
process made them emerge in the first place. This is the topic of this
chapter.
The easiest way to ensure that your network will have a broad degree
distribution is to force it to have it. No fancy mechanics, no emerging
properties. You first establish a degree sequence and then you force
each node to pick a value from the sequence as its degree. This 1
Mark EJ Newman. The structure and
simple idea is at the basis of the configuration model1 . In fact, the function of complex networks. SIAM
configuration model is more general than this. You can use it to review, 45(2):167–256, 2003b
match the degree sequence of any real world graph, regardless of the
simplicity or complexity of its actual degree distribution.
The configuration model starts from the assumption that, if we 2
Michael Molloy and Bruce Reed. A
want to preserve the degree distribution, we can take it as an input of critical point for random graphs with
our network generating process. We know exactly how many nodes a given degree sequence. Random
structures & algorithms, 6(2-3):161–180,
have how many edges. So we forget about the actual connections, 1995
and we have a set of nodes with “stubs” that we have to fill in. Figure 3
Mark EJ Newman, Steven H Strogatz,
15.1 shows an example. and Duncan J Watts. Random graphs
with arbitrary degree distributions and
There’s a relatively simple algorithm to generate a configuration their applications. Physical review E, 64
model network, the Molloy-Reed approach2 , 3 . First, as we saw, you (2):026118, 2001
generating realistic data 211
# nodes
Figure 15.1: In a configuration
model, you start from the de-
gree histogram to determine
2 how many nodes have how
4 many open “edge stubs”.
1 1 1
properly high clustering, but these are rare enough that researchers
needed to modify the configuration model to explicitly include the 10
Mark EJ Newman. Random graphs
generation of triangles10 , 11 . with clustering. Physical review letters,
This is usually achieved by generating a joint degree sequence. 103(5):058701, 2009
Rather than simply specifying the degree of each node, we now have
11
Joel C Miller. Percolation and epi-
demics in random clustered networks.
to fix two values. The first is the number of triangles to which the Physical Review E, 80(2):020901, 2009
node belongs. The second is the number of remaining edges the
node has that are not part of a triangle. One can see that we’re still
generating realistic data 213
(a) (b)
15.2 Communities
One could go even deeper and determine that each pair of nodes
u and v can have its own connection probability. This would generate
an input matrix for the SBM that looks like the one in Figure 15.5(a).
Figure 15.5(b) shows a likely result of the SBM using Figure 15.5(a)
as an input. It’s easy to see why we call this matrix “block diagonal”.
The blocks are.... on the diagonal, man.
In one swoop we obtained what we were looking for: both com-
munities and high clustering. The very dense blocks contribute a lot
to the clustering calculation, more than the sparse areas around the
communities can bring the clustering down. One observation we will
come back to is that, in real world networks, the community sizes dis-
tribute broadly, just like the degree: there are few giant communities
and many small ones. This can be embedded in SBM, since we’re free
to determine the input partition as we please.
If you set pout to a relatively high value, you might make your
communities harder to find, but you gain something else: smaller
diameters. You’re also free to set pin < pout in which case you’d find
a disassortative community structure, where nodes tend to dislike
nodes in the same community. See Section 26.2 to know more about
216 the atlas for the aspiring network scientist
GN Benchmark
17
Michelle Girvan and Mark EJ New-
The GN benchmark is a modification of the cavemen graph and one
man. Community structure in social
of the first network models designed to test community discovery and biological networks. Proceedings of
algorithms17 . The first defining characteristic of this model is set- the national academy of sciences, 99(12):
7821–7826, 2002
ting some of the parameters of the cavemen graph as fixed. In the
benchmark, we have only four caves and each cave contains 32 nodes.
Differently from the caveman graph, the caves are not cliques: each
node has an expected degree of 16, thus it can connect at most to half
of its own cave.
(b) µ = 0.0625
(c) µ = 0.125
(a) µ = 0 (d) µ = 0.1875
LFR Benchmark
The LFR Benchmark was developed to serve as a test case for com- 18
Andrea Lancichinetti, Santo Fortu-
munity discovery algorithms18 . The objective is to generate a large nato, and Filippo Radicchi. Benchmark
number of benchmarks to test a new algorithm such that we know graphs for testing community detection
algorithms. Physical review E, 78(4):
the “true” allegiance of each node. Once an algorithm returns us a 046110, 2008
possible node partition, we can compare its solution with the true
communities.
Since you want to have networks with lots of realistic properties,
some of which are difficult to reproduce organically, the LFR bench-
mark takes lots of input parameters. If you want an LFR network,
you have to specify:
• The k̄ average degree of the nodes – you can also set k min and k max
as the minimum and maximum degree, respectively;
As you can see, the LFR assumes that both the degree distribution
and the size of your communities distribute like a power law. The
(d) Step #5
Figure 15.9: A run through a
Note that, in light of step #4, you have some constraints in your simple LFR model.
choice of parameter. Specifically k min < smin and k max < smax ,
otherwise the nodes with minimum and maximum degree will never
belong to any community. Even with such constraints, sometimes
the combination of parameters will require the generation of an
impossible graph, so the LFR benchmark will always be some sort of
approximation of your desires. In practice, differences are going to be
relatively tiny and insignificant.
Since we’re plugging in a power law degree distribution and
communities, it is obvious that LFR benchmarks will reproduce these
characteristics of real world networks well – although now you’re
actually forced to have a power degree distribution, which in some
cases you might not want. They also respect clustering and small
diameters, making them the most realistic model we have.
Kronecker Graphs
u1 u1 v1 u1 v2 u1 v3
u h i u v u2 v2 u2 v3
u ⊗ v = uv> = 2 v1
2 1
v2 v3 = .
u3 u3 v1 u3 v2 u3 v3
u4 u4 v1 u4 v2 u4 v3
xA xA xA xA
Figure 15.10: An example of
xA xA xA xA Kronecker product. (a) A ma-
xA xA xA xA trix A. (b) The operation we
xA xA xA xA perform to obtain A ⊗ A.
(a) (b)
So, in practice, the outer product of two vectors u and v is a
|u| × |v| matrix, whose i, j entry is the multiplication of ui to v j .
The Kronecker product is the same thing, applied to matrices. Fig-
ure 15.10 shows an example. To calculate A ⊗ B, we’re basically
multiplying each entry of A with B20 . 20
G Zehfuss. Über eine gewisse
determinante. Zeitschrift für Mathematik
When it comes to generating graphs, the matrix we’re multiplying und Physik, 3(1858):298–301, 1858
is the adjacency matrix. We usually multiply it with itself. So we’re
calculating A ⊗ A, as I show in Figure 15.10(b). This generates a new
squared matrix, whose size is the square of the previous size. We can
multiply this new adjacency matrix with our original one once more,
for as many times as we want. We stop when we reach the desired
number of nodes21 , 22 .
21
Jurij Leskovec, Deepayan Chakrabarti,
Jon Kleinberg, and Christos Faloutsos.
Figure 15.11 shows the progression of the Kronecker graph. Fig- Realistic, mathematically tractable
ure 15.11(a) is our seed graph which we multiply to itself (Figure graph generation and evolution, using
kronecker multiplication. In European
15.11(b)) twice (Figure 15.11(c)). Conference on Principles of Data Mining
One small adjustment that is customary to do when generating and Knowledge Discovery, pages 133–145.
a Kronecker graph is to fill the diagonal with ones instead of zeros. Springer, 2005b
22
Jure Leskovec and Christos Faloutsos.
If you remember my linear algebra primer, this means we consider Scalable modeling of real graphs using
every node to have a self-loop to itself. This is because we want kronecker multiplication. In Proceedings
the Kronecker graph to be a block-diagonal matrix, with lots of of the 24th international conference on
Machine learning, pages 497–504. ACM,
connections around the diagonal. This is required if we want them to 2007
show a sort of community partition.
By how the Kronecker product is defined you can see that, if
the seed matrix had an empty diagonal, we would not get a block
plane able to fly directly from once city to the other. We can model 23
Mathew Penrose et al. Random
these constraints using random geometric graphs23 , 24 . geometric graphs, volume 5. Oxford
The concept is simple. You first decide the dimensionality of your university press, 2003
space: is it a 2D plane, three dimensional, or n-dimensional? Then
24
Jesper Dall and Michael Christensen.
Random geometric graphs. Physical
you generate |V | points in this space, by extracting them uniformly at review E, 66(1):016121, 2002
random. Finally, you connect two points if they are at r distance – or
less – from each other. Every point will be at distance zero from itself
but, for the sake of simplicity, we simply ignore self-loops. Note that
you are free to decide how to calculate the distance between points:
you’re not forced to use the Euclidean.
One thing that all models discussed so far have in common is that
they are engineered to have specific properties. These are the prop-
erties we think are salient in real world networks: broad degree
distributions, community structures, etc. But what if we are wrong?
Maybe some of these properties are not the most relevant things
about a network we want to model. Moreover: what if there are other
properties that we aren’t seeing? Edges are dependent on each other,
but these dependencies can be complex and it’s difficult to put them
in simple measures we can then optimize. 27
Yujia Li, Oriol Vinyals, Chris Dyer,
The field of graph generative networks27 , 28 , 29 aims at tackling Razvan Pascanu, and Peter Battaglia.
this problem. Here we want to generate networks that look like Learning deep generative models of
graphs. arXiv preprint arXiv:1803.03324,
specific real world networks, without us knowing what “looking like” 2018
actually means. In other words, we want the generative process to 28
Nicola De Cao and Thomas Kipf.
Molgan: An implicit generative model
“learn” how a real world network looks like, so that it can generate for small molecular graphs. arXiv
synthetic versions at will. preprint arXiv:1805.11973, 2018
The most trivial way you can do this is by feeding the adjacency 29
Aleksandar Bojchevski, Oleksandr
Shchur, Daniel Zügner, and Stephan
matrix of your graph – or a suitably modified version of it – to a Günnemann. Netgan: Generating graphs
standard neural network. The neural network will learn the depen- via random walks. In International
dencies between edges. You can think of this approach as a SBM Conference on Machine Learning, pages
609–618, 2018
process without inferring the communities beforehand. SBM wants
to preserve the community structure and, on this basis, learns edge
probabilities that only depend on the community affiliation. Here,
we want to preserve the general network properties, thus each edge
probability is dependent on the entirety of the adjacency matrix.
However, this approach has two problems. First, it will only gener-
ate a graph with the same number of nodes as the input, while you
might want to vary the size of your synthetic networks. Second, it
can only learn from a single graph at a time. Sometimes, you might
want to model a class of graphs.
These limitations are solved in a variety of ways. Just to give an 30
Jiaxuan You, Rex Ying, Xiang Ren,
example, GraphRNN30 allows for two moves: a graph-level update William Hamilton, and Jure Leskovec.
and an edge-level update. In the first step, GraphRNN adds a new Graphrnn: Generating realistic graphs
with deep auto-regressive models. In
node into the network. Every time a new nodes is added, the edge- International Conference on Machine
level update is triggered, determining to which nodes the new node Learning, pages 5694–5703, 2018
connects. This is achieved by representing the graph as a sequence.
For each node, in order, we list to which of the previous nodes it con-
224 the atlas for the aspiring network scientist
[1, 1, 0, 0, 1, 1, 0, 0, 1, 1]
2 Figure 15.15: A graph and its
1, sequence representation in
4 1, 0, GraphRNN. Each element in
1 the sequence belongs to a node
0, 1, 1, (character color) and records
3 5 0, 0, 1, 1 whether that node connects to a
specific node preceding it in the
sequence (underline color).
nects. For instance, the sequence [1, 1, 0, 0, 1, 1, 0, 0, 1, 1] corresponds to
the graph in Figure 15.15.
To see why, it is useful to break down the sequence in sections,
each one referring to a node, as the figure does. The first node has
no element in the sequence, because it has no preceding nodes.
The second node contributes only one element to the sequence: 0
if it doesn’t connect to the first node, 1 otherwise. The third node
contributes two values, one for its edge with the first node and one
for the second node, and so on.
2 2 2 2
4 4
1 1 1 1 1
3 3 3 5
15.6 Exercises
2 4 2 4
(a) (b)
1200 900
Figure 16.2: The edge swap
900
# Null Models
# Null Models
convert from the z-score into a p-value, provided that you know
whether you’re interested in a one-sided or a two-sided test. The
one-sided test means that your success is exclusively on one side of
the distribution – e.g. you want to score more than average, you’re 3
When discussing discoveries in
not interested whether your score is significantly below average3 (or physics, you’ll hear often the term “five
vice versa). sigma” (5σ) thrown around. This means
a z-score equal to 5. In turn, this can
Of course, if the null model distribution is not pseudo-normal,
be converted to a (one-sided) p-value
estimating the statistical significance is a bit trickier. We don’t need to lower than 10−6 , way lower than the
go into that, because we’re about to learn how to perform this task in p < 0.01 you’ll see in other fields. For
p < 0.01, you’re looking at a z-score a
the “right” way, using ERGMs. bit higher than 2.3. I’m simplifying a lot
here, since this is not – and never will
be – a statistics book.
16.2 Exponential Random Graphs
12 9
Au,v ku kv Au,v ku kv
10
4 1 9 4 .92 9 4
2 8
11
1 9 3 .87 9 3
10
7 1 9 3 .87 9 3 12
3
1 1 9 2 .8 9 2 5
11 1 9 2 .8 9 2 1
6
5
9 1 9 2 .8 9 2
6 3
0 9 1 .7 9 1 4
... ... ... ... ... ... 7
8 2
(b) (c)
(a) (d)
Let’s make an example. Figure 16.4(a) represents our observed Figure 16.4: Step-by-step exam-
graph. I generated it with a configuration model with the degree ple of a simple ERGM proces.
sequence (9, 4, 3, 3, 2, 2, 2, 1, 1, 1, 1, 1), but the ERGM doesn’t know (a) Observed network. (b) Ob-
that. The edge presence, the outcome, is our adjacency matrix A. So served edge table (only first
the edge between u and v is Au,v . I now make an hypothesis: the seven rows). Au,v is one if the
degree of a node influences its likelihood of getting a connection. Or, edge is present, zero otherwise;
in mathematical terms, Au,v = β 1 k u + β 2 k v . The degrees of u and k u and k v are the degrees of
v (k u and k v ) can be used to predict the probability of existence of the two nodes. (c) Result of the
an edge. This is equivalent of running a logit regression on an edge logit regression (only first seven
table like the one in Figure 16.4(b): we’re trying to predict the binary rows). Au,v is the estimated
Au,v variable using the degrees of u and v. probability of the edge existing.
Once the logit model is done, for each (u, v) pair we have a proba- (d) An extracted ERGM from
bility of its existence: Au,v (Figure 16.4(c)). We can now flip a loaded the edge probabilities in (c).
coin for each node pair and add the edge in case of success. Au,v is
determined by the β 1 and β 2 parameters. Since they are the result
of a logit model estimation, they are the ones most likely to describe
the family of random graphs from which we extracted the observed
G. By using Figure 16.4(c) to generate a new graph (e.g. the one in
Figure 16.4(d)), we’re sure to extract a graph from the same family
that generated the original one – at least when it comes to its degree
distribution.
So far, I simplified the process for the sake of intuition. For in-
evaluating statistical significance 231
1
Pr ( A = A0 ) = exp p| E0 | ,
B
evaluating statistical significance 233
which is exactly a Gn,p model: a graph whose edges are all equally
likely to be observed. We can also simulate a stochastic blockmodel
by not reducing all connection probabilities to p but by having mul-
tiple ps for each block (and for inter-block connections). If you have
a directed graph you can represent reciprocity with the probability
p1 of a node to reciprocate the connection, adding a term to the Gn,p
model:
1
Pr ( A = A0 ) = exp p| E0 | + p1 R( A0 ) .
B
Here R( A0 ) is the number of reciprocated ties. You can make 8
Emmanuel Lazega and Marijtje
edges dependent on node attributes as researchers do in p2 models8 , 9 . Van Duijn. Position in formal structure,
Finally, you can also plug higher-order structures in the model, for personal characteristics and choices
of advisors in a law firm: A logistic
instance: regression model for dyadic network
data. Social networks, 19(4):375–397, 1997
1
Pr ( A = A0 ) = exp p| E0 | + τT ( A0 ) ,
9
Marijtje AJ Van Duijn, Tom AB
B Snijders, and Bonne JH Zijlstra. p2:
a random effects model with covariates
where – under the homogeneity assumption – T ( A0 ) is the number for directed graphs. Statistica Neerlandica,
of triangles in A0 . This way, you can also control the transitivity of 58(2):234–254, 2004
the graph.
These models can be very difficult to solve analytically for all
but the simplest networks. Modern techniques rely on Monte Carlo 10
Tom AB Snijders. Markov chain monte
maximum likelihood estimation10 , 11 . We don’t need to go too much carlo estimation of exponential random
into details on how these methods work, but these work in the line of graph models. Journal of Social Structure,
3(2):1–40, 2002
any Markov chain Monte Carlo method12 . However, if your network 11
Tom AB Snijders, Philippa E Pattison,
is dense, your estimation might need to take an exponentially large Garry L Robins, and Mark S Handcock.
number of samples to estimate your βs13 . There are ways to get New specifications for exponential
random graph models. Sociological
around this problem by expanding the ERG model14 , but by now methodology, 36(1):99–153, 2006
we’re already way over my head and I don’t think I can characterize 12
Walter R Gilks, Sylvia Richardson,
this fairly. and David Spiegelhalter. Markov chain
Monte Carlo in practice. Chapman and
How does all of this look like in practice? The result of your model Hall/CRC, 1995
might look like something from Figure 16.5. Here we decide to have 13
Shankar Bhamidi, Guy Bresler, and
four parameters: a single edge (this is always going to be present Allan Sly. Mixing time of exponential
random graphs. In 2008 49th Annual
in any ERGM), a chain of three nodes, a star of four nodes, and IEEE Symposium on Foundations of
a triangle. Each motif has a likelihood parameter: the higher the Computer Science, pages 803–812. IEEE,
2008
parameter the more likely the pattern. Negative values mean that the 14
Arun G Chandrasekhar and
pattern is less likely than chance to appear. Matthew O Jackson. Tractable and
The negative value for simple edges means that the network is consistent random graph models.
Technical report, National Bureau of
sparse: two nodes are unlikely to be connected. The positive value Economic Research, 2014
for the triangle means that triangles tend to close: when you have a
triad, it is more likely than chance to have the third edge. The other
two configurations are not significantly different from zero (you can’t
tell because I omitted the standard errors, but trust me on that). Thus
we should not emphasize their interpretation too much.
On the right side of the figure you can see a potential network
234 the atlas for the aspiring network scientist
β
Figure 16.5: On the left we
have the estimated parameters
-4.27
from the observation for four
patterns, with positive values
1.09 indicating a “more than chance”
occurrence of the pattern, and
-0.67 negative values a “less than
chance”. On the right we have
a likely network extracted from
1.32
the set of ERGM with the given
parameters.
15
John F Padgett and Christopher K
that is very likely to be extracted by this ERGM. In fact, I cheated a
Ansell. Robust action and the rise of the
bit, because that is the network on which I fitted the model. It is the medici, 1400-1434. American journal of
famous graph mapping the business relationship between Florentine sociology, 98(6):1259–1319, 1993
16
Skyler J Cranmer and Bruce A
families in the Renaissance15 .
Desmarais. Inferential network analysis
In this chapter I presented only the simplest of the ERGM forms. with exponential random graph models.
Recent research has shifted to more sophisticated models. A few of Political analysis, 19(1):66–86, 2011
those are:
17
Steve Hanneke, Wenjie Fu, Eric P
Xing, et al. Discrete temporal models
of social networks. Electronic Journal of
• Longitudinal ERGMs16 , which are specialized to deal with net- Statistics, 4:585–605, 2010
works that are evolving over time, for instance co-sponsorship of 18
Pavel N Krivitsky and Mark S Hand-
cock. A separable model for dynamic
bills in the US Congress – two representatives might co-sponsor a networks. Journal of the Royal Statistical
bill in one year, but not in another; Society. Series B, Statistical Methodology,
76(1):29, 2014
• Similarly, TERGMs17 introduce the temporal aspect in ERGMS. 19
Bruce A Desmarais and Skyler J
This contains the “separable” TERGMs18 which works on discrete Cranmer. Statistical inference for
valued-edge networks: The generalized
models, rather than modeling the evolution as happening on a exponential random graph model. PloS
continuous time flow; one, 7(1):e30136, 2012
20
Pavel N Krivitsky. Exponential-family
• ERGMs that can take into account edge weights, initially only random graph models for valued
continuous weights19 , but subsequently also discrete ones20 ; networks. Electronic journal of statistics, 6:
1100, 2012
• ERGMs for multilayer networks21 . 21
Alberto Caimo and Isabella Gollini. A
multilayer exponential random graph
modelling approach for weighted
ERGMs have been successfully applied in many fields. For in-
networks. Computational Statistics & Data
stance, they help in cases in which longitudinal network data col- Analysis, 142:106825, 2020
lection is unfeasible – e.g. informal face to face contacts in certain 22
Pierre-Alexandre Balland, José An-
tonio Belso-Martínez, and Andrea
business clusters22 , or inside firms23 . In economics they are partic-
Morrison. The dynamics of technical
ularly useful because of their ability to estimate structural network and business knowledge networks in
parameters, extending conventional analyses that use, for instance, industrial clusters: Embeddedness,
status, or proximity? Economic Geography,
gravity models24 . To make an example, in migration a gravity model 92(1):35–60, 2016
would say that the number of migrants from country u to v is directly 23
Tom Broekel and Matté Hartog.
proportional to the size – in number of inhabitants – of the two coun- Explaining the structure of inter-
organizational networks using exponen-
tries, and inversely proportional to their distance. ERGMs allow you tial random graph models. Industry and
to model more complex interdependencies25 . Innovation, 20(3):277–295, 2013
evaluating statistical significance 235
16.3 Summary 24
Tom Broekel, Pierre-Alexandre
Balland, Martijn Burger, and Frank van
Oort. Modeling knowledge networks
1. Network shuffling is a way to create a null version of your net- in economic geography: a discussion
work, created through edge swapping. In edge swapping, you pick of four methods. The annals of regional
two pairs of connected nodes and you rewire the edges to connect science, 53(2):423–452, 2014
25
Michael Windzio. The network of
nodes from the other pair. global migration 1990–2013: Using
ergms to test theories of migration
2. Once you generate thousands of null versions of your network, between countries. Social Networks, 53:
you can test a property of interest and obtain an indication of 20–29, 2018
how statistically significant your observation is, by counting the
number of standard deviations between the observation and the
null average.
16.4 Exercises
Spreading Processes
17
Epidemics
So far, we’ve seen some dynamics you can embed in your network. In
Section 4.4 I showed you how to model graphs whose edges might
appear and disappear, while in the previous book part we’ve seen
models of network growth: nodes arrive steadily into the network
and we determine rules to connect them such as in the preferential
attachment model. This part deals with another type of dynamics on
networks. Here, edges don’t change, but nodes can transition into
different states.
The easiest metaphor to understand these processes is disease.
Normally, people are healthy: their bodies are in a homeostatic state
and they go about their merry day. However, they also constantly
enter into contact with pathogens. Most of the times, their immune
systems are competent enough to fend off the invasion. Sometimes
this does not happen. The person transitions into a different state:
they now are sick. Sickness might be permanent, but also temporary.
People can recover from most diseases. In some cases, recovery is
permanent, in others it isn’t.
These are all different states in which any individual might find
themselves at any given time. Like individuals, nodes too can change
state as time goes on. This book part will teach you all the possible
models we have to study these state transitions. In this chapter
we look at three models we defined to study the progression of 1
Romualdo Pastor-Satorras, Claudio
diseases through social networks1 . Note that such models can easily Castellano, Piet Van Mieghem, and
represent other forms of contagion, for instance the spread of viruses Alessandro Vespignani. Epidemic
processes in complex networks. Reviews
in computer and mobile networks2 . of modern physics, 87(3):925, 2015
We’re going to complicate these models in Chapter 18, to see how 2
Pu Wang, Marta C González, César A
different criteria for passing the diseases between friends affect the Hidalgo, and Albert-László Barabási.
Understanding the spreading patterns
final results. Then, in Chapter 19, we’ll see how the same model of mobile phone viruses. Science, 324
can be adapted to describe other network events, such as infrastruc- (5930):1071–1076, 2009
ture failure and word-of-mouth systems to aid a viral marketing
campaign.
Another complication is the one introduced by simplicial spread-
238 the atlas for the aspiring network scientist
17.1 SI
The only thing that can happen in this model is the transition from
the Susceptible to the Infected state: S → I. In this world, the only
possible action is for a healthy person to contract the disease. Noth-
ing else is allowed.
Given that there are only two states (S and I) and only one transi-
tion (S → I), we call this the SI Model. Figure 17.1 shows the schema
fully defining the model. In practice, SI models diseases with no
recovery. An example would be some variants of the herpes virus.
Love goes by, herpes is forever.
There is one assumption underlying the traditional SI Model:
homogenous mixing – keep this in mind because it’s important. In
homogenous mixing, we assume that each susceptible individual has
the same probability to come into contact with an infected person.
This is simply determined by the current fraction of the population in
the infected state. Once the susceptible individual meets an infected,
there is a probability that they will transition into the I state too. This
probability is a parameter of the model, traditionally indicated by β.
If β = 1, any contact with an infected will transmit the disease, while
if β = 0.2, you have an 20% chance to contract the disease.
1
0.9 Figure 17.2: The solution of the
0.8
SI Model for different β values.
Infected Ratio
0.7
0.6
β = 0.175 The plot reports on the y axis
0.5
0.4 β = 0.15
0.3 the share of infected individu-
β = 0.125
0.2
0.1
als (i = | I |/(| I | + |S|)) at a given
β = 0.1
0
0 20 40 60 80 100 120 140 160 180 200
time step (x axis).
Time
Once you have β you can solve the SI Model. Usually, the way
it’s done is assuming that at the first time step you have a set of one
or more patient zeros scattered randomly into society. Then, you
track the ratio of people in the I status as time goes on, which is
| I |/(| I | + |S|). This usually generates a plot like the one in Figure 17.2.
SI models have the same signature. At first, the ratio of infected
individuals grows slowly, because there are few people in the I state.
Then, as soon as I expands a little, we see an exponential growth, as
more and more people have a chance to meet an infected individual.
After a critical point, the growth of I slows down, because there
aren’t many people left in S to infect.
Eventually, all SI models stop when every single individual is in
the set I and so no one else can transition. All SI Models, no matter
the value of β will end up with a complete infection, where S is
empty and I contains the entirety of society. The only thing β affects
240 the atlas for the aspiring network scientist
– as you can see from Figure 17.2 – is the speed of the system: when
the exponential growth of I starts to kick in and when S gets emptied
out.
We can re-tell the story I’ve just exposed in mathematical form.
In our SI model, the probability that an infected individual meets a
susceptible one is simply the number of susceptible individuals over
the total population, because of the homogenous mixing hypothesis:
|S|/|V | (remember |V | is our number of nodes). There are | I | infected
individuals, each with k̄ meetings (the average degree). Thus the total
| I ||S|
number of meetings is k̄ . Since each meeting has a probability
|V |
| I ||S|
β of passing the disease, at each time step there are βk̄ new
|V |
infected people in I.
We can simplify the equation a bit, because | I |/|V | and |S|/|V | are
related. They sum to one, since S and I are the only possible states
in which you can have a node. So, if we say i = | I |/|V |, that is, the
fraction of nodes in I, then |S|/|V | = 1 − i. So our formula becomes: 8
Note that, for simplicity, I’m only
it+1 = βk̄it (1 − it )8 , where t is the current time step. If we integrate including the addition to it+1 , not
over time, we can derive the fraction of infected nodes depending its full composition. So, pedantically,
the correct formula should be it+1 =
solely on the time step9 :
it + βk̄it (1 − it ), but that would make
the discussion harder to follow. This
i0 e βk̄t warning applies to all formulas with the
i= . time subscript.
1 − i0 + i0 e βk̄t 9
Albert-László Barabási et al. Network
This is the mathematical solution to the SI model with homoge- science. Cambridge university press, 2016
nous mixing, generating the plot in Figure 17.2. You can see why you
have an initial exponential growth at the beginning and a flat growth
at the end. If i0 ∼ 0, then the denominator is 1 and the numerator is
dominated by the e βk̄t factor: exponential growth (very slow at the
beginning because multiplied with the small i0 ). When i0 ∼ 1, both
the denominator and the numerator reduce to e βk̄t , which means that,
in the end, i ∼ i0 ∼ 1, so there’s no growth.
Why did we go to the trouble of all this math? Because, at this
point, we have to tear down the homogenous mixing hypothesis. The
formulas will allow to see the difference better.
Homogenous mixing is based on the assumption that the more
people are infected, the more likely you’re going to be infected. In
practice, it assumes everybody is the same. In homogenous mixing,
the global social network is a lattice: a regular grid where each node
is connected only to its immediate neighbors. Figure 17.3 shows an
example of square lattice (Section 4.6): each node connects regularly
to four spatial neighbors. On a lattice, the infection spreads like water
10
This is a useful mental image: https:
filling a surface10 . //[Link]/wikipedia/
We know that real networks are not neat regular lattices. The commons/a/a6/SIR_model_simulated_
using_python.gif.
degree is distributed unevenly, with hubs having thousands of con-
epidemics 241
nections – see Section 6.3. When the infection hits such a hub, it will
accelerate faster through the network. In fact, it is extremely easy to
infect an hub early on. Hubs have more connections, thus they are
more likely to be connected to one of your patient zeros. Those same
connections make them super-spreaders: once infected, the hub will
allow the disease to reach the rest of the network quickly. In fact,
when searching information in a peer-to-peer network, your best 11
Lada A Adamic, Rajan M Lukose,
guess is always to ask your neighbor with highest degree11 . Amit R Puniyani, and Bernardo A
To treat the SI model mathematically you have to first group nodes Huberman. Search in power-law
networks. Physical review E, 64(4):046135,
by their degree. Rather than solving for i – the fraction of infected
2001
nodes –, you solve for ik : the fraction of infected nodes of degree k.
The formula for a network-aware SI model is similar as the one we
saw for the vanilla SI model:
ik,t+1 = βk f k (1 − ik,t ).
The two differences are that: (i) we replace the average degree k̄
with the actual node’s degree k, and (ii) rather than using ik,t we use
f k – a function of the degree k. This is because real world networks
typically have degree correlations: if you have a degree k the degree
of your neighbors is usually not random (see Section 27.1 for more).
If it were random, then we could simply use ik,t , because the number
of infected individuals around you should be proportional to the
current infection rate. But it isn’t: in presence of degree correlations,
if you have k neighbors then there exists a function f k able to predict
how many neighbors they have. Thus the likelihood of a node of
degree k of having infected neighbors is specific to its degree, and not
(only) dependent on ik,t . 12
Romualdo Pastor-Satorras and
If you do the proper derivations12 , you’ll discover that in a Gn,p Alessandro Vespignani. Epidemic dy-
network the dynamics have the same functional form to the ones of namics and endemic states in complex
networks. Physical Review E, 63(6):066117,
the homogeneous mixing, as Figure 17.4 shows. In Gn,p the exponen- 2001a
tial rises faster at the beginning – due to the few outliers with high
degree – and tails off slower at the end – due to the outliers with
low degree – but the rising and falling of the infection rates is still an
242 the atlas for the aspiring network scientist
1
0.9 Figure 17.4: The solution of the
0.8
SI Model for different β values
Infected Ratio
0.7
0.6
β = 0.175 (no net) β = 0.175 (Gn,p)
0.5 in homogeneous mixing (reds)
0.4 β = 0.15 (no net) β = 0.15 (Gn,p)
0.3 and Gn,p graphs (blues). The
β = 0.125 (no net) β = 0.125 (Gn,p)
0.2
0.1 β = 0.1 (Gn,p)
plot reports on the y axis the
β = 0.1 (no net)
0
0 50 100 150 200 250 300
share of infected individuals
Time (i = | I |/(| I | + |S|)) at a given
time step (x axis).
1
0.9 Figure 17.5: The solution of the
β = 0.175 (α = 2) β = 0.125 (α = 3)
0.8
SI Model for different β values
Infected Ratio
exponential warm up any more (in red in Figure 17.5). You know
that, no matter where you started, you’re going to hit the largest hub
of the network at the second time step t = 2, because it is connected
to practically every node. And, since it is connected to practically
every node, at t = 3 you’ll have almost the entire network infected.
In fact, I ran the simulations from Figure 17.5 on imperfect and
finite power law models. Theoretically, if you had a perfect infinite
power law network, infection would be instantaneous for any non-zero
value of β. Meaning that, no matter how infectious a disease is, with
α = 2 it will infect the entire network almost immediately. And
things get even more complicated when you add to the mix the fact 13
Eugenio Valdano, Luca Ferreri, Chiara
that networks evolve over time13 . Scary thought, isn’t it? Poletto, and Vittoria Colizza. Analytical
computation of the epidemic threshold
on temporal networks. Physical Review X,
17.2 SIS 5(2):021005, 2015
Just like in the SI model, also in the SIS model nodes can either be 14
Herbert W Hethcote. Three basic
Susceptible or Infected14 . However, the SIS model adds a transi- epidemiological models. In Applied
tion. Where in SI you could only get infected without possibility of mathematical ecology, pages 119–144.
Springer, 1989
recovery (S → I), in SIS you can heal (I → S).
Thus the SIS model requires a new parameter. The first one,
shared with SI, is β: the probability that you will contract the disease
after meeting an infected individual. Once you’re infected, you also
have a recovery rate: µ. µ is the probability that you will transition
from I to S at each time step. High values of µ mean that recovering
from the disease is fast and easy. Note that recovery puts you back
to the Susceptible state, thus you can catch the disease again in the
future.
S I
Figure 17.6 shows the schema fully defining the model. In practice,
SIS models disease with recovery and relapse. An example would
be the general umbrella of the flu family. Once you heal from a
particular strain of the flu you’re unlikely to fall ill again under the
244 the atlas for the aspiring network scientist
same strain. However, you can easily catch a similar strain, thus
cycling each year between the S and I states.
The presence of µ changes the outcome of the model. SI models
always reach full saturation: eventually, every node will end up in
status I. For SIS models that is not true, because a certain fraction
of nodes – µ – heal at each time step. The interplay between the
recovery rate µ, the infection rate β, and the average degree k̄ deter-
mines the asymptotic size of I: the share of infected nodes as time
approaches infinity (t → ∞). To see how, let’s look at the math again.
% Infected
Figure 17.7: The typical evolu-
Depends on μ and β
tion of an endemic SIS model:
< the equilibrium state is the one
100%
in which a constant fraction
Endemic state i < 1 contracted the disease.
The rate at which infected peo-
ple recover and the infection
rate are perfectly balanced, as
all things should be.
Time
100
-2
10
and k¯2 grows relatively to it, the critical threshold k̄/k¯2 tends to zero.
If we say that you have an endemic value if λ > k̄/k¯2 and k̄/k¯2 = 0,
then any disease, no matter β and µ, will be endemic in a network
with a power law degree distribution. Oops.
% Infected
at time ∞ Figure 17.9: The solution of the
SIS Model for λ. As λ grows
(x axis), I show the share of
infected individuals i at the
λ=β/μ endemic state (t → ∞).
Power Law
Random
17.3 SIR
S I R
had the disease and healed. Figure 17.11 shows such a typical evolu-
tion. At the beginning, everybody is susceptible. Then, people start
getting infected, so I grows. R cannot start growing immediately,
as I is still too small for the recovery parameter µ to significantly
contribute to R size. As I grows, though, there are enough infected
individuals that start being removed. Eventually every I individual
transitions to R.
% of Pop
Figure 17.11: The typical evolu-
Susceptible Removed tion of an SIR model: after an
initial exponential growth of
the infected, the removed ratio
takes over until it occupies the
entire network.
Infected
Time
least an I neighbor.
4. SIS models are like SI models, but nodes can transition back to S
state with a stochastic probability µ at each time step.
17.5 Exercises
4. Extend your SI model to an SIR. With β = 0.2, run the model for
400 steps with µ values of 0.01, 0.02, and 0.04 and plot the share of
nodes in the Removed state for both the networks used in Q1 and
Q2. How quickly does it converge to a full R state network?
18
Complex Contagion
You may or may not have noticed that, in the previous chapter, all
our models of epidemic contagion shared an assumption. Every time
a susceptible individual comes in contact with an infected individual,
they have a chance to become infected as well. If that doesn’t happen,
the healthy person is still in the susceptible pool. The next time step
represents a new occasion for them to contract the disease. And so
on, ad infinitum.
Without that assumption, the models wouldn’t be mathematically
tractable. For instance, if each node gets only one chance to be in-
fected, you can easily see how it is not given that a SI model would
eventually infect the entire network. In fact, it takes any β < 1 to
make that impossible. The first time you fail to infect somebody you
won’t get the chance to try again.
SI, SIS, and SIR models are useful and generated tons of great
insights. But this limitation allows them to model only rather specific
types of outbreak. We usually consider them models of simple
contagion. There are fundamentally two ways to make such models
more complex and realistic. They involve changing two things: (i) the
triggering mechanism, which is the condition regulating the S → I
transition, and (ii) the assumption that each individual gets infinite
chances to infect their neighbors.
We deal with the triggering mechanisms in Section 18.1 and
infection chances in Section 18.2. We also explore the possibilities of
interfering with the outbreak in Section 18.3, dedicated to epidemic
interventions.
18.1 Triggers
Classical
In classical reinforcement you have an independent probability of
being infected for each of your neighbors that are infected. Note that
this is different from the simple contagion of Chapter 17: in there,
you get β chance to transition regardless whether you have one or
more infected neighbors. Here, more infected neighbors mean more
chances of infection.
If you have n sick friends, and you visit them one by one, at each
visit you toss a coin. To calculate the probability you are going to be
infected, it is easier to calculate the probability of not being infected
by any contact, and then invert it. If our parameter β tells us the
probability of being infected by a single contact, then (1 − β) is
the probability of not being infected. Since the coin tosses are all
independent, the probability of never being infected by any of the n
contacts is (1 − β)n . So the probability that at least one contact will
infect us is 1 − ((1 − β)n ).
Figure 18.2 shows a vignette of this process. The healthy individ-
ual has four neighbors, all of which are infected. Thus she has to
make four independent coin tosses, each of which has β chance to
succeed. Thus, the more infected neighbors the more likely she will
contract the disease. The difference with a simple SI model without
reinforcement is that in the simple SI model you always toss a sin-
252 the atlas for the aspiring network scientist
gle coin at each time step, no matter how many infected neighbors
you have – as long as you have at least one. So, at each time step,
in simple SI the infection probability is β if you have 1 or n infected
neighbors. In classical complex SI, you have 1 − ((1 − β)n ) probability
of being infected. The whole difference between the two models is
that the latter depends on n, the number of your friends that are
infected.
The vignette makes clear why, in the classical model, it’s easy to
infect hubs: they have more neighbors. More neighbors mean that
they toss their coins much more often. This is what generates the
super-exponential – theoretically instantaneous – outbreak growth in
power law models with large hubs.
Threshold
Cascade
rules and they are just different phenomena. This needs not to be the
case. The separation is mostly done out of a pedagogical need. In
fact, there is a universal model of spreading dynamics concentrating 6
Baruch Barzel and Albert-László
on hubs6 . We don’t need to delve deep into the details on how this Barabási. Universality in network
model works, but it mostly hinges on a parameter: φ. φ determines dynamics. Nature physics, 9(10):673,
2013b
the interplay between the degree of a node and its propensity of be-
ing part of the epidemics. If φ = 0, the likelihood of contagion of the
node is independent with its degree. If φ > 0 we are in the threshold
scenario: hubs have a stronger impact on the network. With φ < 0, as
you might expect, the opposite holds. Figure 18.4 shows a vignette of
the model.
Sprinkling a bit of economics into the mix, you can relate the
threshold or the cascade parameter with the utility an actor v gets
from playing along or not. Each individual calculates their cost and
benefit from undertaking or not undertaking an action. There is a
cost in adopting a behavior before it gets popular, and in not doing 7
Thomas C Schelling. Hockey helmets,
so after it did7 . Being aware of these effects makes for very effective concealed weapons, and daylight
strategies to make your own decisions while you’re in doubt. You saving: A study of binary choices with
externalities. Journal of Conflict resolution,
can establish a Schelling point which determines whether or not
17(3):381–428, 1973
you’re going to undertake an action8 , which effectively means you 8
[Link]
consciously set your own κv . However, this is getting dangerously posts/Kbm6QnJv9dgWsPHQP/
schelling-fences-on-slippery-slopes
close to a weird blend of economics, philosophy, and game theory. If
you’re interested in learn more, you’d be best served by closing this 9
Herbert Gintis. The bounds of reason:
book and looking elsewhere9 . Game theory and the unification of the
behavioral sciences. Princeton University
Press, 2014
18.2 Limited Infection Chances
you want them to infect themselves with the idea that the product is 10
Pedro Domingos and Matt Richard-
good10 , 11 . son. Mining the network value of
The obvious strategy would be to target hubs, since they have customers. In Proceedings of the seventh
ACM SIGKDD international conference
more connections. However, this heavily depends on your triggering
on Knowledge discovery and data mining,
model, and hubs come with a disadvantage. First, by being promi- pages 57–66. ACM, 2001
nent, hubs are targeted by many things, thus they have a very high 11
Dashun Wang, Zhen Wen, Hanghang
Tong, Ching-Yung Lin, Chaoming Song,
barrier to attention. Second, they have many connections: if the trig-
and Albert-László Barabási. Information
gering mechanism requires reinforcement, most of their connections spreading in context. In Proceedings of
might not get it, thinning out the intervention. A third and final prob- the 20th international conference on World
wide web, pages 735–744. ACM, 2011b
lem might be that you have only one shot at convincing a person. If
you fail, it’s game over forever. If a hub fails, you might not have a
second shot to get to all their peripheral nodes.
persuade me, just not you. This is a crucial difference with regard to
SI models. We know that any SI model will eventually fill the entire 12
Jacob Goldenberg, Barak Libai, and
network. The independent cascade model12 won’t: the nodes we Eitan Muller. Using complex systems
choose to start the infection with are very important to maximize the analysis to advance marketing theory
development: Modeling heterogeneity
reach of our message.
effects on new product growth through
In the simplest model, node u has a probability pu,v of convincing stochastic cellular automata. Academy of
node v. However, the past history of attempts to convince v might Marketing Science Review, 9(3):1–18, 2001
influence this probability, that’s why you should get to hubs when
you’re the most sure you’re going to convince them. So we can
modify that probability as pu,v (S), with S being the set of nodes 13
David Kempe, Jon Kleinberg, and
who already tried to influence v13 , 14 , 15 . The process ends when all Éva Tardos. Maximizing the spread of
infected nodes exhausted all their chances of convincing people so no influence through a social network. In
Proceedings of the ninth ACM SIGKDD
more moves can happen.
international conference on Knowledge
So you get the problem: find the set of cascade initiators I0 such discovery and data mining, pages 137–146.
that, when the infection process ends at time t, the share of infected ACM, 2003
14
David Kempe, Jon Kleinberg, and
nodes in the network it is maximized. Kempe et al. solve the prob- Éva Tardos. Influential nodes in a
lem with a greedy algorithm. We start from an empty I0 . Then we diffusion model for social networks. In
calculate for each node its marginal utility to the cascade. We add International Colloquium on Automata,
Languages, and Programming, pages
the node with the largest utility, meaning the number of potential 1127–1138. Springer, 2005
infected nodes, to I0 and we repeat until we reach the size we can 15
Note that, if we only use pu,v we
afford to infect. Of course, each node we add to I0 changes the ex- call this independent cascade model,
because the previous attempts do not
pected utility of each other node, because they might have common influence future attempts. When we
friends, thus we cannot simply choose the | I0 | nodes with the largest introduce pu,v (S) the cascades are not
independent any more. Specifically, for
initial utility.
the paper I’m citing, we have decreasing
There are many improvements for this algorithm, focused on cascades because, the more people
improving time efficiency, lowering the expected error, and integrat- try, the hardest it is to convince v,
i.e. pu,v (S) < pu,v (S ∪ z). If we did
ing different utility functions. However, things get more interesting the opposite, pu,v (S) > pu,v (S ∪ z),
when you start adding metadata to your network. For instance, Gu- then this model would be practically
equivalent to the threshold model: the
rumine16 is a system that lets you create influence graphs, as I show
more infected neighbors you have, the
in Figure 18.6. You start from a social network (Figure 18.6(a)) and a more likely you’re going to turn.
table of actions (Figure 18.6(b)). You know when a node did what. 16
Amit Goyal, Francesco Bonchi, and
Laks VS Lakshmanan. Learning influ-
You can use the data to infer that node v does action a1 regularly
ence probabilities in social networks. In
after node u performed the same action. In the example, for two Proceedings of the third ACM international
actions a1 and a2 you see node 2 repeating immediately after node conference on Web search and data mining,
pages 241–250. ACM, 2010
1. Since these two nodes are connected, maybe node 1 is influencing
node 2. You can use that to infer pu,v = 0.66 (Figure 18.6(c)) – or, if
you’re really gallant, to infer pu,v (S) by looking at all neighbors of v
performing a1 before it.
Note that node 6 performed the same action at the same time as
node 3. Node 6 could only be influenced by node 2. For node 3 we 17
Eytan Bakshy, Jake M Hofman,
Winter A Mason, and Duncan J Watts.
prefer inferring that node 1 did it, because we know that it influenced Everyone’s an influencer: quantifying
node 2 too, so that’s the most parsimonious hypothesis. The size in influence on twitter. In Proceedings of
number of nodes of these cascades can be approximated by – you the fourth ACM international conference on
Web search and data mining, pages 65–74.
guessed it – a power law17 . ACM, 2011
258 the atlas for the aspiring network scientist
3 1
1 Time Node Action Figure 18.6: (a) The underlying
1 1 a1 social network. (b) The actions
0.66 0.33
nodes made. (c) A possible
2 2 1 a2
4 inferred influence graph.
2 2 a1
2 3
6 3 1 a3
5 3 2 a2
7 0.5
3 3 a1
9 3 6 a1
8 6
(b)
(a) (c)
When running such models on real data you can find funny
things. For instance, I ran it with some co-authors on LastFm data,
a social network recording which user listened to which musical 18
Diego Pennacchioli, Giulio Rossetti,
artist at which time18 – the artist is considered the “action”. In doing Luca Pappalardo, Dino Pedreschi, Fosca
so, we discovered that we could build these influence graphs and Giannotti, and Michele Coscia. The
three dimensions of social prominence.
describe their trade offs. For instance, the more intensely a user was
In International Conference on Social
influenced by a prominent friend – meaning that they listened the Informatics, pages 319–332. Springer,
new artist a lot – the fewer friends the influencer hit. In other words: 2013
the stronger you want to influence people, the fewer people you can
influence.
the speed with which it triggers other people19 , 20 . For instance, in Dow, Jon Michael Kleinberg, and Jure
Figure 18.7 we have two hypothetical cascades with the same number Leskovec. Can cascades be predicted? In
WWW, pages 925–936. ACM, 2014
of shares and the same topological pattern: one re-sharer then two. 20
Justin Cheng, Lada A Adamic, Jon M
However, the fact that the first cascade happened faster is enough for Kleinberg, and Jure Leskovec. Do
us to infer that it’s much more likely to end up being much larger cascades recur? In WWW, pages 671–681.
ACM, 2016
than the second, slower, one. I’m going to explore more in depth this
complex contagion 259
.1 5
.06
.1 1
.08
% Infected
No intervention
100%
x%
With immunization
% Infected
Figure 18.10: The second crite-
100%
rion of immunization success:
No intervention
temporary immunity can delay
propagation.
With immunization
Time
tion for each infected neighbor. In the threshold model you need
at least κ neighbors, independently of your degree. In the cascade
model you need at least a fraction of neighbors.
4. You can estimate which type of infection model a real world out-
break follows by estimating the universality class of the spreading,
via its parameter φ: φ > 0 is similar to a threshold model (positive
correlation between degree and chance of infection), φ < 0 is sim-
ilar to a cascade model (negative correlation between degree and
chance of infection).
18.6 Exercises
% nodes in LCC
Figure 19.2: The probability of
being part of the largest con-
nected component as a function
of the number of failing nodes
in a random Gn,p graph.
x
|R|
ing about the probability of a node being part of the GCC in a Gn,p
model. In that case, the function on the x-axis was the probability
p of establishing an edge between two nodes. In fact, the two are
practically equivalent: if you have a Gn,p graph with failures it is as if
you’re manipulating n and p.
What Figure 19.2 says is that a Gn,p network will withstand small
failures: a few nodes in R will not break the network apart. However,
the failure will start to become serious very quickly, until we reach a
critical value of | R| beyond which the GCC disappears and the net-
work effectively breaks down. Just like the appearance of a GCC for
increasing p in a Gn,p model is not a gradual process, so are random
failures. At some point there is a phase transition, from having to not
having a GCC.
Ah – you say – but we’re not amateurs at this. Who would engineer
a random power grid network? For sure it won’t be a Gn,p graph. Good
point. In fact that’s true: the power grid’s degree distribution is
skewed. For the sake of the argument – and the simplicity of the
math – let’s check the resilience to random failures of a network with
a power law degree distribution.
Good news everybody! Power law random networks are more
resilient than Gn,p networks to random failures. The typical signature
of a power law network under random node disappearances looks
something like Figure 19.3. In the figure you see no trace of the phase
transition. The critical value under which the GCC disappears is
much higher than in the Gn,p case. Of course the size of the largest
connected component goes down, because you’re removing nodes
from the network. However, the nodes that remain in the network
still tend to be able to communicate to each other, even for very high
| R |.
Why would that be the case? The reason is always the power law
degree distribution. If you remember Section 6.3, having a heavy
tailed degree distribution means to have very few gigantic hubs and
268 the atlas for the aspiring network scientist
% nodes in LCC
Figure 19.3: The probability of
being part of the largest con-
nected component as a function
of the number of failing nodes
in a network with a skewed
degree distribution.
|R|
% nodes in LCC
Figure 19.4: The probability of
being part of the largest con-
nected component as a function
of the number of failing nodes
in a network for different α ex-
α=2
ponents of its power law degree
α=4 distribution.
|R|
1
Random
0.9 Targeted Figure 19.5: The probability
0.8
of being part of the largest
Nodes in GCC
0.7
0.6
0.5 connected component as a func-
0.4
0.3 tion of the number of failing
0.2
0.1
nodes in a Gn,p network, for
0
1 10 100
random (blue) and targeted
|R| (red) failures.
1
Random
0.9 Targeted Figure 19.6: The probability of
0.8
being part of the largest con-
Nodes in GCC
0.7
0.6
0.5 nected component as a function
0.4
0.3 of the number of failing nodes
0.2
0.1
in a degree skewed network,
0
1 10 100 1000
for random (blue) and targeted
|R| (red) failures.
0.7
0.6
0.5 connected component as a func-
0.4
0.3 tion of the number of failing
0.2
0.1
nodes in a degree skewed net-
0
1 10 100 1000
work, for different α and k min
|R| combinations.
and the capacity is the maximum amount that it can pass before
congestion happens.
At time t = 1 we shut down a node in the network. Maybe the
traffic light failed and so no one can pass through until we repaired
it. This means that the node transitions to state I. People still need
to do their errands, so we have to redistribute the load of cars that
wanted to pass through that intersection through alternative routes:
the neighbors of that node. However, that means that their load
will increase. If the new load exceeds the capacity of the node, also
this node shuts down due to congestion. So its load has also to be
redistributed to its neighbors and so on and so forth.
Figure 19.8 shows an example of failure propagation with this
load-capacity feature. You can see that the network was built with
some slack in mind: its normal total load is 37 – the sum of all loads
of all nodes – for a maximum capacity of 90 – the sum of all nodes’
capacities. Yet, shutting down the top node whose load was only 6
and redistributing the loads causes a cascade that, eventually, brings
272 the atlas for the aspiring network scientist
6/10 0 /10 0/ 1 0 0/ 1 0
1 / 10 1 / 10 4 /1 0 4/ 10
6/7 4 /1 6 6/7 4 /1 6 8/7 8/16 0/7 10/16
4 / 10 11/10 0/ 1 0 0/ 10
0/7 23/16 0/7 0/ 1 6 0/7 0/16 0/7 0/16
Using this perspective has its own advantages. It makes the failure
propagation model more amenable to analysis. The final size of the
failure cascade depends on the average degree of nodes in this tree
k̄. The critical value here is k̄ = 1. If, on average, the failure of a node
generates another node failure – or more – the cascade will propagate
indefinitely, until all nodes in the network will fail. If, instead, k̄ < 1,
the failure will die out, often rather quickly.
It’s easy to see why if you have the mental picture of a domino
snake: each domino falling will cause the fall of another domino,
until there’s nothing standing. If, however, there is as much as a
single gap in this chain, the rest of the system will be unaffected.
Quick show of hands: how many of you expect the size of a failure
catastrophic failures 273
X X X
X X X
% nodes in LCC
Figure 19.12: The relationship
Gn,p
α=3
between the degree exponent α
α = 2.7 of coupled power law networks
α = 2.3 and the fraction of nodes in R
state (x axis) needed to destroy
the GCC (shades of red). In
blue the equivalent plot for
coupled random Gn,p graphs.
|R|
3. When one node fails, all its load needs to be redistributed to non-
failing nodes. This can and will make the failure propagate on
the network in a cascade event which might end up bringing the
entire network down.
19.6 Exercises
2. Perform the same operation as the one from the previous exercise,
but for the network at [Link]
19/2/[Link]. Can you tell which is the network with a power
law degree distribution and which is the Gn,p network?
Link prediction
20
For Simple Graphs
Link prediction is the branch of network analysis that deals with the
prediction of new links in a network. In link prediction, you see the
network as fundamentally dynamic, it can change its connections.
Suppose you’re at a party. You came there with your friends, and
you’re talking to each other, using the old connections. At some
point, you want to go and get a drink so you detach from the group.
On the way, you could meet a new person, and start talking to them.
This creates a new link in the social network. Link prediction wants
to find a theory to predict these events – in this case, that alcohol is
the main cause of new friendships at parties, or so I’m told –, as I
show the vignette in Figure 20.1.
other. Our best guess is that they will connect soon. In practice, the
probability of connecting two nodes is directly proportional to their
current degree: score(u, v) = k u k v , where k u and k v are u’s and v’s
degrees, respectively.
20.3 Adamic-Adar
1
∑ . The only difference with Adamic-Adar is that the scaling
z∈ Nu ∩ Nv k z
is assumed to be linear rather than logarithmic. Thus, Resource
Allocation punishes the high-degree common neighbors more heavily
than Adamic-Adar. You can see that the difference between k z and
log k z is practically nil for low values of k z , but balloons when k z is
high.
One could make a more complex version of the Resource Alloca-
tion index by assuming that the bandwidth of each node and of each
link is not fixed. Thus the amount of resources u sends can change,
and the amount of resources that can pass through the (u, v) link can
also be different from the one passing through other edges.
STEM
Figure 20.5: An example of a
Physics NASA Hierarchical Random Graph
link prediction. The hierarchy
Radioactivity Human
Computers fits the observed connections,
showing that researchers in the
same field are more likely to
connect. Then HRG looks at
pairs of nodes in the same part
of the hierarchy that are not yet
Likely Unlikely connected, and gives them a
higher score.
In HRG we’re basically saying that communities matter: it is more
likely for nodes in the same community to connect. Thus we fit the
hierarchy and then we say that the likelihood of nodes to connect is
proportional to the edge density of the group in which they both are.
If the nodes are very related, the group containing both nodes might
for simple graphs 283
3
284 the atlas for the aspiring network scientist
With the power of these two rules we know we can close the two
open triads we have and add a new neighbor to each node in the
original data. The end result is on the right. We now have more open
triads and could apply the rules again, which would in turn create
more square patterns and so on and so forth. In fact, one could use
GERM not only as a link predictor but also as a graph generator (and
put it in Chapter 14).
Note that I made two simplifications to GERM that the original
papers don’t make. First, in Figure 20.7, I assumed that each rule
applies with the same priority. I ignored its frequency and its confi-
dence. Of course, that would be sub-optimal, so the papers describe a
way to rank each candidate new edge according to the frequency and
confidence of each rule that would predict it.
Second, in all my examples I always assumed that the rules add a
new edge in the next time step and that all edges in the antecedent
are present at the same time. In reality, GERM allows to have more
for simple graphs 285
complex rules spanning multiple time steps. You could have a rule
saying something like: you have a single edge at time t = 0, you add
a second edge at time t = 1 creating an open triad, and then you close
the triangle at time t = 2. This triad closure rule spans three time
steps, rather than only two.
GERM has a final ace up its sleeve. We can classify new links
coming into a network into three groups: old-old, old-new, and new-
new. We base these groups according to the type of node they attach
to. I show an example in Figure 20.8. An “old-old” link appearing
at time t + 1 connected two nodes that were already present in the
network at time t. These are two “old” nodes. You can expect what
an “old-new” link is: a link connecting an old node with a node
that was not present at time t – a “new” node. New nodes can also
connect to each other in a “new-new” link. If the network represents
paper co-authorships, this would be a new paper published by two or
more individuals who have never published before.
45%
encode the fact that the attribute values 1 and 2 should be considered
“more similar” to each other than 1 and 1, 000. Extensions taking care
of these limitations might be possible, but I’m not aware of one.
This is not a function unique to GERM, though. Many link pre-
diction methods can be extended to take into consideration node
attributes as well. In fact, this is also a key ingredient in some net-
work generating processes. Node attributes are used, for instance,
when modeling exponential random graphs, as we saw in Section
16.2. In this case, differently than GERM, quantitative attributes
represent no issue.
The assumption is that, the more short paths are between u and v,
the more the two nodes are related. This is regulated by the 0 < α <
1 parameter: a lower α penalizes long paths more – because they have
a high l. Here, α plays the exact same role it did in Katz centrality. If
288 the atlas for the aspiring network scientist
4 3
(a) (b)
Another problem of using Hu,v is that the hitting time might
increase even if u and v are nearby in the graph, simply because the
graph has many vertices and edges that can lead the random walkers
astray. To counteract this problem, a solution could be allowing the
random walker to restart from u. This is practically equivalent to
calculate the PageRank, with the difference that you fix the origin of
the random walker. For this reason, since you random walker has a
root (u), it is usually called “rooted PageRank”.
21
Glen Jeh and Jennifer Widom. Scaling
personalized web search. In Proceedings
SimRank21 . As the name suggests, SimRank is based on an idea of
of the 12th international conference on
node similarity. The more similar two nodes are, the more likely they World Wide Web, pages 271–279. Acm,
are to connect. Similarity here is defined recursively: two nodes are 2003
∑ ∑ score( a, b)
a∈ Nu b∈ Nv
score(u, v) = γ .
ku kv
γ is a parameter you can tune. This is surprisingly similar to the
hitting time approach. The expected value of a SimRank score is γl ,
where l is the length of an average random walk from u to v.
22
Elizabeth A Leicht, Petter Holme, and
Vertex similarity22 . The name of this approach should tip you off Mark EJ Newman. Vertex similarity in
regarding its relationship with SimRank. However, it’s actually much networks. Physical Review E, 73(2):026120,
2006
closer to the Jaccard variant of common neighbor. In fact, the only
difference with Jaccard is the denominator. While Jaccard normalizes
the number of common neighbors by the total possible number of
common neighbors – which is the union of the two neighbor sets –
this approach builds an expectation using a random configuration
graph as a null model. This is a definition in line with the philosophy
of the structural equivalence we saw in Section 12.2.
In practice, score(u, v) = | Nu ∩ Nv |/(k u k v ). This is because two
nodes u and v with k u and k v neighbors are expected to have k u k v
common neighbors (multiplied by a constant derived from the aver-
age degree which would not make any difference as it is the same for
all node pairs in the network).
The same authors in the same paper also make a global variant of
this measure. Their inspiration is the Katz link prediction, where they
again provide a correction for a random expectation in a random
graph with the same degree distribution as G. I won’t provide the
full derivation, which you can find in the paper, but their score is:
−1
φA
score(u, v) = 2| E|λ1 D −1 I− D −1 .
λ1
The elements in this formula are the usual suspects: | E| is the
number of edges, λ1 is the leading eigenvalue of the adjacency matrix
A (not the stochastic, as that would be equal to one), D is the degree
matrix, and I is the identity matrix. The only odd thing is 0 < φ <
1, which is a parameter you can set at will. This is similar to the
parameter of Katz: smaller φ give more weight to shorter paths.
23
Weiping Liu and Linyuan Lü. Link
Local and superposed random walks23 . These two methods are a close prediction based on local random walk.
EPL (Europhysics Letters), 89(5):58007,
sibling to the hitting time approach. To determine the similarity
2010
between u and v, we place a random walker on u and we calculate
the probability it will hit node v. Note that, if we were to do infinite
length random walks, this would be the stationary distribution π.
This would be bad, as you know that this only depends on the degree
290 the atlas for the aspiring network scientist
of v, not on your starting point u. For this reason, the authors limit
the length of the random walk, and also add a vector q determining
different starting configurations – namely, giving different sources
different weights.
To sum up, the local random walk method determines score(u, v) =
qu πu,v + qv πv,u . The superposed variant works in the same way, with
the difference that the random walker is constantly brought back to
its starting point u. This tends to give higher scores to nodes closer in
the network.
24
Roger Guimerà and Marta Sales-
Stochastic block models24 . We saw the stochastic block models (SBM) Pardo. Missing and spurious interac-
as a way to generate graphs with community partitions (Section 15.2) tions and the reconstruction of complex
– and we will see them again as a method to detect communities networks. Proceedings of the National
Academy of Sciences, 106(52):22073–22078,
(Section 31.1). In fact, any link prediction approach, in a sense, is a 2009
graph generating model. Given the close relationship of SBMs with
community discovery, this class of solutions is particularly related to
the Hierarchical Random Graph approach.
CAR Index30 . In this index you favor pairs of nodes that are part of
30
Carlo Vittorio Cannistraci, Gregorio
Alanis-Lobato, and Timothy Ravasi.
a local community, i.e. they are embedded in many mutual connec- From link-prediction in brain connec-
tions. This is a variant of the idea of common neighbor: each shared tomes and protein interactomes to the
local-community-paradigm in complex
connection counts not equally, but proportionally more if it also networks. Scientific reports, 3:1613, 2013
shares neighbors with u and v. This basic idea can be implemented
in multiple ways, depending on which of the traditional link predic-
tion methods we want to extend. For instance, if we extend vanilla
common neighbors, you’d say that:
| Nu ∩ Nv ∩ Nz |
score(u, v) = ∑ 1+
2
.
z∈ Nu ∩ Nv
a
Figure 20.13: Comparing the
u v CAR index contribution to
score(u, v) for nodes a and b.
Node size is proportional to the
b contribution.
292 the atlas for the aspiring network scientist
| Nu ∩ Nv ∩ Nz |
score(u, v) = ∑ | Nz |
.
z∈ Nu ∩ Nv
5. Nothing stops you from using all the link prediction methods at
once and then aggregate their results. Really, it’s a free country.
20.9 Exercises
1. What are the ten most likely edges to appear in the network at
[Link] accord-
ing to the preferential attachment index?
2. Compare the top ten edges predicted for the previous question
with the ones predicted by the jaccard, Adamic-Adar, and resource
allocation indexes.
5 3 6
v v v v
Figure 21.3: The 16 templates
of directed signed triangles
u z u z u z u z in social status networks. The
(a) (b) (c) (d) color of the edge determines
v v v v its status: green = positive, red
= negative, blue = the edge we
are trying to predict – can be
u z u z u z u z either positive or negative.
(e) (f) (g) (h)
v v v v
u z u z u z u z
(i) (j) (k) (l)
v v v v
u z u z u z u z
(m) (n) (o) (p)
So far we have only considered the case of two possible edge types.
Moreover, these two types have a clear semantic: one type is positive,
300 the atlas for the aspiring network scientist
Layer Independence
As you might expect, there are tons of ways to face this problem.
The most trivial way to go about it is to apply any of the single layer
link prediction methods from Chapter 20 to each layer separately.
9
Manisha Pujari and Rushed Kanawati.
Then, you can create a single ranking table by merging all these Link prediction in multiplex networks.
predictions9 . NHM, 10(1):17–35, 2015
Figure 21.5 depicts an example for this process. Note that here
I use a rather trivial approach to aggregate, by comparing directly
the various scores. One could also apply to this problem the rank
aggregation measures presented in the previous chapter. In this way,
you could also aggregate different scores using different criteria:
common neighbors, preferential attachment, and so on.
This is practically a baseline: it will work as long as we have an
assumption of independence between the layers. As soon as having a
link in a layer changes the likelihood of connecting into another layer,
we expect to grossly underperform.
for multilayer graphs 301
Blending Layers
A slightly more sophisticated alternative is to consider the multilayer
network as a single structure and perform the estimations on it. For
instance, consider the hitting time method. This is based on the
estimation of the number of steps required for a random walker
starting on u to visit v. We can allow the random walker to, at any
time, use the inter layer coupling links exactly as if they were normal
edges in the network. At that point, a random walker starting from
u in layer l1 can and will visit node v in layer l2 . The creation of
our connection likelihood score is thus well defined for multilayer
networks. Figure 21.6 depicts an example for this process.
b
e
Figure 21.6: A slightly more so-
phisticated way to perform mul-
c
f
a d tilayer link prediction. Given
l1
g the input network, perform the
link prediction procedure on
l2
the full structure. In this case,
the gray arrow simulates a ran-
l3
dom walker going from node g
in layer l3 to node a in the same
layer, passing through node
These paths crossing layers are often called meta-paths. The in- c in layer l2 . The mutlilayer
formation from these meta-paths can be used directly as we just random walker contributes to
saw, informing a multilayer hitting time. Or we can feed them to a the score( g, a, l3 ).
classifier, which is trying to put potential edges in one of two cate-
10
Mahdi Jalili, Yasin Orouskhani,
gories: future existing and future non-existing links. Any classifier
Milad Asgari, Nazanin Alipourfard,
can perform this job once you collect the multilayer information from and Matjaž Perc. Link prediction in
the meta-path: naive Bayes, support vector machines (SVM), and multiplex online social networks. Royal
Society open science, 4(2):160863, 2017
others10 . 11
Darcy Davis, Ryan Lichtenwalter,
Other extensions to handle multilayer networks have been pro- and Nitesh V Chawla. Multi-relational
posed11 . These studies show that multilayer link prediction is indeed link prediction in heterogeneous
information networks. In ASONAM,
an interesting task, as there is a correlation between the neighbor- pages 281–288. IEEE, 2011
hood of the same nodes in different layer. The classical case involves
the prediction of links in a social media platform using information 12
Desislava Hristova, Anastasios
about the two users coming from a different platforms12 . Such layer- Noulas, Chloë Brown, Mirco Musolesi,
layer correlations are not limited to social media, but can also be and Cecilia Mascolo. A multilayer
approach to multiplexity and link
found in infrastructure networks13 . prediction in online geo-social networks.
EPJ Data Science, 5(1):24, 2016
13
Kaj-Kolja Kleineberg, Marián Boguná,
Multilayer Scores M Ángeles Serrano, and Fragkiskos
Papadopoulos. Hidden geometric
The last mentioned strategy is better, but it still doesn’t consider correlations in real multiplex networks.
all the wealth of information a multilayer network can give you. To Nature Physics, 12(11):1076, 2016
302 the atlas for the aspiring network scientist
see why, let’s dust off the concept of layer relevance we introduced
in Section 6.1. That is a way to tell you that a node u has a strong
tendency of connecting through a specific layer. If a layer exclusively
hosts many neighbors of u, that might mean that it is its preferred
channel of connection.
Rule Confidence
neous link prediction is one of the main tasks tackled in this subfield. 21
Yizhou Sun, Rick Barber, Manish
This is actually where metapaths were firstly developed21 . Figure Gupta, Charu C Aggarwal, and Jiawei
21.9 shows examples of possible metapaths in a co-authorship net- Han. Co-author relationship prediction
in heterogeneous bibliographic net-
work. These metapaths form the input of a classifier, which will then works. In 2011 International Conference on
spit out the most likely new metapaths involving specific nodes. Advances in Social Networks Analysis and
Mining, pages 121–128. IEEE, 2011
Other common approaches use a ranking factor graph model22 , 22
Yuxiao Dong, Jie Tang, Sen Wu, Jilei
which searches for common general patterns shared by the various Tian, Nitesh V Chawla, Jinghai Rao, and
layers of the network; or consider link prediction as a matching Huanhuan Cao. Link prediction and
recommendation across heterogeneous
problem23 .
social networks. In 2012 IEEE 12th
By the way, the converse of what I said about GERM and tensor International conference on data mining,
factorization applies also to heterogeneous link predictions. There is pages 181–190. IEEE, 2012
23
Xiangnan Kong, Jiawei Zhang, and
research showing how you can use this class of approaches to predict
Philip S Yu. Inferring anchor links
when an edge will appear, rather than its type24 . across multiple heterogeneous social
I should also mention that link prediction, community discovery, networks. In Proceedings of the 22nd ACM
international conference on Information &
and generating synthetic networks are sides of the same weirdly Knowledge Management, pages 179–188.
triangular coin. This holds also for multilayer networks. There are ACM, 2013
efforts to create models generating multilayer networks than can
24
Yizhou Sun, Jiawei Han, Charu C
Aggarwal, and Nitesh V Chawla. When
then be applied to predict new links on already existing real-world will it happen?: relationship prediction
multilayer networks25 , 26 . in heterogeneous information networks.
In Proceedings of the fifth ACM interna-
In this chapter, as in all chapters of this book, I presented only
tional conference on Web search and data
the most prominent methods to tackle the issue at hand, and the mining, pages 663–672. ACM, 2012
ones I’m most familiar with. The study of a deeper review work27 25
Caterina De Bacco, Eleanor A Power,
Daniel B Larremore, and Cristopher
is necessary if you want to make a living off solving multilayer link
Moore. Community detection, link
prediction. prediction, and layer interdependence
in multilayer networks. Physical Review
E, 95(4):042317, 2017
26
A. Roxana Pamfil, Sam D. Howison,
and Mason A. Porter. Edge correlations
in multilayer networks. arXiv preprint
arXiv:1908.03875, 2019
21.3 Summary 27
Yizhou Sun and Jiawei Han. Mining
heterogeneous information networks: a
structural analysis approach. Acm Sigkdd
1. In multilayer link prediction, besides predicting the appearance of Explorations Newsletter, 14(2):20–28, 2013
a new edge, you also need to guess in which layer the new edge
will appear. It’s not only about whether two nodes will connect, it’s
also about how they will connect.
21.4 Exercises
The Basics
When it comes to evaluate your prediction algorithm, you have to
distinguish between the training and the test datasets. The training
dataset is what your model uses to learn the patterns it is supposed
to predict. For instance, if you’re doing a common neighbor link
predictor, the training dataset is what you use to count the number of
shared connections between two nodes. Once you’re done examining
the input data, you have generated the results of the score(u, v)
function for all possible pairs of u, v inputs.
The test dataset is a set of examples used to assess performance.
When your model is done learning on the training dataset, it is
designing an experiment 307
Edge Score
5 5 Figure 22.1: An example of
1, 2 1
4 6 1, 3 2 4 6 train and test sets for a network.
The information (a) we use to
1, 4 1
8 1, 5 1 8 build the score table (b), using
3 7 3 7 the common neighbor approach.
1, 6 2
I highlight the test edges (c) in
2, 4 2
2 1 2 1 blue.
... ...
(a) Train (c) Test
(b) Scores
In machine learning there is also what you’d call a “validation”
dataset, for the tuning of the parameters, but that usually doesn’t
apply to link prediction. Link predictors usually have no or a trivial
number of parameters, therefore you can safely conflate training and
validation in the same set.
There is a fundamental tenet for making data-driven predictions
that still holds. You can never ever ever use the data that trained
your model to test it. In other words, training and test sets have to
be disjoint. If you test your method on the same data that trained
it, you’re going to overfit: your method is going to learn only the
specifics of the training set and nothing about the general forces that
shaped it the way it is.
What this means in link prediction is that you cannot claim to have
predicted a link that was already in your data. You have to focus
only on those pairs of nodes that were not connected in the training
set. That is why in Figure 22.1(c) the edges that were already in the
training set are gray rather than blue: we won’t make predictions on
those, because we already know they exist.
So now the problem is: how do you do that? If you have a net-
308 the atlas for the aspiring network scientist
609, 066 × .05 ∼ 30, 453 new edges. As we just saw, the number of
potential edges is just above 18B. Putting these two facts together
lets us reach an absurd conclusion: we can build a link prediction
method that will tell us that no new link will ever appear. If we do
so, we would be right 99.999% of the times. We would make 18B
correct predictions – no edge – and we would get it wrong only 30k
times. The accuracy of the “always negative” predictor in Figure 22.3
is ∼ 85%: not bad!
However that’s... kind of not the point? We’re in this business be-
cause we want to predict new links. Returning a negative prediction
for all possible cases is not helpful. The usual fix for this problem 3
Ryan N Lichtenwalter, Jake T Lussier,
is building your test set in a balanced way3 , 4 . Rather than asking and Nitesh V Chawla. New perspectives
about all possible new edges, you create a smaller test set. Half of the and methods in link prediction. In
Proceedings of the 16th ACM SIGKDD
edges in the test set is an actual new edge, and then you sample an international conference on Knowledge
equal number of non-edges. This would make our Internet test set discovery and data mining, pages 243–252,
2010
containing 60k edges, not 18B. 4
Ryan Lichtnwalter and Nitesh V
Chawla. Link prediction: fair and
effective evaluation. In 2012 IEEE/ACM
22.2 Evaluating International Conference on Advances in
Social Networks Analysis and Mining,
pages 376–383. IEEE, 2012
Let’s assume that we have competently built our training and test set.
We made our model learn on the former. We now have two things:
prediction – the result of the model – and reality – the test set. We
want to know how much these two sets overlap.
There are four possible cases:
Confusion Matrix
Humans like single numbers, because seeing a number going up
tingles our pleasure centers (wait, what? You don’t feel inexplicable
arousal while maximizing scores? I question whether you’re in the
right line of work...). However, we should beware of what we call
“fixed threshold metrics”, i.e. everything that boils down a complex
phenomenon to a single number. Usually, to reduce everything to
a single measure you have to make a number of assumptions and
simplifications that may warp your perception of performance.
Actual
Yes No Figure 22.4: The schema of
a confusion matrix for link
prediction. From the top-left
Yes corner, clockwise: true positives,
false positives, true negatives,
false negatives.
Predicted
No
That is why one of the first thing you should look at is a confusion
matrix. A confusion matrix is simply a grid of four cells, putting 5
Stephen V Stehman. Selecting and
interpreting measures of thematic
the four counts I just introduced in a nice pattern5 . You can see an
classification accuracy. Remote sensing of
example in Figure 22.4. Confusion matrices are nice because they Environment, 62(1):77–89, 1997
don’t attempt to reduce complexity, but at the same time you see
information in an easy-to-parse pattern.
By looking at two confusion matrices you can say surprisingly
sophisticated things about two different methods. The one in Figure
22.5(a) does a better job in making sure a positive prediction really
designing an experiment 311
(a) (b)
corresponds to a new link: there are very few false positives (one)
compared to the true positives (15). The one in Figure 22.5(b) mini-
mizes the number of false negatives, with the downside of having a
lot of false positives.
By combining the cells of a confusion matrix, you can easily de-
rive measures like TPR or FPR, or many others. They are simple
operations on the rows and columns.
If you didn’t balance your test set, the confusion matrix can end
up being irrelevant, as the vast majority of your observations will end
up in the true negative cell, obliterating all the rest.
Another disadvantage of the confusion matrix is that you have
to pick a threshold in your score. In other words, you predict the
appearance of a link if it obtains a score higher than the specific
threshold, otherwise you don’t. This is in itself a problematic choice,
thus it is common to show the evolution of your accuracy as you
change that threshold. For high values of the threshold you only
report high confidence predictions, which become less and less
confident as you decrease the threshold. This is the topic explored in
the rest of the chapter.
1
True Positive Rate
( + )
TPR
Highest score 2nd highest Random guess
/( + ) FPR 1
(a) (b)
Figure 22.6: Schema of ROC
true positives, and thus contribute to the y-axis more than they do to curves.
the x-axis. Just like in the confusion matrix, there are multiple ways
for this to happen. We can be very precise at high scores, or at all
scores on average. The two classifiers in Figure 22.7 will be used in
different scenarios with different requirements.
FPR
ROC curves are great – you might even say that they ROC – but,
at the end of the day, you might want to know which of the two
classifiers is better on average. ROC curves can be reduced to a single
number, a fixed threshold metric. Since we just said that the higher
the line on the ROC plot the better, one could calculate the Area
Under the Curve (AUC). The more area under that curve, the better
your classifier is, because for each corresponding FPR value, your
TPR is higher – thus encompassing more area.
You don’t need to know calculus to estimate the area under the
curve, because it’s such a standard metric that any machine learning
package will output it for you. The AUC is 0.5 for the random guess:
that’s the area under the 45 degree line. An AUC of 1 – which you’ll
designing an experiment 313
never see and, if you do, it means you did something wrong – means
a perfect classifier.
Note that ROC curves and AUCs are unaffected if you sample
your test set randomly, namely if you only test potential edges at
random from the set of all potential edges – I discussed before how
this is a common thing to do because of the unmanageable size
of the real test set. However, that is not true if you perform a non-
random sampling. This means choosing potential edges according
to a specific criterion. If your criterion is “good”, meaning that your
sampling method is correlated with the actual edge appearance
likelihood, you’re going to see a different – lower – AUC value. That
is because, if you don’t sample, the vast number of easy-to-predict
false negatives increases your classifier’s accuracy.
( + ) ( + )
You can do a few things with precision and recall. First, you
can transform them into fixed threshold metrics. This is done by
calculating what we call “Precision@n”, defining n as the number of
predictions we want to make. For instance, in Precision@100 we only
consider as an actual prediction the 100 pairs of nodes that have the
314 the atlas for the aspiring network scientist
positives,
at the price
of having
many false
ones.
Recall
P
PP = 10 log10 .
Pr
This is a decibel-like logscale: a PP = 1 implies your predictor
is ten times better than random, while PP = 2 means you are one
hundred times better than random. You can also create PP-curves by
having on the x-axis the share of links you remove from your training
designing an experiment 315
22.3 Summary
3. Since real networks are sparse, there are more non-edges than
edges. Thus a link prediction always predicting non-edge would
have high performance. That is why you should balance your test
sets, having an equal number of edges and non-edges.
22.4 Exercises
2. Draw the ROC curves on the cross validation of the network used
at the previous question, comparing the following link predictors:
preferential attachment, jaccard, Adamic-Adar, and resource alloca-
tion. Which of those has the highest AUC? (Again, scikit-learn
has helper functions for you)
3. Calculate precision, recall, and F1-score for the four link predictors
as used in the previous question. Set up as cutoff point the nineti-
eth percentile, meaning that you predict a link only for the highest
ten percent of the scores in each classifier. Which method performs
best according to these measures? (Note: when scoring with the
scikit-learn function, remember that this is a binary prediction
task)
The Hairball
23
Bipartite Projections
1. Degree distributions;
2. Epidemics spread;
3. Communities.
Many papers have been written on how power law degree distribu-
tions are ubiquitous1 , 2 , 3 , 4 . Chances are that any and all the networks 1
Albert-László Barabási and Réka
you’ll find on your way as a network analyst do not have even a hint Albert. Emergence of scaling in random
networks. science, 286(5439):509–512,
of a power law degree distribution. In the best case scenario you are 1999
going to have shifted power laws, or exponential cutoffs – if you’re 2
Albert-László Barabási and Eric
lucky – (for a refresher on these terms, see Section 6.4). Bonabeau. Scale-free networks. Scientific
american, 288(5):60–69, 2003
My second example is epidemics spread – Figure 23.1. As we 3
Reka Albert. Scale-free networks in cell
saw in Part V, SIS/SIR models tell us exactly when the next node is biology. Journal of cell science, 118(21):
going to be activated. In practice, data about real activation times 4947–4957, 2005
4
Albert-László Barabási. Scale-free
has (a) high levels of noise, (b) many exogenous factors that have as networks: a decade and beyond. science,
much power in influencing how the infection spreads as the network 325(5939):412–413, 2009
connections have.
Third, and more famously, communities. We are not going to dive
deeply into the topic only until Part IX. But, very superficially, when
(a) (b)
There are a few ways in which hairballs arise, which are the focus
of this book part. First, many networks are not observed directly:
they are inferred. If the edge inference process you’re applying does
not fit your data, it will generate edges it shouldn’t. Second, even
if you observe the network directly, your observation is subject to
noise, connections that do not reflect real interactions but appear due
to some random fluctuations. Finally, you might have the opposite
problem: you’re looking at an incomplete sample, and thus missing
crucial information.
Reality Inference You Reality Noise You Reality Data Collection You
2
1 6
3
5 7
4
5
3 8
Figure 23.4: An example of
6
naive bipartite projection,
7 4 2 where we connect nodes of one
8 “Connecting movies type if they have a common
because the same
Users Movies users watch them” neighbor.
bipartite projections 321
are broad. This means that there are going to be some users in your
bipartite user-movie network with a very high degree. These are
power users, people who watched everything. They are a problem:
under the rule we just gave to project the bipartite networks, you’ll
end up with all movies connected to each other. A hairball. The key
lies in recognizing that not all edges have the same importance. Two
movies that are watched by three common users are more related to
each other than two movies that only have one common spectator.
1
Figure 23.5: An example of Sim-
2 1 6
ple Weight bipartite projection,
3
5 7 where we connect nodes of one
4 2 2 type with the number of their
5
2
common neighbors.
3 2 8
6
3
7 4 2
8
Wu,v = |Nu ∩ Nv|
1
Figure 23.7: An example of
2 1 6
Hyperbolic Weight bipartite
3
5 .46 7 projection, where each common
4 .46 neighbor z contributes k− 1
z to
5
.46
the sum of the edge weight.
3 .46 8
6
.79
7 4 2
8 1
Wu,v = ΣzϵN ∩N
u v kz
1
wu,v = ∑ ku kz
.
z∈ Nu ∩ Nv
1
1/2 Figure 23.8: An example of
1/3 2
Resource Allocation bipartite
1/8
1/2 3
1 0.229 projection, where each com-
4 mon neighbor z contributes
5 (k u k z )−1 to the sum of the
6 0.153 2 edge weight. When connecting
node 1 to node 2, from node
7
1’s perspective the edge weight
8 1
Wu,v = ΣzϵN ∩N u v kukz is (1/2 ∗ 1/3) + (1/2 ∗ 1/8),
because the two common
This strategy also works for weighted bipartite networks. If B is neighbors have degree of
your weighted bipartite adjacency matrix, the entries of W are: 3 and 8, respectively, and
node 1 has degree of two.
Buv
wu,v = ∑ ku kz
. However, from node 2’s per-
z∈ Nu ∩ Nv spective, the edge weight is
In practice, you replace the 1 in the numerator with the edge (1/3 ∗ 1/3) + (1/3 ∗ 1/8),
weights connecting z to v and u. Moreover, we can also have node because node 2 has three neigh-
weights, noticing that some nodes might have more resources than bors.
others. Suppose that you have a function f giving each node in the
network a resource weight. After you perform the resource allocation
projection, each node will have a new amount of resources f 0 = W f .
Note that, in this case, W is not symmetric: in the scenario with a
single common neighbor z, u’s score for v would be (k u k z )−1 , while
v’s score would be (k v k z )−1 . If k u 6= k v , then the scores are different.
In many cases, this provides a better representation of the network
than one ignoring asymmetries. You might be the most similar
author to me because I always collaborated with you, but if you also
contributed to many other papers with other people, then I might not
be the author most similar to you.
1 1 1
W has a well-defined diagonal: wu,u = ∑ = ∑ .
z∈ Nu k u k z | Nu | z∈ Nu k z
In fact, this diagonal is the maximum possible similarity value of the
row: only a node v with the very same neighbors and nothing else
can have a weight wu,v = wu,u .
326 the atlas for the aspiring network scientist
1
Figure 23.9: An example of Ran-
2
dom Walks bipartite projection,
3
1 0.049 where the connection strength
4 between u and v is dependent
5 on the stationary distribution
6 0.022 2 π, telling us the probability
of ending in v after a random
7
walk.
8 Wu,v = πAu,v
# Edges
# Edges
# Edges
104
103 103 103
103
102 102 102 102
101 101 101 101
100 100 100 100
100 101 102 103 10-3 10-2 10-1 100 10-2 10-1 100 100
Edge Weight Edge Weight Edge Weight Edge Weight
# Edges
# Edges
# Edges
103 103 103 103
102 102 102 102
101 101 101 101
100 100 100 100
10-6 10-5 10-4 10-3 10-2 10-1 100 10-6 10-5 10-4 10-3 10-2 10-1 10-8 10-6 10-4 10-2 100 10-11 10-9 10-7 10-5 10-3
Edge Weight Edge Weight Edge Weight Edge Weight
(e) Hyperbolic (f) ProbS (g) Hybrid (λ = 0.5) (h) Random Walks
Figure 23.10: The distributions
weight is 2.14. This is very much not the case for other projection of edges weights in the pro-
strategies such as Jaccard (Figure 23.10(b)), where there is no trace jected Twitter network for eight
of a power law. And, in many cases such as cosine and Pearson different projection methods.
(Figures 23.10(c-d)), the highest edge weight is actually the most The plot report the number
common value, rather than being an outlier such as in the hyperbolic of edges (y axis) with a given
projection (Figure 23.10(e)). weight (x axis).
Is the difference exclusively in the shape of the distribution, or
do these approaches disagree on the weights of specific edges? To
answer this question we have to look at a scattergram comparing the
edge weights for two different projection strategies. This is what I do
in Figure 23.11.
I picked three cases to show the full width of possibilities. In
Figure 23.11(a), I compare the cosine projection against the Jaccard
one. This is the pair of projections that, in this dataset, agree the most.
Their correlation is > 0.94. Looking at the figure, it is easy to see that
there isn’t much difference. You can pick either method and you’re
going to have comparable weights. The opposite case compares two
method that are anti-correlated the most. This would be HeatS and
the random walks approach, in Figure 23.11(b). They correlation in a
log-log space is a staggering −0.7. From the figure you can probably
spot a few patterns, but the lesson learned is that the two methods
build fundamentally different projections.
Ok, but these are extreme cases. How does the average case looks
like? To get an idea, I chose a particular pair of measures: HeatS and
ProbS (Figure 23.11(c)). You might expect the two to be more similar
than the average method: after all, one is the transpose of the other.
You’d be very wrong. In this dataset, HeatS and ProbS are actually
anti correlated, at −0.34 in the log-log space. HeatS and ProbS would
be positively correlated if the nodes of type V1 with similar degrees
bipartite projections 329
23.7 Summary
2. In network projection you pick one of the two node types and you
connect the nodes of that type if they have common neighbors
of the other type. Normally you’d count the number of common
neighbors they have (simple weighted) and then evaluate their
330 the atlas for the aspiring network scientist
statistical significance.
23.8 Exercises
24.1 Naive
Discarded
Figure 24.1: A vignette of the
naive thresholding procedure.
Each red bar is an edge in
the network. The bar’s width
is proportional to the edge’s
weight. Here, I sort all edges
Threshold in decreasing weight order. I
then establish a threshold and
Broad Weight Distributions discard everything to its right.
The first problem is that, in real world networks, edge weights dis-
tribute broadly in a fat-tail highly skewed fashion, much like the
degree (Section 6.3). Let’s take a quick look again at the edge weight
distribution we got using the simple projection in the previous chap-
ter for our Twitter network. I show the distribution again in Figure
24.2.
7
10
Figure 24.2: The distributions of
106
edges weights in the projected
105 Twitter network using the sim-
# Edges
In this network, 82% of the edges have weight equal to one. The
smallest possible hard threshold would remove 82% of the network,
without allowing for any nuance. Moreover, since we have a fat
tailed edge weight distribution, it is hard to motivate the choice of a
threshold. Such a highly skewed distribution lacks of a well-defined
average value and has undefined variance. You cannot motivate your
threshold choice by saying that it is “x standard deviations from the
average” or anything resembling this formulation.
around u will have weight equal to one. On the other hand, if u had
shared thousands of URLs, it will likely connect to another user with
similar sharing patterns, because statistically speaking they have high
odds of sharing at least few of the same URLs, even if it happens by
chance. Thus many edges around u will have high weights.
9
Figure 24.3: The average weight
Avg Neighbor Weight
8
7 of edges sharing a node with
6 a focus edge (y axis) against
5 the weight of the focus edge (x
4
axis). Thin lines show the stan-
3
2 dard deviation. One percent
1 sample of the Twitter network.
0
0 1 2
10 10 10
Edge Weight
of columns are the same number. This cheeky proof means that you
cannot apply the doubly stochastic backboning to bipartite networks,
unless |V1 | = |V2 |.
Doubly stochastic matrices have other fun properties. If you
remember Section 8.1, the leading left eigenvector of a stochastic
adjacency matrix is the stationary distribution, while the leading
right eigenvector is a constant – assuming the graph is connected. In
a doubly stochastic matrix, both the left and the right eigenvectors
are equal to a constant or, in other words, the stationary distribution
of a doubly stochastic matrix is constant. This isn’t really a necessary
thing to know while doing network backboning, but I though it was
cool, so do with this information what you will.
The intuition behind the high salience skeleton (HSS) is that a net-
work is a structure facilitating the exchange of information or goods.
Thus, some connections are more important than others because they
keep the network together in a single component. The main impera-
tive is to allow all nodes to reach all other nodes in the most efficient
and high-throughput way possible. Thus you need to interrogate
each node and ask them what are the most efficient paths from their
perspective. This cannot be done repurposing measures such as edge
betweenness – whose objective is also telling us how structurally
important an edge is (Section 11.2) – because these measures adopt
a “global” point of view: they are the salient connections for the
network as a whole, but they might leave some nodes poorly served.
To build an HSS we loop over the nodes and we build their short-
est path tree: a tree originating from a node, touching all other nodes
in the minimum number of hops possible and maximum amount of
edge weight possible. In practice we start exploring the graph with
a BFS and note down the total edge weight of each path. When we
reach a node that we already visited we consider the edge weights of
the two paths and the one with the highest one wins.
Note that we have the constraint of the structure originating from
a node to be a tree. Thus it cannot contain a triangle. Consider Figure
24.6 as an example. In the bottom example, we might want to save
two edges at the same time. Our origin node, the one at the top of
the network connects strongly with one node which also connects
strongly to the node on the left. However, we cannot have both
edges in the shortest path tree, as that would create a cycle. The final
salience skeleton is allowed to have triangles and cycles, because it is
the sum of all the shortest path trees.
We perform this operation for all nodes in the network and we
network backboning 337
180
160 Figure 24.7: A typical “horn”
140 plot for the edge weight distri-
120 bution in HSS. The plot reports
# Edges
100
how many edges (y axis) are
80
part of a given share of shortest
60
40 path trees (x axis).
20
0
0 0.2 0.4 0.6 0.8 1
HSS Score
The HSS makes a lot of sense for networks in which paths are
meaningful, like infrastructure networks. However, it requires a
lot of shortest path calculations – which makes it computationally
expensive. Moreover, the edges are either part of (almost) all trees or
of (almost) none of them. Figure 24.7 shows an example of this edge
weight distribution, showcasing the typical “horn” shape of the HSS
score attached to the original edges. You can see clearly that there
are two peaks: one at zero – the edge is in no shortest path tree –; the
338 the atlas for the aspiring network scientist
2
1 Figure 24.8: An example of
2 induced graph. (a) The original
4 3
graph. I highlight in red the
5 nodes I pick for my induced
9
8 3 graph. (b) The induced graph
6
of (a), including only nodes in
11
7
10 9 red and all connections between
8 them.
12 6
(a) (b)
A network is convex if all its induced subgraph are convex. No
matter which set of nodes you pick: as long as they are part of a
single connected component, they are all going to be convex. This
might look like a weird and difficult to understand concept, but you
can grasp it with the help of elementary building blocks you already
saw in this book.
Figure 24.9 shows the two basic alternatives for a convex network.
In a tree – Figure 24.9(a) – any set of connected nodes is a convex
subgraph. There are no other edges in G you can use to make short-
cuts, because they’d create a cycle and trees cannot contain a cycle.
A clique – Figure 24.9(b) – is a convex network as well: all possible
connections are part of G, so picking any subset V 0 of nodes will
also result in a clique. Since all nodes are connected to each other
in a clique, you have all the shortest paths between them, making it
convex.
network backboning 339
(a) (b)
In this and the following sections, we’re slightly turning the perspec-
tive on network backboning. You could consider these as a different
subclass of the problem. They all apply a general template to solve
the problem of filtering out connections, which relate to the “noise
reduction” application scenario of network backboning. Up until
now, we adopted a purely structural approach which re-weights
nodes according to some topological properties of the graph. Here,
instead, given a weighted graph, we adopt a template composed by
three main steps: (1) define a null model based on node distribution
properties; (2) compute a p-value for every edge to determine the
statistical significance of properties assigned to edges from a given
distribution; (3) filter out all edges having p-value above a chosen
significance level, i.e. keep all edges that are least likely to have
occurred due to random chance.
The disparity filter (DF) is the first example in this class of solu-
tions. It takes a node-centric approach. Each node has a different
threshold to accept or reject its own edges. This is done by modeling
an expected typical “node strength”, for instance the average of its
edge weights. Then we keep only those edges which are higher than
340 the atlas for the aspiring network scientist
Since you need one success out of the two attempts to keep the
edge, you end up with strong hubs connected to the entire net-
work, and few peripheral connections (hub-spoke structure, or core-
periphery, with no communities). In other words, the disparity filter
tends to create networks with high centralization (Section 11.8),
broad degree distributions, and weak communities. In many cases,
that is fine. For some other scenarios, we might want to consider an
alternative.
In summary, this means that DF ignores the weights of the neigh-
bors of a node when deciding whether to keep an edge or not. There 15
Navid Dianati. Unwinding the hairball
is a collection of alternatives15 , 16 that take this additional piece of graph: pruning algorithms for weighted
information into account and are thus less biased. complex networks. Physical Review E, 93
(1):012304, 2016
In this section I explained the disparity filter only in the case of 16
Valerio Gemmetto, Alessio Cardillo,
undirected networks. You can apply the same technique also for di- and Diego Garlaschelli. Irreducible
rected networks. In this case, you need to make sure that you’re prop- network backbones: unbiased graph
filtering via maximum entropy. arXiv
erly accounting for direction in your p-value calculation: the edge preprint arXiv:1706.00230, 2017
must be significant either when compared to the out-connections of
the node sending the edge, or when compared to the in-connection
weights of the node receiving it.
24.6 Noise-Corrected
successes) is the weight of the edge wu,v , the number of trials is the
total sum of edge weights in the network ∑ wu,v , and the probability
u,v
of success is given by:
∑ wu,v0 × ∑ wu0 ,v
v0 ∈ Nu u0 ∈ Nv
pu,v = !2 .
∑ wu0 ,v0
u0 ,v0
6 5
To any person who has ever worked with real world data, it should
come as no surprise that datasets are often disappointing. They
contain glaring errors, incomprehensible omissions, and a number
of other issues that make them borderline useless if you don’t pour
hours of effort into fixing them. In fact, I’d say that 80% of data
science is just about cleaning data, and only 20% about shiny and fun
analysis techniques. This obviously applies to network data as well.
You’ll find edges in your networks that shouldn’t be there, and you’ll
have plenty of missing or unobserved connections.
Admittedly, techniques to clean network data would deserve their
own chapter but, frankly, I don’t know many of them. This section
is awkwardly placed here because intimately related to the initial
assumption of noise-corrected backboning – that connections are
noisy. But I could have placed it in the link prediction part as well,
since it’s not only about throwing away observed connections but
also inferring missing ones. In fact, a survey paper about measure- 18
Dan J Wang, Xiaolin Shi, Daniel A
ment error in network data18 points out that measurement error is McFarland, and Jure Leskovec. Mea-
routinely considered only a problem about missing data, rather than surement error in network data: A
re-classification. Social Networks, 34(4):
the more general framing as uncertainty. 396–409, 2012a
Network data cleaning is thus the lovechild of network backboning
and link prediction, but that’s a rather barren marriage – as far as
344 the atlas for the aspiring network scientist
19
Tiago P Peixoto. Reconstructing
I know. In fact, one of the few papers I know19 delivers the truth networks with unknown and heteroge-
in a brutal and deadpan way: “[in network analysis] the practice of neous errors. Physical Review X, 8(4):
041011, 2018
ignoring measurement error is still mainstream”. I hope this tiny
section will contribute to make things change.
To give you an idea of the significance of the measurement error
blind spot in network science, consider the Zachary Karate Club
network. As I’ll explain in details in Section 46.4, everyone in our
field is madly in love with this toy example. The paper presenting the 20
Wayne W Zachary. An information
network20 has been cited more than 4.5k times – and not everybody flow model for conflict and fission in
using this network cites it. The fun thing about this graph is that we small groups. Journal of anthropological
research, 33(4):452–473, 1977
actually don’t know whether it has 77 or 78 edges. The bewildering
thing about this graph is that almost no one even mentions this problem!
The basic way to go about estimating (and correcting) measure- 21
M Newman. Network recon-
ment error is by measuring network data multiple times21 . This is struction and error estimation with
a way to reconstruct a primary error estimate, i.e. to diagnose how noisy network data. arXiv preprint
arXiv:1803.02427, 2018a
good or bad our data collection is. The paper I cited in the previous
paragraph creates a clever Bayesian framework, which enables a
similar result, but does not require multiple measurements. There are
some works outside network science proper that also cite measure-
ment error as one of the many things you should think about when 22
Arun Advani and Bansi Malde.
working with networked data22 . Empirical methods for networks data:
Social effects, network formation and
measurement error. Technical report, IFS
24.8 Summary Working Papers, 2014
4. In high-salience skeleton, you calculate the short path tree for each
node and you re-weight the edges counting the number of trees
using them. Then you keep the most used edges. This is usually
network backboning 345
computationally expensive.
24.9 Exercises
4. How many edges would you keep if you were to return the dou-
bly stochastic backbone including all nodes in the network in a
single (weakly) connected component with the minimum number
of edges?
25
Network Sampling
25.1 Induced
Node Induced
If you focus on nodes, it means that you are specifying the IDs of a
set of nodes that must be in your sample. Then, usually, what you do
is collecting all their immediate neighbors. The issue here is clearly
deciding the best set of node IDs from which to start your sampling.
There are a few alternatives you could consider.
The first, obvious, one is to choose your node IDs completely
at random. Random sampling is a standard procedure in many
other scenarios, and has its advantages. If the properties you’re
interested in studying are normally distributed in your population,
a large enough random sample will be representative. However,
network sampling 349
(a) (b)
Edge Induced
Another way to generate induced samples is to focus on edges rather
than nodes. This means selecting edges in a network and then crawl
their immediate neighbors. There are a few techniques to do so. One
is the obvious extension of random node induced sampling: random
edge induced sampling. You select edges at random and you collect
350 the atlas for the aspiring network scientist
Snowball
Name k
of your
friends
Figure 25.4: Snowball sampling.
Name k Your sampler (blue) starts
of your from a seed (red) and asks for
friends
k = 3 connections. Red names
their green friends, but not
the gray ones. The interviewer
then recursively asks the same
question to each of the newly
sampled green individuals. If
no one ever mentions the gray
ones, those are not sampled and
won’t be part of the network.
Forest Fire
Name
all your
friends
Figure 25.5: Forest fire sam-
Name pling. Your sampler (blue)
all your starts from a seed (red) and
friends
asks for all the connections a
node. If the probability test
succeeds, the neighbor turns
green and is also explored. If it
fails, the neighbor remains gray
and is not explored further.
The random walk sampling family does exactly what you would
expect it to do given its name: it performs a random walk on the
graph, sampling the nodes it encounters. After all, if random walks
are so powerful and we can use them for ranking nodes (Section
11.4) or projecting bipartite networks (Section 23.5), why can’t we use
them for sampling too? I’ll start by explaining the simplest approach
and its problems, moving into sophisticated variants that address its
downsides.
Vanilla
In Random Walk (RW) sampling, we take an individual and we ask
them to name one of their friends at random. Then we do the same
with her and so on. Figure 25.6 shows the usual vignette applied to
this strategy.
Name 1
of your
friends
Figure 25.6: Random walk
Name 1 sampling. Your sampler (blue)
of your starts from a seed (red) and
friends
asks for all the connections a
node (green + gray). One of the
neighbors is picked at random
and becomes the new seed
(green) and, when asked, will
name another green node to
become the new seed.
Metropolis-Hastings
One way in which we could fix the issues of random walk sampling
is by perform a “random” walk. Meaning that we still pick a neigh-
bor at random to grow our sample, but we become picky about
whether we really want to sample this new node or not.
In the Metropolis-Hastings Random Walk (MHRW), when we
select a neighbor of the currently visited node, we do not accept
it with probability 1. Instead, we look at its degree. If its degree is
higher than the one of the node we are visiting, we have a chance of
rejecting this neighbor and trying a different one. This probability is
the old node’s degree over the new node’s degree. The exact formula
for this decision is p = k v /k u , assuming that we visited v and we’re 20
Daniel Stutzbach, Reza Rejaie, Nick
Duffield, Subhabrata Sen, and Walter
considering u as a potential next step20 , 21 . Willinger. On unbiased sampling for
Thus, if the current node v has degree of 3, and its u neighbor unstructured peer-to-peer networks.
has degree of 100, the probability of transitioning to u is only 3% – IEEE/ACM Transactions on Networking
(TON), 17(2):377–390, 2009
note that this is after we selected u as the next step of the random 21
Balachander Krishnamurthy, Phillipa
walk, thus the visit probability is actually lower than 3%: first you Gill, and Martin Arlitt. A few chirps
have a 1/k v probability of being selected and then a k v /k u probability about twitter. In Proceedings of the first
workshop on Online social networks, pages
of being accepted. If we were, instead, to transition from u to v, we 19–24. ACM, 2008
would always accept the move, because 100/3 > 1, thus the test
always succeeds. In practice, we might refuse to visit a neighbor if
its degree is higher than the currently visited node. The higher this
difference, the less likely we’re going to visit it. A random walk with
network sampling 355
Re-Weighted
In Re-Weighted Random Walk (RWRW) we take a different approach.
We don’t modify the way the random walk is performed. We extract
the sample using a vanilla random walk. What we modify is the way
we look at it. Once we’re done exploring the network, we correct the 22
Matthew J Salganik and Douglas D
result for the property of interest22 , 23 . Say we are interested in the Heckathorn. Sampling and estimation in
hidden populations using respondent-
degree. We want to know the probability of a node to have degree driven sampling. Sociological methodology,
equal to i. We correct the observation with the following formula: 34(1):193–240, 2004
23
Amir Hassan Rasti, Mojtaba Tork-
∑ i −1 jazi, Reza Rejaie, Nick Duffield, Wal-
v∈Vi ter Willinger, and Daniel Stutzbach.
pi = . Respondent-driven sampling for charac-
∑ xv−0 1 terizing unstructured overlays. In IEEE
v 0 ∈V INFOCOM 2009, pages 2701–2705. IEEE,
2009
The formula tells us the probability of a node to have degree equal
to i (pi ). This is the sum of i−1 – the inverse of the value – for all
nodes in the sample with degree i (Vi ), over 1/ degree (xv−0 1 ) of all
nodes in the sample (V). This is also known as Respondent-Driven 24
H Russell Bernard and Harvey Rus-
sell Bernard. Social research methods:
Survey24 , because it is used in sociology to correct for biases in the Qualitative and quantitative approaches.
sample when the properties of interest are rare and non-randomly Sage, 2013
distributed throughout the population. Figure 25.8 attempts to break
down all parts of the formula.
Let’s make an example. Suppose you want to estimate the proba-
bility of a node to have degree i = 2. First, you perform your vanilla
random walk sample. Say you extracted 100 nodes. Twenty of those
nodes have degree equal to two. So your numerator in the formula
will be the sum of i−1 = 1/2 for |Vi | = 20 times: 20 ∗ 1/2 = 10. If
we assume that there were 50 nodes of degree 1, 10 of degree 3, 8 of
356 the atlas for the aspiring network scientist
p of nodes
with value i Figure 25.8: The Re-Weighted
Random Walk formula, esti-
Set of mating the probability pi of
nodes observing the i value in a prop-
with erty of interest, using the set
value i of sampled nodes Vi with that
particular value in the total set
of v sampled nodes.
Value
Set of nodes for v’
in the sample
19
20 Figure 25.9: Neighbor Reser-
voir sampling. Nodes in the
12
11 explored set V 0 are in red.
8 21
Neighbors of V 0 – the reservoir
3
– are in green.
18 9 7 5
2
4 14
6
17 10
1
15 16
13
able. Other unlucky u-v draws are forbidden too. For instance, you
cannot perform the swap if you pick nodes 3 and 12.
Figure 25.10 shows an example of how some of these different
strategies would explore a simple tree. I don’t show RWRW, because
the samples it extracts are indistinguishable from the vanilla random
walk ones. I also don’t include NRS, because it’s too subtle to really
be appreciated in a figure like this one.
106
Figure 25.11: A power law de-
105
gree distribution, showing the
104 count of nodes (y-axis) with
Count
3
10 a given degree (x-axis). The
2
10 colors in the plot represent in
which cases the first API policy
101
described in the text is faster
100
100 101 102 103 than the second (purple) and
k when the second is faster than
the first (blue).
system. Which means that there are going to be trade-offs when
reconstructing the underlying network.
Pagination is often not the only thing you need to worry about.
Other challenges might be sampling a network in presence of hostile 32
Edward Bortnikov, Maxim Gure-
behavior32 . For instance, some hostile nodes will try to lie about their vich, Idit Keidar, Gabriel Kliot, and
connections and it’s your duty to reconstruct the true underlying Alexander Shraer. Brahms: Byzantine
resilient random membership sampling.
structure. Or not: there are reasonable and legit reasons to lie about
Computer Networks, 53(13):2340–2359,
one’s connection, for instance to protect one own privacy. 2009
In another scenario, you might not be interested in the topological
properties of the full network. What’s interesting for you is just
estimating the local properties of one – or more – nodes. In that case, 33
Manos Papagelis, Gautam Das, and
specialized node-centric strategies can be used33 . Nick Koudas. Sampling online social
networks. IEEE Transactions on knowledge
and data engineering, 25(3):662–676, 2013
25.5 Network Completion
if it was a standard BFS. But, since you don’t know this piece of
information, you need a general strategy working regardless of the
shape of the initial sample.
Naively, you might think to just go and probe the nodes with the
highest degree. However, there are a few considerations to make.
First, since – by definition – your sample is incomplete, you don’t
really know the true degree of a node. You only know how many
neighbors it has in your sample. Second, since the node has a high
degree in your sample, there’s some chance you already explored all
its neighbors, thus probing it won’t help you. 34
Sucheta Soundarajan, Tina Eliassi-
The first technique, MaxReach34 , estimates the true degree of a Rad, Brian Gallagher, and Ali Pinar.
node and its clustering coefficient using the information gathered Maxreach: Reducing network incom-
pleteness through node probes. In 2016
so far. It does so with a technique similar to Re-Weighted Random
IEEE/ACM International Conference on
Walk. The difference is that, in RWRW, we only want to know how Advances in Social Networks Analysis
many nodes have a given degree i. In MaxReach, we want to also and Mining (ASONAM), pages 152–157.
IEEE, 2016
know which nodes have that given degree value. At this point, the
score of a node is the difference between its estimated degree and its
degree in the sample. Nodes with higher scores are probed earlier.
After each probe, since we gathered more information in the sample,
MaxReach will recalculate the degree estimates. 35
Sucheta Soundarajan, Tina Eliassi-
e-wgx35 is a more recent alternative. Rad, Brian Gallagher, and Ali Pinar.
ε-wgx: Adaptive edge probing for
enhancing incomplete networks. In
25.6 Summary Proceedings of the 2017 ACM on Web
Science Conference, pages 161–170. ACM,
2017
1. Network sampling is a necessary operation when the network you
need to analyze is too large and/or you need to gather data one
node/edge at a time from a high latency source (e.g. the API of a
social media platform). Sometimes the decision is not up to you
and all you can access is a sample made by somebody else.
6. When sampling from real API systems one has to be careful that
the throughput in edges per second is not necessarily a good
indicator of how quickly you can gather a representative sample.
Due to pagination, high-throughput sources might return smaller
samples.
25.7 Exercises
2. Compare the CCDF of your sample with the one of the original
network by fitting a log-log regression and comparing the ex-
ponents. You can take multiple samples from different seeds to
ensure the robustness of your result.
Mesoscale
26
Homophily
don’t study this fact any more because it’s so boringly obvious. In 7
Yuexin Jiang, Daniel I Bolnick, and
this, we’re truly similar to other animals we often look down to7 . Mark Kirkpatrick. Assortative mating in
Rather than asking whether romantic ties show homophily, it’s animals. The American Naturalist, 181(6):
E125–E138, 2013
more interesting to use the degree of homophily of romantic ties to
compare societies.
8
Salvatore Scellato, Anastasios Noulas,
In Figure 26.3 you see an example of mixed marriage in the United
Renaud Lambiotte, and Cecilia Mascolo.
States. To that diagonally dominated matrix, you have to add the Socio-spatial properties of online
consideration that the United States is probably one of the most location-based social networks. In
ICWSM, 2011
diverse countries in the world. Imagine how this would look like 9
Kerstin Sailer and Ian McCulloh.
elsewhere! Social networks and spatial configura-
tion—how office layouts drive social
interaction. Social networks, 34(1):47–58,
2012
(a) (b)
Once you have an ego network, you can start investigating its
“global” properties such as the degree distribution or its homophily,
and these are not properties of the global network as a whole, but of
the local neighborhood of the ego, the ego network, which lives in
the mesoscale. Ego networks are frequently used in social network 18
Stephen P Borgatti, Ajay Mehra,
analysis18 , 19 , for instance to estimate a person’s social capital20 . Daniel J Brass, and Giuseppe Labianca.
Network analysis in the social sciences.
A consequence of this procedure is that we know that an ego science, 323(5916):892–895, 2009
node is connected to all nodes in its ego network. This is unfortunate 19
Jure Leskovec and Julian J Mcauley.
in some cases, depending on our analytic needs. For instance, all Learning to discover social circles in
ego networks. In Advances in neural
ego networks have a single connected component and will have information processing systems, pages
a diameter of two. If those forced properties are undesirable, one 539–547, 2012
can extract an ego network and then remove the ego and all its 20
Stephen P Borgatti, Candace Jones,
and Martin G Everett. Network
connections. measures of social capital. Connections,
21(2):27–36, 1998
∑ eii − ∑ ai bi
i i
r= ,
1 − ∑ a i bi
i
node with value i, and bi is the probability that an edge has as desti-
nation a node with value i. In an undirected network, the latter two
are equal: ai = bi . This formula takes values between −1 (perfect dis-
assortativity) and 1 (perfect assortativity: each attribute is a separate
component of the network).
2 2
((8 /22)+(12/ 22))−((10/ 22) +(14 / 22) ) Figure 26.6: How to calculate
2 2
1−((10 / 22) +(14 / 22) )
homophily using the formula in
the text.
~0.766
In Figure 26.6 we have two values i: red and green. There are 22
edges in the graph: eight green-green edges – thus the probability
is 8/22 – and 12 red-red edges – thus the corresponding eii value is
12/22. Ten edges originate (or end) in a green node: ai = bi = 10/22;
and 14 originate (or end) in a red node: ai = bi = 14/22. The final
value of homophily is ∼ 0.766. This value is interpretable as a sort of
Pearson correlation coefficient, which means that 0.766 is pretty high.
and absent. The terminology should not fool you. In this case, we
are not referring to the edge’s weight (Section 3.3). This is rather
a categorical difference, more akin to multilayer networks (Section
4.2). A weak tie is established between individuals whose social
circles do not overlap much. A strong tie is the opposite: an edge
between nodes well embedded in the same community. The absent
tie is more of a construct in sociology, which lacks a well defined
counterpart in network science. It can be considered as a potential
connection lurking in the background. For instance, there is an
absent tie between you and that neighbor you always say “hello” to
but never interact beyond that. You could consider an absent tie as
one of the most likely edges to appear next, if you were to perform a
classical link prediction (Part VI).
You can see now that you can have strong, weak, and absent ties
in an unweighted network. We can, of course, expect a correlation
between being a weak tie and having a low weight. However, we can
construct equally valid scenarios in which there is an anti-correlation
instead. For instance, we could weight the edges by their edge be-
tweenness centrality (Section 11.2). A weak tie must have a high edge
betweenness, because by definition it spans across communities and
thus all the shortest paths going from one community to the other
must pass through it.
Note that, notwithstanding their usefulness in favoring informa-
tion spread, weak ties are not the only game in town in a society. The
competing concept of the “strength of strong ties” shows that strong
ties are important as well. They are specifically useful in times of 25
David Krackhardt, N Nohria, and
uncertainty: “Strong ties constitute a base of trust that can reduce B Eccles. The strength of strong ties.
Networks in the knowledge economy, 82,
resistance and provide comfort25 ”. 2003
homophily 371
28
[Link]
drinking28 .
Since humans are social animals and tend to succumb to peer
pressure, homophily can be a channel for behavioral changes. In a
health study, researchers looked at health indicators from thousands
of people in a community over 32 years. They saw that behavior and
health risks that are not contagious actually are. For instance obesity:
if you have an obese friend, the likelihood of you becoming obese 29
Nicholas A Christakis and James H
increases by 57% in the short term29 . This is like the Susceptible- Fowler. The spread of obesity in a
Infected epidemic models we saw, even if obesity is not a biological large social network over 32 years.
New England journal of medicine, 357(4):
virus. It is rather a social type of virus.
370–379, 2007
Same with smoking, although in this case it worked the opposite: 30
Nicholas A Christakis and James H
people were quitting in droves30 . This is due to social pressure and Fowler. The collective dynamics of
homophily: a behavior you might not adopt by yourself is brokered smoking in a large social network. New
England journal of medicine, 358(21):
by your social circle, which you trust because it is made by people
2249–2258, 2008
like you – it speaks to your identity.
31
Lada A Adamic and Natalie Glance.
Another paper shows strong homophily in political blogs31 . In The political blogosphere and the
Figure 26.10 we see a visualization of how people writing online 2004 us election: divided they blog.
In Proceedings of the 3rd international
about politics connect to each other. A common political vision is the
workshop on Link discovery, pages 36–43.
clear driving force behind the creation of an hyperlink from one blog ACM, 2005
to another.
homophily 373
26.5 Summary
26.6 Exercises
kv
possible explanation.
Figure 27.1: A scatter plot we
can use to visualize degree
assortativity. For each edge, we
have the degree of one node on
the x axis and of the other node
u-v on the y axis.
ku
v u
5 3
ku
100 1000
avg(kN )
u
u
100
real world networks: (a) co-
authorship in scientific pub-
1 10 100
10
1 10 100 1000 10000
lishing, (b) P2P network, (c)
ku ku
Internet routers, (d) Slashdot
social network.
(a) (b)
1000 1000
100
avg(kN )
avg(kN )
u
u
100
10
1 10
1 10 100 1000 1 10 100 1000
ku ku
(c) (d)
1 1 1
Figure 27.5: The
2 2 2
(dis)assortativity inducing
4 3 4 3 4 3 model. (a) Select two pairs
of connected nodes (in green
the edges we select). (b) As-
sortativity inducing move. (c)
Disassortativity inducing move.
(a) (b) (c)
Note that this swap doesn’t always change the topology nor
alter the characteristics of the network. For instance, if all nodes
have the same degree, the move would not affect assortativity. But,
after enough trials in a large enough network, you’ll see that these
operations will have the desired effect.
Degree assortativity, as I discussed it so far, is defined for undi-
rected networks. There are straightforward extensions for directed 15
Jacob G Foster, David V Foster, Peter
networks15 . The standard strategy is to look at four correlation coeffi- Grassberger, and Maya Paczuski. Edge
cients: in-degree with in-degree, in-degree with out-degree (and vice direction and the structure of networks.
Proceedings of the National Academy of
versa), and out-degree with out-degree.
Sciences, 107(24):10815–10820, 2010
Obviously, everything I wrote so far on degree assortativity also
380 the atlas for the aspiring network scientist
In other words, the identity line divides the space in two. Above
the identity line we have all the nodes for which, on average, the
neighbor degree is higher than the node’s degree. Below the identity
line it’s the opposite: the node’s degree is higher than the neighbors
degree.
At first glance, the situation seems balanced. There are as many
points above the identity line as there are below. However, remember
that we’re aggregating all nodes with the same degree value in a sin- 16
Scott L Feld. Why your friends have
gle point. We know that the degree has a broad distribution, because more friends than you do. American
Journal of Sociology, 96(6):1464–1477,
we visualized it for the co-authorship network before. Therefore,
1991
there are way more nodes above the identity line than below. 17
Ezra W Zuckerman and John T Jost.
That’s the friendship paradox: your friends are, on average, more What makes you think you’re so pop-
ular? self-evaluation maintenance and
popular than you16 , 17 ! This means that, for the average node, its
the subjective side of the" friendship
degree is lower than the average degree of their neighbors. This is paradox". Social Psychology Quarterly,
actually pretty obvious once you think about it: a node with degree pages 207–223, 2001
quantitative assortativity 381
80
Figure 27.7: The data model
100 of the business to business net-
work. The edge color tells corre-
100
sponds to the node making the
claim about the transaction. For
90
instance, the green node reports
selling 90 to the red node and
buying 80 back.
This trustworthiness score is a quantitative attribute. It is strongly
correlated with the likelihood that the business was in fact cheating
on their taxes, as I have information whether the audited businesses
were fined and, if they were, how much they had to pay.
Simulations show that, with this correction, in a randomly wired
network the score should be disassortative. Instead, in the real ob-
served network, the score is assortative. Figure 27.8 shows the re-
lation: there are more nodes above the identity line than below – it
382 the atlas for the aspiring network scientist
100
Figure 27.8: The average trust-
# Nodes
100
non-neighbors (y axis). The
blue line shows the identity,
10-2 and the point color the number
10
of nodes with the given value
combination.
10-3 -3 1
10 10-2 10-1 100
Avg Neigh Trust Diff
might appear that the opposite is true, but you need to take into ac-
count the color of the dots. Being above the identity line means that
the trust score difference between neighbors is lower than between
non-neighbors: a sign of assortativity.
Given the connection with the actual tax fines, the assortativity
analysis can make us conclude something tangible about business
connections. In this case, that fraudulent and untrustworthy busi-
nesses band together. If I know your customer/suppliers are scam-
mers, I should update my priors on whether you’re a scammer as
well.
A more famous example of quantitative attribute assortativity is
related to the friendship paradox that I just described in the previous
section, and it is ten times more enraging. We humans are social
animals. Notwithstanding introverted cavemen like myself, in general
our level of happiness is correlated with the number of friends we
have. In fact, just like the degree, happiness is assortative in social 19
Johan Bollen, Bruno Gonçalves,
networks: happy people tend to befriend each other19 . Guangchen Ruan, and Huina Mao.
What I just stated is that more friends imply more happiness. And Happiness is assortative in online social
networks. Artificial life, 17(3):237–251,
the friendship paradox tells us that our friends have more friends
2011
than us. Do I mean to tell that, like with friendship, there is also a 20
Johan Bollen, Bruno Gonçalves, Ingrid
happiness paradox? Why, yes there is20 . If you ask people about van de Leemput, and Guangchen Ruan.
their level of happiness in a social network, you will find out that the The happiness paradox: your friends
are happier than you. EPJ Data Science, 6
average happiness level of one’s friends tends to be higher than their
(1):4, 2017
own happiness level. That is probably why you think everyone is
having such a great time on social media. Everyone but you. It’s not
you, it’s the system. Luckily, the researchers behind this discovery
have a few guidelines on how to unplug from social media toxicity 21
Johan Bollen and Bruno Gonçalves.
and live a more fulfilling life21 . Network happiness: How online social
In general, everything that correlates with degree – be it happiness, interactions relate to our well being. In
Complex Spreading Phenomena in Social
income, or tax fraud – will get its own paradox for free.
Systems, pages 257–268. Springer, 2018
quantitative assortativity 383
27.4 Summary
27.5 Exercises
When you obtain a new network dataset and you plot it for the
first time, in the vast majority of cases you will see a blobbed mess.
This is usually due to the fact that raw network data is usually a
hairball, and you need to backbone it, or perform other data cleaning
tasks, as I detailed in Part VII. However, in some cases, there is an
unobjectionable truth. It might be that, deep down, your network
really is a hairball.
Many large scale networks have a common topology: a very
densely connected set of core nodes, and a bunch of casual nodes at-
taching only to few neighbors. This should not be surprising. If you
create a network with a configuration model and you have a broad
degree distribution, the high degree nodes have a high probability
of connecting to each other – see Section 15.1. The surprising part
is that the cores of some empirical networks are even denser than
what you’d anticipate by looking at the degree distribution of the 1
Shi Zhou and Raúl J Mondragón. The
network1 ! rich-club phenomenon in the internet
topology. IEEE Communications Letters, 8
Since everything that departs from null expectation is interesting,
(3):180–182, 2004
this phenomenon in real world networks has attracted the attention
2
Petter Holme. Core-periphery orga-
of network scientists. They gave a couple of names to this special
nization of complex networks. Phys.
meso-scale organization of networks: core-periphery2 , 3 , with the core Rev. E, 72:046111, Oct 2005. doi:
sometimes dubbed as “rich club”4 . 10.1103/PhysRevE.72.046111
28.1 Models
p Nothing
here
lighted in blue, a dense area of
the network with many connec-
tions. In green, a sparser area:
the periphery. Connections only
go to (or from) a core member,
There can be only two types of connections: between core nodes meaning that in the main diago-
– which is the most common edge type, since the core is densely nal in the peripheral area there
connected – and between a core-periphery pair. Peripheral nodes do are no entries larger than zero.
not connect to each other. In the adjacency matrix, which I show in
Figure 28.1, there’s a big area with no connections. This is known
as the “Discrete Model”. It is a very strict one, and rarely real world
networks comply with this standard. A perfect discrete model in
which the core is composed by a single node is a star.
If you want to detect the core-periphery structure using the dis-
crete model, you have a simple quality measure you want to maxi-
mize. This is ∑ Aij ∆uv , with A being the adjacency matrix, and ∆ a
uv
matrix with a value per node pair. An entry in ∆ is equal to one if
either of the two nodes is part of the core.
386 the atlas for the aspiring network scientist
Continuous
Reality rarely conforms with strict expectations. Having only two
classes in which to put nodes is exceedingly restrictive. What if nodes
can be sorted in three classes? What if a semi-periphery exists? This
is an enticing opportunity, until you realize that you could also ask:
why three classes? Why not four? Why not five? Why not... you get
the idea.
Other Approaches
The continuous model is powerful, but it doesn’t really tell you much 7
M Puck Rombach, Mason A Porter,
on how you should build your ci values. Rombach et al.7 propose James H Fowler, and Peter J Mucha.
a way to build such vector, introducing two parameters, α and β. β Core-periphery structure in networks.
SIAM Journal on Applied mathematics, 74
determines the size of the core, from the entirety of the network to
(1):167–190, 2014
an empty core. α regulates the c score difference between the core
classes. If a node u at a specific core level has a score cu , the node v
at the closest highest core level will have cv = α + cu – or, really, any
function taking α as a parameter.
(a)
(b)
However, a traditional community is also sparsely connected to
nodes outside the community. This means that, if nodes u and v
are in different communities, likely Auv = 0. But we just saw that
their ∆uv should be 1 because they have high degree! All of that
score is wasted! Maximizing the quality function would imply to
core-periphery 389
put all nodes in the same core, sacrificing the defining characteristic
of a core: the fact that nodes in it should tend to connect to each
other. Figure 28.4 shows the difference between the two archetypal
meso-scale organizations.
This is problematic since we have evidence that core-periphery
structures are ubiquitous, and so are communities. There are a
couple of explanations we can use to restore our sanity.
The first explanation is realizing that every network lives on a
core-periphery to community structure continuum. The real world
networks we observe distribute through this continuum in such a
way that perfect instances are extremely rare – as Figure 28.5 shows.
You’ll very rarely find a natural discrete model, exactly as rarely
as finding a real world network organizing like a caveman graph
(Section 14.1) – the quintessential community structure. As a conse-
quence, one way to solve this conundrum is to admit that a network
might have multiple cores. Thus, one first performs community dis-
covery to find the multiple cores and then applies the core-periphery 12
Sadamori Kojaku and Naoki Masuda.
detection algorithm independently on each community12 . Core-periphery structure requires
something else in the network. New
Journal of Physics, 20(4):043012, 2018
# Networks
The blend is always different! Figure 28.5: The number of
networks with a pure core-
periphery network and with a
pure community structure is
actually tiny.
13
Jaewon Yang and Jure Leskovec.
the network13 , bringing together the community structure with the Overlapping communities explain core–
core-periphery organization. periphery organization of networks.
Proceedings of the IEEE, 102(12):1892–
Alternative explanations use the power of random walkers to ex-
1902, 2014
plain core-peripheries14 , an approach that is also commonly used in 14
Fabio Della Rossa, Fabio Dercole, and
community discovery. Other models attempt to embed nodes into a Carlo Piccardi. Profiling core-periphery
network structure by random walkers.
spatial dimentions, showing how this can create core-periphery struc- Scientific reports, 3:1467, 2013
tures15 . It is following this example that I’ll try to draw a connection 15
Daniel A Hojman and Adam Szeidl.
between core-periphery and some well studied aspects of economics Core and periphery in networks. Journal
of Economic Theory, 139(1):295–309, 2008
in the next section.
and, since the two vendors offer the same quality, the customers
will go to the closest vendor. Therefore, a rational vendor would
move their stand so that it can capture the people in the middle. The
other vendor would do the same. The solution is an equilibrium in
which vendors concentrate in the middle, even though that means
increasing the walk length for every customer.
28.4 Nestedness
(a) (b)
28.5 Summary
28.6 Exercises
3. [Link] contains
a nested bipartite network. Draw its adjacency matrix, sorting
rows and columns by their degree.
29
Hierarchies
8
Enys Mones, Lilla Vicsek, and Tamás
them in three categories: order, nested, and flow hierarchy8 . I’ll Vicsek. Hierarchy measure for complex
present the three of them in this section, noting how this chapter will networks. PloS one, 7(3):e33799, 2012
then only focus on flow hierarchy. Order and nested hierarchies are
covered elsewhere in this book with different names.
Order
In an order hierarchy, the objective is to determine the order in which
to sort nodes. We want to place each node to its corresponding level,
according to the topology of its connections. Usually, this is achieved
by calculating some sort of centrality score. The most central nodes
are placed on top and the least central on the bottom.
5
8 Figure 29.1: (a) A toy network.
(b) Its order hierarchy. I place
2
9 3 nodes in descending order of
5 betweenness centrality from top
7 3 1
to bottom.
9
6
4
6 7 8 2 4 1
(a)
(b)
Figure 29.1 provides an example of detecting an order hierarchy in
a toy network, using betweenness centrality as the guiding principle.
Node 5 has the highest betweenness centrality, followed by node
3 and then node 9. All other nodes have the same betweenness
centrality – equal to zero.
One can easily see that we already covered this sense of hierarchi-
cal organization of complex networks. The order hierarchy is nothing
more than a different point of view of node centrality. Thus, I refer to
Chapter 11 for a deeper discussion on the topic.
Nested
Nested hierarchy is about finding higher-order structures that fully
contain lower order structures, at different levels ultimately ending
in nodes. In the corporation example, the largest group is the cor-
poration itself, encompassing all workers. We can first subdivide
the corporation into branches, if it is a multinational, they could
be regional offices. Each office can be broken down into different
departments, which have teams and, finally, the workers in each
team.
Figure 29.2 provides an example of detecting a nested hierarchy in
a toy network. This is usually done by detecting smaller and smaller
398 the atlas for the aspiring network scientist
8
Figure 29.2: (a) A toy network.
2 Colored circles delineate nested
9
5 substructures. (b) Its nested
hierarchy, according to the
7 3 1
6 highlighted substructures. Each
node and substructure is con-
8 9 5 6 7 2 3 4 1
4 nected to the substructure it
(b) belongs to.
(a)
densely connected units in the network. Note how, in this case, the
hierarchy does not place nodes on levels, but organizes the detected
substructures.
This is equivalent to performing hierarchical community discovery
on complex networks. Thus, I refer to Chapter 33 for further reading.
Flow
In a flow hierarchy, nodes in a higher level connect to nodes at the
level directly beneath it, and can be seen as managers spreading infor-
mation or messages to the lower levels. We call it a “flow” hierarchy
because you can see the highest level node as the origin of a flow,
which hits first the nodes at the level directly beneath it, and so on
until it reaches the leaves of the network: the nodes at the bottom
layer.
8
5 Figure 29.3: (a) A toy network.
2 (b) Its flow hierarchy.
9
5
8 9 7 3 2
7 3 1
6
6 4 1
4
(b)
(a)
Figure 29.2 provides an example of detecting a flow hierarchy in a
toy network. Note how it tends to have high centrality nodes on top,
like the order hierarchy, but it creates a substantially different orga-
nization. Nodes directly connected to a given level tend to belong to
the level immediately beneath it, no matter their different centrality
values. One could think that the order hierarchy is a special case of
flow hierarchy, but that is incorrect: in a flow hierarchy all nodes
belonging to a level need to connect to the level directly below and
above, while that’s not the case for the order hierarchy. Moreover,
hierarchies 399
29.2 Cycles
1 1
Figure 29.4: (a) A directed net-
2 3 2 3 work. In the figure, I highlight
in blue the edges partaking
2
6 12 13 9 9 in cycles. (b) The condensed
version of (a), where all nodes
4 5 14 15 part of a strongly connected
component are condensed in a
7 8 16 10 11 7 8 16 10 11 node (colored in blue).
(a) (b)
1
|V | − 1 v∑
GRC = LRC MAX − LRCv .
∈V
In practice, you average out its difference with all reach centrality
values in the network. This is an effective way of counteracting the
degeneracy of cycle-based hierarchy measures. In both toy examples
from Figure 29.5, GRC is well behaved, returning values of 0.555 and
0.22, respectively.
However, GRC has a blind spot of its own. Since we’re averaging
the differences between the most central node against all others, we
know we will never get a perfect GRC score if there is more than one
node with non-zero local reach centrality. Consider Figure 29.6. I
don’t know about you, but to me it looks like a pretty darn perfect
hierarchy. Yet, we know that the two nodes connected by the root
don’t have a zero local reach centrality. In fact, the GRC for that
network is 0.89.
So, if cycle-based flow hierarchy is too lenient – every directed
402 the atlas for the aspiring network scientist
29.4 Arborescences
29.5 Agony
A( G, l ) = ∑ max(lv − lu + 1, 0).
(u,v)∈ E
Every time lu < lv , we contribute zero to the sum. Note that agony
404 the atlas for the aspiring network scientist
requires you to specify the lu value for all nodes in the network. This
is not usually something you know beforehand. So the problem is to
find the lu values that will minimize the agony measure. There are 15
Nikolaj Tatti. Hierarchies in directed
efficient algorithms to estimate the agony of a directed graph15 . networks. In 2015 IEEE international
conference on data mining, pages 991–996.
IEEE, 2015
Figure 29.8: Two hierarchies
with different values of agony.
The vertical positioning of each
node determines its level, from
top (lu = 1) to bottom (lu = 4).
(a) (b)
Consider Figure 29.8. In both cases, we have only one edge going
against the flow. Agony, however, ranks these two structures differ-
ently. In Figure 29.8(a), the difference in rank is only of one, thus the
total agony is 2. In Figure 29.8(b), the difference in rank is 3, resulting
in a higher agony. Also a cycle-based flow hierarchy measure takes
different values, as Figure 29.8(b) involves more edges in a cycle (four
edges, versus just two in Figure 29.8(a)).
Ultimately, the resting assumption of agony is the same of the
cycle-based flow hierarchy. Agony considers any directed acyclic
graph as a perfect hierarchy. Thus it will give perfect scores to the
imperfect hierarchies from Figure 29.5.
To wrap up this chapter, note that all these methods have a various
degree of graphical flavor to them. Meaning that you can use them to
create a picture of your hierarchy, which might help you to navigate
the structure. The most rudimentary method is the cycle-based flow
hierarchy, because it just reduces the graph to a DAG, which doesn’t
help you much.
Global reach centrality is better, as you can place nodes on a ver-
tical level according to their local reach centrality value. In Figure
29.9(b) I apply a reach centrality informed layout to the directed
graph from Figure 29.9(a). The reach centrality layout is quite rudi-
mentary, as the resulting picture looks more akin to an order hierar-
chy than a flow hierarchy. In the figure, I had to do a bit of manual
work to make it look more like a flow hierarchy, which you might not
be able to do for larger graphs. Since the method allows you to find
flow hierarchies, this mismatch could be confusing. However, at least
hierarchies 405
it allows you to find out the root of the hierarchy, the node(s) with
the highest reach centrality, which the previous method could not do.
The arborescence approach is a further step up. Since it reduces
the network to an arborescence, one can draw the resulting con-
densed graph, identifying not only the root of the hierarchy, but at
which level each node lies. The same can be said for agony: it assigns
each node to a level, thus you can plot the network by layering nodes
vertically according to their assigned rank. Figure 29.9(c) shows a
possible layout informed by arborescence (or agony).
29.7 Summary
4. We can identify the head of the hierarchy as the node with the
highest reach. Alternatively, arborescences are prefect hierarchies
– directed acyclic graphs with all nodes having in-degree of one,
except the head of the hierarchy having in-degree of zero.
406 the atlas for the aspiring network scientist
29.8 Exercises
4. Perform the null model test you did for exercise 1 also for global
reach centrality and arborescence. Which method is farther from
the average expected hierarchy value?
30
High-Order Dynamics
The first natural way to have higher order memory in your network
analysis is by embedding it into the structure itself. This is a pow-
erful approach, because it allows you to use any non-high-order
algorithm you want. You have the entirety of the network analysis
toolbox at your disposal. The price you have to pay is that you need
to keep track of your operation. You need to reconstruct the original
structure if you want to properly interpret your results.
This practically boils down to performing a pre-processing on your
data structure and a post-processing on your results. Here we focus
mainly on the pre-processing as, hopefully, how to post-process the
results should be straightforward. Unfortunately, the pre-process is
not going to be as simple as I make it to be in Figure 30.1. You cannot
simply re-weight your edges, because the re-weighting would be
dependent on the current position of the agent in the network. Thus
you’d have to have a re-weighting for every node in the network.
Worse still, if you do that you’re just implementing second-order
dynamics: if you need to go to a higher order than that (say, you
need to remember the last two nodes through which you passed)
you’re in no better position than before.
I’m going to specifically focus on a few papers in this line of
research, but hopefully you could see how the general approach
in this category of high-order analysis works. Also note that other
structures described elsewhere in the book can count as high order
structures. For instance, hypergraphs connecting multiple nodes at
the same time, and networks with simplicial complexes, which are
high-order dynamics 409
Memory Network
3-5
3 Figure 30.4: (a) A graph. (b) Its
4-5 linegraph version.
2 4 3-4
5 2-4
1-4
1
1-2
(a)
(b)
Motif Dictionary
The first approach I consider is the one building motif dictionaries. In
this approach, one realizes that there are different motifs of interest
that have an impact on the analysis. For instance, one could focus
specifically on triangles. Once you specify all the motifs you’re
interested in, you take a traditional network measure and you extend
412 the atlas for the aspiring network scientist
| ES,S̄ |
φ(S) = ,
min (| ES |, | ES̄ |)
where S is one of the two groups – i.e. a set of nodes on one side
of the cut –, S̄ is its complement – i.e. S̄ = V − S, the set of nodes
on the other side of the cut. Ex is the set of edges in the set x, and
Ex,y is the set of edges established between a node in x and a node
in y. In practice, let’s find S such that φ(S) is minimum. We do so
by minimizing the number of edges between S and its complement,
normalized by their sizes (so that we don’t find trivial solutions by
cutting off simply a dangling leaf node).
But here we say: No! Not the number of edges! We are interested in
higher order structures! We want to minimize the number of triangles
between groups! How would that look like? Exactly the same. We
just count not the number of the edges, but the number of arbitrary
motifs M spanning between S and non-S:
| MS,S̄ |
φ(S, M ) = .
min (| MS |, | MS̄ |)
Boom. By finding an S minimizing this specific φ(S, M ) we just
made a high-order normalized cut. This can be done exactly like 5
Austin R Benson, David F Gleich, and
finding the “normal” normalized cut, by examining the eigenvectors Jure Leskovec. Tensor spectral clustering
for partitioning higher-order network
of a specially constructed Laplacian5 . Figure 30.5 shows an example.
structures. In Proceedings of the 2015
In the figure, the two groups still have tons of edges going from one SIAM International Conference on Data
group to another, this is hardly a solution for a regular normalized Mining, pages 118–126. SIAM, 2015
high-order dynamics 413
cut. But there are few triangles between S and S̄, thus this is a proper
solution for the higher order normalized cut. The shape of this
solution can be applied to many other problems.
In this category of solutions you find all the approaches who take
the equation of the Markov process and add a temporal factor. For
instance, suppose that we’re observing random walks (Chapter
8). We have a vector of probability pk that tells us the status of the
random walkers at time k, namely the probability of the walkers to
be in a node (pk is a vector of length |V |). Now, if we were doing
a normal random walk and perform one step, we’d know that the
next step is simply given by the stochastic adjacency matrix, i.e.
pk+1 = pk D −1 A (assuming A is the normal adjacency matrix, and
D the degree diagonal matrix). This is a discrete random walk on a
graph.
But what if we want higher order dynamics? What if we want to
look at the result of the process after t steps? Well, the idea is to add
the temporal information to the equation: pk+t = pk Tt . So what’s Tt ?
Tt is a matrix telling us the transition probability from u to v after t
steps. If you followed the linear algebra math in Chapter 8 you know
that, to perform a random walk of length t, you can simply raise
D −1 A to the power of t: ( D −1 A)t .
The limitation here is that time ticks discretely, i.e. the random 6
Michael T Schaub, Renaud Lambiotte,
and Mauricio Barahona. Encoding
walker makes one move per timestep. In many cases, you might want dynamics for multiscale community
to simulate the passage of time as a continuous flow6 . If you want detection: Markov time sweeping for
to do it, you need to derive a different Tt . First, rewrite D −1 A as the map equation. Physical Review E, 86
(2):026112, 2012b
− D −1 L. This changes the random walk from discrete to continuous7 . 7
Renaud Lambiotte, J-C Delvenne, and
This allows us to let time flow and take the result of the random Mauricio Barahona. Laplacian dynamics
−1
walk at time t: Tt = e−tD L . If we set t = 1, we recover exactly the and multiscale modular structure in
networks. arXiv preprint arXiv:0812.1770,
transition probabilities of a one-step random walk. But now we’re 2008
free to change t at will, to get second, third, and any higher order
Markov processes.
414 the atlas for the aspiring network scientist
Other Approaches
There is a whole bunch of other approaches to introduce high order
dynamics into your network structure. More than I can competently
cover. I provide here some examples with brief, but overly simplistic,
explanations.
One way is to use tensors (Section 5.2). In practice, you create a
multidimensional representation of the topological features of the
network. Each added dimension of the tensor represents an extra 8
Michael Chertok and Yosi Keller.
order of relationships between the nodes8 , 9 . Then, by operating on Efficient high order matching. IEEE
this tensor, you can solve any high order problem: in the papers I cite Transactions on Pattern Analysis and
Machine Intelligence, 32(12):2205–2215,
the problem the authors focus on is graph matching.
2010
Another general category of solutions is a collection of techniques 9
Olivier Duchenne, Francis Bach, In-So
to take high order relation data and transform it into an equiva- Kweon, and Jean Ponce. A tensor-
based algorithm for high-order graph
lent first order representation10 , 11 . The first order solution in this matching. IEEE transactions on pattern
structure then translates into the high order one, much like in the analysis and machine intelligence, 33(12):
2383–2395, 2011
HON and memory network approach. The cited techniques are also 10
Hiroshi Ishikawa. Higher-order clique
generally applied to the problem of finding a high order cut in the reduction in binary graph cut. In 2009
network. IEEE Conference on Computer Vision and
Pattern Recognition, pages 2993–3000.
Being able to study high order interactions can help you making
IEEE, 2009
sense of many complex systems. For instance, they have been used 11
Alexander Fix, Aritanan Gruber,
to explain the remarkable stability of biodiversity in complex eco- Endre Boros, and Ramin Zabih. A graph
cut algorithm for higher-order markov
logical systems12 . Other application examples include the study of random fields. In 2011 International
infrastructure networks13 – how to track the high order flow of cars Conference on Computer Vision, pages
in a road graph –, and aiding in solving the problem of controlling 1020–1027. IEEE, 2011
12
Jacopo Grilli, György Barabás,
complex systems14 – which we introduced in Section 18.4. Matthew J Michalska-Smith, and
Stefano Allesina. Higher-order interac-
tions stabilize dynamics in competitive
30.3 Summary network models. Nature, 548(7666):210,
2017
1. Classical network analysis is single-order: only the direct connec- 13
Jan D Wegner, Javier A Montoya-
Zegarra, and Konrad Schindler. A
tions matters. But many phenomena they represent are high-order:
higher-order crf model for road net-
in a flight passenger networks, the nodes you just visited greatly work extraction. In Proceedings of the
influence to which node you’ll move next. IEEE Conference on Computer Vision and
Pattern Recognition, pages 1698–1705,
2013
2. There are two approaches to high-order networks: embedding the 14
Abubakr Muhammad and Magnus
dynamics in the structure, meaning that we modify the network Egerstedt. Control using higher order
data so that it has “memory”; or embedding it in the analysis, laplacians in network topologies. In
Proc. of 17th International Symposium
giving to the algorithm the task of remembering previous moves.
on Mathematical Theory of Networks and
Systems, pages 1024–1038. Citeseer, 2006
3. In High Order Networks you split each node to represent all the
paths that lead to it. In memory networks you instead model
second order dynamics with a line graph, third order dynamics
with the line graph of the line graph, and so on.
the classical Markov equation (e.g. for random walks), and other
approaches.
30.4 Exercises
Communities
31
Graph Partitions
We have reached the part of network analysis that has probably re-
ceived the most attention since the explosion of network science in
the early 90s: community discovery (or detection). To put it bluntly,
community discovery is the subfield of network science that pos-
tulates that the main mesoscale organization of a network is its
partition of nodes into communities. Communities are groups of
nodes that are very related to each other. If two nodes are in the
same community they are more likely to connect than if they are in
two distinct communities. This is an overly simplistic view of the
problem, and we will decompose this assumption when the time
comes, but we need to start from somewhere.
So the community discovery subfield is ginormous. You might
ask: “Why?” Why do we want to find communities? There are many
reasons why community discovery is useful. I can give you a couple
of them. First, this is the equivalent of performing data clustering
in data mining, machine learning, etc. Any reason why you want
to do data clustering also applies to community discovery. Maybe
you want to find similar nodes which would react similarly to your
interventions. Another reason is to condense a complex network into
a simpler view, that could be more amenable to manual analysis or
human understanding.
More generally, decomposing a big messy network into groups
is a useful way to simplify it, making it easier to understand. The
reason why there are so many methods to find communities – which,
as we’ll see, rarely agree with each other – is because there are innu-
merable ways to simplify a network.
It is difficult to give you a perspective of how vast this subfield of
network science is. Probably, one way to do it is by telling you that
there are so many papers proposing a new community discovery
algorithm or discussing some specific aspect of the problem, that
making a review paper is not sufficient any more. We are in need
of making review papers of review papers of community discovery.
418 the atlas for the aspiring network scientist
be a wild ride. 14
Michele Coscia, Fosca Giannotti,
and Dino Pedreschi. A classification
for community discovery methods in
31.1 Stochastic Blockmodels complex networks. SADM, 4(5):512–546,
2011
15
Fragkiskos D Malliaros and Michalis
Classical Community Definition Vazirgiannis. Clustering and community
detection in directed networks: A
When people started mapping complex systems as networks, they survey. Physics Reports, 533(4):95–142,
2013
realized that the edges didn’t distribute randomly across nodes. We 16
Nguyen Xuan Vinh, Julien Epps, and
already saw that deviating from random expectation is a source of James Bailey. Information theoretic
interest when we talked about degree distributions (Section 6.2). measures for clusterings comparison:
Variants, properties, normalization and
On top of the rich-get-richer effect, researchers realized that edges
correction for chance. J. Mach. Learn. Res,
distributed unevenly among different groups of nodes. Especially, 11(Oct):2837–2854, 2010
but not only, in social networks, there were lumps of connections, 17
Vinh-Loc Dao, Cécile Bothorel, and
Philippe Lenca. Estimating the similarity
separated by very sparse areas.
of community detection methods
The ideal scenario is something resembling Figure 31.1(a). The nat- based on cluster size distribution.
ural next step was trying to see if we could separate these lumps of In International Workshop on Complex
Networks and their Applications, pages
connections into coherent groups: communities. This created the first 183–194. Springer, 2018b
and most commonly accepted definition of a network community: 18
Vinh-Loc Dao, Cécile Bothorel, and
Philippe Lenca. Community structure:
Communities are groups of nodes densely connected to each other and A comparative evaluation of community
detection methods. arXiv preprint
sparsely connected to nodes outside the community.
arXiv:1812.06598, 2018a
19
Amir Ghasemian, Homa Hossein-
I would call this the classical definition of a network community. mardi, and Aaron Clauset. Evaluating
This definition can be attacked and deconstructed from multiple overfit and underfit in models of net-
parts but, for now, let’s accept it. Note that, for now, we assume that work community structure. TKDE,
2019
a node can only be part of a single community. We use the term
“partition” to refer to the assignment of nodes to their community.
(b)
(a)
In the early community discovery days20 – before we even had
20
Paul W Holland, Kathryn Blackmond
Laskey, and Samuel Leinhardt. Stochas-
coined the term “community” –, the main approach was using tic blockmodels: First steps. Social
stochastic blockmodels. We already introduced them in Section networks, 5(2):109–137, 1983
15.2 as a method to generate a synthetic graph. How do we apply
them to the problem of finding communities? The first step in our
quest is changing the perspective over the graph from Figure 31.1.
420 the atlas for the aspiring network scientist
Maximum Likelihood
(b)
(a)
Lθ,A = ∑ lθ,A,u,v .
u,v∈ A
These are only a few examples of the many papers using this ap-
proach – you’re going to hear this excuse from me a lot, to save
myself from citing literally everything and making this a book about
community discovery.
101101
1001 110011 Figure 31.5: A binary node ID
schema you’d use to encode
a random walk. Each colored
arrow points to the ID of the
three nodes involved in the
orange walk.
1001101101110011
The most known and best performing of these approaches is usu- 33
Martin Rosvall and Carl T Bergstrom.
ally considered the map equation approach, or Infomap33 , 34 . The Maps of random walks on complex
networks reveal community structure.
map equation is what you use to encode the random walk informa- Proceedings of the National Academy of
tion with the minimum possible number of bits – i.e. minimizing the Sciences, 105(4):1118–1123, 2008
“code length”. Suppose that you give each node a binary ID. Since 34
Ludvig Bohlin, Daniel Edler, Andrea
Lancichinetti, and Martin Rosvall.
we have 36 nodes, we need around 5 bits, but we can save a little if Community detection and visualization
we give shorter codes to central nodes (they are going to be visited of networks with the map equation
more often). Then the cost of describing the random walk is simply framework. In Measuring Scholarly
Impact, pages 3–34. Springer, 2014
the length of the code of the node multiplied by the number of times
we’re going to see it in a random walk, which is given by the station-
ary distribution (Section 8.1). In the example from Figure 31.5, the
orange walk is fully described by the bits in the figure: the node id
sequence.
Infomap saves bits by using community prefixes. Nodes in the
same community get the same prefix. So now we need fewer bits to
uniquely refer to each node, because we can prepend the community
code the first time and then omit it as long as we are in the same
community. Since a community contains, in this case, 9 nodes instead
of 36, we can use shorter codes. We need to add an extra code that
allows us to know we’re jumping out of a community. This is an
overhead, but the assumption here is that a random walker will
spend most of its time in the community, so this community prefix
and jump overhead is rarely used. Figure 31.6 shows this re-labeling
process.
graph partitions 425
00 01
101 Figure 31.6: The large two-digit
100 101101 1100 codes are the IDs of the com-
1001 110011
munities. Each node gets a
new shorter ID, given that IDs
need to be community-unique,
rather than network-unique.
Now the random walk uses
the prefix (in red) to indicate in
Prefix Node IDs in the walk Jumping out
which community it is, then the
01 1100 101 100 1111 new shorter node IDs (in blue)
10 11 Rarely used! and finally adds an extra ID to
indicate it’s jumping out of a
community (in green).
If the partition is good, we can compress the random walk in-
formation by a lot. Consider the example in Figure 31.7. Without
communities we have no overhead, but we need to fully encode our
36 nodes. The path in orange is simply the sequence of node IDs and
can be stored in 72 bits. If we have community partitions, we add
the community prefixes and the jump overhead (for the community
jump in brown), but the node IDs are shorter. The encoding of the
same walk is 56 bits, and we can see that the overhead parts are tiny
compared with the rest.
10 11
1000010001101001010110110110110011001000000010001101000
Without 11001000011100101
72 bits
With 00110111001001110001010000111110100111011000000010101101 56 bits
If my explanation still makes little sense, you can try out an in-
teractive system showing all the mechanics of the map equation 35
[Link]
approach35 . Infomap has been adapted to numerous scenarios. Many [Link]
Evolutionary Clustering
42
Qing Cai, Lijia Ma, Maoguo Gong,
A couple of good review works42 , 43 focus on dynamic community and Dayong Tian. A survey on network
discovery and can help you obtaining a deeper understanding of this community detection based on evolu-
tionary computation. IJBIC, 8(2):84–98,
problem. Let’s explore what can happen to your communities over
2016
time.
Time
Figure 31.9: Two things that can
happen to your communities in
an evolving network: growing
Grow and shrinking.
43
Giulio Rossetti and Rémy Cazabet.
Community discovery in dynamic
networks: a survey. ACM Computing
Surveys (CSUR), 51(2):35, 2018
Shrink
One possibility is that the community will grow: it will attract new
nodes that were previously unobserved. The other side of the coin is
shrinking: nodes that were part of the community disappear from the
network. Figure 31.9 shows visual examples of these events.
Time
Figure 31.10: Two things that
can happen to your commu-
nities in an evolving network:
Merge merging and splitting.
Split
Time
Figure 31.11: Two things that
can happen to your commu-
nities in an evolving network:
Birth birth and death.
Death
lowest number of bits. Let’s say this is its quality function – which is
known as code length (CL).
In evolutionary clustering you don’t just optimize CL. You have
CL as a term in your more general quality function Q. The other
term in Q is consistency. For simplicity sake, let’s just assume it is
some sort of Jaccard coefficient between the partitions at time t and
the partition at time t − 1. To sum up, a very simple evolutionary
clustering evaluates the partition pt at time t as:
Q pt = αCL pt + (1 − α) J pt ,pt−1 .
16
16 10 16
10 10
11 13 14 Figure 31.12: (a) The commu-
13 14 9 13 14
9 11 9 11
15
nity partition of a graph at time
15 12 15
12 12
t. (b) A partition of the graph at
6
time t + 1 exclusively optimizing
3 6 3 3 6
the code length, using Infomap.
2 4 5 7 2 4 5 7 2 4 5 7
(c) A partition of the graph at
1 8 1 8 1 8
time t + 1 balancing a good code
(a) (b) (c) length and consistency with the
Here, α is a parameter you can specify which regulates how much partition at time t + 1.
weight you want to give to your previous partitions. For α = 1 you
have standard clustering, while for α = 0 the new information is
discarded and you only use the partition you found at the previous
time step. Figure 31.12 shows you that maximizing CL pt might yield
significantly different results than maximizing a temporally-aware
Q pt function.
This is only one – the simplest – of the many ways to perform
smoothing, which the other review works I cited describe more
in details. However, all these methods (and the ones that follow)
have something in common: they are all at odds with the classical
definition of community that I gave you earlier. That is because,
at time t + 1, we’re not simply trying to group nodes in the same
community according to the density of their connections. Eventually,
we’re going to end up with a partition with many edges running
between communities, which is against the traditional definition of
community. Together with the ability of SBMs to find disassortative 48
Mark Goldberg, Malik Magdon-
communities, these are yet more cracks appearing in the classical Ismail, Srinivas Nambirajan, and James
Thompson. Tracking and predicting
community detection assumption of assortative communities.
evolution of social communities. In
Smoothing is not necessarily applied to adjacent snapshots: you SocialCom, pages 780–783. IEEE, 2011
can have a longer memory looking at t − 2, t − 3, and so on48 , 49 . In 49
Matteo Morini, Patrick Flandrin,
alternative approaches, you can skip the smoothing altogether. You Eric Fleury, Tommaso Venturini, and
Pablo Jensen. Revealing evolutions in
can identify a “core-node” which is the center of the community and dynamical networks. arXiv preprint
will identify it for all snapshots. You then find communities around arXiv:1707.02114, 2017
graph partitions 431
50
Zhengzhang Chen, Kevin A Wilson,
that node50 , 51 . Ye Jin, William Hendrix, and Nagiza F
Samatova. Detecting and tracking
community dynamics in evolutionary
Other Approaches networks. In ICDMW, pages 318–327.
IEEE, 2010
Alternatives to evolutionary clustering exist. You could find an
51
Yi Wang, Bin Wu, and Xin Pei. Comm-
tracker: A core-based algorithm of
optimal partition only for the very first snapshot of your network. tracking community evolution. In
As you receive a new snapshot, rather that starting from scratch ADMA, pages 229–240. Springer, 2008b
and then smoothing, you can adapt the old communities to the new 52
K Miller and Tina Eliassi-Rad. Contin-
network, whether you do it via global optimization52 , or using a uous time group discovery in dynamic
graphs. Technical report, LLNL, 2010
specific set of rules to update the old communities53 , 54 , 55 . 53
Giulio Rossetti, Luca Pappalardo,
Another approach consists in defining a dynamic null model: a Dino Pedreschi, and Fosca Giannotti.
null version of your evolving network which has no communities56 , Tiles: an online algorithm for com-
munity discovery in dynamic social
much like a random graph. Then you look at deviations from this ex-
networks. Machine Learning, 106(8):
pected null model in the network as the potential sources of dynamic 1213–1241, 2017
communities. 54
Yizhou Sun, Jie Tang, Jiawei Han,
Manish Gupta, and Bo Zhao. Commu-
You can use SBMs in this case as well – you can use SBMs for
nity evolution detection in dynamic
any case, really. SBMs tend to be more principled than evolutionary heterogeneous information networks. In
clustering approaches, because they model temporal communities MLGraphs, pages 137–146. ACM, 2010
55
Lei Tang, Huan Liu, Jianping Zhang,
directly57 , rather than chasing communities around as your network
and Zohreh Nazeri. Community
evolves. As an example, you could use a tensor representation in evolution in dynamic multi-mode
which each slice of the tensor is a snapshot of the network. Since a networks. In SIGKDD, pages 677–685.
ACM, 2008
tensor is nothing more than a high dimensional matrix, and SBMs 56
Danielle S Bassett, Mason A Porter,
understand matrices, you can make a tensor-SBM58 . One nice thing Nicholas F Wymbs, Scott T Grafton,
about the SBM approach is that it allows to estimate from data the Jean M Carlson, and Peter J Mucha. Ro-
bust detection of dynamic community
timescale at which the community structure changes – which is structure in networks. Chaos Journal, 23
inferred by the coupling between snapshots. This is nice because then (1):013142, 2013
you don’t need to decide yourself the granularity of the temporal
57
Tiago P Peixoto. Inferring the
mesoscale structure of layered, edge-
observation. Basing your inferences on data rather than taking a valued, and time-varying networks.
guess is always a plus! Physical Review E, 92(4):042807, 2015
A final approach is not to consider the different snapshots as
58
Marc Tarrés-Deulofeu, Antonia
Godoy-Lorite, Roger Guimera, and
separate, but taking the entire structure of the network as input all at Marta Sales-Pardo. Tensorial and bi-
once, as if it were a single structure. For instance, you can split each partite block models for link prediction
in layered networks and temporal net-
node v into many meta-nodes vt1 , vt2 , ... connected to each other by works. Physical Review E, 99(3):032307,
special edges59 , 60 , 61 , 62 , 63 . This is similar to performing multi-layer 2019
community discovery, which we’ll see later. 59
Jimeng Sun, Christos Faloutsos,
Spiros Papadimitriou, and Philip S Yu.
Graphscope: parameter-free mining of
large time-evolving graphs. In SIGKDD,
31.5 Local Communities pages 687–696. ACM, 2007a
60
Laetitia Gauvin, André Panisson, and
In some cases, you are not interested in grouping every node into Ciro Cattuto. Detecting the community
a community. I’m not just referring to allowing nodes to be part of structure and activity patterns of
temporal networks: a non-negative
no communities – a feature included in many algorithms, regardless tensor factorization approach. PloS one, 9
of their guiding principle. Sometimes, you want to find local com- (1):e86028, 2014
munities: you’re interested in knowing the communities around a
specific (set of) node(s), regardless of the rest of the network. This
432 the atlas for the aspiring network scientist
makes sense if the network is very large and some nodes are just too 61
Leto Peel and Aaron Clauset. De-
far to ever influence the results on your specific objectives. Or you tecting change points in the large-scale
structure of evolving networks. In
cannot analyze it fully because it would take too much memory. Or it Twenty-Ninth AAAI Conference on Artifi-
might take too much time to access the entire network, imposing you cial Intelligence, 2015
to sample it (see Chapter 25).
62
Amir Ghasemian, Pan Zhang, Aaron
Clauset, Cristopher Moore, and Leto
This is usually done by exploring the graph one node at a time, Peel. Detectability thresholds and
putting nodes into different bins according to their exploration status optimal algorithms for community
structure in dynamic networks. Physical
– and their community affiliation. For instance, you start from a seed Review X, 6(3):031005, 2016
node v0 , which by definition is part of your local community C . All of 63
Tiphaine Viard, Matthieu Latapy, and
its neighbors are part of the unexplored node set U . Clémence Magnien. Computing maxi-
mal cliques in link streams. Theoretical
Computer Science, 609:245–252, 2016
collapsed-sbm
birch
ward
agglomerative kmeans
dbscan mixnet bnmtf
crossass pmm
affinity ocg
meanshift kerlin
spectral conclude
kclique code-dense
hrg
vbmod
demon
moses bridgeboundm m s b
fluid
ganxis cme-td oslom leadeig
labelperc
extr
tiles ganet
spinglass
gce conga
agm metis
ilcd moganet
infocentr
edgebetween tabu louvain
bigclam
edgeclust copra
savi cme-bu
svinet lwplocal
fuzzyclust netcarto
slpa
hlc ganet+
linecomms
mlrmcl
rmcl
infomap graclus m s g cliquemod
olc infomap-overlap mcl fastgreedy
bagrowlocal
peacock walktrap
graclus2stage vm
clausetlocal
31.8 Exercises
2. Find the local communities in the same network using the same
algorithm, by only looking at the 2-step neighborhood of nodes 1,
21, and 181.
How do you know if you found a good partition of nodes into com-
munities? Or, if you have two competing partitions, how do you
decide which is best? In this chapter, I present to you a battery of
functions you can use to solve this problem. Why a “battery” of func-
tions? Doesn’t “best” imply that there is some sort of ideal partition?
Not really. What’s “best” depends on what you want to use your
communities for. Different functions privilege different applications.
So we need a quality function per application and you need to care-
fully choose your evaluation strategy to match the problem definition
you’re trying to solve with your communities.
Think about “evaluating your communities” more as a data ex-
ploration task than a quest to find the ultimate truth. Since there is
no one True partition – and not even one True definition of commu-
nity as I suggested in the previous chapter –, there also cannot be
one True quality function. You have, instead, multiple ways to see
different kinds of communities, some of which might be more or less
useful given the network you have and the task you want to perform.
In the first two sections, I start by focusing on functions that only
take into account the topological information of your network. In this
case, the only thing that matters are the nodes and edges – at most
we can consider the direction and/or the weight of an edge.
In the latter two sections I move to a different perspective. First,
we consider the network as essentially dynamic and we use commu-
nities as clues as to which links will appear next, under the assump-
tion the communities tend to densify: it is much more likely that a
new link will appear between nodes in the same community. Finally,
we look at metadata that could be attached to nodes, which might be
providing some sort of “ground truth” for the actual communities in
which nodes are grouped into in the real world.
438 the atlas for the aspiring network scientist
32.1 Modularity
As a Quality Measure
When it comes to functions evaluating the goodness of a commu-
nity partition using exclusively topological information, there is one 1
Mark EJ Newman. Modularity and
undisputed queen: modularity1 . You shouldn’t be fooled by its popu- community structure in networks.
larity: modularity has severe known issues that limits its usefulness. Proceedings of the national academy of
sciences, 103(23):8577–8582, 2006b
We’ll get to those in the second half of this section.
Modularity is a measure following closely the classical definition
of community discovery. It is all about the internal density of your
communities. However, you cannot simply maximize internal density,
as the partition with the highest possible density is a degenerate one,
where you simply have one community per edge – two connected
nodes have, by definition, a density of one.
do. Picking those nodes from Figure 32.1(b) results in finding only 17
edges among them.
The domain of the modularity function is thus defined between
+1 and −0.5, as Figure 32.2 shows. A positive modularity happens
when our partition finds nodes whose number of edges exceeds null
expectation. When expectation exactly matches the number of edges
in our community partition, modularity is zero. You can achieve
negative modularity by trying to group nodes together that connect
to each other less than chance. This can be a reasonable scenario: for
instance, if you have disassortative communities (see Section 26.2).
Note that, in the leftmost graph in Figure 32.2, nodes of the same
color do not connect with each other.
1 kv ku
M=
2| E | ∑ Auv −
2| E |
δ ( c v , c u ),
u,v∈V
which translates into: for every pair of nodes in the same commu-
nity subtract from their observed relation the expected number of
relations given the degree of the two nodes and the total number of
edges in the network, then normalize so that the maximum is 1.
Modularity and Stochastic Blockmodels are related. Optimizing
the community partition following modularity is proven to be equiv-
2
Mark EJ Newman. Equivalence
between modularity optimization and
alent to a special restricted version of SBM2 . Specifically, you need maximum likelihood methods for
to use the degree-correlated SBM – since it fixes the degree distribu- community detection. Physical Review E,
94(5):052315, 2016a
tion just like the configuration model does (which is the null model
on which modularity is defined). Then, you must fix pin and pout
– the probabilities of connecting to nodes inside and outside their
community – to be the same for all nodes.
In general, you can use both to evaluate the quality of your par-
tition, but there are subtle differences. SBM is by nature generative:
it gives you connection probabilities between your nodes. Modu-
larity doesn’t. On the other hand, modularity has this inherent test
against a null graph which you don’t really have in SBMs. In fact,
you can easily extend modularity in such way that you can talk about
3
Brian Karrer, Elizaveta Levina, and
a statistically significant community partition, one that is sufficiently Mark EJ Newman. Robustness of
different from chance3 . community structure in networks.
Physical review E, 77(4):046119, 2008
As a Maximization Target
As I mentioned earlier, modularity can be used in two ways. So far,
we’ve seen the use case of evaluating your partitions. You start from
a graph, you try two algorithms (or the same algorithm twice) and
you get two partitions. The one with the highest modularity is the
preferred one – see Figure 32.4(a).
kout in
1 v ku
M=
2| E | ∑ Auv −
| E|
δ ( c v , c u ).
u,v∈V
1 wvout win
M= ∑ wuv −
u
δ ( c v , c u ).
2| E | u,v∈V ∑ wuv
u,v∈V
Known Issues
But the issues raised so far are only child’s play. Let’s take a look
at the real problematic stuff when it comes to modularity. There are
three main grievances with modularity. The first is that random
fluctuations in the graph structure and/or in your partition can 21
Roger Guimera, Marta Sales-Pardo,
and Luís A Nunes Amaral. Modularity
make your modularity increase21 . However, I already mentioned that
from fluctuations in random graphs and
modularity can be extended to take care of statistical significance. complex networks. Physical Review E, 70
A harder beast to tame is the infamous resolution limit of modu- (2):025101, 2004
M = 0.535
Intuitive → M = 0.902
Figure 32.6: A ring of cliques,
showing another side of the
resolution limit problem of
Best → M = 0.904 modularity.
(a)
(b)
Internal density
The other side of the conductance coin is the internal density mea-
sure. This is exactly what you’d think it is: how many edges are
inside the community over the total possible number of edges the 38
Filippo Radicchi, Claudio Castellano,
community could host38 . Borrowing EC from the previous section: Federico Cecconi, Vittorio Loreto,
and Domenico Parisi. Defining and
| EC | identifying communities in networks.
f (C ) = . Proceedings of the National Academy of
|C |(|C | − 1)/2 Sciences, 101(9):2658–2663, 2004
So you can see that, in this case, both communities in Figure 32.7
have an internal density of 1, since they’re cliques. Thus, internal
density is unable to distinguish between them, which we would like
since community Figure 32.7(b) is clearly “weaker”, given its high
number of external connections.
Cut
Originally, we define the cut ratio as the fraction of all possible edges
leaving the community. The worst case scenario is when every node
in C has a link to a node not in C. There are |C | nodes in C and
(|V | − |C |) nodes outside C, so there can be |C |(|V | − |C |) such links.
Thus:
| EC,B |
f (C ) = .
|C |(|V | − |C |)
This is usually what gets minimized when solving the mincut
problem (Section 8.4). Again, this is a measure easy to game. That is
why we often modify it to be a “normalized” mincut:
| EB,C | | EB,C |
f (C ) = + .
2| EC | + | EB,C | 2(| E| − | EC |) + | EB,C |
The most attentive readers already noticed that the first term in
this equation is conductance. The second term is also a conductance
of sorts. If the first term is the conductance from the community to
the rest of the network, the second term is the conductance from
the rest of the network to the community. The two are not the same,
because the number of edges in C is | EC |, while the number of edges
outside C is | E| − | EC |.
|(u, v) : v 6∈ C |
f (C ) = max .
u∈C ku
The idea here is that, in a good community partition, there
shouldn’t be any node with a significant number of edges point-
ing outside the community. We can tolerate if a node has a large
number of edges pointing out, only if the node is a gigantic hub with
a humongous degree k u .
Requiring that there is absolutely no node with a large out degree
fraction might be a bit too much. So we also have a relaxed Average-
ODF:
1 |(u, v) : v 6∈ C |
f (C ) =
|C | ∑ ku
.
u∈C
In this case, we’re ok if, on average, nodes tend not to connect
relatively much to neighbors outside the cluster. If there is one
node doing so, the presence of many other nodes without external
connections will overwhelm it.
Finally, Flake et al. in their paper propose a further variant of the
same idea:
1
f (C ) = |{u : u ∈ C, |(u, v) : v 6∈ C | < k u /2}|.
|C |
For each node u in C, we count the number of edges pointing
outside the cluster. If it’s more than half of its edges, we mark the
node as “bad”, because it connects more outside the community than
inside. A node shouldn’t do that! The measure tells you the share of
bad nodes in C, which is something you want to minimize.
Figure 32.9 shows an example community, which we can use to
understand the difference between the various ODF variants. In the
Maximum-ODF, we’re looking for the node with the relative highest
out degree. That is node 1 as its degree is just three, and two of
those edges point outside the community. Thus, the Maximum-ODF
is f (C ) = 2/3. Both nodes 2 and 3 have a higher out-community
degree, but they also have a higher degree and thus they don’t count
at all for the community quality. You can see how Maximum-ODF is
a blunt tool which disregards lots of information.
450 the atlas for the aspiring network scientist
3
2
If you have a temporal network, you gain a new way to test the
quality of your communities. After all, communities are dense areas
in the network, thus they tell you something about where you expect
to find new links. In a strong assortative community partition, there
are more links between nodes in the same community than between
40
Or so the classical definition of
nodes in different communities. Otherwise, you communities would
community says. I already started
be weak – or there won’t be communities at all40 . tearing it apart, and I’ll continue doing
Thus you can use your communities to have a prior about where so, but in this specific test you base
your assumption on this classical
the new links will appear in your dynamic network. This sounds definition. If you have a different
familiar because it is: it is literally the definition of the link prediction definition of community, don’t use this
test.
problem (Part VI). In this approach of community evaluation, you
use the community partition as your input. You use it to estimate the
likelihood of connection between any pair of nodes in the network,
and then you can design the experiment (Chapter 22) and use any
link prediction quality measure as your criterion to decide which
community partition is better. The higher your AUC, the better
looking your ROC curve, the better your partition is.
The classical way to create a score(u, v) is having a simple binary
classifier: 1 if u and v are in the same community, 0 otherwise. This
is a bit clunky, so you usually want to add a bit of information: how
well embedded are the nodes in the network? This also works in the
community evaluation 451
Your network might not be temporal, but you could have additional
information about the nodes, besides to which other nodes they
connect (Section 4.5). In this context, node attributes are usually
referred to as “node metadata”. There is a widespread assumption
in community discovery: if you have good node metadata, some of
them have information about the true communities of the network.
Nodes with similar values, following the homophily assumption
(Chapter 26), will tend to connect to each other. Therefore there
should be some sort of agreement between the community partition 41
Jaewon Yang and Jure Leskovec.
of the network and the node metadata41 . Defining and evaluating network
communities based on ground-truth.
For instance, a classical paper42 analyzed a network whose nodes Knowledge and Information Systems, 42(1):
were cellphones, connected together if they made a significant num- 181–213, 2015
ber of calls to each other. The network showed three well-separated
42
Vincent D Blondel, Jean-Loup Guil-
laume, Renaud Lambiotte, and Etienne
communities. Figure 32.10 shows an extreme simplification of that Lefebvre. Fast unfolding of communities
(very large) graph. in large networks. Journal of statistical
mechanics: theory and experiment, 2008
Why was that the case? Why were there gigantic communities? It
(10):P10008, 2008
all becomes clear when I tell you that the country they studied was
452 the atlas for the aspiring network scientist
a NMI ~ 0.09
Figure 32.12: Two random vec-
b AMI ~ -0.22 tors with a positive NMI even if
generated completely indepen-
Consider the vectors in Figure 32.12. I generated them by extract- dently from one another.
ing ten random elements, with three possible values. This is done
uniformly at random and with independent draws – pinky promise!
Yet, if you calculate their NMI values, you’re going to obtain around
0.09: a non-zero mutual information from vectors that literally have
nothing to do with each other. This is not good.
That is why researchers developed a new normalization for mu-
45
Marina Meilă. Comparing cluster-
tual information: Adjusted mutual information (AMI)45 , 46 . In this ings—an information based distance.
case, we subtract from mutual information the amount of bits we Journal of multivariate analysis, 98(5):
would expect to obtain about a vector by pure chance. In this, AMI 873–895, 2007
46
Nguyen Xuan Vinh, Julien Epps, and
is similar to modularity: you’re comparing the observed value with James Bailey. Information theoretic
the one you’d get from some sort of null model. AMI is defined to be measures for clusterings comparison:
equal to zero when you get nothing more than you’d expect by just Variants, properties, normalization and
correction for chance. J. Mach. Learn. Res,
tossing coins. At this point, any positive AMI value starts getting in- 11(Oct):2837–2854, 2010
teresting. AMI can be negative, and it is for the two vectors in Figure
454 the atlas for the aspiring network scientist
32.5 Summary
2. You can also use modularity for something more than evaluating
the communities you found: it can be an optimization target. Your
algorithm will operate on your communities until it cannot find
any additional move that would increase modularity.
32.6 Exercises
Merging
In the merging approach, you start from a condition where all your
nodes are isolated in their own community and you create a criterion
to merge communities. This is a bottom-up approach. It is similar
to the meta-algorithm from earlier, but it’s not really the same. Let’s
take a look at how it works, highlighting where the differences with
the meta-algorithm are.
The template I’m using to describe this approach is the Louvain 5
Vincent D Blondel, Jean-Loup Guil-
algorithm5 . This is one of the many heuristics used to recursively laume, Renaud Lambiotte, and Etienne
Lefebvre. Fast unfolding of communities
merge communities with the aim of maximizing modularity6 , 7 ,
in large networks. Journal of statistical
which happens to be among the fastest and most popular. mechanics: theory and experiment, 2008
The Louvain algorithm starts with each node in its own commu- (10):P10008, 2008
6
Marta Sales-Pardo, Roger Guimera,
nity. It calculates, for each edge, the modularity gain one would get
André A Moreira, and Luís A Nunes
if they were to merge the two nodes in the same community. Then Amaral. Extracting the hierarchical
it merges all edges with a positive modularity gain. Now we have a organization of complex systems.
Proceedings of the National Academy of
different network for which the expensive modularity gains need to Sciences, 104(39):15224–15229, 2007
be recomputed. However, this network is smaller, because of all the 7
Tiago P Peixoto. Hierarchical block
edge merges. You repeat the process until you have all nodes in the structures and high-resolution model
selection in large networks. Physical
same community. Figure 33.3 shows an example of this process. Review X, 4(1):011047, 2014c
High Gain
Figure 33.3: An example of the
first step of the Louvain algo-
rithm. All in-clique edges (like
Low Gain
the representative I highlight
in blue) are merged, while all
out-clique edges (like the repre-
sentative I point to with a gray
arrow) are ignored.
Iterations
Splitting 9
Michelle Girvan and Mark EJ New-
man. Community structure in social
In the splitting approach, you do the opposite of what I described and biological networks. Proceedings of
so far. You start with all nodes in the same community and you use the national academy of sciences, 99(12):
a criterion to split it up in different communities. For instance by 7821–7826, 2002
10
Mark EJ Newman and Michelle
identifying edges to cut. This is a top-down approach. Girvan. Finding and evaluating
Historically speaking, the first algorithm using this approach used community structure in networks.
edge betweenness as its criterion to split communities9 , 10 . That is Physical review E, 69(2):026113, 2004
11
Filippo Radicchi, Claudio Castellano,
not to say there aren’t valid alternatives as your splitting criterion, Federico Cecconi, Vittorio Loreto,
including – but not limiting to – edge clustering11 and information and Domenico Parisi. Defining and
centrality12 . However, given its historical prominence, I’m going to identifying communities in networks.
Proceedings of the National Academy of
allow the edge betweenness Girvan-Newman algorithm to have its Sciences, 101(9):2658–2663, 2004
place under the limelight. 12
Santo Fortunato, Vito Latora, and
The first step of the algorithm is to calculate the edge betweenness Massimo Marchiori. Method to
find community structures based on
of each edge in the network, that is the normalized number of short- information centrality. Physical review E,
est paths passing through it (Section 11.2). The assumption is that 70(5):056104, 2004
hierarchical community discovery 461
The second step of the algorithm is to cut the edge with the high-
est edge betweenness. The final aim is to break the network down
into multiple components. Each component of the network is a com-
munity.
Unfortunately, after each edge deletion you have to recalculate
the betwennesses. Every time you alter the topology of the network
you change the distribution of its shortest paths. This makes edge
betweenness extremely computationally heavy. Calculating the edge
betweenness for all edges takes an operation per node and per edge
(O(|V || E|)) and you have to repeat this for every edge you delete,
resulting in a crazy complexity of O(|V || E|2 ). You cannot apply this
naive algorithm to anything but trivially small networks.
You can now see the parallels with the Louvain method I de-
scribed earlier. The difference is that you are exploring the dendo-
gram of communities from the top down, rather than bottom up.
Each iteration brings you further down in the hierarchy. At the very
top you start with a network with a single connected component. As
you delete edges, you find different connected components. As you
continue, you end up with more and more. At the last iteration, each
node is now isolated.
Differently from the Louvain algorithm, in the Girvan-Newman
method you do not calculate modularity gains as you explore the
dendogram. Thus, the algorithm will normally perform all the
possible splits and returns you the full structure, rather than the
462 the atlas for the aspiring network scientist
cut that maximizes modularity. Thus you will have to calculate the
modularity of each split yourself, something similar to what you see
in Figure 33.6.
Modularity
Figure 33.6: The dendogram
building from the top down
typical of a “splitting” approach
in hierarchical community dis-
covery. The left panel shows
the modularity values of each
possible cut.
Iterations
30 29
Figure 33.7: A graph with hi-
31 32
erarchical communities (node
27
color according to the commu-
33
26 nity partition at one level of the
36
28 34
25
hierarchy).
35
10 15
11
13
9 12 14 16
8
3 18 22
1 6 17 23
4 7
19 21
2 5 20 24
0.1
0.15
1 1 1 1 1 1 1 1 1
1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1
33 35 34 36 31 30 32 29 26 25 27 28 11 9 10 12 1 2 3 4 7 8 5 6 17 20 18 19 21 22 24 23 14 15 16 13
Modularity Modularity
Density Density
Iterations
Iterations
(a) (b)
Modularity Modularity
Density Density
Iterations Iterations
(c) (d)
Figure 33.9: The contrast be-
tween modularity and density
at different cuts in the hierarchi-
33.4 Summary
cal community organization: (a)
1. You can find communities at different scales in a network. Mean- every node in the same commu-
ing that there are communities of nodes, communities of com- nity; (b) sub-optimal high-level
munities, communities of communities of communities, and so partition; (c) optimal low-level
on. The process to find such structures is hierarchical community partition; (d) maximal density
discovery. but low modularity partition.
466 the atlas for the aspiring network scientist
33.5 Exercises
4. Using the algorithm you made for exercise 3, answer these ques-
tions: What is the latest step for which you have the average
internal community edge density equal to 1? What is the modu-
larity at that step? What is the highest modularity you can obtain?
What is the average internal community edge density at that step?
34
Overlapping Coverage
This seems to imply that communities are a clear cut case. Nodes
have a majority of connections to other nodes in their community.
However real world networks do not have to conform to this expecta-
tion, and in fact often they don’t. There are numerous cases in which
nodes belong to multiple communities: to which community does the
center person in Figure 34.1 belong? The red or the blue one?
The classical community definition forces us to make a choice. Re-
gardless of the choice we make – red or blue – it’d not be a satisfying
solution. The more reasonable answer is “she belongs to both”. For
instance, a person can very well be part of one community because
it is composed by the people they went to school with. And she can
be part of a work community too, of people she works with. Some of
these people could be the same, but usually they are not.
468 the atlas for the aspiring network scientist
The problem is that none of the methods seen so far allow for such
a consideration. For instance, the basic stochastic blockmodels only
allows you to plant a node in a community, not multiple. Modularity
also has issues, because of the Kronecker delta: since this is going
to be 1 for multiple communities for a node, there will be double-
counting and the formula breaks down.
This is where the concept of overlapping community discovery
was born. We need to explicitly allow for overlapping communities:
communities that can share nodes. There are many ways to do this, 1
Jierui Xie, Stephen Kelley, and
which have been reviewed in several articles1 , 2 dedicated especially Boleslaw K Szymanski. Overlap-
to this sub problem of community detection (itself a sub problem of ping community detection in networks:
The state-of-the-art and comparative
network analysis: it’s communities all the way down).
study. Acm computing surveys (csur), 45
Here we explore a few of the most popular approaches. (4):43, 2013
2
Alessia Amelio and Clara Pizzuti.
Overlapping community discovery
34.1 Evaluating Overlapping Communities methods: A survey. In Social Networks:
Analysis and Case Studies, pages 105–125.
Springer, 2014
Before we delve deep into overlapping community discovery, let’s
amend Chapter 32 to this new scenario. We can have a few options
when we try to evaluate how well we divided the network into
overlapping communities.
Normalized mutual information expects you to put nodes into
a single category. However, there are ways to make it accept an 3
Andrea Lancichinetti, Santo Fortunato,
overlapping coverage3 , 4 . The obstacle is that NMI wants to compare and János Kertész. Detecting the over-
the vector of metadata with the vector containing the community lapping and hierarchical community
structure in complex networks. New
partition. The vector can only have one value per node but, in an Journal of Physics, 11(3):033015, 2009
overlapping coverage, it can have multiple values. Thus we don’t 4
Aaron F McDaid, Derek Greene, and
compare the vectors directly. We compare two bipartite matrices. Neil Hurley. Normalized mutual
information to evaluate overlapping
Suppose you found C communities, and you have A node at- community finding algorithms. arXiv
tributes. You can describe the overlapping coverage in communities preprint arXiv:1110.2515, 2011
with a |V | × C binary matrix, whose u, C entry is equal to 1 if node
u is part of community C. The node attribute matrix is similarly de-
fined. Figure 34.2 shows an example of this procedure. Now you can
calculate the mutual information between the two matrices by pairing
the columns such that we assign to each column on one side the ones
on the other side that is the most similar to it.
We can normalize this mutual information in different ways. In
fact, the papers I cited earlier propose six alternatives, providing
different motivations for each of those. These overlapping NMIs
share with their original counterpart the issue of non-zero values for
independent vectors – although they try to mitigate the issue with
different strategies. 5
Alexander J Gates, Ian B Wood,
William P Hetrick, and Yong-Yeol Ahn.
Some researchers have pointed out a few biases in the overlapping
Element-centric clustering comparison
extensions of NMI and similar measures5 . They propose a unified unifies overlaps and hierarchy. Scientific
framework that can evaluate disjoint, overlapping, and even hierar- reports, 9(1):8574, 2019
overlapping coverage 469
2
Figure 34.2: (a) A network with
1
three overlapping communities,
4
encoded by the node’s color. (b)
7
6 Transforming the overlapping
3 5
coverage into a binary affilia-
8 tion matrix, which we can use
10 as input for the overlapping
9 version of NMI
13
11 12
(b)
(a)
Overlapping Communities:
Figure 34.3: Comparing overlap-
u: ping and fuzzy clustering for
node u: the size of the square is
u Fuzzy Communities:
proportional to u’s “belonging”
u: coefficient, in share of number
(60%) (40%) of u’s edges connected to the
community.
01 00
100 01 Figure 34.4: The encoding of
two random walks in the over-
00
lapping version of Infomap.
Note that in neither the red nor
11 the blue path we’re crossing
community boundaries, so
101 100
we don’t use the community
11 101 crossing code.
1100100101 010011100
step (from Figure 34.5(a) to 34.5(b)), the two fully red nodes stay
fully red, because they both receive 0.5 of the red label, 0.33 of the
blue and 0.16 of the purple label. The only label clearing the 0.5
threshold is the red one, thus they become fully red. The half red half
purple node becomes red because that’s the only label around it. The
central node is also red, receiving 0.5 red, 0.25 blue and 0.25 orange.
From Figure 34.5(b) to 34.5(c) though, the central node will correctly
split between red and blue, because they both contribute half of its
neighborhood.
Clique Percolation
Clique percolation starts from the observation that communities
should be dense. What is the densest possible subgraph? The clique.
In a clique, all nodes are connected to all other nodes. So the problem
of community discovery more or less reduces to the problem of
finding all cliques in the network. However, this is a bit too strict:
there are subgraphs in the network that, while being very dense and
close to being a clique, are not fully connected. It would be a pity to
split them into many small substructures.
Thus researchers developed the more sophisticated k-clique perco- 18
Imre Derényi, Gergely Palla, and
Tamás Vicsek. Clique percolation in
lation algorithm18 . Clique percolation says that communities must be
random networks. Physical review letters,
cliques of at least k nodes, with k being a parameter you can freely 94(16):160202, 2005
set. In the first step, the algorithm finds all cliques of size k, whether
they are maximal or not. Then, it attempts to merge two communities
in the same community if the two communities share at least a k − 1
clique.
For instance, consider the example in Figure 34.6, setting the
overlapping coverage 473
(c) (d)
parameter k = 5. The blue and green 5-cliques only share two nodes,
so it cannot be a 4-clique. But the green and purple do share a 4-
clique, so they are merged (top row). And there is another purple
5-clique that can now be merged with the green community (bottom
row). 19
Tim S Evans. Clique graphs and
This is generally implemented via the creation of a clique graph19 . overlapping communities. Journal
The nodes of a clique graph are the cliques in the original graph. of Statistical Mechanics: Theory and
Experiment, 2010(12):P12037, 2010
We connect two cliques if they share nodes. For instance, if we only
connect cliques sharing k − 1 nodes, then we can efficiently find
all communities by finding all connected components in the clique
graph.
This algorithm works well in practice. It has been used to study 20
Marta C González, Hans J Herrmann,
overlapping friendship patterns in school systems20 – due to class- J Kertész, and Tamás Vicsek. Commu-
room being quasi-cliques: pupils have rare but significant friendships nity structure and ethnic preferences in
school friendship networks. Physica A,
across classes –, and in metabolic networks21 . However, it has a
379(1):307–316, 2007
couple of downsides. 21
Shihua Zhang, Xuemei Ning, and
First, finding all cliques in a network is computationally expensive. Xiang-Sun Zhang. Identification of
functional modules in a ppi network
One could fix this problem by setting k to be relatively high. If we by clique percolation clustering. Com-
set k = 5 we know that nodes with degree three or less cannot be putational biology and chemistry, 30(6):
in any community, because they need at least four edges to be part 445–451, 2006
Node Splitting
Another approach is to simply recognize that a node is part of multi-
ple communities if it has different identities. This is extremely similar
to the approach of overlapping Infomap. In that case we represented
the two identities of the node by giving it two different codes: one
per community to which it belongs. Here we literally split it in two.
We modify the structure of the network in such a way that, when we
are done, by performing a normal non-overlapping community dis-
covery we recover the overlapping clusters. In the resulting structure
we have multiple nodes all referring to the same original one.
If we want to split nodes, we need to answer two questions: which
nodes do we split and how. First we identify the nodes most likely
to be in between communities. If you remember the definition of
24
Steve Gregory. An algorithm to find
betweenness, you’ll recollect that nodes between communities are the overlapping community structure in
gatekeepers of all shortest paths from one community to the other. So networks. In European Conference on
they are the best candidates to split. There are many ways to perform Principles of Data Mining and Knowledge
Discovery, pages 91–102. Springer, 2007
the split, but I’ll focus on the one that involves calculating a special 25
Steve Gregory. Finding overlapping
betweenness: pair betweenness24 , 25 . Pair betweenness is a measure communities using disjoint community
for a pair of edges: the number of shortest paths that use both of detection algorithms. In Complex
networks, pages 47–61. Springer, 2009
them.
For instance, consider the graph in Figure 34.7(a). The most central
node is node 1. To try and split it, we build its split graph. Meaning
that we remove node 1 and we connect all nodes that were connected
by 1. Each edge has a weight: the number of shortest paths in the
original graph that passed through node 1. In this case, there are
two shortest paths using the (4, 1) and (1, 3) edges: the one going
from node 4 to node 3 and the one going from node 3 to node 4. We
can represent the pair betweenness of all neighbors of node 1 with a
overlapping coverage 475
4-5 2 2 5 2
4 0 4-5 8 2-3 1 1
3 4 3
(b)
(a) (c)
Figure 34.8: (a-b) Merging the
weighted clique (Figure 34.7(b)). nodes connected by the weakest
To find the split we use a simple algorithm: we identify the edges split betweenness edges. (c) The
with the lowest pair betweenness and we merge the nodes connected resulting split in the original
by those edges (Figures 34.8(a-b)). At each merge, we sum up the graph.
pair betweennesses of all edges that got merged together by the merg-
ing of the node. Once we have one remaining edge, the resulting
split is the best one. The reason is that edges with low pair between-
ness are likely to be in the same community. Once you identify the
split (Figure 34.8(c)), it is easy to find disjoint communities and then
merge them into overlapping.
regular SBM. From this moment on, you attempt to find the set 28
Eric P Xing, Wenjie Fu, Le Song,
et al. A state-space mixed membership
of community affiliation vectors and the community-community blockmodel for dynamic network
probability matrix that are most likely to reproduce your observed tomography. The Annals of Applied
Statistics, 4(2):535–566, 2010
data, exactly as you do in SBM. 29
Qirong Ho, Le Song, and Eric Xing.
Just like we saw in Section 31.4, we can have dynamic MMSB, Evolving cluster mixed-membership
adding time to the mix27 , 28 , 29 , 30 : the community affiliation vectors blockmodel for time-evolving networks.
In Proceedings of the Fourteenth Interna-
and the community-community matrix can change over time. There tional Conference on Artificial Intelligence
is also a hierarchical (Chapter 33) variant of MMSB, allowing a nested and Statistics, pages 342–350, 2011
community structure31 . 30
Kevin S Xu and Alfred O Hero.
Dynamic stochastic blockmodels:
Statistical models for time-evolving
Community Affiliation Graph networks. In International conference
on social computing, behavioral-cultural
Affiliations graphs have been often used to describe the overlapping modeling, and prediction, pages 201–210.
Springer, 2013
community structure of real world networks32 . In a community 31
Tracy M Sweet, Andrew C Thomas,
affiliation graph you assume that you can describe your observed and Brian W Junker. Hierarchical mixed
network with a latent bipartite network. In this bipartite network, membership stochastic blockmodels for
multiple networks and experimental
the nodes of one type are the nodes of your observed network. The interventions. Handbook on mixed
other type, the latent nodes, represent your communities. Nodes are membership models and their applications,
pages 463–488, 2014
connected to the communities they belong to. This is the community 32
Jae Dong Noh, Hyeong-Chai Jeong,
affiliation graph, because it describes the affiliations to communities Yong-Yeol Ahn, and Hawoong Jeong.
of your nodes. Growing network model for community
with group structure. Physical Review E,
71(3):036131, 2005
2
1 Figure 34.9: (a) A graph with
4
7 overlapping communities indi-
6 1 2 3 4 5 6 7 8 9 10 11 12 13
3
cated by the colored outlines.
5
8
(b) Its corresponding com-
10
9 munity affiliation graph. The
13
community latent nodes are
11 12 (b)
triangular and their color cor-
responds to the color used in
(a) (a).
Figure 34.9 shows a representation of a community affiliation
graph. Of course, you can build such a graph easily once you already 33
Jaewon Yang and Jure Leskovec.
know to which communities the nodes belong. The hard part is Overlapping community detection at
scale: a nonnegative matrix factorization
finding out the best representation. There are a few ways to do so,
approach. In Proceedings of the sixth ACM
usually relying on the expectation maximization algorithm that is international conference on Web search and
also at the basis of the MMSB. One such approach is BigClam33 , data mining, pages 587–596. ACM, 2013
overlapping coverage 477
Line Graphs
To cluster the edges rather than the nodes we can transform the 34
TS Evans and Renaud Lambiotte. Line
network into its corresponding line graph34 . In a line graph, as we graphs, link partitions, and overlapping
communities. Physical Review E, 80(1):
saw in Section 3.1, the edges become nodes and they are connected
016105, 2009
if they’re incident on the same node. A way to do so is to generate a
weighted line graph.
5 2 4-5 1-4 1-5 1-2 1-3 2-3 Figure 34.10: (a) A simple
1 graph. (b) A bipartite version of
(a) connecting each node to its
4 5 1 2 3
4 3 edges.
(b)
(a)
To create a line graph you first transform the network into bi-
partite connecting the nodes to the edges they are connected to, as
Figure 34.10 shows. Then you project this network over the edges.
The most important thing to define is how to weight the edges in
the line graph. Different weight profiles will steer the community
discovery on the line graph in different directions.
You could use any of the weighting schemes I discussed in Chap-
ter 23, but the researchers proposing this method also have their
suggestions. The reason you might need a special projection is be-
cause you want nodes that are part of an overlap to give their edges
lower weights, because their connections are spread out in different
communities.
At that point, a disjoint community discovery will downplay
478 the atlas for the aspiring network scientist
| Nu ∩ Nv |
S(u,k),(v,k) = .
| Nu ∪ Nv |
1
3 Figure 34.12: A graph and its
8 best link communities. The
2 color of the edge represents its
4 7 community. Nodes are part of
9 all communities of their links.
6 For instance, node 4 belongs to
5 three communities: red, blue,
and purple.
The edges with the highest S value are merged in the same com-
munity. For instance, in Figure 34.12, edges (1, 2) and (1, 3) have a
high S value: the neighborhoods of nodes 2 and 3 are identical, thus
S(1,2),(1,3) = 1. On the other hand, edges (4, 7) and (7, 8) only have
one node in the numerator, thus: S(4,7),(7,8) = 1/6.
Then, the merging happens recursively for lower and lower S
values, building a full dendrogram, as we saw in Chapter 33 for
hierarchical community discovery. We then need a criterion to cut
the dendrogram. We cannot use modularity, because these are link
communities, not node communities.
overlapping coverage 479
| Ec | − (|Vc | − 1)
Dc = ,
|Vc |(|Vc | − 1)/2 − (|Vc | − 1)
which is the number of links in c, normalized by the maximum
number of links possible between those nodes (|Vc |(|Vc | − 1)/2), and
its minimum |Vc | − 1, since we assume that the subgraph induced by
c is connected. Note that, if |Vc | = 2, we simply take Dc = 0. All Dc
scores for all cs in your link partition are aggregated to find the final
partition density, which is the average of Dc weighted by how many
1
links are in c: D = ∑ | Ec | Dc .
| E| c
Ego Networks
Assuming that links exists for one primary reason works usually
well, but it is a problematic assumption. Let’s look back at the case
of work and school communities. What would happen if you were to
end up working in the same company and play in the same team of a
former schoolmate? Is it still fair to say that the link between the two
of you exists for only one predominant reason?
Modeling truly overlapping communities can get rid of this prob-
lem. There are many ways to do it, but we’ll focus on one that is
easy to understand. The starting observation is that networks have
large and messy overlaps. However, just like in the assumption of
clustering links, here we realize that the neighbors of a node usually
are easier to analyze. It is easy for a node to look at a neighbor and
say: “I know this other node for this reason (or set of reasons)”. 37
Michele Coscia, Giulio Rossetti, Fosca
The procedure37 works as follows, and I use Figure 34.13 to guide Giannotti, and Dino Pedreschi. Demon:
you. First, we extract the ego network of a node, removing the ego a local-first discovery method for
overlapping communities. In Proceedings
itself. This creates a simpler network to analyze. In the figure, I of the 18th ACM SIGKDD international
start by looking at node 1 on the top right. This is a graph with two conference on Knowledge discovery and
data mining, pages 615–623. ACM, 2012
connected components: one connecting nodes 2 and 3, the other
connecting nodes 4 and 5.
Then, we apply a disjoint community discovery algorithm to the
ego network. The ego network is easier to analyze and often has
easily distinguishable communities. In the example, these are the
blue and red outlines, which find two trivial communities. We repeat
the process for all nodes in the network, extracting all ego networks:
in the figure I show the ego networks of nodes 2 and 3, with the
communities I find in those cases, the green and purple outlines,
respectively. Note that I omit the ego networks for nodes 4 and 5, but
480 the atlas for the aspiring network scientist
5 2
Figure 34.13: The process of
5 2 community discovery via the
1 breaking down of the network
into ego networks.
4 3
4 3
Apply CD
{2,3}
1 2 {1,2} Merge similar {1,2,3}
{1,3}
3 1
{4,5}
{1,5} Merge similar {1,4,5}
Repeat for all nodes {1,4}
(b)
(a)
If there are nodes in between communities either of two things
will happen. The nodes in the overlap could connect with all nodes
in both communities but not to each other, to maintain the “external
sparsity” condition. But doing so contradicts the “internal density”
part, because the overlap nodes do no connect to each other even
though they belong to the same community. Figure 34.14 provides an
example for this scenario.
In Figure 34.14(a) we have a graph made by four 5-cliques and
four sets of four nodes overlapping between two neighboring cliques.
482 the atlas for the aspiring network scientist
The overlap nodes don’t connect to each other. Figure 34.14(b) shows
how a stochastic blockmodel would interpret such a structure. You
can clearly see that there are “holes” in the communities where the
overlap nodes should be. If the overlap nodes don’t connect to each
other, they have low connection probability, which contradicts the
fact that they are part of the same community.
(b)
(a)
If, on the other hand, we maintain the “internal density” condi-
tion, since these nodes share not one but two communities, then
they are more likely to connect to each other than nodes sharing
only one community. In doing so, we end up with the opposite prob-
lem: breaking external sparsity. The overlap, which by definition is 39
Jaewon Yang and Jure Leskovec.
between the two communities, is denser than the community itself39 ! Community-affiliation graph model
In Figure 34.15(a) we have such a scenario, with a graph similar for overlapping network community
detection. In 2012 IEEE 12th international
to the one from Figure 34.14(a). But here all the overlap nodes are conference on data mining, pages 1170–
connected to each other, and they are connected more strongly than 1175. IEEE, 2012
non-overlap nodes, given that they share more communities with
each other. The corresponding stochastic blockmodel (Figure 34.15(b))
now shows that the communities themselves look weaker than the
overlap.
This is another reason why our golden rule, the standard defi-
nition of communities in complex networks, isn’t as shiny as we
originally thought.
34.8 Summary
34.9 Exercises
3. Implement the ego network algorithm: for each node, extract its
ego minus ego network and apply the label propagation algorithm,
then merge communities with a node Jaccard coefficient higher
than 0.1 (ignoring singletons: communities of a single node). Does
this method return a better NMI than k-clique percolation for
k = 3?
35
Bipartite Community Discovery
Since it has been a looming presence across this entire book part, let’s
start again with modularity, the elephant in the room of community
discovery. Network scientists in the community detection business
love modularity. If there is a scenario in which modularity doesn’t
work, they panic and start amending it to hell, until it works again.
We’ve seen this with directed and overlapping community discovery,
and we’re seeing it again.
There are a couple of alternatives when it comes to define a modu-
larity that works for bipartite networks. If you remember the original
version of the modularity, it hinges on the fact that we want the parti-
tion to divide the network in communities that are denser than what 2
Michael J Barber. Modularity and
we would expect given a null model – the configuration model. Thus, community detection in bipartite
networks. Physical Review E, 76(6):066102,
extending modularity means to find the right formulation of a null 2007
model for bipartite networks2 , 3 . 3
Roger Guimerà, Marta Sales-Pardo,
This is not that difficult, the only thing to keep in mind is that and Luís A Nunes Amaral. Module
identification in bipartite and directed
the expected number of edges in a bipartite network is different networks. Physical Review E, 76(3):036102,
than in a regular network. So, while in the traditional modularity 2007
bipartite community discovery 485
ku kv
the configuration model connection probability was , here it is
2| E |
ku kv
instead , with the added constraints that u and v needs to be
| E|
nodes of unlike type. The sum of modularity is made only across
pairs of nodes of unlike types, otherwise we would have negative
modularity contributions from nodes that cannot be connected,
which would make the modularity estimation incorrect.
To see why this is the case, suppose that we’re checking u and v
and they are of the same type. Since they are of the same type and
we’re in a bipartite network, they cannot connect to each other, so
Auv = 0. But they are both part of the network, thus k u 6= 0 and
ku kv ku kv
k v 6= 0. Thus > 0, meaning that Auv − < 0. Negative
| E| | E|
modularity contribution.
Once you have a proper bipartite modularity you can use any of
the modularity maximization algorithms to find modules in your 4
Stephen J Beckett. Improved com-
network, or even specialized ones4 . munity detection in weighted bipartite
networks. Royal Society open science, 3(1):
140536, 2016
35.2 Via Projection
bipartite networks into its two unipartite versions and then you
analyze them at the same time with specialized techniques. This dual 6
David Melamed. Community struc-
projection approach has been applied to community discovery6 , with tures in bipartite networks: A dual-
encouraging results. projection approach. PloS one, 9(5):
e97823, 2014
Bi-Clique Percolation
The solution is to perform the community discovery directly on the
bipartite structure. Here, we use the concept of bi-clique we saw
earlier. Remember that a clique is a set of nodes in which all possible
edges are present. A bi-clique is the same thing, considering that
some edges in a bipartite network are not possible. For instance, a
5-clique in a unipartite network is a graph with five nodes and ten
edges. In a bipartite network, a 2,3-clique has two nodes of type 1,
three nodes of type 2, and all nodes of type 1 are connected to nodes
bipartite community discovery 487
Two 2,4-cliques
sharing a 0,1-clique
dating will update nodes in a random order. This might impact the
stability of the resulting partition, because the order in which nodes
are updated matters. Two subsequent runs of the algorithm could
yield very different results. Moreover, it might impact convergence
time, because there are many node orders which will still result in
label oscillation. The most sure way to prevent label oscillation is to
update first all nodes of one type and then all nodes of the other.
There are a few other ways to prevent oscillation. First, one could
10
Xin Liu and Tsuyoshi Murata. Com-
integrate label propagation with the modularity approach10 . In munity detection in large-scale bipartite
this scenario, one doesn’t run label propagation until convergence, networks. Transactions of the Japanese
but only for a few steps. Then they would refine the communities Society for Artificial Intelligence, 25(1):
16–24, 2010
by maximizing bipartite modularity. Alternatively, one could put
constraints on how we allow labels to propagate. For instance, we
could force communities to be of comparable sizes in number of
bipartite community discovery 489
11
Michael J Barber and John W Clark.
nodes or edges11 . Detecting network communities by
Finally, we can adapt stochastic blockmodels to the bipartite case, propagating labels under constraints.
Physical Review E, 80(2):026129, 2009
creating a biSBM12 . This has some similarities with the MMSB we 12
Daniel B Larremore, Aaron Clauset,
saw for overlapping community discovery in Section 34.4. First, we and Abigail Z Jacobs. Efficiently
don’t look directly at the |V1 | × |V2 | biadjacency matrix B. It is more inferring community structure in
bipartite networks. Physical Review E, 90
convenient to look at its adjacency matrix equivalent:
(1):012805, 2014
!
0 B
A=
BT 0
The zeros on the main diagonal mean that nodes of the same type
cannot connect to each other, enforcing the bipartite structure (see
Section 5.1).
Then, just like in the overlapping case, we can have a special
community-community matrix that tells us the probability of nodes
in two distinct communities to connect to each other. The special
condition here is that communities grouping nodes of the same type
will have zero probability of connecting to each other, respecting the
bipartite constraint.
Once we have these two special structures in place, one can pro-
ceed finding the most likely blockmodel that explains the observed
data, which is the one with the best community partition, with the
same strategies as in vanilla SBM. Note that this method can be triv-
ially extended to multi-partite networks, modifying the fundamental
structures accordingly.
The biSBM clusters the two modes separately, so you get a mixing
matrix that tells you how the groups in the V1 nodes interact with the
groups in the V2 nodes. In contrast, bipartite modularity and some
other approaches will produce mixed groups, which contain nodes
from both V1 and V2 . This makes biSBM a co-clustering method, like
the ones we’ll see in Section 35.4. The difference with those methods
is that they find communities discovery via neighbor similarity,
which is not the philosophy of biSBM.
There is a related method that works by means of matrix factoriza- 13
Zhong-Yuan Zhang and Yong-Yeol
tion13 . It starts by noticing that zeroes in A have different meanings. Ahn. Community detection in bipartite
The zeroes in the main diagonal block represent impossible connec- networks using weighted symmetric
binary matrix factorization. International
tions, connections that shouldn’t be penalized. The zeroes in the Journal of Modern Physics C, 26(09):
off-diagonal block instead represent edges that could exist. Thus, the 1550096, 2015
authors define a mask matrix M, with the same dimensions as A,
with zeros on the main diagonal blocks and ones in the off diagonal
blocks. By factorizing the product of M and A together with our best
guess at the community organization of A, we obtain a function we
can maximize to find the best community partition, knowing that
we’re only penalizing zeroes corresponding to connections that could
exist.
490 the atlas for the aspiring network scientist
(a) (b)
This needs not to worry us. We can redefine the clustering coeffi-
cient to make sense in a bipartite network. In a unipartite network,
the triangle is the smallest non-trivial cycle, the one that does not
backtrack using the same edge, as you can see in Figure 35.4(a). We
can also have a smallest non-trivial cycle in bipartite networks. It
involves four nodes, as Figure 35.4(b) shows. So we can say that the
local clustering coefficient of a node in a bipartite network is the
number of times such cycles appear in its neighborhood, divided by 14
Peng Zhang, Jinliang Wang, Xiaojia
the number of times they could appear given its degree14 . Li, Menghui Li, Zengru Di, and Ying
Let us assume that we want to know the local square clustering Fan. Clustering coefficient and com-
munity structure of bipartite networks.
coefficient of node z. If we say that nodes u, v, and z are involved Physica A: Statistical Mechanics and its
in suvz squares, then contribution of nodes u and v to the square Applications, 387(27):6869–6875, 2008
clustering coefficient of z is:
suvz
C4u,v (z) = ,
suvz + (k u − ηuvz ) + (k v − ηuvz )
with ηuvz = 1 + suvz . In practice, the number of possible squares
bipartite community discovery 491
(in the denominator) is the number of actual squares plus how many
additional squares you could have given u’s and v’s free edges, edges
not involved in any square. Here, u and v are the nodes of the same
type.
z
Figure 35.5: An example of
bipartite network on which we
u v can calculate the local cluster-
ing coefficient of node z.
b c a d
(a)
(b)
492 the atlas for the aspiring network scientist
1 2 3 4 5 6 7 8 9 10 11 12
9 8 7 6 5 4 3 2 1
(a) (b)
(c)
Figure 35.7: (a) A bipartite net-
The bipartite community discovery introduces another issue work. (b) Its adjacency matrix.
with the standard definition of community based on density. In (c) A 2D spatial representation
n,m-cliques, we have n + m nodes that cannot be connected to each of the circular nodes, using
other because the graph is bipartite. For instance, the community in their adjacencies to determine
Figure 35.8 has many missing links. To be precise, since n = 4 (the the position. Node color is its
triangles) and m = 7 (the circles), we have 4 × 3/2 + 7/2 = 27 missing cluster, as identified by spatial
connections that the classical definition would want. More than the clustering (dashed line).
connections actually there! Nodes of the same type cannot connect in
a bipartite graph – so they have density of zero –, but they can and
will be part of the same community. So again this criterion of internal
density is a bit flaky.
35.5 Summary
35.6 Exercises
The last chapter of community discovery, at least for this book, fo-
cuses on multilayer networks. In multilayer networks, nodes can
belong to different layers and thus they can connect for different
reasons. In multilayer networks we want to find communities that
span across layers. For example, we want to figure out communities
of friends even if your friends are spread across multiple social media
platforms.
There are a few review works you can check out to have a more 1
Jungeun Kim and Jae-Gil Lee. Com-
in-depth exploration of the topic1 , 2 . Here, I go over briefly the main munity detection in multi-layer graphs:
approaches and peculiar problems of community discovery in multi- A survey. ACM SIGMOD Record, 44(3):
37–48, 2015
layer networks. 2
Obaida Hanteer, Roberto Interdonato,
Matteo Magnani, Andrea Tagarelli, and
Luca Rossi. Community detection in
36.1 Flattening multiplex networks, 2019
one – an edge weight. This assumes that every edge type is equally
important. Then you can perform a normal mono-layer community
discovery. Figure 36.1 shows an example.
There are a few choices for your edge weights. The simplest
one could be to simply count the number of layers in which the
connection between the nodes appear. However, you might want to
take into account some interplay between the layers. For instance,
you can count the number of common neighbors that two nodes 4
Jungeun Kim, Jae-Gil Lee, and Sungsu
have and use that as the weight of the layer, under the assumption Lim. Differential flattening: A novel
that a layer where two nodes have many common neighbors should framework for community detection in
multi-layer graphs. ACM Transactions on
count for more when discoverying communities. Or you could use Intelligent Systems and Technology (TIST),
“differential flattening4 ”: flatten the multilayer graph into the single 8(2):27, 2017
496 the atlas for the aspiring network scientist
(a) (b)
each row is a node and each column is the partition assignment for 6
Lei Tang, Xufei Wang, and Huan Liu.
that node in a specific layer6 . This is then a |V | × |C | matrix. Then one Community detection via heteroge-
could perform kMeans on it, finding clusters of nodes that tend to be neous interaction analysis. Data mining
and knowledge discovery, 25(1):1–33, 2012
clustered in the same communities across layers. 7
Michele Berlingerio, Fabio Pinelli, and
A similar approach7 uses frequent pattern mining, a topic we’ll see Francesco Calabrese. Abacus: frequent
more in depth in Section 39.3. For now, suffice to say that we again pattern mining-based community dis-
covery in multidimensional networks.
perform community discovery on each layer separately. Each node
Data Mining and Knowledge Discovery, 27
can then be represented as a simple list of community affiliations. We (3):294–320, 2013c
then look for sets of communities that are frequently together: these
are communities sharing nodes across layers.
Node L1 L2 L3
1 C1L1 C1L2 C1L3
2 C1L1 C1L2 C1L3
3 C1L1 C1L2 C1L3
MLComm SLComms MLComm Nodes
4 C2L1 C1L2 C1L3
MLC1 C1L1, C1L2, C1L3 MLC1 1, 2, 3
5 C2L1 C1L2 C2L3
MLC2 C2L1, C1L2 MLC2 4, 5, 6
6 C2L1 C1L2 C3L3
MLC3 C2L2, C3L3 MLC3 7, 8, 9
7 C1L1 C2L2 C3L3
8 C2L1 C2L2 C3L3 (b) (c)
9 C2L1 C2L2 C3L3
10 C1L1 C2L2 C2L3
(a)
Figure 36.2 shows an example. In Figure 36.2(a) we have the Figure 36.2: (a) The communi-
communities found for each layer for each node. Then we decide that ties found in each layer of each
we want to merge communities if they have at least three nodes in node. (b) The merged multi-
common, i.e. they appear in at least three rows of the table. layer communities. (c) The final
Figure 36.2(b) shows the multilayer communities mapping and node-community affiliation.
there are many interesting things happening. First, we only want
maximal sets, meaning that we aren’t interested in returning C1L1 by
itself if we also find it in a larger set of communities. Second, we are
ok if a community gets merged in different sets – i.e. the multilayer
communities can overlap –: C1L2 is part of two maximal sets, MLC1
and MLC2. Figure 36.2(c) shows the final output: the multilayer com-
munity affiliation. A node is part of a multidimensional community
if it is part of all communities composing it. 8
Arlei Silva, Wagner Meira Jr, and
Mohammed J Zaki. Mining attribute-
Node 10 is an example of a final interesting thing: it is part of structure correlated patterns in large
no multidimensional community because its affiliation is a weird attributed graphs. Proceedings of the
VLDB Endowment, 5(5):466–477, 2012
combination of communities. We can decide to let it be without com- 9
Zhiping Zeng, Jianyong Wang, Lizhu
munity affiliation, or to allow it to be part only of its non-multilayer Zhou, and George Karypis. Coherent
communities. closed quasi-clique discovery from large
dense graph databases. In Proceedings
There are other algorithms solving the same problem and inspired
of the 12th ACM SIGKDD international
by frequent pattern mining8 , 9 . conference on Knowledge discovery and
The last solution for this section is inspired by ensemble cluster- data mining, pages 797–802. ACM, 2006
498 the atlas for the aspiring network scientist
ing. Again, we have a community per layer. Then we use the same
strategy I outlined in Section 31.6: we consider each community
partition in each layer as a valid clustering of the same underlying re- 10
Andrea Tagarelli, Alessia Amelio,
lationship via different datasets10 . We then find the “true” clustering, and Francesco Gullo. Ensemble-based
which is the partition that is the closest one to the combination of all community detection in multilayer
networks. Data Mining and Knowledge
partitions.
Discovery, 31(5):1506–1543, 2017
Aggregating communities across layers has some benefits. For 11
Dane Taylor, Saray Shai, Natalie
instance, it might solve the resolution problem of modularity11 that Stanley, and Peter J Mucha. Enhanced
I discussed in Section 32.1. However, all these methods have the detectability of community structure
in multilayer networks through layer
downside of relying more or less on the same assumption: that the
aggregation. Physical review letters, 116
layers are correlated to each other. While this might not be a bad (22):228301, 2016
assumption to start with12 , disassortative layers exist and might 12
Desislava Hristova, Mirco Musolesi,
and Cecilia Mascolo. Keep your friends
represent a problem. close and your facebook friends closer:
A multiplex network approach to the
analysis of offline and online social ties.
36.3 Multilayer Adaptations In Eighth International AAAI Conference
on Weblogs and Social Media, 2014
Multilayer Modularity
I already mentioned how obsessed networks scientists are with mod- 13
Peter J Mucha, Thomas Richardson,
ularity, so you know what’s coming next: multilayer modularity13 . Kevin Macon, Mason A Porter, and
Suppose we’re using the Louvain method, which grows communities Jukka-Pekka Onnela. Community
structure in time-dependent, multiscale,
node by node. If we found a triangle in a layer, can we extend it by and multiplex networks. science, 328
taking a node in a different layer? Intuitively yes, the edge should (5980):876–878, 2010
count because the node is the same. However, if we were to represent
this as a flat network, the new node is not densely connected to the
rest of the triangle: a node couples only with itself, not with its com-
munity fellows. So the coupling edges have to count in some special
way.
In practice, standard modularity works well in each layer sepa-
rately. Consider Figure 36.3: in modularity, the part testing for the
ku kv
density of the community is Auv − . If we use this same part for
2| E |
the inter-layer coupling, we would end up with a case in which the
community cannot be expanded across layers, because there are only
sparse connections between layers. A node couples only with itself in
a different layer, not connecting to its community members, making a
multi-layer community sparser than it actually is. So we need to add
something that will allow us to count the coupling links, so that we
don’t end up with the trivial result of all mono-layer communities.
The full formulation of multilayer modularity is the following:
1 k vs k us
2(| E| + |C |) ∑ Avus − γs
2| Es |
δsr + Cvsr δuv δ(cus , cvr ).
vusr
Let’s break it down – and you can check Figure 36.4 for a graphical
multilayer community discovery 499
1 k k
∑
2(|E|+|C|) vusr [( 2|E s| )
u,v = Nodes
]
A vus − γ s vs us δ sr +C vsr δuv δ (c us , c vr ) Figure 36.4: The adaptation of
modularity to the multilyer set-
ting. Each part of the formula
s,r = Layers
If s=r → same layer is underlined with a color corre-
Modularity in s sponding to its interpretation.
Importance of s
If u=v → same node
Coupling strength
If node u in layer s is in
the same community as
node v in layer r
Normalized by number
of links and coupling
links
thus s and r must refer to different layers. Both deltas can be zero if
we’re looking at uncoupled nodes in different layers. The final delta,
δ(cus , cvr ) is the same as in standard modularity, it is equal to 1 only
if we are looking at nodes inside the same community, i.e. cus = cvr .
With γs we can regulate how important each layer s is for the
community. In practice, γ is a vector of weights, one per layer s of the
network.
Cvsr is the strength of the coupling link, which is a parameter just
like γ is: you can decide how strong the layer couplings should be.
It matters only when we’re looking at the same node connected by
a coupling link across layer (u = v), and so δuv is 1. In this case,
nothing else matters, because δsr is 0 (because s 6= r), so the standard
modularity part cancels out.
Just like in standard modularity, only nodes in the same commu-
nity contribute to the sum, so when node u in layer s is in the same
community as node v in layer r (meaning that δ(cus , cvr ) = 1). This
is normalized by the number of edges (| E|) across all layers plus all
coupling links (|C |).
(a) (b)
If you decide that your inter layer couplings are very strong,
you’ll end up with “pillar communities” where nodes tend to favor
grouping with themselves across layers: the inter layer couplings
trump any intra-layer regular edge. If your inter layer couplings
are weak (low Cvsr ) then you’ll end up with “flat communities” as
nodes prefer to group with other nodes in the same layer. I show an
example in Figure 36.5.
Instead, γ allows you to indicate some layers as more important
than others, as I show in Figure 36.6. If the purple layer is more
important than the green one, multilayer modularity will group in
the community a node that is not connected with the two nodes in
the green layer. If we flip the γ values to make green more important
than purple, the situation is reversed, and modularity will return
different communities.
An alternative way to adapt modularity maximization to multi-
multilayer community discovery 501
(a) (b)
we could be stuck with label oscillation (Section 35.3), this time across
layers. Second, the authors define a quality function that regulates
the propagation of labels. This is done because there might be layers
that are relevant for a community and layers that are not. We do not
want a community, which is very strong in some layers, to “evaporate
away” just because in most layers the nodes are not related. 20
Nazanin Afsarmanesh and Matteo
Next on the menu is k-clique percolation20 . In this scenario, we Magnani. Finding overlapping com-
need to redefine a couple of concepts, particularly what a clique is munities in multiplex networks. arXiv
preprint arXiv:1602.03746, 2016
in a multilayer network, and how we determine when two multiplex
cliques are adjacent. For the first case, we need to talk about k-l-
cliques: a set of k nodes all connected through a specific set of l
layers. Moreover, there are two ways for nodes to be all connected
via the layers: all pairs of nodes could be connected in all layers at
the same time, or they could be connected in only one layer at a time.
The first type of clique is an k-l-AND-clique, the second type is a
k-l-OR-clique. Figure 36.8 shows an example.
(a) (b)
It becomes clear that two k-l-cliques might share (k − 1) nodes
across different layers. In such a case, we need some care in defining
a parameter to regulate percolation. We need a minimum number 21
Caterina De Bacco, Eleanor A Power,
m of shared layers to allow the percolation. If the two cliques do not Daniel B Larremore, and Cristopher
Moore. Community detection, link
share at least m layers, even if they share k − 1 nodes they are not prediction, and layer interdependence
considered adjacent. Figure 36.9 shows an example. in multilayer networks. Physical Review
E, 95(4):042317, 2017
1 1
Figure 36.9: Two 2,2-cliques
2 4 2 4 sharing a 1,2-clique. The edge
color represents the edge layer.
3 3 If m = 1 (a) does NOT percolate
(a) (b) because the rightmost clique
does not share a layer with
The final adaptation we consider is the stochastic blockmodels21 , 22 .
the leftmost clique; (b) DOES
Just like we saw for overlapping and bipartite SBMs, we need to
percolate, since the two cliques
add an additional matrix into our expectation maximization frame-
share the blue layer.
work. For overlapping and hierarchical SBMs it was a community- 22
Natalie Stanley, Saray Shai, Dane
community matrix telling us how strongly communities connect to Taylor, and Peter J Mucha. Cluster-
ing network layers with the strata
each other. In this case, instead, we have a layer-layer matrix telling
multilayer stochastic block model.
how likely it is for two layers to have the same edges. IEEE transactions on network science and
This is neat, because it allows us to model assortative, disassorta- engineering, 3(2):95–105, 2016
multilayer community discovery 503
1
Figure 36.10: A multilayer
network, with the edge color
encoding the layer in which it
7 6 2 5 4 appears.
In some others, you are ok with looking at all layers to find commu-
nities. You cannot rely on a fixed definition of communities based on
density, because it cannot apply to all scenarios.
Thus, you need to have measures to determine when you are in
one case and when you are in another. I proposed a couple in a paper 26
Michele Berlingerio, Michele Coscia,
of mine26 . We decided to call them “redundancy” and “complemen- and Fosca Giannotti. Finding redundant
tarity”. and complementary communities
in multidimensional networks. In
Redundancy is the easiest of the two. To consider a set of nodes
Proceedings of the 20th ACM international
to be densely connected in a multilayer network, we require that conference on Information and knowledge
the edges appear in all the layers we are interested in. If we have a management, pages 2181–2184, 2011b
(a)
(b)
Let’s see some examples to put all these Greek letters to good
use. Let’s consider Figure 36.11, assuming that the network has a
total of three layers. In Figure 36.11(a) we have a high redundancy
506 the atlas for the aspiring network scientist
case. Since the community includes all layers of the network, the
numerator of redundancy is simply the count of edges: 18. The
denominator is 3 (the number of layers) times the number of node
pairs in the community, which is 10, since we have 5 nodes. Thus the
redundancy is 18/30 = 0.6.
Figure 36.11(b) is instead a high complementary case. Variety is 1
by definition, since the community contains all layers of the network.
Exclusivity is 9/10, because there is one pair of nodes (nodes 2 and 3)
which is connected in two layers, and thus it is not counted. Finally,
the standard deviation of the distribution of the edges in c across
layers is the standard deviation of the vector [5, 1, 5], since there are
five edges in the red and green layer, and only one in the blue layer.
This is ∼ 1.88, which is exactly two thirds of the maximum possible,
leaving us with a total homogeneity of 0.33. Thus, complementarity
is 1 × 0.9 × 0.33 = 0.297, penalized by the low representation of the
blue layer in the community.
and thus leaving out some parts of the network that would make
the community denser.
36.6 Summary
36.7 Exercises
Graph Mining
37
Graph Embeddings
Node Embedding
3
2 1 {0.88, 0.69} Figure 37.1: (a) An example
2 {0.88, 0.69} 4
1 graph. (b) One of the possible
4 3 {0.88, 0.69}
4 {1, 1} 5
embeddings of (a), assigning a
5
5 {0.48, 0.66} 1,2,3 two dimensional vector to each
6 6 {0, 0}
7 node. (c) The scatter plot repre-
7 {0.42, 0.1} 6
9 {0.42, 0.1} 7,8,9
8
8 sentation of (a)’s embeddings.
9 {0.42, 0.1}
(a) (c)
(b)
Node Embedding 4
1 {0.01, 0.01} Figure 37.2: (a) A different
6
2 {0.01, 0.01}
3 {0.01, 0.01}
valid embedding of the graph
5
4 {0.01, 1} in Figure 37.1(a), assigning a
5 {0.33, 0.5} two dimensional vector to each
6 {0.05, 0.9}
7 {0.05, 0.05} 7,8,9
node. (b) The scatter plot repre-
8 {0.05, 0.05} sentation of (a)’s embeddings.
1,2,3
9 {0.05, 0.05}
(a) (b)
1. Spectral: this is the oldest category and uses simple matrix forms
to represent the relationships between nodes. The idea is to take
a matrix, which can be the adjacency matrix, and factorize it
minimizing some objective function.
Spectral
Random Walks
science hard
3
2 Φ = max(p(3,4|5) & p(6,7|5)) Figure 37.4: A styilized exam-
ple of DeepWalk. Blue and
1 Φ = max(p(3 or 1,4|5) & p(6, 7 or 8|5))
4 green arrows show two random
5
walks of length five. After each
random walk, the function Φ
6 representing node 5 is updated.
7
9
8
how much the random walk needs to look like a DFS or a BFS. If you
came to node v from node u, common neighbors between u and v
will have a different exploration probability than neighbors of v that
are not connected to u. This allows you to set parameters to explore
structural equivalence rather than modular structure, as I already 14
Haochen Chen, Bryan Perozzi,
Yifan Hu, and Steven Skiena. Harp:
showed you in Figures 37.1 and 37.2.
Hierarchical representation learning
HARP14 improves over both methods by employing a smarter for networks. In Thirty-Second AAAI
way to initialize the weights of the function summarizing your nodes Conference on Artificial Intelligence, 2018
this book, as it deserves a book on its own. For this reason, I suggest 19
Li Deng, Dong Yu, et al. Deep
you some readings19 , 20 , 21 , which might help you figuring out better learning: methods and applications.
what is going on. Foundations and Trends® in Signal
Processing, 7(3–4):197–387, 2014
The fundamental difference between the methods in this class 20
Yann LeCun, Yoshua Bengio, and
and the ones in the previous class, is that random walk models can Geoffrey Hinton. Deep learning. nature,
be considered as a sort of “shallow” learning: they encode the walk 521(7553):436–444, 2015
information with a simple function. Here, instead, we use deep
21
Ian Goodfellow, Yoshua Bengio, and
Aaron Courville. Deep learning. MIT
learning techniques. In general, this allows you to use more complex press, 2016
functions to encode information at the same time. For instance, in 22
Daixin Wang, Peng Cui, and Wenwu
SDNE22 you can learn the first and second order relationships be- Zhu. Structural deep network em-
tween nodes at the same time, using an autoencoder. An autoencoder, bedding. In Proceedings of the 22nd
ACM SIGKDD international conference
as Figure 37.5 shows, is a type of deep neural network that has the
on Knowledge discovery and data mining,
same number of nodes in its output layer as in its input layer. pages 1225–1234, 2016a
25
Joan Bruna, Wojciech Zaremba,
Arthur Szlam, and Yann LeCun. Spec-
There are many approaches defining different strategies to implement tral networks and locally connected
convolution on a network, for instance: using the spectrum of the networks on graphs. arXiv preprint
Laplacian25 , 26 , approaches inspired by extended-connectivity circular arXiv:1312.6203, 2013
26
Michaël Defferrard, Xavier Bresson,
fingerprints27 , a maximization of mutual information between local and Pierre Vandergheynst. Convolu-
and global graph properties28 , and GraphSAGE29 . tional neural networks on graphs with
fast localized spectral filtering. In NIPS,
pages 3844–3852, 2016
H
In SIGKDD, pages 119–128, 2015
Figure 37.7: A troubling set
of connections in an heteroge-
neous network: a paper with
hundreds of co-authors.
32
Yuxiao Dong, Nitesh V Chawla, and
Ananthram Swami. metapath2vec:
Scalable representation learning for
Figure 37.7 shows a stylized depiction of the issue. The paper heterogeneous networks. In SIGKDD,
reporting the discovery of the Higgs boson has 5, 154 co-authors. pages 135–144, 2017
Every pair of co-authors is a valid path in the co-authorship network.
As a result, the embedding of the node representing the paper is
extremely noisy, as it could be visited by 107 walks of length 2, not
even counting the ones that could lead you back to your origin node.
But by instead focusing on each of the metapaths (Section 21.2), we
will have a much cleaner signal.
graph embeddings 519
+Director
(a) (b)
3
2 1 2
1,2,3
Figure 37.10: (a) An example
1 graph. (b) One of the possible
4 3 4
4
5 5
embeddings of (a). The dashed
5
6
6 circles encapsulate the radius
6 7
7
7,8,9 inside which we consider the
9 9 8 nodes as connected together.
8
(b) (c) (c) The version of graph (a)
(a)
reconstructed through its two
A second natural application of graph embeddings is graph sum- dimensional embedding.
marization. We’re going to examine the problem more in details in 45
Jian Tang, Jingzhou Liu, Ming Zhang,
Chapter 38. For now, suffice to say that, if two nodes have very simi- and Qiaozhu Mei. Visualizing large-
lar embeddings, we could just collapse them into the same node. The scale and high-dimensional data. In
WWW, pages 287–297, 2016
benefit would be to have a smaller structure to analyze – less memory
and time consumption for your algorithm – while still preserving the
general properties of the network as a whole.
One common way is to do the following. First, take your graph
and calculate its node embeddings. These, as we saw, are spatial
representations in d dimensions. Then, calculate the pairwise dis-
tance between all these points. You can use the Euclidean distance
or whatever floats your boat. At this point, you can establish a cer-
tain distance k. Nodes that are closer than k get connected together,
otherwise they remain disconnected. This is a graph reconstructed
from the embeddings. You can compare the reconstructed graph with
the original one. The more similar they are, the better the embedding
worked. Figure 37.10 shows an example of this procedure.
You can consider closer points as the same node and summarize
your graph this way. For instance, nodes 7, 8, 9 have the same embed-
ding and thus they could be considered to be the same node, just like
522 the atlas for the aspiring network scientist
nodes 1, 2, 3.
Other classical applications of graph embeddings are node rank- 46
Namyong Park, Andrey Kan,
ing46 , solving classical combinatorial problems like the traveling Xin Luna Dong, Tong Zhao, and
salesman problem47 , and network alignment48 , the problem of fig- Christos Faloutsos. Estimating node
importance in knowledge graphs using
uring out which nodes from two distinct networks might refer to
graph neural networks. In Proceedings
the same real world entity. This is still limited to the realm of net- of the 25th ACM SIGKDD International
work analysis for network analysis’ sake, but we know we can use Conference on Knowledge Discovery &
Data Mining, pages 596–606, 2019
networks – and, therefore, graph embeddings – to solve many more 47
Elias Khalil, Hanjun Dai, Yuyu Zhang,
problems. Some include natural language processing49 , computer Bistra Dilkina, and Le Song. Learning
vision50 , and bioinformatics51 , to cite a few pointers. combinatorial optimization algorithms
over graphs. In Advances in Neural
Information Processing Systems, pages
6348–6358, 2017
37.5 Summary 48
Mark Heimann, Haoming Shen,
Tara Safavi, and Danai Koutra. Regal:
1. A graph embedding is a low dimensional representation of nodes Representation learning-based graph
alignment. In Proceedings of the 27th
in your network. Most commonly, this means representing a node
ACM International Conference on Informa-
as a vector of length d, with similar nodes being represented by tion and Knowledge Management, pages
spatially close vectors. 117–126, 2018
49
Zhouhan Lin, Minwei Feng, Cicero
Nogueira dos Santos, Mo Yu, Bing
2. Depending on how you build them, embeddings can have multi-
Xiang, Bowen Zhou, and Yoshua
ple meanings and facilitate different analyses. For instance, you Bengio. A structured self-attentive
can use embeddings to determine node communities, or identify sentence embedding. arXiv preprint
arXiv:1703.03130, 2017
structurally equivalent nodes. 50
Sijie Yan, Yuanjun Xiong, and Dahua
Lin. Spatial temporal graph convo-
3. One can build embeddings with different techniques: by factor- lutional networks for skeleton-based
izing the adjacency matrix or its spectrum, by exploring a node’s action recognition. In Thirty-Second
AAAI Conference on Artificial Intelligence,
neighborhood via random walks, or by employing deep learning 2018
strategies. 51
Marinka Zitnik, Monica Agrawal, and
Jure Leskovec. Modeling polypharmacy
4. Embeddings in knowledge graphs are a special case. Knoweldge side effects with graph convolutional
networks. Bioinformatics, 34(13):i457–i466,
graphs are heterogeneous networks expressing relations between 2018
real world concepts. In this scenario, embeddings can help us to
uncover previously unknown meanings – i.e. groups of nodes in
the knowledge graph.
37.6 Exercises
3. Is the NMI you get from the previous question better or worse
than the one you’d get from a classical community discovery like
label propagation?
5. Is the AUC you get from the previous question better or worse
than the one you’d get from a classical link prediction like Jaccard,
Resource Allocation, Preferential Attachment, or Adamic-Adar?
38
Graph Summarization
38.1 Aggregation
5
1 1 2 3 4 5
1 4, 5 1 0 2/3 2/3 1/3 1/3
2 4
2 2/3 0 2/3 1/3 1/3
ization. In this case, one might want to just simplify complex motifs 12
Cody Dunne and Ben Shneiderman.
that would tangle up your visualization12 . Motif simplification: improving network
Since we’re shifting perspectives, we might as well keep shifting visualization readability with fan,
connector, and clique glyphs. In
them. So far, we have assumed that aggregation involves the collapse
Proceedings of the SIGCHI Conference on
of nodes into super nodes. However, we could very well collapse Human Factors in Computing Systems,
edges instead. In this case, the edge is aggregated into what we call pages 3247–3256, 2013
(b)
(a)
38.2 Compression
8
Figure 38.4: (a) An input graph.
2 7 ,8 4,5,6
2,3 1 (b) Its summarization via mini-
5 6
1 mization of description length.
+ (1,5)
3 - (4,8)
I label each super node with the
7 4 list of nodes it contains. On the
(b)
(a) bottom, the additional rules we
need to reconstruct (a).
Figure 38.4 shows the approach. The original graph in Figure
38.4(a) can be compressed in the graph in Figure 38.4(b). However,
the summary is not perfect. It assumes the existence of an edge that
does not really exist, and it misses another edge that exists. We add
these two rules to the model M and now the summary is a perfect
reconstruction of Figure 38.4(a). In Figure 38.4(b) we say that nodes 7
and 8 are connected to nodes 4, 5, and 6. This is mostly accurate, but
not completely correct: we need the additional rule that nodes 4 and
8 do not connect to each other.
The objective now is to find the best combination of summary 14
Paolo Boldi and Sebastiano Vigna. The
and additional rules that uses as little information as possible, given webgraph framework i: compression
techniques. In Proceedings of the 13th
or take a margin of error you can set as parameter. There are many international conference on World Wide
information-theoretic approaches in this category14 , 15 , 16 . Web, pages 595–602, 2004
A natural domain of application for compression-based summa-
15
Sebastian E Ahnert. Power graph
compression reveals dominant relation-
rization is in the description of evolving networks. In practice, one ships in genetic transcription networks.
has many snapshots of the same network and they are trying to re- Molecular BioSystems, 9(11):2681–2685,
2013
construct what the whole network looks like17 . The idea is to find the 16
Danai Koutra, U Kang, Jilles Vreeken,
model that is able to best represent all the snapshots you collected. and Christos Faloutsos. Vog: Summariz-
You might feel like aggregation and compression are basically ing and understanding large graphs. In
Proceedings of the 2014 SIAM international
the same category. After all, if you look at Figure 38.4, what you’re conference on data mining, pages 91–99.
seeing is basically an aggregation strategy. The fundamental differ- SIAM, 2014
ence between the two categories is the existence of the model M. In 17
Neil Shah, Danai Koutra, Tianmin
Zou, Brian Gallagher, and Christos
aggregation, there is no M: we simply brute force our way through Faloutsos. Timecrunch: Interpretable
the graph to save every node or edge we can, regardless whether dynamic graph summarization. In
we’re uncovering common patterns or not. The existence of M in the Proceedings of the 21th ACM SIGKDD
International Conference on Knowledge
compression category, instead, forces us only to perform an aggrega- Discovery and Data Mining, pages
tion if it results in a leaner and more elegant M. As a consequence, 1055–1064, 2015
530 the atlas for the aspiring network scientist
38.3 Simplification
Timur Bekmambetov
Johnny Depp
Figure 38.5: (a) An input graph:
Johnny Depp directors (in blue) and actors
Tim Burton
Helena Bonham (in red) connected if they collab-
Carter
Helena Bonham orated with each other. (b) Its
Eva Green Carter simplification via the selection
Bernardo Bertolucci Eva Green
of only actor-type nodes di-
Daniel Craig (b) rectly connected to Tim Burton.
(a)
Those are not the only valid approaches. If the graph also has
metadata attached to nodes – or edges – we can exploit them. For 18
Zeqian Shen, Kwan-Liu Ma, and
instance, Ontovis18 allows the simplification of the graph via the Tina Eliassi-Rad. Visual analysis of
large heterogeneous social networks
specification of a set of attribute values we’re interested in studying. by semantic and structural abstraction.
For instance, the graph in Figure 38.5(a) can be simplified into the IEEE transactions on visualization and
one in Figure 38.5(b), if we’re only interested in knowing the rela- computer graphics, 12(6):1427–1439, 2006
tionships between actors (node type value) working with Tim Burton
(topology attribute).
Ontovis finds the best way to simplify the graph, primarily fo-
cusing on its visual characteristics when plotted in 2D: it is first
and foremost a visualization-aiding tool. Ontovis focuses on node 19
Cheng-Te Li and Shou-De Lin. Ego-
attributes, but one could also switch their focus to edges19 . centric information abstraction for
heterogeneous social networks. In 2009
Note, also, that another difference with sampling is that, in graph International Conference on Advances
simplification, we are not really interested in preserving any spe- in Social Network Analysis and Mining,
cific property of the original graph. This is, instead, a core focus of pages 255–260. IEEE, 2009
network sampling.
When not necessarily focusing on visualization, the simplifica-
tion approach uses a database metaphor to help the user navigate
20
Surajit Chaudhuri and Umeshwar
Dayal. An overview of data warehousing
between different “views” of the network data. A classical database and olap technology. ACM Sigmod record,
infrastructure is OLAP20 , which stands for OnLine Analytical Pro- 26(1):65–74, 1997
graph summarization 531
Place
Place Place Time
Product Product
Time Time
Product What sold when and
Time and place of
Which products sold where, for a subset Which products
all purchases across
when and where of products, places, a place sold
all products
& times
Visually, this would look like Figure 38.7. The central hubs influ-
ence each other, and each is responsible for influencing their branch.
Thus one could summarize the graph as a clique of interacting tribes. 28
Yasir Mehmood, Nicola Barbieri,
Of course, you don’t have to use GuruMine for this. In some cases, re- Francesco Bonchi, and Antti Ukkonen.
searchers have used special adaptations to estimate community-level Csi: Community-level social influence
analysis. In Joint European Conference
influence28 . on Machine Learning and Knowledge
There’s a completely different way to interpret summarization Discovery in Databases, pages 48–63.
Springer, 2013
by influence preservation. We have seen that the spectrum of the
Laplacian can be used to partition a graph – solving the cut problem
(Section 8.4). This is related to diffusion processes: the reason why
the eigenvectors of the Laplacian help you with cutting is because
they are a sort of simulation of a diffusion process, and the edges to
cut are the bottlenecks though which things cannot flow efficiently. 29
Manish Purohit, B Aditya Prakash,
For now, let’s take this for granted, but we’ll see more about this rela- Chanhyun Kang, Yao Zhang, and
tionship in Section 40.2, where we’ll talk about using the Laplacian to VS Subrahmanian. Fast influence-based
coarsening for large networks. In
estimate distances between sets of nodes by releasing a flow from the Proceedings of the 20th ACM SIGKDD
nodes in the origin to the nodes in the destination. international conference on Knowledge
If we want to summarize the graph to preserve these diffusion discovery and data mining, pages 1296–
1305, 2014
properties, we can use the Laplacian to guide our process. What
30
Michael Mathioudakis, Francesco
we’re after, in the simplest possible terms, is a smaller Laplacian ma- Bonchi, Carlos Castillo, Aristides Gio-
trix, with fewer rows and columns, that has the same eigenvectors29 . nis, and Antti Ukkonen. Sparsification
Among other interesting approaches there is SPINE30 . In SPINE, of influence networks. In Proceedings
of the 17th ACM SIGKDD international
one analyzes many influence events in the network. Then, SPINE conference on Knowledge discovery and
only keeps in the network the edges that are able to better explain data mining, pages 529–537, 2011
the paths of influence you observe. You might realize that an edge
is never used to transport information, and thus you could remove
it from the structure without hampering your explanatory power.
Figure 38.8 shows an example of this.
graph summarization 533
(a) (b)
38.5 Summary
38.6 Exercises
Our experience with modeling real world networks tells us that they
are not random: they are an expression of complex dynamics shaping
their topology. Among other things explored in other chapters, this
also means that networks will tend to have overexpressed connection
patterns. Nodes and edges will form different shapes much more
– or less – often than what you’d expect if the connections were
random. For instance, the clustering coefficient analysis tells us that
you’re going to find more triangles than expected given the number
of nodes or edges.
Frequent subgraph mining is the branch of network science that
attempts to find these overexpressed patterns, when they are more
complex than a simple triangle. Want to know whether a square
with a dangling edge appears more often than chance? You have to
perform subgraph mining! In frequent subgraph mining we have a
wealth of techniques to systematically and efficiently enumerate all
possible subgraphs and finding the ones that occur more often in
your networks.
1
Ron Milo, Shai Shen-Orr, Shalev
39.1 Network Motifs Itzkovitz, Nadav Kashtan, Dmitri
Chklovskii, and Uri Alon. Network
We start by defining the building blocks of complex networks. These motifs: simple building blocks of
complex networks. Science, 298(5594):
are network motifs1 , 2 , 3 , 4 . A network motif is a subgraph of your 824–827, 2002
original network with a given topology. A triangle is a motif, a 2
Shai S Shen-Orr, Ron Milo, Shmoolik
square is a motif, the five nodes with the connection pattern in Figure Mangan, and Uri Alon. Network
motifs in the transcriptional regulation
39.1 is a motif. network of escherichia coli. Nature
Generally, one wants to know which motifs are relevant for a genetics, 31(1):64, 2002
network and which ones aren’t. So the standard technique is to
3
Uri Alon. Network motifs: theory and
experimental approaches. Nature Reviews
follow a relatively simple procedure. First, you count how many Genetics, 8(6):450, 2007
times each motif appears in your network. Second, you define a null 4
Jukka-Pekka Onnela, Jari Saramäki,
model of your network, keeping its relevant properties fixed – maybe János Kertész, and Kimmo Kaski.
Intensity and coherence of motifs in
just the degree distribution. Third, you count the expected number weighted complex networks. Physical
of occurrences of the motifs in the null model. Finally, you compare Review E, 71(6):065103, 2005
536 the atlas for the aspiring network scientist
(b) (c)
(a)
with your observation, so that you can build an idea of the statistical 5
Alex Arenas, Alberto Fernandez,
significance of the motif. Santo Fortunato, and Sergio Gomez.
We use network motifs for many different applications. For in- Motif-based communities in complex
networks. Journal of Physics A: Math-
stance, and I won’t get tired of bringing this up, we use them for
ematical and Theoretical, 41(22):224001,
community discovery5 . Of course nobody forces you to have exclu- 2008a
sively the simple motifs I depict in Figure 39.1. One can extend the
6
Lauri Kovanen, Márton Karsai, Kimmo
Kaski, János Kertész, and Jari Saramäki.
concept of network motifs to encompass temporal networks6 , 7 – so Temporal motifs in time-dependent
time-evolving motifs –, and multilayer networks8 , 9 . networks. Journal of Statistical Mechanics:
Theory and Experiment, 2011(11):P11005,
You might have noticed that the examples I show in Figure 39.1
2011
are all very small. They include only a handful of nodes and edges. 7
Ashwin Paranjape, Austin R Benson,
There’s a reason for that. Finding motifs in a large network is a hard and Jure Leskovec. Motifs in temporal
networks. In Proceedings of the Tenth
problem. There are clever techniques to enumerate specific small ACM International Conference on Web
motifs which are reasonably fast10 , 11 . However, in general, one has to Search and Data Mining, pages 601–610.
ACM, 2017
solve the scary graph isomorphism problem, which is the topic of the 8
Federico Battiston, Vincenzo Nicosia,
next section. Mario Chavez, and Vito Latora. Multi-
layer motif analysis of brain networks.
Chaos: An Interdisciplinary Journal of
Nonlinear Science, 27(4):047404, 2017
39.2 Graph Isomorphism 9
Manlio De Domenico, Vincenzo
Nicosia, Alexandre Arenas, and Vito La-
Colloquially speaking, we can state the graph isomorphism problem tora. Structural reducibility of multilayer
as follows: given two graphs, decide whether they are the same networks. Nature communications, 6:6864,
2015b
graph. Two graphs are the “same” if they have the same topology. 10
Ali Pinar, C Seshadhri, and
More formally, graph isomorphism is the search of a function which Vaidyanathan Vishal. Escape: Effi-
maps each node of a graph to each node of the other graph, such that ciently counting all 5-vertex subgraphs.
In Proceedings of the 26th International
they have the same neighbors – identically mapped nodes12 . Conference on World Wide Web, pages
Are the graphs in Figure 39.2(a-b) isomorphic? Table 39.2(c) at- 1431–1440. International World Wide
Web Conferences Steering Committee,
tempts to answer positively: it relabels nodes from Figure 39.2(a)
2017
into nodes from Figure 39.2(b). Since all nodes are connected to 11
Leo Torres, Pablo Suárez-Serrato, and
their identically labeled neighbors, the answer is yes, the graphs are Tina Eliassi-Rad. Non-backtracking
cycles: length spectrum theory and
isomorphic – in fact they’re both 4-cliques.
graph mining applications. Applied
This example is simple enough, but the problem gets very ugly Network Science, 4(1):41, 2019
very soon when we start considering non-trivial graphs. Subgraph 12
Brendan D McKay et al. Practical
graph isomorphism. Department of Com-
isomorphism is an NP-complete problem13 : a type of problem where
puter Science, Vanderbilt University
a correct solution requires you to try all possible combinations of Tennessee, USA, 1981
labeling. This grows exponentially and requires a time longer than 13
Scott Aaronson. P=?np. Electronic
Colloquium on Computational Complexity
the age of the universe even for simple graphs of a few hundreds (ECCC), 24:4, 2017. URL [Link]
nodes. [Link]/report/2017/004
frequent subgraph mining 537
(c)
(a) (b)
• End loop, case 1: you explored all nodes in G1 and G2, then the
graphs are isomorphic;
• End loop, case 2: you have no more candidate match, then the
graphs are not isomorphic.
Most of the heavy lifting is made in the main loop, when checking
whether a match is successful or not. An illustrated example with a
simple graph would probably be helpful. Consider Figure 39.3. VF2
attempts to explore the tree of all possible node matching (Figure
39.3(c)). It starts from the empty match – the root node.
The first attempted match always succeeds, as any node can be
matched to any other node – in this case matching node 1 with node
a. For the second match to succeed, we need that the two matched
nodes are connected to each other. Since node a connects to node b
and node 1 connects to node 2, then the 1 = a and 2 = b match is a
success.
However, attempting to match 3 = c fails, because while node 2
is connected to node 3, node b (matched with 2) isn’t connected to
c (matched to 3). Thus VF2 backtracks: it undoes the last matches
538 the atlas for the aspiring network scientist
*
Figure 39.3: (a, b) Two graphs,
with their nodes labeled with
10 11
1 their ids. (c) The inner data
b 3 1=a 1=b
structure used by the VF2 algo-
rithm to test for isomorphism.
6
a 2 5 2 9 12
I label each node with the at-
tempted match. The node color
c 2=b 2=c 2=a tells the result of the match
1
(green = successful, red = un-
(a) (b) 3 7
4 8 13 successful). I label the edges to
follow the step progression of
3=c 3=b 3=c the algorithm.
(c)
and starts from the last successful match – provided that there are
possible matches to try. In this case there aren’t , so it backtracks
again.
Trying to set 2 = c and 3 = b fails again, for the same reason as
before. So VF2 has to give up also on the 1 = a match and start from
scratch. Luckily, there’s another possible move: 1 = b. When we go
down the tree all matches are successful, until we touched all nodes
in the graph. At that point, we can safely conclude the two graphs
are isomorphic. Note that Figure 39.3(c) doesn’t include the branches
that VF2 never tries in this case, for instance the 1 = c branch.
As expected, multilayer networks provide another level of dif-
ficulty. One can perform graph isomorphism directly on the full 17
Mikko Kivelä and Mason A Porter.
multilayer structure17 , or give up a bit of the complexity and repre- Isomorphisms in multilayer networks.
sent them as labeled multigraphs18 , 19 . IEEE Transactions on Network Science and
Engineering, 5(3):198–211, 2017
18
Vijay Ingalalli, Dino Ienco, and
39.3 Transactional Graph Mining Pascal Poncelet. Sumgra: Querying
multigraphs via efficient indexing. In
International Conference on Database
So far we’ve been dealing with network motifs on a “top-down” and Expert Systems Applications, pages
approach. We have some motifs of interest and we ask ourselves 387–401. Springer, 2016
whether they are overexpressed or underexpressed. This implies that
19
Giovanni Micale, Alfredo Pulvirenti,
Alfredo Ferro, Rosalba Giugno, and
you have to start with your motifs already in mind. This might not be Dennis Shasha. Fast methods for finding
possible. Sometimes, you need a “bottom-up” approach: you want an significant motifs on labelled multi-
relational networks. Journal of Complex
algorithm telling you the frequencies of all possible simple network Networks, 2019
motifs. This is usually the task of frequent subgraph mining.
We split frequent subgraph mining in two: transactional and sim-
ple graph mining. Transactional graph mining was developed first,
because single graph mining introduces some non-trivial problems.
It’s best to start by explaining transactional graph mining, and we’ll
deal with the additional obstacles of single graph mining later (in
frequent subgraph mining 539
Section 39.4).
Data Frequency
Figure 39.4: An example of
4 3 2 frequent itemset mining. The
5 2 1 original data is on the left, one
3 1 line per set of items (itemset).
We calculate the frequency
2 1
of each itemset, including all
1 3 subsets (right).
1
1
The first thing you do is giving up on the idea of finding all sub-
sets. You only want to find the frequent ones. Thus you establish a
support threshold: if a subset fails to occur in that many sets, then
you don’t want to see it. This allows you to prune the search space.
If subset A is not frequent, then none of its extensions can be: they
24
That is, the support function is anti-
monotonic: it can only stay constant or
have to contain it so they can be at most as frequent as A is24 . Thus, shrink as your set grows in size.
once you rule out subset A, none of its possible extensions should
even be considered, since none can be frequent. This usually allows
to perform much fewer tests than the possible ones, and still return
all frequent subsets.
For instance, in Figure 39.4, the orange circle only occurs once.
If the support threshold is 2, we know we don’t need to check the
red-orange, purple-orange, and red-purple-orange subsets. With one
check, we prevented three.
540 the atlas for the aspiring network scientist
Some Rules
Figure 39.6: An example of as-
75%
sociation rule mining. Assume
60% the frequencies of each itemset
100% are the ones from Figure 39.4.
66% We generate rules recording the
relative frequency of observing
Suppose that, in your data, you see 100 instances of sets containing two itemsets. Note that these
objects a1 , a2 , and a3 . And let’s say that, among them, 80 also contain frequencies are not symmet-
object b. Then you can say, with 80% confidence, that the following ric! While the green item only
rule applies: { a1 , a2 , a3 } → b. The { a1 , a2 , a3 } part is the antecedent of occurs 60% of the times a blue
the rule, while b is the consequent. item occurs, every time green
You can also correct your confidence for chance, if you know b’s occurs we also have the blue
overall frequency in the data, and the size of the dataset. This is item.
the “lift” measure. Let’s say that, in our example, b appears in 120
sets. Also, our dataset contains a total of 400 sets. The lift of the
rule is the relative frequency (support) of { a1 , a2 , a3 , b} (80/400) over
the product of the support of the antecedent and the consequent
frequent subgraph mining 541
Graph DB
Figure 39.7: For each of the
patterns on the left we check
Support
Motif whether a graph in the database
(on top) contains it (green
checkmark) or not (red cross).
3 The number of graphs in the
database containing the motifs
is its support.
3
1
542 the atlas for the aspiring network scientist
a
a Figure 39.8: Three possible DFS
b
a d explorations of the graph on
b c top. Blue arrows show the DFS
b
a exploration and purple dashed
c arrows indicate the backwards
c
edges (pointing to a node we
already explored). Remember
a a Start
a
a a a that DFS backtracks to the last
Start
b b b
a d a d a d explored node, not to where the
b c b c b c backward edges points.
b b b
a a a
c c c Marc Wörlein, Thorsten Meinl,
c c c
30
Motifs Data
Figure 39.9: Two motifs (left)
m1 and our graph data (right).
Motif m1 appears only once.
How many times does motif m2
m2 appears?
Ego Networks
One option is to bring back the problem into familiar territory. One
can split the single graph in many different subgraphs and then
apply any transactional graph mining technique. For example, one
could take the ego networks of all nodes in the network. The support
definition would then be the number of nodes seeing the pattern
around them in the network.
Harmful Overlap
Another option starts from recognizing that the entire problem of
non-monotonicity is due to the fact that motif m2 appears twice only
because we allow the re-use of parts of the data graph when counting
the motif’s occurrences. In Figure 39.9, we use the red node in the
data graph twice to count the support of m2 . In practice, the two
patterns supporting m2 overlap: they have the single red node in
common. We could forbid such overlap: we don’t allow the re-use
of nodes when counting a motif’s occurrences. With such a rule, m2
would appear only once in the data graph. If we applied the rule, 33
Michihiro Kuramochi and George
we would have an anti-monotone support definition33 : larger motifs Karypis. Finding frequent patterns in
a large sparse graph. Data mining and
would only appear fewer times or as many times as the smaller knowledge discovery, 11(3):243–271, 2005
motifs they contain.
To see how, consider Figure 39.10. The motif appears four times,
but each of these four occurrences share at least one node. We can
frequent subgraph mining 545
Motif Data
A B Figure 39.10: From left to right:
A 1 D Simple a pattern, the graph dataset,
Overlap and its corresponding simple
2 3 4 C D and harmful overlap graphs.
B C A B
I label each occurrence of the
5 6 7 Harmful motif with a letter, which also
labels the corresponding node
Overlap
8 9 C D in the overlap graph.
occurs twice in the network – you have two independent sets of size
two (A, C and B, D).
1 8 9 1 3
(a)
(b)
39.5 Summary
1. Network motifs are small simple graphs that you can use to
describe the topology of a larger network. For instance, you can
count the number of times a triangle or a square appears in your
frequent subgraph mining 547
network.
39.6 Exercises
2. How many times do the motifs from the previous question ap-
pear in the network? [Link]
39/1/[Link] is included in [Link]
exercises/39/1/[Link]: is the latter less frequent the former
as we would require in an anti-monotonic counting function?
Network Distances
40
Node Vector Distance
points of one to the interest points of the other. Small amounts will
indicate that the images are similar.
3
Ricardo Hausmann, César A Hidalgo,
• In economics3 , 4 , you can represent products as nodes, connected if Sebastián Bustos, Michele Coscia,
there is a significant number of countries that are able to co-export Alexander Simoes, and Muhammed A
Yildirim. The atlas of economic complexity:
significant quantities of them. A country occupies the products in
Mapping paths to prosperity. Mit Press,
this network it can export. From one year to another, the country 2014
will change its export basket, by shifting its industries to different 4
César A Hidalgo, Bailey Klinger, A-
L Barabási, and Ricardo Hausmann.
products. How dynamic is the country’s export basket? The product space conditions the
• In epidemics5 , 6 , 7 , a disease occupies the nodes in a social network development of nations. Science, 317
(5837):482–487, 2007
it has infected. Across time, the disease will move from a set 5
Vittoria Colizza, Alain Barrat, Marc
of infected individuals to another. Similarly, in viral marketing, Barthélemy, and Alessandro Vespignani.
product adoption can be modeled as a disease. The role of the airline transporta-
tion network in the prediction and
predictability of global epidemics.
All these cases can be represented by the same problem formula- Proceedings of the National Academy of
tion. You have a network G. Then you have two vectors: an origin Sciences of the United States of America,
103(7):2015–2020, 2006a
vector p and a destination vector q. Both p and q tell you how much 6
Ayalvadi Ganesh, Laurent Massoulié,
value there is in each node. pu tells you how much value there is and Don Towsley. The effect of network
in node u at the origin, and qv tells you how much value there is in topology on the spread of epidemics.
In INFOCOM 2005. 24th Annual Joint
node v at the destination.
Conference of the IEEE Computer and
All you want to do is to define a δ( p, q, G ) function. Given the Communications Societies. Proceedings
graph and the vectors of origin and destination, the function will tell IEEE, volume 2, pages 1455–1466. IEEE,
2005
you how far these vectors are. There are many ways to do so, which 7
Romualdo Pastor-Satorras and
are organized in a survey paper8 , on which this chapter is based. Alessandro Vespignani. Epidemic
Before we jump into the network distances, it is probably wise to dynamics and endemic states in com-
plex networks. Physical Review E, 63(6):
have a refresher on non-network distances, since it will allow us to 066117, 2001a
introduce concepts that will be helpful later. 8
Michele Coscia, Andres Gomez-
Lievano, James McNerney, and Frank
Neffke. The node vector distance
40.1 Non-Network Distances problem in complex networks. ACM
Computing Surveys, 2020
How to estimate node vector distances on networks is a new and
difficult problem. Let’s take it easy and first have a quick refresher
on the many ways we can estimate distances of vectors without a
network. The easiest way to do it is by assuming a vector of numbers
just represents a set of coordinates in space. If you’re on Earth, with
three numbers you can establish your latitude, longitude and altitude.
That is enough to place you on a position in a three dimensional
space. Another person might be at a different latitude, longitude and
altitude than you. What is the distance between you and your friend?
Easy! You throw a straight rope between you and your friend and its
length is the distance between you. This is the Euclidean distance.
In Section 5.2 I did my very best to connect this intuitive idea in
real life with linear algebra operations. The pain you felt back then
should pay off now. To sum up, if p and q are the vectors defining
your two positions in space, the Euclidean distance is (( p − q) T I ( p −
552 the atlas for the aspiring network scientist
p = (4, 5)
Figure 40.2: Euclidean distance
in m = 2 dimensions. We build
the special p − q vector to have,
((p - q)T(p - q))1/2
at its ith entry, the difference
p2 - q 2 = 4
between the ith entries of p and
q.
p1 - q 1 = 3
q = (1, 1)
In cosine distance you look at the angle made by the vectors con-
necting the two points, as Figure 40.4 shows with a thick green line.
The distance between them is one minus the cosine of that angle.
This is useful, because the cosine is 1 for angles of zero degrees and 0
for angles at ninety degrees. Two points on the same straight line will
have a distance of zero, even if they’re infinitely farther apart on such
a line. For instance, the two points at the bottom of Figure 40.4 are at
a considerable Euclidean distance, but practically neighbors when it
554 the atlas for the aspiring network scientist
Laplacian
9
Michele Coscia. Generalized euclidean
In the Laplacian version9 , we look at the Laplacian of the adjacency measure to estimate network distances.
matrix. Remember from Section 5.3, that the Laplacian L = D − A, In Proceedings of the International AAAI
Conference on Web and Social Media,
with D being the degree matrix and A the adjacency matrix. Since volume 14, pages 119–129, 2020
the smallest eigenvalue of L is zero, L is positive semi-definite.
We use the graph Laplacian because the Laplace operator de- 10
Lawrence C Evans. Partial differential
scribes mathematically statuses of equilibrium10 . In practice, if you equations and monge-kantorovich
have a liquid on a container making waves – i.e. being out of equi- mass transfer. Current developments in
mathematics, 1997(1):65–126, 1997
librium – the Laplace operator will tell you how the diffusion of the
node vector distance 555
δ( p, q, G ) = (( p − q) T L+ ( p − q))1/2 ,
with L+ being the pseudoinverse of the graph Laplacian of G.
Markov Chain
Random walks are helpful to estimate node-node distances (Section
11.4). If we are in the situation of Figure 40.6, we could estimate the
distance between the red and blue node by simply asking how long it
will take for a random walker to go from one node to the other. Here,
we generalize this idea to groups of nodes.
In the Markov Chain distance we start from the assumption that,
given a starting point p, by looking at G we can construct an ex-
[ Z ]u,u = ∑ σu,v
2
.
v ∈V
Now, the problem of this formulation is that it’d make this dis-
tance not symmetric, because to build σu,v we only used p as the
origin of the diffusion. So, if σu,v is the deviation of the diffusion
from v to u, we can also calculate a deviation of the diffusion from u
to v: σv,u . This is done as above, switching p’s and q’s places. Then
2 and σ2 .
the u, u entry in Z’s diagonal is the sum of σu,v v,u
Finally we can write our distance as:
δp,q,G = (( p − q) T Z −1 ( p − q))1/2 .
Annihilation
Let’s take that last thought a bit further. If you let your p and q vec-
tors to diffuse via random walks for an infinite amount of time, they
will distribute themselves to all nodes of G proportionally to their
degree, because they will both tend to approximate the stationary
distribution. In the wavy water basin I mentioned before, p and q are
simply two different waves conditions, while the stationary distribu-
tion is... well ... the stationary distribution: a waveless basin where
the water is at the same level everywhere. Mathematically, this means
∞ ∞ ∞
that ∑ Ak q and ∑ Ak p are the same thing, or ∑ Ak ( p − q) = 0.
k =0 k =0 k =0
Now, the interesting bit is for which value of k this is true or, put
in another words, how fast will p and q cancel out. If they cancel
each other out quickly, it means that they were already pretty similar
to begin with. In fact, that equation would be true at k = 0 if p = q.
So we’re interested in the speed of that equation. This is given us by
the following formula:
∞
δp,q,G = (( p − q) T ∑ Ak ( p − q))1/2 .
k =0
∞
An efficient way to approximate ∑ Ak is by calculating ( I − ( P −
k =0
P∞ ))−1 .
The solutions based on shortest paths start from the assumption that
the problem of establishing distances between sets of nodes can be
generalized from solving the problem of finding the distance between
pairs of nodes. This is a well understood and solved problem: using a
12
Edsger W Dijkstra. A note on two
shortest path algorithm – for instance Dijkstra’s12 – one can count the problems in connexion with graphs.
number of edges separating node u to v. Numerische mathematik, 1(1):269–271,
1959
Non-Optimized
Here we show a set of possible aggregations of shortest path dis-
tances between the nodes in p and q, by taking hierarchical clustering
as an inspiration. There are of course more strategies than the ones
listed here, but I can’t really list them all – and most haven’t really
been researched yet.
When performing hierarchical clustering, there are three common 13
Gabor J Szekely and Maria L Rizzo.
ways to merge clusters according to their distance13 : single, complete, Hierarchical clustering via joint
and average linkage. Single linkage (green in Figure 40.8) means that between-within distances: Extend-
ing ward’s minimum variance method.
the distance between two clusters is the distance between their two
Journal of classification, 22(2):151–183,
closest points. On the other hand, complete linkage (purple in Figure 2005
40.8) considers the distance of the two farthest points as the cluster
distance. In average linkage (orange in Figure 40.8), one calculates
the average distance between all pairs of points in the two clusters as
the distance between the clusters.
Similarly, our aim is to reach the destination from the origin in the
minimum distance possible. In the single linkage strategy, the “cost”
of reaching a destination node is the distance of it from the closest
possible origin node. First, we need to make sure that ∑ p = ∑ q. If
that isn’t the case, we rescale up the vector with the smallest sum so
that this equation is satisfied. For instance, if q had a lower sum, we
transform it: q0 = (∑ p/ ∑ q)q.
node vector distance 559
∑ ∑ pu qv | Pu,v |
∀v∈q ∀u∈ p
δp,q,G = .
∑p
Here it doesn’t matter what we put in the denominator, since we
already ensured that p and q sum to the same value.
vectors, we would again count each path as contributing one third, i.e.
(18/3)/3 = 2.
In complete linkage, we perform a similar operation as in single
linkage, but looking at the farthest destination for each origin. The
farthest destination is 5 → 1, at three steps; then 8 → 4 and 7 → 9 at 14
Andrew McGregor and Daniel Stubbs.
two steps each. Thus, complete linkage will return 3 + 2 + 2 = 7 as Sketching earth-mover distance on
graph metrics. In Approximation, Random-
distance. If we normalized the vectors, we would again count each ization, and Combinatorial Optimization.
path as contributing one third, i.e. 7/3 = 2.3̄. Algorithms and Techniques, pages 274–
286. Springer, 2013
15
Gaspard Monge. Mémoire sur la
théorie des déblais et des remblais.
Optimized Histoire de l’Académie Royale des Sciences
de Paris, 1781
Here we try to be a bit smarter than the aggregation strategies we 16
Frank L Hitchcock. The distribution
saw so far. In this branch of approaches, we try to optimize this of a product from several sources to
numerous localities. Studies in Applied
aggregation such that the number of edge crossing is minimized. Mathematics, 20(1-4):224–230, 1941
If there are no further constraints in this optimization problem, 17
Ira Assent, Andrea Wenning, and
we are in the realm of the Optimal Transportation Problem (OTP) Thomas Seidl. Approximation tech-
niques for indexing the earth mover’s
on graphs14 . In its original formulation15 , OTP focuses on the dis- distance in multimedia databases. In
tance between two probability distributions without an underlying Data Engineering, 2006. ICDE’06. Proceed-
ings of the 22nd International Conference
network. However, it has been observed how this problem can be on, pages 11–11. IEEE, 2006
applied to transportation through an infrastructure, known as the 18
Matthias Erbar, Martin Rumpf,
multi-commodity network flow16 . Specifically, one has to simply Bernhard Schmitzer, and Stefan Simon.
Computation of optimal transport on
specify how distant two dimensions in the vector are. The distance
discrete metric measure spaces. arXiv
needs to be a metric, and the number of edges in the shortest path preprint arXiv:1707.06859, 2017
between two nodes satisfies the requirement. 19
Montacer Essid and Justin Solomon.
Quadratically-regularized optimal
In its most general form, the assumption is that we have a distri-
transport on graphs. arXiv preprint
bution of weights on the network’s nodes, and we want to estimate arXiv:1704.08200, 2017
the minimal number of edge crossings we have to perform to trans- 20
George Karakostas. Faster ap-
proximation schemes for fractional
form the origin distribution into the destination one. This is a high
multicommodity flow problems. ACM
complexity problem, which has lead to an extensive search for effi- Transactions on Algorithms (TALG), 4(1):
cient approximations17 , 18 , 19 , 20 , 21 , 22 , 23 , 24 . For what concerns us, all 13, 2008
21
Jan Maas. Gradient flows of the
these methods are equivalent: they all solve OTP and the difference entropy for finite markov chains. Journal
between them is how they perform the expensive optimization step. of Functional Analysis, 261(8):2250–2292,
Thus, they all return a very similar distance given p, q and G – plus 2011
22
Ofir Pele and Michael Werman.
or minus some approximation due to their optimization strategy –, A linear time histogram metric for
and fall in the same category. improved sift matching. In European
conference on computer vision, pages
More formally, in OTP we want to find a set of movements M such
495–508. Springer, 2008
that: 23
Ofir Pele and Michael Werman. Fast
and robust earth mover’s distances.
In Computer vision, 2009 IEEE 12th
M = arg min ∑ ∑ m pu ,qv du,v , international conference on, pages 460–467.
m pu ,qv pu qv IEEE, 2009
24
Justin Solomon, Raif Rustamov,
where pu and qv are the weighted entries of p and q, respectively; Leonidas Guibas, and Adrian Butscher.
Continuous-flow graph transportation
m pu ,qv is the amount of weights from pu that we transport into qv ; distances. arXiv preprint arXiv:1603.06927,
and d pu ,qv is the distance between them. Then: 2016
node vector distance 561
∑ ∑ m pu ,qv du,v
pu qv
δp,q,G = ,
∑ ∑ m pu ,qv
pu qv
28
Julio E Godoy, Ioannis Karamouzas,
Stephen J Guy, and Maria Gini. Adap-
Moreover, since in MAPF robots cannot be in the same node at
tive learning for multi-agent navigation.
the same time, you still have a problem. Say that we assigned a In Int Conf on Autonomous Agents and
robot to go from u to v in our preprocessing. If pu 6= qv , then either Multiagent Systems, pages 1577–1585. In-
ternational Foundation for Autonomous
u or v has some unallocated weight. Thus we would need to add Agents and Multiagent Systems, 2015
at least a second robot that can either start in u or terminate in v. 29
Jamie Snape, Jur Van Den Berg,
But this violates MAPF. The way we solve the issue is by running Stephen J Guy, and Dinesh Manocha.
The hybrid reciprocal velocity obstacle.
a sequence of MAPF sessions. In each session, we attempt to move IEEE Transactions on Robotics, 27(4):
all the weights that were left over during the previous session. We 696–706, 2011
keep running smaller and smaller sessions until all weights have
30
Glenn Wagner and Howie Choset. Sub-
dimensional expansion for multirobot
been allocated – which we can guarantee by normalizing either p or path planning. Artificial Intelligence, 219:
q so that they sum to the same value, as we did in the non-optimized 1–24, 2015
solutions.
31
Andrew Dobson, Kiril Solovey, Rahul
Shome, Dan Halperin, and Kostas E
There are many algorithms to solve MAPF28 , 29 , 30 , 31 , 32 , 33 , 34 , 35 , 36 , 37 , 38 , Bekris. Scalable asymptotically-optimal
each of them providing a different solution to NVD with our prepro- multi-robot motion planning. In 2017
International Symposium on Multi-Robot
cessing strategy.
and Multi-Agent Systems (MRS), pages
120–127. IEEE, 2017
8
4
Figure 40.11: Attempting to
2
find a pursue solution from
6 red nodes to green nodes. Blue
5 7 arrows show attempted moves.
1
3 Hang Ma, TK Satish Kumar, and Sven
32
Note that this is not one, but a family of measures. One could
replace the Euclidean distance with any other off-the-shelf measure
(cosine, correlation, etc) to estimate the distance between the filtered
564 the atlas for the aspiring network scientist
5. Finally, you can use signal cleaning techniques. You can see
your network as describing sets of sensors that return correlated
results. Thus, two “signals” are far apart if they are reported by
uncorrelated sensors, which are not connected to each other.
40.6 Exercises
2. Calculate the distance using the same data as the previous ques-
tion, this time with the average linkage shortest path approach.
Normalize the vectors so that they both sum to one.
3. Calculate the distance using the same vectors as the previous ques-
tions, this time on the [Link]
40/3/[Link] network, with both the average linkage shortest
path and the Laplacian approaches. Are these vectors closer or
farther in this network than in the previous one?
41
Topological Distances
By far, the most common and popular way to intend the term “net-
work distance” is as the opposite of the similarity between two
networks. The term “network similarity” is, unfortunately, rather
ambiguous, and you might find papers dealing with very different
problems but using the same terminology. For instance, one could
intend “network similarity” as a measure of how similar two nodes
are (see Section 12.2). Or one could be talking about “similarity net-
works”, which are ways to express the similarities between different
entities by connecting the ones that are the most similar to each other
– something you might do via bipartite projections (Chapter 23).
topological distances 567
# Nodes
Figure 41.1: On the left we have
four graphs, each identified by
the color of its nodes. On the
right, I make a two dimensional
projection by recording each
graph’s node count (y axis)
Edge Density and edge density (x axis). The
similarity between two graphs
The main issue is that we still don’t know which set of network is the inverse of their distance
statistics is sufficient to cover the space of all possible networks. in this space.
Whatever dimensions you use to organize your networks will col-
lapse many – possibly dissimilar – networks into the same place in
your scatter plot. This happens in Figure 41.1, where a star (in red)
is confused with a set of unconnected cliques (in blue). This is not
necessarily a bad thing! If the summary statistics you chose are mean-
ingful to you in some fundamental way, this is a feature. However, if
you’re hunting for “universal” patterns, this approach could mislead
568 the atlas for the aspiring network scientist
you.
(a) (b)
(a) (b)
5
5
(a)
(b)
Figure 41.4 can help you to visualize the process. Here, we want to
know how many operations we need to go from the graph in Figure
41.4(a) to the graph in Figure 41.4(b). Starting from node 1, we need
to change its label (from red to blue) and to add the edge connecting
it to node 6. Node 2 is fine, but node 3 needs to replace its edge to
node 5 with one labeled in green. There are no more edits we need to
topological distances 571
The big caveat for using graph edit distances is that they only
work for specific data generating processes. For instance, remember
the Gn,p uniform random graphs from Chapter 13? Two Gn,p graphs
with the same n and (low) p are similar, in the sense that they are re-
alizations of the same process. However, since edges are independent
and the graphs are sparse, they will have almost no edge in com-
mon. As a consequence, their edit distance is large! So what you’re
looking for when using edit distances is for a generating process that
has strong dependencies between edges: the fact that two nodes are
connected implies the presence/absence of other edges in their neigh-
borhood. You should discard graph edit distance measures as soon as
you think that the edges inside your networks are independent from
each other.
Substructure Comparison
20
Xifeng Yan, Philip S Yu, and Jiawei
Substructure comparison20 , 21 is similar to graph edit distance. In Han. Substructure similarity search
this class of methods, you describe the network as a dictionary of in graph databases. In Proceedings of
the 2005 ACM SIGMOD international
motifs and how they connect to each other. Usually, you’d find
conference on Management of data, pages
the motifs by applying frequent subgraph mining (Chapter 39). In 766–777, 2005
practice, graph edit distance is equivalent to a simple substructure 21
Haichuan Shang, Xuemin Lin, Ying
Zhang, Jeffrey Xu Yu, and Wei Wang.
comparison, where the only substructure you’re focusing on is the
Connected substructure similarity
edge. There is not much to say about this class, given its similarity search. In Proceedings of the 2010 ACM
with the previous one: the same considerations and warnings that SIGMOD International Conference on
Management of data, pages 903–914, 2010
applied there also apply here. 22
Thomas R Hagadone. Molecular
When it comes to applications of substructure similarity, the clas- substructure similarity searching:
sical scenario is estimating compound similarities at the molecular efficient retrieval in two-dimensional
structure databases. Journal of chemical
level in a biological database22 . But there are more fun scenarios, information and computer sciences, 32(5):
such as an analysis of Chinese recipes23 . 515–521, 1992
23
Liping Wang, Qing Li, Na Li, Guozhu
Dong, and Yu Yang. Substructure
Holistic Approaches similarity measurement in chinese
recipes. In Proceedings of the 17th
In the holistic category I group a series of approaches that are a international conference on World Wide
mixture of the four previous strategies. Meaning that they use parts Web, pages 979–988, 2008a
topological distances 573
... ...
24
Michele Berlingerio, Danai Koutra,
Tina Eliassi-Rad, and Christos Faloutsos.
Netsimile: A scalable approach to size-
NetSimile24 is one of the many algorithms in this class. I represent independent network similarity. arXiv
preprint arXiv:1209.2684, 2012
its workflow in Figure 41.6. First, NetSimile calculates seven features 25
Godfrey N Lance and William T
for each node of the graph: degree, local clustering coefficient, aver- Williams. Computer programs for
age neighbor degree, average neighbor local clustering coefficient, hierarchical polythetic classification
(“similarity analyses”). The Computer
number of edges among neighbors, etc. Then, these features are Journal, 9(1):60–64, 1966
aggregated across nodes, i.e. NetSimile calculates their summary 26
Xiaohong Wang, Jun Huan, Aaron
statistics like average, standard deviation, etc. This is the signature Smalter, and Gerald H Lushington.
G-hash: towards fast kernel-based simi-
vector of the graph, which can now be used to compare G with any larity search in large graph databases. In
other graph. Any distance measure discussed so far in the book – Graph Data Management: Techniques and
Applications, pages 176–213. IGI Global,
cosine, Euclidean, ... – can be used to perform the comparison. The
2012c
authors focus specifically on the Camberra distance25 . 27
Michele Berlingerio, Danai Koutra,
Similar approaches are graph hashes26 , designed for optimizing Tina Eliassi-Rad, and Christos Faloutsos.
Network similarity via multiple social
graph similarity searches in a graph database possibly containing theories. In Proceedings of the 2013
thousands of graphs; and approaches that are more rooted in social IEEE/ACM International Conference on
Advances in Social Networks Analysis and
theories27 . The latter case is an evolution of NetSimile. Rather than
Mining, pages 1439–1440, 2013b
including a laundry list of all the measures we think we can use to 28
Thomas Gärtner, Peter Flach, and
compare graphs, we pick the ones that are theoretically motivated. Stefan Wrobel. On graph kernels: Hard-
We define which are the criteria of similarity based on different ness results and efficient alternatives.
In Learning theory and kernel machines,
theories, and we discard the rest. The objective is to be able to better pages 129–143. Springer, 2003
interpret the similarity scores. 29
SVN Vishwanathan, Karsten M
A close cousin of holistic approaches is the one of graph ker- Borgwardt, Nicol N Schraudolph, et al.
Fast computation of graph kernels. In
nels28 , 29 , 30 . Just like in NetSimile and in graph embeddings, a graph NIPS, volume 19, pages 131–138, 2006
kernel is the reduction of a complex high-dimensional graph into 30
U Kang, Hanghang Tong, and Jimeng
a vector of numbers. These vectors are then fed to a machine learn- Sun. Fast random walk graph kernel. In
Proceedings of the 2012 SIAM international
ing algorithm that is able to learn the shape of the space in which conference on data mining, pages 828–838.
these vectors live and thus the similarity between them. Just like with SIAM, 2012
574 the atlas for the aspiring network scientist
Information Theory
A radically different approach works directly with the adjacency
matrix of a graph. The idea here is to generalize the Kullback-Leibler
divergence (KL-divergence) so that it can be applied to determining
the distance between two graphs. The KL-divergence is a cornerstone
of information theory and linked with the concept of information
entropy – see Section 2.8 for a refresher.
The KL-divergence is also known as “relative entropy”. From
Section 2.8, you learned that the information entropy of a vector X is
the number of bits per element you need to encode it. Now, of course
when you try to encode a vector, you try to be as smart as possible.
You create a codebook that is specialized to encode that particular
vector. If there is an element that appears much more often than the
others, you will give it a short code: you will have to use it more
often and, if it is shorter, every time you use it you will save bits. This
is the strategy used by Infomap to solve community discovery – see
Section 31.2.
Now suppose you have another vector, Y. You want to know how
similar Y is to X. One thing you could do is to encode Y using the
code book you optimized to encode X. If X = Y, then the codebook
is as good encoding X as it is encoding Y: you need no extra bits. As
soon as there are differences between X and Y, you will start needing
extra bits to encode Y, because X’s codebook is not perfect for Y any
more. The KL-divergence boils down to the number of extra bits you
need to encode Y using X’s codebook.
X X Y Y
=0 0 =0 0 Figure 41.7: An example of the
9 bits 11 bits
0 10
0
6 values
10
6 values spirit of KL-divergence. The
= 10 10 =9/6 = 10 10 = 11 / 6 code we use for X (a) requires
10 = 1.5 bits/value 10 = 1.83 bits/value additional bits to encode Y (b).
= 11 11 = 11 11
(a) (b)
Figure 41.7 presents a rough outline of the idea behind the KL-
divergence – simplified to help intuition. The X vector in Figure
41.7(a) requires 1.5 bits per element. Using its codebook to encode Y
in Figure 41.7(b) increases the requirement to 11 total bits instead of
the original 9. 31
David J Galas, Gregory Dewey,
In its original formulation, the KL-divergence is defined for pairs James Kunert-Graf, and Nikita A
Sakhanenko. Expansion of the kullback-
of vectors. However, one can expand it to allow it to consider differ- leibler divergence, and a new class of
ent inputs31 . One can say that entries in the vectors are dependent on information metrics. Axioms, 6(2):8, 2017
topological distances 575
36
Oleksii Kuchaiev and Nataša Pržulj.
as many nodes as possible36 . If |V1 | 6= |V2 | you will have to face a Integrative network alignment reveals
choice: either you do not map some nodes or you allow nodes from large regions of global network similar-
ity in yeast and human. Bioinformatics,
one network to map to multiple nodes in the other. This common ap-
27(10):1390–1396, 2011
proach can be extended, for instance, by calculating multiple versions
of this matrix using different measures and then seeking a consensus
matrix which is a combination of all the similarity measures.
6 e
b Figure 41.8: Two graphs of
1
4 which we want to discover the
c d
3 alignment. (c) assigns to each
2 a
node pair from (a) and (b) an
5 f
alignment probability.
(a) (c)
(b)
One could also find just a few node mappings with extremely 37
Giorgos Kollias, Shahin Mohammadi,
high confidence and then expand from that seed37 , assuming that and Ananth Grama. Network similarity
the neighborhoods around these high confidence nodes should look decomposition (nsd): A fast and
scalable approach to network alignment.
alike. Of course, a large portion of network alignment solutions rely
IEEE Transactions on Knowledge and Data
on solving the maximum common subgraph problem: if you find Engineering, 24(12):2232–2243, 2011
isomorphic subgraphs in both networks, chances are that the nodes
inside these subgraphs are the same, and thus should be aligned 38
Gunnar W Klau. A new graph-based
to each other38 . Other approaches rely on the fact that isomorphic method for pairwise global network
graphs have the same spectrum, thus similar values in the eigenvec- alignment. BMC bioinformatics, 10(1):S59,
2009
tors of the Laplacian imply that the nodes are relatively similar39 . 39
Rob Patro and Carl Kingsford. Global
Another approach uses a dictionary of networks motifs. Each network alignment using multiscale
node is described by counting the number of motifs it is part of. We spectral signatures. Bioinformatics, 28(23):
3105–3114, 2012
can then describe the node as a numerical count vector. Two nodes 40
Tijana Milenković, Weng Leong Ng,
with similar vectors are similar40 . This approach has to solve the Wayne Hayes, and Nataša Pržulj. Opti-
graph isomorphism problem as well, but it needs to do so only for mal network alignment with graphlet
degree vectors. Cancer informatics, 9:
small graph motifs rather than for – supposedly – large common
CIN–S4744, 2010
subgraphs. This way, it can be more efficient.
3 2
9
1
1 2 3 3
9 1 9 1 Figure 41.10: An example of
2
3 3
2
3 2 2 4 3 .0
2.0
2 . 5 3. 0
8 2 network fusion: (c) is the re-
4 4 8 4 2 8
2
2
1 1
2 3 2.0
sult of the fusion of (a) and (b).
3
4 7 7
2.0
Edge thickness proportional to
7 2
1
5 5 5 its weight.
2 3 2. 5
3 1 2.0
6 6
6
(b) (c)
(a)
Slowly but surely, network fusion crept into network science and
found applications that go beyond increasing the performance and
applicability of neural networks. For instance, consider genomic data.
You can collect samples of interactions from many individuals. These
578 the atlas for the aspiring network scientist
are similar, but not always the same. You might want to combine 45
Bo Wang, Aziz M Mezlini, Feyyaz
them to create a prototypical interaction network45 . Alternatively, it Demir, Marc Fiume, Zhuowen Tu,
could be that some of these samples are incomplete, and you can use Michael Brudno, Benjamin Haibe-Kains,
and Anna Goldenberg. Similarity
their fusion as the complete genomic data.
network fusion for aggregating data
types on a genomic scale. Nature methods,
11(3):333, 2014a
41.4 Summary
41.5 Exercises
Visualization
42
Node Visual Attributes
12 12
10 10
I II III IV
x y x y x y x y 8 8
y
y
10 8.04 10 9.14 10 7.46 8 6.58
I II
6 6
8 6.95 8 8.14 8 6.77 8 5.76
4 4
13 7.58 13 8.74 13 12.74 8 7.71
4 6 8 10 12 14 16 18 4 6 8 10 12 14 16 18
9 8.81 9 8.77 9 7.11 8 8.84 x x
11 8.33 11 9.26 11 7.81 8 8.47
14 9.96 14 8.1 14 8.84 8 7.04 12 12
III IV
7 4.82 7 7.26 7 6.42 8 7.91 6 6
4 6 8 10 12 14 16 18 4 6 8 10 12 14 16 18
(a) x x
(b)
You cannot simply take away the message that any quantitative
measure of node importance is an equally good choice for your node
size. Some of those measures will not highlight what you want to
highlight. For instance, the degree is not always the right choice.
Consider Figure 42.5(a): would you think to use the node’s degree
as a measure of its size? If you do, you end up with Figure 42.5(b)
where the node playing arguably the strongest role in keeping the
network together almost disappears. A much better choice, in this
case, is betweenness centrality (Figure 42.5(c)).
It shouldn’t surprise you – after all the network analysis we’ve
node visual attributes 585
100
10-1
10-2
p(k)
10-3
10-4
100 101 102 103 104
k
(a)
(b)
Figure 42.6: (a) Degree distri-
To counteract this, you need to apply a quasi-logarithmic scaling.
bution of the Marvel social
If you’re creating your visualizations programmatically you can have
network example. (b) Visualiz-
an actual log scale, although you probably will still have to manually
ing the network with a linear
tweak it a bit to make the result more pleasing. The idea is to have
node size map, where the de-
diminishing returns to the contribution of the degree to the node
gree directly determines the
size. The differences in size from the minimum degree, to the average
node size.
586 the atlas for the aspiring network scientist
degree – which can be quite low – are big but, from that point on, the
contribution to the node’s size plateaus.
(a) (b)
If you do so, you can find new clusters that were previously clut-
tered by the huge nodes, or that had a low degree and so they did
not pop up. You can compare the two hairballs in Figures 42.7(a) and
42.7(b). Note that the visualization is still truthful: we’re never going
to make nodes with lower degree larger than nodes with higher de-
gree. That would be bad and land you in a corner. We’re just making
the visualization more useful.
This is probably a good place to stop and make a disclaimer. Even
if eyes are the highest bandwidth sensors we have, it doesn’t mean
they are flawless. Nor that our monkey brain is able to use the in-
formation they gather in a perfect way. Human perception is flawed
and you cannot expect that something a computer understands will
appear obvious to your viewers as well. In the case of node size this
takes the form of the confusion between radii and areas.
Unless otherwise specified by the software/program of choice,
you are going to decide the radius of the node when determining its
size. This can be trouble if you don’t handle this choice properly. The
reason is that, when you increase the radius, you are substantially
performing a linear increase: you think that, if the degree increases
by one unit, you should increase the radius by one unit. Unfortu-
nately, what a viewer will perceive is you changing the area of the
circle. The crux of the problem is that a radius is a one dimensional
quantity, and it should never be used for controlling a two dimen-
sional one such as an area – which is what your readers perceive.
You think you’re increasing something linearly, but you’re actually
raising that increase by the power of two.
Figure 42.8 shows you why you need to be well aware of the
difference. What you think is a small increase can seem humongous
to your reader.
node visual attributes 587
42.2 Color
The second obvious feature to manage for your nodes is their color. If
we routinely use node sizes for quantitative attributes, we primarily
use node color for qualitative ones. The reason is that, while humans
perceive size as quantitative, color hue is not perceived in the same 11
William S Cleveland and Robert
way. Cleveland and McGill11 distinguish between different data types McGill. Graphical perception: Theory,
and how much different graphical features are effective for each data experimentation, and application to the
development of graphical methods. Jour-
type. Color is good for nominal attributes – categories that cannot nal of the American statistical association,
be compared/sorted, like “apple” vs “orange”. Color could be used 79(387):531–554, 1984
for ordinal attributes, that are still categories but can be compared –
for instance days of the week, Monday comes before Tuesday. Color
is terrible for quantitative attributes, for which areas are a more
effective tool. As always, what follows is based on my experience
and, if you want or need more in-depth explanations, you should
check out the paper.
When it comes to network visualization, this implies that we put
nodes into classes and we use colors to emphasize that different
nodes are in different classes. Classical examples can be the node’s
community – Part IX –, or its role – Section 12. I already mentioned
nodes can have metadata, and these metadata could be categorical.
For instance, in a network connecting online shopping products
because they are co-purchased together you could use the color to
determine their category (outdoors, rather than kitchen, rather than
electrical appliances).
You could still use node color for ordinal attributes, and maybe Figure 42.9: (top) A gradient
palette for diverging quantities
and a meaningful middle point
of the spectrum. (bottom) An
intensity gradient, useful to go
from zero to a maximum value
without a meaningful middle
point.
588 the atlas for the aspiring network scientist
for quantities as well, provided that you have clear and intuitive bins.
The way one would use colors for quantities is by implementing a
gradient. A classical one is a blue-red spectrum for temperatures:
this is a diverging scale that can be useful, e.g., if you have some
sort of correlation data. You have a very precise and semantically
meaningful middle point, and nodes can diverge in either of two
directions, as I show in Figure 42.9 (top). Otherwise, if we’re talking
of a more classical intensity – say how much money a customer spent
in your online shop – you want a simple sequential gradient, just like
the one in Figure 42.9 (bottom).
There are many things you need to take into account when using 12
Samuel Silva, Beatriz Sousa Santos,
colors. One of the trickiest ones is cultural associations12 . When you and Joaquim Madeira. Using color in
visualize something, your visualizations come after centuries – if not visualization: A survey. Computers &
Graphics, 35(2):320–333, 2011
millennia – of other people using colors for different tasks. These
usages ingrained in our mind a quick way to decode information.
For instance, we associate red with danger, yellow with caution,
green with “good to go”. Black is death, and – stereotypical – blue
is for boys and pink is for girls. But blue is also Democrat against
red Republican if we’re talking about elections in the US – which,
interestingly, is the opposite of the left-right wing spectrum for
other countries in which red is communism. It all depends on the
context in which you’re visualizing. Color can aid you in making
your visualization quicker to decode, but if you’re instead using it
differently from a convention it can make things harder.
This doesn’t even take into consideration the deficits in the phys-
ical perception by humans. Just to repeat myself – we are very lim-
ited when it comes to distinguish colors. If you ask your laptop
how many colors there are out there, a popular reaction would be
counting the number of possible RGB combinations and to reply: 16
millions! That would be very wrong for any human with a hint of
common sense. In fact, what I would say is that you should never
use more than nine colors in your visualization, and I’m sure that a
few of my data designer friends are already gasping in horror to the
extent of my liberalism. Nine, for them, is already way too much.
It’s not just about the quantity of colors, though, it is also about
how to choose and use them. How many colors would you say I
used for the nodes in Figure 42.10(a)? If you guessed 16 – which is
the correct answer – you’re very lucky, or you have some Truman
Capote levels of pattern recognition. In the network I highlighted
three groups of nodes. These have different colors, believe me or
not. Few – if any – people would be able to tell without scanning
the figure for more than a handful of seconds. Requiring this level of
effort from your viewer means to lose them.
Why does Figure 42.10(a) fail? Because it assumes that RGB is a
node visual attributes 589
(a)
14
Mark Harrower and Cynthia A
you can use is the Color Brewer interactive tool14 , which will gen- Brewer. Colorbrewer. org: an online tool
erate the palettes for you15 . Color Brewer is embedded in many for selecting colour schemes for maps.
The Cartographic Journal, 40(1):27–37,
software/programming packages that you might already use for
2003
your visualizations, including R16 , Cytoscape (since version 3.7.1, 15
[Link]
for earlier version you need the Color Cast plugin), QGis, Python 16
Erich Neuwirth and R Color Brewer.
(Matplotlib and Seaborn, for instance), and Matlab. Colorbrewer palettes. R package version,
pages 1–1, 2014
Figure 42.11 uses the Color Brewer space and fixes one of the
many problems of Figure 42.10(a). In Figure 42.11 we use also fewer
colors – just nine – which is always good.
Color Brewer and RGB are not the only possible color spaces you
could use. If you are creating visualization for printing, you should
use a CMYK color space. This is similar to RGB, but RGB is an
additive color space, while CMYK is subtractive. Additive color spaces
describe how different wavelengths of light add to each other, which
is how computer screens work. Subtractive color spaces, instead,
describe how ink combines on the page, which is why it’ll show
better how things will look in print. HSV and HSL are alternative
color spaces which transform RGB to be more perceptually-relevant:
we as humans don’t really perceive colors as combinations of red-
blue-green, but as variation in hue, saturation and lightness, which is
what HSL stands for.
To wrap up this chapter, let’s see a few more things you can do
to your nodes. They both stem from the same idea: your nodes
represent something, and so you want to communicate this to your
viewers.
The first strategy involves node labels. If you want the audience to
know something, you simply tell them. You plaster some text on top
of your nodes and you call it a day. In my opinion, this is a desperate
move and it should be avoided if possible. Just as in movies, also
in data visualizations it’s better to “show, not tell”. In other words,
nobody wants to read your network. They want it to speak to them.
That is not to say that sometimes a good choice of node labels
can enhance your visualization. You can practically transform your
network into a glorified word cloud. I don’t love it, but I grudgingly
admit that sometimes it works. An example could be the one in Fig-
ure 42.14 – although in this case one should choose a less saturated
color for the nodes, because the current red goes in the way of the
readability of the label. My rule of the thumb is that the node label
font size should have a one-to-one correspondence to the node size. It
would look weird to have a gigantic label on top of a tiny node, and
vice versa.
The second visual attribute you could play with is the node’s
border. This is an interesting one, because it could be used for quanti-
tative and qualitative attributes at the same time. For the node border,
you can both decide the color and the thickness. Again, you should
really ask yourself whether you really need to do it. Personally I
almost never touch node borders – I’d say that in 99% of my visu-
alization the border is invisible. If you already have node sizes and
592 the atlas for the aspiring network scientist
RWA
BHR SWZ
NPL UGA
KHM BTN MDV
THA
HKG
YEM
SOM
BDI Figure 42.14: A network with
PHL OMN SDN KEN
KWT LKA TZA MWI
PNG
LAO
MMR MYS
DJI
JOR
PAK
node labels conveying informa-
IND
NZL BGD ZMB
BHS IDN LBN
tion about a node’s importance.
AUS
FJI
SAU SYR
TGO
ZAF ZWE
SLB MLI
MNG BEN
SGP KOR KGZ
TJK
AFG NAM
In this trade network, it is
GHA
GIN SEN
CIV GNB BFA
BRN
IRQ
BWA
MOZ
TUR
BLR ARM Product.
BGR
COG KAZ
MRT MDA
TCD LSO LBY MUS
NER
CHE AZE SVK
ROM
ALB
GRC
CHL ARG GAB
CYP
LTU
TUN
NGA POL
BRA
ERI
PER
USA ISR
GBR RUS
CZE LVA
PRY
BOL
CMR ITA HUN MKD
PRT
VEN SUR AUT
IRL
FRA
CAF
DEU
URY PAN HRV EST
ECU
MEX MAR CPV
HTI
COL SWE
BIH
DZA SVN
GUY
ESP NLD
CAN
DOM TTO BLZ
GNQ
BEL NOR
NIC GTM
JAM FIN
ISL DNK
CRI LUX
CUB
HND
SLV
colors, adding a border of a different size and color would just cause
information overload in your reader’s brain. You should only do it if
there are extremely clear patterns in your network, which involve no
more than a handful distinct values, and that can be easily parsed.
For instance, in Figure 42.15, we could have two nodes of same
size and color – perhaps these are two plants in the same country
(color) and employing the same number of people (size). However,
they process different products (border color) and they have differ-
ent throughput in number of products processed per day (border
thickness).
Another strategy is more creative – and for this reason you should
apply tons of caution if you want to go this way. It involves xeno- 18
[Link]
graphic18 . This translates to “weird visualizations”, stuff that has
very specific and almost unique use cases, and thus it’s likely to
choose a style that people haven’t seen before. You can be creative
with what you put on your nodes, as long as you don’t abuse it and
it has a meaningful relationship with your message.
One obvious way you can communicate differences in kind when
it comes to a node would be to represent it not as a dot, but as a fig-
ure. The classical case is by transforming the node’s shape. I already
used this approach in this book for bipartite networks. A classic way
to visualize them is to use one node shape for V1 nodes, and another
node visual attributes 593
chapter.
42.4 Summary
1. The first visual attribute of nodes is their size. Usually, you want
to show quantitative attributes via size – the degree, the capacity,
etc. Be aware that you should always manipulate the area of the
node, which is what your viewer perceives. If your software only
allows you to control a node’s radius, keep in mind that your area
will change quadratically for each linear change of the radius.
2. Second, you can control a node’s color. Usually, this is for qualita-
tive attribute, e.g. community affiliation. Use no more than nine
distinct colors, from a perceptual-aware space (not RGB rainbows!).
42.5 Exercises
the node size. Make sure you scale it logarithmically. This can be
performed entirely via Cytoscape. (The solution will be provided
as a Cytoscape session file)
Size
The equivalent for edges of node size is the thickness. As in the
previous case, this is mostly for quantitative attributes on edges. The
most trivial one is the edge’s weight: heavy edges usually appear to
be more thick. Another common use case is to put edge betweenness
as the determinant of the edge thickness. This works well when used
in conjunction with nodes sizes following the same semantics. It
gives a sense of balance to the visualization, so you can see which
edges are contributing to the node’s centrality. Figure 43.1 shows an
example.
Color
clues can sum and make each other clearer. This is the case for edge
betweenness, determining both color and thickness of the edges in
Figure 43.3(a).
Transparency
Transparency is another aspect in which edges diverge from nodes.
In the previous chapter I mentioned that nodes should be fully
visible, and provided only a single use case in which I believe trans-
parency can add something to the visualization by removing the
nodes from sight. When it comes to edges, I usually abuse trans-
parency lavishly. Most commonly, I make transparencies work to-
gether with colors, to reinforce them. Significant links have darker
edge visual attributes 599
colors and are more opaque. The objective of playing with the alpha
channel for edges is to create a visual hierarchy, where nodes come to
the forefront and edges go to the background.
However, sometimes you can play with edge transparencies even if
you don’t have any attribute to attach to them at all! This is because
of the sheer number of edges: in most real world networks, they
are going to overlap to each other, no matter what. Thus edge trans-
parency, even a fixed value, can highlight structure, because there
are going to be more overlaps in dense areas of the network than in
sparser ones. Thus, you can highlight such clusters even without any
edge metadata.
(a) (b)
Labels
Just like with nodes, also with edges you can be... edgy in how you
visualize them. There are two fundamental aspects I’m going to
mention here: shapes and bends.
The classical edge visualization is as a straight, solid line. This is
what you should do in 99% of the cases. However, in many cases,
you might want to slightly change this shape. The most common
shape change for edges is when you are working with directed con-
nections. In this case, the convention is to add an arrow that indicates
the direction of the edge. The arrow points from the originator of the
edge to the target.
Directed networks are more challenging to visualize than you
might think. The reason is not only that you’re doubling the possible
number of edges, which is true and it is an issue. But the real trouble
is that now you might have a significant number of double edges be-
tween the same two nodes: u → v and u ← v. This might make your
visualization a real mess. One convention you can implement is not
to actually draw the two edges. What you can do is to draw a single
edge and add to it a second arrow pointing in the opposite direction
if that edge is reciprocal. Figure 43.5 shows how this strategy looks
like.
meaning that the (u, v) edge is not a straight line from u to v any
more, but it takes a “detour”. Why would you want to do this?
There are fundamentally two reasons. Figure 43.5(a) provides an
example: since there are two edges between the nodes, we want the
visualization to be more symmetric and pleasant, and thus we bend
the two edges.
More often, edge bends are used to make your network layout
more clear. You bend edges to bundle together the ones going from
nearby nodes to other nearby nodes. Since this is done mostly to
clean up the visualization after you already decided where the nodes
should be placed, I will deal with this topic in the network layout
chapter (Chapter 44).
Let’s recap all the advice I gave you on node and edge visual at-
tributes and see a case of applying each feature one by one to go
from a meaningless hairball to something that conveys at least a little
bit of information. Our starting point is the smudge of edges you al-
ready saw a couple of times: that’s Figure 43.4(a). Note that this isn’t
really the starting point, because we already settled on a network
layout, but that will be the topic of the next chapter.
The usual order I apply to my networks after I settled on a layout
is the following:
1. Edge transparencies;
2. Edge sizes;
3. Edge colors;
4. Node sizes;
5. Node colors.
So let’s do this.
Edge transparencies. In this network, I do have quantitative infor-
mation, that is the edge betweenness of each connection. However,
I think that it’s better if I limit that to the other edge visual features,
so I fix the same edge transparency to all links. The result is Figure
43.4(b).
Edge sizes & colors. We now move on to use edge betweenness. I
merge the two steps of edge size and color into one, because using
simply the thickness does not make a significant difference with the
previous visualization. Compare Figure 43.4(b) with Figure 43.6(a)
and see that not much has changed. So I apply a Color Brewer color
602 the atlas for the aspiring network scientist
(a) (b)
(a) (b)
color palette that Color Brewer provides, I’m partial to Set1 – even if
it is not exactly color blind friendly. See Figure 43.7(b) for the final
result.
Figure 43.7(b) still has a long way to go before we can call it a
good network visualization. But, compared to the starting point in
Figure 43.4(a) we can definitely say more things about its structure.
Which is exactly what network visualization is for.
A way to improve this picture would be to choose a better layout,
and to apply some tweaks to it. That is the topic of next chapter.
43.4 Summary
2. Differently from node colors, edge colors are used mostly for
quantitative attributes. Usually, there are more edges than nodes
in a network, thus it is harder to limit the number of edge colors.
Anyhow, you should not have more than nine different colors in
total, whether they are node or edge colors. Classical qualitative
edge colors choices are layers or link communities.
43.5 Exercises
100
10-1
10-2
p(k)
10-3
10-4
100 101 102 103 104
k
(a)
(b)
Figure 44.1: (a) In a scatter plot,
Changing graphical elements as we saw in the previous chapters is
changing the x-y coordinates
good, but network data has a peculiarity that other data types don’t
of a point is forbidden, because
have. Networks are a particular data type that allow our representa-
that will change the data. (b) In
tions an additional degree of freedom. This influences the network
a network, you can move nodes
visualization. To see what I mean consider that, in a scatter plot, you
around, as long as you do it
cannot move the points around because that would change the data.
with a consistent set of rules to
But in a network you have connections. What count is not the “abso-
all nodes.
lute” position of a node, but its “relative” one. Figure 44.1 shows an
example of what I mean.
network layouts 605
Hierarchical
If force directed and its variants are a good default choice, they are
not the only way to display your networks. As I concluded in the
previous section, they have some pretty limited use cases. What
happens when we go out of those use cases?
(a)
(b)
Circular
A second scenario to consider is the case of an extremely dense
network. The network could be so dense, that the force directed is
not able to pull nodes apart and show structure. In this case, the first
step of the solution involves considering a layout that might not seem
the best for the job, but has a few tricks up its sleeve: the circular
layout. Which is exactly what it sounds: it places nodes on a circle,
equidistant from one another.
The first part of the trick in using circular layouts is not to display
the nodes in a random order, but choosing an appropriate one. Ide-
ally, you want to place nodes in bunches such that most connections
happen across neighbors. Usually this is achieved by identifying the
nodes’ attribute which groups them best. You can also run a custom
algorithm deciding the order and then provide that as the attribute
for the circular layout.
Figure 44.6 shows an example. In the figure you can see that the
layout still works: it shows how most connections remain within
the communities, and clearly points at how many and where the
inter-community connections are. In a force directed layout, these
connections would stretch long and be forced in the background of
denser areas, with the effect of being difficult to appreciate. However,
the real kicker for circular layouts happens when you consider the
610 the atlas for the aspiring network scientist
(a) (b)
3 3
(a) (b)
Seeing that configuration, a viewer would instantly assume that
node 1 is connected to both node 2 and 3, without a connection
between the latter two. This, as we know, is wrong. Worse still,
there’s no way by looking at the configuration to the right to know
what really is going on between these three nodes. It might be that
node 1 is not connected to either of those, and ended up there by
accidents of your layout. Or that could be a squeezed triangle. Or it
might be even true that the three nodes are not connected at all to
each other, and that’s just a long edge connecting two other nodes
out of sight. There are so many ways to lie in a 2D network layout –
whether you do it accidentally or on purpose.
Edge bends can partially save you. In particular, the trick is to use
organic or orthogonal edge routing18 , implemented by yFiles. To
18
Tim Dwyer, Kim Marriott, and
Michael Wybrow. Integrating edge
know what it looks like, consider Figure 44.10. In the figure, there routing into force-directed layout.
are many cases in which the vanilla force directed layout would pass In International Symposium on Graph
Drawing, pages 8–19. Springer, 2006
a straight edge across the nodes. As it happens, some of those cases
were actually two edges with a node in between. The organic edge
router would not change the edge shape in that case. But some edges
network layouts 613
In this section, we break the assumption that nodes are circles and
edges are lines. We try to find weird ways to summarize the network
topology in a way that is more compact and compelling.
Matrix Layouts
What’s the last resource for networks in which the density is too high
even for a circular layout plus edge bundles? If your network is so
dense, then you don’t have a network: you have a matrix and you
should visualize it as such. In these cases, what matters more is not
really which area is denser than which other, but which blocks of If your network is this dense and also
19
nodes have connections with higher and lower weights.19 unweighted, consider changing job.
(a) (b)
Color scales should be chosen depending whether your edge
weights are defined as progressive or divergent. The first case, in
Figure 44.11(a), is the classical case of an edge indicating the intensity
of the connection between two nodes, or the cost of edge traversal.
The second case, in Figure 44.11(b), is a classical correlation network.
This is a default visualization scenario, as correlations are always
defined between any pair of nodes, and thus the network will be
complete.
That is not to say that this is the only use case of a matrix visual-
ization. Even for sparser networks, sometimes a matrix is worth a
thousand nodes. Consider the case of nestedness (Section 28.4), a par-
ticular core-periphery structure for bipartite networks where you can
sort nodes from most to least connected. The most connected node
connects to every node in the network, while the least connected
nodes only connects to the nodes that everyone connects to. This
sort of linear ordering naturally lends itself to a matrix visualization.
614 the atlas for the aspiring network scientist
The aim of these rules is to divide nodes into classes. Nodes in the
same class will be grouped together on the same axis. Then, the user
selects a specific measure, which determines the position of a node
on that axis. The hive plot will then attempt to place the axes in such
a way to minimize edge crossings. If there are connections between
the nodes on the same axis, the axis will be duplicated in order to
avoid confusing loops.
network layouts 615
(a)
(b)
Graph Thumbnails 22
Vahan Yoghourdjian, Tim Dwyer,
Karsten Klein, Kim Marriott, and
Just like hive plots, also graph thumbnails22 aim to provide a de- Michael Wybrow. Graph thumbnails:
terministic layout, where two isomorphic graphs result in the same Identifying and comparing multiple
graphs at a glance. IEEE Transactions on
visualization. In graph thumbnails, we decide to give up the ability
Visualization and Computer Graphics, 24
of analyzing local structures. We are not seeing each individual node: (12):3081–3095, 2018
the visualization is a summary of the graph’s global structure. The
idea is to dissect a graph into its main core components in a hierarchi-
cal fashion. Each core is then visualized as a circle, whose color tells
us its core level. You should use graph thumbnails when you need
to compare a large number of graphs in a compact way and you care
about the high-level organization of the graph as a whole, rather than
the meso-level communities – or the individual nodes.
The decomposition is done via the classical k-core detection – see
Section 11.7. Each connected component of a network is part of a 1-
core. Then, there could be multiple k-cores around the network. Each
k-core is represented as a circle, and it is nestled inside the (k − 1)-core
that contains it.
616 the atlas for the aspiring network scientist
(b)
(a)
Probabilistic Layout
Following the same “we can’t visualize all nodes” philosophy of 24
Christoph Schulz, Arlind Nocaj,
graph thumbnails, we have probabilistic layouts24 . This technique is Jochen Goertler, Oliver Deussen,
handy when you have a generic guess of where the nodes should be, Ulrik Brandes, and Daniel Weiskopf.
Probabilistic graph layout for uncertain
but you cannot draw them all. You should use such layouts especially network visualization. IEEE transactions
for very large graphs that you couldn’t visualize otherwise, because on visualization and computer graphics, 23
(1):531–540, 2016
they have too many nodes and/or edges.
The idea is as follows. First, you sample the nodes in your net-
work, taking only a few of them. Then you calculate their positions
using a deterministic force directed layout. You repeat the procedure
multiple times, obtaining, for each node, a good approximation of
where it should be. If you have nodes that you never sampled, you
can reasonably assume that they are going to be in the area surround-
network layouts 617
Revealing Matrices
One key visualization technique is scatterplot matrices or SPLOMs.
When you have multiple variables in your dataset, you might be
interested in knowing which one correlates with which other. So you
can create a matrix where each row/column is a variable, and each
cell contains the scatter plot of the row variable against the column
variable. Figure 44.16 shows an example.
The same visualization technique can be applied to networks. In 25
Maximilian Schich. Revealing matrices.
revealing matrices25 , each row/column of your matrix is an entity. 2010
Then, each cell of the matrix contains a bipartite network, where the
nodes of one type are the row entity and the nodes of the other type
are the column entity.
One defect of SPLOMs is that the main diagonal of the matrix is a
bit awkward. In it, the row variable and the column variable are the
same. Thus the scatter plot is meaningless, as it is the same variable
on the x and y axes: a straight line. One could modify it by showing
some sort of statistical distribution of the variable, but that would
mean breaking the axis consistency of the SPLOM. For this reason,
the main diagonal of a SPLOM is often omitted.
618 the atlas for the aspiring network scientist
Product Space
26
César A Hidalgo, Bailey Klinger,
The Product Space26 , 27 is a popular example. The Product Space A-L Barabási, and Ricardo Hausmann.
The product space conditions the
is a network in which each node is a product that is traded among
development of nations. Science, 317
countries in the global market. Two products are connected if the sets (5837):482–487, 2007
of countries exporting them have a large overlap. The idea of this 27
Ricardo Hausmann, César A Hidalgo,
Sebastián Bustos, Michele Coscia,
visualization is to show you which products are similar to each other,
Alexander Simoes, and Muhammed A
because if your country can make a given set of products, via the Yildirim. The atlas of economic complexity:
Product Space it can figure out which are the most similar products it Mapping paths to prosperity. Mit Press,
2014
should consider trying to export.
(b)
(a)
Figure 44.17: The Product
The original way to try and visualize the Product Space was a Space. (a) Classical force di-
simple force directed layout, as I show in Figure 44.17(a). However, as rected layout. (b) Manually
I mentioned previously, the force directed layouts have this tendency adjusted linear force directed.
of placing your node in a circle. This happens to work really poorly The node color is a product’s
in the case of the Product Space. The reason is that not all prod- Leamer category.
ucts are the same. Some products are harder to export than others.
This is a key concept in the original research, known as Economic
Complexity.
This means that the Product Space has an inherent “direction”.
Countries want to move from simple to more complex products, as
the latter is a more rewarding category to be able to export. However,
the circle has no direction. It is a loop: you always get back to where
you started. The shape of the Product Space in Figure 44.17(a) does
not allow us to perceive the development path. That is why it is 28
Eerily looking like an angel from
necessary to stretch out the visualization as I do in Figure 44.17(b): Neon Genesis Evangelion, with that
creepy head with multiple green eyes...
now the Product Space is a (complex, multidimensional) line28 Am I the only one seeing it?
and you can see that there is a clear direction going from right to
left, from less to more complex products. It is still a type of force
directed layout, but it needed to be customized to remove its inherent
circularity.
620 the atlas for the aspiring network scientist
Cathedral
29
Stephen Kosack, Michele Coscia,
In another paper of mine, I analyze government networks29 . My Evann Smith, Kim Albrecht, Albert-
nodes are government agencies and I establish edges between them László Barabási, and Ricardo Haus-
mann. Functional structures of us state
if the website of an agency has an hyperlink pointing to the website governments. Proceedings of the National
of another agency. One key question is verifying if this network has Academy of Sciences, 115(46):11748–11753,
2018
a hierarchical organization – see Chapter 29. One obvious way to
explore this question is visualizing the network and see if it looks like
a hierarchy. Unfortunately, the network is relatively large and dense.
So I need to come up with a custom layout. Such layout is useful to
visualize dense hierrchical networks, and thus can be considered as
an enhancement of the classical hierarchical layout presented earlier,
that works only for tree-like structures.
The first step is to group nodes into a 2-level functional classifica-
tion. This means to assign to each agency the function it performs
in the government. For instance, a school is part of the education
system (level 1 function) and of the primary & secondary education
(level 2 function). Or: a city government is part of general adminis-
tration (level 1 function) and of the municipal administration (level 2
function).
There are not many level 2 functions so I can collapse all agencies
into their level 2 function. Then I display these functions in a scatter
44.6 Summary
2. The most common principle is the one of the force directed layout.
Nodes are charges of the same sign repelling each other and edges
are springs trying to keep connected nodes together. This layout
works for sparse networks with communities, whose topology fits
on a circle.
6. In many cases, your network will have a clear and unique message
that has never been visualized before. In those cases, you need to
bend rules and create a unique visualization serving your specific
communication objective.
44.7 Exercises
Useful Resources
45
Network Science Applications
1012 103
Figure 45.1: (a) Total wage sum
1011
(y axis) as a function of a city’s
# Gas Stations
2
10
Total Wage
(a) (b)
7
Hyejin Youn, Deborah Strumsky,
already present so far. This is easy to see especially in patents data7 . Luis MA Bettencourt, and José Lobo.
Every time someone makes a new invention, that new invention Invention as a combinatorial process:
evidence from us patents. Journal of The
can be combined with all the previous inventions to create a new
Royal Society Interface, 12(106):20150272,
one, and so on at infinity. Thus, the knowledge added by a new 2015
invention potentially multiplies itself with the previously accumulated
knowledge, rather than just adding to it.
4 5
10
Shirin Nilizadeh, Apu Kapadia, and
in Figure 45.3. The node 1 in the center of this network might be Yong-Yeol Ahn. Community-enhanced
you. If an attacker has identified some of your connections, they de-anonymization of online social
can say a lot about you10 : the only node in this network that has networks. In Proceedings of the 2014
acm sigsac conference on computer and
exactly {2, 3, 4, 5} as the set of their friends. They might not be able to communications security, pages 537–548,
know your name, but under the assumption of homophily (Chapter 2014
26) they could infer what you like, your sexual orientation, and
11
Elena Zheleva and Lise Getoor. To join
or not to join: the illusion of privacy in
maybe even health issues. This is not solved by making your profile social networks with mixed public and
private11 . In fact, even not having a profile at all on a social media private user profiles. In Proceedings of
the 18th international conference on World
won’t make you safe: the platform can always create a shadow profile wide web, pages 531–540, 2009
of you12 . 12
David Garcia. Leaking privacy
De-anonymizing social networks13 is feasible, under a wide array and shadow profiles in online social
networks. Science advances, 3(8):e1701172,
of different scenarios – whether the attacker is a government agency, 2017
a marketing campaign, or an individual stalker. This is usually done 13
Arvind Narayanan and Vitaly
by creating a certain amount of auxiliary information that can then Shmatikov. De-anonymizing social
networks. In 2009 30th IEEE symposium
be used to recursively de-anonymize more and more nodes in the on security and privacy, pages 173–187.
network. Counter-measures usually adopt the k-anonymity style: IEEE, 2009
making sure that no individual can be identified by obfuscating 14
Kun Liu and Evimaria Terzi. Towards
identity anonymization on graphs. In
enough data to make at least k − 1 other individuals identical to her in Proceedings of the 2008 ACM SIGMOD
some respect. For instance, a network is k-degree anonymous if there international conference on Management of
data, pages 93–106, 2008
are at least k nodes with any given degree value14 .
15
Elena Zheleva and Lise Getoor.
Sometimes, the focus is preventing the disclosure of information Preserving the privacy of sensitive rela-
about a relationship, i.e. to combat link re-identification15 . You might tionships in graph data. In International
not want Facebook to know you are friend with someone, which Workshop on Privacy, Security, and Trust
in KDD, pages 153–171. Springer, 2007
they could do by performing some relatively trivial link prediction 16
JW Scannell, GAPC Burns, CC Hilge-
– see Part VI. In those cases, you might want to hide some of your tag, MA O’Neil, and Malcolm P Young.
relationships, and/or add a few fake connections, to throw off the The connectional organization of the
cortico-thalamic system of the cat.
score function of the link you want to hide. Cerebral Cortex, 9(3):277–299, 1999
17
Quanxin Wang, Olaf Sporns, and
Andreas Burkhalter. Network analysis
45.3 Human Connectome of corticocortical connections reveals
ventral and dorsal processing streams
in mouse visual cortex. Journal of
Quite likely, the most famous and studied network in human history Neuroscience, 32(13):4386–4399, 2012b
is the brain. We have been studying neural networks of many ani- 18
Siming Li, Christopher M Armstrong,
mals, due to their limited size and ease of analysis: cats16 , mices17 , Nicolas Bertin, Hui Ge, Stuart Milstein,
Mike Boxem, Pierre-Olivier Vidalain,
and, of course, the superstar C. Elegans worm18 . However, most of Jing-Dong J Han, Alban Chesneau,
this is done with the big prize as the ultimate objective: the human Tong Hao, et al. A map of the inter-
brain. You might have heard of the Human Connectome Project. Pro- actome network of the metazoan c.
elegans. Science, 303(5657):540–543, 2004
posed in 200519 , its objective was to create a low-level network map 19
Olaf Sporns, Giulio Tononi, and Rolf
of the human brain: a network where nodes are individual neurons Kötter. The human connectome: a
and connections are the synapses between them. structural description of the human
brain. PLoS computational biology, 1(4),
The idea was that applying all the network science artillery to 2005
such a network would help us understanding better how our brains 20
Ed Bullmore and Olaf Sporns. Com-
work20 – or don’t, sometimes. In fact, one of the major lines of re- plex brain networks: graph theoretical
analysis of structural and functional
search is comparing the brain connection patterns between healthy systems. Nature reviews neuroscience, 10
and unhealthy individuals, because network analysis should be (3):186–198, 2009
628 the atlas for the aspiring network scientist
a hierarchical fashion: neurons are part of modules25 , and there are Bassett. Multi-scale brain networks.
Neuroimage, 160:73–83, 2017
modules of modules, and so on – check out Chapters 29 and 33 for 25
Paolo Bonifazi, Miri Goldin, Michel A
a few refreshers on hierarchies. In fact, one of the most appropriate Picardo, Isabel Jorquera, A Cattani,
models of the brain is multilayer networks26 . Gregory Bianconi, Alfonso Represa,
Yehezkel Ben-Ari, and Rosa Cossart.
Gabaergic hub neurons orchestrate
synchrony in developing hippocampal
45.4 Science of Science and of Success networks. Science, 326(5958):1419–1424,
2009
Unsurprisingly, one of the things that interests scientists the most 26
Manlio De Domenico. Multilayer
modeling and analysis of human brain
is... scientists. Network scientists are no exception to this rule. There networks. Giga Science, 6(5):gix004, 2017
is a large and healthy literature in analyzing networks of scientists. 27
According to scientists.
We already saw many examples of two types of science networks: 28
Albert-László Barabási, Chaoming
co-authorship networks, where scientists are connected to each other Song, and Dashun Wang. Publishing:
if they collaborate on the same paper/project; and citation networks, Handful of papers dominates citation.
Nature, 491(7422):40, 2012
connecting papers if one cites another. 29
Dashun Wang, Chaoming Song, and
The two can be combined to try and gather a general picture of Albert-László Barabási. Quantifying
how science gets done. Science is one of the most important human long-term scientific impact. Science, 342
(6154):127–132, 2013
activities27 , because we rely on it to develop new and better ways to 30
Santo Fortunato, Carl T Bergstrom,
improve our everyday life. It’s better to understand how it works, so Katy Börner, James A Evans, Dirk
that we can do it better. This is fundamentally the mission statement Helbing, Staša Milojević, Alexander M
Petersen, Filippo Radicchi, Roberta
of the science of science field28 , 29 , 30 , kickstarted by network scientists Sinatra, Brian Uzzi, et al. Science of
and making extensive use of network analysis tools. science. Science, 359(6379):eaao0185, 2018
network science applications 629
103 103
Figure 45.5: Two examples of
10
2
10
2 career paths of scientists, show-
# Citations
# Citations
(a) (b)
PDF
the distance to the destination
(x axis). (b) Probability of a
trip (y axis) as a function of
the destination’s rank (x axis).
Distance Rank In both cases, different colors
(a) (b) report data from different cities.
57
James P Gleeson, Jonathan A Ward,
Kevin P O’sullivan, and William T Lee.
Competition-induced criticality in a
model of meme popularity. Physical
review letters, 112(4):048701, 2014
More complex topological features, such as communities, are 58
Bjarke Mønsted, Piotr Sapieżyński,
difficult to treat mathematically, but their impact can be studied Emilio Ferrara, and Sune Lehmann.
using real world data. Figure 45.7 shows a toy example of the role of Evidence of complex contagion of infor-
mation in social media: An experiment
communities in meme propagation. Memes originating in the overlap using twitter bots. PloS one, 12(9), 2017
between different communities – in red in Figure 45.7 – have a better 59
Lilian Weng, Filippo Menczer, and
chance to go viral59 , 60 . Being born well embedded in a community Yong-Yeol Ahn. Virality prediction and
community structure in social networks.
– blue in Figure 45.7 – is bad for propagation, because there are not Scientific reports, 3:2522, 2013
many paths leading the meme outside of the community. 60
Lilian Weng, Filippo Menczer, and
In general, there are many empirical studies investigating how Yong-Yeol Ahn. Predicting successful
memes using network and community
information propagates through a social network61 , be it memes, structure. In ICWSM, 2014
rumors, news62 , videos63 , 64 , or photographs65 , 66 . 61
Kristina Lerman and Rumi Ghosh.
Specifically, one study focuses on the dynamics of “following” a Information contagion: An empirical
study of the spread of news on digg
content creator on social media67 . The common sense thing is that and twitter social networks. In ICWSM,
the more people are following you – say on Twitter – the better it is. 2010
You can be more influential if more people listen to you. However,
62
Soroush Vosoughi, Deb Roy, and
Sinan Aral. The spread of true and
experiments show that this is true only to a certain point. What false news online. Science, 359(6380):
matter most is the engagement of the followers. Just increasing the 1146–1151, 2018
network science applications 633
104
103
Total Deaths
Rome to Paris and then to New York, because more and more people
die in the city where they work – and notable people work where
most notable people are. There are other interesting patterns, for in-
stance the fact that the median distance between the birth and death
place is increasing, reflecting technological advancements.
With a similar dataset, researchers built the “notable people portfo- 72
Amy Zhao Yu, Shahar Ronen, Kevin
lio” of cities and nations72 . The idea is to classify all famous people Hu, Tiffany Lu, and César A Hidalgo.
in the area they contributed the most to humanity. Then, one can Pantheon 1.0, a manually verified
dataset of globally famous biographies.
visualize in which areas places specialize73 . For instance, the largest
Scientific data, 3:150075, 2016
profession represented in the United States is actors, while it is 73
[Link]
politicians for Greece. But one could explore other dimensions. For
instance, professions that are over-expressed in a country against the
rest of the world, like chess players in Armenia (6% of all famous
people!). Or explore gender divide: in Canada 27.4% of male famous
people were actors against 55.4% female famous people. Finally, you
can explore time as well. Before 1700 AD, the most common way
to become famous in Italy was to have a career in politics (30.7% of 74
Tom Brughmans. Thinking through
networks: a review of formal network
famous people did). Afterward? You’re better off trying as a soccer methods in archaeology. Journal of
player (21.6%). Archaeological Method and Theory, 20(4):
Other digital humanities applications of network science involve 623–662, 2013
75
Tom Brughmans. Connecting the
archaeology. This mostly involves the use of network visualization dots: towards archaeological network
techniques to make sense of a complex, interconnected, and often analysis. Oxford Journal of Archaeology, 29
largely incomplete set of evidence74 . However, it is not necessary (3):277–303, 2010
76
Barbara J Mills, Jeffery J Clark,
to limit ourselves to this: network analysis can be used as a tool to Matthew A Peeples, W Randall Haas,
explore evidence. For instance, there are studies of social networks John M Roberts, J Brett Hill, Deborah L
in classical Rome75 . Departing from archaeology, the field of social Huntley, Lewis Borck, Ronald L Breiger,
Aaron Clauset, et al. Transformation of
network analysis in a historic76 , religious77 , or anthropological78 social networks in the late pre-hispanic
setting is alive and well. us southwest. Proceedings of the National
Academy of Sciences, 110(15):5785–5790,
And since network analysis endows us with powerful tools to 2013
study hidden preferences – such as homophily and segregation, see 77
Eleanor A Power. Discerning devotion:
Chapter 26 – it is a natural instrument to use in other humanities Testing the signaling theory of religion.
Evolution and Human Behavior, 38(1):
fields, such as gender studies. In particular, there are studies showing 82–91, 2017
unequal gender dynamics when it comes to power relations in online 78
Jessica C Flack, Michelle Girvan,
collaborative tools such as, e.g., Wikipedia. As you might expect, the Frans BM De Waal, and David C
Krakauer. Policing stabilizes construc-
majority of Wikipedia contributors are white men. tion of social niches in primates. Nature,
When it comes to female representation in the content79 , 80 , this 439(7075):426–429, 2006
gender gap shows. The researchers find that women are equally 79
Claudia Wagner, David Garcia,
Mohsen Jadidi, and Markus Strohmaier.
represented in article numbers – at least in the main six language It’s a man’s wikipedia? assessing
editions of Wikipedia. However, it is the way women are portrayed gender inequality in an online ency-
clopedia. In Ninth international AAAI
that is the problem. Women on Wikipedia tend to be more linked to
conference on web and social media, 2015
men than vice versa. Moreover, romantic relationships and family- 80
Claudia Wagner, Eduardo Graells-
related issues are much more frequently discussed on Wikipedia Garrido, David Garcia, and Filippo
articles about women than men. Menczer. Women through the glass ceil-
ing: gender asymmetries in wikipedia.
EPJ Data Science, 5(1):5, 2016
46
Data & Tools
46.1 Libraries
Networkx
1
Aric Hagberg, Pieter Swart, and Daniel
I start by dealing with Networkx1 , 2 . Networkx is a Python library S Chult. Exploring network structure,
implementing a vast array of network algorithms and analyses. Net- dynamics, and function using networkx.
Technical report, Los Alamos National
workx is – as far as I can tell – the most popular choice for students Lab.(LANL), Los Alamos, NM (United
approaching network analysis tasks. It isn’t particularly good at States), 2008
2
anything, and it’s by far the library struggling with computational [Link]
efficiency the most, but it is the most complete and popular. If this
were a chapter about cinema, Networkx would be Steven Spielberg:
everybody knows him, every movie he makes is good but not really
great, and it’s the most boring possible choice as a favorite director.
3
10
iGraph Figure 46.1: The running times
Networkx
2 for different implementations
Running Time (s)
10
of community discovery algo-
101 rithms in Networkx (red) and
iGraph (blue).
100
10-1
LPA FG
Algorithm
Graph Tool
but the ones that are there benefit from having a single mind behind
them. The aforementioned issue I had with the bugs in labeled
multigraph isomorphism was solved by simply using Graph Tool.
Moreover, getting into Tiago’s frame of mind is necessary to use
Graph Tool. You need to understand the way he does things in order
to be able to do them as well. Things like function naming, object
types, parameter passing – what you call the interface of the library –
are not as pythonic and intuitive as in Networkx.
iGraph
6
Gabor Csardi and Tamas Nepusz. The
Among all the alternatives, iGraph6 , 7 is certainly the most versatile igraph software package for complex
tool. It combines the strengths – and weaknesses – of Networkx and network research. InterJournal, Complex
Systems, 1695(5):1–9, 2006
Graph Tool. On the one hand, it is a surprisingly complete tool with 7
[Link]
lots of implemented functions – just like Networkx –, and it is pretty
efficiently written – like Graph Tool. Other advantages reside in the
fact that the library is available on a vast array of platforms: you can
use it both in Python and in R. You can even import it directly as a
C library. Thus, if you are capable of writing in C, you can probably
cook up a customized analysis using the power of iGraph that cannot
data & tools 639
Gee, thank you, it’s refreshing to see that not even who developed
this function knows what the function is doing. Continuing my
movie directors analogy, iGraph is David Lynch: probably the only
one able to do what he is doing, but good luck knowing what’s going
on when you look at something made by him.
The fact that I’m badmouthing iGraph so hard should really
convince you that it is a fundamental tool. I hate it with a passion
and yet I am including it in the book and I use it. If I could live
without it – trust me – I would. But I can’t, because sometimes it is
the only thing that will save you.
Other
The libraries I presented in this section are the three horsemen of
network analysis: most of the times, if you need a library, one of
them will cover you. That is not to say they are the only things you
should know. Here I group a bunch of miscellanea that could come
in handy sooner of later. 8
Jure Leskovec and Rok Sosič. Snap: A
There are a few competing libraries for doing network analysis. general-purpose network analysis and
graph-mining library. ACM Transactions
Stanford Network Analysis Project8 , 9 (SNAP) comes to mind. This on Intelligent Systems and Technology
is another C++ library, thus it competes more directly with iGraph. (TIST), 8(1):1–20, 2016
It is owned by a research group at Stanford, which is both a blessing 9
[Link]
640 the atlas for the aspiring network scientist
46.2 Software
For Visualization 20
Paul Shannon, Andrew Markiel,
Owen Ozier, Nitin S Baliga, Jonathan T
By far, the software I use the most is Cytoscape20 , 21 . Mostly, I use it Wang, Daniel Ramage, Nada Amin,
for network visualization. The visual style of Cytoscape is based on Benno Schwikowski, and Trey Ideker.
the Protovis Java library22 , which is one of the ancestors of D323 , 24 – Cytoscape: a software environment
for integrated models of biomolecular
and it shows. The visual style of Cytoscape is really good, and you interaction networks. Genome research, 13
can customize a large quantity of visual attributes relatively easily. (11):2498–2504, 2003
21
[Link]
Cytoscape supports some basic network analysis. You can calcu- 22
Michael Bostock and Jeffrey Heer.
late a bunch of node, edge, and network statistics, the ones you’d Protovis: A graphical toolkit for visual-
come up first in your exploratory data analysis phase – nothing too ization. IEEE transactions on visualization
and computer graphics, 15(6):1121–1128,
fancy. This analytic capability is mostly there only to allow you to use
2009
node and edge statistical properties to augment your visualization. 23
Michael Bostock, Vadim Ogievetsky,
Since version 3.8, Cytoscape shows fewer plots. For instance you and Jeffrey Heer. D3 data-driven
documents. IEEE transactions on
cannot see any more the distribution of shortest path lengths. More-
visualization and computer graphics, 17
over, I can’t seem to be able to show them in a log-log scale – it was (12):2301–2309, 2011
possible before. However, the graphical quality of the plots greatly 24
[Link]
more general tool, which allows you to do more than what this book
focuses on.
Secondarily, you can use NetLogo for visualizing the effects of
specific network processes. If you follow the link I provide, you can
data & tools 643
For Analysis
There are many pieces of software out there that will allow you
to perform network analysis and are commonly used by network
professionals. They are far more than I can include here. So I will
limit myself to those with which I had some personal experience.
The programs I talk about here are the ones that primarily pro-
vide analytic power. You can visualize networks with them, but
you should not do that. Their visualization capabilities are not the
main focus of the software, and are there mostly for you to get a
quick sense of what sort of analyses you should ask the program to
perform.
I think the program for network analysis I stumble the most upon 33
Wouter De Nooy, Andrej Mrvar, and
Vladimir Batagelj. Exploratory social
in the literature is Pajek33 , 34 . Pajek allows you to perform a vast array network analysis with Pajek. Cambridge
of network analysis, ranging from classical social science ones, to University Press, 2018
more computer science-y ones – like community discovery. Pajek 34
[Link]
networks/pajek/
comes in different versions: Pajek, Pajek XXL, and Pajek 3XL. The
main difference between the versions is the capability of handling
larger and larger networks. The idea is that you would perform
the memory-intense analyses on the XL versions of Pajek and then
import the results for further investigation in the standard version of
the program.
Pajek is such a popular program that its own specific file format
is compatible with most of the software libraries I mentioned earlier.
Both Networkx and iGraph have functions that will allow you to
import networks saved in Pajek’s file format. Pajek is Lars von Trier:
perfect for geeking out every possible detail, but not the prettiest
thing to look at. 35
Stephen P Borgatti, Martin G Everett,
and Linton C Freeman. Ucinet for
A popular alternative from Pajek is UCINET35 , 36 . UCINET’s windows: Software for social network
strength is in its deep dive into the social branch of social network analysis. 2002
analysis. It is possibly the most comprehensive tool for social scien- 36
[Link]
ucinetsoftware/home
tists to use.
As a result, its coverage of the more computer science and physics
branches is less than optimal. UCINET works best with small net-
works, it is not particularly well optimized for large scale analysis,
and will lack some of the typical algorithms you might expect to find
after reading this book. However, my biggest gripe with it is prob-
644 the atlas for the aspiring network scientist
Community Discovery
I create this special subsection to focus exclusively of implementa-
tions of algorithms solving the community discovery problem. This is
easily the largest subfield of network analysis. Thus this subsection
satisfies two needs. First, it gives you an idea about the immense
wealth of code that cannot find space in generic libraries/software.
Second, it contains the necessary references to the algorithms I con-
sider in my algorithm similarity network that was included in Section
31.6.
The way this subsection works is as follows. Now I will list a
bunch of labels that are consistent with Figure 31.15. For each label,
I tell you where to find the implementation I used to build that
figure. The general disclaimer is that, of course, some of these links
are bound to break in the future. I accessed them last time around
November 2018, so the Internet Archive could help.
• mcl: [Link]
• ganxis: [Link]
• conclude: [Link]
data & tools 645
• mlrmcl: [Link]
• metis: [Link]
• pmm: [Link]
• crossass: [Link]
[Link].
• demon: [Link]
• hlc: [Link]
• tiles: [Link]
• oslom: [Link]
software.
• code-dense: [Link]
s10618-014-0373-y.
• gce: [Link]
• ilcd: [Link]
• bnmtf: [Link]
• rmcl: [Link]
[Link].
• OLC: [Link]
zip.
39
Santo Fortunato, Vito Latora, and
• edgeclust: [Link] Massimo Marchiori. Method to
radetal_algorithm.tgz. find community structures based on
information centrality. Physical review E,
• infocentr: my own implementation of the algorithm described in 70(5):056104, 2004
40
Anand Narasimhamurthy, Derek
the original paper39 .
Greene, Neil Hurley, and Pádraig
• msg, vm: [Link] Cunningham. Community finding in
large social networks through problem
• ocg: [Link] decomposition. In Proc. 19th Irish
Conference on Artificial Intelligence and
• savi: [Link] Cognitive Science, AICS, volume 8, 2008
• mixnet: [Link]
mixnet. 41
E Gabasova. The star wars so-
• vbmod: [Link] cial network. Evelina Gabasova’s
Blog. Data available at: [Link]
• bridgebound, bagrowLocal, clausetLocal, lwplocal: https:// com/evelinag/StarWars-social-
network/tree/master/networks, 2015
[Link]/kleinmind/bridge-bounding.
42
Tom AB Snijders, Gerhard G Van de
• netcarto: [Link] Bunt, and Christian EG Steglich. In-
troduction to stochastic actor-based
network-cartography-netcarto/.
models for network dynamics. Social
• infomap, infomap-overlap: [Link] networks, 32(1):44–60, 2010
html#Download-and-compile.
43
Gerhard G Van de Bunt, Marijtje AJ
Van Duijn, and Tom AB Snijders.
• graclus: [Link] Friendship networks through time:
An actor-oriented dynamic statistical
[Link].
network model. Computational &
• graclus2stage: my own implementation of the algorithm described Mathematical Organization Theory, 5(2):
167–192, 1999
in the original paper40 . 44
CJ Rhodes and P Jones. Inferring
• fuzzyclust: [Link] missing links in partially observed
social networks. In OR, Defence and
• linecomms: [Link] (to Security, pages 256–271. Springer, 2015
generate the line graph + igraph’s implementation of Louvain).
45
Karine Descormiers and Carlo
Morselli. Alliances, conflicts, and
contradictions in montreal’s street gang
It’s now time to move to a section dedicated not to code, but to
landscape. International Criminal Justice
data. However, before doing so, I’ll give you a small preview. In Review, 21(3):297–314, 2011
the paper building the algorithm similarity network, I test all these 46
Siva R Sundaresan, Ilya R Fischhoff,
Jonathan Dushoff, and Daniel I Ruben-
algorithms on 819 real world networks that I use as a benchmark.
stein. Network metrics reveal differences
These networks are taken from data kindly shared by the authors in social organization between two
of a bunch of papers41 , 42 , 43 , 44 , 45 , 46 , 47 , 48 , 49 , 50 , 51 , 52 and online web- fission–fusion species, grevy’s zebra
and onager. Oecologia, 151(1):140–149,
pages53 , 54 . 2007
47
Martin W Schein and Milton H
Fohrman. Social dominance relation-
46.3 Data ships in a herd of dairy cattle. The British
Journal of Animal Behaviour, 3(2):45–55,
There’s more to life than just the software running on your computer 1955
48
A Gimenez-Salinas Framis. Illegal
– or so I’m told. Many great resources that can make you a better
networks or criminal organizations:
network analyst – or even just a better person overall – can be found Power, roles and facilitators in four
online. Specifically, here I focus on online network resources concern- cocaine trafficking structures. In Third
Annual Illicit Networks Workshop, 2011
ing the first ingredient of every network paper: network data. You 49
Andrew Beveridge and Jie Shan.
need data to test your models, to run your algorithm, to make a com- Network of thrones. Math Horizons, 23
pelling case of why your paper is important. Often, your study will (4):18–22, 2016
data & tools 647
start from a dataset you already have, or you will collect one specially 50
Wouter De Nooy. A literary play-
tailored for your purposes. In many other cases, you simply need ground: Literary criticism and balance
theory. Poetics, 26(5-6):385–404, 1999
any network you can put your hands on that fulfills some specific 51
Dale F Lott. Dominance relations and
constraints. This section should help you with this task. breeding rate in mature male american
There are many places where you can find networks directly bison. Zeitschrift für Tierpsychologie, 49(4):
418–432, 1979
available for download, but I start with an index: the Colorado 52
Jermain Kaminski, Michael Schober,
Index of Complex Networks55 , 56 (ICON). This is quite possibly the Raymond Albaladejo, Oleksandr Zastu-
most comprehensive index of network datasets from all domains of pailo, and Cesar Hidalgo. Moviegalaxies-
social networks in movies. 2018
network science. Chances are that, if the network data is available 53
[Link]
somewhere, you can find it via ICON. networks/data/bio/foodweb/foodweb.
htm
However, this is an index of network datasets, not a dataset repos- 54
[Link]
itory, like the ones that will follow. This means that ICON is not [Link]
hosting any network data itself. It rather contains the links to those
55
A Clauset, E Tucker, and M Sainz. The
colorado index of complex networks,
datasets. This has advantages and disadvantages. The advantage 2016
is completeness: not all datasets can be moved from their original 56
[Link]
source and hosted somewhere else. ICON can include those datasets,
while the other repositories cannot. The other side of the coin is the
dynamism of the Internet. Resources get moved all the time, and not
everybody does it properly via HTTP redirects – actually almost no
one does it. Thus it is possible to find dead links in ICON, because
the managers of the website cannot possibly constantly check that all
links are working.
ICON will point you to tons of resources from which you can
actually download your data, for instance Pajek’s and UCINET’s
websites. There you can find a collection of network datasets you can
download, which is a nice additional resource to the software. One
issue you might have with this solution is that they distribute data in
their own file formats, so you might need to convert them before you
can use them with another software.
Also SNAP provides network data. While there is a large overlap
between what you can find in Pajek and UCINET, SNAP’s focus
goes decisively more towards computer science. You will find very
large datasets there, sometimes larger than what you can handle –
at the time of writing this chapter, I believe the largest network is
from Friendster, which contains more than 1.8 billion edges. Just like
with the implemented functions in SNAP, also the datasets are very
much focused on the ones the Stanford research group used for their
publications. 57
Jérôme Kunegis. Konect: the koblenz
Another interesting resource is Konect57 , 58 . Konect is also a Mat- network collection. In Proceedings of the
lab package for network analysis. Since I do not use Matlab unless 22nd International Conference on World
Wide Web, pages 1343–1350, 2013
someone is pointing a gun at me, I have no experience with it as 58
[Link]
an analysis tool. However, I used to browse Konect daily to find
and download some interesting network data. The list of available
datasets, as far as I can tell, is a superset of what you can find in the
648 the atlas for the aspiring network scientist
websites of Pajek and UCINET, and more. The web interface is also
well done, and you will be able to tell what are the main character-
istics of a network before downloading it: if it’s bipartite, what its
degree distribution is, etc. I said “used to browse” because recently
Konect moved to a subscription system – it’s not free any more. The
price tag is not something than an individual can afford without an
organization with deep pockets backing them, so it’s a non-starter for
independents. 59
[Link]
Network Repository59 is probably the largest repository with di-
rect network data download – thus excluding ICON. The interface
isn’t as good as Konect, but it’s free and it includes more network
data. Tiago Peixoto of Graph Tool fame also launched his own net- 60
[Link]
work data resource: Netzschleuder60 , which rivals Network Repos-
itory in size. It mostly takes all the network data indexed by ICON
or Konect and provides a direct download in different formats. The
networks can also be directly imported in Graph Tool via a function,
without worrying about having the network file saved on your hard
disk. The interface is snappier although it could do a better job to
highlight where the direct download buttons are.
If you’re specifically interested in multilayer network data, one 61
[Link]
cool resource is the Comune Lab61 . Comune is a project owned by
the same people behind Muxviz and the two can be considered as
closely integrated.
There are some graphs that are so widely used that you don’t really
need to look for them in an online repository. These are the pillars
on which the entire cathedral of network science is founded. They
are often directly included in software and libraries, and all online
self-respecting network data repositories have one or multiple copies
of them. I include a few here.
The first – and by far most popular – of these legendary graphs is 62
Wayne W Zachary. An information
the Zachary Karate Club62 . This is a network of members of a karate flow model for conflict and fission in
club, connecting two members if they sparred against each other. It small groups. Journal of anthropological
research, 33(4):452–473, 1977
is often used because the network focuses on two main nodes: the
coach and the president of the club. The club eventually split due to
a disagreement between the two, and one can reconstruct on which
side each member went by analyzing with whom they sparred. It is
a classical example of community discovery. Figure 46.3 shows this
beauty in all of its glory.
Aaron Clauset told me a fun fact about this network. Zachary’s
original paper contains a figure showing the undirected adjacency
matrix of the Karate Club network, except that it’s not fully undi-
data & tools 649
rected! One edge appears in one direction, but not in the other. This
means that there are technically two Karate Club graphs, depending
on whether this edge is a typo or not, one with | E| = 77 edges and
one with | E| = 78 edges. The latter is the most common you’ll find
around, because it is the one that Mark Newman and Michelle Gir-
van used for their paper, which arguably launched the Karate Club
network in the Olympus of network science.
Network scientists are obsessed with this network. It has its own 63
[Link]
t-shirt63 . They even created the Zachary Karate Club Club64 : the club zachary_karate_club_with_label_
of network scientists who are the first using the Zachary network as t_shirt-235415254499870147, the label
says: “If your method doesn’t work on
an example in their presentation at a network science conference. If
this network, then go home”.
you do so, you become the current holder of the Zachary Karate Club 64
[Link]
Trophy and you are responsible for handing it at the next conference
you attend. This is fiercely competitive, and often you’ll see this prize
awarded at satellites events happening before the conference itself,
because people will use the network as an example as soon as they
can, to get their hands on the trophy.
The Network Science Society hands many prestigious awards: the 65
[Link]
Erdős-Rényi prize65 , to the career of the most outstanding network award-prizes/er-prize
scientist under the age of forty; or the Euler award66 , to the authors 66
[Link]
award-prizes/euler-award
of paradigm-changing publications in network science. But don’t get
fooled. The Zachary Karate Club Trophy is where it’s at.
Another commonly used network is the one obtained from Victor 67
Donald Ervin Knuth. The Stanford
Hugo’s novel Les Miserables67 . In the network, each node is a char- GraphBase: a platform for combinatorial
acter, and two characters are connected together if they appear in computing. AcM Press New York, 1993
the same chapter. Also in this case the classical application is for
community discovery, given that there are sets of characters closely
interacting with each other that never appear in chapters with other
groups of characters. Figure 46.4 shows an example. This is one of
those graphs that even non-network scientist would use for examples
650 the atlas for the aspiring network scientist
68
[Link]
related to other fields, for instance data visualization68 . The likely miserables/
reason is the inclusion of this network in Knuth’s popular book. 69
[Link]
The college football network69 is another network commonly American_College_Football_Network_
used for community discovery – I’m sensing a pattern here. Figure Files/93179
46.5 shows it. The reason it works well is due to the way sports
are organized in the United States. Usually, teams are divided in
conferences and divisions. A team will play with all other teams
in their division, but only with a selected number of teams in the
same conference and almost no team from the other conference. This
creates a nice hierarchical community structure. There is also an
overlap, as the most successful teams will then access to the finals
and thus play a significant number of matches with teams from the
other conference.
Average Path Length: The sum of the lengths of all shortest paths in a
network over the total number of such paths.
glossary 653
Chain: A set of nodes that can be ordered, and each node is con-
nected only to its predecessor – except the first node – and its succes-
sor – except the last node.
Clique: A set of nodes where all possible edges are present, i.e. each
node in a clique is connected with each other node in the same clique.
Cycle: A path in which the starting and ending node is the same.
Hyperedges: Edges that can connect more than two nodes at the
same time.
Identity Matrix: A matrix with ones on the diagonal and zeros every-
where else.
k-core: Set of nodes that have a minimum degree of k, once you recur-
sively remove from the network all nodes that have k − 1 connections
or fewer.
Line Graph: The graph that represents the adjacencies between edges
of an undirected graph: each edge of the original graph is a node in
the line graph, and two nodes in the line graph connect if they have a
node in common in the original graph.
Parallel edges: Two (or more) edges established between the same
pair of nodes.
Star: A set of nodes with one acting as a center connected to all other
nodes in the star. All other nodes have only one connection, to the
star’s center.
A: Adjacency matrix.
E: Set of edges.
H: The hitting time matrix, telling you how long it’ll take for a ran-
dom walker to visit one node when starting from another.
L: Laplacian matrix.
V: Set of nodes.
Lada A Adamic and Eytan Adar. Friends and neighbors on the web.
Social networks, 25(3):211–230, 2003.
Arun Advani and Bansi Malde. Empirical methods for networks data:
Social effects, network formation and measurement error. Technical
report, IFS Working Papers, 2014.
William Aiello, Fan Chung, and Linyuan Lu. A random graph model
for massive graphs. In Proceedings of the thirty-second annual ACM
symposium on Theory of computing, pages 171–180. Acm, 2000.
Alex Arenas, Jordi Duch, Alberto Fernández, and Sergio Gómez. Size
reduction of complex networks preserving modularity. New Journal
of Physics, 9(6):176, 2007.
Wirt Atmar and Bruce D Patterson. The measure of order and disor-
der in the distribution of species in fragmented habitat. Oecologia, 96
(3):373–382, 1993.
Boris C Bernhardt, Zhang Chen, Yong He, Alan C Evans, and Neda
Bernasconi. Graph-theoretical analysis reveals disrupted small-world
organization of cortical thickness correlation networks in temporal
lobe epilepsy. Cerebral cortex, 21(9):2147–2157, 2011.
Alessandro Bessi and Emilio Ferrara. Social bots distort the 2016 us
presidential election online discussion. 2016.
Dirk Brockmann, Lars Hufnagel, and Theo Geisel. The scaling laws
of human travel. Nature, 439(7075):462–465, 2006.
Qing Cai, Lijia Ma, Maoguo Gong, and Dayong Tian. A survey on
network community detection based on evolutionary computation.
IJBIC, 8(2):84–98, 2016.
Shaosheng Cao, Wei Lu, and Qiongkai Xu. Grarep: Learning graph
representations with global structural information. In Proceedings of
the 24th ACM international on conference on information and knowledge
management, pages 891–900, 2015.
Shaosheng Cao, Wei Lu, and Qiongkai Xu. Deep neural networks
for learning graph representations. In Thirtieth AAAI conference on
artificial intelligence, 2016.
bibliography 681
Ciro Cattuto, Wouter Van den Broeck, Alain Barrat, Vittoria Colizza,
Jean-François Pinton, and Alessandro Vespignani. Dynamics of
person-to-person interactions from distributed rfid sensor networks.
PloS one, 5(7), 2010.
Shiyu Chang, Wei Han, Jiliang Tang, Guo-Jun Qi, Charu C Aggar-
wal, and Thomas S Huang. Heterogeneous network embedding via
deep architectures. In SIGKDD, pages 119–128, 2015.
Chen Chen, Xifeng Yan, Feida Zhu, Jiawei Han, and S Yu Philip.
Graph olap: Towards online analytical processing on graphs. In
ICDM, pages 103–112. IEEE, 2008.
Haochen Chen, Bryan Perozzi, Yifan Hu, and Steven Skiena. Harp:
Hierarchical representation learning for networks. In Thirty-Second
AAAI Conference on Artificial Intelligence, 2018.
Sung-Bae Cho and Jin H Kim. Multiple network fusion using fuzzy
logic. IEEE Transactions on Neural Networks, 6(2):497–501, 1995.
Fan Chung and Linyuan Lu. The average distances in random graphs
with given expected degrees. Proceedings of the National Academy of
Sciences, 99(25):15879–15882, 2002a.
Gabor Csardi and Tamas Nepusz. The igraph software package for
complex network research. InterJournal, Complex Systems, 1695(5):1–9,
2006.
Fabio Della Rossa, Fabio Dercole, and Carlo Piccardi. Profiling core-
periphery network structure by random walkers. Scientific reports, 3:
1467, 2013.
Ugur Dogrusoz, Erhan Giral, Ahmet Cetintas, Ali Civril, and Emek
Demir. A layout algorithm for undirected compound graphs.
Information Sciences, 179(7):980–994, 2009.
Yuxiao Dong, Jie Tang, Sen Wu, Jilei Tian, Nitesh V Chawla, Jinghai
Rao, and Huanhuan Cao. Link prediction and recommendation
across heterogeneous social networks. In 2012 IEEE 12th International
conference on data mining, pages 181–190. IEEE, 2012.
Sergey Edunov, Carlos Diuk, Ismail Onur Filiz, Smriti Bhagat, and
Moira Burke. Three and a half degrees of separation. Research at
Facebook, 2016.
694 the atlas for the aspiring network scientist
Paul Erdos and Alfred Renyi. On random matrices. Magyar Tud. Akad.
Mat. Kutató Int. Közl, 8(455-461):1964, 1964.
Scott L Feld. Why your friends have more friends than you do.
American Journal of Sociology, 96(6):1464–1477, 1991.
Ove Frank and David Strauss. Markov graphs. Journal of the american
Statistical association, 81(395):832–842, 1986.
Nir Friedman, Lise Getoor, Daphne Koller, and Avi Pfeffer. Learning
probabilistic relational models. In IJCAI, volume 99, pages 1300–1309,
1999.
Galileo Galilei. Dialogo sopra i due massimi sistemi del mondo, 1632.
Xinbo Gao, Bing Xiao, Dacheng Tao, and Xuelong Li. A survey of
graph edit distance. Pattern Analysis and applications, 13(1):113–129,
2010.
Herbert Gintis. The bounds of reason: Game theory and the unification of
the behavioral sciences. Princeton University Press, 2014.
Chris Godsil and Gordon F Royle. Algebraic graph theory, volume 207.
Springer Science & Business Media, 2013.
Jonathan L Gross and Jay Yellen. Graph theory and its applications.
CRC press, 2005.
Jiawei Han, Jian Pei, and Yiwen Yin. Mining frequent patterns
without candidate generation. In ACM sigmod record, volume 29,
pages 1–12. ACM, 2000.
bibliography 705
Jiawei Han, Hong Cheng, Dong Xin, and Xifeng Yan. Frequent
pattern mining: current status and future directions. Data mining and
knowledge discovery, 15(1):55–86, 2007.
James A Hanley and Barbara J McNeil. The meaning and use of the
area under a receiver operating characteristic (roc) curve. Radiology,
143(1):29–36, 1982.
Petter Holme and Beom Jun Kim. Growing scale-free networks with
tunable clustering. Physical review E, 65(2):026107, 2002.
708 the atlas for the aspiring network scientist
John Hopcroft, Omar Khan, Brian Kulis, and Bart Selman. Tracking
evolving communities in large linked networks. Proceedings of the
National Academy of Sciences, 101(suppl 1):5249–5253, 2004.
Jun Huan, Wei Wang, and Jan Prins. Efficient mining of frequent
subgraphs in the presence of isomorphism. In Third IEEE International
Conference on Data Mining, pages 549–552. IEEE, 2003.
bibliography 709
Jianbin Huang, Heli Sun, Jiawei Han, Hongbo Deng, Yizhou Sun,
and Yaguang Liu. Shrink: a structural clustering algorithm for
detecting hierarchical communities in networks. In Proceedings of
the 19th ACM international conference on Information and knowledge
management, pages 219–228. ACM, 2010.
Jianbin Huang, Heli Sun, Yaguang Liu, Qinbao Song, and Tim
Weninger. Towards online multiresolution community detection in
large-scale networks. PloS one, 6(8):e23829, 2011a.
Tommy R Jensen and Bjarne Toft. Graph coloring problems, volume 39.
John Wiley & Sons, 2011.
Guoliang Ji, Shizhu He, Liheng Xu, Kang Liu, and Jun Zhao. Knowl-
edge graph embedding via dynamic mapping matrix. In IJCNLP,
pages 687–696, 2015.
Kara Joyner and Grace Kao. School racial composition and ado-
lescent racial homophily. Social science quarterly, pages 810–825,
2000.
U Kang, Hanghang Tong, and Jimeng Sun. Fast random walk graph
kernel. In Proceedings of the 2012 SIAM international conference on data
mining, pages 828–838. SIAM, 2012.
Elias Khalil, Hanjun Dai, Yuyu Zhang, Bistra Dilkina, and Le Song.
Learning combinatorial optimization algorithms over graphs. In
Advances in Neural Information Processing Systems, pages 6348–6358,
2017.
Jin Seop Kim, Kwang-Il Goh, Byungnam Kahng, and Doochul Kim.
Fractality and self-similarity in scale-free networks. New Journal of
Physics, 9(6):177, 2007.
Tamara Kolda and Brett Bader. The tophits model for higher-order
web link analysis. In Workshop on link analysis, counterterrorism and
security, volume 7, pages 26–29, 2006.
David Lazer, Alex Sandy Pentland, Lada Adamic, Sinan Aral, Al-
bert Laszlo Barabasi, Devon Brewer, Nicholas Christakis, Noshir
Contractor, James Fowler, Myron Gutmann, et al. Life in the network:
the coming age of computational social science. Science (New York,
NY), 323(5915):721, 2009.
Geng Li, Murat Semerci, Bulent Yener, and Mohammed J Zaki. Graph
classification via topological and label attributes. In Proceedings of the
9th international workshop on mining and learning with graphs (MLG),
San Diego, USA, volume 2, 2011.
Jiaoyang Li, Pavel Surynek, Ariel Felner, Hang Ma, TK Satish Ku-
mar, and Sven Koenig. Multi-agent path finding for large agents. In
Proceedings of the AAAI Conference on Artificial Intelligence, volume 33,
pages 7627–7634, 2019.
Jundong Li, Harsh Dani, Xia Hu, Jiliang Tang, Yi Chang, and Huan
Liu. Attributed network embedding for learning in a dynamic
environment. In Proceedings of the 2017 ACM on Conference on
Information and Knowledge Management, pages 387–396, 2017a.
Menghui Li, Ying Fan, Jiawei Chen, Liang Gao, Zengru Di, and
Jinshan Wu. Weighted networks of scientific communication: the
720 the atlas for the aspiring network scientist
Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. Diffusion convolu-
tional recurrent neural network: Data-driven traffic forecasting. arXiv
preprint arXiv:1707.01926, 2017b.
Yujia Li, Oriol Vinyals, Chris Dyer, Razvan Pascanu, and Peter
Battaglia. Learning deep generative models of graphs. arXiv preprint
arXiv:1803.03324, 2018.
Yankai Lin, Zhiyuan Liu, Maosong Sun, Yang Liu, and Xuan Zhu.
Learning entity and relation embeddings for knowledge graph
completion. In Twenty-ninth AAAI conference on artificial intelligence,
2015.
bibliography 721
Minghua Liu, Hang Ma, Jiaoyang Li, and Sven Koenig. Task and
path planning for multi-agent pickup and delivery. In Int Conf
on Autonomous Agents and MultiAgent Systems, pages 1152–1160.
IFAAMAS, 2019.
Weiping Liu and Linyuan Lü. Link prediction based on local random
walk. EPL (Europhysics Letters), 89(5):58007, 2010.
Yike Liu, Tara Safavi, Abhilash Dighe, and Danai Koutra. Graph
summarization methods and applications: A survey. ACM Computing
Surveys (CSUR), 51(3):1–34, 2018b.
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor
Darrell. Rethinking the value of network pruning. arXiv preprint
arXiv:1810.05270, 2018c.
Can Lu, Jeffrey Xu Yu, Rong-Hua Li, and Hao Wei. Exploring
hierarchies in online social networks. IEEE Transactions on Knowledge
and Data Engineering, 28(8):2086–2100, 2016.
Hao Ma, Haixuan Yang, Michael R Lyu, and Irwin King. Min-
ing social networks using heat diffusion processes for marketing
candidates selection. In Proceedings of the 17th ACM conference on
Information and knowledge management, pages 233–242. ACM, 2008.
Jan Maas. Gradient flows of the entropy for finite markov chains.
Journal of Functional Analysis, 261(8):2250–2292, 2011.
Laurens van der Maaten and Geoffrey Hinton. Visualizing data using
t-sne. Journal of machine learning research, 9(Nov):2579–2605, 2008.
Matteo Magnani and Luca Rossi. The ml-model for multi-layer social
networks. In ASONAM, pages 5–12. IEEE, 2011.
Elizabeth Aura McClintock. When does race matter? race, sex, and
dating at an elite university. Journal of Marriage and Family, 72(1):
45–72, 2010.
Carl D Meyer. Matrix analysis and applied linear algebra, volume 71.
Siam, 2000.
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient
estimation of word representations in vector space. arXiv preprint
arXiv:1301.3781, 2013a.
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff
Dean. Distributed representations of words and phrases and their
compositionality. Advances in neural information processing systems, 26:
3111–3119, 2013b.
Michael Molloy and Bruce Reed. A critical point for random graphs
with a given degree sequence. Random structures & algorithms, 6(2-3):
161–180, 1995.
Michael Molloy and Bruce Reed. The size of the giant component
of a random graph with a given degree sequence. Combinatorics,
probability and computing, 7(3):295–305, 1998.
Enys Mones, Lilla Vicsek, and Tamás Vicsek. Hierarchy measure for
complex networks. PloS one, 7(3):e33799, 2012.
Jacob Levy Moreno, Helen Hall Jennings, and Ernest Stagg Whitin.
Group method and group psychotherapy. Number 5. Beacon House, 1932.
Siegfried Nijssen and Joost N Kok. The gaston tool for frequent
subgraph mining. Electronic Notes in Theoretical Computer Science, 127
(1):77–87, 2005.
Mingdong Ou, Peng Cui, Jian Pei, Ziwei Zhang, and Wenwu Zhu.
Asymmetric transitivity preserving graph embedding. In SIGKDD,
pages 1105–1114, 2016.
John F Padgett and Christopher K Ansell. Robust action and the rise
of the medici, 1400-1434. American journal of sociology, 98(6):1259–1319,
1993.
Raj Kumar Pan and Jari Saramäki. Path lengths, correlations, and
centrality in temporal networks. Physical Review E, 84(1):016105, 2011.
Shirui Pan, Jia Wu, Xingquan Zhu, Chengqi Zhang, and Yang Wang.
Tri-party deep network representation. Network, 11(9):12, 2016.
732 the atlas for the aspiring network scientist
Luca Pappalardo, Giulio Rossetti, and Dino Pedreschi. " how well
do we know each other?" detecting tie strength in multidimensional
social networks. In 2012 IEEE/ACM International Conference on
Advances in Social Networks Analysis and Mining, pages 1040–1045.
IEEE, 2012.
Namyong Park, Andrey Kan, Xin Luna Dong, Tong Zhao, and
Christos Faloutsos. Estimating node importance in knowledge graphs
using graph neural networks. In Proceedings of the 25th ACM SIGKDD
International Conference on Knowledge Discovery & Data Mining, pages
596–606, 2019.
bibliography 733
Judea Pearl and Dana Mackenzie. The book of why: the new science of
cause and effect. Basic Books, 2018.
Tiago P Peixoto. Efficient monte carlo and greedy heuristic for the
inference of stochastic block models. Physical Review E, 89(1):012804,
2014a.
Ofir Pele and Michael Werman. A linear time histogram metric for
improved sift matching. In European conference on computer vision,
pages 495–508. Springer, 2008.
Ofir Pele and Michael Werman. Fast and robust earth mover’s
distances. In Computer vision, 2009 IEEE 12th international conference
on, pages 460–467. IEEE, 2009.
Davi de Castro Reis, Paulo Braz Golgher, Altigran Soares Silva, and
AlbertoF Laender. Automatic web news extraction using tree edit
distance. In Proceedings of the 13th international conference on World
Wide Web, pages 502–511, 2004.
Neil Shah, Danai Koutra, Tianmin Zou, Brian Gallagher, and Chris-
tos Faloutsos. Timecrunch: Interpretable dynamic graph summariza-
tion. In Proceedings of the 21th ACM SIGKDD International Conference
on Knowledge Discovery and Data Mining, pages 1055–1064, 2015.
Haichuan Shang, Xuemin Lin, Ying Zhang, Jeffrey Xu Yu, and Wei
Wang. Connected substructure similarity search. In Proceedings of the
2010 ACM SIGMOD International Conference on Management of data,
pages 903–914, 2010.
Shai S Shen-Orr, Ron Milo, Shmoolik Mangan, and Uri Alon. Net-
work motifs in the transcriptional regulation network of escherichia
coli. Nature genetics, 31(1):64, 2002.
Jianbo Shi and Jitendra Malik. Normalized cuts and image segmen-
tation. Departmental Papers (CIS), page 107, 2000.
Yu Shi, Qi Zhu, Fang Guo, Chao Zhang, and Jiawei Han. Easing
embedding learning by comprehensive transcription of heteroge-
neous information networks. In Proceedings of the 24th ACM SIGKDD
International Conference on Knowledge Discovery & Data Mining, pages
2190–2199, 2018.
bibliography 743
Ben Shneiderman. The eyes have it: A task by data type taxonomy
for information visualizations. In Proceedings 1996 IEEE symposium
on visual languages, pages 336–343. IEEE, 1996.
Jamie Snape, Jur Van Den Berg, Stephen J Guy, and Dinesh
Manocha. The hybrid reciprocal velocity obstacle. IEEE Transac-
tions on Robotics, 27(4):696–706, 2011.
Olaf Sporns, Giulio Tononi, and Rolf Kötter. The human connectome:
a structural description of the human brain. PLoS computational
biology, 1(4), 2005.
Natalie Stanley, Saray Shai, Dane Taylor, and Peter J Mucha. Cluster-
ing network layers with the strata multilayer stochastic block model.
IEEE transactions on network science and engineering, 3(2):95–105, 2016.
Jimeng Sun, Yinglian Xie, Hui Zhang, and Christos Faloutsos. Less
is more: Compact matrix decomposition for large sparse graphs. In
Proceedings of the 2007 SIAM International Conference on Data Mining,
pages 366–377. SIAM, 2007b.
Yizhou Sun, Jie Tang, Jiawei Han, Manish Gupta, and Bo Zhao. Com-
munity evolution detection in dynamic heterogeneous information
networks. In MLGraphs, pages 137–146. ACM, 2010.
Nassim Nicholas Taleb. The black swan: The impact of the highly
improbable, volume 2. Random house, 2007.
Fei Tan, Yongxiang Xia, and Boyao Zhu. Link prediction in complex
networks: a mutual information perspective. PloS one, 9(9):e107056,
2014.
Jian Tang, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, and
Qiaozhu Mei. Line: Large-scale information network embedding. In
Proceedings of the 24th international conference on world wide web, pages
1067–1077. International World Wide Web Conferences Steering
Committee, 2015a.
Jian Tang, Jingzhou Liu, Ming Zhang, and Qiaozhu Mei. Visualizing
large-scale and high-dimensional data. In WWW, pages 287–297,
2016.
Jiliang Tang, Shiyu Chang, Charu Aggarwal, and Huan Liu. Negative
link prediction in social media. In Proceedings of the eighth ACM
international conference on web search and data mining, pages 87–96.
ACM, 2015b.
748 the atlas for the aspiring network scientist
Lei Tang, Huan Liu, Jianping Zhang, and Zohreh Nazeri. Community
evolution in dynamic multi-mode networks. In SIGKDD, pages
677–685. ACM, 2008.
Lei Tang, Xufei Wang, and Huan Liu. Community detection via
heterogeneous interaction analysis. Data mining and knowledge
discovery, 25(1):1–33, 2012.
Jun Tao, Jian Xu, Chaoli Wang, and Nitesh V Chawla. Honvis:
Visualizing and exploring higher-order networks. In 2017 IEEE Pacific
Visualization Symposium (PacificVis), pages 1–10. IEEE, 2017.
Dane Taylor, Saray Shai, Natalie Stanley, and Peter J Mucha. En-
hanced detectability of community structure in multilayer networks
through layer aggregation. Physical review letters, 116(22):228301, 2016.
Ivan Voitalov, Pim van der Hoorn, Remco van der Hofstad, and
Dmitri Krioukov. Scale-free networks well done. arXiv preprint
arXiv:1811.02071, 2018.
Soroush Vosoughi, Deb Roy, and Sinan Aral. The spread of true and
false news online. Science, 359(6380):1146–1151, 2018.
Daixin Wang, Peng Cui, and Wenwu Zhu. Structural deep network
embedding. In Proceedings of the 22nd ACM SIGKDD international
conference on Knowledge discovery and data mining, pages 1225–1234,
2016a.
752 the atlas for the aspiring network scientist
Liping Wang, Qing Li, Na Li, Guozhu Dong, and Yu Yang. Substruc-
ture similarity measurement in chinese recipes. In Proceedings of the
17th international conference on World Wide Web, pages 979–988, 2008a.
Peng Wang, BaoWen Xu, YuRong Wu, and XiaoYu Zhou. Link
prediction in social networks: the state-of-the-art. Science China
Information Sciences, 58(1):1–38, 2015a.
Zhen Wang, Jianwen Zhang, Jianlin Feng, and Zheng Chen. Knowl-
edge graph embedding by translating on hyperplanes. In Twenty-
Eighth AAAI conference on artificial intelligence, 2014b.
Zhen Wang, Lin Wang, Attila Szolnoki, and Matjaž Perc. Evolu-
tionary games on multilayer networks: a colloquium. The European
physical journal B, 88(5):124, 2015b.
Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial temporal graph
convolutional networks for skeleton-based action recognition. In
Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
Jiaxuan You, Rex Ying, Xiang Ren, William Hamilton, and Jure
Leskovec. Graphrnn: Generating realistic graphs with deep auto-
regressive models. In International Conference on Machine Learning,
pages 5694–5703, 2018.
Amy Zhao Yu, Shahar Ronen, Kevin Hu, Tiffany Lu, and César A
Hidalgo. Pantheon 1.0, a manually verified dataset of globally
famous biographies. Scientific data, 3:150075, 2016.
Kai Yu, Wei Chu, Shipeng Yu, Volker Tresp, and Zhao Xu. Stochastic
relational models for discriminative link prediction. In Advances in
neural information processing systems, pages 1553–1560, 2007.
Hongming Zhang, Liwei Qiu, Lingling Yi, and Yangqiu Song. Scal-
able multiplex network embedding. In IJCAI, volume 18, pages
3082–3088, 2018.
Peng Zhang, Jinliang Wang, Xiaojia Li, Menghui Li, Zengru Di, and
Ying Fan. Clustering coefficient and community structure of bipartite
networks. Physica A: Statistical Mechanics and its Applications, 387(27):
6869–6875, 2008.
Peixiang Zhao, Xiaolei Li, Dong Xin, and Jiawei Han. Graph cube:
on warehousing and olap multidimensional networks. In Proceedings
of the 2011 ACM SIGMOD International Conference on Management of
data, pages 853–864, 2011.
Elena Zheleva and Lise Getoor. To join or not to join: the illusion
of privacy in social networks with mixed public and private user
profiles. In Proceedings of the 18th international conference on World wide
web, pages 531–540, 2009.
Jie Zhou, Ganqu Cui, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu,
and Maosong Sun. Graph neural networks: A review of methods and
applications. arXiv preprint arXiv:1812.08434, 2018.
Tao Zhou, Jie Ren, Matúš Medo, and Yi-Cheng Zhang. Bipartite
network projection and personal recommendation. Physical Review
E, 76(4):046115, 2007.
Tao Zhou, Zoltán Kuscsik, Jian-Guo Liu, Matúš Medo, Joseph Rush-
ton Wakeling, and Yi-Cheng Zhang. Solving the apparent diversity-
accuracy dilemma of recommender systems. Proceedings of the
National Academy of Sciences, 107(10):4511–4515, 2010.
Ezra W Zuckerman and John T Jost. What makes you think you’re so
popular? self-evaluation maintenance and the subjective side of the"
friendship paradox". Social Psychology Quarterly, pages 207–223, 2001.