0% found this document useful (0 votes)
15 views23 pages

Understanding Entropy in Neuroscience

Primer on entropy in neuroscience

Uploaded by

javan.lee15
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views23 pages

Understanding Entropy in Neuroscience

Primer on entropy in neuroscience

Uploaded by

javan.lee15
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A primer on entropy in neuroscience

Erik D. Fagerholm, Zalina Dezhina, Rosalyn J. Moran, Federico E. Turkheimer, Robert Leech

Department of Neuroimaging, King’s College London

Corresponding author: [Link]@[Link]

Abstract

Entropy is not just a property of a system — it is a property of a system and an observer.


Specifically, entropy is a measure of the amount of hidden information in a system that arises
due to an observer's limitations. Here we provide an account of entropy from first principles in
statistical mechanics with the aid of toy models of neural systems. Specifically, we describe
the distinction between micro and macrostates in the context of simplified binary-state neurons
and the characteristics of entropy required to capture an associated measure of hidden
information. We discuss the origin of the mathematical form of entropy via the indistinguishable
re-arrangements of discrete-state neurons and show the way in which the arguments are
extended into a phase space description for continuous large-scale neural systems. Finally,
we show the ways in which limitations in neuroimaging resolution, as represented by coarse
graining operations in phase space, lead to an increase in entropy in time as per the second
law of thermodynamics. It is our hope that this primer will support the increasing number of
studies that use entropy as a way of characterising neuroimaging timeseries and of making
inferences about brain states.

Page 1 of 23
Introduction

The mathematical form of entropy was first derived by Boltzmann in 1872 in his seminal work
on statistical mechanics (Perrot, 1998) and later re-derived independently by Shannon in 1942
in the context of information theory (Shannon, 2001). Since then, entropy has been adopted
as a measure of the information processing capacity of neural systems (Bergström and
Nevanlinna, 1972; Borst and Theunissen, 1999; Hancock et al., 2022; Keshmiri, 2020; Saxe
et al., 2018; Strong et al., 1998; Wang et al., 2014; Yuan et al., 2003). Furthermore, entropy
is used to quantify the repertoire of brain states (Ahmed et al., 2011; Monteforte and Wolf,
2010; Song and Zhang, 2016; Tagliazucchi et al., 2014; Vivot et al., 2020), the effect of
disease (Drachman, 2006; Gómez and Hornero, 2010; Jara et al., 2021; Kannathal et al.,
2005; Morabito et al., 2012; Sharanreddy and Kulkarni, 2013; Takahashi, 2013; Thomas et
al., 2015; van Drongelen et al., 2003; Wang et al., 2017; Wu et al., 2017; Zhou et al., 2016),
and more recently the influence of psychedelics (Carhart-Harris, 2018; Carhart-Harris et al.,
2014; Herzog et al., 2020; Lebedev et al., 2016; Viol et al., 2017).

The purpose of the work presented here is to provide support to this growing community of
neuroscientists that employ entropy as a tool in their research. To this end, we use toy models
in order to build a picture of what entropy represents in the context of neural systems.
Specifically, we will appeal to frameworks first defined in statistical mechanics to highlight the
fact that entropy is not just the property of a system — it is the property of a system and an
observer. It is for this reason that we will frequently refer to entropy in terms of the "ignorance"
of an observer throughout this manuscript.

Microstates
Box 1: System states
Let us imagine a simplified type of neuron that
A state is a feature of a dynamical can only be in one of two possible states (see Box
system that tells us everything we 1). The state can be ‘on’ when the neuron is firing
need to know about what the system or ‘off’ when it is not firing. Specifically, let us
does next. For instance, we may consider two of these binary state neurons
know that if a binary neuron is ‘on’ at
connected to one another — we will call these
one point in time then it will be ‘off’ in
the next and vice versa. In this case,
neuron 1 and neuron 2. Together, these two
we call ‘on’ and ‘off’ the states of the neurons can display a total of four configurations:
system, because they allow us to both neurons are on (Fig. 1A); neuron 1 is on (Fig.
predict what will happen next. 1B); neuron 2 is on (Fig. 1C); and neither neuron
is on (Fig. 1D).

In Fig. 1, when displaying the unique states


Fig 1: A) both neurons are on
that the system is capable of producing, we
are implicitly presenting the system from
B) neuron 1 (left) is on
the perspective of an infinitely
knowledgeable observer. This is because
C) neuron 2 (right) is on
we are assuming that we are able to
distinguish between states without
D) neither neuron is on
limitation either in our spatial resolution or
in our ability to keep track of state
transitions in time. When viewed in such a pure form, a system’s state repertoire is known as
its microstates. Going forward, it will be useful to represent the microstates for the two-neuron
system in Fig. 1 as points in a two-dimensional space (Fig. 2).

Page 2 of 23
Fig 2: The two possible states of neurons 1 and 2 are
shown along the x and y-axes, respectively. This
creates a two dimensional space containing four
points representing the microstates in Fig. 1.

Macrostates

However, in reality there is no such thing as an infinitely knowledgeable observer. We are


inevitably prone to compounding errors when attempting to keep track of a system’s evolution
and lack the necessary resolution to perfectly distinguish between microstates. To illustrate
this point, consider a voltmeter that is connected in parallel across the two neurons in Fig. 1
— such a device would be able to produce three possible readouts: ‘2’ when both neurons are
on (Fig. 3A), ‘1’ when one of the neurons is on (Fig. 3B & C); or ‘0’ when neither neuron is on
(Fig. 3D):

These readouts are known as Fig 3: A voltmeter (circle) registers:


macrostates, referring to the
macroscopic measurements A) '2’ when both neurons are on;
among which we as observers
are able to distinguish. We B) ‘1’ when only neuron 1 is on;
now note from Figs. 3B and C
that there is an ambiguity due C) '1' when only neuron 2 is on;
to our voltmeter assigning the
same readout ‘1’ to two D) '0' when neither neuron is on.
different microstates. This is
because the voltmeter can
only tell us that one neuron is
on, but not precisely which neuron. As we will expand upon, the ignorance that
we as observers face due to this type of ambiguity is quantified by the entropy.

We can illustrate the ways in which the four microstates of Fig. 3 (on/on, on/off, off/on, off/off)
are grouped into the three macrostates (voltages 0, 1, 2) by dividing the two-dimensional
space shown in Fig. 2 into regions (Fig. 4).

Fig 4: The same layout as Fig. 2, together with


lines enclosing the macrostates ‘0’,‘1’, ‘2’
corresponding to the voltmeter readings in
Fig. 3. We see that an observer cannot
distinguish between the two microstates
on/off (top left) and off/on (bottom right), as
these are both enclosed within the same
macrostate ‘1’.

Page 3 of 23
A qualitative description of entropy

Having discussed the concepts of micro and macrostates we are now in a position to make
the following qualitative statement:
entropy is related to the number of microstates
that are compatible with a given macrostate

To understand this more clearly, let us go back to the examples in Figs. 1 through 4 and
examine the three possible measurements that we could take of the two-neuron system. Note
that we use the words ‘measurement’ and ‘macrostate’ interchangeably, as both terms refer
to what we as observers can know about the system.

measurement 1: the voltmeter reads ‘0’.


This measurement, or macrostate, is situated within the top-right circle in
Fig. 4 and corresponds to the single microstate situated at the top-right
point in Fig. 2.

measurement 2: the voltmeter reads ‘2’.


This macrostate is situated within the bottom-left circle in Fig. 4 and
corresponds to the single microstate situated at the bottom-left point in
Fig. 2.

The fact that there is just one microstate for each of these measurements (macrostates)
means that we know all there is to know about the system just by virtue of the measurements.
In other words, by knowing the macrostate we automatically know the microstate and hence
know precisely which unique configuration the system is expressing. As observers we
therefore have a minimum level of ignorance when faced with measurements 1 and 2.

measurement 3: the voltmeter reads ‘1’.


This means that the macrostate is situated within both the top-left and
bottom-right connected circles in Fig. 4. This macrostate covers two
microstates situated at the top-left and bottom-right points in Fig. 2.

In this case our knowledge about the system is limited by the fact that the measurement alone
is not able to distinguish between the two contained microstates and as such our ignorance
as observers has increased.

Note that we are using the term ‘ignorance’ frequently when addressing a discrepancy
between numbers of micro and macrostates. This leads us to another qualitative definition of
entropy:
entropy is a measure of the ignorance of an observer

or equivalently, we can express this as:

entropy is a measure of the hidden information within a system

Page 4 of 23
Characteristics of entropy

What mathematical function would serve best to describe entropy, given that we seek a
quantification of hidden information? To answer this question, let us consider the following
three characteristics that such a function must satisfy.

Characteristic I: function input of unity yields zero output

The mathematical form of entropy must yield an output of zero when the input is one microstate
for a given macrostate. In this case, we as observers are not limited by any ambiguities
regarding the underlying system's state and hence we have zero ignorance.

Characteristic II: function is maximized for a uniform distribution

The mathematical form of entropy must be maximized when the input number of microstates
for a given macrostate is equal to the total number of microstates that the system can
represent.

To better understand this characteristic, let us go back to Fig. 3 and consider what would
happen if we downgraded the quality of the voltmeter. Specifically, this new voltmeter is now
insufficiently sensitive to register any signal from fewer than four neurons. This means that we
would obtain readings of ‘0’ for all microstates containing zero, one, or two active neurons, as
these are all below the new device’s threshold (Fig. 5A, B, C, D).

Fig 5: A voltmeter (circle) registers an


output of ‘0’ for every microstate
in Fig. 1, as they all fall below the
device’s threshold.

The four microstates in Fig. 2 are hence all associated with the same macrostate (Fig. 6).

As this macrostate covers the


entire space of the four
microstates that the system Fig 6: The same layout as
can produce, it is not possible Fig. 2. We indicate
for more information to be the single macrostate
hidden from us — i.e., we are '0' enclosed by the
maximally ignorant. This is square, covering all four
the reason why we require possible microstates.
the mathematical form of
entropy to be maximized for
a uniform distribution.

Characteristic III: function is additive/extensive

Let us take the same voltmeter used in Fig. 5 and consider what happens if we now connect
it across a single-neuron system. This neuron can only be in one of two states: on (Fig. 7A)
and off (Fig. 7B), both of which register an output of ‘0’, because the activity lies below the
device’s threshold. We can represent the microstates (on/off) and the one macrostate (‘0’) as

Page 5 of 23
points along a single axis (Fig. 7C). Our ignorance, as quantified by the size of the macrostate
in Fig. 7C, is half the size of that of the two-neuron system in Fig. 6.
Now let us consider what happens to
the number of microstates if we
Fig 7: A single neuron
combine the two-neuron system in Fig.
connected to a voltmeter
2 with the one-neuron system in Fig 7C.
(circle) registers an
We obtain a three-neuron system, for
output of ‘0’ as the
which the microstates can be counted
signal is below the
by adding a third axis to show that there
device's threshold when
is now a total of eight microstates (Fig.
the neuron is both A) on;
8).
and B) off. C) The two
microstates enclosed by
In Fig. 8 we show that a two-neuron
same macrostate '0', as
system with 4 microstates (Fig. 2)
indicated by the
combined with a single-neuron system
rectangle.
with 2 microstates (Fig. 7C) results in a
three-neuron system with 4 × 2 = 8
microstates. In other words, the number of microstates of the two smaller systems multiply to
give the number of microstates of the new larger system.

Fig 8: The two possible states of neurons 1, 2, and 3 are shown


along the x, y, and z axes. This creates a three-dimensional
space, in which we can represent the eight possible microstates
(squares) of the three-neuron system.

Let us then imagine combining the two-neuron system from Figs. 6 with the single-neuron
system from Fig. 7C. Our total ignorance of the new resultant three-neuron system will then
be equal to the sum of our individual ignorance of the constituent two-neuron (Fig. 9A). and
single-neuron (Fig. 9B) systems.

Therefore, from Figs. 8 and 9, Fig 9: The entropy of the combined


the third characteristic of three-neuron system relates
entropy is that it must be a to the sum of the macrostates
function of the sum of the of the constituent: A) two-
number of macrostates, but neuron; and B) single-neuron
multiplicative for the number of systems indicated by the
microstates. A system that outlining rectangles,
possesses this type of containing microstates
property is said to be indicated by squares.
extensive (see Box 2).

Page 6 of 23
Box 2: Extensivity

There are two types of quantities: those that are intensive and those that are extensive.
Intensive quantities do not grow with the size of the system. An example is temperature —
combining two systems that are each at a temperature of 1℃ does not give a larger system
that is at 2℃. Instead, the result is a larger system that also has a temperature of 1℃.

Extensive quantities, on the other hand, grow with the size of the system. An example is
mass — combining two systems that each have a mass of 1kg gives a larger system that has
a mass of 2 kg The observer’s ignorance is of the extensive type. This is because the more
neurons there are, the more potential there is for the observer to lose track of the states. This
means that we need the entropy to grow in proportion to the number of neurons in the system
being observed.

A quantitative description of entropy

There is only one function that satisfies the three criteria outlined in the last section: the
logarithm.

Characteristic I: unity input, zero output: We can see from the graph of the logarithm function
that this characteristic is satisfied, as 𝑙𝑜𝑔(1) = 0 (Fig. 10).

Characteristic II: maximized for uniform


Fig 10: The number of
distribution: We also see from the graph of
microstates for a
the logarithm (Fig. 10) that the function is
given macrostate
strictly monotone increasing§. This means
(𝜇, x-axis) vs. the
that our ignorance is maximized when the
logarithm of 𝜇 (y-
number of microstates is also maximized.
axis).
Characteristic III: additivity/extensivity: Let us
suppose that there are two systems, one of which has 𝐴 microstates and the other has 𝐵
microstates. If we combine these two systems, the new larger system will have a total 𝐴 × 𝐵
microstates, for which the entropy will be 𝑙𝑜𝑔(𝐴 × 𝐵). The log of a product is the sum of the
logs, i.e. 𝑙𝑜𝑔(𝐴 × 𝐵) = 𝑙𝑜𝑔𝐴 + 𝑙𝑜𝑔𝐵 and therefore, multiplicative numbers of microstates yield
an additive number of macrostates.

Furthermore, we know that when there are two neurons as in Fig. 2 then the number of
microstates is 2! = 4. Similarly, if there are three neurons as in Fig. 8, then the number of
microstates increases to 2" = 8.

In general, we can say that: 𝑛𝑜. 𝑚𝑖𝑐𝑟𝑜𝑠𝑡𝑎𝑡𝑒𝑠 = 2#$. #'()$#*

and hence: 𝑙𝑜𝑔(𝑛𝑜. 𝑚𝑖𝑐𝑟𝑜𝑠𝑡𝑎𝑡𝑒𝑠) = 𝑙𝑜𝑔2#$. #'()$#*


= 𝑛𝑜. 𝑛𝑒𝑢𝑟𝑜𝑛𝑠 × 𝑙𝑜𝑔2

Therefore, the logarithm is extensive, as the entropy increases proportionally to the size of the
system (number of neurons). Note that all entropy values presented henceforth are in units of
bits — i.e., log base 2.

§A larger input of a strictly monotone increasing function always results in a larger output — i.e., as
opposed to a smaller or to a constant output.

Page 7 of 23
These three§ characteristics are rigorously defined within the Shannon-Khinchin axioms
(Khinchin, 1957).

We are now in a position to quantitatively define the entropy:

the entropy is the logarithm of the number of microstates


that are compatible with a given macrostate

Entropy from probability distributions

Thus far, we have made an implicit assumption that all microstates are equally likely to occur.
Going back to Fig. 5 we can depict this assumption in terms of a flat probability distribution
(Fig. 11).

Fig 11: The four possible microstates from Fig. 5 are shown along
the x-axis. Each of these microstates occurs with a
probability (y-axis) of 25%.

However, the microstates may not always be equally likely (Fig. 12).

Fig 12: The four possible microstates are shown along the x-axis.
Each of these microstates has a certain probability (y-axis)
of occurring, as indicated by the grey bars.

When dealing with an unequal distribution as in Fig. 12 we must re-formulate the definition of
entropy as follows:
the entropy is the negative average of the log probabilities of
occurrence of every microstate for a given macrostate

§
In principle there is a fourth characteristic, in that entropy depends only on the number of
microstates, as opposed to the properties of the microstates themselves.

Page 8 of 23
See Box 3 for further details.
Box 3: Entropy from probability
We note that assigning equal probabilities to all
The entropy 𝑆 is defined as the microstates as in Fig. 11 is equivalent to invoking
negative average of the log Occam’s razor (Gauch Jr and Gauch, 2003). In
probabilities 𝑝 of every microstate other words, given that we are maximally ignorant,
for a given macrostate: 𝑆 = −〈𝑙𝑜𝑔𝑝〉, we assume as little as possible regarding
where the angled brackets denote individual outcomes (Kang et al., 2021; Watanabe
an average. Given the definition of
et al., 2013).
the average of a variable
〈𝑥〉 = ∑! 𝑝! 𝑥! , where 𝑝! is the Entropy from re-arrangements
probability of observing the value 𝑥! ,
the entropy can also be written as: We might then wonder why entropy is defined in
𝑆 = − , 𝑝! 𝑙𝑜𝑔𝑝! terms of log probabilities. The reason for this can
!
be understood in terms of indistinguishable re-
arrangements of microstates.
where 𝑝! is the probability of
observing the system in its 𝑖"#
microstate. In the case of 𝑁 equally Consider the macroscopic measurement ‘1’
likely microstates, the probability of shown in Fig. 4 in a two-neuron system when just
observing any individual microstate one neuron is 'on'. This situation occurs because
$ we have a device that is sufficiently sensitive to
is %. We can therefore write the
entropy as: detect single neuron activity, but insufficiently
%
sensitive to detect precisely which of the two
1 1 1 1 neurons is responsible for the measurement.
𝑆 = −, 𝑙𝑜𝑔 = −𝑁 𝑙𝑜𝑔
𝑁 𝑁 𝑁 𝑁
!&$
1 Our ignorance in this scenario comes from the fact
= −𝑙𝑜𝑔 = 𝑙𝑜𝑔𝑁 that there are two indistinguishable ways of re-
𝑁
Therefore, the entropy in the case ofarranging the two-neuron system with respect to
equal probabilities is equivalent to the macroscopic measurement: i.e., neurons 1
the logarithm of the number of and 2 can be on/off (Fig. 13A) or off/on (Fig. 13B).
microstates for a given macrostate.
This is the premise from which the probabilistic
form of entropy is derived — i.e., in terms of
indistinguishable combinations that leave the numbers of neurons in each state (the
‘occupation numbers’) unchanged (see Box 4).

This same simple model can be


extended to other metrics commonly Fig 13: A macroscopic measurement
used in neuroscience, such as is made showing that one
conditional entropy (Liu and Fu, neuron is a two-neuron
2020; Wadhera and Kakkar, 2020; system is on. As observers
Żochowski and Dzakpasu, 2004), we obtain this same
mutual information (Belghazi et al., measurement by re-arranging
2018; Cassidy et al., 2014; Michel et the neurons such that either:
al., 2008; Modat et al., 2010; A) Neuron 1 is on and neuron
Salvador et al., 2007), and transfer 2 is off.
entropy (Mao and Shang, 2017; B) Neuron 1 is off and neuron
Shorten et al., 2021; Spinney et al., 2 is on.
2017; Ursino et al., 2020; Wibral et
al., 2014; Wollstadt et al., 2014):
(see Appendix).

Page 9 of 23
Box 4: Combinatorics and entropy

The number of ways 𝑊 of re-arranging 𝑁 neurons, such that 𝑛$ are in state 1, 𝑛' are in state 2
etc., is given by:
𝑁! 𝑁!
𝑊= =
𝑛$ ! 𝑛' ! 𝑛( ! … ∏! 𝑛! !
where 𝑛! is the occupation number of the 𝑖"# state. It is then convenient to deal with 𝑙𝑜𝑔𝑊:
𝑁!
𝑙𝑜𝑔𝑊 = 𝑙𝑜𝑔 6 7
∏! 𝑛! !

= 𝑙𝑜𝑔𝑁! − 𝑙𝑜𝑔 89 𝑛! !: = 𝑙𝑜𝑔𝑁! − , 𝑙𝑜𝑔𝑛! !


! !
When the occupation numbers 𝑛! are large, we can use Stirling’s approximation (Feller): 𝑙𝑜𝑔𝑥! ≈
𝑥𝑙𝑜𝑔𝑥 − 𝑥 , such that:

𝑙𝑜𝑔𝑊 = 𝑁𝑙𝑜𝑔𝑁 − 𝑁 − , 𝑛! 𝑙𝑜𝑔𝑛! + , 𝑛!


! !
where ∑! 𝑛! = 𝑁 , which means that:

𝑙𝑜𝑔𝑊 = 𝑁𝑙𝑜𝑔𝑁 − , 𝑛! 𝑙𝑜𝑔𝑛!


!
We can express the occupation number 𝑛! in terms of the probability 𝑝! of observing a single
neuron in the 𝑖"# state according to: 𝑛! = 𝑁𝑝! , such that:

𝑙𝑜𝑔𝑊 = 𝑁𝑙𝑜𝑔𝑁 − , 𝑁𝑝! 𝑙𝑜𝑔(𝑁𝑝! )


!

= 𝑁𝑙𝑜𝑔𝑁 − 𝑁𝑙𝑜𝑔𝑁 , 𝑝! − 𝑁 , 𝑝! 𝑙𝑜𝑔𝑝!


! !
where ∑! 𝑝! = 1, and so:
𝑙𝑜𝑔𝑊 = −𝑁 , 𝑝! 𝑙𝑜𝑔𝑝!
!
i.e., 𝑙𝑜𝑔𝑊 is the entropy associated with a single neuron: − ∑! 𝑝! 𝑙𝑜𝑔𝑝! multiplied by the total
number of neurons 𝑁, therefore yielding the entropy of the total system.

Phase space

Let us increase the complexity of the model we have been dealing with so far, such that a
neuron can display more than just two states. We begin by adding an intermediate third state
between ‘on’ and ‘off’, which together could represent, for instance, the depolarization,
repolarization, and hyperpolarization stages of an action potential (Fig. 14A).

Fig 14: A) One state between on (left) and off (right)

B) Three states between on (left) and off (right)

C) An infinite (continuous) number of states


between on (left) and off (right)

Page 10 of 23
We can continue by adding two more
Box 5: Discrete vs. continuous states intermediate states for a total of five
(Fig. 14B). We can then (at least
A discrete state is one that can only take certain mathematically) keep adding such
fixed values — e.g., the neurons in Figs. 1 intermediate states to the point where
through 13 can only be ‘on’ or ‘off’, but they they vary continuously (see Box 5).
cannot be anything in between.
Let us then consider a system
A continuous state, on the other hand, is one that composed of three neurons, each of
varies smoothly between its possible values — as
which can express a continuously
opposed to the sharp jumps of the discrete states.
varying number of states as in Fig.
For instance, if we assign a value of 1 to a neural 14C. We can then represent the
region that is fully active and 0 to one that is allowable microstates as a volume
completely inactive, then a continuous state (Fig. 15).
system would allow us to take a measurement
with a readout anywhere between 0 and 1.

Fig. 15: The states for neurons n1, n2, and n3 shown along
the x, y, and z axes vary continuously. The cube
within this three-dimensional space represents every
possible microstate for a given macrostate.

Previously, in dealing with discrete state systems we were able to count the number of
microstates. However, in Fig. 15 we can no longer count the microstates within the shown
phase space§, as neighbouring points are infinitely close to one another such that they form a
continuous shape (Babloyantz et al., 1985; Fang et al., 2015; Faure and Korn, 2001; Iasemidis
et al., 1990; McKenna et al., 1994; Meyer-Lindenberg, 1996; Rabinovich et al., 2012;
Rabinovich and Muezzinoglu, 2010; Sharma and Pachori, 2015). Therefore, instead of
quantifying the entropy in terms of a countable number of microstates, we instead say that:

§ Phase space refers to a depiction in which every point represents a possible state of a system.

Page 11 of 23
entropy is the logarithm of the volume in
Box 6: Entropy from phase space phase space associated with a
macrostate
In Box 4 we saw that the discrete entropy (𝑆)
is defined as: We then note the following two points:
𝑆 = − , 𝑝! 𝑙𝑜𝑔𝑝!
! 1) although we use the term ‘volume’, this
where 𝑝! is the probability of observing the does not always refer to a three-
system in its 𝑖"# state. However, now that we dimensional space. A phase space can
are dealing with neurons that have a have any number of dimensions according
continuous number of states (𝑧), we instead
to the number of neurons and the space
define the differential entropy in terms of a
probability distribution 𝑝(𝑧): therein is always referred to as a ‘volume’.

𝑆 = − @ 𝑝𝑙𝑜𝑔𝑝𝑑𝑧 2) as there are infinite states contained


within any volume in phase space, the
It should be noted, however, that the
differential entropy is not, in general, a
mathematical form of entropy in Box 4
natural extension of discrete entropy. changes from a sum to an integral (see
Box 6).

Conservation of information

Let us suppose that we allow the system in Fig. 15 to evolve for a certain amount of time, after
which we find that the volume in phase space has adopted a new shape (Fig. 16).

Fig 16: The continuously varying states for neurons n1, n2, and
n3 are shown along the x, y, and z axes. The volume in
phase space at one point in time (left) encloses all
microstates compatible with a macroscopic
measurement. We then allow some time to pass
(stopwatch), after which we find the shape has been
stretched in one dimension and compressed in another,
but with a total volume that remains unchanged.

What we find is that if the volume stretches along one dimension, this is compensated by a
compression in a different dimension, or in other words:

statement 1: the volume in phase space remains constant

This is a principle known as Liouville’s theorem (Wallace, 2013), which tells us that the
distinctions between microstates in a certain class of deterministic dynamical systems is never
lost. This is a statement of the conservation of information.

However, we recall the statement from the previous section that:

statement 2: entropy is the logarithm of the volume in phase space

It therefore follows from the above two statements that:

statements 1 & 2 ⟹ the entropy remains constant

Page 12 of 23
This is a point that we will clarify further after discussing the phenomenon of chaos in the
next section.

Chaos

Most systems found in nature, including those in the brain (Faure and Korn, 2001; Freeman,
1995; Rabinovich and Abarbanel, 1998; Schiff et al., 1994), exhibit a phenomenon known as
chaos. To understand what this means, let us simplify the situation shown in Fig. 16, such that
there are just two neurons. We then imagine that the microstates are contained within some
initial area in the two-dimensional phase space. If we then allow some time to pass, we find
that the shape begins to grow ‘branches’, which in time sprout their own smaller branches in
a fractal manner as the shape gradually spreads out in phase space (Fig. 17).

Fig 17: The black shape contains all the microstates


that are compatible with a given macroscopic
measurement. As the system evolves in time,
indicated by the stopwatch and arrows, the
black shape spreads out in phase space, whilst
preserving its total area.

This shows the evolution of a chaotic


Box 7: Lyapunov exponents system, as characterised by points in
phase space that are infinitely close
In a chaotic system, two points in phase space together diverging exponentially in time
that are initially separated by an arbitrarily small at a rate depending on their positive
distance 𝑥 will diverge exponentially in time at a
Lyapunov exponent (Arnold and
rate determined by a positive
Wihstutz, 1986) (see Box 7). It is
Lyapunov exponent 𝜆 such that,
after a time 𝑡, the separation has important to remember that, although
increased to: 𝑥𝑒 )" the shape is branching out in this
manner, the total area remains constant
in time according to Liouville’s theorem.

Coarse graining and the second law

Liouville’s theorem tells us that the volume of the phase space in Fig. 17 does not change in
time and hence neither does the entropy (i.e., the logarithm of the volume). However, there is
an important caveat to this statement, which is that the entropy only remains constant for an
observer with the ability to track the system with infinite precision (Susskind and Hrabovsky,
2014). In other words, if it were possible to know all initial conditions perfectly and to resolve
infinitely close points in phase space, then the entropy would indeed remain constant. We will
refer to this ideal as the ‘fine grain entropy’, as it involves unlimited knowledge of the fine grain
structure of the system in question.

However, there is no such thing as a perfect observer. Regardless of how accurate we are,
there will always be some uncertainty due to limitations in our ability to take measurements.
We can represent such limitations by a process known as 'coarse graining' (Castiglione et al.,
2008; Espanol, 2004; Mazenko, 2000; Meshulam et al., 2019; Ridderbos, 2002). This involves
dividing the phase space into regions that represent the limit of our resolving power. Any two
points that lie within such a region are indistinguishable to an observer.

Page 13 of 23
Let us therefore revisit Fig. 17, but this time view the system with a device that allows for a
maximum resolution set by the size of the squares shown in Fig. 18A. This means that,
although the area of the underlying shape from Fig. 17 remains constant, the number of
squares, and hence the coarse-grained area, increases in time (Fig. 18B).

Therefore, although the fine grained entropy of the perfect observer in Fig. 17 remains
constant, the coarse grained entropy in Fig. 18B increases with time as the imperfect observer
gradually loses track of the distinctions between states (Collell and Fauquet, 2015; Lucia,
2013; Salerian, 2010). This is Gibbs’ interpretation of the second law of thermodynamics which
tells us that, due to compounding errors, our ignorance (entropy) tends to a maximum (Jaynes,
1965).

Fig 18: A) The same system from Fig. 17


now shown overlaid on a grid for
which each square represents the
maximum resolution of an imperfect
observer — i.e., two points that lie
within such a square cannot be
distinguished from one another.

B) As the system evolves in time,


we require an ever-increasing
number of squares to cover the
branching microstates of the black
shape in A).

Discussion

Entropy quantifies the hidden information contained within a system. As such, entropy is a
property of both a system and an observer — a point that was discussed extensively by ET
Jaynes (Jaynes, 1988a; Jaynes, 1957, 1965, 1982; Jaynes, 1985, 1988b).

Entropy is often calculated from neuroimaging timeseries. To begin with, such timeseries are
usually discretized (or 'binned'), which means that at every time point activity within a certain
range is placed into a bin. By counting the number of time points that fall within a bin range
one obtains a probability distribution, from which the entropy is subsequently calculated —
see, for example, the approach used in the context of psychedelics in (Carhart-Harris et al.,
2014).

But what exactly does the entropy tell us when calculated this way? To answer this, we
consider a time course that contains only three time points (Fig. 19A). The activity levels for
two of the time points fall within the first bin and one falls within the second bin, which means
that the associated probabilities of occurrence are 2/3 and 1/3, respectively (Fig. 19B).

Fig 19: A) Activity (y-axis) with two binned values plotted


against three time points (x-axis).

B) The probabilities (y-axis) of 2/3 and 1/3 of


observing the two binned activities from A).

Page 14 of 23
We then calculate the entropy associated with the probability distribution in Fig. 19B:

+ + ! !
𝑆 = − ∑𝑖 𝑝𝑖 𝑙𝑜𝑔𝑝𝑖 = − D" 𝑙𝑜𝑔 " + " 𝑙𝑜𝑔 "E ≈ 0.9 bits

This 0.9 bit ignorance represents the observer's inability to distinguish between different ways
of re-arranging a time course such that it remains consistent with the probability distribution in
Fig. 19B. In other words, the same macrostate in Fig. 19B is compatible with three
indistinguishable microstates that correspond to different ways of re-arranging the underlying
time course (Fig. 20).

Fig 20: The three possible time courses (microstates)


that are compatible with the probability
distribution in Fig. 19B (macrostate).

We see that each of the three re-arrangements in Fig. 20 could have produced the probability
distribution in Fig. 19B — this ambiguity is the source of the calculated entropy of 0.9 bits.

Let us now step backwards in the analysis pipeline of such a study to the stage in which a
continuous time course is binned to produce a discrete time course. For the sake of argument,
let us suppose that a single time point is entered into a certain bin if its activity range falls
somewhere between zero and one (arbitrary units). What this means is that if we observe a
binned time point for the zero-to-one range (macrostate), this could have arisen by virtue of
any continuous number of values between zero and one (microstates). There is a quantifiable
ignorance associated with this ambiguity, which we can calculate by turning to the integral
form of entropy (Box 6), the exact value of which depends on the way in which the activity is
distributed. We therefore obtain a second value of entropy that quantifies how much
information is hidden from us when binning a time point from a continuous time course within
this range.

We can now step back further in the analysis pipeline to the stage at which activity levels are
actually being measured from the neural system. Any measured activity level (macrostate)
originates as one or more configurations of on/off firing patterns (microstates) in the underlying
neural system at smaller scales, yielding a third value of entropy.

Furthermore, it is in principle possible to calculate the entropy of a neural region embedded


within a larger system. For instance, we can consider two regions: 1) A receiver (Fig. 21A);
and 2) a transmitter (Fig. 21B).

Fig 21: A) A receiver in a certain macrostate.

B) A transmitter in a certain microstate


compatible with the macrostate in A).

Note that the receiver therefore effectively takes on the role of the observer in this scenario.
One would thus be in a position to count the number and relative occurrence of transmitted
firing patterns (Fig. 21B, microstates) that all result in the same received firing pattern (Fig.
21A, macrostate). The logarithm of this number of patterns yields the entropy, which would

Page 15 of 23
provide a measure of how much information about the transmitter is lost upon being encoded
by the receiver.

As such, there are many perspectives that can be considered in a neuroimaging experiment,
each with values of entropy that depend on the specific limitations and perspective of the
observer in question. The purpose of these examples is to stress the importance of the
observer when reporting entropy in the context of neuroimaging. The specific scenarios
analysed were chosen to be sufficiently simple to allow for ease of explanation and
interpretation. However, the line of reasoning that we employed could equally well be applied
to a multitude of scenarios within neuroscience, each with a different micro/macrostate
relationship.

It is our hope that the clarification of entropy as presented here will lead to more coherence
across studies seeking to quantify neural information processing within increasingly complex
neuroimaging datasets.

Page 16 of 23
Appendix

Conditional entropy

Let us suppose that, within the context of a two-neuron system, only one of the neurons can
be observed directly. If the observable neuron fires and we are asked to make a guess as to
the state of the unobservable neuron then our guess would depend on what we know about
the connection between the two. For instance, we may not know anything about the extent to
which the two neurons are connected, or even if they are connected at all, in which case we
would have to assign 50/50 odds to the obscured neuron being either on or off, as our
ignorance would be at a maximum (Fig. 22A).

Fig 22: Two connected neurons in which the left is in the ‘on’ state
and the right is obscured from view (indicated by the blank
square). A) We have no information about the connection
between the two neurons. B) We know that the neurons
are connected, but we do not know how strongly. C) We
know that the two neurons are sufficiently strongly
connected to ensure co-activation.

Alternatively, we may know that the two neurons are connected, but not how strongly. In this
case, we would assign higher odds to the obscured neuron also being on due to the influence
of the observable neuron. Compared to Fig. 22A, our ignorance is now somewhat lower (Fig.
22B).

Another possibility is that we know with certainty that the two neurons are very strongly
connected to the point where co-activation is ensured. In this case we would assign 100%
odds to the obscured neuron also being on and thus our ignorance would be zero (Fig. 22C).

This example illustrates the concept of conditional entropy, in which we characterise the
ignorance regarding one part of the system (the obscured neuron) given knowledge about a
different part of the system (the observable neuron and the connection strength):

How ignorant we are What we know about the


The conditional
entropy 𝑒𝑞𝑢𝑎𝑙𝑠 about the state of the 𝑔𝑖𝑣𝑒𝑛 observable neuron and
obscured neuron the connection strength

Mutual information

Mutual information (also called expected information gain) is directly related to conditional
entropy, and it tells us how much an observer’s ignorance of one part of the system (the
obscured neuron) is reduced by knowing something about an entirely different part of the
system — the observable neuron and the connection strength.

The mutual Our ignorance about The conditional


information
𝑒𝑞𝑢𝑎𝑙𝑠 the obscured neuron
𝑚𝑖𝑛𝑢𝑠
entropy

Page 17 of 23
Transfer entropy

In Fig. 22 we considered the influence that one neuron exerts upon another across space via
their connection. However, we can also consider influence across time. For example, if we
observe a single neuron firing, we may want to make a prediction as to that same neuron’s
future state (Fig. 23).

Fig 23: A neuron fires (bottom left). Time passes (stopwatch), after
which we make a guess as to the same neuron’s future state
(question mark).

Our guess as to the future state once again depends upon how much we know about the
system. If our ignorance is maximized then we must assign 50/50 odds to the neuron's future
state as being either being on or off.

On the other hand, we may know something about the temporal behaviour of the neuron —
e.g., its refractory period. If the neuron has finished firing and the amount of time that elapses
between observations is shorter than the refractory period, we would be certain that the future
state will be ‘off’. This corresponds to minimum ignorance (entropy).

We will now consider a more complicated situation, in which we track two spatially connected
neurons in time. Specifically, we are faced with the following situation: neuron 1 fires and we
are asked to guess what the state of neuron 2 is at a future point in time (Fig. 24).

Fig 24: A connected two-neuron system (left column) in which neuron 1


fires. Time passes (stopwatch), after which we make a guess as to 1 1
the state of neuron 2 (question mark).
2 2?

In this case, we would like to know how ignorant we are about the future of neuron 2, given
that we know something about both neurons 1 and 2. The probability that we assign to neuron
2 being ‘on’ at a future point in time depends on what an observer knows about both the spatial
properties of the system (e.g., the connection strengths), as well as its temporal properties
(e.g., the refractory period).
Transfer entropy is a measure of how much the present state
of one system affects the future state of a different system.

Our ignorance of the Our ignorance of the future


The transfer
future state of neuron 2, state of neuron 2, given
entropy between 𝑒𝑞𝑢𝑎𝑙𝑠 given knowledge of the 𝑚𝑖𝑛𝑢𝑠 knowledge of the present states
neurons 1 and 2
present state of neuron 2 of both neurons 1 and 2

Page 18 of 23
Acknowledgements

The authors wish to thank Karl J. Friston for his feedback in preparing this work, as well as
Pedro Mediano for his constructive review.

Funding

The authors acknowledge support from the AI Centre for Value Based Healthcare, the Data
to Early Diagnosis and Precision Medicine Industrial Strategy Challenge Fund, UK Research
and Innovation (UKRI), the National Institute for Health Research (NIHR), the Biomedical
Research Centre at South London, the Maudsley NHS Foundation Trust, the Wellcome, and
King’s College London.

References

Ahmed, M.U., Li, L., Cao, J., Mandic, D.P., 2011. Multivariate multiscale entropy for brain
consciousness analysis, 2011 Annual International Conference of the IEEE Engineering in Medicine
and Biology Society. IEEE, pp. 810-813.

Arnold, L., Wihstutz, V., 1986. Lyapunov Exponents, volume 1186 of. Lecture Notes in Mathematics.

Babloyantz, A., Salazar, J., Nicolis, C., 1985. Evidence of chaotic dynamics of brain activity during the
sleep cycle. Physics letters A 111, 152-156.

Belghazi, M.I., Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., Hjelm, D., 2018. Mutual
information neural estimation, International conference on machine learning. PMLR, pp. 531-540.

Bergström, R., Nevanlinna, O., 1972. An entropy model of primitive neural systems. International
Journal of Neuroscience 4, 171-173.

Borst, A., Theunissen, F.E., 1999. Information theory and neural coding. Nature neuroscience 2, 947-
957.

Carhart-Harris, R.L., 2018. The entropic brain-revisited. Neuropharmacology 142, 167-178.

Carhart-Harris, R.L., Leech, R., Hellyer, P.J., Shanahan, M., Feilding, A., Tagliazucchi, E., Chialvo,
D.R., Nutt, D., 2014. The entropic brain: a theory of conscious states informed by neuroimaging
research with psychedelic drugs. Frontiers in human neuroscience, 20.

Cassidy, B., Rae, C., Solo, V., 2014. Brain activity: Connectivity, sparsity, and mutual information.
IEEE transactions on medical imaging 34, 846-860.

Castiglione, P., Falcioni, M., Lesne, A., Vulpiani, A., 2008. Chaos and coarse graining in statistical
mechanics. Chaos and Coarse Graining in Statistical Mechanics.

Collell, G., Fauquet, J., 2015. Brain activity and cognition: a connection from thermodynamics and
information theory. Frontiers in psychology 6, 818.

Drachman, D.A., 2006. Aging of the brain, entropy, and Alzheimer disease. Neurology 67, 1340-1352.

Espanol, P., 2004. Statistical mechanics of coarse-graining, Novel methods in soft matter simulations.
Springer, pp. 69-115.

Page 19 of 23
Fang, Y., Chen, M., Zheng, X., 2015. Extracting features from phase space of EEG signals in brain–
computer interfaces. Neurocomputing 151, 1477-1485.

Faure, P., Korn, H., 2001. Is there chaos in the brain? I. Concepts of nonlinear dynamics and
methods of investigation. Comptes Rendus de l'Académie des Sciences-Series III-Sciences de la Vie
324, 773-793.

Feller, W., An introduction to probability theory and its applications. 1, 2nd.

Freeman, W.J., 1995. Chaos in the brain: Possible roles in biological intelligence. International
Journal of Intelligent Systems 10, 71-88.

Gauch Jr, H.G., Gauch, H.G., 2003. Scientific method in practice. Cambridge University Press.

Gómez, C., Hornero, R., 2010. Entropy and complexity analyses in Alzheimer’s disease: An MEG
study. The open biomedical engineering journal 4, 223.

Hancock, F., Rosas, F., Mediano, P., Luppi, A., Cabral, J., Dipasquale, O., Turkheimer, F., 2022. May
the 4C’s be with you: An overview of complexity-inspired frameworks for analyzing resting-state
neuroimaging data.

Herzog, R., Mediano, P.A., Rosas, F.E., Carhart-Harris, R., Perl, Y.S., Tagliazucchi, E., Cofre, R.,
2020. A mechanistic model of the neural entropy increase elicited by psychedelic drugs. Scientific
reports 10, 1-12.

Iasemidis, L.D., Chris Sackellares, J., Zaveri, H.P., Williams, W.J., 1990. Phase space topography
and the Lyapunov exponent of electrocorticograms in partial seizures. Brain topography 2, 187-201.

Jara, J., Morales-Rojas, C., Fernández-Muñoz, J., Haunton, V.J., Chacón, M., 2021. Using
complexity–entropy planes to detect Parkinson’s disease from short segments of haemodynamic
signals. Physiological measurement 42, 084002.

Jaynes, E., 1988a. The relation of Bayesian and maximum entropy methods, Maximum-entropy and
Bayesian methods in science and engineering. Springer, pp. 25-29.

Jaynes, E.T., 1957. Information theory and statistical mechanics. Physical review 106, 620.

Jaynes, E.T., 1965. Gibbs vs Boltzmann entropies. American Journal of Physics 33, 391-398.

Jaynes, E.T., 1982. On the rationale of maximum-entropy methods. Proceedings of the IEEE 70, 939-
952.

Jaynes, E.T., 1985. Macroscopic prediction, Complex Systems—Operational Approaches in


Neurobiology, Physics, and Computers. Springer, pp. 254-269.

Jaynes, E.T., 1988b. How does the brain do plausible reasoning?, Maximum-entropy and Bayesian
methods in science and engineering. Springer, pp. 1-24.

Kang, J., Jeong, S.O., Pae, C., Park, H.J., 2021. Bayesian estimation of maximum entropy model for
individualized energy landscape analysis of brain state dynamics. Human Brain Mapping 42, 3411-
3428.

Kannathal, N., Choo, M.L., Acharya, U.R., Sadasivan, P., 2005. Entropies for detection of epilepsy in
EEG. Computer methods and programs in biomedicine 80, 187-194.

Keshmiri, S., 2020. Entropy and the Brain: An overview. Entropy 22, 917.

Khinchin, A.I., 1957. Mathematical foundations of information theory, Vol. 434. Courier Corporation.

Page 20 of 23
Lebedev, A.V., Kaelen, M., Lövdén, M., Nilsson, J., Feilding, A., Nutt, D.J., Carhart‐Harris, R.L., 2016.
LSD‐induced entropic brain activity predicts subsequent personality change. Human brain mapping
37, 3203-3213.

Liu, X., Fu, Z., 2020. A novel recognition strategy for epilepsy EEG signals based on conditional
entropy of ordinal patterns. Entropy 22, 1092.

Lucia, U., 2013. Irreversible human brain. Medical Hypotheses 80, 112-114.

Mao, X., Shang, P., 2017. Transfer entropy between multivariate time series. Communications in
Nonlinear Science and Numerical Simulation 47, 338-347.

Mazenko, G.F., 2000. Equilibrium statistical mechanics.

McKenna, T., McMullen, T., Shlesinger, M., 1994. The brain as a dynamic physical system.
Neuroscience 60, 587-605.

Meshulam, L., Gauthier, J.L., Brody, C.D., Tank, D.W., Bialek, W., 2019. Coarse Graining, Fixed
Points, and Scaling in a Large Population of Neurons. Phys Rev Lett 123, 178103.

Meyer-Lindenberg, A., 1996. The evolution of complexity in human brain development: an EEG study.
Electroencephalography and clinical neurophysiology 99, 405-411.

Michel, V., Damon, C., Thirion, B., 2008. Mutual information-based feature selection enhances fMRI
brain activity classification, 2008 5th IEEE international symposium on biomedical imaging: From
nano to macro. IEEE, pp. 592-595.

Modat, M., Vercauteren, T., Ridgway, G.R., Hawkes, D.J., Fox, N.C., Ourselin, S., 2010.
Diffeomorphic demons using normalized mutual information, evaluation on multimodal brain MR
images, Medical Imaging 2010: Image Processing. SPIE, pp. 800-807.

Monteforte, M., Wolf, F., 2010. Dynamical entropy production in spiking neuron networks in the
balanced state. Physical review letters 105, 268104.

Morabito, F.C., Labate, D., Foresta, F.L., Bramanti, A., Morabito, G., Palamara, I., 2012. Multivariate
multi-scale permutation entropy for complexity analysis of Alzheimer’s disease EEG. Entropy 14,
1186-1202.

Perrot, P., 1998. A to Z of Thermodynamics. Oxford University Press on Demand.

Rabinovich, M., Abarbanel, H., 1998. The role of chaos in neural systems. Neuroscience 87, 5-14.

Rabinovich, M.I., Afraimovich, V.S., Bick, C., Varona, P., 2012. Information flow dynamics in the brain.
Physics of life reviews 9, 51-73.

Rabinovich, M.I., Muezzinoglu, M., 2010. Nonlinear dynamics of the brain: emotion and cognition.
Physics-Uspekhi 53, 357.

Ridderbos, K., 2002. The coarse-graining approach to statistical mechanics: how blissful is our
ignorance? Studies in History and Philosophy of Science Part B: Studies in History and Philosophy of
Modern Physics 33, 65-77.

Salerian, A.J., 2010. Thermodynamic laws apply to brain function. Medical hypotheses 74, 270-274.

Salvador, R., Martinez, A., Pomarol-Clotet, E., Sarró, S., Suckling, J., Bullmore, E., 2007. Frequency
based mutual information measures between clusters of brain regions in functional magnetic
resonance imaging. Neuroimage 35, 83-88.

Page 21 of 23
Saxe, G.N., Calderone, D., Morales, L.J., 2018. Brain entropy and human intelligence: A resting-state
fMRI study. PloS one 13, e0191582.

Schiff, S.J., Jerger, K., Duong, D.H., Chang, T., Spano, M.L., Ditto, W.L., 1994. Controlling chaos in
the brain. Nature 370, 615-620.

Shannon, C.E., 2001. A mathematical theory of communication. ACM SIGMOBILE mobile computing
and communications review 5, 3-55.

Sharanreddy, P., Kulkarni, P., 2013. EEG signal classification for epilepsy seizure detection using
improved approximate entropy. Int J Public Health Sci 2, 23-32.

Sharma, R., Pachori, R.B., 2015. Classification of epileptic seizures in EEG signals based on phase
space representation of intrinsic mode functions. Expert Systems with Applications 42, 1106-1117.

Shorten, D.P., Spinney, R.E., Lizier, J.T., 2021. Estimating transfer entropy in continuous time
between neural spike trains or other event-based data. PLoS computational biology 17, e1008054.

Song, Y., Zhang, J., 2016. Discriminating preictal and interictal brain states in intracranial EEG by
sample entropy and extreme learning machine. Journal of neuroscience methods 257, 45-54.

Spinney, R.E., Prokopenko, M., Lizier, J.T., 2017. Transfer entropy in continuous time, with
applications to jump and neural spiking processes. Physical Review E 95, 032319.

Strong, S.P., Koberle, R., Van Steveninck, R.R.D.R., Bialek, W., 1998. Entropy and information in
neural spike trains. Physical review letters 80, 197.

Susskind, L., Hrabovsky, G., 2014. The theoretical minimum: what you need to know to start doing
physics. Basic Books.

Tagliazucchi, E., Carhart‐Harris, R., Leech, R., Nutt, D., Chialvo, D.R., 2014. Enhanced repertoire of
brain dynamical states during the psychedelic experience. Human brain mapping 35, 5442-5456.

Takahashi, T., 2013. Complexity of spontaneous brain activity in mental disorders. Progress in Neuro-
Psychopharmacology and Biological Psychiatry 45, 258-266.

Thomas, J.B., Brier, M.R., Ortega, M., Benzinger, T.L., Ances, B.M., 2015. Weighted brain networks
in disease: centrality and entropy in human immunodeficiency virus and aging. Neurobiology of aging
36, 401-412.

Ursino, M., Ricci, G., Magosso, E., 2020. Transfer entropy as a measure of brain connectivity: A
critical analysis with the help of neural mass models. Frontiers in computational neuroscience 14, 45.

van Drongelen, W., Nayak, S., Frim, D.M., Kohrman, M.H., Towle, V.L., Lee, H.C., McGee, A.B.,
Chico, M.S., Hecox, K.E., 2003. Seizure anticipation in pediatric epilepsy: use of Kolmogorov entropy.
Pediatric neurology 29, 207-213.

Viol, A., Palhano-Fontes, F., Onias, H., de Araujo, D.B., Viswanathan, G., 2017. Shannon entropy of
brain functional complex networks under the influence of the psychedelic Ayahuasca. Scientific
reports 7, 1-13.

Vivot, R.M., Pallavicini, C., Zamberlan, F., Vigo, D., Tagliazucchi, E., 2020. Meditation increases the
entropy of brain oscillatory activity. Neuroscience 431, 40-51.

Wadhera, T., Kakkar, D., 2020. Conditional entropy approach to analyze cognitive dynamics in autism
spectrum disorder. Neurological Research 42, 869-878.

Wallace, D., 2013. What statistical mechanics actually does.

Page 22 of 23
Wang, B., Niu, Y., Miao, L., Cao, R., Yan, P., Guo, H., Li, D., Guo, Y., Yan, T., Wu, J., 2017.
Decreased complexity in Alzheimer's disease: resting-state fMRI evidence of brain entropy mapping.
Frontiers in aging neuroscience 9, 378.

Wang, Z., Li, Y., Childress, A.R., Detre, J.A., 2014. Brain entropy mapping using fMRI. PloS one 9,
e89948.

Watanabe, T., Hirose, S., Wada, H., Imai, Y., Machida, T., Shirouzu, I., Konishi, S., Miyashita, Y.,
Masuda, N., 2013. A pairwise maximum entropy model accurately describes resting-state human
brain networks. Nature communications 4, 1-10.

Wibral, M., Vicente, R., Lindner, M., 2014. Transfer entropy in neuroscience, Directed information
measures in neuroscience. Springer, pp. 3-36.

Wollstadt, P., Martínez-Zarzuela, M., Vicente, R., Díaz-Pernas, F.J., Wibral, M., 2014. Efficient
transfer entropy analysis of non-stationary neural time series. PloS one 9, e102833.

Wu, Y., Chen, P., Luo, X., Wu, M., Liao, L., Yang, S., Rangayyan, R.M., 2017. Measuring signal
fluctuations in gait rhythm time series of patients with Parkinson's disease using entropy parameters.
Biomedical Signal Processing and Control 31, 265-271.

Yuan, H., Xiong, F., Huai, X., 2003. A method for estimating the number of hidden neurons in feed-
forward neural networks based on information entropy. Computers and Electronics in Agriculture 40,
57-64.

Zhou, F., Zhuang, Y., Gong, H., Zhan, J., Grossman, M., Wang, Z., 2016. Resting state brain entropy
alterations in relapsing remitting multiple sclerosis. PLoS One 11, e0146080.

Żochowski, M., Dzakpasu, R., 2004. Conditional entropies, phase synchronization and changes in the
directionality of information flow in neural systems. Journal of Physics A: Mathematical and General
37, 3823.

Page 23 of 23

Common questions

Powered by AI

Microstates refer to the specific configurations of a system at a detailed level, such as the firing states 'on' or 'off' of neurons within a neural system, as described by an infinitely knowledgeable observer . Macrostates, by contrast, are broader overviews or measurements that an observer can ascertain, such as using a voltmeter that reads a particular value when observing neural behavior but cannot distinguish between different microstates that lead to the same voltage reading . Entropy in this context measures the 'ignorance' or lack of knowledge of the observer about which specific microstate corresponds to an observed macrostate . Thus, entropy reflects the number of microstates compatible with a given macrostate .

Entropy in neural systems and its application in information theory as proposed by Shannon are conceptually connected in that both measure the degree of uncertainty or ignorance associated with a system. In neural systems, entropy quantifies the ignorance about the particular microstates of neurons that correspond to a given macrostate measured by an observer . In Shannon's information theory, entropy is used to characterize the uncertainty in predicting the state of a communication system or a message, representing the amount of disorder or unpredictability involved . Both contexts use entropy to describe the informational complexity and potential for disorder within a system, highlighting fundamental principles of measuring uncertainty .

Macroscopic measurements define macrostates via observable phenomena, such as voltage readings, that can distinguish certain system-level states while remaining unable to resolve down to individual microstate details . This definition implies that each macrostate encompasses an ensemble of potential microstates, leading to intrinsic uncertainty or ignorance in exact system behavior. This, in turn, impacts our understanding of the system's entropy, reflecting the degree of this indeterminacy. Knowing the macrostate provides partial information about the system, but without full microstate visibility, entropy becomes a measure of our informational limitations regarding the system's behavior .

Applying thermodynamic concepts such as entropy to neural data analysis can provide insights into the complexity and variability of brain functioning. Entropy measures correlate with the diversity of brain states, with higher entropy indicating a richer repertoire of configurations, which may be associated with more adaptable or complex functioning such as during cognitive tasks or under conditions like psychedelic experiences . It also aids in understanding disorders characterized by altered brain dynamics, as lower entropy values have been associated with conditions like Alzheimer's disease, indicating loss of complexity and adaptability in brain networks .

The relationship between entropy and indistinguishable re-arrangements of microstates involves the inability to distinguish different configurations that result in the same observable measurement or macrostate. In a neural system, the entropy reflects the measure of such indistinguishability—for example, when a voltmeter reading shows a macrostate of '1', it does not specify whether the configuration is neuron 1 'on' and neuron 2 'off' or vice versa . This indistinguishability contributes to entropy, as the same macrostate can be produced by multiple microstate arrangements, adding ambiguity and increasing entropy .

The concept of observer 'ignorance' in measuring entropy arises from the observer's inability to distinguish between different microstates that may correspond to the same macrostate. This results in ambiguity and a lack of complete information about the system . The ignorance is quantified by entropy, which is related to the number of indistinguishable microstates that exist for a given macrostate reading. For instance, two microstates such as one neuron being 'on' and the other 'off' can result in the same macrostate reading of '1' by a voltmeter, leading to ambiguity for the observer .

Increasing the number of neurons in a neural system tends to increase the entropy of the system because it leads to a higher number of potential microstates and more possible combinations of neuron states. This increase in complexity and potential for observer ignorance results in higher entropy, reflecting the extensive nature of entropy as it grows proportionally with the size of the system being observed .

Entropy is considered an 'extensive' property because it scales with the size of the system, such as the number of neurons in a neural network. As the system size increases, the potential combinations of states and thus the number of microstates increase, leading to higher entropy . This implies that as neural networks scale up, careful modeling is required to account for the increasing complexity and loss of information resolution that occurs with larger systems, impacting how accurately the model can predict or analyze system behavior .

Probability distributions impact entropy calculation by indicating the likelihood of each microstate occurring. Initially, the assumption might be that all microstates are equally likely, leading to a flat probability distribution . However, when microstates are not equally probable, entropy must be calculated using the negative average of the log probabilities of each microstate . The logarithm is used to account for the indistinguishable nature of rearrangements of microstates, providing a way to quantify entropy that scales appropriately with the number of possible configurations of a system .

The assumption of equal probabilities for all microstates in entropy measurement relates to Occam's razor by embodying the principle of minimizing assumptions when dealing with maximum uncertainty or ignorance. By assigning equal likelihood to each microstate, the model assumes the least amount of prior knowledge about the system, hence avoiding unnecessary complexity in the absence of detailed information. This approach aligns with Occam's razor, which suggests that the simplest explanation, or the one with the fewest assumptions, should be preferred when multiple plausible explanations exist .

You might also like